Skip to main content
Migration Notice
We're migrating documentation from the old portal into this one. Some things may look a little different or out of place in the meantime — we know, and we're working to get it right. If something's unclear or doesn't look right, let us know.

System Monitor

The System Monitor page gives real-time visibility into application memory and CPU health, external endpoint reachability, audit ingest performance, and API performance. Use it first when you investigate slowness, high memory usage, or a suspected stalled ingest.

Where: System Operations > System Monitor

Overview​

The page is divided into panels, each covering one aspect of system health. The page refreshes every 5 seconds, and the memory and CPU panel refreshes every 15 seconds. A countdown timer shows the time until the next refresh of that panel.

The panels show the current snapshot only. The exception is the API Call Statistics table, which accumulates counts since the last backend restart.

JVM Memory and CPU​

This panel shows memory and CPU metrics for the running application process.

MetricMeaning
Heap used / committed / maxShown as a bar. Heap used consistently near max puts the process at risk of running out of memory.
Non-heap memoryMemory the JVM uses outside the heap, including code cache and compiler buffers.
Direct buffersOff-heap buffers allocated by NIO channels and similar APIs. High direct buffer usage together with high heap can indicate I/O-heavy operations.
MetaspaceMemory used for class metadata. It grows with the number of loaded classes and does not shrink unless classes are unloaded.
Thread count (total / daemon / peak)Current number of threads, how many are daemon threads, and the peak since startup.
GC run count and timeNumber of garbage collection cycles and the cumulative time spent in them. High GC time relative to uptime can indicate heap pressure.
System memoryTotal and free memory on the host OS, as reported by the JVM.

Active Operation Spans​

This panel lists the operations running now, with the wall-clock time elapsed and the CPU time consumed. Each span is a tracked operation such as an ingest run, a ClickHouse query, or a copy job step.

An operation that stays here for an unusually long time can indicate a stuck job or a slow dependency, such as a network timeout waiting on an external device.

Recent Operation Spans​

This panel lists the most recently completed operations, sorted by completion time, with their wall-clock and CPU duration. Use it to review what the system did in the last few minutes and to spot operations that took unexpectedly long.

Endpoint Reachability​

This table shows whether each configured external endpoint is reachable. Only the endpoints that apply to your deployment are shown. The panel refreshes every 60 seconds, and a refresh button in the panel header refreshes it immediately.

ColumnMeaning
EndpointEndpoint name
TypeBadge: Storage, ClickHouse, or App Database
StatusUP, DOWN, or Not configured
LatencyResponse time in milliseconds
URLEndpoint address
Last checkedTime of the last check

Rows with DOWN status are highlighted in red.

ClickHouse and PostgreSQL health​

These summary panels link to the detailed health pages.

  • ClickHouse Summary — connection status, latency, version, and compressed and uncompressed table size for the ClickHouse analytics database. It links to the ClickHouse Health page.
  • PostgreSQL Summary — connection status, latency, version, active connections, idle connections, pool utilization percentage (green, amber, or red), total connections, maximum connections, and database size for the application PostgreSQL database. It links to the PostgreSQL Database Health page. The panel is hidden when PostgreSQL is not configured.

Audit Ingest Lag​

One lag card per device or cluster shows the time since the most recent ingested audit event from that source. Use these cards to find devices that have stopped delivering events.

Card colorTime since last event
GreenLess than 5 minutes
Amber5 to 60 minutes
RedMore than 60 minutes

Audit Ingest Performance​

This panel shows throughput for completed ingest jobs.

  • Per-type summary pills — average events per second, total events inserted, and job count for the last 10 completed jobs of each ingest type (PowerScale, ECS SSH, ECS NFS).
  • Total Events by Storage Node — a horizontal bar chart of accumulated event counts per source device. Click a bar to filter the throughput chart to that device.
  • Ingest Throughput per Job — a scatter chart of events per second for each completed ingest job. Colors identify the job type: PowerScale is purple, ECS NFS is teal, and ECS SSH is amber. Hover over a dot to see the job ID, device name, event rate, and insert count.

API Call Statistics​

This table lists every API endpoint group with metrics accumulated since the last backend restart.

ColumnMeaning
Call countTotal requests handled by the endpoint group
Error ratePercentage of requests that returned an error response
Average / p95 / max latencyResponse time distribution. p95 is the 95th percentile: 95% of requests were faster than this value.
SparklinesSmall inline charts of the recent call rate

Use the table to find slow endpoints and endpoint groups with elevated error rates. ECS Management API entries are highlighted in sky blue.

Disk I/O​

This panel appears on Linux hosts only. For each block device visible to the operating system, it shows read and write IOPS, I/O latency, and utilization percentage. High utilization on the PVC-backed device can explain slow ingest or slow ClickHouse writes.

Workflows​

Diagnose high memory usage​

  1. In the JVM Memory and CPU panel, check the Heap used bar. If it is near max, memory pressure is the likely cause.
  2. Check the GC run count and time. Frequent GC with a long cumulative time confirms heap pressure.
  3. In Active Operation Spans, look for an operation that has run unusually long. It may hold references to large objects.

Investigate a stalled ingest​

  1. Open the Audit Ingest Lag panel.
  2. Find the card for the affected device and check the lag value.
  3. If the lag is red and growing, check the Endpoint Reachability table for that device. It may be unreachable.
  4. Check Active Operation Spans for an ingest span that appears stuck.

Identify the slowest API endpoint​

  1. Open the API Call Statistics table.
  2. Sort by p95 latency in descending order.
  3. Investigate the top entries. Compare them with Active Operation Spans to see whether a long-running operation corresponds to that endpoint.

Tips​

  • If the heap stays above 80% of max, consider increasing the heap allocation in the deployment configuration before the process runs out of memory.
  • API call statistics reset on each backend restart. After a recent restart, the counts and averages reflect only the current session.