System Monitor
The System Monitor page gives real-time visibility into application memory and CPU health, external endpoint reachability, audit ingest performance, and API performance. Use it first when you investigate slowness, high memory usage, or a suspected stalled ingest.
Where: System Operations > System Monitor
Overview
The page is divided into panels, each covering one aspect of system health. The page refreshes every 5 seconds, and the memory and CPU panel refreshes every 15 seconds. A countdown timer shows the time until the next refresh of that panel.
The panels show the current snapshot only. The exception is the API Call Statistics table, which accumulates counts since the last backend restart.
JVM Memory and CPU
This panel shows memory and CPU metrics for the running application process.
| Metric | Meaning |
|---|---|
| Heap used / committed / max | Shown as a bar. Heap used consistently near max puts the process at risk of running out of memory. |
| Non-heap memory | Memory the JVM uses outside the heap, including code cache and compiler buffers. |
| Direct buffers | Off-heap buffers allocated by NIO channels and similar APIs. High direct buffer usage together with high heap can indicate I/O-heavy operations. |
| Metaspace | Memory used for class metadata. It grows with the number of loaded classes and does not shrink unless classes are unloaded. |
| Thread count (total / daemon / peak) | Current number of threads, how many are daemon threads, and the peak since startup. |
| GC run count and time | Number of garbage collection cycles and the cumulative time spent in them. High GC time relative to uptime can indicate heap pressure. |
| System memory | Total and free memory on the host OS, as reported by the JVM. |
Active Operation Spans
This panel lists the operations running now, with the wall-clock time elapsed and the CPU time consumed. Each span is a tracked operation such as an ingest run, a ClickHouse query, or a copy job step.
An operation that stays here for an unusually long time can indicate a stuck job or a slow dependency, such as a network timeout waiting on an external device.
Recent Operation Spans
This panel lists the most recently completed operations, sorted by completion time, with their wall-clock and CPU duration. Use it to review what the system did in the last few minutes and to spot operations that took unexpectedly long.
Endpoint Reachability
This table shows whether each configured external endpoint is reachable. Only the endpoints that apply to your deployment are shown. The panel refreshes every 60 seconds, and a refresh button in the panel header refreshes it immediately.
| Column | Meaning |
|---|---|
| Endpoint | Endpoint name |
| Type | Badge: Storage, ClickHouse, or App Database |
| Status | UP, DOWN, or Not configured |
| Latency | Response time in milliseconds |
| URL | Endpoint address |
| Last checked | Time of the last check |
Rows with DOWN status are highlighted in red.
ClickHouse and PostgreSQL health
These summary panels link to the detailed health pages.
- ClickHouse Summary — connection status, latency, version, and compressed and uncompressed table size for the ClickHouse analytics database. It links to the ClickHouse Health page.
- PostgreSQL Summary — connection status, latency, version, active connections, idle connections, pool utilization percentage (green, amber, or red), total connections, maximum connections, and database size for the application PostgreSQL database. It links to the PostgreSQL Database Health page. The panel is hidden when PostgreSQL is not configured.
Audit Ingest Lag
One lag card per device or cluster shows the time since the most recent ingested audit event from that source. Use these cards to find devices that have stopped delivering events.
| Card color | Time since last event |
|---|---|
| Green | Less than 5 minutes |
| Amber | 5 to 60 minutes |
| Red | More than 60 minutes |
Audit Ingest Performance
This panel shows throughput for completed ingest jobs.
- Per-type summary pills — average events per second, total events inserted, and job count for the last 10 completed jobs of each ingest type (PowerScale, ECS SSH, ECS NFS).
- Total Events by Storage Node — a horizontal bar chart of accumulated event counts per source device. Click a bar to filter the throughput chart to that device.
- Ingest Throughput per Job — a scatter chart of events per second for each completed ingest job. Colors identify the job type: PowerScale is purple, ECS NFS is teal, and ECS SSH is amber. Hover over a dot to see the job ID, device name, event rate, and insert count.
API Call Statistics
This table lists every API endpoint group with metrics accumulated since the last backend restart.
| Column | Meaning |
|---|---|
| Call count | Total requests handled by the endpoint group |
| Error rate | Percentage of requests that returned an error response |
| Average / p95 / max latency | Response time distribution. p95 is the 95th percentile: 95% of requests were faster than this value. |
| Sparklines | Small inline charts of the recent call rate |
Use the table to find slow endpoints and endpoint groups with elevated error rates. ECS Management API entries are highlighted in sky blue.
Disk I/O
This panel appears on Linux hosts only. For each block device visible to the operating system, it shows read and write IOPS, I/O latency, and utilization percentage. High utilization on the PVC-backed device can explain slow ingest or slow ClickHouse writes.
Workflows
Diagnose high memory usage
- In the JVM Memory and CPU panel, check the Heap used bar. If it is near max, memory pressure is the likely cause.
- Check the GC run count and time. Frequent GC with a long cumulative time confirms heap pressure.
- In Active Operation Spans, look for an operation that has run unusually long. It may hold references to large objects.
Investigate a stalled ingest
- Open the Audit Ingest Lag panel.
- Find the card for the affected device and check the lag value.
- If the lag is red and growing, check the Endpoint Reachability table for that device. It may be unreachable.
- Check Active Operation Spans for an ingest span that appears stuck.
Identify the slowest API endpoint
- Open the API Call Statistics table.
- Sort by p95 latency in descending order.
- Investigate the top entries. Compare them with Active Operation Spans to see whether a long-running operation corresponds to that endpoint.
Tips
- If the heap stays above 80% of max, consider increasing the heap allocation in the deployment configuration before the process runs out of memory.
- API call statistics reset on each backend restart. After a recent restart, the counts and averages reflect only the current session.