Agent Dashboard
The Agent Dashboard tab charts the capacity, throughput, and health of every registered agent over time. Every heartbeat carries a sample of the agent's host and process metrics, which the console stores in ClickHouse. The tab draws that history as one time-series line per agent.
Use it to spot an agent that runs hot, an agent with slow disk, or a JVM agent under heap pressure. You can also use it to confirm fleet health before or after a change.
Where: System Operations > Fleet Management > Agent Dashboard
The tab has two parts, from top to bottom:
- Agent Health Trends — the metric chart, with a metric selector above it.
- All Agents — a grid of every registered agent, with bulk actions (drain, upgrade, max jobs). The bulk actions are available only here, not on the Live Agents tab. See All Agents Grid.
Agent Health Trends
The chart plots one colored line per agent for the selected metric. The legend names each agent. The Y axis shows the unit of the metric (%, MiB, ms, sockets, Mbps, rows, q/min, B/min, rows/min, or s) and the X axis shows time.
Hover over any point to open a tooltip with the exact value for every agent at that time. The tooltip also shows how many jobs the agent had active at that moment. For a pipeline agent, this is the number of files in progress.
A new sample arrives with every heartbeat (every 30 seconds), so the lines fill in continuously while agents are online.
Choosing a Metric
The tabs above the chart select the metric. Only the selected chart renders at a time. The metrics are grouped into Health, Network, and Trino.
Health Metrics
| Metric | Unit | Description | Applies to |
|---|---|---|---|
| Host CPU | % | Host CPU utilization across all cores. | All agents |
| Container CPU | % | CPU utilization inside the container. | Container agents |
| Memory Used | % | Container memory used as a percentage of its limit. | Container agents |
| Memory Used (MiB) | MiB | Absolute container memory. | Container agents |
| JVM Heap Used | % | JVM heap utilization. | Java agents |
| JVM CPU | % | CPU consumed by the JVM process. | Java agents |
| Disk Read Latency | ms | Average read latency on the agent host. | All agents |
| Disk Write Latency | ms | Average write latency on the agent host. | All agents |
| CH Buffer Peak (per heartbeat) | rows | Peak depth of the agent's ClickHouse write buffer between heartbeats. A rising buffer means that the agent generates rows faster than it can write them to ClickHouse. | Java agents |
Network Metrics
| Metric | Unit | Description |
|---|---|---|
| TCP In-Use (Established) | sockets | Established and active TCP sockets. |
| TCP TIME_WAIT | sockets | Sockets in TIME_WAIT (recently closed, awaiting teardown). |
| TCP Orphan | sockets | Orphaned sockets with no file descriptor, awaiting teardown. |
| Egress (Tx) | Mbps | Outbound network throughput. |
| Ingress (Rx) | Mbps | Inbound network throughput. |
Trino Metrics
These metrics are reported by Trino agents only.
| Metric | Unit | Description |
|---|---|---|
| Query Rate | q/min | Queries per minute. |
| Active Queries | queries | Queries currently running. |
| Queued Queries | queries | Queries waiting to run. |
| Failed Queries | queries | Queries that failed. |
| Bytes Read | B/min | Data read per minute. |
| Bytes Written | B/min | Data written per minute. |
| Rows Processed | rows/min | Rows processed per minute. |
| JVM Uptime | s | Uptime of the Trino JVM. |
| JVM Threads | threads | Thread count of the Trino JVM. |
Agent-Type Badges
Some tabs carry a JAVA, CONTAINER, or TRINO badge. The chart header repeats it as a JAVA AGENTS ONLY, CONTAINER AGENTS ONLY, or TRINO AGENTS ONLY pill. The badge shows which agent type reports the metric.
- JAVA metrics (JVM Heap Used, JVM CPU, and CH Buffer Peak) are reported only by Java agents. Python pipeline agents are drawn as flat lines at 0 on these charts.
- CONTAINER metrics (Container CPU, Memory Used, and Memory Used in MiB) are reported only by containerized agents. Java agents that run on bare metal are drawn flat at 0.
- TRINO metrics are reported by Trino agents only. Other agents are drawn flat at 0.
- Tabs with no badge (for example Host CPU, disk latency, TCP, and network) are reported by all agent types.
Every agent with a sample in the selected time range is drawn on every chart, at 0 where the metric does not apply. Flat lines at 0 for the wrong agent type are expected, not a fault.
A chart is empty only when no agent has a sample in the range. The message in the chart area then names the agent type that the metric needs ("No JVM agents in this fleet…" or "No containerized agents reporting yet…"), or says "No data yet — agents write on each heartbeat (every 30 s)."
Fleet Total Line
On the Egress (Tx) and Ingress (Rx) charts, and on the Trino Bytes Read, Bytes Written, and Rows Processed charts, a dashed line labeled Fleet Total appears when more than one agent reports. It is the sum of all agents' throughput at each point in time, so you can read the combined bandwidth of the whole fleet. The chart header notes "— dashed line = fleet total" while the line is shown.
An agent that runs on its host's own network reports the host's traffic, not its own. This applies to every pipeline agent on a single-role host, and to all three agents on a host that runs all three roles on the host's network (any new install). Three agents on one such host each show the same traffic, so Fleet Total counts that host three times.
Time Range
The range selector in the top-right toolbar sets how far back the chart looks:
- 1h
- 6h
- 24h (default)
- 7d
Changing the range reloads the data for that window. Wider windows show longer trends with coarser detail. Samples are averaged into 30-second buckets for 1h, 5-minute buckets for 6h and 24h, and 1-hour buckets for 7d.
Auto-Refresh and Response Time
The controls to the right of the time range are:
- Auto-refresh interval — Auto 10s, Auto 30s (default), Auto 60s, Auto 5m, or Auto Off. While auto-refresh is on, the chart reloads at the chosen interval. The All Agents grid refreshes every 10 seconds regardless, even with Auto Off.
- Refresh in … — a countdown to the next automatic refresh, shown as
Xm Ysonce it is 60 seconds or more. Hidden when auto-refresh is Auto Off. - Response time (for example
142ms) — how long the last metric load took. A climbing value can indicate that ClickHouse is under load. - Manual refresh (circular-arrow icon) — reloads the chart and the grid immediately and resets the countdown.
All Agents Grid
Below the chart, All Agents lists every registered agent in a table:
| Column | Description |
|---|---|
| Name, IP, Status, Version, Registered, Last Seen | Identity and connection details of the agent. |
| Image | For external pipeline agents (Extraction, LC, Classify) only: the Docker image build that is running, shown as a version label and a short checksum fingerprint, for example <version> (0e92dc3dfe91…). Shows a dash for the console and for any agent that has never reported a running image, such as a pipeline pod deployed on Kubernetes. To push a new build, use Agents Upgrade. |
| Max Jobs | How many files the agent accepts at once. default means the agent's own limit. |
| Upgrade, Drain | The upgrade and drain state of each agent (Upgrading…, Pending, Draining). |
Drain All and Undrain All are always available. Select rows to act on them together with Upgrade (N), Cancel Upgrade (N), Drain (N), Undrain (N), and Max jobs…, which sets the job limit on the selected agents.
Workflows
Check Whether an Agent Is Overloaded
- Open the Agent Dashboard tab.
- Select Host CPU. For Java agents, also check JVM Heap Used.
- Set the time range to 24h to see the trend, or 1h to see the current state.
- Look for a line that stays high. Hover over a peak to confirm which agent it is and how many active jobs it had at that moment.
See Total Fleet Bandwidth During a Large Run
- Select the Egress (Tx) tab. The dashed Fleet Total line shows the combined outbound throughput of all agents.
- Select Ingress (Rx) to see inbound throughput.
Check the Pipeline Agents' Containers
- Select Container CPU and Memory Used to confirm that the pipeline agent pods are healthy.
- Check Memory Used (MiB) against the container limit for the stage. Linguistic Coherence (LC) agents in particular run close to their memory limit.
Troubleshooting
- A chart shows only flat lines at 0 — the metric applies to another agent type (see its badge), and other agents are drawn at 0. This is expected.
- A chart is empty with a JAVA, CONTAINER, or TRINO message — no agent has a sample in the selected time range.
- All charts are empty with "No data yet" — no agent has written a heartbeat sample yet. Wait one heartbeat interval (30 seconds) and refresh.
- A red error banner appears — the request for chart data to the console failed. If ClickHouse itself is not configured or not reachable, no banner appears and the charts show their empty message instead. Confirm that ClickHouse is configured and that agents are online and sending heartbeats.
- An agent's line stops — the agent has gone offline. Check its status in the All Agents grid below the chart.
- CH Buffer Peak rises on a Java agent — the agent produces rows faster than it can write them to ClickHouse. Check the agent's ClickHouse connectivity and the health of the database.