Skip to main content
Migration Notice
We're migrating documentation from the old portal into this one. Some things may look a little different or out of place in the meantime — we know, and we're working to get it right. If something's unclear or doesn't look right, let us know.

Agent Dashboard

The Agent Dashboard tab charts the capacity, throughput, and health of every registered agent over time. Every heartbeat carries a sample of the agent's host and process metrics, which the console stores in ClickHouse. The tab draws that history as one time-series line per agent.

Use it to spot an agent that runs hot, an agent with slow disk, or a JVM agent under heap pressure. You can also use it to confirm fleet health before or after a change.

Where: System Operations > Fleet Management > Agent Dashboard

The tab has two parts, from top to bottom:

  1. Agent Health Trends — the metric chart, with a metric selector above it.
  2. All Agents — a grid of every registered agent, with bulk actions (drain, upgrade, max jobs). The bulk actions are available only here, not on the Live Agents tab. See All Agents Grid.

The chart plots one colored line per agent for the selected metric. The legend names each agent. The Y axis shows the unit of the metric (%, MiB, ms, sockets, Mbps, rows, q/min, B/min, rows/min, or s) and the X axis shows time.

Hover over any point to open a tooltip with the exact value for every agent at that time. The tooltip also shows how many jobs the agent had active at that moment. For a pipeline agent, this is the number of files in progress.

A new sample arrives with every heartbeat (every 30 seconds), so the lines fill in continuously while agents are online.

Choosing a Metric​

The tabs above the chart select the metric. Only the selected chart renders at a time. The metrics are grouped into Health, Network, and Trino.

Health Metrics​

MetricUnitDescriptionApplies to
Host CPU%Host CPU utilization across all cores.All agents
Container CPU%CPU utilization inside the container.Container agents
Memory Used%Container memory used as a percentage of its limit.Container agents
Memory Used (MiB)MiBAbsolute container memory.Container agents
JVM Heap Used%JVM heap utilization.Java agents
JVM CPU%CPU consumed by the JVM process.Java agents
Disk Read LatencymsAverage read latency on the agent host.All agents
Disk Write LatencymsAverage write latency on the agent host.All agents
CH Buffer Peak (per heartbeat)rowsPeak depth of the agent's ClickHouse write buffer between heartbeats. A rising buffer means that the agent generates rows faster than it can write them to ClickHouse.Java agents

Network Metrics​

MetricUnitDescription
TCP In-Use (Established)socketsEstablished and active TCP sockets.
TCP TIME_WAITsocketsSockets in TIME_WAIT (recently closed, awaiting teardown).
TCP OrphansocketsOrphaned sockets with no file descriptor, awaiting teardown.
Egress (Tx)MbpsOutbound network throughput.
Ingress (Rx)MbpsInbound network throughput.

Trino Metrics​

These metrics are reported by Trino agents only.

MetricUnitDescription
Query Rateq/minQueries per minute.
Active QueriesqueriesQueries currently running.
Queued QueriesqueriesQueries waiting to run.
Failed QueriesqueriesQueries that failed.
Bytes ReadB/minData read per minute.
Bytes WrittenB/minData written per minute.
Rows Processedrows/minRows processed per minute.
JVM UptimesUptime of the Trino JVM.
JVM ThreadsthreadsThread count of the Trino JVM.

Agent-Type Badges​

Some tabs carry a JAVA, CONTAINER, or TRINO badge. The chart header repeats it as a JAVA AGENTS ONLY, CONTAINER AGENTS ONLY, or TRINO AGENTS ONLY pill. The badge shows which agent type reports the metric.

  • JAVA metrics (JVM Heap Used, JVM CPU, and CH Buffer Peak) are reported only by Java agents. Python pipeline agents are drawn as flat lines at 0 on these charts.
  • CONTAINER metrics (Container CPU, Memory Used, and Memory Used in MiB) are reported only by containerized agents. Java agents that run on bare metal are drawn flat at 0.
  • TRINO metrics are reported by Trino agents only. Other agents are drawn flat at 0.
  • Tabs with no badge (for example Host CPU, disk latency, TCP, and network) are reported by all agent types.

Every agent with a sample in the selected time range is drawn on every chart, at 0 where the metric does not apply. Flat lines at 0 for the wrong agent type are expected, not a fault.

A chart is empty only when no agent has a sample in the range. The message in the chart area then names the agent type that the metric needs ("No JVM agents in this fleet…" or "No containerized agents reporting yet…"), or says "No data yet — agents write on each heartbeat (every 30 s)."

Fleet Total Line​

On the Egress (Tx) and Ingress (Rx) charts, and on the Trino Bytes Read, Bytes Written, and Rows Processed charts, a dashed line labeled Fleet Total appears when more than one agent reports. It is the sum of all agents' throughput at each point in time, so you can read the combined bandwidth of the whole fleet. The chart header notes "— dashed line = fleet total" while the line is shown.

An agent that runs on its host's own network reports the host's traffic, not its own. This applies to every pipeline agent on a single-role host, and to all three agents on a host that runs all three roles on the host's network (any new install). Three agents on one such host each show the same traffic, so Fleet Total counts that host three times.

Time Range​

The range selector in the top-right toolbar sets how far back the chart looks:

  • 1h
  • 6h
  • 24h (default)
  • 7d

Changing the range reloads the data for that window. Wider windows show longer trends with coarser detail. Samples are averaged into 30-second buckets for 1h, 5-minute buckets for 6h and 24h, and 1-hour buckets for 7d.

Auto-Refresh and Response Time​

The controls to the right of the time range are:

  • Auto-refresh interval — Auto 10s, Auto 30s (default), Auto 60s, Auto 5m, or Auto Off. While auto-refresh is on, the chart reloads at the chosen interval. The All Agents grid refreshes every 10 seconds regardless, even with Auto Off.
  • Refresh in … — a countdown to the next automatic refresh, shown as Xm Ys once it is 60 seconds or more. Hidden when auto-refresh is Auto Off.
  • Response time (for example 142ms) — how long the last metric load took. A climbing value can indicate that ClickHouse is under load.
  • Manual refresh (circular-arrow icon) — reloads the chart and the grid immediately and resets the countdown.

All Agents Grid​

Below the chart, All Agents lists every registered agent in a table:

ColumnDescription
Name, IP, Status, Version, Registered, Last SeenIdentity and connection details of the agent.
ImageFor external pipeline agents (Extraction, LC, Classify) only: the Docker image build that is running, shown as a version label and a short checksum fingerprint, for example <version> (0e92dc3dfe91…). Shows a dash for the console and for any agent that has never reported a running image, such as a pipeline pod deployed on Kubernetes. To push a new build, use Agents Upgrade.
Max JobsHow many files the agent accepts at once. default means the agent's own limit.
Upgrade, DrainThe upgrade and drain state of each agent (Upgrading…, Pending, Draining).

Drain All and Undrain All are always available. Select rows to act on them together with Upgrade (N), Cancel Upgrade (N), Drain (N), Undrain (N), and Max jobs…, which sets the job limit on the selected agents.

Workflows​

Check Whether an Agent Is Overloaded​

  1. Open the Agent Dashboard tab.
  2. Select Host CPU. For Java agents, also check JVM Heap Used.
  3. Set the time range to 24h to see the trend, or 1h to see the current state.
  4. Look for a line that stays high. Hover over a peak to confirm which agent it is and how many active jobs it had at that moment.

See Total Fleet Bandwidth During a Large Run​

  1. Select the Egress (Tx) tab. The dashed Fleet Total line shows the combined outbound throughput of all agents.
  2. Select Ingress (Rx) to see inbound throughput.

Check the Pipeline Agents' Containers​

  1. Select Container CPU and Memory Used to confirm that the pipeline agent pods are healthy.
  2. Check Memory Used (MiB) against the container limit for the stage. Linguistic Coherence (LC) agents in particular run close to their memory limit.

Troubleshooting​

  • A chart shows only flat lines at 0 — the metric applies to another agent type (see its badge), and other agents are drawn at 0. This is expected.
  • A chart is empty with a JAVA, CONTAINER, or TRINO message — no agent has a sample in the selected time range.
  • All charts are empty with "No data yet" — no agent has written a heartbeat sample yet. Wait one heartbeat interval (30 seconds) and refresh.
  • A red error banner appears — the request for chart data to the console failed. If ClickHouse itself is not configured or not reachable, no banner appears and the charts show their empty message instead. Confirm that ClickHouse is configured and that agents are online and sending heartbeats.
  • An agent's line stops — the agent has gone offline. Check its status in the All Agents grid below the chart.
  • CH Buffer Peak rises on a Java agent — the agent produces rows faster than it can write them to ClickHouse. Check the agent's ClickHouse connectivity and the health of the database.