Skip to main content
Migration Notice
We're migrating documentation from the old portal into this one. Some things may look a little different or out of place in the meantime — we know, and we're working to get it right. If something's unclear or doesn't look right, let us know.

Fleet Network

The Network tab shows the current TCP socket state of every agent in the fleet in one table. It helps you answer one question: is an agent leaking or piling up network connections? Each agent reports its socket counters on every heartbeat, and the tab shows the latest values side by side so that you can compare the fleet and spot an agent that behaves differently.

Where: System Operations > Fleet Management > Network

The counters are read from the Linux kernel, so they are populated only for Linux agents. macOS and Windows agents report zeros, which the table shows as 0.

An agent on its host's own network sees the whole host's sockets. This applies to every pipeline agent on a single-role host, and to all three agents on a host that runs all three roles on the host's network. Agents that share a host therefore show the same counters, and a leak in one shows on all of them.

Overview​

Copy and classification work opens many short-lived TCP connections, especially against S3 and object endpoints, where each object operation is a separate HTTP connection. Under heavy load, it is normal for sockets to accumulate in TIME_WAIT and for the in-use count to climb.

The problem case is the opposite: counts that stay high while the agent is idle, or an Orphan count that does not return to zero. These patterns point to leaked or stuck connections. They can exhaust the ephemeral port range of the host and stall the agent. This tab lets you catch that early, before jobs fail.

TCP Socket State Table​

The table is titled TCP Socket State — Current, with the note from last agent heartbeat · Linux only · zeros on macOS/Windows agents. It has one row per agent, plus the local console entry, tagged (console), whose counters always read —.

ColumnDescription
AgentThe agent's display name. The console row has a (console) suffix.
StatusThe agent's connection status: Online, Degraded, or Offline.
TCP In-UseSockets in ESTABLISHED and other active states. Java agents add IPv6 sockets. Values above 500 are shown in amber.
TIME_WAITConnections that wait through the 2-MSL timeout after a graceful close. Values above 200 are shown in amber.
OrphanSockets with no file descriptor that are pending teardown. Values above 10 are shown in red.
Last SeenTime since the agent last reported, for example just now, 30s ago, 5m ago, 2h ago, or never. The console row always reads just now, and its Status is always Online.

When a row has no socket data, for example the console row or an agent that has not sent a heartbeat since the feature was deployed, the three counter cells show —. If the TCP In-Use value is missing, the whole row shows as no data.

What the Counters Mean​

A card below the table explains how to read each counter:

  • TCP In-Use — high values during a copy job are normal. High values at idle can indicate leaked connections.
  • TIME_WAIT — a large count (200 or more) under sustained S3 load is expected, because each object operation uses a short-lived HTTP connection. It is a problem only if sockets exhaust the ephemeral port range.
  • Orphan — normally close to zero. Persistent non-zero values can indicate FIN_WAIT1 accumulation from aborted connections.

How Fresh the Data Is​

The values are a snapshot from the agent's last heartbeat, which is about every 30 seconds. They are not a live socket dump. The agent list refreshes every 10 seconds, so the table updates without a manual reload. Use the Last Seen column to confirm that a row is current. A stale Last Seen means that the counters in that row are stale too.

For history instead of a single snapshot, open the Agent Dashboard tab. Under the Network group, it charts the same counters over time as TCP In-Use (Established), TCP TIME_WAIT, and TCP Orphan, next to the Egress (Tx) and Ingress (Rx) bandwidth charts.

Workflows​

Spot an Agent That Leaks Connections​

  1. Open Fleet Management and select the Network tab.
  2. Look in the TCP In-Use column for an amber value (over 500) on an agent that should be idle.
  3. Check Last Seen to confirm that the reading is current.
  4. Open the Agent Dashboard tab and select TCP In-Use (Established). A count that climbs steadily indicates a leak. A spike during a job is normal.

Confirm That a High TIME_WAIT Count Is Harmless​

  1. On the Network tab, find the agent with the amber TIME_WAIT value.
  2. Check whether the agent runs sustained S3 or object workloads. If it does, a high count is expected and is not a fault.
  3. Investigate further only if the agent also reports failures or the host is close to its ephemeral port limit.

Tips​

  • A — in all three counter columns is not an error. The row has no socket data yet, or it is the console row. Only Linux agents populate the counters. macOS and Windows agents show 0.
  • Amber (TCP In-Use and TIME_WAIT) is an attention cue, not an alarm. Red on Orphan needs attention, because a persistently non-zero count points to connections that never finished closing.
  • The console row is a placeholder for the console itself. Its counters are always —.