Agents Upgrade
The Agents Upgrade tab pushes new Docker image builds to your external pipeline agents (Extraction, Linguistic Coherence (LC), and Classify). These agents are the containers that run on your own Docker hosts, not inside the appliance. Use the tab when a new agent build is ready and you want to roll it out to some or all of your registered agents.
Where: System Operations > Fleet Management > Agents Upgrade
Each stage has its own card, because each stage runs its own independent image.
Prod and Staged Builds
Each stage card shows two builds:
| Build | Description |
|---|---|
| Prod (currently deployed) | The officially current build. New installs and re-runs of the installer install this build. It changes only when you click Promote to Prod. The Will install line on Agent Installation shows it too. |
| Staged (pending review) | A build that you uploaded but have not promoted. Uploads always go here, never to prod, and an upload never pushes anything to an agent by itself. |
Until a build is promoted, Prod reads nothing deployed yet. With no build staged, Staged reads — nothing staged —.
Nothing here happens automatically. A successful update does not promote the staged build, and a failed update does not clear it. Prod and staged change only when you click a button.
Upload a New Image
- Choose the
.tarfile for the stage that you are updating. This is that stage's agent image, downloaded from the Superna support site. It is adocker saveof the image, tagged with its real version (for exampleextraction-agent:<version>). - Click Upload. A progress bar tracks the upload. When the console accepts the file, the card says Uploaded <stage>-agent — detected version <version>.
A tar that you build yourself must hold exactly one image under one tag. A host refuses a tar that carries several images or several names.
Give every rebuild a new tag. A host refuses an image that reuses the tag of the build it is running now, or of the build it would roll back to, unless it is that same image. Loading it would replace the bytes that the tag names, and the host would have nothing left to roll back to.
The console validates the tar before it accepts it:
- The file must be a valid Docker image archive.
- The image's embedded tag must start with the prefix that the stage expects:
extraction-agent:,lc-agent:, orclassify-agent:. Anything else is rejected.
You do not enter a version. The console reads it from the tag (the part after the :), so the version shown everywhere in this tab is the one the image was built and tagged with. An accepted upload becomes the new Staged build for the stage.
Promote or Clear the Staged Build
When you are satisfied with a staged build, typically after you update one or more agents and confirm that it succeeds, choose one of these buttons:
- Promote to Prod — replaces prod with the staged build and clears staged. From then on, fresh installs get this build. Use it when you have finished the rollout.
- Clear — discards the staged build without promoting it, for example after a failed test that you abandon instead of fixing. Prod is not changed either way.
If the build you are about to promote already failed and rolled back on one of your agents, Promote to Prod asks you to confirm. The prompt names the agent and the recorded failure reason, and the card shows the warning in amber. The warning is not a hard block. You can confirm if you have fixed the cause or want to roll the build out to your other agents anyway. After you confirm, that build becomes what every fresh install and future agent registration pulls.
Both buttons are greyed out when nothing is staged, and while any agent of the stage has an update in progress (Pending). Wait for the update to finish (Done or Rolled back) before you promote or clear. Upload and both buttons are also disabled while an upload, promote, or clear runs for that stage.
Update Agents
Each stage card lists every registered agent of that type.
| Column | Description |
|---|---|
| Currently Running Image | The build that the agent runs right now. |
| Upgrading To | The build that an update in progress targets. After a failure, it shows the build that most recently failed and rolled back, so you can always tell which build did not work, even if you have since uploaded another build. |
A stage with no registered agents says No <stage> agents registered.
To update agents:
- Select the agents to move to the staged build, or use Select all.
- Click Update Selected (N). The build that it pushes (version and checksum) is shown beside the button and in its tooltip.
Update Selected always targets Staged, never prod directly. It never consumes or clears the staged build. You can select more agents and click Update Selected again later against the same staged build, without uploading anything again, until you promote or clear it.
The update then runs automatically in the background. The agent downloads the new image. A watcher process on its host has the image checked with the console, which confirms that it holds exactly that file, and swaps in the container. This typically takes a couple of minutes.
Each agent's status updates live:
| Status | Meaning |
|---|---|
| Pending | The agent has been told to update and has not picked it up yet, or is in the middle of the swap. |
| Done | The agent runs the new build. |
| Rolled back | The new build failed a health check after the swap, or the host refused to load it (for example because the console no longer holds that exact file). The agent stays on, or returns to, its previous working image. Click the badge to see the recorded failure reason. |
| Outdated / Up to date | Shown for agents that have not been updated yet. The status compares the running build with Prod, not with staged. "Up to date" means that the agent matches the officially current build. |
The agent is stopped while its service swaps images, both while the new image is checked, loaded, and started, and again for a rollback. Expect a short gap in that agent's work during any update attempt.
Cancel an Update
A Pending row has a Cancel button. It withdraws the update: the agent stops being offered the build on its next heartbeat, and its badge returns to Outdated or Up to date. Cancel does not undo a swap that the host has already started.
Failed Updates
If an update fails and rolls back, the agent refuses that same build automatically. It does not retry a known-bad image. To try again, upload a corrected build (a different file) and select the agent again.
An update that could not be checked is different. This happens when the download arrived damaged, the host could not reach the console to confirm the file, or the host had no room to verify it. Nothing is known to be wrong with the build, so it is not marked as bad.
- The agent downloads the build again on its own, up to three attempts in a row. After the third, the update is treated as failed, so that a problem on the host cannot keep a multi-GB download going forever.
- The recorded reason says which case it was. The journal of the agent's service on that host has the detail. See Troubleshooting.
- The host itself must be able to reach the console over HTTPS for this check, not only the agent container.
A build that fails its health check is marked as bad only when the build it rolled back to comes back healthy. A new build counts as healthy only after the console has answered one of its heartbeats. If the console cannot be reached during the check (for example, it is restarting), neither build can pass. The new build is then not marked as bad. The agent stays on its previous build, the update stays Pending, and it is offered again when the console answers. The failure reason is kept in the host's upgrade-failed.txt.
Stopped Services
While an agent's service is stopped on its host (systemctl stop argus-agent@<stage>), the host does not act on an update or a rollback. It waits until the service runs again, so an update never undoes a stop. The agent stays Pending until then.
If you stop the service in the middle of an update, only that attempt is abandoned. The agent starts on its previous build, the build is not marked as failed, and the update stays Pending. The agent downloads the build again once its service runs. Click Cancel to stop the update for good.
A host judges a loaded image only when its own watcher asked for it. A reboot, a crash restart, or a Docker restart never loads a downloaded image by itself. The watcher does it, then waits for the health of the new build.
Agents Deployed on Kubernetes
This mechanism applies only to agents that run on your own external Docker hosts.
- Helm deployment — if your entire deployment was installed with the Kubernetes Helm chart, this tab and Agent Installation are disabled. The tab buttons are greyed out with an explanation, and the underlying actions refuse if called directly. A Helm deployment scales its pipeline agents by changing the pod replica count in your Helm values. Nothing here applies to it.
- Mixed fleet — if you run the OVA appliance and some of your pipeline agents are Kubernetes pods, those agents still appear in the list, because they register the same way. They are greyed out with not eligible and cannot be selected, because a Kubernetes pod cannot apply this kind of update. To update them, run
helm upgradewith a new image tag in your values file. Kubernetes handles the rolling update.
Roll Back Manually
Automatic rollback fires only when a newly pushed build fails its health check. To revert an agent that runs fine, run this command on the agent's own host. You cannot do it from this page.
sudo aisec-agent-upgrade --stage <extraction|lc|classify> --rollback
It reverts that one agent to the image it ran immediately before its last update.
Troubleshooting
Updates Stay on Pending
If an update never reaches Done or Rolled back, the host-side watcher that drives the swap may not be running on that agent's host. The watcher runs as a systemd service outside Docker, under the unprivileged argus-upgrade account. To give up on the update, click Cancel on the agent's row.
To find out why, check the watcher on that host:
sudo systemctl status argus-agent-upgrade-watcher@<stage>.service
sudo journalctl -u argus-agent-upgrade-watcher@<stage>.service -n 50
The watcher logs only to the journal, so journalctl holds its history.
If the swap itself failed rather than the watcher, the agent's service records why. One restart of the service verifies and loads the new image and then starts the container, and root notes what it did for each request.
sudo journalctl -u argus-agent@<stage>.service -n 50 # loading the image and starting the container
cat /opt/argus-agents/status/<stage>.load /opt/argus-agents/status/<stage>.prepare # what root decided, and why
Health Check and Restarts
The watcher waits for the agent's own health check: its start period plus 90 seconds.
- A restart during that wait is normal when the console holds a newer agent code package. The agent applies the package and restarts itself once. Only repeated restarts count as a failed build.
- Only the agent's own health endpoint counts, and only while its service runs. On a single-role host that endpoint is port 9100. Another program listening there (node_exporter uses 9100 by default) makes every update roll back. A fresh install refuses such a host up front, where
ssis installed. On a host converted from the old layout, check the port yourself.
Build Too Old for the Host's Network
On a host that runs all three agents on the host's network, an update or rollback to a build that is too old to share that network is refused. The host keeps the build it runs, and the agent's note says the image "is too old to share this host's network". Upload a current build.
Trigger Held
If the watcher's journal says trigger held, it is waiting on purpose:
- The agent's service is stopped. Start it again.
- The host still has the older docker compose layout. Re-run the full installer to finish converting it.
Watcher Not Running
If the watcher is not running at all, re-run the one-line installer on that host, exactly as you originally ran it, with the same ARGUS_INSTALL_DIR and any other install-time variables. Add the --refresh-watcher flag. It reinstalls only the watcher, the systemd services, and their helper scripts, without touching the running containers, .env, or images. It refreshes the roles already installed on the host. The command names no role, and none is needed. Then trigger the update again.
--refresh-watcher refuses a host that is still on the older docker compose layout, which predates the per-agent systemd services. Run the full installer without the flag on such a host. It converts the host, and its agents keep their identities.
For an install that used the default location, run:
curl -fsSk "<console-url>/api/agents/container/install.sh?token=<enrollment-token>&imageToken=<image-token>" -o install.sh && sudo -E bash install.sh --refresh-watcher
Use sudo -E, not plain sudo. The default env_reset of sudo drops ARGUS_INSTALL_DIR and AISEC_KEK_FILE, so a plain sudo re-run on a host installed with either override silently reverts to the default paths.
If you installed to a custom directory or a custom key path, pass them explicitly:
sudo env ARGUS_INSTALL_DIR=/opt/data/aisec/argus-agents AISEC_KEK_FILE=/opt/data/aisec/kek bash install.sh --refresh-watcher
Keep the -f in the curl command:
- The console returns its rejection reasons (a malformed token or a bad role) as an HTTP 200 whose body prints the reason and exits non-zero, so
-fdoes not discard them. - A rotated token is not refused there. The script runs and fails at the image download with HTTP 401.
- Without
-f, a console that is restarting answers with an nginx 502 page.curlsaves it asinstall.sh, andsudothen runs it as root.