Jobs
The Jobs page shows the history and current state of the system jobs that keep this product's data running. Use it to monitor what the system is doing, investigate failures, and trigger manual runs.
Where: System Operations > System Jobs
The page covers these job types: Kubernetes scans, audit-event ingest from PowerScale (Dell OneFS) and ObjectScale (Dell ECS), ObjectScale (Dell ECS) inventory scans, data exposure risk discovery, and posture analytics.
Overview
The page has two tabs:
- Job Monitor — the job history table, summary stat cards, an Active Jobs strip, filters, and the job detail flyout.
- Schedules — per-job-type schedule configuration: enable toggle, interval picker, and Run Now.
The Job Monitor refreshes automatically every 30 seconds. The header shows a live countdown (auto in Ns), the last load time in milliseconds, and a ↻ button to refresh immediately.
Stat cards
| Card | Meaning |
|---|---|
| Total Jobs | Count of all jobs ever recorded |
| Completed | Jobs that finished successfully |
| Failed | Jobs that ended in an error state |
| Running | Jobs currently executing |
| Avg Duration | Mean wall-clock time across all completed jobs of any type |
Active Jobs strip
Jobs that are running, pending, awaiting report, or auto-paused appear in the Active Jobs strip above the stat cards. Each entry shows:
- The job type and job ID.
- The device or configuration that started the job.
- A
completed / total stepsprogress count. - A Cancel button.
Job table
The table lists jobs newest first, 50 rows per page. Use ← Prev and Next → to page; the footer shows Page X of Y · N total jobs.
Filters
- Job Type — narrow to a specific job type. Only the job types available in your deployment are listed.
- Status — narrow to Completed, Failed, Running, Auto-Paused, Pending, or Cancelled.
- Device / Config — narrow to jobs started by a specific device or configuration.
Columns
Each row shows the job ID, type, device or configuration (with the classification agent name when an agent ran the job), status badge, start time, duration, and a completed / total steps progress count. The Trigger badge shows Auto for scheduled runs and Manual for runs started by a user.
Click View on a row to open the detail flyout. Click View again to close it.
Status values
| Status | Meaning |
|---|---|
| Running | Currently executing |
| Pending | Queued, not yet started |
| Completed | Finished successfully |
| Failed | Ended in an error state |
| Cancelled | Stopped by a user |
| Auto-Paused | Throttled by an HTTP 503 from the storage device; waiting to retry (see Auto-pause on throttling) |
| Awaiting Report | The job was running when the console restarted; waiting for the agent to reconnect and report the outcome |
Job types
| Job type | What it does |
|---|---|
| K8 Scan | Collects the Kubernetes inventory (pods, PVCs, mounts), runs RunAI correlation, and discovers COSI (Container Object Storage Interface) buckets and maps them to the pods that use them. |
| PowerScale Audit Ingest | Pulls SMB and NFS file-audit events from PowerScale clusters into ClickHouse. |
| ECS SSH Ingest | Pulls S3 object-audit events from ObjectScale devices into ClickHouse. |
| ECS Inventory Scan | Lists every namespace and bucket on an ObjectScale device and reconciles the bucket inventory. |
| Data Exposure Risk Processing | Runs the scheduled discovery that feeds the AI risk pipeline (data classification). It scans recent audit activity for new file and object writes and queues them for the classification agents. |
| Posture Analytics | Runs the predictive threat detection computation (see below). |
Header buttons start jobs on demand:
- K8 Scan Now — runs a Kubernetes scan.
- Audit Ingest Now — runs PowerScale and ObjectScale ingest in parallel, using the method each device is configured for, and reports the per-device result inline.
- Run Inventory Scan — on Settings > Storage, starts an ECS Inventory Scan.
Posture Analytics
Each Posture Analytics run:
- Recomputes statistical baselines (mean, standard deviation, median, median absolute deviation (MAD), and exponentially weighted moving average (EWMA)) of PII and content-integrity scores per file, PVC, and namespace over a trailing window.
- Runs the transition detectors against those baselines. The detectors look for PII verdict flips (clean to PII), new PII entity types, count spikes, and content-integrity drops (the encryption or corruption signature).
Detected transitions appear in the Predictive Threat Detection feed, and HIGH and CRITICAL transitions raise alarms. The job step summary reports how many transitions were detected.
The default interval is 30 minutes. Run Now triggers an immediate recompute and detection pass. A scheduled run is skipped while a K8 scan is running.
Job detail flyout
Click View to open the flyout on the right. It polls every 3 seconds while the job runs and stops polling when the job reaches a final state. It shows:
- Summary tiles — Status, Duration, Started, and Completed.
- Error banner — for a failed job, the reason and the step that failed.
- Awaiting Report banner — appears if the console restarted mid-run. The job is waiting for the agent to reconnect and report, and is marked failed automatically if no report arrives within the configured timeout.
- Resume banners — a Resume of Job #N banner on a job created to retry an earlier run, and a Superseded — resumed as Job #N banner on the original. Click either banner to jump to the linked job.
- Execution Steps — every step with its status, message, duration, and any detail badges.
Execution steps
A step with log output has a collapsible N log lines panel. The log for a running step expands automatically and refreshes every 3 seconds.
| Step status | Meaning |
|---|---|
| Running | The step is executing |
| Completed | The step finished successfully |
| Failed | The step ended in an error |
| Pending | The step has not started |
| Skipped | The job deliberately did not run the step |
A skipped step is shown rather than hidden so the full shape of the job stays visible, and it counts as done for the job's progress.
Ingest Summary
For PowerScale Audit Ingest and ECS SSH Ingest jobs, the flyout shows an Ingest Summary panel. The badges update live during a run and stay populated afterward.
| Badge | Meaning |
|---|---|
| Records Decoded | Raw audit records parsed across every file in the run. For PowerScale this is the count before filtering. For ObjectScale it is every line parsed from the compressed log. |
| Records Matched (PowerScale) | Records remaining after the CSI (Container Storage Interface) path filter and the 60-second summarizer. |
| Records Inserted (ObjectScale) | Records inserted. ObjectScale has no filter step. |
| Inserted to CH (ClickHouse) | Rows written to ClickHouse. This can be lower than Records Decoded when the summarizer merges several events into one row, and lower than Records Matched if a batch exhausts its retries. In dry-run mode the badge uses a different color. |
| Ingest Rate | Rows per second over the run's wall-clock duration. |
| Batch Size | Rows per insert for this run, set in Settings > Database > Ingest Batch Size (default 100,000). Smaller batches are gentler on a struggling ClickHouse cluster. Larger batches reduce HTTP overhead. |
| CH Errors | Batches that exhausted all 3 retry attempts and were lost. A non-zero count means those rows are not in ClickHouse. Check ClickHouse health, network connectivity, and ClickHouse disk space. |
| CH Retries | Total retry attempts across all batches. Each batch insert is retried up to 3 times, with 1, 2, and 4 second backoff. CH Retries above zero with CH Errors at zero means the retry layer absorbed transient instability, which is informational only. |
| Skipped | Shown only when above zero. Log files that were skipped because they were already fully ingested. A fresh cluster skips nothing. A re-run on an already ingested cluster skips many files. |
| Threads | Shown only when above 1. Planned total ingest parallelism for the run (node count multiplied by per-node threads). If the per-node value was reduced to stay under the global thread cap, the badge reads Threads (capped at N). |
The panel also shows:
- A Log Rollover Detected callout when a log rollover occurs.
- An Event Type Breakdown of the ingested events.
- A Node / File Details table with per-node, per-file state, records, offset and size, rate, and duration.
Below the badges, the File Ingest State — Last 24 Hours table lists one row per log file: node, file, byte offset, total records, status (IN_PROGRESS or COMPLETE), and last-updated time. It lists only files currently in progress and files completed in the last 24 hours. Click View all files → to open the full Ingest Offsets page.
Inventory Diff
For an ECS Inventory Scan, the Inventory Diff strip shows the outcome of the reconcile step.
| Counter | Meaning |
|---|---|
| Added | Buckets discovered in this scan that were not in the inventory |
| Updated | Known buckets that were re-confirmed. Their group, tag, and ACL assignments are preserved. |
| Removed | Buckets that were in the inventory but are no longer on the device. These are permanently deleted from the inventory. |
| Namespaces Removed | Whole namespaces that no longer exist on the device |
| Total After | Inventory entries remaining for the device after the diff is applied |
The reconcile step log lists each added bucket and removed bucket action. Removals apply only to what the scan fully listed. If listing a namespace fails, that namespace is skipped and its inventory is preserved, so a temporary outage cannot remove valid inventory.
K8 Scan details
The flyout for a K8 Scan shows four headline counts: Pods, PVCs, COSI Buckets, and Bucket→Pod Mappings. It also shows three diagnostic sections, each with a Refresh button and an empty-state message that explains why the section is empty:
- Resolved PVC Paths — the state the path resolver holds after the scan: each PVC's namespace, PowerScale path, PV, mount-cache entries, and natural-key prefix. Use it when audit rows show an empty PVC.
- COSI Bucket Inventory — the COSI
Bucketresources the scan discovered: claim namespace, bucket, device, account ID, credentials secret, and whether the bucket matched an existing inventory entry or was unmatched. - Resolved Bucket→Pod — the bucket-to-pod bindings the resolver knows about: pod, namespace, service account, bucket, account ID, and the kind of Secret reference that matched.
COSI exposes no direct pod reference, so the pod-to-bucket link is inferred by matching the bucket's credentials Secret name against each pod's volume, env, and envFrom Secret references. Two pods that mount the same Secret are both linked. The attribution is a candidate match, not an exact one.
Discovery Summary
For a Data Exposure Risk Processing job, the Discovery Summary shows the exact time range the run scanned on a Window: line, followed by a funnel for PowerScale files and, when ObjectScale is configured, a funnel for ObjectScale objects.
Scan window
Discovery is incremental. Each run scans only the file and object changes since the last successful run. The window start is the time of that last successful run.
- The window start is capped at 24 hours back. After downtime, or on the first run, the run scans at most the most recent day.
- The start point advances only after a fully successful pass over files and objects. A failed run leaves it unchanged, so the next run scans the same window again.
- A manual Run Now also advances the start point, so the next run starts where the manual run finished.
Files (PowerScale) funnel
| Cell | Meaning |
|---|---|
| Time period | The window scanned |
| Records found | Audit records found in the window |
| After exclude list | Records left after the exclude list. Shows how many exclude rules are configured and a − N excluded count for events dropped in this run. Rules match by extension, filename glob, or path substring. |
| Removed (net-deleted) | Files whose final state in the window is deleted |
| Collapsed (redundant touches) | Extra audit events folded into one work item per file |
| Skipped (recently scored) | Files an earlier run already scored |
| Files submitted | Files queued for classification |
A per-device breakdown shows Creates, Modifies, Renames, Excluded, Collapsed, and Net-deleted for each device.
Objects (ObjectScale) funnel
The funnel shows Buckets in scope, Objects found (PUT and POST with a 2xx response), Removed (net-deleted), Collapsed (redundant touches) with the − N excluded count in its sub-label, Skipped, and Objects submitted. The per-device breakdown includes Excluded, Collapsed, and Net-deleted. The exclude list matches the object key by extension, name glob, or path or key substring.
Net-deleted and collapse work the same way for objects, keyed on bucket and object key. An object that is created and then deleted within the window is dropped.
Object discovery covers only buckets that meet both conditions:
- The bucket is matched by COSI (listed as Discovered by COSI and bound to a pod through the COSI BucketAccess credentials Secret).
- The bucket has a read service account granted under Settings > Storage > ObjectScale Bucket Classification Service Account.
A read service account granted on a bucket that COSI did not match is not scanned. When no COSI-matched bucket is granted, the section reads Object discovery skipped. To include a COSI bucket, filter the settings page to Discovered by COSI and use Select all matching to grant read access.
Net-deleted and collapsed counts
Removed (net-deleted) counts files or objects that were created or written and then deleted within the window. Discovery replays each path's sequence of events. If the last event is a delete, the data no longer exists, so it is dropped and never queued. This avoids "No such file" and 404 errors from fetching deleted data.
- A delete followed by a re-create in the same window keeps the re-created item.
- If a delete arrives only in a later window, the item is still queued. The downstream pipeline handles the not-found result.
Collapsed (redundant touches) counts the extra audit events folded away by keeping one work item per file or object key per window. Only the latest write of each path is scored. A high count means a few busy files are touched repeatedly, rather than many distinct files changing.
These counts are independent of each other and are never counted twice. Neither is the same as Skipped (recently scored), which suppresses files an earlier run already scored, or duplicate-content deduplication, which happens later when two different files have identical content.
A callout appears if the funnel narrows unexpectedly, meaning audit events matched but nothing was submitted. It explains whether the exclude list filtered everything out or the files were deduplicated against existing pipeline work items.
Cancel jobs
Cancel appears on any running job, both in the Active Jobs strip and in the bulk toolbar. The job finishes its current unit of work before it stops, so the status can take a short time to change to Cancelled. A cancelled ingest job keeps the events it already wrote to ClickHouse.
Bulk actions
Use the checkbox column to select several jobs. The header checkbox selects all jobs on the current page. When jobs are selected, a toolbar appears with:
- Cancel (N) — sends a cancel signal to the selected jobs that are running, pending, auto-paused, or awaiting report. N is the number of selected jobs that are eligible.
- Clear selection — deselects all jobs.
Jobs that are not eligible are skipped.
Auto-pause on throttling
When a storage device returns an HTTP 503 "reduce your request rate" response, the affected job changes to Auto-Paused. The job waits a random hold-off of up to about 5 minutes and then resumes automatically. The status badge tooltip shows the 503 reason.
To stop waiting, cancel the job from the bulk toolbar.
Schedules tab
The Schedules tab configures the schedulable job types. Each card has an enable toggle, an interval picker, and a Run Now action.
| Job type | Default interval | Minimum interval |
|---|---|---|
| Data Exposure Risk Processing | 60 minutes | 5 minutes |
| Posture Analytics | 30 minutes | — |
Pending preview
Open the Data Exposure Risk Processing card to see a live preview of what a run would queue now. The preview refreshes every few seconds.
The preview uses the same window the run uses: changes since the last successful run, capped at 24 hours. The header reads Scanning changes since the last-run time → now.
The preview has two sections:
- PowerScale — a funnel with Time period, Records found, After exclude list, Removed (net-deleted), Collapsed (redundant touches), Skipped (recently scored), and Would submit. It also lists the active CSI path filters per cluster and a per-device table (Files, Creates, Modifies, Renames, Excluded, Collapsed, Net-deleted).
- ObjectScale / COSI — a funnel with Buckets in scope, Records found (object PUT and POST 2xx events), Removed (net-deleted), Collapsed (redundant touches), Skipped (recently scored), and Would submit. A per-row table is keyed on Device, Namespace, and Bucket (Objects, Creates (PUTs), POSTs, Excluded, Collapsed, Net-deleted). Object discovery covers only buckets that have a read service account granted. When none is granted, the section reads Object discovery skipped — no bucket has a read service account granted.
The preview uses the same counts as a real run, so you can confirm that the exclude list, net-deleted drop, per-path collapse, and recent-score skip behave as expected before you trigger extraction, content-integrity analysis, and classification. Run Now queues both the PowerScale files and the ObjectScale objects shown.
If a file is created and modified within the same one-second timestamp, the preview and the completed-run summary can place that file in different columns of the Creates, Modifies, and Renames split. The totals (Collapsed, distinct files, and Would submit or Files submitted) are always identical.
Workflows
Investigate a failed ingest job
- Set the Status filter to Failed.
- Set the Job Type filter to PowerScale Audit Ingest or ECS SSH Ingest.
- Click View on the failed row.
- Read the error banner and the Ingest Summary (check CH Errors), then read the log of the failing step to find the cause.
Trigger a manual scan or ingest
- On the Job Monitor tab, click K8 Scan Now to run a Kubernetes scan, or Audit Ingest Now to run PowerScale and ObjectScale ingest.
- Watch the new job appear in the Active Jobs strip.
- Click View to follow its steps live.
Confirm a K8 scan found COSI buckets
- Set the Job Type filter to K8 Scan and open the most recent completed run.
- Check the COSI Buckets and Bucket→Pod Mappings counts.
- Open COSI Bucket Inventory and Resolved Bucket→Pod for the per-bucket and per-pair detail. If a section is empty, it explains why, for example that the COSI CRDs are not installed or that no pod references the bucket's Secret.
Tips
- If a job is stuck running, check System Monitor > Audit Ingest Lag for clues about what it is waiting on.
- Avg Duration covers all job types combined. For a baseline per job type, filter the table to that type and review recent completed rows.
- A job in Awaiting Report has not necessarily failed. The agent may still be running it. Wait for the timeout before concluding the job is lost.