Skip to main content
Migration Notice
We're migrating documentation from the old portal into this one. Some things may look a little different or out of place in the meantime — we know, and we're working to get it right. If something's unclear or doesn't look right, let us know.

AI Risk Pipeline

The AI Risk Pipeline page is a live view of the data-classification pipeline that turns raw files and objects into security verdicts. Every file discovered on a PowerScale (Dell OneFS) path or ObjectScale (Dell ECS) bucket flows through three analyzer stages, Extraction, Linguistic Coherence (LC) and Classify (PII), and lands in Analysis Results.

Where: AI Risk Pipeline

Use the page to check whether the pipeline is processing work, which stage is the bottleneck, whether a file was flagged for PII, and why an agent stopped receiving work.

The page is read-mostly. The only controls that change state are Drain (stop sending new work to one agent) and Clear backlog (admin).

Page layout​

From top to bottom, the page contains:

  1. Header — the page title, a status pill, a one-line summary, the Diagnostics and Clear backlog buttons, a next-refresh countdown, a Refresh button and the last-update time.
  2. Live Pipeline Telemetry deck — an animated flow graph of Extraction, LC, Classify and Analysis Results, with files traveling between them.
  3. Dispatch strip — one pill per stage showing work in progress (WIP), WIP cap, pending work and flow state.
  4. Flow rates strip — completed-file rates and latency percentiles for each hop.
  5. Last 24h source mix — the file and object split per stage. This line appears only when there is data.
  6. Stage cards — one card each for Extraction, LC and Classify, with a container roster, CPU and memory charts, a throughput chart and drain controls.

The deck and stage cards refresh every 2 seconds, the dispatch strip every 3 seconds, and the per-container throughput charts every 30 seconds.

Pipeline status​

The status pill next to the title summarizes the whole pipeline.

StatusMeaning
IDLEAgents are registered, but nothing is in flight and there was no throughput in the last 60 seconds.
ACTIVEWork is flowing (WIP in flight or throughput above zero) and all containers are healthy.
DEGRADEDAt least one registered container is stale and not sending heartbeats. Check the stage cards for the dimmed container chip.

The summary line reads N/N containers healthy · N in flight · N/min throughput. When any stage reported failures in the last minute, it also shows N failed/min.

Flow deck​

Files are drawn as small document glyphs that travel along the edges between stages. The animation reflects live activity.

  • Extraction > LC and Extraction > Classify — the number of glyphs tracks the WIP currently held at the destination stage. A rate-based stream is added so a fast stage that empties its WIP quickly still shows movement. The edge label reads WIP n · n/m, the in-flight count and files per minute.
  • LC > Analysis Results and Classify > Analysis Results — glyphs are emitted at the stage's measured throughput. The edge label shows that stage's files per minute.
  • Extraction > Analysis Results — a dashed amber arc below the Classify node. It represents duplicate content that skipped the scan (see Duplicate-content skips and inherited verdicts).

The duplicate-content arc behaves as follows:

  • Its label reads skip n/m, the skipped files per minute over the last 60 seconds.
  • The number of glyphs in flight scales with the skip rate.
  • The arc is dim and quiet when nothing is being skipped.
  • It animates only when Duplicate Content Detection is enabled in Settings > Pipeline Discovery.

Each stage node shows the stage icon, the live WIP count, the stage name, and healthy/total · n/min. A pulsing ring appears around a node only while it holds WIP.

The Analysis Results node shows the lifetime TOTAL of LC and Classify outputs and pulses only while data is arriving.

Hover over a node to see agents online, in-flight WIP and throughput. Click a node, or its magnifier icon, to open its inspection flyout. Right-clicking also opens it.

Dispatch strip​

Each stage pill shows the dispatcher's view of that stage in the form currentWip / targetWipCap · N pending, a WIP utilization bar and a state badge.

StateMeaning
flowingThe dispatcher is assigning work normally.
throttledThe dispatcher is holding work back, typically because the stage is at its WIP cap. Hover over the badge for the reason.
no agentsNo healthy agents are online for the stage, so there is nowhere to send work. This is the most common reason a stage looks stuck. Start the container fleet for that stage.

A pill can also show these indicators:

  • N/N sat — the number of the stage's healthy agents that are at their per-agent WIP cap.
  • N/min push-back — the number of times agents in the stage refused new work because they were full (HTTP 503) in the last 60 seconds.
  • N skipped · n/m — shown on the Extraction pill only. It shows the cumulative count of files and objects whose verdict was inherited from a byte-identical earlier scan, plus the current skips per minute. Hover over it for the file and object breakdown.

The skipped indicator appears only when Duplicate Content Detection is enabled and at least one skip has occurred. A duplicate-content skip is not a failure and is not a work-item Skipped status.

Flow rates​

The flow rates strip lists each hop on which a file completed in the last 60 seconds, in the form Source → Target n/min p50 Nms · p95 Nms.

  • n/min — completed files per minute.
  • p50 and p95 — the processing-latency percentiles for the hop.

If nothing completed in the last minute, the strip reads "No completed files in the last 60 s".

Last 24h source mix​

When the pipeline has processed work in the last 24 hours, a line shows the split between Files (PowerScale) and Objects (ObjectScale) for each stage (ext · lc · cls). The per-agent WIP elsewhere on the page does not separate file work from object work, so this line is the only place that does.

Stage cards​

The Extraction, LC and Classify cards give per-stage detail.

  • Header — the stage name, container count, and avg cpu % · avg mem % · n/min.
  • Container roster — one chip per agent container, showing the short name, CPU % and current WIP. A stale container (one that has stopped sending heartbeats) is dimmed so you can spot a down replica. It is removed automatically after about ten minutes without a report, or immediately when you remove or gracefully shut down that agent. A draining container is tinted orange and tagged drn.
  • Stat tiles — aggregate CPU, MEM, WIP and 60s (files per minute) across the stage's containers.
  • CPU + MEM chart — one solid line per container for CPU and one dashed line per container for memory, over recent history.
  • THROUGHPUT chart — files processed per 60-second window over the last 30 minutes, with one bar series per container.

Drain an agent​

Each container chip has a pause/play button.

  • Pause stops the dispatcher from assigning new work to that agent. Work already in flight completes normally. The chip turns orange and shows drn.
  • Play resumes assigning work to a drained agent.

Use drain to take a container out of rotation for a restart or investigation without dropping its in-flight files.

Inspection flyouts​

Clicking a node on the deck opens a flyout on the right edge of the page. Flyouts refresh every 5 seconds (shown by a countdown ring), close on Esc or when you click the backdrop, and show 50 rows per page.

Every row shows the complete source path of its file or object (sftp://…, s3://bucket/key or /ifs/…). Long paths wrap onto multiple lines instead of being truncated. A copy icon at the start of each path copies the full path to your clipboard, and a green check briefly confirms the copy.

Duplicate-content skips and inherited verdicts​

The pipeline recognizes when a file's bytes are identical to content it has already scanned, for example the same dataset copied into many PVCs, or an object re-uploaded under a new key. When Duplicate Content Detection is enabled in Settings > Pipeline Discovery, such a file does not re-run Extraction, LC and Classify. The pipeline reuses the earlier verdict and skips the scan.

An inherited verdict is genuine. It was derived from a byte-identical copy rather than recomputed for this copy. To see the originating scan, read the Inherited from source on the row.

Duplicate-content skips appear in the Analysis Results explorer:

  • Skip-stats line — the explorer shows Duplicate content skipped: N (last Nh), the cumulative number of files satisfied by an inherited verdict. The count covers a fixed recent look-back window, shown as (last Nh). It does not depend on the explorer's Window filter. The same line appears in the Data Exposure Risk preview flyout.
  • Inherited badge — an inherited row carries an Inherited badge and an Inherited from line naming the source (PVC, bucket or key, and path) that was actually scanned. The path is shown in full and has its own copy button.
  • Inherited only filter — inherited rows are interleaved by time with natively scanned rows by default. Turn on Inherited only to view just the inherited rows.

Both file and object skips appear. An inherited ObjectScale object, matched by its S3 ETag, appears exactly like an inherited PowerScale file, with its Object source badge.

Files with no extractable text​

A binary file, an image or an empty file has no language or PII to score. The pipeline exits early with a terminal no-usable-text, clean verdict and does not run full LC and Classify scoring.

These files produce a result row and count as Completed in the Extraction work-in-progress view. A later byte-identical copy can then inherit the clean verdict. Skipped is reserved for work the pipeline did not run, such as duplicate-content inheritance.

Extraction work-in-progress flyout​

Open this flyout from the Extraction node. It has three tabs: Files, Routing and Extraction. The Extraction tab shows per-agent delivery and disk state.

Files tab​

The Files tab lists extraction work items in a table with the columns Status, Type (File or Object, plus device), Source, Age, ms, Agent and MB. A route-track line under each row shows how the file fanned out to LC and Classify, with the target agent and an OK or error result for each hop.

  • Window — 1 h, 4 h, 24 h or 48 h. It bounds the list and the route-outcome counters on the Routing tab.
  • Filters — Status, Device, and Source (All, Files or Objects).
  • Sort — by enqueued time, status, source, agent or duration.

The status pills at the top reflect the entire extraction backlog, not only the rows on screen. All is the full count, and Pending increases as work queues up. Click a status pill to filter the list, and click it again to reset to All.

Two pills are derived metrics rather than status filters:

  • Submitted — the downstream LC and Classify handoffs across the loaded page.
  • Skipped (dedup) — the count of duplicate-content skips, with a file and object split. It appears only when duplicate-content detection has skipped something.

Routing tab​

The Routing tab aggregates route attempts per stage and target agent over the selected Window. For LC and Classify it draws one bar per target agent, split into successful and failed attempts, with the success rate. Use it to confirm that load is spread across multiple analyzer agents.

Above the bars, each stage shows a tally of routing outcomes over the selected Window. The counts are retained for 48 hours, so they honor the Window filter and survive a console restart.

CounterMeaning
Routed (200)A free slot was available, so the file was routed to the analyzer and delivered. This is the normal outcome.
No worker slots availableEvery eligible analyzer was at its worker-slot cap, so the sender waits for a slot and retries. This is not a failure. A steady stream under load means the analyzers are the bottleneck.
PUT 503A slot was handed out, but the analyzer filled up before the file was delivered, so the sender waits and retries. It is counted once per file, not per retry. This should stay near zero. A high count relative to No worker slots available suggests the console's slot view is stale, so check the heartbeat cadence.
Failed (410)No usable analyzer exists for the stage: none is healthy, all are suppressed for this file, or none has an ingress address. This is a real failure and the stage is failed for that file. A rising count means lost analyzer capacity, so check the stage's dispatch pill and container roster.
QueuedThe number of files pending scoring at the stage that were enqueued within the selected Window. A healthy idle pipeline shows a count near zero.
WaitingThe number of seconds the oldest waiting item has been parked. This is a live gauge and is not limited by the Window.

Read the counters together. Mostly 200 with some No worker slots available and near-zero PUT 503 under a burst indicates a healthy, slot-limited pipeline. A sustained Failed (410) points to a stage with no live analyzer.

Self-healing re-dispatch​

Occasionally the hand-off from an extraction agent to an analyzer is dropped for one leg. The LC or Classify work item exists but is never routed, so the file would stay pending and unscored.

The console detects this stalled hand-off when an LC or Classify item has no route attempt for about 2 minutes while that stage has an idle analyzer. It automatically re-queues the file's extraction so the hand-off is retried, instead of waiting for the one-hour delivery deadline that only marks the file failed.

  • A re-driven file shows a console-redispatch hop in its route-track line.
  • Each stalled leg is re-driven at most once. If it stalls again, normal delivery-deadline handling applies.
  • Re-dispatch applies only while LC is enabled.

Extraction tab​

Each extraction agent reports its delivery queue and staging disk state. Agents are listed by readable name, for example pipeline-extraction-<name>.

  • Delivery queue — files that have been extracted and staged to disk and are waiting for a free LC or Classify slot. Each entry shows how long it has been waiting. A short, fast-draining queue is normal. A queue that keeps growing means the analyzers are saturated, so cross-check the 503 counts on the Routing tab. Queued files are not lost, and they stay parked until a slot frees or the maximum wait deadline is reached.
  • Staging disk — the on-disk area where extracted text is held until it is delivered. It shows the staged text file count, the total staged bytes, the container's free disk space, and the live worker count (how many of the configured extraction workers are running).

The staged file count and bytes cover exactly the delivery files listed by the view N files on disk link. Temporary scratch space used during text extraction is not counted as a staged file, but it reduces disk free.

Click view N files on disk to list the staged files with name, size and time staged, oldest first. A staged file with a large age usually means a delivery stuck waiting for a slot, or one that reached its wait deadline.

Confirm that staged count and bytes are low and stable and that disk headroom is healthy. Rising staged bytes together with a growing delivery queue is the disk-side view of analyzer back-pressure.

LC verdict explorer​

Open this flyout from the LC node. It shows the Linguistic Coherence verdicts the LC analyzer computed in the selected window, with the source path and size of each file.

  • Summary pills — All, Coherent, Anomalous, Low coherence and Structured, each with a count and MB total.
  • Filters — Window, Agent, and a Path substring.
  • Columns — time, source, verdict, MLM (the average masked-language-model score), language, MB, processing ms and agent.
  • Footer — a per-agent breakdown, shown when more than one LC agent has rows.

Inherited duplicate-content verdicts are not listed here, because the LC container never ran on those files. They appear in the Analysis Results explorer, where you can use the Inherited only filter.

Verdict badges​

The verdict badge groups the analyzer's free-form verdict text. Hover over a badge to see the full verdict.

BadgeVerdict text
Anomalous (amber)Mentions manipulation, garbage or gibberish.
Low coherence (rose)Uncertain, insufficient, or low natural language.
Coherent (green)Clean, or no text.

Structured data​

Tables, CSV and Parquet exports, key-value record dumps and document extracts contain little prose, so the language model scores them low. They are not a coherence threat. The analyzer flags them as structured data, and the explorer marks them with a green STRUCT badge. The Structured pill filters to these files.

For a structured file that contains sensitive data, the data-classification (PII) score drives its risk, not the coherence score. Genuinely random or encrypted content is not flagged as structured and still scores Anomalous.

Early-exit potential​

A band below the title reports what a predictive early exit of the LC analyzer would do on your data. Early exit is off and record-only, so these figures are decision support and do not change behavior.

  • Reach capture % — the share of scored files long enough for the predictor to engage.
  • Forecast fires % — how often the predictor produces a confident forecast.
  • Match rate % — how often the early forecast agrees with the file's full-scan verdict.
  • Scoring saved % — the share of LC scoring work an early exit would skip. A low value means early exit would save little on this data. It pays off mainly with many long, clean documents.

These rates cover single-chunk files only. A very large document (over about 256 KiB) is scored in multiple chunks, and its forecast reflects only its opening chunk, so it is excluded. When large documents are excluded, the band shows how many, for example "3 large multi-chunk files excluded".

PII verdict explorer​

Open this flyout from the Classify node. It shows the PII verdicts the Classify analyzer computed in the selected window.

  • Summary pills — All, PII and Clean, each with a count and MB total.
  • Filters — Window, Agent, and a Path substring.
  • Columns — time, source, verdict, Entities, Top, MB and agent.
  • Footer — a per-agent breakdown, shown when more than one Classify agent has rows.

Entities lists the top entity types with counts, for example EMAIL_ADDRESS ×3. Top is the highest detection score. A PII verdict shows pii ×N, where N is the total entity count. A clean verdict shows a green shield.

Inherited duplicate-content verdicts are not listed here, for the same reason as in the LC explorer. They appear in the Analysis Results explorer.

Analysis Results explorer​

Open this flyout from the Analysis Results node. It is the consolidated view of every file's LC and PII verdicts, together with its Kubernetes context.

  • Filters — Window, Device, Lifecycle (All, Running or Terminated pod), Owner, Path substring, PVC (type-ahead from PVCs that have results in the window), Pod (type-ahead by pod UID) and Inherited only.
  • Sorting — click any column header to sort.
  • Total bytes — the filter row shows the total bytes (Σ) across loaded rows.
ColumnContents
PathThe full source path, with a per-row copy button. An inherited row also shows its Inherited from path with its own copy button.
Device / PVC / BucketThe storage device on the first line. The second line is the PVC for PowerScale files, or the object bucket for ObjectScale objects.
Pod · OwnerThe pod and its owner, with a status dot for running or terminated.
PIIThe verdict and top entity types.
LCThe coherence score.
MB and timeFile size and completion time.

Rows that have scan history show an expand chevron. Click it to see every LC and Classify verdict recorded for that physical file across modifications.

The explorer also shows the duplicate-content skip line and the Inherited badge described in Duplicate-content skips and inherited verdicts.

Diagnostics​

The Diagnostics button in the header opens a cross-agent log search. It covers log lines from every pipeline agent (Extraction, LC and Classify), which are retained for 7 days.

Filter the logs by:

  • Window
  • Stage
  • Agent
  • Level — All, Debug, Info, Warn or Error.
  • Component — type-ahead from components currently shipping logs, or type a substring.
  • Search — free text within the message.

Use these controls:

  • Auto-tail re-polls every 5 seconds. Turn it off (paused) to freeze the view while you read a busy window.
  • Each line shows local time, level, stage, agent, component and message. Click a line to expand it and see error_kind, the full message and the formatted raw_json.
  • When Component is set to lc-internal-chunking, a row of quick filters appears for per-chunk LC debugging: Errors only, a One file input (paste a file_id), and a One thread dropdown populated from the thread IDs in the visible rows.

If no lines match and the agent fleet was just restarted, allow 5 to 10 seconds for the first batch of logs to arrive.

Clear backlog​

The Clear backlog button is available to administrators. It deletes every pending and assigned work item from the pipeline queue.

  • Completed and failed rows are kept as an audit trail.
  • Work already claimed by an agent is marked orphaned and cleaned up automatically.
  • A confirmation dialog explains the effect. The action cannot be undone.
  • The result, for example Cleared N of M backlog rows, shows briefly in the header.

Use it for a clean slate after a configuration change has left the queue with stale or misrouted work items.

Common workflows​

Confirm the pipeline is processing work​

  1. Check that the status pill reads ACTIVE and the summary line shows non-zero throughput.
  2. Watch the deck. Glyphs should travel along the edges and the Analysis Results total should increase.
  3. If a stage looks stuck, check its dispatch pill. no agents means start that stage's fleet. throttled means the stage is at its WIP cap.

Find out why a stage is not processing​

  1. Look at the stage's dispatch pill. If it reads no agents, no healthy container is online for that stage, so start the fleet.
  2. If it reads throttled or shows push-back, the agents are full. Check the stage card's WIP and CPU and memory.
  3. Open Diagnostics, set Stage to that stage and Level to Error, and read the recent log lines.

Check whether a file was flagged for PII​

  1. Click the Analysis Results node to open the explorer.
  2. Set the Path filter, or pick the PVC or Pod, and widen the Window if needed.
  3. Read the PII column. A PII verdict with entity pills means sensitive data was detected. Clean means none.
  4. Expand the row to see the file's verdict history across scans.