Pipeline Discovery Settings
The Pipeline Discovery tab controls how the AI Risk Pipeline discovers files for processing. The discovery scheduler reads the file and object audit events and adds new create, write, put, and rename events to the pipeline work queue. The settings on this tab keep the queue focused on files worth running through extraction, linguistic coherence (LC), and classification.
Discovery Scan Window
Discovery is incremental. Each run scans only the file and object audit events that changed since the last successful run. The scheduler tracks this point as a watermark and advances it after every pass in which both files and objects succeed.
- Ingest-safety overlap — the watermark advances to the end of the run minus a 60-minute safety margin. Audit records arrive a little after a file or object is written, so an item written just before a run may not be in the audit data yet. The overlap lets the next run pick up these late records. Re-scanning the overlap does not process anything twice, because a repeated audit event maps to the same work item.
- Steady state — a typical run scans roughly one schedule interval of changes plus the overlap.
- 24-hour recovery cap — the scan window never starts more than 24 hours back. After console downtime, a transient ClickHouse failure, or the first run after a fresh deployment, a run scans at most the most recent day.
- Failure handling — the watermark advances only when a pass fully succeeds. After a failed run, the next run scans the same window again.
- Run Now — a manual Run Now advances the watermark in the same way as a scheduled run.
The preview on the Schedules tab and the Discovery Summary of each completed job show this window as Scanning changes since the last run → now.
The 60-minute margin covers the audit delay seen in practice, which is about 17 minutes. If audit records in your environment can take longer than an hour to become available, raise the margin to match.
Exclude List
Files and objects that match any entry on this list are skipped at discovery and never enter the pipeline. The pipeline can read about 1,500 file formats, so the list only needs to name what you never want scanned. Typical entries are binaries, images, audio and video, archives, encrypted files, noisy paths, and scratch file names.
Entry Types
The match type is inferred from what you type. You can also force a type with a prefix.
| You type | Match type | What it matches |
|---|---|---|
tmp, iso (a bare word) | Extension | Files whose extension is exactly that value, for example report.tmp. |
temp*, *.bak, name:*.part (contains * or ?) | Name glob | The file name or object key leaf. * matches any run of characters and ? matches one character. |
pvc/4a5d2c2ee7, /_staging/, path:/snapshots/ (contains /) | Path contains | A substring anywhere in the full file path or object key. |
- Prefixes —
ext:,name:, andpath:force a match type regardless of the text. For example,path:logsmatches the substringlogseven though it has no slash. - Bare words — a plain word such as
tempis treated as an extension. To match files namedtemp…, usename:temp*. - Case — matching ignores case, so
TMPandtmpare the same rule. - Patterns — regular expressions are not supported. Only literal substrings and the
*and?wildcards are supported.
Entries are matched against the readable file path or object key.
Manage the List
- Add an entry — type it and press Enter or click Add. Each entry is color-coded and labeled with its type: extension, name glob, or path contains.
- Remove an entry — click the trash icon on the entry.
- Reset to defaults — restores the default extension list, which includes
bin,exe,png,mp4, andzip. - Empty list — nothing is excluded. Every discovered file and object is processed, and the extraction agent still skips non-text content when it detects it.
Changes apply at the next discovery run, and no restart is needed. Each entry can be up to 256 characters, and the list can hold up to 200 entries. Entries with control characters or quotes are rejected.
The number of files and objects dropped by this list in each run appears per device in the Excluded column of the Discovery Summary.
Skip Recently Scored Files
When a file was classified within the last N hours, discovery does not add it to the queue again. This avoids repeated work when one file produces several audit events in a short time, such as open, read, and rename, without its content changing.
| Setting | Value |
|---|---|
| Range | 0–168 hours (1 week) |
| Default | 1 hour |
Set the value to 0 to turn this off.
Delivery Slot Wait
After a file is extracted, its text is delivered to the LC and data-classification analyzers for scoring. When the analyzers are at capacity, delivery waits for a free slot and retries instead of failing. Two settings tune this wait.
| Setting | Range | Default | Description |
|---|---|---|---|
| Retry backoff (seconds) | 1–300 | 10 | How often a waiting delivery asks again for a slot when all analyzer slots are busy. |
| Max wait deadline (seconds) | 30–7200 (2 hours) | 1800 (30 minutes) | The total time a file's text waits for a slot before the stage fails for that file. |
- Retry backoff — a lower value picks up a freed slot sooner. Do not set it below the 5-second agent heartbeat interval. Slot availability refreshes only at each heartbeat, so a shorter backoff gives no benefit.
- Max wait deadline — a higher value is more tolerant under sustained load and fails fewer files while a backlog clears, but it holds extracted text in progress for longer. A lower value fails files sooner when the analyzers are saturated.
These settings apply to the delivery from extraction to the LC and classification analyzers, not to discovery. Changes reach the extraction agents at the next heartbeat, with no restart.
Extraction Tuning
Two settings tune the extraction stage. Both reach the extraction agents at the next heartbeat and apply without a restart.
Extraction Worker Count
The number of concurrent extraction workers that each extraction container runs.
| Setting | Value |
|---|---|
| Range | 1–32 |
| Default | 4 |
- Raise the count when extraction is the bottleneck. For example, delivery threads are held by large-file LC or classification back-pressure and starve the parser, or the Extraction stage card shows spare CPU and memory while its work in progress stays high.
- Memory limit — each worker holds a file parse in memory. Keep the extraction heap size multiplied by the worker count at or below the container's memory limit. The container sets its heap size from its own memory limit at startup.
- Lower the count if extraction CPU stays at 100% or memory pressure appears. A count higher than the host can support slows extraction or can cause the container to be stopped for running out of memory.
Max File Size (MB)
Files whose raw size is larger than this value are skipped before extraction and recorded as oversize. This stops one very large file from exhausting the extractor's memory and stopping the container.
| Setting | Value |
|---|---|
| Range | 0–10000 MB |
| Default | 0 |
- 0 — no size limit. Every discovered file is passed to the parser.
- A limit — set one, for example a few hundred MB, when your data includes occasional very large files. A lower limit skips more files, and a higher limit attempts larger files.
Linguistic Coherence
The Linguistic Coherence box groups the false-positive reduction and LC analyzer concurrency panels. A master Enabled / Disabled switch at the top controls the group. LC is on by default. Turning it off keeps the settings but greys out both panels, and turning it on makes them editable again.
- Enabled — files follow the route extraction, LC, classification. LC pipeline agents must be deployed.
- Disabled — the LC stage is removed. Files go directly from extraction to classification and are not sent to an LC analyzer. Use this for a deployment with classification agents only and no LC agent.
False-positive Reduction
These controls set how strongly the content-integrity (Entropy) risk signal ignores benign noise. They can be edited only when the LC feature is enabled. All of them are on by default.
Regardless of these settings, manipulation verdicts are never suppressed, and the sharp content-integrity-drop alarm always fires. Changes apply at the next risk recalculation or discovery run, with no restart.
- Baseline-deviation suppression — when on, a garbage or gibberish verdict is not counted if the file's entropy is normal for that file's own history, such as a recurring data dump that has always had high entropy.
- Min baseline samples — the history a file needs before its baseline is trusted. Default is 5.
- Deviation band K (σ) — a file is normal for itself when its entropy is within K standard deviations of its own mean. Default is 1.0. Only a clear deviation, up or down, counts.
- Stability quorum — the number of garbage scans a file needs, with the file still anomalous on its latest scan, before the verdict counts. This ignores a single short-lived blip. Default is 2. Set it to 1 to count on the first garbage scan. It applies to garbage verdicts only, and manipulation verdicts always count.
- Confidently-clean re-scan skip window — the number of hours within which a file is not analyzed again if its latest verdict is one of the verdicts selected in Skip verdicts. A harmless rewrite of a clean file then does not restart the pipeline. Set it to 0 to turn this off. An anomalous file is always scanned again.
- Operator allowlist — known harmless noisy sources whose anomalies never count toward the score. Enter one entry per line: a PVC or file natural key (SHA-256), or a path glob that contains
*. Globs are applied at discovery. - LC drift early-warning — raises the low-severity slow-drift signal on Posture Analytics. You can set the entropy and semantic-drift floors and a minimum number of samples before the signal is raised.
LC Analyzer Concurrency
These settings control how the LC analyzer splits large extracted-text files into chunks, scores the chunks in parallel, and merges the results into one file verdict. They can be edited only when the LC feature is enabled.
| Setting | Default | Description |
|---|---|---|
| Chunk Bytes | 256 KB | Target size of each chunk, in UTF-8 bytes. Smaller chunks allow more parallelism on multi-MB files. Larger chunks reduce the cost of merging. |
| Chunk Overlap Bytes | 2 KB | Text taken from the end of the previous chunk and added to every chunk after the first, so scoring has context across a cut. |
| Internal Parallelism | 4 | Worker threads inside the analyzer. This matches one thread per core on a 4-vCPU LC container. Each thread uses about 0.25 GB of working memory in addition to the shared 1.5 GB model. |
| Max Failed Chunk Pct | 0.25 | If more than this fraction of chunks fail, the file verdict is error instead of a misleading partial result. |
| GPU Enabled | On | Runs analysis on the GPU when CUDA is available on the LC host, which is much faster. It has no effect on CPU-only hosts. |
The number of files an LC agent accepts at once is set per agent in Fleet Management > Max Jobs, not here. Saved values reach the LC agents at each heartbeat and apply from the next file, with no restart.
Duplicate Content Detection
Before a file or object enters extraction, LC, and classification, the pipeline compares a hash of its contents with files it has already processed. On a match, the console reuses the earlier verdict, including PII, LC, and early-exit outcomes, and skips the item. Two files with identical content are therefore scanned only once.
Each skip is recorded in the pipeline logs as a duplicate-content skip, so the avoided work can be audited and appears in the pipeline skip statistics.
Difference from Skipping Recently Scored Files
| Feature | Matches | Purpose |
|---|---|---|
| Skip recently scored files | The same path within a time window | Avoids re-scanning a file that keeps producing audit events when its content has not changed. |
| Duplicate content detection | Identical content at different paths | Scans content once when it is copied or renamed. Examples are a dataset copied into ten PVCs, a template saved under many names, or an object uploaded again under a new key. The other copies reuse the verdict. |
Controls
- Enable — turns duplicate detection on or off. It is off by default. While it is off, every discovered file runs through the full pipeline.
- Hash algorithm — the hash function used to identify content. Default is SHA-256.
- Batch size — the number of candidate hashes an agent groups into one duplicate check. Range is 1–5000. Default is 500.
Advanced Controls
- Confirm read at discovery — reads the object again at discovery to confirm the match before reusing a verdict. This applies to object stores only and has no effect on files. Default is off.
- Inherit early-exit outcomes — when on, early-exit verdicts, such as a binary file or a file with no usable text, are also reused on a content match, in addition to full PII and LC verdicts. Default is on.
- Inherited-verdict TTL (hours) — the maximum age of a reused verdict. When a verdict is older, the match is ignored and the file is scanned again. Range is 1–87600 hours (10 years). Default is 4320 hours (180 days).
Changes to the enable setting, hash algorithm, and batch size reach the pipeline agents at the next heartbeat.