Skip to main content
Migration Notice
We're migrating documentation from the old portal into this one. Some things may look a little different or out of place in the meantime — we know, and we're working to get it right. If something's unclear or doesn't look right, let us know.

File Activity

The File Activity page lets you browse and query the raw audit event log. It supports two source types, PowerScale (Dell OneFS) file operations and ObjectScale (Dell ECS) S3 object operations. You can filter by time window, path, user, device, content-risk verdict, or Kubernetes context.

Where: Data Auditing > File Activity

The page is one of three views in the Storage Audit area:

  • File Activity — this page.
  • Structured Data Audit (Trino) — audit of structured data access. See Structured Data Audit (Trino).
  • User Actions — the inverse of the Pod Creator column on this page. It starts from a user and lists every Kubernetes API action they took. See User Actions.

Overview​

Audit events are written to the audit database by the ingest pipeline as they arrive from remote devices. The page queries that database directly, so results reflect events that have been fully ingested. All filters are combined with AND, so each added filter narrows the result set. The page is read-only and does not modify any data.

Page layout​

From top to bottom, the page contains:

  1. Header — the page title, a Live indicator, and a one-line summary of indexed events, date range, disk usage, and compression ratio. A link returns to K8 Data Security, a database health badge shows connection status, and a Refresh button appears after you run a query.
  2. Cluster scope chips — one chip per registered Kubernetes cluster, plus All Clusters. Selecting a K8s cluster chip sets the K8 Cluster filter and re-runs the query.
  3. Statistic tiles — Total Events, Writes & Creates, Deletes, Reads, DB Latency, and Last Query duration.
  4. Operation mix — one chip per operation type with its event count. Click a chip to pre-fill the Operation filter, and click it again to clear it.
  5. Source toggle — All Sources, PowerScale · SMB · NFS, or ObjectScale · S3. The selection changes which filters are visible and which result columns appear.
  6. Filters — a grid of dropdowns and inputs, with Run Query, Clear, and, once you have results, Export CSV.

The Results table appears below the filters after you run a query. Failed operations have a tinted row background and successful operations show a green dot. Pagination controls appear above and below the table.

Database status badge​

A badge in the page header shows whether the audit database is reachable. It is refreshed every 30 seconds.

  • Connected (green) — shows the round-trip latency in milliseconds. The same value feeds the DB Latency tile.
  • Error (red) — hover to see the underlying error message. Check the ClickHouse health page for diagnostics.

Source types​

SourceCollected fromEach event records
PowerScaleSMB and NFS file operations on Dell PowerScale clustersOperation type (create, delete, rename, close, read, write), file path, user identity (UID or SID), client IP address, protocol (SMB or NFS), cluster name, and the storage node that served the request
ObjectScaleS3 object operations on Dell ObjectScale (ECS) devicesHTTP method (PUT, GET, DELETE, HEAD, COPY), HTTP status code, bucket name, namespace, S3 user, and object key

Filters​

FilterDescription
Source typeSwitch between PowerScale, ObjectScale, or both.
Time rangeA preset window that bounds every query, including the content-risk filters. The default is Last 24 hours. Other presets are 48 hours and 7, 30, 60, and 90 days. Custom range reveals From and To datetime inputs.
Path or object key prefixFilters PowerScale events by file path prefix or ObjectScale events by object key prefix.
IP addressFilters PowerScale events to a client IP.
Cluster or deviceFilters to a PowerScale cluster or ECS device.
BucketFilters ObjectScale events to a bucket.
NamespaceFilters ObjectScale events to an ECS namespace.
Operation or methodFilters to an operation such as delete, or an HTTP method such as PUT.
ResultSuccess or Failure. See Failure reasons.
K8 ClusterFilters to events from one Kubernetes cluster. The list is built from the clusters present in the data.

Keep a time bound on every query. It keeps queries fast as the audit log grows, so widen the window only as far back as you need.

Failure reasons​

Choosing Failure shows a dependent Failure Reason dropdown. Leave it on Any reason to match every failure, or pick a specific cause.

Failure reasons are the exact status returned by the storage system:

  • PowerScale NTSTATUS codes, for example ACCESS_DENIED — failure(0xC0000022), OBJECT_NAME_COLLISION — failure(0xC0000035), DISK_FULL, and PRIVILEGE_NOT_HELD.
  • ObjectScale S3 HTTP codes, for example HTTP 403 Forbidden.

The dropdown lists only reasons present in the data.

note

Only the event types that your storage audit configuration records appear here. For example, if read auditing is disabled on PowerScale, failed reads are not present.

Kubernetes context filters​

Three dropdowns narrow PowerScale events to the Kubernetes context that generated them:

  • Namespace (K8) — the Kubernetes namespace where the workload runs, for example aisec-test or data-science.
  • PVC (K8) — the PersistentVolumeClaim whose mounted directory contains the file. Entries appear as <namespace>/<pvc_name> with the event count over the last 24 hours. The filter uses a stable PVC key, so it keeps working if the PVC is renamed and replaced.
  • Pod (K8) — the pod that wrote or read the file. Entries appear as <namespace>/<pod_name> with the event count. The filter uses the pod's Kubernetes-assigned UID, so a pod deleted and recreated with the same name does not collide with the original.

Each dropdown is a typeahead. Pick from the top 200 entries by event count, or type any value not in the list, such as a UID copied from kubectl describe. Selections combine with all other filters.

These dropdowns are hidden when the source type is ObjectScale, because Kubernetes context is stamped only on PowerScale rows.

When a PowerScale file event arrives, the ingest pipeline looks up which pod was mounting the PVC at the event time and stamps the namespace, PVC, and pod UID on the event. This lookup is refreshed after every K8 scan.

An empty dropdown means no events in the last 24 hours have been stamped with that context. This usually happens because no K8s mounts are bound to the PowerScale paths the events came from, or because the events were ingested before Kubernetes context stamping was available. You can still type a value to filter manually.

Include legacy rows​

When you filter on a PVC together with a path, for example from the Data Security Posture topology right-click, the search matches the PVC's stable key. This is a fast, indexed lookup that returns exactly that PVC's events. Events ingested before PVC keys were recorded have no key, so the key match does not return them.

Select Include legacy rows, next to Scored only, to also return those older rows by matching the path as well.

  • Leave it off for the fast indexed search.
  • Turn it on only when investigating activity that predates PVC keys. The path match is slower and can also return same-path events from other PVCs.
  • It has no effect unless both a PVC and a path are set.

Other K8 Data Security pages can open File Activity with filters pre-filled. When any of these URL parameters is present, the search runs automatically on load.

ParameterEffect
source_type=powerscale or source_type=objectscalePre-selects the source type.
bucket=<name>Pre-fills the Bucket filter (ObjectScale).
ecs_namespace=<ns>Pre-fills the Namespace filter (ObjectScale).
path=<prefix>Pre-fills the path or object key prefix.
ps_cluster=<name>Pre-fills the PowerScale cluster filter.
pod=<name> and namespace=<ns>Selects that pod in the K8 context dropdown.
from=YYYY-MM-DDTHH:MM:SS and to=YYYY-MM-DDTHH:MM:SSPre-fills the time window.

This supports incident review. For example, right-click a bucket on the Data Security Posture topology with the time slider set to a past day, and land directly on that day's audit events for that bucket.

Pod context and attribution​

Pod context for ObjectScale events​

When the source type is ObjectScale, four additional columns appear in the results: Namespace, Pod, SA (ServiceAccount), and Pod Creator.

The Namespace, Pod, and SA values come from matching each event's S3 user to the Kubernetes pods that reference the bucket's COSI credentials Secret. The pod bindings are refreshed every 60 seconds. Each event is attributed using the latest binding at or before the event time, so events are attributed to the pod that was mounting the bucket at that moment, even if the pod has since been deleted or rotated.

Hover the Pod cell to see the pod UID, which helps you trace a workload across pod re-creations.

A cell shows — in these cases:

  • COSI is not deployed on the Kubernetes cluster, so no bucket access resources exist to map the S3 user to a pod.
  • The event's S3 user is not a COSI-managed identity, for example a user created with the ECS command line.
  • The event predates the first COSI snapshot for that bucket, so no binding existed yet at the event time.

Pod Creator​

Pod Creator is the user or service account that created the resolved pod, taken from the Kubernetes API audit. It answers "which user's workload touched this bucket". The column appears for both PowerScale and ObjectScale rows.

Pods created by a controller such as a Deployment, StatefulSet, or Job are attributed to the person who authored the workload, not to the controller that named the pod. For example, a pod from alice's Deployment shows alice, not the ReplicaSet controller.

The column shows one of four states:

StateMeaning
A person (teal)The human who authored the workload. For a workload that predates ingestion, it is the person who last modified or deleted it, and the entry adds · modified or · deleted. A pod created directly with kubectl run, including by the cluster admin, shows that person.
automation · <identity> (amber)The workload was created by software, such as a GitOps tool (ArgoCD), a CI pipeline, another operator's ServiceAccount, or a node. This is a correct, final answer.
not attributed · predates window (grey)Kubernetes API audit ingestion is running, but the workload has no create, modify, or delete event inside the audit window. Attribution applies only to events from when ingestion began.
not attributed · audit off (grey)Kubernetes API audit ingestion is not configured or not receiving events. Turn it on in Settings > K8s Audit. See Kubernetes API Audit Ingestion.

Workload Owner and Pod Creator​

Two columns together answer "which person was behind this file or object operation":

  • Workload Owner (green) — the RunAI user who submitted the job, typically the data scientist.
  • Pod Creator (teal) — the Kubernetes user who created the pod. This is the right attribution when RunAI is not installed.

The two columns come from independent sources. Pod Creator requires Kubernetes API audit ingestion to be enabled and can show — for pods created before ingestion began. For setup and the concept behind the two identities, see Kubernetes API Audit Ingestion.

Both columns populate on every query, for PowerScale and ObjectScale activity alike. You do not need to arrive from the Data Security Posture topology. A column shows — only when the pod cannot be attributed.

For events that were not tagged when first ingested, the report fills in the pod, PVC, and namespace at search time. This applies, for example, to events recorded just after a restart, before a PVC or bucket was known, or, for ObjectScale, before the bucket's first snapshot. The fill-in uses the current Kubernetes binding as a best effort, so a workload that consistently uses one PVC or bucket is attributed correctly even for its earliest events. A point-in-time attribution made at ingest always takes precedence.

Content risk​

Each audit row shows the content-risk verdict for the file the event touched, so you can tell whether a file was PII, gibberish, or encrypted without leaving File Activity. Two columns appear after the Path / Object column.

PII column​

The file's PII classification verdict is pii, clean, or mixed, shown as a colored chip (red, green, or amber).

Below the verdict, each detected entity type appears as a chip in the form TYPE ×count, for example EMAIL_ADDRESS ×42 · US_SSN ×7 · CREDIT_CARD ×3. Chips are sorted by count, and the cell shows six with a +N more indicator. Hover the indicator to see the rest. Hover the cell for the highest match score and the time the file was classified.

LC column​

The Language-Coherence (LC) verdict is free text rather than a fixed set. Examples include "likely clean", "mixed / uncertain", "likely encrypted or gibberish", and "likely semantic manipulation: internal contradiction".

The chip color depends on keywords in the verdict:

ColorVerdicts
RedEncrypted or gibberish
AmberSemantic manipulation
GreenClean or coherent
GreyAnything else

Underneath the chip, mlm <avg_mlm_score> · H <shannon_entropy> shows the average masked-language-model score and the Shannon entropy, both to two decimals. Hover for the semantic-drift score and the time the file was scored.

A file that has never been scanned shows — in both columns. This is common, because discovery of files outpaces content scanning, and it is not an error.

Latest or as of event time​

The Latest / As of event time toggle in the Content risk filter section controls which scan the PII and LC columns show.

  • Latest (default) — the file's current verdict from the most recent scan, regardless of when the event happened. It answers "is this file risky now?".
  • As of event time — the verdict in effect when the event occurred, which is the nearest scan at or before the event time. It answers "what did we know when this happened?". Each PII and LC cell shows an as-of badge, and its tooltip shows the matching scan time. This mode is slower and is intended for incident review rather than everyday browsing.

The point-in-time verdict comes from the retained history of every LC and PII scan. Scan history is kept for the Analyzer results retention period (default 12 months, set in Settings > Advanced Settings > ClickHouse retention). For a file whose only recorded scan predates scan-history recording, the cell falls back to the file's current verdict.

File scan history​

Each scanned row has a clock button in its own column, shown only when the file has scan data. Click it to open a flyout listing the file's entire scan history, newest first. Each entry shows the verdict, entity count, and MLM score for every LC and PII scan. The same timeline appears on the Pipeline Analysis flyout, so you can follow a file's verdict drift across modifications from either page. Press Esc or click outside the flyout to close it.

Content-risk filters​

The Content risk filter section, below the main filter grid, narrows the results to rows whose file scan matches a verdict or score.

FilterDescription
PII verdictKeep only files classified pii, clean, or mixed.
PII entity typeKeep only files in which a given entity type was found, such as EMAIL_ADDRESS, US_SSN, or CREDIT_CARD. The list contains the common Presidio entity types.
LC verdict containsFree-text match against the LC verdict, for example gibberish, semantic, or clean.
Entropy (H) min / maxKeep files whose Shannon entropy falls in the range. High entropy indicates encrypted or random content.
MLM score min / maxKeep files whose average masked-language-model score falls in the range.
Has any PIIKeep only files with at least one detected PII entity.
LC anomalousKeep only files that the language-coherence analyzer flagged as anomalous, meaning encrypted, gibberish, or high entropy.
Scored onlyDrop rows whose file was never scanned, so every visible row has a PII and LC verdict.
warning

Content-risk filters join the file-scan data, which is slower than a plain audit search. Always pair them with a date range, because a very broad window can be slow on large datasets. A banner reminds you whenever a content-risk filter is active.

Results table​

Results are paginated, with 50, 100, 500, or 1000 rows per page. The Operation column uses color-coded badges so destructive operations stand out. Export CSV downloads the current result set as a comma-separated file.

Rename events​

PowerScale rename events, for SMB and NFS and for files and folders, record both the original and the new path. In the Path column, a rename row appears as source → destination. The common parent directory is removed so only the changed part shows, for example finance-2025 → finance-2025-archived. Hover the cell to see the full source and destination paths.

The destination is recorded for every rename regardless of protocol. Rename events ingested before destination capture was available show only the source path. The record is for auditing and visibility only.

Export CSV writes a rename as two columns, rename_from (the source path) and rename_to (the destination). Both are blank for other operations, so you can sort or filter a downloaded sheet on rename activity. The generic path column is still present for every row.

Summary rows​

To keep the audit database compact under high read workloads, runs of high-volume events from the same client, user, and path or bucket are combined into one row per 60-second window. A summary row shows:

  • An amber N ops badge in place of a single object key or operation. A tooltip shows how many individual operations the row combines.
  • Bytes uploaded and downloaded, summed across all combined operations.
  • The status set to the most frequent code in the window. For PowerScale, error codes are preferred over success, so failed operations stay visible in Failure filters and attack-detection alarms.

Page counts and totals use the number of combined operations rather than the row count. A query that matches a 71,000-event read-heavy file reports 71,000 events even though only about 130 rows back the result.

SourceSummarizedNever summarized
ObjectScaleGET and HEAD object operationsPUT, DELETE, POST, and LIST operations (GET with list-type=…)
PowerScaleread, write, open, and close events, grouped by cluster, node, client IP, user, file path, protocol, and event typeState-changing events: create, delete, rename, and set-security

State-changing events remain individual rows for compliance and forensic queries.

Workflows​

Find all deletes in the last hour​

  1. Set Source type to PowerScale.
  2. Set the time range to cover the last hour.
  3. Set Operation to delete.
  4. Click Run Query.

Download ObjectScale PUT events for a bucket​

  1. Set Source type to ObjectScale.
  2. Set Bucket to the bucket name.
  3. Set Method to PUT.
  4. Set a time range.
  5. Click Run Query, then click Export CSV.

Tips​

  • Audit events are retained for the Storage audit events retention period (default 12 months, set in Settings > Advanced Settings > ClickHouse retention). Older events expire automatically and do not appear in results.
  • Very broad queries, with no time range and no filters, can be slow. Always apply at least a time range on large datasets.
  • Object keys for ObjectScale events are percent-encoded in the audit log. The results decode them for display.