File Activity
The File Activity page lets you browse and query the raw audit event log. It supports two source types, PowerScale (Dell OneFS) file operations and ObjectScale (Dell ECS) S3 object operations. You can filter by time window, path, user, device, content-risk verdict, or Kubernetes context.
Where: Data Auditing > File Activity
The page is one of three views in the Storage Audit area:
- File Activity — this page.
- Structured Data Audit (Trino) — audit of structured data access. See Structured Data Audit (Trino).
- User Actions — the inverse of the Pod Creator column on this page. It starts from a user and lists every Kubernetes API action they took. See User Actions.
Overview
Audit events are written to the audit database by the ingest pipeline as they arrive from remote devices. The page queries that database directly, so results reflect events that have been fully ingested. All filters are combined with AND, so each added filter narrows the result set. The page is read-only and does not modify any data.
Page layout
From top to bottom, the page contains:
- Header — the page title, a Live indicator, and a one-line summary of indexed events, date range, disk usage, and compression ratio. A link returns to K8 Data Security, a database health badge shows connection status, and a Refresh button appears after you run a query.
- Cluster scope chips — one chip per registered Kubernetes cluster, plus All Clusters. Selecting a K8s cluster chip sets the K8 Cluster filter and re-runs the query.
- Statistic tiles — Total Events, Writes & Creates, Deletes, Reads, DB Latency, and Last Query duration.
- Operation mix — one chip per operation type with its event count. Click a chip to pre-fill the Operation filter, and click it again to clear it.
- Source toggle — All Sources, PowerScale · SMB · NFS, or ObjectScale · S3. The selection changes which filters are visible and which result columns appear.
- Filters — a grid of dropdowns and inputs, with Run Query, Clear, and, once you have results, Export CSV.
The Results table appears below the filters after you run a query. Failed operations have a tinted row background and successful operations show a green dot. Pagination controls appear above and below the table.
Database status badge
A badge in the page header shows whether the audit database is reachable. It is refreshed every 30 seconds.
- Connected (green) — shows the round-trip latency in milliseconds. The same value feeds the DB Latency tile.
- Error (red) — hover to see the underlying error message. Check the ClickHouse health page for diagnostics.
Source types
| Source | Collected from | Each event records |
|---|---|---|
| PowerScale | SMB and NFS file operations on Dell PowerScale clusters | Operation type (create, delete, rename, close, read, write), file path, user identity (UID or SID), client IP address, protocol (SMB or NFS), cluster name, and the storage node that served the request |
| ObjectScale | S3 object operations on Dell ObjectScale (ECS) devices | HTTP method (PUT, GET, DELETE, HEAD, COPY), HTTP status code, bucket name, namespace, S3 user, and object key |
Filters
| Filter | Description |
|---|---|
| Source type | Switch between PowerScale, ObjectScale, or both. |
| Time range | A preset window that bounds every query, including the content-risk filters. The default is Last 24 hours. Other presets are 48 hours and 7, 30, 60, and 90 days. Custom range reveals From and To datetime inputs. |
| Path or object key prefix | Filters PowerScale events by file path prefix or ObjectScale events by object key prefix. |
| IP address | Filters PowerScale events to a client IP. |
| Cluster or device | Filters to a PowerScale cluster or ECS device. |
| Bucket | Filters ObjectScale events to a bucket. |
| Namespace | Filters ObjectScale events to an ECS namespace. |
| Operation or method | Filters to an operation such as delete, or an HTTP method such as PUT. |
| Result | Success or Failure. See Failure reasons. |
| K8 Cluster | Filters to events from one Kubernetes cluster. The list is built from the clusters present in the data. |
Keep a time bound on every query. It keeps queries fast as the audit log grows, so widen the window only as far back as you need.
Failure reasons
Choosing Failure shows a dependent Failure Reason dropdown. Leave it on Any reason to match every failure, or pick a specific cause.
Failure reasons are the exact status returned by the storage system:
- PowerScale NTSTATUS codes, for example
ACCESS_DENIED — failure(0xC0000022),OBJECT_NAME_COLLISION — failure(0xC0000035),DISK_FULL, andPRIVILEGE_NOT_HELD. - ObjectScale S3 HTTP codes, for example
HTTP 403 Forbidden.
The dropdown lists only reasons present in the data.
Only the event types that your storage audit configuration records appear here. For example, if read auditing is disabled on PowerScale, failed reads are not present.
Kubernetes context filters
Three dropdowns narrow PowerScale events to the Kubernetes context that generated them:
- Namespace (K8) — the Kubernetes namespace where the workload runs, for example
aisec-testordata-science. - PVC (K8) — the PersistentVolumeClaim whose mounted directory contains the file. Entries appear as
<namespace>/<pvc_name>with the event count over the last 24 hours. The filter uses a stable PVC key, so it keeps working if the PVC is renamed and replaced. - Pod (K8) — the pod that wrote or read the file. Entries appear as
<namespace>/<pod_name>with the event count. The filter uses the pod's Kubernetes-assigned UID, so a pod deleted and recreated with the same name does not collide with the original.
Each dropdown is a typeahead. Pick from the top 200 entries by event count, or type any value not in the list, such as a UID copied from kubectl describe. Selections combine with all other filters.
These dropdowns are hidden when the source type is ObjectScale, because Kubernetes context is stamped only on PowerScale rows.
When a PowerScale file event arrives, the ingest pipeline looks up which pod was mounting the PVC at the event time and stamps the namespace, PVC, and pod UID on the event. This lookup is refreshed after every K8 scan.
An empty dropdown means no events in the last 24 hours have been stamped with that context. This usually happens because no K8s mounts are bound to the PowerScale paths the events came from, or because the events were ingested before Kubernetes context stamping was available. You can still type a value to filter manually.
Include legacy rows
When you filter on a PVC together with a path, for example from the Data Security Posture topology right-click, the search matches the PVC's stable key. This is a fast, indexed lookup that returns exactly that PVC's events. Events ingested before PVC keys were recorded have no key, so the key match does not return them.
Select Include legacy rows, next to Scored only, to also return those older rows by matching the path as well.
- Leave it off for the fast indexed search.
- Turn it on only when investigating activity that predates PVC keys. The path match is slower and can also return same-path events from other PVCs.
- It has no effect unless both a PVC and a path are set.
Deep links
Other K8 Data Security pages can open File Activity with filters pre-filled. When any of these URL parameters is present, the search runs automatically on load.
| Parameter | Effect |
|---|---|
source_type=powerscale or source_type=objectscale | Pre-selects the source type. |
bucket=<name> | Pre-fills the Bucket filter (ObjectScale). |
ecs_namespace=<ns> | Pre-fills the Namespace filter (ObjectScale). |
path=<prefix> | Pre-fills the path or object key prefix. |
ps_cluster=<name> | Pre-fills the PowerScale cluster filter. |
pod=<name> and namespace=<ns> | Selects that pod in the K8 context dropdown. |
from=YYYY-MM-DDTHH:MM:SS and to=YYYY-MM-DDTHH:MM:SS | Pre-fills the time window. |
This supports incident review. For example, right-click a bucket on the Data Security Posture topology with the time slider set to a past day, and land directly on that day's audit events for that bucket.
Pod context and attribution
Pod context for ObjectScale events
When the source type is ObjectScale, four additional columns appear in the results: Namespace, Pod, SA (ServiceAccount), and Pod Creator.
The Namespace, Pod, and SA values come from matching each event's S3 user to the Kubernetes pods that reference the bucket's COSI credentials Secret. The pod bindings are refreshed every 60 seconds. Each event is attributed using the latest binding at or before the event time, so events are attributed to the pod that was mounting the bucket at that moment, even if the pod has since been deleted or rotated.
Hover the Pod cell to see the pod UID, which helps you trace a workload across pod re-creations.
A cell shows — in these cases:
- COSI is not deployed on the Kubernetes cluster, so no bucket access resources exist to map the S3 user to a pod.
- The event's S3 user is not a COSI-managed identity, for example a user created with the ECS command line.
- The event predates the first COSI snapshot for that bucket, so no binding existed yet at the event time.
Pod Creator
Pod Creator is the user or service account that created the resolved pod, taken from the Kubernetes API audit. It answers "which user's workload touched this bucket". The column appears for both PowerScale and ObjectScale rows.
Pods created by a controller such as a Deployment, StatefulSet, or Job are attributed to the person who authored the workload, not to the controller that named the pod. For example, a pod from alice's Deployment shows alice, not the ReplicaSet controller.
The column shows one of four states:
| State | Meaning |
|---|---|
| A person (teal) | The human who authored the workload. For a workload that predates ingestion, it is the person who last modified or deleted it, and the entry adds · modified or · deleted. A pod created directly with kubectl run, including by the cluster admin, shows that person. |
automation · <identity> (amber) | The workload was created by software, such as a GitOps tool (ArgoCD), a CI pipeline, another operator's ServiceAccount, or a node. This is a correct, final answer. |
not attributed · predates window (grey) | Kubernetes API audit ingestion is running, but the workload has no create, modify, or delete event inside the audit window. Attribution applies only to events from when ingestion began. |
not attributed · audit off (grey) | Kubernetes API audit ingestion is not configured or not receiving events. Turn it on in Settings > K8s Audit. See Kubernetes API Audit Ingestion. |
Workload Owner and Pod Creator
Two columns together answer "which person was behind this file or object operation":
- Workload Owner (green) — the RunAI user who submitted the job, typically the data scientist.
- Pod Creator (teal) — the Kubernetes user who created the pod. This is the right attribution when RunAI is not installed.
The two columns come from independent sources. Pod Creator requires Kubernetes API audit ingestion to be enabled and can show — for pods created before ingestion began. For setup and the concept behind the two identities, see Kubernetes API Audit Ingestion.
Both columns populate on every query, for PowerScale and ObjectScale activity alike. You do not need to arrive from the Data Security Posture topology. A column shows — only when the pod cannot be attributed.
For events that were not tagged when first ingested, the report fills in the pod, PVC, and namespace at search time. This applies, for example, to events recorded just after a restart, before a PVC or bucket was known, or, for ObjectScale, before the bucket's first snapshot. The fill-in uses the current Kubernetes binding as a best effort, so a workload that consistently uses one PVC or bucket is attributed correctly even for its earliest events. A point-in-time attribution made at ingest always takes precedence.
Content risk
Each audit row shows the content-risk verdict for the file the event touched, so you can tell whether a file was PII, gibberish, or encrypted without leaving File Activity. Two columns appear after the Path / Object column.
PII column
The file's PII classification verdict is pii, clean, or mixed, shown as a colored chip (red, green, or amber).
Below the verdict, each detected entity type appears as a chip in the form TYPE ×count, for example EMAIL_ADDRESS ×42 · US_SSN ×7 · CREDIT_CARD ×3. Chips are sorted by count, and the cell shows six with a +N more indicator. Hover the indicator to see the rest. Hover the cell for the highest match score and the time the file was classified.
LC column
The Language-Coherence (LC) verdict is free text rather than a fixed set. Examples include "likely clean", "mixed / uncertain", "likely encrypted or gibberish", and "likely semantic manipulation: internal contradiction".
The chip color depends on keywords in the verdict:
| Color | Verdicts |
|---|---|
| Red | Encrypted or gibberish |
| Amber | Semantic manipulation |
| Green | Clean or coherent |
| Grey | Anything else |
Underneath the chip, mlm <avg_mlm_score> · H <shannon_entropy> shows the average masked-language-model score and the Shannon entropy, both to two decimals. Hover for the semantic-drift score and the time the file was scored.
A file that has never been scanned shows — in both columns. This is common, because discovery of files outpaces content scanning, and it is not an error.
Latest or as of event time
The Latest / As of event time toggle in the Content risk filter section controls which scan the PII and LC columns show.
- Latest (default) — the file's current verdict from the most recent scan, regardless of when the event happened. It answers "is this file risky now?".
- As of event time — the verdict in effect when the event occurred, which is the nearest scan at or before the event time. It answers "what did we know when this happened?". Each PII and LC cell shows an as-of badge, and its tooltip shows the matching scan time. This mode is slower and is intended for incident review rather than everyday browsing.
The point-in-time verdict comes from the retained history of every LC and PII scan. Scan history is kept for the Analyzer results retention period (default 12 months, set in Settings > Advanced Settings > ClickHouse retention). For a file whose only recorded scan predates scan-history recording, the cell falls back to the file's current verdict.
File scan history
Each scanned row has a clock button in its own column, shown only when the file has scan data. Click it to open a flyout listing the file's entire scan history, newest first. Each entry shows the verdict, entity count, and MLM score for every LC and PII scan. The same timeline appears on the Pipeline Analysis flyout, so you can follow a file's verdict drift across modifications from either page. Press Esc or click outside the flyout to close it.
Content-risk filters
The Content risk filter section, below the main filter grid, narrows the results to rows whose file scan matches a verdict or score.
| Filter | Description |
|---|---|
| PII verdict | Keep only files classified pii, clean, or mixed. |
| PII entity type | Keep only files in which a given entity type was found, such as EMAIL_ADDRESS, US_SSN, or CREDIT_CARD. The list contains the common Presidio entity types. |
| LC verdict contains | Free-text match against the LC verdict, for example gibberish, semantic, or clean. |
| Entropy (H) min / max | Keep files whose Shannon entropy falls in the range. High entropy indicates encrypted or random content. |
| MLM score min / max | Keep files whose average masked-language-model score falls in the range. |
| Has any PII | Keep only files with at least one detected PII entity. |
| LC anomalous | Keep only files that the language-coherence analyzer flagged as anomalous, meaning encrypted, gibberish, or high entropy. |
| Scored only | Drop rows whose file was never scanned, so every visible row has a PII and LC verdict. |
Content-risk filters join the file-scan data, which is slower than a plain audit search. Always pair them with a date range, because a very broad window can be slow on large datasets. A banner reminds you whenever a content-risk filter is active.
Results table
Results are paginated, with 50, 100, 500, or 1000 rows per page. The Operation column uses color-coded badges so destructive operations stand out. Export CSV downloads the current result set as a comma-separated file.
Rename events
PowerScale rename events, for SMB and NFS and for files and folders, record both the original and the new path. In the Path column, a rename row appears as source → destination. The common parent directory is removed so only the changed part shows, for example finance-2025 → finance-2025-archived. Hover the cell to see the full source and destination paths.
The destination is recorded for every rename regardless of protocol. Rename events ingested before destination capture was available show only the source path. The record is for auditing and visibility only.
Export CSV writes a rename as two columns, rename_from (the source path) and rename_to (the destination). Both are blank for other operations, so you can sort or filter a downloaded sheet on rename activity. The generic path column is still present for every row.
Summary rows
To keep the audit database compact under high read workloads, runs of high-volume events from the same client, user, and path or bucket are combined into one row per 60-second window. A summary row shows:
- An amber N ops badge in place of a single object key or operation. A tooltip shows how many individual operations the row combines.
- Bytes uploaded and downloaded, summed across all combined operations.
- The status set to the most frequent code in the window. For PowerScale, error codes are preferred over success, so failed operations stay visible in Failure filters and attack-detection alarms.
Page counts and totals use the number of combined operations rather than the row count. A query that matches a 71,000-event read-heavy file reports 71,000 events even though only about 130 rows back the result.
| Source | Summarized | Never summarized |
|---|---|---|
| ObjectScale | GET and HEAD object operations | PUT, DELETE, POST, and LIST operations (GET with list-type=…) |
| PowerScale | read, write, open, and close events, grouped by cluster, node, client IP, user, file path, protocol, and event type | State-changing events: create, delete, rename, and set-security |
State-changing events remain individual rows for compliance and forensic queries.
Workflows
Find all deletes in the last hour
- Set Source type to PowerScale.
- Set the time range to cover the last hour.
- Set Operation to delete.
- Click Run Query.
Download ObjectScale PUT events for a bucket
- Set Source type to ObjectScale.
- Set Bucket to the bucket name.
- Set Method to PUT.
- Set a time range.
- Click Run Query, then click Export CSV.
Tips
- Audit events are retained for the Storage audit events retention period (default 12 months, set in Settings > Advanced Settings > ClickHouse retention). Older events expire automatically and do not appear in results.
- Very broad queries, with no time range and no filters, can be slow. Always apply at least a time range on large datasets.
- Object keys for ObjectScale events are percent-encoded in the audit log. The results decode them for display.