Skip to main content
Migration Notice
We're migrating documentation from the old portal into this one. Some things may look a little different or out of place in the meantime — we know, and we're working to get it right. If something's unclear or doesn't look right, let us know.

Structured Behavioral Baselines - Trino Anomaly Detection

Structured Anomaly Detection brings behavioral analytics to your Trino SQL engine. Where the Structured Data Audit (Trino) view lets you browse the query log, this feature learns each Trino user's normal pattern of structured-data access over a rolling window (default 30 days) and flags deviations.

It helps you answer questions such as:

  • Is this user reading tables, schemas, or catalogs they have never touched? (scope expansion or lateral movement)
  • Is the volume of their reads, rows, or bytes far above their own norm? (bulk export or exfiltration)
  • Has a read-mostly analyst started issuing INSERT, DELETE, or DROP? (role drift or destruction)
  • Is someone probing with a burst of failed or access-denied queries? (reconnaissance)
  • Is a normally quiet table suddenly read by many users, or at off-hours? (data-asset anomaly)

The feature tracks two kinds of subject: the user, and the data asset (catalog.schema.table and its parents). The user is the Trino user, or the authentication principal when no Trino user is set.

Where to see anomalies and baselines​

Detected anomalies appear in the Posture Analytics transitions feed. Filter the feed by the trino_* signals.

The learned baselines, which represent each user's normal behavior, are in the Structured Behavioral Baselines view. Open the Data Security Posture page (Posture Analytics) and select the Structured Behavioral Baselines toggle at the top, next to Transitions feed. The view is read-only, and opening it does not recalculate anything.

Structured Behavioral Baselines view​

Select a work zone at the top right: All, Working hours, After hours, or Weekend. Baselines are kept for each work zone because structured-data access varies strongly by time.

Users tab​

The Users tab has one row for each Trino user.

ColumnDescription
UserThe user's identity. This is the Trino user, or the authentication principal when no Trino user is set.
Maturitylearning until the user has enough baseline samples. The sample count (n) and active days are shown. The status changes to mature once the user passes the minimum. Until then, only the peer baseline can flag the user.
Op-mixA proportional bar of read, insert, update, delete, and DDL operations. Hover to see counts.
ScopeThe number of distinct catalogs, schemas, and tables the user touched, shown as c / s / t.
Queries, Rows read, Rows ins., Rows del.The user's baseline averages. Reads, inserts, and deletes are tracked separately, so reading, inserting, and deleting 10 million rows are never confused.
Last anomalyWhen the user most recently produced a flagged transition.

Click a row to open the user detail panel.

User detail panel​

  • Operation mix — the user's distribution across read, insert, update, delete, and DDL operations.
  • Per-zone baselines — for each work zone, a table of every metric with its mean, median, MAD, EWMA, and sample count (n). The magnitude metrics are rows read, rows inserted, rows deleted, max_stmt_rows, and ddl_impact. Click a metric to open its peer-distribution panel.
  • Access set — the catalogs, schemas, and tables the user normally touches, with access counts and last-seen times. The first access to an item that is not in this set is a novelty event. An access to an item seen far below the rarity floor is a rare access.
  • Recent transitions — the anomalies recently produced for this user, with signal, metric, z-score, severity, and age.

Tables tab​

The Tables tab shows, for each table and work zone, the mean, standard deviation, and MAD of the reader count, and the mean read volume, with the sample count. This data supports table-side detection. A normally quiet table that is suddenly read by many or new users, or at off-hours, is flagged independently of any single user.

Peer-distribution panel​

For a chosen metric and work zone, the panel shows the population median ± MAD band, with every user's value and the selected user highlighted. It uses the MAD rather than the mean and standard deviation, so a few heavy ETL or service accounts cannot distort the band. The panel explains why a trino_peer_outlier signal fired.

Configuration​

Tune the detectors in Settings > Structured Anomaly Detection. See Structured Anomaly Detection Settings for each setting, its range, and its default.

The work zones used for baselines are:

  • Working hours — 09:00 to 17:00, Monday to Friday.
  • After hours — outside working hours on weekdays.
  • Weekend — Saturday and Sunday.

Work zones follow the timezone set in the settings.

How the detection job runs​

The Structured Anomaly Detection scheduled job runs every detection interval (default 30 minutes). Each run does two things against two different time windows:

  1. Rebuilds baselines over the rolling baseline window (default 30 days). This is the "normal" against which each user and table is measured.
  2. Scores the most recent 24 hours of activity against those baselines and records any anomalies.

The 24-hour window is a sliding window that ends at the current time on every run. It is not limited to events since the previous run, so consecutive runs examine overlapping activity. Each finding has a deterministic ID, so detecting the same event again does not create a duplicate entry in the feed.

This has two practical effects:

  • An anomaly usually appears within one detection interval of the activity. It is confirmed again on later runs while it stays inside the trailing 24 hours.
  • The detectors compare against the median and MAD, which a single unusual day barely changes, rather than the mean. Because the rebuilt baseline includes the most recent activity, a brand-new access can be absorbed into it in the same run. It then often appears as a rare access rather than a first-seen novelty.

Severity, confidence, and alarms​

Each anomaly that fires appears in the transitions feed with the metric, z-score, baseline values, and an explanation.

  • Off-hours — severity is raised by one level for activity at off-hours.
  • Reads and writes — reads are weighted as possible exfiltration. Writes and DDL are weighted as possible tampering or destruction.
  • Alarms — only HIGH and CRITICAL anomalies raise alarms. MEDIUM and LOW anomalies appear in the feed only.
  • Acknowledge and suppress — these use the same controls as the rest of the transitions feed.

Deviation size​

For volume and write-surge anomalies, severity depends on how far the activity is above the user's own median, not only on the statistical score.

Activity compared with the baseline medianSeverity
About 2 timesMEDIUM
About 10 timesHIGH
About 100 timesCRITICAL

To reduce noise, a volume anomaly must also be at least 1.5 times the baseline median before it fires. The exception is a baseline that is effectively zero. Activity the user has never done before, such as a read-only account issuing its first DELETE, is always flagged regardless of size.

Sensitive data​

Access to a catalog, schema, or table whose name indicates sensitive data raises the severity. Examples of such names are finance, payroll, hr, pii, phi, ssn, and secret.

  • A novel access to a sensitive scope is CRITICAL.
  • A rare access to a sensitive table is HIGH instead of MEDIUM.

The check applies to each part of the name, so postgres.finance.payroll is treated as sensitive even though the postgres catalog is not.