ClickHouse Health
The ClickHouse Health page provides deep diagnostics for the analytics database that stores all audit events, copy job metrics, and ingest telemetry. Use it to investigate database performance issues, stuck mutations, and storage growth.
Overview
ClickHouse is the authoritative store for audit event data. This page shows internals that application-level status checks do not show: mutation state, merge queue depth, active queries, and per-table storage metrics. Review it regularly, particularly after you run data wipe operations.
The page refreshes every 15 seconds by default. Pause in the header freezes the view while you investigate, and Resume restarts the refresh.
Summary cards
The summary cards at the top give a quick snapshot of database health.
| Card | Meaning |
|---|---|
| Uptime | How long the ClickHouse server process has run since its last restart |
| Total rows | Row count across all tables. Use it to track data growth over time. |
| Active queries | Queries currently executing. A persistently high count can indicate a backlog of slow queries. |
| Memory usage | Current memory consumption of the ClickHouse process |
| CPU threads | CPU threads currently in use by ClickHouse |
| Merge queue depth | Part merges waiting to be processed. A high and growing depth together with a high part count on a table can temporarily slow inserts while merges catch up. |
| Replication lag | For replicated setups only. How far replicas are behind the primary. |
Per-Table Health
Each table is listed with these metrics:
| Metric | Meaning |
|---|---|
| Row count | Total rows stored in the table |
| Compressed size | Disk space used after compression |
| Uncompressed size | Logical data size before compression. The ratio to the compressed size shows how well the table's data compresses. |
| Part count | Number of data parts the table is split into |
ClickHouse merges parts in the background. A very high part count means many small inserts arrived faster than ClickHouse can merge them. This is a warning sign, but it usually resolves itself once the insert rate drops.
Active Mutations
This panel lists the ALTER TABLE mutations that are in progress or recently completed.
- done (green) — the mutation completed successfully.
- Any other state (amber) — the mutation needs attention.
A completed mutation that stays registered in the system re-applies to all new data inserted into that table. If a completed mutation is not cleaned up, issue a KILL MUTATION command through the ClickHouse HTTP API immediately.
Slow Query Log
Queries that exceeded the configured slow-query threshold appear in the Recent Queries — Last 1 Hour panel. Click a row to expand the full SQL. Use the panel to find inefficient query patterns or missing filters that cause full-table scans.
The slow-query log is off by default to conserve disk space. When it is disabled, the panel shows a "query log disabled" badge and an explanation instead of a table. The Active Processes panel remains the live view of what is running now. The panel fills automatically if query log retention is later enabled on the ClickHouse server.
Active Processes
This panel lists the queries running now, with the elapsed time and the user that submitted each query. Use it to find runaway queries that consume disproportionate resources. A query that has run for an unexpected length of time can be killed through the ClickHouse HTTP API.
Workflows
Check for stuck mutations after a data wipe
- Open the Active Mutations panel.
- Look for any mutation in a state other than done.
- If a completed mutation is still present, issue a
KILL MUTATIONfor it through the ClickHouse HTTP API. - Refresh the page to confirm the mutation is removed.
Investigate a table with slow inserts
- Open the Per-Table Health panel.
- Find the table and check its part count.
- If the part count is very high, check Merge queue depth in the summary cards.
- If the merge queue is also high, the table is catching up on merges. Inserts stay slow until the merges complete.
Find the cause of high ClickHouse CPU usage
- Open the Active Processes panel.
- Look for queries with a high elapsed time that belong to the application user.
- Expand the Recent Queries — Last 1 Hour panel to read the full SQL of any suspicious query. This panel requires the slow-query log to be enabled.
Tips
- After you run any wipe or delete operation from the application, return to this page and check Active Mutations. Completed mutations that are not killed re-apply silently to every new row inserted into the affected table and corrupt future data.
- High merge queue depth with many parts on a table usually resolves within minutes once insert pressure eases. It needs manual intervention only if it persists for hours.
- Use Pause when you need to read a slow-query entry or examine a mutation without the panel refreshing. Select Resume when you are done.