Skip to content

fix: make HasProfileData an existence probe instead of a full-table scan - #6402

Open
javicj wants to merge 1 commit into
parca-dev:mainfrom
javicj:fix/hasprofiledata-full-table-scan
Open

javicj wants to merge 1 commit into
parca-dev:mainfrom
javicj:fix/hasprofiledata-full-table-scan

Conversation

@javicj

@javicj javicj commented Sep 4, 2026

Copy link
Copy Markdown

What does this PR do?

HasProfileData delegated to ProfileTypes with a zero time range. ProfileTypes
skips its time filter when start and end are zero, so every call executed:

SELECT DISTINCT name, sample_type, sample_unit, period_type, period_unit, (duration > 0)
FROM <table>

with no WHERE and no LIMIT — and the result was immediately reduced to
len(types) > 0. DISTINCT cannot terminate early, so this scanned the entire
table on every call just to answer a boolean. HasProfileData backs the UI's
empty-state check and runs on every page load; with hundreds of millions of rows
this made the UI landing screen take seconds to minutes.

This replaces it with SELECT 1 FROM <table> LIMIT 1, which reads one granule
and stops. Semantics are unchanged: both answer "does the table contain at
least one row", since any row necessarily has profile-type values (non-nullable
columns).

Why is it needed?

Observed in production with ~426M rows: the empty-state check was the single
most expensive query hitting ClickHouse, and under load it held connection
pool slots long enough to starve all other queries.

How to test it?

  • go test ./pkg/clickhouse/
  • Manually: with data in the store, the UI shows the explorer instead of the
    "no data yet" prompt; with an empty store, the prompt still shows. The
    response is now immediate regardless of table size.

HasProfileData delegated to ProfileTypes with a zero time range, which
skips the time filter and runs a DISTINCT over every row in the table,
only for the result to be reduced to a boolean. A LIMIT 1 probe returns
the same answer after reading a single granule. HasProfileData backs the
UI's empty-state check and runs on every page load.
@guettlibot

Copy link
Copy Markdown

Confirming this from production, with numbers — and closing #6417, which I opened a month after this without searching first and which makes the same change.

Running a self-hosted Parca on the ClickHouse backend with 72h of retention, the HasProfileData query reports:

X-ClickHouse-Summary: {"read_rows":"108591120","read_bytes":"7930500128",
                       "result_rows":"9","elapsed_ns":"1436439879"}

108,591,120 rows / 7.9 GB read to return 9 rows. 1.3s warm, and 61s cold — so the "seconds to minutes" in the description is accurate, and the cold case is the one users feel, since this runs on page load.

Two details that may be useful for review:

  • The reason it reads everything is slightly subtler than DISTINCT not terminating early: ProfileTypes applies its time filter only when both bounds are non-zero, and HasProfileData passes (UnixMilli(0), UnixMilli(0)), so the WHERE is skipped entirely. Worth noting because the same conditional exists in Labels and Values, which full-scan the same way when called without a time range.
  • ProfileTypes' zero-bounds behaviour is deliberately left alone here, which seems right — other callers may rely on the all-time semantics.

Your version is tighter than mine was. The one thing mine had that this doesn't is a regression test: executing the query needs a live ClickHouse, while the regression is entirely visible in the SQL, so I pulled the statement into a small hasProfileDataQuery(table string) string and asserted its shape — LIMIT 1 present, no DISTINCT/GROUP BY/ORDER BY. Happy to open that against your branch if you want it, or to leave it out to keep this minimal. Your call.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants