Skip to content

Index the query log for search - #291

Merged
drudge merged 2 commits into
mainfrom
claude/query-log-search-index
Oct 1, 2026
Merged

drudge merged 2 commits into
mainfrom
claude/query-log-search-index

Conversation

@drudge

@drudge drudge commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Searching the query log read every row: the domain, client address, and answer of each one, and then the same again to count the matches. On a million rows a term that matched nothing took over a second. ns1 logged about 2 million queries last month. This indexes the log for search.

How it works

SQLite: an FTS5 table with the trigram tokenizer (content='', contentless_delete=1), so it stores only the index, not a copy of the text.

  • Writes: each batch the log writes is indexed in the same transaction with one INSERT … SELECT. I tried triggers first: indexing row by row roughly doubled the cost of a write batch, and pruning a day held the write lock for 6.7 s.
  • Pruning leaves the index alone. Searches join the index to the log, so a pruned row's entry never shows. A background pass every 15 minutes clears those entries in batches of 5,000.
  • Upgrades: the upgrade that adds the index records the rows already stored, and the background pass indexes them newest first. Searches read every row until that's done. The same happens on the first start after an older Sable wrote to the log after a downgrade.
  • Searches of 3 or more characters use the index: the search box, and the Domain and Client filters. They start from the index and join the log, so a page stops after its rows. Text that's the same in every column, which is nearly always, is one phrase over the whole row instead of one per column. Shorter searches read every row as before.

PostgreSQL: the search SQL doesn't change. pg_trgm GIN indexes on name_key, client_ip_key, and LOWER(answer) serve its LIKE conditions.

  • They're built once per start in the background with CREATE INDEX CONCURRENTLY, so the log isn't locked. An invalid index left by an interrupted build is dropped and rebuilt.
  • If the extension can't be enabled, Sable logs one warning and searches read every row, as before.

The index files are in internal/store/query_log_search.go. The background pass is maintainQueryLogSearch in internal/app. Backups already leave the query log out, so they don't grow.

Measurements

One million realistic rows in SQLite through QueryEvents (M3 Max):

Search Reading every row Index
remarkable (112 matches) 1,341 ms 5 ms
Domain filter remarkable 256 ms 1 ms
icloud (190k matches) 610 ms 326 ms
icloud, AAAA only 763 ms 214 ms
icloud, page 40 595 ms 185 ms
17.253 (667k matches) 986 ms 561 ms
2603:7083 (500k matches, half the log) 445 ms 756 ms

The last row is the one case that got slower. That's counting 500k matches through the index, and I kept the count exact rather than letting it over-count right after a prune.

Cost Main This branch
BenchmarkWriteQueryEventsBatch (256 rows) 36–41 ms 38–45 ms
Prune a day (~33k rows) 1.4 s, index untouched
Clear the pruned entries (background) 2.0 s in batches of 5,000
Index an existing 1M-row log on upgrade (background) 30 s
Index size per million rows about 340 MB

For ns1 that's about 0.7 GB of index and a minute or two of background indexing after the upgrade.

Tests

  • TestQueryLogSearchIndexFindsWhatTheFullReadFinds: the index and the full read return the same rows for long, short, quoted, punctuated, and filtered searches.
  • TestQueryLogSearchIndexesAnUpgradedLog: an upgraded log is indexed across several batches. Rows written and pruned while it waits stay right, and the index ends with exactly the log's rows.
  • TestQueryLogSearchIndexesRowsWrittenWithoutIt: rows an older Sable wrote after a downgrade get indexed.
  • TestQueryLogSearchForgetsPrunedRows: a pruned row leaves search at once and the index at the next pass.
  • TestQueryLogSearchBuildsPostgresTrigramIndexes: runs when SABLE_TEST_POSTGRES_DSN is set. The indexes build and are valid, the planner uses all three, and results are right.
  • TestQueryLogSearchWithoutPgTrgmStillSearches: runs when SABLE_TEST_POSTGRES_PLAIN_DSN is set. A role without CREATE on the database gets one warning per start, and search still works.

I ran both Postgres tests against postgres:17 locally. CI has no Postgres, so they skip there.

Checks: go tool mage verify, gofmt -l ., and go test -race ./internal/store pass.

Upgrading

No configuration changes. On SQLite the first start indexes the stored log in the background, about 30 s per million rows on fast hardware. Searches work the old way until it finishes. The index takes about 340 MB per million queries kept. An older Sable after a downgrade ignores the index. The next upgrade indexes whatever it wrote and clears what it pruned.

drudge added 2 commits October 1, 2026 17:32
Searching the query log read every row: the domain, client address, and
answer of each one, then again to count the matches. On a million rows
a term that matched nothing took over a second, and a month of a busy
network is millions.

On SQLite a trigram FTS5 index now serves those searches, and the
Domain and Client filters. It stores no copy of the text. Each batch
the log writes is indexed in its own transaction with one statement,
which costs about the same as the write did before. Pruning leaves the
index alone: searches join it to the log, so a pruned row's entry never
shows, and a background pass every 15 minutes clears those entries in
small batches. The upgrade that adds the index, and the next start
after an older Sable wrote to the log, index the rows written without it
in the background, newest first; searches read every row until that is
done, and searches under 3 characters always do.

On PostgreSQL the same LIKE searches are served by pg_trgm GIN indexes,
built once in the background without locking the log. A database where
the extension can't be enabled logs a warning and keeps reading every
row.
@drudge
drudge enabled auto-merge (squash) October 1, 2026 21:36
@drudge
drudge merged commit 93e21bb into main Oct 1, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant