Index the query log for search - #291
Merged
Merged
Conversation
Searching the query log read every row: the domain, client address, and answer of each one, then again to count the matches. On a million rows a term that matched nothing took over a second, and a month of a busy network is millions. On SQLite a trigram FTS5 index now serves those searches, and the Domain and Client filters. It stores no copy of the text. Each batch the log writes is indexed in its own transaction with one statement, which costs about the same as the write did before. Pruning leaves the index alone: searches join it to the log, so a pruned row's entry never shows, and a background pass every 15 minutes clears those entries in small batches. The upgrade that adds the index, and the next start after an older Sable wrote to the log, index the rows written without it in the background, newest first; searches read every row until that is done, and searches under 3 characters always do. On PostgreSQL the same LIKE searches are served by pg_trgm GIN indexes, built once in the background without locking the log. A database where the extension can't be enabled logs a warning and keeps reading every row.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Searching the query log read every row: the domain, client address, and answer of each one, and then the same again to count the matches. On a million rows a term that matched nothing took over a second. ns1 logged about 2 million queries last month. This indexes the log for search.
How it works
SQLite: an FTS5 table with the trigram tokenizer (
content='',contentless_delete=1), so it stores only the index, not a copy of the text.INSERT … SELECT. I tried triggers first: indexing row by row roughly doubled the cost of a write batch, and pruning a day held the write lock for 6.7 s.PostgreSQL: the search SQL doesn't change.
pg_trgmGIN indexes onname_key,client_ip_key, andLOWER(answer)serve itsLIKEconditions.CREATE INDEX CONCURRENTLY, so the log isn't locked. An invalid index left by an interrupted build is dropped and rebuilt.The index files are in
internal/store/query_log_search.go. The background pass ismaintainQueryLogSearchininternal/app. Backups already leave the query log out, so they don't grow.Measurements
One million realistic rows in SQLite through
QueryEvents(M3 Max):remarkable(112 matches)remarkableicloud(190k matches)icloud, AAAA onlyicloud, page 4017.253(667k matches)2603:7083(500k matches, half the log)The last row is the one case that got slower. That's counting 500k matches through the index, and I kept the count exact rather than letting it over-count right after a prune.
BenchmarkWriteQueryEventsBatch(256 rows)For ns1 that's about 0.7 GB of index and a minute or two of background indexing after the upgrade.
Tests
TestQueryLogSearchIndexFindsWhatTheFullReadFinds: the index and the full read return the same rows for long, short, quoted, punctuated, and filtered searches.TestQueryLogSearchIndexesAnUpgradedLog: an upgraded log is indexed across several batches. Rows written and pruned while it waits stay right, and the index ends with exactly the log's rows.TestQueryLogSearchIndexesRowsWrittenWithoutIt: rows an older Sable wrote after a downgrade get indexed.TestQueryLogSearchForgetsPrunedRows: a pruned row leaves search at once and the index at the next pass.TestQueryLogSearchBuildsPostgresTrigramIndexes: runs whenSABLE_TEST_POSTGRES_DSNis set. The indexes build and are valid, the planner uses all three, and results are right.TestQueryLogSearchWithoutPgTrgmStillSearches: runs whenSABLE_TEST_POSTGRES_PLAIN_DSNis set. A role withoutCREATEon the database gets one warning per start, and search still works.I ran both Postgres tests against
postgres:17locally. CI has no Postgres, so they skip there.Checks:
go tool mage verify,gofmt -l ., andgo test -race ./internal/storepass.Upgrading
No configuration changes. On SQLite the first start indexes the stored log in the background, about 30 s per million rows on fast hardware. Searches work the old way until it finishes. The index takes about 340 MB per million queries kept. An older Sable after a downgrade ignores the index. The next upgrade indexes whatever it wrote and clears what it pruned.