Tools for cataloguing the ~13,000 commemorative paver bricks removed from Town Square Park (Municipality of Anchorage) during the park's renovation. Every brick was photographed on its warehouse pallet; a vision LLM reads each photo, the read is fuzzy-matched against the Municipality's official brick lists, humans review the leftovers in a hosted click-through page, and the results feed a public search page plus Excel reports for Parks & Recreation — so a brick buyer can learn up front whether their brick survived the move and which pallet it sits on, instead of searching a warehouse.
Live public search page: https://codeforanchorage.org/bricks/
(GitHub Pages, built from docs/ in this repo). The staff review and
verification pages are hosted separately behind basic auth — see
web/DEPLOY_DREAMHOST.md for the full hosting recipe.
The system end to end (run_pipeline.py runs the whole loop in order):
- Photograph — one photo per brick, filed by warehouse pallet
(
photos/pallets/<pallet>/IMG.jpg). - OCR —
single_pipeline.pyreads every photo with the production model Gemini 3.1 Flash-Lite; parallel, resumable, retrying. - Reference build — parse the two official Municipality brick lists
and merge them into one master lookup (
parse_xls_list.py,parse_tsp_list.py→resolve_tsp_rows.py,merge_lists.py). - Match —
match.pyfuzzy-matches each photo read against the master list (scan-confusable folding, phonetic folding, token containment with a uniqueness margin). - Human review —
make_review_page.pyandmake_fp_page.pybuild hosted click-through pages; decisions autosave to a small PHP receiver (web/receiver.php) andapply_decisions.pyfolds them back into the catalogue. - Deliverables —
make_search_page.py(staff and public search pages),make_report.py(per-section Excel workbook),make_pallet_report.py(per-pallet workbook for the warehouse floor), and per-sectionmissing_<S>.csvlost/broken candidate lists.
A pytest regression gate — 26 labeled warehouse photos that must all match with zero false positives — runs before any rebuilt data ships.
| Where | Model | Notes |
|---|---|---|
Production photo OCR (single_pipeline.py) |
gemini-3.1-flash-lite |
validated 26/26 on the labeled set, zero wrong IDs |
| Second opinion / escalation | gemini-3-flash-preview |
scored identically; its hidden thinking tokens bill as output, ~10x the cost in bulk |
Scan-row re-reads (rescan_rows.py) |
gemini-3.1-flash-lite + gemini-3.5-flash-lite |
a new read is adopted only when both models agree |
Unmatched-photo classification (classify_photos.py) |
gemini-3.1-flash-lite |
labels photos single brick / stack / other |
One-time list re-OCR (reocr_tsp_pdf.py, resolve_tsp_rows.py) |
gemini-2.5-flash |
historical builds of the by-name list, already done |
The project started as a bake-off, and pipeline.py still runs any mix
of methods side by side (--methods): PaddleOCR (local, GPU
optional), Claude Sonnet 4.6 (claude-sonnet-4-6), Claude Sonnet
5 (claude-sonnet-5), Gemini 2.5 Pro, Gemini 2.5 Flash
(deprecated by Google, shutdown 2026-10-16), Gemini 3 Flash Preview,
and Gemini 3.1 / 3.5 Flash-Lite.
PaddleOCR detects text line by line; the pipeline then groups those
detections back into bricks by spatial layout (see group_bricks.py), so all
methods produce one row per brick for an apples-to-apples comparison. Pass
--no-group to inspect PaddleOCR's raw per-line detections instead.
- Python 3.11, 3.12, or 3.13.
⚠️ Not Python 3.14 if you want the PaddleOCR method:paddlepaddle(PaddleOCR's runtime) publishes no wheels for 3.14. - A Google Gemini API key — this is all the production path needs.
- Optional: an Anthropic API key (only for the Claude methods in the comparison harness) and an NVIDIA GPU for PaddleOCR acceleration.
Install a supported Python and create a virtual environment:
winget install Python.Python.3.12
py -3.12 -m venv .venv
.\.venv\Scripts\Activate.ps1pip install -r requirements.txtThat installs the LLM methods (anthropic, google-genai, Pillow) and
paddleocr. PaddleOCR also needs the paddlepaddle runtime, installed
separately to match your hardware — pick one:
# CPU build — always works, fine for small test batches:
pip install paddlepaddle
# GPU build (CUDA) — get the exact command for your CUDA version from
# https://www.paddlepaddle.org.cn/en/install/quick , e.g.:
pip install paddlepaddle-gpu==3.0.0 -i https://www.paddlepaddle.org.cn/packages/stable/cu126/GPU note — Blackwell cards (RTX 50-series). PaddlePaddle's published GPU wheels target CUDA 12.6, which predates Blackwell (
sm_120) support, so GPU init may fail on those cards. The pipeline automatically falls back to CPU in that case — fine for small batches. Pass--no-gputo skip the GPU attempt entirely.
The two LLM methods need API keys:
- Anthropic (Claude Sonnet) — from https://console.anthropic.com/
- Google Gemini (Gemini Pro & Flash) — from https://aistudio.google.com/
Recommended — .env file. Put the keys in brick-ocr/.env, one per line:
ANTHROPIC_API_KEY=sk-ant-...
GEMINI_API_KEY=...
pipeline.py loads .env automatically on startup. The file is git-ignored,
stays out of shell history, and is separate from Claude Code's own
ANTHROPIC_API_KEY auth.
Alternative — environment variables. A real environment variable wins: if
a key is already set in the environment, .env does not override it.
$env:ANTHROPIC_API_KEY = "sk-ant-..." # current terminal only
$env:GEMINI_API_KEY = "..."If a method's key is missing the pipeline still runs the others and records an error row for the method that could not run.
Drop a few brick JPEGs into test_images/ to try the comparison harness
before pointing the production runner at a full photo tree.
Run from inside the brick-ocr/ directory:
python pipeline.py --input test_images/ --output output/results.csv| Flag | Default | Purpose |
|---|---|---|
--input |
(required) | Folder of .jpg / .jpeg / .png images |
--output |
(required) | CSV file to write |
--methods |
sonnet,gemini-pro,gemini-flash |
Methods to run; add paddle for the OCR cross-check |
--tiles |
1x1 |
Split each image into an RxC overlapping tile grid (see below) |
--tile-overlap |
0.2 |
Tile overlap as a fraction of tile size |
--no-gpu |
off | Force PaddleOCR onto the CPU |
--cpu-threads |
8 |
PaddleOCR CPU inference threads |
--no-group |
off | Report raw PaddleOCR lines instead of grouped bricks |
--brick-gap |
1.5 |
Line-into-brick grouping sensitivity (see below) |
To run a subset of methods:
python pipeline.py --input test_images/ --output output/results.csv --methods paddle,gemini-proTerminal — a per-image comparison with one stacked block per method, each result tagged with its confidence (PaddleOCR also shows its raw 0-1 score).
CSV — written to --output incrementally (one image at a time, with a
flush, so an interrupted run keeps every finished image), with columns:
| Column | Notes |
|---|---|
filename |
source image |
method |
paddleocr, claude-sonnet, gemini-pro, gemini-lite-31, ... (one per method run) |
brick_inscription |
one row per brick; a brick's lines joined with / |
confidence |
high / medium / low, plus none (nothing found) or error |
(The production runner, single_pipeline.py, writes a richer catalogue —
see below — but this comparison CSV is where new methods get judged.)
A single wide photo spreads its pixels across dozens of bricks, and vision
APIs downscale each image (Anthropic caps it at ~1.15 MP) — so on a wide shot
every brick reaches the model at only ~100 px. --tiles RxC (e.g. --tiles 3x3) crops each photo into an R x C grid and OCRs every tile separately, so
each brick is read at far higher effective resolution; for the LLM methods
each tile is also its own full-detail API call.
python pipeline.py --input test_images/ --output output/results.csv --tiles 3x3Tiles overlap (--tile-overlap, default 0.2) so a brick on a cut line still
appears whole in one tile; the pipeline then removes the duplicate detections
the overlap produces (by box overlap for PaddleOCR, by text similarity for the
LLM methods) and maps everything back to full-image coordinates before grouping.
Notes:
- Choose the grid so each brick is comfortably smaller than the overlap band
(
2 x tile-overlap x tile-size), or boundary bricks may be split. --tiles 3x3means 9 API calls per image, per LLM method — more thorough, ~9x the cost. Size the grid to your images.- If the re-shot photos are tight (just a few bricks per frame),
--tiles 1x1(the default) is fine and tiling is unnecessary.
group_bricks.py reassembles PaddleOCR's individual text detections into
bricks in two spatial passes: same-row detections merge into lines, then
vertically-adjacent, horizontally-aligned lines stack into a brick (normally
1-3 lines). Thresholds are multiples of the median detected text height, so
they scale with image resolution.
The defaults are tuned on a synthetic test grid and will likely need tuning
against real brick photos. The most layout-sensitive knob is --brick-gap
(how large a vertical gap still counts as "same brick"):
- Separate bricks getting merged into one → lower
--brick-gap. - One brick's lines being split apart → raise
--brick-gap.
The pipeline warns when a grouped brick exceeds 3 lines, a likely sign of
over-merging. Use --no-group to see the raw detections while tuning.
brick-ocr/
pipeline.py whole-image comparison pipeline + CSV writer
brick_pipeline.py per-brick pipeline: detect -> crop -> OCR each brick
single_pipeline.py one-photo-per-brick production runner (parallel, resumable)
detect_bricks.py classical-CV paver detector (mortar-joint grid)
parse_xls_list.py parse the source Excel workbook of new brick numbers
parse_brick_list.py parse the by-area brick-list PDF (superseded by the .xls)
parse_tsp_list.py parse the original by-name (all bricks) PDF into a CSV
reocr_tsp_pdf.py re-transcribe the scanned by-name PDF with a vision LLM
resolve_tsp_rows.py settle disputed rows via isolated strip reads -> v2 list
rescan_rows.py re-read troublesome scan rows with two current models;
adopt the new text only when both agree
merge_lists.py merge both lists into the master lookup table
match.py match the OCR catalogue against an official list
make_review_page.py pack the review queue into one offline review.html
(--receiver-url adds autosave-to-server;
--photo-base-url uses hosted derivatives instead of
embedding -- required for big queues)
make_fp_page.py duplicate-claim (false-positive) check page:
photos whose reads disagree but claim the same brick
make_search_page.py build the pickup-counter search page (search.html);
--public builds the visitor variant for GitHub Pages
make_derivatives.py build thumbs/ + zoom/ JPEG trees for photo hosting
make_strips.py render each scanned-list row as an image; review
pages (--master) show it so humans read the PRINT,
not the OCR of it
strip_image.py shared strip-cropping/rendering code
apply_decisions.py fold a reviewer's decisions.csv back into the catalogue
make_report.py build the Parks & Rec Excel workbook (per-section,
reviewer notes, Unofficial-bricks sheet)
make_pallet_report.py
the warehouse workbook: one tab per PALLET, rows in
original-number order -- printed and posted per pallet
classify_photos.py label unmatched photos: single brick / stack / other
run_pipeline.py the whole loop in one command (see workflow below)
hostpaths.py shared photo -> thumbs/zoom derivative path mapping
pagenav.py shared nav bar for the hosted staff pages
web/ Dreamhost hosting: receiver.php + DEPLOY_DREAMHOST.md
docs/ the built PUBLIC search page, served by GitHub Pages
consensus.py fold tables (scan-confusable, phonetic) + text scoring
shared by the matcher and the search pages
ocr_paddle.py PaddleOCR wrapper
ocr_anthropic.py Anthropic provider (Claude)
ocr_google.py Google provider (Gemini)
vision_ocr.py shared prompt + image/JSON handling for the LLM methods
tiling.py splits an image into overlapping tiles
group_bricks.py groups PaddleOCR line detections into bricks
compare.py stacked terminal comparison output
requirements.txt dependencies
reference/ official Municipality brick lists (see below)
tests/ pytest suite incl. the 26-photo regression gate
test_images/ input JPEGs go here
output/ results, crops, and annotated images land here
single_pipeline.py is the runner for the full warehouse batch (one photo =
one brick). It is built to survive a ~13,000-photo run:
python single_pipeline.py --input photos/ --output output/singles.csv --workers 8-
The default method is
gemini-lite-31(gemini-3.1-flash-lite), validated 2026-07-29 on the labeled set: 26/26 matched, 0 wrong IDs — including three worn bricks the previous default left in review. It replacesgemini-flash(gemini-2.5-flash), which Google deprecates on 2026-10-16.gemini-flash-3(Gemini 3 Flash Preview) scored identically and serves as the second-opinion/escalation method. -
--workers N(default 8) OCRs images concurrently; the LLM calls are network-bound, so threads scale nearly linearly (a serial run would take ~8 hours; 8 workers cut it to ~1). -
Every API call retries transient failures (429/5xx, malformed replies) 3 times with backoff inside the provider (
ocr_google.py/ocr_anthropic.py) before anERROR:row is recorded. -
Every row is flushed as it is written -- a crash or Ctrl-C loses nothing.
-
--resumekeeps a previous run's good rows and redoes only missing images andERROR:rows. An empty read is kept (the model really saw no text); anERROR:read is retried. -
Subdirectories are walked, and two folder conventions are read back into the catalogue -- needed both because the warehouse pallet labels are arbitrary (K, H1..H6, ...) and do not encode the park section:
layout when tags recorded photos/pallets/<pallet>/IMG.jpgsection unknown (the usual case) pallet only photos/<SECTION>/<pallet>/IMG.jpgsection known at photo time section + pallet A pallet-only photo simply matches against the full master list, and the match itself reveals the section (every master row carries one) -- the pallet tag is what the pickup counter needs to find the physical brick. The
imagecolumn is the path relative to--input(so repeated camera filenames on different pallets stay distinct rows, including across--resume). A flat folder of photos works exactly as before (both columns empty); a nested folder matching neither convention is reported at startup and its photos get no tag -- rename it before the run, not after. Downstream tools treatimageas an opaque key, so matched CSVs, the review page, and decisions files join up unchanged; givemake_review_page.py --photosthe same root the run used as--input.
So the crash-recovery loop is simply: re-run the same command with --resume
until the end-of-run summary reports no remaining ERROR reads.
python -m pytest tests/No API calls, pure CSV -- safe to run anywhere. Two layers:
- Unit tests pin every measured matching behaviour: the scan-confusable folds, the phonetic (voice-transcription) folds, token-containment scoring and its uniqueness margin, identical-copy handling, the E/F/G original-number ranges, NO-BRICK detection, and the resume bookkeeping.
- The regression gate (
tests/test_regression_labeled.py) replays the 26 labeled warehouse photos' frozen OCR reads (tests/fixtures/, from the production methodgemini-3.1-flash-lite) throughmatch.pyagainst the committedreference/master_list.csvand asserts the validated baseline: 26/26 matched to the verified brick, zero false positives. Any change to the matching layers -- or a master-list rebuild that breaks identification -- fails here first.
One command runs the whole loop (configuration in .env — API keys,
receiver URL/token, staff login, photo base URL):
python run_pipeline.py # ocr(resume) -> merge -> GATE -> match ->
# classify -> pull decisions -> apply ->
# pages -> Excel report
python run_pipeline.py --rescan # also re-read troublesome scan rows (API $)
python run_pipeline.py --only pages,report # regenerate outputs onlyThe regression gate runs before anything ships: if the rebuilt data breaks
validated matching, the pipeline aborts and the previous master list is
kept at reference/master_list.csv.prev. Reviewer decisions are pulled
straight from the hosted receiver (its action=list/action=fetch GET
API), so a hosted review round needs no manual downloads at all. The
pages step also builds fp_review.html (the duplicate-claim check
page) whenever output/fp_candidates.csv exists.
The individual steps, for running by hand:
# 1. OCR the photos (parallel, resumable -- see above)
python single_pipeline.py --input photos/ --output output/singles.csv --workers 8
# 2. Identify each photo against the master list
python match.py --catalog output/singles.csv --reference reference/master_list.csv \
--output output/matched.csv --scan-ocr
# -> matched.csv, review_matched.csv (unmatched queue), duplicates_matched.csv (QA),
# missing_<S>.csv per photographed section (lost/broken candidates)
# 3. Pack the review queue into ONE self-contained page and send it out
python make_review_page.py --review output/review_matched.csv \
--photos photos/ --catalog output/singles.csv --output output/review.htmlreview.html opens in any browser, fully offline: each undecided photo is
embedded next to its top candidates as clickable choices (plus "None of
these" / "Can't read the photo"). Choices autosave in the browser, and one
button downloads decisions.csv — the reviewer emails that single small file
back. Multiple reviewers / sittings produce multiple decisions files; all are
accepted below.
# 4. Fold the human decisions back in
python apply_decisions.py --matched output/matched.csv \
--decisions decisions.csv --output output/matched_final.csv
# 5. Build the deliverable: the Parks & Rec Excel workbook
python make_report.py --master reference/master_list.csv \
--matched output/matched_final.csv --output output/brick_report.xlsxThe workbook has a Summary sheet (per-section totals), an All-bricks sheet,
and one sheet per section, each row a brick with buyer, inscription, both
numbers, review flags, and its photo status — Present rows highlighted. A
blank photo status means not photographed yet, which is not evidence a
brick is missing until its section is photographed in full; the Summary sheet
says so in words, because that distinction is the whole point of the count.
A second workbook answers the warehouse-floor question — what is actually stacked on this pallet? — with one tab per pallet, rows in original-number order, both brick numbers, both buyer spellings, and the photo's own OCR read. It is built for printing and posting at each pallet:
python make_pallet_report.py --matched output/pallets_final.csv \
--master reference/master_list.csv \
--output output/brick_report_by_pallet.xlsxPallet labels are warehouse labels, not park sections: most pallets are ~90% one section but every one carries strays, so the Section column on each tab is the retrieval truth.
Three staff pages — search.html, review.html, fp_review.html — are
hosted as plain static files behind one basic-auth login, next to the
photo derivative trees (thumbs/, zoom/, strips/ from
make_derivatives.py / make_strips.py) and one small PHP script,
web/receiver.php, which stores reviewers' autosaved decisions. Any
static host with PHP works; web/DEPLOY_DREAMHOST.md is the complete
recipe (subdomain, basic auth, token, upload, pulling decisions back).
The public page is a separate build of the same search page
(make_search_page.py --public): visitor-facing help text, no nav bar,
no token — safe to publish. It is committed as docs/index.html and
served by GitHub Pages at https://codeforanchorage.org/bricks/.
Refreshing it is rebuild-then-commit. Photos are not in the repo (far
too big), so --photo-base-url makes the public page hotlink the hosted
derivative trees; omit it for a pure text-search build. Google Analytics
(GA4) is injected on the --public build only — the staff pages carry
no tracking.
The page generators are unit-tested, but the in-browser behaviour is not. After regenerating pages, a two-minute pass over this list catches what pytest can't:
search.html (log in, hard-refresh)
- Type a surname -> results appear as you type;
Escclears. - Type a misspelling ("JOLY COY") -> the right brick still ranks first.
- Type a brick number -> both numbering eras listed, era labelled.
- A photographed brick shows "at pickup site -- pallet X" and a
verifybutton; clicking it shows photo + OCR read + scan row. Clicking the photo or the scan row magnifies it in place (the full zoom image loads only on that click); click again orEsccloses. - If any reviewer confirmed "None of these": search their brick's words -> an "unofficial" entry appears with pallet and note.
review.html 6. Photos and candidate strip images load; hovering a strip magnifies it. 7. Picking a candidate turns the card green; "Clear choice" un-decides it; both survive a page reload (localStorage). 8. With the receiver live: a decision shows "saved to server HH:MM" within a few seconds. 9. Stack photos sit at the bottom under the banner, one-click "Stack / pallet overview photo" first, candidates still available beneath. 10. "Download decisions.csv" produces a file with your reviewer name.
all three pages (search.html, review.html, fp_review.html) 11. The header nav bar clicks through to the other two pages without a second login prompt (one basic-auth realm covers the directory); review.html and fp_review.html show today's date as "built ..." in the nav.
There are two Municipality lists, and they complement each other. Neither alone is enough.
brick_list_xls.csv (by area) |
tsp_brick_list_v2.csv (by name) |
|
|---|---|---|
| source | TSP Bricks All.xls, the renovation-era source workbook |
TSP Bricks ALL - OG List by Name - OCR.pdf, a scan |
| rows | 8,281 | 13,389 |
| coverage | areas A–D and H–K only | all bricks, including E, F, G |
| gives you | the section + grid position (Column/Row) | the brick # and the buyer's name |
| text quality | clean (digital source) | two noisy transcriptions, cross-checked (see below) |
The by-name list's canonical form is v2 (resolve_tsp_rows.py): each row
carries the coordinate parse of the scan's embedded text (full_name) and
a vision-model re-read (alt_name) — whole-page (reocr_tsp_pdf.py) where
the two agree, an isolated single-row strip where they disputed. The two
transcriptions err in complementary ways (the parse has right rows / noisy
glyphs; the model has clean glyphs / occasional wrong-row slips), so
merge_lists.py and match.py score against both and keep the better —
neither ever replaces the other. Strip renumberings are accepted only when
they fill an unclaimed, in-range brick number; everything else stays flagged
(verified / flag columns) for the review pile. tsp_brick_list.csv (the
parse alone) is kept as the v2 build input.
brick_list_xls.csv (from parse_xls_list.py) supersedes brick_list.csv,
the older parse of the printed PDF (ABCDHIJK.pdf, a.k.a. "ABCDHIJK New
Brick Numbers.pdf"): the workbook is what that PDF was printed from, and the
PDF text-extraction had lost 407 rows (mostly area C), truncated 68
inscriptions, and mis-sectioned 111 boundary bricks.
A separate data caveat that applies to both lists: some entries were
transcribed by voice, so the list can spell a name phonetically while the
brick spells it properly (BRIAN vs BRYAN, CATHY vs KATHY). Matching folds the
classic phonetic equivalences (consensus.phonetic_fold) so those pairs
compare as equal.
Why two lists exist (per the 2009 How to Find Your Brick brochure and muni.org): the 2009 renovation relocated ~8,000 bricks and renumbered them — those are areas A–D/H–K and the by-area list ("new brick numbers"). Areas E, F and G were not moved in 2009, so they kept their original (certificate) numbers and never appeared in the by-area list. The by-name list's Brick# is the original number for every brick. The original numbers of the unmoved areas are contiguous ranges, so for them the number alone gives the section:
| original # | area | 2009 fate |
|---|---|---|
| 1–3,377 | F | unmoved — original # still valid |
| 8,279–9,126 | G | unmoved — original # still valid |
| 9,127–10,070 | E | unmoved — original # still valid |
| everything else | A–D, H–K | relocated + renumbered — look up by text |
For a moved brick the two numbers are unrelated ("Travis E Williams" is
#3770 originally, #7305 in area I now) — join on inscription text, never
on the id.
merge_lists.py combines both lists into reference/master_list.csv — one
row per original brick with orig_id, new_id, section, moved, status,
buyer, both inscriptions, and two review columns: og_verified (the v2
list's per-row trust tier: agreed / strip / parse) and flag, a
;-joined list of v2's row flags (number?:…, page0, …) plus two checks
added at merge time — dup_orig (the same original number appears on more
than one OG row: a residual scan id collision, so inside E/F/G two rows claim
one physical brick) and orig_range (an impossible certificate number,
outside 1–13,344). Flagged rows keep their best-effort assignment; the flag
routes them to human review, it does not withhold the data. Current counts:
1,374 rows carry a flag (308 dup-id, 3 out-of-range, 1,203 carried from v2). Unmoved bricks get their section from the
number ranges; moved bricks are text-joined to the by-area list (word-blocked
fuzzy match at ≥0.80, plus a stricter rescue pass — lower score but a clear
margin over the best differently-inscribed runner-up and buyer-surname
corroboration; identical-copy batches are then assigned one-to-one, strongest
join first, so a weak rescue can never displace a rightful 1.00 owner —
contested bricks go to the best-scoring claimant and the loser drops to
review). Current yield with the v2 OG list and
the .xls source: 12,902 of 13,389 rows fully resolved, 468 flagged unjoined
for review, 19 sales recorded as "NO BRICK NO INSCRIPTION". By-area rows no OG
row claimed land in master_unclaimed.csv.
python merge_lists.py --og reference/tsp_brick_list_v2.csv \
--new reference/brick_list_xls.csv --output reference/master_list.csvmatch.py scores each candidate two ways and reports which won in the
match_basis column: text (whole-string similarity, threshold --min-score)
and tokens (containment — each read word scored against its best counterpart
in the inscription, for worn bricks whose read is a noisy subset like
"GRAND BEATY BUDDY" for "HAROLD G BEATY 1938-1991 MY 'BUDDY'"). Because a
generic word set fits many bricks, a tokens match must also beat the best
differently-inscribed candidate by a clear margin; in exchange it gets a
slightly lower score bar. Hallucinated reads tie across several bricks and are
rejected by the margin; distinctive subsets stand alone and pass. A tokens
winner that fails its own gate does not drag the photo to review when a
whole-string candidate clears the normal bar on its own — a noisy containment
tie can outscore the true text match by a hair (measured on the pallet-K test:
"JOLY COY" hit a 0.94 tokens tie while the real Joey Coy stood at text
0.93), so the matcher falls back to the best text-basis candidate.
When the catalogue carries a section column (single_pipeline.py fills it
from the pallet folder convention), match.py scopes each photo to its own
section's bricks — plus the master rows with no section assigned, which could
be anywhere — since the pallet says where the brick was lifted from. That
both shrinks the candidate pool and disambiguates identical inscriptions
sold in different sections. The whole list stays the fallback: an accepted
cross-section match is kept but flagged off-section in the section_check
column (mis-sorted brick or mis-tagged folder — either way a human should
glance at it), and the end-of-run summary counts them. Review-queue
candidates for a tagged photo come from its section too. The global
--section flag overrides the per-row tags.
The section tags also drive the missing-brick reports: every section with at
least one tagged photo gets a missing_<S>.csv — its official bricks that no
photo (from any section) has matched, keyed on (section, id) since bare ids
collide across the two numbering eras. A section with no photos yet gets no
report, because it would just list itself in full. As always, a missing list
is only a real lost/broken list once its section is photographed in full;
with --section the report covers exactly that one section instead.
match.py accepts the master list directly, so a warehouse photo resolves to
section + current number + buyer in one step:
python match.py --catalog output/singles.csv \
--reference reference/master_list.csv \
--output output/matched.csv --scan-ocrBricks that stay unmatched are the human-review queue: alongside the matched
CSV, match.py writes review_<output name> listing each unmatched brick's
--top N best candidates (default 5), ranked best-first, one row per
candidate — identical-inscription copies collapsed to one slot, scored
against the whole reference (not the word-blocked pool, since a badly
misread brick may share no words with its own inscription). A reviewer sees
the photo plus its five most plausible bricks instead of re-searching the
list; --top 0 disables it.
As continuous QA, match.py also writes duplicates_<output name> whenever
more than one photo claims the same official brick. That is either a
duplicate photo (harmless) or a false positive — and the file's copies
column (how many identical copies of that inscription exist in the reference)
tells them apart: n_claims > copies means at least one claim is wrong.
This audits the matcher's zero-false-positive record for free on every batch.
This is the backbone of the pickup workflow: a visitor (or staff) searches the
master list by surname or inscription → does the brick exist? (status) →
which section / pallet group? (section); photo-matching against pallets
then confirms present / broken per brick; the review file catches the rest.
Because the by-name list is a scan, match against it (or the master list)
with --scan-ocr, which folds the confusable letter groups (I/L/T/1/J, E/F,
O/D/Q/0) on both sides. Rebuilding the by-name list from the PDF, in order:
python parse_tsp_list.py --pdf "TSP Bricks ALL - OG List by Name - OCR.pdf" \
--output reference/tsp_brick_list.csv # coordinate parse
python reocr_tsp_pdf.py --pdf "TSP Bricks ALL - OG List by Name - OCR.pdf" \
--pages all --output reference/tsp_reocr_full.csv \
--compare reference/tsp_brick_list.csv # whole-page re-OCR (~$2)
python resolve_tsp_rows.py --pdf "TSP Bricks ALL - OG List by Name - OCR.pdf" \
--reocr reference/tsp_reocr_full.csv \
--output reference/tsp_brick_list_v2.csv # strip-resolve disputes (~$1)--section does not apply to the by-name list (it has no section column).
Photography is complete: every warehouse pallet (park sections A–K)
has been shot, one photo per brick. Of ~11,500 brick photos, ~11,200
matched an official brick automatically and ~10,600 of those are
human-confirmed; a few hundred remain in the review queue. Every section
has its missing_<S>.csv lost/broken candidate list. The public search
page and the staff pages are live, and the Excel workbooks are the
running deliverable to Parks & Recreation. Numbers shift as review
rounds land — the CSVs in output/ (not committed) are the source of
truth.