Skip to content

perf(manifest): project scan columns during reads - #1972

Merged
laskoviymishka merged 10 commits into
apache:mainfrom
fallintoplace:perf/project-scan-manifest-columns
Sep 8, 2026
Merged

laskoviymishka merged 10 commits into
apache:mainfrom
fallintoplace:perf/project-scan-manifest-columns

Conversation

@fallintoplace

@fallintoplace fallintoplace commented Aug 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Project manifest reads during local Arrow scan planning. PlanFiles() still returns full metadata.
  • Keep data-file metrics for row filtering and equality-delete pruning. Positional deletes and deletion vectors do not force data-file stats to be decoded.
  • Drop transient stats after filtering and delete matching. Reuse stripped delete files across tasks.
  • Cache projected Avro schemas by writer schema and projection.

Projected-read API

  • Adds ManifestEntryProjection, EntriesWithProjection, and NewManifestReaderWithProjection for scan planning.
  • IncludePruningStats includes value/null/NaN counts and lower/upper bounds. It still omits column_sizes and deprecated distinct_counts.
  • Adds DataFileWithoutColumnStats and ManifestEntryWithoutColumnStats for dropping transient metadata.
  • Prevent lossy rewrites: manifest writers and the DataFile Avro codec reject projected or stripped files with ErrInvalidArgument.

Benchmark

Command: go test . -run '^$' -bench BenchmarkManifestEntryProjection -benchtime=1x -count=1

Apple M1 Pro, 10,000 entries per case. These measurements cover manifest reading; task construction and delete matching are outside this benchmark:

Stats fields Full bytes Scan bytes Full allocs Scan allocs
10 32.6 MB 15.6 MB 443,928 65,797
100 260.2 MB 100.5 MB 2,257,763 77,930
1,000 2.93 GB 1.45 GB 20,512,582 332,865

Validation

  • make test
  • Focused race tests for projected reads, stat stripping, delete pruning, and codec rejection.
  • golangci-lint run --timeout=10m --allow-parallel-runners (v2.12.2)
  • Regression tests reproduce silent rewrites before the fix.
  • Coverage includes v1/v2/v3 data manifests, equality deletes, deletion vectors, and writer recovery after a rejected entry.

@fallintoplace
fallintoplace force-pushed the perf/project-scan-manifest-columns branch 2 times, most recently from 2662e4e to 589d91a Compare August 30, 2026 22:14

@zeroshade zeroshade left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The projection keeps every field its consumers need with one exception that only became a problem while this PR was open. Flagging it before it bites on rebase.

Major — dropped data-file stats defeat #1960's equality-delete pruning (table/scanner.go, projection/stat-drop branch)

Stats are retained for data manifests only when rowFilter != AlwaysTrue. So ToArrowRecords with the default AlwaysTrue filter drops ValueCounts, NullCounts, NaNCounts, and the lower/upper bounds.

Those are exactly the statistics that #1960 (perf(table): prune equality deletes by data-file metrics) reads to decide an equality-delete file cannot match a data file. Without them, equalityDeleteCanContainData falls back to its conservative answer and attaches every equality delete to every candidate data file. The unfiltered full-scan case is precisely where that pruning pays off most, so the optimisation is silently neutralised in the case it was written for.

To be clear about the impact: this is a performance regression, not a correctness bug. Conservative means more deletes get attached than necessary — results stay correct and deletes still apply. Nothing returns wrong rows.

Also to be fair about the timing: #1960 merged today, almost certainly after this PR's base. The projection logic was correct against the tree you wrote it on; the interaction only becomes live once this rebases onto current main. This isn't an oversight on your part, it's two changes meeting.

Suggested fix: include data-file stats in the projection whenever the scan has equality deletes attached, independent of whether a row filter is present. The filter presence turns out to be the wrong signal for "are stats needed" now that a second consumer exists.

Minor — projection test matrix is thin

Coverage is v3 data manifests and v2 data manifests only. There's no v1 case, and no delete-manifest or DV-manifest case. Projection correctness is version-sensitive — field sets genuinely differ across v1/v2/v3 — so a matrix across versions and manifest kinds would be worth having, especially for delete manifests where a missing field silently changes delete application rather than erroring.

What I verified is intact

The reassuring half — every other consumer still receives its required fields:

  • Partition values, for partition filtering and residuals.
  • Task basics: file_path, file_format, record_count, file_size_in_bytes.
  • Delete-file fields including content, equality_ids, and — importantly — referenced_data_file, content_offset, and content_size_in_bytes, all three of which deletion vectors require.
  • Row lineage: first_row_id, plus entry-level sequence_number and file_sequence_number that gate whether equality deletes apply at all.
  • status and snapshot_id.

CI green. Benchmark evidence is present.


This review was drafted by an AI-assisted tool and confirmed by an Apache Iceberg Go maintainer. The findings cite the project's review criteria; if you think one of them is mis-applied, please reply on the PR and a maintainer will weigh in.

More on how to contribute to Apache Iceberg Go: CONTRIBUTING.md

@fallintoplace
fallintoplace force-pushed the perf/project-scan-manifest-columns branch from 589d91a to 0a4f4e1 Compare September 1, 2026 05:30

@laskoviymishka laskoviymishka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Second pass, catching up to the three commits since my last review.

The two things I was holding on are resolved. The partition-copy race is fixed: DataFileWithoutColumnStats now calls initPartitionData() before the clone shares fieldIDToPartitionData by reference, with a new concurrency test hammering it. And the delete path is in good shape now, which is exactly where I had it backwards last round. I'd suggested dropping delete-manifest stats under an AlwaysTrue filter as wasted work, but that would have broken equality-delete pruning, which needs the data file's bounds and the equality-delete file's bounds on either side of the range comparison. The new retainDataFileStats / manifestProjectionForManifest / manifestProjectionRetainsDataStats logic keeps data-file stats whenever any delete manifest is present, and the comment spells out why delete manifests always keep stats: the manifest-list metadata doesn't distinguish equality from positional deletes, so you can't safely narrow it. That's the right conservative call, and a correctness fix I'd missed. TestManifestProjectionRetainsDataFileStatsForDeleteScans and TestManifestEntryProjectionSupportsManifestVersionsAndDeletes (v1/v2/v3 data, v2 equality delete, v3 DV) now pin the decision logic and the reader across versions.

DataFileWithoutColumnStats still returns non-*dataFile inputs unchanged, but the doc comment now states the reason, so I'm happy with that being documented rather than changed.

What's left is non-blocking polish, noted inline: the projection cache still calls writerSchema.String() on every lookup and stores it as the key, which quietly gives back some of the allocation win this PR is going for; the projected schema node shares Props / Aliases with the writer schema by reference; and the field whitelist still zeroes block_size_in_bytes on the projected path versus a full read. None of these block merge.

Approve with nits from me. Nice work on the equality-delete pruning fix.

Comment thread manifest_projection.go Outdated
projection ManifestEntryProjection,
) (*avro.Schema, error) {
key := manifestEntryProjectionCacheKey{
writerSchema: writerSchema.String(),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This one survived the rework, and it's the most worthwhile of what's left. We build the key with writerSchema.String() before the Get, so every lookup pays a full JSON serialization of the writer schema even on a hit. A scan opening a thousand manifests that share one schema does ~999 serializations of something that can run to tens of KB, and the same string is what we store as the key, so 256 entries can pin a few MB on a wide schema. For a PR whose whole point is cutting planning-time allocations, this quietly gives some of that back.

Could we key on a hash of the schema string, or the pointer identity of the *avro.Schema from reader.Schema(), and only touch String() on the miss path? wdyt?

Comment thread manifest_projection.go
}

root := writerSchema.Root()
projectedRoot := *root

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

projectedRoot := *root shallow-copies the node and we clone Fields, but Props and Aliases still alias the writer schema's, which is cached in ocf.Reader and can be shared across goroutines. CI is green so twmb/avro isn't mutating those in Schema() today, but it's an unstated assumption. A one-line comment noting we rely on Schema() being non-mutating would be enough. wdyt?

Comment thread manifest_projection.go
return true
case "value_counts", "null_value_counts", "nan_value_counts", "lower_bounds", "upper_bounds":
return includeColumnStats
default:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new version/delete-type test is good coverage for the fields that are kept, so this is softer than last round. The remaining edge: default: return false still silently zeroes block_size_in_bytes on the projected path while a full NewManifestReader carries its real value (a required long in v1). It's deprecated and unused for planning so it's harmless today, but the two readers returning different DataFiles for the same manifest is the kind of thing that bites a future field. A test cross-checking this whitelist against the avro-tagged fields on dataFile would make an omission fail loudly instead of vanishing. Non-blocking. wdyt?

@laskoviymishka laskoviymishka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM,

What's left is non-blocking polish, noted inline: the projection cache still calls writerSchema.String() on every lookup and stores it as the key, which quietly gives back some of the allocation win this PR is going for; the projected schema node shares Props / Aliases with the writer schema by reference; and the field whitelist still zeroes block_size_in_bytes on the projected path versus a full read. None of these block merge.

Approve with nits from me. Nice work on the equality-delete pruning fix.

Comment thread manifest_projection.go Outdated
projection ManifestEntryProjection,
) (*avro.Schema, error) {
key := manifestEntryProjectionCacheKey{
writerSchema: writerSchema.String(),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This one survived the rework, and it's the most worthwhile of what's left. We build the key with writerSchema.String() before the Get, so every lookup pays a full JSON serialization of the writer schema even on a hit. A scan opening a thousand manifests that share one schema does ~999 serializations of something that can run to tens of KB, and the same string is what we store as the key, so 256 entries can pin a few MB on a wide schema. For a PR whose whole point is cutting planning-time allocations, this quietly gives some of that back.

Could we key on a hash of the schema string, or the pointer identity of the *avro.Schema from reader.Schema(), and only touch String() on the miss path? wdyt?

Comment thread manifest_projection.go
}

root := writerSchema.Root()
projectedRoot := *root

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

projectedRoot := *root shallow-copies the node and we clone Fields, but Props and Aliases still alias the writer schema's, which is cached in ocf.Reader and can be shared across goroutines. CI is green so twmb/avro isn't mutating those in Schema() today, but it's an unstated assumption. A one-line comment noting we rely on Schema() being non-mutating would be enough. wdyt?

Comment thread manifest_projection.go
return true
case "value_counts", "null_value_counts", "nan_value_counts", "lower_bounds", "upper_bounds":
return includeColumnStats
default:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new version/delete-type test is good coverage for the fields that are kept, so this is softer than last round. The remaining edge: default: return false still silently zeroes block_size_in_bytes on the projected path while a full NewManifestReader carries its real value (a required long in v1). It's deprecated and unused for planning so it's harmless today, but the two readers returning different DataFiles for the same manifest is the kind of thing that bites a future field. A test cross-checking this whitelist against the avro-tagged fields on dataFile would make an omission fail loudly instead of vanishing. Non-blocking. wdyt?

Signed-off-by: Minh Vu <vuhoangminh97@gmail.com>

@zeroshade zeroshade left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The equality-delete regression I flagged last round is genuinely fixed — I verified it end to end rather than at the decision function. But the fix commit introduced a measured allocation regression on narrow MOR tables, and there's a dead function. Both are small.

First: the red macos-latest go1.25.9 check is not yours. It's catalog/rest → TestWaitForPlanBoundsSlowServerSideCancel, and git diff --name-only fc0e6fca 35db939d -- catalog/rest is empty. That test's most recent commit is literally 37ee0276 test(rest): fix flaky TestWaitForPlanBoundsSlowServerSideCancel (#1475). It passes 8/8 locally under -race, and macos go1.26.1 plus both ubuntu jobs are green on the same head. Timing flake.

Blocking

1. table/scanner.go:1123 — collectManifestEntriesWithSchemaOptions is dead code. Defined by this PR, called by nothing — not production, not tests, not benchmarks. It's invisible to CI because .golangci.yml enables staticcheck but not unused. planFilesLocal calls collectManifestEntriesWithSchemaMinSequenceNum directly at :1635. Delete it.

2. table/scanner.go:1382-1389 — the per-task stat strip is gated on projectScanColumns, not on whether stats were actually decoded, and that's a net allocation regression on narrow MOR tables.

DataFileWithoutColumnStats costs a measured 320 B, 1 alloc, ~750 ns per call (200,000×, count=6: already_statless 726–853 ns/320 B/1 alloc; with_stats 665–764 ns/320 B/1 alloc — the reflect-based cloneDataFileAvroFields dominates, not the nil-ing). Two problems:

  • No deletes + AlwaysTrue (the headline case): includeColumnStats=false, so the decoded *dataFile already has every stats map nil. The clone nils already-nil fields — pure waste, handing back 18.7% of the byte win (320 of 1,709 B/entry) and 23% of the time win (750 of 3,260 ns/entry) the projection just earned.
  • Any delete manifest present: retainDataFileStats forces includeColumnStats=true on every data manifest, so the read saving collapses to 190 B/entry (3,266 → 3,076, benchstat −5.84%, p=0.000). The 320 B clone alone exceeds it — net +130 B/entry before a single delete reference is touched. The win only returns at wide schemas (100 fields: 1,833 B/entry saved; 1,000 fields: 16,423 B/entry).

Ten-column MOR tables are a completely ordinary shape, and the benchmark never exercises the task-build path, so this cost is entirely unmeasured. Fix: return the per-manifest includeColumnStats from manifestProjectionForManifest and guard on that. When stats were never decoded there's nothing to strip.

3. table/scanner.go:1386-1388 + 1728-1739 — dataFilesWithoutColumnStats de-shares every delete-file reference. posDeleteIndex.forDataFile / eqDeleteIndex.forDataFile return fresh slices whose elements are shared DataFile pointers (positional_delete_index.go:119-124, equality_delete_index.go:550-556). Cloning per task destroys that sharing:

unprojected: 3 tasks, 3 eq-delete refs, 1 DISTINCT delete object
projected  : 3 tasks, 3 eq-delete refs, 3 DISTINCT delete objects

A partition-scoped delete covering 100,000 data files becomes 100,000 clones of one object — ~32 MB and ~75 ms of pure overhead, in exactly the case where the read-side win is 6%. Stripping the delete files is legitimate (the tasks are the only thing keeping those maps reachable once the indexes are nilled at :1667-1670), but it must happen once per unique delete file. A map[iceberg.DataFile]iceberg.DataFile memo threaded through dataFilesWithoutColumnStats preserves sharing in ~5 lines.

Major

1. table/scanner.go:436 + :1385 — double clone on the row-filter path. When dropColumnStats is true (data manifest, no delete manifests, non-trivial row filter — reachable, my probe hit it), streamManifest calls ManifestEntryWithoutColumnStats (which clones entry and DataFile), then the task closure clones the already-stripped DataFile again at :1385. That's 2 dataFile clones + 1 entry clone where one would do. Once the guard in Blocking 2 is in, one of the two becomes redundant.

2. manifest_projection.go:172-198 — DataFileWithoutColumnStats mutates its argument, and its doc says it returns a copy. d.initPartitionData() at :182 writes d.fieldIDToPartitionData and fires d.initPartition on the input. Safe today (every other reader funnels through the same sync.Once, and the clone's re-init allocates a fresh map — your new concurrency test covers it), but a function documented as returning a copy silently forcing eager lazy-init on its input is a trap. Say so in the doc, or hoist the call to the caller.

3. The fix for my prior finding is pinned only at the decision-function level. TestManifestProjectionRetainsDataFileStatsForDeleteScans unit-tests manifestProjectionForManifest in isolation. I mutation-tested it: reverting the fix leaves every existing end-to-end eq-delete test green — including TestEqualityDeleteReadPrunesNonOverlappingDeleteFiles — because they all call scan.PlanFiles(ctx), which never projects. So if the strip ordering rather than the decision function regressed, CI would not notice. Please add the permanent end-to-end equivalent: plan the same fixture through the projected path and assert len(task.EqualityDeleteFiles) == 1 per data file.

4. Five new exported symbols with a silent-data-loss footgun and no warning. ManifestEntryProjection, EntriesWithProjection, NewManifestReaderWithProjection, DataFileWithoutColumnStats, ManifestEntryWithoutColumnStats. Round-trip probe: feeding any of their outputs to ManifestWriter.Add/Existing silently drops six statistics fields. The current doc says "Callers that need the complete DataFile metadata should use ManifestFile.Entries", which is not the same as "never write a projected entry back into a manifest." For an Apache-governed public API that needs an explicit prohibition on all five. (Related: referencedDataFilePath at conflict_validation.go:983-996 returns "" on a stripped delete file and the delete becomes partition-scoped — degrades conservatively, so safe, but silent.)

What checks out

  • Field-by-field task parity. Every non-statistic FileScanTask field is byte-identical between PlanFiles and the projected path — path, content, format, record_count, file_size, spec_id, partition, split_offsets, key_metadata, sort_order_id, first_row_id, data_sequence_number, Start, Length, Residual, delete-list lengths — for both AlwaysTrue and a row filter. Only the 7 stats maps differ.
  • Nothing reads a projected-away field and gets a lie. Complete accessor audit against the whitelist: the only always-dropped fields are column_sizes and distinct_counts. Reading a dropped field yields an empty map, never a panic. Every consumer on the projected path checked individually.
  • Round-trip loss is real but unreachable in-repo. I traced every writer path — RewriteManifests, compaction, rewrite_data_files, conflict_validation all obtain entries via ManifestFile.Entries/ReadManifest/PlanFiles, all unprojected. Local plans aren't cached, so projected tasks can't leak into a later PlanFiles(). No in-repo silent data loss.
  • v1 fallback path unaffected — hand-built a v1 OCF with a plain-long snapshot_id to trigger isFallback; projected and full agree on everything but stats. Correct by construction, since isFallback/getFieldIDMap are computed from the writer schema.
  • Shared cached *avro.Schema is concurrency-safe — 64 goroutines, cold cache, alternating IncludeColumnStats, -race -count=5 clean. avro.SchemaField.Type is a value SchemaNode, so the clone can't write through, and avro.Resolve never mutates its arguments.
  • Both new tests have teeth (mutation-tested): removing key_metadata from the whitelist produces 4 distinct failures; reverting retainDataFileStats fails the projection test.
  • Benchmark table is accurate, reproduced within ~3% on x86/Linux. Interleaved benchstat, 10 stats fields: no-deletes -59.80% sec/op, -52.66% B/op, -85.42% allocs/op. The win is an allocation/materialization win, not an I/O win (Avro can't skip bytes on the wire) — the description makes no I/O claim, so no mismatch.
  • No existing test was weakened; three new test files, none modified.

Minor

  • manifest_projection.go:39,:163-164 — IncludeColumnStats: true does not include column_sizes (verified across v1/v2/v3). Deliberate and pinned, but the name is misleading for exported API; IncludeColumnMetrics would say what it means.
  • manifest_projection.go:128-130 — a data_file declared as a named-type reference rather than an inline record makes ToArrowRecords fail outright on a manifest PlanFiles reads fine. I could not construct a real-world manifest that triggers this, so I report it as unverified robustness only — falling back to an unprojected read would be strictly safer than erroring.
  • table/scanner.go:1667-1670 — four dead stores; Go's precise stack maps already make those collectable. If they do help on your platform, PlanFiles has identical structure and deserves the same treatment.
  • table/scanner.go:373/:1267 — openManifest and planDataManifestTasks are now test-only wrappers. Worth a one-line note.
  • manifest_projection.go:104-110 — @laskoviymishka's cache-key nit: the added comment is factually correct (twmb/avro@v1.8.0/schema.go:585 returns the stored s.full, no serialization), so the "~999 serializations" premise was wrong. The retention half stands: 256 wide JSON strings pinned in a process-global LRU, plus a whole-JSON hash per lookup. Keying on a fingerprint would be cheaper.
  • table/scanner.go:1207-1221 — dropColumnStats == true is reachable but has no decision-level test; the existing test only uses AlwaysTrue, so that assertion is always false.
  • manifest_projection_test.go:40-48 — the optionalFields allow-list is hand-maintained; adding a field to both it and the schema passes silently. Fine, but worth a comment saying what adding a name there means.
  • License headers present on all three new files; range-over-int clean; intrange enabled and passing.

Prior items

Mine (CHANGES_REQUESTED 2026-09-01):

  1. Dropped stats defeat #1960's equality-delete pruning → Fixed, verified end to end. Planned a real 2-data-file / 2-equality-delete fixture through scan.planFiles(ctx, true) — the exact path ToArrowRecords takes — and diffed against PlanFiles. Each data file gets exactly 1 of the 2 equality deletes on both paths, so pruning still fires, and the returned tasks correctly carry no stats. Mutation-confirmed load-bearing. The cost of the chosen fix is Blocking 2/3.
  2. Thin projection test matrix → Fixed. v1/v2/v3 data, v2 equality delete, v3 DV now covered. Still missing: no case runs the version matrix with IncludeColumnStats: true (the mode delete manifests actually use).

@laskoviymishka: cache key String() cost → partially fixed, premise was wrong (above). Props/Aliases aliasing → Fixed, and I confirmed the sharing can't bite. block_size_in_bytes zeroed → Fixed, pinned, and probe-confirmed on both a repo-written and a hand-built v1 fallback manifest. Whitelist cross-check test → Fixed, with teeth. Partition-copy race → Fixed, independently re-verified under -race -count=5.

Description

  • "Keeps metric maps only when row pruning or positional-delete matching needs them" is stale. The actual rule after 0a4f4e12 is: keep them whenever any delete manifest exists anywhere in the snapshot's manifest list, or the row filter is non-trivial. And the motivating consumer is equality-delete pruning (#1960), not positional. A reader of that bullet would not predict that a single DV manifest forces full stats decoding for every data manifest in the table.
  • Omits the delete-path cost — the benchmark covers only the manifest read, never the task-build path this PR also changed, which is where the 320 B/entry clone and the O(tasks × deletes) delete clones live.
  • Omits the five new exported symbols. For a release-governed module, new public API belongs in the description.
  • Omits that IncludeColumnStats: true still drops column_sizes, and the dead collectManifestEntriesWithSchemaOptions.

This review was drafted by an AI-assisted tool and confirmed by an Iceberg Go maintainer. The findings cite the project's review criteria; if you think one is mis-applied, please reply and a maintainer will weigh in.

@zeroshade zeroshade left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All three prior findings are genuinely fixed (verified by mutation and by reading avro@v1.8.0 source), and the projection provably does not change file selection, but the newly exported projected-read API silently drops column_sizes even when IncludeColumnStats is true and re-serializes that loss through ManifestWriter without error.

Re-review verification: 3 of 3 prior findings confirmed fixed at 1088c9c (each verified by mutating the fix and observing the suite go red, not by taking the claim on trust).

Verification performed
go build ./... (OK); go vet . ./table (clean); go vet ./catalog/... ./cmd/... (clean, exit 0); go test . -count=1 (ok 0.374s); go test ./table -count=1 (ok 5.453s); go test ./catalog/rest ./catalog/glue -count=1 (ok, disproving the pi-lens '[setup failed]' reports); go test . -run TestDataFileWithoutColumnStats -race (ok); go test ./table -run 'TestOpenManifestWithProjectionDropsStatsAfterFiltering|TestDataFilesWithoutColumnStatsReusesSharedFiles' -race (ok); 3 mutations (whitelist split_offsets removal; retainDataFileStats disjunct; ManifestContentDeletes disjunct) each restored via git checkout; 2 differential probes (40-file x 7-row-filter projected-vs-unprojected planning parity, and per-field DataFile comparison) plus 2 rewrite-hazard probes, all removed afterwards. Final `git status --porcelain` empty.

This review was drafted by an AI-assisted tool and confirmed by an Apache Iceberg Go maintainer. The findings below are observations, not blockers; an Apache Iceberg Go maintainer — a real person — will take the next look at the PR. If you think a finding is mis-applied, please reply on the PR and a maintainer will weigh in.

More on how Apache Iceberg Go handles maintainer review: CONTRIBUTING.md.

Comment thread manifest_projection.go
"github.com/twmb/avro"
)

// ManifestEntryProjection selects the optional data-file fields decoded while

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

major — IncludeColumnStats=true still drops column_sizes, and rewriting a projected entry silently loses it with no error

ManifestEntryProjection.IncludeColumnStats reads as an all-stats toggle, but manifestScanDataFileField (manifest_projection.go:157-170) returns false for column_sizes and distinct_counts unconditionally. The exported surface added by this PR -- EntriesWithProjection (manifest_projection.go:63), NewManifestReaderWithProjection (manifest.go:715), DataFileWithoutColumnStats, ManifestEntryWithoutColumnStats -- can therefore hand a lossy DataFile to a ManifestWriter, which re-serializes it via cloneDataFileAvroFields (manifest.go:1973) and persists the loss without any error or guard. The 'must not be passed to a ManifestWriter' contract is documentation-only and pinned by no test. This is NOT reachable on the shipped internal path (only ToArrowRecords calls planFiles(ctx,true), and compaction/transaction.go/row_delta.go all use unprojected PlanFiles), so it is not blocking -- but it is permanent public API. Suggest unexporting these helpers, or having the writer reject entries flagged as projected, or renaming the field to reflect that column_sizes is never included.

Evidence
Probe pr1972_probe_test.go (root pkg), run and then removed: 'unprojected read: valueCounts=map[1:10] colSizes=map[1:512]'; 'projected read (IncludeColumnStats=true): valueCounts=map[1:10] colSizes=map[]'; 'after rewrite of projected entry: colSizes=map[]'; 'CONFIRMED SILENT LOSS: column_sizes map[1:512] -> map[] with no error'. WriteManifest returned nil error on the projected entry.

Comment thread table/scanner.go Outdated
func (scan *Scan) manifestProjectionForManifest(
manifest iceberg.ManifestFile,
retainDataFileStats bool,
) (iceberg.ManifestEntryProjection, bool) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

minor — ManifestContentDeletes clause and manifestProjectionRetainsDataStats are tautological at every call site

In manifestProjectionForManifest the manifest.ManifestContent() == iceberg.ManifestContentDeletes disjunct can never decide the result. Call site 1 (collectManifestEntriesWithSchemaMinSequenceNum, scanner.go:1637) is only reached under if len(deleteManifests) > 0 with an all-Deletes list, so retainDataFileStats = manifestProjectionRetainsDataStats(deleteManifests) is already unconditionally true. Call site 2 (planDataManifestTasksWithOptions) only ever receives dataManifests, which splitManifestList guarantees contain no Deletes manifests. So manifestProjectionRetainsDataStats always returns true where it is used. TestManifestProjectionRetainsDataFileStatsForDeleteScans pins a mixed [dataManifest, deleteManifest] list that cannot occur in production, so it does not compensate. Consider collapsing the condition or documenting why the defensive clause is kept.

…anifest-columns

Signed-off-by: Minh Vu <vuhoangminh97@gmail.com>
Signed-off-by: Minh Vu <vuhoangminh97@gmail.com>

@zeroshade zeroshade left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both prior findings are genuinely fixed: the projected-write guard is pinned by four independent mutation tests that each go red, and a reflective differential probe proves a projected read diverges from a full read in ColumnSizes only.

Re-review verification: 2 of 2 prior findings confirmed fixed at 0faff65 (each verified by mutating the fix and observing the suite go red, not by taking the claim on trust).

Verification performed
go build ./... (OK); go vet ./... (clean); gofmt -l (only table/snapshot_producers.go, pre-existing at merge base 79125e9, inherited from upstream main); go test -race -timeout=900s . ./codec ./table -> all ok (2.5s / 1.7s / 24.7s). Mutation testing: 4 guard mechanisms individually removed, each turned a named test RED; a 5th mutation relocating the guard after w.writer.Encode produced an actual manifest-corruption failure. Throwaway reflective differential probe over every DataFile accessor on v1/v2/v3 confirmed ColumnSizes is the sole divergence; probe deleted, worktree restored to `git status --porcelain` empty.

This review was drafted by an AI-assisted tool and confirmed by an Apache Iceberg Go maintainer. The maintainer approving this PR has read the findings and signed off. If something feels off, please reply on the PR and a maintainer will follow up.

More on how Apache Iceberg Go handles maintainer review: CONTRIBUTING.md.

Comment thread manifest.go
tmp = tmp.(*fallbackManifestEntry).toEntry()
}
tmp.DataFile().(*dataFile).projected = c.projected
switch tmp.Status() {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit — Unchecked type assertion inconsistent with comma-ok form 28 lines below

tmp.DataFile().(*dataFile).projected = c.projected asserts unconditionally, while manifest.go:967 uses if df, ok := tmp.DataFile().(*dataFile); ok on the identical expression. Cannot panic today because tmp is always constructed locally as &manifestEntry{Data: &dataFile{}} (or the fallback wrapping the same), so this is a style/robustness inconsistency rather than a live defect.

Comment thread manifest.go
Content ManifestEntryContent `avro:"content"`
Content ManifestEntryContent `avro:"content"`
projected bool
Path string `avro:"file_path"`

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit — Non-wire projected field placed mid-struct inside the avro-tagged block

projected bool sits between Content and Path in the middle of the avro-tagged wire-shape struct. It is safe because avroFieldIndexes (data_file_codec.go:185) discovers indexes by tag lookup rather than position, so cloneDataFileAvroFields neither shifts fields nor launders the flag. Grouping non-wire state at the end of the struct would remove the positional-correspondence trap for future readers.

Comment thread table/scanner.go
manifestEntries, err := openManifest(fs, mf, partEval, metricsEval)
var projection *iceberg.ManifestEntryProjection
if projectScanColumns {
// Projected collection reads delete manifests before classification.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit — Hardcoded IncludePruningStats:true assumes a delete-only manifest list

collectManifestEntriesWithSchemaMinSequenceNum accepts an arbitrary manifestList but, when projectScanColumns is true, unconditionally reads pruning stats. Correct today (the only projected caller at scanner.go:1615 passes deleteManifests, and the comment says so), but a future caller passing data manifests would silently lose the projection benefit. Perf pessimization only, never a correctness bug.

@laskoviymishka
laskoviymishka merged commit 5b5d11d into apache:main Sep 8, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants