* fix(regions): serve invisible (Tr 3) OCR text layers instead of needs_ocr
Scanned pages (archive.org-style digitizations) carry their text as an
invisible render-mode-3 layer behind the page raster. extract_text_in_regions
extracted only visible items, so such pages yielded nothing but
'[Image: ...]' placeholders and every region fell back to OCR — while the
markdown path already includes the invisible layer for Mixed PDFs. The two
extractors disagreed about the same page.
Mirror the markdown path's gate, page-scoped: when the visible pass is
effectively textless (<40 non-placeholder alphanumerics), retry with
invisible text included and adopt the retry only when it contributes real,
non-garbage text. Pages with real visible text never retry, so double-layer
PDFs (visible text plus an invisible accessibility copy) are unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: strict zero-visible adoption gate, whole-layer garbage check, doc placement, discriminating tests
- Adoption now requires ZERO visible text on the page: the invisible pass
returns visible items too, so adopting alongside any visible text would
duplicate it. Strict gate instead of fuzzy dedupe; the dead
inv_alnum > visible_alnum condition goes with it.
- Garbage check judges the whole recovered layer, not the first 200 items.
- Constant/helper moved above the doc block so rustdoc stays attached to
extract_text_in_regions_mem.
- Tests: any-visible-text-blocks-adoption case (single short visible line)
with exactly-once assertions; mode-0 guard asserts occurrence count.
- Re-verified against real scanned-book pages: all recover (9.2-11.7K chars,
needs_ocr=false) — pure OCR-layer scans carry no visible text.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: gate adoption on visible-item presence, not alphanumeric mass
Punctuation-only visible text (zero alphanumerics) could still adopt the
invisible pass; its OCR twin in the layer would duplicate the glyphs. The
gate is now item-presence: any non-image item with non-whitespace text
blocks adoption (whitespace-only artifacts still tolerated). Pinned by a
punctuation-only fixture test; real scanned-book pages re-verified — all
recover unchanged (they carry no visible items at all).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: punctuation-gate test pins preservation, comment states fixture scope
The fixture's visible punctuation is a separate block, not mirrored in the
invisible layer — so it pins the gate, not the duplication scenario. The
doc comment now says so, and a positive assertion checks the punctuation
survives exactly once.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: retry only when invisible text was actually skipped; negative gate tests
- extract_page_text_items now reports skipped_invisible (4th return): the
Tr-3 suppression sites set it, so the region extractor retries only when
a recoverable layer exists. Blank pages and image-only scans without an
OCR layer — the common scanned case — no longer pay a second
content-stream parse.
- Negative tests: below-floor watermark layer and symbol-garbage layer are
both rejected (region keeps only the raster placeholder). needs_ocr
semantics for placeholder-only regions deliberately unchanged — that
contract predates this PR and downstream pipelines handle it.
- Real scanned-book pages re-verified: recovery unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: flag the ' show-text suppression path; require nonempty strings
- P1: invisible layers shown via the ' operator never set skipped_invisible,
so those pages kept the placeholder and fell back to GPU OCR. The '
suppression path now flags too — pinned by a fixture whose layer is shown
entirely via ' (nothing through Tj/TJ).
- P3: all three flag sites (Tj, TJ, ') require a nonempty string operand —
numeric-only TJ kerning arrays and empty shows no longer trigger the
invisible reparse.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>