111 Commits
Author SHA1 Message Date
Abimael Martell 841513d3fb fix(extract): pass page lines to per-page markdown conversion (#434)
CI / Test (push) Has been cancelled
CI / Format (push) Has been cancelled
CI / Clippy (push) Has been cancelled
CI / Build (macos-latest) (push) Has been cancelled
CI / Build (ubuntu-latest) (push) Has been cancelled
CI / OCR (macos-latest) (push) Has been cancelled
CI / OCR (ubuntu-latest) (push) Has been cancelled
CI / OCR (windows-latest) (push) Has been cancelled
CI / OCR runtime smoke (push) Has been cancelled
CI / WebAssembly (push) Has been cancelled
* fix(layout): keep relative-valley column detection on table pages

The page_has_table guard predates item claiming: by grouping time the
detected table's items have already left the flow, so a table cannot
fake a gutter in the histogram this guard protects, and the prose gate
inside relative-valley acceptance rejects any residual table-shaped
split. Without the fallback, the prose remainder of a table-bearing
two-column page fell to single-column Y-sorting and its columns
interleaved line by line.

* fix(layout): lower the relative-valley floor to sparse pages

The 100-item floor guarded against shallow histogram dips on sparse
pages, but OCR'd multi-column pages produce few long line-runs (a
two-column French academic page arrives as ~60 items) and were falling
to single-column Y-sorting, weaving their columns line by line. The
prose gate inside relative-valley acceptance is the real defense
against spurious dips; 30 items is enough for it to judge.

* fix(layout): re-guard the ungated XY-cut fallthrough on table pages

Removing page_has_table from the relative-valley condition also
unlocked the XY-cut fallback inside that block, which has no prose
gate — a table page whose valley candidate was just rejected could
take an unvalidated split. The guard is restored on that call
specifically; the relative-valley path keeps its prose-gated access.
Also rewrite the stale dense-page comment for the 30-item floor.
Bench and corpus unchanged (462/884, corpus byte-identical).

* fix(extract): pass page lines to per-page markdown conversion

The per-page extraction path (extract_pages_markdown_mem_impl, used by the
pages API and every --ocr auto run) partitioned each page's rects but
passed an empty slice for PDF lines to markdown conversion, while the
extracted all_lines sat unused. Every table only the line-based detector
finds (text-anchor rule tables, ruled grids) was silently dropped in that
mode, even though the whole-document path emitted it fine.

Partition all_lines per page like rects and pass them through. Regression
test: a rule-anchored table fixture must survive the pages API.
2026-08-19 09:45:21 -07:00
Abimael Martell 5eb6a13860 feat(layout): banded region segmentation for vertically-changing column layouts (#426)
* feat(layout): banded region segmentation for vertically-changing column layouts

Pages whose column structure changes down the page (newsletter bands,
figure-split flows, a three-column strip inside a two-column page) cannot
be represented by one full-height column set: the projection profile
either finds nothing and Y-interleaves the columns, or weaves the odd
band's text into the wrong buckets.

- Split pages into horizontal bands at full-width whitespace gaps (wide
  spanning items are excluded from occupancy — separators sit inside the
  very gaps being sought), detect columns independently per band, and
  re-merge consecutive bands with matching gutters across empty gaps so
  figure floats keep flowing down their columns while headline-separated
  bands stay independent.
- Engage only on contradicting evidence: a prose-validated band whose
  column count differs from the page-level structure. Pages the flat
  column model already explains keep their current ordering.
- Read short prose columns (5-14 lines, >=60% width fill on >=60% of
  lines in every column) as newspaper instead of Y-interleaving them as
  tabular. Kept as a standalone reading-order refinement so the table
  pipeline's is_newspaper_layout veto is unaffected.
- Split ordering entry points: order_multi_column_region keeps every
  page-level defense; order_validated_band (banded planner only) trusts
  validated bands, skipping line-count minimums and straggler splitting
  that would misfire on Y-cohesive bands.

* docs(layout): record why the band wide-item test is per-item

Assembling same-baseline fragments into runs before the wide test was
implemented and measured against the reading-order benchmark: word-gap
and gutter-gap distributions overlap in real documents, so assembled
runs fused narrow-guttered column pairs into page-wide lines, emptied
the occupancy, and disengaged banding on pages it rescues — a measured
regression with no measured win. Keep the per-item test (a fragmented
separator can suppress a cut, which only misses an engagement) and
document the boundary for future attempts.

* fix(layout): keep figure placeholders out of band whitespace probes

An image placeholder sitting between two matching column bands is the
very figure float whose flow-through the band merge exists for, yet it
read as content twice: its glyph box filled the occupancy gap (blocking
the cut) and the merge probe counted it as separator content (blocking
the merge). Both probes now see text layout items only.

* fix(layout): anchor band-merge matching on the founding band's columns

The merge comparison ran against the widened union, whose gutter is the
intersection of its constituents' gutters. Across a chain of
one-directionally drifting bands that intersection can walk past
GUTTER_TOLERANCE and reject a band identical to the run's own first
member. Compare candidates against the founding band's raw columns
instead: the run's column system is defined by its founder, so
drift can no longer accumulate in either direction. Union widening is
kept for item bucketing only.

* test(layout): differential coverage for the band-merge anchor rule

- banded_layout_rejects_creeping_drift: a band within tolerance of the
  moving union but 36pt from the founder must not join the run — the
  case the anchor rule exists for; the pre-anchor union admitted it.
- Reword the founder-anchor chain test as the invariant lock it is.
- Note at the merge site why a reject-overlapping-unions guard is
  unimplementable: detect_columns returns contiguous partitions whose
  adjacent regions share boundary coordinates, so the check degenerates
  to exact-equality matching and rejects every legitimate merge; the
  boundary-disagreement zone is bounded by GUTTER_TOLERANCE and split
  proportionally by greatest-overlap bucketing.
2026-08-18 11:00:34 -07:00
Abimael Martell 0f9b5fa1c6 feat(bindings): expose OCR in Node and Python (#405)
* feat(bindings): expose selective OCR

* fix(bindings): address review feedback
2026-08-17 09:25:56 -07:00
Abimael Martell 264a1c8372 refactor(vision): finalize OCR API (#403) 2026-08-16 23:55:56 -07:00
Abimael Martell 828a68c03b feat(vision): adaptively fuse native and OCR text (#394)
* feat(vision): adaptively fuse native and OCR text

* fix(vision): preserve adaptive OCR fallbacks
2026-08-16 23:55:55 -07:00
Abimael Martell 99069ce3d9 feat(vision): recover credible PDFium text layers (#393)
* feat(vision): recover credible PDFium text layers

* fix(vision): validate native recovery coverage
2026-08-16 23:55:55 -07:00
Abimael Martell 926720f8ff perf(vision): reuse OCR runtime sessions (#391)
* perf(vision): reuse OCR runtime sessions

* fix(vision): harden OCR engine caching
2026-08-16 23:55:54 -07:00
Abimael Martell 2cebb3c95f feat(vision): expose OCR pipeline (#360)
* feat(vision): expose OCR pipeline

* fix(vision): harden OCR API

* refactor(vision): expose OCR API
2026-08-16 23:55:53 -07:00
Abimael Martell 12d30b43b0 feat(vision): add OAR OCR engine (#357)
* feat(vision): add OAR OCR engine

* fix(vision): harden OAR runtime loading

* refactor(vision): use OCR engine terminology
2026-08-16 23:55:51 -07:00
Abimael Martell d9b83993df feat(render): add optional PDFium page rendering (#348)
* feat(render): add optional PDFium page rendering

* fix(render): honor PDFium row stride

* docs(render): use OCR terminology
2026-08-16 23:55:51 -07:00
Abimael Martell 2543abe371 feat(markdown): rejoin words hyphenated at line breaks (#388)
Justified print breaks words at syllables; after paragraph lines are joined
with spaces those breaks survive as "de- fendant" — thousands of them in a
long document — and, when an emphasis span was split with the word, as
"Bap-** **tist".

Whether the hyphen belongs in the word cannot be decided locally ("de-
fendant" is one word, "Third- Party" is a hyphenated compound), so the
document is used as its own dictionary. For each break:

  1. fragments appear joined elsewhere ("defendant")      -> join plain
  2. appear hyphenated elsewhere ("six-month"), or the
     continuation is capitalized ("Hinds- Radix")         -> keep the hyphen
  3. both fragments are words the document uses and the
     continuation has 4+ letters ("commercial- type")     -> keep the hyphen
  4. no evidence                                          -> leave untouched

The policy contains zero hard-coded words: vocabulary evidence,
capitalization, and two length invariants. The 4-letter floor keeps
suspended hyphens intact in any language ("mid- and long-term", "klein- und
mittelgroß", "kuva- tai video") because conjunctions are near-universally
1-3 letters. Fragments over 40 combined characters are fused reading-order
noise and are never joined. The vocabulary is collected after scrubbing the
break pairs themselves and excludes fenced code blocks; table rows and code
blocks are never rewritten. Split emphasis spans rejoin inside their
markers. Runs under the existing fix_hyphenation option (default on).

On a 1,370-page justified legal reporter this rejoins ~8,000 broken words
(98.8% of breaks; evidence-less ones stay visibly intact); word recall
against a reference extraction rises from 97.8% to 99.1%. No "six-month" ->
"sixmonth" class errors, and no fused-column corruption by construction:
no rule joins without evidence.

Regression-checked against a ~200-document corpus with semantic scoring
against an OCR baseline: zero regressions. Three in-repo fixture snapshots
regenerated with each diff inspected. 17 unit tests cover every rule, the
vocabulary scrubbing and code-block exclusion, the length gates, chained
breaks, mismatched emphasis markers, accented and Cyrillic words, German
and Finnish suspended hyphens, and the table/code skips.
2026-08-14 18:52:22 -07:00
Abimael MartellandClaude Fable 5 f65b25906c fix(regions): serve invisible (Tr 3) OCR text layers instead of needs_ocr (#351)
* fix(regions): serve invisible (Tr 3) OCR text layers instead of needs_ocr

Scanned pages (archive.org-style digitizations) carry their text as an
invisible render-mode-3 layer behind the page raster. extract_text_in_regions
extracted only visible items, so such pages yielded nothing but
'[Image: ...]' placeholders and every region fell back to OCR — while the
markdown path already includes the invisible layer for Mixed PDFs. The two
extractors disagreed about the same page.

Mirror the markdown path's gate, page-scoped: when the visible pass is
effectively textless (<40 non-placeholder alphanumerics), retry with
invisible text included and adopt the retry only when it contributes real,
non-garbage text. Pages with real visible text never retry, so double-layer
PDFs (visible text plus an invisible accessibility copy) are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: strict zero-visible adoption gate, whole-layer garbage check, doc placement, discriminating tests

- Adoption now requires ZERO visible text on the page: the invisible pass
  returns visible items too, so adopting alongside any visible text would
  duplicate it. Strict gate instead of fuzzy dedupe; the dead
  inv_alnum > visible_alnum condition goes with it.
- Garbage check judges the whole recovered layer, not the first 200 items.
- Constant/helper moved above the doc block so rustdoc stays attached to
  extract_text_in_regions_mem.
- Tests: any-visible-text-blocks-adoption case (single short visible line)
  with exactly-once assertions; mode-0 guard asserts occurrence count.
- Re-verified against real scanned-book pages: all recover (9.2-11.7K chars,
  needs_ocr=false) — pure OCR-layer scans carry no visible text.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: gate adoption on visible-item presence, not alphanumeric mass

Punctuation-only visible text (zero alphanumerics) could still adopt the
invisible pass; its OCR twin in the layer would duplicate the glyphs. The
gate is now item-presence: any non-image item with non-whitespace text
blocks adoption (whitespace-only artifacts still tolerated). Pinned by a
punctuation-only fixture test; real scanned-book pages re-verified — all
recover unchanged (they carry no visible items at all).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: punctuation-gate test pins preservation, comment states fixture scope

The fixture's visible punctuation is a separate block, not mirrored in the
invisible layer — so it pins the gate, not the duplication scenario. The
doc comment now says so, and a positive assertion checks the punctuation
survives exactly once.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: retry only when invisible text was actually skipped; negative gate tests

- extract_page_text_items now reports skipped_invisible (4th return): the
  Tr-3 suppression sites set it, so the region extractor retries only when
  a recoverable layer exists. Blank pages and image-only scans without an
  OCR layer — the common scanned case — no longer pay a second
  content-stream parse.
- Negative tests: below-floor watermark layer and symbol-garbage layer are
  both rejected (region keeps only the raster placeholder). needs_ocr
  semantics for placeholder-only regions deliberately unchanged — that
  contract predates this PR and downstream pipelines handle it.
- Real scanned-book pages re-verified: recovery unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: flag the ' show-text suppression path; require nonempty strings

- P1: invisible layers shown via the ' operator never set skipped_invisible,
  so those pages kept the placeholder and fell back to GPU OCR. The '
  suppression path now flags too — pinned by a fixture whose layer is shown
  entirely via ' (nothing through Tj/TJ).
- P3: all three flag sites (Tj, TJ, ') require a nonempty string operand —
  numeric-only TJ kerning arrays and empty shows no longer trigger the
  invisible reparse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 14:11:31 -07:00
Abimael MartellandClaude Fable 5 a67ee03269 feat(bindings): expose TextItem.mcid and structure-tree element extraction (#346)
Tagged PDFs carry a structure tree with real heading roles (H1..H6), and
the core already parses it (structure_tree::StructTree) and threads MCIDs
onto TextItem — but neither surfaced through the bindings.

- Expose TextItem.mcid (Option<i64>) through the napi and pyo3 bindings,
  matching the core field added with the marked-content extractor.
- Add StructRole::name(), the inverse of from_name, so roles have a
  stable string form.
- Add extract_structure_elements / extract_structure_elements_mem to the
  core: one (page, mcid, role) entry per marked-content reference, sorted
  by (page, mcid), empty for untagged PDFs. Pages are 1-indexed to match
  TextItem.page, so results join directly against
  extract_text_with_positions output.
- Bind it as extractStructureElements (napi) and
  extract_structure_elements / extract_structure_elements_bytes (pyo3),
  with type-stub updates in pdf_inspector.pyi.
- Cover the join in Rust integration tests, napi test.mjs, and pytest,
  using the existing firecrawl_docs_tagged.pdf fixture (tagged) and
  thermo-freon12.pdf (untagged).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 20:03:07 -07:00
585d36e6a6 fix: recover from a corrupted startxref pointer (#230)
* fix: recover from a corrupted startxref pointer

Fixes #228.

A PDF whose startxref pointer has been corrupted to point at the wrong
byte offset — a single flipped digit, which is what damaged writers
emit in the wild — was entirely unprocessable: every entry point
(classify_pdf, extract_pages_markdown, process_pdf) raised "Invalid
PDF structure", even though the file's object data, real xref table,
and trailer were all completely intact just past the wrong pointer.
Both pypdf and pdfium recover from this by locating the real table
directly instead of trusting the pointer; lopdf doesn't.

Added a new repair candidate (alongside the existing
missing-%%EOF-marker and stripped-leading-bytes repairs in
repair_pdf_container_candidates): scan the buffer for the real,
standalone `xref` keyword and append a corrected trailing
`startxref`/`%%EOF` block. lopdf's own get_xref_start always reads the
*last* `%%EOF` in the final 512 bytes of the buffer and the
`startxref` value immediately before it, so the appended block
transparently supersedes the corrupted one already in the file — no
in-place byte surgery on content the original writer produced.

Scoped to classic (non-stream) xref tables, matching the reported
repro and the common case; a corrupted pointer into a cross-reference
*stream* (`N 0 obj << /Type /XRef ...>>`, some PDF 1.5+ writers) would
need the containing object's number, not just a byte offset — out of
scope here.

Verified against the issue's exact repro (a valid one-page PDF with a
single corrupted byte in its startxref offset): before this fix,
process_pdf/classify_pdf/extract_pages_markdown all raised "Invalid
PDF structure"; after, both the page count and the real extracted text
("Order Detail Report by Account", "WIDGET ASSEMBLY", the dollar
amount) come back correctly. New regression test added.

Full suite (859 tests, 1 new) passes; cargo clippy --all-targets
-- -D warnings unchanged at 28 pre-existing/unrelated errors.

* fix: validate xref table shape and scan in a single reverse pass

Addresses cubic-dev-ai's review of #230.

- P2 (correctness/safety): the recovery candidate trusted the last
  standalone "xref" token unconditionally, without confirming it's
  actually a cross-reference table. A coincidental "xref" substring
  inside unrelated content — a stream, a string, uncompressed
  metadata — could get "repaired" against a bogus offset, letting
  lopdf load successfully against garbage instead of returning a
  clean error: a real failure turned into silent data corruption on
  the fallback path. Added looks_like_xref_subsection_header, which
  confirms a plausible classic xref subsection header (`<start-id>
  <count>`, e.g. "0 6" — the shape every real classic table starts
  with) actually follows the candidate token before accepting it.
  find_last_valid_xref_table_start now walks backward from the end of
  the buffer until it finds a token that both stands alone *and*
  validates, rather than accepting the first (rightmost) standalone
  match unconditionally.

- P2 (performance): the old scan re-invoked
  `buf[..search_end].windows(4).rposition(...)` on a shrinking prefix
  every time a candidate token failed the boundary check, which is
  quadratic on a pathological buffer with many non-standalone "xref"
  occurrences. Rewrote as a single reverse byte-index walk — O(n)
  regardless of how many false candidates it has to reject along the
  way.

Added direct unit tests on the byte-level scan (more precise than
constructing adversarial full PDFs, and the coincidental-match
scenario can't be represented in an integration-test fixture anyway
since reportlab compresses page content by default): a coincidental
standalone "xref" with no subsection header is rejected; a real
classic table is found; a coincidental match positioned *after* the
real table in the buffer doesn't shadow it; "xref" as a substring of
"startxref" still doesn't match. The original #228 repro (corrupted
startxref pointer, real table otherwise intact) is unaffected —
verified manually in addition to the existing integration test.

Full suite (863 tests, 5 new) passes; cargo clippy --all-targets
-- -D warnings unchanged at 28 pre-existing/unrelated errors.

* fix: reject xref subsection count runs with trailing garbage

looks_like_xref_subsection_header validated that a count run of digits
followed the whitespace separator, but never checked what came after
it. A coincidental "xref\n0 6garbage" in stream/literal content would
still validate as a real subsection header shape and get repaired
against a bogus offset.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Abimael Martell <1450169+abimaelmartell@users.noreply.github.com>
2026-08-05 16:43:11 -07:00
371de80b14 fix: extract_pages_markdown's needs_ocr now agrees with classify_pdf (#231)
* fix: extract_pages_markdown's needs_ocr now agrees with classify_pdf

Fixes #227.

extract_pages_markdown_mem computed its per-page needs_ocr entirely
from text-quality signals: decoding/garble issues, empty markdown, GID
fonts, garbage-text ratio. It had no awareness of the page's image
content at all — so a page that is fundamentally a full-page scan
with a little genuine native text drawn over it (a header, a stamp, a
cover-sheet annotation) extracts that text cleanly, trips none of the
text-quality checks, and reports needs_ocr=false — while
classify_pdf/detect_pdf_type correctly see the dominant background
image and flag the same page as needing OCR. Two public APIs
answering the same question, silently disagreeing, in the unsafe
direction (skipping OCR on a page that needs it).

Exposed detector::analyze_page_images at crate visibility (was
private) and call it per page in extract_pages_markdown_mem's loop —
the same "large background image" signal (>50% page coverage) that
already powers has_template_image in classify_pdf/detect_pdf_type,
rather than reimplementing image-area detection a second time with
its own thresholds that could drift out of sync again. When it's
true, the page is flagged needs_ocr (with OCR_REASON_SCANNED added
to ocr_reasons_by_page, matching how the same signal is already
reported elsewhere) and its markdown is blanked, exactly like the
existing text-quality-triggered needs_ocr paths already do — no
special-casing added for "cleanly-extracted-but-still-a-scan" text.

Verified against the issue's exact repro (a full-page raster with one
native text line drawn over it, built via reportlab/pillow): before
this fix, extract_pages_markdown_bytes reported page 0
needs_ocr=False with the header line as markdown while
classify_pdf_bytes correctly flagged pages_needing_ocr=[0]; after,
both agree needs_ocr=True and the page's markdown is empty. Confirmed
no regression on a normal text-based fixture (nexo-price-en.pdf:
needs_ocr stays False, full markdown returned). New Rust regression
test added exercising both APIs against the same fixture.

Full suite (860 tests, 1 new) passes; cargo clippy --all-targets
-- -D warnings unchanged at 28 pre-existing/unrelated errors.

* fix: gate has_template_image behind the same OCR signals classify_pdf uses

extract_pages_markdown_mem was treating has_template_image alone as
sufficient to force needs_ocr=true and discard the page's markdown, but
classify_pdf/detect_pdf_type never treats that raw signal alone as
needing OCR. A text page with a full-bleed watermark, letterhead, or
large figure would get its clean markdown wrongly blanked and routed
to OCR.

Added page_template_image_needs_ocr(), mirroring the two distinct
signals classify_pdf actually uses to decide a template-image page
needs OCR: the looks_like_scan gate (image_count <= 1, few text ops,
low alphanumeric diversity) used for Mixed-type routing, and the
insufficient-text-volume signal (text_operator_count < 10) that routes
a page with a dominant background image and only a couple of native
text calls to PdfType::ImageBased independent of looks_like_scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: match per-page OCR threshold and add missing vector-text signal

Two follow-up findings on the has_template_image gate added in the
previous commit:

1. insufficient_text used a hard-coded threshold of 10 text operators,
   but Mixed-type per-page routing (the actual per-page decision this
   function tries to agree with) uses config.min_text_ops_per_page
   (default 3). The higher 10 threshold was borrowed from a *different*
   classify_pdf code path — the effective_min_ops floor used only for
   whole-document ImageBased/Scanned classification, a cross-page
   aggregate this per-page function can't replicate anyway. Using the
   lower per-page threshold removes a real disagreement window
   (3-9 text ops with high alphanumeric diversity) without breaking
   the #227 regression fixture (text_ops=1, still well under 3).

2. extract_pages_markdown_mem never checked has_vector_text at all,
   even though Mixed-type per-page routing always sends
   vector-outlined-text pages to OCR (outlined glyphs can't be
   extracted as text). A page with massive path ops plus a short
   genuine caption could extract that caption cleanly, slipping past
   the existing empty/garbage-text checks. Added
   page_has_vector_text() and wired it into needs_ocr the same way
   has_template_image is.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* perf: compute template-image and vector-text OCR signals in one pass

page_template_image_needs_ocr and page_has_vector_text each called
analyze_page_content independently, so every requested page's content
streams (page + XObjects) and image coverage were decompressed and
scanned twice per page with one result discarded each time.
detect_from_document avoids this by caching its per-page PageAnalysis;
extract_pages_markdown_mem had no such cache.

Merged both into page_ocr_signals(), a single analyze_page_content
call returning both signals as a tuple.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Abimael Martell <1450169+abimaelmartell@users.noreply.github.com>
2026-08-05 15:41:44 -07:00
Abimael Martell 12e9a655e3 fix(extractor): supply built-in metrics for non-embedded base-14 fonts (#241)
* fix(extractor): supply built-in metrics for non-embedded base-14 fonts

PDFs may legally omit /Widths for non-embedded standard fonts (Times,
Helvetica, Courier, Symbol, ZapfDingbats) — the spec requires the reader
to supply the metrics. We returned None, so every glyph advanced 0 and
each text item got width 0, silently breaking every gap-based heuristic
downstream: space synthesis, sub/superscript detection, table column
detection, heading merging.

- src/extractor/base14.rs: Adobe Core-14 AFM width tables keyed by
  Unicode char, plus the standard Symbol/ZapfDingbats encoding vectors
  (their glyphs sit at byte positions unrelated to Latin text, so widths
  must resolve through the built-in encoding, not cp1252)
- Width resolution order: Differences -> built-in encoding -> the same
  cp1252-style fallback the text decoder uses, so a code's advance always
  matches the character we emit for it
- Type3 visual sizing: PK bitmap fonts (dvips) use FontMatrix
  [1 0 0 -1 0 0] with nominal sizes like 0.12pt; scale by FontBBox height
  x |matrix_y|. Applied in the page-stream and Form XObject paths.
  Indirect numeric array elements are resolved before use.

Effect on Shannon's 'A Mathematical Theory of Communication' (1998
dvips/Distiller, the reported case): glued sentences 95 -> 5. Corpus
impact: 12 of 184 eval documents, e.g. Data-Processing-Agreement
recovers a paragraph that a phantom table had shredded into cells.

Layout heuristics tuned on the same document (indent-based paragraph
breaks, heading reclassification, table script filtering) are held back
for a separate PR — they change ~98 further documents and need to be
justified against the corpus, not against one PDF.

* review: narrow Type3 rescaling to self-inconsistent fonts; dedup + test all width tables

Addresses cubic review on #241, plus a follow-up from a local cubic run.

- Type3 visual scaling was applied to every Type3 font whose FontBBox
  height x |matrix_y| deviated >5% from 1.0. FontBBox is the glyph box,
  not the em box, so a conventional 1/1000-matrix font with a
  descender..ascender bbox (~700 units) computed 0.7 and had every
  reported size shrunk by 30% — corrupting the drop-cap, heading-tier,
  sub/superscript and table heuristics this is meant to fix.

  First attempt gated on the matrix being unit-scale, but a local cubic
  run pointed out that wrongly excludes valid non-standard matrices (a
  0.005 matrix with a full-em bbox legitimately needs a 5x scale). The
  product is the right discriminator, not the matrix: a self-consistent
  font lands near 1.0 because the matrix is the reciprocal of the
  glyph-space em, so only a wildly inconsistent one (dvips/PK bitmap
  fonts sit at ~159) is renormalized. Band widened to [0.25, 4.0].

  Corpus effect: 12 -> 7 documents change. The 5 that drop out were
  being wrongly rescaled — including Data-Processing-Agreement, whose
  phantom-table fix turned out to come from this bug rather than from
  the width fallback, so it is correctly given up.

- base14: all 14 width tables now covered by the sort-invariant test via
  an ALL_TABLES registry, not a hand-picked subset.
- base14: identical tables share one static (all four Courier variants
  are monospace 600; the oblique Helvetica variants match their upright
  forms), removing 5 duplicate copies.

* test: refresh Shannon snapshot after merging main

CI checks out a merge of the PR head with main, and main advanced 8
commits since this branch was cut — including #201 (contextual digit
runs), #240 and #253 (markdown fixes). Those change extraction output,
so a snapshot generated on the unmerged branch could not match; the
Test job failed on the merge commit while passing on the branch itself.

The merged behaviour is better: the footnote marker '2' before
'Hartley, R. V. L.' is now recovered instead of dropped.

950 tests pass on the merged tree, clippy clean.
2026-08-04 18:00:43 -07:00
Abimael Martell 3fb545284b fix(layout): preserve contextual digit runs (#201)
* fix(layout): preserve contextual digit runs

* fix(layout): harden page folio filtering

* fix(markdown): distinguish folios from contextual numbers

* fix(layout): distinguish running folios from contextual digits

* fix(layout): tighten running folio evidence

* test(layout): guard the folio evidence floor

* fix(layout): preserve page-number decisions across partitions

* fix(layout): harden document-level folio filtering

* fix(markdown): carry folio context across public APIs

* fix(layout): isolate folios from layout metadata

* fix(markdown): preserve table cells in per-page extraction

* fix(layout): preserve folio context with page filters

* fix(layout): handle contextual folio sequences

* fix(layout): tighten adjacent folio evidence

* fix(layout): constrain folio context inference

* fix(forms): resolve widget pages from annotations
2026-08-03 16:35:52 -07:00
Abimael Martell b084769fda feat(markdown): add fidelity output profile (#168) 2026-07-15 17:18:16 -07:00
Abimael MartellandClaude Fable 5 f741e49dec fix(headings): all-bold single words qualify as headings when standalone (#167)
Single-word bold section headings ('Replace', 'Trash', 'Instructions')
required a paragraph break before AND after, but headings hug their
section's first paragraph — the break-after almost never exists. A
standalone all-bold single word (>=4 chars, paragraph break before or
page top) now classifies; mixed bold lead-ins ('Note: ...') stay
excluded via all_bold.

opendataloader-bench: 0.8567 -> 0.8575, MHS 0.773 -> 0.776; docs 145
+0.118, 069 +0.112 (net of one cover-page layout shuffle at -0.066
where the new output is semantically closer to GT). pdf-evals: 66
snapshots, composite 0.5864 -> 0.5883, sole >0.02 mover positive.
p1244/thermo fixture snapshots regenerated ('Instructions' un-fuses
from its body paragraph — the intended behavior).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 14:44:46 -07:00
Abimael MartellandClaude Fable 5 0f898a18fb fix(headings): digit-only lines do not define heading tiers (#166)
* fix(headings): digit-only lines do not define heading tiers

A large bold page number (14pt folio over 11pt body) claimed tier 0:
every real heading demoted one level document-wide, and the bold-size
fallback (which requires an empty tier list) was blocked for documents
whose headings match body size.

Bench-neutral (MHS scores relative hierarchy); pdf-evals: 18 docs get
their heading levels back (#### -> ###), semantic composite +0.0006,
no percentile down.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(headings): exclude digit-only lines from the bold fallback tier pass too

The exclusion in the main pass wasn't enough: with the page-number
tier gone, the bold fallback re-collected the same bold folio. Also
regenerates the thermo-freon12 snapshot (cosmetic churn on the
scrambled legend fixture).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:42:41 -07:00
Abimael MartellandClaude Fable 5 4aa2c0c208 fix(headings): rescue wrapped bold headings on interleaved column pages (#164)
* fix(headings): rescue wrapped bold headings on interleaved column pages

Report pages whose columns can't be detected (6pt gutters) interleave
both columns' lines, which breaks every whitespace signal the bold
heading heuristic relies on: para_threshold inflates to ~3x line
height and a wrapped heading's own internal line gap defeats isolation.
'9.5. Adapting to the New Normal: Changing / Business Models' merged
into the following paragraph.

Four changes:
- merge_wrapped_bold_heading_groups: 2-3 consecutive all-bold
  body-size lines merge into one line when the group is isolated
  (column-locally, judged by x-overlapping lines only) or starts with
  a section number.
- Section-numbered all-bold lines ('9.5. ...') classify as headings
  without the standalone/isolation score gate.
- Line unfusing extends to uppercase-start continuations, gated on a
  bold-style mismatch between the runs (a bold heading beside regular
  body text) — same-style label rows stay joined.
- The unfuse line-side wordiness requirement drops to 2 words so a
  wrapped heading's short last line ('Business Models') still splits
  from the neighboring column.

opendataloader-bench: overall 0.8554 -> 0.8576, MHS 0.769 -> 0.777;
docs 037 +0.161, 111 +0.157, 039 +0.091, 198 +0.028, none down.
pdf-evals: 63 snapshots, composite 0.5952 -> 0.5964, sole >0.02 mover
positive. thermo-freon12 snapshot regenerated (cosmetic churn on an
already-scrambled 3-column legend).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(headings): review follow-ups — multi-component section numbers, wholly-bold line gate

Single '1. ' prefixes are ordered list items and no longer bypass
isolation; the uppercase unfuse requires the whole line bold (a
heading), not merely its last run, so mixed bold-label/value rows
stay joined.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:00:54 -07:00
Abimael MartellandClaude Fable 5 39c31a8404 fix(underline): rescue snug-owned underlines from the table-ruling filters (#143)
* fix(underline): rescue snug-owned underlines from the table-ruling filters

Documents that underline many full-width lines (dense CJK business docs,
legal redlines, 10-K section links) produce span-similar rules at 3+
y-levels — exactly what the repeated-ruling filter treats as table
rulings, so every semantic underline on such pages was discarded. Three
changes fix detection without re-marking real tables:

1. Snug-owner rescue: a rule survives the repeated-ruling filter when
   the union of touching text runs on its baseline row owns it (rule
   contained within the union's span +0.75em, runs cover >=60% of the
   rule, no column-sized gaps between runs). Table row separators fail
   ownership: they overshoot their cells' text or match gapped items.
   Same-row segmented rules (column-header separators) always stay
   discarded, and a rule enclosed by a drawn cell-sized box (rect-grid
   tables) is never rescued.

2. Vertical window widened 0.35em -> 0.72em below the baseline: CJK
   layouts draw underlines under the full em box, measured at ~0.67em.

3. Prose-table guard in the positions-path suppressor: a detected
   'table' whose cells hold flowing prose (>=30% of cells over 100
   chars) is a detection artifact of boxed callouts + stacked rules,
   not a real table — suppressing there erased every underline on the
   page.

Also fixes cluster_x_positions fabricating phantom table columns from
style-split continuation runs (touching items, gap <2pt, now feed one
column start) — the fix that keeps rect-grid table shapes stable while
underlined links inside cells are correctly marked.

Snapshot updates are underline gains on regulation/form fixtures and one
empty spacer-column change in a subscripted header.

Corpus (508-doc public bench sweep): text output byte-identical on all
docs; underlined items +224/-0; strikeout now fires on redline docs.
Item-level GT coverage: is_underline 86->151/405, is_strikeout 0->10/44.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(underline): fraction-bar guard + subscript merge across underline marks

Found by a 202-doc real-world corpus diff (pdf-evals) that exercises the
full markdown pipeline, which the bench-corpus item sweep does not:

1. Math fraction bars and lattice grid lines are underline geometry —
   short horizontal rules under digits. Guard: a narrow rule (<=60pt)
   with bar-sized text hanging just below it (denominator) never marks.
   The below-text width bound matters: tightly-leaded REAL underlines
   have a full-width next line below, which must not trip the guard.

2. merge_subscript_items refused to merge when the parent was underlined
   but the tiny digit was not (the drawn rule easily misses the digit's
   own overlap window) — losing the merge broke subscript tokens inside
   table cells (b+2 no longer became b₂). Strikeout boundaries still
   block the merge in both directions; only parent-underlined/digit-bare
   merges, absorbing with the parent's flags.

Corpus after refinement: underlined items +220/-2 (the 2 are fraction
bars the old code wrongly marked), GT rule-text coverage 149/405
underline + 10/44 strikeout, text output identical on all 508 bench
docs and word-count-identical on the 202 pdf-evals docs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(underline): address review — grid-evidence veto, strikeout-safe fraction guard, bounded gaps

- Cell-box veto now requires GRID EVIDENCE (a vertically abutting
  neighbor rect with x-overlap) instead of a height window: multiline
  table cells taller than the old 90pt ceiling veto again, and isolated
  filled callout panels (which legitimately contain underlines) no
  longer veto at all.
- The fraction guard gates only UNDERLINE marking; rule_strikes_item
  still evaluates, so short strikeouts near lower text survive.
- Fraction hug distance tightened to 0.3em so a short last-line at
  normal leading is not mistaken for a denominator.
- Continuation-run suppression bounds the negative gap (-4pt): text
  overhanging from an adjacent cell keeps its own column start.

Corpus after review fixes: underlined items +222/-2, GT coverage
150/405 underline + 10/44 strikeout, bench text output identical on
all 508 docs, pdf-evals word loss bounded at equation-reflow noise.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* chore: appease clippy (redundant closure)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 08:47:45 -07:00
Abimael MartellandClaude Fable 5 d8eb33e390 feat: per-page OCR routing reasons (#139)
* feat: per-page OCR routing reasons (scanned/no_text/vector_text/garbled)

Replaces the single suspected_garbled_text signal with a per-page
explanation for why each OCR-flagged page needs OCR. The detector
classifies each page in pages_needing_ocr from its content analysis:

- scanned            — no usable text, image-backed page
- no_text            — no text and no image (blank/unreachable)
- vector_text        — text drawn as vector outlines, not extractable
- suspected_garbled_text — undecodable Identity-H/Type3 fonts

Exposed on PdfTypeResult.ocr_reasons_by_page and surfaced through
PdfProcessResult and the detect-pdf CLI (JSON + human output). Reasons
only ever explain pages already flagged for OCR — a text page with an
embedded logo stays TextBased, so this doesn't widen the OCR net.
Markdown output is byte-identical across the regression corpus.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: cache freshly-analyzed pages so OCR reasons aren't lost

Under a sampling ScanStrategy, the Mixed per-page loop (Phase 2) and the
garbled-font check (Phase 3) analyze non-sampled pages but dropped the
PageAnalysis after flagging them. The reason-classification pass then
missed the cache and defaulted those pages to "scanned", masking the
real vector_text / suspected_garbled_text cause. Insert the fresh
analyses into analysis_cache so the reason pass classifies them
correctly. No change under the default full-sampling strategy (all
pages are already cached).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 13:23:37 -07:00
Abimael Martell a38efcf142 feat: --password support for encrypted PDFs (#138) 2026-07-11 11:55:36 -07:00
Abimael MartellandClaude Fable 5 dd8aba9e20 fix(markdown): keep isolated headings on sparse pages; gate tagged roles (#132)
* fix(markdown): keep isolated headings on sparse pages; gate tagged roles

The isolated-line density guard wiped every isolated line on a page
where they exceeded 25% of lines. On sparse pages (covers, ToC pages
with a lone "CONTENTS" title, section-divider pages) a single heading
is trivially >25%, so the guard erased exactly the line it exists to
find. Require the page to have >=10 lines before the guard runs — the
25% ratio only signals a multi-column misfire on a dense page.

That let more isolated lines through, exposing that the visual heading
heuristic could promote lines already tagged with a non-heading struct
role (list item, blockquote, code, caption, ToC) or set in a monospace
font. Gate the heuristic on those in both converter paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: extend non-heading role gate; centralize on StructRole method

Move the non-heading-role check to StructRole::is_non_heading_content
and extend it to the content roles the inline allowlist missed: Quote,
Index, Note, Reference, BibEntry, Formula, Form (in addition to the
existing list/quote/caption/toc/code roles).

Figure is deliberately excluded: cover and banner pages routinely tag
the document title inside a Figure next to a seal/logo, and that title
is a real heading — including Figure demoted the LA County protocol
cover title from headings to bold. Verified against the reference.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: block table roles from heading promotion too

Add Table/TR/TH/TD/THead/TBody/TFoot to is_non_heading_content. When
table reconstruction falls back and cells reach the line loop as plain
text, a short isolated cell (a TH column header in particular) could be
promoted to a heading. Defensive: no change across either regression
corpus, so pure hardening for the fallback path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 22:25:42 -07:00
Abimael MartellandClaude Fable 5 3a389f079e fix(markdown): ToC-page suppression, wrapped bold headings, math fragments (#131)
* fix(markdown): ToC-page suppression, wrapped bold headings, math fragments

Three heading-classification improvements:

- After emitting a "Contents"/"Table of Contents" heading, suppress
  heading promotion for the rest of that page: ToC entries are section
  titles that look exactly like headings ("1. Overview of OCR Pack")
  and whole contents pages came out as stacks of ##.
- merge_heading_lines only merged font-size-tier and struct-tree
  headings, so bold-at-body-size headings that wrap emitted two
  separate ## lines. Merge a fully-bold line into the previous
  fully-bold line when it reads as a wrap continuation (starts
  lowercase, tiny Y gap, no terminal punctuation on the previous line).
- Reject display-math fragments from the bold/rarity heading heuristic:
  equations ending in an equation number ("S = kB ln W, (2)") and
  lead-ins referencing one ("Rearranging Equation (8) gives:"). A bare
  trailing colon is deliberately NOT a signal — real headings often end
  with colons ("Procedure:").

The p1244 snapshot change is the bold-merge working as intended:
stacked form labels "**Subtotals** **from pages**" now read
"**Subtotals from pages**".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: guard tier path, protect prev headings in merge, narrow (N) rule

All three review findings applied, calibrated against the corpora:

- is_heading_fragment now gates the font-size-tier path too, not just
  the rarity heuristic.
- The bold wrap-merge requires the previous line to be tier-less as
  suggested; corpus diff confirmed the old behavior was absorbing a
  wrapped list-item fragment into a real heading.
- The bare "(N)" suffix rule suppressed real headings ("Nicaea (325)",
  appendix numbering). It now requires math evidence: an operator
  (=, <=, <<, ...) in the line or ,/: immediately before the number.
  Page-of-total running headers ("PM 2 (10)") get an explicit rule
  since the old blanket suffix check had been catching them only by
  accident.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 21:40:18 -07:00
Abimael MartellandClaude Fable 5 57335f8bcf feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection (#125)
* feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection

Two style-recall gaps, both invisible to the existing name-based
heuristics:

1. Subset fonts with opaque BaseFont names ("Tc1", "AAAAAB+Amplitude")
   defeat is_italic_font/is_bold_font. New descriptor_style_flags reads
   the FontDescriptor (ItalicAngle beyond 4 degrees, Flags bit 7 Italic,
   bit 19 ForceBold) and, when the descriptor claims upright, falls back
   to the embedded font file: ttf-parser's OS/2 fsSelection + post
   italicAngle for sfnt fonts, and the CFF Name INDEX PostScript name
   for bare-CFF FontFile3 (descriptor rewritten to ItalicAngle 0 while
   embedding "Amplitude-LightItalic" was observed in the wild).
   ORed into is_bold/is_italic at item creation (content streams and
   form XObjects).

2. No strikeout signal existed. New is_strikeout on TextItem, detected
   in the same pass as underline: same rules pipeline (stroked lines /
   thin filled rects, table-ruling suppression), different vertical
   window — a rule crossing the glyphs at 12-55% of the em above the
   baseline instead of sitting at it. Exposed through napi and python
   bindings and pdf2md --items-json.

Verified on public ParseBench corpus docs: previously-missed italic
council titles and bold CJK itinerary headings now flagged (render-
checked); 24/508 docs gain flags, none lose any; 35 strikeout items
detected corpus-wide, disjoint from underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): quote-op advance width, Ts text rise, doc-level font style cache (PR #125 review)

Address three valid findings from review:

- The ' (move-to-next-line-and-show-text) operator emitted zero-width
  items and never advanced the text matrix, so geometric underline/
  strikeout detection (which requires width > 0) could never mark its
  text, and following show ops overlapped it. Reuse Tj's advance-width
  computation and matrix advance.

- Ts (text rise) was dropped entirely: raised/lowered runs kept the
  unshifted baseline, so rules drawn at the risen glyph position missed
  the strike/underline windows. Track rise in the text state (saved and
  restored with q/Q) and shift the rendering position through the text
  matrix's y column; advances stay on the unshifted matrix per spec.

- descriptor_style_flags re-decompressed and re-parsed the same embedded
  font program on every page whenever the descriptor left a style flag
  unset (the common case). Add a document-scoped FontStyleCache keyed by
  the FontFile2/FontFile3 object id, threaded through page and form
  extraction alongside the existing CMapDecisionCache.

The fourth finding (Form XObject rules never reach geometric detection)
is real but pre-existing for underline and needs the form walker to grow
path/paint tracking plus a new return type; deferred as a follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): ActualText items render at their glyphs' text rise (PR #125 review)

The EMC-built ActualText item used the captured text matrix without the
rise adjustment the ordinary Tj/TJ/' emission sites apply, so a tagged
run shown with Ts landed on the unshifted baseline — off the strikeout/
underline windows and inconsistent with untagged runs. The rise is
captured together with the first-glyph matrix (and at BDC for the
entry-position fallback): the item must render at the rise of its
GLYPHS, not whatever rise is set by EMC time.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): capture ActualText glyph position after the quote op's line move (PR #125 review)

The `'` handler skipped the entire suppressed-extraction block, so a
tagged span whose show op is `'` never captured its glyph matrix/rise —
the EMC item fell back to the BDC-entry matrix, which sits on the
PREVIOUS line (the `'` line move happens after BDC) with no rise. The
capture now happens right after the line move, matching the Tj/TJ
paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): style-boundary gate on subscript merge + strikeout suppression coverage (PR #125 review)

merge_subscript_items absorbed a script digit into its parent
regardless of underline/strikeout flags — dropping the digit's own mark
or widening the parent's over it. The merged item carries one flag, so
differing marks now break the merge, mirroring merge_text_items'
style-boundary rule (pre-existing for underline as well).

Also extends the table-suppression test to assert is_strikeout is
cleared alongside is_underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 17:30:47 -07:00
Abimael MartellandClaude Fable 5 6e5e5849c8 Detect substitution-cipher garbled text from broken ToUnicode CMaps (#120)
* fix(lib): detect substitution-cipher garbled text from broken ToUnicode CMaps

ParseBench text_simple__att10k.pdf (issue #118) ships Type0/Identity-H
fonts whose ToUnicode CMaps are authored garbled: every bfrange maps with
a wrong constant delta, so text extracts as pure-ASCII ciphertext
("Certificate" -> "8VceZWZTReV"). The embedded subset font has no cmap
table and no glyph names, so no decode source can recover the real text
(poppler and mupdf emit the same ciphertext). The only correct behavior
is to flag the page for OCR instead of serving the garbage silently --
but the text is 100% printable ASCII with word-like tokens, so it slipped
past is_garbage_text and detect_encoding_issues.

Add CipherGarbleStats, a letter-statistics discriminator that flags a
Latin-dominant sample (>=200 ASCII letters) when vowels are starved
(<=30% of letters) AND either:
- lowercase->uppercase transitions inside words exceed 10% of letter
  bigrams (a shifted lowercase alphabet straddles the ASCII uppercase
  block), or
- the letter histogram's cosine similarity against English letter
  frequencies drops below 0.60 (catches shifts that stay within case
  blocks).

Wired into analyze_text_quality (per-page, item-level) and
detect_encoding_issues (markdown-level), so extract_pages_markdown
reports needs_ocr + suspected_garbled_text and suppresses the garbage.

Thresholds validated against the 380-document pdf-evals snapshot corpus
(Swedish, Finnish, Turkish, German, romaji, schematics, all-caps and
camelCase-heavy docs): zero false positives, and byte-identical eval
output vs main. Garbled page measures vowel ratio 0.245 / case-shift
rate 0.225 / cosine 0.532; closest legitimate document on each axis is
0.264 / 0.021 / 0.801.

Fixes #118

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump pdf-inspector to 0.1.4, npm package to 1.9.11

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(lib): exempt uniform-case structured content from cipher detection

Address PR review (cubic P2): the frequency branch (english_cosine < 0.60)
fired on any Latin-dominant, low-vowel letter distribution unlike English,
so non-linguistic ASCII — DNA/protein sequences, ticker symbols, hex dumps —
could be suppressed and routed to OCR despite not being garbled. Measured:
DNA cosine 0.428 / vowel ratio 0.260, protein 0.738, tickers 0.747, hex
0.549 — all would have flagged.

Add a mixed-case guard to looks_garbled: garbled English is a permutation of
natural language and carries sentence capitalization (block-straddling shifts
invert the ratio — att10k is 60% uppercase; in-case Caesar shifts preserve it
at ~3%), so both keep some of each case. The exempted structured content is
uniform case (all upper or all lower). Requiring the minority case to be >=1%
of ASCII letters exempts single-case sequences while preserving both garble
signals, including the in-case-shift scenario the frequency branch exists for.

Strictly tightens the detector: it can only remove flags, so the eval corpus
stays at zero false positives (verified byte-identical to a baseline main
binary across all 185 PDFs) and att10k remains flagged. Adds regression tests
for DNA, protein, tickers, and an in-case Caesar shift.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(lib): make cipher detection case-agnostic via sorted-histogram shape

Address PR review follow-up: the mixed-case guard from the previous commit
returned before the vowel/frequency checks, creating a blind spot — a
uniform-case (all-lower or all-upper) substitution cipher is a plausible
broken-CMap output and would bypass OCR entirely.

Replace the case proxy with the actual invariant. A substitution cipher is
a bijection over a real language's alphabet, so it preserves the frequency
SHAPE (the sorted histogram) while scrambling letter POSITIONS (the unsorted
histogram). Signal 2 now flags when english_cosine < 0.60 (positions unlike
English) AND english_shape_cosine >= 0.90 (profile is still English-shaped).
This is independent of case, so it catches all-lower, all-upper, and
case-straddling shifts alike.

The exempted structured content fails one half: DNA/hex dumps have too steep
a profile (shape cosine 0.74 / 0.81 < 0.90), while protein sequences, ticker
symbols and base64 are not sufficiently unlike English in position (unsorted
cosine 0.74 / 0.75 / 0.77 >= 0.60). All stay out of OCR.

Still strictly corpus-safe: every real Latin document scores unsorted cosine
>= 0.70 (min 0.80), far above the 0.60 gate, so none can reach Signal 2.
Re-verified byte-identical to a baseline main binary across all 185 eval
PDFs; att10k remains flagged. Drops the now-unused case counters and adds
all-lowercase / all-uppercase shifted-prose regression tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: source Python package version from Cargo.toml via maturin

Address PR review (cubic P2): pyproject.toml pinned version = "0.1.0",
which overrides Cargo.toml, so a maturin build produced a 0.1.0 Python
artifact regardless of the crate version (it had drifted since the PyO3
bindings were added). Switch to dynamic = ["version"] so maturin sources
the version from Cargo.toml [package] version and the two can no longer
diverge. No workflow auto-publishes the Python package, so this is metadata
hygiene rather than a release-path fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 12:44:32 -07:00
Abimael MartellandCursor b375d6f102 feat(markdown): underline emission, Unicode scripts, style-preserving merges (#117)
* feat(markdown): underline emission, Unicode scripts, style-preserving merges (ENG-5015 2b)

Three formatting losses in the direct-extraction markdown path:

1. text_with_formatting gains <u> run emission (detect_underline option,
   default on) using the geometric is_underline flag from 1.9.9.
   Underline runs stay free of nested bold/italic markers — consumers
   match tag content literally. Heading lines keep plain text for
   bold/italic but preserve <u>: the tag carries meaning `#` doesn't.
2. merge_subscript_items now maps absorbed digit scripts to Unicode
   sub/superscript forms with direction from the baseline offset
   ("H"+"2" -> "H₂", "word"+raised "2" -> "word²", "m"+"3" -> "m³").
   NFKC/NFKD folds these back to plain digits so text matching
   downstream is unaffected; renderers keep the script semantics.
3. merge_text_items no longer merges across bold/italic boundaries —
   absorbing a styled run into a plain neighbor erased the styling
   before markdown emission ever saw it. On eval docs this recovers
   20-82 italic runs per document that previously emitted as plain.

Snapshots regenerated (diffs are the features: CCl₂F₂, m³, underlined
legal section headings, finer bold runs). pdf-evals regression suite:
202/202 real PDFs pass. napi 1.9.9 -> 1.9.10.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): break merges at underline boundaries too (review)

OR-merging underline stretched the eventual <u> span over neighboring
plain fragments. Merge runs now break on any style-flag change, the
redundant accumulator is gone, and format_list_item learned to move
bullet markers outside <u> wrappers so fully-underlined bullet lines
still render as markdown lists. td9264 snapshot regenerated — spans are
tighter (trailing periods correctly outside the tag).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(markdown): strip stray spaces before sentence punctuation (review)

Style-boundary item splits can strand a trailing period in its own
fragment, and multiple assembly paths join fragments with spaces,
yielding "word ." artifacts. Rather than chasing every join site, a
postprocess pass removes a space before `.`/`,`/`;` when the mark ends
its token (whitespace, cell boundary `|`, or end of text follows).
Dot leaders/ellipses and mid-token periods are untouched.

Fixes the td9264 "companies ." artifacts and two pre-existing
"armoring ," artifacts in the 2013-app2 snapshot. pdf-evals: zero
markdown diffs across all 203 corpus PDFs vs committed baselines.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(tables): trim spaces inside parenthetical cell fragments

* fix(tables): reject sparse prose row-stripe tables

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 23:40:08 -07:00
Abimael MartellandCursor 422a2ff118 feat(extractor): geometric underline detection on TextItem (#116)
* feat(extractor): geometric underline detection on TextItem (ENG-5015)

PDFs carry no underline font flag — underlines are stroked horizontal
lines or thin filled rects drawn under the baseline. Correlate those
graphics (already parsed from the content stream) with text items in a
post-pass: a rule within ~0.35em below the baseline covering >=60% of
an item's width marks is_underline.

Exposed through the napi and python bindings. Verified on real docs:
4/4 underlined sentences flagged on a Japanese report, links/headings
flagged on 8 of 10 underline-bearing eval docs, zero flags on docs
without underlines. Known FP source (table cell borders) documented —
downstream applies inline styling only to plain-text regions.

napi 1.9.8 -> 1.9.9.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): underline rules only from painted rects, normalized extents (review)

Two review fixes: (1) normalize rect extents before the thickness/width
checks — `re` operands pass through the CTM so width/height can be
negative, which missed negative-width rules and let negative-height
bands pass as thin; (2) only feed painted rects to underline detection —
`re` rects now wait in a pending list until a paint operator (S/s, f/F/
f*, B/B*/b/b*) confirms them, and `re W n` clip-only paths are discarded
at `n`, so invisible clip boundaries no longer underline nearby text.
Marking moved into content_stream where paint state lives (pre-rotation,
consistent device space).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): harden underline detection

* feat(cli): export positioned text item json

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 10:20:18 -07:00
Abimael Martell 1a5ba6f1e9 feat(api): expose OCR reason signal (#110) 2026-06-23 15:51:55 -07:00
Abimael Martell 8b63ceb084 emit ItemType::Image bboxes for Image XObjects (was: silently dropped) (#94)
* emit ItemType::Image bboxes for Image XObjects (was: silently dropped)

Background. ItemType::Image, MarkdownOptions::include_images, and the
markdown emitter's image-collection path have all been in the tree
for a while, but no producer ever populated them — content_stream.rs
explicitly `// Skip images — text extraction only` at the Do
operator, and the nested Form-XObject walker in xobjects.rs only
matched XObjectType::Form, silently dropping Image entries. The
declared types were dead code.

This PR lights them up. At every Do that resolves to an Image
XObject (both top-level and nested inside Form XObjects), we now
compute the page-space bbox from the current CTM via a new
`image_bbox_from_ctm` helper — handling both axis-aligned and
rotated/sheared placements via 4-corner AABB — and emit a TextItem
with `item_type: ItemType::Image` and the legacy `[Image: <name>]`
text payload that the markdown emitter already knows how to render.

Callers can now find raster figures via `extract_text_with_positions`
(and the `_mem` variant, newly re-exported at the crate root) without
needing to re-parse the PDF or run a vision/layout model. The intended
consumer is layout-aware text pipelines that want to crop figures and
caption them out-of-band.

Two backstops to avoid silent breakage for existing callers:

  1. `MarkdownOptions::include_images` default flipped `true → false`.
     If it stayed at `true`, every existing user of
     `extract_pages_markdown` would suddenly see `![Image: Im0](image)`
     placeholders inserted throughout their output the moment they
     upgraded. Image data is still available structurally via
     `extract_text_with_positions`; rendering it into markdown is now
     an opt-in. New regression test asserts `extract_pages_markdown`
     output is unchanged for the image-bearing fixture.

  2. Image items now also skip the layout heuristics
     (`detect_columns`, `detect_tables_from_rects`) via a new
     `is_text_layout_item` predicate. Without this filter, an image's
     left edge would land in the column-projection profile and skew
     table column detection — surfaced by
     `vector_grid_tests::upstage_key_functions_four_cols` going from 4
     detected columns to 5 in CI before the filter was added.

Re-exporting `extract_text_with_positions_mem` at the crate root —
strictly additive; mirrors how `extract_pages_markdown_mem` is already
available there.

Tests:

  - test_extract_text_with_positions_emits_image_bboxes — minimal PDF
    with one 200×100 image at (50, 600); asserts one Image item with
    correct bbox + page + text.
  - test_image_xobject_bbox_handles_rotated_ctm — 90° rotated image
    via shear-component CTM; asserts AABB is correct (handles non-
    axis-aligned placements via 4-corner clamp).
  - test_image_emission_does_not_change_default_markdown — asserts no
    `Image:` token leaks into default markdown output, regression
    guard for the include_images flip.
  - test_markdown_options_default_has_include_images_false — explicit
    sentinel so anyone flipping it back catches it in CI.

* Bump version from 1.8.15 to 1.9.0
2026-05-20 13:43:05 -07:00
Abimael MartellandClaude Opus 4.7 b5b91470db extract_tables: try vector-grid detectors before text heuristic (#85)
* extract_tables: try vector-grid detectors before text heuristic

extract_tables_in_regions_mem previously ran only the text-only
heuristic detector (tables::detect_tables) on the items inside each
region, discarding the rects and lines that extract_page_text_items
returned. That left the rect-backed and line-backed detectors
(detect_tables_from_rects, detect_tables_from_lines) unused by the
public region-scoped extraction path — they only ran through
detect_vector_grid_in_region_mem, which most callers don't use.

Keep the rects and lines, filter them to each region, and try in
order: rect detector → line detector → heuristic. Each candidate's
markdown is quality-gated by the existing needs_ocr checks
(is_garbage_text, is_cid_garbage, detect_encoding_issues,
looks_like_partial_table_ex); only the first clean output wins.
If all three produce empty or noisy output we still return
needs_ocr=true, matching prior behavior.

Effect on real prod-shape inputs from shadow logs:

  Full-page ruled ledger, 6 cols x ~15 rows:
    before: heuristic emits a 355-char two-row fragment
    after:  line detector emits the full 6520-char table

  Multi-row key/value layout with paragraph values:
    before: heuristic emits a 188-char header-only fragment
    after:  rect detector emits the full 1733-char table including
            the multi-bullet description cell

Existing fixtures that already passed via the heuristic continue to
pass: the quality gate rejects partial vector-grid output and falls
through, so the heuristic still wins where it produced the cleaner
result.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Bump version from 1.8.9 to 1.8.10

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 19:25:08 -04:00
Abimael MartellandClaude Opus 4.7 96c0f2102a tables/detect_rects: prefer rect-grid edges over text-cluster on N≥3 column tables (#82)
When a wire-bordered table has headers centered/right-aligned in their
cells but data left-aligned, cluster_x_positions can both merge adjacent
data columns (when the data-to-data gap is below the clamped threshold)
and drop the header-only x-positions in its singleton-filter pass. The
cell-rect fallback then used text-cluster column edges and lost a column
or fragmented neighbor cells.

Prefer rect-derived column edges when the rect grid has 3+ columns and
every rect column holds multiple text items. The all-cols-populated
check protects against decorative or background rects (prose laid out
in a frame, cell-fill rects with extra borders) that would otherwise
split a logical column into spurious sub-columns. The existing
prose-in-frame, well-distributed-columns, and wireless-prose guards
still fire for the cases they were built for.

Bump napi version 1.8.7 → 1.8.8.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 16:44:26 -04:00
Abimael Martell 6086577d11 tables/detect_rects: emit grid for multiline indented cells (#79) 2026-05-07 11:44:53 -07:00
Abimael MartellandCursor f2186ec1aa tables/detect_rects: don't accept relaxed grid on wireless prose (#78)
Require rect-derived column evidence before relaxing prose checks for two-column cell-rect fallbacks, so text-position alignment alone cannot synthesize a vector grid on wireless content.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-07 09:30:29 -07:00
Abimael MartellandClaude Opus 4.7 59b17f372a tables: tighten prose-in-frame rejection (#77)
* tables: lift detection on shaded-header + alt-row tables (#wired-grids)

Production telemetry on `wired_high_confidence`-classified table regions
showed `detect_vector_grid_in_region_mem` returning a usable grid only
~27% of the time, with the rest falling through to GLM-OCR. Three
surgical fixes target the dominant production shapes:

* Path-fill cell backgrounds: when the page has no `re` rects but draws
  cell backgrounds via `m`/`l`/`h`/`f*` sequences, prefer the fill-derived
  rects over the few section-level `W*` clip paths that previously won
  the priority gate. Activated when fill rects outnumber clip rects ≥3×.

* Dedup-induced cluster splits: page-background rects could pose as
  containers in the sub-rect dedup and evict a slightly smaller
  table-frame rect, breaking adjacency between column-cell groups so each
  column became its own cluster. Origin-anchored containers are now
  disqualified from sub-rect dedup. A separate exact-duplicate pass
  collapses the cell-padding/text-bg/cell-border triple emissions some
  PDFs produce, preserving original order to avoid reshuffling table
  output on multi-table pages.

* Prose-words rejection: the `cell-rect` fallback's whole-grid prose
  threshold also rejected real tables that include a description column.
  Now relaxed when content is well-distributed (≥75% of cols filled),
  while keeping the original strictness for prose-in-a-frame layouts.

Two regression fixtures from the opendataloader-bench corpus, covering
the dominant production failure categories:

* `greencomp_competence.pdf` — 2-col shaded-header + plain-body glossary.
  Mirrors production crops #1 (Contractions glossary) and #6 (BIO 350
  course header).
* `upstage_key_functions.pdf` — 4-col shaded-header + alt-row backgrounds
  + merged left column. Mirrors production crops #2 (Parameter/Value
  alt-row), #7 (Spanish XML schema), and #8 (Córdoba multi-row header).

Existing fixtures stay green (doc 51 wrapped-label, doc 128 forecast
six-cols, td9264 snapshot). 133 unit + integration tests pass; clippy
clean.

Bumps napi/package.json 1.8.4 → 1.8.5.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* tables: tighten prose-in-frame rejection — fixes pdf-evals #30 regression

PR #76's shaded-header detection lift surfaced a regression on
accessory_building_permit_application_1 (TEDS 0.10 → 0.05): a
paragraph of legal text laid out in a 2-column justified block was
being admitted as a 10×2 fake table where every cell holds a
sentence fragment ("I agree to comply...", "I", "It is the property
owner's responsibility..."). Per pdf-evals PR #30 review, this is
the kind of regression production users will notice — the markdown
is structurally and semantically misleading.

Root cause: PR #76's prose-rejection only fires for `num_cols >= 4`,
so the 2-col prose-in-a-frame case slipped past it entirely. The new
fill-priority + dedup changes started producing rects for this layout
that 1.8.4 correctly ignored.

Fix: tighten the prose-in-frame check.
- Lower the column-count guard from `>= 4` to `>= 2`.
- Add a content-length signal as the primary discriminator: when the
  prose-words trigger fires AND mean non-empty cell length exceeds
  65 chars, reject regardless of column distribution.

The 65-char threshold cleanly separates observed cases:
  accessory_building (prose-in-frame): mean 74 chars  → REJECT
  upstage_key_functions (real 4-col table): mean 53   → admit
  greencomp_competence (real 2-col glossary): mean 20 → admit
  accessory_building (real 5×3 form data): mean 10    → admit

The well-distributed-cols relaxation that PR #76 added stays —
"label / value / description / benefit" tables (#7, #8 from the
production crops) still pass, but only when their mean cell length
stays below the prose threshold.

New regression test `accessory_building_rejects_prose_in_frame` asserts
both that the real 5×3 form data table survives AND the 10×2 prose
block is rejected. Snapshot test `test_snapshot_td9264` updated to
match new output — old snapshot captured the same prose-in-frame bug
on regulatory text (paragraphs emitted as 3-col `||text||` fake-table
rows). New snapshot emits clean prose paragraphs, which is correct.

Verification:
- cargo test --all: 424 lib + 133 integration + 2 doc tests pass
- cargo fmt --check clean
- cargo clippy -- -D warnings clean (lib-level; pre-existing
  test-level clippy issues on the wired-grids branch unaffected)
- Existing fixtures stay green: forecast_table_chart_six_cols (PR
  #72), bits_pilani_* (PR #73), greencomp_competence_two_cols and
  upstage_key_functions_four_cols (PR #76).

This branch is based on abimaelmartell/wired-grids so it includes
PR #76's commits plus this fix on top. Suggest merging this and
closing #76, OR rebasing #76 to incorporate this fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* tables: drop early-dedup atom — caused broad TOC + matrix corruption

Bisected PR #76's 4 atoms against the SEC 10-K
0001104659-25-093871_183e44ac.pdf which appeared as a TEDS regression
in pdf-evals PR #30. Result:

  atom                       | TOC | perf-graph | qualifications
  ---------------------------|-----|------------|---------------
  fill-priority              |  ✓  |     ✓      |       ✓
  early-dedup                |  ✗  |     ✗      |       ✗
  page-bg disqualification   |  ✓  |     ✓      |       ✓
  prose-relaxation           |  ✓  |     ✓      |       ✓

Early-dedup was the SOLE source of all three regressions on this doc.
Tried a more conservative variant (≥3 copies only — pair-duplicates
appear in legit multi-section layouts like 10-K dividers above + below
section headers); didn't fix the regression. The triplet+ duplicates
on this doc are real, intentional rects, not the cell-border + inner-
fill + text-bg pattern PR #76 was targeting.

Drop early-dedup. Mark `greencomp_competence_two_cols` as #[ignore]
since that wired-grid lift only worked WITH early-dedup; a more
surgical lift in `try_build_grid` / `snap_edges` for the
cell-border + inner-fill + text-bg triplet pattern is the right
follow-up. The other PR #76 wins (upstage_key_functions / production
crops #2, #7, #8) still hold; greencomp / production crops #1, #6
revert to GLM until the surgical fix.

Validation on the regression doc:
  0001104659 TOC PART II markers:    4 (matches main, was 2 with PR#76)
  0001104659 perf-graph data row:    2 (matches main, was 1)
  0001104659 qualifications rows:    9 (matches main, was 6)
Validation on the prose-frame doc:
  accessory_building fake-table:     0 (matches main, was 1 with PR#76)
  accessory_building prose intact:   1 (matches main)

cargo test --all clean, cargo fmt --check clean, cargo clippy --lib
-- -D warnings clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 13:37:08 -07:00
Abimael MartellandClaude Opus 4.7 cfc080f79a extractor: drop Latin-1 mojibake on Type0/CID fonts; tokenize wide TSR items (#75)
* extractor: drop Latin-1 mojibake on Type0/CID fonts; tokenize wide TSR items

Two text-extraction failure modes surfaced by table-candidate shadow
data; both also affect the existing TableFormer / vector-grid paths
since they share `extract_tables_with_structure_*_mem`'s downstream
cell-fill.

1. CJK / multi-byte mojibake. The bottom Latin-1 fallback in
   `extract_text_from_operand` ran unconditionally. For a Type0/CID
   (Identity-H) font whose ToUnicode CMap fails to parse, the bytes
   are CIDs (font-internal indices), not character codes — per-byte
   Latin-1 produces mojibake (e.g. 2-byte CID 0xCDD9 surfaces as "ÍÙ").
   Gate that fallback on `FontWidthInfo.is_cid` (set by
   `parse_type0_widths` for `/Subtype /Type0`). For Type0 fonts with
   any non-ASCII byte, emit one U+FFFD per CID instead so
   `detect_encoding_issues` still trips and the page is flagged for
   OCR — preserving the existing OCR-routing path that the
   high-Latin-1 garbage used to satisfy by accident. Type1 / TrueType
   simple fonts retain the per-byte Latin-1 round-trip (it IS the
   canonical interpretation for them; verified against an existing
   pdf-evals fixture where bytes like 0xB6 are legitimate Latin-1).

   Threaded `font_widths: &PageFontWidths` through
   `extract_text_from_operand` and its 5 call sites in
   `content_stream.rs` / `xobjects.rs`.

2. Dense-cell text collapse in `extract_tables_with_structure_cells_mem`.
   Stage-1 routing did per-item assignment — each TextItem went into the
   single cell whose bbox contained its center. When a row's text is
   rendered as one wide Tj (e.g. "Marshall Islands 0.9 0.9 0.9"), the
   whole row parks in one cell and the rest of the row stays empty.

   New `split_item_into_token_subitems` helper splits each item into
   per-token virtual sub-items with x positions estimated from
   `effective_width / char_count` and the token's character offset.
   Stage 1 then routes per-token. Single-token items collapse to a
   one-element vector (no behavior change). Multi-token items spanning
   multiple cells distribute correctly. Stage-2 orphan recovery now
   operates on token-grain orphans rather than re-trying whole items.

Tests:
 - `cid_font_with_unparseable_cmap_does_not_emit_latin1_mojibake` (unit)
   exercises the Type0/CID + unparseable-CMap fallback path.
 - `simple_font_latin1_fallback_passes_high_bytes_through` (unit)
   guards the false-positive case where a Type1 font's `/ToUnicode`
   reference is set but bytes are legitimate Latin-1 character codes.
 - `test_extract_tables_with_structure_distributes_wide_item_across_cells`
   (integration) builds a synthetic PDF with one wide Tj and asserts
   each token lands in its own cell.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* tests: pin CID mojibake fix mechanism with FFFD assertions

Two complementary tests for the Type0/CID Latin-1-fallback guard:

1. Tighten `test_identity_h_no_tounicode_suppresses_garbage` on the
   existing real-PDF fixture `shinagawa_identity_h.pdf` to also assert
   the pre-suppression text contains U+FFFD and contains no high-Latin-1
   chars. Pins down WHICH mechanism is suppressing the garbage so a
   future regression that re-enables Latin-1 mojibake fails loudly here
   instead of silently switching the suppression chain back to
   `is_cid_garbage` + high-Latin-1 detection.

2. Add `test_synthetic_type0_broken_tounicode_emits_fffd_not_latin1_mojibake`
   with a fully-synthetic Type0 / Identity-H PDF built in process. We
   control the malformed ToUnicode contents, the descendant CIDFontType2
   shape (just enough for `parse_type0_widths` to set `is_cid=true`,
   which is what the new guard keys off of), and the Tj byte stream.
   No fixture file or external license needed. Reproduces the exact
   "Type0 + non-ASCII bytes + unparseable ToUnicode" code path that
   produced the production mojibake samples.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-02 21:28:39 -07:00
Abimael Martell 1f28e523fd tables: keep wrapped labels in TSR output (#73) 2026-04-30 10:56:05 -07:00
Abimael Martell 97fc32ac70 tables: prefer rect edges for cell-grid fallback (#72)
* tables: prefer rect edges in cell-grid fallback

* Bump version from 1.8.1 to 1.8.2
2026-04-29 21:27:27 -07:00
Abimael Martell 5b1fe30c66 tables: expand multi-row cells in-place when fallback heuristic is empty (#71)
* tables: expand multi-row TSR cells in place

Recover row-under-counted TSR tables by splitting overstuffed cells with native PDF text bands before falling back to heuristic extraction.

Made-with: Cursor

* docs: note multi-row expansion scope

Clarify that the row-band cap intentionally keeps v1 focused on common small row-loss cases while larger compressions continue to use heuristic fallback.

Made-with: Cursor
2026-04-29 08:18:09 -07:00
Abimael Martell c186a036fc tables: add detectVectorGridInRegion napi export for region-scoped vector grid detection (#70)
* feat: add vector grid region detector napi export

Expose region-scoped vector PDF grid detection so TSR callers can reuse native geometry before model fallback.

Made-with: Cursor

* fix: address vector grid review feedback

Return null for rotated vector grids until the coordinate transform has coverage and reject out-of-crop cell boxes surfaced by real-PDF smoke testing.

Made-with: Cursor

* test: add crop bbox plausibility coverage

Cover in-crop, out-of-crop, slack-boundary, and non-positive DPI behavior for vector grid cell bbox validation.

Made-with: Cursor
2026-04-28 17:20:15 -07:00
Abimael MartellandClaude Opus 4.7 d196d435d1 fix: TSR auto-fallback bugs found in review, v1.7.2 (#68)
Three fixes to extract_tables_with_structure_auto_mem (added in
1.7.1) caught by external review:

1. multi_row_in_cell over-triggered on legitimate multi-line cells.
   The previous threshold (item span > 1.3× either smallest cell or
   tallest item height) fires on any cell with 2+ y-separated text
   items — including rowspan>1 cells, wrapped descriptions, and
   superscript/subscript runs. Replaced with two gates:
   - skip cells whose declared rowspan > 1 (intentional multi-line)
   - require an actual whitespace gap (>~half a line height)
     between the bottom of one item and the top of the next, in
     PDF-native y-coordinates. Same-line items with tall glyphs or
     superscripts have negative or near-zero gap; truly separate
     visual rows have gap ≈ leading − line-height.
   FNBO regression test still passes; new test covers a rowspan=2
   cell with two visible text lines and verifies no fallback fires.

2. Heuristic returning empty silently replaced TSR markdown with
   "". The auto wrapper now keeps the TSR markdown when the
   heuristic markdown is empty/whitespace and tags fallback_reason
   with `_heuristic_empty` suffix (e.g.
   `multi_row_in_cell_heuristic_empty`). Worst case we ship the
   same wrong-but-non-empty TSR output we'd have shipped before
   1.7.1; we never replace useful output with literally nothing.

3. One bad input blanked the whole batch. Errors from
   detect_tsr_quality_issue or extract_tables_in_regions_mem now
   stay scoped to the single input — that input falls through to
   raw TSR markdown with a `_error` reason label so callers can
   metric on it. Other inputs in the batch are unaffected.

3 new integration tests:
- test_auto_does_not_fire_on_legit_rowspan_cell
- test_auto_keeps_tsr_markdown_when_heuristic_returns_empty
- test_auto_isolates_per_input_failures

All 6 auto tests + full 123-test suite pass. FNBO local replay
still triggers fallback (phantom_empty_row signal in this run) and
emits correct Shawnee/BVP/Sonoma rows with correct census tracts.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-27 11:23:03 -07:00
Abimael MartellandClaude Opus 4.7 fbab84fc20 feat: TSR auto-fallback to heuristic on quality issues, v1.7.1 (#67)
Adds extract_tables_with_structure_auto_mem (Rust) /
extractTablesWithStructureAuto (napi). Returns
TableExtractionResult { markdown, fallback_reason } per input.

The wrapper runs the existing TSR-hybrid path then checks the
resulting cells for two known SLANet detection pathologies:

* phantom_empty_row: empty row sandwiched between non-empty rows
  (cheap, cell-metadata only).
* multi_row_in_cell: re-reads PDF text items, flags any cell whose
  contained items span >1.3× either the smallest cell height or
  the tallest contained item height. Catches the FNBO failure mode
  where a tall TSR cell absorbs two adjacent PDF rows.

When either fires, extract_tables_in_regions_mem runs over the same
crop bbox and its markdown replaces the TSR markdown.
fallback_reason carries the diagnostic label so callers can emit
metrics and watch each pathology independently.

Validated on FNBO branches PDF page 1 (Kansas region):
- TSR-only: merges Shawnee into BVP, wrong census tract on Sonoma.
- Auto fallback (phantom_empty_row): each row separate, correct tracts.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-27 10:13:25 -07:00
Abimael Martell 8a0f98dee7 repair malformed PDF containers (#65) 2026-04-27 09:38:25 -07:00
Abimael MartellandClaude Opus 4.7 3f8fb645c9 Fix TSR cell assignment for overlapping table boxes (#62)
* feat: TSR-aware table extraction (extract_tables_with_structure_mem)

New public function that consumes raw structure-recovery output (HTML
structure tokens + per-cell bboxes from a model like SLANet) and assembles
markdown tables by pulling cell text from the native PDF — no OCR, no
geometry inference.

Why: the existing extract_tables_in_regions_mem infers grid geometry from
text positions only and can't distinguish merged cells from multiple narrow
columns. Pairing structure recovery from a layout/TSR model with native
PDF text gets perfect text quality with proper row/col/span structure.

- New module src/tables/structured.rs: token state machine, polygon→AABB,
  crop-px→page-pt, rowspan/colspan-aware cell layout, markdown emitter.
  Accepts both 4-element rects and 8-element 4-corner polygons.
- New public extract_tables_with_structure_mem in src/lib.rs that reuses
  extract_page_text_items, region_overlaps_item, and the shared region
  text-collection helper. No existing public function modified.
- napi binding extractTablesWithStructure mirroring the existing
  extractTablesInRegions shape (f64 in JS → f32 internally).
- 14 unit tests + 5 integration tests, including a real-PDF gold-standard
  match against bits_pilani_feedback.pdf.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* TSR follow-ups: header-aware separator, cells API, v1.6.0

- cells_to_markdown emits the separator after the LAST row that contains
  is_header=true cells, falling back to "after row 0" when no header is
  flagged. Multi-row theads now render correctly. Three new unit tests
  cover: multi-row header, header not on row 0, no headers (fallback).
- New public extract_tables_with_structure_cells_mem returning
  Vec<Vec<StructuredCell>> so callers can drive their own rendering or
  debug overlays without re-doing the parse + extraction. The markdown
  variant now wraps it. The previously-unused page_pt_bbox field is
  surfaced through this API.
- New napi binding extractTablesWithStructureCells + StructuredCellJs.
- Bump @firecrawl/pdf-inspector to 1.6.0.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix TSR cell text assignment for overlapping bboxes

Made-with: Cursor

* bump npm package version to 1.6.1

Made-with: Cursor

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 16:35:19 -07:00
Abimael MartellandClaude Opus 4.7 f6d5e214f1 feat: TSR-aware table extraction (extract_tables_with_structure_mem) (#61)
* feat: TSR-aware table extraction (extract_tables_with_structure_mem)

New public function that consumes raw structure-recovery output (HTML
structure tokens + per-cell bboxes from a model like SLANet) and assembles
markdown tables by pulling cell text from the native PDF — no OCR, no
geometry inference.

Why: the existing extract_tables_in_regions_mem infers grid geometry from
text positions only and can't distinguish merged cells from multiple narrow
columns. Pairing structure recovery from a layout/TSR model with native
PDF text gets perfect text quality with proper row/col/span structure.

- New module src/tables/structured.rs: token state machine, polygon→AABB,
  crop-px→page-pt, rowspan/colspan-aware cell layout, markdown emitter.
  Accepts both 4-element rects and 8-element 4-corner polygons.
- New public extract_tables_with_structure_mem in src/lib.rs that reuses
  extract_page_text_items, region_overlaps_item, and the shared region
  text-collection helper. No existing public function modified.
- napi binding extractTablesWithStructure mirroring the existing
  extractTablesInRegions shape (f64 in JS → f32 internally).
- 14 unit tests + 5 integration tests, including a real-PDF gold-standard
  match against bits_pilani_feedback.pdf.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* TSR follow-ups: header-aware separator, cells API, v1.6.0

- cells_to_markdown emits the separator after the LAST row that contains
  is_header=true cells, falling back to "after row 0" when no header is
  flagged. Multi-row theads now render correctly. Three new unit tests
  cover: multi-row header, header not on row 0, no headers (fallback).
- New public extract_tables_with_structure_cells_mem returning
  Vec<Vec<StructuredCell>> so callers can drive their own rendering or
  debug overlays without re-doing the parse + extraction. The markdown
  variant now wraps it. The previously-unused page_pt_bbox field is
  surfaced through this API.
- New napi binding extractTablesWithStructureCells + StructuredCellJs.
- Bump @firecrawl/pdf-inspector to 1.6.0.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 00:55:39 -07:00
Abimael MartellandClaude Opus 4.7 6819852541 fix: reject cell-rect "tables" that are actually prose in a framed box (#58)
The rect-based cell fallback in detect_row_stripe_table_from_cell_rects
derives columns purely from text X-position clustering. When prose
wraps inside a bounding-box rect (chat transcripts, stylized figures),
the word-boundary gaps cluster into many spurious columns, producing
a multi-column "table" that is just fragmented prose.

Count cells containing common English function words (articles,
prepositions, pronouns, common verbs) and reject the fallback when
20%+ of non-empty cells contain any such word. Real tabular data —
labels, units, numbers, short identifiers — rarely contains these.

Update the td9264 snapshot: the government document section that
previously rendered as a malformed table now renders as cleaner
prose + a proper CFR list.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 10:23:28 -07:00
Abimael MartellandClaude Opus 4.7 2876fa4b3e fix: tighten rect-row span check in propagate_merged_cells (#57)
* fix: don't reclassify wrapped bold list leads as headings

When a numbered/bulleted list item's bold lead phrase wraps onto a
second visual line, that line is all_bold + standalone, which scored
above the rarity heading threshold and was emitted as #### in the
middle of the item. That reset in_list, so the body continuation
below picked up a stray `- ` bullet via the struct-tree LI path,
shattering a single item into heading + stray bullets.

Guard the font heuristic: when already inside a list, skip heading
classification for lines at the list continuation indent with a Y
gap within para_threshold. Structure-tree headings still win.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: tighten rect-row span check in propagate_merged_cells

propagate_merged_cells used an overlap-based predicate with ±tol
slop that returned true at shared row boundaries — a rect whose
top exactly equals row N's bottom lies entirely below the row, yet
the predicate considered it to span row N. When multiple background
rects aligned on a shared Y edge (e.g. consecutive row-stripe
shading), each adjacent rect would over-reach by one row, cascading
labels and data from unrelated rows into a single merged cell.

Replace the overlap predicate with a containment check: rect bottom
at or below row bottom, rect top at or above row top (each within
tol). Genuine merged-cell rects fully contain the rows they span;
tangent rects do not.

Update two snapshots that were encoding the old buggy output.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 09:33:43 -07:00
Abimael MartellandClaude Opus 4.7 4b5ae91f54 feat: expose per-page markdown extraction to Python and Node (#53)
* feat: expose per-page markdown extraction to Python and Node (#49)

Implements the feature requested in issue #49: a list-of-pages markdown output
from the Python API. Matching the existing project pattern, the feature lives
in the Rust core and is surfaced through every binding.

- Rust core: `extract_pages_markdown` (path) and `extract_pages_markdown_mem`
  (bytes) now take `Option<&[u32]>` — `None` returns every page in document
  order; a slice restricts and preserves caller order.
- Python: new `extract_pages_markdown(path, pages=None)` and
  `extract_pages_markdown_bytes(data, pages=None)` functions plus
  `PageMarkdown` / `PagesExtractionResult` classes; stub file updated.
- Node: `extractPagesMarkdown(buffer, pages?)` — `pages` is now optional.
- Tests: 2 new Rust integration tests, 9 new Python tests, 2 new Node
  assertions. All 372 unit + 107 integration + 53 Python tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump version from 1.3.0 to 1.4.0

Minor bump for the new per-page markdown extraction API exposed through
the Python and Node bindings.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 17:19:03 -07:00