9bb134975914b2ba0cccd5537b496e59a51923dc
86
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5eb6a13860 |
feat(layout): banded region segmentation for vertically-changing column layouts (#426)
* feat(layout): banded region segmentation for vertically-changing column layouts Pages whose column structure changes down the page (newsletter bands, figure-split flows, a three-column strip inside a two-column page) cannot be represented by one full-height column set: the projection profile either finds nothing and Y-interleaves the columns, or weaves the odd band's text into the wrong buckets. - Split pages into horizontal bands at full-width whitespace gaps (wide spanning items are excluded from occupancy — separators sit inside the very gaps being sought), detect columns independently per band, and re-merge consecutive bands with matching gutters across empty gaps so figure floats keep flowing down their columns while headline-separated bands stay independent. - Engage only on contradicting evidence: a prose-validated band whose column count differs from the page-level structure. Pages the flat column model already explains keep their current ordering. - Read short prose columns (5-14 lines, >=60% width fill on >=60% of lines in every column) as newspaper instead of Y-interleaving them as tabular. Kept as a standalone reading-order refinement so the table pipeline's is_newspaper_layout veto is unaffected. - Split ordering entry points: order_multi_column_region keeps every page-level defense; order_validated_band (banded planner only) trusts validated bands, skipping line-count minimums and straggler splitting that would misfire on Y-cohesive bands. * docs(layout): record why the band wide-item test is per-item Assembling same-baseline fragments into runs before the wide test was implemented and measured against the reading-order benchmark: word-gap and gutter-gap distributions overlap in real documents, so assembled runs fused narrow-guttered column pairs into page-wide lines, emptied the occupancy, and disengaged banding on pages it rescues — a measured regression with no measured win. Keep the per-item test (a fragmented separator can suppress a cut, which only misses an engagement) and document the boundary for future attempts. * fix(layout): keep figure placeholders out of band whitespace probes An image placeholder sitting between two matching column bands is the very figure float whose flow-through the band merge exists for, yet it read as content twice: its glyph box filled the occupancy gap (blocking the cut) and the merge probe counted it as separator content (blocking the merge). Both probes now see text layout items only. * fix(layout): anchor band-merge matching on the founding band's columns The merge comparison ran against the widened union, whose gutter is the intersection of its constituents' gutters. Across a chain of one-directionally drifting bands that intersection can walk past GUTTER_TOLERANCE and reject a band identical to the run's own first member. Compare candidates against the founding band's raw columns instead: the run's column system is defined by its founder, so drift can no longer accumulate in either direction. Union widening is kept for item bucketing only. * test(layout): differential coverage for the band-merge anchor rule - banded_layout_rejects_creeping_drift: a band within tolerance of the moving union but 36pt from the founder must not join the run — the case the anchor rule exists for; the pre-anchor union admitted it. - Reword the founder-anchor chain test as the invariant lock it is. - Note at the merge site why a reject-overlapping-unions guard is unimplementable: detect_columns returns contiguous partitions whose adjacent regions share boundary coordinates, so the check degenerates to exact-equality matching and rejects every legitimate merge; the boundary-disagreement zone is bounded by GUTTER_TOLERANCE and split proportionally by greatest-overlap bucketing. |
||
|
|
74ebce430c |
fix(layout): accept figure-diluted prose columns via a sustained full-line run (#417)
* fix(layout): accept figure-diluted prose columns via a sustained full-line run The columns_have_prose gate (added against table/TOC/checklist false-splits in relative-valley detection) requires >=40% of a column's lines to span >=45% of its width. A genuine prose column hosting a figure and caption dilutes that global ratio below the bar — measured 0.38 on a two-column research page — so the valley is rejected, the page falls back to single-column, and same-baseline items from both columns merge across the gutter into woven sentences. Accept a column that contains a sustained paragraph block instead: at least 6 consecutive full-width lines. The scattered layouts the gate exists to reject cannot produce an unbroken run of full lines, and the other guards (column width, minimum lines, items-per-line) still apply. Tests pin both directions; regression corpus is unchanged. * docs(layout): restore columns_have_prose rustdoc displaced by hoisted const The hoisted LINE_FILL_THRESHOLD landed between the gate's rustdoc block and the items that followed, absorbing the function's documentation onto the constant. The docs return to the function, extended to cover the sustained-run acceptance path. |
||
|
|
84789459b1 |
feat(extractor): stamp items with the font family name, not the resource tag (#415)
* feat(extractor): stamp items with the font family name, not the resource tag
TextItem::font carried the page's font resource name ("F2", "T22") —
an arbitrary per-page tag — even though both content-stream parsers
already resolve the /BaseFont family name for bold/italic detection at
every item-creation site. Stamp that resolved family name instead
("ABCDEF+CMMI10", "Courier"), from a single item_font_name helper so
the two parsers cannot drift.
One deliberate carve-out, documented on the helper: resource names
using Distiller's CID convention (C2_0, C0_1) are kept as-is, because
text_utils::is_cid_font keys on that prefix for micro-gap joining and
the family name carries no CID marker to replace it.
Consumers that match on font names start working against real names:
- Code detection (is_monospace_font) previously never fired against
opaque resource tags. It now does — so line classification also moves
from any-item matching to a majority-by-characters rule
(line_is_monospace): code lines are wholly monospace, while a lone
URL or identifier styled in a mono face inside a prose line must not
fence the surrounding sentence.
- Heading/body font grouping now merges resource aliases of the same
family instead of treating them as distinct fonts.
- Positioned-item output (--items-json and the bindings) reports real
face names.
Regression corpus: code-heavy manuals improve substantially (assembly
and C snippets previously emitted as prose now fence with line
structure preserved); remaining churn reviewed as improvements.
* fix(markdown): address review of font-name consumers
- Monotype is a foundry prefix on proportional faces (Monotype Corsiva,
Monotype Garamond); it must not satisfy is_monospace_font's generic
"mono" token. Regression tests pin both directions.
- Flush the pending code block before inserting a positioned table or
image, so a block that falls between two code lines cannot be emitted
ahead of code that precedes it in reading order; a code line after
the block reopens a new fence naturally.
* fix(markdown): emit sub-3-char mono fragments as plain text, not fences
A lone registered-trademark glyph or stray bullet set in a mono face is
not code; a fenced block containing one character reads as noise.
* fix(markdown): font-based code blocks open only at paragraph boundaries
HTML-to-PDF producers smear an inline code literal's mono style across
whole wrapped lines, so a prose paragraph can alternate body and mono
fonts line by line. Fencing those lines cut sentences in three: prose
head, fenced middle, prose tail. A mono-set line that continues an open
prose paragraph now stays prose; font-based blocks open at paragraph
boundaries (or continue an open block), and struct-tree Code roles are
honored unconditionally.
* refactor(markdown): drop paragraph-flush branch made unreachable by the boundary gate
The enclosing guard proves in_paragraph is false, so the nested flush
could never run; the guard and mono check collapse into one condition.
|
||
|
|
076183e2e4 |
fix(extractor): cap content-stream decode before allocating operators (#373)
* fix(extractor): cap content-stream decode before allocating operators The 1M operation limit ran after lopdf materialized the full vector, so a compact page of q/Q pairs could still abort under memory pressure. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(extractor): treat NUL and form-feed as PDF whitespace in op counting Names must stop on the full PDF whitespace set so a following operator is not absorbed into /Name, which would undercount and skip the decode cap. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(extractor): scan inline-image EI with the full PDF whitespace set A missed EI terminator used to consume the rest of the stream and drop later operators from the decode cap. If EI is absent, keep scanning. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
ec6e54afb8 |
fix(extractor): merge small-caps runs so they stop reading as table columns (#371)
* fix(xobjects): track text line matrix and handle T*/TL/'/"/Tc/Tw in Form XObjects The Form XObject text extractor in xobjects.rs is a separate hand-rolled implementation of the operator state machine in content_stream.rs, and it had drifted well out of parity: - No text line matrix (TLM). `Td`/`TD` were applied to the text matrix already advanced by `Tj`/`TJ`, so every line began where the previous line *ended* instead of at the line start. Lines marched off the right edge and were dropped as off-page. - `T*` was not handled at all, so it never advanced to the next line. - `TL`, `'` and `"` were missing, and `TD` never set the leading as a side effect. - `Tc`/`Tw` were hardcoded to 0.0 when computing advance widths, drifting positions and inserting spurious spaces. - Text state (Tc/Tw/TL/Tf) is part of the graphics state but was not saved or restored by `q`/`Q`. This matters well beyond an edge case: producers that emit a page stream of just `q /X Do Q` and put all content in a Form XObject are common in print-to-PDF and typesetting workflows, so this parser is on the hot path for whole classes of real documents. Measured on nycourts.gov 199AD3d.pdf (1370 pages, PDFlib producer, every page wrapped in a Form XObject, 1331 pages using T*), word recall against a pdftotext reference goes from 19.2% to 97.7% — 116k extracted words to 582k against a 578k-word reference. On a 10-page subset, sequence similarity goes from 27.3% to 97.7%, against 99.8% for Mistral OCR. pdf-evals: 195 passed / 7 failed, byte-identical to the origin/main baseline with the same failure list — no regressions. Adds 7 unit tests covering the line-matrix-relative `Td`, `T*`, `TD` setting leading, `'`, `"`, `Tc` advance widths, and `q`/`Q` text-state restore, all driven through a page whose content is only `q /X1 Do Q`. * fix(xobjects): restore fill colour across q/Q in Form XObjects A white fill set inside a q/Q pair leaked past the Q, so all subsequent text was treated as invisible and dropped. Save and restore fill_is_white with the rest of the graphics state. This was the cause of several long-standing extraction failures where whole passages went missing or degraded into per-character garbage: cambridge_excerpt (+8.5KB of recovered text), MTUAeroEngines (+7.4KB), 2025_findings-acl_668 (+1.4KB), HTM_02-01_Part_A (+1.3KB), and HuttoISDWorkPerks / ebgt7isj04ophcq, which both went from exploded per-character tables to clean prose. Adds a regression test that fails without the restore (the text after Q is dropped entirely). Reported by cubic on #369. * fix(extractor): merge small-caps runs so they stop reading as table columns Typesetters render small caps as a full-size capital immediately followed by shrunken capitals in the same font — `(R) Tj` at 9.98pt, then `(OLANDO) Tj` at 6.74pt, touching. The 20% font-size band in merge_text_items split those into separate items, and since table column boundaries cluster on item *start* positions (find_column_boundaries never consults widths), each fragment started far enough from the last to become its own column. The result was garbled pseudo-tables. From 199AD3d.pdf p.5: |OLANDO|COSTA|NIL|INGH|| |---|---|---|---|---| |IANNE|ENWICK|ETER|OULTON|| |R|T. A|, P.J.|A|C. S| |---|---|---|---|---| now: |ROLANDO T. ACOSTA, P.J.|ANIL C. SINGH| |---|---| |DIANNE T. RENWICK|PETER H. MOULTON| Adds is_small_caps_continuation, gated tightly enough to exclude the other reasons a smaller run follows a larger one: superscripts and footnote markers (requires an uppercase *letter*, so digits never qualify), drop caps (the following body text is mixed case), and adjacent cells or separate words (requires the runs to be visually contiguous). A small-caps junction is mid-word, so it also suppresses the space that would otherwise be inserted ("T. A" + "COSTA" -> "T. ACOSTA", not "T. A COSTA"). On 199AD3d.pdf this drops false-positive tables from 65 to 52 and lifts word recall against a pdftotext reference from 97.7% to 97.8%. Verified by running both binaries over all 203 eval PDFs and diffing outputs: 19 differ, 184 byte-identical. Beyond the reporter volume the merge also fixes: Waters-Edge two more garbled heading-tables — "##### B. C R S" plus "||OMPLIANCE|EPORTING YSTEM|" became "##### B. COMPLIANCE REPORTING SYSTEM" cn-student-handbook six TOC entries had collapsed to their initials; "U M S H..." is now "UNIVERSITY MISSION AND STUDENT HANDBOOK..." 546403 "(FAA)" + "ADVISORY CIRCULARS ( )" + a stray "CONT" became "(FAA) ADVISORY CIRCULARS (CONT)" ERP-2025 "T ABLE B-1" -> "TABLE B-1" DMP-Keypad "THINLINE" + stray "TM KEYPADS" -> "THINLINETM KEYPADS" zhaw / Stijn subscripted math variables: "*T* *G*" -> "*TG*" One known regression, called out rather than hidden: PA_PVEM_Sen_Waldo_Fernández (+1689 bytes). Its running footer genuinely is small caps, so the merge correctly assembles the text (400 items -> 275), but the now-contiguous footer lines align well enough that the heuristic detector turns the wrapped document title into a 4-column table where it previously rendered as bold prose. Attempting to fix that in the detector was a dead end and is not included here: gating the `num_cols >= 3` bypass in has_table_like_content on "no cell has >= 12 words" removed 143 tables across 44 files, including correct ones — MCF5235RM went 505 -> 483 and its register glossary degraded into a fused column plus a ~100-word cell, because rejecting a good candidate lets a worse fallback win. No wordiness threshold separates the two cases: PA_PVEM's longest cell is 14 words while MCF5235RM's legitimate cells are <= 10. A positional signal (running-header/footer bands, or cross-page repetition) is the way in, and belongs in its own change. Adds 7 unit tests: the two-column small-caps row from the reporter volume, and rejection of superscript digits, drop caps, separate words, lowercase continuations, and out-of-band size ratios. * fix(extractor): tighten small-caps gate to cross-band junctions only Three findings from cubic on #371, all valid. 1. Space suppression was too broad. `is_small_caps_continuation` accepted size ratios up to 0.92, which is *inside* the 20% band that merge_text_items already treats as the same size. Two similarly-sized uppercase words with a small real word gap therefore had their space suppressed even though the normal path would have merged them correctly with a space ("SEE" + "ALSO" -> "SEEALSO"). The helper now requires the junction to *cross* the band, since rescuing junctions the band would break is its only purpose; within-band pairs keep the normal word-spacing logic. The band is now a shared MERGE_FONT_SIZE_BAND constant so the helper and its caller cannot drift. 2. A trailing digit was skipped. The backward search for "the capital we are continuing" skipped non-alphabetic characters, so text ending in a footnote marker ("ANGELA M. MAZZARELLI1") found the earlier uppercase letter and accepted the join. It now checks the actual trailing character. Rejecting every trailing digit outright turned out to cost real quality, so this keeps one narrow exception. In meetings_in_mass_july_2023 the source reads "TUESDAY, JULY 4TH" with the ordinal suffix set as a smaller run; a blanket digit rejection reverted that to "TUESDAY, JULY 4" plus a stray "TH" leaking onto the next line, which is what main produces and what pdftotext shows is wrong. Only the four English ordinal suffixes (TH/ST/ND/RD) may follow a digit; anything else after one is treated as a footnote marker and rejected, so the case cubic raised stays blocked. 3. A test's name did not match its data. `small_caps_merge_does_not_swallow_a_ second_column` claimed to exercise the 72pt column gap but contained only the second column's items, so it just re-tested the happy-path merge. It now holds the full nine-item row and asserts exactly two merged results, "ROLANDO T. ACOSTA, P.J." and "ANIL C. SINGH", which genuinely exercises the gap. Both new guards were verified to be load-bearing: removing either one makes its test fail. The document that motivated the change is unaffected — small caps there run at a 0.675 ratio, far outside the band — and 199AD3d.pdf p.5 still produces the correct two-column justices table. 900 unit tests pass (10 covering small caps), fmt and clippy clean. |
||
|
|
89dd20d02c |
fix(xobjects): track text line matrix and handle T*/TL/'/"/Tc/Tw in Form XObjects (#369)
* fix(xobjects): track text line matrix and handle T*/TL/'/"/Tc/Tw in Form XObjects The Form XObject text extractor in xobjects.rs is a separate hand-rolled implementation of the operator state machine in content_stream.rs, and it had drifted well out of parity: - No text line matrix (TLM). `Td`/`TD` were applied to the text matrix already advanced by `Tj`/`TJ`, so every line began where the previous line *ended* instead of at the line start. Lines marched off the right edge and were dropped as off-page. - `T*` was not handled at all, so it never advanced to the next line. - `TL`, `'` and `"` were missing, and `TD` never set the leading as a side effect. - `Tc`/`Tw` were hardcoded to 0.0 when computing advance widths, drifting positions and inserting spurious spaces. - Text state (Tc/Tw/TL/Tf) is part of the graphics state but was not saved or restored by `q`/`Q`. This matters well beyond an edge case: producers that emit a page stream of just `q /X Do Q` and put all content in a Form XObject are common in print-to-PDF and typesetting workflows, so this parser is on the hot path for whole classes of real documents. Measured on nycourts.gov 199AD3d.pdf (1370 pages, PDFlib producer, every page wrapped in a Form XObject, 1331 pages using T*), word recall against a pdftotext reference goes from 19.2% to 97.7% — 116k extracted words to 582k against a 578k-word reference. On a 10-page subset, sequence similarity goes from 27.3% to 97.7%, against 99.8% for Mistral OCR. pdf-evals: 195 passed / 7 failed, byte-identical to the origin/main baseline with the same failure list — no regressions. Adds 7 unit tests covering the line-matrix-relative `Td`, `T*`, `TD` setting leading, `'`, `"`, `Tc` advance widths, and `q`/`Q` text-state restore, all driven through a page whose content is only `q /X1 Do Q`. * fix(xobjects): restore fill colour across q/Q in Form XObjects A white fill set inside a q/Q pair leaked past the Q, so all subsequent text was treated as invisible and dropped. Save and restore fill_is_white with the rest of the graphics state. This was the cause of several long-standing extraction failures where whole passages went missing or degraded into per-character garbage: cambridge_excerpt (+8.5KB of recovered text), MTUAeroEngines (+7.4KB), 2025_findings-acl_668 (+1.4KB), HTM_02-01_Part_A (+1.3KB), and HuttoISDWorkPerks / ebgt7isj04ophcq, which both went from exploded per-character tables to clean prose. Adds a regression test that fails without the restore (the text after Q is dropped entirely). Reported by cubic on #369. |
||
|
|
3d33ff3dbd |
fix(extractor): bound CID /W range expansion (#372)
Type0 /W parsing and the Unicode-CID heuristic expanded every CID in every range. Repeating a full-width [0 65535 w] entry therefore re-materialized the same 65,536-key domain on every copy, growing a temporary vector and HashMap work without bound. Cap expansion at the 16-bit CID domain, collect unique CIDs for the median heuristic, and stop width assignment once that many entries have been written. Legitimate compact /W arrays are unchanged. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
75e9b09593 |
fix(extractor): bound Form XObject expansion per page (#370)
* fix(extractor): bound Form XObject expansion with invocation and operation budgets Nested Form XObjects were only limited by recursion depth (5). An acyclic graph where each form invokes the next N times still expands to N^depth work, so a small PDF can force millions of nested /Do evaluations. Share a per-page FormWalkBudget that caps 10,000 Form invocations and 1,000,000 operations walked across those expansions. Extraction stops when either cap is hit. Legitimate shallow nesting is unchanged. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(extractor): share Form XObject budget across invisible-layer retry extract_page_text_items created a fresh FormWalkBudget on every call, so the invisible-text retry could consume a second full expansion budget for the same page. Own the budget at the call site and pass it into both passes. Also charge form operations independently of the invocation cap so a form that was already admitted can finish its stream. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
f65b25906c |
fix(regions): serve invisible (Tr 3) OCR text layers instead of needs_ocr (#351)
* fix(regions): serve invisible (Tr 3) OCR text layers instead of needs_ocr Scanned pages (archive.org-style digitizations) carry their text as an invisible render-mode-3 layer behind the page raster. extract_text_in_regions extracted only visible items, so such pages yielded nothing but '[Image: ...]' placeholders and every region fell back to OCR — while the markdown path already includes the invisible layer for Mixed PDFs. The two extractors disagreed about the same page. Mirror the markdown path's gate, page-scoped: when the visible pass is effectively textless (<40 non-placeholder alphanumerics), retry with invisible text included and adopt the retry only when it contributes real, non-garbage text. Pages with real visible text never retry, so double-layer PDFs (visible text plus an invisible accessibility copy) are unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review: strict zero-visible adoption gate, whole-layer garbage check, doc placement, discriminating tests - Adoption now requires ZERO visible text on the page: the invisible pass returns visible items too, so adopting alongside any visible text would duplicate it. Strict gate instead of fuzzy dedupe; the dead inv_alnum > visible_alnum condition goes with it. - Garbage check judges the whole recovered layer, not the first 200 items. - Constant/helper moved above the doc block so rustdoc stays attached to extract_text_in_regions_mem. - Tests: any-visible-text-blocks-adoption case (single short visible line) with exactly-once assertions; mode-0 guard asserts occurrence count. - Re-verified against real scanned-book pages: all recover (9.2-11.7K chars, needs_ocr=false) — pure OCR-layer scans carry no visible text. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review: gate adoption on visible-item presence, not alphanumeric mass Punctuation-only visible text (zero alphanumerics) could still adopt the invisible pass; its OCR twin in the layer would duplicate the glyphs. The gate is now item-presence: any non-image item with non-whitespace text blocks adoption (whitespace-only artifacts still tolerated). Pinned by a punctuation-only fixture test; real scanned-book pages re-verified — all recover unchanged (they carry no visible items at all). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review: punctuation-gate test pins preservation, comment states fixture scope The fixture's visible punctuation is a separate block, not mirrored in the invisible layer — so it pins the gate, not the duplication scenario. The doc comment now says so, and a positive assertion checks the punctuation survives exactly once. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review: retry only when invisible text was actually skipped; negative gate tests - extract_page_text_items now reports skipped_invisible (4th return): the Tr-3 suppression sites set it, so the region extractor retries only when a recoverable layer exists. Blank pages and image-only scans without an OCR layer — the common scanned case — no longer pay a second content-stream parse. - Negative tests: below-floor watermark layer and symbol-garbage layer are both rejected (region keeps only the raster placeholder). needs_ocr semantics for placeholder-only regions deliberately unchanged — that contract predates this PR and downstream pipelines handle it. - Real scanned-book pages re-verified: recovery unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review: flag the ' show-text suppression path; require nonempty strings - P1: invisible layers shown via the ' operator never set skipped_invisible, so those pages kept the placeholder and fell back to GPU OCR. The ' suppression path now flags too — pinned by a fixture whose layer is shown entirely via ' (nothing through Tj/TJ). - P3: all three flag sites (Tj, TJ, ') require a nonempty string operand — numeric-only TJ kerning arrays and empty shows no longer trigger the invisible reparse. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1f28c00a13 |
Bound column-detection histogram and harden coordinate handling (#328)
* Bound column-detection histogram size Derive the projection histogram from a clamped bin count and skip non-finite page widths. Extreme or malformed text-item coordinates (from the content-stream text matrix) could otherwise drive a very large allocation. 65,536 bins is ~9x the largest legal page, so real layouts are unaffected. Adds regression tests. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * Exclude non-finite coordinates from page bounds Items at NaN/inf positions are now skipped when folding the page bounds, so a malformed coordinate can no longer escape as a ColumnRegion boundary, and an all-non-finite page returns no columns. Bad items are dropped individually rather than failing the page, so one stray glyph does not disable column detection. Addresses review feedback on the finite-width guard. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * Trim far-outlier coordinates from page bounds Gutter margins, spanning-item width and the XY-cut margin are all fractions of page_width, so a single far-but-finite item (x=50_000 is enough) set the scale for the whole page: real gutters fell inside the rejected margin band and a genuine two-column page collapsed to one region. When the span exceeds one legal page (14_400 units), re-derive the bounds from items clustered around the median x. Outliers keep their text because column assignment buckets by nearest overlap. The MAX_BINS ceiling stays as an allocation bound that does not depend on this heuristic. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * Harden bounds trimming against widths and wide layouts Check both item edges when trimming: a malformed width at an ordinary position poisoned x_max just as a malformed position poisoned x_min, so a huge width still collapsed a two-column page to one region. Only trim when the far items are a small minority (<=10%). A genuinely large-format page has content spread across its full width, so it now keeps its true bounds instead of being reduced to the median cluster. Correct the MAX_PAGE_EXTENT comment: 14_400 units is the traditional Acrobat architectural limit, not a format cap. PDF 2.0 sets no page-size limit and UserUnit scales physical size, so this is a heuristic. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * Scale bin width so the histogram spans the whole page Clamping the bin count alone left anything past MAX_BINS * BIN_WIDTH (~131k points) outside the histogram, folded into the final bin. A page wide enough to hit that lost real gutters: with a visible gutter inside the covered range the XY-cut fallback never runs, so a three-column layout silently reported two. Derive bin_width from page_width instead, keeping the same allocation ceiling and degrading only resolution. Also anchor the trimming median on the same finite left/right items that bounds() accepts, so a malformed width cannot shift which items count as strays. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * Require geometric evidence before detaching far content The median-window trim narrowed any content more than one page from the centre, so a valid large page with a sparse far sidebar lost the sidebar from its bounds and its text fell into column 0. An item-count minority rule cannot tell that layout from malformed coordinates. Group content into clusters separated by more than a whole page of continuous emptiness, and only drop a cluster that is both detached by such a void and a small minority of items. Real content does not leave a gap that large; a stray coordinate sits alone beyond one. A single run wider than one page is treated as a malformed width, which also covers the huge-width case the cluster sweep cannot see (such an item spans everything and leaves no gap). Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * Judge run width against page content, not a fixed extent Treating any run wider than 14_400 units as malformed penalised valid large pages: one made entirely of such runs reported no columns at all, and a mixed page lost the right edge of every long run. Judge width relative to the page's own content instead. Positions cannot be inflated by a bogus width, so the spread of the core cluster is a sound scale: a run wider than that spread plus one page is malformed. A genuinely large page keeps its genuinely long runs, while a 1e12-wide run beside ordinary text is still rejected. Cluster on positions rather than filled intervals, so a bogus width can no longer merge everything into one cluster, and keep ordinary pages on an O(n) fast path that skips the sort entirely. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> |
||
|
|
493fed498e |
fix(links): prevent stack-overflow DoS from AcroForm /Kids self-cycle (#314)
* fix(links): guard AcroForm /Kids traversal against cycles and huge trees A crafted PDF whose AcroForm field lists itself (or another ancestor) in /Kids caused walk_form_fields to recurse indefinitely, overflowing the stack and aborting pdf2md (exit 134) — an application-level DoS from a ~730-byte input. Track visited field object IDs to break /Kids cycles, and cap total field-node traversal at 100k nodes to bound pathologically large trees. Adds regression tests for self-cycle and mutual-cycle field graphs. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * fix(links): cap AcroForm /Kids recursion depth to stop deep-chain overflow The visited-set guard stops cyclic /Kids graphs, but a long *acyclic* chain of distinct fields still recurses to the chain length and overflows the stack (a ~1.6MB PDF with 20k linked fields aborts pdf2md, exit 134) before the 100k node budget is reached. Add an explicit recursion depth cap (100 levels — far above any legitimate form hierarchy) so stack usage is bounded independently of node count. Adds a deep-acyclic-chain regression test. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * fix(links): enforce form-field node budget before insertion The node-budget guard inserted each field ID into the visited set before checking the budget, so the check triggered an early return but never actually capped the set. A field with a huge /Kids array kept inserting post-budget IDs, letting visited (memory and work) grow with the crafted input rather than stopping at MAX_FORM_FIELD_NODES. Check depth and budget before inserting, so visited can never exceed the cap. Adds a wide-tree regression test. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * fix(links): stop /Fields and /Kids iteration once node budget is spent Checking the budget before insertion capped the visited set, but callers still iterated every remaining entry of a wide /Fields or /Kids array after the budget was exhausted — each walk returned immediately, yet the O(N) sibling iteration let a single multi-million-entry array burn extraction CPU unbounded. Break out of both the top-level and recursive loops once visited reaches the cap, making the budget a true traversal-work cap. Adds a top-level wide-/Fields regression test. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * fix(links): charge examined entries against the field-node budget The budget counted only distinct visited nodes, so /Fields or /Kids arrays full of invalid (non-reference) or duplicate entries never grew visited and ran to completion regardless of size — the node budget did not actually cap traversal work. Introduce FieldWalkBudget tracking both visited nodes and total entries examined; charge every array entry (valid, invalid, or duplicate) and stop once either hits MAX_FORM_FIELD_NODES. Adds a regression test with a huge /Kids array of duplicate + null entries. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * fix(links): iterate /Fields and /Kids arrays by borrow, not clone Both arrays were cloned in full before the budget check, so a crafted oversized /Fields or /Kids array forced an O(n) allocation and copy regardless of the cap. resolve_array already returns a borrow tied to the document and the walker only needs a shared &Document, so iterate the borrowed arrays directly — the early break now bounds how many entries are even touched, before any per-array allocation. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> * docs(links): correct wide-array test comments to match range assertions The two wide-array tests assert item counts within a range near the budget, not an exact value (charging entries in the entry guard shifts the boundary by one or two). Fix the stale comments that claimed exact counts. Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com> |
||
|
|
12e9a655e3 |
fix(extractor): supply built-in metrics for non-embedded base-14 fonts (#241)
* fix(extractor): supply built-in metrics for non-embedded base-14 fonts PDFs may legally omit /Widths for non-embedded standard fonts (Times, Helvetica, Courier, Symbol, ZapfDingbats) — the spec requires the reader to supply the metrics. We returned None, so every glyph advanced 0 and each text item got width 0, silently breaking every gap-based heuristic downstream: space synthesis, sub/superscript detection, table column detection, heading merging. - src/extractor/base14.rs: Adobe Core-14 AFM width tables keyed by Unicode char, plus the standard Symbol/ZapfDingbats encoding vectors (their glyphs sit at byte positions unrelated to Latin text, so widths must resolve through the built-in encoding, not cp1252) - Width resolution order: Differences -> built-in encoding -> the same cp1252-style fallback the text decoder uses, so a code's advance always matches the character we emit for it - Type3 visual sizing: PK bitmap fonts (dvips) use FontMatrix [1 0 0 -1 0 0] with nominal sizes like 0.12pt; scale by FontBBox height x |matrix_y|. Applied in the page-stream and Form XObject paths. Indirect numeric array elements are resolved before use. Effect on Shannon's 'A Mathematical Theory of Communication' (1998 dvips/Distiller, the reported case): glued sentences 95 -> 5. Corpus impact: 12 of 184 eval documents, e.g. Data-Processing-Agreement recovers a paragraph that a phantom table had shredded into cells. Layout heuristics tuned on the same document (indent-based paragraph breaks, heading reclassification, table script filtering) are held back for a separate PR — they change ~98 further documents and need to be justified against the corpus, not against one PDF. * review: narrow Type3 rescaling to self-inconsistent fonts; dedup + test all width tables Addresses cubic review on #241, plus a follow-up from a local cubic run. - Type3 visual scaling was applied to every Type3 font whose FontBBox height x |matrix_y| deviated >5% from 1.0. FontBBox is the glyph box, not the em box, so a conventional 1/1000-matrix font with a descender..ascender bbox (~700 units) computed 0.7 and had every reported size shrunk by 30% — corrupting the drop-cap, heading-tier, sub/superscript and table heuristics this is meant to fix. First attempt gated on the matrix being unit-scale, but a local cubic run pointed out that wrongly excludes valid non-standard matrices (a 0.005 matrix with a full-em bbox legitimately needs a 5x scale). The product is the right discriminator, not the matrix: a self-consistent font lands near 1.0 because the matrix is the reciprocal of the glyph-space em, so only a wildly inconsistent one (dvips/PK bitmap fonts sit at ~159) is renormalized. Band widened to [0.25, 4.0]. Corpus effect: 12 -> 7 documents change. The 5 that drop out were being wrongly rescaled — including Data-Processing-Agreement, whose phantom-table fix turned out to come from this bug rather than from the width fallback, so it is correctly given up. - base14: all 14 width tables now covered by the sort-invariant test via an ALL_TABLES registry, not a hand-picked subset. - base14: identical tables share one static (all four Courier variants are monospace 600; the oblique Helvetica variants match their upright forms), removing 5 duplicate copies. * test: refresh Shannon snapshot after merging main CI checks out a merge of the PR head with main, and main advanced 8 commits since this branch was cut — including #201 (contextual digit runs), #240 and #253 (markdown fixes). Those change extraction output, so a snapshot generated on the unmerged branch could not match; the Test job failed on the merge commit while passing on the branch itself. The merged behaviour is better: the footnote marker '2' before 'Hartley, R. V. L.' is now recovered instead of dropped. 950 tests pass on the merged tree, clippy clean. |
||
|
|
bfd6c3eabb |
fix(extractor): make the comment stripper escape-aware (#259)
strip_pdf_comments tracked parenthesis nesting to protect string literals, but ignored backslash escapes. An escaped \) desynced the depth counter, after which a % glyph inside a string was stripped as a top-level comment, corrupting the stream for Content::decode and silently truncating the page's text. Treat \ inside a string literal as escaping the next byte, so \(, \), and \\ never touch the nesting depth. |
||
|
|
04abab951f | Support password-protected PDF item JSON extraction (#245) | ||
|
|
7747b3a086 |
fix(markdown): reject non-text strikeout rules (#253)
* fix(markdown): reject non-text strikeout rules * fix(markdown): address strikeout ownership edge cases * fix(markdown): reject connected filled strike rules * fix(markdown): group drifted strikeout runs |
||
|
|
3fb545284b |
fix(layout): preserve contextual digit runs (#201)
* fix(layout): preserve contextual digit runs * fix(layout): harden page folio filtering * fix(markdown): distinguish folios from contextual numbers * fix(layout): distinguish running folios from contextual digits * fix(layout): tighten running folio evidence * test(layout): guard the folio evidence floor * fix(layout): preserve page-number decisions across partitions * fix(layout): harden document-level folio filtering * fix(markdown): carry folio context across public APIs * fix(layout): isolate folios from layout metadata * fix(markdown): preserve table cells in per-page extraction * fix(layout): preserve folio context with page filters * fix(layout): handle contextual folio sequences * fix(layout): tighten adjacent folio evidence * fix(layout): constrain folio context inference * fix(forms): resolve widget pages from annotations |
||
|
|
b31e4b1727 |
fix(extractor): don't flag gid Differences names covered by ToUnicode (#186)
* fix(extractor): don't flag gid Differences names covered by ToUnicode Pages were marked as having unresolvable gid-encoded fonts whenever any font's /Differences array used gidNNNN glyph names, and when every page carried such a font the whole document's markdown was suppressed. LibreOffice exports do exactly this: subset fonts get /gidNNNN names in Differences alongside a complete ToUnicode CMap that decodes them, so ordinary text documents lost their entire markdown output even though extraction decoded every glyph. Track the character codes behind the gid names and only flag the font when its ToUnicode CMap addresses none of them. Partially mapped codes stay unflagged: an emoji ZWJ sequence maps whole on its first code, and the remaining component-glyph codes are subset leftovers, not damage. Fonts without ToUnicode, or whose CMap ignores the gid codes, are flagged as before, and the downstream garbage/encoding checks still catch partial breakage. * fix(extractor): require a usable ToUnicode mapping to clear the gid flag A mapping to U+FFFD (or an empty string) is rejected by extraction as an invalid CMap result, so it must not count as decodable when deciding whether gid-named Differences codes are resolvable. |
||
|
|
5b287341a0 |
feat(wasm): add browser bindings (#180)
* feat(wasm): add browser bindings * fix(wasm): address review feedback * fix(wasm): preserve numeric plain text * chore(wasm): prepare 0.1.2 release |
||
|
|
64a0930f9f |
feat(layout): order image-anchored regions (#174)
* feat(layout): order image-anchored regions * fix(layout): preserve region flow boundaries * fix(layout): gate image-backed column flows |
||
|
|
c908e33b39 |
fix(fonts): remap misnamed TeXCMMathsSymbols glyphs to their true symbols (#165)
IntechOpen-family academic PDFs embed Computer Modern math symbol subsets whose glyphs are misnamed after Latin lookalikes (equal → /onequarter, plus → /thorn, parens → /eth //Thorn) — and the generated ToUnicode faithfully propagates the wrong names, so formulas decode as 'S ¼ kB þ 1' instead of 'S = kB + 1'. Remap the observed misnames, gated strictly on the TeXCMMathsSymbols base font (subset prefix stripped) so genuine fractions and thorns in text fonts are untouched. Known limitation: a sibling subset misnames the slash as /onequarter too, so an occasional '/' renders as '=' — still strictly better than the previous mojibake. Affects 4 bench PDFs (028/031 +0.001-0.004 NID) and zero pdf-evals snapshots. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4aa2c0c208 |
fix(headings): rescue wrapped bold headings on interleaved column pages (#164)
* fix(headings): rescue wrapped bold headings on interleaved column pages
Report pages whose columns can't be detected (6pt gutters) interleave
both columns' lines, which breaks every whitespace signal the bold
heading heuristic relies on: para_threshold inflates to ~3x line
height and a wrapped heading's own internal line gap defeats isolation.
'9.5. Adapting to the New Normal: Changing / Business Models' merged
into the following paragraph.
Four changes:
- merge_wrapped_bold_heading_groups: 2-3 consecutive all-bold
body-size lines merge into one line when the group is isolated
(column-locally, judged by x-overlapping lines only) or starts with
a section number.
- Section-numbered all-bold lines ('9.5. ...') classify as headings
without the standalone/isolation score gate.
- Line unfusing extends to uppercase-start continuations, gated on a
bold-style mismatch between the runs (a bold heading beside regular
body text) — same-style label rows stay joined.
- The unfuse line-side wordiness requirement drops to 2 words so a
wrapped heading's short last line ('Business Models') still splits
from the neighboring column.
opendataloader-bench: overall 0.8554 -> 0.8576, MHS 0.769 -> 0.777;
docs 037 +0.161, 111 +0.157, 039 +0.091, 198 +0.028, none down.
pdf-evals: 63 snapshots, composite 0.5952 -> 0.5964, sole >0.02 mover
positive. thermo-freon12 snapshot regenerated (cosmetic churn on an
already-scrambled 3-column legend).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(headings): review follow-ups — multi-component section numbers, wholly-bold line gate
Single '1. ' prefixes are ordered list items and no longer bypass
isolation; the uppercase unfuse requires the whole line bold (a
heading), not merely its last run, so mixed bold-label/value rows
stay joined.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
31918ff62f |
fix(layout): unfuse independent column runs sharing a baseline (#163)
* fix(layout): unfuse independent column runs sharing a baseline
Two-column report pages with charts fused headings into the adjacent
column's body text: the columns' ~6pt gutter is below what histogram
valley detection can safely use, so the page grouped single-column and
same-baseline items from both columns joined into one line ('6.2.
Expectations for Re-Hiring Employees' + mid-sentence text, killing
MHS and NID on the whole survey-report doc family).
Three changes:
- Line grouping splits same-baseline runs separated by a wide void
(>3x font size, >=30pt) when the incoming run starts lowercase
(mid-sentence continuation from another column) and both sides are
multi-word prose. TOC page numbers, dot leaders, and table cells
(numbered/capitalized) stay joined.
- Column detection is blind to chart-region text (tight 2pt bounds —
wider padding ate rows adjacent to charts), via a chart-aware line
grouping variant wired from the markdown pipeline.
- validate_and_build_columns computes its vertical span from
histogram-eligible items only, so full-width captions no longer sink
the overlap ratio for partial-page column regions.
opendataloader-bench: overall 0.8532 -> 0.8554, MHS 0.761 -> 0.769;
doc 038 +0.434, no regressions. pdf-evals: 34 snapshots change,
semantic composite wash (0.5749 -> 0.5748), no per-doc mover >0.015.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(layout): review follow-ups — chart-aware band-split grouping, single chart scan per page
Band-split pages now route through the chart-aware grouping too, and
the band loop reuses the precomputed page_chart_map instead of
re-scanning the rect list per page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
7d108cdff1 |
fix(extractor): use center-y for off-box link filtering (#161)
Link items carry an annotation rect, so y is a box edge — unlike text items, where y is a baseline. Testing rect-bottom dropped partially visible links whose bottom edge dipped past the tolerance. Follow-up to a #160 review comment. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c4cbf49f44 |
fix(extractor): clip page content to the visible page box (#160)
* fix(extractor): clip page content to the visible page box Single-page extracts and imposed spreads keep neighboring pages' content in the stream, positioned outside the CropBox. Extracting it appends invisible sections to the page, scrambles NID, and poisons font statistics (heading tiers built from off-page text). Clip items (by center), and — only when off-page text was actually found — rects and lines (by overlap) to CropBox-else-MediaBox, walking page-tree inheritance. Rotated pages are left unclipped: their item coordinates are already transformed out of box space. Degenerate boxes (<1 inch) are ignored. opendataloader-bench: overall 0.8445 -> 0.8537, NID +0.008, MHS +0.013; six docs up (best +0.426), none down. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extractor): guard page-box clipping with coherence and straddle checks Two real-document counterexamples: curved display text leaves short glyph fragments with artifact coordinates outside the box (judge by character mass, not item count), and some PDFs compute inflated coordinates for visible body text (an off-page item continuing an on-page baseline means our transform model is wrong there — skip). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extractor): clip off-box link annotations when page text was clipped Review follow-up: annotations from the neighboring page bypassed the filter. Form fields are left as-is — they're document-scoped and rare on imposed spreads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
39c31a8404 |
fix(underline): rescue snug-owned underlines from the table-ruling filters (#143)
* fix(underline): rescue snug-owned underlines from the table-ruling filters Documents that underline many full-width lines (dense CJK business docs, legal redlines, 10-K section links) produce span-similar rules at 3+ y-levels — exactly what the repeated-ruling filter treats as table rulings, so every semantic underline on such pages was discarded. Three changes fix detection without re-marking real tables: 1. Snug-owner rescue: a rule survives the repeated-ruling filter when the union of touching text runs on its baseline row owns it (rule contained within the union's span +0.75em, runs cover >=60% of the rule, no column-sized gaps between runs). Table row separators fail ownership: they overshoot their cells' text or match gapped items. Same-row segmented rules (column-header separators) always stay discarded, and a rule enclosed by a drawn cell-sized box (rect-grid tables) is never rescued. 2. Vertical window widened 0.35em -> 0.72em below the baseline: CJK layouts draw underlines under the full em box, measured at ~0.67em. 3. Prose-table guard in the positions-path suppressor: a detected 'table' whose cells hold flowing prose (>=30% of cells over 100 chars) is a detection artifact of boxed callouts + stacked rules, not a real table — suppressing there erased every underline on the page. Also fixes cluster_x_positions fabricating phantom table columns from style-split continuation runs (touching items, gap <2pt, now feed one column start) — the fix that keeps rect-grid table shapes stable while underlined links inside cells are correctly marked. Snapshot updates are underline gains on regulation/form fixtures and one empty spacer-column change in a subscripted header. Corpus (508-doc public bench sweep): text output byte-identical on all docs; underlined items +224/-0; strikeout now fires on redline docs. Item-level GT coverage: is_underline 86->151/405, is_strikeout 0->10/44. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB * fix(underline): fraction-bar guard + subscript merge across underline marks Found by a 202-doc real-world corpus diff (pdf-evals) that exercises the full markdown pipeline, which the bench-corpus item sweep does not: 1. Math fraction bars and lattice grid lines are underline geometry — short horizontal rules under digits. Guard: a narrow rule (<=60pt) with bar-sized text hanging just below it (denominator) never marks. The below-text width bound matters: tightly-leaded REAL underlines have a full-width next line below, which must not trip the guard. 2. merge_subscript_items refused to merge when the parent was underlined but the tiny digit was not (the drawn rule easily misses the digit's own overlap window) — losing the merge broke subscript tokens inside table cells (b+2 no longer became b₂). Strikeout boundaries still block the merge in both directions; only parent-underlined/digit-bare merges, absorbing with the parent's flags. Corpus after refinement: underlined items +220/-2 (the 2 are fraction bars the old code wrongly marked), GT rule-text coverage 149/405 underline + 10/44 strikeout, text output identical on all 508 bench docs and word-count-identical on the 202 pdf-evals docs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB * fix(underline): address review — grid-evidence veto, strikeout-safe fraction guard, bounded gaps - Cell-box veto now requires GRID EVIDENCE (a vertically abutting neighbor rect with x-overlap) instead of a height window: multiline table cells taller than the old 90pt ceiling veto again, and isolated filled callout panels (which legitimately contain underlines) no longer veto at all. - The fraction guard gates only UNDERLINE marking; rule_strikes_item still evaluates, so short strikeouts near lower text survive. - Fraction hug distance tightened to 0.3em so a short last-line at normal leading is not mistaken for a denominator. - Continuation-run suppression bounds the negative gap (-4pt): text overhanging from an adjacent cell keeps its own column start. Corpus after review fixes: underlined items +222/-2, GT coverage 150/405 underline + 10/44 strikeout, bench text output identical on all 508 docs, pdf-evals word loss bounded at equation-reflow noise. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB * chore: appease clippy (redundant closure) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fed3b90d37 |
feat(extractor): run-local space floor for tracked (letter-spaced) glyph runs (#133)
* feat(extractor): run-local space floor for tracked (letter-spaced) glyph runs
Display type set with tracking renders one glyph per show op; the merge
loop's fixed space thresholds (0.08-0.13 em) then read every letter gap
as a word boundary and emit "H O W" / "F U R T H E R" instead of
"HOW" / "FURTHER". The page-level Canva fixer can't help: it requires
>=50% of the page's items to be letter-spaced, and these docs track
only their display headings.
merge_text_items now pre-scans each run of consecutive single-glyph
items (same size band, same style, mergeable gaps — the loop's own
break conditions) and, when the run is tracked, derives the space floor
from the run's own gap distribution:
- runs with >=4 gaps qualify when the median gap clears the fixed
threshold; word gaps, if present, form a second mode — split at the
largest relative jump (>=1.4x), else the run is a single word
("I T I S I M P O R T A N T" -> "IT IS IMPORTANT")
- short runs (2-3 gaps: "H O W") additionally demand uniform gaps and
ALL-CAPS or CJK — a genuine spaced sequence of single letters
("x y z" variables) has the same gap count, and display tracking is
a caps convention; CJK never wants inter-glyph spaces
Corpus sweep (708 opendataloader + ParseBench text PDFs) vs main: 9
docs change — the tracked display titles ("HOW CAN YOU HELP?",
"LUNCHTIME MENU", a tracked email address), and CJK glyph-per-item
docs whose spurious inter-glyph spaces now collapse (GT for those docs
is unspaced CJK; should_join_items already treats no-space CJK as
correct on its path). No other doc moves.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(review): convention gate on both tiers; Han/Kana floor always infinite (PR #133 review)
- Long lowercase spaced-single runs ("a b c d e") had the tracked gap
shape in the >=4-gap tier with no convention guard — word boundaries
lost. The caps/CJK/title-case gate now applies to BOTH tiers; a
title-case single word ("B u f f a l o") also qualifies.
- Han/Kana runs skipped straight to the bimodal split, so a nonuniform
gap distribution (justification, punctuation spacing) could
manufacture a word boundary. Han/Kana now always floors at infinity;
Hangul deliberately keeps word-boundary handling — Korean spaces
between words (is_spaceless_cjk excludes it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(extractor): preserve mixed-case glyph boundaries
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
57335f8bcf |
feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection (#125)
* feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection
Two style-recall gaps, both invisible to the existing name-based
heuristics:
1. Subset fonts with opaque BaseFont names ("Tc1", "AAAAAB+Amplitude")
defeat is_italic_font/is_bold_font. New descriptor_style_flags reads
the FontDescriptor (ItalicAngle beyond 4 degrees, Flags bit 7 Italic,
bit 19 ForceBold) and, when the descriptor claims upright, falls back
to the embedded font file: ttf-parser's OS/2 fsSelection + post
italicAngle for sfnt fonts, and the CFF Name INDEX PostScript name
for bare-CFF FontFile3 (descriptor rewritten to ItalicAngle 0 while
embedding "Amplitude-LightItalic" was observed in the wild).
ORed into is_bold/is_italic at item creation (content streams and
form XObjects).
2. No strikeout signal existed. New is_strikeout on TextItem, detected
in the same pass as underline: same rules pipeline (stroked lines /
thin filled rects, table-ruling suppression), different vertical
window — a rule crossing the glyphs at 12-55% of the em above the
baseline instead of sitting at it. Exposed through napi and python
bindings and pdf2md --items-json.
Verified on public ParseBench corpus docs: previously-missed italic
council titles and bold CJK itinerary headings now flagged (render-
checked); 24/508 docs gain flags, none lose any; 35 strikeout items
detected corpus-wide, disjoint from underline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(review): quote-op advance width, Ts text rise, doc-level font style cache (PR #125 review)
Address three valid findings from review:
- The ' (move-to-next-line-and-show-text) operator emitted zero-width
items and never advanced the text matrix, so geometric underline/
strikeout detection (which requires width > 0) could never mark its
text, and following show ops overlapped it. Reuse Tj's advance-width
computation and matrix advance.
- Ts (text rise) was dropped entirely: raised/lowered runs kept the
unshifted baseline, so rules drawn at the risen glyph position missed
the strike/underline windows. Track rise in the text state (saved and
restored with q/Q) and shift the rendering position through the text
matrix's y column; advances stay on the unshifted matrix per spec.
- descriptor_style_flags re-decompressed and re-parsed the same embedded
font program on every page whenever the descriptor left a style flag
unset (the common case). Add a document-scoped FontStyleCache keyed by
the FontFile2/FontFile3 object id, threaded through page and form
extraction alongside the existing CMapDecisionCache.
The fourth finding (Form XObject rules never reach geometric detection)
is real but pre-existing for underline and needs the form walker to grow
path/paint tracking plus a new return type; deferred as a follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(review): ActualText items render at their glyphs' text rise (PR #125 review)
The EMC-built ActualText item used the captured text matrix without the
rise adjustment the ordinary Tj/TJ/' emission sites apply, so a tagged
run shown with Ts landed on the unshifted baseline — off the strikeout/
underline windows and inconsistent with untagged runs. The rise is
captured together with the first-glyph matrix (and at BDC for the
entry-position fallback): the item must render at the rise of its
GLYPHS, not whatever rise is set by EMC time.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(review): capture ActualText glyph position after the quote op's line move (PR #125 review)
The `'` handler skipped the entire suppressed-extraction block, so a
tagged span whose show op is `'` never captured its glyph matrix/rise —
the EMC item fell back to the BDC-entry matrix, which sits on the
PREVIOUS line (the `'` line move happens after BDC) with no rise. The
capture now happens right after the line move, matching the Tj/TJ
paths.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(review): style-boundary gate on subscript merge + strikeout suppression coverage (PR #125 review)
merge_subscript_items absorbed a script digit into its parent
regardless of underline/strikeout flags — dropping the digit's own mark
or widening the parent's over it. The merged item carries one flag, so
differing marks now break the merge, mirroring merge_text_items'
style-boundary rule (pre-existing for underline as well).
Also extends the table-suppression test to assert is_strikeout is
cleared alongside is_underline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
b375d6f102 |
feat(markdown): underline emission, Unicode scripts, style-preserving merges (#117)
* feat(markdown): underline emission, Unicode scripts, style-preserving merges (ENG-5015 2b)
Three formatting losses in the direct-extraction markdown path:
1. text_with_formatting gains <u> run emission (detect_underline option,
default on) using the geometric is_underline flag from 1.9.9.
Underline runs stay free of nested bold/italic markers — consumers
match tag content literally. Heading lines keep plain text for
bold/italic but preserve <u>: the tag carries meaning `#` doesn't.
2. merge_subscript_items now maps absorbed digit scripts to Unicode
sub/superscript forms with direction from the baseline offset
("H"+"2" -> "H₂", "word"+raised "2" -> "word²", "m"+"3" -> "m³").
NFKC/NFKD folds these back to plain digits so text matching
downstream is unaffected; renderers keep the script semantics.
3. merge_text_items no longer merges across bold/italic boundaries —
absorbing a styled run into a plain neighbor erased the styling
before markdown emission ever saw it. On eval docs this recovers
20-82 italic runs per document that previously emitted as plain.
Snapshots regenerated (diffs are the features: CCl₂F₂, m³, underlined
legal section headings, finer bold runs). pdf-evals regression suite:
202/202 real PDFs pass. napi 1.9.9 -> 1.9.10.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(extractor): break merges at underline boundaries too (review)
OR-merging underline stretched the eventual <u> span over neighboring
plain fragments. Merge runs now break on any style-flag change, the
redundant accumulator is gone, and format_list_item learned to move
bullet markers outside <u> wrappers so fully-underlined bullet lines
still render as markdown lists. td9264 snapshot regenerated — spans are
tighter (trailing periods correctly outside the tag).
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(markdown): strip stray spaces before sentence punctuation (review)
Style-boundary item splits can strand a trailing period in its own
fragment, and multiple assembly paths join fragments with spaces,
yielding "word ." artifacts. Rather than chasing every join site, a
postprocess pass removes a space before `.`/`,`/`;` when the mark ends
its token (whitespace, cell boundary `|`, or end of text follows).
Dot leaders/ellipses and mid-token periods are untouched.
Fixes the td9264 "companies ." artifacts and two pre-existing
"armoring ," artifacts in the 2013-app2 snapshot. pdf-evals: zero
markdown diffs across all 203 corpus PDFs vs committed baselines.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(tables): trim spaces inside parenthetical cell fragments
* fix(tables): reject sparse prose row-stripe tables
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
|
||
|
|
422a2ff118 |
feat(extractor): geometric underline detection on TextItem (#116)
* feat(extractor): geometric underline detection on TextItem (ENG-5015) PDFs carry no underline font flag — underlines are stroked horizontal lines or thin filled rects drawn under the baseline. Correlate those graphics (already parsed from the content stream) with text items in a post-pass: a rule within ~0.35em below the baseline covering >=60% of an item's width marks is_underline. Exposed through the napi and python bindings. Verified on real docs: 4/4 underlined sentences flagged on a Japanese report, links/headings flagged on 8 of 10 underline-bearing eval docs, zero flags on docs without underlines. Known FP source (table cell borders) documented — downstream applies inline styling only to plain-text regions. napi 1.9.8 -> 1.9.9. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(extractor): underline rules only from painted rects, normalized extents (review) Two review fixes: (1) normalize rect extents before the thickness/width checks — `re` operands pass through the CTM so width/height can be negative, which missed negative-width rules and let negative-height bands pass as thin; (2) only feed painted rects to underline detection — `re` rects now wait in a pending list until a paint operator (S/s, f/F/ f*, B/B*/b/b*) confirms them, and `re W n` clip-only paths are discarded at `n`, so invisible clip boundaries no longer underline nearby text. Marking moved into content_stream where paint state lives (pre-rotation, consistent device space). Co-authored-by: Cursor <cursoragent@cursor.com> * fix(extractor): harden underline detection * feat(cli): export positioned text item json --------- Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
30eddade77 | fix(extractor): preserve tagged overlapping text order (#114) | ||
|
|
1b2e2c76d6 | fix(extractor): make trace previews unicode-safe (#113) | ||
|
|
ce49794719 |
fix(extractor): reduce garbled OCR false positives (#112)
* fix(extractor): reduce garbled OCR false positives * fix(extractor): tighten garbled text OCR routing * fix(extractor): decode UTF-16 ToUnicode destinations * fix(extractor): narrow ToUnicode destination cleanup * fix(extractor): decode Aptos private ff ligature |
||
|
|
f25808e0a7 |
fix(extractor): restore CID font state for Chinese text (#106)
* fix Chinese CID text decoding * bump package versions |
||
|
|
85890648c9 | chore: use crates.io lopdf (#101) | ||
|
|
8b63ceb084 |
emit ItemType::Image bboxes for Image XObjects (was: silently dropped) (#94)
* emit ItemType::Image bboxes for Image XObjects (was: silently dropped)
Background. ItemType::Image, MarkdownOptions::include_images, and the
markdown emitter's image-collection path have all been in the tree
for a while, but no producer ever populated them — content_stream.rs
explicitly `// Skip images — text extraction only` at the Do
operator, and the nested Form-XObject walker in xobjects.rs only
matched XObjectType::Form, silently dropping Image entries. The
declared types were dead code.
This PR lights them up. At every Do that resolves to an Image
XObject (both top-level and nested inside Form XObjects), we now
compute the page-space bbox from the current CTM via a new
`image_bbox_from_ctm` helper — handling both axis-aligned and
rotated/sheared placements via 4-corner AABB — and emit a TextItem
with `item_type: ItemType::Image` and the legacy `[Image: <name>]`
text payload that the markdown emitter already knows how to render.
Callers can now find raster figures via `extract_text_with_positions`
(and the `_mem` variant, newly re-exported at the crate root) without
needing to re-parse the PDF or run a vision/layout model. The intended
consumer is layout-aware text pipelines that want to crop figures and
caption them out-of-band.
Two backstops to avoid silent breakage for existing callers:
1. `MarkdownOptions::include_images` default flipped `true → false`.
If it stayed at `true`, every existing user of
`extract_pages_markdown` would suddenly see ``
placeholders inserted throughout their output the moment they
upgraded. Image data is still available structurally via
`extract_text_with_positions`; rendering it into markdown is now
an opt-in. New regression test asserts `extract_pages_markdown`
output is unchanged for the image-bearing fixture.
2. Image items now also skip the layout heuristics
(`detect_columns`, `detect_tables_from_rects`) via a new
`is_text_layout_item` predicate. Without this filter, an image's
left edge would land in the column-projection profile and skew
table column detection — surfaced by
`vector_grid_tests::upstage_key_functions_four_cols` going from 4
detected columns to 5 in CI before the filter was added.
Re-exporting `extract_text_with_positions_mem` at the crate root —
strictly additive; mirrors how `extract_pages_markdown_mem` is already
available there.
Tests:
- test_extract_text_with_positions_emits_image_bboxes — minimal PDF
with one 200×100 image at (50, 600); asserts one Image item with
correct bbox + page + text.
- test_image_xobject_bbox_handles_rotated_ctm — 90° rotated image
via shear-component CTM; asserts AABB is correct (handles non-
axis-aligned placements via 4-corner clamp).
- test_image_emission_does_not_change_default_markdown — asserts no
`Image:` token leaks into default markdown output, regression
guard for the include_images flip.
- test_markdown_options_default_has_include_images_false — explicit
sentinel so anyone flipping it back catches it in CI.
* Bump version from 1.8.15 to 1.9.0
|
||
|
|
59b17f372a |
tables: tighten prose-in-frame rejection (#77)
* tables: lift detection on shaded-header + alt-row tables (#wired-grids) Production telemetry on `wired_high_confidence`-classified table regions showed `detect_vector_grid_in_region_mem` returning a usable grid only ~27% of the time, with the rest falling through to GLM-OCR. Three surgical fixes target the dominant production shapes: * Path-fill cell backgrounds: when the page has no `re` rects but draws cell backgrounds via `m`/`l`/`h`/`f*` sequences, prefer the fill-derived rects over the few section-level `W*` clip paths that previously won the priority gate. Activated when fill rects outnumber clip rects ≥3×. * Dedup-induced cluster splits: page-background rects could pose as containers in the sub-rect dedup and evict a slightly smaller table-frame rect, breaking adjacency between column-cell groups so each column became its own cluster. Origin-anchored containers are now disqualified from sub-rect dedup. A separate exact-duplicate pass collapses the cell-padding/text-bg/cell-border triple emissions some PDFs produce, preserving original order to avoid reshuffling table output on multi-table pages. * Prose-words rejection: the `cell-rect` fallback's whole-grid prose threshold also rejected real tables that include a description column. Now relaxed when content is well-distributed (≥75% of cols filled), while keeping the original strictness for prose-in-a-frame layouts. Two regression fixtures from the opendataloader-bench corpus, covering the dominant production failure categories: * `greencomp_competence.pdf` — 2-col shaded-header + plain-body glossary. Mirrors production crops #1 (Contractions glossary) and #6 (BIO 350 course header). * `upstage_key_functions.pdf` — 4-col shaded-header + alt-row backgrounds + merged left column. Mirrors production crops #2 (Parameter/Value alt-row), #7 (Spanish XML schema), and #8 (Córdoba multi-row header). Existing fixtures stay green (doc 51 wrapped-label, doc 128 forecast six-cols, td9264 snapshot). 133 unit + integration tests pass; clippy clean. Bumps napi/package.json 1.8.4 → 1.8.5. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * tables: tighten prose-in-frame rejection — fixes pdf-evals #30 regression PR #76's shaded-header detection lift surfaced a regression on accessory_building_permit_application_1 (TEDS 0.10 → 0.05): a paragraph of legal text laid out in a 2-column justified block was being admitted as a 10×2 fake table where every cell holds a sentence fragment ("I agree to comply...", "I", "It is the property owner's responsibility..."). Per pdf-evals PR #30 review, this is the kind of regression production users will notice — the markdown is structurally and semantically misleading. Root cause: PR #76's prose-rejection only fires for `num_cols >= 4`, so the 2-col prose-in-a-frame case slipped past it entirely. The new fill-priority + dedup changes started producing rects for this layout that 1.8.4 correctly ignored. Fix: tighten the prose-in-frame check. - Lower the column-count guard from `>= 4` to `>= 2`. - Add a content-length signal as the primary discriminator: when the prose-words trigger fires AND mean non-empty cell length exceeds 65 chars, reject regardless of column distribution. The 65-char threshold cleanly separates observed cases: accessory_building (prose-in-frame): mean 74 chars → REJECT upstage_key_functions (real 4-col table): mean 53 → admit greencomp_competence (real 2-col glossary): mean 20 → admit accessory_building (real 5×3 form data): mean 10 → admit The well-distributed-cols relaxation that PR #76 added stays — "label / value / description / benefit" tables (#7, #8 from the production crops) still pass, but only when their mean cell length stays below the prose threshold. New regression test `accessory_building_rejects_prose_in_frame` asserts both that the real 5×3 form data table survives AND the 10×2 prose block is rejected. Snapshot test `test_snapshot_td9264` updated to match new output — old snapshot captured the same prose-in-frame bug on regulatory text (paragraphs emitted as 3-col `||text||` fake-table rows). New snapshot emits clean prose paragraphs, which is correct. Verification: - cargo test --all: 424 lib + 133 integration + 2 doc tests pass - cargo fmt --check clean - cargo clippy -- -D warnings clean (lib-level; pre-existing test-level clippy issues on the wired-grids branch unaffected) - Existing fixtures stay green: forecast_table_chart_six_cols (PR #72), bits_pilani_* (PR #73), greencomp_competence_two_cols and upstage_key_functions_four_cols (PR #76). This branch is based on abimaelmartell/wired-grids so it includes PR #76's commits plus this fix on top. Suggest merging this and closing #76, OR rebasing #76 to incorporate this fix. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * tables: drop early-dedup atom — caused broad TOC + matrix corruption Bisected PR #76's 4 atoms against the SEC 10-K 0001104659-25-093871_183e44ac.pdf which appeared as a TEDS regression in pdf-evals PR #30. Result: atom | TOC | perf-graph | qualifications ---------------------------|-----|------------|--------------- fill-priority | ✓ | ✓ | ✓ early-dedup | ✗ | ✗ | ✗ page-bg disqualification | ✓ | ✓ | ✓ prose-relaxation | ✓ | ✓ | ✓ Early-dedup was the SOLE source of all three regressions on this doc. Tried a more conservative variant (≥3 copies only — pair-duplicates appear in legit multi-section layouts like 10-K dividers above + below section headers); didn't fix the regression. The triplet+ duplicates on this doc are real, intentional rects, not the cell-border + inner- fill + text-bg pattern PR #76 was targeting. Drop early-dedup. Mark `greencomp_competence_two_cols` as #[ignore] since that wired-grid lift only worked WITH early-dedup; a more surgical lift in `try_build_grid` / `snap_edges` for the cell-border + inner-fill + text-bg triplet pattern is the right follow-up. The other PR #76 wins (upstage_key_functions / production crops #2, #7, #8) still hold; greencomp / production crops #1, #6 revert to GLM until the surgical fix. Validation on the regression doc: 0001104659 TOC PART II markers: 4 (matches main, was 2 with PR#76) 0001104659 perf-graph data row: 2 (matches main, was 1) 0001104659 qualifications rows: 9 (matches main, was 6) Validation on the prose-frame doc: accessory_building fake-table: 0 (matches main, was 1 with PR#76) accessory_building prose intact: 1 (matches main) cargo test --all clean, cargo fmt --check clean, cargo clippy --lib -- -D warnings clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
cfc080f79a |
extractor: drop Latin-1 mojibake on Type0/CID fonts; tokenize wide TSR items (#75)
* extractor: drop Latin-1 mojibake on Type0/CID fonts; tokenize wide TSR items Two text-extraction failure modes surfaced by table-candidate shadow data; both also affect the existing TableFormer / vector-grid paths since they share `extract_tables_with_structure_*_mem`'s downstream cell-fill. 1. CJK / multi-byte mojibake. The bottom Latin-1 fallback in `extract_text_from_operand` ran unconditionally. For a Type0/CID (Identity-H) font whose ToUnicode CMap fails to parse, the bytes are CIDs (font-internal indices), not character codes — per-byte Latin-1 produces mojibake (e.g. 2-byte CID 0xCDD9 surfaces as "ÍÙ"). Gate that fallback on `FontWidthInfo.is_cid` (set by `parse_type0_widths` for `/Subtype /Type0`). For Type0 fonts with any non-ASCII byte, emit one U+FFFD per CID instead so `detect_encoding_issues` still trips and the page is flagged for OCR — preserving the existing OCR-routing path that the high-Latin-1 garbage used to satisfy by accident. Type1 / TrueType simple fonts retain the per-byte Latin-1 round-trip (it IS the canonical interpretation for them; verified against an existing pdf-evals fixture where bytes like 0xB6 are legitimate Latin-1). Threaded `font_widths: &PageFontWidths` through `extract_text_from_operand` and its 5 call sites in `content_stream.rs` / `xobjects.rs`. 2. Dense-cell text collapse in `extract_tables_with_structure_cells_mem`. Stage-1 routing did per-item assignment — each TextItem went into the single cell whose bbox contained its center. When a row's text is rendered as one wide Tj (e.g. "Marshall Islands 0.9 0.9 0.9"), the whole row parks in one cell and the rest of the row stays empty. New `split_item_into_token_subitems` helper splits each item into per-token virtual sub-items with x positions estimated from `effective_width / char_count` and the token's character offset. Stage 1 then routes per-token. Single-token items collapse to a one-element vector (no behavior change). Multi-token items spanning multiple cells distribute correctly. Stage-2 orphan recovery now operates on token-grain orphans rather than re-trying whole items. Tests: - `cid_font_with_unparseable_cmap_does_not_emit_latin1_mojibake` (unit) exercises the Type0/CID + unparseable-CMap fallback path. - `simple_font_latin1_fallback_passes_high_bytes_through` (unit) guards the false-positive case where a Type1 font's `/ToUnicode` reference is set but bytes are legitimate Latin-1 character codes. - `test_extract_tables_with_structure_distributes_wide_item_across_cells` (integration) builds a synthetic PDF with one wide Tj and asserts each token lands in its own cell. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * tests: pin CID mojibake fix mechanism with FFFD assertions Two complementary tests for the Type0/CID Latin-1-fallback guard: 1. Tighten `test_identity_h_no_tounicode_suppresses_garbage` on the existing real-PDF fixture `shinagawa_identity_h.pdf` to also assert the pre-suppression text contains U+FFFD and contains no high-Latin-1 chars. Pins down WHICH mechanism is suppressing the garbage so a future regression that re-enables Latin-1 mojibake fails loudly here instead of silently switching the suppression chain back to `is_cid_garbage` + high-Latin-1 detection. 2. Add `test_synthetic_type0_broken_tounicode_emits_fffd_not_latin1_mojibake` with a fully-synthetic Type0 / Identity-H PDF built in process. We control the malformed ToUnicode contents, the descendant CIDFontType2 shape (just enough for `parse_type0_widths` to set `is_cid=true`, which is what the new guard keys off of), and the Tj byte stream. No fixture file or external license needed. Reproduces the exact "Type0 + non-ASCII bytes + unparseable ToUnicode" code path that produced the production mojibake samples. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
8a0f98dee7 | repair malformed PDF containers (#65) | ||
|
|
f44509ac3a |
fix: don't treat bullet-marker columns as real text columns (#55)
On PDFs where every list item starts with ● at the left margin and content at a fixed offset, histogram column detection sees the gap between marker and content as a gutter and splits each line across two phantom "columns," scrambling the reading order (Anthropic's Mythos system card p.73–74). - layout: reject gutter candidates where the smaller side is ≥80% standalone bullet-marker glyphs (•, ●, ○, ◦, ▪, ▫, ◆, ◇, ■, □) - markdown/classify: add starts_with_bullet_marker helper (narrower than is_list_item — excludes numbered/lettered patterns like 1. and a) so numbered section headings stay as headings) - markdown/convert: skip heuristic heading detection on lines that start with a bullet marker - markdown/classify: strip a leading bullet wrapped in a bold/italic run (e.g. "**● Label:**" → "- **Label:**") — some PDFs put the marker inside the same bold run as the label Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
c409e3b8ca |
fix: use first-glyph position for ActualText ligature items (#43)
* fix: use first-glyph position for ActualText ligature items The content stream parser emitted ActualText items (from BDC/Span marked content) at the text matrix position captured at BDC entry. When the content stream contains Td operators between BDC and the first glyph Tj — common in PDFs generated by Google Docs, Figma, and web-to-PDF tools — the BDC-entry position is on the previous visual line while the actual glyph renders on the correct line. This caused ligature glyphs (fi, fl, ff, etc.) wrapped in ActualText spans to be positioned one text-leading offset above their surrounding text. Downstream merge logic couldn't reconnect them (different y-band), producing broken words like "fi" + "ndings" instead of "findings" throughout the document. Fix: capture the text matrix at the first Tj/TJ operation inside the BDC block (after any Td repositioning), and use that as the ActualText item's rendering position at EMC time. Falls back to the BDC-entry position when no glyph ops occur inside the block. One new variable, five insertion points, zero behavior change on PDFs that don't use ActualText marked content. * chore: allow clippy::collapsible_match for Rust 1.95 Rust 1.95 introduced the collapsible_match lint which flags `if` blocks inside match arms that could be converted to match guards. The content-stream parsers use this pattern extensively for readability (match on PDF operator name, then check preconditions like `in_text_block && !op.operands.is_empty()`). Allow crate-wide rather than refactoring 21 match arms across the parser files. |
||
|
|
10dd7e2881 |
feat: XY-cut fallback for column detection on asymmetric layouts
When the histogram-based column detector finds no valleys (common with sidebar/asymmetric layouts), fall back to a simplified XY-cut: find the largest horizontal gap between item edges and split there if both sides have enough items with vertical overlap. Inspired by opendataloader's XY-Cut++ algorithm but implemented as a single-level fallback rather than full recursive segmentation. Doc 156: NID 0.545→0.966, Doc 157: NID 0.564→0.962. NID-S +0.007, TEDS-S +0.066 across 200 docs. No regressions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
9c463b9c95 |
fix: prevent multi-column text from being misdetected as tables
On pages where column detection finds 2+ columns, skip body-font heuristic table detection in the merged-band retry path. This prevents sidebar/two-column prose from being formatted as markdown tables. The fix is targeted: per-band heuristic detection still runs (bands are scoped to single columns), so real tables within columns are still detected. Only the merged-band retry (which sees all items across columns) is gated. Also relaxes column validation to accept asymmetric layouts (sidebars) where one side has fewer items, and tries center-based item assignment before edge-based to improve column splitting for asymmetric layouts. Benchmark: NID 0.865→0.869, NID-S 0.798→0.805, overall +0.002. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
f664f3056d |
fix: harden hybrid OCR region extraction path (#23)
Align region filtering with rotated-page coordinate rewrites, switch region text assembly to the shared line-grouping pipeline, and retain edge-overlap text to avoid false empty regions that incorrectly trigger OCR fallback. Also make Python region inputs fail fast with clear ValueError messages for malformed boxes. Made-with: Cursor |
||
|
|
fba0a644ef |
fix: replace partial_cmp with total_cmp to prevent NaN sort panics (#20)
* fix sort panics on NaN values from bogus PDF font metrics Replace all `partial_cmp(...).unwrap_or(Ordering::Equal)` and bare `partial_cmp(...).unwrap()` with `total_cmp()` across the codebase. `partial_cmp` returns `None` for NaN, and mapping that to `Equal` violates total ordering: `a == NaN` and `NaN == b` but `a != b`. Rust 1.81+ detects this and panics in sort_by. `total_cmp` handles NaN deterministically (sorts to end) and guarantees total ordering. The critical crash was in `extract_text_in_regions` (lib.rs:478) where PDFs with bogus font ascent/descent values produced NaN in text item coordinates, causing process abort via NAPI. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix missed partial_cmp in layout.rs and restore napi exports - Convert two remaining b.y.partial_cmp(&a.y) calls to total_cmp in group_single_column and column layout sorting - Restore missing napi exports: detectPdf, extractText, extractTextWithPositions, processPdf Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
263ed0aa42 |
add region-based text extraction for hybrid OCR pipelines
Add `extract_text_in_regions_mem` that takes layout-detected bounding boxes (top-left origin, PDF points) and returns text within each region in reading order. Designed for pipelines where a layout model detects regions and text-based pages can skip GPU OCR by extracting text from the PDF structure directly. Also add `classify_pdf_mem` for lightweight PDF type classification returning 0-indexed pages_needing_ocr. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
4d52d7af52 |
fix(extractor): content stream comment parsing + CJK mojibake detection (#18)
* fix(extractor): strip PDF comments that break lopdf content stream parsing Some PDF generators (notably PD4ML used by school districts) embed comments (% to end of line) in content streams. lopdf's Content::decode parser fails to parse operators that follow comments, silently dropping ET (end text) and Q (restore graphics state) operators. This caused entire pages to produce 0 text items despite having valid text. Fix: pre-process content streams to strip comments before parsing. Comments inside string literals (parentheses) and hex strings are preserved. The comment is replaced with a space to maintain token separation. Impact: fixes 13+ school district PDFs and similar PD4ML-generated documents that were producing near-empty output (454 → 31,955 chars for a 22-page school improvement plan). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(detect): flag sparse-extraction pages as needing OCR When a TEXT-BASED PDF produces <50 chars/page average with <500 total chars, flag all pages as needing OCR. This catches PDFs where the extractable text is minimal (form templates, image-heavy layouts) and the bulk of content requires OCR to access. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(detect): improve CID mojibake detection for Japanese/CJK PDFs Extend is_cid_garbage to detect CID-as-Latin-1 mojibake: when ≥40% of characters are high Latin-1 (U+00A0-00FF) and <33% are ASCII letters, the text is likely CID values misinterpreted as Latin-1 characters (common in Japanese/CJK PDFs with broken ToUnicode CMaps). Also add sparse-extraction OCR flagging: TEXT-BASED PDFs with <50 chars/page and <500 total chars get all pages flagged for OCR. Impact: Softbank Japanese PDFs now produce empty output with pages_needing_ocr=all instead of mojibake garbage. Korean PDFs with valid extraction (nexo-price-en) remain unaffected. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(detect): sparse extraction check only when markdown is generated The sparse-extraction OCR check was triggering in Analyze mode where markdown is not generated (md_len=0), causing false OCR flags on every PDF processed via detect-pdf --analyze. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
95aac6a7cd |
feat(layout): relative valley column detection for justified text (#15)
* feat(layout): relative valley column detection for justified text Add fallback column detection using relative valley analysis for PDFs with justified text where item widths extend past gutter boundaries. The absolute valley detector fails on these layouts because gutter bins are at ~40% of peak (well above the 15% noise threshold). The relative valley detector smooths the histogram with a 5-bin moving average, finds local minima where contrast < 0.60 of surrounding peaks, and validates with peak balance >= 0.40. Limited to single best valley (max 2 columns) and requires >= 100 items per page. Tested on IRS Publication 17 (2002), a 289-page 2-column justified text document: column detection went from ~40 pages to 165 pages. 190 passed, 0 regressions across 191 eval PDFs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(layout): tighten relative valley thresholds to reduce false positives Reduce PEAK_WINDOW from 40 to 25 bins (50pt) so valleys are only validated against nearby peaks, not distant ones. Add MIN_PEAK_HEIGHT of 20 (smoothed) to reject sparse pages where histogram peaks are too low to indicate dense two-column text. Previous thresholds caused 13 regressions across the eval suite by splitting tables, TOCs, checklists, and forms. Now: 188 passed, 0 regressions (2 minor metadata-only diffs on IRS P17 and 9978293). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(layout): skip relative valley detection on pages with tables Table column gaps in the histogram look identical to text column gutters but the table pipeline already handles reading order for those pages. Pass page_has_table flag through detect_columns to suppress the relative valley fallback on pages where tables were detected. This eliminates all remaining regressions from relative valley detection: 190 passed, 0 regressions across 191 eval PDFs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(layout): prose density validation for relative valley detection Add columns_have_prose() to validate relative valley column splits. Checks that both sides of a proposed split contain paragraph-like content (fill ratio >= 40%, avg items/line <= 3.5) before committing to a column split. Combined with the table-page guard, this prevents false column splits on financial statements, forms, and tabular layouts where long labels or dot leaders fill the column width. Also tightens find_relative_valleys() thresholds (PEAK_WINDOW 40->25, MIN_PEAK_HEIGHT 5->20) to reduce false positive valley candidates. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
ad10f670bf |
chore: migrate from deprecated load_mem_with_password to load_mem_with_options
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
fa94314760 |
fix: skip page extraction when operation count exceeds 1M (#12)
* fix: skip page extraction when operation count exceeds 1M Vector-heavy architectural PDFs can have 10-26M path operators per page, causing ~60 GB allocation during per-op processing. The decode itself is tolerable but the subsequent loop amplifies memory 2-5x with text items, rects, paths, and state tracking. Check operation count after Content::decode() and return empty extraction for pages exceeding the limit, with a warning log. * test: add unit test for excessive operations guard Constructs a synthetic PDF with 1.1M path operators to verify that pages exceeding the operation limit return empty extraction. |
||
|
|
8c4181434f |
fix(text): handle Tc/Tw character and word spacing (#11)
* fix(text): handle Tc/Tw character and word spacing in text width computation PDFs using Tc (character spacing) and Tw (word spacing) operators for text justification had words incorrectly split across TextItems. The computed advance width didn't account for these spacing parameters, causing spurious spaces mid-word (e.g. "deve lopers" instead of "developers"). - Add Tc/Tw operator handling and graphics state save/restore - Incorporate char_spacing and word_spacing into compute_string_width_ts - Add adaptive merge threshold: tighter for lowercase→lowercase junctions, wider before joining punctuation - Add unit tests for Tc/Tw width computation and merge behavior Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * test(fonts): add large Tc width computation test Verifies that large character spacing values are applied in full without any artificial cap, matching PDF spec behavior. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(text): guard against Tc/Tw-inflated widths in merge and join paths Two targeted fixes to prevent character-spacing (Tc) and word-spacing (Tw) inflation from causing data quality regressions: 1. should_join_items: reject large negative gaps (< -font_size) that arise when Tc/Tw inflate item widths past adjacent items. Fixes FY_2015 merged numbers (e.g. "239.696.0" → "239.69 6.0"). 2. merge_text_items: cap effective width for gap computation when Tw inflates space-containing items beyond 0.85× font_size per char. Prevents column-level gaps from collapsing into merge range, recovering table detection for Baldwin-Edwards and similar PDFs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |