Compare commits

...
Author SHA1 Message Date
Abimael MartellandCursor 23a29f8894 docs(extractor): distinguish /W insert vs unique-key CID caps
Encoding cidrange and /W width assignment count every insert; the /W
unicode heuristic caps unique CIDs with the same 65,536 bound.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 17:07:35 -07:00
Abimael MartellandCursor 261da949f3 docs(extractor): clarify Encoding cidrange cap counts insert operations
The bound includes overwrites so repeated full-width ranges cannot keep
working after the map is full. Unique-key coverage alone would re-open
the CPU blow-up.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:24:16 -07:00
Abimael MartellandCursor 00287b293d fix(extractor): bound Encoding CMap cidrange expansion
Repeating full-width begincidrange declarations re-inserted the entire
16-bit domain on every copy. Stop after 65,536 assignments, matching the
existing /W cap.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:15:56 -07:00
Abimael MartellandCursor 076183e2e4 fix(extractor): cap content-stream decode before allocating operators (#373)
* fix(extractor): cap content-stream decode before allocating operators

The 1M operation limit ran after lopdf materialized the full vector, so a
compact page of q/Q pairs could still abort under memory pressure.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): treat NUL and form-feed as PDF whitespace in op counting

Names must stop on the full PDF whitespace set so a following operator is
not absorbed into /Name, which would undercount and skip the decode cap.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): scan inline-image EI with the full PDF whitespace set

A missed EI terminator used to consume the rest of the stream and drop
later operators from the decode cap. If EI is absent, keep scanning.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:13:12 -07:00
Abimael Martell ec6e54afb8 fix(extractor): merge small-caps runs so they stop reading as table columns (#371)
* fix(xobjects): track text line matrix and handle T*/TL/'/"/Tc/Tw in Form XObjects

The Form XObject text extractor in xobjects.rs is a separate hand-rolled
implementation of the operator state machine in content_stream.rs, and it had
drifted well out of parity:

- No text line matrix (TLM). `Td`/`TD` were applied to the text matrix already
  advanced by `Tj`/`TJ`, so every line began where the previous line *ended*
  instead of at the line start. Lines marched off the right edge and were
  dropped as off-page.
- `T*` was not handled at all, so it never advanced to the next line.
- `TL`, `'` and `"` were missing, and `TD` never set the leading as a side
  effect.
- `Tc`/`Tw` were hardcoded to 0.0 when computing advance widths, drifting
  positions and inserting spurious spaces.
- Text state (Tc/Tw/TL/Tf) is part of the graphics state but was not saved or
  restored by `q`/`Q`.

This matters well beyond an edge case: producers that emit a page stream of
just `q /X Do Q` and put all content in a Form XObject are common in
print-to-PDF and typesetting workflows, so this parser is on the hot path for
whole classes of real documents.

Measured on nycourts.gov 199AD3d.pdf (1370 pages, PDFlib producer, every page
wrapped in a Form XObject, 1331 pages using T*), word recall against a
pdftotext reference goes from 19.2% to 97.7% — 116k extracted words to 582k
against a 578k-word reference. On a 10-page subset, sequence similarity goes
from 27.3% to 97.7%, against 99.8% for Mistral OCR.

pdf-evals: 195 passed / 7 failed, byte-identical to the origin/main baseline
with the same failure list — no regressions.

Adds 7 unit tests covering the line-matrix-relative `Td`, `T*`, `TD` setting
leading, `'`, `"`, `Tc` advance widths, and `q`/`Q` text-state restore, all
driven through a page whose content is only `q /X1 Do Q`.

* fix(xobjects): restore fill colour across q/Q in Form XObjects

A white fill set inside a q/Q pair leaked past the Q, so all subsequent
text was treated as invisible and dropped. Save and restore fill_is_white
with the rest of the graphics state.

This was the cause of several long-standing extraction failures where
whole passages went missing or degraded into per-character garbage:
cambridge_excerpt (+8.5KB of recovered text), MTUAeroEngines (+7.4KB),
2025_findings-acl_668 (+1.4KB), HTM_02-01_Part_A (+1.3KB), and
HuttoISDWorkPerks / ebgt7isj04ophcq, which both went from exploded
per-character tables to clean prose.

Adds a regression test that fails without the restore (the text after Q
is dropped entirely).

Reported by cubic on #369.

* fix(extractor): merge small-caps runs so they stop reading as table columns

Typesetters render small caps as a full-size capital immediately followed by
shrunken capitals in the same font — `(R) Tj` at 9.98pt, then `(OLANDO) Tj`
at 6.74pt, touching. The 20% font-size band in merge_text_items split those
into separate items, and since table column boundaries cluster on item *start*
positions (find_column_boundaries never consults widths), each fragment
started far enough from the last to become its own column.

The result was garbled pseudo-tables. From 199AD3d.pdf p.5:

  |OLANDO|COSTA|NIL|INGH||
  |---|---|---|---|---|
  |IANNE|ENWICK|ETER|OULTON||

  |R|T. A|, P.J.|A|C. S|
  |---|---|---|---|---|

now:

  |ROLANDO T. ACOSTA, P.J.|ANIL C. SINGH|
  |---|---|
  |DIANNE T. RENWICK|PETER H. MOULTON|

Adds is_small_caps_continuation, gated tightly enough to exclude the other
reasons a smaller run follows a larger one: superscripts and footnote markers
(requires an uppercase *letter*, so digits never qualify), drop caps (the
following body text is mixed case), and adjacent cells or separate words
(requires the runs to be visually contiguous). A small-caps junction is
mid-word, so it also suppresses the space that would otherwise be inserted
("T. A" + "COSTA" -> "T. ACOSTA", not "T. A COSTA").

On 199AD3d.pdf this drops false-positive tables from 65 to 52 and lifts word
recall against a pdftotext reference from 97.7% to 97.8%.

Verified by running both binaries over all 203 eval PDFs and diffing outputs:
19 differ, 184 byte-identical. Beyond the reporter volume the merge also fixes:

  Waters-Edge          two more garbled heading-tables — "##### B. C R S" plus
                       "||OMPLIANCE|EPORTING YSTEM|" became
                       "##### B. COMPLIANCE REPORTING SYSTEM"
  cn-student-handbook  six TOC entries had collapsed to their initials;
                       "U M S H..." is now
                       "UNIVERSITY MISSION AND STUDENT HANDBOOK..."
  546403               "(FAA)" + "ADVISORY CIRCULARS ( )" + a stray "CONT"
                       became "(FAA) ADVISORY CIRCULARS (CONT)"
  ERP-2025             "T ABLE B-1" -> "TABLE B-1"
  DMP-Keypad           "THINLINE" + stray "TM KEYPADS" -> "THINLINETM KEYPADS"
  zhaw / Stijn         subscripted math variables: "*T* *G*" -> "*TG*"

One known regression, called out rather than hidden:
PA_PVEM_Sen_Waldo_Fernández (+1689 bytes). Its running footer genuinely is
small caps, so the merge correctly assembles the text (400 items -> 275), but
the now-contiguous footer lines align well enough that the heuristic detector
turns the wrapped document title into a 4-column table where it previously
rendered as bold prose.

Attempting to fix that in the detector was a dead end and is not included
here: gating the `num_cols >= 3` bypass in has_table_like_content on "no cell
has >= 12 words" removed 143 tables across 44 files, including correct ones —
MCF5235RM went 505 -> 483 and its register glossary degraded into a fused
column plus a ~100-word cell, because rejecting a good candidate lets a worse
fallback win. No wordiness threshold separates the two cases: PA_PVEM's
longest cell is 14 words while MCF5235RM's legitimate cells are <= 10. A
positional signal (running-header/footer bands, or cross-page repetition) is
the way in, and belongs in its own change.

Adds 7 unit tests: the two-column small-caps row from the reporter volume, and
rejection of superscript digits, drop caps, separate words, lowercase
continuations, and out-of-band size ratios.

* fix(extractor): tighten small-caps gate to cross-band junctions only

Three findings from cubic on #371, all valid.

1. Space suppression was too broad. `is_small_caps_continuation` accepted size
   ratios up to 0.92, which is *inside* the 20% band that merge_text_items
   already treats as the same size. Two similarly-sized uppercase words with a
   small real word gap therefore had their space suppressed even though the
   normal path would have merged them correctly with a space ("SEE" + "ALSO" ->
   "SEEALSO"). The helper now requires the junction to *cross* the band, since
   rescuing junctions the band would break is its only purpose; within-band
   pairs keep the normal word-spacing logic. The band is now a shared
   MERGE_FONT_SIZE_BAND constant so the helper and its caller cannot drift.

2. A trailing digit was skipped. The backward search for "the capital we are
   continuing" skipped non-alphabetic characters, so text ending in a footnote
   marker ("ANGELA M. MAZZARELLI1") found the earlier uppercase letter and
   accepted the join. It now checks the actual trailing character.

   Rejecting every trailing digit outright turned out to cost real quality, so
   this keeps one narrow exception. In meetings_in_mass_july_2023 the source
   reads "TUESDAY, JULY 4TH" with the ordinal suffix set as a smaller run; a
   blanket digit rejection reverted that to "TUESDAY, JULY 4" plus a stray "TH"
   leaking onto the next line, which is what main produces and what pdftotext
   shows is wrong. Only the four English ordinal suffixes (TH/ST/ND/RD) may
   follow a digit; anything else after one is treated as a footnote marker and
   rejected, so the case cubic raised stays blocked.

3. A test's name did not match its data. `small_caps_merge_does_not_swallow_a_
   second_column` claimed to exercise the 72pt column gap but contained only
   the second column's items, so it just re-tested the happy-path merge. It now
   holds the full nine-item row and asserts exactly two merged results,
   "ROLANDO T. ACOSTA, P.J." and "ANIL C. SINGH", which genuinely exercises the
   gap.

Both new guards were verified to be load-bearing: removing either one makes its
test fail.

The document that motivated the change is unaffected — small caps there run at
a 0.675 ratio, far outside the band — and 199AD3d.pdf p.5 still produces the
correct two-column justices table.

900 unit tests pass (10 covering small caps), fmt and clippy clean.
2026-08-12 15:33:18 -07:00
Abimael Martell 89dd20d02c fix(xobjects): track text line matrix and handle T*/TL/'/"/Tc/Tw in Form XObjects (#369)
* fix(xobjects): track text line matrix and handle T*/TL/'/"/Tc/Tw in Form XObjects

The Form XObject text extractor in xobjects.rs is a separate hand-rolled
implementation of the operator state machine in content_stream.rs, and it had
drifted well out of parity:

- No text line matrix (TLM). `Td`/`TD` were applied to the text matrix already
  advanced by `Tj`/`TJ`, so every line began where the previous line *ended*
  instead of at the line start. Lines marched off the right edge and were
  dropped as off-page.
- `T*` was not handled at all, so it never advanced to the next line.
- `TL`, `'` and `"` were missing, and `TD` never set the leading as a side
  effect.
- `Tc`/`Tw` were hardcoded to 0.0 when computing advance widths, drifting
  positions and inserting spurious spaces.
- Text state (Tc/Tw/TL/Tf) is part of the graphics state but was not saved or
  restored by `q`/`Q`.

This matters well beyond an edge case: producers that emit a page stream of
just `q /X Do Q` and put all content in a Form XObject are common in
print-to-PDF and typesetting workflows, so this parser is on the hot path for
whole classes of real documents.

Measured on nycourts.gov 199AD3d.pdf (1370 pages, PDFlib producer, every page
wrapped in a Form XObject, 1331 pages using T*), word recall against a
pdftotext reference goes from 19.2% to 97.7% — 116k extracted words to 582k
against a 578k-word reference. On a 10-page subset, sequence similarity goes
from 27.3% to 97.7%, against 99.8% for Mistral OCR.

pdf-evals: 195 passed / 7 failed, byte-identical to the origin/main baseline
with the same failure list — no regressions.

Adds 7 unit tests covering the line-matrix-relative `Td`, `T*`, `TD` setting
leading, `'`, `"`, `Tc` advance widths, and `q`/`Q` text-state restore, all
driven through a page whose content is only `q /X1 Do Q`.

* fix(xobjects): restore fill colour across q/Q in Form XObjects

A white fill set inside a q/Q pair leaked past the Q, so all subsequent
text was treated as invisible and dropped. Save and restore fill_is_white
with the rest of the graphics state.

This was the cause of several long-standing extraction failures where
whole passages went missing or degraded into per-character garbage:
cambridge_excerpt (+8.5KB of recovered text), MTUAeroEngines (+7.4KB),
2025_findings-acl_668 (+1.4KB), HTM_02-01_Part_A (+1.3KB), and
HuttoISDWorkPerks / ebgt7isj04ophcq, which both went from exploded
per-character tables to clean prose.

Adds a regression test that fails without the restore (the text after Q
is dropped entirely).

Reported by cubic on #369.
2026-08-12 14:32:50 -07:00
Abimael MartellandCursor 3d33ff3dbd fix(extractor): bound CID /W range expansion (#372)
Type0 /W parsing and the Unicode-CID heuristic expanded every CID in every
range. Repeating a full-width [0 65535 w] entry therefore re-materialized
the same 65,536-key domain on every copy, growing a temporary vector and
HashMap work without bound.

Cap expansion at the 16-bit CID domain, collect unique CIDs for the
median heuristic, and stop width assignment once that many entries have
been written. Legitimate compact /W arrays are unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 14:27:29 -07:00
Abimael MartellandCursor 75e9b09593 fix(extractor): bound Form XObject expansion per page (#370)
* fix(extractor): bound Form XObject expansion with invocation and operation budgets

Nested Form XObjects were only limited by recursion depth (5). An acyclic
graph where each form invokes the next N times still expands to N^depth
work, so a small PDF can force millions of nested /Do evaluations.

Share a per-page FormWalkBudget that caps 10,000 Form invocations and
1,000,000 operations walked across those expansions. Extraction stops
when either cap is hit. Legitimate shallow nesting is unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): share Form XObject budget across invisible-layer retry

extract_page_text_items created a fresh FormWalkBudget on every call, so
the invisible-text retry could consume a second full expansion budget for
the same page. Own the budget at the call site and pass it into both
passes.

Also charge form operations independently of the invocation cap so a form
that was already admitted can finish its stream.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 14:03:19 -07:00
Abimael MartellandClaude Fable 5 f4aab3b36f chore(release): bump package versions to 1.14.1 (#354)
Releases fix(regions) #351 — invisible (Tr 3) OCR text layers served from
the region extractor instead of falling back to GPU OCR.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 14:18:41 -07:00
Abimael MartellandClaude Fable 5 f65b25906c fix(regions): serve invisible (Tr 3) OCR text layers instead of needs_ocr (#351)
* fix(regions): serve invisible (Tr 3) OCR text layers instead of needs_ocr

Scanned pages (archive.org-style digitizations) carry their text as an
invisible render-mode-3 layer behind the page raster. extract_text_in_regions
extracted only visible items, so such pages yielded nothing but
'[Image: ...]' placeholders and every region fell back to OCR — while the
markdown path already includes the invisible layer for Mixed PDFs. The two
extractors disagreed about the same page.

Mirror the markdown path's gate, page-scoped: when the visible pass is
effectively textless (<40 non-placeholder alphanumerics), retry with
invisible text included and adopt the retry only when it contributes real,
non-garbage text. Pages with real visible text never retry, so double-layer
PDFs (visible text plus an invisible accessibility copy) are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: strict zero-visible adoption gate, whole-layer garbage check, doc placement, discriminating tests

- Adoption now requires ZERO visible text on the page: the invisible pass
  returns visible items too, so adopting alongside any visible text would
  duplicate it. Strict gate instead of fuzzy dedupe; the dead
  inv_alnum > visible_alnum condition goes with it.
- Garbage check judges the whole recovered layer, not the first 200 items.
- Constant/helper moved above the doc block so rustdoc stays attached to
  extract_text_in_regions_mem.
- Tests: any-visible-text-blocks-adoption case (single short visible line)
  with exactly-once assertions; mode-0 guard asserts occurrence count.
- Re-verified against real scanned-book pages: all recover (9.2-11.7K chars,
  needs_ocr=false) — pure OCR-layer scans carry no visible text.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: gate adoption on visible-item presence, not alphanumeric mass

Punctuation-only visible text (zero alphanumerics) could still adopt the
invisible pass; its OCR twin in the layer would duplicate the glyphs. The
gate is now item-presence: any non-image item with non-whitespace text
blocks adoption (whitespace-only artifacts still tolerated). Pinned by a
punctuation-only fixture test; real scanned-book pages re-verified — all
recover unchanged (they carry no visible items at all).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: punctuation-gate test pins preservation, comment states fixture scope

The fixture's visible punctuation is a separate block, not mirrored in the
invisible layer — so it pins the gate, not the duplication scenario. The
doc comment now says so, and a positive assertion checks the punctuation
survives exactly once.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: retry only when invisible text was actually skipped; negative gate tests

- extract_page_text_items now reports skipped_invisible (4th return): the
  Tr-3 suppression sites set it, so the region extractor retries only when
  a recoverable layer exists. Blank pages and image-only scans without an
  OCR layer — the common scanned case — no longer pay a second
  content-stream parse.
- Negative tests: below-floor watermark layer and symbol-garbage layer are
  both rejected (region keeps only the raster placeholder). needs_ocr
  semantics for placeholder-only regions deliberately unchanged — that
  contract predates this PR and downstream pipelines handle it.
- Real scanned-book pages re-verified: recovery unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: flag the ' show-text suppression path; require nonempty strings

- P1: invisible layers shown via the ' operator never set skipped_invisible,
  so those pages kept the placeholder and fell back to GPU OCR. The '
  suppression path now flags too — pinned by a fixture whose layer is shown
  entirely via ' (nothing through Tj/TJ).
- P3: all three flag sites (Tj, TJ, ') require a nonempty string operand —
  numeric-only TJ kerning arrays and empty shows no longer trigger the
  invisible reparse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 14:11:31 -07:00
Abimael Martell 9947485a92 docs: document structure-element extraction and TextItem.mcid (#349) 2026-08-11 12:32:56 -07:00
Abimael Martell 7054d6aa69 chore(release): unify package versions (#344)
* chore(release): unify package versions

* fix(release): harden version synchronization
2026-08-11 09:26:19 -07:00
Abimael MartellandClaude Fable 5 a67ee03269 feat(bindings): expose TextItem.mcid and structure-tree element extraction (#346)
Tagged PDFs carry a structure tree with real heading roles (H1..H6), and
the core already parses it (structure_tree::StructTree) and threads MCIDs
onto TextItem — but neither surfaced through the bindings.

- Expose TextItem.mcid (Option<i64>) through the napi and pyo3 bindings,
  matching the core field added with the marked-content extractor.
- Add StructRole::name(), the inverse of from_name, so roles have a
  stable string form.
- Add extract_structure_elements / extract_structure_elements_mem to the
  core: one (page, mcid, role) entry per marked-content reference, sorted
  by (page, mcid), empty for untagged PDFs. Pages are 1-indexed to match
  TextItem.page, so results join directly against
  extract_text_with_positions output.
- Bind it as extractStructureElements (napi) and
  extract_structure_elements / extract_structure_elements_bytes (pyo3),
  with type-stub updates in pdf_inspector.pyi.
- Cover the join in Rust integration tests, napi test.mjs, and pytest,
  using the existing firecrawl_docs_tagged.pdf fixture (tagged) and
  thermo-freon12.pdf (untagged).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 20:03:07 -07:00
Abimael Martell 965dc65f1b chore(release): bump package versions (#343) 2026-08-10 14:33:40 -07:00
36dd5fa426 fix(structure-tree): bound recursive /K parsing with cycle detection and a node budget (#322)
* fix(structure-tree): bound tagged /K parsing against alias/cycle DoS

A struct element that references itself (or an ancestor) through /K — e.g.
/K [n 0 R n 0 R] — made parse_struct_element_dict branch exponentially:
the depth cap (64) alone still permits 2^depth materialized nodes, so a
~830-byte PDF exhausts memory (OOM, exit 134).

Add a StructWalk carrying (1) an active-path set of object IDs so a node
that references itself/an ancestor is not re-expanded (breaks self- and
mutual-reference cycles cheaply), and (2) a global node budget
(MAX_STRUCT_NODES) that caps total materialization for aliased/DAG-shaped
graphs of distinct objects the path guard cannot catch.

Adds regression tests for self-alias, mutual-alias, and the aliased-DAG
budget cap.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): charge /K content refs against the node budget

The per-node budget only covered materialized struct elements and child
recursion; bare MCIDs and MCR dicts in a /K array append to content_refs
without charging it, so one element with a very wide /K array could still
allocate content_refs without bound. Charge every /K array item before
handling it, and stop the top-level /K loop once the budget is spent, so
content refs and loop work are bounded too. Adds a wide-MCID-array test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): charge /K budget per materialized item, not per array entry

Charging every /K array item double-counted structural children (charged
here and again at their node entry) and charged cycle-skipped references
that materialize nothing, draining the budget up to ~2x faster than the
per-node semantics and risking early truncation of large legitimate trees.
Charge only the unbounded content-ref items (bare MCIDs and MCR dicts);
structural children remain charged once at their node entry.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* refactor(structure-tree): charge every content ref uniformly via helper

Route all budget charges through StructWalk::charge() so every
marked-content reference is charged once, including the single-value /K
branches (bare integer and MCR dict) that previously appended without
charging. charge() also guards against underflow, so charging after the
node-entry charge (which can leave the budget at 0) is safe. Makes the
documented per-item budget contract hold uniformly across all branches.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* feat(structure-tree): log once when the node budget truncates parsing

Add a one-shot truncation flag on StructWalk, set the first time the
budget is exhausted, and emit a single warn! after parsing so an operator
can tell when a (very large or malformed) tagged tree was cut off. Avoids
per-item log spam; negligible overhead on the normal path.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): flag truncation at budget guards, not just in charge

The truncation flag was only set inside charge() on the budget==0 branch,
but the dominant skip paths use budget==0 guards that break/return before
charge() is ever called with an empty budget, so the flag (and the warn!)
almost never fired. Route those guards through a new exhausted() that sets
the flag when it skips remaining work. Adds a parser-level test that would
have caught the missed warning.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): flag cycle/depth skips and charge bare MCIDs fully

Two review follow-ups:
- Cycle-broken and depth-capped /K skips dropped tagged content without
  setting the truncation flag, so the one-shot warning never fired for
  malformed/over-deep trees. Mark those skips via note_skipped() and
  broaden the warning to cover non-budget truncation.
- A bare /K MCID materializes a wrapper node AND a content reference but
  charged only one budget unit, allowing ~2x the advertised budget for
  such content; charge both.

Adds tests: cycle-skip flags truncation, and bare MCID charges two units.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): charge MCR-dict wrappers the same two units as bare MCIDs

A top-level MCR /K dict flows through parse_kid -> parse_struct_element_dict
and materializes a Span node + one content ref (two items) but was charged
only one unit at node entry, while the bare-MCID path charges two. Charge
the content reference in the MCR branch too so the per-item budget is
uniform across both wrapper paths. Adds a symmetric test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): reserve leaf-wrapper budget units atomically

A leaf MCID wrapper (bare MCID or MCR dict) materializes a node + one
content ref and charged the two units via separate charge() calls. At the
last unit the first charge succeeded and the second failed, consuming a
unit without emitting the wrapper and denying it to a later element that
would have fit. Add charge_n() to reserve both units atomically (or
neither), and detect MCR before the node charge so it reserves both up
front. Adds a boundary test asserting the leftover unit is preserved.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): stop scanning wide /K once a leaf reservation stalls

The atomic charge_n(2) left budget nonzero (==1) when it failed, so
exhausted() (budget==0) never broke the root /K loop and a crafted wide
array of leaf wrappers was scanned in full after no leaf could fit. Add a
stalled flag set on an insufficient reservation and fold it into
exhausted(); charge()-based (one-unit) loops are unaffected since they
reach budget 0 exactly. Adds a test that a one-unit budget still allows a
one-unit item but a failed two-unit reservation stops the scan.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): add traversal budget and stop charging non-materializing dicts

Two review follow-ups on budget accounting:
- Wide /K arrays of non-materializing items (unsupported value types, OBJR
  dicts, cycle back-edges) consumed no node budget, so the loop scanned the
  whole array. Add a separate work budget charged per examined /K item and
  break the loops when it is spent, bounding traversal even when nothing
  materializes.
- OBJR dicts and dicts without a valid /S were charged the node budget before
  being recognized and skipped, draining the shared budget and truncating
  real content later. Hoist the OBJR check and /S validation above the node
  charge so only materializing nodes consume it (matching the MCR hoisting).

Adds tests for the work-budget bound, wide unsupported /K, and non-materializing
dicts not charging the node budget.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* docs(structure-tree): mention traversal budget in truncation warning

The one-shot truncation warning listed the node budget, cycle, and depth
as causes but not the new traversal (work) budget, so a work-budget
truncation printed a misleading message. Include MAX_STRUCT_WORK so
malformed-PDF debugging identifies the actual limit hit.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 13:36:40 -07:00
fabec0aec3 feat(napi): add async variants that keep the Node event loop free (#337)
* feat(napi): add processPdfAsync, classifyPdfAsync, extractPagesMarkdownAsync

The Node bindings are synchronous, so every call parses on the event
loop thread — up to hundreds of milliseconds of dead loop per document
in a server. Add additive AsyncTask-based variants that run the same
shared implementations on the libuv thread pool and return promises.

The existing synchronous exports keep their names, signatures, and
behaviour; each sync/async pair shares one implementation. Panics in
compute() are caught and surfaced as rejections, matching the sync
error contract.

Closes #336

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(napi): read async task buffers in place instead of copying

Review feedback on #337: buffer.to_vec() copied the whole PDF on the
event loop before the task was queued, so large inputs still stalled
the loop and doubled peak memory. The tasks now hold the napi Buffer
itself — its ref pins the JS allocation for the task's lifetime and
the backing store is stable, so compute() reads it directly from the
worker thread. Callers must not mutate the buffer until the promise
settles (same contract as Node's async fs APIs); documented on each
export and in the README.

The suggested removal of ts_return_type was checked and rejected:
without it napi-rs generates Promise<unknown> for AsyncTask returns.
A comment now records that finding.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(napi): copy async task input on the JS thread for soundness

Review feedback on #337: holding the napi Buffer and reading it from
the libuv worker was unsound. Buffer derefs straight to the JS-side
allocation, so a caller mutating it before the promise settled would
race the worker's reads — undefined behavior, not a recoverable error,
and the documented don't-mutate contract was unenforceable. Deferring
the copy to compute() would not help: any off-thread read races the
same way. The JS thread is the only race-free place to take the copy,
because JS is single-threaded and nothing can mutate the buffer during
the synchronous part of the call.

Revert to an owned Vec<u8> copied at call time. The cost is one memcpy,
negligible next to the parse the async variants exist to unblock. Docs
now state the buffer may be reused or mutated immediately, and a test
locks in the copy semantics by mutating the input while a parse is in
flight.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 12:29:23 -07:00
1f28c00a13 Bound column-detection histogram and harden coordinate handling (#328)
* Bound column-detection histogram size

Derive the projection histogram from a clamped bin count and skip
non-finite page widths. Extreme or malformed text-item coordinates
(from the content-stream text matrix) could otherwise drive a very
large allocation. 65,536 bins is ~9x the largest legal page, so real
layouts are unaffected. Adds regression tests.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Exclude non-finite coordinates from page bounds

Items at NaN/inf positions are now skipped when folding the page bounds,
so a malformed coordinate can no longer escape as a ColumnRegion
boundary, and an all-non-finite page returns no columns. Bad items are
dropped individually rather than failing the page, so one stray glyph
does not disable column detection.

Addresses review feedback on the finite-width guard.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Trim far-outlier coordinates from page bounds

Gutter margins, spanning-item width and the XY-cut margin are all
fractions of page_width, so a single far-but-finite item (x=50_000 is
enough) set the scale for the whole page: real gutters fell inside the
rejected margin band and a genuine two-column page collapsed to one
region. When the span exceeds one legal page (14_400 units), re-derive
the bounds from items clustered around the median x. Outliers keep their
text because column assignment buckets by nearest overlap.

The MAX_BINS ceiling stays as an allocation bound that does not depend
on this heuristic.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Harden bounds trimming against widths and wide layouts

Check both item edges when trimming: a malformed width at an ordinary
position poisoned x_max just as a malformed position poisoned x_min, so
a huge width still collapsed a two-column page to one region.

Only trim when the far items are a small minority (<=10%). A genuinely
large-format page has content spread across its full width, so it now
keeps its true bounds instead of being reduced to the median cluster.

Correct the MAX_PAGE_EXTENT comment: 14_400 units is the traditional
Acrobat architectural limit, not a format cap. PDF 2.0 sets no page-size
limit and UserUnit scales physical size, so this is a heuristic.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Scale bin width so the histogram spans the whole page

Clamping the bin count alone left anything past MAX_BINS * BIN_WIDTH
(~131k points) outside the histogram, folded into the final bin. A page
wide enough to hit that lost real gutters: with a visible gutter inside
the covered range the XY-cut fallback never runs, so a three-column
layout silently reported two. Derive bin_width from page_width instead,
keeping the same allocation ceiling and degrading only resolution.

Also anchor the trimming median on the same finite left/right items that
bounds() accepts, so a malformed width cannot shift which items count as
strays.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Require geometric evidence before detaching far content

The median-window trim narrowed any content more than one page from the
centre, so a valid large page with a sparse far sidebar lost the sidebar
from its bounds and its text fell into column 0. An item-count minority
rule cannot tell that layout from malformed coordinates.

Group content into clusters separated by more than a whole page of
continuous emptiness, and only drop a cluster that is both detached by
such a void and a small minority of items. Real content does not leave a
gap that large; a stray coordinate sits alone beyond one.

A single run wider than one page is treated as a malformed width, which
also covers the huge-width case the cluster sweep cannot see (such an
item spans everything and leaves no gap).

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Judge run width against page content, not a fixed extent

Treating any run wider than 14_400 units as malformed penalised valid
large pages: one made entirely of such runs reported no columns at all,
and a mixed page lost the right edge of every long run.

Judge width relative to the page's own content instead. Positions cannot
be inflated by a bogus width, so the spread of the core cluster is a
sound scale: a run wider than that spread plus one page is malformed.
A genuinely large page keeps its genuinely long runs, while a 1e12-wide
run beside ordinary text is still rejected.

Cluster on positions rather than filled intervals, so a bogus width can
no longer merge everything into one cluster, and keep ordinary pages on
an O(n) fast path that skips the sort entirely.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 12:29:12 -07:00
f4b8c9e854 Clarify SECURITY.md reporting channels (#329)
* Update SECURITY.md reporting channels

Clarify that email is the only required channel and point the
alternative at Firecrawl's Bugcrowd disclosure engagement instead of
the private-advisory link, which is not enabled on this repo.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Make Bugcrowd the preferred reporting channel

Bugcrowd's disclosure engagement is the primary channel; email to
help@firecrawl.dev is offered as the alternative.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 16:27:50 -07:00
69039f2728 Fix char-boundary panic in hex_to_unicode_string (#320)
Use hex.get(i..i+2) instead of &hex[i..i+2] so a non-hex, non-ASCII
destination in a /ToUnicode CMap can no longer trigger a UTF-8
char-boundary panic. An even byte length does not guarantee the byte
offset falls on a char boundary; get() returns None on a non-boundary
or out-of-range index, folding cleanly into the existing flow.

Add regression tests covering a multi-byte destination char and a
replacement-char byte.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 08:21:31 -07:00
3cca6446bd fix(glyph_names): handle non-ASCII input in uniXXXX glyph name parsing (#321)
Use str::get instead of a byte-length check plus slice when parsing the
uniXXXX glyph-name form. The byte-length guard only proved the index was
in bounds, not on a UTF-8 char boundary, so a glyph name containing
non-ASCII bytes could cause a slice on a non-boundary index. Switch to a
checked slice that folds into the existing Option flow, and add tests.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 08:21:22 -07:00
493fed498e fix(links): prevent stack-overflow DoS from AcroForm /Kids self-cycle (#314)
* fix(links): guard AcroForm /Kids traversal against cycles and huge trees

A crafted PDF whose AcroForm field lists itself (or another ancestor) in
/Kids caused walk_form_fields to recurse indefinitely, overflowing the
stack and aborting pdf2md (exit 134) — an application-level DoS from a
~730-byte input.

Track visited field object IDs to break /Kids cycles, and cap total
field-node traversal at 100k nodes to bound pathologically large trees.

Adds regression tests for self-cycle and mutual-cycle field graphs.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): cap AcroForm /Kids recursion depth to stop deep-chain overflow

The visited-set guard stops cyclic /Kids graphs, but a long *acyclic*
chain of distinct fields still recurses to the chain length and overflows
the stack (a ~1.6MB PDF with 20k linked fields aborts pdf2md, exit 134)
before the 100k node budget is reached.

Add an explicit recursion depth cap (100 levels — far above any legitimate
form hierarchy) so stack usage is bounded independently of node count.

Adds a deep-acyclic-chain regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): enforce form-field node budget before insertion

The node-budget guard inserted each field ID into the visited set before
checking the budget, so the check triggered an early return but never
actually capped the set. A field with a huge /Kids array kept inserting
post-budget IDs, letting visited (memory and work) grow with the crafted
input rather than stopping at MAX_FORM_FIELD_NODES.

Check depth and budget before inserting, so visited can never exceed the
cap. Adds a wide-tree regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): stop /Fields and /Kids iteration once node budget is spent

Checking the budget before insertion capped the visited set, but callers
still iterated every remaining entry of a wide /Fields or /Kids array
after the budget was exhausted — each walk returned immediately, yet the
O(N) sibling iteration let a single multi-million-entry array burn
extraction CPU unbounded. Break out of both the top-level and recursive
loops once visited reaches the cap, making the budget a true
traversal-work cap. Adds a top-level wide-/Fields regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): charge examined entries against the field-node budget

The budget counted only distinct visited nodes, so /Fields or /Kids
arrays full of invalid (non-reference) or duplicate entries never grew
visited and ran to completion regardless of size — the node budget did
not actually cap traversal work.

Introduce FieldWalkBudget tracking both visited nodes and total entries
examined; charge every array entry (valid, invalid, or duplicate) and
stop once either hits MAX_FORM_FIELD_NODES. Adds a regression test with a
huge /Kids array of duplicate + null entries.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): iterate /Fields and /Kids arrays by borrow, not clone

Both arrays were cloned in full before the budget check, so a crafted
oversized /Fields or /Kids array forced an O(n) allocation and copy
regardless of the cap. resolve_array already returns a borrow tied to the
document and the walker only needs a shared &Document, so iterate the
borrowed arrays directly — the early break now bounds how many entries
are even touched, before any per-array allocation.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* docs(links): correct wide-array test comments to match range assertions

The two wide-array tests assert item counts within a range near the
budget, not an exact value (charging entries in the entry guard shifts
the boundary by one or two). Fix the stale comments that claimed exact
counts.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-08 23:43:33 -07:00
40 changed files with 4633 additions and 276 deletions
+3
View File
@@ -24,6 +24,9 @@ jobs:
- name: Cache cargo
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Run tests
run: cargo test --verbose
+3
View File
@@ -31,6 +31,9 @@ jobs:
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check if version changed
id: check
run: |
+3
View File
@@ -28,6 +28,9 @@ jobs:
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check if version changed
id: check
run: |
+3
View File
@@ -28,6 +28,9 @@ jobs:
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check package version
id: check
run: |
+3
View File
@@ -28,6 +28,9 @@ jobs:
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check if version changed
id: check
run: |
+1 -2
View File
@@ -31,9 +31,8 @@ Thumbs.db
napi/index.js
napi/index.d.ts
# Local samples and scripts
# Local samples
samples/
scripts/
# Test output
test_output/
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector"
version = "0.1.7"
version = "1.14.1"
edition = "2021"
autobins = false
authors = ["Firecrawl Team"]
+1 -1
View File
@@ -110,7 +110,7 @@ Or add it manually:
```toml
[dependencies]
pdf-inspector = "0.1"
pdf-inspector = "1"
```
```rust
+4 -3
View File
@@ -5,14 +5,15 @@
If you believe you've found a security vulnerability in pdf-inspector, please
report it privately so we can fix it before public disclosure.
**Preferred:** Email **help@firecrawl.dev** with:
**Preferred:** Submit through Firecrawl's Bugcrowd vulnerability disclosure
program at <https://bugcrowd.com/engagements/firecrawl-vdp-ess>. Please include:
- A description of the issue and its impact
- Steps to reproduce (a minimal PDF or input that triggers the bug is ideal)
- The version or commit hash of pdf-inspector you tested against
**Alternative:** Use GitHub's private vulnerability reporting under the
[Security tab](https://github.com/firecrawl/pdf-inspector/security/advisories/new).
**Alternative:** If you'd rather not use Bugcrowd, email
**help@firecrawl.dev** with the same details.
We'll acknowledge your report in a timely manner and keep you updated on
remediation progress. Please do not open a public GitHub issue for security
+43 -24
View File
@@ -1,38 +1,57 @@
# Publishing
The Rust crate is published to [crates.io](https://crates.io/crates/pdf-inspector) with trusted publishing from GitHub Actions. The first release was published manually; future releases publish from `.github/workflows/publish-crate.yml` when a `Cargo.toml` version change lands on `main`.
Every pdf-inspector distribution uses one shared semantic version:
## crates.io Trusted Publisher
- Rust crate: `pdf-inspector`
- Python package: `pdf-inspector`
- Node package: `@firecrawl/pdf-inspector` and its platform packages
- Browser package: `@firecrawl/pdf-inspector-wasm`
- Internal NAPI and WASM Rust crates
Configure the trusted publisher for the `pdf-inspector` crate with:
`Cargo.toml` is the canonical version source. Update every manifest and lockfile
with:
- Repository: `firecrawl/pdf-inspector`
- Workflow: `publish-crate.yml`
- Environment: `crates-io`
```bash
python3 scripts/version.py <version>
```
The workflow uses `rust-lang/crates-io-auth-action@v1` to exchange GitHub's OIDC token for a short-lived crates.io token, then passes it to `cargo publish`.
Verify that nothing has diverged with:
## Release Steps
```bash
python3 scripts/version.py --check
```
1. Update `version` in `Cargo.toml`.
2. Merge the version bump to `main`.
3. The publish workflow compares the new `Cargo.toml` version with `HEAD~1`, runs `cargo publish --dry-run`, then publishes if that version is not already on crates.io.
CI and every publishing workflow run this check before building or publishing.
If `Cargo.toml` changes without a package version bump, the workflow exits without publishing.
## Release steps
## Browser WebAssembly package
1. Choose the next shared semantic version and run `scripts/version.py`.
2. Review the manifest and lockfile changes in the version-bump pull request.
3. Merge the pull request to `main`.
4. The crates.io, PyPI, Node, and WASM workflows independently build and
publish that version from the same commit.
5. After all registries succeed, create one `v<version>` GitHub release that
links to each package and describes changes since the previous shared tag.
The browser package is published as `@firecrawl/pdf-inspector-wasm`. Its version lives in `wasm/Cargo.toml`, and `.github/workflows/publish-wasm.yml` builds the `web` target with `wasm-pack` before publishing the generated package.
The independent workflows are intentionally idempotent. A manual dispatch from
`main` can repair a partial release, and already-published artifacts are skipped.
The npm package must exist before a trusted publisher can be configured. For the first release only:
## Trusted publishers
1. Build with `wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release`.
2. Inspect with `npm pack --dry-run ./wasm/pkg`.
3. Publish with `npm publish ./wasm/pkg --access public` from an authorized maintainer session.
4. In the package settings on npm, configure the GitHub Actions trusted publisher:
- Organization: `firecrawl`
- Repository: `pdf-inspector`
- Workflow: `publish-wasm.yml`
- Allowed action: `npm publish`
The repositories use GitHub Actions OIDC instead of long-lived registry tokens.
Configure each registry's trusted publisher for `firecrawl/pdf-inspector` and
its corresponding workflow:
After that one-time bootstrap, bumping the version in `wasm/Cargo.toml` and merging it to `main` publishes through OIDC. Until the package exists, the workflow exits cleanly without attempting an unauthenticated first publish. See npm's [trusted publishing documentation](https://docs.npmjs.com/trusted-publishers/) for the registry-side setup.
- crates.io: `publish-crate.yml`, environment `crates-io`
- PyPI: `publish-pypi.yml`, environment `pypi`
- npm Node package: `publish.yml`
- npm WASM package: `publish-wasm.yml`
The WASM package must exist before npm trusted publishing can be configured. If
it ever needs to be bootstrapped again, build and inspect it before publishing:
```bash
wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release
npm pack --dry-run ./wasm/pkg
npm publish ./wasm/pkg --access public
```
+19
View File
@@ -80,6 +80,17 @@ for page in result.pages:
# Restrict to specific 0-indexed pages (preserves caller order)
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
# Structure-tree elements from tagged PDFs (empty list when untagged).
# Pages are 1-indexed to match TextItem.page, so (page, mcid) joins directly
# against extract_text_with_positions — e.g. to recover real heading levels:
elements = pdf_inspector.extract_structure_elements("tagged.pdf")
roles = {(e.page, e.mcid): e.role for e in elements}
headings = [
item.text
for item in pdf_inspector.extract_text_with_positions("tagged.pdf")
if item.mcid is not None and roles.get((item.page, item.mcid), "").startswith("H")
]
```
## API reference
@@ -100,6 +111,8 @@ result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes |
| `extract_pages_markdown(path, pages=None)` | Per-page Markdown + layout metadata (all pages by default) |
| `extract_pages_markdown_bytes(data, pages=None)` | Per-page Markdown from bytes |
| `extract_structure_elements(path, pages=None)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
| `extract_structure_elements_bytes(data, pages=None)` | Structure-tree elements from bytes |
## Types
@@ -144,6 +157,12 @@ class TextItem: # extract_text_with_positions
is_underline: bool
is_strikeout: bool
item_type: str
mcid: int | None # marked-content ID for tagged PDFs (None otherwise)
class StructureElement: # extract_structure_elements
page: int # 1-indexed (matches TextItem.page)
mcid: int
role: str # "H1".."H6", "P", "Table", ... (resolved via /RoleMap)
class RegionText: # extract_text_in_regions
text: str
+32 -1
View File
@@ -138,6 +138,34 @@ for page in &result.pages {
println!("Complex layout? {}", result.is_complex);
```
Extract structure-tree elements from tagged PDFs, and join them against
`extract_text_with_positions` to attach semantic roles (heading levels,
paragraphs, table cells) to extracted text:
```rust
use pdf_inspector::{extract_structure_elements, extract_text_with_positions};
use std::collections::HashMap;
// One entry per marked-content reference, sorted by (page, mcid); empty for
// untagged PDFs. Pages are 1-indexed to match `TextItem::page`, so the
// (page, mcid) pair is a direct join key.
let elements = extract_structure_elements("tagged.pdf", None)?;
let roles: HashMap<(u32, i64), &str> = elements
.iter()
.map(|e| ((e.page, e.mcid), e.role.as_str()))
.collect();
for item in extract_text_with_positions("tagged.pdf")? {
if let Some(mcid) = item.mcid {
if let Some(role) = roles.get(&(item.page, mcid)) {
if role.starts_with('H') {
println!("{}: {}", role, item.text);
}
}
}
}
```
## Processing modes
| Mode | What it does | Returns |
@@ -163,6 +191,8 @@ println!("Complex layout? {}", result.is_complex);
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
| `extract_pages_markdown(path, pages)` | Per-page Markdown + layout metadata (file) |
| `extract_pages_markdown_mem(bytes, pages)` | Per-page Markdown from bytes |
| `extract_structure_elements(path, pages)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
| `extract_structure_elements_mem(bytes, pages)` | Structure-tree elements from bytes |
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
@@ -178,7 +208,8 @@ Low-level detection functions are also available via the `detector` module (`det
| `DetectionConfig` | Configuration for detection: scan strategy, thresholds |
| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` |
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
| `TextItem` | Text with position, font info, and page number |
| `TextItem` | Text with position, font info, page number, and optional structure-tree `mcid` |
| `StructureElement` | Tagged-PDF structure reference: page (1-indexed), mcid, role (`"H1"`..`"H6"`, `"P"`, …) |
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
+2 -2
View File
@@ -851,7 +851,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
[[package]]
name = "pdf-inspector"
version = "0.1.7"
version = "1.14.1"
dependencies = [
"env_logger",
"include_dir",
@@ -867,7 +867,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-napi"
version = "0.2.2"
version = "1.14.1"
dependencies = [
"napi",
"napi-build",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-napi"
version = "0.2.2"
version = "1.14.1"
edition = "2021"
[lib]
+16
View File
@@ -83,6 +83,22 @@ for (const region of result[0].regions) {
}
```
### Async variants
`processPdf`, `classifyPdf`, and `extractPagesMarkdown` are synchronous and parse on the calling thread — in Node, that's the event loop. For a one-off call in a script that's fine, but in a server a large document can hold the loop for tens to hundreds of milliseconds.
`processPdfAsync`, `classifyPdfAsync`, and `extractPagesMarkdownAsync` take the same arguments and produce the same results, but run the parse on the libuv thread pool and return a promise, keeping the event loop free. The input buffer is copied before the call returns, so it's safe to reuse or mutate immediately:
```typescript
import { classifyPdfAsync, extractPagesMarkdownAsync } from '@firecrawl/pdf-inspector'
const classification = await classifyPdfAsync(pdf)
if (classification.pdfType === 'TextBased') {
const { pages } = await extractPagesMarkdownAsync(pdf)
// ...
}
```
## Types
```typescript
+6 -6
View File
@@ -8,12 +8,12 @@
"@napi-rs/cli": "^3.4.1",
},
"optionalDependencies": {
"@firecrawl/pdf-inspector-darwin-arm64": "1.12.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.12.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.12.0",
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.12.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.12.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.12.0",
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.1",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.1",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.1",
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.1",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.1",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.1",
},
},
},
+7 -7
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.12.0",
"version": "1.14.1",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
@@ -52,11 +52,11 @@
"@napi-rs/cli": "^3.4.1"
},
"optionalDependencies": {
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.12.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.12.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.12.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.12.0",
"@firecrawl/pdf-inspector-darwin-arm64": "1.12.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.12.0"
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.1",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.1",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.1",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.1",
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.1",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.1"
}
}
+236 -41
View File
@@ -89,6 +89,11 @@ pub struct TextItem {
pub item_type: ItemType,
/// URL for link items, `None` for other types.
pub link_url: Option<String>,
/// Marked Content ID from the content stream's BDC/BMC operator, `None`
/// when the text is not part of marked content. Join with the
/// `page`/`mcid` pairs from [`extractStructureElements`] to attach
/// structure-tree roles (headings, paragraphs, …) in tagged PDFs.
pub mcid: Option<i64>,
}
/// A page's regions for text extraction: (page_index_0based, bboxes).
@@ -153,9 +158,7 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
}
}
fn to_napi_page_ocr_reasons(
reasons: Vec<pdf_inspector::PageOcrReasons>,
) -> Vec<PageOcrReasons> {
fn to_napi_page_ocr_reasons(reasons: Vec<pdf_inspector::PageOcrReasons>) -> Vec<PageOcrReasons> {
reasons
.into_iter()
.map(|reason| PageOcrReasons {
@@ -202,6 +205,31 @@ where
}
}
// ---------------------------------------------------------------------------
// Shared implementations (single body behind sync and async entry points)
// ---------------------------------------------------------------------------
fn process_pdf_impl(bytes: &[u8], pages: Option<Vec<u32>>) -> Result<PdfResult> {
let mut opts = pdf_inspector::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = pdf_inspector::process_pdf_mem_with_options(bytes, opts)
.map_err(|e| to_napi_err(e, "process_pdf"))?;
Ok(to_napi_result(result))
}
fn classify_pdf_impl(bytes: &[u8]) -> Result<PdfClassification> {
let result =
pdf_inspector::classify_pdf_mem(bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
Ok(PdfClassification {
pdf_type: convert_pdf_type(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
}
// ---------------------------------------------------------------------------
// Public NAPI API
// ---------------------------------------------------------------------------
@@ -210,15 +238,7 @@ where
#[napi]
pub fn process_pdf(buffer: Buffer, pages: Option<Vec<u32>>) -> Result<PdfResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("process_pdf", move || {
let mut opts = pdf_inspector::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = pdf_inspector::process_pdf_mem_with_options(&bytes, opts)
.map_err(|e| to_napi_err(e, "process_pdf"))?;
Ok(to_napi_result(result))
})
catch_panic("process_pdf", move || process_pdf_impl(&bytes, pages))
}
/// Fast detection only — no text extraction or markdown.
@@ -238,16 +258,7 @@ pub fn detect_pdf(buffer: Buffer) -> Result<PdfResult> {
#[napi]
pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("classify_pdf", move || {
let result =
pdf_inspector::classify_pdf_mem(&bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
Ok(PdfClassification {
pdf_type: convert_pdf_type(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
})
catch_panic("classify_pdf", move || classify_pdf_impl(&bytes))
}
/// Extract plain text from a PDF Buffer.
@@ -300,12 +311,61 @@ pub fn extract_text_with_positions(
is_strikeout: item.is_strikeout,
item_type,
link_url,
mcid: item.mcid,
}
})
.collect())
})
}
/// One structure-tree element reference from a tagged PDF.
#[napi(object)]
pub struct StructureElementJs {
/// 1-indexed page number (matches `TextItem.page`).
pub page: u32,
/// Marked Content ID from the page's content stream (matches
/// `TextItem.mcid`).
pub mcid: i64,
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
/// Custom tags are resolved through the document's role map; tags with
/// no standard mapping are returned verbatim.
pub role: String,
}
/// Extract structure-tree element references from a tagged PDF.
///
/// Parses the document's structure tree (when present) and returns one
/// entry per marked-content reference, resolved to its 1-indexed page,
/// MCID, and structure type name. Returns an empty array when the PDF is
/// not tagged.
///
/// Join `(page, mcid)` against the `page`/`mcid` fields from
/// [`extractTextWithPositions`] to attach heading levels (H1..H6) and other
/// semantic roles to extracted text.
///
/// Pass 1-indexed page numbers (matching `TextItem.page`) to restrict
/// output; omit `pages` for the whole document. Entries are sorted by
/// `(page, mcid)`.
#[napi]
pub fn extract_structure_elements(
buffer: Buffer,
pages: Option<Vec<u32>>,
) -> Result<Vec<StructureElementJs>> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_structure_elements", move || {
let elements = pdf_inspector::extract_structure_elements_mem(&bytes, pages.as_deref())
.map_err(|e| to_napi_err(e, "extract_structure_elements"))?;
Ok(elements
.into_iter()
.map(|e| StructureElementJs {
page: e.page,
mcid: e.mcid,
role: e.role,
})
.collect())
})
}
/// Extract text within bounding-box regions from a PDF.
///
/// For hybrid OCR: layout model detects regions in rendered images,
@@ -633,25 +693,32 @@ pub fn extract_pages_markdown(
) -> Result<PagesExtractionResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_pages_markdown", move || {
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, pages.as_deref())
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
.pages
.into_iter()
.map(|r| PageMarkdownResult {
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
is_complex: result.is_complex,
})
extract_pages_markdown_impl(&bytes, pages.as_deref())
})
}
fn extract_pages_markdown_impl(
bytes: &[u8],
pages: Option<&[u32]>,
) -> Result<PagesExtractionResult> {
let result = pdf_inspector::extract_pages_markdown_mem(bytes, pages)
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
.pages
.into_iter()
.map(|r| PageMarkdownResult {
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
is_complex: result.is_complex,
})
}
@@ -692,3 +759,131 @@ fn to_page_region_texts(results: Vec<pdf_inspector::PageRegionResult>) -> Vec<Pa
})
.collect()
}
// ---------------------------------------------------------------------------
// Async variants (libuv thread pool via AsyncTask)
//
// The synchronous exports above parse on the calling thread, which in Node is
// the event loop. These `*Async` variants run the same shared implementations
// on the libuv thread pool and hand JavaScript a promise, so servers under
// concurrent load keep answering requests while a document parses. The sync
// exports keep their names, signatures, and behaviour.
//
// Each factory copies the input Buffer to an owned `Vec<u8>` on the calling
// (JS) thread — deliberately. JS execution is single-threaded, so no JS code
// can mutate the buffer while the synchronous part of the call copies it.
// Holding the napi `Buffer` and reading it from the worker instead would be
// zero-copy, but a caller mutating the buffer before the promise settles
// would then race the worker's reads — undefined behavior, not a recoverable
// error (a known napi-rs soundness hazard with cross-thread Buffer access).
// The copy is a one-time memcpy, negligible next to the parse it unblocks.
// ---------------------------------------------------------------------------
pub struct ProcessPdfTask {
bytes: Vec<u8>,
pages: Option<Vec<u32>>,
}
impl Task for ProcessPdfTask {
type Output = PdfResult;
type JsValue = PdfResult;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
let pages = self.pages.take();
// AssertUnwindSafe: `bytes`/`pages` are moved into the closure and
// dropped on unwind — no shared state can be observed broken.
catch_panic(
"process_pdf",
panic::AssertUnwindSafe(move || process_pdf_impl(&bytes, pages)),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`processPdf`]: same result, but the parse runs on the
/// libuv thread pool instead of the event loop and the call returns a
/// promise. The buffer is copied before the call returns, so it may be
/// reused or mutated immediately.
// ts_return_type is required: napi-rs emits `Promise<unknown>` for
// `AsyncTask<T>` returns without it.
#[napi(ts_return_type = "Promise<PdfResult>")]
pub fn process_pdf_async(buffer: Buffer, pages: Option<Vec<u32>>) -> AsyncTask<ProcessPdfTask> {
AsyncTask::new(ProcessPdfTask {
bytes: buffer.to_vec(),
pages,
})
}
pub struct ClassifyPdfTask {
bytes: Vec<u8>,
}
impl Task for ClassifyPdfTask {
type Output = PdfClassification;
type JsValue = PdfClassification;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
catch_panic(
"classify_pdf",
panic::AssertUnwindSafe(move || classify_pdf_impl(&bytes)),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`classifyPdf`]: same result, but the classification runs
/// on the libuv thread pool instead of the event loop and the call returns a
/// promise. The buffer is copied before the call returns, so it may be
/// reused or mutated immediately.
#[napi(ts_return_type = "Promise<PdfClassification>")]
pub fn classify_pdf_async(buffer: Buffer) -> AsyncTask<ClassifyPdfTask> {
AsyncTask::new(ClassifyPdfTask {
bytes: buffer.to_vec(),
})
}
pub struct ExtractPagesMarkdownTask {
bytes: Vec<u8>,
pages: Option<Vec<u32>>,
}
impl Task for ExtractPagesMarkdownTask {
type Output = PagesExtractionResult;
type JsValue = PagesExtractionResult;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
let pages = self.pages.take();
catch_panic(
"extract_pages_markdown",
panic::AssertUnwindSafe(move || extract_pages_markdown_impl(&bytes, pages.as_deref())),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`extractPagesMarkdown`]: same result, but the extraction
/// runs on the libuv thread pool instead of the event loop and the call
/// returns a promise. The buffer is copied before the call returns, so it
/// may be reused or mutated immediately.
#[napi(ts_return_type = "Promise<PagesExtractionResult>")]
pub fn extract_pages_markdown_async(
buffer: Buffer,
pages: Option<Vec<u32>>,
) -> AsyncTask<ExtractPagesMarkdownTask> {
AsyncTask::new(ExtractPagesMarkdownTask {
bytes: buffer.to_vec(),
pages,
})
}
+110
View File
@@ -2,16 +2,21 @@ import { readFileSync } from 'fs';
import { strict as assert } from 'assert';
import {
processPdf,
processPdfAsync,
detectPdf,
classifyPdf,
classifyPdfAsync,
extractText,
extractTextWithPositions,
extractStructureElements,
extractTextInRegions,
detectVectorGridInRegion,
extractPagesMarkdown,
extractPagesMarkdownAsync,
} from './index.js';
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
const taggedFixture = readFileSync('../tests/fixtures/firecrawl_docs_tagged.pdf');
// --- processPdf ---
console.log('Testing processPdf...');
@@ -79,6 +84,46 @@ assert.ok(page1Items.length > 0);
assert.ok(page1Items.every(i => i.page === 1));
console.log(' extractTextWithPositions with pages: OK');
// mcid: undefined on untagged PDFs, numeric on tagged marked content
assert.ok(items.every(i => i.mcid === undefined || typeof i.mcid === 'number'));
const taggedItems = extractTextWithPositions(taggedFixture);
assert.ok(
taggedItems.some(i => typeof i.mcid === 'number'),
'tagged PDF text items should carry Marked Content IDs',
);
console.log(' extractTextWithPositions mcid: OK');
// --- extractStructureElements ---
console.log('Testing extractStructureElements...');
const structureElements = extractStructureElements(taggedFixture);
assert.ok(structureElements.length > 0);
assert.ok(structureElements.every(e => typeof e.page === 'number'));
assert.ok(structureElements.every(e => typeof e.mcid === 'number'));
assert.ok(structureElements.every(e => typeof e.role === 'string' && e.role.length > 0));
assert.ok(
structureElements.some(e => e.role === 'H1'),
'tagged fixture should surface H1 heading roles',
);
// (page, mcid) joins against extractTextWithPositions to recover heading text
const h1Refs = new Set(
structureElements.filter(e => e.role === 'H1').map(e => `${e.page}:${e.mcid}`),
);
const h1Text = taggedItems
.filter(i => typeof i.mcid === 'number' && h1Refs.has(`${i.page}:${i.mcid}`))
.map(i => i.text)
.join('');
assert.ok(h1Text.trim().length > 0, 'H1 join should recover heading text');
// pages filter is 1-indexed, matching TextItem.page
const page1Elements = extractStructureElements(taggedFixture, [1]);
assert.ok(page1Elements.length > 0);
assert.ok(page1Elements.every(e => e.page === 1));
// untagged PDFs yield an empty array
assert.deepEqual(extractStructureElements(fixture), []);
console.log(' extractStructureElements: OK');
// --- extractTextInRegions ---
console.log('Testing extractTextInRegions...');
const regionResults = extractTextInRegions(fixture, [
@@ -124,10 +169,75 @@ assert.equal(picked.pages[0].page, 2);
assert.equal(picked.pages[1].page, 0);
console.log(' extractPagesMarkdown with pages: OK');
// --- Async variants ---
console.log('Testing async variants...');
// processPdfAsync returns a promise and matches the sync result
const asyncResultPromise = processPdfAsync(fixture);
assert.ok(asyncResultPromise instanceof Promise);
const asyncResult = await asyncResultPromise;
assert.equal(asyncResult.pdfType, result.pdfType);
assert.equal(asyncResult.pageCount, result.pageCount);
assert.equal(asyncResult.markdown, result.markdown);
console.log(' processPdfAsync: OK');
// processPdfAsync with pages
const asyncResult2 = await processPdfAsync(fixture, [1]);
assert.equal(asyncResult2.markdown, result2.markdown);
console.log(' processPdfAsync with pages: OK');
// classifyPdfAsync matches the sync result
const asyncClassified = await classifyPdfAsync(fixture);
assert.equal(asyncClassified.pdfType, classified.pdfType);
assert.equal(asyncClassified.pageCount, classified.pageCount);
assert.equal(asyncClassified.confidence, classified.confidence);
assert.deepEqual(asyncClassified.pagesNeedingOcr, classified.pagesNeedingOcr);
console.log(' classifyPdfAsync: OK');
// extractPagesMarkdownAsync matches the sync result
const asyncAllPages = await extractPagesMarkdownAsync(fixture);
assert.equal(asyncAllPages.pages.length, allPages.pages.length);
assert.deepEqual(
asyncAllPages.pages.map(p => p.markdown),
allPages.pages.map(p => p.markdown),
);
assert.equal(asyncAllPages.isComplex, allPages.isComplex);
console.log(' extractPagesMarkdownAsync: OK');
// selected pages preserve caller order
const asyncPicked = await extractPagesMarkdownAsync(fixture, [2, 0]);
assert.equal(asyncPicked.pages.length, 2);
assert.equal(asyncPicked.pages[0].page, 2);
assert.equal(asyncPicked.pages[1].page, 0);
console.log(' extractPagesMarkdownAsync with pages: OK');
// input buffer is copied at call time: mutating it immediately after the
// call must not affect the in-flight parse
const scratch = Buffer.from(fixture);
const inFlight = processPdfAsync(scratch);
scratch.fill(0);
const fromMutated = await inFlight;
assert.equal(fromMutated.markdown, result.markdown);
console.log(' processPdfAsync input copied at call time: OK');
// concurrent async calls all settle
const [c1, c2, c3] = await Promise.all([
processPdfAsync(fixture),
classifyPdfAsync(fixture),
extractPagesMarkdownAsync(fixture),
]);
assert.equal(c1.pdfType, 'TextBased');
assert.equal(c2.pdfType, 'TextBased');
assert.equal(c3.pages.length, 3);
console.log(' concurrent async calls: OK');
// --- Error handling ---
console.log('Testing error handling...');
assert.throws(() => processPdf(Buffer.from('not a pdf')), /process_pdf/);
assert.throws(() => classifyPdf(Buffer.from('')), /classify_pdf/);
await assert.rejects(processPdfAsync(Buffer.from('not a pdf')), /process_pdf/);
await assert.rejects(classifyPdfAsync(Buffer.from('')), /classify_pdf/);
await assert.rejects(extractPagesMarkdownAsync(Buffer.from('')), /extract_pages_markdown/);
console.log(' error handling: OK');
console.log('\nAll NAPI tests passed!');
+35
View File
@@ -51,6 +51,20 @@ class TextItem:
is_underline: bool
is_strikeout: bool
item_type: str
mcid: Optional[int]
"""Marked Content ID from the content stream's BDC/BMC operator, None when
the text is not part of marked content. Join with the (page, mcid) pairs
from extract_structure_elements to attach structure-tree roles in tagged
PDFs."""
class StructureElement:
"""One structure-tree element reference from a tagged PDF."""
page: int
"""1-indexed page number (matches TextItem.page)."""
mcid: int
"""Marked Content ID from the page's content stream (matches TextItem.mcid)."""
role: str
"""Standard structure type name ("H1".."H6", "P", "Table", "TD", ...)."""
class RegionText:
"""Extracted text for a single region."""
@@ -132,6 +146,27 @@ def extract_text_with_positions_bytes(data: bytes, pages: Optional[list[int]] =
"""Extract text with position information from bytes."""
...
def extract_structure_elements(path: str, pages: Optional[list[int]] = None) -> list[StructureElement]:
"""Extract structure-tree element references from a tagged PDF file.
Returns one entry per marked-content reference, resolved to its 1-indexed
page, MCID, and structure type name ("H1".."H6", "P", "Table", ...), sorted
by (page, mcid). Returns an empty list when the PDF is not tagged.
Args:
path: Path to the PDF file.
pages: Optional list of 1-indexed pages (matching ``TextItem.page``).
When ``None`` (default), the whole document is returned.
"""
...
def extract_structure_elements_bytes(data: bytes, pages: Optional[list[int]] = None) -> list[StructureElement]:
"""Extract structure-tree element references from tagged PDF bytes.
See :func:`extract_structure_elements` for details.
"""
...
def extract_text_in_regions(
path: str,
page_regions: list[tuple[int, list[list[float]]]],
+3 -3
View File
@@ -4,9 +4,9 @@ build-backend = "maturin"
[project]
name = "pdf-inspector"
# Bump this to publish to PyPI — CI publishes automatically when the version
# changes on main (same flow as napi/package.json for npm).
version = "0.2.6"
# Keep package versions in sync with `python3 scripts/version.py <version>`.
# CI publishes automatically when the synchronized change lands on main.
version = "1.14.1"
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
readme = "docs/python.md"
license = { text = "MIT" }
+111
View File
@@ -0,0 +1,111 @@
import json
import sys
import tempfile
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from version import PLATFORM_PACKAGES, check_versions, set_versions
class VersionTests(unittest.TestCase):
def setUp(self):
self.temporary = tempfile.TemporaryDirectory()
self.root = Path(self.temporary.name)
(self.root / "napi").mkdir()
(self.root / "site").mkdir()
(self.root / "wasm").mkdir()
self._write_manifest("Cargo.toml", "package", "0.1.0")
self._write_manifest("pyproject.toml", "project", "0.1.0")
self._write_manifest("napi/Cargo.toml", "package", "0.1.0")
self._write_manifest("wasm/Cargo.toml", "package", "0.1.0")
package = {
"name": "@firecrawl/pdf-inspector",
"version": "0.1.0",
"optionalDependencies": {
dependency: "0.1.0" for dependency in PLATFORM_PACKAGES
},
}
(self.root / "napi/package.json").write_text(
json.dumps(package), encoding="utf-8"
)
(self.root / "napi/bun.lock").write_text(
"\n".join(
f' "{dependency}": "0.1.0",'
for dependency in PLATFORM_PACKAGES
)
+ "\n",
encoding="utf-8",
)
(self.root / "site/index.html").write_text(
'https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.0/'
'pdf_inspector_wasm.js\n',
encoding="utf-8",
)
self._write_lock(
"napi/Cargo.lock", ("pdf-inspector", "pdf-inspector-napi")
)
self._write_lock(
"wasm/Cargo.lock", ("pdf-inspector", "pdf-inspector-wasm")
)
def tearDown(self):
self.temporary.cleanup()
def _write_manifest(self, relative, section, version):
(self.root / relative).write_text(
f'[{section}]\nname = "fixture"\nversion = "{version}"\n',
encoding="utf-8",
)
def _write_lock(self, relative, packages):
content = "\n".join(
f'[[package]]\nname = "{package}"\nversion = "0.1.0"\n'
for package in packages
)
(self.root / relative).write_text(content, encoding="utf-8")
def test_updates_every_version_location(self):
set_versions("1.14.0", self.root)
self.assertEqual(check_versions(self.root), "1.14.0")
def test_reports_a_divergent_package(self):
self._write_manifest("wasm/Cargo.toml", "package", "0.2.0")
with self.assertRaisesRegex(ValueError, "WASM package: 0.2.0"):
check_versions(self.root)
def test_rejects_an_invalid_version(self):
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
set_versions("next", self.root)
def test_rejects_numeric_prerelease_with_leading_zero(self):
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
set_versions("1.2.3-01", self.root)
self.assertEqual(
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
)
def test_preflight_failure_does_not_partially_update(self):
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
(self.root / "site/index.html").write_text(
"missing module URL\n", encoding="utf-8"
)
with self.assertRaisesRegex(ValueError, "Missing pinned WASM package URL"):
set_versions("1.14.0", self.root)
self.assertEqual(
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
)
if __name__ == "__main__":
unittest.main()
+262
View File
@@ -0,0 +1,262 @@
#!/usr/bin/env python3
"""Keep every pdf-inspector package on one release version."""
from __future__ import annotations
import argparse
import json
import re
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
PRERELEASE_IDENTIFIER = (
r"(?:0|[1-9]\d*|[0-9A-Za-z-]*[A-Za-z-][0-9A-Za-z-]*)"
)
SEMVER = re.compile(
r"^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)"
rf"(?:-{PRERELEASE_IDENTIFIER}(?:\.{PRERELEASE_IDENTIFIER})*)?"
r"(?:\+[0-9A-Za-z-]+(?:\.[0-9A-Za-z-]+)*)?$"
)
VERSION_LINE = re.compile(r'^(\s*version\s*=\s*")[^"]+(".*)$')
SECTION_LINE = re.compile(r"^\s*\[([^]]+)]\s*$")
PLATFORM_PACKAGES = (
"@firecrawl/pdf-inspector-linux-x64-gnu",
"@firecrawl/pdf-inspector-linux-x64-musl",
"@firecrawl/pdf-inspector-linux-arm64-gnu",
"@firecrawl/pdf-inspector-linux-arm64-musl",
"@firecrawl/pdf-inspector-darwin-arm64",
"@firecrawl/pdf-inspector-win32-x64-msvc",
)
TOML_VERSIONS = (
("Rust crate", Path("Cargo.toml"), "package"),
("Python package", Path("pyproject.toml"), "project"),
("NAPI crate", Path("napi/Cargo.toml"), "package"),
("WASM package", Path("wasm/Cargo.toml"), "package"),
)
LOCK_VERSIONS = (
("NAPI lock: core", Path("napi/Cargo.lock"), "pdf-inspector"),
("NAPI lock: binding", Path("napi/Cargo.lock"), "pdf-inspector-napi"),
("WASM lock: core", Path("wasm/Cargo.lock"), "pdf-inspector"),
("WASM lock: binding", Path("wasm/Cargo.lock"), "pdf-inspector-wasm"),
)
SITE_WASM_VERSION = re.compile(
r"(@firecrawl/pdf-inspector-wasm@)([^/\"]+)(/pdf_inspector_wasm\.js)"
)
def _read_section_version(path: Path, section: str) -> str:
active = False
for line in path.read_text(encoding="utf-8").splitlines():
section_match = SECTION_LINE.match(line)
if section_match:
active = section_match.group(1) == section
elif active:
version_match = VERSION_LINE.match(line)
if version_match:
return line.split('"', 2)[1]
raise ValueError(f"No version found in [{section}] of {path}")
def _write_section_version(path: Path, section: str, version: str) -> None:
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
active = False
for index, line in enumerate(lines):
section_match = SECTION_LINE.match(line)
if section_match:
active = section_match.group(1) == section
elif active:
version_match = VERSION_LINE.match(line)
if version_match:
newline = "\n" if line.endswith("\n") else ""
replacement = (
f"{version_match.group(1)}{version}"
f"{version_match.group(2).rstrip()}"
)
lines[index] = (
f"{replacement}{newline}"
)
path.write_text("".join(lines), encoding="utf-8")
return
raise ValueError(f"No version found in [{section}] of {path}")
def _package_block(lines: list[str], package: str) -> tuple[int, int]:
for start, line in enumerate(lines):
if line.strip() != "[[package]]":
continue
end = next(
(
index
for index in range(start + 1, len(lines))
if lines[index].strip() == "[[package]]"
),
len(lines),
)
if any(line.strip() == f'name = "{package}"' for line in lines[start:end]):
return start, end
raise ValueError(f"No lockfile entry found for {package}")
def _read_lock_version(path: Path, package: str) -> str:
lines = path.read_text(encoding="utf-8").splitlines()
start, end = _package_block(lines, package)
for line in lines[start:end]:
version_match = VERSION_LINE.match(line)
if version_match:
return line.split('"', 2)[1]
raise ValueError(f"No version found for {package} in {path}")
def _write_lock_version(path: Path, package: str, version: str) -> None:
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
start, end = _package_block(lines, package)
for index in range(start, end):
version_match = VERSION_LINE.match(lines[index])
if version_match:
newline = "\n" if lines[index].endswith("\n") else ""
lines[index] = (
f'{version_match.group(1)}{version}{version_match.group(2).rstrip()}'
f"{newline}"
)
path.write_text("".join(lines), encoding="utf-8")
return
raise ValueError(f"No version found for {package} in {path}")
def _node_versions(root: Path) -> dict[str, str]:
package = json.loads((root / "napi/package.json").read_text(encoding="utf-8"))
versions = {"Node package": package["version"]}
optional = package.get("optionalDependencies", {})
for dependency in PLATFORM_PACKAGES:
if dependency not in optional:
raise ValueError(f"Missing Node optional dependency: {dependency}")
versions[f"Node optional dependency: {dependency}"] = optional[dependency]
return versions
def _bun_versions(root: Path) -> dict[str, str]:
text = (root / "napi/bun.lock").read_text(encoding="utf-8")
versions = {}
for dependency in PLATFORM_PACKAGES:
match = re.search(
rf'"{re.escape(dependency)}": "([^"]+)"[,]', text
)
if not match:
raise ValueError(f"Missing Bun lock dependency: {dependency}")
versions[f"Bun lock: {dependency}"] = match.group(1)
return versions
def _site_wasm_version(root: Path) -> str:
text = (root / "site/index.html").read_text(encoding="utf-8")
match = SITE_WASM_VERSION.search(text)
if not match:
raise ValueError("Missing pinned WASM package URL in site/index.html")
return match.group(2)
def package_versions(root: Path = ROOT) -> dict[str, str]:
versions = {
label: _read_section_version(root / relative, section)
for label, relative, section in TOML_VERSIONS
}
versions.update(_node_versions(root))
versions.update(_bun_versions(root))
versions["Website WASM module"] = _site_wasm_version(root)
versions.update(
{
label: _read_lock_version(root / relative, package)
for label, relative, package in LOCK_VERSIONS
}
)
return versions
def check_versions(root: Path = ROOT) -> str:
versions = package_versions(root)
expected = versions["Rust crate"]
if not SEMVER.fullmatch(expected):
raise ValueError(f"Rust crate has an invalid semantic version: {expected}")
mismatches = {
label: version for label, version in versions.items() if version != expected
}
if mismatches:
details = "\n".join(
f" - {label}: {version}" for label, version in mismatches.items()
)
raise ValueError(f"Expected every package to use {expected}:\n{details}")
return expected
def set_versions(version: str, root: Path = ROOT) -> None:
if not SEMVER.fullmatch(version):
raise ValueError(f"Invalid semantic version: {version}")
# Validate every expected location before writing the first file. This
# prevents a stale manifest or generated file from leaving a partial bump.
package_versions(root)
for _, relative, section in TOML_VERSIONS:
_write_section_version(root / relative, section, version)
package_path = root / "napi/package.json"
package = json.loads(package_path.read_text(encoding="utf-8"))
package["version"] = version
optional = package.get("optionalDependencies", {})
for dependency in PLATFORM_PACKAGES:
if dependency not in optional:
raise ValueError(f"Missing Node optional dependency: {dependency}")
optional[dependency] = version
package_path.write_text(json.dumps(package, indent=2) + "\n", encoding="utf-8")
bun_path = root / "napi/bun.lock"
bun_text = bun_path.read_text(encoding="utf-8")
for dependency in PLATFORM_PACKAGES:
pattern = rf'("{re.escape(dependency)}": ")[^"]+("[,])'
bun_text, count = re.subn(
pattern, rf"\g<1>{version}\g<2>", bun_text, count=1
)
if count != 1:
raise ValueError(f"Missing Bun lock dependency: {dependency}")
bun_path.write_text(bun_text, encoding="utf-8")
site_path = root / "site/index.html"
site_text = site_path.read_text(encoding="utf-8")
site_text, count = SITE_WASM_VERSION.subn(
rf"\g<1>{version}\g<3>", site_text, count=1
)
if count != 1:
raise ValueError("Missing pinned WASM package URL in site/index.html")
site_path.write_text(site_text, encoding="utf-8")
for _, relative, package_name in LOCK_VERSIONS:
_write_lock_version(root / relative, package_name, version)
check_versions(root)
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("version", nargs="?", help="new shared semantic version")
parser.add_argument(
"--check", action="store_true", help="fail if package versions have diverged"
)
arguments = parser.parse_args()
if arguments.check == bool(arguments.version):
parser.error("provide either a version or --check")
try:
if arguments.check:
version = check_versions()
print(f"All packages use {version}")
else:
set_versions(arguments.version)
print(f"Updated all packages to {arguments.version}")
except ValueError as error:
parser.exit(1, f"{error}\n")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+1 -1
View File
@@ -975,7 +975,7 @@ result = pdf_inspector.<span class="fn">process_pdf</span>(<span class="str">"do
<script>
(() => {
const MAX_FILE_SIZE = 25 * 1024 * 1024;
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.1/pdf_inspector_wasm.js";
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@1.14.1/pdf_inspector_wasm.js";
const input = document.querySelector("#pdf-input");
const dropZone = document.querySelector("#drop-zone");
const filePanel = document.querySelector("#demo-file");
+328
View File
@@ -0,0 +1,328 @@
//! Bounded content-stream decoding.
//!
//! `lopdf::content::Content::decode` materializes every operator before any
//! caller can apply a limit. A compact page of `q Q` pairs can therefore
//! allocate hundreds of megabytes and abort. Count operators first (without
//! allocating `Operation` objects) and skip decode when the cap is exceeded.
use crate::PdfError;
use lopdf::content::Content;
/// Maximum content-stream operators decoded for a page or a single Form
/// XObject. Matches the previous post-decode skip threshold.
pub(crate) const MAX_PAGE_OPERATIONS: usize = 1_000_000;
/// Decode `data` unless it contains more than `max_operations` operators.
///
/// Returns `Ok(None)` when the stream exceeds the cap, so callers can skip
/// extraction without first allocating the operation vector.
pub(crate) fn decode_content_bounded(
data: &[u8],
max_operations: usize,
) -> Result<Option<Content>, PdfError> {
if content_exceeds_operation_limit(data, max_operations) {
return Ok(None);
}
Content::decode(data)
.map(Some)
.map_err(|e| PdfError::Parse(e.to_string()))
}
fn content_exceeds_operation_limit(data: &[u8], max_operations: usize) -> bool {
count_content_operators(data, max_operations.saturating_add(1)) > max_operations
}
/// Count operators using the same token rules as lopdf's content parser,
/// stopping at `limit`. Does not allocate `Operation` / `Object` values.
fn count_content_operators(data: &[u8], limit: usize) -> usize {
let mut i = 0;
let mut count = 0;
while i < data.len() && count < limit {
skip_content_space(data, &mut i);
if i >= data.len() {
break;
}
if data[i] == b'%' {
skip_comment(data, &mut i);
continue;
}
match data[i] {
b'(' => i = skip_literal_string(data, i),
b'<' => {
if data.get(i + 1) == Some(&b'<') {
i += 2;
} else {
i = skip_hex_string(data, i);
}
}
b'>' => {
i += 1;
if data.get(i) == Some(&b'>') {
i += 1;
}
}
b'[' | b']' => i += 1,
b'/' => skip_name(data, &mut i),
b'+' | b'-' | b'.' => skip_number(data, &mut i),
b if b.is_ascii_digit() => skip_number(data, &mut i),
b if is_operator_byte(b) => {
let start = i;
i += 1;
while i < data.len() && is_operator_byte(data[i]) {
i += 1;
}
let token = &data[start..i];
if token == b"true" || token == b"false" || token == b"null" {
continue;
}
count += 1;
if token == b"BI" && (i >= data.len() || is_content_space(data[i])) {
i = skip_inline_image_after_bi(data, i);
}
}
_ => i += 1,
}
}
count
}
fn is_content_space(b: u8) -> bool {
// PDF whitespace (ISO 32000): NUL, tab, LF, FF, CR, space. Names must
// stop on these so a following operator is not absorbed into `/Name`.
matches!(b, b'\0' | b'\t' | b'\n' | b'\x0c' | b'\r' | b' ')
}
fn is_operator_byte(b: u8) -> bool {
b.is_ascii_alphabetic() || matches!(b, b'*' | b'\'' | b'"')
}
fn is_delimiter(b: u8) -> bool {
matches!(
b,
b'(' | b')' | b'<' | b'>' | b'[' | b']' | b'{' | b'}' | b'/' | b'%'
)
}
fn skip_content_space(data: &[u8], i: &mut usize) {
while *i < data.len() && is_content_space(data[*i]) {
*i += 1;
}
}
fn skip_comment(data: &[u8], i: &mut usize) {
while *i < data.len() && data[*i] != b'\n' && data[*i] != b'\r' {
*i += 1;
}
}
fn skip_literal_string(data: &[u8], mut i: usize) -> usize {
let mut depth = 1i32;
i += 1;
while i < data.len() && depth > 0 {
match data[i] {
b'\\' => {
i += 1;
if i < data.len() {
i += 1;
}
}
b'(' => {
depth += 1;
i += 1;
}
b')' => {
depth -= 1;
i += 1;
}
_ => i += 1,
}
}
i
}
fn skip_hex_string(data: &[u8], mut i: usize) -> usize {
i += 1;
while i < data.len() && data[i] != b'>' {
i += 1;
}
if i < data.len() {
i += 1;
}
i
}
fn skip_name(data: &[u8], i: &mut usize) {
*i += 1;
while *i < data.len() && !is_content_space(data[*i]) && !is_delimiter(data[*i]) {
*i += 1;
}
}
fn skip_number(data: &[u8], i: &mut usize) {
if *i < data.len() && matches!(data[*i], b'+' | b'-') {
*i += 1;
}
while *i < data.len() && data[*i].is_ascii_digit() {
*i += 1;
}
if *i < data.len() && data[*i] == b'.' {
*i += 1;
while *i < data.len() && data[*i].is_ascii_digit() {
*i += 1;
}
}
}
/// After a `BI` operator, skip inline-image data through `EI`.
/// Uses the same PDF whitespace set as `is_content_space`. If `EI` is not
/// found, leave the cursor in place so later operators are still counted
/// (undercounting would let decode allocate the full vector).
fn skip_inline_image_after_bi(data: &[u8], mut i: usize) -> usize {
skip_content_space(data, &mut i);
let rest = &data[i..];
if let Some(pos) = rest.windows(4).position(|w| {
is_content_space(w[0]) && w[1] == b'E' && w[2] == b'I' && is_content_space(w[3])
}) {
return i + pos + 3;
}
i
}
#[cfg(test)]
mod tests {
use super::*;
fn lopdf_op_count(data: &[u8]) -> usize {
Content::decode(data)
.map(|c| c.operations.len())
.unwrap_or(0)
}
/// DoS safety: never report fewer operators than lopdf would allocate.
/// Overcount is acceptable (skip a page); undercount would re-open decode.
fn assert_count_does_not_undercount(data: &[u8]) {
let ours = count_content_operators(data, usize::MAX);
match Content::decode(data) {
Ok(content) => assert!(
ours >= content.operations.len(),
"undercount: ours={ours} lopdf={} for {:?}",
content.operations.len(),
String::from_utf8_lossy(data)
),
Err(_) => {}
}
}
#[test]
fn operator_count_matches_lopdf_for_typical_streams() {
let samples: &[&[u8]] = &[
b"q 1 0 0 1 0 0 cm BT /F1 12 Tf 72 720 Td (Hello) Tj ET Q",
b"q Q q Q",
b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (first) Tj (struck) ' ET",
b"1 0 0 rg 0 0 10 10 re f",
b"true false null q",
b"% comment\nq Q\n",
b"[ (a) 1 (b) ] TJ",
b"1 0 0 1 0 0 cm /Im0 Do",
];
for data in samples {
assert_eq!(
count_content_operators(data, usize::MAX),
lopdf_op_count(data),
"count mismatch for {}",
String::from_utf8_lossy(data)
);
}
}
#[test]
fn strings_and_comments_are_not_operators() {
let data = b"(q Q Tj) Tj % q Q\nET";
assert_eq!(
count_content_operators(data, usize::MAX),
lopdf_op_count(data)
);
assert_eq!(count_content_operators(data, usize::MAX), 2); // Tj, ET
}
#[test]
fn inline_image_counts_as_one_operator() {
let data = b"BI /W 2 /H 2 /CS /RGB /BPC 8 ID \x00\x01\x02\x03 EI q";
assert_eq!(
count_content_operators(data, usize::MAX),
lopdf_op_count(data)
);
assert_eq!(count_content_operators(data, usize::MAX), 2); // BI, q
}
#[test]
fn inline_image_ei_accepts_pdf_whitespace() {
let tab = b"BI /W 1 /H 1 ID \xff\tEI\t q Q";
let nul = b"BI /W 1 /H 1 ID \xff\x00EI\x00 q Q";
let ff = b"BI /W 1 /H 1 ID \xff\x0cEI\x0c q Q";
for data in [tab.as_slice(), nul.as_slice(), ff.as_slice()] {
assert_count_does_not_undercount(data);
assert!(
count_content_operators(data, usize::MAX) >= 3,
"BI plus following q Q must remain visible after EI, got {} for {:?}",
count_content_operators(data, usize::MAX),
String::from_utf8_lossy(data)
);
}
}
#[test]
fn decode_is_skipped_when_operator_cap_is_exceeded() {
let mut data = Vec::new();
for _ in 0..20 {
data.extend_from_slice(b"q Q\n");
}
assert!(decode_content_bounded(&data, 10).unwrap().is_none());
let decoded = decode_content_bounded(&data, 50).unwrap().unwrap();
assert_eq!(decoded.operations.len(), 40);
}
#[test]
fn name_whitespace_does_not_swallow_following_operator() {
// NUL / form-feed end a name (PDF whitespace). Absorbing `q` into
// `/x` would undercount and let decode allocate the operator vector.
let mut nul_sep = Vec::new();
let mut ff_sep = Vec::new();
for _ in 0..8_000 {
nul_sep.extend_from_slice(b"/x\x00q");
ff_sep.extend_from_slice(b"/x\x0cq");
}
assert_count_does_not_undercount(&nul_sep);
assert_count_does_not_undercount(&ff_sep);
assert!(count_content_operators(&ff_sep, usize::MAX) >= 8_000);
}
#[test]
fn edge_streams_do_not_undercount_vs_lopdf() {
let samples: &[&[u8]] = &[
b".5 0 0 .5 0 0 cm",
b"+1 -2 3.0 rg",
b"<0041> Tj",
b"(unbalanced",
b"BI /W 1 /H 1 ID \xff\xff no EI here q Q q Q",
b"q\x00Q\x00q\x00Q",
b"/F1\x0c12 Tf (Hi) Tj",
b"{ 1 2 add } cvx",
];
for data in samples {
assert_count_does_not_undercount(data);
}
}
#[test]
fn million_q_pairs_are_rejected_without_decode() {
let mut data = Vec::with_capacity((MAX_PAGE_OPERATIONS + 1) * 2);
for _ in 0..=MAX_PAGE_OPERATIONS {
data.extend_from_slice(b"q\n");
}
assert!(content_exceeds_operation_limit(&data, MAX_PAGE_OPERATIONS));
assert!(decode_content_bounded(&data, MAX_PAGE_OPERATIONS)
.unwrap()
.is_none());
}
}
+87 -22
View File
@@ -19,7 +19,7 @@ use super::fonts::{
CMapDecisionCache, FontStyleCache,
};
use super::underline::UnderlineLine;
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, FormWalkBudget, XObjectType};
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
/// Strip PDF comments (% to end of line) from content stream bytes.
@@ -137,8 +137,11 @@ fn rise_adjusted(tm: &[f32; 6], rise: f32) -> [f32; 6] {
]
}
/// Returns `(page_extraction, has_gid_fonts)` where `has_gid_fonts` indicates
/// the page uses fonts with unresolvable gid-encoded glyphs.
/// Returns `(page_extraction, has_gid_fonts, coords_rotated, skipped_invisible)`
/// where `has_gid_fonts` indicates the page uses fonts with unresolvable
/// gid-encoded glyphs and `skipped_invisible` reports that invisible (Tr 3)
/// text was present but suppressed — callers can use it to decide whether an
/// `include_invisible` retry could recover anything at all.
pub(crate) fn extract_page_text_items(
doc: &Document,
page_id: ObjectId,
@@ -146,9 +149,8 @@ pub(crate) fn extract_page_text_items(
font_cmaps: &FontCMaps,
include_invisible: bool,
style_cache: &mut FontStyleCache,
) -> Result<(PageExtraction, bool, bool), PdfError> {
use lopdf::content::Content;
form_budget: &mut FormWalkBudget,
) -> Result<(PageExtraction, bool, bool, bool), PdfError> {
let mut items = Vec::new();
let mut rects: Vec<PdfRect> = Vec::new();
let mut clip_rects: Vec<PdfRect> = Vec::new();
@@ -252,22 +254,27 @@ pub(crate) fn extract_page_text_items(
// Content::decode parser, causing it to skip operators like ET and Q.
let content_data = strip_pdf_comments(&content_data);
let content = Content::decode(&content_data).map_err(|e| PdfError::Parse(e.to_string()))?;
const MAX_OPERATIONS: usize = 1_000_000;
if content.operations.len() > MAX_OPERATIONS {
log::warn!(
"page {}: skipping extraction — {} operations exceeds limit ({})",
page_num,
content.operations.len(),
MAX_OPERATIONS
);
return Ok(((Vec::new(), Vec::new(), Vec::new()), false, false));
}
let content = match super::content_decode::decode_content_bounded(
&content_data,
super::content_decode::MAX_PAGE_OPERATIONS,
)? {
Some(content) => content,
None => {
log::warn!(
"page {}: skipping extraction — content stream exceeds {} operations",
page_num,
super::content_decode::MAX_PAGE_OPERATIONS
);
return Ok(((Vec::new(), Vec::new(), Vec::new()), false, false, false));
}
};
// Graphics state tracking
let mut ctm = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0]; // Current Transformation Matrix
let mut text_rendering_mode: i32 = 0; // 0=fill, 1=stroke, 2=fill+stroke, 3=invisible
// Invisible (Tr 3) text was present but suppressed — reported to callers
// so an include_invisible retry is attempted only when it can recover.
let mut skipped_invisible = false;
let mut line_width: f32 = 1.0;
#[derive(Clone)]
struct SavedGraphicsState {
@@ -494,6 +501,14 @@ pub(crate) fn extract_page_text_items(
// For Mixed/template PDFs, include_invisible=true extracts
// the OCR text layer that sits behind scanned images.
if text_rendering_mode == 3 && !include_invisible {
if op
.operands
.first()
.and_then(get_operand_bytes)
.is_some_and(|raw| !raw.is_empty())
{
skipped_invisible = true;
}
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
@@ -565,6 +580,16 @@ pub(crate) fn extract_page_text_items(
if in_text_block && !op.operands.is_empty() {
if let Ok(array) = op.operands[0].as_array() {
let font_info = font_widths.get(&current_font);
// Numeric-only TJ arrays (pure kerning) show no
// text — they must not trigger the invisible retry.
if text_rendering_mode == 3
&& !include_invisible
&& array
.iter()
.any(|el| get_operand_bytes(el).is_some_and(|raw| !raw.is_empty()))
{
skipped_invisible = true;
}
let is_invisible = (text_rendering_mode == 3 && !include_invisible)
|| suppress_glyph_extraction;
// Capture first-glyph position for ActualText
@@ -770,6 +795,16 @@ pub(crate) fn extract_page_text_items(
)
})
});
if text_rendering_mode == 3
&& !include_invisible
&& op
.operands
.first()
.and_then(get_operand_bytes)
.is_some_and(|raw| !raw.is_empty())
{
skipped_invisible = true;
}
if !((text_rendering_mode == 3 && !include_invisible)
|| suppress_glyph_extraction
|| op.operands.is_empty())
@@ -882,6 +917,7 @@ pub(crate) fn extract_page_text_items(
&ctm,
&mut cmap_decisions,
style_cache,
form_budget,
);
items.extend(form_items);
}
@@ -1251,6 +1287,12 @@ pub(crate) fn extract_page_text_items(
}
}
if form_budget.was_truncated() {
log::warn!(
"page {page_num}: Form XObject expansion truncated (invocation or operation budget reached); nested form text may be incomplete"
);
}
// Underline detection reads only painted ink: `re` rects confirmed by
// a paint operator plus filled-subpath rects — never clip-only rects,
// which draw nothing.
@@ -1300,7 +1342,12 @@ pub(crate) fn extract_page_text_items(
let items = super::merge_text_items(items);
let items = super::merge_subscript_items(items);
Ok(((items, rects, lines), has_gid_fonts, coords_rotated))
Ok((
(items, rects, lines),
has_gid_fonts,
coords_rotated,
skipped_invisible,
))
}
/// Counts of text operators with horizontal vs rotated combined matrices.
@@ -1498,13 +1545,14 @@ mod tests {
let (doc, page_id) = simple_doc_with_content(content);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) = extract_page_text_items(
let ((items, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
items
@@ -1734,9 +1782,10 @@ BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
let ((items, rects, lines), _has_gid, _coords_rotated) = result;
let ((items, rects, lines), _has_gid, _coords_rotated, _skipped_invisible) = result;
assert!(items.is_empty());
assert!(rects.is_empty());
assert!(lines.is_empty());
@@ -1817,13 +1866,14 @@ BT 30 700 Tm <41> Tj ET";
doc.trailer.set("Root", Object::Reference(catalog_id));
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) = extract_page_text_items(
let ((items, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
let text = items
@@ -1887,4 +1937,19 @@ BT 30 700 Tm <41> Tj ET";
let output = strip_pdf_comments(input);
assert_eq!(output, b"(x\\\\) Tj \nET\n");
}
#[test]
fn oversized_content_stream_skips_extraction() {
let mut content =
Vec::with_capacity((super::super::content_decode::MAX_PAGE_OPERATIONS + 1) * 2);
for _ in 0..=super::super::content_decode::MAX_PAGE_OPERATIONS {
content.extend_from_slice(b"q\n");
}
content.extend_from_slice(b"BT /F1 12 Tf 72 720 Td (Hello) Tj ET\n");
let items = extract_simple_items(&content);
assert!(
items.is_empty(),
"pages over the operator cap must not be decoded"
);
}
}
+107 -16
View File
@@ -482,7 +482,11 @@ pub(crate) fn parse_cid_w_array(
widths: &mut HashMap<u16, u16>,
) {
let mut i = 0;
let mut assigned = 0usize;
while i < w_array.len() {
if assigned >= crate::tounicode::MAX_CID_W_EXPANSION {
return;
}
let start_cid = match &w_array[i] {
Object::Integer(n) => *n as u16,
Object::Real(n) => *n as u16,
@@ -501,12 +505,14 @@ pub(crate) fn parse_cid_w_array(
Object::Array(arr) => {
// [c [w1 w2 ...]] — consecutive widths starting at c
for (j, w_obj) in arr.iter().enumerate() {
let w = match w_obj {
Object::Integer(n) => *n as u16,
Object::Real(n) => *n as u16,
_ => continue,
};
widths.insert(start_cid + j as u16, w);
if !assign_cid_width(
widths,
start_cid.wrapping_add(j as u16),
w_obj,
&mut assigned,
) {
return;
}
}
i += 1;
}
@@ -514,12 +520,14 @@ pub(crate) fn parse_cid_w_array(
// Could be a reference to an array
if let Ok(Object::Array(arr)) = doc.get_object(*r) {
for (j, w_obj) in arr.iter().enumerate() {
let w = match w_obj {
Object::Integer(n) => *n as u16,
Object::Real(n) => *n as u16,
_ => continue,
};
widths.insert(start_cid + j as u16, w);
if !assign_cid_width(
widths,
start_cid.wrapping_add(j as u16),
w_obj,
&mut assigned,
) {
return;
}
}
i += 1;
} else {
@@ -542,8 +550,8 @@ pub(crate) fn parse_cid_w_array(
continue;
}
};
for cid in start_cid..=end {
widths.insert(cid, w);
if !assign_cid_width_range(widths, start_cid, end, w, &mut assigned) {
return;
}
i += 1;
}
@@ -561,8 +569,8 @@ pub(crate) fn parse_cid_w_array(
continue;
}
};
for cid in start_cid..=end {
widths.insert(cid, w);
if !assign_cid_width_range(widths, start_cid, end, w, &mut assigned) {
return;
}
i += 1;
}
@@ -573,6 +581,45 @@ pub(crate) fn parse_cid_w_array(
}
}
fn assign_cid_width(
widths: &mut HashMap<u16, u16>,
cid: u16,
w_obj: &Object,
assigned: &mut usize,
) -> bool {
let w = match w_obj {
Object::Integer(n) => *n as u16,
Object::Real(n) => *n as u16,
_ => return true,
};
if *assigned >= crate::tounicode::MAX_CID_W_EXPANSION {
return false;
}
widths.insert(cid, w);
*assigned += 1;
true
}
fn assign_cid_width_range(
widths: &mut HashMap<u16, u16>,
start: u16,
end: u16,
w: u16,
assigned: &mut usize,
) -> bool {
if start > end {
return true;
}
for cid in start..=end {
if *assigned >= crate::tounicode::MAX_CID_W_EXPANSION {
return false;
}
widths.insert(cid, w);
*assigned += 1;
}
true
}
/// Compute the width of a string in text space units,
/// given raw bytes and font width info.
/// Returns width in text space units (font_units * units_scale * font_size).
@@ -2313,4 +2360,48 @@ end",
// invalid CMap result — so it must not clear the gid flag.
assert!(gid_flagged(Some("<01> <FFFD>\n<02> <FFFD>")));
}
#[test]
fn parse_cid_w_array_range_and_consecutive() {
use super::parse_cid_w_array;
use lopdf::{Document, Object};
use std::collections::HashMap;
let doc = Document::new();
let mut widths = HashMap::new();
let w = vec![
Object::Integer(10),
Object::Integer(12),
Object::Integer(500),
Object::Integer(20),
Object::Array(vec![Object::Integer(100), Object::Integer(200)]),
];
parse_cid_w_array(&doc, &w, &mut widths);
assert_eq!(widths.get(&10), Some(&500));
assert_eq!(widths.get(&11), Some(&500));
assert_eq!(widths.get(&12), Some(&500));
assert_eq!(widths.get(&20), Some(&100));
assert_eq!(widths.get(&21), Some(&200));
}
#[test]
fn parse_cid_w_array_repeated_full_ranges_stay_bounded() {
use super::parse_cid_w_array;
use crate::tounicode::MAX_CID_W_EXPANSION;
use lopdf::{Document, Object};
use std::collections::HashMap;
let doc = Document::new();
let mut widths = HashMap::new();
let mut w = Vec::new();
for _ in 0..5_000 {
w.push(Object::Integer(0));
w.push(Object::Integer(65535));
w.push(Object::Integer(500));
}
parse_cid_w_array(&doc, &w, &mut widths);
assert!(widths.len() <= MAX_CID_W_EXPANSION);
assert_eq!(widths.get(&0), Some(&500));
assert_eq!(widths.get(&65535), Some(&500));
}
}
+327 -17
View File
@@ -41,15 +41,110 @@ pub(crate) fn detect_columns(
}
debug!("page {}: detect_columns: {} items", page, page_items.len());
// Find page bounds
let x_min = page_items.iter().map(|i| i.x).fold(f32::INFINITY, f32::min);
let x_max = page_items
.iter()
.map(|i| i.x + effective_width(i))
.fold(f32::NEG_INFINITY, f32::max);
// The width of one ordinary page, used three ways below: as the largest
// credible width for a single text run, as the size of empty gap that marks
// content as detached, and as the span past which those checks run at all.
// This is a heuristic, not a format rule: PDF 2.0 sets no page-size limit,
// and since PDF 1.6 `UserUnit` scales a page's physical size independently
// of its coordinates. 14_400 units (200in at the default 1/72in unit) is
// the traditional Acrobat architectural limit, which makes it a reasonable
// "wider than any ordinary page" mark in coordinate space.
const MAX_PAGE_EXTENT: f32 = 14_400.0;
// A detached cluster is only dropped if it also holds a small minority of
// the items, so a genuine two-part layout keeps its full bounds even when
// the halves are far apart.
const MAX_TRIM_FRACTION: f32 = 0.10;
// Position and width of each item, skipping only non-finite geometry.
let finite_span = |i: &&TextItem| -> Option<(f32, f32)> {
let (left, width) = (i.x, effective_width(i));
(left.is_finite() && (left + width).is_finite()).then_some((left, width))
};
let (min_left, max_right, total) = page_items.iter().filter_map(finite_span).fold(
(f32::INFINITY, f32::NEG_INFINITY, 0usize),
|(lo, hi, n), (left, width)| (lo.min(left), hi.max(left + width), n + 1),
);
// No item had usable geometry, so there is no layout to report.
if total == 0 {
return vec![];
}
// Every threshold below (gutter margins, spanning-item width, the XY-cut
// margin) is a fraction of the page width, so a far item can set the scale
// for the whole page and shrink the effective detection window to a
// rounding error — real gutters then fall inside the margin band and a
// genuine multi-column page collapses to one region.
//
// Anything inside one page extent is ordinary, so the common case keeps the
// plain bounds and skips the work below entirely.
let (x_min, x_max) = if max_right - min_left <= MAX_PAGE_EXTENT {
(min_left, max_right)
} else {
// Discarding content needs positive evidence that it is not part of the
// layout, because a count-based rule alone cannot tell a stray from a
// sparse far sidebar. The evidence is geometric: positions are grouped
// into clusters separated by more than a whole page of continuous
// emptiness. Real content, however sparse, does not leave a void that
// large; a malformed coordinate sits alone beyond one.
let mut spans: Vec<(f32, f32)> = page_items.iter().filter_map(finite_span).collect();
spans.sort_by(|a, b| a.0.total_cmp(&b.0));
let mut core: Option<std::ops::Range<usize>> = None;
let mut start = 0usize;
for i in 1..=spans.len() {
if i < spans.len() && spans[i].0 - spans[i - 1].0 <= MAX_PAGE_EXTENT {
continue;
}
if core.as_ref().is_none_or(|best| i - start > best.len()) {
core = Some(start..i);
}
start = i;
}
let mut core = core.unwrap_or(0..spans.len());
// Only drop the detached clusters when they are a small minority, so a
// genuine two-part layout keeps its full bounds.
let dropped = spans.len() - core.len();
if dropped as f32 > spans.len() as f32 * MAX_TRIM_FRACTION {
core = 0..spans.len();
}
let core = &spans[core];
// Positions cannot be inflated by a bogus width, so the spread of the
// content is a sound scale for judging one. A run much wider than the
// page's own content is a malformed width — the test is relative, so a
// genuinely large page keeps its genuinely long runs.
let (lo, widest_left) = (core[0].0, core[core.len() - 1].0);
let max_run_width = (widest_left - lo) + MAX_PAGE_EXTENT;
let hi = core
.iter()
.filter(|&&(_, width)| width <= max_run_width)
.map(|&(left, width)| left + width)
.fold(widest_left, f32::max);
if lo != min_left || hi != max_right {
debug!(
"page {page}: bounds {min_left}..{max_right} exceed one page; \
dropped {dropped}/{} detached item(s), using {lo}..{hi}",
spans.len()
);
}
(lo, hi)
};
// Hard ceiling on the histogram size, independent of the trimming above:
// the bounds are attacker-influenced, so an unclamped
// `page_width / BIN_WIDTH` lets a crafted PDF force an arbitrarily large
// `vec![0u32; num_bins]` allocation. 65_536 bins covers ~128k points at
// BIN_WIDTH 2.0 — roughly 9x the largest legal page — so this never binds
// on a real layout. Kept as a bound that does not depend on the outlier
// heuristic staying correct.
const MAX_BINS: usize = 65_536;
let page_width = x_max - x_min;
if page_width < 200.0 {
if !page_width.is_finite() || page_width < 200.0 {
return vec![ColumnRegion { x_min, x_max }];
}
@@ -57,13 +152,20 @@ pub(crate) fn detect_columns(
return vec![ColumnRegion { x_min, x_max }];
}
// Widen the bins rather than dropping the tail of the page. Clamping the
// count alone would leave anything past MAX_BINS * BIN_WIDTH outside the
// histogram, folded into the last bin, which places gutters at the wrong
// coordinates. Scaling keeps full coverage under the same allocation
// ceiling; only the resolution degrades, and only beyond ~131k points.
let bin_width = BIN_WIDTH.max(page_width / MAX_BINS as f32);
// Build occupancy histogram.
// Exclude items wider than 60% of page width — these are spanning items
// (titles, full-width paragraphs) that would fill the gutter and prevent
// detection of partial-page column layouts (e.g. two-column abstracts on
// a page that also has single-column introduction text).
let wide_threshold = page_width * 0.6;
let num_bins = ((page_width / BIN_WIDTH).ceil() as usize).max(1);
let num_bins = ((page_width / bin_width).ceil() as usize).clamp(1, MAX_BINS);
let mut histogram = vec![0u32; num_bins];
for item in &page_items {
@@ -71,8 +173,8 @@ pub(crate) fn detect_columns(
if w > wide_threshold {
continue;
}
let left = ((item.x - x_min) / BIN_WIDTH).floor() as usize;
let right = (((item.x + w) - x_min) / BIN_WIDTH).ceil() as usize;
let left = ((item.x - x_min) / bin_width).floor() as usize;
let right = (((item.x + w) - x_min) / bin_width).ceil() as usize;
let left = left.min(num_bins);
let right = right.min(num_bins);
for count in histogram.iter_mut().take(right).skip(left) {
@@ -109,12 +211,12 @@ pub(crate) fn detect_columns(
let valleys: Vec<(usize, usize)> = valleys
.into_iter()
.filter(|&(start, end)| {
let width_pts = (end - start) as f32 * BIN_WIDTH;
let width_pts = (end - start) as f32 * bin_width;
if width_pts < MIN_GUTTER_WIDTH {
return false;
}
// Valley center must not be within 5% of page edges
let center_pts = ((start + end) as f32 / 2.0) * BIN_WIDTH;
let center_pts = ((start + end) as f32 / 2.0) * bin_width;
center_pts > margin_threshold && center_pts < (page_width - margin_threshold)
})
.collect();
@@ -132,7 +234,7 @@ pub(crate) fn detect_columns(
&histogram,
num_bins,
x_min,
BIN_WIDTH,
bin_width,
page_width,
margin_threshold,
);
@@ -141,7 +243,7 @@ pub(crate) fn detect_columns(
&rel_valleys,
&page_items,
x_min,
BIN_WIDTH,
bin_width,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
@@ -182,7 +284,7 @@ pub(crate) fn detect_columns(
&valleys,
&page_items,
x_min,
BIN_WIDTH,
bin_width,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
@@ -196,7 +298,7 @@ pub(crate) fn detect_columns(
&valleys,
&page_items,
x_min,
BIN_WIDTH,
bin_width,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
@@ -1827,7 +1929,7 @@ fn split_column_stragglers(lines: Vec<TextLine>) -> (Vec<TextLine>, Vec<TextLine
.unwrap();
let (cs, ce) = segments[core_seg];
let mut core = Vec::with_capacity(ce - cs);
let mut core = Vec::with_capacity(ce.saturating_sub(cs));
let mut stragglers = Vec::new();
for (i, line) in lines.into_iter().enumerate() {
if i >= cs && i < ce {
@@ -2533,6 +2635,214 @@ mod tests {
);
}
#[test]
fn extreme_far_coordinate_does_not_allocate_unboundedly() {
// A crafted PDF can place a text run at an arbitrary coordinate via the
// text matrix. The derived page width must not drive an unbounded
// histogram allocation (previously `page_width / BIN_WIDTH` bins with no
// upper bound would try to reserve terabytes and abort the process).
let mut items = Vec::new();
for i in 0..24 {
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
}
// Item placed 1e12 points away — 5e11 bins if left unclamped.
items.push(make_item(1, 1e12, 700.0, "Z"));
// Must return without aborting; content is preserved as a single region.
let cols = detect_columns(&items, 1, false);
assert!(!cols.is_empty());
}
#[test]
fn non_finite_coordinates_never_leak_into_region_bounds() {
// An inf/NaN coordinate must not escape as a column boundary: callers
// treat these as page/column edges.
for bad_x in [f32::INFINITY, f32::NEG_INFINITY, f32::NAN] {
let mut items = Vec::new();
for i in 0..24 {
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
}
items.push(make_item(1, bad_x, 700.0, "Z"));
for col in detect_columns(&items, 1, false) {
assert!(
col.x_min.is_finite() && col.x_max.is_finite(),
"bad_x {bad_x} leaked bounds {}..{}",
col.x_min,
col.x_max
);
}
}
}
#[test]
fn all_non_finite_coordinates_yield_no_columns() {
let items: Vec<TextItem> = (0..24)
.map(|i| make_item(1, f32::NAN, 700.0 - i as f32 * 5.0, "A"))
.collect();
assert!(detect_columns(&items, 1, false).is_empty());
}
#[test]
fn one_bad_item_does_not_disable_column_detection() {
// A single stray item should not collapse a clean two-column page to
// one region. Every gutter threshold is a fraction of the page width,
// so an untrimmed outlier pushes real gutters inside the rejected
// margin band. A malformed *width* at an ordinary position poisons the
// bounds just as a malformed position does.
for (label, bad_x, bad_width) in [
("nan position", f32::NAN, 0.0),
("inf position", f32::INFINITY, 0.0),
("far position", 50_000.0, 0.0),
("very far position", 1e12, 0.0),
("huge width", 100.0, 1e12),
("inf width", 100.0, f32::INFINITY),
] {
let mut items = Vec::new();
items.extend(fill_zone(1, 30.0, 280.0, 750.0, 50.0));
items.extend(fill_zone(1, 320.0, 570.0, 750.0, 50.0));
let mut bad = make_item(1, bad_x, 400.0, "Z");
bad.width = bad_width;
items.push(bad);
let cols = detect_columns(&items, 1, false);
assert_eq!(
cols.len(),
2,
"{label}: expected 2 columns, got {}",
cols.len()
);
for col in &cols {
assert!(
col.x_max - col.x_min <= MAX_PAGE_EXTENT_FOR_TEST,
"{label}: region {}..{} exceeds one page",
col.x_min,
col.x_max
);
}
}
}
/// Mirrors `MAX_PAGE_EXTENT` in `detect_columns`.
const MAX_PAGE_EXTENT_FOR_TEST: f32 = 14_400.0;
#[test]
fn very_wide_page_keeps_full_histogram_coverage() {
// Beyond MAX_BINS * BIN_WIDTH (~131k points) the bins must widen rather
// than stop covering the page. Three zones: the first gutter is inside
// the old coverage limit, the second is past it. Because the first
// gutter is found, the XY-cut fallback never runs, so a truncated
// histogram silently reports two columns instead of three.
let mut items = Vec::new();
items.extend(fill_zone(1, 0.0, 60_000.0, 750.0, 700.0));
items.extend(fill_zone(1, 70_000.0, 140_000.0, 750.0, 700.0));
items.extend(fill_zone(1, 160_000.0, 200_000.0, 750.0, 700.0));
let cols = detect_columns(&items, 1, false);
assert_eq!(
cols.len(),
3,
"Expected 3 columns across a 200k-wide page, got {}",
cols.len()
);
assert!(
(140_000.0..=160_000.0).contains(&cols[1].x_max),
"second gutter at {}, expected inside the real 140k..160k gap",
cols[1].x_max
);
}
#[test]
fn large_page_with_legitimately_long_runs_is_kept() {
// On a very large page, individual runs can exceed one ordinary page's
// width. They are real content, so they must not be judged malformed:
// the page keeps its columns and its full right edge.
let mut items = Vec::new();
for row in 0..30 {
let y = 750.0 - row as f32 * 14.0;
let mut left = make_item(1, 0.0, y, "Left run");
left.width = 20_000.0;
let mut right = make_item(1, 25_000.0, y, "Right run");
right.width = 20_000.0;
items.extend([left, right]);
}
let cols = detect_columns(&items, 1, false);
assert!(
!cols.is_empty(),
"a page of long-but-valid runs must still report a layout"
);
let right_edge = cols
.iter()
.map(|c| c.x_max)
.fold(f32::NEG_INFINITY, f32::max);
assert!(
right_edge > 44_000.0,
"long runs were treated as malformed: right edge {right_edge}, expected ~45_000"
);
}
#[test]
fn sparse_far_sidebar_on_a_large_page_is_kept() {
// A large-format page with a thin, sparsely-populated sidebar far from
// the main block. The sidebar is a small minority of the items, so an
// item-count rule alone would discard it — but nothing about its
// geometry says it is invalid, so its bounds must survive.
let mut items = Vec::new();
items.extend(fill_zone(1, 0.0, 12_000.0, 750.0, 500.0));
for i in 0..12 {
items.push(make_item(1, 24_000.0, 750.0 - i as f32 * 14.0, "Sidebar"));
}
let cols = detect_columns(&items, 1, false);
let right_edge = cols
.iter()
.map(|c| c.x_max)
.fold(f32::NEG_INFINITY, f32::max);
assert!(
right_edge > 24_000.0,
"sidebar was trimmed away: right edge {right_edge}, expected >24_000"
);
}
#[test]
fn genuinely_wide_layout_keeps_its_true_bounds() {
// A large-format page whose content really is spread beyond one
// ordinary page must not be trimmed to the median cluster: its far
// items are the majority, not strays.
let mut items = Vec::new();
items.extend(fill_zone(1, 100.0, 20_000.0, 750.0, 600.0));
items.extend(fill_zone(1, 22_000.0, 40_000.0, 750.0, 600.0));
let cols = detect_columns(&items, 1, false);
let widest = cols
.iter()
.map(|c| c.x_max)
.fold(f32::NEG_INFINITY, f32::max);
assert!(
widest > 35_000.0,
"wide layout was trimmed: right edge {widest}, expected ~40_000"
);
}
#[test]
fn oversized_but_legal_page_is_not_trimmed() {
// A wide-format page well inside the 14_400pt spec limit must keep its
// real bounds — outlier trimming is only for spans beyond a legal page.
let mut items = Vec::new();
items.extend(fill_zone(1, 100.0, 4_000.0, 750.0, 400.0));
items.extend(fill_zone(1, 4_400.0, 8_000.0, 750.0, 400.0));
let cols = detect_columns(&items, 1, false);
assert_eq!(cols.len(), 2, "Expected 2 columns, got {}", cols.len());
assert!(
cols[1].x_max > 7_000.0,
"right column should keep its true extent, got {}",
cols[1].x_max
);
}
#[test]
fn two_column_regression_guard() {
// Standard 2-column layout with clear gutter at center
+311 -5
View File
@@ -2,11 +2,51 @@
use crate::types::{ItemType, TextItem};
use lopdf::{Document, Object, ObjectId};
use std::collections::HashMap;
use std::collections::{HashMap, HashSet};
use super::fonts::{resolve_array, resolve_dict};
use super::get_number;
/// Upper bound on the number of form-field nodes visited during a single
/// `extract_form_fields` pass. A crafted PDF can chain thousands of distinct
/// `/Kids` fields to blow the stack even without an outright reference cycle,
/// so we cap total traversal work in addition to detecting cycles.
const MAX_FORM_FIELD_NODES: usize = 100_000;
/// Upper bound on `/Kids` recursion depth. Real AcroForm hierarchies are only
/// a few levels deep (fields → child fields → widgets); a crafted PDF can chain
/// tens of thousands of distinct fields into a linear `/Kids` list that would
/// overflow the stack via depth-first recursion long before the node budget is
/// reached. This depth cap bounds the stack independently of total node count.
const MAX_FORM_FIELD_DEPTH: usize = 100;
/// Traversal budget for the AcroForm field walk. Bounds both the number of
/// distinct nodes visited *and* the total number of `/Fields`/`/Kids` entries
/// examined.
///
/// Counting `visited` alone is not enough: invalid entries (non-references) and
/// duplicate references never grow `visited`, so an oversized array full of them
/// would iterate to completion no matter how large. Charging every examined
/// entry against the same budget makes it a real cap on traversal work.
pub(crate) struct FieldWalkBudget {
visited: HashSet<ObjectId>,
examined: usize,
}
impl FieldWalkBudget {
fn new() -> Self {
Self {
visited: HashSet::new(),
examined: 0,
}
}
/// True once the budget is spent; callers must stop iterating and recursing.
fn exhausted(&self) -> bool {
self.visited.len() >= MAX_FORM_FIELD_NODES || self.examined >= MAX_FORM_FIELD_NODES
}
}
pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> Vec<TextItem> {
let mut links = Vec::new();
@@ -146,9 +186,12 @@ pub(crate) fn extract_form_fields(
Err(_) => return items,
};
// Borrow the array rather than cloning it: a crafted `/Fields` can be huge,
// and cloning would pay an O(n) allocation/copy before the budget check
// below can stop the work.
let fields = match acroform.get(b"Fields") {
Ok(obj) => match resolve_array(doc, obj) {
Some(arr) => arr.clone(),
Some(arr) => arr,
None => return items,
},
Err(_) => return items,
@@ -158,7 +201,19 @@ pub(crate) fn extract_form_fields(
}
let annotation_pages = annotation_page_map(doc, page_map);
for field_obj in &fields {
// Bound the walk so a crafted PDF cannot send us into unbounded recursion
// via a `/Kids` cycle, a deep chain, or an oversized array of invalid or
// duplicate entries.
let mut budget = FieldWalkBudget::new();
for field_obj in fields {
// Stop once the budget is spent so a `/Fields` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op. Charge
// every entry (including invalid ones) against the budget.
if budget.exhausted() {
break;
}
budget.examined += 1;
if let Ok(field_ref) = field_obj.as_reference() {
walk_form_fields(
doc,
@@ -168,6 +223,8 @@ pub(crate) fn extract_form_fields(
page_map,
&annotation_pages,
&mut items,
&mut budget,
0,
);
}
}
@@ -202,6 +259,7 @@ fn annotation_page_map(
}
/// Recursively walk the form field tree, extracting leaf field values.
#[allow(clippy::too_many_arguments)]
pub(crate) fn walk_form_fields(
doc: &Document,
field_id: ObjectId,
@@ -210,7 +268,22 @@ pub(crate) fn walk_form_fields(
page_map: &HashMap<ObjectId, u32>,
annotation_pages: &HashMap<ObjectId, u32>,
items: &mut Vec<TextItem>,
budget: &mut FieldWalkBudget,
depth: usize,
) {
// Guard against `/Kids` cycles and pathologically large field trees.
// Exceeding the depth cap means the chain is too deep to be a legitimate
// form (and would overflow the stack); an exhausted budget means the tree is
// too large. Both checks run *before* inserting so the visited set can never
// grow past the budget.
if depth > MAX_FORM_FIELD_DEPTH || budget.exhausted() {
return;
}
// Revisiting an object ID means we hit a `/Kids` cycle.
if !budget.visited.insert(field_id) {
return;
}
let field_dict = match doc.get_dictionary(field_id) {
Ok(d) => d,
Err(_) => return,
@@ -241,9 +314,19 @@ pub(crate) fn walk_form_fields(
// Check for /Kids — if present, recurse into children
if let Ok(kids_obj) = field_dict.get(b"Kids") {
// Iterate the borrowed array directly — cloning a crafted, oversized
// `/Kids` would allocate and copy every entry before the budget check
// below could stop the work.
if let Some(kids) = resolve_array(doc, kids_obj) {
let kids = kids.clone();
for kid in &kids {
for kid in kids {
// Stop once the budget is spent so a `/Kids` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op.
// Charge every entry (including invalid/duplicate ones) against
// the budget so this is a true traversal-work cap.
if budget.exhausted() {
break;
}
budget.examined += 1;
if let Ok(kid_ref) = kid.as_reference() {
walk_form_fields(
doc,
@@ -253,6 +336,8 @@ pub(crate) fn walk_form_fields(
page_map,
annotation_pages,
items,
budget,
depth + 1,
);
}
}
@@ -411,4 +496,225 @@ mod tests {
assert_eq!(items[0].page, 2);
assert_eq!(items[0].text, "customer: Alice");
}
#[test]
fn kids_self_cycle_does_not_overflow_stack() {
// A crafted AcroForm field that lists itself in `/Kids` must not send
// the traversal into unbounded recursion.
let mut doc = Document::new();
let field_id = doc.new_object_id();
doc.set_object(
field_id,
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("loop"),
"Kids" => vec![Object::Reference(field_id)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
// Completes (rather than overflowing the stack) and yields no items.
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn kids_mutual_cycle_terminates() {
// Two fields that reference each other via `/Kids` form a cycle that
// must also terminate.
let mut doc = Document::new();
let field_a = doc.new_object_id();
let field_b = doc.new_object_id();
doc.set_object(
field_a,
dictionary! {
"T" => Object::string_literal("a"),
"Kids" => vec![Object::Reference(field_b)],
},
);
doc.set_object(
field_b,
dictionary! {
"T" => Object::string_literal("b"),
"Kids" => vec![Object::Reference(field_a)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_a)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn deep_acyclic_kids_chain_does_not_overflow_stack() {
// A long chain of *distinct* fields (no cycle) must also terminate:
// the visited set alone would still recurse to the chain length, so
// the depth cap is what prevents a stack overflow here.
let mut doc = Document::new();
let n = MAX_FORM_FIELD_DEPTH * 500;
let ids: Vec<ObjectId> = (0..=n).map(|_| doc.new_object_id()).collect();
for i in 0..n {
doc.set_object(
ids[i],
dictionary! {
"FT" => "Tx",
"Kids" => vec![Object::Reference(ids[i + 1])],
},
);
}
// Leaf carries a value; it sits far below the depth cap so it is never
// reached, proving traversal stops early rather than crashing.
doc.set_object(
ids[n],
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("leaf"),
"V" => Object::string_literal("x"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(ids[0])],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn wide_tree_traversal_stops_at_node_budget() {
// A single field with a `/Kids` array wider than the node budget must
// stop traversal at the cap rather than growing `visited` (and the work)
// without bound. Each processed leaf emits one item, so the item count
// is bounded by the budget and reaches right up to it (a couple of
// slots go to the root and the boundary node charged against the cap).
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let kids: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
// Extraction stops at the budget: bounded above by the cap, and it gets
// right up to it (allowing a small delta for the root/boundary nodes
// charged against the budget).
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn wide_top_level_fields_stop_at_node_budget() {
// A top-level `/Fields` array wider than the budget must also stop at
// the cap: the item count is bounded by the budget and reaches right up
// to it.
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let fields: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => fields,
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn duplicate_and_invalid_kids_entries_stop_at_budget() {
// Duplicate references and non-reference junk never grow `visited`, so
// without charging examined entries against the budget an oversized
// array of them would iterate to completion. The walk must still
// terminate and extract the single real leaf exactly once.
let mut doc = Document::new();
let leaf_id = doc.new_object_id();
doc.set_object(
leaf_id,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
// A `/Kids` array far wider than the budget: half duplicate references
// to the same leaf, half invalid (null) entries.
let mut kids: Vec<Object> = Vec::new();
for i in 0..(MAX_FORM_FIELD_NODES * 2) {
if i % 2 == 0 {
kids.push(Object::Reference(leaf_id));
} else {
kids.push(Object::Null);
}
}
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert_eq!(items.len(), 1);
}
}
+253 -14
View File
@@ -3,6 +3,7 @@
//! This module extracts text with position information for structure detection.
mod base14;
mod content_decode;
pub(crate) mod content_stream;
mod fonts;
mod layout;
@@ -37,6 +38,7 @@ pub(crate) use layout::group_prefiltered_items_into_lines_with_thresholds_and_re
pub(crate) use layout::is_newspaper_layout;
pub(crate) use layout::ColumnRegion;
pub use layout::{group_into_lines, group_into_lines_preserving_all_text};
pub(crate) use xobjects::FormWalkBudget;
// ---------------------------------------------------------------------------
// Public API
@@ -282,18 +284,22 @@ fn extract_positioned_text_impl(
font_cmaps,
include_invisible,
&mut style_cache,
&mut FormWalkBudget::new(),
);
let ((mut items, mut rects, mut lines), has_gid_fonts, coords_rotated) = match page_result {
Ok(extraction) => extraction,
Err(error) if required_pages.is_some_and(|required| !required.contains(page_num)) => {
debug!(
"page {}: skipping context-only extraction error: {}",
page_num, error
);
continue;
}
Err(error) => return Err(error),
};
let ((mut items, mut rects, mut lines), has_gid_fonts, coords_rotated, _skipped_invisible) =
match page_result {
Ok(extraction) => extraction,
Err(error)
if required_pages.is_some_and(|required| !required.contains(page_num)) =>
{
debug!(
"page {}: skipping context-only extraction error: {}",
page_num, error
);
continue;
}
Err(error) => return Err(error),
};
// Clip to the visible page box: single-page extracts and imposed
// spreads keep neighboring pages' content in the stream, positioned
// outside the CropBox. Extracting it interleaves invisible text into
@@ -884,6 +890,92 @@ fn tracked_run_space_floor(group: &[&TextItem], start: usize) -> Option<(usize,
Some((end, floor * fs))
}
/// Fractional font-size band within which `merge_text_items` treats two runs as
/// the same size. Shared with `is_small_caps_continuation`, which exists only to
/// rescue junctions this band would otherwise break.
const MERGE_FONT_SIZE_BAND: f32 = 0.20;
/// Detect a small-caps continuation: typesetters render small caps as a
/// full-size capital immediately followed by shrunken capitals in the same
/// font (`(R) Tj` at 9.98pt, then `(OLANDO) Tj` at 6.74pt). Those runs are one
/// word, but the font-size band in `merge_text_items` would split them,
/// leaving "R" and "OLANDO" as separate items — which then read as separate
/// table columns, since column boundaries cluster on item start positions.
///
/// Gated tightly so it cannot absorb the other reasons a smaller run follows a
/// larger one:
/// - runs the size band already accepts — excluded by requiring the junction
/// to *cross* the band, so within-band pairs keep the normal word-spacing
/// logic instead of having their space suppressed
/// - superscripts / footnote markers — excluded by requiring an uppercase
/// *letter* on both sides, so digits never qualify
/// - drop caps — excluded because the body text that follows is mixed case
/// - adjacent table cells or separate words — excluded by requiring the runs
/// to be visually contiguous (essentially no gap)
fn is_small_caps_continuation(
text_so_far: &str,
first: &TextItem,
next: &TextItem,
gap: f32,
) -> bool {
// Must shrink. Real small caps sit near 0.7-0.8 of the full cap height;
// anything smaller is a superscript or a different run entirely.
if first.font_size <= 0.0 || next.font_size >= first.font_size {
return false;
}
// Only rescue junctions the size band would have broken. Within-band pairs
// merge on their own, and suppressing their space would swallow real word
// gaps between two similarly-sized uppercase words.
if (next.font_size - first.font_size).abs() <= first.font_size * MERGE_FONT_SIZE_BAND {
return false;
}
if next.font_size / first.font_size < 0.55 {
return false;
}
// Visually contiguous: the capital and its small caps touch. A real word
// space or a column gap disqualifies.
if !(-first.font_size * 0.2..=first.font_size * 0.15).contains(&gap) {
return false;
}
// The continuation must be all-uppercase letters (digits and lowercase
// both disqualify), and must contain at least one letter.
let mut saw_letter = false;
for ch in next.text.chars() {
if ch.is_alphabetic() {
saw_letter = true;
if !ch.is_uppercase() {
return false;
}
} else if ch.is_numeric() {
return false;
}
}
if !saw_letter {
return false;
}
// What we are continuing must itself end in a capital. Check the actual
// trailing character rather than skipping back to the nearest letter: after
// "ANGELA M. MAZZARELLI1" the run to continue is the footnote marker, not
// the "I" before it.
let trimmed = text_so_far.trim_end();
if trimmed.chars().last().is_some_and(|c| c.is_numeric()) {
// One legitimate exception: an ordinal suffix set as a smaller run,
// e.g. "JULY 4" + "TH". Only the four English suffixes qualify —
// anything else after a digit is a footnote marker or numeric suffix.
return matches!(trimmed_suffix(next), "TH" | "ST" | "ND" | "RD");
}
trimmed
.chars()
.rev()
.find(|c| c.is_alphabetic())
.is_some_and(|c| c.is_uppercase())
}
/// The continuation run's text, trimmed — used to spot ordinal suffixes.
fn trimmed_suffix(next: &TextItem) -> &str {
next.text.trim()
}
pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
if items.is_empty() {
return items;
@@ -942,8 +1034,16 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
let mut j = i + 1;
while j < group.len() {
let next = group[j];
// Must be similar font size (within 20%)
if (next.font_size - first.font_size).abs() > first.font_size * 0.20 {
// A small-caps junction is mid-word: it both survives the
// font-size band below and must never take a space.
let small_caps_join =
is_small_caps_continuation(&text, first, next, next.x - end_x);
// Must be similar font size, except for genuine small-caps
// runs, where the shrunken capitals are the same word as the
// full-size initial (see helper).
if (next.font_size - first.font_size).abs() > first.font_size * MERGE_FONT_SIZE_BAND
&& !small_caps_join
{
break;
}
// Never merge across style boundaries: the merged item
@@ -997,7 +1097,7 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
Some((run_end, floor)) if j <= run_end => floor,
_ => threshold,
};
if needs_bullet_space || gap > effective_threshold {
if !small_caps_join && (needs_bullet_space || gap > effective_threshold) {
text.push(' ');
}
text.push_str(&next.text);
@@ -3001,6 +3101,145 @@ mod tests {
}
}
/// Small caps as typesetters emit them: a full-size capital at 9.98pt
/// immediately followed by shrunken capitals at 6.74pt, touching.
/// Modelled on `199AD3d.pdf` p.5 ("ROLANDO T. ACOSTA, P.J.").
#[test]
fn small_caps_run_merges_into_one_word() {
let items = vec![
make_item_fs("R", 144.36, 581.84, 7.20, 9.98),
make_item_fs("OLANDO", 151.56, 581.84, 30.56, 6.74),
make_item_fs("T. A", 185.45, 581.84, 17.58, 9.98),
make_item_fs("COSTA", 203.94, 581.84, 23.15, 6.74),
make_item_fs(", P.J.", 227.09, 581.84, 22.56, 9.98),
];
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1, "got {:?}", merged);
assert_eq!(merged[0].text, "ROLANDO T. ACOSTA, P.J.");
}
/// The full two-column row: both names must merge independently and the
/// 72pt column gap between them must survive as an item boundary.
#[test]
fn small_caps_merge_does_not_swallow_a_second_column() {
let items = vec![
// Column 1: "ROLANDO T. ACOSTA, P.J." ending at x=249.65
make_item_fs("R", 144.36, 581.84, 7.20, 9.98),
make_item_fs("OLANDO", 151.56, 581.84, 30.56, 6.74),
make_item_fs("T. A", 185.45, 581.84, 17.58, 9.98),
make_item_fs("COSTA", 203.94, 581.84, 23.15, 6.74),
make_item_fs(", P.J.", 227.09, 581.84, 22.56, 9.98),
// Column 2 starts at x=321.96 — a 72pt gap.
make_item_fs("A", 321.96, 581.84, 7.20, 9.98),
make_item_fs("NIL", 329.17, 581.84, 12.72, 6.74),
make_item_fs("C. S", 345.04, 581.84, 19.59, 9.98),
make_item_fs("INGH", 364.62, 581.84, 19.08, 6.74),
];
let merged = merge_text_items(items);
let texts: Vec<&str> = merged.iter().map(|i| i.text.as_str()).collect();
assert_eq!(
texts,
vec!["ROLANDO T. ACOSTA, P.J.", "ANIL C. SINGH"],
"column gap should keep the two names apart"
);
}
#[test]
fn small_caps_merge_keeps_word_space_between_same_size_capitals() {
// Two uppercase words at sizes the merge band already accepts (9.98 and
// 9.0, a 10% drop) separated by a real word gap. The small-caps path
// must not claim this junction and swallow the space.
let items = vec![
make_item_fs("SEE", 100.0, 500.0, 18.0, 9.98),
make_item_fs("ALSO", 119.2, 500.0, 24.0, 9.0),
];
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1, "got {:?}", merged);
assert_eq!(merged[0].text, "SEE ALSO");
}
#[test]
fn trailing_digit_is_not_a_capital_awaiting_small_caps() {
// "...MAZZARELLI1" ends in a footnote marker; the backward search for an
// uppercase letter must not skip the digit and glue the next run.
assert!(!is_small_caps_continuation(
"ANGELA M. MAZZARELLI1",
&make_item_fs("ANGELA", 100.0, 500.0, 40.0, 9.98),
&make_item_fs("SHULMAN", 140.0, 500.0, 30.0, 6.74),
0.0,
));
}
#[test]
fn ordinal_suffix_after_a_digit_still_merges() {
// "TUESDAY, JULY 4" + "TH" is one word in the source; the digit guard
// must not block the four English ordinal suffixes.
for suffix in ["TH", "ST", "ND", "RD"] {
assert!(
is_small_caps_continuation(
"TUESDAY, JULY 4",
&make_item_fs("JULY", 100.0, 500.0, 30.0, 12.0),
&make_item_fs(suffix, 130.0, 500.0, 8.0, 8.0),
0.0,
),
"{suffix} should merge after a digit"
);
}
}
#[test]
fn superscript_footnote_marker_is_not_a_small_caps_continuation() {
// A digit must never qualify — otherwise footnote markers get glued on
// without the superscript handling.
assert!(!is_small_caps_continuation(
"MAZZARELLI",
&make_item_fs("MAZZARELLI", 100.0, 500.0, 50.0, 9.98),
&make_item_fs("1", 150.0, 503.0, 3.0, 6.74),
0.0,
));
}
#[test]
fn drop_cap_is_not_a_small_caps_continuation() {
// Mixed-case body text after a large initial is a drop cap, not small
// caps.
assert!(!is_small_caps_continuation(
"T",
&make_item_fs("T", 100.0, 500.0, 20.0, 30.0),
&make_item_fs("he court held", 120.0, 500.0, 60.0, 10.0),
0.0,
));
}
#[test]
fn separate_word_is_not_a_small_caps_continuation() {
// A real word space disqualifies even when both runs are uppercase.
let first = make_item_fs("SEE", 100.0, 500.0, 20.0, 9.98);
let next = make_item_fs("ALSO", 128.0, 500.0, 25.0, 6.74);
assert!(!is_small_caps_continuation("SEE", &first, &next, 8.0));
}
#[test]
fn lowercase_continuation_is_not_small_caps() {
assert!(!is_small_caps_continuation(
"SMALL",
&make_item_fs("SMALL", 100.0, 500.0, 30.0, 9.98),
&make_item_fs("caps", 130.0, 500.0, 20.0, 6.74),
0.0,
));
}
#[test]
fn too_small_a_ratio_is_not_small_caps() {
// 0.4 ratio is a superscript/sub-run, outside the small-caps band.
assert!(!is_small_caps_continuation(
"A",
&make_item_fs("A", 100.0, 500.0, 7.0, 10.0),
&make_item_fs("BC", 107.0, 500.0, 8.0, 4.0),
0.0,
));
}
#[test]
fn test_merge_subscript_items_chemical_formula() {
// NH₃: "NH" at fs=8 followed by subscript "3" at fs=4.7
+600 -21
View File
@@ -16,6 +16,76 @@ use super::{get_number, image_bbox_from_ctm, multiply_matrices};
const MAX_FORM_XOBJECT_DEPTH: u8 = 5;
/// Upper bound on Form XObject invocations during a single page extraction.
/// Depth alone is not enough: an acyclic DAG where each form invokes the next
/// N times expands to N^depth work before the depth cap is reached.
const MAX_FORM_XOBJECT_INVOCATIONS: usize = 10_000;
/// Upper bound on content-stream operations walked across all Form XObject
/// expansions for a page. Nested forms are decoded independently of the
/// page-level operation cap, so this keeps total form work in the same
/// ballpark as that page cap.
const MAX_FORM_XOBJECT_OPERATIONS: usize = 1_000_000;
/// Shared budget for Form XObject expansion on a page. Bounds both nested DAG
/// expansion and repeated sibling `/Do` invocations of the same form.
pub(crate) struct FormWalkBudget {
invocations: usize,
operations: usize,
max_invocations: usize,
max_operations: usize,
truncated: bool,
}
impl FormWalkBudget {
pub(crate) fn new() -> Self {
Self::with_limits(MAX_FORM_XOBJECT_INVOCATIONS, MAX_FORM_XOBJECT_OPERATIONS)
}
fn with_limits(max_invocations: usize, max_operations: usize) -> Self {
Self {
invocations: 0,
operations: 0,
max_invocations,
max_operations,
truncated: false,
}
}
fn exhausted(&mut self) -> bool {
if self.invocations >= self.max_invocations || self.operations >= self.max_operations {
self.truncated = true;
true
} else {
false
}
}
fn charge_invocation(&mut self) -> bool {
if self.exhausted() {
return false;
}
self.invocations += 1;
true
}
/// Charge one walked content-stream operator. Independent of the
/// invocation cap so a form that was already admitted can finish its
/// stream (up to the operation cap).
fn charge_operation(&mut self) -> bool {
if self.operations >= self.max_operations {
self.truncated = true;
return false;
}
self.operations += 1;
true
}
pub(crate) fn was_truncated(&self) -> bool {
self.truncated
}
}
pub(crate) enum XObjectType {
Image,
Form(ObjectId),
@@ -109,6 +179,7 @@ fn collect_xobjects_from_dict(
}
/// Extract text items from a Form XObject
#[allow(clippy::too_many_arguments)]
pub(crate) fn extract_form_xobject_text(
doc: &Document,
form_id: ObjectId,
@@ -117,6 +188,7 @@ pub(crate) fn extract_form_xobject_text(
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
budget: &mut FormWalkBudget,
) -> Vec<TextItem> {
extract_form_xobject_text_inner(
doc,
@@ -127,6 +199,7 @@ pub(crate) fn extract_form_xobject_text(
cmap_decisions,
style_cache,
0,
budget,
)
}
@@ -140,11 +213,14 @@ fn extract_form_xobject_text_inner(
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
depth: u8,
budget: &mut FormWalkBudget,
) -> Vec<TextItem> {
use lopdf::content::Content;
let mut items = Vec::new();
if !budget.charge_invocation() {
return items;
}
// Get the Form XObject stream
let Ok(Object::Stream(stream)) = doc.get_object(form_id) else {
return items;
@@ -156,8 +232,12 @@ fn extract_form_xobject_text_inner(
Err(_) => stream.content.clone(),
};
// Decode the content stream
let Ok(content) = Content::decode(&content_data) else {
// Decode the content stream. Cap before lopdf materializes the operator
// vector — the walk budget cannot help if decode itself allocates first.
let Ok(Some(content)) = super::content_decode::decode_content_bounded(
&content_data,
super::content_decode::MAX_PAGE_OPERATIONS,
) else {
return items;
};
@@ -246,19 +326,55 @@ fn extract_form_xobject_text_inner(
let mut current_font = String::new();
let mut current_font_size: f32 = 12.0;
let mut text_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
// Text line matrix (TLM) — Td/TD/T* move relative to the start of the
// current line, not to the position left by the last show operator.
let mut line_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut text_leading: f32 = 0.0; // TL parameter (text-space units)
let mut char_spacing: f32 = 0.0; // Tc parameter
let mut word_spacing: f32 = 0.0; // Tw parameter
let mut in_text_block = false;
let mut fill_is_white = false;
let mut ctm = base_ctm;
let mut ctm_stack: Vec<[f32; 6]> = Vec::new();
// Text state (Tc/Tw/TL/Tf) and the fill colour are part of the graphics
// state and must be saved/restored by q/Q alongside the CTM.
#[derive(Clone)]
struct GraphicsState {
ctm: [f32; 6],
char_spacing: f32,
word_spacing: f32,
text_leading: f32,
current_font: String,
current_font_size: f32,
fill_is_white: bool,
}
let mut ctm_stack: Vec<GraphicsState> = Vec::new();
for op in &content.operations {
if !budget.charge_operation() {
break;
}
match op.operator.as_str() {
"q" => {
ctm_stack.push(ctm);
ctm_stack.push(GraphicsState {
ctm,
char_spacing,
word_spacing,
text_leading,
current_font: current_font.clone(),
current_font_size,
fill_is_white,
});
}
"Q" => {
if let Some(saved) = ctm_stack.pop() {
ctm = saved;
ctm = saved.ctm;
char_spacing = saved.char_spacing;
word_spacing = saved.word_spacing;
text_leading = saved.text_leading;
current_font = saved.current_font;
current_font_size = saved.current_font_size;
fill_is_white = saved.fill_is_white;
}
}
"cm" => {
@@ -276,7 +392,7 @@ fn extract_form_xobject_text_inner(
let xobj_name = String::from_utf8_lossy(name).to_string();
match form_xobjects.get(&xobj_name) {
Some(XObjectType::Form(nested_id)) => {
if depth < MAX_FORM_XOBJECT_DEPTH {
if depth < MAX_FORM_XOBJECT_DEPTH && !budget.exhausted() {
let nested_items = extract_form_xobject_text_inner(
doc,
*nested_id,
@@ -286,6 +402,7 @@ fn extract_form_xobject_text_inner(
cmap_decisions,
style_cache,
depth + 1,
budget,
);
items.extend(nested_items);
}
@@ -321,6 +438,7 @@ fn extract_form_xobject_text_inner(
"BT" => {
in_text_block = true;
text_matrix = [1.0, 0.0, 0.0, 1.0, 0.0, 0.0];
line_matrix = text_matrix;
}
"ET" => {
in_text_block = false;
@@ -333,12 +451,33 @@ fn extract_form_xobject_text_inner(
current_font_size = get_number(&op.operands[1]).unwrap_or(12.0);
}
}
"TL" => {
// Set text leading (used by T*, ', and ")
if let Some(tl) = op.operands.first().and_then(get_number) {
text_leading = tl;
}
}
"Tc" => {
if let Some(tc) = op.operands.first().and_then(get_number) {
char_spacing = tc;
}
}
"Tw" => {
if let Some(tw) = op.operands.first().and_then(get_number) {
word_spacing = tw;
}
}
"Td" | "TD" => {
// Move text position: TLM = T(tx,ty) x TLM; Tm = TLM
if op.operands.len() >= 2 {
let tx = get_number(&op.operands[0]).unwrap_or(0.0);
let ty = get_number(&op.operands[1]).unwrap_or(0.0);
text_matrix[4] += tx * text_matrix[0] + ty * text_matrix[2];
text_matrix[5] += tx * text_matrix[1] + ty * text_matrix[3];
line_matrix[4] += tx * line_matrix[0] + ty * line_matrix[2];
line_matrix[5] += tx * line_matrix[1] + ty * line_matrix[3];
text_matrix = line_matrix;
if op.operator == "TD" {
text_leading = -ty;
}
}
}
"Tm" => {
@@ -347,8 +486,20 @@ fn extract_form_xobject_text_inner(
text_matrix[i] =
get_number(operand).unwrap_or(if i == 0 || i == 3 { 1.0 } else { 0.0 });
}
line_matrix = text_matrix;
}
}
"T*" => {
// Move to start of next line: equivalent to `0 -TL Td`
let tl = if text_leading != 0.0 {
text_leading
} else {
current_font_size * 1.2
};
line_matrix[4] += (-tl) * line_matrix[2];
line_matrix[5] += (-tl) * line_matrix[3];
text_matrix = line_matrix;
}
"g" => {
if let Some(gray) = op.operands.first().and_then(get_number) {
fill_is_white = gray > 0.95;
@@ -384,17 +535,33 @@ fn extract_form_xobject_text_inner(
_ => fill_is_white = false,
}
}
"Tj" => {
if in_text_block && !op.operands.is_empty() {
"Tj" | "'" | "\"" => {
// `'` = move to next line then show; `"` = set word/char spacing,
// move to next line, then show (string is the last operand).
if op.operator != "Tj" {
if op.operator == "\"" && op.operands.len() >= 3 {
word_spacing = get_number(&op.operands[0]).unwrap_or(word_spacing);
char_spacing = get_number(&op.operands[1]).unwrap_or(char_spacing);
}
let tl = if text_leading != 0.0 {
text_leading
} else {
current_font_size * 1.2
};
line_matrix[4] += (-tl) * line_matrix[2];
line_matrix[5] += (-tl) * line_matrix[3];
text_matrix = line_matrix;
}
if let (true, Some(show_operand)) = (in_text_block, op.operands.last()) {
if fill_is_white {
if let Some(font_info) = font_widths.get(&current_font) {
if let Some(raw_bytes) = get_operand_bytes(&op.operands[0]) {
if let Some(raw_bytes) = get_operand_bytes(show_operand) {
let w_ts = compute_string_width_ts(
raw_bytes,
font_info,
current_font_size,
0.0,
0.0,
char_spacing,
word_spacing,
);
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
@@ -403,7 +570,7 @@ fn extract_form_xobject_text_inner(
continue;
}
if let Some(text) = extract_text_from_operand(
&op.operands[0],
show_operand,
&current_font,
font_base_names.get(&current_font).map(|s| s.as_str()),
font_cmaps,
@@ -419,13 +586,13 @@ fn extract_form_xobject_text_inner(
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
let width = if let Some(font_info) = font_widths.get(&current_font) {
if let Some(raw_bytes) = get_operand_bytes(&op.operands[0]) {
if let Some(raw_bytes) = get_operand_bytes(show_operand) {
let w_ts = compute_string_width_ts(
raw_bytes,
font_info,
current_font_size,
0.0,
0.0,
char_spacing,
word_spacing,
);
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
@@ -548,8 +715,8 @@ fn extract_form_xobject_text_inner(
raw_bytes,
fi,
current_font_size,
0.0,
0.0,
char_spacing,
word_spacing,
);
}
}
@@ -683,3 +850,415 @@ pub(crate) fn get_form_fonts<'a>(
fonts
}
#[cfg(test)]
mod tests {
use super::*;
use crate::extractor::content_stream::extract_page_text_items;
use lopdf::{dictionary, Dictionary, Stream};
/// Build an acyclic Form XObject DAG: `levels` form objects, each non-leaf
/// invoking the next form `branches` times. The leaf draws a single `(X)`.
/// Returns `(doc, root_form_id)`.
fn form_dag(branches: usize, levels: usize) -> (Document, ObjectId) {
assert!(levels >= 2);
let mut doc = Document::new();
let font_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
});
let ids: Vec<ObjectId> = (0..levels).map(|_| doc.new_object_id()).collect();
for level in 0..levels {
let stream = if level + 1 == levels {
Stream::new(
dictionary! {
"Type" => "XObject",
"Subtype" => "Form",
"BBox" => vec![0.into(), 0.into(), 100.into(), 100.into()],
"Resources" => dictionary! {
"Font" => dictionary! {
"F1" => Object::Reference(font_id),
},
},
},
b"BT /F1 10 Tf 10 10 Td (X) Tj ET\n".to_vec(),
)
} else {
let next_name = format!("Fm{}", level + 1);
let content = format!("/{next_name} Do\n").repeat(branches);
let mut xobjects = Dictionary::new();
xobjects.set(next_name, Object::Reference(ids[level + 1]));
let mut resources = Dictionary::new();
resources.set("XObject", Object::Dictionary(xobjects));
let mut dict = dictionary! {
"Type" => "XObject",
"Subtype" => "Form",
"BBox" => vec![0.into(), 0.into(), 100.into(), 100.into()],
};
dict.set("Resources", Object::Dictionary(resources));
Stream::new(dict, content.into_bytes())
};
doc.set_object(ids[level], Object::Stream(stream));
}
(doc, ids[0])
}
fn page_invoking_form(mut doc: Document, form_id: ObjectId) -> (Document, ObjectId) {
let content_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
b"/Fm0 Do\n".to_vec(),
)));
let page_id = doc.add_object(dictionary! {
"Type" => "Page",
"Contents" => Object::Reference(content_id),
"Resources" => dictionary! {
"XObject" => dictionary! {
"Fm0" => Object::Reference(form_id),
},
},
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
});
let pages_id = doc.add_object(dictionary! {
"Type" => "Pages",
"Count" => Object::Integer(1),
"Kids" => vec![Object::Reference(page_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => Object::Reference(pages_id),
});
doc.trailer.set("Root", Object::Reference(catalog_id));
(doc, page_id)
}
fn extract_form(
doc: &Document,
form_id: ObjectId,
budget: &mut FormWalkBudget,
) -> Vec<TextItem> {
extract_form_xobject_text(
doc,
form_id,
1,
&FontCMaps::from_doc(doc),
&[1.0, 0.0, 0.0, 1.0, 0.0, 0.0],
&mut CMapDecisionCache::new(),
&mut FontStyleCache::new(),
budget,
)
}
#[test]
fn nested_form_still_extracts_leaf_text() {
let (doc, root) = form_dag(1, 3);
let items = extract_form(&doc, root, &mut FormWalkBudget::new());
assert_eq!(items.len(), 1);
assert_eq!(items[0].text, "X");
}
#[test]
fn acyclic_form_dag_within_budget_keeps_all_leaves() {
// 4 sibling invocations across 4 nested levels → 4^4 leaf drawings.
// Default budgets are far above 256, so legitimate nesting is intact.
let (doc, root) = form_dag(4, 5);
let items = extract_form(&doc, root, &mut FormWalkBudget::new());
assert_eq!(items.len(), 4usize.pow(4));
assert!(items.iter().all(|item| item.text == "X"));
}
#[test]
fn acyclic_form_dag_stops_at_invocation_budget() {
// Same DAG as above would draw 256 leaves; a tiny invocation cap must
// stop expansion rather than walking the full tree.
let (doc, root) = form_dag(4, 5);
let mut budget = FormWalkBudget::with_limits(20, MAX_FORM_XOBJECT_OPERATIONS);
let items = extract_form(&doc, root, &mut budget);
assert!(
items.len() < 4usize.pow(4),
"invocation budget must truncate DAG expansion; got {} items",
items.len()
);
assert!(budget.was_truncated());
}
#[test]
fn form_operations_stop_at_budget() {
let mut doc = Document::new();
let font_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
});
let mut content = b"q Q\n".repeat(50);
content.extend_from_slice(b"BT /F1 10 Tf 10 10 Td (X) Tj ET\n");
let form_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {
"Type" => "XObject",
"Subtype" => "Form",
"BBox" => vec![0.into(), 0.into(), 100.into(), 100.into()],
"Resources" => dictionary! {
"Font" => dictionary! {
"F1" => Object::Reference(font_id),
},
},
},
content,
)));
let mut budget = FormWalkBudget::with_limits(MAX_FORM_XOBJECT_INVOCATIONS, 10);
let items = extract_form(&doc, form_id, &mut budget);
assert!(
items.is_empty(),
"operation budget must stop before the trailing text show"
);
assert!(budget.was_truncated());
}
#[test]
fn page_level_form_dag_stays_within_production_budget() {
// A page-level `/Do` of an 8-wide, 6-level Form DAG would expand to
// 8^5 = 32_768 leaf drawings without a budget. The production
// invocation cap must keep extraction bounded.
let (doc, root) = form_dag(8, 6);
let (doc, page_id) = page_invoking_form(doc, root);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
assert!(
items.len() <= MAX_FORM_XOBJECT_INVOCATIONS,
"page-level Form expansion must stay within the invocation cap; got {}",
items.len()
);
assert!(
!items.is_empty(),
"budget must still allow some nested form text through"
);
}
#[test]
fn shared_form_budget_spans_two_extraction_passes() {
// The invisible-layer retry calls extract_page_text_items twice for
// the same page; both passes must share one budget.
let (doc, root) = form_dag(1, 2);
let (doc, page_id) = page_invoking_form(doc, root);
let font_cmaps = FontCMaps::from_doc(&doc);
// Root + leaf = 2 invocations on the first pass.
let mut budget = FormWalkBudget::with_limits(2, MAX_FORM_XOBJECT_OPERATIONS);
let ((first, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut budget,
)
.unwrap();
assert_eq!(first.iter().filter(|item| item.text == "X").count(), 1);
assert!(!budget.was_truncated());
let ((second, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
true,
&mut FontStyleCache::new(),
&mut budget,
)
.unwrap();
assert!(
second.iter().all(|item| item.text != "X"),
"second pass must not get a fresh invocation budget"
);
assert!(budget.was_truncated());
}
/// Build a document whose page draws *all* of its content through a single
/// Form XObject — the shape emitted by print-to-PDF producers like PDFlib,
/// where the page stream itself is only `q /X1 Do Q`.
fn doc_with_form_content(form_content: &[u8]) -> (Document, ObjectId) {
let mut doc = Document::new();
let widths: Vec<Object> = (0..=255).map(|_| 600.into()).collect();
let font_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
"FirstChar" => 0,
"LastChar" => 255,
"Widths" => Object::Array(widths),
});
let form_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {
"Type" => "XObject",
"Subtype" => "Form",
"BBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
"Resources" => dictionary! {
"Font" => dictionary! { "F1" => Object::Reference(font_id) },
},
},
form_content.to_vec(),
)));
let content_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
b"q /X1 Do Q".to_vec(),
)));
let page_id = doc.add_object(dictionary! {
"Type" => "Page",
"Contents" => Object::Reference(content_id),
"Resources" => dictionary! {
"XObject" => dictionary! { "X1" => Object::Reference(form_id) },
},
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
});
let pages_id = doc.add_object(dictionary! {
"Type" => "Pages",
"Count" => Object::Integer(1),
"Kids" => vec![Object::Reference(page_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => Object::Reference(pages_id),
});
doc.trailer.set("Root", Object::Reference(catalog_id));
(doc, page_id)
}
fn form_items(form_content: &[u8]) -> Vec<TextItem> {
let (doc, page_id) = doc_with_form_content(form_content);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
items
}
fn find<'a>(items: &'a [TextItem], text: &str) -> &'a TextItem {
items
.iter()
.find(|item| item.text == text)
.unwrap_or_else(|| {
let found: Vec<&String> = items.iter().map(|i| &i.text).collect();
panic!("no item {text:?} in {found:?}")
})
}
#[test]
fn t_star_inside_form_moves_to_next_line() {
// T* was previously unhandled inside Form XObjects, so every line after
// the first piled onto the preceding baseline and drifted right.
let items =
form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm (first) Tj T* (second) Tj ET");
let first = find(&items, "first");
let second = find(&items, "second");
assert!((first.y - 700.0).abs() < 0.1, "first y = {}", first.y);
assert!((second.y - 688.0).abs() < 0.1, "second y = {}", second.y);
assert!((second.x - 100.0).abs() < 0.1, "second x = {}", second.x);
}
#[test]
fn td_inside_form_is_relative_to_line_start_not_shown_text() {
// Td moves relative to the text *line* matrix. Applying it to the
// matrix already advanced by Tj marched each line off the right edge.
let items = form_items(b"BT /F1 12 Tf 1 0 0 1 100 700 Tm (AAAAA) Tj 0 -12 Td (B) Tj ET");
let b = find(&items, "B");
assert!((b.x - 100.0).abs() < 0.1, "B x = {} (expected 100)", b.x);
assert!((b.y - 688.0).abs() < 0.1, "B y = {}", b.y);
}
#[test]
fn td_inside_form_sets_leading_for_later_t_star() {
// `TD` sets the leading to -ty as a side effect; a following T* must
// reuse it.
let items = form_items(
b"BT /F1 12 Tf 1 0 0 1 100 700 Tm (one) Tj 0 -15 TD (two) Tj T* (three) Tj ET",
);
assert!((find(&items, "two").y - 685.0).abs() < 0.1);
let three = find(&items, "three");
assert!((three.y - 670.0).abs() < 0.1, "three y = {}", three.y);
assert!((three.x - 100.0).abs() < 0.1, "three x = {}", three.x);
}
#[test]
fn quote_operator_inside_form_moves_to_next_line() {
let items = form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm (first) Tj (second) ' ET");
let second = find(&items, "second");
assert!((second.y - 688.0).abs() < 0.1, "second y = {}", second.y);
assert!((second.x - 100.0).abs() < 0.1, "second x = {}", second.x);
}
#[test]
fn double_quote_operator_inside_form_sets_spacing_and_moves() {
// `aw ac (string) "` — set word spacing and char spacing, then T* and show.
let items =
form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm (first) Tj 0 0 (second) \" ET");
let second = find(&items, "second");
assert!((second.y - 688.0).abs() < 0.1, "second y = {}", second.y);
assert!((second.x - 100.0).abs() < 0.1, "second x = {}", second.x);
}
#[test]
fn char_spacing_inside_form_widens_advance() {
// Tc was hardcoded to 0 in the form parser, so advance widths drifted.
// 2 glyphs x 600/1000 x 12pt = 14.4, plus 2 x Tc(2.0) = 18.4.
let items = form_items(b"BT /F1 12 Tf 1 0 0 1 100 700 Tm 2 Tc (AB) Tj ET");
let ab = find(&items, "AB");
assert!((ab.width - 18.4).abs() < 0.1, "AB width = {}", ab.width);
}
#[test]
fn q_restores_fill_colour_inside_form() {
// A white fill set inside q/Q must not leak past the Q — otherwise the
// following black text is treated as invisible and dropped entirely.
let items = form_items(
b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm q 1 g (hidden) Tj Q T* (visible) Tj ET",
);
assert!(
items.iter().any(|item| item.text == "visible"),
"text after Q was dropped: {:?}",
items.iter().map(|i| &i.text).collect::<Vec<_>>()
);
assert!(
!items.iter().any(|item| item.text == "hidden"),
"white-filled text should still be suppressed"
);
}
#[test]
fn q_restores_text_state_inside_form() {
// Tc/TL live in the graphics state; `Q` must roll them back.
let items =
form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm q 30 TL (a) Tj Q T* (b) Tj ET");
let b = find(&items, "b");
assert!(
(b.y - 688.0).abs() < 0.1,
"b y = {} (leading should restore to 12)",
b.y
);
}
}
+40 -3
View File
@@ -4566,9 +4566,13 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
}
}
// Try to parse uniXXXX format
if name.starts_with("uni") && name.len() >= 7 {
if let Ok(code) = u32::from_str_radix(&name[3..7], 16) {
// Try to parse uniXXXX format.
// Use `get` rather than a byte-length check + slice: `name` can contain
// non-ASCII bytes (e.g. U+FFFD from lossy UTF-8 decoding of an attacker
// controlled /Differences name), so byte index 7 may not be a char
// boundary and `&name[3..7]` would panic.
if let Some(hex) = name.strip_prefix("uni").and_then(|rest| rest.get(..4)) {
if let Ok(code) = u32::from_str_radix(hex, 16) {
// Strip PUA F000 offset: uniF0XX → U+00XX (Windows Symbol encoding convention)
let code = if (0xF000..=0xF0FF).contains(&code) {
code - 0xF000
@@ -4588,3 +4592,36 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
None
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn uni_hex_parsing() {
assert_eq!(glyph_to_char("uni0041"), Some('A'));
assert_eq!(glyph_to_char("uni00e9"), Some('\u{00e9}'));
// PUA F0xx symbol-encoding offset is stripped.
assert_eq!(glyph_to_char("uniF041"), Some('A'));
}
#[test]
fn u_hex_parsing() {
assert_eq!(glyph_to_char("u0041"), Some('A'));
assert_eq!(glyph_to_char("u1F600"), Some('\u{1F600}'));
}
#[test]
fn non_ascii_uni_name_does_not_panic() {
// A crafted /Differences name like `/uni#80#80#80#80` decodes via
// from_utf8_lossy into "uni" followed by four U+FFFD replacements.
// Byte index 7 lands mid-character, so a naive `&name[3..7]` slice
// would panic. It must be handled gracefully instead.
let crafted = format!("uni{0}{0}{0}{0}", '\u{FFFD}');
assert_eq!(glyph_to_char(&crafted), None);
// Assorted non-ASCII bytes right after the "uni" prefix.
assert_eq!(glyph_to_char("uni\u{FFFD}bc"), None);
assert_eq!(glyph_to_char("uni\u{00e9}00"), None);
}
}
+173 -24
View File
@@ -657,6 +657,80 @@ pub fn extract_pages_markdown<P: AsRef<Path>>(
extract_pages_markdown_mem(&buffer, pages)
}
// =========================================================================
// Structure-tree element extraction (tagged PDFs)
// =========================================================================
/// One structure-tree element reference from a tagged PDF, resolved to a
/// page and Marked Content ID.
///
/// Join `(page, mcid)` against [`TextItem::page`] / [`TextItem::mcid`] from
/// [`extract_text_with_positions`] to attach semantic roles (heading levels,
/// paragraphs, table cells, …) to extracted text.
#[derive(Debug, Clone)]
pub struct StructureElement {
/// 1-indexed page number (matches [`TextItem::page`]).
pub page: u32,
/// Marked Content ID from the page's content stream (matches
/// [`TextItem::mcid`]).
pub mcid: i64,
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
/// Custom tags are resolved through the document's `/RoleMap`; tags
/// with no standard mapping are returned verbatim.
pub role: String,
}
/// Extract structure-tree element references from a tagged PDF in memory.
///
/// Parses `/StructTreeRoot` (when present) and returns one entry per
/// marked-content reference, resolved to its 1-indexed page, MCID, and
/// structure type name. Returns an empty list when the PDF is not tagged.
///
/// Pass `Some(&[...])` with 1-indexed page numbers (matching
/// [`TextItem::page`]) to restrict output to those pages; pass `None` for
/// the whole document. Entries are sorted by `(page, mcid)`.
pub fn extract_structure_elements_mem(
buffer: &[u8],
pages: Option<&[u32]>,
) -> Result<Vec<StructureElement>, PdfError> {
validate_pdf_bytes(buffer)?;
let (doc, _page_count) = load_document_from_mem(buffer)?;
let Some(tree) = structure_tree::StructTree::from_doc(&doc) else {
return Ok(Vec::new());
};
let page_ids = doc.get_pages();
let roles = tree.mcid_to_roles(&page_ids);
let page_filter: Option<HashSet<u32>> = pages.map(|p| p.iter().copied().collect());
let mut elements: Vec<StructureElement> = roles
.into_iter()
.filter(|(page, _)| page_filter.as_ref().is_none_or(|f| f.contains(page)))
.flat_map(|(page, mcids)| {
mcids.into_iter().map(move |(mcid, role)| StructureElement {
page,
mcid,
role: role.name().to_string(),
})
})
.collect();
elements.sort_unstable_by_key(|e| (e.page, e.mcid));
Ok(elements)
}
/// Path-based wrapper for [`extract_structure_elements_mem`].
///
/// Reads the PDF from disk and extracts structure-tree element references.
/// Pass `None` for `pages` to return the whole document, or `Some(&[...])`
/// to restrict to specific 1-indexed pages.
pub fn extract_structure_elements<P: AsRef<Path>>(
path: P,
pages: Option<&[u32]>,
) -> Result<Vec<StructureElement>, PdfError> {
validate_pdf_file(&path)?;
let buffer = std::fs::read(path.as_ref())?;
extract_structure_elements_mem(&buffer, pages)
}
// =========================================================================
// Region-based text extraction (for hybrid OCR pipelines)
// =========================================================================
@@ -683,6 +757,23 @@ pub struct PageRegionResult {
pub regions: Vec<RegionText>,
}
/// Minimum alphanumeric mass an invisible (Tr 3) text layer must carry for
/// the OCR-layer fallback in [`extract_text_in_regions_mem`] to adopt it. A
/// real OCR layer carries far more; a stray watermark or artifact does not.
const OCR_LAYER_MIN_ALNUM: usize = 40;
/// Alphanumeric mass of extracted items, ignoring raster placeholders.
/// `[Image: ...]` items (ItemType::Image) are synthesized for image
/// XObjects — they mark that pixels exist, not that text was read, so they
/// must not count as coverage.
fn non_placeholder_alnum(items: &[TextItem]) -> usize {
items
.iter()
.filter(|it| !matches!(it.item_type, types::ItemType::Image))
.map(|it| it.text.chars().filter(|c| c.is_alphanumeric()).count())
.sum()
}
/// Extract text within bounding-box regions from a PDF in memory.
///
/// This is designed for hybrid OCR pipelines: a layout model detects regions
@@ -735,8 +826,11 @@ pub fn extract_text_in_regions_mem(
let height = get_page_height(&doc, page_id).unwrap_or(792.0);
page_heights.insert(*page_num, height);
// Extract text items for this page
let ((mut items, _rects, _lines), has_gid, coords_rotated) =
// Extract text items for this page. The Form XObject budget is shared
// with the invisible-layer retry below so one page cannot consume two
// full expansion budgets.
let mut form_budget = extractor::FormWalkBudget::new();
let ((mut items, _rects, _lines), mut has_gid, mut coords_rotated, skipped_invisible) =
extractor::content_stream::extract_page_text_items(
&doc,
page_id,
@@ -744,7 +838,54 @@ pub fn extract_text_in_regions_mem(
&font_cmaps,
false,
&mut style_cache,
&mut form_budget,
)?;
// OCR-layer fallback: scanned pages often carry their text as an
// invisible (Tr 3) layer behind the page raster. The visible-only
// pass sees nothing there but `[Image: ...]` placeholders, so every
// region on the page reports needs_ocr even though the exact text is
// embedded in the PDF — and this extractor then disagrees with the
// markdown path, which already retries Mixed PDFs with the invisible
// layer included. Retry page-scoped, and only when (a) the first
// pass actually SKIPPED invisible text — blank pages and image-only
// scans without an OCR layer must not pay a second content-stream
// parse (review catch) — and (b) the page has NO visible text item
// at all (punctuation counts, whitespace-only artifacts don't): an
// invisible OCR layer transcribes the raster, so any visible glyph
// has an invisible twin there and adoption would duplicate it
// (review catches — strict gate, no fuzzy dedupe). Adopt the retry
// only when it contributes real, non-garbage text.
let has_visible_text = items.iter().any(|it| {
!matches!(it.item_type, types::ItemType::Image) && !it.text.trim().is_empty()
});
if skipped_invisible && !has_visible_text {
if let Ok(((inv_items, _inv_rects, _inv_lines), inv_gid, inv_rotated, _)) =
extractor::content_stream::extract_page_text_items(
&doc,
page_id,
*page_num,
&font_cmaps,
true,
&mut style_cache,
&mut form_budget,
)
{
let inv_alnum = non_placeholder_alnum(&inv_items);
// Judge the WHOLE recovered layer, not a prefix — a broken
// OCR layer can hide its garbage past any fixed sample size
// (review catch).
let sample: String = inv_items
.iter()
.filter(|it| !matches!(it.item_type, types::ItemType::Image))
.map(|it| it.text.as_str())
.collect();
if inv_alnum >= OCR_LAYER_MIN_ALNUM && !is_garbage_text(&sample) {
items = inv_items;
has_gid = inv_gid;
coords_rotated = inv_rotated;
}
}
}
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
page_thresholds.insert(*page_num, threshold);
@@ -901,7 +1042,7 @@ pub fn extract_tables_in_regions_mem(
let height = get_page_height(&doc, page_id).unwrap_or(792.0);
page_heights.insert(*page_num, height);
let ((mut items, rects, lines), has_gid, coords_rotated) =
let ((mut items, rects, lines), has_gid, coords_rotated, _skipped_invisible) =
extractor::content_stream::extract_page_text_items(
&doc,
page_id,
@@ -909,6 +1050,7 @@ pub fn extract_tables_in_regions_mem(
&font_cmaps,
false,
&mut style_cache,
&mut extractor::FormWalkBudget::new(),
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -1212,7 +1354,7 @@ pub fn detect_vector_grid_in_region_mem(
let needed_pages = HashSet::from([page_1idx]);
let font_cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed_pages));
let page_h = get_page_height(&doc, page_id).unwrap_or(792.0);
let ((mut items, rects, lines), _has_gid, coords_rotated) =
let ((mut items, rects, lines), _has_gid, coords_rotated, _skipped_invisible) =
extractor::content_stream::extract_page_text_items(
&doc,
page_id,
@@ -1220,6 +1362,7 @@ pub fn detect_vector_grid_in_region_mem(
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
&mut extractor::FormWalkBudget::new(),
)?;
text_utils::fix_letterspaced_items(&mut items);
@@ -1406,15 +1549,17 @@ mod vector_grid_tests {
let &page_id = pages.get(&1).unwrap();
let needed: HashSet<u32> = HashSet::from([1]);
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
&doc,
page_id,
1,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
)
.unwrap();
let ((items, rects, _lines), _has_gid, _rotated, _skipped_invisible) =
extract_page_text_items(
&doc,
page_id,
1,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
&mut crate::extractor::FormWalkBudget::new(),
)
.unwrap();
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, 1);
assert_eq!(rect_tables.len(), 1, "expected one rect-detected table");
@@ -1448,15 +1593,17 @@ mod vector_grid_tests {
let &page_id = pages.get(&page_num).unwrap();
let needed: HashSet<u32> = HashSet::from([page_num]);
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
&doc,
page_id,
page_num,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
)
.unwrap();
let ((items, rects, _lines), _has_gid, _rotated, _skipped_invisible) =
extract_page_text_items(
&doc,
page_id,
page_num,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
&mut crate::extractor::FormWalkBudget::new(),
)
.unwrap();
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, page_num);
rect_tables
@@ -2184,7 +2331,7 @@ pub fn extract_tables_with_structure_cells_mem(
let height = get_page_height(&doc, page_id).unwrap_or(792.0);
page_heights.insert(*page_num, height);
let ((mut items, _rects, _lines), _has_gid, coords_rotated) =
let ((mut items, _rects, _lines), _has_gid, coords_rotated, _skipped_invisible) =
extractor::content_stream::extract_page_text_items(
&doc,
page_id,
@@ -2192,6 +2339,7 @@ pub fn extract_tables_with_structure_cells_mem(
&font_cmaps,
false,
&mut style_cache,
&mut extractor::FormWalkBudget::new(),
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -2986,7 +3134,7 @@ fn detect_tsr_quality_issue(
let mut needed: HashSet<u32> = HashSet::new();
needed.insert(page_1idx);
let font_cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
let ((mut items, _rects, _lines), _has_gid, coords_rotated) =
let ((mut items, _rects, _lines), _has_gid, coords_rotated, _skipped_invisible) =
extractor::content_stream::extract_page_text_items(
&doc,
page_id,
@@ -2994,6 +3142,7 @@ fn detect_tsr_quality_issue(
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
&mut extractor::FormWalkBudget::new(),
)?;
let adaptive_threshold = text_utils::fix_letterspaced_items(&mut items);
let coords = if coords_rotated {
+89
View File
@@ -271,6 +271,12 @@ pub struct PyTextItem {
pub is_strikeout: bool,
#[pyo3(get)]
pub item_type: String,
/// Marked Content ID from the content stream's BDC/BMC operator, None
/// when the text is not part of marked content. Join with the
/// (page, mcid) pairs from extract_structure_elements to attach
/// structure-tree roles (headings, paragraphs, ...) in tagged PDFs.
#[pyo3(get)]
pub mcid: Option<i64>,
}
#[pymethods]
@@ -286,6 +292,32 @@ impl PyTextItem {
}
}
/// One structure-tree element reference from a tagged PDF.
#[pyclass(name = "StructureElement")]
#[derive(Clone)]
pub struct PyStructureElement {
/// 1-indexed page number (matches TextItem.page).
#[pyo3(get)]
pub page: u32,
/// Marked Content ID from the page's content stream (matches
/// TextItem.mcid).
#[pyo3(get)]
pub mcid: i64,
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", ...).
#[pyo3(get)]
pub role: String,
}
#[pymethods]
impl PyStructureElement {
fn __repr__(&self) -> String {
format!(
"StructureElement(page={}, mcid={}, role='{}')",
self.page, self.mcid, self.role
)
}
}
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
@@ -356,6 +388,18 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item_type_str(&item.item_type),
mcid: item.mcid,
})
.collect()
}
fn convert_structure_elements(elements: Vec<crate::StructureElement>) -> Vec<PyStructureElement> {
elements
.into_iter()
.map(|e| PyStructureElement {
page: e.page,
mcid: e.mcid,
role: e.role,
})
.collect()
}
@@ -613,6 +657,48 @@ fn extract_pages_markdown_bytes(
Ok(to_py_pages_result(result))
}
/// Extract structure-tree element references from a tagged PDF file.
///
/// Parses the document's structure tree (when present) and returns one
/// entry per marked-content reference, resolved to its 1-indexed page,
/// MCID, and structure type name ("H1".."H6", "P", "Table", ...). Returns
/// an empty list when the PDF is not tagged.
///
/// Join (page, mcid) against the page/mcid attributes from
/// [`extract_text_with_positions`] to attach heading levels and other
/// semantic roles to extracted text.
///
/// Args:
/// path: Path to the PDF file.
/// pages: Optional list of 1-indexed pages (matching TextItem.page).
/// When None (default), the whole document is returned.
///
/// Returns:
/// List of StructureElement sorted by (page, mcid).
#[pyfunction]
#[pyo3(signature = (path, pages=None))]
fn extract_structure_elements(
path: &str,
pages: Option<Vec<u32>>,
) -> PyResult<Vec<PyStructureElement>> {
let elements = crate::extract_structure_elements(path, pages.as_deref()).map_err(to_py_err)?;
Ok(convert_structure_elements(elements))
}
/// Extract structure-tree element references from tagged PDF bytes.
///
/// See [`extract_structure_elements`] for details.
#[pyfunction]
#[pyo3(signature = (data, pages=None))]
fn extract_structure_elements_bytes(
data: &[u8],
pages: Option<Vec<u32>>,
) -> PyResult<Vec<PyStructureElement>> {
let elements =
crate::extract_structure_elements_mem(data, pages.as_deref()).map_err(to_py_err)?;
Ok(convert_structure_elements(elements))
}
/// Python module definition.
#[pymodule]
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
@@ -620,6 +706,7 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_class::<PyPageOcrReasons>()?;
m.add_class::<PyPdfClassification>()?;
m.add_class::<PyTextItem>()?;
m.add_class::<PyStructureElement>()?;
m.add_class::<PyRegionText>()?;
m.add_class::<PyPageRegionTexts>()?;
m.add_class::<PyPageMarkdown>()?;
@@ -634,6 +721,8 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_function(wrap_pyfunction!(extract_text_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_with_positions, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_with_positions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_structure_elements, m)?)?;
m.add_function(wrap_pyfunction!(extract_structure_elements_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_pages_markdown, m)?)?;
+819 -38
View File
File diff suppressed because it is too large Load Diff
+170 -20
View File
@@ -594,7 +594,7 @@ fn hex_to_unicode_string(hex: &str) -> Option<String> {
let bytes: Option<Vec<u8>> = (0..hex.len())
.step_by(2)
.map(|i| u8::from_str_radix(&hex[i..i + 2], 16).ok())
.map(|i| u8::from_str_radix(hex.get(i..i + 2)?, 16).ok())
.collect();
let bytes = bytes?;
@@ -1561,23 +1561,31 @@ fn parse_encoding_cmap_stream(data: &[u8]) -> Option<EncodingCMap> {
}
let mut map = HashMap::new();
let mut assigned = 0usize;
let mut pos = 0;
while let Some(start) = text[pos..].find("begincidchar") {
let section_start = pos + start + "begincidchar".len();
if let Some(end) = text[section_start..].find("endcidchar") {
let section = &text[section_start..section_start + end];
parse_cidchar_section(section, &mut map, &mut src_hex_lengths);
if !parse_cidchar_section(section, &mut map, &mut src_hex_lengths, &mut assigned) {
break;
}
pos = section_start + end;
} else {
break;
}
}
pos = 0;
while let Some(start) = text[pos..].find("begincidrange") {
while assigned < MAX_CID_W_EXPANSION {
let Some(start) = text[pos..].find("begincidrange") else {
break;
};
let section_start = pos + start + "begincidrange".len();
if let Some(end) = text[section_start..].find("endcidrange") {
let section = &text[section_start..section_start + end];
parse_cidrange_section(section, &mut map, &mut src_hex_lengths);
if !parse_cidrange_section(section, &mut map, &mut src_hex_lengths, &mut assigned) {
break;
}
pos = section_start + end;
} else {
break;
@@ -1612,7 +1620,8 @@ fn parse_cidchar_section(
section: &str,
map: &mut HashMap<u16, u16>,
src_hex_lengths: &mut Vec<usize>,
) {
assigned: &mut usize,
) -> bool {
let mut chars = section.chars().peekable();
loop {
while chars.peek().is_some_and(|c| c.is_whitespace()) {
@@ -1643,16 +1652,20 @@ fn parse_cidchar_section(
}
}
if let (Some(code), Ok(cid)) = (parse_hex_u16(&src_hex), cid_str.parse::<u16>()) {
map.insert(code, cid);
if !assign_encoding_cid(map, code, cid, assigned) {
return false;
}
}
}
true
}
fn parse_cidrange_section(
section: &str,
map: &mut HashMap<u16, u16>,
src_hex_lengths: &mut Vec<usize>,
) {
assigned: &mut usize,
) -> bool {
let mut chars = section.chars().peekable();
loop {
while chars.peek().is_some_and(|c| c.is_whitespace()) {
@@ -1703,12 +1716,34 @@ fn parse_cidrange_section(
) else {
continue;
};
if start > end {
continue;
}
let mut cid = start_cid;
for code in start..=end {
map.insert(code, cid);
if !assign_encoding_cid(map, code, cid, assigned) {
return false;
}
cid = cid.saturating_add(1);
}
}
true
}
fn assign_encoding_cid(
map: &mut HashMap<u16, u16>,
code: u16,
cid: u16,
assigned: &mut usize,
) -> bool {
// Count overwrites: unique-key coverage alone would not stop a repeated
// full-width range from re-inserting all 65,536 codes.
if *assigned >= MAX_CID_W_EXPANSION {
return false;
}
map.insert(code, cid);
*assigned += 1;
true
}
fn parse_binary_cmap_encoding(data: &[u8]) -> Result<EncodingCMap, String> {
@@ -1832,6 +1867,13 @@ fn merge_cmaps(mut base: ToUnicodeCMap, overlay: ToUnicodeCMap) -> ToUnicodeCMap
base
}
/// Shared 16-bit CID expansion cap (65,536).
/// Encoding `begincidrange` and `/W` width assignment count every insert,
/// including overwrites, so a repeated full-width range cannot keep working
/// after the domain is filled. The `/W` unicode heuristic caps unique CIDs
/// with the same number.
pub(crate) const MAX_CID_W_EXPANSION: usize = 65_536;
/// Check if a CIDFont's /W (widths) array contains CID values that look like
/// Unicode codepoints rather than low-value GIDs.
///
@@ -1843,20 +1885,23 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
_ => return false,
};
// The /W array format: [cid [w1 w2 ...]] or [cid_start cid_end w]
// We extract all CID values (the first element of each group).
let mut cids: Vec<u16> = Vec::new();
// The /W array format: [cid [w1 w2 ...]] or [cid_start cid_end w].
// Collect unique CIDs only: repeating a full-width range must not grow a
// temporary vector (or the sort) with the range length on every copy.
let mut seen = HashSet::new();
let mut i = 0;
while i < w_arr.len() {
while i < w_arr.len() && seen.len() < MAX_CID_W_EXPANSION {
if let Ok(cid) = w_arr[i].as_i64() {
cids.push(cid as u16);
// Skip the width data
let start = cid as u16;
if i + 1 < w_arr.len() {
match &w_arr[i + 1] {
Object::Array(widths) => {
// [cid [w1 w2 ...]] — CIDs are cid, cid+1, ..., cid+len-1
for j in 1..widths.len() {
cids.push((cid as u16).wrapping_add(j as u16));
for j in 0..widths.len() {
if seen.len() >= MAX_CID_W_EXPANSION {
break;
}
seen.insert(start.wrapping_add(j as u16));
}
i += 2;
}
@@ -1864,9 +1909,7 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
// [cid_start cid_end w] — range of CIDs
if i + 2 < w_arr.len() {
if let Ok(cid_end) = w_arr[i + 1].as_i64() {
for c in (cid as u16)..=(cid_end as u16) {
cids.push(c);
}
record_unique_cid_range(start, cid_end as u16, &mut seen);
}
i += 3;
} else {
@@ -1875,6 +1918,7 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
}
}
} else {
seen.insert(start);
i += 1;
}
} else {
@@ -1882,10 +1926,11 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
}
}
if cids.is_empty() {
if seen.is_empty() {
return false;
}
let mut cids: Vec<u16> = seen.into_iter().collect();
cids.sort_unstable();
let median = cids[cids.len() / 2];
// Unicode text CIDs are typically >= 0x20 (space) with letters at 0x41+.
@@ -1894,6 +1939,18 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
median >= 0x41
}
fn record_unique_cid_range(start: u16, end: u16, seen: &mut HashSet<u16>) {
if start > end {
return;
}
for cid in start..=end {
if seen.len() >= MAX_CID_W_EXPANSION {
return;
}
seen.insert(cid);
}
}
/// Build a ToUnicodeCMap from predefined CID→Unicode mapping based on CIDSystemInfo.
///
/// Supports Adobe-Korea1 (Korean) character collection. Can be extended for
@@ -2606,6 +2663,24 @@ endcmap
assert_eq!(cmap.lookup(0x0025), Some("B".to_string()));
}
#[test]
fn test_hex_to_unicode_non_ascii_no_panic() {
// A destination containing a multi-byte char makes the byte length even
// while a byte offset can land inside a char. Slicing must not panic;
// it should be rejected gracefully.
assert_eq!(hex_to_unicode_string("XéY"), None);
assert_eq!(hex_to_unicode_string("\u{fffd}0"), None);
}
#[test]
fn test_parse_bfchar_non_ascii_destination_no_panic() {
// Crafted /ToUnicode CMap: a non-hex, non-ASCII destination previously
// triggered a char-boundary panic in hex_to_unicode_string.
let cmap_content = "beginbfchar <0041> <XéY> endbfchar";
// Must not panic; the malformed entry is simply skipped.
let _ = ToUnicodeCMap::parse(cmap_content.as_bytes());
}
#[test]
fn test_parse_bfchar_1byte() {
// This is the pattern that caused the CJK bug: codespace is <0000><FFFF>
@@ -3280,4 +3355,79 @@ endbfrange
"An indirect /Subtype naming CIDFontType2 must still reach the remap"
);
}
#[test]
fn cid_values_look_like_unicode_letter_range() {
let mut dict = lopdf::Dictionary::new();
dict.set(
"W",
Object::Array(vec![
Object::Integer(0x41),
Object::Integer(0x5A),
Object::Integer(500),
]),
);
assert!(cid_values_look_like_unicode(&dict));
}
#[test]
fn cid_values_look_like_unicode_low_gids() {
let mut dict = lopdf::Dictionary::new();
dict.set(
"W",
Object::Array(vec![
Object::Integer(0),
Object::Array(vec![Object::Integer(500); 10]),
]),
);
assert!(!cid_values_look_like_unicode(&dict));
}
#[test]
fn cid_values_look_like_unicode_repeated_full_ranges_stay_bounded() {
// Repeating `[0 65535 w]` must not materialize 65,536 CIDs per copy.
let mut w = Vec::new();
for _ in 0..5_000 {
w.push(Object::Integer(0));
w.push(Object::Integer(65535));
w.push(Object::Integer(500));
}
let mut dict = lopdf::Dictionary::new();
dict.set("W", Object::Array(w));
assert!(cid_values_look_like_unicode(&dict));
}
#[test]
fn encoding_cidrange_maps_a_normal_range() {
let data = b"1 begincodespacerange\n<0000> <FFFF>\nendcodespacerange\n\
1 begincidrange\n<0041> <0043> 65\nendcidrange\n";
let enc = parse_encoding_cmap_stream(data).unwrap();
assert_eq!(enc.map.get(&0x41), Some(&65));
assert_eq!(enc.map.get(&0x42), Some(&66));
assert_eq!(enc.map.get(&0x43), Some(&67));
assert_eq!(enc.map.len(), 3);
assert_eq!(enc.code_byte_length, 2);
}
#[test]
fn encoding_cidrange_repeated_full_ranges_stay_bounded() {
// 5,000 copies of `<0000> <ffff> 0` must not re-expand the 16-bit
// domain on every declaration.
let mut body = String::new();
let mut remaining = 5_000usize;
while remaining > 0 {
let n = remaining.min(100);
body.push_str(&format!("{n} begincidrange\n"));
for _ in 0..n {
body.push_str("<0000> <ffff> 0\n");
}
body.push_str("endcidrange\n");
remaining -= n;
}
let data = format!("1 begincodespacerange\n<0000> <FFFF>\nendcodespacerange\n{body}");
let enc = parse_encoding_cmap_stream(data.as_bytes()).unwrap();
assert!(enc.map.len() <= MAX_CID_W_EXPANSION);
assert_eq!(enc.map.get(&0), Some(&0));
assert_eq!(enc.map.get(&65535), Some(&65535));
}
}
+347
View File
@@ -1429,6 +1429,77 @@ fn test_firecrawl_tagged_pdf_struct_tree() {
assert_eq!(fence_count % 2, 0, "Code fences should be balanced");
}
#[test]
fn test_tagged_pdf_text_items_carry_mcid() {
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
assert!(
items.iter().any(|i| i.mcid.is_some()),
"Tagged PDF text items should carry Marked Content IDs"
);
}
#[test]
fn test_extract_structure_elements_tagged_pdf() {
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
assert!(!elements.is_empty(), "Tagged PDF should yield elements");
assert!(
elements.iter().any(|e| e.role == "H1"),
"Should surface H1 heading roles"
);
assert!(
elements.iter().all(|e| !e.role.is_empty()),
"Every element should carry a role name"
);
// Sorted by (page, mcid) for deterministic output
assert!(
elements
.windows(2)
.all(|w| (w[0].page, w[0].mcid) <= (w[1].page, w[1].mcid)),
"Elements should be sorted by (page, mcid)"
);
// The advertised join: (page, mcid) pairs must line up with the
// mcid-carrying TextItems from positioned extraction, and joining the
// H1 entries must recover non-empty heading text.
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
let h1_refs: std::collections::HashSet<(u32, i64)> = elements
.iter()
.filter(|e| e.role == "H1")
.map(|e| (e.page, e.mcid))
.collect();
let h1_text: String = items
.iter()
.filter(|i| i.mcid.is_some_and(|mcid| h1_refs.contains(&(i.page, mcid))))
.map(|i| i.text.as_str())
.collect();
assert!(
!h1_text.trim().is_empty(),
"Joining H1 structure elements to text items should recover heading text"
);
// Page filter is 1-indexed (matching TextItem.page) and equals the
// corresponding subset of the full document result.
let page1 = pdf_inspector::extract_structure_elements_mem(&buf, Some(&[1])).unwrap();
assert!(!page1.is_empty(), "Page 1 should have elements");
assert!(page1.iter().all(|e| e.page == 1));
let full_page1_count = elements.iter().filter(|e| e.page == 1).count();
assert_eq!(page1.len(), full_page1_count);
}
#[test]
fn test_extract_structure_elements_untagged_pdf_empty() {
let buf = std::fs::read("tests/fixtures/thermo-freon12.pdf").unwrap();
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
assert!(
elements.is_empty(),
"Untagged PDF should yield no structure elements, got {:?}",
elements
);
}
#[test]
fn test_identity_h_no_tounicode_suppresses_garbage() {
// shinagawa_identity_h.pdf uses YuGothic with Identity-H encoding and no
@@ -1583,6 +1654,282 @@ fn test_extract_regions_mem_basic_text_pdf() {
assert_eq!(regions[0].page, 0);
}
/// Build a synthetic "scanned page" PDF: a full-page image XObject with a
/// text layer drawn in the given render mode (3 = invisible OCR overlay,
/// 0 = normal visible fill). `visible_extra` optionally adds a normally
/// rendered line so double-layer behavior can be tested; `layer_lines`
/// overrides the layer content (default: three pangram lines);
/// `quote_ops` shows every layer line via the `'` operator instead of Tj
/// (both are standard show-text encodings for OCR layers).
fn make_pdf_with_custom_text_layer(
text_render_mode: i32,
visible_extra: Option<&str>,
layer_lines: Option<&[&str]>,
quote_ops: bool,
) -> Vec<u8> {
let mut pdf = b"%PDF-1.4\n".to_vec();
let mut offsets = vec![0usize];
fn add_object(pdf: &mut Vec<u8>, offsets: &mut Vec<usize>, id: usize, body: &str) {
offsets.push(pdf.len());
pdf.extend_from_slice(format!("{id} 0 obj\n").as_bytes());
pdf.extend_from_slice(body.as_bytes());
pdf.extend_from_slice(b"\nendobj\n");
}
fn add_stream_object(
pdf: &mut Vec<u8>,
offsets: &mut Vec<usize>,
id: usize,
dict: &str,
stream_bytes: &[u8],
) {
offsets.push(pdf.len());
pdf.extend_from_slice(format!("{id} 0 obj\n").as_bytes());
pdf.extend_from_slice(
format!("<< {} /Length {} >>\nstream\n", dict, stream_bytes.len()).as_bytes(),
);
pdf.extend_from_slice(stream_bytes);
pdf.extend_from_slice(b"\nendstream\nendobj\n");
}
add_object(
&mut pdf,
&mut offsets,
1,
"<< /Type /Catalog /Pages 2 0 R >>",
);
add_object(
&mut pdf,
&mut offsets,
2,
"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
);
add_object(
&mut pdf,
&mut offsets,
3,
"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] \
/Resources << /Font << /F1 5 0 R >> /XObject << /Im0 6 0 R >> >> \
/Contents 4 0 R >>",
);
// Full-page raster, then the text layer in the requested render mode —
// several lines so the OCR-layer gate's alnum floor (40) is well cleared.
let mut content = String::from("q 612 0 0 792 0 0 cm /Im0 Do Q\n");
let default_layer = [
"The quick brown fox jumps over the lazy dog",
"Pack my box with five dozen liquor jugs tonight",
"Sphinx of black quartz judge my vow carefully",
];
let layer: &[&str] = layer_lines.unwrap_or(&default_layer);
if quote_ops {
// Every line shown via `'` (move-to-next-line + show) — nothing on
// this layer goes through Tj, pinning the `'` suppression path.
content.push_str(&format!(
"BT /F1 12 Tf {text_render_mode} Tr 16 TL 72 716 Td "
));
for line in layer {
content.push_str(&format!("({line}) ' "));
}
} else {
content.push_str(&format!("BT /F1 12 Tf {text_render_mode} Tr 72 700 Td "));
for (i, line) in layer.iter().enumerate() {
if i > 0 {
content.push_str("0 -16 Td ");
}
content.push_str(&format!("({line}) Tj "));
}
}
content.push_str("ET\n");
if let Some(extra) = visible_extra {
content.push_str(&format!("BT /F1 12 Tf 0 Tr 72 500 Td ({extra}) Tj ET\n"));
}
add_stream_object(&mut pdf, &mut offsets, 4, "", content.as_bytes());
add_object(
&mut pdf,
&mut offsets,
5,
"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica >>",
);
let image_pixel = [128u8];
add_stream_object(
&mut pdf,
&mut offsets,
6,
"/Type /XObject /Subtype /Image /Width 1 /Height 1 \
/ColorSpace /DeviceGray /BitsPerComponent 8",
&image_pixel,
);
let xref_start = pdf.len();
pdf.extend_from_slice(format!("xref\n0 {}\n", offsets.len()).as_bytes());
pdf.extend_from_slice(b"0000000000 65535 f \n");
for offset in offsets.iter().skip(1) {
pdf.extend_from_slice(format!("{offset:010} 00000 n \n").as_bytes());
}
pdf.extend_from_slice(
format!(
"trailer\n<< /Size {} /Root 1 0 R >>\nstartxref\n{}\n%%EOF",
offsets.len(),
xref_start
)
.as_bytes(),
);
pdf
}
fn make_pdf_with_text_layer(text_render_mode: i32, visible_extra: Option<&str>) -> Vec<u8> {
make_pdf_with_custom_text_layer(text_render_mode, visible_extra, None, false)
}
/// A scanned page whose only text is an invisible (Tr 3) OCR layer behind
/// the raster must serve that layer from the region extractor instead of
/// reporting the region as needs_ocr — the exact text is already in the PDF.
#[test]
fn test_extract_regions_mem_recovers_invisible_ocr_layer() {
let buf = make_pdf_with_text_layer(3, None);
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
assert_eq!(regions.len(), 1);
let region = &regions[0].regions[0];
assert!(
region.text.contains("quick brown fox"),
"invisible OCR layer should be served as region text, got: {:?}",
region.text
);
assert!(
!region.needs_ocr,
"recovered OCR layer must not fall back to GPU OCR"
);
}
/// ANY visible text on the page — even a single short line — must block the
/// invisible-layer adoption entirely: the invisible pass returns visible
/// items too, so adopting it alongside visible text would duplicate the
/// visible words. Strict zero-visible gate, no fuzzy dedupe.
#[test]
fn test_extract_regions_mem_visible_text_blocks_invisible_layer() {
let buf = make_pdf_with_text_layer(3, Some("Folio 142"));
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
let region = &regions[0].regions[0];
assert!(
region.text.contains("Folio 142"),
"visible text should be extracted, got: {:?}",
region.text
);
assert!(
!region.text.contains("quick brown fox"),
"invisible layer must not be adopted when any visible text exists, got: {:?}",
region.text
);
assert_eq!(
region.text.matches("Folio 142").count(),
1,
"visible text must appear exactly once, got: {:?}",
region.text
);
}
/// An invisible OCR layer shown entirely via the `'` show-text operator
/// (move-to-next-line + show) must also be recovered — the skipped_invisible
/// signal has to fire on every show-text path, not just Tj/TJ.
#[test]
fn test_extract_regions_mem_recovers_quote_operator_layer() {
let buf = make_pdf_with_custom_text_layer(3, None, None, true);
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
let region = &regions[0].regions[0];
assert!(
region.text.contains("quick brown fox"),
"'-operator OCR layer should be recovered, got: {:?}",
region.text
);
assert!(!region.needs_ocr);
}
/// An invisible layer below the 40-alnum floor (a stray watermark line)
/// must NOT be adopted — the region keeps its needs_ocr fallback.
#[test]
fn test_extract_regions_mem_tiny_invisible_layer_not_adopted() {
let buf = make_pdf_with_custom_text_layer(3, None, Some(&["Scanned by ACME"]), false);
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
let region = &regions[0].regions[0];
assert!(
!region.text.contains("Scanned by ACME"),
"below-floor invisible layer must not be adopted, got: {:?}",
region.text
);
// Only the raster placeholder remains — needs_ocr stays whatever main
// reports for placeholder-only regions (false today; downstream
// pipelines route placeholder-only text to OCR themselves, and this PR
// deliberately does not change that contract).
assert!(
region.text.trim().starts_with("[Image:"),
"region should hold only the raster placeholder, got: {:?}",
region.text
);
}
/// An invisible layer that clears the alnum floor but is mostly symbol
/// garbage (a broken OCR run) must be rejected by the garbage gate.
#[test]
fn test_extract_regions_mem_garbage_invisible_layer_not_adopted() {
// Each line: 5 alphanumerics among 15 symbol chars. Ten lines clear the
// 40-alnum floor (50 alnum) while staying well under the half-alnum
// ratio is_garbage_text requires.
let garbage_lines: Vec<&str> = vec!["a@@b%%c&&d==e~~"; 10];
let buf = make_pdf_with_custom_text_layer(3, None, Some(&garbage_lines), false);
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
let region = &regions[0].regions[0];
assert!(
!region.text.contains("a@@b"),
"garbage invisible layer must not be adopted, got: {:?}",
region.text
);
assert!(
region.text.trim().starts_with("[Image:"),
"region should hold only the raster placeholder, got: {:?}",
region.text
);
}
/// Punctuation-only visible text (zero alphanumerics) must ALSO block
/// adoption — the gate is item-presence, not alphanumeric mass. (Real-world
/// rationale: an invisible OCR layer transcribes the raster, so visible
/// glyphs typically have invisible twins there; this fixture's layers are
/// disjoint, so it pins the gate itself, not the duplication scenario.)
#[test]
fn test_extract_regions_mem_punctuation_visible_blocks_invisible_layer() {
let buf = make_pdf_with_text_layer(3, Some("... --- ..."));
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
let region = &regions[0].regions[0];
assert!(
!region.text.contains("quick brown fox"),
"invisible layer must not be adopted over punctuation-only visible text, got: {:?}",
region.text
);
assert_eq!(
region.text.matches("... --- ...").count(),
1,
"visible punctuation must be preserved exactly once, got: {:?}",
region.text
);
}
/// Regression guard: a normal visible-text page (render mode 0) is served
/// once and only once — if the fallback ever mis-fired here and merged a
/// second pass, the phrase would duplicate.
#[test]
fn test_extract_regions_mem_visible_layer_unchanged() {
let buf = make_pdf_with_text_layer(0, None);
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
let region = &regions[0].regions[0];
assert_eq!(
region.text.matches("quick brown fox").count(),
1,
"visible text must appear exactly once, got: {:?}",
region.text
);
assert!(!region.needs_ocr);
}
#[test]
fn test_extract_regions_mem_identity_h_needs_ocr() {
let buf = std::fs::read("tests/fixtures/shinagawa_identity_h.pdf").unwrap();
+73
View File
@@ -203,6 +203,79 @@ class TestExtractTextWithPositions:
assert len(items) > 0
assert all(item.page == 1 for item in items)
def test_mcid(self):
# Untagged fixture: mcid is None or int, never anything else
items = pdf_inspector.extract_text_with_positions(
fixture_path("thermo-freon12.pdf")
)
assert all(item.mcid is None or isinstance(item.mcid, int) for item in items)
# Tagged fixture: marked content carries MCIDs
tagged = pdf_inspector.extract_text_with_positions(
fixture_path("firecrawl_docs_tagged.pdf")
)
assert any(item.mcid is not None for item in tagged)
# ---------------------------------------------------------------------------
# extract_structure_elements / extract_structure_elements_bytes
# ---------------------------------------------------------------------------
class TestExtractStructureElements:
def test_tagged_file(self):
elements = pdf_inspector.extract_structure_elements(
fixture_path("firecrawl_docs_tagged.pdf")
)
assert len(elements) > 0
assert all(isinstance(e.page, int) for e in elements)
assert all(isinstance(e.mcid, int) for e in elements)
assert all(isinstance(e.role, str) and len(e.role) > 0 for e in elements)
assert any(e.role == "H1" for e in elements)
def test_join_with_text_items(self):
# (page, mcid) joins against extract_text_with_positions to recover
# heading text
path = fixture_path("firecrawl_docs_tagged.pdf")
elements = pdf_inspector.extract_structure_elements(path)
items = pdf_inspector.extract_text_with_positions(path)
h1_refs = {(e.page, e.mcid) for e in elements if e.role == "H1"}
h1_text = "".join(
item.text
for item in items
if item.mcid is not None and (item.page, item.mcid) in h1_refs
)
assert len(h1_text.strip()) > 0
def test_with_pages(self):
# pages filter is 1-indexed, matching TextItem.page
elements = pdf_inspector.extract_structure_elements(
fixture_path("firecrawl_docs_tagged.pdf"), pages=[1]
)
assert len(elements) > 0
assert all(e.page == 1 for e in elements)
def test_bytes(self):
data = fixture_bytes("firecrawl_docs_tagged.pdf")
elements = pdf_inspector.extract_structure_elements_bytes(data)
assert len(elements) > 0
assert any(e.role == "H1" for e in elements)
def test_untagged_returns_empty(self):
elements = pdf_inspector.extract_structure_elements(
fixture_path("thermo-freon12.pdf")
)
assert elements == []
def test_repr(self):
elements = pdf_inspector.extract_structure_elements(
fixture_path("firecrawl_docs_tagged.pdf")
)
assert "StructureElement" in repr(elements[0])
def test_not_a_pdf(self):
with pytest.raises(ValueError):
pdf_inspector.extract_structure_elements_bytes(b"not a pdf")
# ---------------------------------------------------------------------------
# extract_text_in_regions / extract_text_in_regions_bytes
+2 -2
View File
@@ -724,7 +724,7 @@ checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
[[package]]
name = "pdf-inspector"
version = "0.1.7"
version = "1.14.1"
dependencies = [
"env_logger",
"include_dir",
@@ -740,7 +740,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-wasm"
version = "0.1.3"
version = "1.14.1"
dependencies = [
"console_error_panic_hook",
"js-sys",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-wasm"
version = "0.1.3"
version = "1.14.1"
edition = "2021"
authors = ["Firecrawl Team"]
description = "Browser WebAssembly bindings for pdf-inspector"