Compare commits

..
Author SHA1 Message Date
Abimael Martell 0b0ca64ffb docs: document structure-element extraction and TextItem.mcid 2026-08-11 10:55:15 -07:00
Abimael Martell 7054d6aa69 chore(release): unify package versions (#344)
* chore(release): unify package versions

* fix(release): harden version synchronization
2026-08-11 09:26:19 -07:00
Abimael MartellandClaude Fable 5 a67ee03269 feat(bindings): expose TextItem.mcid and structure-tree element extraction (#346)
Tagged PDFs carry a structure tree with real heading roles (H1..H6), and
the core already parses it (structure_tree::StructTree) and threads MCIDs
onto TextItem — but neither surfaced through the bindings.

- Expose TextItem.mcid (Option<i64>) through the napi and pyo3 bindings,
  matching the core field added with the marked-content extractor.
- Add StructRole::name(), the inverse of from_name, so roles have a
  stable string form.
- Add extract_structure_elements / extract_structure_elements_mem to the
  core: one (page, mcid, role) entry per marked-content reference, sorted
  by (page, mcid), empty for untagged PDFs. Pages are 1-indexed to match
  TextItem.page, so results join directly against
  extract_text_with_positions output.
- Bind it as extractStructureElements (napi) and
  extract_structure_elements / extract_structure_elements_bytes (pyo3),
  with type-stub updates in pdf_inspector.pyi.
- Cover the join in Rust integration tests, napi test.mjs, and pytest,
  using the existing firecrawl_docs_tagged.pdf fixture (tagged) and
  thermo-freon12.pdf (untagged).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 20:03:07 -07:00
Abimael Martell 965dc65f1b chore(release): bump package versions (#343) 2026-08-10 14:33:40 -07:00
36dd5fa426 fix(structure-tree): bound recursive /K parsing with cycle detection and a node budget (#322)
* fix(structure-tree): bound tagged /K parsing against alias/cycle DoS

A struct element that references itself (or an ancestor) through /K — e.g.
/K [n 0 R n 0 R] — made parse_struct_element_dict branch exponentially:
the depth cap (64) alone still permits 2^depth materialized nodes, so a
~830-byte PDF exhausts memory (OOM, exit 134).

Add a StructWalk carrying (1) an active-path set of object IDs so a node
that references itself/an ancestor is not re-expanded (breaks self- and
mutual-reference cycles cheaply), and (2) a global node budget
(MAX_STRUCT_NODES) that caps total materialization for aliased/DAG-shaped
graphs of distinct objects the path guard cannot catch.

Adds regression tests for self-alias, mutual-alias, and the aliased-DAG
budget cap.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): charge /K content refs against the node budget

The per-node budget only covered materialized struct elements and child
recursion; bare MCIDs and MCR dicts in a /K array append to content_refs
without charging it, so one element with a very wide /K array could still
allocate content_refs without bound. Charge every /K array item before
handling it, and stop the top-level /K loop once the budget is spent, so
content refs and loop work are bounded too. Adds a wide-MCID-array test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): charge /K budget per materialized item, not per array entry

Charging every /K array item double-counted structural children (charged
here and again at their node entry) and charged cycle-skipped references
that materialize nothing, draining the budget up to ~2x faster than the
per-node semantics and risking early truncation of large legitimate trees.
Charge only the unbounded content-ref items (bare MCIDs and MCR dicts);
structural children remain charged once at their node entry.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* refactor(structure-tree): charge every content ref uniformly via helper

Route all budget charges through StructWalk::charge() so every
marked-content reference is charged once, including the single-value /K
branches (bare integer and MCR dict) that previously appended without
charging. charge() also guards against underflow, so charging after the
node-entry charge (which can leave the budget at 0) is safe. Makes the
documented per-item budget contract hold uniformly across all branches.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* feat(structure-tree): log once when the node budget truncates parsing

Add a one-shot truncation flag on StructWalk, set the first time the
budget is exhausted, and emit a single warn! after parsing so an operator
can tell when a (very large or malformed) tagged tree was cut off. Avoids
per-item log spam; negligible overhead on the normal path.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): flag truncation at budget guards, not just in charge

The truncation flag was only set inside charge() on the budget==0 branch,
but the dominant skip paths use budget==0 guards that break/return before
charge() is ever called with an empty budget, so the flag (and the warn!)
almost never fired. Route those guards through a new exhausted() that sets
the flag when it skips remaining work. Adds a parser-level test that would
have caught the missed warning.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): flag cycle/depth skips and charge bare MCIDs fully

Two review follow-ups:
- Cycle-broken and depth-capped /K skips dropped tagged content without
  setting the truncation flag, so the one-shot warning never fired for
  malformed/over-deep trees. Mark those skips via note_skipped() and
  broaden the warning to cover non-budget truncation.
- A bare /K MCID materializes a wrapper node AND a content reference but
  charged only one budget unit, allowing ~2x the advertised budget for
  such content; charge both.

Adds tests: cycle-skip flags truncation, and bare MCID charges two units.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): charge MCR-dict wrappers the same two units as bare MCIDs

A top-level MCR /K dict flows through parse_kid -> parse_struct_element_dict
and materializes a Span node + one content ref (two items) but was charged
only one unit at node entry, while the bare-MCID path charges two. Charge
the content reference in the MCR branch too so the per-item budget is
uniform across both wrapper paths. Adds a symmetric test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): reserve leaf-wrapper budget units atomically

A leaf MCID wrapper (bare MCID or MCR dict) materializes a node + one
content ref and charged the two units via separate charge() calls. At the
last unit the first charge succeeded and the second failed, consuming a
unit without emitting the wrapper and denying it to a later element that
would have fit. Add charge_n() to reserve both units atomically (or
neither), and detect MCR before the node charge so it reserves both up
front. Adds a boundary test asserting the leftover unit is preserved.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): stop scanning wide /K once a leaf reservation stalls

The atomic charge_n(2) left budget nonzero (==1) when it failed, so
exhausted() (budget==0) never broke the root /K loop and a crafted wide
array of leaf wrappers was scanned in full after no leaf could fit. Add a
stalled flag set on an insufficient reservation and fold it into
exhausted(); charge()-based (one-unit) loops are unaffected since they
reach budget 0 exactly. Adds a test that a one-unit budget still allows a
one-unit item but a failed two-unit reservation stops the scan.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): add traversal budget and stop charging non-materializing dicts

Two review follow-ups on budget accounting:
- Wide /K arrays of non-materializing items (unsupported value types, OBJR
  dicts, cycle back-edges) consumed no node budget, so the loop scanned the
  whole array. Add a separate work budget charged per examined /K item and
  break the loops when it is spent, bounding traversal even when nothing
  materializes.
- OBJR dicts and dicts without a valid /S were charged the node budget before
  being recognized and skipped, draining the shared budget and truncating
  real content later. Hoist the OBJR check and /S validation above the node
  charge so only materializing nodes consume it (matching the MCR hoisting).

Adds tests for the work-budget bound, wide unsupported /K, and non-materializing
dicts not charging the node budget.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* docs(structure-tree): mention traversal budget in truncation warning

The one-shot truncation warning listed the node budget, cycle, and depth
as causes but not the new traversal (work) budget, so a work-budget
truncation printed a misleading message. Include MAX_STRUCT_WORK so
malformed-PDF debugging identifies the actual limit hit.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 13:36:40 -07:00
fabec0aec3 feat(napi): add async variants that keep the Node event loop free (#337)
* feat(napi): add processPdfAsync, classifyPdfAsync, extractPagesMarkdownAsync

The Node bindings are synchronous, so every call parses on the event
loop thread — up to hundreds of milliseconds of dead loop per document
in a server. Add additive AsyncTask-based variants that run the same
shared implementations on the libuv thread pool and return promises.

The existing synchronous exports keep their names, signatures, and
behaviour; each sync/async pair shares one implementation. Panics in
compute() are caught and surfaced as rejections, matching the sync
error contract.

Closes #336

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(napi): read async task buffers in place instead of copying

Review feedback on #337: buffer.to_vec() copied the whole PDF on the
event loop before the task was queued, so large inputs still stalled
the loop and doubled peak memory. The tasks now hold the napi Buffer
itself — its ref pins the JS allocation for the task's lifetime and
the backing store is stable, so compute() reads it directly from the
worker thread. Callers must not mutate the buffer until the promise
settles (same contract as Node's async fs APIs); documented on each
export and in the README.

The suggested removal of ts_return_type was checked and rejected:
without it napi-rs generates Promise<unknown> for AsyncTask returns.
A comment now records that finding.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(napi): copy async task input on the JS thread for soundness

Review feedback on #337: holding the napi Buffer and reading it from
the libuv worker was unsound. Buffer derefs straight to the JS-side
allocation, so a caller mutating it before the promise settled would
race the worker's reads — undefined behavior, not a recoverable error,
and the documented don't-mutate contract was unenforceable. Deferring
the copy to compute() would not help: any off-thread read races the
same way. The JS thread is the only race-free place to take the copy,
because JS is single-threaded and nothing can mutate the buffer during
the synchronous part of the call.

Revert to an owned Vec<u8> copied at call time. The cost is one memcpy,
negligible next to the parse the async variants exist to unblock. Docs
now state the buffer may be reused or mutated immediately, and a test
locks in the copy semantics by mutating the input while a parse is in
flight.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 12:29:23 -07:00
1f28c00a13 Bound column-detection histogram and harden coordinate handling (#328)
* Bound column-detection histogram size

Derive the projection histogram from a clamped bin count and skip
non-finite page widths. Extreme or malformed text-item coordinates
(from the content-stream text matrix) could otherwise drive a very
large allocation. 65,536 bins is ~9x the largest legal page, so real
layouts are unaffected. Adds regression tests.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Exclude non-finite coordinates from page bounds

Items at NaN/inf positions are now skipped when folding the page bounds,
so a malformed coordinate can no longer escape as a ColumnRegion
boundary, and an all-non-finite page returns no columns. Bad items are
dropped individually rather than failing the page, so one stray glyph
does not disable column detection.

Addresses review feedback on the finite-width guard.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Trim far-outlier coordinates from page bounds

Gutter margins, spanning-item width and the XY-cut margin are all
fractions of page_width, so a single far-but-finite item (x=50_000 is
enough) set the scale for the whole page: real gutters fell inside the
rejected margin band and a genuine two-column page collapsed to one
region. When the span exceeds one legal page (14_400 units), re-derive
the bounds from items clustered around the median x. Outliers keep their
text because column assignment buckets by nearest overlap.

The MAX_BINS ceiling stays as an allocation bound that does not depend
on this heuristic.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Harden bounds trimming against widths and wide layouts

Check both item edges when trimming: a malformed width at an ordinary
position poisoned x_max just as a malformed position poisoned x_min, so
a huge width still collapsed a two-column page to one region.

Only trim when the far items are a small minority (<=10%). A genuinely
large-format page has content spread across its full width, so it now
keeps its true bounds instead of being reduced to the median cluster.

Correct the MAX_PAGE_EXTENT comment: 14_400 units is the traditional
Acrobat architectural limit, not a format cap. PDF 2.0 sets no page-size
limit and UserUnit scales physical size, so this is a heuristic.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Scale bin width so the histogram spans the whole page

Clamping the bin count alone left anything past MAX_BINS * BIN_WIDTH
(~131k points) outside the histogram, folded into the final bin. A page
wide enough to hit that lost real gutters: with a visible gutter inside
the covered range the XY-cut fallback never runs, so a three-column
layout silently reported two. Derive bin_width from page_width instead,
keeping the same allocation ceiling and degrading only resolution.

Also anchor the trimming median on the same finite left/right items that
bounds() accepts, so a malformed width cannot shift which items count as
strays.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Require geometric evidence before detaching far content

The median-window trim narrowed any content more than one page from the
centre, so a valid large page with a sparse far sidebar lost the sidebar
from its bounds and its text fell into column 0. An item-count minority
rule cannot tell that layout from malformed coordinates.

Group content into clusters separated by more than a whole page of
continuous emptiness, and only drop a cluster that is both detached by
such a void and a small minority of items. Real content does not leave a
gap that large; a stray coordinate sits alone beyond one.

A single run wider than one page is treated as a malformed width, which
also covers the huge-width case the cluster sweep cannot see (such an
item spans everything and leaves no gap).

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Judge run width against page content, not a fixed extent

Treating any run wider than 14_400 units as malformed penalised valid
large pages: one made entirely of such runs reported no columns at all,
and a mixed page lost the right edge of every long run.

Judge width relative to the page's own content instead. Positions cannot
be inflated by a bogus width, so the spread of the core cluster is a
sound scale: a run wider than that spread plus one page is malformed.
A genuinely large page keeps its genuinely long runs, while a 1e12-wide
run beside ordinary text is still rejected.

Cluster on positions rather than filled intervals, so a bogus width can
no longer merge everything into one cluster, and keep ordinary pages on
an O(n) fast path that skips the sort entirely.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 12:29:12 -07:00
f4b8c9e854 Clarify SECURITY.md reporting channels (#329)
* Update SECURITY.md reporting channels

Clarify that email is the only required channel and point the
alternative at Firecrawl's Bugcrowd disclosure engagement instead of
the private-advisory link, which is not enabled on this repo.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Make Bugcrowd the preferred reporting channel

Bugcrowd's disclosure engagement is the primary channel; email to
help@firecrawl.dev is offered as the alternative.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 16:27:50 -07:00
69039f2728 Fix char-boundary panic in hex_to_unicode_string (#320)
Use hex.get(i..i+2) instead of &hex[i..i+2] so a non-hex, non-ASCII
destination in a /ToUnicode CMap can no longer trigger a UTF-8
char-boundary panic. An even byte length does not guarantee the byte
offset falls on a char boundary; get() returns None on a non-boundary
or out-of-range index, folding cleanly into the existing flow.

Add regression tests covering a multi-byte destination char and a
replacement-char byte.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 08:21:31 -07:00
3cca6446bd fix(glyph_names): handle non-ASCII input in uniXXXX glyph name parsing (#321)
Use str::get instead of a byte-length check plus slice when parsing the
uniXXXX glyph-name form. The byte-length guard only proved the index was
in bounds, not on a UTF-8 char boundary, so a glyph name containing
non-ASCII bytes could cause a slice on a non-boundary index. Switch to a
checked slice that folds into the existing Option flow, and add tests.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 08:21:22 -07:00
493fed498e fix(links): prevent stack-overflow DoS from AcroForm /Kids self-cycle (#314)
* fix(links): guard AcroForm /Kids traversal against cycles and huge trees

A crafted PDF whose AcroForm field lists itself (or another ancestor) in
/Kids caused walk_form_fields to recurse indefinitely, overflowing the
stack and aborting pdf2md (exit 134) — an application-level DoS from a
~730-byte input.

Track visited field object IDs to break /Kids cycles, and cap total
field-node traversal at 100k nodes to bound pathologically large trees.

Adds regression tests for self-cycle and mutual-cycle field graphs.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): cap AcroForm /Kids recursion depth to stop deep-chain overflow

The visited-set guard stops cyclic /Kids graphs, but a long *acyclic*
chain of distinct fields still recurses to the chain length and overflows
the stack (a ~1.6MB PDF with 20k linked fields aborts pdf2md, exit 134)
before the 100k node budget is reached.

Add an explicit recursion depth cap (100 levels — far above any legitimate
form hierarchy) so stack usage is bounded independently of node count.

Adds a deep-acyclic-chain regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): enforce form-field node budget before insertion

The node-budget guard inserted each field ID into the visited set before
checking the budget, so the check triggered an early return but never
actually capped the set. A field with a huge /Kids array kept inserting
post-budget IDs, letting visited (memory and work) grow with the crafted
input rather than stopping at MAX_FORM_FIELD_NODES.

Check depth and budget before inserting, so visited can never exceed the
cap. Adds a wide-tree regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): stop /Fields and /Kids iteration once node budget is spent

Checking the budget before insertion capped the visited set, but callers
still iterated every remaining entry of a wide /Fields or /Kids array
after the budget was exhausted — each walk returned immediately, yet the
O(N) sibling iteration let a single multi-million-entry array burn
extraction CPU unbounded. Break out of both the top-level and recursive
loops once visited reaches the cap, making the budget a true
traversal-work cap. Adds a top-level wide-/Fields regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): charge examined entries against the field-node budget

The budget counted only distinct visited nodes, so /Fields or /Kids
arrays full of invalid (non-reference) or duplicate entries never grew
visited and ran to completion regardless of size — the node budget did
not actually cap traversal work.

Introduce FieldWalkBudget tracking both visited nodes and total entries
examined; charge every array entry (valid, invalid, or duplicate) and
stop once either hits MAX_FORM_FIELD_NODES. Adds a regression test with a
huge /Kids array of duplicate + null entries.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): iterate /Fields and /Kids arrays by borrow, not clone

Both arrays were cloned in full before the budget check, so a crafted
oversized /Fields or /Kids array forced an O(n) allocation and copy
regardless of the cap. resolve_array already returns a borrow tied to the
document and the walker only needs a shared &Document, so iterate the
borrowed arrays directly — the early break now bounds how many entries
are even touched, before any per-array allocation.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* docs(links): correct wide-array test comments to match range assertions

The two wide-array tests assert item counts within a range near the
budget, not an exact value (charging entries in the entry guard shifts
the boundary by one or two). Fix the stale comments that claimed exact
counts.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-08 23:43:33 -07:00
Abimael Martell 436af97038 fix(tables): exclude script attachments and tiny numeric fragments from detection (#264)
* fix(tables): exclude script attachments and tiny numeric fragments from detection

Split out of #242 (draft) so it can be reviewed on its own evidence.

Display equations with sub/superscripts form phantom small-font table
regions: the subscripts cluster with nearby small text (footnotes, axis
labels) into fake multi-column grids. Two guards:

- Script attachment: a small-font item horizontally adjacent to a
  larger-font item at a genuine baseline offset is a sub/superscript,
  not a table cell, and is excluded from candidates. A real baseline
  offset is required so a small cell beside a larger same-baseline
  label is never filtered. Attachment targets are indexed by Y and
  scanned through a bounded window rather than a full-page sweep.
  The body-font pass applies the same exclusion but only for
  heading-sized anchors (>= 1.15x base), so body-size table cells
  beside slightly larger labels are untouched.
- Tiny numeric fragments: a <=2-row grid whose every cell is a bare
  1-2 digit number carries no tabular information. Restricted to the
  small-font pass, where the pattern is overwhelmingly exponent
  clusters; body-font numeric grids are unaffected.

Corpus impact: 20 of 186 documents, measured against a control build of
main so main's own drift is excluded. Table rows fall in 17 of 18
inspected documents and no content is lost — 2103_07786 drops all 21
rows, every one a math fragment ('|X 42 43|1|||'); Stijn_SB_doc drops
173 rows of footnote text that had been shredded into cells, with word
count slightly UP and footnote markers intact. M2019_mordeste gains 11
rows from a 3-column grid re-detected as 2-column, neither clearly
better nor worse.

Note the 20 documents is far more than the 3 that per-change ablation
suggested: that figure measured sole-cause attribution inside the
original combined PR, where other heuristics changed the same files and
masked this one. Reach and sole-cause are different measurements.

805 unit + 148 integration tests pass, clippy clean, in-repo snapshots
unchanged.

* fix(tables): suppress script column evidence instead of dropping candidates

Reworked after reviewing the corpus diffs: the first approach removed
sub/superscripts from table candidates entirely, which had two failure
modes beyond the intended fix.

- Legitimate cell content was displaced. citizen-sr-282 is a calculator
  manual whose engineering-notation table lists M = 10^6, k = 10^3.
  Those exponents are superscripts, so they were dropped from the table
  and resurfaced elsewhere in the reading order ('9 mega 6 kilo = 10 3
  milli').
- Removing items changed the candidate geometry, so different spurious
  structure could form from what remained.

Scripts are now kept as candidates and excluded only from the geometry:
they cannot create a column (find_column_boundaries), cannot qualify a
region on their own (find_table_regions / _strict), but are still
assigned to cells. Column alignment is validated against ALL items
including scripts — validating only the non-script subset would let a
region manufacture alignment by ignoring its awkward items, which is
what block-diagram pages did.

Corpus: 20 of 186 documents, net -359 table rows, no content lost.
Token-level comparison shows the only text changes are merges in the
right direction: 'X' + '10' becomes 'X10', 'L' + 'g' becomes 'Lg' —
subscripts joining their base instead of floating free.

Remaining artifact: MCF5235RM (and its _nxp duplicate) gains a small
spurious table from a block-diagram label line, and M2019_mordeste
gains 11 rows from a 3-column grid re-detected as 2-column. Both are
borderline regions where the previous output was also wrong; documented
rather than tuned away.

805 unit + 148 integration tests pass, clippy clean.

* fix(tables): use a heading-anchored script mask in the body-font pass

Cubic review of #264: a single script mask computed with a 0.0 anchor
was applied to both passes, including body-font region qualification and
geometry. The body pass is supposed to require a heading-sized anchor —
that distinction existed before the geometry rework and was lost in it.

Why it matters: body-pass candidates are themselves body-sized
(0.85..1.05x base). A cell at the low end of that band, say 8.5pt,
sitting beside a 10.5pt label clears the inherent 'anchor >= 1.2x cell'
rule (10.2) and so was flagged as a script attachment. At body sizes a
slightly larger neighbour is a bold label or column header, not the base
of a superscript, so flagging it stripped real cells out of the region
evidence and column geometry and could lose the table entirely.

Two masks now: the small-font pass keeps the 0.0 anchor, the body pass
requires >= 1.15x base. Note the threshold only bites below base size —
for a cell at base, 1.2x-of-cell already exceeds 1.15x-of-base — which
is exactly the 0.85..1.0x band cubic identified.

Corpus: 20 documents, net -367 table rows (was -359 with the single
mask), so the body pass now keeps 8 rows of real table it had been
discarding. 958 tests pass, clippy clean.

* fix(markdown): reject headings that end on a relational verb

A heading candidate ending in 'equals', 'denotes', 'implies' and the
like is the first half of a sentence, not a title. This shows up when a
block dissolves and strands its lead-in ahead of the formula it
introduced — opendataloader 01030000000144 produced

    ## Note that the exact error equals
    M - Q(h) = e - 2.7525... = -0.0342....

Deliberately a very short list. Broader variants were tried and
measured, then rejected:

- Function words (of/and/for/the): a heading that WRAPS across lines
  ends on exactly those. Destroyed real IRS Publication 17 headings —
  'Casualty and' -> 'Casualty and Theft Losses', 'Rule 10. You Must Be
  at' -> '... At Least Age 25'. 52 documents affected, -619 headings.
- Copulas and auxiliaries (is/are/be/have): same failure. 'Rule 15.
  Your AGI Must Be', 'What Medical Expenses Are' and 'When Can a Roth
  IRA Be' are real wrapped headings, while 'the tax burden should be'
  is a genuine fragment. The trailing word cannot separate them; that
  needs the next line's context, which this text-only predicate lacks.

The verbs kept never end a heading in any register, so they are safe
without context. Standalone the guard is a no-op on both benchmarks
(0 documents on opendataloader, 4 on pdf-evals with no net heading
change) — its value is as a companion to the table filter in this PR,
which is what strands these lead-ins.

Combined effect on opendataloader (200 docs, vs a control build of
main), where the table filter alone regressed:

                 table filter    + this guard
    overall        -0.0003          +0.0003
    mhs            -0.0019          +0.0003
    doc ...144     -0.063           +0.053
    doc ...144 mhs -0.203           +0.028

* review: gate the dangling-verb veto on sentence case, drop 'yields'

Cubic review of a5a6e8f — both findings valid.

1. 'yields' is also a plural noun. 'Bond Yields', 'Crop Yields' and
   'Dividend Yields' are real section titles in financial documents,
   which this corpus contains. Removed from the list; my claim that
   these verbs 'never end a heading in any register' was wrong for it.

2. A wrapped title-case heading whose first line ends on one of these
   verbs would be suppressed if the heading preprocessor failed to
   merge it.

Both are fixed by the same gate, which is the discriminator I was
missing: case. A heading is title case ('Bond Yields', 'The Theorem
Implies'); a stranded lead-in is sentence case ('Note that the exact
error equals', 'the method yields'). The veto now applies only when
every content word is NOT capitalized, so titles are spared regardless
of their final word.

This is also why the earlier function-word and copula variants failed:
they had no way to tell 'Rule 15. Your AGI Must Be' from 'the tax
burden should be'. Case separates those two as well.

No measured cost. opendataloader is unchanged from the previous
revision — overall +0.0003, mhs +0.0003, doc 01030000000144 still
0.732 -> 0.785 — and pdf-evals still 20 documents. 963 tests pass,
clippy clean.

* review: exempt section-numbered lines from the dangling-verb veto

Valid ordering bug. heading.rs consults is_heading_fragment at line 282
and only applies its numbered-prefix allowance at line 288, so the veto
pre-empted it: '1. What the model implies' is sentence case and ends on
a listed verb, so it was discarded before numbering could vouch for it.

Numbering is independent evidence of a heading, so the veto now skips
any line opening with a section number.

Acceptance is deliberately a little broader than heading::parse_numbering
(which requires a trailing delimiter) because '2.3 Section Title' is
written without one, and being permissive in a veto exemption can only
avoid suppressing headings. Two guards keep it from swallowing prose:

- a bare single number needs a delimiter ('1.' yes, '3 apples' no)
- roman numerals always need one, since a leading 'I' is the pronoun far
  more often than a section number

Not reused from convert::starts_with_section_number, which deliberately
demands two components because it bypasses isolation checks — that would
reject the reviewer's single-'1.' case.

No measured change: opendataloader still overall +0.0003 / mhs +0.0003
with doc 01030000000144 at 0.732 -> 0.785, pdf-evals still 20 documents,
target case still suppressed. 964 tests pass, clippy clean.

* review: share roman_value so the veto exemption matches the parser

Valid. My numbering predicate accepted tokens heading::parse_numbering
rejects — lowercase 'iv)', alphabetical 'd)', over-long 'MMMM.' — because
it case-folded and allowed D and M. Anything the parser rejects is not
numbering, so exempting it let ordinary list items bypass the
dangling-verb veto and reach font-based heading promotion.

Rather than restate the grammar, roman_value is now pub(super) and the
exemption calls it, so the two cannot drift. Its rules apply as written:
uppercase I/V/X/L/C only, at most 8 characters, positive total.

Decimal numbering keeps its slightly broader acceptance (bare '2.3' with
no trailing delimiter), which is deliberate and documented — that form is
common in real headings and being permissive in a veto exemption cannot
manufacture a heading, only decline to suppress one. The roman case is
different because single letters collide with alphabetical list markers.

No measured change: opendataloader overall +0.0003 / mhs +0.0003, doc
01030000000144 still 0.732 -> 0.785. 964 tests pass, clippy clean.

* test: cover the roman length bound with a nine-character token

Valid P3. The 'MMMM.' case fails on the unsupported M, not on length, so
the 8-character bound in roman_value had no coverage and could regress
silently. Added a nine-'I' token, which is rejected only by the bound,
plus an eight-'I' token that must stay exempt to pin the boundary from
both sides.

* fix(tables): stop dropping body-band scripts from the candidate set

Valid: the body-font pass filtered scripts out of body_candidates
itself, so body_script_flags and its two downstream uses were dead. The
mask filters region_evidence and feeds detect_table_in_region's is_script
closure, but neither ever saw a script item because the candidate set no
longer contained any.

Consequences: a body-band sub/superscript attached to a heading-sized
anchor was dropped from the table outright rather than assigned to a
cell, so its text was lost — the opposite of what both the
body_script_flags comment ('they stay candidates') and the
detect_table_in_region docstring ('they remain eligible for cell
assignment') describe, and inconsistent with the small-font pass.

Root cause: the geometry rework removed the candidate-level filter from
the small-font pass, but the body one had been reflowed onto a single
line by rustfmt so the same edit missed it. Adding body_script_flags in
a later review then wired a mask that the surviving filter made
unreachable.

No measured change on either benchmark — pdf-evals still 20 documents
and -367 table rows, opendataloader still overall +0.0003 / mhs +0.0003
with one document changed — because the combination it affects (a
body-sized script attached to a heading-sized anchor) does not occur in
either corpus. The fix is for correctness and consistency between the
two passes, not for a score.

964 tests pass, clippy clean.
2026-08-07 09:41:46 -07:00
Abimael Martell f731e1191c fix(site): refresh benchmark results (#289) 2026-08-06 12:52:00 -07:00
Andrew Barnes 54a1e9ab74 Preserve page-prefixed Markdown content (#284) 2026-08-06 12:14:27 -07:00
m-naoki-mandClaude Opus 5 fabbb63521 fix(tounicode): skip subset GID remap for CIDFontType0 (CFF) descendants (#209)
* fix(tounicode): skip subset GID remap for CIDFontType0 (CFF) descendants

The sequential-GID repair in try_remap_subset_cmap assumes CIDs are glyph
indices that a subsetter can renumber. That holds for CIDFontType2
(TrueType) but not for CIDFontType0 (CFF), where CIDs are resolved through
the CFF charset, so a valid ToUnicode CMap stays valid after subsetting.

For CFF fonts the corrupting path was unavoidable: CIDToGIDMap is
CIDFontType2-only (PDF 32000-1:2008, 9.7.4.2), so the branch that repairs
the CMap correctly can never be taken, and any CFF font whose /W array
starts at a low CID fell through into remap_to_sequential. Japanese
Adobe-Japan1 documents extracted as long runs of a single unrelated kanji.

Guard both repair paths on a CIDFontType2 descendant, placed before the
CIDToGIDMap branch so a CIDToGIDMap wrongly attached to a CFF font by a
malformed producer is ignored too.

On a National Diet Library proceedings PDF: 1233 U+FFFD in 69099 chars
before, 0 in 68331 after; character 3-gram recall against a hand-written
ground truth 0.354 -> 0.605. The PDF from #118 (CIDFontType2) extracts
byte-identically before and after.

Two existing tests build descendant dicts without a /Subtype and set
CIDToGIDMap, which is CIDFontType2-only, so the fixtures now say what they
already meant. Without that, test_try_remap_skipped_when_w_covers_cmap
would keep passing while no longer exercising the W-coverage logic.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tounicode): resolve indirect /Subtype, skip only explicit non-CIDFontType2

Addresses review feedback on the guard added in the previous commit.

- /Subtype may be an indirect reference, and as_name() does not dereference
  it. A genuine CIDFontType2 font storing /Subtype indirectly would have been
  read as "not CIDFontType2", returning early and losing the repair it needs —
  reintroducing the corruption this PR fixes, for those fonts. Resolve the
  reference through the document before comparing.

- Bail out only when /Subtype is explicitly a non-CIDFontType2 name. A missing
  or unresolvable /Subtype now keeps the pre-existing behaviour instead of
  silently disabling the repair. As a result the two existing tests no longer
  need fixture changes, and this commit reverts those; the diff against main
  is now additive only.

- The CFF regression test now attaches a real CIDToGIDMap stream rather than
  /Identity, which get_cid_to_gid_map treats as "no map". With the stream, the
  test also fails if the guard is moved back below the CIDToGIDMap branch —
  verified by moving it and watching it fail.

- Added test_try_remap_resolves_indirect_subtype.

cargo fmt --check, cargo clippy -- -D warnings and cargo test (862 tests) pass.
The Diet PDF still extracts with 0 U+FFFD and the #118 PDF is still unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 12:02:27 -07:00
585d36e6a6 fix: recover from a corrupted startxref pointer (#230)
* fix: recover from a corrupted startxref pointer

Fixes #228.

A PDF whose startxref pointer has been corrupted to point at the wrong
byte offset — a single flipped digit, which is what damaged writers
emit in the wild — was entirely unprocessable: every entry point
(classify_pdf, extract_pages_markdown, process_pdf) raised "Invalid
PDF structure", even though the file's object data, real xref table,
and trailer were all completely intact just past the wrong pointer.
Both pypdf and pdfium recover from this by locating the real table
directly instead of trusting the pointer; lopdf doesn't.

Added a new repair candidate (alongside the existing
missing-%%EOF-marker and stripped-leading-bytes repairs in
repair_pdf_container_candidates): scan the buffer for the real,
standalone `xref` keyword and append a corrected trailing
`startxref`/`%%EOF` block. lopdf's own get_xref_start always reads the
*last* `%%EOF` in the final 512 bytes of the buffer and the
`startxref` value immediately before it, so the appended block
transparently supersedes the corrupted one already in the file — no
in-place byte surgery on content the original writer produced.

Scoped to classic (non-stream) xref tables, matching the reported
repro and the common case; a corrupted pointer into a cross-reference
*stream* (`N 0 obj << /Type /XRef ...>>`, some PDF 1.5+ writers) would
need the containing object's number, not just a byte offset — out of
scope here.

Verified against the issue's exact repro (a valid one-page PDF with a
single corrupted byte in its startxref offset): before this fix,
process_pdf/classify_pdf/extract_pages_markdown all raised "Invalid
PDF structure"; after, both the page count and the real extracted text
("Order Detail Report by Account", "WIDGET ASSEMBLY", the dollar
amount) come back correctly. New regression test added.

Full suite (859 tests, 1 new) passes; cargo clippy --all-targets
-- -D warnings unchanged at 28 pre-existing/unrelated errors.

* fix: validate xref table shape and scan in a single reverse pass

Addresses cubic-dev-ai's review of #230.

- P2 (correctness/safety): the recovery candidate trusted the last
  standalone "xref" token unconditionally, without confirming it's
  actually a cross-reference table. A coincidental "xref" substring
  inside unrelated content — a stream, a string, uncompressed
  metadata — could get "repaired" against a bogus offset, letting
  lopdf load successfully against garbage instead of returning a
  clean error: a real failure turned into silent data corruption on
  the fallback path. Added looks_like_xref_subsection_header, which
  confirms a plausible classic xref subsection header (`<start-id>
  <count>`, e.g. "0 6" — the shape every real classic table starts
  with) actually follows the candidate token before accepting it.
  find_last_valid_xref_table_start now walks backward from the end of
  the buffer until it finds a token that both stands alone *and*
  validates, rather than accepting the first (rightmost) standalone
  match unconditionally.

- P2 (performance): the old scan re-invoked
  `buf[..search_end].windows(4).rposition(...)` on a shrinking prefix
  every time a candidate token failed the boundary check, which is
  quadratic on a pathological buffer with many non-standalone "xref"
  occurrences. Rewrote as a single reverse byte-index walk — O(n)
  regardless of how many false candidates it has to reject along the
  way.

Added direct unit tests on the byte-level scan (more precise than
constructing adversarial full PDFs, and the coincidental-match
scenario can't be represented in an integration-test fixture anyway
since reportlab compresses page content by default): a coincidental
standalone "xref" with no subsection header is rejected; a real
classic table is found; a coincidental match positioned *after* the
real table in the buffer doesn't shadow it; "xref" as a substring of
"startxref" still doesn't match. The original #228 repro (corrupted
startxref pointer, real table otherwise intact) is unaffected —
verified manually in addition to the existing integration test.

Full suite (863 tests, 5 new) passes; cargo clippy --all-targets
-- -D warnings unchanged at 28 pre-existing/unrelated errors.

* fix: reject xref subsection count runs with trailing garbage

looks_like_xref_subsection_header validated that a count run of digits
followed the whitespace separator, but never checked what came after
it. A coincidental "xref\n0 6garbage" in stream/literal content would
still validate as a real subsection header shape and get repaired
against a bogus offset.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Abimael Martell <1450169+abimaelmartell@users.noreply.github.com>
2026-08-05 16:43:11 -07:00
371de80b14 fix: extract_pages_markdown's needs_ocr now agrees with classify_pdf (#231)
* fix: extract_pages_markdown's needs_ocr now agrees with classify_pdf

Fixes #227.

extract_pages_markdown_mem computed its per-page needs_ocr entirely
from text-quality signals: decoding/garble issues, empty markdown, GID
fonts, garbage-text ratio. It had no awareness of the page's image
content at all — so a page that is fundamentally a full-page scan
with a little genuine native text drawn over it (a header, a stamp, a
cover-sheet annotation) extracts that text cleanly, trips none of the
text-quality checks, and reports needs_ocr=false — while
classify_pdf/detect_pdf_type correctly see the dominant background
image and flag the same page as needing OCR. Two public APIs
answering the same question, silently disagreeing, in the unsafe
direction (skipping OCR on a page that needs it).

Exposed detector::analyze_page_images at crate visibility (was
private) and call it per page in extract_pages_markdown_mem's loop —
the same "large background image" signal (>50% page coverage) that
already powers has_template_image in classify_pdf/detect_pdf_type,
rather than reimplementing image-area detection a second time with
its own thresholds that could drift out of sync again. When it's
true, the page is flagged needs_ocr (with OCR_REASON_SCANNED added
to ocr_reasons_by_page, matching how the same signal is already
reported elsewhere) and its markdown is blanked, exactly like the
existing text-quality-triggered needs_ocr paths already do — no
special-casing added for "cleanly-extracted-but-still-a-scan" text.

Verified against the issue's exact repro (a full-page raster with one
native text line drawn over it, built via reportlab/pillow): before
this fix, extract_pages_markdown_bytes reported page 0
needs_ocr=False with the header line as markdown while
classify_pdf_bytes correctly flagged pages_needing_ocr=[0]; after,
both agree needs_ocr=True and the page's markdown is empty. Confirmed
no regression on a normal text-based fixture (nexo-price-en.pdf:
needs_ocr stays False, full markdown returned). New Rust regression
test added exercising both APIs against the same fixture.

Full suite (860 tests, 1 new) passes; cargo clippy --all-targets
-- -D warnings unchanged at 28 pre-existing/unrelated errors.

* fix: gate has_template_image behind the same OCR signals classify_pdf uses

extract_pages_markdown_mem was treating has_template_image alone as
sufficient to force needs_ocr=true and discard the page's markdown, but
classify_pdf/detect_pdf_type never treats that raw signal alone as
needing OCR. A text page with a full-bleed watermark, letterhead, or
large figure would get its clean markdown wrongly blanked and routed
to OCR.

Added page_template_image_needs_ocr(), mirroring the two distinct
signals classify_pdf actually uses to decide a template-image page
needs OCR: the looks_like_scan gate (image_count <= 1, few text ops,
low alphanumeric diversity) used for Mixed-type routing, and the
insufficient-text-volume signal (text_operator_count < 10) that routes
a page with a dominant background image and only a couple of native
text calls to PdfType::ImageBased independent of looks_like_scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: match per-page OCR threshold and add missing vector-text signal

Two follow-up findings on the has_template_image gate added in the
previous commit:

1. insufficient_text used a hard-coded threshold of 10 text operators,
   but Mixed-type per-page routing (the actual per-page decision this
   function tries to agree with) uses config.min_text_ops_per_page
   (default 3). The higher 10 threshold was borrowed from a *different*
   classify_pdf code path — the effective_min_ops floor used only for
   whole-document ImageBased/Scanned classification, a cross-page
   aggregate this per-page function can't replicate anyway. Using the
   lower per-page threshold removes a real disagreement window
   (3-9 text ops with high alphanumeric diversity) without breaking
   the #227 regression fixture (text_ops=1, still well under 3).

2. extract_pages_markdown_mem never checked has_vector_text at all,
   even though Mixed-type per-page routing always sends
   vector-outlined-text pages to OCR (outlined glyphs can't be
   extracted as text). A page with massive path ops plus a short
   genuine caption could extract that caption cleanly, slipping past
   the existing empty/garbage-text checks. Added
   page_has_vector_text() and wired it into needs_ocr the same way
   has_template_image is.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* perf: compute template-image and vector-text OCR signals in one pass

page_template_image_needs_ocr and page_has_vector_text each called
analyze_page_content independently, so every requested page's content
streams (page + XObjects) and image coverage were decompressed and
scanned twice per page with one result discarded each time.
detect_from_document avoids this by caching its per-page PageAnalysis;
extract_pages_markdown_mem had no such cache.

Merged both into page_ocr_signals(), a single analyze_page_content
call returning both signals as a tuple.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Abimael Martell <1450169+abimaelmartell@users.noreply.github.com>
2026-08-05 15:41:44 -07:00
Abimael Martell ede48099c0 fix(layout): preserve ruled tables and chart prose order (#262)
* fix(layout): preserve ruled tables and chart prose order

* fix(layout): harden chart region detection

* fix(layout): tighten chart geometry guards

* fix(layout): bound chart inference

* fix(layout): tighten chart claim bounds

* fix(layout): tighten chart evidence

* fix(layout): preserve edge-adjacent chart labels

* fix(layout): require external chart label overlap
2026-08-05 10:22:13 -07:00
Abimael Martell 12e9a655e3 fix(extractor): supply built-in metrics for non-embedded base-14 fonts (#241)
* fix(extractor): supply built-in metrics for non-embedded base-14 fonts

PDFs may legally omit /Widths for non-embedded standard fonts (Times,
Helvetica, Courier, Symbol, ZapfDingbats) — the spec requires the reader
to supply the metrics. We returned None, so every glyph advanced 0 and
each text item got width 0, silently breaking every gap-based heuristic
downstream: space synthesis, sub/superscript detection, table column
detection, heading merging.

- src/extractor/base14.rs: Adobe Core-14 AFM width tables keyed by
  Unicode char, plus the standard Symbol/ZapfDingbats encoding vectors
  (their glyphs sit at byte positions unrelated to Latin text, so widths
  must resolve through the built-in encoding, not cp1252)
- Width resolution order: Differences -> built-in encoding -> the same
  cp1252-style fallback the text decoder uses, so a code's advance always
  matches the character we emit for it
- Type3 visual sizing: PK bitmap fonts (dvips) use FontMatrix
  [1 0 0 -1 0 0] with nominal sizes like 0.12pt; scale by FontBBox height
  x |matrix_y|. Applied in the page-stream and Form XObject paths.
  Indirect numeric array elements are resolved before use.

Effect on Shannon's 'A Mathematical Theory of Communication' (1998
dvips/Distiller, the reported case): glued sentences 95 -> 5. Corpus
impact: 12 of 184 eval documents, e.g. Data-Processing-Agreement
recovers a paragraph that a phantom table had shredded into cells.

Layout heuristics tuned on the same document (indent-based paragraph
breaks, heading reclassification, table script filtering) are held back
for a separate PR — they change ~98 further documents and need to be
justified against the corpus, not against one PDF.

* review: narrow Type3 rescaling to self-inconsistent fonts; dedup + test all width tables

Addresses cubic review on #241, plus a follow-up from a local cubic run.

- Type3 visual scaling was applied to every Type3 font whose FontBBox
  height x |matrix_y| deviated >5% from 1.0. FontBBox is the glyph box,
  not the em box, so a conventional 1/1000-matrix font with a
  descender..ascender bbox (~700 units) computed 0.7 and had every
  reported size shrunk by 30% — corrupting the drop-cap, heading-tier,
  sub/superscript and table heuristics this is meant to fix.

  First attempt gated on the matrix being unit-scale, but a local cubic
  run pointed out that wrongly excludes valid non-standard matrices (a
  0.005 matrix with a full-em bbox legitimately needs a 5x scale). The
  product is the right discriminator, not the matrix: a self-consistent
  font lands near 1.0 because the matrix is the reciprocal of the
  glyph-space em, so only a wildly inconsistent one (dvips/PK bitmap
  fonts sit at ~159) is renormalized. Band widened to [0.25, 4.0].

  Corpus effect: 12 -> 7 documents change. The 5 that drop out were
  being wrongly rescaled — including Data-Processing-Agreement, whose
  phantom-table fix turned out to come from this bug rather than from
  the width fallback, so it is correctly given up.

- base14: all 14 width tables now covered by the sort-invariant test via
  an ALL_TABLES registry, not a hand-picked subset.
- base14: identical tables share one static (all four Courier variants
  are monospace 600; the oblique Helvetica variants match their upright
  forms), removing 5 duplicate copies.

* test: refresh Shannon snapshot after merging main

CI checks out a merge of the PR head with main, and main advanced 8
commits since this branch was cut — including #201 (contextual digit
runs), #240 and #253 (markdown fixes). Those change extraction output,
so a snapshot generated on the unmerged branch could not match; the
Test job failed on the merge commit while passing on the branch itself.

The merged behaviour is better: the footnote marker '2' before
'Hartley, R. V. L.' is now recovered instead of dropped.

950 tests pass on the merged tree, clippy clean.
2026-08-04 18:00:43 -07:00
Abimael MartellandClaude Fable 5 1d134e26aa docs: sync AGENTS.md with CLAUDE.md, refresh eval workflow guidance (#243)
AGENTS.md was stale (179+ PDFs, missing the semantic-quality bullet).
Both files now match: ~200-PDF corpus, and iteration guidance to prefer
subset runs (bench.py test -q / -s <name>) with the full suite as the
final pre-commit check.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 17:26:04 -07:00
Shubham Mathur bfd6c3eabb fix(extractor): make the comment stripper escape-aware (#259)
strip_pdf_comments tracked parenthesis nesting to protect string
literals, but ignored backslash escapes. An escaped \) desynced the
depth counter, after which a % glyph inside a string was stripped as a
top-level comment, corrupting the stream for Content::decode and
silently truncating the page's text.

Treat \ inside a string literal as escaping the next byte, so \(,
\), and \\ never touch the nesting depth.
2026-08-04 16:22:25 -07:00
Sheroy Cooper 04abab951f Support password-protected PDF item JSON extraction (#245) 2026-08-04 12:41:22 -07:00
Sunil a410d5aa08 fix: sync Python type stubs (.pyi) with actual bindings and fix doc bugs (#250)
## Bug 1: Python type stubs missing 5 fields/classes

The pdf_inspector.pyi file was out of sync with the actual Python
bindings exposed via #[pyo3(get)] in src/python.rs. This breaks
IDE autocompletion and type checking (PyCharm, VS Code/Pylance, mypy)
for all Python users.

Added:
- PdfResult.ocr_reasons_by_page (python.rs:35)
- PageOcrReasons class with page and 
easons fields (python.rs:67-86)
- RegionText.ocr_reason (python.rs:136)
- PageMarkdown.ocr_reason (python.rs:193)
- PagesExtractionResult.ocr_reasons_by_page (python.rs:226)

## Bug 2: PdfResult.pages_needing_ocr indexing undocumented

PdfResult.pages_needing_ocr is 1-indexed (per python.rs:30) but
neither the .pyi stubs nor docs/python.md annotated this, while the
same field on PdfClassification was annotated as 0-indexed. Users
mixing both APIs would get wrong page numbers.

## Bug 3: README.md duplicate bullet character

The Markdown features table listed * twice in bullet prefixes.
The first should be ullet (U+2022), matching the actual source code in
src/markdown/mod.rs:34 which uses: ullet, dash, sterisk, circle, illed-circle, open-circle.

## Bug 4: docs/python.md missing fields in type reference

The Types section was missing PageOcrReasons, RegionText class
definition, ocr_reason fields, and ocr_reasons_by_page fields.

## Evidence

Cross-referenced every #[pyo3(get)] attribute in src/python.rs
against the .pyi declarations and docs/python.md type reference.
2026-08-04 12:36:30 -07:00
Abimael Martell 7747b3a086 fix(markdown): reject non-text strikeout rules (#253)
* fix(markdown): reject non-text strikeout rules

* fix(markdown): address strikeout ownership edge cases

* fix(markdown): reject connected filled strike rules

* fix(markdown): group drifted strikeout runs
2026-08-04 11:31:58 -07:00
59 changed files with 7332 additions and 242 deletions
+3
View File
@@ -24,6 +24,9 @@ jobs:
- name: Cache cargo
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Run tests
run: cargo test --verbose
+3
View File
@@ -31,6 +31,9 @@ jobs:
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check if version changed
id: check
run: |
+3
View File
@@ -28,6 +28,9 @@ jobs:
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check if version changed
id: check
run: |
+3
View File
@@ -28,6 +28,9 @@ jobs:
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check package version
id: check
run: |
+3
View File
@@ -28,6 +28,9 @@ jobs:
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check if version changed
id: check
run: |
+2 -2
View File
@@ -31,12 +31,12 @@ Thumbs.db
napi/index.js
napi/index.d.ts
# Local samples and scripts
# Local samples
samples/
scripts/
# Test output
test_output/
.firecrawl/
# Python
__pycache__/
+2 -1
View File
@@ -61,7 +61,8 @@ src/
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
- **Regression suite**: sibling repo `pdf-evals` with 179+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
- **Regression suite**: sibling repo `pdf-evals` with ~200 snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing. While iterating, prefer a subset run (`bench.py test -q` for the quick set, or `-s <name>` for a named test set) and save the full `bench.py test` for the final pre-commit check.
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
## Debugging
+1 -1
View File
@@ -61,7 +61,7 @@ src/
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
- **Regression suite**: sibling repo `pdf-evals` with 187+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
- **Regression suite**: sibling repo `pdf-evals` with ~200 snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing. While iterating, prefer a subset run (`bench.py test -q` for the quick set, or `-s <name>` for a named test set) and save the full `bench.py test` for the final pre-commit check.
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
## Debugging
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector"
version = "0.1.7"
version = "1.14.0"
edition = "2021"
autobins = false
authors = ["Firecrawl Team"]
+2 -2
View File
@@ -110,7 +110,7 @@ Or add it manually:
```toml
[dependencies]
pdf-inspector = "0.1"
pdf-inspector = "1"
```
```rust
@@ -238,7 +238,7 @@ The converter handles:
|---|---|
| Headings (H1-H4) | Font size tiers relative to body text, with 0.5pt clustering |
| Bold/italic | Font name patterns (Bold, Italic, Oblique) |
| Bullet lists | `*`, `-`, `*`, `○`, `●`, `◦` prefixes |
| Bullet lists | ``, `-`, `*`, `○`, `●`, `◦` prefixes |
| Numbered lists | `1.`, `1)`, `(1)` patterns |
| Letter lists | `a.`, `a)`, `(a)` patterns |
| Code blocks | Monospace fonts (Courier, Consolas, Monaco, Menlo, Fira Code, JetBrains Mono) and keyword detection |
+4 -3
View File
@@ -5,14 +5,15 @@
If you believe you've found a security vulnerability in pdf-inspector, please
report it privately so we can fix it before public disclosure.
**Preferred:** Email **help@firecrawl.dev** with:
**Preferred:** Submit through Firecrawl's Bugcrowd vulnerability disclosure
program at <https://bugcrowd.com/engagements/firecrawl-vdp-ess>. Please include:
- A description of the issue and its impact
- Steps to reproduce (a minimal PDF or input that triggers the bug is ideal)
- The version or commit hash of pdf-inspector you tested against
**Alternative:** Use GitHub's private vulnerability reporting under the
[Security tab](https://github.com/firecrawl/pdf-inspector/security/advisories/new).
**Alternative:** If you'd rather not use Bugcrowd, email
**help@firecrawl.dev** with the same details.
We'll acknowledge your report in a timely manner and keep you updated on
remediation progress. Please do not open a public GitHub issue for security
+43 -24
View File
@@ -1,38 +1,57 @@
# Publishing
The Rust crate is published to [crates.io](https://crates.io/crates/pdf-inspector) with trusted publishing from GitHub Actions. The first release was published manually; future releases publish from `.github/workflows/publish-crate.yml` when a `Cargo.toml` version change lands on `main`.
Every pdf-inspector distribution uses one shared semantic version:
## crates.io Trusted Publisher
- Rust crate: `pdf-inspector`
- Python package: `pdf-inspector`
- Node package: `@firecrawl/pdf-inspector` and its platform packages
- Browser package: `@firecrawl/pdf-inspector-wasm`
- Internal NAPI and WASM Rust crates
Configure the trusted publisher for the `pdf-inspector` crate with:
`Cargo.toml` is the canonical version source. Update every manifest and lockfile
with:
- Repository: `firecrawl/pdf-inspector`
- Workflow: `publish-crate.yml`
- Environment: `crates-io`
```bash
python3 scripts/version.py <version>
```
The workflow uses `rust-lang/crates-io-auth-action@v1` to exchange GitHub's OIDC token for a short-lived crates.io token, then passes it to `cargo publish`.
Verify that nothing has diverged with:
## Release Steps
```bash
python3 scripts/version.py --check
```
1. Update `version` in `Cargo.toml`.
2. Merge the version bump to `main`.
3. The publish workflow compares the new `Cargo.toml` version with `HEAD~1`, runs `cargo publish --dry-run`, then publishes if that version is not already on crates.io.
CI and every publishing workflow run this check before building or publishing.
If `Cargo.toml` changes without a package version bump, the workflow exits without publishing.
## Release steps
## Browser WebAssembly package
1. Choose the next shared semantic version and run `scripts/version.py`.
2. Review the manifest and lockfile changes in the version-bump pull request.
3. Merge the pull request to `main`.
4. The crates.io, PyPI, Node, and WASM workflows independently build and
publish that version from the same commit.
5. After all registries succeed, create one `v<version>` GitHub release that
links to each package and describes changes since the previous shared tag.
The browser package is published as `@firecrawl/pdf-inspector-wasm`. Its version lives in `wasm/Cargo.toml`, and `.github/workflows/publish-wasm.yml` builds the `web` target with `wasm-pack` before publishing the generated package.
The independent workflows are intentionally idempotent. A manual dispatch from
`main` can repair a partial release, and already-published artifacts are skipped.
The npm package must exist before a trusted publisher can be configured. For the first release only:
## Trusted publishers
1. Build with `wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release`.
2. Inspect with `npm pack --dry-run ./wasm/pkg`.
3. Publish with `npm publish ./wasm/pkg --access public` from an authorized maintainer session.
4. In the package settings on npm, configure the GitHub Actions trusted publisher:
- Organization: `firecrawl`
- Repository: `pdf-inspector`
- Workflow: `publish-wasm.yml`
- Allowed action: `npm publish`
The repositories use GitHub Actions OIDC instead of long-lived registry tokens.
Configure each registry's trusted publisher for `firecrawl/pdf-inspector` and
its corresponding workflow:
After that one-time bootstrap, bumping the version in `wasm/Cargo.toml` and merging it to `main` publishes through OIDC. Until the package exists, the workflow exits cleanly without attempting an unauthenticated first publish. See npm's [trusted publishing documentation](https://docs.npmjs.com/trusted-publishers/) for the registry-side setup.
- crates.io: `publish-crate.yml`, environment `crates-io`
- PyPI: `publish-pypi.yml`, environment `pypi`
- npm Node package: `publish.yml`
- npm WASM package: `publish-wasm.yml`
The WASM package must exist before npm trusted publishing can be configured. If
it ever needs to be bootstrapped again, build and inspect it before publishing:
```bash
wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release
npm pack --dry-run ./wasm/pkg
npm publish ./wasm/pkg --access public
```
+33 -3
View File
@@ -80,6 +80,17 @@ for page in result.pages:
# Restrict to specific 0-indexed pages (preserves caller order)
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
# Structure-tree elements from tagged PDFs (empty list when untagged).
# Pages are 1-indexed to match TextItem.page, so (page, mcid) joins directly
# against extract_text_with_positions — e.g. to recover real heading levels:
elements = pdf_inspector.extract_structure_elements("tagged.pdf")
roles = {(e.page, e.mcid): e.role for e in elements}
headings = [
item.text
for item in pdf_inspector.extract_text_with_positions("tagged.pdf")
if item.mcid is not None and roles.get((item.page, item.mcid), "").startswith("H")
]
```
## API reference
@@ -100,6 +111,8 @@ result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes |
| `extract_pages_markdown(path, pages=None)` | Per-page Markdown + layout metadata (all pages by default) |
| `extract_pages_markdown_bytes(data, pages=None)` | Per-page Markdown from bytes |
| `extract_structure_elements(path, pages=None)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
| `extract_structure_elements_bytes(data, pages=None)` | Structure-tree elements from bytes |
## Types
@@ -111,7 +124,8 @@ class PdfResult: # process_pdf / detect_pdf
markdown: str | None # extracted Markdown (None for detect_pdf)
page_count: int
processing_time_ms: int
pages_needing_ocr: list[int]
pages_needing_ocr: list[int] # 1-indexed
ocr_reasons_by_page: list[PageOcrReasons]
title: str | None
confidence: float # 0.0 - 1.0
is_complex_layout: bool
@@ -119,6 +133,10 @@ class PdfResult: # process_pdf / detect_pdf
pages_with_columns: list[int]
has_encoding_issues: bool # broken font encodings — consider OCR fallback
class PageOcrReasons: # per-page OCR diagnostics
page: int # 1-indexed
reasons: list[str] # machine-readable reason identifiers
class PdfClassification: # classify_pdf
pdf_type: str
page_count: int
@@ -139,15 +157,27 @@ class TextItem: # extract_text_with_positions
is_underline: bool
is_strikeout: bool
item_type: str
mcid: int | None # marked-content ID for tagged PDFs (None otherwise)
class StructureElement: # extract_structure_elements
page: int # 1-indexed (matches TextItem.page)
mcid: int
role: str # "H1".."H6", "P", "Table", ... (resolved via /RoleMap)
class RegionText: # extract_text_in_regions
text: str
needs_ocr: bool
ocr_reason: str | None # machine-readable OCR reason
class PageRegionTexts: # extract_text_in_regions
page: int # 0-indexed
regions: list[RegionText] # RegionText: text: str, needs_ocr: bool
regions: list[RegionText]
class PagesExtractionResult: # extract_pages_markdown
pages: list[PageMarkdown] # PageMarkdown: page (0-indexed), markdown, needs_ocr
pages: list[PageMarkdown] # PageMarkdown: page (0-indexed), markdown, needs_ocr, ocr_reason
pages_with_tables: list[int] # 1-indexed
pages_with_columns: list[int] # 1-indexed
pages_needing_ocr: list[int] # 1-indexed
ocr_reasons_by_page: list[PageOcrReasons]
is_complex: bool # any page has tables or multi-column layout
```
+32 -1
View File
@@ -138,6 +138,34 @@ for page in &result.pages {
println!("Complex layout? {}", result.is_complex);
```
Extract structure-tree elements from tagged PDFs, and join them against
`extract_text_with_positions` to attach semantic roles (heading levels,
paragraphs, table cells) to extracted text:
```rust
use pdf_inspector::{extract_structure_elements, extract_text_with_positions};
use std::collections::HashMap;
// One entry per marked-content reference, sorted by (page, mcid); empty for
// untagged PDFs. Pages are 1-indexed to match `TextItem::page`, so the
// (page, mcid) pair is a direct join key.
let elements = extract_structure_elements("tagged.pdf", None)?;
let roles: HashMap<(u32, i64), &str> = elements
.iter()
.map(|e| ((e.page, e.mcid), e.role.as_str()))
.collect();
for item in extract_text_with_positions("tagged.pdf")? {
if let Some(mcid) = item.mcid {
if let Some(role) = roles.get(&(item.page, mcid)) {
if role.starts_with('H') {
println!("{}: {}", role, item.text);
}
}
}
}
```
## Processing modes
| Mode | What it does | Returns |
@@ -163,6 +191,8 @@ println!("Complex layout? {}", result.is_complex);
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
| `extract_pages_markdown(path, pages)` | Per-page Markdown + layout metadata (file) |
| `extract_pages_markdown_mem(bytes, pages)` | Per-page Markdown from bytes |
| `extract_structure_elements(path, pages)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
| `extract_structure_elements_mem(bytes, pages)` | Structure-tree elements from bytes |
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
@@ -178,7 +208,8 @@ Low-level detection functions are also available via the `detector` module (`det
| `DetectionConfig` | Configuration for detection: scan strategy, thresholds |
| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` |
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
| `TextItem` | Text with position, font info, and page number |
| `TextItem` | Text with position, font info, page number, and optional structure-tree `mcid` |
| `StructureElement` | Tagged-PDF structure reference: page (1-indexed), mcid, role (`"H1"`..`"H6"`, `"P"`, …) |
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
+2 -2
View File
@@ -851,7 +851,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
[[package]]
name = "pdf-inspector"
version = "0.1.7"
version = "1.14.0"
dependencies = [
"env_logger",
"include_dir",
@@ -867,7 +867,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-napi"
version = "0.2.2"
version = "1.14.0"
dependencies = [
"napi",
"napi-build",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-napi"
version = "0.2.2"
version = "1.14.0"
edition = "2021"
[lib]
+16
View File
@@ -83,6 +83,22 @@ for (const region of result[0].regions) {
}
```
### Async variants
`processPdf`, `classifyPdf`, and `extractPagesMarkdown` are synchronous and parse on the calling thread — in Node, that's the event loop. For a one-off call in a script that's fine, but in a server a large document can hold the loop for tens to hundreds of milliseconds.
`processPdfAsync`, `classifyPdfAsync`, and `extractPagesMarkdownAsync` take the same arguments and produce the same results, but run the parse on the libuv thread pool and return a promise, keeping the event loop free. The input buffer is copied before the call returns, so it's safe to reuse or mutate immediately:
```typescript
import { classifyPdfAsync, extractPagesMarkdownAsync } from '@firecrawl/pdf-inspector'
const classification = await classifyPdfAsync(pdf)
if (classification.pdfType === 'TextBased') {
const { pages } = await extractPagesMarkdownAsync(pdf)
// ...
}
```
## Types
```typescript
+6 -6
View File
@@ -8,12 +8,12 @@
"@napi-rs/cli": "^3.4.1",
},
"optionalDependencies": {
"@firecrawl/pdf-inspector-darwin-arm64": "1.12.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.12.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.12.0",
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.12.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.12.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.12.0",
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0",
},
},
},
+7 -7
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.12.0",
"version": "1.14.0",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
@@ -52,11 +52,11 @@
"@napi-rs/cli": "^3.4.1"
},
"optionalDependencies": {
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.12.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.12.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.12.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.12.0",
"@firecrawl/pdf-inspector-darwin-arm64": "1.12.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.12.0"
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0"
}
}
+236 -41
View File
@@ -89,6 +89,11 @@ pub struct TextItem {
pub item_type: ItemType,
/// URL for link items, `None` for other types.
pub link_url: Option<String>,
/// Marked Content ID from the content stream's BDC/BMC operator, `None`
/// when the text is not part of marked content. Join with the
/// `page`/`mcid` pairs from [`extractStructureElements`] to attach
/// structure-tree roles (headings, paragraphs, …) in tagged PDFs.
pub mcid: Option<i64>,
}
/// A page's regions for text extraction: (page_index_0based, bboxes).
@@ -153,9 +158,7 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
}
}
fn to_napi_page_ocr_reasons(
reasons: Vec<pdf_inspector::PageOcrReasons>,
) -> Vec<PageOcrReasons> {
fn to_napi_page_ocr_reasons(reasons: Vec<pdf_inspector::PageOcrReasons>) -> Vec<PageOcrReasons> {
reasons
.into_iter()
.map(|reason| PageOcrReasons {
@@ -202,6 +205,31 @@ where
}
}
// ---------------------------------------------------------------------------
// Shared implementations (single body behind sync and async entry points)
// ---------------------------------------------------------------------------
fn process_pdf_impl(bytes: &[u8], pages: Option<Vec<u32>>) -> Result<PdfResult> {
let mut opts = pdf_inspector::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = pdf_inspector::process_pdf_mem_with_options(bytes, opts)
.map_err(|e| to_napi_err(e, "process_pdf"))?;
Ok(to_napi_result(result))
}
fn classify_pdf_impl(bytes: &[u8]) -> Result<PdfClassification> {
let result =
pdf_inspector::classify_pdf_mem(bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
Ok(PdfClassification {
pdf_type: convert_pdf_type(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
}
// ---------------------------------------------------------------------------
// Public NAPI API
// ---------------------------------------------------------------------------
@@ -210,15 +238,7 @@ where
#[napi]
pub fn process_pdf(buffer: Buffer, pages: Option<Vec<u32>>) -> Result<PdfResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("process_pdf", move || {
let mut opts = pdf_inspector::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = pdf_inspector::process_pdf_mem_with_options(&bytes, opts)
.map_err(|e| to_napi_err(e, "process_pdf"))?;
Ok(to_napi_result(result))
})
catch_panic("process_pdf", move || process_pdf_impl(&bytes, pages))
}
/// Fast detection only — no text extraction or markdown.
@@ -238,16 +258,7 @@ pub fn detect_pdf(buffer: Buffer) -> Result<PdfResult> {
#[napi]
pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("classify_pdf", move || {
let result =
pdf_inspector::classify_pdf_mem(&bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
Ok(PdfClassification {
pdf_type: convert_pdf_type(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
})
catch_panic("classify_pdf", move || classify_pdf_impl(&bytes))
}
/// Extract plain text from a PDF Buffer.
@@ -300,12 +311,61 @@ pub fn extract_text_with_positions(
is_strikeout: item.is_strikeout,
item_type,
link_url,
mcid: item.mcid,
}
})
.collect())
})
}
/// One structure-tree element reference from a tagged PDF.
#[napi(object)]
pub struct StructureElementJs {
/// 1-indexed page number (matches `TextItem.page`).
pub page: u32,
/// Marked Content ID from the page's content stream (matches
/// `TextItem.mcid`).
pub mcid: i64,
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
/// Custom tags are resolved through the document's role map; tags with
/// no standard mapping are returned verbatim.
pub role: String,
}
/// Extract structure-tree element references from a tagged PDF.
///
/// Parses the document's structure tree (when present) and returns one
/// entry per marked-content reference, resolved to its 1-indexed page,
/// MCID, and structure type name. Returns an empty array when the PDF is
/// not tagged.
///
/// Join `(page, mcid)` against the `page`/`mcid` fields from
/// [`extractTextWithPositions`] to attach heading levels (H1..H6) and other
/// semantic roles to extracted text.
///
/// Pass 1-indexed page numbers (matching `TextItem.page`) to restrict
/// output; omit `pages` for the whole document. Entries are sorted by
/// `(page, mcid)`.
#[napi]
pub fn extract_structure_elements(
buffer: Buffer,
pages: Option<Vec<u32>>,
) -> Result<Vec<StructureElementJs>> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_structure_elements", move || {
let elements = pdf_inspector::extract_structure_elements_mem(&bytes, pages.as_deref())
.map_err(|e| to_napi_err(e, "extract_structure_elements"))?;
Ok(elements
.into_iter()
.map(|e| StructureElementJs {
page: e.page,
mcid: e.mcid,
role: e.role,
})
.collect())
})
}
/// Extract text within bounding-box regions from a PDF.
///
/// For hybrid OCR: layout model detects regions in rendered images,
@@ -633,25 +693,32 @@ pub fn extract_pages_markdown(
) -> Result<PagesExtractionResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_pages_markdown", move || {
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, pages.as_deref())
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
.pages
.into_iter()
.map(|r| PageMarkdownResult {
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
is_complex: result.is_complex,
})
extract_pages_markdown_impl(&bytes, pages.as_deref())
})
}
fn extract_pages_markdown_impl(
bytes: &[u8],
pages: Option<&[u32]>,
) -> Result<PagesExtractionResult> {
let result = pdf_inspector::extract_pages_markdown_mem(bytes, pages)
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
.pages
.into_iter()
.map(|r| PageMarkdownResult {
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
is_complex: result.is_complex,
})
}
@@ -692,3 +759,131 @@ fn to_page_region_texts(results: Vec<pdf_inspector::PageRegionResult>) -> Vec<Pa
})
.collect()
}
// ---------------------------------------------------------------------------
// Async variants (libuv thread pool via AsyncTask)
//
// The synchronous exports above parse on the calling thread, which in Node is
// the event loop. These `*Async` variants run the same shared implementations
// on the libuv thread pool and hand JavaScript a promise, so servers under
// concurrent load keep answering requests while a document parses. The sync
// exports keep their names, signatures, and behaviour.
//
// Each factory copies the input Buffer to an owned `Vec<u8>` on the calling
// (JS) thread — deliberately. JS execution is single-threaded, so no JS code
// can mutate the buffer while the synchronous part of the call copies it.
// Holding the napi `Buffer` and reading it from the worker instead would be
// zero-copy, but a caller mutating the buffer before the promise settles
// would then race the worker's reads — undefined behavior, not a recoverable
// error (a known napi-rs soundness hazard with cross-thread Buffer access).
// The copy is a one-time memcpy, negligible next to the parse it unblocks.
// ---------------------------------------------------------------------------
pub struct ProcessPdfTask {
bytes: Vec<u8>,
pages: Option<Vec<u32>>,
}
impl Task for ProcessPdfTask {
type Output = PdfResult;
type JsValue = PdfResult;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
let pages = self.pages.take();
// AssertUnwindSafe: `bytes`/`pages` are moved into the closure and
// dropped on unwind — no shared state can be observed broken.
catch_panic(
"process_pdf",
panic::AssertUnwindSafe(move || process_pdf_impl(&bytes, pages)),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`processPdf`]: same result, but the parse runs on the
/// libuv thread pool instead of the event loop and the call returns a
/// promise. The buffer is copied before the call returns, so it may be
/// reused or mutated immediately.
// ts_return_type is required: napi-rs emits `Promise<unknown>` for
// `AsyncTask<T>` returns without it.
#[napi(ts_return_type = "Promise<PdfResult>")]
pub fn process_pdf_async(buffer: Buffer, pages: Option<Vec<u32>>) -> AsyncTask<ProcessPdfTask> {
AsyncTask::new(ProcessPdfTask {
bytes: buffer.to_vec(),
pages,
})
}
pub struct ClassifyPdfTask {
bytes: Vec<u8>,
}
impl Task for ClassifyPdfTask {
type Output = PdfClassification;
type JsValue = PdfClassification;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
catch_panic(
"classify_pdf",
panic::AssertUnwindSafe(move || classify_pdf_impl(&bytes)),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`classifyPdf`]: same result, but the classification runs
/// on the libuv thread pool instead of the event loop and the call returns a
/// promise. The buffer is copied before the call returns, so it may be
/// reused or mutated immediately.
#[napi(ts_return_type = "Promise<PdfClassification>")]
pub fn classify_pdf_async(buffer: Buffer) -> AsyncTask<ClassifyPdfTask> {
AsyncTask::new(ClassifyPdfTask {
bytes: buffer.to_vec(),
})
}
pub struct ExtractPagesMarkdownTask {
bytes: Vec<u8>,
pages: Option<Vec<u32>>,
}
impl Task for ExtractPagesMarkdownTask {
type Output = PagesExtractionResult;
type JsValue = PagesExtractionResult;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
let pages = self.pages.take();
catch_panic(
"extract_pages_markdown",
panic::AssertUnwindSafe(move || extract_pages_markdown_impl(&bytes, pages.as_deref())),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`extractPagesMarkdown`]: same result, but the extraction
/// runs on the libuv thread pool instead of the event loop and the call
/// returns a promise. The buffer is copied before the call returns, so it
/// may be reused or mutated immediately.
#[napi(ts_return_type = "Promise<PagesExtractionResult>")]
pub fn extract_pages_markdown_async(
buffer: Buffer,
pages: Option<Vec<u32>>,
) -> AsyncTask<ExtractPagesMarkdownTask> {
AsyncTask::new(ExtractPagesMarkdownTask {
bytes: buffer.to_vec(),
pages,
})
}
+110
View File
@@ -2,16 +2,21 @@ import { readFileSync } from 'fs';
import { strict as assert } from 'assert';
import {
processPdf,
processPdfAsync,
detectPdf,
classifyPdf,
classifyPdfAsync,
extractText,
extractTextWithPositions,
extractStructureElements,
extractTextInRegions,
detectVectorGridInRegion,
extractPagesMarkdown,
extractPagesMarkdownAsync,
} from './index.js';
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
const taggedFixture = readFileSync('../tests/fixtures/firecrawl_docs_tagged.pdf');
// --- processPdf ---
console.log('Testing processPdf...');
@@ -79,6 +84,46 @@ assert.ok(page1Items.length > 0);
assert.ok(page1Items.every(i => i.page === 1));
console.log(' extractTextWithPositions with pages: OK');
// mcid: undefined on untagged PDFs, numeric on tagged marked content
assert.ok(items.every(i => i.mcid === undefined || typeof i.mcid === 'number'));
const taggedItems = extractTextWithPositions(taggedFixture);
assert.ok(
taggedItems.some(i => typeof i.mcid === 'number'),
'tagged PDF text items should carry Marked Content IDs',
);
console.log(' extractTextWithPositions mcid: OK');
// --- extractStructureElements ---
console.log('Testing extractStructureElements...');
const structureElements = extractStructureElements(taggedFixture);
assert.ok(structureElements.length > 0);
assert.ok(structureElements.every(e => typeof e.page === 'number'));
assert.ok(structureElements.every(e => typeof e.mcid === 'number'));
assert.ok(structureElements.every(e => typeof e.role === 'string' && e.role.length > 0));
assert.ok(
structureElements.some(e => e.role === 'H1'),
'tagged fixture should surface H1 heading roles',
);
// (page, mcid) joins against extractTextWithPositions to recover heading text
const h1Refs = new Set(
structureElements.filter(e => e.role === 'H1').map(e => `${e.page}:${e.mcid}`),
);
const h1Text = taggedItems
.filter(i => typeof i.mcid === 'number' && h1Refs.has(`${i.page}:${i.mcid}`))
.map(i => i.text)
.join('');
assert.ok(h1Text.trim().length > 0, 'H1 join should recover heading text');
// pages filter is 1-indexed, matching TextItem.page
const page1Elements = extractStructureElements(taggedFixture, [1]);
assert.ok(page1Elements.length > 0);
assert.ok(page1Elements.every(e => e.page === 1));
// untagged PDFs yield an empty array
assert.deepEqual(extractStructureElements(fixture), []);
console.log(' extractStructureElements: OK');
// --- extractTextInRegions ---
console.log('Testing extractTextInRegions...');
const regionResults = extractTextInRegions(fixture, [
@@ -124,10 +169,75 @@ assert.equal(picked.pages[0].page, 2);
assert.equal(picked.pages[1].page, 0);
console.log(' extractPagesMarkdown with pages: OK');
// --- Async variants ---
console.log('Testing async variants...');
// processPdfAsync returns a promise and matches the sync result
const asyncResultPromise = processPdfAsync(fixture);
assert.ok(asyncResultPromise instanceof Promise);
const asyncResult = await asyncResultPromise;
assert.equal(asyncResult.pdfType, result.pdfType);
assert.equal(asyncResult.pageCount, result.pageCount);
assert.equal(asyncResult.markdown, result.markdown);
console.log(' processPdfAsync: OK');
// processPdfAsync with pages
const asyncResult2 = await processPdfAsync(fixture, [1]);
assert.equal(asyncResult2.markdown, result2.markdown);
console.log(' processPdfAsync with pages: OK');
// classifyPdfAsync matches the sync result
const asyncClassified = await classifyPdfAsync(fixture);
assert.equal(asyncClassified.pdfType, classified.pdfType);
assert.equal(asyncClassified.pageCount, classified.pageCount);
assert.equal(asyncClassified.confidence, classified.confidence);
assert.deepEqual(asyncClassified.pagesNeedingOcr, classified.pagesNeedingOcr);
console.log(' classifyPdfAsync: OK');
// extractPagesMarkdownAsync matches the sync result
const asyncAllPages = await extractPagesMarkdownAsync(fixture);
assert.equal(asyncAllPages.pages.length, allPages.pages.length);
assert.deepEqual(
asyncAllPages.pages.map(p => p.markdown),
allPages.pages.map(p => p.markdown),
);
assert.equal(asyncAllPages.isComplex, allPages.isComplex);
console.log(' extractPagesMarkdownAsync: OK');
// selected pages preserve caller order
const asyncPicked = await extractPagesMarkdownAsync(fixture, [2, 0]);
assert.equal(asyncPicked.pages.length, 2);
assert.equal(asyncPicked.pages[0].page, 2);
assert.equal(asyncPicked.pages[1].page, 0);
console.log(' extractPagesMarkdownAsync with pages: OK');
// input buffer is copied at call time: mutating it immediately after the
// call must not affect the in-flight parse
const scratch = Buffer.from(fixture);
const inFlight = processPdfAsync(scratch);
scratch.fill(0);
const fromMutated = await inFlight;
assert.equal(fromMutated.markdown, result.markdown);
console.log(' processPdfAsync input copied at call time: OK');
// concurrent async calls all settle
const [c1, c2, c3] = await Promise.all([
processPdfAsync(fixture),
classifyPdfAsync(fixture),
extractPagesMarkdownAsync(fixture),
]);
assert.equal(c1.pdfType, 'TextBased');
assert.equal(c2.pdfType, 'TextBased');
assert.equal(c3.pages.length, 3);
console.log(' concurrent async calls: OK');
// --- Error handling ---
console.log('Testing error handling...');
assert.throws(() => processPdf(Buffer.from('not a pdf')), /process_pdf/);
assert.throws(() => classifyPdf(Buffer.from('')), /classify_pdf/);
await assert.rejects(processPdfAsync(Buffer.from('not a pdf')), /process_pdf/);
await assert.rejects(classifyPdfAsync(Buffer.from('')), /classify_pdf/);
await assert.rejects(extractPagesMarkdownAsync(Buffer.from('')), /extract_pages_markdown/);
console.log(' error handling: OK');
console.log('\nAll NAPI tests passed!');
+51
View File
@@ -10,6 +10,9 @@ class PdfResult:
page_count: int
processing_time_ms: int
pages_needing_ocr: list[int]
"""1-indexed page numbers that need OCR."""
ocr_reasons_by_page: list["PageOcrReasons"]
"""Machine-readable OCR reasons by 1-indexed page."""
title: Optional[str]
confidence: float
is_complex_layout: bool
@@ -17,6 +20,13 @@ class PdfResult:
pages_with_columns: list[int]
has_encoding_issues: bool
class PageOcrReasons:
"""OCR reasons for a single 1-indexed page."""
page: int
"""1-indexed page number."""
reasons: list[str]
"""Machine-readable OCR reason identifiers."""
class PdfClassification:
"""Lightweight PDF classification result."""
pdf_type: str
@@ -41,12 +51,28 @@ class TextItem:
is_underline: bool
is_strikeout: bool
item_type: str
mcid: Optional[int]
"""Marked Content ID from the content stream's BDC/BMC operator, None when
the text is not part of marked content. Join with the (page, mcid) pairs
from extract_structure_elements to attach structure-tree roles in tagged
PDFs."""
class StructureElement:
"""One structure-tree element reference from a tagged PDF."""
page: int
"""1-indexed page number (matches TextItem.page)."""
mcid: int
"""Marked Content ID from the page's content stream (matches TextItem.mcid)."""
role: str
"""Standard structure type name ("H1".."H6", "P", "Table", "TD", ...)."""
class RegionText:
"""Extracted text for a single region."""
text: str
needs_ocr: bool
"""True when the text should not be trusted."""
ocr_reason: Optional[str]
"""Machine-readable OCR reason when the cause is known."""
class PageRegionTexts:
"""Extracted text for one page's regions."""
@@ -62,6 +88,8 @@ class PageMarkdown:
"""Formatted markdown for this page (empty string when needs_ocr is True)."""
needs_ocr: bool
"""True when text on this page is unreliable and OCR should be used instead."""
ocr_reason: Optional[str]
"""Machine-readable OCR reason when the cause is known."""
class PagesExtractionResult:
"""Per-page markdown output with document-wide layout classification."""
@@ -73,6 +101,8 @@ class PagesExtractionResult:
"""1-indexed pages where multi-column layout was detected."""
pages_needing_ocr: list[int]
"""1-indexed pages that need OCR."""
ocr_reasons_by_page: list[PageOcrReasons]
"""Machine-readable OCR reasons by 1-indexed page."""
is_complex: bool
"""True if any page has tables or multi-column layout."""
@@ -116,6 +146,27 @@ def extract_text_with_positions_bytes(data: bytes, pages: Optional[list[int]] =
"""Extract text with position information from bytes."""
...
def extract_structure_elements(path: str, pages: Optional[list[int]] = None) -> list[StructureElement]:
"""Extract structure-tree element references from a tagged PDF file.
Returns one entry per marked-content reference, resolved to its 1-indexed
page, MCID, and structure type name ("H1".."H6", "P", "Table", ...), sorted
by (page, mcid). Returns an empty list when the PDF is not tagged.
Args:
path: Path to the PDF file.
pages: Optional list of 1-indexed pages (matching ``TextItem.page``).
When ``None`` (default), the whole document is returned.
"""
...
def extract_structure_elements_bytes(data: bytes, pages: Optional[list[int]] = None) -> list[StructureElement]:
"""Extract structure-tree element references from tagged PDF bytes.
See :func:`extract_structure_elements` for details.
"""
...
def extract_text_in_regions(
path: str,
page_regions: list[tuple[int, list[list[float]]]],
+3 -3
View File
@@ -4,9 +4,9 @@ build-backend = "maturin"
[project]
name = "pdf-inspector"
# Bump this to publish to PyPI — CI publishes automatically when the version
# changes on main (same flow as napi/package.json for npm).
version = "0.2.6"
# Keep package versions in sync with `python3 scripts/version.py <version>`.
# CI publishes automatically when the synchronized change lands on main.
version = "1.14.0"
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
readme = "docs/python.md"
license = { text = "MIT" }
+111
View File
@@ -0,0 +1,111 @@
import json
import sys
import tempfile
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from version import PLATFORM_PACKAGES, check_versions, set_versions
class VersionTests(unittest.TestCase):
def setUp(self):
self.temporary = tempfile.TemporaryDirectory()
self.root = Path(self.temporary.name)
(self.root / "napi").mkdir()
(self.root / "site").mkdir()
(self.root / "wasm").mkdir()
self._write_manifest("Cargo.toml", "package", "0.1.0")
self._write_manifest("pyproject.toml", "project", "0.1.0")
self._write_manifest("napi/Cargo.toml", "package", "0.1.0")
self._write_manifest("wasm/Cargo.toml", "package", "0.1.0")
package = {
"name": "@firecrawl/pdf-inspector",
"version": "0.1.0",
"optionalDependencies": {
dependency: "0.1.0" for dependency in PLATFORM_PACKAGES
},
}
(self.root / "napi/package.json").write_text(
json.dumps(package), encoding="utf-8"
)
(self.root / "napi/bun.lock").write_text(
"\n".join(
f' "{dependency}": "0.1.0",'
for dependency in PLATFORM_PACKAGES
)
+ "\n",
encoding="utf-8",
)
(self.root / "site/index.html").write_text(
'https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.0/'
'pdf_inspector_wasm.js\n',
encoding="utf-8",
)
self._write_lock(
"napi/Cargo.lock", ("pdf-inspector", "pdf-inspector-napi")
)
self._write_lock(
"wasm/Cargo.lock", ("pdf-inspector", "pdf-inspector-wasm")
)
def tearDown(self):
self.temporary.cleanup()
def _write_manifest(self, relative, section, version):
(self.root / relative).write_text(
f'[{section}]\nname = "fixture"\nversion = "{version}"\n',
encoding="utf-8",
)
def _write_lock(self, relative, packages):
content = "\n".join(
f'[[package]]\nname = "{package}"\nversion = "0.1.0"\n'
for package in packages
)
(self.root / relative).write_text(content, encoding="utf-8")
def test_updates_every_version_location(self):
set_versions("1.14.0", self.root)
self.assertEqual(check_versions(self.root), "1.14.0")
def test_reports_a_divergent_package(self):
self._write_manifest("wasm/Cargo.toml", "package", "0.2.0")
with self.assertRaisesRegex(ValueError, "WASM package: 0.2.0"):
check_versions(self.root)
def test_rejects_an_invalid_version(self):
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
set_versions("next", self.root)
def test_rejects_numeric_prerelease_with_leading_zero(self):
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
set_versions("1.2.3-01", self.root)
self.assertEqual(
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
)
def test_preflight_failure_does_not_partially_update(self):
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
(self.root / "site/index.html").write_text(
"missing module URL\n", encoding="utf-8"
)
with self.assertRaisesRegex(ValueError, "Missing pinned WASM package URL"):
set_versions("1.14.0", self.root)
self.assertEqual(
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
)
if __name__ == "__main__":
unittest.main()
+262
View File
@@ -0,0 +1,262 @@
#!/usr/bin/env python3
"""Keep every pdf-inspector package on one release version."""
from __future__ import annotations
import argparse
import json
import re
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
PRERELEASE_IDENTIFIER = (
r"(?:0|[1-9]\d*|[0-9A-Za-z-]*[A-Za-z-][0-9A-Za-z-]*)"
)
SEMVER = re.compile(
r"^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)"
rf"(?:-{PRERELEASE_IDENTIFIER}(?:\.{PRERELEASE_IDENTIFIER})*)?"
r"(?:\+[0-9A-Za-z-]+(?:\.[0-9A-Za-z-]+)*)?$"
)
VERSION_LINE = re.compile(r'^(\s*version\s*=\s*")[^"]+(".*)$')
SECTION_LINE = re.compile(r"^\s*\[([^]]+)]\s*$")
PLATFORM_PACKAGES = (
"@firecrawl/pdf-inspector-linux-x64-gnu",
"@firecrawl/pdf-inspector-linux-x64-musl",
"@firecrawl/pdf-inspector-linux-arm64-gnu",
"@firecrawl/pdf-inspector-linux-arm64-musl",
"@firecrawl/pdf-inspector-darwin-arm64",
"@firecrawl/pdf-inspector-win32-x64-msvc",
)
TOML_VERSIONS = (
("Rust crate", Path("Cargo.toml"), "package"),
("Python package", Path("pyproject.toml"), "project"),
("NAPI crate", Path("napi/Cargo.toml"), "package"),
("WASM package", Path("wasm/Cargo.toml"), "package"),
)
LOCK_VERSIONS = (
("NAPI lock: core", Path("napi/Cargo.lock"), "pdf-inspector"),
("NAPI lock: binding", Path("napi/Cargo.lock"), "pdf-inspector-napi"),
("WASM lock: core", Path("wasm/Cargo.lock"), "pdf-inspector"),
("WASM lock: binding", Path("wasm/Cargo.lock"), "pdf-inspector-wasm"),
)
SITE_WASM_VERSION = re.compile(
r"(@firecrawl/pdf-inspector-wasm@)([^/\"]+)(/pdf_inspector_wasm\.js)"
)
def _read_section_version(path: Path, section: str) -> str:
active = False
for line in path.read_text(encoding="utf-8").splitlines():
section_match = SECTION_LINE.match(line)
if section_match:
active = section_match.group(1) == section
elif active:
version_match = VERSION_LINE.match(line)
if version_match:
return line.split('"', 2)[1]
raise ValueError(f"No version found in [{section}] of {path}")
def _write_section_version(path: Path, section: str, version: str) -> None:
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
active = False
for index, line in enumerate(lines):
section_match = SECTION_LINE.match(line)
if section_match:
active = section_match.group(1) == section
elif active:
version_match = VERSION_LINE.match(line)
if version_match:
newline = "\n" if line.endswith("\n") else ""
replacement = (
f"{version_match.group(1)}{version}"
f"{version_match.group(2).rstrip()}"
)
lines[index] = (
f"{replacement}{newline}"
)
path.write_text("".join(lines), encoding="utf-8")
return
raise ValueError(f"No version found in [{section}] of {path}")
def _package_block(lines: list[str], package: str) -> tuple[int, int]:
for start, line in enumerate(lines):
if line.strip() != "[[package]]":
continue
end = next(
(
index
for index in range(start + 1, len(lines))
if lines[index].strip() == "[[package]]"
),
len(lines),
)
if any(line.strip() == f'name = "{package}"' for line in lines[start:end]):
return start, end
raise ValueError(f"No lockfile entry found for {package}")
def _read_lock_version(path: Path, package: str) -> str:
lines = path.read_text(encoding="utf-8").splitlines()
start, end = _package_block(lines, package)
for line in lines[start:end]:
version_match = VERSION_LINE.match(line)
if version_match:
return line.split('"', 2)[1]
raise ValueError(f"No version found for {package} in {path}")
def _write_lock_version(path: Path, package: str, version: str) -> None:
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
start, end = _package_block(lines, package)
for index in range(start, end):
version_match = VERSION_LINE.match(lines[index])
if version_match:
newline = "\n" if lines[index].endswith("\n") else ""
lines[index] = (
f'{version_match.group(1)}{version}{version_match.group(2).rstrip()}'
f"{newline}"
)
path.write_text("".join(lines), encoding="utf-8")
return
raise ValueError(f"No version found for {package} in {path}")
def _node_versions(root: Path) -> dict[str, str]:
package = json.loads((root / "napi/package.json").read_text(encoding="utf-8"))
versions = {"Node package": package["version"]}
optional = package.get("optionalDependencies", {})
for dependency in PLATFORM_PACKAGES:
if dependency not in optional:
raise ValueError(f"Missing Node optional dependency: {dependency}")
versions[f"Node optional dependency: {dependency}"] = optional[dependency]
return versions
def _bun_versions(root: Path) -> dict[str, str]:
text = (root / "napi/bun.lock").read_text(encoding="utf-8")
versions = {}
for dependency in PLATFORM_PACKAGES:
match = re.search(
rf'"{re.escape(dependency)}": "([^"]+)"[,]', text
)
if not match:
raise ValueError(f"Missing Bun lock dependency: {dependency}")
versions[f"Bun lock: {dependency}"] = match.group(1)
return versions
def _site_wasm_version(root: Path) -> str:
text = (root / "site/index.html").read_text(encoding="utf-8")
match = SITE_WASM_VERSION.search(text)
if not match:
raise ValueError("Missing pinned WASM package URL in site/index.html")
return match.group(2)
def package_versions(root: Path = ROOT) -> dict[str, str]:
versions = {
label: _read_section_version(root / relative, section)
for label, relative, section in TOML_VERSIONS
}
versions.update(_node_versions(root))
versions.update(_bun_versions(root))
versions["Website WASM module"] = _site_wasm_version(root)
versions.update(
{
label: _read_lock_version(root / relative, package)
for label, relative, package in LOCK_VERSIONS
}
)
return versions
def check_versions(root: Path = ROOT) -> str:
versions = package_versions(root)
expected = versions["Rust crate"]
if not SEMVER.fullmatch(expected):
raise ValueError(f"Rust crate has an invalid semantic version: {expected}")
mismatches = {
label: version for label, version in versions.items() if version != expected
}
if mismatches:
details = "\n".join(
f" - {label}: {version}" for label, version in mismatches.items()
)
raise ValueError(f"Expected every package to use {expected}:\n{details}")
return expected
def set_versions(version: str, root: Path = ROOT) -> None:
if not SEMVER.fullmatch(version):
raise ValueError(f"Invalid semantic version: {version}")
# Validate every expected location before writing the first file. This
# prevents a stale manifest or generated file from leaving a partial bump.
package_versions(root)
for _, relative, section in TOML_VERSIONS:
_write_section_version(root / relative, section, version)
package_path = root / "napi/package.json"
package = json.loads(package_path.read_text(encoding="utf-8"))
package["version"] = version
optional = package.get("optionalDependencies", {})
for dependency in PLATFORM_PACKAGES:
if dependency not in optional:
raise ValueError(f"Missing Node optional dependency: {dependency}")
optional[dependency] = version
package_path.write_text(json.dumps(package, indent=2) + "\n", encoding="utf-8")
bun_path = root / "napi/bun.lock"
bun_text = bun_path.read_text(encoding="utf-8")
for dependency in PLATFORM_PACKAGES:
pattern = rf'("{re.escape(dependency)}": ")[^"]+("[,])'
bun_text, count = re.subn(
pattern, rf"\g<1>{version}\g<2>", bun_text, count=1
)
if count != 1:
raise ValueError(f"Missing Bun lock dependency: {dependency}")
bun_path.write_text(bun_text, encoding="utf-8")
site_path = root / "site/index.html"
site_text = site_path.read_text(encoding="utf-8")
site_text, count = SITE_WASM_VERSION.subn(
rf"\g<1>{version}\g<3>", site_text, count=1
)
if count != 1:
raise ValueError("Missing pinned WASM package URL in site/index.html")
site_path.write_text(site_text, encoding="utf-8")
for _, relative, package_name in LOCK_VERSIONS:
_write_lock_version(root / relative, package_name, version)
check_versions(root)
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("version", nargs="?", help="new shared semantic version")
parser.add_argument(
"--check", action="store_true", help="fail if package versions have diverged"
)
arguments = parser.parse_args()
if arguments.check == bool(arguments.version):
parser.error("provide either a version or --check")
try:
if arguments.check:
version = check_versions()
print(f"All packages use {version}")
else:
set_versions(arguments.version)
print(f"Updated all packages to {arguments.version}")
except ValueError as error:
parser.exit(1, f"{error}\n")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+8 -8
View File
@@ -858,22 +858,22 @@
<p>Evaluated on the <a class="text-link" href="https://github.com/opendataloader-project/opendataloader-bench">opendataloader-bench</a> corpus of 200 PDFs. This comparison covers local engines without model-based PDF parsing, with OCR disabled. Higher scores are better.</p>
</div>
<div class="benchmark-card">
<div class="benchmark-top"><span><strong>200 PDFs</strong> · OpenDataLoader benchmark</span><span>Apple M4 Pro · median of 3 runs</span></div>
<div class="benchmark-top"><span><strong>200 PDFs</strong> · OpenDataLoader benchmark</span><span>Apple M4 Pro · median of 5 runs</span></div>
<div class="table-scroll">
<table aria-label="PDF extraction benchmark results">
<thead>
<tr><th>Engine</th><th>Overall</th><th>Reading order</th><th>Tables</th><th>Headings</th><th>Complete run</th></tr>
</thead>
<tbody>
<tr class="highlight"><td>pdf-inspector</td><td>0.875</td><td>0.915</td><td>0.814</td><td>0.788</td><td>2.8s</td></tr>
<tr><td>LiteParse</td><td>0.870</td><td>0.908</td><td>0.693</td><td>0.811</td><td>13.9s</td></tr>
<tr><td>OpenDataLoader</td><td>0.843</td><td>0.912</td><td>0.489</td><td>0.760</td><td>9.8s</td></tr>
<tr><td>PyMuPDF4LLM</td><td>0.735</td><td>0.886</td><td>0.401</td><td>0.424</td><td>15.5s</td></tr>
<tr><td>MarkItDown</td><td>0.583</td><td>0.879</td><td>0.000</td><td>0.000</td><td>6.7s</td></tr>
<tr class="highlight"><td>pdf-inspector</td><td>0.875</td><td>0.915</td><td>0.814</td><td>0.788</td><td>0.470s</td></tr>
<tr><td>LiteParse</td><td>0.873</td><td>0.913</td><td>0.693</td><td>0.811</td><td>0.750s</td></tr>
<tr><td>OpenDataLoader</td><td>0.831</td><td>0.902</td><td>0.489</td><td>0.739</td><td>2.569s</td></tr>
<tr><td>PyMuPDF4LLM</td><td>0.735</td><td>0.886</td><td>0.401</td><td>0.424</td><td>17.117s</td></tr>
<tr><td>MarkItDown</td><td>0.589</td><td>0.844</td><td>0.273</td><td>0.000</td><td>16.165s</td></tr>
</tbody>
</table>
</div>
<div class="benchmark-note">Refreshed July 16, 2026. Scores use the benchmarks NID, TEDS, and MHS evaluators.</div>
<div class="benchmark-note">Refreshed July 31, 2026. Scores use the benchmarks NID, TEDS, and MHS evaluators; speed is the median of five alternating or rotating complete corpus runs after an excluded warm-up. <a class="text-link" href="https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results">Versions and raw artifacts</a>.</div>
</div>
<div class="best-fit">
<strong>Best fit</strong>
@@ -975,7 +975,7 @@ result = pdf_inspector.<span class="fn">process_pdf</span>(<span class="str">"do
<script>
(() => {
const MAX_FILE_SIZE = 25 * 1024 * 1024;
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.1/pdf_inspector_wasm.js";
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@1.14.0/pdf_inspector_wasm.js";
const input = document.querySelector("#pdf-input");
const dropZone = document.querySelector("#drop-zone");
const filePanel = document.querySelector("#demo-file");
+32 -5
View File
@@ -2,8 +2,8 @@
use pdf_inspector::extractor::ItemType;
use pdf_inspector::{
extract_text_with_positions_pages, process_pdf_with_options, LayoutComplexity, PdfOptions,
PdfType, ProcessMode, TextItem,
extract_text_with_positions_pages_with_password, process_pdf_with_options, LayoutComplexity,
PdfOptions, PdfType, ProcessMode, TextItem,
};
use std::collections::HashSet;
use std::env;
@@ -103,9 +103,18 @@ fn format_items_json(items: &[TextItem]) -> String {
)
}
fn extract_items_json(
pdf_path: &str,
page_filter: Option<&HashSet<u32>>,
password: Option<&str>,
) -> Result<String, pdf_inspector::PdfError> {
extract_text_with_positions_pages_with_password(pdf_path, page_filter, password)
.map(|items| format_items_json(&items))
}
#[cfg(test)]
mod tests {
use super::format_items_json;
use super::{extract_items_json, format_items_json};
use pdf_inspector::extractor::ItemType;
use pdf_inspector::TextItem;
@@ -137,6 +146,24 @@ mod tests {
assert!(json.contains(r#""item_type":"text""#));
assert!(json.contains(r#""mcid":7"#));
}
#[test]
fn items_json_uses_supplied_pdf_password() {
let path = "tests/fixtures/encrypted-secret123.pdf";
let without_password = extract_items_json(path, None, None);
assert!(
without_password.is_err(),
"encrypted fixture unexpectedly extracted without a password"
);
let json = extract_items_json(path, None, Some("secret123"))
.expect("correct password should decrypt positioned text");
assert!(
json.contains("Procurement"),
"decrypted item JSON should contain fixture text, got {json}"
);
}
}
/// Parse a page specification like "1,3,5-10,20" into a HashSet of page numbers.
@@ -257,8 +284,8 @@ fn main() {
});
if items_json_output {
match extract_text_with_positions_pages(pdf_path, page_filter.as_ref()) {
Ok(items) => println!("{}", format_items_json(&items)),
match extract_items_json(pdf_path, page_filter.as_ref(), password.as_deref()) {
Ok(json) => println!("{}", json),
Err(e) => {
println!(r#"{{"error":"{}"}}"#, json_escape(&e.to_string()));
process::exit(1);
+65 -1
View File
@@ -1659,7 +1659,15 @@ fn hex_val(b: u8) -> Option<u8> {
/// Standard page: 612x792 points (US Letter) = ~485,000 sq points
/// At 2x resolution that's ~1.9M pixels, so we use 250K pixels as threshold
/// (accounting for varying DPI and page sizes)
fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
/// Returns `(has_images, total_image_area, has_template_image)` for a page.
/// `has_template_image` means a single large (>50% page coverage)
/// background image — the signal `classify_pdf`/`detect_pdf_type` uses to
/// route a page to OCR regardless of any incidental native text drawn over
/// it. Exposed at crate visibility so extraction-side per-page `needs_ocr`
/// computation (`extract_pages_markdown_mem`) can consult the same signal
/// instead of maintaining its own, independent notion of "needs OCR" that
/// can silently disagree with detection — see #227.
pub(crate) fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
// Threshold: image covering roughly half a page at 150+ DPI
// 612 * 792 / 2 * (150/72)^2 ≈ 1M pixels, but we'll be conservative
const TEMPLATE_IMAGE_THRESHOLD: u64 = 500_000; // 500K pixels
@@ -1741,6 +1749,62 @@ fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
(has_images, total_area, has_template_image)
}
/// Computes both `(needs_ocr_for_template_image, has_vector_text)` for a
/// page from a single shared `analyze_page_content` pass — that call
/// decompresses and scans every content stream (page + XObjects) plus
/// image coverage, so `extract_pages_markdown_mem` must not invoke it
/// twice per page (once per signal) the way `detect_from_document` avoids
/// by caching its per-page `PageAnalysis`.
///
/// `needs_ocr_for_template_image` is true when a page's template image
/// should be treated as a scan needing OCR — a single full-page background
/// image with little/no real text — rather than a text page that happens
/// to carry a watermark, letterhead, or figure. Mirrors the two distinct
/// signals classification uses to route a template-image page to OCR:
///
/// 1. `looks_like_scan`: image_count <= 1, few text operators (<50), and
/// low alphanumeric diversity in raw string operands (unless decodable
/// CID/ToUnicode fonts explain that away) — the gate used for
/// `pages_with_template_images` and Mixed-type per-page routing.
/// 2. Insufficient real text volume, using `DetectionConfig::default()`'s
/// `min_text_ops_per_page` (3) — the same threshold Mixed-type per-page
/// routing applies via `text_operator_count < config.min_text_ops_per_page
/// && has_images` (simplified here since a template image implies
/// `has_images`). Deliberately *not* the higher `effective_min_ops`
/// floor (`min_text_ops_per_page.max(10)`) that whole-document
/// `PdfType::ImageBased`/`Scanned` classification uses for
/// `pages_with_text` — that's a cross-page aggregate decision this
/// per-page function has no way to replicate exactly, and the lower
/// per-page threshold is the one a single page's own signals can
/// actually agree with.
///
/// `has_vector_text` is true when a page has vector-outlined text (glyphs
/// drawn as paths rather than shown via text-showing operators) —
/// `detect_from_document`'s Mixed-type per-page routing always sends
/// these pages to OCR, independent of any template-image check, since
/// outlined glyphs can't be extracted as text at all.
///
/// Exposed at crate visibility so `extract_pages_markdown_mem` can apply
/// the same gates classification needs elsewhere instead of treating the
/// raw signals alone as sufficient — see #227/#231.
pub(crate) fn page_ocr_signals(doc: &Document, page_id: ObjectId) -> (bool, bool) {
let analysis = analyze_page_content(doc, page_id);
let needs_ocr_for_template_image = if !analysis.has_template_image {
false
} else {
let alphanum_low = analysis.unique_alphanum_chars < 10
&& !(analysis.has_decodable_text_fonts && analysis.text_operator_count >= 10);
let looks_like_scan =
analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_low;
let insufficient_text =
analysis.text_operator_count < DetectionConfig::default().min_text_ops_per_page;
looks_like_scan || insufficient_text
};
(needs_ocr_for_template_image, analysis.has_vector_text)
}
/// Recursively collect image dimensions from XObject resources,
/// including images nested inside Form XObjects.
fn collect_images_from_resources(
+649
View File
@@ -0,0 +1,649 @@
//! Built-in glyph metrics for the 14 standard PDF fonts.
//!
//! PDFs may omit `/Widths` for non-embedded base-14 fonts (Times, Helvetica,
//! Courier, Symbol, ZapfDingbats); per the PDF spec the reader must supply
//! the metrics. Without them every text item gets width 0, which breaks
//! space synthesis, sub/superscript detection, and table column detection
//! (common in 1990s dvips/Distiller output).
//!
//! Tables are generated from the Adobe Core 14 AFM files (via reportlab's
//! `_fontdata`), keyed by Unicode char, sorted for binary search.
//! Generator: scratchpad/gen_base14.py (session tooling, not checked in).
/// Width in 1000ths of an em for `c` in the given base-14 font, or `None`
/// if the font is not one of the base 14 (after name normalization) or the
/// char has no glyph in its AFM.
pub(crate) fn base14_char_width(base_font: &str, c: char) -> Option<u16> {
let table = base14_table(base_font)?;
// AFM tables key visible glyphs only; alias the invisible variants the
// cp1252 fallback can produce so they get the metric of their visible
// counterpart instead of the generic default.
let c = match c {
'\u{00A0}' => ' ', // no-break space -> space
'\u{00AD}' => '-', // soft hyphen -> hyphen
_ => c,
};
table
.binary_search_by_key(&c, |&(ch, _)| ch)
.ok()
.map(|i| table[i].1)
}
/// True when the base font name normalizes to one of the standard 14 fonts.
pub(crate) fn is_base14_font(base_font: &str) -> bool {
base14_table(base_font).is_some()
}
/// Code → Unicode through the font's BUILT-IN encoding, for the base-14
/// fonts whose repertoire is not Latin (Symbol, ZapfDingbats). Their glyphs
/// live at byte positions that have nothing to do with cp1252 (Symbol 0x61
/// renders α, Zapf 0x21 renders ✁), so advance widths must be resolved
/// through this mapping — the renderer draws these glyphs regardless of how
/// the text decoder transliterates them. Returns `None` for the Latin text
/// fonts, which follow standard single-byte encodings.
pub(crate) fn builtin_encoding_char(base_font: &str, code: u8) -> Option<char> {
let table = base14_table(base_font)?;
let enc: &[(u8, char)] = if std::ptr::eq(table, SYMBOL) {
SYMBOL_ENCODING
} else if std::ptr::eq(table, ZAPFDINGBATS) {
ZAPFDINGBATS_ENCODING
} else {
return None;
};
enc.binary_search_by_key(&code, |&(b, _)| b)
.ok()
.map(|i| enc[i].1)
}
/// Map a BaseFont name (possibly subset-prefixed, e.g. "ABCDEF+Times-Bold",
/// or a common alias like "Arial" / "TimesNewRomanPSMT") to its width table.
fn base14_table(base_font: &str) -> Option<&'static [(char, u16)]> {
// Strip subset prefix "ABCDEF+"
let name = match base_font.split_once('+') {
Some((prefix, rest))
if prefix.len() == 6 && prefix.chars().all(|c| c.is_ascii_uppercase()) =>
{
rest
}
_ => base_font,
};
let lower = name.to_ascii_lowercase();
let bold = lower.contains("bold");
let italic = lower.contains("italic") || lower.contains("oblique");
if lower.contains("courier") {
return Some(match (bold, italic) {
(false, false) => COURIER,
(true, false) => COURIER_BOLD,
(false, true) => COURIER_OBLIQUE,
(true, true) => COURIER_BOLDOBLIQUE,
});
}
if lower.contains("helvetica") || lower.contains("arial") {
return Some(match (bold, italic) {
(false, false) => HELVETICA,
(true, false) => HELVETICA_BOLD,
(false, true) => HELVETICA_OBLIQUE,
(true, true) => HELVETICA_BOLDOBLIQUE,
});
}
if lower.contains("times") {
return Some(match (bold, italic) {
(false, false) => TIMES_ROMAN,
(true, false) => TIMES_BOLD,
(false, true) => TIMES_ITALIC,
(true, true) => TIMES_BOLDITALIC,
});
}
// Symbol and ZapfDingbats have unique glyph repertoires, so only exact
// names (plus the common MT/ITC aliases) qualify — a custom font that
// merely mentions "Symbol" in its name must not get these metrics.
match lower.as_str() {
"zapfdingbats" | "dingbats" | "itczapfdingbats" | "zapfdingbatsitc" => {
return Some(ZAPFDINGBATS)
}
"symbol" | "symbolmt" | "symbolitc" => return Some(SYMBOL),
_ => {}
}
None
}
#[rustfmt::skip]
static COURIER: &[(char, u16)] = &[
(' ', 600), ('!', 600), ('"', 600), ('#', 600), ('$', 600), ('%', 600),
('&', 600), ('\'', 600), ('(', 600), (')', 600), ('*', 600), ('+', 600),
(',', 600), ('-', 600), ('.', 600), ('/', 600), ('0', 600), ('1', 600),
('2', 600), ('3', 600), ('4', 600), ('5', 600), ('6', 600), ('7', 600),
('8', 600), ('9', 600), (':', 600), (';', 600), ('<', 600), ('=', 600),
('>', 600), ('?', 600), ('@', 600), ('A', 600), ('B', 600), ('C', 600),
('D', 600), ('E', 600), ('F', 600), ('G', 600), ('H', 600), ('I', 600),
('J', 600), ('K', 600), ('L', 600), ('M', 600), ('N', 600), ('O', 600),
('P', 600), ('Q', 600), ('R', 600), ('S', 600), ('T', 600), ('U', 600),
('V', 600), ('W', 600), ('X', 600), ('Y', 600), ('Z', 600), ('[', 600),
('\\', 600), (']', 600), ('^', 600), ('_', 600), ('`', 600), ('a', 600),
('b', 600), ('c', 600), ('d', 600), ('e', 600), ('f', 600), ('g', 600),
('h', 600), ('i', 600), ('j', 600), ('k', 600), ('l', 600), ('m', 600),
('n', 600), ('o', 600), ('p', 600), ('q', 600), ('r', 600), ('s', 600),
('t', 600), ('u', 600), ('v', 600), ('w', 600), ('x', 600), ('y', 600),
('z', 600), ('{', 600), ('|', 600), ('}', 600), ('~', 600), ('\u{00A1}', 600),
('\u{00A2}', 600), ('\u{00A3}', 600), ('\u{00A4}', 600), ('\u{00A5}', 600), ('\u{00A6}', 600), ('\u{00A7}', 600),
('\u{00A8}', 600), ('\u{00A9}', 600), ('\u{00AA}', 600), ('\u{00AB}', 600), ('\u{00AC}', 600), ('\u{00AE}', 600),
('\u{00AF}', 600), ('\u{00B0}', 600), ('\u{00B1}', 600), ('\u{00B2}', 600), ('\u{00B3}', 600), ('\u{00B4}', 600),
('\u{00B5}', 600), ('\u{00B6}', 600), ('\u{00B7}', 600), ('\u{00B8}', 600), ('\u{00B9}', 600), ('\u{00BA}', 600),
('\u{00BB}', 600), ('\u{00BC}', 600), ('\u{00BD}', 600), ('\u{00BE}', 600), ('\u{00BF}', 600), ('\u{00C0}', 600),
('\u{00C1}', 600), ('\u{00C2}', 600), ('\u{00C3}', 600), ('\u{00C4}', 600), ('\u{00C5}', 600), ('\u{00C6}', 600),
('\u{00C7}', 600), ('\u{00C8}', 600), ('\u{00C9}', 600), ('\u{00CA}', 600), ('\u{00CB}', 600), ('\u{00CC}', 600),
('\u{00CD}', 600), ('\u{00CE}', 600), ('\u{00CF}', 600), ('\u{00D0}', 600), ('\u{00D1}', 600), ('\u{00D2}', 600),
('\u{00D3}', 600), ('\u{00D4}', 600), ('\u{00D5}', 600), ('\u{00D6}', 600), ('\u{00D7}', 600), ('\u{00D8}', 600),
('\u{00D9}', 600), ('\u{00DA}', 600), ('\u{00DB}', 600), ('\u{00DC}', 600), ('\u{00DD}', 600), ('\u{00DE}', 600),
('\u{00DF}', 600), ('\u{00E0}', 600), ('\u{00E1}', 600), ('\u{00E2}', 600), ('\u{00E3}', 600), ('\u{00E4}', 600),
('\u{00E5}', 600), ('\u{00E6}', 600), ('\u{00E7}', 600), ('\u{00E8}', 600), ('\u{00E9}', 600), ('\u{00EA}', 600),
('\u{00EB}', 600), ('\u{00EC}', 600), ('\u{00ED}', 600), ('\u{00EE}', 600), ('\u{00EF}', 600), ('\u{00F0}', 600),
('\u{00F1}', 600), ('\u{00F2}', 600), ('\u{00F3}', 600), ('\u{00F4}', 600), ('\u{00F5}', 600), ('\u{00F6}', 600),
('\u{00F7}', 600), ('\u{00F8}', 600), ('\u{00F9}', 600), ('\u{00FA}', 600), ('\u{00FB}', 600), ('\u{00FC}', 600),
('\u{00FD}', 600), ('\u{00FE}', 600), ('\u{00FF}', 600), ('\u{0131}', 600), ('\u{0141}', 600), ('\u{0142}', 600),
('\u{0152}', 600), ('\u{0153}', 600), ('\u{0160}', 600), ('\u{0161}', 600), ('\u{0178}', 600), ('\u{017D}', 600),
('\u{017E}', 600), ('\u{0192}', 600), ('\u{02C6}', 600), ('\u{02C7}', 600), ('\u{02D8}', 600), ('\u{02D9}', 600),
('\u{02DA}', 600), ('\u{02DB}', 600), ('\u{02DC}', 600), ('\u{02DD}', 600), ('\u{2013}', 600), ('\u{2014}', 600),
('\u{2018}', 600), ('\u{2019}', 600), ('\u{201A}', 600), ('\u{201C}', 600), ('\u{201D}', 600), ('\u{201E}', 600),
('\u{2020}', 600), ('\u{2021}', 600), ('\u{2022}', 600), ('\u{2026}', 600), ('\u{2030}', 600), ('\u{2039}', 600),
('\u{203A}', 600), ('\u{2044}', 600), ('\u{20AC}', 600), ('\u{2122}', 600), ('\u{2212}', 600), ('\u{FB01}', 600),
('\u{FB02}', 600),
];
static COURIER_BOLD: &[(char, u16)] = COURIER;
static COURIER_OBLIQUE: &[(char, u16)] = COURIER;
static COURIER_BOLDOBLIQUE: &[(char, u16)] = COURIER;
#[rustfmt::skip]
static HELVETICA: &[(char, u16)] = &[
(' ', 278), ('!', 278), ('"', 355), ('#', 556), ('$', 556), ('%', 889),
('&', 667), ('\'', 191), ('(', 333), (')', 333), ('*', 389), ('+', 584),
(',', 278), ('-', 333), ('.', 278), ('/', 278), ('0', 556), ('1', 556),
('2', 556), ('3', 556), ('4', 556), ('5', 556), ('6', 556), ('7', 556),
('8', 556), ('9', 556), (':', 278), (';', 278), ('<', 584), ('=', 584),
('>', 584), ('?', 556), ('@', 1015), ('A', 667), ('B', 667), ('C', 722),
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 722), ('I', 278),
('J', 500), ('K', 667), ('L', 556), ('M', 833), ('N', 722), ('O', 778),
('P', 667), ('Q', 778), ('R', 722), ('S', 667), ('T', 611), ('U', 722),
('V', 667), ('W', 944), ('X', 667), ('Y', 667), ('Z', 611), ('[', 278),
('\\', 278), (']', 278), ('^', 469), ('_', 556), ('`', 333), ('a', 556),
('b', 556), ('c', 500), ('d', 556), ('e', 556), ('f', 278), ('g', 556),
('h', 556), ('i', 222), ('j', 222), ('k', 500), ('l', 222), ('m', 833),
('n', 556), ('o', 556), ('p', 556), ('q', 556), ('r', 333), ('s', 500),
('t', 278), ('u', 556), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
('z', 500), ('{', 334), ('|', 260), ('}', 334), ('~', 584), ('\u{00A1}', 333),
('\u{00A2}', 556), ('\u{00A3}', 556), ('\u{00A4}', 556), ('\u{00A5}', 556), ('\u{00A6}', 260), ('\u{00A7}', 556),
('\u{00A8}', 333), ('\u{00A9}', 737), ('\u{00AA}', 370), ('\u{00AB}', 556), ('\u{00AC}', 584), ('\u{00AE}', 737),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 584), ('\u{00B2}', 333), ('\u{00B3}', 333), ('\u{00B4}', 333),
('\u{00B5}', 556), ('\u{00B6}', 537), ('\u{00B7}', 278), ('\u{00B8}', 333), ('\u{00B9}', 333), ('\u{00BA}', 365),
('\u{00BB}', 556), ('\u{00BC}', 834), ('\u{00BD}', 834), ('\u{00BE}', 834), ('\u{00BF}', 611), ('\u{00C0}', 667),
('\u{00C1}', 667), ('\u{00C2}', 667), ('\u{00C3}', 667), ('\u{00C4}', 667), ('\u{00C5}', 667), ('\u{00C6}', 1000),
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 278),
('\u{00CD}', 278), ('\u{00CE}', 278), ('\u{00CF}', 278), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 584), ('\u{00D8}', 778),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 667), ('\u{00DE}', 667),
('\u{00DF}', 611), ('\u{00E0}', 556), ('\u{00E1}', 556), ('\u{00E2}', 556), ('\u{00E3}', 556), ('\u{00E4}', 556),
('\u{00E5}', 556), ('\u{00E6}', 889), ('\u{00E7}', 500), ('\u{00E8}', 556), ('\u{00E9}', 556), ('\u{00EA}', 556),
('\u{00EB}', 556), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 556),
('\u{00F1}', 556), ('\u{00F2}', 556), ('\u{00F3}', 556), ('\u{00F4}', 556), ('\u{00F5}', 556), ('\u{00F6}', 556),
('\u{00F7}', 584), ('\u{00F8}', 611), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
('\u{00FD}', 500), ('\u{00FE}', 556), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 556), ('\u{0142}', 222),
('\u{0152}', 1000), ('\u{0153}', 944), ('\u{0160}', 667), ('\u{0161}', 500), ('\u{0178}', 667), ('\u{017D}', 611),
('\u{017E}', 500), ('\u{0192}', 556), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 556), ('\u{2014}', 1000),
('\u{2018}', 222), ('\u{2019}', 222), ('\u{201A}', 222), ('\u{201C}', 333), ('\u{201D}', 333), ('\u{201E}', 333),
('\u{2020}', 556), ('\u{2021}', 556), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 556), ('\u{2122}', 1000), ('\u{2212}', 584), ('\u{FB01}', 500),
('\u{FB02}', 500),
];
#[rustfmt::skip]
static HELVETICA_BOLD: &[(char, u16)] = &[
(' ', 278), ('!', 333), ('"', 474), ('#', 556), ('$', 556), ('%', 889),
('&', 722), ('\'', 238), ('(', 333), (')', 333), ('*', 389), ('+', 584),
(',', 278), ('-', 333), ('.', 278), ('/', 278), ('0', 556), ('1', 556),
('2', 556), ('3', 556), ('4', 556), ('5', 556), ('6', 556), ('7', 556),
('8', 556), ('9', 556), (':', 333), (';', 333), ('<', 584), ('=', 584),
('>', 584), ('?', 611), ('@', 975), ('A', 722), ('B', 722), ('C', 722),
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 722), ('I', 278),
('J', 556), ('K', 722), ('L', 611), ('M', 833), ('N', 722), ('O', 778),
('P', 667), ('Q', 778), ('R', 722), ('S', 667), ('T', 611), ('U', 722),
('V', 667), ('W', 944), ('X', 667), ('Y', 667), ('Z', 611), ('[', 333),
('\\', 278), (']', 333), ('^', 584), ('_', 556), ('`', 333), ('a', 556),
('b', 611), ('c', 556), ('d', 611), ('e', 556), ('f', 333), ('g', 611),
('h', 611), ('i', 278), ('j', 278), ('k', 556), ('l', 278), ('m', 889),
('n', 611), ('o', 611), ('p', 611), ('q', 611), ('r', 389), ('s', 556),
('t', 333), ('u', 611), ('v', 556), ('w', 778), ('x', 556), ('y', 556),
('z', 500), ('{', 389), ('|', 280), ('}', 389), ('~', 584), ('\u{00A1}', 333),
('\u{00A2}', 556), ('\u{00A3}', 556), ('\u{00A4}', 556), ('\u{00A5}', 556), ('\u{00A6}', 280), ('\u{00A7}', 556),
('\u{00A8}', 333), ('\u{00A9}', 737), ('\u{00AA}', 370), ('\u{00AB}', 556), ('\u{00AC}', 584), ('\u{00AE}', 737),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 584), ('\u{00B2}', 333), ('\u{00B3}', 333), ('\u{00B4}', 333),
('\u{00B5}', 611), ('\u{00B6}', 556), ('\u{00B7}', 278), ('\u{00B8}', 333), ('\u{00B9}', 333), ('\u{00BA}', 365),
('\u{00BB}', 556), ('\u{00BC}', 834), ('\u{00BD}', 834), ('\u{00BE}', 834), ('\u{00BF}', 611), ('\u{00C0}', 722),
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 1000),
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 278),
('\u{00CD}', 278), ('\u{00CE}', 278), ('\u{00CF}', 278), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 584), ('\u{00D8}', 778),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 667), ('\u{00DE}', 667),
('\u{00DF}', 611), ('\u{00E0}', 556), ('\u{00E1}', 556), ('\u{00E2}', 556), ('\u{00E3}', 556), ('\u{00E4}', 556),
('\u{00E5}', 556), ('\u{00E6}', 889), ('\u{00E7}', 556), ('\u{00E8}', 556), ('\u{00E9}', 556), ('\u{00EA}', 556),
('\u{00EB}', 556), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 611),
('\u{00F1}', 611), ('\u{00F2}', 611), ('\u{00F3}', 611), ('\u{00F4}', 611), ('\u{00F5}', 611), ('\u{00F6}', 611),
('\u{00F7}', 584), ('\u{00F8}', 611), ('\u{00F9}', 611), ('\u{00FA}', 611), ('\u{00FB}', 611), ('\u{00FC}', 611),
('\u{00FD}', 556), ('\u{00FE}', 611), ('\u{00FF}', 556), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
('\u{0152}', 1000), ('\u{0153}', 944), ('\u{0160}', 667), ('\u{0161}', 556), ('\u{0178}', 667), ('\u{017D}', 611),
('\u{017E}', 500), ('\u{0192}', 556), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 556), ('\u{2014}', 1000),
('\u{2018}', 278), ('\u{2019}', 278), ('\u{201A}', 278), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
('\u{2020}', 556), ('\u{2021}', 556), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 556), ('\u{2122}', 1000), ('\u{2212}', 584), ('\u{FB01}', 611),
('\u{FB02}', 611),
];
static HELVETICA_OBLIQUE: &[(char, u16)] = HELVETICA;
static HELVETICA_BOLDOBLIQUE: &[(char, u16)] = HELVETICA_BOLD;
#[rustfmt::skip]
static TIMES_ROMAN: &[(char, u16)] = &[
(' ', 250), ('!', 333), ('"', 408), ('#', 500), ('$', 500), ('%', 833),
('&', 778), ('\'', 180), ('(', 333), (')', 333), ('*', 500), ('+', 564),
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
('8', 500), ('9', 500), (':', 278), (';', 278), ('<', 564), ('=', 564),
('>', 564), ('?', 444), ('@', 921), ('A', 722), ('B', 667), ('C', 667),
('D', 722), ('E', 611), ('F', 556), ('G', 722), ('H', 722), ('I', 333),
('J', 389), ('K', 722), ('L', 611), ('M', 889), ('N', 722), ('O', 722),
('P', 556), ('Q', 722), ('R', 667), ('S', 556), ('T', 611), ('U', 722),
('V', 722), ('W', 944), ('X', 722), ('Y', 722), ('Z', 611), ('[', 333),
('\\', 278), (']', 333), ('^', 469), ('_', 500), ('`', 333), ('a', 444),
('b', 500), ('c', 444), ('d', 500), ('e', 444), ('f', 333), ('g', 500),
('h', 500), ('i', 278), ('j', 278), ('k', 500), ('l', 278), ('m', 778),
('n', 500), ('o', 500), ('p', 500), ('q', 500), ('r', 333), ('s', 389),
('t', 278), ('u', 500), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
('z', 444), ('{', 480), ('|', 200), ('}', 480), ('~', 541), ('\u{00A1}', 333),
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 200), ('\u{00A7}', 500),
('\u{00A8}', 333), ('\u{00A9}', 760), ('\u{00AA}', 276), ('\u{00AB}', 500), ('\u{00AC}', 564), ('\u{00AE}', 760),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 564), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
('\u{00B5}', 500), ('\u{00B6}', 453), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 310),
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 444), ('\u{00C0}', 722),
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 889),
('\u{00C7}', 667), ('\u{00C8}', 611), ('\u{00C9}', 611), ('\u{00CA}', 611), ('\u{00CB}', 611), ('\u{00CC}', 333),
('\u{00CD}', 333), ('\u{00CE}', 333), ('\u{00CF}', 333), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 722),
('\u{00D3}', 722), ('\u{00D4}', 722), ('\u{00D5}', 722), ('\u{00D6}', 722), ('\u{00D7}', 564), ('\u{00D8}', 722),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 722), ('\u{00DE}', 556),
('\u{00DF}', 500), ('\u{00E0}', 444), ('\u{00E1}', 444), ('\u{00E2}', 444), ('\u{00E3}', 444), ('\u{00E4}', 444),
('\u{00E5}', 444), ('\u{00E6}', 667), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
('\u{00F1}', 500), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
('\u{00F7}', 564), ('\u{00F8}', 500), ('\u{00F9}', 500), ('\u{00FA}', 500), ('\u{00FB}', 500), ('\u{00FC}', 500),
('\u{00FD}', 500), ('\u{00FE}', 500), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
('\u{0152}', 889), ('\u{0153}', 722), ('\u{0160}', 556), ('\u{0161}', 389), ('\u{0178}', 722), ('\u{017D}', 611),
('\u{017E}', 444), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 1000),
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 444), ('\u{201D}', 444), ('\u{201E}', 444),
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 980), ('\u{2212}', 564), ('\u{FB01}', 556),
('\u{FB02}', 556),
];
#[rustfmt::skip]
static TIMES_BOLD: &[(char, u16)] = &[
(' ', 250), ('!', 333), ('"', 555), ('#', 500), ('$', 500), ('%', 1000),
('&', 833), ('\'', 278), ('(', 333), (')', 333), ('*', 500), ('+', 570),
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
('8', 500), ('9', 500), (':', 333), (';', 333), ('<', 570), ('=', 570),
('>', 570), ('?', 500), ('@', 930), ('A', 722), ('B', 667), ('C', 722),
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 778), ('I', 389),
('J', 500), ('K', 778), ('L', 667), ('M', 944), ('N', 722), ('O', 778),
('P', 611), ('Q', 778), ('R', 722), ('S', 556), ('T', 667), ('U', 722),
('V', 722), ('W', 1000), ('X', 722), ('Y', 722), ('Z', 667), ('[', 333),
('\\', 278), (']', 333), ('^', 581), ('_', 500), ('`', 333), ('a', 500),
('b', 556), ('c', 444), ('d', 556), ('e', 444), ('f', 333), ('g', 500),
('h', 556), ('i', 278), ('j', 333), ('k', 556), ('l', 278), ('m', 833),
('n', 556), ('o', 500), ('p', 556), ('q', 556), ('r', 444), ('s', 389),
('t', 333), ('u', 556), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
('z', 444), ('{', 394), ('|', 220), ('}', 394), ('~', 520), ('\u{00A1}', 333),
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 220), ('\u{00A7}', 500),
('\u{00A8}', 333), ('\u{00A9}', 747), ('\u{00AA}', 300), ('\u{00AB}', 500), ('\u{00AC}', 570), ('\u{00AE}', 747),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 570), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
('\u{00B5}', 556), ('\u{00B6}', 540), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 330),
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 500), ('\u{00C0}', 722),
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 1000),
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 389),
('\u{00CD}', 389), ('\u{00CE}', 389), ('\u{00CF}', 389), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 570), ('\u{00D8}', 778),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 722), ('\u{00DE}', 611),
('\u{00DF}', 556), ('\u{00E0}', 500), ('\u{00E1}', 500), ('\u{00E2}', 500), ('\u{00E3}', 500), ('\u{00E4}', 500),
('\u{00E5}', 500), ('\u{00E6}', 722), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
('\u{00F1}', 556), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
('\u{00F7}', 570), ('\u{00F8}', 500), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
('\u{00FD}', 500), ('\u{00FE}', 556), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 667), ('\u{0142}', 278),
('\u{0152}', 1000), ('\u{0153}', 722), ('\u{0160}', 556), ('\u{0161}', 389), ('\u{0178}', 722), ('\u{017D}', 667),
('\u{017E}', 444), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 1000),
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 1000), ('\u{2212}', 570), ('\u{FB01}', 556),
('\u{FB02}', 556),
];
#[rustfmt::skip]
static TIMES_ITALIC: &[(char, u16)] = &[
(' ', 250), ('!', 333), ('"', 420), ('#', 500), ('$', 500), ('%', 833),
('&', 778), ('\'', 214), ('(', 333), (')', 333), ('*', 500), ('+', 675),
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
('8', 500), ('9', 500), (':', 333), (';', 333), ('<', 675), ('=', 675),
('>', 675), ('?', 500), ('@', 920), ('A', 611), ('B', 611), ('C', 667),
('D', 722), ('E', 611), ('F', 611), ('G', 722), ('H', 722), ('I', 333),
('J', 444), ('K', 667), ('L', 556), ('M', 833), ('N', 667), ('O', 722),
('P', 611), ('Q', 722), ('R', 611), ('S', 500), ('T', 556), ('U', 722),
('V', 611), ('W', 833), ('X', 611), ('Y', 556), ('Z', 556), ('[', 389),
('\\', 278), (']', 389), ('^', 422), ('_', 500), ('`', 333), ('a', 500),
('b', 500), ('c', 444), ('d', 500), ('e', 444), ('f', 278), ('g', 500),
('h', 500), ('i', 278), ('j', 278), ('k', 444), ('l', 278), ('m', 722),
('n', 500), ('o', 500), ('p', 500), ('q', 500), ('r', 389), ('s', 389),
('t', 278), ('u', 500), ('v', 444), ('w', 667), ('x', 444), ('y', 444),
('z', 389), ('{', 400), ('|', 275), ('}', 400), ('~', 541), ('\u{00A1}', 389),
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 275), ('\u{00A7}', 500),
('\u{00A8}', 333), ('\u{00A9}', 760), ('\u{00AA}', 276), ('\u{00AB}', 500), ('\u{00AC}', 675), ('\u{00AE}', 760),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 675), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
('\u{00B5}', 500), ('\u{00B6}', 523), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 310),
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 500), ('\u{00C0}', 611),
('\u{00C1}', 611), ('\u{00C2}', 611), ('\u{00C3}', 611), ('\u{00C4}', 611), ('\u{00C5}', 611), ('\u{00C6}', 889),
('\u{00C7}', 667), ('\u{00C8}', 611), ('\u{00C9}', 611), ('\u{00CA}', 611), ('\u{00CB}', 611), ('\u{00CC}', 333),
('\u{00CD}', 333), ('\u{00CE}', 333), ('\u{00CF}', 333), ('\u{00D0}', 722), ('\u{00D1}', 667), ('\u{00D2}', 722),
('\u{00D3}', 722), ('\u{00D4}', 722), ('\u{00D5}', 722), ('\u{00D6}', 722), ('\u{00D7}', 675), ('\u{00D8}', 722),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 556), ('\u{00DE}', 611),
('\u{00DF}', 500), ('\u{00E0}', 500), ('\u{00E1}', 500), ('\u{00E2}', 500), ('\u{00E3}', 500), ('\u{00E4}', 500),
('\u{00E5}', 500), ('\u{00E6}', 667), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
('\u{00F1}', 500), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
('\u{00F7}', 675), ('\u{00F8}', 500), ('\u{00F9}', 500), ('\u{00FA}', 500), ('\u{00FB}', 500), ('\u{00FC}', 500),
('\u{00FD}', 444), ('\u{00FE}', 500), ('\u{00FF}', 444), ('\u{0131}', 278), ('\u{0141}', 556), ('\u{0142}', 278),
('\u{0152}', 944), ('\u{0153}', 667), ('\u{0160}', 500), ('\u{0161}', 389), ('\u{0178}', 556), ('\u{017D}', 556),
('\u{017E}', 389), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 889),
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 556), ('\u{201D}', 556), ('\u{201E}', 556),
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 889), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 980), ('\u{2212}', 675), ('\u{FB01}', 500),
('\u{FB02}', 500),
];
#[rustfmt::skip]
static TIMES_BOLDITALIC: &[(char, u16)] = &[
(' ', 250), ('!', 389), ('"', 555), ('#', 500), ('$', 500), ('%', 833),
('&', 778), ('\'', 278), ('(', 333), (')', 333), ('*', 500), ('+', 570),
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
('8', 500), ('9', 500), (':', 333), (';', 333), ('<', 570), ('=', 570),
('>', 570), ('?', 500), ('@', 832), ('A', 667), ('B', 667), ('C', 667),
('D', 722), ('E', 667), ('F', 667), ('G', 722), ('H', 778), ('I', 389),
('J', 500), ('K', 667), ('L', 611), ('M', 889), ('N', 722), ('O', 722),
('P', 611), ('Q', 722), ('R', 667), ('S', 556), ('T', 611), ('U', 722),
('V', 667), ('W', 889), ('X', 667), ('Y', 611), ('Z', 611), ('[', 333),
('\\', 278), (']', 333), ('^', 570), ('_', 500), ('`', 333), ('a', 500),
('b', 500), ('c', 444), ('d', 500), ('e', 444), ('f', 333), ('g', 500),
('h', 556), ('i', 278), ('j', 278), ('k', 500), ('l', 278), ('m', 778),
('n', 556), ('o', 500), ('p', 500), ('q', 500), ('r', 389), ('s', 389),
('t', 278), ('u', 556), ('v', 444), ('w', 667), ('x', 500), ('y', 444),
('z', 389), ('{', 348), ('|', 220), ('}', 348), ('~', 570), ('\u{00A1}', 389),
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 220), ('\u{00A7}', 500),
('\u{00A8}', 333), ('\u{00A9}', 747), ('\u{00AA}', 266), ('\u{00AB}', 500), ('\u{00AC}', 606), ('\u{00AE}', 747),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 570), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
('\u{00B5}', 576), ('\u{00B6}', 500), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 300),
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 500), ('\u{00C0}', 667),
('\u{00C1}', 667), ('\u{00C2}', 667), ('\u{00C3}', 667), ('\u{00C4}', 667), ('\u{00C5}', 667), ('\u{00C6}', 944),
('\u{00C7}', 667), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 389),
('\u{00CD}', 389), ('\u{00CE}', 389), ('\u{00CF}', 389), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 722),
('\u{00D3}', 722), ('\u{00D4}', 722), ('\u{00D5}', 722), ('\u{00D6}', 722), ('\u{00D7}', 570), ('\u{00D8}', 722),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 611), ('\u{00DE}', 611),
('\u{00DF}', 500), ('\u{00E0}', 500), ('\u{00E1}', 500), ('\u{00E2}', 500), ('\u{00E3}', 500), ('\u{00E4}', 500),
('\u{00E5}', 500), ('\u{00E6}', 722), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
('\u{00F1}', 556), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
('\u{00F7}', 570), ('\u{00F8}', 500), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
('\u{00FD}', 444), ('\u{00FE}', 500), ('\u{00FF}', 444), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
('\u{0152}', 944), ('\u{0153}', 722), ('\u{0160}', 556), ('\u{0161}', 389), ('\u{0178}', 611), ('\u{017D}', 611),
('\u{017E}', 389), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 1000),
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 1000), ('\u{2212}', 606), ('\u{FB01}', 556),
('\u{FB02}', 556),
];
#[rustfmt::skip]
static SYMBOL: &[(char, u16)] = &[
(' ', 250), ('!', 333), ('#', 500), ('%', 833), ('&', 778), ('(', 333),
(')', 333), ('+', 549), (',', 250), ('.', 250), ('/', 278), ('0', 500),
('1', 500), ('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500),
('7', 500), ('8', 500), ('9', 500), (':', 278), (';', 278), ('<', 549),
('=', 549), ('>', 549), ('?', 444), ('[', 333), (']', 333), ('_', 500),
('{', 480), ('|', 200), ('}', 480), ('\u{00AC}', 713), ('\u{00B0}', 400), ('\u{00B1}', 549),
('\u{00B5}', 576), ('\u{00D7}', 549), ('\u{00F7}', 549), ('\u{0192}', 500), ('\u{0391}', 722), ('\u{0392}', 667),
('\u{0393}', 603), ('\u{0395}', 611), ('\u{0396}', 611), ('\u{0397}', 722), ('\u{0398}', 741), ('\u{0399}', 333),
('\u{039A}', 722), ('\u{039B}', 686), ('\u{039C}', 889), ('\u{039D}', 722), ('\u{039E}', 645), ('\u{039F}', 722),
('\u{03A0}', 768), ('\u{03A1}', 556), ('\u{03A3}', 592), ('\u{03A4}', 611), ('\u{03A5}', 690), ('\u{03A6}', 763),
('\u{03A7}', 722), ('\u{03A8}', 795), ('\u{03B1}', 631), ('\u{03B2}', 549), ('\u{03B3}', 411), ('\u{03B4}', 494),
('\u{03B5}', 439), ('\u{03B6}', 494), ('\u{03B7}', 603), ('\u{03B8}', 521), ('\u{03B9}', 329), ('\u{03BA}', 549),
('\u{03BB}', 549), ('\u{03BD}', 521), ('\u{03BE}', 493), ('\u{03BF}', 549), ('\u{03C0}', 549), ('\u{03C1}', 549),
('\u{03C2}', 439), ('\u{03C3}', 603), ('\u{03C4}', 439), ('\u{03C5}', 576), ('\u{03C6}', 521), ('\u{03C7}', 549),
('\u{03C8}', 686), ('\u{03C9}', 686), ('\u{03D1}', 631), ('\u{03D2}', 620), ('\u{03D5}', 603), ('\u{03D6}', 713),
('\u{2022}', 460), ('\u{2026}', 1000), ('\u{2032}', 247), ('\u{2033}', 411), ('\u{2044}', 167), ('\u{20AC}', 750),
('\u{2111}', 686), ('\u{2118}', 987), ('\u{211C}', 795), ('\u{2126}', 768), ('\u{2135}', 823), ('\u{2190}', 987),
('\u{2191}', 603), ('\u{2192}', 987), ('\u{2193}', 603), ('\u{2194}', 1042), ('\u{21B5}', 658), ('\u{21D0}', 987),
('\u{21D1}', 603), ('\u{21D2}', 987), ('\u{21D3}', 603), ('\u{21D4}', 1042), ('\u{2200}', 713), ('\u{2202}', 494),
('\u{2203}', 549), ('\u{2205}', 823), ('\u{2206}', 612), ('\u{2207}', 713), ('\u{2208}', 713), ('\u{2209}', 713),
('\u{220B}', 439), ('\u{220F}', 823), ('\u{2211}', 713), ('\u{2212}', 549), ('\u{2217}', 500), ('\u{221A}', 549),
('\u{221D}', 713), ('\u{221E}', 713), ('\u{2220}', 768), ('\u{2227}', 603), ('\u{2228}', 603), ('\u{2229}', 768),
('\u{222A}', 768), ('\u{222B}', 274), ('\u{2234}', 863), ('\u{223C}', 549), ('\u{2245}', 549), ('\u{2248}', 549),
('\u{2260}', 549), ('\u{2261}', 549), ('\u{2264}', 549), ('\u{2265}', 549), ('\u{2282}', 713), ('\u{2283}', 713),
('\u{2284}', 713), ('\u{2286}', 713), ('\u{2287}', 713), ('\u{2295}', 768), ('\u{2297}', 768), ('\u{22A5}', 658),
('\u{22C5}', 250), ('\u{2320}', 686), ('\u{2321}', 686), ('\u{2329}', 329), ('\u{232A}', 329), ('\u{25CA}', 494),
('\u{2660}', 753), ('\u{2663}', 753), ('\u{2665}', 753), ('\u{2666}', 753), ('\u{F6D9}', 790), ('\u{F6DA}', 790),
('\u{F6DB}', 890), ('\u{F8E5}', 500), ('\u{F8E6}', 603), ('\u{F8E7}', 1000), ('\u{F8E8}', 790), ('\u{F8E9}', 790),
('\u{F8EA}', 786), ('\u{F8EB}', 384), ('\u{F8EC}', 384), ('\u{F8ED}', 384), ('\u{F8EE}', 384), ('\u{F8EF}', 384),
('\u{F8F0}', 384), ('\u{F8F1}', 494), ('\u{F8F2}', 494), ('\u{F8F3}', 494), ('\u{F8F4}', 494), ('\u{F8F5}', 686),
('\u{F8F6}', 384), ('\u{F8F7}', 384), ('\u{F8F8}', 384), ('\u{F8F9}', 384), ('\u{F8FA}', 384), ('\u{F8FB}', 384),
('\u{F8FC}', 494), ('\u{F8FD}', 494), ('\u{F8FE}', 494), ('\u{F8FF}', 790),
];
#[rustfmt::skip]
static ZAPFDINGBATS: &[(char, u16)] = &[
(' ', 278), ('\u{2192}', 838), ('\u{2194}', 1016), ('\u{2195}', 458), ('\u{2460}', 788), ('\u{2461}', 788),
('\u{2462}', 788), ('\u{2463}', 788), ('\u{2464}', 788), ('\u{2465}', 788), ('\u{2466}', 788), ('\u{2467}', 788),
('\u{2468}', 788), ('\u{2469}', 788), ('\u{25A0}', 761), ('\u{25B2}', 892), ('\u{25BC}', 892), ('\u{25C6}', 788),
('\u{25CF}', 791), ('\u{25D7}', 438), ('\u{2605}', 816), ('\u{260E}', 719), ('\u{261B}', 960), ('\u{261E}', 939),
('\u{2660}', 626), ('\u{2663}', 776), ('\u{2665}', 694), ('\u{2666}', 595), ('\u{2701}', 974), ('\u{2702}', 961),
('\u{2703}', 974), ('\u{2704}', 980), ('\u{2706}', 789), ('\u{2707}', 790), ('\u{2708}', 791), ('\u{2709}', 690),
('\u{270C}', 549), ('\u{270D}', 855), ('\u{270E}', 911), ('\u{270F}', 933), ('\u{2710}', 911), ('\u{2711}', 945),
('\u{2712}', 974), ('\u{2713}', 755), ('\u{2714}', 846), ('\u{2715}', 762), ('\u{2716}', 761), ('\u{2717}', 571),
('\u{2718}', 677), ('\u{2719}', 763), ('\u{271A}', 760), ('\u{271B}', 759), ('\u{271C}', 754), ('\u{271D}', 494),
('\u{271E}', 552), ('\u{271F}', 537), ('\u{2720}', 577), ('\u{2721}', 692), ('\u{2722}', 786), ('\u{2723}', 788),
('\u{2724}', 788), ('\u{2725}', 790), ('\u{2726}', 793), ('\u{2727}', 794), ('\u{2729}', 823), ('\u{272A}', 789),
('\u{272B}', 841), ('\u{272C}', 823), ('\u{272D}', 833), ('\u{272E}', 816), ('\u{272F}', 831), ('\u{2730}', 923),
('\u{2731}', 744), ('\u{2732}', 723), ('\u{2733}', 749), ('\u{2734}', 790), ('\u{2735}', 792), ('\u{2736}', 695),
('\u{2737}', 776), ('\u{2738}', 768), ('\u{2739}', 792), ('\u{273A}', 759), ('\u{273B}', 707), ('\u{273C}', 708),
('\u{273D}', 682), ('\u{273E}', 701), ('\u{273F}', 826), ('\u{2740}', 815), ('\u{2741}', 789), ('\u{2742}', 789),
('\u{2743}', 707), ('\u{2744}', 687), ('\u{2745}', 696), ('\u{2746}', 689), ('\u{2747}', 786), ('\u{2748}', 787),
('\u{2749}', 713), ('\u{274A}', 791), ('\u{274B}', 785), ('\u{274D}', 873), ('\u{274F}', 762), ('\u{2750}', 762),
('\u{2751}', 759), ('\u{2752}', 759), ('\u{2756}', 784), ('\u{2758}', 138), ('\u{2759}', 277), ('\u{275A}', 415),
('\u{275B}', 392), ('\u{275C}', 392), ('\u{275D}', 668), ('\u{275E}', 668), ('\u{2761}', 732), ('\u{2762}', 544),
('\u{2763}', 544), ('\u{2764}', 910), ('\u{2765}', 667), ('\u{2766}', 760), ('\u{2767}', 760), ('\u{2768}', 390),
('\u{2769}', 390), ('\u{276A}', 317), ('\u{276B}', 317), ('\u{276C}', 276), ('\u{276D}', 276), ('\u{276E}', 509),
('\u{276F}', 509), ('\u{2770}', 410), ('\u{2771}', 410), ('\u{2772}', 234), ('\u{2773}', 234), ('\u{2774}', 334),
('\u{2775}', 334), ('\u{2776}', 788), ('\u{2777}', 788), ('\u{2778}', 788), ('\u{2779}', 788), ('\u{277A}', 788),
('\u{277B}', 788), ('\u{277C}', 788), ('\u{277D}', 788), ('\u{277E}', 788), ('\u{277F}', 788), ('\u{2780}', 788),
('\u{2781}', 788), ('\u{2782}', 788), ('\u{2783}', 788), ('\u{2784}', 788), ('\u{2785}', 788), ('\u{2786}', 788),
('\u{2787}', 788), ('\u{2788}', 788), ('\u{2789}', 788), ('\u{278A}', 788), ('\u{278B}', 788), ('\u{278C}', 788),
('\u{278D}', 788), ('\u{278E}', 788), ('\u{278F}', 788), ('\u{2790}', 788), ('\u{2791}', 788), ('\u{2792}', 788),
('\u{2793}', 788), ('\u{2794}', 894), ('\u{2798}', 748), ('\u{2799}', 924), ('\u{279A}', 748), ('\u{279B}', 918),
('\u{279C}', 927), ('\u{279D}', 928), ('\u{279E}', 928), ('\u{279F}', 834), ('\u{27A0}', 873), ('\u{27A1}', 828),
('\u{27A2}', 924), ('\u{27A3}', 924), ('\u{27A4}', 917), ('\u{27A5}', 930), ('\u{27A6}', 931), ('\u{27A7}', 463),
('\u{27A8}', 883), ('\u{27A9}', 836), ('\u{27AA}', 836), ('\u{27AB}', 867), ('\u{27AC}', 867), ('\u{27AD}', 696),
('\u{27AE}', 696), ('\u{27AF}', 874), ('\u{27B1}', 874), ('\u{27B2}', 760), ('\u{27B3}', 946), ('\u{27B4}', 771),
('\u{27B5}', 865), ('\u{27B6}', 771), ('\u{27B7}', 888), ('\u{27B8}', 967), ('\u{27B9}', 888), ('\u{27BA}', 831),
('\u{27BB}', 873), ('\u{27BC}', 927), ('\u{27BD}', 970), ('\u{27BE}', 918),
];
#[rustfmt::skip]
static SYMBOL_ENCODING: &[(u8, char)] = &[
(0x20, ' '), (0x21, '!'), (0x22, '\u{2200}'), (0x23, '#'), (0x24, '\u{2203}'), (0x25, '%'),
(0x26, '&'), (0x27, '\u{220B}'), (0x28, '('), (0x29, ')'), (0x2A, '\u{2217}'), (0x2B, '+'),
(0x2C, ','), (0x2D, '\u{2212}'), (0x2E, '.'), (0x2F, '/'), (0x30, '0'), (0x31, '1'),
(0x32, '2'), (0x33, '3'), (0x34, '4'), (0x35, '5'), (0x36, '6'), (0x37, '7'),
(0x38, '8'), (0x39, '9'), (0x3A, ':'), (0x3B, ';'), (0x3C, '<'), (0x3D, '='),
(0x3E, '>'), (0x3F, '?'), (0x40, '\u{2245}'), (0x41, '\u{0391}'), (0x42, '\u{0392}'), (0x43, '\u{03A7}'),
(0x44, '\u{2206}'), (0x45, '\u{0395}'), (0x46, '\u{03A6}'), (0x47, '\u{0393}'), (0x48, '\u{0397}'), (0x49, '\u{0399}'),
(0x4A, '\u{03D1}'), (0x4B, '\u{039A}'), (0x4C, '\u{039B}'), (0x4D, '\u{039C}'), (0x4E, '\u{039D}'), (0x4F, '\u{039F}'),
(0x50, '\u{03A0}'), (0x51, '\u{0398}'), (0x52, '\u{03A1}'), (0x53, '\u{03A3}'), (0x54, '\u{03A4}'), (0x55, '\u{03A5}'),
(0x56, '\u{03C2}'), (0x57, '\u{2126}'), (0x58, '\u{039E}'), (0x59, '\u{03A8}'), (0x5A, '\u{0396}'), (0x5B, '['),
(0x5C, '\u{2234}'), (0x5D, ']'), (0x5E, '\u{22A5}'), (0x5F, '_'), (0x60, '\u{F8E5}'), (0x61, '\u{03B1}'),
(0x62, '\u{03B2}'), (0x63, '\u{03C7}'), (0x64, '\u{03B4}'), (0x65, '\u{03B5}'), (0x66, '\u{03C6}'), (0x67, '\u{03B3}'),
(0x68, '\u{03B7}'), (0x69, '\u{03B9}'), (0x6A, '\u{03D5}'), (0x6B, '\u{03BA}'), (0x6C, '\u{03BB}'), (0x6D, '\u{00B5}'),
(0x6E, '\u{03BD}'), (0x6F, '\u{03BF}'), (0x70, '\u{03C0}'), (0x71, '\u{03B8}'), (0x72, '\u{03C1}'), (0x73, '\u{03C3}'),
(0x74, '\u{03C4}'), (0x75, '\u{03C5}'), (0x76, '\u{03D6}'), (0x77, '\u{03C9}'), (0x78, '\u{03BE}'), (0x79, '\u{03C8}'),
(0x7A, '\u{03B6}'), (0x7B, '{'), (0x7C, '|'), (0x7D, '}'), (0x7E, '\u{223C}'), (0xA0, '\u{20AC}'),
(0xA1, '\u{03D2}'), (0xA2, '\u{2032}'), (0xA3, '\u{2264}'), (0xA4, '\u{2044}'), (0xA5, '\u{221E}'), (0xA6, '\u{0192}'),
(0xA7, '\u{2663}'), (0xA8, '\u{2666}'), (0xA9, '\u{2665}'), (0xAA, '\u{2660}'), (0xAB, '\u{2194}'), (0xAC, '\u{2190}'),
(0xAD, '\u{2191}'), (0xAE, '\u{2192}'), (0xAF, '\u{2193}'), (0xB0, '\u{00B0}'), (0xB1, '\u{00B1}'), (0xB2, '\u{2033}'),
(0xB3, '\u{2265}'), (0xB4, '\u{00D7}'), (0xB5, '\u{221D}'), (0xB6, '\u{2202}'), (0xB7, '\u{2022}'), (0xB8, '\u{00F7}'),
(0xB9, '\u{2260}'), (0xBA, '\u{2261}'), (0xBB, '\u{2248}'), (0xBC, '\u{2026}'), (0xBD, '\u{F8E6}'), (0xBE, '\u{F8E7}'),
(0xBF, '\u{21B5}'), (0xC0, '\u{2135}'), (0xC1, '\u{2111}'), (0xC2, '\u{211C}'), (0xC3, '\u{2118}'), (0xC4, '\u{2297}'),
(0xC5, '\u{2295}'), (0xC6, '\u{2205}'), (0xC7, '\u{2229}'), (0xC8, '\u{222A}'), (0xC9, '\u{2283}'), (0xCA, '\u{2287}'),
(0xCB, '\u{2284}'), (0xCC, '\u{2282}'), (0xCD, '\u{2286}'), (0xCE, '\u{2208}'), (0xCF, '\u{2209}'), (0xD0, '\u{2220}'),
(0xD1, '\u{2207}'), (0xD2, '\u{F6DA}'), (0xD3, '\u{F6D9}'), (0xD4, '\u{F6DB}'), (0xD5, '\u{220F}'), (0xD6, '\u{221A}'),
(0xD7, '\u{22C5}'), (0xD8, '\u{00AC}'), (0xD9, '\u{2227}'), (0xDA, '\u{2228}'), (0xDB, '\u{21D4}'), (0xDC, '\u{21D0}'),
(0xDD, '\u{21D1}'), (0xDE, '\u{21D2}'), (0xDF, '\u{21D3}'), (0xE0, '\u{25CA}'), (0xE1, '\u{2329}'), (0xE2, '\u{F8E8}'),
(0xE3, '\u{F8E9}'), (0xE4, '\u{F8EA}'), (0xE5, '\u{2211}'), (0xE6, '\u{F8EB}'), (0xE7, '\u{F8EC}'), (0xE8, '\u{F8ED}'),
(0xE9, '\u{F8EE}'), (0xEA, '\u{F8EF}'), (0xEB, '\u{F8F0}'), (0xEC, '\u{F8F1}'), (0xED, '\u{F8F2}'), (0xEE, '\u{F8F3}'),
(0xEF, '\u{F8F4}'), (0xF1, '\u{232A}'), (0xF2, '\u{222B}'), (0xF3, '\u{2320}'), (0xF4, '\u{F8F5}'), (0xF5, '\u{2321}'),
(0xF6, '\u{F8F6}'), (0xF7, '\u{F8F7}'), (0xF8, '\u{F8F8}'), (0xF9, '\u{F8F9}'), (0xFA, '\u{F8FA}'), (0xFB, '\u{F8FB}'),
(0xFC, '\u{F8FC}'), (0xFD, '\u{F8FD}'), (0xFE, '\u{F8FE}'),
];
#[rustfmt::skip]
static ZAPFDINGBATS_ENCODING: &[(u8, char)] = &[
(0x20, ' '), (0x21, '\u{2701}'), (0x22, '\u{2702}'), (0x23, '\u{2703}'), (0x24, '\u{2704}'), (0x25, '\u{260E}'),
(0x26, '\u{2706}'), (0x27, '\u{2707}'), (0x28, '\u{2708}'), (0x29, '\u{2709}'), (0x2A, '\u{261B}'), (0x2B, '\u{261E}'),
(0x2C, '\u{270C}'), (0x2D, '\u{270D}'), (0x2E, '\u{270E}'), (0x2F, '\u{270F}'), (0x30, '\u{2710}'), (0x31, '\u{2711}'),
(0x32, '\u{2712}'), (0x33, '\u{2713}'), (0x34, '\u{2714}'), (0x35, '\u{2715}'), (0x36, '\u{2716}'), (0x37, '\u{2717}'),
(0x38, '\u{2718}'), (0x39, '\u{2719}'), (0x3A, '\u{271A}'), (0x3B, '\u{271B}'), (0x3C, '\u{271C}'), (0x3D, '\u{271D}'),
(0x3E, '\u{271E}'), (0x3F, '\u{271F}'), (0x40, '\u{2720}'), (0x41, '\u{2721}'), (0x42, '\u{2722}'), (0x43, '\u{2723}'),
(0x44, '\u{2724}'), (0x45, '\u{2725}'), (0x46, '\u{2726}'), (0x47, '\u{2727}'), (0x48, '\u{2605}'), (0x49, '\u{2729}'),
(0x4A, '\u{272A}'), (0x4B, '\u{272B}'), (0x4C, '\u{272C}'), (0x4D, '\u{272D}'), (0x4E, '\u{272E}'), (0x4F, '\u{272F}'),
(0x50, '\u{2730}'), (0x51, '\u{2731}'), (0x52, '\u{2732}'), (0x53, '\u{2733}'), (0x54, '\u{2734}'), (0x55, '\u{2735}'),
(0x56, '\u{2736}'), (0x57, '\u{2737}'), (0x58, '\u{2738}'), (0x59, '\u{2739}'), (0x5A, '\u{273A}'), (0x5B, '\u{273B}'),
(0x5C, '\u{273C}'), (0x5D, '\u{273D}'), (0x5E, '\u{273E}'), (0x5F, '\u{273F}'), (0x60, '\u{2740}'), (0x61, '\u{2741}'),
(0x62, '\u{2742}'), (0x63, '\u{2743}'), (0x64, '\u{2744}'), (0x65, '\u{2745}'), (0x66, '\u{2746}'), (0x67, '\u{2747}'),
(0x68, '\u{2748}'), (0x69, '\u{2749}'), (0x6A, '\u{274A}'), (0x6B, '\u{274B}'), (0x6C, '\u{25CF}'), (0x6D, '\u{274D}'),
(0x6E, '\u{25A0}'), (0x6F, '\u{274F}'), (0x70, '\u{2750}'), (0x71, '\u{2751}'), (0x72, '\u{2752}'), (0x73, '\u{25B2}'),
(0x74, '\u{25BC}'), (0x75, '\u{25C6}'), (0x76, '\u{2756}'), (0x77, '\u{25D7}'), (0x78, '\u{2758}'), (0x79, '\u{2759}'),
(0x7A, '\u{275A}'), (0x7B, '\u{275B}'), (0x7C, '\u{275C}'), (0x7D, '\u{275D}'), (0x7E, '\u{275E}'), (0x80, '\u{2768}'),
(0x81, '\u{2769}'), (0x82, '\u{276A}'), (0x83, '\u{276B}'), (0x84, '\u{276C}'), (0x85, '\u{276D}'), (0x86, '\u{276E}'),
(0x87, '\u{276F}'), (0x88, '\u{2770}'), (0x89, '\u{2771}'), (0x8A, '\u{2772}'), (0x8B, '\u{2773}'), (0x8C, '\u{2774}'),
(0x8D, '\u{2775}'), (0xA1, '\u{2761}'), (0xA2, '\u{2762}'), (0xA3, '\u{2763}'), (0xA4, '\u{2764}'), (0xA5, '\u{2765}'),
(0xA6, '\u{2766}'), (0xA7, '\u{2767}'), (0xA8, '\u{2663}'), (0xA9, '\u{2666}'), (0xAA, '\u{2665}'), (0xAB, '\u{2660}'),
(0xAC, '\u{2460}'), (0xAD, '\u{2461}'), (0xAE, '\u{2462}'), (0xAF, '\u{2463}'), (0xB0, '\u{2464}'), (0xB1, '\u{2465}'),
(0xB2, '\u{2466}'), (0xB3, '\u{2467}'), (0xB4, '\u{2468}'), (0xB5, '\u{2469}'), (0xB6, '\u{2776}'), (0xB7, '\u{2777}'),
(0xB8, '\u{2778}'), (0xB9, '\u{2779}'), (0xBA, '\u{277A}'), (0xBB, '\u{277B}'), (0xBC, '\u{277C}'), (0xBD, '\u{277D}'),
(0xBE, '\u{277E}'), (0xBF, '\u{277F}'), (0xC0, '\u{2780}'), (0xC1, '\u{2781}'), (0xC2, '\u{2782}'), (0xC3, '\u{2783}'),
(0xC4, '\u{2784}'), (0xC5, '\u{2785}'), (0xC6, '\u{2786}'), (0xC7, '\u{2787}'), (0xC8, '\u{2788}'), (0xC9, '\u{2789}'),
(0xCA, '\u{278A}'), (0xCB, '\u{278B}'), (0xCC, '\u{278C}'), (0xCD, '\u{278D}'), (0xCE, '\u{278E}'), (0xCF, '\u{278F}'),
(0xD0, '\u{2790}'), (0xD1, '\u{2791}'), (0xD2, '\u{2792}'), (0xD3, '\u{2793}'), (0xD4, '\u{2794}'), (0xD5, '\u{2192}'),
(0xD6, '\u{2194}'), (0xD7, '\u{2195}'), (0xD8, '\u{2798}'), (0xD9, '\u{2799}'), (0xDA, '\u{279A}'), (0xDB, '\u{279B}'),
(0xDC, '\u{279C}'), (0xDD, '\u{279D}'), (0xDE, '\u{279E}'), (0xDF, '\u{279F}'), (0xE0, '\u{27A0}'), (0xE1, '\u{27A1}'),
(0xE2, '\u{27A2}'), (0xE3, '\u{27A3}'), (0xE4, '\u{27A4}'), (0xE5, '\u{27A5}'), (0xE6, '\u{27A6}'), (0xE7, '\u{27A7}'),
(0xE8, '\u{27A8}'), (0xE9, '\u{27A9}'), (0xEA, '\u{27AA}'), (0xEB, '\u{27AB}'), (0xEC, '\u{27AC}'), (0xED, '\u{27AD}'),
(0xEE, '\u{27AE}'), (0xEF, '\u{27AF}'), (0xF1, '\u{27B1}'), (0xF2, '\u{27B2}'), (0xF3, '\u{27B3}'), (0xF4, '\u{27B4}'),
(0xF5, '\u{27B5}'), (0xF6, '\u{27B6}'), (0xF7, '\u{27B7}'), (0xF8, '\u{27B8}'), (0xF9, '\u{27B9}'), (0xFA, '\u{27BA}'),
(0xFB, '\u{27BB}'), (0xFC, '\u{27BC}'), (0xFD, '\u{27BD}'), (0xFE, '\u{27BE}'),
];
/// Every width table, for exhaustive testing.
#[cfg(test)]
static ALL_TABLES: &[(&str, &[(char, u16)])] = &[
("COURIER", COURIER),
("COURIER_BOLD", COURIER_BOLD),
("COURIER_OBLIQUE", COURIER_OBLIQUE),
("COURIER_BOLDOBLIQUE", COURIER_BOLDOBLIQUE),
("HELVETICA", HELVETICA),
("HELVETICA_BOLD", HELVETICA_BOLD),
("HELVETICA_OBLIQUE", HELVETICA_OBLIQUE),
("HELVETICA_BOLDOBLIQUE", HELVETICA_BOLDOBLIQUE),
("TIMES_ROMAN", TIMES_ROMAN),
("TIMES_BOLD", TIMES_BOLD),
("TIMES_ITALIC", TIMES_ITALIC),
("TIMES_BOLDITALIC", TIMES_BOLDITALIC),
("SYMBOL", SYMBOL),
("ZAPFDINGBATS", ZAPFDINGBATS),
];
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn times_roman_ascii_widths() {
assert_eq!(base14_char_width("Times-Roman", ' '), Some(250));
assert_eq!(base14_char_width("Times-Roman", 'M'), Some(889));
assert_eq!(base14_char_width("Times-Roman", 'i'), Some(278));
}
#[test]
fn subset_prefix_and_aliases_normalize() {
assert_eq!(
base14_char_width("ABCDEF+Times-Bold", ' '),
base14_char_width("Times-Bold", ' ')
);
assert!(base14_char_width("ArialMT", 'a').is_some());
assert!(base14_char_width("TimesNewRomanPSMT", 'a').is_some());
}
#[test]
fn non_base14_returns_none() {
assert_eq!(base14_char_width("DejaVuSans", 'a'), None);
assert!(!is_base14_font("Garamond"));
}
#[test]
fn builtin_encoding_resolves_symbol_and_zapf_codes() {
// Symbol 0x61 renders alpha; Zapf 0x21 renders U+2701.
assert_eq!(builtin_encoding_char("Symbol", 0x61), Some('\u{03B1}'));
assert_eq!(builtin_encoding_char("Symbol", 0xA5), Some('\u{221E}'));
assert_eq!(
builtin_encoding_char("ZapfDingbats", 0x21),
Some('\u{2701}')
);
// Latin text fonts follow standard encodings — no builtin override.
assert_eq!(builtin_encoding_char("Times-Roman", 0x61), None);
// The resolved chars have real AFM widths.
let alpha_w = base14_char_width("Symbol", '\u{03B1}');
assert!(alpha_w.is_some() && alpha_w != Some(500));
}
#[test]
fn encoding_tables_are_sorted_for_binary_search() {
for table in [SYMBOL_ENCODING, ZAPFDINGBATS_ENCODING] {
assert!(table.windows(2).all(|w| w[0].0 < w[1].0));
}
}
#[test]
fn tables_are_sorted_for_binary_search() {
// Every table is queried by binary search, so all of them must be
// sorted — not just a sample.
for (name, table) in ALL_TABLES {
assert!(
table.windows(2).all(|w| w[0].0 < w[1].0),
"{name} is not sorted"
);
}
}
}
+45 -7
View File
@@ -14,9 +14,9 @@ use lopdf::{Document, Encoding, Object, ObjectId};
use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, descriptor_style_flags,
extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
FontStyleCache,
build_font_encodings, build_font_widths, build_type3_scales, compute_string_width_ts,
descriptor_style_flags, extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes,
CMapDecisionCache, FontStyleCache,
};
use super::underline::UnderlineLine;
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
@@ -41,6 +41,17 @@ fn strip_pdf_comments(data: &[u8]) -> Vec<u8> {
while i < data.len() {
let b = data[i];
match b {
// Inside a string literal, a backslash escapes the next byte —
// `\(`, `\)`, and `\\` must not touch the nesting depth, or a
// later `%` glyph inside a string gets stripped as a comment,
// corrupting the stream.
b'\\' if in_string > 0 => {
result.push(b);
if let Some(&next) = data.get(i + 1) {
result.push(next);
i += 1;
}
}
b'(' if !in_hex_string => {
in_string += 1;
result.push(b);
@@ -166,6 +177,7 @@ pub(crate) fn extract_page_text_items(
// Build font width info for accurate text positioning
let font_widths = build_font_widths(doc, &fonts);
let type3_scales = build_type3_scales(doc, &fonts);
// Build maps of font resource names to their base font names and ToUnicode object refs
let mut font_base_names: std::collections::HashMap<String, String> =
@@ -502,7 +514,8 @@ pub(crate) fn extract_page_text_items(
) {
let combined =
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
@@ -673,7 +686,8 @@ pub(crate) fn extract_page_text_items(
} else {
rotation_votes.rotated += 1;
}
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let base_font = font_base_names
.get(&current_font)
.map(|s| s.as_str())
@@ -780,7 +794,8 @@ pub(crate) fn extract_page_text_items(
} else {
rotation_votes.rotated += 1;
}
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
let width = w_ts_opt
.map(|w_ts| {
@@ -932,7 +947,8 @@ pub(crate) fn extract_page_text_items(
} else {
rotation_votes.rotated += 1;
}
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
// Width in device space from text matrix delta
let delta_ts = text_matrix[4] - start_tm[4];
@@ -1849,4 +1865,26 @@ BT 30 700 Tm <41> Tj ET";
"ET should be preserved after comment stripping"
);
}
#[test]
fn test_strip_pdf_comments_escaped_parens() {
// An escaped `\)` must not close the string: the `%` after it is
// still string content, not a comment (subset fonts routinely map
// glyphs to `%` and to escaped parens in the same TJ array).
let input = b"[ (a\\)b) 1 (%) 1 (c) ] TJ\n";
let output = strip_pdf_comments(input);
assert_eq!(output, input.to_vec());
// Same for an escaped `\(` — must not open a phantom string that
// shields a real comment.
let input = b"(x\\(y) Tj % real comment\nET\n";
let output = strip_pdf_comments(input);
assert_eq!(output, b"(x\\(y) Tj \nET\n");
// Escaped backslash before a real close-paren: `\\` ends the escape,
// the `)` does close the string, and the comment is stripped.
let input = b"(x\\\\) Tj % comment\nET\n";
let output = strip_pdf_comments(input);
assert_eq!(output, b"(x\\\\) Tj \nET\n");
}
}
+234 -1
View File
@@ -138,6 +138,93 @@ pub(crate) fn build_font_widths(
widths
}
/// Visual-size scale factors for Type3 fonts, keyed by resource name.
///
/// A Type3 font's glyph space maps to text space through FontMatrix, so the
/// visual height of its glyphs is `nominal_size × |matrix_y| × FontBBox
/// height`. For a well-behaved font (matrix 0.001, bbox ≈ 1000 units) that
/// factor is ≈ 1.0 and the nominal size is already right. TeX PK bitmap
/// fonts (dvips → Distiller) instead use FontMatrix [1 0 0 -1 0 0] with
/// nominal sizes like 0.12, which makes every downstream font-size heuristic
/// (drop caps, sub/superscripts, small-font tables, line heights) see
/// nonsense. Fonts without a usable FontBBox are omitted (treated as 1.0).
pub(crate) fn build_type3_scales(
doc: &Document,
fonts: &std::collections::BTreeMap<Vec<u8>, &lopdf::Dictionary>,
) -> HashMap<String, f32> {
let mut scales = HashMap::new();
for (font_name, font_dict) in fonts {
let is_type3 = font_dict
.get(b"Subtype")
.ok()
.and_then(|o| o.as_name().ok())
.is_some_and(|n| n == b"Type3");
if !is_type3 {
continue;
}
// Array elements may themselves be indirect references per PDF
// syntax — resolve before reading the numeric value.
let num = |o: &Object| {
let resolved = match o {
Object::Reference(r) => match doc.get_object(*r) {
Ok(inner) => inner,
Err(_) => return 0.0,
},
other => other,
};
match resolved {
Object::Integer(i) => *i as f32,
Object::Real(r) => *r,
_ => 0.0,
}
};
let Some(matrix) = font_dict
.get(b"FontMatrix")
.ok()
.and_then(|o| resolve_array(doc, o))
else {
continue;
};
let Some(bbox) = font_dict
.get(b"FontBBox")
.ok()
.and_then(|o| resolve_array(doc, o))
else {
continue;
};
if matrix.len() < 4 || bbox.len() < 4 {
continue;
}
let scale_y = (num(&matrix[2]).powi(2) + num(&matrix[3]).powi(2)).sqrt();
let bbox_h = (num(&bbox[3]) - num(&bbox[1])).abs();
let scale = bbox_h * scale_y;
// `scale` is the glyph box measured in text-space units. For a
// self-consistent font it lands near 1.0 — the FontMatrix is the
// reciprocal of the glyph-space em by construction — so the Tf
// operand is already the rendered size and must be left alone.
// A modest deviation is normal and must NOT trigger rescaling:
// FontBBox is the glyph bounding box, not the em box, so it is
// routinely somewhat smaller (descender..ascender ≈ 0.7) or larger
// (tall accents > 1.0).
//
// Only a wildly inconsistent font gets renormalized. dvips/PK
// bitmap fonts declare [1 0 0 -1 0 0] with glyphs spanning
// hundreds of units, giving scale ≈ 159 against a nominal size of
// 0.12pt — there the declared size carries no information. The
// band is deliberately wide so that only that class qualifies,
// while any matrix scale (including non-standard ones like 0.005
// with a full-em bbox, scale = 5.0) is judged on the product
// rather than on the matrix alone.
const CONSISTENT_LO: f32 = 0.25;
const CONSISTENT_HI: f32 = 4.0;
if scale.is_finite() && scale > 0.0 && !(CONSISTENT_LO..=CONSISTENT_HI).contains(&scale) {
scales.insert(String::from_utf8_lossy(font_name).to_string(), scale);
}
}
scales
}
/// Parse font widths from a font dictionary, dispatching by Subtype
pub(crate) fn parse_font_widths(
doc: &Document,
@@ -149,11 +236,71 @@ pub(crate) fn parse_font_widths(
match subtype_name {
b"Type0" => parse_type0_widths(doc, font_dict),
b"Type1" | b"TrueType" | b"MMType1" | b"Type3" => parse_simple_font_widths(doc, font_dict),
b"Type1" | b"TrueType" | b"MMType1" => parse_simple_font_widths(doc, font_dict)
.or_else(|| base14_fallback_widths(doc, font_dict)),
b"Type3" => parse_simple_font_widths(doc, font_dict),
_ => None,
}
}
/// Fallback metrics for non-embedded base-14 fonts whose dictionary omits
/// `/FirstChar`/`/Widths` (legal per the PDF spec — the reader must supply
/// standard-font metrics). Without this, every glyph advances 0 and all
/// downstream gap-based logic (space synthesis, script detection, table
/// columns) collapses — common in 1990s dvips/Distiller PDFs.
///
/// Widths are resolved per code through the font's Differences encoding when
/// present, falling back to the same single-byte decode the text extractor
/// uses (cp1252-style smart punctuation for 0x80..=0x9F, Latin-1 elsewhere) —
/// so the width of a code always matches the char we extract for it.
fn base14_fallback_widths(doc: &Document, font_dict: &lopdf::Dictionary) -> Option<FontWidthInfo> {
let base_font = font_dict
.get(b"BaseFont")
.ok()
.and_then(|o| o.as_name().ok())
.map(|n| String::from_utf8_lossy(n).to_string())?;
if !crate::extractor::base14::is_base14_font(&base_font) {
return None;
}
let enc_map = parse_font_encoding(doc, font_dict)
.map(|r| r.map)
.unwrap_or_default();
let mut widths = HashMap::new();
for code in 0u16..=255 {
// Resolution order: Differences override, then the font's BUILT-IN
// encoding (Symbol/ZapfDingbats glyphs live at positions unrelated
// to cp1252 — the renderer draws α for Symbol 0x61 no matter how
// the text decoder transliterates it, so the advance must be α's),
// then the cp1252-style fallback used by the text decoder.
let ch = enc_map
.get(&(code as u8))
.copied()
.or_else(|| crate::extractor::base14::builtin_encoding_char(&base_font, code as u8))
.unwrap_or_else(|| decode_single_byte_fallback_char(code as u8, true));
if let Some(w) = crate::extractor::base14::base14_char_width(&base_font, ch) {
widths.insert(code, w);
}
}
let space_width = widths.get(&32).copied().unwrap_or(250);
debug!(
" base14 fallback widths for {} ({} codes mapped)",
base_font,
widths.len()
);
Some(FontWidthInfo {
widths,
default_width: 500,
space_width,
is_cid: false,
units_scale: 0.001,
wmode: 0,
})
}
/// Parse widths for simple fonts (Type1, TrueType, MMType1, Type3)
/// Reads FirstChar, LastChar, and Widths array.
/// For Type3 fonts, reads FontMatrix to determine the correct units_scale.
@@ -1470,6 +1617,92 @@ fn score_text(text: &str) -> i32 {
#[cfg(test)]
mod tests {
#[test]
fn type3_scale_resolves_indirect_matrix_and_bbox_numbers() {
use lopdf::{dictionary, Document, Object};
// FontMatrix/FontBBox elements may be indirect references per PDF
// syntax; the scale must use their resolved values, not zero.
let mut doc = Document::with_version("1.4");
let matrix_d = doc.add_object(Object::Real(-1.0));
let bbox_top = doc.add_object(Object::Integer(3));
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "Type3",
"FontMatrix" => vec![
Object::Integer(1),
Object::Integer(0),
Object::Integer(0),
Object::Reference(matrix_d),
Object::Integer(0),
Object::Integer(0),
],
"FontBBox" => vec![
Object::Integer(1),
Object::Integer(-156),
Object::Integer(37),
Object::Reference(bbox_top),
],
};
let mut fonts = std::collections::BTreeMap::new();
fonts.insert(b"T2".to_vec(), &font_dict);
let scales = super::build_type3_scales(&doc, &fonts);
let scale = scales.get("T2").copied().unwrap_or(1.0);
// bbox height 159 x |matrix_y| 1.0
assert!(
(scale - 159.0).abs() < 0.5,
"scale should use resolved indirect values, got {scale}"
);
}
/// Build a one-font Type3 document and return its computed scale, if any.
#[cfg(test)]
fn type3_scale_for(matrix_y: f32, bbox_lo: i64, bbox_hi: i64) -> Option<f32> {
use lopdf::{dictionary, Document, Object};
let doc = Document::with_version("1.4");
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "Type3",
"FontMatrix" => vec![
Object::Real(matrix_y), Object::Integer(0), Object::Integer(0),
Object::Real(matrix_y), Object::Integer(0), Object::Integer(0),
],
"FontBBox" => vec![
Object::Integer(0), Object::Integer(bbox_lo),
Object::Integer(600), Object::Integer(bbox_hi),
],
};
let mut fonts = std::collections::BTreeMap::new();
fonts.insert(b"T9".to_vec(), &font_dict);
super::build_type3_scales(&doc, &fonts).get("T9").copied()
}
#[test]
fn type3_scale_skips_self_consistent_fonts() {
// Conventional 1/1000 matrix with a descender..ascender bbox of 700
// units: scale 0.7. The Tf operand is already the rendered size, so
// renormalizing would report every size at 0.7x.
assert_eq!(type3_scale_for(0.001, -200, 500), None);
// Tall-accent bbox slightly over the em (1100 units, scale 1.1).
assert_eq!(type3_scale_for(0.001, -100, 1000), None);
}
#[test]
fn type3_scale_applies_to_inconsistent_fonts_at_any_matrix_scale() {
// Non-standard but valid matrix (0.005) with a full-em bbox:
// scale 5.0, so the declared size is off by 5x and must be fixed.
let s = type3_scale_for(0.005, 0, 1000).expect("0.005 matrix should rescale");
assert!((s - 5.0).abs() < 0.01, "got {s}");
// dvips/PK bitmap pattern: unit matrix, glyphs spanning ~159 units.
let s = type3_scale_for(1.0, -156, 3).expect("PK pattern should rescale");
assert!((s - 159.0).abs() < 0.5, "got {s}");
}
#[test]
fn type3_scale_ignores_degenerate_bbox() {
// [0 0 0 0] is legal and carries no size information.
assert_eq!(type3_scale_for(0.001, 0, 0), None);
}
#[test]
fn texcm_math_symbols_remap() {
assert_eq!(
+327 -17
View File
@@ -41,15 +41,110 @@ pub(crate) fn detect_columns(
}
debug!("page {}: detect_columns: {} items", page, page_items.len());
// Find page bounds
let x_min = page_items.iter().map(|i| i.x).fold(f32::INFINITY, f32::min);
let x_max = page_items
.iter()
.map(|i| i.x + effective_width(i))
.fold(f32::NEG_INFINITY, f32::max);
// The width of one ordinary page, used three ways below: as the largest
// credible width for a single text run, as the size of empty gap that marks
// content as detached, and as the span past which those checks run at all.
// This is a heuristic, not a format rule: PDF 2.0 sets no page-size limit,
// and since PDF 1.6 `UserUnit` scales a page's physical size independently
// of its coordinates. 14_400 units (200in at the default 1/72in unit) is
// the traditional Acrobat architectural limit, which makes it a reasonable
// "wider than any ordinary page" mark in coordinate space.
const MAX_PAGE_EXTENT: f32 = 14_400.0;
// A detached cluster is only dropped if it also holds a small minority of
// the items, so a genuine two-part layout keeps its full bounds even when
// the halves are far apart.
const MAX_TRIM_FRACTION: f32 = 0.10;
// Position and width of each item, skipping only non-finite geometry.
let finite_span = |i: &&TextItem| -> Option<(f32, f32)> {
let (left, width) = (i.x, effective_width(i));
(left.is_finite() && (left + width).is_finite()).then_some((left, width))
};
let (min_left, max_right, total) = page_items.iter().filter_map(finite_span).fold(
(f32::INFINITY, f32::NEG_INFINITY, 0usize),
|(lo, hi, n), (left, width)| (lo.min(left), hi.max(left + width), n + 1),
);
// No item had usable geometry, so there is no layout to report.
if total == 0 {
return vec![];
}
// Every threshold below (gutter margins, spanning-item width, the XY-cut
// margin) is a fraction of the page width, so a far item can set the scale
// for the whole page and shrink the effective detection window to a
// rounding error — real gutters then fall inside the margin band and a
// genuine multi-column page collapses to one region.
//
// Anything inside one page extent is ordinary, so the common case keeps the
// plain bounds and skips the work below entirely.
let (x_min, x_max) = if max_right - min_left <= MAX_PAGE_EXTENT {
(min_left, max_right)
} else {
// Discarding content needs positive evidence that it is not part of the
// layout, because a count-based rule alone cannot tell a stray from a
// sparse far sidebar. The evidence is geometric: positions are grouped
// into clusters separated by more than a whole page of continuous
// emptiness. Real content, however sparse, does not leave a void that
// large; a malformed coordinate sits alone beyond one.
let mut spans: Vec<(f32, f32)> = page_items.iter().filter_map(finite_span).collect();
spans.sort_by(|a, b| a.0.total_cmp(&b.0));
let mut core: Option<std::ops::Range<usize>> = None;
let mut start = 0usize;
for i in 1..=spans.len() {
if i < spans.len() && spans[i].0 - spans[i - 1].0 <= MAX_PAGE_EXTENT {
continue;
}
if core.as_ref().is_none_or(|best| i - start > best.len()) {
core = Some(start..i);
}
start = i;
}
let mut core = core.unwrap_or(0..spans.len());
// Only drop the detached clusters when they are a small minority, so a
// genuine two-part layout keeps its full bounds.
let dropped = spans.len() - core.len();
if dropped as f32 > spans.len() as f32 * MAX_TRIM_FRACTION {
core = 0..spans.len();
}
let core = &spans[core];
// Positions cannot be inflated by a bogus width, so the spread of the
// content is a sound scale for judging one. A run much wider than the
// page's own content is a malformed width — the test is relative, so a
// genuinely large page keeps its genuinely long runs.
let (lo, widest_left) = (core[0].0, core[core.len() - 1].0);
let max_run_width = (widest_left - lo) + MAX_PAGE_EXTENT;
let hi = core
.iter()
.filter(|&&(_, width)| width <= max_run_width)
.map(|&(left, width)| left + width)
.fold(widest_left, f32::max);
if lo != min_left || hi != max_right {
debug!(
"page {page}: bounds {min_left}..{max_right} exceed one page; \
dropped {dropped}/{} detached item(s), using {lo}..{hi}",
spans.len()
);
}
(lo, hi)
};
// Hard ceiling on the histogram size, independent of the trimming above:
// the bounds are attacker-influenced, so an unclamped
// `page_width / BIN_WIDTH` lets a crafted PDF force an arbitrarily large
// `vec![0u32; num_bins]` allocation. 65_536 bins covers ~128k points at
// BIN_WIDTH 2.0 — roughly 9x the largest legal page — so this never binds
// on a real layout. Kept as a bound that does not depend on the outlier
// heuristic staying correct.
const MAX_BINS: usize = 65_536;
let page_width = x_max - x_min;
if page_width < 200.0 {
if !page_width.is_finite() || page_width < 200.0 {
return vec![ColumnRegion { x_min, x_max }];
}
@@ -57,13 +152,20 @@ pub(crate) fn detect_columns(
return vec![ColumnRegion { x_min, x_max }];
}
// Widen the bins rather than dropping the tail of the page. Clamping the
// count alone would leave anything past MAX_BINS * BIN_WIDTH outside the
// histogram, folded into the last bin, which places gutters at the wrong
// coordinates. Scaling keeps full coverage under the same allocation
// ceiling; only the resolution degrades, and only beyond ~131k points.
let bin_width = BIN_WIDTH.max(page_width / MAX_BINS as f32);
// Build occupancy histogram.
// Exclude items wider than 60% of page width — these are spanning items
// (titles, full-width paragraphs) that would fill the gutter and prevent
// detection of partial-page column layouts (e.g. two-column abstracts on
// a page that also has single-column introduction text).
let wide_threshold = page_width * 0.6;
let num_bins = ((page_width / BIN_WIDTH).ceil() as usize).max(1);
let num_bins = ((page_width / bin_width).ceil() as usize).clamp(1, MAX_BINS);
let mut histogram = vec![0u32; num_bins];
for item in &page_items {
@@ -71,8 +173,8 @@ pub(crate) fn detect_columns(
if w > wide_threshold {
continue;
}
let left = ((item.x - x_min) / BIN_WIDTH).floor() as usize;
let right = (((item.x + w) - x_min) / BIN_WIDTH).ceil() as usize;
let left = ((item.x - x_min) / bin_width).floor() as usize;
let right = (((item.x + w) - x_min) / bin_width).ceil() as usize;
let left = left.min(num_bins);
let right = right.min(num_bins);
for count in histogram.iter_mut().take(right).skip(left) {
@@ -109,12 +211,12 @@ pub(crate) fn detect_columns(
let valleys: Vec<(usize, usize)> = valleys
.into_iter()
.filter(|&(start, end)| {
let width_pts = (end - start) as f32 * BIN_WIDTH;
let width_pts = (end - start) as f32 * bin_width;
if width_pts < MIN_GUTTER_WIDTH {
return false;
}
// Valley center must not be within 5% of page edges
let center_pts = ((start + end) as f32 / 2.0) * BIN_WIDTH;
let center_pts = ((start + end) as f32 / 2.0) * bin_width;
center_pts > margin_threshold && center_pts < (page_width - margin_threshold)
})
.collect();
@@ -132,7 +234,7 @@ pub(crate) fn detect_columns(
&histogram,
num_bins,
x_min,
BIN_WIDTH,
bin_width,
page_width,
margin_threshold,
);
@@ -141,7 +243,7 @@ pub(crate) fn detect_columns(
&rel_valleys,
&page_items,
x_min,
BIN_WIDTH,
bin_width,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
@@ -182,7 +284,7 @@ pub(crate) fn detect_columns(
&valleys,
&page_items,
x_min,
BIN_WIDTH,
bin_width,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
@@ -196,7 +298,7 @@ pub(crate) fn detect_columns(
&valleys,
&page_items,
x_min,
BIN_WIDTH,
bin_width,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
@@ -1827,7 +1929,7 @@ fn split_column_stragglers(lines: Vec<TextLine>) -> (Vec<TextLine>, Vec<TextLine
.unwrap();
let (cs, ce) = segments[core_seg];
let mut core = Vec::with_capacity(ce - cs);
let mut core = Vec::with_capacity(ce.saturating_sub(cs));
let mut stragglers = Vec::new();
for (i, line) in lines.into_iter().enumerate() {
if i >= cs && i < ce {
@@ -2533,6 +2635,214 @@ mod tests {
);
}
#[test]
fn extreme_far_coordinate_does_not_allocate_unboundedly() {
// A crafted PDF can place a text run at an arbitrary coordinate via the
// text matrix. The derived page width must not drive an unbounded
// histogram allocation (previously `page_width / BIN_WIDTH` bins with no
// upper bound would try to reserve terabytes and abort the process).
let mut items = Vec::new();
for i in 0..24 {
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
}
// Item placed 1e12 points away — 5e11 bins if left unclamped.
items.push(make_item(1, 1e12, 700.0, "Z"));
// Must return without aborting; content is preserved as a single region.
let cols = detect_columns(&items, 1, false);
assert!(!cols.is_empty());
}
#[test]
fn non_finite_coordinates_never_leak_into_region_bounds() {
// An inf/NaN coordinate must not escape as a column boundary: callers
// treat these as page/column edges.
for bad_x in [f32::INFINITY, f32::NEG_INFINITY, f32::NAN] {
let mut items = Vec::new();
for i in 0..24 {
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
}
items.push(make_item(1, bad_x, 700.0, "Z"));
for col in detect_columns(&items, 1, false) {
assert!(
col.x_min.is_finite() && col.x_max.is_finite(),
"bad_x {bad_x} leaked bounds {}..{}",
col.x_min,
col.x_max
);
}
}
}
#[test]
fn all_non_finite_coordinates_yield_no_columns() {
let items: Vec<TextItem> = (0..24)
.map(|i| make_item(1, f32::NAN, 700.0 - i as f32 * 5.0, "A"))
.collect();
assert!(detect_columns(&items, 1, false).is_empty());
}
#[test]
fn one_bad_item_does_not_disable_column_detection() {
// A single stray item should not collapse a clean two-column page to
// one region. Every gutter threshold is a fraction of the page width,
// so an untrimmed outlier pushes real gutters inside the rejected
// margin band. A malformed *width* at an ordinary position poisons the
// bounds just as a malformed position does.
for (label, bad_x, bad_width) in [
("nan position", f32::NAN, 0.0),
("inf position", f32::INFINITY, 0.0),
("far position", 50_000.0, 0.0),
("very far position", 1e12, 0.0),
("huge width", 100.0, 1e12),
("inf width", 100.0, f32::INFINITY),
] {
let mut items = Vec::new();
items.extend(fill_zone(1, 30.0, 280.0, 750.0, 50.0));
items.extend(fill_zone(1, 320.0, 570.0, 750.0, 50.0));
let mut bad = make_item(1, bad_x, 400.0, "Z");
bad.width = bad_width;
items.push(bad);
let cols = detect_columns(&items, 1, false);
assert_eq!(
cols.len(),
2,
"{label}: expected 2 columns, got {}",
cols.len()
);
for col in &cols {
assert!(
col.x_max - col.x_min <= MAX_PAGE_EXTENT_FOR_TEST,
"{label}: region {}..{} exceeds one page",
col.x_min,
col.x_max
);
}
}
}
/// Mirrors `MAX_PAGE_EXTENT` in `detect_columns`.
const MAX_PAGE_EXTENT_FOR_TEST: f32 = 14_400.0;
#[test]
fn very_wide_page_keeps_full_histogram_coverage() {
// Beyond MAX_BINS * BIN_WIDTH (~131k points) the bins must widen rather
// than stop covering the page. Three zones: the first gutter is inside
// the old coverage limit, the second is past it. Because the first
// gutter is found, the XY-cut fallback never runs, so a truncated
// histogram silently reports two columns instead of three.
let mut items = Vec::new();
items.extend(fill_zone(1, 0.0, 60_000.0, 750.0, 700.0));
items.extend(fill_zone(1, 70_000.0, 140_000.0, 750.0, 700.0));
items.extend(fill_zone(1, 160_000.0, 200_000.0, 750.0, 700.0));
let cols = detect_columns(&items, 1, false);
assert_eq!(
cols.len(),
3,
"Expected 3 columns across a 200k-wide page, got {}",
cols.len()
);
assert!(
(140_000.0..=160_000.0).contains(&cols[1].x_max),
"second gutter at {}, expected inside the real 140k..160k gap",
cols[1].x_max
);
}
#[test]
fn large_page_with_legitimately_long_runs_is_kept() {
// On a very large page, individual runs can exceed one ordinary page's
// width. They are real content, so they must not be judged malformed:
// the page keeps its columns and its full right edge.
let mut items = Vec::new();
for row in 0..30 {
let y = 750.0 - row as f32 * 14.0;
let mut left = make_item(1, 0.0, y, "Left run");
left.width = 20_000.0;
let mut right = make_item(1, 25_000.0, y, "Right run");
right.width = 20_000.0;
items.extend([left, right]);
}
let cols = detect_columns(&items, 1, false);
assert!(
!cols.is_empty(),
"a page of long-but-valid runs must still report a layout"
);
let right_edge = cols
.iter()
.map(|c| c.x_max)
.fold(f32::NEG_INFINITY, f32::max);
assert!(
right_edge > 44_000.0,
"long runs were treated as malformed: right edge {right_edge}, expected ~45_000"
);
}
#[test]
fn sparse_far_sidebar_on_a_large_page_is_kept() {
// A large-format page with a thin, sparsely-populated sidebar far from
// the main block. The sidebar is a small minority of the items, so an
// item-count rule alone would discard it — but nothing about its
// geometry says it is invalid, so its bounds must survive.
let mut items = Vec::new();
items.extend(fill_zone(1, 0.0, 12_000.0, 750.0, 500.0));
for i in 0..12 {
items.push(make_item(1, 24_000.0, 750.0 - i as f32 * 14.0, "Sidebar"));
}
let cols = detect_columns(&items, 1, false);
let right_edge = cols
.iter()
.map(|c| c.x_max)
.fold(f32::NEG_INFINITY, f32::max);
assert!(
right_edge > 24_000.0,
"sidebar was trimmed away: right edge {right_edge}, expected >24_000"
);
}
#[test]
fn genuinely_wide_layout_keeps_its_true_bounds() {
// A large-format page whose content really is spread beyond one
// ordinary page must not be trimmed to the median cluster: its far
// items are the majority, not strays.
let mut items = Vec::new();
items.extend(fill_zone(1, 100.0, 20_000.0, 750.0, 600.0));
items.extend(fill_zone(1, 22_000.0, 40_000.0, 750.0, 600.0));
let cols = detect_columns(&items, 1, false);
let widest = cols
.iter()
.map(|c| c.x_max)
.fold(f32::NEG_INFINITY, f32::max);
assert!(
widest > 35_000.0,
"wide layout was trimmed: right edge {widest}, expected ~40_000"
);
}
#[test]
fn oversized_but_legal_page_is_not_trimmed() {
// A wide-format page well inside the 14_400pt spec limit must keep its
// real bounds — outlier trimming is only for spans beyond a legal page.
let mut items = Vec::new();
items.extend(fill_zone(1, 100.0, 4_000.0, 750.0, 400.0));
items.extend(fill_zone(1, 4_400.0, 8_000.0, 750.0, 400.0));
let cols = detect_columns(&items, 1, false);
assert_eq!(cols.len(), 2, "Expected 2 columns, got {}", cols.len());
assert!(
cols[1].x_max > 7_000.0,
"right column should keep its true extent, got {}",
cols[1].x_max
);
}
#[test]
fn two_column_regression_guard() {
// Standard 2-column layout with clear gutter at center
+311 -5
View File
@@ -2,11 +2,51 @@
use crate::types::{ItemType, TextItem};
use lopdf::{Document, Object, ObjectId};
use std::collections::HashMap;
use std::collections::{HashMap, HashSet};
use super::fonts::{resolve_array, resolve_dict};
use super::get_number;
/// Upper bound on the number of form-field nodes visited during a single
/// `extract_form_fields` pass. A crafted PDF can chain thousands of distinct
/// `/Kids` fields to blow the stack even without an outright reference cycle,
/// so we cap total traversal work in addition to detecting cycles.
const MAX_FORM_FIELD_NODES: usize = 100_000;
/// Upper bound on `/Kids` recursion depth. Real AcroForm hierarchies are only
/// a few levels deep (fields → child fields → widgets); a crafted PDF can chain
/// tens of thousands of distinct fields into a linear `/Kids` list that would
/// overflow the stack via depth-first recursion long before the node budget is
/// reached. This depth cap bounds the stack independently of total node count.
const MAX_FORM_FIELD_DEPTH: usize = 100;
/// Traversal budget for the AcroForm field walk. Bounds both the number of
/// distinct nodes visited *and* the total number of `/Fields`/`/Kids` entries
/// examined.
///
/// Counting `visited` alone is not enough: invalid entries (non-references) and
/// duplicate references never grow `visited`, so an oversized array full of them
/// would iterate to completion no matter how large. Charging every examined
/// entry against the same budget makes it a real cap on traversal work.
pub(crate) struct FieldWalkBudget {
visited: HashSet<ObjectId>,
examined: usize,
}
impl FieldWalkBudget {
fn new() -> Self {
Self {
visited: HashSet::new(),
examined: 0,
}
}
/// True once the budget is spent; callers must stop iterating and recursing.
fn exhausted(&self) -> bool {
self.visited.len() >= MAX_FORM_FIELD_NODES || self.examined >= MAX_FORM_FIELD_NODES
}
}
pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> Vec<TextItem> {
let mut links = Vec::new();
@@ -146,9 +186,12 @@ pub(crate) fn extract_form_fields(
Err(_) => return items,
};
// Borrow the array rather than cloning it: a crafted `/Fields` can be huge,
// and cloning would pay an O(n) allocation/copy before the budget check
// below can stop the work.
let fields = match acroform.get(b"Fields") {
Ok(obj) => match resolve_array(doc, obj) {
Some(arr) => arr.clone(),
Some(arr) => arr,
None => return items,
},
Err(_) => return items,
@@ -158,7 +201,19 @@ pub(crate) fn extract_form_fields(
}
let annotation_pages = annotation_page_map(doc, page_map);
for field_obj in &fields {
// Bound the walk so a crafted PDF cannot send us into unbounded recursion
// via a `/Kids` cycle, a deep chain, or an oversized array of invalid or
// duplicate entries.
let mut budget = FieldWalkBudget::new();
for field_obj in fields {
// Stop once the budget is spent so a `/Fields` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op. Charge
// every entry (including invalid ones) against the budget.
if budget.exhausted() {
break;
}
budget.examined += 1;
if let Ok(field_ref) = field_obj.as_reference() {
walk_form_fields(
doc,
@@ -168,6 +223,8 @@ pub(crate) fn extract_form_fields(
page_map,
&annotation_pages,
&mut items,
&mut budget,
0,
);
}
}
@@ -202,6 +259,7 @@ fn annotation_page_map(
}
/// Recursively walk the form field tree, extracting leaf field values.
#[allow(clippy::too_many_arguments)]
pub(crate) fn walk_form_fields(
doc: &Document,
field_id: ObjectId,
@@ -210,7 +268,22 @@ pub(crate) fn walk_form_fields(
page_map: &HashMap<ObjectId, u32>,
annotation_pages: &HashMap<ObjectId, u32>,
items: &mut Vec<TextItem>,
budget: &mut FieldWalkBudget,
depth: usize,
) {
// Guard against `/Kids` cycles and pathologically large field trees.
// Exceeding the depth cap means the chain is too deep to be a legitimate
// form (and would overflow the stack); an exhausted budget means the tree is
// too large. Both checks run *before* inserting so the visited set can never
// grow past the budget.
if depth > MAX_FORM_FIELD_DEPTH || budget.exhausted() {
return;
}
// Revisiting an object ID means we hit a `/Kids` cycle.
if !budget.visited.insert(field_id) {
return;
}
let field_dict = match doc.get_dictionary(field_id) {
Ok(d) => d,
Err(_) => return,
@@ -241,9 +314,19 @@ pub(crate) fn walk_form_fields(
// Check for /Kids — if present, recurse into children
if let Ok(kids_obj) = field_dict.get(b"Kids") {
// Iterate the borrowed array directly — cloning a crafted, oversized
// `/Kids` would allocate and copy every entry before the budget check
// below could stop the work.
if let Some(kids) = resolve_array(doc, kids_obj) {
let kids = kids.clone();
for kid in &kids {
for kid in kids {
// Stop once the budget is spent so a `/Kids` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op.
// Charge every entry (including invalid/duplicate ones) against
// the budget so this is a true traversal-work cap.
if budget.exhausted() {
break;
}
budget.examined += 1;
if let Ok(kid_ref) = kid.as_reference() {
walk_form_fields(
doc,
@@ -253,6 +336,8 @@ pub(crate) fn walk_form_fields(
page_map,
annotation_pages,
items,
budget,
depth + 1,
);
}
}
@@ -411,4 +496,225 @@ mod tests {
assert_eq!(items[0].page, 2);
assert_eq!(items[0].text, "customer: Alice");
}
#[test]
fn kids_self_cycle_does_not_overflow_stack() {
// A crafted AcroForm field that lists itself in `/Kids` must not send
// the traversal into unbounded recursion.
let mut doc = Document::new();
let field_id = doc.new_object_id();
doc.set_object(
field_id,
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("loop"),
"Kids" => vec![Object::Reference(field_id)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
// Completes (rather than overflowing the stack) and yields no items.
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn kids_mutual_cycle_terminates() {
// Two fields that reference each other via `/Kids` form a cycle that
// must also terminate.
let mut doc = Document::new();
let field_a = doc.new_object_id();
let field_b = doc.new_object_id();
doc.set_object(
field_a,
dictionary! {
"T" => Object::string_literal("a"),
"Kids" => vec![Object::Reference(field_b)],
},
);
doc.set_object(
field_b,
dictionary! {
"T" => Object::string_literal("b"),
"Kids" => vec![Object::Reference(field_a)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_a)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn deep_acyclic_kids_chain_does_not_overflow_stack() {
// A long chain of *distinct* fields (no cycle) must also terminate:
// the visited set alone would still recurse to the chain length, so
// the depth cap is what prevents a stack overflow here.
let mut doc = Document::new();
let n = MAX_FORM_FIELD_DEPTH * 500;
let ids: Vec<ObjectId> = (0..=n).map(|_| doc.new_object_id()).collect();
for i in 0..n {
doc.set_object(
ids[i],
dictionary! {
"FT" => "Tx",
"Kids" => vec![Object::Reference(ids[i + 1])],
},
);
}
// Leaf carries a value; it sits far below the depth cap so it is never
// reached, proving traversal stops early rather than crashing.
doc.set_object(
ids[n],
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("leaf"),
"V" => Object::string_literal("x"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(ids[0])],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn wide_tree_traversal_stops_at_node_budget() {
// A single field with a `/Kids` array wider than the node budget must
// stop traversal at the cap rather than growing `visited` (and the work)
// without bound. Each processed leaf emits one item, so the item count
// is bounded by the budget and reaches right up to it (a couple of
// slots go to the root and the boundary node charged against the cap).
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let kids: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
// Extraction stops at the budget: bounded above by the cap, and it gets
// right up to it (allowing a small delta for the root/boundary nodes
// charged against the budget).
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn wide_top_level_fields_stop_at_node_budget() {
// A top-level `/Fields` array wider than the budget must also stop at
// the cap: the item count is bounded by the budget and reaches right up
// to it.
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let fields: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => fields,
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn duplicate_and_invalid_kids_entries_stop_at_budget() {
// Duplicate references and non-reference junk never grow `visited`, so
// without charging examined entries against the budget an oversized
// array of them would iterate to completion. The walk must still
// terminate and extract the single real leaf exactly once.
let mut doc = Document::new();
let leaf_id = doc.new_object_id();
doc.set_object(
leaf_id,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
// A `/Kids` array far wider than the budget: half duplicate references
// to the same leaf, half invalid (null) entries.
let mut kids: Vec<Object> = Vec::new();
for i in 0..(MAX_FORM_FIELD_NODES * 2) {
if i % 2 == 0 {
kids.push(Object::Reference(leaf_id));
} else {
kids.push(Object::Null);
}
}
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert_eq!(items.len(), 1);
}
}
+21 -4
View File
@@ -2,6 +2,7 @@
//!
//! This module extracts text with position information for structure detection.
mod base14;
pub(crate) mod content_stream;
mod fonts;
mod layout;
@@ -84,17 +85,33 @@ pub fn extract_text_with_positions_pages<P: AsRef<Path>>(
path: P,
page_filter: Option<&HashSet<u32>>,
) -> Result<Vec<TextItem>, PdfError> {
let (items, _rects, _lines) = extract_text_with_positions_and_rects(path, page_filter)?;
let (items, _rects, _lines) =
extract_text_with_positions_and_rects_with_password(path, page_filter, None)?;
Ok(items)
}
/// Extract text with positions and rectangles from a file.
pub(crate) fn extract_text_with_positions_and_rects<P: AsRef<Path>>(
/// Extract text with positions from a file, limited to specific pages and
/// decrypting with `password` when the PDF is encrypted.
///
/// `page_filter` is an optional set of 1-indexed page numbers to process.
/// When `None`, all pages are processed.
pub fn extract_text_with_positions_pages_with_password<P: AsRef<Path>>(
path: P,
page_filter: Option<&HashSet<u32>>,
password: Option<&str>,
) -> Result<Vec<TextItem>, PdfError> {
let (items, _rects, _lines) =
extract_text_with_positions_and_rects_with_password(path, page_filter, password)?;
Ok(items)
}
pub(crate) fn extract_text_with_positions_and_rects_with_password<P: AsRef<Path>>(
path: P,
page_filter: Option<&HashSet<u32>>,
password: Option<&str>,
) -> Result<PageExtraction, PdfError> {
crate::validate_pdf_file(&path)?;
let (doc, _) = crate::load_document_from_path(&path)?;
let (doc, _) = crate::load_document_from_path_with_password(&path, password)?;
let font_cmaps = FontCMaps::from_doc(&doc);
let (extraction, _thresholds, _gid_pages) =
extract_positioned_text_from_doc(&doc, &font_cmaps, page_filter)?;
+8 -4
View File
@@ -8,8 +8,9 @@ use lopdf::{Document, Encoding, Object, ObjectId};
use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache, FontStyleCache,
build_font_encodings, build_font_widths, build_type3_scales, compute_string_width_ts,
extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
FontStyleCache,
};
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
@@ -166,6 +167,7 @@ fn extract_form_xobject_text_inner(
// Build font width info for the form
let font_widths = build_font_widths(doc, &form_fonts);
let type3_scales = build_type3_scales(doc, &form_fonts);
// Build font base names and ToUnicode refs for the form
let mut font_base_names: HashMap<String, String> = HashMap::new();
@@ -413,7 +415,8 @@ fn extract_form_xobject_text_inner(
&font_widths,
) {
let combined = multiply_matrices(&text_matrix, &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
let width = if let Some(font_info) = font_widths.get(&current_font) {
if let Some(raw_bytes) = get_operand_bytes(&op.operands[0]) {
@@ -572,7 +575,8 @@ fn extract_form_xobject_text_inner(
}
if !sub_items.is_empty() {
let combined = multiply_matrices(&text_matrix, &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let base_font = font_base_names
.get(&current_font)
.map(|s| s.as_str())
+40 -3
View File
@@ -4566,9 +4566,13 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
}
}
// Try to parse uniXXXX format
if name.starts_with("uni") && name.len() >= 7 {
if let Ok(code) = u32::from_str_radix(&name[3..7], 16) {
// Try to parse uniXXXX format.
// Use `get` rather than a byte-length check + slice: `name` can contain
// non-ASCII bytes (e.g. U+FFFD from lossy UTF-8 decoding of an attacker
// controlled /Differences name), so byte index 7 may not be a char
// boundary and `&name[3..7]` would panic.
if let Some(hex) = name.strip_prefix("uni").and_then(|rest| rest.get(..4)) {
if let Ok(code) = u32::from_str_radix(hex, 16) {
// Strip PUA F000 offset: uniF0XX → U+00XX (Windows Symbol encoding convention)
let code = if (0xF000..=0xF0FF).contains(&code) {
code - 0xF000
@@ -4588,3 +4592,36 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
None
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn uni_hex_parsing() {
assert_eq!(glyph_to_char("uni0041"), Some('A'));
assert_eq!(glyph_to_char("uni00e9"), Some('\u{00e9}'));
// PUA F0xx symbol-encoding offset is stripped.
assert_eq!(glyph_to_char("uniF041"), Some('A'));
}
#[test]
fn u_hex_parsing() {
assert_eq!(glyph_to_char("u0041"), Some('A'));
assert_eq!(glyph_to_char("u1F600"), Some('\u{1F600}'));
}
#[test]
fn non_ascii_uni_name_does_not_panic() {
// A crafted /Differences name like `/uni#80#80#80#80` decodes via
// from_utf8_lossy into "uni" followed by four U+FFFD replacements.
// Byte index 7 lands mid-character, so a naive `&name[3..7]` slice
// would panic. It must be handled gracefully instead.
let crafted = format!("uni{0}{0}{0}{0}", '\u{FFFD}');
assert_eq!(glyph_to_char(&crafted), None);
// Assorted non-ASCII bytes right after the "uni" prefix.
assert_eq!(glyph_to_char("uni\u{FFFD}bc"), None);
assert_eq!(glyph_to_char("uni\u{00e9}00"), None);
}
}
+386 -8
View File
@@ -50,7 +50,7 @@ pub use detector::{
};
pub use extractor::{
extract_text, extract_text_with_positions, extract_text_with_positions_mem,
extract_text_with_positions_pages,
extract_text_with_positions_pages, extract_text_with_positions_pages_with_password,
};
pub use markdown::{
to_markdown, to_markdown_from_items, to_markdown_from_items_with_rects,
@@ -491,7 +491,14 @@ pub fn extract_pages_markdown_mem(
// Tables need the original numeric cells; columns use folio-cleaned
// evidence so removed page numbers cannot create false layout metadata.
let complexity = compute_layout_complexity(&all_items, &filtered_items, &all_rects, &all_lines);
let chart_regions = markdown::chart_regions_by_page(&all_items, &all_rects, &all_lines);
let complexity = compute_layout_complexity_with_chart_regions(
&all_items,
&filtered_items,
&all_rects,
&all_lines,
&chart_regions,
);
// Compute font stats from full document (cross-page consistency).
let font_stats = markdown::analysis::calculate_font_stats_from_items(&filtered_items);
@@ -509,6 +516,7 @@ pub fn extract_pages_markdown_mem(
let mut results = Vec::with_capacity(pages_slice.len());
let mut pages_needing_ocr = Vec::new();
let mut ocr_reasons_by_page = BTreeMap::new();
let lopdf_pages = doc.get_pages();
for &page_0idx in pages_slice {
// Out-of-range pages → empty + needs_ocr
@@ -542,6 +550,25 @@ pub fn extract_pages_markdown_mem(
let has_gid = gid_pages.contains(&page_1idx);
let has_text_quality_issue = text_quality.pages_needing_ocr.contains(&page_1idx);
// A page can extract cleanly (no decoding issues, non-empty text)
// while still being fundamentally a scan: a full-page raster with
// a little genuine native text drawn over it (a header, a stamp, a
// cover-sheet annotation). Text-quality signals alone can't see
// that — consult the same "large background image" signal
// classify_pdf/detect_pdf_type already uses, so the two APIs can't
// silently disagree on whether a page needs OCR. See #227.
// Also covers vector-outlined text (glyphs drawn as paths, not
// shown via a text-showing operator): a hybrid page with real
// embedded-font body text elsewhere would otherwise still extract
// non-empty, non-garbled markdown and miss OCR routing entirely.
// detect_from_document's Mixed-type per-page routing always sends
// these pages to OCR; mirror that here too. Both signals share one
// analyze_page_content pass — see page_ocr_signals's doc comment.
let (has_template_image, has_vector_text) = lopdf_pages
.get(&page_1idx)
.map(|&page_id| detector::page_ocr_signals(&doc, page_id))
.unwrap_or((false, false));
// Build markdown with document-wide font stats
let options = MarkdownOptions {
base_font_size: Some(font_stats.most_common_size),
@@ -565,6 +592,7 @@ pub fn extract_pages_markdown_mem(
page_count,
prefiltered_page_number_pages: Some(&removed_page_number_pages),
prefiltered_page_number_mask: Some(&page_number_removal_mask),
precomputed_chart_regions: Some(&chart_regions),
},
)
};
@@ -578,10 +606,20 @@ pub fn extract_pages_markdown_mem(
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
if has_template_image {
add_ocr_reason(&mut ocr_reasons_by_page, page_1idx, OCR_REASON_SCANNED);
}
if has_vector_text {
add_ocr_reason(&mut ocr_reasons_by_page, page_1idx, OCR_REASON_VECTOR_TEXT);
}
let ocr_reason = page_ocr_reason(&ocr_reasons_by_page, page_1idx);
let needs_ocr =
ocr_reason.is_some() || md.trim().is_empty() || has_gid || is_garbage_text(&md);
let needs_ocr = ocr_reason.is_some()
|| md.trim().is_empty()
|| has_gid
|| is_garbage_text(&md)
|| has_template_image
|| has_vector_text;
if needs_ocr {
pages_needing_ocr.push(page_1idx);
@@ -619,6 +657,80 @@ pub fn extract_pages_markdown<P: AsRef<Path>>(
extract_pages_markdown_mem(&buffer, pages)
}
// =========================================================================
// Structure-tree element extraction (tagged PDFs)
// =========================================================================
/// One structure-tree element reference from a tagged PDF, resolved to a
/// page and Marked Content ID.
///
/// Join `(page, mcid)` against [`TextItem::page`] / [`TextItem::mcid`] from
/// [`extract_text_with_positions`] to attach semantic roles (heading levels,
/// paragraphs, table cells, …) to extracted text.
#[derive(Debug, Clone)]
pub struct StructureElement {
/// 1-indexed page number (matches [`TextItem::page`]).
pub page: u32,
/// Marked Content ID from the page's content stream (matches
/// [`TextItem::mcid`]).
pub mcid: i64,
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
/// Custom tags are resolved through the document's `/RoleMap`; tags
/// with no standard mapping are returned verbatim.
pub role: String,
}
/// Extract structure-tree element references from a tagged PDF in memory.
///
/// Parses `/StructTreeRoot` (when present) and returns one entry per
/// marked-content reference, resolved to its 1-indexed page, MCID, and
/// structure type name. Returns an empty list when the PDF is not tagged.
///
/// Pass `Some(&[...])` with 1-indexed page numbers (matching
/// [`TextItem::page`]) to restrict output to those pages; pass `None` for
/// the whole document. Entries are sorted by `(page, mcid)`.
pub fn extract_structure_elements_mem(
buffer: &[u8],
pages: Option<&[u32]>,
) -> Result<Vec<StructureElement>, PdfError> {
validate_pdf_bytes(buffer)?;
let (doc, _page_count) = load_document_from_mem(buffer)?;
let Some(tree) = structure_tree::StructTree::from_doc(&doc) else {
return Ok(Vec::new());
};
let page_ids = doc.get_pages();
let roles = tree.mcid_to_roles(&page_ids);
let page_filter: Option<HashSet<u32>> = pages.map(|p| p.iter().copied().collect());
let mut elements: Vec<StructureElement> = roles
.into_iter()
.filter(|(page, _)| page_filter.as_ref().is_none_or(|f| f.contains(page)))
.flat_map(|(page, mcids)| {
mcids.into_iter().map(move |(mcid, role)| StructureElement {
page,
mcid,
role: role.name().to_string(),
})
})
.collect();
elements.sort_unstable_by_key(|e| (e.page, e.mcid));
Ok(elements)
}
/// Path-based wrapper for [`extract_structure_elements_mem`].
///
/// Reads the PDF from disk and extracts structure-tree element references.
/// Pass `None` for `pages` to return the whole document, or `Some(&[...])`
/// to restrict to specific 1-indexed pages.
pub fn extract_structure_elements<P: AsRef<Path>>(
path: P,
pages: Option<&[u32]>,
) -> Result<Vec<StructureElement>, PdfError> {
validate_pdf_file(&path)?;
let buffer = std::fs::read(path.as_ref())?;
extract_structure_elements_mem(&buffer, pages)
}
// =========================================================================
// Region-based text extraction (for hybrid OCR pipelines)
// =========================================================================
@@ -3515,6 +3627,7 @@ fn repair_pdf_container_candidates(buf: &[u8]) -> Vec<Vec<u8>> {
let mut candidates = Vec::new();
add_repair_candidate(&mut candidates, append_missing_eof_marker(buf), buf);
add_repair_candidate(&mut candidates, recover_startxref_pointer(buf), buf);
let stripped = strip_leading_pdf_container_bytes(buf);
if let Some(stripped_buf) = stripped.as_deref() {
@@ -3524,11 +3637,112 @@ fn repair_pdf_container_candidates(buf: &[u8]) -> Vec<Vec<u8>> {
append_missing_eof_marker(stripped_buf),
buf,
);
add_repair_candidate(
&mut candidates,
recover_startxref_pointer(stripped_buf),
buf,
);
}
candidates
}
/// Some PDF writers emit a `startxref` pointer that doesn't actually point
/// at the cross-reference table — a single corrupted byte in the offset is
/// enough. lopdf trusts that pointer outright and fails to load rather than
/// searching for the real table, unlike pypdf/pdfium which both recover by
/// locating it directly. This finds the real (classic, non-stream) `xref`
/// table by scanning for the keyword — validating that a plausible
/// subsection header follows, not just any standalone "xref" token, since
/// this crate processes untrusted input and a coincidental match inside
/// unrelated stream/string content must not get "repaired" against a bogus
/// offset (lopdf would then load successfully against garbage instead of
/// returning a clean error) — and appends a corrected trailing
/// `startxref`/`%%EOF` block. lopdf's own `get_xref_start` always uses the
/// *last* `%%EOF` in the final 512 bytes of the buffer, so ours
/// transparently supersedes the broken one without needing to touch
/// anything already in the file.
///
/// Doesn't cover cross-reference *streams* (`N 0 obj << /Type /XRef ...`,
/// used by some PDF 1.5+ writers instead of a classic table) — recovering
/// those needs the containing object's number, not just a byte offset.
fn recover_startxref_pointer(buf: &[u8]) -> Option<Vec<u8>> {
let xref_pos = find_last_valid_xref_table_start(buf)?;
let mut repaired = Vec::with_capacity(buf.len() + 32);
repaired.extend_from_slice(buf);
if !repaired.ends_with(b"\n") {
repaired.push(b'\n');
}
repaired.extend_from_slice(format!("startxref\n{xref_pos}\n%%EOF\n").as_bytes());
Some(repaired)
}
/// Finds the last standalone `xref` token in `buf` that is immediately
/// followed by a plausible classic cross-reference subsection header
/// (`<start-id> <count>`, e.g. "0 6") — the shape every real classic xref
/// table starts with. A single reverse byte scan: O(n) even on a
/// pathological buffer with many non-matching or non-standalone "xref"
/// occurrences, unlike repeatedly re-searching a shrinking prefix.
fn find_last_valid_xref_table_start(buf: &[u8]) -> Option<usize> {
const KEYWORD: &[u8] = b"xref";
if buf.len() < KEYWORD.len() {
return None;
}
let mut pos = buf.len() - KEYWORD.len();
loop {
if &buf[pos..pos + KEYWORD.len()] == KEYWORD {
let before_ok = pos == 0 || buf[pos - 1].is_ascii_whitespace();
let after_ok = buf
.get(pos + KEYWORD.len())
.is_none_or(|c| c.is_ascii_whitespace());
if before_ok && after_ok && looks_like_xref_subsection_header(buf, pos + KEYWORD.len())
{
return Some(pos);
}
}
if pos == 0 {
return None;
}
pos -= 1;
}
}
/// Checks that `buf[pos..]` starts (after whitespace) with two
/// whitespace-separated runs of ASCII digits — `<start-id> <count>`, the
/// first subsection header of a classic PDF cross-reference table.
fn looks_like_xref_subsection_header(buf: &[u8], pos: usize) -> bool {
fn skip_ws(buf: &[u8], mut pos: usize) -> usize {
while buf.get(pos).is_some_and(u8::is_ascii_whitespace) {
pos += 1;
}
pos
}
fn skip_digits(buf: &[u8], mut pos: usize) -> usize {
while buf.get(pos).is_some_and(u8::is_ascii_digit) {
pos += 1;
}
pos
}
let pos = skip_ws(buf, pos);
let after_first_digits = skip_digits(buf, pos);
if after_first_digits == pos {
return false; // no start-id
}
let sep = skip_ws(buf, after_first_digits);
if sep == after_first_digits {
return false; // start-id and count must be whitespace-separated
}
let after_count = skip_digits(buf, sep);
if after_count == sep {
return false; // no count
}
// The count run must end at whitespace/buffer-end, not run into trailing
// garbage (e.g. a coincidental "xref\n0 6garbage" in stream content).
buf.get(after_count).is_none_or(u8::is_ascii_whitespace)
}
fn add_repair_candidate(
candidates: &mut Vec<Vec<u8>>,
candidate: Option<Vec<u8>>,
@@ -3816,7 +4030,14 @@ fn process_document(
let text_quality = analyze_text_quality(&items);
merge_ocr_reasons(&mut ocr_reasons_by_page, text_quality.reasons_by_page);
let layout = compute_layout_complexity(&items, &layout_items, &rects, &lines);
let chart_regions = markdown::chart_regions_by_page(&items, &rects, &lines);
let layout = compute_layout_complexity_with_chart_regions(
&items,
&layout_items,
&rects,
&lines,
&chart_regions,
);
let md = if options.mode == ProcessMode::Analyze {
None
@@ -3833,6 +4054,7 @@ fn process_document(
page_count,
prefiltered_page_number_pages: Some(&removed_pages),
prefiltered_page_number_mask: Some(removal_mask.as_slice()),
precomputed_chart_regions: Some(&chart_regions),
},
))
};
@@ -5633,11 +5855,29 @@ fn select_items_with_document_folio_context(
}
/// Analyse extracted items and rects for layout complexity.
#[cfg(test)]
fn compute_layout_complexity(
items: &[types::TextItem],
column_items: &[types::TextItem],
rects: &[types::PdfRect],
lines: &[types::PdfLine],
) -> LayoutComplexity {
let page_chart_regions = markdown::chart_regions_by_page(items, rects, lines);
compute_layout_complexity_with_chart_regions(
items,
column_items,
rects,
lines,
&page_chart_regions,
)
}
fn compute_layout_complexity_with_chart_regions(
items: &[types::TextItem],
column_items: &[types::TextItem],
rects: &[types::PdfRect],
lines: &[types::PdfLine],
page_chart_regions: &markdown::PageChartRegions,
) -> LayoutComplexity {
use markdown::analysis::calculate_font_stats_from_items;
@@ -5659,6 +5899,10 @@ fn compute_layout_complexity(
let owned_items: Vec<types::TextItem> = page_items.iter().map(|i| (*i).clone()).collect();
let page_content_width = tables::content_width(&owned_items);
let bands = markdown::split_side_by_side(&owned_items);
let chart_regions = page_chart_regions
.get(&page)
.map(Vec::as_slice)
.unwrap_or_default();
let band_ranges: Vec<(f32, f32)> = if bands.is_empty() {
// Single region — use sentinel range that includes everything
@@ -5673,7 +5917,8 @@ fn compute_layout_complexity(
let band_items: Vec<types::TextItem> = owned_items
.iter()
.filter(|item| {
x_lo == f32::MIN || (item.x >= x_lo - margin && item.x < x_hi + margin)
(x_lo == f32::MIN || (item.x >= x_lo - margin && item.x < x_hi + margin))
&& !markdown::item_is_in_chart_region(item, chart_regions)
})
.cloned()
.collect();
@@ -5725,8 +5970,20 @@ fn compute_layout_complexity(
}
let mut pages_with_columns: Vec<u32> = Vec::new();
for page in seen_pages {
let cols = extractor::detect_columns(column_items, page, pages_with_tables.contains(&page));
for &page in &seen_pages {
let chart_regions = page_chart_regions
.get(&page)
.map(Vec::as_slice)
.unwrap_or_default();
let page_column_items: Vec<types::TextItem> = column_items
.iter()
.filter(|item| {
item.page == page && !markdown::item_is_in_chart_region(item, chart_regions)
})
.cloned()
.collect();
let cols =
extractor::detect_columns(&page_column_items, page, pages_with_tables.contains(&page));
if cols.len() >= 2 {
pages_with_columns.push(page);
}
@@ -5955,6 +6212,66 @@ mod tests {
assert!(filtered.pages_with_columns.is_empty());
}
#[test]
fn dense_chart_panel_is_not_reported_as_a_table() {
let mut items: Vec<TextItem> = (0..8)
.flat_map(|row| {
(0..6).map(move |column| {
test_item(
&format!("{}", row * 10 + column),
105.0 + column as f32 * 35.0,
525.0 - row as f32 * 15.0,
24.0,
10.0,
)
})
})
.collect();
for row in 0..6 {
items.push(test_item(
"Left column prose continues here",
80.0,
320.0 - row as f32 * 15.0,
160.0,
10.0,
));
items.push(test_item(
"Right column prose continues here",
300.0,
320.0 - row as f32 * 15.0,
160.0,
10.0,
));
}
let mut lines: Vec<PdfLine> = (0..30)
.map(|column| PdfLine {
x1: 100.0 + column as f32 * 8.0,
y1: 400.0,
x2: 100.0 + column as f32 * 8.0,
y2: 550.0,
page: 1,
})
.collect();
lines.extend((0..6).map(|row| PdfLine {
x1: 100.0,
y1: 400.0 + row as f32 * 30.0,
x2: 332.0,
y2: 400.0 + row as f32 * 30.0,
page: 1,
}));
let rects = vec![PdfRect {
x: 80.0,
y: 350.0,
width: 280.0,
height: 240.0,
page: 1,
}];
let complexity = compute_layout_complexity(&items, &items, &rects, &lines);
assert!(complexity.pages_with_tables.is_empty());
}
#[test]
fn page_selection_keeps_document_wide_folio_layout_decisions() {
let mut items = Vec::new();
@@ -6958,4 +7275,65 @@ mod tests {
// Pre-filled cell was not touched.
assert_eq!(cells[1].text, "Pre-filled");
}
// -- recover_startxref_pointer / find_last_valid_xref_table_start ------
//
// Direct unit tests on the byte-level scan, addressing review feedback
// on #230: a coincidental standalone "xref" token that isn't actually
// followed by a subsection header (start-id + count) must not be
// treated as a real table — accepting it would let lopdf "succeed"
// against a bogus offset and silently return garbled/empty content
// instead of a clean error.
#[test]
fn find_xref_rejects_standalone_token_without_subsection_header() {
// "xref" appears as a real standalone word, but nothing that looks
// like "<start-id> <count>" follows it.
let buf = b"Please refer to the xref appendix for details.";
assert_eq!(find_last_valid_xref_table_start(buf), None);
}
#[test]
fn find_xref_accepts_real_classic_table_header() {
let buf = b"garbage\nxref\n0 6\n0000000000 65535 f \n%%EOF";
let pos = find_last_valid_xref_table_start(buf).expect("should find the real table");
assert_eq!(&buf[pos..pos + 4], b"xref");
assert_eq!(&buf[pos..], b"xref\n0 6\n0000000000 65535 f \n%%EOF");
}
#[test]
fn find_xref_skips_coincidental_match_and_finds_real_table_before_it() {
// A coincidental "xref" (no subsection header) appears *after* the
// real table in the buffer — the scan must not stop at the first
// (rightmost) standalone token it finds; it must keep looking
// backward until one actually validates.
let buf = b"xref\n0 3\n0000000000 65535 f \ntrailer\nsee the xref\n";
let pos = find_last_valid_xref_table_start(buf).expect("should find the real table");
assert_eq!(pos, 0);
}
#[test]
fn find_xref_rejects_substring_of_startxref() {
// "xref" is a substring of "startxref" but isn't a standalone
// token there (not preceded by whitespace) — must not match, even
// though a number immediately follows it.
let buf = b"startxref\n1234\n%%EOF";
assert_eq!(find_last_valid_xref_table_start(buf), None);
}
#[test]
fn find_xref_rejects_count_run_with_trailing_garbage() {
// "xref\n0 6garbage" has the right shape (digits, whitespace,
// digits) but the count run doesn't end at whitespace/EOF — it
// runs straight into non-digit garbage, so this must not be
// accepted as a real subsection header.
let buf = b"xref\n0 6garbage\n%%EOF";
assert_eq!(find_last_valid_xref_table_start(buf), None);
}
#[test]
fn recover_startxref_pointer_returns_none_without_a_valid_table() {
let buf = b"Please refer to the xref appendix for details.";
assert!(recover_startxref_pointer(buf).is_none());
}
}
+192
View File
@@ -171,6 +171,74 @@ pub(crate) fn is_toc_marker_heading(text: &str) -> bool {
/// equation and absent from name-plus-number headings. A bare trailing colon
/// is NOT a fragment signal either: real headings frequently end with colons
/// ("Procedure:", "Steps for Using the Microscope:").
/// True when the line opens with a section number ("3.", "2.1.4", "IV)").
///
/// Mirrors the acceptance of `heading::parse_numbering` rather than the
/// stricter `convert::starts_with_section_number`, which deliberately
/// requires two components because it bypasses isolation checks. Here a
/// single "1." counts: numbering is independent evidence of a heading, and
/// `heading.rs` applies its numbered-prefix allowance *after* consulting
/// `is_heading_fragment`, so without this exemption a numbered
/// sentence-case heading would be vetoed before that allowance can run.
fn starts_with_numbering_prefix(t: &str) -> bool {
let Some(first) = t.split_whitespace().next() else {
return false;
};
let has_delimiter = first.ends_with(['.', ')', ':']);
let token = first.trim_end_matches(['.', ')', ':']);
if token.is_empty() {
return false;
}
let parts: Vec<&str> = token.split('.').collect();
let decimal = parts
.iter()
.all(|p| !p.is_empty() && p.len() <= 3 && p.chars().all(|c| c.is_ascii_digit()));
if decimal {
// "1." / "2.1." carry a delimiter; "2.3 Title" is written without
// one, so a multi-component number is accepted bare. A bare single
// number ("3 apples") is not — that is ordinary prose.
return has_delimiter || parts.len() >= 2;
}
// Roman numerals go through the heading parser's own grammar so the two
// agree: uppercase I/V/X/L/C only, at most 8 characters. A looser rule
// here would exempt markers the parser rejects — "iv)" or "d)" from an
// alphabetical list — letting an ordinary list item bypass the veto and
// reach heading promotion.
//
// A delimiter is also required: a bare leading "I" is the pronoun far
// more often than a section number.
has_delimiter && crate::markdown::heading::roman_value(token).is_some()
}
/// True when the line reads as a title rather than a sentence: every
/// content word (ignoring minor words) starts uppercase. Used to spare real
/// headings from the dangling-verb veto — "Bond Yields" is a section title,
/// "the method yields" is a stranded clause, and only the casing tells them
/// apart.
fn looks_title_case(t: &str) -> bool {
const MINOR: &[&str] = &[
"a", "an", "the", "of", "and", "or", "for", "to", "in", "on", "at", "by", "with", "from",
"as", "is", "are", "that", "than", "into",
];
let mut content = 0usize;
let mut capitalized = 0usize;
for w in t.split_whitespace() {
let cleaned: String = w.chars().filter(|c| c.is_alphabetic()).collect();
if cleaned.is_empty() {
continue;
}
if MINOR.contains(&cleaned.to_lowercase().as_str()) {
continue;
}
content += 1;
if cleaned.chars().next().is_some_and(char::is_uppercase) {
capitalized += 1;
}
}
// A single content word ("Yields") is a title by default.
content == 0 || capitalized == content
}
pub(crate) fn is_heading_fragment(text: &str) -> bool {
let t = text.trim_end();
@@ -244,9 +312,133 @@ pub(crate) fn is_heading_fragment(text: &str) -> bool {
if t.ends_with(':') && t.split_whitespace().any(is_equation_number) {
return true;
}
// Dangling clause: a stranded sentence lead-in ends on a relational
// verb with no terminal punctuation — "Note that the exact error equals"
// left ahead of its formula when a phantom table dissolved.
//
// Gated on the line reading as prose rather than a title. Case is the
// discriminator the trailing word alone cannot provide: a heading is
// title case ("Bond Yields", "The Method Yields") while a stranded
// lead-in is sentence case ("the method yields"). Without this gate the
// veto eats real headings — "Bond Yields", "Crop Yields" and any wrapped
// title-case heading the preprocessor failed to merge.
if !t.ends_with(['.', '!', '?', ':', ';', ')', ']'])
&& !looks_title_case(t)
&& !starts_with_numbering_prefix(t)
{
if let Some(last) = t.split_whitespace().next_back() {
let word: String = last
.trim_matches(|c: char| !c.is_alphanumeric())
.to_lowercase();
// Relational verbs only, and only those with no common noun
// sense. "yields" was dropped for exactly that reason: "Bond
// Yields" is a real section title. Function words, copulas and
// auxiliaries were measured and rejected outright — a heading
// that wraps across lines ends on those, and suppressing them
// destroyed real IRS Publication 17 headings.
const DANGLING_TAIL: &[&str] =
&["equals", "denotes", "implies", "satisfies", "signifies"];
if DANGLING_TAIL.contains(&word.as_str()) {
return true;
}
}
}
false
}
#[cfg(test)]
mod fragment_heading_tests {
use super::is_heading_fragment;
#[test]
fn dangling_tail_marks_stranded_clause() {
// opendataloader 01030000000144: left behind when a phantom table
// dissolved, ahead of its formula on the next line.
assert!(is_heading_fragment("Note that the exact error equals"));
assert!(is_heading_fragment("The remainder term satisfies"));
assert!(is_heading_fragment("we conclude that the sum equals"));
}
#[test]
fn real_headings_survive() {
assert!(!is_heading_fragment("Introduction"));
assert!(!is_heading_fragment("Error Analysis"));
assert!(!is_heading_fragment("Materials and Methods"));
assert!(!is_heading_fragment("Results"));
assert!(!is_heading_fragment("3.2 Richardson Extrapolation"));
assert!(!is_heading_fragment("Discussion and Conclusions"));
// Terminal punctuation means the clause is complete.
assert!(!is_heading_fragment("What is a Derivative?"));
assert!(!is_heading_fragment("Procedure:"));
assert!(!is_heading_fragment("Note that this is important."));
}
#[test]
fn title_case_headings_ending_in_a_verb_survive() {
// "yields" is also a plural noun; these are real section titles.
assert!(!is_heading_fragment("Bond Yields"));
assert!(!is_heading_fragment("Crop Yields"));
assert!(!is_heading_fragment("Dividend Yields"));
assert!(!is_heading_fragment("Yields"));
// A wrapped title-case heading whose first line ends on a listed
// verb must survive even if the preprocessor failed to merge it.
assert!(!is_heading_fragment("The Theorem Implies"));
assert!(!is_heading_fragment("What This Denotes"));
}
#[test]
fn numbered_sentence_case_headings_survive() {
// heading.rs consults is_heading_fragment BEFORE applying its
// numbered-prefix allowance, so the veto must not pre-empt it.
assert!(!is_heading_fragment("1. What the model implies"));
assert!(!is_heading_fragment("2.3 How the estimator satisfies"));
assert!(!is_heading_fragment("IV) What this denotes"));
// Without numbering the same wording is still a stranded clause.
assert!(is_heading_fragment("What the model implies"));
// A bare leading number or pronoun is prose, not numbering.
assert!(is_heading_fragment("3 apples and what that implies"));
assert!(is_heading_fragment("I think the model implies"));
// Markers heading::parse_numbering rejects must not be exempted
// either, or an ordinary list item bypasses the veto: lowercase
// roman, alphabetical markers, and over-long tokens.
assert!(is_heading_fragment("iv) the estimator satisfies"));
assert!(is_heading_fragment("d) the value implies"));
// Unsupported character (M is outside the parser's I/V/X/L/C set).
assert!(is_heading_fragment("MMMM. the value implies"));
// Over-long token: nine valid characters, so this exercises the
// 8-character bound rather than the character set.
assert!(is_heading_fragment("IIIIIIIII. the value implies"));
// Eight is still within the bound and stays exempt.
assert!(!is_heading_fragment("IIIIIIII. What this implies"));
// Uppercase roman within the parser's grammar is still exempt.
assert!(!is_heading_fragment("IV. What this denotes"));
assert!(!is_heading_fragment("XII) What this implies"));
}
#[test]
fn wrapped_headings_are_not_fragments() {
// A heading that wraps across lines ends on a function word. These
// are real headings from IRS Publication 17 and must survive.
assert!(!is_heading_fragment("Casualty and"));
assert!(!is_heading_fragment("Rule 10. You Must Be at"));
assert!(!is_heading_fragment("Higher Standard Deduction for"));
assert!(!is_heading_fragment("Qualifying Child of"));
assert!(!is_heading_fragment("When Can I Withdraw or"));
// Copulas and auxiliaries also end real wrapped headings.
assert!(!is_heading_fragment("Rule 15. Your AGI Must Be"));
assert!(!is_heading_fragment("What Medical Expenses Are"));
assert!(!is_heading_fragment("Rule 13. You Must Have"));
assert!(!is_heading_fragment("When Can a Roth IRA Be"));
}
#[test]
fn dangling_check_is_case_insensitive() {
// All-caps is not sentence case, so the veto must not fire there.
assert!(!is_heading_fragment("THE REMAINDER EQUALS"));
}
}
/// Compute the Y-gap threshold for paragraph break detection.
///
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
+3 -1
View File
@@ -127,7 +127,9 @@ fn visual_style(line: &TextLine) -> Option<VisualStyle> {
})
}
fn roman_value(token: &str) -> Option<u32> {
/// Shared with `analysis::starts_with_numbering_prefix` so the veto
/// exemption and the heading parser agree on what a roman numeral is.
pub(super) fn roman_value(token: &str) -> Option<u32> {
if token.is_empty() || token.len() > 8 {
return None;
}
+85 -12
View File
@@ -69,7 +69,7 @@ fn is_chart_adjacent_label(item: &TextItem, region: (f32, f32, f32, f32)) -> boo
|| (mostly_inside_chart_width && close_to_chart_edge && category_sized))
}
fn item_is_in_chart_region(item: &TextItem, regions: &[(f32, f32, f32, f32)]) -> bool {
pub(crate) fn item_is_in_chart_region(item: &TextItem, regions: &[(f32, f32, f32, f32)]) -> bool {
regions.iter().any(|&(x0, y0, x1, y1)| {
let cx = item.x + item.width / 2.0;
let within_padded_x = cx >= x0 - CHART_REGION_PAD && cx <= x1 + CHART_REGION_PAD;
@@ -92,6 +92,72 @@ fn items_outside_chart_regions(
.collect()
}
pub(crate) fn merge_chart_regions(
regions: impl IntoIterator<Item = (f32, f32, f32, f32)>,
) -> Vec<(f32, f32, f32, f32)> {
const MERGE_TOLERANCE: f32 = 3.0;
let mut merged: Vec<(f32, f32, f32, f32)> = Vec::new();
for (x0, y0, x1, y1) in regions {
let mut current = (x0.min(x1), y0.min(y1), x0.max(x1), y0.max(y1));
let mut index = 0;
while index < merged.len() {
let candidate = merged[index];
let overlaps = current.2 + MERGE_TOLERANCE >= candidate.0
&& candidate.2 + MERGE_TOLERANCE >= current.0
&& current.3 + MERGE_TOLERANCE >= candidate.1
&& candidate.3 + MERGE_TOLERANCE >= current.1;
if overlaps {
current = (
current.0.min(candidate.0),
current.1.min(candidate.1),
current.2.max(candidate.2),
current.3.max(candidate.3),
);
merged.swap_remove(index);
} else {
index += 1;
}
}
merged.push(current);
}
merged
}
pub(crate) type PageChartRegions = HashMap<u32, Vec<(f32, f32, f32, f32)>>;
/// Compute the chart masks used by both layout analysis and Markdown output.
///
/// Keeping the rect-backed and dense-line heuristics behind one entry point
/// ensures metadata and extraction cannot drift when either detector changes.
pub(crate) fn chart_regions_by_page(
items: &[TextItem],
rects: &[PdfRect],
lines: &[PdfLine],
) -> PageChartRegions {
let mut page_items: HashMap<u32, Vec<TextItem>> = HashMap::new();
for item in items.iter().filter(|item| {
matches!(
&item.item_type,
crate::types::ItemType::Text | crate::types::ItemType::FormField
)
}) {
page_items.entry(item.page).or_default().push(item.clone());
}
page_items
.into_iter()
.filter_map(|(page, items)| {
let rect_regions = crate::tables::detect_chart_regions(&items, rects, page);
let line_regions = crate::tables::detect_dense_line_chart_regions(lines, rects, page)
.into_iter()
.filter(|&region| chart_region_separates_prose_columns(&items, region));
let regions = merge_chart_regions(rect_regions.into_iter().chain(line_regions));
(!regions.is_empty()).then_some((page, regions))
})
.collect()
}
/// Detect side-by-side table layout by finding a significant X-position gap.
///
/// Returns X-band boundaries `[(x_min, split_x), (split_x, x_max)]` when a
@@ -329,6 +395,15 @@ fn chart_spans_prose_split(region: (f32, f32, f32, f32), split_x: f32) -> bool {
split_x - left >= MIN_CHART_WIDTH_PER_SIDE && right - split_x >= MIN_CHART_WIDTH_PER_SIDE
}
pub(crate) fn chart_region_separates_prose_columns(
items: &[TextItem],
region: (f32, f32, f32, f32),
) -> bool {
let outside = items_outside_chart_regions(items, &[region]);
chart_page_prose_column_split(&outside)
.is_some_and(|split_x| chart_spans_prose_split(region, split_x))
}
/// True when adjacent physical rows form an unterminated, lowercase prose
/// continuation in the same projected column.
fn is_cross_row_prose_continuation(previous: &str, current: &str) -> bool {
@@ -1004,6 +1079,7 @@ pub fn to_markdown_from_items_with_rects_and_page_count(
page_count: document_page_count,
prefiltered_page_number_pages: None,
prefiltered_page_number_mask: None,
precomputed_chart_regions: None,
},
)
}
@@ -1021,6 +1097,9 @@ pub(crate) struct MarkdownDocumentContext<'a> {
/// Table detection consumes the original items; the mask is applied only
/// after table claims have been established.
pub(crate) prefiltered_page_number_mask: Option<&'a [bool]>,
/// Optional chart masks shared with layout analysis so the geometry is
/// detected once and interpreted identically by both pipelines.
pub(crate) precomputed_chart_regions: Option<&'a PageChartRegions>,
}
/// Convert positioned text items to markdown, using rectangles and line segments for table detection.
@@ -1047,6 +1126,7 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
page_count: document_page_count,
prefiltered_page_number_pages,
prefiltered_page_number_mask,
precomputed_chart_regions,
} = context;
if items.is_empty() {
@@ -1119,17 +1199,9 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
// Chart regions per page: their text must not steer column detection
// during line grouping (it fills the gutter and fuses two-column lines).
let mut page_chart_map: HashMap<u32, Vec<(f32, f32, f32, f32)>> = HashMap::new();
for &page in page_groups.keys() {
let page_items_ref: Vec<TextItem> = page_groups[&page]
.iter()
.map(|(_, item)| (*item).clone())
.collect();
let regions = crate::tables::detect_chart_regions(&page_items_ref, rects, page);
if !regions.is_empty() {
page_chart_map.insert(page, regions);
}
}
let page_chart_map = precomputed_chart_regions
.cloned()
.unwrap_or_else(|| chart_regions_by_page(&text_items, rects, pdf_lines));
let mut pages: Vec<u32> = page_groups.keys().copied().collect();
pages.sort();
@@ -2051,6 +2123,7 @@ mod tests {
page_count: 1,
prefiltered_page_number_pages: Some(&removed_pages),
prefiltered_page_number_mask: Some(&removal_mask),
precomputed_chart_regions: None,
},
);
+9
View File
@@ -464,6 +464,7 @@ mod tests {
assert!(!is_page_number_line("Hello World"));
assert!(!is_page_number_line("Chapter 1"));
assert!(!is_page_number_line("Total: 500"));
assert!(!is_page_number_line("PAGE0-PARA2-END-MARKER-0"));
}
#[test]
@@ -507,6 +508,14 @@ mod tests {
assert!(result.contains("End"));
}
#[test]
fn test_remove_page_numbers_preserves_page_prefixed_content() {
let input = "PAGE0-PARA2-START substantive report text PAGE0-PARA2-END-MARKER-0";
let result = remove_page_numbers(input);
assert_eq!(result, input);
}
#[test]
fn test_remove_page_numbers_multiple_patterns() {
let input = "\n1\n\nContent\n\n2\n\n---\nMore\n\n3\n";
+89
View File
@@ -271,6 +271,12 @@ pub struct PyTextItem {
pub is_strikeout: bool,
#[pyo3(get)]
pub item_type: String,
/// Marked Content ID from the content stream's BDC/BMC operator, None
/// when the text is not part of marked content. Join with the
/// (page, mcid) pairs from extract_structure_elements to attach
/// structure-tree roles (headings, paragraphs, ...) in tagged PDFs.
#[pyo3(get)]
pub mcid: Option<i64>,
}
#[pymethods]
@@ -286,6 +292,32 @@ impl PyTextItem {
}
}
/// One structure-tree element reference from a tagged PDF.
#[pyclass(name = "StructureElement")]
#[derive(Clone)]
pub struct PyStructureElement {
/// 1-indexed page number (matches TextItem.page).
#[pyo3(get)]
pub page: u32,
/// Marked Content ID from the page's content stream (matches
/// TextItem.mcid).
#[pyo3(get)]
pub mcid: i64,
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", ...).
#[pyo3(get)]
pub role: String,
}
#[pymethods]
impl PyStructureElement {
fn __repr__(&self) -> String {
format!(
"StructureElement(page={}, mcid={}, role='{}')",
self.page, self.mcid, self.role
)
}
}
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
@@ -356,6 +388,18 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item_type_str(&item.item_type),
mcid: item.mcid,
})
.collect()
}
fn convert_structure_elements(elements: Vec<crate::StructureElement>) -> Vec<PyStructureElement> {
elements
.into_iter()
.map(|e| PyStructureElement {
page: e.page,
mcid: e.mcid,
role: e.role,
})
.collect()
}
@@ -613,6 +657,48 @@ fn extract_pages_markdown_bytes(
Ok(to_py_pages_result(result))
}
/// Extract structure-tree element references from a tagged PDF file.
///
/// Parses the document's structure tree (when present) and returns one
/// entry per marked-content reference, resolved to its 1-indexed page,
/// MCID, and structure type name ("H1".."H6", "P", "Table", ...). Returns
/// an empty list when the PDF is not tagged.
///
/// Join (page, mcid) against the page/mcid attributes from
/// [`extract_text_with_positions`] to attach heading levels and other
/// semantic roles to extracted text.
///
/// Args:
/// path: Path to the PDF file.
/// pages: Optional list of 1-indexed pages (matching TextItem.page).
/// When None (default), the whole document is returned.
///
/// Returns:
/// List of StructureElement sorted by (page, mcid).
#[pyfunction]
#[pyo3(signature = (path, pages=None))]
fn extract_structure_elements(
path: &str,
pages: Option<Vec<u32>>,
) -> PyResult<Vec<PyStructureElement>> {
let elements = crate::extract_structure_elements(path, pages.as_deref()).map_err(to_py_err)?;
Ok(convert_structure_elements(elements))
}
/// Extract structure-tree element references from tagged PDF bytes.
///
/// See [`extract_structure_elements`] for details.
#[pyfunction]
#[pyo3(signature = (data, pages=None))]
fn extract_structure_elements_bytes(
data: &[u8],
pages: Option<Vec<u32>>,
) -> PyResult<Vec<PyStructureElement>> {
let elements =
crate::extract_structure_elements_mem(data, pages.as_deref()).map_err(to_py_err)?;
Ok(convert_structure_elements(elements))
}
/// Python module definition.
#[pymodule]
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
@@ -620,6 +706,7 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_class::<PyPageOcrReasons>()?;
m.add_class::<PyPdfClassification>()?;
m.add_class::<PyTextItem>()?;
m.add_class::<PyStructureElement>()?;
m.add_class::<PyRegionText>()?;
m.add_class::<PyPageRegionTexts>()?;
m.add_class::<PyPageMarkdown>()?;
@@ -634,6 +721,8 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_function(wrap_pyfunction!(extract_text_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_with_positions, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_with_positions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_structure_elements, m)?)?;
m.add_function(wrap_pyfunction!(extract_structure_elements_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_pages_markdown, m)?)?;
+819 -38
View File
File diff suppressed because it is too large Load Diff
+308 -10
View File
@@ -435,6 +435,72 @@ fn revised_table_cell_indices(
.collect()
}
/// Index of candidate "body" items (larger-font attachment targets) sorted by
/// Y, so script-attachment checks scan a narrow Y window instead of the whole
/// page per candidate.
struct ScriptBodyIndex<'a> {
/// (y, item), sorted ascending by y
by_y: Vec<(f32, &'a TextItem)>,
/// widest vertical attachment window any body item can produce
max_window: f32,
}
impl<'a> ScriptBodyIndex<'a> {
fn new(items: &'a [TextItem]) -> Self {
// Smallest table-candidate font is 6pt, so any possible attachment
// target is at least 6 x 1.2 pt.
let mut by_y: Vec<(f32, &TextItem)> = items
.iter()
.filter(|i| i.font_size >= 6.0 * 1.2)
.map(|i| (i.y, i))
.collect();
by_y.sort_by(|a, b| a.0.total_cmp(&b.0));
let max_window = by_y
.iter()
.map(|(_, i)| i.font_size * 0.8)
.fold(0.0f32, f32::max);
Self { by_y, max_window }
}
/// True when a small-font item is horizontally attached to a larger-font
/// item at a script baseline offset — a sub/superscript in running text
/// or math (equation subscripts, footnote markers). Script attachments
/// are not table cells; without this filter, display equations with
/// sub/superscripts form phantom small-font table regions (e.g. TeX
/// papers where log subscripts cluster with footnote lines into a fake
/// 3-column table). A genuine baseline offset is required so same-line
/// table neighbours (a small cell beside a larger label cell) are never
/// classified as scripts.
///
/// `min_anchor_size` additionally constrains what counts as an
/// attachment target: the small-font pass accepts any sufficiently
/// larger item (0.0), while the body-font pass requires a heading-sized
/// anchor so a body-size table cell beside a slightly larger label with
/// baseline jitter is never treated as a script.
fn is_script_attachment(&self, small: &TextItem, min_anchor_size: f32) -> bool {
let attach_gap = small.font_size.max(4.0) * 0.6;
let lo = self
.by_y
.partition_point(|(y, _)| *y < small.y - self.max_window);
self.by_y[lo..]
.iter()
.take_while(|(y, _)| *y <= small.y + self.max_window)
.any(|(_, body)| {
let dy = (small.y - body.y).abs();
body.font_size >= small.font_size * 1.2
&& body.font_size >= min_anchor_size
&& dy > body.font_size * 0.05
&& dy <= body.font_size * 0.8
&& {
let gap_after_body = small.x - (body.x + body.width);
let gap_before_body = body.x - (small.x + small.width);
(-attach_gap..=attach_gap).contains(&gap_after_body)
|| (-attach_gap..=attach_gap).contains(&gap_before_body)
}
})
}
}
/// Detect tables in a set of text items from a single page
pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bool) -> Vec<Table> {
detect_tables_with_page_width(items, base_font_size, skip_body_font, content_width(items))
@@ -483,6 +549,27 @@ pub(crate) fn detect_tables_with_page_width(
// === Pass 1: Small-font tables (existing behavior) ===
let table_font_threshold = base_font_size * 0.90;
// Mark sub/superscript attachments once per pass. They stay candidates —
// the masks only remove them from region qualification and column/row
// geometry.
//
// The two passes need different anchor thresholds. In the small-font pass
// any sufficiently larger neighbour is a plausible base for a script. In
// the body-font pass the candidates are themselves body-sized
// (0.85..1.05x), so a merely "slightly larger" neighbour is usually a bold
// label or an adjacent column header, not the base of a superscript —
// treating it as one would strip real cells out of the geometry and lose
// the table. Requiring a heading-sized anchor (>= 1.15x base) keeps the
// body pass to genuine scripts hanging off headings.
let script_index = ScriptBodyIndex::new(items);
let script_flags: Vec<bool> = items
.iter()
.map(|item| script_index.is_script_attachment(item, 0.0))
.collect();
let body_script_flags: Vec<bool> = items
.iter()
.map(|item| script_index.is_script_attachment(item, base_font_size * 1.15))
.collect();
let table_candidates: Vec<(usize, &TextItem)> = items
.iter()
.enumerate()
@@ -494,7 +581,14 @@ pub(crate) fn detect_tables_with_page_width(
.collect();
if table_candidates.len() >= 6 {
let regions = find_table_regions(&table_candidates);
// Qualify regions from non-script items: a cluster of sub/superscripts
// must not, on its own, mark out a table region.
let region_evidence: Vec<(usize, &TextItem)> = table_candidates
.iter()
.filter(|(idx, _)| !script_flags[*idx])
.cloned()
.collect();
let regions = find_table_regions(&region_evidence);
for (y_min, y_max) in regions {
let region_items: Vec<(usize, &TextItem)> = table_candidates
@@ -508,7 +602,9 @@ pub(crate) fn detect_tables_with_page_width(
}
if let Some(mut table) =
detect_table_in_region(&region_items, TableDetectionMode::SmallFont)
detect_table_in_region(&region_items, TableDetectionMode::SmallFont, &|i| {
script_flags[i]
})
{
// Try to recover body-font header row above the small-font table
recover_header_row(&mut table, items, table_font_threshold);
@@ -553,8 +649,20 @@ pub(crate) fn detect_tables_with_page_width(
body_font_low,
body_font_high,
);
// Scripts are NOT filtered out of the candidate set here, mirroring
// the small-font pass: they must stay eligible for cell assignment so
// a sub/superscript that belongs inside a table cell keeps its text.
// The heading-anchored `body_script_flags` mask removes them from
// geometry only.
if body_candidates.len() >= 6 {
let regions = find_table_regions_strict(&body_candidates);
// Same reasoning as the small-font pass: scripts do not qualify
// regions, but remain available for cell assignment within one.
let region_evidence: Vec<(usize, &TextItem)> = body_candidates
.iter()
.filter(|(idx, _)| !body_script_flags[*idx])
.cloned()
.collect();
let regions = find_table_regions_strict(&region_evidence);
log::debug!("body-font: {} strict regions found", regions.len());
for (y_min, y_max, _x_min, _x_max) in &regions {
@@ -580,7 +688,9 @@ pub(crate) fn detect_tables_with_page_width(
}
if let Some(table) =
detect_table_in_region(&region_items, TableDetectionMode::BodyFont)
detect_table_in_region(&region_items, TableDetectionMode::BodyFont, &|i| {
body_script_flags[i]
})
{
tables.push(table);
}
@@ -808,10 +918,30 @@ fn find_table_regions_strict(items: &[(usize, &TextItem)]) -> Vec<(f32, f32, f32
regions
}
/// Detect a table within a specific region
fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode) -> Option<Table> {
// Find column boundaries
let columns = find_column_boundaries(items, mode);
/// Detect a table within a specific region.
///
/// `is_script` marks items that are sub/superscript attachments. Those are
/// excluded from the *geometry* — they must not be able to create a column,
/// which is how equation subscript clusters used to fabricate phantom grids —
/// but they remain eligible for cell assignment, so legitimate cell content
/// (exponents in an engineering-notation table, footnote markers) stays in
/// the cell it belongs to instead of leaking out into the reading order.
fn detect_table_in_region(
items: &[(usize, &TextItem)],
mode: TableDetectionMode,
is_script: &dyn Fn(usize) -> bool,
) -> Option<Table> {
// Column geometry from non-script items only.
let geometry_items: Vec<(usize, &TextItem)> = items
.iter()
.filter(|(idx, _)| !is_script(*idx))
.cloned()
.collect();
// A region that is *entirely* scripts has no table structure at all.
if geometry_items.is_empty() {
return None;
}
let columns = find_column_boundaries(&geometry_items, mode);
let min_cols = 2;
if columns.len() < min_cols || columns.len() > 25 {
log::debug!(
@@ -822,8 +952,8 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
return None;
}
// Find row boundaries
let rows = find_row_boundaries(items);
// Find row boundaries (geometry items only, same reasoning)
let rows = find_row_boundaries(&geometry_items);
let min_rows = 2;
if rows.len() < min_rows {
log::debug!(
@@ -842,6 +972,11 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
);
// Verify this looks like a table: multiple items should align to columns
// Validate against ALL items, including scripts. Columns are derived from
// non-script geometry so scripts cannot *create* a column, but excluding
// them from validation too would let a region manufacture alignment: drop
// the awkward items and whatever remains looks like a tidy grid. Block
// diagrams did exactly that. Everything in the region must fit.
let col_alignment = check_column_alignment(items, &columns, mode);
let min_alignment = match mode {
TableDetectionMode::SmallFont => 0.5,
@@ -912,6 +1047,29 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
cells.push(row_cells);
}
// Validation 0 (small-font pass only): reject tiny all-numeric
// fragments. A <=2-row grid whose every cell is a bare 1-2 digit number
// carries no tabular information — in practice these are
// exponent/subscript clusters from display math that happen to align.
// Body-font tables are not subject to this veto: their cells cannot be
// script glyphs.
if matches!(mode, TableDetectionMode::SmallFont) {
let nonempty_cells: Vec<&String> =
cells.iter().flatten().filter(|c| !c.is_empty()).collect();
if rows.len() <= 2
&& !nonempty_cells.is_empty()
&& nonempty_cells
.iter()
.all(|c| c.len() <= 2 && c.chars().all(|ch| ch.is_ascii_digit()))
{
log::debug!(
" validation 0 fail: tiny all-numeric fragment ({} cells)",
nonempty_cells.len()
);
return None;
}
}
// Validation 1: some rows should have content in first column.
// Use a lower threshold (25%) for tables with wrapped cells where
// continuation lines leave the first column empty.
@@ -1977,6 +2135,146 @@ fn try_add_label_column(
#[cfg(test)]
mod tests {
fn make_item(text: &str, x: f32, y: f32, font_size: f32, width: f32) -> TextItem {
TextItem {
text: text.to_string(),
x,
y,
width,
height: font_size,
font: "TestFont".to_string(),
font_size,
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
}
#[test]
fn script_attachment_detects_subscript_after_body_text() {
let body = make_item("log", 100.0, 500.0, 10.0, 15.0);
let sub = make_item("10", 115.5, 497.0, 7.0, 7.0);
let items = vec![body, sub.clone()];
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sub, 0.0));
}
#[test]
fn script_attachment_detects_superscript_footnote_marker() {
let body = make_item("Hartley", 200.0, 500.0, 10.0, 35.0);
let sup = make_item("2", 235.8, 504.0, 6.6, 3.5);
let items = vec![body, sup.clone()];
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sup, 0.0));
}
#[test]
fn script_attachment_ignores_small_cell_far_from_body_text() {
let body = make_item("Revenue", 100.0, 500.0, 10.0, 40.0);
let cell = make_item("1,234", 180.0, 500.0, 7.0, 20.0);
let items = vec![body, cell.clone()];
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
}
#[test]
fn body_pass_anchor_spares_cells_beside_slightly_larger_labels() {
// A body-font table cell (10pt) sitting beside a slightly larger,
// NON-heading label (12.5pt) with a little baseline jitter. The
// small-font pass treats any larger neighbour as a possible script
// base, but the body pass must not: at body sizes a slightly larger
// neighbour is a bold label or column header, and flagging the cell
// would strip it out of the table geometry and lose the table.
// Cell at the low end of the body band (0.85x base) beside a 10.5pt
// label. 10.5 clears the inherent 1.2x-of-cell rule (10.2) but falls
// below the body pass's heading anchor (11.5), which is exactly the
// band where the two masks must disagree.
let label = make_item("Revenue", 100.0, 500.0, 10.5, 40.0);
let cell = make_item("1,234", 141.0, 496.5, 8.5, 22.0);
let items = vec![label, cell.clone()];
let index = ScriptBodyIndex::new(&items);
let base = 10.0;
assert!(
index.is_script_attachment(&cell, 0.0),
"small-font pass anchor should still see this as an attachment"
);
assert!(
!index.is_script_attachment(&cell, base * 1.15),
"body pass must not treat a cell beside a slightly larger label \
as a script that removes real cells from the geometry"
);
// A genuine heading-sized anchor still qualifies in the body pass.
let heading = make_item("Section", 100.0, 500.0, 20.0, 60.0);
let sup = make_item("3", 161.0, 508.0, 10.0, 5.0);
let h_items = vec![heading, sup.clone()];
assert!(
ScriptBodyIndex::new(&h_items).is_script_attachment(&sup, base * 1.15),
"script hanging off a heading must still be excluded in the body pass"
);
}
#[test]
fn script_attachment_ignores_same_baseline_neighbor_cell() {
// A small cell beside a larger label on the SAME baseline is a table
// layout, not a subscript — a genuine baseline offset is required.
let label = make_item("Total", 100.0, 500.0, 10.0, 25.0);
let cell = make_item("42", 127.0, 500.0, 7.5, 9.0);
let items = vec![label, cell.clone()];
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
}
#[test]
fn script_attachment_ignores_neighbor_on_different_line() {
let body = make_item("Header", 100.0, 500.0, 10.0, 30.0);
let cell = make_item("42", 131.0, 486.0, 7.0, 10.0);
let items = vec![body, cell.clone()];
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
}
/// Equation-subscript + footnote layout from Shannon entropy.pdf page 1,
/// with real coordinates. Without the larger-font anchors the small items
/// alone DO form a phantom table — proving the layout reaches detection —
/// and adding the anchors must suppress it.
fn shannon_page1_small_items() -> Vec<TextItem> {
vec![
make_item("2", 267.4, 133.9, 7.4, 3.7),
make_item("10", 306.2, 133.9, 7.4, 7.4),
make_item("10", 342.7, 133.9, 7.4, 7.4),
make_item("10", 325.0, 118.9, 7.4, 7.4),
make_item("Bell System Technical Journal,", 295.7, 101.9, 8.0, 95.0),
make_item(
"April 1924, p. 324; Certain Topics in",
396.7,
101.9,
8.0,
130.0,
),
make_item("v. 47, April 1928, p. 617.", 250.9, 92.5, 8.0, 90.0),
make_item("Bell System Technical Journal,", 264.2, 82.6, 8.0, 95.0),
make_item("July 1928, p. 535.", 364.3, 82.6, 8.0, 65.0),
]
}
#[test]
fn equation_scripts_do_not_form_phantom_table() {
let bare = shannon_page1_small_items();
assert!(
!detect_tables(&bare, 10.0, false).is_empty(),
"test layout must form a phantom table when the filter cannot fire"
);
let mut items = shannon_page1_small_items();
items.push(make_item("log", 253.0, 137.0, 10.0, 13.5));
items.push(make_item("log", 291.5, 137.0, 10.0, 13.5));
items.push(make_item("log", 328.0, 137.0, 10.0, 13.5));
items.push(make_item("log", 310.3, 122.0, 10.0, 13.5));
let tables = detect_tables(&items, 10.0, false);
assert!(
tables.is_empty(),
"equation scripts + footnotes must not become a table: {tables:?}"
);
}
use super::*;
use crate::types::ItemType;
+450 -2
View File
@@ -4,10 +4,10 @@
//! gridlines. Many IRS forms and government PDFs use these instead of
//! `re` (rectangle) operators.
use std::collections::HashSet;
use std::collections::{HashMap, HashSet};
use crate::tables::Table;
use crate::types::{PdfLine, TextItem};
use crate::types::{PdfLine, PdfRect, TextItem};
use super::detect_rects::{assign_items_to_grid, snap_edges};
@@ -15,11 +15,33 @@ const RULE_Y_TOLERANCE: f32 = 2.0;
const RULE_JOIN_GAP: f32 = 6.0;
const RULE_SPAN_TOLERANCE: f32 = 8.0;
const TEXT_ROW_TOLERANCE: f32 = 2.5;
const DENSE_CHART_MIN_VERTICAL_EDGES: usize = 27;
const DENSE_CHART_LABEL_PAD: f32 = 20.0;
const DENSE_CHART_MAX_SHARED_PANEL_GRIDS: usize = 4;
type HorizontalRule = (f32, f32, f32); // (y, x_min, x_max)
type VerticalRule = (f32, f32, f32); // (x, y_min, y_max)
type AnchoredRow<'a> = (f32, Vec<(usize, &'a TextItem)>);
fn dense_chart_grids_are_co_located(
left: (f32, f32, f32, f32),
right: (f32, f32, f32, f32),
) -> bool {
let left_width = left.2 - left.0;
let right_width = right.2 - right.0;
let left_height = left.3 - left.1;
let right_height = right.3 - right.1;
let horizontal_overlap = (left.2.min(right.2) - left.0.max(right.0)).max(0.0);
let vertical_overlap = (left.3.min(right.3) - left.1.max(right.1)).max(0.0);
let horizontal_gap = (left.0.max(right.0) - left.2.min(right.2)).max(0.0);
let vertical_gap = (left.1.max(right.1) - left.3.min(right.3)).max(0.0);
(vertical_overlap >= left_height.min(right_height) * 0.5
&& horizontal_gap <= left_width.min(right_width) * 0.5)
|| (horizontal_overlap >= left_width.min(right_width) * 0.5
&& vertical_gap <= left_height.min(right_height) * 0.5)
}
#[derive(Debug)]
struct TextAnchorTable {
table: Table,
@@ -1187,6 +1209,279 @@ pub fn detect_tables_from_lines(items: &[TextItem], lines: &[PdfLine], page: u32
detect_tables_from_lines_inner(items, lines, page, true, true)
}
/// Bounding boxes of chart panels backed by a very dense vector grid.
///
/// Tables support at most 25 columns, so a panel with at least 27 distinct,
/// long vertical coordinates plus repeated horizontal rules is treated as
/// chart geometry. When the grid is enclosed by a painted panel rectangle,
/// the region expands to that rectangle so axis labels, legends, and source
/// notes remain part of the figure instead of forming a heuristic table.
pub(crate) fn detect_dense_line_chart_regions(
lines: &[PdfLine],
rects: &[PdfRect],
page: u32,
) -> Vec<(f32, f32, f32, f32)> {
const ANGLE_TOLERANCE: f32 = 0.035;
const MIN_GRID_LINE_LENGTH: f32 = 40.0;
const EXTENT_TOLERANCE: f32 = 6.0;
let mut verticals = Vec::new();
let mut horizontals = Vec::new();
for line in lines.iter().filter(|line| line.page == page) {
let dx = (line.x2 - line.x1).abs();
let dy = (line.y2 - line.y1).abs();
let length = dx.hypot(dy);
if length < MIN_GRID_LINE_LENGTH {
continue;
}
if dy > 0.01 && dx / dy <= ANGLE_TOLERANCE {
verticals.push((
(line.x1 + line.x2) / 2.0,
line.y1.min(line.y2),
line.y1.max(line.y2),
));
} else if dx > 0.01 && dy / dx <= ANGLE_TOLERANCE {
horizontals.push((
(line.y1 + line.y2) / 2.0,
line.x1.min(line.x2),
line.x1.max(line.x2),
));
}
}
if verticals.len() < DENSE_CHART_MIN_VERTICAL_EDGES || horizontals.len() < 3 {
return Vec::new();
}
// Group similar vertical extents once. Neighboring buckets are consulted
// below so coordinates that straddle a bucket boundary still form one
// family, while each line participates in only a constant number of
// candidates instead of being re-scanned for every vertical anchor.
let extent_key = |value: f32| (value / EXTENT_TOLERANCE).round() as i32;
let mut extent_buckets: HashMap<(i32, i32), Vec<VerticalRule>> = HashMap::new();
for vertical in verticals {
extent_buckets
.entry((extent_key(vertical.1), extent_key(vertical.2)))
.or_default()
.push(vertical);
}
let mut grid_regions = Vec::new();
let extent_keys: Vec<(i32, i32)> = extent_buckets.keys().copied().collect();
for key in extent_keys {
let anchor_family = &extent_buckets[&key];
let anchor_bottom = anchor_family.iter().map(|vertical| vertical.1).sum::<f32>()
/ anchor_family.len() as f32;
let anchor_top = anchor_family.iter().map(|vertical| vertical.2).sum::<f32>()
/ anchor_family.len() as f32;
let mut family = Vec::new();
for bottom_offset in -1..=1 {
for top_offset in -1..=1 {
if let Some(bucket) =
extent_buckets.get(&(key.0 + bottom_offset, key.1 + top_offset))
{
family.extend(bucket.iter().copied().filter(|vertical| {
(vertical.1 - anchor_bottom).abs() <= EXTENT_TOLERANCE
&& (vertical.2 - anchor_top).abs() <= EXTENT_TOLERANCE
}));
}
}
}
let xs = snap_edges(&family.iter().map(|&(x, _, _)| x).collect::<Vec<_>>(), 3.0);
if xs.len() < DENSE_CHART_MIN_VERTICAL_EDGES {
continue;
}
let grid_bottom =
family.iter().map(|vertical| vertical.1).sum::<f32>() / family.len() as f32;
let grid_top = family.iter().map(|vertical| vertical.2).sum::<f32>() / family.len() as f32;
if grid_top - grid_bottom < 60.0 {
continue;
}
// A horizontal rule must support the same contiguous dense run of
// vertical coordinates. Splitting at sparse X gaps prevents a shared
// rule from joining a chart to a neighboring ruled table. Keying by
// the covered X-index range also lets multiple chart panels sharing
// the same Y extents produce independent regions.
let mut supported_spans: HashMap<(usize, usize), Vec<f32>> = HashMap::new();
for &(y, line_left, line_right) in &horizontals {
if y < grid_bottom - EXTENT_TOLERANCE || y > grid_top + EXTENT_TOLERANCE {
continue;
}
let start = xs.partition_point(|&x| x < line_left - EXTENT_TOLERANCE);
let end = xs.partition_point(|&x| x <= line_right + EXTENT_TOLERANCE);
if end - start < DENSE_CHART_MIN_VERTICAL_EDGES {
continue;
}
let mut gaps: Vec<f32> = xs[start..end]
.windows(2)
.map(|pair| pair[1] - pair[0])
.collect();
gaps.sort_by(f32::total_cmp);
let dense_gap = gaps[gaps.len() / 4];
let run_break = (dense_gap * 3.0).max(12.0);
let locally_dense_gap_limit = (dense_gap * 1.5).max(6.0);
let locally_dense_gaps = gaps
.iter()
.filter(|&&gap| gap <= locally_dense_gap_limit)
.count();
let mut run_start = start;
let mut retained_dense_run = false;
for index in start..end - 1 {
if xs[index + 1] - xs[index] <= run_break {
continue;
}
let run_end = index + 1;
if run_end - run_start >= DENSE_CHART_MIN_VERTICAL_EDGES
&& xs[run_end - 1] - xs[run_start] >= 120.0
{
supported_spans
.entry((run_start, run_end))
.or_default()
.push(y);
retained_dense_run = true;
}
run_start = run_end;
}
if end - run_start >= DENSE_CHART_MIN_VERTICAL_EDGES
&& xs[end - 1] - xs[run_start] >= 120.0
{
supported_spans.entry((run_start, end)).or_default().push(y);
retained_dense_run = true;
}
// One or two wider category gaps may split an otherwise dense
// chart into sub-threshold runs. Keep the full family only when
// its total width remains close to the expected dense spacing;
// a neighboring sparse table makes this ratio much larger.
let span_width = xs[end - 1] - xs[start];
let expected_dense_width = dense_gap * (end - start - 1) as f32;
if !retained_dense_run
&& span_width >= 120.0
&& span_width <= expected_dense_width * 1.35
&& gaps.len().saturating_sub(locally_dense_gaps) <= 2
{
supported_spans.entry((start, end)).or_default().push(y);
}
}
for ((start, end), ys) in supported_spans {
if snap_edges(&ys, 3.0).len() >= 3 {
grid_regions.push((xs[start], grid_bottom, xs[end - 1], grid_top));
}
}
}
// Prefer the smallest qualifying region when a broad rule happens to
// cover a denser nested panel, and retain every non-overlapping panel.
grid_regions.sort_by(|left, right| {
let left_area = (left.2 - left.0) * (left.3 - left.1);
let right_area = (right.2 - right.0) * (right.3 - right.1);
left_area.total_cmp(&right_area)
});
let mut selected_regions: Vec<(f32, f32, f32, f32)> = Vec::new();
for region in grid_regions {
let area = (region.2 - region.0) * (region.3 - region.1);
let duplicates_existing = selected_regions.iter().any(|existing| {
let overlap_width = (region.2.min(existing.2) - region.0.max(existing.0)).max(0.0);
let overlap_height = (region.3.min(existing.3) - region.1.max(existing.1)).max(0.0);
let overlap_area = overlap_width * overlap_height;
let existing_area = (existing.2 - existing.0) * (existing.3 - existing.1);
overlap_area >= area.min(existing_area) * 0.8
});
if !duplicates_existing {
selected_regions.push(region);
}
}
let all_grid_regions = selected_regions.clone();
let mut regions: Vec<_> = selected_regions
.into_iter()
.map(|grid_region| {
let (grid_left, grid_bottom, grid_right, grid_top) = grid_region;
let enclosing_panel = rects
.iter()
.filter(|rect| rect.page == page)
.filter_map(|rect| {
let (left, width) = if rect.width < 0.0 {
(rect.x + rect.width, -rect.width)
} else {
(rect.x, rect.width)
};
let (bottom, height) = if rect.height < 0.0 {
(rect.y + rect.height, -rect.height)
} else {
(rect.y, rect.height)
};
let right = left + width;
let top = bottom + height;
let enclosed_grids: Vec<_> = all_grid_regions
.iter()
.filter(|&&(other_left, other_bottom, other_right, other_top)| {
left <= other_left + EXTENT_TOLERANCE
&& right >= other_right - EXTENT_TOLERANCE
&& bottom <= other_bottom + EXTENT_TOLERANCE
&& top >= other_top - EXTENT_TOLERANCE
})
.copied()
.collect();
if enclosed_grids.len() > DENSE_CHART_MAX_SHARED_PANEL_GRIDS
|| enclosed_grids.iter().any(|&other| {
other != grid_region
&& !dense_chart_grids_are_co_located(grid_region, other)
})
{
return None;
}
let enclosed_grid_bounds =
enclosed_grids.into_iter().reduce(|bounds, other| {
(
bounds.0.min(other.0),
bounds.1.min(other.1),
bounds.2.max(other.2),
bounds.3.max(other.3),
)
})?;
let enclosed_width = enclosed_grid_bounds.2 - enclosed_grid_bounds.0;
let enclosed_height = enclosed_grid_bounds.3 - enclosed_grid_bounds.1;
(left <= grid_left + EXTENT_TOLERANCE
&& right >= grid_right - EXTENT_TOLERANCE
&& bottom <= grid_bottom + EXTENT_TOLERANCE
&& top >= grid_top - EXTENT_TOLERANCE
&& width <= enclosed_width * 2.0
&& height <= enclosed_height * 4.0
&& !(left < 5.0 && bottom < 5.0))
.then_some(((left, bottom, right, top), width * height))
})
.min_by(|left, right| left.1.total_cmp(&right.1))
.map(|(region, _)| region);
enclosing_panel.unwrap_or((
grid_left - DENSE_CHART_LABEL_PAD,
grid_bottom - DENSE_CHART_LABEL_PAD,
grid_right + DENSE_CHART_LABEL_PAD,
grid_top + DENSE_CHART_LABEL_PAD,
))
})
.collect();
regions.sort_by(|left, right| {
left.0
.total_cmp(&right.0)
.then_with(|| left.1.total_cmp(&right.1))
});
regions.dedup_by(|left, right| {
(left.0 - right.0).abs() <= EXTENT_TOLERANCE
&& (left.1 - right.1).abs() <= EXTENT_TOLERANCE
&& (left.2 - right.2).abs() <= EXTENT_TOLERANCE
&& (left.3 - right.3).abs() <= EXTENT_TOLERANCE
});
regions
}
/// Detect only tables whose cell grid is backed by explicit vector geometry.
///
/// Region-level TSR callers need physical cell boundaries for crop bboxes, so
@@ -1621,6 +1916,159 @@ mod tests {
}
}
#[test]
fn dense_vector_grid_expands_to_enclosing_chart_panel() {
let mut lines: Vec<PdfLine> = (0..30)
.map(|column| make_vline(100.0 + column as f32 * 8.0, 400.0, 550.0, 1))
.collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 332.0, 1)));
let rects = vec![PdfRect {
x: 80.0,
y: 350.0,
width: 280.0,
height: 240.0,
page: 1,
}];
assert_eq!(
detect_dense_line_chart_regions(&lines, &rects, 1),
vec![(80.0, 350.0, 360.0, 590.0)]
);
}
#[test]
fn frameless_dense_vector_grid_includes_label_padding() {
let mut lines: Vec<PdfLine> = (0..30)
.map(|column| make_vline(100.0 + column as f32 * 8.0, 400.0, 550.0, 1))
.collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 332.0, 1)));
assert_eq!(
detect_dense_line_chart_regions(&lines, &[], 1),
vec![(80.0, 380.0, 352.0, 570.0)]
);
}
#[test]
fn multiple_dense_vector_panels_are_retained() {
let mut lines = Vec::new();
for panel_left in [60.0, 380.0] {
lines.extend(
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
);
lines.extend((0..6).map(|row| {
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
}));
}
assert_eq!(
detect_dense_line_chart_regions(&lines, &[], 1),
vec![(40.0, 380.0, 312.0, 570.0), (360.0, 380.0, 632.0, 570.0),]
);
}
#[test]
fn multiple_dense_vector_panels_use_shared_enclosing_panel() {
let mut lines = Vec::new();
for panel_left in [60.0, 380.0] {
lines.extend(
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
);
lines.extend((0..6).map(|row| {
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
}));
}
let rects = vec![PdfRect {
x: 40.0,
y: 350.0,
width: 592.0,
height: 240.0,
page: 1,
}];
assert_eq!(
detect_dense_line_chart_regions(&lines, &rects, 1),
vec![(40.0, 350.0, 632.0, 590.0)]
);
}
#[test]
fn shared_rules_do_not_join_dense_chart_to_adjacent_table() {
let mut lines: Vec<PdfLine> = (0..30)
.map(|column| make_vline(60.0 + column as f32 * 8.0, 400.0, 550.0, 1))
.collect();
lines.extend(
[330.0, 390.0, 450.0, 510.0, 570.0, 630.0]
.into_iter()
.map(|x| make_vline(x, 400.0, 550.0, 1)),
);
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 630.0, 1)));
assert_eq!(
detect_dense_line_chart_regions(&lines, &[], 1),
vec![(40.0, 380.0, 312.0, 570.0)]
);
}
#[test]
fn uneven_dense_spacing_keeps_the_complete_chart_region() {
let mut xs: Vec<f32> = (0..15).map(|column| 60.0 + column as f32 * 8.0).collect();
xs.extend((0..15).map(|column| 212.0 + column as f32 * 8.0));
let mut lines: Vec<PdfLine> = xs.iter().map(|&x| make_vline(x, 400.0, 550.0, 1)).collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 324.0, 1)));
assert_eq!(
detect_dense_line_chart_regions(&lines, &[], 1),
vec![(40.0, 380.0, 344.0, 570.0)]
);
}
#[test]
fn subthreshold_dense_run_does_not_absorb_adjacent_sparse_grid() {
let mut xs: Vec<f32> = (0..21).map(|column| 60.0 + column as f32 * 8.0).collect();
xs.extend((0..6).map(|column| 248.0 + column as f32 * 18.0));
let mut lines: Vec<PdfLine> = xs.iter().map(|&x| make_vline(x, 400.0, 550.0, 1)).collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 338.0, 1)));
assert!(detect_dense_line_chart_regions(&lines, &[], 1).is_empty());
}
#[test]
fn broad_frame_does_not_merge_distant_dense_grids() {
let mut lines = Vec::new();
for panel_left in [60.0, 700.0] {
lines.extend(
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
);
lines.extend((0..6).map(|row| {
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
}));
}
let rects = vec![PdfRect {
x: 40.0,
y: 350.0,
width: 912.0,
height: 240.0,
page: 1,
}];
assert_eq!(
detect_dense_line_chart_regions(&lines, &rects, 1),
vec![(40.0, 380.0, 312.0, 570.0), (680.0, 380.0, 952.0, 570.0)]
);
}
#[test]
fn supported_width_vector_table_is_not_a_dense_chart() {
let mut lines: Vec<PdfLine> = (0..26)
.map(|column| make_vline(100.0 + column as f32 * 10.0, 400.0, 550.0, 1))
.collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 350.0, 1)));
assert!(detect_dense_line_chart_regions(&lines, &[], 1).is_empty());
}
#[test]
fn test_basic_grid_detection() {
// 3x2 grid with horizontal lines at y=500, 480, 460 and vertical at x=100, 200, 300
+454 -10
View File
@@ -1996,17 +1996,249 @@ fn without_dominant_page_backgrounds(rects: &[(f32, f32, f32, f32)]) -> Vec<(f32
.collect()
}
/// Detect a table from cell-background rects that failed grid detection.
/// Repeated rows of touching cell rectangles are stronger table evidence
/// than the bar-length variation used by the chart detector.
///
/// Uses rect Y-edges for row boundaries and text X-position clustering for
/// columns. Handles tables with cell backgrounds that don't form a clean
/// X-edge grid (variable column widths, decorative fills).
/// Chart-bar signature: ≥3 rects sharing an aligned bottom edge (the axis),
/// with similar widths (bars) but strongly varying heights (data-driven),
/// holding at most a single numeric data label each. Bar charts drawn as
/// filled rects otherwise read as cell rects and grid their axis labels
/// into a phantom table. The mirrored check catches horizontal bar charts.
fn is_chart_bar_cluster(
/// Ruled tables with wrapped labels naturally have variable row heights, and
/// numeric-heavy cells can otherwise resemble horizontal or vertical bars.
/// Require several rows to repeat a shared edge schema before overriding the
/// chart hypothesis so sparse plots and independent bars remain unaffected.
fn is_repeated_cell_grid(group_rects: &[(f32, f32, f32, f32)]) -> bool {
type RowGroup = (f32, f32, Vec<(f32, f32)>);
const ROW_EDGE_TOLERANCE: f32 = 3.0;
const MIN_GRID_ROWS: usize = 4;
const MIN_CELLS_PER_ROW: usize = 3;
if group_rects.len() < MIN_GRID_ROWS * MIN_CELLS_PER_ROW {
return false;
}
let mut row_groups: Vec<RowGroup> = Vec::new();
for &(x, y, width, height) in group_rects {
if width < 5.0 || height < 5.0 {
continue;
}
let top = y + height;
if let Some((_, _, cells)) = row_groups.iter_mut().find(|(bottom, row_top, _)| {
(y - *bottom).abs() <= ROW_EDGE_TOLERANCE
&& (top - *row_top).abs() <= ROW_EDGE_TOLERANCE
}) {
cells.push((x, x + width));
} else {
row_groups.push((y, top, vec![(x, x + width)]));
}
}
let mut row_schemas = Vec::new();
for (_, _, mut cells) in row_groups {
if cells.len() < MIN_CELLS_PER_ROW {
continue;
}
let mut widths: Vec<f32> = cells.iter().map(|&(left, right)| right - left).collect();
widths.sort_by(f32::total_cmp);
let median_width = widths[widths.len() / 2];
cells.retain(|&(left, right)| right - left <= median_width * 2.5);
cells.sort_by(|left, right| {
left.0
.total_cmp(&right.0)
.then_with(|| left.1.total_cmp(&right.1))
});
cells.dedup_by(|left, right| {
(left.0 - right.0).abs() <= ROW_EDGE_TOLERANCE
&& (left.1 - right.1).abs() <= ROW_EDGE_TOLERANCE
});
if cells.len() < MIN_CELLS_PER_ROW
|| cells
.windows(2)
.any(|pair| pair[1].0 > pair[0].1 + ROW_EDGE_TOLERANCE)
{
continue;
}
let edges: Vec<f32> = cells
.iter()
.flat_map(|&(left, right)| [left, right])
.collect();
let schema = snap_edges(&edges, ROW_EDGE_TOLERANCE);
if schema.len() > MIN_CELLS_PER_ROW {
row_schemas.push(schema);
}
}
if row_schemas.len() < MIN_GRID_ROWS {
return false;
}
let reference = row_schemas
.iter()
.max_by_key(|schema| schema.len())
.expect("grid rows are non-empty");
row_schemas
.iter()
.filter(|schema| {
let comparable_edges = reference.len().min(schema.len());
let matched_edges = schema
.iter()
.filter(|edge| {
reference
.iter()
.any(|reference_edge| (*edge - *reference_edge).abs() <= ROW_EDGE_TOLERANCE)
})
.count();
matched_edges > MIN_CELLS_PER_ROW && matched_edges * 4 >= comparable_edges * 3
})
.count()
>= MIN_GRID_ROWS
}
fn repeated_cell_grid_overrides_bar_hypothesis(group_rects: &[(f32, f32, f32, f32)]) -> bool {
is_repeated_cell_grid(group_rects)
&& without_dominant_page_backgrounds(group_rects).len() == group_rects.len()
}
/// Detect horizontal segmented stacks from aligned rows of touching rects.
///
/// Category rows must have visible gutters and data-varying internal segment
/// boundaries, unlike the stable boundaries of a ruled table.
struct SegmentedBarGeometry {
bounds: (f32, f32, f32, f32),
row_bands: Vec<(f32, f32)>,
}
fn segmented_stacked_bar_geometry(
group_rects: &[(f32, f32, f32, f32)],
) -> Option<SegmentedBarGeometry> {
type BarRow = (f32, f32, Vec<(f32, f32)>);
const EDGE_TOLERANCE: f32 = 3.0;
const MIN_ROWS: usize = 4;
const MIN_SEGMENTS: usize = 3;
let mut rows: Vec<BarRow> = Vec::new();
for &(x, y, width, height) in group_rects {
if width < 5.0 || height < 5.0 {
continue;
}
let top = y + height;
if let Some((_, _, segments)) = rows.iter_mut().find(|(bottom, row_top, _)| {
(y - *bottom).abs() <= EDGE_TOLERANCE && (top - *row_top).abs() <= EDGE_TOLERANCE
}) {
segments.push((x, x + width));
} else {
rows.push((y, top, vec![(x, x + width)]));
}
}
rows.retain_mut(|(_, _, segments)| {
segments.sort_by(|left, right| left.0.total_cmp(&right.0));
segments.len() >= MIN_SEGMENTS
&& segments
.windows(2)
.all(|pair| (pair[1].0 - pair[0].1).abs() <= EDGE_TOLERANCE)
});
if rows.len() < MIN_ROWS {
return None;
}
rows.sort_by(|left, right| left.0.total_cmp(&right.0));
// Table rows normally share borders. Horizontal stacked bars instead
// leave a visible gutter between category rows.
if rows.windows(2).any(|pair| {
let shorter_height = (pair[0].1 - pair[0].0).min(pair[1].1 - pair[1].0);
pair[1].0 - pair[0].1 < (shorter_height * 0.25).max(2.0)
}) {
return None;
}
// At least two rows must move an internal segment boundary. Stable
// boundaries across every row are stronger evidence for a ruled table.
let reference_edges: Vec<f32> = rows[0]
.2
.iter()
.take(rows[0].2.len() - 1)
.map(|segment| segment.1)
.collect();
let drifting_rows = rows
.iter()
.skip(1)
.filter(|(_, _, segments)| {
let edges: Vec<f32> = segments
.iter()
.take(segments.len() - 1)
.map(|segment| segment.1)
.collect();
edges.len() == reference_edges.len()
&& edges
.iter()
.zip(&reference_edges)
.any(|(edge, reference)| (edge - reference).abs() > EDGE_TOLERANCE)
})
.count();
if drifting_rows < 2 {
return None;
}
let left = rows
.iter()
.flat_map(|row| &row.2)
.map(|segment| segment.0)
.reduce(f32::min)?;
let right = rows
.iter()
.flat_map(|row| &row.2)
.map(|segment| segment.1)
.reduce(f32::max)?;
let bottom = rows.iter().map(|row| row.0).reduce(f32::min)?;
let top = rows.iter().map(|row| row.1).reduce(f32::max)?;
let row_bands = rows.iter().map(|row| (row.0, row.1)).collect();
Some(SegmentedBarGeometry {
bounds: (left, bottom, right, top),
row_bands,
})
}
/// Category labels beside multiple bar rows are independent chart evidence:
/// numeric table text stays inside its cells, regardless of whether the table
/// has an outer border or extra padding.
fn has_external_segmented_bar_labels(
items: &[TextItem],
page: u32,
geometry: &SegmentedBarGeometry,
) -> bool {
const LABEL_EDGE_TOLERANCE: f32 = 3.0;
const LABEL_CLAIM_PAD: f32 = 20.0;
let (content_left, _, content_right, _) = geometry.bounds;
let labeled_rows = geometry
.row_bands
.iter()
.filter(|&&(row_bottom, row_top)| {
items.iter().any(|item| {
if item.page != page || item.text.trim().is_empty() {
return false;
}
let item_left = item.x.min(item.x + item.width);
let item_right = item.x.max(item.x + item.width);
let item_center_x = (item_left + item_right) / 2.0;
let item_center_y = item.y + item.height / 2.0;
let beside_stack = (item_center_x <= content_left + LABEL_EDGE_TOLERANCE
&& item_center_x >= content_left - LABEL_CLAIM_PAD
&& item_left < content_left)
|| (item_center_x >= content_right - LABEL_EDGE_TOLERANCE
&& item_center_x <= content_right + LABEL_CLAIM_PAD
&& item_right > content_right);
beside_stack
&& item_center_y >= row_bottom - LABEL_EDGE_TOLERANCE
&& item_center_y <= row_top + LABEL_EDGE_TOLERANCE
})
})
.count();
labeled_rows >= 2 && labeled_rows * 2 >= geometry.row_bands.len()
}
/// Recognize filled vertical or horizontal bars whose geometry and labels are
/// data-driven rather than uniform table cells.
fn has_chart_bar_signature(
items: &[TextItem],
group_rects: &[(f32, f32, f32, f32)],
page: u32,
@@ -2113,6 +2345,30 @@ fn is_chart_bar_cluster(
|| bar_family(|r| r.1, |r| r.3, |r| r.2, |r| r.0)
}
fn is_chart_bar_cluster(
items: &[TextItem],
group_rects: &[(f32, f32, f32, f32)],
page: u32,
) -> bool {
let has_bar_signature = has_chart_bar_signature(items, group_rects, page);
// A segmented horizontal chart can share most of its edges across rows.
// Row-aligned category labels outside the stack distinguish it from a
// numeric table without depending on whether either shape has a frame.
if has_bar_signature {
if let Some(geometry) = segmented_stacked_bar_geometry(group_rects) {
if has_external_segmented_bar_labels(items, page, &geometry) {
return true;
}
}
}
if repeated_cell_grid_overrides_bar_hypothesis(group_rects) {
return false;
}
has_bar_signature
}
fn detect_row_stripe_table_from_cell_rects(
items: &[TextItem],
group_rects: &[(f32, f32, f32, f32)],
@@ -3083,6 +3339,194 @@ mod tests {
assert!(detect_chart_regions(&items, &rects, 1).is_empty());
}
#[test]
fn variable_height_ruled_grid_overrides_bar_hypothesis() {
let edge_sets = [
[80.0, 140.0, 200.0, 260.0, 320.0, 380.0, 440.0, 500.0, 560.0],
[80.0, 140.0, 210.0, 260.0, 320.0, 380.0, 450.0, 500.0, 560.0],
];
let heights = [20.0, 34.0, 26.0, 42.0, 20.0, 34.0];
let edge_variants = [0, 0, 0, 0, 1, 1];
let mut rects = Vec::new();
let mut y = 650.0;
for (row, height) in heights.into_iter().enumerate() {
let edges = edge_sets[edge_variants[row]];
rects.extend(
edges
.windows(2)
.map(|edge| (edge[0], y, edge[1] - edge[0], height)),
);
y -= height;
}
assert!(is_repeated_cell_grid(&rects));
assert!(has_chart_bar_signature(&[], &rects, 1));
assert!(repeated_cell_grid_overrides_bar_hypothesis(&rects));
assert!(segmented_stacked_bar_geometry(&rects).is_none());
assert!(!is_chart_bar_cluster(&[], &rects, 1));
let mut with_page_fills =
vec![(0.0, 0.0, 600.0, 800.0); DOMINANT_PAGE_BACKGROUND_MIN_REPETITIONS];
with_page_fills.extend(rects);
assert!(!repeated_cell_grid_overrides_bar_hypothesis(
&with_page_fills
));
}
#[test]
fn touching_segments_with_spaced_rows_remain_a_chart() {
let row_edges = [
[100.0, 140.0, 180.0, 220.0, 260.0],
[100.0, 140.0, 180.0, 228.0, 260.0],
[100.0, 140.0, 180.0, 214.0, 260.0],
[100.0, 140.0, 180.0, 232.0, 260.0],
];
let mut raw_rects = vec![(90.0, 530.0, 190.0, 100.0)];
for (row, edges) in row_edges.into_iter().enumerate() {
let y = 540.0 + row as f32 * 20.0;
raw_rects.extend(
edges
.windows(2)
.map(|edge| (edge[0], y, edge[1] - edge[0], 12.0)),
);
}
let items: Vec<TextItem> = (0..4)
.map(|row| make_item("Category", 62.0, 541.0 + row as f32 * 20.0, 9.0))
.collect();
assert!(is_repeated_cell_grid(&raw_rects));
assert!(has_chart_bar_signature(&items, &raw_rects, 1));
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented stack");
assert!(has_external_segmented_bar_labels(&items, 1, &geometry));
assert!(is_chart_bar_cluster(&items, &raw_rects, 1));
let numeric_items: Vec<TextItem> = (0..4)
.map(|row| make_item("2024", 80.0, 541.0 + row as f32 * 18.0, 9.0))
.collect();
assert!(has_external_segmented_bar_labels(
&numeric_items,
1,
&geometry
));
assert!(is_chart_bar_cluster(&numeric_items, &raw_rects, 1));
let edge_adjacent_items: Vec<TextItem> = (0..4)
.map(|row| make_item("2024", 92.0, 541.0 + row as f32 * 18.0, 9.0))
.collect();
assert!(has_external_segmented_bar_labels(
&edge_adjacent_items,
1,
&geometry
));
assert!(is_chart_bar_cluster(&edge_adjacent_items, &raw_rects, 1));
let far_items: Vec<TextItem> = (0..4)
.map(|row| make_item("Category", 20.0, 541.0 + row as f32 * 18.0, 9.0))
.collect();
assert!(!has_external_segmented_bar_labels(&far_items, 1, &geometry));
assert!(!is_chart_bar_cluster(&far_items, &raw_rects, 1));
let rects: Vec<PdfRect> = raw_rects
.into_iter()
.map(|(x, y, width, height)| PdfRect {
x,
y,
width,
height,
page: 1,
})
.collect();
assert_eq!(detect_chart_regions(&items, &rects, 1).len(), 1);
let (tables, hints) = detect_tables_from_rects(&items, &rects, 1);
assert!(tables.is_empty());
assert!(hints.is_empty());
}
#[test]
fn padded_numeric_grid_frame_remains_a_table() {
let row_edges = [
[100.0, 140.0, 180.0, 220.0, 260.0],
[100.0, 140.0, 180.0, 228.0, 260.0],
[100.0, 140.0, 180.0, 214.0, 260.0],
[100.0, 140.0, 180.0, 232.0, 260.0],
];
let mut raw_rects = vec![(96.0, 536.0, 168.0, 80.0)];
let mut items = Vec::new();
for (row, edges) in row_edges.into_iter().enumerate() {
let y = 540.0 + row as f32 * 20.0;
for edge in edges.windows(2) {
raw_rects.push((edge[0], y, edge[1] - edge[0], 12.0));
items.push(make_item("42", edge[0] + 8.0, y + 1.0, 9.0));
}
}
assert!(is_repeated_cell_grid(&raw_rects));
assert!(has_chart_bar_signature(&items, &raw_rects, 1));
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented rows");
assert!(!has_external_segmented_bar_labels(&items, 1, &geometry));
assert!(!is_chart_bar_cluster(&items, &raw_rects, 1));
let flush_items: Vec<TextItem> = (0..4)
.map(|row| make_item("1", 100.0, 541.0 + row as f32 * 20.0, 9.0))
.collect();
assert!(!has_external_segmented_bar_labels(
&flush_items,
1,
&geometry
));
assert!(!is_chart_bar_cluster(&flush_items, &raw_rects, 1));
let rects: Vec<PdfRect> = raw_rects
.into_iter()
.map(|(x, y, width, height)| PdfRect {
x,
y,
width,
height,
page: 1,
})
.collect();
assert!(detect_chart_regions(&items, &rects, 1).is_empty());
assert!(!detect_tables_from_rects(&items, &rects, 1).0.is_empty());
}
#[test]
fn frameless_segmented_chart_with_category_labels_remains_a_chart() {
let row_edges = [
[100.0, 140.0, 180.0, 220.0, 260.0],
[100.0, 140.0, 180.0, 228.0, 260.0],
[100.0, 140.0, 180.0, 214.0, 260.0],
[100.0, 140.0, 180.0, 232.0, 260.0],
];
let mut raw_rects = Vec::new();
let mut items = Vec::new();
for (row, edges) in row_edges.into_iter().enumerate() {
let y = 540.0 + row as f32 * 18.0;
raw_rects.extend(
edges
.windows(2)
.map(|edge| (edge[0], y, edge[1] - edge[0], 12.0)),
);
items.push(make_item("Category", 62.0, y + 1.0, 9.0));
}
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented stack");
assert!(has_external_segmented_bar_labels(&items, 1, &geometry));
assert!(is_chart_bar_cluster(&items, &raw_rects, 1));
let rects: Vec<PdfRect> = raw_rects
.into_iter()
.map(|(x, y, width, height)| PdfRect {
x,
y,
width,
height,
page: 1,
})
.collect();
assert_eq!(detect_chart_regions(&items, &rects, 1).len(), 1);
}
// --- detect_stacked_box_table ---
/// N stacked boxes at x=100, w=300, h=22, top-to-bottom from y=600.
+3 -1
View File
@@ -16,7 +16,9 @@ pub(crate) use detect_heuristic::{
content_width, detect_tables_with_page_width, is_table_of_contents,
};
pub use detect_lines::detect_tables_from_lines;
pub(crate) use detect_lines::detect_vector_grid_tables_from_lines;
pub(crate) use detect_lines::{
detect_dense_line_chart_regions, detect_vector_grid_tables_from_lines,
};
pub(crate) use detect_rects::cluster_rects;
pub use detect_rects::{detect_chart_regions, detect_tables_from_rects, RectHintRegion};
pub use detect_struct::detect_tables_from_struct_tree;
+10 -3
View File
@@ -70,10 +70,17 @@ pub(crate) fn is_page_number_line(text: &str) -> bool {
let lowercase = text.trim().to_ascii_lowercase();
lowercase.strip_prefix("page").is_some_and(|rest| {
rest.trim_start()
.chars()
.next()
let mut characters = rest.trim_start().chars().peekable();
let mut has_page_number = false;
while characters
.peek()
.is_some_and(|character| character.is_ascii_digit())
{
has_page_number = true;
characters.next();
}
has_page_number && characters.next().is_none_or(char::is_whitespace)
})
}
+140 -1
View File
@@ -594,7 +594,7 @@ fn hex_to_unicode_string(hex: &str) -> Option<String> {
let bytes: Option<Vec<u8>> = (0..hex.len())
.step_by(2)
.map(|i| u8::from_str_radix(&hex[i..i + 2], 16).ok())
.map(|i| u8::from_str_radix(hex.get(i..i + 2)?, 16).ok())
.collect();
let bytes = bytes?;
@@ -880,6 +880,25 @@ fn try_remap_subset_cmap(
None => return (cmap, None),
};
// Both repair paths below assume CIDs are glyph indices that a subsetter can
// renumber, which is only true for CIDFontType2 (TrueType). For CIDFontType0
// (CFF), CIDs are resolved through the CFF charset, so a valid CMap stays valid
// after subsetting and renumbering it corrupts otherwise-correct text.
// CIDToGIDMap is likewise CIDFontType2-only (PDF 32000-1:2008, 9.7.4.2), so this
// also ignores a CIDToGIDMap that a malformed producer attached to a CFF font.
// /Subtype may be an indirect reference, so resolve it through the document.
// Only bail out when the descendant is *explicitly* something other than
// CIDFontType2: a missing or unresolvable /Subtype keeps the previous
// behaviour rather than silently disabling the repair.
let subtype = cid_font_dict.get(b"Subtype").ok().and_then(|o| match o {
Object::Reference(r) => doc.get_object(*r).ok().and_then(|o| o.as_name().ok()),
other => other.as_name().ok(),
});
if subtype.is_some_and(|name| name != b"CIDFontType2") {
debug!("Subset remap skipped for obj={obj_num}: descendant is not CIDFontType2");
return (cmap, None);
}
// If there's an explicit CIDToGIDMap, build a repaired CMap using it.
if let Some(cid_to_gid) = get_cid_to_gid_map(cid_font_dict, doc) {
if let Some(repaired) = build_cmap_with_cid_to_gid_map(&cmap, &cid_to_gid) {
@@ -2587,6 +2606,24 @@ endcmap
assert_eq!(cmap.lookup(0x0025), Some("B".to_string()));
}
#[test]
fn test_hex_to_unicode_non_ascii_no_panic() {
// A destination containing a multi-byte char makes the byte length even
// while a byte offset can land inside a char. Slicing must not panic;
// it should be rejected gracefully.
assert_eq!(hex_to_unicode_string("XéY"), None);
assert_eq!(hex_to_unicode_string("\u{fffd}0"), None);
}
#[test]
fn test_parse_bfchar_non_ascii_destination_no_panic() {
// Crafted /ToUnicode CMap: a non-hex, non-ASCII destination previously
// triggered a char-boundary panic in hex_to_unicode_string.
let cmap_content = "beginbfchar <0041> <XéY> endbfchar";
// Must not panic; the malformed entry is simply skipped.
let _ = ToUnicodeCMap::parse(cmap_content.as_bytes());
}
#[test]
fn test_parse_bfchar_1byte() {
// This is the pattern that caused the CJK bug: codespace is <0000><FFFF>
@@ -3159,4 +3196,106 @@ endbfrange
"Remap must fire when CMap's CIDs are outside W array coverage"
);
}
#[test]
fn test_try_remap_skipped_for_cid_font_type0() {
// Same W/CMap mismatch as the CIDFontType2 case above, but the descendant is
// CIDFontType0 (CFF). There CIDs are resolved through the CFF charset, so the
// ToUnicode CIDs stay valid after subsetting and must not be renumbered.
// Real-world case: Japanese Adobe-Japan1 PDFs (e.g. National Diet Library
// minutes) where remapping turned correct text into unrelated glyphs.
let cmap_content = r#"
1 begincodespacerange
<0000><FFFF>
endcodespacerange
1 beginbfrange
<0200> <0220> <0410>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
let mut doc = Document::new();
// CIDToGIDMap is CIDFontType2-only, but a malformed producer can still emit
// one on a CFF font. Use a real stream (not /Identity, which is treated as
// "no map") so this also fails if the guard is moved back below the
// CIDToGIDMap branch: cid 1 -> gid 0x0200, which the CMap resolves.
let mut cid_to_gid = vec![0u8; 68];
cid_to_gid[2] = 0x02;
cid_to_gid[3] = 0x00;
let cid_to_gid_id =
doc.add_object(lopdf::Stream::new(lopdf::Dictionary::new(), cid_to_gid));
let mut cid_font = lopdf::Dictionary::new();
cid_font.set("Subtype", lopdf::Object::Name(b"CIDFontType0".to_vec()));
cid_font.set("CIDToGIDMap", lopdf::Object::Reference(cid_to_gid_id));
cid_font.set(
"W",
lopdf::Object::Array(vec![
lopdf::Object::Integer(1),
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
]),
);
let cid_font_id = doc.add_object(cid_font);
let mut font_dict = lopdf::Dictionary::new();
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
font_dict.set(
"DescendantFonts",
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
);
let (primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 789);
assert!(
remapped.is_none(),
"Remap must be skipped for CIDFontType0 (CFF) descendants, including a \
CIDToGIDMap a malformed producer attached to one"
);
// The original CMap must still resolve its own CIDs.
assert_eq!(primary.lookup(0x0200), Some("\u{0410}".to_string()));
}
#[test]
fn test_try_remap_resolves_indirect_subtype() {
// /Subtype may be stored as an indirect reference. A genuine CIDFontType2
// font must still get the repair, so the guard has to dereference it rather
// than treat the unresolved value as "not CIDFontType2".
let cmap_content = r#"
1 begincodespacerange
<0000><FFFF>
endcodespacerange
1 beginbfrange
<0200> <0220> <0410>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
let mut doc = Document::new();
let subtype_id = doc.add_object(lopdf::Object::Name(b"CIDFontType2".to_vec()));
let mut cid_font = lopdf::Dictionary::new();
cid_font.set("Subtype", lopdf::Object::Reference(subtype_id));
cid_font.set("CIDToGIDMap", lopdf::Object::Name(b"Identity".to_vec()));
cid_font.set(
"W",
lopdf::Object::Array(vec![
lopdf::Object::Integer(0),
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
]),
);
let cid_font_id = doc.add_object(cid_font);
let mut font_dict = lopdf::Dictionary::new();
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
font_dict.set(
"DescendantFonts",
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
);
let (_primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 790);
assert!(
remapped.is_some(),
"An indirect /Subtype naming CIDFontType2 must still reach the remap"
);
}
}
+68
View File
@@ -0,0 +1,68 @@
%PDF-1.3
%“Œ‹ž ReportLab Generated PDF document (opensource)
1 0 obj
<<
/F1 2 0 R
>>
endobj
2 0 obj
<<
/BaseFont /Helvetica /Encoding /WinAnsiEncoding /Name /F1 /Subtype /Type1 /Type /Font
>>
endobj
3 0 obj
<<
/Contents 7 0 R /MediaBox [ 0 0 612 792 ] /Parent 6 0 R /Resources <<
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>> /Rotate 0 /Trans <<
>>
/Type /Page
>>
endobj
4 0 obj
<<
/PageMode /UseNone /Pages 6 0 R /Type /Catalog
>>
endobj
5 0 obj
<<
/Author (anonymous) /CreationDate (D:20260803112923+00'00') /Creator (anonymous) /Keywords () /ModDate (D:20260803112923+00'00') /Producer (ReportLab PDF Library - \(opensource\))
/Subject (unspecified) /Title (untitled) /Trapped /False
>>
endobj
6 0 obj
<<
/Count 1 /Kids [ 3 0 R ] /Type /Pages
>>
endobj
7 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ] /Length 202
>>
stream
GarW05mr9@&;9NOME,dW.,B;'jAjYq0S4Z`*D9aMA;]$5J)A/$3lESen1?F)ZJsa4$4&N%-%cs)#qW5EVhhbPiRDrAV>MC%.spto@CU"ZdipR'TtFiMR_%m*Hm$N%qL7a"ckkp9T/s[N2"Og377mP*M^akb2XQZ@'l*qT(9bVtDb5+)S&Q.#%E)<]Ao`TSk2AE'/E\fn~>endstream
endobj
xref
0 8
0000000000 65535 f
0000000061 00000 n
0000000092 00000 n
0000000199 00000 n
0000000392 00000 n
0000000460 00000 n
0000000721 00000 n
0000000780 00000 n
trailer
<<
/ID
[<6d7ea1213c5974c78613d5d2a08423b5><6d7ea1213c5974c78613d5d2a08423b5>]
% ReportLab generated PDF document -- digest (opensource)
/Info 5 0 R
/Root 4 0 R
/Size 8
>>
startxref
9072
%%EOF
File diff suppressed because one or more lines are too long
Binary file not shown.
Binary file not shown.
File diff suppressed because it is too large Load Diff
+199
View File
@@ -1310,6 +1310,18 @@ fn test_snapshot_2013_app2() {
assert_snapshot("2013-app2");
}
/// First two pages of Shannon's "A Mathematical Theory of Communication"
/// (1998 dvips 5.58 → Distiller 3 retypesetting). Canonical legacy-TeX PDF:
/// non-embedded base-14 fonts with no /Widths (exercises the built-in AFM
/// metrics fallback), Type3 PK bitmap math fonts with FontMatrix
/// [1 0 0 -1 0 0] (exercises visual-size scaling), a two-line embedded drop
/// cap, indent-only paragraph breaks, and display math that must not be
/// detected as tables or headings.
#[test]
fn test_snapshot_shannon_entropy() {
assert_snapshot("shannon-entropy-p1-2");
}
// ============================================================================
// Pages Needing OCR Tests
// ============================================================================
@@ -1417,6 +1429,77 @@ fn test_firecrawl_tagged_pdf_struct_tree() {
assert_eq!(fence_count % 2, 0, "Code fences should be balanced");
}
#[test]
fn test_tagged_pdf_text_items_carry_mcid() {
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
assert!(
items.iter().any(|i| i.mcid.is_some()),
"Tagged PDF text items should carry Marked Content IDs"
);
}
#[test]
fn test_extract_structure_elements_tagged_pdf() {
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
assert!(!elements.is_empty(), "Tagged PDF should yield elements");
assert!(
elements.iter().any(|e| e.role == "H1"),
"Should surface H1 heading roles"
);
assert!(
elements.iter().all(|e| !e.role.is_empty()),
"Every element should carry a role name"
);
// Sorted by (page, mcid) for deterministic output
assert!(
elements
.windows(2)
.all(|w| (w[0].page, w[0].mcid) <= (w[1].page, w[1].mcid)),
"Elements should be sorted by (page, mcid)"
);
// The advertised join: (page, mcid) pairs must line up with the
// mcid-carrying TextItems from positioned extraction, and joining the
// H1 entries must recover non-empty heading text.
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
let h1_refs: std::collections::HashSet<(u32, i64)> = elements
.iter()
.filter(|e| e.role == "H1")
.map(|e| (e.page, e.mcid))
.collect();
let h1_text: String = items
.iter()
.filter(|i| i.mcid.is_some_and(|mcid| h1_refs.contains(&(i.page, mcid))))
.map(|i| i.text.as_str())
.collect();
assert!(
!h1_text.trim().is_empty(),
"Joining H1 structure elements to text items should recover heading text"
);
// Page filter is 1-indexed (matching TextItem.page) and equals the
// corresponding subset of the full document result.
let page1 = pdf_inspector::extract_structure_elements_mem(&buf, Some(&[1])).unwrap();
assert!(!page1.is_empty(), "Page 1 should have elements");
assert!(page1.iter().all(|e| e.page == 1));
let full_page1_count = elements.iter().filter(|e| e.page == 1).count();
assert_eq!(page1.len(), full_page1_count);
}
#[test]
fn test_extract_structure_elements_untagged_pdf_empty() {
let buf = std::fs::read("tests/fixtures/thermo-freon12.pdf").unwrap();
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
assert!(
elements.is_empty(),
"Untagged PDF should yield no structure elements, got {:?}",
elements
);
}
#[test]
fn test_identity_h_no_tounicode_suppresses_garbage() {
// shinagawa_identity_h.pdf uses YuGothic with Identity-H encoding and no
@@ -3914,6 +3997,65 @@ fn encrypted_pdf_decrypts_with_correct_password() {
);
}
/// Regression for the #231 review finding: `extract_pages_markdown`'s
/// `has_template_image` check must be gated the same way
/// `classify_pdf`/`detect_pdf_type` gates it (image_count <= 1, few text
/// ops, low alphanumeric diversity) — not treated as sufficient on its
/// own. The fixture is a real text page with substantial, richly varied
/// body text (>=50 Tj ops) drawn over a full-bleed background image
/// (e.g. letterhead/watermark). Before the fix, has_template_image alone
/// forced needs_ocr=true and discarded the page's clean markdown; now the
/// page must extract normally.
#[test]
fn test_extract_pages_markdown_does_not_ocr_text_page_with_watermark_image() {
let buf = std::fs::read("tests/fixtures/text_page_with_watermark_image.pdf").unwrap();
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
let page = &ext.pages[0];
assert!(
!page.needs_ocr,
"a text page with substantial real text should not be routed to OCR \
just because it has a background image"
);
assert!(
page.markdown.contains("watermark"),
"expected the page's real body text to be preserved, got: {:?}",
page.markdown
);
}
/// Regression for the #231 review finding: `extract_pages_markdown` never
/// checked `has_vector_text` at all, even though `detect_from_document`'s
/// Mixed-type per-page routing always sends vector-outlined-text pages to
/// OCR (outlined glyphs can't be extracted as text). A page with massive
/// path ops (outlined decorative text) plus a short genuine caption would
/// extract that caption cleanly — non-empty, non-garbled — so the
/// existing empty/garbage-text checks alone couldn't catch it.
#[test]
fn test_extract_pages_markdown_ocrs_page_with_vector_outlined_text() {
let buf = std::fs::read("tests/fixtures/vector_outlined_text_with_caption.pdf").unwrap();
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
assert!(
cls.pages_needing_ocr.contains(&1),
"classify_pdf should flag page 1 as needing OCR (vector-outlined text), got: {:?}",
cls.pages_needing_ocr
);
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
let page = &ext.pages[0];
assert!(
page.needs_ocr,
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
);
assert!(
page.markdown.is_empty(),
"a page flagged needs_ocr must not return markdown as if extraction were \
trustworthy, got: {:?}",
page.markdown
);
}
#[test]
fn pdf_options_debug_redacts_password() {
let opts = PdfOptions::new().password("secret123");
@@ -3924,3 +4066,60 @@ fn pdf_options_debug_redacts_password() {
);
assert!(dbg.contains("REDACTED"), "expected redaction marker: {dbg}");
}
/// Regression for #228: a `startxref` pointer corrupted to point at the
/// wrong byte offset (a single flipped digit — a real, common writer bug)
/// must not make the whole file unprocessable. The real classic xref table
/// is still present and findable by scanning for the `xref` keyword; both
/// pypdf and pdfium recover the same way. Before this fix, every entry
/// point raised "Invalid PDF structure" on a file whose object data was
/// otherwise completely intact.
#[test]
fn test_process_pdf_recovers_corrupted_startxref_pointer() {
let result = process_pdf_with_options(
"tests/fixtures/broken_startxref_pointer.pdf",
PdfOptions::new(),
)
.expect("a corrupted startxref pointer should be recoverable, like pypdf/pdfium");
assert_eq!(result.page_count, 1);
let md = result.markdown.unwrap_or_default();
assert!(
md.contains("Order Detail Report by Account") && md.contains("WIDGET ASSEMBLY"),
"recovered document should extract its real text, got: {md:?}"
);
}
/// Regression for #227: `extract_pages_markdown`'s per-page `needs_ocr`
/// must agree with `classify_pdf`/`detect_pdf_type` on the same page. The
/// fixture is a full-page raster "scan" with a single line of genuine
/// native text drawn over it (a header) — the native text extracts
/// perfectly cleanly (no decoding issues, non-empty), so a needs_ocr
/// computation based on text-quality signals alone says `false`, while
/// detection correctly sees a dominant background image and says the page
/// needs OCR. Both must now agree it needs OCR, and the markdown must not
/// be returned as if the extraction were trustworthy.
#[test]
fn test_extract_pages_markdown_agrees_with_classify_on_scan_with_native_header() {
let buf = std::fs::read("tests/fixtures/scan_with_native_header_text.pdf").unwrap();
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
assert!(
cls.pages_needing_ocr.contains(&1),
"classify_pdf should flag page 1 as needing OCR (image-dominated), got: {:?}",
cls.pages_needing_ocr
);
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
let page = &ext.pages[0];
assert!(
page.needs_ocr,
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
);
assert!(
page.markdown.is_empty(),
"a page flagged needs_ocr must not return markdown as if extraction were \
trustworthy, got: {:?}",
page.markdown
);
}
+43
View File
@@ -0,0 +1,43 @@
Reprinted with corrections from *The Bell System Technical Journal,* Vol. 27, pp. 379423, 623656, July, October, 1948.
## A Mathematical Theory of Communication
### By C. E. SHANNON
INTRODUCTION
HE recent development of various methods of modulation such as PCM and PPM which exchange
# Tbandwidth for signal-to-noise ratio has intensified the interest in a general theory of communication. A
basis for such a theory is contained in the important papers of Nyquist¹ and Hartley² on this subject. In the present paper we will extend the theory to include a number of new factors, in particular the effect of noise in the channel, and the savings possible due to the statistical structure of the original message and due to the nature of the final destination of the information. The fundamental problem of communication is that of reproducing at one point either exactly or ap- proximately a message selected at another point. Frequently the messages have *meaning*; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one *selected from a set* of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design. If the number of messages in the set is finite then this number or any monotonic function of this number can be regarded as a measure of the information produced when one message is chosen from the set, all choices being equally likely. As was pointed out by Hartley the most natural choice is the logarithmic function. Although this definition must be generalized considerably when we consider the influence of the statistics of the message and when we have a continuous range of messages, we will in all cases use an essentially logarithmic measure. The logarithmic measure is more convenient for various reasons:
1. It is practically more useful. Parameters of engineering importance such as time, bandwidth, number of relays, etc., tend to vary linearly with the logarithm of the number of possibilities. For example, adding one relay to a group doubles the number of possible states of the relays. It adds 1 to the base 2 logarithm of this number. Doubling the time roughly squares the number of possible messages, or doubles the logarithm, etc.
2. It is nearer to our intuitive feeling as to the proper measure. This is closely related to (1) since we in- tuitively measures entities by linear comparison with common standards. One feels, for example, that two punched cards should have twice the capacity of one for information storage, and two identical channels twice the capacity of one for transmitting information.
3. It is mathematically more suitable. Many of the limiting operations are simple in terms of the loga- rithm but would require clumsy restatement in terms of the number of possibilities. The choice of a logarithmic base corresponds to the choice of a unit for measuring information. If the
base 2 is used the resulting units may be called binary digits, or more briefly *bits,* a word suggested by
J. W. Tukey. A device with two stable positions, such as a relay or a flip-flop circuit, can store one bit of information. *N* such devices can store*N* bits, since the total number of possible states is 2
*N* and log₂2 *N* = *N*. If the base 10 is used the units may be called decimal digits. Since
log₂*M* = log₁₀*M*= log₁₀2 = 3:32 log₁₀*M*;
1 Nyquist, H., “Certain Factors Affecting Telegraph Speed,” *Bell System Technical Journal,* April 1924, p. 324; “Certain Topics in Telegraph Transmission Theory,” *A.I.E.E. Trans.,* v. 47, April 1928, p. 617. 2 Hartley, R. V. L., “Transmission of Information,” *Bell System Technical Journal,* July 1928, p. 535.
INFORMATION SOURCE TRANSMITTER RECEIVER DESTINATION
SIGNAL RECEIVED SIGNAL MESSAGE MESSAGE
NOISE SOURCE
Fig. 1 — Schematic diagram of a general communication system.
a decimal digit is about 3 13 bits. A digit wheel on a desk computing machine has ten stable positions and therefore has a storage capacity of one decimal digit. In analytical work where integration and differentiation are involved the base *e* is sometimes useful. The resulting units of information will be called natural units. Change from the base *a* to base *b* merely requires multiplication by log*ba*. By a communication system we will mean a system of the type indicated schematically in Fig. 1. It consists of essentially five parts:
1. An *information source* which produces a message or sequence of messages to be communicated to the receiving terminal. The message may be of various types: (a) A sequence of letters as in a telegraph of teletype system; (b) A single function of time *f* (*t*) as in radio or telephony; (c) A function of time and other variables as in black and white television — here the message may be thought of as a function *f* (*x*; *y*;*t*) of two space coordinates and time, the light intensity at point (*x*; *y*) and time *t* on a pickup tube plate; (d) Two or more functions of time, say *f* (*t*), *g*(*t*), *h*(*t*) — this is the case in “three- dimensional” sound transmission or if the system is intended to service several individual channels in multiplex; (e) Several functions of several variables — in color television the message consists of three functions *f* (*x*; *y*;*t*), *g*(*x*; *y*;*t*), *h*(*x*; *y*;*t*) defined in a three-dimensional continuum — we may also think of these three functions as components of a vector field defined in the region — similarly, several black and white television sources would produce “messages” consisting of a number of functions of three variables; (f) Various combinations also occur, for example in television with an associated audio channel.
2. A *transmitter* which operates on the message in some way to produce a signal suitable for trans- mission over the channel. In telephony this operation consists merely of changing sound pressure into a proportional electrical current. In telegraphy we have an encoding operation which produces a sequence of dots, dashes and spaces on the channel corresponding to the message. In a multiplex PCM system the different speech functions must be sampled, compressed, quantized and encoded, and finally interleaved properly to construct the signal. Vocoder systems, television and frequency modulation are other examples of complex operations applied to the message to obtain the signal.
3. The *channel* is merely the medium used to transmit the signal from transmitter to receiver. It may be a pair of wires, a coaxial cable, a band of radio frequencies, a beam of light, etc.
4. The *receiver* ordinarily performs the inverse operation of that done by the transmitter, reconstructing the message from the signal.
5. The *destination* is the person (or thing) for whom the message is intended. We wish to consider certain general problems involving communication systems. To do this it is first
necessary to represent the various elements involved as mathematical entities, suitably idealized from their
+73
View File
@@ -203,6 +203,79 @@ class TestExtractTextWithPositions:
assert len(items) > 0
assert all(item.page == 1 for item in items)
def test_mcid(self):
# Untagged fixture: mcid is None or int, never anything else
items = pdf_inspector.extract_text_with_positions(
fixture_path("thermo-freon12.pdf")
)
assert all(item.mcid is None or isinstance(item.mcid, int) for item in items)
# Tagged fixture: marked content carries MCIDs
tagged = pdf_inspector.extract_text_with_positions(
fixture_path("firecrawl_docs_tagged.pdf")
)
assert any(item.mcid is not None for item in tagged)
# ---------------------------------------------------------------------------
# extract_structure_elements / extract_structure_elements_bytes
# ---------------------------------------------------------------------------
class TestExtractStructureElements:
def test_tagged_file(self):
elements = pdf_inspector.extract_structure_elements(
fixture_path("firecrawl_docs_tagged.pdf")
)
assert len(elements) > 0
assert all(isinstance(e.page, int) for e in elements)
assert all(isinstance(e.mcid, int) for e in elements)
assert all(isinstance(e.role, str) and len(e.role) > 0 for e in elements)
assert any(e.role == "H1" for e in elements)
def test_join_with_text_items(self):
# (page, mcid) joins against extract_text_with_positions to recover
# heading text
path = fixture_path("firecrawl_docs_tagged.pdf")
elements = pdf_inspector.extract_structure_elements(path)
items = pdf_inspector.extract_text_with_positions(path)
h1_refs = {(e.page, e.mcid) for e in elements if e.role == "H1"}
h1_text = "".join(
item.text
for item in items
if item.mcid is not None and (item.page, item.mcid) in h1_refs
)
assert len(h1_text.strip()) > 0
def test_with_pages(self):
# pages filter is 1-indexed, matching TextItem.page
elements = pdf_inspector.extract_structure_elements(
fixture_path("firecrawl_docs_tagged.pdf"), pages=[1]
)
assert len(elements) > 0
assert all(e.page == 1 for e in elements)
def test_bytes(self):
data = fixture_bytes("firecrawl_docs_tagged.pdf")
elements = pdf_inspector.extract_structure_elements_bytes(data)
assert len(elements) > 0
assert any(e.role == "H1" for e in elements)
def test_untagged_returns_empty(self):
elements = pdf_inspector.extract_structure_elements(
fixture_path("thermo-freon12.pdf")
)
assert elements == []
def test_repr(self):
elements = pdf_inspector.extract_structure_elements(
fixture_path("firecrawl_docs_tagged.pdf")
)
assert "StructureElement" in repr(elements[0])
def test_not_a_pdf(self):
with pytest.raises(ValueError):
pdf_inspector.extract_structure_elements_bytes(b"not a pdf")
# ---------------------------------------------------------------------------
# extract_text_in_regions / extract_text_in_regions_bytes
+2 -2
View File
@@ -724,7 +724,7 @@ checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
[[package]]
name = "pdf-inspector"
version = "0.1.7"
version = "1.14.0"
dependencies = [
"env_logger",
"include_dir",
@@ -740,7 +740,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-wasm"
version = "0.1.3"
version = "1.14.0"
dependencies = [
"console_error_panic_hook",
"js-sys",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-wasm"
version = "0.1.3"
version = "1.14.0"
edition = "2021"
authors = ["Firecrawl Team"]
description = "Browser WebAssembly bindings for pdf-inspector"