Compare commits

..
Author SHA1 Message Date
Abimael Martell 0b0ca64ffb docs: document structure-element extraction and TextItem.mcid 2026-08-11 10:55:15 -07:00
Abimael Martell 7054d6aa69 chore(release): unify package versions (#344)
* chore(release): unify package versions

* fix(release): harden version synchronization
2026-08-11 09:26:19 -07:00
Abimael MartellandClaude Fable 5 a67ee03269 feat(bindings): expose TextItem.mcid and structure-tree element extraction (#346)
Tagged PDFs carry a structure tree with real heading roles (H1..H6), and
the core already parses it (structure_tree::StructTree) and threads MCIDs
onto TextItem — but neither surfaced through the bindings.

- Expose TextItem.mcid (Option<i64>) through the napi and pyo3 bindings,
  matching the core field added with the marked-content extractor.
- Add StructRole::name(), the inverse of from_name, so roles have a
  stable string form.
- Add extract_structure_elements / extract_structure_elements_mem to the
  core: one (page, mcid, role) entry per marked-content reference, sorted
  by (page, mcid), empty for untagged PDFs. Pages are 1-indexed to match
  TextItem.page, so results join directly against
  extract_text_with_positions output.
- Bind it as extractStructureElements (napi) and
  extract_structure_elements / extract_structure_elements_bytes (pyo3),
  with type-stub updates in pdf_inspector.pyi.
- Cover the join in Rust integration tests, napi test.mjs, and pytest,
  using the existing firecrawl_docs_tagged.pdf fixture (tagged) and
  thermo-freon12.pdf (untagged).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 20:03:07 -07:00
Abimael Martell 965dc65f1b chore(release): bump package versions (#343) 2026-08-10 14:33:40 -07:00
36dd5fa426 fix(structure-tree): bound recursive /K parsing with cycle detection and a node budget (#322)
* fix(structure-tree): bound tagged /K parsing against alias/cycle DoS

A struct element that references itself (or an ancestor) through /K — e.g.
/K [n 0 R n 0 R] — made parse_struct_element_dict branch exponentially:
the depth cap (64) alone still permits 2^depth materialized nodes, so a
~830-byte PDF exhausts memory (OOM, exit 134).

Add a StructWalk carrying (1) an active-path set of object IDs so a node
that references itself/an ancestor is not re-expanded (breaks self- and
mutual-reference cycles cheaply), and (2) a global node budget
(MAX_STRUCT_NODES) that caps total materialization for aliased/DAG-shaped
graphs of distinct objects the path guard cannot catch.

Adds regression tests for self-alias, mutual-alias, and the aliased-DAG
budget cap.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): charge /K content refs against the node budget

The per-node budget only covered materialized struct elements and child
recursion; bare MCIDs and MCR dicts in a /K array append to content_refs
without charging it, so one element with a very wide /K array could still
allocate content_refs without bound. Charge every /K array item before
handling it, and stop the top-level /K loop once the budget is spent, so
content refs and loop work are bounded too. Adds a wide-MCID-array test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): charge /K budget per materialized item, not per array entry

Charging every /K array item double-counted structural children (charged
here and again at their node entry) and charged cycle-skipped references
that materialize nothing, draining the budget up to ~2x faster than the
per-node semantics and risking early truncation of large legitimate trees.
Charge only the unbounded content-ref items (bare MCIDs and MCR dicts);
structural children remain charged once at their node entry.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* refactor(structure-tree): charge every content ref uniformly via helper

Route all budget charges through StructWalk::charge() so every
marked-content reference is charged once, including the single-value /K
branches (bare integer and MCR dict) that previously appended without
charging. charge() also guards against underflow, so charging after the
node-entry charge (which can leave the budget at 0) is safe. Makes the
documented per-item budget contract hold uniformly across all branches.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* feat(structure-tree): log once when the node budget truncates parsing

Add a one-shot truncation flag on StructWalk, set the first time the
budget is exhausted, and emit a single warn! after parsing so an operator
can tell when a (very large or malformed) tagged tree was cut off. Avoids
per-item log spam; negligible overhead on the normal path.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): flag truncation at budget guards, not just in charge

The truncation flag was only set inside charge() on the budget==0 branch,
but the dominant skip paths use budget==0 guards that break/return before
charge() is ever called with an empty budget, so the flag (and the warn!)
almost never fired. Route those guards through a new exhausted() that sets
the flag when it skips remaining work. Adds a parser-level test that would
have caught the missed warning.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): flag cycle/depth skips and charge bare MCIDs fully

Two review follow-ups:
- Cycle-broken and depth-capped /K skips dropped tagged content without
  setting the truncation flag, so the one-shot warning never fired for
  malformed/over-deep trees. Mark those skips via note_skipped() and
  broaden the warning to cover non-budget truncation.
- A bare /K MCID materializes a wrapper node AND a content reference but
  charged only one budget unit, allowing ~2x the advertised budget for
  such content; charge both.

Adds tests: cycle-skip flags truncation, and bare MCID charges two units.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): charge MCR-dict wrappers the same two units as bare MCIDs

A top-level MCR /K dict flows through parse_kid -> parse_struct_element_dict
and materializes a Span node + one content ref (two items) but was charged
only one unit at node entry, while the bare-MCID path charges two. Charge
the content reference in the MCR branch too so the per-item budget is
uniform across both wrapper paths. Adds a symmetric test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): reserve leaf-wrapper budget units atomically

A leaf MCID wrapper (bare MCID or MCR dict) materializes a node + one
content ref and charged the two units via separate charge() calls. At the
last unit the first charge succeeded and the second failed, consuming a
unit without emitting the wrapper and denying it to a later element that
would have fit. Add charge_n() to reserve both units atomically (or
neither), and detect MCR before the node charge so it reserves both up
front. Adds a boundary test asserting the leftover unit is preserved.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): stop scanning wide /K once a leaf reservation stalls

The atomic charge_n(2) left budget nonzero (==1) when it failed, so
exhausted() (budget==0) never broke the root /K loop and a crafted wide
array of leaf wrappers was scanned in full after no leaf could fit. Add a
stalled flag set on an insufficient reservation and fold it into
exhausted(); charge()-based (one-unit) loops are unaffected since they
reach budget 0 exactly. Adds a test that a one-unit budget still allows a
one-unit item but a failed two-unit reservation stops the scan.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(structure-tree): add traversal budget and stop charging non-materializing dicts

Two review follow-ups on budget accounting:
- Wide /K arrays of non-materializing items (unsupported value types, OBJR
  dicts, cycle back-edges) consumed no node budget, so the loop scanned the
  whole array. Add a separate work budget charged per examined /K item and
  break the loops when it is spent, bounding traversal even when nothing
  materializes.
- OBJR dicts and dicts without a valid /S were charged the node budget before
  being recognized and skipped, draining the shared budget and truncating
  real content later. Hoist the OBJR check and /S validation above the node
  charge so only materializing nodes consume it (matching the MCR hoisting).

Adds tests for the work-budget bound, wide unsupported /K, and non-materializing
dicts not charging the node budget.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* docs(structure-tree): mention traversal budget in truncation warning

The one-shot truncation warning listed the node budget, cycle, and depth
as causes but not the new traversal (work) budget, so a work-budget
truncation printed a misleading message. Include MAX_STRUCT_WORK so
malformed-PDF debugging identifies the actual limit hit.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 13:36:40 -07:00
fabec0aec3 feat(napi): add async variants that keep the Node event loop free (#337)
* feat(napi): add processPdfAsync, classifyPdfAsync, extractPagesMarkdownAsync

The Node bindings are synchronous, so every call parses on the event
loop thread — up to hundreds of milliseconds of dead loop per document
in a server. Add additive AsyncTask-based variants that run the same
shared implementations on the libuv thread pool and return promises.

The existing synchronous exports keep their names, signatures, and
behaviour; each sync/async pair shares one implementation. Panics in
compute() are caught and surfaced as rejections, matching the sync
error contract.

Closes #336

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(napi): read async task buffers in place instead of copying

Review feedback on #337: buffer.to_vec() copied the whole PDF on the
event loop before the task was queued, so large inputs still stalled
the loop and doubled peak memory. The tasks now hold the napi Buffer
itself — its ref pins the JS allocation for the task's lifetime and
the backing store is stable, so compute() reads it directly from the
worker thread. Callers must not mutate the buffer until the promise
settles (same contract as Node's async fs APIs); documented on each
export and in the README.

The suggested removal of ts_return_type was checked and rejected:
without it napi-rs generates Promise<unknown> for AsyncTask returns.
A comment now records that finding.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(napi): copy async task input on the JS thread for soundness

Review feedback on #337: holding the napi Buffer and reading it from
the libuv worker was unsound. Buffer derefs straight to the JS-side
allocation, so a caller mutating it before the promise settled would
race the worker's reads — undefined behavior, not a recoverable error,
and the documented don't-mutate contract was unenforceable. Deferring
the copy to compute() would not help: any off-thread read races the
same way. The JS thread is the only race-free place to take the copy,
because JS is single-threaded and nothing can mutate the buffer during
the synchronous part of the call.

Revert to an owned Vec<u8> copied at call time. The cost is one memcpy,
negligible next to the parse the async variants exist to unblock. Docs
now state the buffer may be reused or mutated immediately, and a test
locks in the copy semantics by mutating the input while a parse is in
flight.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 12:29:23 -07:00
1f28c00a13 Bound column-detection histogram and harden coordinate handling (#328)
* Bound column-detection histogram size

Derive the projection histogram from a clamped bin count and skip
non-finite page widths. Extreme or malformed text-item coordinates
(from the content-stream text matrix) could otherwise drive a very
large allocation. 65,536 bins is ~9x the largest legal page, so real
layouts are unaffected. Adds regression tests.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Exclude non-finite coordinates from page bounds

Items at NaN/inf positions are now skipped when folding the page bounds,
so a malformed coordinate can no longer escape as a ColumnRegion
boundary, and an all-non-finite page returns no columns. Bad items are
dropped individually rather than failing the page, so one stray glyph
does not disable column detection.

Addresses review feedback on the finite-width guard.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Trim far-outlier coordinates from page bounds

Gutter margins, spanning-item width and the XY-cut margin are all
fractions of page_width, so a single far-but-finite item (x=50_000 is
enough) set the scale for the whole page: real gutters fell inside the
rejected margin band and a genuine two-column page collapsed to one
region. When the span exceeds one legal page (14_400 units), re-derive
the bounds from items clustered around the median x. Outliers keep their
text because column assignment buckets by nearest overlap.

The MAX_BINS ceiling stays as an allocation bound that does not depend
on this heuristic.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Harden bounds trimming against widths and wide layouts

Check both item edges when trimming: a malformed width at an ordinary
position poisoned x_max just as a malformed position poisoned x_min, so
a huge width still collapsed a two-column page to one region.

Only trim when the far items are a small minority (<=10%). A genuinely
large-format page has content spread across its full width, so it now
keeps its true bounds instead of being reduced to the median cluster.

Correct the MAX_PAGE_EXTENT comment: 14_400 units is the traditional
Acrobat architectural limit, not a format cap. PDF 2.0 sets no page-size
limit and UserUnit scales physical size, so this is a heuristic.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Scale bin width so the histogram spans the whole page

Clamping the bin count alone left anything past MAX_BINS * BIN_WIDTH
(~131k points) outside the histogram, folded into the final bin. A page
wide enough to hit that lost real gutters: with a visible gutter inside
the covered range the XY-cut fallback never runs, so a three-column
layout silently reported two. Derive bin_width from page_width instead,
keeping the same allocation ceiling and degrading only resolution.

Also anchor the trimming median on the same finite left/right items that
bounds() accepts, so a malformed width cannot shift which items count as
strays.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Require geometric evidence before detaching far content

The median-window trim narrowed any content more than one page from the
centre, so a valid large page with a sparse far sidebar lost the sidebar
from its bounds and its text fell into column 0. An item-count minority
rule cannot tell that layout from malformed coordinates.

Group content into clusters separated by more than a whole page of
continuous emptiness, and only drop a cluster that is both detached by
such a void and a small minority of items. Real content does not leave a
gap that large; a stray coordinate sits alone beyond one.

A single run wider than one page is treated as a malformed width, which
also covers the huge-width case the cluster sweep cannot see (such an
item spans everything and leaves no gap).

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Judge run width against page content, not a fixed extent

Treating any run wider than 14_400 units as malformed penalised valid
large pages: one made entirely of such runs reported no columns at all,
and a mixed page lost the right edge of every long run.

Judge width relative to the page's own content instead. Positions cannot
be inflated by a bogus width, so the spread of the core cluster is a
sound scale: a run wider than that spread plus one page is malformed.
A genuinely large page keeps its genuinely long runs, while a 1e12-wide
run beside ordinary text is still rejected.

Cluster on positions rather than filled intervals, so a bogus width can
no longer merge everything into one cluster, and keep ordinary pages on
an O(n) fast path that skips the sort entirely.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 12:29:12 -07:00
f4b8c9e854 Clarify SECURITY.md reporting channels (#329)
* Update SECURITY.md reporting channels

Clarify that email is the only required channel and point the
alternative at Firecrawl's Bugcrowd disclosure engagement instead of
the private-advisory link, which is not enabled on this repo.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* Make Bugcrowd the preferred reporting channel

Bugcrowd's disclosure engagement is the primary channel; email to
help@firecrawl.dev is offered as the alternative.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 16:27:50 -07:00
69039f2728 Fix char-boundary panic in hex_to_unicode_string (#320)
Use hex.get(i..i+2) instead of &hex[i..i+2] so a non-hex, non-ASCII
destination in a /ToUnicode CMap can no longer trigger a UTF-8
char-boundary panic. An even byte length does not guarantee the byte
offset falls on a char boundary; get() returns None on a non-boundary
or out-of-range index, folding cleanly into the existing flow.

Add regression tests covering a multi-byte destination char and a
replacement-char byte.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 08:21:31 -07:00
3cca6446bd fix(glyph_names): handle non-ASCII input in uniXXXX glyph name parsing (#321)
Use str::get instead of a byte-length check plus slice when parsing the
uniXXXX glyph-name form. The byte-length guard only proved the index was
in bounds, not on a UTF-8 char boundary, so a glyph name containing
non-ASCII bytes could cause a slice on a non-boundary index. Switch to a
checked slice that folds into the existing Option flow, and add tests.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 08:21:22 -07:00
493fed498e fix(links): prevent stack-overflow DoS from AcroForm /Kids self-cycle (#314)
* fix(links): guard AcroForm /Kids traversal against cycles and huge trees

A crafted PDF whose AcroForm field lists itself (or another ancestor) in
/Kids caused walk_form_fields to recurse indefinitely, overflowing the
stack and aborting pdf2md (exit 134) — an application-level DoS from a
~730-byte input.

Track visited field object IDs to break /Kids cycles, and cap total
field-node traversal at 100k nodes to bound pathologically large trees.

Adds regression tests for self-cycle and mutual-cycle field graphs.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): cap AcroForm /Kids recursion depth to stop deep-chain overflow

The visited-set guard stops cyclic /Kids graphs, but a long *acyclic*
chain of distinct fields still recurses to the chain length and overflows
the stack (a ~1.6MB PDF with 20k linked fields aborts pdf2md, exit 134)
before the 100k node budget is reached.

Add an explicit recursion depth cap (100 levels — far above any legitimate
form hierarchy) so stack usage is bounded independently of node count.

Adds a deep-acyclic-chain regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): enforce form-field node budget before insertion

The node-budget guard inserted each field ID into the visited set before
checking the budget, so the check triggered an early return but never
actually capped the set. A field with a huge /Kids array kept inserting
post-budget IDs, letting visited (memory and work) grow with the crafted
input rather than stopping at MAX_FORM_FIELD_NODES.

Check depth and budget before inserting, so visited can never exceed the
cap. Adds a wide-tree regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): stop /Fields and /Kids iteration once node budget is spent

Checking the budget before insertion capped the visited set, but callers
still iterated every remaining entry of a wide /Fields or /Kids array
after the budget was exhausted — each walk returned immediately, yet the
O(N) sibling iteration let a single multi-million-entry array burn
extraction CPU unbounded. Break out of both the top-level and recursive
loops once visited reaches the cap, making the budget a true
traversal-work cap. Adds a top-level wide-/Fields regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): charge examined entries against the field-node budget

The budget counted only distinct visited nodes, so /Fields or /Kids
arrays full of invalid (non-reference) or duplicate entries never grew
visited and ran to completion regardless of size — the node budget did
not actually cap traversal work.

Introduce FieldWalkBudget tracking both visited nodes and total entries
examined; charge every array entry (valid, invalid, or duplicate) and
stop once either hits MAX_FORM_FIELD_NODES. Adds a regression test with a
huge /Kids array of duplicate + null entries.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): iterate /Fields and /Kids arrays by borrow, not clone

Both arrays were cloned in full before the budget check, so a crafted
oversized /Fields or /Kids array forced an O(n) allocation and copy
regardless of the cap. resolve_array already returns a borrow tied to the
document and the walker only needs a shared &Document, so iterate the
borrowed arrays directly — the early break now bounds how many entries
are even touched, before any per-array allocation.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* docs(links): correct wide-array test comments to match range assertions

The two wide-array tests assert item counts within a range near the
budget, not an exact value (charging entries in the entry guard shifts
the boundary by one or two). Fix the stale comments that claimed exact
counts.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-08 23:43:33 -07:00
Abimael Martell 436af97038 fix(tables): exclude script attachments and tiny numeric fragments from detection (#264)
* fix(tables): exclude script attachments and tiny numeric fragments from detection

Split out of #242 (draft) so it can be reviewed on its own evidence.

Display equations with sub/superscripts form phantom small-font table
regions: the subscripts cluster with nearby small text (footnotes, axis
labels) into fake multi-column grids. Two guards:

- Script attachment: a small-font item horizontally adjacent to a
  larger-font item at a genuine baseline offset is a sub/superscript,
  not a table cell, and is excluded from candidates. A real baseline
  offset is required so a small cell beside a larger same-baseline
  label is never filtered. Attachment targets are indexed by Y and
  scanned through a bounded window rather than a full-page sweep.
  The body-font pass applies the same exclusion but only for
  heading-sized anchors (>= 1.15x base), so body-size table cells
  beside slightly larger labels are untouched.
- Tiny numeric fragments: a <=2-row grid whose every cell is a bare
  1-2 digit number carries no tabular information. Restricted to the
  small-font pass, where the pattern is overwhelmingly exponent
  clusters; body-font numeric grids are unaffected.

Corpus impact: 20 of 186 documents, measured against a control build of
main so main's own drift is excluded. Table rows fall in 17 of 18
inspected documents and no content is lost — 2103_07786 drops all 21
rows, every one a math fragment ('|X 42 43|1|||'); Stijn_SB_doc drops
173 rows of footnote text that had been shredded into cells, with word
count slightly UP and footnote markers intact. M2019_mordeste gains 11
rows from a 3-column grid re-detected as 2-column, neither clearly
better nor worse.

Note the 20 documents is far more than the 3 that per-change ablation
suggested: that figure measured sole-cause attribution inside the
original combined PR, where other heuristics changed the same files and
masked this one. Reach and sole-cause are different measurements.

805 unit + 148 integration tests pass, clippy clean, in-repo snapshots
unchanged.

* fix(tables): suppress script column evidence instead of dropping candidates

Reworked after reviewing the corpus diffs: the first approach removed
sub/superscripts from table candidates entirely, which had two failure
modes beyond the intended fix.

- Legitimate cell content was displaced. citizen-sr-282 is a calculator
  manual whose engineering-notation table lists M = 10^6, k = 10^3.
  Those exponents are superscripts, so they were dropped from the table
  and resurfaced elsewhere in the reading order ('9 mega 6 kilo = 10 3
  milli').
- Removing items changed the candidate geometry, so different spurious
  structure could form from what remained.

Scripts are now kept as candidates and excluded only from the geometry:
they cannot create a column (find_column_boundaries), cannot qualify a
region on their own (find_table_regions / _strict), but are still
assigned to cells. Column alignment is validated against ALL items
including scripts — validating only the non-script subset would let a
region manufacture alignment by ignoring its awkward items, which is
what block-diagram pages did.

Corpus: 20 of 186 documents, net -359 table rows, no content lost.
Token-level comparison shows the only text changes are merges in the
right direction: 'X' + '10' becomes 'X10', 'L' + 'g' becomes 'Lg' —
subscripts joining their base instead of floating free.

Remaining artifact: MCF5235RM (and its _nxp duplicate) gains a small
spurious table from a block-diagram label line, and M2019_mordeste
gains 11 rows from a 3-column grid re-detected as 2-column. Both are
borderline regions where the previous output was also wrong; documented
rather than tuned away.

805 unit + 148 integration tests pass, clippy clean.

* fix(tables): use a heading-anchored script mask in the body-font pass

Cubic review of #264: a single script mask computed with a 0.0 anchor
was applied to both passes, including body-font region qualification and
geometry. The body pass is supposed to require a heading-sized anchor —
that distinction existed before the geometry rework and was lost in it.

Why it matters: body-pass candidates are themselves body-sized
(0.85..1.05x base). A cell at the low end of that band, say 8.5pt,
sitting beside a 10.5pt label clears the inherent 'anchor >= 1.2x cell'
rule (10.2) and so was flagged as a script attachment. At body sizes a
slightly larger neighbour is a bold label or column header, not the base
of a superscript, so flagging it stripped real cells out of the region
evidence and column geometry and could lose the table entirely.

Two masks now: the small-font pass keeps the 0.0 anchor, the body pass
requires >= 1.15x base. Note the threshold only bites below base size —
for a cell at base, 1.2x-of-cell already exceeds 1.15x-of-base — which
is exactly the 0.85..1.0x band cubic identified.

Corpus: 20 documents, net -367 table rows (was -359 with the single
mask), so the body pass now keeps 8 rows of real table it had been
discarding. 958 tests pass, clippy clean.

* fix(markdown): reject headings that end on a relational verb

A heading candidate ending in 'equals', 'denotes', 'implies' and the
like is the first half of a sentence, not a title. This shows up when a
block dissolves and strands its lead-in ahead of the formula it
introduced — opendataloader 01030000000144 produced

    ## Note that the exact error equals
    M - Q(h) = e - 2.7525... = -0.0342....

Deliberately a very short list. Broader variants were tried and
measured, then rejected:

- Function words (of/and/for/the): a heading that WRAPS across lines
  ends on exactly those. Destroyed real IRS Publication 17 headings —
  'Casualty and' -> 'Casualty and Theft Losses', 'Rule 10. You Must Be
  at' -> '... At Least Age 25'. 52 documents affected, -619 headings.
- Copulas and auxiliaries (is/are/be/have): same failure. 'Rule 15.
  Your AGI Must Be', 'What Medical Expenses Are' and 'When Can a Roth
  IRA Be' are real wrapped headings, while 'the tax burden should be'
  is a genuine fragment. The trailing word cannot separate them; that
  needs the next line's context, which this text-only predicate lacks.

The verbs kept never end a heading in any register, so they are safe
without context. Standalone the guard is a no-op on both benchmarks
(0 documents on opendataloader, 4 on pdf-evals with no net heading
change) — its value is as a companion to the table filter in this PR,
which is what strands these lead-ins.

Combined effect on opendataloader (200 docs, vs a control build of
main), where the table filter alone regressed:

                 table filter    + this guard
    overall        -0.0003          +0.0003
    mhs            -0.0019          +0.0003
    doc ...144     -0.063           +0.053
    doc ...144 mhs -0.203           +0.028

* review: gate the dangling-verb veto on sentence case, drop 'yields'

Cubic review of a5a6e8f — both findings valid.

1. 'yields' is also a plural noun. 'Bond Yields', 'Crop Yields' and
   'Dividend Yields' are real section titles in financial documents,
   which this corpus contains. Removed from the list; my claim that
   these verbs 'never end a heading in any register' was wrong for it.

2. A wrapped title-case heading whose first line ends on one of these
   verbs would be suppressed if the heading preprocessor failed to
   merge it.

Both are fixed by the same gate, which is the discriminator I was
missing: case. A heading is title case ('Bond Yields', 'The Theorem
Implies'); a stranded lead-in is sentence case ('Note that the exact
error equals', 'the method yields'). The veto now applies only when
every content word is NOT capitalized, so titles are spared regardless
of their final word.

This is also why the earlier function-word and copula variants failed:
they had no way to tell 'Rule 15. Your AGI Must Be' from 'the tax
burden should be'. Case separates those two as well.

No measured cost. opendataloader is unchanged from the previous
revision — overall +0.0003, mhs +0.0003, doc 01030000000144 still
0.732 -> 0.785 — and pdf-evals still 20 documents. 963 tests pass,
clippy clean.

* review: exempt section-numbered lines from the dangling-verb veto

Valid ordering bug. heading.rs consults is_heading_fragment at line 282
and only applies its numbered-prefix allowance at line 288, so the veto
pre-empted it: '1. What the model implies' is sentence case and ends on
a listed verb, so it was discarded before numbering could vouch for it.

Numbering is independent evidence of a heading, so the veto now skips
any line opening with a section number.

Acceptance is deliberately a little broader than heading::parse_numbering
(which requires a trailing delimiter) because '2.3 Section Title' is
written without one, and being permissive in a veto exemption can only
avoid suppressing headings. Two guards keep it from swallowing prose:

- a bare single number needs a delimiter ('1.' yes, '3 apples' no)
- roman numerals always need one, since a leading 'I' is the pronoun far
  more often than a section number

Not reused from convert::starts_with_section_number, which deliberately
demands two components because it bypasses isolation checks — that would
reject the reviewer's single-'1.' case.

No measured change: opendataloader still overall +0.0003 / mhs +0.0003
with doc 01030000000144 at 0.732 -> 0.785, pdf-evals still 20 documents,
target case still suppressed. 964 tests pass, clippy clean.

* review: share roman_value so the veto exemption matches the parser

Valid. My numbering predicate accepted tokens heading::parse_numbering
rejects — lowercase 'iv)', alphabetical 'd)', over-long 'MMMM.' — because
it case-folded and allowed D and M. Anything the parser rejects is not
numbering, so exempting it let ordinary list items bypass the
dangling-verb veto and reach font-based heading promotion.

Rather than restate the grammar, roman_value is now pub(super) and the
exemption calls it, so the two cannot drift. Its rules apply as written:
uppercase I/V/X/L/C only, at most 8 characters, positive total.

Decimal numbering keeps its slightly broader acceptance (bare '2.3' with
no trailing delimiter), which is deliberate and documented — that form is
common in real headings and being permissive in a veto exemption cannot
manufacture a heading, only decline to suppress one. The roman case is
different because single letters collide with alphabetical list markers.

No measured change: opendataloader overall +0.0003 / mhs +0.0003, doc
01030000000144 still 0.732 -> 0.785. 964 tests pass, clippy clean.

* test: cover the roman length bound with a nine-character token

Valid P3. The 'MMMM.' case fails on the unsupported M, not on length, so
the 8-character bound in roman_value had no coverage and could regress
silently. Added a nine-'I' token, which is rejected only by the bound,
plus an eight-'I' token that must stay exempt to pin the boundary from
both sides.

* fix(tables): stop dropping body-band scripts from the candidate set

Valid: the body-font pass filtered scripts out of body_candidates
itself, so body_script_flags and its two downstream uses were dead. The
mask filters region_evidence and feeds detect_table_in_region's is_script
closure, but neither ever saw a script item because the candidate set no
longer contained any.

Consequences: a body-band sub/superscript attached to a heading-sized
anchor was dropped from the table outright rather than assigned to a
cell, so its text was lost — the opposite of what both the
body_script_flags comment ('they stay candidates') and the
detect_table_in_region docstring ('they remain eligible for cell
assignment') describe, and inconsistent with the small-font pass.

Root cause: the geometry rework removed the candidate-level filter from
the small-font pass, but the body one had been reflowed onto a single
line by rustfmt so the same edit missed it. Adding body_script_flags in
a later review then wired a mask that the surviving filter made
unreachable.

No measured change on either benchmark — pdf-evals still 20 documents
and -367 table rows, opendataloader still overall +0.0003 / mhs +0.0003
with one document changed — because the combination it affects (a
body-sized script attached to a heading-sized anchor) does not occur in
either corpus. The fix is for correctness and consistency between the
two passes, not for a score.

964 tests pass, clippy clean.
2026-08-07 09:41:46 -07:00
Abimael Martell f731e1191c fix(site): refresh benchmark results (#289) 2026-08-06 12:52:00 -07:00
Andrew Barnes 54a1e9ab74 Preserve page-prefixed Markdown content (#284) 2026-08-06 12:14:27 -07:00
m-naoki-mandClaude Opus 5 fabbb63521 fix(tounicode): skip subset GID remap for CIDFontType0 (CFF) descendants (#209)
* fix(tounicode): skip subset GID remap for CIDFontType0 (CFF) descendants

The sequential-GID repair in try_remap_subset_cmap assumes CIDs are glyph
indices that a subsetter can renumber. That holds for CIDFontType2
(TrueType) but not for CIDFontType0 (CFF), where CIDs are resolved through
the CFF charset, so a valid ToUnicode CMap stays valid after subsetting.

For CFF fonts the corrupting path was unavoidable: CIDToGIDMap is
CIDFontType2-only (PDF 32000-1:2008, 9.7.4.2), so the branch that repairs
the CMap correctly can never be taken, and any CFF font whose /W array
starts at a low CID fell through into remap_to_sequential. Japanese
Adobe-Japan1 documents extracted as long runs of a single unrelated kanji.

Guard both repair paths on a CIDFontType2 descendant, placed before the
CIDToGIDMap branch so a CIDToGIDMap wrongly attached to a CFF font by a
malformed producer is ignored too.

On a National Diet Library proceedings PDF: 1233 U+FFFD in 69099 chars
before, 0 in 68331 after; character 3-gram recall against a hand-written
ground truth 0.354 -> 0.605. The PDF from #118 (CIDFontType2) extracts
byte-identically before and after.

Two existing tests build descendant dicts without a /Subtype and set
CIDToGIDMap, which is CIDFontType2-only, so the fixtures now say what they
already meant. Without that, test_try_remap_skipped_when_w_covers_cmap
would keep passing while no longer exercising the W-coverage logic.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tounicode): resolve indirect /Subtype, skip only explicit non-CIDFontType2

Addresses review feedback on the guard added in the previous commit.

- /Subtype may be an indirect reference, and as_name() does not dereference
  it. A genuine CIDFontType2 font storing /Subtype indirectly would have been
  read as "not CIDFontType2", returning early and losing the repair it needs —
  reintroducing the corruption this PR fixes, for those fonts. Resolve the
  reference through the document before comparing.

- Bail out only when /Subtype is explicitly a non-CIDFontType2 name. A missing
  or unresolvable /Subtype now keeps the pre-existing behaviour instead of
  silently disabling the repair. As a result the two existing tests no longer
  need fixture changes, and this commit reverts those; the diff against main
  is now additive only.

- The CFF regression test now attaches a real CIDToGIDMap stream rather than
  /Identity, which get_cid_to_gid_map treats as "no map". With the stream, the
  test also fails if the guard is moved back below the CIDToGIDMap branch —
  verified by moving it and watching it fail.

- Added test_try_remap_resolves_indirect_subtype.

cargo fmt --check, cargo clippy -- -D warnings and cargo test (862 tests) pass.
The Diet PDF still extracts with 0 U+FFFD and the #118 PDF is still unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 12:02:27 -07:00
585d36e6a6 fix: recover from a corrupted startxref pointer (#230)
* fix: recover from a corrupted startxref pointer

Fixes #228.

A PDF whose startxref pointer has been corrupted to point at the wrong
byte offset — a single flipped digit, which is what damaged writers
emit in the wild — was entirely unprocessable: every entry point
(classify_pdf, extract_pages_markdown, process_pdf) raised "Invalid
PDF structure", even though the file's object data, real xref table,
and trailer were all completely intact just past the wrong pointer.
Both pypdf and pdfium recover from this by locating the real table
directly instead of trusting the pointer; lopdf doesn't.

Added a new repair candidate (alongside the existing
missing-%%EOF-marker and stripped-leading-bytes repairs in
repair_pdf_container_candidates): scan the buffer for the real,
standalone `xref` keyword and append a corrected trailing
`startxref`/`%%EOF` block. lopdf's own get_xref_start always reads the
*last* `%%EOF` in the final 512 bytes of the buffer and the
`startxref` value immediately before it, so the appended block
transparently supersedes the corrupted one already in the file — no
in-place byte surgery on content the original writer produced.

Scoped to classic (non-stream) xref tables, matching the reported
repro and the common case; a corrupted pointer into a cross-reference
*stream* (`N 0 obj << /Type /XRef ...>>`, some PDF 1.5+ writers) would
need the containing object's number, not just a byte offset — out of
scope here.

Verified against the issue's exact repro (a valid one-page PDF with a
single corrupted byte in its startxref offset): before this fix,
process_pdf/classify_pdf/extract_pages_markdown all raised "Invalid
PDF structure"; after, both the page count and the real extracted text
("Order Detail Report by Account", "WIDGET ASSEMBLY", the dollar
amount) come back correctly. New regression test added.

Full suite (859 tests, 1 new) passes; cargo clippy --all-targets
-- -D warnings unchanged at 28 pre-existing/unrelated errors.

* fix: validate xref table shape and scan in a single reverse pass

Addresses cubic-dev-ai's review of #230.

- P2 (correctness/safety): the recovery candidate trusted the last
  standalone "xref" token unconditionally, without confirming it's
  actually a cross-reference table. A coincidental "xref" substring
  inside unrelated content — a stream, a string, uncompressed
  metadata — could get "repaired" against a bogus offset, letting
  lopdf load successfully against garbage instead of returning a
  clean error: a real failure turned into silent data corruption on
  the fallback path. Added looks_like_xref_subsection_header, which
  confirms a plausible classic xref subsection header (`<start-id>
  <count>`, e.g. "0 6" — the shape every real classic table starts
  with) actually follows the candidate token before accepting it.
  find_last_valid_xref_table_start now walks backward from the end of
  the buffer until it finds a token that both stands alone *and*
  validates, rather than accepting the first (rightmost) standalone
  match unconditionally.

- P2 (performance): the old scan re-invoked
  `buf[..search_end].windows(4).rposition(...)` on a shrinking prefix
  every time a candidate token failed the boundary check, which is
  quadratic on a pathological buffer with many non-standalone "xref"
  occurrences. Rewrote as a single reverse byte-index walk — O(n)
  regardless of how many false candidates it has to reject along the
  way.

Added direct unit tests on the byte-level scan (more precise than
constructing adversarial full PDFs, and the coincidental-match
scenario can't be represented in an integration-test fixture anyway
since reportlab compresses page content by default): a coincidental
standalone "xref" with no subsection header is rejected; a real
classic table is found; a coincidental match positioned *after* the
real table in the buffer doesn't shadow it; "xref" as a substring of
"startxref" still doesn't match. The original #228 repro (corrupted
startxref pointer, real table otherwise intact) is unaffected —
verified manually in addition to the existing integration test.

Full suite (863 tests, 5 new) passes; cargo clippy --all-targets
-- -D warnings unchanged at 28 pre-existing/unrelated errors.

* fix: reject xref subsection count runs with trailing garbage

looks_like_xref_subsection_header validated that a count run of digits
followed the whitespace separator, but never checked what came after
it. A coincidental "xref\n0 6garbage" in stream/literal content would
still validate as a real subsection header shape and get repaired
against a bogus offset.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Abimael Martell <1450169+abimaelmartell@users.noreply.github.com>
2026-08-05 16:43:11 -07:00
371de80b14 fix: extract_pages_markdown's needs_ocr now agrees with classify_pdf (#231)
* fix: extract_pages_markdown's needs_ocr now agrees with classify_pdf

Fixes #227.

extract_pages_markdown_mem computed its per-page needs_ocr entirely
from text-quality signals: decoding/garble issues, empty markdown, GID
fonts, garbage-text ratio. It had no awareness of the page's image
content at all — so a page that is fundamentally a full-page scan
with a little genuine native text drawn over it (a header, a stamp, a
cover-sheet annotation) extracts that text cleanly, trips none of the
text-quality checks, and reports needs_ocr=false — while
classify_pdf/detect_pdf_type correctly see the dominant background
image and flag the same page as needing OCR. Two public APIs
answering the same question, silently disagreeing, in the unsafe
direction (skipping OCR on a page that needs it).

Exposed detector::analyze_page_images at crate visibility (was
private) and call it per page in extract_pages_markdown_mem's loop —
the same "large background image" signal (>50% page coverage) that
already powers has_template_image in classify_pdf/detect_pdf_type,
rather than reimplementing image-area detection a second time with
its own thresholds that could drift out of sync again. When it's
true, the page is flagged needs_ocr (with OCR_REASON_SCANNED added
to ocr_reasons_by_page, matching how the same signal is already
reported elsewhere) and its markdown is blanked, exactly like the
existing text-quality-triggered needs_ocr paths already do — no
special-casing added for "cleanly-extracted-but-still-a-scan" text.

Verified against the issue's exact repro (a full-page raster with one
native text line drawn over it, built via reportlab/pillow): before
this fix, extract_pages_markdown_bytes reported page 0
needs_ocr=False with the header line as markdown while
classify_pdf_bytes correctly flagged pages_needing_ocr=[0]; after,
both agree needs_ocr=True and the page's markdown is empty. Confirmed
no regression on a normal text-based fixture (nexo-price-en.pdf:
needs_ocr stays False, full markdown returned). New Rust regression
test added exercising both APIs against the same fixture.

Full suite (860 tests, 1 new) passes; cargo clippy --all-targets
-- -D warnings unchanged at 28 pre-existing/unrelated errors.

* fix: gate has_template_image behind the same OCR signals classify_pdf uses

extract_pages_markdown_mem was treating has_template_image alone as
sufficient to force needs_ocr=true and discard the page's markdown, but
classify_pdf/detect_pdf_type never treats that raw signal alone as
needing OCR. A text page with a full-bleed watermark, letterhead, or
large figure would get its clean markdown wrongly blanked and routed
to OCR.

Added page_template_image_needs_ocr(), mirroring the two distinct
signals classify_pdf actually uses to decide a template-image page
needs OCR: the looks_like_scan gate (image_count <= 1, few text ops,
low alphanumeric diversity) used for Mixed-type routing, and the
insufficient-text-volume signal (text_operator_count < 10) that routes
a page with a dominant background image and only a couple of native
text calls to PdfType::ImageBased independent of looks_like_scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: match per-page OCR threshold and add missing vector-text signal

Two follow-up findings on the has_template_image gate added in the
previous commit:

1. insufficient_text used a hard-coded threshold of 10 text operators,
   but Mixed-type per-page routing (the actual per-page decision this
   function tries to agree with) uses config.min_text_ops_per_page
   (default 3). The higher 10 threshold was borrowed from a *different*
   classify_pdf code path — the effective_min_ops floor used only for
   whole-document ImageBased/Scanned classification, a cross-page
   aggregate this per-page function can't replicate anyway. Using the
   lower per-page threshold removes a real disagreement window
   (3-9 text ops with high alphanumeric diversity) without breaking
   the #227 regression fixture (text_ops=1, still well under 3).

2. extract_pages_markdown_mem never checked has_vector_text at all,
   even though Mixed-type per-page routing always sends
   vector-outlined-text pages to OCR (outlined glyphs can't be
   extracted as text). A page with massive path ops plus a short
   genuine caption could extract that caption cleanly, slipping past
   the existing empty/garbage-text checks. Added
   page_has_vector_text() and wired it into needs_ocr the same way
   has_template_image is.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* perf: compute template-image and vector-text OCR signals in one pass

page_template_image_needs_ocr and page_has_vector_text each called
analyze_page_content independently, so every requested page's content
streams (page + XObjects) and image coverage were decompressed and
scanned twice per page with one result discarded each time.
detect_from_document avoids this by caching its per-page PageAnalysis;
extract_pages_markdown_mem had no such cache.

Merged both into page_ocr_signals(), a single analyze_page_content
call returning both signals as a tuple.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Abimael Martell <1450169+abimaelmartell@users.noreply.github.com>
2026-08-05 15:41:44 -07:00
Abimael Martell ede48099c0 fix(layout): preserve ruled tables and chart prose order (#262)
* fix(layout): preserve ruled tables and chart prose order

* fix(layout): harden chart region detection

* fix(layout): tighten chart geometry guards

* fix(layout): bound chart inference

* fix(layout): tighten chart claim bounds

* fix(layout): tighten chart evidence

* fix(layout): preserve edge-adjacent chart labels

* fix(layout): require external chart label overlap
2026-08-05 10:22:13 -07:00
Abimael Martell 12e9a655e3 fix(extractor): supply built-in metrics for non-embedded base-14 fonts (#241)
* fix(extractor): supply built-in metrics for non-embedded base-14 fonts

PDFs may legally omit /Widths for non-embedded standard fonts (Times,
Helvetica, Courier, Symbol, ZapfDingbats) — the spec requires the reader
to supply the metrics. We returned None, so every glyph advanced 0 and
each text item got width 0, silently breaking every gap-based heuristic
downstream: space synthesis, sub/superscript detection, table column
detection, heading merging.

- src/extractor/base14.rs: Adobe Core-14 AFM width tables keyed by
  Unicode char, plus the standard Symbol/ZapfDingbats encoding vectors
  (their glyphs sit at byte positions unrelated to Latin text, so widths
  must resolve through the built-in encoding, not cp1252)
- Width resolution order: Differences -> built-in encoding -> the same
  cp1252-style fallback the text decoder uses, so a code's advance always
  matches the character we emit for it
- Type3 visual sizing: PK bitmap fonts (dvips) use FontMatrix
  [1 0 0 -1 0 0] with nominal sizes like 0.12pt; scale by FontBBox height
  x |matrix_y|. Applied in the page-stream and Form XObject paths.
  Indirect numeric array elements are resolved before use.

Effect on Shannon's 'A Mathematical Theory of Communication' (1998
dvips/Distiller, the reported case): glued sentences 95 -> 5. Corpus
impact: 12 of 184 eval documents, e.g. Data-Processing-Agreement
recovers a paragraph that a phantom table had shredded into cells.

Layout heuristics tuned on the same document (indent-based paragraph
breaks, heading reclassification, table script filtering) are held back
for a separate PR — they change ~98 further documents and need to be
justified against the corpus, not against one PDF.

* review: narrow Type3 rescaling to self-inconsistent fonts; dedup + test all width tables

Addresses cubic review on #241, plus a follow-up from a local cubic run.

- Type3 visual scaling was applied to every Type3 font whose FontBBox
  height x |matrix_y| deviated >5% from 1.0. FontBBox is the glyph box,
  not the em box, so a conventional 1/1000-matrix font with a
  descender..ascender bbox (~700 units) computed 0.7 and had every
  reported size shrunk by 30% — corrupting the drop-cap, heading-tier,
  sub/superscript and table heuristics this is meant to fix.

  First attempt gated on the matrix being unit-scale, but a local cubic
  run pointed out that wrongly excludes valid non-standard matrices (a
  0.005 matrix with a full-em bbox legitimately needs a 5x scale). The
  product is the right discriminator, not the matrix: a self-consistent
  font lands near 1.0 because the matrix is the reciprocal of the
  glyph-space em, so only a wildly inconsistent one (dvips/PK bitmap
  fonts sit at ~159) is renormalized. Band widened to [0.25, 4.0].

  Corpus effect: 12 -> 7 documents change. The 5 that drop out were
  being wrongly rescaled — including Data-Processing-Agreement, whose
  phantom-table fix turned out to come from this bug rather than from
  the width fallback, so it is correctly given up.

- base14: all 14 width tables now covered by the sort-invariant test via
  an ALL_TABLES registry, not a hand-picked subset.
- base14: identical tables share one static (all four Courier variants
  are monospace 600; the oblique Helvetica variants match their upright
  forms), removing 5 duplicate copies.

* test: refresh Shannon snapshot after merging main

CI checks out a merge of the PR head with main, and main advanced 8
commits since this branch was cut — including #201 (contextual digit
runs), #240 and #253 (markdown fixes). Those change extraction output,
so a snapshot generated on the unmerged branch could not match; the
Test job failed on the merge commit while passing on the branch itself.

The merged behaviour is better: the footnote marker '2' before
'Hartley, R. V. L.' is now recovered instead of dropped.

950 tests pass on the merged tree, clippy clean.
2026-08-04 18:00:43 -07:00
Abimael MartellandClaude Fable 5 1d134e26aa docs: sync AGENTS.md with CLAUDE.md, refresh eval workflow guidance (#243)
AGENTS.md was stale (179+ PDFs, missing the semantic-quality bullet).
Both files now match: ~200-PDF corpus, and iteration guidance to prefer
subset runs (bench.py test -q / -s <name>) with the full suite as the
final pre-commit check.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 17:26:04 -07:00
Shubham Mathur bfd6c3eabb fix(extractor): make the comment stripper escape-aware (#259)
strip_pdf_comments tracked parenthesis nesting to protect string
literals, but ignored backslash escapes. An escaped \) desynced the
depth counter, after which a % glyph inside a string was stripped as a
top-level comment, corrupting the stream for Content::decode and
silently truncating the page's text.

Treat \ inside a string literal as escaping the next byte, so \(,
\), and \\ never touch the nesting depth.
2026-08-04 16:22:25 -07:00
Sheroy Cooper 04abab951f Support password-protected PDF item JSON extraction (#245) 2026-08-04 12:41:22 -07:00
Sunil a410d5aa08 fix: sync Python type stubs (.pyi) with actual bindings and fix doc bugs (#250)
## Bug 1: Python type stubs missing 5 fields/classes

The pdf_inspector.pyi file was out of sync with the actual Python
bindings exposed via #[pyo3(get)] in src/python.rs. This breaks
IDE autocompletion and type checking (PyCharm, VS Code/Pylance, mypy)
for all Python users.

Added:
- PdfResult.ocr_reasons_by_page (python.rs:35)
- PageOcrReasons class with page and 
easons fields (python.rs:67-86)
- RegionText.ocr_reason (python.rs:136)
- PageMarkdown.ocr_reason (python.rs:193)
- PagesExtractionResult.ocr_reasons_by_page (python.rs:226)

## Bug 2: PdfResult.pages_needing_ocr indexing undocumented

PdfResult.pages_needing_ocr is 1-indexed (per python.rs:30) but
neither the .pyi stubs nor docs/python.md annotated this, while the
same field on PdfClassification was annotated as 0-indexed. Users
mixing both APIs would get wrong page numbers.

## Bug 3: README.md duplicate bullet character

The Markdown features table listed * twice in bullet prefixes.
The first should be ullet (U+2022), matching the actual source code in
src/markdown/mod.rs:34 which uses: ullet, dash, sterisk, circle, illed-circle, open-circle.

## Bug 4: docs/python.md missing fields in type reference

The Types section was missing PageOcrReasons, RegionText class
definition, ocr_reason fields, and ocr_reasons_by_page fields.

## Evidence

Cross-referenced every #[pyo3(get)] attribute in src/python.rs
against the .pyi declarations and docs/python.md type reference.
2026-08-04 12:36:30 -07:00
Abimael Martell 7747b3a086 fix(markdown): reject non-text strikeout rules (#253)
* fix(markdown): reject non-text strikeout rules

* fix(markdown): address strikeout ownership edge cases

* fix(markdown): reject connected filled strike rules

* fix(markdown): group drifted strikeout runs
2026-08-04 11:31:58 -07:00
ae6246ba0c fix(markdown): preserve redline edits in prose (#240)
* fix(markdown): preserve redline edits in prose

* fix(tables): preserve live evidence near redlines

* fix(tables): scope redline table suppression

* fix(tables): preserve revised cell evidence

* fix(tables): retain page context for redlines

* fix(tables): propagate revised column evidence

* fix(tables): preserve redline boundaries in headers

* fix(tables): separate redline suppression spans

* fix(tables): tighten revised cell evidence

---------

Co-authored-by: Bryan Nathan <bryan@users.noreply.github.com>
Co-authored-by: Abimael Martell <1450169+abimaelmartell@users.noreply.github.com>
2026-08-03 20:08:39 -07:00
Abimael Martell 3fb545284b fix(layout): preserve contextual digit runs (#201)
* fix(layout): preserve contextual digit runs

* fix(layout): harden page folio filtering

* fix(markdown): distinguish folios from contextual numbers

* fix(layout): distinguish running folios from contextual digits

* fix(layout): tighten running folio evidence

* test(layout): guard the folio evidence floor

* fix(layout): preserve page-number decisions across partitions

* fix(layout): harden document-level folio filtering

* fix(markdown): carry folio context across public APIs

* fix(layout): isolate folios from layout metadata

* fix(markdown): preserve table cells in per-page extraction

* fix(layout): preserve folio context with page filters

* fix(layout): handle contextual folio sequences

* fix(layout): tighten adjacent folio evidence

* fix(layout): constrain folio context inference

* fix(forms): resolve widget pages from annotations
2026-08-03 16:35:52 -07:00
Jase-Omeileo West 1c32e4bd69 fix(deps): bump lopdf to 0.42.0 for nesting-depth DoS (#198)
Bumps lopdf from 0.41.0 to 0.42.0 to fix RUSTSEC-2026-0187, preventing deeply nested PDFs from causing an unrecoverable stack-overflow process abort.
2026-08-03 11:08:34 -07:00
Abimael Martell 8121ae97ce chore(ci): bump GitHub Actions to current majors (#237)
* chore(ci): bump GitHub Actions to current majors

Node 20 action runtimes are deprecated on GitHub runners; bump every
first-party action to its latest major across all workflows:

- actions/checkout v4/v6 -> v7
- actions/cache v4 -> v6
- actions/upload-artifact v4 -> v7, download-artifact v4 -> v8
- actions/setup-node v6 -> v7, setup-python v5 -> v7
- actions/upload-pages-artifact v3 -> v5, deploy-pages v4 -> v5

Third-party pins (dtolnay/rust-toolchain, Swatinem/rust-cache,
setup-zig, setup-bun, taiki-e/install-action, maturin-action) are
already on their latest majors.

* chore(ci): pin all actions to full commit SHAs

Mutable @vN tags can be retagged; in the publish workflows that code
runs with OIDC credentials before npm/PyPI/crates.io publishes. Pin
every action (first- and third-party) to its release commit SHA with
the version in a trailing comment.

dtolnay/rust-toolchain infers the toolchain from its ref name, so the
SHA-pinned invocations pass an explicit toolchain: stable input.
2026-08-03 11:07:36 -07:00
Abimael Martell 98990cc550 feat(npm): add Linux musl and ARM64 N-API binaries (#233)
Adds x86_64-unknown-linux-musl, aarch64-unknown-linux-gnu, and
aarch64-unknown-linux-musl to the napi build targets so
@firecrawl/pdf-inspector works on Alpine and ARM64 Linux deployments.

- gnu arm64 cross-compiles with --use-napi-cross (old-glibc sysroot),
  musl targets with -x (zig + cargo-zigbuild), per the napi-rs template
- new platform packages carry npm libc metadata (glibc/musl)
- smoke-test job runs napi/test.mjs on all six targets before publish
  (Alpine containers for musl, ubuntu-24.04-arm runners for ARM64)
- bump to 1.12.0 to trigger publishing of all platform packages

Closes #216
2026-08-03 10:09:23 -07:00
Abimael Martell a15ec2d68d docs(readme): refresh local parser benchmark results (#192)
* docs(readme): refresh parser benchmark results

* docs(readme): include stored parser results

* docs(readme): refresh remaining parser results

* docs: sync benchmark references
2026-07-31 22:58:46 -06:00
Abimael Martell 7b7960ee73 chore(pypi): bump pdf-inspector to 0.2.6 (#188) 2026-07-31 17:12:34 -06:00
Abimael Martell 7188667045 chore(wasm): bump @firecrawl/pdf-inspector-wasm to 0.1.3 (#190) 2026-07-31 16:43:52 -06:00
Abimael Martell 6ff104409a chore(npm): bump @firecrawl/pdf-inspector to 1.11.2 (#189) 2026-07-31 16:37:50 -06:00
Abimael Martell 5e8f1570f6 chore(crate): bump pdf-inspector to 0.1.7 (#187) 2026-07-31 16:25:57 -06:00
tomsideguide b31e4b1727 fix(extractor): don't flag gid Differences names covered by ToUnicode (#186)
* fix(extractor): don't flag gid Differences names covered by ToUnicode

Pages were marked as having unresolvable gid-encoded fonts whenever any
font's /Differences array used gidNNNN glyph names, and when every page
carried such a font the whole document's markdown was suppressed.
LibreOffice exports do exactly this: subset fonts get /gidNNNN names in
Differences alongside a complete ToUnicode CMap that decodes them, so
ordinary text documents lost their entire markdown output even though
extraction decoded every glyph.

Track the character codes behind the gid names and only flag the font
when its ToUnicode CMap addresses none of them. Partially mapped codes
stay unflagged: an emoji ZWJ sequence maps whole on its first code, and
the remaining component-glyph codes are subset leftovers, not damage.
Fonts without ToUnicode, or whose CMap ignores the gid codes, are
flagged as before, and the downstream garbage/encoding checks still
catch partial breakage.

* fix(extractor): require a usable ToUnicode mapping to clear the gid flag

A mapping to U+FFFD (or an empty string) is rejected by extraction as
an invalid CMap result, so it must not count as decodable when deciding
whether gid-named Differences codes are resolvable.
2026-07-31 15:54:19 -06:00
Abimael Martell 3c6eb8bf6b feat(site): add local WebAssembly demo (#181)
* feat(site): add local WebAssembly demo

* feat(site): parse PDFs on selection
2026-07-17 15:00:47 -07:00
Abimael Martell 5b287341a0 feat(wasm): add browser bindings (#180)
* feat(wasm): add browser bindings

* fix(wasm): address review feedback

* fix(wasm): preserve numeric plain text

* chore(wasm): prepare 0.1.2 release
2026-07-17 14:09:43 -07:00
Abimael Martell 55d50ad1c4 docs(site): redesign open-source project page (#179)
* docs(site): redesign open-source project page

* docs(site): lead with node and cli

* docs(site): emphasize package registries

* docs(site): align footer wordmark

* docs(site): remove hero terminal scrollbar

* docs(site): keep hero install command inline

* docs(site): feature rust core in hero

* docs(site): add hero language playground

* Revert "docs(site): add hero language playground"

This reverts commit 16df5fcf07.
2026-07-16 22:55:51 -07:00
Abimael Martell a910b7df1d docs(benchmark): refresh parser comparison (#178)
* docs: refresh benchmark comparison

* docs(site): refresh benchmark section

* docs: reframe benchmark positioning

* docs: focus benchmark positioning on best fit
2026-07-16 17:43:53 -07:00
Abimael Martell 15c0b22093 test(bench): probe optional backend evidence (#176)
* test(bench): probe optional backend evidence

* fix(bench): accept native stext pages
2026-07-16 15:25:37 -07:00
Abimael Martell 0c06dac976 test(bench): compare OpenDataLoader builds (#175)
* test(bench): compare OpenDataLoader builds

* docs(bench): keep reference comparisons generic

* fix(bench): keep regression gates complete

* fix(bench): clarify missing reference gates

* fix(bench): validate nonnegative limits

* fix(bench): isolate prediction runs

* chore(bench): refresh review
2026-07-16 14:25:25 -07:00
Abimael Martell 64a0930f9f feat(layout): order image-anchored regions (#174)
* feat(layout): order image-anchored regions

* fix(layout): preserve region flow boundaries

* fix(layout): gate image-backed column flows
2026-07-16 12:37:54 -07:00
Abimael Martell c726a92435 feat(tables): score competing line-table candidates (#177)
* feat(tables): score competing line-table candidates

* fix(tables): validate competing hypotheses

* fix(tables): preserve multiline open-edge headers

* fix(tables): recover two-column open-edge grids

* fix(tables): retain multiple open-edge grids
2026-07-16 11:32:37 -07:00
Abimael Martell 1b5ec414a4 feat(tables): compete normalized table and chart evidence (#172)
* feat(tables): compete normalized table and chart evidence

* fix(tables): preserve normalized hypotheses
2026-07-16 03:15:05 -07:00
Abimael Martell a0a6a445bf feat(tables): infer columns from horizontal-rule text anchors (#171)
* feat(tables): infer columns from horizontal-rule text anchors

* fix(tables): scope sparse-rule vertical evidence

* fix(tables): validate sparse-rule evidence by content

* fix(tables): preserve independent line-table regions

* fix(tables): isolate sparse-rule regions
2026-07-16 01:58:21 -07:00
Abimael Martell 4c71feb2f1 feat(markdown): classify document heading sequences (#170)
* feat(markdown): classify document heading sequences

* fix(markdown): harden heading sequence evidence

* fix(markdown): validate fixed-size sidebar evidence

* fix(markdown): harden sidebar sequence guards
2026-07-15 23:07:59 -07:00
Abimael Martell 660ffe4a0c feat(markdown): gate chart pages before table detection (#169)
* feat(markdown): add chart-aware layout gating

* fix(markdown): keep chart labels in separator zones

* fix(markdown): harden chart layout gating

* fix(markdown): order chart page blocks by stream

* fix(markdown): narrow parallel prose rejection

* fix(markdown): require cross-row prose evidence

* fix(markdown): classify chart labels by geometry

* fix(markdown): harden chart block ordering

* fix(markdown): cache chart blocks and retry tables

* fix(markdown): preserve chart order with physical bands
2026-07-15 21:08:58 -07:00
Abimael Martell b084769fda feat(markdown): add fidelity output profile (#168) 2026-07-15 17:18:16 -07:00
Abimael MartellandClaude Fable 5 f741e49dec fix(headings): all-bold single words qualify as headings when standalone (#167)
Single-word bold section headings ('Replace', 'Trash', 'Instructions')
required a paragraph break before AND after, but headings hug their
section's first paragraph — the break-after almost never exists. A
standalone all-bold single word (>=4 chars, paragraph break before or
page top) now classifies; mixed bold lead-ins ('Note: ...') stay
excluded via all_bold.

opendataloader-bench: 0.8567 -> 0.8575, MHS 0.773 -> 0.776; docs 145
+0.118, 069 +0.112 (net of one cover-page layout shuffle at -0.066
where the new output is semantically closer to GT). pdf-evals: 66
snapshots, composite 0.5864 -> 0.5883, sole >0.02 mover positive.
p1244/thermo fixture snapshots regenerated ('Instructions' un-fuses
from its body paragraph — the intended behavior).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 14:44:46 -07:00
Abimael MartellandClaude Fable 5 0f898a18fb fix(headings): digit-only lines do not define heading tiers (#166)
* fix(headings): digit-only lines do not define heading tiers

A large bold page number (14pt folio over 11pt body) claimed tier 0:
every real heading demoted one level document-wide, and the bold-size
fallback (which requires an empty tier list) was blocked for documents
whose headings match body size.

Bench-neutral (MHS scores relative hierarchy); pdf-evals: 18 docs get
their heading levels back (#### -> ###), semantic composite +0.0006,
no percentile down.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(headings): exclude digit-only lines from the bold fallback tier pass too

The exclusion in the main pass wasn't enough: with the page-number
tier gone, the bold fallback re-collected the same bold folio. Also
regenerates the thermo-freon12 snapshot (cosmetic churn on the
scrambled legend fixture).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:42:41 -07:00
Abimael MartellandClaude Fable 5 c908e33b39 fix(fonts): remap misnamed TeXCMMathsSymbols glyphs to their true symbols (#165)
IntechOpen-family academic PDFs embed Computer Modern math symbol
subsets whose glyphs are misnamed after Latin lookalikes (equal →
/onequarter, plus → /thorn, parens → /eth //Thorn) — and the generated
ToUnicode faithfully propagates the wrong names, so formulas decode as
'S ¼ kB þ 1' instead of 'S = kB + 1'.

Remap the observed misnames, gated strictly on the TeXCMMathsSymbols
base font (subset prefix stripped) so genuine fractions and thorns in
text fonts are untouched. Known limitation: a sibling subset misnames
the slash as /onequarter too, so an occasional '/' renders as '=' —
still strictly better than the previous mojibake.

Affects 4 bench PDFs (028/031 +0.001-0.004 NID) and zero pdf-evals
snapshots.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:19:34 -07:00
Abimael MartellandClaude Fable 5 4aa2c0c208 fix(headings): rescue wrapped bold headings on interleaved column pages (#164)
* fix(headings): rescue wrapped bold headings on interleaved column pages

Report pages whose columns can't be detected (6pt gutters) interleave
both columns' lines, which breaks every whitespace signal the bold
heading heuristic relies on: para_threshold inflates to ~3x line
height and a wrapped heading's own internal line gap defeats isolation.
'9.5. Adapting to the New Normal: Changing / Business Models' merged
into the following paragraph.

Four changes:
- merge_wrapped_bold_heading_groups: 2-3 consecutive all-bold
  body-size lines merge into one line when the group is isolated
  (column-locally, judged by x-overlapping lines only) or starts with
  a section number.
- Section-numbered all-bold lines ('9.5. ...') classify as headings
  without the standalone/isolation score gate.
- Line unfusing extends to uppercase-start continuations, gated on a
  bold-style mismatch between the runs (a bold heading beside regular
  body text) — same-style label rows stay joined.
- The unfuse line-side wordiness requirement drops to 2 words so a
  wrapped heading's short last line ('Business Models') still splits
  from the neighboring column.

opendataloader-bench: overall 0.8554 -> 0.8576, MHS 0.769 -> 0.777;
docs 037 +0.161, 111 +0.157, 039 +0.091, 198 +0.028, none down.
pdf-evals: 63 snapshots, composite 0.5952 -> 0.5964, sole >0.02 mover
positive. thermo-freon12 snapshot regenerated (cosmetic churn on an
already-scrambled 3-column legend).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(headings): review follow-ups — multi-component section numbers, wholly-bold line gate

Single '1. ' prefixes are ordered list items and no longer bypass
isolation; the uppercase unfuse requires the whole line bold (a
heading), not merely its last run, so mixed bold-label/value rows
stay joined.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:00:54 -07:00
Abimael MartellandClaude Fable 5 31918ff62f fix(layout): unfuse independent column runs sharing a baseline (#163)
* fix(layout): unfuse independent column runs sharing a baseline

Two-column report pages with charts fused headings into the adjacent
column's body text: the columns' ~6pt gutter is below what histogram
valley detection can safely use, so the page grouped single-column and
same-baseline items from both columns joined into one line ('6.2.
Expectations for Re-Hiring Employees' + mid-sentence text, killing
MHS and NID on the whole survey-report doc family).

Three changes:
- Line grouping splits same-baseline runs separated by a wide void
  (>3x font size, >=30pt) when the incoming run starts lowercase
  (mid-sentence continuation from another column) and both sides are
  multi-word prose. TOC page numbers, dot leaders, and table cells
  (numbered/capitalized) stay joined.
- Column detection is blind to chart-region text (tight 2pt bounds —
  wider padding ate rows adjacent to charts), via a chart-aware line
  grouping variant wired from the markdown pipeline.
- validate_and_build_columns computes its vertical span from
  histogram-eligible items only, so full-width captions no longer sink
  the overlap ratio for partial-page column regions.

opendataloader-bench: overall 0.8532 -> 0.8554, MHS 0.761 -> 0.769;
doc 038 +0.434, no regressions. pdf-evals: 34 snapshots change,
semantic composite wash (0.5749 -> 0.5748), no per-doc mover >0.015.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(layout): review follow-ups — chart-aware band-split grouping, single chart scan per page

Band-split pages now route through the chart-aware grouping too, and
the band loop reuses the precomputed page_chart_map instead of
re-scanning the rect list per page.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 11:58:14 -07:00
Abimael MartellandClaude Fable 5 4c3330f93a fix(headings): bold-size fallback tiers when nothing clears the ratio gate (#162)
* fix(headings): bold-size fallback tiers when nothing clears the ratio gate

Books often set section headings barely above body size (11pt bold
over 10pt text). Nothing cleared the 1.2x heading-tier ratio gate, so
tiers stayed empty and every bold heading defaulted to H2 — H1 was
unreachable for the whole document.

When no size clears the gate, build tiers from bold line sizes >=1.05x
body, and let tier matches through detect_header_level down to that
ratio. Documents with real (>=1.2x) tiers are untouched.

Bench-neutral by construction (the MHS metric scores relative
hierarchy, not absolute levels); pdf-evals semantic composite +0.004
on the 11 affected docs with all percentiles up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(headings): require boldness for sub-gate tier matches

Review follow-up: fallback tiers come from bold lines, so honoring
them for non-bold text at the same size would promote captions.
detect_header_level now takes is_bold and only matches tiers below
the 1.2x gate for bold lines; >=1.2x matches stay bold-agnostic.
Also restores the >=1.2x tier-match loop the refactor dropped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(headings): judge line boldness by character mass

Review follow-up: a heading with an unbold section-number prefix
('4. ' + bold title) failed the first-item boldness test. Judge the
line by bold character mass instead, at all three call sites.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 10:14:37 -07:00
Abimael MartellandClaude Fable 5 7d108cdff1 fix(extractor): use center-y for off-box link filtering (#161)
Link items carry an annotation rect, so y is a box edge — unlike text
items, where y is a baseline. Testing rect-bottom dropped partially
visible links whose bottom edge dipped past the tolerance. Follow-up
to a #160 review comment.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:14:38 -07:00
Abimael MartellandClaude Fable 5 c4cbf49f44 fix(extractor): clip page content to the visible page box (#160)
* fix(extractor): clip page content to the visible page box

Single-page extracts and imposed spreads keep neighboring pages'
content in the stream, positioned outside the CropBox. Extracting it
appends invisible sections to the page, scrambles NID, and poisons
font statistics (heading tiers built from off-page text).

Clip items (by center), and — only when off-page text was actually
found — rects and lines (by overlap) to CropBox-else-MediaBox, walking
page-tree inheritance. Rotated pages are left unclipped: their item
coordinates are already transformed out of box space. Degenerate boxes
(<1 inch) are ignored.

opendataloader-bench: overall 0.8445 -> 0.8537, NID +0.008,
MHS +0.013; six docs up (best +0.426), none down.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(extractor): guard page-box clipping with coherence and straddle checks

Two real-document counterexamples: curved display text leaves short
glyph fragments with artifact coordinates outside the box (judge by
character mass, not item count), and some PDFs compute inflated
coordinates for visible body text (an off-page item continuing an
on-page baseline means our transform model is wrong there — skip).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(extractor): clip off-box link annotations when page text was clipped

Review follow-up: annotations from the neighboring page bypassed the
filter. Form fields are left as-is — they're document-scoped and rare
on imposed spreads.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:07:52 -07:00
Abimael MartellandClaude Fable 5 2ed152e49b fix(tables): recognize chart-bar clusters and mask their text from table detection (#159)
Bar charts drawn as filled rects read as cell rects or aligned text:
the cell-rect fallback gridded their axis labels into phantom tables,
and when rect paths rejected them, hint regions and the gap-histogram
heuristic re-gridded the same text. On survey-report pages this
scrambled reading order and swallowed section headings.

Detection (is_chart_bar_cluster): the dominant equal-breadth rect
family arranged in >=2 spaced positions (bars; cell rects touch),
data-driven extent variation (>=1.3x), no same-offset/same-extent
partners across positions (grid rows pair up, chart segments don't),
and only numeric labels inside. Mirrored predicate covers horizontal
bar charts.

Chart clusters are skipped in detect_tables_from_rects (no table, no
hint), and a new detect_chart_regions pass lets the markdown pipeline
pre-claim chart items so heuristic/line/column detection and the
merged-band retry all skip them — the text flows out as plain lines.

opendataloader-bench: overall 0.8389 -> 0.8446, TEDS 0.699 -> 0.708,
MHS 0.742 -> 0.750, NID 0.889 -> 0.894; 7 docs up (best +0.459), none
down. pdf-evals changed-set composite +0.013, TEDS +0.036.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 00:06:17 -07:00
Abimael MartellandClaude Fable 5 e48e34dfe9 docs(tables): document the stacked-box 6-rect precision gate (#158)
* docs(tables): document the stacked-box 6-rect precision gate

Review follow-up on #157: 3-5 box stacks never reach the stacked-box
fallback (the main loop needs >=6-rect clusters on a >=6-rect page).
Routing smaller clusters through the detector was implemented and
measured: zero opendataloader-bench movement and four pdf-evals
regressions (striped bullet lists, wrapped regulation text, stats-table
columns) across three guard iterations — with 3-5 boxes the anti-prose
guards have too little signal. Keep the gate, document it at the call
site, and pin the behavior with an end-to-end test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: exercise the cluster gate, not just the page gate

Review follow-up: the pinned test's 3-rect page exited at the 6-rect
page gate before reaching the cluster minimum it documents. Scattered
unrelated rects now push the page past the page gate while the 3-box
stack stays below the cluster minimum.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 22:52:49 -07:00
Abimael MartellandClaude Fable 5 3244547b3f feat(tables): detect single-column stacked-box tables (#157)
A framework list drawn as a vertical stack of boxes (one short title
per box) scored TEDS 0 and worse, ran into the surrounding prose as a
single paragraph: the grid path rejects one-column rect structures by
design (needs >=3 x-edges).

Add a stacked-box fallback after grid and row-stripe detection: >=3
x-aligned, same-width, same-height boxes forming a contiguous vertical
stack, each holding one short text run, become a single-column table.

Guards against striped prose and grid fragments (all unit-tested):
- boxes flanked by rects or text at their y-level are one column of a
  wider structure — bail to the grid/cell-rect paths
- multiple separated text runs per box = striped multi-column content
- prose rows: function-word-dense cells averaging >60 chars
- sentence continuation across rows (trailing comma / open + lowercase)
- numbered/lettered list items stay lists

opendataloader-bench: overall 0.8362 -> 0.8389, TEDS 0.675 -> 0.699,
target doc +0.553, no other doc moved.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 22:26:22 -07:00
Abimael MartellandClaude Fable 5 2f5a8923dd fix(layout): veto side-by-side split that cleaves a rect table (#156)
A 4-column compliance table (labels + Small/Medium/Large) was emitted
as two separate tables: split_side_by_side read the text gap between
the ruled and remaining columns as a page-layout gutter, and each band
then detected its own fragment.

Before accepting a side-by-side split, check rect clusters near the
boundary: if a table-shaped cluster spans it (or ends at it with
cell-like text row-aligned beyond), and those table rows account for
the majority of far-side text, the split runs through a table — veto
it. The majority guard keeps legitimate splits on pages where a figure
spans two prose columns.

opendataloader-bench: overall 0.8306 -> 0.8362, TEDS 0.656 -> 0.675,
target doc +0.506, no regressions.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 21:46:38 -07:00
Abimael MartellandClaude Fable 5 20bb22aa3c fix(tables): flatten page-number tables of contents instead of gridding (#142)
* fix(tables): flatten page-number tables of contents instead of gridding

A title-based contents page ("About the Publisher  vii", "Experiment #1
… 3") with no dot leaders and no section numbers was detected as a
2-column data table and rendered as a markdown grid, scrambling the
linear reading order (a top cause of NID loss on affected docs) and
scoring 0 on table structure.

Add is_page_number_toc: a narrow (2-3 col) list whose last column is
mostly page numbers (short integers or roman numerals) that are mostly
non-decreasing, with a text-title first column and NO header row (a
TOC's first row is already an entry). Such tables now route through the
existing flat-list TOC renderer.

The no-header + narrow-width + monotonic guards keep real data tables
intact — e.g. a 4-column regional table, or a 2-column "Mineral | CEC"
table with a header row and ascending values.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: bump reading-order (NID) benchmark to 0.89

Reflects the phantom-TOC fix in this PR: NID 0.88 -> 0.89 on the
200-doc benchmark. Other cells are unchanged at 2-decimal precision.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: canonical roman validation, real-first-row header check, roman page cells

- page_number_value now requires a *canonical* roman numeral (re-encode
  and compare), so words like "civil"/"mix"/"ill" are no longer parsed
  as page numbers.
- The no-header guard checks the actual first row's last cell instead of
  the first non-empty one, so a blank header cell ("Category | ") still
  rejects the TOC heuristic.
- format::is_page_number_cell recognizes canonical roman numerals, so
  roman front-matter pages (vii, ix) get proper title/page separation in
  the flat TOC list.
- Fix the non-monotonic test to use 5 rows so it exercises the
  monotonicity guard rather than the row-count early return; add
  roman-lookalike and blank-header rejection tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: share roman helper, widen length to 8, require page-span for TOC

- Extract canonical_roman_value + to_roman_lower into tables/mod.rs and
  use them from both the TOC detector and the formatter, removing the
  duplicated mapping/loop and keeping them in sync. The shared helper
  accepts ≤8 chars, so longer front-matter numerals (xxxviii) flatten
  consistently on both sides.
- Add a page-span guard to is_page_number_toc: real page numbers skip
  through the document (range >> entry count), so a dense consecutive
  ordinal/rank/ID column (1,2,3,…) is rejected — monotonicity alone did
  not separate those data tables from contents.

Costs ~0.001 aggregate on the benchmark (NID 0.888->0.887) for the added
precision; still a clear win over baseline (NID 0.883, TEDS unchanged).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: recover consecutive-page TOCs via a title signal

The strict page-span rule rejected legitimate one-page-per-entry TOCs
(range ~= entry count). Relax it: accept any page sequence with a gap
(nearly all real contents). Only a *perfectly dense* consecutive run —
which rank/ID/ordinal columns produce, but a chapter-per-page TOC can
too — falls back to a title signal: flatten when the first-column
entries average multi-word headings, keep as a table when they are the
short single-word labels typical of leaderboards/ID lists.

Recovers the ~0.001 the range-only rule cost (NID back to 0.888) while
still rejecting dense ordinal data tables.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 21:09:45 -07:00
Abimael MartellandClaude Fable 5 673fbe998f docs(registries): add Features and benchmark to crates.io/PyPI/npm pages (#155)
Concise Features list and the opendataloader-bench comparison table on
each registry readme, adapted per ecosystem. Bump all three versions
(crate 0.1.6, python 0.2.5, npm 1.11.1) to republish the pages.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 20:32:29 -07:00
Abimael MartellandClaude Fable 5 c80bedf4bd feat(npm): split platform binaries into optionalDependencies (1.11.0) (#154)
* feat(npm): split platform binaries into optionalDependencies (1.11.0)

The single package bundled all three .node binaries (17.6 MB unpacked)
so every install downloaded every platform. Publish one package per
platform (@firecrawl/pdf-inspector-{linux-x64-gnu,darwin-arm64,
win32-x64-msvc}) holding just its binary; the napi-generated loader
already falls back to exactly these names. Main package drops *.node
from files (8.5 kB tarball) and pins the platform packages as
optionalDependencies, re-stamped to the exact version at publish time.

Publish workflow gains a workflow_dispatch fallback and per-package
already-published checks so partial releases can be retried.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(npm): document Windows support and platform packages; drop stale napi.package.name

Review follow-ups: the README claimed only linux-x64 and macOS ARM64
despite the win32-x64-msvc binary shipping, and napi.package.name
(@firecrawl/pdf-inspector-js) contradicts the real platform package
prefix — the loader and workflow derive it from the root package name.
Verified the generated loader is unchanged without the config.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 20:11:19 -07:00
Abimael MartellandClaude Fable 5 cffe253d1d fix(pypi): slim sdist inherited from crate allowlist; keep type stub (0.2.4) (#153)
The 0.2.3 sdist was 10.09 MiB (packaged tests/fixtures). maturin derives
the sdist file list from Cargo's include allowlist, so it's now 1.35 MiB
— but the allowlist dropped pdf_inspector.pyi, which would strip type
hints from wheels built from the sdist. Add it back and bump to 0.2.4.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:33:39 -07:00
Abimael MartellandClaude Fable 5 2cd1cf1b23 fix(crate): allowlist package contents to fit crates.io size cap (#152)
cargo publish of 0.1.5 failed with 413: the crate packaged everything
(260 files, 10.1MiB compressed) and tests/fixtures alone is 10.2MB.
Add an explicit include list (src, external/bcmaps which tounicode.rs
loads at runtime, readme, license) — 1.3MiB compressed.

Also add a workflow_dispatch fallback to publish-crate.yml so a failed
publish can be retried without a version bump (0.1.5 is already on
main, so a re-push won't register as a version change).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:24:40 -07:00
Abimael MartellandClaude Fable 5 2c97f4979e docs(crate): Rust-specific readme for crates.io (0.1.5) (#151)
crates.io showed the repo README, which leads with Python/Node quick
starts and repo-relative links. Point the crate readme at
docs/rust-api.md, refreshed with an intro, crates.io install, and CLI
install instructions. Bump to 0.1.5 to republish.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:15:42 -07:00
Abimael MartellandClaude Fable 5 60cb953284 docs: add PyPI version badge to README (#150)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:04:00 -07:00
Abimael MartellandClaude Fable 5 0d80efa84a docs(pypi): reformat Types section as stub-style code block (0.2.3) (#149)
The bold-label + comma-list paragraphs render as cramped walls of
inline code on PyPI. A python code block mirroring pdf_inspector.pyi
renders cleanly everywhere and adds field types plus the missing
is_underline/is_strikeout TextItem fields.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 18:54:39 -07:00
Abimael MartellandClaude Fable 5 f3e3129a9a feat(pypi): add package readme and project URLs (0.2.2) (#148)
PyPI showed an empty description because pyproject.toml declared no
readme. Point it at docs/python.md (refreshed with pip install now that
wheels exist) and add sidebar URLs. Bump to 0.2.2 to republish.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 18:36:45 -07:00
Abimael MartellandClaude Fable 5 d7e697fc5e chore(napi): bump @firecrawl/pdf-inspector to 1.10.4 (#147)
Ships #145: exclusive item->region assignment in extract_text_in_regions
(overlapping layout regions no longer double-extract shared items —
duplicated lines on 21% of a 2,078-doc bench corpus, with occasional
content loss when downstream dedup kept the wrong variant).


Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 18:26:32 -07:00
Abimael MartellandClaude Fable 5 bcaadd52fd fix(regions): exclusive item→region assignment in extract_text_in_regions (#145)
* fix(regions): exclusive item->region assignment in extract_text_in_regions

Overlapping layout regions used to extract shared items into EVERY
region they touched (the 1.5pt inclusion margin makes borders generous),
duplicating whole lines in the final markdown on 21% of a 2,078-doc
bench corpus — and downstream duplicate-handling sometimes dropped the
variant holding a sentence tail, turning duplication into content loss.

Each item is now pre-assigned to the single region with the largest
overlap area (same margin as the boolean test); the per-region filter
uses the assignment. Items are partitioned, never suppressed, so no
content can vanish that was previously extracted.

Paired with fire-pdf assembly fixes (neighbor-local sweep dedup +
remainder salvage); verified together on the repro doc: duplicate lines
6 -> 0, the audit's lost sentence recovered (fuzz 78 -> 87.5). Batch
over the worst duplication docs: 185 -> 59 total, 6 of 8 docs to zero.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(regions): review round — single-pass bucketing, no-OCR for lost-to-neighbor empties, shared margin constant

- Assignment and materialization now happen in ONE pass over items
  (clone bucketed at argmax time) instead of a second O(items x regions)
  traversal.
- A region whose only overlapping items were assigned to a
  better-overlapping neighbor no longer flags needs_ocr: the pixels it
  would re-read belong to that neighbor, and OCR would reintroduce the
  duplication exclusivity removed. Matches the pre-change OCR load
  (these regions were non-empty native before).
- REGION_MARGIN hoisted to a module const shared by the boolean
  predicates and the area score — they must stay in sync or an item
  passing the guard could score zero area.

Repro re-verified after fixes: fuzz 87.5, duplicates 0; 756 tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(regions): lost-to-neighbor requires zero items assigned to the region

had_candidates records overlap, not assignment loss: a region whose own
assigned items materialize to empty text (whitespace-only items,
collector filtering) was indistinguishable from one that lost everything
to a neighbor, and wrongly skipped its OCR fallback. The suppression now
also requires assigned_count == 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 18:22:30 -07:00
Abimael MartellandClaude Fable 5 6f75873807 fix(ci): move x86_64 macOS wheel build to macos-15-intel (#146)
macos-13 runners were retired by GitHub, so the x86_64-apple-darwin
build job queued forever and the publish never ran.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 18:22:06 -07:00
Abimael MartellandClaude Fable 5 3ed30d01e1 ci: add PyPI trusted publishing (abi3 wheels, v0.2.1) (#123)
* ci: add PyPI trusted publishing, abi3 wheels, bump to 0.2.1

Adds publish-pypi.yml mirroring the npm/crates.io pattern: triggers on
Cargo.toml version change, builds wheels for 5 platforms via maturin,
publishes with OIDC trusted publishing (no tokens). workflow_dispatch
serves as a manual fallback for the first run after the PyPI project
transfer.

Enables pyo3 abi3-py38 so one wheel per platform covers CPython >=3.8
(previous manual uploads were cp312-only). Bumps version to 0.2.1 since
PyPI already has 0.2.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(pypi): guard dispatch to main, support partial-release repair

Review feedback: trusted publishing doesn't match on branch, so
workflow_dispatch needed an explicit main-ref guard. Manual dispatch now
always rebuilds and publishes with skip-existing so a release that
failed after uploading only some wheels can be completed by re-running.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(pypi): version PyPI package from pyproject.toml, not Cargo.toml

Decouple the Python package version from the crate version, matching
how npm publishing keys off napi/package.json: bump [project] version
in pyproject.toml manually and CI publishes on merge. Reverts the
Cargo.toml bump so this PR no longer triggers a crates.io release.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: remove accidentally committed uv.lock

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): tolerate missing version key in parent pyproject.toml

The first merge of this workflow has a parent commit where pyproject.toml
still used dynamic = ["version"], so the old-version read would KeyError
and the auto-publish would never fire. Treat a missing key as a change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 17:52:13 -07:00
Abimael Martell eac8af0df8 chore(napi): bump @firecrawl/pdf-inspector to 1.10.3 (#144) 2026-07-13 08:58:14 -07:00
Abimael MartellandClaude Fable 5 39c31a8404 fix(underline): rescue snug-owned underlines from the table-ruling filters (#143)
* fix(underline): rescue snug-owned underlines from the table-ruling filters

Documents that underline many full-width lines (dense CJK business docs,
legal redlines, 10-K section links) produce span-similar rules at 3+
y-levels — exactly what the repeated-ruling filter treats as table
rulings, so every semantic underline on such pages was discarded. Three
changes fix detection without re-marking real tables:

1. Snug-owner rescue: a rule survives the repeated-ruling filter when
   the union of touching text runs on its baseline row owns it (rule
   contained within the union's span +0.75em, runs cover >=60% of the
   rule, no column-sized gaps between runs). Table row separators fail
   ownership: they overshoot their cells' text or match gapped items.
   Same-row segmented rules (column-header separators) always stay
   discarded, and a rule enclosed by a drawn cell-sized box (rect-grid
   tables) is never rescued.

2. Vertical window widened 0.35em -> 0.72em below the baseline: CJK
   layouts draw underlines under the full em box, measured at ~0.67em.

3. Prose-table guard in the positions-path suppressor: a detected
   'table' whose cells hold flowing prose (>=30% of cells over 100
   chars) is a detection artifact of boxed callouts + stacked rules,
   not a real table — suppressing there erased every underline on the
   page.

Also fixes cluster_x_positions fabricating phantom table columns from
style-split continuation runs (touching items, gap <2pt, now feed one
column start) — the fix that keeps rect-grid table shapes stable while
underlined links inside cells are correctly marked.

Snapshot updates are underline gains on regulation/form fixtures and one
empty spacer-column change in a subscripted header.

Corpus (508-doc public bench sweep): text output byte-identical on all
docs; underlined items +224/-0; strikeout now fires on redline docs.
Item-level GT coverage: is_underline 86->151/405, is_strikeout 0->10/44.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(underline): fraction-bar guard + subscript merge across underline marks

Found by a 202-doc real-world corpus diff (pdf-evals) that exercises the
full markdown pipeline, which the bench-corpus item sweep does not:

1. Math fraction bars and lattice grid lines are underline geometry —
   short horizontal rules under digits. Guard: a narrow rule (<=60pt)
   with bar-sized text hanging just below it (denominator) never marks.
   The below-text width bound matters: tightly-leaded REAL underlines
   have a full-width next line below, which must not trip the guard.

2. merge_subscript_items refused to merge when the parent was underlined
   but the tiny digit was not (the drawn rule easily misses the digit's
   own overlap window) — losing the merge broke subscript tokens inside
   table cells (b+2 no longer became b₂). Strikeout boundaries still
   block the merge in both directions; only parent-underlined/digit-bare
   merges, absorbing with the parent's flags.

Corpus after refinement: underlined items +220/-2 (the 2 are fraction
bars the old code wrongly marked), GT rule-text coverage 149/405
underline + 10/44 strikeout, text output identical on all 508 bench
docs and word-count-identical on the 202 pdf-evals docs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(underline): address review — grid-evidence veto, strikeout-safe fraction guard, bounded gaps

- Cell-box veto now requires GRID EVIDENCE (a vertically abutting
  neighbor rect with x-overlap) instead of a height window: multiline
  table cells taller than the old 90pt ceiling veto again, and isolated
  filled callout panels (which legitimately contain underlines) no
  longer veto at all.
- The fraction guard gates only UNDERLINE marking; rule_strikes_item
  still evaluates, so short strikeouts near lower text survive.
- Fraction hug distance tightened to 0.3em so a short last-line at
  normal leading is not mistaken for a denominator.
- Continuation-run suppression bounds the negative gap (-4pt): text
  overhanging from an adjacent cell keeps its own column start.

Corpus after review fixes: underlined items +222/-2, GT coverage
150/405 underline + 10/44 strikeout, bench text output identical on
all 508 docs, pdf-evals word loss bounded at equation-reflow noise.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* chore: appease clippy (redundant closure)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 08:47:45 -07:00
Abimael MartellandClaude Fable 5 cb3906e7b8 chore(napi): bump @firecrawl/pdf-inspector to 1.10.2 (#141)
Releases since 1.10.1: lowercase-fragment heading gate (#140,
opendataloader MHS spurious-heading class), per-page OCR routing
reasons (#139), --password support for encrypted PDFs (#138), docs
refresh (#137).


Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 13:30:49 -07:00
Abimael MartellandClaude Fable 5 ebfd096b78 fix(markdown): lowercase-initial short fragments are not headings (#140)
A lowercase-initial one-or-two-word "heading" is a mid-sentence
fragment beside display math ("or inversely", "and therefore") — real
headings that short start uppercase. Measured as spurious headings on
academic docs (opendataloader MHS via fire-pdf's coverage_fallback
path, ENG-5029).

Extends is_heading_fragment, so both bold-heading call sites get the
gate. Corpus sweep (708 opendataloader + ParseBench PDFs): 33 docs
change, all lowercase-fragment demotions from `##`/`#` to plain or
bold ("## caldera" -> "**caldera**", "## of quorum.\"" ->
"**of quorum.\"**") — no real heading is lowercase-initial and that
short in either corpus.


Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 13:24:01 -07:00
Abimael MartellandClaude Fable 5 d8eb33e390 feat: per-page OCR routing reasons (#139)
* feat: per-page OCR routing reasons (scanned/no_text/vector_text/garbled)

Replaces the single suspected_garbled_text signal with a per-page
explanation for why each OCR-flagged page needs OCR. The detector
classifies each page in pages_needing_ocr from its content analysis:

- scanned            — no usable text, image-backed page
- no_text            — no text and no image (blank/unreachable)
- vector_text        — text drawn as vector outlines, not extractable
- suspected_garbled_text — undecodable Identity-H/Type3 fonts

Exposed on PdfTypeResult.ocr_reasons_by_page and surfaced through
PdfProcessResult and the detect-pdf CLI (JSON + human output). Reasons
only ever explain pages already flagged for OCR — a text page with an
embedded logo stays TextBased, so this doesn't widen the OCR net.
Markdown output is byte-identical across the regression corpus.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: cache freshly-analyzed pages so OCR reasons aren't lost

Under a sampling ScanStrategy, the Mixed per-page loop (Phase 2) and the
garbled-font check (Phase 3) analyze non-sampled pages but dropped the
PageAnalysis after flagging them. The reason-classification pass then
missed the cache and defaulted those pages to "scanned", masking the
real vector_text / suspected_garbled_text cause. Insert the fresh
analyses into analysis_cache so the reason pass classifies them
correctly. No change under the default full-sampling strategy (all
pages are already cached).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 13:23:37 -07:00
Abimael Martell a38efcf142 feat: --password support for encrypted PDFs (#138) 2026-07-11 11:55:36 -07:00
Abimael MartellandClaude Fable 5 42f0e52987 feat(site): GitHub Pages landing page (#135)
* feat(site): add GitHub Pages landing page

Self-contained landing page (single index.html, no build step) plus a
Pages deploy workflow that publishes site/ on push to main. Covers the
pitch, install commands for all three registries, feature grid,
benchmark, and tabbed quick-start for Rust/Python/Node/CLI.

Benchmark numbers mirror the README's current published table; both
should be refreshed together in a follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(site): add Firecrawl Parse upsell to closing section

Replace the single CTA with a two-path split: run pdf-inspector locally
(OSS) vs. hand scanned/OCR/at-scale documents to Firecrawl Parse
(hosted). Links the OSS library back to the paid product for the cases
local parsing can't cover.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(site): add Firecrawl branding (flame mark + wordmark)

Firecrawl flame mark anchors the hosted-parse card; charcoal wordmark
in the footer credit. Brand SVGs referenced as-is (exact colors).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(site): refresh benchmark with current main numbers

pdf-inspector row updated from a fresh run on latest main: overall
0.78→0.83, tables 0.59→0.66, headings 0.57→0.74. Now within 0.01 of
opendataloader overall, best tables of the group, headings on par.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 01:03:45 -07:00
Abimael MartellandClaude Fable 5 0ee3cd7d10 docs: refresh benchmark table with current numbers (#137)
Re-ran opendataloader-bench on current main. pdf-inspector improved
across the board since the last table: overall 0.78→0.83, tables
0.59→0.66, headings 0.57→0.74 (competitor rows unchanged). Updated the
prose — heading detection no longer lags opendataloader, and overall is
now within 0.01 of it at ~2.5× the speed.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 00:58:43 -07:00
Abimael Martell f5be40143c chore(napi): bump @firecrawl/pdf-inspector to 1.10.1 (#136) 2026-07-11 00:42:27 -07:00
Abimael MartellandClaude Fable 5 fed3b90d37 feat(extractor): run-local space floor for tracked (letter-spaced) glyph runs (#133)
* feat(extractor): run-local space floor for tracked (letter-spaced) glyph runs

Display type set with tracking renders one glyph per show op; the merge
loop's fixed space thresholds (0.08-0.13 em) then read every letter gap
as a word boundary and emit "H O W" / "F U R T H E R" instead of
"HOW" / "FURTHER". The page-level Canva fixer can't help: it requires
>=50% of the page's items to be letter-spaced, and these docs track
only their display headings.

merge_text_items now pre-scans each run of consecutive single-glyph
items (same size band, same style, mergeable gaps — the loop's own
break conditions) and, when the run is tracked, derives the space floor
from the run's own gap distribution:

- runs with >=4 gaps qualify when the median gap clears the fixed
  threshold; word gaps, if present, form a second mode — split at the
  largest relative jump (>=1.4x), else the run is a single word
  ("I T I S I M P O R T A N T" -> "IT IS IMPORTANT")
- short runs (2-3 gaps: "H O W") additionally demand uniform gaps and
  ALL-CAPS or CJK — a genuine spaced sequence of single letters
  ("x y z" variables) has the same gap count, and display tracking is
  a caps convention; CJK never wants inter-glyph spaces

Corpus sweep (708 opendataloader + ParseBench text PDFs) vs main: 9
docs change — the tracked display titles ("HOW CAN YOU HELP?",
"LUNCHTIME MENU", a tracked email address), and CJK glyph-per-item
docs whose spurious inter-glyph spaces now collapse (GT for those docs
is unspaced CJK; should_join_items already treats no-space CJK as
correct on its path). No other doc moves.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): convention gate on both tiers; Han/Kana floor always infinite (PR #133 review)

- Long lowercase spaced-single runs ("a b c d e") had the tracked gap
  shape in the >=4-gap tier with no convention guard — word boundaries
  lost. The caps/CJK/title-case gate now applies to BOTH tiers; a
  title-case single word ("B u f f a l o") also qualifies.
- Han/Kana runs skipped straight to the bimodal split, so a nonuniform
  gap distribution (justification, punctuation spacing) could
  manufacture a word boundary. Han/Kana now always floors at infinity;
  Hangul deliberately keeps word-boundary handling — Korean spaces
  between words (is_spaceless_cjk excludes it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(extractor): preserve mixed-case glyph boundaries

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 00:28:22 -07:00
Abimael MartellandClaude Fable 5 dd8aba9e20 fix(markdown): keep isolated headings on sparse pages; gate tagged roles (#132)
* fix(markdown): keep isolated headings on sparse pages; gate tagged roles

The isolated-line density guard wiped every isolated line on a page
where they exceeded 25% of lines. On sparse pages (covers, ToC pages
with a lone "CONTENTS" title, section-divider pages) a single heading
is trivially >25%, so the guard erased exactly the line it exists to
find. Require the page to have >=10 lines before the guard runs — the
25% ratio only signals a multi-column misfire on a dense page.

That let more isolated lines through, exposing that the visual heading
heuristic could promote lines already tagged with a non-heading struct
role (list item, blockquote, code, caption, ToC) or set in a monospace
font. Gate the heuristic on those in both converter paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: extend non-heading role gate; centralize on StructRole method

Move the non-heading-role check to StructRole::is_non_heading_content
and extend it to the content roles the inline allowlist missed: Quote,
Index, Note, Reference, BibEntry, Formula, Form (in addition to the
existing list/quote/caption/toc/code roles).

Figure is deliberately excluded: cover and banner pages routinely tag
the document title inside a Figure next to a seal/logo, and that title
is a real heading — including Figure demoted the LA County protocol
cover title from headings to bold. Verified against the reference.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: block table roles from heading promotion too

Add Table/TR/TH/TD/THead/TBody/TFoot to is_non_heading_content. When
table reconstruction falls back and cells reach the line loop as plain
text, a short isolated cell (a TH column header in particular) could be
promoted to a heading. Defensive: no change across either regression
corpus, so pure hardening for the fallback path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 22:25:42 -07:00
Abimael MartellandClaude Fable 5 2dcba76d6f refactor(lib): extract text-quality detectors into a dedicated module (#121)
The garbage/encoding detectors had accreted as ~490 lines of free
functions scattered through lib.rs (flagged in the #120 review). Move the
whole cluster into src/text_quality.rs behind a module doc that maps the
surface: the two interfaces (markdown-level vs item/span-level) and the
detection classes (replacement runs, private-use/C1 runs, dollar-as-space,
non-alphanumeric dominance, substitution-cipher statistics).

Moved verbatim: detect_encoding_issues, is_garbage_text, is_cid_garbage,
analyze_text_quality, region_items_have_decoding_issue and their helpers,
CipherGarbleStats, and the TextQuality* types. The OCR-reason aggregation
plumbing (add_ocr_reason, merge_ocr_reasons, page_ocr_reason*) stays in
lib.rs since it is shared by the main extraction loops, not detection.

Pure code motion — function bodies are unchanged; only visibility keywords
were added (pub(crate) on the six items lib.rs consumes; add_ocr_reason is
now pub(crate) so the module can call it). Behavior is provably unchanged:
same test counts (565 unit + 139 integration), and release output is
byte-identical to merged main across all 185 eval PDFs. Detector unit tests
stay in lib.rs for now because they share test helpers with the table and
layout tests there.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 22:13:08 -07:00
Abimael MartellandClaude Fable 5 3a389f079e fix(markdown): ToC-page suppression, wrapped bold headings, math fragments (#131)
* fix(markdown): ToC-page suppression, wrapped bold headings, math fragments

Three heading-classification improvements:

- After emitting a "Contents"/"Table of Contents" heading, suppress
  heading promotion for the rest of that page: ToC entries are section
  titles that look exactly like headings ("1. Overview of OCR Pack")
  and whole contents pages came out as stacks of ##.
- merge_heading_lines only merged font-size-tier and struct-tree
  headings, so bold-at-body-size headings that wrap emitted two
  separate ## lines. Merge a fully-bold line into the previous
  fully-bold line when it reads as a wrap continuation (starts
  lowercase, tiny Y gap, no terminal punctuation on the previous line).
- Reject display-math fragments from the bold/rarity heading heuristic:
  equations ending in an equation number ("S = kB ln W, (2)") and
  lead-ins referencing one ("Rearranging Equation (8) gives:"). A bare
  trailing colon is deliberately NOT a signal — real headings often end
  with colons ("Procedure:").

The p1244 snapshot change is the bold-merge working as intended:
stacked form labels "**Subtotals** **from pages**" now read
"**Subtotals from pages**".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: guard tier path, protect prev headings in merge, narrow (N) rule

All three review findings applied, calibrated against the corpora:

- is_heading_fragment now gates the font-size-tier path too, not just
  the rarity heuristic.
- The bold wrap-merge requires the previous line to be tier-less as
  suggested; corpus diff confirmed the old behavior was absorbing a
  wrapped list-item fragment into a real heading.
- The bare "(N)" suffix rule suppressed real headings ("Nicaea (325)",
  appendix numbering). It now requires math evidence: an operator
  (=, <=, <<, ...) in the line or ,/: immediately before the number.
  Page-of-total running headers ("PM 2 (10)") get an explicit rule
  since the old blanket suffix check had been catching them only by
  accident.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 21:40:18 -07:00
Abimael MartellandClaude Fable 5 83754ddad0 fix(tables): reject row-stripe grids that swallow body text (#130)
* fix(tables): reject row-stripe grids that swallow body text

Charts (bar graphs, axis gridlines) emit fields of drawing rects that
pass the row-stripe shape test; the resulting phantom table then
captures the page's prose in scrambled reading order — losing headings
and paragraph flow with it. The existing max-cell-length gate only
fires for tables with <4 non-empty rows, which these grids exceed.

Add has_dominant_prose_cell: reject when one cell holds >=60 words AND
at least a third of the table's total words. Real tables never
concentrate that much text in a single cell, even with a description
column.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: extend prose-cell guard to merged-cluster path, add boundary test

The merged-cluster fallback had the same unguarded 4+-row gap as
row-stripe; apply has_dominant_prose_cell there too. The cell-rect path
already runs its own function-word prose check and is left unchanged.

Also add the boundary test from review: a 4+-row data table with one
60-word note cell stays accepted because the 1/3-of-total-words
denominator scales with table size.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: document intended small-table boundary of prose-cell guard

Investigated the suggested row-count exemption: adding a second-prose-
cell requirement (or a non_empty_rows >= 4 exemption) resurrects
verified phantom grids in 7 corpus documents — scrambled body text and
chart/figure regions, one of which is a 4-row grid. Every observed
single-dominant-cell grid in the corpora is swallowed prose, never a
real note table, and rejection degrades gracefully to prose while
acceptance scrambles reading order. Keep the guard unconditional,
document the rationale, and pin the boundary with a test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 20:02:23 -07:00
Abimael MartellandClaude Fable 5 f2f49bdac2 fix(markdown): one-word bold headings, block ToC entries from headings (#129)
Two heading-classification fixes:

- Accept single-word headings ("IMPLEMENTATION", "CONTENTS") when the
  line is all-bold and isolated; the word_count >= 2 gate rejected them
  unconditionally.
- Add is_toc_entry_line: a line ending in a dot-leader group plus page
  number ("Measurement Lab worksheet ... 3") is a table-of-contents
  entry, never a heading. has_dot_leaders misses single-group leaders,
  so entire ToC pages were being promoted to ## headings.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:30:33 -07:00
Abimael MartellandClaude Fable 5 d28594ef94 fix(fonts): recognize URW -Medi suffix as bold (#128)
URW Type 1 fonts abbreviate the Medium weight as "Medi" in the font
name (NimbusRomNo9L-Medi is the Times-Bold substitute embedded by most
LaTeX toolchains; -MediItal is bold italic). is_bold_font only matched
the full word "medium", so bold ran undetected across LaTeX-produced
PDFs — dropping ** emphasis and starving bold-based heading detection
of its signal.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:30:14 -07:00
Abimael MartellandClaude Fable 5 b4d401241e chore(napi): bump @firecrawl/pdf-inspector to 1.10.0 (#127)
New in this release (#125): isStrikeout on TextItem (geometric
detection sharing the underline rules pipeline), descriptor/embedded-
font bold+italic recall for subset fonts (FontDescriptor flags,
ttf-parser OS/2+post, bare-CFF Name INDEX), quote-operator advance
width, Ts text-rise handling, ActualText rise/position fixes, and a
document-scoped font style cache.


Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 17:41:46 -07:00
Abimael MartellandClaude Fable 5 57335f8bcf feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection (#125)
* feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection

Two style-recall gaps, both invisible to the existing name-based
heuristics:

1. Subset fonts with opaque BaseFont names ("Tc1", "AAAAAB+Amplitude")
   defeat is_italic_font/is_bold_font. New descriptor_style_flags reads
   the FontDescriptor (ItalicAngle beyond 4 degrees, Flags bit 7 Italic,
   bit 19 ForceBold) and, when the descriptor claims upright, falls back
   to the embedded font file: ttf-parser's OS/2 fsSelection + post
   italicAngle for sfnt fonts, and the CFF Name INDEX PostScript name
   for bare-CFF FontFile3 (descriptor rewritten to ItalicAngle 0 while
   embedding "Amplitude-LightItalic" was observed in the wild).
   ORed into is_bold/is_italic at item creation (content streams and
   form XObjects).

2. No strikeout signal existed. New is_strikeout on TextItem, detected
   in the same pass as underline: same rules pipeline (stroked lines /
   thin filled rects, table-ruling suppression), different vertical
   window — a rule crossing the glyphs at 12-55% of the em above the
   baseline instead of sitting at it. Exposed through napi and python
   bindings and pdf2md --items-json.

Verified on public ParseBench corpus docs: previously-missed italic
council titles and bold CJK itinerary headings now flagged (render-
checked); 24/508 docs gain flags, none lose any; 35 strikeout items
detected corpus-wide, disjoint from underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): quote-op advance width, Ts text rise, doc-level font style cache (PR #125 review)

Address three valid findings from review:

- The ' (move-to-next-line-and-show-text) operator emitted zero-width
  items and never advanced the text matrix, so geometric underline/
  strikeout detection (which requires width > 0) could never mark its
  text, and following show ops overlapped it. Reuse Tj's advance-width
  computation and matrix advance.

- Ts (text rise) was dropped entirely: raised/lowered runs kept the
  unshifted baseline, so rules drawn at the risen glyph position missed
  the strike/underline windows. Track rise in the text state (saved and
  restored with q/Q) and shift the rendering position through the text
  matrix's y column; advances stay on the unshifted matrix per spec.

- descriptor_style_flags re-decompressed and re-parsed the same embedded
  font program on every page whenever the descriptor left a style flag
  unset (the common case). Add a document-scoped FontStyleCache keyed by
  the FontFile2/FontFile3 object id, threaded through page and form
  extraction alongside the existing CMapDecisionCache.

The fourth finding (Form XObject rules never reach geometric detection)
is real but pre-existing for underline and needs the form walker to grow
path/paint tracking plus a new return type; deferred as a follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): ActualText items render at their glyphs' text rise (PR #125 review)

The EMC-built ActualText item used the captured text matrix without the
rise adjustment the ordinary Tj/TJ/' emission sites apply, so a tagged
run shown with Ts landed on the unshifted baseline — off the strikeout/
underline windows and inconsistent with untagged runs. The rise is
captured together with the first-glyph matrix (and at BDC for the
entry-position fallback): the item must render at the rise of its
GLYPHS, not whatever rise is set by EMC time.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): capture ActualText glyph position after the quote op's line move (PR #125 review)

The `'` handler skipped the entire suppressed-extraction block, so a
tagged span whose show op is `'` never captured its glyph matrix/rise —
the EMC item fell back to the BDC-entry matrix, which sits on the
PREVIOUS line (the `'` line move happens after BDC) with no rise. The
capture now happens right after the line move, matching the Tj/TJ
paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): style-boundary gate on subscript merge + strikeout suppression coverage (PR #125 review)

merge_subscript_items absorbed a script digit into its parent
regardless of underline/strikeout flags — dropping the digit's own mark
or widening the parent's over it. The merged item carries one flag, so
differing marks now break the merge, mirroring merge_text_items'
style-boundary rule (pre-existing for underline as well).

Also extends the table-suppression test to assert is_strikeout is
cleared alongside is_underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 17:30:47 -07:00
Abimael MartellandClaude Fable 5 15bc7894a4 docs: add MIT LICENSE file and license badge (#124)
Cargo.toml and pyproject.toml already declare MIT but the repo had no
LICENSE file, so GitHub and package registries couldn't display it.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 14:25:08 -07:00
Abimael MartellandClaude Fable 5 6e5e5849c8 Detect substitution-cipher garbled text from broken ToUnicode CMaps (#120)
* fix(lib): detect substitution-cipher garbled text from broken ToUnicode CMaps

ParseBench text_simple__att10k.pdf (issue #118) ships Type0/Identity-H
fonts whose ToUnicode CMaps are authored garbled: every bfrange maps with
a wrong constant delta, so text extracts as pure-ASCII ciphertext
("Certificate" -> "8VceZWZTReV"). The embedded subset font has no cmap
table and no glyph names, so no decode source can recover the real text
(poppler and mupdf emit the same ciphertext). The only correct behavior
is to flag the page for OCR instead of serving the garbage silently --
but the text is 100% printable ASCII with word-like tokens, so it slipped
past is_garbage_text and detect_encoding_issues.

Add CipherGarbleStats, a letter-statistics discriminator that flags a
Latin-dominant sample (>=200 ASCII letters) when vowels are starved
(<=30% of letters) AND either:
- lowercase->uppercase transitions inside words exceed 10% of letter
  bigrams (a shifted lowercase alphabet straddles the ASCII uppercase
  block), or
- the letter histogram's cosine similarity against English letter
  frequencies drops below 0.60 (catches shifts that stay within case
  blocks).

Wired into analyze_text_quality (per-page, item-level) and
detect_encoding_issues (markdown-level), so extract_pages_markdown
reports needs_ocr + suspected_garbled_text and suppresses the garbage.

Thresholds validated against the 380-document pdf-evals snapshot corpus
(Swedish, Finnish, Turkish, German, romaji, schematics, all-caps and
camelCase-heavy docs): zero false positives, and byte-identical eval
output vs main. Garbled page measures vowel ratio 0.245 / case-shift
rate 0.225 / cosine 0.532; closest legitimate document on each axis is
0.264 / 0.021 / 0.801.

Fixes #118

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump pdf-inspector to 0.1.4, npm package to 1.9.11

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(lib): exempt uniform-case structured content from cipher detection

Address PR review (cubic P2): the frequency branch (english_cosine < 0.60)
fired on any Latin-dominant, low-vowel letter distribution unlike English,
so non-linguistic ASCII — DNA/protein sequences, ticker symbols, hex dumps —
could be suppressed and routed to OCR despite not being garbled. Measured:
DNA cosine 0.428 / vowel ratio 0.260, protein 0.738, tickers 0.747, hex
0.549 — all would have flagged.

Add a mixed-case guard to looks_garbled: garbled English is a permutation of
natural language and carries sentence capitalization (block-straddling shifts
invert the ratio — att10k is 60% uppercase; in-case Caesar shifts preserve it
at ~3%), so both keep some of each case. The exempted structured content is
uniform case (all upper or all lower). Requiring the minority case to be >=1%
of ASCII letters exempts single-case sequences while preserving both garble
signals, including the in-case-shift scenario the frequency branch exists for.

Strictly tightens the detector: it can only remove flags, so the eval corpus
stays at zero false positives (verified byte-identical to a baseline main
binary across all 185 PDFs) and att10k remains flagged. Adds regression tests
for DNA, protein, tickers, and an in-case Caesar shift.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(lib): make cipher detection case-agnostic via sorted-histogram shape

Address PR review follow-up: the mixed-case guard from the previous commit
returned before the vowel/frequency checks, creating a blind spot — a
uniform-case (all-lower or all-upper) substitution cipher is a plausible
broken-CMap output and would bypass OCR entirely.

Replace the case proxy with the actual invariant. A substitution cipher is
a bijection over a real language's alphabet, so it preserves the frequency
SHAPE (the sorted histogram) while scrambling letter POSITIONS (the unsorted
histogram). Signal 2 now flags when english_cosine < 0.60 (positions unlike
English) AND english_shape_cosine >= 0.90 (profile is still English-shaped).
This is independent of case, so it catches all-lower, all-upper, and
case-straddling shifts alike.

The exempted structured content fails one half: DNA/hex dumps have too steep
a profile (shape cosine 0.74 / 0.81 < 0.90), while protein sequences, ticker
symbols and base64 are not sufficiently unlike English in position (unsorted
cosine 0.74 / 0.75 / 0.77 >= 0.60). All stay out of OCR.

Still strictly corpus-safe: every real Latin document scores unsorted cosine
>= 0.70 (min 0.80), far above the 0.60 gate, so none can reach Signal 2.
Re-verified byte-identical to a baseline main binary across all 185 eval
PDFs; att10k remains flagged. Drops the now-unused case counters and adds
all-lowercase / all-uppercase shifted-prose regression tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: source Python package version from Cargo.toml via maturin

Address PR review (cubic P2): pyproject.toml pinned version = "0.1.0",
which overrides Cargo.toml, so a maturin build produced a 0.1.0 Python
artifact regardless of the crate version (it had drifted since the PyO3
bindings were added). Switch to dynamic = ["version"] so maturin sources
the version from Cargo.toml [package] version and the two can no longer
diverge. No workflow auto-publishes the Python package, so this is metadata
hygiene rather than a release-path fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 12:44:32 -07:00
Abimael MartellandCursor b375d6f102 feat(markdown): underline emission, Unicode scripts, style-preserving merges (#117)
* feat(markdown): underline emission, Unicode scripts, style-preserving merges (ENG-5015 2b)

Three formatting losses in the direct-extraction markdown path:

1. text_with_formatting gains <u> run emission (detect_underline option,
   default on) using the geometric is_underline flag from 1.9.9.
   Underline runs stay free of nested bold/italic markers — consumers
   match tag content literally. Heading lines keep plain text for
   bold/italic but preserve <u>: the tag carries meaning `#` doesn't.
2. merge_subscript_items now maps absorbed digit scripts to Unicode
   sub/superscript forms with direction from the baseline offset
   ("H"+"2" -> "H₂", "word"+raised "2" -> "word²", "m"+"3" -> "m³").
   NFKC/NFKD folds these back to plain digits so text matching
   downstream is unaffected; renderers keep the script semantics.
3. merge_text_items no longer merges across bold/italic boundaries —
   absorbing a styled run into a plain neighbor erased the styling
   before markdown emission ever saw it. On eval docs this recovers
   20-82 italic runs per document that previously emitted as plain.

Snapshots regenerated (diffs are the features: CCl₂F₂, m³, underlined
legal section headings, finer bold runs). pdf-evals regression suite:
202/202 real PDFs pass. napi 1.9.9 -> 1.9.10.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): break merges at underline boundaries too (review)

OR-merging underline stretched the eventual <u> span over neighboring
plain fragments. Merge runs now break on any style-flag change, the
redundant accumulator is gone, and format_list_item learned to move
bullet markers outside <u> wrappers so fully-underlined bullet lines
still render as markdown lists. td9264 snapshot regenerated — spans are
tighter (trailing periods correctly outside the tag).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(markdown): strip stray spaces before sentence punctuation (review)

Style-boundary item splits can strand a trailing period in its own
fragment, and multiple assembly paths join fragments with spaces,
yielding "word ." artifacts. Rather than chasing every join site, a
postprocess pass removes a space before `.`/`,`/`;` when the mark ends
its token (whitespace, cell boundary `|`, or end of text follows).
Dot leaders/ellipses and mid-token periods are untouched.

Fixes the td9264 "companies ." artifacts and two pre-existing
"armoring ," artifacts in the 2013-app2 snapshot. pdf-evals: zero
markdown diffs across all 203 corpus PDFs vs committed baselines.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(tables): trim spaces inside parenthetical cell fragments

* fix(tables): reject sparse prose row-stripe tables

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 23:40:08 -07:00
Abimael MartellandCursor 422a2ff118 feat(extractor): geometric underline detection on TextItem (#116)
* feat(extractor): geometric underline detection on TextItem (ENG-5015)

PDFs carry no underline font flag — underlines are stroked horizontal
lines or thin filled rects drawn under the baseline. Correlate those
graphics (already parsed from the content stream) with text items in a
post-pass: a rule within ~0.35em below the baseline covering >=60% of
an item's width marks is_underline.

Exposed through the napi and python bindings. Verified on real docs:
4/4 underlined sentences flagged on a Japanese report, links/headings
flagged on 8 of 10 underline-bearing eval docs, zero flags on docs
without underlines. Known FP source (table cell borders) documented —
downstream applies inline styling only to plain-text regions.

napi 1.9.8 -> 1.9.9.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): underline rules only from painted rects, normalized extents (review)

Two review fixes: (1) normalize rect extents before the thickness/width
checks — `re` operands pass through the CTM so width/height can be
negative, which missed negative-width rules and let negative-height
bands pass as thin; (2) only feed painted rects to underline detection —
`re` rects now wait in a pending list until a paint operator (S/s, f/F/
f*, B/B*/b/b*) confirms them, and `re W n` clip-only paths are discarded
at `n`, so invisible clip boundaries no longer underline nearby text.
Marking moved into content_stream where paint state lives (pre-rotation,
consistent device space).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): harden underline detection

* feat(cli): export positioned text item json

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 10:20:18 -07:00
Abimael Martell 30eddade77 fix(extractor): preserve tagged overlapping text order (#114) 2026-06-24 10:26:52 -07:00
Abimael Martell 1b2e2c76d6 fix(extractor): make trace previews unicode-safe (#113) 2026-06-24 01:27:27 -06:00
Abimael Martell ce49794719 fix(extractor): reduce garbled OCR false positives (#112)
* fix(extractor): reduce garbled OCR false positives

* fix(extractor): tighten garbled text OCR routing

* fix(extractor): decode UTF-16 ToUnicode destinations

* fix(extractor): narrow ToUnicode destination cleanup

* fix(extractor): decode Aptos private ff ligature
2026-06-24 01:13:16 -06:00
Abimael Martell 1a5ba6f1e9 feat(api): expose OCR reason signal (#110) 2026-06-23 15:51:55 -07:00
Abimael Martell 57b98c6a5d fix(extractor): flag garbled text spans for OCR (#108)
* fix(extractor): flag garbled text spans for OCR

* fix(extractor): apply text quality checks to regions

* chore(napi): bump npm package version
2026-06-23 13:12:10 -07:00
Abimael Martell f25808e0a7 fix(extractor): restore CID font state for Chinese text (#106)
* fix Chinese CID text decoding

* bump package versions
2026-06-20 19:43:27 -06:00
Abimael Martell 9360c8464d ci: add crates trusted publishing (#103) 2026-06-05 11:21:55 -07:00
Abimael Martell 252d87ac58 docs(readme): add package badges and install docs (#102)
* docs: add crates.io install instructions

* docs: add npm badge
2026-06-05 11:01:56 -07:00
Abimael Martell 85890648c9 chore: use crates.io lopdf (#101) 2026-06-05 10:50:05 -07:00
Abimael Martell 6e55e38b55 fix(markdown): handle wrapped bold abstracts (#100)
* fix(markdown): handle wrapped bold abstracts

* chore: bump napi package version
2026-06-01 14:59:18 -07:00
Abimael Martell 42befcea57 fix(pdf-inspector): recover wrapped key-value tables (#99)
* fix(pdf-inspector): recover wrapped key-value tables

* fix(pdf-inspector): satisfy clippy
2026-06-01 10:10:55 -07:00
Abimael Martell e547f616f9 fix(pdf-inspector): recover key-value region tables (#98) 2026-05-29 18:55:38 -07:00
Abimael Martell 455dfe5a74 fix(pdf-inspector): recover borderless region tables (#97) 2026-05-28 10:20:06 -07:00
90 changed files with 30496 additions and 1001 deletions
+61 -11
View File
@@ -14,44 +14,57 @@ jobs:
name: Test
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Install Rust
uses: dtolnay/rust-toolchain@stable
uses: dtolnay/rust-toolchain@4cda84d5c5c54efe2404f9d843567869ab1699d4 # stable
with:
toolchain: stable
- name: Cache cargo
uses: Swatinem/rust-cache@v2
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Run tests
run: cargo test --verbose
- name: Test developer scripts
run: python3 -m unittest discover -s scripts/tests
fmt:
name: Format
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Install Rust
uses: dtolnay/rust-toolchain@stable
uses: dtolnay/rust-toolchain@4cda84d5c5c54efe2404f9d843567869ab1699d4 # stable
with:
toolchain: stable
components: rustfmt
- name: Check formatting
run: cargo fmt --all -- --check
- name: Check WASM formatting
run: cargo fmt --manifest-path wasm/Cargo.toml -- --check
clippy:
name: Clippy
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Install Rust
uses: dtolnay/rust-toolchain@stable
uses: dtolnay/rust-toolchain@4cda84d5c5c54efe2404f9d843567869ab1699d4 # stable
with:
toolchain: stable
components: clippy
- name: Cache cargo
uses: Swatinem/rust-cache@v2
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
with:
key: clippy
@@ -65,15 +78,52 @@ jobs:
matrix:
os: [ubuntu-latest, macos-latest]
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Install Rust
uses: dtolnay/rust-toolchain@stable
uses: dtolnay/rust-toolchain@4cda84d5c5c54efe2404f9d843567869ab1699d4 # stable
with:
toolchain: stable
- name: Cache cargo
uses: Swatinem/rust-cache@v2
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
with:
key: build
- name: Build
run: cargo build --release --verbose
wasm:
name: WebAssembly
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Install Rust
uses: dtolnay/rust-toolchain@4cda84d5c5c54efe2404f9d843567869ab1699d4 # stable
with:
toolchain: stable
targets: wasm32-unknown-unknown
components: clippy
- name: Cache cargo
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
with:
workspaces: |
wasm -> target
key: wasm
- name: Check WebAssembly bindings
run: cargo check --manifest-path wasm/Cargo.toml --target wasm32-unknown-unknown
- name: Check root package for WebAssembly
run: cargo check --target wasm32-unknown-unknown
- name: Lint WebAssembly bindings
run: cargo clippy --manifest-path wasm/Cargo.toml --target wasm32-unknown-unknown -- -D warnings
- name: Install wasm-pack
run: cargo install wasm-pack --version 0.15.0 --locked
- name: Test WebAssembly package
run: wasm-pack test --node --release wasm
+36
View File
@@ -0,0 +1,36 @@
name: Deploy landing page
on:
push:
branches: [main]
paths: ['site/**', '.github/workflows/pages.yml']
workflow_dispatch:
permissions:
contents: read
pages: write
id-token: write
# Allow one concurrent deployment; don't cancel an in-progress production deploy.
concurrency:
group: pages
cancel-in-progress: false
jobs:
deploy:
name: Build & deploy to GitHub Pages
runs-on: ubuntu-latest
environment:
name: github-pages
url: ${{ steps.deploy.outputs.page_url }}
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Upload site artifact
uses: actions/upload-pages-artifact@fc324d3547104276b827a68afc52ff2a11cc49c9 # v5.0.0
with:
path: site
- name: Deploy to GitHub Pages
id: deploy
uses: actions/deploy-pages@cd2ce8fcbc39b97be8ca5fce6e763baed58fa128 # v5.0.0
+108
View File
@@ -0,0 +1,108 @@
name: Publish Rust crate
on:
push:
branches: [main]
paths: ['Cargo.toml']
# Manual fallback: retry a publish that failed after the version was
# already merged (a plain re-push won't register as a version change).
workflow_dispatch:
permissions:
contents: read
env:
CARGO_TERM_COLOR: always
jobs:
check-version:
name: Check version change
# Guard manual dispatches: crates.io trusted publishing matches
# repo+workflow+environment but NOT branch, so without this a
# workflow_dispatch from any branch could publish unmerged code.
if: github.ref == 'refs/heads/main'
runs-on: ubuntu-latest
outputs:
changed: ${{ steps.check.outputs.changed }}
published: ${{ steps.check.outputs.published }}
version: ${{ steps.check.outputs.version }}
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check if version changed
id: check
run: |
NEW_VERSION=$(python3 -c 'import pathlib, tomllib; print(tomllib.loads(pathlib.Path("Cargo.toml").read_text())["package"]["version"])')
echo "version=$NEW_VERSION" >> "$GITHUB_OUTPUT"
if [ "${{ github.event_name }}" = "workflow_dispatch" ]; then
# Manual dispatch publishes the current version regardless of the
# previous commit; the crates.io check below still prevents
# double-publishing an already-released version.
echo "manual dispatch: publishing v$NEW_VERSION"
echo "changed=true" >> "$GITHUB_OUTPUT"
else
OLD_VERSION=$(git show HEAD~1:Cargo.toml | python3 -c 'import sys, tomllib; print(tomllib.loads(sys.stdin.read())["package"]["version"])')
echo "old=$OLD_VERSION new=$NEW_VERSION"
if [ "$NEW_VERSION" = "$OLD_VERSION" ]; then
echo "changed=false" >> "$GITHUB_OUTPUT"
echo "published=false" >> "$GITHUB_OUTPUT"
exit 0
fi
echo "changed=true" >> "$GITHUB_OUTPUT"
fi
HTTP_STATUS=$(curl --silent --show-error --output /tmp/crate-version.json --write-out "%{http_code}" \
-H "User-Agent: firecrawl/pdf-inspector publish workflow (https://github.com/firecrawl/pdf-inspector)" \
"https://crates.io/api/v1/crates/pdf-inspector/$NEW_VERSION")
case "$HTTP_STATUS" in
200)
echo "published=true" >> "$GITHUB_OUTPUT"
echo "pdf-inspector v$NEW_VERSION is already published"
;;
404)
echo "published=false" >> "$GITHUB_OUTPUT"
;;
*)
cat /tmp/crate-version.json
echo "Unexpected crates.io response: $HTTP_STATUS" >&2
exit 1
;;
esac
publish:
name: Publish to crates.io
needs: check-version
if: needs.check-version.outputs.changed == 'true' && needs.check-version.outputs.published == 'false'
runs-on: ubuntu-latest
environment: crates-io
permissions:
contents: read
id-token: write
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Install Rust
uses: dtolnay/rust-toolchain@4cda84d5c5c54efe2404f9d843567869ab1699d4 # stable
with:
toolchain: stable
- name: Verify package
run: cargo publish --dry-run
- name: Authenticate with crates.io
id: auth
uses: rust-lang/crates-io-auth-action@c6f97d42243bad5fab37ca0427f495c86d5b1a18 # v1.0.5
- name: Publish crate
run: cargo publish
env:
CARGO_REGISTRY_TOKEN: ${{ steps.auth.outputs.token }}
+169
View File
@@ -0,0 +1,169 @@
name: Publish Python package
on:
push:
branches: [main]
paths: ['pyproject.toml']
# Manual fallback: re-publish the current version without a version bump
# (e.g. first run after PyPI trusted publishing is configured).
workflow_dispatch:
permissions:
contents: read
jobs:
check-version:
name: Check version change
# Guard manual dispatches too: PyPI trusted publishing matches
# repo+workflow+environment but NOT branch, so without this a
# workflow_dispatch from any branch could publish unmerged code.
if: github.ref == 'refs/heads/main'
runs-on: ubuntu-latest
outputs:
changed: ${{ steps.check.outputs.changed }}
published: ${{ steps.check.outputs.published }}
version: ${{ steps.check.outputs.version }}
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check if version changed
id: check
run: |
NEW_VERSION=$(python3 -c 'import pathlib, tomllib; print(tomllib.loads(pathlib.Path("pyproject.toml").read_text())["project"]["version"])')
echo "version=$NEW_VERSION" >> "$GITHUB_OUTPUT"
if [ "${{ github.event_name }}" = "workflow_dispatch" ]; then
# Manual dispatch always rebuilds and publishes. Combined with
# skip-existing on the publish step, this repairs partial releases
# (PyPI's version endpoint returns 200 even when only some of the
# expected wheels were uploaded).
echo "manual dispatch: publishing v$NEW_VERSION (skip-existing handles uploaded files)"
echo "changed=true" >> "$GITHUB_OUTPUT"
echo "published=false" >> "$GITHUB_OUTPUT"
exit 0
fi
# .get(): the parent commit may predate the static version field
# (pyproject.toml used dynamic = ["version"]) — treat that as a change
# so the very first merge of this workflow publishes.
OLD_VERSION=$(git show HEAD~1:pyproject.toml | python3 -c 'import sys, tomllib; print(tomllib.loads(sys.stdin.read())["project"].get("version", ""))')
echo "old=$OLD_VERSION new=$NEW_VERSION"
if [ "$NEW_VERSION" = "$OLD_VERSION" ]; then
echo "changed=false" >> "$GITHUB_OUTPUT"
echo "published=false" >> "$GITHUB_OUTPUT"
exit 0
fi
echo "changed=true" >> "$GITHUB_OUTPUT"
HTTP_STATUS=$(curl --silent --show-error --output /tmp/pypi-version.json --write-out "%{http_code}" \
"https://pypi.org/pypi/pdf-inspector/$NEW_VERSION/json")
case "$HTTP_STATUS" in
200)
echo "published=true" >> "$GITHUB_OUTPUT"
echo "pdf-inspector v$NEW_VERSION is already published to PyPI"
;;
404)
echo "published=false" >> "$GITHUB_OUTPUT"
;;
*)
cat /tmp/pypi-version.json
echo "Unexpected PyPI response: $HTTP_STATUS" >&2
exit 1
;;
esac
build:
needs: check-version
if: needs.check-version.outputs.changed == 'true' && needs.check-version.outputs.published == 'false'
name: Build ${{ matrix.target }}
runs-on: ${{ matrix.os }}
strategy:
matrix:
include:
- os: ubuntu-latest
target: x86_64-unknown-linux-gnu
- os: ubuntu-latest
target: aarch64-unknown-linux-gnu
# macos-13 was retired by GitHub; macos-15-intel is the remaining
# Intel runner label (available through 2027).
- os: macos-15-intel
target: x86_64-apple-darwin
- os: macos-14
target: aarch64-apple-darwin
- os: windows-latest
target: x86_64-pc-windows-msvc
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: '3.12'
- name: Build wheel
uses: PyO3/maturin-action@e83996d129638aa358a18fbd1dfb82f0b0fb5d3b # v1.51.0
with:
target: ${{ matrix.target }}
args: --release --out dist
manylinux: auto
- name: Upload wheel
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: wheels-${{ matrix.target }}
path: dist/*.whl
if-no-files-found: error
sdist:
needs: check-version
if: needs.check-version.outputs.changed == 'true' && needs.check-version.outputs.published == 'false'
name: Build sdist
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Build sdist
uses: PyO3/maturin-action@e83996d129638aa358a18fbd1dfb82f0b0fb5d3b # v1.51.0
with:
command: sdist
args: --out dist
- name: Upload sdist
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: sdist
path: dist/*.tar.gz
if-no-files-found: error
publish:
name: Publish to PyPI
needs: [check-version, build, sdist]
runs-on: ubuntu-latest
environment: pypi
permissions:
contents: read
id-token: write
steps:
- name: Download all artifacts
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
with:
path: dist
merge-multiple: true
- name: List artifacts
run: ls -la dist/
- name: Publish to PyPI
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
with:
packages-dir: dist
# Tolerate already-uploaded files so a manual re-run can complete
# a release that previously failed partway through.
skip-existing: true
+115
View File
@@ -0,0 +1,115 @@
name: Publish WebAssembly package
on:
push:
branches: [main]
paths: ['wasm/Cargo.toml']
workflow_dispatch:
permissions:
contents: read
id-token: write
env:
CARGO_TERM_COLOR: always
jobs:
check-version:
name: Check version change
if: github.ref == 'refs/heads/main'
runs-on: ubuntu-latest
outputs:
changed: ${{ steps.check.outputs.changed }}
package_exists: ${{ steps.check.outputs.package_exists }}
published: ${{ steps.check.outputs.published }}
version: ${{ steps.check.outputs.version }}
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check package version
id: check
run: |
NEW_VERSION=$(python3 -c 'import pathlib, tomllib; print(tomllib.loads(pathlib.Path("wasm/Cargo.toml").read_text())["package"]["version"])')
echo "version=$NEW_VERSION" >> "$GITHUB_OUTPUT"
if [ "${{ github.event_name }}" = "workflow_dispatch" ]; then
echo "changed=true" >> "$GITHUB_OUTPUT"
elif git cat-file -e HEAD~1:wasm/Cargo.toml 2>/dev/null; then
OLD_VERSION=$(git show HEAD~1:wasm/Cargo.toml | python3 -c 'import sys, tomllib; print(tomllib.loads(sys.stdin.read())["package"]["version"])')
if [ "$NEW_VERSION" = "$OLD_VERSION" ]; then
echo "changed=false" >> "$GITHUB_OUTPUT"
echo "published=false" >> "$GITHUB_OUTPUT"
exit 0
fi
echo "changed=true" >> "$GITHUB_OUTPUT"
else
echo "changed=true" >> "$GITHUB_OUTPUT"
fi
if ! npm view "@firecrawl/pdf-inspector-wasm" name >/dev/null 2>&1; then
echo "package_exists=false" >> "$GITHUB_OUTPUT"
echo "published=false" >> "$GITHUB_OUTPUT"
echo "The initial package must be published once before trusted publishing can be configured."
exit 0
fi
echo "package_exists=true" >> "$GITHUB_OUTPUT"
if npm view "@firecrawl/pdf-inspector-wasm@$NEW_VERSION" version >/dev/null 2>&1; then
echo "published=true" >> "$GITHUB_OUTPUT"
else
echo "published=false" >> "$GITHUB_OUTPUT"
fi
publish:
name: Build and publish
needs: check-version
if: needs.check-version.outputs.changed == 'true' && needs.check-version.outputs.package_exists == 'true' && needs.check-version.outputs.published == 'false'
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Install Rust
uses: dtolnay/rust-toolchain@4cda84d5c5c54efe2404f9d843567869ab1699d4 # stable
with:
toolchain: stable
targets: wasm32-unknown-unknown
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: '24'
registry-url: 'https://registry.npmjs.org'
- name: Install wasm-pack
run: cargo install wasm-pack --version 0.15.0 --locked
- name: Build browser package
run: wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release
- name: Prepare package metadata
run: |
node -e '
const fs = require("fs")
const path = "wasm/pkg/package.json"
const pkg = JSON.parse(fs.readFileSync(path, "utf8"))
pkg.name = "@firecrawl/pdf-inspector-wasm"
pkg.description = "Browser WebAssembly bindings for the pdf-inspector Rust PDF parser"
pkg.keywords = ["pdf", "pdf-parser", "webassembly", "wasm", "markdown", "rust", "firecrawl"]
pkg.repository = { type: "git", url: "https://github.com/firecrawl/pdf-inspector" }
pkg.homepage = "https://github.com/firecrawl/pdf-inspector/tree/main/wasm"
pkg.publishConfig = { access: "public" }
fs.writeFileSync(path, JSON.stringify(pkg, null, 2) + "\n")
'
- name: Inspect package contents
run: npm pack --dry-run ./wasm/pkg
- name: Publish package
run: npm publish ./wasm/pkg --provenance --access public
+184 -19
View File
@@ -4,6 +4,9 @@ on:
push:
branches: [main]
paths: ['napi/package.json']
# Manual fallback: retry a publish that failed partway (per-package
# already-published checks make re-runs idempotent).
workflow_dispatch:
permissions:
contents: read
@@ -12,24 +15,41 @@ permissions:
jobs:
check-version:
name: Check version change
# Guard manual dispatches: npm trusted publishing matches
# repo+workflow+environment but NOT branch, so without this a
# workflow_dispatch from any branch could publish unmerged code.
if: github.ref == 'refs/heads/main'
runs-on: ubuntu-latest
outputs:
changed: ${{ steps.check.outputs.changed }}
version: ${{ steps.check.outputs.version }}
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 2
- name: Check package version sync
run: python3 scripts/version.py --check
- name: Check if version changed
id: check
run: |
NEW_VERSION=$(node -p "require('./napi/package.json').version")
echo "version=$NEW_VERSION" >> "$GITHUB_OUTPUT"
if [ "${{ github.event_name }}" = "workflow_dispatch" ]; then
# Manual dispatch rebuilds and publishes the current version; the
# per-package already-published checks in the publish job skip
# anything that made it out in a previous partial run.
echo "manual dispatch: publishing v$NEW_VERSION"
echo "changed=true" >> "$GITHUB_OUTPUT"
exit 0
fi
OLD_VERSION=$(git show HEAD~1:napi/package.json | node -p "JSON.parse(require('fs').readFileSync('/dev/stdin','utf8')).version")
echo "old=$OLD_VERSION new=$NEW_VERSION"
if [ "$NEW_VERSION" != "$OLD_VERSION" ]; then
echo "changed=true" >> "$GITHUB_OUTPUT"
echo "version=$NEW_VERSION" >> "$GITHUB_OUTPUT"
else
echo "changed=false" >> "$GITHUB_OUTPUT"
fi
@@ -44,31 +64,61 @@ jobs:
include:
- os: ubuntu-latest
target: x86_64-unknown-linux-gnu
# napi-cross builds gnu targets against an old glibc sysroot for
# broad distro compatibility; musl targets cross-compile with
# zig via cargo-zigbuild (napi's -x flag).
- os: ubuntu-latest
target: aarch64-unknown-linux-gnu
build-flags: --use-napi-cross
- os: ubuntu-latest
target: x86_64-unknown-linux-musl
build-flags: -x
- os: ubuntu-latest
target: aarch64-unknown-linux-musl
build-flags: -x
- os: macos-14
target: aarch64-apple-darwin
- os: windows-latest
target: x86_64-pc-windows-msvc
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Install Rust
uses: dtolnay/rust-toolchain@stable
uses: dtolnay/rust-toolchain@4cda84d5c5c54efe2404f9d843567869ab1699d4 # stable
with:
toolchain: stable
targets: ${{ matrix.target }}
- uses: oven-sh/setup-bun@v2
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2.2.0
with:
bun-version: latest
- name: Install zig
if: contains(matrix.target, 'musl')
uses: mlugg/setup-zig@d1434d08867e3ee9daa34448df10607b98908d29 # v2.2.1
with:
version: 0.14.1
- name: Install cargo-zigbuild
if: contains(matrix.target, 'musl')
uses: taiki-e/install-action@67729d5c413db75907f0ad1e39bb04b9c868ff60 # v2.85.7
env:
GITHUB_TOKEN: ${{ github.token }}
with:
tool: cargo-zigbuild
- name: Cache cargo
uses: actions/cache@v4
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
with:
path: |
~/.cargo/registry/index/
~/.cargo/registry/cache/
~/.cargo/git/db/
~/.napi-rs
napi/target/
key: ${{ runner.os }}-cargo-napi-${{ hashFiles('**/Cargo.lock') }}
key: ${{ matrix.target }}-cargo-napi-${{ hashFiles('**/Cargo.lock') }}
restore-keys: |
${{ runner.os }}-cargo-napi-
${{ matrix.target }}-cargo-napi-
- name: Install dependencies
working-directory: napi
@@ -76,10 +126,10 @@ jobs:
- name: Build native addon
working-directory: napi
run: bunx napi build --platform --release
run: bunx napi build --platform --release --target ${{ matrix.target }} ${{ matrix.build-flags }}
- name: Upload native binary
uses: actions/upload-artifact@v4
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: bindings-${{ matrix.target }}
path: napi/*.node
@@ -87,7 +137,7 @@ jobs:
- name: Upload generated JS bindings
if: matrix.target == 'x86_64-unknown-linux-gnu'
uses: actions/upload-artifact@v4
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: js-bindings
path: |
@@ -95,34 +145,149 @@ jobs:
napi/index.d.ts
if-no-files-found: error
smoke-test:
name: Smoke test ${{ matrix.target }}
needs: [check-version, build]
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
include:
- os: ubuntu-latest
target: x86_64-unknown-linux-gnu
- os: ubuntu-latest
target: x86_64-unknown-linux-musl
- os: ubuntu-24.04-arm
target: aarch64-unknown-linux-gnu
- os: ubuntu-24.04-arm
target: aarch64-unknown-linux-musl
- os: macos-14
target: aarch64-apple-darwin
- os: windows-latest
target: x86_64-pc-windows-msvc
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Download native binary
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
with:
name: bindings-${{ matrix.target }}
path: napi
- name: Download generated JS bindings
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
with:
name: js-bindings
path: napi
# musl binaries must load under a real musl libc, so run inside Alpine.
- name: Run smoke test (Alpine)
if: contains(matrix.target, 'musl')
run: docker run --rm -v "$PWD:/repo" -w /repo/napi node:24-alpine node test.mjs
- name: Run smoke test
if: ${{ !contains(matrix.target, 'musl') }}
working-directory: napi
run: node test.mjs
publish:
name: Publish to npm
needs: [check-version, build]
needs: [check-version, build, smoke-test]
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write
steps:
- uses: actions/checkout@v6
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- uses: actions/setup-node@v6
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: '24'
registry-url: 'https://registry.npmjs.org'
- name: Download all artifacts
uses: actions/download-artifact@v4
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
with:
path: napi/artifacts
- name: Collect binaries and publish
- name: Publish platform packages
working-directory: napi
run: |
cp artifacts/bindings-*/*.node .
VERSION="${{ needs.check-version.outputs.version }}"
for node_file in artifacts/bindings-*/pdf-inspector.*.node; do
base=$(basename "$node_file")
suffix=${base#pdf-inspector.}
suffix=${suffix%.node}
pkg="@firecrawl/pdf-inspector-$suffix"
if npm view "$pkg@$VERSION" version >/dev/null 2>&1; then
echo "$pkg@$VERSION already published — skipping"
continue
fi
dir="npm-dist/$suffix"
mkdir -p "$dir"
cp "$node_file" "$dir/"
node -e '
const [suffix, version] = process.argv.slice(1)
const meta = {
"linux-x64-gnu": { os: ["linux"], cpu: ["x64"], libc: ["glibc"] },
"linux-x64-musl": { os: ["linux"], cpu: ["x64"], libc: ["musl"] },
"linux-arm64-gnu": { os: ["linux"], cpu: ["arm64"], libc: ["glibc"] },
"linux-arm64-musl": { os: ["linux"], cpu: ["arm64"], libc: ["musl"] },
"darwin-arm64": { os: ["darwin"], cpu: ["arm64"] },
"win32-x64-msvc": { os: ["win32"], cpu: ["x64"] },
}[suffix]
if (!meta) {
console.error(`unknown platform suffix: ${suffix} — add it to the meta map`)
process.exit(1)
}
const pkg = {
name: `@firecrawl/pdf-inspector-${suffix}`,
version,
description: `Prebuilt ${suffix} binary for @firecrawl/pdf-inspector`,
main: `pdf-inspector.${suffix}.node`,
files: [`pdf-inspector.${suffix}.node`],
license: "MIT",
engines: { node: ">= 10" },
repository: { type: "git", url: "https://github.com/firecrawl/pdf-inspector" },
publishConfig: { access: "public" },
...meta,
}
require("fs").writeFileSync(`npm-dist/${suffix}/package.json`, JSON.stringify(pkg, null, 2) + "\n")
' "$suffix" "$VERSION"
echo "=== $pkg@$VERSION ==="
ls -la "$dir"
(cd "$dir" && npm publish --provenance --access public)
done
- name: Publish main package
working-directory: napi
run: |
VERSION="${{ needs.check-version.outputs.version }}"
if npm view "@firecrawl/pdf-inspector@$VERSION" version >/dev/null 2>&1; then
echo "@firecrawl/pdf-inspector@$VERSION already published — skipping"
exit 0
fi
cp artifacts/js-bindings/index.js .
cp artifacts/js-bindings/index.d.ts .
echo "=== Package contents ==="
ls -la *.node index.js index.d.ts
# Stamp optionalDependencies to this exact version so the platform
# pins can never drift from the main package version.
node -e '
const fs = require("fs")
const pkg = JSON.parse(fs.readFileSync("package.json", "utf8"))
for (const dep of Object.keys(pkg.optionalDependencies ?? {})) {
pkg.optionalDependencies[dep] = pkg.version
}
fs.writeFileSync("package.json", JSON.stringify(pkg, null, 2) + "\n")
'
echo "=== Main package contents ==="
npm pack --dry-run
npm publish --provenance --access public
+5 -3
View File
@@ -1,10 +1,13 @@
# Rust build artifacts
/target/
/wasm/target/
/wasm/pkg/
debug/
*.pdb
# Cargo lock (optional for libraries)
Cargo.lock
!/wasm/Cargo.lock
# IDE
.idea/
@@ -28,15 +31,14 @@ Thumbs.db
napi/index.js
napi/index.d.ts
# Local samples and scripts
# Local samples
samples/
scripts/
# Test output
test_output/
.firecrawl/
# Python
__pycache__/
*.pyc
.pytest_cache/
+2 -1
View File
@@ -61,7 +61,8 @@ src/
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
- **Regression suite**: sibling repo `pdf-evals` with 179+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
- **Regression suite**: sibling repo `pdf-evals` with ~200 snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing. While iterating, prefer a subset run (`bench.py test -q` for the quick set, or `-s <name>` for a named test set) and save the full `bench.py test` for the final pre-commit check.
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
## Debugging
+1 -1
View File
@@ -61,7 +61,7 @@ src/
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
- **Regression suite**: sibling repo `pdf-evals` with 187+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
- **Regression suite**: sibling repo `pdf-evals` with ~200 snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing. While iterating, prefer a subset run (`bench.py test -q` for the quick set, or `-s <name>` for a named test set) and save the full `bench.py test` for the final pre-commit check.
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
## Debugging
+28 -9
View File
@@ -1,12 +1,25 @@
[package]
name = "pdf-inspector"
version = "0.1.0"
version = "1.14.0"
edition = "2021"
autobins = false
authors = ["Firecrawl Team"]
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
license = "MIT"
repository = "https://github.com/firecrawl/pdf-inspector"
readme = "docs/rust-api.md"
# Explicit allowlist: crates.io caps uploads at 10 MiB and tests/fixtures
# alone exceeds that. external/bcmaps ships in the crate — tounicode.rs
# loads it at runtime relative to CARGO_MANIFEST_DIR.
include = [
"/src/**",
"/external/bcmaps/**",
"/docs/rust-api.md",
"/LICENSE",
# maturin derives the sdist file list from this allowlist; the stub must
# ship so wheels built from the sdist keep their type hints.
"/pdf_inspector.pyi",
]
[lib]
name = "pdf_inspector"
@@ -14,20 +27,13 @@ crate-type = ["lib", "cdylib"]
[dependencies]
# Python bindings
pyo3 = { version = "0.25", features = ["extension-module"], optional = true }
# PDF parsing
lopdf = { git = "https://github.com/J-F-Liu/lopdf", rev = "7a05512d831415b1f2b1ce522391d6beab8a1284", features = ["rayon"] }
pyo3 = { version = "0.25", features = ["extension-module", "abi3-py38"], optional = true }
# Error handling
thiserror = "2.0"
# Parallel processing
rayon = "1.10"
# Logging
log = "0.4"
env_logger = "0.11"
# Text processing
regex = "1.10"
@@ -37,6 +43,19 @@ unicode-normalization = "0.1"
# TrueType font parsing (for Identity-H CID font cmap extraction)
ttf-parser = "0.25"
# Native builds keep lopdf's parallel parser and CLI logging. Browser WASM is
# deliberately single-threaded so it works without cross-origin isolation.
[target.'cfg(not(target_arch = "wasm32"))'.dependencies]
lopdf = { version = "0.42.0", features = ["rayon"] }
rayon = "1.10"
env_logger = "0.11"
# Browser builds use JavaScript randomness for encrypted PDFs and embed the
# bundled CMaps because there is no filesystem at runtime.
[target.'cfg(target_arch = "wasm32")'.dependencies]
lopdf = { version = "0.42.0", default-features = false, features = ["wasm_js"] }
include_dir = "0.7"
[dev-dependencies]
tempfile = "3.3"
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Firecrawl
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+69 -20
View File
@@ -1,6 +1,11 @@
# pdf-inspector
Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for [Python](docs/python.md) and [Node.js](napi/README.md).
[![Crates.io](https://img.shields.io/crates/v/pdf-inspector.svg)](https://crates.io/crates/pdf-inspector)
[![npm](https://img.shields.io/npm/v/@firecrawl/pdf-inspector.svg)](https://www.npmjs.com/package/@firecrawl/pdf-inspector)
[![PyPI](https://img.shields.io/pypi/v/pdf-inspector.svg)](https://pypi.org/project/pdf-inspector/)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for [Python](docs/python.md), [Node.js](napi/README.md), and [browser WebAssembly](wasm/README.md).
Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.
@@ -14,24 +19,28 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
- **Multi-column layout** — Automatic detection of newspaper-style columns, sequential reading order, and RTL text support.
- **Encoding issue detection** — Automatically flags broken font encodings so callers can fall back to OCR.
- **Single document load** — The document is parsed once and shared between detection and extraction, avoiding redundant I/O.
- **Browser WebAssembly** — Run the same Rust parser locally in browsers and Web Workers, with embedded CMaps and no server round trip.
- **Lightweight** — Pure Rust, no ML models, no external services. Single dependency on `lopdf` for PDF parsing.
## Benchmark
Evaluated on the [opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs). Only direct text extraction engines are shown — no OCR, no ML models. Scores are 0-1, higher is better.
Evaluated on the [opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs). Only local engines without model-based PDF parsing are shown; OCR was disabled. Scores are 0-1, higher is better.
| Engine | Overall | Reading Order (NID) | Tables (TEDS) | Headings (MHS) | Speed (200 docs) |
|---|---|---|---|---|---|
| pdf-inspector | 0.78 | 0.87 | 0.59 | 0.57 | 4s |
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
| markitdown | 0.58 | 0.88 | 0.00 | 0.00 | 8s |
| pdf-inspector | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
For context, engines that use OCR/ML (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus.
Results were refreshed on July 31, 2026, on an Apple M4 Pro. Engine versions were pdf-inspector 0.2.6, LiteParse 2.10.1, OpenDataLoader 2.2.1, PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.5. Speed is the median of five alternating or rotating complete corpus runs after an excluded warm-up run, with each parser processing documents sequentially in a single process.
**Where we do well:** Speed (fastest of all engines), reading order, table detection vs other direct-text tools.
The complete parser configuration, per-document predictions, evaluator output, and generated charts are available in the [reproducible results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
**Where we lag:** Heading detection trails opendataloader — many PDFs use bold text at body font size for headings, or headings that are only slightly larger than body text. Table detection trails OCR-based engines that can see visual table structure.
**Best fit:** Native-text PDFs where speed, reading order, and table structure matter. In this comparison, pdf-inspector delivered the higher overall, reading-order, and table scores, along with the fastest complete run. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.
Use the [paired benchmark harness](docs/benchmarking.md) to compare two local builds against the exact same corpus and evaluator revision.
## Quick start
@@ -69,11 +78,39 @@ console.log(result.markdown); // Markdown string or null
> Full API reference: [napi/README.md](napi/README.md)
### Browser WebAssembly
```bash
npm install @firecrawl/pdf-inspector-wasm
```
```javascript
import init, { processPdf } from '@firecrawl/pdf-inspector-wasm';
await init();
const response = await fetch('/document.pdf');
const pdf = new Uint8Array(await response.arrayBuffer());
const result = processPdf(pdf);
console.log(result.pdfType);
console.log(result.markdown);
```
> Full API reference: [wasm/README.md](wasm/README.md)
### Rust
Install from [crates.io](https://crates.io/crates/pdf-inspector):
```bash
cargo add pdf-inspector
```
Or add it manually:
```toml
[dependencies]
pdf-inspector = { git = "https://github.com/firecrawl/pdf-inspector" }
pdf-inspector = "1"
```
```rust
@@ -91,29 +128,40 @@ if let Some(markdown) = &result.markdown {
### CLI
```bash
# Install the CLI tools
cargo install pdf-inspector
# Convert PDF to Markdown
cargo run --bin pdf2md -- document.pdf
pdf2md document.pdf
# JSON output (for piping)
cargo run --bin pdf2md -- document.pdf --json
pdf2md document.pdf --json
# Positioned TextItem JSON, including is_underline metadata
pdf2md document.pdf --items-json
# Raw markdown only (no headers)
cargo run --bin pdf2md -- document.pdf --raw
pdf2md document.pdf --raw
# Token-efficient output (collapses long dot leaders and similar source padding)
pdf2md document.pdf --compact
# Insert page break markers (<!-- Page N -->)
cargo run --bin pdf2md -- document.pdf --pages
pdf2md document.pdf --pages
# Process only specific pages
cargo run --bin pdf2md -- document.pdf --select-pages 1,3,5-10
pdf2md document.pdf --select-pages 1,3,5-10
# Detection only (no extraction)
cargo run --bin detect-pdf -- document.pdf
cargo run --bin detect-pdf -- document.pdf --json
detect-pdf document.pdf
detect-pdf document.pdf --json
# Detection + layout analysis (tables, columns)
cargo run --bin detect-pdf -- document.pdf --analyze --json
detect-pdf document.pdf --analyze --json
```
From a source checkout, use `cargo run --bin pdf2md -- document.pdf` or `cargo run --bin detect-pdf -- document.pdf` instead.
## Architecture
```
@@ -161,6 +209,7 @@ src/
markdown/ — Markdown conversion and structure detection
bin/ — CLI tools (pdf2md, detect_pdf)
napi/ — Node.js/Bun bindings (napi-rs)
wasm/ — Browser bindings (wasm-bindgen)
```
## How classification works
@@ -189,7 +238,7 @@ The converter handles:
|---|---|
| Headings (H1-H4) | Font size tiers relative to body text, with 0.5pt clustering |
| Bold/italic | Font name patterns (Bold, Italic, Oblique) |
| Bullet lists | `*`, `-`, `*`, `○`, `●`, `◦` prefixes |
| Bullet lists | ``, `-`, `*`, `○`, `●`, `◦` prefixes |
| Numbered lists | `1.`, `1)`, `(1)` patterns |
| Letter lists | `a.`, `a)`, `(a)` patterns |
| Code blocks | Monospace fonts (Courier, Consolas, Monaco, Menlo, Fira Code, JetBrains Mono) and keyword detection |
@@ -223,4 +272,4 @@ See [docs/debugging.md](docs/debugging.md) for `RUST_LOG` environment variable u
## License
MIT
[MIT](LICENSE)
+4 -3
View File
@@ -5,14 +5,15 @@
If you believe you've found a security vulnerability in pdf-inspector, please
report it privately so we can fix it before public disclosure.
**Preferred:** Email **help@firecrawl.dev** with:
**Preferred:** Submit through Firecrawl's Bugcrowd vulnerability disclosure
program at <https://bugcrowd.com/engagements/firecrawl-vdp-ess>. Please include:
- A description of the issue and its impact
- Steps to reproduce (a minimal PDF or input that triggers the bug is ideal)
- The version or commit hash of pdf-inspector you tested against
**Alternative:** Use GitHub's private vulnerability reporting under the
[Security tab](https://github.com/firecrawl/pdf-inspector/security/advisories/new).
**Alternative:** If you'd rather not use Bugcrowd, email
**help@firecrawl.dev** with the same details.
We'll acknowledge your report in a timely manner and keep you updated on
remediation progress. Please do not open a public GitHub issue for security
+62
View File
@@ -0,0 +1,62 @@
# Benchmarking against OpenDataLoader
The paired harness runs two `pdf2md` binaries through the same local
OpenDataLoader corpus, evaluates both outputs, and reports aggregate and
per-document deltas. This avoids comparing results produced from different
corpus revisions or evaluator versions.
Build a candidate and provide a released or worktree build as the baseline:
```bash
cargo build --release
python3 scripts/bench_opendataloader.py \
--bench-dir ../opendataloader-bench \
--baseline ../pdf-inspector-main/target/release/pdf2md \
--candidate target/release/pdf2md \
--max-document-regression 0.02 \
--json-output /tmp/pdf-inspector-benchmark.json
```
Pass `--reference-evaluation path/to/evaluation.json` to report the candidate
delta against another evaluation, and add `--require-reference-lead` to make a
negative reference delta fail the run. By default, the candidate must not
regress the baseline overall score or introduce missing predictions. Use
`--min-overall-delta` to require a specific aggregate gain.
The OpenDataLoader repository is external and keeps its normal
`prediction/pdf-inspector` output. Paired evaluation copies each run into a
temporary directory before evaluating it, so the baseline and candidate cannot
overwrite one another.
## Published comparison protocol
The public benchmark table was refreshed on July 31, 2026, on an Apple M4 Pro
using pdf-inspector 0.2.6, LiteParse 2.10.1, OpenDataLoader 2.2.1,
PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.5. Every engine processed the same 200
PDFs sequentially in a single process with OCR disabled. Reported speed is the
median of five alternating or rotating complete corpus runs after an excluded
warm-up run; quality scores come from the benchmark evaluator over all 200
outputs. Raw timings, predictions, evaluations, and charts are available in the
[results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
## Optional backend evidence probe
The evidence probe compares positioned `pdf2md` items with MuPDF structured
text on the same pages. It is intended to find deterministic extraction or
layout evidence that could justify a future native implementation; it does not
merge MuPDF output into Markdown, invoke OCR, or add a runtime dependency.
Install MuPDF's `mutool`, build `pdf2md`, then run:
```bash
python3 scripts/probe_backend_evidence.py document.pdf \
--pdf2md target/release/pdf2md \
--json-output /tmp/backend-evidence.json
```
The report flags pages when MuPDF exposes a material net token gain, repeated
alignment anchors absent from local evidence, or additional image blocks. The
JSON includes bounded token samples and page-level counts so promising cases
can be inspected without treating backend disagreement as automatically
correct. Thresholds are configurable with `--min-token-gain`,
`--min-alternate-only-ratio`, and `--min-anchor-gain`.
+57
View File
@@ -0,0 +1,57 @@
# Publishing
Every pdf-inspector distribution uses one shared semantic version:
- Rust crate: `pdf-inspector`
- Python package: `pdf-inspector`
- Node package: `@firecrawl/pdf-inspector` and its platform packages
- Browser package: `@firecrawl/pdf-inspector-wasm`
- Internal NAPI and WASM Rust crates
`Cargo.toml` is the canonical version source. Update every manifest and lockfile
with:
```bash
python3 scripts/version.py <version>
```
Verify that nothing has diverged with:
```bash
python3 scripts/version.py --check
```
CI and every publishing workflow run this check before building or publishing.
## Release steps
1. Choose the next shared semantic version and run `scripts/version.py`.
2. Review the manifest and lockfile changes in the version-bump pull request.
3. Merge the pull request to `main`.
4. The crates.io, PyPI, Node, and WASM workflows independently build and
publish that version from the same commit.
5. After all registries succeed, create one `v<version>` GitHub release that
links to each package and describes changes since the previous shared tag.
The independent workflows are intentionally idempotent. A manual dispatch from
`main` can repair a partial release, and already-published artifacts are skipped.
## Trusted publishers
The repositories use GitHub Actions OIDC instead of long-lived registry tokens.
Configure each registry's trusted publisher for `firecrawl/pdf-inspector` and
its corresponding workflow:
- crates.io: `publish-crate.yml`, environment `crates-io`
- PyPI: `publish-pypi.yml`, environment `pypi`
- npm Node package: `publish.yml`
- npm WASM package: `publish-wasm.yml`
The WASM package must exist before npm trusted publishing can be configured. If
it ever needs to be bootstrapped again, build and inspect it before publishing:
```bash
wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release
npm pack --dry-run ./wasm/pkg
npm publish ./wasm/pkg --access public
```
+104 -9
View File
@@ -1,9 +1,39 @@
# Python API
# pdf-inspector
Python bindings via [PyO3](https://pyo3.rs). Requires Rust toolchain for building from source.
Fast PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Python bindings via [PyO3](https://pyo3.rs) for the [pdf-inspector](https://github.com/firecrawl/pdf-inspector) Rust library.
Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.
## Features
- **Smart classification** — `text_based` / `scanned` / `image_based` / `mixed` in ~1050ms, with a confidence score and per-page OCR routing.
- **Markdown conversion** — headings, lists, code blocks, bold/italic, URL linking, and dual-mode table detection (PDF drawing ops + text-alignment heuristics).
- **Layout-aware extraction** — multi-column reading order, position and font info per text item, RTL support.
- **Robust text decoding** — CID/Type0 fonts via ToUnicode CMaps, plus automatic flagging of broken encodings so callers can fall back to OCR.
- **Lightweight** — native Rust core, no ML models, no external services; ships type stubs.
## Benchmark
[opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 01, higher is better:
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|---|---|---|---|---|---|
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark), with raw timings and artifacts in the [results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
## Install
```bash
pip install pdf-inspector
```
Prebuilt wheels cover CPython ≥3.8 on Linux (x86_64, aarch64), macOS (Intel, Apple Silicon), and Windows (x64). Other platforms build from source, which requires a Rust toolchain. For local development in a repo checkout:
```bash
pip install maturin
maturin develop --release
@@ -50,6 +80,17 @@ for page in result.pages:
# Restrict to specific 0-indexed pages (preserves caller order)
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
# Structure-tree elements from tagged PDFs (empty list when untagged).
# Pages are 1-indexed to match TextItem.page, so (page, mcid) joins directly
# against extract_text_with_positions — e.g. to recover real heading levels:
elements = pdf_inspector.extract_structure_elements("tagged.pdf")
roles = {(e.page, e.mcid): e.role for e in elements}
headings = [
item.text
for item in pdf_inspector.extract_text_with_positions("tagged.pdf")
if item.mcid is not None and roles.get((item.page, item.mcid), "").startswith("H")
]
```
## API reference
@@ -70,19 +111,73 @@ result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes |
| `extract_pages_markdown(path, pages=None)` | Per-page Markdown + layout metadata (all pages by default) |
| `extract_pages_markdown_bytes(data, pages=None)` | Per-page Markdown from bytes |
| `extract_structure_elements(path, pages=None)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
| `extract_structure_elements_bytes(data, pages=None)` | Structure-tree elements from bytes |
## Types
**`PdfResult` fields:** `pdf_type`, `markdown`, `page_count`, `processing_time_ms`, `pages_needing_ocr`, `title`, `confidence`, `is_complex_layout`, `pages_with_tables`, `pages_with_columns`, `has_encoding_issues`
Type stubs (`pdf_inspector.pyi`) ship with the package. Result types at a glance:
**`PdfClassification` fields:** `pdf_type`, `page_count`, `pages_needing_ocr` (0-indexed), `confidence`
```python
class PdfResult: # process_pdf / detect_pdf
pdf_type: str # "text_based" | "scanned" | "image_based" | "mixed"
markdown: str | None # extracted Markdown (None for detect_pdf)
page_count: int
processing_time_ms: int
pages_needing_ocr: list[int] # 1-indexed
ocr_reasons_by_page: list[PageOcrReasons]
title: str | None
confidence: float # 0.0 - 1.0
is_complex_layout: bool
pages_with_tables: list[int]
pages_with_columns: list[int]
has_encoding_issues: bool # broken font encodings — consider OCR fallback
**`TextItem` fields:** `text`, `x`, `y`, `width`, `height`, `font`, `font_size`, `page`, `is_bold`, `is_italic`, `item_type`
class PageOcrReasons: # per-page OCR diagnostics
page: int # 1-indexed
reasons: list[str] # machine-readable reason identifiers
**`RegionText` fields:** `text`, `needs_ocr`
class PdfClassification: # classify_pdf
pdf_type: str
page_count: int
pages_needing_ocr: list[int] # 0-indexed
confidence: float
**`PageRegionTexts` fields:** `page` (0-indexed), `regions` (list of RegionText)
class TextItem: # extract_text_with_positions
text: str
x: float
y: float
width: float
height: float
font: str
font_size: float
page: int
is_bold: bool
is_italic: bool
is_underline: bool
is_strikeout: bool
item_type: str
mcid: int | None # marked-content ID for tagged PDFs (None otherwise)
**`PageMarkdown` fields:** `page` (0-indexed), `markdown`, `needs_ocr`
class StructureElement: # extract_structure_elements
page: int # 1-indexed (matches TextItem.page)
mcid: int
role: str # "H1".."H6", "P", "Table", ... (resolved via /RoleMap)
**`PagesExtractionResult` fields:** `pages` (list of PageMarkdown), `pages_with_tables` (1-indexed), `pages_with_columns` (1-indexed), `pages_needing_ocr` (1-indexed), `is_complex`
class RegionText: # extract_text_in_regions
text: str
needs_ocr: bool
ocr_reason: str | None # machine-readable OCR reason
class PageRegionTexts: # extract_text_in_regions
page: int # 0-indexed
regions: list[RegionText]
class PagesExtractionResult: # extract_pages_markdown
pages: list[PageMarkdown] # PageMarkdown: page (0-indexed), markdown, needs_ocr, ocr_reason
pages_with_tables: list[int] # 1-indexed
pages_with_columns: list[int] # 1-indexed
pages_needing_ocr: list[int] # 1-indexed
ocr_reasons_by_page: list[PageOcrReasons]
is_complex: bool # any page has tables or multi-column layout
```
+72 -3
View File
@@ -1,12 +1,50 @@
# Rust API
# pdf-inspector
Add to your `Cargo.toml`:
Fast PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Pure Rust, no ML models, no external services; the only PDF dependency is [lopdf](https://crates.io/crates/lopdf). Also available for [Python](https://pypi.org/project/pdf-inspector/) and [Node.js](https://www.npmjs.com/package/@firecrawl/pdf-inspector).
Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.
## Features
- **Smart classification** — TextBased / Scanned / ImageBased / Mixed in ~1050ms, with a confidence score and per-page OCR routing.
- **Markdown conversion** — headings, lists, code blocks, bold/italic, URL linking, and dual-mode table detection (PDF drawing ops + text-alignment heuristics).
- **Layout-aware extraction** — multi-column reading order, position and font info per text item, RTL support.
- **Robust text decoding** — CID/Type0 fonts via ToUnicode CMaps, plus automatic flagging of broken encodings so callers can fall back to OCR.
- **Lightweight** — pure Rust, no ML models, no external services; single PDF dependency ([lopdf](https://crates.io/crates/lopdf)).
## Benchmark
[opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 01, higher is better:
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|---|---|---|---|---|---|
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark), with raw timings and artifacts in the [results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
## Install
```bash
cargo add pdf-inspector
```
For the latest unreleased changes, use the git dependency instead:
```toml
[dependencies]
pdf-inspector = { git = "https://github.com/firecrawl/pdf-inspector" }
```
The crate also ships CLI binaries — `pdf2md` (PDF → Markdown, with `--json`, `--pages`, `--select-pages`, and the opt-in token-saving `--compact` profile) and `detect-pdf` (classification, with `--analyze --json`):
```bash
cargo install pdf-inspector
```
## Usage
Detect and extract in one call:
@@ -100,6 +138,34 @@ for page in &result.pages {
println!("Complex layout? {}", result.is_complex);
```
Extract structure-tree elements from tagged PDFs, and join them against
`extract_text_with_positions` to attach semantic roles (heading levels,
paragraphs, table cells) to extracted text:
```rust
use pdf_inspector::{extract_structure_elements, extract_text_with_positions};
use std::collections::HashMap;
// One entry per marked-content reference, sorted by (page, mcid); empty for
// untagged PDFs. Pages are 1-indexed to match `TextItem::page`, so the
// (page, mcid) pair is a direct join key.
let elements = extract_structure_elements("tagged.pdf", None)?;
let roles: HashMap<(u32, i64), &str> = elements
.iter()
.map(|e| ((e.page, e.mcid), e.role.as_str()))
.collect();
for item in extract_text_with_positions("tagged.pdf")? {
if let Some(mcid) = item.mcid {
if let Some(role) = roles.get(&(item.page, mcid)) {
if role.starts_with('H') {
println!("{}: {}", role, item.text);
}
}
}
}
```
## Processing modes
| Mode | What it does | Returns |
@@ -125,6 +191,8 @@ println!("Complex layout? {}", result.is_complex);
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
| `extract_pages_markdown(path, pages)` | Per-page Markdown + layout metadata (file) |
| `extract_pages_markdown_mem(bytes, pages)` | Per-page Markdown from bytes |
| `extract_structure_elements(path, pages)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
| `extract_structure_elements_mem(bytes, pages)` | Structure-tree elements from bytes |
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
@@ -140,7 +208,8 @@ Low-level detection functions are also available via the `detector` module (`det
| `DetectionConfig` | Configuration for detection: scan strategy, thresholds |
| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` |
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
| `TextItem` | Text with position, font info, and page number |
| `TextItem` | Text with position, font info, page number, and optional structure-tree `mcid` |
| `StructureElement` | Tagged-PDF structure reference: page (1-indexed), mcid, role (`"H1"`..`"H6"`, `"P"`, …) |
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
+27 -4
View File
@@ -499,11 +499,13 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0de51e6874e94e7bf76d726fc5d13ba782deca734ff60d5bb2fb2607c7406555"
dependencies = [
"cfg-if",
"js-sys",
"libc",
"r-efi",
"rand_core",
"wasip2",
"wasip3",
"wasm-bindgen",
]
[[package]]
@@ -557,6 +559,25 @@ version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3d3067d79b975e8844ca9eb072e16b31c3c1c36928edf9c6789548c524d0d954"
[[package]]
name = "include_dir"
version = "0.7.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "923d117408f1e49d914f1a379a309cffe4f18c05cf4e3d12e613a15fc81bd0dd"
dependencies = [
"include_dir_macros",
]
[[package]]
name = "include_dir_macros"
version = "0.7.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7cab85a7ed0bd5f0e76d93846e0147172bed2e2d3f859bcc33a8d9699cad1a75"
dependencies = [
"proc-macro2",
"quote",
]
[[package]]
name = "indexmap"
version = "2.13.0"
@@ -672,8 +693,9 @@ checksum = "5e5032e24019045c762d3c0f28f5b6b8bbf38563a65908389bf7978758920897"
[[package]]
name = "lopdf"
version = "0.40.0"
source = "git+https://github.com/J-F-Liu/lopdf?rev=7a05512d831415b1f2b1ce522391d6beab8a1284#7a05512d831415b1f2b1ce522391d6beab8a1284"
version = "0.42.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "25aab26d99567469098e64a02f42679f8965c6401263eefa31d8f2dcc37a221c"
dependencies = [
"aes",
"bitflags",
@@ -829,9 +851,10 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
[[package]]
name = "pdf-inspector"
version = "0.1.0"
version = "1.14.0"
dependencies = [
"env_logger",
"include_dir",
"log",
"lopdf",
"once_cell",
@@ -844,7 +867,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-napi"
version = "0.2.0"
version = "1.14.0"
dependencies = [
"napi",
"napi-build",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-napi"
version = "0.2.0"
version = "1.14.0"
edition = "2021"
[lib]
+51 -6
View File
@@ -4,6 +4,28 @@ Fast PDF classification and region-based text extraction for Node.js/Bun. Native
Built by [Firecrawl](https://firecrawl.dev) for hybrid OCR pipelines — extract text from PDF structure where possible, fall back to OCR only when needed.
## Features
- **Smart classification** — text-based / scanned / image-based / mixed in ~1050ms, with a confidence score and per-page OCR routing.
- **Region-based extraction** — pull text from bounding boxes with per-region quality checks (`needsOcr`).
- **Layout-aware** — multi-column reading order, position and font info per text item, RTL support.
- **Robust text decoding** — CID/Type0 fonts via ToUnicode CMaps, plus automatic flagging of broken encodings so callers can fall back to OCR.
- **Lightweight** — native Rust core via napi-rs, no ML models, no external services; ~56 MB platform binary, TypeScript definitions included.
## Benchmark
[opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 01, higher is better:
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|---|---|---|---|---|---|
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark), with raw timings and artifacts in the [results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
## Install
```bash
@@ -12,7 +34,7 @@ npm install @firecrawl/pdf-inspector
bun add @firecrawl/pdf-inspector
```
Prebuilt binaries included for **linux-x64** and **macOS ARM64**. No Rust toolchain needed.
Prebuilt binaries for **Linux x64/ARM64** (glibc and musl/Alpine), **macOS ARM64**, and **Windows x64** — npm installs only the one matching your platform. No Rust toolchain needed.
## API
@@ -37,7 +59,7 @@ console.log(result.confidence) // 0.875
Extract text within bounding-box regions from a PDF. Designed for hybrid OCR pipelines where a layout model detects regions in rendered page images, and this function extracts text from the PDF structure for text-based pages — skipping GPU OCR.
Each region result includes a `needsOcr` flag that signals unreliable extraction (empty text, GID-encoded fonts, garbage text, encoding issues).
Each region result includes a `needsOcr` flag that signals unreliable extraction (empty text, GID-encoded fonts, garbage text, encoding issues). When the cause is a suspected garbled text layer, `ocrReason` is set to `"suspected_garbled_text"`.
```typescript
import { extractTextInRegions } from '@firecrawl/pdf-inspector'
@@ -61,6 +83,22 @@ for (const region of result[0].regions) {
}
```
### Async variants
`processPdf`, `classifyPdf`, and `extractPagesMarkdown` are synchronous and parse on the calling thread — in Node, that's the event loop. For a one-off call in a script that's fine, but in a server a large document can hold the loop for tens to hundreds of milliseconds.
`processPdfAsync`, `classifyPdfAsync`, and `extractPagesMarkdownAsync` take the same arguments and produce the same results, but run the parse on the libuv thread pool and return a promise, keeping the event loop free. The input buffer is copied before the call returns, so it's safe to reuse or mutate immediately:
```typescript
import { classifyPdfAsync, extractPagesMarkdownAsync } from '@firecrawl/pdf-inspector'
const classification = await classifyPdfAsync(pdf)
if (classification.pdfType === 'TextBased') {
const { pages } = await extractPagesMarkdownAsync(pdf)
// ...
}
```
## Types
```typescript
@@ -84,15 +122,22 @@ interface PageRegionTexts {
interface RegionText {
text: string
needsOcr: boolean // true when text is unreliable
ocrReason?: string // "suspected_garbled_text" when known
}
```
## Platforms
| Platform | Architecture | Supported |
|----------|-------------|-----------|
| Linux | x64 | Yes |
| macOS | ARM64 | Yes |
Prebuilt binaries ship as platform-specific packages installed automatically via `optionalDependencies`:
| Platform | Architecture | Package |
|----------|-------------|---------|
| Linux | x64 (glibc) | `@firecrawl/pdf-inspector-linux-x64-gnu` |
| Linux | x64 (musl/Alpine) | `@firecrawl/pdf-inspector-linux-x64-musl` |
| Linux | ARM64 (glibc) | `@firecrawl/pdf-inspector-linux-arm64-gnu` |
| Linux | ARM64 (musl/Alpine) | `@firecrawl/pdf-inspector-linux-arm64-musl` |
| macOS | ARM64 | `@firecrawl/pdf-inspector-darwin-arm64` |
| Windows | x64 | `@firecrawl/pdf-inspector-win32-x64-msvc` |
## License
+8
View File
@@ -7,6 +7,14 @@
"devDependencies": {
"@napi-rs/cli": "^3.4.1",
},
"optionalDependencies": {
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0",
},
},
},
"packages": {
+13 -6
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.9.2",
"version": "1.14.0",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
@@ -22,7 +22,6 @@
"files": [
"index.js",
"index.d.ts",
"*.node",
"bin/",
"README.md"
],
@@ -38,12 +37,12 @@
"binaryName": "pdf-inspector",
"targets": [
"x86_64-unknown-linux-gnu",
"x86_64-unknown-linux-musl",
"aarch64-unknown-linux-gnu",
"aarch64-unknown-linux-musl",
"aarch64-apple-darwin",
"x86_64-pc-windows-msvc"
],
"package": {
"name": "@firecrawl/pdf-inspector-js"
}
]
},
"scripts": {
"build": "napi build --platform --release",
@@ -51,5 +50,13 @@
},
"devDependencies": {
"@napi-rs/cli": "^3.4.1"
},
"optionalDependencies": {
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0"
}
}
+270 -36
View File
@@ -40,6 +40,8 @@ pub struct PdfResult {
pub processing_time_ms: u32,
/// 1-indexed page numbers that need OCR.
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
pub ocr_reasons_by_page: Vec<PageOcrReasons>,
pub title: Option<String>,
pub confidence: f64,
pub is_complex_layout: bool,
@@ -48,6 +50,13 @@ pub struct PdfResult {
pub has_encoding_issues: bool,
}
/// OCR reasons for a single 1-indexed page.
#[napi(object)]
pub struct PageOcrReasons {
pub page: u32,
pub reasons: Vec<String>,
}
/// Lightweight PDF classification result.
#[napi(object)]
pub struct PdfClassification {
@@ -71,9 +80,20 @@ pub struct TextItem {
pub page: u32,
pub is_bold: bool,
pub is_italic: bool,
/// Underline detected geometrically (drawn rule/thin rect under the
/// baseline) — PDFs carry no underline font flag.
pub is_underline: bool,
/// Strikeout detected geometrically (rule crossing the glyphs at mid
/// x-height).
pub is_strikeout: bool,
pub item_type: ItemType,
/// URL for link items, `None` for other types.
pub link_url: Option<String>,
/// Marked Content ID from the content stream's BDC/BMC operator, `None`
/// when the text is not part of marked content. Join with the
/// `page`/`mcid` pairs from [`extractStructureElements`] to attach
/// structure-tree roles (headings, paragraphs, …) in tagged PDFs.
pub mcid: Option<i64>,
}
/// A page's regions for text extraction: (page_index_0based, bboxes).
@@ -90,6 +110,8 @@ pub struct RegionText {
pub text: String,
/// `true` when the text should not be trusted (empty, GID fonts, garbage, encoding issues).
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
pub ocr_reason: Option<String>,
}
/// Extracted text for one page's regions.
@@ -126,6 +148,7 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
page_count: r.page_count,
processing_time_ms: r.processing_time_ms as u32,
pages_needing_ocr: r.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(r.ocr_reasons_by_page),
title: r.title,
confidence: r.confidence as f64,
is_complex_layout: r.layout.is_complex,
@@ -135,6 +158,16 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
}
}
fn to_napi_page_ocr_reasons(reasons: Vec<pdf_inspector::PageOcrReasons>) -> Vec<PageOcrReasons> {
reasons
.into_iter()
.map(|reason| PageOcrReasons {
page: reason.page,
reasons: reason.reasons,
})
.collect()
}
fn convert_item_type(t: &pdf_inspector::types::ItemType) -> (ItemType, Option<String>) {
match t {
pdf_inspector::types::ItemType::Text => (ItemType::Text, None),
@@ -172,6 +205,31 @@ where
}
}
// ---------------------------------------------------------------------------
// Shared implementations (single body behind sync and async entry points)
// ---------------------------------------------------------------------------
fn process_pdf_impl(bytes: &[u8], pages: Option<Vec<u32>>) -> Result<PdfResult> {
let mut opts = pdf_inspector::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = pdf_inspector::process_pdf_mem_with_options(bytes, opts)
.map_err(|e| to_napi_err(e, "process_pdf"))?;
Ok(to_napi_result(result))
}
fn classify_pdf_impl(bytes: &[u8]) -> Result<PdfClassification> {
let result =
pdf_inspector::classify_pdf_mem(bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
Ok(PdfClassification {
pdf_type: convert_pdf_type(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
}
// ---------------------------------------------------------------------------
// Public NAPI API
// ---------------------------------------------------------------------------
@@ -180,15 +238,7 @@ where
#[napi]
pub fn process_pdf(buffer: Buffer, pages: Option<Vec<u32>>) -> Result<PdfResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("process_pdf", move || {
let mut opts = pdf_inspector::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = pdf_inspector::process_pdf_mem_with_options(&bytes, opts)
.map_err(|e| to_napi_err(e, "process_pdf"))?;
Ok(to_napi_result(result))
})
catch_panic("process_pdf", move || process_pdf_impl(&bytes, pages))
}
/// Fast detection only — no text extraction or markdown.
@@ -208,16 +258,7 @@ pub fn detect_pdf(buffer: Buffer) -> Result<PdfResult> {
#[napi]
pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("classify_pdf", move || {
let result =
pdf_inspector::classify_pdf_mem(&bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
Ok(PdfClassification {
pdf_type: convert_pdf_type(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
})
catch_panic("classify_pdf", move || classify_pdf_impl(&bytes))
}
/// Extract plain text from a PDF Buffer.
@@ -266,14 +307,65 @@ pub fn extract_text_with_positions(
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type,
link_url,
mcid: item.mcid,
}
})
.collect())
})
}
/// One structure-tree element reference from a tagged PDF.
#[napi(object)]
pub struct StructureElementJs {
/// 1-indexed page number (matches `TextItem.page`).
pub page: u32,
/// Marked Content ID from the page's content stream (matches
/// `TextItem.mcid`).
pub mcid: i64,
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
/// Custom tags are resolved through the document's role map; tags with
/// no standard mapping are returned verbatim.
pub role: String,
}
/// Extract structure-tree element references from a tagged PDF.
///
/// Parses the document's structure tree (when present) and returns one
/// entry per marked-content reference, resolved to its 1-indexed page,
/// MCID, and structure type name. Returns an empty array when the PDF is
/// not tagged.
///
/// Join `(page, mcid)` against the `page`/`mcid` fields from
/// [`extractTextWithPositions`] to attach heading levels (H1..H6) and other
/// semantic roles to extracted text.
///
/// Pass 1-indexed page numbers (matching `TextItem.page`) to restrict
/// output; omit `pages` for the whole document. Entries are sorted by
/// `(page, mcid)`.
#[napi]
pub fn extract_structure_elements(
buffer: Buffer,
pages: Option<Vec<u32>>,
) -> Result<Vec<StructureElementJs>> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_structure_elements", move || {
let elements = pdf_inspector::extract_structure_elements_mem(&bytes, pages.as_deref())
.map_err(|e| to_napi_err(e, "extract_structure_elements"))?;
Ok(elements
.into_iter()
.map(|e| StructureElementJs {
page: e.page,
mcid: e.mcid,
role: e.role,
})
.collect())
})
}
/// Extract text within bounding-box regions from a PDF.
///
/// For hybrid OCR: layout model detects regions in rendered images,
@@ -563,6 +655,8 @@ pub struct PageMarkdownResult {
pub markdown: String,
/// `true` when text on this page is unreliable.
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
pub ocr_reason: Option<String>,
}
/// Combined per-page markdown extraction and layout classification result.
@@ -576,6 +670,8 @@ pub struct PagesExtractionResult {
pub pages_with_columns: Vec<u32>,
/// 1-indexed pages that need OCR (scanned/image-based).
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
pub ocr_reasons_by_page: Vec<PageOcrReasons>,
/// True if any page has tables or columns.
pub is_complex: bool,
}
@@ -597,23 +693,32 @@ pub fn extract_pages_markdown(
) -> Result<PagesExtractionResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_pages_markdown", move || {
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, pages.as_deref())
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
.pages
.into_iter()
.map(|r| PageMarkdownResult {
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
is_complex: result.is_complex,
})
extract_pages_markdown_impl(&bytes, pages.as_deref())
})
}
fn extract_pages_markdown_impl(
bytes: &[u8],
pages: Option<&[u32]>,
) -> Result<PagesExtractionResult> {
let result = pdf_inspector::extract_pages_markdown_mem(bytes, pages)
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
.pages
.into_iter()
.map(|r| PageMarkdownResult {
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
is_complex: result.is_complex,
})
}
@@ -648,8 +753,137 @@ fn to_page_region_texts(results: Vec<pdf_inspector::PageRegionResult>) -> Vec<Pa
.map(|r| RegionText {
text: r.text,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
})
.collect()
}
// ---------------------------------------------------------------------------
// Async variants (libuv thread pool via AsyncTask)
//
// The synchronous exports above parse on the calling thread, which in Node is
// the event loop. These `*Async` variants run the same shared implementations
// on the libuv thread pool and hand JavaScript a promise, so servers under
// concurrent load keep answering requests while a document parses. The sync
// exports keep their names, signatures, and behaviour.
//
// Each factory copies the input Buffer to an owned `Vec<u8>` on the calling
// (JS) thread — deliberately. JS execution is single-threaded, so no JS code
// can mutate the buffer while the synchronous part of the call copies it.
// Holding the napi `Buffer` and reading it from the worker instead would be
// zero-copy, but a caller mutating the buffer before the promise settles
// would then race the worker's reads — undefined behavior, not a recoverable
// error (a known napi-rs soundness hazard with cross-thread Buffer access).
// The copy is a one-time memcpy, negligible next to the parse it unblocks.
// ---------------------------------------------------------------------------
pub struct ProcessPdfTask {
bytes: Vec<u8>,
pages: Option<Vec<u32>>,
}
impl Task for ProcessPdfTask {
type Output = PdfResult;
type JsValue = PdfResult;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
let pages = self.pages.take();
// AssertUnwindSafe: `bytes`/`pages` are moved into the closure and
// dropped on unwind — no shared state can be observed broken.
catch_panic(
"process_pdf",
panic::AssertUnwindSafe(move || process_pdf_impl(&bytes, pages)),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`processPdf`]: same result, but the parse runs on the
/// libuv thread pool instead of the event loop and the call returns a
/// promise. The buffer is copied before the call returns, so it may be
/// reused or mutated immediately.
// ts_return_type is required: napi-rs emits `Promise<unknown>` for
// `AsyncTask<T>` returns without it.
#[napi(ts_return_type = "Promise<PdfResult>")]
pub fn process_pdf_async(buffer: Buffer, pages: Option<Vec<u32>>) -> AsyncTask<ProcessPdfTask> {
AsyncTask::new(ProcessPdfTask {
bytes: buffer.to_vec(),
pages,
})
}
pub struct ClassifyPdfTask {
bytes: Vec<u8>,
}
impl Task for ClassifyPdfTask {
type Output = PdfClassification;
type JsValue = PdfClassification;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
catch_panic(
"classify_pdf",
panic::AssertUnwindSafe(move || classify_pdf_impl(&bytes)),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`classifyPdf`]: same result, but the classification runs
/// on the libuv thread pool instead of the event loop and the call returns a
/// promise. The buffer is copied before the call returns, so it may be
/// reused or mutated immediately.
#[napi(ts_return_type = "Promise<PdfClassification>")]
pub fn classify_pdf_async(buffer: Buffer) -> AsyncTask<ClassifyPdfTask> {
AsyncTask::new(ClassifyPdfTask {
bytes: buffer.to_vec(),
})
}
pub struct ExtractPagesMarkdownTask {
bytes: Vec<u8>,
pages: Option<Vec<u32>>,
}
impl Task for ExtractPagesMarkdownTask {
type Output = PagesExtractionResult;
type JsValue = PagesExtractionResult;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
let pages = self.pages.take();
catch_panic(
"extract_pages_markdown",
panic::AssertUnwindSafe(move || extract_pages_markdown_impl(&bytes, pages.as_deref())),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`extractPagesMarkdown`]: same result, but the extraction
/// runs on the libuv thread pool instead of the event loop and the call
/// returns a promise. The buffer is copied before the call returns, so it
/// may be reused or mutated immediately.
#[napi(ts_return_type = "Promise<PagesExtractionResult>")]
pub fn extract_pages_markdown_async(
buffer: Buffer,
pages: Option<Vec<u32>>,
) -> AsyncTask<ExtractPagesMarkdownTask> {
AsyncTask::new(ExtractPagesMarkdownTask {
bytes: buffer.to_vec(),
pages,
})
}
+110
View File
@@ -2,16 +2,21 @@ import { readFileSync } from 'fs';
import { strict as assert } from 'assert';
import {
processPdf,
processPdfAsync,
detectPdf,
classifyPdf,
classifyPdfAsync,
extractText,
extractTextWithPositions,
extractStructureElements,
extractTextInRegions,
detectVectorGridInRegion,
extractPagesMarkdown,
extractPagesMarkdownAsync,
} from './index.js';
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
const taggedFixture = readFileSync('../tests/fixtures/firecrawl_docs_tagged.pdf');
// --- processPdf ---
console.log('Testing processPdf...');
@@ -79,6 +84,46 @@ assert.ok(page1Items.length > 0);
assert.ok(page1Items.every(i => i.page === 1));
console.log(' extractTextWithPositions with pages: OK');
// mcid: undefined on untagged PDFs, numeric on tagged marked content
assert.ok(items.every(i => i.mcid === undefined || typeof i.mcid === 'number'));
const taggedItems = extractTextWithPositions(taggedFixture);
assert.ok(
taggedItems.some(i => typeof i.mcid === 'number'),
'tagged PDF text items should carry Marked Content IDs',
);
console.log(' extractTextWithPositions mcid: OK');
// --- extractStructureElements ---
console.log('Testing extractStructureElements...');
const structureElements = extractStructureElements(taggedFixture);
assert.ok(structureElements.length > 0);
assert.ok(structureElements.every(e => typeof e.page === 'number'));
assert.ok(structureElements.every(e => typeof e.mcid === 'number'));
assert.ok(structureElements.every(e => typeof e.role === 'string' && e.role.length > 0));
assert.ok(
structureElements.some(e => e.role === 'H1'),
'tagged fixture should surface H1 heading roles',
);
// (page, mcid) joins against extractTextWithPositions to recover heading text
const h1Refs = new Set(
structureElements.filter(e => e.role === 'H1').map(e => `${e.page}:${e.mcid}`),
);
const h1Text = taggedItems
.filter(i => typeof i.mcid === 'number' && h1Refs.has(`${i.page}:${i.mcid}`))
.map(i => i.text)
.join('');
assert.ok(h1Text.trim().length > 0, 'H1 join should recover heading text');
// pages filter is 1-indexed, matching TextItem.page
const page1Elements = extractStructureElements(taggedFixture, [1]);
assert.ok(page1Elements.length > 0);
assert.ok(page1Elements.every(e => e.page === 1));
// untagged PDFs yield an empty array
assert.deepEqual(extractStructureElements(fixture), []);
console.log(' extractStructureElements: OK');
// --- extractTextInRegions ---
console.log('Testing extractTextInRegions...');
const regionResults = extractTextInRegions(fixture, [
@@ -124,10 +169,75 @@ assert.equal(picked.pages[0].page, 2);
assert.equal(picked.pages[1].page, 0);
console.log(' extractPagesMarkdown with pages: OK');
// --- Async variants ---
console.log('Testing async variants...');
// processPdfAsync returns a promise and matches the sync result
const asyncResultPromise = processPdfAsync(fixture);
assert.ok(asyncResultPromise instanceof Promise);
const asyncResult = await asyncResultPromise;
assert.equal(asyncResult.pdfType, result.pdfType);
assert.equal(asyncResult.pageCount, result.pageCount);
assert.equal(asyncResult.markdown, result.markdown);
console.log(' processPdfAsync: OK');
// processPdfAsync with pages
const asyncResult2 = await processPdfAsync(fixture, [1]);
assert.equal(asyncResult2.markdown, result2.markdown);
console.log(' processPdfAsync with pages: OK');
// classifyPdfAsync matches the sync result
const asyncClassified = await classifyPdfAsync(fixture);
assert.equal(asyncClassified.pdfType, classified.pdfType);
assert.equal(asyncClassified.pageCount, classified.pageCount);
assert.equal(asyncClassified.confidence, classified.confidence);
assert.deepEqual(asyncClassified.pagesNeedingOcr, classified.pagesNeedingOcr);
console.log(' classifyPdfAsync: OK');
// extractPagesMarkdownAsync matches the sync result
const asyncAllPages = await extractPagesMarkdownAsync(fixture);
assert.equal(asyncAllPages.pages.length, allPages.pages.length);
assert.deepEqual(
asyncAllPages.pages.map(p => p.markdown),
allPages.pages.map(p => p.markdown),
);
assert.equal(asyncAllPages.isComplex, allPages.isComplex);
console.log(' extractPagesMarkdownAsync: OK');
// selected pages preserve caller order
const asyncPicked = await extractPagesMarkdownAsync(fixture, [2, 0]);
assert.equal(asyncPicked.pages.length, 2);
assert.equal(asyncPicked.pages[0].page, 2);
assert.equal(asyncPicked.pages[1].page, 0);
console.log(' extractPagesMarkdownAsync with pages: OK');
// input buffer is copied at call time: mutating it immediately after the
// call must not affect the in-flight parse
const scratch = Buffer.from(fixture);
const inFlight = processPdfAsync(scratch);
scratch.fill(0);
const fromMutated = await inFlight;
assert.equal(fromMutated.markdown, result.markdown);
console.log(' processPdfAsync input copied at call time: OK');
// concurrent async calls all settle
const [c1, c2, c3] = await Promise.all([
processPdfAsync(fixture),
classifyPdfAsync(fixture),
extractPagesMarkdownAsync(fixture),
]);
assert.equal(c1.pdfType, 'TextBased');
assert.equal(c2.pdfType, 'TextBased');
assert.equal(c3.pages.length, 3);
console.log(' concurrent async calls: OK');
// --- Error handling ---
console.log('Testing error handling...');
assert.throws(() => processPdf(Buffer.from('not a pdf')), /process_pdf/);
assert.throws(() => classifyPdf(Buffer.from('')), /classify_pdf/);
await assert.rejects(processPdfAsync(Buffer.from('not a pdf')), /process_pdf/);
await assert.rejects(classifyPdfAsync(Buffer.from('')), /classify_pdf/);
await assert.rejects(extractPagesMarkdownAsync(Buffer.from('')), /extract_pages_markdown/);
console.log(' error handling: OK');
console.log('\nAll NAPI tests passed!');
+53
View File
@@ -10,6 +10,9 @@ class PdfResult:
page_count: int
processing_time_ms: int
pages_needing_ocr: list[int]
"""1-indexed page numbers that need OCR."""
ocr_reasons_by_page: list["PageOcrReasons"]
"""Machine-readable OCR reasons by 1-indexed page."""
title: Optional[str]
confidence: float
is_complex_layout: bool
@@ -17,6 +20,13 @@ class PdfResult:
pages_with_columns: list[int]
has_encoding_issues: bool
class PageOcrReasons:
"""OCR reasons for a single 1-indexed page."""
page: int
"""1-indexed page number."""
reasons: list[str]
"""Machine-readable OCR reason identifiers."""
class PdfClassification:
"""Lightweight PDF classification result."""
pdf_type: str
@@ -38,13 +48,31 @@ class TextItem:
page: int
is_bold: bool
is_italic: bool
is_underline: bool
is_strikeout: bool
item_type: str
mcid: Optional[int]
"""Marked Content ID from the content stream's BDC/BMC operator, None when
the text is not part of marked content. Join with the (page, mcid) pairs
from extract_structure_elements to attach structure-tree roles in tagged
PDFs."""
class StructureElement:
"""One structure-tree element reference from a tagged PDF."""
page: int
"""1-indexed page number (matches TextItem.page)."""
mcid: int
"""Marked Content ID from the page's content stream (matches TextItem.mcid)."""
role: str
"""Standard structure type name ("H1".."H6", "P", "Table", "TD", ...)."""
class RegionText:
"""Extracted text for a single region."""
text: str
needs_ocr: bool
"""True when the text should not be trusted."""
ocr_reason: Optional[str]
"""Machine-readable OCR reason when the cause is known."""
class PageRegionTexts:
"""Extracted text for one page's regions."""
@@ -60,6 +88,8 @@ class PageMarkdown:
"""Formatted markdown for this page (empty string when needs_ocr is True)."""
needs_ocr: bool
"""True when text on this page is unreliable and OCR should be used instead."""
ocr_reason: Optional[str]
"""Machine-readable OCR reason when the cause is known."""
class PagesExtractionResult:
"""Per-page markdown output with document-wide layout classification."""
@@ -71,6 +101,8 @@ class PagesExtractionResult:
"""1-indexed pages where multi-column layout was detected."""
pages_needing_ocr: list[int]
"""1-indexed pages that need OCR."""
ocr_reasons_by_page: list[PageOcrReasons]
"""Machine-readable OCR reasons by 1-indexed page."""
is_complex: bool
"""True if any page has tables or multi-column layout."""
@@ -114,6 +146,27 @@ def extract_text_with_positions_bytes(data: bytes, pages: Optional[list[int]] =
"""Extract text with position information from bytes."""
...
def extract_structure_elements(path: str, pages: Optional[list[int]] = None) -> list[StructureElement]:
"""Extract structure-tree element references from a tagged PDF file.
Returns one entry per marked-content reference, resolved to its 1-indexed
page, MCID, and structure type name ("H1".."H6", "P", "Table", ...), sorted
by (page, mcid). Returns an empty list when the PDF is not tagged.
Args:
path: Path to the PDF file.
pages: Optional list of 1-indexed pages (matching ``TextItem.page``).
When ``None`` (default), the whole document is returned.
"""
...
def extract_structure_elements_bytes(data: bytes, pages: Optional[list[int]] = None) -> list[StructureElement]:
"""Extract structure-tree element references from tagged PDF bytes.
See :func:`extract_structure_elements` for details.
"""
...
def extract_text_in_regions(
path: str,
page_regions: list[tuple[int, list[list[float]]]],
+9 -1
View File
@@ -4,8 +4,11 @@ build-backend = "maturin"
[project]
name = "pdf-inspector"
version = "0.1.0"
# Keep package versions in sync with `python3 scripts/version.py <version>`.
# CI publishes automatically when the synchronized change lands on main.
version = "1.14.0"
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
readme = "docs/python.md"
license = { text = "MIT" }
requires-python = ">=3.8"
classifiers = [
@@ -17,5 +20,10 @@ classifiers = [
"Topic :: Text Processing",
]
[project.urls]
Homepage = "https://github.com/firecrawl/pdf-inspector"
Repository = "https://github.com/firecrawl/pdf-inspector"
Documentation = "https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md"
[tool.maturin]
features = ["python"]
+351
View File
@@ -0,0 +1,351 @@
#!/usr/bin/env python3
"""Run a paired pdf-inspector OpenDataLoader benchmark and report deltas."""
from __future__ import annotations
import argparse
import json
import math
import os
import shutil
import subprocess
import sys
import tempfile
from pathlib import Path
from typing import Any
SCORE_KEYS = (
"overall_mean",
"nid_mean",
"nid_s_mean",
"teds_mean",
"teds_s_mean",
"mhs_mean",
"mhs_s_mean",
)
def _non_negative_int(value: str) -> int:
parsed = int(value)
if parsed < 0:
raise argparse.ArgumentTypeError("must be non-negative")
return parsed
def _non_negative_float(value: str) -> float:
parsed = float(value)
if not math.isfinite(parsed) or parsed < 0.0:
raise argparse.ArgumentTypeError("must be finite and non-negative")
return parsed
def _finite_float(value: str) -> float:
parsed = float(value)
if not math.isfinite(parsed):
raise argparse.ArgumentTypeError("must be finite")
return parsed
def _scores(evaluation: dict[str, Any]) -> dict[str, float]:
score = evaluation.get("metrics", {}).get("score", {})
return {key: float(score[key]) for key in SCORE_KEYS if score.get(key) is not None}
def _documents(evaluation: dict[str, Any]) -> dict[str, float]:
documents: dict[str, float] = {}
for document in evaluation.get("documents", []):
overall = document.get("scores", {}).get("overall")
if overall is not None:
documents[str(document["document_id"])] = float(overall)
return documents
def compare_evaluations(
baseline: dict[str, Any],
candidate: dict[str, Any],
reference: dict[str, Any] | None = None,
*,
top: int = 10,
) -> dict[str, Any]:
"""Build aggregate and per-document deltas from evaluator JSON payloads."""
baseline_scores = _scores(baseline)
candidate_scores = _scores(candidate)
metric_deltas = {
key: candidate_scores[key] - baseline_scores[key]
for key in SCORE_KEYS
if key in baseline_scores and key in candidate_scores
}
baseline_documents = _documents(baseline)
candidate_documents = _documents(candidate)
shared = sorted(baseline_documents.keys() & candidate_documents.keys())
document_deltas = [
{
"document_id": document_id,
"baseline": baseline_documents[document_id],
"candidate": candidate_documents[document_id],
"delta": candidate_documents[document_id] - baseline_documents[document_id],
}
for document_id in shared
]
epsilon = 1e-12
improvements = sorted(document_deltas, key=lambda item: item["delta"], reverse=True)
regressions = sorted(document_deltas, key=lambda item: item["delta"])
result: dict[str, Any] = {
"baseline": baseline_scores,
"candidate": candidate_scores,
"deltas": metric_deltas,
"missing_predictions": {
"baseline": int(baseline.get("metrics", {}).get("missing_predictions", 0)),
"candidate": int(candidate.get("metrics", {}).get("missing_predictions", 0)),
},
"documents": {
"shared": len(shared),
"improved": sum(item["delta"] > epsilon for item in document_deltas),
"regressed": sum(item["delta"] < -epsilon for item in document_deltas),
"unchanged": sum(abs(item["delta"]) <= epsilon for item in document_deltas),
"largest_improvements": [
item for item in improvements if item["delta"] > epsilon
][:top],
"largest_regressions": [
item for item in regressions if item["delta"] < -epsilon
][:top],
"worst_regression": next(
(item for item in regressions if item["delta"] < -epsilon), None
),
},
}
if reference is not None:
reference_scores = _scores(reference)
result["reference"] = reference_scores
result["candidate_vs_reference"] = {
key: candidate_scores[key] - reference_scores[key]
for key in SCORE_KEYS
if key in candidate_scores and key in reference_scores
}
return result
def evaluate_gates(
comparison: dict[str, Any],
*,
min_overall_delta: float,
max_document_regression: float | None,
max_missing: int,
require_reference_lead: bool,
) -> list[str]:
"""Return human-readable gate failures; an empty list means pass."""
failures: list[str] = []
overall_delta = comparison["deltas"].get("overall_mean")
if overall_delta is None or overall_delta < min_overall_delta:
failures.append(
f"overall delta {overall_delta!r} is below {min_overall_delta:+.6f}"
)
candidate_missing = comparison["missing_predictions"]["candidate"]
if candidate_missing > max_missing:
failures.append(
f"candidate has {candidate_missing} missing predictions (maximum {max_missing})"
)
if max_document_regression is not None:
regression = comparison["documents"].get("worst_regression")
if regression is not None and regression["delta"] < -max_document_regression:
failures.append(
"largest document regression "
f"{regression['document_id']}={regression['delta']:+.6f} "
f"exceeds {-max_document_regression:+.6f}"
)
if require_reference_lead:
reference_delta = comparison.get("candidate_vs_reference", {}).get("overall_mean")
if reference_delta is None:
failures.append("reference overall score is unavailable")
elif reference_delta < 0.0:
failures.append(
f"candidate trails reference overall by {reference_delta!r}"
)
return failures
def _run(command: list[str], *, cwd: Path, env: dict[str, str] | None = None) -> None:
print("+", " ".join(command), flush=True)
subprocess.run(command, cwd=cwd, env=env, check=True)
def _run_engine(
*,
bench_dir: Path,
python: Path,
binary: Path,
label: str,
scratch_root: Path,
) -> dict[str, Any]:
env = os.environ.copy()
env["PDF_INSPECTOR_BINARY"] = str(binary)
source = bench_dir / "prediction" / "pdf-inspector"
if source.exists():
if source.is_dir():
shutil.rmtree(source)
else:
source.unlink()
_run(
[
str(python),
"src/pdf_parser.py",
"--engine",
"pdf-inspector",
"--log-level",
"WARNING",
],
cwd=bench_dir,
env=env,
)
if not source.is_dir():
raise RuntimeError(f"parser did not produce predictions: {source}")
destination = scratch_root / label
shutil.copytree(source, destination)
_run(
[
str(python),
"src/evaluator.py",
"--prediction-root",
str(scratch_root),
"--engine",
label,
"--log-level",
"WARNING",
],
cwd=bench_dir,
)
with (destination / "evaluation.json").open(encoding="utf-8") as handle:
return json.load(handle)
def _print_report(comparison: dict[str, Any]) -> None:
print("\nMetric baseline candidate delta")
print("-------------------- ---------- ---------- ----------")
for key in SCORE_KEYS:
if key not in comparison["deltas"]:
continue
print(
f"{key:<20} {comparison['baseline'][key]:>10.6f} "
f"{comparison['candidate'][key]:>10.6f} "
f"{comparison['deltas'][key]:>+10.6f}"
)
if "reference" in comparison:
delta = comparison["candidate_vs_reference"].get("overall_mean")
reference = comparison["reference"].get("overall_mean")
reference_display = f"{reference:.6f}" if reference is not None else "n/a"
delta_display = f"{delta:+.6f}" if delta is not None else "n/a"
print(f"\nReference overall: {reference_display}; candidate delta: {delta_display}")
documents = comparison["documents"]
print(
"\nDocuments: "
f"{documents['improved']} improved, {documents['regressed']} regressed, "
f"{documents['unchanged']} unchanged ({documents['shared']} shared)"
)
for heading, key in (
("Largest improvements", "largest_improvements"),
("Largest regressions", "largest_regressions"),
):
print(f"\n{heading}:")
rows = documents[key]
if not rows:
print(" none")
for row in rows:
print(
f" {row['document_id']}: {row['delta']:+.6f} "
f"({row['baseline']:.6f} -> {row['candidate']:.6f})"
)
def _arguments(argv: list[str] | None = None) -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--bench-dir", type=Path, required=True)
parser.add_argument("--baseline", type=Path, required=True)
parser.add_argument("--candidate", type=Path, required=True)
parser.add_argument("--python", type=Path)
parser.add_argument("--reference-evaluation", type=Path)
parser.add_argument("--json-output", type=Path)
parser.add_argument("--top", type=_non_negative_int, default=10)
parser.add_argument("--min-overall-delta", type=_finite_float, default=0.0)
parser.add_argument("--max-document-regression", type=_non_negative_float)
parser.add_argument("--max-missing", type=_non_negative_int, default=0)
parser.add_argument("--require-reference-lead", action="store_true")
return parser.parse_args(argv)
def main(argv: list[str] | None = None) -> int:
args = _arguments(argv)
bench_dir = args.bench_dir.resolve()
baseline = args.baseline.resolve()
candidate = args.candidate.resolve()
# Keep the virtualenv launcher path intact. Resolving its symlink would
# invoke the underlying system interpreter without the benchmark's site
# packages.
python = (args.python or bench_dir / ".venv" / "bin" / "python").absolute()
for path, description in (
(bench_dir / "src" / "pdf_parser.py", "OpenDataLoader parser"),
(bench_dir / "src" / "evaluator.py", "OpenDataLoader evaluator"),
(baseline, "baseline binary"),
(candidate, "candidate binary"),
(python, "Python interpreter"),
):
if not path.exists():
raise SystemExit(f"{description} not found: {path}")
with tempfile.TemporaryDirectory(prefix="pdf-inspector-opendataloader-") as temporary:
scratch_root = Path(temporary)
baseline_evaluation = _run_engine(
bench_dir=bench_dir,
python=python,
binary=baseline,
label="baseline",
scratch_root=scratch_root,
)
candidate_evaluation = _run_engine(
bench_dir=bench_dir,
python=python,
binary=candidate,
label="candidate",
scratch_root=scratch_root,
)
reference = None
if args.reference_evaluation is not None:
with args.reference_evaluation.resolve().open(encoding="utf-8") as handle:
reference = json.load(handle)
comparison = compare_evaluations(
baseline_evaluation,
candidate_evaluation,
reference,
top=args.top,
)
_print_report(comparison)
if args.json_output is not None:
args.json_output.resolve().write_text(
json.dumps(comparison, indent=2) + "\n", encoding="utf-8"
)
failures = evaluate_gates(
comparison,
min_overall_delta=args.min_overall_delta,
max_document_regression=args.max_document_regression,
max_missing=args.max_missing,
require_reference_lead=args.require_reference_lead,
)
if failures:
print("\nBenchmark gate failed:", file=sys.stderr)
for failure in failures:
print(f" - {failure}", file=sys.stderr)
return 1
print("\nBenchmark gate passed.")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+351
View File
@@ -0,0 +1,351 @@
#!/usr/bin/env python3
"""Compare pdf-inspector evidence with optional MuPDF structured text.
This is an experiment and diagnostic tool, not an extraction fallback. It runs
MuPDF's deterministic ``stext.json`` backend without OCR and highlights pages
where that backend exposes materially different text or layout evidence.
"""
from __future__ import annotations
import argparse
from collections import Counter
import json
from pathlib import Path
import re
import shutil
import subprocess
import sys
from typing import Any, Iterable
TOKEN_PATTERN = re.compile(r"[^\W_]+(?:[\u2019'][^\W_]+)*", re.UNICODE)
def _tokens(texts: Iterable[str]) -> Counter[str]:
tokens: Counter[str] = Counter()
for text in texts:
for token in TOKEN_PATTERN.findall(text.casefold()):
# Lone letters are frequently bullets, chart labels, or fragmented
# glyphs. Digits remain useful even when they are one character.
if len(token) > 1 or token.isdigit():
tokens[token] += 1
return tokens
def _repeated_x_anchors(xs: Iterable[float], *, tolerance: float = 4.0) -> int:
buckets = Counter(round(float(x) / tolerance) for x in xs)
return sum(count >= 3 for count in buckets.values())
def local_pages(payload: dict[str, Any]) -> dict[int, dict[str, Any]]:
"""Summarize positioned ``pdf2md --items-json`` evidence by page."""
pages: dict[int, dict[str, Any]] = {}
for item in payload.get("items", []):
page_number = int(item["page"])
page = pages.setdefault(
page_number,
{"texts": [], "xs": [], "text_items": 0, "image_items": 0},
)
if item.get("item_type") == "image":
page["image_items"] += 1
continue
text = str(item.get("text", ""))
if text.strip():
page["texts"].append(text)
page["xs"].append(float(item.get("x", 0.0)))
page["text_items"] += 1
return pages
def alternate_pages(
payload: dict[str, Any] | list[dict[str, Any]],
) -> dict[int, dict[str, Any]]:
"""Summarize MuPDF ``stext.json`` evidence by page."""
pages: dict[int, dict[str, Any]] = {}
raw_pages = payload if isinstance(payload, list) else payload.get("pages", [])
for index, raw_page in enumerate(raw_pages, start=1):
page_number = int(raw_page.get("number", index))
page = {
"texts": [],
"xs": [],
"text_blocks": 0,
"text_lines": 0,
"image_blocks": 0,
}
for block in raw_page.get("blocks", []):
if block.get("type") == "image":
page["image_blocks"] += 1
continue
if block.get("type") != "text":
continue
page["text_blocks"] += 1
for line in block.get("lines", []):
text = str(line.get("text", ""))
if text.strip():
page["texts"].append(text)
bbox = line.get("bbox", {})
page["xs"].append(float(bbox.get("x", line.get("x", 0.0))))
page["text_lines"] += 1
pages[page_number] = page
return pages
def compare_page(
local: dict[str, Any],
alternate: dict[str, Any],
*,
min_token_gain: int,
min_alternate_only_ratio: float,
min_anchor_gain: int,
) -> dict[str, Any]:
"""Compare semantic and coarse layout evidence for one page."""
local_tokens = _tokens(local.get("texts", []))
alternate_tokens = _tokens(alternate.get("texts", []))
shared = local_tokens & alternate_tokens
alternate_only = alternate_tokens - local_tokens
local_only = local_tokens - alternate_tokens
local_total = sum(local_tokens.values())
alternate_total = sum(alternate_tokens.values())
shared_total = sum(shared.values())
alternate_only_total = sum(alternate_only.values())
local_only_total = sum(local_only.values())
net_token_gain = alternate_total - local_total
alternate_only_ratio = alternate_only_total / max(alternate_total, 1)
local_anchors = _repeated_x_anchors(local.get("xs", []))
alternate_anchors = _repeated_x_anchors(alternate.get("xs", []))
anchor_gain = alternate_anchors - local_anchors
image_gain = int(alternate.get("image_blocks", 0)) - int(
local.get("image_items", 0)
)
reasons: list[str] = []
if local_total == 0 and alternate_total >= max(5, min_token_gain // 2):
reasons.append("local_text_empty")
elif (
net_token_gain >= min_token_gain
and alternate_only_ratio >= min_alternate_only_ratio
):
reasons.append("alternate_has_more_text")
if anchor_gain >= min_anchor_gain:
reasons.append("alternate_has_more_alignment_anchors")
if image_gain > 0:
reasons.append("alternate_has_more_image_blocks")
if reasons:
classification = "investigate_alternate_evidence"
elif local_total - alternate_total >= min_token_gain:
classification = "local_has_more_text"
elif alternate_only_total + local_only_total:
classification = "different_segmentation_or_decoding"
else:
classification = "equivalent_text_evidence"
return {
"classification": classification,
"reasons": reasons,
"tokens": {
"local": local_total,
"alternate": alternate_total,
"shared": shared_total,
"net_alternate_gain": net_token_gain,
"alternate_only": alternate_only_total,
"local_only": local_only_total,
"alternate_only_ratio": alternate_only_ratio,
"alternate_only_sample": sorted(alternate_only)[:12],
"local_only_sample": sorted(local_only)[:12],
},
"layout": {
"local_text_items": int(local.get("text_items", 0)),
"local_image_items": int(local.get("image_items", 0)),
"local_repeated_x_anchors": local_anchors,
"alternate_text_blocks": int(alternate.get("text_blocks", 0)),
"alternate_text_lines": int(alternate.get("text_lines", 0)),
"alternate_image_blocks": int(alternate.get("image_blocks", 0)),
"alternate_repeated_x_anchors": alternate_anchors,
},
}
def compare_documents(
local_payload: dict[str, Any],
alternate_payload: dict[str, Any] | list[dict[str, Any]],
*,
min_token_gain: int = 20,
min_alternate_only_ratio: float = 0.15,
min_anchor_gain: int = 2,
) -> dict[str, Any]:
"""Return a page-level evidence report for already extracted payloads."""
local = local_pages(local_payload)
alternate = alternate_pages(alternate_payload)
page_numbers = sorted(local.keys() | alternate.keys())
pages = []
for page_number in page_numbers:
result = compare_page(
local.get(page_number, {}),
alternate.get(page_number, {}),
min_token_gain=min_token_gain,
min_alternate_only_ratio=min_alternate_only_ratio,
min_anchor_gain=min_anchor_gain,
)
result["page"] = page_number
pages.append(result)
flagged = [
page
for page in pages
if page["classification"] == "investigate_alternate_evidence"
]
return {
"summary": {
"pages": len(pages),
"flagged_pages": len(flagged),
"flagged_page_numbers": [page["page"] for page in flagged],
"local_tokens": sum(page["tokens"]["local"] for page in pages),
"alternate_tokens": sum(page["tokens"]["alternate"] for page in pages),
"alternate_only_tokens": sum(
page["tokens"]["alternate_only"] for page in pages
),
},
"pages": pages,
}
def _json_command(command: list[str]) -> Any:
try:
completed = subprocess.run(
command,
check=True,
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
text=True,
)
except subprocess.CalledProcessError as error:
detail = error.stderr.strip() or error.stdout.strip() or "no diagnostic output"
raise RuntimeError(f"command failed: {' '.join(command)}\n{detail}") from error
try:
return json.loads(completed.stdout)
except json.JSONDecodeError as error:
raise RuntimeError(
f"command did not return JSON: {' '.join(command)}: {error}"
) from error
def probe_pdf(
pdf: Path,
*,
pdf2md: Path,
mutool: Path,
min_token_gain: int,
min_alternate_only_ratio: float,
min_anchor_gain: int,
) -> dict[str, Any]:
local_payload = _json_command([str(pdf2md), str(pdf), "--items-json"])
# `stext.json` is MuPDF's native structured text output. The OCR formats
# are intentionally not used so this remains a deterministic no-model
# comparison.
alternate_payload = _json_command(
[str(mutool), "draw", "-q", "-F", "stext.json", "-o", "-", str(pdf)]
)
report = compare_documents(
local_payload,
alternate_payload,
min_token_gain=min_token_gain,
min_alternate_only_ratio=min_alternate_only_ratio,
min_anchor_gain=min_anchor_gain,
)
report["pdf"] = str(pdf)
return report
def _arguments(argv: list[str] | None = None) -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("pdf", type=Path, nargs="+")
parser.add_argument("--pdf2md", type=Path, default=Path("target/release/pdf2md"))
parser.add_argument("--mutool", type=Path)
parser.add_argument("--json-output", type=Path)
parser.add_argument("--min-token-gain", type=int, default=20)
parser.add_argument("--min-alternate-only-ratio", type=float, default=0.15)
parser.add_argument("--min-anchor-gain", type=int, default=2)
return parser.parse_args(argv)
def _print_report(result: dict[str, Any]) -> None:
summary = result["summary"]
print(f"\n{result['pdf']}")
print(
f" {summary['flagged_pages']}/{summary['pages']} pages flagged; "
f"tokens local={summary['local_tokens']} alternate={summary['alternate_tokens']} "
f"alternate-only={summary['alternate_only_tokens']}"
)
for page in result["pages"]:
if page["classification"] != "investigate_alternate_evidence":
continue
reasons = ", ".join(page["reasons"])
tokens = page["tokens"]
print(
f" page {page['page']}: {reasons}; "
f"net tokens={tokens['net_alternate_gain']:+d}, "
f"alternate-only={tokens['alternate_only']}"
)
def main(argv: list[str] | None = None) -> int:
args = _arguments(argv)
pdf2md = args.pdf2md.absolute()
mutool = args.mutool or (Path(found) if (found := shutil.which("mutool")) else None)
if not pdf2md.is_file():
print(f"error: pdf2md binary not found: {pdf2md}", file=sys.stderr)
return 2
if mutool is None or not mutool.is_file():
print("error: mutool not found; install MuPDF or pass --mutool", file=sys.stderr)
return 2
if (
args.min_token_gain < 0
or args.min_anchor_gain < 0
or not 0.0 <= args.min_alternate_only_ratio <= 1.0
):
print("error: thresholds must be non-negative and ratio must be in [0, 1]", file=sys.stderr)
return 2
results = []
for pdf in args.pdf:
path = pdf.absolute()
if not path.is_file():
print(f"error: PDF not found: {path}", file=sys.stderr)
return 2
try:
result = probe_pdf(
path,
pdf2md=pdf2md,
mutool=mutool,
min_token_gain=args.min_token_gain,
min_alternate_only_ratio=args.min_alternate_only_ratio,
min_anchor_gain=args.min_anchor_gain,
)
except RuntimeError as error:
print(f"error: {error}", file=sys.stderr)
return 1
results.append(result)
_print_report(result)
payload = {
"schema_version": 1,
"experiment": "optional_mupdf_stext_evidence",
"ocr": False,
"thresholds": {
"min_token_gain": args.min_token_gain,
"min_alternate_only_ratio": args.min_alternate_only_ratio,
"min_anchor_gain": args.min_anchor_gain,
},
"documents": results,
}
if args.json_output:
args.json_output.parent.mkdir(parents=True, exist_ok=True)
args.json_output.write_text(json.dumps(payload, indent=2) + "\n", encoding="utf-8")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+203
View File
@@ -0,0 +1,203 @@
import io
import json
import sys
import tempfile
import unittest
from contextlib import redirect_stderr, redirect_stdout
from pathlib import Path
from unittest.mock import patch
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from bench_opendataloader import (
_arguments,
_print_report,
_run_engine,
compare_evaluations,
evaluate_gates,
)
def evaluation(overall, documents, *, missing=0):
return {
"metrics": {
"score": {
"overall_mean": overall,
"nid_mean": overall + 0.01,
},
"missing_predictions": missing,
},
"documents": [
{
"document_id": document_id,
"scores": {"overall": score},
}
for document_id, score in documents.items()
],
}
class ComparisonTests(unittest.TestCase):
def test_reports_metric_and_document_deltas(self):
baseline = evaluation(0.80, {"a": 0.8, "b": 0.6, "c": 0.7})
candidate = evaluation(0.82, {"a": 0.9, "b": 0.5, "c": 0.7})
result = compare_evaluations(baseline, candidate, top=1)
self.assertAlmostEqual(result["deltas"]["overall_mean"], 0.02)
self.assertEqual(result["documents"]["improved"], 1)
self.assertEqual(result["documents"]["regressed"], 1)
self.assertEqual(result["documents"]["unchanged"], 1)
self.assertEqual(
result["documents"]["largest_improvements"][0]["document_id"], "a"
)
self.assertEqual(
result["documents"]["largest_regressions"][0]["document_id"], "b"
)
def test_reference_delta_is_reported(self):
baseline = evaluation(0.80, {})
candidate = evaluation(0.82, {})
reference = evaluation(0.81, {})
result = compare_evaluations(baseline, candidate, reference)
self.assertAlmostEqual(
result["candidate_vs_reference"]["overall_mean"], 0.01
)
def test_gates_cover_aggregate_document_missing_and_reference(self):
comparison = compare_evaluations(
evaluation(0.80, {"a": 0.8}),
evaluation(0.79, {"a": 0.7}, missing=1),
evaluation(0.81, {}),
)
failures = evaluate_gates(
comparison,
min_overall_delta=0.0,
max_document_regression=0.05,
max_missing=0,
require_reference_lead=True,
)
self.assertEqual(len(failures), 4)
def test_regression_gate_is_independent_of_report_limit(self):
comparison = compare_evaluations(
evaluation(0.80, {"a": 0.8}),
evaluation(0.80, {"a": 0.7}),
top=0,
)
failures = evaluate_gates(
comparison,
min_overall_delta=0.0,
max_document_regression=0.05,
max_missing=0,
require_reference_lead=False,
)
self.assertEqual(len(failures), 1)
self.assertIn("largest document regression", failures[0])
def test_report_handles_reference_without_overall_score(self):
result = compare_evaluations(
evaluation(0.80, {}),
evaluation(0.82, {}),
{"metrics": {"score": {"nid_mean": 0.81}}},
)
output = io.StringIO()
with redirect_stdout(output):
_print_report(result)
self.assertIn("Reference overall: n/a; candidate delta: n/a", output.getvalue())
def test_reference_gate_reports_missing_score_as_unavailable(self):
comparison = compare_evaluations(
evaluation(0.80, {}),
evaluation(0.82, {}),
)
failures = evaluate_gates(
comparison,
min_overall_delta=0.0,
max_document_regression=None,
max_missing=0,
require_reference_lead=True,
)
self.assertEqual(failures, ["reference overall score is unavailable"])
def test_arguments_reject_negative_counts_and_allow_zero_top(self):
required = [
"--bench-dir",
".",
"--baseline",
"baseline",
"--candidate",
"candidate",
]
self.assertEqual(_arguments(required + ["--top", "0"]).top, 0)
for option in ("--top", "--max-document-regression", "--max-missing"):
with self.subTest(option=option), redirect_stderr(io.StringIO()):
with self.assertRaises(SystemExit):
_arguments(required + [option, "-1"])
def test_arguments_reject_nonfinite_float_thresholds(self):
required = [
"--bench-dir",
".",
"--baseline",
"baseline",
"--candidate",
"candidate",
]
for option in ("--min-overall-delta", "--max-document-regression"):
for value in ("nan", "inf", "-inf"):
with self.subTest(option=option, value=value), redirect_stderr(
io.StringIO()
):
with self.assertRaises(SystemExit):
_arguments(required + [option, value])
def test_run_engine_clears_stale_predictions_before_parser(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
bench_dir = root / "bench"
source = bench_dir / "prediction" / "pdf-inspector"
source.mkdir(parents=True)
(source / "stale.md").write_text("stale", encoding="utf-8")
scratch = root / "scratch"
scratch.mkdir()
def fake_run(command, *, cwd, env=None):
if any(part.endswith("pdf_parser.py") for part in command):
self.assertFalse(source.exists())
(source / "markdown").mkdir(parents=True)
(source / "markdown" / "new.md").write_text(
"new", encoding="utf-8"
)
else:
destination = scratch / "candidate"
(destination / "evaluation.json").write_text(
json.dumps(evaluation(0.82, {})), encoding="utf-8"
)
with patch("bench_opendataloader._run", side_effect=fake_run):
result = _run_engine(
bench_dir=bench_dir,
python=Path("python"),
binary=Path("pdf2md"),
label="candidate",
scratch_root=scratch,
)
self.assertEqual(result["metrics"]["score"]["overall_mean"], 0.82)
self.assertFalse((source / "stale.md").exists())
self.assertFalse((scratch / "candidate" / "stale.md").exists())
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,103 @@
import sys
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from probe_backend_evidence import compare_documents
def local_payload(items):
return {"items": items}
def item(page, text, x=10, item_type="text"):
return {"page": page, "text": text, "x": x, "item_type": item_type}
def alternate_payload(pages):
return {"pages": pages}
def page(lines, *, images=0):
blocks = [
{
"type": "text",
"lines": [
{"text": text, "bbox": {"x": x, "y": index * 10, "w": 80, "h": 8}}
for index, (text, x) in enumerate(lines)
],
}
]
blocks.extend({"type": "image"} for _ in range(images))
return {"blocks": blocks}
class EvidenceComparisonTests(unittest.TestCase):
def test_accepts_real_top_level_page_array(self):
local = local_payload([item(1, "alpha beta")])
alternate = [page([("alpha beta gamma", 10)])]
result = compare_documents(local, alternate)["pages"][0]
self.assertEqual(result["tokens"]["alternate"], 3)
self.assertEqual(result["tokens"]["net_alternate_gain"], 1)
def test_flags_material_alternate_text_gain(self):
local = local_payload([item(1, "alpha beta")])
alternate = alternate_payload(
[page([("alpha beta gamma delta epsilon zeta", 10)])]
)
report = compare_documents(
local,
alternate,
min_token_gain=3,
min_alternate_only_ratio=0.2,
)
result = report["pages"][0]
self.assertEqual(result["classification"], "investigate_alternate_evidence")
self.assertIn("alternate_has_more_text", result["reasons"])
self.assertEqual(result["tokens"]["net_alternate_gain"], 4)
def test_repeated_alignment_and_image_evidence_are_reported(self):
local = local_payload([item(1, "one two", 10)])
alternate = alternate_payload(
[
page(
[
("one two", 10),
("row three", 100),
("row four", 100),
("row five", 100),
],
images=1,
)
]
)
result = compare_documents(
local,
alternate,
min_token_gain=99,
min_anchor_gain=1,
)["pages"][0]
self.assertIn("alternate_has_more_alignment_anchors", result["reasons"])
self.assertIn("alternate_has_more_image_blocks", result["reasons"])
self.assertEqual(result["layout"]["alternate_repeated_x_anchors"], 1)
def test_token_segmentation_difference_does_not_imply_more_evidence(self):
local = local_payload([item(1, "Revenue 2025")])
alternate = alternate_payload([page([("Revenue 2024", 10)])])
result = compare_documents(local, alternate, min_token_gain=2)["pages"][0]
self.assertEqual(result["classification"], "different_segmentation_or_decoding")
self.assertEqual(result["reasons"], [])
self.assertEqual(result["tokens"]["alternate_only_sample"], ["2024"])
if __name__ == "__main__":
unittest.main()
+111
View File
@@ -0,0 +1,111 @@
import json
import sys
import tempfile
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from version import PLATFORM_PACKAGES, check_versions, set_versions
class VersionTests(unittest.TestCase):
def setUp(self):
self.temporary = tempfile.TemporaryDirectory()
self.root = Path(self.temporary.name)
(self.root / "napi").mkdir()
(self.root / "site").mkdir()
(self.root / "wasm").mkdir()
self._write_manifest("Cargo.toml", "package", "0.1.0")
self._write_manifest("pyproject.toml", "project", "0.1.0")
self._write_manifest("napi/Cargo.toml", "package", "0.1.0")
self._write_manifest("wasm/Cargo.toml", "package", "0.1.0")
package = {
"name": "@firecrawl/pdf-inspector",
"version": "0.1.0",
"optionalDependencies": {
dependency: "0.1.0" for dependency in PLATFORM_PACKAGES
},
}
(self.root / "napi/package.json").write_text(
json.dumps(package), encoding="utf-8"
)
(self.root / "napi/bun.lock").write_text(
"\n".join(
f' "{dependency}": "0.1.0",'
for dependency in PLATFORM_PACKAGES
)
+ "\n",
encoding="utf-8",
)
(self.root / "site/index.html").write_text(
'https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.0/'
'pdf_inspector_wasm.js\n',
encoding="utf-8",
)
self._write_lock(
"napi/Cargo.lock", ("pdf-inspector", "pdf-inspector-napi")
)
self._write_lock(
"wasm/Cargo.lock", ("pdf-inspector", "pdf-inspector-wasm")
)
def tearDown(self):
self.temporary.cleanup()
def _write_manifest(self, relative, section, version):
(self.root / relative).write_text(
f'[{section}]\nname = "fixture"\nversion = "{version}"\n',
encoding="utf-8",
)
def _write_lock(self, relative, packages):
content = "\n".join(
f'[[package]]\nname = "{package}"\nversion = "0.1.0"\n'
for package in packages
)
(self.root / relative).write_text(content, encoding="utf-8")
def test_updates_every_version_location(self):
set_versions("1.14.0", self.root)
self.assertEqual(check_versions(self.root), "1.14.0")
def test_reports_a_divergent_package(self):
self._write_manifest("wasm/Cargo.toml", "package", "0.2.0")
with self.assertRaisesRegex(ValueError, "WASM package: 0.2.0"):
check_versions(self.root)
def test_rejects_an_invalid_version(self):
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
set_versions("next", self.root)
def test_rejects_numeric_prerelease_with_leading_zero(self):
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
set_versions("1.2.3-01", self.root)
self.assertEqual(
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
)
def test_preflight_failure_does_not_partially_update(self):
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
(self.root / "site/index.html").write_text(
"missing module URL\n", encoding="utf-8"
)
with self.assertRaisesRegex(ValueError, "Missing pinned WASM package URL"):
set_versions("1.14.0", self.root)
self.assertEqual(
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
)
if __name__ == "__main__":
unittest.main()
+262
View File
@@ -0,0 +1,262 @@
#!/usr/bin/env python3
"""Keep every pdf-inspector package on one release version."""
from __future__ import annotations
import argparse
import json
import re
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
PRERELEASE_IDENTIFIER = (
r"(?:0|[1-9]\d*|[0-9A-Za-z-]*[A-Za-z-][0-9A-Za-z-]*)"
)
SEMVER = re.compile(
r"^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)"
rf"(?:-{PRERELEASE_IDENTIFIER}(?:\.{PRERELEASE_IDENTIFIER})*)?"
r"(?:\+[0-9A-Za-z-]+(?:\.[0-9A-Za-z-]+)*)?$"
)
VERSION_LINE = re.compile(r'^(\s*version\s*=\s*")[^"]+(".*)$')
SECTION_LINE = re.compile(r"^\s*\[([^]]+)]\s*$")
PLATFORM_PACKAGES = (
"@firecrawl/pdf-inspector-linux-x64-gnu",
"@firecrawl/pdf-inspector-linux-x64-musl",
"@firecrawl/pdf-inspector-linux-arm64-gnu",
"@firecrawl/pdf-inspector-linux-arm64-musl",
"@firecrawl/pdf-inspector-darwin-arm64",
"@firecrawl/pdf-inspector-win32-x64-msvc",
)
TOML_VERSIONS = (
("Rust crate", Path("Cargo.toml"), "package"),
("Python package", Path("pyproject.toml"), "project"),
("NAPI crate", Path("napi/Cargo.toml"), "package"),
("WASM package", Path("wasm/Cargo.toml"), "package"),
)
LOCK_VERSIONS = (
("NAPI lock: core", Path("napi/Cargo.lock"), "pdf-inspector"),
("NAPI lock: binding", Path("napi/Cargo.lock"), "pdf-inspector-napi"),
("WASM lock: core", Path("wasm/Cargo.lock"), "pdf-inspector"),
("WASM lock: binding", Path("wasm/Cargo.lock"), "pdf-inspector-wasm"),
)
SITE_WASM_VERSION = re.compile(
r"(@firecrawl/pdf-inspector-wasm@)([^/\"]+)(/pdf_inspector_wasm\.js)"
)
def _read_section_version(path: Path, section: str) -> str:
active = False
for line in path.read_text(encoding="utf-8").splitlines():
section_match = SECTION_LINE.match(line)
if section_match:
active = section_match.group(1) == section
elif active:
version_match = VERSION_LINE.match(line)
if version_match:
return line.split('"', 2)[1]
raise ValueError(f"No version found in [{section}] of {path}")
def _write_section_version(path: Path, section: str, version: str) -> None:
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
active = False
for index, line in enumerate(lines):
section_match = SECTION_LINE.match(line)
if section_match:
active = section_match.group(1) == section
elif active:
version_match = VERSION_LINE.match(line)
if version_match:
newline = "\n" if line.endswith("\n") else ""
replacement = (
f"{version_match.group(1)}{version}"
f"{version_match.group(2).rstrip()}"
)
lines[index] = (
f"{replacement}{newline}"
)
path.write_text("".join(lines), encoding="utf-8")
return
raise ValueError(f"No version found in [{section}] of {path}")
def _package_block(lines: list[str], package: str) -> tuple[int, int]:
for start, line in enumerate(lines):
if line.strip() != "[[package]]":
continue
end = next(
(
index
for index in range(start + 1, len(lines))
if lines[index].strip() == "[[package]]"
),
len(lines),
)
if any(line.strip() == f'name = "{package}"' for line in lines[start:end]):
return start, end
raise ValueError(f"No lockfile entry found for {package}")
def _read_lock_version(path: Path, package: str) -> str:
lines = path.read_text(encoding="utf-8").splitlines()
start, end = _package_block(lines, package)
for line in lines[start:end]:
version_match = VERSION_LINE.match(line)
if version_match:
return line.split('"', 2)[1]
raise ValueError(f"No version found for {package} in {path}")
def _write_lock_version(path: Path, package: str, version: str) -> None:
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
start, end = _package_block(lines, package)
for index in range(start, end):
version_match = VERSION_LINE.match(lines[index])
if version_match:
newline = "\n" if lines[index].endswith("\n") else ""
lines[index] = (
f'{version_match.group(1)}{version}{version_match.group(2).rstrip()}'
f"{newline}"
)
path.write_text("".join(lines), encoding="utf-8")
return
raise ValueError(f"No version found for {package} in {path}")
def _node_versions(root: Path) -> dict[str, str]:
package = json.loads((root / "napi/package.json").read_text(encoding="utf-8"))
versions = {"Node package": package["version"]}
optional = package.get("optionalDependencies", {})
for dependency in PLATFORM_PACKAGES:
if dependency not in optional:
raise ValueError(f"Missing Node optional dependency: {dependency}")
versions[f"Node optional dependency: {dependency}"] = optional[dependency]
return versions
def _bun_versions(root: Path) -> dict[str, str]:
text = (root / "napi/bun.lock").read_text(encoding="utf-8")
versions = {}
for dependency in PLATFORM_PACKAGES:
match = re.search(
rf'"{re.escape(dependency)}": "([^"]+)"[,]', text
)
if not match:
raise ValueError(f"Missing Bun lock dependency: {dependency}")
versions[f"Bun lock: {dependency}"] = match.group(1)
return versions
def _site_wasm_version(root: Path) -> str:
text = (root / "site/index.html").read_text(encoding="utf-8")
match = SITE_WASM_VERSION.search(text)
if not match:
raise ValueError("Missing pinned WASM package URL in site/index.html")
return match.group(2)
def package_versions(root: Path = ROOT) -> dict[str, str]:
versions = {
label: _read_section_version(root / relative, section)
for label, relative, section in TOML_VERSIONS
}
versions.update(_node_versions(root))
versions.update(_bun_versions(root))
versions["Website WASM module"] = _site_wasm_version(root)
versions.update(
{
label: _read_lock_version(root / relative, package)
for label, relative, package in LOCK_VERSIONS
}
)
return versions
def check_versions(root: Path = ROOT) -> str:
versions = package_versions(root)
expected = versions["Rust crate"]
if not SEMVER.fullmatch(expected):
raise ValueError(f"Rust crate has an invalid semantic version: {expected}")
mismatches = {
label: version for label, version in versions.items() if version != expected
}
if mismatches:
details = "\n".join(
f" - {label}: {version}" for label, version in mismatches.items()
)
raise ValueError(f"Expected every package to use {expected}:\n{details}")
return expected
def set_versions(version: str, root: Path = ROOT) -> None:
if not SEMVER.fullmatch(version):
raise ValueError(f"Invalid semantic version: {version}")
# Validate every expected location before writing the first file. This
# prevents a stale manifest or generated file from leaving a partial bump.
package_versions(root)
for _, relative, section in TOML_VERSIONS:
_write_section_version(root / relative, section, version)
package_path = root / "napi/package.json"
package = json.loads(package_path.read_text(encoding="utf-8"))
package["version"] = version
optional = package.get("optionalDependencies", {})
for dependency in PLATFORM_PACKAGES:
if dependency not in optional:
raise ValueError(f"Missing Node optional dependency: {dependency}")
optional[dependency] = version
package_path.write_text(json.dumps(package, indent=2) + "\n", encoding="utf-8")
bun_path = root / "napi/bun.lock"
bun_text = bun_path.read_text(encoding="utf-8")
for dependency in PLATFORM_PACKAGES:
pattern = rf'("{re.escape(dependency)}": ")[^"]+("[,])'
bun_text, count = re.subn(
pattern, rf"\g<1>{version}\g<2>", bun_text, count=1
)
if count != 1:
raise ValueError(f"Missing Bun lock dependency: {dependency}")
bun_path.write_text(bun_text, encoding="utf-8")
site_path = root / "site/index.html"
site_text = site_path.read_text(encoding="utf-8")
site_text, count = SITE_WASM_VERSION.subn(
rf"\g<1>{version}\g<3>", site_text, count=1
)
if count != 1:
raise ValueError("Missing pinned WASM package URL in site/index.html")
site_path.write_text(site_text, encoding="utf-8")
for _, relative, package_name in LOCK_VERSIONS:
_write_lock_version(root / relative, package_name, version)
check_versions(root)
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("version", nargs="?", help="new shared semantic version")
parser.add_argument(
"--check", action="store_true", help="fail if package versions have diverged"
)
arguments = parser.parse_args()
if arguments.check == bool(arguments.version):
parser.error("provide either a version or --check")
try:
if arguments.check:
version = check_versions()
print(f"All packages use {version}")
else:
set_versions(arguments.version)
print(f"Updated all packages to {arguments.version}")
except ValueError as error:
parser.exit(1, f"{error}\n")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+3
View File
@@ -0,0 +1,3 @@
<svg width="200" height="284" viewBox="0 0 200 284" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M166.862 90.7716C155.812 94.0514 147.483 101.471 141.383 109.53C140.073 111.26 137.343 109.96 137.863 107.841C149.543 59.8136 134.113 19.896 86.0157 0.247269C83.5758 -0.752669 81.0359 1.43719 81.6759 3.99704C103.555 91.8416 11.5294 84.432 23.1588 184.016C23.3588 185.726 21.4389 186.896 20.039 185.896C15.6792 182.766 10.8095 176.236 7.46963 171.647C6.48968 170.297 4.36978 170.677 3.9198 172.287C1.25994 181.906 0 190.965 0 199.965C0 234.963 17.9891 265.771 45.2177 283.63C46.7777 284.65 48.7776 283.19 48.2476 281.4C46.8477 276.7 46.0577 271.74 45.9977 266.611C45.9977 263.461 46.1977 260.241 46.6877 257.241C47.8276 249.702 50.4475 242.522 54.8473 235.983C69.9365 213.334 100.185 191.455 95.3552 161.747C95.0453 159.867 97.2651 158.627 98.6651 159.917C119.974 179.386 124.194 205.575 120.694 229.063C120.394 231.103 122.954 232.193 124.244 230.593C127.504 226.513 131.483 222.933 135.813 220.244C136.893 219.574 138.333 220.084 138.743 221.284C141.153 228.293 144.733 234.873 148.113 241.452C152.152 249.362 154.302 258.391 153.962 267.951C153.792 272.6 153.022 277.1 151.732 281.38C151.182 283.19 153.162 284.7 154.752 283.66C182.001 265.801 200 234.993 200 199.975C200 187.806 197.87 175.876 193.84 164.697C185.391 141.248 163.952 123.64 169.372 93.0815C169.632 91.6216 168.282 90.3517 166.862 90.7716Z" fill="#FA5D19" style="fill:#FA5D19;fill:color(display-p3 0.9816 0.3634 0.0984);fill-opacity:1;"/>
</svg>

After

Width:  |  Height:  |  Size: 1.5 KiB

+12
View File
@@ -0,0 +1,12 @@
<svg width="172" height="40" viewBox="0 0 172 40" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M23.3606 12.8281C21.8137 13.2873 20.6476 14.3261 19.7936 15.4544C19.6102 15.6966 19.228 15.5146 19.3008 15.2178C20.936 8.49401 18.7759 2.90556 12.0422 0.154735C11.7006 0.0147436 11.345 0.321324 11.4346 0.679702C14.4977 12.9779 1.61412 11.9406 3.24224 25.8823C3.27024 26.1217 3.00145 26.2855 2.80546 26.1455C2.19509 25.7073 1.51332 24.7932 1.04575 24.1506C0.908555 23.9616 0.611769 24.0148 0.548773 24.2402C0.176391 25.5869 0 26.8553 0 28.1152C0 33.0149 2.51847 37.328 6.33048 39.8283C6.54887 39.9711 6.82886 39.7667 6.75466 39.5161C6.55867 38.8581 6.44808 38.1638 6.43968 37.4456C6.43968 37.0046 6.46768 36.5539 6.53627 36.1339C6.69587 35.0784 7.06265 34.0732 7.67862 33.1577C9.79111 29.9869 14.0259 26.9239 13.3497 22.7647C13.3063 22.5015 13.6171 22.328 13.8131 22.5085C16.7964 25.2342 17.3871 28.9005 16.8972 32.1889C16.8552 32.4745 17.2135 32.6271 17.3941 32.4031C17.8505 31.832 18.4077 31.3308 19.0138 30.9542C19.165 30.8604 19.3666 30.9318 19.424 31.0998C19.7614 32.0811 20.2626 33.0023 20.7358 33.9234C21.3013 35.0308 21.6023 36.2949 21.5547 37.6332C21.5309 38.2842 21.4231 38.9141 21.2425 39.5133C21.1655 39.7667 21.4427 39.9781 21.6653 39.8325C25.4801 37.3322 28 33.0191 28 28.1166C28 26.4129 27.7018 24.7428 27.1376 23.1777C25.9547 19.8949 22.9533 17.4297 23.712 13.1515C23.7484 12.9471 23.5594 12.7693 23.3606 12.8281Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M41 34.0521V10.9618H55.7586V14.3264H44.7969V21.0226H53.8436V24.2882H44.7969V34.0521H41Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M59.9569 14.7882C58.7352 14.7882 57.7777 13.8976 57.7777 12.6441C57.7777 11.3906 58.7352 10.5 59.9569 10.5C61.1785 10.5 62.136 11.3906 62.136 12.6441C62.136 13.8976 61.1785 14.7882 59.9569 14.7882ZM58.1409 34.0521V17.1632H61.7068V34.0521H58.1409Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M73.5885 17.1632H74.3809V20.4948H72.796C69.6264 20.4948 68.6029 22.9687 68.6029 25.5747V34.0521H65.0371V17.1632H68.2067L68.6029 19.7031C69.4613 18.2847 70.815 17.1632 73.5885 17.1632Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M83.632 34.25C78.3163 34.25 74.9816 30.8194 74.9816 25.6406C74.9816 20.4288 78.3163 16.9653 83.3019 16.9653C88.1884 16.9653 91.457 20.066 91.5561 25.0139C91.5561 25.4427 91.5231 25.9045 91.457 26.3663H78.7125V26.5972C78.8116 29.467 80.6275 31.3472 83.4339 31.3472C85.613 31.3472 87.1979 30.2587 87.6931 28.3785H91.2589C90.6646 31.7101 87.8252 34.25 83.632 34.25ZM78.8446 23.7604H87.8582C87.561 21.2535 85.8112 19.8351 83.3349 19.8351C81.0567 19.8351 79.1087 21.3524 78.8446 23.7604Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M102.033 34.25C96.9151 34.25 93.6465 30.9184 93.6465 25.6406C93.6465 20.4288 97.0142 16.9653 102.132 16.9653C106.49 16.9653 109.197 19.3733 109.891 23.1997H106.16C105.698 21.2205 104.278 20 102.066 20C99.1933 20 97.3113 22.309 97.3113 25.6406C97.3113 28.9392 99.1933 31.2153 102.066 31.2153C104.245 31.2153 105.698 29.9618 106.127 28.0156H109.891C109.23 31.842 106.358 34.25 102.033 34.25Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M121.006 17.1632H121.799V20.4948H120.214C117.044 20.4948 116.021 22.9687 116.021 25.5747V34.0521H112.455V17.1632H115.625L116.021 19.7031C116.879 18.2847 118.233 17.1632 121.006 17.1632Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M130.614 16.9653C135.104 16.9653 137.679 19.1094 137.679 23.1007V34.0521H134.576L134.279 31.6441C133.123 33.1615 131.505 34.25 128.831 34.25C125.133 34.25 122.657 32.4358 122.657 29.3021C122.657 25.8385 125.166 23.8924 129.92 23.8924H134.147V22.8698C134.147 20.9896 132.793 19.8351 130.449 19.8351C128.336 19.8351 126.916 20.8247 126.652 22.309H123.152C123.515 19.0104 126.355 16.9653 130.614 16.9653ZM129.425 31.4792C132.397 31.4792 134.114 29.7309 134.147 27.125V26.5312H129.722C127.51 26.5312 126.289 27.3559 126.289 29.0712C126.289 30.4896 127.477 31.4792 129.425 31.4792Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M144.653 34.0521L139.139 17.1632H142.903L146.766 30.0937L150.629 17.1632H153.897L157.595 30.0937L161.59 17.1632H165.222L159.609 34.0521H155.779L152.214 22.5729L148.516 34.0521H144.653Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M166.934 34.0521V10.9618H170.5V34.0521H166.934Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
</svg>

After

Width:  |  Height:  |  Size: 4.8 KiB

+1172
View File
File diff suppressed because it is too large Load Diff
+44 -2
View File
@@ -11,6 +11,37 @@ use std::process;
use std::time::Instant;
/// Escape a string for embedding in a JSON string value.
fn format_detector_ocr_reasons(reasons: &std::collections::BTreeMap<u32, Vec<String>>) -> String {
reasons
.iter()
.map(|(page, page_reasons)| {
let reasons_json = page_reasons
.iter()
.map(|reason| format!(r#""{}""#, json_escape(reason)))
.collect::<Vec<_>>()
.join(",");
format!(r#"{{"page":{},"reasons":[{}]}}"#, page, reasons_json)
})
.collect::<Vec<_>>()
.join(",")
}
fn format_ocr_reasons_by_page(reasons: &[pdf_inspector::PageOcrReasons]) -> String {
reasons
.iter()
.map(|entry| {
let reasons_json = entry
.reasons
.iter()
.map(|reason| format!(r#""{}""#, json_escape(reason)))
.collect::<Vec<_>>()
.join(",");
format!(r#"{{"page":{},"reasons":[{}]}}"#, entry.page, reasons_json)
})
.collect::<Vec<_>>()
.join(",")
}
fn json_escape(s: &str) -> String {
let mut out = String::with_capacity(s.len() + 16);
for ch in s.chars() {
@@ -32,6 +63,7 @@ fn json_escape(s: &str) -> String {
}
fn main() {
#[cfg(not(target_arch = "wasm32"))]
env_logger::init();
let args: Vec<String> = env::args().collect();
@@ -117,11 +149,13 @@ fn run_analyze(pdf_path: &str, json_output: bool, start: Instant) {
.iter()
.map(|p| p.to_string())
.collect();
let ocr_reasons = format_ocr_reasons_by_page(&result.ocr_reasons_by_page);
println!(
r#"{{"pdf_type":"{}","page_count":{},"pages_needing_ocr":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"detection_time_ms":{}}}"#,
r#"{{"pdf_type":"{}","page_count":{},"pages_needing_ocr":[{}],"ocr_reasons_by_page":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"detection_time_ms":{}}}"#,
pdf_type_str(&result.pdf_type),
result.page_count,
ocr_pages.join(","),
ocr_reasons,
result.layout.is_complex,
table_pages.join(","),
col_pages.join(","),
@@ -144,6 +178,9 @@ fn run_analyze(pdf_path: &str, json_output: bool, start: Instant) {
println!("Page count: {}", result.page_count);
if !result.pages_needing_ocr.is_empty() {
println!("Pages needing OCR: {:?}", result.pages_needing_ocr);
for entry in &result.ocr_reasons_by_page {
println!(" page {}: {}", entry.page, entry.reasons.join(", "));
}
}
println!();
if result.layout.is_complex {
@@ -184,8 +221,9 @@ fn run_detect_only(pdf_path: &str, json_output: bool, start: Instant) {
.iter()
.map(|p| p.to_string())
.collect();
let ocr_reasons = format_detector_ocr_reasons(&result.ocr_reasons_by_page);
println!(
r#"{{"pdf_type":"{}","page_count":{},"pages_sampled":{},"pages_with_text":{},"confidence":{:.2},"title":{},"ocr_recommended":{},"pages_needing_ocr":[{}],"detection_time_ms":{}}}"#,
r#"{{"pdf_type":"{}","page_count":{},"pages_sampled":{},"pages_with_text":{},"confidence":{:.2},"title":{},"ocr_recommended":{},"pages_needing_ocr":[{}],"ocr_reasons_by_page":[{}],"detection_time_ms":{}}}"#,
pdf_type_str(&result.pdf_type),
result.page_count,
result.pages_sampled,
@@ -198,6 +236,7 @@ fn run_detect_only(pdf_path: &str, json_output: bool, start: Instant) {
.unwrap_or_else(|| "null".to_string()),
result.ocr_recommended,
ocr_pages.join(","),
ocr_reasons,
elapsed.as_millis()
);
} else {
@@ -232,6 +271,9 @@ fn run_detect_only(pdf_path: &str, json_output: bool, start: Instant) {
result.pages_needing_ocr, result.page_count
);
}
for (page, reasons) in &result.ocr_reasons_by_page {
println!(" page {}: {}", page, reasons.join(", "));
}
}
if let Some(title) = &result.title {
println!("Title: {}", title);
+176 -3
View File
@@ -1,6 +1,10 @@
//! CLI tool for PDF to Markdown conversion
use pdf_inspector::{process_pdf_with_options, LayoutComplexity, PdfOptions, PdfType, ProcessMode};
use pdf_inspector::extractor::ItemType;
use pdf_inspector::{
extract_text_with_positions_pages_with_password, process_pdf_with_options, LayoutComplexity,
PdfOptions, PdfType, ProcessMode, TextItem,
};
use std::collections::HashSet;
use std::env;
use std::fmt::Write;
@@ -31,6 +35,137 @@ fn json_escape(s: &str) -> String {
out
}
fn format_ocr_reasons_by_page(reasons: &[pdf_inspector::PageOcrReasons]) -> String {
reasons
.iter()
.map(|entry| {
let reasons_json = entry
.reasons
.iter()
.map(|reason| format!(r#""{}""#, json_escape(reason)))
.collect::<Vec<_>>()
.join(",");
format!(r#"{{"page":{},"reasons":[{}]}}"#, entry.page, reasons_json)
})
.collect::<Vec<_>>()
.join(",")
}
fn item_type_label(item_type: &ItemType) -> &'static str {
match item_type {
ItemType::Text => "text",
ItemType::Image => "image",
ItemType::Link(_) => "link",
ItemType::FormField => "form_field",
}
}
fn format_items_json(items: &[TextItem]) -> String {
let underlined_count = items.iter().filter(|item| item.is_underline).count();
let items_json = items
.iter()
.map(|item| {
let mcid = item
.mcid
.map(|value| value.to_string())
.unwrap_or_else(|| "null".to_string());
let link_url = match &item.item_type {
ItemType::Link(url) => format!(r#","url":"{}""#, json_escape(url)),
_ => String::new(),
};
format!(
r#"{{"text":"{}","page":{},"x":{:.2},"y":{:.2},"width":{:.2},"height":{:.2},"font":"{}","font_size":{:.2},"is_bold":{},"is_italic":{},"is_underline":{},"is_strikeout":{},"item_type":"{}","mcid":{}{}}}"#,
json_escape(&item.text),
item.page,
item.x,
item.y,
item.width,
item.height,
json_escape(&item.font),
item.font_size,
item.is_bold,
item.is_italic,
item.is_underline,
item.is_strikeout,
item_type_label(&item.item_type),
mcid,
link_url,
)
})
.collect::<Vec<_>>()
.join(",");
format!(
r#"{{"total_items":{},"underlined_count":{},"items":[{}]}}"#,
items.len(),
underlined_count,
items_json
)
}
fn extract_items_json(
pdf_path: &str,
page_filter: Option<&HashSet<u32>>,
password: Option<&str>,
) -> Result<String, pdf_inspector::PdfError> {
extract_text_with_positions_pages_with_password(pdf_path, page_filter, password)
.map(|items| format_items_json(&items))
}
#[cfg(test)]
mod tests {
use super::{extract_items_json, format_items_json};
use pdf_inspector::extractor::ItemType;
use pdf_inspector::TextItem;
#[test]
fn items_json_includes_position_and_underline_metadata() {
let items = vec![TextItem {
text: "A \"quoted\" item".to_string(),
x: 12.345,
y: 67.891,
width: 23.456,
height: 9.876,
font: "F1".to_string(),
font_size: 10.0,
page: 2,
is_bold: false,
is_italic: true,
is_underline: true,
is_strikeout: true,
item_type: ItemType::Text,
mcid: Some(7),
}];
let json = format_items_json(&items);
assert!(json.contains(r#""text":"A \"quoted\" item""#));
assert!(json.contains(r#""page":2"#));
assert!(json.contains(r#""x":12.35"#));
assert!(json.contains(r#""is_underline":true"#));
assert!(json.contains(r#""item_type":"text""#));
assert!(json.contains(r#""mcid":7"#));
}
#[test]
fn items_json_uses_supplied_pdf_password() {
let path = "tests/fixtures/encrypted-secret123.pdf";
let without_password = extract_items_json(path, None, None);
assert!(
without_password.is_err(),
"encrypted fixture unexpectedly extracted without a password"
);
let json = extract_items_json(path, None, Some("secret123"))
.expect("correct password should decrypt positioned text");
assert!(
json.contains("Procurement"),
"decrypted item JSON should contain fixture text, got {json}"
);
}
}
/// Parse a page specification like "1,3,5-10,20" into a HashSet of page numbers.
fn parse_page_spec(spec: &str) -> Result<HashSet<u32>, String> {
let mut pages = HashSet::new();
@@ -82,12 +217,14 @@ fn print_layout_info(layout: &LayoutComplexity) {
}
fn main() {
#[cfg(not(target_arch = "wasm32"))]
env_logger::init();
let args: Vec<String> = env::args().collect();
if args.len() < 2 {
eprintln!("Usage: {} <pdf_file> [output_file]", args[0]);
eprintln!(" {} <pdf_file> --json", args[0]);
eprintln!(" {} <pdf_file> --items-json", args[0]);
eprintln!(" {} <pdf_file> --raw", args[0]);
eprintln!();
eprintln!("Converts PDF to Markdown with smart type detection.");
@@ -95,9 +232,14 @@ fn main() {
eprintln!();
eprintln!("Options:");
eprintln!(" --json Output result as JSON");
eprintln!(" --items-json Output positioned TextItem JSON");
eprintln!(" --raw Output only markdown (no headers)");
eprintln!(
" --compact Collapse token-heavy source formatting such as dot leaders"
);
eprintln!(" --pages Insert page break markers (<!-- Page N -->)");
eprintln!(" --select-pages N Only process specified pages (e.g. 1,3,5-10)");
eprintln!(" --password PW Password for an encrypted PDF");
eprintln!(" --detect-only Only detect PDF type (no extraction)");
eprintln!(" --analyze Detect + extract + layout analysis (no markdown)");
process::exit(1);
@@ -105,11 +247,23 @@ fn main() {
let pdf_path = &args[1];
let json_output = args.iter().any(|a| a == "--json");
let items_json_output = args.iter().any(|a| a == "--items-json");
let raw_output = args.iter().any(|a| a == "--raw");
let compact_output = args.iter().any(|a| a == "--compact");
let page_numbers = args.iter().any(|a| a == "--pages");
let detect_only = args.iter().any(|a| a == "--detect-only");
let analyze = args.iter().any(|a| a == "--analyze");
// Parse --password value
let password = args.iter().position(|a| a == "--password").map(|i| {
args.get(i + 1)
.unwrap_or_else(|| {
eprintln!("Error: --password requires a value");
process::exit(1);
})
.clone()
});
// Parse --select-pages value
let page_filter = args
.iter()
@@ -129,6 +283,17 @@ fn main() {
})
});
if items_json_output {
match extract_items_json(pdf_path, page_filter.as_ref(), password.as_deref()) {
Ok(json) => println!("{}", json),
Err(e) => {
println!(r#"{{"error":"{}"}}"#, json_escape(&e.to_string()));
process::exit(1);
}
}
return;
}
let output_file = args
.get(2)
.filter(|a| !a.starts_with("--"))
@@ -143,10 +308,14 @@ fn main() {
};
let mut options = PdfOptions::new().mode(process_mode);
if compact_output {
options.markdown.profile = pdf_inspector::MarkdownProfile::Compact;
}
options.markdown.include_page_numbers = page_numbers;
if let Some(pages) = page_filter {
options.page_filter = Some(pages);
}
options.password = password;
match process_pdf_with_options(pdf_path, options) {
Ok(result) => {
@@ -177,12 +346,14 @@ fn main() {
.iter()
.map(|p| p.to_string())
.collect();
let ocr_reasons = format_ocr_reasons_by_page(&result.ocr_reasons_by_page);
println!(
r#"{{"pdf_type":"{}","page_count":{},"processing_time_ms":{},"pages_needing_ocr":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"has_encoding_issues":{}}}"#,
r#"{{"pdf_type":"{}","page_count":{},"processing_time_ms":{},"pages_needing_ocr":[{}],"ocr_reasons_by_page":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"has_encoding_issues":{}}}"#,
pdf_type_str,
result.page_count,
result.processing_time_ms,
ocr_pages.join(","),
ocr_reasons,
result.layout.is_complex,
table_pages.join(","),
col_pages.join(","),
@@ -223,8 +394,9 @@ fn main() {
.iter()
.map(|p| p.to_string())
.collect();
let ocr_reasons = format_ocr_reasons_by_page(&result.ocr_reasons_by_page);
println!(
r#"{{"pdf_type":"{}","page_count":{},"has_text":{},"processing_time_ms":{},"markdown_length":{},"pages_needing_ocr":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"has_encoding_issues":{},"markdown":"{}"}}"#,
r#"{{"pdf_type":"{}","page_count":{},"has_text":{},"processing_time_ms":{},"markdown_length":{},"pages_needing_ocr":[{}],"ocr_reasons_by_page":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"has_encoding_issues":{},"markdown":"{}"}}"#,
match result.pdf_type {
PdfType::TextBased => "text_based",
PdfType::Scanned => "scanned",
@@ -236,6 +408,7 @@ fn main() {
result.processing_time_ms,
result.markdown.as_ref().map(|m| m.len()).unwrap_or(0),
ocr_pages.join(","),
ocr_reasons,
result.layout.is_complex,
table_pages.join(","),
col_pages.join(","),
+176 -3
View File
@@ -60,6 +60,10 @@ pub struct PdfTypeResult {
/// 1-indexed page numbers that need OCR (image-only or insufficient text).
/// Empty for TextBased. All pages for Scanned/ImageBased. Specific pages for Mixed.
pub pages_needing_ocr: Vec<u32>,
/// Per-page explanation for `pages_needing_ocr`: 1-indexed page → reason
/// codes (`scanned`, `no_text`, `vector_text`, `suspected_garbled_text`).
/// Only contains pages that need OCR.
pub ocr_reasons_by_page: std::collections::BTreeMap<u32, Vec<String>>,
}
/// Configuration for PDF type detection
@@ -382,7 +386,12 @@ pub(crate) fn detect_from_document(
let analysis = if let Some(cached) = analysis_cache.get(&page_num) {
cached.clone()
} else if let Some(&page_id) = pages.get(&page_num) {
analyze_page_content(doc, page_id)
// Cache the fresh analysis so the reason-classification pass
// below sees the real signals (vector_text, etc.) instead of
// defaulting to "scanned".
let a = analyze_page_content(doc, page_id);
analysis_cache.insert(page_num, a.clone());
a
} else {
continue;
};
@@ -429,6 +438,9 @@ pub(crate) fn detect_from_document(
let analysis = analyze_page_content(doc, page_id);
if analysis.has_identity_h_no_tounicode || analysis.has_only_type3_fonts {
pages_needing_ocr.push(page_num);
// Cache so the reason pass reports suspected_garbled_text
// rather than defaulting to "scanned".
analysis_cache.insert(page_num, analysis);
}
}
}
@@ -436,6 +448,19 @@ pub(crate) fn detect_from_document(
pages_needing_ocr.sort();
pages_needing_ocr.dedup();
// Explain each OCR-flagged page. Pages we analyzed get a signal-derived
// reason; pages flagged only by whole-document classification (unsampled
// pages of a Scanned/ImageBased doc) default to `scanned`.
let mut ocr_reasons_by_page: std::collections::BTreeMap<u32, Vec<String>> =
std::collections::BTreeMap::new();
for &page_num in &pages_needing_ocr {
let reasons = match analysis_cache.get(&page_num) {
Some(analysis) => page_ocr_reasons(analysis),
None => vec![crate::OCR_REASON_SCANNED],
};
ocr_reasons_by_page.insert(page_num, reasons.into_iter().map(String::from).collect());
}
// Try to get title from metadata
let title = get_document_title(doc);
@@ -448,6 +473,7 @@ pub(crate) fn detect_from_document(
title,
ocr_recommended,
pages_needing_ocr,
ocr_reasons_by_page,
})
}
@@ -487,7 +513,7 @@ fn distribute_pages(n: u32, total: u32) -> Vec<u32> {
}
/// Page content analysis result
#[derive(Clone)]
#[derive(Clone, Default)]
struct PageAnalysis {
text_operator_count: u32,
has_images: bool,
@@ -523,6 +549,31 @@ struct PageAnalysis {
has_decodable_text_fonts: bool,
}
/// Explain *why* a page needs OCR, from its content analysis. Priority:
/// undecodable fonts (`suspected_garbled_text`) and vector-outlined text
/// (`vector_text`) come first because they persist even when a text layer is
/// present; otherwise a page with no extractable text is `scanned` when an
/// image backs it or `no_text` when nothing does.
fn page_ocr_reasons(a: &PageAnalysis) -> Vec<&'static str> {
let mut reasons = Vec::new();
if a.has_identity_h_no_tounicode || a.has_only_type3_fonts {
reasons.push(crate::OCR_REASON_SUSPECTED_GARBLED_TEXT);
}
if a.has_vector_text {
reasons.push(crate::OCR_REASON_VECTOR_TEXT);
}
if reasons.is_empty() {
let has_extractable_text = a.text_operator_count > 0 && a.unique_text_chars > 0;
if !has_extractable_text && !a.has_images && !a.has_template_image {
reasons.push(crate::OCR_REASON_NO_TEXT);
} else {
// Image-backed with no usable text, or too little text to trust.
reasons.push(crate::OCR_REASON_SCANNED);
}
}
reasons
}
/// Extracted font information from a Resource dictionary entry.
/// Stores the properties needed for decodability/identity-h checks
/// without holding a reference to the document.
@@ -1608,7 +1659,15 @@ fn hex_val(b: u8) -> Option<u8> {
/// Standard page: 612x792 points (US Letter) = ~485,000 sq points
/// At 2x resolution that's ~1.9M pixels, so we use 250K pixels as threshold
/// (accounting for varying DPI and page sizes)
fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
/// Returns `(has_images, total_image_area, has_template_image)` for a page.
/// `has_template_image` means a single large (>50% page coverage)
/// background image — the signal `classify_pdf`/`detect_pdf_type` uses to
/// route a page to OCR regardless of any incidental native text drawn over
/// it. Exposed at crate visibility so extraction-side per-page `needs_ocr`
/// computation (`extract_pages_markdown_mem`) can consult the same signal
/// instead of maintaining its own, independent notion of "needs OCR" that
/// can silently disagree with detection — see #227.
pub(crate) fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
// Threshold: image covering roughly half a page at 150+ DPI
// 612 * 792 / 2 * (150/72)^2 ≈ 1M pixels, but we'll be conservative
const TEMPLATE_IMAGE_THRESHOLD: u64 = 500_000; // 500K pixels
@@ -1690,6 +1749,62 @@ fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
(has_images, total_area, has_template_image)
}
/// Computes both `(needs_ocr_for_template_image, has_vector_text)` for a
/// page from a single shared `analyze_page_content` pass — that call
/// decompresses and scans every content stream (page + XObjects) plus
/// image coverage, so `extract_pages_markdown_mem` must not invoke it
/// twice per page (once per signal) the way `detect_from_document` avoids
/// by caching its per-page `PageAnalysis`.
///
/// `needs_ocr_for_template_image` is true when a page's template image
/// should be treated as a scan needing OCR — a single full-page background
/// image with little/no real text — rather than a text page that happens
/// to carry a watermark, letterhead, or figure. Mirrors the two distinct
/// signals classification uses to route a template-image page to OCR:
///
/// 1. `looks_like_scan`: image_count <= 1, few text operators (<50), and
/// low alphanumeric diversity in raw string operands (unless decodable
/// CID/ToUnicode fonts explain that away) — the gate used for
/// `pages_with_template_images` and Mixed-type per-page routing.
/// 2. Insufficient real text volume, using `DetectionConfig::default()`'s
/// `min_text_ops_per_page` (3) — the same threshold Mixed-type per-page
/// routing applies via `text_operator_count < config.min_text_ops_per_page
/// && has_images` (simplified here since a template image implies
/// `has_images`). Deliberately *not* the higher `effective_min_ops`
/// floor (`min_text_ops_per_page.max(10)`) that whole-document
/// `PdfType::ImageBased`/`Scanned` classification uses for
/// `pages_with_text` — that's a cross-page aggregate decision this
/// per-page function has no way to replicate exactly, and the lower
/// per-page threshold is the one a single page's own signals can
/// actually agree with.
///
/// `has_vector_text` is true when a page has vector-outlined text (glyphs
/// drawn as paths rather than shown via text-showing operators) —
/// `detect_from_document`'s Mixed-type per-page routing always sends
/// these pages to OCR, independent of any template-image check, since
/// outlined glyphs can't be extracted as text at all.
///
/// Exposed at crate visibility so `extract_pages_markdown_mem` can apply
/// the same gates classification needs elsewhere instead of treating the
/// raw signals alone as sufficient — see #227/#231.
pub(crate) fn page_ocr_signals(doc: &Document, page_id: ObjectId) -> (bool, bool) {
let analysis = analyze_page_content(doc, page_id);
let needs_ocr_for_template_image = if !analysis.has_template_image {
false
} else {
let alphanum_low = analysis.unique_alphanum_chars < 10
&& !(analysis.has_decodable_text_fonts && analysis.text_operator_count >= 10);
let looks_like_scan =
analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_low;
let insufficient_text =
analysis.text_operator_count < DetectionConfig::default().min_text_ops_per_page;
looks_like_scan || insufficient_text
};
(needs_ocr_for_template_image, analysis.has_vector_text)
}
/// Recursively collect image dimensions from XObject resources,
/// including images nested inside Form XObjects.
fn collect_images_from_resources(
@@ -1809,6 +1924,64 @@ fn get_document_title(doc: &Document) -> Option<String> {
mod tests {
use super::*;
#[test]
fn page_ocr_reasons_classify() {
// Scanned: no text, full-page image.
let scanned = PageAnalysis {
has_template_image: true,
..Default::default()
};
assert_eq!(page_ocr_reasons(&scanned), vec![crate::OCR_REASON_SCANNED]);
// Image-only page (no template flag, but has an image).
let image_only = PageAnalysis {
has_images: true,
..Default::default()
};
assert_eq!(
page_ocr_reasons(&image_only),
vec![crate::OCR_REASON_SCANNED]
);
// No text, no image → no_text.
let blank = PageAnalysis::default();
assert_eq!(page_ocr_reasons(&blank), vec![crate::OCR_REASON_NO_TEXT]);
// Vector-outlined text.
let vector = PageAnalysis {
has_vector_text: true,
..Default::default()
};
assert_eq!(
page_ocr_reasons(&vector),
vec![crate::OCR_REASON_VECTOR_TEXT]
);
// Undecodable fonts → garbled, and it wins over the fall-through.
let garbled = PageAnalysis {
has_identity_h_no_tounicode: true,
has_images: true,
..Default::default()
};
assert_eq!(
page_ocr_reasons(&garbled),
vec![crate::OCR_REASON_SUSPECTED_GARBLED_TEXT]
);
// A page with real extractable text and an image is not flagged here
// as scanned/no_text (only reached for pages already needing OCR).
let text_with_image = PageAnalysis {
text_operator_count: 40,
unique_text_chars: 120,
has_images: true,
..Default::default()
};
assert_eq!(
page_ocr_reasons(&text_with_image),
vec![crate::OCR_REASON_SCANNED]
);
}
#[test]
fn test_scan_content_operators() {
let mut uchars = HashSet::new();
+649
View File
@@ -0,0 +1,649 @@
//! Built-in glyph metrics for the 14 standard PDF fonts.
//!
//! PDFs may omit `/Widths` for non-embedded base-14 fonts (Times, Helvetica,
//! Courier, Symbol, ZapfDingbats); per the PDF spec the reader must supply
//! the metrics. Without them every text item gets width 0, which breaks
//! space synthesis, sub/superscript detection, and table column detection
//! (common in 1990s dvips/Distiller output).
//!
//! Tables are generated from the Adobe Core 14 AFM files (via reportlab's
//! `_fontdata`), keyed by Unicode char, sorted for binary search.
//! Generator: scratchpad/gen_base14.py (session tooling, not checked in).
/// Width in 1000ths of an em for `c` in the given base-14 font, or `None`
/// if the font is not one of the base 14 (after name normalization) or the
/// char has no glyph in its AFM.
pub(crate) fn base14_char_width(base_font: &str, c: char) -> Option<u16> {
let table = base14_table(base_font)?;
// AFM tables key visible glyphs only; alias the invisible variants the
// cp1252 fallback can produce so they get the metric of their visible
// counterpart instead of the generic default.
let c = match c {
'\u{00A0}' => ' ', // no-break space -> space
'\u{00AD}' => '-', // soft hyphen -> hyphen
_ => c,
};
table
.binary_search_by_key(&c, |&(ch, _)| ch)
.ok()
.map(|i| table[i].1)
}
/// True when the base font name normalizes to one of the standard 14 fonts.
pub(crate) fn is_base14_font(base_font: &str) -> bool {
base14_table(base_font).is_some()
}
/// Code → Unicode through the font's BUILT-IN encoding, for the base-14
/// fonts whose repertoire is not Latin (Symbol, ZapfDingbats). Their glyphs
/// live at byte positions that have nothing to do with cp1252 (Symbol 0x61
/// renders α, Zapf 0x21 renders ✁), so advance widths must be resolved
/// through this mapping — the renderer draws these glyphs regardless of how
/// the text decoder transliterates them. Returns `None` for the Latin text
/// fonts, which follow standard single-byte encodings.
pub(crate) fn builtin_encoding_char(base_font: &str, code: u8) -> Option<char> {
let table = base14_table(base_font)?;
let enc: &[(u8, char)] = if std::ptr::eq(table, SYMBOL) {
SYMBOL_ENCODING
} else if std::ptr::eq(table, ZAPFDINGBATS) {
ZAPFDINGBATS_ENCODING
} else {
return None;
};
enc.binary_search_by_key(&code, |&(b, _)| b)
.ok()
.map(|i| enc[i].1)
}
/// Map a BaseFont name (possibly subset-prefixed, e.g. "ABCDEF+Times-Bold",
/// or a common alias like "Arial" / "TimesNewRomanPSMT") to its width table.
fn base14_table(base_font: &str) -> Option<&'static [(char, u16)]> {
// Strip subset prefix "ABCDEF+"
let name = match base_font.split_once('+') {
Some((prefix, rest))
if prefix.len() == 6 && prefix.chars().all(|c| c.is_ascii_uppercase()) =>
{
rest
}
_ => base_font,
};
let lower = name.to_ascii_lowercase();
let bold = lower.contains("bold");
let italic = lower.contains("italic") || lower.contains("oblique");
if lower.contains("courier") {
return Some(match (bold, italic) {
(false, false) => COURIER,
(true, false) => COURIER_BOLD,
(false, true) => COURIER_OBLIQUE,
(true, true) => COURIER_BOLDOBLIQUE,
});
}
if lower.contains("helvetica") || lower.contains("arial") {
return Some(match (bold, italic) {
(false, false) => HELVETICA,
(true, false) => HELVETICA_BOLD,
(false, true) => HELVETICA_OBLIQUE,
(true, true) => HELVETICA_BOLDOBLIQUE,
});
}
if lower.contains("times") {
return Some(match (bold, italic) {
(false, false) => TIMES_ROMAN,
(true, false) => TIMES_BOLD,
(false, true) => TIMES_ITALIC,
(true, true) => TIMES_BOLDITALIC,
});
}
// Symbol and ZapfDingbats have unique glyph repertoires, so only exact
// names (plus the common MT/ITC aliases) qualify — a custom font that
// merely mentions "Symbol" in its name must not get these metrics.
match lower.as_str() {
"zapfdingbats" | "dingbats" | "itczapfdingbats" | "zapfdingbatsitc" => {
return Some(ZAPFDINGBATS)
}
"symbol" | "symbolmt" | "symbolitc" => return Some(SYMBOL),
_ => {}
}
None
}
#[rustfmt::skip]
static COURIER: &[(char, u16)] = &[
(' ', 600), ('!', 600), ('"', 600), ('#', 600), ('$', 600), ('%', 600),
('&', 600), ('\'', 600), ('(', 600), (')', 600), ('*', 600), ('+', 600),
(',', 600), ('-', 600), ('.', 600), ('/', 600), ('0', 600), ('1', 600),
('2', 600), ('3', 600), ('4', 600), ('5', 600), ('6', 600), ('7', 600),
('8', 600), ('9', 600), (':', 600), (';', 600), ('<', 600), ('=', 600),
('>', 600), ('?', 600), ('@', 600), ('A', 600), ('B', 600), ('C', 600),
('D', 600), ('E', 600), ('F', 600), ('G', 600), ('H', 600), ('I', 600),
('J', 600), ('K', 600), ('L', 600), ('M', 600), ('N', 600), ('O', 600),
('P', 600), ('Q', 600), ('R', 600), ('S', 600), ('T', 600), ('U', 600),
('V', 600), ('W', 600), ('X', 600), ('Y', 600), ('Z', 600), ('[', 600),
('\\', 600), (']', 600), ('^', 600), ('_', 600), ('`', 600), ('a', 600),
('b', 600), ('c', 600), ('d', 600), ('e', 600), ('f', 600), ('g', 600),
('h', 600), ('i', 600), ('j', 600), ('k', 600), ('l', 600), ('m', 600),
('n', 600), ('o', 600), ('p', 600), ('q', 600), ('r', 600), ('s', 600),
('t', 600), ('u', 600), ('v', 600), ('w', 600), ('x', 600), ('y', 600),
('z', 600), ('{', 600), ('|', 600), ('}', 600), ('~', 600), ('\u{00A1}', 600),
('\u{00A2}', 600), ('\u{00A3}', 600), ('\u{00A4}', 600), ('\u{00A5}', 600), ('\u{00A6}', 600), ('\u{00A7}', 600),
('\u{00A8}', 600), ('\u{00A9}', 600), ('\u{00AA}', 600), ('\u{00AB}', 600), ('\u{00AC}', 600), ('\u{00AE}', 600),
('\u{00AF}', 600), ('\u{00B0}', 600), ('\u{00B1}', 600), ('\u{00B2}', 600), ('\u{00B3}', 600), ('\u{00B4}', 600),
('\u{00B5}', 600), ('\u{00B6}', 600), ('\u{00B7}', 600), ('\u{00B8}', 600), ('\u{00B9}', 600), ('\u{00BA}', 600),
('\u{00BB}', 600), ('\u{00BC}', 600), ('\u{00BD}', 600), ('\u{00BE}', 600), ('\u{00BF}', 600), ('\u{00C0}', 600),
('\u{00C1}', 600), ('\u{00C2}', 600), ('\u{00C3}', 600), ('\u{00C4}', 600), ('\u{00C5}', 600), ('\u{00C6}', 600),
('\u{00C7}', 600), ('\u{00C8}', 600), ('\u{00C9}', 600), ('\u{00CA}', 600), ('\u{00CB}', 600), ('\u{00CC}', 600),
('\u{00CD}', 600), ('\u{00CE}', 600), ('\u{00CF}', 600), ('\u{00D0}', 600), ('\u{00D1}', 600), ('\u{00D2}', 600),
('\u{00D3}', 600), ('\u{00D4}', 600), ('\u{00D5}', 600), ('\u{00D6}', 600), ('\u{00D7}', 600), ('\u{00D8}', 600),
('\u{00D9}', 600), ('\u{00DA}', 600), ('\u{00DB}', 600), ('\u{00DC}', 600), ('\u{00DD}', 600), ('\u{00DE}', 600),
('\u{00DF}', 600), ('\u{00E0}', 600), ('\u{00E1}', 600), ('\u{00E2}', 600), ('\u{00E3}', 600), ('\u{00E4}', 600),
('\u{00E5}', 600), ('\u{00E6}', 600), ('\u{00E7}', 600), ('\u{00E8}', 600), ('\u{00E9}', 600), ('\u{00EA}', 600),
('\u{00EB}', 600), ('\u{00EC}', 600), ('\u{00ED}', 600), ('\u{00EE}', 600), ('\u{00EF}', 600), ('\u{00F0}', 600),
('\u{00F1}', 600), ('\u{00F2}', 600), ('\u{00F3}', 600), ('\u{00F4}', 600), ('\u{00F5}', 600), ('\u{00F6}', 600),
('\u{00F7}', 600), ('\u{00F8}', 600), ('\u{00F9}', 600), ('\u{00FA}', 600), ('\u{00FB}', 600), ('\u{00FC}', 600),
('\u{00FD}', 600), ('\u{00FE}', 600), ('\u{00FF}', 600), ('\u{0131}', 600), ('\u{0141}', 600), ('\u{0142}', 600),
('\u{0152}', 600), ('\u{0153}', 600), ('\u{0160}', 600), ('\u{0161}', 600), ('\u{0178}', 600), ('\u{017D}', 600),
('\u{017E}', 600), ('\u{0192}', 600), ('\u{02C6}', 600), ('\u{02C7}', 600), ('\u{02D8}', 600), ('\u{02D9}', 600),
('\u{02DA}', 600), ('\u{02DB}', 600), ('\u{02DC}', 600), ('\u{02DD}', 600), ('\u{2013}', 600), ('\u{2014}', 600),
('\u{2018}', 600), ('\u{2019}', 600), ('\u{201A}', 600), ('\u{201C}', 600), ('\u{201D}', 600), ('\u{201E}', 600),
('\u{2020}', 600), ('\u{2021}', 600), ('\u{2022}', 600), ('\u{2026}', 600), ('\u{2030}', 600), ('\u{2039}', 600),
('\u{203A}', 600), ('\u{2044}', 600), ('\u{20AC}', 600), ('\u{2122}', 600), ('\u{2212}', 600), ('\u{FB01}', 600),
('\u{FB02}', 600),
];
static COURIER_BOLD: &[(char, u16)] = COURIER;
static COURIER_OBLIQUE: &[(char, u16)] = COURIER;
static COURIER_BOLDOBLIQUE: &[(char, u16)] = COURIER;
#[rustfmt::skip]
static HELVETICA: &[(char, u16)] = &[
(' ', 278), ('!', 278), ('"', 355), ('#', 556), ('$', 556), ('%', 889),
('&', 667), ('\'', 191), ('(', 333), (')', 333), ('*', 389), ('+', 584),
(',', 278), ('-', 333), ('.', 278), ('/', 278), ('0', 556), ('1', 556),
('2', 556), ('3', 556), ('4', 556), ('5', 556), ('6', 556), ('7', 556),
('8', 556), ('9', 556), (':', 278), (';', 278), ('<', 584), ('=', 584),
('>', 584), ('?', 556), ('@', 1015), ('A', 667), ('B', 667), ('C', 722),
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 722), ('I', 278),
('J', 500), ('K', 667), ('L', 556), ('M', 833), ('N', 722), ('O', 778),
('P', 667), ('Q', 778), ('R', 722), ('S', 667), ('T', 611), ('U', 722),
('V', 667), ('W', 944), ('X', 667), ('Y', 667), ('Z', 611), ('[', 278),
('\\', 278), (']', 278), ('^', 469), ('_', 556), ('`', 333), ('a', 556),
('b', 556), ('c', 500), ('d', 556), ('e', 556), ('f', 278), ('g', 556),
('h', 556), ('i', 222), ('j', 222), ('k', 500), ('l', 222), ('m', 833),
('n', 556), ('o', 556), ('p', 556), ('q', 556), ('r', 333), ('s', 500),
('t', 278), ('u', 556), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
('z', 500), ('{', 334), ('|', 260), ('}', 334), ('~', 584), ('\u{00A1}', 333),
('\u{00A2}', 556), ('\u{00A3}', 556), ('\u{00A4}', 556), ('\u{00A5}', 556), ('\u{00A6}', 260), ('\u{00A7}', 556),
('\u{00A8}', 333), ('\u{00A9}', 737), ('\u{00AA}', 370), ('\u{00AB}', 556), ('\u{00AC}', 584), ('\u{00AE}', 737),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 584), ('\u{00B2}', 333), ('\u{00B3}', 333), ('\u{00B4}', 333),
('\u{00B5}', 556), ('\u{00B6}', 537), ('\u{00B7}', 278), ('\u{00B8}', 333), ('\u{00B9}', 333), ('\u{00BA}', 365),
('\u{00BB}', 556), ('\u{00BC}', 834), ('\u{00BD}', 834), ('\u{00BE}', 834), ('\u{00BF}', 611), ('\u{00C0}', 667),
('\u{00C1}', 667), ('\u{00C2}', 667), ('\u{00C3}', 667), ('\u{00C4}', 667), ('\u{00C5}', 667), ('\u{00C6}', 1000),
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 278),
('\u{00CD}', 278), ('\u{00CE}', 278), ('\u{00CF}', 278), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 584), ('\u{00D8}', 778),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 667), ('\u{00DE}', 667),
('\u{00DF}', 611), ('\u{00E0}', 556), ('\u{00E1}', 556), ('\u{00E2}', 556), ('\u{00E3}', 556), ('\u{00E4}', 556),
('\u{00E5}', 556), ('\u{00E6}', 889), ('\u{00E7}', 500), ('\u{00E8}', 556), ('\u{00E9}', 556), ('\u{00EA}', 556),
('\u{00EB}', 556), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 556),
('\u{00F1}', 556), ('\u{00F2}', 556), ('\u{00F3}', 556), ('\u{00F4}', 556), ('\u{00F5}', 556), ('\u{00F6}', 556),
('\u{00F7}', 584), ('\u{00F8}', 611), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
('\u{00FD}', 500), ('\u{00FE}', 556), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 556), ('\u{0142}', 222),
('\u{0152}', 1000), ('\u{0153}', 944), ('\u{0160}', 667), ('\u{0161}', 500), ('\u{0178}', 667), ('\u{017D}', 611),
('\u{017E}', 500), ('\u{0192}', 556), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 556), ('\u{2014}', 1000),
('\u{2018}', 222), ('\u{2019}', 222), ('\u{201A}', 222), ('\u{201C}', 333), ('\u{201D}', 333), ('\u{201E}', 333),
('\u{2020}', 556), ('\u{2021}', 556), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 556), ('\u{2122}', 1000), ('\u{2212}', 584), ('\u{FB01}', 500),
('\u{FB02}', 500),
];
#[rustfmt::skip]
static HELVETICA_BOLD: &[(char, u16)] = &[
(' ', 278), ('!', 333), ('"', 474), ('#', 556), ('$', 556), ('%', 889),
('&', 722), ('\'', 238), ('(', 333), (')', 333), ('*', 389), ('+', 584),
(',', 278), ('-', 333), ('.', 278), ('/', 278), ('0', 556), ('1', 556),
('2', 556), ('3', 556), ('4', 556), ('5', 556), ('6', 556), ('7', 556),
('8', 556), ('9', 556), (':', 333), (';', 333), ('<', 584), ('=', 584),
('>', 584), ('?', 611), ('@', 975), ('A', 722), ('B', 722), ('C', 722),
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 722), ('I', 278),
('J', 556), ('K', 722), ('L', 611), ('M', 833), ('N', 722), ('O', 778),
('P', 667), ('Q', 778), ('R', 722), ('S', 667), ('T', 611), ('U', 722),
('V', 667), ('W', 944), ('X', 667), ('Y', 667), ('Z', 611), ('[', 333),
('\\', 278), (']', 333), ('^', 584), ('_', 556), ('`', 333), ('a', 556),
('b', 611), ('c', 556), ('d', 611), ('e', 556), ('f', 333), ('g', 611),
('h', 611), ('i', 278), ('j', 278), ('k', 556), ('l', 278), ('m', 889),
('n', 611), ('o', 611), ('p', 611), ('q', 611), ('r', 389), ('s', 556),
('t', 333), ('u', 611), ('v', 556), ('w', 778), ('x', 556), ('y', 556),
('z', 500), ('{', 389), ('|', 280), ('}', 389), ('~', 584), ('\u{00A1}', 333),
('\u{00A2}', 556), ('\u{00A3}', 556), ('\u{00A4}', 556), ('\u{00A5}', 556), ('\u{00A6}', 280), ('\u{00A7}', 556),
('\u{00A8}', 333), ('\u{00A9}', 737), ('\u{00AA}', 370), ('\u{00AB}', 556), ('\u{00AC}', 584), ('\u{00AE}', 737),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 584), ('\u{00B2}', 333), ('\u{00B3}', 333), ('\u{00B4}', 333),
('\u{00B5}', 611), ('\u{00B6}', 556), ('\u{00B7}', 278), ('\u{00B8}', 333), ('\u{00B9}', 333), ('\u{00BA}', 365),
('\u{00BB}', 556), ('\u{00BC}', 834), ('\u{00BD}', 834), ('\u{00BE}', 834), ('\u{00BF}', 611), ('\u{00C0}', 722),
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 1000),
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 278),
('\u{00CD}', 278), ('\u{00CE}', 278), ('\u{00CF}', 278), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 584), ('\u{00D8}', 778),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 667), ('\u{00DE}', 667),
('\u{00DF}', 611), ('\u{00E0}', 556), ('\u{00E1}', 556), ('\u{00E2}', 556), ('\u{00E3}', 556), ('\u{00E4}', 556),
('\u{00E5}', 556), ('\u{00E6}', 889), ('\u{00E7}', 556), ('\u{00E8}', 556), ('\u{00E9}', 556), ('\u{00EA}', 556),
('\u{00EB}', 556), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 611),
('\u{00F1}', 611), ('\u{00F2}', 611), ('\u{00F3}', 611), ('\u{00F4}', 611), ('\u{00F5}', 611), ('\u{00F6}', 611),
('\u{00F7}', 584), ('\u{00F8}', 611), ('\u{00F9}', 611), ('\u{00FA}', 611), ('\u{00FB}', 611), ('\u{00FC}', 611),
('\u{00FD}', 556), ('\u{00FE}', 611), ('\u{00FF}', 556), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
('\u{0152}', 1000), ('\u{0153}', 944), ('\u{0160}', 667), ('\u{0161}', 556), ('\u{0178}', 667), ('\u{017D}', 611),
('\u{017E}', 500), ('\u{0192}', 556), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 556), ('\u{2014}', 1000),
('\u{2018}', 278), ('\u{2019}', 278), ('\u{201A}', 278), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
('\u{2020}', 556), ('\u{2021}', 556), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 556), ('\u{2122}', 1000), ('\u{2212}', 584), ('\u{FB01}', 611),
('\u{FB02}', 611),
];
static HELVETICA_OBLIQUE: &[(char, u16)] = HELVETICA;
static HELVETICA_BOLDOBLIQUE: &[(char, u16)] = HELVETICA_BOLD;
#[rustfmt::skip]
static TIMES_ROMAN: &[(char, u16)] = &[
(' ', 250), ('!', 333), ('"', 408), ('#', 500), ('$', 500), ('%', 833),
('&', 778), ('\'', 180), ('(', 333), (')', 333), ('*', 500), ('+', 564),
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
('8', 500), ('9', 500), (':', 278), (';', 278), ('<', 564), ('=', 564),
('>', 564), ('?', 444), ('@', 921), ('A', 722), ('B', 667), ('C', 667),
('D', 722), ('E', 611), ('F', 556), ('G', 722), ('H', 722), ('I', 333),
('J', 389), ('K', 722), ('L', 611), ('M', 889), ('N', 722), ('O', 722),
('P', 556), ('Q', 722), ('R', 667), ('S', 556), ('T', 611), ('U', 722),
('V', 722), ('W', 944), ('X', 722), ('Y', 722), ('Z', 611), ('[', 333),
('\\', 278), (']', 333), ('^', 469), ('_', 500), ('`', 333), ('a', 444),
('b', 500), ('c', 444), ('d', 500), ('e', 444), ('f', 333), ('g', 500),
('h', 500), ('i', 278), ('j', 278), ('k', 500), ('l', 278), ('m', 778),
('n', 500), ('o', 500), ('p', 500), ('q', 500), ('r', 333), ('s', 389),
('t', 278), ('u', 500), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
('z', 444), ('{', 480), ('|', 200), ('}', 480), ('~', 541), ('\u{00A1}', 333),
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 200), ('\u{00A7}', 500),
('\u{00A8}', 333), ('\u{00A9}', 760), ('\u{00AA}', 276), ('\u{00AB}', 500), ('\u{00AC}', 564), ('\u{00AE}', 760),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 564), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
('\u{00B5}', 500), ('\u{00B6}', 453), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 310),
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 444), ('\u{00C0}', 722),
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 889),
('\u{00C7}', 667), ('\u{00C8}', 611), ('\u{00C9}', 611), ('\u{00CA}', 611), ('\u{00CB}', 611), ('\u{00CC}', 333),
('\u{00CD}', 333), ('\u{00CE}', 333), ('\u{00CF}', 333), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 722),
('\u{00D3}', 722), ('\u{00D4}', 722), ('\u{00D5}', 722), ('\u{00D6}', 722), ('\u{00D7}', 564), ('\u{00D8}', 722),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 722), ('\u{00DE}', 556),
('\u{00DF}', 500), ('\u{00E0}', 444), ('\u{00E1}', 444), ('\u{00E2}', 444), ('\u{00E3}', 444), ('\u{00E4}', 444),
('\u{00E5}', 444), ('\u{00E6}', 667), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
('\u{00F1}', 500), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
('\u{00F7}', 564), ('\u{00F8}', 500), ('\u{00F9}', 500), ('\u{00FA}', 500), ('\u{00FB}', 500), ('\u{00FC}', 500),
('\u{00FD}', 500), ('\u{00FE}', 500), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
('\u{0152}', 889), ('\u{0153}', 722), ('\u{0160}', 556), ('\u{0161}', 389), ('\u{0178}', 722), ('\u{017D}', 611),
('\u{017E}', 444), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 1000),
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 444), ('\u{201D}', 444), ('\u{201E}', 444),
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 980), ('\u{2212}', 564), ('\u{FB01}', 556),
('\u{FB02}', 556),
];
#[rustfmt::skip]
static TIMES_BOLD: &[(char, u16)] = &[
(' ', 250), ('!', 333), ('"', 555), ('#', 500), ('$', 500), ('%', 1000),
('&', 833), ('\'', 278), ('(', 333), (')', 333), ('*', 500), ('+', 570),
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
('8', 500), ('9', 500), (':', 333), (';', 333), ('<', 570), ('=', 570),
('>', 570), ('?', 500), ('@', 930), ('A', 722), ('B', 667), ('C', 722),
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 778), ('I', 389),
('J', 500), ('K', 778), ('L', 667), ('M', 944), ('N', 722), ('O', 778),
('P', 611), ('Q', 778), ('R', 722), ('S', 556), ('T', 667), ('U', 722),
('V', 722), ('W', 1000), ('X', 722), ('Y', 722), ('Z', 667), ('[', 333),
('\\', 278), (']', 333), ('^', 581), ('_', 500), ('`', 333), ('a', 500),
('b', 556), ('c', 444), ('d', 556), ('e', 444), ('f', 333), ('g', 500),
('h', 556), ('i', 278), ('j', 333), ('k', 556), ('l', 278), ('m', 833),
('n', 556), ('o', 500), ('p', 556), ('q', 556), ('r', 444), ('s', 389),
('t', 333), ('u', 556), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
('z', 444), ('{', 394), ('|', 220), ('}', 394), ('~', 520), ('\u{00A1}', 333),
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 220), ('\u{00A7}', 500),
('\u{00A8}', 333), ('\u{00A9}', 747), ('\u{00AA}', 300), ('\u{00AB}', 500), ('\u{00AC}', 570), ('\u{00AE}', 747),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 570), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
('\u{00B5}', 556), ('\u{00B6}', 540), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 330),
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 500), ('\u{00C0}', 722),
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 1000),
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 389),
('\u{00CD}', 389), ('\u{00CE}', 389), ('\u{00CF}', 389), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 570), ('\u{00D8}', 778),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 722), ('\u{00DE}', 611),
('\u{00DF}', 556), ('\u{00E0}', 500), ('\u{00E1}', 500), ('\u{00E2}', 500), ('\u{00E3}', 500), ('\u{00E4}', 500),
('\u{00E5}', 500), ('\u{00E6}', 722), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
('\u{00F1}', 556), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
('\u{00F7}', 570), ('\u{00F8}', 500), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
('\u{00FD}', 500), ('\u{00FE}', 556), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 667), ('\u{0142}', 278),
('\u{0152}', 1000), ('\u{0153}', 722), ('\u{0160}', 556), ('\u{0161}', 389), ('\u{0178}', 722), ('\u{017D}', 667),
('\u{017E}', 444), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 1000),
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 1000), ('\u{2212}', 570), ('\u{FB01}', 556),
('\u{FB02}', 556),
];
#[rustfmt::skip]
static TIMES_ITALIC: &[(char, u16)] = &[
(' ', 250), ('!', 333), ('"', 420), ('#', 500), ('$', 500), ('%', 833),
('&', 778), ('\'', 214), ('(', 333), (')', 333), ('*', 500), ('+', 675),
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
('8', 500), ('9', 500), (':', 333), (';', 333), ('<', 675), ('=', 675),
('>', 675), ('?', 500), ('@', 920), ('A', 611), ('B', 611), ('C', 667),
('D', 722), ('E', 611), ('F', 611), ('G', 722), ('H', 722), ('I', 333),
('J', 444), ('K', 667), ('L', 556), ('M', 833), ('N', 667), ('O', 722),
('P', 611), ('Q', 722), ('R', 611), ('S', 500), ('T', 556), ('U', 722),
('V', 611), ('W', 833), ('X', 611), ('Y', 556), ('Z', 556), ('[', 389),
('\\', 278), (']', 389), ('^', 422), ('_', 500), ('`', 333), ('a', 500),
('b', 500), ('c', 444), ('d', 500), ('e', 444), ('f', 278), ('g', 500),
('h', 500), ('i', 278), ('j', 278), ('k', 444), ('l', 278), ('m', 722),
('n', 500), ('o', 500), ('p', 500), ('q', 500), ('r', 389), ('s', 389),
('t', 278), ('u', 500), ('v', 444), ('w', 667), ('x', 444), ('y', 444),
('z', 389), ('{', 400), ('|', 275), ('}', 400), ('~', 541), ('\u{00A1}', 389),
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 275), ('\u{00A7}', 500),
('\u{00A8}', 333), ('\u{00A9}', 760), ('\u{00AA}', 276), ('\u{00AB}', 500), ('\u{00AC}', 675), ('\u{00AE}', 760),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 675), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
('\u{00B5}', 500), ('\u{00B6}', 523), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 310),
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 500), ('\u{00C0}', 611),
('\u{00C1}', 611), ('\u{00C2}', 611), ('\u{00C3}', 611), ('\u{00C4}', 611), ('\u{00C5}', 611), ('\u{00C6}', 889),
('\u{00C7}', 667), ('\u{00C8}', 611), ('\u{00C9}', 611), ('\u{00CA}', 611), ('\u{00CB}', 611), ('\u{00CC}', 333),
('\u{00CD}', 333), ('\u{00CE}', 333), ('\u{00CF}', 333), ('\u{00D0}', 722), ('\u{00D1}', 667), ('\u{00D2}', 722),
('\u{00D3}', 722), ('\u{00D4}', 722), ('\u{00D5}', 722), ('\u{00D6}', 722), ('\u{00D7}', 675), ('\u{00D8}', 722),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 556), ('\u{00DE}', 611),
('\u{00DF}', 500), ('\u{00E0}', 500), ('\u{00E1}', 500), ('\u{00E2}', 500), ('\u{00E3}', 500), ('\u{00E4}', 500),
('\u{00E5}', 500), ('\u{00E6}', 667), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
('\u{00F1}', 500), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
('\u{00F7}', 675), ('\u{00F8}', 500), ('\u{00F9}', 500), ('\u{00FA}', 500), ('\u{00FB}', 500), ('\u{00FC}', 500),
('\u{00FD}', 444), ('\u{00FE}', 500), ('\u{00FF}', 444), ('\u{0131}', 278), ('\u{0141}', 556), ('\u{0142}', 278),
('\u{0152}', 944), ('\u{0153}', 667), ('\u{0160}', 500), ('\u{0161}', 389), ('\u{0178}', 556), ('\u{017D}', 556),
('\u{017E}', 389), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 889),
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 556), ('\u{201D}', 556), ('\u{201E}', 556),
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 889), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 980), ('\u{2212}', 675), ('\u{FB01}', 500),
('\u{FB02}', 500),
];
#[rustfmt::skip]
static TIMES_BOLDITALIC: &[(char, u16)] = &[
(' ', 250), ('!', 389), ('"', 555), ('#', 500), ('$', 500), ('%', 833),
('&', 778), ('\'', 278), ('(', 333), (')', 333), ('*', 500), ('+', 570),
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
('8', 500), ('9', 500), (':', 333), (';', 333), ('<', 570), ('=', 570),
('>', 570), ('?', 500), ('@', 832), ('A', 667), ('B', 667), ('C', 667),
('D', 722), ('E', 667), ('F', 667), ('G', 722), ('H', 778), ('I', 389),
('J', 500), ('K', 667), ('L', 611), ('M', 889), ('N', 722), ('O', 722),
('P', 611), ('Q', 722), ('R', 667), ('S', 556), ('T', 611), ('U', 722),
('V', 667), ('W', 889), ('X', 667), ('Y', 611), ('Z', 611), ('[', 333),
('\\', 278), (']', 333), ('^', 570), ('_', 500), ('`', 333), ('a', 500),
('b', 500), ('c', 444), ('d', 500), ('e', 444), ('f', 333), ('g', 500),
('h', 556), ('i', 278), ('j', 278), ('k', 500), ('l', 278), ('m', 778),
('n', 556), ('o', 500), ('p', 500), ('q', 500), ('r', 389), ('s', 389),
('t', 278), ('u', 556), ('v', 444), ('w', 667), ('x', 500), ('y', 444),
('z', 389), ('{', 348), ('|', 220), ('}', 348), ('~', 570), ('\u{00A1}', 389),
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 220), ('\u{00A7}', 500),
('\u{00A8}', 333), ('\u{00A9}', 747), ('\u{00AA}', 266), ('\u{00AB}', 500), ('\u{00AC}', 606), ('\u{00AE}', 747),
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 570), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
('\u{00B5}', 576), ('\u{00B6}', 500), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 300),
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 500), ('\u{00C0}', 667),
('\u{00C1}', 667), ('\u{00C2}', 667), ('\u{00C3}', 667), ('\u{00C4}', 667), ('\u{00C5}', 667), ('\u{00C6}', 944),
('\u{00C7}', 667), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 389),
('\u{00CD}', 389), ('\u{00CE}', 389), ('\u{00CF}', 389), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 722),
('\u{00D3}', 722), ('\u{00D4}', 722), ('\u{00D5}', 722), ('\u{00D6}', 722), ('\u{00D7}', 570), ('\u{00D8}', 722),
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 611), ('\u{00DE}', 611),
('\u{00DF}', 500), ('\u{00E0}', 500), ('\u{00E1}', 500), ('\u{00E2}', 500), ('\u{00E3}', 500), ('\u{00E4}', 500),
('\u{00E5}', 500), ('\u{00E6}', 722), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
('\u{00F1}', 556), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
('\u{00F7}', 570), ('\u{00F8}', 500), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
('\u{00FD}', 444), ('\u{00FE}', 500), ('\u{00FF}', 444), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
('\u{0152}', 944), ('\u{0153}', 722), ('\u{0160}', 556), ('\u{0161}', 389), ('\u{0178}', 611), ('\u{017D}', 611),
('\u{017E}', 389), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 1000),
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 1000), ('\u{2212}', 606), ('\u{FB01}', 556),
('\u{FB02}', 556),
];
#[rustfmt::skip]
static SYMBOL: &[(char, u16)] = &[
(' ', 250), ('!', 333), ('#', 500), ('%', 833), ('&', 778), ('(', 333),
(')', 333), ('+', 549), (',', 250), ('.', 250), ('/', 278), ('0', 500),
('1', 500), ('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500),
('7', 500), ('8', 500), ('9', 500), (':', 278), (';', 278), ('<', 549),
('=', 549), ('>', 549), ('?', 444), ('[', 333), (']', 333), ('_', 500),
('{', 480), ('|', 200), ('}', 480), ('\u{00AC}', 713), ('\u{00B0}', 400), ('\u{00B1}', 549),
('\u{00B5}', 576), ('\u{00D7}', 549), ('\u{00F7}', 549), ('\u{0192}', 500), ('\u{0391}', 722), ('\u{0392}', 667),
('\u{0393}', 603), ('\u{0395}', 611), ('\u{0396}', 611), ('\u{0397}', 722), ('\u{0398}', 741), ('\u{0399}', 333),
('\u{039A}', 722), ('\u{039B}', 686), ('\u{039C}', 889), ('\u{039D}', 722), ('\u{039E}', 645), ('\u{039F}', 722),
('\u{03A0}', 768), ('\u{03A1}', 556), ('\u{03A3}', 592), ('\u{03A4}', 611), ('\u{03A5}', 690), ('\u{03A6}', 763),
('\u{03A7}', 722), ('\u{03A8}', 795), ('\u{03B1}', 631), ('\u{03B2}', 549), ('\u{03B3}', 411), ('\u{03B4}', 494),
('\u{03B5}', 439), ('\u{03B6}', 494), ('\u{03B7}', 603), ('\u{03B8}', 521), ('\u{03B9}', 329), ('\u{03BA}', 549),
('\u{03BB}', 549), ('\u{03BD}', 521), ('\u{03BE}', 493), ('\u{03BF}', 549), ('\u{03C0}', 549), ('\u{03C1}', 549),
('\u{03C2}', 439), ('\u{03C3}', 603), ('\u{03C4}', 439), ('\u{03C5}', 576), ('\u{03C6}', 521), ('\u{03C7}', 549),
('\u{03C8}', 686), ('\u{03C9}', 686), ('\u{03D1}', 631), ('\u{03D2}', 620), ('\u{03D5}', 603), ('\u{03D6}', 713),
('\u{2022}', 460), ('\u{2026}', 1000), ('\u{2032}', 247), ('\u{2033}', 411), ('\u{2044}', 167), ('\u{20AC}', 750),
('\u{2111}', 686), ('\u{2118}', 987), ('\u{211C}', 795), ('\u{2126}', 768), ('\u{2135}', 823), ('\u{2190}', 987),
('\u{2191}', 603), ('\u{2192}', 987), ('\u{2193}', 603), ('\u{2194}', 1042), ('\u{21B5}', 658), ('\u{21D0}', 987),
('\u{21D1}', 603), ('\u{21D2}', 987), ('\u{21D3}', 603), ('\u{21D4}', 1042), ('\u{2200}', 713), ('\u{2202}', 494),
('\u{2203}', 549), ('\u{2205}', 823), ('\u{2206}', 612), ('\u{2207}', 713), ('\u{2208}', 713), ('\u{2209}', 713),
('\u{220B}', 439), ('\u{220F}', 823), ('\u{2211}', 713), ('\u{2212}', 549), ('\u{2217}', 500), ('\u{221A}', 549),
('\u{221D}', 713), ('\u{221E}', 713), ('\u{2220}', 768), ('\u{2227}', 603), ('\u{2228}', 603), ('\u{2229}', 768),
('\u{222A}', 768), ('\u{222B}', 274), ('\u{2234}', 863), ('\u{223C}', 549), ('\u{2245}', 549), ('\u{2248}', 549),
('\u{2260}', 549), ('\u{2261}', 549), ('\u{2264}', 549), ('\u{2265}', 549), ('\u{2282}', 713), ('\u{2283}', 713),
('\u{2284}', 713), ('\u{2286}', 713), ('\u{2287}', 713), ('\u{2295}', 768), ('\u{2297}', 768), ('\u{22A5}', 658),
('\u{22C5}', 250), ('\u{2320}', 686), ('\u{2321}', 686), ('\u{2329}', 329), ('\u{232A}', 329), ('\u{25CA}', 494),
('\u{2660}', 753), ('\u{2663}', 753), ('\u{2665}', 753), ('\u{2666}', 753), ('\u{F6D9}', 790), ('\u{F6DA}', 790),
('\u{F6DB}', 890), ('\u{F8E5}', 500), ('\u{F8E6}', 603), ('\u{F8E7}', 1000), ('\u{F8E8}', 790), ('\u{F8E9}', 790),
('\u{F8EA}', 786), ('\u{F8EB}', 384), ('\u{F8EC}', 384), ('\u{F8ED}', 384), ('\u{F8EE}', 384), ('\u{F8EF}', 384),
('\u{F8F0}', 384), ('\u{F8F1}', 494), ('\u{F8F2}', 494), ('\u{F8F3}', 494), ('\u{F8F4}', 494), ('\u{F8F5}', 686),
('\u{F8F6}', 384), ('\u{F8F7}', 384), ('\u{F8F8}', 384), ('\u{F8F9}', 384), ('\u{F8FA}', 384), ('\u{F8FB}', 384),
('\u{F8FC}', 494), ('\u{F8FD}', 494), ('\u{F8FE}', 494), ('\u{F8FF}', 790),
];
#[rustfmt::skip]
static ZAPFDINGBATS: &[(char, u16)] = &[
(' ', 278), ('\u{2192}', 838), ('\u{2194}', 1016), ('\u{2195}', 458), ('\u{2460}', 788), ('\u{2461}', 788),
('\u{2462}', 788), ('\u{2463}', 788), ('\u{2464}', 788), ('\u{2465}', 788), ('\u{2466}', 788), ('\u{2467}', 788),
('\u{2468}', 788), ('\u{2469}', 788), ('\u{25A0}', 761), ('\u{25B2}', 892), ('\u{25BC}', 892), ('\u{25C6}', 788),
('\u{25CF}', 791), ('\u{25D7}', 438), ('\u{2605}', 816), ('\u{260E}', 719), ('\u{261B}', 960), ('\u{261E}', 939),
('\u{2660}', 626), ('\u{2663}', 776), ('\u{2665}', 694), ('\u{2666}', 595), ('\u{2701}', 974), ('\u{2702}', 961),
('\u{2703}', 974), ('\u{2704}', 980), ('\u{2706}', 789), ('\u{2707}', 790), ('\u{2708}', 791), ('\u{2709}', 690),
('\u{270C}', 549), ('\u{270D}', 855), ('\u{270E}', 911), ('\u{270F}', 933), ('\u{2710}', 911), ('\u{2711}', 945),
('\u{2712}', 974), ('\u{2713}', 755), ('\u{2714}', 846), ('\u{2715}', 762), ('\u{2716}', 761), ('\u{2717}', 571),
('\u{2718}', 677), ('\u{2719}', 763), ('\u{271A}', 760), ('\u{271B}', 759), ('\u{271C}', 754), ('\u{271D}', 494),
('\u{271E}', 552), ('\u{271F}', 537), ('\u{2720}', 577), ('\u{2721}', 692), ('\u{2722}', 786), ('\u{2723}', 788),
('\u{2724}', 788), ('\u{2725}', 790), ('\u{2726}', 793), ('\u{2727}', 794), ('\u{2729}', 823), ('\u{272A}', 789),
('\u{272B}', 841), ('\u{272C}', 823), ('\u{272D}', 833), ('\u{272E}', 816), ('\u{272F}', 831), ('\u{2730}', 923),
('\u{2731}', 744), ('\u{2732}', 723), ('\u{2733}', 749), ('\u{2734}', 790), ('\u{2735}', 792), ('\u{2736}', 695),
('\u{2737}', 776), ('\u{2738}', 768), ('\u{2739}', 792), ('\u{273A}', 759), ('\u{273B}', 707), ('\u{273C}', 708),
('\u{273D}', 682), ('\u{273E}', 701), ('\u{273F}', 826), ('\u{2740}', 815), ('\u{2741}', 789), ('\u{2742}', 789),
('\u{2743}', 707), ('\u{2744}', 687), ('\u{2745}', 696), ('\u{2746}', 689), ('\u{2747}', 786), ('\u{2748}', 787),
('\u{2749}', 713), ('\u{274A}', 791), ('\u{274B}', 785), ('\u{274D}', 873), ('\u{274F}', 762), ('\u{2750}', 762),
('\u{2751}', 759), ('\u{2752}', 759), ('\u{2756}', 784), ('\u{2758}', 138), ('\u{2759}', 277), ('\u{275A}', 415),
('\u{275B}', 392), ('\u{275C}', 392), ('\u{275D}', 668), ('\u{275E}', 668), ('\u{2761}', 732), ('\u{2762}', 544),
('\u{2763}', 544), ('\u{2764}', 910), ('\u{2765}', 667), ('\u{2766}', 760), ('\u{2767}', 760), ('\u{2768}', 390),
('\u{2769}', 390), ('\u{276A}', 317), ('\u{276B}', 317), ('\u{276C}', 276), ('\u{276D}', 276), ('\u{276E}', 509),
('\u{276F}', 509), ('\u{2770}', 410), ('\u{2771}', 410), ('\u{2772}', 234), ('\u{2773}', 234), ('\u{2774}', 334),
('\u{2775}', 334), ('\u{2776}', 788), ('\u{2777}', 788), ('\u{2778}', 788), ('\u{2779}', 788), ('\u{277A}', 788),
('\u{277B}', 788), ('\u{277C}', 788), ('\u{277D}', 788), ('\u{277E}', 788), ('\u{277F}', 788), ('\u{2780}', 788),
('\u{2781}', 788), ('\u{2782}', 788), ('\u{2783}', 788), ('\u{2784}', 788), ('\u{2785}', 788), ('\u{2786}', 788),
('\u{2787}', 788), ('\u{2788}', 788), ('\u{2789}', 788), ('\u{278A}', 788), ('\u{278B}', 788), ('\u{278C}', 788),
('\u{278D}', 788), ('\u{278E}', 788), ('\u{278F}', 788), ('\u{2790}', 788), ('\u{2791}', 788), ('\u{2792}', 788),
('\u{2793}', 788), ('\u{2794}', 894), ('\u{2798}', 748), ('\u{2799}', 924), ('\u{279A}', 748), ('\u{279B}', 918),
('\u{279C}', 927), ('\u{279D}', 928), ('\u{279E}', 928), ('\u{279F}', 834), ('\u{27A0}', 873), ('\u{27A1}', 828),
('\u{27A2}', 924), ('\u{27A3}', 924), ('\u{27A4}', 917), ('\u{27A5}', 930), ('\u{27A6}', 931), ('\u{27A7}', 463),
('\u{27A8}', 883), ('\u{27A9}', 836), ('\u{27AA}', 836), ('\u{27AB}', 867), ('\u{27AC}', 867), ('\u{27AD}', 696),
('\u{27AE}', 696), ('\u{27AF}', 874), ('\u{27B1}', 874), ('\u{27B2}', 760), ('\u{27B3}', 946), ('\u{27B4}', 771),
('\u{27B5}', 865), ('\u{27B6}', 771), ('\u{27B7}', 888), ('\u{27B8}', 967), ('\u{27B9}', 888), ('\u{27BA}', 831),
('\u{27BB}', 873), ('\u{27BC}', 927), ('\u{27BD}', 970), ('\u{27BE}', 918),
];
#[rustfmt::skip]
static SYMBOL_ENCODING: &[(u8, char)] = &[
(0x20, ' '), (0x21, '!'), (0x22, '\u{2200}'), (0x23, '#'), (0x24, '\u{2203}'), (0x25, '%'),
(0x26, '&'), (0x27, '\u{220B}'), (0x28, '('), (0x29, ')'), (0x2A, '\u{2217}'), (0x2B, '+'),
(0x2C, ','), (0x2D, '\u{2212}'), (0x2E, '.'), (0x2F, '/'), (0x30, '0'), (0x31, '1'),
(0x32, '2'), (0x33, '3'), (0x34, '4'), (0x35, '5'), (0x36, '6'), (0x37, '7'),
(0x38, '8'), (0x39, '9'), (0x3A, ':'), (0x3B, ';'), (0x3C, '<'), (0x3D, '='),
(0x3E, '>'), (0x3F, '?'), (0x40, '\u{2245}'), (0x41, '\u{0391}'), (0x42, '\u{0392}'), (0x43, '\u{03A7}'),
(0x44, '\u{2206}'), (0x45, '\u{0395}'), (0x46, '\u{03A6}'), (0x47, '\u{0393}'), (0x48, '\u{0397}'), (0x49, '\u{0399}'),
(0x4A, '\u{03D1}'), (0x4B, '\u{039A}'), (0x4C, '\u{039B}'), (0x4D, '\u{039C}'), (0x4E, '\u{039D}'), (0x4F, '\u{039F}'),
(0x50, '\u{03A0}'), (0x51, '\u{0398}'), (0x52, '\u{03A1}'), (0x53, '\u{03A3}'), (0x54, '\u{03A4}'), (0x55, '\u{03A5}'),
(0x56, '\u{03C2}'), (0x57, '\u{2126}'), (0x58, '\u{039E}'), (0x59, '\u{03A8}'), (0x5A, '\u{0396}'), (0x5B, '['),
(0x5C, '\u{2234}'), (0x5D, ']'), (0x5E, '\u{22A5}'), (0x5F, '_'), (0x60, '\u{F8E5}'), (0x61, '\u{03B1}'),
(0x62, '\u{03B2}'), (0x63, '\u{03C7}'), (0x64, '\u{03B4}'), (0x65, '\u{03B5}'), (0x66, '\u{03C6}'), (0x67, '\u{03B3}'),
(0x68, '\u{03B7}'), (0x69, '\u{03B9}'), (0x6A, '\u{03D5}'), (0x6B, '\u{03BA}'), (0x6C, '\u{03BB}'), (0x6D, '\u{00B5}'),
(0x6E, '\u{03BD}'), (0x6F, '\u{03BF}'), (0x70, '\u{03C0}'), (0x71, '\u{03B8}'), (0x72, '\u{03C1}'), (0x73, '\u{03C3}'),
(0x74, '\u{03C4}'), (0x75, '\u{03C5}'), (0x76, '\u{03D6}'), (0x77, '\u{03C9}'), (0x78, '\u{03BE}'), (0x79, '\u{03C8}'),
(0x7A, '\u{03B6}'), (0x7B, '{'), (0x7C, '|'), (0x7D, '}'), (0x7E, '\u{223C}'), (0xA0, '\u{20AC}'),
(0xA1, '\u{03D2}'), (0xA2, '\u{2032}'), (0xA3, '\u{2264}'), (0xA4, '\u{2044}'), (0xA5, '\u{221E}'), (0xA6, '\u{0192}'),
(0xA7, '\u{2663}'), (0xA8, '\u{2666}'), (0xA9, '\u{2665}'), (0xAA, '\u{2660}'), (0xAB, '\u{2194}'), (0xAC, '\u{2190}'),
(0xAD, '\u{2191}'), (0xAE, '\u{2192}'), (0xAF, '\u{2193}'), (0xB0, '\u{00B0}'), (0xB1, '\u{00B1}'), (0xB2, '\u{2033}'),
(0xB3, '\u{2265}'), (0xB4, '\u{00D7}'), (0xB5, '\u{221D}'), (0xB6, '\u{2202}'), (0xB7, '\u{2022}'), (0xB8, '\u{00F7}'),
(0xB9, '\u{2260}'), (0xBA, '\u{2261}'), (0xBB, '\u{2248}'), (0xBC, '\u{2026}'), (0xBD, '\u{F8E6}'), (0xBE, '\u{F8E7}'),
(0xBF, '\u{21B5}'), (0xC0, '\u{2135}'), (0xC1, '\u{2111}'), (0xC2, '\u{211C}'), (0xC3, '\u{2118}'), (0xC4, '\u{2297}'),
(0xC5, '\u{2295}'), (0xC6, '\u{2205}'), (0xC7, '\u{2229}'), (0xC8, '\u{222A}'), (0xC9, '\u{2283}'), (0xCA, '\u{2287}'),
(0xCB, '\u{2284}'), (0xCC, '\u{2282}'), (0xCD, '\u{2286}'), (0xCE, '\u{2208}'), (0xCF, '\u{2209}'), (0xD0, '\u{2220}'),
(0xD1, '\u{2207}'), (0xD2, '\u{F6DA}'), (0xD3, '\u{F6D9}'), (0xD4, '\u{F6DB}'), (0xD5, '\u{220F}'), (0xD6, '\u{221A}'),
(0xD7, '\u{22C5}'), (0xD8, '\u{00AC}'), (0xD9, '\u{2227}'), (0xDA, '\u{2228}'), (0xDB, '\u{21D4}'), (0xDC, '\u{21D0}'),
(0xDD, '\u{21D1}'), (0xDE, '\u{21D2}'), (0xDF, '\u{21D3}'), (0xE0, '\u{25CA}'), (0xE1, '\u{2329}'), (0xE2, '\u{F8E8}'),
(0xE3, '\u{F8E9}'), (0xE4, '\u{F8EA}'), (0xE5, '\u{2211}'), (0xE6, '\u{F8EB}'), (0xE7, '\u{F8EC}'), (0xE8, '\u{F8ED}'),
(0xE9, '\u{F8EE}'), (0xEA, '\u{F8EF}'), (0xEB, '\u{F8F0}'), (0xEC, '\u{F8F1}'), (0xED, '\u{F8F2}'), (0xEE, '\u{F8F3}'),
(0xEF, '\u{F8F4}'), (0xF1, '\u{232A}'), (0xF2, '\u{222B}'), (0xF3, '\u{2320}'), (0xF4, '\u{F8F5}'), (0xF5, '\u{2321}'),
(0xF6, '\u{F8F6}'), (0xF7, '\u{F8F7}'), (0xF8, '\u{F8F8}'), (0xF9, '\u{F8F9}'), (0xFA, '\u{F8FA}'), (0xFB, '\u{F8FB}'),
(0xFC, '\u{F8FC}'), (0xFD, '\u{F8FD}'), (0xFE, '\u{F8FE}'),
];
#[rustfmt::skip]
static ZAPFDINGBATS_ENCODING: &[(u8, char)] = &[
(0x20, ' '), (0x21, '\u{2701}'), (0x22, '\u{2702}'), (0x23, '\u{2703}'), (0x24, '\u{2704}'), (0x25, '\u{260E}'),
(0x26, '\u{2706}'), (0x27, '\u{2707}'), (0x28, '\u{2708}'), (0x29, '\u{2709}'), (0x2A, '\u{261B}'), (0x2B, '\u{261E}'),
(0x2C, '\u{270C}'), (0x2D, '\u{270D}'), (0x2E, '\u{270E}'), (0x2F, '\u{270F}'), (0x30, '\u{2710}'), (0x31, '\u{2711}'),
(0x32, '\u{2712}'), (0x33, '\u{2713}'), (0x34, '\u{2714}'), (0x35, '\u{2715}'), (0x36, '\u{2716}'), (0x37, '\u{2717}'),
(0x38, '\u{2718}'), (0x39, '\u{2719}'), (0x3A, '\u{271A}'), (0x3B, '\u{271B}'), (0x3C, '\u{271C}'), (0x3D, '\u{271D}'),
(0x3E, '\u{271E}'), (0x3F, '\u{271F}'), (0x40, '\u{2720}'), (0x41, '\u{2721}'), (0x42, '\u{2722}'), (0x43, '\u{2723}'),
(0x44, '\u{2724}'), (0x45, '\u{2725}'), (0x46, '\u{2726}'), (0x47, '\u{2727}'), (0x48, '\u{2605}'), (0x49, '\u{2729}'),
(0x4A, '\u{272A}'), (0x4B, '\u{272B}'), (0x4C, '\u{272C}'), (0x4D, '\u{272D}'), (0x4E, '\u{272E}'), (0x4F, '\u{272F}'),
(0x50, '\u{2730}'), (0x51, '\u{2731}'), (0x52, '\u{2732}'), (0x53, '\u{2733}'), (0x54, '\u{2734}'), (0x55, '\u{2735}'),
(0x56, '\u{2736}'), (0x57, '\u{2737}'), (0x58, '\u{2738}'), (0x59, '\u{2739}'), (0x5A, '\u{273A}'), (0x5B, '\u{273B}'),
(0x5C, '\u{273C}'), (0x5D, '\u{273D}'), (0x5E, '\u{273E}'), (0x5F, '\u{273F}'), (0x60, '\u{2740}'), (0x61, '\u{2741}'),
(0x62, '\u{2742}'), (0x63, '\u{2743}'), (0x64, '\u{2744}'), (0x65, '\u{2745}'), (0x66, '\u{2746}'), (0x67, '\u{2747}'),
(0x68, '\u{2748}'), (0x69, '\u{2749}'), (0x6A, '\u{274A}'), (0x6B, '\u{274B}'), (0x6C, '\u{25CF}'), (0x6D, '\u{274D}'),
(0x6E, '\u{25A0}'), (0x6F, '\u{274F}'), (0x70, '\u{2750}'), (0x71, '\u{2751}'), (0x72, '\u{2752}'), (0x73, '\u{25B2}'),
(0x74, '\u{25BC}'), (0x75, '\u{25C6}'), (0x76, '\u{2756}'), (0x77, '\u{25D7}'), (0x78, '\u{2758}'), (0x79, '\u{2759}'),
(0x7A, '\u{275A}'), (0x7B, '\u{275B}'), (0x7C, '\u{275C}'), (0x7D, '\u{275D}'), (0x7E, '\u{275E}'), (0x80, '\u{2768}'),
(0x81, '\u{2769}'), (0x82, '\u{276A}'), (0x83, '\u{276B}'), (0x84, '\u{276C}'), (0x85, '\u{276D}'), (0x86, '\u{276E}'),
(0x87, '\u{276F}'), (0x88, '\u{2770}'), (0x89, '\u{2771}'), (0x8A, '\u{2772}'), (0x8B, '\u{2773}'), (0x8C, '\u{2774}'),
(0x8D, '\u{2775}'), (0xA1, '\u{2761}'), (0xA2, '\u{2762}'), (0xA3, '\u{2763}'), (0xA4, '\u{2764}'), (0xA5, '\u{2765}'),
(0xA6, '\u{2766}'), (0xA7, '\u{2767}'), (0xA8, '\u{2663}'), (0xA9, '\u{2666}'), (0xAA, '\u{2665}'), (0xAB, '\u{2660}'),
(0xAC, '\u{2460}'), (0xAD, '\u{2461}'), (0xAE, '\u{2462}'), (0xAF, '\u{2463}'), (0xB0, '\u{2464}'), (0xB1, '\u{2465}'),
(0xB2, '\u{2466}'), (0xB3, '\u{2467}'), (0xB4, '\u{2468}'), (0xB5, '\u{2469}'), (0xB6, '\u{2776}'), (0xB7, '\u{2777}'),
(0xB8, '\u{2778}'), (0xB9, '\u{2779}'), (0xBA, '\u{277A}'), (0xBB, '\u{277B}'), (0xBC, '\u{277C}'), (0xBD, '\u{277D}'),
(0xBE, '\u{277E}'), (0xBF, '\u{277F}'), (0xC0, '\u{2780}'), (0xC1, '\u{2781}'), (0xC2, '\u{2782}'), (0xC3, '\u{2783}'),
(0xC4, '\u{2784}'), (0xC5, '\u{2785}'), (0xC6, '\u{2786}'), (0xC7, '\u{2787}'), (0xC8, '\u{2788}'), (0xC9, '\u{2789}'),
(0xCA, '\u{278A}'), (0xCB, '\u{278B}'), (0xCC, '\u{278C}'), (0xCD, '\u{278D}'), (0xCE, '\u{278E}'), (0xCF, '\u{278F}'),
(0xD0, '\u{2790}'), (0xD1, '\u{2791}'), (0xD2, '\u{2792}'), (0xD3, '\u{2793}'), (0xD4, '\u{2794}'), (0xD5, '\u{2192}'),
(0xD6, '\u{2194}'), (0xD7, '\u{2195}'), (0xD8, '\u{2798}'), (0xD9, '\u{2799}'), (0xDA, '\u{279A}'), (0xDB, '\u{279B}'),
(0xDC, '\u{279C}'), (0xDD, '\u{279D}'), (0xDE, '\u{279E}'), (0xDF, '\u{279F}'), (0xE0, '\u{27A0}'), (0xE1, '\u{27A1}'),
(0xE2, '\u{27A2}'), (0xE3, '\u{27A3}'), (0xE4, '\u{27A4}'), (0xE5, '\u{27A5}'), (0xE6, '\u{27A6}'), (0xE7, '\u{27A7}'),
(0xE8, '\u{27A8}'), (0xE9, '\u{27A9}'), (0xEA, '\u{27AA}'), (0xEB, '\u{27AB}'), (0xEC, '\u{27AC}'), (0xED, '\u{27AD}'),
(0xEE, '\u{27AE}'), (0xEF, '\u{27AF}'), (0xF1, '\u{27B1}'), (0xF2, '\u{27B2}'), (0xF3, '\u{27B3}'), (0xF4, '\u{27B4}'),
(0xF5, '\u{27B5}'), (0xF6, '\u{27B6}'), (0xF7, '\u{27B7}'), (0xF8, '\u{27B8}'), (0xF9, '\u{27B9}'), (0xFA, '\u{27BA}'),
(0xFB, '\u{27BB}'), (0xFC, '\u{27BC}'), (0xFD, '\u{27BD}'), (0xFE, '\u{27BE}'),
];
/// Every width table, for exhaustive testing.
#[cfg(test)]
static ALL_TABLES: &[(&str, &[(char, u16)])] = &[
("COURIER", COURIER),
("COURIER_BOLD", COURIER_BOLD),
("COURIER_OBLIQUE", COURIER_OBLIQUE),
("COURIER_BOLDOBLIQUE", COURIER_BOLDOBLIQUE),
("HELVETICA", HELVETICA),
("HELVETICA_BOLD", HELVETICA_BOLD),
("HELVETICA_OBLIQUE", HELVETICA_OBLIQUE),
("HELVETICA_BOLDOBLIQUE", HELVETICA_BOLDOBLIQUE),
("TIMES_ROMAN", TIMES_ROMAN),
("TIMES_BOLD", TIMES_BOLD),
("TIMES_ITALIC", TIMES_ITALIC),
("TIMES_BOLDITALIC", TIMES_BOLDITALIC),
("SYMBOL", SYMBOL),
("ZAPFDINGBATS", ZAPFDINGBATS),
];
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn times_roman_ascii_widths() {
assert_eq!(base14_char_width("Times-Roman", ' '), Some(250));
assert_eq!(base14_char_width("Times-Roman", 'M'), Some(889));
assert_eq!(base14_char_width("Times-Roman", 'i'), Some(278));
}
#[test]
fn subset_prefix_and_aliases_normalize() {
assert_eq!(
base14_char_width("ABCDEF+Times-Bold", ' '),
base14_char_width("Times-Bold", ' ')
);
assert!(base14_char_width("ArialMT", 'a').is_some());
assert!(base14_char_width("TimesNewRomanPSMT", 'a').is_some());
}
#[test]
fn non_base14_returns_none() {
assert_eq!(base14_char_width("DejaVuSans", 'a'), None);
assert!(!is_base14_font("Garamond"));
}
#[test]
fn builtin_encoding_resolves_symbol_and_zapf_codes() {
// Symbol 0x61 renders alpha; Zapf 0x21 renders U+2701.
assert_eq!(builtin_encoding_char("Symbol", 0x61), Some('\u{03B1}'));
assert_eq!(builtin_encoding_char("Symbol", 0xA5), Some('\u{221E}'));
assert_eq!(
builtin_encoding_char("ZapfDingbats", 0x21),
Some('\u{2701}')
);
// Latin text fonts follow standard encodings — no builtin override.
assert_eq!(builtin_encoding_char("Times-Roman", 0x61), None);
// The resolved chars have real AFM widths.
let alpha_w = base14_char_width("Symbol", '\u{03B1}');
assert!(alpha_w.is_some() && alpha_w != Some(500));
}
#[test]
fn encoding_tables_are_sorted_for_binary_search() {
for table in [SYMBOL_ENCODING, ZAPFDINGBATS_ENCODING] {
assert!(table.windows(2).all(|w| w[0].0 < w[1].0));
}
}
#[test]
fn tables_are_sorted_for_binary_search() {
// Every table is queried by binary search, so all of them must be
// sorted — not just a sample.
for (name, table) in ALL_TABLES {
assert!(
table.windows(2).all(|w| w[0].0 < w[1].0),
"{name} is not sorted"
);
}
}
}
+610 -40
View File
@@ -14,9 +14,11 @@ use lopdf::{Document, Encoding, Object, ObjectId};
use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
build_font_encodings, build_font_widths, build_type3_scales, compute_string_width_ts,
descriptor_style_flags, extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes,
CMapDecisionCache, FontStyleCache,
};
use super::underline::UnderlineLine;
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
@@ -39,6 +41,17 @@ fn strip_pdf_comments(data: &[u8]) -> Vec<u8> {
while i < data.len() {
let b = data[i];
match b {
// Inside a string literal, a backslash escapes the next byte —
// `\(`, `\)`, and `\\` must not touch the nesting depth, or a
// later `%` glyph inside a string gets stripped as a comment,
// corrupting the stream.
b'\\' if in_string > 0 => {
result.push(b);
if let Some(&next) = data.get(i + 1) {
result.push(next);
i += 1;
}
}
b'(' if !in_hex_string => {
in_string += 1;
result.push(b);
@@ -74,6 +87,56 @@ fn strip_pdf_comments(data: &[u8]) -> Vec<u8> {
result
}
fn transform_path_point(x: f32, y: f32, ctm: &[f32; 6]) -> (f32, f32) {
(
x * ctm[0] + y * ctm[2] + ctm[4],
x * ctm[1] + y * ctm[3] + ctm[5],
)
}
fn transformed_stroke_width(
line_width: f32,
ctm: &[f32; 6],
x1: f32,
y1: f32,
x2: f32,
y2: f32,
) -> f32 {
let user_width = line_width.abs();
let dx = x2 - x1;
let dy = y2 - y1;
let len = (dx * dx + dy * dy).sqrt();
if len <= f32::EPSILON {
return user_width;
}
// PDF stroke width scales perpendicular to the path direction.
let nx = -dy / len;
let ny = dx / len;
let ndx = nx * ctm[0] + ny * ctm[2];
let ndy = nx * ctm[1] + ny * ctm[3];
user_width * (ndx * ndx + ndy * ndy).sqrt()
}
/// Text rise (Ts) displaces the glyph origin by (0, rise) in unscaled text
/// space — per the rendering-matrix definition it sits left of Tm, so the
/// offset maps through the text matrix's y column. Rise never contributes
/// to the advance, so callers apply it only to the rendering position and
/// keep advancing the unshifted text matrix.
fn rise_adjusted(tm: &[f32; 6], rise: f32) -> [f32; 6] {
if rise == 0.0 {
return *tm;
}
[
tm[0],
tm[1],
tm[2],
tm[3],
tm[4] + rise * tm[2],
tm[5] + rise * tm[3],
]
}
/// Returns `(page_extraction, has_gid_fonts)` where `has_gid_fonts` indicates
/// the page uses fonts with unresolvable gid-encoded glyphs.
pub(crate) fn extract_page_text_items(
@@ -82,6 +145,7 @@ pub(crate) fn extract_page_text_items(
page_num: u32,
font_cmaps: &FontCMaps,
include_invisible: bool,
style_cache: &mut FontStyleCache,
) -> Result<(PageExtraction, bool, bool), PdfError> {
use lopdf::content::Content;
@@ -89,6 +153,7 @@ pub(crate) fn extract_page_text_items(
let mut rects: Vec<PdfRect> = Vec::new();
let mut clip_rects: Vec<PdfRect> = Vec::new();
let mut lines: Vec<PdfLine> = Vec::new();
let mut underline_lines: Vec<UnderlineLine> = Vec::new();
// Path construction state for m/l/h → S/s line extraction
let mut path_subpath_start: Option<(f32, f32)> = None;
@@ -97,15 +162,22 @@ pub(crate) fn extract_page_text_items(
// Completed subpaths (each a vec of line segments) for f/f* rect extraction
let mut pending_subpaths: Vec<Vec<(f32, f32, f32, f32)>> = Vec::new();
let mut fill_rects: Vec<PdfRect> = Vec::new();
// `re` rects awaiting a paint operator. Underline detection must only
// see painted rects: a `re W n` clip path or `re n` no-op draws nothing
// on the page, so treating every `re` as ink would underline text that
// merely sits near an invisible clip boundary.
let mut pending_re_rects: Vec<PdfRect> = Vec::new();
let mut painted_rects: Vec<PdfRect> = Vec::new();
// Get fonts for encoding
let fonts = doc.get_page_fonts(page_id).unwrap_or_default();
// Build font encoding maps from Differences arrays
let (font_encodings, has_gid_fonts) = build_font_encodings(doc, &fonts);
let (font_encodings, has_gid_fonts) = build_font_encodings(doc, &fonts, font_cmaps);
// Build font width info for accurate text positioning
let font_widths = build_font_widths(doc, &fonts);
let type3_scales = build_type3_scales(doc, &fonts);
// Build maps of font resource names to their base font names and ToUnicode object refs
let mut font_base_names: std::collections::HashMap<String, String> =
@@ -114,6 +186,8 @@ pub(crate) fn extract_page_text_items(
std::collections::HashMap::new();
let mut inline_cmaps: std::collections::HashMap<String, crate::tounicode::CMapEntry> =
std::collections::HashMap::new();
let mut font_style_flags: std::collections::HashMap<String, (bool, bool)> =
std::collections::HashMap::new();
for (font_name, font_dict) in &fonts {
let resource_name = String::from_utf8_lossy(font_name).to_string();
if let Ok(base_font) = font_dict.get(b"BaseFont") {
@@ -122,6 +196,12 @@ pub(crate) fn extract_page_text_items(
font_base_names.insert(resource_name.clone(), base_name);
}
}
// Descriptor style flags rescue subset fonts whose BaseFont names
// are opaque tags the name heuristics can't read.
let style = descriptor_style_flags(doc, font_dict, style_cache);
if style != (false, false) {
font_style_flags.insert(resource_name.clone(), style);
}
// Track ToUnicode object reference, with FontFile2 fallback for Identity-H/V.
// Also handle inline ToUnicode streams.
match font_dict.get(b"ToUnicode") {
@@ -188,7 +268,20 @@ pub(crate) fn extract_page_text_items(
// Graphics state tracking
let mut ctm = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0]; // Current Transformation Matrix
let mut text_rendering_mode: i32 = 0; // 0=fill, 1=stroke, 2=fill+stroke, 3=invisible
let mut gstate_stack: Vec<([f32; 6], i32, f32, f32)> = Vec::new();
let mut line_width: f32 = 1.0;
#[derive(Clone)]
struct SavedGraphicsState {
ctm: [f32; 6],
text_rendering_mode: i32,
line_width: f32,
char_spacing: f32,
word_spacing: f32,
text_rise: f32,
text_leading: f32,
current_font: String,
current_font_size: f32,
}
let mut gstate_stack: Vec<SavedGraphicsState> = Vec::new();
// Text state tracking
let mut current_font = String::new();
@@ -196,6 +289,7 @@ pub(crate) fn extract_page_text_items(
let mut text_leading: f32 = 0.0; // TL parameter (in text-space units)
let mut char_spacing: f32 = 0.0; // Tc parameter (extra spacing per character, unscaled)
let mut word_spacing: f32 = 0.0; // Tw parameter (extra spacing per space char, unscaled)
let mut text_rise: f32 = 0.0; // Ts parameter (baseline shift for super/subscripts, unscaled)
let mut text_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut line_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut in_text_block = false;
@@ -217,6 +311,10 @@ pub(crate) fn extract_page_text_items(
let mut suppress_glyph_extraction = false;
let mut actual_text_start_tm: Option<[f32; 6]> = None; // text matrix at BDC entry
let mut actual_text_glyph_tm: Option<[f32; 6]> = None; // text matrix at first glyph inside BDC
// Text rise in effect at each captured matrix — the item must render at
// the rise of its GLYPHS, not whatever rise is set by EMC time.
let mut actual_text_start_rise: f32 = 0.0;
let mut actual_text_glyph_rise: Option<f32> = None;
/// Get the innermost MCID from the marked content stack.
fn current_mcid(stack: &[MarkedContentEntry]) -> Option<i64> {
stack.iter().rev().find_map(|e| e.mcid)
@@ -227,15 +325,30 @@ pub(crate) fn extract_page_text_items(
match op.operator.as_str() {
"q" => {
// Save graphics state
gstate_stack.push((ctm, text_rendering_mode, char_spacing, word_spacing));
gstate_stack.push(SavedGraphicsState {
ctm,
text_rendering_mode,
line_width,
char_spacing,
word_spacing,
text_rise,
text_leading,
current_font: current_font.clone(),
current_font_size,
});
}
"Q" => {
// Restore graphics state
if let Some((saved_ctm, saved_tr, saved_tc, saved_tw)) = gstate_stack.pop() {
ctm = saved_ctm;
text_rendering_mode = saved_tr;
char_spacing = saved_tc;
word_spacing = saved_tw;
if let Some(saved) = gstate_stack.pop() {
ctm = saved.ctm;
text_rendering_mode = saved.text_rendering_mode;
line_width = saved.line_width;
char_spacing = saved.char_spacing;
word_spacing = saved.word_spacing;
text_rise = saved.text_rise;
text_leading = saved.text_leading;
current_font = saved.current_font;
current_font_size = saved.current_font_size;
}
}
"cm" => {
@@ -252,6 +365,11 @@ pub(crate) fn extract_page_text_items(
ctm = multiply_matrices(&new_matrix, &ctm);
}
}
"w" => {
if let Some(width) = op.operands.first().and_then(get_number) {
line_width = width;
}
}
"BT" => {
// Begin text block
in_text_block = true;
@@ -300,6 +418,12 @@ pub(crate) fn extract_page_text_items(
word_spacing = tw;
}
}
"Ts" => {
// Set text rise (baseline shift for superscripts/subscripts)
if let Some(ts) = op.operands.first().and_then(get_number) {
text_rise = ts;
}
}
"Td" | "TD" => {
// Move text position: TLM = T(tx,ty) × TLM; Tm = TLM
// tx,ty are in text space — must be scaled by the text line matrix
@@ -358,6 +482,7 @@ pub(crate) fn extract_page_text_items(
if suppress_glyph_extraction {
if actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
@@ -387,8 +512,10 @@ pub(crate) fn extract_page_text_items(
&mut cmap_decisions,
&font_widths,
) {
let combined = multiply_matrices(&text_matrix, &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
let combined =
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
@@ -409,6 +536,10 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
@@ -418,8 +549,10 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -437,6 +570,7 @@ pub(crate) fn extract_page_text_items(
// Capture first-glyph position for ActualText
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
// Compute space threshold based on font metrics when available
@@ -552,11 +686,16 @@ pub(crate) fn extract_page_text_items(
} else {
rotation_votes.rotated += 1;
}
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let base_font = font_base_names
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
let scale_x = text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2];
for (text, start_w, end_w) in &sub_items {
let offset_tm = [
@@ -567,7 +706,8 @@ pub(crate) fn extract_page_text_items(
text_matrix[4] + start_w * text_matrix[0],
text_matrix[5] + start_w * text_matrix[1],
];
let combined = multiply_matrices(&offset_tm, &ctm);
let combined =
multiply_matrices(&rise_adjusted(&offset_tm, text_rise), &ctm);
let (x, y) = (combined[4], combined[5]);
let width = if font_info.is_some() {
((end_w - start_w) * scale_x).abs()
@@ -583,8 +723,10 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -608,6 +750,26 @@ pub(crate) fn extract_page_text_items(
line_matrix[4] += (-tl) * line_matrix[2];
line_matrix[5] += (-tl) * line_matrix[3];
text_matrix = line_matrix;
// Capture first-glyph position for ActualText AFTER the
// line move — the BDC-entry matrix is on the previous line.
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
// Advance width, as for Tj — without it the item stays
// zero-width and geometric underline/strikeout detection
// rejects it (`is_underline_candidate` needs width > 0).
let w_ts_opt = font_widths.get(&current_font).and_then(|fi| {
op.operands.first().and_then(get_operand_bytes).map(|raw| {
compute_string_width_ts(
raw,
fi,
current_font_size,
char_spacing,
word_spacing,
)
})
});
if !((text_rendering_mode == 3 && !include_invisible)
|| suppress_glyph_extraction
|| op.operands.is_empty())
@@ -625,35 +787,55 @@ pub(crate) fn extract_page_text_items(
&font_widths,
) {
if !text.trim().is_empty() {
let combined = multiply_matrices(&text_matrix, &ctm);
let combined =
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
} else {
rotation_votes.rotated += 1;
}
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
let width = w_ts_opt
.map(|w_ts| {
(w_ts * (text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2]))
.abs()
})
.unwrap_or(0.0);
let base_font = font_base_names
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
y,
width: 0.0,
width,
height: rendered_size,
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
}
}
}
// Advance regardless of visibility so later show-text
// operators on the same line stay positioned (as for Tj).
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
}
}
"Do" => {
// XObject invocation - could be an image or form
@@ -684,6 +866,8 @@ pub(crate) fn extract_page_text_items(
page: page_num,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Image,
mcid: current_mcid(&marked_content_stack),
});
@@ -697,6 +881,7 @@ pub(crate) fn extract_page_text_items(
font_cmaps,
&ctm,
&mut cmap_decisions,
style_cache,
);
items.extend(form_items);
}
@@ -737,7 +922,9 @@ pub(crate) fn extract_page_text_items(
if actual_text.is_some() {
suppress_glyph_extraction = true;
actual_text_start_tm = Some(text_matrix);
actual_text_start_rise = text_rise;
actual_text_glyph_tm = None; // reset — will be captured at first Tj/TJ
actual_text_glyph_rise = None;
}
marked_content_stack.push(MarkedContentEntry { actual_text, mcid });
}
@@ -750,15 +937,18 @@ pub(crate) fn extract_page_text_items(
// Tj may have moved the text position to the correct line —
// the BDC-entry position can be on the previous line.
let glyph_tm = actual_text_glyph_tm.take();
let glyph_rise = actual_text_glyph_rise.take();
let entry_tm = actual_text_start_tm.take();
if let Some(start_tm) = glyph_tm.or(entry_tm) {
let combined = multiply_matrices(&start_tm, &ctm);
let rise = glyph_rise.unwrap_or(actual_text_start_rise);
let combined = multiply_matrices(&rise_adjusted(&start_tm, rise), &ctm);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
} else {
rotation_votes.rotated += 1;
}
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
// Width in device space from text matrix delta
let delta_ts = text_matrix[4] - start_tm[4];
@@ -769,6 +959,10 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&at),
x,
@@ -778,8 +972,10 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: entry
.mcid
@@ -804,13 +1000,19 @@ pub(crate) fn extract_page_text_items(
let y_dev = rx * ctm[1] + ry * ctm[3] + ctm[5];
let w_dev = rw * ctm[0];
let h_dev = rh * ctm[3];
rects.push(PdfRect {
let rect = PdfRect {
x: x_dev,
y: y_dev,
width: w_dev,
height: h_dev,
page: page_num,
});
};
// Underline detection must only see rects that are
// actually painted — a `re` used purely as a clip path
// (`re W n`) or discarded (`re n`) draws nothing. Hold
// the rect as pending until a paint operator confirms it.
pending_re_rects.push(rect.clone());
rects.push(rect);
}
}
// ── Path construction operators ──────────────────────
@@ -860,10 +1062,8 @@ pub(crate) fn extract_page_text_items(
}
}
for (x1, y1, x2, y2) in pending_lines.drain(..) {
let x1d = x1 * ctm[0] + y1 * ctm[2] + ctm[4];
let y1d = x1 * ctm[1] + y1 * ctm[3] + ctm[5];
let x2d = x2 * ctm[0] + y2 * ctm[2] + ctm[4];
let y2d = x2 * ctm[1] + y2 * ctm[3] + ctm[5];
let (x1d, y1d) = transform_path_point(x1, y1, &ctm);
let (x2d, y2d) = transform_path_point(x2, y2, &ctm);
lines.push(PdfLine {
x1: x1d,
y1: y1d,
@@ -871,7 +1071,16 @@ pub(crate) fn extract_page_text_items(
y2: y2d,
page: page_num,
});
underline_lines.push(UnderlineLine {
x1: x1d,
y1: y1d,
x2: x2d,
y2: y2d,
stroke_width: transformed_stroke_width(line_width, &ctm, x1, y1, x2, y2),
page: page_num,
});
}
painted_rects.append(&mut pending_re_rects);
pending_subpaths.clear();
path_subpath_start = None;
path_current = None;
@@ -887,10 +1096,8 @@ pub(crate) fn extract_page_text_items(
}
}
for (x1, y1, x2, y2) in pending_lines.drain(..) {
let x1d = x1 * ctm[0] + y1 * ctm[2] + ctm[4];
let y1d = x1 * ctm[1] + y1 * ctm[3] + ctm[5];
let x2d = x2 * ctm[0] + y2 * ctm[2] + ctm[4];
let y2d = x2 * ctm[1] + y2 * ctm[3] + ctm[5];
let (x1d, y1d) = transform_path_point(x1, y1, &ctm);
let (x2d, y2d) = transform_path_point(x2, y2, &ctm);
lines.push(PdfLine {
x1: x1d,
y1: y1d,
@@ -898,7 +1105,16 @@ pub(crate) fn extract_page_text_items(
y2: y2d,
page: page_num,
});
underline_lines.push(UnderlineLine {
x1: x1d,
y1: y1d,
x2: x2d,
y2: y2d,
stroke_width: transformed_stroke_width(line_width, &ctm, x1, y1, x2, y2),
page: page_num,
});
}
painted_rects.append(&mut pending_re_rects);
pending_subpaths.clear();
path_subpath_start = None;
path_current = None;
@@ -956,6 +1172,7 @@ pub(crate) fn extract_page_text_items(
}
}
}
painted_rects.append(&mut pending_re_rects);
pending_lines.clear();
path_subpath_start = None;
path_current = None;
@@ -1021,7 +1238,10 @@ pub(crate) fn extract_page_text_items(
// Do NOT clear pending_lines — the following `n` does that
}
"n" => {
// end path (no-op): discard
// end path (no-op): discard — including any `re` rects that
// were only ever part of a clip path (`re W n`), which draw
// no ink and must not feed underline detection.
pending_re_rects.clear();
pending_lines.clear();
pending_subpaths.clear();
path_subpath_start = None;
@@ -1031,6 +1251,12 @@ pub(crate) fn extract_page_text_items(
}
}
// Underline detection reads only painted ink: `re` rects confirmed by
// a paint operator plus filled-subpath rects — never clip-only rects,
// which draw nothing.
let mut underline_rects = painted_rects;
underline_rects.extend(fill_rects.iter().cloned());
// Only use clip/fill rects when no `re` rects exist on this page.
// Clip rects take priority over fill rects, but first we deduplicate
// them: some PDFs wrap every text block in a full-page W* clip path,
@@ -1060,8 +1286,17 @@ pub(crate) fn extract_page_text_items(
// Some PDFs embed landscape content in portrait pages using a rotated text
// matrix (e.g. [0, b, -b, 0, tx, ty] for 90° CCW). The layout engine
// assumes x=horizontal, y=vertical — so we swap coordinates to match.
let (items, rects, lines, coords_rotated) =
let (mut items, rects, lines, coords_rotated) =
correct_rotated_page(items, rects, lines, &rotation_votes);
if coords_rotated {
rotate_underline_graphics(&mut underline_rects, &mut underline_lines);
}
super::underline::mark_underlined_items(
&mut items,
&underline_rects,
&underline_lines,
page_num,
);
let items = super::merge_text_items(items);
let items = super::merge_subscript_items(items);
@@ -1147,6 +1382,27 @@ fn correct_rotated_page(
(items, rects, lines, true)
}
fn rotate_underline_graphics(rects: &mut [PdfRect], lines: &mut [UnderlineLine]) {
for rect in rects {
let new_x = rect.y;
let new_y = -(rect.x + rect.width.abs());
rect.x = new_x;
rect.y = new_y;
std::mem::swap(&mut rect.width, &mut rect.height);
}
for line in lines {
let new_x1 = line.y1;
let new_y1 = -line.x1;
let new_x2 = line.y2;
let new_y2 = -line.x2;
line.x1 = new_x1;
line.y1 = new_y1;
line.x2 = new_x2;
line.y2 = new_y2;
}
}
/// Remove near-duplicate rects (same coordinates within 0.5 pt tolerance).
/// Some PDFs emit a full-page clip path for every text block, producing
/// thousands of identical rects. After dedup these collapse to one rect,
@@ -1196,6 +1452,64 @@ mod tests {
}
}
fn simple_doc_with_content(content: &[u8]) -> (lopdf::Document, lopdf::ObjectId) {
use lopdf::{dictionary, Object, Stream};
let mut doc = lopdf::Document::new();
let widths: Vec<Object> = (0..=255).map(|_| 600.into()).collect();
let font_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
"FirstChar" => 0,
"LastChar" => 255,
"Widths" => Object::Array(widths),
});
let content_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
content.to_vec(),
)));
let page_id = doc.add_object(dictionary! {
"Type" => "Page",
"Contents" => Object::Reference(content_id),
"Resources" => dictionary! {
"Font" => dictionary! {
"F1" => Object::Reference(font_id),
},
},
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
});
let pages_id = doc.add_object(dictionary! {
"Type" => "Pages",
"Count" => Object::Integer(1),
"Kids" => vec![Object::Reference(page_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => Object::Reference(pages_id),
});
doc.trailer.set("Root", Object::Reference(catalog_id));
(doc, page_id)
}
fn extract_simple_items(content: &[u8]) -> Vec<TextItem> {
use crate::tounicode::FontCMaps;
let (doc, page_id) = simple_doc_with_content(content);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
items
}
#[test]
fn test_dedup_rects_identical() {
let mut rects = vec![rect(0.0, 0.0, 612.0, 792.0, 1); 3759];
@@ -1245,6 +1559,140 @@ mod tests {
assert_eq!(single.len(), 1);
}
#[test]
fn thick_stroked_rule_does_not_mark_underline() {
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm (THICK) Tj ET
4 w
100 498 m 170 498 l S
BT /F1 12 Tf 1 0 0 1 100 480 Tm (THIN) Tj ET
1 w
100 478 m 160 478 l S";
let items = extract_simple_items(content);
let thick = items.iter().find(|item| item.text == "THICK").unwrap();
let thin = items.iter().find(|item| item.text == "THIN").unwrap();
assert!(!thick.is_underline);
assert!(thin.is_underline);
}
#[test]
fn rotated_page_underline_is_detected_after_coordinate_correction() {
let content = b"BT /F1 12 Tf 0 1 -1 0 200 100 Tm (HELLO) Tj ET
BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
1 w
202 100 m 202 170 l S";
let items = extract_simple_items(content);
let hello = items.iter().find(|item| item.text == "HELLO").unwrap();
let world = items.iter().find(|item| item.text == "WORLD").unwrap();
assert!(hello.is_underline);
assert!(!world.is_underline);
}
#[test]
fn quote_operator_text_carries_advance_width() {
// `'` (move-to-next-line-and-show-text) must retain the string's
// advance width like Tj — zero-width items are invisible to
// geometric underline/strikeout detection.
let content = b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (first) Tj (struck) ' ET
1 w
99 503 m 145 503 l S";
let items = extract_simple_items(content);
let struck = items.iter().find(|item| item.text == "struck").unwrap();
// 6 glyphs x 600/1000 x 12pt = 43.2pt, drawn one leading below Tm.
assert!((struck.width - 43.2).abs() < 0.1);
assert!((struck.y - 500.0).abs() < 0.1);
assert!(struck.is_strikeout);
assert!(!struck.is_underline);
}
#[test]
fn quote_operator_advances_text_matrix() {
// Text shown after `'` on the same line must start past the shown
// string: "CD" lands at x=114.4 (2 glyphs x 600/1000 x 12pt after
// x=100), flush against "AB", so the merge pass joins them. Without
// the advance "CD" overlaps "AB" at x=100 and the items stay apart.
let content = b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (AB) ' (CD) Tj ET";
let items = extract_simple_items(content);
let merged = items.iter().find(|item| item.text == "ABCD").unwrap();
assert!((merged.x - 100.0).abs() < 0.1);
assert!((merged.width - 28.8).abs() < 0.1);
assert!((merged.y - 500.0).abs() < 0.1);
}
#[test]
fn text_rise_shifts_item_baseline() {
// Ts displaces the glyph origin vertically without touching the
// advance; the next run at rise 0 must return to the original
// baseline and follow the raised run horizontally.
let content =
b"BT /F1 12 Tf 1 0 0 1 100 500 Tm (base) Tj 5 Ts (super) Tj 0 Ts (after) Tj ET";
let items = extract_simple_items(content);
let base = items.iter().find(|item| item.text == "base").unwrap();
let raised = items.iter().find(|item| item.text == "super").unwrap();
let after = items.iter().find(|item| item.text == "after").unwrap();
assert!((base.y - 500.0).abs() < 0.1);
assert!((raised.y - 505.0).abs() < 0.1);
assert!((after.y - 500.0).abs() < 0.1);
assert!(after.x > raised.x);
}
#[test]
fn actual_text_item_uses_glyph_rise() {
// The ActualText replacement item must render at the rise in
// effect when its glyphs were drawn — not the unshifted BDC
// baseline, and not whatever rise is set by EMC time.
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm \
/Span <</ActualText (super) >> BDC 5 Ts (sup) Tj 0 Ts EMC (after) Tj ET";
let items = extract_simple_items(content);
let sup = items.iter().find(|item| item.text == "super").unwrap();
let after = items.iter().find(|item| item.text == "after").unwrap();
assert!((sup.y - 505.0).abs() < 0.1);
assert!((after.y - 500.0).abs() < 0.1);
}
#[test]
fn actual_text_shown_with_quote_op_uses_moved_risen_baseline() {
// When the tagged span's show op is `'`, the glyph position is
// only known AFTER its line move — falling back to the BDC-entry
// matrix would place the item on the previous line, unrisen.
let content = b"BT /F1 12 Tf 14 TL 1 0 0 1 100 500 Tm \
/Span <</ActualText (replaced) >> BDC 3 Ts (raw) ' 0 Ts EMC ET";
let items = extract_simple_items(content);
let item = items.iter().find(|item| item.text == "replaced").unwrap();
// Line move: 500 - 14 = 486; rise: +3 -> 489.
assert!((item.y - 489.0).abs() < 0.1);
assert!(item.width > 0.0);
}
#[test]
fn strikeout_detected_on_risen_text() {
// The rule crosses the glyphs at their risen position; without the
// rise in item.y the strike window sits 4pt too low and misses.
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm 4 Ts (struck) Tj ET
1 w
99 507 m 145 507 l S";
let items = extract_simple_items(content);
let struck = items.iter().find(|item| item.text == "struck").unwrap();
assert!((struck.y - 504.0).abs() < 0.1);
assert!(struck.is_strikeout);
assert!(!struck.is_underline);
}
#[test]
fn test_skip_excessive_operations() {
use crate::tounicode::FontCMaps;
@@ -1279,13 +1727,113 @@ mod tests {
doc.add_object(catalog);
let font_cmaps = FontCMaps::from_doc(&doc);
let result = extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let result = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
let ((items, rects, lines), _has_gid, _coords_rotated) = result;
assert!(items.is_empty());
assert!(rects.is_empty());
assert!(lines.is_empty());
}
#[test]
fn test_q_restores_current_font_for_text_decoding() {
use crate::tounicode::FontCMaps;
use lopdf::{dictionary, Object, Stream};
fn cmap_stream(dst_hex: &str) -> Stream {
let cmap = format!(
r#"/CIDInit /ProcSet findresource begin
12 dict begin
begincmap
/CIDSystemInfo << /Registry (Adobe) /Ordering (UCS) /Supplement 0 >> def
/CMapName /Test-UCS def
/CMapType 2 def
1 begincodespacerange
<00> <FF>
endcodespacerange
1 beginbfchar
<41> <{dst_hex}>
endbfchar
endcmap
CMapName currentdict /CMap defineresource pop
end
end"#
);
Stream::new(dictionary! {}, cmap.into_bytes())
}
let mut doc = lopdf::Document::new();
let f1_cmap = doc.add_object(Object::Stream(cmap_stream("0058"))); // X
let f2_cmap = doc.add_object(Object::Stream(cmap_stream("0059"))); // Y
let f1 = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
"ToUnicode" => Object::Reference(f1_cmap),
});
let f2 = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
"ToUnicode" => Object::Reference(f2_cmap),
});
let content = b"BT /F1 12 Tf 10 700 Tm <41> Tj ET
q
BT /F2 12 Tf 20 700 Tm <41> Tj ET
Q
BT 30 700 Tm <41> Tj ET";
let content_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
content.to_vec(),
)));
let page_id = doc.add_object(dictionary! {
"Type" => "Page",
"Contents" => Object::Reference(content_id),
"Resources" => dictionary! {
"Font" => dictionary! {
"F1" => Object::Reference(f1),
"F2" => Object::Reference(f2),
},
},
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
});
let pages_id = doc.add_object(dictionary! {
"Type" => "Pages",
"Count" => Object::Integer(1),
"Kids" => vec![Object::Reference(page_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => Object::Reference(pages_id),
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
let text = items
.iter()
.map(|item| item.text.as_str())
.collect::<String>();
assert_eq!(text, "XYX");
}
#[test]
fn test_strip_pdf_comments() {
// Basic comment stripping
@@ -1317,4 +1865,26 @@ mod tests {
"ET should be preserved after comment stripping"
);
}
#[test]
fn test_strip_pdf_comments_escaped_parens() {
// An escaped `\)` must not close the string: the `%` after it is
// still string content, not a comment (subset fonts routinely map
// glyphs to `%` and to escaped parens in the same TJ array).
let input = b"[ (a\\)b) 1 (%) 1 (c) ] TJ\n";
let output = strip_pdf_comments(input);
assert_eq!(output, input.to_vec());
// Same for an escaped `\(` — must not open a phantom string that
// shields a real comment.
let input = b"(x\\(y) Tj % real comment\nET\n";
let output = strip_pdf_comments(input);
assert_eq!(output, b"(x\\(y) Tj \nET\n");
// Escaped backslash before a real close-paren: `\\` ends the escape,
// the `)` does close the string, and the comment is stripped.
let input = b"(x\\\\) Tj % comment\nET\n";
let output = strip_pdf_comments(input);
assert_eq!(output, b"(x\\\\) Tj \nET\n");
}
}
+1039 -60
View File
File diff suppressed because it is too large Load Diff
+1338 -41
View File
File diff suppressed because it is too large Load Diff
+405 -7
View File
@@ -2,11 +2,51 @@
use crate::types::{ItemType, TextItem};
use lopdf::{Document, Object, ObjectId};
use std::collections::HashMap;
use std::collections::{HashMap, HashSet};
use super::fonts::{resolve_array, resolve_dict};
use super::get_number;
/// Upper bound on the number of form-field nodes visited during a single
/// `extract_form_fields` pass. A crafted PDF can chain thousands of distinct
/// `/Kids` fields to blow the stack even without an outright reference cycle,
/// so we cap total traversal work in addition to detecting cycles.
const MAX_FORM_FIELD_NODES: usize = 100_000;
/// Upper bound on `/Kids` recursion depth. Real AcroForm hierarchies are only
/// a few levels deep (fields → child fields → widgets); a crafted PDF can chain
/// tens of thousands of distinct fields into a linear `/Kids` list that would
/// overflow the stack via depth-first recursion long before the node budget is
/// reached. This depth cap bounds the stack independently of total node count.
const MAX_FORM_FIELD_DEPTH: usize = 100;
/// Traversal budget for the AcroForm field walk. Bounds both the number of
/// distinct nodes visited *and* the total number of `/Fields`/`/Kids` entries
/// examined.
///
/// Counting `visited` alone is not enough: invalid entries (non-references) and
/// duplicate references never grow `visited`, so an oversized array full of them
/// would iterate to completion no matter how large. Charging every examined
/// entry against the same budget makes it a real cap on traversal work.
pub(crate) struct FieldWalkBudget {
visited: HashSet<ObjectId>,
examined: usize,
}
impl FieldWalkBudget {
fn new() -> Self {
Self {
visited: HashSet::new(),
examined: 0,
}
}
/// True once the budget is spent; callers must stop iterating and recursing.
fn exhausted(&self) -> bool {
self.visited.len() >= MAX_FORM_FIELD_NODES || self.examined >= MAX_FORM_FIELD_NODES
}
}
pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> Vec<TextItem> {
let mut links = Vec::new();
@@ -78,6 +118,8 @@ pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> V
page: page_num,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Link(url),
mcid: None,
});
@@ -144,32 +186,104 @@ pub(crate) fn extract_form_fields(
Err(_) => return items,
};
// Borrow the array rather than cloning it: a crafted `/Fields` can be huge,
// and cloning would pay an O(n) allocation/copy before the budget check
// below can stop the work.
let fields = match acroform.get(b"Fields") {
Ok(obj) => match resolve_array(doc, obj) {
Some(arr) => arr.clone(),
Some(arr) => arr,
None => return items,
},
Err(_) => return items,
};
if fields.is_empty() {
return items;
}
let annotation_pages = annotation_page_map(doc, page_map);
for field_obj in &fields {
// Bound the walk so a crafted PDF cannot send us into unbounded recursion
// via a `/Kids` cycle, a deep chain, or an oversized array of invalid or
// duplicate entries.
let mut budget = FieldWalkBudget::new();
for field_obj in fields {
// Stop once the budget is spent so a `/Fields` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op. Charge
// every entry (including invalid ones) against the budget.
if budget.exhausted() {
break;
}
budget.examined += 1;
if let Ok(field_ref) = field_obj.as_reference() {
walk_form_fields(doc, field_ref, None, "", page_map, &mut items);
walk_form_fields(
doc,
field_ref,
None,
"",
page_map,
&annotation_pages,
&mut items,
&mut budget,
0,
);
}
}
items
}
/// Map widget annotation objects back to the page whose `/Annots` array owns
/// them. Some valid widgets omit `/P`, so the page tree is the only reliable
/// ownership signal available for page-filtered extraction.
fn annotation_page_map(
doc: &Document,
page_map: &HashMap<ObjectId, u32>,
) -> HashMap<ObjectId, u32> {
let mut annotation_pages = HashMap::new();
for (&page_id, &page_num) in page_map {
let Some(annotations) = doc
.get_dictionary(page_id)
.ok()
.and_then(|page| page.get(b"Annots").ok())
.and_then(|annotations| resolve_array(doc, annotations))
else {
continue;
};
for annotation in annotations {
if let Ok(annotation_id) = annotation.as_reference() {
annotation_pages.insert(annotation_id, page_num);
}
}
}
annotation_pages
}
/// Recursively walk the form field tree, extracting leaf field values.
#[allow(clippy::too_many_arguments)]
pub(crate) fn walk_form_fields(
doc: &Document,
field_id: ObjectId,
parent_ft: Option<&[u8]>,
parent_name: &str,
page_map: &HashMap<ObjectId, u32>,
annotation_pages: &HashMap<ObjectId, u32>,
items: &mut Vec<TextItem>,
budget: &mut FieldWalkBudget,
depth: usize,
) {
// Guard against `/Kids` cycles and pathologically large field trees.
// Exceeding the depth cap means the chain is too deep to be a legitimate
// form (and would overflow the stack); an exhausted budget means the tree is
// too large. Both checks run *before* inserting so the visited set can never
// grow past the budget.
if depth > MAX_FORM_FIELD_DEPTH || budget.exhausted() {
return;
}
// Revisiting an object ID means we hit a `/Kids` cycle.
if !budget.visited.insert(field_id) {
return;
}
let field_dict = match doc.get_dictionary(field_id) {
Ok(d) => d,
Err(_) => return,
@@ -200,11 +314,31 @@ pub(crate) fn walk_form_fields(
// Check for /Kids — if present, recurse into children
if let Ok(kids_obj) = field_dict.get(b"Kids") {
// Iterate the borrowed array directly — cloning a crafted, oversized
// `/Kids` would allocate and copy every entry before the budget check
// below could stop the work.
if let Some(kids) = resolve_array(doc, kids_obj) {
let kids = kids.clone();
for kid in &kids {
for kid in kids {
// Stop once the budget is spent so a `/Kids` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op.
// Charge every entry (including invalid/duplicate ones) against
// the budget so this is a true traversal-work cap.
if budget.exhausted() {
break;
}
budget.examined += 1;
if let Ok(kid_ref) = kid.as_reference() {
walk_form_fields(doc, kid_ref, ft, &full_name, page_map, items);
walk_form_fields(
doc,
kid_ref,
ft,
&full_name,
page_map,
annotation_pages,
items,
budget,
depth + 1,
);
}
}
return;
@@ -297,6 +431,7 @@ pub(crate) fn walk_form_fields(
.ok()
.and_then(|o| o.as_reference().ok())
.and_then(|p| page_map.get(&p).copied())
.or_else(|| annotation_pages.get(&field_id).copied())
.unwrap_or(1);
let text = if full_name.is_empty() {
@@ -316,7 +451,270 @@ pub(crate) fn walk_form_fields(
page: page_num,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::FormField,
mcid: None,
});
}
#[cfg(test)]
mod tests {
use super::*;
use lopdf::{dictionary, Object};
#[test]
fn widget_without_page_reference_uses_owning_page_annotation() {
let mut doc = Document::new();
let widget_id = doc.add_object(dictionary! {
"Type" => "Annot",
"Subtype" => "Widget",
"FT" => "Tx",
"T" => Object::string_literal("customer"),
"V" => Object::string_literal("Alice"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
});
let page_one_id = doc.add_object(dictionary! {
"Type" => "Page",
});
let page_two_id = doc.add_object(dictionary! {
"Type" => "Page",
"Annots" => vec![Object::Reference(widget_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(widget_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::from([(page_one_id, 1), (page_two_id, 2)]);
let items = extract_form_fields(&doc, &page_map);
assert_eq!(items.len(), 1);
assert_eq!(items[0].page, 2);
assert_eq!(items[0].text, "customer: Alice");
}
#[test]
fn kids_self_cycle_does_not_overflow_stack() {
// A crafted AcroForm field that lists itself in `/Kids` must not send
// the traversal into unbounded recursion.
let mut doc = Document::new();
let field_id = doc.new_object_id();
doc.set_object(
field_id,
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("loop"),
"Kids" => vec![Object::Reference(field_id)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
// Completes (rather than overflowing the stack) and yields no items.
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn kids_mutual_cycle_terminates() {
// Two fields that reference each other via `/Kids` form a cycle that
// must also terminate.
let mut doc = Document::new();
let field_a = doc.new_object_id();
let field_b = doc.new_object_id();
doc.set_object(
field_a,
dictionary! {
"T" => Object::string_literal("a"),
"Kids" => vec![Object::Reference(field_b)],
},
);
doc.set_object(
field_b,
dictionary! {
"T" => Object::string_literal("b"),
"Kids" => vec![Object::Reference(field_a)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_a)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn deep_acyclic_kids_chain_does_not_overflow_stack() {
// A long chain of *distinct* fields (no cycle) must also terminate:
// the visited set alone would still recurse to the chain length, so
// the depth cap is what prevents a stack overflow here.
let mut doc = Document::new();
let n = MAX_FORM_FIELD_DEPTH * 500;
let ids: Vec<ObjectId> = (0..=n).map(|_| doc.new_object_id()).collect();
for i in 0..n {
doc.set_object(
ids[i],
dictionary! {
"FT" => "Tx",
"Kids" => vec![Object::Reference(ids[i + 1])],
},
);
}
// Leaf carries a value; it sits far below the depth cap so it is never
// reached, proving traversal stops early rather than crashing.
doc.set_object(
ids[n],
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("leaf"),
"V" => Object::string_literal("x"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(ids[0])],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn wide_tree_traversal_stops_at_node_budget() {
// A single field with a `/Kids` array wider than the node budget must
// stop traversal at the cap rather than growing `visited` (and the work)
// without bound. Each processed leaf emits one item, so the item count
// is bounded by the budget and reaches right up to it (a couple of
// slots go to the root and the boundary node charged against the cap).
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let kids: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
// Extraction stops at the budget: bounded above by the cap, and it gets
// right up to it (allowing a small delta for the root/boundary nodes
// charged against the budget).
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn wide_top_level_fields_stop_at_node_budget() {
// A top-level `/Fields` array wider than the budget must also stop at
// the cap: the item count is bounded by the budget and reaches right up
// to it.
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let fields: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => fields,
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn duplicate_and_invalid_kids_entries_stop_at_budget() {
// Duplicate references and non-reference junk never grow `visited`, so
// without charging examined entries against the budget an oversized
// array of them would iterate to completion. The walk must still
// terminate and extract the single real leaf exactly once.
let mut doc = Document::new();
let leaf_id = doc.new_object_id();
doc.set_object(
leaf_id,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
// A `/Kids` array far wider than the budget: half duplicate references
// to the same leaf, half invalid (null) entries.
let mut kids: Vec<Object> = Vec::new();
for i in 0..(MAX_FORM_FIELD_NODES * 2) {
if i % 2 == 0 {
kids.push(Object::Reference(leaf_id));
} else {
kids.push(Object::Null);
}
}
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert_eq!(items.len(), 1);
}
}
+1779 -35
View File
File diff suppressed because it is too large Load Diff
+591
View File
@@ -0,0 +1,591 @@
//! Region-graph evidence for page reading order.
//!
//! Whole-page column histograms fail when images or spanning captions occupy
//! only part of a page. This module turns image geometry and repeated row
//! gutters into a small directed acyclic graph: content above a local column
//! band, the left flow, the right flow, and content below it. The graph is
//! deliberately evidence-gated; ordinary pages keep the established layout
//! path.
use crate::text_utils::{effective_width, is_cjk_char, is_rtl_text};
use crate::types::TextItem;
const MIN_IMAGE_WIDTH: f32 = 60.0;
const MIN_IMAGE_HEIGHT: f32 = 40.0;
const MIN_ROW_GUTTER: f32 = 8.0;
const SPLIT_CLUSTER_TOLERANCE: f32 = 20.0;
const MIN_ALIGNED_ROWS: usize = 4;
pub(crate) type ImageRegion = (f32, f32, f32, f32);
#[derive(Debug, Clone, Copy, PartialEq)]
pub(crate) struct ColumnFlowBand {
pub(crate) split_x: f32,
pub(crate) y_bottom: f32,
pub(crate) y_top: f32,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub(crate) enum RegionKind {
FullWidth,
Column,
}
#[derive(Debug)]
pub(crate) struct RegionNode {
pub(crate) kind: RegionKind,
pub(crate) items: Vec<TextItem>,
}
#[derive(Debug)]
struct Row<'a> {
y: f32,
items: Vec<&'a TextItem>,
}
fn page_x_bounds(items: &[TextItem], images: &[ImageRegion]) -> Option<(f32, f32)> {
let text_min = items
.iter()
.map(|item| item.x)
.fold(f32::INFINITY, f32::min);
let text_max = items
.iter()
.map(|item| item.x + effective_width(item))
.fold(f32::NEG_INFINITY, f32::max);
let image_min = images
.iter()
.map(|region| region.0.min(region.2))
.fold(f32::INFINITY, f32::min);
let image_max = images
.iter()
.map(|region| region.0.max(region.2))
.fold(f32::NEG_INFINITY, f32::max);
let x_min = text_min.min(image_min);
let x_max = text_max.max(image_max);
(x_min.is_finite() && x_max.is_finite() && x_max > x_min).then_some((x_min, x_max))
}
fn group_rows(items: &[TextItem]) -> Vec<Row<'_>> {
const Y_TOLERANCE: f32 = 3.0;
let mut sorted: Vec<&TextItem> = items.iter().collect();
sorted.sort_by(|left, right| right.y.total_cmp(&left.y));
let mut rows: Vec<Row<'_>> = Vec::new();
for item in sorted {
if let Some(row) = rows
.last_mut()
.filter(|row| (row.y - item.y).abs() <= Y_TOLERANCE)
{
row.items.push(item);
row.y = row.items.iter().map(|member| member.y).sum::<f32>() / row.items.len() as f32;
} else {
rows.push(Row {
y: item.y,
items: vec![item],
});
}
}
for row in &mut rows {
row.items.sort_by(|left, right| left.x.total_cmp(&right.x));
}
rows
}
fn side_is_prose(items: &[&TextItem]) -> bool {
let text = items
.iter()
.map(|item| item.text.trim())
.collect::<Vec<_>>()
.join(" ");
let alphabetic_count = text
.chars()
.filter(|character| character.is_alphabetic())
.count();
let cjk_count = text
.chars()
.filter(|character| is_cjk_char(*character))
.count();
(text.split_whitespace().count() >= 3 || cjk_count >= 10) && alphabetic_count >= 10
}
fn aligned_row_split(row: &Row<'_>, x_min: f32, x_max: f32) -> Option<f32> {
if row.items.len() < 2 {
return None;
}
let page_width = x_max - x_min;
let center_low = x_min + page_width * 0.25;
let center_high = x_min + page_width * 0.75;
row.items
.windows(2)
.filter_map(|pair| {
let left_end = pair[0].x + effective_width(pair[0]);
let right_start = pair[1].x;
let gap = right_start - left_end;
let split_x = (left_end + right_start) / 2.0;
if gap < MIN_ROW_GUTTER || split_x < center_low || split_x > center_high {
return None;
}
let left: Vec<&TextItem> = row
.items
.iter()
.copied()
.filter(|item| item.x + effective_width(item) / 2.0 < split_x)
.collect();
let right: Vec<&TextItem> = row
.items
.iter()
.copied()
.filter(|item| item.x + effective_width(item) / 2.0 >= split_x)
.collect();
(side_is_prose(&left) && side_is_prose(&right)).then_some((split_x, gap))
})
.max_by(|left, right| left.1.total_cmp(&right.1))
.map(|candidate| candidate.0)
}
fn local_flow_below_full_width_image(
items: &[TextItem],
images: &[ImageRegion],
x_min: f32,
x_max: f32,
) -> Option<ColumnFlowBand> {
let page_width = x_max - x_min;
let full_width_images: Vec<ImageRegion> = images
.iter()
.copied()
.filter(|&(x0, y0, x1, y1)| {
let width = (x1 - x0).abs();
let height = (y1 - y0).abs();
width >= page_width * 0.65 && height >= 60.0
})
.collect();
// A local column flow below an image is only unambiguous for a single,
// nearly square hero/figure. Wide report banners and full-page artwork
// frequently sit above unrelated page furniture whose aligned labels can
// mimic prose columns.
if full_width_images.len() != 1 {
return None;
}
let (image_x0, _, image_x1, _) = full_width_images[0];
let anchor_width = (image_x1 - image_x0).abs();
let anchor_height = (full_width_images[0].3 - full_width_images[0].1).abs();
if anchor_width < page_width * 0.85
|| anchor_height < anchor_width * 0.85
|| anchor_height > anchor_width * 1.2
{
return None;
}
let image_bottom = full_width_images
.iter()
.map(|&(_, y0, _, y1)| y0.min(y1))
.fold(f32::NEG_INFINITY, f32::max);
if !image_bottom.is_finite() {
return None;
}
let below: Vec<TextItem> = items
.iter()
.filter(|item| item.y < image_bottom && item.y >= image_bottom - 220.0)
.cloned()
.collect();
let candidates: Vec<(f32, f32)> = group_rows(&below)
.into_iter()
.filter_map(|row| aligned_row_split(&row, x_min, x_max).map(|split| (split, row.y)))
.collect();
if candidates.len() < MIN_ALIGNED_ROWS {
return None;
}
let mut clusters: Vec<Vec<(f32, f32)>> = Vec::new();
for candidate in candidates {
if let Some(cluster) = clusters.iter_mut().find(|cluster| {
let mean = cluster.iter().map(|entry| entry.0).sum::<f32>() / cluster.len() as f32;
(mean - candidate.0).abs() <= SPLIT_CLUSTER_TOLERANCE
}) {
cluster.push(candidate);
} else {
clusters.push(vec![candidate]);
}
}
let dominant = clusters.into_iter().max_by_key(Vec::len)?;
if dominant.len() < MIN_ALIGNED_ROWS {
return None;
}
let split_x = dominant.iter().map(|entry| entry.0).sum::<f32>() / dominant.len() as f32;
let y_top = dominant
.iter()
.map(|entry| entry.1)
.fold(f32::NEG_INFINITY, f32::max)
+ 3.0;
let image_gap = image_bottom - y_top;
if !(60.0..=120.0).contains(&image_gap) {
return None;
}
let y_bottom = dominant
.iter()
.map(|entry| entry.1)
.fold(f32::INFINITY, f32::min)
- 3.0;
if y_top - y_bottom > 130.0 {
return None;
}
log::debug!(
"page {}: full-width image flow images={} aligned_rows={} split={:.1} page=[{:.1}..{:.1}] image_bottom={:.1} y=[{:.1}..{:.1}] full_width={:?}",
items.first().map_or(0, |item| item.page),
images.len(),
dominant.len(),
split_x,
x_min,
x_max,
image_bottom,
y_bottom,
y_top,
full_width_images
);
Some(ColumnFlowBand {
split_x,
y_bottom,
y_top,
})
}
fn paired_column_images(
items: &[TextItem],
images: &[ImageRegion],
split_x: f32,
x_min: f32,
x_max: f32,
) -> Option<ColumnFlowBand> {
let page_width = x_max - x_min;
if split_x < x_min + page_width * 0.4 || split_x > x_min + page_width * 0.6 {
return None;
}
let qualifying: Vec<ImageRegion> = images
.iter()
.copied()
.filter(|&(x0, y0, x1, y1)| {
let image_left = x0.min(x1);
let image_right = x0.max(x1);
let confined_to_one_column = image_right <= split_x || image_left >= split_x;
confined_to_one_column
&& (x1 - x0).abs() >= MIN_IMAGE_WIDTH
&& (y1 - y0).abs() >= MIN_IMAGE_HEIGHT
})
.collect();
let wide_images: Vec<ImageRegion> = qualifying
.iter()
.copied()
.filter(|(x0, _, x1, _)| (x1 - x0).abs() >= page_width * 0.35)
.collect();
if qualifying.len() < 3 || wide_images.len() < 3 {
return None;
}
let has_left = qualifying
.iter()
.any(|&(x0, _, x1, _)| (x0 + x1) / 2.0 < split_x);
let has_right = qualifying
.iter()
.any(|&(x0, _, x1, _)| (x0 + x1) / 2.0 >= split_x);
if !has_left || !has_right {
return None;
}
// A meaningful image-backed column flow spans multiple vertical panels.
// Three same-row header/logo images can otherwise satisfy the image count
// and send an ordinary asymmetric page through sequential column order.
let image_y_min = wide_images
.iter()
.map(|region| region.1.min(region.3))
.fold(f32::INFINITY, f32::min);
let image_y_max = wide_images
.iter()
.map(|region| region.1.max(region.3))
.fold(f32::NEG_INFINITY, f32::max);
let has_vertical_stack = wide_images.iter().enumerate().any(|(index, left)| {
wide_images.iter().skip(index + 1).any(|right| {
let same_side =
((left.0 + left.2) / 2.0 < split_x) == ((right.0 + right.2) / 2.0 < split_x);
let left_center = (left.1 + left.3) / 2.0;
let right_center = (right.1 + right.3) / 2.0;
let left_height = (left.3 - left.1).abs();
let right_height = (right.3 - right.1).abs();
let vertical_gap = if left.1.max(left.3) < right.1.min(right.3) {
right.1.min(right.3) - left.1.max(left.3)
} else if right.1.max(right.3) < left.1.min(left.3) {
left.1.min(left.3) - right.1.max(right.3)
} else {
0.0
};
same_side
&& (left_center - right_center).abs() >= left_height.min(right_height) * 0.5
&& vertical_gap <= left_height.max(right_height) * 0.5
})
});
if image_y_max - image_y_min < page_width * 0.45 || !has_vertical_stack {
return None;
}
let y_top = qualifying
.iter()
.map(|region| region.1.max(region.3))
.fold(f32::NEG_INFINITY, f32::max)
+ 3.0;
// Only column-confined text proves the lower extent of the flow. A
// spanning heading or caption below the columns must become the trailing
// full-width node rather than stretching the column band to the page foot.
let y_bottom = items
.iter()
.filter(|item| {
let item_right = item.x + effective_width(item);
item.y <= y_top && (item_right <= split_x || item.x >= split_x)
})
.map(|item| item.y)
.fold(f32::INFINITY, f32::min)
- 3.0;
if !y_bottom.is_finite() {
return None;
}
let distinct_rows = |right: bool| {
let mut ys: Vec<f32> = items
.iter()
.filter(|item| {
item.y <= y_top && (item.x + effective_width(item) / 2.0 >= split_x) == right
})
.map(|item| item.y)
.collect();
ys.sort_by(|left, right| left.total_cmp(right));
ys.dedup_by(|left, right| (*left - *right).abs() <= 3.0);
ys.len()
};
let left_rows = distinct_rows(false);
let right_rows = distinct_rows(true);
let line_balance = left_rows.min(right_rows) as f32 / left_rows.max(right_rows).max(1) as f32;
(left_rows >= 5 && right_rows >= 5 && line_balance < 0.55).then(|| {
log::debug!(
"page {}: paired-image flow qualifying_images={} rows={}/{} split={:.1} page=[{:.1}..{:.1}] y=[{:.1}..{:.1}] images={:?}",
items.first().map_or(0, |item| item.page),
qualifying.len(),
left_rows,
right_rows,
split_x,
x_min,
x_max,
y_bottom,
y_top,
qualifying
);
ColumnFlowBand {
split_x,
y_bottom,
y_top,
}
})
}
pub(crate) fn infer_image_anchored_flow(
items: &[TextItem],
images: &[ImageRegion],
detected_split: Option<f32>,
) -> Option<ColumnFlowBand> {
if items.is_empty() || images.is_empty() {
return None;
}
let (x_min, x_max) = page_x_bounds(items, images)?;
detected_split
.and_then(|split_x| paired_column_images(items, images, split_x, x_min, x_max))
.or_else(|| local_flow_below_full_width_image(items, images, x_min, x_max))
}
/// Partition a page into the topological order `above -> left -> right -> below`.
/// These edges encode the reading-order DAG; empty nodes are omitted.
pub(crate) fn build_region_graph(items: Vec<TextItem>, band: ColumnFlowBand) -> Vec<RegionNode> {
let mut above = Vec::new();
let mut left = Vec::new();
let mut right = Vec::new();
let mut below = Vec::new();
for item in items {
if item.y > band.y_top {
above.push(item);
} else if item.y < band.y_bottom {
below.push(item);
} else if item.x + effective_width(&item) / 2.0 < band.split_x {
left.push(item);
} else {
right.push(item);
}
}
let rtl = is_rtl_text(left.iter().chain(right.iter()).map(|item| &item.text));
let mut ordered = vec![(RegionKind::FullWidth, above)];
if rtl {
ordered.push((RegionKind::Column, right));
ordered.push((RegionKind::Column, left));
} else {
ordered.push((RegionKind::Column, left));
ordered.push((RegionKind::Column, right));
}
ordered.push((RegionKind::FullWidth, below));
ordered
.into_iter()
.filter_map(|(kind, items)| (!items.is_empty()).then_some(RegionNode { kind, items }))
.collect()
}
#[cfg(test)]
mod tests {
use super::*;
use crate::types::ItemType;
fn item(text: &str, x: f32, y: f32, width: f32) -> TextItem {
TextItem {
text: text.into(),
x,
y,
width,
height: 11.0,
font: "F1".into(),
font_size: 11.0,
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
}
#[test]
fn full_width_image_anchors_local_two_column_flow() {
let mut items = vec![
item("A full width caption", 55.0, 230.0, 430.0),
item("A trailing full width heading", 55.0, 80.0, 430.0),
];
for index in 0..5 {
let y = 170.0 - index as f32 * 14.0;
items.push(item("left column prose words", 55.0, y, 210.0));
items.push(item("right column prose words", 280.0, y, 210.0));
}
let images = vec![(55.0, 250.0, 490.0, 680.0)];
let band = infer_image_anchored_flow(&items, &images, None).unwrap();
assert!((band.split_x - 272.5).abs() < 2.0);
let graph = build_region_graph(items, band);
assert_eq!(graph.len(), 4);
assert_eq!(graph[0].kind, RegionKind::FullWidth);
assert_eq!(graph[1].kind, RegionKind::Column);
assert_eq!(graph[2].kind, RegionKind::Column);
assert_eq!(graph[3].kind, RegionKind::FullWidth);
assert_eq!(graph[3].items[0].text, "A trailing full width heading");
}
#[test]
fn full_width_image_anchors_cjk_column_flow() {
let mut items = Vec::new();
for index in 0..5 {
let y = 170.0 - index as f32 * 14.0;
items.push(item("左栏这是没有空格的正文内容", 55.0, y, 210.0));
items.push(item("右栏这是没有空格的正文内容", 280.0, y, 210.0));
}
let images = vec![(55.0, 250.0, 490.0, 680.0)];
assert!(infer_image_anchored_flow(&items, &images, None).is_some());
}
#[test]
fn paired_images_anchor_unbalanced_column_flows() {
let mut items = vec![
item("running header", 55.0, 700.0, 430.0),
item("trailing full width caption", 55.0, 300.0, 430.0),
];
for index in 0..5 {
items.push(item(
"left prose words",
55.0,
500.0 - index as f32 * 14.0,
200.0,
));
items.push(item(
"right prose words",
280.0,
520.0 - index as f32 * 14.0,
200.0,
));
}
for index in 5..12 {
items.push(item(
"right continuation prose words",
280.0,
520.0 - index as f32 * 14.0,
200.0,
));
}
let images = vec![
(55.0, 530.0, 255.0, 680.0),
(55.0, 380.0, 255.0, 530.0),
(280.0, 560.0, 490.0, 680.0),
];
let band = infer_image_anchored_flow(&items, &images, Some(270.0)).unwrap();
let graph = build_region_graph(items, band);
assert_eq!(graph[0].kind, RegionKind::FullWidth);
assert_eq!(graph[1].kind, RegionKind::Column);
assert_eq!(graph[2].kind, RegionKind::Column);
assert_eq!(graph[3].kind, RegionKind::FullWidth);
assert_eq!(graph[3].items[0].text, "trailing full width caption");
}
#[test]
fn rtl_region_graph_reads_right_column_first() {
let items = vec![
item("A long English report header", 55.0, 250.0, 430.0),
item("نص العمود الأيسر", 55.0, 150.0, 180.0),
item("نص العمود الأيمن", 300.0, 150.0, 180.0),
];
let graph = build_region_graph(
items,
ColumnFlowBand {
split_x: 270.0,
y_bottom: 100.0,
y_top: 200.0,
},
);
assert_eq!(graph.len(), 3);
assert_eq!(graph[0].kind, RegionKind::FullWidth);
assert!(graph[1].items[0].x > graph[2].items[0].x);
}
#[test]
fn paired_header_logos_do_not_anchor_page_columns() {
let mut items = Vec::new();
for index in 0..7 {
items.push(item(
"left prose words",
55.0,
700.0 - index as f32 * 14.0,
200.0,
));
}
for index in 0..30 {
items.push(item(
"right prose words",
280.0,
700.0 - index as f32 * 14.0,
200.0,
));
}
let images = vec![
(55.0, 720.0, 205.0, 770.0),
(60.0, 718.0, 210.0, 768.0),
(280.0, 720.0, 450.0, 770.0),
];
assert!(infer_image_anchored_flow(&items, &images, Some(270.0)).is_none());
}
#[test]
fn wide_banner_does_not_anchor_local_columns() {
let mut items = Vec::new();
for index in 0..7 {
let y = 270.0 - index as f32 * 14.0;
items.push(item("left column prose words", 55.0, y, 210.0));
items.push(item("right column prose words", 280.0, y, 210.0));
}
let images = vec![(55.0, 310.0, 490.0, 550.0)];
assert!(infer_image_anchored_flow(&items, &images, None).is_none());
}
}
File diff suppressed because it is too large Load Diff
+38 -9
View File
@@ -1,5 +1,6 @@
//! Form XObject and image XObject extraction.
use super::fonts::descriptor_style_flags;
use crate::text_utils::{effective_font_size, expand_ligatures, is_bold_font, is_italic_font};
use crate::tounicode::FontCMaps;
use crate::types::{ItemType, TextItem};
@@ -7,8 +8,9 @@ use lopdf::{Document, Encoding, Object, ObjectId};
use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
build_font_encodings, build_font_widths, build_type3_scales, compute_string_width_ts,
extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
FontStyleCache,
};
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
@@ -114,6 +116,7 @@ pub(crate) fn extract_form_xobject_text(
font_cmaps: &FontCMaps,
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
) -> Vec<TextItem> {
extract_form_xobject_text_inner(
doc,
@@ -122,10 +125,12 @@ pub(crate) fn extract_form_xobject_text(
font_cmaps,
parent_ctm,
cmap_decisions,
style_cache,
0,
)
}
#[allow(clippy::too_many_arguments)]
fn extract_form_xobject_text_inner(
doc: &Document,
form_id: ObjectId,
@@ -133,6 +138,7 @@ fn extract_form_xobject_text_inner(
font_cmaps: &FontCMaps,
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
depth: u8,
) -> Vec<TextItem> {
use lopdf::content::Content;
@@ -157,16 +163,18 @@ fn extract_form_xobject_text_inner(
// Get fonts from the Form's Resources
let form_fonts = get_form_fonts(doc, &stream.dict);
let (font_encodings, _has_gid_fonts) = build_font_encodings(doc, &form_fonts);
let (font_encodings, _has_gid_fonts) = build_font_encodings(doc, &form_fonts, font_cmaps);
// Build font width info for the form
let font_widths = build_font_widths(doc, &form_fonts);
let type3_scales = build_type3_scales(doc, &form_fonts);
// Build font base names and ToUnicode refs for the form
let mut font_base_names: HashMap<String, String> = HashMap::new();
let mut font_tounicode_refs: HashMap<String, u32> = HashMap::new();
let mut inline_cmaps: HashMap<String, crate::tounicode::CMapEntry> = HashMap::new();
let mut font_style_flags: HashMap<String, (bool, bool)> = HashMap::new();
for (font_name, font_dict) in &form_fonts {
let resource_name = String::from_utf8_lossy(font_name).to_string();
if let Ok(base_font) = font_dict.get(b"BaseFont") {
@@ -175,6 +183,10 @@ fn extract_form_xobject_text_inner(
font_base_names.insert(resource_name.clone(), base_name);
}
}
let style = descriptor_style_flags(doc, font_dict, style_cache);
if style != (false, false) {
font_style_flags.insert(resource_name.clone(), style);
}
match font_dict.get(b"ToUnicode") {
Ok(tounicode) => {
if let Ok(obj_ref) = tounicode.as_reference() {
@@ -272,6 +284,7 @@ fn extract_form_xobject_text_inner(
font_cmaps,
&ctm,
cmap_decisions,
style_cache,
depth + 1,
);
items.extend(nested_items);
@@ -294,6 +307,8 @@ fn extract_form_xobject_text_inner(
page: page_num,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Image,
mcid: None,
});
@@ -400,7 +415,8 @@ fn extract_form_xobject_text_inner(
&font_widths,
) {
let combined = multiply_matrices(&text_matrix, &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
let width = if let Some(font_info) = font_widths.get(&current_font) {
if let Some(raw_bytes) = get_operand_bytes(&op.operands[0]) {
@@ -427,6 +443,10 @@ fn extract_form_xobject_text_inner(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
@@ -436,8 +456,10 @@ fn extract_form_xobject_text_inner(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -553,11 +575,16 @@ fn extract_form_xobject_text_inner(
}
if !sub_items.is_empty() {
let combined = multiply_matrices(&text_matrix, &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
let rendered_size = effective_font_size(current_font_size, &combined)
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let base_font = font_base_names
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
let scale_x = text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2];
for (text, start_w, end_w) in &sub_items {
let offset_tm = [
@@ -584,8 +611,10 @@ fn extract_form_xobject_text_inner(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
+40 -3
View File
@@ -4566,9 +4566,13 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
}
}
// Try to parse uniXXXX format
if name.starts_with("uni") && name.len() >= 7 {
if let Ok(code) = u32::from_str_radix(&name[3..7], 16) {
// Try to parse uniXXXX format.
// Use `get` rather than a byte-length check + slice: `name` can contain
// non-ASCII bytes (e.g. U+FFFD from lossy UTF-8 decoding of an attacker
// controlled /Differences name), so byte index 7 may not be a char
// boundary and `&name[3..7]` would panic.
if let Some(hex) = name.strip_prefix("uni").and_then(|rest| rest.get(..4)) {
if let Ok(code) = u32::from_str_radix(hex, 16) {
// Strip PUA F000 offset: uniF0XX → U+00XX (Windows Symbol encoding convention)
let code = if (0xF000..=0xF0FF).contains(&code) {
code - 0xF000
@@ -4588,3 +4592,36 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
None
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn uni_hex_parsing() {
assert_eq!(glyph_to_char("uni0041"), Some('A'));
assert_eq!(glyph_to_char("uni00e9"), Some('\u{00e9}'));
// PUA F0xx symbol-encoding offset is stripped.
assert_eq!(glyph_to_char("uniF041"), Some('A'));
}
#[test]
fn u_hex_parsing() {
assert_eq!(glyph_to_char("u0041"), Some('A'));
assert_eq!(glyph_to_char("u1F600"), Some('\u{1F600}'));
}
#[test]
fn non_ascii_uni_name_does_not_panic() {
// A crafted /Differences name like `/uni#80#80#80#80` decodes via
// from_utf8_lossy into "uni" followed by four U+FFFD replacements.
// Byte index 7 lands mid-character, so a naive `&name[3..7]` slice
// would panic. It must be handled gracefully instead.
let crafted = format!("uni{0}{0}{0}{0}", '\u{FFFD}');
assert_eq!(glyph_to_char(&crafted), None);
// Assorted non-ASCII bytes right after the "uni" prefix.
assert_eq!(glyph_to_char("uni\u{FFFD}bc"), None);
assert_eq!(glyph_to_char("uni\u{00e9}00"), None);
}
}
+1379 -198
View File
File diff suppressed because it is too large Load Diff
+499
View File
@@ -130,6 +130,315 @@ pub(crate) fn has_dot_leaders(text: &str) -> bool {
dot_groups >= 2
}
/// Detect a table-of-contents entry: a line ending in a page number preceded by
/// a dot-leader group (e.g. "Measurement Lab worksheet ... 3"). `has_dot_leaders`
/// misses single-group leaders ("..."), but a trailing "<dots> <number>" is a
/// strong TOC signal on its own. Such lines must never be promoted to headings.
pub(crate) fn is_toc_entry_line(text: &str) -> bool {
let trimmed = text.trim_end();
let digits = trimmed
.chars()
.rev()
.take_while(|c| c.is_ascii_digit())
.count();
if digits == 0 || digits > 4 {
return false;
}
let before_number = trimmed[..trimmed.len() - digits].trim_end();
let dots = before_number
.chars()
.rev()
.take_while(|c| *c == '.')
.count();
dots >= 3
}
/// A heading that announces a table of contents ("Contents", "Table of
/// Contents"). Lines after it on the same page are ToC entries — section
/// titles that look exactly like headings but must not be promoted.
pub(crate) fn is_toc_marker_heading(text: &str) -> bool {
let t = text.trim().trim_end_matches(':').trim().to_lowercase();
matches!(t.as_str(), "contents" | "table of contents")
}
/// Lines that resemble headings structurally but are display-math fragments:
/// equations ending in an equation number ("S = kB ln W, (2)") or equation
/// lead-ins ("Rearranging Equation (8) gives:"). Both carry an "(N)" equation
/// reference — but a trailing "(N)" alone is not enough: real headings end
/// with parenthesized numbers too ("Nicaea (325)", appendix numbering), so
/// the suffix form additionally requires math evidence — an "=" in the line
/// or a comma immediately before the number, both present in every display
/// equation and absent from name-plus-number headings. A bare trailing colon
/// is NOT a fragment signal either: real headings frequently end with colons
/// ("Procedure:", "Steps for Using the Microscope:").
/// True when the line opens with a section number ("3.", "2.1.4", "IV)").
///
/// Mirrors the acceptance of `heading::parse_numbering` rather than the
/// stricter `convert::starts_with_section_number`, which deliberately
/// requires two components because it bypasses isolation checks. Here a
/// single "1." counts: numbering is independent evidence of a heading, and
/// `heading.rs` applies its numbered-prefix allowance *after* consulting
/// `is_heading_fragment`, so without this exemption a numbered
/// sentence-case heading would be vetoed before that allowance can run.
fn starts_with_numbering_prefix(t: &str) -> bool {
let Some(first) = t.split_whitespace().next() else {
return false;
};
let has_delimiter = first.ends_with(['.', ')', ':']);
let token = first.trim_end_matches(['.', ')', ':']);
if token.is_empty() {
return false;
}
let parts: Vec<&str> = token.split('.').collect();
let decimal = parts
.iter()
.all(|p| !p.is_empty() && p.len() <= 3 && p.chars().all(|c| c.is_ascii_digit()));
if decimal {
// "1." / "2.1." carry a delimiter; "2.3 Title" is written without
// one, so a multi-component number is accepted bare. A bare single
// number ("3 apples") is not — that is ordinary prose.
return has_delimiter || parts.len() >= 2;
}
// Roman numerals go through the heading parser's own grammar so the two
// agree: uppercase I/V/X/L/C only, at most 8 characters. A looser rule
// here would exempt markers the parser rejects — "iv)" or "d)" from an
// alphabetical list — letting an ordinary list item bypass the veto and
// reach heading promotion.
//
// A delimiter is also required: a bare leading "I" is the pronoun far
// more often than a section number.
has_delimiter && crate::markdown::heading::roman_value(token).is_some()
}
/// True when the line reads as a title rather than a sentence: every
/// content word (ignoring minor words) starts uppercase. Used to spare real
/// headings from the dangling-verb veto — "Bond Yields" is a section title,
/// "the method yields" is a stranded clause, and only the casing tells them
/// apart.
fn looks_title_case(t: &str) -> bool {
const MINOR: &[&str] = &[
"a", "an", "the", "of", "and", "or", "for", "to", "in", "on", "at", "by", "with", "from",
"as", "is", "are", "that", "than", "into",
];
let mut content = 0usize;
let mut capitalized = 0usize;
for w in t.split_whitespace() {
let cleaned: String = w.chars().filter(|c| c.is_alphabetic()).collect();
if cleaned.is_empty() {
continue;
}
if MINOR.contains(&cleaned.to_lowercase().as_str()) {
continue;
}
content += 1;
if cleaned.chars().next().is_some_and(char::is_uppercase) {
capitalized += 1;
}
}
// A single content word ("Yields") is a title by default.
content == 0 || capitalized == content
}
pub(crate) fn is_heading_fragment(text: &str) -> bool {
let t = text.trim_end();
// A lowercase-initial one-or-two-word "heading" is a mid-sentence
// fragment beside display math ("or inversely", "and therefore") —
// real headings that short start uppercase. Measured as spurious
// headings on academic docs (fire-pdf ENG-5029 / opendataloader MHS).
{
let words: Vec<&str> = t.split_whitespace().collect();
if words.len() <= 2 {
if let Some(first_alpha) = t.chars().find(|c| c.is_alphabetic()) {
if first_alpha.is_lowercase() {
return true;
}
}
}
}
fn is_equation_number(s: &str) -> bool {
s.strip_prefix('(')
.and_then(|r| r.strip_suffix(')'))
.is_some_and(|inner| {
!inner.is_empty() && inner.len() <= 3 && inner.chars().all(|c| c.is_ascii_digit())
})
}
// Equation-number suffix with math evidence: "S = kB ln W, (2)"
let mut rev = t.rsplit(' ');
let last = rev.next().unwrap_or("");
if is_equation_number(last) {
// Page-of-total running headers: "LIVSMEDELSVERKET PM 2 (10)"
if let Some(prev_word) = t.rsplit(' ').nth(1) {
if let (Ok(page), Some(total)) = (
prev_word.parse::<u32>(),
last.trim_start_matches('(')
.trim_end_matches(')')
.parse::<u32>()
.ok(),
) {
if page <= total {
return true;
}
}
}
let punct_before = rev
.next()
.is_some_and(|w| w.ends_with(',') || w.ends_with(':'));
let has_math_op = t.chars().any(|c| {
matches!(
c,
'=' | '<'
| '>'
| '≤'
| '≥'
| '≪'
| '≫'
| '≈'
| '≠'
| '±'
| '∑'
| '∫'
| '√'
| '∝'
)
});
if punct_before || has_math_op {
return true;
}
}
// Lead-in: ends with a colon AND references an equation number inline
if t.ends_with(':') && t.split_whitespace().any(is_equation_number) {
return true;
}
// Dangling clause: a stranded sentence lead-in ends on a relational
// verb with no terminal punctuation — "Note that the exact error equals"
// left ahead of its formula when a phantom table dissolved.
//
// Gated on the line reading as prose rather than a title. Case is the
// discriminator the trailing word alone cannot provide: a heading is
// title case ("Bond Yields", "The Method Yields") while a stranded
// lead-in is sentence case ("the method yields"). Without this gate the
// veto eats real headings — "Bond Yields", "Crop Yields" and any wrapped
// title-case heading the preprocessor failed to merge.
if !t.ends_with(['.', '!', '?', ':', ';', ')', ']'])
&& !looks_title_case(t)
&& !starts_with_numbering_prefix(t)
{
if let Some(last) = t.split_whitespace().next_back() {
let word: String = last
.trim_matches(|c: char| !c.is_alphanumeric())
.to_lowercase();
// Relational verbs only, and only those with no common noun
// sense. "yields" was dropped for exactly that reason: "Bond
// Yields" is a real section title. Function words, copulas and
// auxiliaries were measured and rejected outright — a heading
// that wraps across lines ends on those, and suppressing them
// destroyed real IRS Publication 17 headings.
const DANGLING_TAIL: &[&str] =
&["equals", "denotes", "implies", "satisfies", "signifies"];
if DANGLING_TAIL.contains(&word.as_str()) {
return true;
}
}
}
false
}
#[cfg(test)]
mod fragment_heading_tests {
use super::is_heading_fragment;
#[test]
fn dangling_tail_marks_stranded_clause() {
// opendataloader 01030000000144: left behind when a phantom table
// dissolved, ahead of its formula on the next line.
assert!(is_heading_fragment("Note that the exact error equals"));
assert!(is_heading_fragment("The remainder term satisfies"));
assert!(is_heading_fragment("we conclude that the sum equals"));
}
#[test]
fn real_headings_survive() {
assert!(!is_heading_fragment("Introduction"));
assert!(!is_heading_fragment("Error Analysis"));
assert!(!is_heading_fragment("Materials and Methods"));
assert!(!is_heading_fragment("Results"));
assert!(!is_heading_fragment("3.2 Richardson Extrapolation"));
assert!(!is_heading_fragment("Discussion and Conclusions"));
// Terminal punctuation means the clause is complete.
assert!(!is_heading_fragment("What is a Derivative?"));
assert!(!is_heading_fragment("Procedure:"));
assert!(!is_heading_fragment("Note that this is important."));
}
#[test]
fn title_case_headings_ending_in_a_verb_survive() {
// "yields" is also a plural noun; these are real section titles.
assert!(!is_heading_fragment("Bond Yields"));
assert!(!is_heading_fragment("Crop Yields"));
assert!(!is_heading_fragment("Dividend Yields"));
assert!(!is_heading_fragment("Yields"));
// A wrapped title-case heading whose first line ends on a listed
// verb must survive even if the preprocessor failed to merge it.
assert!(!is_heading_fragment("The Theorem Implies"));
assert!(!is_heading_fragment("What This Denotes"));
}
#[test]
fn numbered_sentence_case_headings_survive() {
// heading.rs consults is_heading_fragment BEFORE applying its
// numbered-prefix allowance, so the veto must not pre-empt it.
assert!(!is_heading_fragment("1. What the model implies"));
assert!(!is_heading_fragment("2.3 How the estimator satisfies"));
assert!(!is_heading_fragment("IV) What this denotes"));
// Without numbering the same wording is still a stranded clause.
assert!(is_heading_fragment("What the model implies"));
// A bare leading number or pronoun is prose, not numbering.
assert!(is_heading_fragment("3 apples and what that implies"));
assert!(is_heading_fragment("I think the model implies"));
// Markers heading::parse_numbering rejects must not be exempted
// either, or an ordinary list item bypasses the veto: lowercase
// roman, alphabetical markers, and over-long tokens.
assert!(is_heading_fragment("iv) the estimator satisfies"));
assert!(is_heading_fragment("d) the value implies"));
// Unsupported character (M is outside the parser's I/V/X/L/C set).
assert!(is_heading_fragment("MMMM. the value implies"));
// Over-long token: nine valid characters, so this exercises the
// 8-character bound rather than the character set.
assert!(is_heading_fragment("IIIIIIIII. the value implies"));
// Eight is still within the bound and stays exempt.
assert!(!is_heading_fragment("IIIIIIII. What this implies"));
// Uppercase roman within the parser's grammar is still exempt.
assert!(!is_heading_fragment("IV. What this denotes"));
assert!(!is_heading_fragment("XII) What this implies"));
}
#[test]
fn wrapped_headings_are_not_fragments() {
// A heading that wraps across lines ends on a function word. These
// are real headings from IRS Publication 17 and must survive.
assert!(!is_heading_fragment("Casualty and"));
assert!(!is_heading_fragment("Rule 10. You Must Be at"));
assert!(!is_heading_fragment("Higher Standard Deduction for"));
assert!(!is_heading_fragment("Qualifying Child of"));
assert!(!is_heading_fragment("When Can I Withdraw or"));
// Copulas and auxiliaries also end real wrapped headings.
assert!(!is_heading_fragment("Rule 15. Your AGI Must Be"));
assert!(!is_heading_fragment("What Medical Expenses Are"));
assert!(!is_heading_fragment("Rule 13. You Must Have"));
assert!(!is_heading_fragment("When Can a Roth IRA Be"));
}
#[test]
fn dangling_check_is_case_insensitive() {
// All-caps is not sentence case, so the veto must not fire there.
assert!(!is_heading_fragment("THE REMAINDER EQUALS"));
}
}
/// Compute the Y-gap threshold for paragraph break detection.
///
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
@@ -257,6 +566,15 @@ pub(crate) fn compute_heading_tiers(lines: &[TextLine], base_size: f32) -> Vec<f
for line in lines {
if let Some(first) = line.items.first() {
if first.font_size / base_size >= 1.2 {
// Digit-only lines (page numbers, issue numbers) must not
// define heading tiers: a large bold folio claims tier 0 and
// blocks the bold-size fallback for the document's real
// same-size headings.
let text = line.text();
let t = text.trim();
if !t.is_empty() && t.chars().all(|c| !c.is_alphabetic()) {
continue;
}
heading_sizes.push(first.font_size);
}
}
@@ -274,11 +592,45 @@ pub(crate) fn compute_heading_tiers(lines: &[TextLine], base_size: f32) -> Vec<f
}
}
// Books often set section headings barely above body size (e.g. 11pt
// bold over 10pt text). When nothing clears the 1.2x ratio gate, fall
// back to bold lines modestly larger than body so those documents still
// get an H1 instead of every bold heading defaulting to H2.
if tiers.is_empty() {
let mut bold_sizes: Vec<f32> = lines
.iter()
.filter(|line| {
let text = line.text();
let t = text.trim();
!t.is_empty() && t.chars().any(|c| c.is_alphabetic())
})
.filter_map(|line| line.items.first())
.filter(|it| it.is_bold && it.font_size / base_size >= 1.05)
.map(|it| it.font_size)
.collect();
bold_sizes.sort_by(|a, b| b.total_cmp(a));
for size in bold_sizes {
if !tiers.iter().any(|&t| (t - size).abs() < 0.5) {
tiers.push(size);
}
}
}
// Cap at 4 tiers
tiers.truncate(4);
tiers
}
/// Boldness of a line judged by character mass, so a heading with an
/// unbold section-number prefix ("4. " + bold title) still counts as bold.
pub(crate) fn line_is_mostly_bold(line: &TextLine) -> bool {
let (bold, total) = line.items.iter().fold((0usize, 0usize), |(b, t), it| {
let n = it.text.trim().chars().count();
(b + if it.is_bold { n } else { 0 }, t + n)
});
total > 0 && bold * 2 >= total
}
/// Detect header level from font size using document-specific heading tiers.
/// When tiers are available, maps tier 0→H1, tier 1→H2, etc.
/// Falls back to ratio-based thresholds when no tiers exist.
@@ -286,9 +638,21 @@ pub(crate) fn detect_header_level(
font_size: f32,
base_size: f32,
heading_tiers: &[f32],
is_bold: bool,
) -> Option<usize> {
let ratio = font_size / base_size;
// Tier matches are trusted below the 1.2x gate (down to 1.05x) only for
// bold lines: sub-gate tiers come from the bold fallback, and honoring
// them for non-bold text at the same size would promote captions.
if (1.05..1.2).contains(&ratio) && is_bold && !heading_tiers.is_empty() {
for (i, &tier_size) in heading_tiers.iter().enumerate() {
if (font_size - tier_size).abs() < 0.5 {
return Some(i + 1); // tier 0 → H1, tier 1 → H2, etc.
}
}
}
if ratio < 1.2 {
return None; // Regular text
}
@@ -320,3 +684,138 @@ pub(crate) fn detect_header_level(
Some(4)
}
}
#[cfg(test)]
mod tests {
use super::*;
fn line_of(text: &str, font_size: f32, bold: bool, y: f32) -> crate::types::TextLine {
let item = crate::types::TextItem {
text: text.into(),
x: 72.0,
y,
width: text.len() as f32 * font_size * 0.5,
height: font_size,
font: "Test".into(),
font_size,
page: 1,
is_bold: bold,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
};
crate::types::TextLine {
items: vec![item],
y,
page: 1,
adaptive_threshold: 0.10,
}
}
#[test]
fn digit_only_lines_do_not_define_tiers() {
// A 14pt bold page number must not claim tier 0 — that both demotes
// every real heading a level and blocks the bold-size fallback.
let lines = vec![
line_of("76", 14.0, true, 760.0),
line_of("Replace", 11.0, true, 700.0),
line_of("body text at eleven points", 11.0, false, 680.0),
];
let tiers = compute_heading_tiers(&lines, 11.0);
assert!(tiers.is_empty(), "page number claimed a tier: {tiers:?}");
}
#[test]
fn bold_fallback_tiers_when_nothing_clears_ratio_gate() {
// 10pt body, 11pt bold section headings (book-style): no size clears
// 1.2x, so bold sizes modestly above body form the tiers.
let lines = vec![
line_of("4. Entropy", 11.0, true, 700.0),
line_of("body text about entropy", 10.0, false, 680.0),
line_of("5. The dynamics", 11.0, true, 500.0),
];
let tiers = compute_heading_tiers(&lines, 10.0);
assert_eq!(tiers, vec![11.0]);
assert_eq!(detect_header_level(11.0, 10.0, &tiers, true), Some(1));
// Non-bold text at the fallback size must not become a heading.
assert_eq!(detect_header_level(11.0, 10.0, &tiers, false), None);
// Non-tier body text stays regular.
assert_eq!(detect_header_level(10.0, 10.0, &tiers, true), None);
}
#[test]
fn bold_fallback_skipped_when_real_tiers_exist() {
let lines = vec![
line_of("Chapter One", 18.0, false, 700.0),
line_of("bold label", 11.0, true, 600.0),
line_of("body", 10.0, false, 580.0),
];
let tiers = compute_heading_tiers(&lines, 10.0);
assert_eq!(tiers, vec![18.0]);
// The 11pt bold label does not match any tier and stays non-heading.
assert_eq!(detect_header_level(11.0, 10.0, &tiers, true), None);
}
#[test]
fn toc_entry_with_single_dot_group() {
assert!(is_toc_entry_line("Measurement Lab worksheet ... 3"));
assert!(is_toc_entry_line("Results ........ 12"));
assert!(is_toc_entry_line("Appendix B...42"));
}
#[test]
fn non_toc_lines_pass() {
assert!(!is_toc_entry_line(
"6.2. Expectations for Re-Hiring Employees"
));
assert!(!is_toc_entry_line("What happened in 2020"));
assert!(!is_toc_entry_line("IMPLEMENTATION"));
// Ellipsis without a trailing page number
assert!(!is_toc_entry_line("and so it goes ..."));
// Long numbers are data, not page refs
assert!(!is_toc_entry_line("ISBN ... 97814"));
}
#[test]
fn toc_marker_headings() {
assert!(is_toc_marker_heading("Contents"));
assert!(is_toc_marker_heading("CONTENTS"));
assert!(is_toc_marker_heading("Table of Contents"));
assert!(is_toc_marker_heading("Table of contents:"));
assert!(!is_toc_marker_heading("Contents of the Shipment"));
assert!(!is_toc_marker_heading("Introduction"));
}
#[test]
fn heading_fragments() {
// Equation lead-ins: colon ending + inline equation reference
assert!(is_heading_fragment("or inversely"));
assert!(is_heading_fragment("and therefore"));
assert!(!is_heading_fragment("Introduction"));
assert!(!is_heading_fragment("iPhone Sales Strategy Overview")); // 4 words, exempt
assert!(is_heading_fragment("Rearranging Equation (8) gives:"));
// Display-equation neighbours ending in an equation number
assert!(is_heading_fragment("S = kB ln W, (2)"));
assert!(is_heading_fragment("E = mc2 (12)"));
assert!(is_heading_fragment("x + y = z, (3)"));
// Page-of-total running headers
assert!(is_heading_fragment("LIVSMEDELSVERKET PM 2 (10)"));
// Comparison-operator evidence and colon-before-number
assert!(is_heading_fragment(
"PLL\u{fe} PHH\u{226a} PLH\u{fe} PHL: (12)"
));
// Real headings pass — including name-plus-number and colon-ended ones
assert!(!is_heading_fragment("Nicaea (325)"));
assert!(!is_heading_fragment(
"\u{627}\u{644}\u{645}\u{644}\u{62d}\u{642} \u{631}\u{642}\u{645} (1)"
));
assert!(!is_heading_fragment("4. Entropy"));
assert!(!is_heading_fragment("Procedure:"));
assert!(!is_heading_fragment("Steps for Using the Microscope:"));
assert!(!is_heading_fragment("Changing objectives:"));
assert!(!is_heading_fragment("Sales by Region (2024)"));
assert!(!is_heading_fragment("Results (preliminary)"));
}
}
+12 -4
View File
@@ -131,10 +131,11 @@ pub(crate) fn format_list_item(text: &str) -> String {
if let Some(rest) = trimmed.strip_prefix(*bullet) {
return format!("- {}", rest.trim_start());
}
// Bullet inside a leading bold/italic run (e.g. "**● Label:** rest").
// The run wraps both the marker and the following label because both
// use a bold font in the PDF.
for wrapper in ["**", "*"] {
// Bullet inside a leading style run (e.g. "**● Label:** rest" or
// "<u>● Label</u>"). The run wraps both the marker and the following
// label because both carry the style in the PDF. The marker must move
// outside the wrapper so markdown still sees a list item.
for wrapper in ["**", "*", "<u>"] {
if let Some(after_open) = trimmed.strip_prefix(wrapper) {
if let Some(rest) = after_open.strip_prefix(*bullet) {
return format!("- {}{}", wrapper, rest.trim_start());
@@ -235,6 +236,13 @@ mod tests {
assert_eq!(format_list_item("• Item"), "- Item");
}
#[test]
fn format_list_item_bullet_inside_underline() {
// Fully-underlined bullet line: the marker must move outside the
// <u> wrapper so markdown still renders a list item.
assert_eq!(format_list_item("<u>● Item text</u>"), "- <u>Item text</u>");
}
#[test]
fn format_list_item_bullet_inside_bold() {
// PDF that uses bold font for both the marker and the label produces
+840 -98
View File
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+1452 -99
View File
File diff suppressed because it is too large Load Diff
+136 -67
View File
@@ -2,12 +2,16 @@
use regex::Regex;
use super::MarkdownOptions;
use super::{MarkdownOptions, MarkdownProfile};
use crate::text_utils::is_page_number_line;
/// Clean up markdown output with post-processing
pub(crate) fn clean_markdown(mut text: String, options: &MarkdownOptions) -> String {
// Collapse dot leaders (e.g. TOC entries: "Introduction...............................1")
text = collapse_dot_leaders(&text);
if options.profile == MarkdownProfile::Compact {
// Dot-leader collapse saves tokens but changes source text, so it is
// reserved for the explicit compact profile.
text = collapse_dot_leaders(&text);
}
// Fix hyphenation first (before other processing)
if options.fix_hyphenation {
@@ -29,6 +33,8 @@ pub(crate) fn clean_markdown(mut text: String, options: &MarkdownOptions) -> Str
// text item, which combine with gap-based space insertion to produce
// double spaces ("Vice President" instead of "Vice President").
collapse_consecutive_spaces(&mut text);
remove_spaces_before_closing_brackets(&mut text);
remove_spaces_before_sentence_punctuation(&mut text);
// Remove excessive newlines (more than 2 in a row)
while text.contains("\n\n\n") {
@@ -71,6 +77,46 @@ fn collapse_consecutive_spaces(text: &mut String) {
*text = result;
}
/// Remove spaces before closing square brackets.
/// Unit markers and markdown links occasionally pick up a gap-inserted space
/// before `]` (e.g. `[kg/m3 ]`), which is cosmetic padding.
fn remove_spaces_before_closing_brackets(text: &mut String) {
let mut result = String::with_capacity(text.len());
for ch in text.chars() {
if ch == ']' && result.ends_with(' ') {
result.pop();
}
result.push(ch);
}
*text = result;
}
/// Remove a stray space before sentence punctuation ("word ." → "word.").
/// Style-boundary item splits (bold/italic/underline runs) can strand a
/// trailing period or comma in its own fragment, and several assembly paths
/// join fragments with spaces. Only fires when the punctuation ends the
/// token (followed by whitespace or end of text), so decimals ("3 .14" stays
/// untouched — no such input exists, but the guard is cheap) and dot leaders
/// (" ... ") are unaffected.
fn remove_spaces_before_sentence_punctuation(text: &mut String) {
let chars: Vec<char> = text.chars().collect();
let mut result = String::with_capacity(text.len());
for (i, &ch) in chars.iter().enumerate() {
if matches!(ch, '.' | ',' | ';') && result.ends_with(' ') {
let next = chars.get(i + 1);
// `|` counts as a token end so table cells get the same fix.
let token_ends = next.is_none_or(|c| c.is_whitespace() || *c == '|');
// Never touch runs of dots (ellipsis / dot leaders).
let in_dot_run = ch == '.' && next == Some(&'.');
if token_ends && !in_dot_run {
result.pop();
}
}
result.push(ch);
}
*text = result;
}
/// Collapse dot leaders (runs of 4+ dots) into " ... "
/// Common in tables of contents: "Introduction...............................1" -> "Introduction ... 1"
fn collapse_dot_leaders(text: &str) -> String {
@@ -100,7 +146,7 @@ fn fix_hyphenation(text: &str) -> String {
result
}
/// Remove standalone page numbers (lines that are just 1-4 digit numbers)
/// Remove isolated page-number expressions from Markdown.
fn remove_page_numbers(text: &str) -> String {
let mut result = Vec::new();
let lines: Vec<&str> = text.lines().collect();
@@ -138,69 +184,6 @@ fn remove_page_numbers(text: &str) -> String {
result.join("\n")
}
/// Check if a line looks like a page number
fn is_page_number_line(trimmed: &str) -> bool {
// Empty lines are not page numbers
if trimmed.is_empty() {
return false;
}
// Pattern 1: Just a number (1-4 digits)
if trimmed.len() <= 4 && trimmed.chars().all(|c| c.is_ascii_digit()) {
return true;
}
// Pattern 2: "Page X of Y" or "Page X" or "Page of" (placeholder)
let lower = trimmed.to_lowercase();
if let Some(rest) = lower.strip_prefix("page") {
let rest = rest.trim();
// "Page of" (empty page numbers)
if rest == "of" || rest.starts_with("of ") {
return true;
}
// "Page X" or "Page X of Y"
if rest
.chars()
.next()
.map(|c| c.is_ascii_digit())
.unwrap_or(false)
{
return true;
}
// Just "Page" followed by whitespace and maybe "of"
if rest.is_empty()
|| rest
.split_whitespace()
.all(|w| w == "of" || w.chars().all(|c| c.is_ascii_digit()))
{
return true;
}
}
// Pattern 3: "X of Y" where X and Y are numbers
if let Some(of_idx) = trimmed.find(" of ") {
let before = trimmed[..of_idx].trim();
let after = trimmed[of_idx + 4..].trim();
if before.chars().all(|c| c.is_ascii_digit())
&& after.chars().all(|c| c.is_ascii_digit())
&& !before.is_empty()
&& !after.is_empty()
{
return true;
}
}
// Pattern 4: "- X -" centered page number
if trimmed.len() >= 3 && trimmed.starts_with('-') && trimmed.ends_with('-') {
let inner = trimmed[1..trimmed.len() - 1].trim();
if inner.chars().all(|c| c.is_ascii_digit()) && !inner.is_empty() {
return true;
}
}
false
}
/// Convert URLs to markdown links
fn format_urls(text: &str) -> String {
use once_cell::sync::Lazy;
@@ -313,6 +296,23 @@ fn format_urls(text: &str) -> String {
mod tests {
use super::*;
#[test]
fn fidelity_profile_preserves_dot_leaders() {
let input = "Introduction............................1".to_string();
let result = clean_markdown(input.clone(), &MarkdownOptions::default());
assert_eq!(result, format!("{input}\n"));
}
#[test]
fn compact_profile_collapses_dot_leaders() {
let input = "Introduction............................1".to_string();
let options = MarkdownOptions {
profile: MarkdownProfile::Compact,
..MarkdownOptions::default()
};
assert_eq!(clean_markdown(input, &options), "Introduction ... 1\n");
}
// --- collapse_dot_leaders ---
#[test]
@@ -342,6 +342,48 @@ mod tests {
assert!(result.contains("Chapter 2 ... 20"));
}
// --- remove_spaces_before_closing_brackets ---
#[test]
fn test_remove_spaces_before_closing_brackets() {
let mut input = "Density [kg/m3 ] and [linked text ](https://example.com)".to_string();
remove_spaces_before_closing_brackets(&mut input);
assert_eq!(
input,
"Density [kg/m3] and [linked text](https://example.com)"
);
}
// --- remove_spaces_before_sentence_punctuation ---
#[test]
fn strips_space_before_trailing_period() {
let mut t = "Foreign insurance companies . The provisions".to_string();
remove_spaces_before_sentence_punctuation(&mut t);
assert_eq!(t, "Foreign insurance companies. The provisions");
}
#[test]
fn strips_space_before_period_at_cell_boundary() {
let mut t = "|Applicability date .|This section|".to_string();
remove_spaces_before_sentence_punctuation(&mut t);
assert_eq!(t, "|Applicability date.|This section|");
}
#[test]
fn keeps_dot_leaders_and_ellipses() {
let mut t = "Introduction ... 1".to_string();
remove_spaces_before_sentence_punctuation(&mut t);
assert_eq!(t, "Introduction ... 1");
}
#[test]
fn keeps_mid_token_periods() {
let mut t = "version 3 .14 released".to_string();
remove_spaces_before_sentence_punctuation(&mut t);
assert_eq!(t, "version 3 .14 released");
}
// --- fix_hyphenation ---
#[test]
@@ -385,12 +427,14 @@ mod tests {
fn test_is_page_number_page_x() {
assert!(is_page_number_line("Page 5"));
assert!(is_page_number_line("page 12"));
assert!(is_page_number_line("Page123"));
}
#[test]
fn test_is_page_number_page_x_of_y() {
assert!(is_page_number_line("Page 3 of 10"));
assert!(is_page_number_line("page 1 of 5"));
assert!(is_page_number_line("Page 3 of 10 Report header"));
}
#[test]
@@ -420,6 +464,13 @@ mod tests {
assert!(!is_page_number_line("Hello World"));
assert!(!is_page_number_line("Chapter 1"));
assert!(!is_page_number_line("Total: 500"));
assert!(!is_page_number_line("PAGE0-PARA2-END-MARKER-0"));
}
#[test]
fn test_is_page_number_labeled_running_header() {
assert!(is_page_number_line("Page 42 Chapter 5"));
assert!(is_page_number_line("Page 42 explains the result"));
}
// --- remove_page_numbers ---
@@ -447,6 +498,24 @@ mod tests {
assert!(result.contains("42"));
}
#[test]
fn test_remove_page_numbers_labeled_header_with_content() {
let input = "Content\n\nPage 42 explains the result\n---\nEnd";
let result = remove_page_numbers(input);
assert!(!result.contains("Page 42 explains the result"));
assert!(result.contains("Content"));
assert!(result.contains("End"));
}
#[test]
fn test_remove_page_numbers_preserves_page_prefixed_content() {
let input = "PAGE0-PARA2-START substantive report text PAGE0-PARA2-END-MARKER-0";
let result = remove_page_numbers(input);
assert_eq!(result, input);
}
#[test]
fn test_remove_page_numbers_multiple_patterns() {
let input = "\n1\n\nContent\n\n2\n\n---\nMore\n\n3\n";
+108 -2
View File
@@ -3,7 +3,7 @@
use std::collections::{HashMap, HashSet};
use crate::structure_tree::StructRole;
use crate::types::TextLine;
use crate::types::{TextItem, TextLine};
use super::analysis::detect_header_level;
@@ -42,7 +42,12 @@ fn effective_heading_level(
// Fall back to font-size heuristic
let font = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
detect_header_level(font, base_size, heading_tiers)
detect_header_level(
font,
base_size,
heading_tiers,
crate::markdown::analysis::line_is_mostly_bold(line),
)
}
/// Merge consecutive heading lines at the same level into a single line.
@@ -87,6 +92,41 @@ pub(crate) fn merge_heading_lines(
false
};
// Bold headings at body font size never reach a tier, so wrapped ones
// split into two output headings ("…of wood pellets and cost" /
// "structure in Japan"). Merge a fully-bold line into the previous
// fully-bold line when it reads as a wrap continuation: starts
// lowercase, tiny Y gap, and the previous line has no terminal
// punctuation. Kept deliberately narrow — bold list labels and bold
// sentences start with markers or capitals and are unaffected.
let should_merge = should_merge
|| if let Some(prev) = result.last() {
let all_bold = |l: &TextLine| {
!l.items.is_empty() && l.items.iter().all(|i: &TextItem| i.is_bold)
};
let prev_text = prev.text();
let prev_trim = prev_text.trim_end();
let curr_text = line.text();
let curr_trim = curr_text.trim();
let y_gap = prev.y - line.y;
// Both lines must be tier-less: a tiered/tagged bold heading
// followed by bold body text must not absorb it.
line_level.is_none()
&& effective_heading_level(prev, base_size, heading_tiers, struct_roles)
.is_none()
&& prev.page == line.page
&& all_bold(prev)
&& all_bold(&line)
&& y_gap > 0.0
&& y_gap < line_font * 1.6
&& curr_trim.chars().next().is_some_and(|c| c.is_lowercase())
&& !prev_trim.ends_with(['.', ':', ';', '!', '?'])
&& prev_trim.split_whitespace().count() + curr_trim.split_whitespace().count()
<= 20
} else {
false
};
if should_merge {
// Append this line's items to the previous line
let prev = result.last_mut().unwrap();
@@ -542,6 +582,8 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid,
}
@@ -683,4 +725,68 @@ mod tests {
.unwrap();
assert_eq!(first_header.page, 1, "first occurrence should be on page 1");
}
fn make_bold_line(text: &str, page: u32, y: f32) -> TextLine {
let mut item = make_item(text, 12.0, None);
item.is_bold = true;
TextLine {
items: vec![item],
y,
page,
adaptive_threshold: 0.10,
}
}
#[test]
fn merge_wrapped_bold_heading_lowercase_continuation() {
// Bold-at-body-size heading wrapped across two lines: the second line
// starts lowercase and must merge into the first.
let lines = vec![
make_bold_line(
"3. Perspective of supply and demand balance and cost",
1,
700.0,
),
make_bold_line("structure in Japan", 1, 686.0),
make_line("Body text paragraph follows here.", 12.0, 1, 660.0, None),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "wrapped bold heading should merge");
assert!(result[0].text().contains("cost structure in Japan"));
}
#[test]
fn no_merge_for_bold_sentences_or_new_headings() {
// Second bold line starts with a capital — a new heading or label,
// not a wrap continuation.
let lines = vec![
make_bold_line("Replace", 1, 700.0),
make_bold_line("Trash", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "distinct bold lines must not merge");
// Previous line ends a sentence — continuation must not merge.
let lines = vec![
make_bold_line("This is a bold sentence.", 1, 700.0),
make_bold_line("another bold line", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "sentence-final bold line must not merge");
}
#[test]
fn tiered_bold_heading_does_not_absorb_bold_body() {
// Previous line is a tier-level bold heading (16pt vs 12pt body);
// a following lowercase bold body line must NOT merge into it.
let mut heading = make_bold_line("Section Title", 1, 700.0);
heading.items[0].font_size = 16.0;
heading.items[0].height = 16.0;
let lines = vec![
heading,
make_bold_line("emphasized body text continues here", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[16.0], None);
assert_eq!(result.len(), 2, "tiered heading must not absorb bold body");
}
}
+144
View File
@@ -30,6 +30,9 @@ pub struct PyPdfResult {
/// 1-indexed page numbers that need OCR.
#[pyo3(get)]
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
#[pyo3(get)]
pub ocr_reasons_by_page: Vec<PyPageOcrReasons>,
/// Title from PDF metadata.
#[pyo3(get)]
pub title: Option<String>,
@@ -60,6 +63,28 @@ impl PyPdfResult {
}
}
/// OCR reasons for a single 1-indexed page.
#[pyclass(name = "PageOcrReasons")]
#[derive(Clone)]
pub struct PyPageOcrReasons {
/// 1-indexed page number.
#[pyo3(get)]
pub page: u32,
/// Machine-readable OCR reason identifiers.
#[pyo3(get)]
pub reasons: Vec<String>,
}
#[pymethods]
impl PyPageOcrReasons {
fn __repr__(&self) -> String {
format!(
"PageOcrReasons(page={}, reasons={:?})",
self.page, self.reasons
)
}
}
// ---------------------------------------------------------------------------
// Classification wrapper (lightweight)
// ---------------------------------------------------------------------------
@@ -106,6 +131,9 @@ pub struct PyRegionText {
/// True when the text should not be trusted (empty, GID fonts, garbage, encoding issues).
#[pyo3(get)]
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
#[pyo3(get)]
pub ocr_reason: Option<String>,
}
#[pymethods]
@@ -160,6 +188,9 @@ pub struct PyPageMarkdown {
/// encoding issues, garbage text, or empty extraction).
#[pyo3(get)]
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
#[pyo3(get)]
pub ocr_reason: Option<String>,
}
#[pymethods]
@@ -190,6 +221,9 @@ pub struct PyPagesExtractionResult {
/// 1-indexed pages that need OCR (scanned/image-based or unreliable text).
#[pyo3(get)]
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
#[pyo3(get)]
pub ocr_reasons_by_page: Vec<PyPageOcrReasons>,
/// True if any page has tables or columns.
#[pyo3(get)]
pub is_complex: bool,
@@ -232,7 +266,17 @@ pub struct PyTextItem {
#[pyo3(get)]
pub is_italic: bool,
#[pyo3(get)]
pub is_underline: bool,
#[pyo3(get)]
pub is_strikeout: bool,
#[pyo3(get)]
pub item_type: String,
/// Marked Content ID from the content stream's BDC/BMC operator, None
/// when the text is not part of marked content. Join with the
/// (page, mcid) pairs from extract_structure_elements to attach
/// structure-tree roles (headings, paragraphs, ...) in tagged PDFs.
#[pyo3(get)]
pub mcid: Option<i64>,
}
#[pymethods]
@@ -248,6 +292,32 @@ impl PyTextItem {
}
}
/// One structure-tree element reference from a tagged PDF.
#[pyclass(name = "StructureElement")]
#[derive(Clone)]
pub struct PyStructureElement {
/// 1-indexed page number (matches TextItem.page).
#[pyo3(get)]
pub page: u32,
/// Marked Content ID from the page's content stream (matches
/// TextItem.mcid).
#[pyo3(get)]
pub mcid: i64,
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", ...).
#[pyo3(get)]
pub role: String,
}
#[pymethods]
impl PyStructureElement {
fn __repr__(&self) -> String {
format!(
"StructureElement(page={}, mcid={}, role='{}')",
self.page, self.mcid, self.role
)
}
}
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
@@ -268,6 +338,7 @@ fn to_py_result(r: crate::PdfProcessResult) -> PyPdfResult {
page_count: r.page_count,
processing_time_ms: r.processing_time_ms,
pages_needing_ocr: r.pages_needing_ocr,
ocr_reasons_by_page: to_py_page_ocr_reasons(r.ocr_reasons_by_page),
title: r.title,
confidence: r.confidence,
is_complex_layout: r.layout.is_complex,
@@ -277,6 +348,16 @@ fn to_py_result(r: crate::PdfProcessResult) -> PyPdfResult {
}
}
fn to_py_page_ocr_reasons(reasons: Vec<crate::PageOcrReasons>) -> Vec<PyPageOcrReasons> {
reasons
.into_iter()
.map(|reason| PyPageOcrReasons {
page: reason.page,
reasons: reason.reasons,
})
.collect()
}
fn to_py_err(e: crate::PdfError) -> PyErr {
PyValueError::new_err(e.to_string())
}
@@ -304,7 +385,21 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item_type_str(&item.item_type),
mcid: item.mcid,
})
.collect()
}
fn convert_structure_elements(elements: Vec<crate::StructureElement>) -> Vec<PyStructureElement> {
elements
.into_iter()
.map(|e| PyStructureElement {
page: e.page,
mcid: e.mcid,
role: e.role,
})
.collect()
}
@@ -350,11 +445,13 @@ fn to_py_pages_result(r: crate::PagesExtractionResult) -> PyPagesExtractionResul
page: p.page,
markdown: p.markdown,
needs_ocr: p.needs_ocr,
ocr_reason: p.ocr_reason,
})
.collect(),
pages_with_tables: r.pages_with_tables,
pages_with_columns: r.pages_with_columns,
pages_needing_ocr: r.pages_needing_ocr,
ocr_reasons_by_page: to_py_page_ocr_reasons(r.ocr_reasons_by_page),
is_complex: r.is_complex,
}
}
@@ -370,6 +467,7 @@ fn convert_region_results(results: Vec<crate::PageRegionResult>) -> Vec<PyPageRe
.map(|r| PyRegionText {
text: r.text,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
})
@@ -559,12 +657,56 @@ fn extract_pages_markdown_bytes(
Ok(to_py_pages_result(result))
}
/// Extract structure-tree element references from a tagged PDF file.
///
/// Parses the document's structure tree (when present) and returns one
/// entry per marked-content reference, resolved to its 1-indexed page,
/// MCID, and structure type name ("H1".."H6", "P", "Table", ...). Returns
/// an empty list when the PDF is not tagged.
///
/// Join (page, mcid) against the page/mcid attributes from
/// [`extract_text_with_positions`] to attach heading levels and other
/// semantic roles to extracted text.
///
/// Args:
/// path: Path to the PDF file.
/// pages: Optional list of 1-indexed pages (matching TextItem.page).
/// When None (default), the whole document is returned.
///
/// Returns:
/// List of StructureElement sorted by (page, mcid).
#[pyfunction]
#[pyo3(signature = (path, pages=None))]
fn extract_structure_elements(
path: &str,
pages: Option<Vec<u32>>,
) -> PyResult<Vec<PyStructureElement>> {
let elements = crate::extract_structure_elements(path, pages.as_deref()).map_err(to_py_err)?;
Ok(convert_structure_elements(elements))
}
/// Extract structure-tree element references from tagged PDF bytes.
///
/// See [`extract_structure_elements`] for details.
#[pyfunction]
#[pyo3(signature = (data, pages=None))]
fn extract_structure_elements_bytes(
data: &[u8],
pages: Option<Vec<u32>>,
) -> PyResult<Vec<PyStructureElement>> {
let elements =
crate::extract_structure_elements_mem(data, pages.as_deref()).map_err(to_py_err)?;
Ok(convert_structure_elements(elements))
}
/// Python module definition.
#[pymodule]
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_class::<PyPdfResult>()?;
m.add_class::<PyPageOcrReasons>()?;
m.add_class::<PyPdfClassification>()?;
m.add_class::<PyTextItem>()?;
m.add_class::<PyStructureElement>()?;
m.add_class::<PyRegionText>()?;
m.add_class::<PyPageRegionTexts>()?;
m.add_class::<PyPageMarkdown>()?;
@@ -579,6 +721,8 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_function(wrap_pyfunction!(extract_text_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_with_positions, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_with_positions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_structure_elements, m)?)?;
m.add_function(wrap_pyfunction!(extract_structure_elements_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_pages_markdown, m)?)?;
+913 -38
View File
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+2563 -23
View File
File diff suppressed because it is too large Load Diff
+1453 -14
View File
File diff suppressed because it is too large Load Diff
+2
View File
@@ -586,6 +586,8 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid,
}
+2
View File
@@ -108,6 +108,8 @@ pub(crate) fn try_split_financial_item(item: &TextItem) -> Option<Vec<TextItem>>
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
+62 -2
View File
@@ -123,6 +123,7 @@ fn format_toc_as_list(cells: &[Vec<String>], footnotes: &[String]) -> String {
/// True when the cell looks like a page number. Accepts:
/// - plain digit tokens: "42", "86 86"
/// - canonical roman numerals (front-matter pages): "vii", "ix", "xii"
/// - dashed section-page IDs: "5-21", "A-1", "B--3", "TC-2" (common in
/// technical manuals)
fn is_page_number_cell(cell: &str) -> bool {
@@ -138,6 +139,9 @@ fn is_page_number_cell(cell: &str) -> bool {
if all_digits {
return t.len() <= 4;
}
if super::canonical_roman_value(t).is_some() {
return true;
}
// Section-page form: uppercase letters, digits, dashes; at least
// one digit present.
t.chars()
@@ -184,6 +188,21 @@ fn starts_with_numbered_label(cell: &str) -> bool {
.is_some_and(|c| matches!(c, '.' | ')' | '-' | ':'))
}
fn starts_with_hierarchical_numbered_label(cell: &str) -> bool {
let token = cell
.split_whitespace()
.next()
.unwrap_or("")
.trim_end_matches(['.', ')', ':', '-']);
let levels: Vec<&str> = token.split('.').collect();
(2..=4).contains(&levels.len())
&& levels.iter().all(|level| {
!level.is_empty()
&& level.len() <= 3
&& level.chars().all(|character| character.is_ascii_digit())
})
}
fn alpha_word_count(cell: &str) -> usize {
cell.split_whitespace()
.filter(|word| word.chars().any(|c| c.is_alphabetic()))
@@ -337,11 +356,12 @@ fn clean_table_cells(cells: &[Vec<String>]) -> (Vec<Vec<String>>, Vec<String>) {
// mid-sentence/lowercase ("continued text here", "with 3.5%...") or
// carry lowercase fragments in the later cells, so keep those mergeable.
let looks_like_hierarchical_subrow = first_cell.is_empty()
&& row.len() >= 3
&& first_non_empty_col == Some(1)
&& looks_like_compact_entry_label(first_non_empty_cell)
&& ((non_first_cells.len() >= 2 && title_like_later_cells > 0)
&& ((row.len() == 2 && starts_with_hierarchical_numbered_label(first_non_empty_cell))
|| (row.len() >= 3 && non_first_cells.len() >= 2 && title_like_later_cells > 0)
|| (non_first_cells.len() == 1
&& row.len() >= 3
&& prev_first_cell_empty
&& alpha_word_count(first_non_empty_cell) >= 2));
let looks_like_new_first_column_entry = !first_cell.is_empty()
@@ -684,6 +704,46 @@ mod tests {
assert_eq!(cleaned[4][1], "Model training");
}
#[test]
fn test_clean_table_cells_two_column_numbered_subrows_not_merged() {
let cells = vec![
vec!["Area".into(), "Competence".into()],
vec![
"1. Embodying sustainability values".into(),
"1.1 Valuing sustainability".into(),
],
vec!["".into(), "1.2 Supporting fairness".into()],
vec!["".into(), "1.3 Promoting nature".into()],
vec![
"2. Embracing complexity".into(),
"2.1 Systems thinking".into(),
],
vec!["".into(), "2.2 Critical thinking".into()],
];
let (cleaned, _) = clean_table_cells(&cells);
assert_eq!(cleaned.len(), 6);
assert_eq!(cleaned[2], vec!["", "1.2 Supporting fairness"]);
assert_eq!(cleaned[3], vec!["", "1.3 Promoting nature"]);
assert_eq!(cleaned[5], vec!["", "2.2 Critical thinking"]);
}
#[test]
fn test_clean_table_cells_two_column_numbered_continuation_merges() {
let cells = vec![
vec!["Area".into(), "Requirement".into()],
vec!["Safety".into(), "The program includes".into()],
vec!["".into(), "1. First requirement for every operator".into()],
];
let (cleaned, _) = clean_table_cells(&cells);
assert_eq!(cleaned.len(), 2);
assert_eq!(
cleaned[1][1],
"The program includes 1. First requirement for every operator"
);
}
#[test]
fn test_clean_table_cells_partial_hierarchical_subrow_not_merged() {
let cells = vec![
+51 -1
View File
@@ -369,6 +369,10 @@ pub(crate) fn join_cell_items(items: &[&TextItem]) -> String {
let prev_ends_with_hyphen = result.ends_with('-');
let curr_is_hyphen = text == "-";
let curr_starts_with_hyphen = text.starts_with('-');
let prev_ends_with_open_delimiter =
result.ends_with('(') || result.ends_with('[') || result.ends_with('{');
let curr_starts_with_close_delimiter =
text.starts_with(')') || text.starts_with(']') || text.starts_with('}');
// Detect subscript/superscript: smaller font size and/or Y offset
let font_ratio = item.font_size / prev_item.font_size;
@@ -385,6 +389,8 @@ pub(crate) fn join_cell_items(items: &[&TextItem]) -> String {
|| curr_starts_with_hyphen
|| is_sub_super
|| was_sub_super
|| prev_ends_with_open_delimiter
|| curr_starts_with_close_delimiter
{
result.push_str(text);
} else {
@@ -430,7 +436,8 @@ pub(crate) fn recover_header_row(
.iter()
.enumerate()
.filter(|(_, item)| {
item.font_size > small_font_threshold
!item.is_strikeout
&& item.font_size > small_font_threshold
&& item.y > first_row_y
&& item.y <= first_row_y + row_gap_limit
})
@@ -514,6 +521,8 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -727,6 +736,18 @@ mod tests {
assert_eq!(join_cell_items(&[&a, &b, &c]), "pre-fix");
}
#[test]
fn test_join_cell_items_parenthetical_no_inner_spaces() {
let a = make_item("The first sentence", 100.0, 500.0, 10.0);
let b = make_item("(", 190.0, 500.0, 10.0);
let c = make_item("twice", 195.0, 500.0, 10.0);
let d = make_item(")", 220.0, 500.0, 10.0);
assert_eq!(
join_cell_items(&[&a, &b, &c, &d]),
"The first sentence (twice)"
);
}
#[test]
fn test_join_cell_items_subscript_no_space() {
let a = make_item("H", 100.0, 500.0, 12.0);
@@ -784,6 +805,31 @@ mod tests {
assert_eq!(table.rows.len(), rows_before);
}
#[test]
fn test_recover_header_row_skips_strikeout_candidates() {
let mut old_col1 = make_item("Old Col1", 100.0, 520.0, 12.0);
old_col1.is_strikeout = true;
let mut old_col2 = make_item("Old Col2", 200.0, 520.0, 12.0);
old_col2.is_strikeout = true;
let all_items = vec![
old_col1,
old_col2,
make_item("A", 100.0, 500.0, 8.0),
make_item("B", 200.0, 500.0, 8.0),
];
let mut table = Table {
columns: vec![100.0, 200.0],
rows: vec![500.0, 480.0],
cells: vec![vec!["A".into(), "B".into()], vec!["C".into(), "D".into()]],
item_indices: vec![2, 3],
kind: TableKind::Data,
};
recover_header_row(&mut table, &all_items, 9.0);
assert_eq!(table.rows.len(), 2);
assert_eq!(table.cells[0], vec!["A", "B"]);
}
#[test]
fn test_recover_header_row_too_far_above() {
let all_items = vec![
@@ -866,6 +912,8 @@ mod tests {
font: String::new(),
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
@@ -902,6 +950,8 @@ mod tests {
font: String::new(),
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
+1185 -3
View File
File diff suppressed because it is too large Load Diff
+520
View File
@@ -0,0 +1,520 @@
//! Text-quality detection: deciding when an extracted text layer is too broken
//! to serve and a page should fall back to OCR.
//!
//! Extraction can produce plausible-looking bytes that are actually garbage —
//! failed CID→Unicode mappings, broken ToUnicode CMaps, mojibake. These
//! detectors catch that and let callers set `needs_ocr`. They come in two
//! layers, sharing the same primitives:
//!
//! - **Markdown-level** ([`detect_encoding_issues`], [`is_garbage_text`],
//! [`is_cid_garbage`]) run on a page's final markdown string. Used as a
//! backstop on the region-extraction and whole-document paths.
//! - **Item/span-level** ([`analyze_text_quality`],
//! [`region_items_have_decoding_issue`]) run on individual `TextItem`s and
//! accumulate per-page evidence, so localized garbled spans on an otherwise
//! clean page are caught without a single span having to condemn the page.
//!
//! Detection classes, roughly by signal:
//! - **Replacement runs**: U+FFFD clusters ([`has_replacement_text_run`]).
//! - **Private-use / C1-control runs**: CID passthrough landing in PUA or the
//! C1 block ([`has_private_use_text_run`], [`has_cid_control_token`]).
//! - **Dollar-as-space**: `Word$Word$Word` from broken CMaps
//! ([`has_dollar_as_space_pattern`]).
//! - **Non-alphanumeric dominance**: symbol soup ([`is_garbage_text`]).
//! - **Substitution-cipher letter statistics**: pure-ASCII output whose letter
//! distribution is a permutation of natural language ([`CipherGarbleStats`]).
use crate::types::TextItem;
use crate::{add_ocr_reason, OCR_REASON_SUSPECTED_GARBLED_TEXT};
use std::collections::BTreeMap;
/// Detect broken font encodings in extracted markdown text.
///
/// Two heuristics:
/// 1. **U+FFFD**: Any replacement character indicates decode failures.
/// 2. **Dollar-as-space**: Pattern like `Word$Word$Word` where `$` is used as a
/// word separator due to broken ToUnicode CMaps. Triggers when either:
/// - More than 50% of `$` are between letters (clear substitution pattern), OR
/// - More than 20 letter-dollar-letter occurrences (even if some `$` are also
/// used as trailing/leading separators, 20+ is far beyond normal financial text).
pub(crate) fn detect_encoding_issues(markdown: &str) -> bool {
// Heuristic 1: U+FFFD replacement characters
if markdown.contains('\u{FFFD}') {
return true;
}
// Heuristic 2: dollar-as-space pattern
if has_dollar_as_space_pattern(markdown) {
return true;
}
// Heuristic 3: substitution-cipher letter statistics (broken ToUnicode)
let mut stats = CipherGarbleStats::default();
stats.add_text(markdown);
stats.looks_garbled()
}
fn has_dollar_as_space_pattern(markdown: &str) -> bool {
let total_dollars = markdown.matches('$').count();
if total_dollars > 10 {
let bytes = markdown.as_bytes();
let mut letter_dollar_letter = 0usize;
for i in 1..bytes.len().saturating_sub(1) {
if bytes[i] == b'$'
&& bytes[i - 1].is_ascii_alphabetic()
&& bytes[i + 1].is_ascii_alphabetic()
{
letter_dollar_letter += 1;
}
}
if letter_dollar_letter > 20 || letter_dollar_letter * 2 > total_dollars {
return true;
}
}
false
}
/// English letter frequencies (percent, az). Used as a natural-language
/// reference: every Latin-script language in the eval corpus (Swedish,
/// Finnish, Turkish, German, romaji) scores ≥ 0.80 cosine similarity against
/// it, while substitution-cipher text scores ~0.53.
const ENGLISH_LETTER_FREQ: [f64; 26] = [
8.2, 1.5, 2.8, 4.3, 12.7, 2.2, 2.0, 6.1, 7.0, 0.15, 0.8, 4.0, 2.4, 6.7, 7.5, 1.9, 0.1, 6.0,
6.3, 9.1, 2.8, 1.0, 2.4, 0.15, 2.0, 0.07,
];
/// Letter statistics for detecting substitution-cipher garbling: broken
/// ToUnicode CMaps that shift every character by a per-range constant (e.g.
/// `Certificate` extracted as `8VceZWZTReV`). Such text is 100% printable
/// ASCII with word-like token lengths, so it defeats `is_garbage_text` and
/// produces no replacement characters — it needs its own discriminator.
#[derive(Debug, Default)]
struct CipherGarbleStats {
/// Case-folded ASCII letter histogram.
letter_counts: [u32; 26],
ascii_letters: usize,
ascii_vowels: usize,
/// Accented Latin letters (Latin-1 Supplement through Latin Extended-B,
/// plus Latin Extended Additional). Count toward Latin dominance only.
latin_ext_letters: usize,
non_latin_letters: usize,
/// Adjacent ASCII-letter pairs, and how many of them switch from
/// lowercase straight to uppercase mid-word.
letter_bigrams: usize,
case_shift_bigrams: usize,
}
impl CipherGarbleStats {
fn add_text(&mut self, text: &str) {
let mut prev: Option<char> = None;
for ch in text.chars() {
if ch.is_ascii_alphabetic() {
let idx = (ch.to_ascii_lowercase() as u8 - b'a') as usize;
self.letter_counts[idx] += 1;
self.ascii_letters += 1;
if matches!(ch.to_ascii_lowercase(), 'a' | 'e' | 'i' | 'o' | 'u') {
self.ascii_vowels += 1;
}
if let Some(p) = prev {
self.letter_bigrams += 1;
if p.is_ascii_lowercase() && ch.is_ascii_uppercase() {
self.case_shift_bigrams += 1;
}
}
prev = Some(ch);
} else {
if ch.is_alphabetic() {
if matches!(ch as u32, 0xC0..=0x24F | 0x1E00..=0x1EFF) {
self.latin_ext_letters += 1;
} else {
self.non_latin_letters += 1;
}
}
prev = None;
}
}
}
/// Cosine similarity between the observed letter histogram and English
/// letter frequencies. A shifted alphabet permutes the histogram, which
/// destroys the similarity regardless of the shift amount.
fn english_cosine(&self) -> f64 {
if self.ascii_letters == 0 {
return 1.0;
}
let n = self.ascii_letters as f64;
let mut dot = 0.0;
let mut norm_obs = 0.0;
for (count, freq) in self.letter_counts.iter().zip(ENGLISH_LETTER_FREQ) {
let p = *count as f64 / n;
dot += p * freq;
norm_obs += p * p;
}
let norm_en = ENGLISH_LETTER_FREQ
.iter()
.map(|f| f * f)
.sum::<f64>()
.sqrt();
dot / (norm_obs.sqrt() * norm_en)
}
/// Cosine similarity between the observed histogram and English
/// frequencies after sorting BOTH descending — i.e. comparing the *shape*
/// of the frequency profile, ignoring which letter sits where. A
/// substitution cipher is a bijection, so it preserves this shape exactly
/// (att10k 0.97, arbitrary shifts 0.99) regardless of case or offset.
/// Non-linguistic ASCII has a different profile: a small alphabet is far
/// steeper (random DNA 0.74, hex dumps 0.81), so the shape diverges.
fn english_shape_cosine(&self) -> f64 {
if self.ascii_letters == 0 {
return 1.0;
}
let n = self.ascii_letters as f64;
let mut obs: [f64; 26] = std::array::from_fn(|i| self.letter_counts[i] as f64 / n);
obs.sort_unstable_by(|a, b| b.total_cmp(a));
let mut en = ENGLISH_LETTER_FREQ;
en.sort_unstable_by(|a, b| b.total_cmp(a));
let dot: f64 = obs.iter().zip(en).map(|(o, e)| o * e).sum();
let norm_obs = obs.iter().map(|o| o * o).sum::<f64>().sqrt();
let norm_en = en.iter().map(|e| e * e).sum::<f64>().sqrt();
dot / (norm_obs * norm_en)
}
/// Thresholds validated against the 380-document pdf-evals snapshot
/// corpus (0 false positives) and the garbled ParseBench `att10k` page
/// (vowel ratio 0.245, case-shift rate 0.225, cosine 0.532). Closest
/// legitimate document on each axis: vowel ratio 0.264 (circuit
/// schematic), case-shift rate 0.021, cosine 0.801.
fn looks_garbled(&self) -> bool {
// Need a statistically meaningful, Latin-dominant sample.
if self.ascii_letters < 200
|| self.non_latin_letters > self.ascii_letters + self.latin_ext_letters
{
return false;
}
// Real Latin-script text keeps vowels above ~30% of letters even in
// acronym- and part-number-heavy documents; shifted text starves them.
let vowel_ratio = self.ascii_vowels as f64 / self.ascii_letters as f64;
if vowel_ratio > 0.30 {
return false;
}
// Signal 1: lowercase→uppercase transitions inside words. A shifted
// lowercase alphabet straddles the ASCII uppercase block ('i'→'Z',
// 't'→'e'), so garbled words flip case constantly. Real documents
// stay ≤ 0.02 even with camelCase identifiers.
let case_shifts = self.letter_bigrams >= 100
&& self.case_shift_bigrams as f64 >= self.letter_bigrams as f64 * 0.10;
// Signal 2: the histogram is a permutation of natural language — an
// English-like frequency SHAPE (sorted cosine high) but with letters
// in the wrong POSITIONS (unsorted cosine low). This is the signature
// of a substitution cipher and is case-independent, so it catches
// all-lowercase and all-uppercase shifts as well as case-straddling
// ones. Genuinely non-linguistic ASCII that is merely "unlike English"
// fails one of the two halves: DNA/hex dumps have too steep a profile
// (shape cosine < 0.90), while protein sequences, ticker symbols and
// base64 are not sufficiently unlike English in position (unsorted
// cosine ≥ 0.60) — so none of them are routed to OCR.
let permuted_language = self.english_cosine() < 0.60 && self.english_shape_cosine() >= 0.90;
case_shifts || permuted_language
}
}
#[derive(Debug, Default)]
pub(crate) struct TextQualityReport {
pub(crate) pages_needing_ocr: Vec<u32>,
pub(crate) has_encoding_issues: bool,
pub(crate) reasons_by_page: BTreeMap<u32, Vec<String>>,
}
#[derive(Debug, Default)]
struct PageTextQualityEvidence {
chars: usize,
replacement_chars: usize,
replacement_spans: usize,
longest_replacement_run: usize,
cipher_garble: CipherGarbleStats,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum TextSpanIssueKind {
Replacement,
Strong,
}
pub(crate) fn analyze_text_quality(items: &[TextItem]) -> TextQualityReport {
let mut reasons_by_page = BTreeMap::new();
let mut evidence_by_page = BTreeMap::<u32, PageTextQualityEvidence>::new();
for item in items {
if !matches!(item.item_type, crate::types::ItemType::Text) {
continue;
}
let evidence = evidence_by_page.entry(item.page).or_default();
evidence.chars += item.text.chars().filter(|ch| !ch.is_whitespace()).count();
evidence.cipher_garble.add_text(&item.text);
match text_span_decoding_issue_kind(&item.text) {
Some(TextSpanIssueKind::Strong) => {
add_ocr_reason(
&mut reasons_by_page,
item.page,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
Some(TextSpanIssueKind::Replacement) => {
let stats = replacement_text_stats(&item.text);
evidence.replacement_chars += stats.0;
evidence.replacement_spans += 1;
evidence.longest_replacement_run = evidence.longest_replacement_run.max(stats.1);
}
None => {}
}
}
for (page, evidence) in evidence_by_page {
if reasons_by_page.contains_key(&page) {
continue;
}
if page_replacement_evidence_needs_ocr(&evidence) || evidence.cipher_garble.looks_garbled()
{
add_ocr_reason(
&mut reasons_by_page,
page,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
}
let pages_needing_ocr: Vec<u32> = reasons_by_page.keys().copied().collect();
TextQualityReport {
has_encoding_issues: !pages_needing_ocr.is_empty(),
pages_needing_ocr,
reasons_by_page,
}
}
pub(crate) fn region_items_have_decoding_issue(items: &[TextItem]) -> bool {
items.iter().any(|item| {
matches!(item.item_type, crate::types::ItemType::Text)
&& text_span_has_decoding_issue(&item.text)
})
}
fn text_span_has_decoding_issue(text: &str) -> bool {
text_span_decoding_issue_kind(text).is_some()
}
fn text_span_decoding_issue_kind(text: &str) -> Option<TextSpanIssueKind> {
let text = text.trim();
if text.is_empty() {
return None;
}
if has_dollar_as_space_pattern(text)
|| has_private_use_text_run(text)
|| is_cid_garbage(text)
|| has_cid_control_token(text)
{
return Some(TextSpanIssueKind::Strong);
}
if has_replacement_text_run(text) {
return Some(TextSpanIssueKind::Replacement);
}
None
}
fn replacement_text_stats(text: &str) -> (usize, usize) {
let mut replacement = 0usize;
let mut current_run = 0usize;
let mut longest_run = 0usize;
for ch in text.chars() {
if ch == '\u{FFFD}' {
replacement += 1;
current_run += 1;
longest_run = longest_run.max(current_run);
} else {
current_run = 0;
}
}
(replacement, longest_run)
}
fn page_replacement_evidence_needs_ocr(evidence: &PageTextQualityEvidence) -> bool {
if evidence.replacement_chars == 0 || evidence.chars == 0 {
return false;
}
// If the entire page is only a short broken text layer, even a short
// replacement run is enough evidence. On otherwise text-heavy pages,
// require density so math formulas do not force full-page OCR.
if evidence.chars <= 80 && evidence.longest_replacement_run >= 2 {
return true;
}
let replacement_density_bps = evidence.replacement_chars * 10_000 / evidence.chars;
let enough_bad_text = evidence.replacement_chars >= 12 && replacement_density_bps >= 500;
let repeated_bad_spans = evidence.replacement_spans >= 3 && replacement_density_bps >= 250;
let long_bad_run = evidence.longest_replacement_run >= 8 && replacement_density_bps >= 250;
enough_bad_text || repeated_bad_spans || long_bad_run
}
fn has_replacement_text_run(text: &str) -> bool {
let (replacement, longest_run) = replacement_text_stats(text);
longest_run >= 2 || replacement >= 3
}
fn has_private_use_text_run(text: &str) -> bool {
let mut total = 0usize;
let mut private_use = 0usize;
let mut current_run = 0usize;
let mut longest_run = 0usize;
for ch in text.chars() {
if ch.is_whitespace() {
current_run = 0;
continue;
}
total += 1;
if is_private_use_char(ch) {
private_use += 1;
current_run += 1;
longest_run = longest_run.max(current_run);
} else {
current_run = 0;
}
}
if private_use == 0 {
return false;
}
longest_run >= 3 || (total >= 5 && private_use >= 2 && private_use * 2 >= total)
}
fn has_cid_control_token(text: &str) -> bool {
text.split_whitespace().any(token_has_cid_control)
}
fn token_has_cid_control(token: &str) -> bool {
let mut total = 0usize;
let mut c1_control = 0usize;
for ch in token.chars() {
total += 1;
if ('\u{0080}'..='\u{009F}').contains(&ch) {
c1_control += 1;
}
}
total >= 5 && c1_control >= 2 && c1_control * 20 >= total
}
fn is_private_use_char(ch: char) -> bool {
matches!(
ch as u32,
0xE000..=0xF8FF | 0xF0000..=0xFFFFD | 0x100000..=0x10FFFD
)
}
/// Check if extracted text is predominantly garbage (non-alphanumeric).
///
/// Broken font encodings produce text like "----1-.-.-.___ --.-. .._ I_---."
/// where most characters are punctuation/symbols. Real text in any language
/// has >50% alphanumeric characters.
pub(crate) fn is_garbage_text(markdown: &str) -> bool {
let mut alphanum = 0usize;
let mut non_alphanum = 0usize;
let chars: Vec<char> = markdown.chars().collect();
let mut i = 0usize;
while i < chars.len() {
let ch = chars[i];
let mut run_end = i + 1;
while run_end < chars.len() && chars[run_end] == ch {
run_end += 1;
}
let is_decorative_leader = matches!(ch, '.' | '_' | '·') && run_end - i >= 3;
if !is_decorative_leader {
for &run_ch in &chars[i..run_end] {
if run_ch.is_whitespace() {
continue;
}
// Skip markdown syntax chars that we add (not from the PDF)
if matches!(run_ch, '#' | '*' | '|' | '-' | '\n') {
continue;
}
if run_ch.is_alphanumeric() {
alphanum += 1;
} else {
non_alphanum += 1;
}
}
}
i = run_end;
}
let total = alphanum + non_alphanum;
total >= 50 && alphanum * 2 < total
}
/// Detect garbage from failed CID-to-Unicode mapping on Identity-H fonts.
///
/// When CID values don't correspond to Unicode codepoints, the raw bytes often
/// produce characters in the C1 control range (U+0080U+009F) or Private Use
/// Area, mixed with random Latin Extended characters. Valid text in any
/// language almost never contains C1 controls. We also fall back to the
/// general `is_garbage_text` check for non-alphanumeric-heavy patterns.
pub(crate) fn is_cid_garbage(text: &str) -> bool {
if is_garbage_text(text) {
return true;
}
let mut total = 0usize;
let mut c1_control = 0usize;
let mut high_latin = 0usize;
for ch in text.chars() {
if ch.is_whitespace() {
continue;
}
total += 1;
// C1 control characters (U+0080U+009F) — almost never in real text
if ch == '·' {
continue;
}
if ('\u{0080}'..='\u{009F}').contains(&ch) {
c1_control += 1;
}
// High Latin-1 (U+00A0U+00FF) — legitimate in Western European text
// but when combined with ASCII in CID passthrough, indicates mojibake
// from CID values being misinterpreted as Latin-1 characters.
if ('\u{00A0}'..='\u{00FF}').contains(&ch) {
high_latin += 1;
}
}
if total < 5 {
return false;
}
// If ≥5% of non-whitespace chars are C1 controls, it's garbage
if c1_control >= 2 && c1_control * 20 >= total {
return true;
}
// If ≥40% of non-whitespace chars are high Latin-1 AND the text has few
// ASCII letters, it's likely CID-as-Latin-1 mojibake (Japanese/CJK PDFs
// where CID values 0x80-0xFF become accented Latin characters). Keep a
// minimum length so short math tokens like "2×()×" do not route a clean
// page to OCR.
let ascii_letters = text.chars().filter(|c| c.is_ascii_alphabetic()).count();
total >= 20 && high_latin * 5 >= total * 2 && ascii_letters * 3 < total
}
+97
View File
@@ -7,6 +7,83 @@
use crate::types::TextItem;
use unicode_normalization::UnicodeNormalization;
/// Return whether text is an explicit page-number expression.
///
/// This strict form is suitable before layout, where removing one numeric item
/// from substantive text such as `Page 42 explains the result` would lose data.
pub(crate) fn is_explicit_page_number_expression(text: &str) -> bool {
let trimmed = text.trim();
if trimmed.is_empty() {
return false;
}
let is_number = |value: &str| {
!value.is_empty() && value.chars().all(|character| character.is_ascii_digit())
};
if trimmed.len() <= 4 && is_number(trimmed) {
return true;
}
if trimmed.len() >= 3 && trimmed.starts_with('-') && trimmed.ends_with('-') {
let inner = trimmed[1..trimmed.len() - 1].trim();
if is_number(inner) {
return true;
}
}
let lowercase = trimmed.to_ascii_lowercase();
if let Some(rest) = lowercase.strip_prefix("page") {
let words: Vec<&str> = rest.split_whitespace().collect();
if words.len() >= 3 && is_number(words[0]) && words[1] == "of" && is_number(words[2]) {
return true;
}
if words.len() >= 2 && words[0] == "of" && is_number(words[1]) {
return true;
}
return match words.as_slice() {
[] | ["of"] => true,
[number] => is_number(number),
["of", total] => is_number(total),
[number, "of", total] => is_number(number) && is_number(total),
_ => false,
};
}
let words: Vec<&str> = lowercase.split_whitespace().collect();
match words.as_slice() {
[number, "of", total] => is_number(number) && is_number(total),
_ => false,
}
}
/// Return whether a completed Markdown line looks like a page number or a
/// labeled running header.
///
/// At this stage the complete line and surrounding breaks are available, so a
/// leading `Page N` remains compatible with the existing header cleanup even
/// when the PDF appends a chapter or document title.
pub(crate) fn is_page_number_line(text: &str) -> bool {
if is_explicit_page_number_expression(text) {
return true;
}
let lowercase = text.trim().to_ascii_lowercase();
lowercase.strip_prefix("page").is_some_and(|rest| {
let mut characters = rest.trim_start().chars().peekable();
let mut has_page_number = false;
while characters
.peek()
.is_some_and(|character| character.is_ascii_digit())
{
has_page_number = true;
characters.next();
}
has_page_number && characters.next().is_none_or(char::is_whitespace)
})
}
/// Check if a character is CJK (Chinese, Japanese, Korean).
/// CJK languages don't use spaces between words, so word-boundary
/// heuristics should not apply when CJK characters are involved.
@@ -93,6 +170,9 @@ pub fn is_bold_font(font_name: &str) -> bool {
|| lower.contains("extrabold")
|| lower.contains("ultrabold")
|| lower.contains("medium") && !lower.contains("mediumitalic") // Some fonts use Medium for semi-bold
// URW Type 1 fonts abbreviate Medium as "Medi" (e.g. NimbusRomNo9L-Medi,
// the Times-Bold substitute in LaTeX documents; -MediItal is bold italic).
|| lower.contains("-medi") && !lower.contains("mediumital")
}
/// Detect if a font name indicates italic/oblique style
@@ -762,6 +842,17 @@ mod tests {
use super::*;
use crate::types::ItemType;
#[test]
fn bold_font_urw_medi_abbreviation() {
// URW Type 1 fonts (LaTeX default Times) abbreviate Medium as "Medi"
assert!(is_bold_font("NROFIU+NimbusRomNo9L-Medi"));
assert!(is_bold_font("NimbusRomNo9L-MediItal"));
assert!(!is_bold_font("DSSZWN+NimbusRomNo9L-Regu"));
assert!(!is_bold_font("NimbusRomNo9L-ReguItal"));
// Medium-Italic exclusion still holds
assert!(!is_bold_font("Foo-MediumItalic"));
}
#[test]
fn strip_soft_hyphen() {
assert_eq!(expand_ligatures("con\u{00AD}tent"), "content");
@@ -883,6 +974,8 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1002,6 +1095,8 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -1078,6 +1173,8 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+330 -33
View File
@@ -4,11 +4,17 @@
use log::{debug, warn};
use lopdf::{Document, Object, ObjectId};
use std::borrow::Cow;
use std::collections::{HashMap, HashSet};
#[cfg(not(target_arch = "wasm32"))]
use std::path::{Path, PathBuf};
use crate::glyph_names::glyph_to_char;
#[cfg(target_arch = "wasm32")]
static BUILTIN_CMAPS: include_dir::Dir<'_> =
include_dir::include_dir!("$CARGO_MANIFEST_DIR/external/bcmaps");
/// A parsed ToUnicode CMap mapping CIDs to Unicode strings
#[derive(Debug, Default, Clone)]
pub struct ToUnicodeCMap {
@@ -329,7 +335,7 @@ impl ToUnicodeCMap {
if let (Some(start), Some(end), Some(base)) = (
parse_hex_u16(&start_hex),
parse_hex_u16(&end_hex),
parse_hex_u32(&base_hex),
hex_to_unicode_scalar(&base_hex),
) {
self.ranges.push((start, end, base));
}
@@ -575,32 +581,86 @@ fn parse_hex_u16(hex: &str) -> Option<u16> {
u16::from_str_radix(hex.trim(), 16).ok()
}
/// Parse a hex string to u32
fn parse_hex_u32(hex: &str) -> Option<u32> {
u32::from_str_radix(hex.trim(), 16).ok()
}
/// Convert a hex string to a Unicode string
/// Handles both 2-byte (BMP) and 4-byte (supplementary) codepoints
/// Convert a ToUnicode destination hex string to Unicode.
///
/// PDF ToUnicode destinations are UTF-16BE strings. Supplementary-plane
/// characters are encoded as surrogate pairs, so treating each 4-hex chunk as
/// a scalar drops emoji like D83CDF1F.
fn hex_to_unicode_string(hex: &str) -> Option<String> {
let hex = hex.trim();
let mut result = String::new();
// Process 4 hex digits at a time
let mut i = 0;
while i + 4 <= hex.len() {
if let Ok(cp) = u32::from_str_radix(&hex[i..i + 4], 16) {
if let Some(c) = char::from_u32(cp) {
result.push(c);
}
}
i += 4;
let hex: String = hex.chars().filter(|ch| !ch.is_ascii_whitespace()).collect();
if hex.is_empty() || !hex.len().is_multiple_of(2) {
return None;
}
if result.is_empty() {
None
let bytes: Option<Vec<u8>> = (0..hex.len())
.step_by(2)
.map(|i| u8::from_str_radix(hex.get(i..i + 2)?, 16).ok())
.collect();
let bytes = bytes?;
if bytes.len().is_multiple_of(2) {
let units: Vec<u16> = bytes
.chunks_exact(2)
.map(|chunk| u16::from_be_bytes([chunk[0], chunk[1]]))
.collect();
if let Ok(result) = String::from_utf16(&units) {
if !result.is_empty() {
return Some(normalize_tounicode_destination(result));
}
}
}
// Be permissive for non-standard one-byte destinations.
if bytes.len() == 1 {
let ch = bytes[0] as char;
if !ch.is_control() || ch == '\t' || ch == '\n' {
return Some(ch.to_string());
}
}
None
}
fn normalize_tounicode_destination(text: String) -> String {
let is_multi_char = text.chars().nth(1).is_some();
// Some malformed producer CMaps put a list of alternative whitespace or
// hyphen codepoints into one destination. Keep ordinary multi-character
// mappings intact unless that malformed signature is present.
if is_multi_char
&& text.chars().all(char::is_whitespace)
&& text.chars().any(|ch| matches!(ch, '\t' | '\n' | '\r'))
{
return if text.contains('\t') {
"\t".to_string()
} else {
" ".to_string()
};
}
if is_multi_char
&& text.contains('\u{00ad}')
&& text.chars().all(|ch| {
matches!(
ch,
'-' | '\u{00ad}' | '\u{2010}' | '\u{2011}' | '\u{2012}' | '\u{2013}' | '\u{2212}'
)
})
{
return "-".to_string();
}
text
}
fn hex_to_unicode_scalar(hex: &str) -> Option<u32> {
let text = hex_to_unicode_string(hex)?;
let mut chars = text.chars();
let ch = chars.next()?;
if chars.next().is_none() {
Some(ch as u32)
} else {
Some(result)
None
}
}
@@ -820,6 +880,25 @@ fn try_remap_subset_cmap(
None => return (cmap, None),
};
// Both repair paths below assume CIDs are glyph indices that a subsetter can
// renumber, which is only true for CIDFontType2 (TrueType). For CIDFontType0
// (CFF), CIDs are resolved through the CFF charset, so a valid CMap stays valid
// after subsetting and renumbering it corrupts otherwise-correct text.
// CIDToGIDMap is likewise CIDFontType2-only (PDF 32000-1:2008, 9.7.4.2), so this
// also ignores a CIDToGIDMap that a malformed producer attached to a CFF font.
// /Subtype may be an indirect reference, so resolve it through the document.
// Only bail out when the descendant is *explicitly* something other than
// CIDFontType2: a missing or unresolvable /Subtype keeps the previous
// behaviour rather than silently disabling the repair.
let subtype = cid_font_dict.get(b"Subtype").ok().and_then(|o| match o {
Object::Reference(r) => doc.get_object(*r).ok().and_then(|o| o.as_name().ok()),
other => other.as_name().ok(),
});
if subtype.is_some_and(|name| name != b"CIDFontType2") {
debug!("Subset remap skipped for obj={obj_num}: descendant is not CIDFontType2");
return (cmap, None);
}
// If there's an explicit CIDToGIDMap, build a repaired CMap using it.
if let Some(cid_to_gid) = get_cid_to_gid_map(cid_font_dict, doc) {
if let Some(repaired) = build_cmap_with_cid_to_gid_map(&cmap, &cid_to_gid) {
@@ -1095,9 +1174,7 @@ fn build_gid_to_unicode(face: &ttf_parser::Face<'_>) -> Option<HashMap<u16, char
/// Build a ToUnicodeCMap from pdf.js built-in binary CMaps (bcmaps).
fn build_cmap_from_builtin_cmap(ordering: &str) -> Option<ToUnicodeCMap> {
let name = format!("Adobe-{}-UCS2.bcmap", ordering);
let dir = find_bcmaps_dir()?;
let path = dir.join(name);
let data = std::fs::read(&path).ok()?;
let data = read_builtin_cmap_file(&name)?;
let mut cmap = parse_binary_cmap(&data).ok()?;
if cmap.char_map.is_empty() && cmap.ranges.is_empty() {
return None;
@@ -1105,13 +1182,14 @@ fn build_cmap_from_builtin_cmap(ordering: &str) -> Option<ToUnicodeCMap> {
cmap.code_byte_length = 2;
debug!(
"Built-in CMap {}: char_map={} ranges={}",
path.display(),
name,
cmap.char_map.len(),
cmap.ranges.len()
);
Some(cmap)
}
#[cfg(not(target_arch = "wasm32"))]
fn find_bcmaps_dir() -> Option<PathBuf> {
if let Ok(dir) = std::env::var("PDF_INSPECTOR_BCMAPS_DIR") {
let p = PathBuf::from(dir);
@@ -1128,6 +1206,18 @@ fn find_bcmaps_dir() -> Option<PathBuf> {
None
}
#[cfg(not(target_arch = "wasm32"))]
fn read_builtin_cmap_file(name: &str) -> Option<Cow<'static, [u8]>> {
let path = find_bcmaps_dir()?.join(name);
std::fs::read(path).ok().map(Cow::Owned)
}
#[cfg(target_arch = "wasm32")]
fn read_builtin_cmap_file(name: &str) -> Option<Cow<'static, [u8]>> {
let file = BUILTIN_CMAPS.get_file(name)?;
Some(Cow::Borrowed(file.contents()))
}
fn parse_binary_cmap(data: &[u8]) -> Result<ToUnicodeCMap, String> {
let mut stream = BinaryCMapStream::new(data);
let _header = stream.read_byte().ok_or("unexpected EOF in bcmap header")?;
@@ -1438,9 +1528,7 @@ fn parse_encoding_cmap_object(obj: &Object, doc: &Document) -> Option<EncodingCM
}
fn load_builtin_encoding_cmap(name: &str) -> Option<EncodingCMap> {
let dir = find_bcmaps_dir()?;
let path = dir.join(format!("{}.bcmap", name));
let data = std::fs::read(&path).ok()?;
let data = read_builtin_cmap_file(&format!("{}.bcmap", name))?;
parse_binary_cmap_encoding(&data).ok()
}
@@ -1725,9 +1813,7 @@ fn load_builtin_cmap_by_name(name: &str) -> Option<ToUnicodeCMap> {
if !name.ends_with("UCS2") {
return None;
}
let dir = find_bcmaps_dir()?;
let path = dir.join(format!("{}.bcmap", name));
let data = std::fs::read(&path).ok()?;
let data = read_builtin_cmap_file(&format!("{}.bcmap", name))?;
let mut cmap = parse_binary_cmap(&data).ok()?;
if cmap.char_map.is_empty() && cmap.ranges.is_empty() {
return None;
@@ -2520,6 +2606,24 @@ endcmap
assert_eq!(cmap.lookup(0x0025), Some("B".to_string()));
}
#[test]
fn test_hex_to_unicode_non_ascii_no_panic() {
// A destination containing a multi-byte char makes the byte length even
// while a byte offset can land inside a char. Slicing must not panic;
// it should be rejected gracefully.
assert_eq!(hex_to_unicode_string("XéY"), None);
assert_eq!(hex_to_unicode_string("\u{fffd}0"), None);
}
#[test]
fn test_parse_bfchar_non_ascii_destination_no_panic() {
// Crafted /ToUnicode CMap: a non-hex, non-ASCII destination previously
// triggered a char-boundary panic in hex_to_unicode_string.
let cmap_content = "beginbfchar <0041> <XéY> endbfchar";
// Must not panic; the malformed entry is simply skipped.
let _ = ToUnicodeCMap::parse(cmap_content.as_bytes());
}
#[test]
fn test_parse_bfchar_1byte() {
// This is the pattern that caused the CJK bug: codespace is <0000><FFFF>
@@ -2607,6 +2711,97 @@ endbfrange
assert_eq!(cmap.lookup(0x0005), Some("C".to_string()));
}
#[test]
fn test_parse_bfchar_surrogate_pair_emoji() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
2 beginbfchar
<16> <D83CDF1F>
<9D> <D83CDFAD>
endbfchar
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.code_byte_length, 1);
assert_eq!(cmap.lookup(0x16), Some("🌟".to_string()));
assert_eq!(cmap.lookup(0x9D), Some("🎭".to_string()));
}
#[test]
fn test_parse_bfrange_surrogate_pair_base() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
1 beginbfrange
<C8> <C9> <D83CDFD8>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.code_byte_length, 1);
assert_eq!(cmap.lookup(0xC8), Some("🏘".to_string()));
assert_eq!(cmap.lookup(0xC9), Some("🏙".to_string()));
}
#[test]
fn test_parse_bfrange_preserves_single_hyphen_like_base() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
1 beginbfrange
<21> <22> <2013>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.lookup(0x21), Some("".to_string()));
assert_eq!(cmap.lookup(0x22), Some("".to_string()));
}
#[test]
fn test_parse_spaced_destination_hex_without_control_noise() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
3 beginbfchar
<21> < 0009 000d 0020 00a0 >
<22> < 002d 00ad 2010 >
<23> <00a0>
endbfchar
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.lookup(0x21), Some("\t".to_string()));
assert_eq!(cmap.lookup(0x22), Some("-".to_string()));
assert_eq!(cmap.lookup(0x23), Some("\u{00a0}".to_string()));
}
#[test]
fn test_parse_preserves_valid_multi_character_destinations() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
4 beginbfchar
<21> <002d002d>
<22> <20132013>
<23> <002000a0>
<24> <00660069>
endbfchar
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.lookup(0x21), Some("--".to_string()));
assert_eq!(cmap.lookup(0x22), Some("––".to_string()));
assert_eq!(cmap.lookup(0x23), Some(" \u{00a0}".to_string()));
assert_eq!(cmap.lookup(0x24), Some("fi".to_string()));
}
#[test]
fn test_remap_to_sequential() {
// Simulate a broken CMap where GIDs are from pre-subsetting:
@@ -3001,4 +3196,106 @@ endbfrange
"Remap must fire when CMap's CIDs are outside W array coverage"
);
}
#[test]
fn test_try_remap_skipped_for_cid_font_type0() {
// Same W/CMap mismatch as the CIDFontType2 case above, but the descendant is
// CIDFontType0 (CFF). There CIDs are resolved through the CFF charset, so the
// ToUnicode CIDs stay valid after subsetting and must not be renumbered.
// Real-world case: Japanese Adobe-Japan1 PDFs (e.g. National Diet Library
// minutes) where remapping turned correct text into unrelated glyphs.
let cmap_content = r#"
1 begincodespacerange
<0000><FFFF>
endcodespacerange
1 beginbfrange
<0200> <0220> <0410>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
let mut doc = Document::new();
// CIDToGIDMap is CIDFontType2-only, but a malformed producer can still emit
// one on a CFF font. Use a real stream (not /Identity, which is treated as
// "no map") so this also fails if the guard is moved back below the
// CIDToGIDMap branch: cid 1 -> gid 0x0200, which the CMap resolves.
let mut cid_to_gid = vec![0u8; 68];
cid_to_gid[2] = 0x02;
cid_to_gid[3] = 0x00;
let cid_to_gid_id =
doc.add_object(lopdf::Stream::new(lopdf::Dictionary::new(), cid_to_gid));
let mut cid_font = lopdf::Dictionary::new();
cid_font.set("Subtype", lopdf::Object::Name(b"CIDFontType0".to_vec()));
cid_font.set("CIDToGIDMap", lopdf::Object::Reference(cid_to_gid_id));
cid_font.set(
"W",
lopdf::Object::Array(vec![
lopdf::Object::Integer(1),
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
]),
);
let cid_font_id = doc.add_object(cid_font);
let mut font_dict = lopdf::Dictionary::new();
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
font_dict.set(
"DescendantFonts",
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
);
let (primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 789);
assert!(
remapped.is_none(),
"Remap must be skipped for CIDFontType0 (CFF) descendants, including a \
CIDToGIDMap a malformed producer attached to one"
);
// The original CMap must still resolve its own CIDs.
assert_eq!(primary.lookup(0x0200), Some("\u{0410}".to_string()));
}
#[test]
fn test_try_remap_resolves_indirect_subtype() {
// /Subtype may be stored as an indirect reference. A genuine CIDFontType2
// font must still get the repair, so the guard has to dereference it rather
// than treat the unresolved value as "not CIDFontType2".
let cmap_content = r#"
1 begincodespacerange
<0000><FFFF>
endcodespacerange
1 beginbfrange
<0200> <0220> <0410>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
let mut doc = Document::new();
let subtype_id = doc.add_object(lopdf::Object::Name(b"CIDFontType2".to_vec()));
let mut cid_font = lopdf::Dictionary::new();
cid_font.set("Subtype", lopdf::Object::Reference(subtype_id));
cid_font.set("CIDToGIDMap", lopdf::Object::Name(b"Identity".to_vec()));
cid_font.set(
"W",
lopdf::Object::Array(vec![
lopdf::Object::Integer(0),
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
]),
);
let cid_font_id = doc.add_object(cid_font);
let mut font_dict = lopdf::Dictionary::new();
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
font_dict.set(
"DescendantFonts",
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
);
let (_primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 790);
assert!(
remapped.is_some(),
"An indirect /Subtype naming CIDFontType2 must still reach the remap"
);
}
}
+139 -7
View File
@@ -116,6 +116,14 @@ pub struct TextItem {
pub is_bold: bool,
/// Whether the font is italic
pub is_italic: bool,
/// Whether the text is underlined (drawn rule/thin rect under the
/// baseline — PDFs have no underline font flag, so this is detected
/// geometrically after extraction; see `extractor::underline`).
pub is_underline: bool,
/// Whether the text is struck out (drawn rule/thin rect crossing the
/// glyphs at mid x-height). Same geometric detection as underline,
/// different vertical window; see `extractor::underline`.
pub is_strikeout: bool,
/// Type of item (text, image, link)
pub item_type: ItemType,
/// Marked Content ID from the content stream's BDC/BMC operator.
@@ -137,12 +145,20 @@ pub struct TextLine {
impl TextLine {
pub fn text(&self) -> String {
self.text_with_formatting(false, false)
self.text_with_formatting(false, false, false)
}
/// Get text with optional bold/italic markdown formatting
pub fn text_with_formatting(&self, format_bold: bool, format_italic: bool) -> String {
if !format_bold && !format_italic {
/// Get text with optional bold/italic/decorative markdown formatting.
///
/// `format_decorations` enables both geometrically detected source
/// decorations: underline (`<u>`) and strikeout (`<s>`).
pub fn text_with_formatting(
&self,
format_bold: bool,
format_italic: bool,
format_decorations: bool,
) -> String {
if !format_bold && !format_italic && !format_decorations {
return self.text_plain();
}
@@ -151,6 +167,8 @@ impl TextLine {
let mut result = String::new();
let mut current_bold = false;
let mut current_italic = false;
let mut current_underline = false;
let mut current_strikeout = false;
for (i, item) in self.items.iter().enumerate() {
let text = item.text.as_str();
@@ -176,9 +194,16 @@ impl TextLine {
// we push text_trimmed below (which strips it).
let has_leading_space = text.starts_with(' ');
// Check for style changes
let item_bold = format_bold && item.is_bold;
let item_italic = format_italic && item.is_italic;
// Check for style changes. Source decorations are exclusive:
// `<u>`/`<s>` content stays free of `**`/`*` markers — consumers
// (and the eval harnesses this feeds) match tag content literally,
// and mixed nesting breaks that. A struck-and-underlined item is
// emitted as struck text because deletion is the stronger semantic
// distinction in redline documents.
let item_strikeout = format_decorations && item.is_strikeout;
let item_underline = format_decorations && item.is_underline && !item_strikeout;
let item_bold = format_bold && item.is_bold && !item_underline && !item_strikeout;
let item_italic = format_italic && item.is_italic && !item_underline && !item_strikeout;
// Close previous styles if they change
if current_italic && !item_italic {
@@ -189,6 +214,14 @@ impl TextLine {
result.push_str("**");
current_bold = false;
}
if current_underline && !item_underline {
result.push_str("</u>");
current_underline = false;
}
if current_strikeout && !item_strikeout {
result.push_str("</s>");
current_strikeout = false;
}
// Add space: either from spacing logic or preserved from item text
if needs_space || (has_leading_space && !result.is_empty() && !result.ends_with(' ')) {
@@ -196,6 +229,14 @@ impl TextLine {
}
// Open new styles
if item_underline && !current_underline {
result.push_str("<u>");
current_underline = true;
}
if item_strikeout && !current_strikeout {
result.push_str("<s>");
current_strikeout = true;
}
if item_bold && !current_bold {
result.push_str("**");
current_bold = true;
@@ -215,6 +256,12 @@ impl TextLine {
if current_bold {
result.push_str("**");
}
if current_underline {
result.push_str("</u>");
}
if current_strikeout {
result.push_str("</s>");
}
result
}
@@ -280,3 +327,88 @@ impl TextLine {
|| space_already_exists)
}
}
#[cfg(test)]
mod formatting_tests {
use super::{ItemType, TextItem, TextLine};
fn item(text: &str, x: f32, width: f32, strikeout: bool) -> TextItem {
TextItem {
text: text.to_string(),
x,
y: 100.0,
width,
height: 12.0,
font: "F1".to_string(),
font_size: 12.0,
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: strikeout,
item_type: ItemType::Text,
mcid: None,
}
}
fn line(items: Vec<TextItem>) -> TextLine {
TextLine {
items,
y: 100.0,
page: 1,
adaptive_threshold: 0.1,
}
}
#[test]
fn formatting_emits_semantic_strikeout() {
let line = line(vec![item("deleted", 10.0, 42.0, true)]);
assert_eq!(
line.text_with_formatting(true, true, true),
"<s>deleted</s>"
);
}
#[test]
fn formatting_closes_strikeout_before_live_text() {
let line = line(vec![
item("keep", 10.0, 24.0, false),
item("remove", 40.0, 42.0, true),
item("keep", 88.0, 24.0, false),
]);
assert_eq!(
line.text_with_formatting(true, true, true),
"keep <s>remove</s> keep"
);
}
#[test]
fn formatting_coalesces_adjacent_struck_items() {
let line = line(vec![
item("deleted", 10.0, 42.0, true),
item("words", 58.0, 30.0, true),
]);
assert_eq!(
line.text_with_formatting(true, true, true),
"<s>deleted words</s>"
);
}
#[test]
fn strikeout_takes_precedence_over_other_styles() {
let mut decorated = item("deleted", 10.0, 42.0, true);
decorated.is_bold = true;
decorated.is_italic = true;
decorated.is_underline = true;
let line = line(vec![decorated]);
assert_eq!(
line.text_with_formatting(true, true, true),
"<s>deleted</s>"
);
assert_eq!(line.text(), "deleted");
}
}
+68
View File
@@ -0,0 +1,68 @@
%PDF-1.3
%“Œ‹ž ReportLab Generated PDF document (opensource)
1 0 obj
<<
/F1 2 0 R
>>
endobj
2 0 obj
<<
/BaseFont /Helvetica /Encoding /WinAnsiEncoding /Name /F1 /Subtype /Type1 /Type /Font
>>
endobj
3 0 obj
<<
/Contents 7 0 R /MediaBox [ 0 0 612 792 ] /Parent 6 0 R /Resources <<
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>> /Rotate 0 /Trans <<
>>
/Type /Page
>>
endobj
4 0 obj
<<
/PageMode /UseNone /Pages 6 0 R /Type /Catalog
>>
endobj
5 0 obj
<<
/Author (anonymous) /CreationDate (D:20260803112923+00'00') /Creator (anonymous) /Keywords () /ModDate (D:20260803112923+00'00') /Producer (ReportLab PDF Library - \(opensource\))
/Subject (unspecified) /Title (untitled) /Trapped /False
>>
endobj
6 0 obj
<<
/Count 1 /Kids [ 3 0 R ] /Type /Pages
>>
endobj
7 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ] /Length 202
>>
stream
GarW05mr9@&;9NOME,dW.,B;'jAjYq0S4Z`*D9aMA;]$5J)A/$3lESen1?F)ZJsa4$4&N%-%cs)#qW5EVhhbPiRDrAV>MC%.spto@CU"ZdipR'TtFiMR_%m*Hm$N%qL7a"ckkp9T/s[N2"Og377mP*M^akb2XQZ@'l*qT(9bVtDb5+)S&Q.#%E)<]Ao`TSk2AE'/E\fn~>endstream
endobj
xref
0 8
0000000000 65535 f
0000000061 00000 n
0000000092 00000 n
0000000199 00000 n
0000000392 00000 n
0000000460 00000 n
0000000721 00000 n
0000000780 00000 n
trailer
<<
/ID
[<6d7ea1213c5974c78613d5d2a08423b5><6d7ea1213c5974c78613d5d2a08423b5>]
% ReportLab generated PDF document -- digest (opensource)
/Info 5 0 R
/Root 4 0 R
/Size 8
>>
startxref
9072
%%EOF
Binary file not shown.
File diff suppressed because one or more lines are too long
Binary file not shown.
Binary file not shown.
Binary file not shown.
File diff suppressed because it is too large Load Diff
+552 -4
View File
@@ -8,12 +8,13 @@ use pdf_inspector::{
detect_pdf_type, detect_vector_grid_in_region_mem, extract_pages_markdown,
extract_pages_markdown_mem, extract_tables_in_regions_mem, extract_text,
extract_text_in_regions_mem, extract_text_with_positions, extract_text_with_positions_mem,
process_pdf_mem, process_pdf_with_options, to_markdown, MarkdownOptions, PdfError, PdfOptions,
process_pdf_mem, process_pdf_mem_with_options, process_pdf_with_options, to_markdown,
to_markdown_from_items_with_rects_and_page_count, MarkdownOptions, PdfError, PdfOptions,
PdfType, TextItem,
};
use std::collections::HashSet;
fn make_minimal_text_pdf() -> Vec<u8> {
fn make_text_pdf(content: &str, media_box: &str) -> Vec<u8> {
let mut pdf = b"%PDF-1.4\n".to_vec();
let mut offsets = vec![0usize];
@@ -40,10 +41,11 @@ fn make_minimal_text_pdf() -> Vec<u8> {
&mut pdf,
&mut offsets,
3,
"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] /Resources << /Font << /F1 5 0 R >> >> /Contents 4 0 R >>",
&format!(
"<< /Type /Page /Parent 2 0 R /MediaBox [{media_box}] /Resources << /Font << /F1 5 0 R >> >> /Contents 4 0 R >>"
),
);
let content = "BT /F1 12 Tf 100 700 Td (Hello World) Tj 0 -14 Td (Second Line) Tj 0 -14 Td (Third Line) Tj ET";
add_object(
&mut pdf,
&mut offsets,
@@ -79,6 +81,186 @@ fn make_minimal_text_pdf() -> Vec<u8> {
pdf
}
fn make_recurring_contextual_folio_pdf() -> Vec<u8> {
let mut pdf = b"%PDF-1.4\n".to_vec();
let mut offsets = vec![0usize];
fn add_object(pdf: &mut Vec<u8>, offsets: &mut Vec<usize>, id: usize, body: &str) {
offsets.push(pdf.len());
pdf.extend_from_slice(format!("{id} 0 obj\n").as_bytes());
pdf.extend_from_slice(body.as_bytes());
pdf.extend_from_slice(b"\nendobj\n");
}
add_object(
&mut pdf,
&mut offsets,
1,
"<< /Type /Catalog /Pages 2 0 R >>",
);
add_object(
&mut pdf,
&mut offsets,
2,
"<< /Type /Pages /Kids [3 0 R 5 0 R 7 0 R 9 0 R] /Count 4 >>",
);
for page_index in 0..4 {
let page_id = 3 + page_index * 2;
let content_id = page_id + 1;
add_object(
&mut pdf,
&mut offsets,
page_id,
&format!(
"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] /Resources << /Font << /F1 11 0 R >> >> /Contents {content_id} 0 R >>"
),
);
let page_number = page_index + 1;
let content = format!(
"BT /F1 12 Tf 1 0 0 1 25 30 Tm ({page_number}) Tj 1 0 0 1 41 30 Tm (Company report footer) Tj 1 0 0 1 72 700 Tm (Body page {page_number}) Tj ET"
);
add_object(
&mut pdf,
&mut offsets,
content_id,
&format!(
"<< /Length {} >>\nstream\n{}\nendstream",
content.len(),
content
),
);
}
add_object(
&mut pdf,
&mut offsets,
11,
"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica >>",
);
let xref_start = pdf.len();
pdf.extend_from_slice(format!("xref\n0 {}\n", offsets.len()).as_bytes());
pdf.extend_from_slice(b"0000000000 65535 f \n");
for offset in offsets.iter().skip(1) {
pdf.extend_from_slice(format!("{offset:010} 00000 n \n").as_bytes());
}
pdf.extend_from_slice(
format!(
"trailer\n<< /Size {} /Root 1 0 R >>\nstartxref\n{}\n%%EOF",
offsets.len(),
xref_start
)
.as_bytes(),
);
pdf
}
fn make_pdf_with_malformed_unselected_page() -> Vec<u8> {
let mut pdf = b"%PDF-1.4\n".to_vec();
let mut offsets = vec![0usize];
fn add_object(pdf: &mut Vec<u8>, offsets: &mut Vec<usize>, id: usize, body: &str) {
offsets.push(pdf.len());
pdf.extend_from_slice(format!("{id} 0 obj\n").as_bytes());
pdf.extend_from_slice(body.as_bytes());
pdf.extend_from_slice(b"\nendobj\n");
}
add_object(
&mut pdf,
&mut offsets,
1,
"<< /Type /Catalog /Pages 2 0 R >>",
);
add_object(
&mut pdf,
&mut offsets,
2,
"<< /Type /Pages /Kids [3 0 R 5 0 R] /Count 2 >>",
);
add_object(
&mut pdf,
&mut offsets,
3,
"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] /Resources << /Font << /F1 7 0 R >> >> /Contents 4 0 R >>",
);
let content = "BT /F1 12 Tf 1 0 0 1 25 30 Tm (1) Tj 1 0 0 1 41 30 Tm (Company report footer) Tj 1 0 0 1 72 700 Tm (Selected page text) Tj 0 -16 Td (More selected text) Tj 0 -16 Td (Still selected text) Tj ET";
add_object(
&mut pdf,
&mut offsets,
4,
&format!(
"<< /Length {} >>\nstream\n{}\nendstream",
content.len(),
content
),
);
add_object(
&mut pdf,
&mut offsets,
5,
"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] /Resources << /Font << /F1 7 0 R >> >> /Contents 6 0 R >>",
);
add_object(
&mut pdf,
&mut offsets,
6,
"<< /Length 3 >>\nstream\nBI \nendstream",
);
add_object(
&mut pdf,
&mut offsets,
7,
"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica >>",
);
let xref_start = pdf.len();
pdf.extend_from_slice(format!("xref\n0 {}\n", offsets.len()).as_bytes());
pdf.extend_from_slice(b"0000000000 65535 f \n");
for offset in offsets.iter().skip(1) {
pdf.extend_from_slice(format!("{offset:010} 00000 n \n").as_bytes());
}
pdf.extend_from_slice(
format!(
"trailer\n<< /Size {} /Root 1 0 R >>\nstartxref\n{}\n%%EOF",
offsets.len(),
xref_start
)
.as_bytes(),
);
pdf
}
fn make_minimal_text_pdf() -> Vec<u8> {
make_text_pdf(
"BT /F1 12 Tf 100 700 Td (Hello World) Tj 0 -14 Td (Second Line) Tj 0 -14 Td (Third Line) Tj ET",
"0 0 612 792",
)
}
fn make_digit_run_repro_pdf() -> Vec<u8> {
let content = r#"BT
/F1 12 Tf
1 0 0 1 72 780 Tm (A\)) Tj
1 0 0 1 96 780 Tm (The) Tj
1 0 0 1 126 780 Tm (total) Tj
1 0 0 1 166 780 Tm (of) Tj
1 0 0 1 186 780 Tm (730) Tj
1 0 0 1 220 780 Tm (seats) Tj
1 0 0 1 262 780 Tm (was) Tj
1 0 0 1 296 780 Tm (approved.) Tj
1 0 0 1 72 755 Tm (B\)) Tj
1 0 0 1 96 755 Tm (let) Tj
1 0 0 1 120 755 Tm (log) Tj
1 0 0 1 150 755 Tm (2) Tj
1 0 0 1 164 755 Tm (=) Tj
1 0 0 1 180 755 Tm (a) Tj
1 0 0 1 72 720 Tm (C\) Control: The total of 730 seats was approved. let log 2 = a) Tj
ET"#;
make_text_pdf(content, "0 0 595 842")
}
fn truncate_eof_marker(mut pdf: Vec<u8>) -> Vec<u8> {
assert!(pdf.ends_with(b"%%EOF"));
pdf.pop();
@@ -104,6 +286,8 @@ fn make_text_item(text: &str, x: f32, y: f32, font_size: f32, page: u32) -> Text
page,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -129,6 +313,8 @@ fn make_text_item_with_font(
page,
is_bold: is_bold_font(font),
is_italic: is_italic_font(font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -326,6 +512,21 @@ fn test_group_into_lines_sorting_by_x() {
assert_eq!(lines[0].text(), "First Second Third");
}
#[test]
fn test_digit_only_text_runs_are_preserved_in_markdown() {
let pdf = make_digit_run_repro_pdf();
let items = extract_text_with_positions_mem(&pdf).expect("extract positioned text");
assert!(items.iter().any(|item| item.text == "730"));
assert!(items.iter().any(|item| item.text == "2"));
let result = process_pdf_mem(&pdf).expect("convert PDF to markdown");
assert_eq!(
result.markdown.expect("markdown output").trim(),
"A) The total of 730 seats was approved.\nB) let log 2 = a\nC) Control: The total of 730 seats was approved. let log 2 = a"
);
}
// ============================================================================
// MarkdownOptions Tests
// ============================================================================
@@ -333,6 +534,7 @@ fn test_group_into_lines_sorting_by_x() {
#[test]
fn test_markdown_options_default() {
let opts = MarkdownOptions::default();
assert_eq!(opts.profile, pdf_inspector::MarkdownProfile::Fidelity);
assert!(opts.detect_headers);
assert!(opts.detect_lists);
assert!(opts.detect_code);
@@ -342,6 +544,7 @@ fn test_markdown_options_default() {
#[test]
fn test_markdown_options_custom() {
let opts = MarkdownOptions {
profile: pdf_inspector::MarkdownProfile::Compact,
detect_headers: false,
detect_lists: true,
detect_code: false,
@@ -357,6 +560,7 @@ fn test_markdown_options_custom() {
..Default::default()
};
assert!(!opts.detect_headers);
assert_eq!(opts.profile, pdf_inspector::MarkdownProfile::Compact);
assert!(opts.detect_lists);
assert!(!opts.detect_code);
assert_eq!(opts.base_font_size, Some(14.0));
@@ -540,6 +744,30 @@ fn test_markdown_from_items_page_breaks() {
assert!(md.contains("Content on second page"));
}
#[test]
fn test_markdown_page_count_overload_includes_trailing_blank_pages_in_folio_coverage() {
let mut items = Vec::new();
for (page, value) in [(1, "1"), (2, "2"), (3, "3"), (4, "4")] {
items.push(make_text_item(value, 25.0, 30.0, 12.0, page));
items.push(make_text_item(
"Company report footer",
41.0,
30.0,
12.0,
page,
));
}
let options = MarkdownOptions {
strip_headers_footers: false,
..MarkdownOptions::default()
};
let md = to_markdown_from_items_with_rects_and_page_count(items, options, &[], 20);
assert!(md.contains("1 Company report footer"));
assert!(md.contains("4 Company report footer"));
}
// ============================================================================
// Markdown From Lines Tests
// ============================================================================
@@ -1082,6 +1310,18 @@ fn test_snapshot_2013_app2() {
assert_snapshot("2013-app2");
}
/// First two pages of Shannon's "A Mathematical Theory of Communication"
/// (1998 dvips 5.58 → Distiller 3 retypesetting). Canonical legacy-TeX PDF:
/// non-embedded base-14 fonts with no /Widths (exercises the built-in AFM
/// metrics fallback), Type3 PK bitmap math fonts with FontMatrix
/// [1 0 0 -1 0 0] (exercises visual-size scaling), a two-line embedded drop
/// cap, indent-only paragraph breaks, and display math that must not be
/// detected as tables or headings.
#[test]
fn test_snapshot_shannon_entropy() {
assert_snapshot("shannon-entropy-p1-2");
}
// ============================================================================
// Pages Needing OCR Tests
// ============================================================================
@@ -1098,6 +1338,7 @@ fn test_pages_needing_ocr_field_accessible() {
title: None,
ocr_recommended: false,
pages_needing_ocr: Vec::new(),
ocr_reasons_by_page: std::collections::BTreeMap::new(),
};
assert!(detection_result.pages_needing_ocr.is_empty());
@@ -1107,6 +1348,7 @@ fn test_pages_needing_ocr_field_accessible() {
page_count: 1,
processing_time_ms: 0,
pages_needing_ocr: vec![1, 3],
ocr_reasons_by_page: Vec::new(),
title: None,
confidence: 1.0,
layout: pdf_inspector::LayoutComplexity::default(),
@@ -1187,6 +1429,77 @@ fn test_firecrawl_tagged_pdf_struct_tree() {
assert_eq!(fence_count % 2, 0, "Code fences should be balanced");
}
#[test]
fn test_tagged_pdf_text_items_carry_mcid() {
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
assert!(
items.iter().any(|i| i.mcid.is_some()),
"Tagged PDF text items should carry Marked Content IDs"
);
}
#[test]
fn test_extract_structure_elements_tagged_pdf() {
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
assert!(!elements.is_empty(), "Tagged PDF should yield elements");
assert!(
elements.iter().any(|e| e.role == "H1"),
"Should surface H1 heading roles"
);
assert!(
elements.iter().all(|e| !e.role.is_empty()),
"Every element should carry a role name"
);
// Sorted by (page, mcid) for deterministic output
assert!(
elements
.windows(2)
.all(|w| (w[0].page, w[0].mcid) <= (w[1].page, w[1].mcid)),
"Elements should be sorted by (page, mcid)"
);
// The advertised join: (page, mcid) pairs must line up with the
// mcid-carrying TextItems from positioned extraction, and joining the
// H1 entries must recover non-empty heading text.
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
let h1_refs: std::collections::HashSet<(u32, i64)> = elements
.iter()
.filter(|e| e.role == "H1")
.map(|e| (e.page, e.mcid))
.collect();
let h1_text: String = items
.iter()
.filter(|i| i.mcid.is_some_and(|mcid| h1_refs.contains(&(i.page, mcid))))
.map(|i| i.text.as_str())
.collect();
assert!(
!h1_text.trim().is_empty(),
"Joining H1 structure elements to text items should recover heading text"
);
// Page filter is 1-indexed (matching TextItem.page) and equals the
// corresponding subset of the full document result.
let page1 = pdf_inspector::extract_structure_elements_mem(&buf, Some(&[1])).unwrap();
assert!(!page1.is_empty(), "Page 1 should have elements");
assert!(page1.iter().all(|e| e.page == 1));
let full_page1_count = elements.iter().filter(|e| e.page == 1).count();
assert_eq!(page1.len(), full_page1_count);
}
#[test]
fn test_extract_structure_elements_untagged_pdf_empty() {
let buf = std::fs::read("tests/fixtures/thermo-freon12.pdf").unwrap();
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
assert!(
elements.is_empty(),
"Untagged PDF should yield no structure elements, got {:?}",
elements
);
}
#[test]
fn test_identity_h_no_tounicode_suppresses_garbage() {
// shinagawa_identity_h.pdf uses YuGothic with Identity-H encoding and no
@@ -1353,6 +1666,32 @@ fn test_extract_regions_mem_identity_h_needs_ocr() {
);
}
/// ParseBench `text_simple__att10k.pdf` (issue #118): the producer authored a
/// broken ToUnicode CMap that shifts every character by a per-range constant,
/// and the embedded subset font has no `cmap` table to recover from. The
/// resulting ciphertext is 100% printable ASCII, so it must be caught by the
/// substitution-cipher statistics and routed to OCR instead of served silently.
#[test]
fn test_extract_pages_mem_shifted_cipher_tounicode_needs_ocr() {
let buf = std::fs::read("tests/fixtures/shifted_cipher_tounicode.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, None).unwrap();
assert_eq!(result.pages.len(), 1);
assert!(
result.pages[0].needs_ocr,
"shifted-cipher garbled page should be flagged needs_ocr"
);
assert!(
result.pages[0].markdown.is_empty(),
"garbled markdown should be suppressed"
);
assert_eq!(result.pages_needing_ocr, vec![1]);
assert_eq!(
result.pages[0].ocr_reason.as_deref(),
Some("suspected_garbled_text")
);
}
#[test]
fn test_extract_regions_mem_multiple_regions_per_page() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
@@ -2880,6 +3219,57 @@ fn test_extract_pages_markdown_basic() {
assert!(!result.pages[0].needs_ocr);
}
#[test]
fn test_extract_pages_markdown_uses_document_wide_folio_context() {
let pdf = make_recurring_contextual_folio_pdf();
let result = extract_pages_markdown_mem(&pdf, None).unwrap();
assert_eq!(result.pages.len(), 4);
for (index, page) in result.pages.iter().enumerate() {
assert!(page.markdown.contains("Company report footer"));
assert!(
!page
.markdown
.contains(&format!("{} Company report footer", index + 1)),
"recurring contextual folio survived on page {}: {}",
index + 1,
page.markdown
);
}
}
#[test]
fn test_process_pdf_page_filter_uses_document_wide_folio_context() {
let pdf = make_recurring_contextual_folio_pdf();
let result = process_pdf_mem_with_options(&pdf, PdfOptions::new().pages([1])).unwrap();
let markdown = result.markdown.unwrap();
assert!(markdown.contains("Company report footer"));
assert!(!markdown.contains("1 Company report footer"), "{markdown}");
assert!(markdown.contains("Body page 1"));
assert!(!markdown.contains("Body page 2"));
}
#[test]
fn test_selected_page_ignores_context_only_extraction_failure() {
let pdf = make_pdf_with_malformed_unselected_page();
let pages = extract_pages_markdown_mem(&pdf, Some(&[0])).unwrap();
assert_eq!(pages.pages.len(), 1);
assert!(pages.pages[0].markdown.contains("Selected page text"));
let result = process_pdf_mem_with_options(&pdf, PdfOptions::new().pages([1])).unwrap();
let markdown = result.markdown.unwrap();
assert!(markdown.contains("Selected page text"));
}
#[test]
fn test_requested_page_extraction_failure_remains_fatal() {
let pdf = make_pdf_with_malformed_unselected_page();
assert!(extract_pages_markdown_mem(&pdf, Some(&[1])).is_err());
}
#[test]
fn test_extract_pages_markdown_page_ordering() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
@@ -3575,3 +3965,161 @@ fn test_markdown_options_default_has_include_images_false() {
let opts = MarkdownOptions::default();
assert!(!opts.include_images);
}
#[test]
fn encrypted_pdf_decrypts_with_correct_password() {
let path = "tests/fixtures/encrypted-secret123.pdf";
// No password: the file is encrypted and can't be read.
let no_pw = process_pdf_with_options(path, PdfOptions::new());
assert!(
matches!(no_pw, Err(PdfError::Encrypted)),
"expected Encrypted without a password, got {no_pw:?}"
);
// Wrong password: still rejected.
let wrong = process_pdf_with_options(path, PdfOptions::new().password("wrong"));
assert!(
matches!(wrong, Err(PdfError::Encrypted)),
"expected Encrypted with a wrong password, got {wrong:?}"
);
// Correct password: decrypts and extracts real content.
let ok = process_pdf_with_options(path, PdfOptions::new().password("secret123"))
.expect("correct password should decrypt");
let md = ok.markdown.unwrap_or_default();
// Assert a stable fixture token so a garbled-but-long extraction (the
// encrypted-stream regression this guards) still fails the test.
assert!(
md.contains("Procurement"),
"decrypted markdown should contain the fixture's real text, got {} chars",
md.len()
);
}
/// Regression for the #231 review finding: `extract_pages_markdown`'s
/// `has_template_image` check must be gated the same way
/// `classify_pdf`/`detect_pdf_type` gates it (image_count <= 1, few text
/// ops, low alphanumeric diversity) — not treated as sufficient on its
/// own. The fixture is a real text page with substantial, richly varied
/// body text (>=50 Tj ops) drawn over a full-bleed background image
/// (e.g. letterhead/watermark). Before the fix, has_template_image alone
/// forced needs_ocr=true and discarded the page's clean markdown; now the
/// page must extract normally.
#[test]
fn test_extract_pages_markdown_does_not_ocr_text_page_with_watermark_image() {
let buf = std::fs::read("tests/fixtures/text_page_with_watermark_image.pdf").unwrap();
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
let page = &ext.pages[0];
assert!(
!page.needs_ocr,
"a text page with substantial real text should not be routed to OCR \
just because it has a background image"
);
assert!(
page.markdown.contains("watermark"),
"expected the page's real body text to be preserved, got: {:?}",
page.markdown
);
}
/// Regression for the #231 review finding: `extract_pages_markdown` never
/// checked `has_vector_text` at all, even though `detect_from_document`'s
/// Mixed-type per-page routing always sends vector-outlined-text pages to
/// OCR (outlined glyphs can't be extracted as text). A page with massive
/// path ops (outlined decorative text) plus a short genuine caption would
/// extract that caption cleanly — non-empty, non-garbled — so the
/// existing empty/garbage-text checks alone couldn't catch it.
#[test]
fn test_extract_pages_markdown_ocrs_page_with_vector_outlined_text() {
let buf = std::fs::read("tests/fixtures/vector_outlined_text_with_caption.pdf").unwrap();
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
assert!(
cls.pages_needing_ocr.contains(&1),
"classify_pdf should flag page 1 as needing OCR (vector-outlined text), got: {:?}",
cls.pages_needing_ocr
);
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
let page = &ext.pages[0];
assert!(
page.needs_ocr,
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
);
assert!(
page.markdown.is_empty(),
"a page flagged needs_ocr must not return markdown as if extraction were \
trustworthy, got: {:?}",
page.markdown
);
}
#[test]
fn pdf_options_debug_redacts_password() {
let opts = PdfOptions::new().password("secret123");
let dbg = format!("{opts:?}");
assert!(
!dbg.contains("secret123"),
"password leaked in Debug: {dbg}"
);
assert!(dbg.contains("REDACTED"), "expected redaction marker: {dbg}");
}
/// Regression for #228: a `startxref` pointer corrupted to point at the
/// wrong byte offset (a single flipped digit — a real, common writer bug)
/// must not make the whole file unprocessable. The real classic xref table
/// is still present and findable by scanning for the `xref` keyword; both
/// pypdf and pdfium recover the same way. Before this fix, every entry
/// point raised "Invalid PDF structure" on a file whose object data was
/// otherwise completely intact.
#[test]
fn test_process_pdf_recovers_corrupted_startxref_pointer() {
let result = process_pdf_with_options(
"tests/fixtures/broken_startxref_pointer.pdf",
PdfOptions::new(),
)
.expect("a corrupted startxref pointer should be recoverable, like pypdf/pdfium");
assert_eq!(result.page_count, 1);
let md = result.markdown.unwrap_or_default();
assert!(
md.contains("Order Detail Report by Account") && md.contains("WIDGET ASSEMBLY"),
"recovered document should extract its real text, got: {md:?}"
);
}
/// Regression for #227: `extract_pages_markdown`'s per-page `needs_ocr`
/// must agree with `classify_pdf`/`detect_pdf_type` on the same page. The
/// fixture is a full-page raster "scan" with a single line of genuine
/// native text drawn over it (a header) — the native text extracts
/// perfectly cleanly (no decoding issues, non-empty), so a needs_ocr
/// computation based on text-quality signals alone says `false`, while
/// detection correctly sees a dominant background image and says the page
/// needs OCR. Both must now agree it needs OCR, and the markdown must not
/// be returned as if the extraction were trustworthy.
#[test]
fn test_extract_pages_markdown_agrees_with_classify_on_scan_with_native_header() {
let buf = std::fs::read("tests/fixtures/scan_with_native_header_text.pdf").unwrap();
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
assert!(
cls.pages_needing_ocr.contains(&1),
"classify_pdf should flag page 1 as needing OCR (image-dominated), got: {:?}",
cls.pages_needing_ocr
);
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
let page = &ext.pages[0];
assert!(
page.needs_ocr,
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
);
assert!(
page.markdown.is_empty(),
"a page flagged needs_ocr must not return markdown as if extraction were \
trustworthy, got: {:?}",
page.markdown
);
}
+2 -2
View File
@@ -188,7 +188,7 @@
|156|23/7|General renovation works to Block B at Belonie Secondary School|MOE|Belvedere Builders|SR869,505.75|
|157|30/7|Procurement of Engine Block and Crankshaft for Engine A11|PUC|Ras Tek Pvt Ltd|Euro798,650.00|
|158|30/7|procurement of Wartsila Engine spares|PUC|Wartsila Eastern Africa ltd|Euro158,424.00|
|159|30/7|Proposed walkway, Drain, rock armoring , road and Bridge widening at Anse Talbot( Ex-Golden Egg)|SLTA|G&S Enterpise|SR1,113,010.00|
|159|30/7|Proposed walkway, Drain, rock armoring, road and Bridge widening at Anse Talbot( Ex-Golden Egg)|SLTA|G&S Enterpise|SR1,113,010.00|
|160|30/7|Procurement of transfer pump control panel|PUC|CA Engineering Consultancy Pte Ltd|SGD14,600.00|
|161|30/7|Consultancy service for North to South Victoria Bye- Pass road and utilities organisation|MLUH|Sonnel Seychelles LTD|SR1,332,000.00|
|162 AUG|30/7|Procurement of the supply of sodium cardonate|PUC|HPL Chemical LTD|USD42,600.00|
@@ -237,7 +237,7 @@
|201|24/9|Procurement of vehicle x 2|SLTA|Abhaye Valabhji Pty Ltd|SR1000.000.00|
||OCT|||||
|202|1/10|Proposed new traffic lane to 5th June Avenue|SLTA|Divy Constrution|SR2,864,589.00|
|203|1/10||Proposed Walkway, Drain, rock armoring , road and Bridge widening at Anse Talbot( Ex-Golden Egg) - Variations SLTA|G & S Enterprise|SR200,448.00|
|203|1/10||Proposed Walkway, Drain, rock armoring, road and Bridge widening at Anse Talbot(Ex-Golden Egg) - Variations SLTA|G & S Enterprise|SR200,448.00|
|204|1/10|Proposed Reconstrcution of Burnt House-Au Cap|MLUH|Furui Construction|SR946,130.00|
|205|1/10|Variation on the project associated with the procurement of seven 100m3/day containerised plant|PUC|Tornado Group (UAE)|USD172,500.00|
|206|1/10|Works on the breaker system at Bel Omber desalination plant|PUC|United Concrete Products (Sey)Ltd|SR1,998,993.11|
+18 -17
View File
@@ -10,7 +10,7 @@ Department of the Treasury **Internal Revenue Service**
### This publication contains:
**Form 4070A, Employees Daily Record of** Tips **Form 4070, Employees Report of Tips to** Employer
**Form 4070A,** Employees Daily Record of Tips **Form 4070,** Employees Report of Tips to Employer
For the period
@@ -22,7 +22,9 @@ Name and address of employee
**Publication 1244 (Rev. 7-96)** Cat. No. 44472W
**Instructions** You must keep sufficient proof to show the amount of your tip income for the year. A daily record of your tip income is considered sufficient proof. Keep a daily record for each workday showing the amount of cash and credit card tips received directly from customers or other employees. Also keep a record of the amount of tips, if any, you paid to other employees through tip sharing, tip pooling or other arrangements, and the names of employees to whom you paid tips. Show the date that each entry is made. This date should be on or near the date you received the tip income. You may use Form 4070A, Employees Daily Record of Tips, or any other daily record to record your tips. **Reporting Tips to Your Employer.—If you** receive tips that total $20 or more for any month while working for one employer, you must report the tips to your employer. Tips include cash left by customers, tips customers add to credit card charges, and tips you receive from other employees. You must report your tips for any one month by the 10th day of the next month. If the 10th day falls on a Saturday, Sunday, or legal holiday, you may give the report to your employer on the next business day that is not a Saturday, Sunday, or legal holiday. You must report tips that total $20 or more every month regardless of your total wages and tips for the year. You may use Form 4070, Employees Report of Tips to Employer, to report your tips to your employer. See the instructions on the back of Form 4070. You must include all tips, including tips not reported to your employer, as wages on your income tax return. You may use the last page of this publication to total your tips for the year. Your employer must withhold income, social security, and Medicare (or railroad retirement) taxes on tips you report. Your employer usually deducts the withholding due on tips from your regular wages.
### Instructions
You must keep sufficient proof to show the amount of your tip income for the year. A daily record of your tip income is considered sufficient proof. Keep a daily record for each workday showing the amount of cash and credit card tips received directly from customers or other employees. Also keep a record of the amount of tips, if any, you paid to other employees through tip sharing, tip pooling or other arrangements, and the names of employees to whom you paid tips. Show the date that each entry is made. This date should be on or near the date you received the tip income. You may use **Form 4070A**, Employees Daily Record of Tips, or any other daily record to record your tips. **Reporting Tips to Your Employer.—**If you receive tips that total $20 or more for any month while working for one employer, you must report the tips to your employer. Tips include cash left by customers, tips customers add to credit card charges, and tips you receive from other employees. You must report your tips for any one month by the 10th day of the next month. If the 10th day falls on a Saturday, Sunday, or legal holiday, you may give the report to your employer on the next business day that is not a Saturday, Sunday, or legal holiday. You must report tips that total $20 or more every month regardless of your total wages and tips for the year. You may use **Form 4070**, Employees Report of Tips to Employer, to report your tips to your employer. See the instructions on the back of Form 4070. You must include all tips, including tips not reported to your employer, as wages on your income tax return. You may use the last page of this publication to total your tips for the year. Your employer must withhold income, social security, and Medicare (or railroad retirement) taxes on tips you report. Your employer usually deducts the withholding due on tips from your regular wages.
*(continued on inside of back cover)*
@@ -30,14 +32,14 @@ Form **4070A** Employees Daily Record of Tips (Rev. July 1996) **This is a vo
Establishment name (if different)
Date Date **a. Tips received**
Date Date **a.** Tips received
**b. Credit card tips c. Tips paid out to d. Names of employees to whom you**
**b.** Credit card tips **c.** Tips paid out to **d.** Names of employees to whom you
tips of directly from customers received other employees paid tips recd. entry and other employees 1 2 3 4 5 **Subtotals** **For Paperwork Reduction Act Notice, see Instructions on the back of Form 4070. Page 1**
Date Date **a. Tips received**
Date Date **a.** Tips received
**b. Credit card tips c. Tips paid out to d. Names of employees to whom you**
**b.** Credit card tips **c.** Tips paid out to **d.** Names of employees to whom you
tips of directly from customers received other employees paid tips recd. entry and other employees
7 8 9 10 11 12 13 14 15 **Subtotals**
@@ -48,13 +50,13 @@ tips of directly from customers received other employees paid tips recd. entr
**Page 3**
27 28 29 30 31 **Subtotals** **from pages** **1, 2, and 3** **Totals**
27 28 29 30 31 **Subtotals from pages** **1, 2, and 3** **Totals**
**1.** Report total cash tips (col. a) on Form 4070, line 1.
**2.** Report total credit card tips (col. b) on Form 4070, line 2.
**3.** Report total tips paid out (col. c) on Form 4070, line 3. **Page 4**
**1.** Report total cash tips (col. **a**) on Form 4070, line **1.**
**2.** Report total credit card tips (col. **b**) on Form 4070, line **2.**
**3.** Report total tips paid out (col. **c**) on Form 4070, line **3.** **Page 4**
Form Employees Report (Rev. July 1996)
Form **4070** Employees Report (Rev. July 1996)
## of Tips to EmployerOMB No. 1545-0065
@@ -66,17 +68,16 @@ Employers name and address (include establishment name, if different) **1** C
**3** Tips paid out
Month or shorter period in which tips were received **4** Net tips (lines 1 + 2 - 3) from, 19, to, 19 Signature Date
Month or shorter period in which tips were received **4** Net tips (lines **1 + 2 - 3**) from, 19, to, 19 Signature Date
**Paperwork Reduction Act Notice.—We ask for the** information on these forms to carry out the Internal Revenue laws of the United States. You are required to give us the information. We need it to ensure that you are complying with these laws and to allow us to figure and collect the right amount of tax. You are not required to provide the information requested on a form that is subject to the Paperwork Reduction Act unless the form displays a valid OMB control number. Books or records relating to a form or its instructions must be retained as long as their contents may become material in the administration of any Internal Revenue law. Generally, tax returns and return information are confidential, as required by Code section 6103. The time needed to complete Forms 4070 and 4070A will vary depending on individual circumstances. The estimated average times are: Recordkeeping—Form 4070, 7 min.; Form 4070A, 3 hr. and 23 min.; Learning **about the law—each form, 2 min.; Preparing Form 4070,** 13 min.; Form 4070A, 55 min.; and Copying and **providing Form 4070, 10 min.; Form 4070A, 14 min.** If you have comments concerning the accuracy of these time estimates or suggestions for making these
**Paperwork Reduction Act Notice.—**We ask for the information on these forms to carry out the Internal Revenue laws of the United States. You are required to give us the information. We need it to ensure that you are complying with these laws and to allow us to figure and collect the right amount of tax. You are not required to provide the information requested on a form that is subject to the Paperwork Reduction Act unless the form displays a valid OMB control number. Books or records relating to a form or its instructions must be retained as long as their contents may become material in the administration of any Internal Revenue law. Generally, tax returns and return information are confidential, as required by Code section 6103. The time needed to complete Forms 4070 and 4070A will vary depending on individual circumstances. The estimated average times are: **Recordkeeping**—Form 4070, 7 min.; Form 4070A, 3 hr. and 23 min.; **Learning** **about the law**—each form, 2 min.; **Preparing** Form 4070, 13 min.; Form 4070A, 55 min.; and **Copying and** **providing** Form 4070, 10 min.; Form 4070A, 14 min. If you have comments concerning the accuracy of these time estimates or suggestions for making these
forms simpler, we would be happy to hear from you. You can write to the Tax Forms Committee, Western Area Distribution Center, Rancho Cordova, CA 95743-0001. **Purpose.—Use this form to report tips you receive to** your employer. This includes cash tips, tips you receive from other employees, and credit card tips. You must report tips every month regardless of your total wages and tips for the year. However, you do not have to report tips to your employer for any month you received less than $20 in tips while working for that employer. Report tips by the 10th day of the month following the month that you receive them. If the 10th day is a Saturday, Sunday, or legal holiday, report tips by the next day that is not a Saturday, Sunday, or legal holiday. See Pub. 531, Reporting Tip Income, for more information. You can get additional copies of Pub. 1244, Employees Daily Record of Tips and Report to Employer, which contains both Forms 4070A and 4070, by calling 1-800-TAX-FORM (1-800-829-3676).
forms simpler, we would be happy to hear from you. You can write to the Tax Forms Committee, Western Area Distribution Center, Rancho Cordova, CA 95743-0001. **Purpose.—**Use this form to report tips you receive to your employer. This includes cash tips, tips you receive from other employees, and credit card tips. You must report tips every month regardless of your total wages and tips for the year. However, you do not have to report tips to your employer for any month you received less than $20 in tips while working for that employer. Report tips by the 10th day of the month following the month that you receive them. If the 10th day is a Saturday, Sunday, or legal holiday, report tips by the next day that is not a Saturday, Sunday, or legal holiday. See **Pub. 531**, Reporting Tip Income, for more information. You can get additional copies of **Pub. 1244**, Employees Daily Record of Tips and Report to Employer, which contains both Forms 4070A and 4070, by calling 1-800-TAX-FORM (1-800-829-3676).
**Instructions (continued)**
<u>Instructions (continued)</u>
**Unreported Tips.—If you received tips of $20 or** more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you must use Form 1040 and Form 4137, Social Security and Medicare Tax on Unreported Tip Income, to report them. You may not use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act cannot use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—Get Pub. 531, Reporting** Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—If you do not keep a daily** record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
**Unreported Tips.—**If you received tips of $20 or more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you **must** use Form 1040 and **Form 4137,** Social Security and Medicare Tax on Unreported Tip Income, to report them. You may **not** use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act **cannot** use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—**Get **Pub. 531,** Reporting Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—**If you do not keep a daily record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
### Instructions (continued)
Use this space to total your tips for the year
+5 -3
View File
@@ -6,7 +6,7 @@
8 4 Z E L L / L U R I E R E A L E S T A T E C E N T E R
**Table I: Cap rate correlations** **Cap Rate Correlation With:*** **BBB Corp** **10-Year Bond Yield S&P Dividend** **Treasury (10-15 yr) Yield** Multifamily 0.187 0.771 0.068 Industrial-0.221 0.748-0.307 CBD Office-0.449 0.694-0.458 Retail-0.181 0.649-02.58
**Table I:** Cap rate correlations **Cap Rate Correlation With:*** **BBB Corp** **10-Year Bond Yield S&P Dividend** **Treasury (10-15 yr) Yield** Multifamily 0.187 0.771 0.068 Industrial-0.221 0.748-0.307 CBD Office-0.449 0.694-0.458 Retail-0.181 0.649-02.58
* Based on 25 years of data for the 10-yrT & S&P DivYld; and 14 years for BBB.
**Figure 1:** NCREIF cap rates vs. 10-yearTreasury
@@ -20,7 +20,9 @@ R E V I E W 8 5
**Figure 2:** Capratespreadsover10-yearTreasury
**Basis Points -200** -400
**Basis Points** -200
-400
-600
@@ -32,7 +34,7 @@ R E V I E W 8 5
1982 1986 1990 1994 1998 2002 2006
**Table II: Correlationsofspreadsbypropertytype** **Correlation of Cap Rate Spreads Over Treasury** **Multifamily Industrial CBD Office**
**Table II:** Correlationsofspreadsbypropertytype **Correlation of Cap Rate Spreads Over Treasury** **Multifamily Industrial CBD Office**
||Multifamily|Industrial|CBD Office|
|---|---|---|---|
+43
View File
@@ -0,0 +1,43 @@
Reprinted with corrections from *The Bell System Technical Journal,* Vol. 27, pp. 379423, 623656, July, October, 1948.
## A Mathematical Theory of Communication
### By C. E. SHANNON
INTRODUCTION
HE recent development of various methods of modulation such as PCM and PPM which exchange
# Tbandwidth for signal-to-noise ratio has intensified the interest in a general theory of communication. A
basis for such a theory is contained in the important papers of Nyquist¹ and Hartley² on this subject. In the present paper we will extend the theory to include a number of new factors, in particular the effect of noise in the channel, and the savings possible due to the statistical structure of the original message and due to the nature of the final destination of the information. The fundamental problem of communication is that of reproducing at one point either exactly or ap- proximately a message selected at another point. Frequently the messages have *meaning*; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one *selected from a set* of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design. If the number of messages in the set is finite then this number or any monotonic function of this number can be regarded as a measure of the information produced when one message is chosen from the set, all choices being equally likely. As was pointed out by Hartley the most natural choice is the logarithmic function. Although this definition must be generalized considerably when we consider the influence of the statistics of the message and when we have a continuous range of messages, we will in all cases use an essentially logarithmic measure. The logarithmic measure is more convenient for various reasons:
1. It is practically more useful. Parameters of engineering importance such as time, bandwidth, number of relays, etc., tend to vary linearly with the logarithm of the number of possibilities. For example, adding one relay to a group doubles the number of possible states of the relays. It adds 1 to the base 2 logarithm of this number. Doubling the time roughly squares the number of possible messages, or doubles the logarithm, etc.
2. It is nearer to our intuitive feeling as to the proper measure. This is closely related to (1) since we in- tuitively measures entities by linear comparison with common standards. One feels, for example, that two punched cards should have twice the capacity of one for information storage, and two identical channels twice the capacity of one for transmitting information.
3. It is mathematically more suitable. Many of the limiting operations are simple in terms of the loga- rithm but would require clumsy restatement in terms of the number of possibilities. The choice of a logarithmic base corresponds to the choice of a unit for measuring information. If the
base 2 is used the resulting units may be called binary digits, or more briefly *bits,* a word suggested by
J. W. Tukey. A device with two stable positions, such as a relay or a flip-flop circuit, can store one bit of information. *N* such devices can store*N* bits, since the total number of possible states is 2
*N* and log₂2 *N* = *N*. If the base 10 is used the units may be called decimal digits. Since
log₂*M* = log₁₀*M*= log₁₀2 = 3:32 log₁₀*M*;
1 Nyquist, H., “Certain Factors Affecting Telegraph Speed,” *Bell System Technical Journal,* April 1924, p. 324; “Certain Topics in Telegraph Transmission Theory,” *A.I.E.E. Trans.,* v. 47, April 1928, p. 617. 2 Hartley, R. V. L., “Transmission of Information,” *Bell System Technical Journal,* July 1928, p. 535.
INFORMATION SOURCE TRANSMITTER RECEIVER DESTINATION
SIGNAL RECEIVED SIGNAL MESSAGE MESSAGE
NOISE SOURCE
Fig. 1 — Schematic diagram of a general communication system.
a decimal digit is about 3 13 bits. A digit wheel on a desk computing machine has ten stable positions and therefore has a storage capacity of one decimal digit. In analytical work where integration and differentiation are involved the base *e* is sometimes useful. The resulting units of information will be called natural units. Change from the base *a* to base *b* merely requires multiplication by log*ba*. By a communication system we will mean a system of the type indicated schematically in Fig. 1. It consists of essentially five parts:
1. An *information source* which produces a message or sequence of messages to be communicated to the receiving terminal. The message may be of various types: (a) A sequence of letters as in a telegraph of teletype system; (b) A single function of time *f* (*t*) as in radio or telephony; (c) A function of time and other variables as in black and white television — here the message may be thought of as a function *f* (*x*; *y*;*t*) of two space coordinates and time, the light intensity at point (*x*; *y*) and time *t* on a pickup tube plate; (d) Two or more functions of time, say *f* (*t*), *g*(*t*), *h*(*t*) — this is the case in “three- dimensional” sound transmission or if the system is intended to service several individual channels in multiplex; (e) Several functions of several variables — in color television the message consists of three functions *f* (*x*; *y*;*t*), *g*(*x*; *y*;*t*), *h*(*x*; *y*;*t*) defined in a three-dimensional continuum — we may also think of these three functions as components of a vector field defined in the region — similarly, several black and white television sources would produce “messages” consisting of a number of functions of three variables; (f) Various combinations also occur, for example in television with an associated audio channel.
2. A *transmitter* which operates on the message in some way to produce a signal suitable for trans- mission over the channel. In telephony this operation consists merely of changing sound pressure into a proportional electrical current. In telegraphy we have an encoding operation which produces a sequence of dots, dashes and spaces on the channel corresponding to the message. In a multiplex PCM system the different speech functions must be sampled, compressed, quantized and encoded, and finally interleaved properly to construct the signal. Vocoder systems, television and frequency modulation are other examples of complex operations applied to the message to obtain the signal.
3. The *channel* is merely the medium used to transmit the signal from transmitter to receiver. It may be a pair of wires, a coaxial cable, a band of radio frequencies, a beam of light, etc.
4. The *receiver* ordinarily performs the inverse operation of that done by the transmitter, reconstructing the message from the signal.
5. The *destination* is the person (or thing) for whom the message is intended. We wish to consider certain general problems involving communication systems. To do this it is first
necessary to represent the various elements involved as mathematical entities, suitably idealized from their
+19 -28
View File
@@ -1,17 +1,8 @@
(e) [Reserved]. For further guidance, see §1.1563-3T(e)(1). Par. 50. Section 1.1563-3T is added to read as follows:
§1.1563-3T Rules for determining stock ownership (temporary).
(a) through (d)(2)(iii) [Reserved]. For further guidance, see §1.1563-3(a)
through (d)(2)(iii). (iv) Statement. If the application of paragraph (d)(2)(ii) or (iii) of §1.1563-3 does not result in a corporation being treated as a component member of only one controlled group of corporations on a December 31, then such corporation will be treated as a component member of only one such group on such date. Such corporation may elect the group in which it is to be included by including on or with its income tax return a statement entitled, “STATEMENT TO ELECT CONTROLLED GROUP PURSUANT TO §1.1563-3T(d)(2)(iv).” The statement must include--
(A) A description of each of the controlled groups in which the corporation
could be included. The description must include the name and employer identification number of each component member of each such group and the stock ownership of the component members of each such group; and
(B) The following representation: [INSERT NAME AND EMPLOYER
IDENTIFICATION NUMBER OF CORPORATION] ELECTS TO BE TREATED AS A COMPONENT MEMBER OF THE [INSERT DESIGNATION OF GROUP].
(v) Election-- (A) Election filed. An election filed under paragraph (d)(2)(iv) of
this section is irrevocable and effective until paragraph (d)(2)(ii) or (iii) of §1.1563-3 applies or until a change in the stock ownership of the corporation results in
||||(e) [Reserved]. For further guidance, see §1.1563-3T(e)(1). Par. 50. Section 1.1563-3T is added to read as follows: §1.1563-3T Rules for determining stock ownership (temporary). (a) through (d)(2)(iii) [Reserved]. For further guidance, see §1.1563-3(a)|
|---|---|---|---|
||through (d)(2)(iii).|||
||(iv)|Statement|. If the application of paragraph (d)(2)(ii) or (iii) of §1.1563-3 does not result in a corporation being treated as a component member of only one controlled group of corporations on a December 31, then such corporation will be treated as a component member of only one such group on such date. Such corporation may elect the group in which it is to be included by including on or with its income tax return a statement entitled, “STATEMENT TO ELECT CONTROLLED GROUP PURSUANT TO §1.1563-3T(d)(2)(iv).” The statement must include-- (A) A description of each of the controlled groups in which the corporation could be included. The description must include the name and employer identification number of each component member of each such group and the stock ownership of the component members of each such group; and (B) The following representation: [INSERT NAME AND EMPLOYER IDENTIFICATION NUMBER OF CORPORATION] ELECTS TO BE TREATED AS A COMPONENT MEMBER OF THE [INSERT DESIGNATION OF GROUP].|
||(v)|Election|-- (A) Election filed. An election filed under paragraph (d)(2)(iv) of this section is irrevocable and effective until paragraph (d)(2)(ii) or (iii) of §1.1563-3 applies or until a change in the stock ownership of the corporation results in|
|termination of membership in the controlled group in which such corporation has||
|---|---|
@@ -29,48 +20,48 @@ this section is irrevocable and effective until paragraph (d)(2)(ii) or (iii) of
Federal income tax return (including any amended return filed on or before the due date (including extensions) of such original return) timely filed on or after May 30,
2006.
(2) Expiration date. The applicability of this section will expire on May 26,
2009. Par. 51. Section 1.6012-2 is amended by revising paragraph (c) and adding paragraph (k) to read as follows: §1.6012-2 Corporations required to make returns of income.
(2) <u>Expiration date</u>. The applicability of this section will expire on May 26,
2009. Par. 51. Section 1.6012-2 is amended by revising paragraph (c) and adding paragraph (k) to read as follows: <u>§1.6012-2 Corporations required to make returns of income</u>.
* * * * *
(c) [Reserved]. For further guidance, see §1.6012-2T(c).
* * * * *
(k) [Reserved]. For further guidance, see §1.6012-2T(k)(1).
Par. 52. Section 1.6012-2T is added to read as follows: §1.6012-2T Corporations required to make returns of income (temporary).
Par. 52. Section 1.6012-2T is added to read as follows: <u>§1.6012-2T Corporations required to make returns of income (temporary)</u>.
(a) through (b) [Reserved]. For further guidance, see §1.6012-2(a) through
(b).
(c) Insurance companies-- (1) Domestic life insurance companies-- (i) In
general. A life insurance company subject to tax under section 801 shall make a return on Form 1120L. Except as provided in paragraph (c)(4) of this section, such company shall file with its return--
<u>general</u>. A life insurance company subject to tax under section 801 shall make a return on Form 1120L. Except as provided in paragraph (c)(4) of this section, such company shall file with its return--
(A) A copy of its annual statement which shows the reserves used by the
company in computing the taxable income reported on its return; and
(B) A copy of Schedule A (real estate) and of Schedule D (bonds and stocks),
or any successor thereto, of such annual statement. (ii) Mutual savings banks. Mutual savings banks conducting life insurance business and meeting the requirements of section 594 are subject to partial tax computed on Form 1120 and partial tax computed on Form 1120L. The Form 1120L is attached as a schedule to Form 1120, together with the annual statement and schedules required to be filed with Form 1120L.
or any successor thereto, of such annual statement. (ii) <u>Mutual savings banks</u>. Mutual savings banks conducting life insurance business and meeting the requirements of section 594 are subject to partial tax computed on Form 1120 and partial tax computed on Form 1120L. The Form 1120L is attached as a schedule to Form 1120, together with the annual statement and schedules required to be filed with Form 1120L.
(2) Domestic nonlife insurance companies. Every domestic insurance
(2) <u>Domestic nonlife insurance companies</u>. Every domestic insurance
company other than a life insurance company shall make a return on Form 1120PC. This includes organizations described in section 501(m)(1) that provide commercial- type insurance and organizations described in section 833. Except as provided in paragraph (c)(4) of this section, such company shall file with its return a copy of its
annual statement (or a pro forma annual statement), including the underwriting and investment exhibit for the year covered by such return.
(3) Foreign insurance companies. The provisions of paragraphs (c)(1) and
(3) <u>Foreign insurance companies</u>. The provisions of paragraphs (c)(1) and
(c)(2) of this section concerning the returns and statements of insurance companies subject to tax under section 801 or section 831 also apply to foreign insurance companies subject to tax under those sections, except that the copy of the annual statement required to be submitted with the return shall, in the case of a foreign insurance company that is not required to file an annual statement, be a copy of the pro forma annual statement relating to the United States business of such company.
(4) Exception for insurance companies filing their Federal income tax returns
electronically. If an insurance company described in paragraph (c)(1), (c)(2), or
(4) <u>Exception for insurance companies filing their Federal income tax returns</u>
<u>electronically</u>. If an insurance company described in paragraph (c)(1), (c)(2), or
(c)(3) of this section files its Federal income tax return electronically, it should not include on or with such return its annual statement (or pro forma annual statement), or any portion thereof. Such statement must be available at all times for inspection by authorized Internal Revenue Service officers or employees and retained for so long as such statements may be material in the administration of any internal revenue law. See §1.6001-1(e).
(5) Definition. For purposes of this section, the term annual statement means
(5) <u>Definition</u>. For purposes of this section, the term <u>annual statement</u> means
the annual statement, the form of which is approved by the National Association of Insurance Commissioners (NAIC), which is filed by an insurance company for the year with the insurance departments of States, Territories, and the District of
Columbia. The term annual statement also includes a pro forma annual statement if the insurance company is not required to file the NAIC annual statement.
(d) through (j) [Reserved]. For further guidance, see §1.6012-2(d) through (j).
(k) Effective date-- (1) Applicability date. This section applies to any original
(k) <u>Effective date</u>-- (1) <u>Applicability date</u>. This section applies to any original
Federal income tax return (including any amended return filed on or before the due date (including extensions) of such original return) timely filed on or after May 30,
2006.
(2) Expiration date. The applicability of this section will expire on May 26,
(2) <u>Expiration date</u>. The applicability of this section will expire on May 26,
2009.
|||Par. 53. For each entry in the “Location” column of the following table,|
@@ -165,7 +156,7 @@ section and paragraph
PART 602--OMB CONTROL NUMBERS UNDER THE PAPERWORK REDUCTION ACT Par. 54. The authority citation for part 602 continues to read as follows: Authority: 26 U.S.C. 7805. Par. 55. In §602.101, paragraph (b) is amended to read as follows:
1. The following entries to the table are removed:
§602.101 OMB Control numbers.
<u>§602.101 OMB Control numbers</u>.
* * * * *
(b) * * *
@@ -180,7 +171,7 @@ CFR part or section where Current OMB identified or described control No.
1.1081-11………………………………………………………………. 1545-2019
* * * * * **______________________________________________________________**
2. The following entries are added in numerical order to the table:
§602.101 OMB Control numbers.
<u>§602.101 OMB Control numbers</u>.
* * * * *
(b) * * *
+11 -9
View File
@@ -6,27 +6,29 @@
#### Thermodynamic Properties
**of**
®
**of** ®
# Freon 12
**(R-12)** **Technical Information** **Technical Information**
##### (R-12)
##### Technical Information Technical Information
**®** **Thermodynamic Properties of Freon 12 Refrigerant** **(R-12)** **SI Units**
Tables of the thermodynamic **Units** properties of R-12 have been developed and are presented here. P = Pressure in kPa. Absolute This information is based on values calculated using the NIST REFPROP T = Temperature in Celcius Database (McLinden, M.O., Klein,
S.A., Lemmon, E.W., and Peskin, Vf = Fluid (liquid) specific volume
A.P., NIST Standard Reference in cubic meters per kilogram Database 23, NIST thermodynamic and transport properties of Vg = Vapour (gas) specific volume refrigerants and refrigerant in cubic meters per kilogram mixtures REFPROP version 6.01, Standard Reference Data Program, df and dg = Fluid and Vapour National Institute of Standards and (respectively) densities in Technology, 1998). kilograms per cubic meter
A.P., NIST Standard Reference in cubic meters per kilogram Database 23, NIST thermodynamic and transport properties of Vg = Vapour (gas) specific volume refrigerants and refrigerant in cubic meters per kilogram mixtures REFPROP version 6.01, Standard Reference Data Program, df and dg = Fluid and Vapour National Institute of Standards and (respectively) densities in Technology, 1998).
kilograms per cubic meter
##### H = Enthalpy (kJ/kg)
##### S = Entropy (kJ/kg.K)
##### Physical Properties
|Chemical Formula|CCl2F2|
|Chemical Formula|CCl₂F₂|
|---|---|
|Molecular mass|120.91|
|Boiling Point At one atmosphere|-29.75°C|
@@ -43,9 +45,9 @@ l
**Freon** **®** **12 Saturation Properties-Temperature Table**
|Temp|Pressure||Volume|||Density||Enthalpy|||Entropy|Temp|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
|°C|[kPa]|[m3 Liquid v f|/kg]|Vapour v g|Liquid d f|[kg/m3] Vapour d g|Liquid H f|[kJ/kg] Latent H fg|Vapour H g|Liquid S f|[kJ/K-kg] Vapour S g|°C|
|Temp|Pressure||Volume||Density||Enthalpy|||Entropy|Temp|
|---|---|---|---|---|---|---|---|---|---|---|---|
|°C|[kPa]|[m³ Liquid v f|/kg] Vapour v g|[kg/m³ Liquid d f|] Vapour d g|Liquid H f|[kJ/kg] Latent H fg|Vapour H g|Liquid S f|[kJ/K-kg] Vapour S g|°C|
|-100|1.2|0.0006|10.0000|1679.0|0.100|113.3|192.8|306.1|0.6077|1.7210|-100|
|---|---|---|---|---|---|---|---|---|---|---|---|
+73
View File
@@ -203,6 +203,79 @@ class TestExtractTextWithPositions:
assert len(items) > 0
assert all(item.page == 1 for item in items)
def test_mcid(self):
# Untagged fixture: mcid is None or int, never anything else
items = pdf_inspector.extract_text_with_positions(
fixture_path("thermo-freon12.pdf")
)
assert all(item.mcid is None or isinstance(item.mcid, int) for item in items)
# Tagged fixture: marked content carries MCIDs
tagged = pdf_inspector.extract_text_with_positions(
fixture_path("firecrawl_docs_tagged.pdf")
)
assert any(item.mcid is not None for item in tagged)
# ---------------------------------------------------------------------------
# extract_structure_elements / extract_structure_elements_bytes
# ---------------------------------------------------------------------------
class TestExtractStructureElements:
def test_tagged_file(self):
elements = pdf_inspector.extract_structure_elements(
fixture_path("firecrawl_docs_tagged.pdf")
)
assert len(elements) > 0
assert all(isinstance(e.page, int) for e in elements)
assert all(isinstance(e.mcid, int) for e in elements)
assert all(isinstance(e.role, str) and len(e.role) > 0 for e in elements)
assert any(e.role == "H1" for e in elements)
def test_join_with_text_items(self):
# (page, mcid) joins against extract_text_with_positions to recover
# heading text
path = fixture_path("firecrawl_docs_tagged.pdf")
elements = pdf_inspector.extract_structure_elements(path)
items = pdf_inspector.extract_text_with_positions(path)
h1_refs = {(e.page, e.mcid) for e in elements if e.role == "H1"}
h1_text = "".join(
item.text
for item in items
if item.mcid is not None and (item.page, item.mcid) in h1_refs
)
assert len(h1_text.strip()) > 0
def test_with_pages(self):
# pages filter is 1-indexed, matching TextItem.page
elements = pdf_inspector.extract_structure_elements(
fixture_path("firecrawl_docs_tagged.pdf"), pages=[1]
)
assert len(elements) > 0
assert all(e.page == 1 for e in elements)
def test_bytes(self):
data = fixture_bytes("firecrawl_docs_tagged.pdf")
elements = pdf_inspector.extract_structure_elements_bytes(data)
assert len(elements) > 0
assert any(e.role == "H1" for e in elements)
def test_untagged_returns_empty(self):
elements = pdf_inspector.extract_structure_elements(
fixture_path("thermo-freon12.pdf")
)
assert elements == []
def test_repr(self):
elements = pdf_inspector.extract_structure_elements(
fixture_path("firecrawl_docs_tagged.pdf")
)
assert "StructureElement" in repr(elements[0])
def test_not_a_pdf(self):
with pytest.raises(ValueError):
pdf_inspector.extract_structure_elements_bytes(b"not a pdf")
# ---------------------------------------------------------------------------
# extract_text_in_regions / extract_text_in_regions_bytes
+1304
View File
File diff suppressed because it is too large Load Diff
+37
View File
@@ -0,0 +1,37 @@
[package]
name = "pdf-inspector-wasm"
version = "1.14.0"
edition = "2021"
authors = ["Firecrawl Team"]
description = "Browser WebAssembly bindings for pdf-inspector"
license = "MIT"
repository = "https://github.com/firecrawl/pdf-inspector"
homepage = "https://github.com/firecrawl/pdf-inspector"
readme = "README.md"
publish = false
[lib]
crate-type = ["cdylib", "rlib"]
[dependencies]
console_error_panic_hook = "0.1"
js-sys = "0.3"
pdf-inspector = { path = ".." }
serde = { version = "1", features = ["derive"] }
serde-wasm-bindgen = "0.6"
wasm-bindgen = "0.2"
[dev-dependencies]
wasm-bindgen-test = "0.3"
[profile.release]
codegen-units = 1
lto = true
opt-level = "s"
strip = true
[package.metadata.wasm-pack.profile.release]
# Rust 1.95 emits bulk-memory instructions that the binaryen bundled with
# wasm-pack 0.15.0 does not yet validate. rustc still performs the release,
# size, and LTO optimizations above.
wasm-opt = false
+58
View File
@@ -0,0 +1,58 @@
MIT License
Copyright (c) 2026 Firecrawl
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
Third-party notices
===================
Adobe CMaps
-----------
The WebAssembly binary embeds binary CMaps derived from Adobe CMap resources.
Copyright 1990-2009 Adobe Systems Incorporated.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
Redistributions of source code must retain the above copyright notice, this
list of conditions and the following disclaimer.
Redistributions in binary form must reproduce the above copyright notice,
this list of conditions and the following disclaimer in the documentation
and/or other materials provided with the distribution.
Neither the name of Adobe Systems Incorporated nor the names of its
contributors may be used to endorse or promote products derived from this
software without specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE
LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR
CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF
SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS
INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN
CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE)
ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE
POSSIBILITY OF SUCH DAMAGE.
+60
View File
@@ -0,0 +1,60 @@
# @firecrawl/pdf-inspector-wasm
Browser WebAssembly bindings for [pdf-inspector](https://github.com/firecrawl/pdf-inspector). Classify PDFs and extract structured Markdown locally from a `Uint8Array`, using the same Rust core as the native Node.js, Python, and Rust packages.
## Install
```bash
npm install @firecrawl/pdf-inspector-wasm
```
## Usage
```ts
import init, { processPdf } from "@firecrawl/pdf-inspector-wasm";
await init();
const response = await fetch("/annual-report.pdf");
const pdf = new Uint8Array(await response.arrayBuffer());
const result = processPdf(pdf);
console.log(result.pdfType);
console.log(result.markdown);
```
Pass options when you need selected pages or compact Markdown:
```ts
const result = processPdf(pdf, {
pages: [1, 3, 5],
profile: "compact",
includePageMarkers: true,
});
```
The package also exports:
- `detectPdf(pdf, options?)` for detection without extraction.
- `classifyPdf(pdf)` for the lightweight result shape shared with the native Node.js API.
- `extractText(pdf)` for plain text.
- `version()` for the WASM package version.
## Browser behavior
- Parsing runs locally. PDF bytes are not uploaded anywhere.
- The build is single-threaded and does not require cross-origin isolation.
- CMaps are embedded so CJK font decoding does not depend on a filesystem.
- Extraction is synchronous after `init()`. For large documents, call it from a Web Worker to keep the UI responsive.
- Image-only documents still require a separate OCR step.
## Build from source
```bash
cargo install wasm-pack --version 0.15.0 --locked
wasm-pack build wasm --target web --scope firecrawl --release
```
## License
MIT
+440
View File
@@ -0,0 +1,440 @@
use pdf_inspector::{
LayoutComplexity, MarkdownProfile, PageOcrReasons, PdfOptions, PdfProcessResult, PdfType,
ProcessMode,
};
use serde::{Deserialize, Serialize};
use wasm_bindgen::prelude::*;
#[wasm_bindgen(typescript_custom_section)]
const TYPESCRIPT_TYPES: &str = r#"
export type PdfType = "TextBased" | "Scanned" | "ImageBased" | "Mixed";
export type MarkdownProfile = "fidelity" | "compact";
export interface ProcessOptions {
/** Restrict extraction to these 1-indexed page numbers. */
pages?: number[];
/** Password for an encrypted PDF. */
password?: string;
/** Source-faithful output by default, or compact output for fewer tokens. */
profile?: MarkdownProfile;
/** Insert `<!-- Page N -->` markers between pages. */
includePageMarkers?: boolean;
/** Include image placeholders in Markdown output. */
includeImages?: boolean;
}
export interface PageOcrReasons {
/** 1-indexed page number. */
page: number;
reasons: string[];
}
export interface LayoutComplexity {
isComplex: boolean;
/** 1-indexed page numbers. */
pagesWithTables: number[];
/** 1-indexed page numbers. */
pagesWithColumns: number[];
}
export interface PdfProcessResult {
pdfType: PdfType;
markdown?: string;
pageCount: number;
processingTimeMs: number;
/** 1-indexed page numbers. */
pagesNeedingOcr: number[];
ocrReasonsByPage: PageOcrReasons[];
title?: string;
confidence: number;
layout: LayoutComplexity;
hasEncodingIssues: boolean;
}
export interface PdfClassification {
pdfType: PdfType;
pageCount: number;
/** 0-indexed page numbers, matching the native Node.js API. */
pagesNeedingOcr: number[];
confidence: number;
}
export function processPdf(data: Uint8Array, options?: ProcessOptions): PdfProcessResult;
export function detectPdf(data: Uint8Array, options?: Pick<ProcessOptions, "password">): PdfProcessResult;
export function classifyPdf(data: Uint8Array): PdfClassification;
export function extractText(data: Uint8Array): string;
export function version(): string;
"#;
#[derive(Debug, Default, Deserialize)]
#[serde(default, rename_all = "camelCase", deny_unknown_fields)]
struct WasmProcessOptions {
pages: Option<Vec<u32>>,
password: Option<String>,
profile: Option<WasmMarkdownProfile>,
include_page_markers: Option<bool>,
include_images: Option<bool>,
}
#[derive(Debug, Deserialize)]
#[serde(rename_all = "lowercase")]
enum WasmMarkdownProfile {
Fidelity,
Compact,
}
#[derive(Serialize)]
#[serde(rename_all = "camelCase")]
struct WasmPageOcrReasons {
page: u32,
reasons: Vec<String>,
}
impl From<PageOcrReasons> for WasmPageOcrReasons {
fn from(value: PageOcrReasons) -> Self {
Self {
page: value.page,
reasons: value.reasons,
}
}
}
#[derive(Serialize)]
#[serde(rename_all = "camelCase")]
struct WasmLayoutComplexity {
is_complex: bool,
pages_with_tables: Vec<u32>,
pages_with_columns: Vec<u32>,
}
impl From<LayoutComplexity> for WasmLayoutComplexity {
fn from(value: LayoutComplexity) -> Self {
Self {
is_complex: value.is_complex,
pages_with_tables: value.pages_with_tables,
pages_with_columns: value.pages_with_columns,
}
}
}
#[derive(Serialize)]
#[serde(rename_all = "camelCase")]
struct WasmPdfProcessResult {
pdf_type: &'static str,
markdown: Option<String>,
page_count: u32,
processing_time_ms: f64,
pages_needing_ocr: Vec<u32>,
ocr_reasons_by_page: Vec<WasmPageOcrReasons>,
title: Option<String>,
confidence: f64,
layout: WasmLayoutComplexity,
has_encoding_issues: bool,
}
impl From<PdfProcessResult> for WasmPdfProcessResult {
fn from(value: PdfProcessResult) -> Self {
Self {
pdf_type: pdf_type_name(value.pdf_type),
markdown: value.markdown,
page_count: value.page_count,
processing_time_ms: value.processing_time_ms as f64,
pages_needing_ocr: value.pages_needing_ocr,
ocr_reasons_by_page: value
.ocr_reasons_by_page
.into_iter()
.map(Into::into)
.collect(),
title: value.title,
confidence: value.confidence as f64,
layout: value.layout.into(),
has_encoding_issues: value.has_encoding_issues,
}
}
}
#[derive(Serialize)]
#[serde(rename_all = "camelCase")]
struct WasmPdfClassification {
pdf_type: &'static str,
page_count: u32,
pages_needing_ocr: Vec<u32>,
confidence: f64,
}
fn pdf_type_name(pdf_type: PdfType) -> &'static str {
match pdf_type {
PdfType::TextBased => "TextBased",
PdfType::Scanned => "Scanned",
PdfType::ImageBased => "ImageBased",
PdfType::Mixed => "Mixed",
}
}
fn js_error(context: &str, error: impl std::fmt::Display) -> JsValue {
js_sys::Error::new(&format!("{context}: {error}")).into()
}
fn deserialize_options(value: JsValue) -> Result<WasmProcessOptions, JsValue> {
if value.is_undefined() || value.is_null() {
return Ok(WasmProcessOptions::default());
}
serde_wasm_bindgen::from_value(value).map_err(|error| js_error("invalid options", error))
}
fn build_options(value: JsValue, mode: ProcessMode) -> Result<PdfOptions, JsValue> {
let options = deserialize_options(value)?;
if options
.pages
.as_ref()
.is_some_and(|pages| pages.contains(&0))
{
return Err(js_error(
"invalid options",
"pages are 1-indexed; page 0 is invalid",
));
}
let mut result = PdfOptions::new().mode(mode);
if let Some(pages) = options.pages {
result = result.pages(pages);
}
if let Some(password) = options.password {
result = result.password(password);
}
if let Some(profile) = options.profile {
result.markdown.profile = match profile {
WasmMarkdownProfile::Fidelity => MarkdownProfile::Fidelity,
WasmMarkdownProfile::Compact => MarkdownProfile::Compact,
};
}
if let Some(include_page_markers) = options.include_page_markers {
result.markdown.include_page_numbers = include_page_markers;
}
if let Some(include_images) = options.include_images {
result.markdown.include_images = include_images;
}
Ok(result)
}
fn serialize<T: Serialize>(value: &T) -> Result<JsValue, JsValue> {
serde_wasm_bindgen::to_value(value).map_err(|error| js_error("serialize result", error))
}
fn initialize() {
console_error_panic_hook::set_once();
}
/// Process PDF bytes entirely inside WebAssembly.
#[wasm_bindgen(js_name = processPdf, skip_typescript)]
pub fn process_pdf(data: &[u8], options: JsValue) -> Result<JsValue, JsValue> {
initialize();
let options = build_options(options, ProcessMode::Full)?;
let started = js_sys::Date::now();
let mut result = pdf_inspector::process_pdf_mem_with_options(data, options)
.map_err(|error| js_error("process PDF", error))?;
result.processing_time_ms = (js_sys::Date::now() - started).max(0.0) as u64;
serialize(&WasmPdfProcessResult::from(result))
}
/// Classify PDF bytes without extracting text or producing Markdown.
#[wasm_bindgen(js_name = detectPdf, skip_typescript)]
pub fn detect_pdf(data: &[u8], options: JsValue) -> Result<JsValue, JsValue> {
initialize();
let options = build_options(options, ProcessMode::DetectOnly)?;
let started = js_sys::Date::now();
let mut result = pdf_inspector::process_pdf_mem_with_options(data, options)
.map_err(|error| js_error("detect PDF", error))?;
result.processing_time_ms = (js_sys::Date::now() - started).max(0.0) as u64;
serialize(&WasmPdfProcessResult::from(result))
}
/// Return the lightweight classification shape used by the native Node API.
#[wasm_bindgen(js_name = classifyPdf, skip_typescript)]
pub fn classify_pdf(data: &[u8]) -> Result<JsValue, JsValue> {
initialize();
let result =
pdf_inspector::classify_pdf_mem(data).map_err(|error| js_error("classify PDF", error))?;
serialize(&WasmPdfClassification {
pdf_type: pdf_type_name(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
}
/// Extract plain text from PDF bytes without Markdown conversion.
#[wasm_bindgen(js_name = extractText, skip_typescript)]
pub fn extract_text(data: &[u8]) -> Result<String, JsValue> {
initialize();
let items = pdf_inspector::extractor::extract_text_with_positions_mem(data)
.map_err(|error| js_error("extract text", error))?;
Ok(
pdf_inspector::extractor::group_into_lines_preserving_all_text(items)
.into_iter()
.map(|line| line.text())
.filter(|line| !line.trim().is_empty())
.collect::<Vec<_>>()
.join("\n"),
)
}
/// Return the WebAssembly package version.
#[wasm_bindgen(skip_typescript)]
pub fn version() -> String {
env!("CARGO_PKG_VERSION").to_string()
}
#[cfg(all(test, target_arch = "wasm32"))]
mod tests {
use super::*;
use js_sys::Reflect;
use wasm_bindgen_test::*;
const TEXT_PDF: &[u8] = include_bytes!("../../tests/fixtures/thermo-freon12.pdf");
const ENCRYPTED_PDF: &[u8] = include_bytes!("../../tests/fixtures/encrypted-secret123.pdf");
fn synthetic_korea1_pdf() -> Vec<u8> {
let mut pdf = b"%PDF-1.4\n".to_vec();
let mut offsets = vec![0usize];
fn add_object(pdf: &mut Vec<u8>, offsets: &mut Vec<usize>, id: usize, body: &str) {
offsets.push(pdf.len());
pdf.extend_from_slice(format!("{id} 0 obj\n").as_bytes());
pdf.extend_from_slice(body.as_bytes());
pdf.extend_from_slice(b"\nendobj\n");
}
add_object(
&mut pdf,
&mut offsets,
1,
"<< /Type /Catalog /Pages 2 0 R >>",
);
add_object(
&mut pdf,
&mut offsets,
2,
"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
);
add_object(
&mut pdf,
&mut offsets,
3,
"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] /Resources << /Font << /F0 5 0 R >> >> /Contents 4 0 R >>",
);
// Adobe-Korea1 CID 1086 (0x043E) maps to U+AC00 (Korean syllable GA).
// There is deliberately no ToUnicode stream: decoding must use the
// embedded predefined CMap rather than lopdf's plain-text fallback.
// Korea1 CIDs 21 and 19 map to ASCII "4" and "2". Place them near
// the bottom edge so they look exactly like a numeric page footer.
let content = "BT /F0 12 Tf 50 100 Td <043E> Tj 0 -60 Td <00150013> Tj ET";
add_object(
&mut pdf,
&mut offsets,
4,
&format!(
"<< /Length {} >>\nstream\n{}\nendstream",
content.len(),
content
),
);
add_object(
&mut pdf,
&mut offsets,
5,
"<< /Type /Font /Subtype /Type0 /BaseFont /SyntheticKorea1 /Encoding /Identity-H /DescendantFonts [6 0 R] >>",
);
add_object(
&mut pdf,
&mut offsets,
6,
"<< /Type /Font /Subtype /CIDFontType2 /BaseFont /SyntheticKorea1 /CIDSystemInfo << /Registry (Adobe) /Ordering (Korea1) /Supplement 2 >> /FontDescriptor 7 0 R /DW 1000 >>",
);
add_object(
&mut pdf,
&mut offsets,
7,
"<< /Type /FontDescriptor /FontName /SyntheticKorea1 /Flags 4 /FontBBox [-100 -200 1000 900] /ItalicAngle 0 /Ascent 800 /Descent -200 /CapHeight 700 /StemV 80 >>",
);
let xref_start = pdf.len();
pdf.extend_from_slice(format!("xref\n0 {}\n", offsets.len()).as_bytes());
pdf.extend_from_slice(b"0000000000 65535 f \n");
for offset in offsets.iter().skip(1) {
pdf.extend_from_slice(format!("{offset:010} 00000 n \n").as_bytes());
}
pdf.extend_from_slice(
format!(
"trailer\n<< /Size {} /Root 1 0 R >>\nstartxref\n{}\n%%EOF",
offsets.len(),
xref_start
)
.as_bytes(),
);
pdf
}
#[wasm_bindgen_test]
fn processes_pdf_to_markdown() {
let result = process_pdf(TEXT_PDF, JsValue::UNDEFINED).expect("process PDF");
let pdf_type = Reflect::get(&result, &JsValue::from_str("pdfType"))
.expect("pdfType")
.as_string()
.expect("pdfType string");
let markdown = Reflect::get(&result, &JsValue::from_str("markdown"))
.expect("markdown")
.as_string()
.expect("markdown string");
assert_eq!(pdf_type, "TextBased");
assert!(!markdown.is_empty());
}
#[wasm_bindgen_test]
fn rejects_non_pdf_bytes() {
assert!(process_pdf(b"not a PDF", JsValue::UNDEFINED).is_err());
}
#[wasm_bindgen_test]
fn classifies_and_extracts_plain_text() {
let classification = classify_pdf(TEXT_PDF).expect("classify PDF");
let pdf_type = Reflect::get(&classification, &JsValue::from_str("pdfType"))
.expect("pdfType")
.as_string()
.expect("pdfType string");
let text = extract_text(TEXT_PDF).expect("extract text");
assert_eq!(pdf_type, "TextBased");
assert!(!text.is_empty());
}
#[wasm_bindgen_test]
fn extracts_cjk_and_preserves_numeric_page_footer() {
let text = extract_text(&synthetic_korea1_pdf()).expect("extract predefined CMap text");
assert_eq!(text, "\n42");
}
#[wasm_bindgen_test]
fn opens_encrypted_pdf_with_password() {
assert!(process_pdf(ENCRYPTED_PDF, JsValue::UNDEFINED).is_err());
let options = js_sys::Object::new();
Reflect::set(
&options,
&JsValue::from_str("password"),
&JsValue::from_str("secret123"),
)
.expect("set password");
let result = process_pdf(ENCRYPTED_PDF, options.into()).expect("process encrypted PDF");
let markdown = Reflect::get(&result, &JsValue::from_str("markdown"))
.expect("markdown")
.as_string()
.expect("markdown string");
assert!(!markdown.is_empty());
}
}