Compare commits

...
Author SHA1 Message Date
Abimael MartellandClaude Opus 4.7 a25d510fd9 Bump version from 1.8.9 to 1.8.10
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 19:19:09 -04:00
Abimael MartellandClaude Opus 4.7 43bb62647e extract_tables: try vector-grid detectors before text heuristic
extract_tables_in_regions_mem previously ran only the text-only
heuristic detector (tables::detect_tables) on the items inside each
region, discarding the rects and lines that extract_page_text_items
returned. That left the rect-backed and line-backed detectors
(detect_tables_from_rects, detect_tables_from_lines) unused by the
public region-scoped extraction path — they only ran through
detect_vector_grid_in_region_mem, which most callers don't use.

Keep the rects and lines, filter them to each region, and try in
order: rect detector → line detector → heuristic. Each candidate's
markdown is quality-gated by the existing needs_ocr checks
(is_garbage_text, is_cid_garbage, detect_encoding_issues,
looks_like_partial_table_ex); only the first clean output wins.
If all three produce empty or noisy output we still return
needs_ocr=true, matching prior behavior.

Effect on real prod-shape inputs from shadow logs:

  Full-page ruled ledger, 6 cols x ~15 rows:
    before: heuristic emits a 355-char two-row fragment
    after:  line detector emits the full 6520-char table

  Multi-row key/value layout with paragraph values:
    before: heuristic emits a 188-char header-only fragment
    after:  rect detector emits the full 1733-char table including
            the multi-bullet description cell

Existing fixtures that already passed via the heuristic continue to
pass: the quality gate rejects partial vector-grid output and falls
through, so the heuristic still wins where it produced the cleaner
result.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 19:13:14 -04:00
Abimael Martell f26efe6673 Bump version from 1.8.8 to 1.8.9 2026-05-12 18:40:12 -04:00
Abimael MartellandClaude Opus 4.7 06ccaf5732 detect_rects: accept multi-row tables with long-content cells (#84)
Three rejection sites in detect_rects.rs killed any candidate grid
where a single cell exceeded 500 chars:

  - detect_row_stripe_table (line 1399)
  - detect_row_stripe_table_from_cell_rects (line 1729)
  - detect_merged_cluster_table (line 2154)

The intent was to skip layout-background rects — sidebars, banners,
section bands — where one big rectangle wraps a paragraph of prose.
Those almost always present as ≤3 row stripes (header / body / footer
or single big block).

Multi-row key/value tables with paragraph-length values in one column
present the same cell-length signal but are legitimate tables.
Gating the rejection on `non_empty_rows < 4` preserves the
layout-background guard for narrow stripe layouts while letting
through multi-row tables with descriptive content.

Verified against a 9-row × 2-column key/value layout where the value
column has multi-line content (~1.4KB in the longest cell). Before:
rejected with `max cell length 1384 > 500`. After: accepted with 89%
density. The existing `test_row_stripe_rejects_layout_background_long_cells`
regression test for narrow stripe layouts still passes.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 18:35:56 -04:00
Abimael MartellandClaude Opus 4.7 5a84c29ea4 detect_lines: accept full-page tables with internal grid (#83)
The page-spanning-frame guard rejected any line set whose bounding box
exceeded ~90% of a standard A4/Letter page in both axes. The intent was
to skip decorative outer borders, but it also threw away every real
full-page table — common in governmental ledgers, financial filings,
and dense report layouts.

Decorative borders have just 4 edges (top/bottom/left/right). Real
full-page tables have many internal row and column rules. Gate the
rejection on `horizontals.len() <= 4 && verticals.len() <= 4` so the
guard still catches bare frames but lets through line sets with real
internal grid structure.

Tested against a regione.lazio.it Estrazione-provvedimenti page (full
A4-width table, 14 rows × 6 cols): now ACCEPTED with 211/211 items
captured. Bare-frame regression test added.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 18:28:08 -04:00
Abimael MartellandClaude Opus 4.7 96c0f2102a tables/detect_rects: prefer rect-grid edges over text-cluster on N≥3 column tables (#82)
When a wire-bordered table has headers centered/right-aligned in their
cells but data left-aligned, cluster_x_positions can both merge adjacent
data columns (when the data-to-data gap is below the clamped threshold)
and drop the header-only x-positions in its singleton-filter pass. The
cell-rect fallback then used text-cluster column edges and lost a column
or fragmented neighbor cells.

Prefer rect-derived column edges when the rect grid has 3+ columns and
every rect column holds multiple text items. The all-cols-populated
check protects against decorative or background rects (prose laid out
in a frame, cell-fill rects with extra borders) that would otherwise
split a logical column into spurious sub-columns. The existing
prose-in-frame, well-distributed-columns, and wireless-prose guards
still fire for the cases they were built for.

Bump napi version 1.8.7 → 1.8.8.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 16:44:26 -04:00
Abimael Martell 6086577d11 tables/detect_rects: emit grid for multiline indented cells (#79) 2026-05-07 11:44:53 -07:00
Abimael MartellandCursor f2186ec1aa tables/detect_rects: don't accept relaxed grid on wireless prose (#78)
Require rect-derived column evidence before relaxing prose checks for two-column cell-rect fallbacks, so text-position alignment alone cannot synthesize a vector grid on wireless content.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-07 09:30:29 -07:00
Abimael MartellandClaude Opus 4.7 59b17f372a tables: tighten prose-in-frame rejection (#77)
* tables: lift detection on shaded-header + alt-row tables (#wired-grids)

Production telemetry on `wired_high_confidence`-classified table regions
showed `detect_vector_grid_in_region_mem` returning a usable grid only
~27% of the time, with the rest falling through to GLM-OCR. Three
surgical fixes target the dominant production shapes:

* Path-fill cell backgrounds: when the page has no `re` rects but draws
  cell backgrounds via `m`/`l`/`h`/`f*` sequences, prefer the fill-derived
  rects over the few section-level `W*` clip paths that previously won
  the priority gate. Activated when fill rects outnumber clip rects ≥3×.

* Dedup-induced cluster splits: page-background rects could pose as
  containers in the sub-rect dedup and evict a slightly smaller
  table-frame rect, breaking adjacency between column-cell groups so each
  column became its own cluster. Origin-anchored containers are now
  disqualified from sub-rect dedup. A separate exact-duplicate pass
  collapses the cell-padding/text-bg/cell-border triple emissions some
  PDFs produce, preserving original order to avoid reshuffling table
  output on multi-table pages.

* Prose-words rejection: the `cell-rect` fallback's whole-grid prose
  threshold also rejected real tables that include a description column.
  Now relaxed when content is well-distributed (≥75% of cols filled),
  while keeping the original strictness for prose-in-a-frame layouts.

Two regression fixtures from the opendataloader-bench corpus, covering
the dominant production failure categories:

* `greencomp_competence.pdf` — 2-col shaded-header + plain-body glossary.
  Mirrors production crops #1 (Contractions glossary) and #6 (BIO 350
  course header).
* `upstage_key_functions.pdf` — 4-col shaded-header + alt-row backgrounds
  + merged left column. Mirrors production crops #2 (Parameter/Value
  alt-row), #7 (Spanish XML schema), and #8 (Córdoba multi-row header).

Existing fixtures stay green (doc 51 wrapped-label, doc 128 forecast
six-cols, td9264 snapshot). 133 unit + integration tests pass; clippy
clean.

Bumps napi/package.json 1.8.4 → 1.8.5.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* tables: tighten prose-in-frame rejection — fixes pdf-evals #30 regression

PR #76's shaded-header detection lift surfaced a regression on
accessory_building_permit_application_1 (TEDS 0.10 → 0.05): a
paragraph of legal text laid out in a 2-column justified block was
being admitted as a 10×2 fake table where every cell holds a
sentence fragment ("I agree to comply...", "I", "It is the property
owner's responsibility..."). Per pdf-evals PR #30 review, this is
the kind of regression production users will notice — the markdown
is structurally and semantically misleading.

Root cause: PR #76's prose-rejection only fires for `num_cols >= 4`,
so the 2-col prose-in-a-frame case slipped past it entirely. The new
fill-priority + dedup changes started producing rects for this layout
that 1.8.4 correctly ignored.

Fix: tighten the prose-in-frame check.
- Lower the column-count guard from `>= 4` to `>= 2`.
- Add a content-length signal as the primary discriminator: when the
  prose-words trigger fires AND mean non-empty cell length exceeds
  65 chars, reject regardless of column distribution.

The 65-char threshold cleanly separates observed cases:
  accessory_building (prose-in-frame): mean 74 chars  → REJECT
  upstage_key_functions (real 4-col table): mean 53   → admit
  greencomp_competence (real 2-col glossary): mean 20 → admit
  accessory_building (real 5×3 form data): mean 10    → admit

The well-distributed-cols relaxation that PR #76 added stays —
"label / value / description / benefit" tables (#7, #8 from the
production crops) still pass, but only when their mean cell length
stays below the prose threshold.

New regression test `accessory_building_rejects_prose_in_frame` asserts
both that the real 5×3 form data table survives AND the 10×2 prose
block is rejected. Snapshot test `test_snapshot_td9264` updated to
match new output — old snapshot captured the same prose-in-frame bug
on regulatory text (paragraphs emitted as 3-col `||text||` fake-table
rows). New snapshot emits clean prose paragraphs, which is correct.

Verification:
- cargo test --all: 424 lib + 133 integration + 2 doc tests pass
- cargo fmt --check clean
- cargo clippy -- -D warnings clean (lib-level; pre-existing
  test-level clippy issues on the wired-grids branch unaffected)
- Existing fixtures stay green: forecast_table_chart_six_cols (PR
  #72), bits_pilani_* (PR #73), greencomp_competence_two_cols and
  upstage_key_functions_four_cols (PR #76).

This branch is based on abimaelmartell/wired-grids so it includes
PR #76's commits plus this fix on top. Suggest merging this and
closing #76, OR rebasing #76 to incorporate this fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* tables: drop early-dedup atom — caused broad TOC + matrix corruption

Bisected PR #76's 4 atoms against the SEC 10-K
0001104659-25-093871_183e44ac.pdf which appeared as a TEDS regression
in pdf-evals PR #30. Result:

  atom                       | TOC | perf-graph | qualifications
  ---------------------------|-----|------------|---------------
  fill-priority              |  ✓  |     ✓      |       ✓
  early-dedup                |  ✗  |     ✗      |       ✗
  page-bg disqualification   |  ✓  |     ✓      |       ✓
  prose-relaxation           |  ✓  |     ✓      |       ✓

Early-dedup was the SOLE source of all three regressions on this doc.
Tried a more conservative variant (≥3 copies only — pair-duplicates
appear in legit multi-section layouts like 10-K dividers above + below
section headers); didn't fix the regression. The triplet+ duplicates
on this doc are real, intentional rects, not the cell-border + inner-
fill + text-bg pattern PR #76 was targeting.

Drop early-dedup. Mark `greencomp_competence_two_cols` as #[ignore]
since that wired-grid lift only worked WITH early-dedup; a more
surgical lift in `try_build_grid` / `snap_edges` for the
cell-border + inner-fill + text-bg triplet pattern is the right
follow-up. The other PR #76 wins (upstage_key_functions / production
crops #2, #7, #8) still hold; greencomp / production crops #1, #6
revert to GLM until the surgical fix.

Validation on the regression doc:
  0001104659 TOC PART II markers:    4 (matches main, was 2 with PR#76)
  0001104659 perf-graph data row:    2 (matches main, was 1)
  0001104659 qualifications rows:    9 (matches main, was 6)
Validation on the prose-frame doc:
  accessory_building fake-table:     0 (matches main, was 1 with PR#76)
  accessory_building prose intact:   1 (matches main)

cargo test --all clean, cargo fmt --check clean, cargo clippy --lib
-- -D warnings clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 13:37:08 -07:00
Abimael Martell 88844d18be Bump version from 1.8.3 to 1.8.4 2026-05-02 21:29:51 -07:00
Abimael MartellandClaude Opus 4.7 cfc080f79a extractor: drop Latin-1 mojibake on Type0/CID fonts; tokenize wide TSR items (#75)
* extractor: drop Latin-1 mojibake on Type0/CID fonts; tokenize wide TSR items

Two text-extraction failure modes surfaced by table-candidate shadow
data; both also affect the existing TableFormer / vector-grid paths
since they share `extract_tables_with_structure_*_mem`'s downstream
cell-fill.

1. CJK / multi-byte mojibake. The bottom Latin-1 fallback in
   `extract_text_from_operand` ran unconditionally. For a Type0/CID
   (Identity-H) font whose ToUnicode CMap fails to parse, the bytes
   are CIDs (font-internal indices), not character codes — per-byte
   Latin-1 produces mojibake (e.g. 2-byte CID 0xCDD9 surfaces as "ÍÙ").
   Gate that fallback on `FontWidthInfo.is_cid` (set by
   `parse_type0_widths` for `/Subtype /Type0`). For Type0 fonts with
   any non-ASCII byte, emit one U+FFFD per CID instead so
   `detect_encoding_issues` still trips and the page is flagged for
   OCR — preserving the existing OCR-routing path that the
   high-Latin-1 garbage used to satisfy by accident. Type1 / TrueType
   simple fonts retain the per-byte Latin-1 round-trip (it IS the
   canonical interpretation for them; verified against an existing
   pdf-evals fixture where bytes like 0xB6 are legitimate Latin-1).

   Threaded `font_widths: &PageFontWidths` through
   `extract_text_from_operand` and its 5 call sites in
   `content_stream.rs` / `xobjects.rs`.

2. Dense-cell text collapse in `extract_tables_with_structure_cells_mem`.
   Stage-1 routing did per-item assignment — each TextItem went into the
   single cell whose bbox contained its center. When a row's text is
   rendered as one wide Tj (e.g. "Marshall Islands 0.9 0.9 0.9"), the
   whole row parks in one cell and the rest of the row stays empty.

   New `split_item_into_token_subitems` helper splits each item into
   per-token virtual sub-items with x positions estimated from
   `effective_width / char_count` and the token's character offset.
   Stage 1 then routes per-token. Single-token items collapse to a
   one-element vector (no behavior change). Multi-token items spanning
   multiple cells distribute correctly. Stage-2 orphan recovery now
   operates on token-grain orphans rather than re-trying whole items.

Tests:
 - `cid_font_with_unparseable_cmap_does_not_emit_latin1_mojibake` (unit)
   exercises the Type0/CID + unparseable-CMap fallback path.
 - `simple_font_latin1_fallback_passes_high_bytes_through` (unit)
   guards the false-positive case where a Type1 font's `/ToUnicode`
   reference is set but bytes are legitimate Latin-1 character codes.
 - `test_extract_tables_with_structure_distributes_wide_item_across_cells`
   (integration) builds a synthetic PDF with one wide Tj and asserts
   each token lands in its own cell.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* tests: pin CID mojibake fix mechanism with FFFD assertions

Two complementary tests for the Type0/CID Latin-1-fallback guard:

1. Tighten `test_identity_h_no_tounicode_suppresses_garbage` on the
   existing real-PDF fixture `shinagawa_identity_h.pdf` to also assert
   the pre-suppression text contains U+FFFD and contains no high-Latin-1
   chars. Pins down WHICH mechanism is suppressing the garbage so a
   future regression that re-enables Latin-1 mojibake fails loudly here
   instead of silently switching the suppression chain back to
   `is_cid_garbage` + high-Latin-1 detection.

2. Add `test_synthetic_type0_broken_tounicode_emits_fffd_not_latin1_mojibake`
   with a fully-synthetic Type0 / Identity-H PDF built in process. We
   control the malformed ToUnicode contents, the descendant CIDFontType2
   shape (just enough for `parse_type0_widths` to set `is_cid=true`,
   which is what the new guard keys off of), and the Tj byte stream.
   No fixture file or external license needed. Reproduces the exact
   "Type0 + non-ASCII bytes + unparseable ToUnicode" code path that
   produced the production mojibake samples.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-02 21:28:39 -07:00
Abimael Martell 1f28e523fd tables: keep wrapped labels in TSR output (#73) 2026-04-30 10:56:05 -07:00
Abimael Martell 97fc32ac70 tables: prefer rect edges for cell-grid fallback (#72)
* tables: prefer rect edges in cell-grid fallback

* Bump version from 1.8.1 to 1.8.2
2026-04-29 21:27:27 -07:00
Abimael Martell a4161c8392 Bump version from 1.8.0 to 1.8.1 2026-04-29 08:23:05 -07:00
Abimael Martell 5b1fe30c66 tables: expand multi-row cells in-place when fallback heuristic is empty (#71)
* tables: expand multi-row TSR cells in place

Recover row-under-counted TSR tables by splitting overstuffed cells with native PDF text bands before falling back to heuristic extraction.

Made-with: Cursor

* docs: note multi-row expansion scope

Clarify that the row-band cap intentionally keeps v1 focused on common small row-loss cases while larger compressions continue to use heuristic fallback.

Made-with: Cursor
2026-04-29 08:18:09 -07:00
Abimael Martell 63b5573133 Bump version from 1.7.2 to 1.8.0 2026-04-28 17:21:44 -07:00
Abimael Martell c186a036fc tables: add detectVectorGridInRegion napi export for region-scoped vector grid detection (#70)
* feat: add vector grid region detector napi export

Expose region-scoped vector PDF grid detection so TSR callers can reuse native geometry before model fallback.

Made-with: Cursor

* fix: address vector grid review feedback

Return null for rotated vector grids until the coordinate transform has coverage and reject out-of-crop cell boxes surfaced by real-PDF smoke testing.

Made-with: Cursor

* test: add crop bbox plausibility coverage

Cover in-crop, out-of-crop, slack-boundary, and non-positive DPI behavior for vector grid cell bbox validation.

Made-with: Cursor
2026-04-28 17:20:15 -07:00
Abimael MartellandClaude Opus 4.7 d196d435d1 fix: TSR auto-fallback bugs found in review, v1.7.2 (#68)
Three fixes to extract_tables_with_structure_auto_mem (added in
1.7.1) caught by external review:

1. multi_row_in_cell over-triggered on legitimate multi-line cells.
   The previous threshold (item span > 1.3× either smallest cell or
   tallest item height) fires on any cell with 2+ y-separated text
   items — including rowspan>1 cells, wrapped descriptions, and
   superscript/subscript runs. Replaced with two gates:
   - skip cells whose declared rowspan > 1 (intentional multi-line)
   - require an actual whitespace gap (>~half a line height)
     between the bottom of one item and the top of the next, in
     PDF-native y-coordinates. Same-line items with tall glyphs or
     superscripts have negative or near-zero gap; truly separate
     visual rows have gap ≈ leading − line-height.
   FNBO regression test still passes; new test covers a rowspan=2
   cell with two visible text lines and verifies no fallback fires.

2. Heuristic returning empty silently replaced TSR markdown with
   "". The auto wrapper now keeps the TSR markdown when the
   heuristic markdown is empty/whitespace and tags fallback_reason
   with `_heuristic_empty` suffix (e.g.
   `multi_row_in_cell_heuristic_empty`). Worst case we ship the
   same wrong-but-non-empty TSR output we'd have shipped before
   1.7.1; we never replace useful output with literally nothing.

3. One bad input blanked the whole batch. Errors from
   detect_tsr_quality_issue or extract_tables_in_regions_mem now
   stay scoped to the single input — that input falls through to
   raw TSR markdown with a `_error` reason label so callers can
   metric on it. Other inputs in the batch are unaffected.

3 new integration tests:
- test_auto_does_not_fire_on_legit_rowspan_cell
- test_auto_keeps_tsr_markdown_when_heuristic_returns_empty
- test_auto_isolates_per_input_failures

All 6 auto tests + full 123-test suite pass. FNBO local replay
still triggers fallback (phantom_empty_row signal in this run) and
emits correct Shawnee/BVP/Sonoma rows with correct census tracts.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-27 11:23:03 -07:00
Abimael MartellandClaude Opus 4.7 fbab84fc20 feat: TSR auto-fallback to heuristic on quality issues, v1.7.1 (#67)
Adds extract_tables_with_structure_auto_mem (Rust) /
extractTablesWithStructureAuto (napi). Returns
TableExtractionResult { markdown, fallback_reason } per input.

The wrapper runs the existing TSR-hybrid path then checks the
resulting cells for two known SLANet detection pathologies:

* phantom_empty_row: empty row sandwiched between non-empty rows
  (cheap, cell-metadata only).
* multi_row_in_cell: re-reads PDF text items, flags any cell whose
  contained items span >1.3× either the smallest cell height or
  the tallest contained item height. Catches the FNBO failure mode
  where a tall TSR cell absorbs two adjacent PDF rows.

When either fires, extract_tables_in_regions_mem runs over the same
crop bbox and its markdown replaces the TSR markdown.
fallback_reason carries the diagnostic label so callers can emit
metrics and watch each pathology independently.

Validated on FNBO branches PDF page 1 (Kansas region):
- TSR-only: merges Shawnee into BVP, wrong census tract on Sonoma.
- Auto fallback (phantom_empty_row): each row separate, correct tracts.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-27 10:13:25 -07:00
Abimael Martell bdea4f345a Bump version from 1.6.4 to 1.7.0 2026-04-27 10:01:30 -07:00
Abimael Martell 8a0f98dee7 repair malformed PDF containers (#65) 2026-04-27 09:38:25 -07:00
Abimael MartellandClaude Opus 4.7 d8894326e8 Stage 1 exclusive item-to-cell assignment, v1.6.4 (#66)
Closes the row-merge pattern observed against FNBO branch list at the
College/Fairway boundary: when SLANet emits cells whose y-extents
overlap between consecutive rows, an item whose center fell in the
overlap region got pulled into BOTH cells, producing run-on cells
(e.g. "Kansas Kansas" + concatenated addresses).

Cause: stage 1's "for each cell, gather items inside" loop allowed an
item to match multiple cells. `normalize_cell_bands` reduces overlap
but is biased when cell-mean-center is offset from actual text baseline
(SLANet bboxes are typically taller than their text content), so the
midpoint-clamp can land on the wrong side of the row boundary, and
items at the boundary still match two cells.

Fix: invert the matching. For each PDF text item, find candidate cells
(those whose bbox satisfies tsr_region_contains_item — center inside
OR >=60% overlap on both axes), and assign to the cell whose CENTER
is geometrically closest. Build per-cell text from the assigned items.
Stage 2 (orphan recovery) is unchanged.

Exclusivity prevents item duplication across cells. The closest-center
rule disambiguates the cell-overlap case naturally without aggressive
band clamping. normalize_cell_bands stays — it tightens cells before
matching (smaller overlap → fewer ambiguous candidates) but is no
longer load-bearing for correctness of the overlap case.

Local replay against FNBO via api/scripts/local-tsr-replay.ts (which
exercises the full layout-pod → table-pod → pdf-inspector chain
in-process):

  Pre-1.6.4 (deployed 1.6.3):
    |LITH West|Illinois Illinois|11700 S. IL Route 47, Huntley IL...
    |College|||0534.03|
    |Fairway|Kansas Kansas|4650 College Blvd... 2828 Shawnee Mission...

  Post-1.6.4 (this change):
    |Huntley|Illinois|11700 S. IL Route 47, Huntley IL...|8711.15|
    |LITH West|Illinois|4520 W Algonquin Rd, Lake in the Hills...
    |College|Kansas|4650 College Blvd, Overland Park KS...|0532.01|
    |Fairway|Kansas|2828 Shawnee Mission Pkwy, Fairway KS...|0500.00|

One row (Shawnee in the Kansas section) can still get dropped under
SLANet detection variance — the model occasionally under-detects rows
and emits N structure rows for N+1 PDF rows. The squeezed row's text
gets routed to the structurally-nearest existing cell. This is a
SLANet limitation, not addressable in pdf-inspector without
synthesizing rows from PDF text geometry; deferred.

Tests: 416 lib + 120 integration + 2 doctests pass. clippy + fmt clean.

Bump @firecrawl/pdf-inspector to 1.6.4 (patch — refines stage 1's
matching strategy; no API changes).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-27 09:22:58 -07:00
Abimael MartellandClaude Opus 4.7 9cce4dd161 TSR stage 2: reject cross-line orphan stacking, v1.6.3 (#64)
Closes the residual run-on-cell pattern observed on FNBO branch list
after 1.6.2 deployed:

  |Shawnee Blue Valley Parkway|Kansas|6301 Pflumm, ...|0523.04|
  |Sonoma Plaza|Kansas Kansas|<addr1> <addr2>|0531.05|
  |Mitchell Woonsocket|South Dakota|<addr1>|9628.01|

Cause: when col-N cells across multiple consecutive rows are y-shifted
the same way (a local SLANet drift), the stage 2 orphan pass sees
multiple orphans qualifying for the same nearest empty cell. The cell
gets all of them appended in order, producing "RowA-text RowB-text".

Fix: track the y-coordinate of the first orphan that lands in each
cell. Subsequent orphans only join that cell if their y is within
half-a-row-height of the first orphan's y (same line). Cross-line
orphans skip that cell and look for the next-nearest empty cell on
their own line.

Same-line slack preserves multi-token branch names like
"Blue Valley Parkway" (3 PDF text items at the same y) — all three
stack into the same cell. Cross-row stacking is what gets rejected.

Two new tests:

  - stage2_rejects_cross_line_stacking_into_same_cell
    Two orphans on different rows, both equidistant to the same empty
    cell. First wins; second routes to its own row's cell.

  - stage2_allows_same_line_orphans_to_stack_into_one_cell
    Three same-line orphans (multi-token branch name) all land in the
    same empty cell, joined by spaces.

The existing four stage-2 tests + the cell-bleed regression test from
PR #62 + #63 all pass: 416 lib + 115 integration + 2 doctests.
clippy + fmt clean.

Bump @firecrawl/pdf-inspector to 1.6.3 (patch — refines 1.6.2; no API
changes).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-27 08:11:17 -07:00
Abimael MartellandClaude Opus 4.7 f61d139710 Recover orphan text after band normalization (TSR stage 2), v1.6.2 (#63)
PR #62 (1.6.1) closed the cell-bleed regression by clamping SLANet's
loose cell bboxes into non-overlapping row/column bands and tightening
text membership to center-containment OR >=60% overlap. That worked,
but exposed the opposite failure: legitimate native PDF text whose
center fell just outside the *clamped* cell bbox now had nowhere to go.

Two distinct failure modes observed against the FNBO branch-list PDF
after 1.6.1 deployed:

  Symptom A — header text positioned at the LEFT of a column whose band
  was derived from data-cell centers farther right. Header "Address"
  PDF text at x=331..375 fell outside the clamped col band starting
  at x=410. Strict membership rejected it (0% x-overlap, center
  outside).

  Symptom B — local SLANet row drift in col 0 over a 5-row stretch.
  Cell bboxes sat just above the actual branch-name text items
  (x-overlap 100% but y-overlap ~30-43%, below the 60% threshold).

Both share one root: post-normalization bboxes are too tight, and the
strict rule has no escape valve for legitimate edge text.

Fix: add a stage-2 orphan-recovery pass after the strict fill. Items
that NO cell claimed in stage 1 get re-assigned to their nearest
*empty* cell, distance-capped by `(median_col_width, median_row_height)`
so a far-orphan figure title can't get pulled into a faraway empty
cell. Stage 2 only fills empties — never overwrites stage 1 — so the
cell-bleed case PR #62 closed cannot regress.

Three new lib tests cover the bug shapes:
  - stage2_recovers_left_aligned_header_text_outside_data_band
    (Symptom A: header text left-of-band, data cells already filled,
     stage 2 fills only the header)
  - stage2_recovers_y_shifted_col0_in_consecutive_rows
    (Symptom B: 3 col-0 cells shifted vs text, all 3 recovered)
  - stage2_does_not_overwrite_filled_cells_or_admit_far_orphans
    (cap rejects far figure titles; pre-filled cells untouched)

Plus tests for the cap-derivation helper:
  - tsr_assignment_caps_uses_median_geometry
  - tsr_assignment_caps_floor_protects_degenerate_input

The existing dense-overlapping-rows regression test (added in #62) still
passes, confirming no regression on the cell-bleed case.

Test results: 414 lib + 115 integration + 2 doctests pass. clippy + fmt
clean.

Bump @firecrawl/pdf-inspector to 1.6.2 (patch — fixes a regression
introduced by 1.6.1; no API changes).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 23:36:43 -07:00
Abimael MartellandClaude Opus 4.7 3f8fb645c9 Fix TSR cell assignment for overlapping table boxes (#62)
* feat: TSR-aware table extraction (extract_tables_with_structure_mem)

New public function that consumes raw structure-recovery output (HTML
structure tokens + per-cell bboxes from a model like SLANet) and assembles
markdown tables by pulling cell text from the native PDF — no OCR, no
geometry inference.

Why: the existing extract_tables_in_regions_mem infers grid geometry from
text positions only and can't distinguish merged cells from multiple narrow
columns. Pairing structure recovery from a layout/TSR model with native
PDF text gets perfect text quality with proper row/col/span structure.

- New module src/tables/structured.rs: token state machine, polygon→AABB,
  crop-px→page-pt, rowspan/colspan-aware cell layout, markdown emitter.
  Accepts both 4-element rects and 8-element 4-corner polygons.
- New public extract_tables_with_structure_mem in src/lib.rs that reuses
  extract_page_text_items, region_overlaps_item, and the shared region
  text-collection helper. No existing public function modified.
- napi binding extractTablesWithStructure mirroring the existing
  extractTablesInRegions shape (f64 in JS → f32 internally).
- 14 unit tests + 5 integration tests, including a real-PDF gold-standard
  match against bits_pilani_feedback.pdf.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* TSR follow-ups: header-aware separator, cells API, v1.6.0

- cells_to_markdown emits the separator after the LAST row that contains
  is_header=true cells, falling back to "after row 0" when no header is
  flagged. Multi-row theads now render correctly. Three new unit tests
  cover: multi-row header, header not on row 0, no headers (fallback).
- New public extract_tables_with_structure_cells_mem returning
  Vec<Vec<StructuredCell>> so callers can drive their own rendering or
  debug overlays without re-doing the parse + extraction. The markdown
  variant now wraps it. The previously-unused page_pt_bbox field is
  surfaced through this API.
- New napi binding extractTablesWithStructureCells + StructuredCellJs.
- Bump @firecrawl/pdf-inspector to 1.6.0.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix TSR cell text assignment for overlapping bboxes

Made-with: Cursor

* bump npm package version to 1.6.1

Made-with: Cursor

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 16:35:19 -07:00
Abimael MartellandClaude Opus 4.7 f6d5e214f1 feat: TSR-aware table extraction (extract_tables_with_structure_mem) (#61)
* feat: TSR-aware table extraction (extract_tables_with_structure_mem)

New public function that consumes raw structure-recovery output (HTML
structure tokens + per-cell bboxes from a model like SLANet) and assembles
markdown tables by pulling cell text from the native PDF — no OCR, no
geometry inference.

Why: the existing extract_tables_in_regions_mem infers grid geometry from
text positions only and can't distinguish merged cells from multiple narrow
columns. Pairing structure recovery from a layout/TSR model with native
PDF text gets perfect text quality with proper row/col/span structure.

- New module src/tables/structured.rs: token state machine, polygon→AABB,
  crop-px→page-pt, rowspan/colspan-aware cell layout, markdown emitter.
  Accepts both 4-element rects and 8-element 4-corner polygons.
- New public extract_tables_with_structure_mem in src/lib.rs that reuses
  extract_page_text_items, region_overlaps_item, and the shared region
  text-collection helper. No existing public function modified.
- napi binding extractTablesWithStructure mirroring the existing
  extractTablesInRegions shape (f64 in JS → f32 internally).
- 14 unit tests + 5 integration tests, including a real-PDF gold-standard
  match against bits_pilani_feedback.pdf.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* TSR follow-ups: header-aware separator, cells API, v1.6.0

- cells_to_markdown emits the separator after the LAST row that contains
  is_header=true cells, falling back to "after row 0" when no header is
  flagged. Multi-row theads now render correctly. Three new unit tests
  cover: multi-row header, header not on row 0, no headers (fallback).
- New public extract_tables_with_structure_cells_mem returning
  Vec<Vec<StructuredCell>> so callers can drive their own rendering or
  debug overlays without re-doing the parse + extraction. The markdown
  variant now wraps it. The previously-unused page_pt_bbox field is
  surfaced through this API.
- New napi binding extractTablesWithStructureCells + StructuredCellJs.
- Bump @firecrawl/pdf-inspector to 1.6.0.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 00:55:39 -07:00
Abimael Martell 5c4c6e8d33 Table improvements (v1.5.0) 2026-04-22 09:52:15 -07:00
Abimael Martell d7fb493a61 fix tagged table header recovery (#60) 2026-04-22 09:51:15 -07:00
Abimael MartellandClaude Opus 4.7 6819852541 fix: reject cell-rect "tables" that are actually prose in a framed box (#58)
The rect-based cell fallback in detect_row_stripe_table_from_cell_rects
derives columns purely from text X-position clustering. When prose
wraps inside a bounding-box rect (chat transcripts, stylized figures),
the word-boundary gaps cluster into many spurious columns, producing
a multi-column "table" that is just fragmented prose.

Count cells containing common English function words (articles,
prepositions, pronouns, common verbs) and reject the fallback when
20%+ of non-empty cells contain any such word. Real tabular data —
labels, units, numbers, short identifiers — rarely contains these.

Update the td9264 snapshot: the government document section that
previously rendered as a malformed table now renders as cleaner
prose + a proper CFR list.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 10:23:28 -07:00
Abimael MartellandClaude Opus 4.7 2876fa4b3e fix: tighten rect-row span check in propagate_merged_cells (#57)
* fix: don't reclassify wrapped bold list leads as headings

When a numbered/bulleted list item's bold lead phrase wraps onto a
second visual line, that line is all_bold + standalone, which scored
above the rarity heading threshold and was emitted as #### in the
middle of the item. That reset in_list, so the body continuation
below picked up a stray `- ` bullet via the struct-tree LI path,
shattering a single item into heading + stray bullets.

Guard the font heuristic: when already inside a list, skip heading
classification for lines at the list continuation indent with a Y
gap within para_threshold. Structure-tree headings still win.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: tighten rect-row span check in propagate_merged_cells

propagate_merged_cells used an overlap-based predicate with ±tol
slop that returned true at shared row boundaries — a rect whose
top exactly equals row N's bottom lies entirely below the row, yet
the predicate considered it to span row N. When multiple background
rects aligned on a shared Y edge (e.g. consecutive row-stripe
shading), each adjacent rect would over-reach by one row, cascading
labels and data from unrelated rows into a single merged cell.

Replace the overlap predicate with a containment check: rect bottom
at or below row bottom, rect top at or above row top (each within
tol). Genuine merged-cell rects fully contain the rows they span;
tangent rects do not.

Update two snapshots that were encoding the old buggy output.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 09:33:43 -07:00
Abimael MartellandClaude Opus 4.7 5859ed651b fix: don't reclassify wrapped bold list leads as headings (#56)
When a numbered/bulleted list item's bold lead phrase wraps onto a
second visual line, that line is all_bold + standalone, which scored
above the rarity heading threshold and was emitted as #### in the
middle of the item. That reset in_list, so the body continuation
below picked up a stray `- ` bullet via the struct-tree LI path,
shattering a single item into heading + stray bullets.

Guard the font heuristic: when already inside a list, skip heading
classification for lines at the list continuation indent with a Y
gap within para_threshold. Structure-tree headings still win.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 08:47:47 -07:00
Abimael Martell df023474f1 update eval instructions 2026-04-20 22:50:57 -07:00
Abimael MartellandClaude Opus 4.7 f44509ac3a fix: don't treat bullet-marker columns as real text columns (#55)
On PDFs where every list item starts with ● at the left margin and content
at a fixed offset, histogram column detection sees the gap between marker
and content as a gutter and splits each line across two phantom "columns,"
scrambling the reading order (Anthropic's Mythos system card p.73–74).

- layout: reject gutter candidates where the smaller side is ≥80%
  standalone bullet-marker glyphs (•, ●, ○, ◦, ▪, ▫, ◆, ◇, ■, □)
- markdown/classify: add starts_with_bullet_marker helper (narrower than
  is_list_item — excludes numbered/lettered patterns like 1. and a) so
  numbered section headings stay as headings)
- markdown/convert: skip heuristic heading detection on lines that start
  with a bullet marker
- markdown/classify: strip a leading bullet wrapped in a bold/italic run
  (e.g. "**● Label:**" → "- **Label:**") — some PDFs put the marker inside
  the same bold run as the label

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 22:47:34 -07:00
Abimael MartellandClaude Opus 4.7 2e933ba8c1 fix: don't split tagged-LI continuation lines into separate bullets (#54)
Some tagged PDFs use a flat style where each wrapped visual line of a list
item gets its own MCID tagged directly under /LI. The LI branch was
unconditionally prefixing every such line with "- ", turning continuation
lines into their own bullets. Only emit a new bullet when we're not already
inside a list; otherwise fall through to the existing continuation logic so
the text is appended to the previous item.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 21:59:53 -07:00
Abimael MartellandClaude Opus 4.7 4b5ae91f54 feat: expose per-page markdown extraction to Python and Node (#53)
* feat: expose per-page markdown extraction to Python and Node (#49)

Implements the feature requested in issue #49: a list-of-pages markdown output
from the Python API. Matching the existing project pattern, the feature lives
in the Rust core and is surfaced through every binding.

- Rust core: `extract_pages_markdown` (path) and `extract_pages_markdown_mem`
  (bytes) now take `Option<&[u32]>` — `None` returns every page in document
  order; a slice restricts and preserves caller order.
- Python: new `extract_pages_markdown(path, pages=None)` and
  `extract_pages_markdown_bytes(data, pages=None)` functions plus
  `PageMarkdown` / `PagesExtractionResult` classes; stub file updated.
- Node: `extractPagesMarkdown(buffer, pages?)` — `pages` is now optional.
- Tests: 2 new Rust integration tests, 9 new Python tests, 2 new Node
  assertions. All 372 unit + 107 integration + 53 Python tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump version from 1.3.0 to 1.4.0

Minor bump for the new per-page markdown extraction API exposed through
the Python and Node bindings.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 17:19:03 -07:00
Abimael MartellandClaude Opus 4.7 b4c34ba4b7 refactor: classify Table as Data vs Toc once at construction (#52)
The TOC/data-table distinction was being recomputed at every consumer:
- format.rs::table_to_markdown ran is_table_of_contents to decide between
  flat-list and markdown-table rendering.
- compute_layout_complexity ran it again to filter TOCs out of
  pages_with_tables.
- detect_heuristic validations used it to decide whether to relax val 1/9.

Each caller had to remember tables can be either kind, which leaks the
TOC concept across the codebase.

Add `TableKind { Data, Toc }` and a `Table::new` constructor that classifies
once from the cells. All five detectors (heuristic, rect, line, struct,
columns) now go through `Table::new`. Consumers match on `kind` instead of
re-running classification.

Pure refactor — no behavior change. Verified: pdf-evals output is byte-for-
byte identical (0 changed snapshots).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 14:04:50 -07:00
Abimael MartellandClaude Opus 4.7 462a7fdfde fix: exclude TOC pages from pages_with_tables metadata (#51)
* fix: detect hierarchical-indent TOCs as tables

Validation 1 (≥25% of rows have first col) and validation 9 (paragraph
content) were rejecting TOC pages where top-level chapters indent at the
leftmost X but subsections cascade further right. The leftmost column ends
up sparse (only chapter rows land there) and most cells are empty, but the
structure is still an unambiguous TOC.

Skip both validations when cells form a TOC pattern AND the table is narrow
(≤5 cols). The width cap preserves existing handling of wide multi-column
TOCs (e.g. back-of-book 2-up indices) where format_toc_as_list would mash
adjacent visual entries together.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: exclude TOC pages from pages_with_tables metadata

TOC detection routes through the table pipeline (detected as a table, then
formatted as a flat list with tab-aligned page numbers via
format_toc_as_list). That made TOC pages appear in LayoutComplexity's
pages_with_tables, which:

- Misleads downstream consumers that read this field as "this page has a
  data table".
- Trips the table-page guard in column detection, which switches to a
  different valley-detection threshold for table pages.

Add an is_table_of_contents check in compute_layout_complexity so TOC-shaped
detections don't count toward the table flag. The TOC still renders correctly
as a flat list — only the metadata classification changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 13:48:36 -07:00
Abimael MartellandClaude Opus 4.7 9abcbbb359 fix: detect hierarchical-indent TOCs as tables (#50)
Validation 1 (≥25% of rows have first col) and validation 9 (paragraph
content) were rejecting TOC pages where top-level chapters indent at the
leftmost X but subsections cascade further right. The leftmost column ends
up sparse (only chapter rows land there) and most cells are empty, but the
structure is still an unambiguous TOC.

Skip both validations when cells form a TOC pattern AND the table is narrow
(≤5 cols). The width cap preserves existing handling of wide multi-column
TOCs (e.g. back-of-book 2-up indices) where format_toc_as_list would mash
adjacent visual entries together.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 13:41:59 -07:00
Abimael Martell 0b3ba8784c Bump version from 1.2.0 to 1.3.0 2026-04-20 11:09:16 -07:00
Abimael MartellandClaude Opus 4.7 3cd9795fe1 fix: don't fire subset-GID remap when W array covers CMap CIDs (#47)
The subset-GID-mismatch heuristic tripped whenever the CIDFont's W array
started at CID 0 and the ToUnicode CMap's minimum source CID was > 2.
That's the normal sparse layout for a correctly-subsetted Identity-H
font with a .notdef at CID 0 and actual glyphs at high CIDs (Cyrillic,
Arabic, etc.). The spurious remap to sequential CIDs, combined with
score_text penalizing non-Latin letters as "other", caused the garbled
CMap to win over the correct primary — spaces turned into %, digits
shifted into punctuation, and Latin letters got rewritten to Cyrillic
codepoints.

Check whether the W array actually covers the CMap's maximum source CID.
If it does, the font and CMap agree, no renumbering happened, and the
remap stays off. True mismatches (CMap CIDs outside W coverage) still
trigger the remap.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-19 09:02:23 -07:00
Abimael MartellandClaude Opus 4.7 c0607ab152 feat(npm): add Windows binary to publish matrix (#46)
Adds x86_64-pc-windows-msvc target to the napi package and CI build
matrix so the published @firecrawl/pdf-inspector npm package ships
native Windows binaries alongside Linux and macOS.

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-17 13:58:17 -07:00
Abimael Martell cf2faa3221 Bump version to 1.1.0 in package.json 2026-04-17 11:59:14 -07:00
Abimael MartellandClaude Opus 4.7 e043c4414d fix: reject dot-less TOC as heuristic table false-positive (#45)
* fix: reject dot-less tables of contents as heuristic table false-positive

The body-font heuristic table detector was firing on tagged PDF TOCs
whose entries are laid out as multi-column rows (section number, title,
page number) without dot leaders. Extend is_table_of_contents to flag
this pattern by requiring:
- 2+ columns and 4+ rows
- >=60% of rows start with a dotted section number (e.g. "4.3.1")
- >=70% of filled last-column cells are all-digit page numbers

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* format TOC tables as flat per-row list instead of dropping them

Keeping the previous approach (rejecting TOCs at detection time) meant
dot-less TOCs fell back to the column-aware text reader, which stacked
every section title first and dumped all page numbers at the end of the
region — so a reader saw a wall of titles followed by a wall of numbers.

Keep the detected Table, and in the formatter interleave the cells into
one line per row with the page number appended after a tab.  The raw
(pre-clean) cells are used for the TOC check because clean_table_cells
can collapse genuine data tables into a TOC-looking shape.

Tightened is_table_of_contents to require both dot leaders AND page
numbers (previously dot-ratio alone was enough, which misfired on
register tables that use "..." as a continuation marker).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* split TOC detection into dot-leader and tabular patterns

Previous is_table_of_contents rejected every TOC-shaped table at detect
time, dropping dot-leader TOCs that render well as flat-list output and
misclassifying data tables with trailing "..." cells as TOCs.

Now is_dot_leader_toc (structural + inline-leader) keeps per-row flat
layouts for the formatter, while only wide inline-leader indices are
rejected at detect time.  row_cell_is_page_number accepts dashed
section-page IDs ("A-1", "5-21") and rejects decimals/thousands
separators; cell_has_trailing_leader requires alphabetic content so
numeric data-row labels ("1973 ... ") no longer qualify.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-17 11:58:41 -07:00
Abimael Martell 6466e59271 update node pacakge paths in docs (add agents file too) 2026-04-17 08:33:37 -07:00
Abimael Martell adcaaa1b3a npm package under firecrawl organization 2026-04-17 08:31:18 -07:00
Abimael MartellandClaude Opus 4.6 b17e8e3443 feat: add pdf-inspector CLI to npm package (#44)
* feat: add CLI bin to npm package

Installing `firecrawl-pdf-inspector` now provides a `pdf-inspector` CLI command.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: bump npm package version to 0.8.0

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: bump npm package version to 1.0.0

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-17 07:23:08 -07:00
Abimael Martell c409e3b8ca fix: use first-glyph position for ActualText ligature items (#43)
* fix: use first-glyph position for ActualText ligature items

The content stream parser emitted ActualText items (from BDC/Span
marked content) at the text matrix position captured at BDC entry.
When the content stream contains Td operators between BDC and the
first glyph Tj — common in PDFs generated by Google Docs, Figma,
and web-to-PDF tools — the BDC-entry position is on the previous
visual line while the actual glyph renders on the correct line.

This caused ligature glyphs (fi, fl, ff, etc.) wrapped in
ActualText spans to be positioned one text-leading offset above
their surrounding text. Downstream merge logic couldn't reconnect
them (different y-band), producing broken words like "fi" + "ndings"
instead of "findings" throughout the document.

Fix: capture the text matrix at the first Tj/TJ operation inside
the BDC block (after any Td repositioning), and use that as the
ActualText item's rendering position at EMC time. Falls back to the
BDC-entry position when no glyph ops occur inside the block.

One new variable, five insertion points, zero behavior change on
PDFs that don't use ActualText marked content.

* chore: allow clippy::collapsible_match for Rust 1.95

Rust 1.95 introduced the collapsible_match lint which flags `if`
blocks inside match arms that could be converted to match guards.
The content-stream parsers use this pattern extensively for
readability (match on PDF operator name, then check preconditions
like `in_text_block && !op.operands.is_empty()`). Allow crate-wide
rather than refactoring 21 match arms across the parser files.
2026-04-16 17:14:00 -07:00
Abimael Martell 704ca588f5 Update package name in README for pdf inspector 2026-04-15 15:43:45 -07:00
Abimael Martell 6f4a523a06 fix: correct classifier false flags for CID-encoded text and supplementary fonts (0.7.4) (#41)
* fix(detector): correct false flags for CID-encoded text and supplementary fonts (0.7.4)

The page classifier was over-aggressively flagging Mixed-PDF pages as
needing OCR in three distinct cases. Each is fixed at the root in
analyze_page_content / page_has_identity_h_no_tounicode / the
looks_like_scan check.

1. has_vector_text false positives on dense layouts
   path_ops > text_ops*200 fired on pages with decorative paths
   (column borders, dividers) alongside real selectable text. Added
   a unique_alphanum_chars < 30 guard: real outlined-text pages have
   very few unique alphanum chars (each glyph is a path), while
   pages with real text + decorations have many.

2. Identity-H without ToUnicode flagged whole pages on supplementary fonts
   page_has_identity_h_no_tounicode would flag a page if any single
   Type0 font lacked ToUnicode and had no fallback CMap, even when
   the page's actual text came from other decodable fonts (Type1
   with ToUnicode, etc.). Rewrote to track both undecodable
   Identity-H fonts AND other decodable fonts, only flagging when
   no decodable text font is present.

3. CID-encoded text with ToUnicode misclassified as scan
   looks_like_scan checked unique_alphanum_chars < 10 on raw string
   operand bytes. CID-encoded fonts (Type0 with ToUnicode) emit
   2-byte CID values that aren't ASCII alphanum, so the metric is
   blind to them even when the text is fully decodable. Added a
   has_decodable_text_fonts signal: when a page has decodable fonts
   AND >= 10 text ops, the low alphanum count is treated as a CID
   encoding artifact rather than evidence of a scan.

Validated against a broad PDF corpus:
- 6 known false-positive pages now correctly classified as text
- 22 previously-missed scan pages (cover/blank/photo) now correctly
  flagged for OCR
- 0 regressions on truly-scanned PDFs (61/61 pages stay flagged)
- All 437 existing tests pass; clippy clean

Bumps NAPI package to 0.7.4.

* test(detector): add unit tests for the three classifier fixes

Adds 10 unit tests covering the heuristic changes:

- has_vector_text alphanum guard
  - real text + decorative paths → not flagged
  - true outlined glyphs (low alphanum) → still flagged

- page_has_identity_h_no_tounicode supplementary-font handling
  - undecodable Identity-H + decodable Type1 → not flagged (new)
  - undecodable Identity-H alone → still flagged (regression)

- page_has_decodable_text_fonts (new helper)
  - Type1 → true
  - Type0 with ToUnicode → true
  - undecodable Identity-H only → false

- looks_like_scan with has_decodable_text_fonts override
  - CID-encoded decodable text → not flagged as scan
  - same metrics with no decodable fonts → still flagged
  - decodable fonts but text_ops < 10 (page-number overlay) → still flagged

* fix(detector): make decodable-font checks usage-based and XObject-aware

Addresses two reviewer concerns on the previous heuristic fix:

P1 — resource-based check could create an inverse bug
  page_has_identity_h_no_tounicode and page_has_decodable_text_fonts
  iterated all fonts in the page Resources dict, including unused fonts.
  A page whose actual text was rendered exclusively in an undecodable
  Identity-H font but whose Resources also listed an unused decodable
  Type1 would be wrongly unflagged.

  Fix: parse Tf operator operands during content stream scanning to
  collect the set of font names actually referenced. The font checks
  now filter to only USED fonts via a new used_fonts_have_*
  family of functions operating on (used_font_names, font_map).

P2 — checks didn't follow text into Form XObjects
  analyze_page_content correctly recurses through Form XObjects via
  scan_xobjects_in_resources, but the font checks only looked at the
  page's top-level Resources/Font. Pages that render text through Form
  XObjects (corporate templates, header/footer overlays) had their
  XObject font resources missed entirely.

  Fix: scan_xobjects_in_resources now propagates the used_font_names
  set AND collects fonts from each Form XObject's own Resources into
  the shared font_map. The usage-based check sees the full picture:
  page-level fonts + every nested XObject's fonts, intersected with
  fonts actually referenced by Tf operators anywhere in the content.

Implementation:
- New extract_font_name_before_tf helper (parses /Name immediately
  preceding Tf).
- New FontInfo struct caches font properties per-name.
- New collect_fonts_from_resource_dict + new used_fonts_have_*
  functions are pure filters over (used_names, font_map).
- analyze_page_content threads used_font_names + font_map through
  page content scan and XObject recursion, then runs the new checks.
- Old resource-based functions kept as #[cfg(test)] for the existing
  unit-test interface.
- Phase 3 uncached-page loop now goes through analyze_page_content
  so it also gets the usage-based + XObject-aware behavior.

Tests added (8):
  - extract_font_name_before_tf basic + long-name parsing
  - scan_content_for_text_operators collects used font names
  - P1 — unused decodable font in Resources doesn't save a page
    whose used font is undecodable
  - P1 — both fonts used → decodable font correctly prevents flag
  - P2 — decodable font inside Form XObject correctly unflags
  - P2 — undecodable font only in XObject still flags even with
    unused decodable font at page level
  - P2 — has_decodable_text_fonts populated from XObject fonts

Validation:
- 349 lib + 104 integration + 2 doc tests pass (was 341)
- cargo clippy --lib --bin detect-pdf -- -D warnings: clean
- External eval: 9/9 PDFs pass, 6/6 false positives resolved,
  0 regressions, 61/61 scanned pages still correctly flagged
- No eval delta — confirms previous fix wasn't relying on the
  resource-based bug for any of the eval PDFs

* fix(detector): scope font lookups by ObjectId + handle indirect Form Resources

Addresses two more reviewer findings on the previous decodable-font commit.

P1 — Resource-name scoping bug
  The previous fix keyed used_font_names and font_map by raw resource
  names like b"F1". PDF resource names are scoped to each resource
  dictionary: a Form XObject can legally define its own /F1 that points
  to a completely different font from the page's /F1. Because
  collect_fonts_from_resource_dict skipped duplicates with
  `if font_map.contains_key(name)`, the first definition won and later
  Tf /F1 usages in different scopes resolved against the wrong font.
  This could reintroduce both the undecodable-Identity-H false flag
  and the decodable-CID false unflag depending on which side of the
  collision happened to be inserted first.

  Fix: switch the lookup mechanism from font names to font ObjectIds.
    - font_map: HashMap<ObjectId, FontInfo>  (was Vec<u8> keys)
    - used_font_ids: HashSet<ObjectId>       (was Vec<u8> names)
    - new resolve_font_names_to_ids() runs immediately after each
      content scan, against the resource dict in scope, to translate
      the per-scope name set into ObjectIds.
  Each Form XObject's content stream now resolves /F1 against THAT
  XObject's own Resources, so name collisions are impossible by design.
  Inline (no-ID) font dicts are skipped — extremely rare in practice
  and have no stable key.

P2 — Indirect Form /Resources skipped
  scan_xobjects_in_resources used `.as_dict()` on the Form's /Resources
  entry, which returns None for indirect references. PDFs frequently
  store /Resources as `X 0 R`, in which case font collection and
  recursion were both skipped — even though the Tf usages inside the
  XObject content had already been recorded.

  Fix: handle Object::Reference(r) in addition to Object::Dictionary(d)
  by resolving via doc.get_dictionary. Audited the rest of the file —
  the other /Resources access points (analyze_page_images,
  collect_images_from_resources) already handled both cases.

Tests added (4):
  - P1 same-name-different-font (page undecodable, XObject decodable):
    must NOT flag — XObject's text is decodable in its own scope.
  - P1 inverse (page decodable, XObject undecodable, content uses
    XObject /F1): MUST flag — undecodable text exists in real scope.
  - P2 indirect Form /Resources: font discovery must still work when
    /Resources is a `X 0 R` reference rather than inline.
  - Combined regression: indirect Resources + name collision.

Validation:
  - cargo test --release: 459 tests pass (353 lib + 104 integration + 2 doc)
  - cargo clippy --lib --bin detect-pdf -- -D warnings: clean
  - external eval (9 PDFs): 9/9 pass, 6/6 false positives resolved,
    0 regressions, 61/61 truly-scanned pages still flagged

The behavior on the eval set is identical — confirms the correctness
fix isn't masking any change in classifier outcomes.

* fix(detector): respect resource shadowing when resolving page-content fonts

The previous ObjectId-based fix correctly scoped Form XObject fonts
but still violated PDF resource inheritance for page content. When a
page overrides /F1 from a parent /Pages node (different font dict for
the same name), get_page_resources returns the page's own /Resources
plus all ancestor /Resources dicts. The old code called
resolve_font_names_to_ids on each one and added every match to
used_font_ids — both font ObjectIds ended up in the used set even
though only the page's /F1 is actually visible to that page's content.

Per ISO 32000-1 §7.7.3.4, resource names are inherited with
shadowing semantics: the most-specific (deepest, closest to the page)
definition wins.

Fix:
- New lookup_font_id helper resolves a single name in a single dict.
- New resolve_with_shadowing iterates names, checking the page's own
  /Resources first, then walking ancestors in most-specific-first
  order (which is the order lopdf's get_page_resources returns).
  First hit wins via a labeled `continue 'name` — subsequent
  ancestors are skipped for that name.
- analyze_page_content's flat resolution loop replaced with one call
  to resolve_with_shadowing.

Audit:
- XObject path is correct: each Form XObject already resolves names
  against its OWN /Resources (XObjects don't inherit from page tree).
- font_map population is correct: keyed by ObjectId, so collecting
  from all dicts builds the full available-fonts catalog. The bug
  was only in the used-set resolution.
- Confirmed lopdf returns ancestors in most-specific-first order
  (page → parent → grandparent → root), matching the shadowing
  direction used here.

Tests added (3):
  - page /F1 undecodable shadows parent's decodable /F1 → MUST flag
  - page /F1 decodable shadows parent's undecodable /F1 → MUST NOT flag
  - no override: page inherits parent's decodable /F1 → MUST NOT flag

Validation:
  - cargo test --release: 462 tests pass (356 lib + 104 integration + 2 doc)
  - cargo clippy --lib --bin detect-pdf -- -D warnings: clean
  - external eval: 9/9, 0 regressions, 6/6 false positives resolved,
    61/61 scanned pages still correctly flagged
2026-04-15 13:47:06 -07:00
Abimael MartellandClaude Opus 4.6 2f23f07f6e fix: use AND logic for looks_like_scan heuristic in detector (#39)
The looks_like_scan check incorrectly used OR logic, causing any single
condition (image_count <= 1, text_ops < 50, alphanum < 10) to flag a page
as a scan. A real scan has ALL three: single full-page image AND low text
AND low alphanum. Text pages with one figure were falsely flagged for OCR.

Bump napi to 0.7.3.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 22:05:22 -07:00
Abimael MartellandClaude Opus 4.6 0a9c120a6b fix: reduce false OCR recommendations for text PDFs with figures (#38)
* fix: reduce false OCR recommendations for text PDFs with figure images

Two fixes in the detector:

1. Fix Tf operator parsing: some PDFs concatenate Tf directly with the
   next operator (e.g. "25 Tf[<01>...") without whitespace. The scanner
   now accepts [, (, <, / as valid followers, fixing font_changes being
   reported as 0.

2. Distinguish text-with-figures from scanned-with-OCR: pages with
   multiple images (image_count > 1) and strong text signals (text_ops
   >= 50, alphanum >= 10) are recognized as text pages with figures,
   not scanned templates. Scanned PDFs have exactly 1 full-page image.

This prevents academic papers, reports with charts, and similar PDFs
from being incorrectly classified as Mixed/OCR-needed when their text
is perfectly extractable.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: bump napi version to 0.7.2

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: remove template image influence from page classification

Template images (large background/figure images) no longer affect
pages_needing_ocr. In the region-based pipeline, text regions are
extracted independently from image regions, and per-region needs_ocr
quality checks handle scanned-with-OCR garbage text.

Also makes the invisible text retry (for OCR text layers) trigger on
text quality rather than PDF type, so it works regardless of
classification.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Revert "fix: remove template image influence from page classification"

This reverts commit 100cbe5453.

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 21:42:19 -07:00
Abimael MartellandClaude Opus 4.6 7c8b09be67 fix: improve table detection for numeric columns and multi-line headers (#35)
* fix: improve heuristic table detection for numeric columns and multi-line headers

Two fixes for tables that have clean extractable text but fail heuristic
structure detection:

1. Numeric column merge pass (grid.rs): After initial X-position
   clustering, adjacent clusters are merged when one is sparse (header
   text) and the other is dense with >50% numeric items (data column).
   Multi-line wrapped headers often land slightly offset from their
   data column — the merge closes gaps within 1.5× the clustering
   threshold. New is_numeric_text() helper matches decimals, percentages,
   negative numbers, and comma-separated thousands.

2. Duplicate-header skip (detect_heuristic.rs): Spanning super-headers
   like "First Degree | First Degree | Higher Degree" contain duplicate
   cells that trigger looks_like_partial_table_ex rejection. Now skips
   rows with duplicate cells when a better header candidate exists
   within the next 3 rows (higher fill ratio or numeric cells).

Tested on BITS Pilani university report (430 pages, 314 table pages).
Page 4 (multi-line header + numeric data) previously returned
needs_ocr=true; now correctly detects the table structure.

Eval: 197 PDFs, zero regressions, all 104+ tests pass, zero clippy.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* bump version to 0.7.1

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 17:27:59 -07:00
Abimael MartellandClaude Opus 4.6 35445c3208 Auto-publish npm package when version changes in package.json (#33)
Replace tag-based trigger with push-to-main trigger that detects
version changes in napi/package.json, removing the need for manual
git tags to publish new npm releases.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 13:12:55 -07:00
Abimael MartellandClaude Opus 4.6 20f24d1f8d extractPagesMarkdown: return classification metadata (0.7.0) (#32)
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Combine per-page markdown extraction with layout classification into a
single parse. extractPagesMarkdown now returns PagesExtractionResult with
pages_with_tables, pages_with_columns, pages_needing_ocr, and is_complex
alongside the per-page markdown — eliminating redundant PDF parses for
callers that need both.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 13:08:07 -07:00
Abimael MartellandClaude Opus 4.6 abb0b925fb Add extractPagesMarkdown for per-page markdown extraction (#31)
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
* add extract_pages_markdown_mem for per-page markdown extraction

Enables hybrid OCR pipelines to skip GPU render+layout for simple text
pages by providing per-page markdown with needs_ocr flags. Font stats
are computed document-wide for consistent header detection.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* bump napi package version to 0.6.0

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 12:37:32 -07:00
Abimael MartellandClaude Opus 4.6 00c5c18e2a napi: use string enums for PdfType and ItemType (0.5.0) (#29)
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Replace stringly-typed pdf_type and item_type fields with
#[napi(string_enum)] enums for proper TypeScript type checking.
Add link_url field to TextItem instead of encoding URL in the
item_type string.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 11:25:26 -07:00
Abimael MartellandClaude Opus 4.6 5159abe9c2 fix clippy warnings: prefix unused page_has_gid, cfg(test) wrapper
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 19:53:28 -07:00
Abimael MartellandClaude Opus 4.6 843a745460 relax table extraction validation for layout-assisted regions (0.4.3)
Two changes that reduce false needsOcr rejections without hurting quality:

1. Per-region GID check instead of per-page blanket rejection.
   Previously, if ANY font on the page used GID-encoded glyphs (common
   in logos, decorative fonts), ALL table and text regions on that page
   were forced to GPU OCR via needsOcr=true. Now the page-level bail is
   removed; per-region text quality checks (is_garbage_text, is_cid_garbage,
   detect_encoding_issues) catch actual GID corruption in the extracted
   content. Tables whose text is clean pass through even if an unrelated
   font elsewhere on the page is GID-encoded.

2. Relaxed looks_like_partial_table for layout-assisted extraction.
   When the layout model already identified a region as a table (i.e.,
   extract_tables_in_regions_mem), boundary-detection heuristics are
   less necessary — we're not guessing "is this a table?" anymore, only
   "can we extract it correctly?". Relaxations:
   - Numeric first header cell accepted (e.g., year "2024")
   - 1 empty header cell allowed in 3+ column tables (merged headers)
   - Sparse first data row threshold relaxed from 33% to 50%
   Paragraph detection and duplicate-header checks remain strict.

Eval: 196/196 pass (full regression suite), 91/91 Rust tests pass
including 7 new layout-assisted validation tests. Zero regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 19:46:56 -07:00
Abimael MartellandClaude Opus 4.6 780efdb955 extract_tables_in_regions: detect paragraph-as-table misreads (0.4.2)
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Adds a 5th failure-mode check to looks_like_partial_table: when the
heuristic mis-detects text-wrapped paragraph prose as a multi-column
table, cells in the same column tend to start with lowercase letters
or continuation punctuation (commas, closing quotes) — because they're
actually sentence fragments. Real tables almost never have most data
cells starting lowercase.

Trigger: ≥2 cols, ≥4 data rows, ≥60% of non-empty data cells start
with lowercase or continuation punctuation → return needs_ocr=true.

Caught in the eval as the next-largest failure mode after the 0.4.1 fix:
PDFs 088, 182, 090 — heuristic produced "tables" like:

  |Approval is needed from the|Acquisitions of|
  |Treasurer if the acquisition|residential and|
  |constitutes a "significant|agricultural|
  |action," including acquiring an|land by foreign|

Reading column 1 top-to-bottom: "Approval is needed from the Treasurer
if the acquisition constitutes a 'significant action,' including
acquiring an interest..." — a paragraph, not tabular data.

Tests: 2 new tests (the 088-style failure case + a real multi-word
table that must NOT be flagged). All 11 looks_like_partial_table tests
pass; 323 unit + 91 integration tests still green.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 15:09:08 -07:00
Abimael MartellandClaude Opus 4.6 d0dd067e70 extract_tables_in_regions: needs_ocr on suspicious table structure (0.4.1)
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
When the heuristic returns markdown that looks like a partial / mis-detected
table, set needs_ocr=true so the caller falls back to GPU OCR. Previously the
same cases returned the broken table with needs_ocr=false, which produced
real-world TEDS=0 scores in fire-pdf evals (heuristic-built table didn't
match ground truth structure at all, but caller had no signal to fall back).

Four failure modes detected, all observed in opendataloader-bench eval losses:

1. **Header looks like a data row** — first cell of header is a bare number
   (e.g. `|2|...`), suggesting the actual header row was skipped. Real
   headers almost never start with just a number.

2. **Empty header cells in a multi-column table** — ≥3 cols, ≥1 empty cell
   in the header row. Indicates poor column boundary detection.

3. **Duplicate header cells** — same non-empty value appearing twice in the
   header (e.g. "Administration|Administration"). Means a multi-line header
   was collapsed wrong.

4. **Sparse first data row** — ≥3 cols and ≥1/3 of first-data-row cells are
   empty. Multi-row headers in the source PDF get smashed into header +
   sparse data row by the heuristic; this catches that.

Tests: 9 new unit tests in `looks_like_partial_table_tests` cover each
failure mode plus realistic non-failures (well-formed table, single-column
list, two-col with a single empty cell). All 91 existing tests still pass.

Bumps `napi/package.json` to 0.4.1 since this changes the function's return
behaviour for callers (some inputs that returned needs_ocr=false now return
true). The output text field is also cleared on the new fallback path so
callers don't accidentally use the broken markdown.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 14:10:44 -07:00
Abimael Martell 8282c2f8ee bump napi package
Publish npm package / Publish to npm (push) Has been cancelled
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
2026-04-13 12:43:28 -07:00
Abimael MartellandClaude Opus 4.6 8e8ab4a19d feat: add extractTablesInRegions NAPI binding for region-based table extraction (#27)
Adds a new function that takes a PDF buffer and page+bbox regions (same interface
as extractTextInRegions), runs heuristic table detection on items within each region,
and returns markdown pipe-tables. Falls back to needs_ocr=true when no table
structure is found or text quality is suspect.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 23:32:02 -07:00
Abimael MartellandClaude Opus 4.6 8e3084183c fix: suppress overused struct tree heading tags for better paragraph detection (#28)
Some PDFs (e.g. British Academy grant guidance, Carter BloodCare privacy
policy) have structure trees that incorrectly tag body text as H2 headings.
This caused every line within numbered paragraphs to render as a separate
## heading instead of being joined into flowing paragraph text.

Added detect_overused_struct_heading_levels() which pre-scans heading tag
frequency and suppresses levels appearing on >15% of tagged lines, allowing
those lines to fall through to normal paragraph joining.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 23:31:42 -07:00
Abimael MartellandClaude Opus 4.6 d8bb0f5898 chore: bump npm version to 0.3.6
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 23:16:20 -07:00
Abimael MartellandClaude Opus 4.6 1b4f1f4640 fix: reduce false OCR flags for Identity-H fonts with fallback decoding (#26)
The detector flagged pages for OCR whenever any font was Identity-H
without ToUnicode, even when the extraction pipeline could decode the
font via fallback paths (CID-as-Unicode passthrough or embedded TrueType
cmap). This caused false positives on PDFs from Chromium, wkhtmltopdf,
and other generators that use Identity-H with Unicode CID values.

Now checks DescendantFonts W array and embedded font cmap before
flagging. Fonts that are genuinely undecodable (stripped cmap, low GID
CIDs) are still correctly flagged.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 18:22:11 -07:00
Abimael MartellandClaude Opus 4.6 2455f1437b chore: bump npm version to 0.3.5
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 16:44:45 -07:00
Abimael MartellandClaude Opus 4.6 640cdaaa13 fix: stop counting Do operators as images in content stream scanner (#25)
Do invokes any XObject (Form or Image), but scan_content_for_text_operators
was counting every Do as an image. PDFs with Form XObjects (e.g. ACS
publisher watermark pages) were misclassified as ImageBased because the
inflated image_count raised the min text ops threshold above the actual
text operator count.

Image detection is already correctly handled by scan_xobjects_in_resources
(checks Subtype) and analyze_page_images (measures pixel area).

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 16:41:40 -07:00
Abimael MartellandClaude Opus 4.6 be313cdb81 fix: require strong signal for rarity-based heading detection (#24)
In multi-column PDFs, column switches break paragraph continuity,
making body text lines appear "standalone". Combined with moderate
font-size rarity from minor size variation between columns, this
caused hundreds of false heading classifications (e.g. 281 false ##
headings on a single academic paper).

Non-bold, non-isolated lines now require very high rarity (≥0.97)
and short word count (≤8) to qualify as headings via the rarity path.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 10:45:51 -07:00
Abimael MartellandClaude Opus 4.6 6a9ff170dc fix: simplify region text extraction to trust layout model ordering
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
When fire-pdf sends pre-segmented bboxes from the layout model,
pdf-inspector no longer runs column detection, stream-order heuristics,
or newspaper/tabular mode detection within the region. These heuristics
conflict with the layout model's decisions and cause wrong reading order.

Region extraction now simply: Y-sorts items, groups into lines, and
sorts within each line by X position. The heavy heuristics remain
available for standalone full-page extraction.

Eval showed pure OCR (0.2875 NED) beating native+heuristics (0.2916)
across all categories, especially multi-column (-0.08) and newspaper
(-0.16). This change should close that gap.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 21:58:05 -07:00
Abimael MartellandClaude Opus 4.6 0db9863919 feat: lookahead-based isolated line heading detection
Pre-scan lines to identify "isolated" ones — short lines (1-6 words)
with paragraph breaks both before AND after. These are heading
candidates even at body font size, common in academic papers
("Acknowledgements", "Limitations", "B.3 Prompt Engineering").

Inspired by opendataloader's HeadingProcessor which passes prevNode
and nextNode context to the heading probability scorer.

The isolated signal (+0.3) combines with rarity/bold/standalone
signals. A per-page density guard prevents false positives on
multi-column pages where many lines appear isolated. Continuation
word detection (ending in "the", "and", etc.) filters wrapped
paragraph lines.

MHS=0 docs: 18→13. MHS-S +0.004. No regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 12:32:31 -07:00
Abimael MartellandClaude Opus 4.6 14154ee5ee docs: update benchmark scores
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 12:19:38 -07:00
Abimael MartellandClaude Opus 4.6 10dd7e2881 feat: XY-cut fallback for column detection on asymmetric layouts
When the histogram-based column detector finds no valleys (common with
sidebar/asymmetric layouts), fall back to a simplified XY-cut: find the
largest horizontal gap between item edges and split there if both sides
have enough items with vertical overlap.

Inspired by opendataloader's XY-Cut++ algorithm but implemented as a
single-level fallback rather than full recursive segmentation.

Doc 156: NID 0.545→0.966, Doc 157: NID 0.564→0.962.
NID-S +0.007, TEDS-S +0.066 across 200 docs. No regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 12:11:45 -07:00
Abimael MartellandClaude Opus 4.6 cf7e6b895d Bump napi package version to 0.3.3
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 11:57:32 -07:00
Abimael MartellandClaude Opus 4.6 af34bbb6b5 feat: synthesize lines from thin filled rects as last-resort table detection
PDFs that draw table borders as thin filled rectangles (< 2pt, common in
spreadsheet exports) were invisible to both rect-based and line-based
table detectors. Now converts these thin rects to PdfLine objects and
runs line-based detection, but ONLY as a last resort after all other
methods (rect, line, heuristic, column-based) found nothing.

This avoids the regression from the earlier attempt which ran synthesis
at step 2, preempting the heuristic detector on PDFs where it worked
better.

Also relaxes uniform row spacing threshold (CV 0.05→0.02) to accept
spreadsheet-exported tables with even row heights.

Benchmark: TEDS 0.519→0.586 (+0.067), overall +0.006, no regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 11:44:11 -07:00
Abimael MartellandClaude Opus 4.6 dbc2de3e9b revert: undo table splitting changes that caused TEDS regression
Reverts commits 937311c, 1dcb0c6, 0300e96, 999f9a2. The thin-rect-to-line
synthesis and stacked table splitting improved extraction for specific
government PDFs but caused -0.05 TEDS regression on the benchmark by
preempting the heuristic detector with worse line-based grids.

These features need more targeted guards before re-enabling.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 11:20:05 -07:00
Abimael MartellandClaude Opus 4.6 999f9a2f2d feat: split stacked tables at rows without vertical borders
Rows that sit between horizontal rules but lack vertical border coverage
are not table cells — they're freestanding text (e.g. "Note: The cutoff
mark is out of 120"). These rows now split the grid into separate
sub-tables, with the unbounded text emitted as plain text between them.

Single-cell "tables" (from the split) render as plain text instead of
a degenerate 1x1 markdown table.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 10:24:28 -07:00
Abimael MartellandClaude Opus 4.6 0300e96b7d fix: don't merge column header rows as continuations
Rows with 3+ short-valued cells (avg ≤10 chars) and an empty first cell
are column headers (e.g. "UR | SC | ST | OBC | EWS"), not text overflow
from the previous row. Prevents them from being merged into the
preceding section title row.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 02:24:52 -07:00
Abimael MartellandClaude Opus 4.6 1dcb0c6feb fix: improve table row splitting for stacked sub-tables
Three fixes for better table extraction from spreadsheet-exported PDFs:

1. Convert thin filled rects (< 2pt) to PdfLine objects before line-based
   table detection. Many PDFs draw table borders as narrow filled rectangles
   instead of stroked paths — these were invisible to the line detector.

2. Relax uniform row spacing rejection (CV 0.05 → 0.02). Spreadsheet
   exports have very even row heights that were being rejected as "chart
   grids".

3. Fix continuation row merging: don't merge rows where the only non-first
   cell content is a long label (section headers like "Category No. 03").
   Don't merge first-cell-only rows with long text ("Note: ...").

Also adds multi-Y row splitting in line-based detection and column-aware
table detection skipping for multi-column pages.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 02:20:42 -07:00
Abimael MartellandClaude Opus 4.6 937311ce27 feat: convert thin filled rects to lines for table border detection
Many PDFs (especially spreadsheet exports) draw table borders as thin
filled rectangles (height/width < 2pt) instead of stroked paths. These
were invisible to our line-based table detector since only stroke
operations produced PdfLine objects.

Now synthesizes PdfLine from thin rects before line-based detection,
enabling table detection on border-drawn PDFs like government forms.

Also relaxes the uniform row spacing rejection threshold (CV 0.05→0.02)
to accept spreadsheet-exported tables with even row heights.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 02:12:05 -07:00
Abimael MartellandClaude Opus 4.6 9c463b9c95 fix: prevent multi-column text from being misdetected as tables
On pages where column detection finds 2+ columns, skip body-font
heuristic table detection in the merged-band retry path. This prevents
sidebar/two-column prose from being formatted as markdown tables.

The fix is targeted: per-band heuristic detection still runs (bands
are scoped to single columns), so real tables within columns are
still detected. Only the merged-band retry (which sees all items
across columns) is gated.

Also relaxes column validation to accept asymmetric layouts (sidebars)
where one side has fewer items, and tries center-based item assignment
before edge-based to improve column splitting for asymmetric layouts.

Benchmark: NID 0.865→0.869, NID-S 0.798→0.805, overall +0.002.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 02:01:24 -07:00
Abimael MartellandClaude Opus 4.6 1c48a014f7 docs: update benchmark MHS score after rarity-based detection
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 01:38:21 -07:00
Abimael MartellandClaude Opus 4.6 6e5abd0b48 feat: rarity-based heading detection inspired by opendataloader
Replace ad-hoc bold/ratio heading checks with a unified scoring system
based on font size rarity. For each line, compute:
  score = font_rarity * 0.5 + bold * 0.3 + standalone * 0.2

Font rarity measures how infrequently a font size appears across the
document — heading fonts are rare while body text is common. This
approach (from opendataloader's ModeWeightStatistics) naturally adapts
to each document's font distribution instead of relying on fixed
thresholds.

Guards: require font_size >= 0.95 * base_size (no small-font headings),
word_count >= 3, and standalone (paragraph break before).

Benchmark improvement: MHS 0.56→0.58, MHS-S 0.66→0.70, overall +0.003.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 01:38:03 -07:00
Abimael MartellandClaude Opus 4.6 b7ec80097b chore: update lopdf to latest commit
7a05512d831415b1f2b1ce522391d6beab8a1284

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 01:27:32 -07:00
Abimael MartellandClaude Opus 4.6 895e47b094 docs: add opendataloader-bench results to README
Compare pdf-inspector against other direct text extraction engines
(no OCR/ML) on the opendataloader-bench corpus.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 01:25:08 -07:00
Abimael MartellandClaude Opus 4.6 2f9390e4bf feat: detect slightly-larger-than-body text as headings
Lines with font size 1.10-1.20x body text that are standalone and short
(1-8 words) are promoted to headings. This catches academic paper
headings where the font is only ~10% larger than body text, below the
previous 1.2x threshold.

Also syncs the simpler to_markdown_from_lines path to match the
table-aware path (removes stale colon exclusion).

Benchmark improvement: MHS 0.54→0.56, overall 0.761→0.766.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 01:18:31 -07:00
Abimael MartellandClaude Opus 4.6 b25a122655 fix: require digit after Table/Figure prefix in caption detection
Caption detection was incorrectly classifying "Table of Contents" as a
caption because it starts with "Table ". Now "Table" and "Figure"
prefixes require a digit, parenthesis, or hash after them — matching
actual captions like "Table 1", "Figure 3.2" but not titles.

Also removes debug logging left from previous iteration.

Benchmark improvement: MHS 0.52→0.54, overall 0.757→0.761.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 01:13:17 -07:00
Abimael MartellandClaude Opus 4.6 103995c200 feat: remove colon exclusion from bold heading detection
The colon exclusion was preventing legitimate headings like "Steps for
Using the Microscope:" and "Changing objectives:" from being detected.
The single edge case it was protecting (chart sub-headers) is less
impactful than the many headings it was blocking.

Benchmark improvement: MHS 0.51→0.52, MHS-S 0.61→0.62.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 01:09:19 -07:00
Abimael MartellandClaude Opus 4.6 28823bd849 feat: improve heuristic table detection for small body-font tables
- Lower minimum item count for body-font table candidates from 9 to 6,
  allowing small 2-3 row tables to be detected.
- Allow 2-column body-font tables with short cells (avg ≤25 chars) to
  bypass the "table-like content" validation. This catches text-only
  definition/category tables (e.g., species lists) without false-positiving
  on 2-column paragraph text (which has longer cells).

Benchmark improvement: TEDS 0.498→0.519, overall 0.750→0.754.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 01:07:08 -07:00
Abimael MartellandClaude Opus 4.6 830b8d955b feat: detect bold-only lines as section headings
Bold lines at body font size that are standalone (preceded by a paragraph
break) and have ≥3 words are promoted to headings. This catches the
common pattern in academic/technical PDFs where section headings use
bold text at the same size as body text.

Guards against false positives: minimum word count, colon-ending
exclusion (labels like "Table I:").

Benchmark improvement: MHS 0.37→0.50, overall 0.71→0.75.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 00:56:18 -07:00
Abimael Martell f664f3056d fix: harden hybrid OCR region extraction path (#23)
Align region filtering with rotated-page coordinate rewrites, switch region text assembly to the shared line-grouping pipeline, and retain edge-overlap text to avoid false empty regions that incorrectly trigger OCR fallback. Also make Python region inputs fail fast with clear ValueError messages for malformed boxes.

Made-with: Cursor
2026-04-03 22:48:13 -07:00
Abimael Martell ebb01ab6a3 Merge pull request #22 from firecrawl/perf/skip-truetype-fallback-in-region-extract
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
perf: skip TrueType fallback in extract_text_in_regions (22,000x faster font parsing)
2026-04-02 18:14:05 -07:00
Abimael MartellandClaude Opus 4.6 53eff94147 Bump napi package version to 0.3.2
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 18:13:20 -07:00
Abimael MartellandClaude Opus 4.6 30244078d8 fix: sort panic and missing CID garbage check in extract_text_in_regions_mem
Two bugs in collect_text_in_region / extract_text_in_regions_mem:

1. The threshold-based sort comparator in collect_text_in_region was not
   transitive, causing Rust's sort to panic on certain PDFs. Replaced with
   strict total_cmp ordering — the line-grouping phase already handles
   fuzzy Y matching via threshold.

2. The needs_ocr check was missing is_cid_garbage, so Identity-H fonts
   with CID garbage (C1 control chars, high Latin mojibake) could pass
   all quality checks and be served as real text with needs_ocr=false.

Also adds 7 integration tests for extract_text_in_regions_mem (previously
had zero coverage): basic extraction, Identity-H needs_ocr, multiple
regions, nonexistent page, empty region, invalid input, and a fast-vs-normal
comparison test across all text-based fixtures.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 18:13:11 -07:00
Abimael MartellandClaude Opus 4.6 57673ebb69 perf: skip expensive TrueType fallback in extract_text_in_regions
FontCMaps::from_doc can spend 4.5+ seconds decompressing and parsing
large embedded TrueType fonts for CID fonts with sparse ToUnicode
CMaps. For extract_text_in_regions (hybrid OCR pipeline), this is
unnecessary — fonts that can't be decoded cheaply will produce
empty/garbage text, triggering needs_ocr=true and GPU OCR fallback.

Changes:
- Add FontCMaps::from_doc_pages_fast() that skips TrueType font
  fallback parsing (build_fallback_cmap_for_type0) and Identity-H/V
  second pass entirely
- Add FontCMaps::from_doc_pages() for filtered page sets
- extract_text_in_regions_mem uses fast mode
- Restructure fallback chain: try cheap fallbacks first, only attempt
  expensive TrueType parsing when needed and not in fast mode

Benchmark on nihms-1771367.pdf (19-page chemistry paper):
- FontCMaps fast:  201µs
- FontCMaps slow:  4.47s
- 22,000x speedup on font parsing

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 18:13:11 -07:00
Abimael MartellandClaude Opus 4.6 2d1c6f9ff8 Bump napi package version to 0.3.1
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 15:12:47 -07:00
Abimael Martell 55b32d542e Merge pull request #21 from firecrawl/fix/catch-unwind-napi-panics
fix: catch_unwind in NAPI layer to prevent process abort on panic
2026-04-02 15:12:28 -07:00
Abimael MartellandClaude Opus 4.6 79e53f779d wrap all NAPI functions in catch_unwind to prevent process abort on panic
Rust panics in NAPI modules abort the Node.js process with no chance
to report errors. This wraps every exported function in catch_unwind,
converting panics into JS Error exceptions that can be caught and
reported to Sentry.

Buffer data is extracted to Vec<u8> before the catch_unwind boundary
to satisfy UnwindSafe requirements.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 15:10:15 -07:00
Abimael MartellandClaude Opus 4.6 6dfa370fd4 Bump napi package version to 0.3.0
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 15:01:59 -07:00
Abimael MartellandClaude Opus 4.6 d3d76d0660 Fix js-bindings artifact path in publish workflow
upload-artifact strips the common napi/ prefix, so files are at
artifacts/js-bindings/index.js not artifacts/js-bindings/napi/index.js.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 14:58:00 -07:00
Abimael MartellandClaude Opus 4.6 2eb78b8a56 Upload generated JS bindings from build step instead of napi pre-publish
napi pre-publish expects a version-bump commit message convention.
Instead, upload index.js and index.d.ts generated by napi build
as artifacts and copy them into the publish step.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 14:54:56 -07:00
Abimael MartellandClaude Opus 4.6 a802874f2b Fix napi pre-publish command in publish workflow
The command is `napi pre-publish` (hyphenated), not `napi prepublish`,
and `--skip-gh-release` is not a valid flag.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 14:51:03 -07:00
Abimael MartellandClaude Opus 4.6 176d1ff2a3 Remove auto-generated napi files from version control
index.js and index.d.ts are generated by napi-rs. Generate them
at publish time via `napi prepublish` instead of checking them in.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 14:44:49 -07:00
Abimael MartellandClaude Opus 4.6 fba0a644ef fix: replace partial_cmp with total_cmp to prevent NaN sort panics (#20)
* fix sort panics on NaN values from bogus PDF font metrics

Replace all `partial_cmp(...).unwrap_or(Ordering::Equal)` and bare
`partial_cmp(...).unwrap()` with `total_cmp()` across the codebase.

`partial_cmp` returns `None` for NaN, and mapping that to `Equal`
violates total ordering: `a == NaN` and `NaN == b` but `a != b`.
Rust 1.81+ detects this and panics in sort_by. `total_cmp` handles
NaN deterministically (sorts to end) and guarantees total ordering.

The critical crash was in `extract_text_in_regions` (lib.rs:478)
where PDFs with bogus font ascent/descent values produced NaN in
text item coordinates, causing process abort via NAPI.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix missed partial_cmp in layout.rs and restore napi exports

- Convert two remaining b.y.partial_cmp(&a.y) calls to total_cmp
  in group_single_column and column layout sorting
- Restore missing napi exports: detectPdf, extractText,
  extractTextWithPositions, processPdf

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 14:39:42 -07:00
Abimael MartellandClaude Opus 4.6 732b1b1359 Mention Python and Node.js bindings in README intro
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 12:46:06 -07:00
Abimael MartellandClaude Opus 4.6 c66819c165 Reorganize README: move API references to docs/
Move detailed Python, Rust, and debugging docs into docs/ to keep
the main README focused on overview and quick start examples.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 12:28:17 -07:00
Abimael Martell 86476eb6b2 Merge pull request #2 from Jacqkues/feat/python-bindings
feat: add Python bindings via PyO3
2026-04-02 12:18:53 -07:00
Abimael MartellandClaude Opus 4.6 9868c02f6e add npm package description, keywords, and README
Publish npm package / Build aarch64-apple-darwin (push) Has been cancelled
Publish npm package / Build x86_64-unknown-linux-gnu (push) Has been cancelled
Publish npm package / Publish to npm (push) Has been cancelled
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 12:18:34 -07:00
Abimael MartellandClaude Opus 4.6 2e874f6cfd update napi index.js version strings to match 0.2.2
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 12:18:06 -07:00
Abimael Martell 4cba66ec94 Merge remote-tracking branch 'origin/main' into feat/python-bindings 2026-04-02 12:17:07 -07:00
Abimael Martell c9c65b78a7 Merge remote-tracking branch 'origin/main' into feat/python-bindings 2026-04-02 11:58:26 -07:00
Abimael MartellandClaude Opus 4.6 506b2a0c70 unify NAPI and Python binding APIs for consistent surface
Both bindings now expose the same 6 function families: process, detect,
classify, extractText, extractTextWithPositions, and extractTextInRegions.
Bumps PyO3 from 0.22 to 0.25 for Python 3.14 support.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 11:48:18 -07:00
Abimael Martell 5f829f5258 Merge branch 'main' into feat/python-bindings 2026-04-02 11:35:18 -07:00
Jacques DumoraandClaude Opus 4.6 0bf5463a1e feat: add Python bindings via PyO3
Expose the pdf-inspector Rust library as a Python package using PyO3 + maturin.
Python users can now `pip install` and use `import pdf_inspector` for PDF
classification, text extraction, and markdown conversion with native Rust speed.

Adds:
- src/python.rs: PyO3 bindings (process_pdf, detect_pdf, extract_text, etc.)
- pyproject.toml: maturin build configuration
- pdf_inspector.pyi: type stubs for IDE support
- tests/test_python.py: 21 pytest tests covering all Python API functions
- examples/basic_usage.py: example script demonstrating all features
- Updated README with Python quick start and API reference

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 17:38:53 +01:00
63 changed files with 16254 additions and 1546 deletions
+44 -4
View File
@@ -2,14 +2,41 @@ name: Publish npm package
on:
push:
tags: ['v*']
branches: [main]
paths: ['napi/package.json']
permissions:
contents: read
id-token: write
jobs:
check-version:
name: Check version change
runs-on: ubuntu-latest
outputs:
changed: ${{ steps.check.outputs.changed }}
version: ${{ steps.check.outputs.version }}
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 2
- name: Check if version changed
id: check
run: |
NEW_VERSION=$(node -p "require('./napi/package.json').version")
OLD_VERSION=$(git show HEAD~1:napi/package.json | node -p "JSON.parse(require('fs').readFileSync('/dev/stdin','utf8')).version")
echo "old=$OLD_VERSION new=$NEW_VERSION"
if [ "$NEW_VERSION" != "$OLD_VERSION" ]; then
echo "changed=true" >> "$GITHUB_OUTPUT"
echo "version=$NEW_VERSION" >> "$GITHUB_OUTPUT"
else
echo "changed=false" >> "$GITHUB_OUTPUT"
fi
build:
needs: check-version
if: needs.check-version.outputs.changed == 'true'
name: Build ${{ matrix.target }}
runs-on: ${{ matrix.os }}
strategy:
@@ -19,6 +46,8 @@ jobs:
target: x86_64-unknown-linux-gnu
- os: macos-14
target: aarch64-apple-darwin
- os: windows-latest
target: x86_64-pc-windows-msvc
steps:
- uses: actions/checkout@v4
@@ -49,16 +78,26 @@ jobs:
working-directory: napi
run: bunx napi build --platform --release
- name: Upload artifact
- name: Upload native binary
uses: actions/upload-artifact@v4
with:
name: bindings-${{ matrix.target }}
path: napi/*.node
if-no-files-found: error
- name: Upload generated JS bindings
if: matrix.target == 'x86_64-unknown-linux-gnu'
uses: actions/upload-artifact@v4
with:
name: js-bindings
path: |
napi/index.js
napi/index.d.ts
if-no-files-found: error
publish:
name: Publish to npm
needs: build
needs: [check-version, build]
runs-on: ubuntu-latest
permissions:
contents: read
@@ -79,8 +118,9 @@ jobs:
- name: Collect binaries and publish
working-directory: napi
run: |
# Copy all .node binaries into the package directory
cp artifacts/bindings-*/*.node .
cp artifacts/js-bindings/index.js .
cp artifacts/js-bindings/index.d.ts .
echo "=== Package contents ==="
ls -la *.node index.js index.d.ts
+9
View File
@@ -24,6 +24,10 @@ Thumbs.db
# Build cache
**/*.rs.bk
# NAPI generated (regenerated by `napi prepublish` during CI)
napi/index.js
napi/index.d.ts
# Local samples and scripts
samples/
scripts/
@@ -31,3 +35,8 @@ scripts/
# Test output
test_output/
# Python
__pycache__/
*.pyc
.pytest_cache/
+79
View File
@@ -0,0 +1,79 @@
# pdf-inspector
Fast PDF text extraction to structured Markdown. CLI binary: `pdf2md`. Detection binary: `detect-pdf`.
## Build & Test
```bash
cargo fmt # format
cargo clippy -- -D warnings # lint (enforced, zero warnings)
cargo test # unit + integration tests (267+ unit, 73+ integration)
cargo build --release # release binary for benchmarks
```
All three must pass before committing.
## Binaries
- `pdf2md` — extract PDF → Markdown. Supports `--json` for structured output.
- `detect-pdf` — classify PDF type (TextBased/Scanned/Mixed/ImageBased). Supports `--analyze --json`.
## Architecture
```
src/
lib.rs public API, process_pdf_with_options, encoding issue detection
detector.rs PDF type classification, tiled-scan detection, page sampling
types.rs TextItem, TextLine, PdfRect, PdfLine
tounicode.rs CMap/ToUnicode parsing, CID decoding
text_utils.rs CJK/RTL handling, Otsu threshold, ligature expansion, NFKC
extractor/
mod.rs top-level extraction orchestrator
content_stream.rs PDF operator state machine (Tj/TJ/Td/Tm/q/Q)
fonts.rs font width/encoding, CMapDecisionCache, TrueType cmap fallback
layout.rs column detection (histogram), newspaper/tabular classification,
spanning-line pre-masking, sidebar detection
tables/
detect_rects.rs rect-based table detection (union-find clustering)
detect_heuristic.rs heuristic table detection (gap-histogram, body-font tables)
detect_lines.rs line-based table detection (H/V line grids)
grid.rs column/row boundaries, cell assignment
format.rs table→Markdown formatting, continuation row merging
markdown/
convert.rs core line→Markdown loop, struct-tree role support
analysis.rs font stats, heading tiers, paragraph thresholds
classify.rs line classification (header, list, code, caption)
preprocess.rs drop cap merging, heading line merging
postprocess.rs dot leaders, hyphenation, page numbers, URL formatting
```
## Key design decisions
- **Primary audience is AI agents.** Output optimized for token efficiency and semantic quality, not visual formatting. No cosmetic padding.
- **Three table detection strategies** run in priority order: rect-based → line-based → heuristic. First valid result wins.
- **Column detection** uses horizontal projection histograms with valley detection. Multi-item spanning lines (titles, headers) are pre-masked using column-aware thresholds before column assignment.
- **Newspaper vs tabular** classification determines reading order: newspaper reads columns sequentially, tabular Y-interleaves them.
- **Tiled-scan detection** catches scanned PDFs with JBIG2/strip images where no single tile exceeds the template threshold but aggregate area does (≥2M pixels).
- **Garbage text upgrade** reclassifies Mixed PDFs as Scanned when extracted text is <50% alphanumeric.
- **Tagged PDF support** uses structure tree roles (H1-H6, P, L, Code, BlockQuote) when available, falling back to font-size heuristics.
## Testing
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
- **Regression suite**: sibling repo `pdf-evals` with 179+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
## Debugging
```bash
RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- file.pdf
RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- file.pdf
RUST_LOG=pdf_inspector::detector=debug cargo run --release --bin detect-pdf -- file.pdf
```
## Conventions
- Clippy: use `is_some_and(...)` not `map_or(false, ...)`
- lopdf quirk: `ParseError` is private — match by string for `InvalidFileHeader`
- Column limit for tables: 25 (wide statistical tables)
- `propagate_merged_cells` skipped for >10 columns (spanning rects = background fills)
+2 -1
View File
@@ -61,7 +61,8 @@ src/
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
- **Regression suite**: sibling repo `pdf-evals` with 179+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
- **Regression suite**: sibling repo `pdf-evals` with 187+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
## Debugging
+9 -3
View File
@@ -8,9 +8,16 @@ description = "Fast PDF inspection, classification, and text extraction with sma
license = "MIT"
repository = "https://github.com/firecrawl/pdf-inspector"
[lib]
name = "pdf_inspector"
crate-type = ["lib", "cdylib"]
[dependencies]
# Python bindings
pyo3 = { version = "0.25", features = ["extension-module"], optional = true }
# PDF parsing
lopdf = { git = "https://github.com/J-F-Liu/lopdf", rev = "052674053814a9f4897af94f0b8e46a545c9b329", features = ["rayon"] }
lopdf = { git = "https://github.com/J-F-Liu/lopdf", rev = "7a05512d831415b1f2b1ce522391d6beab8a1284", features = ["rayon"] }
# Error handling
thiserror = "2.0"
@@ -35,6 +42,7 @@ tempfile = "3.3"
[features]
default = []
python = ["pyo3"]
[[bin]]
name = "pdf2md"
@@ -47,5 +55,3 @@ path = "src/bin/detect_pdf.rs"
[[bin]]
name = "dump_ops"
path = "src/bin/dump_ops.rs"
+61 -139
View File
@@ -1,6 +1,6 @@
# pdf-inspector
Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR.
Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for [Python](docs/python.md) and [Node.js](napi/README.md).
Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.
@@ -12,90 +12,81 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
- **Table detection** — Dual-mode: rectangle-based detection from PDF drawing ops, plus heuristic detection from text alignment. Handles financial tables, footnotes, and continuation tables across pages.
- **CID font support** — ToUnicode CMap decoding for Type0/Identity-H fonts, UTF-16BE, UTF-8, and Latin-1 encodings.
- **Multi-column layout** — Automatic detection of newspaper-style columns, sequential reading order, and RTL text support.
- **Encoding issue detection** — Automatically flags broken font encodings (garbled text, replacement characters) so callers can fall back to OCR.
- **Encoding issue detection** — Automatically flags broken font encodings so callers can fall back to OCR.
- **Single document load** — The document is parsed once and shared between detection and extraction, avoiding redundant I/O.
- **Lightweight** — Pure Rust, no ML models, no external services. Single dependency on `lopdf` for PDF parsing.
## Benchmark
Evaluated on the [opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs). Only direct text extraction engines are shown — no OCR, no ML models. Scores are 0-1, higher is better.
| Engine | Overall | Reading Order (NID) | Tables (TEDS) | Headings (MHS) | Speed (200 docs) |
|---|---|---|---|---|---|
| pdf-inspector | 0.78 | 0.87 | 0.59 | 0.57 | 4s |
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
| markitdown | 0.58 | 0.88 | 0.00 | 0.00 | 8s |
For context, engines that use OCR/ML (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus.
**Where we do well:** Speed (fastest of all engines), reading order, table detection vs other direct-text tools.
**Where we lag:** Heading detection trails opendataloader — many PDFs use bold text at body font size for headings, or headings that are only slightly larger than body text. Table detection trails OCR-based engines that can see visual table structure.
## Quick start
### As a library
### Python
Add to your `Cargo.toml`:
```bash
pip install maturin
maturin develop --release
```
```python
import pdf_inspector
result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type) # "text_based", "scanned", "image_based", "mixed"
print(result.markdown) # Markdown string or None
```
> Full API reference: [docs/python.md](docs/python.md)
### Node.js
```bash
npm install @firecrawl/pdf-inspector
```
```javascript
import { readFileSync } from 'fs';
import { processPdf, classifyPdf } from '@firecrawl/pdf-inspector';
const result = processPdf(readFileSync('document.pdf'));
console.log(result.pdfType); // "TextBased", "Scanned", "ImageBased", "Mixed"
console.log(result.markdown); // Markdown string or null
```
> Full API reference: [napi/README.md](napi/README.md)
### Rust
```toml
[dependencies]
pdf-inspector = { git = "https://github.com/firecrawl/pdf-inspector" }
```
Detect and extract in one call:
```rust
use pdf_inspector::process_pdf;
let result = process_pdf("document.pdf")?;
println!("Type: {:?}", result.pdf_type); // TextBased, Scanned, ImageBased, Mixed
println!("Confidence: {:.0}%", result.confidence * 100.0);
println!("Pages: {}", result.page_count);
println!("Type: {:?}", result.pdf_type);
if let Some(markdown) = &result.markdown {
println!("{}", markdown);
}
```
Fast metadata-only detection (no text extraction or markdown generation):
```rust
use pdf_inspector::detect_pdf;
let info = detect_pdf("document.pdf")?;
match info.pdf_type {
pdf_inspector::PdfType::TextBased => {
// Extract locally — fast and free
}
_ => {
// Route to OCR service
// info.pages_needing_ocr tells you exactly which pages
}
}
```
Customize processing with `PdfOptions`:
```rust
use pdf_inspector::{process_pdf_with_options, PdfOptions, ProcessMode, DetectionConfig, ScanStrategy};
// Analyze layout without generating markdown
let result = process_pdf_with_options(
"document.pdf",
PdfOptions::new().mode(ProcessMode::Analyze),
)?;
// Full extraction with custom detection strategy
let result = process_pdf_with_options(
"large.pdf",
PdfOptions::new().detection(DetectionConfig {
strategy: ScanStrategy::Sample(5),
..Default::default()
}),
)?;
// Process only specific pages
let result = process_pdf_with_options(
"document.pdf",
PdfOptions::new().pages([1, 3, 5]),
)?;
```
Process from a byte buffer (no filesystem needed):
```rust
use pdf_inspector::process_pdf_mem;
let bytes = std::fs::read("document.pdf")?;
let result = process_pdf_mem(&bytes)?;
```
> Full API reference: [docs/rust-api.md](docs/rust-api.md)
### CLI
@@ -120,7 +111,6 @@ cargo run --bin detect-pdf -- document.pdf
cargo run --bin detect-pdf -- document.pdf --json
# Detection + layout analysis (tables, columns)
cargo run --bin detect-pdf -- document.pdf --analyze
cargo run --bin detect-pdf -- document.pdf --analyze --json
```
@@ -159,6 +149,7 @@ The document is loaded **once** via `load_document_from_path` / `load_document_f
```
src/
lib.rs — Public API, PdfOptions builder, convenience functions
python.rs — PyO3 Python bindings
types.rs — Shared types: TextItem, TextLine, PdfRect, ItemType
text_utils.rs — Character/text helpers (CJK, RTL, ligatures, bold/italic)
process_mode.rs — ProcessMode enum (DetectOnly, Analyze, Full)
@@ -169,6 +160,7 @@ src/
tables/ — Table detection and formatting
markdown/ — Markdown conversion and structure detection
bin/ — CLI tools (pdf2md, detect_pdf)
napi/ — Node.js/Bun bindings (napi-rs)
```
## How classification works
@@ -189,50 +181,6 @@ This detects 300+ page PDFs in milliseconds. The result includes `pages_needing_
| `Sample(n)` | Sample `n` evenly distributed pages (first, last, middle) | Very large PDFs where speed matters more than precision |
| `Pages(vec)` | Only scan specific 1-indexed page numbers | When the caller knows which pages to check |
## API
### Processing modes
| Mode | What it does | Returns |
|---|---|---|
| `ProcessMode::Full` (default) | Detect + extract + convert to Markdown | Everything populated |
| `ProcessMode::Analyze` | Detect + extract + layout analysis (no Markdown) | `markdown` is `None`, `layout` is populated |
| `ProcessMode::DetectOnly` | Classification only (fastest) | `markdown` is `None`, `layout` is default |
### Functions
| Function | Description |
|---|---|
| `process_pdf(path)` | Full processing with defaults |
| `detect_pdf(path)` | Fast metadata-only detection (no extraction) |
| `process_pdf_with_options(path, options)` | Process with custom `PdfOptions` |
| `process_pdf_mem(bytes)` | Full processing from a byte buffer |
| `detect_pdf_mem(bytes)` | Fast detection from a byte buffer |
| `process_pdf_mem_with_options(bytes, options)` | Process from bytes with custom options |
| `extract_text(path)` | Plain text extraction |
| `extract_text_with_positions(path)` | Text with X/Y coordinates and font info |
| `to_markdown(text, options)` | Convert plain text to Markdown |
| `to_markdown_from_items(items, options)` | Markdown from pre-extracted `TextItem`s |
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
### Types
| Type | Description |
|---|---|
| `PdfOptions` | Builder for processing configuration (mode, detection, markdown, page filter) |
| `ProcessMode` | `DetectOnly`, `Analyze`, `Full` |
| `PdfType` | `TextBased`, `Scanned`, `ImageBased`, `Mixed` |
| `PdfProcessResult` | Full result: pdf_type, markdown, page_count, confidence, layout, has_encoding_issues, timing |
| `PdfTypeResult` | Low-level detection result: type, confidence, page count, pages needing OCR |
| `DetectionConfig` | Configuration for detection: scan strategy, thresholds |
| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` |
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
| `TextItem` | Text with position, font info, and page number |
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
| `PdfError` | `Io`, `Parse`, `Encrypted`, `InvalidStructure`, `NotAPdf` |
## Markdown output
The converter handles:
@@ -255,36 +203,6 @@ The converter handles:
| Drop caps | Large initial letters merged with following text |
| Dot leaders | TOC-style dots collapsed to " ... " |
## Debugging with RUST_LOG
Structured logging via `RUST_LOG` replaces the former debug binaries. Set the environment variable to control which sections emit debug output on stderr:
```bash
# Raw PDF content stream operators (replaces dump_ops)
RUST_LOG=pdf_inspector::extractor::content_stream=trace cargo run --bin pdf2md -- file.pdf > /dev/null
# Font metadata, encodings, ligatures (replaces debug_fonts / debug_ligatures)
RUST_LOG=pdf_inspector::extractor::fonts=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# ToUnicode CMap parsing
RUST_LOG=pdf_inspector::tounicode=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Text items per page with x/y/width (replaces debug_spaces / debug_pages)
RUST_LOG=pdf_inspector::extractor=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Column detection and reading order (replaces debug_order)
RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Y-gap analysis and paragraph thresholds (replaces debug_ygaps)
RUST_LOG=pdf_inspector::markdown::analysis=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Table detection
RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Everything
RUST_LOG=pdf_inspector=debug cargo run --bin pdf2md -- file.pdf > /dev/null
```
## Use case: smart PDF routing
pdf-inspector was built for pipelines that process PDFs at scale. Instead of sending every PDF through OCR:
@@ -299,6 +217,10 @@ PDF arrives
This saves cost and latency for the majority of PDFs that are already text-based (reports, papers, invoices, legal docs).
## Debugging
See [docs/debugging.md](docs/debugging.md) for `RUST_LOG` environment variable usage.
## License
MIT
+29
View File
@@ -0,0 +1,29 @@
# Debugging with RUST_LOG
Structured logging via `RUST_LOG` replaces the former debug binaries. Set the environment variable to control which sections emit debug output on stderr:
```bash
# Raw PDF content stream operators (replaces dump_ops)
RUST_LOG=pdf_inspector::extractor::content_stream=trace cargo run --bin pdf2md -- file.pdf > /dev/null
# Font metadata, encodings, ligatures (replaces debug_fonts / debug_ligatures)
RUST_LOG=pdf_inspector::extractor::fonts=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# ToUnicode CMap parsing
RUST_LOG=pdf_inspector::tounicode=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Text items per page with x/y/width (replaces debug_spaces / debug_pages)
RUST_LOG=pdf_inspector::extractor=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Column detection and reading order (replaces debug_order)
RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Y-gap analysis and paragraph thresholds (replaces debug_ygaps)
RUST_LOG=pdf_inspector::markdown::analysis=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Table detection
RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- file.pdf > /dev/null
# Everything
RUST_LOG=pdf_inspector=debug cargo run --bin pdf2md -- file.pdf > /dev/null
```
+88
View File
@@ -0,0 +1,88 @@
# Python API
Python bindings via [PyO3](https://pyo3.rs). Requires Rust toolchain for building from source.
## Install
```bash
pip install maturin
maturin develop --release
```
## Usage
```python
import pdf_inspector
# Full processing: detect + extract + convert to Markdown
result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type) # "text_based", "scanned", "image_based", "mixed"
print(result.confidence) # 0.0 - 1.0
print(result.page_count) # number of pages
print(result.markdown) # Markdown string or None
# Process specific pages only
result = pdf_inspector.process_pdf("document.pdf", pages=[1, 3, 5])
# Process from bytes (no filesystem needed)
with open("document.pdf", "rb") as f:
result = pdf_inspector.process_pdf_bytes(f.read())
# Fast detection only (no text extraction)
result = pdf_inspector.detect_pdf("document.pdf")
if result.pdf_type == "text_based":
print("Can extract locally!")
else:
print(f"Pages needing OCR: {result.pages_needing_ocr}")
# Plain text extraction
text = pdf_inspector.extract_text("document.pdf")
# Positioned text items with font info
items = pdf_inspector.extract_text_with_positions("document.pdf")
for item in items[:5]:
print(f"'{item.text}' at ({item.x:.0f}, {item.y:.0f}) size={item.font_size}")
# Per-page markdown (one Markdown string per page, plus layout metadata)
result = pdf_inspector.extract_pages_markdown("document.pdf")
for page in result.pages:
print(f"Page {page.page}: {len(page.markdown)} chars, needs_ocr={page.needs_ocr}")
# Restrict to specific 0-indexed pages (preserves caller order)
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
```
## API reference
| Function | Description |
|---|---|
| `process_pdf(path, pages=None)` | Full processing (detect + extract + markdown) |
| `process_pdf_bytes(data, pages=None)` | Full processing from bytes |
| `detect_pdf(path)` | Fast detection only (returns PdfResult) |
| `detect_pdf_bytes(data)` | Fast detection from bytes |
| `classify_pdf(path)` | Lightweight classification (returns PdfClassification) |
| `classify_pdf_bytes(data)` | Lightweight classification from bytes |
| `extract_text(path)` | Plain text extraction |
| `extract_text_bytes(data)` | Plain text extraction from bytes |
| `extract_text_with_positions(path, pages=None)` | Text with X/Y coords and font info |
| `extract_text_with_positions_bytes(data, pages=None)` | Text with positions from bytes |
| `extract_text_in_regions(path, page_regions)` | Extract text in bounding-box regions |
| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes |
| `extract_pages_markdown(path, pages=None)` | Per-page Markdown + layout metadata (all pages by default) |
| `extract_pages_markdown_bytes(data, pages=None)` | Per-page Markdown from bytes |
## Types
**`PdfResult` fields:** `pdf_type`, `markdown`, `page_count`, `processing_time_ms`, `pages_needing_ocr`, `title`, `confidence`, `is_complex_layout`, `pages_with_tables`, `pages_with_columns`, `has_encoding_issues`
**`PdfClassification` fields:** `pdf_type`, `page_count`, `pages_needing_ocr` (0-indexed), `confidence`
**`TextItem` fields:** `text`, `x`, `y`, `width`, `height`, `font`, `font_size`, `page`, `is_bold`, `is_italic`, `item_type`
**`RegionText` fields:** `text`, `needs_ocr`
**`PageRegionTexts` fields:** `page` (0-indexed), `regions` (list of RegionText)
**`PageMarkdown` fields:** `page` (0-indexed), `markdown`, `needs_ocr`
**`PagesExtractionResult` fields:** `pages` (list of PageMarkdown), `pages_with_tables` (1-indexed), `pages_with_columns` (1-indexed), `pages_needing_ocr` (1-indexed), `is_complex`
+147
View File
@@ -0,0 +1,147 @@
# Rust API
Add to your `Cargo.toml`:
```toml
[dependencies]
pdf-inspector = { git = "https://github.com/firecrawl/pdf-inspector" }
```
## Usage
Detect and extract in one call:
```rust
use pdf_inspector::process_pdf;
let result = process_pdf("document.pdf")?;
println!("Type: {:?}", result.pdf_type); // TextBased, Scanned, ImageBased, Mixed
println!("Confidence: {:.0}%", result.confidence * 100.0);
println!("Pages: {}", result.page_count);
if let Some(markdown) = &result.markdown {
println!("{}", markdown);
}
```
Fast metadata-only detection (no text extraction or markdown generation):
```rust
use pdf_inspector::detect_pdf;
let info = detect_pdf("document.pdf")?;
match info.pdf_type {
pdf_inspector::PdfType::TextBased => {
// Extract locally — fast and free
}
_ => {
// Route to OCR service
// info.pages_needing_ocr tells you exactly which pages
}
}
```
Customize processing with `PdfOptions`:
```rust
use pdf_inspector::{process_pdf_with_options, PdfOptions, ProcessMode, DetectionConfig, ScanStrategy};
// Analyze layout without generating markdown
let result = process_pdf_with_options(
"document.pdf",
PdfOptions::new().mode(ProcessMode::Analyze),
)?;
// Full extraction with custom detection strategy
let result = process_pdf_with_options(
"large.pdf",
PdfOptions::new().detection(DetectionConfig {
strategy: ScanStrategy::Sample(5),
..Default::default()
}),
)?;
// Process only specific pages
let result = process_pdf_with_options(
"document.pdf",
PdfOptions::new().pages([1, 3, 5]),
)?;
```
Process from a byte buffer (no filesystem needed):
```rust
use pdf_inspector::process_pdf_mem;
let bytes = std::fs::read("document.pdf")?;
let result = process_pdf_mem(&bytes)?;
```
Extract per-page Markdown (one string per page, plus document-wide layout
metadata):
```rust
use pdf_inspector::extract_pages_markdown;
// Pass `None` for every page in document order, or a slice of 0-indexed
// pages to restrict the output (caller-supplied order is preserved).
let result = extract_pages_markdown("document.pdf", None)?;
for page in &result.pages {
if page.needs_ocr {
// Route this page to OCR
} else {
println!("Page {}: {}", page.page, page.markdown);
}
}
println!("Complex layout? {}", result.is_complex);
```
## Processing modes
| Mode | What it does | Returns |
|---|---|---|
| `ProcessMode::Full` (default) | Detect + extract + convert to Markdown | Everything populated |
| `ProcessMode::Analyze` | Detect + extract + layout analysis (no Markdown) | `markdown` is `None`, `layout` is populated |
| `ProcessMode::DetectOnly` | Classification only (fastest) | `markdown` is `None`, `layout` is default |
## Functions
| Function | Description |
|---|---|
| `process_pdf(path)` | Full processing with defaults |
| `detect_pdf(path)` | Fast metadata-only detection (no extraction) |
| `process_pdf_with_options(path, options)` | Process with custom `PdfOptions` |
| `process_pdf_mem(bytes)` | Full processing from a byte buffer |
| `detect_pdf_mem(bytes)` | Fast detection from a byte buffer |
| `process_pdf_mem_with_options(bytes, options)` | Process from bytes with custom options |
| `extract_text(path)` | Plain text extraction |
| `extract_text_with_positions(path)` | Text with X/Y coordinates and font info |
| `to_markdown(text, options)` | Convert plain text to Markdown |
| `to_markdown_from_items(items, options)` | Markdown from pre-extracted `TextItem`s |
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
| `extract_pages_markdown(path, pages)` | Per-page Markdown + layout metadata (file) |
| `extract_pages_markdown_mem(bytes, pages)` | Per-page Markdown from bytes |
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
## Types
| Type | Description |
|---|---|
| `PdfOptions` | Builder for processing configuration (mode, detection, markdown, page filter) |
| `ProcessMode` | `DetectOnly`, `Analyze`, `Full` |
| `PdfType` | `TextBased`, `Scanned`, `ImageBased`, `Mixed` |
| `PdfProcessResult` | Full result: pdf_type, markdown, page_count, confidence, layout, has_encoding_issues, timing |
| `PdfTypeResult` | Low-level detection result: type, confidence, page count, pages needing OCR |
| `DetectionConfig` | Configuration for detection: scan strategy, thresholds |
| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` |
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
| `TextItem` | Text with position, font info, and page number |
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
| `PdfError` | `Io`, `Parse`, `Encrypted`, `InvalidStructure`, `NotAPdf` |
+98
View File
@@ -0,0 +1,98 @@
"""Basic usage examples for pdf-inspector Python library."""
import sys
import pdf_inspector
def main():
if len(sys.argv) < 2:
print("Usage: python basic_usage.py <path-to-pdf>")
sys.exit(1)
path = sys.argv[1]
# 1. Full processing: detect + extract + markdown
print("=" * 60)
print("Full processing")
print("=" * 60)
result = pdf_inspector.process_pdf(path)
print(f"Type: {result.pdf_type}")
print(f"Pages: {result.page_count}")
print(f"Confidence: {result.confidence:.0%}")
print(f"Time: {result.processing_time_ms}ms")
print(f"Title: {result.title}")
print(f"Complex: {result.is_complex_layout}")
print(f"Tables on: {result.pages_with_tables}")
print(f"Columns on: {result.pages_with_columns}")
print(f"Encoding: {'issues detected' if result.has_encoding_issues else 'ok'}")
print(f"OCR needed: {result.pages_needing_ocr or 'none'}")
if result.markdown:
print(f"\n--- Markdown ({len(result.markdown)} chars) ---")
print(result.markdown[:500])
if len(result.markdown) > 500:
print(f"\n... ({len(result.markdown) - 500} more chars)")
# 2. Fast detection only
print("\n" + "=" * 60)
print("Detection only")
print("=" * 60)
info = pdf_inspector.detect_pdf(path)
print(f"Type: {info.pdf_type}")
print(f"Confidence: {info.confidence:.0%}")
print(f"Time: {info.processing_time_ms}ms")
# 3. From bytes
print("\n" + "=" * 60)
print("From bytes")
print("=" * 60)
with open(path, "rb") as f:
data = f.read()
result = pdf_inspector.process_pdf_bytes(data)
print(f"Type: {result.pdf_type}, Pages: {result.page_count}")
# 4. Plain text
print("\n" + "=" * 60)
print("Plain text extraction")
print("=" * 60)
text = pdf_inspector.extract_text(path)
print(text[:300])
# 5. Positioned items
print("\n" + "=" * 60)
print("Positioned text items (first 10)")
print("=" * 60)
items = pdf_inspector.extract_text_with_positions(path, pages=[1])
for item in items[:10]:
bold = " [B]" if item.is_bold else ""
italic = " [I]" if item.is_italic else ""
print(
f" p{item.page} ({item.x:6.1f}, {item.y:6.1f}) "
f"size={item.font_size:5.1f}{bold}{italic} "
f"'{item.text}'"
)
# 6. Lightweight classification
print("\n" + "=" * 60)
print("Lightweight classification")
print("=" * 60)
cls = pdf_inspector.classify_pdf(path)
print(f"Type: {cls.pdf_type}")
print(f"Pages: {cls.page_count}")
print(f"Confidence: {cls.confidence:.0%}")
print(f"OCR pages: {cls.pages_needing_ocr or 'none'} (0-indexed)")
# 7. Region-based text extraction
print("\n" + "=" * 60)
print("Region-based text extraction (page 0, top region)")
print("=" * 60)
regions = pdf_inspector.extract_text_in_regions(
path, [(0, [[0.0, 0.0, 600.0, 200.0]])]
)
for page_result in regions:
for i, region in enumerate(page_result.regions):
print(f" Region {i}: needs_ocr={region.needs_ocr}")
print(f" Text: {region.text[:200]}")
if __name__ == "__main__":
main()
+1 -19
View File
@@ -129,12 +129,6 @@ version = "3.20.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5d20789868f4b01b2f2caec9f5c4e0213b41e3e5702a50157d699ae31ced2fcb"
[[package]]
name = "bytecount"
version = "0.6.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "175812e0be2bccb6abe50bb8d566126198344f707e304f45c648fd8f2cc0365e"
[[package]]
name = "cbc"
version = "0.1.2"
@@ -679,7 +673,7 @@ checksum = "5e5032e24019045c762d3c0f28f5b6b8bbf38563a65908389bf7978758920897"
[[package]]
name = "lopdf"
version = "0.40.0"
source = "git+https://github.com/J-F-Liu/lopdf?rev=052674053814a9f4897af94f0b8e46a545c9b329#052674053814a9f4897af94f0b8e46a545c9b329"
source = "git+https://github.com/J-F-Liu/lopdf?rev=7a05512d831415b1f2b1ce522391d6beab8a1284#7a05512d831415b1f2b1ce522391d6beab8a1284"
dependencies = [
"aes",
"bitflags",
@@ -695,7 +689,6 @@ dependencies = [
"log",
"md-5",
"nom",
"nom_locate",
"rand",
"rangemap",
"rayon",
@@ -807,17 +800,6 @@ dependencies = [
"memchr",
]
[[package]]
name = "nom_locate"
version = "5.0.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0b577e2d69827c4740cba2b52efaad1c4cc7c73042860b199710b3575c68438d"
dependencies = [
"bytecount",
"memchr",
"nom",
]
[[package]]
name = "num-conv"
version = "0.2.1"
+99
View File
@@ -0,0 +1,99 @@
# PDF Inspector
Fast PDF classification and region-based text extraction for Node.js/Bun. Native Rust performance via [napi-rs](https://napi.rs).
Built by [Firecrawl](https://firecrawl.dev) for hybrid OCR pipelines — extract text from PDF structure where possible, fall back to OCR only when needed.
## Install
```bash
npm install @firecrawl/pdf-inspector
# or
bun add @firecrawl/pdf-inspector
```
Prebuilt binaries included for **linux-x64** and **macOS ARM64**. No Rust toolchain needed.
## API
### `classifyPdf(buffer: Buffer): PdfClassification`
Classify a PDF as TextBased, Scanned, Mixed, or ImageBased (~10-50ms). Returns which pages need OCR.
```typescript
import { classifyPdf } from '@firecrawl/pdf-inspector'
import { readFileSync } from 'fs'
const pdf = readFileSync('document.pdf')
const result = classifyPdf(pdf)
console.log(result.pdfType) // "TextBased" | "Scanned" | "Mixed" | "ImageBased"
console.log(result.pageCount) // 42
console.log(result.pagesNeedingOcr) // [5, 12, 15] (0-indexed)
console.log(result.confidence) // 0.875
```
### `extractTextInRegions(buffer: Buffer, pageRegions: PageRegions[]): PageRegionTexts[]`
Extract text within bounding-box regions from a PDF. Designed for hybrid OCR pipelines where a layout model detects regions in rendered page images, and this function extracts text from the PDF structure for text-based pages — skipping GPU OCR.
Each region result includes a `needsOcr` flag that signals unreliable extraction (empty text, GID-encoded fonts, garbage text, encoding issues).
```typescript
import { extractTextInRegions } from '@firecrawl/pdf-inspector'
const result = extractTextInRegions(pdf, [
{
page: 0, // 0-indexed
regions: [
[0, 0, 300, 400], // [x1, y1, x2, y2] in PDF points, top-left origin
[300, 0, 612, 400],
]
}
])
for (const region of result[0].regions) {
if (region.needsOcr) {
// Unreliable text — send this region to OCR instead
} else {
console.log(region.text) // Extracted text in reading order
}
}
```
## Types
```typescript
interface PdfClassification {
pdfType: string // "TextBased" | "Scanned" | "Mixed" | "ImageBased"
pageCount: number
pagesNeedingOcr: number[] // 0-indexed page numbers
confidence: number // 0.0 - 1.0
}
interface PageRegions {
page: number // 0-indexed
regions: number[][] // [[x1, y1, x2, y2], ...] in PDF points, top-left origin
}
interface PageRegionTexts {
page: number
regions: RegionText[]
}
interface RegionText {
text: string
needsOcr: boolean // true when text is unreliable
}
```
## Platforms
| Platform | Architecture | Supported |
|----------|-------------|-----------|
| Linux | x64 | Yes |
| macOS | ARM64 | Yes |
## License
MIT
+131
View File
@@ -0,0 +1,131 @@
#!/usr/bin/env node
import { readFileSync, writeFileSync } from "fs";
import { createRequire } from "module";
const require = createRequire(import.meta.url);
const { version } = require("../package.json");
const HELP = `pdf-inspector v${version} — Fast PDF text extraction to Markdown
Usage:
pdf-inspector <file> Extract markdown (default)
pdf-inspector detect <file> Classify PDF type
Options:
--json Output as JSON
--pages <pages> Comma-separated page numbers (e.g. 1,3,5)
-o, --output <file> Write output to file instead of stdout
-h, --help Show this help
-v, --version Show version
Examples:
pdf-inspector document.pdf
pdf-inspector document.pdf --json
pdf-inspector document.pdf --pages 1,2,3
pdf-inspector detect document.pdf --json
cat document.pdf | pdf-inspector -`;
function die(msg) {
process.stderr.write(`error: ${msg}\n`);
process.exit(1);
}
function parseArgs(argv) {
const opts = { json: false, pages: null, output: null, file: null, command: "extract" };
let i = 0;
// Check for subcommand
if (argv[0] === "detect") {
opts.command = "detect";
i = 1;
}
while (i < argv.length) {
const arg = argv[i];
if (arg === "-h" || arg === "--help") {
process.stdout.write(HELP + "\n");
process.exit(0);
} else if (arg === "-v" || arg === "--version") {
process.stdout.write(`${version}\n`);
process.exit(0);
} else if (arg === "--json") {
opts.json = true;
} else if (arg === "--pages") {
i++;
if (!argv[i]) die("--pages requires a value (e.g. 1,3,5)");
opts.pages = argv[i].split(",").map((p) => {
const n = parseInt(p.trim(), 10);
if (Number.isNaN(n) || n < 1) die(`invalid page number: ${p}`);
return n;
});
} else if (arg === "-o" || arg === "--output") {
i++;
if (!argv[i]) die("-o requires a filename");
opts.output = argv[i];
} else if (arg === "-" || !arg.startsWith("-")) {
if (opts.file) die(`unexpected argument: ${arg}`);
opts.file = arg;
} else {
die(`unknown option: ${arg}`);
}
i++;
}
return opts;
}
function readInput(file) {
if (file === "-") {
return readFileSync(0); // stdin fd
}
try {
return readFileSync(file);
} catch (err) {
if (err.code === "ENOENT") die(`file not found: ${file}`);
die(err.message);
}
}
function output(text, outputPath) {
if (outputPath) {
writeFileSync(outputPath, text);
} else {
process.stdout.write(text);
}
}
// ---- main ----
const opts = parseArgs(process.argv.slice(2));
if (!opts.file) {
// Check if stdin is piped
if (process.stdin.isTTY !== false) {
process.stderr.write(HELP + "\n");
process.exit(1);
}
opts.file = "-";
}
const { processPdf, classifyPdf } = await import("../index.js");
const buffer = readInput(opts.file);
if (opts.command === "detect") {
const result = classifyPdf(buffer);
if (opts.json) {
output(JSON.stringify(result, null, 2) + "\n", opts.output);
} else {
const ocr = result.pagesNeedingOcr.length > 0
? `, ${result.pagesNeedingOcr.length} pages need OCR`
: "";
output(`${result.pdfType} (${result.pageCount} pages, confidence: ${result.confidence.toFixed(2)}${ocr})\n`, opts.output);
}
} else {
const result = processPdf(buffer, opts.pages ?? undefined);
if (opts.json) {
output(JSON.stringify(result, null, 2) + "\n", opts.output);
} else {
output((result.markdown ?? "") + "\n", opts.output);
}
}
+1 -1
View File
@@ -1,3 +1,3 @@
fn main() {
napi_build::setup();
napi_build::setup();
}
-49
View File
@@ -1,49 +0,0 @@
/* auto-generated by NAPI-RS */
/* eslint-disable */
/**
* Classify a PDF: detect type (TextBased/Scanned/Mixed/ImageBased),
* page count, and which pages need OCR. Takes PDF bytes as Buffer.
*/
export declare function classifyPdf(buffer: Buffer): PdfClassification
/**
* Extract text within bounding-box regions from a PDF.
*
* For hybrid OCR: layout model detects regions in rendered images,
* this extracts PDF text within those regions — skipping GPU OCR
* for text-based pages.
*
* Each region result includes `needs_ocr` — set when the extracted text
* is unreliable (empty, GID-encoded fonts, garbage, encoding issues).
*
* Coordinates are PDF points with top-left origin.
*/
export declare function extractTextInRegions(buffer: Buffer, pageRegions: Array<PageRegions>): Array<PageRegionTexts>
/** A page's regions for text extraction: (page_index_0based, bboxes). */
export interface PageRegions {
page: number
/** Each bbox is [x1, y1, x2, y2] in PDF points, top-left origin. */
regions: Array<Array<number>>
}
/** Extracted text for one page's regions. */
export interface PageRegionTexts {
page: number
regions: Array<RegionText>
}
/** Lightweight PDF classification result. */
export interface PdfClassification {
pdfType: string
pageCount: number
pagesNeedingOcr: Array<number>
confidence: number
}
/** Extracted text for a single region. */
export interface RegionText {
text: string
/** `true` when the text should not be trusted (empty, GID fonts, garbage, encoding issues). */
needsOcr: boolean
}
-580
View File
@@ -1,580 +0,0 @@
// prettier-ignore
/* eslint-disable */
// @ts-nocheck
/* auto-generated by NAPI-RS */
const { readFileSync } = require('node:fs')
let nativeBinding = null
const loadErrors = []
const isMusl = () => {
let musl = false
if (process.platform === 'linux') {
musl = isMuslFromFilesystem()
if (musl === null) {
musl = isMuslFromReport()
}
if (musl === null) {
musl = isMuslFromChildProcess()
}
}
return musl
}
const isFileMusl = (f) => f.includes('libc.musl-') || f.includes('ld-musl-')
const isMuslFromFilesystem = () => {
try {
return readFileSync('/usr/bin/ldd', 'utf-8').includes('musl')
} catch {
return null
}
}
const isMuslFromReport = () => {
let report = null
if (typeof process.report?.getReport === 'function') {
process.report.excludeNetwork = true
report = process.report.getReport()
}
if (!report) {
return null
}
if (report.header && report.header.glibcVersionRuntime) {
return false
}
if (Array.isArray(report.sharedObjects)) {
if (report.sharedObjects.some(isFileMusl)) {
return true
}
}
return false
}
const isMuslFromChildProcess = () => {
try {
return require('child_process').execSync('ldd --version', { encoding: 'utf8' }).includes('musl')
} catch (e) {
// If we reach this case, we don't know if the system is musl or not, so is better to just fallback to false
return false
}
}
function requireNative() {
if (process.env.NAPI_RS_NATIVE_LIBRARY_PATH) {
try {
return require(process.env.NAPI_RS_NATIVE_LIBRARY_PATH);
} catch (err) {
loadErrors.push(err)
}
} else if (process.platform === 'android') {
if (process.arch === 'arm64') {
try {
return require('./pdf-inspector.android-arm64.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-android-arm64')
const bindingPackageVersion = require('firecrawl-pdf-inspector-android-arm64/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else if (process.arch === 'arm') {
try {
return require('./pdf-inspector.android-arm-eabi.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-android-arm-eabi')
const bindingPackageVersion = require('firecrawl-pdf-inspector-android-arm-eabi/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
loadErrors.push(new Error(`Unsupported architecture on Android ${process.arch}`))
}
} else if (process.platform === 'win32') {
if (process.arch === 'x64') {
if (process.config?.variables?.shlib_suffix === 'dll.a' || process.config?.variables?.node_target_type === 'shared_library') {
try {
return require('./pdf-inspector.win32-x64-gnu.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-win32-x64-gnu')
const bindingPackageVersion = require('firecrawl-pdf-inspector-win32-x64-gnu/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
try {
return require('./pdf-inspector.win32-x64-msvc.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-win32-x64-msvc')
const bindingPackageVersion = require('firecrawl-pdf-inspector-win32-x64-msvc/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
}
} else if (process.arch === 'ia32') {
try {
return require('./pdf-inspector.win32-ia32-msvc.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-win32-ia32-msvc')
const bindingPackageVersion = require('firecrawl-pdf-inspector-win32-ia32-msvc/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else if (process.arch === 'arm64') {
try {
return require('./pdf-inspector.win32-arm64-msvc.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-win32-arm64-msvc')
const bindingPackageVersion = require('firecrawl-pdf-inspector-win32-arm64-msvc/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
loadErrors.push(new Error(`Unsupported architecture on Windows: ${process.arch}`))
}
} else if (process.platform === 'darwin') {
try {
return require('./pdf-inspector.darwin-universal.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-darwin-universal')
const bindingPackageVersion = require('firecrawl-pdf-inspector-darwin-universal/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
if (process.arch === 'x64') {
try {
return require('./pdf-inspector.darwin-x64.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-darwin-x64')
const bindingPackageVersion = require('firecrawl-pdf-inspector-darwin-x64/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else if (process.arch === 'arm64') {
try {
return require('./pdf-inspector.darwin-arm64.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-darwin-arm64')
const bindingPackageVersion = require('firecrawl-pdf-inspector-darwin-arm64/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
loadErrors.push(new Error(`Unsupported architecture on macOS: ${process.arch}`))
}
} else if (process.platform === 'freebsd') {
if (process.arch === 'x64') {
try {
return require('./pdf-inspector.freebsd-x64.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-freebsd-x64')
const bindingPackageVersion = require('firecrawl-pdf-inspector-freebsd-x64/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else if (process.arch === 'arm64') {
try {
return require('./pdf-inspector.freebsd-arm64.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-freebsd-arm64')
const bindingPackageVersion = require('firecrawl-pdf-inspector-freebsd-arm64/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
loadErrors.push(new Error(`Unsupported architecture on FreeBSD: ${process.arch}`))
}
} else if (process.platform === 'linux') {
if (process.arch === 'x64') {
if (isMusl()) {
try {
return require('./pdf-inspector.linux-x64-musl.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-x64-musl')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-x64-musl/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
try {
return require('./pdf-inspector.linux-x64-gnu.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-x64-gnu')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-x64-gnu/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
}
} else if (process.arch === 'arm64') {
if (isMusl()) {
try {
return require('./pdf-inspector.linux-arm64-musl.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-arm64-musl')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-arm64-musl/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
try {
return require('./pdf-inspector.linux-arm64-gnu.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-arm64-gnu')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-arm64-gnu/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
}
} else if (process.arch === 'arm') {
if (isMusl()) {
try {
return require('./pdf-inspector.linux-arm-musleabihf.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-arm-musleabihf')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-arm-musleabihf/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
try {
return require('./pdf-inspector.linux-arm-gnueabihf.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-arm-gnueabihf')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-arm-gnueabihf/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
}
} else if (process.arch === 'loong64') {
if (isMusl()) {
try {
return require('./pdf-inspector.linux-loong64-musl.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-loong64-musl')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-loong64-musl/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
try {
return require('./pdf-inspector.linux-loong64-gnu.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-loong64-gnu')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-loong64-gnu/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
}
} else if (process.arch === 'riscv64') {
if (isMusl()) {
try {
return require('./pdf-inspector.linux-riscv64-musl.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-riscv64-musl')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-riscv64-musl/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
try {
return require('./pdf-inspector.linux-riscv64-gnu.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-riscv64-gnu')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-riscv64-gnu/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
}
} else if (process.arch === 'ppc64') {
try {
return require('./pdf-inspector.linux-ppc64-gnu.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-ppc64-gnu')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-ppc64-gnu/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else if (process.arch === 's390x') {
try {
return require('./pdf-inspector.linux-s390x-gnu.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-linux-s390x-gnu')
const bindingPackageVersion = require('firecrawl-pdf-inspector-linux-s390x-gnu/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
loadErrors.push(new Error(`Unsupported architecture on Linux: ${process.arch}`))
}
} else if (process.platform === 'openharmony') {
if (process.arch === 'arm64') {
try {
return require('./pdf-inspector.openharmony-arm64.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-openharmony-arm64')
const bindingPackageVersion = require('firecrawl-pdf-inspector-openharmony-arm64/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else if (process.arch === 'x64') {
try {
return require('./pdf-inspector.openharmony-x64.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-openharmony-x64')
const bindingPackageVersion = require('firecrawl-pdf-inspector-openharmony-x64/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else if (process.arch === 'arm') {
try {
return require('./pdf-inspector.openharmony-arm.node')
} catch (e) {
loadErrors.push(e)
}
try {
const binding = require('firecrawl-pdf-inspector-openharmony-arm')
const bindingPackageVersion = require('firecrawl-pdf-inspector-openharmony-arm/package.json').version
if (bindingPackageVersion !== '0.2.0' && process.env.NAPI_RS_ENFORCE_VERSION_CHECK && process.env.NAPI_RS_ENFORCE_VERSION_CHECK !== '0') {
throw new Error(`Native binding package version mismatch, expected 0.2.0 but got ${bindingPackageVersion}. You can reinstall dependencies to fix this issue.`)
}
return binding
} catch (e) {
loadErrors.push(e)
}
} else {
loadErrors.push(new Error(`Unsupported architecture on OpenHarmony: ${process.arch}`))
}
} else {
loadErrors.push(new Error(`Unsupported OS: ${process.platform}, architecture: ${process.arch}`))
}
}
nativeBinding = requireNative()
if (!nativeBinding || process.env.NAPI_RS_FORCE_WASI) {
let wasiBinding = null
let wasiBindingError = null
try {
wasiBinding = require('./pdf-inspector.wasi.cjs')
nativeBinding = wasiBinding
} catch (err) {
if (process.env.NAPI_RS_FORCE_WASI) {
wasiBindingError = err
}
}
if (!nativeBinding || process.env.NAPI_RS_FORCE_WASI) {
try {
wasiBinding = require('firecrawl-pdf-inspector-wasm32-wasi')
nativeBinding = wasiBinding
} catch (err) {
if (process.env.NAPI_RS_FORCE_WASI) {
if (!wasiBindingError) {
wasiBindingError = err
} else {
wasiBindingError.cause = err
}
loadErrors.push(err)
}
}
}
if (process.env.NAPI_RS_FORCE_WASI === 'error' && !wasiBinding) {
const error = new Error('WASI binding not found and NAPI_RS_FORCE_WASI is set to error')
error.cause = wasiBindingError
throw error
}
}
if (!nativeBinding) {
if (loadErrors.length > 0) {
throw new Error(
`Cannot find native binding. ` +
`npm has a bug related to optional dependencies (https://github.com/npm/cli/issues/4828). ` +
'Please try `npm i` again after removing both package-lock.json and node_modules directory.',
{
cause: loadErrors.reduce((err, cur) => {
cur.cause = err
return cur
}),
},
)
}
throw new Error(`Failed to load native binding`)
}
module.exports = nativeBinding
module.exports.classifyPdf = nativeBinding.classifyPdf
module.exports.extractTextInRegions = nativeBinding.extractTextInRegions
+23 -4
View File
@@ -1,18 +1,36 @@
{
"name": "firecrawl-pdf-inspector",
"version": "0.2.2",
"name": "@firecrawl/pdf-inspector",
"version": "1.8.10",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
"bin": {
"pdf-inspector": "bin/pdf-inspector.mjs"
},
"license": "MIT",
"keywords": [
"pdf",
"pdf-extraction",
"pdf-parser",
"text-extraction",
"ocr",
"pdf-classification",
"napi",
"rust",
"firecrawl"
],
"files": [
"index.js",
"index.d.ts",
"*.node"
"*.node",
"bin/",
"README.md"
],
"repository": {
"type": "git",
"url": "https://github.com/firecrawl/pdf-inspector"
},
"homepage": "https://github.com/firecrawl/pdf-inspector",
"publishConfig": {
"access": "public"
},
@@ -20,7 +38,8 @@
"binaryName": "pdf-inspector",
"targets": [
"x86_64-unknown-linux-gnu",
"aarch64-apple-darwin"
"aarch64-apple-darwin",
"x86_64-pc-windows-msvc"
],
"package": {
"name": "@firecrawl/pdf-inspector-js"
+29
View File
@@ -0,0 +1,29 @@
import { readFileSync } from "node:fs";
import { createRequire } from "node:module";
const require = createRequire(import.meta.url);
const { detectVectorGridInRegion } = require("./index.js");
const pdfPath =
process.argv[2] ?? "/tmp/pdf_inspector_indent_fixtures/cis_edge_benchmark.pdf";
const pdf = readFileSync(pdfPath);
const dpi = Number(process.argv[3] ?? 200);
const crops = [
{ pageIdx: 29, box: [0, 0, 612, 792], label: "page30-full" },
{ pageIdx: 16, box: [0, 0, 612, 792], label: "page17-full" },
{ pageIdx: 23, box: [0, 0, 612, 792], label: "page24-full" },
];
for (const { pageIdx, box, label } of crops) {
const result = detectVectorGridInRegion(pdf, pageIdx, box, dpi);
if (!result) {
console.log(`${label}: null`);
continue;
}
const rows = result.structureTokens.filter((token) => token === "<tr>").length;
const cols = rows > 0 ? result.cellBboxes.length / rows : 0;
console.log(
`${label}: cells=${result.cellBboxes.length} rows=${rows} cols=${cols}`,
);
}
+608 -69
View File
@@ -2,58 +2,276 @@
use napi::bindgen_prelude::*;
use napi_derive::napi;
use std::collections::HashSet;
use std::panic;
// ---------------------------------------------------------------------------
// Enums
// ---------------------------------------------------------------------------
/// PDF document type classification.
#[napi(string_enum)]
pub enum PdfType {
TextBased,
Scanned,
ImageBased,
Mixed,
}
/// Type of a positioned text item.
#[napi(string_enum)]
pub enum ItemType {
Text,
Image,
Link,
FormField,
}
// ---------------------------------------------------------------------------
// Result types
// ---------------------------------------------------------------------------
/// Full PDF processing result with markdown and metadata.
#[napi(object)]
pub struct PdfResult {
pub pdf_type: PdfType,
pub markdown: Option<String>,
pub page_count: u32,
pub processing_time_ms: u32,
/// 1-indexed page numbers that need OCR.
pub pages_needing_ocr: Vec<u32>,
pub title: Option<String>,
pub confidence: f64,
pub is_complex_layout: bool,
pub pages_with_tables: Vec<u32>,
pub pages_with_columns: Vec<u32>,
pub has_encoding_issues: bool,
}
/// Lightweight PDF classification result.
#[napi(object)]
pub struct PdfClassification {
pub pdf_type: String,
pub page_count: u32,
pub pages_needing_ocr: Vec<u32>,
pub confidence: f64,
pub pdf_type: PdfType,
pub page_count: u32,
/// 0-indexed page numbers that need OCR.
pub pages_needing_ocr: Vec<u32>,
pub confidence: f64,
}
/// A positioned text item extracted from a PDF.
#[napi(object)]
pub struct TextItem {
pub text: String,
pub x: f64,
pub y: f64,
pub width: f64,
pub height: f64,
pub font: String,
pub font_size: f64,
pub page: u32,
pub is_bold: bool,
pub is_italic: bool,
pub item_type: ItemType,
/// URL for link items, `None` for other types.
pub link_url: Option<String>,
}
/// A page's regions for text extraction: (page_index_0based, bboxes).
#[napi(object)]
pub struct PageRegions {
pub page: u32,
/// Each bbox is [x1, y1, x2, y2] in PDF points, top-left origin.
pub regions: Vec<Vec<f64>>,
pub page: u32,
/// Each bbox is [x1, y1, x2, y2] in PDF points, top-left origin.
pub regions: Vec<Vec<f64>>,
}
/// Extracted text for a single region.
#[napi(object)]
pub struct RegionText {
pub text: String,
/// `true` when the text should not be trusted (empty, GID fonts, garbage, encoding issues).
pub needs_ocr: bool,
pub text: String,
/// `true` when the text should not be trusted (empty, GID fonts, garbage, encoding issues).
pub needs_ocr: bool,
}
/// Extracted text for one page's regions.
#[napi(object)]
pub struct PageRegionTexts {
pub page: u32,
pub regions: Vec<RegionText>,
pub page: u32,
pub regions: Vec<RegionText>,
}
/// Classify a PDF: detect type (TextBased/Scanned/Mixed/ImageBased),
/// page count, and which pages need OCR. Takes PDF bytes as Buffer.
/// Vector-grid detection result compatible with `extractTablesWithStructure*`.
#[napi(object)]
pub struct VectorGridDetectionJs {
pub structure_tokens: Vec<String>,
pub cell_bboxes: Vec<Vec<f64>>,
}
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
fn convert_pdf_type(t: pdf_inspector::PdfType) -> PdfType {
match t {
pdf_inspector::PdfType::TextBased => PdfType::TextBased,
pdf_inspector::PdfType::Scanned => PdfType::Scanned,
pdf_inspector::PdfType::ImageBased => PdfType::ImageBased,
pdf_inspector::PdfType::Mixed => PdfType::Mixed,
}
}
fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
PdfResult {
pdf_type: convert_pdf_type(r.pdf_type),
markdown: r.markdown,
page_count: r.page_count,
processing_time_ms: r.processing_time_ms as u32,
pages_needing_ocr: r.pages_needing_ocr,
title: r.title,
confidence: r.confidence as f64,
is_complex_layout: r.layout.is_complex,
pages_with_tables: r.layout.pages_with_tables,
pages_with_columns: r.layout.pages_with_columns,
has_encoding_issues: r.has_encoding_issues,
}
}
fn convert_item_type(t: &pdf_inspector::types::ItemType) -> (ItemType, Option<String>) {
match t {
pdf_inspector::types::ItemType::Text => (ItemType::Text, None),
pdf_inspector::types::ItemType::Image => (ItemType::Image, None),
pdf_inspector::types::ItemType::Link(url) => (ItemType::Link, Some(url.clone())),
pdf_inspector::types::ItemType::FormField => (ItemType::FormField, None),
}
}
fn to_napi_err(e: impl std::fmt::Display, ctx: &str) -> Error {
Error::new(Status::GenericFailure, format!("{ctx}: {e}"))
}
/// Run a closure, catching any Rust panic and converting it to a NAPI error.
/// Prevents process abort from unwind panics in the native module.
fn catch_panic<F, T>(ctx: &str, f: F) -> Result<T>
where
F: FnOnce() -> Result<T> + panic::UnwindSafe,
{
match panic::catch_unwind(f) {
Ok(result) => result,
Err(payload) => {
let msg = if let Some(s) = payload.downcast_ref::<&str>() {
s.to_string()
} else if let Some(s) = payload.downcast_ref::<String>() {
s.clone()
} else {
"unknown panic".to_string()
};
Err(Error::new(
Status::GenericFailure,
format!("{ctx}: Rust panic: {msg}"),
))
}
}
}
// ---------------------------------------------------------------------------
// Public NAPI API
// ---------------------------------------------------------------------------
/// Process a PDF from a Buffer: detect type, extract text, and convert to Markdown.
#[napi]
pub fn process_pdf(buffer: Buffer, pages: Option<Vec<u32>>) -> Result<PdfResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("process_pdf", move || {
let mut opts = pdf_inspector::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = pdf_inspector::process_pdf_mem_with_options(&bytes, opts)
.map_err(|e| to_napi_err(e, "process_pdf"))?;
Ok(to_napi_result(result))
})
}
/// Fast detection only — no text extraction or markdown.
#[napi]
pub fn detect_pdf(buffer: Buffer) -> Result<PdfResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("detect_pdf", move || {
let result =
pdf_inspector::detect_pdf_mem(&bytes).map_err(|e| to_napi_err(e, "detect_pdf"))?;
Ok(to_napi_result(result))
})
}
/// Lightweight PDF classification — returns type, page count, and OCR pages.
/// Faster than detectPdf as it skips building the full PdfResult.
/// Pages in pagesNeedingOcr are 0-indexed.
#[napi]
pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
let result = pdf_inspector::classify_pdf_mem(&buffer).map_err(|e| {
Error::new(Status::GenericFailure, format!("classify_pdf failed: {e}"))
})?;
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("classify_pdf", move || {
let result =
pdf_inspector::classify_pdf_mem(&bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
Ok(PdfClassification {
pdf_type: convert_pdf_type(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
})
}
Ok(PdfClassification {
pdf_type: match result.pdf_type {
pdf_inspector::PdfType::TextBased => "TextBased".to_string(),
pdf_inspector::PdfType::Scanned => "Scanned".to_string(),
pdf_inspector::PdfType::ImageBased => "ImageBased".to_string(),
pdf_inspector::PdfType::Mixed => "Mixed".to_string(),
},
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
/// Extract plain text from a PDF Buffer.
#[napi]
pub fn extract_text(buffer: Buffer) -> Result<String> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_text", move || {
pdf_inspector::extractor::extract_text_mem(&bytes)
.map_err(|e| to_napi_err(e, "extract_text"))
})
}
/// Extract text with position information from a PDF Buffer.
#[napi]
pub fn extract_text_with_positions(
buffer: Buffer,
pages: Option<Vec<u32>>,
) -> Result<Vec<TextItem>> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_text_with_positions", move || {
let items = match pages {
Some(p) => {
let page_set: HashSet<u32> = p.into_iter().collect();
pdf_inspector::extractor::extract_text_with_positions_mem_pages(
&bytes,
Some(&page_set),
)
.map_err(|e| to_napi_err(e, "extract_text_with_positions"))?
}
None => pdf_inspector::extractor::extract_text_with_positions_mem(&bytes)
.map_err(|e| to_napi_err(e, "extract_text_with_positions"))?,
};
Ok(items
.into_iter()
.map(|item| {
let (item_type, link_url) = convert_item_type(&item.item_type);
TextItem {
text: item.text,
x: item.x as f64,
y: item.y as f64,
width: item.width as f64,
height: item.height as f64,
font: item.font,
font_size: item.font_size as f64,
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
item_type,
link_url,
}
})
.collect())
})
}
/// Extract text within bounding-box regions from a PDF.
@@ -62,55 +280,376 @@ pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
/// this extracts PDF text within those regions — skipping GPU OCR
/// for text-based pages.
///
/// Each region result includes `needs_ocr` — set when the extracted text
/// Each region result includes `needsOcr` — set when the extracted text
/// is unreliable (empty, GID-encoded fonts, garbage, encoding issues).
///
/// Coordinates are PDF points with top-left origin.
#[napi]
pub fn extract_text_in_regions(
buffer: Buffer,
page_regions: Vec<PageRegions>,
buffer: Buffer,
page_regions: Vec<PageRegions>,
) -> Result<Vec<PageRegionTexts>> {
// Convert from napi types to the Rust API's expected format
let regions: Vec<(u32, Vec<[f32; 4]>)> = page_regions
.iter()
.map(|pr| {
let bboxes: Vec<[f32; 4]> = pr
.regions
.iter()
.map(|r| {
if r.len() != 4 {
[0.0, 0.0, 0.0, 0.0]
} else {
[r[0] as f32, r[1] as f32, r[2] as f32, r[3] as f32]
}
})
.collect();
(pr.page, bboxes)
let bytes: Vec<u8> = buffer.to_vec();
let regions = parse_page_regions(&page_regions);
catch_panic("extract_text_in_regions", move || {
let results = pdf_inspector::extract_text_in_regions_mem(&bytes, &regions)
.map_err(|e| to_napi_err(e, "extract_text_in_regions"))?;
Ok(to_page_region_texts(results))
})
.collect();
}
let results = pdf_inspector::extract_text_in_regions_mem(&buffer, &regions).map_err(|e| {
Error::new(
Status::GenericFailure,
format!("extract_text_in_regions failed: {e}"),
)
})?;
/// Extract markdown tables within bounding-box regions from a PDF.
///
/// Like `extractTextInRegions` but runs table detection on items within each
/// region and returns markdown pipe-tables instead of flat text.
///
/// When table structure is detected, `text` contains a markdown pipe-table and
/// `needsOcr` is `false`. When no table is found, `text` is empty and
/// `needsOcr` is `true` so the caller can fall back to GPU OCR.
///
/// Coordinates are PDF points with top-left origin.
#[napi]
pub fn extract_tables_in_regions(
buffer: Buffer,
page_regions: Vec<PageRegions>,
) -> Result<Vec<PageRegionTexts>> {
let bytes: Vec<u8> = buffer.to_vec();
let regions = parse_page_regions(&page_regions);
Ok(
catch_panic("extract_tables_in_regions", move || {
let results = pdf_inspector::extract_tables_in_regions_mem(&bytes, &regions)
.map_err(|e| to_napi_err(e, "extract_tables_in_regions"))?;
Ok(to_page_region_texts(results))
})
}
/// Detect a vector ruled-line / rectangle grid inside one page region.
///
/// Returns TSR-compatible structure tokens plus crop-pixel cell bboxes, or
/// `null` when the region does not contain a valid vector grid.
///
/// `pageIdx` is 0-indexed. `regionPdfPtBbox` is `[x1,y1,x2,y2]` in PDF
/// points with top-left origin. `renderDpi` is the DPI of the crop image that
/// will consume the returned cell bboxes.
#[napi]
pub fn detect_vector_grid_in_region(
buffer: Buffer,
page_idx: u32,
region_pdf_pt_bbox: Vec<f64>,
render_dpi: f64,
) -> Result<Option<VectorGridDetectionJs>> {
let bytes: Vec<u8> = buffer.to_vec();
let region = if region_pdf_pt_bbox.len() == 4 {
[
region_pdf_pt_bbox[0] as f32,
region_pdf_pt_bbox[1] as f32,
region_pdf_pt_bbox[2] as f32,
region_pdf_pt_bbox[3] as f32,
]
} else {
[0.0, 0.0, 0.0, 0.0]
};
catch_panic("detect_vector_grid_in_region", move || {
let result = pdf_inspector::detect_vector_grid_in_region_mem(
&bytes,
page_idx,
region,
render_dpi as f32,
)
.map_err(|e| to_napi_err(e, "detect_vector_grid_in_region"))?;
Ok(result.map(|r| VectorGridDetectionJs {
structure_tokens: r.structure_tokens,
cell_bboxes: r
.cell_bboxes
.into_iter()
.map(|bbox| bbox.into_iter().map(|v| v as f64).collect())
.collect(),
}))
})
}
/// One cropped table region plus its raw structure-recovery output, for
/// `extractTablesWithStructure`.
///
/// `structureTokens` and `cellBboxes` are typically produced by an external
/// table-structure recognition model (e.g. SLANet on PaddleOCR) running on
/// a rendered crop of the page. pdf-inspector uses the structure to lay out
/// the cells and pulls the cell text from the native PDF — no OCR involved.
#[napi(object)]
pub struct TsrTableInputJs {
/// 0-indexed page number where the crop was taken from.
pub page: u32,
/// Crop bbox on the page, `[x1, y1, x2, y2]` in PDF points with
/// top-left origin.
pub crop_pdf_pt_bbox: Vec<f64>,
/// DPI the crop image was rendered at (e.g. `200.0`).
pub render_dpi: f64,
/// Raw structure tokens emitted by the TSR model, in document order.
pub structure_tokens: Vec<String>,
/// One bbox per cell (in document order). May be 4-element
/// `[x1,y1,x2,y2]` or 8-element 4-corner polygon, in crop image-pixel
/// space.
pub cell_bboxes: Vec<Vec<f64>>,
}
/// Extract markdown tables using externally-supplied structure recovery.
///
/// For each input, pairs structure tokens with cell bboxes (rowspan/colspan
/// aware), converts each cell bbox from crop image-pixels into page PDF
/// points, pulls the cell's text from the native PDF, and emits a markdown
/// pipe-table.
///
/// Returns one markdown string per input, in input order.
#[napi]
pub fn extract_tables_with_structure(
buffer: Buffer,
inputs: Vec<TsrTableInputJs>,
) -> Result<Vec<String>> {
let bytes: Vec<u8> = buffer.to_vec();
let parsed = parse_tsr_inputs(&inputs);
catch_panic("extract_tables_with_structure", move || {
pdf_inspector::extract_tables_with_structure_mem(&bytes, &parsed)
.map_err(|e| to_napi_err(e, "extract_tables_with_structure"))
})
}
/// One resolved cell from `extractTablesWithStructureCells`.
#[napi(object)]
pub struct StructuredCellJs {
/// 0-indexed grid row.
pub row: u32,
/// 0-indexed grid column.
pub col: u32,
/// 1 for a normal cell.
pub rowspan: u32,
/// 1 for a normal cell.
pub colspan: u32,
/// `true` when the cell is a `<th>` or sits inside `<thead>`.
pub is_header: bool,
/// Text extracted from the native PDF for this cell (may be empty).
pub text: String,
/// Axis-aligned bbox `[x1, y1, x2, y2]` in page PDF-points, top-left
/// origin. Useful for debug overlays or per-cell post-processing.
pub page_pt_bbox: Vec<f64>,
}
/// Extract structured cells using externally-supplied structure recovery.
///
/// Lower-level sibling of [`extractTablesWithStructure`]: instead of
/// rendering markdown, returns the resolved cells (row, col, rowspan,
/// colspan, isHeader, text, pagePtBbox) so callers can drive their own
/// rendering, debug overlays, or per-cell post-processing.
///
/// Returns one `Array<StructuredCellJs>` per input, in input order.
#[napi]
pub fn extract_tables_with_structure_cells(
buffer: Buffer,
inputs: Vec<TsrTableInputJs>,
) -> Result<Vec<Vec<StructuredCellJs>>> {
let bytes: Vec<u8> = buffer.to_vec();
let parsed = parse_tsr_inputs(&inputs);
catch_panic("extract_tables_with_structure_cells", move || {
let result = pdf_inspector::extract_tables_with_structure_cells_mem(&bytes, &parsed)
.map_err(|e| to_napi_err(e, "extract_tables_with_structure_cells"))?;
Ok(result
.into_iter()
.map(|cells| {
cells
.into_iter()
.map(|c| StructuredCellJs {
row: c.row as u32,
col: c.col as u32,
rowspan: c.rowspan as u32,
colspan: c.colspan as u32,
is_header: c.is_header,
text: c.text,
page_pt_bbox: c.page_pt_bbox.iter().map(|v| *v as f64).collect(),
})
.collect()
})
.collect())
})
}
/// One result from `extractTablesWithStructureAuto` — markdown plus a
/// diagnostic flag identifying which path produced it.
///
/// `fallbackReason` is `null` when the TSR-hybrid path produced the
/// markdown directly. When stage 1's quality check fires (the cells
/// look like a SLANet detection pathology — phantom rows or multi-row
/// content in a single cell), the auto path may expand the TSR cells
/// in-place or run the heuristic table extractor on the same region.
/// `fallbackReason` carries the diagnostic label (for example
/// `"multi_row_in_cell_expanded"` or `"phantom_empty_row"`).
#[napi(object)]
pub struct TableExtractionResultJs {
pub markdown: String,
pub fallback_reason: Option<String>,
}
/// Auto-fallback variant of [`extractTablesWithStructure`].
///
/// Runs the TSR-hybrid path, checks the resulting cells for known
/// SLANet detection pathologies, expands multi-row cells in-place when
/// possible, and otherwise falls back to the heuristic
/// `extractTablesInRegions` for inputs where the TSR path looks
/// compromised.
///
/// On clean inputs this returns identical markdown to
/// `extractTablesWithStructure`; on flagged inputs `fallbackReason` is
/// set to the recovery path that produced the result.
#[napi]
pub fn extract_tables_with_structure_auto(
buffer: Buffer,
inputs: Vec<TsrTableInputJs>,
) -> Result<Vec<TableExtractionResultJs>> {
let bytes: Vec<u8> = buffer.to_vec();
let parsed = parse_tsr_inputs(&inputs);
catch_panic("extract_tables_with_structure_auto", move || {
let result = pdf_inspector::extract_tables_with_structure_auto_mem(&bytes, &parsed)
.map_err(|e| to_napi_err(e, "extract_tables_with_structure_auto"))?;
Ok(result
.into_iter()
.map(|r| TableExtractionResultJs {
markdown: r.markdown,
fallback_reason: r.fallback_reason,
})
.collect())
})
}
fn parse_tsr_inputs(inputs: &[TsrTableInputJs]) -> Vec<pdf_inspector::TsrTableInput> {
inputs
.iter()
.map(|i| {
let crop = if i.crop_pdf_pt_bbox.len() == 4 {
[
i.crop_pdf_pt_bbox[0] as f32,
i.crop_pdf_pt_bbox[1] as f32,
i.crop_pdf_pt_bbox[2] as f32,
i.crop_pdf_pt_bbox[3] as f32,
]
} else {
[0.0, 0.0, 0.0, 0.0]
};
let cell_bboxes: Vec<Vec<f32>> = i
.cell_bboxes
.iter()
.map(|bb| bb.iter().map(|v| *v as f32).collect())
.collect();
pdf_inspector::TsrTableInput {
page: i.page,
crop_pdf_pt_bbox: crop,
render_dpi: i.render_dpi as f32,
structure_tokens: i.structure_tokens.clone(),
cell_bboxes,
}
})
.collect()
}
/// Per-page markdown extraction result.
#[napi(object)]
pub struct PageMarkdownResult {
/// 0-indexed page number.
pub page: u32,
/// Formatted markdown for this page.
pub markdown: String,
/// `true` when text on this page is unreliable.
pub needs_ocr: bool,
}
/// Combined per-page markdown extraction and layout classification result.
#[napi(object)]
pub struct PagesExtractionResult {
/// Per-page markdown results.
pub pages: Vec<PageMarkdownResult>,
/// 1-indexed pages where tables were detected.
pub pages_with_tables: Vec<u32>,
/// 1-indexed pages where multi-column layout was detected.
pub pages_with_columns: Vec<u32>,
/// 1-indexed pages that need OCR (scanned/image-based).
pub pages_needing_ocr: Vec<u32>,
/// True if any page has tables or columns.
pub is_complex: bool,
}
/// Extract formatted markdown for pages of a PDF, with layout classification
/// metadata.
///
/// Returns per-page markdown and classification data (tables, columns,
/// OCR needs) from a single parse. Font statistics are computed from the
/// full document so header detection is consistent across pages.
///
/// Omit `pages` (or pass `undefined`) to return every page in document
/// order. Pass an array of 0-indexed page numbers to restrict output to
/// those pages, in caller-supplied order.
#[napi]
pub fn extract_pages_markdown(
buffer: Buffer,
pages: Option<Vec<u32>>,
) -> Result<PagesExtractionResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_pages_markdown", move || {
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, pages.as_deref())
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
.pages
.into_iter()
.map(|r| PageMarkdownResult {
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
is_complex: result.is_complex,
})
})
}
fn parse_page_regions(page_regions: &[PageRegions]) -> Vec<(u32, Vec<[f32; 4]>)> {
page_regions
.iter()
.map(|pr| {
let bboxes: Vec<[f32; 4]> = pr
.regions
.iter()
.map(|r| {
if r.len() != 4 {
[0.0, 0.0, 0.0, 0.0]
} else {
[r[0] as f32, r[1] as f32, r[2] as f32, r[3] as f32]
}
})
.collect();
(pr.page, bboxes)
})
.collect()
}
fn to_page_region_texts(results: Vec<pdf_inspector::PageRegionResult>) -> Vec<PageRegionTexts> {
results
.into_iter()
.map(|page_result| PageRegionTexts {
page: page_result.page,
regions: page_result
.regions
.into_iter()
.map(|r| RegionText {
text: r.text,
needs_ocr: r.needs_ocr,
})
.collect(),
})
.collect(),
)
.into_iter()
.map(|page_result| PageRegionTexts {
page: page_result.page,
regions: page_result
.regions
.into_iter()
.map(|r| RegionText {
text: r.text,
needs_ocr: r.needs_ocr,
})
.collect(),
})
.collect()
}
+133
View File
@@ -0,0 +1,133 @@
import { readFileSync } from 'fs';
import { strict as assert } from 'assert';
import {
processPdf,
detectPdf,
classifyPdf,
extractText,
extractTextWithPositions,
extractTextInRegions,
detectVectorGridInRegion,
extractPagesMarkdown,
} from './index.js';
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
// --- processPdf ---
console.log('Testing processPdf...');
const result = processPdf(fixture);
assert.equal(result.pdfType, 'TextBased');
assert.equal(result.pageCount, 3);
assert.ok(result.confidence > 0);
assert.ok(result.markdown && result.markdown.length > 0);
assert.equal(typeof result.isComplexLayout, 'boolean');
assert.ok(Array.isArray(result.pagesWithTables));
assert.ok(Array.isArray(result.pagesWithColumns));
assert.equal(typeof result.hasEncodingIssues, 'boolean');
console.log(' processPdf: OK');
// processPdf with pages
const result2 = processPdf(fixture, [1]);
assert.ok(result2.markdown && result2.markdown.length > 0);
console.log(' processPdf with pages: OK');
// --- detectPdf ---
console.log('Testing detectPdf...');
const detected = detectPdf(fixture);
assert.equal(detected.pdfType, 'TextBased');
assert.equal(detected.pageCount, 3);
assert.equal(detected.markdown, undefined);
console.log(' detectPdf: OK');
// --- classifyPdf ---
console.log('Testing classifyPdf...');
const classified = classifyPdf(fixture);
assert.equal(classified.pdfType, 'TextBased');
assert.equal(classified.pageCount, 3);
assert.ok(classified.confidence > 0);
assert.ok(Array.isArray(classified.pagesNeedingOcr));
console.log(' classifyPdf: OK');
// --- extractText ---
console.log('Testing extractText...');
const text = extractText(fixture);
assert.equal(typeof text, 'string');
assert.ok(text.length > 0);
console.log(' extractText: OK');
// --- extractTextWithPositions ---
console.log('Testing extractTextWithPositions...');
const items = extractTextWithPositions(fixture);
assert.ok(items.length > 0);
const item = items[0];
assert.equal(typeof item.text, 'string');
assert.equal(typeof item.x, 'number');
assert.equal(typeof item.y, 'number');
assert.equal(typeof item.width, 'number');
assert.equal(typeof item.height, 'number');
assert.equal(typeof item.font, 'string');
assert.equal(typeof item.fontSize, 'number');
assert.equal(typeof item.page, 'number');
assert.equal(typeof item.isBold, 'boolean');
assert.equal(typeof item.isItalic, 'boolean');
assert.equal(typeof item.itemType, 'string');
console.log(' extractTextWithPositions: OK');
// with pages filter
const page1Items = extractTextWithPositions(fixture, [1]);
assert.ok(page1Items.length > 0);
assert.ok(page1Items.every(i => i.page === 1));
console.log(' extractTextWithPositions with pages: OK');
// --- extractTextInRegions ---
console.log('Testing extractTextInRegions...');
const regionResults = extractTextInRegions(fixture, [
{ page: 0, regions: [[0, 0, 600, 100]] },
]);
assert.equal(regionResults.length, 1);
assert.equal(regionResults[0].page, 0);
assert.equal(regionResults[0].regions.length, 1);
assert.equal(typeof regionResults[0].regions[0].text, 'string');
assert.equal(typeof regionResults[0].regions[0].needsOcr, 'boolean');
console.log(' extractTextInRegions: OK');
// --- detectVectorGridInRegion ---
console.log('Testing detectVectorGridInRegion...');
const vectorGrid = detectVectorGridInRegion(fixture, 0, [0, 0, 600, 800], 72);
assert.ok(vectorGrid === null || typeof vectorGrid === 'object');
if (vectorGrid) {
assert.ok(Array.isArray(vectorGrid.structureTokens));
assert.ok(Array.isArray(vectorGrid.cellBboxes));
assert.ok(vectorGrid.cellBboxes.every(bbox => Array.isArray(bbox) && bbox.length === 4));
}
console.log(' detectVectorGridInRegion: OK');
// --- extractPagesMarkdown ---
console.log('Testing extractPagesMarkdown...');
// omit pages → every page in document order
const allPages = extractPagesMarkdown(fixture);
assert.equal(allPages.pages.length, 3);
assert.deepEqual(allPages.pages.map(p => p.page), [0, 1, 2]);
assert.ok(typeof allPages.pages[0].markdown === 'string');
assert.equal(typeof allPages.pages[0].needsOcr, 'boolean');
assert.ok(Array.isArray(allPages.pagesWithTables));
assert.ok(Array.isArray(allPages.pagesWithColumns));
assert.ok(Array.isArray(allPages.pagesNeedingOcr));
assert.equal(typeof allPages.isComplex, 'boolean');
console.log(' extractPagesMarkdown (no pages arg): OK');
// selected pages preserve caller order
const picked = extractPagesMarkdown(fixture, [2, 0]);
assert.equal(picked.pages.length, 2);
assert.equal(picked.pages[0].page, 2);
assert.equal(picked.pages[1].page, 0);
console.log(' extractPagesMarkdown with pages: OK');
// --- Error handling ---
console.log('Testing error handling...');
assert.throws(() => processPdf(Buffer.from('not a pdf')), /process_pdf/);
assert.throws(() => classifyPdf(Buffer.from('')), /classify_pdf/);
console.log(' error handling: OK');
console.log('\nAll NAPI tests passed!');
+167
View File
@@ -0,0 +1,167 @@
"""Type stubs for pdf_inspector."""
from typing import Optional
class PdfResult:
"""Result of processing a PDF file."""
pdf_type: str
"""'text_based', 'scanned', 'image_based', or 'mixed'."""
markdown: Optional[str]
page_count: int
processing_time_ms: int
pages_needing_ocr: list[int]
title: Optional[str]
confidence: float
is_complex_layout: bool
pages_with_tables: list[int]
pages_with_columns: list[int]
has_encoding_issues: bool
class PdfClassification:
"""Lightweight PDF classification result."""
pdf_type: str
"""'text_based', 'scanned', 'image_based', or 'mixed'."""
page_count: int
pages_needing_ocr: list[int]
"""0-indexed page numbers that need OCR."""
confidence: float
class TextItem:
"""A positioned text item extracted from a PDF."""
text: str
x: float
y: float
width: float
height: float
font: str
font_size: float
page: int
is_bold: bool
is_italic: bool
item_type: str
class RegionText:
"""Extracted text for a single region."""
text: str
needs_ocr: bool
"""True when the text should not be trusted."""
class PageRegionTexts:
"""Extracted text for one page's regions."""
page: int
"""0-indexed page number."""
regions: list[RegionText]
class PageMarkdown:
"""Per-page markdown extraction result."""
page: int
"""0-indexed page number."""
markdown: str
"""Formatted markdown for this page (empty string when needs_ocr is True)."""
needs_ocr: bool
"""True when text on this page is unreliable and OCR should be used instead."""
class PagesExtractionResult:
"""Per-page markdown output with document-wide layout classification."""
pages: list[PageMarkdown]
"""Per-page markdown results, in the order requested."""
pages_with_tables: list[int]
"""1-indexed pages where tables were detected."""
pages_with_columns: list[int]
"""1-indexed pages where multi-column layout was detected."""
pages_needing_ocr: list[int]
"""1-indexed pages that need OCR."""
is_complex: bool
"""True if any page has tables or multi-column layout."""
def process_pdf(path: str, pages: Optional[list[int]] = None) -> PdfResult:
"""Process a PDF: detect type, extract text, convert to Markdown."""
...
def process_pdf_bytes(data: bytes, pages: Optional[list[int]] = None) -> PdfResult:
"""Process a PDF from bytes in memory."""
...
def detect_pdf(path: str) -> PdfResult:
"""Fast detection only — no text extraction."""
...
def detect_pdf_bytes(data: bytes) -> PdfResult:
"""Fast detection from bytes."""
...
def classify_pdf(path: str) -> PdfClassification:
"""Lightweight classification — type, page count, and OCR pages (0-indexed)."""
...
def classify_pdf_bytes(data: bytes) -> PdfClassification:
"""Lightweight classification from bytes."""
...
def extract_text(path: str) -> str:
"""Extract plain text from a PDF."""
...
def extract_text_bytes(data: bytes) -> str:
"""Extract plain text from PDF bytes."""
...
def extract_text_with_positions(path: str, pages: Optional[list[int]] = None) -> list[TextItem]:
"""Extract text with position information."""
...
def extract_text_with_positions_bytes(data: bytes, pages: Optional[list[int]] = None) -> list[TextItem]:
"""Extract text with position information from bytes."""
...
def extract_text_in_regions(
path: str,
page_regions: list[tuple[int, list[list[float]]]],
) -> list[PageRegionTexts]:
"""Extract text within bounding-box regions from a PDF file.
Args:
path: Path to the PDF file.
page_regions: List of (page_0indexed, [[x1, y1, x2, y2], ...]) tuples.
"""
...
def extract_text_in_regions_bytes(
data: bytes,
page_regions: list[tuple[int, list[list[float]]]],
) -> list[PageRegionTexts]:
"""Extract text within bounding-box regions from PDF bytes.
Args:
data: PDF file contents as bytes.
page_regions: List of (page_0indexed, [[x1, y1, x2, y2], ...]) tuples.
"""
...
def extract_pages_markdown(
path: str,
pages: Optional[list[int]] = None,
) -> PagesExtractionResult:
"""Extract formatted markdown for pages of a PDF, with layout classification.
Args:
path: Path to the PDF file.
pages: Optional list of 0-indexed pages. When ``None`` (default), every
page is returned in document order. Otherwise, output matches the
caller-supplied order.
Returns:
PagesExtractionResult with per-page markdown and document-wide layout
classification (tables, columns, OCR needs).
"""
...
def extract_pages_markdown_bytes(
data: bytes,
pages: Optional[list[int]] = None,
) -> PagesExtractionResult:
"""Extract formatted markdown for pages of a PDF from bytes.
See :func:`extract_pages_markdown` for details.
"""
...
+21
View File
@@ -0,0 +1,21 @@
[build-system]
requires = ["maturin>=1.0,<2.0"]
build-backend = "maturin"
[project]
name = "pdf-inspector"
version = "0.1.0"
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
license = { text = "MIT" }
requires-python = ">=3.8"
classifiers = [
"Programming Language :: Rust",
"Programming Language :: Python :: Implementation :: CPython",
"Programming Language :: Python :: 3",
"License :: OSI Approved :: MIT License",
"Operating System :: OS Independent",
"Topic :: Text Processing",
]
[tool.maturin]
features = ["python"]
+33 -11
View File
@@ -1,8 +1,12 @@
//! CLI tool for detecting PDF type (text-based vs scanned)
use pdf_inspector::{detect_pdf_type, process_pdf_with_options, PdfOptions, PdfType, ProcessMode};
use pdf_inspector::{
detect_pdf_type, detector::estimate_page_count_from_bytes, process_pdf_with_options,
PdfOptions, PdfType, ProcessMode,
};
use std::env;
use std::fmt::Write;
use std::fs;
use std::process;
use std::time::Instant;
@@ -64,6 +68,32 @@ fn pdf_type_str(pdf_type: &PdfType) -> &'static str {
}
}
fn page_count_hint(pdf_path: &str) -> Option<u32> {
fs::read(pdf_path)
.ok()
.map(|bytes| estimate_page_count_from_bytes(&bytes))
.filter(|&count| count > 0)
}
fn print_error(e: &pdf_inspector::PdfError, pdf_path: &str, json_output: bool) {
if json_output {
if let Some(count) = page_count_hint(pdf_path) {
println!(
r#"{{"error":"{}","page_count_hint":{}}}"#,
json_escape(&e.to_string()),
count
);
} else {
println!(r#"{{"error":"{}"}}"#, json_escape(&e.to_string()));
}
} else {
eprintln!("Error: {}", e);
if let Some(count) = page_count_hint(pdf_path) {
eprintln!("Page count hint: {}", count);
}
}
}
fn run_analyze(pdf_path: &str, json_output: bool, start: Instant) {
match process_pdf_with_options(pdf_path, PdfOptions::new().mode(ProcessMode::Analyze)) {
Ok(result) => {
@@ -135,11 +165,7 @@ fn run_analyze(pdf_path: &str, json_output: bool, start: Instant) {
}
}
Err(e) => {
if json_output {
println!(r#"{{"error":"{}"}}"#, e);
} else {
eprintln!("Error: {}", e);
}
print_error(&e, pdf_path, json_output);
process::exit(1);
}
}
@@ -236,11 +262,7 @@ fn run_detect_only(pdf_path: &str, json_output: bool, start: Instant) {
}
}
Err(e) => {
if json_output {
println!(r#"{{"error":"{}"}}"#, e);
} else {
eprintln!("Error: {}", e);
}
print_error(&e, pdf_path, json_output);
process::exit(1);
}
}
+2238 -133
View File
File diff suppressed because it is too large Load Diff
+43 -13
View File
@@ -82,7 +82,7 @@ pub(crate) fn extract_page_text_items(
page_num: u32,
font_cmaps: &FontCMaps,
include_invisible: bool,
) -> Result<(PageExtraction, bool), PdfError> {
) -> Result<(PageExtraction, bool, bool), PdfError> {
use lopdf::content::Content;
let mut items = Vec::new();
@@ -182,7 +182,7 @@ pub(crate) fn extract_page_text_items(
content.operations.len(),
MAX_OPERATIONS
);
return Ok(((Vec::new(), Vec::new(), Vec::new()), false));
return Ok(((Vec::new(), Vec::new(), Vec::new()), false, false));
}
// Graphics state tracking
@@ -216,6 +216,7 @@ pub(crate) fn extract_page_text_items(
let mut marked_content_stack: Vec<MarkedContentEntry> = Vec::new();
let mut suppress_glyph_extraction = false;
let mut actual_text_start_tm: Option<[f32; 6]> = None; // text matrix at BDC entry
let mut actual_text_glyph_tm: Option<[f32; 6]> = None; // text matrix at first glyph inside BDC
/// Get the innermost MCID from the marked content stack.
fn current_mcid(stack: &[MarkedContentEntry]) -> Option<i64> {
stack.iter().rev().find_map(|e| e.mcid)
@@ -349,8 +350,15 @@ pub(crate) fn extract_page_text_items(
)
})
});
// ActualText: suppress glyph extraction, just advance text matrix
// ActualText: suppress glyph extraction, just advance text matrix.
// Capture the FIRST glyph's text matrix as the rendering position
// for the ActualText item. Td ops between BDC and the first Tj
// may have moved the position to the correct line — the BDC-entry
// position (actual_text_start_tm) can be on the previous line.
if suppress_glyph_extraction {
if actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
}
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
@@ -377,6 +385,7 @@ pub(crate) fn extract_page_text_items(
&font_encodings,
&encoding_cache,
&mut cmap_decisions,
&font_widths,
) {
let combined = multiply_matrices(&text_matrix, &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
@@ -425,6 +434,10 @@ pub(crate) fn extract_page_text_items(
let font_info = font_widths.get(&current_font);
let is_invisible = (text_rendering_mode == 3 && !include_invisible)
|| suppress_glyph_extraction;
// Capture first-glyph position for ActualText
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
}
// Compute space threshold based on font metrics when available
let space_threshold = if let Some(font_info) = font_info {
@@ -521,6 +534,7 @@ pub(crate) fn extract_page_text_items(
&font_encodings,
&encoding_cache,
&mut cmap_decisions,
&font_widths,
) {
current_text.push_str(&text);
}
@@ -608,6 +622,7 @@ pub(crate) fn extract_page_text_items(
&font_encodings,
&encoding_cache,
&mut cmap_decisions,
&font_widths,
) {
if !text.trim().is_empty() {
let combined = multiply_matrices(&text_matrix, &ctm);
@@ -700,6 +715,7 @@ pub(crate) fn extract_page_text_items(
if actual_text.is_some() {
suppress_glyph_extraction = true;
actual_text_start_tm = Some(text_matrix);
actual_text_glyph_tm = None; // reset — will be captured at first Tj/TJ
}
marked_content_stack.push(MarkedContentEntry { actual_text, mcid });
}
@@ -707,8 +723,13 @@ pub(crate) fn extract_page_text_items(
// End Marked Content — emit ActualText item with correct width
if let Some(entry) = marked_content_stack.pop() {
if let Some(at) = entry.actual_text {
// Compute width from text matrix advancement during BDC..EMC
if let Some(start_tm) = actual_text_start_tm.take() {
// Use the first-glyph position (if available) instead of the
// BDC-entry position. Td operators between BDC and the first
// Tj may have moved the text position to the correct line —
// the BDC-entry position can be on the previous line.
let glyph_tm = actual_text_glyph_tm.take();
let entry_tm = actual_text_start_tm.take();
if let Some(start_tm) = glyph_tm.or(entry_tm) {
let combined = multiply_matrices(&start_tm, &ctm);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
@@ -994,9 +1015,17 @@ pub(crate) fn extract_page_text_items(
// producing thousands of identical rects that yield a degenerate grid.
// After dedup, if too few unique clip rects remain we fall through to
// fill rects (explicitly drawn visible rectangles).
//
// When fill rects substantially outnumber clip rects, the clips are
// typically section-level wrappers and the fills are the actual table
// cell backgrounds (e.g. shaded-header tables drawn with `m`/`l`/`h`/`f*`
// sequences). In that case, prefer fills.
if rects.is_empty() {
dedup_rects(&mut clip_rects);
if clip_rects.len() >= 4 {
let prefer_fills = !fill_rects.is_empty() && fill_rects.len() >= clip_rects.len() * 3;
if prefer_fills {
rects = fill_rects;
} else if clip_rects.len() >= 4 {
rects = clip_rects;
} else if !fill_rects.is_empty() {
rects = fill_rects;
@@ -1009,11 +1038,12 @@ pub(crate) fn extract_page_text_items(
// Some PDFs embed landscape content in portrait pages using a rotated text
// matrix (e.g. [0, b, -b, 0, tx, ty] for 90° CCW). The layout engine
// assumes x=horizontal, y=vertical — so we swap coordinates to match.
let (items, rects, lines) = correct_rotated_page(items, rects, lines, &rotation_votes);
let (items, rects, lines, coords_rotated) =
correct_rotated_page(items, rects, lines, &rotation_votes);
let items = super::merge_text_items(items);
let items = super::merge_subscript_items(items);
Ok(((items, rects, lines), has_gid_fonts))
Ok(((items, rects, lines), has_gid_fonts, coords_rotated))
}
/// Counts of text operators with horizontal vs rotated combined matrices.
@@ -1030,9 +1060,9 @@ fn correct_rotated_page(
mut rects: Vec<PdfRect>,
mut lines: Vec<PdfLine>,
votes: &RotationVotes,
) -> (Vec<TextItem>, Vec<PdfRect>, Vec<PdfLine>) {
) -> (Vec<TextItem>, Vec<PdfRect>, Vec<PdfLine>, bool) {
if items.len() < 2 {
return (items, rects, lines);
return (items, rects, lines, false);
}
// Use the combined-matrix direction votes collected during extraction.
@@ -1041,7 +1071,7 @@ fn correct_rotated_page(
let total_votes = votes.horizontal + votes.rotated;
if total_votes == 0 || votes.rotated * 3 < total_votes * 2 {
// Less than ~67% of text operators are rotated → not a rotated page
return (items, rects, lines);
return (items, rects, lines, false);
}
log::debug!(
@@ -1092,7 +1122,7 @@ fn correct_rotated_page(
line.y2 = new_y2;
}
(items, rects, lines)
(items, rects, lines, true)
}
/// Remove near-duplicate rects (same coordinates within 0.5 pt tolerance).
@@ -1228,7 +1258,7 @@ mod tests {
let font_cmaps = FontCMaps::from_doc(&doc);
let result = extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let ((items, rects, lines), _has_gid) = result;
let ((items, rects, lines), _has_gid, _coords_rotated) = result;
assert!(items.is_empty());
assert!(rects.is_empty());
assert!(lines.is_empty());
+122 -1
View File
@@ -720,7 +720,11 @@ pub(crate) fn extract_text_from_operand(
font_encodings: &PageFontEncodings,
encoding_cache: &HashMap<String, Encoding<'_>>,
cmap_decisions: &mut CMapDecisionCache,
font_widths: &PageFontWidths,
) -> Option<String> {
let is_type0_cid_font = font_widths
.get(current_font)
.is_some_and(|info| info.is_cid);
let result = (|| -> Option<String> {
if let Object::String(bytes, _) = obj {
let mut decode_with_entry = |entry: &crate::tounicode::CMapEntry| -> Option<String> {
@@ -962,7 +966,31 @@ pub(crate) fn extract_text_from_operand(
return Some(symbol_text);
}
// Latin-1 fallback
// Latin-1 fallback. Safe ONLY for fonts that use single-byte
// encodings — for these, an unmapped byte is a valid character
// code in Latin-1/WinAnsi space. CID fonts (Type0 / Identity-H)
// emit multi-byte CIDs that aren't characters; per-byte Latin-1
// produces mojibake (e.g. 2-byte CID 0xCDD9 → "ÍÙ" for the
// production scrape_id 019de78c-... samples).
//
// For a CID font (has_cmap is set OR a /ToUnicode reference
// exists) with any non-ASCII bytes, emit a single U+FFFD per
// CID instead. This both replaces the mojibake with a proper
// "decode failed" marker AND keeps `detect_encoding_issues`
// tripping so the page is flagged for OCR — the existing
// garbage-detection path that the high-Latin-1 mojibake used
// to satisfy by accident.
if is_type0_cid_font && bytes.iter().any(|&b| b > 0x7F) {
// 2-byte CIDs (Identity-H) are by far the common case; for
// an odd byte count we still emit at least one marker so
// detection downstream fires.
let cid_count = (bytes.len() / 2).max(1);
return Some("\u{FFFD}".repeat(cid_count));
}
// Pure ASCII bytes round-trip safely (Latin-1 == ASCII for
// 0x00..=0x7F), and non-CID (Type1 / TrueType / Type3) fonts
// use single-byte encodings where Latin-1 fallback is the
// canonical interpretation.
Some(bytes.iter().map(|&b| b as char).collect())
} else {
None
@@ -1213,4 +1241,97 @@ mod tests {
let bad = "###!!!@@@$$$";
assert!(score_text(good) > score_text(bad));
}
#[test]
fn cid_font_with_unparseable_cmap_does_not_emit_latin1_mojibake() {
// Type0/CID font (font_widths reports `is_cid=true`) where the
// ToUnicode CMap couldn't be parsed (FontCMaps doesn't have the
// obj_num). Bytes are a 2-byte CID stream containing high bytes
// that aren't valid UTF-8 — exactly the case in the production
// samples (Identity-H text where the ToUnicode CMap was missing
// or malformed, scrape_id 019de78c-..., e.g. "Í Ù Z)¿").
//
// Without the guard, the function falls through to the byte-by-byte
// Latin-1 fallback and produces "ÍÙ" (U+00CD U+00D9). The correct
// behavior is to emit U+FFFD per CID so downstream
// `detect_encoding_issues` flags the page for OCR.
let bytes = vec![0xCD_u8, 0xD9, 0xCD, 0xD9];
let obj = Object::String(bytes, lopdf::StringFormat::Hexadecimal);
let font_cmaps = FontCMaps::default();
let mut font_tounicode_refs: HashMap<String, u32> = HashMap::new();
font_tounicode_refs.insert("F0".to_string(), 999);
let inline_cmaps = HashMap::new();
let font_encodings: PageFontEncodings = HashMap::new();
let encoding_cache: HashMap<String, Encoding<'_>> = HashMap::new();
let mut decisions = CMapDecisionCache::new();
let mut font_widths: PageFontWidths = HashMap::new();
font_widths.insert("F0".to_string(), make_font_info(&[], 1000, true));
let result = extract_text_from_operand(
&obj,
"F0",
None,
&font_cmaps,
&font_tounicode_refs,
&inline_cmaps,
&font_encodings,
&encoding_cache,
&mut decisions,
&font_widths,
);
let text = result.expect("CID font fallback should still emit a marker");
assert!(
!text.contains('\u{00CD}') && !text.contains('\u{00D9}'),
"CID font with unparseable CMap leaked Latin-1 mojibake: {text:?}"
);
assert!(
text.contains('\u{FFFD}'),
"CID font with unparseable CMap should emit U+FFFD so detect_encoding_issues fires: {text:?}"
);
}
#[test]
fn simple_font_latin1_fallback_passes_high_bytes_through() {
// A Type1/TrueType simple font (is_cid=false) with a `/ToUnicode`
// reference but no usable CMap and no `/Differences` map.
// Per-byte Latin-1 IS the canonical interpretation here — these
// bytes are character codes, not CIDs. The CID guard must NOT
// strip them. Reproduces the false positive that an earlier
// version of the guard introduced for fonts in PDFs like
// pdf-evals/Navigating-Artificial-Intelligence-..., where bytes
// like 0xB6 are legitimate Latin-1 character codes.
let bytes = vec![0x24_u8, 0x47, 0xB6, 0x56]; // "$G¶V"
let obj = Object::String(bytes, lopdf::StringFormat::Hexadecimal);
let font_cmaps = FontCMaps::default();
let mut font_tounicode_refs: HashMap<String, u32> = HashMap::new();
font_tounicode_refs.insert("F1".to_string(), 999);
let inline_cmaps = HashMap::new();
let font_encodings: PageFontEncodings = HashMap::new();
let encoding_cache: HashMap<String, Encoding<'_>> = HashMap::new();
let mut decisions = CMapDecisionCache::new();
let mut font_widths: PageFontWidths = HashMap::new();
font_widths.insert("F1".to_string(), make_font_info(&[], 1000, false));
let text = extract_text_from_operand(
&obj,
"F1",
None,
&font_cmaps,
&font_tounicode_refs,
&inline_cmaps,
&font_encodings,
&encoding_cache,
&mut decisions,
&font_widths,
)
.expect("simple font should round-trip Latin-1 bytes");
assert_eq!(text, "$G\u{00B6}V");
assert!(
!text.contains('\u{FFFD}'),
"simple font fallback must not stamp FFFD over legitimate bytes: {text:?}"
);
}
}
+299 -25
View File
@@ -35,6 +35,7 @@ pub(crate) fn detect_columns(
if page_items.is_empty() {
return vec![];
}
debug!("page {}: detect_columns: {} items", page, page_items.len());
// Find page bounds
let x_min = page_items.iter().map(|i| i.x).fold(f32::INFINITY, f32::min);
@@ -163,10 +164,17 @@ pub(crate) fn detect_columns(
}
}
}
// Try XY-cut fallback before giving up
if let Some(columns) = try_xy_cut_split(&page_items, x_min, x_max, page) {
return columns;
}
return vec![ColumnRegion { x_min, x_max }];
}
return validate_and_build_columns(
// Try center-based assignment first (handles asymmetric layouts / sidebars
// better than edge-based). Fall back to edge-based if center produces
// a degenerate split (one side empty).
let result = validate_and_build_columns(
&valleys,
&page_items,
x_min,
@@ -175,8 +183,177 @@ pub(crate) fn detect_columns(
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
page,
false, // edge-based assignment for absolute valleys
true, // center-based assignment
);
if result.len() > 1 {
return result;
}
let result = validate_and_build_columns(
&valleys,
&page_items,
x_min,
BIN_WIDTH,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
page,
false, // edge-based fallback
);
if result.len() > 1 {
return result;
}
// Fallback: XY-cut style gap detection. When the histogram finds no
// clear valleys (common with asymmetric/sidebar layouts), look for the
// largest horizontal gap between item edges. This is a simplified
// single-level XY-cut inspired by opendataloader's XY-Cut++ algorithm.
if page_items.len() >= 20 && !page_has_table {
if let Some(columns) = try_xy_cut_split(&page_items, x_min, x_max, page) {
return columns;
}
}
vec![ColumnRegion { x_min, x_max }]
}
/// Simplified single-level XY-cut: find the largest horizontal gap between
/// item right-edges and left-edges. If the gap is wide enough and both sides
/// have sufficient items with vertical overlap, split into two columns.
///
/// Inspired by opendataloader's XY-Cut++ algorithm but without full recursion.
/// Handles asymmetric layouts (sidebars) that the histogram misses because
/// the narrow column has too few items to register in the occupancy profile.
fn try_xy_cut_split(
page_items: &[&TextItem],
page_x_min: f32,
page_x_max: f32,
page: u32,
) -> Option<Vec<ColumnRegion>> {
const MIN_GAP: f32 = 15.0; // minimum gap to consider a split
const MIN_ITEMS_MAJOR: usize = 10; // major column must have ≥10 items
const MIN_ITEMS_MINOR: usize = 3; // minor column (sidebar) must have ≥3
let page_width = page_x_max - page_x_min;
if page_width < 200.0 {
return None;
}
// Collect all item edges: (right_edge, left_edge) pairs sorted by right_edge
// The gap between one item's right edge and the next item's left edge
// reveals column gutters.
let mut edges: Vec<(f32, f32)> = page_items
.iter()
.map(|i| (i.x, i.x + effective_width(i)))
.collect();
edges.sort_by(|a, b| a.0.total_cmp(&b.0));
// Find the largest gap between consecutive items (by left edge).
// Use a sweep: sort left edges, find max gap between sorted right edges
// of items to the left and left edges of items to the right.
let mut left_edges: Vec<f32> = page_items.iter().map(|i| i.x).collect();
left_edges.sort_by(|a, b| a.total_cmp(b));
// Build prefix max of right edges (for items sorted by left edge)
let mut sorted_by_left: Vec<(f32, f32)> = page_items
.iter()
.map(|i| (i.x, i.x + effective_width(i)))
.collect();
sorted_by_left.sort_by(|a, b| a.0.total_cmp(&b.0));
let mut best_gap = 0.0f32;
let mut best_split = 0.0f32;
let mut max_right_so_far = f32::NEG_INFINITY;
for i in 0..sorted_by_left.len() - 1 {
let (_, right) = sorted_by_left[i];
max_right_so_far = max_right_so_far.max(right);
let (next_left, _) = sorted_by_left[i + 1];
let gap = next_left - max_right_so_far;
if gap > best_gap {
best_gap = gap;
best_split = (max_right_so_far + next_left) / 2.0;
}
}
if best_gap < MIN_GAP {
return None;
}
// Don't split at page margins (within 10% of edges)
let margin = page_width * 0.10;
if best_split - page_x_min < margin || page_x_max - best_split < margin {
return None;
}
// Count items on each side
let left_count = page_items
.iter()
.filter(|i| i.x + effective_width(i) / 2.0 <= best_split)
.count();
let right_count = page_items
.iter()
.filter(|i| i.x + effective_width(i) / 2.0 > best_split)
.count();
let (minor, major) = if left_count <= right_count {
(left_count, right_count)
} else {
(right_count, left_count)
};
if major < MIN_ITEMS_MAJOR || minor < MIN_ITEMS_MINOR {
return None;
}
// Check vertical overlap — both sides should span a meaningful Y range
let left_items: Vec<&&TextItem> = page_items
.iter()
.filter(|i| i.x + effective_width(i) / 2.0 <= best_split)
.collect();
let right_items: Vec<&&TextItem> = page_items
.iter()
.filter(|i| i.x + effective_width(i) / 2.0 > best_split)
.collect();
let l_y_min = left_items.iter().map(|i| i.y).fold(f32::INFINITY, f32::min);
let l_y_max = left_items
.iter()
.map(|i| i.y)
.fold(f32::NEG_INFINITY, f32::max);
let r_y_min = right_items
.iter()
.map(|i| i.y)
.fold(f32::INFINITY, f32::min);
let r_y_max = right_items
.iter()
.map(|i| i.y)
.fold(f32::NEG_INFINITY, f32::max);
let overlap_min = l_y_min.max(r_y_min);
let overlap_max = l_y_max.min(r_y_max);
let overlap = (overlap_max - overlap_min).max(0.0);
let y_range = (l_y_max.max(r_y_max) - l_y_min.min(r_y_min)).max(1.0);
if overlap / y_range < 0.20 {
return None;
}
debug!(
"page {}: XY-cut split at x={:.1} (gap={:.1}pt, left={}, right={})",
page, best_split, best_gap, left_count, right_count
);
Some(vec![
ColumnRegion {
x_min: page_x_min,
x_max: best_split,
},
ColumnRegion {
x_min: best_split,
x_max: page_x_max,
},
])
}
/// Check whether each proposed column contains paragraph-like content.
@@ -218,7 +395,7 @@ fn columns_have_prose(columns: &[ColumnRegion], items: &[&TextItem]) -> bool {
// Sort by Y descending (top of page = higher Y in PDF coords)
let mut sorted: Vec<&TextItem> = col_items;
sorted.sort_by(|a, b| b.y.partial_cmp(&a.y).unwrap_or(std::cmp::Ordering::Equal));
sorted.sort_by(|a, b| b.y.total_cmp(&a.y));
// Group into lines by Y-proximity and measure fill + item count
let mut full_lines = 0usize;
@@ -449,6 +626,33 @@ fn find_relative_valleys(
valleys
}
/// Detect whether a side of a gutter consists predominantly of list-marker
/// glyphs (•, ●, ○, ◦, ▪, ▫, ◆, ◇). A column of bullets on the left margin
/// creates a spurious histogram valley between the bullet and the content.
/// Treating it as a real column splits each list item's text across two
/// "columns," so we reject these candidates.
fn is_list_marker_column(items: &[&&TextItem]) -> bool {
const LIST_MARKERS: &[char] = &['•', '●', '○', '◦', '▪', '▫', '◆', '◇', '■', '□'];
if items.is_empty() {
return false;
}
let marker_count = items
.iter()
.filter(|i| {
let t = i.text.trim();
let mut chars = t.chars();
match (chars.next(), chars.next()) {
(Some(c), None) => LIST_MARKERS.contains(&c),
_ => false,
}
})
.count();
// Require ≥80% of items on this side to be standalone markers. A handful
// of non-marker items (stray page numbers, footnote refs) shouldn't
// defeat the check.
marker_count as f32 / items.len() as f32 >= 0.8
}
/// Validate valley candidates with vertical consistency checks and build column regions.
///
/// When `center_assign` is true, items are assigned to columns based on their
@@ -505,7 +709,28 @@ fn validate_and_build_columns(
})
.collect();
if left_items.len() < min_items || right_items.len() < min_items {
// Require both sides to have items. Symmetric layout needs min_items
// on each side. Asymmetric layouts (sidebars) are accepted when the
// dominant side has ≥ min_items and the smaller side has ≥ 3 items.
let (smaller, larger) = if left_items.len() <= right_items.len() {
(left_items.len(), right_items.len())
} else {
(right_items.len(), left_items.len())
};
if larger < min_items || smaller < 3 {
continue;
}
// Reject valleys where the smaller side is just a column of list
// markers (bullets aligned at the left margin). This is a common
// pattern in PDFs where ● starts each list item: histogram detection
// sees the gap between bullet and content as a gutter.
let smaller_items: &[&&TextItem] = if left_items.len() <= right_items.len() {
&left_items
} else {
&right_items
};
if is_list_marker_column(smaller_items) {
continue;
}
@@ -616,7 +841,7 @@ fn identify_spanning_lines(items: &[TextItem], columns: &[ColumnRegion]) -> Vec<
// Build (original_index, y) pairs sorted by Y descending for grouping
let mut indexed: Vec<(usize, f32)> =
items.iter().enumerate().map(|(i, it)| (i, it.y)).collect();
indexed.sort_by(|a, b| b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal));
indexed.sort_by(|a, b| b.1.total_cmp(&a.1));
// Group by Y-proximity into rough lines (as index sets)
let mut groups: Vec<Vec<usize>> = Vec::new();
@@ -778,7 +1003,7 @@ pub(crate) fn is_newspaper_layout(
return 0.0;
}
let mut ys: Vec<f32> = lines.iter().map(|l| l.y).collect();
ys.sort_by(|a, b| a.partial_cmp(b).unwrap());
ys.sort_by(|a, b| a.total_cmp(b));
let span = ys.last().unwrap() - ys.first().unwrap();
span / (lines.len() as f32 - 1.0)
};
@@ -845,7 +1070,7 @@ fn split_column_stragglers(lines: Vec<TextLine>) -> (Vec<TextLine>, Vec<TextLine
// Median gap = typical line spacing
let mut sorted_gaps = gaps.clone();
sorted_gaps.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
sorted_gaps.sort_by(|a, b| a.total_cmp(b));
let median_gap = sorted_gaps[sorted_gaps.len() / 2];
// A gap > 3× median (min 30pt) indicates a break between content clusters
@@ -1078,9 +1303,8 @@ pub(crate) fn group_into_lines_with_thresholds(
}
}
above.sort_by(|a, b| b.y.partial_cmp(&a.y).unwrap_or(std::cmp::Ordering::Equal));
below_spanning
.sort_by(|a, b| b.y.partial_cmp(&a.y).unwrap_or(std::cmp::Ordering::Equal));
above.sort_by(|a, b| b.y.total_cmp(&a.y));
below_spanning.sort_by(|a, b| b.y.total_cmp(&a.y));
all_lines.extend(above);
for col in core_columns {
@@ -1101,16 +1325,13 @@ pub(crate) fn group_into_lines_with_thresholds(
// Sort by Y descending (top-first), then by X for same-Y lines
all_page_lines.sort_by(|a, b| {
b.y.partial_cmp(&a.y)
.unwrap_or(std::cmp::Ordering::Equal)
.then(
a.items
.first()
.map(|i| i.x)
.unwrap_or(0.0)
.partial_cmp(&b.items.first().map(|i| i.x).unwrap_or(0.0))
.unwrap_or(std::cmp::Ordering::Equal),
)
b.y.total_cmp(&a.y).then(
a.items
.first()
.map(|i| i.x)
.unwrap_or(0.0)
.total_cmp(&b.items.first().map(|i| i.x).unwrap_or(0.0)),
)
});
// Merge lines at the same Y (within tolerance) into single lines
@@ -1185,11 +1406,7 @@ fn group_single_column(items: Vec<TextItem>, adaptive_threshold: f32) -> Vec<Tex
let items = if use_y_sorting {
// Sort by Y descending (top to bottom in PDF coords)
let mut sorted = items;
sorted.sort_by(|a, b| {
b.y.partial_cmp(&a.y)
.unwrap_or(std::cmp::Ordering::Equal)
.then(a.x.partial_cmp(&b.x).unwrap_or(std::cmp::Ordering::Equal))
});
sorted.sort_by(|a, b| b.y.total_cmp(&a.y).then(a.x.total_cmp(&b.x)));
sorted
} else {
items
@@ -1603,6 +1820,63 @@ mod tests {
);
}
#[test]
fn bullet_marker_column_not_detected_as_column() {
// Pattern: every line is `● <content>`, with ● at x=90 and content
// starting at x=104. Histogram detection sees a gutter between them
// and would split the page into a "bullet column" and "content column",
// scrambling every list item.
let mut items = Vec::new();
for i in 0..15 {
let y = 750.0 - i as f32 * 30.0;
items.push(make_item(1, 90.0, y, ""));
items.push(make_item(
1,
104.0,
y,
"FullContentLineTextHere________________",
));
}
// Pad with content to satisfy min item count for column detection.
for i in 0..15 {
let y = 300.0 - i as f32 * 14.0;
items.push(make_item(1, 72.0, y, "FootnoteText_____________________"));
}
let cols = detect_columns(&items, 1, false);
assert_eq!(
cols.len(),
1,
"Bullet markers aligned at left margin should not be treated as their own column"
);
}
#[test]
fn is_list_marker_column_detects_bullets() {
let items = vec![
make_item(1, 90.0, 100.0, ""),
make_item(1, 90.0, 114.0, ""),
make_item(1, 90.0, 128.0, ""),
make_item(1, 90.0, 142.0, ""),
];
let refs: Vec<&TextItem> = items.iter().collect();
let wrapped: Vec<&&TextItem> = refs.iter().collect();
assert!(is_list_marker_column(&wrapped));
}
#[test]
fn is_list_marker_column_rejects_prose() {
let items = vec![
make_item(1, 30.0, 100.0, "Regular prose line"),
make_item(1, 30.0, 114.0, "Another sentence"),
make_item(1, 30.0, 128.0, "Third line"),
make_item(1, 30.0, 142.0, "Fourth line"),
];
let refs: Vec<&TextItem> = items.iter().collect();
let wrapped: Vec<&&TextItem> = refs.iter().collect();
assert!(!is_list_marker_column(&wrapped));
}
#[test]
fn premask_narrow_line_not_masked() {
// Items that form a line spanning only ~40% of column width → not masked
+9 -36
View File
@@ -36,26 +36,14 @@ pub(crate) use layout::ColumnRegion;
/// Extract text from PDF file as plain string
pub fn extract_text<P: AsRef<Path>>(path: P) -> Result<String, PdfError> {
crate::validate_pdf_file(&path)?;
let doc = match Document::load(&path) {
Ok(d) => d,
Err(ref e) if crate::is_encrypted_lopdf_error(e) => {
Document::load_with_password(&path, "")?
}
Err(e) => return Err(e.into()),
};
let (doc, _) = crate::load_document_from_path(&path)?;
extract_text_from_doc(&doc)
}
/// Extract text from PDF memory buffer
pub fn extract_text_mem(buffer: &[u8]) -> Result<String, PdfError> {
crate::validate_pdf_bytes(buffer)?;
let doc = match Document::load_mem(buffer) {
Ok(d) => d,
Err(ref e) if crate::is_encrypted_lopdf_error(e) => {
Document::load_mem_with_options(buffer, lopdf::LoadOptions::with_password(""))?
}
Err(e) => return Err(e.into()),
};
let (doc, _) = crate::load_document_from_mem(buffer)?;
extract_text_from_doc(&doc)
}
@@ -91,13 +79,7 @@ pub(crate) fn extract_text_with_positions_and_rects<P: AsRef<Path>>(
page_filter: Option<&HashSet<u32>>,
) -> Result<PageExtraction, PdfError> {
crate::validate_pdf_file(&path)?;
let doc = match Document::load(&path) {
Ok(d) => d,
Err(ref e) if crate::is_encrypted_lopdf_error(e) => {
Document::load_with_password(&path, "")?
}
Err(e) => return Err(e.into()),
};
let (doc, _) = crate::load_document_from_path(&path)?;
let font_cmaps = FontCMaps::from_doc(&doc);
let (extraction, _thresholds, _gid_pages) =
extract_positioned_text_from_doc(&doc, &font_cmaps, page_filter)?;
@@ -124,13 +106,7 @@ pub(crate) fn extract_text_with_positions_mem_and_rects(
page_filter: Option<&HashSet<u32>>,
) -> Result<PageExtraction, PdfError> {
crate::validate_pdf_bytes(buffer)?;
let doc = match Document::load_mem(buffer) {
Ok(d) => d,
Err(ref e) if crate::is_encrypted_lopdf_error(e) => {
Document::load_mem_with_options(buffer, lopdf::LoadOptions::with_password(""))?
}
Err(e) => return Err(e.into()),
};
let (doc, _) = crate::load_document_from_mem(buffer)?;
let font_cmaps = FontCMaps::from_doc(&doc);
let (extraction, _thresholds, _gid_pages) =
extract_positioned_text_from_doc(&doc, &font_cmaps, page_filter)?;
@@ -188,7 +164,7 @@ fn extract_positioned_text_impl(
continue;
}
}
let ((mut items, rects, lines), has_gid_fonts) =
let ((mut items, rects, lines), has_gid_fonts, _coords_rotated) =
extract_page_text_items(doc, page_id, *page_num, font_cmaps, include_invisible)?;
if has_gid_fonts {
gid_encoded_pages.insert(*page_num);
@@ -334,17 +310,14 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
for (_, _, group) in &mut line_groups {
let rtl = is_rtl_text(group.iter().map(|i| &i.text));
if rtl {
group.sort_by(|a, b| b.x.partial_cmp(&a.x).unwrap_or(std::cmp::Ordering::Equal));
group.sort_by(|a, b| b.x.total_cmp(&a.x));
} else {
group.sort_by(|a, b| a.x.partial_cmp(&b.x).unwrap_or(std::cmp::Ordering::Equal));
group.sort_by(|a, b| a.x.total_cmp(&b.x));
}
}
// Sort groups by page then Y descending (top of page first)
line_groups.sort_by(|a, b| {
a.0.cmp(&b.0)
.then_with(|| b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal))
});
line_groups.sort_by(|a, b| a.0.cmp(&b.0).then_with(|| b.1.total_cmp(&a.1)));
let mut merged = Vec::new();
@@ -452,7 +425,7 @@ pub(crate) fn merge_subscript_items(items: Vec<TextItem>) -> Vec<TextItem> {
for (_, _, mut group) in line_groups {
// Sort by X position
group.sort_by(|a, b| a.x.partial_cmp(&b.x).unwrap_or(std::cmp::Ordering::Equal));
group.sort_by(|a, b| a.x.total_cmp(&b.x));
// Find the dominant (most common) font size in this group
let max_fs = group.iter().map(|i| i.font_size).fold(0.0_f32, f32::max);
+2
View File
@@ -373,6 +373,7 @@ fn extract_form_xobject_text_inner(
&font_encodings,
&encoding_cache,
cmap_decisions,
&font_widths,
) {
let combined = multiply_matrices(&text_matrix, &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
@@ -517,6 +518,7 @@ fn extract_form_xobject_text_inner(
&font_encodings,
&encoding_cache,
cmap_decisions,
&font_widths,
) {
current_text.push_str(&text);
}
+3665 -70
View File
File diff suppressed because it is too large Load Diff
+46 -4
View File
@@ -8,6 +8,24 @@ use log::debug;
/// Font statistics for a document
pub(crate) struct FontStats {
pub(crate) most_common_size: f32,
/// Font size frequency distribution (size_key → line count).
/// Used for rarity-based heading detection.
pub(crate) size_counts: HashMap<i32, usize>,
/// Total number of lines counted.
pub(crate) total_lines: usize,
}
/// Compute how rare a font size is in the document (0.0 = most common, 1.0 = unique).
/// Mirrors opendataloader's font rarity boosting approach: heading fonts appear on
/// far fewer lines than body text, so their percentile rank is high.
pub(crate) fn font_size_rarity(font_size: f32, stats: &FontStats) -> f32 {
if stats.total_lines == 0 {
return 0.0;
}
let key = (font_size * 10.0) as i32;
let count = stats.size_counts.get(&key).copied().unwrap_or(0);
// Rarity = 1 - (frequency ratio). A size used on 1/100 lines has rarity ~0.99.
1.0 - (count as f32 / stats.total_lines as f32)
}
/// Calculate font stats directly from items (before grouping into lines)
@@ -21,6 +39,8 @@ pub(crate) fn calculate_font_stats_from_items(items: &[TextItem]) -> FontStats {
}
}
let total_lines = size_counts.values().sum();
// Break ties by preferring the smaller font size for deterministic output
let most_common_size = size_counts
.iter()
@@ -30,7 +50,11 @@ pub(crate) fn calculate_font_stats_from_items(items: &[TextItem]) -> FontStats {
.map(|(size, _)| *size as f32 / 10.0)
.unwrap_or(12.0);
FontStats { most_common_size }
FontStats {
most_common_size,
size_counts,
total_lines,
}
}
/// Calculate font stats from grouped lines
@@ -48,6 +72,8 @@ pub(crate) fn calculate_font_stats(lines: &[TextLine]) -> FontStats {
}
}
let total_lines = size_counts.values().sum();
// Break ties by preferring the smaller font size for deterministic output
let most_common_size = size_counts
.iter()
@@ -57,7 +83,23 @@ pub(crate) fn calculate_font_stats(lines: &[TextLine]) -> FontStats {
.map(|(size, _)| *size as f32 / 10.0)
.unwrap_or(12.0);
FontStats { most_common_size }
FontStats {
most_common_size,
size_counts,
total_lines,
}
}
/// Determine the heading level for a bold-only line that didn't meet the font-size
/// threshold. These are common in academic papers where section headings are bold
/// at the same size as body text.
///
/// Returns a level below the lowest font-size tier (or H2 when no tiers exist).
pub(crate) fn bold_heading_level(heading_tiers: &[f32]) -> usize {
let level = heading_tiers.len() + 1;
// Clamp to 1..=6 — if no font-size tiers, bold headings become H2
// (H1 is reserved for titles which are typically larger)
level.clamp(2, 6)
}
/// Detect TOC-style lines that contain dot leaders (e.g., "Section Name .... 42").
@@ -121,7 +163,7 @@ pub(crate) fn compute_paragraph_threshold(lines: &[TextLine], base_size: f32) ->
return fallback;
}
gaps.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
gaps.sort_by(|a, b| a.total_cmp(b));
let median = gaps[gaps.len() / 2];
@@ -221,7 +263,7 @@ pub(crate) fn compute_heading_tiers(lines: &[TextLine], base_size: f32) -> Vec<f
}
// Sort descending
heading_sizes.sort_by(|a, b| b.partial_cmp(a).unwrap_or(std::cmp::Ordering::Equal));
heading_sizes.sort_by(|a, b| b.total_cmp(a));
// Cluster sizes within 0.5pt into same tier (use first value as representative)
let mut tiers: Vec<f32> = Vec::new();
+94 -9
View File
@@ -4,13 +4,11 @@
pub(crate) fn is_caption_line(text: &str) -> bool {
let trimmed = text.trim();
// Common caption prefixes in multiple languages
let caption_prefixes = [
"Figure ",
// Caption prefixes that always match (always followed by identifiers)
let always_prefixes = [
"Figura ",
"Fig. ",
"Fig ",
"Table ",
"Tabela ",
"Source:",
"Fonte:",
@@ -27,23 +25,61 @@ pub(crate) fn is_caption_line(text: &str) -> bool {
"Photo ",
"Foto ",
];
// Check if line starts with a caption prefix
for prefix in &caption_prefixes {
for prefix in &always_prefixes {
if trimmed.starts_with(prefix) {
return true;
}
}
// Check case-insensitive patterns
// "Figure" and "Table" need a digit/reference after them to distinguish
// captions ("Table 1", "Figure 3.2") from headings ("Table of Contents")
for prefix in ["Figure ", "Table "] {
if let Some(rest) = trimmed.strip_prefix(prefix) {
if rest
.trim_start()
.starts_with(|c: char| c.is_ascii_digit() || c == '(' || c == '#')
{
return true;
}
}
}
// Check case-insensitive patterns — require digit or punctuation after
// prefix to avoid matching "Table of Contents" or "Figure drawing" etc.
let lower = trimmed.to_lowercase();
if lower.starts_with("figure ") || lower.starts_with("table ") || lower.starts_with("source:") {
for pfx in ["figure ", "table "] {
if let Some(rest) = lower.strip_prefix(pfx) {
if rest
.trim_start()
.starts_with(|c: char| c.is_ascii_digit() || c == '(' || c == '#')
{
return true;
}
}
}
if lower.starts_with("source:") {
return true;
}
false
}
/// Check if text starts with an unambiguous bullet marker (●, •, ○, ◦).
///
/// Narrower than [`is_list_item`]: it excludes numbered/lettered patterns
/// like `1.` or `a)`, which legitimately appear as section headings in many
/// documents. Used by the heading classifier to reject bullet lines without
/// also demoting numbered headings.
pub(crate) fn starts_with_bullet_marker(text: &str) -> bool {
let trimmed = text.trim_start();
trimmed.starts_with("")
|| trimmed.starts_with("")
|| trimmed.starts_with("")
|| trimmed.starts_with("")
|| trimmed.starts_with("- ")
|| trimmed.starts_with("* ")
}
/// Check if text looks like a list item
pub(crate) fn is_list_item(text: &str) -> bool {
let trimmed = text.trim_start();
@@ -95,6 +131,16 @@ pub(crate) fn format_list_item(text: &str) -> String {
if let Some(rest) = trimmed.strip_prefix(*bullet) {
return format!("- {}", rest.trim_start());
}
// Bullet inside a leading bold/italic run (e.g. "**● Label:** rest").
// The run wraps both the marker and the following label because both
// use a bold font in the PDF.
for wrapper in ["**", "*"] {
if let Some(after_open) = trimmed.strip_prefix(wrapper) {
if let Some(rest) = after_open.strip_prefix(*bullet) {
return format!("- {}{}", wrapper, rest.trim_start());
}
}
}
}
if trimmed.starts_with("- ") || trimmed.starts_with("* ") {
@@ -178,3 +224,42 @@ pub(crate) fn is_monospace_font(font_name: &str) -> bool {
patterns.iter().any(|p| lower.contains(p))
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn format_list_item_plain_bullet() {
assert_eq!(format_list_item("● Item"), "- Item");
assert_eq!(format_list_item("• Item"), "- Item");
}
#[test]
fn format_list_item_bullet_inside_bold() {
// PDF that uses bold font for both the marker and the label produces
// a single bold run like "**● Label:** rest"; the bullet must still
// be stripped and the bold wrapper preserved on the label.
assert_eq!(
format_list_item("**● Fraud: Willing cooperation;**"),
"- **Fraud: Willing cooperation;**"
);
assert_eq!(
format_list_item("**● Label:** rest of line"),
"- **Label:** rest of line"
);
assert_eq!(format_list_item("*● Italic:* rest"), "- *Italic:* rest");
}
#[test]
fn format_list_item_already_dash() {
assert_eq!(format_list_item("- existing"), "- existing");
}
#[test]
fn is_list_item_with_bullet_space() {
assert!(is_list_item("● Item"));
assert!(is_list_item("• Item"));
assert!(is_list_item("- Item"));
}
}
+492 -11
View File
@@ -1,19 +1,154 @@
//! Core line-to-markdown conversion loop with table/image interleaving.
use std::collections::HashSet;
use std::collections::{HashMap, HashSet};
use crate::structure_tree::StructRole;
use crate::types::TextLine;
use super::analysis::{
calculate_font_stats, compute_heading_tiers, compute_paragraph_threshold, detect_header_level,
has_dot_leaders,
bold_heading_level, calculate_font_stats, compute_heading_tiers, compute_paragraph_threshold,
detect_header_level, font_size_rarity, has_dot_leaders,
};
use super::classify::{
format_list_item, is_caption_line, is_list_item, is_monospace_font, starts_with_bullet_marker,
};
use super::classify::{format_list_item, is_caption_line, is_list_item, is_monospace_font};
use super::postprocess::clean_markdown;
use super::preprocess::{merge_drop_caps, merge_heading_lines};
use super::MarkdownOptions;
/// Pre-scan struct heading tags to find levels that are overused — i.e., tagged on
/// so many lines that they clearly represent body text, not real headings.
/// Returns the set of heading levels (16) that should be suppressed.
///
/// Some PDFs (e.g. British Academy grant guidance) tag every numbered paragraph
/// line as H2, producing hundreds of false headings. We detect this by checking
/// if any heading level accounts for >25% of tagged lines.
fn detect_overused_struct_heading_levels(
lines: &[TextLine],
struct_roles: Option<
&std::collections::HashMap<u32, std::collections::HashMap<i64, StructRole>>,
>,
) -> HashSet<usize> {
let mut overused = HashSet::new();
let Some(roles) = struct_roles else {
return overused;
};
let mut level_counts: HashMap<usize, usize> = HashMap::new();
let mut total = 0usize;
for line in lines {
if let Some(role) = resolve_line_struct_role(line, roles) {
total += 1;
if let Some(level) = struct_role_heading_level(&role) {
*level_counts.entry(level).or_insert(0) += 1;
}
}
}
if total < 20 {
return overused;
}
for (&level, &count) in &level_counts {
let ratio = count as f32 / total as f32;
if ratio > 0.15 {
log::debug!(
"struct heading H{} overused: {}/{} lines ({:.0}%), suppressing",
level,
count,
total,
ratio * 100.0
);
overused.insert(level);
}
}
overused
}
/// Pre-scan lines to find "isolated" ones: short lines with paragraph breaks both
/// before and after. These are heading candidates even at body font size — common
/// in academic papers ("Acknowledgements", "B.3 Prompt Engineering").
fn find_isolated_lines(lines: &[TextLine], base_size: f32, para_threshold: f32) -> HashSet<usize> {
let mut set = HashSet::new();
for i in 0..lines.len() {
let line = &lines[i];
let plain = line.text();
let trimmed = plain.trim();
let word_count = trimmed.split_whitespace().count();
if !(1..=6).contains(&word_count) || trimmed.len() <= 3 {
continue;
}
let font_size = line.items.first().map(|it| it.font_size).unwrap_or(0.0);
if font_size < base_size * 0.95 {
continue;
}
if is_list_item(trimmed) || is_caption_line(trimmed) {
continue;
}
// Reject lines that look like wrapped paragraph text:
// ends with hyphen, comma, preposition, or lowercase continuation
let last_char = trimmed.chars().last().unwrap_or(' ');
if last_char == '-' || last_char == ',' || last_char == ';' {
continue;
}
// Last word is a common continuation word → wrapped paragraph
let last_word = trimmed.split_whitespace().last().unwrap_or("");
let continuation_words = [
"the", "a", "an", "and", "or", "of", "in", "to", "for", "with", "by", "on", "at",
"from", "as", "is", "are", "was", "were", "be", "that", "this", "their", "its", "our",
"your", "has", "have", "had", "not",
];
if continuation_words.contains(&last_word.to_lowercase().as_str()) {
continue;
}
// Paragraph break BEFORE
let break_before = if i == 0 {
true
} else {
let prev = &lines[i - 1];
prev.page != line.page || (prev.y - line.y).abs() > para_threshold
};
// Paragraph break AFTER
let break_after = if i + 1 >= lines.len() {
true
} else {
let next = &lines[i + 1];
next.page != line.page || (line.y - next.y).abs() > para_threshold
};
if !break_before || !break_after {
continue;
}
set.insert(i);
}
// Density guard: if too many lines on a page are "isolated", they're
// all paragraph lines in a multi-column layout, not headings. Real
// headings are rare — at most ~20% of lines on a page.
let mut page_line_counts: HashMap<u32, (usize, usize)> = HashMap::new(); // (total, isolated)
for (i, line) in lines.iter().enumerate() {
let entry = page_line_counts.entry(line.page).or_insert((0, 0));
entry.0 += 1;
if set.contains(&i) {
entry.1 += 1;
}
}
for (&page, &(total, isolated)) in &page_line_counts {
if total > 0 && isolated as f32 / total as f32 > 0.25 {
// Too many isolated lines on this page — remove them all
set.retain(|&i| lines[i].page != page);
}
}
set
}
/// Resolve the dominant structure role for a text line by looking up its items' MCIDs.
///
/// Returns the first non-container role found (skipping Document/Part/Sect/Div/NonStruct/Span).
@@ -256,6 +391,16 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
// threshold and cause every line to be treated as a paragraph break.
let para_threshold = compute_paragraph_threshold(&lines, base_size);
// Pre-scan: identify isolated lines (paragraph break before AND after).
// These are heading candidates even without bold/large font — common in
// academic papers where section titles like "Acknowledgements" sit alone
// between paragraphs at body font size. Inspired by opendataloader's
// lookahead in HeadingProcessor (prevNode/nextNode context).
let isolated_lines = find_isolated_lines(&lines, base_size, para_threshold);
// Detect struct heading levels that are overused (body text mistagged as headings)
let overused_heading_levels = detect_overused_struct_heading_levels(&lines, struct_roles);
let mut output = String::new();
let mut current_page = 0u32;
let mut prev_y = f32::MAX;
@@ -277,7 +422,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
all_content_pages.sort();
all_content_pages.dedup();
for line in lines {
for (line_idx, line) in lines.iter().enumerate() {
// Page break
if line.page != current_page {
// Flush current page's remaining tables and images
@@ -405,7 +550,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
// Detect figure/table captions and source citations
// These should be on their own line followed by a paragraph break
let struct_role = struct_roles.and_then(|roles| resolve_line_struct_role(&line, roles));
let struct_role = struct_roles.and_then(|roles| resolve_line_struct_role(line, roles));
// Determine if this line is code (struct-tree or font-based) for block accumulation
let is_code_line = struct_role
@@ -437,13 +582,72 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
// Structure roles ADD headings (e.g. same-size text tagged H2) but do NOT
// suppress headings that the font heuristic would detect (some tagged PDFs
// mark obvious headings as P or Span).
let struct_heading = struct_role.as_ref().and_then(struct_role_heading_level);
let struct_heading = struct_role
.as_ref()
.and_then(struct_role_heading_level)
.filter(|level| !overused_heading_levels.contains(level));
// Protect wrapped list items: when inside a list, a visually-continuing
// line (same indent, line-wrap spacing) must not be reclassified as a
// heading by the font heuristic — PDFs often bold the lead phrase of a
// list item across multiple wrap lines, and an all-bold middle line
// would otherwise split one item into a heading + stray body text.
// We gate on the document's paragraph threshold so genuine section
// headings that follow a numbered paragraph (y_gap > para_threshold)
// remain detectable.
let looks_like_list_continuation = in_list
&& match (last_list_x, line.items.first().map(|i| i.x)) {
(Some(list_x), Some(curr_x)) => {
let x_ok = curr_x >= list_x - 5.0 && curr_x <= list_x + 50.0;
let y_ok = y_gap >= 0.0 && y_gap <= para_threshold;
x_ok && y_ok && !is_list_item(plain_trimmed)
}
_ => false,
};
let heuristic_heading = if options.detect_headers
&& !looks_like_list_continuation
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !starts_with_bullet_marker(plain_trimmed)
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
detect_header_level(line_font_size, base_size, &heading_tiers)
detect_header_level(line_font_size, base_size, &heading_tiers).or_else(|| {
// Rarity-based heading detection (inspired by opendataloader).
// Heading probability scoring with lookahead context.
// Score = rarity * 0.5 + bold * 0.3 + standalone * 0.2
// + isolated * 0.3 (paragraph break before AND after)
// Only consider lines at or above body font size.
if line_font_size < base_size * 0.95 {
return None;
}
let word_count = plain_trimmed.split_whitespace().count();
if !(1..=15).contains(&word_count) {
return None;
}
let rarity = font_size_rarity(line_font_size, &font_stats);
let all_bold = !line.items.is_empty() && line.items.iter().all(|i| i.is_bold);
let standalone = !in_paragraph;
let isolated = isolated_lines.contains(&line_idx);
let score = rarity * 0.5
+ if all_bold { 0.3 } else { 0.0 }
+ if standalone { 0.2 } else { 0.0 }
+ if isolated { 0.3 } else { 0.0 };
// Require standalone + at least one strong signal.
// Non-bold, non-isolated lines need very high rarity (≥0.97)
// to avoid classifying ordinary body text as headings in
// multi-column layouts where column switches break
// paragraph continuity and minor font-size variation
// inflates rarity scores.
let has_strong_signal = all_bold || isolated || (rarity >= 0.97 && word_count <= 8);
if score >= 0.5 && standalone && word_count >= 2 && has_strong_signal {
Some(bold_heading_level(&heading_tiers))
} else {
None
}
})
} else {
None
};
@@ -460,11 +664,16 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
continue;
}
// Structure-tree list item (LI only — LBody is a continuation, not a new item)
// Structure-tree list item (LI only — LBody is a continuation, not a new item).
// Some tagged PDFs use a "flat" style where every wrapped line in a list item
// gets its own MCID tagged directly under LI. When we're already inside a list
// and the line has no visible bullet marker, treat it as a continuation (falls
// through to the continuation logic below) rather than a new list item.
if struct_role
.as_ref()
.is_some_and(|r| matches!(r, StructRole::LI))
&& !is_list_item(plain_trimmed)
&& !in_list
{
if in_paragraph {
output.push_str("\n\n");
@@ -626,6 +835,8 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
// Compute the typical line spacing for paragraph break detection
let para_threshold = compute_paragraph_threshold(&lines, base_size);
let isolated_lines = find_isolated_lines(&lines, base_size, para_threshold);
let mut output = String::new();
let mut current_page = 0u32;
let mut prev_y = f32::MAX;
@@ -634,7 +845,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
let mut last_list_x: Option<f32> = None;
let mut prev_had_dot_leaders = false;
for line in lines {
for (line_idx, line) in lines.iter().enumerate() {
// Page break
if line.page != current_page {
if current_page > 0 {
@@ -699,7 +910,27 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
if let Some(header_level) =
detect_header_level(line_font_size, base_size, &heading_tiers)
detect_header_level(line_font_size, base_size, &heading_tiers).or_else(|| {
if line_font_size < base_size * 0.95 {
return None;
}
let word_count = plain_trimmed.split_whitespace().count();
if !(1..=15).contains(&word_count) {
return None;
}
let rarity = font_size_rarity(line_font_size, &font_stats);
let all_bold = !line.items.is_empty() && line.items.iter().all(|i| i.is_bold);
let standalone = !in_paragraph;
let isolated = isolated_lines.contains(&line_idx);
let score = rarity * 0.5
+ if all_bold { 0.3 } else { 0.0 }
+ if standalone { 0.2 } else { 0.0 }
+ if isolated { 0.3 } else { 0.0 };
if score >= 0.5 && standalone && word_count >= 2 {
return Some(bold_heading_level(&heading_tiers));
}
None
})
{
if in_paragraph {
output.push_str("\n\n");
@@ -887,6 +1118,59 @@ mod tests {
);
}
#[test]
fn test_struct_role_li_flat_continuation_lines_merge() {
// Regression: some tagged PDFs put each wrapped visual line of a list
// item under its own MCID, all tagged directly as LI. Continuation
// lines (no bullet marker) must merge into the bulleted parent item,
// not each become their own list item.
let make = |text: &str, mcid: i64, x: f32, y: f32| {
let mut item = make_item(text, 1, Some(mcid));
item.x = x;
item.y = y;
item
};
let lines = vec![
make_line(vec![make("● First item that wraps onto", 0, 90.0, 322.0)]),
make_line(vec![make("a continuation line.", 1, 108.0, 306.0)]),
make_line(vec![make("● Second bullet also wraps", 2, 90.0, 290.0)]),
make_line(vec![make("to a second line here.", 3, 108.0, 274.0)]),
];
let mut page_roles = HashMap::new();
for mcid in 0..4 {
page_roles.insert(mcid, StructRole::LI);
}
let mut roles = HashMap::new();
roles.insert(1u32, page_roles);
let md = to_markdown_from_lines_with_tables_and_images(
lines,
MarkdownOptions::default(),
HashMap::new(),
HashMap::new(),
&std::collections::HashSet::new(),
Some(&roles),
);
assert!(
md.contains("- First item that wraps onto a continuation line."),
"continuation should merge into first bullet: {md}"
);
assert!(
md.contains("- Second bullet also wraps to a second line here."),
"continuation should merge into second bullet: {md}"
);
assert!(
!md.contains("- a continuation line."),
"continuation line should not get its own bullet: {md}"
);
assert!(
!md.contains("- to a second line here."),
"continuation line should not get its own bullet: {md}"
);
}
#[test]
fn test_struct_role_blockquote() {
let lines = vec![make_line(vec![make_item("Quoted text", 1, Some(0))])];
@@ -1016,6 +1300,62 @@ mod tests {
);
}
#[test]
fn test_rarity_heading_requires_strong_signal() {
// Simulate a two-column academic paper where body text lines become
// "standalone" due to column switches. Body text at the same font
// size as most of the document should NOT be classified as headings
// just because of moderate rarity + standalone.
//
// Regression: previously, lines with rarity ~0.62 and standalone=true
// scored 0.51 (>=0.5 threshold), producing hundreds of false ## headings.
// Create many body-text lines at font_size=10.9 (most common)
let mut lines = Vec::new();
for i in 0..20 {
let mut item = make_item("This is ordinary body text in a paragraph.", 1, None);
item.font_size = 10.9;
item.y = 700.0 - i as f32 * 14.0;
lines.push(make_line(vec![item]));
}
// A few lines at a slightly different size (simulating column B text)
for i in 0..10 {
let mut item = make_item("Another body text line from the second column.", 1, None);
item.font_size = 11.0; // slightly different → non-zero rarity
item.y = 700.0 - i as f32 * 14.0;
item.x = 320.0; // right column
lines.push(make_line(vec![item]));
}
// One genuine bold heading
let mut heading_item = make_item("3 Philosophical Perspectives", 1, None);
heading_item.font_size = 10.9;
heading_item.is_bold = true;
heading_item.y = 200.0;
lines.push(make_line(vec![heading_item]));
let md = to_markdown_from_lines_with_tables_and_images(
lines,
MarkdownOptions::default(),
HashMap::new(),
HashMap::new(),
&std::collections::HashSet::new(),
None,
);
// The bold heading should be detected
assert!(
md.contains("## 3 Philosophical Perspectives"),
"Bold heading should be detected: {md}"
);
// Body text lines should NOT be headings
let heading_count = md.lines().filter(|l| l.starts_with("##")).count();
assert!(
heading_count <= 2,
"Expected at most 2 headings but found {heading_count} in:\n{md}"
);
}
#[test]
fn test_struct_role_code_multiline_accumulation() {
let mut line1 = make_item("fn main() {", 1, Some(0));
@@ -1058,4 +1398,145 @@ mod tests {
"Should not have adjacent close/open fences: {md}"
);
}
#[test]
fn test_overused_struct_heading_suppressed() {
// Simulate a PDF where H2 is mistagged on body text lines.
// 30 lines total: 5 tagged H1 (real headings), 20 tagged H2 (mistagged body),
// 5 tagged P.
let mut lines = Vec::new();
let mut page_roles = HashMap::new();
let mut mcid = 0i64;
for i in 0..30 {
let mut item = make_item(&format!("Line {i}"), 1, Some(mcid));
item.y = 700.0 - (i as f32 * 15.0);
lines.push(make_line(vec![item]));
let role = if i < 5 {
StructRole::H1
} else if i < 25 {
StructRole::H2
} else {
StructRole::P
};
page_roles.insert(mcid, role);
mcid += 1;
}
let mut roles = HashMap::new();
roles.insert(1u32, page_roles);
let overused = detect_overused_struct_heading_levels(&lines, Some(&roles));
// H2 is on 20/30 = 67% of lines — should be suppressed
assert!(
overused.contains(&2),
"H2 should be detected as overused: {:?}",
overused
);
// H1 is on 5/30 = 17% — should also be suppressed at >15% threshold
assert!(
overused.contains(&1),
"H1 at 17% should also be suppressed: {:?}",
overused
);
}
#[test]
fn test_normal_struct_headings_not_suppressed() {
// Normal document: a few headings, mostly body text
let mut lines = Vec::new();
let mut page_roles = HashMap::new();
let mut mcid = 0i64;
for i in 0..50 {
let mut item = make_item(&format!("Line {i}"), 1, Some(mcid));
item.y = 700.0 - (i as f32 * 14.0);
lines.push(make_line(vec![item]));
let role = if i % 10 == 0 {
StructRole::H1 // 5 headings out of 50 = 10%
} else {
StructRole::P
};
page_roles.insert(mcid, role);
mcid += 1;
}
let mut roles = HashMap::new();
roles.insert(1u32, page_roles);
let overused = detect_overused_struct_heading_levels(&lines, Some(&roles));
assert!(
overused.is_empty(),
"No heading level should be overused: {:?}",
overused
);
}
#[test]
fn test_wrapped_bold_lead_in_list_item_not_heading() {
// Regression: numbered-list items whose bold "lead" phrase wraps onto
// a second line (e.g. definitions in system cards) must not have the
// wrapped line reclassified as a heading. The middle line is
// all_bold + standalone (in_paragraph=false while in_list), which
// previously tripped the rarity heuristic and emitted #### in the
// middle of the item, splitting the body into stray bullets.
let make = |text: &str, x: f32, y: f32, bold: bool| {
let mut item = make_item(text, 1, None);
item.x = x;
item.y = y;
item.is_bold = bold;
item
};
let lines = vec![
// "1. **bold lead phrase start**"
make_line(vec![
make("1. ", 72.0, 700.0, false),
make(
"Chemical and biological weapons threat model 1 (CB-1): Non-novel",
90.0,
700.0,
true,
),
]),
// wrapped continuation of the bold lead — all_bold, same indent
make_line(vec![make(
"chemical/biological weapons production capabilities: A model has CB-1",
90.0,
686.0,
true,
)]),
// body text of the same list item
make_line(vec![make(
"capabilities if it has the ability to significantly help individuals.",
90.0,
672.0,
false,
)]),
];
let md = to_markdown_from_lines_with_tables_and_images(
lines,
MarkdownOptions::default(),
HashMap::new(),
HashMap::new(),
&std::collections::HashSet::new(),
None,
);
assert!(
!md.contains("#### "),
"wrapped bold lead must not become a heading: {md}"
);
assert!(
md.lines().filter(|l| l.starts_with("- ")).count() == 0,
"continuation body must not become a stray bullet: {md}"
);
assert!(
md.contains("1. ") && md.contains("A model has CB-1"),
"numbered list item should remain intact: {md}"
);
}
}
+74 -9
View File
@@ -42,7 +42,7 @@ pub(crate) fn split_side_by_side(items: &[TextItem]) -> Vec<(f32, f32)> {
// Sort items by left edge
let mut xs: Vec<f32> = items.iter().map(|i| i.x).collect();
xs.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
xs.sort_by(|a, b| a.total_cmp(b));
// Find all candidate gaps: ≥30pt, in the middle 60% of the X range,
// with ≥20 items on each side.
@@ -118,7 +118,7 @@ pub(crate) fn split_side_by_side(items: &[TextItem]) -> Vec<(f32, f32)> {
})
.copied()
.collect();
balanced_positions.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
balanced_positions.sort_by(|a, b| a.total_cmp(b));
balanced_positions.dedup_by(|a, b| (*a - *b).abs() < 50.0);
if balanced_positions.len() > 1 {
return vec![];
@@ -209,7 +209,7 @@ fn split_from_hint_regions(items: &[TextItem], rects: &[PdfRect], page: u32) ->
// Width outlier filter (same as detect_tables_from_rects)
let mut widths: Vec<f32> = page_rects.iter().map(|&(_, _, w, _)| w).collect();
widths.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
widths.sort_by(|a, b| a.total_cmp(b));
let median_width = widths[widths.len() / 2];
page_rects.retain(|&(_, _, w, _)| w <= median_width * 10.0);
@@ -601,6 +601,15 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
let group = page_groups.get(&page).unwrap();
let page_items: Vec<TextItem> = group.iter().map(|(_, item)| (*item).clone()).collect();
// Detect columns early — on multi-column pages, the merged-band retry
// should skip body-font heuristic table detection (which mistakes column
// text for tables). Individual band heuristic detection is left enabled
// because bands are scoped to single columns.
let page_has_columns = {
let cols = crate::extractor::detect_columns(&page_items, page, false);
cols.len() >= 2
};
// Check for side-by-side layout (e.g. two tables placed left and right)
let mut bands = split_side_by_side(&page_items);
// Fallback: use rect hint regions to detect side-by-side layout
@@ -873,10 +882,7 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
run_heuristic(&unclaimed_items, &unclaimed_map, 6);
}
// 4. Column-based table detection: last resort for borderless tabular
// layouts (e.g. exam/reference grids) when ALL structural methods
// found nothing. Only runs when no rects/lines exist (truly borderless)
// and no other detection method found tables in this band.
// 4. Column-based table detection for borderless tabular layouts.
let band_has_tables = band_items.iter().enumerate().any(|(idx, _)| {
band_index_map
.get(idx)
@@ -903,6 +909,65 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
}
}
// 5. Thin-rect border synthesis: last resort for PDFs that draw table
// borders as thin filled rectangles (common in spreadsheet exports).
// Only runs when ALL other methods found nothing on this page.
if !page_tables.contains_key(&page) {
let page_rects: Vec<&crate::types::PdfRect> =
rects.iter().filter(|r| r.page == page).collect();
let mut synth_lines: Vec<crate::types::PdfLine> = Vec::new();
for r in &page_rects {
let (mut w, mut h) = (r.width, r.height);
let (mut x, mut y) = (r.x, r.y);
if w < 0.0 {
x += w;
w = -w;
}
if h < 0.0 {
y += h;
h = -h;
}
if h < 2.0 && w >= 10.0 {
let mid_y = y + h / 2.0;
synth_lines.push(crate::types::PdfLine {
x1: x,
y1: mid_y,
x2: x + w,
y2: mid_y,
page,
});
} else if w < 2.0 && h >= 10.0 {
let mid_x = x + w / 2.0;
synth_lines.push(crate::types::PdfLine {
x1: mid_x,
y1: y,
x2: mid_x,
y2: y + h,
page,
});
}
}
if synth_lines.len() >= 10 {
let page_text: Vec<TextItem> = text_items
.iter()
.filter(|i| i.page == page)
.cloned()
.collect();
let line_tables = detect_tables_from_lines(&page_text, &synth_lines, page);
for table in &line_tables {
for &idx in &table.item_indices {
table_items.insert(idx);
}
let table_y = table.rows.first().copied().unwrap_or(0.0);
let table_md = table_to_markdown(table);
page_tables
.entry(page)
.or_default()
.push((table_y, table_md));
}
}
}
// Merged-band retry: if we split into bands but found no tables in
// any band, retry heuristic detection with all items as a single band.
// This catches borderless tables whose text-column alignment was
@@ -915,7 +980,7 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
band_items.len(),
was_split
);
let heuristic_tables = detect_tables(band_items, base_size, false);
let heuristic_tables = detect_tables(band_items, base_size, page_has_columns);
for table in &heuristic_tables {
for &idx in &table.item_indices {
if let Some(&page_idx) = band_index_map.get(idx) {
@@ -1037,7 +1102,7 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
}
// Sort by Y descending (top to bottom) so left and right
// band lines interleave in visual reading order.
page_lines.sort_by(|a, b| b.y.partial_cmp(&a.y).unwrap_or(std::cmp::Ordering::Equal));
page_lines.sort_by(|a, b| b.y.total_cmp(&a.y));
all_lines.extend(page_lines);
}
all_lines
+1 -1
View File
@@ -275,7 +275,7 @@ pub(crate) fn strip_repeated_lines(lines: Vec<TextLine>, page_count: u32) -> Vec
page_sorted_ys.entry(line.page).or_default().push(line.y);
}
for ys in page_sorted_ys.values_mut() {
ys.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
ys.sort_by(|a, b| a.total_cmp(b));
ys.dedup();
}
+587
View File
@@ -0,0 +1,587 @@
//! PyO3 Python bindings for pdf-inspector.
use pyo3::exceptions::PyValueError;
use pyo3::prelude::*;
use std::collections::HashSet;
use crate::detector::PdfType;
use crate::types::ItemType;
// ---------------------------------------------------------------------------
// Result wrapper
// ---------------------------------------------------------------------------
/// Result of processing a PDF file.
#[pyclass(name = "PdfResult")]
#[derive(Clone)]
pub struct PyPdfResult {
/// The detected PDF type: "text_based", "scanned", "image_based", or "mixed".
#[pyo3(get)]
pub pdf_type: String,
/// Markdown output (None if detect-only or scanned PDF).
#[pyo3(get)]
pub markdown: Option<String>,
/// Total number of pages.
#[pyo3(get)]
pub page_count: u32,
/// Processing time in milliseconds.
#[pyo3(get)]
pub processing_time_ms: u64,
/// 1-indexed page numbers that need OCR.
#[pyo3(get)]
pub pages_needing_ocr: Vec<u32>,
/// Title from PDF metadata.
#[pyo3(get)]
pub title: Option<String>,
/// Detection confidence (0.0-1.0).
#[pyo3(get)]
pub confidence: f32,
/// Whether the layout is complex (tables/columns detected).
#[pyo3(get)]
pub is_complex_layout: bool,
/// Pages with tables detected.
#[pyo3(get)]
pub pages_with_tables: Vec<u32>,
/// Pages with multi-column layout.
#[pyo3(get)]
pub pages_with_columns: Vec<u32>,
/// Whether encoding issues were detected.
#[pyo3(get)]
pub has_encoding_issues: bool,
}
#[pymethods]
impl PyPdfResult {
fn __repr__(&self) -> String {
format!(
"PdfResult(pdf_type='{}', pages={}, confidence={:.2})",
self.pdf_type, self.page_count, self.confidence
)
}
}
// ---------------------------------------------------------------------------
// Classification wrapper (lightweight)
// ---------------------------------------------------------------------------
/// Lightweight PDF classification result.
#[pyclass(name = "PdfClassification")]
#[derive(Clone)]
pub struct PyPdfClassification {
/// The detected PDF type: "text_based", "scanned", "image_based", or "mixed".
#[pyo3(get)]
pub pdf_type: String,
/// Total number of pages.
#[pyo3(get)]
pub page_count: u32,
/// 0-indexed page numbers that need OCR.
#[pyo3(get)]
pub pages_needing_ocr: Vec<u32>,
/// Detection confidence (0.0-1.0).
#[pyo3(get)]
pub confidence: f32,
}
#[pymethods]
impl PyPdfClassification {
fn __repr__(&self) -> String {
format!(
"PdfClassification(pdf_type='{}', pages={}, confidence={:.2})",
self.pdf_type, self.page_count, self.confidence
)
}
}
// ---------------------------------------------------------------------------
// Region extraction wrappers
// ---------------------------------------------------------------------------
/// Extracted text for a single region.
#[pyclass(name = "RegionText")]
#[derive(Clone)]
pub struct PyRegionText {
/// Extracted text content.
#[pyo3(get)]
pub text: String,
/// True when the text should not be trusted (empty, GID fonts, garbage, encoding issues).
#[pyo3(get)]
pub needs_ocr: bool,
}
#[pymethods]
impl PyRegionText {
fn __repr__(&self) -> String {
format!(
"RegionText(text='{}', needs_ocr={})",
self.text.chars().take(40).collect::<String>(),
self.needs_ocr
)
}
}
/// Extracted text for one page's regions.
#[pyclass(name = "PageRegionTexts")]
#[derive(Clone)]
pub struct PyPageRegionTexts {
/// 0-indexed page number.
#[pyo3(get)]
pub page: u32,
/// Per-region results, parallel to the input regions.
#[pyo3(get)]
pub regions: Vec<PyRegionText>,
}
#[pymethods]
impl PyPageRegionTexts {
fn __repr__(&self) -> String {
format!(
"PageRegionTexts(page={}, regions={})",
self.page,
self.regions.len()
)
}
}
// ---------------------------------------------------------------------------
// Text item wrapper
// ---------------------------------------------------------------------------
/// Per-page markdown extraction result.
#[pyclass(name = "PageMarkdown")]
#[derive(Clone)]
pub struct PyPageMarkdown {
/// 0-indexed page number.
#[pyo3(get)]
pub page: u32,
/// Formatted markdown for this page.
#[pyo3(get)]
pub markdown: String,
/// True when text on this page is unreliable (GID-encoded fonts,
/// encoding issues, garbage text, or empty extraction).
#[pyo3(get)]
pub needs_ocr: bool,
}
#[pymethods]
impl PyPageMarkdown {
fn __repr__(&self) -> String {
format!(
"PageMarkdown(page={}, markdown='{}', needs_ocr={})",
self.page,
self.markdown.chars().take(40).collect::<String>(),
self.needs_ocr
)
}
}
/// Combined per-page markdown extraction and layout classification result.
#[pyclass(name = "PagesExtractionResult")]
#[derive(Clone)]
pub struct PyPagesExtractionResult {
/// Per-page markdown results, in the order requested.
#[pyo3(get)]
pub pages: Vec<PyPageMarkdown>,
/// 1-indexed pages where tables were detected.
#[pyo3(get)]
pub pages_with_tables: Vec<u32>,
/// 1-indexed pages where multi-column layout was detected.
#[pyo3(get)]
pub pages_with_columns: Vec<u32>,
/// 1-indexed pages that need OCR (scanned/image-based or unreliable text).
#[pyo3(get)]
pub pages_needing_ocr: Vec<u32>,
/// True if any page has tables or columns.
#[pyo3(get)]
pub is_complex: bool,
}
#[pymethods]
impl PyPagesExtractionResult {
fn __repr__(&self) -> String {
format!(
"PagesExtractionResult(pages={}, pages_with_tables={:?}, is_complex={})",
self.pages.len(),
self.pages_with_tables,
self.is_complex
)
}
}
/// A positioned text item extracted from a PDF.
#[pyclass(name = "TextItem")]
#[derive(Clone)]
pub struct PyTextItem {
#[pyo3(get)]
pub text: String,
#[pyo3(get)]
pub x: f32,
#[pyo3(get)]
pub y: f32,
#[pyo3(get)]
pub width: f32,
#[pyo3(get)]
pub height: f32,
#[pyo3(get)]
pub font: String,
#[pyo3(get)]
pub font_size: f32,
#[pyo3(get)]
pub page: u32,
#[pyo3(get)]
pub is_bold: bool,
#[pyo3(get)]
pub is_italic: bool,
#[pyo3(get)]
pub item_type: String,
}
#[pymethods]
impl PyTextItem {
fn __repr__(&self) -> String {
format!(
"TextItem(text='{}', page={}, x={:.1}, y={:.1})",
self.text.chars().take(40).collect::<String>(),
self.page,
self.x,
self.y,
)
}
}
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
fn pdf_type_str(t: PdfType) -> String {
match t {
PdfType::TextBased => "text_based".into(),
PdfType::Scanned => "scanned".into(),
PdfType::ImageBased => "image_based".into(),
PdfType::Mixed => "mixed".into(),
}
}
fn to_py_result(r: crate::PdfProcessResult) -> PyPdfResult {
PyPdfResult {
pdf_type: pdf_type_str(r.pdf_type),
markdown: r.markdown,
page_count: r.page_count,
processing_time_ms: r.processing_time_ms,
pages_needing_ocr: r.pages_needing_ocr,
title: r.title,
confidence: r.confidence,
is_complex_layout: r.layout.is_complex,
pages_with_tables: r.layout.pages_with_tables,
pages_with_columns: r.layout.pages_with_columns,
has_encoding_issues: r.has_encoding_issues,
}
}
fn to_py_err(e: crate::PdfError) -> PyErr {
PyValueError::new_err(e.to_string())
}
fn item_type_str(t: &ItemType) -> String {
match t {
ItemType::Text => "text".into(),
ItemType::Image => "image".into(),
ItemType::Link(url) => format!("link:{url}"),
ItemType::FormField => "form_field".into(),
}
}
fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
items
.into_iter()
.map(|item| PyTextItem {
text: item.text,
x: item.x,
y: item.y,
width: item.width,
height: item.height,
font: item.font,
font_size: item.font_size,
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
item_type: item_type_str(&item.item_type),
})
.collect()
}
fn parse_page_regions(
page_regions: Vec<(u32, Vec<Vec<f64>>)>,
) -> PyResult<Vec<(u32, Vec<[f32; 4]>)>> {
page_regions
.into_iter()
.map(|(page, regions)| {
let mut bboxes: Vec<[f32; 4]> = Vec::with_capacity(regions.len());
for (idx, region) in regions.into_iter().enumerate() {
if region.len() != 4 {
return Err(PyValueError::new_err(format!(
"Invalid region at page {page}, index {idx}: expected [x1, y1, x2, y2], got {} values",
region.len()
)));
}
let [x1, y1, x2, y2] = [region[0], region[1], region[2], region[3]];
if !(x1.is_finite() && y1.is_finite() && x2.is_finite() && y2.is_finite()) {
return Err(PyValueError::new_err(format!(
"Invalid region at page {page}, index {idx}: coordinates must be finite numbers"
)));
}
if x2 < x1 || y2 < y1 {
return Err(PyValueError::new_err(format!(
"Invalid region at page {page}, index {idx}: expected x2>=x1 and y2>=y1, got [{x1}, {y1}, {x2}, {y2}]"
)));
}
bboxes.push([x1 as f32, y1 as f32, x2 as f32, y2 as f32]);
}
Ok((page, bboxes))
})
.collect()
}
fn to_py_pages_result(r: crate::PagesExtractionResult) -> PyPagesExtractionResult {
PyPagesExtractionResult {
pages: r
.pages
.into_iter()
.map(|p| PyPageMarkdown {
page: p.page,
markdown: p.markdown,
needs_ocr: p.needs_ocr,
})
.collect(),
pages_with_tables: r.pages_with_tables,
pages_with_columns: r.pages_with_columns,
pages_needing_ocr: r.pages_needing_ocr,
is_complex: r.is_complex,
}
}
fn convert_region_results(results: Vec<crate::PageRegionResult>) -> Vec<PyPageRegionTexts> {
results
.into_iter()
.map(|page_result| PyPageRegionTexts {
page: page_result.page,
regions: page_result
.regions
.into_iter()
.map(|r| PyRegionText {
text: r.text,
needs_ocr: r.needs_ocr,
})
.collect(),
})
.collect()
}
// ---------------------------------------------------------------------------
// Public Python API
// ---------------------------------------------------------------------------
/// Process a PDF file: detect type, extract text, and convert to Markdown.
#[pyfunction]
#[pyo3(signature = (path, pages=None))]
fn process_pdf(path: &str, pages: Option<Vec<u32>>) -> PyResult<PyPdfResult> {
let mut opts = crate::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = crate::process_pdf_with_options(path, opts).map_err(to_py_err)?;
Ok(to_py_result(result))
}
/// Process a PDF from bytes in memory.
#[pyfunction]
#[pyo3(signature = (data, pages=None))]
fn process_pdf_bytes(data: &[u8], pages: Option<Vec<u32>>) -> PyResult<PyPdfResult> {
let mut opts = crate::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = crate::process_pdf_mem_with_options(data, opts).map_err(to_py_err)?;
Ok(to_py_result(result))
}
/// Fast detection only — no text extraction or markdown.
#[pyfunction]
fn detect_pdf(path: &str) -> PyResult<PyPdfResult> {
let result = crate::detect_pdf(path).map_err(to_py_err)?;
Ok(to_py_result(result))
}
/// Fast detection from bytes — no text extraction or markdown.
#[pyfunction]
fn detect_pdf_bytes(data: &[u8]) -> PyResult<PyPdfResult> {
let result = crate::detect_pdf_mem(data).map_err(to_py_err)?;
Ok(to_py_result(result))
}
/// Lightweight PDF classification — returns type, page count, and OCR pages.
/// Faster than detect_pdf as it skips building the full PdfProcessResult.
/// Pages in pages_needing_ocr are 0-indexed.
#[pyfunction]
fn classify_pdf(path: &str) -> PyResult<PyPdfClassification> {
let data = std::fs::read(path).map_err(|e| PyValueError::new_err(e.to_string()))?;
classify_pdf_bytes(&data)
}
/// Lightweight PDF classification from bytes.
/// Pages in pages_needing_ocr are 0-indexed.
#[pyfunction]
fn classify_pdf_bytes(data: &[u8]) -> PyResult<PyPdfClassification> {
let result = crate::classify_pdf_mem(data).map_err(to_py_err)?;
Ok(PyPdfClassification {
pdf_type: pdf_type_str(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence,
})
}
/// Extract plain text from a PDF file.
#[pyfunction]
fn extract_text(path: &str) -> PyResult<String> {
crate::extract_text(path).map_err(to_py_err)
}
/// Extract plain text from PDF bytes.
#[pyfunction]
fn extract_text_bytes(data: &[u8]) -> PyResult<String> {
crate::extractor::extract_text_mem(data).map_err(to_py_err)
}
/// Extract text with position information from a file.
#[pyfunction]
#[pyo3(signature = (path, pages=None))]
fn extract_text_with_positions(path: &str, pages: Option<Vec<u32>>) -> PyResult<Vec<PyTextItem>> {
let items = match pages {
Some(p) => {
let page_set: HashSet<u32> = p.into_iter().collect();
crate::extract_text_with_positions_pages(path, Some(&page_set)).map_err(to_py_err)?
}
None => crate::extract_text_with_positions(path).map_err(to_py_err)?,
};
Ok(convert_text_items(items))
}
/// Extract text with position information from bytes.
#[pyfunction]
#[pyo3(signature = (data, pages=None))]
fn extract_text_with_positions_bytes(
data: &[u8],
pages: Option<Vec<u32>>,
) -> PyResult<Vec<PyTextItem>> {
let items = match pages {
Some(p) => {
let page_set: HashSet<u32> = p.into_iter().collect();
crate::extractor::extract_text_with_positions_mem_pages(data, Some(&page_set))
.map_err(to_py_err)?
}
None => crate::extractor::extract_text_with_positions_mem(data).map_err(to_py_err)?,
};
Ok(convert_text_items(items))
}
/// Extract text within bounding-box regions from a PDF file.
///
/// Args:
/// path: Path to the PDF file.
/// page_regions: List of (page_0indexed, [[x1, y1, x2, y2], ...]) tuples.
/// Coordinates are PDF points with top-left origin.
///
/// Returns:
/// List of PageRegionTexts with per-region text and needs_ocr flag.
#[pyfunction]
fn extract_text_in_regions(
path: &str,
page_regions: Vec<(u32, Vec<Vec<f64>>)>,
) -> PyResult<Vec<PyPageRegionTexts>> {
let data = std::fs::read(path).map_err(|e| PyValueError::new_err(e.to_string()))?;
extract_text_in_regions_bytes(&data, page_regions)
}
/// Extract text within bounding-box regions from PDF bytes.
///
/// Args:
/// data: PDF file contents as bytes.
/// page_regions: List of (page_0indexed, [[x1, y1, x2, y2], ...]) tuples.
/// Coordinates are PDF points with top-left origin.
///
/// Returns:
/// List of PageRegionTexts with per-region text and needs_ocr flag.
#[pyfunction]
fn extract_text_in_regions_bytes(
data: &[u8],
page_regions: Vec<(u32, Vec<Vec<f64>>)>,
) -> PyResult<Vec<PyPageRegionTexts>> {
let regions = parse_page_regions(page_regions)?;
let results = crate::extract_text_in_regions_mem(data, &regions).map_err(to_py_err)?;
Ok(convert_region_results(results))
}
/// Extract formatted markdown for pages of a PDF file, with layout
/// classification metadata.
///
/// Returns per-page markdown and classification data (tables, columns,
/// OCR needs) from a single parse. Font statistics are computed from the
/// full document so header detection is consistent across pages.
///
/// Args:
/// path: Path to the PDF file.
/// pages: Optional list of 0-indexed pages. When None (default), every
/// page is returned in document order. When provided, output
/// matches the caller-supplied order.
///
/// Returns:
/// PagesExtractionResult with per-page markdown and classification data.
#[pyfunction]
#[pyo3(signature = (path, pages=None))]
fn extract_pages_markdown(
path: &str,
pages: Option<Vec<u32>>,
) -> PyResult<PyPagesExtractionResult> {
let result = crate::extract_pages_markdown(path, pages.as_deref()).map_err(to_py_err)?;
Ok(to_py_pages_result(result))
}
/// Extract formatted markdown for pages of a PDF from bytes.
///
/// See [`extract_pages_markdown`] for details.
#[pyfunction]
#[pyo3(signature = (data, pages=None))]
fn extract_pages_markdown_bytes(
data: &[u8],
pages: Option<Vec<u32>>,
) -> PyResult<PyPagesExtractionResult> {
let result = crate::extract_pages_markdown_mem(data, pages.as_deref()).map_err(to_py_err)?;
Ok(to_py_pages_result(result))
}
/// Python module definition.
#[pymodule]
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_class::<PyPdfResult>()?;
m.add_class::<PyPdfClassification>()?;
m.add_class::<PyTextItem>()?;
m.add_class::<PyRegionText>()?;
m.add_class::<PyPageRegionTexts>()?;
m.add_class::<PyPageMarkdown>()?;
m.add_class::<PyPagesExtractionResult>()?;
m.add_function(wrap_pyfunction!(process_pdf, m)?)?;
m.add_function(wrap_pyfunction!(process_pdf_bytes, m)?)?;
m.add_function(wrap_pyfunction!(detect_pdf, m)?)?;
m.add_function(wrap_pyfunction!(detect_pdf_bytes, m)?)?;
m.add_function(wrap_pyfunction!(classify_pdf, m)?)?;
m.add_function(wrap_pyfunction!(classify_pdf_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_with_positions, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_with_positions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_pages_markdown, m)?)?;
m.add_function(wrap_pyfunction!(extract_pages_markdown_bytes, m)?)?;
Ok(())
}
+578 -76
View File
@@ -48,7 +48,7 @@ pub(crate) fn merge_adjacent_items(items: &[TextItem]) -> (Vec<TextItem>, Vec<Ve
}
// Sort groups by Y descending (top of page first)
line_groups.sort_by(|a, b| b.0.partial_cmp(&a.0).unwrap_or(std::cmp::Ordering::Equal));
line_groups.sort_by(|a, b| b.0.total_cmp(&a.0));
let mut merged_items = Vec::new();
let mut index_map: Vec<Vec<usize>> = Vec::new();
@@ -219,7 +219,7 @@ pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bo
body_font_low,
body_font_high,
);
if body_candidates.len() >= 9 {
if body_candidates.len() >= 6 {
let regions = find_table_regions_strict(&body_candidates);
log::debug!("body-font: {} strict regions found", regions.len());
@@ -241,7 +241,7 @@ pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bo
body_candidates.len()
);
if region_items.len() < 9 {
if region_items.len() < 6 {
continue;
}
@@ -284,7 +284,7 @@ fn find_table_regions(items: &[(usize, &TextItem)]) -> Vec<(f32, f32)> {
}
let mut y_positions: Vec<f32> = items.iter().map(|(_, i)| i.y).collect();
y_positions.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
y_positions.sort_by(|a, b| a.total_cmp(b));
// Find clusters of Y positions (table regions)
let mut regions = Vec::new();
@@ -347,7 +347,7 @@ fn find_table_regions_strict(items: &[(usize, &TextItem)]) -> Vec<(f32, f32, f32
let mut qualifying_rows: Vec<(f32, Vec<f32>)> = Vec::new(); // (y, cluster_starts)
for (y, x_positions) in &row_groups {
let mut sorted_xs = x_positions.clone();
sorted_xs.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
sorted_xs.sort_by(|a, b| a.total_cmp(b));
if sorted_xs.is_empty() {
continue;
@@ -379,14 +379,14 @@ fn find_table_regions_strict(items: &[(usize, &TextItem)]) -> Vec<(f32, f32, f32
// Step 3: Find contiguous runs of qualifying rows.
// Use adaptive gap: median spacing × 3 (handles wrapped cells where
// qualifying rows are spaced further apart), with a floor of 25pt.
qualifying_rows.sort_by(|a, b| a.0.partial_cmp(&b.0).unwrap_or(std::cmp::Ordering::Equal));
qualifying_rows.sort_by(|a, b| a.0.total_cmp(&b.0));
let max_gap = if qualifying_rows.len() >= 3 {
let mut gaps: Vec<f32> = qualifying_rows
.windows(2)
.map(|w| (w[1].0 - w[0].0).abs())
.collect();
gaps.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
gaps.sort_by(|a, b| a.total_cmp(b));
let median_gap = gaps[gaps.len() / 2];
(median_gap * 3.0).max(25.0)
} else {
@@ -566,11 +566,9 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
// Sort by X position (direction-aware)
let rtl = is_rtl_text(col_items.iter().map(|i| &i.text));
if rtl {
col_items
.sort_by(|a, b| b.x.partial_cmp(&a.x).unwrap_or(std::cmp::Ordering::Equal));
col_items.sort_by(|a, b| b.x.total_cmp(&a.x));
} else {
col_items
.sort_by(|a, b| a.x.partial_cmp(&b.x).unwrap_or(std::cmp::Ordering::Equal));
col_items.sort_by(|a, b| a.x.total_cmp(&b.x));
}
// Join items with subscript-aware spacing
@@ -583,8 +581,14 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
// Validation 1: some rows should have content in first column.
// Use a lower threshold (25%) for tables with wrapped cells where
// continuation lines leave the first column empty.
// Skip when cells form a narrow TOC pattern: hierarchical entries indented
// across multiple X levels leave the leftmost column sparse (only top-level
// chapters land there) but the structure is still a valid TOC. Narrow only
// (<=5 cols) — wide multi-column TOCs (e.g. 2-up indices) would render
// poorly through format_toc_as_list, which assumes one entry per row.
let rows_with_first_col = cells.iter().filter(|row| !row[0].is_empty()).count();
if rows_with_first_col < rows.len() / 4 {
let is_narrow_toc = columns.len() <= 5 && is_table_of_contents(&cells);
if rows_with_first_col < rows.len() / 4 && !is_narrow_toc {
log::debug!(
" validation 1 fail: {}/{} rows have first col",
rows_with_first_col,
@@ -655,15 +659,24 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
return None;
}
// Validation 8: Check for Table of Contents pattern
if is_table_of_contents(&cells) {
log::debug!(" validation 8 fail: table of contents");
// Validation 8: Reject paragraph-like content falsely detected as tables.
// TOC pages with deep indentation (top-level chapters in col 0, subsections
// in cols 1-3, page numbers in last col) leave most cells empty and trip
// the paragraph heuristic; TOC shape is a safer signal here. Narrow only
// — see narrow-TOC rationale at validation 1.
if is_paragraph_content(&cells) && !is_narrow_toc {
log::debug!(" validation 9 fail: paragraph content");
return None;
}
// Validation 9: Reject paragraph-like content falsely detected as tables
if is_paragraph_content(&cells) {
log::debug!(" validation 9 fail: paragraph content");
// Validation 9: Reject wide "index" layouts where every cell carries a
// full "label ... page" fragment (back-of-book IRS-style indices).
// These render poorly in any structured form; text flow is the best
// fallback. Narrow dot-leader TOCs (2-3 cols) are kept so format.rs
// can emit them as a per-row flat list with titles tab-joined to page
// numbers.
if is_inline_leader_index(&cells) {
log::debug!(" validation 9 fail: inline-leader index");
return None;
}
@@ -674,12 +687,7 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
item_indices.len()
);
Some(Table {
columns,
rows,
cells,
item_indices,
})
Some(Table::new(columns, rows, cells, item_indices))
}
/// Check if this looks like a key-value pair layout rather than a table
@@ -803,7 +811,25 @@ fn has_table_like_content(cells: &[Vec<String>], mode: TableDetectionMode) -> bo
// Bypass content check for wide tables (3+ columns) — text-only tables
// (category lists, program descriptions) are legitimate if they passed
// all structural validations (alignment, consistency, not key-value).
pct_data > min_pct || num_cols >= 3
// Also bypass for 2-column body-font tables with short cells (avg ≤40 chars),
// which are likely definition/category lists, not paragraph text.
if pct_data > min_pct || num_cols >= 3 {
return true;
}
if num_cols == 2 && matches!(mode, TableDetectionMode::BodyFont) {
let non_empty: Vec<usize> = cells
.iter()
.skip(1)
.flat_map(|row| row.iter())
.filter(|c| !c.trim().is_empty())
.map(|c| c.trim().len())
.collect();
if !non_empty.is_empty() {
let avg_len = non_empty.iter().sum::<usize>() / non_empty.len();
return avg_len <= 25;
}
}
false
}
/// Check if a cell value looks like table data
@@ -882,77 +908,247 @@ fn looks_like_number(s: &str) -> bool {
&& s.chars().any(|c| c.is_ascii_digit())
}
/// Check if this looks like a Table of Contents
/// TOCs have characteristic patterns: leader dots, page numbers, section names
fn is_table_of_contents(cells: &[Vec<String>]) -> bool {
/// Check if this looks like a Table of Contents (either style).
///
/// Used by format.rs to render TOCs as flat lists instead of markdown tables.
pub fn is_table_of_contents(cells: &[Vec<String>]) -> bool {
is_dot_leader_toc(cells) || is_tabular_toc(cells)
}
/// Dot-leader TOC: any "Chapter 1 ........ 42" style with explicit leader
/// dots. Covers both narrow 2-3 col TOCs (where the leader is a dedicated
/// cell) and wide indices (where each cell encodes a full "label ... page"
/// fragment). Used by format.rs to render as a flat list.
pub(super) fn is_dot_leader_toc(cells: &[Vec<String>]) -> bool {
has_structural_dot_leader(cells) || is_inline_leader_index(cells)
}
/// Rows with a dedicated dots-only cell flanked by label + number (2-3 col
/// TOC layout). Format.rs handles these well via per-row flat-list
/// rendering; they should NOT be rejected at detect time.
fn has_structural_dot_leader(cells: &[Vec<String>]) -> bool {
if cells.is_empty() {
return false;
}
let structural_rows = cells.iter().filter(|row| row_has_dot_leader(row)).count();
structural_rows as f32 / cells.len() as f32 >= 0.3
}
let num_cols = cells[0].len();
let mut dot_cells = 0;
let mut page_number_cells = 0;
let mut total_cells = 0;
// Track which columns contain dots vs numbers to distinguish
// TOC (dots span middle, page number at end) from data tables
// (dots only in label column, many number columns).
let mut dot_cols = vec![0u32; num_cols];
let mut numeric_cols = vec![0u32; num_cols];
/// Wide index layout: each cell holds a full "label ... page" fragment
/// because the column detector kept multi-column indices as single cells.
/// These render poorly both as markdown tables (column boundaries are
/// arbitrary) and as flat lists (each row holds 3+ separate index
/// entries). Reject these at detect time so they fall back to the page's
/// normal text flow.
pub(super) fn is_inline_leader_index(cells: &[Vec<String>]) -> bool {
let mut inline_cells = 0;
let mut total_nonempty = 0;
for row in cells {
for (ci, cell) in row.iter().enumerate() {
for cell in row {
let trimmed = cell.trim();
if trimmed.is_empty() {
continue;
}
total_cells += 1;
// Check for leader dots (sequences of periods)
// TOCs often have "........" or ". . . ." patterns
let dot_count = trimmed.chars().filter(|&c| c == '.').count();
let is_mostly_dots = dot_count > trimmed.len() / 2 && dot_count >= 3;
if is_mostly_dots {
dot_cells += 1;
if ci < num_cols {
dot_cols[ci] += 1;
}
}
// Check for standalone page numbers (1-4 digits, possibly with spaces)
let digits_only: String = trimmed.chars().filter(|c| !c.is_whitespace()).collect();
if digits_only.len() <= 4
&& !digits_only.is_empty()
&& digits_only.chars().all(|c| c.is_ascii_digit())
{
page_number_cells += 1;
if ci < num_cols {
numeric_cols[ci] += 1;
}
total_nonempty += 1;
if cell_is_inline_leader(trimmed) {
inline_cells += 1;
}
}
}
total_nonempty >= 4 && inline_cells as f32 / total_nonempty as f32 >= 0.25
}
if total_cells == 0 {
/// A row with a dot-leader. Accepts two layouts:
/// 1. A dedicated dots-only cell ("....") with a text label somewhere
/// to its left and a page number somewhere to its right.
/// 2. A "title ... " cell (trailing leader dots glued to the title)
/// with a page number elsewhere in the same row.
fn row_has_dot_leader(row: &[String]) -> bool {
let has_page_number = row.iter().any(|c| row_cell_is_page_number(c));
for (ci, cell) in row.iter().enumerate() {
let trimmed = cell.trim();
// Pattern 1: dedicated dots-only cell.
let dot_count = trimmed.chars().filter(|&c| c == '.').count();
let is_mostly_dots = dot_count >= 3
&& dot_count > trimmed.len() / 2
&& trimmed.chars().all(|c| c == '.' || c.is_whitespace());
if is_mostly_dots {
let has_label_left = row[..ci].iter().any(|c| {
let t = c.trim();
!t.is_empty() && t.chars().any(|ch| ch.is_alphabetic())
});
if has_label_left && has_page_number {
return true;
}
continue;
}
// Pattern 2: cell ends with a trailing " ... " run after a label.
if has_page_number && cell_has_trailing_leader(trimmed) {
return true;
}
}
false
}
/// Cell ends with a run of ≥3 dots preceded by alphabetic text and a
/// space — the "Title ... " layout where the leader is glued to the name.
/// Alphabetic (not alphanumeric) so that data-table row labels like
/// "1973 ... " do not register as titles.
fn cell_has_trailing_leader(cell: &str) -> bool {
let trimmed = cell.trim_end();
if !trimmed.ends_with('.') {
return false;
}
let without_dots = trimmed.trim_end_matches('.');
let dot_run = trimmed.len() - without_dots.len();
if dot_run < 3 {
return false;
}
// Require a space before the dot run (rules out "etc..." / "Mr...") and
// at least one alphabetic char (rules out "1973 ... " data-row labels).
without_dots.ends_with(' ') && without_dots.trim().chars().any(|c| c.is_alphabetic())
}
/// Page-number shape: single ≤4-digit integer, a ", "-separated list of
/// ≤4-digit integers ("18, 36, 107"), or a dashed section-page ID
/// ("A-1", "5-21"). Rejects decimal cells ("4. 0"), thousands-separated
/// values ("189,164"), and other long numeric data that appears in
/// statistical tables.
fn row_cell_is_page_number(cell: &str) -> bool {
let t = cell.trim();
if t.is_empty() {
return false;
}
if looks_like_section_page_id(t) {
return true;
}
// Page list: ", " separator (with space) distinguishes real page lists
// from thousands-separated numbers like "189,164".
let parts: Vec<&str> = t.split(", ").collect();
parts
.iter()
.all(|p| !p.is_empty() && p.len() <= 4 && p.chars().all(|c| c.is_ascii_digit()))
}
/// A cell shaped like an index leader fragment. Accepts two forms:
/// - "text ... number" — label + dots + page number in one cell
/// - "... number" — bare leader + number (row where the label
/// landed in a separate column)
///
/// Both only count if followed by pure numeric content (optionally
/// comma-separated page lists like "127, 213").
fn cell_is_inline_leader(cell: &str) -> bool {
let cell = cell.trim();
// Find the first "..." run. Surrounding-whitespace checks below
// reject intra-word ellipses ("etc...").
let idx = match cell.match_indices("...").next() {
Some((i, _)) => i,
None => return false,
};
let before = &cell[..idx];
let after_dots = &cell[idx + 3..];
// Allow extra dots (e.g. "....") by skipping any additional '.'
let after = after_dots.trim_start_matches('.');
// Require space (or start-of-cell) before the dots and space/digit
// after — blocks intra-word ellipses.
let before_ok = before.is_empty() || before.ends_with(' ');
let after_ok = after.starts_with(' ') || after.is_empty();
if !before_ok || !after_ok {
return false;
}
// Data tables with dot leaders (e.g. "1973....") have dots concentrated
// in one column (the label column) while many other columns contain numbers.
// True TOCs have dots spanning the middle and one page-number column at the end.
// If dots are confined to ≤1 column AND there are ≥3 columns with numbers,
// this is a data table, not a TOC.
let cols_with_dots = dot_cols.iter().filter(|&&c| c >= 2).count();
let cols_with_numbers = numeric_cols.iter().filter(|&&c| c >= 2).count();
if cols_with_dots <= 1 && cols_with_numbers >= 3 {
let after_trim = after.trim();
if after_trim.is_empty() {
return false;
}
// Tail must be purely numeric/page-list content.
let tail_numeric = after_trim
.chars()
.all(|c| c.is_ascii_digit() || matches!(c, ',' | ' ' | '.' | '-' | '$'))
&& after_trim.chars().any(|c| c.is_ascii_digit());
if !tail_numeric {
return false;
}
// If a significant portion of cells are dots or page numbers, it's likely a TOC
let dot_ratio = dot_cells as f32 / total_cells as f32;
let page_num_ratio = page_number_cells as f32 / total_cells as f32;
// Either we have a label before, or the leader is bare (starts the cell)
// — both are legitimate index fragments.
before.chars().any(|c| c.is_alphabetic()) || before.trim().is_empty()
}
// TOC typically has >15% dot cells and >10% page number cells
dot_ratio > 0.15 || (dot_ratio > 0.05 && page_num_ratio > 0.15)
/// Dot-less tabular TOC: tagged PDFs emit entries as rows where the first
/// column starts with a dotted section number (e.g. "4.3.1 Something") and
/// the last column is one or more page numbers. These have no leader dots
/// and benefit from flat-list formatting (page numbers aligned to titles).
pub(super) fn is_tabular_toc(cells: &[Vec<String>]) -> bool {
if cells.is_empty() {
return false;
}
let num_cols = cells[0].len();
if num_cols < 2 || cells.len() < 4 {
return false;
}
let section_rows = cells
.iter()
.filter(|row| {
row.iter()
.find(|c| !c.trim().is_empty())
.is_some_and(|c| starts_with_section_number(c.trim()))
})
.count();
let last_col = num_cols - 1;
let (last_filled, last_page_num) = cells.iter().fold((0u32, 0u32), |(f, n), row| {
let cell = row.get(last_col).map(|s| s.trim()).unwrap_or("");
if cell.is_empty() {
return (f, n);
}
let is_page_nums = cell
.split_whitespace()
.all(|tok| !tok.is_empty() && tok.chars().all(|c| c.is_ascii_digit()));
(f + 1, n + if is_page_nums { 1 } else { 0 })
});
let section_ratio = section_rows as f32 / cells.len() as f32;
let page_num_last_ratio = if last_filled > 0 {
last_page_num as f32 / last_filled as f32
} else {
0.0
};
section_ratio >= 0.6 && last_filled >= 3 && page_num_last_ratio >= 0.7
}
/// Matches dashed section-page identifiers used in technical manuals:
/// "5-21", "A-1", "B--3", "TC-2". At least one ASCII digit is required.
fn looks_like_section_page_id(s: &str) -> bool {
let ok = s
.chars()
.all(|c| c.is_ascii_digit() || c.is_ascii_uppercase() || c == '-');
ok && s.chars().any(|c| c.is_ascii_digit())
}
/// Returns true when the leading token looks like a dotted section number:
/// "1", "1.2", "1.2.3", "4.3.1.2" — integer components joined by dots,
/// with at least one dot (single-number prefixes are too ambiguous).
fn starts_with_section_number(s: &str) -> bool {
let Some(first) = s.split_whitespace().next() else {
return false;
};
let first = first.trim_end_matches('.');
let parts: Vec<&str> = first.split('.').collect();
if parts.len() < 2 || parts.len() > 6 {
return false;
}
parts
.iter()
.all(|p| !p.is_empty() && p.len() <= 3 && p.chars().all(|c| c.is_ascii_digit()))
}
/// Check if detected "table" cells are actually paragraph text fragments.
@@ -1128,6 +1324,33 @@ pub(crate) fn find_first_table_row(
continue;
}
// Skip rows that have duplicate non-empty cells. These are spanning
// super-headers (e.g., "First Degree | First Degree | Higher Degree")
// that sit above the real column header row. Using them as the markdown
// header produces duplicate column names that downstream validation
// rejects. Only skip if a subsequent row looks like a better header
// (denser fill or has data).
if filled_count >= 2 && !has_data {
let mut text_counts: std::collections::HashMap<&str, usize> =
std::collections::HashMap::new();
for cell in &filled_cells {
*text_counts.entry(cell.trim()).or_insert(0) += 1;
}
let has_duplicates = text_counts.values().any(|&count| count >= 2);
if has_duplicates {
// Check if a later row is a better header candidate
let has_better_below = cells.iter().skip(row_idx + 1).take(3).any(|r| {
let next_filled = r.iter().filter(|c| !c.trim().is_empty()).count();
let next_fill = next_filled as f32 / total_cols as f32;
let next_numeric = r.iter().filter(|c| looks_like_number(c.trim())).count();
next_fill >= 0.4 || next_numeric >= 2
});
if has_better_below {
continue;
}
}
}
// Data rows are definitely table content
if has_data {
first_table_row = row_idx;
@@ -1382,4 +1605,283 @@ mod tests {
"data table with dot-leader labels should not be rejected as TOC"
);
}
#[test]
fn is_table_of_contents_accepts_hierarchical_indented_toc() {
// Mythos system card pages 4-5: top-level chapters indent at col 0,
// subsections at cols 1-2, leaving col 0 mostly empty (only ~10% of
// rows). Validation 1 was rejecting these even though the structure
// is unambiguously a TOC.
let cells = vec![
vec!["Abstract".to_string(), String::new(), "3".to_string()],
vec![
"1 Introduction".to_string(),
String::new(),
"10".to_string(),
],
vec![
String::new(),
"1.1 Model training".to_string(),
"11".to_string(),
],
vec![
String::new(),
"1.1.1 Training data".to_string(),
"11".to_string(),
],
vec![
String::new(),
"1.1.2 Crowd workers".to_string(),
"12".to_string(),
],
vec![
String::new(),
"1.2 Release decision".to_string(),
"13".to_string(),
],
vec![
"2 RSP evaluations".to_string(),
String::new(),
"16".to_string(),
],
vec![
String::new(),
"2.1 RSP risk assessment".to_string(),
"16".to_string(),
],
vec![String::new(), "2.1.1 Context".to_string(), "16".to_string()],
vec![
String::new(),
"2.2 CB evaluations".to_string(),
"20".to_string(),
],
];
assert!(
is_table_of_contents(&cells),
"hierarchical TOC with sparse col 0 should still be detected"
);
}
#[test]
fn is_table_of_contents_rejects_dotless_toc() {
// Tabular TOC without leader dots: first column starts with dotted
// section numbers, last column is page numbers. Pattern from
// Mythos system card pages 6-8.
let cells = vec![
vec![
"4.3 Case studies and targeted evaluations".to_string(),
String::new(),
"86".to_string(),
],
vec![
"4.3.1 Destructive or reckless actions".to_string(),
"4.3.1.1 Synthetic-backend evaluation".to_string(),
"86 86".to_string(),
],
vec![
"4.3.2 Adherence to constitution".to_string(),
"4.3.2.1 Overview".to_string(),
"89 89".to_string(),
],
vec![
"4.3.3 Honesty and hallucinations".to_string(),
"4.3.3.1 Factual hallucinations".to_string(),
"93 94".to_string(),
],
vec![
"4.4 Capability evaluations".to_string(),
String::new(),
"101".to_string(),
],
];
assert!(
is_table_of_contents(&cells),
"dot-less TOC with section numbers + page numbers should be rejected"
);
}
#[test]
fn dot_leader_toc_accepts_short_inline_leaders() {
// Index-style cells where the full "label ... number" pattern is
// preserved in a single cell (IRS Publication 17 back-of-book index).
let cells = vec![
vec!["Child tax credit ... 235".to_string(), String::new()],
vec!["Church employee ... 252".to_string(), String::new()],
vec!["Citizens outside the U.S ... 6".to_string(), String::new()],
vec![
"Claim for refund ... 18, 36, 107".to_string(),
String::new(),
],
vec!["Clergy ... 7, 52".to_string(), String::new()],
];
assert!(is_dot_leader_toc(&cells));
}
#[test]
fn dot_leader_toc_allows_ellipsis_data_table() {
// Data tables using "..." as a row-omission marker must not be
// mistaken for dot-leader TOCs. Based on MCF5235RM QSPI RAM layout.
let cells = vec![
vec![
"0x00".to_string(),
"QTR0".to_string(),
"Transmit RAM".to_string(),
],
vec!["0x01".to_string(), "QTR1".to_string(), String::new()],
vec![
"...".to_string(),
"...".to_string(),
"16 bits wide".to_string(),
],
vec!["0x0F".to_string(), "QTR15".to_string(), String::new()],
vec![
"0x10".to_string(),
"QRR0".to_string(),
"Receive RAM".to_string(),
],
vec!["0x11".to_string(), "QRR1".to_string(), String::new()],
vec![
"...".to_string(),
"...".to_string(),
"16 bits wide".to_string(),
],
vec!["0x1F".to_string(), "QRR15".to_string(), String::new()],
];
assert!(
!is_dot_leader_toc(&cells),
"ellipsis markers in a data table should not match TOC detection"
);
}
#[test]
fn dot_leader_toc_rejects_year_row_data_table() {
// ERP-2025 economic data tables: year labels with trailing " ... ",
// a final " ... " column, and decimal-looking numeric cells. The
// detection previously classified these as dot-leader TOCs and
// routed them through flat-list formatting, destroying the grid.
let cells = vec![
vec![
"1973 ... ".to_string(),
"4. 0".to_string(),
"1. 8".to_string(),
"0. 4".to_string(),
"3. 2".to_string(),
" ... ".to_string(),
],
vec![
"1974 ... ".to_string(),
"1. 9".to_string(),
"1. 6".to_string(),
"5. 6".to_string(),
"2. 4".to_string(),
" ... ".to_string(),
],
vec![
"1975 ... ".to_string(),
"2. 6".to_string(),
"5. 1".to_string(),
"6. 1".to_string(),
"4. 1".to_string(),
" ... ".to_string(),
],
vec![
"1976 ... ".to_string(),
"4. 3".to_string(),
"5. 4".to_string(),
"6. 4".to_string(),
"4. 5".to_string(),
" ... ".to_string(),
],
];
assert!(
!is_dot_leader_toc(&cells),
"year-indexed data tables with decimal cells must not match TOC detection"
);
}
#[test]
fn dot_leader_toc_rejects_monthly_data_table() {
// ERP-2025 Table B-22: monthly labor-force rows with "Jan ... ",
// "Feb ... " labels and thousands-separated cells ("189,164").
// Previously matched TOC detection because "Jan ..." has alphabetic
// text and "189,164" passed the page-number shape check.
let cells = vec![
vec![
"2023: Jan ... ".to_string(),
"265,962".to_string(),
"165,871".to_string(),
"160,152".to_string(),
"62. 4".to_string(),
],
vec![
"Feb ... ".to_string(),
"266,112".to_string(),
"166,263".to_string(),
"160,301".to_string(),
"62. 5".to_string(),
],
vec![
"Mar ... ".to_string(),
"266,272".to_string(),
"166,690".to_string(),
"160,824".to_string(),
"62. 6".to_string(),
],
vec![
"Apr ... ".to_string(),
"266,443".to_string(),
"166,678".to_string(),
"160,962".to_string(),
"62. 6".to_string(),
],
];
assert!(
!is_dot_leader_toc(&cells),
"monthly labor-force rows with thousands-separated data must not match TOC detection"
);
}
#[test]
fn tabular_toc_requires_section_numbers_and_pages() {
// Dot-less tabular TOC matches is_tabular_toc but not dot-leader.
let cells = vec![
vec![
"4.3 Case studies".to_string(),
String::new(),
"86".to_string(),
],
vec![
"4.3.1 Destructive actions".to_string(),
String::new(),
"86".to_string(),
],
vec![
"4.3.2 Adherence".to_string(),
String::new(),
"89".to_string(),
],
vec!["4.3.3 Honesty".to_string(), String::new(), "93".to_string()],
];
assert!(is_tabular_toc(&cells));
assert!(!is_dot_leader_toc(&cells));
}
#[test]
fn starts_with_section_number_matches_dotted() {
assert!(starts_with_section_number("1.2"));
assert!(starts_with_section_number("4.3.1"));
assert!(starts_with_section_number("4.3.1.2"));
assert!(starts_with_section_number("4.3 Case studies"));
assert!(starts_with_section_number("2.2.5.1 Expert red teaming"));
}
#[test]
fn starts_with_section_number_rejects_non_sections() {
assert!(!starts_with_section_number("Chapter 1"));
assert!(!starts_with_section_number("1973"));
assert!(!starts_with_section_number("1.5M"));
assert!(!starts_with_section_number("10.0%"));
assert!(!starts_with_section_number(""));
assert!(!starts_with_section_number("Hello world"));
}
}
+90 -13
View File
@@ -110,15 +110,21 @@ pub fn detect_tables_from_lines(items: &[TextItem], lines: &[PdfLine], page: u32
return Vec::new();
}
// Reject page-spanning frames: if the grid covers >90% of a standard page
// dimension in both axes, it's a border frame, not a table.
// Reject page-spanning frames: a decorative outer border has just 4
// edges (top/bottom/left/right). Real full-page tables — common in
// governmental ledgers, financial reports, etc. — span the same A4 /
// Letter dimensions but have many internal row/column rules. Only
// reject when the line set looks like a bare frame, not a grid.
// Standard pages are ~595×842 (A4) or ~612×792 (Letter).
if table_width > 500.0 && table_height > 700.0 {
if table_width > 500.0 && table_height > 700.0 && horizontals.len() <= 4 && verticals.len() <= 4
{
log::debug!(
"detect_lines p{}: rejected — page-spanning frame ({:.0}×{:.0})",
"detect_lines p{}: rejected — page-spanning frame ({:.0}×{:.0}, {} h + {} v)",
page,
table_width,
table_height
table_height,
horizontals.len(),
verticals.len()
);
return Vec::new();
}
@@ -167,7 +173,7 @@ pub fn detect_tables_from_lines(items: &[TextItem], lines: &[PdfLine], page: u32
// Row edges need to be in descending order (top of page = higher Y first)
let mut row_edges_desc = row_edges;
row_edges_desc.sort_by(|a, b| b.partial_cmp(a).unwrap_or(std::cmp::Ordering::Equal));
row_edges_desc.sort_by(|a, b| b.total_cmp(a));
log::debug!(
"detect_lines p{}: {} row_edges, {} col_edges, table=({:.0},{:.0})-({:.0},{:.0}), spanning_h={}, spanning_v={}",
@@ -243,9 +249,11 @@ pub fn detect_tables_from_lines(items: &[TextItem], lines: &[PdfLine], page: u32
.map(|s| (s - mean_spacing).powi(2))
.sum::<f32>()
/ spacings.len() as f32;
let cv = variance.sqrt() / mean_spacing; // coefficient of variation
// CV < 0.05 means nearly identical spacing — chart grid
if cv < 0.05 {
let cv = variance.sqrt() / mean_spacing;
// CV < 0.02 means nearly identical spacing — likely chart grid.
// Spreadsheet-exported tables often have uniform rows (CV 0.03-0.05),
// so we use a tighter threshold to avoid false negatives.
if cv < 0.02 {
return Vec::new();
}
}
@@ -263,12 +271,12 @@ pub fn detect_tables_from_lines(items: &[TextItem], lines: &[PdfLine], page: u32
page, num_rows, num_cols, item_indices.len(), page_item_count, non_empty_rows, cols_with_content
);
vec![Table {
columns: col_edges,
rows: row_edges_desc[..num_rows].to_vec(),
vec![Table::new(
col_edges,
row_edges_desc[..num_rows].to_vec(),
cells,
item_indices,
}]
)]
}
#[cfg(test)]
@@ -408,6 +416,75 @@ mod tests {
assert!(tables.is_empty());
}
#[test]
fn test_page_spanning_bare_frame_rejected() {
// Just an outer A4-sized rectangle: 2 horizontals + 2 verticals.
// No internal structure → decorative border, not a table.
let lines = vec![
make_hline(20.0, 20.0, 575.0, 1), // top
make_hline(820.0, 20.0, 575.0, 1), // bottom
make_vline(20.0, 20.0, 820.0, 1), // left
make_vline(575.0, 20.0, 820.0, 1), // right
];
let items = vec![
make_item("title", 100.0, 100.0, 1),
make_item("body", 100.0, 200.0, 1),
];
let tables = detect_tables_from_lines(&items, &lines, 1);
assert!(
tables.is_empty(),
"Page-sized 4-edge frame should be rejected as decoration"
);
}
#[test]
fn test_page_spanning_grid_with_internal_lines_accepted() {
// Full-page table (governmental-ledger pattern): A4-sized grid
// that previously hit the "page-spanning frame" early reject
// before downstream validation could even look at it.
// Verticals span the full table height so we isolate the
// frame-vs-grid decision under test.
let mut lines = Vec::new();
// 13 horizontal rules: header + 12 row separators
let h_ys = [
22.5, 37.9, 95.5, 144.5, 184.9, 233.9, 291.7, 340.7, 415.8, 499.6, 574.7, 623.7, 698.8,
];
for &y in &h_ys {
lines.push(make_hline(y, 22.6, 566.6, 1));
}
// 7 column dividers spanning full table height.
let v_xs = [22.6, 66.3, 116.3, 186.6, 263.1, 493.5, 566.5];
for &x in &v_xs {
lines.push(make_vline(x, 22.5, 698.8, 1));
}
// Populate every cell so the capture-ratio + density checks pass.
let mut items = Vec::new();
for r in 0..(h_ys.len() - 1) {
let row_y = (h_ys[r] + h_ys[r + 1]) / 2.0;
for c in 0..(v_xs.len() - 1) {
let col_x = (v_xs[c] + v_xs[c + 1]) / 2.0;
items.push(make_item("x", col_x, row_y, 1));
}
}
let tables = detect_tables_from_lines(&items, &lines, 1);
assert_eq!(
tables.len(),
1,
"Full-page table with internal grid should be accepted"
);
let t = &tables[0];
assert!(
t.cells.len() >= 6,
"expected ≥6 rows, got {}",
t.cells.len()
);
assert!(
t.cells[0].len() >= 3,
"expected ≥3 columns, got {}",
t.cells[0].len()
);
}
#[test]
fn test_single_column_rejected() {
// Only 2 col edges (1 column) — not a table even with verticals
File diff suppressed because it is too large Load Diff
+846 -62
View File
@@ -4,15 +4,385 @@
//! elements linked to MCIDs, this module builds `Table` structs directly from
//! the semantic hierarchy — no geometry heuristics needed.
use std::collections::HashMap;
use std::collections::{HashMap, HashSet};
use log::debug;
use crate::structure_tree::StructTable;
use crate::structure_tree::{StructTable, StructTableRow};
use crate::types::TextItem;
use super::Table;
#[derive(Debug, Clone)]
struct MatchedCell {
text: String,
item_indices: Vec<usize>,
x: Option<f32>,
y: Option<f32>,
}
fn legacy_column_positions(
page_rows: &[&StructTableRow],
mcid_to_items: &HashMap<i64, Vec<usize>>,
items: &[TextItem],
page: u32,
num_cols: usize,
) -> Vec<f32> {
let mut col_positions: Vec<f32> = vec![0.0; num_cols];
for (col, col_pos) in col_positions.iter_mut().enumerate() {
for row in page_rows {
if col < row.cells.len() {
if let Some(x) = row.cells[col]
.mcids
.iter()
.filter(|(_, p)| *p == page)
.filter_map(|(mcid, _)| mcid_to_items.get(mcid))
.flatten()
.map(|&idx| items[idx].x)
.reduce(f32::min)
{
*col_pos = x;
break;
}
}
}
}
col_positions
}
fn infer_column_positions(
raw_rows: &[Vec<MatchedCell>],
fallback_positions: &[f32],
num_cols: usize,
) -> Vec<f32> {
const SAME_COLUMN_TOLERANCE: f32 = 18.0;
let mut anchors = raw_rows
.iter()
.max_by_key(|row| row.iter().filter(|cell| cell.x.is_some()).count())
.map(|row| row.iter().filter_map(|cell| cell.x).collect::<Vec<_>>())
.unwrap_or_default();
if anchors.len() > num_cols {
anchors.truncate(num_cols);
}
let mut additional_positions: Vec<f32> = raw_rows
.iter()
.flat_map(|row| row.iter().filter_map(|cell| cell.x))
.collect();
additional_positions.sort_by(|a, b| a.total_cmp(b));
for x in additional_positions {
if anchors.len() >= num_cols {
break;
}
if anchors
.iter()
.all(|existing| (x - *existing).abs() > SAME_COLUMN_TOLERANCE)
{
anchors.push(x);
anchors.sort_by(|a, b| a.total_cmp(b));
}
}
if anchors.len() < num_cols {
for &x in fallback_positions {
if anchors.len() >= num_cols {
break;
}
if anchors
.iter()
.all(|existing| (x - *existing).abs() > SAME_COLUMN_TOLERANCE)
{
anchors.push(x);
anchors.sort_by(|a, b| a.total_cmp(b));
}
}
}
if anchors.is_empty() {
return fallback_positions.to_vec();
}
while anchors.len() < num_cols {
anchors.push(*anchors.last().unwrap());
}
anchors
}
fn align_positions_to_columns(cell_xs: &[f32], columns: &[f32]) -> Vec<usize> {
if cell_xs.is_empty() || columns.is_empty() {
return Vec::new();
}
if cell_xs.len() >= columns.len() {
return (0..cell_xs.len().min(columns.len())).collect();
}
let mut dp = vec![vec![f32::INFINITY; columns.len() + 1]; cell_xs.len() + 1];
let mut take = vec![vec![false; columns.len() + 1]; cell_xs.len() + 1];
for value in &mut dp[0] {
*value = 0.0;
}
for i in 1..=cell_xs.len() {
for j in 1..=columns.len() {
let skip_cost = dp[i][j - 1];
let take_cost = dp[i - 1][j - 1] + (cell_xs[i - 1] - columns[j - 1]).abs();
if take_cost <= skip_cost {
dp[i][j] = take_cost;
take[i][j] = true;
} else {
dp[i][j] = skip_cost;
}
}
}
let mut assignments_rev = Vec::with_capacity(cell_xs.len());
let mut i = cell_xs.len();
let mut j = columns.len();
while i > 0 && j > 0 {
if take[i][j] {
assignments_rev.push(j - 1);
i -= 1;
j -= 1;
} else {
j -= 1;
}
}
assignments_rev.reverse();
assignments_rev
}
fn align_struct_rows(
raw_rows: &[Vec<MatchedCell>],
col_positions: &[f32],
) -> (Vec<Vec<String>>, Vec<f32>, Vec<usize>) {
let mut cells: Vec<Vec<String>> = Vec::with_capacity(raw_rows.len());
let mut row_positions: Vec<f32> = Vec::with_capacity(raw_rows.len());
let mut all_item_indices: Vec<usize> = Vec::new();
for row in raw_rows {
let present_cells: Vec<&MatchedCell> = row
.iter()
.filter(|cell| {
!cell.item_indices.is_empty() || !cell.text.is_empty() || cell.x.is_some()
})
.collect();
let cell_xs: Vec<f32> = present_cells.iter().filter_map(|cell| cell.x).collect();
let assignments = if cell_xs.len() == present_cells.len() {
align_positions_to_columns(&cell_xs, col_positions)
} else {
(0..present_cells.len().min(col_positions.len())).collect()
};
let mut row_cells = vec![String::new(); col_positions.len()];
for (cell, &col_idx) in present_cells.iter().zip(assignments.iter()) {
if !cell.text.is_empty() {
if !row_cells[col_idx].is_empty() {
row_cells[col_idx].push(' ');
}
row_cells[col_idx].push_str(&cell.text);
}
all_item_indices.extend(cell.item_indices.iter().copied());
}
let row_y = row
.iter()
.filter_map(|cell| cell.y)
.reduce(f32::max)
.unwrap_or(0.0);
cells.push(row_cells);
row_positions.push(row_y);
}
(cells, row_positions, all_item_indices)
}
fn left_align_struct_rows(
raw_rows: &[Vec<MatchedCell>],
num_cols: usize,
) -> (Vec<Vec<String>>, Vec<f32>, Vec<usize>) {
let mut cells: Vec<Vec<String>> = Vec::with_capacity(raw_rows.len());
let mut row_positions: Vec<f32> = Vec::with_capacity(raw_rows.len());
let mut all_item_indices: Vec<usize> = Vec::new();
for row in raw_rows {
let mut row_cells: Vec<String> = row.iter().map(|cell| cell.text.clone()).collect();
row_cells.truncate(num_cols);
while row_cells.len() < num_cols {
row_cells.push(String::new());
}
cells.push(row_cells);
all_item_indices.extend(
row.iter()
.flat_map(|cell| cell.item_indices.iter().copied()),
);
row_positions.push(
row.iter()
.filter_map(|cell| cell.y)
.reduce(f32::max)
.unwrap_or(0.0),
);
}
(cells, row_positions, all_item_indices)
}
fn recover_unclaimed_header_row(table: &mut Table, items: &[TextItem], has_ragged_rows: bool) {
if !has_ragged_rows || table.rows.is_empty() || table.columns.len() < 3 {
return;
}
const MAX_HEADER_DISTANCE: f32 = 90.0;
const MAX_GAP_TO_TABLE: f32 = 35.0;
const MAX_INTER_HEADER_GAP: f32 = 25.0;
const MAX_HEADER_ROWS: usize = 3;
const Y_TOLERANCE: f32 = 5.0;
let top_row_y = table.rows[0];
let x_min = table.columns.first().copied().unwrap_or(0.0) - 25.0;
let x_max = table.columns.last().copied().unwrap_or(0.0) + 120.0;
let claimed: HashSet<usize> = table.item_indices.iter().copied().collect();
let mut candidate_rows: Vec<(f32, Vec<(usize, &TextItem)>)> = Vec::new();
for (idx, item) in items.iter().enumerate() {
if claimed.contains(&idx)
|| item.text.trim().is_empty()
|| item.y <= top_row_y
|| item.y - top_row_y > MAX_HEADER_DISTANCE
|| item.x < x_min
|| item.x > x_max
{
continue;
}
if let Some((_, row_items)) = candidate_rows
.iter_mut()
.find(|(row_y, _)| (item.y - *row_y).abs() < Y_TOLERANCE)
{
row_items.push((idx, item));
} else {
candidate_rows.push((item.y, vec![(idx, item)]));
}
}
if candidate_rows.is_empty() {
return;
}
for (_, row_items) in &mut candidate_rows {
row_items.sort_by(|a, b| a.1.x.total_cmp(&b.1.x));
}
candidate_rows.sort_by(|a, b| a.0.total_cmp(&b.0));
if candidate_rows[0].0 - top_row_y > MAX_GAP_TO_TABLE {
return;
}
let mut candidate_iter = candidate_rows.into_iter();
let Some(first_row) = candidate_iter.next() else {
return;
};
let mut selected_rows: Vec<(f32, Vec<(usize, &TextItem)>)> = vec![first_row];
let mut prev_y = selected_rows[0].0;
for (row_y, row_items) in candidate_iter {
if selected_rows.len() >= MAX_HEADER_ROWS {
break;
}
if row_y - prev_y > MAX_INTER_HEADER_GAP {
break;
}
prev_y = row_y;
selected_rows.push((row_y, row_items));
}
if selected_rows.is_empty() {
return;
}
let mut assigned_rows: Vec<(f32, Vec<String>, Vec<usize>)> = Vec::new();
let mut closest_row_populated = 0usize;
let mut combined_cols: HashSet<usize> = HashSet::new();
for (row_idx, (row_y, row_items)) in selected_rows.iter().enumerate() {
if row_items.len() > table.columns.len() {
return;
}
let row_xs: Vec<f32> = row_items.iter().map(|(_, item)| item.x).collect();
let assignments = align_positions_to_columns(&row_xs, &table.columns);
if assignments.len() != row_items.len() {
return;
}
let mut row_cells = vec![String::new(); table.columns.len()];
let mut row_indices = Vec::with_capacity(row_items.len());
let mut populated_cols: HashSet<usize> = HashSet::new();
for ((idx, item), &col_idx) in row_items.iter().zip(assignments.iter()) {
let text = item.text.trim();
if text.is_empty() {
continue;
}
if !row_cells[col_idx].is_empty() {
row_cells[col_idx].push(' ');
}
row_cells[col_idx].push_str(text);
row_indices.push(*idx);
populated_cols.insert(col_idx);
}
if row_idx == 0 {
closest_row_populated = populated_cols.len();
}
combined_cols.extend(populated_cols.iter().copied());
assigned_rows.push((*row_y, row_cells, row_indices));
}
let required_cols = if table.columns.len() <= 4 {
table.columns.len()
} else {
table.columns.len() - 1
};
if closest_row_populated < 2 || combined_cols.len() < required_cols {
return;
}
let mut header_cells = vec![String::new(); table.columns.len()];
let mut header_indices = Vec::new();
for (_, row_cells, row_indices) in assigned_rows.iter().rev() {
for (col_idx, cell_text) in row_cells.iter().enumerate() {
if cell_text.is_empty() {
continue;
}
if !header_cells[col_idx].is_empty() {
header_cells[col_idx].push(' ');
}
header_cells[col_idx].push_str(cell_text);
}
header_indices.extend(row_indices.iter().copied());
}
table.rows.insert(
0,
assigned_rows
.iter()
.map(|(row_y, _, _)| *row_y)
.reduce(f32::max)
.unwrap_or(top_row_y),
);
table.cells.insert(0, header_cells);
table.item_indices.extend(header_indices);
table.item_indices.sort_unstable();
table.item_indices.dedup();
}
/// Build tables from structure-tree table descriptors by matching MCIDs to TextItems.
///
/// Returns tables for the given page. Tables where fewer than 50% of cells
@@ -67,18 +437,14 @@ pub fn detect_tables_from_struct_tree(
continue;
}
// Build cell text and collect item indices
let mut cells: Vec<Vec<String>> = Vec::new();
let mut all_item_indices: Vec<usize> = Vec::new();
// Build cell text and geometry for alignment and header recovery.
let mut raw_rows: Vec<Vec<MatchedCell>> = Vec::new();
let mut total_cells = 0u32;
let mut matched_cells = 0u32;
for row in &page_rows {
let mut row_cells = Vec::with_capacity(num_cols);
for (col_idx, cell) in row.cells.iter().enumerate() {
if col_idx >= num_cols {
break;
}
let mut row_cells = Vec::with_capacity(row.cells.len());
for cell in &row.cells {
total_cells += 1;
// Collect all items for this cell's MCIDs
@@ -115,18 +481,18 @@ pub fn detect_tables_from_struct_tree(
.collect::<Vec<_>>()
.join(" ");
for (idx, _) in &cell_items {
all_item_indices.push(*idx);
}
let item_indices = cell_items.iter().map(|(idx, _)| *idx).collect::<Vec<_>>();
let x = cell_items.iter().map(|(_, item)| item.x).reduce(f32::min);
let y = cell_items.iter().map(|(_, item)| item.y).reduce(f32::max);
row_cells.push(text);
row_cells.push(MatchedCell {
text,
item_indices,
x,
y,
});
}
// Pad to num_cols
while row_cells.len() < num_cols {
row_cells.push(String::new());
}
cells.push(row_cells);
raw_rows.push(row_cells);
}
// Reject if too few cells matched (stale structure tree)
@@ -148,51 +514,54 @@ pub fn detect_tables_from_struct_tree(
continue;
}
// Derive row/column positions from item geometry
let mut row_positions: Vec<f32> = Vec::new();
for row in &page_rows {
let y = row
.cells
.iter()
.flat_map(|c| c.mcids.iter())
.filter(|(_, p)| *p == page)
.filter_map(|(mcid, _)| mcid_to_items.get(mcid))
.flatten()
.map(|&idx| items[idx].y)
.reduce(f32::max)
.unwrap_or(0.0);
row_positions.push(y);
}
let has_ragged_rows = raw_rows
.iter()
.any(|row| row.iter().filter(|cell| cell.x.is_some()).count() < num_cols);
let first_row_has_tagged_header = page_rows.first().is_some_and(|row| {
let header_cells = row.cells.iter().filter(|cell| cell.is_header).count();
header_cells * 2 >= row.cells.len()
});
let fallback_col_positions =
legacy_column_positions(&page_rows, &mcid_to_items, items, page, num_cols);
let (legacy_cells, legacy_row_positions, mut legacy_item_indices) =
left_align_struct_rows(&raw_rows, num_cols);
legacy_item_indices.sort_unstable();
legacy_item_indices.dedup();
let legacy_table = Table::new(
fallback_col_positions.clone(),
legacy_row_positions,
legacy_cells,
legacy_item_indices,
);
// Column positions: use X positions of first non-empty cell in each column
let mut col_positions: Vec<f32> = vec![0.0; num_cols];
for (col, col_pos) in col_positions.iter_mut().enumerate() {
for row in &page_rows {
if col < row.cells.len() {
if let Some(x) = row.cells[col]
.mcids
.iter()
.filter(|(_, p)| *p == page)
.filter_map(|(mcid, _)| mcid_to_items.get(mcid))
.flatten()
.map(|&idx| items[idx].x)
.reduce(f32::min)
{
*col_pos = x;
break;
}
}
}
}
let col_positions = infer_column_positions(&raw_rows, &fallback_col_positions, num_cols);
let (aligned_cells, aligned_row_positions, mut aligned_item_indices) =
align_struct_rows(&raw_rows, &col_positions);
aligned_item_indices.sort_unstable();
aligned_item_indices.dedup();
all_item_indices.sort_unstable();
all_item_indices.dedup();
let mut aligned_table = Table::new(
col_positions,
aligned_row_positions,
aligned_cells,
aligned_item_indices,
);
let item_count_before_header = aligned_table.item_indices.len();
let row_count_before_header = aligned_table.cells.len();
recover_unclaimed_header_row(
&mut aligned_table,
items,
has_ragged_rows && !first_row_has_tagged_header,
);
tables.push(Table {
columns: col_positions,
rows: row_positions,
cells,
item_indices: all_item_indices,
let recovered_header = aligned_table.item_indices.len() > item_count_before_header
|| aligned_table.cells.len() > row_count_before_header;
let prefer_aligned = recovered_header;
tables.push(if prefer_aligned {
aligned_table
} else {
legacy_table
});
}
@@ -374,4 +743,419 @@ mod tests {
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 2);
assert_eq!(tables.len(), 1);
}
#[test]
fn realigns_ragged_rows_and_recovers_untagged_header() {
let items = vec![
make_item("Category", 50.0, 120.0, 1, None),
make_item("Potentially", 150.0, 120.0, 1, None),
make_item("Summary", 250.0, 120.0, 1, None),
make_item("Most commonly", 350.0, 120.0, 1, None),
make_item("concerning aspect", 150.0, 110.0, 1, None),
make_item("suggested", 350.0, 110.0, 1, None),
make_item("of circumstances", 150.0, 100.0, 1, None),
make_item("intervention", 350.0, 100.0, 1, None),
make_item("Existence of red-teaming", 150.0, 80.0, 1, Some(10)),
make_item("Important for safety", 250.0, 80.0, 1, Some(11)),
make_item("Ensure welfare interviews", 350.0, 80.0, 1, Some(12)),
make_item("Identity & self-knowledge", 50.0, 60.0, 1, Some(20)),
make_item("Lack of knowledge", 150.0, 60.0, 1, Some(21)),
make_item("Overall negative", 250.0, 60.0, 1, Some(22)),
make_item("Describe training process", 350.0, 60.0, 1, Some(23)),
make_item("Uncertainty around other copies", 150.0, 40.0, 1, Some(30)),
make_item("High uncertainty", 250.0, 40.0, 1, Some(31)),
make_item("No intervention suggested", 350.0, 40.0, 1, Some(32)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(10, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(11, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(12, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(22, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(23, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(30, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(31, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(32, 1)],
},
],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells.len(), 4);
assert_eq!(
table.cells[0],
vec![
"Category",
"Potentially concerning aspect of circumstances",
"Summary",
"Most commonly suggested intervention",
]
);
assert_eq!(table.cells[1][0], "");
assert_eq!(table.cells[1][1], "Existence of red-teaming");
assert_eq!(table.cells[2][0], "Identity & self-knowledge");
assert_eq!(table.cells[3][0], "");
assert_eq!(table.columns.len(), 4);
assert!(table.columns.windows(2).all(|w| w[0] < w[1]));
assert_eq!(table.item_indices.len(), items.len());
}
#[test]
fn does_not_absorb_caption_above_ragged_struct_table() {
let items = vec![
make_item("Table 5-7: Summary of responses", 50.0, 120.0, 1, None),
make_item("Aspect one", 150.0, 80.0, 1, Some(10)),
make_item("Summary one", 250.0, 80.0, 1, Some(11)),
make_item("Category", 50.0, 60.0, 1, Some(20)),
make_item("Aspect two", 150.0, 60.0, 1, Some(21)),
make_item("Summary two", 250.0, 60.0, 1, Some(22)),
make_item("Aspect three", 150.0, 40.0, 1, Some(30)),
make_item("Summary three", 250.0, 40.0, 1, Some(31)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(10, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(11, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(22, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(30, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(31, 1)],
},
],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells.len(), 3);
assert!(
table
.cells
.iter()
.flatten()
.all(|cell| !cell.contains("Table 5-7")),
"caption must stay outside the table"
);
assert!(!table.item_indices.contains(&0));
}
#[test]
fn keeps_existing_tagged_header_without_absorbing_intro_or_caption() {
let items = vec![
make_item(
"Eighteen people left other comments regarding the Project.",
50.0,
130.0,
1,
None,
),
make_item("Table 5-1:", 220.0, 130.0, 1, None),
make_item("Other Comments", 350.0, 130.0, 1, None),
make_item("Theme", 50.0, 110.0, 1, Some(10)),
make_item("Specific Concern/Inquiry", 200.0, 110.0, 1, Some(11)),
make_item("Response", 420.0, 110.0, 1, Some(12)),
make_item("Traffic", 50.0, 90.0, 1, Some(20)),
make_item("Road conditions", 200.0, 90.0, 1, Some(21)),
make_item("Maintenance response", 420.0, 90.0, 1, Some(22)),
make_item("Noise", 50.0, 70.0, 1, Some(30)),
make_item("Dust concerns", 200.0, 70.0, 1, Some(31)),
make_item("Mitigation response", 420.0, 70.0, 1, Some(32)),
make_item("Resource Use", 50.0, 50.0, 1, Some(40)),
make_item("Snowmobile trails", 200.0, 50.0, 1, Some(41)),
make_item("Access response", 420.0, 50.0, 1, Some(42)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![
StructTableCell {
is_header: true,
mcids: vec![(10, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(11, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(12, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(22, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(30, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(31, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(32, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(40, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(41, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(42, 1)],
},
],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells.len(), 4);
assert_eq!(
table.cells[0],
vec!["Theme", "Specific Concern/Inquiry", "Response"]
);
assert!(!table.item_indices.contains(&0));
assert!(!table.item_indices.contains(&1));
assert!(!table.item_indices.contains(&2));
}
#[test]
fn does_not_recover_header_for_narrow_two_column_table() {
let items = vec![
make_item("Alpha", 50.0, 120.0, 1, None),
make_item("Beta", 200.0, 120.0, 1, None),
make_item("First value", 200.0, 80.0, 1, Some(10)),
make_item("Only labeled row", 50.0, 60.0, 1, Some(20)),
make_item("Second value", 200.0, 60.0, 1, Some(21)),
make_item("Third value", 200.0, 40.0, 1, Some(30)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![StructTableCell {
is_header: false,
mcids: vec![(10, 1)],
}],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
],
},
StructTableRow {
cells: vec![StructTableCell {
is_header: false,
mcids: vec![(30, 1)],
}],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells.len(), 3);
assert_eq!(table.cells[0], vec!["First value", ""]);
assert_eq!(table.cells[1], vec!["Only labeled row", "Second value"]);
assert_eq!(table.cells[2], vec!["Third value", ""]);
assert!(!table.item_indices.contains(&0));
assert!(!table.item_indices.contains(&1));
}
#[test]
fn ragged_rows_without_recovered_header_keep_legacy_alignment() {
let items = vec![
make_item("Date", 150.0, 120.0, 1, Some(10)),
make_item("Title", 250.0, 120.0, 1, Some(11)),
make_item("PE", 350.0, 120.0, 1, Some(12)),
make_item("Bidder", 450.0, 120.0, 1, Some(13)),
make_item("Amount", 550.0, 120.0, 1, Some(14)),
make_item("1", 50.0, 100.0, 1, Some(20)),
make_item("8/1", 150.0, 100.0, 1, Some(21)),
make_item("Procurement", 250.0, 100.0, 1, Some(22)),
make_item("PUC", 350.0, 100.0, 1, Some(23)),
make_item("Vendor", 450.0, 100.0, 1, Some(24)),
make_item("SR1", 550.0, 100.0, 1, Some(25)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![
StructTableCell {
is_header: true,
mcids: vec![(10, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(11, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(12, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(13, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(14, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(22, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(23, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(24, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(25, 1)],
},
],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells[0][0], "Date");
assert_eq!(table.cells[0][4], "Amount");
assert_eq!(table.cells[0][5], "");
}
}
+226 -1
View File
@@ -1,12 +1,22 @@
//! Table-to-markdown formatting and cell cleanup.
use super::Table;
use super::{Table, TableKind};
pub fn table_to_markdown(table: &Table) -> String {
if table.cells.is_empty() || table.cells[0].is_empty() {
return String::new();
}
// TOCs render poorly as markdown tables — emit a flat per-row text list
// instead so the page numbers stay aligned with their section titles
// rather than drifting to a separate column. Format from raw cells
// because continuation-row merging in clean_table_cells collapses
// separate TOC entries (e.g. "6.2 Contamination" + "6.2.1 SWE-bench")
// into one line where sub-entries leave column 0 empty.
if table.kind == TableKind::Toc {
return format_toc_as_list(&table.cells, &[]);
}
// Clean up the table: merge continuation rows, extract footnotes, remove empty rows
let (cleaned_cells, footnotes) = clean_table_cells(&table.cells);
@@ -49,6 +59,107 @@ pub fn table_to_markdown(table: &Table) -> String {
output
}
/// Render a table-of-contents as a flat per-row text block.
///
/// Each row becomes one line: non-empty cells joined with spaces, and the
/// last cell (typically a page number) is separated by a tab so the page
/// numbers stay aligned with their titles instead of being pulled into a
/// separate column by the column-aware reader.
fn format_toc_as_list(cells: &[Vec<String>], footnotes: &[String]) -> String {
let mut output = String::new();
for row in cells {
let trimmed: Vec<&str> = row.iter().map(|c| c.trim()).collect();
let last_idx = trimmed.iter().rposition(|c| !c.is_empty());
let Some(last_idx) = last_idx else {
continue;
};
let last_cell = trimmed[last_idx];
let last_is_page = is_page_number_cell(last_cell);
let (title_cells, trailing) = if last_is_page && last_idx > 0 {
(&trimmed[..last_idx], Some(last_cell))
} else {
(&trimmed[..=last_idx], None)
};
// Skip dots-only cells when joining the title — in a detected TOC
// layout, a "...." cell is a leader separator, not part of the
// entry name.
let title = title_cells
.iter()
.filter(|c| !c.is_empty() && !is_dots_only(c))
.copied()
.collect::<Vec<_>>()
.join(" ");
if title.is_empty() && trailing.is_none() {
continue;
}
if !title.is_empty() {
output.push_str(&title);
}
if let Some(page) = trailing {
if !title.is_empty() {
output.push('\t');
}
output.push_str(page);
}
output.push('\n');
}
if !footnotes.is_empty() {
output.push('\n');
for footnote in footnotes {
output.push_str(footnote);
output.push('\n');
}
}
output
}
/// True when the cell looks like a page number. Accepts:
/// - plain digit tokens: "42", "86 86"
/// - dashed section-page IDs: "5-21", "A-1", "B--3", "TC-2" (common in
/// technical manuals)
fn is_page_number_cell(cell: &str) -> bool {
let tokens: Vec<&str> = cell.split_whitespace().collect();
if tokens.is_empty() {
return false;
}
tokens.iter().all(|t| {
if t.is_empty() || t.len() > 8 {
return false;
}
let all_digits = t.chars().all(|c| c.is_ascii_digit());
if all_digits {
return t.len() <= 4;
}
// Section-page form: uppercase letters, digits, dashes; at least
// one digit present.
t.chars()
.all(|c| c.is_ascii_digit() || c.is_ascii_uppercase() || c == '-')
&& t.chars().any(|c| c.is_ascii_digit())
})
}
/// True when the cell is purely leader dots (any length ≥ 3) with optional
/// whitespace.
fn is_dots_only(cell: &str) -> bool {
let t = cell.trim();
let dots = t.chars().filter(|&c| c == '.').count();
dots >= 3 && t.chars().all(|c| c == '.' || c.is_whitespace())
}
fn starts_with_uppercase_word(cell: &str) -> bool {
cell.chars()
.find(|c| c.is_alphanumeric())
.is_some_and(|c| c.is_uppercase())
}
/// Clean up table cells: merge continuation rows, extract footnotes, remove empty rows
fn clean_table_cells(cells: &[Vec<String>]) -> (Vec<Vec<String>>, Vec<String>) {
let mut cleaned: Vec<Vec<String>> = Vec::new();
@@ -107,11 +218,20 @@ fn clean_table_cells(cells: &[Vec<String>]) -> (Vec<Vec<String>>, Vec<String>) {
let looks_like_data_row = non_first_cells.len() >= 2
&& avg_cell_len <= 10.0
&& numeric_cells > non_first_cells.len() / 2;
let uppercase_leading_cells = non_first_cells
.iter()
.filter(|cell| starts_with_uppercase_word(cell))
.count();
let looks_like_spanning_first_column_row = first_cell.is_empty()
&& row.len() >= 4
&& non_first_cells.len() == row.len().saturating_sub(1)
&& uppercase_leading_cells >= non_first_cells.len().saturating_sub(1);
// Classic continuation: first cell empty, content in other cells
let is_classic_continuation = first_cell.is_empty()
&& !non_first_cells.is_empty()
&& !is_short_subheader
&& !looks_like_data_row
&& !looks_like_spanning_first_column_row
&& cleaned.len() > 1;
// Wrapped-cell continuation: row has fewer filled cells than the header
@@ -141,6 +261,7 @@ fn clean_table_cells(cells: &[Vec<String>]) -> (Vec<Vec<String>>, Vec<String>) {
&& filled_cells <= max_filled_for_merge
&& prev_filled > filled_cells
&& !looks_like_data_row
&& !looks_like_spanning_first_column_row
&& !is_short_subheader;
let is_continuation = is_classic_continuation || is_wrapped_continuation;
@@ -310,6 +431,65 @@ mod tests {
assert_eq!(cleaned.len(), 3);
}
#[test]
fn test_clean_table_cells_spanning_first_column_row_not_merged() {
let cells = vec![
vec![
"Category".into(),
"Potentially concerning aspect".into(),
"Summary".into(),
"Intervention".into(),
],
vec![
"Identity & self-knowledge".into(),
"Lack of knowledge".into(),
"Overall negative".into(),
"Describe training".into(),
],
vec![
"".into(),
"Uncertainty around other copies".into(),
"High uncertainty".into(),
"No intervention suggested".into(),
],
];
let (cleaned, _) = clean_table_cells(&cells);
assert_eq!(cleaned.len(), 3);
assert_eq!(cleaned[2][0], "");
assert_eq!(cleaned[2][1], "Uncertainty around other copies");
}
#[test]
fn test_clean_table_cells_full_width_continuation_row_still_merges_when_lowercase() {
let cells = vec![
vec![
"Classification".into(),
"Before tax".into(),
"After tax".into(),
"Standard equipment".into(),
"Options".into(),
],
vec![
"Exclusive Special".into(),
"83,500,000".into(),
"79,275,000".into(),
"Standard equipment".into(),
"Option A".into(),
],
vec![
"".into(),
"with 3.5% individual consumption tax applied".into(),
"with 3.5% individual consumption tax applied".into(),
"lighting(crash pad)".into(),
"sound system".into(),
],
];
let (cleaned, _) = clean_table_cells(&cells);
assert_eq!(cleaned.len(), 2);
assert!(cleaned[1][1].contains("83,500,000"));
assert!(cleaned[1][1].contains("with 3.5%"));
}
#[test]
fn test_clean_table_cells_header_row_not_merged() {
// Continuation requires cleaned.len() > 1 (don't merge into header)
@@ -358,6 +538,7 @@ mod tests {
vec!["Bob".into(), "25".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
assert!(md.contains("|Name|"));
@@ -373,6 +554,7 @@ mod tests {
rows: vec![500.0],
cells: vec![vec!["Only".into(), "Row".into()]],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
assert!(md.contains("|Only|"));
@@ -386,6 +568,7 @@ mod tests {
rows: vec![],
cells: vec![],
item_indices: vec![],
kind: TableKind::Data,
};
assert_eq!(table_to_markdown(&table), "");
}
@@ -401,6 +584,7 @@ mod tests {
vec!["(1)".into(), "Footnote text".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
assert!(md.contains("(1) Footnote text"));
@@ -416,6 +600,7 @@ mod tests {
vec!["太郎".into(), "25".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
assert!(md.contains("名前"));
@@ -429,7 +614,47 @@ mod tests {
rows: vec![500.0],
cells: vec![vec![]],
item_indices: vec![],
kind: TableKind::Data,
};
assert_eq!(table_to_markdown(&table), "");
}
#[test]
fn test_table_to_markdown_toc_renders_as_flat_list() {
// A TOC-shaped table with section numbers in col 0 and page numbers
// in the last column should render as a flat list, not a markdown
// table, so the page numbers stay on the same line as their titles.
let table = Table::new(
vec![50.0, 80.0, 300.0],
vec![500.0; 5],
vec![
vec![
"4.3".into(),
"Case studies and targeted evaluations".into(),
"86".into(),
],
vec![
"4.3.1".into(),
"Destructive or reckless actions".into(),
"86".into(),
],
vec![
"4.3.2".into(),
"Adherence to its constitution".into(),
"89".into(),
],
vec!["4.4".into(), "Capability evaluations".into(), "101".into()],
vec!["4.5".into(), "White-box analyses".into(), "113".into()],
],
vec![],
);
assert_eq!(table.kind, TableKind::Toc);
let md = table_to_markdown(&table);
assert!(
!md.contains("|---|"),
"TOC should not render as a markdown table: {md}"
);
assert!(md.contains("4.3 Case studies and targeted evaluations\t86"));
assert!(md.contains("4.5 White-box analyses\t113"));
}
}
+142 -16
View File
@@ -9,7 +9,7 @@ pub(crate) fn find_column_boundaries(
mode: TableDetectionMode,
) -> Vec<f32> {
let mut x_positions: Vec<f32> = items.iter().map(|(_, i)| i.x).collect();
x_positions.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
x_positions.sort_by(|a, b| a.total_cmp(b));
if x_positions.is_empty() {
return vec![];
@@ -43,7 +43,7 @@ pub(crate) fn find_column_boundaries(
.collect();
if consec_gaps.len() > 2 {
consec_gaps.sort_by(|a, b| a.partial_cmp(b).unwrap());
consec_gaps.sort_by(|a, b| a.total_cmp(b));
// Find the biggest jump in the sorted gap sequence — natural break
// between within-column jitter and between-column spacing.
// Require at least 3 values on each side to avoid outlier-dominated
@@ -82,33 +82,42 @@ pub(crate) fn find_column_boundaries(
}
}
let mut columns = Vec::new();
let mut cluster_items: Vec<f32> = vec![x_positions[0]];
// Track cluster membership: for each cluster, store the list of x positions
let mut cluster_xs: Vec<Vec<f32>> = vec![vec![x_positions[0]]];
for &x in &x_positions[1..] {
let last_cluster = cluster_xs.last().unwrap();
// For dense columns (gap-histogram triggered), use edge-based clustering:
// compare with the last item to avoid center-drift that merges adjacent
// narrow columns. For normal tables, use center-based (original behavior).
let reference = if use_edge_clustering {
*cluster_items.last().unwrap()
*last_cluster.last().unwrap()
} else {
cluster_items.iter().sum::<f32>() / cluster_items.len() as f32
last_cluster.iter().sum::<f32>() / last_cluster.len() as f32
};
if x - reference > cluster_threshold {
let cluster_center = cluster_items.iter().sum::<f32>() / cluster_items.len() as f32;
columns.push(cluster_center);
cluster_items = vec![x];
cluster_xs.push(vec![x]);
} else {
cluster_items.push(x);
cluster_xs.last_mut().unwrap().push(x);
}
}
// Don't forget last cluster
if !cluster_items.is_empty() {
columns.push(cluster_items.iter().sum::<f32>() / cluster_items.len() as f32);
// Numeric column merge pass: when a sparse cluster (few items, typically
// header text) is adjacent to a dense numeric cluster and within 1.5×
// threshold, merge them. This fixes tables where multi-line wrapped
// headers have slightly different X positions than the data columns,
// causing the header and data to split into separate clusters.
let columns_before_merge = cluster_xs.len();
if columns_before_merge >= 3 {
cluster_xs = merge_numeric_adjacent_clusters(cluster_xs, items, cluster_threshold);
}
let columns: Vec<f32> = cluster_xs
.iter()
.map(|xs| xs.iter().sum::<f32>() / xs.len() as f32)
.collect();
// Filter columns - each should have multiple items
let min_items_per_col = (items.len() / columns.len().max(1) / 4).max(2);
let columns: Vec<f32> = columns
@@ -123,8 +132,9 @@ pub(crate) fn find_column_boundaries(
.collect();
log::debug!(
" find_column_boundaries: {} columns before filter, threshold={:.1}, {} items",
" find_column_boundaries: {} columns (merged from {}), threshold={:.1}, {} items",
columns.len(),
columns_before_merge,
cluster_threshold,
items.len()
);
@@ -148,10 +158,120 @@ pub(crate) fn find_column_boundaries(
columns
}
/// Check if a text string looks like a number (digits, decimals, sign, comma).
fn is_numeric_text(s: &str) -> bool {
let s = s.trim();
if s.is_empty() {
return false;
}
// Match patterns like: 8.23, -1.05, 9.99, 7.12, 100, 3,456.78, +5%, ---
// But NOT: BIO, Department, Core Courses
s.chars()
.all(|c| c.is_ascii_digit() || c == '.' || c == ',' || c == '-' || c == '+' || c == '%')
&& s.chars().any(|c| c.is_ascii_digit())
}
/// Merge adjacent X-position clusters when one is a sparse header cluster
/// and the other is a dense numeric data cluster. This prevents multi-line
/// wrapped headers from splitting a logical column into two clusters.
fn merge_numeric_adjacent_clusters(
mut clusters: Vec<Vec<f32>>,
items: &[(usize, &TextItem)],
threshold: f32,
) -> Vec<Vec<f32>> {
// For each cluster, compute: center, item count, numeric fraction
struct ClusterInfo {
center: f32,
count: usize,
numeric_frac: f32,
}
let compute_info = |xs: &[f32]| -> ClusterInfo {
let center = xs.iter().sum::<f32>() / xs.len() as f32;
// Count items and numeric fraction for items near this cluster center
let mut total = 0;
let mut numeric = 0;
for (_, item) in items {
if (item.x - center).abs() < threshold {
total += 1;
if is_numeric_text(&item.text) {
numeric += 1;
}
}
}
ClusterInfo {
center,
count: total,
numeric_frac: if total > 0 {
numeric as f32 / total as f32
} else {
0.0
},
}
};
// Merge distance: allow merging clusters that are slightly beyond the
// original threshold. Use 1.5× threshold to catch header-vs-data splits.
let merge_dist = threshold * 1.5;
// Iterate and merge adjacent pairs. Use a simple left-to-right scan.
let mut merged = true;
while merged {
merged = false;
let mut i = 0;
while i + 1 < clusters.len() {
let info_a = compute_info(&clusters[i]);
let info_b = compute_info(&clusters[i + 1]);
let dist = (info_b.center - info_a.center).abs();
if dist > merge_dist {
i += 1;
continue;
}
// Determine if one cluster is sparse (header) and the other
// is dense and numeric (data). A cluster is "sparse" if it has
// significantly fewer items than the other.
let (sparse, dense) = if info_a.count < info_b.count {
(&info_a, &info_b)
} else {
(&info_b, &info_a)
};
// Merge if the dense cluster is predominantly numeric (>50%)
// and the sparse cluster has at most 1/3 the items of the dense one.
let should_merge =
dense.numeric_frac > 0.50 && sparse.count <= dense.count / 2 && sparse.count <= 5;
if should_merge {
log::debug!(
" merging column clusters: center {:.1} ({} items, {:.0}% numeric) + {:.1} ({} items, {:.0}% numeric), dist={:.1}",
info_a.center,
info_a.count,
info_a.numeric_frac * 100.0,
info_b.center,
info_b.count,
info_b.numeric_frac * 100.0,
dist,
);
// Merge cluster i+1 into cluster i
let next = clusters.remove(i + 1);
clusters[i].extend(next);
merged = true;
// Don't increment i — check if the merged cluster can merge further
} else {
i += 1;
}
}
}
clusters
}
/// Find row boundaries by clustering Y positions
pub(crate) fn find_row_boundaries(items: &[(usize, &TextItem)]) -> Vec<f32> {
let mut y_positions: Vec<f32> = items.iter().map(|(_, i)| i.y).collect();
y_positions.sort_by(|a, b| b.partial_cmp(a).unwrap_or(std::cmp::Ordering::Equal)); // Descending
y_positions.sort_by(|a, b| b.total_cmp(a)); // Descending
if y_positions.is_empty() {
return vec![];
@@ -162,7 +282,7 @@ pub(crate) fn find_row_boundaries(items: &[(usize, &TextItem)]) -> Vec<f32> {
// inter-row gaps (≥1× font size), preventing row merging in uniform-spaced PDFs.
let cluster_threshold = {
let mut font_sizes: Vec<f32> = items.iter().map(|(_, i)| i.font_size).collect();
font_sizes.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
font_sizes.sort_by(|a, b| a.total_cmp(b));
let median_font = font_sizes[font_sizes.len() / 2];
(median_font * 0.8).max(4.0)
};
@@ -379,6 +499,7 @@ pub(crate) fn recover_header_row(
#[cfg(test)]
mod tests {
use super::*;
use crate::tables::TableKind;
use crate::types::ItemType;
fn make_item(text: &str, x: f32, y: f32, font_size: f32) -> TextItem {
@@ -636,6 +757,7 @@ mod tests {
rows: vec![500.0, 480.0],
cells: vec![vec!["A".into(), "B".into()], vec!["C".into(), "D".into()]],
item_indices: vec![2, 3],
kind: TableKind::Data,
};
recover_header_row(&mut table, &all_items, 9.0);
@@ -654,6 +776,7 @@ mod tests {
rows: vec![500.0],
cells: vec![vec!["A".into(), "B".into()]],
item_indices: vec![0, 1],
kind: TableKind::Data,
};
let rows_before = table.rows.len();
@@ -674,6 +797,7 @@ mod tests {
rows: vec![500.0, 480.0],
cells: vec![vec!["A".into(), "B".into()], vec!["C".into(), "D".into()]],
item_indices: vec![2, 3],
kind: TableKind::Data,
};
let rows_before = table.rows.len();
@@ -694,6 +818,7 @@ mod tests {
rows: vec![500.0],
cells: vec![vec!["A".into(), "B".into()]],
item_indices: vec![1, 2],
kind: TableKind::Data,
};
let rows_before = table.rows.len();
@@ -709,6 +834,7 @@ mod tests {
rows: vec![],
cells: vec![],
item_indices: vec![],
kind: TableKind::Data,
};
recover_header_row(&mut table, &all_items, 9.0);
+57 -17
View File
@@ -9,13 +9,16 @@ mod detect_struct;
mod financial;
mod format;
mod grid;
pub mod structured;
pub use detect_heuristic::detect_tables;
pub(crate) use detect_heuristic::is_table_of_contents;
pub use detect_lines::detect_tables_from_lines;
pub(crate) use detect_rects::cluster_rects;
pub use detect_rects::{detect_tables_from_rects, RectHintRegion};
pub use detect_struct::detect_tables_from_struct_tree;
pub use format::table_to_markdown;
pub use structured::{cells_to_markdown, StructuredCell};
use crate::types::TextItem;
@@ -33,7 +36,7 @@ pub(crate) fn try_build_rect_guided_table(
// 1. Derive column boundaries from rect X positions (snapped to 2pt tolerance)
let mut x_lefts: Vec<f32> = cluster_rects.iter().map(|&(x, _, _, _)| x).collect();
x_lefts.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
x_lefts.sort_by(|a, b| a.total_cmp(b));
// Snap: deduplicate within 2pt tolerance
let mut col_boundaries: Vec<f32> = Vec::new();
for x in &x_lefts {
@@ -54,7 +57,7 @@ pub(crate) fn try_build_rect_guided_table(
// boundaries so every day gets a column.
if col_boundaries.len() >= 2 {
let mut spacings: Vec<f32> = col_boundaries.windows(2).map(|w| w[1] - w[0]).collect();
spacings.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
spacings.sort_by(|a, b| a.total_cmp(b));
let median_spacing = spacings[spacings.len() / 2];
let threshold = median_spacing * 1.5;
@@ -87,7 +90,7 @@ pub(crate) fn try_build_rect_guided_table(
// 3. Derive row boundaries from item Y positions (5pt tolerance)
let mut y_values: Vec<f32> = expanded_items.iter().map(|(item, _)| item.y).collect();
y_values.sort_by(|a, b| b.partial_cmp(a).unwrap_or(std::cmp::Ordering::Equal)); // descending
y_values.sort_by(|a, b| b.total_cmp(a)); // descending
let mut row_boundaries: Vec<f32> = Vec::new();
for y in &y_values {
if row_boundaries
@@ -166,12 +169,12 @@ pub(crate) fn try_build_rect_guided_table(
used_indices.sort_unstable();
used_indices.dedup();
Some(Table {
columns: col_boundaries,
rows: row_boundaries,
Some(Table::new(
col_boundaries,
row_boundaries,
cells,
item_indices: used_indices,
})
used_indices,
))
}
/// Split a TextItem whose text contains multiple whitespace-separated tokens
@@ -296,7 +299,7 @@ pub(crate) fn try_build_table_from_columns(items: &[TextItem], page: u32) -> Opt
// Find the top-most row with items in multiple columns (likely the header)
let mut ys: Vec<f32> = page_items.iter().map(|i| i.y).collect();
ys.sort_by(|a, b| b.partial_cmp(a).unwrap_or(std::cmp::Ordering::Equal));
ys.sort_by(|a, b| b.total_cmp(a));
ys.dedup_by(|a, b| (*a - *b).abs() < y_tol);
for &header_y in ys.iter().take(5) {
@@ -318,7 +321,7 @@ pub(crate) fn try_build_table_from_columns(items: &[TextItem], page: u32) -> Opt
if col_items.len() >= 2 {
// Sort by X and find the split point
let mut sorted: Vec<f32> = col_items.iter().map(|i| i.x).collect();
sorted.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
sorted.sort_by(|a, b| a.total_cmp(b));
// Split at the midpoint between the two items
let split_x = (sorted[0]
+ col_items.iter().find(|i| i.x == sorted[0]).unwrap().width
@@ -411,7 +414,7 @@ pub(crate) fn try_build_table_from_columns(items: &[TextItem], page: u32) -> Opt
}
}
}
row_ys.sort_by(|a, b| b.partial_cmp(a).unwrap_or(std::cmp::Ordering::Equal));
row_ys.sort_by(|a, b| b.total_cmp(a));
if row_ys.len() < 3 || row_ys.len() > 40 {
return None;
@@ -541,12 +544,22 @@ pub(crate) fn try_build_table_from_columns(items: &[TextItem], page: u32) -> Opt
multi_col_rows
);
Some(Table {
columns: col_xs,
rows: row_ys,
cells,
item_indices,
})
Some(Table::new(col_xs, row_ys, cells, item_indices))
}
/// What kind of structure a detected `Table` represents. Classification is
/// computed once at construction so consumers don't have to re-analyze the
/// cells (and stay consistent across detection backends).
#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
pub enum TableKind {
/// A real data table — renders as markdown table syntax.
#[default]
Data,
/// A table of contents — renders as a flat list with tab-aligned page
/// numbers via `format_toc_as_list`. Detected through the table pipeline
/// because TOCs share row/column structure with tables, but they are not
/// data tables and shouldn't appear in `pages_with_tables` etc.
Toc,
}
/// A detected table.
@@ -560,6 +573,31 @@ pub struct Table {
pub cells: Vec<Vec<String>>,
/// Items that belong to this table
pub item_indices: Vec<usize>,
/// Data table vs TOC. Set by `Table::new` from `cells`.
pub kind: TableKind,
}
impl Table {
/// Build a table and classify it (data vs TOC) from its cells.
pub fn new(
columns: Vec<f32>,
rows: Vec<f32>,
cells: Vec<Vec<String>>,
item_indices: Vec<usize>,
) -> Self {
let kind = if is_table_of_contents(&cells) {
TableKind::Toc
} else {
TableKind::Data
};
Self {
columns,
rows,
cells,
item_indices,
kind,
}
}
}
#[cfg(test)]
@@ -642,6 +680,7 @@ mod tests {
vec!["Cell 1".into(), "Cell 2".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
@@ -1002,6 +1041,7 @@ mod tests {
vec!["3".into(), "5/2".into(), "Item C".into(), "300".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
+972
View File
@@ -0,0 +1,972 @@
//! Structure-recovery-aware (TSR) table assembly.
//!
//! Consumes the raw output of an external table-structure recognition model
//! (e.g. SLANet on PaddleOCR): a flat list of HTML structure tokens plus a
//! parallel list of per-cell bboxes. Pairs each cell open-tag with its bbox
//! in document order, tracks row/column position with rowspan/colspan
//! awareness, and emits a markdown pipe-table.
//!
//! No real HTML parser is needed — the token grammar is restricted (see
//! [`parse_structure`]), so a small state machine is enough.
//!
//! Cell text is supplied separately by the caller (typically by overlap-
//! testing PDF text items against each cell's page-PDF-pt bbox).
use std::collections::{HashMap, HashSet};
/// A single resolved cell, with both structural metadata and its bbox in
/// page PDF-points (top-left origin).
#[derive(Debug, Clone)]
pub struct StructuredCell {
/// 0-indexed grid row.
pub row: usize,
/// 0-indexed grid column.
pub col: usize,
/// 1 for a normal cell.
pub rowspan: usize,
/// 1 for a normal cell.
pub colspan: usize,
/// `true` when the cell is a `<th>` or sits inside `<thead>`.
pub is_header: bool,
/// Cell text (filled in by the caller after overlap-testing PDF items).
pub text: String,
/// Axis-aligned bbox `[x1, y1, x2, y2]` in page PDF-points, top-left origin.
pub page_pt_bbox: [f32; 4],
}
/// Intermediate parse result before the caller fills in text + page coords.
#[derive(Debug, Clone)]
pub(crate) struct CellSlot {
pub row: usize,
pub col: usize,
pub rowspan: usize,
pub colspan: usize,
pub is_header: bool,
/// Index into the parallel `cell_bboxes` array.
pub bbox_idx: usize,
}
/// Parse a sequence of SLANet structure tokens into ordered cell slots.
///
/// Token grammar (no real HTML parsing required):
/// - Section markers: `<thead>`, `</thead>`, `<tbody>`, `</tbody>` and
/// wrapper tokens (`<html>`, `<body>`, `<table>`, plus closing variants)
/// are tracked or skipped.
/// - Row markers: `<tr>` opens a new row, `</tr>` is informational.
/// - Empty cell, single token: `<td></td>` or `<th></th>`.
/// - Cell with attributes, multi-token sequence: `<td` (or `<th`), then
/// attribute fragments like ` colspan="4"`, then `>`, then later `</td>`
/// (or `</th>`). Cells get paired with the next bbox in document order.
///
/// Cells inside `<thead>` and any `<th>` cells are flagged as headers.
/// rowspan/colspan attributes are honoured and prior-row rowspans push
/// later-row cells to the right.
pub(crate) fn parse_structure(tokens: &[String]) -> Vec<CellSlot> {
let mut slots: Vec<CellSlot> = Vec::new();
let mut occupied: HashSet<(usize, usize)> = HashSet::new();
let mut row: usize = 0;
let mut col: usize = 0;
let mut bbox_idx: usize = 0;
let mut in_thead = false;
let mut started_first_row = false;
let mut i = 0;
while i < tokens.len() {
let tok = tokens[i].trim();
match tok {
"<thead>" => {
in_thead = true;
}
"</thead>" => {
in_thead = false;
}
"<tr>" => {
if started_first_row {
row += 1;
}
col = 0;
started_first_row = true;
}
"<td></td>" | "<th></th>" => {
let is_th = tok == "<th></th>";
while occupied.contains(&(row, col)) {
col += 1;
}
slots.push(CellSlot {
row,
col,
rowspan: 1,
colspan: 1,
is_header: in_thead || is_th,
bbox_idx,
});
bbox_idx += 1;
col += 1;
}
"<td" | "<th" => {
let is_th = tok == "<th";
let mut rowspan: usize = 1;
let mut colspan: usize = 1;
// Consume attribute fragments until we hit ">".
i += 1;
while i < tokens.len() && tokens[i].trim() != ">" {
let attr = tokens[i].as_str();
if let Some(v) = parse_int_attr(attr, "rowspan") {
rowspan = v.max(1);
} else if let Some(v) = parse_int_attr(attr, "colspan") {
colspan = v.max(1);
}
i += 1;
}
// i now points at the `>` token (or off the end if malformed).
while occupied.contains(&(row, col)) {
col += 1;
}
slots.push(CellSlot {
row,
col,
rowspan,
colspan,
is_header: in_thead || is_th,
bbox_idx,
});
for r in row..row + rowspan {
for c in col..col + colspan {
occupied.insert((r, c));
}
}
bbox_idx += 1;
col += colspan;
}
// Wrapper / informational tokens — no-op.
_ => {}
}
i += 1;
}
slots
}
/// Parse an attribute fragment like ` colspan="4"` or `rowspan='2'`.
///
/// Tolerates leading whitespace and either single or double quotes.
fn parse_int_attr(s: &str, name: &str) -> Option<usize> {
let trimmed = s.trim();
if !trimmed.starts_with(name) {
return None;
}
let rest = trimmed[name.len()..].trim_start();
let rest = rest.strip_prefix('=')?.trim_start();
let value = rest
.trim_start_matches(['"', '\''])
.trim_end_matches(['"', '\'']);
value.parse().ok()
}
/// Convert a SLANet polygon (4 or 8 elements) into an axis-aligned
/// `[x1, y1, x2, y2]` rect.
///
/// 8-element form: `[x1,y1, x2,y1, x2,y2, x1,y2]` (4 corners). We ignore the
/// implicit corner order and just take min/max so rotated polygons collapse
/// to a sane bounding box.
///
/// 4-element form: `[x1, y1, x2, y2]` (axis-aligned, older SLANet variants).
pub(crate) fn polygon_to_aabb(coords: &[f32]) -> Option<[f32; 4]> {
match coords.len() {
4 => {
let x1 = coords[0].min(coords[2]);
let y1 = coords[1].min(coords[3]);
let x2 = coords[0].max(coords[2]);
let y2 = coords[1].max(coords[3]);
Some([x1, y1, x2, y2])
}
8 => {
let xs = [coords[0], coords[2], coords[4], coords[6]];
let ys = [coords[1], coords[3], coords[5], coords[7]];
let x1 = xs.iter().copied().fold(f32::INFINITY, f32::min);
let y1 = ys.iter().copied().fold(f32::INFINITY, f32::min);
let x2 = xs.iter().copied().fold(f32::NEG_INFINITY, f32::max);
let y2 = ys.iter().copied().fold(f32::NEG_INFINITY, f32::max);
if x1.is_finite() && y1.is_finite() && x2.is_finite() && y2.is_finite() {
Some([x1, y1, x2, y2])
} else {
None
}
}
_ => None,
}
}
/// Convert a cell rect from crop image-pixel space to page PDF-points
/// (top-left origin), given the crop's PDF-point offset on the page and the
/// DPI the crop image was rendered at.
pub(crate) fn cell_px_to_page_pt(
cell_px: [f32; 4],
render_dpi: f32,
crop_origin_pt: [f32; 2],
) -> [f32; 4] {
let pt_per_px = if render_dpi > 0.0 {
72.0 / render_dpi
} else {
1.0
};
let [x_off, y_off] = crop_origin_pt;
[
cell_px[0] * pt_per_px + x_off,
cell_px[1] * pt_per_px + y_off,
cell_px[2] * pt_per_px + x_off,
cell_px[3] * pt_per_px + y_off,
]
}
/// Refine TSR cell bboxes into non-overlapping row/column bands.
///
/// SLANet-style bboxes are often plausible but too tall on dense borderless
/// tables. Native PDF text assignment is more reliable when each parsed row
/// owns the band between neighboring row centers instead of the full model box.
pub(crate) fn normalize_cell_bands(cells: &mut [StructuredCell]) {
if cells.len() < 2 {
return;
}
let row_bands = derive_axis_bands(cells, Axis::Y);
let col_bands = derive_axis_bands(cells, Axis::X);
for cell in cells {
let row_end = cell.row + cell.rowspan.max(1).saturating_sub(1);
if let (Some(&(y1, _)), Some(&(_, y2))) =
(row_bands.get(&cell.row), row_bands.get(&row_end))
{
let clamped_y1 = cell.page_pt_bbox[1].max(y1);
let clamped_y2 = cell.page_pt_bbox[3].min(y2);
if clamped_y1 < clamped_y2 {
cell.page_pt_bbox[1] = clamped_y1;
cell.page_pt_bbox[3] = clamped_y2;
}
}
let col_end = cell.col + cell.colspan.max(1).saturating_sub(1);
if let (Some(&(x1, _)), Some(&(_, x2))) =
(col_bands.get(&cell.col), col_bands.get(&col_end))
{
let clamped_x1 = cell.page_pt_bbox[0].max(x1);
let clamped_x2 = cell.page_pt_bbox[2].min(x2);
if clamped_x1 < clamped_x2 {
cell.page_pt_bbox[0] = clamped_x1;
cell.page_pt_bbox[2] = clamped_x2;
}
}
}
}
#[derive(Clone, Copy)]
enum Axis {
X,
Y,
}
fn derive_axis_bands(cells: &[StructuredCell], axis: Axis) -> HashMap<usize, (f32, f32)> {
let mut by_index: HashMap<usize, Vec<(f32, f32)>> = HashMap::new();
// Prefer non-spanning cells so colspan/rowspan boxes do not skew a single
// column/row center. If an axis has no non-spanning examples for an index,
// fall back to anchored cells below.
for cell in cells {
let span = match axis {
Axis::X => cell.colspan.max(1),
Axis::Y => cell.rowspan.max(1),
};
if span == 1 {
let idx = match axis {
Axis::X => cell.col,
Axis::Y => cell.row,
};
by_index
.entry(idx)
.or_default()
.push(axis_bounds(cell.page_pt_bbox, axis));
}
}
for cell in cells {
let idx = match axis {
Axis::X => cell.col,
Axis::Y => cell.row,
};
if !by_index.contains_key(&idx) {
by_index
.entry(idx)
.or_default()
.push(axis_bounds(cell.page_pt_bbox, axis));
}
}
let mut rows: Vec<(usize, f32, f32, f32)> = by_index
.into_iter()
.filter_map(|(idx, bounds)| {
let mut min_edge = f32::INFINITY;
let mut max_edge = f32::NEG_INFINITY;
let mut center_sum = 0.0;
let mut count = 0usize;
for (lo, hi) in bounds {
if lo.is_finite() && hi.is_finite() && lo < hi {
min_edge = min_edge.min(lo);
max_edge = max_edge.max(hi);
center_sum += (lo + hi) * 0.5;
count += 1;
}
}
(count > 0).then_some((idx, center_sum / count as f32, min_edge, max_edge))
})
.collect();
if rows.len() < 2 {
return rows
.into_iter()
.map(|(idx, _center, lo, hi)| (idx, (lo, hi)))
.collect();
}
rows.sort_by_key(|(idx, _, _, _)| *idx);
let mut bands = HashMap::new();
for i in 0..rows.len() {
let (idx, _center, min_edge, max_edge) = rows[i];
let lo = if i == 0 {
min_edge
} else {
(rows[i - 1].1 + rows[i].1) * 0.5
};
let hi = if i + 1 == rows.len() {
max_edge
} else {
(rows[i].1 + rows[i + 1].1) * 0.5
};
if lo.is_finite() && hi.is_finite() && lo < hi {
bands.insert(idx, (lo, hi));
}
}
bands
}
fn axis_bounds(bbox: [f32; 4], axis: Axis) -> (f32, f32) {
match axis {
Axis::X => (bbox[0].min(bbox[2]), bbox[0].max(bbox[2])),
Axis::Y => (bbox[1].min(bbox[3]), bbox[1].max(bbox[3])),
}
}
/// Sanitize cell text for inclusion in a markdown pipe-table cell:
/// collapse whitespace runs, drop newlines/tabs (cells must be one line),
/// and escape pipes that would otherwise break the table.
fn sanitize_cell(text: &str) -> String {
let mut s = String::with_capacity(text.len());
let mut prev_space = false;
for c in text.chars() {
match c {
'|' => {
s.push_str("\\|");
prev_space = false;
}
'\n' | '\r' | '\t' | ' ' => {
if !prev_space {
s.push(' ');
}
prev_space = true;
}
other => {
s.push(other);
prev_space = false;
}
}
}
s.trim().to_string()
}
/// Render a list of explicitly-positioned cells as a markdown pipe-table.
///
/// Grid dimensions are inferred from the cells' (row, col, rowspan, colspan)
/// extents. A cell with colspan/rowspan > 1 is rendered in its top-left
/// position; the absorbed grid positions are emitted as empty cells so the
/// markdown stays a valid rectangular grid that downstream readers can
/// column-count correctly.
///
/// The separator row (`|---|...|`) is emitted after the **last** row that
/// contains a header cell (`is_header == true`). When no cells are flagged
/// as headers — e.g. the upstream TSR model didn't emit `<thead>`/`<th>` —
/// the separator falls back to "after row 0" so the output is still a
/// valid pipe-table.
pub fn cells_to_markdown(cells: &[StructuredCell]) -> String {
if cells.is_empty() {
return String::new();
}
let num_rows = cells
.iter()
.map(|c| c.row + c.rowspan.max(1))
.max()
.unwrap_or(0);
let num_cols = cells
.iter()
.map(|c| c.col + c.colspan.max(1))
.max()
.unwrap_or(0);
if num_rows == 0 || num_cols == 0 {
return String::new();
}
// Separator goes after the last header row, falling back to row 0 when
// no header cells exist. Clamped into range so a malformed cell with
// row >= num_rows can't push it past the table.
let separator_after_row = cells
.iter()
.filter(|c| c.is_header)
.map(|c| c.row)
.max()
.unwrap_or(0)
.min(num_rows.saturating_sub(1));
let mut grid: Vec<Vec<String>> = vec![vec![String::new(); num_cols]; num_rows];
for cell in cells {
if cell.row < num_rows && cell.col < num_cols {
grid[cell.row][cell.col] = sanitize_cell(&cell.text);
}
}
let mut output = String::new();
for (row_idx, row) in grid.iter().enumerate() {
output.push('|');
for cell in row {
output.push_str(cell);
output.push('|');
}
output.push('\n');
if row_idx == separator_after_row {
output.push('|');
for _ in 0..num_cols {
output.push_str("---|");
}
output.push('\n');
}
}
output
}
#[cfg(test)]
mod tests {
use super::*;
fn t(s: &str) -> String {
s.to_string()
}
/// Tokens for the synthetic 3×3 grid example (one colspan-4 row + two
/// data rows of 4 cells each = 9 cells total, 3 rows × 4 cols).
fn synthetic_3x3_tokens() -> Vec<String> {
vec![
"<html>",
"<body>",
"<table>",
"<tbody>",
"<tr>",
"<td",
" colspan=\"4\"",
">",
"</td>",
"</tr>",
"<tr>",
"<td></td>",
"<td></td>",
"<td></td>",
"<td></td>",
"</tr>",
"<tr>",
"<td></td>",
"<td></td>",
"<td></td>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
"</body>",
"</html>",
]
.into_iter()
.map(t)
.collect()
}
/// Bboxes for the synthetic 3×3 grid (8-element polygon form), all
/// within a 400×120 px crop.
fn synthetic_3x3_bboxes() -> Vec<Vec<f32>> {
vec![
vec![3.0, 2.0, 395.0, 2.0, 396.0, 59.0, 3.0, 59.0],
vec![26.0, 62.0, 140.0, 62.0, 141.0, 120.0, 26.0, 120.0],
vec![149.0, 64.0, 248.0, 64.0, 248.0, 119.0, 149.0, 119.0],
vec![257.0, 64.0, 350.0, 64.0, 350.0, 119.0, 257.0, 119.0],
vec![359.0, 64.0, 395.0, 64.0, 395.0, 119.0, 359.0, 119.0],
vec![26.0, 122.0, 140.0, 122.0, 140.0, 178.0, 26.0, 178.0],
vec![149.0, 124.0, 248.0, 124.0, 248.0, 179.0, 149.0, 179.0],
vec![257.0, 124.0, 350.0, 124.0, 350.0, 179.0, 257.0, 179.0],
vec![359.0, 124.0, 395.0, 124.0, 395.0, 179.0, 359.0, 179.0],
]
}
#[test]
fn parse_structure_synthetic_3x3() {
let tokens = synthetic_3x3_tokens();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 9, "should parse 9 cells");
// Cell 0: row 0 col 0, colspan 4
assert_eq!(slots[0].row, 0);
assert_eq!(slots[0].col, 0);
assert_eq!(slots[0].colspan, 4);
assert_eq!(slots[0].rowspan, 1);
// Cells 1..5: row 1, cols 0..3
for (i, slot) in slots.iter().enumerate().skip(1).take(4) {
assert_eq!(slot.row, 1, "cell {i}: row should be 1");
assert_eq!(slot.col, i - 1, "cell {i}: col should be {}", i - 1);
assert_eq!(slot.colspan, 1);
assert_eq!(slot.rowspan, 1);
}
// Cells 5..9: row 2, cols 0..3
for (i, slot) in slots.iter().enumerate().skip(5).take(4) {
assert_eq!(slot.row, 2, "cell {i}: row should be 2");
assert_eq!(slot.col, i - 5);
assert_eq!(slot.colspan, 1);
}
}
#[test]
fn polygon_to_aabb_8elt() {
// Synthetic cell bbox 0
let coords = vec![3.0, 2.0, 395.0, 2.0, 396.0, 59.0, 3.0, 59.0];
let aabb = polygon_to_aabb(&coords).unwrap();
assert_eq!(aabb, [3.0, 2.0, 396.0, 59.0]);
}
#[test]
fn polygon_to_aabb_4elt() {
let coords = vec![5.0, 10.0, 50.0, 60.0];
let aabb = polygon_to_aabb(&coords).unwrap();
assert_eq!(aabb, [5.0, 10.0, 50.0, 60.0]);
}
#[test]
fn polygon_to_aabb_4elt_unordered() {
// Caller may pass corners in any order; min/max should normalise.
let coords = vec![50.0, 60.0, 5.0, 10.0];
let aabb = polygon_to_aabb(&coords).unwrap();
assert_eq!(aabb, [5.0, 10.0, 50.0, 60.0]);
}
#[test]
fn polygon_to_aabb_invalid_len() {
assert!(polygon_to_aabb(&[1.0, 2.0, 3.0]).is_none());
assert!(polygon_to_aabb(&[1.0; 6]).is_none());
assert!(polygon_to_aabb(&[]).is_none());
}
#[test]
fn synthetic_3x3_aabbs_inside_crop() {
// All 9 bboxes should produce valid (x1<x2, y1<y2) rects within the
// crop bounds (400 wide, ~180 tall by inspection of the fixture).
let bboxes = synthetic_3x3_bboxes();
assert_eq!(bboxes.len(), 9);
for (i, bb) in bboxes.iter().enumerate() {
let aabb = polygon_to_aabb(bb).unwrap_or_else(|| panic!("bbox {i} invalid"));
assert!(aabb[0] < aabb[2], "bbox {i}: x1 < x2");
assert!(aabb[1] < aabb[3], "bbox {i}: y1 < y2");
assert!(aabb[0] >= 0.0 && aabb[2] <= 500.0, "bbox {i}: within crop");
assert!(aabb[1] >= 0.0 && aabb[3] <= 200.0, "bbox {i}: within crop");
}
}
#[test]
fn normalize_cell_bands_splits_overlapping_slanet_rows() {
let mut cells = vec![
StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 1,
is_header: true,
text: String::new(),
page_pt_bbox: [10.0, 100.0, 90.0, 120.0],
},
StructuredCell {
row: 0,
col: 1,
rowspan: 1,
colspan: 1,
is_header: true,
text: String::new(),
page_pt_bbox: [90.0, 100.0, 170.0, 120.0],
},
StructuredCell {
row: 1,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [10.0, 116.0, 90.0, 136.0],
},
StructuredCell {
row: 1,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [90.0, 116.0, 170.0, 136.0],
},
];
normalize_cell_bands(&mut cells);
assert_eq!(cells[0].page_pt_bbox[3], cells[2].page_pt_bbox[1]);
assert_eq!(cells[1].page_pt_bbox[3], cells[3].page_pt_bbox[1]);
assert!(
(cells[0].page_pt_bbox[3] - 118.0).abs() < 0.01,
"row separator should be midpoint between row centers: {:?}",
cells
);
}
#[test]
fn normalize_cell_bands_preserves_colspan_extent() {
let mut cells = vec![
StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 2,
is_header: true,
text: String::new(),
page_pt_bbox: [8.0, 80.0, 172.0, 98.0],
},
StructuredCell {
row: 1,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [10.0, 96.0, 90.0, 114.0],
},
StructuredCell {
row: 1,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [88.0, 96.0, 170.0, 114.0],
},
];
normalize_cell_bands(&mut cells);
assert!(
cells[0].page_pt_bbox[0] <= cells[1].page_pt_bbox[0],
"spanning cell should retain the first column's left edge"
);
assert!(
cells[0].page_pt_bbox[2] >= cells[2].page_pt_bbox[2],
"spanning cell should retain the last column's right edge"
);
}
#[test]
fn parse_int_attr_basic() {
assert_eq!(parse_int_attr(" colspan=\"4\"", "colspan"), Some(4));
assert_eq!(parse_int_attr(" rowspan=\"2\"", "rowspan"), Some(2));
assert_eq!(parse_int_attr("colspan='3'", "colspan"), Some(3));
assert_eq!(parse_int_attr(" colspan=\"4\"", "rowspan"), None);
assert_eq!(parse_int_attr(" class=\"foo\"", "colspan"), None);
}
#[test]
fn parse_structure_rowspan_pushes_next_row_right() {
// <tr><td rowspan="2">A</td><td>B</td></tr><tr><td>C</td></tr>
// Expected: A at (0,0), B at (0,1), C at (1,1) — col 0 of row 1
// is occupied by A's rowspan.
let tokens: Vec<String> = vec![
"<table>",
"<tbody>",
"<tr>",
"<td",
" rowspan=\"2\"",
">",
"</td>",
"<td></td>",
"</tr>",
"<tr>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(t)
.collect();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 3);
assert_eq!((slots[0].row, slots[0].col), (0, 0));
assert_eq!(slots[0].rowspan, 2);
assert_eq!((slots[1].row, slots[1].col), (0, 1));
// C should be at (1, 1) because (1, 0) is occupied by A's rowspan.
assert_eq!((slots[2].row, slots[2].col), (1, 1));
}
#[test]
fn parse_structure_thead_marks_headers() {
// <thead><tr><th>H1</th><th>H2</th></tr></thead>
// <tbody><tr><td>D1</td><td>D2</td></tr></tbody>
let tokens: Vec<String> = vec![
"<table>",
"<thead>",
"<tr>",
"<th></th>",
"<th></th>",
"</tr>",
"</thead>",
"<tbody>",
"<tr>",
"<td></td>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(t)
.collect();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 4);
assert!(slots[0].is_header && slots[1].is_header);
assert!(!slots[2].is_header && !slots[3].is_header);
}
#[test]
fn parse_structure_th_outside_thead_still_header() {
// A row-header style: leading <th> in tbody.
let tokens: Vec<String> = vec![
"<table>",
"<tbody>",
"<tr>",
"<th></th>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(t)
.collect();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 2);
assert!(slots[0].is_header);
assert!(!slots[1].is_header);
}
#[test]
fn parse_structure_th_with_attrs() {
let tokens: Vec<String> = vec![
"<table>",
"<thead>",
"<tr>",
"<th",
" colspan=\"2\"",
">",
"</th>",
"</tr>",
"</thead>",
"</table>",
]
.into_iter()
.map(t)
.collect();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 1);
assert_eq!(slots[0].colspan, 2);
assert!(slots[0].is_header);
}
#[test]
fn cells_to_markdown_synthetic_3x3() {
// Build the cells the parser would produce for the synthetic grid,
// and provide some sample text so we can sanity-check output.
let cells = vec![
StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 4,
is_header: false,
text: "Title".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 1,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: "a".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 1,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: "b".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 1,
col: 2,
rowspan: 1,
colspan: 1,
is_header: false,
text: "c".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 1,
col: 3,
rowspan: 1,
colspan: 1,
is_header: false,
text: "d".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
];
let md = cells_to_markdown(&cells);
// Header row contains the spanning cell text in col 0 and pads to 4 cols.
// Absorbed-by-colspan positions render as empty cells (no padding).
assert!(md.starts_with("|Title||||\n"), "got: {md}");
assert!(md.contains("|---|---|---|---|\n"));
assert!(md.contains("|a|b|c|d|\n"));
}
#[test]
fn cells_to_markdown_escapes_pipes() {
let cells = vec![
StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: "a|b".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 0,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: "x".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
];
let md = cells_to_markdown(&cells);
assert!(md.contains("|a\\|b|x|"));
}
#[test]
fn cells_to_markdown_collapses_whitespace_and_newlines() {
let cells = vec![StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: "foo \n bar\tbaz".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
}];
let md = cells_to_markdown(&cells);
assert!(md.contains("|foo bar baz|"));
}
fn cell(row: usize, col: usize, is_header: bool, text: &str) -> StructuredCell {
StructuredCell {
row,
col,
rowspan: 1,
colspan: 1,
is_header,
text: text.into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
}
}
#[test]
fn cells_to_markdown_separator_after_last_header_row() {
// Two-row header (a multi-row thead), then two body rows. Separator
// should land after row 1 (the LAST header row), not after row 0.
let cells = vec![
cell(0, 0, true, "H0a"),
cell(0, 1, true, "H0b"),
cell(1, 0, true, "H1a"),
cell(1, 1, true, "H1b"),
cell(2, 0, false, "d0a"),
cell(2, 1, false, "d0b"),
cell(3, 0, false, "d1a"),
cell(3, 1, false, "d1b"),
];
let md = cells_to_markdown(&cells);
let expected = "|H0a|H0b|\n|H1a|H1b|\n|---|---|\n|d0a|d0b|\n|d1a|d1b|\n";
assert_eq!(md, expected, "got: {md}");
}
#[test]
fn cells_to_markdown_separator_when_row_0_not_header() {
// Row 0 is not flagged as a header but row 1 is. Separator should
// follow row 1 (the header), demonstrating that we don't blindly
// emit after row 0.
let cells = vec![
cell(0, 0, false, "x0a"),
cell(0, 1, false, "x0b"),
cell(1, 0, true, "Hdr1"),
cell(1, 1, true, "Hdr2"),
cell(2, 0, false, "data1"),
cell(2, 1, false, "data2"),
];
let md = cells_to_markdown(&cells);
// Confirm the separator is NOT after row 0.
assert!(!md.starts_with("|x0a|x0b|\n|---|"), "got: {md}");
// Confirm it IS after row 1.
assert!(
md.contains("|Hdr1|Hdr2|\n|---|---|\n|data1|data2|"),
"got: {md}"
);
}
#[test]
fn cells_to_markdown_no_headers_falls_back_to_row_0() {
// No header cells at all — fallback: separator after row 0 so the
// output is still a valid markdown pipe-table.
let cells = vec![
cell(0, 0, false, "a"),
cell(0, 1, false, "b"),
cell(1, 0, false, "c"),
cell(1, 1, false, "d"),
];
let md = cells_to_markdown(&cells);
assert_eq!(md, "|a|b|\n|---|---|\n|c|d|\n");
}
}
+4 -4
View File
@@ -68,9 +68,9 @@ where
pub(crate) fn sort_line_items(items: &mut [TextItem]) {
let rtl = is_rtl_text(items.iter().map(|i| &i.text));
if rtl {
items.sort_by(|a, b| b.x.partial_cmp(&a.x).unwrap_or(std::cmp::Ordering::Equal));
items.sort_by(|a, b| b.x.total_cmp(&a.x));
} else {
items.sort_by(|a, b| a.x.partial_cmp(&b.x).unwrap_or(std::cmp::Ordering::Equal));
items.sort_by(|a, b| a.x.total_cmp(&b.x));
}
}
@@ -376,7 +376,7 @@ fn compute_canva_join_threshold(items: &[TextItem]) -> f32 {
}
let mut sorted: Vec<f32> = ratios;
sorted.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
sorted.sort_by(|a, b| a.total_cmp(b));
if sorted[sorted.len() - 1] < 0.40 || sorted[0] < 0.40 {
return DEFAULT;
@@ -478,7 +478,7 @@ fn compute_single_char_join_threshold(items: &[TextItem]) -> f32 {
return DEFAULT;
}
ratios.sort_by(|a, b| a.partial_cmp(b).unwrap_or(std::cmp::Ordering::Equal));
ratios.sort_by(|a, b| a.total_cmp(b));
// If all gaps are tight (max < 0.40), use default — normal PDF
let max_ratio = ratios[ratios.len() - 1];
+366 -12
View File
@@ -520,6 +520,18 @@ impl ToUnicodeCMap {
}
}
/// Get the maximum source CID across all mappings (char_map + ranges).
fn max_source_cid(&self) -> Option<u16> {
let char_max = self.char_map.keys().copied().max();
let range_max = self.ranges.iter().map(|&(_, end, _)| end).max();
match (char_max, range_max) {
(Some(a), Some(b)) => Some(a.max(b)),
(a @ Some(_), None) => a,
(None, b @ Some(_)) => b,
(None, None) => None,
}
}
/// Remap a CMap that references pre-subsetting GIDs to sequential post-subsetting GIDs.
/// Collects all source CIDs, sorts them, and reassigns to 1, 2, 3, ...
pub fn remap_to_sequential(&self) -> ToUnicodeCMap {
@@ -657,6 +669,81 @@ fn get_w_array_start_cid(cid_font_dict: &lopdf::Dictionary, doc: &Document) -> O
}
}
/// Return true if the CIDFont's W (widths) array explicitly covers the given CID.
///
/// The W array uses two formats (PDF 32000-1:2008, §9.7.4.3):
/// 1. `c [w1 w2 ... wn]` — widths for CIDs c, c+1, ..., c+n-1
/// 2. `c_first c_last w` — CIDs c_first..c_last all have width w
fn w_array_covers_cid(cid_font_dict: &lopdf::Dictionary, doc: &Document, target: u16) -> bool {
let Ok(w_obj) = cid_font_dict.get(b"W") else {
return false;
};
let arr = match w_obj {
Object::Array(arr) => arr,
Object::Reference(r) => match doc.get_object(*r) {
Ok(Object::Array(arr)) => arr,
_ => return false,
},
_ => return false,
};
let resolve_int = |o: &Object| -> Option<i64> {
match o {
Object::Integer(n) => Some(*n),
Object::Reference(r) => match doc.get_object(*r) {
Ok(Object::Integer(n)) => Some(*n),
_ => None,
},
_ => None,
}
};
let resolve_arr = |o: &Object| -> Option<Vec<Object>> {
match o {
Object::Array(a) => Some(a.clone()),
Object::Reference(r) => match doc.get_object(*r) {
Ok(Object::Array(a)) => Some(a.clone()),
_ => None,
},
_ => None,
}
};
let target = target as i64;
let mut i = 0usize;
while i < arr.len() {
let Some(first) = resolve_int(&arr[i]) else {
break;
};
i += 1;
if i >= arr.len() {
break;
}
// Peek at arr[i] to decide format.
if let Some(widths) = resolve_arr(&arr[i]) {
// Format 1: c [w1 ... wn]
let last = first + widths.len() as i64 - 1;
if target >= first && target <= last {
return true;
}
i += 1;
} else if let Some(last) = resolve_int(&arr[i]) {
// Format 2: c_first c_last w
i += 1;
if i < arr.len() {
i += 1; // skip the width value
}
if target >= first && target <= last {
return true;
}
} else {
// Unknown token — abort parsing safely
break;
}
}
false
}
/// Extract CIDToGIDMap as a vector of GIDs (u16) indexed by CID.
fn get_cid_to_gid_map(cid_font_dict: &lopdf::Dictionary, doc: &Document) -> Option<Vec<u16>> {
let obj = cid_font_dict.get(b"CIDToGIDMap").ok()?;
@@ -752,6 +839,20 @@ fn try_remap_subset_cmap(
_ => return (cmap, None),
};
// If the W array actually covers the CMap's max source CID, the CMap is
// aligned with the font — no sequential renumbering happened. A sparse W
// array starting at CID 0 (for .notdef) with additional high-CID entries
// matching the CMap is the normal subset layout, not a mismatch.
if let Some(max_cid) = cmap.max_source_cid() {
if w_array_covers_cid(cid_font_dict, doc, max_cid) {
debug!(
"Subset remap skipped for obj={}: W array covers CMap max CID {}",
obj_num, max_cid
);
return (cmap, None);
}
}
debug!(
"Subset GID mismatch detected for obj={}: W starts at CID {}, CMap min CID {}. Remapping to sequential.",
obj_num, w_start, min_cid
@@ -1650,7 +1751,7 @@ fn merge_cmaps(mut base: ToUnicodeCMap, overlay: ToUnicodeCMap) -> ToUnicodeCMap
///
/// Returns true if the median CID is >= 0x41 (letter 'A'), indicating
/// the PDF generator likely used Unicode codepoints as CIDs.
fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) -> bool {
pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) -> bool {
let w_arr = match cid_font_dict.get(b"W").ok() {
Some(Object::Array(arr)) => arr,
_ => return false,
@@ -1771,15 +1872,49 @@ impl FontCMaps {
/// Iterates every page, collects fonts (including Form XObject fonts),
/// and parses any `/ToUnicode` streams via lopdf's decompression.
pub fn from_doc(doc: &Document) -> Self {
Self::from_doc_pages(doc, None)
}
/// Build FontCMaps for specific pages only. Pass `None` for all pages.
pub fn from_doc_pages(doc: &Document, page_filter: Option<&HashSet<u32>>) -> Self {
Self::from_doc_pages_inner(doc, page_filter, false)
}
/// Build FontCMaps in fast mode: skip expensive TrueType font fallback
/// parsing. Fonts that can't be decoded from their ToUnicode CMap alone
/// will be missing, causing text extraction to produce empty/garbage text
/// which triggers `needs_ocr` fallback. This is ideal for hybrid OCR
/// pipelines where GPU OCR is always available as a fallback.
pub fn from_doc_pages_fast(doc: &Document, page_filter: Option<&HashSet<u32>>) -> Self {
Self::from_doc_pages_inner(doc, page_filter, true)
}
fn from_doc_pages_inner(
doc: &Document,
page_filter: Option<&HashSet<u32>>,
skip_truetype_fallback: bool,
) -> Self {
let mut by_obj_num: HashMap<u32, CMapEntry> = HashMap::new();
for (_page_num, &page_id) in doc.get_pages().iter() {
for (page_num, &page_id) in doc.get_pages().iter() {
if let Some(filter) = page_filter {
if !filter.contains(page_num) {
continue;
}
}
// Page-level fonts (includes inherited parent resources)
let fonts = doc.get_page_fonts(page_id).unwrap_or_default();
Self::collect_cmaps_from_fonts(&fonts, doc, &mut by_obj_num);
Self::collect_cmaps_from_fonts_inner(
&fonts,
doc,
&mut by_obj_num,
skip_truetype_fallback,
);
// Fonts inside Form XObjects referenced by this page
Self::collect_cmaps_from_xobjects(doc, page_id, &mut by_obj_num);
if !skip_truetype_fallback {
// Fonts inside Form XObjects referenced by this page
Self::collect_cmaps_from_xobjects(doc, page_id, &mut by_obj_num);
}
}
FontCMaps { by_obj_num }
@@ -1792,6 +1927,15 @@ impl FontCMaps {
fonts: &std::collections::BTreeMap<Vec<u8>, &lopdf::Dictionary>,
doc: &Document,
by_obj_num: &mut HashMap<u32, CMapEntry>,
) {
Self::collect_cmaps_from_fonts_inner(fonts, doc, by_obj_num, false);
}
fn collect_cmaps_from_fonts_inner(
fonts: &std::collections::BTreeMap<Vec<u8>, &lopdf::Dictionary>,
doc: &Document,
by_obj_num: &mut HashMap<u32, CMapEntry>,
skip_truetype_fallback: bool,
) {
// First pass: collect ToUnicode CMaps
for font_dict in fonts.values() {
@@ -1825,13 +1969,32 @@ impl FontCMaps {
);
let (mut primary, mut remapped) =
try_remap_subset_cmap(cmap, font_dict, doc, obj_num);
let mut fallback = build_fallback_tounicode_from_encoding(font_dict, doc)
.or_else(|| build_fallback_cmap_for_type0(font_dict, doc))
.or_else(|| build_fallback_cmap_for_simple(font_dict, doc));
// If the ToUnicode map is extremely sparse, prefer the fallback
// (often a better mapping for Symbol/Wingdings/Arabic CID fonts).
// Only build expensive fallbacks when the primary CMap is sparse.
// build_fallback_cmap_for_type0 can take seconds on large embedded
// TrueType fonts (decompressing + parsing 100K+ byte font files).
// Skip entirely when the primary CMap is sufficient.
let primary_entries = primary.char_map.len() + primary.ranges.len();
let mut fallback = if primary_entries < 10 && !skip_truetype_fallback {
// Try cheap fallback first; only attempt expensive TrueType
// parsing if cheap fallbacks don't yield results.
let cheap = build_fallback_tounicode_from_encoding(font_dict, doc)
.or_else(|| build_fallback_cmap_for_simple(font_dict, doc));
if cheap.is_some() {
cheap
} else {
build_fallback_cmap_for_type0(font_dict, doc)
}
} else if primary_entries < 10 {
// Fast mode: only try cheap fallbacks, skip TrueType parsing.
// Regions using this font will get needs_ocr=true.
build_fallback_tounicode_from_encoding(font_dict, doc)
.or_else(|| build_fallback_cmap_for_simple(font_dict, doc))
} else {
// Primary is rich enough; only try the cheap encoding fallback
build_fallback_tounicode_from_encoding(font_dict, doc)
};
if primary_entries < 10 {
if let Some(fb) = fallback.take() {
debug!(
@@ -1852,8 +2015,12 @@ impl FontCMaps {
);
} else {
// ToUnicode present but parse failed; try fallbacks to avoid empty decoding.
let fallback = build_fallback_cmap_for_type0(font_dict, doc)
.or_else(|| build_fallback_cmap_for_simple(font_dict, doc));
let fallback = if skip_truetype_fallback {
build_fallback_cmap_for_simple(font_dict, doc)
} else {
build_fallback_cmap_for_type0(font_dict, doc)
.or_else(|| build_fallback_cmap_for_simple(font_dict, doc))
};
if let Some(fb) = fallback {
debug!(
"ToUnicode CMap obj={} parse failed; using fallback (entries={})",
@@ -1874,6 +2041,10 @@ impl FontCMaps {
// Second pass: Identity-H/V fonts without ToUnicode
// Try: (1) embedded TrueType/OpenType cmap, (2) predefined CID→Unicode mapping
// Skip entirely in fast mode — these fonts require expensive TrueType parsing.
if skip_truetype_fallback {
return;
}
for font_dict in fonts.values() {
if font_dict.get(b"ToUnicode").is_ok() {
continue;
@@ -2647,4 +2818,187 @@ endbfchar
assert_eq!(remapped.unwrap().char_map.len(), 50);
assert_eq!(fallback.unwrap().char_map.len(), 10);
}
#[test]
fn test_max_source_cid() {
let cmap_content = r#"
1 begincodespacerange
<0000><FFFF>
endcodespacerange
2 beginbfchar
<0003> <0020>
<0031> <004E>
endbfchar
1 beginbfrange
<0208> <0227> <0430>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.min_source_cid(), Some(0x0003));
assert_eq!(cmap.max_source_cid(), Some(0x0227));
}
/// Helper: build a minimal CIDFont dict with a W array and check coverage.
fn cid_font_dict_with_w(w_items: Vec<lopdf::Object>) -> lopdf::Dictionary {
let mut d = lopdf::Dictionary::new();
d.set("W", lopdf::Object::Array(w_items));
d
}
#[test]
fn test_w_array_covers_cid_format1() {
// Format 1: `c [w1 w2 ... wn]` — widths for CIDs c..c+n-1.
// Mimics the 16.pdf Tahoma W array: 0[1000] 3[313] 5[401] 11[383 383] 16[363 303 382]
let doc = Document::new();
let d = cid_font_dict_with_w(vec![
lopdf::Object::Integer(0),
lopdf::Object::Array(vec![lopdf::Object::Integer(1000)]),
lopdf::Object::Integer(3),
lopdf::Object::Array(vec![lopdf::Object::Integer(313)]),
lopdf::Object::Integer(5),
lopdf::Object::Array(vec![lopdf::Object::Integer(401)]),
lopdf::Object::Integer(11),
lopdf::Object::Array(vec![
lopdf::Object::Integer(383),
lopdf::Object::Integer(383),
]),
lopdf::Object::Integer(16),
lopdf::Object::Array(vec![
lopdf::Object::Integer(363),
lopdf::Object::Integer(303),
lopdf::Object::Integer(382),
]),
lopdf::Object::Integer(570),
lopdf::Object::Array(vec![lopdf::Object::Integer(667); 26]),
]);
assert!(w_array_covers_cid(&d, &doc, 0));
assert!(w_array_covers_cid(&d, &doc, 3));
assert!(w_array_covers_cid(&d, &doc, 5));
assert!(w_array_covers_cid(&d, &doc, 11));
assert!(w_array_covers_cid(&d, &doc, 12));
assert!(w_array_covers_cid(&d, &doc, 16));
assert!(w_array_covers_cid(&d, &doc, 18));
assert!(w_array_covers_cid(&d, &doc, 570));
assert!(w_array_covers_cid(&d, &doc, 595));
// Gaps are NOT covered
assert!(!w_array_covers_cid(&d, &doc, 1));
assert!(!w_array_covers_cid(&d, &doc, 4));
assert!(!w_array_covers_cid(&d, &doc, 19));
assert!(!w_array_covers_cid(&d, &doc, 596));
}
#[test]
fn test_w_array_covers_cid_format2() {
// Format 2: `c_first c_last w` — CIDs c_first..c_last all have width w.
let doc = Document::new();
let d = cid_font_dict_with_w(vec![
lopdf::Object::Integer(100),
lopdf::Object::Integer(120),
lopdf::Object::Integer(500),
]);
assert!(w_array_covers_cid(&d, &doc, 100));
assert!(w_array_covers_cid(&d, &doc, 110));
assert!(w_array_covers_cid(&d, &doc, 120));
assert!(!w_array_covers_cid(&d, &doc, 99));
assert!(!w_array_covers_cid(&d, &doc, 121));
}
#[test]
fn test_w_array_covers_cid_missing_w() {
let doc = Document::new();
let d = lopdf::Dictionary::new();
assert!(!w_array_covers_cid(&d, &doc, 3));
}
#[test]
fn test_try_remap_skipped_when_w_covers_cmap() {
// Simulates 16.pdf: CMap's max source CID (0x0279 = 633) is explicitly
// in the W array, so no subset-renumbering happened — remap must NOT fire.
let cmap_content = r#"
1 begincodespacerange
<0000><FFFF>
endcodespacerange
2 beginbfchar
<0003> <0020>
<0031> <004E>
endbfchar
2 beginbfrange
<023A> <0253> <0410>
<0255> <0279> <042B>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
let mut doc = Document::new();
// Build a CIDFont dict with Identity CIDToGIDMap and a W array that
// covers CID 633 via `597 [widths...]`.
let mut cid_font = lopdf::Dictionary::new();
cid_font.set("CIDToGIDMap", lopdf::Object::Name(b"Identity".to_vec()));
cid_font.set(
"W",
lopdf::Object::Array(vec![
lopdf::Object::Integer(0),
lopdf::Object::Array(vec![lopdf::Object::Integer(750)]),
lopdf::Object::Integer(597),
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 37]), // 597..633
]),
);
let cid_font_id = doc.add_object(cid_font);
// Build the Type0 font dict with Identity-H + DescendantFonts ref.
let mut font_dict = lopdf::Dictionary::new();
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
font_dict.set(
"DescendantFonts",
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
);
let (primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 123);
assert!(
remapped.is_none(),
"Remap must be skipped when W covers CMap max CID (this is 16.pdf)"
);
assert_eq!(primary.lookup(0x0003), Some(" ".to_string()));
}
#[test]
fn test_try_remap_fires_for_true_subset_mismatch() {
// True mismatch: CMap has high CIDs (512-544) but W only lists low sequential CIDs.
let cmap_content = r#"
1 begincodespacerange
<0000><FFFF>
endcodespacerange
1 beginbfrange
<0200> <0220> <0410>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
let mut doc = Document::new();
let mut cid_font = lopdf::Dictionary::new();
cid_font.set("CIDToGIDMap", lopdf::Object::Name(b"Identity".to_vec()));
cid_font.set(
"W",
lopdf::Object::Array(vec![
lopdf::Object::Integer(0),
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]), // 0..33
]),
);
let cid_font_id = doc.add_object(cid_font);
let mut font_dict = lopdf::Dictionary::new();
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
font_dict.set(
"DescendantFonts",
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
);
let (_primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 456);
assert!(
remapped.is_some(),
"Remap must fire when CMap's CIDs are outside W array coverage"
);
}
}
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
File diff suppressed because it is too large Load Diff
+4 -4
View File
@@ -1,9 +1,9 @@
본 가격표는 국내 거주 중인 외국인을 위한 한국어 가격표의 비공식 번역본입니다. ※ The post-tax benefit sales price is provided for your reference only, reflecting the current tax benefits and eco-friendly vehicle individual consumption tax reductions. 본 가격표와 한국어 가격표의 내용이 상이한 경우 한국어 가격표의 내용이 우선하므로, 반드시 한국어 가격표의 내용을 확인하십시오. The final sales price may vary depending on the addition of optional items and whether the eco-friendly vehicle criteria are met, so please be sure to check the quotation. This price list is an unofficial translation of the Korean price list for the convenience of foreign residents in South Korea. ※ Please check the Korean price list for information on colors, details, and fuel consumption for each model. If the price list differs from the Korean price list, please check the contents of the Korean price list first. ※ All optional item prices are listed based on pre-tax reduction amounts. The actual sales price, which reflects the total individual consumption tax reduction including optional items, may differ depending on applicable tax benefits. ※ The items (specifications, colors, etc.) and prices listed in this pricing table are subject to change without prior notice depending on the holding of new car launch events, improvements made in automobile performance, introduction of related laws and regulations, and changes in company circumstances. The all-new NEXO Release Date: June 10, 2025 / (Unit: KRW)
|Classification Exclusive Exclusive|Selling price before tax benefit Supply value(surtax) 80,509,000 73,190,000(7,319,000)|Selling price after tax benefit 76,435,000|Standard equipment • Powertrain/Performance: Fuel cell system(150kW drive motor, lithium-ion battery, and reducer), Regenerative braking system, Column-Type Shift By Wire(vibration warning), Drive mode select • Safety: 9 airbag system(1st-row advanced/center side airbags, 1st/2nd-row side airbags, and rollover-resistant curtain airbags), Multi-Collision Brake System, Active hood system(for pedestrian protection), Safety unlock function, Artificial engine sound(for pedestrian protection), Child seat fastening system (2 in 2nd-row), Fire extinguisher for vehicles, Pedal Misapplication Safety Assist • Smart Safety Technology: Forward Collision-avoidance Assist(vehicles/ pedestrians/two-wheeled vehicles/junction turning/front oncoming), Smart Cruise Control with Stop & Go, Lane Keeping Assist, Lane Following Assist 2, Blind-spot Collision Warning(driving), Blind-spot Collision-avoidance Assist(forward exit), Rear Cross-traffic Collision-avoidance Assist, Safety Exit Assist, Driver Attention Warning, High Beam Assist, Advanced Rear Occupant Alert, Intelligent Speed Limit Assist, Hands-On Detection, Highway Driving Assist, Navigation-based Smart Cruise Control(safety speed zone/curve control), Vibration warning steering wheel • Exterior: Full LED headlamps(projection type), LED turn signal lamps (front and rear), LED Daytime Running Lights, LED positioning lights, LED rear combination lamps, LED third brake lights, 18-inch alloy wheels & tires, Solar glass(windshield), Double-glazed soundproof glass(windshield, and 1st/ 2nd-row doors), Outside mirror(heating, power-folding, power adjustment,|Options (before tax benefit) ▶ Hi-pass(e hi-pass) [200,000]|
|Classification|Selling price before tax benefit Supply value(surtax)|Selling price after tax benefit|Standard equipment|Options (before tax benefit)|
|---|---|---|---|---|
|Special|with 3.5% individual consumption tax applied 79,287,000 83,500,000 75,909,091(7,590,909) with 3.5% individual consumption tax applied 82,232,000|with 3.5% individual consumption tax applied 76,435,000 79,275,000 with 3.5% individual consumption tax applied 79,275,000|and LED turn signal lamps), Auto flush door handles, Black door garnish • Interior: Panoramic curved display, 12.3-inch color LCD cluster, Leather- upholstered steering wheel(with heating, two-tone color, and Interactive Pixel Lights), LED interior lamp (map lamp, personal lamp, sun visor lamp, and luggage lamp), Metallic door scuff plate • Seat: Synthetic leather seats, 1st-row manual seats, Heated 1st-row seats, 2nd-row 60/40-split folding seats(reclining) • Convenience: Proximity key with push-button start, Smart key remote start, Electronic Parking Brake(with automatic vehicle hold), Paddle shift(regenerative control), Dual-zone full automatic air conditioning(with high-performance antibacterial combination filter, auto defog system, fine dust sensor, air cleaning mode, and after-blow function), 2nd-row seat air vent, Auto light control system, USB Type-C Ports(1×27W switchable charging/data port in 1st-row, and 2×100W charging ports in both 1st and 2nd-row), ECM room mirror(frameless), Rain sensor, Power windows with pinch protection(1st/2nd-row), Power outlet (1 in 1st-row), Parking Distance Warning-Forward/Reverse, Rear View Monitor, Wireless phone charger(single), Walk-away lock, Route planner, Hyundai AI Assistant • Infotainment: 12.3-inch navigation(Bluelink, phone projection, Bluetooth hands-free, and In-car Payment), Audio system(6 speakers), Over-The-Air navigation updates Standard equipment of Exclusive plus • Smart Safety Technology: Forward Collision-avoidance Assist(intersection crossing/changing lanes in oncoming traffic/approaching from either side/ evasive steering assist), Highway Driving Assist 2, Navigation-based Smart Cruise Control(access road) • Exterior: Roof rack • Interior: Metallic pedal, Driving mode-dependent ambient mood lighting(crash pad, 1st/2nd-row door trim) • Seat: Synthetic leather seats(patch applied), Power-adjustable driver's|▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [950,000] Parking Assist ▶ [1,150,000] Audio by BANG & OLUFSEN|
|Prestige|87,893,000 79,902,727(7,990,273) with 3.5% individual consumption tax applied 86,559,000|83,445,000 with 3.5% individual consumption tax applied 83,445,000|seat(8-way, lumbar support, and Integrated Memory System(driver's seat and outside mirror connected)), Power-adjustable front passengers seat(8-way), Ventilated 1st-row seats, Heated 2nd-row seats • Convenience: Hi-pass(e hi-pass), In-car fingerprint authentication system(personalization, startup, payment, and etc.), Smart power tailgate ▶ Standard equipment of Exclusive Special plus • Smart Safety Technology: Remote Smart Parking Assist 2, Parking Collison- avoidance Assist(front/side/rear) • Exterior: Intelligent Front-Lighting System(IFS), Dynamic welcome/escort lighting(1 type), Sequential turn signals(front and rear), Ambient lighting auto flush door handles, Two-tone door garnish, Glossy black rear diffuserInterior: Recycled PET suede interior materials(headlining/sunvisor), Fabric upholstered crash pad • Seat: BIO-processed natural leather seats(metal patch applied, embossed design punching), Passenger's seat walk-in device, 1st-row relaxation comfort seats(leg rest included), Ventilated 2nd-row seats • Convenience: Parking Distance Warning-Side, Head-Up Display, Digital key 2, Wireless phone charger(dual), Surround View Monitor, Blind-spot View Monitor, LED reverse light guide • Infotainment: Audio by BANG & OLUFSEN sound system(14 speakers, including external amp), Active road noise control, Active Sound Design|sound system ▶ [250,000] 19-inch alloy wheels & tires ▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [900,000] Vision roof ▶ [1,380,000] Digital side mirror ▶ [750,000] Camera package ▶ [250,000] 19-inch alloy wheels & tires|
|Exclusive|80,509,000 73,190,000(7,319,000) with 3.5% individual consumption tax applied 79,287,000|76,435,000 with 3.5% individual consumption tax applied 76,435,000|• Powertrain/Performance: Fuel cell system(150kW drive motor, lithium-ion battery, and reducer), Regenerative braking system, Column-Type Shift By Wire(vibration warning), Drive mode select • Safety: 9 airbag system(1st-row advanced/center side airbags, 1st/2nd-row side airbags, and rollover-resistant curtain airbags), Multi-Collision Brake System, Active hood system(for pedestrian protection), Safety unlock function, Artificial engine sound(for pedestrian protection), Child seat fastening system (2 in 2nd-row), Fire extinguisher for vehicles, Pedal Misapplication Safety Assist • Smart Safety Technology: Forward Collision-avoidance Assist(vehicles/ pedestrians/two-wheeled vehicles/junction turning/front oncoming), Smart Cruise Control with Stop & Go, Lane Keeping Assist, Lane Following Assist 2, Blind-spot Collision Warning(driving), Blind-spot Collision-avoidance Assist(forward exit), Rear Cross-traffic Collision-avoidance Assist, Safety Exit Assist, Driver Attention Warning, High Beam Assist, Advanced Rear Occupant Alert, Intelligent Speed Limit Assist, Hands-On Detection, Highway Driving Assist, Navigation-based Smart Cruise Control(safety speed zone/curve control), Vibration warning steering wheel • Exterior: Full LED headlamps(projection type), LED turn signal lamps (front and rear), LED Daytime Running Lights, LED positioning lights, LED rear combination lamps, LED third brake lights, 18-inch alloy wheels & tires, Solar glass(windshield), Double-glazed soundproof glass(windshield, and 1st/ 2nd-row doors), Outside mirror(heating, power-folding, power adjustment, and LED turn signal lamps), Auto flush door handles, Black door garnish • Interior: Panoramic curved display, 12.3-inch color LCD cluster, Leather- upholstered steering wheel(with heating, two-tone color, and Interactive Pixel Lights), LED interior lamp (map lamp, personal lamp, sun visor lamp, and luggage lamp), Metallic door scuff plate • Seat: Synthetic leather seats, 1st-row manual seats, Heated 1st-row seats, 2nd-row 60/40-split folding seats(reclining) • Convenience: Proximity key with push-button start, Smart key remote start, Electronic Parking Brake(with automatic vehicle hold), Paddle shift(regenerative control), Dual-zone full automatic air conditioning(with high-performance antibacterial combination filter, auto defog system, fine dust sensor, air cleaning mode, and after-blow function), 2nd-row seat air vent, Auto light control system, USB Type-C Ports(1×27W switchable charging/data port in 1st-row, and 2×100W charging ports in both 1st and 2nd-row), ECM room mirror(frameless), Rain sensor, Power windows with pinch protection(1st/2nd-row), Power outlet (1 in 1st-row), Parking Distance Warning-Forward/Reverse, Rear View Monitor, Wireless phone charger(single), Walk-away lock, Route planner, Hyundai AI Assistant • Infotainment: 12.3-inch navigation(Bluelink, phone projection, Bluetooth hands-free, and In-car Payment), Audio system(6 speakers), Over-The-Air navigation updates|Hi-pass(e hi-pass) [200,000]|
|Exclusive Special|83,500,000 75,909,091(7,590,909) with 3.5% individual consumption tax applied 82,232,000|79,275,000 with 3.5% individual consumption tax applied 79,275,000|▶ Standard equipment of Exclusive plus • Smart Safety Technology: Forward Collision-avoidance Assist(intersection crossing/changing lanes in oncoming traffic/approaching from either side/ evasive steering assist), Highway Driving Assist 2, Navigation-based Smart Cruise Control(access road)Exterior: Roof rack • Interior: Metallic pedal, Driving mode-dependent ambient mood lighting(crash pad, 1st/2nd-row door trim) • Seat: Synthetic leather seats(patch applied), Power-adjustable driver's seat(8-way, lumbar support, and Integrated Memory System(driver's seat and outside mirror connected)), Power-adjustable front passengers seat(8-way), Ventilated 1st-row seats, Heated 2nd-row seats • Convenience: Hi-pass(e hi-pass), In-car fingerprint authentication system(personalization, startup, payment, and etc.), Smart power tailgate|▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [950,000] Parking Assist ▶ [1,150,000] Audio by BANG & OLUFSEN sound system ▶ [250,000] 19-inch alloy wheels & tires|
|Prestige|87,893,000 79,902,727(7,990,273) with 3.5% individual consumption tax applied 86,559,000|83,445,000 with 3.5% individual consumption tax applied 83,445,000|▶ Standard equipment of Exclusive Special plus • Smart Safety Technology: Remote Smart Parking Assist 2, Parking Collison- avoidance Assist(front/side/rear) • Exterior: Intelligent Front-Lighting System(IFS), Dynamic welcome/escort lighting(1 type), Sequential turn signals(front and rear), Ambient lighting auto flush door handles, Two-tone door garnish, Glossy black rear diffuser • Interior: Recycled PET suede interior materials(headlining/sunvisor), Fabric upholstered crash pad • Seat: BIO-processed natural leather seats(metal patch applied, embossed design punching), Passenger's seat walk-in device, 1st-row relaxation comfort seats(leg rest included), Ventilated 2nd-row seats • Convenience: Parking Distance Warning-Side, Head-Up Display, Digital key 2, Wireless phone charger(dual), Surround View Monitor, Blind-spot View Monitor, LED reverse light guide • Infotainment: Audio by BANG & OLUFSEN sound system(14 speakers, including external amp), Active road noise control, Active Sound Design|▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [900,000] Vision roof ▶ [1,380,000] Digital side mirror ▶ [750,000] Camera package ▶ [250,000] 19-inch alloy wheels & tires|
**Classification Details** **Indoor/outdoor V2L** Indoor V2L, Outdoor V2L(connectorless type) **Parking Assist** Surround View Monitor, Blind-spot View Monitor, Parking Distance Warning-Side, Parking Collison-avoidance Assist-Rear **Audio by BANG & OLUFSEN** Audio by BANG & OLUFSEN sound system(14 speakers, including external amp.), Active road noise control, Active Sound Design **sound system** **Camera package** Digital center mirror(with camera sensor cleaning system), Driver monitoring system THE ALL-NEW NEXO /// ECO-FRIENDLY CAR
+4 -2
View File
@@ -8,7 +8,9 @@ Department of the Treasury **Internal Revenue Service**
# and Report to Employer
**This publication contains:** **Form 4070A, Employees Daily Record of** Tips **Form 4070, Employees Report of Tips to** Employer
### This publication contains:
**Form 4070A, Employees Daily Record of** Tips **Form 4070, Employees Report of Tips to** Employer
For the period
@@ -74,7 +76,7 @@ forms simpler, we would be happy to hear from you. You can write to the Tax Form
**Unreported Tips.—If you received tips of $20 or** more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you must use Form 1040 and Form 4137, Social Security and Medicare Tax on Unreported Tip Income, to report them. You may not use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act cannot use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—Get Pub. 531, Reporting** Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—If you do not keep a daily** record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
**Instructions (continued)**
### Instructions (continued)
Use this space to total your tips for the year
+2 -6
View File
@@ -6,9 +6,7 @@
8 4 Z E L L / L U R I E R E A L E S T A T E C E N T E R
**Table I: Cap rate correlations**
**Cap Rate Correlation With:*** **BBB Corp** **10-Year Bond Yield S&P Dividend** **Treasury (10-15 yr) Yield** Multifamily 0.187 0.771 0.068 Industrial-0.221 0.748-0.307 CBD Office-0.449 0.694-0.458 Retail-0.181 0.649-02.58
**Table I: Cap rate correlations** **Cap Rate Correlation With:*** **BBB Corp** **10-Year Bond Yield S&P Dividend** **Treasury (10-15 yr) Yield** Multifamily 0.187 0.771 0.068 Industrial-0.221 0.748-0.307 CBD Office-0.449 0.694-0.458 Retail-0.181 0.649-02.58
* Based on 25 years of data for the 10-yrT & S&P DivYld; and 14 years for BBB.
**Figure 1:** NCREIF cap rates vs. 10-yearTreasury
@@ -34,9 +32,7 @@ R E V I E W 8 5
1982 1986 1990 1994 1998 2002 2006
**Table II: Correlationsofspreadsbypropertytype**
**Correlation of Cap Rate Spreads Over Treasury** **Multifamily Industrial CBD Office**
**Table II: Correlationsofspreadsbypropertytype** **Correlation of Cap Rate Spreads Over Treasury** **Multifamily Industrial CBD Office**
||Multifamily|Industrial|CBD Office|
|---|---|---|---|
+55 -42
View File
@@ -13,8 +13,10 @@ IDENTIFICATION NUMBER OF CORPORATION] ELECTS TO BE TREATED AS A COMPONENT MEMBER
(v) Election-- (A) Election filed. An election filed under paragraph (d)(2)(iv) of
this section is irrevocable and effective until paragraph (d)(2)(ii) or (iii) of §1.1563-3 applies or until a change in the stock ownership of the corporation results in
|termination of membership in the controlled group in which such corporation has been included. (B) Election not filed.|In the event no election is filed in accordance with the|
|termination of membership in the controlled group in which such corporation has||
|---|---|
|been included.||
|(B) Election not filed.|In the event no election is filed in accordance with the|
|provisions of paragraph (d)(2)(iv) of this section, then the Internal Revenue Service||
|will determine the group in which such corporation is to be included. Such||
|determination will be binding for all subsequent years unless the corporation files a||
@@ -52,9 +54,14 @@ company other than a life insurance company shall make a return on Form 1120PC.
annual statement (or a pro forma annual statement), including the underwriting and investment exhibit for the year covered by such return.
||(3) Foreign insurance companies. The provisions of paragraphs (c)(1) and|
|---|---|
||(c)(2) of this section concerning the returns and statements of insurance companies subject to tax under section 801 or section 831 also apply to foreign insurance companies subject to tax under those sections, except that the copy of the annual statement required to be submitted with the return shall, in the case of a foreign insurance company that is not required to file an annual statement, be a copy of the pro forma annual statement relating to the United States business of such company. (4) Exception for insurance companies filing their Federal income tax returns electronically. If an insurance company described in paragraph (c)(1), (c)(2), or (c)(3) of this section files its Federal income tax return electronically, it should not include on or with such return its annual statement (or pro forma annual statement), or any portion thereof. Such statement must be available at all times for inspection by authorized Internal Revenue Service officers or employees and retained for so long as such statements may be material in the administration of any internal revenue law. See §1.6001-1(e). (5) Definition. For purposes of this section, the term annual statement means the annual statement, the form of which is approved by the National Association of Insurance Commissioners (NAIC), which is filed by an insurance company for the year with the insurance departments of States, Territories, and the District of|
(3) Foreign insurance companies. The provisions of paragraphs (c)(1) and
(c)(2) of this section concerning the returns and statements of insurance companies subject to tax under section 801 or section 831 also apply to foreign insurance companies subject to tax under those sections, except that the copy of the annual statement required to be submitted with the return shall, in the case of a foreign insurance company that is not required to file an annual statement, be a copy of the pro forma annual statement relating to the United States business of such company.
(4) Exception for insurance companies filing their Federal income tax returns
electronically. If an insurance company described in paragraph (c)(1), (c)(2), or
(c)(3) of this section files its Federal income tax return electronically, it should not include on or with such return its annual statement (or pro forma annual statement), or any portion thereof. Such statement must be available at all times for inspection by authorized Internal Revenue Service officers or employees and retained for so long as such statements may be material in the administration of any internal revenue law. See §1.6001-1(e).
(5) Definition. For purposes of this section, the term annual statement means
the annual statement, the form of which is approved by the National Association of Insurance Commissioners (NAIC), which is filed by an insurance company for the year with the insurance departments of States, Territories, and the District of
Columbia. The term annual statement also includes a pro forma annual statement if the insurance company is not required to file the NAIC annual statement.
@@ -101,45 +108,32 @@ section and paragraph
(c)(4), and (c)(5) of this section, and paragraph
(c)(2) of §1.382-8T
||§1.382-8(g), Example|
|---|---|
||The first sentence of §1.382-8(g), Example §1.382-8(g), Example §1.382-8(g), Example|
(2)(c)
(2)(e)
(3)(b)
(3)(c)(1)(B)
||The second sentence of|
|---|---|
||§1.382-8(g), Example The second sentence of §1.382-8(g), Example The first sentence of §1.1502-32(b)(4)(v)(A) The first sentence of §1.1502-32(b)(4)(v)(B)|
(4)(c)
(5)(c)
|The fifth sentence of|paragraph (c) of this|paragraphs (c)(1), (c)(3),|
|---|---|---|
|§1.382-8(f)|section|(c)(4), and (c)(5) of this section, and paragraph (c)(2) of §1.382-8T|
|§1.382-8(g), Example|paragraph (c) of this|paragraphs (c)(1), (c)(3), section, and paragraph (c)(2) of §1.382-8T|
|The second sentence of|paragraph (c) of this|paragraphs (c)(1), (c)(3),|
|§1.382-8(g), Example|section|(c)(4), and (c)(5) of this|
|§1.382-8(g), Example|section paragraph (c)(2) of this section paragraph (c)(2) of this section paragraph (c)(2) of this section paragraphs (c)(1) and (2) of this section paragraph (c)(2) of this section paragraph (c)(2) of this section paragraph (b)(4)(iv) of this section paragraph (b)(4)(iv) of this section|(c)(4), and (c)(5) of this (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (c)(1) of this section and paragraph (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (b)(4)(iv) of §1.1502-32T paragraph (b)(4)(iv) of §1.1502-32T|
(1)(b)(2) section (c)(4), and (c)(5) of this
(1)(c) section, and paragraph
(c)(2) of §1.382-8T
|§1.382-8(g), Example|paragraph (c)(2) of this|paragraph (c)(2) of|
|---|---|---|
|The first sentence of|paragraph (c)(2) of this|paragraph (c)(2) of|
|§1.382-8(g), Example|section|§1.382-8T|
|§1.382-8(g), Example|paragraph (c)(2) of this|paragraph (c)(2) of|
|§1.382-8(g), Example|paragraphs (c)(1) and (2)|paragraph (c)(1) of this|
(2)(c) section §1.382-8T
(2)(e)
(3)(b) section §1.382-8T
(3)(c)(1)(B) of this section section and paragraph
(c)(2) of §1.382-8T
|The second sentence of|paragraph (c)(2) of this|paragraph (c)(2) of|
|---|---|---|
|§1.382-8(g), Example|section|§1.382-8T|
|The second sentence of|paragraph (c)(2) of this|paragraph (c)(2) of|
|§1.382-8(g), Example|section|§1.382-8T|
|The first sentence of|paragraph (b)(4)(iv) of|paragraph (b)(4)(iv) of|
|§1.1502-32(b)(4)(v)(A)|this section|§1.1502-32T|
|The first sentence of|paragraph (b)(4)(iv) of|paragraph (b)(4)(iv) of|
|§1.1502-32(b)(4)(v)(B)|this section|§1.1502-32T|
(4)(c)
(5)(c)
|§1.1502-35(c)(4)(ii)(B)|§1.1502-76(b)(2)(ii)(D)|§1.1502-76T(b)(2)(ii)(D)|
|---|---|---|
|§1.1502-76(b)(2)(ii)(A)(2)|paragraph (b)(2)(ii)(D) of this section|paragraph (b)(2)(ii)(D) of §1.1502-76T|
@@ -168,14 +162,33 @@ section and paragraph
|§1.6043-2(a)|or 1.1081-11|3T(a), or §1.1081-11T|
|The first sentence of §301.6011-5T(a) (twice)|§1.6012-2|paragraphs (a), (b) and (d) through (j) of §1.6012- 2, and paragraph (c) of §1.6012-2T|
|||PART 602--OMB CONTROL NUMBERS UNDER THE PAPERWORK||
|---|---|---|---|
||REDUCTION ACT Authority: 26 U.S.C. 7805. 1. The following entries to the table are removed: §602.101 OMB Control numbers.|Par. 54. The authority citation for part 602 continues to read as follows: Par. 55. In §602.101, paragraph (b) is amended to read as follows:||
|* * * * *|(b) * * * CFR part or section where identified or described||Current OMB control No.|
|* * * * *|1.332-6………………………………………………………………….|1.382-11……………………………………………………………….. 1545-2019 1.351-3…………………………………………………………………. 1545-2019 1.355-5…………………………………………………………………. 1545-2019 1.368-3…………………………………………………………………. 1545-2019 1.1081-11………………………………………………………………. 1545-2019|1545-2019|
|* * * * *|§602.101 OMB Control numbers.|______________________________________________________________ 2. The following entries are added in numerical order to the table:||
|* * * * *|(b) * * * CFR part or section where identified or described||Current OMB control No.|
|* * * * *|1.302-2T………………………………………………………………… 1545 1.302-4T………………………………………………………………… 1545||-2019 -2019|
PART 602--OMB CONTROL NUMBERS UNDER THE PAPERWORK REDUCTION ACT Par. 54. The authority citation for part 602 continues to read as follows: Authority: 26 U.S.C. 7805. Par. 55. In §602.101, paragraph (b) is amended to read as follows:
1. The following entries to the table are removed:
§602.101 OMB Control numbers.
* * * * *
(b) * * *
CFR part or section where Current OMB identified or described control No.
* * * * *
1.332-6…………………………………………………………………. 1545-2019
1.382-11……………………………………………………………….. 1545-2019
1.351-3…………………………………………………………………. 1545-2019
1.355-5…………………………………………………………………. 1545-2019
1.368-3…………………………………………………………………. 1545-2019
1.1081-11………………………………………………………………. 1545-2019
* * * * * **______________________________________________________________**
2. The following entries are added in numerical order to the table:
§602.101 OMB Control numbers.
* * * * *
(b) * * *
CFR part or section where Current OMB identified or described control No.
* * * * *
1.302-2T………………………………………………………………… 1545-2019
1.302-4T………………………………………………………………… 1545-2019
|1.331-1T………………………………………………………………… 1545|-2019|
|---|---|
+14 -17
View File
@@ -1,8 +1,8 @@
**Technical Information**
##### Technical Information
## l T-12 SI
DuPont Fluorochemicals
##### DuPont Fluorochemicals
#### Thermodynamic Properties
@@ -20,25 +20,22 @@ Tables of the thermodynamic **Units** properties of R-12 have been developed and
S.A., Lemmon, E.W., and Peskin, Vf = Fluid (liquid) specific volume
A.P., NIST Standard Reference in cubic meters per kilogram Database 23, NIST thermodynamic and transport properties of Vg = Vapour (gas) specific volume refrigerants and refrigerant in cubic meters per kilogram mixtures REFPROP version 6.01, Standard Reference Data Program, df and dg = Fluid and Vapour National Institute of Standards and (respectively) densities in Technology, 1998). kilograms per cubic meter
H = Enthalpy (kJ/kg)
##### H = Enthalpy (kJ/kg)
S = Entropy (kJ/kg.K)
##### S = Entropy (kJ/kg.K)
**Physical Properties**
##### Physical Properties
Chemical Formula CCl2F2
|Chemical Formula|CCl2F2|
|---|---|
|Molecular mass|120.91|
|Boiling Point At one atmosphere|-29.75°C|
|Critical Temperature|111.97°C|
|Critical Pressure|4136 kPa|
|Critical Density|565.0 kg/m|
|Critical Volume|0.0018 m|
Molecular mass 120.91
Boiling Point-29.75°C At one atmosphere
Critical Temperature 111.97°C
Critical Pressure 4136 kPa
3 Critical Density 565.0 kg/m
Critical Volume 0.0018 m /kg
/kg
l
+403
View File
@@ -0,0 +1,403 @@
"""Tests for the pdf_inspector Python bindings."""
import os
import pytest
import pdf_inspector
FIXTURES_DIR = os.path.join(os.path.dirname(__file__), "fixtures")
def fixture_path(name: str) -> str:
return os.path.join(FIXTURES_DIR, name)
def fixture_bytes(name: str) -> bytes:
with open(fixture_path(name), "rb") as f:
return f.read()
# ---------------------------------------------------------------------------
# process_pdf
# ---------------------------------------------------------------------------
class TestProcessPdf:
def test_basic(self):
result = pdf_inspector.process_pdf(fixture_path("thermo-freon12.pdf"))
assert result.pdf_type == "text_based"
assert result.page_count == 3
assert result.confidence > 0.0
assert result.markdown is not None
assert len(result.markdown) > 0
def test_result_repr(self):
result = pdf_inspector.process_pdf(fixture_path("thermo-freon12.pdf"))
r = repr(result)
assert "PdfResult" in r
assert "text_based" in r
def test_with_pages(self):
result = pdf_inspector.process_pdf(
fixture_path("thermo-freon12.pdf"), pages=[1]
)
assert result.page_count == 3 # total pages in doc
assert result.markdown is not None
def test_result_fields(self):
result = pdf_inspector.process_pdf(fixture_path("thermo-freon12.pdf"))
# All fields should be accessible
assert isinstance(result.pdf_type, str)
assert isinstance(result.page_count, int)
assert isinstance(result.processing_time_ms, int)
assert isinstance(result.pages_needing_ocr, list)
assert isinstance(result.confidence, float)
assert isinstance(result.is_complex_layout, bool)
assert isinstance(result.pages_with_tables, list)
assert isinstance(result.pages_with_columns, list)
assert isinstance(result.has_encoding_issues, bool)
# title can be None or str
assert result.title is None or isinstance(result.title, str)
# ---------------------------------------------------------------------------
# process_pdf_bytes
# ---------------------------------------------------------------------------
class TestProcessPdfBytes:
def test_basic(self):
data = fixture_bytes("thermo-freon12.pdf")
result = pdf_inspector.process_pdf_bytes(data)
assert result.pdf_type == "text_based"
assert result.markdown is not None
def test_with_pages(self):
data = fixture_bytes("thermo-freon12.pdf")
result = pdf_inspector.process_pdf_bytes(data, pages=[1, 2])
assert result.markdown is not None
# ---------------------------------------------------------------------------
# detect_pdf / detect_pdf_bytes
# ---------------------------------------------------------------------------
class TestDetectPdf:
def test_detect_file(self):
result = pdf_inspector.detect_pdf(fixture_path("thermo-freon12.pdf"))
assert result.pdf_type == "text_based"
assert result.markdown is None # detect only — no markdown
assert result.page_count == 3
def test_detect_bytes(self):
data = fixture_bytes("thermo-freon12.pdf")
result = pdf_inspector.detect_pdf_bytes(data)
assert result.pdf_type == "text_based"
assert result.markdown is None
# ---------------------------------------------------------------------------
# classify_pdf / classify_pdf_bytes
# ---------------------------------------------------------------------------
class TestClassifyPdf:
def test_classify_file(self):
result = pdf_inspector.classify_pdf(fixture_path("thermo-freon12.pdf"))
assert result.pdf_type == "text_based"
assert result.page_count == 3
assert result.confidence > 0.0
assert isinstance(result.pages_needing_ocr, list)
def test_classify_bytes(self):
data = fixture_bytes("thermo-freon12.pdf")
result = pdf_inspector.classify_pdf_bytes(data)
assert result.pdf_type == "text_based"
assert result.page_count == 3
assert result.confidence > 0.0
def test_classify_repr(self):
result = pdf_inspector.classify_pdf(fixture_path("thermo-freon12.pdf"))
r = repr(result)
assert "PdfClassification" in r
assert "text_based" in r
def test_classify_fields(self):
result = pdf_inspector.classify_pdf(fixture_path("thermo-freon12.pdf"))
assert isinstance(result.pdf_type, str)
assert isinstance(result.page_count, int)
assert isinstance(result.pages_needing_ocr, list)
assert isinstance(result.confidence, float)
# ---------------------------------------------------------------------------
# extract_text / extract_text_bytes
# ---------------------------------------------------------------------------
class TestExtractText:
def test_basic(self):
text = pdf_inspector.extract_text(fixture_path("thermo-freon12.pdf"))
assert isinstance(text, str)
assert len(text) > 0
def test_bytes(self):
data = fixture_bytes("thermo-freon12.pdf")
text = pdf_inspector.extract_text_bytes(data)
assert isinstance(text, str)
assert len(text) > 0
def test_bytes_matches_file(self):
text_file = pdf_inspector.extract_text(fixture_path("thermo-freon12.pdf"))
text_bytes = pdf_inspector.extract_text_bytes(fixture_bytes("thermo-freon12.pdf"))
assert text_file == text_bytes
# ---------------------------------------------------------------------------
# extract_text_with_positions / extract_text_with_positions_bytes
# ---------------------------------------------------------------------------
class TestExtractTextWithPositions:
def test_basic(self):
items = pdf_inspector.extract_text_with_positions(
fixture_path("thermo-freon12.pdf")
)
assert len(items) > 0
item = items[0]
assert isinstance(item.text, str)
assert isinstance(item.x, float)
assert isinstance(item.y, float)
assert isinstance(item.width, float)
assert isinstance(item.height, float)
assert isinstance(item.font, str)
assert isinstance(item.font_size, float)
assert isinstance(item.page, int)
assert isinstance(item.is_bold, bool)
assert isinstance(item.is_italic, bool)
assert isinstance(item.item_type, str)
def test_with_pages(self):
items = pdf_inspector.extract_text_with_positions(
fixture_path("thermo-freon12.pdf"), pages=[1]
)
assert len(items) > 0
assert all(item.page == 1 for item in items)
def test_repr(self):
items = pdf_inspector.extract_text_with_positions(
fixture_path("thermo-freon12.pdf")
)
r = repr(items[0])
assert "TextItem" in r
def test_bytes(self):
data = fixture_bytes("thermo-freon12.pdf")
items = pdf_inspector.extract_text_with_positions_bytes(data)
assert len(items) > 0
assert isinstance(items[0].text, str)
def test_bytes_with_pages(self):
data = fixture_bytes("thermo-freon12.pdf")
items = pdf_inspector.extract_text_with_positions_bytes(data, pages=[1])
assert len(items) > 0
assert all(item.page == 1 for item in items)
# ---------------------------------------------------------------------------
# extract_text_in_regions / extract_text_in_regions_bytes
# ---------------------------------------------------------------------------
class TestExtractTextInRegions:
def test_file(self):
results = pdf_inspector.extract_text_in_regions(
fixture_path("thermo-freon12.pdf"),
[(0, [[0.0, 0.0, 600.0, 100.0]])],
)
assert len(results) == 1
assert results[0].page == 0
assert len(results[0].regions) == 1
assert isinstance(results[0].regions[0].text, str)
assert isinstance(results[0].regions[0].needs_ocr, bool)
def test_bytes(self):
data = fixture_bytes("thermo-freon12.pdf")
results = pdf_inspector.extract_text_in_regions_bytes(
data,
[(0, [[0.0, 0.0, 600.0, 100.0]])],
)
assert len(results) == 1
assert results[0].page == 0
assert len(results[0].regions) == 1
assert isinstance(results[0].regions[0].text, str)
def test_repr(self):
results = pdf_inspector.extract_text_in_regions(
fixture_path("thermo-freon12.pdf"),
[(0, [[0.0, 0.0, 600.0, 100.0]])],
)
r = repr(results[0])
assert "PageRegionTexts" in r
r2 = repr(results[0].regions[0])
assert "RegionText" in r2
def test_multiple_regions(self):
results = pdf_inspector.extract_text_in_regions(
fixture_path("thermo-freon12.pdf"),
[(0, [[0.0, 0.0, 300.0, 100.0], [300.0, 0.0, 600.0, 100.0]])],
)
assert len(results) == 1
assert len(results[0].regions) == 2
def test_multiple_pages(self):
results = pdf_inspector.extract_text_in_regions(
fixture_path("thermo-freon12.pdf"),
[
(0, [[0.0, 0.0, 600.0, 100.0]]),
(1, [[0.0, 0.0, 600.0, 100.0]]),
],
)
assert len(results) == 2
assert results[0].page == 0
assert results[1].page == 1
def test_malformed_region_raises_value_error(self):
with pytest.raises(ValueError, match="Invalid region"):
pdf_inspector.extract_text_in_regions(
fixture_path("thermo-freon12.pdf"),
[(0, [[0.0, 0.0, 600.0]])],
)
# ---------------------------------------------------------------------------
# extract_pages_markdown / extract_pages_markdown_bytes
# ---------------------------------------------------------------------------
class TestExtractPagesMarkdown:
def test_default_returns_all_pages(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf")
)
assert len(result.pages) == 3
assert [p.page for p in result.pages] == [0, 1, 2]
assert all(isinstance(p.markdown, str) for p in result.pages)
def test_bytes_default_returns_all_pages(self):
data = fixture_bytes("thermo-freon12.pdf")
result = pdf_inspector.extract_pages_markdown_bytes(data)
assert len(result.pages) == 3
def test_selected_pages_preserve_order(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf"), pages=[2, 0]
)
assert [p.page for p in result.pages] == [2, 0]
def test_bytes_selected_pages_preserve_order(self):
data = fixture_bytes("thermo-freon12.pdf")
result = pdf_inspector.extract_pages_markdown_bytes(data, pages=[1])
assert len(result.pages) == 1
assert result.pages[0].page == 1
def test_page_fields(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf"), pages=[0]
)
page = result.pages[0]
assert isinstance(page.page, int)
assert isinstance(page.markdown, str)
assert isinstance(page.needs_ocr, bool)
assert not page.needs_ocr # text-based fixture
assert len(page.markdown) > 0
def test_result_fields(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf")
)
assert isinstance(result.pages, list)
assert isinstance(result.pages_with_tables, list)
assert isinstance(result.pages_with_columns, list)
assert isinstance(result.pages_needing_ocr, list)
assert isinstance(result.is_complex, bool)
def test_out_of_range_page_marks_needs_ocr(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf"), pages=[9999]
)
assert len(result.pages) == 1
assert result.pages[0].needs_ocr
assert result.pages[0].markdown == ""
def test_repr(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf"), pages=[0]
)
assert "PagesExtractionResult" in repr(result)
assert "PageMarkdown" in repr(result.pages[0])
def test_not_a_pdf(self):
with pytest.raises(ValueError):
pdf_inspector.extract_pages_markdown_bytes(b"not a pdf")
# ---------------------------------------------------------------------------
# Error handling
# ---------------------------------------------------------------------------
class TestErrors:
def test_nonexistent_file(self):
with pytest.raises(ValueError):
pdf_inspector.process_pdf("/nonexistent/file.pdf")
def test_not_a_pdf(self):
with pytest.raises(ValueError):
pdf_inspector.process_pdf_bytes(b"this is not a pdf")
def test_empty_bytes(self):
with pytest.raises(ValueError):
pdf_inspector.process_pdf_bytes(b"")
def test_classify_not_a_pdf(self):
with pytest.raises(ValueError):
pdf_inspector.classify_pdf_bytes(b"not a pdf")
def test_classify_nonexistent(self):
with pytest.raises((ValueError, OSError)):
pdf_inspector.classify_pdf("/nonexistent/file.pdf")
def test_extract_text_bytes_not_a_pdf(self):
with pytest.raises(ValueError):
pdf_inspector.extract_text_bytes(b"not a pdf")
def test_regions_not_a_pdf(self):
with pytest.raises(ValueError):
pdf_inspector.extract_text_in_regions_bytes(
b"not a pdf", [(0, [[0.0, 0.0, 100.0, 100.0]])]
)
# ---------------------------------------------------------------------------
# Multiple fixtures
# ---------------------------------------------------------------------------
class TestMultipleFixtures:
"""Run basic processing on all available test fixtures."""
@pytest.mark.parametrize(
"filename",
[f for f in os.listdir(FIXTURES_DIR) if f.endswith(".pdf")],
)
def test_process_all_fixtures(self, filename):
result = pdf_inspector.process_pdf(fixture_path(filename))
assert result.pdf_type in (
"text_based",
"scanned",
"image_based",
"mixed",
)
assert result.page_count > 0
assert result.confidence >= 0.0