Closes the residual run-on-cell pattern observed on FNBO branch list
after 1.6.2 deployed:
|Shawnee Blue Valley Parkway|Kansas|6301 Pflumm, ...|0523.04|
|Sonoma Plaza|Kansas Kansas|<addr1> <addr2>|0531.05|
|Mitchell Woonsocket|South Dakota|<addr1>|9628.01|
Cause: when col-N cells across multiple consecutive rows are y-shifted
the same way (a local SLANet drift), the stage 2 orphan pass sees
multiple orphans qualifying for the same nearest empty cell. The cell
gets all of them appended in order, producing "RowA-text RowB-text".
Fix: track the y-coordinate of the first orphan that lands in each
cell. Subsequent orphans only join that cell if their y is within
half-a-row-height of the first orphan's y (same line). Cross-line
orphans skip that cell and look for the next-nearest empty cell on
their own line.
Same-line slack preserves multi-token branch names like
"Blue Valley Parkway" (3 PDF text items at the same y) — all three
stack into the same cell. Cross-row stacking is what gets rejected.
Two new tests:
- stage2_rejects_cross_line_stacking_into_same_cell
Two orphans on different rows, both equidistant to the same empty
cell. First wins; second routes to its own row's cell.
- stage2_allows_same_line_orphans_to_stack_into_one_cell
Three same-line orphans (multi-token branch name) all land in the
same empty cell, joined by spaces.
The existing four stage-2 tests + the cell-bleed regression test from
PR #62 + #63 all pass: 416 lib + 115 integration + 2 doctests.
clippy + fmt clean.
Bump @firecrawl/pdf-inspector to 1.6.3 (patch — refines 1.6.2; no API
changes).
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
PR #62 (1.6.1) closed the cell-bleed regression by clamping SLANet's
loose cell bboxes into non-overlapping row/column bands and tightening
text membership to center-containment OR >=60% overlap. That worked,
but exposed the opposite failure: legitimate native PDF text whose
center fell just outside the *clamped* cell bbox now had nowhere to go.
Two distinct failure modes observed against the FNBO branch-list PDF
after 1.6.1 deployed:
Symptom A — header text positioned at the LEFT of a column whose band
was derived from data-cell centers farther right. Header "Address"
PDF text at x=331..375 fell outside the clamped col band starting
at x=410. Strict membership rejected it (0% x-overlap, center
outside).
Symptom B — local SLANet row drift in col 0 over a 5-row stretch.
Cell bboxes sat just above the actual branch-name text items
(x-overlap 100% but y-overlap ~30-43%, below the 60% threshold).
Both share one root: post-normalization bboxes are too tight, and the
strict rule has no escape valve for legitimate edge text.
Fix: add a stage-2 orphan-recovery pass after the strict fill. Items
that NO cell claimed in stage 1 get re-assigned to their nearest
*empty* cell, distance-capped by `(median_col_width, median_row_height)`
so a far-orphan figure title can't get pulled into a faraway empty
cell. Stage 2 only fills empties — never overwrites stage 1 — so the
cell-bleed case PR #62 closed cannot regress.
Three new lib tests cover the bug shapes:
- stage2_recovers_left_aligned_header_text_outside_data_band
(Symptom A: header text left-of-band, data cells already filled,
stage 2 fills only the header)
- stage2_recovers_y_shifted_col0_in_consecutive_rows
(Symptom B: 3 col-0 cells shifted vs text, all 3 recovered)
- stage2_does_not_overwrite_filled_cells_or_admit_far_orphans
(cap rejects far figure titles; pre-filled cells untouched)
Plus tests for the cap-derivation helper:
- tsr_assignment_caps_uses_median_geometry
- tsr_assignment_caps_floor_protects_degenerate_input
The existing dense-overlapping-rows regression test (added in #62) still
passes, confirming no regression on the cell-bleed case.
Test results: 414 lib + 115 integration + 2 doctests pass. clippy + fmt
clean.
Bump @firecrawl/pdf-inspector to 1.6.2 (patch — fixes a regression
introduced by 1.6.1; no API changes).
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: TSR-aware table extraction (extract_tables_with_structure_mem)
New public function that consumes raw structure-recovery output (HTML
structure tokens + per-cell bboxes from a model like SLANet) and assembles
markdown tables by pulling cell text from the native PDF — no OCR, no
geometry inference.
Why: the existing extract_tables_in_regions_mem infers grid geometry from
text positions only and can't distinguish merged cells from multiple narrow
columns. Pairing structure recovery from a layout/TSR model with native
PDF text gets perfect text quality with proper row/col/span structure.
- New module src/tables/structured.rs: token state machine, polygon→AABB,
crop-px→page-pt, rowspan/colspan-aware cell layout, markdown emitter.
Accepts both 4-element rects and 8-element 4-corner polygons.
- New public extract_tables_with_structure_mem in src/lib.rs that reuses
extract_page_text_items, region_overlaps_item, and the shared region
text-collection helper. No existing public function modified.
- napi binding extractTablesWithStructure mirroring the existing
extractTablesInRegions shape (f64 in JS → f32 internally).
- 14 unit tests + 5 integration tests, including a real-PDF gold-standard
match against bits_pilani_feedback.pdf.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* TSR follow-ups: header-aware separator, cells API, v1.6.0
- cells_to_markdown emits the separator after the LAST row that contains
is_header=true cells, falling back to "after row 0" when no header is
flagged. Multi-row theads now render correctly. Three new unit tests
cover: multi-row header, header not on row 0, no headers (fallback).
- New public extract_tables_with_structure_cells_mem returning
Vec<Vec<StructuredCell>> so callers can drive their own rendering or
debug overlays without re-doing the parse + extraction. The markdown
variant now wraps it. The previously-unused page_pt_bbox field is
surfaced through this API.
- New napi binding extractTablesWithStructureCells + StructuredCellJs.
- Bump @firecrawl/pdf-inspector to 1.6.0.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix TSR cell text assignment for overlapping bboxes
Made-with: Cursor
* bump npm package version to 1.6.1
Made-with: Cursor
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: TSR-aware table extraction (extract_tables_with_structure_mem)
New public function that consumes raw structure-recovery output (HTML
structure tokens + per-cell bboxes from a model like SLANet) and assembles
markdown tables by pulling cell text from the native PDF — no OCR, no
geometry inference.
Why: the existing extract_tables_in_regions_mem infers grid geometry from
text positions only and can't distinguish merged cells from multiple narrow
columns. Pairing structure recovery from a layout/TSR model with native
PDF text gets perfect text quality with proper row/col/span structure.
- New module src/tables/structured.rs: token state machine, polygon→AABB,
crop-px→page-pt, rowspan/colspan-aware cell layout, markdown emitter.
Accepts both 4-element rects and 8-element 4-corner polygons.
- New public extract_tables_with_structure_mem in src/lib.rs that reuses
extract_page_text_items, region_overlaps_item, and the shared region
text-collection helper. No existing public function modified.
- napi binding extractTablesWithStructure mirroring the existing
extractTablesInRegions shape (f64 in JS → f32 internally).
- 14 unit tests + 5 integration tests, including a real-PDF gold-standard
match against bits_pilani_feedback.pdf.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* TSR follow-ups: header-aware separator, cells API, v1.6.0
- cells_to_markdown emits the separator after the LAST row that contains
is_header=true cells, falling back to "after row 0" when no header is
flagged. Multi-row theads now render correctly. Three new unit tests
cover: multi-row header, header not on row 0, no headers (fallback).
- New public extract_tables_with_structure_cells_mem returning
Vec<Vec<StructuredCell>> so callers can drive their own rendering or
debug overlays without re-doing the parse + extraction. The markdown
variant now wraps it. The previously-unused page_pt_bbox field is
surfaced through this API.
- New napi binding extractTablesWithStructureCells + StructuredCellJs.
- Bump @firecrawl/pdf-inspector to 1.6.0.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: expose per-page markdown extraction to Python and Node (#49)
Implements the feature requested in issue #49: a list-of-pages markdown output
from the Python API. Matching the existing project pattern, the feature lives
in the Rust core and is surfaced through every binding.
- Rust core: `extract_pages_markdown` (path) and `extract_pages_markdown_mem`
(bytes) now take `Option<&[u32]>` — `None` returns every page in document
order; a slice restricts and preserves caller order.
- Python: new `extract_pages_markdown(path, pages=None)` and
`extract_pages_markdown_bytes(data, pages=None)` functions plus
`PageMarkdown` / `PagesExtractionResult` classes; stub file updated.
- Node: `extractPagesMarkdown(buffer, pages?)` — `pages` is now optional.
- Tests: 2 new Rust integration tests, 9 new Python tests, 2 new Node
assertions. All 372 unit + 107 integration + 53 Python tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: bump version from 1.3.0 to 1.4.0
Minor bump for the new per-page markdown extraction API exposed through
the Python and Node bindings.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds x86_64-pc-windows-msvc target to the napi package and CI build
matrix so the published @firecrawl/pdf-inspector npm package ships
native Windows binaries alongside Linux and macOS.
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* feat: add CLI bin to npm package
Installing `firecrawl-pdf-inspector` now provides a `pdf-inspector` CLI command.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore: bump npm package version to 0.8.0
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore: bump npm package version to 1.0.0
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(detector): correct false flags for CID-encoded text and supplementary fonts (0.7.4)
The page classifier was over-aggressively flagging Mixed-PDF pages as
needing OCR in three distinct cases. Each is fixed at the root in
analyze_page_content / page_has_identity_h_no_tounicode / the
looks_like_scan check.
1. has_vector_text false positives on dense layouts
path_ops > text_ops*200 fired on pages with decorative paths
(column borders, dividers) alongside real selectable text. Added
a unique_alphanum_chars < 30 guard: real outlined-text pages have
very few unique alphanum chars (each glyph is a path), while
pages with real text + decorations have many.
2. Identity-H without ToUnicode flagged whole pages on supplementary fonts
page_has_identity_h_no_tounicode would flag a page if any single
Type0 font lacked ToUnicode and had no fallback CMap, even when
the page's actual text came from other decodable fonts (Type1
with ToUnicode, etc.). Rewrote to track both undecodable
Identity-H fonts AND other decodable fonts, only flagging when
no decodable text font is present.
3. CID-encoded text with ToUnicode misclassified as scan
looks_like_scan checked unique_alphanum_chars < 10 on raw string
operand bytes. CID-encoded fonts (Type0 with ToUnicode) emit
2-byte CID values that aren't ASCII alphanum, so the metric is
blind to them even when the text is fully decodable. Added a
has_decodable_text_fonts signal: when a page has decodable fonts
AND >= 10 text ops, the low alphanum count is treated as a CID
encoding artifact rather than evidence of a scan.
Validated against a broad PDF corpus:
- 6 known false-positive pages now correctly classified as text
- 22 previously-missed scan pages (cover/blank/photo) now correctly
flagged for OCR
- 0 regressions on truly-scanned PDFs (61/61 pages stay flagged)
- All 437 existing tests pass; clippy clean
Bumps NAPI package to 0.7.4.
* test(detector): add unit tests for the three classifier fixes
Adds 10 unit tests covering the heuristic changes:
- has_vector_text alphanum guard
- real text + decorative paths → not flagged
- true outlined glyphs (low alphanum) → still flagged
- page_has_identity_h_no_tounicode supplementary-font handling
- undecodable Identity-H + decodable Type1 → not flagged (new)
- undecodable Identity-H alone → still flagged (regression)
- page_has_decodable_text_fonts (new helper)
- Type1 → true
- Type0 with ToUnicode → true
- undecodable Identity-H only → false
- looks_like_scan with has_decodable_text_fonts override
- CID-encoded decodable text → not flagged as scan
- same metrics with no decodable fonts → still flagged
- decodable fonts but text_ops < 10 (page-number overlay) → still flagged
* fix(detector): make decodable-font checks usage-based and XObject-aware
Addresses two reviewer concerns on the previous heuristic fix:
P1 — resource-based check could create an inverse bug
page_has_identity_h_no_tounicode and page_has_decodable_text_fonts
iterated all fonts in the page Resources dict, including unused fonts.
A page whose actual text was rendered exclusively in an undecodable
Identity-H font but whose Resources also listed an unused decodable
Type1 would be wrongly unflagged.
Fix: parse Tf operator operands during content stream scanning to
collect the set of font names actually referenced. The font checks
now filter to only USED fonts via a new used_fonts_have_*
family of functions operating on (used_font_names, font_map).
P2 — checks didn't follow text into Form XObjects
analyze_page_content correctly recurses through Form XObjects via
scan_xobjects_in_resources, but the font checks only looked at the
page's top-level Resources/Font. Pages that render text through Form
XObjects (corporate templates, header/footer overlays) had their
XObject font resources missed entirely.
Fix: scan_xobjects_in_resources now propagates the used_font_names
set AND collects fonts from each Form XObject's own Resources into
the shared font_map. The usage-based check sees the full picture:
page-level fonts + every nested XObject's fonts, intersected with
fonts actually referenced by Tf operators anywhere in the content.
Implementation:
- New extract_font_name_before_tf helper (parses /Name immediately
preceding Tf).
- New FontInfo struct caches font properties per-name.
- New collect_fonts_from_resource_dict + new used_fonts_have_*
functions are pure filters over (used_names, font_map).
- analyze_page_content threads used_font_names + font_map through
page content scan and XObject recursion, then runs the new checks.
- Old resource-based functions kept as #[cfg(test)] for the existing
unit-test interface.
- Phase 3 uncached-page loop now goes through analyze_page_content
so it also gets the usage-based + XObject-aware behavior.
Tests added (8):
- extract_font_name_before_tf basic + long-name parsing
- scan_content_for_text_operators collects used font names
- P1 — unused decodable font in Resources doesn't save a page
whose used font is undecodable
- P1 — both fonts used → decodable font correctly prevents flag
- P2 — decodable font inside Form XObject correctly unflags
- P2 — undecodable font only in XObject still flags even with
unused decodable font at page level
- P2 — has_decodable_text_fonts populated from XObject fonts
Validation:
- 349 lib + 104 integration + 2 doc tests pass (was 341)
- cargo clippy --lib --bin detect-pdf -- -D warnings: clean
- External eval: 9/9 PDFs pass, 6/6 false positives resolved,
0 regressions, 61/61 scanned pages still correctly flagged
- No eval delta — confirms previous fix wasn't relying on the
resource-based bug for any of the eval PDFs
* fix(detector): scope font lookups by ObjectId + handle indirect Form Resources
Addresses two more reviewer findings on the previous decodable-font commit.
P1 — Resource-name scoping bug
The previous fix keyed used_font_names and font_map by raw resource
names like b"F1". PDF resource names are scoped to each resource
dictionary: a Form XObject can legally define its own /F1 that points
to a completely different font from the page's /F1. Because
collect_fonts_from_resource_dict skipped duplicates with
`if font_map.contains_key(name)`, the first definition won and later
Tf /F1 usages in different scopes resolved against the wrong font.
This could reintroduce both the undecodable-Identity-H false flag
and the decodable-CID false unflag depending on which side of the
collision happened to be inserted first.
Fix: switch the lookup mechanism from font names to font ObjectIds.
- font_map: HashMap<ObjectId, FontInfo> (was Vec<u8> keys)
- used_font_ids: HashSet<ObjectId> (was Vec<u8> names)
- new resolve_font_names_to_ids() runs immediately after each
content scan, against the resource dict in scope, to translate
the per-scope name set into ObjectIds.
Each Form XObject's content stream now resolves /F1 against THAT
XObject's own Resources, so name collisions are impossible by design.
Inline (no-ID) font dicts are skipped — extremely rare in practice
and have no stable key.
P2 — Indirect Form /Resources skipped
scan_xobjects_in_resources used `.as_dict()` on the Form's /Resources
entry, which returns None for indirect references. PDFs frequently
store /Resources as `X 0 R`, in which case font collection and
recursion were both skipped — even though the Tf usages inside the
XObject content had already been recorded.
Fix: handle Object::Reference(r) in addition to Object::Dictionary(d)
by resolving via doc.get_dictionary. Audited the rest of the file —
the other /Resources access points (analyze_page_images,
collect_images_from_resources) already handled both cases.
Tests added (4):
- P1 same-name-different-font (page undecodable, XObject decodable):
must NOT flag — XObject's text is decodable in its own scope.
- P1 inverse (page decodable, XObject undecodable, content uses
XObject /F1): MUST flag — undecodable text exists in real scope.
- P2 indirect Form /Resources: font discovery must still work when
/Resources is a `X 0 R` reference rather than inline.
- Combined regression: indirect Resources + name collision.
Validation:
- cargo test --release: 459 tests pass (353 lib + 104 integration + 2 doc)
- cargo clippy --lib --bin detect-pdf -- -D warnings: clean
- external eval (9 PDFs): 9/9 pass, 6/6 false positives resolved,
0 regressions, 61/61 truly-scanned pages still flagged
The behavior on the eval set is identical — confirms the correctness
fix isn't masking any change in classifier outcomes.
* fix(detector): respect resource shadowing when resolving page-content fonts
The previous ObjectId-based fix correctly scoped Form XObject fonts
but still violated PDF resource inheritance for page content. When a
page overrides /F1 from a parent /Pages node (different font dict for
the same name), get_page_resources returns the page's own /Resources
plus all ancestor /Resources dicts. The old code called
resolve_font_names_to_ids on each one and added every match to
used_font_ids — both font ObjectIds ended up in the used set even
though only the page's /F1 is actually visible to that page's content.
Per ISO 32000-1 §7.7.3.4, resource names are inherited with
shadowing semantics: the most-specific (deepest, closest to the page)
definition wins.
Fix:
- New lookup_font_id helper resolves a single name in a single dict.
- New resolve_with_shadowing iterates names, checking the page's own
/Resources first, then walking ancestors in most-specific-first
order (which is the order lopdf's get_page_resources returns).
First hit wins via a labeled `continue 'name` — subsequent
ancestors are skipped for that name.
- analyze_page_content's flat resolution loop replaced with one call
to resolve_with_shadowing.
Audit:
- XObject path is correct: each Form XObject already resolves names
against its OWN /Resources (XObjects don't inherit from page tree).
- font_map population is correct: keyed by ObjectId, so collecting
from all dicts builds the full available-fonts catalog. The bug
was only in the used-set resolution.
- Confirmed lopdf returns ancestors in most-specific-first order
(page → parent → grandparent → root), matching the shadowing
direction used here.
Tests added (3):
- page /F1 undecodable shadows parent's decodable /F1 → MUST flag
- page /F1 decodable shadows parent's undecodable /F1 → MUST NOT flag
- no override: page inherits parent's decodable /F1 → MUST NOT flag
Validation:
- cargo test --release: 462 tests pass (356 lib + 104 integration + 2 doc)
- cargo clippy --lib --bin detect-pdf -- -D warnings: clean
- external eval: 9/9, 0 regressions, 6/6 false positives resolved,
61/61 scanned pages still correctly flagged
The looks_like_scan check incorrectly used OR logic, causing any single
condition (image_count <= 1, text_ops < 50, alphanum < 10) to flag a page
as a scan. A real scan has ALL three: single full-page image AND low text
AND low alphanum. Text pages with one figure were falsely flagged for OCR.
Bump napi to 0.7.3.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: reduce false OCR recommendations for text PDFs with figure images
Two fixes in the detector:
1. Fix Tf operator parsing: some PDFs concatenate Tf directly with the
next operator (e.g. "25 Tf[<01>...") without whitespace. The scanner
now accepts [, (, <, / as valid followers, fixing font_changes being
reported as 0.
2. Distinguish text-with-figures from scanned-with-OCR: pages with
multiple images (image_count > 1) and strong text signals (text_ops
>= 50, alphanum >= 10) are recognized as text pages with figures,
not scanned templates. Scanned PDFs have exactly 1 full-page image.
This prevents academic papers, reports with charts, and similar PDFs
from being incorrectly classified as Mixed/OCR-needed when their text
is perfectly extractable.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore: bump napi version to 0.7.2
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: remove template image influence from page classification
Template images (large background/figure images) no longer affect
pages_needing_ocr. In the region-based pipeline, text regions are
extracted independently from image regions, and per-region needs_ocr
quality checks handle scanned-with-OCR garbage text.
Also makes the invisible text retry (for OCR text layers) trigger on
text quality rather than PDF type, so it works regardless of
classification.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Revert "fix: remove template image influence from page classification"
This reverts commit 100cbe5453.
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: improve heuristic table detection for numeric columns and multi-line headers
Two fixes for tables that have clean extractable text but fail heuristic
structure detection:
1. Numeric column merge pass (grid.rs): After initial X-position
clustering, adjacent clusters are merged when one is sparse (header
text) and the other is dense with >50% numeric items (data column).
Multi-line wrapped headers often land slightly offset from their
data column — the merge closes gaps within 1.5× the clustering
threshold. New is_numeric_text() helper matches decimals, percentages,
negative numbers, and comma-separated thousands.
2. Duplicate-header skip (detect_heuristic.rs): Spanning super-headers
like "First Degree | First Degree | Higher Degree" contain duplicate
cells that trigger looks_like_partial_table_ex rejection. Now skips
rows with duplicate cells when a better header candidate exists
within the next 3 rows (higher fill ratio or numeric cells).
Tested on BITS Pilani university report (430 pages, 314 table pages).
Page 4 (multi-line header + numeric data) previously returned
needs_ocr=true; now correctly detects the table structure.
Eval: 197 PDFs, zero regressions, all 104+ tests pass, zero clippy.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* bump version to 0.7.1
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Combine per-page markdown extraction with layout classification into a
single parse. extractPagesMarkdown now returns PagesExtractionResult with
pages_with_tables, pages_with_columns, pages_needing_ocr, and is_complex
alongside the per-page markdown — eliminating redundant PDF parses for
callers that need both.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* add extract_pages_markdown_mem for per-page markdown extraction
Enables hybrid OCR pipelines to skip GPU render+layout for simple text
pages by providing per-page markdown with needs_ocr flags. Font stats
are computed document-wide for consistent header detection.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* bump napi package version to 0.6.0
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace stringly-typed pdf_type and item_type fields with
#[napi(string_enum)] enums for proper TypeScript type checking.
Add link_url field to TextItem instead of encoding URL in the
item_type string.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Two changes that reduce false needsOcr rejections without hurting quality:
1. Per-region GID check instead of per-page blanket rejection.
Previously, if ANY font on the page used GID-encoded glyphs (common
in logos, decorative fonts), ALL table and text regions on that page
were forced to GPU OCR via needsOcr=true. Now the page-level bail is
removed; per-region text quality checks (is_garbage_text, is_cid_garbage,
detect_encoding_issues) catch actual GID corruption in the extracted
content. Tables whose text is clean pass through even if an unrelated
font elsewhere on the page is GID-encoded.
2. Relaxed looks_like_partial_table for layout-assisted extraction.
When the layout model already identified a region as a table (i.e.,
extract_tables_in_regions_mem), boundary-detection heuristics are
less necessary — we're not guessing "is this a table?" anymore, only
"can we extract it correctly?". Relaxations:
- Numeric first header cell accepted (e.g., year "2024")
- 1 empty header cell allowed in 3+ column tables (merged headers)
- Sparse first data row threshold relaxed from 33% to 50%
Paragraph detection and duplicate-header checks remain strict.
Eval: 196/196 pass (full regression suite), 91/91 Rust tests pass
including 7 new layout-assisted validation tests. Zero regressions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds a 5th failure-mode check to looks_like_partial_table: when the
heuristic mis-detects text-wrapped paragraph prose as a multi-column
table, cells in the same column tend to start with lowercase letters
or continuation punctuation (commas, closing quotes) — because they're
actually sentence fragments. Real tables almost never have most data
cells starting lowercase.
Trigger: ≥2 cols, ≥4 data rows, ≥60% of non-empty data cells start
with lowercase or continuation punctuation → return needs_ocr=true.
Caught in the eval as the next-largest failure mode after the 0.4.1 fix:
PDFs 088, 182, 090 — heuristic produced "tables" like:
|Approval is needed from the|Acquisitions of|
|Treasurer if the acquisition|residential and|
|constitutes a "significant|agricultural|
|action," including acquiring an|land by foreign|
Reading column 1 top-to-bottom: "Approval is needed from the Treasurer
if the acquisition constitutes a 'significant action,' including
acquiring an interest..." — a paragraph, not tabular data.
Tests: 2 new tests (the 088-style failure case + a real multi-word
table that must NOT be flagged). All 11 looks_like_partial_table tests
pass; 323 unit + 91 integration tests still green.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When the heuristic returns markdown that looks like a partial / mis-detected
table, set needs_ocr=true so the caller falls back to GPU OCR. Previously the
same cases returned the broken table with needs_ocr=false, which produced
real-world TEDS=0 scores in fire-pdf evals (heuristic-built table didn't
match ground truth structure at all, but caller had no signal to fall back).
Four failure modes detected, all observed in opendataloader-bench eval losses:
1. **Header looks like a data row** — first cell of header is a bare number
(e.g. `|2|...`), suggesting the actual header row was skipped. Real
headers almost never start with just a number.
2. **Empty header cells in a multi-column table** — ≥3 cols, ≥1 empty cell
in the header row. Indicates poor column boundary detection.
3. **Duplicate header cells** — same non-empty value appearing twice in the
header (e.g. "Administration|Administration"). Means a multi-line header
was collapsed wrong.
4. **Sparse first data row** — ≥3 cols and ≥1/3 of first-data-row cells are
empty. Multi-row headers in the source PDF get smashed into header +
sparse data row by the heuristic; this catches that.
Tests: 9 new unit tests in `looks_like_partial_table_tests` cover each
failure mode plus realistic non-failures (well-formed table, single-column
list, two-col with a single empty cell). All 91 existing tests still pass.
Bumps `napi/package.json` to 0.4.1 since this changes the function's return
behaviour for callers (some inputs that returned needs_ocr=false now return
true). The output text field is also cleared on the new fallback path so
callers don't accidentally use the broken markdown.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds a new function that takes a PDF buffer and page+bbox regions (same interface
as extractTextInRegions), runs heuristic table detection on items within each region,
and returns markdown pipe-tables. Falls back to needs_ocr=true when no table
structure is found or text quality is suspect.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When fire-pdf sends pre-segmented bboxes from the layout model,
pdf-inspector no longer runs column detection, stream-order heuristics,
or newspaper/tabular mode detection within the region. These heuristics
conflict with the layout model's decisions and cause wrong reading order.
Region extraction now simply: Y-sorts items, groups into lines, and
sorts within each line by X position. The heavy heuristics remain
available for standalone full-page extraction.
Eval showed pure OCR (0.2875 NED) beating native+heuristics (0.2916)
across all categories, especially multi-column (-0.08) and newspaper
(-0.16). This change should close that gap.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Rust panics in NAPI modules abort the Node.js process with no chance
to report errors. This wraps every exported function in catch_unwind,
converting panics into JS Error exceptions that can be caught and
reported to Sentry.
Buffer data is extracted to Vec<u8> before the catch_unwind boundary
to satisfy UnwindSafe requirements.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
index.js and index.d.ts are generated by napi-rs. Generate them
at publish time via `napi prepublish` instead of checking them in.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix sort panics on NaN values from bogus PDF font metrics
Replace all `partial_cmp(...).unwrap_or(Ordering::Equal)` and bare
`partial_cmp(...).unwrap()` with `total_cmp()` across the codebase.
`partial_cmp` returns `None` for NaN, and mapping that to `Equal`
violates total ordering: `a == NaN` and `NaN == b` but `a != b`.
Rust 1.81+ detects this and panics in sort_by. `total_cmp` handles
NaN deterministically (sorts to end) and guarantees total ordering.
The critical crash was in `extract_text_in_regions` (lib.rs:478)
where PDFs with bogus font ascent/descent values produced NaN in
text item coordinates, causing process abort via NAPI.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix missed partial_cmp in layout.rs and restore napi exports
- Convert two remaining b.y.partial_cmp(&a.y) calls to total_cmp
in group_single_column and column layout sorting
- Restore missing napi exports: detectPdf, extractText,
extractTextWithPositions, processPdf
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Drop platform-specific optional deps — ship all .node binaries in one
package (~4 MB total). Simplifies publishing and consumer install.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Skip napi prepublish/artifacts commands that require GitHub API auth.
Instead, create platform package.json files and copy binaries directly.
Add optionalDependencies to main package so npm/bun auto-selects the
right binary.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Both bindings now expose the same 6 function families: process, detect,
classify, extractText, extractTextWithPositions, and extractTextInRegions.
Bumps PyO3 from 0.22 to 0.25 for Python 3.14 support.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Move the napi bridge from fire-pdf into pdf-inspector as `napi/`.
Package name: @firecrawl/pdf-inspector-js, published to GitHub Packages
(npm.pkg.github.com) as a public package on v* tags.
Exposes two functions:
- classifyPdf(buffer) → type, page count, pages needing OCR
- extractTextInRegions(buffer, pageRegions) → per-region text with
needsOcr quality flag (GID fonts, garbage, encoding issues)
Includes publish workflow that builds linux-x64-gnu + darwin-arm64
binaries and publishes main + platform-specific packages.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>