Files
pdf-inspector/napi/package.json
T
Abimael MartellandClaude Opus 4.7 5845cb62f5 Recover orphan text after band normalization (TSR stage 2), v1.6.2
PR #62 (1.6.1) closed the cell-bleed regression by clamping SLANet's
loose cell bboxes into non-overlapping row/column bands and tightening
text membership to center-containment OR >=60% overlap. That worked,
but exposed the opposite failure: legitimate native PDF text whose
center fell just outside the *clamped* cell bbox now had nowhere to go.

Two distinct failure modes observed against the FNBO branch-list PDF
after 1.6.1 deployed:

  Symptom A — header text positioned at the LEFT of a column whose band
  was derived from data-cell centers farther right. Header "Address"
  PDF text at x=331..375 fell outside the clamped col band starting
  at x=410. Strict membership rejected it (0% x-overlap, center
  outside).

  Symptom B — local SLANet row drift in col 0 over a 5-row stretch.
  Cell bboxes sat just above the actual branch-name text items
  (x-overlap 100% but y-overlap ~30-43%, below the 60% threshold).

Both share one root: post-normalization bboxes are too tight, and the
strict rule has no escape valve for legitimate edge text.

Fix: add a stage-2 orphan-recovery pass after the strict fill. Items
that NO cell claimed in stage 1 get re-assigned to their nearest
*empty* cell, distance-capped by `(median_col_width, median_row_height)`
so a far-orphan figure title can't get pulled into a faraway empty
cell. Stage 2 only fills empties — never overwrites stage 1 — so the
cell-bleed case PR #62 closed cannot regress.

Three new lib tests cover the bug shapes:
  - stage2_recovers_left_aligned_header_text_outside_data_band
    (Symptom A: header text left-of-band, data cells already filled,
     stage 2 fills only the header)
  - stage2_recovers_y_shifted_col0_in_consecutive_rows
    (Symptom B: 3 col-0 cells shifted vs text, all 3 recovered)
  - stage2_does_not_overwrite_filled_cells_or_admit_far_orphans
    (cap rejects far figure titles; pre-filled cells untouched)

Plus tests for the cap-derivation helper:
  - tsr_assignment_caps_uses_median_geometry
  - tsr_assignment_caps_floor_protects_degenerate_input

The existing dense-overlapping-rows regression test (added in #62) still
passes, confirming no regression on the cell-bleed case.

Test results: 414 lib + 115 integration + 2 doctests pass. clippy + fmt
clean.

Bump @firecrawl/pdf-inspector to 1.6.2 (patch — fixes a regression
introduced by 1.6.1; no API changes).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 21:34:25 -07:00

56 lines
1.2 KiB
JSON

{
"name": "@firecrawl/pdf-inspector",
"version": "1.6.2",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
"bin": {
"pdf-inspector": "bin/pdf-inspector.mjs"
},
"license": "MIT",
"keywords": [
"pdf",
"pdf-extraction",
"pdf-parser",
"text-extraction",
"ocr",
"pdf-classification",
"napi",
"rust",
"firecrawl"
],
"files": [
"index.js",
"index.d.ts",
"*.node",
"bin/",
"README.md"
],
"repository": {
"type": "git",
"url": "https://github.com/firecrawl/pdf-inspector"
},
"homepage": "https://github.com/firecrawl/pdf-inspector",
"publishConfig": {
"access": "public"
},
"napi": {
"binaryName": "pdf-inspector",
"targets": [
"x86_64-unknown-linux-gnu",
"aarch64-apple-darwin",
"x86_64-pc-windows-msvc"
],
"package": {
"name": "@firecrawl/pdf-inspector-js"
}
},
"scripts": {
"build": "napi build --platform --release",
"build:debug": "napi build --platform"
},
"devDependencies": {
"@napi-rs/cli": "^3.4.1"
}
}