PR #62 (1.6.1) closed the cell-bleed regression by clamping SLANet's loose cell bboxes into non-overlapping row/column bands and tightening text membership to center-containment OR >=60% overlap. That worked, but exposed the opposite failure: legitimate native PDF text whose center fell just outside the *clamped* cell bbox now had nowhere to go. Two distinct failure modes observed against the FNBO branch-list PDF after 1.6.1 deployed: Symptom A — header text positioned at the LEFT of a column whose band was derived from data-cell centers farther right. Header "Address" PDF text at x=331..375 fell outside the clamped col band starting at x=410. Strict membership rejected it (0% x-overlap, center outside). Symptom B — local SLANet row drift in col 0 over a 5-row stretch. Cell bboxes sat just above the actual branch-name text items (x-overlap 100% but y-overlap ~30-43%, below the 60% threshold). Both share one root: post-normalization bboxes are too tight, and the strict rule has no escape valve for legitimate edge text. Fix: add a stage-2 orphan-recovery pass after the strict fill. Items that NO cell claimed in stage 1 get re-assigned to their nearest *empty* cell, distance-capped by `(median_col_width, median_row_height)` so a far-orphan figure title can't get pulled into a faraway empty cell. Stage 2 only fills empties — never overwrites stage 1 — so the cell-bleed case PR #62 closed cannot regress. Three new lib tests cover the bug shapes: - stage2_recovers_left_aligned_header_text_outside_data_band (Symptom A: header text left-of-band, data cells already filled, stage 2 fills only the header) - stage2_recovers_y_shifted_col0_in_consecutive_rows (Symptom B: 3 col-0 cells shifted vs text, all 3 recovered) - stage2_does_not_overwrite_filled_cells_or_admit_far_orphans (cap rejects far figure titles; pre-filled cells untouched) Plus tests for the cap-derivation helper: - tsr_assignment_caps_uses_median_geometry - tsr_assignment_caps_floor_protects_degenerate_input The existing dense-overlapping-rows regression test (added in #62) still passes, confirming no regression on the cell-bleed case. Test results: 414 lib + 115 integration + 2 doctests pass. clippy + fmt clean. Bump @firecrawl/pdf-inspector to 1.6.2 (patch — fixes a regression introduced by 1.6.1; no API changes). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
56 lines
1.2 KiB
JSON
56 lines
1.2 KiB
JSON
{
|
|
"name": "@firecrawl/pdf-inspector",
|
|
"version": "1.6.2",
|
|
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
|
"main": "index.js",
|
|
"types": "index.d.ts",
|
|
"bin": {
|
|
"pdf-inspector": "bin/pdf-inspector.mjs"
|
|
},
|
|
"license": "MIT",
|
|
"keywords": [
|
|
"pdf",
|
|
"pdf-extraction",
|
|
"pdf-parser",
|
|
"text-extraction",
|
|
"ocr",
|
|
"pdf-classification",
|
|
"napi",
|
|
"rust",
|
|
"firecrawl"
|
|
],
|
|
"files": [
|
|
"index.js",
|
|
"index.d.ts",
|
|
"*.node",
|
|
"bin/",
|
|
"README.md"
|
|
],
|
|
"repository": {
|
|
"type": "git",
|
|
"url": "https://github.com/firecrawl/pdf-inspector"
|
|
},
|
|
"homepage": "https://github.com/firecrawl/pdf-inspector",
|
|
"publishConfig": {
|
|
"access": "public"
|
|
},
|
|
"napi": {
|
|
"binaryName": "pdf-inspector",
|
|
"targets": [
|
|
"x86_64-unknown-linux-gnu",
|
|
"aarch64-apple-darwin",
|
|
"x86_64-pc-windows-msvc"
|
|
],
|
|
"package": {
|
|
"name": "@firecrawl/pdf-inspector-js"
|
|
}
|
|
},
|
|
"scripts": {
|
|
"build": "napi build --platform --release",
|
|
"build:debug": "napi build --platform"
|
|
},
|
|
"devDependencies": {
|
|
"@napi-rs/cli": "^3.4.1"
|
|
}
|
|
}
|