PR #62 (1.6.1) closed the cell-bleed regression by clamping SLANet's loose cell bboxes into non-overlapping row/column bands and tightening text membership to center-containment OR >=60% overlap. That worked, but exposed the opposite failure: legitimate native PDF text whose center fell just outside the *clamped* cell bbox now had nowhere to go. Two distinct failure modes observed against the FNBO branch-list PDF after 1.6.1 deployed: Symptom A — header text positioned at the LEFT of a column whose band was derived from data-cell centers farther right. Header "Address" PDF text at x=331..375 fell outside the clamped col band starting at x=410. Strict membership rejected it (0% x-overlap, center outside). Symptom B — local SLANet row drift in col 0 over a 5-row stretch. Cell bboxes sat just above the actual branch-name text items (x-overlap 100% but y-overlap ~30-43%, below the 60% threshold). Both share one root: post-normalization bboxes are too tight, and the strict rule has no escape valve for legitimate edge text. Fix: add a stage-2 orphan-recovery pass after the strict fill. Items that NO cell claimed in stage 1 get re-assigned to their nearest *empty* cell, distance-capped by `(median_col_width, median_row_height)` so a far-orphan figure title can't get pulled into a faraway empty cell. Stage 2 only fills empties — never overwrites stage 1 — so the cell-bleed case PR #62 closed cannot regress. Three new lib tests cover the bug shapes: - stage2_recovers_left_aligned_header_text_outside_data_band (Symptom A: header text left-of-band, data cells already filled, stage 2 fills only the header) - stage2_recovers_y_shifted_col0_in_consecutive_rows (Symptom B: 3 col-0 cells shifted vs text, all 3 recovered) - stage2_does_not_overwrite_filled_cells_or_admit_far_orphans (cap rejects far figure titles; pre-filled cells untouched) Plus tests for the cap-derivation helper: - tsr_assignment_caps_uses_median_geometry - tsr_assignment_caps_floor_protects_degenerate_input The existing dense-overlapping-rows regression test (added in #62) still passes, confirming no regression on the cell-bleed case. Test results: 414 lib + 115 integration + 2 doctests pass. clippy + fmt clean. Bump @firecrawl/pdf-inspector to 1.6.2 (patch — fixes a regression introduced by 1.6.1; no API changes). Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
PDF Inspector
Fast PDF classification and region-based text extraction for Node.js/Bun. Native Rust performance via napi-rs.
Built by Firecrawl for hybrid OCR pipelines — extract text from PDF structure where possible, fall back to OCR only when needed.
Install
npm install @firecrawl/pdf-inspector
# or
bun add @firecrawl/pdf-inspector
Prebuilt binaries included for linux-x64 and macOS ARM64. No Rust toolchain needed.
API
classifyPdf(buffer: Buffer): PdfClassification
Classify a PDF as TextBased, Scanned, Mixed, or ImageBased (~10-50ms). Returns which pages need OCR.
import { classifyPdf } from '@firecrawl/pdf-inspector'
import { readFileSync } from 'fs'
const pdf = readFileSync('document.pdf')
const result = classifyPdf(pdf)
console.log(result.pdfType) // "TextBased" | "Scanned" | "Mixed" | "ImageBased"
console.log(result.pageCount) // 42
console.log(result.pagesNeedingOcr) // [5, 12, 15] (0-indexed)
console.log(result.confidence) // 0.875
extractTextInRegions(buffer: Buffer, pageRegions: PageRegions[]): PageRegionTexts[]
Extract text within bounding-box regions from a PDF. Designed for hybrid OCR pipelines where a layout model detects regions in rendered page images, and this function extracts text from the PDF structure for text-based pages — skipping GPU OCR.
Each region result includes a needsOcr flag that signals unreliable extraction (empty text, GID-encoded fonts, garbage text, encoding issues).
import { extractTextInRegions } from '@firecrawl/pdf-inspector'
const result = extractTextInRegions(pdf, [
{
page: 0, // 0-indexed
regions: [
[0, 0, 300, 400], // [x1, y1, x2, y2] in PDF points, top-left origin
[300, 0, 612, 400],
]
}
])
for (const region of result[0].regions) {
if (region.needsOcr) {
// Unreliable text — send this region to OCR instead
} else {
console.log(region.text) // Extracted text in reading order
}
}
Types
interface PdfClassification {
pdfType: string // "TextBased" | "Scanned" | "Mixed" | "ImageBased"
pageCount: number
pagesNeedingOcr: number[] // 0-indexed page numbers
confidence: number // 0.0 - 1.0
}
interface PageRegions {
page: number // 0-indexed
regions: number[][] // [[x1, y1, x2, y2], ...] in PDF points, top-left origin
}
interface PageRegionTexts {
page: number
regions: RegionText[]
}
interface RegionText {
text: string
needsOcr: boolean // true when text is unreliable
}
Platforms
| Platform | Architecture | Supported |
|---|---|---|
| Linux | x64 | Yes |
| macOS | ARM64 | Yes |
License
MIT