Files
pdf-inspector/napi/package.json
T
Abimael Martell f4bdb35935 fix(detector): correct false flags for CID-encoded text and supplementary fonts (0.7.4)
The page classifier was over-aggressively flagging Mixed-PDF pages as
needing OCR in three distinct cases. Each is fixed at the root in
analyze_page_content / page_has_identity_h_no_tounicode / the
looks_like_scan check.

1. has_vector_text false positives on dense layouts
   path_ops > text_ops*200 fired on pages with decorative paths
   (column borders, dividers) alongside real selectable text. Added
   a unique_alphanum_chars < 30 guard: real outlined-text pages have
   very few unique alphanum chars (each glyph is a path), while
   pages with real text + decorations have many.

2. Identity-H without ToUnicode flagged whole pages on supplementary fonts
   page_has_identity_h_no_tounicode would flag a page if any single
   Type0 font lacked ToUnicode and had no fallback CMap, even when
   the page's actual text came from other decodable fonts (Type1
   with ToUnicode, etc.). Rewrote to track both undecodable
   Identity-H fonts AND other decodable fonts, only flagging when
   no decodable text font is present.

3. CID-encoded text with ToUnicode misclassified as scan
   looks_like_scan checked unique_alphanum_chars < 10 on raw string
   operand bytes. CID-encoded fonts (Type0 with ToUnicode) emit
   2-byte CID values that aren't ASCII alphanum, so the metric is
   blind to them even when the text is fully decodable. Added a
   has_decodable_text_fonts signal: when a page has decodable fonts
   AND >= 10 text ops, the low alphanum count is treated as a CID
   encoding artifact rather than evidence of a scan.

Validated against a broad PDF corpus:
- 6 known false-positive pages now correctly classified as text
- 22 previously-missed scan pages (cover/blank/photo) now correctly
  flagged for OCR
- 0 regressions on truly-scanned PDFs (61/61 pages stay flagged)
- All 437 existing tests pass; clippy clean

Bumps NAPI package to 0.7.4.
2026-04-15 12:35:59 -07:00

51 lines
1.1 KiB
JSON

{
"name": "firecrawl-pdf-inspector",
"version": "0.7.4",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
"license": "MIT",
"keywords": [
"pdf",
"pdf-extraction",
"pdf-parser",
"text-extraction",
"ocr",
"pdf-classification",
"napi",
"rust",
"firecrawl"
],
"files": [
"index.js",
"index.d.ts",
"*.node",
"README.md"
],
"repository": {
"type": "git",
"url": "https://github.com/firecrawl/pdf-inspector"
},
"homepage": "https://github.com/firecrawl/pdf-inspector",
"publishConfig": {
"access": "public"
},
"napi": {
"binaryName": "pdf-inspector",
"targets": [
"x86_64-unknown-linux-gnu",
"aarch64-apple-darwin"
],
"package": {
"name": "@firecrawl/pdf-inspector-js"
}
},
"scripts": {
"build": "napi build --platform --release",
"build:debug": "napi build --platform"
},
"devDependencies": {
"@napi-rs/cli": "^3.4.1"
}
}