The page classifier was over-aggressively flagging Mixed-PDF pages as needing OCR in three distinct cases. Each is fixed at the root in analyze_page_content / page_has_identity_h_no_tounicode / the looks_like_scan check. 1. has_vector_text false positives on dense layouts path_ops > text_ops*200 fired on pages with decorative paths (column borders, dividers) alongside real selectable text. Added a unique_alphanum_chars < 30 guard: real outlined-text pages have very few unique alphanum chars (each glyph is a path), while pages with real text + decorations have many. 2. Identity-H without ToUnicode flagged whole pages on supplementary fonts page_has_identity_h_no_tounicode would flag a page if any single Type0 font lacked ToUnicode and had no fallback CMap, even when the page's actual text came from other decodable fonts (Type1 with ToUnicode, etc.). Rewrote to track both undecodable Identity-H fonts AND other decodable fonts, only flagging when no decodable text font is present. 3. CID-encoded text with ToUnicode misclassified as scan looks_like_scan checked unique_alphanum_chars < 10 on raw string operand bytes. CID-encoded fonts (Type0 with ToUnicode) emit 2-byte CID values that aren't ASCII alphanum, so the metric is blind to them even when the text is fully decodable. Added a has_decodable_text_fonts signal: when a page has decodable fonts AND >= 10 text ops, the low alphanum count is treated as a CID encoding artifact rather than evidence of a scan. Validated against a broad PDF corpus: - 6 known false-positive pages now correctly classified as text - 22 previously-missed scan pages (cover/blank/photo) now correctly flagged for OCR - 0 regressions on truly-scanned PDFs (61/61 pages stay flagged) - All 437 existing tests pass; clippy clean Bumps NAPI package to 0.7.4.
51 lines
1.1 KiB
JSON
51 lines
1.1 KiB
JSON
{
|
|
"name": "firecrawl-pdf-inspector",
|
|
"version": "0.7.4",
|
|
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
|
"main": "index.js",
|
|
"types": "index.d.ts",
|
|
"license": "MIT",
|
|
"keywords": [
|
|
"pdf",
|
|
"pdf-extraction",
|
|
"pdf-parser",
|
|
"text-extraction",
|
|
"ocr",
|
|
"pdf-classification",
|
|
"napi",
|
|
"rust",
|
|
"firecrawl"
|
|
],
|
|
"files": [
|
|
"index.js",
|
|
"index.d.ts",
|
|
"*.node",
|
|
"README.md"
|
|
],
|
|
"repository": {
|
|
"type": "git",
|
|
"url": "https://github.com/firecrawl/pdf-inspector"
|
|
},
|
|
"homepage": "https://github.com/firecrawl/pdf-inspector",
|
|
"publishConfig": {
|
|
"access": "public"
|
|
},
|
|
"napi": {
|
|
"binaryName": "pdf-inspector",
|
|
"targets": [
|
|
"x86_64-unknown-linux-gnu",
|
|
"aarch64-apple-darwin"
|
|
],
|
|
"package": {
|
|
"name": "@firecrawl/pdf-inspector-js"
|
|
}
|
|
},
|
|
"scripts": {
|
|
"build": "napi build --platform --release",
|
|
"build:debug": "napi build --platform"
|
|
},
|
|
"devDependencies": {
|
|
"@napi-rs/cli": "^3.4.1"
|
|
}
|
|
}
|