fix: reduce false OCR recommendations for text PDFs with figures (#38)
* fix: reduce false OCR recommendations for text PDFs with figure images
Two fixes in the detector:
1. Fix Tf operator parsing: some PDFs concatenate Tf directly with the
next operator (e.g. "25 Tf[<01>...") without whitespace. The scanner
now accepts [, (, <, / as valid followers, fixing font_changes being
reported as 0.
2. Distinguish text-with-figures from scanned-with-OCR: pages with
multiple images (image_count > 1) and strong text signals (text_ops
>= 50, alphanum >= 10) are recognized as text pages with figures,
not scanned templates. Scanned PDFs have exactly 1 full-page image.
This prevents academic papers, reports with charts, and similar PDFs
from being incorrectly classified as Mixed/OCR-needed when their text
is perfectly extractable.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore: bump napi version to 0.7.2
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: remove template image influence from page classification
Template images (large background/figure images) no longer affect
pages_needing_ocr. In the region-based pipeline, text regions are
extracted independently from image regions, and per-region needs_ocr
quality checks handle scanned-with-OCR garbage text.
Also makes the invisible text retry (for OCR text layers) trigger on
text quality rather than PDF type, so it works regardless of
classification.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Revert "fix: remove template image influence from page classification"
This reverts commit 100cbe5453.
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.6
parent
7c8b09be67
commit
0a9c120a6b
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "firecrawl-pdf-inspector",
|
||||
"version": "0.7.1",
|
||||
"version": "0.7.2",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
|
||||
Reference in New Issue
Block a user