extract_tables_in_regions: needs_ocr on suspicious table structure (0.4.1)
When the heuristic returns markdown that looks like a partial / mis-detected table, set needs_ocr=true so the caller falls back to GPU OCR. Previously the same cases returned the broken table with needs_ocr=false, which produced real-world TEDS=0 scores in fire-pdf evals (heuristic-built table didn't match ground truth structure at all, but caller had no signal to fall back). Four failure modes detected, all observed in opendataloader-bench eval losses: 1. **Header looks like a data row** — first cell of header is a bare number (e.g. `|2|...`), suggesting the actual header row was skipped. Real headers almost never start with just a number. 2. **Empty header cells in a multi-column table** — ≥3 cols, ≥1 empty cell in the header row. Indicates poor column boundary detection. 3. **Duplicate header cells** — same non-empty value appearing twice in the header (e.g. "Administration|Administration"). Means a multi-line header was collapsed wrong. 4. **Sparse first data row** — ≥3 cols and ≥1/3 of first-data-row cells are empty. Multi-row headers in the source PDF get smashed into header + sparse data row by the heuristic; this catches that. Tests: 9 new unit tests in `looks_like_partial_table_tests` cover each failure mode plus realistic non-failures (well-formed table, single-column list, two-col with a single empty cell). All 91 existing tests still pass. Bumps `napi/package.json` to 0.4.1 since this changes the function's return behaviour for callers (some inputs that returned needs_ocr=false now return true). The output text field is also cleared on the new fallback path so callers don't accidentally use the broken markdown. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.6
parent
8282c2f8ee
commit
d0dd067e70
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "firecrawl-pdf-inspector",
|
||||
"version": "0.4.0",
|
||||
"version": "0.4.1",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
|
||||
Reference in New Issue
Block a user