detect_lines: derive column edges from horizontal-segment endpoints (#86)
Some catalog and archival-finding-aid tables draw each row's horizontal rule as N segments (one segment per cell) with no vertical lines at all. The previous detector rejected these outright at `verticals.len() < 2`, even though the segment break points encoded the column boundaries unambiguously. When the vertical-line count is below the existing threshold, walk the horizontal-segment x-endpoints and cluster them with the same snap_edges path used for vertical-line columns. Accept the derived edges only when ≥3 distinct x-positions each appear on ≥50% of the unique horizontal-line rows — that consistency guard distinguishes real per-cell segments from decorative rules with varying widths (which never share endpoints across many rows). When columns come from segment endpoints, skip the downstream "spanning_v / partial_v" gate (there are no vertical lines to validate against). All other gates — horizontal-span coverage, content density, capture ratio, multi-column distribution, the uniform-spacing chart-grid rejector — still apply. Verified on a 7-row × 3-col archival catalog page that previously extracted 98 chars (a 2-row fragment via the heuristic fallback); now extracts 1754 chars with all rows + multi-line cells. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
parent
b5b91470db
commit
79d75dbdca
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "@firecrawl/pdf-inspector",
|
||||
"version": "1.8.10",
|
||||
"version": "1.8.11",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
|
||||
Reference in New Issue
Block a user