The rect-based cell fallback in detect_row_stripe_table_from_cell_rects
derives columns purely from text X-position clustering. When prose
wraps inside a bounding-box rect (chat transcripts, stylized figures),
the word-boundary gaps cluster into many spurious columns, producing
a multi-column "table" that is just fragmented prose.
Count cells containing common English function words (articles,
prepositions, pronouns, common verbs) and reject the fallback when
20%+ of non-empty cells contain any such word. Real tabular data —
labels, units, numbers, short identifiers — rarely contains these.
Update the td9264 snapshot: the government document section that
previously rendered as a malformed table now renders as cleaner
prose + a proper CFR list.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: don't reclassify wrapped bold list leads as headings
When a numbered/bulleted list item's bold lead phrase wraps onto a
second visual line, that line is all_bold + standalone, which scored
above the rarity heading threshold and was emitted as #### in the
middle of the item. That reset in_list, so the body continuation
below picked up a stray `- ` bullet via the struct-tree LI path,
shattering a single item into heading + stray bullets.
Guard the font heuristic: when already inside a list, skip heading
classification for lines at the list continuation indent with a Y
gap within para_threshold. Structure-tree headings still win.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: tighten rect-row span check in propagate_merged_cells
propagate_merged_cells used an overlap-based predicate with ±tol
slop that returned true at shared row boundaries — a rect whose
top exactly equals row N's bottom lies entirely below the row, yet
the predicate considered it to span row N. When multiple background
rects aligned on a shared Y edge (e.g. consecutive row-stripe
shading), each adjacent rect would over-reach by one row, cascading
labels and data from unrelated rows into a single merged cell.
Replace the overlap predicate with a containment check: rect bottom
at or below row bottom, rect top at or above row top (each within
tol). Genuine merged-cell rects fully contain the rows they span;
tangent rects do not.
Update two snapshots that were encoding the old buggy output.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Pre-scan lines to identify "isolated" ones — short lines (1-6 words)
with paragraph breaks both before AND after. These are heading
candidates even at body font size, common in academic papers
("Acknowledgements", "Limitations", "B.3 Prompt Engineering").
Inspired by opendataloader's HeadingProcessor which passes prevNode
and nextNode context to the heading probability scorer.
The isolated signal (+0.3) combines with rarity/bold/standalone
signals. A per-page density guard prevents false positives on
multi-column pages where many lines appear isolated. Continuation
word detection (ending in "the", "and", etc.) filters wrapped
paragraph lines.
MHS=0 docs: 18→13. MHS-S +0.004. No regressions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Reverts commits 937311c, 1dcb0c6, 0300e96, 999f9a2. The thin-rect-to-line
synthesis and stacked table splitting improved extraction for specific
government PDFs but caused -0.05 TEDS regression on the benchmark by
preempting the heuristic detector with worse line-based grids.
These features need more targeted guards before re-enabling.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Rows with 3+ short-valued cells (avg ≤10 chars) and an empty first cell
are column headers (e.g. "UR | SC | ST | OBC | EWS"), not text overflow
from the previous row. Prevents them from being merged into the
preceding section title row.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three fixes for better table extraction from spreadsheet-exported PDFs:
1. Convert thin filled rects (< 2pt) to PdfLine objects before line-based
table detection. Many PDFs draw table borders as narrow filled rectangles
instead of stroked paths — these were invisible to the line detector.
2. Relax uniform row spacing rejection (CV 0.05 → 0.02). Spreadsheet
exports have very even row heights that were being rejected as "chart
grids".
3. Fix continuation row merging: don't merge rows where the only non-first
cell content is a long label (section headers like "Category No. 03").
Don't merge first-cell-only rows with long text ("Note: ...").
Also adds multi-Y row splitting in line-based detection and column-aware
table detection skipping for multi-column pages.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace ad-hoc bold/ratio heading checks with a unified scoring system
based on font size rarity. For each line, compute:
score = font_rarity * 0.5 + bold * 0.3 + standalone * 0.2
Font rarity measures how infrequently a font size appears across the
document — heading fonts are rare while body text is common. This
approach (from opendataloader's ModeWeightStatistics) naturally adapts
to each document's font distribution instead of relying on fixed
thresholds.
Guards: require font_size >= 0.95 * base_size (no small-font headings),
word_count >= 3, and standalone (paragraph break before).
Benchmark improvement: MHS 0.56→0.58, MHS-S 0.66→0.70, overall +0.003.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Lines with font size 1.10-1.20x body text that are standalone and short
(1-8 words) are promoted to headings. This catches academic paper
headings where the font is only ~10% larger than body text, below the
previous 1.2x threshold.
Also syncs the simpler to_markdown_from_lines path to match the
table-aware path (removes stale colon exclusion).
Benchmark improvement: MHS 0.54→0.56, overall 0.761→0.766.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Caption detection was incorrectly classifying "Table of Contents" as a
caption because it starts with "Table ". Now "Table" and "Figure"
prefixes require a digit, parenthesis, or hash after them — matching
actual captions like "Table 1", "Figure 3.2" but not titles.
Also removes debug logging left from previous iteration.
Benchmark improvement: MHS 0.52→0.54, overall 0.757→0.761.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The colon exclusion was preventing legitimate headings like "Steps for
Using the Microscope:" and "Changing objectives:" from being detected.
The single edge case it was protecting (chart sub-headers) is less
impactful than the many headings it was blocking.
Benchmark improvement: MHS 0.51→0.52, MHS-S 0.61→0.62.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Lower minimum item count for body-font table candidates from 9 to 6,
allowing small 2-3 row tables to be detected.
- Allow 2-column body-font tables with short cells (avg ≤25 chars) to
bypass the "table-like content" validation. This catches text-only
definition/category tables (e.g., species lists) without false-positiving
on 2-column paragraph text (which has longer cells).
Benchmark improvement: TEDS 0.498→0.519, overall 0.750→0.754.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Bold lines at body font size that are standalone (preceded by a paragraph
break) and have ≥3 words are promoted to headings. This catches the
common pattern in academic/technical PDFs where section headings use
bold text at the same size as body text.
Guards against false positives: minimum word count, colon-ending
exclusion (labels like "Table I:").
Benchmark improvement: MHS 0.37→0.50, overall 0.71→0.75.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(tables): improve heuristic detection for borderless wrapped-cell tables
Three changes to the body-font heuristic detector:
1. Adaptive Y-gap in find_table_regions_strict: use median qualifying-row
spacing × 3 instead of fixed 25pt. Tables with wrapped cells have
larger gaps between qualifying rows (those with 3+ X-clusters).
2. Y-only region filtering: use full X range when collecting region items.
The strict X bounds from qualifying rows excluded continuation lines
in wrapped cells, starving find_column_boundaries of items.
3. Merged-band retry: when split_side_by_side splits a page into bands
but no band produces a table, retry heuristic detection on all items
merged as a single band.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(tables): gap-histogram column detection for small tables + lower avg_cells
Two changes to fix PDF 045 (borderless table with narrow "No." column):
1. Extend gap-histogram column threshold to small tables: when the gap
between within-column jitter and between-column spacing is >10pt
(unambiguous bimodal signal), use the detected threshold even with
fewer than 500 items. Previously only triggered for dense tables.
2. Lower BodyFont avg_cells_per_row minimum from 2.5 to 2.0 to handle
tables with wrapped multi-line cells where continuation lines have
only 1 filled cell.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(tables): trim empty outer columns + relax partial H-line validation
- Rect detection: trim empty first/last columns instead of rejecting
the whole table. Rect edges often extend beyond text boundaries.
- Line detection: accept tables with 6+ partial horizontal lines
(>15% width) when <3 full-spanning lines exist. Handles tables
with column-level separators instead of full-width rules.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(tables): cell-rect fallback for tables with variable-width backgrounds
When rect clustering produces a grid that fails validation (empty
interior columns from variable-width cell backgrounds), fall through
to a new strategy: use rect Y-edges for row boundaries and text
X-position clustering for columns. This handles tables like the
opendataloader-bench 088-090 comparison tables where each cell has
its own background rect at different widths.
Also widen failed-cluster hint width cap for large clusters (≥30 rects)
to allow page-spanning table regions.
TEDS score on opendataloader-bench: 0.300 → 0.353.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(tables): relax vertical line spanning validation for partial borders
Accept tables with 4+ partial vertical lines (>10% table height) when
fewer than 2 span >30%. Handles tables like opendataloader-bench 053
with column-level vertical separators that don't extend the full height.
TEDS: 0.353 → 0.377 on opendataloader-bench.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(tables): lower cell-rect density threshold + add validation logging
- Lower cell-rect density minimum from 25% to 15% to accept sparser
tables with decorative backgrounds (fixes 147).
- Add debug logging to all heuristic validation paths for diagnosability.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(tables): relax body-font validations for text-only and 2-column tables
Four fixes closing 71% of the TEDS gap vs opendataloader:
1. Validation 7 (table-like content): bypass numeric content requirement
for tables with 3+ columns that passed all structural checks. Text-only
tables (category lists, program descriptions) are legitimate.
2. Qualifying row threshold: lower from 3+ to 2+ X-clusters per row.
Enables 2-column body-font table detection (fixes 166).
3. Row-stripe max cell length: raise from 500 to 2000 for 3+ column
tables. Tables with paragraph descriptions in one column are valid
(fixes 121).
4. Row-stripe empty-column trimming: apply the same outer-column trim
as grid detection (fixes 121 column-0 rejection).
TEDS: 0.377 → 0.438 on opendataloader-bench (gap: -0.056 vs odl).
TEDS=0 docs: 14 → 9.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(tables): enable 2-column body-font tables + lower all minimums
- Lower BodyFont minimum columns from 3 to 2 in detect_table_in_region
- Lower BodyFont minimum rows from 3 to 2
- Lower avg_cells_per_row minimum from 2.0 to 1.5 (handles wrapped cells
in 2-column tables)
- Apply empty-outer-column trimming to row-stripe detection (not just grid)
TEDS: 0.438 → 0.468 on opendataloader-bench (gap: -0.027 vs odl).
TEDS=0 docs: 9 → 8. 86% of original gap closed.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(tables): text-based row fallback + fix 120 flow-chart and 188 leaderboard
Three changes that push TEDS past opendataloader:
1. Cell-rect Y-edge fallback: when rects have too few Y-edges for row
structure, derive rows from text Y-position clustering within the
rect bounding box. Fixes flow-chart tables (120) and column-header-
only rects (188).
2. Lower cell-rect minimum from 20 to 6 rects to catch smaller tables.
3. Relax validation 1 (first-column presence) from 50% to 25% of rows.
Tables with wrapped model names have continuation lines without first
column content.
TEDS: 0.468 → 0.508 on opendataloader-bench.
Now BEATS opendataloader (0.508 vs 0.494, gap=+0.014).
TEDS=0 docs: 8 → 5.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(tables): wrapped-cell continuation row merging
Merge rows that have fewer filled cells than the header row into the
previous row. Handles wrapped multi-line cells where text overflow
creates extra rows (e.g., "Direct" + "communications" → "Direct
communications").
Conditions: fewer filled cells than header, more than previous row had,
not a data row (numeric), not a short subheader label.
TEDS: 0.508 → 0.522 on opendataloader-bench (now +0.028 vs odl).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(tables): tighten cell-rect validation + fix continuation-row merging
Cell-rect false positives:
- Raise density threshold back to 25% (from 15%)
- Add max cell length check (500 chars) to reject paragraph content
- Reject disproportionate grids (>20 rows, <4 cols)
Continuation-row merging:
- Wide tables (5+ cols): only merge rows with ≤50% header cells
- Narrow tables (2-4 cols): merge rows with fewer cells than header
- Prevents merging normal data rows in large tables (6_KE_Chart)
while keeping wrapped-cell merging for narrow tables (178)
TEDS: 0.498 on opendataloader-bench (still +0.004 vs odl).
pdf-evals: 191/192 passed, 0 regressions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Struct-tree tables with incomplete page tagging (e.g., only 22 of 50+
rows tagged on a page) would claim items and block rect detection,
leaving unclaimed items as loose text. Now require struct-tree tables
to capture ≥50% of band items before using them; incomplete trees
fall through to geometry-based detection.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Skip rect-detected tables that overlap with items already claimed by
struct-tree detection. Previously both strategies emitted separate
tables for the same content, doubling the output on tagged PDFs.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When a PDF has a well-formed structure tree with /Table > /TR > /TD|TH
elements linked to MCIDs, build tables directly from the semantic
hierarchy. Runs as highest-priority detection (step 0) before rect-based,
line-based, and heuristic strategies.
- Add StructTree::extract_tables() to walk the tree and collect table
descriptors with row/cell/MCID info
- Add detect_tables_from_struct_tree() to match MCIDs to TextItems
- Reject tables with <30% MCID cell coverage (stale structure trees)
- Update 2013-app2 snapshot (struct-tree gives valid but different
column ordering)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(text): handle Tc/Tw character and word spacing in text width computation
PDFs using Tc (character spacing) and Tw (word spacing) operators for text
justification had words incorrectly split across TextItems. The computed
advance width didn't account for these spacing parameters, causing spurious
spaces mid-word (e.g. "deve lopers" instead of "developers").
- Add Tc/Tw operator handling and graphics state save/restore
- Incorporate char_spacing and word_spacing into compute_string_width_ts
- Add adaptive merge threshold: tighter for lowercase→lowercase junctions,
wider before joining punctuation
- Add unit tests for Tc/Tw width computation and merge behavior
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* test(fonts): add large Tc width computation test
Verifies that large character spacing values are applied in full
without any artificial cap, matching PDF spec behavior.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(text): guard against Tc/Tw-inflated widths in merge and join paths
Two targeted fixes to prevent character-spacing (Tc) and word-spacing
(Tw) inflation from causing data quality regressions:
1. should_join_items: reject large negative gaps (< -font_size) that
arise when Tc/Tw inflate item widths past adjacent items. Fixes
FY_2015 merged numbers (e.g. "239.696.0" → "239.69 6.0").
2. merge_text_items: cap effective width for gap computation when Tw
inflates space-containing items beyond 0.85× font_size per char.
Prevents column-level gaps from collapsing into merge range,
recovering table detection for Baldwin-Edwards and similar PDFs.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* fix(fonts): CID-as-Unicode passthrough and subscript/superscript merging
Add smart CID-as-Unicode passthrough for Identity-H fonts without
ToUnicode maps. Uses /W array median CID heuristic to distinguish
Unicode-CID PDFs (Chromium-generated) from GID-based subsets.
Add merge_subscript_items() pass that merges small-font items (<75%
of dominant font size, ≤4 chars, tightly adjacent) into parent items.
Fixes chemical formulas (NH3, H2O, KClO3), footnote references, and
subscript notation (vf, Hfg, m3/kg) that were previously orphaned
as separate text items.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(subscripts): restrict merge to purely numeric text only
Tighten subscript merging to only merge items containing ASCII digits
(0-9). This avoids false positives with ordinal indicators (º), letter
subscripts (sol, vf), and small bullet characters (▶) that caused
table restructuring regressions.
Numeric-only keeps the primary wins: chemical formulas (NH3, H2O),
footnote references, and unit notation (m2, m3).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(subscripts): restrict merge to parent text ending with a letter
Only merge numeric subscripts when the parent item's text ends with
an alphabetic character. Prevents false merges like "33" + "1" in
fractions (33 1/3%), table credit numbers after spaces, and footnote
refs after punctuation (land.1 → land. 1). Chemical formulas (NH3,
H2O, KClO3) still merge correctly since parent ends with a letter.
Reduces pdf-eval regressions from 13 to 2 (both are correct reversions
of over-aggressive footnote merging from the prior commit).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Document titles, column headers, and newsletter mastheads that repeat
on every page were being fully stripped by strip_repeated_lines. Now
the first page's occurrence is preserved so the content appears once.
Fixes missing titles like "VOICE OF SOUTH MARION" in tax certificate
PDFs and column headers in IRS forms.
OCR text layers and some PDF producers emit trailing spaces on each
text item, which combine with gap-based joining to produce double
spaces ("Vice President"). Now collapses runs of 2+ spaces to single
space within lines, preserving leading indentation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Drop column-width padding from table output. Saves ~37% bytes on
table-heavy PDFs. Markdown renders identically — padding was purely
cosmetic. Optimizes token efficiency for AI agent consumers.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Tables with very wide cells (e.g. 230+ chars from merged schedule columns)
caused massive whitespace padding in every row. Capping alignment at 40
characters reduces output size significantly (franklin: 77KB→24KB,
Closed-Business: 571KB→502KB) without affecting content.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Only skip text with rendering mode 3 (truly invisible). Removes
fill color tracking that was incorrectly hiding white-on-dark text
in presentations and brochures.
Strip leading/trailing digit sequences (page numbers) before frequency
comparison so headers like "Page 5" and "Page 6" are treated as identical.
Group TextLines at the same Y position into Y-bands and propagate removal
to all siblings when any member is stripped. Increase EDGE_LINE_COUNT from
4 to 5 to cover 5-row form column headers (e.g., IRS p1244).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Bumps EDGE_LINE_COUNT from 3 to 4 to strip repeated form column headers
that sit just inside the page margin (4th-from-edge Y position).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add strip_repeated_lines() preprocessing step that removes running
headers and footers before markdown conversion. Uses multiple guard
rails to avoid false positives:
- Edge-line detection: only considers lines among the first/last 3
distinct Y positions on each page (not percentage-based margins)
- Y-position consistency: requires low variance across pages (stddev
< 5% of page span) to distinguish headers from scattered table content
- Position-aware removal: only strips instances at page edges, preserving
body content that happens to match a header/footer text
- Skips structural lines (headings, list items), short lines (<10 chars),
and decorative separators (repeated single characters)
- Frequency threshold: >= max(3, page_count * 30%) distinct pages
New strip_headers_footers option on MarkdownOptions (default: true).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Unit test verifying short sub-header rows (e.g. month names) are not
merged as continuation rows in table formatting. Snapshot test for
2013_app2.pdf pinning the full row-stripe table output.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When a PDF page has a graph directly above a table (e.g. Figure 2 above
Table II in real-estate-pricing.pdf), the heuristic table detector would
merge graph axis labels and legend items into the table, producing a
corrupted result.
Add rect hint regions: when a small cluster of cell-border rects (4-6)
fails full grid validation (e.g. only row borders, no column dividers),
extract their Y bounding box as a "hint region". The heuristic detector
then runs separately on items inside vs outside hint regions, preventing
unrelated content from being merged into tables.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Rects spanning multiple grid rows (e.g. classification labels) now have
their text consolidated into the first sub-row via propagate_merged_cells().
Also prevent continuation-row merge from folding body data into the header.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add 5 stripped public-domain PDF fixtures and golden markdown snapshots
for CI regression testing. Fix non-deterministic output caused by
HashMap iteration order in font stats, table heuristics, and rect
clustering by adding deterministic tie-breaking.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>