Commit Graph
14 Commits
Author SHA1 Message Date
Abimael MartellandClaude Opus 4.7 b4c34ba4b7 refactor: classify Table as Data vs Toc once at construction (#52)
The TOC/data-table distinction was being recomputed at every consumer:
- format.rs::table_to_markdown ran is_table_of_contents to decide between
  flat-list and markdown-table rendering.
- compute_layout_complexity ran it again to filter TOCs out of
  pages_with_tables.
- detect_heuristic validations used it to decide whether to relax val 1/9.

Each caller had to remember tables can be either kind, which leaks the
TOC concept across the codebase.

Add `TableKind { Data, Toc }` and a `Table::new` constructor that classifies
once from the cells. All five detectors (heuristic, rect, line, struct,
columns) now go through `Table::new`. Consumers match on `kind` instead of
re-running classification.

Pure refactor — no behavior change. Verified: pdf-evals output is byte-for-
byte identical (0 changed snapshots).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 14:04:50 -07:00
Abimael MartellandClaude Opus 4.7 e043c4414d fix: reject dot-less TOC as heuristic table false-positive (#45)
* fix: reject dot-less tables of contents as heuristic table false-positive

The body-font heuristic table detector was firing on tagged PDF TOCs
whose entries are laid out as multi-column rows (section number, title,
page number) without dot leaders. Extend is_table_of_contents to flag
this pattern by requiring:
- 2+ columns and 4+ rows
- >=60% of rows start with a dotted section number (e.g. "4.3.1")
- >=70% of filled last-column cells are all-digit page numbers

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* format TOC tables as flat per-row list instead of dropping them

Keeping the previous approach (rejecting TOCs at detection time) meant
dot-less TOCs fell back to the column-aware text reader, which stacked
every section title first and dumped all page numbers at the end of the
region — so a reader saw a wall of titles followed by a wall of numbers.

Keep the detected Table, and in the formatter interleave the cells into
one line per row with the page number appended after a tab.  The raw
(pre-clean) cells are used for the TOC check because clean_table_cells
can collapse genuine data tables into a TOC-looking shape.

Tightened is_table_of_contents to require both dot leaders AND page
numbers (previously dot-ratio alone was enough, which misfired on
register tables that use "..." as a continuation marker).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* split TOC detection into dot-leader and tabular patterns

Previous is_table_of_contents rejected every TOC-shaped table at detect
time, dropping dot-leader TOCs that render well as flat-list output and
misclassifying data tables with trailing "..." cells as TOCs.

Now is_dot_leader_toc (structural + inline-leader) keeps per-row flat
layouts for the formatter, while only wide inline-leader indices are
rejected at detect time.  row_cell_is_page_number accepts dashed
section-page IDs ("A-1", "5-21") and rejects decimals/thousands
separators; cell_has_trailing_leader requires alphabetic content so
numeric data-row labels ("1973 ... ") no longer qualify.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-17 11:58:41 -07:00
Abimael MartellandClaude Opus 4.6 dbc2de3e9b revert: undo table splitting changes that caused TEDS regression
Reverts commits 937311c, 1dcb0c6, 0300e96, 999f9a2. The thin-rect-to-line
synthesis and stacked table splitting improved extraction for specific
government PDFs but caused -0.05 TEDS regression on the benchmark by
preempting the heuristic detector with worse line-based grids.

These features need more targeted guards before re-enabling.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 11:20:05 -07:00
Abimael MartellandClaude Opus 4.6 999f9a2f2d feat: split stacked tables at rows without vertical borders
Rows that sit between horizontal rules but lack vertical border coverage
are not table cells — they're freestanding text (e.g. "Note: The cutoff
mark is out of 120"). These rows now split the grid into separate
sub-tables, with the unbounded text emitted as plain text between them.

Single-cell "tables" (from the split) render as plain text instead of
a degenerate 1x1 markdown table.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 10:24:28 -07:00
Abimael MartellandClaude Opus 4.6 0300e96b7d fix: don't merge column header rows as continuations
Rows with 3+ short-valued cells (avg ≤10 chars) and an empty first cell
are column headers (e.g. "UR | SC | ST | OBC | EWS"), not text overflow
from the previous row. Prevents them from being merged into the
preceding section title row.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 02:24:52 -07:00
Abimael MartellandClaude Opus 4.6 1dcb0c6feb fix: improve table row splitting for stacked sub-tables
Three fixes for better table extraction from spreadsheet-exported PDFs:

1. Convert thin filled rects (< 2pt) to PdfLine objects before line-based
   table detection. Many PDFs draw table borders as narrow filled rectangles
   instead of stroked paths — these were invisible to the line detector.

2. Relax uniform row spacing rejection (CV 0.05 → 0.02). Spreadsheet
   exports have very even row heights that were being rejected as "chart
   grids".

3. Fix continuation row merging: don't merge rows where the only non-first
   cell content is a long label (section headers like "Category No. 03").
   Don't merge first-cell-only rows with long text ("Note: ...").

Also adds multi-Y row splitting in line-based detection and column-aware
table detection skipping for multi-column pages.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 02:20:42 -07:00
Abimael MartellandClaude Opus 4.6 44092bcc9e feat(tables): borderless table detection improvements (#17)
* feat(tables): improve heuristic detection for borderless wrapped-cell tables

Three changes to the body-font heuristic detector:

1. Adaptive Y-gap in find_table_regions_strict: use median qualifying-row
   spacing × 3 instead of fixed 25pt. Tables with wrapped cells have
   larger gaps between qualifying rows (those with 3+ X-clusters).

2. Y-only region filtering: use full X range when collecting region items.
   The strict X bounds from qualifying rows excluded continuation lines
   in wrapped cells, starving find_column_boundaries of items.

3. Merged-band retry: when split_side_by_side splits a page into bands
   but no band produces a table, retry heuristic detection on all items
   merged as a single band.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tables): gap-histogram column detection for small tables + lower avg_cells

Two changes to fix PDF 045 (borderless table with narrow "No." column):

1. Extend gap-histogram column threshold to small tables: when the gap
   between within-column jitter and between-column spacing is >10pt
   (unambiguous bimodal signal), use the detected threshold even with
   fewer than 500 items. Previously only triggered for dense tables.

2. Lower BodyFont avg_cells_per_row minimum from 2.5 to 2.0 to handle
   tables with wrapped multi-line cells where continuation lines have
   only 1 filled cell.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tables): trim empty outer columns + relax partial H-line validation

- Rect detection: trim empty first/last columns instead of rejecting
  the whole table. Rect edges often extend beyond text boundaries.
- Line detection: accept tables with 6+ partial horizontal lines
  (>15% width) when <3 full-spanning lines exist. Handles tables
  with column-level separators instead of full-width rules.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tables): cell-rect fallback for tables with variable-width backgrounds

When rect clustering produces a grid that fails validation (empty
interior columns from variable-width cell backgrounds), fall through
to a new strategy: use rect Y-edges for row boundaries and text
X-position clustering for columns. This handles tables like the
opendataloader-bench 088-090 comparison tables where each cell has
its own background rect at different widths.

Also widen failed-cluster hint width cap for large clusters (≥30 rects)
to allow page-spanning table regions.

TEDS score on opendataloader-bench: 0.300 → 0.353.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tables): relax vertical line spanning validation for partial borders

Accept tables with 4+ partial vertical lines (>10% table height) when
fewer than 2 span >30%. Handles tables like opendataloader-bench 053
with column-level vertical separators that don't extend the full height.

TEDS: 0.353 → 0.377 on opendataloader-bench.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tables): lower cell-rect density threshold + add validation logging

- Lower cell-rect density minimum from 25% to 15% to accept sparser
  tables with decorative backgrounds (fixes 147).
- Add debug logging to all heuristic validation paths for diagnosability.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tables): relax body-font validations for text-only and 2-column tables

Four fixes closing 71% of the TEDS gap vs opendataloader:

1. Validation 7 (table-like content): bypass numeric content requirement
   for tables with 3+ columns that passed all structural checks. Text-only
   tables (category lists, program descriptions) are legitimate.

2. Qualifying row threshold: lower from 3+ to 2+ X-clusters per row.
   Enables 2-column body-font table detection (fixes 166).

3. Row-stripe max cell length: raise from 500 to 2000 for 3+ column
   tables. Tables with paragraph descriptions in one column are valid
   (fixes 121).

4. Row-stripe empty-column trimming: apply the same outer-column trim
   as grid detection (fixes 121 column-0 rejection).

TEDS: 0.377 → 0.438 on opendataloader-bench (gap: -0.056 vs odl).
TEDS=0 docs: 14 → 9.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tables): enable 2-column body-font tables + lower all minimums

- Lower BodyFont minimum columns from 3 to 2 in detect_table_in_region
- Lower BodyFont minimum rows from 3 to 2
- Lower avg_cells_per_row minimum from 2.0 to 1.5 (handles wrapped cells
  in 2-column tables)
- Apply empty-outer-column trimming to row-stripe detection (not just grid)

TEDS: 0.438 → 0.468 on opendataloader-bench (gap: -0.027 vs odl).
TEDS=0 docs: 9 → 8. 86% of original gap closed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tables): text-based row fallback + fix 120 flow-chart and 188 leaderboard

Three changes that push TEDS past opendataloader:

1. Cell-rect Y-edge fallback: when rects have too few Y-edges for row
   structure, derive rows from text Y-position clustering within the
   rect bounding box. Fixes flow-chart tables (120) and column-header-
   only rects (188).

2. Lower cell-rect minimum from 20 to 6 rects to catch smaller tables.

3. Relax validation 1 (first-column presence) from 50% to 25% of rows.
   Tables with wrapped model names have continuation lines without first
   column content.

TEDS: 0.468 → 0.508 on opendataloader-bench.
Now BEATS opendataloader (0.508 vs 0.494, gap=+0.014).
TEDS=0 docs: 8 → 5.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tables): wrapped-cell continuation row merging

Merge rows that have fewer filled cells than the header row into the
previous row. Handles wrapped multi-line cells where text overflow
creates extra rows (e.g., "Direct" + "communications" → "Direct
communications").

Conditions: fewer filled cells than header, more than previous row had,
not a data row (numeric), not a short subheader label.

TEDS: 0.508 → 0.522 on opendataloader-bench (now +0.028 vs odl).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tables): tighten cell-rect validation + fix continuation-row merging

Cell-rect false positives:
- Raise density threshold back to 25% (from 15%)
- Add max cell length check (500 chars) to reject paragraph content
- Reject disproportionate grids (>20 rows, <4 cols)

Continuation-row merging:
- Wide tables (5+ cols): only merge rows with ≤50% header cells
- Narrow tables (2-4 cols): merge rows with fewer cells than header
- Prevents merging normal data rows in large tables (6_KE_Chart)
  while keeping wrapped-cell merging for narrow tables (178)

TEDS: 0.498 on opendataloader-bench (still +0.004 vs odl).
pdf-evals: 191/192 passed, 0 regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 16:57:51 -07:00
Abimael MartellandClaude Opus 4.6 80a3b81ff9 perf(tables): remove cosmetic padding from markdown tables
Drop column-width padding from table output. Saves ~37% bytes on
table-heavy PDFs. Markdown renders identically — padding was purely
cosmetic. Optimizes token efficiency for AI agent consumers.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 11:19:30 -07:00
Abimael MartellandClaude Opus 4.6 f2e3d51e49 fix(tables): cap column alignment width at 40 chars to reduce whitespace bloat
Tables with very wide cells (e.g. 230+ chars from merged schedule columns)
caused massive whitespace padding in every row. Capping alignment at 40
characters reduces output size significantly (franklin: 77KB→24KB,
Closed-Business: 571KB→502KB) without affecting content.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 18:02:49 -07:00
Abimael MartellandClaude Opus 4.6 6b8f4d0295 test: add inline unit tests to postprocess, format, grid, and detect_rects
Add 48 new tests (140 → 188 total) covering previously untested private
functions with synthetic data. No PDF fixtures needed.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 21:49:47 -07:00
Abimael MartellandClaude Opus 4.6 f7afc0c439 feat: improve rect-based table detection for wide statistical tables
- Add width-based outlier filter to remove page-spanning clipping paths
  without losing row-stripe background rects
- Deduplicate sub-rects (cell-internal decorations) to prevent spurious
  Y-edge splits, with height constraint to preserve row-stripe patterns
- Raise column limit from 12 to 25 for statistical lookup tables (MWU,
  chi-square)
- Skip propagate_merged_cells for wide tables (>10 cols) where spanning
  rects are background fills, not true merged cells
- Add numeric cell check to continuation-row heuristic so short text
  labels (e.g. "Liquid", "Vapour") are still merged while numeric data
  rows are kept separate
- Raise row-stripe content density threshold from 25% to 40% to reject
  false-positive tables from alternating-shade prose sections

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 11:05:58 -08:00
Abimael MartellandClaude Opus 4.6 dcbea3f652 Add row-stripe rect fallback for table detection
When PDF tables use full-width alternating row shading (row-stripe rects),
the normal rect grid detection fails because all rects share the same X/width,
collapsing to ~1 column. Add a fallback that uses rect Y-edges for rows and
text X-position clustering for columns, with a lower 15pt threshold to
separate narrow columns like row numbers and dates.

Also fix continuation-row merging in table formatting to not merge short
single-cell rows (≤5 chars) that are section sub-headers (e.g. month names).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:39:25 -08:00
Abimael MartellandClaude Opus 4.6 b6828fff8c fix(tables): Handle vertically-merged cells and protect header from continuation merge
Rects spanning multiple grid rows (e.g. classification labels) now have
their text consolidated into the first sub-row via propagate_merged_cells().
Also prevent continuation-row merge from folding body data into the header.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 16:40:57 -08:00
Abimael Martell 7328af91bd chore(refactor): Better organization of the codebase, split modules, update docs, add AGENTS.md 2026-02-18 10:16:11 -08:00