fix(tables): flatten page-number tables of contents instead of gridding (#142)

* fix(tables): flatten page-number tables of contents instead of gridding

A title-based contents page ("About the Publisher  vii", "Experiment #1
… 3") with no dot leaders and no section numbers was detected as a
2-column data table and rendered as a markdown grid, scrambling the
linear reading order (a top cause of NID loss on affected docs) and
scoring 0 on table structure.

Add is_page_number_toc: a narrow (2-3 col) list whose last column is
mostly page numbers (short integers or roman numerals) that are mostly
non-decreasing, with a text-title first column and NO header row (a
TOC's first row is already an entry). Such tables now route through the
existing flat-list TOC renderer.

The no-header + narrow-width + monotonic guards keep real data tables
intact — e.g. a 4-column regional table, or a 2-column "Mineral | CEC"
table with a header row and ascending values.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: bump reading-order (NID) benchmark to 0.89

Reflects the phantom-TOC fix in this PR: NID 0.88 -> 0.89 on the
200-doc benchmark. Other cells are unchanged at 2-decimal precision.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: canonical roman validation, real-first-row header check, roman page cells

- page_number_value now requires a *canonical* roman numeral (re-encode
  and compare), so words like "civil"/"mix"/"ill" are no longer parsed
  as page numbers.
- The no-header guard checks the actual first row's last cell instead of
  the first non-empty one, so a blank header cell ("Category | ") still
  rejects the TOC heuristic.
- format::is_page_number_cell recognizes canonical roman numerals, so
  roman front-matter pages (vii, ix) get proper title/page separation in
  the flat TOC list.
- Fix the non-monotonic test to use 5 rows so it exercises the
  monotonicity guard rather than the row-count early return; add
  roman-lookalike and blank-header rejection tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: share roman helper, widen length to 8, require page-span for TOC

- Extract canonical_roman_value + to_roman_lower into tables/mod.rs and
  use them from both the TOC detector and the formatter, removing the
  duplicated mapping/loop and keeping them in sync. The shared helper
  accepts ≤8 chars, so longer front-matter numerals (xxxviii) flatten
  consistently on both sides.
- Add a page-span guard to is_page_number_toc: real page numbers skip
  through the document (range >> entry count), so a dense consecutive
  ordinal/rank/ID column (1,2,3,…) is rejected — monotonicity alone did
  not separate those data tables from contents.

Costs ~0.001 aggregate on the benchmark (NID 0.888->0.887) for the added
precision; still a clear win over baseline (NID 0.883, TEDS unchanged).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: recover consecutive-page TOCs via a title signal

The strict page-span rule rejected legitimate one-page-per-entry TOCs
(range ~= entry count). Relax it: accept any page sequence with a gap
(nearly all real contents). Only a *perfectly dense* consecutive run —
which rank/ID/ordinal columns produce, but a chapter-per-page TOC can
too — falls back to a title signal: flatten when the first-column
entries average multi-word headings, keep as a table when they are the
short single-word labels typical of leaderboards/ID lists.

Recovers the ~0.001 the range-only rule cost (NID back to 0.888) while
still rejecting dense ordinal data tables.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Abimael Martell
2026-07-14 21:09:45 -07:00
committed by GitHub
co-authored by Claude Fable 5
parent 673fbe998f
commit 20bb22aa3c
5 changed files with 342 additions and 3 deletions
+1 -1
View File
@@ -27,7 +27,7 @@ Evaluated on the [opendataloader-bench](https://github.com/opendataloader-projec
| Engine | Overall | Reading Order (NID) | Tables (TEDS) | Headings (MHS) | Speed (200 docs) |
|---|---|---|---|---|---|
| pdf-inspector | 0.83 | 0.88 | 0.66 | 0.74 | 4s |
| pdf-inspector | 0.83 | 0.89 | 0.66 | 0.74 | 4s |
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
| markitdown | 0.58 | 0.88 | 0.00 | 0.00 | 8s |
+1 -1
View File
@@ -310,7 +310,7 @@
<tr><th>Engine</th><th>Overall</th><th>Reading order</th><th>Tables</th><th>Headings</th><th>200 docs</th></tr>
</thead>
<tbody>
<tr class="us"><td>pdf-inspector</td><td>0.83</td><td>0.88</td><td>0.66</td><td>0.74</td><td>4s</td></tr>
<tr class="us"><td>pdf-inspector</td><td>0.83</td><td>0.89</td><td>0.66</td><td>0.74</td><td>4s</td></tr>
<tr><td>opendataloader</td><td>0.84</td><td>0.91</td><td>0.49</td><td>0.74</td><td>11s</td></tr>
<tr><td>pymupdf4llm</td><td>0.73</td><td>0.89</td><td>0.40</td><td>0.41</td><td>18s</td></tr>
<tr><td>markitdown</td><td>0.58</td><td>0.88</td><td>0.00</td><td>0.00</td><td>8s</td></tr>
+282 -1
View File
@@ -914,7 +914,126 @@ fn looks_like_number(s: &str) -> bool {
///
/// Used by format.rs to render TOCs as flat lists instead of markdown tables.
pub fn is_table_of_contents(cells: &[Vec<String>]) -> bool {
is_dot_leader_toc(cells) || is_tabular_toc(cells)
is_dot_leader_toc(cells) || is_tabular_toc(cells) || is_page_number_toc(cells)
}
/// Parse a page-number-like token: a short arabic integer (≤4 digits) or a
/// canonical roman numeral (front-matter pages: i, ii, …, xxxviii). Roman
/// parsing is shared with the formatter via `super::canonical_roman_value` so
/// the two stay in sync.
fn page_number_value(token: &str) -> Option<u32> {
let t = token.trim();
if t.is_empty() {
return None;
}
if t.chars().all(|c| c.is_ascii_digit()) && t.len() <= 4 {
return t.parse().ok();
}
super::canonical_roman_value(t)
}
/// Page-number-column TOC: title-based contents with no dot leaders and no
/// section numbers (e.g. "About the Publisher vii", "Experiment #1 … 3").
/// The signature is a text-title first column and a last column that is almost
/// entirely page numbers whose values are *mostly non-decreasing* — the
/// monotonic run is what separates a real TOC from an incidental 2-column
/// numeric data table.
pub(super) fn is_page_number_toc(cells: &[Vec<String>]) -> bool {
let num_cols = cells.first().map(|r| r.len()).unwrap_or(0);
// A page-number TOC is a narrow list (title + page, optionally a leader
// column). Wider grids are data tables, not contents.
if !(2..=3).contains(&num_cols) || cells.len() < 5 {
return false;
}
let last = num_cols - 1;
// No header row: a TOC's first row is already an entry, so its last cell is
// a page number. A data table's first row is a column header (non-numeric,
// or an empty units cell like "Category | ") — the tell that separates
// "Mineral | CEC" tables from real contents. Check the actual first row,
// not the first non-empty one, so a blank header cell still rejects.
let first_last = cells[0].get(last).map(|s| s.trim()).unwrap_or("");
if page_number_value(first_last).is_none() {
return false;
}
// Last column: page numbers on ≥70% of filled rows; collect their values.
let mut filled = 0u32;
let mut page_vals: Vec<u32> = Vec::new();
for row in cells {
let cell = row.get(last).map(|s| s.trim()).unwrap_or("");
if cell.is_empty() {
continue;
}
filled += 1;
if let Some(v) = page_number_value(cell) {
page_vals.push(v);
}
}
if filled < 4 || (page_vals.len() as f32) < 0.7 * filled as f32 {
return false;
}
// First column: mostly text titles (has alphabetic content). This rejects
// numeric-vs-numeric grids.
let text_first = cells
.iter()
.filter(|row| {
row.first()
.is_some_and(|c| c.chars().any(|ch| ch.is_alphabetic()))
})
.count();
if (text_first as f32) < 0.6 * cells.len() as f32 {
return false;
}
// Page numbers mostly ascend (allow front-matter→body resets and noise).
if page_vals.len() < 2 {
return false;
}
let non_decreasing = page_vals.windows(2).filter(|w| w[1] >= w[0]).count();
if (non_decreasing as f32) < 0.7 * (page_vals.len() - 1) as f32 {
return false;
}
// Stronger TOC signal. Real page numbers SPAN the document — entries skip
// (3, 6, 13, 24, …) so their range exceeds the entry count. A rank / ID /
// ordinal column is instead a *perfectly dense* consecutive run (1,2,3,… or
// 100,101,102,…). Accept anything with page gaps; for a dense run — which a
// one-page-per-entry TOC can also produce — fall back to a title signal:
// real contents entries are multi-word headings, rank labels are short.
let min = *page_vals.iter().min().unwrap();
let max = *page_vals.iter().max().unwrap();
let span = max.saturating_sub(min);
if span > page_vals.len() as u32 {
return true;
}
let dense_consecutive = (span as usize) + 1 == page_vals.len() && {
let mut sorted = page_vals.clone();
sorted.sort_unstable();
sorted.dedup();
sorted.len() == page_vals.len()
};
if !dense_consecutive {
// Narrow range but with a gap or repeat — still contents-like.
return true;
}
// Dense counter: only a TOC if the titles read like headings, not the
// short single-word labels typical of rank/leaderboard/ID tables.
let (total_words, titled_rows) = cells
.iter()
.filter_map(|row| row.first())
.filter(|c| c.chars().any(|ch| ch.is_alphabetic()))
.fold((0usize, 0usize), |(w, n), c| {
(
w + c
.split_whitespace()
.filter(|t| t.chars().any(|ch| ch.is_alphabetic()))
.count(),
n + 1,
)
});
titled_rows > 0 && (total_words as f32) / titled_rows as f32 >= 1.8
}
/// Dot-leader TOC: any "Chapter 1 ........ 42" style with explicit leader
@@ -1886,4 +2005,166 @@ mod tests {
assert!(!starts_with_section_number(""));
assert!(!starts_with_section_number("Hello world"));
}
#[test]
fn page_number_value_rejects_roman_lookalike_words() {
// Ordinary words made only of {i,v,x,l,c} are not page numbers.
assert!(page_number_value("civil").is_none());
assert!(page_number_value("mix").is_none());
assert!(page_number_value("ill").is_none());
assert!(page_number_value("lil").is_none());
// Canonical roman numerals still parse.
assert_eq!(page_number_value("vii"), Some(7));
assert_eq!(page_number_value("ix"), Some(9));
assert_eq!(page_number_value("xii"), Some(12));
assert_eq!(page_number_value("42"), Some(42));
}
#[test]
fn page_number_toc_matches_consecutive_pages_with_titles() {
// A short chapter-per-page contents: pages are a dense 1..n run, but
// the multi-word titles mark it as a real TOC (recovered by the title
// signal rather than rejected for lacking page gaps).
let cells: Vec<Vec<String>> = vec![
vec!["Introduction to the Study".into(), "1".into()],
vec!["Materials and Methods".into(), "2".into()],
vec!["Results and Discussion".into(), "3".into()],
vec!["Summary of Findings".into(), "4".into()],
vec!["References and Notes".into(), "5".into()],
];
assert!(is_page_number_toc(&cells));
}
#[test]
fn page_number_toc_rejects_dense_ordinal_column() {
// Headerless title | rank table: values are a consecutive 1..n
// sequence (monotonic, no header, text first column) but their range
// ~= the row count, so it is data, not a table of contents.
let cells: Vec<Vec<String>> = vec![
vec!["Alice".into(), "1".into()],
vec!["Bob".into(), "2".into()],
vec!["Carol".into(), "3".into()],
vec!["Dave".into(), "4".into()],
vec!["Erin".into(), "5".into()],
vec!["Frank".into(), "6".into()],
];
assert!(!is_page_number_toc(&cells));
}
#[test]
fn page_number_toc_rejects_blank_header_cell() {
// First row is a header whose last cell is blank ("Category | ");
// must not be flattened even though later rows look TOC-like.
let cells = vec![
vec!["Category".into(), "".into()],
vec!["Alpha".into(), "3".into()],
vec!["Beta".into(), "9".into()],
vec!["Gamma".into(), "14".into()],
vec!["Delta".into(), "20".into()],
];
assert!(!is_page_number_toc(&cells));
}
#[test]
fn page_number_toc_matches_title_based_contents() {
// Title-left, page-number-right, no dot leaders, no section numbers.
let cells = vec![
vec!["About the Publisher".into(), "vii".into()],
vec!["About This Project".into(), "ix".into()],
vec!["Acknowledgments".into(), "xi".into()],
vec!["Experiment #1: Hydrostatic Pressure".into(), "3".into()],
vec!["Experiment #2: Bernoulli's Theorem".into(), "13".into()],
vec![
"Experiment #3: Energy Loss in Pipe Fittings".into(),
"24".into(),
],
];
assert!(is_page_number_toc(&cells));
assert!(is_table_of_contents(&cells));
}
#[test]
fn page_number_toc_rejects_numeric_data_table() {
// Real 2-col data table: numeric first column, non-monotonic values.
let cells = vec![
vec!["101".into(), "45".into()],
vec!["102".into(), "12".into()],
vec!["103".into(), "88".into()],
vec!["104".into(), "7".into()],
vec!["105".into(), "63".into()],
];
assert!(!is_page_number_toc(&cells));
}
#[test]
fn page_number_toc_rejects_non_monotonic_pages() {
// Text labels but the "page" column jumps around — a small data table,
// not a contents listing. 5 rows so the row-count guard passes and the
// monotonicity check is what does the rejecting.
let cells: Vec<Vec<String>> = vec![
vec!["Apples".into(), "42".into()],
vec!["Oranges".into(), "7".into()],
vec!["Pears".into(), "91".into()],
vec!["Plums".into(), "3".into()],
vec!["Grapes".into(), "60".into()],
];
// Sanity: this input clears the row-count and header guards, so a
// failure here is genuinely the monotonicity check.
assert!(cells.len() >= 5 && page_number_value(cells[0][1].trim()).is_some());
assert!(!is_page_number_toc(&cells));
}
#[test]
fn page_number_toc_rejects_header_row_data_table() {
// Real 2-col data table with a header row ("Mineral | CEC") and
// ascending values that mimic page numbers — the header tells us it
// is data, not contents.
let cells = vec![
vec![
"Mineral or colloid type".into(),
"CEC of pure colloid".into(),
],
vec!["kaolinite".into(), "10".into()],
vec!["illite".into(), "30".into()],
vec!["montmorillonite".into(), "100".into()],
vec!["vermiculite".into(), "150".into()],
];
assert!(!is_page_number_toc(&cells));
}
#[test]
fn page_number_toc_rejects_wide_data_grid() {
// A 4-column regional data table must not be read as a TOC even with a
// text first column and integer last column.
let cells = vec![
vec![
"REGIONS".into(),
"2007".into(),
"2010".into(),
"2016".into(),
],
vec![
"National Capital Region".into(),
"9".into(),
"8".into(),
"5".into(),
],
vec!["Cordillera".into(), "1".into(), "2".into(), "1".into()],
vec!["Ilocos Region".into(), "1".into(), "5".into(), "4".into()],
vec!["Cagayan Valley".into(), "1".into(), "3".into(), "5".into()],
];
assert!(!is_page_number_toc(&cells));
}
#[test]
fn page_number_toc_needs_page_number_last_column() {
// Last column is prose, not page numbers.
let cells = vec![
vec!["Section A".into(), "see appendix".into()],
vec!["Section B".into(), "see notes".into()],
vec!["Section C".into(), "later".into()],
vec!["Section D".into(), "TBD".into()],
];
assert!(!is_page_number_toc(&cells));
}
}
+4
View File
@@ -123,6 +123,7 @@ fn format_toc_as_list(cells: &[Vec<String>], footnotes: &[String]) -> String {
/// True when the cell looks like a page number. Accepts:
/// - plain digit tokens: "42", "86 86"
/// - canonical roman numerals (front-matter pages): "vii", "ix", "xii"
/// - dashed section-page IDs: "5-21", "A-1", "B--3", "TC-2" (common in
/// technical manuals)
fn is_page_number_cell(cell: &str) -> bool {
@@ -138,6 +139,9 @@ fn is_page_number_cell(cell: &str) -> bool {
if all_digits {
return t.len() <= 4;
}
if super::canonical_roman_value(t).is_some() {
return true;
}
// Section-page form: uppercase letters, digits, dashes; at least
// one digit present.
t.chars()
+54
View File
@@ -177,6 +177,60 @@ pub(crate) fn try_build_rect_guided_table(
))
}
/// Canonical lowercase roman numeral for `n` (the i/v/x/l/c range).
pub(super) fn to_roman_lower(mut n: u32) -> String {
const TABLE: [(u32, &str); 9] = [
(100, "c"),
(90, "xc"),
(50, "l"),
(40, "xl"),
(10, "x"),
(9, "ix"),
(5, "v"),
(4, "iv"),
(1, "i"),
];
let mut out = String::new();
for (val, sym) in TABLE {
while n >= val {
out.push_str(sym);
n -= val;
}
}
out
}
/// Parse a *canonical* roman numeral (i/v/x/l/c range, ≤8 chars) to its value.
/// Returns `None` for non-canonical strings, so ordinary words made of those
/// letters — "civil", "mix", "ill" — are not mistaken for numbers. Shared by
/// the TOC detector and the TOC formatter so the two stay in sync.
pub(super) fn canonical_roman_value(token: &str) -> Option<u32> {
let lower = token.trim().to_ascii_lowercase();
if lower.is_empty() || lower.len() > 8 || !lower.chars().all(|c| "ivxlc".contains(c)) {
return None;
}
let mut total = 0i32;
let mut prev = 0i32;
for c in lower.chars().rev() {
let v = match c {
'i' => 1,
'v' => 5,
'x' => 10,
'l' => 50,
'c' => 100,
_ => return None,
};
if v < prev {
total -= v;
} else {
total += v;
prev = v;
}
}
let value = u32::try_from(total).ok().filter(|&n| n > 0)?;
(to_roman_lower(value) == lower).then_some(value)
}
/// Split a TextItem whose text contains multiple whitespace-separated tokens
/// (like "10 11 12 ... 31") into individual TextItems, each assigned to the
/// nearest column boundary.