Compare commits

..
Author SHA1 Message Date
Abimael MartellandClaude Fable 5 573bc3cc90 review: block table roles from heading promotion too
Add Table/TR/TH/TD/THead/TBody/TFoot to is_non_heading_content. When
table reconstruction falls back and cells reach the line loop as plain
text, a short isolated cell (a TH column header in particular) could be
promoted to a heading. Defensive: no change across either regression
corpus, so pure hardening for the fallback path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 22:22:12 -07:00
Abimael MartellandClaude Fable 5 8951d90958 review: extend non-heading role gate; centralize on StructRole method
Move the non-heading-role check to StructRole::is_non_heading_content
and extend it to the content roles the inline allowlist missed: Quote,
Index, Note, Reference, BibEntry, Formula, Form (in addition to the
existing list/quote/caption/toc/code roles).

Figure is deliberately excluded: cover and banner pages routinely tag
the document title inside a Figure next to a seal/logo, and that title
is a real heading — including Figure demoted the LA County protocol
cover title from headings to bold. Verified against the reference.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 22:14:24 -07:00
Abimael MartellandClaude Fable 5 2349474432 fix(markdown): keep isolated headings on sparse pages; gate tagged roles
The isolated-line density guard wiped every isolated line on a page
where they exceeded 25% of lines. On sparse pages (covers, ToC pages
with a lone "CONTENTS" title, section-divider pages) a single heading
is trivially >25%, so the guard erased exactly the line it exists to
find. Require the page to have >=10 lines before the guard runs — the
25% ratio only signals a multi-column misfire on a dense page.

That let more isolated lines through, exposing that the visual heading
heuristic could promote lines already tagged with a non-heading struct
role (list item, blockquote, code, caption, ToC) or set in a monospace
font. Gate the heuristic on those in both converter paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 22:02:07 -07:00
Abimael MartellandClaude Fable 5 3a389f079e fix(markdown): ToC-page suppression, wrapped bold headings, math fragments (#131)
* fix(markdown): ToC-page suppression, wrapped bold headings, math fragments

Three heading-classification improvements:

- After emitting a "Contents"/"Table of Contents" heading, suppress
  heading promotion for the rest of that page: ToC entries are section
  titles that look exactly like headings ("1. Overview of OCR Pack")
  and whole contents pages came out as stacks of ##.
- merge_heading_lines only merged font-size-tier and struct-tree
  headings, so bold-at-body-size headings that wrap emitted two
  separate ## lines. Merge a fully-bold line into the previous
  fully-bold line when it reads as a wrap continuation (starts
  lowercase, tiny Y gap, no terminal punctuation on the previous line).
- Reject display-math fragments from the bold/rarity heading heuristic:
  equations ending in an equation number ("S = kB ln W, (2)") and
  lead-ins referencing one ("Rearranging Equation (8) gives:"). A bare
  trailing colon is deliberately NOT a signal — real headings often end
  with colons ("Procedure:").

The p1244 snapshot change is the bold-merge working as intended:
stacked form labels "**Subtotals** **from pages**" now read
"**Subtotals from pages**".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: guard tier path, protect prev headings in merge, narrow (N) rule

All three review findings applied, calibrated against the corpora:

- is_heading_fragment now gates the font-size-tier path too, not just
  the rarity heuristic.
- The bold wrap-merge requires the previous line to be tier-less as
  suggested; corpus diff confirmed the old behavior was absorbing a
  wrapped list-item fragment into a real heading.
- The bare "(N)" suffix rule suppressed real headings ("Nicaea (325)",
  appendix numbering). It now requires math evidence: an operator
  (=, <=, <<, ...) in the line or ,/: immediately before the number.
  Page-of-total running headers ("PM 2 (10)") get an explicit rule
  since the old blanket suffix check had been catching them only by
  accident.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 21:40:18 -07:00
Abimael MartellandClaude Fable 5 83754ddad0 fix(tables): reject row-stripe grids that swallow body text (#130)
* fix(tables): reject row-stripe grids that swallow body text

Charts (bar graphs, axis gridlines) emit fields of drawing rects that
pass the row-stripe shape test; the resulting phantom table then
captures the page's prose in scrambled reading order — losing headings
and paragraph flow with it. The existing max-cell-length gate only
fires for tables with <4 non-empty rows, which these grids exceed.

Add has_dominant_prose_cell: reject when one cell holds >=60 words AND
at least a third of the table's total words. Real tables never
concentrate that much text in a single cell, even with a description
column.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: extend prose-cell guard to merged-cluster path, add boundary test

The merged-cluster fallback had the same unguarded 4+-row gap as
row-stripe; apply has_dominant_prose_cell there too. The cell-rect path
already runs its own function-word prose check and is left unchanged.

Also add the boundary test from review: a 4+-row data table with one
60-word note cell stays accepted because the 1/3-of-total-words
denominator scales with table size.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: document intended small-table boundary of prose-cell guard

Investigated the suggested row-count exemption: adding a second-prose-
cell requirement (or a non_empty_rows >= 4 exemption) resurrects
verified phantom grids in 7 corpus documents — scrambled body text and
chart/figure regions, one of which is a 4-row grid. Every observed
single-dominant-cell grid in the corpora is swallowed prose, never a
real note table, and rejection degrades gracefully to prose while
acceptance scrambles reading order. Keep the guard unconditional,
document the rationale, and pin the boundary with a test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 20:02:23 -07:00
Abimael MartellandClaude Fable 5 f2f49bdac2 fix(markdown): one-word bold headings, block ToC entries from headings (#129)
Two heading-classification fixes:

- Accept single-word headings ("IMPLEMENTATION", "CONTENTS") when the
  line is all-bold and isolated; the word_count >= 2 gate rejected them
  unconditionally.
- Add is_toc_entry_line: a line ending in a dot-leader group plus page
  number ("Measurement Lab worksheet ... 3") is a table-of-contents
  entry, never a heading. has_dot_leaders misses single-group leaders,
  so entire ToC pages were being promoted to ## headings.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:30:33 -07:00
Abimael MartellandClaude Fable 5 d28594ef94 fix(fonts): recognize URW -Medi suffix as bold (#128)
URW Type 1 fonts abbreviate the Medium weight as "Medi" in the font
name (NimbusRomNo9L-Medi is the Times-Bold substitute embedded by most
LaTeX toolchains; -MediItal is bold italic). is_bold_font only matched
the full word "medium", so bold ran undetected across LaTeX-produced
PDFs — dropping ** emphasis and starving bold-based heading detection
of its signal.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:30:14 -07:00
6 changed files with 555 additions and 9 deletions
+164
View File
@@ -130,6 +130,108 @@ pub(crate) fn has_dot_leaders(text: &str) -> bool {
dot_groups >= 2
}
/// Detect a table-of-contents entry: a line ending in a page number preceded by
/// a dot-leader group (e.g. "Measurement Lab worksheet ... 3"). `has_dot_leaders`
/// misses single-group leaders ("..."), but a trailing "<dots> <number>" is a
/// strong TOC signal on its own. Such lines must never be promoted to headings.
pub(crate) fn is_toc_entry_line(text: &str) -> bool {
let trimmed = text.trim_end();
let digits = trimmed
.chars()
.rev()
.take_while(|c| c.is_ascii_digit())
.count();
if digits == 0 || digits > 4 {
return false;
}
let before_number = trimmed[..trimmed.len() - digits].trim_end();
let dots = before_number
.chars()
.rev()
.take_while(|c| *c == '.')
.count();
dots >= 3
}
/// A heading that announces a table of contents ("Contents", "Table of
/// Contents"). Lines after it on the same page are ToC entries — section
/// titles that look exactly like headings but must not be promoted.
pub(crate) fn is_toc_marker_heading(text: &str) -> bool {
let t = text.trim().trim_end_matches(':').trim().to_lowercase();
matches!(t.as_str(), "contents" | "table of contents")
}
/// Lines that resemble headings structurally but are display-math fragments:
/// equations ending in an equation number ("S = kB ln W, (2)") or equation
/// lead-ins ("Rearranging Equation (8) gives:"). Both carry an "(N)" equation
/// reference — but a trailing "(N)" alone is not enough: real headings end
/// with parenthesized numbers too ("Nicaea (325)", appendix numbering), so
/// the suffix form additionally requires math evidence — an "=" in the line
/// or a comma immediately before the number, both present in every display
/// equation and absent from name-plus-number headings. A bare trailing colon
/// is NOT a fragment signal either: real headings frequently end with colons
/// ("Procedure:", "Steps for Using the Microscope:").
pub(crate) fn is_heading_fragment(text: &str) -> bool {
let t = text.trim_end();
fn is_equation_number(s: &str) -> bool {
s.strip_prefix('(')
.and_then(|r| r.strip_suffix(')'))
.is_some_and(|inner| {
!inner.is_empty() && inner.len() <= 3 && inner.chars().all(|c| c.is_ascii_digit())
})
}
// Equation-number suffix with math evidence: "S = kB ln W, (2)"
let mut rev = t.rsplit(' ');
let last = rev.next().unwrap_or("");
if is_equation_number(last) {
// Page-of-total running headers: "LIVSMEDELSVERKET PM 2 (10)"
if let Some(prev_word) = t.rsplit(' ').nth(1) {
if let (Ok(page), Some(total)) = (
prev_word.parse::<u32>(),
last.trim_start_matches('(')
.trim_end_matches(')')
.parse::<u32>()
.ok(),
) {
if page <= total {
return true;
}
}
}
let punct_before = rev
.next()
.is_some_and(|w| w.ends_with(',') || w.ends_with(':'));
let has_math_op = t.chars().any(|c| {
matches!(
c,
'=' | '<'
| '>'
| '≤'
| '≥'
| '≪'
| '≫'
| '≈'
| '≠'
| '±'
| '∑'
| '∫'
| '√'
| '∝'
)
});
if punct_before || has_math_op {
return true;
}
}
// Lead-in: ends with a colon AND references an equation number inline
if t.ends_with(':') && t.split_whitespace().any(is_equation_number) {
return true;
}
false
}
/// Compute the Y-gap threshold for paragraph break detection.
///
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
@@ -320,3 +422,65 @@ pub(crate) fn detect_header_level(
Some(4)
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn toc_entry_with_single_dot_group() {
assert!(is_toc_entry_line("Measurement Lab worksheet ... 3"));
assert!(is_toc_entry_line("Results ........ 12"));
assert!(is_toc_entry_line("Appendix B...42"));
}
#[test]
fn non_toc_lines_pass() {
assert!(!is_toc_entry_line(
"6.2. Expectations for Re-Hiring Employees"
));
assert!(!is_toc_entry_line("What happened in 2020"));
assert!(!is_toc_entry_line("IMPLEMENTATION"));
// Ellipsis without a trailing page number
assert!(!is_toc_entry_line("and so it goes ..."));
// Long numbers are data, not page refs
assert!(!is_toc_entry_line("ISBN ... 97814"));
}
#[test]
fn toc_marker_headings() {
assert!(is_toc_marker_heading("Contents"));
assert!(is_toc_marker_heading("CONTENTS"));
assert!(is_toc_marker_heading("Table of Contents"));
assert!(is_toc_marker_heading("Table of contents:"));
assert!(!is_toc_marker_heading("Contents of the Shipment"));
assert!(!is_toc_marker_heading("Introduction"));
}
#[test]
fn heading_fragments() {
// Equation lead-ins: colon ending + inline equation reference
assert!(is_heading_fragment("Rearranging Equation (8) gives:"));
// Display-equation neighbours ending in an equation number
assert!(is_heading_fragment("S = kB ln W, (2)"));
assert!(is_heading_fragment("E = mc2 (12)"));
assert!(is_heading_fragment("x + y = z, (3)"));
// Page-of-total running headers
assert!(is_heading_fragment("LIVSMEDELSVERKET PM 2 (10)"));
// Comparison-operator evidence and colon-before-number
assert!(is_heading_fragment(
"PLL\u{fe} PHH\u{226a} PLH\u{fe} PHL: (12)"
));
// Real headings pass — including name-plus-number and colon-ended ones
assert!(!is_heading_fragment("Nicaea (325)"));
assert!(!is_heading_fragment(
"\u{627}\u{644}\u{645}\u{644}\u{62d}\u{642} \u{631}\u{642}\u{645} (1)"
));
assert!(!is_heading_fragment("4. Entropy"));
assert!(!is_heading_fragment("Procedure:"));
assert!(!is_heading_fragment("Steps for Using the Microscope:"));
assert!(!is_heading_fragment("Changing objectives:"));
assert!(!is_heading_fragment("Sales by Region (2024)"));
assert!(!is_heading_fragment("Results (preliminary)"));
}
}
+74 -5
View File
@@ -7,7 +7,8 @@ use crate::types::TextLine;
use super::analysis::{
bold_heading_level, calculate_font_stats, compute_heading_tiers, compute_paragraph_threshold,
detect_header_level, font_size_rarity, has_dot_leaders,
detect_header_level, font_size_rarity, has_dot_leaders, is_heading_fragment, is_toc_entry_line,
is_toc_marker_heading,
};
use super::classify::{
format_list_item, is_caption_line, is_list_item, is_monospace_font, starts_with_bullet_marker,
@@ -140,8 +141,11 @@ fn find_isolated_lines(lines: &[TextLine], base_size: f32, para_threshold: f32)
}
}
for (&page, &(total, isolated)) in &page_line_counts {
if total > 0 && isolated as f32 / total as f32 > 0.25 {
// Too many isolated lines on this page — remove them all
// The ratio only means something on pages dense enough for a
// multi-column misfire; on sparse pages (covers, ToC pages with a
// lone title) one isolated line is 25%+ of the page and exactly the
// line isolation exists to find.
if total >= 10 && isolated as f32 / total as f32 > 0.25 {
set.retain(|&i| lines[i].page != page);
}
}
@@ -486,6 +490,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
let mut in_code_block = false;
let mut prev_had_dot_leaders = false;
let mut paragraph_in_wrapped_bold_run = false;
let mut toc_suppress_page: Option<u32> = None;
let mut inserted_tables: HashSet<(u32, usize)> = HashSet::new();
let mut inserted_images: HashSet<(u32, usize)> = HashSet::new();
@@ -698,11 +703,22 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
_ => false,
};
// Lines explicitly tagged with a non-heading content role must never
// be promoted by the visual heuristic — a tagged list item, quote, or
// code line can look exactly like a heading (short, isolated).
let non_heading_role = struct_role
.as_ref()
.is_some_and(StructRole::is_non_heading_content);
let heuristic_heading = if options.detect_headers
&& !non_heading_role
&& !is_code_line
&& !looks_like_list_continuation
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !starts_with_bullet_marker(plain_trimmed)
&& !is_toc_entry_line(plain_trimmed)
&& !is_heading_fragment(plain_trimmed)
&& toc_suppress_page != Some(line.page)
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
detect_header_level(line_font_size, base_size, &heading_tiers).or_else(|| {
@@ -738,7 +754,11 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
// paragraph continuity and minor font-size variation
// inflates rarity scores.
let has_strong_signal = all_bold || isolated || (rarity >= 0.97 && word_count <= 8);
if score >= 0.5 && standalone && word_count >= 2 && has_strong_signal {
// Single-word headings ("IMPLEMENTATION", "CONTENTS") are common;
// accept them only with the strongest signal combination.
let enough_words =
word_count >= 2 || (all_bold && isolated && plain_trimmed.len() >= 4);
if score >= 0.5 && standalone && enough_words && has_strong_signal {
Some(bold_heading_level(&heading_tiers))
} else {
None
@@ -763,6 +783,9 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
plain_text.clone()
};
output.push_str(&format!("{} {}\n\n", prefix, heading_text.trim()));
if is_toc_marker_heading(plain_trimmed) {
toc_suppress_page = Some(line.page);
}
in_list = false;
continue;
}
@@ -959,6 +982,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
let mut last_list_x: Option<f32> = None;
let mut prev_had_dot_leaders = false;
let mut paragraph_in_wrapped_bold_run = false;
let mut toc_suppress_page: Option<u32> = None;
for (line_idx, line) in lines.iter().enumerate() {
// Page break
@@ -1037,6 +1061,10 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
if options.detect_headers
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !is_toc_entry_line(plain_trimmed)
&& !is_heading_fragment(plain_trimmed)
&& toc_suppress_page != Some(line.page)
&& !(options.detect_code && line.items.iter().any(|i| is_monospace_font(&i.font)))
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
if let Some(header_level) =
@@ -1059,7 +1087,9 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
+ if all_bold { 0.3 } else { 0.0 }
+ if standalone { 0.2 } else { 0.0 }
+ if isolated { 0.3 } else { 0.0 };
if score >= 0.5 && standalone && word_count >= 2 {
let enough_words =
word_count >= 2 || (all_bold && isolated && plain_trimmed.len() >= 4);
if score >= 0.5 && standalone && enough_words {
return Some(bold_heading_level(&heading_tiers));
}
None
@@ -1078,6 +1108,9 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
plain_text.clone()
};
output.push_str(&format!("{} {}\n\n", prefix, heading_text.trim()));
if is_toc_marker_heading(plain_trimmed) {
toc_suppress_page = Some(line.page);
}
in_list = false;
continue;
}
@@ -1206,6 +1239,42 @@ mod tests {
}
}
fn line_at(text: &str, page: u32, y: f32) -> TextLine {
let mut item = make_item(text, page, None);
item.y = y;
make_line(vec![item])
}
#[test]
fn isolated_lines_kept_on_sparse_pages() {
// A ToC page with a lone title and one entry far below: the density
// ratio is 50% but the page is too sparse for the multi-column
// misfire the guard targets — the title must stay isolated.
let lines = vec![
line_at("CONTENTS", 1, 700.0),
line_at("Chapter One 5", 1, 500.0),
];
let isolated = find_isolated_lines(&lines, 12.0, 20.0);
assert!(
isolated.contains(&0),
"sparse-page title must stay isolated"
);
}
#[test]
fn isolated_lines_wiped_on_dense_pages() {
// 12 short lines all with paragraph gaps — the multi-column misfire
// shape. The guard must clear them all.
let lines: Vec<TextLine> = (0..12)
.map(|i| line_at("Short column line", 1, 700.0 - i as f32 * 50.0))
.collect();
let isolated = find_isolated_lines(&lines, 12.0, 20.0);
assert!(
isolated.is_empty(),
"dense page of isolated lines must be wiped"
);
}
#[test]
fn test_struct_role_heading() {
let lines = vec![
+100 -1
View File
@@ -3,7 +3,7 @@
use std::collections::{HashMap, HashSet};
use crate::structure_tree::StructRole;
use crate::types::TextLine;
use crate::types::{TextItem, TextLine};
use super::analysis::detect_header_level;
@@ -87,6 +87,41 @@ pub(crate) fn merge_heading_lines(
false
};
// Bold headings at body font size never reach a tier, so wrapped ones
// split into two output headings ("…of wood pellets and cost" /
// "structure in Japan"). Merge a fully-bold line into the previous
// fully-bold line when it reads as a wrap continuation: starts
// lowercase, tiny Y gap, and the previous line has no terminal
// punctuation. Kept deliberately narrow — bold list labels and bold
// sentences start with markers or capitals and are unaffected.
let should_merge = should_merge
|| if let Some(prev) = result.last() {
let all_bold = |l: &TextLine| {
!l.items.is_empty() && l.items.iter().all(|i: &TextItem| i.is_bold)
};
let prev_text = prev.text();
let prev_trim = prev_text.trim_end();
let curr_text = line.text();
let curr_trim = curr_text.trim();
let y_gap = prev.y - line.y;
// Both lines must be tier-less: a tiered/tagged bold heading
// followed by bold body text must not absorb it.
line_level.is_none()
&& effective_heading_level(prev, base_size, heading_tiers, struct_roles)
.is_none()
&& prev.page == line.page
&& all_bold(prev)
&& all_bold(&line)
&& y_gap > 0.0
&& y_gap < line_font * 1.6
&& curr_trim.chars().next().is_some_and(|c| c.is_lowercase())
&& !prev_trim.ends_with(['.', ':', ';', '!', '?'])
&& prev_trim.split_whitespace().count() + curr_trim.split_whitespace().count()
<= 20
} else {
false
};
if should_merge {
// Append this line's items to the previous line
let prev = result.last_mut().unwrap();
@@ -685,4 +720,68 @@ mod tests {
.unwrap();
assert_eq!(first_header.page, 1, "first occurrence should be on page 1");
}
fn make_bold_line(text: &str, page: u32, y: f32) -> TextLine {
let mut item = make_item(text, 12.0, None);
item.is_bold = true;
TextLine {
items: vec![item],
y,
page,
adaptive_threshold: 0.10,
}
}
#[test]
fn merge_wrapped_bold_heading_lowercase_continuation() {
// Bold-at-body-size heading wrapped across two lines: the second line
// starts lowercase and must merge into the first.
let lines = vec![
make_bold_line(
"3. Perspective of supply and demand balance and cost",
1,
700.0,
),
make_bold_line("structure in Japan", 1, 686.0),
make_line("Body text paragraph follows here.", 12.0, 1, 660.0, None),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "wrapped bold heading should merge");
assert!(result[0].text().contains("cost structure in Japan"));
}
#[test]
fn no_merge_for_bold_sentences_or_new_headings() {
// Second bold line starts with a capital — a new heading or label,
// not a wrap continuation.
let lines = vec![
make_bold_line("Replace", 1, 700.0),
make_bold_line("Trash", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "distinct bold lines must not merge");
// Previous line ends a sentence — continuation must not merge.
let lines = vec![
make_bold_line("This is a bold sentence.", 1, 700.0),
make_bold_line("another bold line", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "sentence-final bold line must not merge");
}
#[test]
fn tiered_bold_heading_does_not_absorb_bold_body() {
// Previous line is a tier-level bold heading (16pt vs 12pt body);
// a following lowercase bold body line must NOT merge into it.
let mut heading = make_bold_line("Section Title", 1, 700.0);
heading.items[0].font_size = 16.0;
heading.items[0].height = 16.0;
let lines = vec![
heading,
make_bold_line("emphasized body text continues here", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[16.0], None);
assert_eq!(result.len(), 2, "tiered heading must not absorb bold body");
}
}
+94
View File
@@ -76,6 +76,52 @@ pub enum StructRole {
}
impl StructRole {
/// Content roles whose text must never be promoted to a heading by the
/// visual heuristic. These carry an explicit non-heading meaning in the
/// struct tree (lists, quotes, notes, references, captions, formulas,
/// forms, ToC entries), yet their text is often short and visually
/// isolated — exactly what the heuristic keys on. Heading roles (H, H1H6)
/// and generic container/flow roles (P, Div, Sect, Span, …) are excluded
/// so the heuristic can still fire there.
///
/// `Figure` is deliberately NOT in this set: cover/banner pages routinely
/// tag the document title inside a Figure (alongside a seal or logo), and
/// that title is a real heading. `Formula` and `Form` stay — a line
/// explicitly tagged as an equation or form field is never a heading.
///
/// Table roles (Table/TR/TH/TD/THead/TBody/TFoot) are included so that
/// when table reconstruction falls back and cells reach the line loop as
/// plain text, a short isolated cell — a `TH` column header especially —
/// is not promoted to a heading.
pub(crate) fn is_non_heading_content(&self) -> bool {
matches!(
self,
Self::L
| Self::LI
| Self::Lbl
| Self::LBody
| Self::BlockQuote
| Self::Quote
| Self::Caption
| Self::TOC
| Self::TOCI
| Self::Index
| Self::Note
| Self::Reference
| Self::BibEntry
| Self::Code
| Self::Formula
| Self::Form
| Self::Table
| Self::TR
| Self::TH
| Self::TD
| Self::THead
| Self::TBody
| Self::TFoot
)
}
fn from_name(name: &str) -> Self {
match name {
"Document" => Self::Document,
@@ -856,6 +902,54 @@ fn contains_bytes(haystack: &[u8], needle: &[u8]) -> bool {
mod tests {
use super::*;
#[test]
fn non_heading_content_roles() {
for r in [
StructRole::L,
StructRole::LI,
StructRole::BlockQuote,
StructRole::Quote,
StructRole::Caption,
StructRole::TOC,
StructRole::TOCI,
StructRole::Index,
StructRole::Note,
StructRole::Reference,
StructRole::BibEntry,
StructRole::Code,
StructRole::Formula,
StructRole::Form,
StructRole::Table,
StructRole::TR,
StructRole::TH,
StructRole::TD,
StructRole::THead,
StructRole::TBody,
StructRole::TFoot,
] {
assert!(
r.is_non_heading_content(),
"{r:?} should block heading promotion"
);
}
// Heading and generic container/flow roles must NOT block promotion
for r in [
StructRole::H,
StructRole::H1,
StructRole::H3,
StructRole::P,
StructRole::Div,
StructRole::Sect,
StructRole::Span,
StructRole::Figure,
] {
assert!(
!r.is_non_heading_content(),
"{r:?} should allow heading promotion"
);
}
}
#[test]
fn test_struct_role_from_name() {
assert_eq!(StructRole::from_name("H1"), StructRole::H1);
+121
View File
@@ -1477,6 +1477,10 @@ fn detect_row_stripe_table(
debug!(" row-stripe rejected: sparse outline/prose continuation shape");
return None;
}
if has_dominant_prose_cell(&cells) {
debug!(" row-stripe rejected: dominant prose cell (chart/figure region over body text)");
return None;
}
let column_centers: Vec<f32> = (0..num_cols)
.map(|c| (col_edges[c] + col_edges[c + 1]) / 2.0)
@@ -1495,6 +1499,34 @@ fn detect_row_stripe_table(
Some(Table::new(column_centers, row_centers, cells, item_indices))
}
/// Detect a grid that swallowed body text instead of tabular data.
///
/// Charts (bar graphs, axis gridlines) emit fields of drawing rects that can
/// pass the row-stripe shape test; the resulting "table" then captures the
/// page's prose. The signature: one cell holds an entire paragraph — ≥60 words
/// AND at least a third of all words in the table.
///
/// There is deliberately no row-count exemption. A small table whose single
/// long cell dominates its word count is indistinguishable by content from a
/// phantom grid over body text, and across the regression corpora every such
/// grid observed has been swallowed prose, never a real note table. The costs
/// are also asymmetric: rejecting a real table degrades it to readable prose,
/// while accepting a phantom scrambles the page into Y-interleaved cells.
/// Larger legitimate tables are safe because the one-third-of-total threshold
/// scales with table size.
fn has_dominant_prose_cell(cells: &[Vec<String>]) -> bool {
let mut total_words = 0usize;
let mut max_cell_words = 0usize;
for row in cells {
for cell in row {
let words = cell.split_whitespace().count();
total_words += words;
max_cell_words = max_cell_words.max(words);
}
}
max_cell_words >= 60 && max_cell_words * 3 >= total_words
}
fn row_stripe_is_sparse_prose_outline(cells: &[Vec<String>]) -> bool {
let Some(num_cols) = cells.first().map(|row| row.len()) else {
return false;
@@ -2294,6 +2326,12 @@ fn detect_merged_cluster_table(
);
return None;
}
if has_dominant_prose_cell(&cells) {
debug!(
" merged-cluster rejected: dominant prose cell (chart/figure region over body text)"
);
return None;
}
// No empty columns
for col in 0..num_cols {
@@ -2398,6 +2436,89 @@ mod tests {
}
}
// --- has_dominant_prose_cell ---
fn cells_of(rows: &[&[&str]]) -> Vec<Vec<String>> {
rows.iter()
.map(|r| r.iter().map(|c| c.to_string()).collect())
.collect()
}
#[test]
fn dominant_prose_cell_rejects_swallowed_paragraph() {
// Two cells hold paragraphs (the shape every observed phantom grid
// has: swallowed body text spans multiple cells), rest are chart labels
let para = ["word"; 70].join(" ");
let para2 = ["word"; 35].join(" ");
let cells = cells_of(&[
&[para.as_str(), "81", "76"],
&[para2.as_str(), "56", "9"],
&["2019", "2020", ""],
]);
assert!(has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_rejects_small_table_dominated_by_one_cell() {
// Boundary case, documented as INTENDED: a small grid whose single
// long cell dominates the word count is rejected even at 4+ rows.
// By content alone this shape is indistinguishable from a phantom
// grid over body text, and every observed instance in the regression
// corpora was swallowed prose (chart/figure regions), not a real
// note table. Rejection degrades gracefully — the text is still
// extracted as prose — while accepting a phantom scrambles reading
// order.
let note = ["word"; 70].join(" ");
let cells = cells_of(&[
&["Purpose", note.as_str()],
&["Owner", "Facilities team"],
&["Date", "2024-06-01"],
&["Status", "Active"],
]);
assert!(has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_allows_description_column() {
// Long-ish description cells, but text is spread across the table
let desc = ["word"; 25].join(" ");
let cells = cells_of(&[
&["Item A", desc.as_str(), "100"],
&["Item B", desc.as_str(), "200"],
&["Item C", desc.as_str(), "300"],
&["Item D", desc.as_str(), "400"],
]);
assert!(!has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_allows_short_tables() {
let cells = cells_of(&[&["Name", "Value"], &["Total", "42"]]);
assert!(!has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_allows_data_table_with_long_note() {
// A real 4+ row table with one verbose remark cell: the note is ≥60
// words but the table's other content carries more than 2× its word
// count, so concentration stays below the 1/3 threshold. The
// denominator scales with table size — this is what keeps large
// legitimate tables safe where a bare length cap would not.
let note = ["word"; 60].join(" ");
let row_text = ["data"; 12].join(" ");
let mut rows: Vec<Vec<String>> = (0..11)
.map(|i| {
vec![
format!("Item {i}"),
row_text.clone(),
format!("{}", i * 100),
]
})
.collect();
rows.push(vec!["Note".into(), note, String::new()]);
assert!(!has_dominant_prose_cell(&rows));
}
// --- rects_overlap ---
#[test]
+2 -3
View File
@@ -48,7 +48,7 @@ tips of directly from customers received other employees paid tips recd. entr
**Page 3**
27 28 29 30 31 **Subtotals** **from pages** **1, 2, and 3** **Totals**
27 28 29 30 31 **Subtotals from pages** **1, 2, and 3** **Totals**
**1.** Report total cash tips (col. **a**) on Form 4070, line **1.**
**2.** Report total credit card tips (col. **b**) on Form 4070, line **2.**
@@ -76,7 +76,6 @@ forms simpler, we would be happy to hear from you. You can write to the Tax Form
**Unreported Tips.—**If you received tips of $20 or more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you **must** use Form 1040 and **Form 4137,** Social Security and Medicare Tax on Unreported Tip Income, to report them. You may **not** use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act **cannot** use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—**Get **Pub. 531,** Reporting Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—**If you do not keep a daily record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
**Instructions** *(continued)*
### Instructions (continued)
Use this space to total your tips for the year