Compare commits

..
Author SHA1 Message Date
Abimael MartellandClaude Fable 5 4b1c7b21b6 review: guard tier path, protect prev headings in merge, narrow (N) rule
All three review findings applied, calibrated against the corpora:

- is_heading_fragment now gates the font-size-tier path too, not just
  the rarity heuristic.
- The bold wrap-merge requires the previous line to be tier-less as
  suggested; corpus diff confirmed the old behavior was absorbing a
  wrapped list-item fragment into a real heading.
- The bare "(N)" suffix rule suppressed real headings ("Nicaea (325)",
  appendix numbering). It now requires math evidence: an operator
  (=, <=, <<, ...) in the line or ,/: immediately before the number.
  Page-of-total running headers ("PM 2 (10)") get an explicit rule
  since the old blanket suffix check had been catching them only by
  accident.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 20:38:48 -07:00
Abimael MartellandClaude Fable 5 da1323a044 fix(markdown): ToC-page suppression, wrapped bold headings, math fragments
Three heading-classification improvements:

- After emitting a "Contents"/"Table of Contents" heading, suppress
  heading promotion for the rest of that page: ToC entries are section
  titles that look exactly like headings ("1. Overview of OCR Pack")
  and whole contents pages came out as stacks of ##.
- merge_heading_lines only merged font-size-tier and struct-tree
  headings, so bold-at-body-size headings that wrap emitted two
  separate ## lines. Merge a fully-bold line into the previous
  fully-bold line when it reads as a wrap continuation (starts
  lowercase, tiny Y gap, no terminal punctuation on the previous line).
- Reject display-math fragments from the bold/rarity heading heuristic:
  equations ending in an equation number ("S = kB ln W, (2)") and
  lead-ins referencing one ("Rearranging Equation (8) gives:"). A bare
  trailing colon is deliberately NOT a signal — real headings often end
  with colons ("Procedure:").

The p1244 snapshot change is the bold-merge working as intended:
stacked form labels "**Subtotals** **from pages**" now read
"**Subtotals from pages**".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 20:18:04 -07:00
Abimael MartellandClaude Fable 5 83754ddad0 fix(tables): reject row-stripe grids that swallow body text (#130)
* fix(tables): reject row-stripe grids that swallow body text

Charts (bar graphs, axis gridlines) emit fields of drawing rects that
pass the row-stripe shape test; the resulting phantom table then
captures the page's prose in scrambled reading order — losing headings
and paragraph flow with it. The existing max-cell-length gate only
fires for tables with <4 non-empty rows, which these grids exceed.

Add has_dominant_prose_cell: reject when one cell holds >=60 words AND
at least a third of the table's total words. Real tables never
concentrate that much text in a single cell, even with a description
column.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: extend prose-cell guard to merged-cluster path, add boundary test

The merged-cluster fallback had the same unguarded 4+-row gap as
row-stripe; apply has_dominant_prose_cell there too. The cell-rect path
already runs its own function-word prose check and is left unchanged.

Also add the boundary test from review: a 4+-row data table with one
60-word note cell stays accepted because the 1/3-of-total-words
denominator scales with table size.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: document intended small-table boundary of prose-cell guard

Investigated the suggested row-count exemption: adding a second-prose-
cell requirement (or a non_empty_rows >= 4 exemption) resurrects
verified phantom grids in 7 corpus documents — scrambled body text and
chart/figure regions, one of which is a 4-row grid. Every observed
single-dominant-cell grid in the corpora is swallowed prose, never a
real note table, and rejection degrades gracefully to prose while
acceptance scrambles reading order. Keep the guard unconditional,
document the rationale, and pin the boundary with a test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 20:02:23 -07:00
Abimael MartellandClaude Fable 5 f2f49bdac2 fix(markdown): one-word bold headings, block ToC entries from headings (#129)
Two heading-classification fixes:

- Accept single-word headings ("IMPLEMENTATION", "CONTENTS") when the
  line is all-bold and isolated; the word_count >= 2 gate rejected them
  unconditionally.
- Add is_toc_entry_line: a line ending in a dot-leader group plus page
  number ("Measurement Lab worksheet ... 3") is a table-of-contents
  entry, never a heading. has_dot_leaders misses single-group leaders,
  so entire ToC pages were being promoted to ## headings.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:30:33 -07:00
Abimael MartellandClaude Fable 5 d28594ef94 fix(fonts): recognize URW -Medi suffix as bold (#128)
URW Type 1 fonts abbreviate the Medium weight as "Medi" in the font
name (NimbusRomNo9L-Medi is the Times-Bold substitute embedded by most
LaTeX toolchains; -MediItal is bold italic). is_bold_font only matched
the full word "medium", so bold ran undetected across LaTeX-produced
PDFs — dropping ** emphasis and starving bold-based heading detection
of its signal.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:30:14 -07:00
Abimael MartellandClaude Fable 5 b4d401241e chore(napi): bump @firecrawl/pdf-inspector to 1.10.0 (#127)
New in this release (#125): isStrikeout on TextItem (geometric
detection sharing the underline rules pipeline), descriptor/embedded-
font bold+italic recall for subset fonts (FontDescriptor flags,
ttf-parser OS/2+post, bare-CFF Name INDEX), quote-operator advance
width, Ts text-rise handling, ActualText rise/position fixes, and a
document-scoped font style cache.


Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 17:41:46 -07:00
Abimael MartellandClaude Fable 5 57335f8bcf feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection (#125)
* feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection

Two style-recall gaps, both invisible to the existing name-based
heuristics:

1. Subset fonts with opaque BaseFont names ("Tc1", "AAAAAB+Amplitude")
   defeat is_italic_font/is_bold_font. New descriptor_style_flags reads
   the FontDescriptor (ItalicAngle beyond 4 degrees, Flags bit 7 Italic,
   bit 19 ForceBold) and, when the descriptor claims upright, falls back
   to the embedded font file: ttf-parser's OS/2 fsSelection + post
   italicAngle for sfnt fonts, and the CFF Name INDEX PostScript name
   for bare-CFF FontFile3 (descriptor rewritten to ItalicAngle 0 while
   embedding "Amplitude-LightItalic" was observed in the wild).
   ORed into is_bold/is_italic at item creation (content streams and
   form XObjects).

2. No strikeout signal existed. New is_strikeout on TextItem, detected
   in the same pass as underline: same rules pipeline (stroked lines /
   thin filled rects, table-ruling suppression), different vertical
   window — a rule crossing the glyphs at 12-55% of the em above the
   baseline instead of sitting at it. Exposed through napi and python
   bindings and pdf2md --items-json.

Verified on public ParseBench corpus docs: previously-missed italic
council titles and bold CJK itinerary headings now flagged (render-
checked); 24/508 docs gain flags, none lose any; 35 strikeout items
detected corpus-wide, disjoint from underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): quote-op advance width, Ts text rise, doc-level font style cache (PR #125 review)

Address three valid findings from review:

- The ' (move-to-next-line-and-show-text) operator emitted zero-width
  items and never advanced the text matrix, so geometric underline/
  strikeout detection (which requires width > 0) could never mark its
  text, and following show ops overlapped it. Reuse Tj's advance-width
  computation and matrix advance.

- Ts (text rise) was dropped entirely: raised/lowered runs kept the
  unshifted baseline, so rules drawn at the risen glyph position missed
  the strike/underline windows. Track rise in the text state (saved and
  restored with q/Q) and shift the rendering position through the text
  matrix's y column; advances stay on the unshifted matrix per spec.

- descriptor_style_flags re-decompressed and re-parsed the same embedded
  font program on every page whenever the descriptor left a style flag
  unset (the common case). Add a document-scoped FontStyleCache keyed by
  the FontFile2/FontFile3 object id, threaded through page and form
  extraction alongside the existing CMapDecisionCache.

The fourth finding (Form XObject rules never reach geometric detection)
is real but pre-existing for underline and needs the form walker to grow
path/paint tracking plus a new return type; deferred as a follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): ActualText items render at their glyphs' text rise (PR #125 review)

The EMC-built ActualText item used the captured text matrix without the
rise adjustment the ordinary Tj/TJ/' emission sites apply, so a tagged
run shown with Ts landed on the unshifted baseline — off the strikeout/
underline windows and inconsistent with untagged runs. The rise is
captured together with the first-glyph matrix (and at BDC for the
entry-position fallback): the item must render at the rise of its
GLYPHS, not whatever rise is set by EMC time.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): capture ActualText glyph position after the quote op's line move (PR #125 review)

The `'` handler skipped the entire suppressed-extraction block, so a
tagged span whose show op is `'` never captured its glyph matrix/rise —
the EMC item fell back to the BDC-entry matrix, which sits on the
PREVIOUS line (the `'` line move happens after BDC) with no rise. The
capture now happens right after the line move, matching the Tj/TJ
paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): style-boundary gate on subscript merge + strikeout suppression coverage (PR #125 review)

merge_subscript_items absorbed a script digit into its parent
regardless of underline/strikeout flags — dropping the digit's own mark
or widening the parent's over it. The merged item carries one flag, so
differing marks now break the merge, mirroring merge_text_items'
style-boundary rule (pre-existing for underline as well).

Also extends the table-suppression test to assert is_strikeout is
cleared alongside is_underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 17:30:47 -07:00
Abimael MartellandClaude Fable 5 15bc7894a4 docs: add MIT LICENSE file and license badge (#124)
Cargo.toml and pyproject.toml already declare MIT but the repo had no
LICENSE file, so GitHub and package registries couldn't display it.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 14:25:08 -07:00
28 changed files with 1239 additions and 43 deletions
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.9.11",
"version": "1.10.0",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
+4
View File
@@ -83,6 +83,9 @@ pub struct TextItem {
/// Underline detected geometrically (drawn rule/thin rect under the
/// baseline) — PDFs carry no underline font flag.
pub is_underline: bool,
/// Strikeout detected geometrically (rule crossing the glyphs at mid
/// x-height).
pub is_strikeout: bool,
pub item_type: ItemType,
/// URL for link items, `None` for other types.
pub link_url: Option<String>,
@@ -294,6 +297,7 @@ pub fn extract_text_with_positions(
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type,
link_url,
}
+1
View File
@@ -39,6 +39,7 @@ class TextItem:
is_bold: bool
is_italic: bool
is_underline: bool
is_strikeout: bool
item_type: str
class RegionText:
+3 -1
View File
@@ -74,7 +74,7 @@ fn format_items_json(items: &[TextItem]) -> String {
_ => String::new(),
};
format!(
r#"{{"text":"{}","page":{},"x":{:.2},"y":{:.2},"width":{:.2},"height":{:.2},"font":"{}","font_size":{:.2},"is_bold":{},"is_italic":{},"is_underline":{},"item_type":"{}","mcid":{}{}}}"#,
r#"{{"text":"{}","page":{},"x":{:.2},"y":{:.2},"width":{:.2},"height":{:.2},"font":"{}","font_size":{:.2},"is_bold":{},"is_italic":{},"is_underline":{},"is_strikeout":{},"item_type":"{}","mcid":{}{}}}"#,
json_escape(&item.text),
item.page,
item.x,
@@ -86,6 +86,7 @@ fn format_items_json(items: &[TextItem]) -> String {
item.is_bold,
item.is_italic,
item.is_underline,
item.is_strikeout,
item_type_label(&item.item_type),
mcid,
link_url,
@@ -122,6 +123,7 @@ mod tests {
is_bold: false,
is_italic: true,
is_underline: true,
is_strikeout: true,
item_type: ItemType::Text,
mcid: Some(7),
}];
+250 -20
View File
@@ -14,8 +14,9 @@ use lopdf::{Document, Encoding, Object, ObjectId};
use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
build_font_encodings, build_font_widths, compute_string_width_ts, descriptor_style_flags,
extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
FontStyleCache,
};
use super::underline::UnderlineLine;
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
@@ -106,6 +107,25 @@ fn transformed_stroke_width(
user_width * (ndx * ndx + ndy * ndy).sqrt()
}
/// Text rise (Ts) displaces the glyph origin by (0, rise) in unscaled text
/// space — per the rendering-matrix definition it sits left of Tm, so the
/// offset maps through the text matrix's y column. Rise never contributes
/// to the advance, so callers apply it only to the rendering position and
/// keep advancing the unshifted text matrix.
fn rise_adjusted(tm: &[f32; 6], rise: f32) -> [f32; 6] {
if rise == 0.0 {
return *tm;
}
[
tm[0],
tm[1],
tm[2],
tm[3],
tm[4] + rise * tm[2],
tm[5] + rise * tm[3],
]
}
/// Returns `(page_extraction, has_gid_fonts)` where `has_gid_fonts` indicates
/// the page uses fonts with unresolvable gid-encoded glyphs.
pub(crate) fn extract_page_text_items(
@@ -114,6 +134,7 @@ pub(crate) fn extract_page_text_items(
page_num: u32,
font_cmaps: &FontCMaps,
include_invisible: bool,
style_cache: &mut FontStyleCache,
) -> Result<(PageExtraction, bool, bool), PdfError> {
use lopdf::content::Content;
@@ -153,6 +174,8 @@ pub(crate) fn extract_page_text_items(
std::collections::HashMap::new();
let mut inline_cmaps: std::collections::HashMap<String, crate::tounicode::CMapEntry> =
std::collections::HashMap::new();
let mut font_style_flags: std::collections::HashMap<String, (bool, bool)> =
std::collections::HashMap::new();
for (font_name, font_dict) in &fonts {
let resource_name = String::from_utf8_lossy(font_name).to_string();
if let Ok(base_font) = font_dict.get(b"BaseFont") {
@@ -161,6 +184,12 @@ pub(crate) fn extract_page_text_items(
font_base_names.insert(resource_name.clone(), base_name);
}
}
// Descriptor style flags rescue subset fonts whose BaseFont names
// are opaque tags the name heuristics can't read.
let style = descriptor_style_flags(doc, font_dict, style_cache);
if style != (false, false) {
font_style_flags.insert(resource_name.clone(), style);
}
// Track ToUnicode object reference, with FontFile2 fallback for Identity-H/V.
// Also handle inline ToUnicode streams.
match font_dict.get(b"ToUnicode") {
@@ -235,6 +264,7 @@ pub(crate) fn extract_page_text_items(
line_width: f32,
char_spacing: f32,
word_spacing: f32,
text_rise: f32,
text_leading: f32,
current_font: String,
current_font_size: f32,
@@ -247,6 +277,7 @@ pub(crate) fn extract_page_text_items(
let mut text_leading: f32 = 0.0; // TL parameter (in text-space units)
let mut char_spacing: f32 = 0.0; // Tc parameter (extra spacing per character, unscaled)
let mut word_spacing: f32 = 0.0; // Tw parameter (extra spacing per space char, unscaled)
let mut text_rise: f32 = 0.0; // Ts parameter (baseline shift for super/subscripts, unscaled)
let mut text_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut line_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut in_text_block = false;
@@ -268,6 +299,10 @@ pub(crate) fn extract_page_text_items(
let mut suppress_glyph_extraction = false;
let mut actual_text_start_tm: Option<[f32; 6]> = None; // text matrix at BDC entry
let mut actual_text_glyph_tm: Option<[f32; 6]> = None; // text matrix at first glyph inside BDC
// Text rise in effect at each captured matrix — the item must render at
// the rise of its GLYPHS, not whatever rise is set by EMC time.
let mut actual_text_start_rise: f32 = 0.0;
let mut actual_text_glyph_rise: Option<f32> = None;
/// Get the innermost MCID from the marked content stack.
fn current_mcid(stack: &[MarkedContentEntry]) -> Option<i64> {
stack.iter().rev().find_map(|e| e.mcid)
@@ -284,6 +319,7 @@ pub(crate) fn extract_page_text_items(
line_width,
char_spacing,
word_spacing,
text_rise,
text_leading,
current_font: current_font.clone(),
current_font_size,
@@ -297,6 +333,7 @@ pub(crate) fn extract_page_text_items(
line_width = saved.line_width;
char_spacing = saved.char_spacing;
word_spacing = saved.word_spacing;
text_rise = saved.text_rise;
text_leading = saved.text_leading;
current_font = saved.current_font;
current_font_size = saved.current_font_size;
@@ -369,6 +406,12 @@ pub(crate) fn extract_page_text_items(
word_spacing = tw;
}
}
"Ts" => {
// Set text rise (baseline shift for superscripts/subscripts)
if let Some(ts) = op.operands.first().and_then(get_number) {
text_rise = ts;
}
}
"Td" | "TD" => {
// Move text position: TLM = T(tx,ty) × TLM; Tm = TLM
// tx,ty are in text space — must be scaled by the text line matrix
@@ -427,6 +470,7 @@ pub(crate) fn extract_page_text_items(
if suppress_glyph_extraction {
if actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
@@ -456,7 +500,8 @@ pub(crate) fn extract_page_text_items(
&mut cmap_decisions,
&font_widths,
) {
let combined = multiply_matrices(&text_matrix, &ctm);
let combined =
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
let (x, y) = (combined[4], combined[5]);
if combined[0].abs() >= combined[1].abs() {
@@ -478,6 +523,10 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
@@ -487,9 +536,10 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -507,6 +557,7 @@ pub(crate) fn extract_page_text_items(
// Capture first-glyph position for ActualText
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
// Compute space threshold based on font metrics when available
@@ -627,6 +678,10 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
let scale_x = text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2];
for (text, start_w, end_w) in &sub_items {
let offset_tm = [
@@ -637,7 +692,8 @@ pub(crate) fn extract_page_text_items(
text_matrix[4] + start_w * text_matrix[0],
text_matrix[5] + start_w * text_matrix[1],
];
let combined = multiply_matrices(&offset_tm, &ctm);
let combined =
multiply_matrices(&rise_adjusted(&offset_tm, text_rise), &ctm);
let (x, y) = (combined[4], combined[5]);
let width = if font_info.is_some() {
((end_w - start_w) * scale_x).abs()
@@ -653,9 +709,10 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -679,6 +736,26 @@ pub(crate) fn extract_page_text_items(
line_matrix[4] += (-tl) * line_matrix[2];
line_matrix[5] += (-tl) * line_matrix[3];
text_matrix = line_matrix;
// Capture first-glyph position for ActualText AFTER the
// line move — the BDC-entry matrix is on the previous line.
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
// Advance width, as for Tj — without it the item stays
// zero-width and geometric underline/strikeout detection
// rejects it (`is_underline_candidate` needs width > 0).
let w_ts_opt = font_widths.get(&current_font).and_then(|fi| {
op.operands.first().and_then(get_operand_bytes).map(|raw| {
compute_string_width_ts(
raw,
fi,
current_font_size,
char_spacing,
word_spacing,
)
})
});
if !((text_rendering_mode == 3 && !include_invisible)
|| suppress_glyph_extraction
|| op.operands.is_empty())
@@ -696,7 +773,8 @@ pub(crate) fn extract_page_text_items(
&font_widths,
) {
if !text.trim().is_empty() {
let combined = multiply_matrices(&text_matrix, &ctm);
let combined =
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
} else {
@@ -704,28 +782,45 @@ pub(crate) fn extract_page_text_items(
}
let rendered_size = effective_font_size(current_font_size, &combined);
let (x, y) = (combined[4], combined[5]);
let width = w_ts_opt
.map(|w_ts| {
(w_ts * (text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2]))
.abs()
})
.unwrap_or(0.0);
let base_font = font_base_names
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
y,
width: 0.0,
width,
height: rendered_size,
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
}
}
}
// Advance regardless of visibility so later show-text
// operators on the same line stay positioned (as for Tj).
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
}
}
"Do" => {
// XObject invocation - could be an image or form
@@ -757,6 +852,7 @@ pub(crate) fn extract_page_text_items(
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Image,
mcid: current_mcid(&marked_content_stack),
});
@@ -770,6 +866,7 @@ pub(crate) fn extract_page_text_items(
font_cmaps,
&ctm,
&mut cmap_decisions,
style_cache,
);
items.extend(form_items);
}
@@ -810,7 +907,9 @@ pub(crate) fn extract_page_text_items(
if actual_text.is_some() {
suppress_glyph_extraction = true;
actual_text_start_tm = Some(text_matrix);
actual_text_start_rise = text_rise;
actual_text_glyph_tm = None; // reset — will be captured at first Tj/TJ
actual_text_glyph_rise = None;
}
marked_content_stack.push(MarkedContentEntry { actual_text, mcid });
}
@@ -823,9 +922,11 @@ pub(crate) fn extract_page_text_items(
// Tj may have moved the text position to the correct line —
// the BDC-entry position can be on the previous line.
let glyph_tm = actual_text_glyph_tm.take();
let glyph_rise = actual_text_glyph_rise.take();
let entry_tm = actual_text_start_tm.take();
if let Some(start_tm) = glyph_tm.or(entry_tm) {
let combined = multiply_matrices(&start_tm, &ctm);
let rise = glyph_rise.unwrap_or(actual_text_start_rise);
let combined = multiply_matrices(&rise_adjusted(&start_tm, rise), &ctm);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
} else {
@@ -842,6 +943,10 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&at),
x,
@@ -851,9 +956,10 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: entry
.mcid
@@ -1376,8 +1482,15 @@ mod tests {
let (doc, page_id) = simple_doc_with_content(content);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) =
extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let ((items, _, _), _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
items
}
@@ -1462,6 +1575,108 @@ BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
assert!(!world.is_underline);
}
#[test]
fn quote_operator_text_carries_advance_width() {
// `'` (move-to-next-line-and-show-text) must retain the string's
// advance width like Tj — zero-width items are invisible to
// geometric underline/strikeout detection.
let content = b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (first) Tj (struck) ' ET
1 w
99 503 m 145 503 l S";
let items = extract_simple_items(content);
let struck = items.iter().find(|item| item.text == "struck").unwrap();
// 6 glyphs x 600/1000 x 12pt = 43.2pt, drawn one leading below Tm.
assert!((struck.width - 43.2).abs() < 0.1);
assert!((struck.y - 500.0).abs() < 0.1);
assert!(struck.is_strikeout);
assert!(!struck.is_underline);
}
#[test]
fn quote_operator_advances_text_matrix() {
// Text shown after `'` on the same line must start past the shown
// string: "CD" lands at x=114.4 (2 glyphs x 600/1000 x 12pt after
// x=100), flush against "AB", so the merge pass joins them. Without
// the advance "CD" overlaps "AB" at x=100 and the items stay apart.
let content = b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (AB) ' (CD) Tj ET";
let items = extract_simple_items(content);
let merged = items.iter().find(|item| item.text == "ABCD").unwrap();
assert!((merged.x - 100.0).abs() < 0.1);
assert!((merged.width - 28.8).abs() < 0.1);
assert!((merged.y - 500.0).abs() < 0.1);
}
#[test]
fn text_rise_shifts_item_baseline() {
// Ts displaces the glyph origin vertically without touching the
// advance; the next run at rise 0 must return to the original
// baseline and follow the raised run horizontally.
let content =
b"BT /F1 12 Tf 1 0 0 1 100 500 Tm (base) Tj 5 Ts (super) Tj 0 Ts (after) Tj ET";
let items = extract_simple_items(content);
let base = items.iter().find(|item| item.text == "base").unwrap();
let raised = items.iter().find(|item| item.text == "super").unwrap();
let after = items.iter().find(|item| item.text == "after").unwrap();
assert!((base.y - 500.0).abs() < 0.1);
assert!((raised.y - 505.0).abs() < 0.1);
assert!((after.y - 500.0).abs() < 0.1);
assert!(after.x > raised.x);
}
#[test]
fn actual_text_item_uses_glyph_rise() {
// The ActualText replacement item must render at the rise in
// effect when its glyphs were drawn — not the unshifted BDC
// baseline, and not whatever rise is set by EMC time.
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm \
/Span <</ActualText (super) >> BDC 5 Ts (sup) Tj 0 Ts EMC (after) Tj ET";
let items = extract_simple_items(content);
let sup = items.iter().find(|item| item.text == "super").unwrap();
let after = items.iter().find(|item| item.text == "after").unwrap();
assert!((sup.y - 505.0).abs() < 0.1);
assert!((after.y - 500.0).abs() < 0.1);
}
#[test]
fn actual_text_shown_with_quote_op_uses_moved_risen_baseline() {
// When the tagged span's show op is `'`, the glyph position is
// only known AFTER its line move — falling back to the BDC-entry
// matrix would place the item on the previous line, unrisen.
let content = b"BT /F1 12 Tf 14 TL 1 0 0 1 100 500 Tm \
/Span <</ActualText (replaced) >> BDC 3 Ts (raw) ' 0 Ts EMC ET";
let items = extract_simple_items(content);
let item = items.iter().find(|item| item.text == "replaced").unwrap();
// Line move: 500 - 14 = 486; rise: +3 -> 489.
assert!((item.y - 489.0).abs() < 0.1);
assert!(item.width > 0.0);
}
#[test]
fn strikeout_detected_on_risen_text() {
// The rule crosses the glyphs at their risen position; without the
// rise in item.y the strike window sits 4pt too low and misses.
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm 4 Ts (struck) Tj ET
1 w
99 507 m 145 507 l S";
let items = extract_simple_items(content);
let struck = items.iter().find(|item| item.text == "struck").unwrap();
assert!((struck.y - 504.0).abs() < 0.1);
assert!(struck.is_strikeout);
assert!(!struck.is_underline);
}
#[test]
fn test_skip_excessive_operations() {
use crate::tounicode::FontCMaps;
@@ -1496,7 +1711,15 @@ BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
doc.add_object(catalog);
let font_cmaps = FontCMaps::from_doc(&doc);
let result = extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let result = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
let ((items, rects, lines), _has_gid, _coords_rotated) = result;
assert!(items.is_empty());
assert!(rects.is_empty());
@@ -1578,8 +1801,15 @@ BT 30 700 Tm <41> Tj ET";
doc.trailer.set("Root", Object::Reference(catalog_id));
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) =
extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let ((items, _, _), _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
let text = items
.iter()
.map(|item| item.text.as_str())
+332 -1
View File
@@ -4,7 +4,7 @@ use crate::glyph_names::glyph_to_char;
use crate::tounicode::FontCMaps;
use crate::types::{FontEncodingMap, FontWidthInfo, PageFontEncodings, PageFontWidths};
use log::debug;
use lopdf::{Document, Encoding, Object};
use lopdf::{Document, Encoding, Object, ObjectId};
use std::collections::HashMap;
#[derive(Debug, Copy, Clone, PartialEq, Eq)]
@@ -739,6 +739,171 @@ pub(crate) fn get_font_file2_obj_num(doc: &Document, font_dict: &lopdf::Dictiona
.map(|r| r.0)
}
/// Document-scoped memo of embedded-font style flags, keyed by the
/// FontFile2/FontFile3 stream's object id. The same font program is
/// referenced from every page that uses the font, and decompressing +
/// parsing it dominates `descriptor_style_flags` — without the memo that
/// cost repeats per page whenever the descriptor leaves a flag unset
/// (the common case: regular fonts report neither italic nor bold).
#[derive(Debug, Default)]
pub(crate) struct FontStyleCache {
by_font_file: HashMap<ObjectId, (bool, bool)>,
}
impl FontStyleCache {
pub(crate) fn new() -> Self {
Self::default()
}
}
/// Style flags from the FontDescriptor, which survive subset fonts whose
/// BaseFont names are opaque tags ("Tc1", "ABCDEF+F1") that defeat the
/// name-based bold/italic heuristics.
///
/// Italic: `ItalicAngle` beyond a few degrees, or Flags bit 7 (Italic,
/// value 64). Bold: Flags bit 19 (ForceBold, value 1<<18). The small
/// ItalicAngle threshold skips fonts that declare a token slant.
pub(crate) fn descriptor_style_flags(
doc: &Document,
font_dict: &lopdf::Dictionary,
style_cache: &mut FontStyleCache,
) -> (bool, bool) {
let descriptor = font_dict
.get(b"FontDescriptor")
.ok()
.and_then(|obj| resolve_dict(doc, obj))
.or_else(|| {
// Type0 fonts hang the descriptor off DescendantFonts[0].
let desc_fonts = font_dict.get(b"DescendantFonts").ok()?;
let desc_fonts = resolve_array(doc, desc_fonts)?;
let cid_font_dict = resolve_dict(doc, desc_fonts.first()?)?;
resolve_dict(doc, cid_font_dict.get(b"FontDescriptor").ok()?)
});
let Some(descriptor) = descriptor else {
return (false, false);
};
let italic_angle = descriptor
.get(b"ItalicAngle")
.ok()
.and_then(|obj| match obj {
Object::Integer(i) => Some(*i as f32),
Object::Real(r) => Some(*r),
_ => None,
})
.unwrap_or(0.0);
let flags = descriptor
.get(b"Flags")
.ok()
.and_then(|obj| obj.as_i64().ok())
.unwrap_or(0);
let mut italic = italic_angle.abs() >= 4.0 || flags & (1 << 6) != 0;
let mut bold = flags & (1 << 18) != 0;
// Descriptors lie: subset generators write ItalicAngle 0 for genuinely
// italic faces. The embedded font file keeps the truth — OS/2
// fsSelection (via `Face::is_italic`) and the post table's italicAngle.
if !italic || !bold {
if let Some(ff_ref) = font_file_ref(descriptor) {
let (emb_italic, emb_bold) = *style_cache
.by_font_file
.entry(ff_ref)
.or_insert_with(|| embedded_style_flags(doc, ff_ref));
italic = italic || emb_italic;
bold = bold || emb_bold;
}
}
(italic, bold)
}
/// Style flags parsed from an embedded font program stream.
fn embedded_style_flags(doc: &Document, ff_ref: ObjectId) -> (bool, bool) {
let Some(data) = font_file_data(doc, ff_ref) else {
return (false, false);
};
if let Ok(face) = ttf_parser::Face::parse(&data, 0) {
(
face.is_italic() || face.italic_angle().abs() >= 4.0,
face.is_bold(),
)
} else if let Some(name) = cff_font_name(&data) {
// FontFile3 is bare CFF (no sfnt container) — ttf_parser
// can't open it, but the CFF Name INDEX keeps the real
// PostScript name ("XXXXXX+Amplitude-LightItalic") even
// when the descriptor was rewritten to claim upright.
(
crate::text_utils::is_italic_font(&name),
crate::text_utils::is_bold_font(&name),
)
} else {
(false, false)
}
}
/// First PostScript name from a bare CFF font's Name INDEX (CFF spec §7).
fn cff_font_name(data: &[u8]) -> Option<String> {
// Header: major(1) minor(1) hdrSize(1) offSize(1); major must be 1.
if data.len() < 4 || data[0] != 1 {
return None;
}
let hdr_size = data[2] as usize;
// Name INDEX: count(u16) offSize(u8) offsets[count+1] data
let count = u16::from_be_bytes([*data.get(hdr_size)?, *data.get(hdr_size + 1)?]) as usize;
if count == 0 {
return None;
}
let off_size = *data.get(hdr_size + 2)? as usize;
if !(1..=4).contains(&off_size) {
return None;
}
let read_offset = |idx: usize| -> Option<usize> {
let at = hdr_size + 3 + idx * off_size;
let bytes = data.get(at..at + off_size)?;
let mut v = 0usize;
for b in bytes {
v = (v << 8) | *b as usize;
}
Some(v)
};
let start = read_offset(0)?;
let end = read_offset(1)?;
if start == 0 || end < start {
return None;
}
// Offsets are 1-based from the byte before the object data.
let objects_base = hdr_size + 3 + (count + 1) * off_size - 1;
let name = data.get(objects_base + start..objects_base + end)?;
Some(String::from_utf8_lossy(name).to_string())
}
/// FontFile2/FontFile3 stream reference from a FontDescriptor.
fn font_file_ref(descriptor: &lopdf::Dictionary) -> Option<ObjectId> {
descriptor
.get(b"FontFile2")
.ok()
.and_then(|o| o.as_reference().ok())
.or_else(|| {
descriptor
.get(b"FontFile3")
.ok()
.and_then(|o| o.as_reference().ok())
})
}
/// Decompressed embedded font program bytes.
fn font_file_data(doc: &Document, ff_ref: ObjectId) -> Option<Vec<u8>> {
let stream = doc
.get_object(ff_ref)
.and_then(lopdf::Object::as_stream)
.ok()?;
Some(
stream
.decompressed_content()
.unwrap_or_else(|_| stream.content.clone()),
)
}
/// Decode text from a PDF string operand using font CMaps, encodings, and fallbacks.
#[allow(clippy::too_many_arguments)]
pub(crate) fn extract_text_from_operand(
@@ -1262,6 +1427,172 @@ mod tests {
}
}
fn doc_with_descriptor(descriptor: lopdf::Dictionary) -> (Document, lopdf::Dictionary) {
let mut doc = Document::with_version("1.4");
let desc_id = doc.add_object(descriptor);
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "TrueType",
"BaseFont" => "Tc1",
"FontDescriptor" => desc_id,
};
(doc, font_dict)
}
#[test]
fn descriptor_italic_angle_sets_italic() {
// Subset font with an opaque BaseFont name ("Tc1") — the name
// heuristic sees nothing, the descriptor carries the truth.
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => -12,
"Flags" => 32,
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(true, false)
);
}
#[test]
fn descriptor_italic_flag_bit_sets_italic() {
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => 0,
"Flags" => 64, // bit 7: Italic
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(true, false)
);
}
#[test]
fn descriptor_force_bold_flag_sets_bold() {
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => 0,
"Flags" => 1 << 18, // ForceBold
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, true)
);
}
#[test]
fn tiny_italic_angle_is_not_italic() {
// A token 1-degree slant is optical correction, not italic.
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => lopdf::Object::Real(-1.0),
"Flags" => 32,
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, false)
);
}
#[test]
fn missing_descriptor_yields_no_flags() {
let doc = Document::with_version("1.4");
let font_dict = dictionary! { "Type" => "Font", "BaseFont" => "Tc1" };
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, false)
);
}
#[test]
fn type0_descendant_descriptor_is_resolved() {
let mut doc = Document::with_version("1.4");
let desc_id = doc.add_object(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "ABCDEF+F1",
"ItalicAngle" => -15,
});
let cid_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "CIDFontType2",
"FontDescriptor" => desc_id,
});
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "Type0",
"BaseFont" => "ABCDEF+F1",
"DescendantFonts" => vec![lopdf::Object::Reference(cid_id)],
};
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(true, false)
);
}
/// Bare CFF: header + Name INDEX only — enough for `cff_font_name`.
fn bare_cff_with_name(name: &str) -> Vec<u8> {
let mut data = vec![1, 0, 4, 1]; // major, minor, hdrSize, offSize
data.extend_from_slice(&1u16.to_be_bytes()); // Name INDEX count
data.push(1); // offSize
data.push(1); // offset of first name
data.push(1 + name.len() as u8); // offset past last name
data.extend_from_slice(name.as_bytes());
data
}
#[test]
fn embedded_font_style_is_cached_by_font_file_object() {
use lopdf::{Object, Stream};
let mut doc = Document::with_version("1.4");
let ff_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
bare_cff_with_name("ABCDEF+Test-BoldItalic"),
)));
let desc_id = doc.add_object(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "ABCDEF+Test-BoldItalic",
"ItalicAngle" => 0,
"Flags" => 32,
"FontFile3" => ff_id,
});
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Tc1",
"FontDescriptor" => desc_id,
};
let mut cache = FontStyleCache::new();
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut cache),
(true, true)
);
assert_eq!(cache.by_font_file.len(), 1);
// Replace the font program with garbage: a repeat call must serve
// the memo instead of re-reading the stream — repeated per-page
// decompression is exactly what the cache exists to avoid.
doc.objects.insert(
ff_id,
Object::Stream(Stream::new(dictionary! {}, vec![0u8; 4])),
);
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut cache),
(true, true)
);
// A cold cache parses the (now garbage) stream, proving the warm
// call above answered from the memo.
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, false)
);
}
#[test]
fn compute_string_width_ts_no_tc_tw() {
// Without Tc/Tw (both 0), width = glyph widths only
+2
View File
@@ -1495,6 +1495,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1625,6 +1626,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
+2
View File
@@ -79,6 +79,7 @@ pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> V
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Link(url),
mcid: None,
});
@@ -318,6 +319,7 @@ pub(crate) fn walk_form_fields(
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::FormField,
mcid: None,
});
+71 -4
View File
@@ -24,6 +24,7 @@ use links::{extract_form_fields, extract_page_links};
// Re-export public types so existing `crate::extractor::X` paths keep working.
pub use crate::text_utils::{is_bold_font, is_italic_font};
pub use crate::types::{ItemType, TextLine};
pub(crate) use fonts::FontStyleCache;
pub(crate) use layout::detect_columns;
pub use layout::group_into_lines;
pub(crate) use layout::group_into_lines_with_thresholds;
@@ -161,6 +162,9 @@ fn extract_positioned_text_impl(
let mut all_lines = Vec::new();
let mut page_thresholds: PageThresholds = HashMap::new();
let mut gid_encoded_pages: HashSet<u32> = HashSet::new();
// Embedded-font style flags are document-scoped: the same font program
// is shared across pages, so parse it once, not once per page.
let mut style_cache = FontStyleCache::new();
// Build page ObjectId → page number map for form field extraction
let page_id_to_num: HashMap<ObjectId, u32> =
@@ -172,8 +176,14 @@ fn extract_positioned_text_impl(
continue;
}
}
let ((mut items, rects, lines), has_gid_fonts, _coords_rotated) =
extract_page_text_items(doc, page_id, *page_num, font_cmaps, include_invisible)?;
let ((mut items, rects, lines), has_gid_fonts, _coords_rotated) = extract_page_text_items(
doc,
page_id,
*page_num,
font_cmaps,
include_invisible,
&mut style_cache,
)?;
if has_gid_fonts {
gid_encoded_pages.insert(*page_num);
}
@@ -234,7 +244,10 @@ fn suppress_table_underlines(
lines: &[PdfLine],
page: u32,
) {
if !items.iter().any(|item| item.is_underline) {
if !items
.iter()
.any(|item| item.is_underline || item.is_strikeout)
{
return;
}
@@ -256,6 +269,7 @@ fn suppress_table_underlines(
for index in table_item_indices {
if let Some(item) = items.get_mut(index) {
item.is_underline = false;
item.is_strikeout = false;
}
}
}
@@ -576,6 +590,7 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
if next.is_bold != first.is_bold
|| next.is_italic != first.is_italic
|| next.is_underline != first.is_underline
|| next.is_strikeout != first.is_strikeout
{
break;
}
@@ -638,6 +653,7 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
is_bold: first.is_bold,
is_italic: first.is_italic,
is_underline: first.is_underline,
is_strikeout: first.is_strikeout,
item_type: first.item_type.clone(),
mcid: first.mcid,
});
@@ -715,7 +731,9 @@ pub(crate) fn merge_subscript_items(items: Vec<TextItem>) -> Vec<TextItem> {
.chars()
.last()
.is_some_and(|c| c.is_alphabetic());
if parent.font_size >= sub_threshold && ends_with_letter {
let same_marks = parent.is_underline == item.is_underline
&& parent.is_strikeout == item.is_strikeout;
if parent.font_size >= sub_threshold && ends_with_letter && same_marks {
let parent_right = parent.x + parent.width;
let gap = item.x - parent_right;
// Subscripts must be tightly adjacent (within ~1pt)
@@ -789,6 +807,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -991,6 +1010,7 @@ mod tests {
items[3].y = 470.0;
for item in &mut items {
item.is_underline = true;
item.is_strikeout = true;
}
let lines = vec![
make_line(100.0, 500.0, 300.0, 500.0),
@@ -1004,6 +1024,30 @@ mod tests {
suppress_table_underlines(&mut items, &[], &lines, 1);
assert!(items.iter().all(|item| !item.is_underline));
assert!(items.iter().all(|item| !item.is_strikeout));
}
#[test]
fn subscript_digit_with_different_marks_is_not_absorbed() {
// A struck-out word followed by an unmarked footnote digit: merging
// would widen the parent's strikeout claim over the digit (and the
// reverse would drop the digit's own mark). Style boundaries break
// the merge, as in merge_text_items.
let mut word = make_merge_item("word", 100.0, 24.0);
word.font_size = 10.0;
word.is_strikeout = true;
let mut digit = make_merge_item("2", 124.5, 4.0);
digit.font_size = 6.0;
digit.y = word.y + 3.0;
let merged = merge_subscript_items(vec![word.clone(), digit.clone()]);
assert_eq!(merged.len(), 2);
// Same marks still merge (footnote ref inside the strike).
digit.is_strikeout = true;
let merged = merge_subscript_items(vec![word, digit]);
assert_eq!(merged.len(), 1);
assert!(merged[0].text.starts_with("word"));
}
#[test]
@@ -1021,6 +1065,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1036,6 +1081,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1051,6 +1097,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1106,6 +1153,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1121,6 +1169,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1136,6 +1185,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1162,6 +1212,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1177,6 +1228,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1192,6 +1244,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1220,6 +1273,7 @@ mod tests {
is_bold: true,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1255,6 +1309,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1291,6 +1346,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1306,6 +1362,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1321,6 +1378,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1344,6 +1402,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1457,6 +1516,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1472,6 +1532,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1497,6 +1558,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1512,6 +1574,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1553,6 +1616,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1598,6 +1662,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1643,6 +1708,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1681,6 +1747,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+64 -1
View File
@@ -231,8 +231,27 @@ fn rule_matches_item(rule: &Rule, item: &TextItem) -> bool {
overlap >= min_overlap
}
/// Strikeout window: a rule crossing the glyphs. Strikethroughs sit at
/// roughly 20-35% of the em above the baseline (about half the x-height);
/// accept a band well inside the glyph body so baseline underlines and
/// overlines never qualify.
fn rule_strikes_item(rule: &Rule, item: &TextItem) -> bool {
let y_min = item.y + item.font_size * 0.12;
let y_max = item.y + item.font_size * 0.55;
if rule.y < y_min || rule.y > y_max {
return false;
}
let ix1 = item.x;
let ix2 = item.x + item.width;
let min_overlap = item.width * MIN_X_OVERLAP;
let overlap = rule.x2.min(ix2) - rule.x1.max(ix1);
overlap >= min_overlap
}
/// Mark `is_underline` on text items that have a horizontal rule just
/// below their baseline. `items`, `rects`, and `lines` are a single
/// below their baseline, and `is_strikeout` on items whose glyphs a rule
/// crosses at mid x-height. `items`, `rects`, and `lines` are a single
/// page's extraction output (all in PDF coordinates, y-up, where
/// `TextItem::y` is the text baseline).
pub(crate) fn mark_underlined_items(
@@ -258,6 +277,11 @@ pub(crate) fn mark_underlined_items(
}
if rule_matches_item(rule, item) {
item.is_underline = true;
}
if rule_strikes_item(rule, item) {
item.is_strikeout = true;
}
if item.is_underline && item.is_strikeout {
break;
}
}
@@ -282,6 +306,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -358,6 +383,44 @@ mod tests {
assert!(!items[0].is_underline);
}
#[test]
fn mid_glyph_rule_marks_strikeout_not_underline() {
// Rule at ~30% of the em above the baseline crosses the glyphs.
let mut items = vec![item("struck out", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 503.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items[0].is_strikeout);
assert!(!items[0].is_underline);
}
#[test]
fn baseline_rule_marks_underline_not_strikeout() {
let mut items = vec![item("underlined", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 498.5)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items[0].is_underline);
assert!(!items[0].is_strikeout);
}
#[test]
fn overline_is_neither_underline_nor_strikeout() {
// Rule just above the cap height (overline / next line's rule).
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 507.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(!items[0].is_underline);
assert!(!items[0].is_strikeout);
}
#[test]
fn thin_filled_rect_at_mid_glyph_marks_strikeout() {
let mut items = vec![item("struck out", 100.0, 500.0, 60.0, 10.0)];
let rects = vec![thin_rect(100.0, 502.6, 60.0)];
mark_underlined_items(&mut items, &rects, &[], 1);
assert!(items[0].is_strikeout);
assert!(!items[0].is_underline);
}
#[test]
fn line_above_baseline_is_not_an_underline() {
// Strikethrough / overline geometry must not mark.
+27 -5
View File
@@ -1,5 +1,6 @@
//! Form XObject and image XObject extraction.
use super::fonts::descriptor_style_flags;
use crate::text_utils::{effective_font_size, expand_ligatures, is_bold_font, is_italic_font};
use crate::tounicode::FontCMaps;
use crate::types::{ItemType, TextItem};
@@ -8,7 +9,7 @@ use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache, FontStyleCache,
};
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
@@ -114,6 +115,7 @@ pub(crate) fn extract_form_xobject_text(
font_cmaps: &FontCMaps,
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
) -> Vec<TextItem> {
extract_form_xobject_text_inner(
doc,
@@ -122,10 +124,12 @@ pub(crate) fn extract_form_xobject_text(
font_cmaps,
parent_ctm,
cmap_decisions,
style_cache,
0,
)
}
#[allow(clippy::too_many_arguments)]
fn extract_form_xobject_text_inner(
doc: &Document,
form_id: ObjectId,
@@ -133,6 +137,7 @@ fn extract_form_xobject_text_inner(
font_cmaps: &FontCMaps,
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
depth: u8,
) -> Vec<TextItem> {
use lopdf::content::Content;
@@ -167,6 +172,7 @@ fn extract_form_xobject_text_inner(
let mut font_tounicode_refs: HashMap<String, u32> = HashMap::new();
let mut inline_cmaps: HashMap<String, crate::tounicode::CMapEntry> = HashMap::new();
let mut font_style_flags: HashMap<String, (bool, bool)> = HashMap::new();
for (font_name, font_dict) in &form_fonts {
let resource_name = String::from_utf8_lossy(font_name).to_string();
if let Ok(base_font) = font_dict.get(b"BaseFont") {
@@ -175,6 +181,10 @@ fn extract_form_xobject_text_inner(
font_base_names.insert(resource_name.clone(), base_name);
}
}
let style = descriptor_style_flags(doc, font_dict, style_cache);
if style != (false, false) {
font_style_flags.insert(resource_name.clone(), style);
}
match font_dict.get(b"ToUnicode") {
Ok(tounicode) => {
if let Ok(obj_ref) = tounicode.as_reference() {
@@ -272,6 +282,7 @@ fn extract_form_xobject_text_inner(
font_cmaps,
&ctm,
cmap_decisions,
style_cache,
depth + 1,
);
items.extend(nested_items);
@@ -295,6 +306,7 @@ fn extract_form_xobject_text_inner(
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Image,
mcid: None,
});
@@ -428,6 +440,10 @@ fn extract_form_xobject_text_inner(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
@@ -437,9 +453,10 @@ fn extract_form_xobject_text_inner(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -560,6 +577,10 @@ fn extract_form_xobject_text_inner(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
let scale_x = text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2];
for (text, start_w, end_w) in &sub_items {
let offset_tm = [
@@ -586,9 +607,10 @@ fn extract_form_xobject_text_inner(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
+29 -4
View File
@@ -579,6 +579,7 @@ pub fn extract_text_in_regions_mem(
let mut gid_pages: HashSet<u32> = HashSet::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
let mut style_cache = extractor::FontStyleCache::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
@@ -597,6 +598,7 @@ pub fn extract_text_in_regions_mem(
*page_num,
&font_cmaps,
false,
&mut style_cache,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -699,6 +701,7 @@ pub fn extract_tables_in_regions_mem(
let mut gid_pages: HashSet<u32> = HashSet::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
let mut style_cache = extractor::FontStyleCache::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
@@ -714,6 +717,7 @@ pub fn extract_tables_in_regions_mem(
*page_num,
&font_cmaps,
false,
&mut style_cache,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -1019,6 +1023,7 @@ pub fn detect_vector_grid_in_region_mem(
page_1idx,
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
)?;
text_utils::fix_letterspaced_items(&mut items);
@@ -1205,8 +1210,15 @@ mod vector_grid_tests {
let &page_id = pages.get(&1).unwrap();
let needed: HashSet<u32> = HashSet::from([1]);
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
let ((items, rects, _lines), _has_gid, _rotated) =
extract_page_text_items(&doc, page_id, 1, &cmaps, false).unwrap();
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
&doc,
page_id,
1,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
)
.unwrap();
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, 1);
assert_eq!(rect_tables.len(), 1, "expected one rect-detected table");
@@ -1240,8 +1252,15 @@ mod vector_grid_tests {
let &page_id = pages.get(&page_num).unwrap();
let needed: HashSet<u32> = HashSet::from([page_num]);
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
let ((items, rects, _lines), _has_gid, _rotated) =
extract_page_text_items(&doc, page_id, page_num, &cmaps, false).unwrap();
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
&doc,
page_id,
page_num,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
)
.unwrap();
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, page_num);
rect_tables
@@ -1966,6 +1985,7 @@ pub fn extract_tables_with_structure_cells_mem(
let mut page_heights: HashMap<u32, f32> = HashMap::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
let mut style_cache = extractor::FontStyleCache::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
@@ -1981,6 +2001,7 @@ pub fn extract_tables_with_structure_cells_mem(
*page_num,
&font_cmaps,
false,
&mut style_cache,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -2782,6 +2803,7 @@ fn detect_tsr_quality_issue(
page_1idx,
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
)?;
let adaptive_threshold = text_utils::fix_letterspaced_items(&mut items);
let coords = if coords_rotated {
@@ -5056,6 +5078,7 @@ mod text_cluster_column_undercount_tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -5331,6 +5354,7 @@ mod table_candidate_selection_tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -6078,6 +6102,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+164
View File
@@ -130,6 +130,108 @@ pub(crate) fn has_dot_leaders(text: &str) -> bool {
dot_groups >= 2
}
/// Detect a table-of-contents entry: a line ending in a page number preceded by
/// a dot-leader group (e.g. "Measurement Lab worksheet ... 3"). `has_dot_leaders`
/// misses single-group leaders ("..."), but a trailing "<dots> <number>" is a
/// strong TOC signal on its own. Such lines must never be promoted to headings.
pub(crate) fn is_toc_entry_line(text: &str) -> bool {
let trimmed = text.trim_end();
let digits = trimmed
.chars()
.rev()
.take_while(|c| c.is_ascii_digit())
.count();
if digits == 0 || digits > 4 {
return false;
}
let before_number = trimmed[..trimmed.len() - digits].trim_end();
let dots = before_number
.chars()
.rev()
.take_while(|c| *c == '.')
.count();
dots >= 3
}
/// A heading that announces a table of contents ("Contents", "Table of
/// Contents"). Lines after it on the same page are ToC entries — section
/// titles that look exactly like headings but must not be promoted.
pub(crate) fn is_toc_marker_heading(text: &str) -> bool {
let t = text.trim().trim_end_matches(':').trim().to_lowercase();
matches!(t.as_str(), "contents" | "table of contents")
}
/// Lines that resemble headings structurally but are display-math fragments:
/// equations ending in an equation number ("S = kB ln W, (2)") or equation
/// lead-ins ("Rearranging Equation (8) gives:"). Both carry an "(N)" equation
/// reference — but a trailing "(N)" alone is not enough: real headings end
/// with parenthesized numbers too ("Nicaea (325)", appendix numbering), so
/// the suffix form additionally requires math evidence — an "=" in the line
/// or a comma immediately before the number, both present in every display
/// equation and absent from name-plus-number headings. A bare trailing colon
/// is NOT a fragment signal either: real headings frequently end with colons
/// ("Procedure:", "Steps for Using the Microscope:").
pub(crate) fn is_heading_fragment(text: &str) -> bool {
let t = text.trim_end();
fn is_equation_number(s: &str) -> bool {
s.strip_prefix('(')
.and_then(|r| r.strip_suffix(')'))
.is_some_and(|inner| {
!inner.is_empty() && inner.len() <= 3 && inner.chars().all(|c| c.is_ascii_digit())
})
}
// Equation-number suffix with math evidence: "S = kB ln W, (2)"
let mut rev = t.rsplit(' ');
let last = rev.next().unwrap_or("");
if is_equation_number(last) {
// Page-of-total running headers: "LIVSMEDELSVERKET PM 2 (10)"
if let Some(prev_word) = t.rsplit(' ').nth(1) {
if let (Ok(page), Some(total)) = (
prev_word.parse::<u32>(),
last.trim_start_matches('(')
.trim_end_matches(')')
.parse::<u32>()
.ok(),
) {
if page <= total {
return true;
}
}
}
let punct_before = rev
.next()
.is_some_and(|w| w.ends_with(',') || w.ends_with(':'));
let has_math_op = t.chars().any(|c| {
matches!(
c,
'=' | '<'
| '>'
| '≤'
| '≥'
| '≪'
| '≫'
| '≈'
| '≠'
| '±'
| '∑'
| '∫'
| '√'
| '∝'
)
});
if punct_before || has_math_op {
return true;
}
}
// Lead-in: ends with a colon AND references an equation number inline
if t.ends_with(':') && t.split_whitespace().any(is_equation_number) {
return true;
}
false
}
/// Compute the Y-gap threshold for paragraph break detection.
///
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
@@ -320,3 +422,65 @@ pub(crate) fn detect_header_level(
Some(4)
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn toc_entry_with_single_dot_group() {
assert!(is_toc_entry_line("Measurement Lab worksheet ... 3"));
assert!(is_toc_entry_line("Results ........ 12"));
assert!(is_toc_entry_line("Appendix B...42"));
}
#[test]
fn non_toc_lines_pass() {
assert!(!is_toc_entry_line(
"6.2. Expectations for Re-Hiring Employees"
));
assert!(!is_toc_entry_line("What happened in 2020"));
assert!(!is_toc_entry_line("IMPLEMENTATION"));
// Ellipsis without a trailing page number
assert!(!is_toc_entry_line("and so it goes ..."));
// Long numbers are data, not page refs
assert!(!is_toc_entry_line("ISBN ... 97814"));
}
#[test]
fn toc_marker_headings() {
assert!(is_toc_marker_heading("Contents"));
assert!(is_toc_marker_heading("CONTENTS"));
assert!(is_toc_marker_heading("Table of Contents"));
assert!(is_toc_marker_heading("Table of contents:"));
assert!(!is_toc_marker_heading("Contents of the Shipment"));
assert!(!is_toc_marker_heading("Introduction"));
}
#[test]
fn heading_fragments() {
// Equation lead-ins: colon ending + inline equation reference
assert!(is_heading_fragment("Rearranging Equation (8) gives:"));
// Display-equation neighbours ending in an equation number
assert!(is_heading_fragment("S = kB ln W, (2)"));
assert!(is_heading_fragment("E = mc2 (12)"));
assert!(is_heading_fragment("x + y = z, (3)"));
// Page-of-total running headers
assert!(is_heading_fragment("LIVSMEDELSVERKET PM 2 (10)"));
// Comparison-operator evidence and colon-before-number
assert!(is_heading_fragment(
"PLL\u{fe} PHH\u{226a} PLH\u{fe} PHL: (12)"
));
// Real headings pass — including name-plus-number and colon-ended ones
assert!(!is_heading_fragment("Nicaea (325)"));
assert!(!is_heading_fragment(
"\u{627}\u{644}\u{645}\u{644}\u{62d}\u{642} \u{631}\u{642}\u{645} (1)"
));
assert!(!is_heading_fragment("4. Entropy"));
assert!(!is_heading_fragment("Procedure:"));
assert!(!is_heading_fragment("Steps for Using the Microscope:"));
assert!(!is_heading_fragment("Changing objectives:"));
assert!(!is_heading_fragment("Sales by Region (2024)"));
assert!(!is_heading_fragment("Results (preliminary)"));
}
}
+25 -3
View File
@@ -7,7 +7,8 @@ use crate::types::TextLine;
use super::analysis::{
bold_heading_level, calculate_font_stats, compute_heading_tiers, compute_paragraph_threshold,
detect_header_level, font_size_rarity, has_dot_leaders,
detect_header_level, font_size_rarity, has_dot_leaders, is_heading_fragment, is_toc_entry_line,
is_toc_marker_heading,
};
use super::classify::{
format_list_item, is_caption_line, is_list_item, is_monospace_font, starts_with_bullet_marker,
@@ -486,6 +487,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
let mut in_code_block = false;
let mut prev_had_dot_leaders = false;
let mut paragraph_in_wrapped_bold_run = false;
let mut toc_suppress_page: Option<u32> = None;
let mut inserted_tables: HashSet<(u32, usize)> = HashSet::new();
let mut inserted_images: HashSet<(u32, usize)> = HashSet::new();
@@ -703,6 +705,9 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !starts_with_bullet_marker(plain_trimmed)
&& !is_toc_entry_line(plain_trimmed)
&& !is_heading_fragment(plain_trimmed)
&& toc_suppress_page != Some(line.page)
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
detect_header_level(line_font_size, base_size, &heading_tiers).or_else(|| {
@@ -738,7 +743,11 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
// paragraph continuity and minor font-size variation
// inflates rarity scores.
let has_strong_signal = all_bold || isolated || (rarity >= 0.97 && word_count <= 8);
if score >= 0.5 && standalone && word_count >= 2 && has_strong_signal {
// Single-word headings ("IMPLEMENTATION", "CONTENTS") are common;
// accept them only with the strongest signal combination.
let enough_words =
word_count >= 2 || (all_bold && isolated && plain_trimmed.len() >= 4);
if score >= 0.5 && standalone && enough_words && has_strong_signal {
Some(bold_heading_level(&heading_tiers))
} else {
None
@@ -763,6 +772,9 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
plain_text.clone()
};
output.push_str(&format!("{} {}\n\n", prefix, heading_text.trim()));
if is_toc_marker_heading(plain_trimmed) {
toc_suppress_page = Some(line.page);
}
in_list = false;
continue;
}
@@ -959,6 +971,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
let mut last_list_x: Option<f32> = None;
let mut prev_had_dot_leaders = false;
let mut paragraph_in_wrapped_bold_run = false;
let mut toc_suppress_page: Option<u32> = None;
for (line_idx, line) in lines.iter().enumerate() {
// Page break
@@ -1037,6 +1050,9 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
if options.detect_headers
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !is_toc_entry_line(plain_trimmed)
&& !is_heading_fragment(plain_trimmed)
&& toc_suppress_page != Some(line.page)
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
if let Some(header_level) =
@@ -1059,7 +1075,9 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
+ if all_bold { 0.3 } else { 0.0 }
+ if standalone { 0.2 } else { 0.0 }
+ if isolated { 0.3 } else { 0.0 };
if score >= 0.5 && standalone && word_count >= 2 {
let enough_words =
word_count >= 2 || (all_bold && isolated && plain_trimmed.len() >= 4);
if score >= 0.5 && standalone && enough_words {
return Some(bold_heading_level(&heading_tiers));
}
None
@@ -1078,6 +1096,9 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
plain_text.clone()
};
output.push_str(&format!("{} {}\n\n", prefix, heading_text.trim()));
if is_toc_marker_heading(plain_trimmed) {
toc_suppress_page = Some(line.page);
}
in_list = false;
continue;
}
@@ -1189,6 +1210,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid,
}
+1
View File
@@ -1221,6 +1221,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
}
+101 -1
View File
@@ -3,7 +3,7 @@
use std::collections::{HashMap, HashSet};
use crate::structure_tree::StructRole;
use crate::types::TextLine;
use crate::types::{TextItem, TextLine};
use super::analysis::detect_header_level;
@@ -87,6 +87,41 @@ pub(crate) fn merge_heading_lines(
false
};
// Bold headings at body font size never reach a tier, so wrapped ones
// split into two output headings ("…of wood pellets and cost" /
// "structure in Japan"). Merge a fully-bold line into the previous
// fully-bold line when it reads as a wrap continuation: starts
// lowercase, tiny Y gap, and the previous line has no terminal
// punctuation. Kept deliberately narrow — bold list labels and bold
// sentences start with markers or capitals and are unaffected.
let should_merge = should_merge
|| if let Some(prev) = result.last() {
let all_bold = |l: &TextLine| {
!l.items.is_empty() && l.items.iter().all(|i: &TextItem| i.is_bold)
};
let prev_text = prev.text();
let prev_trim = prev_text.trim_end();
let curr_text = line.text();
let curr_trim = curr_text.trim();
let y_gap = prev.y - line.y;
// Both lines must be tier-less: a tiered/tagged bold heading
// followed by bold body text must not absorb it.
line_level.is_none()
&& effective_heading_level(prev, base_size, heading_tiers, struct_roles)
.is_none()
&& prev.page == line.page
&& all_bold(prev)
&& all_bold(&line)
&& y_gap > 0.0
&& y_gap < line_font * 1.6
&& curr_trim.chars().next().is_some_and(|c| c.is_lowercase())
&& !prev_trim.ends_with(['.', ':', ';', '!', '?'])
&& prev_trim.split_whitespace().count() + curr_trim.split_whitespace().count()
<= 20
} else {
false
};
if should_merge {
// Append this line's items to the previous line
let prev = result.last_mut().unwrap();
@@ -543,6 +578,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid,
}
@@ -684,4 +720,68 @@ mod tests {
.unwrap();
assert_eq!(first_header.page, 1, "first occurrence should be on page 1");
}
fn make_bold_line(text: &str, page: u32, y: f32) -> TextLine {
let mut item = make_item(text, 12.0, None);
item.is_bold = true;
TextLine {
items: vec![item],
y,
page,
adaptive_threshold: 0.10,
}
}
#[test]
fn merge_wrapped_bold_heading_lowercase_continuation() {
// Bold-at-body-size heading wrapped across two lines: the second line
// starts lowercase and must merge into the first.
let lines = vec![
make_bold_line(
"3. Perspective of supply and demand balance and cost",
1,
700.0,
),
make_bold_line("structure in Japan", 1, 686.0),
make_line("Body text paragraph follows here.", 12.0, 1, 660.0, None),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "wrapped bold heading should merge");
assert!(result[0].text().contains("cost structure in Japan"));
}
#[test]
fn no_merge_for_bold_sentences_or_new_headings() {
// Second bold line starts with a capital — a new heading or label,
// not a wrap continuation.
let lines = vec![
make_bold_line("Replace", 1, 700.0),
make_bold_line("Trash", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "distinct bold lines must not merge");
// Previous line ends a sentence — continuation must not merge.
let lines = vec![
make_bold_line("This is a bold sentence.", 1, 700.0),
make_bold_line("another bold line", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "sentence-final bold line must not merge");
}
#[test]
fn tiered_bold_heading_does_not_absorb_bold_body() {
// Previous line is a tier-level bold heading (16pt vs 12pt body);
// a following lowercase bold body line must NOT merge into it.
let mut heading = make_bold_line("Section Title", 1, 700.0);
heading.items[0].font_size = 16.0;
heading.items[0].height = 16.0;
let lines = vec![
heading,
make_bold_line("emphasized body text continues here", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[16.0], None);
assert_eq!(result.len(), 2, "tiered heading must not absorb bold body");
}
}
+3
View File
@@ -268,6 +268,8 @@ pub struct PyTextItem {
#[pyo3(get)]
pub is_underline: bool,
#[pyo3(get)]
pub is_strikeout: bool,
#[pyo3(get)]
pub item_type: String,
}
@@ -352,6 +354,7 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item_type_str(&item.item_type),
})
.collect()
+1
View File
@@ -105,6 +105,7 @@ pub(crate) fn merge_adjacent_items(items: &[TextItem]) -> (Vec<TextItem>, Vec<Ve
is_bold: first_item.is_bold,
is_italic: first_item.is_italic,
is_underline: first_item.is_underline,
is_strikeout: first_item.is_strikeout,
item_type: first_item.item_type.clone(),
mcid: first_item.mcid,
});
+1
View File
@@ -394,6 +394,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+124
View File
@@ -1477,6 +1477,10 @@ fn detect_row_stripe_table(
debug!(" row-stripe rejected: sparse outline/prose continuation shape");
return None;
}
if has_dominant_prose_cell(&cells) {
debug!(" row-stripe rejected: dominant prose cell (chart/figure region over body text)");
return None;
}
let column_centers: Vec<f32> = (0..num_cols)
.map(|c| (col_edges[c] + col_edges[c + 1]) / 2.0)
@@ -1495,6 +1499,34 @@ fn detect_row_stripe_table(
Some(Table::new(column_centers, row_centers, cells, item_indices))
}
/// Detect a grid that swallowed body text instead of tabular data.
///
/// Charts (bar graphs, axis gridlines) emit fields of drawing rects that can
/// pass the row-stripe shape test; the resulting "table" then captures the
/// page's prose. The signature: one cell holds an entire paragraph — ≥60 words
/// AND at least a third of all words in the table.
///
/// There is deliberately no row-count exemption. A small table whose single
/// long cell dominates its word count is indistinguishable by content from a
/// phantom grid over body text, and across the regression corpora every such
/// grid observed has been swallowed prose, never a real note table. The costs
/// are also asymmetric: rejecting a real table degrades it to readable prose,
/// while accepting a phantom scrambles the page into Y-interleaved cells.
/// Larger legitimate tables are safe because the one-third-of-total threshold
/// scales with table size.
fn has_dominant_prose_cell(cells: &[Vec<String>]) -> bool {
let mut total_words = 0usize;
let mut max_cell_words = 0usize;
for row in cells {
for cell in row {
let words = cell.split_whitespace().count();
total_words += words;
max_cell_words = max_cell_words.max(words);
}
}
max_cell_words >= 60 && max_cell_words * 3 >= total_words
}
fn row_stripe_is_sparse_prose_outline(cells: &[Vec<String>]) -> bool {
let Some(num_cols) = cells.first().map(|row| row.len()) else {
return false;
@@ -2294,6 +2326,12 @@ fn detect_merged_cluster_table(
);
return None;
}
if has_dominant_prose_cell(&cells) {
debug!(
" merged-cluster rejected: dominant prose cell (chart/figure region over body text)"
);
return None;
}
// No empty columns
for col in 0..num_cols {
@@ -2392,11 +2430,95 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
}
// --- has_dominant_prose_cell ---
fn cells_of(rows: &[&[&str]]) -> Vec<Vec<String>> {
rows.iter()
.map(|r| r.iter().map(|c| c.to_string()).collect())
.collect()
}
#[test]
fn dominant_prose_cell_rejects_swallowed_paragraph() {
// Two cells hold paragraphs (the shape every observed phantom grid
// has: swallowed body text spans multiple cells), rest are chart labels
let para = ["word"; 70].join(" ");
let para2 = ["word"; 35].join(" ");
let cells = cells_of(&[
&[para.as_str(), "81", "76"],
&[para2.as_str(), "56", "9"],
&["2019", "2020", ""],
]);
assert!(has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_rejects_small_table_dominated_by_one_cell() {
// Boundary case, documented as INTENDED: a small grid whose single
// long cell dominates the word count is rejected even at 4+ rows.
// By content alone this shape is indistinguishable from a phantom
// grid over body text, and every observed instance in the regression
// corpora was swallowed prose (chart/figure regions), not a real
// note table. Rejection degrades gracefully — the text is still
// extracted as prose — while accepting a phantom scrambles reading
// order.
let note = ["word"; 70].join(" ");
let cells = cells_of(&[
&["Purpose", note.as_str()],
&["Owner", "Facilities team"],
&["Date", "2024-06-01"],
&["Status", "Active"],
]);
assert!(has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_allows_description_column() {
// Long-ish description cells, but text is spread across the table
let desc = ["word"; 25].join(" ");
let cells = cells_of(&[
&["Item A", desc.as_str(), "100"],
&["Item B", desc.as_str(), "200"],
&["Item C", desc.as_str(), "300"],
&["Item D", desc.as_str(), "400"],
]);
assert!(!has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_allows_short_tables() {
let cells = cells_of(&[&["Name", "Value"], &["Total", "42"]]);
assert!(!has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_allows_data_table_with_long_note() {
// A real 4+ row table with one verbose remark cell: the note is ≥60
// words but the table's other content carries more than 2× its word
// count, so concentration stays below the 1/3 threshold. The
// denominator scales with table size — this is what keeps large
// legitimate tables safe where a bare length cap would not.
let note = ["word"; 60].join(" ");
let row_text = ["data"; 12].join(" ");
let mut rows: Vec<Vec<String>> = (0..11)
.map(|i| {
vec![
format!("Item {i}"),
row_text.clone(),
format!("{}", i * 100),
]
})
.collect();
rows.push(vec!["Note".into(), note, String::new()]);
assert!(!has_dominant_prose_cell(&rows));
}
// --- rects_overlap ---
#[test]
@@ -3366,6 +3488,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
});
@@ -3676,6 +3799,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
});
+1
View File
@@ -587,6 +587,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid,
}
+1
View File
@@ -109,6 +109,7 @@ pub(crate) fn try_split_financial_item(item: &TextItem) -> Option<Vec<TextItem>>
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
+3
View File
@@ -521,6 +521,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -886,6 +887,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
@@ -923,6 +925,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
+4
View File
@@ -236,6 +236,7 @@ fn split_merged_numbers(item: &TextItem, col_boundaries: &[f32]) -> Vec<TextItem
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
@@ -257,6 +258,7 @@ fn split_merged_numbers(item: &TextItem, col_boundaries: &[f32]) -> Vec<TextItem
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
@@ -1432,6 +1434,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1450,6 +1453,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+17
View File
@@ -93,6 +93,9 @@ pub fn is_bold_font(font_name: &str) -> bool {
|| lower.contains("extrabold")
|| lower.contains("ultrabold")
|| lower.contains("medium") && !lower.contains("mediumitalic") // Some fonts use Medium for semi-bold
// URW Type 1 fonts abbreviate Medium as "Medi" (e.g. NimbusRomNo9L-Medi,
// the Times-Bold substitute in LaTeX documents; -MediItal is bold italic).
|| lower.contains("-medi") && !lower.contains("mediumital")
}
/// Detect if a font name indicates italic/oblique style
@@ -762,6 +765,17 @@ mod tests {
use super::*;
use crate::types::ItemType;
#[test]
fn bold_font_urw_medi_abbreviation() {
// URW Type 1 fonts (LaTeX default Times) abbreviate Medium as "Medi"
assert!(is_bold_font("NROFIU+NimbusRomNo9L-Medi"));
assert!(is_bold_font("NimbusRomNo9L-MediItal"));
assert!(!is_bold_font("DSSZWN+NimbusRomNo9L-Regu"));
assert!(!is_bold_font("NimbusRomNo9L-ReguItal"));
// Medium-Italic exclusion still holds
assert!(!is_bold_font("Foo-MediumItalic"));
}
#[test]
fn strip_soft_hyphen() {
assert_eq!(expand_ligatures("con\u{00AD}tent"), "content");
@@ -884,6 +898,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1004,6 +1019,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -1081,6 +1097,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+4
View File
@@ -120,6 +120,10 @@ pub struct TextItem {
/// baseline — PDFs have no underline font flag, so this is detected
/// geometrically after extraction; see `extractor::underline`).
pub is_underline: bool,
/// Whether the text is struck out (drawn rule/thin rect crossing the
/// glyphs at mid x-height). Same geometric detection as underline,
/// different vertical window; see `extractor::underline`.
pub is_strikeout: bool,
/// Type of item (text, image, link)
pub item_type: ItemType,
/// Marked Content ID from the content stream's BDC/BMC operator.
+2
View File
@@ -105,6 +105,7 @@ fn make_text_item(text: &str, x: f32, y: f32, font_size: f32, page: u32) -> Text
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -131,6 +132,7 @@ fn make_text_item_with_font(
is_bold: is_bold_font(font),
is_italic: is_italic_font(font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+1 -2
View File
@@ -48,7 +48,7 @@ tips of directly from customers received other employees paid tips recd. entr
**Page 3**
27 28 29 30 31 **Subtotals** **from pages** **1, 2, and 3** **Totals**
27 28 29 30 31 **Subtotals from pages** **1, 2, and 3** **Totals**
**1.** Report total cash tips (col. **a**) on Form 4070, line **1.**
**2.** Report total credit card tips (col. **b**) on Form 4070, line **2.**
@@ -79,4 +79,3 @@ forms simpler, we would be happy to hear from you. You can write to the Tax Form
**Instructions** *(continued)*
Use this space to total your tips for the year