Compare commits

..
Author SHA1 Message Date
Abimael MartellandClaude Fable 5 1470dc162a fix(markdown): one-word bold headings, block ToC entries from headings
Two heading-classification fixes:

- Accept single-word headings ("IMPLEMENTATION", "CONTENTS") when the
  line is all-bold and isolated; the word_count >= 2 gate rejected them
  unconditionally.
- Add is_toc_entry_line: a line ending in a dot-leader group plus page
  number ("Measurement Lab worksheet ... 3") is a table-of-contents
  entry, never a heading. has_dot_leaders misses single-group leaders,
  so entire ToC pages were being promoted to ## headings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:14:32 -07:00
Abimael MartellandClaude Fable 5 b4d401241e chore(napi): bump @firecrawl/pdf-inspector to 1.10.0 (#127)
New in this release (#125): isStrikeout on TextItem (geometric
detection sharing the underline rules pipeline), descriptor/embedded-
font bold+italic recall for subset fonts (FontDescriptor flags,
ttf-parser OS/2+post, bare-CFF Name INDEX), quote-operator advance
width, Ts text-rise handling, ActualText rise/position fixes, and a
document-scoped font style cache.


Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 17:41:46 -07:00
Abimael MartellandClaude Fable 5 57335f8bcf feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection (#125)
* feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection

Two style-recall gaps, both invisible to the existing name-based
heuristics:

1. Subset fonts with opaque BaseFont names ("Tc1", "AAAAAB+Amplitude")
   defeat is_italic_font/is_bold_font. New descriptor_style_flags reads
   the FontDescriptor (ItalicAngle beyond 4 degrees, Flags bit 7 Italic,
   bit 19 ForceBold) and, when the descriptor claims upright, falls back
   to the embedded font file: ttf-parser's OS/2 fsSelection + post
   italicAngle for sfnt fonts, and the CFF Name INDEX PostScript name
   for bare-CFF FontFile3 (descriptor rewritten to ItalicAngle 0 while
   embedding "Amplitude-LightItalic" was observed in the wild).
   ORed into is_bold/is_italic at item creation (content streams and
   form XObjects).

2. No strikeout signal existed. New is_strikeout on TextItem, detected
   in the same pass as underline: same rules pipeline (stroked lines /
   thin filled rects, table-ruling suppression), different vertical
   window — a rule crossing the glyphs at 12-55% of the em above the
   baseline instead of sitting at it. Exposed through napi and python
   bindings and pdf2md --items-json.

Verified on public ParseBench corpus docs: previously-missed italic
council titles and bold CJK itinerary headings now flagged (render-
checked); 24/508 docs gain flags, none lose any; 35 strikeout items
detected corpus-wide, disjoint from underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): quote-op advance width, Ts text rise, doc-level font style cache (PR #125 review)

Address three valid findings from review:

- The ' (move-to-next-line-and-show-text) operator emitted zero-width
  items and never advanced the text matrix, so geometric underline/
  strikeout detection (which requires width > 0) could never mark its
  text, and following show ops overlapped it. Reuse Tj's advance-width
  computation and matrix advance.

- Ts (text rise) was dropped entirely: raised/lowered runs kept the
  unshifted baseline, so rules drawn at the risen glyph position missed
  the strike/underline windows. Track rise in the text state (saved and
  restored with q/Q) and shift the rendering position through the text
  matrix's y column; advances stay on the unshifted matrix per spec.

- descriptor_style_flags re-decompressed and re-parsed the same embedded
  font program on every page whenever the descriptor left a style flag
  unset (the common case). Add a document-scoped FontStyleCache keyed by
  the FontFile2/FontFile3 object id, threaded through page and form
  extraction alongside the existing CMapDecisionCache.

The fourth finding (Form XObject rules never reach geometric detection)
is real but pre-existing for underline and needs the form walker to grow
path/paint tracking plus a new return type; deferred as a follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): ActualText items render at their glyphs' text rise (PR #125 review)

The EMC-built ActualText item used the captured text matrix without the
rise adjustment the ordinary Tj/TJ/' emission sites apply, so a tagged
run shown with Ts landed on the unshifted baseline — off the strikeout/
underline windows and inconsistent with untagged runs. The rise is
captured together with the first-glyph matrix (and at BDC for the
entry-position fallback): the item must render at the rise of its
GLYPHS, not whatever rise is set by EMC time.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): capture ActualText glyph position after the quote op's line move (PR #125 review)

The `'` handler skipped the entire suppressed-extraction block, so a
tagged span whose show op is `'` never captured its glyph matrix/rise —
the EMC item fell back to the BDC-entry matrix, which sits on the
PREVIOUS line (the `'` line move happens after BDC) with no rise. The
capture now happens right after the line move, matching the Tj/TJ
paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): style-boundary gate on subscript merge + strikeout suppression coverage (PR #125 review)

merge_subscript_items absorbed a script digit into its parent
regardless of underline/strikeout flags — dropping the digit's own mark
or widening the parent's over it. The merged item carries one flag, so
differing marks now break the merge, mirroring merge_text_items'
style-boundary rule (pre-existing for underline as well).

Also extends the table-suppression test to assert is_strikeout is
cleared alongside is_underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 17:30:47 -07:00
Abimael MartellandClaude Fable 5 15bc7894a4 docs: add MIT LICENSE file and license badge (#124)
Cargo.toml and pyproject.toml already declare MIT but the repo had no
LICENSE file, so GitHub and package registries couldn't display it.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 14:25:08 -07:00
Abimael MartellandClaude Fable 5 6e5e5849c8 Detect substitution-cipher garbled text from broken ToUnicode CMaps (#120)
* fix(lib): detect substitution-cipher garbled text from broken ToUnicode CMaps

ParseBench text_simple__att10k.pdf (issue #118) ships Type0/Identity-H
fonts whose ToUnicode CMaps are authored garbled: every bfrange maps with
a wrong constant delta, so text extracts as pure-ASCII ciphertext
("Certificate" -> "8VceZWZTReV"). The embedded subset font has no cmap
table and no glyph names, so no decode source can recover the real text
(poppler and mupdf emit the same ciphertext). The only correct behavior
is to flag the page for OCR instead of serving the garbage silently --
but the text is 100% printable ASCII with word-like tokens, so it slipped
past is_garbage_text and detect_encoding_issues.

Add CipherGarbleStats, a letter-statistics discriminator that flags a
Latin-dominant sample (>=200 ASCII letters) when vowels are starved
(<=30% of letters) AND either:
- lowercase->uppercase transitions inside words exceed 10% of letter
  bigrams (a shifted lowercase alphabet straddles the ASCII uppercase
  block), or
- the letter histogram's cosine similarity against English letter
  frequencies drops below 0.60 (catches shifts that stay within case
  blocks).

Wired into analyze_text_quality (per-page, item-level) and
detect_encoding_issues (markdown-level), so extract_pages_markdown
reports needs_ocr + suspected_garbled_text and suppresses the garbage.

Thresholds validated against the 380-document pdf-evals snapshot corpus
(Swedish, Finnish, Turkish, German, romaji, schematics, all-caps and
camelCase-heavy docs): zero false positives, and byte-identical eval
output vs main. Garbled page measures vowel ratio 0.245 / case-shift
rate 0.225 / cosine 0.532; closest legitimate document on each axis is
0.264 / 0.021 / 0.801.

Fixes #118

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump pdf-inspector to 0.1.4, npm package to 1.9.11

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(lib): exempt uniform-case structured content from cipher detection

Address PR review (cubic P2): the frequency branch (english_cosine < 0.60)
fired on any Latin-dominant, low-vowel letter distribution unlike English,
so non-linguistic ASCII — DNA/protein sequences, ticker symbols, hex dumps —
could be suppressed and routed to OCR despite not being garbled. Measured:
DNA cosine 0.428 / vowel ratio 0.260, protein 0.738, tickers 0.747, hex
0.549 — all would have flagged.

Add a mixed-case guard to looks_garbled: garbled English is a permutation of
natural language and carries sentence capitalization (block-straddling shifts
invert the ratio — att10k is 60% uppercase; in-case Caesar shifts preserve it
at ~3%), so both keep some of each case. The exempted structured content is
uniform case (all upper or all lower). Requiring the minority case to be >=1%
of ASCII letters exempts single-case sequences while preserving both garble
signals, including the in-case-shift scenario the frequency branch exists for.

Strictly tightens the detector: it can only remove flags, so the eval corpus
stays at zero false positives (verified byte-identical to a baseline main
binary across all 185 PDFs) and att10k remains flagged. Adds regression tests
for DNA, protein, tickers, and an in-case Caesar shift.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(lib): make cipher detection case-agnostic via sorted-histogram shape

Address PR review follow-up: the mixed-case guard from the previous commit
returned before the vowel/frequency checks, creating a blind spot — a
uniform-case (all-lower or all-upper) substitution cipher is a plausible
broken-CMap output and would bypass OCR entirely.

Replace the case proxy with the actual invariant. A substitution cipher is
a bijection over a real language's alphabet, so it preserves the frequency
SHAPE (the sorted histogram) while scrambling letter POSITIONS (the unsorted
histogram). Signal 2 now flags when english_cosine < 0.60 (positions unlike
English) AND english_shape_cosine >= 0.90 (profile is still English-shaped).
This is independent of case, so it catches all-lower, all-upper, and
case-straddling shifts alike.

The exempted structured content fails one half: DNA/hex dumps have too steep
a profile (shape cosine 0.74 / 0.81 < 0.90), while protein sequences, ticker
symbols and base64 are not sufficiently unlike English in position (unsorted
cosine 0.74 / 0.75 / 0.77 >= 0.60). All stay out of OCR.

Still strictly corpus-safe: every real Latin document scores unsorted cosine
>= 0.70 (min 0.80), far above the 0.60 gate, so none can reach Signal 2.
Re-verified byte-identical to a baseline main binary across all 185 eval
PDFs; att10k remains flagged. Drops the now-unused case counters and adds
all-lowercase / all-uppercase shifted-prose regression tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: source Python package version from Cargo.toml via maturin

Address PR review (cubic P2): pyproject.toml pinned version = "0.1.0",
which overrides Cargo.toml, so a maturin build produced a 0.1.0 Python
artifact regardless of the crate version (it had drifted since the PyO3
bindings were added). Switch to dynamic = ["version"] so maturin sources
the version from Cargo.toml [package] version and the two can no longer
diverge. No workflow auto-publishes the Python package, so this is metadata
hygiene rather than a release-path fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 12:44:32 -07:00
Abimael MartellandCursor b375d6f102 feat(markdown): underline emission, Unicode scripts, style-preserving merges (#117)
* feat(markdown): underline emission, Unicode scripts, style-preserving merges (ENG-5015 2b)

Three formatting losses in the direct-extraction markdown path:

1. text_with_formatting gains <u> run emission (detect_underline option,
   default on) using the geometric is_underline flag from 1.9.9.
   Underline runs stay free of nested bold/italic markers — consumers
   match tag content literally. Heading lines keep plain text for
   bold/italic but preserve <u>: the tag carries meaning `#` doesn't.
2. merge_subscript_items now maps absorbed digit scripts to Unicode
   sub/superscript forms with direction from the baseline offset
   ("H"+"2" -> "H₂", "word"+raised "2" -> "word²", "m"+"3" -> "m³").
   NFKC/NFKD folds these back to plain digits so text matching
   downstream is unaffected; renderers keep the script semantics.
3. merge_text_items no longer merges across bold/italic boundaries —
   absorbing a styled run into a plain neighbor erased the styling
   before markdown emission ever saw it. On eval docs this recovers
   20-82 italic runs per document that previously emitted as plain.

Snapshots regenerated (diffs are the features: CCl₂F₂, m³, underlined
legal section headings, finer bold runs). pdf-evals regression suite:
202/202 real PDFs pass. napi 1.9.9 -> 1.9.10.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): break merges at underline boundaries too (review)

OR-merging underline stretched the eventual <u> span over neighboring
plain fragments. Merge runs now break on any style-flag change, the
redundant accumulator is gone, and format_list_item learned to move
bullet markers outside <u> wrappers so fully-underlined bullet lines
still render as markdown lists. td9264 snapshot regenerated — spans are
tighter (trailing periods correctly outside the tag).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(markdown): strip stray spaces before sentence punctuation (review)

Style-boundary item splits can strand a trailing period in its own
fragment, and multiple assembly paths join fragments with spaces,
yielding "word ." artifacts. Rather than chasing every join site, a
postprocess pass removes a space before `.`/`,`/`;` when the mark ends
its token (whitespace, cell boundary `|`, or end of text follows).
Dot leaders/ellipses and mid-token periods are untouched.

Fixes the td9264 "companies ." artifacts and two pre-existing
"armoring ," artifacts in the 2013-app2 snapshot. pdf-evals: zero
markdown diffs across all 203 corpus PDFs vs committed baselines.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(tables): trim spaces inside parenthetical cell fragments

* fix(tables): reject sparse prose row-stripe tables

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 23:40:08 -07:00
33 changed files with 1255 additions and 46 deletions
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector"
version = "0.1.3"
version = "0.1.4"
edition = "2021"
autobins = false
authors = ["Firecrawl Team"]
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Firecrawl
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+2 -1
View File
@@ -2,6 +2,7 @@
[![Crates.io](https://img.shields.io/crates/v/pdf-inspector.svg)](https://crates.io/crates/pdf-inspector)
[![npm](https://img.shields.io/npm/v/@firecrawl/pdf-inspector.svg)](https://www.npmjs.com/package/@firecrawl/pdf-inspector)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for [Python](docs/python.md) and [Node.js](napi/README.md).
@@ -242,4 +243,4 @@ See [docs/debugging.md](docs/debugging.md) for `RUST_LOG` environment variable u
## License
MIT
[MIT](LICENSE)
+1 -1
View File
@@ -830,7 +830,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
[[package]]
name = "pdf-inspector"
version = "0.1.3"
version = "0.1.4"
dependencies = [
"env_logger",
"log",
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.9.10",
"version": "1.10.0",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
+4
View File
@@ -83,6 +83,9 @@ pub struct TextItem {
/// Underline detected geometrically (drawn rule/thin rect under the
/// baseline) — PDFs carry no underline font flag.
pub is_underline: bool,
/// Strikeout detected geometrically (rule crossing the glyphs at mid
/// x-height).
pub is_strikeout: bool,
pub item_type: ItemType,
/// URL for link items, `None` for other types.
pub link_url: Option<String>,
@@ -294,6 +297,7 @@ pub fn extract_text_with_positions(
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type,
link_url,
}
+1
View File
@@ -39,6 +39,7 @@ class TextItem:
is_bold: bool
is_italic: bool
is_underline: bool
is_strikeout: bool
item_type: str
class RegionText:
+3 -1
View File
@@ -4,7 +4,9 @@ build-backend = "maturin"
[project]
name = "pdf-inspector"
version = "0.1.0"
# Version is sourced from Cargo.toml [package] version by maturin so the Python
# artifact always tracks the crate release instead of drifting on its own.
dynamic = ["version"]
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
license = { text = "MIT" }
requires-python = ">=3.8"
+3 -1
View File
@@ -74,7 +74,7 @@ fn format_items_json(items: &[TextItem]) -> String {
_ => String::new(),
};
format!(
r#"{{"text":"{}","page":{},"x":{:.2},"y":{:.2},"width":{:.2},"height":{:.2},"font":"{}","font_size":{:.2},"is_bold":{},"is_italic":{},"is_underline":{},"item_type":"{}","mcid":{}{}}}"#,
r#"{{"text":"{}","page":{},"x":{:.2},"y":{:.2},"width":{:.2},"height":{:.2},"font":"{}","font_size":{:.2},"is_bold":{},"is_italic":{},"is_underline":{},"is_strikeout":{},"item_type":"{}","mcid":{}{}}}"#,
json_escape(&item.text),
item.page,
item.x,
@@ -86,6 +86,7 @@ fn format_items_json(items: &[TextItem]) -> String {
item.is_bold,
item.is_italic,
item.is_underline,
item.is_strikeout,
item_type_label(&item.item_type),
mcid,
link_url,
@@ -122,6 +123,7 @@ mod tests {
is_bold: false,
is_italic: true,
is_underline: true,
is_strikeout: true,
item_type: ItemType::Text,
mcid: Some(7),
}];
+250 -20
View File
@@ -14,8 +14,9 @@ use lopdf::{Document, Encoding, Object, ObjectId};
use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
build_font_encodings, build_font_widths, compute_string_width_ts, descriptor_style_flags,
extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
FontStyleCache,
};
use super::underline::UnderlineLine;
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
@@ -106,6 +107,25 @@ fn transformed_stroke_width(
user_width * (ndx * ndx + ndy * ndy).sqrt()
}
/// Text rise (Ts) displaces the glyph origin by (0, rise) in unscaled text
/// space — per the rendering-matrix definition it sits left of Tm, so the
/// offset maps through the text matrix's y column. Rise never contributes
/// to the advance, so callers apply it only to the rendering position and
/// keep advancing the unshifted text matrix.
fn rise_adjusted(tm: &[f32; 6], rise: f32) -> [f32; 6] {
if rise == 0.0 {
return *tm;
}
[
tm[0],
tm[1],
tm[2],
tm[3],
tm[4] + rise * tm[2],
tm[5] + rise * tm[3],
]
}
/// Returns `(page_extraction, has_gid_fonts)` where `has_gid_fonts` indicates
/// the page uses fonts with unresolvable gid-encoded glyphs.
pub(crate) fn extract_page_text_items(
@@ -114,6 +134,7 @@ pub(crate) fn extract_page_text_items(
page_num: u32,
font_cmaps: &FontCMaps,
include_invisible: bool,
style_cache: &mut FontStyleCache,
) -> Result<(PageExtraction, bool, bool), PdfError> {
use lopdf::content::Content;
@@ -153,6 +174,8 @@ pub(crate) fn extract_page_text_items(
std::collections::HashMap::new();
let mut inline_cmaps: std::collections::HashMap<String, crate::tounicode::CMapEntry> =
std::collections::HashMap::new();
let mut font_style_flags: std::collections::HashMap<String, (bool, bool)> =
std::collections::HashMap::new();
for (font_name, font_dict) in &fonts {
let resource_name = String::from_utf8_lossy(font_name).to_string();
if let Ok(base_font) = font_dict.get(b"BaseFont") {
@@ -161,6 +184,12 @@ pub(crate) fn extract_page_text_items(
font_base_names.insert(resource_name.clone(), base_name);
}
}
// Descriptor style flags rescue subset fonts whose BaseFont names
// are opaque tags the name heuristics can't read.
let style = descriptor_style_flags(doc, font_dict, style_cache);
if style != (false, false) {
font_style_flags.insert(resource_name.clone(), style);
}
// Track ToUnicode object reference, with FontFile2 fallback for Identity-H/V.
// Also handle inline ToUnicode streams.
match font_dict.get(b"ToUnicode") {
@@ -235,6 +264,7 @@ pub(crate) fn extract_page_text_items(
line_width: f32,
char_spacing: f32,
word_spacing: f32,
text_rise: f32,
text_leading: f32,
current_font: String,
current_font_size: f32,
@@ -247,6 +277,7 @@ pub(crate) fn extract_page_text_items(
let mut text_leading: f32 = 0.0; // TL parameter (in text-space units)
let mut char_spacing: f32 = 0.0; // Tc parameter (extra spacing per character, unscaled)
let mut word_spacing: f32 = 0.0; // Tw parameter (extra spacing per space char, unscaled)
let mut text_rise: f32 = 0.0; // Ts parameter (baseline shift for super/subscripts, unscaled)
let mut text_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut line_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut in_text_block = false;
@@ -268,6 +299,10 @@ pub(crate) fn extract_page_text_items(
let mut suppress_glyph_extraction = false;
let mut actual_text_start_tm: Option<[f32; 6]> = None; // text matrix at BDC entry
let mut actual_text_glyph_tm: Option<[f32; 6]> = None; // text matrix at first glyph inside BDC
// Text rise in effect at each captured matrix — the item must render at
// the rise of its GLYPHS, not whatever rise is set by EMC time.
let mut actual_text_start_rise: f32 = 0.0;
let mut actual_text_glyph_rise: Option<f32> = None;
/// Get the innermost MCID from the marked content stack.
fn current_mcid(stack: &[MarkedContentEntry]) -> Option<i64> {
stack.iter().rev().find_map(|e| e.mcid)
@@ -284,6 +319,7 @@ pub(crate) fn extract_page_text_items(
line_width,
char_spacing,
word_spacing,
text_rise,
text_leading,
current_font: current_font.clone(),
current_font_size,
@@ -297,6 +333,7 @@ pub(crate) fn extract_page_text_items(
line_width = saved.line_width;
char_spacing = saved.char_spacing;
word_spacing = saved.word_spacing;
text_rise = saved.text_rise;
text_leading = saved.text_leading;
current_font = saved.current_font;
current_font_size = saved.current_font_size;
@@ -369,6 +406,12 @@ pub(crate) fn extract_page_text_items(
word_spacing = tw;
}
}
"Ts" => {
// Set text rise (baseline shift for superscripts/subscripts)
if let Some(ts) = op.operands.first().and_then(get_number) {
text_rise = ts;
}
}
"Td" | "TD" => {
// Move text position: TLM = T(tx,ty) × TLM; Tm = TLM
// tx,ty are in text space — must be scaled by the text line matrix
@@ -427,6 +470,7 @@ pub(crate) fn extract_page_text_items(
if suppress_glyph_extraction {
if actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
@@ -456,7 +500,8 @@ pub(crate) fn extract_page_text_items(
&mut cmap_decisions,
&font_widths,
) {
let combined = multiply_matrices(&text_matrix, &ctm);
let combined =
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
let (x, y) = (combined[4], combined[5]);
if combined[0].abs() >= combined[1].abs() {
@@ -478,6 +523,10 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
@@ -487,9 +536,10 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -507,6 +557,7 @@ pub(crate) fn extract_page_text_items(
// Capture first-glyph position for ActualText
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
// Compute space threshold based on font metrics when available
@@ -627,6 +678,10 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
let scale_x = text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2];
for (text, start_w, end_w) in &sub_items {
let offset_tm = [
@@ -637,7 +692,8 @@ pub(crate) fn extract_page_text_items(
text_matrix[4] + start_w * text_matrix[0],
text_matrix[5] + start_w * text_matrix[1],
];
let combined = multiply_matrices(&offset_tm, &ctm);
let combined =
multiply_matrices(&rise_adjusted(&offset_tm, text_rise), &ctm);
let (x, y) = (combined[4], combined[5]);
let width = if font_info.is_some() {
((end_w - start_w) * scale_x).abs()
@@ -653,9 +709,10 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -679,6 +736,26 @@ pub(crate) fn extract_page_text_items(
line_matrix[4] += (-tl) * line_matrix[2];
line_matrix[5] += (-tl) * line_matrix[3];
text_matrix = line_matrix;
// Capture first-glyph position for ActualText AFTER the
// line move — the BDC-entry matrix is on the previous line.
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
// Advance width, as for Tj — without it the item stays
// zero-width and geometric underline/strikeout detection
// rejects it (`is_underline_candidate` needs width > 0).
let w_ts_opt = font_widths.get(&current_font).and_then(|fi| {
op.operands.first().and_then(get_operand_bytes).map(|raw| {
compute_string_width_ts(
raw,
fi,
current_font_size,
char_spacing,
word_spacing,
)
})
});
if !((text_rendering_mode == 3 && !include_invisible)
|| suppress_glyph_extraction
|| op.operands.is_empty())
@@ -696,7 +773,8 @@ pub(crate) fn extract_page_text_items(
&font_widths,
) {
if !text.trim().is_empty() {
let combined = multiply_matrices(&text_matrix, &ctm);
let combined =
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
} else {
@@ -704,28 +782,45 @@ pub(crate) fn extract_page_text_items(
}
let rendered_size = effective_font_size(current_font_size, &combined);
let (x, y) = (combined[4], combined[5]);
let width = w_ts_opt
.map(|w_ts| {
(w_ts * (text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2]))
.abs()
})
.unwrap_or(0.0);
let base_font = font_base_names
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
y,
width: 0.0,
width,
height: rendered_size,
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
}
}
}
// Advance regardless of visibility so later show-text
// operators on the same line stay positioned (as for Tj).
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
}
}
"Do" => {
// XObject invocation - could be an image or form
@@ -757,6 +852,7 @@ pub(crate) fn extract_page_text_items(
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Image,
mcid: current_mcid(&marked_content_stack),
});
@@ -770,6 +866,7 @@ pub(crate) fn extract_page_text_items(
font_cmaps,
&ctm,
&mut cmap_decisions,
style_cache,
);
items.extend(form_items);
}
@@ -810,7 +907,9 @@ pub(crate) fn extract_page_text_items(
if actual_text.is_some() {
suppress_glyph_extraction = true;
actual_text_start_tm = Some(text_matrix);
actual_text_start_rise = text_rise;
actual_text_glyph_tm = None; // reset — will be captured at first Tj/TJ
actual_text_glyph_rise = None;
}
marked_content_stack.push(MarkedContentEntry { actual_text, mcid });
}
@@ -823,9 +922,11 @@ pub(crate) fn extract_page_text_items(
// Tj may have moved the text position to the correct line —
// the BDC-entry position can be on the previous line.
let glyph_tm = actual_text_glyph_tm.take();
let glyph_rise = actual_text_glyph_rise.take();
let entry_tm = actual_text_start_tm.take();
if let Some(start_tm) = glyph_tm.or(entry_tm) {
let combined = multiply_matrices(&start_tm, &ctm);
let rise = glyph_rise.unwrap_or(actual_text_start_rise);
let combined = multiply_matrices(&rise_adjusted(&start_tm, rise), &ctm);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
} else {
@@ -842,6 +943,10 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&at),
x,
@@ -851,9 +956,10 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: entry
.mcid
@@ -1376,8 +1482,15 @@ mod tests {
let (doc, page_id) = simple_doc_with_content(content);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) =
extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let ((items, _, _), _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
items
}
@@ -1462,6 +1575,108 @@ BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
assert!(!world.is_underline);
}
#[test]
fn quote_operator_text_carries_advance_width() {
// `'` (move-to-next-line-and-show-text) must retain the string's
// advance width like Tj — zero-width items are invisible to
// geometric underline/strikeout detection.
let content = b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (first) Tj (struck) ' ET
1 w
99 503 m 145 503 l S";
let items = extract_simple_items(content);
let struck = items.iter().find(|item| item.text == "struck").unwrap();
// 6 glyphs x 600/1000 x 12pt = 43.2pt, drawn one leading below Tm.
assert!((struck.width - 43.2).abs() < 0.1);
assert!((struck.y - 500.0).abs() < 0.1);
assert!(struck.is_strikeout);
assert!(!struck.is_underline);
}
#[test]
fn quote_operator_advances_text_matrix() {
// Text shown after `'` on the same line must start past the shown
// string: "CD" lands at x=114.4 (2 glyphs x 600/1000 x 12pt after
// x=100), flush against "AB", so the merge pass joins them. Without
// the advance "CD" overlaps "AB" at x=100 and the items stay apart.
let content = b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (AB) ' (CD) Tj ET";
let items = extract_simple_items(content);
let merged = items.iter().find(|item| item.text == "ABCD").unwrap();
assert!((merged.x - 100.0).abs() < 0.1);
assert!((merged.width - 28.8).abs() < 0.1);
assert!((merged.y - 500.0).abs() < 0.1);
}
#[test]
fn text_rise_shifts_item_baseline() {
// Ts displaces the glyph origin vertically without touching the
// advance; the next run at rise 0 must return to the original
// baseline and follow the raised run horizontally.
let content =
b"BT /F1 12 Tf 1 0 0 1 100 500 Tm (base) Tj 5 Ts (super) Tj 0 Ts (after) Tj ET";
let items = extract_simple_items(content);
let base = items.iter().find(|item| item.text == "base").unwrap();
let raised = items.iter().find(|item| item.text == "super").unwrap();
let after = items.iter().find(|item| item.text == "after").unwrap();
assert!((base.y - 500.0).abs() < 0.1);
assert!((raised.y - 505.0).abs() < 0.1);
assert!((after.y - 500.0).abs() < 0.1);
assert!(after.x > raised.x);
}
#[test]
fn actual_text_item_uses_glyph_rise() {
// The ActualText replacement item must render at the rise in
// effect when its glyphs were drawn — not the unshifted BDC
// baseline, and not whatever rise is set by EMC time.
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm \
/Span <</ActualText (super) >> BDC 5 Ts (sup) Tj 0 Ts EMC (after) Tj ET";
let items = extract_simple_items(content);
let sup = items.iter().find(|item| item.text == "super").unwrap();
let after = items.iter().find(|item| item.text == "after").unwrap();
assert!((sup.y - 505.0).abs() < 0.1);
assert!((after.y - 500.0).abs() < 0.1);
}
#[test]
fn actual_text_shown_with_quote_op_uses_moved_risen_baseline() {
// When the tagged span's show op is `'`, the glyph position is
// only known AFTER its line move — falling back to the BDC-entry
// matrix would place the item on the previous line, unrisen.
let content = b"BT /F1 12 Tf 14 TL 1 0 0 1 100 500 Tm \
/Span <</ActualText (replaced) >> BDC 3 Ts (raw) ' 0 Ts EMC ET";
let items = extract_simple_items(content);
let item = items.iter().find(|item| item.text == "replaced").unwrap();
// Line move: 500 - 14 = 486; rise: +3 -> 489.
assert!((item.y - 489.0).abs() < 0.1);
assert!(item.width > 0.0);
}
#[test]
fn strikeout_detected_on_risen_text() {
// The rule crosses the glyphs at their risen position; without the
// rise in item.y the strike window sits 4pt too low and misses.
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm 4 Ts (struck) Tj ET
1 w
99 507 m 145 507 l S";
let items = extract_simple_items(content);
let struck = items.iter().find(|item| item.text == "struck").unwrap();
assert!((struck.y - 504.0).abs() < 0.1);
assert!(struck.is_strikeout);
assert!(!struck.is_underline);
}
#[test]
fn test_skip_excessive_operations() {
use crate::tounicode::FontCMaps;
@@ -1496,7 +1711,15 @@ BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
doc.add_object(catalog);
let font_cmaps = FontCMaps::from_doc(&doc);
let result = extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let result = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
let ((items, rects, lines), _has_gid, _coords_rotated) = result;
assert!(items.is_empty());
assert!(rects.is_empty());
@@ -1578,8 +1801,15 @@ BT 30 700 Tm <41> Tj ET";
doc.trailer.set("Root", Object::Reference(catalog_id));
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) =
extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let ((items, _, _), _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
let text = items
.iter()
.map(|item| item.text.as_str())
+332 -1
View File
@@ -4,7 +4,7 @@ use crate::glyph_names::glyph_to_char;
use crate::tounicode::FontCMaps;
use crate::types::{FontEncodingMap, FontWidthInfo, PageFontEncodings, PageFontWidths};
use log::debug;
use lopdf::{Document, Encoding, Object};
use lopdf::{Document, Encoding, Object, ObjectId};
use std::collections::HashMap;
#[derive(Debug, Copy, Clone, PartialEq, Eq)]
@@ -739,6 +739,171 @@ pub(crate) fn get_font_file2_obj_num(doc: &Document, font_dict: &lopdf::Dictiona
.map(|r| r.0)
}
/// Document-scoped memo of embedded-font style flags, keyed by the
/// FontFile2/FontFile3 stream's object id. The same font program is
/// referenced from every page that uses the font, and decompressing +
/// parsing it dominates `descriptor_style_flags` — without the memo that
/// cost repeats per page whenever the descriptor leaves a flag unset
/// (the common case: regular fonts report neither italic nor bold).
#[derive(Debug, Default)]
pub(crate) struct FontStyleCache {
by_font_file: HashMap<ObjectId, (bool, bool)>,
}
impl FontStyleCache {
pub(crate) fn new() -> Self {
Self::default()
}
}
/// Style flags from the FontDescriptor, which survive subset fonts whose
/// BaseFont names are opaque tags ("Tc1", "ABCDEF+F1") that defeat the
/// name-based bold/italic heuristics.
///
/// Italic: `ItalicAngle` beyond a few degrees, or Flags bit 7 (Italic,
/// value 64). Bold: Flags bit 19 (ForceBold, value 1<<18). The small
/// ItalicAngle threshold skips fonts that declare a token slant.
pub(crate) fn descriptor_style_flags(
doc: &Document,
font_dict: &lopdf::Dictionary,
style_cache: &mut FontStyleCache,
) -> (bool, bool) {
let descriptor = font_dict
.get(b"FontDescriptor")
.ok()
.and_then(|obj| resolve_dict(doc, obj))
.or_else(|| {
// Type0 fonts hang the descriptor off DescendantFonts[0].
let desc_fonts = font_dict.get(b"DescendantFonts").ok()?;
let desc_fonts = resolve_array(doc, desc_fonts)?;
let cid_font_dict = resolve_dict(doc, desc_fonts.first()?)?;
resolve_dict(doc, cid_font_dict.get(b"FontDescriptor").ok()?)
});
let Some(descriptor) = descriptor else {
return (false, false);
};
let italic_angle = descriptor
.get(b"ItalicAngle")
.ok()
.and_then(|obj| match obj {
Object::Integer(i) => Some(*i as f32),
Object::Real(r) => Some(*r),
_ => None,
})
.unwrap_or(0.0);
let flags = descriptor
.get(b"Flags")
.ok()
.and_then(|obj| obj.as_i64().ok())
.unwrap_or(0);
let mut italic = italic_angle.abs() >= 4.0 || flags & (1 << 6) != 0;
let mut bold = flags & (1 << 18) != 0;
// Descriptors lie: subset generators write ItalicAngle 0 for genuinely
// italic faces. The embedded font file keeps the truth — OS/2
// fsSelection (via `Face::is_italic`) and the post table's italicAngle.
if !italic || !bold {
if let Some(ff_ref) = font_file_ref(descriptor) {
let (emb_italic, emb_bold) = *style_cache
.by_font_file
.entry(ff_ref)
.or_insert_with(|| embedded_style_flags(doc, ff_ref));
italic = italic || emb_italic;
bold = bold || emb_bold;
}
}
(italic, bold)
}
/// Style flags parsed from an embedded font program stream.
fn embedded_style_flags(doc: &Document, ff_ref: ObjectId) -> (bool, bool) {
let Some(data) = font_file_data(doc, ff_ref) else {
return (false, false);
};
if let Ok(face) = ttf_parser::Face::parse(&data, 0) {
(
face.is_italic() || face.italic_angle().abs() >= 4.0,
face.is_bold(),
)
} else if let Some(name) = cff_font_name(&data) {
// FontFile3 is bare CFF (no sfnt container) — ttf_parser
// can't open it, but the CFF Name INDEX keeps the real
// PostScript name ("XXXXXX+Amplitude-LightItalic") even
// when the descriptor was rewritten to claim upright.
(
crate::text_utils::is_italic_font(&name),
crate::text_utils::is_bold_font(&name),
)
} else {
(false, false)
}
}
/// First PostScript name from a bare CFF font's Name INDEX (CFF spec §7).
fn cff_font_name(data: &[u8]) -> Option<String> {
// Header: major(1) minor(1) hdrSize(1) offSize(1); major must be 1.
if data.len() < 4 || data[0] != 1 {
return None;
}
let hdr_size = data[2] as usize;
// Name INDEX: count(u16) offSize(u8) offsets[count+1] data
let count = u16::from_be_bytes([*data.get(hdr_size)?, *data.get(hdr_size + 1)?]) as usize;
if count == 0 {
return None;
}
let off_size = *data.get(hdr_size + 2)? as usize;
if !(1..=4).contains(&off_size) {
return None;
}
let read_offset = |idx: usize| -> Option<usize> {
let at = hdr_size + 3 + idx * off_size;
let bytes = data.get(at..at + off_size)?;
let mut v = 0usize;
for b in bytes {
v = (v << 8) | *b as usize;
}
Some(v)
};
let start = read_offset(0)?;
let end = read_offset(1)?;
if start == 0 || end < start {
return None;
}
// Offsets are 1-based from the byte before the object data.
let objects_base = hdr_size + 3 + (count + 1) * off_size - 1;
let name = data.get(objects_base + start..objects_base + end)?;
Some(String::from_utf8_lossy(name).to_string())
}
/// FontFile2/FontFile3 stream reference from a FontDescriptor.
fn font_file_ref(descriptor: &lopdf::Dictionary) -> Option<ObjectId> {
descriptor
.get(b"FontFile2")
.ok()
.and_then(|o| o.as_reference().ok())
.or_else(|| {
descriptor
.get(b"FontFile3")
.ok()
.and_then(|o| o.as_reference().ok())
})
}
/// Decompressed embedded font program bytes.
fn font_file_data(doc: &Document, ff_ref: ObjectId) -> Option<Vec<u8>> {
let stream = doc
.get_object(ff_ref)
.and_then(lopdf::Object::as_stream)
.ok()?;
Some(
stream
.decompressed_content()
.unwrap_or_else(|_| stream.content.clone()),
)
}
/// Decode text from a PDF string operand using font CMaps, encodings, and fallbacks.
#[allow(clippy::too_many_arguments)]
pub(crate) fn extract_text_from_operand(
@@ -1262,6 +1427,172 @@ mod tests {
}
}
fn doc_with_descriptor(descriptor: lopdf::Dictionary) -> (Document, lopdf::Dictionary) {
let mut doc = Document::with_version("1.4");
let desc_id = doc.add_object(descriptor);
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "TrueType",
"BaseFont" => "Tc1",
"FontDescriptor" => desc_id,
};
(doc, font_dict)
}
#[test]
fn descriptor_italic_angle_sets_italic() {
// Subset font with an opaque BaseFont name ("Tc1") — the name
// heuristic sees nothing, the descriptor carries the truth.
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => -12,
"Flags" => 32,
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(true, false)
);
}
#[test]
fn descriptor_italic_flag_bit_sets_italic() {
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => 0,
"Flags" => 64, // bit 7: Italic
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(true, false)
);
}
#[test]
fn descriptor_force_bold_flag_sets_bold() {
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => 0,
"Flags" => 1 << 18, // ForceBold
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, true)
);
}
#[test]
fn tiny_italic_angle_is_not_italic() {
// A token 1-degree slant is optical correction, not italic.
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => lopdf::Object::Real(-1.0),
"Flags" => 32,
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, false)
);
}
#[test]
fn missing_descriptor_yields_no_flags() {
let doc = Document::with_version("1.4");
let font_dict = dictionary! { "Type" => "Font", "BaseFont" => "Tc1" };
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, false)
);
}
#[test]
fn type0_descendant_descriptor_is_resolved() {
let mut doc = Document::with_version("1.4");
let desc_id = doc.add_object(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "ABCDEF+F1",
"ItalicAngle" => -15,
});
let cid_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "CIDFontType2",
"FontDescriptor" => desc_id,
});
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "Type0",
"BaseFont" => "ABCDEF+F1",
"DescendantFonts" => vec![lopdf::Object::Reference(cid_id)],
};
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(true, false)
);
}
/// Bare CFF: header + Name INDEX only — enough for `cff_font_name`.
fn bare_cff_with_name(name: &str) -> Vec<u8> {
let mut data = vec![1, 0, 4, 1]; // major, minor, hdrSize, offSize
data.extend_from_slice(&1u16.to_be_bytes()); // Name INDEX count
data.push(1); // offSize
data.push(1); // offset of first name
data.push(1 + name.len() as u8); // offset past last name
data.extend_from_slice(name.as_bytes());
data
}
#[test]
fn embedded_font_style_is_cached_by_font_file_object() {
use lopdf::{Object, Stream};
let mut doc = Document::with_version("1.4");
let ff_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
bare_cff_with_name("ABCDEF+Test-BoldItalic"),
)));
let desc_id = doc.add_object(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "ABCDEF+Test-BoldItalic",
"ItalicAngle" => 0,
"Flags" => 32,
"FontFile3" => ff_id,
});
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Tc1",
"FontDescriptor" => desc_id,
};
let mut cache = FontStyleCache::new();
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut cache),
(true, true)
);
assert_eq!(cache.by_font_file.len(), 1);
// Replace the font program with garbage: a repeat call must serve
// the memo instead of re-reading the stream — repeated per-page
// decompression is exactly what the cache exists to avoid.
doc.objects.insert(
ff_id,
Object::Stream(Stream::new(dictionary! {}, vec![0u8; 4])),
);
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut cache),
(true, true)
);
// A cold cache parses the (now garbage) stream, proving the warm
// call above answered from the memo.
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, false)
);
}
#[test]
fn compute_string_width_ts_no_tc_tw() {
// Without Tc/Tw (both 0), width = glyph widths only
+2
View File
@@ -1495,6 +1495,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1625,6 +1626,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
+2
View File
@@ -79,6 +79,7 @@ pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> V
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Link(url),
mcid: None,
});
@@ -318,6 +319,7 @@ pub(crate) fn walk_form_fields(
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::FormField,
mcid: None,
});
+71 -4
View File
@@ -24,6 +24,7 @@ use links::{extract_form_fields, extract_page_links};
// Re-export public types so existing `crate::extractor::X` paths keep working.
pub use crate::text_utils::{is_bold_font, is_italic_font};
pub use crate::types::{ItemType, TextLine};
pub(crate) use fonts::FontStyleCache;
pub(crate) use layout::detect_columns;
pub use layout::group_into_lines;
pub(crate) use layout::group_into_lines_with_thresholds;
@@ -161,6 +162,9 @@ fn extract_positioned_text_impl(
let mut all_lines = Vec::new();
let mut page_thresholds: PageThresholds = HashMap::new();
let mut gid_encoded_pages: HashSet<u32> = HashSet::new();
// Embedded-font style flags are document-scoped: the same font program
// is shared across pages, so parse it once, not once per page.
let mut style_cache = FontStyleCache::new();
// Build page ObjectId → page number map for form field extraction
let page_id_to_num: HashMap<ObjectId, u32> =
@@ -172,8 +176,14 @@ fn extract_positioned_text_impl(
continue;
}
}
let ((mut items, rects, lines), has_gid_fonts, _coords_rotated) =
extract_page_text_items(doc, page_id, *page_num, font_cmaps, include_invisible)?;
let ((mut items, rects, lines), has_gid_fonts, _coords_rotated) = extract_page_text_items(
doc,
page_id,
*page_num,
font_cmaps,
include_invisible,
&mut style_cache,
)?;
if has_gid_fonts {
gid_encoded_pages.insert(*page_num);
}
@@ -234,7 +244,10 @@ fn suppress_table_underlines(
lines: &[PdfLine],
page: u32,
) {
if !items.iter().any(|item| item.is_underline) {
if !items
.iter()
.any(|item| item.is_underline || item.is_strikeout)
{
return;
}
@@ -256,6 +269,7 @@ fn suppress_table_underlines(
for index in table_item_indices {
if let Some(item) = items.get_mut(index) {
item.is_underline = false;
item.is_strikeout = false;
}
}
}
@@ -576,6 +590,7 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
if next.is_bold != first.is_bold
|| next.is_italic != first.is_italic
|| next.is_underline != first.is_underline
|| next.is_strikeout != first.is_strikeout
{
break;
}
@@ -638,6 +653,7 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
is_bold: first.is_bold,
is_italic: first.is_italic,
is_underline: first.is_underline,
is_strikeout: first.is_strikeout,
item_type: first.item_type.clone(),
mcid: first.mcid,
});
@@ -715,7 +731,9 @@ pub(crate) fn merge_subscript_items(items: Vec<TextItem>) -> Vec<TextItem> {
.chars()
.last()
.is_some_and(|c| c.is_alphabetic());
if parent.font_size >= sub_threshold && ends_with_letter {
let same_marks = parent.is_underline == item.is_underline
&& parent.is_strikeout == item.is_strikeout;
if parent.font_size >= sub_threshold && ends_with_letter && same_marks {
let parent_right = parent.x + parent.width;
let gap = item.x - parent_right;
// Subscripts must be tightly adjacent (within ~1pt)
@@ -789,6 +807,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -991,6 +1010,7 @@ mod tests {
items[3].y = 470.0;
for item in &mut items {
item.is_underline = true;
item.is_strikeout = true;
}
let lines = vec![
make_line(100.0, 500.0, 300.0, 500.0),
@@ -1004,6 +1024,30 @@ mod tests {
suppress_table_underlines(&mut items, &[], &lines, 1);
assert!(items.iter().all(|item| !item.is_underline));
assert!(items.iter().all(|item| !item.is_strikeout));
}
#[test]
fn subscript_digit_with_different_marks_is_not_absorbed() {
// A struck-out word followed by an unmarked footnote digit: merging
// would widen the parent's strikeout claim over the digit (and the
// reverse would drop the digit's own mark). Style boundaries break
// the merge, as in merge_text_items.
let mut word = make_merge_item("word", 100.0, 24.0);
word.font_size = 10.0;
word.is_strikeout = true;
let mut digit = make_merge_item("2", 124.5, 4.0);
digit.font_size = 6.0;
digit.y = word.y + 3.0;
let merged = merge_subscript_items(vec![word.clone(), digit.clone()]);
assert_eq!(merged.len(), 2);
// Same marks still merge (footnote ref inside the strike).
digit.is_strikeout = true;
let merged = merge_subscript_items(vec![word, digit]);
assert_eq!(merged.len(), 1);
assert!(merged[0].text.starts_with("word"));
}
#[test]
@@ -1021,6 +1065,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1036,6 +1081,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1051,6 +1097,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1106,6 +1153,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1121,6 +1169,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1136,6 +1185,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1162,6 +1212,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1177,6 +1228,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1192,6 +1244,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1220,6 +1273,7 @@ mod tests {
is_bold: true,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1255,6 +1309,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1291,6 +1346,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1306,6 +1362,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1321,6 +1378,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1344,6 +1402,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1457,6 +1516,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1472,6 +1532,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1497,6 +1558,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1512,6 +1574,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1553,6 +1616,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1598,6 +1662,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1643,6 +1708,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1681,6 +1747,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+64 -1
View File
@@ -231,8 +231,27 @@ fn rule_matches_item(rule: &Rule, item: &TextItem) -> bool {
overlap >= min_overlap
}
/// Strikeout window: a rule crossing the glyphs. Strikethroughs sit at
/// roughly 20-35% of the em above the baseline (about half the x-height);
/// accept a band well inside the glyph body so baseline underlines and
/// overlines never qualify.
fn rule_strikes_item(rule: &Rule, item: &TextItem) -> bool {
let y_min = item.y + item.font_size * 0.12;
let y_max = item.y + item.font_size * 0.55;
if rule.y < y_min || rule.y > y_max {
return false;
}
let ix1 = item.x;
let ix2 = item.x + item.width;
let min_overlap = item.width * MIN_X_OVERLAP;
let overlap = rule.x2.min(ix2) - rule.x1.max(ix1);
overlap >= min_overlap
}
/// Mark `is_underline` on text items that have a horizontal rule just
/// below their baseline. `items`, `rects`, and `lines` are a single
/// below their baseline, and `is_strikeout` on items whose glyphs a rule
/// crosses at mid x-height. `items`, `rects`, and `lines` are a single
/// page's extraction output (all in PDF coordinates, y-up, where
/// `TextItem::y` is the text baseline).
pub(crate) fn mark_underlined_items(
@@ -258,6 +277,11 @@ pub(crate) fn mark_underlined_items(
}
if rule_matches_item(rule, item) {
item.is_underline = true;
}
if rule_strikes_item(rule, item) {
item.is_strikeout = true;
}
if item.is_underline && item.is_strikeout {
break;
}
}
@@ -282,6 +306,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -358,6 +383,44 @@ mod tests {
assert!(!items[0].is_underline);
}
#[test]
fn mid_glyph_rule_marks_strikeout_not_underline() {
// Rule at ~30% of the em above the baseline crosses the glyphs.
let mut items = vec![item("struck out", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 503.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items[0].is_strikeout);
assert!(!items[0].is_underline);
}
#[test]
fn baseline_rule_marks_underline_not_strikeout() {
let mut items = vec![item("underlined", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 498.5)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items[0].is_underline);
assert!(!items[0].is_strikeout);
}
#[test]
fn overline_is_neither_underline_nor_strikeout() {
// Rule just above the cap height (overline / next line's rule).
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 507.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(!items[0].is_underline);
assert!(!items[0].is_strikeout);
}
#[test]
fn thin_filled_rect_at_mid_glyph_marks_strikeout() {
let mut items = vec![item("struck out", 100.0, 500.0, 60.0, 10.0)];
let rects = vec![thin_rect(100.0, 502.6, 60.0)];
mark_underlined_items(&mut items, &rects, &[], 1);
assert!(items[0].is_strikeout);
assert!(!items[0].is_underline);
}
#[test]
fn line_above_baseline_is_not_an_underline() {
// Strikethrough / overline geometry must not mark.
+27 -5
View File
@@ -1,5 +1,6 @@
//! Form XObject and image XObject extraction.
use super::fonts::descriptor_style_flags;
use crate::text_utils::{effective_font_size, expand_ligatures, is_bold_font, is_italic_font};
use crate::tounicode::FontCMaps;
use crate::types::{ItemType, TextItem};
@@ -8,7 +9,7 @@ use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache, FontStyleCache,
};
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
@@ -114,6 +115,7 @@ pub(crate) fn extract_form_xobject_text(
font_cmaps: &FontCMaps,
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
) -> Vec<TextItem> {
extract_form_xobject_text_inner(
doc,
@@ -122,10 +124,12 @@ pub(crate) fn extract_form_xobject_text(
font_cmaps,
parent_ctm,
cmap_decisions,
style_cache,
0,
)
}
#[allow(clippy::too_many_arguments)]
fn extract_form_xobject_text_inner(
doc: &Document,
form_id: ObjectId,
@@ -133,6 +137,7 @@ fn extract_form_xobject_text_inner(
font_cmaps: &FontCMaps,
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
depth: u8,
) -> Vec<TextItem> {
use lopdf::content::Content;
@@ -167,6 +172,7 @@ fn extract_form_xobject_text_inner(
let mut font_tounicode_refs: HashMap<String, u32> = HashMap::new();
let mut inline_cmaps: HashMap<String, crate::tounicode::CMapEntry> = HashMap::new();
let mut font_style_flags: HashMap<String, (bool, bool)> = HashMap::new();
for (font_name, font_dict) in &form_fonts {
let resource_name = String::from_utf8_lossy(font_name).to_string();
if let Ok(base_font) = font_dict.get(b"BaseFont") {
@@ -175,6 +181,10 @@ fn extract_form_xobject_text_inner(
font_base_names.insert(resource_name.clone(), base_name);
}
}
let style = descriptor_style_flags(doc, font_dict, style_cache);
if style != (false, false) {
font_style_flags.insert(resource_name.clone(), style);
}
match font_dict.get(b"ToUnicode") {
Ok(tounicode) => {
if let Ok(obj_ref) = tounicode.as_reference() {
@@ -272,6 +282,7 @@ fn extract_form_xobject_text_inner(
font_cmaps,
&ctm,
cmap_decisions,
style_cache,
depth + 1,
);
items.extend(nested_items);
@@ -295,6 +306,7 @@ fn extract_form_xobject_text_inner(
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Image,
mcid: None,
});
@@ -428,6 +440,10 @@ fn extract_form_xobject_text_inner(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
@@ -437,9 +453,10 @@ fn extract_form_xobject_text_inner(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -560,6 +577,10 @@ fn extract_form_xobject_text_inner(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
let scale_x = text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2];
for (text, start_w, end_w) in &sub_items {
let offset_tm = [
@@ -586,9 +607,10 @@ fn extract_form_xobject_text_inner(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
+356 -6
View File
@@ -579,6 +579,7 @@ pub fn extract_text_in_regions_mem(
let mut gid_pages: HashSet<u32> = HashSet::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
let mut style_cache = extractor::FontStyleCache::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
@@ -597,6 +598,7 @@ pub fn extract_text_in_regions_mem(
*page_num,
&font_cmaps,
false,
&mut style_cache,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -699,6 +701,7 @@ pub fn extract_tables_in_regions_mem(
let mut gid_pages: HashSet<u32> = HashSet::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
let mut style_cache = extractor::FontStyleCache::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
@@ -714,6 +717,7 @@ pub fn extract_tables_in_regions_mem(
*page_num,
&font_cmaps,
false,
&mut style_cache,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -1019,6 +1023,7 @@ pub fn detect_vector_grid_in_region_mem(
page_1idx,
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
)?;
text_utils::fix_letterspaced_items(&mut items);
@@ -1205,8 +1210,15 @@ mod vector_grid_tests {
let &page_id = pages.get(&1).unwrap();
let needed: HashSet<u32> = HashSet::from([1]);
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
let ((items, rects, _lines), _has_gid, _rotated) =
extract_page_text_items(&doc, page_id, 1, &cmaps, false).unwrap();
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
&doc,
page_id,
1,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
)
.unwrap();
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, 1);
assert_eq!(rect_tables.len(), 1, "expected one rect-detected table");
@@ -1240,8 +1252,15 @@ mod vector_grid_tests {
let &page_id = pages.get(&page_num).unwrap();
let needed: HashSet<u32> = HashSet::from([page_num]);
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
let ((items, rects, _lines), _has_gid, _rotated) =
extract_page_text_items(&doc, page_id, page_num, &cmaps, false).unwrap();
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
&doc,
page_id,
page_num,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
)
.unwrap();
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, page_num);
rect_tables
@@ -1966,6 +1985,7 @@ pub fn extract_tables_with_structure_cells_mem(
let mut page_heights: HashMap<u32, f32> = HashMap::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
let mut style_cache = extractor::FontStyleCache::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
@@ -1981,6 +2001,7 @@ pub fn extract_tables_with_structure_cells_mem(
*page_num,
&font_cmaps,
false,
&mut style_cache,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -2782,6 +2803,7 @@ fn detect_tsr_quality_issue(
page_1idx,
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
)?;
let adaptive_threshold = text_utils::fix_letterspaced_items(&mut items);
let coords = if coords_rotated {
@@ -3697,7 +3719,14 @@ fn detect_encoding_issues(markdown: &str) -> bool {
}
// Heuristic 2: dollar-as-space pattern
has_dollar_as_space_pattern(markdown)
if has_dollar_as_space_pattern(markdown) {
return true;
}
// Heuristic 3: substitution-cipher letter statistics (broken ToUnicode)
let mut stats = CipherGarbleStats::default();
stats.add_text(markdown);
stats.looks_garbled()
}
fn has_dollar_as_space_pattern(markdown: &str) -> bool {
@@ -3721,6 +3750,156 @@ fn has_dollar_as_space_pattern(markdown: &str) -> bool {
false
}
/// English letter frequencies (percent, az). Used as a natural-language
/// reference: every Latin-script language in the eval corpus (Swedish,
/// Finnish, Turkish, German, romaji) scores ≥ 0.80 cosine similarity against
/// it, while substitution-cipher text scores ~0.53.
const ENGLISH_LETTER_FREQ: [f64; 26] = [
8.2, 1.5, 2.8, 4.3, 12.7, 2.2, 2.0, 6.1, 7.0, 0.15, 0.8, 4.0, 2.4, 6.7, 7.5, 1.9, 0.1, 6.0,
6.3, 9.1, 2.8, 1.0, 2.4, 0.15, 2.0, 0.07,
];
/// Letter statistics for detecting substitution-cipher garbling: broken
/// ToUnicode CMaps that shift every character by a per-range constant (e.g.
/// `Certificate` extracted as `8VceZWZTReV`). Such text is 100% printable
/// ASCII with word-like token lengths, so it defeats `is_garbage_text` and
/// produces no replacement characters — it needs its own discriminator.
#[derive(Debug, Default)]
struct CipherGarbleStats {
/// Case-folded ASCII letter histogram.
letter_counts: [u32; 26],
ascii_letters: usize,
ascii_vowels: usize,
/// Accented Latin letters (Latin-1 Supplement through Latin Extended-B,
/// plus Latin Extended Additional). Count toward Latin dominance only.
latin_ext_letters: usize,
non_latin_letters: usize,
/// Adjacent ASCII-letter pairs, and how many of them switch from
/// lowercase straight to uppercase mid-word.
letter_bigrams: usize,
case_shift_bigrams: usize,
}
impl CipherGarbleStats {
fn add_text(&mut self, text: &str) {
let mut prev: Option<char> = None;
for ch in text.chars() {
if ch.is_ascii_alphabetic() {
let idx = (ch.to_ascii_lowercase() as u8 - b'a') as usize;
self.letter_counts[idx] += 1;
self.ascii_letters += 1;
if matches!(ch.to_ascii_lowercase(), 'a' | 'e' | 'i' | 'o' | 'u') {
self.ascii_vowels += 1;
}
if let Some(p) = prev {
self.letter_bigrams += 1;
if p.is_ascii_lowercase() && ch.is_ascii_uppercase() {
self.case_shift_bigrams += 1;
}
}
prev = Some(ch);
} else {
if ch.is_alphabetic() {
if matches!(ch as u32, 0xC0..=0x24F | 0x1E00..=0x1EFF) {
self.latin_ext_letters += 1;
} else {
self.non_latin_letters += 1;
}
}
prev = None;
}
}
}
/// Cosine similarity between the observed letter histogram and English
/// letter frequencies. A shifted alphabet permutes the histogram, which
/// destroys the similarity regardless of the shift amount.
fn english_cosine(&self) -> f64 {
if self.ascii_letters == 0 {
return 1.0;
}
let n = self.ascii_letters as f64;
let mut dot = 0.0;
let mut norm_obs = 0.0;
for (count, freq) in self.letter_counts.iter().zip(ENGLISH_LETTER_FREQ) {
let p = *count as f64 / n;
dot += p * freq;
norm_obs += p * p;
}
let norm_en = ENGLISH_LETTER_FREQ
.iter()
.map(|f| f * f)
.sum::<f64>()
.sqrt();
dot / (norm_obs.sqrt() * norm_en)
}
/// Cosine similarity between the observed histogram and English
/// frequencies after sorting BOTH descending — i.e. comparing the *shape*
/// of the frequency profile, ignoring which letter sits where. A
/// substitution cipher is a bijection, so it preserves this shape exactly
/// (att10k 0.97, arbitrary shifts 0.99) regardless of case or offset.
/// Non-linguistic ASCII has a different profile: a small alphabet is far
/// steeper (random DNA 0.74, hex dumps 0.81), so the shape diverges.
fn english_shape_cosine(&self) -> f64 {
if self.ascii_letters == 0 {
return 1.0;
}
let n = self.ascii_letters as f64;
let mut obs: [f64; 26] = std::array::from_fn(|i| self.letter_counts[i] as f64 / n);
obs.sort_unstable_by(|a, b| b.total_cmp(a));
let mut en = ENGLISH_LETTER_FREQ;
en.sort_unstable_by(|a, b| b.total_cmp(a));
let dot: f64 = obs.iter().zip(en).map(|(o, e)| o * e).sum();
let norm_obs = obs.iter().map(|o| o * o).sum::<f64>().sqrt();
let norm_en = en.iter().map(|e| e * e).sum::<f64>().sqrt();
dot / (norm_obs * norm_en)
}
/// Thresholds validated against the 380-document pdf-evals snapshot
/// corpus (0 false positives) and the garbled ParseBench `att10k` page
/// (vowel ratio 0.245, case-shift rate 0.225, cosine 0.532). Closest
/// legitimate document on each axis: vowel ratio 0.264 (circuit
/// schematic), case-shift rate 0.021, cosine 0.801.
fn looks_garbled(&self) -> bool {
// Need a statistically meaningful, Latin-dominant sample.
if self.ascii_letters < 200
|| self.non_latin_letters > self.ascii_letters + self.latin_ext_letters
{
return false;
}
// Real Latin-script text keeps vowels above ~30% of letters even in
// acronym- and part-number-heavy documents; shifted text starves them.
let vowel_ratio = self.ascii_vowels as f64 / self.ascii_letters as f64;
if vowel_ratio > 0.30 {
return false;
}
// Signal 1: lowercase→uppercase transitions inside words. A shifted
// lowercase alphabet straddles the ASCII uppercase block ('i'→'Z',
// 't'→'e'), so garbled words flip case constantly. Real documents
// stay ≤ 0.02 even with camelCase identifiers.
let case_shifts = self.letter_bigrams >= 100
&& self.case_shift_bigrams as f64 >= self.letter_bigrams as f64 * 0.10;
// Signal 2: the histogram is a permutation of natural language — an
// English-like frequency SHAPE (sorted cosine high) but with letters
// in the wrong POSITIONS (unsorted cosine low). This is the signature
// of a substitution cipher and is case-independent, so it catches
// all-lowercase and all-uppercase shifts as well as case-straddling
// ones. Genuinely non-linguistic ASCII that is merely "unlike English"
// fails one of the two halves: DNA/hex dumps have too steep a profile
// (shape cosine < 0.90), while protein sequences, ticker symbols and
// base64 are not sufficiently unlike English in position (unsorted
// cosine ≥ 0.60) — so none of them are routed to OCR.
let permuted_language = self.english_cosine() < 0.60 && self.english_shape_cosine() >= 0.90;
case_shifts || permuted_language
}
}
#[derive(Debug, Default)]
struct TextQualityReport {
pages_needing_ocr: Vec<u32>,
@@ -3734,6 +3913,7 @@ struct PageTextQualityEvidence {
replacement_chars: usize,
replacement_spans: usize,
longest_replacement_run: usize,
cipher_garble: CipherGarbleStats,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
@@ -3753,6 +3933,7 @@ fn analyze_text_quality(items: &[TextItem]) -> TextQualityReport {
let evidence = evidence_by_page.entry(item.page).or_default();
evidence.chars += item.text.chars().filter(|ch| !ch.is_whitespace()).count();
evidence.cipher_garble.add_text(&item.text);
match text_span_decoding_issue_kind(&item.text) {
Some(TextSpanIssueKind::Strong) => {
@@ -3776,7 +3957,8 @@ fn analyze_text_quality(items: &[TextItem]) -> TextQualityReport {
if reasons_by_page.contains_key(&page) {
continue;
}
if page_replacement_evidence_needs_ocr(&evidence) {
if page_replacement_evidence_needs_ocr(&evidence) || evidence.cipher_garble.looks_garbled()
{
add_ocr_reason(
&mut reasons_by_page,
page,
@@ -4896,6 +5078,7 @@ mod text_cluster_column_undercount_tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -5171,6 +5354,7 @@ mod table_candidate_selection_tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -5918,6 +6102,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -5965,6 +6150,171 @@ mod tests {
assert!(!detect_encoding_issues(text));
}
/// Real garbled output from ParseBench `text_simple__att10k.pdf`: a broken
/// ToUnicode CMap shifts every character by a per-range constant, so
/// "Certificate of Designations with respect to Series C" extracts as
/// pure-ASCII ciphertext (issue #118).
const SHIFTED_CIPHER_TEXT: &str =
"8VceZWZTReVW9VdZXReZdhZeYcVdaVTeeHVcZVd8EcVWVccVUHeT:iYZSZe-(,e2'WZ]V \
;VScfRcj2&,*,*$ .'R CZdecfVehYZTYUVWZVdeYVcZXYedWY]UVcdW]X'eVcUVSeYVcVXZdecReRUR]]ed \
TCBGC@ZUReVUdfSdZUZRcZVdZdWZ]VUYVcVhZeYafcdfReeVXf]ReZV0*S$.$ZZZ$6$&ViTVae \
WceYVZdecfVedcVWVccVUe.'S&.'T&.'U&.'V&.'WSV]h(EfcdfReeYZdcVXf]ReZYV \
cVXZdecReYVcVSjRXcVVdeWfcZdYRTajWRjdfTYZdecfVeWZ]VUYVcVhZeYeYVH:8 \
cVbfVde( .'S fRcRejWTVceRZZXReZTZWZT7V]]IV]VaYV8(RUHfeYhVdeVc7V]]IV]VaYV8 \
:iYZSZe.'TeYVaVcZUVUZX9VTVSVc-&,*$";
#[test]
fn test_detect_encoding_issues_shifted_cipher_text() {
assert!(detect_encoding_issues(SHIFTED_CIPHER_TEXT));
}
#[test]
fn test_shifted_cipher_below_sample_threshold_not_flagged() {
// Fewer than 200 ASCII letters: not enough evidence to condemn.
assert!(!detect_encoding_issues("8VceZWZTReVW9VdZXReZdhZeY"));
}
#[test]
fn test_clean_prose_not_flagged_as_cipher() {
let prose = "Certificate of Designations with respect to Series C Preferred Stock, \
filed February 18, 2020. No instrument which defines the rights of holders of \
long-term debt of the registrant and all of its consolidated subsidiaries is \
filed herewith pursuant to Regulation S-K, Item 601. Pursuant to this regulation, \
the registrant hereby agrees to furnish a copy of any such instrument to the SEC \
upon request.";
assert!(!detect_encoding_issues(prose));
}
#[test]
fn test_camel_case_code_not_flagged_as_cipher() {
let code = "The getElementById and querySelectorAll methods return DOM nodes. Use \
addEventListener with removeEventListener, requestAnimationFrame with \
cancelAnimationFrame, and setTimeout with clearTimeout. The XMLHttpRequest \
object exposes onreadystatechange, responseText and getAllResponseHeaders. \
Prefer createElement, appendChild, insertBefore and replaceChild for DOM \
manipulation, and getBoundingClientRect for layout measurement.";
assert!(!detect_encoding_issues(code));
}
#[test]
fn test_accented_european_text_not_flagged_as_cipher() {
let swedish = "Regeringen föreslår att riksdagen antar förslaget till lag om ändring \
i skatteförfarandelagen. Bestämmelserna föreslås träda i kraft den första januari. \
Förslaget innebär att företag med säte i utlandet måste lämna särskilda uppgifter \
till Skatteverket varje kvartal, och att avgiften höjs för överträdelser av de nya \
bestämmelserna om rapporteringsskyldighet för gränsöverskridande arrangemang.";
assert!(!detect_encoding_issues(swedish));
}
#[test]
fn test_all_caps_text_not_flagged_as_cipher() {
let caps = "EXHIBIT INDEX PURSUANT TO ITEM 601 OF REGULATION SK CERTIFICATE OF \
DESIGNATIONS WITH RESPECT TO SERIES C PREFERRED STOCK FILED FEBRUARY EIGHTEEN \
TWENTY TWENTY AND INCORPORATED HEREIN BY REFERENCE TO THE ANNUAL REPORT ON FORM \
TENK FOR THE PERIOD ENDED DECEMBER THIRTYFIRST TWENTY NINETEEN AS AMENDED";
assert!(!detect_encoding_issues(caps));
}
// A Caesar shift of prose that stays within a single case block does not
// trigger the case-shift signal, but it permutes the letter histogram: the
// frequency SHAPE stays English-like while letter POSITIONS scramble, so
// the permutation signal catches it regardless of case.
const CAESAR_PROSE: &str =
"The registrant hereby agrees to furnish a copy of any such instrument to the \
Commission upon request. This certificate of designations was filed February with \
respect to Series Preferred Stock and incorporated herein by reference to the annual \
report on form for the period ended December as amended and restated thereafter.";
fn caesar_shift(text: &str, k: u8) -> String {
text.chars()
.map(|c| match c {
'a'..='z' => (((c as u8 - b'a' + k) % 26) + b'a') as char,
'A'..='Z' => (((c as u8 - b'A' + k) % 26) + b'A') as char,
_ => c,
})
.collect()
}
#[test]
fn test_mixed_case_caesar_shift_flagged() {
assert!(detect_encoding_issues(&caesar_shift(CAESAR_PROSE, 3)));
}
#[test]
fn test_all_lowercase_caesar_shift_flagged() {
// Uniform all-lowercase garbled prose: the earlier mixed-case guard
// would have exempted this, so it must be caught by the case-agnostic
// permutation signal instead.
assert!(detect_encoding_issues(&caesar_shift(
&CAESAR_PROSE.to_lowercase(),
5
)));
}
#[test]
fn test_all_uppercase_caesar_shift_flagged() {
assert!(detect_encoding_issues(&caesar_shift(
&CAESAR_PROSE.to_uppercase(),
7
)));
}
#[test]
fn test_dna_sequence_not_flagged_as_cipher() {
// Non-linguistic ASCII: unlike English (low vowel ratio, low cosine)
// but not garbled. Its 4-letter alphabet makes the frequency profile
// too steep, so the shape cosine falls below the permutation threshold.
let dna = "ACGT".repeat(120);
assert!(!detect_encoding_issues(&dna));
}
#[test]
fn test_protein_sequence_not_flagged_as_cipher() {
let protein =
"MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEKAVQVKVKALPDAQFEVVHSLAKWKR"
.repeat(6);
assert!(!detect_encoding_issues(&protein));
}
#[test]
fn test_ticker_list_not_flagged_as_cipher() {
let tickers = "AAPL MSFT GOOG TSLA NVDA AMZN META NFLX AMD INTC CSCO ORCL CRM ADBE QCOM \
TXN AVGO MU LRCX KLAC ASML SNPS CDNS FTNT PANW "
.repeat(4);
assert!(!detect_encoding_issues(&tickers));
}
#[test]
fn test_text_quality_flags_shifted_cipher_page() {
let items = vec![
test_text_item_on_page(1, SHIFTED_CIPHER_TEXT),
test_text_item_on_page(2, "A clean second page should not be routed to OCR."),
];
let quality = analyze_text_quality(&items);
assert!(quality.has_encoding_issues);
assert_eq!(quality.pages_needing_ocr, vec![1]);
assert_eq!(
quality.reasons_by_page.get(&1).cloned(),
Some(vec![OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()])
);
}
#[test]
fn test_text_quality_cipher_stats_accumulate_across_items() {
// The garbled page arrives as many short spans; no single span has
// enough letters to flag on its own.
let items: Vec<TextItem> = SHIFTED_CIPHER_TEXT
.split_whitespace()
.map(|chunk| test_text_item_on_page(1, chunk))
.collect();
let quality = analyze_text_quality(&items);
assert_eq!(quality.pages_needing_ocr, vec![1]);
}
#[test]
fn test_text_quality_flags_localized_cid_mojibake_span() {
let items = vec![
+48
View File
@@ -130,6 +130,29 @@ pub(crate) fn has_dot_leaders(text: &str) -> bool {
dot_groups >= 2
}
/// Detect a table-of-contents entry: a line ending in a page number preceded by
/// a dot-leader group (e.g. "Measurement Lab worksheet ... 3"). `has_dot_leaders`
/// misses single-group leaders ("..."), but a trailing "<dots> <number>" is a
/// strong TOC signal on its own. Such lines must never be promoted to headings.
pub(crate) fn is_toc_entry_line(text: &str) -> bool {
let trimmed = text.trim_end();
let digits = trimmed
.chars()
.rev()
.take_while(|c| c.is_ascii_digit())
.count();
if digits == 0 || digits > 4 {
return false;
}
let before_number = trimmed[..trimmed.len() - digits].trim_end();
let dots = before_number
.chars()
.rev()
.take_while(|c| *c == '.')
.count();
dots >= 3
}
/// Compute the Y-gap threshold for paragraph break detection.
///
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
@@ -320,3 +343,28 @@ pub(crate) fn detect_header_level(
Some(4)
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn toc_entry_with_single_dot_group() {
assert!(is_toc_entry_line("Measurement Lab worksheet ... 3"));
assert!(is_toc_entry_line("Results ........ 12"));
assert!(is_toc_entry_line("Appendix B...42"));
}
#[test]
fn non_toc_lines_pass() {
assert!(!is_toc_entry_line(
"6.2. Expectations for Re-Hiring Employees"
));
assert!(!is_toc_entry_line("What happened in 2020"));
assert!(!is_toc_entry_line("IMPLEMENTATION"));
// Ellipsis without a trailing page number
assert!(!is_toc_entry_line("and so it goes ..."));
// Long numbers are data, not page refs
assert!(!is_toc_entry_line("ISBN ... 97814"));
}
}
+12 -3
View File
@@ -7,7 +7,7 @@ use crate::types::TextLine;
use super::analysis::{
bold_heading_level, calculate_font_stats, compute_heading_tiers, compute_paragraph_threshold,
detect_header_level, font_size_rarity, has_dot_leaders,
detect_header_level, font_size_rarity, has_dot_leaders, is_toc_entry_line,
};
use super::classify::{
format_list_item, is_caption_line, is_list_item, is_monospace_font, starts_with_bullet_marker,
@@ -703,6 +703,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !starts_with_bullet_marker(plain_trimmed)
&& !is_toc_entry_line(plain_trimmed)
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
detect_header_level(line_font_size, base_size, &heading_tiers).or_else(|| {
@@ -738,7 +739,11 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
// paragraph continuity and minor font-size variation
// inflates rarity scores.
let has_strong_signal = all_bold || isolated || (rarity >= 0.97 && word_count <= 8);
if score >= 0.5 && standalone && word_count >= 2 && has_strong_signal {
// Single-word headings ("IMPLEMENTATION", "CONTENTS") are common;
// accept them only with the strongest signal combination.
let enough_words =
word_count >= 2 || (all_bold && isolated && plain_trimmed.len() >= 4);
if score >= 0.5 && standalone && enough_words && has_strong_signal {
Some(bold_heading_level(&heading_tiers))
} else {
None
@@ -1037,6 +1042,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
if options.detect_headers
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !is_toc_entry_line(plain_trimmed)
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
if let Some(header_level) =
@@ -1059,7 +1065,9 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
+ if all_bold { 0.3 } else { 0.0 }
+ if standalone { 0.2 } else { 0.0 }
+ if isolated { 0.3 } else { 0.0 };
if score >= 0.5 && standalone && word_count >= 2 {
let enough_words =
word_count >= 2 || (all_bold && isolated && plain_trimmed.len() >= 4);
if score >= 0.5 && standalone && enough_words {
return Some(bold_heading_level(&heading_tiers));
}
None
@@ -1189,6 +1197,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid,
}
+1
View File
@@ -1221,6 +1221,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
}
+1
View File
@@ -543,6 +543,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid,
}
+3
View File
@@ -268,6 +268,8 @@ pub struct PyTextItem {
#[pyo3(get)]
pub is_underline: bool,
#[pyo3(get)]
pub is_strikeout: bool,
#[pyo3(get)]
pub item_type: String,
}
@@ -352,6 +354,7 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item_type_str(&item.item_type),
})
.collect()
+1
View File
@@ -105,6 +105,7 @@ pub(crate) fn merge_adjacent_items(items: &[TextItem]) -> (Vec<TextItem>, Vec<Ve
is_bold: first_item.is_bold,
is_italic: first_item.is_italic,
is_underline: first_item.is_underline,
is_strikeout: first_item.is_strikeout,
item_type: first_item.item_type.clone(),
mcid: first_item.mcid,
});
+1
View File
@@ -394,6 +394,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+3
View File
@@ -2392,6 +2392,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -3366,6 +3367,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
});
@@ -3676,6 +3678,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
});
+1
View File
@@ -587,6 +587,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid,
}
+1
View File
@@ -109,6 +109,7 @@ pub(crate) fn try_split_financial_item(item: &TextItem) -> Option<Vec<TextItem>>
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
+3
View File
@@ -521,6 +521,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -886,6 +887,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
@@ -923,6 +925,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
+4
View File
@@ -236,6 +236,7 @@ fn split_merged_numbers(item: &TextItem, col_boundaries: &[f32]) -> Vec<TextItem
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
@@ -257,6 +258,7 @@ fn split_merged_numbers(item: &TextItem, col_boundaries: &[f32]) -> Vec<TextItem
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
@@ -1432,6 +1434,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1450,6 +1453,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+3
View File
@@ -884,6 +884,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1004,6 +1005,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -1081,6 +1083,7 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+4
View File
@@ -120,6 +120,10 @@ pub struct TextItem {
/// baseline — PDFs have no underline font flag, so this is detected
/// geometrically after extraction; see `extractor::underline`).
pub is_underline: bool,
/// Whether the text is struck out (drawn rule/thin rect crossing the
/// glyphs at mid x-height). Same geometric detection as underline,
/// different vertical window; see `extractor::underline`.
pub is_strikeout: bool,
/// Type of item (text, image, link)
pub item_type: ItemType,
/// Marked Content ID from the content stream's BDC/BMC operator.
Binary file not shown.
+28
View File
@@ -105,6 +105,7 @@ fn make_text_item(text: &str, x: f32, y: f32, font_size: f32, page: u32) -> Text
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -131,6 +132,7 @@ fn make_text_item_with_font(
is_bold: is_bold_font(font),
is_italic: is_italic_font(font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1356,6 +1358,32 @@ fn test_extract_regions_mem_identity_h_needs_ocr() {
);
}
/// ParseBench `text_simple__att10k.pdf` (issue #118): the producer authored a
/// broken ToUnicode CMap that shifts every character by a per-range constant,
/// and the embedded subset font has no `cmap` table to recover from. The
/// resulting ciphertext is 100% printable ASCII, so it must be caught by the
/// substitution-cipher statistics and routed to OCR instead of served silently.
#[test]
fn test_extract_pages_mem_shifted_cipher_tounicode_needs_ocr() {
let buf = std::fs::read("tests/fixtures/shifted_cipher_tounicode.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, None).unwrap();
assert_eq!(result.pages.len(), 1);
assert!(
result.pages[0].needs_ocr,
"shifted-cipher garbled page should be flagged needs_ocr"
);
assert!(
result.pages[0].markdown.is_empty(),
"garbled markdown should be suppressed"
);
assert_eq!(result.pages_needing_ocr, vec![1]);
assert_eq!(
result.pages[0].ocr_reason.as_deref(),
Some("suspected_garbled_text")
);
}
#[test]
fn test_extract_regions_mem_multiple_regions_per_page() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();