Compare commits

..
Author SHA1 Message Date
Abimael MartellandCursor 4d72eea009 fix(extractor): flag printable ASCII mojibake for OCR
Detect large prose-like pages whose printable text has implausibly low vowel and common-word rates, preventing broken ToUnicode output from being served as valid extraction. Add the issue fixture regression and bump Rust/N-API/npm patch versions.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 10:39:22 -07:00
41 changed files with 561 additions and 3004 deletions
-36
View File
@@ -1,36 +0,0 @@
name: Deploy landing page
on:
push:
branches: [main]
paths: ['site/**', '.github/workflows/pages.yml']
workflow_dispatch:
permissions:
contents: read
pages: write
id-token: write
# Allow one concurrent deployment; don't cancel an in-progress production deploy.
concurrency:
group: pages
cancel-in-progress: false
jobs:
deploy:
name: Build & deploy to GitHub Pages
runs-on: ubuntu-latest
environment:
name: github-pages
url: ${{ steps.deploy.outputs.page_url }}
steps:
- uses: actions/checkout@v4
- name: Upload site artifact
uses: actions/upload-pages-artifact@v3
with:
path: site
- name: Deploy to GitHub Pages
id: deploy
uses: actions/deploy-pages@v4
-21
View File
@@ -1,21 +0,0 @@
MIT License
Copyright (c) 2026 Firecrawl
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+5 -6
View File
@@ -2,7 +2,6 @@
[![Crates.io](https://img.shields.io/crates/v/pdf-inspector.svg)](https://crates.io/crates/pdf-inspector)
[![npm](https://img.shields.io/npm/v/@firecrawl/pdf-inspector.svg)](https://www.npmjs.com/package/@firecrawl/pdf-inspector)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for [Python](docs/python.md) and [Node.js](napi/README.md).
@@ -26,16 +25,16 @@ Evaluated on the [opendataloader-bench](https://github.com/opendataloader-projec
| Engine | Overall | Reading Order (NID) | Tables (TEDS) | Headings (MHS) | Speed (200 docs) |
|---|---|---|---|---|---|
| pdf-inspector | 0.83 | 0.88 | 0.66 | 0.74 | 4s |
| pdf-inspector | 0.78 | 0.87 | 0.59 | 0.57 | 4s |
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
| markitdown | 0.58 | 0.88 | 0.00 | 0.00 | 8s |
For context, engines that use OCR/ML (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus — pdf-inspector reaches the low end of that range without any OCR, in 4 seconds.
For context, engines that use OCR/ML (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus.
**Where we do well:** Speed (fastest of all engines), the best table detection of any engine shown, and heading detection now on par with opendataloader. Overall lands within 0.01 of opendataloader at roughly 2.5× the speed.
**Where we do well:** Speed (fastest of all engines), reading order, table detection vs other direct-text tools.
**Where we lag:** Reading order still trails opendataloader slightly, and table structure trails OCR-based engines that can see visual layout.
**Where we lag:** Heading detection trails opendataloader — many PDFs use bold text at body font size for headings, or headings that are only slightly larger than body text. Table detection trails OCR-based engines that can see visual table structure.
## Quick start
@@ -243,4 +242,4 @@ See [docs/debugging.md](docs/debugging.md) for `RUST_LOG` environment variable u
## License
[MIT](LICENSE)
MIT
+1 -1
View File
@@ -845,7 +845,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-napi"
version = "0.2.2"
version = "0.2.3"
dependencies = [
"napi",
"napi-build",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-napi"
version = "0.2.2"
version = "0.2.3"
edition = "2021"
[lib]
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.10.1",
"version": "1.9.11",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
-4
View File
@@ -83,9 +83,6 @@ pub struct TextItem {
/// Underline detected geometrically (drawn rule/thin rect under the
/// baseline) — PDFs carry no underline font flag.
pub is_underline: bool,
/// Strikeout detected geometrically (rule crossing the glyphs at mid
/// x-height).
pub is_strikeout: bool,
pub item_type: ItemType,
/// URL for link items, `None` for other types.
pub link_url: Option<String>,
@@ -297,7 +294,6 @@ pub fn extract_text_with_positions(
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type,
link_url,
}
-1
View File
@@ -39,7 +39,6 @@ class TextItem:
is_bold: bool
is_italic: bool
is_underline: bool
is_strikeout: bool
item_type: str
class RegionText:
+1 -3
View File
@@ -4,9 +4,7 @@ build-backend = "maturin"
[project]
name = "pdf-inspector"
# Version is sourced from Cargo.toml [package] version by maturin so the Python
# artifact always tracks the crate release instead of drifting on its own.
dynamic = ["version"]
version = "0.1.0"
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
license = { text = "MIT" }
requires-python = ">=3.8"
-3
View File
@@ -1,3 +0,0 @@
<svg width="200" height="284" viewBox="0 0 200 284" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M166.862 90.7716C155.812 94.0514 147.483 101.471 141.383 109.53C140.073 111.26 137.343 109.96 137.863 107.841C149.543 59.8136 134.113 19.896 86.0157 0.247269C83.5758 -0.752669 81.0359 1.43719 81.6759 3.99704C103.555 91.8416 11.5294 84.432 23.1588 184.016C23.3588 185.726 21.4389 186.896 20.039 185.896C15.6792 182.766 10.8095 176.236 7.46963 171.647C6.48968 170.297 4.36978 170.677 3.9198 172.287C1.25994 181.906 0 190.965 0 199.965C0 234.963 17.9891 265.771 45.2177 283.63C46.7777 284.65 48.7776 283.19 48.2476 281.4C46.8477 276.7 46.0577 271.74 45.9977 266.611C45.9977 263.461 46.1977 260.241 46.6877 257.241C47.8276 249.702 50.4475 242.522 54.8473 235.983C69.9365 213.334 100.185 191.455 95.3552 161.747C95.0453 159.867 97.2651 158.627 98.6651 159.917C119.974 179.386 124.194 205.575 120.694 229.063C120.394 231.103 122.954 232.193 124.244 230.593C127.504 226.513 131.483 222.933 135.813 220.244C136.893 219.574 138.333 220.084 138.743 221.284C141.153 228.293 144.733 234.873 148.113 241.452C152.152 249.362 154.302 258.391 153.962 267.951C153.792 272.6 153.022 277.1 151.732 281.38C151.182 283.19 153.162 284.7 154.752 283.66C182.001 265.801 200 234.993 200 199.975C200 187.806 197.87 175.876 193.84 164.697C185.391 141.248 163.952 123.64 169.372 93.0815C169.632 91.6216 168.282 90.3517 166.862 90.7716Z" fill="#FA5D19" style="fill:#FA5D19;fill:color(display-p3 0.9816 0.3634 0.0984);fill-opacity:1;"/>
</svg>

Before

Width:  |  Height:  |  Size: 1.5 KiB

-12
View File
@@ -1,12 +0,0 @@
<svg width="172" height="40" viewBox="0 0 172 40" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M23.3606 12.8281C21.8137 13.2873 20.6476 14.3261 19.7936 15.4544C19.6102 15.6966 19.228 15.5146 19.3008 15.2178C20.936 8.49401 18.7759 2.90556 12.0422 0.154735C11.7006 0.0147436 11.345 0.321324 11.4346 0.679702C14.4977 12.9779 1.61412 11.9406 3.24224 25.8823C3.27024 26.1217 3.00145 26.2855 2.80546 26.1455C2.19509 25.7073 1.51332 24.7932 1.04575 24.1506C0.908555 23.9616 0.611769 24.0148 0.548773 24.2402C0.176391 25.5869 0 26.8553 0 28.1152C0 33.0149 2.51847 37.328 6.33048 39.8283C6.54887 39.9711 6.82886 39.7667 6.75466 39.5161C6.55867 38.8581 6.44808 38.1638 6.43968 37.4456C6.43968 37.0046 6.46768 36.5539 6.53627 36.1339C6.69587 35.0784 7.06265 34.0732 7.67862 33.1577C9.79111 29.9869 14.0259 26.9239 13.3497 22.7647C13.3063 22.5015 13.6171 22.328 13.8131 22.5085C16.7964 25.2342 17.3871 28.9005 16.8972 32.1889C16.8552 32.4745 17.2135 32.6271 17.3941 32.4031C17.8505 31.832 18.4077 31.3308 19.0138 30.9542C19.165 30.8604 19.3666 30.9318 19.424 31.0998C19.7614 32.0811 20.2626 33.0023 20.7358 33.9234C21.3013 35.0308 21.6023 36.2949 21.5547 37.6332C21.5309 38.2842 21.4231 38.9141 21.2425 39.5133C21.1655 39.7667 21.4427 39.9781 21.6653 39.8325C25.4801 37.3322 28 33.0191 28 28.1166C28 26.4129 27.7018 24.7428 27.1376 23.1777C25.9547 19.8949 22.9533 17.4297 23.712 13.1515C23.7484 12.9471 23.5594 12.7693 23.3606 12.8281Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M41 34.0521V10.9618H55.7586V14.3264H44.7969V21.0226H53.8436V24.2882H44.7969V34.0521H41Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M59.9569 14.7882C58.7352 14.7882 57.7777 13.8976 57.7777 12.6441C57.7777 11.3906 58.7352 10.5 59.9569 10.5C61.1785 10.5 62.136 11.3906 62.136 12.6441C62.136 13.8976 61.1785 14.7882 59.9569 14.7882ZM58.1409 34.0521V17.1632H61.7068V34.0521H58.1409Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M73.5885 17.1632H74.3809V20.4948H72.796C69.6264 20.4948 68.6029 22.9687 68.6029 25.5747V34.0521H65.0371V17.1632H68.2067L68.6029 19.7031C69.4613 18.2847 70.815 17.1632 73.5885 17.1632Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M83.632 34.25C78.3163 34.25 74.9816 30.8194 74.9816 25.6406C74.9816 20.4288 78.3163 16.9653 83.3019 16.9653C88.1884 16.9653 91.457 20.066 91.5561 25.0139C91.5561 25.4427 91.5231 25.9045 91.457 26.3663H78.7125V26.5972C78.8116 29.467 80.6275 31.3472 83.4339 31.3472C85.613 31.3472 87.1979 30.2587 87.6931 28.3785H91.2589C90.6646 31.7101 87.8252 34.25 83.632 34.25ZM78.8446 23.7604H87.8582C87.561 21.2535 85.8112 19.8351 83.3349 19.8351C81.0567 19.8351 79.1087 21.3524 78.8446 23.7604Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M102.033 34.25C96.9151 34.25 93.6465 30.9184 93.6465 25.6406C93.6465 20.4288 97.0142 16.9653 102.132 16.9653C106.49 16.9653 109.197 19.3733 109.891 23.1997H106.16C105.698 21.2205 104.278 20 102.066 20C99.1933 20 97.3113 22.309 97.3113 25.6406C97.3113 28.9392 99.1933 31.2153 102.066 31.2153C104.245 31.2153 105.698 29.9618 106.127 28.0156H109.891C109.23 31.842 106.358 34.25 102.033 34.25Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M121.006 17.1632H121.799V20.4948H120.214C117.044 20.4948 116.021 22.9687 116.021 25.5747V34.0521H112.455V17.1632H115.625L116.021 19.7031C116.879 18.2847 118.233 17.1632 121.006 17.1632Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M130.614 16.9653C135.104 16.9653 137.679 19.1094 137.679 23.1007V34.0521H134.576L134.279 31.6441C133.123 33.1615 131.505 34.25 128.831 34.25C125.133 34.25 122.657 32.4358 122.657 29.3021C122.657 25.8385 125.166 23.8924 129.92 23.8924H134.147V22.8698C134.147 20.9896 132.793 19.8351 130.449 19.8351C128.336 19.8351 126.916 20.8247 126.652 22.309H123.152C123.515 19.0104 126.355 16.9653 130.614 16.9653ZM129.425 31.4792C132.397 31.4792 134.114 29.7309 134.147 27.125V26.5312H129.722C127.51 26.5312 126.289 27.3559 126.289 29.0712C126.289 30.4896 127.477 31.4792 129.425 31.4792Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M144.653 34.0521L139.139 17.1632H142.903L146.766 30.0937L150.629 17.1632H153.897L157.595 30.0937L161.59 17.1632H165.222L159.609 34.0521H155.779L152.214 22.5729L148.516 34.0521H144.653Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
<path d="M166.934 34.0521V10.9618H170.5V34.0521H166.934Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
</svg>

Before

Width:  |  Height:  |  Size: 4.8 KiB

-428
View File
@@ -1,428 +0,0 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>pdf-inspector — PDF classification &amp; text extraction, no OCR</title>
<meta name="description" content="Fast Rust library that classifies PDFs (text-based vs scanned) and extracts clean Markdown — no OCR, no ML models. Bindings for Rust, Python, and Node.js.">
<meta property="og:title" content="pdf-inspector">
<meta property="og:description" content="Classify PDFs and extract clean Markdown in milliseconds. No OCR. No ML. Pure Rust.">
<meta property="og:type" content="website">
<link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 100 100'%3E%3Ctext y='.9em' font-size='90'%3E%F0%9F%93%84%3C/text%3E%3C/svg%3E">
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Bricolage+Grotesque:opsz,wght@12..96,400;12..96,600;12..96,800&family=Hanken+Grotesk:wght@400;500;600&family=JetBrains+Mono:wght@400;500;700&display=swap" rel="stylesheet">
<style>
:root {
--paper: #f4efe4;
--paper-2: #eee7d8;
--ink: #1b1712;
--ink-soft: #4a433a;
--muted: #8b8375;
--line: #d9cfbb;
--accent: #dd3f22;
--accent-deep: #b32d15;
--card: #faf6ec;
--display: "Bricolage Grotesque", serif;
--body: "Hanken Grotesk", sans-serif;
--mono: "JetBrains Mono", monospace;
}
* { box-sizing: border-box; margin: 0; padding: 0; }
html { scroll-behavior: smooth; }
body {
background: var(--paper);
color: var(--ink);
font-family: var(--body);
font-size: 17px;
line-height: 1.6;
-webkit-font-smoothing: antialiased;
overflow-x: hidden;
background-image:
radial-gradient(circle at 1px 1px, rgba(27,23,18,0.05) 1px, transparent 0);
background-size: 22px 22px;
}
::selection { background: var(--accent); color: var(--paper); }
a { color: inherit; text-decoration: none; }
.wrap { max-width: 1120px; margin: 0 auto; padding: 0 28px; }
/* ── nav ── */
nav {
position: sticky; top: 0; z-index: 50;
background: rgba(244,239,228,0.82);
backdrop-filter: blur(10px);
border-bottom: 1px solid var(--line);
}
.nav-in { display: flex; align-items: center; gap: 22px; height: 60px; }
.brand { font-family: var(--mono); font-weight: 700; font-size: 15px; letter-spacing: -0.02em; display: flex; align-items: center; gap: 9px; }
.brand .dot { width: 9px; height: 9px; background: var(--accent); border-radius: 50%; box-shadow: 0 0 0 3px rgba(221,63,34,0.18); }
.nav-links { margin-left: auto; display: flex; gap: 24px; align-items: center; font-size: 14.5px; font-weight: 500; }
.nav-links a { color: var(--ink-soft); transition: color .15s; }
.nav-links a:hover { color: var(--accent); }
.nav-gh { border: 1px solid var(--ink); border-radius: 999px; padding: 6px 15px; color: var(--ink) !important; transition: all .15s; }
.nav-gh:hover { background: var(--ink); color: var(--paper) !important; }
@media (max-width: 680px) { .nav-links .hide-sm { display: none; } }
/* ── hero ── */
header { padding: 74px 0 40px; position: relative; }
.eyebrow { font-family: var(--mono); font-size: 12.5px; letter-spacing: 0.16em; text-transform: uppercase; color: var(--accent-deep); margin-bottom: 22px; }
h1 {
font-family: var(--display);
font-weight: 800;
font-size: clamp(2.9rem, 8vw, 6.1rem);
line-height: 0.96;
letter-spacing: -0.035em;
max-width: 15ch;
}
h1 .em { color: var(--accent); font-style: normal; position: relative; }
h1 .strike { position: relative; white-space: nowrap; }
h1 .strike::after { content: ""; position: absolute; left: -2%; right: -2%; top: 54%; height: 0.09em; background: var(--accent); transform: rotate(-3deg); }
.lede { margin-top: 28px; font-size: clamp(1.05rem, 2.2vw, 1.32rem); color: var(--ink-soft); max-width: 46ch; line-height: 1.5; }
.lede b { color: var(--ink); font-weight: 600; }
.hero-grid { display: grid; grid-template-columns: 1.15fr 0.85fr; gap: 48px; align-items: end; }
@media (max-width: 880px) { .hero-grid { grid-template-columns: 1fr; gap: 40px; } }
/* readout card */
.readout {
background: var(--ink); color: var(--paper);
border-radius: 14px; padding: 22px 22px 20px;
font-family: var(--mono); font-size: 13px;
box-shadow: 14px 14px 0 rgba(27,23,18,0.09);
position: relative;
}
.readout .rlabel { color: #b8ad98; font-size: 11px; letter-spacing: 0.14em; text-transform: uppercase; margin-bottom: 16px; display: flex; justify-content: space-between; }
.readout .rrow { display: flex; justify-content: space-between; align-items: center; padding: 9px 0; border-top: 1px solid rgba(255,255,255,0.09); }
.readout .rrow:first-of-type { border-top: none; }
.readout .k { color: #cfc6b3; }
.readout .v { font-weight: 700; }
.readout .v.hot { color: var(--accent); }
.bar { height: 6px; background: rgba(255,255,255,0.1); border-radius: 3px; overflow: hidden; margin-top: 3px; width: 96px; }
.bar > i { display: block; height: 100%; background: var(--accent); border-radius: 3px; }
/* ── install row ── */
.installs { display: grid; grid-template-columns: repeat(3,1fr); gap: 14px; margin-top: 54px; }
@media (max-width: 720px) { .installs { grid-template-columns: 1fr; } }
.inst {
background: var(--card); border: 1px solid var(--line); border-radius: 11px;
padding: 15px 17px; transition: transform .16s, border-color .16s, box-shadow .16s;
cursor: pointer;
}
.inst:hover { transform: translateY(-3px); border-color: var(--accent); box-shadow: 0 8px 22px rgba(27,23,18,0.07); }
.inst .reg { font-family: var(--mono); font-size: 11px; letter-spacing: 0.12em; text-transform: uppercase; color: var(--muted); margin-bottom: 8px; display: flex; justify-content: space-between; }
.inst code { font-family: var(--mono); font-size: 14px; color: var(--ink); font-weight: 500; }
.inst .arrow { color: var(--accent); opacity: 0; transition: opacity .16s; }
.inst:hover .arrow { opacity: 1; }
/* ── section scaffold ── */
section { padding: 66px 0; border-top: 1px solid var(--line); }
.sec-head { display: flex; align-items: baseline; gap: 16px; margin-bottom: 40px; flex-wrap: wrap; }
.sec-num { font-family: var(--mono); font-size: 13px; color: var(--accent); font-weight: 700; }
.sec-title { font-family: var(--display); font-weight: 600; font-size: clamp(1.7rem, 4vw, 2.6rem); letter-spacing: -0.025em; }
.sec-sub { color: var(--ink-soft); max-width: 52ch; font-size: 1.02rem; }
/* ── features ── */
.feat-grid { display: grid; grid-template-columns: repeat(4,1fr); gap: 1px; background: var(--line); border: 1px solid var(--line); border-radius: 14px; overflow: hidden; }
@media (max-width: 900px) { .feat-grid { grid-template-columns: repeat(2,1fr); } }
@media (max-width: 560px) { .feat-grid { grid-template-columns: 1fr; } }
.feat { background: var(--card); padding: 24px 22px; transition: background .16s; }
.feat:hover { background: #fff; }
.feat .fn { font-family: var(--mono); font-size: 12px; color: var(--accent); font-weight: 700; }
.feat h3 { font-family: var(--display); font-weight: 600; font-size: 1.16rem; margin: 12px 0 8px; letter-spacing: -0.01em; }
.feat p { font-size: 14.5px; color: var(--ink-soft); line-height: 1.5; }
/* ── benchmark ── */
.bench {
border: 1px solid var(--line); border-radius: 14px; overflow: hidden;
background: var(--card);
}
table { width: 100%; border-collapse: collapse; font-size: 15px; }
thead th { font-family: var(--mono); font-size: 11px; letter-spacing: 0.08em; text-transform: uppercase; color: var(--muted); text-align: right; padding: 15px 18px; background: var(--paper-2); border-bottom: 1px solid var(--line); font-weight: 500; }
thead th:first-child { text-align: left; }
tbody td { padding: 14px 18px; text-align: right; font-family: var(--mono); border-bottom: 1px solid var(--line); }
tbody td:first-child { text-align: left; font-family: var(--body); font-weight: 500; }
tbody tr:last-child td { border-bottom: none; }
tbody tr.us { background: rgba(221,63,34,0.06); }
tbody tr.us td:first-child { color: var(--accent-deep); font-weight: 700; }
tbody tr.us td:first-child::before { content: "▸ "; color: var(--accent); }
.bench-foot { padding: 15px 18px; font-size: 13.5px; color: var(--ink-soft); background: var(--paper-2); border-top: 1px solid var(--line); }
.bench-wrap { overflow-x: auto; }
.callouts { display: grid; grid-template-columns: 1fr 1fr; gap: 16px; margin-top: 22px; }
@media (max-width: 640px) { .callouts { grid-template-columns: 1fr; } }
.callout { border-left: 3px solid var(--accent); padding: 4px 0 4px 16px; }
.callout .ct { font-family: var(--mono); font-size: 11px; letter-spacing: 0.1em; text-transform: uppercase; color: var(--accent-deep); margin-bottom: 5px; }
.callout p { font-size: 14.5px; color: var(--ink-soft); }
/* ── quickstart tabs ── */
.tabs input { position: absolute; opacity: 0; pointer-events: none; }
.tablist { display: flex; gap: 6px; margin-bottom: 0; }
.tablist label {
font-family: var(--mono); font-size: 13px; font-weight: 500;
padding: 10px 18px; cursor: pointer; color: var(--muted);
border: 1px solid var(--line); border-bottom: none;
border-radius: 9px 9px 0 0; background: var(--paper-2); transition: all .15s;
}
.tablist label:hover { color: var(--ink); }
.panel { display: none; }
.code {
background: var(--ink); border-radius: 0 12px 12px 12px;
padding: 22px 24px; overflow-x: auto;
font-family: var(--mono); font-size: 13.5px; line-height: 1.7;
color: #e9e2d3;
box-shadow: 12px 12px 0 rgba(27,23,18,0.07);
}
.code .cm { color: #8a8069; }
.code .kw { color: #ff9f7a; }
.code .st { color: #cbb78a; }
.code .fn { color: #f4efe4; font-weight: 700; }
#t-rust:checked ~ .tablist label[for=t-rust],
#t-py:checked ~ .tablist label[for=t-py],
#t-node:checked ~ .tablist label[for=t-node],
#t-cli:checked ~ .tablist label[for=t-cli] {
background: var(--ink); color: var(--paper); border-color: var(--ink);
}
#t-rust:checked ~ .panels #p-rust,
#t-py:checked ~ .panels #p-py,
#t-node:checked ~ .panels #p-node,
#t-cli:checked ~ .panels #p-cli { display: block; }
.code a.ref { color: #ff9f7a; border-bottom: 1px dotted #ff9f7a; }
/* ── closing split (OSS vs hosted) ── */
.split { display: grid; grid-template-columns: 1fr 1fr; gap: 18px; }
@media (max-width: 780px) { .split { grid-template-columns: 1fr; } }
.path { border: 1px solid var(--line); border-radius: 16px; padding: 32px 30px; background: var(--card); display: flex; flex-direction: column; }
.path .ptag { font-family: var(--mono); font-size: 11px; letter-spacing: 0.12em; text-transform: uppercase; color: var(--muted); margin-bottom: 15px; }
.path h3 { font-family: var(--display); font-weight: 600; font-size: 1.5rem; letter-spacing: -0.02em; line-height: 1.05; margin-bottom: 12px; }
.path p { color: var(--ink-soft); font-size: 15px; line-height: 1.5; flex: 1; margin-bottom: 24px; }
.pbtns { display: flex; gap: 12px; flex-wrap: wrap; }
.path-pro { background: var(--ink); border-color: var(--ink); box-shadow: 14px 14px 0 rgba(27,23,18,0.09); }
.path-pro .ptag { color: #b8ad98; }
.path-pro .ptag b { color: var(--accent); font-weight: 700; }
.path-pro h3 { color: var(--paper); }
.path-pro p { color: #cfc6b3; }
.btn { font-family: var(--mono); font-size: 14px; font-weight: 500; padding: 13px 24px; border-radius: 999px; transition: all .15s; border: 1px solid var(--ink); }
.btn-p { background: var(--accent); border-color: var(--accent); color: var(--paper); }
.btn-p:hover { background: var(--accent-deep); border-color: var(--accent-deep); }
.btn-s:hover { background: var(--ink); color: var(--paper); }
.btn-pro { background: var(--accent); border-color: var(--accent); color: var(--paper); }
.btn-pro:hover { background: var(--accent-deep); border-color: var(--accent-deep); }
.btn-ghost { color: var(--paper); border-color: rgba(255,255,255,0.3); }
.btn-ghost:hover { border-color: var(--paper); background: rgba(255,255,255,0.08); }
.fc-mark { height: 34px; width: auto; display: block; margin-bottom: 20px; }
.fc-wordmark { height: 15px; width: auto; vertical-align: -2px; transition: opacity .15s; }
.fc-wordmark:hover { opacity: 0.65; }
footer { border-top: 1px solid var(--line); padding: 34px 0; font-size: 14px; color: var(--muted); }
.foot-in { display: flex; justify-content: space-between; align-items: center; gap: 18px; flex-wrap: wrap; }
.foot-in a { color: var(--ink-soft); }
.foot-in a:hover { color: var(--accent); }
.foot-links { display: flex; gap: 20px; font-family: var(--mono); font-size: 13px; }
/* ── load animation ── */
.reveal { opacity: 0; transform: translateY(14px); animation: rise .7s cubic-bezier(.2,.7,.3,1) forwards; }
@keyframes rise { to { opacity: 1; transform: none; } }
.d1 { animation-delay: .05s; } .d2 { animation-delay: .15s; } .d3 { animation-delay: .25s; }
.d4 { animation-delay: .35s; } .d5 { animation-delay: .45s; } .d6 { animation-delay: .55s; }
@media (prefers-reduced-motion: reduce) { .reveal { animation: none; opacity: 1; transform: none; } }
</style>
</head>
<body>
<nav>
<div class="wrap nav-in">
<a href="#top" class="brand"><span class="dot"></span>pdf-inspector</a>
<div class="nav-links">
<a href="#features" class="hide-sm">Features</a>
<a href="#benchmark" class="hide-sm">Benchmark</a>
<a href="#start">Quick start</a>
<a class="nav-gh" href="https://github.com/firecrawl/pdf-inspector">GitHub ↗</a>
</div>
</div>
</nav>
<header id="top">
<div class="wrap hero-grid">
<div>
<div class="eyebrow reveal d1">Rust · Python · Node · CLI</div>
<h1 class="reveal d2">Classify PDFs. Extract Markdown. <span class="strike">No OCR.</span></h1>
<p class="lede reveal d3">A fast Rust library that tells text-based PDFs from scanned ones, then extracts position-aware text and clean Markdown — <b>locally, in milliseconds</b>. Skip the OCR bill for the ~54% of PDFs that never needed it.</p>
</div>
<div class="readout reveal d4" aria-hidden="true">
<div class="rlabel"><span>classify_pdf()</span><span>~12ms</span></div>
<div class="rrow"><span class="k">type</span><span class="v hot">TextBased</span></div>
<div class="rrow"><span class="k">confidence</span><span class="v">0.98</span></div>
<div class="rrow"><span class="k">needs_ocr</span><span class="v">false</span></div>
<div class="rrow"><span class="k">route</span><span class="v">local&nbsp;&nbsp;md</span></div>
<div class="rrow" style="border-top:1px solid rgba(255,255,255,.09);padding-top:13px">
<span class="k">signal</span>
<span class="bar"><i style="width:98%"></i></span>
</div>
</div>
</div>
<div class="wrap">
<div class="installs">
<a class="inst reveal d4" href="https://crates.io/crates/pdf-inspector">
<div class="reg"><span>crates.io</span><span class="arrow"></span></div>
<code>cargo add pdf-inspector</code>
</a>
<a class="inst reveal d5" href="https://pypi.org/project/pdf-inspector/">
<div class="reg"><span>PyPI</span><span class="arrow"></span></div>
<code>pip install pdf-inspector</code>
</a>
<a class="inst reveal d6" href="https://www.npmjs.com/package/@firecrawl/pdf-inspector">
<div class="reg"><span>npm</span><span class="arrow"></span></div>
<code>npm i @firecrawl/pdf-inspector</code>
</a>
</div>
</div>
</header>
<section id="features">
<div class="wrap">
<div class="sec-head">
<span class="sec-num">01</span>
<h2 class="sec-title">Built for routing, not just reading</h2>
</div>
<div class="feat-grid">
<div class="feat"><div class="fn">01</div><h3>Smart classification</h3><p>TextBased, Scanned, ImageBased, or Mixed in ~1050ms by sampling content streams. Returns a confidence score and per-page OCR routing.</p></div>
<div class="feat"><div class="fn">02</div><h3>Position-aware text</h3><p>Extraction with font info, X/Y coordinates, and automatic multi-column reading order.</p></div>
<div class="feat"><div class="fn">03</div><h3>Markdown conversion</h3><p>Headings, bullet/numbered lists, code blocks, tables, bold/italic, URL linking, and page breaks.</p></div>
<div class="feat"><div class="fn">04</div><h3>Table detection</h3><p>Rectangle-based detection from drawing ops plus heuristic alignment detection. Financial tables, footnotes, and cross-page continuations.</p></div>
<div class="feat"><div class="fn">05</div><h3>CID font support</h3><p>ToUnicode CMap decoding for Type0/Identity-H fonts, with UTF-16BE, UTF-8, and Latin-1 encodings.</p></div>
<div class="feat"><div class="fn">06</div><h3>Multi-column layout</h3><p>Newspaper-style column detection, sequential reading order, and right-to-left text support.</p></div>
<div class="feat"><div class="fn">07</div><h3>Encoding checks</h3><p>Flags broken font encodings automatically so callers can fall back to OCR only when it's actually needed.</p></div>
<div class="feat"><div class="fn">08</div><h3>Lightweight</h3><p>Pure Rust. No ML models, no external services. A single parse shared between detection and extraction.</p></div>
</div>
</div>
</section>
<section id="benchmark">
<div class="wrap">
<div class="sec-head">
<span class="sec-num">02</span>
<h2 class="sec-title">Fastest of the direct-text engines</h2>
<p class="sec-sub">Evaluated on the <a href="https://github.com/opendataloader-project/opendataloader-bench" style="color:var(--accent-deep);border-bottom:1px solid var(--line)">opendataloader-bench</a> corpus (200 PDFs). Direct text-extraction engines only — no OCR, no ML. Higher is better.</p>
</div>
<div class="bench">
<div class="bench-wrap">
<table>
<thead>
<tr><th>Engine</th><th>Overall</th><th>Reading order</th><th>Tables</th><th>Headings</th><th>200 docs</th></tr>
</thead>
<tbody>
<tr class="us"><td>pdf-inspector</td><td>0.83</td><td>0.88</td><td>0.66</td><td>0.74</td><td>4s</td></tr>
<tr><td>opendataloader</td><td>0.84</td><td>0.91</td><td>0.49</td><td>0.74</td><td>11s</td></tr>
<tr><td>pymupdf4llm</td><td>0.73</td><td>0.89</td><td>0.40</td><td>0.41</td><td>18s</td></tr>
<tr><td>markitdown</td><td>0.58</td><td>0.88</td><td>0.00</td><td>0.00</td><td>8s</td></tr>
</tbody>
</table>
</div>
<div class="bench-foot">OCR/ML engines (docling, marker, mineru) score 0.830.88 overall — but take 2180 minutes on the same corpus.</div>
</div>
<div class="callouts">
<div class="callout"><div class="ct">Where we win</div><p>Fastest engine measured, the best table detection of any engine here, and heading quality now on par with opendataloader — at ~2.5× its speed.</p></div>
<div class="callout"><div class="ct">Where we're working</div><p>Reading order still trails opendataloader slightly, and tables that need the visual structure only an OCR engine can see.</p></div>
</div>
</div>
</section>
<section id="start">
<div class="wrap">
<div class="sec-head">
<span class="sec-num">03</span>
<h2 class="sec-title">Three lines to Markdown</h2>
</div>
<div class="tabs">
<input type="radio" name="tab" id="t-rust" checked>
<input type="radio" name="tab" id="t-py">
<input type="radio" name="tab" id="t-node">
<input type="radio" name="tab" id="t-cli">
<div class="tablist">
<label for="t-rust">Rust</label>
<label for="t-py">Python</label>
<label for="t-node">Node.js</label>
<label for="t-cli">CLI</label>
</div>
<div class="panels">
<div class="panel" id="p-rust"><pre class="code"><span class="kw">use</span> pdf_inspector::process_pdf;
<span class="kw">let</span> result = <span class="fn">process_pdf</span>(<span class="st">"document.pdf"</span>)?;
<span class="fn">println!</span>(<span class="st">"Type: {:?}"</span>, result.pdf_type);
<span class="kw">if let</span> <span class="kw">Some</span>(markdown) = &amp;result.markdown {
<span class="fn">println!</span>(<span class="st">"{}"</span>, markdown);
}
<span class="cm">// full reference → </span><a class="ref" href="https://github.com/firecrawl/pdf-inspector/blob/main/docs/rust-api.md">docs/rust-api.md</a></pre></div>
<div class="panel" id="p-py"><pre class="code"><span class="kw">import</span> pdf_inspector
result = pdf_inspector.<span class="fn">process_pdf</span>(<span class="st">"document.pdf"</span>)
<span class="fn">print</span>(result.pdf_type) <span class="cm"># "text_based" | "scanned" | "image_based" | "mixed"</span>
<span class="fn">print</span>(result.markdown) <span class="cm"># Markdown string or None</span>
<span class="cm"># full reference → </span><a class="ref" href="https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md">docs/python.md</a></pre></div>
<div class="panel" id="p-node"><pre class="code"><span class="kw">import</span> { readFileSync } <span class="kw">from</span> <span class="st">'fs'</span>;
<span class="kw">import</span> { processPdf } <span class="kw">from</span> <span class="st">'@firecrawl/pdf-inspector'</span>;
<span class="kw">const</span> result = <span class="fn">processPdf</span>(<span class="fn">readFileSync</span>(<span class="st">'document.pdf'</span>));
console.<span class="fn">log</span>(result.pdfType); <span class="cm">// "TextBased" | "Scanned" | ...</span>
console.<span class="fn">log</span>(result.markdown); <span class="cm">// Markdown string or null</span>
<span class="cm">// full reference → </span><a class="ref" href="https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md">napi/README.md</a></pre></div>
<div class="panel" id="p-cli"><pre class="code"><span class="cm"># install the CLI tools</span>
cargo <span class="fn">install</span> pdf-inspector
<span class="cm"># convert a PDF to Markdown</span>
<span class="fn">pdf2md</span> document.pdf
<span class="cm"># classify only — is it scanned?</span>
<span class="fn">detect-pdf</span> document.pdf --analyze --json
<span class="cm"># structured output for pipelines</span>
<span class="fn">pdf2md</span> document.pdf --json</pre></div>
</div>
</div>
</div>
</section>
<section class="close">
<div class="wrap">
<div class="sec-head">
<span class="sec-num">04</span>
<h2 class="sec-title">Two ways to parse</h2>
<p class="sec-sub">Run the classifier locally for text-based PDFs; hand the scanned, OCR, and at-scale work to Firecrawl.</p>
</div>
<div class="split">
<div class="path">
<div class="ptag">Open source · runs local</div>
<h3>Use pdf-inspector yourself</h3>
<p>Pure-Rust library and CLI. Classify and extract text-based PDFs on your own machine in milliseconds — no external calls, no OCR bill, MIT licensed.</p>
<div class="pbtns">
<a class="btn btn-p" href="https://github.com/firecrawl/pdf-inspector">Get started</a>
<a class="btn btn-s" href="https://crates.io/crates/pdf-inspector">crates.io</a>
</div>
</div>
<div class="path path-pro">
<img class="fc-mark" src="assets/firecrawl-mark.svg" alt="Firecrawl" width="24" height="34">
<div class="ptag"><b>Firecrawl Parse</b> · hosted API</div>
<h3>Or let Firecrawl handle the hard ones</h3>
<p>Scanned documents, OCR, DOCX / XLSX / HTML, and parsing at scale — clean, LLM-ready Markdown from one API call. The downstream route for everything local parsing can't reach.</p>
<div class="pbtns">
<a class="btn btn-pro" href="https://docs.firecrawl.dev/api-reference/endpoint/parse">Firecrawl Parse ↗</a>
<a class="btn btn-ghost" href="https://firecrawl.dev">firecrawl.dev</a>
</div>
</div>
</div>
</div>
</section>
<footer>
<div class="wrap foot-in">
<div style="display:flex;align-items:center;gap:7px">Built by <a href="https://firecrawl.dev"><img class="fc-wordmark" src="assets/firecrawl-wordmark.svg" alt="Firecrawl"></a> · MIT licensed</div>
<div class="foot-links">
<a href="https://github.com/firecrawl/pdf-inspector">GitHub</a>
<a href="https://crates.io/crates/pdf-inspector">crates.io</a>
<a href="https://pypi.org/project/pdf-inspector/">PyPI</a>
<a href="https://www.npmjs.com/package/@firecrawl/pdf-inspector">npm</a>
</div>
</div>
</footer>
</body>
</html>
+1 -15
View File
@@ -74,7 +74,7 @@ fn format_items_json(items: &[TextItem]) -> String {
_ => String::new(),
};
format!(
r#"{{"text":"{}","page":{},"x":{:.2},"y":{:.2},"width":{:.2},"height":{:.2},"font":"{}","font_size":{:.2},"is_bold":{},"is_italic":{},"is_underline":{},"is_strikeout":{},"item_type":"{}","mcid":{}{}}}"#,
r#"{{"text":"{}","page":{},"x":{:.2},"y":{:.2},"width":{:.2},"height":{:.2},"font":"{}","font_size":{:.2},"is_bold":{},"is_italic":{},"is_underline":{},"item_type":"{}","mcid":{}{}}}"#,
json_escape(&item.text),
item.page,
item.x,
@@ -86,7 +86,6 @@ fn format_items_json(items: &[TextItem]) -> String {
item.is_bold,
item.is_italic,
item.is_underline,
item.is_strikeout,
item_type_label(&item.item_type),
mcid,
link_url,
@@ -123,7 +122,6 @@ mod tests {
is_bold: false,
is_italic: true,
is_underline: true,
is_strikeout: true,
item_type: ItemType::Text,
mcid: Some(7),
}];
@@ -208,7 +206,6 @@ fn main() {
eprintln!(" --raw Output only markdown (no headers)");
eprintln!(" --pages Insert page break markers (<!-- Page N -->)");
eprintln!(" --select-pages N Only process specified pages (e.g. 1,3,5-10)");
eprintln!(" --password PW Password for an encrypted PDF");
eprintln!(" --detect-only Only detect PDF type (no extraction)");
eprintln!(" --analyze Detect + extract + layout analysis (no markdown)");
process::exit(1);
@@ -222,16 +219,6 @@ fn main() {
let detect_only = args.iter().any(|a| a == "--detect-only");
let analyze = args.iter().any(|a| a == "--analyze");
// Parse --password value
let password = args.iter().position(|a| a == "--password").map(|i| {
args.get(i + 1)
.unwrap_or_else(|| {
eprintln!("Error: --password requires a value");
process::exit(1);
})
.clone()
});
// Parse --select-pages value
let page_filter = args
.iter()
@@ -280,7 +267,6 @@ fn main() {
if let Some(pages) = page_filter {
options.page_filter = Some(pages);
}
options.password = password;
match process_pdf_with_options(pdf_path, options) {
Ok(result) => {
+20 -250
View File
@@ -14,9 +14,8 @@ use lopdf::{Document, Encoding, Object, ObjectId};
use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, descriptor_style_flags,
extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
FontStyleCache,
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
};
use super::underline::UnderlineLine;
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
@@ -107,25 +106,6 @@ fn transformed_stroke_width(
user_width * (ndx * ndx + ndy * ndy).sqrt()
}
/// Text rise (Ts) displaces the glyph origin by (0, rise) in unscaled text
/// space — per the rendering-matrix definition it sits left of Tm, so the
/// offset maps through the text matrix's y column. Rise never contributes
/// to the advance, so callers apply it only to the rendering position and
/// keep advancing the unshifted text matrix.
fn rise_adjusted(tm: &[f32; 6], rise: f32) -> [f32; 6] {
if rise == 0.0 {
return *tm;
}
[
tm[0],
tm[1],
tm[2],
tm[3],
tm[4] + rise * tm[2],
tm[5] + rise * tm[3],
]
}
/// Returns `(page_extraction, has_gid_fonts)` where `has_gid_fonts` indicates
/// the page uses fonts with unresolvable gid-encoded glyphs.
pub(crate) fn extract_page_text_items(
@@ -134,7 +114,6 @@ pub(crate) fn extract_page_text_items(
page_num: u32,
font_cmaps: &FontCMaps,
include_invisible: bool,
style_cache: &mut FontStyleCache,
) -> Result<(PageExtraction, bool, bool), PdfError> {
use lopdf::content::Content;
@@ -174,8 +153,6 @@ pub(crate) fn extract_page_text_items(
std::collections::HashMap::new();
let mut inline_cmaps: std::collections::HashMap<String, crate::tounicode::CMapEntry> =
std::collections::HashMap::new();
let mut font_style_flags: std::collections::HashMap<String, (bool, bool)> =
std::collections::HashMap::new();
for (font_name, font_dict) in &fonts {
let resource_name = String::from_utf8_lossy(font_name).to_string();
if let Ok(base_font) = font_dict.get(b"BaseFont") {
@@ -184,12 +161,6 @@ pub(crate) fn extract_page_text_items(
font_base_names.insert(resource_name.clone(), base_name);
}
}
// Descriptor style flags rescue subset fonts whose BaseFont names
// are opaque tags the name heuristics can't read.
let style = descriptor_style_flags(doc, font_dict, style_cache);
if style != (false, false) {
font_style_flags.insert(resource_name.clone(), style);
}
// Track ToUnicode object reference, with FontFile2 fallback for Identity-H/V.
// Also handle inline ToUnicode streams.
match font_dict.get(b"ToUnicode") {
@@ -264,7 +235,6 @@ pub(crate) fn extract_page_text_items(
line_width: f32,
char_spacing: f32,
word_spacing: f32,
text_rise: f32,
text_leading: f32,
current_font: String,
current_font_size: f32,
@@ -277,7 +247,6 @@ pub(crate) fn extract_page_text_items(
let mut text_leading: f32 = 0.0; // TL parameter (in text-space units)
let mut char_spacing: f32 = 0.0; // Tc parameter (extra spacing per character, unscaled)
let mut word_spacing: f32 = 0.0; // Tw parameter (extra spacing per space char, unscaled)
let mut text_rise: f32 = 0.0; // Ts parameter (baseline shift for super/subscripts, unscaled)
let mut text_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut line_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut in_text_block = false;
@@ -299,10 +268,6 @@ pub(crate) fn extract_page_text_items(
let mut suppress_glyph_extraction = false;
let mut actual_text_start_tm: Option<[f32; 6]> = None; // text matrix at BDC entry
let mut actual_text_glyph_tm: Option<[f32; 6]> = None; // text matrix at first glyph inside BDC
// Text rise in effect at each captured matrix — the item must render at
// the rise of its GLYPHS, not whatever rise is set by EMC time.
let mut actual_text_start_rise: f32 = 0.0;
let mut actual_text_glyph_rise: Option<f32> = None;
/// Get the innermost MCID from the marked content stack.
fn current_mcid(stack: &[MarkedContentEntry]) -> Option<i64> {
stack.iter().rev().find_map(|e| e.mcid)
@@ -319,7 +284,6 @@ pub(crate) fn extract_page_text_items(
line_width,
char_spacing,
word_spacing,
text_rise,
text_leading,
current_font: current_font.clone(),
current_font_size,
@@ -333,7 +297,6 @@ pub(crate) fn extract_page_text_items(
line_width = saved.line_width;
char_spacing = saved.char_spacing;
word_spacing = saved.word_spacing;
text_rise = saved.text_rise;
text_leading = saved.text_leading;
current_font = saved.current_font;
current_font_size = saved.current_font_size;
@@ -406,12 +369,6 @@ pub(crate) fn extract_page_text_items(
word_spacing = tw;
}
}
"Ts" => {
// Set text rise (baseline shift for superscripts/subscripts)
if let Some(ts) = op.operands.first().and_then(get_number) {
text_rise = ts;
}
}
"Td" | "TD" => {
// Move text position: TLM = T(tx,ty) × TLM; Tm = TLM
// tx,ty are in text space — must be scaled by the text line matrix
@@ -470,7 +427,6 @@ pub(crate) fn extract_page_text_items(
if suppress_glyph_extraction {
if actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
@@ -500,8 +456,7 @@ pub(crate) fn extract_page_text_items(
&mut cmap_decisions,
&font_widths,
) {
let combined =
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
let combined = multiply_matrices(&text_matrix, &ctm);
let rendered_size = effective_font_size(current_font_size, &combined);
let (x, y) = (combined[4], combined[5]);
if combined[0].abs() >= combined[1].abs() {
@@ -523,10 +478,6 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
@@ -536,10 +487,9 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -557,7 +507,6 @@ pub(crate) fn extract_page_text_items(
// Capture first-glyph position for ActualText
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
// Compute space threshold based on font metrics when available
@@ -678,10 +627,6 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
let scale_x = text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2];
for (text, start_w, end_w) in &sub_items {
let offset_tm = [
@@ -692,8 +637,7 @@ pub(crate) fn extract_page_text_items(
text_matrix[4] + start_w * text_matrix[0],
text_matrix[5] + start_w * text_matrix[1],
];
let combined =
multiply_matrices(&rise_adjusted(&offset_tm, text_rise), &ctm);
let combined = multiply_matrices(&offset_tm, &ctm);
let (x, y) = (combined[4], combined[5]);
let width = if font_info.is_some() {
((end_w - start_w) * scale_x).abs()
@@ -709,10 +653,9 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -736,26 +679,6 @@ pub(crate) fn extract_page_text_items(
line_matrix[4] += (-tl) * line_matrix[2];
line_matrix[5] += (-tl) * line_matrix[3];
text_matrix = line_matrix;
// Capture first-glyph position for ActualText AFTER the
// line move — the BDC-entry matrix is on the previous line.
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
actual_text_glyph_tm = Some(text_matrix);
actual_text_glyph_rise = Some(text_rise);
}
// Advance width, as for Tj — without it the item stays
// zero-width and geometric underline/strikeout detection
// rejects it (`is_underline_candidate` needs width > 0).
let w_ts_opt = font_widths.get(&current_font).and_then(|fi| {
op.operands.first().and_then(get_operand_bytes).map(|raw| {
compute_string_width_ts(
raw,
fi,
current_font_size,
char_spacing,
word_spacing,
)
})
});
if !((text_rendering_mode == 3 && !include_invisible)
|| suppress_glyph_extraction
|| op.operands.is_empty())
@@ -773,8 +696,7 @@ pub(crate) fn extract_page_text_items(
&font_widths,
) {
if !text.trim().is_empty() {
let combined =
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
let combined = multiply_matrices(&text_matrix, &ctm);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
} else {
@@ -782,45 +704,28 @@ pub(crate) fn extract_page_text_items(
}
let rendered_size = effective_font_size(current_font_size, &combined);
let (x, y) = (combined[4], combined[5]);
let width = w_ts_opt
.map(|w_ts| {
(w_ts * (text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2]))
.abs()
})
.unwrap_or(0.0);
let base_font = font_base_names
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
y,
width,
width: 0.0,
height: rendered_size,
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
}
}
}
// Advance regardless of visibility so later show-text
// operators on the same line stay positioned (as for Tj).
if let Some(w_ts) = w_ts_opt {
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
}
}
"Do" => {
// XObject invocation - could be an image or form
@@ -852,7 +757,6 @@ pub(crate) fn extract_page_text_items(
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Image,
mcid: current_mcid(&marked_content_stack),
});
@@ -866,7 +770,6 @@ pub(crate) fn extract_page_text_items(
font_cmaps,
&ctm,
&mut cmap_decisions,
style_cache,
);
items.extend(form_items);
}
@@ -907,9 +810,7 @@ pub(crate) fn extract_page_text_items(
if actual_text.is_some() {
suppress_glyph_extraction = true;
actual_text_start_tm = Some(text_matrix);
actual_text_start_rise = text_rise;
actual_text_glyph_tm = None; // reset — will be captured at first Tj/TJ
actual_text_glyph_rise = None;
}
marked_content_stack.push(MarkedContentEntry { actual_text, mcid });
}
@@ -922,11 +823,9 @@ pub(crate) fn extract_page_text_items(
// Tj may have moved the text position to the correct line —
// the BDC-entry position can be on the previous line.
let glyph_tm = actual_text_glyph_tm.take();
let glyph_rise = actual_text_glyph_rise.take();
let entry_tm = actual_text_start_tm.take();
if let Some(start_tm) = glyph_tm.or(entry_tm) {
let rise = glyph_rise.unwrap_or(actual_text_start_rise);
let combined = multiply_matrices(&rise_adjusted(&start_tm, rise), &ctm);
let combined = multiply_matrices(&start_tm, &ctm);
if combined[0].abs() >= combined[1].abs() {
rotation_votes.horizontal += 1;
} else {
@@ -943,10 +842,6 @@ pub(crate) fn extract_page_text_items(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&at),
x,
@@ -956,10 +851,9 @@ pub(crate) fn extract_page_text_items(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: entry
.mcid
@@ -1482,15 +1376,8 @@ mod tests {
let (doc, page_id) = simple_doc_with_content(content);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
let ((items, _, _), _, _) =
extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
items
}
@@ -1575,108 +1462,6 @@ BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
assert!(!world.is_underline);
}
#[test]
fn quote_operator_text_carries_advance_width() {
// `'` (move-to-next-line-and-show-text) must retain the string's
// advance width like Tj — zero-width items are invisible to
// geometric underline/strikeout detection.
let content = b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (first) Tj (struck) ' ET
1 w
99 503 m 145 503 l S";
let items = extract_simple_items(content);
let struck = items.iter().find(|item| item.text == "struck").unwrap();
// 6 glyphs x 600/1000 x 12pt = 43.2pt, drawn one leading below Tm.
assert!((struck.width - 43.2).abs() < 0.1);
assert!((struck.y - 500.0).abs() < 0.1);
assert!(struck.is_strikeout);
assert!(!struck.is_underline);
}
#[test]
fn quote_operator_advances_text_matrix() {
// Text shown after `'` on the same line must start past the shown
// string: "CD" lands at x=114.4 (2 glyphs x 600/1000 x 12pt after
// x=100), flush against "AB", so the merge pass joins them. Without
// the advance "CD" overlaps "AB" at x=100 and the items stay apart.
let content = b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (AB) ' (CD) Tj ET";
let items = extract_simple_items(content);
let merged = items.iter().find(|item| item.text == "ABCD").unwrap();
assert!((merged.x - 100.0).abs() < 0.1);
assert!((merged.width - 28.8).abs() < 0.1);
assert!((merged.y - 500.0).abs() < 0.1);
}
#[test]
fn text_rise_shifts_item_baseline() {
// Ts displaces the glyph origin vertically without touching the
// advance; the next run at rise 0 must return to the original
// baseline and follow the raised run horizontally.
let content =
b"BT /F1 12 Tf 1 0 0 1 100 500 Tm (base) Tj 5 Ts (super) Tj 0 Ts (after) Tj ET";
let items = extract_simple_items(content);
let base = items.iter().find(|item| item.text == "base").unwrap();
let raised = items.iter().find(|item| item.text == "super").unwrap();
let after = items.iter().find(|item| item.text == "after").unwrap();
assert!((base.y - 500.0).abs() < 0.1);
assert!((raised.y - 505.0).abs() < 0.1);
assert!((after.y - 500.0).abs() < 0.1);
assert!(after.x > raised.x);
}
#[test]
fn actual_text_item_uses_glyph_rise() {
// The ActualText replacement item must render at the rise in
// effect when its glyphs were drawn — not the unshifted BDC
// baseline, and not whatever rise is set by EMC time.
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm \
/Span <</ActualText (super) >> BDC 5 Ts (sup) Tj 0 Ts EMC (after) Tj ET";
let items = extract_simple_items(content);
let sup = items.iter().find(|item| item.text == "super").unwrap();
let after = items.iter().find(|item| item.text == "after").unwrap();
assert!((sup.y - 505.0).abs() < 0.1);
assert!((after.y - 500.0).abs() < 0.1);
}
#[test]
fn actual_text_shown_with_quote_op_uses_moved_risen_baseline() {
// When the tagged span's show op is `'`, the glyph position is
// only known AFTER its line move — falling back to the BDC-entry
// matrix would place the item on the previous line, unrisen.
let content = b"BT /F1 12 Tf 14 TL 1 0 0 1 100 500 Tm \
/Span <</ActualText (replaced) >> BDC 3 Ts (raw) ' 0 Ts EMC ET";
let items = extract_simple_items(content);
let item = items.iter().find(|item| item.text == "replaced").unwrap();
// Line move: 500 - 14 = 486; rise: +3 -> 489.
assert!((item.y - 489.0).abs() < 0.1);
assert!(item.width > 0.0);
}
#[test]
fn strikeout_detected_on_risen_text() {
// The rule crosses the glyphs at their risen position; without the
// rise in item.y the strike window sits 4pt too low and misses.
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm 4 Ts (struck) Tj ET
1 w
99 507 m 145 507 l S";
let items = extract_simple_items(content);
let struck = items.iter().find(|item| item.text == "struck").unwrap();
assert!((struck.y - 504.0).abs() < 0.1);
assert!(struck.is_strikeout);
assert!(!struck.is_underline);
}
#[test]
fn test_skip_excessive_operations() {
use crate::tounicode::FontCMaps;
@@ -1711,15 +1496,7 @@ BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
doc.add_object(catalog);
let font_cmaps = FontCMaps::from_doc(&doc);
let result = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
let result = extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let ((items, rects, lines), _has_gid, _coords_rotated) = result;
assert!(items.is_empty());
assert!(rects.is_empty());
@@ -1801,15 +1578,8 @@ BT 30 700 Tm <41> Tj ET";
doc.trailer.set("Root", Object::Reference(catalog_id));
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
)
.unwrap();
let ((items, _, _), _, _) =
extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let text = items
.iter()
.map(|item| item.text.as_str())
+1 -332
View File
@@ -4,7 +4,7 @@ use crate::glyph_names::glyph_to_char;
use crate::tounicode::FontCMaps;
use crate::types::{FontEncodingMap, FontWidthInfo, PageFontEncodings, PageFontWidths};
use log::debug;
use lopdf::{Document, Encoding, Object, ObjectId};
use lopdf::{Document, Encoding, Object};
use std::collections::HashMap;
#[derive(Debug, Copy, Clone, PartialEq, Eq)]
@@ -739,171 +739,6 @@ pub(crate) fn get_font_file2_obj_num(doc: &Document, font_dict: &lopdf::Dictiona
.map(|r| r.0)
}
/// Document-scoped memo of embedded-font style flags, keyed by the
/// FontFile2/FontFile3 stream's object id. The same font program is
/// referenced from every page that uses the font, and decompressing +
/// parsing it dominates `descriptor_style_flags` — without the memo that
/// cost repeats per page whenever the descriptor leaves a flag unset
/// (the common case: regular fonts report neither italic nor bold).
#[derive(Debug, Default)]
pub(crate) struct FontStyleCache {
by_font_file: HashMap<ObjectId, (bool, bool)>,
}
impl FontStyleCache {
pub(crate) fn new() -> Self {
Self::default()
}
}
/// Style flags from the FontDescriptor, which survive subset fonts whose
/// BaseFont names are opaque tags ("Tc1", "ABCDEF+F1") that defeat the
/// name-based bold/italic heuristics.
///
/// Italic: `ItalicAngle` beyond a few degrees, or Flags bit 7 (Italic,
/// value 64). Bold: Flags bit 19 (ForceBold, value 1<<18). The small
/// ItalicAngle threshold skips fonts that declare a token slant.
pub(crate) fn descriptor_style_flags(
doc: &Document,
font_dict: &lopdf::Dictionary,
style_cache: &mut FontStyleCache,
) -> (bool, bool) {
let descriptor = font_dict
.get(b"FontDescriptor")
.ok()
.and_then(|obj| resolve_dict(doc, obj))
.or_else(|| {
// Type0 fonts hang the descriptor off DescendantFonts[0].
let desc_fonts = font_dict.get(b"DescendantFonts").ok()?;
let desc_fonts = resolve_array(doc, desc_fonts)?;
let cid_font_dict = resolve_dict(doc, desc_fonts.first()?)?;
resolve_dict(doc, cid_font_dict.get(b"FontDescriptor").ok()?)
});
let Some(descriptor) = descriptor else {
return (false, false);
};
let italic_angle = descriptor
.get(b"ItalicAngle")
.ok()
.and_then(|obj| match obj {
Object::Integer(i) => Some(*i as f32),
Object::Real(r) => Some(*r),
_ => None,
})
.unwrap_or(0.0);
let flags = descriptor
.get(b"Flags")
.ok()
.and_then(|obj| obj.as_i64().ok())
.unwrap_or(0);
let mut italic = italic_angle.abs() >= 4.0 || flags & (1 << 6) != 0;
let mut bold = flags & (1 << 18) != 0;
// Descriptors lie: subset generators write ItalicAngle 0 for genuinely
// italic faces. The embedded font file keeps the truth — OS/2
// fsSelection (via `Face::is_italic`) and the post table's italicAngle.
if !italic || !bold {
if let Some(ff_ref) = font_file_ref(descriptor) {
let (emb_italic, emb_bold) = *style_cache
.by_font_file
.entry(ff_ref)
.or_insert_with(|| embedded_style_flags(doc, ff_ref));
italic = italic || emb_italic;
bold = bold || emb_bold;
}
}
(italic, bold)
}
/// Style flags parsed from an embedded font program stream.
fn embedded_style_flags(doc: &Document, ff_ref: ObjectId) -> (bool, bool) {
let Some(data) = font_file_data(doc, ff_ref) else {
return (false, false);
};
if let Ok(face) = ttf_parser::Face::parse(&data, 0) {
(
face.is_italic() || face.italic_angle().abs() >= 4.0,
face.is_bold(),
)
} else if let Some(name) = cff_font_name(&data) {
// FontFile3 is bare CFF (no sfnt container) — ttf_parser
// can't open it, but the CFF Name INDEX keeps the real
// PostScript name ("XXXXXX+Amplitude-LightItalic") even
// when the descriptor was rewritten to claim upright.
(
crate::text_utils::is_italic_font(&name),
crate::text_utils::is_bold_font(&name),
)
} else {
(false, false)
}
}
/// First PostScript name from a bare CFF font's Name INDEX (CFF spec §7).
fn cff_font_name(data: &[u8]) -> Option<String> {
// Header: major(1) minor(1) hdrSize(1) offSize(1); major must be 1.
if data.len() < 4 || data[0] != 1 {
return None;
}
let hdr_size = data[2] as usize;
// Name INDEX: count(u16) offSize(u8) offsets[count+1] data
let count = u16::from_be_bytes([*data.get(hdr_size)?, *data.get(hdr_size + 1)?]) as usize;
if count == 0 {
return None;
}
let off_size = *data.get(hdr_size + 2)? as usize;
if !(1..=4).contains(&off_size) {
return None;
}
let read_offset = |idx: usize| -> Option<usize> {
let at = hdr_size + 3 + idx * off_size;
let bytes = data.get(at..at + off_size)?;
let mut v = 0usize;
for b in bytes {
v = (v << 8) | *b as usize;
}
Some(v)
};
let start = read_offset(0)?;
let end = read_offset(1)?;
if start == 0 || end < start {
return None;
}
// Offsets are 1-based from the byte before the object data.
let objects_base = hdr_size + 3 + (count + 1) * off_size - 1;
let name = data.get(objects_base + start..objects_base + end)?;
Some(String::from_utf8_lossy(name).to_string())
}
/// FontFile2/FontFile3 stream reference from a FontDescriptor.
fn font_file_ref(descriptor: &lopdf::Dictionary) -> Option<ObjectId> {
descriptor
.get(b"FontFile2")
.ok()
.and_then(|o| o.as_reference().ok())
.or_else(|| {
descriptor
.get(b"FontFile3")
.ok()
.and_then(|o| o.as_reference().ok())
})
}
/// Decompressed embedded font program bytes.
fn font_file_data(doc: &Document, ff_ref: ObjectId) -> Option<Vec<u8>> {
let stream = doc
.get_object(ff_ref)
.and_then(lopdf::Object::as_stream)
.ok()?;
Some(
stream
.decompressed_content()
.unwrap_or_else(|_| stream.content.clone()),
)
}
/// Decode text from a PDF string operand using font CMaps, encodings, and fallbacks.
#[allow(clippy::too_many_arguments)]
pub(crate) fn extract_text_from_operand(
@@ -1427,172 +1262,6 @@ mod tests {
}
}
fn doc_with_descriptor(descriptor: lopdf::Dictionary) -> (Document, lopdf::Dictionary) {
let mut doc = Document::with_version("1.4");
let desc_id = doc.add_object(descriptor);
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "TrueType",
"BaseFont" => "Tc1",
"FontDescriptor" => desc_id,
};
(doc, font_dict)
}
#[test]
fn descriptor_italic_angle_sets_italic() {
// Subset font with an opaque BaseFont name ("Tc1") — the name
// heuristic sees nothing, the descriptor carries the truth.
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => -12,
"Flags" => 32,
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(true, false)
);
}
#[test]
fn descriptor_italic_flag_bit_sets_italic() {
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => 0,
"Flags" => 64, // bit 7: Italic
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(true, false)
);
}
#[test]
fn descriptor_force_bold_flag_sets_bold() {
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => 0,
"Flags" => 1 << 18, // ForceBold
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, true)
);
}
#[test]
fn tiny_italic_angle_is_not_italic() {
// A token 1-degree slant is optical correction, not italic.
let (doc, font_dict) = doc_with_descriptor(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "Tc1",
"ItalicAngle" => lopdf::Object::Real(-1.0),
"Flags" => 32,
});
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, false)
);
}
#[test]
fn missing_descriptor_yields_no_flags() {
let doc = Document::with_version("1.4");
let font_dict = dictionary! { "Type" => "Font", "BaseFont" => "Tc1" };
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, false)
);
}
#[test]
fn type0_descendant_descriptor_is_resolved() {
let mut doc = Document::with_version("1.4");
let desc_id = doc.add_object(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "ABCDEF+F1",
"ItalicAngle" => -15,
});
let cid_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "CIDFontType2",
"FontDescriptor" => desc_id,
});
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "Type0",
"BaseFont" => "ABCDEF+F1",
"DescendantFonts" => vec![lopdf::Object::Reference(cid_id)],
};
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(true, false)
);
}
/// Bare CFF: header + Name INDEX only — enough for `cff_font_name`.
fn bare_cff_with_name(name: &str) -> Vec<u8> {
let mut data = vec![1, 0, 4, 1]; // major, minor, hdrSize, offSize
data.extend_from_slice(&1u16.to_be_bytes()); // Name INDEX count
data.push(1); // offSize
data.push(1); // offset of first name
data.push(1 + name.len() as u8); // offset past last name
data.extend_from_slice(name.as_bytes());
data
}
#[test]
fn embedded_font_style_is_cached_by_font_file_object() {
use lopdf::{Object, Stream};
let mut doc = Document::with_version("1.4");
let ff_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
bare_cff_with_name("ABCDEF+Test-BoldItalic"),
)));
let desc_id = doc.add_object(dictionary! {
"Type" => "FontDescriptor",
"FontName" => "ABCDEF+Test-BoldItalic",
"ItalicAngle" => 0,
"Flags" => 32,
"FontFile3" => ff_id,
});
let font_dict = dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Tc1",
"FontDescriptor" => desc_id,
};
let mut cache = FontStyleCache::new();
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut cache),
(true, true)
);
assert_eq!(cache.by_font_file.len(), 1);
// Replace the font program with garbage: a repeat call must serve
// the memo instead of re-reading the stream — repeated per-page
// decompression is exactly what the cache exists to avoid.
doc.objects.insert(
ff_id,
Object::Stream(Stream::new(dictionary! {}, vec![0u8; 4])),
);
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut cache),
(true, true)
);
// A cold cache parses the (now garbage) stream, proving the warm
// call above answered from the memo.
assert_eq!(
descriptor_style_flags(&doc, &font_dict, &mut FontStyleCache::new()),
(false, false)
);
}
#[test]
fn compute_string_width_ts_no_tc_tw() {
// Without Tc/Tw (both 0), width = glyph widths only
-2
View File
@@ -1495,7 +1495,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1626,7 +1625,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
-2
View File
@@ -79,7 +79,6 @@ pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> V
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Link(url),
mcid: None,
});
@@ -319,7 +318,6 @@ pub(crate) fn walk_form_fields(
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::FormField,
mcid: None,
});
+6 -316
View File
@@ -9,7 +9,7 @@ mod links;
pub(crate) mod underline;
mod xobjects;
use crate::text_utils::{is_cjk_char, is_rtl_text};
use crate::text_utils::is_rtl_text;
use crate::tounicode::FontCMaps;
use crate::types::{PageExtraction, PdfLine, PdfRect, TextItem};
use crate::PdfError;
@@ -24,7 +24,6 @@ use links::{extract_form_fields, extract_page_links};
// Re-export public types so existing `crate::extractor::X` paths keep working.
pub use crate::text_utils::{is_bold_font, is_italic_font};
pub use crate::types::{ItemType, TextLine};
pub(crate) use fonts::FontStyleCache;
pub(crate) use layout::detect_columns;
pub use layout::group_into_lines;
pub(crate) use layout::group_into_lines_with_thresholds;
@@ -162,9 +161,6 @@ fn extract_positioned_text_impl(
let mut all_lines = Vec::new();
let mut page_thresholds: PageThresholds = HashMap::new();
let mut gid_encoded_pages: HashSet<u32> = HashSet::new();
// Embedded-font style flags are document-scoped: the same font program
// is shared across pages, so parse it once, not once per page.
let mut style_cache = FontStyleCache::new();
// Build page ObjectId → page number map for form field extraction
let page_id_to_num: HashMap<ObjectId, u32> =
@@ -176,14 +172,8 @@ fn extract_positioned_text_impl(
continue;
}
}
let ((mut items, rects, lines), has_gid_fonts, _coords_rotated) = extract_page_text_items(
doc,
page_id,
*page_num,
font_cmaps,
include_invisible,
&mut style_cache,
)?;
let ((mut items, rects, lines), has_gid_fonts, _coords_rotated) =
extract_page_text_items(doc, page_id, *page_num, font_cmaps, include_invisible)?;
if has_gid_fonts {
gid_encoded_pages.insert(*page_num);
}
@@ -244,10 +234,7 @@ fn suppress_table_underlines(
lines: &[PdfLine],
page: u32,
) {
if !items
.iter()
.any(|item| item.is_underline || item.is_strikeout)
{
if !items.iter().any(|item| item.is_underline) {
return;
}
@@ -269,7 +256,6 @@ fn suppress_table_underlines(
for index in table_item_indices {
if let Some(item) = items.get_mut(index) {
item.is_underline = false;
item.is_strikeout = false;
}
}
}
@@ -527,136 +513,6 @@ fn should_preserve_overlapping_stream_order(group: &[&TextItem]) -> bool {
saw_backtrack
}
/// Detect a tracked (letter-spaced) run of single-glyph items and derive its
/// run-local space floor.
///
/// Display type set with tracking renders one glyph per show op; the merge
/// loop's fixed thresholds (0.08-0.13 em) then read every letter gap as a
/// word boundary and emit "H O W" instead of "HOW". Within such a run the
/// gaps carry the real signal: letter gaps cluster tightly just above the
/// fixed threshold, word gaps sit clearly higher. Returns (run_end_index,
/// space_floor) when the run starting at `start` is tracked — spaces are
/// then inserted only at gaps above the floor (infinity = single word).
/// Normal text (multi-char items, or single-char runs with sub-threshold
/// gaps) returns None and keeps the existing behavior.
/// Han/Kana scripts write without inter-word spaces. Hangul (Korean) DOES
/// space between words and deliberately stays out of this set — a Korean
/// tracked run keeps normal word-boundary handling.
fn is_spaceless_cjk(c: char) -> bool {
matches!(c,
'\u{3000}'..='\u{303F}' // CJK Symbols and Punctuation
| '\u{3040}'..='\u{309F}' // Hiragana
| '\u{30A0}'..='\u{30FF}' // Katakana
| '\u{4E00}'..='\u{9FFF}' // CJK Unified Ideographs
| '\u{F900}'..='\u{FAFF}' // CJK Compatibility Ideographs
| '\u{FF00}'..='\u{FFEF}' // Halfwidth and Fullwidth Forms
)
}
fn tracked_run_space_floor(group: &[&TextItem], start: usize) -> Option<(usize, f32)> {
const MIN_GAPS: usize = 4;
let first = group[start];
if first.text.trim().chars().count() != 1 {
return None;
}
let fs = first.font_size;
if fs <= 0.0 {
return None;
}
// Walk the run under the SAME break conditions as the merge loop
// (size band, style equality, mergeable gap) so indices stay aligned.
let mut gaps: Vec<f32> = Vec::new();
let mut end_x = first.x + effective_merge_width(first);
let mut end = start;
for (offset, next) in group[start + 1..].iter().enumerate() {
if next.text.trim().chars().count() != 1 {
break;
}
if (next.font_size - fs).abs() > fs * 0.20 {
break;
}
if next.is_bold != first.is_bold
|| next.is_italic != first.is_italic
|| next.is_underline != first.is_underline
|| next.is_strikeout != first.is_strikeout
{
break;
}
let gap = next.x - end_x;
if gap > fs * 0.5 || gap < -fs * 0.5 {
break;
}
gaps.push(gap / fs);
end_x = next.x + effective_merge_width(next);
end = start + 1 + offset;
}
if gaps.len() < 2 {
return None;
}
// Tracked signature: the run's TYPICAL gap clears the fixed space
// threshold (0.08) — the merge loop would break almost every letter
// pair into "words". Short runs (2-3 gaps: "H O W") demand a stricter
// shape — clearly wide, uniform, ALL-CAPS — because a genuine spaced
// sequence of single letters ("x y z" variables) has the same gap
// count; display tracking is a caps convention.
let mut sorted = gaps.clone();
sorted.sort_by(|a, b| a.total_cmp(b));
let median = sorted[sorted.len() / 2];
// Typographic convention gate, both tiers: display tracking is an
// all-caps convention, and Han/Kana never space between glyphs. Mixed-
// or lowercase Latin runs keep their boundaries because geometry alone
// cannot distinguish spaced singles ("A b c d e") from a tracked
// title-case word ("B u f f a l o").
let run_chars = || {
group[start..=end]
.iter()
.flat_map(|it| it.text.trim().chars())
};
let spaceless_cjk = run_chars().all(|c| is_spaceless_cjk(c) || !c.is_alphanumeric())
&& run_chars().any(is_spaceless_cjk);
let all_caps = run_chars().all(|c| c.is_uppercase() || is_cjk_char(c) || !c.is_alphabetic());
if !(spaceless_cjk || all_caps) {
return None;
}
if gaps.len() >= MIN_GAPS {
if median <= 0.075 {
return None;
}
} else {
let uniform = sorted[sorted.len() - 1] <= sorted[0].max(0.01) * 1.4;
if median < 0.09 || !uniform {
return None;
}
}
// Han/Kana: no inter-glyph spaces, period — a nonuniform gap
// distribution (punctuation spacing, justification) must not
// manufacture word boundaries.
if spaceless_cjk {
return Some((end, f32::INFINITY));
}
// Word gaps, if present, form a second mode above the letter-gap
// cluster: split at the largest relative jump. Unimodal → one word.
let mut best_jump = 1.0f32;
let mut floor = f32::INFINITY;
for pair in sorted.windows(2) {
let (lo, hi) = (pair[0].max(0.01), pair[1].max(0.01));
let jump = hi / lo;
if jump > best_jump {
best_jump = jump;
floor = (lo + hi) / 2.0;
}
}
if best_jump < 1.4 {
floor = f32::INFINITY;
}
Some((end, floor * fs))
}
pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
if items.is_empty() {
return items;
@@ -704,14 +560,6 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
let mut text = first.text.clone();
let mut end_x = first.x + effective_merge_width(first);
// Tracked display text: run-local space floor overrides the
// fixed thresholds for this run's junctions (see helper).
let tracked = if *preserve_stream_order {
None
} else {
tracked_run_space_floor(group, i)
};
let mut j = i + 1;
while j < group.len() {
let next = group[j];
@@ -728,7 +576,6 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
if next.is_bold != first.is_bold
|| next.is_italic != first.is_italic
|| next.is_underline != first.is_underline
|| next.is_strikeout != first.is_strikeout
{
break;
}
@@ -766,11 +613,7 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
let needs_bullet_space = *preserve_stream_order
&& is_standalone_bullet_text(&text)
&& !next.text.trim().is_empty();
let effective_threshold = match tracked {
Some((run_end, floor)) if j <= run_end => floor,
_ => threshold,
};
if needs_bullet_space || gap > effective_threshold {
if needs_bullet_space || gap > threshold {
text.push(' ');
}
text.push_str(&next.text);
@@ -795,7 +638,6 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
is_bold: first.is_bold,
is_italic: first.is_italic,
is_underline: first.is_underline,
is_strikeout: first.is_strikeout,
item_type: first.item_type.clone(),
mcid: first.mcid,
});
@@ -873,9 +715,7 @@ pub(crate) fn merge_subscript_items(items: Vec<TextItem>) -> Vec<TextItem> {
.chars()
.last()
.is_some_and(|c| c.is_alphabetic());
let same_marks = parent.is_underline == item.is_underline
&& parent.is_strikeout == item.is_strikeout;
if parent.font_size >= sub_threshold && ends_with_letter && same_marks {
if parent.font_size >= sub_threshold && ends_with_letter {
let parent_right = parent.x + parent.width;
let gap = item.x - parent_right;
// Subscripts must be tightly adjacent (within ~1pt)
@@ -936,107 +776,6 @@ mod tests {
use crate::types::{ItemType, PdfLine, TextLine};
use layout::{detect_columns, is_newspaper_layout, ColumnRegion};
/// Glyph-per-item run at `fs`=12 with the given inter-glyph gap (pt).
fn glyph_run(chars: &str, start_x: f32, glyph_w: f32, gap: f32) -> Vec<TextItem> {
let mut x = start_x;
let mut out = Vec::new();
for c in chars.chars() {
out.push(make_merge_item(&c.to_string(), x, glyph_w));
x += glyph_w + gap;
}
out
}
#[test]
fn tracked_caps_run_collapses_to_word() {
// Display tracking: every letter gap (0.19 em) clears the fixed
// space threshold — without the run-local floor this reads "H O W".
let items = glyph_run("HOW", 100.0, 10.0, 2.3);
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "HOW");
}
#[test]
fn tracked_run_keeps_word_gaps_bimodal() {
// Letters at 0.19 em, word gaps at 0.42 em (below the 0.5 em item
// break): the split must land between the modes. Needs >=4 gaps to
// enter the bimodal tier — short runs use the strict uniform gate.
let mut items = glyph_run("ITISOK", 100.0, 8.0, 2.3);
for i in 2..6 {
items[i].x += 2.8; // word gap at T|I
}
for i in 4..6 {
items[i].x += 2.8; // word gap at S|O
}
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "IT IS OK");
}
#[test]
fn lowercase_spaced_singles_stay_words() {
// "x y z" variables: same gap shape but lowercase — the short-run
// caps requirement keeps genuine spaced singles apart.
let items = glyph_run("xyz", 100.0, 6.0, 2.3);
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "x y z");
}
#[test]
fn kerned_singles_unaffected() {
// Tiny kerning gaps never triggered spaces before and still don't.
let items = glyph_run("WORD", 100.0, 8.0, 0.3);
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "WORD");
}
#[test]
fn long_lowercase_spaced_singles_keep_boundaries() {
// Review: a 5+ single-letter lowercase list has the tracked gap
// shape at any length — the convention gate must protect it in
// the >=4-gap tier too.
let items = glyph_run("abcde", 100.0, 6.0, 2.3);
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "a b c d e");
}
#[test]
fn han_run_with_nonuniform_gaps_never_gains_spaces() {
// Review: a bimodal gap distribution (justification, punctuation
// spacing) must not manufacture word boundaries in Han text.
let mut items = glyph_run("北京时事快报", 100.0, 12.0, 1.4);
for item in items.iter_mut().skip(3) {
item.x += 3.0; // wide gap after the third glyph
}
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "北京时事快报");
}
#[test]
fn uppercase_leading_spaced_singles_keep_boundaries() {
// "A b c d e" is indistinguishable from a title-case tracked word
// without reliable tracking metadata, so preserve its boundaries.
let items = glyph_run("Abcde", 100.0, 7.0, 2.3);
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "A b c d e");
}
#[test]
fn cjk_glyph_run_collapses_without_spaces() {
// CJK sets one glyph per item with loose gaps; CJK uses no spaces,
// and the non-alphabetic run passes the caps gate.
let items = glyph_run("北京时事", 100.0, 12.0, 1.4);
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "北京时事");
}
fn make_merge_item(text: &str, x: f32, width: f32) -> TextItem {
TextItem {
text: text.into(),
@@ -1050,7 +789,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1253,7 +991,6 @@ mod tests {
items[3].y = 470.0;
for item in &mut items {
item.is_underline = true;
item.is_strikeout = true;
}
let lines = vec![
make_line(100.0, 500.0, 300.0, 500.0),
@@ -1267,30 +1004,6 @@ mod tests {
suppress_table_underlines(&mut items, &[], &lines, 1);
assert!(items.iter().all(|item| !item.is_underline));
assert!(items.iter().all(|item| !item.is_strikeout));
}
#[test]
fn subscript_digit_with_different_marks_is_not_absorbed() {
// A struck-out word followed by an unmarked footnote digit: merging
// would widen the parent's strikeout claim over the digit (and the
// reverse would drop the digit's own mark). Style boundaries break
// the merge, as in merge_text_items.
let mut word = make_merge_item("word", 100.0, 24.0);
word.font_size = 10.0;
word.is_strikeout = true;
let mut digit = make_merge_item("2", 124.5, 4.0);
digit.font_size = 6.0;
digit.y = word.y + 3.0;
let merged = merge_subscript_items(vec![word.clone(), digit.clone()]);
assert_eq!(merged.len(), 2);
// Same marks still merge (footnote ref inside the strike).
digit.is_strikeout = true;
let merged = merge_subscript_items(vec![word, digit]);
assert_eq!(merged.len(), 1);
assert!(merged[0].text.starts_with("word"));
}
#[test]
@@ -1308,7 +1021,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1324,7 +1036,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1340,7 +1051,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1396,7 +1106,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1412,7 +1121,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1428,7 +1136,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1455,7 +1162,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1471,7 +1177,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1487,7 +1192,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1516,7 +1220,6 @@ mod tests {
is_bold: true,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1552,7 +1255,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1589,7 +1291,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1605,7 +1306,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1621,7 +1321,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1645,7 +1344,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1759,7 +1457,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1775,7 +1472,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1801,7 +1497,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1817,7 +1512,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1859,7 +1553,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1905,7 +1598,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1951,7 +1643,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1990,7 +1681,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
+1 -64
View File
@@ -231,27 +231,8 @@ fn rule_matches_item(rule: &Rule, item: &TextItem) -> bool {
overlap >= min_overlap
}
/// Strikeout window: a rule crossing the glyphs. Strikethroughs sit at
/// roughly 20-35% of the em above the baseline (about half the x-height);
/// accept a band well inside the glyph body so baseline underlines and
/// overlines never qualify.
fn rule_strikes_item(rule: &Rule, item: &TextItem) -> bool {
let y_min = item.y + item.font_size * 0.12;
let y_max = item.y + item.font_size * 0.55;
if rule.y < y_min || rule.y > y_max {
return false;
}
let ix1 = item.x;
let ix2 = item.x + item.width;
let min_overlap = item.width * MIN_X_OVERLAP;
let overlap = rule.x2.min(ix2) - rule.x1.max(ix1);
overlap >= min_overlap
}
/// Mark `is_underline` on text items that have a horizontal rule just
/// below their baseline, and `is_strikeout` on items whose glyphs a rule
/// crosses at mid x-height. `items`, `rects`, and `lines` are a single
/// below their baseline. `items`, `rects`, and `lines` are a single
/// page's extraction output (all in PDF coordinates, y-up, where
/// `TextItem::y` is the text baseline).
pub(crate) fn mark_underlined_items(
@@ -277,11 +258,6 @@ pub(crate) fn mark_underlined_items(
}
if rule_matches_item(rule, item) {
item.is_underline = true;
}
if rule_strikes_item(rule, item) {
item.is_strikeout = true;
}
if item.is_underline && item.is_strikeout {
break;
}
}
@@ -306,7 +282,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -383,44 +358,6 @@ mod tests {
assert!(!items[0].is_underline);
}
#[test]
fn mid_glyph_rule_marks_strikeout_not_underline() {
// Rule at ~30% of the em above the baseline crosses the glyphs.
let mut items = vec![item("struck out", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 503.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items[0].is_strikeout);
assert!(!items[0].is_underline);
}
#[test]
fn baseline_rule_marks_underline_not_strikeout() {
let mut items = vec![item("underlined", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 498.5)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items[0].is_underline);
assert!(!items[0].is_strikeout);
}
#[test]
fn overline_is_neither_underline_nor_strikeout() {
// Rule just above the cap height (overline / next line's rule).
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 507.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(!items[0].is_underline);
assert!(!items[0].is_strikeout);
}
#[test]
fn thin_filled_rect_at_mid_glyph_marks_strikeout() {
let mut items = vec![item("struck out", 100.0, 500.0, 60.0, 10.0)];
let rects = vec![thin_rect(100.0, 502.6, 60.0)];
mark_underlined_items(&mut items, &rects, &[], 1);
assert!(items[0].is_strikeout);
assert!(!items[0].is_underline);
}
#[test]
fn line_above_baseline_is_not_an_underline() {
// Strikethrough / overline geometry must not mark.
+5 -27
View File
@@ -1,6 +1,5 @@
//! Form XObject and image XObject extraction.
use super::fonts::descriptor_style_flags;
use crate::text_utils::{effective_font_size, expand_ligatures, is_bold_font, is_italic_font};
use crate::tounicode::FontCMaps;
use crate::types::{ItemType, TextItem};
@@ -9,7 +8,7 @@ use std::collections::HashMap;
use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache, FontStyleCache,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
};
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
@@ -115,7 +114,6 @@ pub(crate) fn extract_form_xobject_text(
font_cmaps: &FontCMaps,
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
) -> Vec<TextItem> {
extract_form_xobject_text_inner(
doc,
@@ -124,12 +122,10 @@ pub(crate) fn extract_form_xobject_text(
font_cmaps,
parent_ctm,
cmap_decisions,
style_cache,
0,
)
}
#[allow(clippy::too_many_arguments)]
fn extract_form_xobject_text_inner(
doc: &Document,
form_id: ObjectId,
@@ -137,7 +133,6 @@ fn extract_form_xobject_text_inner(
font_cmaps: &FontCMaps,
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
depth: u8,
) -> Vec<TextItem> {
use lopdf::content::Content;
@@ -172,7 +167,6 @@ fn extract_form_xobject_text_inner(
let mut font_tounicode_refs: HashMap<String, u32> = HashMap::new();
let mut inline_cmaps: HashMap<String, crate::tounicode::CMapEntry> = HashMap::new();
let mut font_style_flags: HashMap<String, (bool, bool)> = HashMap::new();
for (font_name, font_dict) in &form_fonts {
let resource_name = String::from_utf8_lossy(font_name).to_string();
if let Ok(base_font) = font_dict.get(b"BaseFont") {
@@ -181,10 +175,6 @@ fn extract_form_xobject_text_inner(
font_base_names.insert(resource_name.clone(), base_name);
}
}
let style = descriptor_style_flags(doc, font_dict, style_cache);
if style != (false, false) {
font_style_flags.insert(resource_name.clone(), style);
}
match font_dict.get(b"ToUnicode") {
Ok(tounicode) => {
if let Ok(obj_ref) = tounicode.as_reference() {
@@ -282,7 +272,6 @@ fn extract_form_xobject_text_inner(
font_cmaps,
&ctm,
cmap_decisions,
style_cache,
depth + 1,
);
items.extend(nested_items);
@@ -306,7 +295,6 @@ fn extract_form_xobject_text_inner(
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Image,
mcid: None,
});
@@ -440,10 +428,6 @@ fn extract_form_xobject_text_inner(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
items.push(TextItem {
text: expand_ligatures(&text),
x,
@@ -453,10 +437,9 @@ fn extract_form_xobject_text_inner(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -577,10 +560,6 @@ fn extract_form_xobject_text_inner(
.get(&current_font)
.map(|s| s.as_str())
.unwrap_or(&current_font);
let (desc_italic, desc_bold) = font_style_flags
.get(&current_font)
.copied()
.unwrap_or((false, false));
let scale_x = text_matrix[0] * ctm[0] + text_matrix[1] * ctm[2];
for (text, start_w, end_w) in &sub_items {
let offset_tm = [
@@ -607,10 +586,9 @@ fn extract_form_xobject_text_inner(
font: current_font.clone(),
font_size: rendered_size,
page: page_num,
is_bold: is_bold_font(base_font) || desc_bold,
is_italic: is_italic_font(base_font) || desc_italic,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
+473 -273
View File
@@ -39,7 +39,6 @@ pub mod markdown;
pub mod process_mode;
pub mod structure_tree;
pub mod tables;
mod text_quality;
pub mod text_utils;
pub mod tounicode;
pub mod types;
@@ -61,10 +60,6 @@ pub use types::{LayoutComplexity, PdfLine, PdfRect, TextItem};
use lopdf::Document;
use std::collections::{BTreeMap, HashMap, HashSet};
use std::path::Path;
use text_quality::{
analyze_text_quality, detect_encoding_issues, is_cid_garbage, is_garbage_text,
region_items_have_decoding_issue,
};
use tounicode::FontCMaps;
/// OCR reason emitted when the extracted text layer appears garbled due to
@@ -125,7 +120,7 @@ pub struct PdfProcessResult {
/// .mode(ProcessMode::Analyze)
/// .pages([1, 3, 5]);
/// ```
#[derive(Clone)]
#[derive(Debug, Clone)]
pub struct PdfOptions {
/// How far the pipeline should run (default: [`ProcessMode::Full`]).
pub mode: ProcessMode,
@@ -135,23 +130,6 @@ pub struct PdfOptions {
pub markdown: MarkdownOptions,
/// Optional set of 1-indexed pages to process. `None` = all pages.
pub page_filter: Option<HashSet<u32>>,
/// Password for decrypting an encrypted PDF. `None` falls back to the
/// empty password (owner-only encryption).
pub password: Option<String>,
}
// Manual `Debug` so the password is never leaked through debug logging or a
// panic that formats the options; it renders as `Some("[REDACTED]")`.
impl std::fmt::Debug for PdfOptions {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("PdfOptions")
.field("mode", &self.mode)
.field("detection", &self.detection)
.field("markdown", &self.markdown)
.field("page_filter", &self.page_filter)
.field("password", &self.password.as_ref().map(|_| "[REDACTED]"))
.finish()
}
}
impl Default for PdfOptions {
@@ -161,7 +139,6 @@ impl Default for PdfOptions {
detection: DetectionConfig::default(),
markdown: MarkdownOptions::default(),
page_filter: None,
password: None,
}
}
}
@@ -203,12 +180,6 @@ impl PdfOptions {
self.page_filter = Some(pages.into_iter().collect());
self
}
/// Set the password used to decrypt an encrypted PDF.
pub fn password(mut self, password: impl Into<String>) -> Self {
self.password = Some(password.into());
self
}
}
// =========================================================================
@@ -241,8 +212,7 @@ pub fn process_pdf_with_options<P: AsRef<Path>>(
validate_pdf_file(&path)?;
// Load the document once — shared by detection AND extraction.
let (doc, page_count) =
load_document_from_path_with_password(&path, options.password.as_deref())?;
let (doc, page_count) = load_document_from_path(&path)?;
process_document(doc, page_count, options, start)
}
@@ -267,8 +237,7 @@ pub fn process_pdf_mem_with_options(
let start = std::time::Instant::now();
validate_pdf_bytes(buffer)?;
let (doc, page_count) =
load_document_from_mem_with_password(buffer, options.password.as_deref())?;
let (doc, page_count) = load_document_from_mem(buffer)?;
process_document(doc, page_count, options, start)
}
@@ -610,7 +579,6 @@ pub fn extract_text_in_regions_mem(
let mut gid_pages: HashSet<u32> = HashSet::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
let mut style_cache = extractor::FontStyleCache::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
@@ -629,7 +597,6 @@ pub fn extract_text_in_regions_mem(
*page_num,
&font_cmaps,
false,
&mut style_cache,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -732,7 +699,6 @@ pub fn extract_tables_in_regions_mem(
let mut gid_pages: HashSet<u32> = HashSet::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
let mut style_cache = extractor::FontStyleCache::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
@@ -748,7 +714,6 @@ pub fn extract_tables_in_regions_mem(
*page_num,
&font_cmaps,
false,
&mut style_cache,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -1054,7 +1019,6 @@ pub fn detect_vector_grid_in_region_mem(
page_1idx,
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
)?;
text_utils::fix_letterspaced_items(&mut items);
@@ -1241,15 +1205,8 @@ mod vector_grid_tests {
let &page_id = pages.get(&1).unwrap();
let needed: HashSet<u32> = HashSet::from([1]);
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
&doc,
page_id,
1,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
)
.unwrap();
let ((items, rects, _lines), _has_gid, _rotated) =
extract_page_text_items(&doc, page_id, 1, &cmaps, false).unwrap();
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, 1);
assert_eq!(rect_tables.len(), 1, "expected one rect-detected table");
@@ -1283,15 +1240,8 @@ mod vector_grid_tests {
let &page_id = pages.get(&page_num).unwrap();
let needed: HashSet<u32> = HashSet::from([page_num]);
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
&doc,
page_id,
page_num,
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
)
.unwrap();
let ((items, rects, _lines), _has_gid, _rotated) =
extract_page_text_items(&doc, page_id, page_num, &cmaps, false).unwrap();
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, page_num);
rect_tables
@@ -2016,7 +1966,6 @@ pub fn extract_tables_with_structure_cells_mem(
let mut page_heights: HashMap<u32, f32> = HashMap::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
let mut style_cache = extractor::FontStyleCache::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
@@ -2032,7 +1981,6 @@ pub fn extract_tables_with_structure_cells_mem(
*page_num,
&font_cmaps,
false,
&mut style_cache,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -2834,7 +2782,6 @@ fn detect_tsr_quality_issue(
page_1idx,
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
)?;
let adaptive_threshold = text_utils::fix_letterspaced_items(&mut items);
let coords = if coords_rotated {
@@ -3293,40 +3240,24 @@ fn tsr_region_contains_item(item: &TextItem, bounds: RegionBounds) -> bool {
/// page count from it directly to avoid the metadata-only round-trip.
pub(crate) fn load_document_from_path<P: AsRef<Path>>(
path: P,
) -> Result<(Document, u32), PdfError> {
load_document_from_path_with_password(path, None)
}
/// Load a PDF file, decrypting with `password` if the file is encrypted.
pub(crate) fn load_document_from_path_with_password<P: AsRef<Path>>(
path: P,
password: Option<&str>,
) -> Result<(Document, u32), PdfError> {
let buffer = std::fs::read(&path)?;
load_document_from_mem_with_password(&buffer, password)
load_document_from_mem(&buffer)
}
/// Load a PDF from a memory buffer.
pub(crate) fn load_document_from_mem(buffer: &[u8]) -> Result<(Document, u32), PdfError> {
load_document_from_mem_with_password(buffer, None)
}
/// Load a PDF from a memory buffer, decrypting with `password` if encrypted.
pub(crate) fn load_document_from_mem_with_password(
buffer: &[u8],
password: Option<&str>,
) -> Result<(Document, u32), PdfError> {
// Fix malformed struct element names before parsing. Some PDF generators
// write bare names (/S Code) instead of proper PDF names (/S /Code), which
// causes lopdf to silently drop the entire object.
let fixed = structure_tree::fix_bare_struct_names(buffer);
let buf = fixed.as_ref();
let doc = match load_document_bytes(buf, password) {
let doc = match load_document_bytes(buf) {
Ok(doc) => doc,
Err(first_err) => {
for repaired in repair_pdf_container_candidates(buf) {
match load_document_bytes(&repaired, password) {
match load_document_bytes(&repaired) {
Ok(doc) => {
log::debug!("loaded PDF after repairing malformed container bytes");
let page_count = doc.get_pages().len() as u32;
@@ -3346,31 +3277,13 @@ pub(crate) fn load_document_from_mem_with_password(
Ok((doc, page_count))
}
fn load_document_bytes(buf: &[u8], password: Option<&str>) -> Result<Document, lopdf::Error> {
fn load_document_bytes(buf: &[u8]) -> Result<Document, lopdf::Error> {
match Document::load_mem(buf) {
// Some encrypted PDFs load structurally but leave their streams
// encrypted (`is_encrypted()` stays true); reading them yields garbage
// until we re-load with a password. Others fail load_mem outright with
// an encryption error. Handle both by re-loading with the password.
Ok(doc) if doc.is_encrypted() => decrypt_document_bytes(buf, password),
Ok(doc) => Ok(doc),
Err(ref e) if is_encrypted_lopdf_error(e) => decrypt_document_bytes(buf, password),
Err(e) => Err(e),
}
}
/// Re-load an encrypted PDF, decrypting with `password`. Falls back to the
/// empty password (owner-only encryption, the common "protected" case) when a
/// non-empty password was supplied but rejected.
fn decrypt_document_bytes(buf: &[u8], password: Option<&str>) -> Result<Document, lopdf::Error> {
let pw = password.unwrap_or("");
match Document::load_mem_with_options(buf, lopdf::LoadOptions::with_password(pw)) {
Ok(doc) => Ok(doc),
Err(inner) if !pw.is_empty() => {
Err(ref e) if is_encrypted_lopdf_error(e) => {
Document::load_mem_with_options(buf, lopdf::LoadOptions::with_password(""))
.map_err(|_| inner)
}
Err(inner) => Err(inner),
Err(e) => Err(e),
}
}
@@ -3768,15 +3681,130 @@ fn process_document(
// Internal helpers
// =========================================================================
/// Detect broken font encodings in extracted markdown text.
///
/// Two heuristics:
/// 1. **U+FFFD**: Any replacement character indicates decode failures.
/// 2. **Dollar-as-space**: Pattern like `Word$Word$Word` where `$` is used as a
/// word separator due to broken ToUnicode CMaps. Triggers when either:
/// - More than 50% of `$` are between letters (clear substitution pattern), OR
/// - More than 20 letter-dollar-letter occurrences (even if some `$` are also
/// used as trailing/leading separators, 20+ is far beyond normal financial text).
fn detect_encoding_issues(markdown: &str) -> bool {
// Heuristic 1: U+FFFD replacement characters
if markdown.contains('\u{FFFD}') {
return true;
}
// Heuristic 2: dollar-as-space pattern
has_dollar_as_space_pattern(markdown)
}
fn has_dollar_as_space_pattern(markdown: &str) -> bool {
let total_dollars = markdown.matches('$').count();
if total_dollars > 10 {
let bytes = markdown.as_bytes();
let mut letter_dollar_letter = 0usize;
for i in 1..bytes.len().saturating_sub(1) {
if bytes[i] == b'$'
&& bytes[i - 1].is_ascii_alphabetic()
&& bytes[i + 1].is_ascii_alphabetic()
{
letter_dollar_letter += 1;
}
}
if letter_dollar_letter > 20 || letter_dollar_letter * 2 > total_dollars {
return true;
}
}
false
}
#[derive(Debug, Default)]
struct TextQualityReport {
pages_needing_ocr: Vec<u32>,
has_encoding_issues: bool,
reasons_by_page: BTreeMap<u32, Vec<String>>,
}
#[derive(Debug, Default)]
struct PageTextQualityEvidence {
chars: usize,
replacement_chars: usize,
replacement_spans: usize,
longest_replacement_run: usize,
ascii_letters: usize,
ascii_vowels: usize,
ascii_word_tokens: usize,
ascii_common_word_hits: usize,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum TextSpanIssueKind {
Replacement,
Strong,
}
fn analyze_text_quality(items: &[TextItem]) -> TextQualityReport {
let mut reasons_by_page = BTreeMap::new();
let mut evidence_by_page = BTreeMap::<u32, PageTextQualityEvidence>::new();
for item in items {
if !matches!(item.item_type, crate::types::ItemType::Text) {
continue;
}
let evidence = evidence_by_page.entry(item.page).or_default();
evidence.chars += item.text.chars().filter(|ch| !ch.is_whitespace()).count();
record_ascii_language_evidence(evidence, &item.text);
match text_span_decoding_issue_kind(&item.text) {
Some(TextSpanIssueKind::Strong) => {
add_ocr_reason(
&mut reasons_by_page,
item.page,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
Some(TextSpanIssueKind::Replacement) => {
let stats = replacement_text_stats(&item.text);
evidence.replacement_chars += stats.0;
evidence.replacement_spans += 1;
evidence.longest_replacement_run = evidence.longest_replacement_run.max(stats.1);
}
None => {}
}
}
for (page, evidence) in evidence_by_page {
if reasons_by_page.contains_key(&page) {
continue;
}
if page_replacement_evidence_needs_ocr(&evidence)
|| printable_ascii_mojibake_needs_ocr(&evidence)
{
add_ocr_reason(
&mut reasons_by_page,
page,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
}
let pages_needing_ocr: Vec<u32> = reasons_by_page.keys().copied().collect();
TextQualityReport {
has_encoding_issues: !pages_needing_ocr.is_empty(),
pages_needing_ocr,
reasons_by_page,
}
}
fn suspected_garbled_reason() -> String {
OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()
}
pub(crate) fn add_ocr_reason(
reasons_by_page: &mut BTreeMap<u32, Vec<String>>,
page: u32,
reason: &str,
) {
fn add_ocr_reason(reasons_by_page: &mut BTreeMap<u32, Vec<String>>, page: u32, reason: &str) {
let reasons = reasons_by_page.entry(page).or_default();
if !reasons.iter().any(|existing| existing == reason) {
reasons.push(reason.to_string());
@@ -3808,6 +3836,321 @@ fn page_ocr_reasons_vec(reasons_by_page: BTreeMap<u32, Vec<String>>) -> Vec<Page
.collect()
}
fn region_items_have_decoding_issue(items: &[TextItem]) -> bool {
items.iter().any(|item| {
matches!(item.item_type, crate::types::ItemType::Text)
&& text_span_has_decoding_issue(&item.text)
})
}
fn text_span_has_decoding_issue(text: &str) -> bool {
text_span_decoding_issue_kind(text).is_some()
}
fn text_span_decoding_issue_kind(text: &str) -> Option<TextSpanIssueKind> {
let text = text.trim();
if text.is_empty() {
return None;
}
if has_dollar_as_space_pattern(text)
|| has_private_use_text_run(text)
|| is_cid_garbage(text)
|| has_cid_control_token(text)
{
return Some(TextSpanIssueKind::Strong);
}
if has_replacement_text_run(text) {
return Some(TextSpanIssueKind::Replacement);
}
None
}
fn replacement_text_stats(text: &str) -> (usize, usize) {
let mut replacement = 0usize;
let mut current_run = 0usize;
let mut longest_run = 0usize;
for ch in text.chars() {
if ch == '\u{FFFD}' {
replacement += 1;
current_run += 1;
longest_run = longest_run.max(current_run);
} else {
current_run = 0;
}
}
(replacement, longest_run)
}
fn record_ascii_language_evidence(evidence: &mut PageTextQualityEvidence, text: &str) {
let mut word = String::new();
for ch in text.chars() {
if ch.is_ascii_alphabetic() {
evidence.ascii_letters += 1;
if matches!(ch.to_ascii_lowercase(), 'a' | 'e' | 'i' | 'o' | 'u') {
evidence.ascii_vowels += 1;
}
word.push(ch.to_ascii_lowercase());
} else {
flush_ascii_word_evidence(evidence, &mut word);
}
}
flush_ascii_word_evidence(evidence, &mut word);
}
fn flush_ascii_word_evidence(evidence: &mut PageTextQualityEvidence, word: &mut String) {
const COMMON_WORDS: [&str; 46] = [
"the",
"and",
"that",
"for",
"with",
"from",
"this",
"are",
"not",
"our",
"was",
"were",
"which",
"will",
"shall",
"has",
"have",
"had",
"its",
"their",
"these",
"those",
"into",
"over",
"under",
"between",
"during",
"page",
"exhibit",
"certificate",
"company",
"agreement",
"securities",
"dated",
"period",
"ending",
"december",
"february",
"registered",
"holders",
"obligations",
"guarantee",
"deferred",
"compensation",
"telephone",
"pursuant",
];
if word.len() >= 3 {
evidence.ascii_word_tokens += 1;
if COMMON_WORDS.contains(&word.as_str()) {
evidence.ascii_common_word_hits += 1;
}
}
word.clear();
}
fn printable_ascii_mojibake_needs_ocr(evidence: &PageTextQualityEvidence) -> bool {
// A broken ToUnicode bfrange can still produce entirely printable ASCII,
// so character-validity checks alone cannot catch it. Require a large,
// prose-sized sample and combine two independent language signals to keep
// short labels, identifiers, formulas, and non-Latin pages out of scope.
if evidence.ascii_letters < 600 || evidence.ascii_word_tokens < 80 {
return false;
}
if evidence.ascii_letters * 2 < evidence.chars {
return false;
}
let vowel_ratio_bps = evidence.ascii_vowels * 10_000 / evidence.ascii_letters;
let common_word_hit_bps = evidence.ascii_common_word_hits * 10_000 / evidence.ascii_word_tokens;
vowel_ratio_bps <= 3_000 && common_word_hit_bps <= 300
}
fn page_replacement_evidence_needs_ocr(evidence: &PageTextQualityEvidence) -> bool {
if evidence.replacement_chars == 0 || evidence.chars == 0 {
return false;
}
// If the entire page is only a short broken text layer, even a short
// replacement run is enough evidence. On otherwise text-heavy pages,
// require density so math formulas do not force full-page OCR.
if evidence.chars <= 80 && evidence.longest_replacement_run >= 2 {
return true;
}
let replacement_density_bps = evidence.replacement_chars * 10_000 / evidence.chars;
let enough_bad_text = evidence.replacement_chars >= 12 && replacement_density_bps >= 500;
let repeated_bad_spans = evidence.replacement_spans >= 3 && replacement_density_bps >= 250;
let long_bad_run = evidence.longest_replacement_run >= 8 && replacement_density_bps >= 250;
enough_bad_text || repeated_bad_spans || long_bad_run
}
fn has_replacement_text_run(text: &str) -> bool {
let (replacement, longest_run) = replacement_text_stats(text);
longest_run >= 2 || replacement >= 3
}
fn has_private_use_text_run(text: &str) -> bool {
let mut total = 0usize;
let mut private_use = 0usize;
let mut current_run = 0usize;
let mut longest_run = 0usize;
for ch in text.chars() {
if ch.is_whitespace() {
current_run = 0;
continue;
}
total += 1;
if is_private_use_char(ch) {
private_use += 1;
current_run += 1;
longest_run = longest_run.max(current_run);
} else {
current_run = 0;
}
}
if private_use == 0 {
return false;
}
longest_run >= 3 || (total >= 5 && private_use >= 2 && private_use * 2 >= total)
}
fn has_cid_control_token(text: &str) -> bool {
text.split_whitespace().any(token_has_cid_control)
}
fn token_has_cid_control(token: &str) -> bool {
let mut total = 0usize;
let mut c1_control = 0usize;
for ch in token.chars() {
total += 1;
if ('\u{0080}'..='\u{009F}').contains(&ch) {
c1_control += 1;
}
}
total >= 5 && c1_control >= 2 && c1_control * 20 >= total
}
fn is_private_use_char(ch: char) -> bool {
matches!(
ch as u32,
0xE000..=0xF8FF | 0xF0000..=0xFFFFD | 0x100000..=0x10FFFD
)
}
/// Check if extracted text is predominantly garbage (non-alphanumeric).
///
/// Broken font encodings produce text like "----1-.-.-.___ --.-. .._ I_---."
/// where most characters are punctuation/symbols. Real text in any language
/// has >50% alphanumeric characters.
fn is_garbage_text(markdown: &str) -> bool {
let mut alphanum = 0usize;
let mut non_alphanum = 0usize;
let chars: Vec<char> = markdown.chars().collect();
let mut i = 0usize;
while i < chars.len() {
let ch = chars[i];
let mut run_end = i + 1;
while run_end < chars.len() && chars[run_end] == ch {
run_end += 1;
}
let is_decorative_leader = matches!(ch, '.' | '_' | '·') && run_end - i >= 3;
if !is_decorative_leader {
for &run_ch in &chars[i..run_end] {
if run_ch.is_whitespace() {
continue;
}
// Skip markdown syntax chars that we add (not from the PDF)
if matches!(run_ch, '#' | '*' | '|' | '-' | '\n') {
continue;
}
if run_ch.is_alphanumeric() {
alphanum += 1;
} else {
non_alphanum += 1;
}
}
}
i = run_end;
}
let total = alphanum + non_alphanum;
total >= 50 && alphanum * 2 < total
}
/// Detect garbage from failed CID-to-Unicode mapping on Identity-H fonts.
///
/// When CID values don't correspond to Unicode codepoints, the raw bytes often
/// produce characters in the C1 control range (U+0080U+009F) or Private Use
/// Area, mixed with random Latin Extended characters. Valid text in any
/// language almost never contains C1 controls. We also fall back to the
/// general `is_garbage_text` check for non-alphanumeric-heavy patterns.
fn is_cid_garbage(text: &str) -> bool {
if is_garbage_text(text) {
return true;
}
let mut total = 0usize;
let mut c1_control = 0usize;
let mut high_latin = 0usize;
for ch in text.chars() {
if ch.is_whitespace() {
continue;
}
total += 1;
// C1 control characters (U+0080U+009F) — almost never in real text
if ch == '·' {
continue;
}
if ('\u{0080}'..='\u{009F}').contains(&ch) {
c1_control += 1;
}
// High Latin-1 (U+00A0U+00FF) — legitimate in Western European text
// but when combined with ASCII in CID passthrough, indicates mojibake
// from CID values being misinterpreted as Latin-1 characters.
if ('\u{00A0}'..='\u{00FF}').contains(&ch) {
high_latin += 1;
}
}
if total < 5 {
return false;
}
// If ≥5% of non-whitespace chars are C1 controls, it's garbage
if c1_control >= 2 && c1_control * 20 >= total {
return true;
}
// If ≥40% of non-whitespace chars are high Latin-1 AND the text has few
// ASCII letters, it's likely CID-as-Latin-1 mojibake (Japanese/CJK PDFs
// where CID values 0x80-0xFF become accented Latin characters). Keep a
// minimum length so short math tokens like "2×()×" do not route a clean
// page to OCR.
let ascii_letters = text.chars().filter(|c| c.is_ascii_alphabetic()).count();
total >= 20 && high_latin * 5 >= total * 2 && ascii_letters * 3 < total
}
/// Detect markdown tables with suspicious structure that suggest the heuristic
/// missed/mangled rows or columns. Returns true when the caller should treat
/// the result as `needs_ocr` and fall back to GPU OCR.
@@ -4656,7 +4999,6 @@ mod text_cluster_column_undercount_tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -4932,7 +5274,6 @@ mod table_candidate_selection_tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -5680,7 +6021,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -5728,171 +6068,6 @@ mod tests {
assert!(!detect_encoding_issues(text));
}
/// Real garbled output from ParseBench `text_simple__att10k.pdf`: a broken
/// ToUnicode CMap shifts every character by a per-range constant, so
/// "Certificate of Designations with respect to Series C" extracts as
/// pure-ASCII ciphertext (issue #118).
const SHIFTED_CIPHER_TEXT: &str =
"8VceZWZTReVW9VdZXReZdhZeYcVdaVTeeHVcZVd8EcVWVccVUHeT:iYZSZe-(,e2'WZ]V \
;VScfRcj2&,*,*$ .'R CZdecfVehYZTYUVWZVdeYVcZXYedWY]UVcdW]X'eVcUVSeYVcVXZdecReRUR]]ed \
TCBGC@ZUReVUdfSdZUZRcZVdZdWZ]VUYVcVhZeYafcdfReeVXf]ReZV0*S$.$ZZZ$6$&ViTVae \
WceYVZdecfVedcVWVccVUe.'S&.'T&.'U&.'V&.'WSV]h(EfcdfReeYZdcVXf]ReZYV \
cVXZdecReYVcVSjRXcVVdeWfcZdYRTajWRjdfTYZdecfVeWZ]VUYVcVhZeYeYVH:8 \
cVbfVde( .'S fRcRejWTVceRZZXReZTZWZT7V]]IV]VaYV8(RUHfeYhVdeVc7V]]IV]VaYV8 \
:iYZSZe.'TeYVaVcZUVUZX9VTVSVc-&,*$";
#[test]
fn test_detect_encoding_issues_shifted_cipher_text() {
assert!(detect_encoding_issues(SHIFTED_CIPHER_TEXT));
}
#[test]
fn test_shifted_cipher_below_sample_threshold_not_flagged() {
// Fewer than 200 ASCII letters: not enough evidence to condemn.
assert!(!detect_encoding_issues("8VceZWZTReVW9VdZXReZdhZeY"));
}
#[test]
fn test_clean_prose_not_flagged_as_cipher() {
let prose = "Certificate of Designations with respect to Series C Preferred Stock, \
filed February 18, 2020. No instrument which defines the rights of holders of \
long-term debt of the registrant and all of its consolidated subsidiaries is \
filed herewith pursuant to Regulation S-K, Item 601. Pursuant to this regulation, \
the registrant hereby agrees to furnish a copy of any such instrument to the SEC \
upon request.";
assert!(!detect_encoding_issues(prose));
}
#[test]
fn test_camel_case_code_not_flagged_as_cipher() {
let code = "The getElementById and querySelectorAll methods return DOM nodes. Use \
addEventListener with removeEventListener, requestAnimationFrame with \
cancelAnimationFrame, and setTimeout with clearTimeout. The XMLHttpRequest \
object exposes onreadystatechange, responseText and getAllResponseHeaders. \
Prefer createElement, appendChild, insertBefore and replaceChild for DOM \
manipulation, and getBoundingClientRect for layout measurement.";
assert!(!detect_encoding_issues(code));
}
#[test]
fn test_accented_european_text_not_flagged_as_cipher() {
let swedish = "Regeringen föreslår att riksdagen antar förslaget till lag om ändring \
i skatteförfarandelagen. Bestämmelserna föreslås träda i kraft den första januari. \
Förslaget innebär att företag med säte i utlandet måste lämna särskilda uppgifter \
till Skatteverket varje kvartal, och att avgiften höjs för överträdelser av de nya \
bestämmelserna om rapporteringsskyldighet för gränsöverskridande arrangemang.";
assert!(!detect_encoding_issues(swedish));
}
#[test]
fn test_all_caps_text_not_flagged_as_cipher() {
let caps = "EXHIBIT INDEX PURSUANT TO ITEM 601 OF REGULATION SK CERTIFICATE OF \
DESIGNATIONS WITH RESPECT TO SERIES C PREFERRED STOCK FILED FEBRUARY EIGHTEEN \
TWENTY TWENTY AND INCORPORATED HEREIN BY REFERENCE TO THE ANNUAL REPORT ON FORM \
TENK FOR THE PERIOD ENDED DECEMBER THIRTYFIRST TWENTY NINETEEN AS AMENDED";
assert!(!detect_encoding_issues(caps));
}
// A Caesar shift of prose that stays within a single case block does not
// trigger the case-shift signal, but it permutes the letter histogram: the
// frequency SHAPE stays English-like while letter POSITIONS scramble, so
// the permutation signal catches it regardless of case.
const CAESAR_PROSE: &str =
"The registrant hereby agrees to furnish a copy of any such instrument to the \
Commission upon request. This certificate of designations was filed February with \
respect to Series Preferred Stock and incorporated herein by reference to the annual \
report on form for the period ended December as amended and restated thereafter.";
fn caesar_shift(text: &str, k: u8) -> String {
text.chars()
.map(|c| match c {
'a'..='z' => (((c as u8 - b'a' + k) % 26) + b'a') as char,
'A'..='Z' => (((c as u8 - b'A' + k) % 26) + b'A') as char,
_ => c,
})
.collect()
}
#[test]
fn test_mixed_case_caesar_shift_flagged() {
assert!(detect_encoding_issues(&caesar_shift(CAESAR_PROSE, 3)));
}
#[test]
fn test_all_lowercase_caesar_shift_flagged() {
// Uniform all-lowercase garbled prose: the earlier mixed-case guard
// would have exempted this, so it must be caught by the case-agnostic
// permutation signal instead.
assert!(detect_encoding_issues(&caesar_shift(
&CAESAR_PROSE.to_lowercase(),
5
)));
}
#[test]
fn test_all_uppercase_caesar_shift_flagged() {
assert!(detect_encoding_issues(&caesar_shift(
&CAESAR_PROSE.to_uppercase(),
7
)));
}
#[test]
fn test_dna_sequence_not_flagged_as_cipher() {
// Non-linguistic ASCII: unlike English (low vowel ratio, low cosine)
// but not garbled. Its 4-letter alphabet makes the frequency profile
// too steep, so the shape cosine falls below the permutation threshold.
let dna = "ACGT".repeat(120);
assert!(!detect_encoding_issues(&dna));
}
#[test]
fn test_protein_sequence_not_flagged_as_cipher() {
let protein =
"MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEKAVQVKVKALPDAQFEVVHSLAKWKR"
.repeat(6);
assert!(!detect_encoding_issues(&protein));
}
#[test]
fn test_ticker_list_not_flagged_as_cipher() {
let tickers = "AAPL MSFT GOOG TSLA NVDA AMZN META NFLX AMD INTC CSCO ORCL CRM ADBE QCOM \
TXN AVGO MU LRCX KLAC ASML SNPS CDNS FTNT PANW "
.repeat(4);
assert!(!detect_encoding_issues(&tickers));
}
#[test]
fn test_text_quality_flags_shifted_cipher_page() {
let items = vec![
test_text_item_on_page(1, SHIFTED_CIPHER_TEXT),
test_text_item_on_page(2, "A clean second page should not be routed to OCR."),
];
let quality = analyze_text_quality(&items);
assert!(quality.has_encoding_issues);
assert_eq!(quality.pages_needing_ocr, vec![1]);
assert_eq!(
quality.reasons_by_page.get(&1).cloned(),
Some(vec![OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()])
);
}
#[test]
fn test_text_quality_cipher_stats_accumulate_across_items() {
// The garbled page arrives as many short spans; no single span has
// enough letters to flag on its own.
let items: Vec<TextItem> = SHIFTED_CIPHER_TEXT
.split_whitespace()
.map(|chunk| test_text_item_on_page(1, chunk))
.collect();
let quality = analyze_text_quality(&items);
assert_eq!(quality.pages_needing_ocr, vec![1]);
}
#[test]
fn test_text_quality_flags_localized_cid_mojibake_span() {
let items = vec![
@@ -5938,6 +6113,31 @@ mod tests {
);
}
#[test]
fn test_text_quality_flags_printable_ascii_mojibake() {
let shifted = "8VceZWZTReVW9VdZXReZdhZeYcVdaVTeeHVcZVd8EcVWVccVUHeT:iYZSZe VScfRcj CZdecfVehYZTYUVWZVdeYVcZXYedWY]UVcdW]X'eVcUVSeYVcVXZdecReRUR]]ed";
let items = vec![test_text_item_on_page(1, &shifted.repeat(12))];
let quality = analyze_text_quality(&items);
assert_eq!(quality.pages_needing_ocr, vec![1]);
assert_eq!(
quality.reasons_by_page.get(&1).cloned(),
Some(vec![OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()])
);
}
#[test]
fn test_text_quality_allows_large_printable_ascii_prose() {
let prose = "The company has entered into an agreement with registered holders during the period ending in December. ";
let items = vec![test_text_item_on_page(1, &prose.repeat(20))];
let quality = analyze_text_quality(&items);
assert!(!quality.has_encoding_issues);
assert!(quality.pages_needing_ocr.is_empty());
}
#[test]
fn test_text_quality_allows_clean_multilingual_and_latin1_text() {
let items = vec![
-183
View File
@@ -130,123 +130,6 @@ pub(crate) fn has_dot_leaders(text: &str) -> bool {
dot_groups >= 2
}
/// Detect a table-of-contents entry: a line ending in a page number preceded by
/// a dot-leader group (e.g. "Measurement Lab worksheet ... 3"). `has_dot_leaders`
/// misses single-group leaders ("..."), but a trailing "<dots> <number>" is a
/// strong TOC signal on its own. Such lines must never be promoted to headings.
pub(crate) fn is_toc_entry_line(text: &str) -> bool {
let trimmed = text.trim_end();
let digits = trimmed
.chars()
.rev()
.take_while(|c| c.is_ascii_digit())
.count();
if digits == 0 || digits > 4 {
return false;
}
let before_number = trimmed[..trimmed.len() - digits].trim_end();
let dots = before_number
.chars()
.rev()
.take_while(|c| *c == '.')
.count();
dots >= 3
}
/// A heading that announces a table of contents ("Contents", "Table of
/// Contents"). Lines after it on the same page are ToC entries — section
/// titles that look exactly like headings but must not be promoted.
pub(crate) fn is_toc_marker_heading(text: &str) -> bool {
let t = text.trim().trim_end_matches(':').trim().to_lowercase();
matches!(t.as_str(), "contents" | "table of contents")
}
/// Lines that resemble headings structurally but are display-math fragments:
/// equations ending in an equation number ("S = kB ln W, (2)") or equation
/// lead-ins ("Rearranging Equation (8) gives:"). Both carry an "(N)" equation
/// reference — but a trailing "(N)" alone is not enough: real headings end
/// with parenthesized numbers too ("Nicaea (325)", appendix numbering), so
/// the suffix form additionally requires math evidence — an "=" in the line
/// or a comma immediately before the number, both present in every display
/// equation and absent from name-plus-number headings. A bare trailing colon
/// is NOT a fragment signal either: real headings frequently end with colons
/// ("Procedure:", "Steps for Using the Microscope:").
pub(crate) fn is_heading_fragment(text: &str) -> bool {
let t = text.trim_end();
// A lowercase-initial one-or-two-word "heading" is a mid-sentence
// fragment beside display math ("or inversely", "and therefore") —
// real headings that short start uppercase. Measured as spurious
// headings on academic docs (fire-pdf ENG-5029 / opendataloader MHS).
{
let words: Vec<&str> = t.split_whitespace().collect();
if words.len() <= 2 {
if let Some(first_alpha) = t.chars().find(|c| c.is_alphabetic()) {
if first_alpha.is_lowercase() {
return true;
}
}
}
}
fn is_equation_number(s: &str) -> bool {
s.strip_prefix('(')
.and_then(|r| r.strip_suffix(')'))
.is_some_and(|inner| {
!inner.is_empty() && inner.len() <= 3 && inner.chars().all(|c| c.is_ascii_digit())
})
}
// Equation-number suffix with math evidence: "S = kB ln W, (2)"
let mut rev = t.rsplit(' ');
let last = rev.next().unwrap_or("");
if is_equation_number(last) {
// Page-of-total running headers: "LIVSMEDELSVERKET PM 2 (10)"
if let Some(prev_word) = t.rsplit(' ').nth(1) {
if let (Ok(page), Some(total)) = (
prev_word.parse::<u32>(),
last.trim_start_matches('(')
.trim_end_matches(')')
.parse::<u32>()
.ok(),
) {
if page <= total {
return true;
}
}
}
let punct_before = rev
.next()
.is_some_and(|w| w.ends_with(',') || w.ends_with(':'));
let has_math_op = t.chars().any(|c| {
matches!(
c,
'=' | '<'
| '>'
| '≤'
| '≥'
| '≪'
| '≫'
| '≈'
| '≠'
| '±'
| '∑'
| '∫'
| '√'
| '∝'
)
});
if punct_before || has_math_op {
return true;
}
}
// Lead-in: ends with a colon AND references an equation number inline
if t.ends_with(':') && t.split_whitespace().any(is_equation_number) {
return true;
}
false
}
/// Compute the Y-gap threshold for paragraph break detection.
///
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
@@ -437,69 +320,3 @@ pub(crate) fn detect_header_level(
Some(4)
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn toc_entry_with_single_dot_group() {
assert!(is_toc_entry_line("Measurement Lab worksheet ... 3"));
assert!(is_toc_entry_line("Results ........ 12"));
assert!(is_toc_entry_line("Appendix B...42"));
}
#[test]
fn non_toc_lines_pass() {
assert!(!is_toc_entry_line(
"6.2. Expectations for Re-Hiring Employees"
));
assert!(!is_toc_entry_line("What happened in 2020"));
assert!(!is_toc_entry_line("IMPLEMENTATION"));
// Ellipsis without a trailing page number
assert!(!is_toc_entry_line("and so it goes ..."));
// Long numbers are data, not page refs
assert!(!is_toc_entry_line("ISBN ... 97814"));
}
#[test]
fn toc_marker_headings() {
assert!(is_toc_marker_heading("Contents"));
assert!(is_toc_marker_heading("CONTENTS"));
assert!(is_toc_marker_heading("Table of Contents"));
assert!(is_toc_marker_heading("Table of contents:"));
assert!(!is_toc_marker_heading("Contents of the Shipment"));
assert!(!is_toc_marker_heading("Introduction"));
}
#[test]
fn heading_fragments() {
// Equation lead-ins: colon ending + inline equation reference
assert!(is_heading_fragment("or inversely"));
assert!(is_heading_fragment("and therefore"));
assert!(!is_heading_fragment("Introduction"));
assert!(!is_heading_fragment("iPhone Sales Strategy Overview")); // 4 words, exempt
assert!(is_heading_fragment("Rearranging Equation (8) gives:"));
// Display-equation neighbours ending in an equation number
assert!(is_heading_fragment("S = kB ln W, (2)"));
assert!(is_heading_fragment("E = mc2 (12)"));
assert!(is_heading_fragment("x + y = z, (3)"));
// Page-of-total running headers
assert!(is_heading_fragment("LIVSMEDELSVERKET PM 2 (10)"));
// Comparison-operator evidence and colon-before-number
assert!(is_heading_fragment(
"PLL\u{fe} PHH\u{226a} PLH\u{fe} PHL: (12)"
));
// Real headings pass — including name-plus-number and colon-ended ones
assert!(!is_heading_fragment("Nicaea (325)"));
assert!(!is_heading_fragment(
"\u{627}\u{644}\u{645}\u{644}\u{62d}\u{642} \u{631}\u{642}\u{645} (1)"
));
assert!(!is_heading_fragment("4. Entropy"));
assert!(!is_heading_fragment("Procedure:"));
assert!(!is_heading_fragment("Steps for Using the Microscope:"));
assert!(!is_heading_fragment("Changing objectives:"));
assert!(!is_heading_fragment("Sales by Region (2024)"));
assert!(!is_heading_fragment("Results (preliminary)"));
}
}
+5 -75
View File
@@ -7,8 +7,7 @@ use crate::types::TextLine;
use super::analysis::{
bold_heading_level, calculate_font_stats, compute_heading_tiers, compute_paragraph_threshold,
detect_header_level, font_size_rarity, has_dot_leaders, is_heading_fragment, is_toc_entry_line,
is_toc_marker_heading,
detect_header_level, font_size_rarity, has_dot_leaders,
};
use super::classify::{
format_list_item, is_caption_line, is_list_item, is_monospace_font, starts_with_bullet_marker,
@@ -141,11 +140,8 @@ fn find_isolated_lines(lines: &[TextLine], base_size: f32, para_threshold: f32)
}
}
for (&page, &(total, isolated)) in &page_line_counts {
// The ratio only means something on pages dense enough for a
// multi-column misfire; on sparse pages (covers, ToC pages with a
// lone title) one isolated line is 25%+ of the page and exactly the
// line isolation exists to find.
if total >= 10 && isolated as f32 / total as f32 > 0.25 {
if total > 0 && isolated as f32 / total as f32 > 0.25 {
// Too many isolated lines on this page — remove them all
set.retain(|&i| lines[i].page != page);
}
}
@@ -490,7 +486,6 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
let mut in_code_block = false;
let mut prev_had_dot_leaders = false;
let mut paragraph_in_wrapped_bold_run = false;
let mut toc_suppress_page: Option<u32> = None;
let mut inserted_tables: HashSet<(u32, usize)> = HashSet::new();
let mut inserted_images: HashSet<(u32, usize)> = HashSet::new();
@@ -703,22 +698,11 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
_ => false,
};
// Lines explicitly tagged with a non-heading content role must never
// be promoted by the visual heuristic — a tagged list item, quote, or
// code line can look exactly like a heading (short, isolated).
let non_heading_role = struct_role
.as_ref()
.is_some_and(StructRole::is_non_heading_content);
let heuristic_heading = if options.detect_headers
&& !non_heading_role
&& !is_code_line
&& !looks_like_list_continuation
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !starts_with_bullet_marker(plain_trimmed)
&& !is_toc_entry_line(plain_trimmed)
&& !is_heading_fragment(plain_trimmed)
&& toc_suppress_page != Some(line.page)
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
detect_header_level(line_font_size, base_size, &heading_tiers).or_else(|| {
@@ -754,11 +738,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
// paragraph continuity and minor font-size variation
// inflates rarity scores.
let has_strong_signal = all_bold || isolated || (rarity >= 0.97 && word_count <= 8);
// Single-word headings ("IMPLEMENTATION", "CONTENTS") are common;
// accept them only with the strongest signal combination.
let enough_words =
word_count >= 2 || (all_bold && isolated && plain_trimmed.len() >= 4);
if score >= 0.5 && standalone && enough_words && has_strong_signal {
if score >= 0.5 && standalone && word_count >= 2 && has_strong_signal {
Some(bold_heading_level(&heading_tiers))
} else {
None
@@ -783,9 +763,6 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
plain_text.clone()
};
output.push_str(&format!("{} {}\n\n", prefix, heading_text.trim()));
if is_toc_marker_heading(plain_trimmed) {
toc_suppress_page = Some(line.page);
}
in_list = false;
continue;
}
@@ -982,7 +959,6 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
let mut last_list_x: Option<f32> = None;
let mut prev_had_dot_leaders = false;
let mut paragraph_in_wrapped_bold_run = false;
let mut toc_suppress_page: Option<u32> = None;
for (line_idx, line) in lines.iter().enumerate() {
// Page break
@@ -1061,10 +1037,6 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
if options.detect_headers
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !is_toc_entry_line(plain_trimmed)
&& !is_heading_fragment(plain_trimmed)
&& toc_suppress_page != Some(line.page)
&& !(options.detect_code && line.items.iter().any(|i| is_monospace_font(&i.font)))
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
if let Some(header_level) =
@@ -1087,9 +1059,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
+ if all_bold { 0.3 } else { 0.0 }
+ if standalone { 0.2 } else { 0.0 }
+ if isolated { 0.3 } else { 0.0 };
let enough_words =
word_count >= 2 || (all_bold && isolated && plain_trimmed.len() >= 4);
if score >= 0.5 && standalone && enough_words {
if score >= 0.5 && standalone && word_count >= 2 {
return Some(bold_heading_level(&heading_tiers));
}
None
@@ -1108,9 +1078,6 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
plain_text.clone()
};
output.push_str(&format!("{} {}\n\n", prefix, heading_text.trim()));
if is_toc_marker_heading(plain_trimmed) {
toc_suppress_page = Some(line.page);
}
in_list = false;
continue;
}
@@ -1222,7 +1189,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid,
}
@@ -1239,42 +1205,6 @@ mod tests {
}
}
fn line_at(text: &str, page: u32, y: f32) -> TextLine {
let mut item = make_item(text, page, None);
item.y = y;
make_line(vec![item])
}
#[test]
fn isolated_lines_kept_on_sparse_pages() {
// A ToC page with a lone title and one entry far below: the density
// ratio is 50% but the page is too sparse for the multi-column
// misfire the guard targets — the title must stay isolated.
let lines = vec![
line_at("CONTENTS", 1, 700.0),
line_at("Chapter One 5", 1, 500.0),
];
let isolated = find_isolated_lines(&lines, 12.0, 20.0);
assert!(
isolated.contains(&0),
"sparse-page title must stay isolated"
);
}
#[test]
fn isolated_lines_wiped_on_dense_pages() {
// 12 short lines all with paragraph gaps — the multi-column misfire
// shape. The guard must clear them all.
let lines: Vec<TextLine> = (0..12)
.map(|i| line_at("Short column line", 1, 700.0 - i as f32 * 50.0))
.collect();
let isolated = find_isolated_lines(&lines, 12.0, 20.0);
assert!(
isolated.is_empty(),
"dense page of isolated lines must be wiped"
);
}
#[test]
fn test_struct_role_heading() {
let lines = vec![
-1
View File
@@ -1221,7 +1221,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
}
+1 -101
View File
@@ -3,7 +3,7 @@
use std::collections::{HashMap, HashSet};
use crate::structure_tree::StructRole;
use crate::types::{TextItem, TextLine};
use crate::types::TextLine;
use super::analysis::detect_header_level;
@@ -87,41 +87,6 @@ pub(crate) fn merge_heading_lines(
false
};
// Bold headings at body font size never reach a tier, so wrapped ones
// split into two output headings ("…of wood pellets and cost" /
// "structure in Japan"). Merge a fully-bold line into the previous
// fully-bold line when it reads as a wrap continuation: starts
// lowercase, tiny Y gap, and the previous line has no terminal
// punctuation. Kept deliberately narrow — bold list labels and bold
// sentences start with markers or capitals and are unaffected.
let should_merge = should_merge
|| if let Some(prev) = result.last() {
let all_bold = |l: &TextLine| {
!l.items.is_empty() && l.items.iter().all(|i: &TextItem| i.is_bold)
};
let prev_text = prev.text();
let prev_trim = prev_text.trim_end();
let curr_text = line.text();
let curr_trim = curr_text.trim();
let y_gap = prev.y - line.y;
// Both lines must be tier-less: a tiered/tagged bold heading
// followed by bold body text must not absorb it.
line_level.is_none()
&& effective_heading_level(prev, base_size, heading_tiers, struct_roles)
.is_none()
&& prev.page == line.page
&& all_bold(prev)
&& all_bold(&line)
&& y_gap > 0.0
&& y_gap < line_font * 1.6
&& curr_trim.chars().next().is_some_and(|c| c.is_lowercase())
&& !prev_trim.ends_with(['.', ':', ';', '!', '?'])
&& prev_trim.split_whitespace().count() + curr_trim.split_whitespace().count()
<= 20
} else {
false
};
if should_merge {
// Append this line's items to the previous line
let prev = result.last_mut().unwrap();
@@ -578,7 +543,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid,
}
@@ -720,68 +684,4 @@ mod tests {
.unwrap();
assert_eq!(first_header.page, 1, "first occurrence should be on page 1");
}
fn make_bold_line(text: &str, page: u32, y: f32) -> TextLine {
let mut item = make_item(text, 12.0, None);
item.is_bold = true;
TextLine {
items: vec![item],
y,
page,
adaptive_threshold: 0.10,
}
}
#[test]
fn merge_wrapped_bold_heading_lowercase_continuation() {
// Bold-at-body-size heading wrapped across two lines: the second line
// starts lowercase and must merge into the first.
let lines = vec![
make_bold_line(
"3. Perspective of supply and demand balance and cost",
1,
700.0,
),
make_bold_line("structure in Japan", 1, 686.0),
make_line("Body text paragraph follows here.", 12.0, 1, 660.0, None),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "wrapped bold heading should merge");
assert!(result[0].text().contains("cost structure in Japan"));
}
#[test]
fn no_merge_for_bold_sentences_or_new_headings() {
// Second bold line starts with a capital — a new heading or label,
// not a wrap continuation.
let lines = vec![
make_bold_line("Replace", 1, 700.0),
make_bold_line("Trash", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "distinct bold lines must not merge");
// Previous line ends a sentence — continuation must not merge.
let lines = vec![
make_bold_line("This is a bold sentence.", 1, 700.0),
make_bold_line("another bold line", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[], None);
assert_eq!(result.len(), 2, "sentence-final bold line must not merge");
}
#[test]
fn tiered_bold_heading_does_not_absorb_bold_body() {
// Previous line is a tier-level bold heading (16pt vs 12pt body);
// a following lowercase bold body line must NOT merge into it.
let mut heading = make_bold_line("Section Title", 1, 700.0);
heading.items[0].font_size = 16.0;
heading.items[0].height = 16.0;
let lines = vec![
heading,
make_bold_line("emphasized body text continues here", 1, 686.0),
];
let result = merge_heading_lines(lines, 12.0, &[16.0], None);
assert_eq!(result.len(), 2, "tiered heading must not absorb bold body");
}
}
-3
View File
@@ -268,8 +268,6 @@ pub struct PyTextItem {
#[pyo3(get)]
pub is_underline: bool,
#[pyo3(get)]
pub is_strikeout: bool,
#[pyo3(get)]
pub item_type: String,
}
@@ -354,7 +352,6 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item_type_str(&item.item_type),
})
.collect()
-94
View File
@@ -76,52 +76,6 @@ pub enum StructRole {
}
impl StructRole {
/// Content roles whose text must never be promoted to a heading by the
/// visual heuristic. These carry an explicit non-heading meaning in the
/// struct tree (lists, quotes, notes, references, captions, formulas,
/// forms, ToC entries), yet their text is often short and visually
/// isolated — exactly what the heuristic keys on. Heading roles (H, H1H6)
/// and generic container/flow roles (P, Div, Sect, Span, …) are excluded
/// so the heuristic can still fire there.
///
/// `Figure` is deliberately NOT in this set: cover/banner pages routinely
/// tag the document title inside a Figure (alongside a seal or logo), and
/// that title is a real heading. `Formula` and `Form` stay — a line
/// explicitly tagged as an equation or form field is never a heading.
///
/// Table roles (Table/TR/TH/TD/THead/TBody/TFoot) are included so that
/// when table reconstruction falls back and cells reach the line loop as
/// plain text, a short isolated cell — a `TH` column header especially —
/// is not promoted to a heading.
pub(crate) fn is_non_heading_content(&self) -> bool {
matches!(
self,
Self::L
| Self::LI
| Self::Lbl
| Self::LBody
| Self::BlockQuote
| Self::Quote
| Self::Caption
| Self::TOC
| Self::TOCI
| Self::Index
| Self::Note
| Self::Reference
| Self::BibEntry
| Self::Code
| Self::Formula
| Self::Form
| Self::Table
| Self::TR
| Self::TH
| Self::TD
| Self::THead
| Self::TBody
| Self::TFoot
)
}
fn from_name(name: &str) -> Self {
match name {
"Document" => Self::Document,
@@ -902,54 +856,6 @@ fn contains_bytes(haystack: &[u8], needle: &[u8]) -> bool {
mod tests {
use super::*;
#[test]
fn non_heading_content_roles() {
for r in [
StructRole::L,
StructRole::LI,
StructRole::BlockQuote,
StructRole::Quote,
StructRole::Caption,
StructRole::TOC,
StructRole::TOCI,
StructRole::Index,
StructRole::Note,
StructRole::Reference,
StructRole::BibEntry,
StructRole::Code,
StructRole::Formula,
StructRole::Form,
StructRole::Table,
StructRole::TR,
StructRole::TH,
StructRole::TD,
StructRole::THead,
StructRole::TBody,
StructRole::TFoot,
] {
assert!(
r.is_non_heading_content(),
"{r:?} should block heading promotion"
);
}
// Heading and generic container/flow roles must NOT block promotion
for r in [
StructRole::H,
StructRole::H1,
StructRole::H3,
StructRole::P,
StructRole::Div,
StructRole::Sect,
StructRole::Span,
StructRole::Figure,
] {
assert!(
!r.is_non_heading_content(),
"{r:?} should allow heading promotion"
);
}
}
#[test]
fn test_struct_role_from_name() {
assert_eq!(StructRole::from_name("H1"), StructRole::H1);
-1
View File
@@ -105,7 +105,6 @@ pub(crate) fn merge_adjacent_items(items: &[TextItem]) -> (Vec<TextItem>, Vec<Ve
is_bold: first_item.is_bold,
is_italic: first_item.is_italic,
is_underline: first_item.is_underline,
is_strikeout: first_item.is_strikeout,
item_type: first_item.item_type.clone(),
mcid: first_item.mcid,
});
-1
View File
@@ -394,7 +394,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
-124
View File
@@ -1477,10 +1477,6 @@ fn detect_row_stripe_table(
debug!(" row-stripe rejected: sparse outline/prose continuation shape");
return None;
}
if has_dominant_prose_cell(&cells) {
debug!(" row-stripe rejected: dominant prose cell (chart/figure region over body text)");
return None;
}
let column_centers: Vec<f32> = (0..num_cols)
.map(|c| (col_edges[c] + col_edges[c + 1]) / 2.0)
@@ -1499,34 +1495,6 @@ fn detect_row_stripe_table(
Some(Table::new(column_centers, row_centers, cells, item_indices))
}
/// Detect a grid that swallowed body text instead of tabular data.
///
/// Charts (bar graphs, axis gridlines) emit fields of drawing rects that can
/// pass the row-stripe shape test; the resulting "table" then captures the
/// page's prose. The signature: one cell holds an entire paragraph — ≥60 words
/// AND at least a third of all words in the table.
///
/// There is deliberately no row-count exemption. A small table whose single
/// long cell dominates its word count is indistinguishable by content from a
/// phantom grid over body text, and across the regression corpora every such
/// grid observed has been swallowed prose, never a real note table. The costs
/// are also asymmetric: rejecting a real table degrades it to readable prose,
/// while accepting a phantom scrambles the page into Y-interleaved cells.
/// Larger legitimate tables are safe because the one-third-of-total threshold
/// scales with table size.
fn has_dominant_prose_cell(cells: &[Vec<String>]) -> bool {
let mut total_words = 0usize;
let mut max_cell_words = 0usize;
for row in cells {
for cell in row {
let words = cell.split_whitespace().count();
total_words += words;
max_cell_words = max_cell_words.max(words);
}
}
max_cell_words >= 60 && max_cell_words * 3 >= total_words
}
fn row_stripe_is_sparse_prose_outline(cells: &[Vec<String>]) -> bool {
let Some(num_cols) = cells.first().map(|row| row.len()) else {
return false;
@@ -2326,12 +2294,6 @@ fn detect_merged_cluster_table(
);
return None;
}
if has_dominant_prose_cell(&cells) {
debug!(
" merged-cluster rejected: dominant prose cell (chart/figure region over body text)"
);
return None;
}
// No empty columns
for col in 0..num_cols {
@@ -2430,95 +2392,11 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
}
// --- has_dominant_prose_cell ---
fn cells_of(rows: &[&[&str]]) -> Vec<Vec<String>> {
rows.iter()
.map(|r| r.iter().map(|c| c.to_string()).collect())
.collect()
}
#[test]
fn dominant_prose_cell_rejects_swallowed_paragraph() {
// Two cells hold paragraphs (the shape every observed phantom grid
// has: swallowed body text spans multiple cells), rest are chart labels
let para = ["word"; 70].join(" ");
let para2 = ["word"; 35].join(" ");
let cells = cells_of(&[
&[para.as_str(), "81", "76"],
&[para2.as_str(), "56", "9"],
&["2019", "2020", ""],
]);
assert!(has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_rejects_small_table_dominated_by_one_cell() {
// Boundary case, documented as INTENDED: a small grid whose single
// long cell dominates the word count is rejected even at 4+ rows.
// By content alone this shape is indistinguishable from a phantom
// grid over body text, and every observed instance in the regression
// corpora was swallowed prose (chart/figure regions), not a real
// note table. Rejection degrades gracefully — the text is still
// extracted as prose — while accepting a phantom scrambles reading
// order.
let note = ["word"; 70].join(" ");
let cells = cells_of(&[
&["Purpose", note.as_str()],
&["Owner", "Facilities team"],
&["Date", "2024-06-01"],
&["Status", "Active"],
]);
assert!(has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_allows_description_column() {
// Long-ish description cells, but text is spread across the table
let desc = ["word"; 25].join(" ");
let cells = cells_of(&[
&["Item A", desc.as_str(), "100"],
&["Item B", desc.as_str(), "200"],
&["Item C", desc.as_str(), "300"],
&["Item D", desc.as_str(), "400"],
]);
assert!(!has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_allows_short_tables() {
let cells = cells_of(&[&["Name", "Value"], &["Total", "42"]]);
assert!(!has_dominant_prose_cell(&cells));
}
#[test]
fn dominant_prose_cell_allows_data_table_with_long_note() {
// A real 4+ row table with one verbose remark cell: the note is ≥60
// words but the table's other content carries more than 2× its word
// count, so concentration stays below the 1/3 threshold. The
// denominator scales with table size — this is what keeps large
// legitimate tables safe where a bare length cap would not.
let note = ["word"; 60].join(" ");
let row_text = ["data"; 12].join(" ");
let mut rows: Vec<Vec<String>> = (0..11)
.map(|i| {
vec![
format!("Item {i}"),
row_text.clone(),
format!("{}", i * 100),
]
})
.collect();
rows.push(vec!["Note".into(), note, String::new()]);
assert!(!has_dominant_prose_cell(&rows));
}
// --- rects_overlap ---
#[test]
@@ -3488,7 +3366,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
});
@@ -3799,7 +3676,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: crate::types::ItemType::Text,
mcid: None,
});
-1
View File
@@ -587,7 +587,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid,
}
-1
View File
@@ -109,7 +109,6 @@ pub(crate) fn try_split_financial_item(item: &TextItem) -> Option<Vec<TextItem>>
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
-3
View File
@@ -521,7 +521,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -887,7 +886,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
@@ -925,7 +923,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
-4
View File
@@ -236,7 +236,6 @@ fn split_merged_numbers(item: &TextItem, col_boundaries: &[f32]) -> Vec<TextItem
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
@@ -258,7 +257,6 @@ fn split_merged_numbers(item: &TextItem, col_boundaries: &[f32]) -> Vec<TextItem
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
is_strikeout: item.is_strikeout,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
@@ -1434,7 +1432,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1453,7 +1450,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
-520
View File
@@ -1,520 +0,0 @@
//! Text-quality detection: deciding when an extracted text layer is too broken
//! to serve and a page should fall back to OCR.
//!
//! Extraction can produce plausible-looking bytes that are actually garbage —
//! failed CID→Unicode mappings, broken ToUnicode CMaps, mojibake. These
//! detectors catch that and let callers set `needs_ocr`. They come in two
//! layers, sharing the same primitives:
//!
//! - **Markdown-level** ([`detect_encoding_issues`], [`is_garbage_text`],
//! [`is_cid_garbage`]) run on a page's final markdown string. Used as a
//! backstop on the region-extraction and whole-document paths.
//! - **Item/span-level** ([`analyze_text_quality`],
//! [`region_items_have_decoding_issue`]) run on individual `TextItem`s and
//! accumulate per-page evidence, so localized garbled spans on an otherwise
//! clean page are caught without a single span having to condemn the page.
//!
//! Detection classes, roughly by signal:
//! - **Replacement runs**: U+FFFD clusters ([`has_replacement_text_run`]).
//! - **Private-use / C1-control runs**: CID passthrough landing in PUA or the
//! C1 block ([`has_private_use_text_run`], [`has_cid_control_token`]).
//! - **Dollar-as-space**: `Word$Word$Word` from broken CMaps
//! ([`has_dollar_as_space_pattern`]).
//! - **Non-alphanumeric dominance**: symbol soup ([`is_garbage_text`]).
//! - **Substitution-cipher letter statistics**: pure-ASCII output whose letter
//! distribution is a permutation of natural language ([`CipherGarbleStats`]).
use crate::types::TextItem;
use crate::{add_ocr_reason, OCR_REASON_SUSPECTED_GARBLED_TEXT};
use std::collections::BTreeMap;
/// Detect broken font encodings in extracted markdown text.
///
/// Two heuristics:
/// 1. **U+FFFD**: Any replacement character indicates decode failures.
/// 2. **Dollar-as-space**: Pattern like `Word$Word$Word` where `$` is used as a
/// word separator due to broken ToUnicode CMaps. Triggers when either:
/// - More than 50% of `$` are between letters (clear substitution pattern), OR
/// - More than 20 letter-dollar-letter occurrences (even if some `$` are also
/// used as trailing/leading separators, 20+ is far beyond normal financial text).
pub(crate) fn detect_encoding_issues(markdown: &str) -> bool {
// Heuristic 1: U+FFFD replacement characters
if markdown.contains('\u{FFFD}') {
return true;
}
// Heuristic 2: dollar-as-space pattern
if has_dollar_as_space_pattern(markdown) {
return true;
}
// Heuristic 3: substitution-cipher letter statistics (broken ToUnicode)
let mut stats = CipherGarbleStats::default();
stats.add_text(markdown);
stats.looks_garbled()
}
fn has_dollar_as_space_pattern(markdown: &str) -> bool {
let total_dollars = markdown.matches('$').count();
if total_dollars > 10 {
let bytes = markdown.as_bytes();
let mut letter_dollar_letter = 0usize;
for i in 1..bytes.len().saturating_sub(1) {
if bytes[i] == b'$'
&& bytes[i - 1].is_ascii_alphabetic()
&& bytes[i + 1].is_ascii_alphabetic()
{
letter_dollar_letter += 1;
}
}
if letter_dollar_letter > 20 || letter_dollar_letter * 2 > total_dollars {
return true;
}
}
false
}
/// English letter frequencies (percent, az). Used as a natural-language
/// reference: every Latin-script language in the eval corpus (Swedish,
/// Finnish, Turkish, German, romaji) scores ≥ 0.80 cosine similarity against
/// it, while substitution-cipher text scores ~0.53.
const ENGLISH_LETTER_FREQ: [f64; 26] = [
8.2, 1.5, 2.8, 4.3, 12.7, 2.2, 2.0, 6.1, 7.0, 0.15, 0.8, 4.0, 2.4, 6.7, 7.5, 1.9, 0.1, 6.0,
6.3, 9.1, 2.8, 1.0, 2.4, 0.15, 2.0, 0.07,
];
/// Letter statistics for detecting substitution-cipher garbling: broken
/// ToUnicode CMaps that shift every character by a per-range constant (e.g.
/// `Certificate` extracted as `8VceZWZTReV`). Such text is 100% printable
/// ASCII with word-like token lengths, so it defeats `is_garbage_text` and
/// produces no replacement characters — it needs its own discriminator.
#[derive(Debug, Default)]
struct CipherGarbleStats {
/// Case-folded ASCII letter histogram.
letter_counts: [u32; 26],
ascii_letters: usize,
ascii_vowels: usize,
/// Accented Latin letters (Latin-1 Supplement through Latin Extended-B,
/// plus Latin Extended Additional). Count toward Latin dominance only.
latin_ext_letters: usize,
non_latin_letters: usize,
/// Adjacent ASCII-letter pairs, and how many of them switch from
/// lowercase straight to uppercase mid-word.
letter_bigrams: usize,
case_shift_bigrams: usize,
}
impl CipherGarbleStats {
fn add_text(&mut self, text: &str) {
let mut prev: Option<char> = None;
for ch in text.chars() {
if ch.is_ascii_alphabetic() {
let idx = (ch.to_ascii_lowercase() as u8 - b'a') as usize;
self.letter_counts[idx] += 1;
self.ascii_letters += 1;
if matches!(ch.to_ascii_lowercase(), 'a' | 'e' | 'i' | 'o' | 'u') {
self.ascii_vowels += 1;
}
if let Some(p) = prev {
self.letter_bigrams += 1;
if p.is_ascii_lowercase() && ch.is_ascii_uppercase() {
self.case_shift_bigrams += 1;
}
}
prev = Some(ch);
} else {
if ch.is_alphabetic() {
if matches!(ch as u32, 0xC0..=0x24F | 0x1E00..=0x1EFF) {
self.latin_ext_letters += 1;
} else {
self.non_latin_letters += 1;
}
}
prev = None;
}
}
}
/// Cosine similarity between the observed letter histogram and English
/// letter frequencies. A shifted alphabet permutes the histogram, which
/// destroys the similarity regardless of the shift amount.
fn english_cosine(&self) -> f64 {
if self.ascii_letters == 0 {
return 1.0;
}
let n = self.ascii_letters as f64;
let mut dot = 0.0;
let mut norm_obs = 0.0;
for (count, freq) in self.letter_counts.iter().zip(ENGLISH_LETTER_FREQ) {
let p = *count as f64 / n;
dot += p * freq;
norm_obs += p * p;
}
let norm_en = ENGLISH_LETTER_FREQ
.iter()
.map(|f| f * f)
.sum::<f64>()
.sqrt();
dot / (norm_obs.sqrt() * norm_en)
}
/// Cosine similarity between the observed histogram and English
/// frequencies after sorting BOTH descending — i.e. comparing the *shape*
/// of the frequency profile, ignoring which letter sits where. A
/// substitution cipher is a bijection, so it preserves this shape exactly
/// (att10k 0.97, arbitrary shifts 0.99) regardless of case or offset.
/// Non-linguistic ASCII has a different profile: a small alphabet is far
/// steeper (random DNA 0.74, hex dumps 0.81), so the shape diverges.
fn english_shape_cosine(&self) -> f64 {
if self.ascii_letters == 0 {
return 1.0;
}
let n = self.ascii_letters as f64;
let mut obs: [f64; 26] = std::array::from_fn(|i| self.letter_counts[i] as f64 / n);
obs.sort_unstable_by(|a, b| b.total_cmp(a));
let mut en = ENGLISH_LETTER_FREQ;
en.sort_unstable_by(|a, b| b.total_cmp(a));
let dot: f64 = obs.iter().zip(en).map(|(o, e)| o * e).sum();
let norm_obs = obs.iter().map(|o| o * o).sum::<f64>().sqrt();
let norm_en = en.iter().map(|e| e * e).sum::<f64>().sqrt();
dot / (norm_obs * norm_en)
}
/// Thresholds validated against the 380-document pdf-evals snapshot
/// corpus (0 false positives) and the garbled ParseBench `att10k` page
/// (vowel ratio 0.245, case-shift rate 0.225, cosine 0.532). Closest
/// legitimate document on each axis: vowel ratio 0.264 (circuit
/// schematic), case-shift rate 0.021, cosine 0.801.
fn looks_garbled(&self) -> bool {
// Need a statistically meaningful, Latin-dominant sample.
if self.ascii_letters < 200
|| self.non_latin_letters > self.ascii_letters + self.latin_ext_letters
{
return false;
}
// Real Latin-script text keeps vowels above ~30% of letters even in
// acronym- and part-number-heavy documents; shifted text starves them.
let vowel_ratio = self.ascii_vowels as f64 / self.ascii_letters as f64;
if vowel_ratio > 0.30 {
return false;
}
// Signal 1: lowercase→uppercase transitions inside words. A shifted
// lowercase alphabet straddles the ASCII uppercase block ('i'→'Z',
// 't'→'e'), so garbled words flip case constantly. Real documents
// stay ≤ 0.02 even with camelCase identifiers.
let case_shifts = self.letter_bigrams >= 100
&& self.case_shift_bigrams as f64 >= self.letter_bigrams as f64 * 0.10;
// Signal 2: the histogram is a permutation of natural language — an
// English-like frequency SHAPE (sorted cosine high) but with letters
// in the wrong POSITIONS (unsorted cosine low). This is the signature
// of a substitution cipher and is case-independent, so it catches
// all-lowercase and all-uppercase shifts as well as case-straddling
// ones. Genuinely non-linguistic ASCII that is merely "unlike English"
// fails one of the two halves: DNA/hex dumps have too steep a profile
// (shape cosine < 0.90), while protein sequences, ticker symbols and
// base64 are not sufficiently unlike English in position (unsorted
// cosine ≥ 0.60) — so none of them are routed to OCR.
let permuted_language = self.english_cosine() < 0.60 && self.english_shape_cosine() >= 0.90;
case_shifts || permuted_language
}
}
#[derive(Debug, Default)]
pub(crate) struct TextQualityReport {
pub(crate) pages_needing_ocr: Vec<u32>,
pub(crate) has_encoding_issues: bool,
pub(crate) reasons_by_page: BTreeMap<u32, Vec<String>>,
}
#[derive(Debug, Default)]
struct PageTextQualityEvidence {
chars: usize,
replacement_chars: usize,
replacement_spans: usize,
longest_replacement_run: usize,
cipher_garble: CipherGarbleStats,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum TextSpanIssueKind {
Replacement,
Strong,
}
pub(crate) fn analyze_text_quality(items: &[TextItem]) -> TextQualityReport {
let mut reasons_by_page = BTreeMap::new();
let mut evidence_by_page = BTreeMap::<u32, PageTextQualityEvidence>::new();
for item in items {
if !matches!(item.item_type, crate::types::ItemType::Text) {
continue;
}
let evidence = evidence_by_page.entry(item.page).or_default();
evidence.chars += item.text.chars().filter(|ch| !ch.is_whitespace()).count();
evidence.cipher_garble.add_text(&item.text);
match text_span_decoding_issue_kind(&item.text) {
Some(TextSpanIssueKind::Strong) => {
add_ocr_reason(
&mut reasons_by_page,
item.page,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
Some(TextSpanIssueKind::Replacement) => {
let stats = replacement_text_stats(&item.text);
evidence.replacement_chars += stats.0;
evidence.replacement_spans += 1;
evidence.longest_replacement_run = evidence.longest_replacement_run.max(stats.1);
}
None => {}
}
}
for (page, evidence) in evidence_by_page {
if reasons_by_page.contains_key(&page) {
continue;
}
if page_replacement_evidence_needs_ocr(&evidence) || evidence.cipher_garble.looks_garbled()
{
add_ocr_reason(
&mut reasons_by_page,
page,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
}
let pages_needing_ocr: Vec<u32> = reasons_by_page.keys().copied().collect();
TextQualityReport {
has_encoding_issues: !pages_needing_ocr.is_empty(),
pages_needing_ocr,
reasons_by_page,
}
}
pub(crate) fn region_items_have_decoding_issue(items: &[TextItem]) -> bool {
items.iter().any(|item| {
matches!(item.item_type, crate::types::ItemType::Text)
&& text_span_has_decoding_issue(&item.text)
})
}
fn text_span_has_decoding_issue(text: &str) -> bool {
text_span_decoding_issue_kind(text).is_some()
}
fn text_span_decoding_issue_kind(text: &str) -> Option<TextSpanIssueKind> {
let text = text.trim();
if text.is_empty() {
return None;
}
if has_dollar_as_space_pattern(text)
|| has_private_use_text_run(text)
|| is_cid_garbage(text)
|| has_cid_control_token(text)
{
return Some(TextSpanIssueKind::Strong);
}
if has_replacement_text_run(text) {
return Some(TextSpanIssueKind::Replacement);
}
None
}
fn replacement_text_stats(text: &str) -> (usize, usize) {
let mut replacement = 0usize;
let mut current_run = 0usize;
let mut longest_run = 0usize;
for ch in text.chars() {
if ch == '\u{FFFD}' {
replacement += 1;
current_run += 1;
longest_run = longest_run.max(current_run);
} else {
current_run = 0;
}
}
(replacement, longest_run)
}
fn page_replacement_evidence_needs_ocr(evidence: &PageTextQualityEvidence) -> bool {
if evidence.replacement_chars == 0 || evidence.chars == 0 {
return false;
}
// If the entire page is only a short broken text layer, even a short
// replacement run is enough evidence. On otherwise text-heavy pages,
// require density so math formulas do not force full-page OCR.
if evidence.chars <= 80 && evidence.longest_replacement_run >= 2 {
return true;
}
let replacement_density_bps = evidence.replacement_chars * 10_000 / evidence.chars;
let enough_bad_text = evidence.replacement_chars >= 12 && replacement_density_bps >= 500;
let repeated_bad_spans = evidence.replacement_spans >= 3 && replacement_density_bps >= 250;
let long_bad_run = evidence.longest_replacement_run >= 8 && replacement_density_bps >= 250;
enough_bad_text || repeated_bad_spans || long_bad_run
}
fn has_replacement_text_run(text: &str) -> bool {
let (replacement, longest_run) = replacement_text_stats(text);
longest_run >= 2 || replacement >= 3
}
fn has_private_use_text_run(text: &str) -> bool {
let mut total = 0usize;
let mut private_use = 0usize;
let mut current_run = 0usize;
let mut longest_run = 0usize;
for ch in text.chars() {
if ch.is_whitespace() {
current_run = 0;
continue;
}
total += 1;
if is_private_use_char(ch) {
private_use += 1;
current_run += 1;
longest_run = longest_run.max(current_run);
} else {
current_run = 0;
}
}
if private_use == 0 {
return false;
}
longest_run >= 3 || (total >= 5 && private_use >= 2 && private_use * 2 >= total)
}
fn has_cid_control_token(text: &str) -> bool {
text.split_whitespace().any(token_has_cid_control)
}
fn token_has_cid_control(token: &str) -> bool {
let mut total = 0usize;
let mut c1_control = 0usize;
for ch in token.chars() {
total += 1;
if ('\u{0080}'..='\u{009F}').contains(&ch) {
c1_control += 1;
}
}
total >= 5 && c1_control >= 2 && c1_control * 20 >= total
}
fn is_private_use_char(ch: char) -> bool {
matches!(
ch as u32,
0xE000..=0xF8FF | 0xF0000..=0xFFFFD | 0x100000..=0x10FFFD
)
}
/// Check if extracted text is predominantly garbage (non-alphanumeric).
///
/// Broken font encodings produce text like "----1-.-.-.___ --.-. .._ I_---."
/// where most characters are punctuation/symbols. Real text in any language
/// has >50% alphanumeric characters.
pub(crate) fn is_garbage_text(markdown: &str) -> bool {
let mut alphanum = 0usize;
let mut non_alphanum = 0usize;
let chars: Vec<char> = markdown.chars().collect();
let mut i = 0usize;
while i < chars.len() {
let ch = chars[i];
let mut run_end = i + 1;
while run_end < chars.len() && chars[run_end] == ch {
run_end += 1;
}
let is_decorative_leader = matches!(ch, '.' | '_' | '·') && run_end - i >= 3;
if !is_decorative_leader {
for &run_ch in &chars[i..run_end] {
if run_ch.is_whitespace() {
continue;
}
// Skip markdown syntax chars that we add (not from the PDF)
if matches!(run_ch, '#' | '*' | '|' | '-' | '\n') {
continue;
}
if run_ch.is_alphanumeric() {
alphanum += 1;
} else {
non_alphanum += 1;
}
}
}
i = run_end;
}
let total = alphanum + non_alphanum;
total >= 50 && alphanum * 2 < total
}
/// Detect garbage from failed CID-to-Unicode mapping on Identity-H fonts.
///
/// When CID values don't correspond to Unicode codepoints, the raw bytes often
/// produce characters in the C1 control range (U+0080U+009F) or Private Use
/// Area, mixed with random Latin Extended characters. Valid text in any
/// language almost never contains C1 controls. We also fall back to the
/// general `is_garbage_text` check for non-alphanumeric-heavy patterns.
pub(crate) fn is_cid_garbage(text: &str) -> bool {
if is_garbage_text(text) {
return true;
}
let mut total = 0usize;
let mut c1_control = 0usize;
let mut high_latin = 0usize;
for ch in text.chars() {
if ch.is_whitespace() {
continue;
}
total += 1;
// C1 control characters (U+0080U+009F) — almost never in real text
if ch == '·' {
continue;
}
if ('\u{0080}'..='\u{009F}').contains(&ch) {
c1_control += 1;
}
// High Latin-1 (U+00A0U+00FF) — legitimate in Western European text
// but when combined with ASCII in CID passthrough, indicates mojibake
// from CID values being misinterpreted as Latin-1 characters.
if ('\u{00A0}'..='\u{00FF}').contains(&ch) {
high_latin += 1;
}
}
if total < 5 {
return false;
}
// If ≥5% of non-whitespace chars are C1 controls, it's garbage
if c1_control >= 2 && c1_control * 20 >= total {
return true;
}
// If ≥40% of non-whitespace chars are high Latin-1 AND the text has few
// ASCII letters, it's likely CID-as-Latin-1 mojibake (Japanese/CJK PDFs
// where CID values 0x80-0xFF become accented Latin characters). Keep a
// minimum length so short math tokens like "2×()×" do not route a clean
// page to OCR.
let ascii_letters = text.chars().filter(|c| c.is_ascii_alphabetic()).count();
total >= 20 && high_latin * 5 >= total * 2 && ascii_letters * 3 < total
}
-17
View File
@@ -93,9 +93,6 @@ pub fn is_bold_font(font_name: &str) -> bool {
|| lower.contains("extrabold")
|| lower.contains("ultrabold")
|| lower.contains("medium") && !lower.contains("mediumitalic") // Some fonts use Medium for semi-bold
// URW Type 1 fonts abbreviate Medium as "Medi" (e.g. NimbusRomNo9L-Medi,
// the Times-Bold substitute in LaTeX documents; -MediItal is bold italic).
|| lower.contains("-medi") && !lower.contains("mediumital")
}
/// Detect if a font name indicates italic/oblique style
@@ -765,17 +762,6 @@ mod tests {
use super::*;
use crate::types::ItemType;
#[test]
fn bold_font_urw_medi_abbreviation() {
// URW Type 1 fonts (LaTeX default Times) abbreviate Medium as "Medi"
assert!(is_bold_font("NROFIU+NimbusRomNo9L-Medi"));
assert!(is_bold_font("NimbusRomNo9L-MediItal"));
assert!(!is_bold_font("DSSZWN+NimbusRomNo9L-Regu"));
assert!(!is_bold_font("NimbusRomNo9L-ReguItal"));
// Medium-Italic exclusion still holds
assert!(!is_bold_font("Foo-MediumItalic"));
}
#[test]
fn strip_soft_hyphen() {
assert_eq!(expand_ligatures("con\u{00AD}tent"), "content");
@@ -898,7 +884,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1019,7 +1004,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -1097,7 +1081,6 @@ mod tests {
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
-4
View File
@@ -120,10 +120,6 @@ pub struct TextItem {
/// baseline — PDFs have no underline font flag, so this is detected
/// geometrically after extraction; see `extractor::underline`).
pub is_underline: bool,
/// Whether the text is struck out (drawn rule/thin rect crossing the
/// glyphs at mid x-height). Same geometric detection as underline,
/// different vertical window; see `extractor::underline`.
pub is_strikeout: bool,
/// Type of item (text, image, link)
pub item_type: ItemType,
/// Marked Content ID from the content stream's BDC/BMC operator.
Binary file not shown.
+36 -71
View File
@@ -9,7 +9,7 @@ use pdf_inspector::{
extract_pages_markdown_mem, extract_tables_in_regions_mem, extract_text,
extract_text_in_regions_mem, extract_text_with_positions, extract_text_with_positions_mem,
process_pdf_mem, process_pdf_with_options, to_markdown, MarkdownOptions, PdfError, PdfOptions,
PdfType, TextItem,
PdfType, TextItem, OCR_REASON_SUSPECTED_GARBLED_TEXT,
};
use std::collections::HashSet;
@@ -105,7 +105,6 @@ fn make_text_item(text: &str, x: f32, y: f32, font_size: f32, page: u32) -> Text
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -132,7 +131,6 @@ fn make_text_item_with_font(
is_bold: is_bold_font(font),
is_italic: is_italic_font(font),
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1358,32 +1356,6 @@ fn test_extract_regions_mem_identity_h_needs_ocr() {
);
}
/// ParseBench `text_simple__att10k.pdf` (issue #118): the producer authored a
/// broken ToUnicode CMap that shifts every character by a per-range constant,
/// and the embedded subset font has no `cmap` table to recover from. The
/// resulting ciphertext is 100% printable ASCII, so it must be caught by the
/// substitution-cipher statistics and routed to OCR instead of served silently.
#[test]
fn test_extract_pages_mem_shifted_cipher_tounicode_needs_ocr() {
let buf = std::fs::read("tests/fixtures/shifted_cipher_tounicode.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, None).unwrap();
assert_eq!(result.pages.len(), 1);
assert!(
result.pages[0].needs_ocr,
"shifted-cipher garbled page should be flagged needs_ocr"
);
assert!(
result.pages[0].markdown.is_empty(),
"garbled markdown should be suppressed"
);
assert_eq!(result.pages_needing_ocr, vec![1]);
assert_eq!(
result.pages[0].ocr_reason.as_deref(),
Some("suspected_garbled_text")
);
}
#[test]
fn test_extract_regions_mem_multiple_regions_per_page() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
@@ -2970,6 +2942,41 @@ fn test_extract_pages_markdown_gid_pages_need_ocr() {
assert!(result.pages_needing_ocr.contains(&1)); // 1-indexed
}
#[test]
fn test_extract_pages_markdown_flags_printable_ascii_mojibake() {
let buf = std::fs::read("tests/fixtures/text_simple__att10k.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&[0])).unwrap();
assert_eq!(result.pages.len(), 1);
assert!(result.pages[0].needs_ocr);
assert_eq!(
result.pages[0].ocr_reason.as_deref(),
Some(OCR_REASON_SUSPECTED_GARBLED_TEXT)
);
assert!(
result.pages[0].markdown.is_empty(),
"garbled direct text should be suppressed when OCR is needed"
);
assert_eq!(result.pages_needing_ocr, vec![1]);
assert_eq!(result.ocr_reasons_by_page.len(), 1);
assert_eq!(result.ocr_reasons_by_page[0].page, 1);
assert_eq!(
result.ocr_reasons_by_page[0].reasons,
vec![OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()]
);
let full = process_pdf_mem(&buf).unwrap();
assert_eq!(full.pages_needing_ocr, vec![1]);
assert!(full.has_encoding_issues);
assert_eq!(full.ocr_reasons_by_page.len(), 1);
assert_eq!(full.ocr_reasons_by_page[0].page, 1);
assert_eq!(
full.ocr_reasons_by_page[0].reasons,
vec![OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()]
);
}
#[test]
fn test_extract_pages_markdown_classification_with_tables() {
// nexo-price-en.pdf is known to have tables
@@ -3606,45 +3613,3 @@ fn test_markdown_options_default_has_include_images_false() {
let opts = MarkdownOptions::default();
assert!(!opts.include_images);
}
#[test]
fn encrypted_pdf_decrypts_with_correct_password() {
let path = "tests/fixtures/encrypted-secret123.pdf";
// No password: the file is encrypted and can't be read.
let no_pw = process_pdf_with_options(path, PdfOptions::new());
assert!(
matches!(no_pw, Err(PdfError::Encrypted)),
"expected Encrypted without a password, got {no_pw:?}"
);
// Wrong password: still rejected.
let wrong = process_pdf_with_options(path, PdfOptions::new().password("wrong"));
assert!(
matches!(wrong, Err(PdfError::Encrypted)),
"expected Encrypted with a wrong password, got {wrong:?}"
);
// Correct password: decrypts and extracts real content.
let ok = process_pdf_with_options(path, PdfOptions::new().password("secret123"))
.expect("correct password should decrypt");
let md = ok.markdown.unwrap_or_default();
// Assert a stable fixture token so a garbled-but-long extraction (the
// encrypted-stream regression this guards) still fails the test.
assert!(
md.contains("Procurement"),
"decrypted markdown should contain the fixture's real text, got {} chars",
md.len()
);
}
#[test]
fn pdf_options_debug_redacts_password() {
let opts = PdfOptions::new().password("secret123");
let dbg = format!("{opts:?}");
assert!(
!dbg.contains("secret123"),
"password leaked in Debug: {dbg}"
);
assert!(dbg.contains("REDACTED"), "expected redaction marker: {dbg}");
}
+3 -2
View File
@@ -48,7 +48,7 @@ tips of directly from customers received other employees paid tips recd. entr
**Page 3**
27 28 29 30 31 **Subtotals from pages** **1, 2, and 3** **Totals**
27 28 29 30 31 **Subtotals** **from pages** **1, 2, and 3** **Totals**
**1.** Report total cash tips (col. **a**) on Form 4070, line **1.**
**2.** Report total credit card tips (col. **b**) on Form 4070, line **2.**
@@ -76,6 +76,7 @@ forms simpler, we would be happy to hear from you. You can write to the Tax Form
**Unreported Tips.—**If you received tips of $20 or more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you **must** use Form 1040 and **Form 4137,** Social Security and Medicare Tax on Unreported Tip Income, to report them. You may **not** use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act **cannot** use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—**Get **Pub. 531,** Reporting Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—**If you do not keep a daily record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
### Instructions (continued)
**Instructions** *(continued)*
Use this space to total your tips for the year