Compare commits

...
Author SHA1 Message Date
Abimael Martell e16808c0c4 fix(tables): reject sparse prose row-stripe tables 2026-07-08 22:58:07 -07:00
Abimael Martell 6eb874a69f fix(tables): trim spaces inside parenthetical cell fragments 2026-07-08 13:12:50 -07:00
Abimael MartellandCursor e96f408d14 fix(markdown): strip stray spaces before sentence punctuation (review)
Style-boundary item splits can strand a trailing period in its own
fragment, and multiple assembly paths join fragments with spaces,
yielding "word ." artifacts. Rather than chasing every join site, a
postprocess pass removes a space before `.`/`,`/`;` when the mark ends
its token (whitespace, cell boundary `|`, or end of text follows).
Dot leaders/ellipses and mid-token periods are untouched.

Fixes the td9264 "companies ." artifacts and two pre-existing
"armoring ," artifacts in the 2013-app2 snapshot. pdf-evals: zero
markdown diffs across all 203 corpus PDFs vs committed baselines.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 12:36:44 -07:00
Abimael MartellandCursor af25252640 fix(extractor): break merges at underline boundaries too (review)
OR-merging underline stretched the eventual <u> span over neighboring
plain fragments. Merge runs now break on any style-flag change, the
redundant accumulator is gone, and format_list_item learned to move
bullet markers outside <u> wrappers so fully-underlined bullet lines
still render as markdown lists. td9264 snapshot regenerated — spans are
tighter (trailing periods correctly outside the tag).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 12:02:51 -07:00
Abimael MartellandCursor eebc105211 feat(markdown): underline emission, Unicode scripts, style-preserving merges (ENG-5015 2b)
Three formatting losses in the direct-extraction markdown path:

1. text_with_formatting gains <u> run emission (detect_underline option,
   default on) using the geometric is_underline flag from 1.9.9.
   Underline runs stay free of nested bold/italic markers — consumers
   match tag content literally. Heading lines keep plain text for
   bold/italic but preserve <u>: the tag carries meaning `#` doesn't.
2. merge_subscript_items now maps absorbed digit scripts to Unicode
   sub/superscript forms with direction from the baseline offset
   ("H"+"2" -> "H₂", "word"+raised "2" -> "word²", "m"+"3" -> "m³").
   NFKC/NFKD folds these back to plain digits so text matching
   downstream is unaffected; renderers keep the script semantics.
3. merge_text_items no longer merges across bold/italic boundaries —
   absorbing a styled run into a plain neighbor erased the styling
   before markdown emission ever saw it. On eval docs this recovers
   20-82 italic runs per document that previously emitted as plain.

Snapshots regenerated (diffs are the features: CCl₂F₂, m³, underlined
legal section headings, finer bold runs). pdf-evals regression suite:
202/202 real PDFs pass. napi 1.9.9 -> 1.9.10.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 11:43:48 -07:00
Abimael MartellandCursor 422a2ff118 feat(extractor): geometric underline detection on TextItem (#116)
* feat(extractor): geometric underline detection on TextItem (ENG-5015)

PDFs carry no underline font flag — underlines are stroked horizontal
lines or thin filled rects drawn under the baseline. Correlate those
graphics (already parsed from the content stream) with text items in a
post-pass: a rule within ~0.35em below the baseline covering >=60% of
an item's width marks is_underline.

Exposed through the napi and python bindings. Verified on real docs:
4/4 underlined sentences flagged on a Japanese report, links/headings
flagged on 8 of 10 underline-bearing eval docs, zero flags on docs
without underlines. Known FP source (table cell borders) documented —
downstream applies inline styling only to plain-text regions.

napi 1.9.8 -> 1.9.9.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): underline rules only from painted rects, normalized extents (review)

Two review fixes: (1) normalize rect extents before the thickness/width
checks — `re` operands pass through the CTM so width/height can be
negative, which missed negative-width rules and let negative-height
bands pass as thin; (2) only feed painted rects to underline detection —
`re` rects now wait in a pending list until a paint operator (S/s, f/F/
f*, B/B*/b/b*) confirms them, and `re W n` clip-only paths are discarded
at `n`, so invisible clip boundaries no longer underline nearby text.
Marking moved into content_stream where paint state lives (pre-rotation,
consistent device space).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): harden underline detection

* feat(cli): export positioned text item json

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 10:20:18 -07:00
Abimael Martell 30eddade77 fix(extractor): preserve tagged overlapping text order (#114) 2026-06-24 10:26:52 -07:00
Abimael Martell 1b2e2c76d6 fix(extractor): make trace previews unicode-safe (#113) 2026-06-24 01:27:27 -06:00
Abimael Martell ce49794719 fix(extractor): reduce garbled OCR false positives (#112)
* fix(extractor): reduce garbled OCR false positives

* fix(extractor): tighten garbled text OCR routing

* fix(extractor): decode UTF-16 ToUnicode destinations

* fix(extractor): narrow ToUnicode destination cleanup

* fix(extractor): decode Aptos private ff ligature
2026-06-24 01:13:16 -06:00
Abimael Martell 1a5ba6f1e9 feat(api): expose OCR reason signal (#110) 2026-06-23 15:51:55 -07:00
Abimael Martell 57b98c6a5d fix(extractor): flag garbled text spans for OCR (#108)
* fix(extractor): flag garbled text spans for OCR

* fix(extractor): apply text quality checks to regions

* chore(napi): bump npm package version
2026-06-23 13:12:10 -07:00
Abimael Martell f25808e0a7 fix(extractor): restore CID font state for Chinese text (#106)
* fix Chinese CID text decoding

* bump package versions
2026-06-20 19:43:27 -06:00
Abimael Martell 9360c8464d ci: add crates trusted publishing (#103) 2026-06-05 11:21:55 -07:00
Abimael Martell 252d87ac58 docs(readme): add package badges and install docs (#102)
* docs: add crates.io install instructions

* docs: add npm badge
2026-06-05 11:01:56 -07:00
Abimael Martell 85890648c9 chore: use crates.io lopdf (#101) 2026-06-05 10:50:05 -07:00
Abimael Martell 6e55e38b55 fix(markdown): handle wrapped bold abstracts (#100)
* fix(markdown): handle wrapped bold abstracts

* chore: bump napi package version
2026-06-01 14:59:18 -07:00
Abimael Martell 42befcea57 fix(pdf-inspector): recover wrapped key-value tables (#99)
* fix(pdf-inspector): recover wrapped key-value tables

* fix(pdf-inspector): satisfy clippy
2026-06-01 10:10:55 -07:00
41 changed files with 3816 additions and 290 deletions
+87
View File
@@ -0,0 +1,87 @@
name: Publish Rust crate
on:
push:
branches: [main]
paths: ['Cargo.toml']
permissions:
contents: read
env:
CARGO_TERM_COLOR: always
jobs:
check-version:
name: Check version change
runs-on: ubuntu-latest
outputs:
changed: ${{ steps.check.outputs.changed }}
published: ${{ steps.check.outputs.published }}
version: ${{ steps.check.outputs.version }}
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 2
- name: Check if version changed
id: check
run: |
NEW_VERSION=$(python3 -c 'import pathlib, tomllib; print(tomllib.loads(pathlib.Path("Cargo.toml").read_text())["package"]["version"])')
OLD_VERSION=$(git show HEAD~1:Cargo.toml | python3 -c 'import sys, tomllib; print(tomllib.loads(sys.stdin.read())["package"]["version"])')
echo "old=$OLD_VERSION new=$NEW_VERSION"
echo "version=$NEW_VERSION" >> "$GITHUB_OUTPUT"
if [ "$NEW_VERSION" = "$OLD_VERSION" ]; then
echo "changed=false" >> "$GITHUB_OUTPUT"
echo "published=false" >> "$GITHUB_OUTPUT"
exit 0
fi
echo "changed=true" >> "$GITHUB_OUTPUT"
HTTP_STATUS=$(curl --silent --show-error --output /tmp/crate-version.json --write-out "%{http_code}" \
-H "User-Agent: firecrawl/pdf-inspector publish workflow (https://github.com/firecrawl/pdf-inspector)" \
"https://crates.io/api/v1/crates/pdf-inspector/$NEW_VERSION")
case "$HTTP_STATUS" in
200)
echo "published=true" >> "$GITHUB_OUTPUT"
echo "pdf-inspector v$NEW_VERSION is already published"
;;
404)
echo "published=false" >> "$GITHUB_OUTPUT"
;;
*)
cat /tmp/crate-version.json
echo "Unexpected crates.io response: $HTTP_STATUS" >&2
exit 1
;;
esac
publish:
name: Publish to crates.io
needs: check-version
if: needs.check-version.outputs.changed == 'true' && needs.check-version.outputs.published == 'false'
runs-on: ubuntu-latest
environment: crates-io
permissions:
contents: read
id-token: write
steps:
- uses: actions/checkout@v4
- name: Install Rust
uses: dtolnay/rust-toolchain@stable
- name: Verify package
run: cargo publish --dry-run
- name: Authenticate with crates.io
id: auth
uses: rust-lang/crates-io-auth-action@v1
- name: Publish crate
run: cargo publish
env:
CARGO_REGISTRY_TOKEN: ${{ steps.auth.outputs.token }}
+2 -2
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector"
version = "0.1.0"
version = "0.1.3"
edition = "2021"
autobins = false
authors = ["Firecrawl Team"]
@@ -17,7 +17,7 @@ crate-type = ["lib", "cdylib"]
pyo3 = { version = "0.25", features = ["extension-module"], optional = true }
# PDF parsing
lopdf = { git = "https://github.com/J-F-Liu/lopdf", rev = "7a05512d831415b1f2b1ce522391d6beab8a1284", features = ["rayon"] }
lopdf = { version = "0.41.0", features = ["rayon"] }
# Error handling
thiserror = "2.0"
+28 -9
View File
@@ -1,5 +1,8 @@
# pdf-inspector
[![Crates.io](https://img.shields.io/crates/v/pdf-inspector.svg)](https://crates.io/crates/pdf-inspector)
[![npm](https://img.shields.io/npm/v/@firecrawl/pdf-inspector.svg)](https://www.npmjs.com/package/@firecrawl/pdf-inspector)
Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for [Python](docs/python.md) and [Node.js](napi/README.md).
Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.
@@ -71,9 +74,17 @@ console.log(result.markdown); // Markdown string or null
### Rust
Install from [crates.io](https://crates.io/crates/pdf-inspector):
```bash
cargo add pdf-inspector
```
Or add it manually:
```toml
[dependencies]
pdf-inspector = { git = "https://github.com/firecrawl/pdf-inspector" }
pdf-inspector = "0.1"
```
```rust
@@ -91,29 +102,37 @@ if let Some(markdown) = &result.markdown {
### CLI
```bash
# Install the CLI tools
cargo install pdf-inspector
# Convert PDF to Markdown
cargo run --bin pdf2md -- document.pdf
pdf2md document.pdf
# JSON output (for piping)
cargo run --bin pdf2md -- document.pdf --json
pdf2md document.pdf --json
# Positioned TextItem JSON, including is_underline metadata
pdf2md document.pdf --items-json
# Raw markdown only (no headers)
cargo run --bin pdf2md -- document.pdf --raw
pdf2md document.pdf --raw
# Insert page break markers (<!-- Page N -->)
cargo run --bin pdf2md -- document.pdf --pages
pdf2md document.pdf --pages
# Process only specific pages
cargo run --bin pdf2md -- document.pdf --select-pages 1,3,5-10
pdf2md document.pdf --select-pages 1,3,5-10
# Detection only (no extraction)
cargo run --bin detect-pdf -- document.pdf
cargo run --bin detect-pdf -- document.pdf --json
detect-pdf document.pdf
detect-pdf document.pdf --json
# Detection + layout analysis (tables, columns)
cargo run --bin detect-pdf -- document.pdf --analyze --json
detect-pdf document.pdf --analyze --json
```
From a source checkout, use `cargo run --bin pdf2md -- document.pdf` or `cargo run --bin detect-pdf -- document.pdf` instead.
## Architecture
```
+21
View File
@@ -0,0 +1,21 @@
# Publishing
The Rust crate is published to [crates.io](https://crates.io/crates/pdf-inspector) with trusted publishing from GitHub Actions. The first release was published manually; future releases publish from `.github/workflows/publish-crate.yml` when a `Cargo.toml` version change lands on `main`.
## crates.io Trusted Publisher
Configure the trusted publisher for the `pdf-inspector` crate with:
- Repository: `firecrawl/pdf-inspector`
- Workflow: `publish-crate.yml`
- Environment: `crates-io`
The workflow uses `rust-lang/crates-io-auth-action@v1` to exchange GitHub's OIDC token for a short-lived crates.io token, then passes it to `cargo publish`.
## Release Steps
1. Update `version` in `Cargo.toml`.
2. Merge the version bump to `main`.
3. The publish workflow compares the new `Cargo.toml` version with `HEAD~1`, runs `cargo publish --dry-run`, then publishes if that version is not already on crates.io.
If `Cargo.toml` changes without a package version bump, the workflow exits without publishing.
+5 -4
View File
@@ -672,8 +672,9 @@ checksum = "5e5032e24019045c762d3c0f28f5b6b8bbf38563a65908389bf7978758920897"
[[package]]
name = "lopdf"
version = "0.40.0"
source = "git+https://github.com/J-F-Liu/lopdf?rev=7a05512d831415b1f2b1ce522391d6beab8a1284#7a05512d831415b1f2b1ce522391d6beab8a1284"
version = "0.41.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "67513274c50a2b51e5f75d9e682fcf4ab064a8a9c9ae2c3c59309084882bb24d"
dependencies = [
"aes",
"bitflags",
@@ -829,7 +830,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
[[package]]
name = "pdf-inspector"
version = "0.1.0"
version = "0.1.3"
dependencies = [
"env_logger",
"log",
@@ -844,7 +845,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-napi"
version = "0.2.0"
version = "0.2.2"
dependencies = [
"napi",
"napi-build",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-napi"
version = "0.2.0"
version = "0.2.2"
edition = "2021"
[lib]
+2 -1
View File
@@ -37,7 +37,7 @@ console.log(result.confidence) // 0.875
Extract text within bounding-box regions from a PDF. Designed for hybrid OCR pipelines where a layout model detects regions in rendered page images, and this function extracts text from the PDF structure for text-based pages — skipping GPU OCR.
Each region result includes a `needsOcr` flag that signals unreliable extraction (empty text, GID-encoded fonts, garbage text, encoding issues).
Each region result includes a `needsOcr` flag that signals unreliable extraction (empty text, GID-encoded fonts, garbage text, encoding issues). When the cause is a suspected garbled text layer, `ocrReason` is set to `"suspected_garbled_text"`.
```typescript
import { extractTextInRegions } from '@firecrawl/pdf-inspector'
@@ -84,6 +84,7 @@ interface PageRegionTexts {
interface RegionText {
text: string
needsOcr: boolean // true when text is unreliable
ocrReason?: string // "suspected_garbled_text" when known
}
```
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.9.3",
"version": "1.9.10",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
+35
View File
@@ -40,6 +40,8 @@ pub struct PdfResult {
pub processing_time_ms: u32,
/// 1-indexed page numbers that need OCR.
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
pub ocr_reasons_by_page: Vec<PageOcrReasons>,
pub title: Option<String>,
pub confidence: f64,
pub is_complex_layout: bool,
@@ -48,6 +50,13 @@ pub struct PdfResult {
pub has_encoding_issues: bool,
}
/// OCR reasons for a single 1-indexed page.
#[napi(object)]
pub struct PageOcrReasons {
pub page: u32,
pub reasons: Vec<String>,
}
/// Lightweight PDF classification result.
#[napi(object)]
pub struct PdfClassification {
@@ -71,6 +80,9 @@ pub struct TextItem {
pub page: u32,
pub is_bold: bool,
pub is_italic: bool,
/// Underline detected geometrically (drawn rule/thin rect under the
/// baseline) — PDFs carry no underline font flag.
pub is_underline: bool,
pub item_type: ItemType,
/// URL for link items, `None` for other types.
pub link_url: Option<String>,
@@ -90,6 +102,8 @@ pub struct RegionText {
pub text: String,
/// `true` when the text should not be trusted (empty, GID fonts, garbage, encoding issues).
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
pub ocr_reason: Option<String>,
}
/// Extracted text for one page's regions.
@@ -126,6 +140,7 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
page_count: r.page_count,
processing_time_ms: r.processing_time_ms as u32,
pages_needing_ocr: r.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(r.ocr_reasons_by_page),
title: r.title,
confidence: r.confidence as f64,
is_complex_layout: r.layout.is_complex,
@@ -135,6 +150,18 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
}
}
fn to_napi_page_ocr_reasons(
reasons: Vec<pdf_inspector::PageOcrReasons>,
) -> Vec<PageOcrReasons> {
reasons
.into_iter()
.map(|reason| PageOcrReasons {
page: reason.page,
reasons: reason.reasons,
})
.collect()
}
fn convert_item_type(t: &pdf_inspector::types::ItemType) -> (ItemType, Option<String>) {
match t {
pdf_inspector::types::ItemType::Text => (ItemType::Text, None),
@@ -266,6 +293,7 @@ pub fn extract_text_with_positions(
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
item_type,
link_url,
}
@@ -563,6 +591,8 @@ pub struct PageMarkdownResult {
pub markdown: String,
/// `true` when text on this page is unreliable.
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
pub ocr_reason: Option<String>,
}
/// Combined per-page markdown extraction and layout classification result.
@@ -576,6 +606,8 @@ pub struct PagesExtractionResult {
pub pages_with_columns: Vec<u32>,
/// 1-indexed pages that need OCR (scanned/image-based).
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
pub ocr_reasons_by_page: Vec<PageOcrReasons>,
/// True if any page has tables or columns.
pub is_complex: bool,
}
@@ -607,11 +639,13 @@ pub fn extract_pages_markdown(
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
is_complex: result.is_complex,
})
})
@@ -648,6 +682,7 @@ fn to_page_region_texts(results: Vec<pdf_inspector::PageRegionResult>) -> Vec<Pa
.map(|r| RegionText {
text: r.text,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
})
+1
View File
@@ -38,6 +38,7 @@ class TextItem:
page: int
is_bold: bool
is_italic: bool
is_underline: bool
item_type: str
class RegionText:
+127 -3
View File
@@ -1,6 +1,10 @@
//! CLI tool for PDF to Markdown conversion
use pdf_inspector::{process_pdf_with_options, LayoutComplexity, PdfOptions, PdfType, ProcessMode};
use pdf_inspector::extractor::ItemType;
use pdf_inspector::{
extract_text_with_positions_pages, process_pdf_with_options, LayoutComplexity, PdfOptions,
PdfType, ProcessMode, TextItem,
};
use std::collections::HashSet;
use std::env;
use std::fmt::Write;
@@ -31,6 +35,108 @@ fn json_escape(s: &str) -> String {
out
}
fn format_ocr_reasons_by_page(reasons: &[pdf_inspector::PageOcrReasons]) -> String {
reasons
.iter()
.map(|entry| {
let reasons_json = entry
.reasons
.iter()
.map(|reason| format!(r#""{}""#, json_escape(reason)))
.collect::<Vec<_>>()
.join(",");
format!(r#"{{"page":{},"reasons":[{}]}}"#, entry.page, reasons_json)
})
.collect::<Vec<_>>()
.join(",")
}
fn item_type_label(item_type: &ItemType) -> &'static str {
match item_type {
ItemType::Text => "text",
ItemType::Image => "image",
ItemType::Link(_) => "link",
ItemType::FormField => "form_field",
}
}
fn format_items_json(items: &[TextItem]) -> String {
let underlined_count = items.iter().filter(|item| item.is_underline).count();
let items_json = items
.iter()
.map(|item| {
let mcid = item
.mcid
.map(|value| value.to_string())
.unwrap_or_else(|| "null".to_string());
let link_url = match &item.item_type {
ItemType::Link(url) => format!(r#","url":"{}""#, json_escape(url)),
_ => String::new(),
};
format!(
r#"{{"text":"{}","page":{},"x":{:.2},"y":{:.2},"width":{:.2},"height":{:.2},"font":"{}","font_size":{:.2},"is_bold":{},"is_italic":{},"is_underline":{},"item_type":"{}","mcid":{}{}}}"#,
json_escape(&item.text),
item.page,
item.x,
item.y,
item.width,
item.height,
json_escape(&item.font),
item.font_size,
item.is_bold,
item.is_italic,
item.is_underline,
item_type_label(&item.item_type),
mcid,
link_url,
)
})
.collect::<Vec<_>>()
.join(",");
format!(
r#"{{"total_items":{},"underlined_count":{},"items":[{}]}}"#,
items.len(),
underlined_count,
items_json
)
}
#[cfg(test)]
mod tests {
use super::format_items_json;
use pdf_inspector::extractor::ItemType;
use pdf_inspector::TextItem;
#[test]
fn items_json_includes_position_and_underline_metadata() {
let items = vec![TextItem {
text: "A \"quoted\" item".to_string(),
x: 12.345,
y: 67.891,
width: 23.456,
height: 9.876,
font: "F1".to_string(),
font_size: 10.0,
page: 2,
is_bold: false,
is_italic: true,
is_underline: true,
item_type: ItemType::Text,
mcid: Some(7),
}];
let json = format_items_json(&items);
assert!(json.contains(r#""text":"A \"quoted\" item""#));
assert!(json.contains(r#""page":2"#));
assert!(json.contains(r#""x":12.35"#));
assert!(json.contains(r#""is_underline":true"#));
assert!(json.contains(r#""item_type":"text""#));
assert!(json.contains(r#""mcid":7"#));
}
}
/// Parse a page specification like "1,3,5-10,20" into a HashSet of page numbers.
fn parse_page_spec(spec: &str) -> Result<HashSet<u32>, String> {
let mut pages = HashSet::new();
@@ -88,6 +194,7 @@ fn main() {
if args.len() < 2 {
eprintln!("Usage: {} <pdf_file> [output_file]", args[0]);
eprintln!(" {} <pdf_file> --json", args[0]);
eprintln!(" {} <pdf_file> --items-json", args[0]);
eprintln!(" {} <pdf_file> --raw", args[0]);
eprintln!();
eprintln!("Converts PDF to Markdown with smart type detection.");
@@ -95,6 +202,7 @@ fn main() {
eprintln!();
eprintln!("Options:");
eprintln!(" --json Output result as JSON");
eprintln!(" --items-json Output positioned TextItem JSON");
eprintln!(" --raw Output only markdown (no headers)");
eprintln!(" --pages Insert page break markers (<!-- Page N -->)");
eprintln!(" --select-pages N Only process specified pages (e.g. 1,3,5-10)");
@@ -105,6 +213,7 @@ fn main() {
let pdf_path = &args[1];
let json_output = args.iter().any(|a| a == "--json");
let items_json_output = args.iter().any(|a| a == "--items-json");
let raw_output = args.iter().any(|a| a == "--raw");
let page_numbers = args.iter().any(|a| a == "--pages");
let detect_only = args.iter().any(|a| a == "--detect-only");
@@ -129,6 +238,17 @@ fn main() {
})
});
if items_json_output {
match extract_text_with_positions_pages(pdf_path, page_filter.as_ref()) {
Ok(items) => println!("{}", format_items_json(&items)),
Err(e) => {
println!(r#"{{"error":"{}"}}"#, json_escape(&e.to_string()));
process::exit(1);
}
}
return;
}
let output_file = args
.get(2)
.filter(|a| !a.starts_with("--"))
@@ -177,12 +297,14 @@ fn main() {
.iter()
.map(|p| p.to_string())
.collect();
let ocr_reasons = format_ocr_reasons_by_page(&result.ocr_reasons_by_page);
println!(
r#"{{"pdf_type":"{}","page_count":{},"processing_time_ms":{},"pages_needing_ocr":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"has_encoding_issues":{}}}"#,
r#"{{"pdf_type":"{}","page_count":{},"processing_time_ms":{},"pages_needing_ocr":[{}],"ocr_reasons_by_page":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"has_encoding_issues":{}}}"#,
pdf_type_str,
result.page_count,
result.processing_time_ms,
ocr_pages.join(","),
ocr_reasons,
result.layout.is_complex,
table_pages.join(","),
col_pages.join(","),
@@ -223,8 +345,9 @@ fn main() {
.iter()
.map(|p| p.to_string())
.collect();
let ocr_reasons = format_ocr_reasons_by_page(&result.ocr_reasons_by_page);
println!(
r#"{{"pdf_type":"{}","page_count":{},"has_text":{},"processing_time_ms":{},"markdown_length":{},"pages_needing_ocr":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"has_encoding_issues":{},"markdown":"{}"}}"#,
r#"{{"pdf_type":"{}","page_count":{},"has_text":{},"processing_time_ms":{},"markdown_length":{},"pages_needing_ocr":[{}],"ocr_reasons_by_page":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"has_encoding_issues":{},"markdown":"{}"}}"#,
match result.pdf_type {
PdfType::TextBased => "text_based",
PdfType::Scanned => "scanned",
@@ -236,6 +359,7 @@ fn main() {
result.processing_time_ms,
result.markdown.as_ref().map(|m| m.len()).unwrap_or(0),
ocr_pages.join(","),
ocr_reasons,
result.layout.is_complex,
table_pages.join(","),
col_pages.join(","),
+321 -19
View File
@@ -17,6 +17,7 @@ use super::fonts::{
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
};
use super::underline::UnderlineLine;
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
@@ -74,6 +75,37 @@ fn strip_pdf_comments(data: &[u8]) -> Vec<u8> {
result
}
fn transform_path_point(x: f32, y: f32, ctm: &[f32; 6]) -> (f32, f32) {
(
x * ctm[0] + y * ctm[2] + ctm[4],
x * ctm[1] + y * ctm[3] + ctm[5],
)
}
fn transformed_stroke_width(
line_width: f32,
ctm: &[f32; 6],
x1: f32,
y1: f32,
x2: f32,
y2: f32,
) -> f32 {
let user_width = line_width.abs();
let dx = x2 - x1;
let dy = y2 - y1;
let len = (dx * dx + dy * dy).sqrt();
if len <= f32::EPSILON {
return user_width;
}
// PDF stroke width scales perpendicular to the path direction.
let nx = -dy / len;
let ny = dx / len;
let ndx = nx * ctm[0] + ny * ctm[2];
let ndy = nx * ctm[1] + ny * ctm[3];
user_width * (ndx * ndx + ndy * ndy).sqrt()
}
/// Returns `(page_extraction, has_gid_fonts)` where `has_gid_fonts` indicates
/// the page uses fonts with unresolvable gid-encoded glyphs.
pub(crate) fn extract_page_text_items(
@@ -89,6 +121,7 @@ pub(crate) fn extract_page_text_items(
let mut rects: Vec<PdfRect> = Vec::new();
let mut clip_rects: Vec<PdfRect> = Vec::new();
let mut lines: Vec<PdfLine> = Vec::new();
let mut underline_lines: Vec<UnderlineLine> = Vec::new();
// Path construction state for m/l/h → S/s line extraction
let mut path_subpath_start: Option<(f32, f32)> = None;
@@ -97,6 +130,12 @@ pub(crate) fn extract_page_text_items(
// Completed subpaths (each a vec of line segments) for f/f* rect extraction
let mut pending_subpaths: Vec<Vec<(f32, f32, f32, f32)>> = Vec::new();
let mut fill_rects: Vec<PdfRect> = Vec::new();
// `re` rects awaiting a paint operator. Underline detection must only
// see painted rects: a `re W n` clip path or `re n` no-op draws nothing
// on the page, so treating every `re` as ink would underline text that
// merely sits near an invisible clip boundary.
let mut pending_re_rects: Vec<PdfRect> = Vec::new();
let mut painted_rects: Vec<PdfRect> = Vec::new();
// Get fonts for encoding
let fonts = doc.get_page_fonts(page_id).unwrap_or_default();
@@ -188,7 +227,19 @@ pub(crate) fn extract_page_text_items(
// Graphics state tracking
let mut ctm = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0]; // Current Transformation Matrix
let mut text_rendering_mode: i32 = 0; // 0=fill, 1=stroke, 2=fill+stroke, 3=invisible
let mut gstate_stack: Vec<([f32; 6], i32, f32, f32)> = Vec::new();
let mut line_width: f32 = 1.0;
#[derive(Clone)]
struct SavedGraphicsState {
ctm: [f32; 6],
text_rendering_mode: i32,
line_width: f32,
char_spacing: f32,
word_spacing: f32,
text_leading: f32,
current_font: String,
current_font_size: f32,
}
let mut gstate_stack: Vec<SavedGraphicsState> = Vec::new();
// Text state tracking
let mut current_font = String::new();
@@ -227,15 +278,28 @@ pub(crate) fn extract_page_text_items(
match op.operator.as_str() {
"q" => {
// Save graphics state
gstate_stack.push((ctm, text_rendering_mode, char_spacing, word_spacing));
gstate_stack.push(SavedGraphicsState {
ctm,
text_rendering_mode,
line_width,
char_spacing,
word_spacing,
text_leading,
current_font: current_font.clone(),
current_font_size,
});
}
"Q" => {
// Restore graphics state
if let Some((saved_ctm, saved_tr, saved_tc, saved_tw)) = gstate_stack.pop() {
ctm = saved_ctm;
text_rendering_mode = saved_tr;
char_spacing = saved_tc;
word_spacing = saved_tw;
if let Some(saved) = gstate_stack.pop() {
ctm = saved.ctm;
text_rendering_mode = saved.text_rendering_mode;
line_width = saved.line_width;
char_spacing = saved.char_spacing;
word_spacing = saved.word_spacing;
text_leading = saved.text_leading;
current_font = saved.current_font;
current_font_size = saved.current_font_size;
}
}
"cm" => {
@@ -252,6 +316,11 @@ pub(crate) fn extract_page_text_items(
ctm = multiply_matrices(&new_matrix, &ctm);
}
}
"w" => {
if let Some(width) = op.operands.first().and_then(get_number) {
line_width = width;
}
}
"BT" => {
// Begin text block
in_text_block = true;
@@ -420,6 +489,7 @@ pub(crate) fn extract_page_text_items(
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -585,6 +655,7 @@ pub(crate) fn extract_page_text_items(
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -648,6 +719,7 @@ pub(crate) fn extract_page_text_items(
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
item_type: ItemType::Text,
mcid: current_mcid(&marked_content_stack),
});
@@ -684,6 +756,7 @@ pub(crate) fn extract_page_text_items(
page: page_num,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Image,
mcid: current_mcid(&marked_content_stack),
});
@@ -780,6 +853,7 @@ pub(crate) fn extract_page_text_items(
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
item_type: ItemType::Text,
mcid: entry
.mcid
@@ -804,13 +878,19 @@ pub(crate) fn extract_page_text_items(
let y_dev = rx * ctm[1] + ry * ctm[3] + ctm[5];
let w_dev = rw * ctm[0];
let h_dev = rh * ctm[3];
rects.push(PdfRect {
let rect = PdfRect {
x: x_dev,
y: y_dev,
width: w_dev,
height: h_dev,
page: page_num,
});
};
// Underline detection must only see rects that are
// actually painted — a `re` used purely as a clip path
// (`re W n`) or discarded (`re n`) draws nothing. Hold
// the rect as pending until a paint operator confirms it.
pending_re_rects.push(rect.clone());
rects.push(rect);
}
}
// ── Path construction operators ──────────────────────
@@ -860,10 +940,8 @@ pub(crate) fn extract_page_text_items(
}
}
for (x1, y1, x2, y2) in pending_lines.drain(..) {
let x1d = x1 * ctm[0] + y1 * ctm[2] + ctm[4];
let y1d = x1 * ctm[1] + y1 * ctm[3] + ctm[5];
let x2d = x2 * ctm[0] + y2 * ctm[2] + ctm[4];
let y2d = x2 * ctm[1] + y2 * ctm[3] + ctm[5];
let (x1d, y1d) = transform_path_point(x1, y1, &ctm);
let (x2d, y2d) = transform_path_point(x2, y2, &ctm);
lines.push(PdfLine {
x1: x1d,
y1: y1d,
@@ -871,7 +949,16 @@ pub(crate) fn extract_page_text_items(
y2: y2d,
page: page_num,
});
underline_lines.push(UnderlineLine {
x1: x1d,
y1: y1d,
x2: x2d,
y2: y2d,
stroke_width: transformed_stroke_width(line_width, &ctm, x1, y1, x2, y2),
page: page_num,
});
}
painted_rects.append(&mut pending_re_rects);
pending_subpaths.clear();
path_subpath_start = None;
path_current = None;
@@ -887,10 +974,8 @@ pub(crate) fn extract_page_text_items(
}
}
for (x1, y1, x2, y2) in pending_lines.drain(..) {
let x1d = x1 * ctm[0] + y1 * ctm[2] + ctm[4];
let y1d = x1 * ctm[1] + y1 * ctm[3] + ctm[5];
let x2d = x2 * ctm[0] + y2 * ctm[2] + ctm[4];
let y2d = x2 * ctm[1] + y2 * ctm[3] + ctm[5];
let (x1d, y1d) = transform_path_point(x1, y1, &ctm);
let (x2d, y2d) = transform_path_point(x2, y2, &ctm);
lines.push(PdfLine {
x1: x1d,
y1: y1d,
@@ -898,7 +983,16 @@ pub(crate) fn extract_page_text_items(
y2: y2d,
page: page_num,
});
underline_lines.push(UnderlineLine {
x1: x1d,
y1: y1d,
x2: x2d,
y2: y2d,
stroke_width: transformed_stroke_width(line_width, &ctm, x1, y1, x2, y2),
page: page_num,
});
}
painted_rects.append(&mut pending_re_rects);
pending_subpaths.clear();
path_subpath_start = None;
path_current = None;
@@ -956,6 +1050,7 @@ pub(crate) fn extract_page_text_items(
}
}
}
painted_rects.append(&mut pending_re_rects);
pending_lines.clear();
path_subpath_start = None;
path_current = None;
@@ -1021,7 +1116,10 @@ pub(crate) fn extract_page_text_items(
// Do NOT clear pending_lines — the following `n` does that
}
"n" => {
// end path (no-op): discard
// end path (no-op): discard — including any `re` rects that
// were only ever part of a clip path (`re W n`), which draw
// no ink and must not feed underline detection.
pending_re_rects.clear();
pending_lines.clear();
pending_subpaths.clear();
path_subpath_start = None;
@@ -1031,6 +1129,12 @@ pub(crate) fn extract_page_text_items(
}
}
// Underline detection reads only painted ink: `re` rects confirmed by
// a paint operator plus filled-subpath rects — never clip-only rects,
// which draw nothing.
let mut underline_rects = painted_rects;
underline_rects.extend(fill_rects.iter().cloned());
// Only use clip/fill rects when no `re` rects exist on this page.
// Clip rects take priority over fill rects, but first we deduplicate
// them: some PDFs wrap every text block in a full-page W* clip path,
@@ -1060,8 +1164,17 @@ pub(crate) fn extract_page_text_items(
// Some PDFs embed landscape content in portrait pages using a rotated text
// matrix (e.g. [0, b, -b, 0, tx, ty] for 90° CCW). The layout engine
// assumes x=horizontal, y=vertical — so we swap coordinates to match.
let (items, rects, lines, coords_rotated) =
let (mut items, rects, lines, coords_rotated) =
correct_rotated_page(items, rects, lines, &rotation_votes);
if coords_rotated {
rotate_underline_graphics(&mut underline_rects, &mut underline_lines);
}
super::underline::mark_underlined_items(
&mut items,
&underline_rects,
&underline_lines,
page_num,
);
let items = super::merge_text_items(items);
let items = super::merge_subscript_items(items);
@@ -1147,6 +1260,27 @@ fn correct_rotated_page(
(items, rects, lines, true)
}
fn rotate_underline_graphics(rects: &mut [PdfRect], lines: &mut [UnderlineLine]) {
for rect in rects {
let new_x = rect.y;
let new_y = -(rect.x + rect.width.abs());
rect.x = new_x;
rect.y = new_y;
std::mem::swap(&mut rect.width, &mut rect.height);
}
for line in lines {
let new_x1 = line.y1;
let new_y1 = -line.x1;
let new_x2 = line.y2;
let new_y2 = -line.x2;
line.x1 = new_x1;
line.y1 = new_y1;
line.x2 = new_x2;
line.y2 = new_y2;
}
}
/// Remove near-duplicate rects (same coordinates within 0.5 pt tolerance).
/// Some PDFs emit a full-page clip path for every text block, producing
/// thousands of identical rects. After dedup these collapse to one rect,
@@ -1196,6 +1330,57 @@ mod tests {
}
}
fn simple_doc_with_content(content: &[u8]) -> (lopdf::Document, lopdf::ObjectId) {
use lopdf::{dictionary, Object, Stream};
let mut doc = lopdf::Document::new();
let widths: Vec<Object> = (0..=255).map(|_| 600.into()).collect();
let font_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
"FirstChar" => 0,
"LastChar" => 255,
"Widths" => Object::Array(widths),
});
let content_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
content.to_vec(),
)));
let page_id = doc.add_object(dictionary! {
"Type" => "Page",
"Contents" => Object::Reference(content_id),
"Resources" => dictionary! {
"Font" => dictionary! {
"F1" => Object::Reference(font_id),
},
},
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
});
let pages_id = doc.add_object(dictionary! {
"Type" => "Pages",
"Count" => Object::Integer(1),
"Kids" => vec![Object::Reference(page_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => Object::Reference(pages_id),
});
doc.trailer.set("Root", Object::Reference(catalog_id));
(doc, page_id)
}
fn extract_simple_items(content: &[u8]) -> Vec<TextItem> {
use crate::tounicode::FontCMaps;
let (doc, page_id) = simple_doc_with_content(content);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) =
extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
items
}
#[test]
fn test_dedup_rects_identical() {
let mut rects = vec![rect(0.0, 0.0, 612.0, 792.0, 1); 3759];
@@ -1245,6 +1430,38 @@ mod tests {
assert_eq!(single.len(), 1);
}
#[test]
fn thick_stroked_rule_does_not_mark_underline() {
let content = b"BT /F1 12 Tf 1 0 0 1 100 500 Tm (THICK) Tj ET
4 w
100 498 m 170 498 l S
BT /F1 12 Tf 1 0 0 1 100 480 Tm (THIN) Tj ET
1 w
100 478 m 160 478 l S";
let items = extract_simple_items(content);
let thick = items.iter().find(|item| item.text == "THICK").unwrap();
let thin = items.iter().find(|item| item.text == "THIN").unwrap();
assert!(!thick.is_underline);
assert!(thin.is_underline);
}
#[test]
fn rotated_page_underline_is_detected_after_coordinate_correction() {
let content = b"BT /F1 12 Tf 0 1 -1 0 200 100 Tm (HELLO) Tj ET
BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
1 w
202 100 m 202 170 l S";
let items = extract_simple_items(content);
let hello = items.iter().find(|item| item.text == "HELLO").unwrap();
let world = items.iter().find(|item| item.text == "WORLD").unwrap();
assert!(hello.is_underline);
assert!(!world.is_underline);
}
#[test]
fn test_skip_excessive_operations() {
use crate::tounicode::FontCMaps;
@@ -1286,6 +1503,91 @@ mod tests {
assert!(lines.is_empty());
}
#[test]
fn test_q_restores_current_font_for_text_decoding() {
use crate::tounicode::FontCMaps;
use lopdf::{dictionary, Object, Stream};
fn cmap_stream(dst_hex: &str) -> Stream {
let cmap = format!(
r#"/CIDInit /ProcSet findresource begin
12 dict begin
begincmap
/CIDSystemInfo << /Registry (Adobe) /Ordering (UCS) /Supplement 0 >> def
/CMapName /Test-UCS def
/CMapType 2 def
1 begincodespacerange
<00> <FF>
endcodespacerange
1 beginbfchar
<41> <{dst_hex}>
endbfchar
endcmap
CMapName currentdict /CMap defineresource pop
end
end"#
);
Stream::new(dictionary! {}, cmap.into_bytes())
}
let mut doc = lopdf::Document::new();
let f1_cmap = doc.add_object(Object::Stream(cmap_stream("0058"))); // X
let f2_cmap = doc.add_object(Object::Stream(cmap_stream("0059"))); // Y
let f1 = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
"ToUnicode" => Object::Reference(f1_cmap),
});
let f2 = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
"ToUnicode" => Object::Reference(f2_cmap),
});
let content = b"BT /F1 12 Tf 10 700 Tm <41> Tj ET
q
BT /F2 12 Tf 20 700 Tm <41> Tj ET
Q
BT 30 700 Tm <41> Tj ET";
let content_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
content.to_vec(),
)));
let page_id = doc.add_object(dictionary! {
"Type" => "Page",
"Contents" => Object::Reference(content_id),
"Resources" => dictionary! {
"Font" => dictionary! {
"F1" => Object::Reference(f1),
"F2" => Object::Reference(f2),
},
},
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
});
let pages_id = doc.add_object(dictionary! {
"Type" => "Pages",
"Count" => Object::Integer(1),
"Kids" => vec![Object::Reference(page_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => Object::Reference(pages_id),
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _) =
extract_page_text_items(&doc, page_id, 1, &font_cmaps, false).unwrap();
let text = items
.iter()
.map(|item| item.text.as_str())
.collect::<String>();
assert_eq!(text, "XYX");
}
#[test]
fn test_strip_pdf_comments() {
// Basic comment stripping
+277 -46
View File
@@ -526,6 +526,11 @@ pub(crate) fn parse_font_encoding(
font_dict: &lopdf::Dictionary,
) -> Option<EncodingResult> {
let encoding_obj = font_dict.get(b"Encoding").ok()?;
let base_font_name = font_dict
.get(b"BaseFont")
.ok()
.and_then(|o| o.as_name().ok())
.map(|n| String::from_utf8_lossy(n).to_string());
// Encoding can be a name or a dictionary
match encoding_obj {
@@ -538,12 +543,14 @@ pub(crate) fn parse_font_encoding(
Object::Reference(obj_ref) => {
// Reference to encoding dictionary
if let Ok(enc_dict) = doc.get_dictionary(*obj_ref) {
parse_encoding_dictionary(doc, enc_dict)
parse_encoding_dictionary(doc, enc_dict, base_font_name.as_deref())
} else {
None
}
}
Object::Dictionary(enc_dict) => parse_encoding_dictionary(doc, enc_dict),
Object::Dictionary(enc_dict) => {
parse_encoding_dictionary(doc, enc_dict, base_font_name.as_deref())
}
_ => None,
}
}
@@ -562,6 +569,7 @@ pub(crate) struct EncodingResult {
pub(crate) fn parse_encoding_dictionary(
doc: &Document,
enc_dict: &lopdf::Dictionary,
base_font_name: Option<&str>,
) -> Option<EncodingResult> {
let differences = enc_dict.get(b"Differences").ok()?;
@@ -591,11 +599,9 @@ pub(crate) fn parse_encoding_dictionary(
Object::Name(name) => {
// Map current code to glyph name -> Unicode
let glyph_name = String::from_utf8_lossy(&name).to_string();
if glyph_name == "fi"
|| glyph_name == "fl"
|| glyph_name == "ffi"
|| glyph_name == "ffl"
{
let mapped_char = glyph_to_char(&glyph_name)
.or_else(|| private_glyph_to_char(&glyph_name, base_font_name));
if mapped_char.is_some_and(is_ligature_char) {
debug!(
" Differences: code=0x{:02X} glyph={:?} (ligature)",
current_code, glyph_name
@@ -610,7 +616,7 @@ pub(crate) fn parse_encoding_dictionary(
{
gid_glyph_count += 1;
}
if let Some(ch) = glyph_to_char(&glyph_name) {
if let Some(ch) = mapped_char {
encoding_map.insert(current_code, ch);
} else {
debug!(
@@ -645,6 +651,31 @@ pub(crate) fn parse_encoding_dictionary(
})
}
fn private_glyph_to_char(glyph_name: &str, base_font_name: Option<&str>) -> Option<char> {
let base_font_name = strip_subset_prefix(base_font_name?);
// Aptos CFF subsets from Office PDFs can expose the ff ligature as /g431
// without a ToUnicode map. Keep this font-scoped because /gNNN names are private.
if base_font_name.eq_ignore_ascii_case("Aptos") && glyph_name == "g431" {
Some('\u{FB00}')
} else {
None
}
}
fn strip_subset_prefix(font_name: &str) -> &str {
font_name
.split_once('+')
.map_or(font_name, |(_, stripped)| stripped)
}
fn is_ligature_char(ch: char) -> bool {
matches!(
ch,
'\u{FB00}' | '\u{FB01}' | '\u{FB02}' | '\u{FB03}' | '\u{FB04}'
)
}
/// Get the CMap lookup key for an Identity-H/V CID font without ToUnicode.
/// Returns the object number used by `collect_cmaps_from_fonts` to store the CMap:
/// - FontFile2 or FontFile3 obj_num (for embedded font cmap)
@@ -725,6 +756,8 @@ pub(crate) fn extract_text_from_operand(
let is_type0_cid_font = font_widths
.get(current_font)
.is_some_and(|info| info.is_cid);
let use_cp1252_fallback =
should_use_cp1252_single_byte_fallback(base_font_name, is_type0_cid_font);
let result = (|| -> Option<String> {
if let Object::String(bytes, _) = obj {
let mut decode_with_entry = |entry: &crate::tounicode::CMapEntry| -> Option<String> {
@@ -755,9 +788,12 @@ pub(crate) fn extract_text_from_operand(
return Some(ch.to_string());
}
}
// 4. Printable ASCII/Latin-1 fallback
// 4. Printable single-byte fallback
if b >= 0x20 {
return Some((b as char).to_string());
return Some(
decode_single_byte_fallback_char(b, use_cp1252_fallback)
.to_string(),
);
}
None
})
@@ -858,6 +894,13 @@ pub(crate) fn extract_text_from_operand(
// unmapped. Don't fall through to text-interpretation fallbacks
// (Latin-1, UTF-16, etc.) which would misinterpret CID bytes as
// character codes (e.g. CID 0x01A9 → Latin-1 "©").
if is_type0_cid_font && bytes.iter().any(|&b| b > 0x7F) {
// 2-byte CIDs (Identity-H) are by far the common case; for
// an odd byte count we still emit at least one marker so
// detection downstream fires.
let cid_count = (bytes.len() / 2).max(1);
return Some("\u{FFFD}".repeat(cid_count));
}
// Try our custom encoding map from Differences arrays.
// The Differences array overrides specific codes in a base encoding (typically
@@ -873,8 +916,9 @@ pub(crate) fn extract_text_from_operand(
Some(ch)
} else if b >= 0x20 {
// Base encoding fallback for printable bytes.
// For codes 0x20-0x7E this matches all standard PDF encodings.
Some(b as char)
// Most PDFs with simple fonts use WinAnsi/PDFDocEncoding
// semantics, not ISO-8859-1 C1 controls.
Some(decode_single_byte_fallback_char(b, use_cp1252_fallback))
} else {
None // Skip unmapped control characters
}
@@ -930,6 +974,7 @@ pub(crate) fn extract_text_from_operand(
// Try to decode using cached font encoding from lopdf
if let Some(encoding) = encoding_cache.get(current_font) {
if let Ok(text) = Document::decode_text(encoding, bytes) {
let text = normalize_cp1252_controls(text, use_cp1252_fallback);
if text.contains('\u{FFFD}') {
debug!(
"decode_text produced replacement for font={} bytes_len={}",
@@ -966,37 +1011,119 @@ pub(crate) fn extract_text_from_operand(
return Some(symbol_text);
}
// Latin-1 fallback. Safe ONLY for fonts that use single-byte
// encodings — for these, an unmapped byte is a valid character
// code in Latin-1/WinAnsi space. CID fonts (Type0 / Identity-H)
// emit multi-byte CIDs that aren't characters; per-byte Latin-1
// produces mojibake (e.g. 2-byte CID 0xCDD9 → "ÍÙ" for the
// production scrape_id 019de78c-... samples).
//
// For a CID font (has_cmap is set OR a /ToUnicode reference
// exists) with any non-ASCII bytes, emit a single U+FFFD per
// CID instead. This both replaces the mojibake with a proper
// "decode failed" marker AND keeps `detect_encoding_issues`
// tripping so the page is flagged for OCR — the existing
// garbage-detection path that the high-Latin-1 mojibake used
// to satisfy by accident.
if is_type0_cid_font && bytes.iter().any(|&b| b > 0x7F) {
// 2-byte CIDs (Identity-H) are by far the common case; for
// an odd byte count we still emit at least one marker so
// detection downstream fires.
let cid_count = (bytes.len() / 2).max(1);
return Some("\u{FFFD}".repeat(cid_count));
}
// Pure ASCII bytes round-trip safely (Latin-1 == ASCII for
// 0x00..=0x7F), and non-CID (Type1 / TrueType / Type3) fonts
// use single-byte encodings where Latin-1 fallback is the
// canonical interpretation.
Some(bytes.iter().map(|&b| b as char).collect())
// Non-CID (Type1 / TrueType / Type3) fonts use single-byte
// encodings. In practice the fallback should follow WinAnsi for
// 0x80..=0x9F so bytes like 0x92 become smart punctuation instead
// of C1 controls that look like CID mojibake.
Some(decode_single_byte_fallback(bytes, use_cp1252_fallback))
} else {
None
}
})();
result.map(clean_symbol_pua)
result.map(|text| {
let text = clean_symbol_pua(text);
normalize_cp1252_controls(text, use_cp1252_fallback)
})
}
fn decode_single_byte_fallback(bytes: &[u8], use_cp1252_fallback: bool) -> String {
bytes
.iter()
.map(|&b| decode_single_byte_fallback_char(b, use_cp1252_fallback))
.collect()
}
fn decode_single_byte_fallback_char(byte: u8, use_cp1252_fallback: bool) -> char {
if !use_cp1252_fallback {
return byte as char;
}
match byte {
0x80 => '\u{20AC}',
0x82 => '\u{201A}',
0x83 => '\u{0192}',
0x84 => '\u{201E}',
0x85 => '\u{2026}',
0x86 => '\u{2020}',
0x87 => '\u{2021}',
0x88 => '\u{02C6}',
0x89 => '\u{2030}',
0x8A => '\u{0160}',
0x8B => '\u{2039}',
0x8C => '\u{0152}',
0x8E => '\u{017D}',
0x91 => '\u{2018}',
0x92 => '\u{2019}',
0x93 => '\u{201C}',
0x94 => '\u{201D}',
0x95 => '\u{2022}',
0x96 => '\u{2013}',
0x97 => '\u{2014}',
0x98 => '\u{02DC}',
0x99 => '\u{2122}',
0x9A => '\u{0161}',
0x9B => '\u{203A}',
0x9C => '\u{0153}',
0x9E => '\u{017E}',
0x9F => '\u{0178}',
_ => byte as char,
}
}
fn normalize_cp1252_controls(text: String, use_cp1252_fallback: bool) -> String {
if !use_cp1252_fallback {
return text;
}
if !text
.chars()
.any(|ch| ('\u{0080}'..='\u{009F}').contains(&ch))
{
return text;
}
text.chars()
.map(|ch| {
if ('\u{0080}'..='\u{009F}').contains(&ch) {
decode_single_byte_fallback_char(ch as u8, true)
} else {
ch
}
})
.collect()
}
fn should_use_cp1252_single_byte_fallback(
base_font_name: Option<&str>,
is_type0_cid_font: bool,
) -> bool {
if is_type0_cid_font {
return false;
}
let Some(base_font_name) = base_font_name else {
return true;
};
let font_name = base_font_name
.rsplit_once('+')
.map_or(base_font_name, |(_, stripped)| stripped)
.to_ascii_lowercase();
// TeX/Computer Modern and math/symbol fonts often place ligatures or
// symbols in the C1 byte range. Treating those bytes as Windows-1252 makes
// words like "deficiente" become "de…ciente" and "fluid" become "‡uid".
let non_cp1252_prefixes = [
"cmr", "cmb", "cmmi", "cmsy", "cmex", "cmtt", "cmss", "cmti", "ecrm", "ecbx", "ecti",
"tcrm", "tctt", "msam", "msbm", "ttdc",
];
if non_cp1252_prefixes
.iter()
.any(|prefix| font_name.starts_with(prefix))
{
return false;
}
let non_cp1252_names = ["math", "symbol", "dingbat", "emoji"];
!non_cp1252_names.iter().any(|name| font_name.contains(name))
}
/// Replace PUA characters in the F000-F0FF range with standard Unicode equivalents.
@@ -1118,6 +1245,7 @@ fn score_text(text: &str) -> i32 {
#[cfg(test)]
mod tests {
use super::*;
use lopdf::dictionary;
fn make_font_info(widths: &[(u16, u16)], default_width: u16, is_cid: bool) -> FontWidthInfo {
FontWidthInfo {
@@ -1242,6 +1370,51 @@ mod tests {
assert!(score_text(good) > score_text(bad));
}
fn doc_with_private_differences() -> (Document, lopdf::ObjectId) {
let mut doc = Document::with_version("1.7");
let encoding_id = doc.add_object(dictionary! {
"Differences" => Object::Array(vec![
Object::Integer(0x88),
Object::Name(b"g431".to_vec()),
Object::Name(b"fi".to_vec()),
Object::Integer(0xAD),
Object::Name(b"fl".to_vec()),
]),
});
(doc, encoding_id)
}
#[test]
fn aptos_private_g431_maps_to_ff_ligature() {
let (doc, encoding_id) = doc_with_private_differences();
let font_dict = dictionary! {
"BaseFont" => Object::Name(b"NJEQOD+Aptos".to_vec()),
"Encoding" => Object::Reference(encoding_id),
};
let result = parse_font_encoding(&doc, &font_dict).expect("encoding should parse");
assert_eq!(result.map.get(&0x88u8), Some(&'\u{FB00}'));
assert_eq!(result.map.get(&0x89u8), Some(&'\u{FB01}'));
assert_eq!(result.map.get(&0xADu8), Some(&'\u{FB02}'));
}
#[test]
fn private_g431_does_not_map_for_unrelated_fonts() {
let (doc, encoding_id) = doc_with_private_differences();
let font_dict = dictionary! {
"BaseFont" => Object::Name(b"ABCDEF+OtherFont".to_vec()),
"Encoding" => Object::Reference(encoding_id),
};
let result = parse_font_encoding(&doc, &font_dict).expect("encoding should parse");
assert!(!result.map.contains_key(&0x88u8));
assert_eq!(result.map.get(&0x89u8), Some(&'\u{FB01}'));
assert_eq!(result.map.get(&0xADu8), Some(&'\u{FB02}'));
}
#[test]
fn cid_font_with_unparseable_cmap_does_not_emit_latin1_mojibake() {
// Type0/CID font (font_widths reports `is_cid=true`) where the
@@ -1293,15 +1466,15 @@ mod tests {
}
#[test]
fn simple_font_latin1_fallback_passes_high_bytes_through() {
fn simple_font_single_byte_fallback_passes_high_bytes_through() {
// A Type1/TrueType simple font (is_cid=false) with a `/ToUnicode`
// reference but no usable CMap and no `/Differences` map.
// Per-byte Latin-1 IS the canonical interpretation here — these
// bytes are character codes, not CIDs. The CID guard must NOT
// strip them. Reproduces the false positive that an earlier
// version of the guard introduced for fonts in PDFs like
// pdf-evals/Navigating-Artificial-Intelligence-..., where bytes
// like 0xB6 are legitimate Latin-1 character codes.
// Per-byte fallback is the canonical interpretation here — these
// bytes are character codes, not CIDs. The CID guard must NOT strip
// them. Reproduces the false positive that an earlier version of the
// guard introduced for fonts in PDFs like pdf-evals/Navigating-
// Artificial-Intelligence-..., where bytes like 0xB6 are legitimate
// single-byte character codes.
let bytes = vec![0x24_u8, 0x47, 0xB6, 0x56]; // "$G¶V"
let obj = Object::String(bytes, lopdf::StringFormat::Hexadecimal);
@@ -1334,4 +1507,62 @@ mod tests {
"simple font fallback must not stamp FFFD over legitimate bytes: {text:?}"
);
}
#[test]
fn simple_font_single_byte_fallback_maps_cp1252_punctuation() {
let bytes = vec![b'l', 0x92_u8, b'a', b'c', b'a', b'd'];
let obj = Object::String(bytes, lopdf::StringFormat::Hexadecimal);
let font_cmaps = FontCMaps::default();
let font_tounicode_refs: HashMap<String, u32> = HashMap::new();
let inline_cmaps = HashMap::new();
let font_encodings: PageFontEncodings = HashMap::new();
let encoding_cache: HashMap<String, Encoding<'_>> = HashMap::new();
let mut decisions = CMapDecisionCache::new();
let font_widths: PageFontWidths = HashMap::new();
let text = extract_text_from_operand(
&obj,
"F1",
None,
&font_cmaps,
&font_tounicode_refs,
&inline_cmaps,
&font_encodings,
&encoding_cache,
&mut decisions,
&font_widths,
)
.expect("simple font should decode CP1252 punctuation");
assert_eq!(text, "lacad");
}
#[test]
fn cached_encoding_decode_normalizes_cp1252_controls() {
let text = normalize_cp1252_controls("d\u{92}un \u{96} test".to_string(), true);
assert_eq!(text, "dun test");
}
#[test]
fn tex_font_decode_keeps_c1_ligature_bytes_unmodified() {
let text = normalize_cp1252_controls("de\u{85}ciente \u{87}uid".to_string(), false);
assert_eq!(text, "de\u{85}ciente \u{87}uid");
assert!(!should_use_cp1252_single_byte_fallback(
Some("TTdcr10"),
false
));
assert!(!should_use_cp1252_single_byte_fallback(
Some("cmr10"),
false
));
}
#[test]
fn winansi_text_font_uses_cp1252_fallback() {
assert!(should_use_cp1252_single_byte_fallback(
Some("BJPQNQ+Times-Roman"),
false
));
}
}
+3 -5
View File
@@ -1234,11 +1234,7 @@ pub(crate) fn group_into_lines_with_thresholds(
ci,
item.x,
item.y,
if item.text.len() > 60 {
&item.text[..60]
} else {
&item.text
}
super::trace_text_preview(&item.text, 60)
);
}
}
@@ -1498,6 +1494,7 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1627,6 +1624,7 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
});
+2
View File
@@ -78,6 +78,7 @@ pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> V
page: page_num,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Link(url),
mcid: None,
});
@@ -316,6 +317,7 @@ pub(crate) fn walk_form_fields(
page: page_num,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::FormField,
mcid: None,
});
+448 -19
View File
@@ -6,11 +6,12 @@ pub(crate) mod content_stream;
mod fonts;
mod layout;
mod links;
pub(crate) mod underline;
mod xobjects;
use crate::text_utils::is_rtl_text;
use crate::tounicode::FontCMaps;
use crate::types::{PageExtraction, TextItem};
use crate::types::{PageExtraction, PdfLine, PdfRect, TextItem};
use crate::PdfError;
use log::debug;
use lopdf::{Document, Object, ObjectId};
@@ -33,6 +34,13 @@ pub(crate) use layout::ColumnRegion;
// Public API
// ---------------------------------------------------------------------------
pub(crate) fn trace_text_preview(text: &str, max_chars: usize) -> &str {
match text.char_indices().nth(max_chars) {
Some((idx, _)) => &text[..idx],
None => text,
}
}
/// Extract text from PDF file as plain string
pub fn extract_text<P: AsRef<Path>>(path: P) -> Result<String, PdfError> {
crate::validate_pdf_file(&path)?;
@@ -173,6 +181,7 @@ fn extract_positioned_text_impl(
if threshold > 0.10 {
page_thresholds.insert(*page_num, threshold);
}
suppress_table_underlines(&mut items, &rects, &lines, *page_num);
debug!(
"page {}: {} text items, {} rects, {} lines{}",
page_num,
@@ -195,11 +204,7 @@ fn extract_positioned_text_impl(
item.width,
item.font_size,
item.font,
if item.text.len() > 80 {
&item.text[..80]
} else {
&item.text
}
trace_text_preview(&item.text, 80)
);
}
}
@@ -223,6 +228,38 @@ fn extract_positioned_text_impl(
))
}
fn suppress_table_underlines(
items: &mut [TextItem],
rects: &[PdfRect],
lines: &[PdfLine],
page: u32,
) {
if !items.iter().any(|item| item.is_underline) {
return;
}
let mut table_item_indices: HashSet<usize> = HashSet::new();
if !rects.is_empty() {
let (rect_tables, _) = crate::tables::detect_tables_from_rects(items, rects, page);
for table in rect_tables {
table_item_indices.extend(table.item_indices);
}
}
if !lines.is_empty() {
for table in crate::tables::detect_tables_from_lines(items, lines, page) {
table_item_indices.extend(table.item_indices);
}
}
for index in table_item_indices {
if let Some(item) = items.get_mut(index) {
item.is_underline = false;
}
}
}
// ---------------------------------------------------------------------------
// Shared helpers (used by submodules via `super::`)
// ---------------------------------------------------------------------------
@@ -349,6 +386,133 @@ fn effective_merge_width(item: &TextItem) -> f32 {
}
}
fn is_standalone_bullet_text(text: &str) -> bool {
matches!(text.trim(), "" | "" | "" | "")
}
fn first_text_char(text: &str) -> Option<char> {
text.trim_start().chars().next()
}
fn is_short_alpha_fragment(text: &str) -> bool {
let trimmed = text.trim();
let char_count = trimmed.chars().count();
(1..=4).contains(&char_count) && trimmed.chars().all(char::is_alphabetic)
}
fn has_phrase_continuation_shape(text: &str) -> bool {
let trimmed = text.trim_start();
trimmed
.chars()
.take(24)
.any(|ch| ch.is_whitespace() || matches!(ch, '-'))
}
fn should_preserve_overlapping_stream_order(group: &[&TextItem]) -> bool {
if group.len() < 3 {
return false;
}
let Some(first) = group.iter().find(|item| !item.text.trim().is_empty()) else {
return false;
};
if group.iter().all(|item| item.mcid.is_none()) {
return false;
}
let mut nonempty_count = 0;
let mut saw_backtrack = false;
let mut nonspace_chars = 0;
let mut math_symbol_chars = 0;
let mut max_font_size = first.font_size;
for item in group {
if !item.text.trim().is_empty() {
nonempty_count += 1;
}
if (item.font_size - first.font_size).abs() > first.font_size * 0.25 {
return false;
}
max_font_size = max_font_size.max(item.font_size);
for ch in item.text.chars().filter(|ch| !ch.is_whitespace()) {
nonspace_chars += 1;
if matches!(
ch,
'*' | 'ˆ' | '^' | '=' | '+' | '_' | '[' | ']' | '{' | '}' | '|' | '<' | '>'
) {
math_symbol_chars += 1;
}
}
}
if nonempty_count < 2 {
return false;
}
if nonspace_chars > 0 && math_symbol_chars * 4 > nonspace_chars {
return false;
}
let mut sorted_by_x = group.to_vec();
sorted_by_x.sort_by(|a, b| a.x.total_cmp(&b.x));
let cluster_start = sorted_by_x[0].x;
let mut cluster_end = cluster_start + effective_merge_width(sorted_by_x[0]);
for item in sorted_by_x.iter().skip(1) {
let gap = item.x - cluster_end;
if gap > max_font_size * 2.5 {
return false;
}
cluster_end = cluster_end.max(item.x + effective_merge_width(item));
}
if cluster_end - cluster_start > max_font_size * 36.0 {
return false;
}
for index in 0..group.len() - 1 {
let previous = group[index];
let next = group[index + 1];
let font_size = previous.font_size.max(next.font_size);
let backtrack_threshold = font_size * 0.25;
let previous_start = previous.x;
let next_start = next.x;
let next_end = next.x + effective_merge_width(next);
if next_start < previous_start - backtrack_threshold
&& next_end > previous_start + backtrack_threshold
{
let has_near_prefix = group[..=index].iter().rev().take(4).any(|item| {
is_short_alpha_fragment(&item.text)
&& item.x >= next_start - font_size * 0.5
&& item.x <= next_start + font_size * 4.0
});
let starts_lowercase = first_text_char(&next.text).is_some_and(char::is_lowercase);
let phrase_continuation = has_phrase_continuation_shape(&next.text);
let has_near_bullet = group[..=index]
.iter()
.position(|item| {
is_standalone_bullet_text(&item.text) && next_start <= item.x + font_size * 3.0
})
.is_some_and(|bullet_index| {
if bullet_index >= index {
return false;
}
group[bullet_index + 1..=index]
.iter()
.rev()
.find(|item| !item.text.trim().is_empty())
.is_some_and(|item| {
item.text.trim().chars().count() <= 8
&& has_phrase_continuation_shape(&next.text)
})
});
if (has_near_prefix && starts_lowercase && phrase_continuation) || has_near_bullet {
saw_backtrack = true;
break;
}
}
}
saw_backtrack
}
pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
if items.is_empty() {
return items;
@@ -369,28 +533,32 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
}
}
// Sort each group by X position (direction-aware)
for (_, _, group) in &mut line_groups {
let mut ordered_line_groups: Vec<(u32, f32, Vec<&TextItem>, bool)> = Vec::new();
// Sort each group by X position (direction-aware), except for lines whose
// content stream intentionally backtracks to overlay ActualText fragments.
for (page, y, mut group) in line_groups {
let rtl = is_rtl_text(group.iter().map(|i| &i.text));
let preserve_stream_order = !rtl && should_preserve_overlapping_stream_order(&group);
if rtl {
group.sort_by(|a, b| b.x.total_cmp(&a.x));
} else {
} else if !preserve_stream_order {
group.sort_by(|a, b| a.x.total_cmp(&b.x));
}
ordered_line_groups.push((page, y, group, preserve_stream_order));
}
// Sort groups by page then Y descending (top of page first)
line_groups.sort_by(|a, b| a.0.cmp(&b.0).then_with(|| b.1.total_cmp(&a.1)));
ordered_line_groups.sort_by(|a, b| a.0.cmp(&b.0).then_with(|| b.1.total_cmp(&a.1)));
let mut merged = Vec::new();
for (_, _, group) in &line_groups {
for (_, _, group, preserve_stream_order) in &ordered_line_groups {
let mut i = 0;
while i < group.len() {
let first = group[i];
let mut text = first.text.clone();
let mut end_x = first.x + effective_merge_width(first);
let x_gap_max = first.font_size * 0.5;
let mut j = i + 1;
while j < group.len() {
@@ -399,11 +567,28 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
if (next.font_size - first.font_size).abs() > first.font_size * 0.20 {
break;
}
// Never merge across style boundaries: the merged item
// carries `first`'s flags, so absorbing a styled run into a
// plain neighbor (or vice versa) silently erases the styling
// that markdown emission and downstream inline-styling need —
// and OR-ing underline instead would stretch `<u>` spans over
// neighboring plain text.
if next.is_bold != first.is_bold
|| next.is_italic != first.is_italic
|| next.is_underline != first.is_underline
{
break;
}
let gap = next.x - end_x;
let x_gap_max = if *preserve_stream_order && is_standalone_bullet_text(&text) {
first.font_size * 1.2
} else {
first.font_size * 0.5
};
if gap > x_gap_max {
break;
}
if gap < -first.font_size * 0.5 {
if gap < -first.font_size * 0.5 && !preserve_stream_order {
break;
}
// Insert space at word boundaries.
@@ -425,11 +610,19 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
first.font_size * 0.08
}
};
if gap > threshold {
let needs_bullet_space = *preserve_stream_order
&& is_standalone_bullet_text(&text)
&& !next.text.trim().is_empty();
if needs_bullet_space || gap > threshold {
text.push(' ');
}
text.push_str(&next.text);
end_x = next.x + effective_merge_width(next);
let next_end = next.x + effective_merge_width(next);
end_x = if *preserve_stream_order {
end_x.max(next_end)
} else {
next_end
};
j += 1;
}
@@ -444,6 +637,7 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
page: first.page,
is_bold: first.is_bold,
is_italic: first.is_italic,
is_underline: first.is_underline,
item_type: first.item_type.clone(),
mcid: first.mcid,
});
@@ -526,7 +720,17 @@ pub(crate) fn merge_subscript_items(items: Vec<TextItem>) -> Vec<TextItem> {
let gap = item.x - parent_right;
// Subscripts must be tightly adjacent (within ~1pt)
if gap < parent.font_size * 0.2 && gap > -parent.font_size * 0.3 {
parent.text.push_str(&item.text);
// Preserve the script when absorbing it: map the
// digits to Unicode sub/superscript forms so the
// raised/lowered rendering survives in extracted
// text ("H"+"2" → "H₂", "word"+"2" → "word²").
// NFKC/NFKD normalization folds these back to
// plain digits, so text matching downstream is
// unaffected. Direction from the baseline offset
// (y-up here): raised → superscript (footnote
// refs), lowered/level → subscript (chemistry).
let raised = item.y > parent.y + parent.font_size * 0.1;
parent.text.push_str(&map_script_digits(&item.text, raised));
parent.width = (item.x + item.width) - parent.x;
continue;
}
@@ -541,6 +745,21 @@ pub(crate) fn merge_subscript_items(items: Vec<TextItem>) -> Vec<TextItem> {
result
}
/// Map ASCII digits to their Unicode superscript (`raised`) or subscript
/// forms. Callers guarantee digit-only input (see `merge_subscript_items`);
/// anything else passes through unchanged.
fn map_script_digits(text: &str, raised: bool) -> String {
const SUP: [char; 10] = ['⁰', '¹', '²', '³', '⁴', '⁵', '⁶', '⁷', '⁸', '⁹'];
const SUB: [char; 10] = ['₀', '₁', '₂', '₃', '₄', '₅', '₆', '₇', '₈', '₉'];
text.chars()
.map(|c| match c.to_digit(10) {
Some(d) if raised => SUP[d as usize],
Some(d) => SUB[d as usize],
None => c,
})
.collect()
}
/// Helper to get f32 from Object
pub(crate) fn get_number(obj: &Object) -> Option<f32> {
match obj {
@@ -554,7 +773,7 @@ pub(crate) fn get_number(obj: &Object) -> Option<f32> {
mod tests {
use super::*;
use crate::text_utils::{is_cjk_char, is_rtl_char, is_rtl_text, sort_line_items};
use crate::types::{ItemType, TextLine};
use crate::types::{ItemType, PdfLine, TextLine};
use layout::{detect_columns, is_newspaper_layout, ColumnRegion};
fn make_merge_item(text: &str, x: f32, width: f32) -> TextItem {
@@ -569,11 +788,59 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
}
fn with_mcid(mut item: TextItem) -> TextItem {
item.mcid = Some(1);
item
}
fn make_line(x1: f32, y1: f32, x2: f32, y2: f32) -> PdfLine {
PdfLine {
x1,
y1,
x2,
y2,
page: 1,
}
}
#[test]
fn trace_text_preview_truncates_on_char_boundary() {
let text = format!("{}{}tail", "a".repeat(79), '\u{FFFD}');
let preview = trace_text_preview(&text, 80);
assert_eq!(preview.chars().count(), 80);
assert!(text.is_char_boundary(preview.len()));
assert!(preview.ends_with('\u{FFFD}'));
}
#[test]
fn merge_items_breaks_at_style_boundaries() {
// A styled run adjacent to plain text must stay a separate item —
// merging would erase the flags (italic) or stretch the span
// (underline) before markdown emission sees them.
let mut italic = make_merge_item("emphasis", 150.0, 40.0);
italic.is_italic = true;
let mut underlined = make_merge_item("term", 195.0, 20.0);
underlined.is_underline = true;
let items = vec![
make_merge_item("plain lead", 100.0, 48.0),
italic,
underlined,
make_merge_item("plain tail", 218.0, 45.0),
];
let merged = merge_text_items(items);
assert_eq!(merged.len(), 4);
assert!(merged[1].is_italic && !merged[1].is_underline);
assert!(merged[2].is_underline && !merged[2].is_italic);
assert!(!merged[3].is_underline && !merged[3].is_italic);
}
#[test]
fn merge_items_no_space_before_period() {
// Simulate Tc/Tw-adjusted width: "date" width is smaller than the gap
@@ -612,6 +879,133 @@ mod tests {
assert_eq!(merged[0].text, "hello world");
}
#[test]
fn merge_items_preserves_underline_from_later_fragment() {
// Fragments with differing underline stay separate items — OR-merging
// would stretch the eventual `<u>` span over the plain fragment.
// Line-level text assembly still joins them without a space (tight
// gap), so the rendered word is unchanged: `pre<u>fix</u>`.
let mut items = vec![
make_merge_item("pre", 100.0, 18.0),
make_merge_item("fix", 119.0, 18.0),
];
items[1].is_underline = true;
let merged = merge_text_items(items);
assert_eq!(merged.len(), 2);
assert_eq!(merged[0].text, "pre");
assert!(!merged[0].is_underline);
assert_eq!(merged[1].text, "fix");
assert!(merged[1].is_underline);
}
#[test]
fn merge_items_preserves_stream_order_for_backtracking_heading() {
// Some tagged PDFs emit first-letter ActualText fragments, then reset
// the text matrix and draw the rest of the word from the line start.
let items = vec![
with_mcid(make_merge_item("F", 79.4, 4.5)),
with_mcid(make_merge_item("r", 83.9, 3.3)),
with_mcid(make_merge_item("om tables to data-", 79.4, 89.7)),
with_mcid(make_merge_item("", 168.9, 33.9)),
with_mcid(make_merge_item("analytics-", 168.9, 75.5)),
with_mcid(make_merge_item("ready content", 210.5, 60.8)),
];
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(
merged[0].text,
"From tables to data-analytics-ready content"
);
}
#[test]
fn merge_items_preserves_stream_order_for_reset_word_prefix() {
let items = vec![
with_mcid(make_merge_item("N", 68.0, 7.0)),
with_mcid(make_merge_item("e", 75.1, 4.0)),
with_mcid(make_merge_item("w fields created", 68.0, 82.0)),
];
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "New fields created");
}
#[test]
fn merge_items_uses_x_order_for_untagged_backtracking_text() {
let items = vec![
make_merge_item("N", 68.0, 7.0),
make_merge_item("e", 75.1, 4.0),
make_merge_item("w fields created", 68.2, 82.0),
];
let merged = merge_text_items(items);
let texts: Vec<_> = merged.iter().map(|item| item.text.as_str()).collect();
assert_eq!(texts, vec!["N", "w fields created", "e"]);
}
#[test]
fn merge_items_preserves_bullet_stream_order_with_backtracking() {
let items = vec![
with_mcid(make_merge_item("", 79.4, 5.0)),
with_mcid(make_merge_item("The MS", 91.0, 32.6)),
with_mcid(make_merge_item("A LoS project", 84.4, 70.0)),
];
let merged = merge_text_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "• The MSA LoS project");
}
#[test]
fn merge_items_keeps_normal_bullet_gap_limit_without_stream_order() {
let items = vec![
make_merge_item("", 79.4, 5.0),
make_merge_item("Distant item", 91.0, 60.0),
];
let merged = merge_text_items(items);
let texts: Vec<_> = merged.iter().map(|item| item.text.as_str()).collect();
assert_eq!(texts, vec!["", "Distant item"]);
}
#[test]
fn suppress_table_underlines_clears_line_detected_table_items() {
let mut items = vec![
make_merge_item("H1", 125.0, 20.0),
make_merge_item("H2", 225.0, 20.0),
make_merge_item("A", 125.0, 20.0),
make_merge_item("B", 225.0, 20.0),
];
items[0].y = 490.0;
items[1].y = 490.0;
items[2].y = 470.0;
items[3].y = 470.0;
for item in &mut items {
item.is_underline = true;
}
let lines = vec![
make_line(100.0, 500.0, 300.0, 500.0),
make_line(100.0, 480.0, 300.0, 480.0),
make_line(100.0, 460.0, 300.0, 460.0),
make_line(100.0, 460.0, 100.0, 500.0),
make_line(200.0, 460.0, 200.0, 500.0),
make_line(300.0, 460.0, 300.0, 500.0),
];
suppress_table_underlines(&mut items, &[], &lines, 1);
assert!(items.iter().all(|item| !item.is_underline));
}
#[test]
fn test_group_into_lines() {
let items = vec![
@@ -626,6 +1020,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -640,6 +1035,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -654,6 +1050,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -708,6 +1105,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -722,6 +1120,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -736,6 +1135,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -761,6 +1161,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -775,6 +1176,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -789,6 +1191,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -816,6 +1219,7 @@ mod tests {
page: 1,
is_bold: true,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -850,6 +1254,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -885,6 +1290,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -899,6 +1305,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -913,6 +1320,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -935,6 +1343,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1047,6 +1456,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1061,6 +1471,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1085,6 +1496,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1099,6 +1511,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
},
@@ -1139,6 +1552,7 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1183,6 +1597,7 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1227,6 +1642,7 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}],
@@ -1264,6 +1680,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1279,7 +1696,8 @@ mod tests {
];
let merged = merge_subscript_items(items);
assert_eq!(merged.len(), 2);
assert_eq!(merged[0].text, "NH3");
// Lowered baseline → Unicode subscript form (NFKC folds back to "NH3")
assert_eq!(merged[0].text, "NH₃");
assert_eq!(merged[1].text, "Cl");
}
@@ -1293,10 +1711,21 @@ mod tests {
];
let merged = merge_subscript_items(items);
assert_eq!(merged.len(), 2);
assert_eq!(merged[0].text, "H2");
assert_eq!(merged[0].text, "H₂");
assert_eq!(merged[1].text, "O");
}
#[test]
fn test_merge_subscript_items_raised_marker_becomes_superscript() {
// Footnote reference: "word" followed by a RAISED small "2" → word²
let mut marker = make_item_fs("2", 90.0, 502.5, 2.3, 4.7);
marker.y = 502.5; // raised above the 499.0 parent baseline
let items = vec![make_item_fs("word", 78.0, 499.0, 12.0, 8.0), marker];
let merged = merge_subscript_items(items);
assert_eq!(merged.len(), 1);
assert_eq!(merged[0].text, "word²");
}
#[test]
fn test_merge_subscript_items_no_merge_far_gap() {
// Subscript-sized item that's far from the parent should NOT merge
+500
View File
@@ -0,0 +1,500 @@
//! Geometric underline detection.
//!
//! PDFs have no underline font flag — underlines are drawn as separate
//! graphics: stroked horizontal lines (`l`/`S` operators) or thin filled
//! rectangles (`re`/`f`). This pass correlates those graphics with text
//! items after extraction: an item is underlined when a horizontal
//! line/thin rect sits just below its baseline and covers most of its
//! horizontal extent.
//!
//! Repeated same-span rules are treated as table/form rulings rather than
//! underlines, which avoids marking every cell in ruled tables.
use std::collections::HashSet;
use crate::types::{ItemType, PdfRect, TextItem};
/// Max thickness (pt) for a stroked line / filled rect to count as an
/// underline rule rather than a border or decorative band.
const MAX_RULE_THICKNESS: f32 = 2.0;
/// Fraction of the item's width that the rule must cover horizontally.
const MIN_X_OVERLAP: f32 = 0.6;
/// Same-span rules repeated at this many y-levels are usually table/form
/// rulings, not semantic underlines.
const MIN_REPEATED_RULE_LEVELS: usize = 3;
/// Vertical tolerance for considering two rules to be on the same row edge.
const RULE_Y_DEDUP_EPS: f32 = 2.0;
/// Horizontal span similarity required when clustering repeated rulings.
const RULE_SPAN_OVERLAP_RATIO: f32 = 0.8;
const RULE_SPAN_WIDTH_RATIO: f32 = 1.5;
/// Multiple separated rule segments on one row are usually per-column table
/// header/body separators.
const MIN_SEGMENTED_ROW_RULES: usize = 3;
const MIN_SEGMENTED_ROW_GAPS: usize = 2;
const SEGMENTED_ROW_GAP_MIN: f32 = 12.0;
/// A single rule under several widely separated items is usually a table
/// header/body separator, not a sentence underline.
const MIN_TABULAR_RULE_ITEMS: usize = 3;
const MIN_TABULAR_RULE_GAPS: usize = 2;
const TABULAR_RULE_GAP_EM: f32 = 2.0;
#[derive(Clone)]
pub(crate) struct UnderlineLine {
pub(crate) x1: f32,
pub(crate) y1: f32,
pub(crate) x2: f32,
pub(crate) y2: f32,
pub(crate) stroke_width: f32,
pub(crate) page: u32,
}
/// A horizontal rule candidate in page coordinates (PDF y-up).
#[derive(Clone)]
struct Rule {
x1: f32,
x2: f32,
y: f32,
}
impl Rule {
fn width(&self) -> f32 {
self.x2 - self.x1
}
}
fn rules_from_graphics(rects: &[PdfRect], lines: &[UnderlineLine], page: u32) -> Vec<Rule> {
let mut rules: Vec<Rule> = Vec::new();
for l in lines {
if l.page != page {
continue;
}
// Horizontal stroked line (tolerate slight skew).
if l.stroke_width <= MAX_RULE_THICKNESS && (l.y1 - l.y2).abs() <= MAX_RULE_THICKNESS {
let (x1, x2) = if l.x1 <= l.x2 {
(l.x1, l.x2)
} else {
(l.x2, l.x1)
};
if x2 - x1 > 1.0 {
rules.push(Rule {
x1,
x2,
y: (l.y1 + l.y2) / 2.0,
});
}
}
}
for r in rects {
if r.page != page {
continue;
}
// Thin filled rect used as an underline rule. Extents are
// normalized first: `re` operands pass through the CTM, so
// width/height can be negative (flipped axes / negative scale) —
// without normalization negative-width rules are missed and
// negative-height bands sneak past the thickness check.
let (x1, x2) = if r.width >= 0.0 {
(r.x, r.x + r.width)
} else {
(r.x + r.width, r.x)
};
if r.height.abs() <= MAX_RULE_THICKNESS && x2 - x1 > 1.0 {
rules.push(Rule {
x1,
x2,
y: r.y + r.height / 2.0,
});
}
}
rules
}
fn discard_repeated_ruling_rules(rules: Vec<Rule>) -> Vec<Rule> {
if rules.len() < MIN_REPEATED_RULE_LEVELS {
return rules;
}
rules
.iter()
.filter(|rule| {
!is_repeated_ruling_rule(rule, &rules) && !is_segmented_row_ruling_rule(rule, &rules)
})
.cloned()
.collect()
}
fn is_repeated_ruling_rule(rule: &Rule, rules: &[Rule]) -> bool {
let mut y_levels: Vec<f32> = rules
.iter()
.filter(|other| has_similar_span(rule, other))
.map(|other| other.y)
.collect();
y_levels.sort_by(|a, b| a.total_cmp(b));
y_levels.dedup_by(|a, b| (*a - *b).abs() <= RULE_Y_DEDUP_EPS);
y_levels.len() >= MIN_REPEATED_RULE_LEVELS
}
fn is_segmented_row_ruling_rule(rule: &Rule, rules: &[Rule]) -> bool {
let mut row_rules: Vec<&Rule> = rules
.iter()
.filter(|other| (other.y - rule.y).abs() <= RULE_Y_DEDUP_EPS)
.collect();
if row_rules.len() < MIN_SEGMENTED_ROW_RULES {
return false;
}
row_rules.sort_by(|a, b| a.x1.total_cmp(&b.x1));
let large_gaps = row_rules
.windows(2)
.filter(|pair| pair[1].x1 - pair[0].x2 > SEGMENTED_ROW_GAP_MIN)
.count();
large_gaps >= MIN_SEGMENTED_ROW_GAPS
}
fn has_similar_span(a: &Rule, b: &Rule) -> bool {
let a_width = a.width();
let b_width = b.width();
if a_width <= 1.0 || b_width <= 1.0 {
return false;
}
let width_ratio = a_width.max(b_width) / a_width.min(b_width);
if width_ratio > RULE_SPAN_WIDTH_RATIO {
return false;
}
let overlap = a.x2.min(b.x2) - a.x1.max(b.x1);
overlap >= a_width.min(b_width) * RULE_SPAN_OVERLAP_RATIO
}
fn tabular_row_separator_rule_indices(rules: &[Rule], items: &[TextItem]) -> HashSet<usize> {
let mut tabular_rules = HashSet::new();
for (rule_idx, rule) in rules.iter().enumerate() {
let mut matched_items: Vec<&TextItem> = items
.iter()
.filter(|item| is_underline_candidate(item) && rule_matches_item(rule, item))
.collect();
if matched_items.len() < MIN_TABULAR_RULE_ITEMS {
continue;
}
matched_items.sort_by(|a, b| a.x.total_cmp(&b.x));
let large_gaps = matched_items
.windows(2)
.filter(|pair| {
let left = pair[0];
let right = pair[1];
let gap = right.x - (left.x + left.width);
let font_size = left.font_size.max(right.font_size).max(1.0);
gap > font_size * TABULAR_RULE_GAP_EM
})
.count();
if large_gaps >= MIN_TABULAR_RULE_GAPS {
tabular_rules.insert(rule_idx);
}
}
tabular_rules
}
fn is_underline_candidate(item: &TextItem) -> bool {
matches!(item.item_type, ItemType::Text) && !item.text.trim().is_empty() && item.width > 0.0
}
fn rule_matches_item(rule: &Rule, item: &TextItem) -> bool {
// Vertical window: underlines sit at or slightly below the baseline.
// Fonts draw them at roughly 5-15% of the em below; allow up to 35%
// (min 3pt) below and 1pt above for rounding.
let below = (item.font_size * 0.35).max(3.0);
let y_min = item.y - below;
let y_max = item.y + 1.0;
if rule.y < y_min || rule.y > y_max {
return false;
}
let ix1 = item.x;
let ix2 = item.x + item.width;
let min_overlap = item.width * MIN_X_OVERLAP;
let overlap = rule.x2.min(ix2) - rule.x1.max(ix1);
overlap >= min_overlap
}
/// Mark `is_underline` on text items that have a horizontal rule just
/// below their baseline. `items`, `rects`, and `lines` are a single
/// page's extraction output (all in PDF coordinates, y-up, where
/// `TextItem::y` is the text baseline).
pub(crate) fn mark_underlined_items(
items: &mut [TextItem],
rects: &[PdfRect],
lines: &[UnderlineLine],
page: u32,
) {
let rules = discard_repeated_ruling_rules(rules_from_graphics(rects, lines, page));
if rules.is_empty() {
return;
}
let tabular_rules = tabular_row_separator_rule_indices(&rules, items);
for item in items.iter_mut() {
if !is_underline_candidate(item) {
continue;
}
for (rule_idx, rule) in rules.iter().enumerate() {
if tabular_rules.contains(&rule_idx) {
continue;
}
if rule_matches_item(rule, item) {
item.is_underline = true;
break;
}
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::types::ItemType;
fn item(text: &str, x: f32, y: f32, width: f32, font_size: f32) -> TextItem {
TextItem {
text: text.to_string(),
x,
y,
width,
height: font_size,
font: "F1".to_string(),
font_size,
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
}
fn hline(x1: f32, x2: f32, y: f32) -> UnderlineLine {
UnderlineLine {
x1,
y1: y,
x2,
y2: y,
stroke_width: 1.0,
page: 1,
}
}
fn thin_rect(x: f32, y: f32, width: f32) -> PdfRect {
PdfRect {
x,
y,
width,
height: 0.8,
page: 1,
}
}
#[test]
fn stroked_line_under_baseline_marks_underline() {
let mut items = vec![item("underlined", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(99.0, 161.0, 498.5)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items[0].is_underline);
}
#[test]
fn thin_filled_rect_under_baseline_marks_underline() {
let mut items = vec![item("underlined", 100.0, 500.0, 60.0, 10.0)];
let rects = vec![thin_rect(100.0, 497.8, 60.0)];
mark_underlined_items(&mut items, &rects, &[], 1);
assert!(items[0].is_underline);
}
#[test]
fn long_rule_under_multiple_items_marks_each() {
// One underline drawn under a whole sentence: every overlapped
// item gets the flag.
let mut items = vec![
item("first", 100.0, 500.0, 40.0, 10.0),
item("second", 145.0, 500.0, 50.0, 10.0),
];
let lines = vec![hline(98.0, 200.0, 498.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items[0].is_underline);
assert!(items[1].is_underline);
}
#[test]
fn line_far_below_baseline_is_not_an_underline() {
// A horizontal rule 30pt below (section divider) must not mark.
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(90.0, 300.0, 470.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(!items[0].is_underline);
}
#[test]
fn thick_stroked_line_is_not_an_underline() {
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let mut line = hline(99.0, 161.0, 498.5);
line.stroke_width = 4.0;
mark_underlined_items(&mut items, &[], &[line], 1);
assert!(!items[0].is_underline);
}
#[test]
fn line_above_baseline_is_not_an_underline() {
// Strikethrough / overline geometry must not mark.
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![hline(90.0, 300.0, 505.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(!items[0].is_underline);
}
#[test]
fn insufficient_horizontal_overlap_is_not_an_underline() {
// Rule under only a quarter of the item (e.g. neighboring cell
// border) must not mark.
let mut items = vec![item("wide text item", 100.0, 500.0, 100.0, 10.0)];
let lines = vec![hline(100.0, 125.0, 498.5)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(!items[0].is_underline);
}
#[test]
fn negative_width_rect_is_normalized_and_marks_underline() {
// A CTM with negative x-scale (or negative `re` operands) produces
// rects whose width is negative; the rule extents must normalize.
let mut items = vec![item("underlined", 100.0, 500.0, 60.0, 10.0)];
let rects = vec![PdfRect {
x: 160.0,
y: 497.8,
width: -60.0,
height: 0.8,
page: 1,
}];
mark_underlined_items(&mut items, &rects, &[], 1);
assert!(items[0].is_underline);
}
#[test]
fn negative_height_band_is_not_an_underline() {
// A 14pt band expressed with negative height must not pass the
// thickness check via sign trickery.
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let rects = vec![PdfRect {
x: 95.0,
y: 509.0,
width: 80.0,
height: -14.0,
page: 1,
}];
mark_underlined_items(&mut items, &rects, &[], 1);
assert!(!items[0].is_underline);
}
#[test]
fn thick_band_is_not_an_underline() {
// A highlight bar / filled cell background (tall rect) must not mark.
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let rects = vec![PdfRect {
x: 95.0,
y: 495.0,
width: 80.0,
height: 14.0,
page: 1,
}];
mark_underlined_items(&mut items, &rects, &[], 1);
assert!(!items[0].is_underline);
}
#[test]
fn vertical_line_is_not_an_underline() {
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let lines = vec![UnderlineLine {
x1: 120.0,
y1: 498.0,
x2: 120.0,
y2: 400.0,
stroke_width: 1.0,
page: 1,
}];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(!items[0].is_underline);
}
#[test]
fn other_pages_graphics_do_not_mark() {
let mut items = vec![item("text", 100.0, 500.0, 60.0, 10.0)];
let mut line = hline(99.0, 161.0, 498.5);
line.page = 2;
mark_underlined_items(&mut items, &[], &[line], 1);
assert!(!items[0].is_underline);
}
#[test]
fn repeated_table_row_rules_do_not_mark_cell_text() {
let mut items = vec![
item("A", 110.0, 500.0, 20.0, 10.0),
item("B", 110.0, 480.0, 20.0, 10.0),
item("C", 110.0, 460.0, 20.0, 10.0),
];
let lines = vec![
hline(100.0, 150.0, 498.0),
hline(100.0, 150.0, 478.0),
hline(100.0, 150.0, 458.0),
];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items.iter().all(|item| !item.is_underline));
}
#[test]
fn row_separator_under_spaced_column_labels_is_not_an_underline() {
let mut items = vec![
item("Date", 100.0, 500.0, 25.0, 10.0),
item("Rate", 200.0, 500.0, 25.0, 10.0),
item("Yield", 300.0, 500.0, 30.0, 10.0),
];
let lines = vec![hline(90.0, 340.0, 498.0)];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items.iter().all(|item| !item.is_underline));
}
#[test]
fn same_row_spaced_rule_segments_do_not_mark_column_labels() {
let mut items = vec![
item("Date", 100.0, 500.0, 25.0, 10.0),
item("Rate", 200.0, 500.0, 25.0, 10.0),
item("Yield", 300.0, 500.0, 30.0, 10.0),
];
let lines = vec![
hline(98.0, 128.0, 498.0),
hline(198.0, 228.0, 498.0),
hline(298.0, 333.0, 498.0),
];
mark_underlined_items(&mut items, &[], &lines, 1);
assert!(items.iter().all(|item| !item.is_underline));
}
}
+3
View File
@@ -294,6 +294,7 @@ fn extract_form_xobject_text_inner(
page: page_num,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Image,
mcid: None,
});
@@ -438,6 +439,7 @@ fn extract_form_xobject_text_inner(
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -586,6 +588,7 @@ fn extract_form_xobject_text_inner(
page: page_num,
is_bold: is_bold_font(base_font),
is_italic: is_italic_font(base_font),
is_underline: false,
item_type: ItemType::Text,
mcid: None,
});
+601 -48
View File
@@ -62,10 +62,23 @@ use std::collections::{BTreeMap, HashMap, HashSet};
use std::path::Path;
use tounicode::FontCMaps;
/// OCR reason emitted when the extracted text layer appears garbled due to
/// broken font decoding or mojibake.
pub const OCR_REASON_SUSPECTED_GARBLED_TEXT: &str = "suspected_garbled_text";
// =========================================================================
// Result type
// =========================================================================
/// OCR reasons for a single 1-indexed page.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct PageOcrReasons {
/// 1-indexed page number.
pub page: u32,
/// Machine-readable OCR reason identifiers.
pub reasons: Vec<String>,
}
/// High-level PDF processing result.
#[derive(Debug)]
pub struct PdfProcessResult {
@@ -79,6 +92,8 @@ pub struct PdfProcessResult {
pub processing_time_ms: u64,
/// 1-indexed page numbers that need OCR.
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
pub ocr_reasons_by_page: Vec<PageOcrReasons>,
/// Title from PDF metadata (if available).
pub title: Option<String>,
/// Detection confidence score (0.01.0).
@@ -322,6 +337,8 @@ pub struct PageMarkdown {
/// `true` when text on this page is unreliable (GID-encoded fonts,
/// encoding issues, garbage text, or empty extraction).
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
pub ocr_reason: Option<String>,
}
/// Combined per-page markdown extraction and layout classification result.
@@ -335,6 +352,8 @@ pub struct PagesExtractionResult {
pub pages_with_columns: Vec<u32>,
/// 1-indexed pages that need OCR (scanned/image-based).
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
pub ocr_reasons_by_page: Vec<PageOcrReasons>,
/// True if any page has tables or columns.
pub is_complex: bool,
}
@@ -368,6 +387,7 @@ pub fn extract_pages_markdown_mem(
// Extract ALL pages to get accurate, document-wide font stats.
let ((all_items, all_rects, all_lines), page_thresholds, gid_pages) =
extractor::extract_positioned_text_from_doc(&doc, &font_cmaps, None)?;
let text_quality = analyze_text_quality(&all_items);
// Compute layout complexity from full document (near-zero cost).
let complexity = compute_layout_complexity(&all_items, &all_rects, &all_lines);
@@ -387,6 +407,7 @@ pub fn extract_pages_markdown_mem(
let mut results = Vec::with_capacity(pages_slice.len());
let mut pages_needing_ocr = Vec::new();
let mut ocr_reasons_by_page = BTreeMap::new();
for &page_0idx in pages_slice {
// Out-of-range pages → empty + needs_ocr
@@ -396,6 +417,7 @@ pub fn extract_pages_markdown_mem(
page: page_0idx,
markdown: String::new(),
needs_ocr: true,
ocr_reason: None,
});
continue;
}
@@ -416,6 +438,7 @@ pub fn extract_pages_markdown_mem(
.collect();
let has_gid = gid_pages.contains(&page_1idx);
let has_text_quality_issue = text_quality.pages_needing_ocr.contains(&page_1idx);
// Build markdown with document-wide font stats
let options = MarkdownOptions {
@@ -425,21 +448,33 @@ pub fn extract_pages_markdown_mem(
..MarkdownOptions::default()
};
let md = markdown::to_markdown_from_items_with_rects_and_lines(
page_items,
options,
&page_rects,
&[],
&page_thresholds,
None,
&[],
);
let md = if has_text_quality_issue {
String::new()
} else {
markdown::to_markdown_from_items_with_rects_and_lines(
page_items,
options,
&page_rects,
&[],
&page_thresholds,
None,
&[],
)
};
let needs_ocr = md.trim().is_empty()
|| has_gid
|| is_garbage_text(&md)
|| is_cid_garbage(&md)
|| detect_encoding_issues(&md);
let has_decoding_issue = has_text_quality_issue
|| (!md.is_empty() && (is_cid_garbage(&md) || detect_encoding_issues(&md)));
if has_decoding_issue {
add_ocr_reason(
&mut ocr_reasons_by_page,
page_1idx,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
let ocr_reason = page_ocr_reason(&ocr_reasons_by_page, page_1idx);
let needs_ocr =
ocr_reason.is_some() || md.trim().is_empty() || has_gid || is_garbage_text(&md);
if needs_ocr {
pages_needing_ocr.push(page_1idx);
@@ -449,6 +484,7 @@ pub fn extract_pages_markdown_mem(
page: page_0idx,
markdown: if needs_ocr { String::new() } else { md },
needs_ocr,
ocr_reason,
});
}
@@ -457,6 +493,7 @@ pub fn extract_pages_markdown_mem(
pages_with_tables: complexity.pages_with_tables,
pages_with_columns: complexity.pages_with_columns,
pages_needing_ocr,
ocr_reasons_by_page: page_ocr_reasons_vec(ocr_reasons_by_page),
is_complex: complexity.is_complex,
})
}
@@ -488,6 +525,8 @@ pub struct RegionText {
/// Set when: the region is empty, the page uses GID-encoded fonts, or the
/// extracted text fails garbage/encoding checks.
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
pub ocr_reason: Option<String>,
}
/// Result for a page's region extractions.
@@ -592,29 +631,36 @@ pub fn extract_text_in_regions_mem(
for rect in regions {
let [rx1, ry1, rx2, ry2] = *rect;
let text = match items {
Some(items) => collect_text_in_region_with_options(
items,
rx1,
ry1,
rx2,
ry2,
page_h,
coords,
adaptive_threshold,
),
None => String::new(),
let bounds = region_bounds(rx1, ry1, rx2, ry2, page_h, coords);
let matched: Vec<TextItem> = match items {
Some(items) => items
.iter()
.filter(|item| region_overlaps_item(item, bounds))
.cloned()
.collect(),
None => Vec::new(),
};
let has_text_quality_issue = region_items_have_decoding_issue(&matched);
let text = collect_text_from_matched_items(matched, adaptive_threshold);
let has_cid_issue = is_cid_garbage(&text);
let has_encoding_issue = detect_encoding_issues(&text);
let ocr_reason = if has_text_quality_issue || has_cid_issue || has_encoding_issue {
Some(suspected_garbled_reason())
} else {
None
};
// Check per-region text quality instead of blanket page-level
// GID rejection. A GID font in a logo elsewhere on the page
// shouldn't force GPU OCR for clean text regions.
let needs_ocr = text.trim().is_empty()
|| is_garbage_text(&text)
|| is_cid_garbage(&text)
|| detect_encoding_issues(&text);
let needs_ocr =
ocr_reason.is_some() || text.trim().is_empty() || is_garbage_text(&text);
page_results.push(RegionText { text, needs_ocr });
page_results.push(RegionText {
text,
needs_ocr,
ocr_reason,
});
}
results.push(PageRegionResult {
@@ -725,6 +771,16 @@ pub fn extract_tables_in_regions_mem(
page_results.push(RegionText {
text: String::new(),
needs_ocr: true,
ocr_reason: None,
});
continue;
}
if region_items_have_decoding_issue(&matched) {
page_results.push(RegionText {
text: String::new(),
needs_ocr: true,
ocr_reason: Some(suspected_garbled_reason()),
});
continue;
}
@@ -899,10 +955,12 @@ pub fn extract_tables_in_regions_mem(
Some(candidate) => page_results.push(RegionText {
text: candidate.markdown.clone(),
needs_ocr: false,
ocr_reason: None,
}),
None => page_results.push(RegionText {
text: String::new(),
needs_ocr: true,
ocr_reason: None,
}),
}
}
@@ -1230,6 +1288,19 @@ mod vector_grid_tests {
);
}
#[test]
fn td9264_insurance_prose_not_rect_table() {
let tables = detect_rect_tables_in_fixture_page("tests/fixtures/td9264.pdf", 4);
assert!(
tables.is_empty(),
"expected no rect-detected tables for the insurance-company prose; got {:?}",
tables
.iter()
.map(|t| (t.rows.len(), t.columns.len(), t.cells.clone()))
.collect::<Vec<_>>()
);
}
/// Wireless table regression: decorative/text-region rects may provide row
/// bands, but without a real rect-derived column scaffold they must not be
/// accepted as a vector grid.
@@ -3316,6 +3387,7 @@ fn process_document(
page_count,
processing_time_ms: start.elapsed().as_millis() as u64,
pages_needing_ocr,
ocr_reasons_by_page: Vec::new(),
title,
confidence,
layout: LayoutComplexity::default(),
@@ -3331,6 +3403,7 @@ fn process_document(
page_count,
processing_time_ms: start.elapsed().as_millis() as u64,
pages_needing_ocr,
ocr_reasons_by_page: Vec::new(),
title,
confidence,
layout: LayoutComplexity::default(),
@@ -3401,8 +3474,17 @@ fn process_document(
})
.unwrap_or((None, Vec::new()));
let (markdown, layout, has_encoding_issues, gid_pages) = match extracted {
let (
markdown,
layout,
has_encoding_issues,
gid_pages,
text_quality_pages,
text_quality_reasons_by_page,
) = match extracted {
Some(((items, rects, lines), page_thresholds, gid_encoded_pages)) => {
let mut ocr_reasons_by_page = BTreeMap::new();
// For TextBased PDFs with pages flagged for OCR (Identity-H or
// Type3 fonts without ToUnicode), check whether the CID-as-Unicode
// passthrough actually produced readable text. If a page's text
@@ -3435,6 +3517,13 @@ fn process_document(
"suppressing garbage text from OCR-flagged pages: {:?}",
garbage_pages
);
for page in &garbage_pages {
add_ocr_reason(
&mut ocr_reasons_by_page,
*page,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
let items: Vec<_> = items
.into_iter()
.filter(|i| !garbage_pages.contains(&i.page))
@@ -3451,6 +3540,8 @@ fn process_document(
}
};
let text_quality = analyze_text_quality(&items);
merge_ocr_reasons(&mut ocr_reasons_by_page, text_quality.reasons_by_page);
let layout = compute_layout_complexity(&items, &rects, &lines);
let md = if options.mode == ProcessMode::Analyze {
@@ -3467,14 +3558,25 @@ fn process_document(
))
};
let enc = md.as_ref().is_some_and(|m| detect_encoding_issues(m));
(md, layout, enc, gid_encoded_pages)
let enc = !ocr_reasons_by_page.is_empty()
|| text_quality.has_encoding_issues
|| md.as_ref().is_some_and(|m| detect_encoding_issues(m));
(
md,
layout,
enc,
gid_encoded_pages,
text_quality.pages_needing_ocr,
ocr_reasons_by_page,
)
}
None => (
None,
LayoutComplexity::default(),
false,
std::collections::HashSet::new(),
Vec::new(),
BTreeMap::new(),
),
};
@@ -3516,6 +3618,19 @@ fn process_document(
}
pages_needing_ocr.sort_unstable();
}
if !text_quality_pages.is_empty() {
log::debug!(
"pages with OCR reason {} (need OCR): {:?}",
OCR_REASON_SUSPECTED_GARBLED_TEXT,
text_quality_pages
);
for page in text_quality_pages {
if !pages_needing_ocr.contains(&page) {
pages_needing_ocr.push(page);
}
}
pages_needing_ocr.sort_unstable();
}
// Detect sparse extraction: when a TEXT-BASED PDF produces very few
// characters per page, the text is likely embedded in images/forms
@@ -3554,6 +3669,7 @@ fn process_document(
page_count,
processing_time_ms: start.elapsed().as_millis() as u64,
pages_needing_ocr,
ocr_reasons_by_page: page_ocr_reasons_vec(text_quality_reasons_by_page),
title,
confidence,
layout,
@@ -3581,6 +3697,10 @@ fn detect_encoding_issues(markdown: &str) -> bool {
}
// Heuristic 2: dollar-as-space pattern
has_dollar_as_space_pattern(markdown)
}
fn has_dollar_as_space_pattern(markdown: &str) -> bool {
let total_dollars = markdown.matches('$').count();
if total_dollars > 10 {
let bytes = markdown.as_bytes();
@@ -3601,6 +3721,242 @@ fn detect_encoding_issues(markdown: &str) -> bool {
false
}
#[derive(Debug, Default)]
struct TextQualityReport {
pages_needing_ocr: Vec<u32>,
has_encoding_issues: bool,
reasons_by_page: BTreeMap<u32, Vec<String>>,
}
#[derive(Debug, Default)]
struct PageTextQualityEvidence {
chars: usize,
replacement_chars: usize,
replacement_spans: usize,
longest_replacement_run: usize,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum TextSpanIssueKind {
Replacement,
Strong,
}
fn analyze_text_quality(items: &[TextItem]) -> TextQualityReport {
let mut reasons_by_page = BTreeMap::new();
let mut evidence_by_page = BTreeMap::<u32, PageTextQualityEvidence>::new();
for item in items {
if !matches!(item.item_type, crate::types::ItemType::Text) {
continue;
}
let evidence = evidence_by_page.entry(item.page).or_default();
evidence.chars += item.text.chars().filter(|ch| !ch.is_whitespace()).count();
match text_span_decoding_issue_kind(&item.text) {
Some(TextSpanIssueKind::Strong) => {
add_ocr_reason(
&mut reasons_by_page,
item.page,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
Some(TextSpanIssueKind::Replacement) => {
let stats = replacement_text_stats(&item.text);
evidence.replacement_chars += stats.0;
evidence.replacement_spans += 1;
evidence.longest_replacement_run = evidence.longest_replacement_run.max(stats.1);
}
None => {}
}
}
for (page, evidence) in evidence_by_page {
if reasons_by_page.contains_key(&page) {
continue;
}
if page_replacement_evidence_needs_ocr(&evidence) {
add_ocr_reason(
&mut reasons_by_page,
page,
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
}
let pages_needing_ocr: Vec<u32> = reasons_by_page.keys().copied().collect();
TextQualityReport {
has_encoding_issues: !pages_needing_ocr.is_empty(),
pages_needing_ocr,
reasons_by_page,
}
}
fn suspected_garbled_reason() -> String {
OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()
}
fn add_ocr_reason(reasons_by_page: &mut BTreeMap<u32, Vec<String>>, page: u32, reason: &str) {
let reasons = reasons_by_page.entry(page).or_default();
if !reasons.iter().any(|existing| existing == reason) {
reasons.push(reason.to_string());
}
}
fn merge_ocr_reasons(
reasons_by_page: &mut BTreeMap<u32, Vec<String>>,
extra_reasons_by_page: BTreeMap<u32, Vec<String>>,
) {
for (page, reasons) in extra_reasons_by_page {
for reason in reasons {
add_ocr_reason(reasons_by_page, page, &reason);
}
}
}
fn page_ocr_reason(reasons_by_page: &BTreeMap<u32, Vec<String>>, page: u32) -> Option<String> {
reasons_by_page
.get(&page)
.and_then(|reasons| reasons.first())
.cloned()
}
fn page_ocr_reasons_vec(reasons_by_page: BTreeMap<u32, Vec<String>>) -> Vec<PageOcrReasons> {
reasons_by_page
.into_iter()
.map(|(page, reasons)| PageOcrReasons { page, reasons })
.collect()
}
fn region_items_have_decoding_issue(items: &[TextItem]) -> bool {
items.iter().any(|item| {
matches!(item.item_type, crate::types::ItemType::Text)
&& text_span_has_decoding_issue(&item.text)
})
}
fn text_span_has_decoding_issue(text: &str) -> bool {
text_span_decoding_issue_kind(text).is_some()
}
fn text_span_decoding_issue_kind(text: &str) -> Option<TextSpanIssueKind> {
let text = text.trim();
if text.is_empty() {
return None;
}
if has_dollar_as_space_pattern(text)
|| has_private_use_text_run(text)
|| is_cid_garbage(text)
|| has_cid_control_token(text)
{
return Some(TextSpanIssueKind::Strong);
}
if has_replacement_text_run(text) {
return Some(TextSpanIssueKind::Replacement);
}
None
}
fn replacement_text_stats(text: &str) -> (usize, usize) {
let mut replacement = 0usize;
let mut current_run = 0usize;
let mut longest_run = 0usize;
for ch in text.chars() {
if ch == '\u{FFFD}' {
replacement += 1;
current_run += 1;
longest_run = longest_run.max(current_run);
} else {
current_run = 0;
}
}
(replacement, longest_run)
}
fn page_replacement_evidence_needs_ocr(evidence: &PageTextQualityEvidence) -> bool {
if evidence.replacement_chars == 0 || evidence.chars == 0 {
return false;
}
// If the entire page is only a short broken text layer, even a short
// replacement run is enough evidence. On otherwise text-heavy pages,
// require density so math formulas do not force full-page OCR.
if evidence.chars <= 80 && evidence.longest_replacement_run >= 2 {
return true;
}
let replacement_density_bps = evidence.replacement_chars * 10_000 / evidence.chars;
let enough_bad_text = evidence.replacement_chars >= 12 && replacement_density_bps >= 500;
let repeated_bad_spans = evidence.replacement_spans >= 3 && replacement_density_bps >= 250;
let long_bad_run = evidence.longest_replacement_run >= 8 && replacement_density_bps >= 250;
enough_bad_text || repeated_bad_spans || long_bad_run
}
fn has_replacement_text_run(text: &str) -> bool {
let (replacement, longest_run) = replacement_text_stats(text);
longest_run >= 2 || replacement >= 3
}
fn has_private_use_text_run(text: &str) -> bool {
let mut total = 0usize;
let mut private_use = 0usize;
let mut current_run = 0usize;
let mut longest_run = 0usize;
for ch in text.chars() {
if ch.is_whitespace() {
current_run = 0;
continue;
}
total += 1;
if is_private_use_char(ch) {
private_use += 1;
current_run += 1;
longest_run = longest_run.max(current_run);
} else {
current_run = 0;
}
}
if private_use == 0 {
return false;
}
longest_run >= 3 || (total >= 5 && private_use >= 2 && private_use * 2 >= total)
}
fn has_cid_control_token(text: &str) -> bool {
text.split_whitespace().any(token_has_cid_control)
}
fn token_has_cid_control(token: &str) -> bool {
let mut total = 0usize;
let mut c1_control = 0usize;
for ch in token.chars() {
total += 1;
if ('\u{0080}'..='\u{009F}').contains(&ch) {
c1_control += 1;
}
}
total >= 5 && c1_control >= 2 && c1_control * 20 >= total
}
fn is_private_use_char(ch: char) -> bool {
matches!(
ch as u32,
0xE000..=0xF8FF | 0xF0000..=0xFFFFD | 0x100000..=0x10FFFD
)
}
/// Check if extracted text is predominantly garbage (non-alphanumeric).
///
/// Broken font encodings produce text like "----1-.-.-.___ --.-. .._ I_---."
@@ -3609,20 +3965,36 @@ fn detect_encoding_issues(markdown: &str) -> bool {
fn is_garbage_text(markdown: &str) -> bool {
let mut alphanum = 0usize;
let mut non_alphanum = 0usize;
for ch in markdown.chars() {
if ch.is_whitespace() {
continue;
let chars: Vec<char> = markdown.chars().collect();
let mut i = 0usize;
while i < chars.len() {
let ch = chars[i];
let mut run_end = i + 1;
while run_end < chars.len() && chars[run_end] == ch {
run_end += 1;
}
// Skip markdown syntax chars that we add (not from the PDF)
if matches!(ch, '#' | '*' | '|' | '-' | '\n') {
continue;
}
if ch.is_alphanumeric() {
alphanum += 1;
} else {
non_alphanum += 1;
let is_decorative_leader = matches!(ch, '.' | '_' | '·') && run_end - i >= 3;
if !is_decorative_leader {
for &run_ch in &chars[i..run_end] {
if run_ch.is_whitespace() {
continue;
}
// Skip markdown syntax chars that we add (not from the PDF)
if matches!(run_ch, '#' | '*' | '|' | '-' | '\n') {
continue;
}
if run_ch.is_alphanumeric() {
alphanum += 1;
} else {
non_alphanum += 1;
}
}
}
i = run_end;
}
let total = alphanum + non_alphanum;
total >= 50 && alphanum * 2 < total
}
@@ -3647,6 +4019,9 @@ fn is_cid_garbage(text: &str) -> bool {
}
total += 1;
// C1 control characters (U+0080U+009F) — almost never in real text
if ch == '·' {
continue;
}
if ('\u{0080}'..='\u{009F}').contains(&ch) {
c1_control += 1;
}
@@ -3661,14 +4036,16 @@ fn is_cid_garbage(text: &str) -> bool {
return false;
}
// If ≥5% of non-whitespace chars are C1 controls, it's garbage
if c1_control * 20 >= total {
if c1_control >= 2 && c1_control * 20 >= total {
return true;
}
// If ≥40% of non-whitespace chars are high Latin-1 AND the text has few
// ASCII letters, it's likely CID-as-Latin-1 mojibake (Japanese/CJK PDFs
// where CID values 0x80-0xFF become accented Latin characters).
// where CID values 0x80-0xFF become accented Latin characters). Keep a
// minimum length so short math tokens like "2×()×" do not route a clean
// page to OCR.
let ascii_letters = text.chars().filter(|c| c.is_ascii_alphabetic()).count();
high_latin * 5 >= total * 2 && ascii_letters * 3 < total
total >= 20 && high_latin * 5 >= total * 2 && ascii_letters * 3 < total
}
/// Detect markdown tables with suspicious structure that suggest the heuristic
@@ -4518,6 +4895,7 @@ mod text_cluster_column_undercount_tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -4792,6 +5170,7 @@ mod table_candidate_selection_tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -5538,11 +5917,19 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
}
fn test_text_item_on_page(page: u32, text: &str) -> TextItem {
TextItem {
page,
..test_item(text, 10.0, 10.0, text.len() as f32 * 5.0, 12.0)
}
}
#[test]
fn test_detect_encoding_issues_fffd() {
assert!(detect_encoding_issues(
@@ -5578,6 +5965,160 @@ mod tests {
assert!(!detect_encoding_issues(text));
}
#[test]
fn test_text_quality_flags_localized_cid_mojibake_span() {
let items = vec![
test_text_item_on_page(
1,
"Waiting Period 等待期 Maternity and newborn infant care benefit",
),
test_text_item_on_page(
1,
"Inpatient and Day-care Benefits DÂB\u{009B}A4gÉ9¶0ÅDÂB\u{009B}Ê(D>öBÑ9¯",
),
test_text_item_on_page(1, "Covered up to annual maximum. 赔付至年度最高保额。"),
test_text_item_on_page(2, "A clean second page should not be routed to OCR."),
];
let quality = analyze_text_quality(&items);
assert!(quality.has_encoding_issues);
assert_eq!(quality.pages_needing_ocr, vec![1]);
assert_eq!(
quality.reasons_by_page.get(&1).cloned(),
Some(vec![OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()])
);
}
#[test]
fn test_text_quality_flags_replacement_and_private_use_runs() {
let items = vec![
test_text_item_on_page(1, "broken \u{FFFD}\u{FFFD} text"),
test_text_item_on_page(3, "\u{E000}\u{E001}\u{E002}"),
];
let quality = analyze_text_quality(&items);
assert_eq!(quality.pages_needing_ocr, vec![1, 3]);
assert_eq!(
quality.reasons_by_page.get(&1).cloned(),
Some(vec![OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()])
);
assert_eq!(
quality.reasons_by_page.get(&3).cloned(),
Some(vec![OCR_REASON_SUSPECTED_GARBLED_TEXT.to_string()])
);
}
#[test]
fn test_text_quality_allows_clean_multilingual_and_latin1_text() {
let items = vec![
test_text_item_on_page(1, "你好世界,这是一段正常的中文文本。"),
test_text_item_on_page(1, "Résumé déjà vu: façade, São Paulo, año 2026."),
test_text_item_on_page(1, "A single icon \u{E000} should not force OCR."),
];
let quality = analyze_text_quality(&items);
assert!(!quality.has_encoding_issues);
assert!(quality.pages_needing_ocr.is_empty());
assert!(quality.reasons_by_page.is_empty());
}
#[test]
fn test_text_quality_allows_toc_leaders_form_rules_and_short_math() {
let items = vec![
test_text_item_on_page(
1,
"Feature Overview ........................................................................................................ 1-5",
),
test_text_item_on_page(1, "Signature __________________________________________"),
test_text_item_on_page(1, "__________________________________________________"),
test_text_item_on_page(1, "2×()×"),
];
let quality = analyze_text_quality(&items);
assert!(!quality.has_encoding_issues);
assert!(quality.pages_needing_ocr.is_empty());
}
#[test]
fn test_text_quality_allows_isolated_replacement_character() {
let items = vec![
test_text_item_on_page(1, "\u{FFFD}2026 FINRA"),
test_text_item_on_page(1, "A mostly clean page should not be sent to OCR."),
];
let quality = analyze_text_quality(&items);
assert!(!quality.has_encoding_issues);
assert!(quality.pages_needing_ocr.is_empty());
}
#[test]
fn test_text_quality_allows_formula_replacement_on_clean_page() {
let items = vec![
test_text_item_on_page(
1,
"The LCOE of a power plant can be decomposed into three parts and described in prose.",
),
test_text_item_on_page(
1,
"This page has enough normal text that a damaged equation should not force OCR.",
),
test_text_item_on_page(1, "\u{FFFD}\u{FFFD}\u{FFFD}\u{FFFD} = x + y"),
test_text_item_on_page(1, "More normal explanatory text follows after the formula."),
];
let quality = analyze_text_quality(&items);
assert!(!quality.has_encoding_issues);
assert!(quality.pages_needing_ocr.is_empty());
}
#[test]
fn test_text_quality_flags_dense_replacement_text_page() {
let items = vec![
test_text_item_on_page(1, "\u{FFFD}\u{FFFD}\u{FFFD}\u{FFFD} broken layer"),
test_text_item_on_page(1, "more \u{FFFD}\u{FFFD}\u{FFFD}\u{FFFD} broken text"),
test_text_item_on_page(1, "\u{FFFD}\u{FFFD}\u{FFFD}\u{FFFD}"),
];
let quality = analyze_text_quality(&items);
assert_eq!(quality.pages_needing_ocr, vec![1]);
assert!(quality.has_encoding_issues);
}
#[test]
fn test_text_quality_allows_tex_ligature_c1_controls_in_words() {
let items = vec![test_text_item_on_page(
1,
"Especialmente de\u{85}ciente es nuestro conocimiento del control. The amniotic \u{87}uid is important.",
)];
let quality = analyze_text_quality(&items);
assert!(!quality.has_encoding_issues);
assert!(quality.pages_needing_ocr.is_empty());
}
#[test]
fn test_region_text_quality_is_scoped_to_matched_items() {
let clean_region = vec![
test_text_item_on_page(1, "Clean native text"),
test_text_item_on_page(1, "Résumé déjà vu"),
];
let garbled_region = vec![
test_text_item_on_page(1, "Clean prefix"),
test_text_item_on_page(1, "DÂB\u{009B}A4gÉ9¶0ÅDÂB\u{009B}Ê(D>öBÑ9¯"),
];
assert!(!region_items_have_decoding_issue(&clean_region));
assert!(region_items_have_decoding_issue(&garbled_region));
}
#[test]
fn test_garbage_text_detection() {
// Simulates garbage output from Identity-H fonts without ToUnicode.
@@ -5621,6 +6162,18 @@ mod tests {
!is_cid_garbage(japanese),
"Valid Japanese text should not be flagged as garbage"
);
let japanese_toc = "第1章 市政経営方針の位置づけ ···································· 1";
assert!(
!is_cid_garbage(japanese_toc),
"Japanese TOC dot leaders should not be flagged as garbage"
);
let tex_ligature = "amniotic \u{87}uid volume regulation";
assert!(
!is_cid_garbage(tex_ligature),
"A single TeX ligature byte inside a word should not be CID garbage"
);
}
#[test]
+12 -4
View File
@@ -131,10 +131,11 @@ pub(crate) fn format_list_item(text: &str) -> String {
if let Some(rest) = trimmed.strip_prefix(*bullet) {
return format!("- {}", rest.trim_start());
}
// Bullet inside a leading bold/italic run (e.g. "**● Label:** rest").
// The run wraps both the marker and the following label because both
// use a bold font in the PDF.
for wrapper in ["**", "*"] {
// Bullet inside a leading style run (e.g. "**● Label:** rest" or
// "<u>● Label</u>"). The run wraps both the marker and the following
// label because both carry the style in the PDF. The marker must move
// outside the wrapper so markdown still sees a list item.
for wrapper in ["**", "*", "<u>"] {
if let Some(after_open) = trimmed.strip_prefix(wrapper) {
if let Some(rest) = after_open.strip_prefix(*bullet) {
return format!("- {}{}", wrapper, rest.trim_start());
@@ -235,6 +236,13 @@ mod tests {
assert_eq!(format_list_item("• Item"), "- Item");
}
#[test]
fn format_list_item_bullet_inside_underline() {
// Fully-underlined bullet line: the marker must move outside the
// <u> wrapper so markdown still renders a list item.
assert_eq!(format_list_item("<u>● Item text</u>"), "- <u>Item text</u>");
}
#[test]
fn format_list_item_bullet_inside_bold() {
// PDF that uses bold font for both the marker and the label produces
+258 -8
View File
@@ -149,6 +149,79 @@ fn find_isolated_lines(lines: &[TextLine], base_size: f32, para_threshold: f32)
set
}
/// Pre-scan body-size all-bold runs that are too long to be headings.
///
/// Some academic PDFs use an all-bold abstract/summary paragraph immediately
/// after the author block. A line-local bold heading heuristic sees each
/// wrapped visual line as "standalone" once the first line is misclassified,
/// producing a stack of `##` headings. Multi-line body-size bold runs with a
/// paragraph-sized word count should stay paragraph text.
fn find_wrapped_bold_paragraph_lines(
lines: &[TextLine],
base_size: f32,
para_threshold: f32,
) -> HashSet<usize> {
let mut set = HashSet::new();
let mut i = 0usize;
while i < lines.len() {
if !is_body_size_all_bold_line(&lines[i], base_size) {
i += 1;
continue;
}
let start = i;
let mut end = i;
let mut word_count = lines[i].text().split_whitespace().count();
while end + 1 < lines.len()
&& is_body_size_all_bold_line(&lines[end + 1], base_size)
&& is_wrapped_same_style_line(&lines[end], &lines[end + 1], para_threshold)
{
end += 1;
word_count += lines[end].text().split_whitespace().count();
}
let line_count = end - start + 1;
if line_count >= 3 && word_count > 20 {
for idx in start..=end {
set.insert(idx);
}
}
i = end + 1;
}
set
}
fn is_body_size_all_bold_line(line: &TextLine, base_size: f32) -> bool {
let Some(first) = line.items.first() else {
return false;
};
first.font_size >= base_size * 0.95
&& first.font_size < base_size * 1.2
&& line
.items
.iter()
.all(|item| item.is_bold && (item.font_size - first.font_size).abs() < 0.5)
}
fn is_wrapped_same_style_line(prev: &TextLine, next: &TextLine, para_threshold: f32) -> bool {
if prev.page != next.page {
return false;
}
let y_gap = prev.y - next.y;
if !(y_gap > 0.0 && y_gap <= para_threshold) {
return false;
}
let prev_x = prev.items.first().map(|item| item.x).unwrap_or(0.0);
let next_x = next.items.first().map(|item| item.x).unwrap_or(0.0);
(prev_x - next_x).abs() <= 40.0
}
/// Resolve the dominant structure role for a text line by looking up its items' MCIDs.
///
/// Returns the first non-container role found (skipping Document/Part/Sect/Div/NonStruct/Span).
@@ -397,6 +470,8 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
// between paragraphs at body font size. Inspired by opendataloader's
// lookahead in HeadingProcessor (prevNode/nextNode context).
let isolated_lines = find_isolated_lines(&lines, base_size, para_threshold);
let wrapped_bold_paragraph_lines =
find_wrapped_bold_paragraph_lines(&lines, base_size, para_threshold);
// Detect struct heading levels that are overused (body text mistagged as headings)
let overused_heading_levels = detect_overused_struct_heading_levels(&lines, struct_roles);
@@ -410,6 +485,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
let mut last_list_x: Option<f32> = None;
let mut in_code_block = false;
let mut prev_had_dot_leaders = false;
let mut paragraph_in_wrapped_bold_run = false;
let mut inserted_tables: HashSet<(u32, usize)> = HashSet::new();
let mut inserted_images: HashSet<(u32, usize)> = HashSet::new();
@@ -475,6 +551,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
current_page = line.page;
prev_y = f32::MAX;
prev_x = 0.0;
paragraph_in_wrapped_bold_run = false;
if options.include_page_numbers {
output.push_str(&format!("<!-- Page {} -->\n\n", current_page));
@@ -489,6 +566,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
output.push('\n');
output.push_str(table_md);
@@ -506,6 +584,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
output.push('\n');
output.push_str(image_md);
@@ -527,9 +606,18 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
&& y_gap.abs() <= para_threshold
&& (prev_x - line_x).abs() > 50.0
&& prev_y < f32::MAX;
if (is_para_break || is_band_switch) && in_paragraph {
let line_all_bold = !line.items.is_empty() && line.items.iter().all(|item| item.is_bold);
let line_in_wrapped_bold_run = wrapped_bold_paragraph_lines.contains(&line_idx);
let is_bold_to_regular_break = in_paragraph
&& paragraph_in_wrapped_bold_run
&& !line_in_wrapped_bold_run
&& !line_all_bold
&& y_gap > base_size * 1.2
&& y_gap <= para_threshold;
if (is_para_break || is_band_switch || is_bold_to_regular_break) && in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
// Don't immediately end list on paragraph break
// Let the continuation check below decide if we're still in a list
@@ -537,7 +625,11 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
prev_x = line_x;
// Get text with optional bold/italic formatting
let text = line.text_with_formatting(options.detect_bold, options.detect_italic);
let text = line.text_with_formatting(
options.detect_bold,
options.detect_italic,
options.detect_underline,
);
let trimmed = text.trim();
// Also get plain text for pattern matching (list detection, captions, etc.)
@@ -572,6 +664,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
output.push_str(trimmed);
output.push_str("\n\n");
@@ -625,6 +718,9 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
if !(1..=15).contains(&word_count) {
return None;
}
if wrapped_bold_paragraph_lines.contains(&line_idx) {
return None;
}
let rarity = font_size_rarity(line_font_size, &font_stats);
let all_bold = !line.items.is_empty() && line.items.iter().all(|i| i.is_bold);
let standalone = !in_paragraph;
@@ -656,10 +752,17 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
let prefix = "#".repeat(level);
// Use plain text for headers to avoid redundant formatting
output.push_str(&format!("{} {}\n\n", prefix, plain_trimmed));
// Plain text for headers (no redundant bold/italic inside `#`),
// but underline is preserved: `<u>` carries meaning `#` doesn't.
let heading_text = if options.detect_underline {
line.text_with_formatting(false, false, true)
} else {
plain_text.clone()
};
output.push_str(&format!("{} {}\n\n", prefix, heading_text.trim()));
in_list = false;
continue;
}
@@ -678,6 +781,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
output.push_str(&format!("- {}", trimmed));
output.push('\n');
@@ -691,6 +795,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
let formatted = format_list_item(trimmed);
output.push_str(&formatted);
@@ -737,6 +842,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
output.push_str(&format!("> {}\n", trimmed));
continue;
@@ -747,6 +853,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
if !in_code_block {
output.push_str("```\n");
@@ -767,6 +874,11 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
}
}
output.push_str(trimmed);
paragraph_in_wrapped_bold_run = if in_paragraph {
paragraph_in_wrapped_bold_run || line_in_wrapped_bold_run
} else {
line_in_wrapped_bold_run
};
in_paragraph = true;
prev_had_dot_leaders = cur_dot_leaders;
}
@@ -836,6 +948,8 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
let para_threshold = compute_paragraph_threshold(&lines, base_size);
let isolated_lines = find_isolated_lines(&lines, base_size, para_threshold);
let wrapped_bold_paragraph_lines =
find_wrapped_bold_paragraph_lines(&lines, base_size, para_threshold);
let mut output = String::new();
let mut current_page = 0u32;
@@ -844,6 +958,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
let mut in_paragraph = false;
let mut last_list_x: Option<f32> = None;
let mut prev_had_dot_leaders = false;
let mut paragraph_in_wrapped_bold_run = false;
for (line_idx, line) in lines.iter().enumerate() {
// Page break
@@ -860,6 +975,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
in_list = false;
last_list_x = None;
prev_had_dot_leaders = false;
paragraph_in_wrapped_bold_run = false;
if options.include_page_numbers {
output.push_str(&format!("<!-- Page {} -->\n\n", current_page));
@@ -870,16 +986,29 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
// (newspaper columns emitted sequentially on the same page).
let y_gap = prev_y - line.y;
let is_para_break = y_gap.abs() > para_threshold;
if is_para_break && in_paragraph {
let line_all_bold = !line.items.is_empty() && line.items.iter().all(|item| item.is_bold);
let line_in_wrapped_bold_run = wrapped_bold_paragraph_lines.contains(&line_idx);
let is_bold_to_regular_break = in_paragraph
&& paragraph_in_wrapped_bold_run
&& !line_in_wrapped_bold_run
&& !line_all_bold
&& y_gap > base_size * 1.2
&& y_gap <= para_threshold;
if (is_para_break || is_bold_to_regular_break) && in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
// Don't immediately end list on paragraph break
// Let the continuation check below decide if we're still in a list
prev_y = line.y;
// Get text with optional bold/italic formatting
let text = line.text_with_formatting(options.detect_bold, options.detect_italic);
let text = line.text_with_formatting(
options.detect_bold,
options.detect_italic,
options.detect_underline,
);
let trimmed = text.trim();
// Also get plain text for pattern matching
@@ -896,6 +1025,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
output.push_str(trimmed);
output.push_str("\n\n");
@@ -918,6 +1048,9 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
if !(1..=15).contains(&word_count) {
return None;
}
if wrapped_bold_paragraph_lines.contains(&line_idx) {
return None;
}
let rarity = font_size_rarity(line_font_size, &font_stats);
let all_bold = !line.items.is_empty() && line.items.iter().all(|i| i.is_bold);
let standalone = !in_paragraph;
@@ -935,10 +1068,16 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
let prefix = "#".repeat(header_level);
// Use plain text for headers to avoid redundant formatting
output.push_str(&format!("{} {}\n\n", prefix, plain_trimmed));
// Plain text for headers, except underline (see above).
let heading_text = if options.detect_underline {
line.text_with_formatting(false, false, true)
} else {
plain_text.clone()
};
output.push_str(&format!("{} {}\n\n", prefix, heading_text.trim()));
in_list = false;
continue;
}
@@ -949,6 +1088,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
let formatted = format_list_item(trimmed);
output.push_str(&formatted);
@@ -993,6 +1133,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
if in_paragraph {
output.push_str("\n\n");
in_paragraph = false;
paragraph_in_wrapped_bold_run = false;
}
// Use plain text for code blocks
output.push_str(&format!("```\n{}\n```\n", plain_trimmed));
@@ -1010,6 +1151,11 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
}
}
output.push_str(trimmed);
paragraph_in_wrapped_bold_run = if in_paragraph {
paragraph_in_wrapped_bold_run || line_in_wrapped_bold_run
} else {
line_in_wrapped_bold_run
};
in_paragraph = true;
prev_had_dot_leaders = cur_dot_leaders;
}
@@ -1042,6 +1188,7 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: crate::types::ItemType::Text,
mcid,
}
@@ -1356,6 +1503,109 @@ mod tests {
);
}
#[test]
fn test_wrapped_bold_abstract_is_not_split_into_headings() {
// Regression for arXiv 1107.1353: the opening abstract paragraph is
// entirely bold at body size. The first wrapped lines used to become
// separate H2 headings, and the following body paragraph was joined to
// the bold abstract because the paragraph gap is modest.
let make = |text: &str, y: f32, font_size: f32, bold: bool| {
let mut item = make_item(text, 1, None);
item.y = y;
item.font_size = font_size;
item.height = font_size;
item.is_bold = bold;
item
};
let lines = vec![
make_line(vec![make(
"Quantum Nature of Light Measured With a Single Detector",
747.7,
25.0,
true,
)]),
make_line(vec![make(
"Gesine A. Steudle1*, Stefan Schietinger1, David Höckel1",
651.1,
11.0,
false,
)]),
make_line(vec![make(
"Zwiller2, and Oliver Benson1",
638.5,
11.0,
false,
)]),
make_line(vec![make(
"The introduction of light quanta by Einstein in 1905 triggered strong efforts to",
607.5,
11.0,
true,
)]),
make_line(vec![make(
"demonstrate the quantum properties of light directly, without involving matter",
594.8,
11.0,
true,
)]),
make_line(vec![make(
"quantization. It however took more than seven decades for the quantum granularity",
582.2,
11.0,
true,
)]),
make_line(vec![make(
"of light to be observed in the fluorescence of single atoms. Single atoms emit",
569.5,
11.0,
true,
)]),
make_line(vec![make(
"photons one at a time, this is typically demonstrated with a Hanbury-Brown-Twiss",
556.9,
11.0,
true,
)]),
make_line(vec![make(
"Our work significantly simplifies a widely used photon-correlation technique.",
544.2,
11.0,
true,
)]),
make_line(vec![make(
"A photon is a single excitation of a mode of the electromagnetic field.",
528.7,
11.0,
false,
)]),
];
let md = to_markdown_from_lines_with_tables_and_images(
lines,
MarkdownOptions::default(),
HashMap::new(),
HashMap::new(),
&std::collections::HashSet::new(),
None,
);
assert!(
md.contains("# Quantum Nature of Light Measured With a Single Detector"),
"title should remain a heading: {md}"
);
assert!(
!md.contains("## The introduction")
&& !md.contains("## demonstrate")
&& !md.contains("## quantization"),
"bold abstract lines should not become headings: {md}"
);
assert!(
md.contains("technique.**\n\nA photon is a single excitation"),
"body paragraph should be separated from bold abstract: {md}"
);
}
#[test]
fn test_struct_role_code_multiline_accumulation() {
let mut line1 = make_item("fn main() {", 1, Some(0));
+4
View File
@@ -400,6 +400,8 @@ pub struct MarkdownOptions {
pub detect_bold: bool,
/// Detect and format italic text from font names
pub detect_italic: bool,
/// Emit `<u>` runs for text with a geometrically-detected underline
pub detect_underline: bool,
/// Include image placeholders in output
pub include_images: bool,
/// Include extracted hyperlinks
@@ -422,6 +424,7 @@ impl Default for MarkdownOptions {
fix_hyphenation: true,
detect_bold: true,
detect_italic: true,
detect_underline: true,
// `include_images: false` is intentional. The content-stream walker
// now emits `ItemType::Image` `TextItem`s for every Image XObject
// it encounters (see `extractor/content_stream.rs`). If we rendered
@@ -1217,6 +1220,7 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: crate::types::ItemType::Text,
mcid: None,
}
+84
View File
@@ -29,6 +29,8 @@ pub(crate) fn clean_markdown(mut text: String, options: &MarkdownOptions) -> Str
// text item, which combine with gap-based space insertion to produce
// double spaces ("Vice President" instead of "Vice President").
collapse_consecutive_spaces(&mut text);
remove_spaces_before_closing_brackets(&mut text);
remove_spaces_before_sentence_punctuation(&mut text);
// Remove excessive newlines (more than 2 in a row)
while text.contains("\n\n\n") {
@@ -71,6 +73,46 @@ fn collapse_consecutive_spaces(text: &mut String) {
*text = result;
}
/// Remove spaces before closing square brackets.
/// Unit markers and markdown links occasionally pick up a gap-inserted space
/// before `]` (e.g. `[kg/m3 ]`), which is cosmetic padding.
fn remove_spaces_before_closing_brackets(text: &mut String) {
let mut result = String::with_capacity(text.len());
for ch in text.chars() {
if ch == ']' && result.ends_with(' ') {
result.pop();
}
result.push(ch);
}
*text = result;
}
/// Remove a stray space before sentence punctuation ("word ." → "word.").
/// Style-boundary item splits (bold/italic/underline runs) can strand a
/// trailing period or comma in its own fragment, and several assembly paths
/// join fragments with spaces. Only fires when the punctuation ends the
/// token (followed by whitespace or end of text), so decimals ("3 .14" stays
/// untouched — no such input exists, but the guard is cheap) and dot leaders
/// (" ... ") are unaffected.
fn remove_spaces_before_sentence_punctuation(text: &mut String) {
let chars: Vec<char> = text.chars().collect();
let mut result = String::with_capacity(text.len());
for (i, &ch) in chars.iter().enumerate() {
if matches!(ch, '.' | ',' | ';') && result.ends_with(' ') {
let next = chars.get(i + 1);
// `|` counts as a token end so table cells get the same fix.
let token_ends = next.is_none_or(|c| c.is_whitespace() || *c == '|');
// Never touch runs of dots (ellipsis / dot leaders).
let in_dot_run = ch == '.' && next == Some(&'.');
if token_ends && !in_dot_run {
result.pop();
}
}
result.push(ch);
}
*text = result;
}
/// Collapse dot leaders (runs of 4+ dots) into " ... "
/// Common in tables of contents: "Introduction...............................1" -> "Introduction ... 1"
fn collapse_dot_leaders(text: &str) -> String {
@@ -342,6 +384,48 @@ mod tests {
assert!(result.contains("Chapter 2 ... 20"));
}
// --- remove_spaces_before_closing_brackets ---
#[test]
fn test_remove_spaces_before_closing_brackets() {
let mut input = "Density [kg/m3 ] and [linked text ](https://example.com)".to_string();
remove_spaces_before_closing_brackets(&mut input);
assert_eq!(
input,
"Density [kg/m3] and [linked text](https://example.com)"
);
}
// --- remove_spaces_before_sentence_punctuation ---
#[test]
fn strips_space_before_trailing_period() {
let mut t = "Foreign insurance companies . The provisions".to_string();
remove_spaces_before_sentence_punctuation(&mut t);
assert_eq!(t, "Foreign insurance companies. The provisions");
}
#[test]
fn strips_space_before_period_at_cell_boundary() {
let mut t = "|Applicability date .|This section|".to_string();
remove_spaces_before_sentence_punctuation(&mut t);
assert_eq!(t, "|Applicability date.|This section|");
}
#[test]
fn keeps_dot_leaders_and_ellipses() {
let mut t = "Introduction ... 1".to_string();
remove_spaces_before_sentence_punctuation(&mut t);
assert_eq!(t, "Introduction ... 1");
}
#[test]
fn keeps_mid_token_periods() {
let mut t = "version 3 .14 released".to_string();
remove_spaces_before_sentence_punctuation(&mut t);
assert_eq!(t, "version 3 .14 released");
}
// --- fix_hyphenation ---
#[test]
+1
View File
@@ -542,6 +542,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid,
}
+52
View File
@@ -30,6 +30,9 @@ pub struct PyPdfResult {
/// 1-indexed page numbers that need OCR.
#[pyo3(get)]
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
#[pyo3(get)]
pub ocr_reasons_by_page: Vec<PyPageOcrReasons>,
/// Title from PDF metadata.
#[pyo3(get)]
pub title: Option<String>,
@@ -60,6 +63,28 @@ impl PyPdfResult {
}
}
/// OCR reasons for a single 1-indexed page.
#[pyclass(name = "PageOcrReasons")]
#[derive(Clone)]
pub struct PyPageOcrReasons {
/// 1-indexed page number.
#[pyo3(get)]
pub page: u32,
/// Machine-readable OCR reason identifiers.
#[pyo3(get)]
pub reasons: Vec<String>,
}
#[pymethods]
impl PyPageOcrReasons {
fn __repr__(&self) -> String {
format!(
"PageOcrReasons(page={}, reasons={:?})",
self.page, self.reasons
)
}
}
// ---------------------------------------------------------------------------
// Classification wrapper (lightweight)
// ---------------------------------------------------------------------------
@@ -106,6 +131,9 @@ pub struct PyRegionText {
/// True when the text should not be trusted (empty, GID fonts, garbage, encoding issues).
#[pyo3(get)]
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
#[pyo3(get)]
pub ocr_reason: Option<String>,
}
#[pymethods]
@@ -160,6 +188,9 @@ pub struct PyPageMarkdown {
/// encoding issues, garbage text, or empty extraction).
#[pyo3(get)]
pub needs_ocr: bool,
/// Machine-readable OCR reason when the cause is known.
#[pyo3(get)]
pub ocr_reason: Option<String>,
}
#[pymethods]
@@ -190,6 +221,9 @@ pub struct PyPagesExtractionResult {
/// 1-indexed pages that need OCR (scanned/image-based or unreliable text).
#[pyo3(get)]
pub pages_needing_ocr: Vec<u32>,
/// Machine-readable OCR reasons by 1-indexed page.
#[pyo3(get)]
pub ocr_reasons_by_page: Vec<PyPageOcrReasons>,
/// True if any page has tables or columns.
#[pyo3(get)]
pub is_complex: bool,
@@ -232,6 +266,8 @@ pub struct PyTextItem {
#[pyo3(get)]
pub is_italic: bool,
#[pyo3(get)]
pub is_underline: bool,
#[pyo3(get)]
pub item_type: String,
}
@@ -268,6 +304,7 @@ fn to_py_result(r: crate::PdfProcessResult) -> PyPdfResult {
page_count: r.page_count,
processing_time_ms: r.processing_time_ms,
pages_needing_ocr: r.pages_needing_ocr,
ocr_reasons_by_page: to_py_page_ocr_reasons(r.ocr_reasons_by_page),
title: r.title,
confidence: r.confidence,
is_complex_layout: r.layout.is_complex,
@@ -277,6 +314,16 @@ fn to_py_result(r: crate::PdfProcessResult) -> PyPdfResult {
}
}
fn to_py_page_ocr_reasons(reasons: Vec<crate::PageOcrReasons>) -> Vec<PyPageOcrReasons> {
reasons
.into_iter()
.map(|reason| PyPageOcrReasons {
page: reason.page,
reasons: reason.reasons,
})
.collect()
}
fn to_py_err(e: crate::PdfError) -> PyErr {
PyValueError::new_err(e.to_string())
}
@@ -304,6 +351,7 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
item_type: item_type_str(&item.item_type),
})
.collect()
@@ -350,11 +398,13 @@ fn to_py_pages_result(r: crate::PagesExtractionResult) -> PyPagesExtractionResul
page: p.page,
markdown: p.markdown,
needs_ocr: p.needs_ocr,
ocr_reason: p.ocr_reason,
})
.collect(),
pages_with_tables: r.pages_with_tables,
pages_with_columns: r.pages_with_columns,
pages_needing_ocr: r.pages_needing_ocr,
ocr_reasons_by_page: to_py_page_ocr_reasons(r.ocr_reasons_by_page),
is_complex: r.is_complex,
}
}
@@ -370,6 +420,7 @@ fn convert_region_results(results: Vec<crate::PageRegionResult>) -> Vec<PyPageRe
.map(|r| PyRegionText {
text: r.text,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
})
@@ -563,6 +614,7 @@ fn extract_pages_markdown_bytes(
#[pymodule]
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_class::<PyPdfResult>()?;
m.add_class::<PyPageOcrReasons>()?;
m.add_class::<PyPdfClassification>()?;
m.add_class::<PyTextItem>()?;
m.add_class::<PyRegionText>()?;
+1
View File
@@ -104,6 +104,7 @@ pub(crate) fn merge_adjacent_items(items: &[TextItem]) -> (Vec<TextItem>, Vec<Ve
page: first_item.page,
is_bold: first_item.is_bold,
is_italic: first_item.is_italic,
is_underline: first_item.is_underline,
item_type: first_item.item_type.clone(),
mcid: first_item.mcid,
});
+1
View File
@@ -393,6 +393,7 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
+96 -1
View File
@@ -1124,12 +1124,13 @@ pub(crate) fn assign_items_to_grid(
.unwrap_or(std::cmp::Ordering::Equal)
})
});
let text: String = col_items
let text = col_items
.iter()
.map(|(_, item)| item.text.trim())
.filter(|t| !t.is_empty())
.collect::<Vec<_>>()
.join(" ");
let text = remove_inner_delimiter_spaces(&text);
row_cells.push(text);
}
cells.push(row_cells);
@@ -1138,6 +1139,27 @@ pub(crate) fn assign_items_to_grid(
(cells, indices)
}
fn remove_inner_delimiter_spaces(text: &str) -> String {
let chars: Vec<char> = text.chars().collect();
let mut result = String::with_capacity(text.len());
for (i, &ch) in chars.iter().enumerate() {
if ch == ' ' {
let after_open =
result.ends_with('(') || result.ends_with('[') || result.ends_with('{');
let before_close = chars
.get(i + 1)
.is_some_and(|next| matches!(next, ')' | ']' | '}'));
if after_open || before_close {
continue;
}
}
result.push(ch);
}
result
}
/// Consolidate text in vertically-merged cells.
///
/// When a single rect spans multiple grid rows (e.g. a "Classification" label
@@ -1451,6 +1473,10 @@ fn detect_row_stripe_table(
(col_edges, cells)
};
let num_cols = col_edges.len() - 1;
if row_stripe_is_sparse_prose_outline(&cells) {
debug!(" row-stripe rejected: sparse outline/prose continuation shape");
return None;
}
let column_centers: Vec<f32> = (0..num_cols)
.map(|c| (col_edges[c] + col_edges[c + 1]) / 2.0)
@@ -1469,6 +1495,57 @@ fn detect_row_stripe_table(
Some(Table::new(column_centers, row_centers, cells, item_indices))
}
fn row_stripe_is_sparse_prose_outline(cells: &[Vec<String>]) -> bool {
let Some(num_cols) = cells.first().map(|row| row.len()) else {
return false;
};
if num_cols != 2 || cells.len() < 4 {
return false;
}
let non_empty_rows = cells
.iter()
.filter(|row| row.iter().any(|cell| !cell.trim().is_empty()))
.count();
if non_empty_rows < 4 {
return false;
}
let mut col_counts = [0usize; 2];
for row in cells {
for (idx, cell) in row.iter().enumerate() {
if !cell.trim().is_empty() {
col_counts[idx] += 1;
}
}
}
let (sparse_col, dense_col) = if col_counts[0] <= col_counts[1] {
(0usize, 1usize)
} else {
(1usize, 0usize)
};
let sparse_count = col_counts[sparse_col];
let dense_count = col_counts[dense_col];
if sparse_count * 2 >= non_empty_rows || dense_count * 3 < non_empty_rows * 2 {
return false;
}
let blank_sparse_dense_rows = cells
.iter()
.filter(|row| row[sparse_col].trim().is_empty() && !row[dense_col].trim().is_empty())
.count();
if blank_sparse_dense_rows * 2 < non_empty_rows {
return false;
}
let long_dense_cells = cells
.iter()
.filter(|row| row[dense_col].split_whitespace().count() >= 6)
.count();
long_dense_cells * 2 >= dense_count
}
/// Detect a table from cell-background rects that failed grid detection.
///
/// Uses rect Y-edges for row boundaries and text X-position clustering for
@@ -2314,6 +2391,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -2523,6 +2601,21 @@ mod tests {
assert!(cells[0][0].contains("World"));
}
#[test]
fn test_assign_items_parenthetical_no_inner_spaces() {
let items = vec![
make_item("The first sentence", 15.0, 85.0, 10.0),
make_item("(", 90.0, 85.0, 10.0),
make_item("twice", 95.0, 85.0, 10.0),
make_item(")", 120.0, 85.0, 10.0),
];
let col_edges = vec![10.0, 150.0];
let row_edges = vec![90.0, 70.0];
let (cells, indices) = assign_items_to_grid(&items, &col_edges, &row_edges, 1);
assert_eq!(indices.len(), 4);
assert_eq!(cells[0][0], "The first sentence (twice)");
}
#[test]
fn test_assign_items_boundary_tolerance() {
// Item right at edge with ±2pt tolerance
@@ -3272,6 +3365,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: crate::types::ItemType::Text,
mcid: None,
});
@@ -3581,6 +3675,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: crate::types::ItemType::Text,
mcid: None,
});
+1
View File
@@ -586,6 +586,7 @@ mod tests {
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid,
}
+1
View File
@@ -108,6 +108,7 @@ pub(crate) fn try_split_financial_item(item: &TextItem) -> Option<Vec<TextItem>>
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
+21
View File
@@ -369,6 +369,10 @@ pub(crate) fn join_cell_items(items: &[&TextItem]) -> String {
let prev_ends_with_hyphen = result.ends_with('-');
let curr_is_hyphen = text == "-";
let curr_starts_with_hyphen = text.starts_with('-');
let prev_ends_with_open_delimiter =
result.ends_with('(') || result.ends_with('[') || result.ends_with('{');
let curr_starts_with_close_delimiter =
text.starts_with(')') || text.starts_with(']') || text.starts_with('}');
// Detect subscript/superscript: smaller font size and/or Y offset
let font_ratio = item.font_size / prev_item.font_size;
@@ -385,6 +389,8 @@ pub(crate) fn join_cell_items(items: &[&TextItem]) -> String {
|| curr_starts_with_hyphen
|| is_sub_super
|| was_sub_super
|| prev_ends_with_open_delimiter
|| curr_starts_with_close_delimiter
{
result.push_str(text);
} else {
@@ -514,6 +520,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -727,6 +734,18 @@ mod tests {
assert_eq!(join_cell_items(&[&a, &b, &c]), "pre-fix");
}
#[test]
fn test_join_cell_items_parenthetical_no_inner_spaces() {
let a = make_item("The first sentence", 100.0, 500.0, 10.0);
let b = make_item("(", 190.0, 500.0, 10.0);
let c = make_item("twice", 195.0, 500.0, 10.0);
let d = make_item(")", 220.0, 500.0, 10.0);
assert_eq!(
join_cell_items(&[&a, &b, &c, &d]),
"The first sentence (twice)"
);
}
#[test]
fn test_join_cell_items_subscript_no_space() {
let a = make_item("H", 100.0, 500.0, 12.0);
@@ -866,6 +885,7 @@ mod tests {
font: String::new(),
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
@@ -902,6 +922,7 @@ mod tests {
font: String::new(),
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
page: 1,
+573 -51
View File
@@ -235,6 +235,7 @@ fn split_merged_numbers(item: &TextItem, col_boundaries: &[f32]) -> Vec<TextItem
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
@@ -255,6 +256,7 @@ fn split_merged_numbers(item: &TextItem, col_boundaries: &[f32]) -> Vec<TextItem
page: item.page,
is_bold: item.is_bold,
is_italic: item.is_italic,
is_underline: item.is_underline,
item_type: item.item_type.clone(),
mcid: item.mcid,
});
@@ -566,7 +568,7 @@ pub(crate) fn try_build_key_value_table_from_rows(items: &[TextItem], page: u32)
})
.collect();
if page_items.len() < 4 {
if page_items.len() < 2 {
return None;
}
@@ -575,15 +577,12 @@ pub(crate) fn try_build_key_value_table_from_rows(items: &[TextItem], page: u32)
.max(1.0);
let y_tol = (median_font_size * 0.75).clamp(4.0, 9.0);
let rows = group_key_value_visual_rows(page_items, y_tol);
if rows.len() < 2 || rows.len() > 80 {
if rows.is_empty() || rows.len() > 80 {
return None;
}
let split_x = infer_key_value_split_x(&rows, median_font_size)?;
let mut kv_rows: Vec<KeyValueRow> = Vec::new();
let mut paired_rows = 0usize;
let mut section_rows = 0usize;
let mut left_label_like = 0usize;
let mut left_starts = Vec::new();
let mut right_starts = Vec::new();
@@ -609,18 +608,12 @@ pub(crate) fn try_build_key_value_table_from_rows(items: &[TextItem], page: u32)
item_indices.dedup();
if !left.is_empty() && !right.is_empty() {
paired_rows += 1;
if looks_like_key_value_label(&left) {
left_label_like += 1;
}
if let Some(x) = left_items.first().map(|ri| ri.item.x) {
left_starts.push(x);
}
if let Some(x) = right_items.first().map(|ri| ri.item.x) {
right_starts.push(x);
}
} else if !left.is_empty() {
section_rows += 1;
}
kv_rows.push(KeyValueRow {
@@ -631,11 +624,75 @@ pub(crate) fn try_build_key_value_table_from_rows(items: &[TextItem], page: u32)
});
}
if kv_rows.len() < 2 || paired_rows < 2 {
if kv_rows.is_empty() {
return None;
}
let raw_left_only_rows = kv_rows
.iter()
.filter(|row| !row.left.is_empty() && row.right.is_empty())
.count();
let raw_right_only_rows = kv_rows
.iter()
.filter(|row| row.left.is_empty() && !row.right.is_empty())
.count();
let edgar_tag_rows = key_value_rows_look_like_edgar_tags(&kv_rows);
if edgar_tag_rows {
kv_rows.retain(|row| !row.right.is_empty() || !is_edgar_table_boundary_cell(&row.left));
}
let header_inferred = !edgar_tag_rows && key_value_first_pair_is_header(&kv_rows);
kv_rows = normalize_key_value_rows(kv_rows, header_inferred);
let paired_rows = kv_rows
.iter()
.filter(|row| !row.left.is_empty() && !row.right.is_empty())
.count();
let section_rows = kv_rows
.iter()
.filter(|row| !row.left.is_empty() && row.right.is_empty())
.count();
let dangling_right_rows = kv_rows
.iter()
.filter(|row| row.left.is_empty() && !row.right.is_empty())
.count();
let left_label_like = kv_rows
.iter()
.filter(|row| !row.left.is_empty() && !row.right.is_empty())
.filter(|row| looks_like_key_value_label(&row.left))
.count();
if paired_rows < 1 {
return None;
}
if dangling_right_rows > 0 {
return None;
}
let left_x = median_f32(left_starts).unwrap_or_else(|| {
rows.iter()
.flat_map(|row| row.items.iter().map(|ri| ri.item.x))
.fold(f32::INFINITY, f32::min)
});
let right_x = median_f32(right_starts).unwrap_or(split_x);
if !left_x.is_finite() || !right_x.is_finite() || right_x - left_x < 40.0 {
return None;
}
let single_pair_allowed = key_value_single_pair_allowed(
KeyValueSinglePairStats {
paired_rows,
section_rows,
raw_left_only_rows,
raw_right_only_rows,
},
&kv_rows,
header_inferred,
left_x,
right_x,
);
if (kv_rows.len() < 2 || paired_rows < 2) && !single_pair_allowed {
return None;
}
let header_inferred = key_value_first_pair_is_header(&kv_rows);
let data_pairs = if header_inferred {
paired_rows.saturating_sub(1)
} else {
@@ -645,7 +702,7 @@ pub(crate) fn try_build_key_value_table_from_rows(items: &[TextItem], page: u32)
return None;
}
if section_rows > paired_rows * 2 + 2 {
if section_rows > paired_rows * 2 + 2 && !single_pair_allowed {
return None;
}
@@ -659,28 +716,25 @@ pub(crate) fn try_build_key_value_table_from_rows(items: &[TextItem], page: u32)
} else {
left_label_like
};
if label_rows_for_score >= 2 && label_like_for_score * 2 < label_rows_for_score {
return None;
}
let left_x = median_f32(left_starts).unwrap_or_else(|| {
rows.iter()
.flat_map(|row| row.items.iter().map(|ri| ri.item.x))
.fold(f32::INFINITY, f32::min)
});
let right_x = median_f32(right_starts).unwrap_or(split_x);
if !left_x.is_finite() || !right_x.is_finite() || right_x - left_x < 40.0 {
return None;
}
let right_cluster_count = significant_side_x_clusters(&rows, split_x, false);
let marker_rows = marker_matrix_value_rows(&kv_rows);
if (right_cluster_count >= 5 && paired_rows >= 3)
|| (right_cluster_count >= 3 && marker_rows >= 3 && marker_rows * 2 >= paired_rows)
if !header_inferred
&& !edgar_tag_rows
&& label_rows_for_score >= 2
&& label_like_for_score * 2 < label_rows_for_score
{
return None;
}
if key_value_rows_look_like_prose(&kv_rows, header_inferred) {
let right_cluster_count = significant_side_x_clusters(&rows, split_x, false);
let marker_rows = marker_matrix_value_rows(&kv_rows);
if !single_pair_allowed
&& !edgar_tag_rows
&& ((right_cluster_count >= 5 && paired_rows >= 3)
|| (right_cluster_count >= 3 && marker_rows >= 3 && marker_rows * 2 >= paired_rows))
{
return None;
}
if !edgar_tag_rows && key_value_rows_look_like_prose(&kv_rows, header_inferred) {
return None;
}
@@ -765,6 +819,82 @@ struct KeyValueRow {
item_indices: Vec<usize>,
}
#[derive(Debug, Clone, Copy)]
struct KeyValueSinglePairStats {
paired_rows: usize,
section_rows: usize,
raw_left_only_rows: usize,
raw_right_only_rows: usize,
}
fn normalize_key_value_rows(rows: Vec<KeyValueRow>, header_inferred: bool) -> Vec<KeyValueRow> {
let mut normalized: Vec<KeyValueRow> = Vec::with_capacity(rows.len());
for row in rows {
if row.left.is_empty() && row.right.is_empty() {
continue;
}
if row.left.is_empty() && !row.right.is_empty() {
if let Some(last) = normalized.last_mut() {
if !last.right.is_empty() {
append_key_value_text(&mut last.right, &row.right);
last.item_indices.extend(row.item_indices);
continue;
}
}
normalized.push(row);
continue;
}
if !row.left.is_empty() && row.right.is_empty() {
let normalized_len = normalized.len();
if let Some(last) = normalized.last_mut() {
let last_is_header = header_inferred && normalized_len == 1;
if !last_is_header
&& !last.left.is_empty()
&& !last.right.is_empty()
&& key_value_left_continuation_allowed(&last.left, &row.left)
{
append_key_value_text(&mut last.left, &row.left);
last.item_indices.extend(row.item_indices);
continue;
}
}
}
normalized.push(row);
}
normalized
}
fn append_key_value_text(target: &mut String, addition: &str) {
let addition = addition.trim();
if addition.is_empty() {
return;
}
if !target.trim().is_empty() {
target.push(' ');
}
target.push_str(addition);
}
fn key_value_left_continuation_allowed(previous_left: &str, continuation: &str) -> bool {
let trimmed = continuation.trim();
if trimmed.is_empty() || looks_like_key_value_section_label(trimmed) {
return false;
}
let previous = previous_left.trim_end();
let continuation_chars = trimmed.chars().count();
let continuation_words = word_count_simple(trimmed);
previous.ends_with(['-', '/', ',', ';', ':'])
|| first_alpha_is_lowercase(trimmed)
|| continuation_chars > 28
|| continuation_words > 4
}
fn group_key_value_visual_rows(mut items: Vec<RowItem>, y_tol: f32) -> Vec<VisualRow> {
items.sort_by(|a, b| {
b.item
@@ -828,6 +958,13 @@ fn infer_key_value_split_x(rows: &[VisualRow], median_font_size: f32) -> Option<
}
if splits.len() < 2 {
let paired_visual_rows = rows.iter().filter(|row| row.items.len() >= 2).count();
if splits.len() == 1
&& paired_visual_rows == 1
&& (rows.len() == 1 || rows.iter().all(|row| row.items.len() <= 2))
{
return splits.into_iter().next();
}
return None;
}
@@ -902,16 +1039,169 @@ fn looks_like_key_value_label(cell: &str) -> bool {
trimmed.chars().any(|c| c.is_alphabetic())
}
fn key_value_rows_look_like_edgar_tags(rows: &[KeyValueRow]) -> bool {
let paired_rows = rows
.iter()
.filter(|row| !row.left.is_empty() && !row.right.is_empty())
.count();
if paired_rows < 2 {
return false;
}
let tag_pairs = rows
.iter()
.filter(|row| !row.left.is_empty() && !row.right.is_empty())
.filter(|row| is_edgar_tag_cell(&row.left))
.count();
let first_marker = rows.first().is_some_and(|row| {
row.left.eq_ignore_ascii_case("<S>") && row.right.eq_ignore_ascii_case("<C>")
});
tag_pairs >= 3 || (first_marker && tag_pairs >= 2)
}
fn is_edgar_tag_cell(cell: &str) -> bool {
let trimmed = cell.trim();
let Some(inner) = trimmed.strip_prefix('<').and_then(|s| s.strip_suffix('>')) else {
return false;
};
!inner.is_empty()
&& inner.len() <= 48
&& inner
.chars()
.all(|ch| ch.is_ascii_uppercase() || ch.is_ascii_digit() || matches!(ch, '-' | '_'))
}
fn is_edgar_table_boundary_cell(cell: &str) -> bool {
let trimmed = cell.trim();
trimmed.eq_ignore_ascii_case("<TABLE>") || trimmed.eq_ignore_ascii_case("</TABLE>")
}
fn key_value_single_pair_allowed(
stats: KeyValueSinglePairStats,
rows: &[KeyValueRow],
header_inferred: bool,
left_x: f32,
right_x: f32,
) -> bool {
if header_inferred || stats.paired_rows != 1 || stats.section_rows != 0 || rows.len() != 1 {
return false;
}
if right_x - left_x < 60.0 {
return false;
}
let Some(row) = rows
.iter()
.find(|row| !row.left.is_empty() && !row.right.is_empty())
else {
return false;
};
let left_chars = row.left.chars().count();
let right_chars = row.right.chars().count();
if !(2..=120).contains(&left_chars) || right_chars == 0 {
return false;
}
if key_value_cell_looks_like_sentence(&row.left) {
return false;
}
if stats.raw_left_only_rows == 0
&& stats.raw_right_only_rows >= 2
&& left_chars <= 70
&& right_chars <= 1_500
&& looks_like_key_value_label(&row.left)
{
return true;
}
if right_chars > 80 {
return false;
}
if key_value_cell_looks_like_sentence(&row.right) && !compact_key_value_scalar(&row.right) {
return false;
}
(looks_like_key_value_label(&row.left) || left_chars <= 90)
&& compact_key_value_scalar(&row.right)
}
fn compact_key_value_scalar(cell: &str) -> bool {
let trimmed = cell.trim();
let chars = trimmed.chars().count();
let words = word_count_simple(trimmed);
if trimmed.is_empty() || chars > 60 || words > 6 || trimmed.ends_with(['.', '!', '?']) {
return false;
}
let lower = trimmed.to_ascii_lowercase();
trimmed.chars().any(|ch| ch.is_ascii_digit())
|| matches!(
lower.as_str(),
"yes" | "no" | "true" | "false" | "none" | "n/a" | "na"
)
|| words <= 4
}
fn looks_like_key_value_section_label(cell: &str) -> bool {
let trimmed = cell.trim();
let chars = trimmed.chars().count();
let words = word_count_simple(trimmed);
if !(1..=5).contains(&words) || !(2..=48).contains(&chars) {
return false;
}
if trimmed.ends_with(['.', ',', ';', ':']) || first_alpha_is_lowercase(trimmed) {
return false;
}
if trimmed
.chars()
.any(|ch| matches!(ch, '.' | ',' | ';' | '(' | ')' | '[' | ']'))
{
return false;
}
trimmed.chars().any(|ch| ch.is_alphabetic())
}
fn first_alpha_is_lowercase(cell: &str) -> bool {
cell.chars()
.find(|ch| ch.is_alphabetic())
.is_some_and(|ch| ch.is_lowercase())
}
fn key_value_cell_looks_like_sentence(cell: &str) -> bool {
let trimmed = cell.trim();
let chars = trimmed.chars().count();
chars > 90
|| word_count_simple(trimmed) > 12
|| (chars > 42 && trimmed.ends_with(['.', '!', '?']))
}
fn key_value_rows_look_like_prose(rows: &[KeyValueRow], header_inferred: bool) -> bool {
let mut long_sentence_cells = 0usize;
let mut total_cells = 0usize;
let mut total_chars = 0usize;
let mut left_cells = 0usize;
let mut left_prose_cells = 0usize;
let mut left_label_like = 0usize;
let mut total_left_chars = 0usize;
let mut paired_rows = 0usize;
let mut paired_sentence_rows = 0usize;
let mut solo_prose_rows = 0usize;
for row in rows.iter().skip(usize::from(header_inferred)) {
if !row.left.is_empty() && !row.right.is_empty() {
paired_rows += 1;
let left = row.left.trim();
let right = row.right.trim();
let left_prose = key_value_cell_looks_like_sentence(left);
let right_prose = key_value_cell_looks_like_sentence(right);
left_cells += 1;
total_left_chars += left.chars().count();
if looks_like_key_value_label(left) {
left_label_like += 1;
}
if left_prose {
left_prose_cells += 1;
}
if left_prose && right_prose {
paired_sentence_rows += 1;
}
} else {
let solo = if row.left.is_empty() {
row.right.trim()
@@ -925,30 +1215,23 @@ fn key_value_rows_look_like_prose(rows: &[KeyValueRow], header_inferred: bool) -
solo_prose_rows += 1;
}
}
for cell in [&row.left, &row.right] {
let trimmed = cell.trim();
if trimmed.is_empty() {
continue;
}
total_cells += 1;
total_chars += trimmed.chars().count();
if trimmed.chars().count() > 100
|| (trimmed.chars().count() > 55 && trimmed.ends_with(['.', '!', '?']))
{
long_sentence_cells += 1;
}
}
}
if paired_rows < 1 || total_cells == 0 {
if paired_rows < 1 || left_cells == 0 {
return true;
}
if solo_prose_rows >= 3 {
return true;
}
if paired_rows >= 2 && paired_sentence_rows * 2 >= paired_rows {
return true;
}
if !header_inferred && left_prose_cells * 2 >= left_cells {
return true;
}
let avg_chars = total_chars as f32 / total_cells as f32;
avg_chars > 75.0 || long_sentence_cells * 2 >= total_cells
let avg_left_chars = total_left_chars as f32 / left_cells as f32;
!header_inferred && avg_left_chars > 70.0 && left_label_like * 2 < left_cells
}
fn marker_matrix_value_rows(rows: &[KeyValueRow]) -> usize {
@@ -1148,6 +1431,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1165,6 +1449,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1354,6 +1639,243 @@ mod tests {
assert!(md.contains("|Engine Code|1ZR-FAE|"), "{md}");
}
#[test]
fn test_key_value_builder_merges_wrapped_value_continuations() {
let items = vec![
make_char("Storage", 80.0, 700.0, 9.0, 42.0),
make_char(
"Store under normal conditions in dry rooms.",
250.0,
700.0,
9.0,
210.0,
),
make_char(
"Protect from heat and humidity in the original packaging material.",
250.0,
686.0,
9.0,
315.0,
),
make_char("Shelf Life", 80.0, 668.0, 9.0, 48.0),
make_char(
"To obtain best performance use within 24 months.",
250.0,
668.0,
9.0,
255.0,
),
make_char("Technical Information", 80.0, 650.0, 9.0, 104.0),
make_char(
"The product is designed for repeated industrial use and long service life.",
250.0,
650.0,
9.0,
340.0,
),
make_char(
"Additional details are provided for compatibility and installation planning.",
250.0,
636.0,
9.0,
350.0,
),
];
let table = try_build_key_value_table_from_rows(&items, 1).unwrap();
let md = table_to_markdown(&table);
assert!(md.contains("|Field|Value|"), "{md}");
assert!(
md.contains(
"|Storage|Store under normal conditions in dry rooms. Protect from heat and humidity in the original packaging material.|"
),
"{md}"
);
assert!(
md.contains(
"|Technical Information|The product is designed for repeated industrial use and long service life. Additional details are provided for compatibility and installation planning.|"
),
"{md}"
);
}
#[test]
fn test_key_value_builder_merges_wrapped_left_labels() {
let items = vec![
make_char("Title/Description", 76.0, 700.0, 9.0, 86.0),
make_char("Instances", 350.0, 700.0, 9.0, 48.0),
make_char("RE: Homes Gerald Ford lived in.", 76.0, 682.0, 9.0, 150.0),
make_char("Box 7", 350.0, 682.0, 9.0, 28.0),
make_char(
"Grand Rapids Remembers Gerald R. Ford issue. Grand",
76.0,
664.0,
9.0,
245.0,
),
make_char("Box 7", 350.0, 664.0, 9.0, 28.0),
make_char(
"Rapids Magazine, September 1987, p. 65.",
76.0,
650.0,
9.0,
196.0,
),
make_char(
"A Workhorse not a show horse: Gerald Ford remembered as humble.",
76.0,
632.0,
9.0,
290.0,
),
make_char("Box 7", 350.0, 632.0, 9.0, 28.0),
make_char(
"not flashy during his public life.",
76.0,
618.0,
9.0,
150.0,
),
];
let table = try_build_key_value_table_from_rows(&items, 1).unwrap();
let md = table_to_markdown(&table);
assert!(md.starts_with("|Title/Description|Instances|"), "{md}");
assert!(
md.contains(
"|Grand Rapids Remembers Gerald R. Ford issue. Grand Rapids Magazine, September 1987, p. 65.|Box 7|"
),
"{md}"
);
assert!(
md.contains(
"|A Workhorse not a show horse: Gerald Ford remembered as humble. not flashy during his public life.|Box 7|"
),
"{md}"
);
}
#[test]
fn test_key_value_builder_allows_tiny_two_cell_region() {
let items = vec![
make_char(
"3M E-A-R Classic Small Earplug Uncorded",
80.0,
700.0,
9.0,
210.0,
),
make_char("02/05/24", 360.0, 700.0, 9.0, 42.0),
];
let table = try_build_key_value_table_from_rows(&items, 1).unwrap();
let md = table_to_markdown(&table);
assert!(md.contains("|Field|Value|"), "{md}");
assert!(
md.contains("|3M E-A-R Classic Small Earplug Uncorded|02/05/24|"),
"{md}"
);
}
#[test]
fn test_key_value_builder_allows_single_wrapped_value_region() {
let items = vec![
make_char("Intrinsic Safety", 42.0, 174.0, 9.0, 60.0),
make_char(
"The powered air purifying respirator has been tested and classified",
311.0,
174.0,
9.0,
260.0,
),
make_char(
"for intrinsic safety in hazardous locations by Underwriters Laboratory",
311.0,
160.0,
9.0,
270.0,
),
make_char(
"for the following classes, divisions, groups, and temperature ratings.",
311.0,
146.0,
9.0,
275.0,
),
];
let table = try_build_key_value_table_from_rows(&items, 1).unwrap();
let md = table_to_markdown(&table);
assert!(md.contains("|Field|Value|"), "{md}");
assert!(
md.contains(
"|Intrinsic Safety|The powered air purifying respirator has been tested and classified for intrinsic safety in hazardous locations by Underwriters Laboratory for the following classes, divisions, groups, and temperature ratings.|"
),
"{md}"
);
}
#[test]
fn test_key_value_builder_rejects_leading_value_only_prose() {
let items = vec![
make_char(
"3rd Party Authorization documenting the reason for the hardship.",
260.0,
714.0,
9.0,
310.0,
),
make_char("Borrower", 80.0, 696.0, 9.0, 44.0),
make_char(
"Homeowner has adequate income to support modified payments.",
260.0,
696.0,
9.0,
300.0,
),
make_char("Servicer", 80.0, 678.0, 9.0, 42.0),
make_char(
"Collects documentation and reviews hardship status.",
260.0,
678.0,
9.0,
260.0,
),
];
assert!(try_build_key_value_table_from_rows(&items, 1).is_none());
}
#[test]
fn test_key_value_builder_recovers_edgar_tag_value_rows() {
let items = vec![
make_char("<S>", 70.0, 700.0, 9.0, 18.0),
make_char("<C>", 240.0, 700.0, 9.0, 18.0),
make_char("<PERIOD-TYPE>", 70.0, 684.0, 9.0, 78.0),
make_char("3-MOS", 240.0, 684.0, 9.0, 30.0),
make_char("<FISCAL-YEAR-END>", 70.0, 668.0, 9.0, 104.0),
make_char("DEC-31-2000", 240.0, 668.0, 9.0, 66.0),
make_char("<PERIOD-END>", 70.0, 652.0, 9.0, 76.0),
make_char("MAR-31-2000", 240.0, 652.0, 9.0, 66.0),
make_char("<CASH>", 70.0, 636.0, 9.0, 38.0),
make_char("214", 240.0, 636.0, 9.0, 18.0),
make_char("</TABLE>", 70.0, 620.0, 9.0, 46.0),
];
let table = try_build_key_value_table_from_rows(&items, 1).unwrap();
let md = table_to_markdown(&table);
assert!(md.starts_with("|Field|Value|"), "{md}");
assert!(md.contains("|<S>|<C>|"), "{md}");
assert!(md.contains("|<FISCAL-YEAR-END>|DEC-31-2000|"), "{md}");
assert!(md.contains("|<CASH>|214|"), "{md}");
assert!(!md.contains("</TABLE>"), "{md}");
}
#[test]
fn test_key_value_builder_rejects_split_prose() {
let items = vec![
+3
View File
@@ -883,6 +883,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1002,6 +1003,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
});
@@ -1078,6 +1080,7 @@ mod tests {
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
+168 -23
View File
@@ -329,7 +329,7 @@ impl ToUnicodeCMap {
if let (Some(start), Some(end), Some(base)) = (
parse_hex_u16(&start_hex),
parse_hex_u16(&end_hex),
parse_hex_u32(&base_hex),
hex_to_unicode_scalar(&base_hex),
) {
self.ranges.push((start, end, base));
}
@@ -575,32 +575,86 @@ fn parse_hex_u16(hex: &str) -> Option<u16> {
u16::from_str_radix(hex.trim(), 16).ok()
}
/// Parse a hex string to u32
fn parse_hex_u32(hex: &str) -> Option<u32> {
u32::from_str_radix(hex.trim(), 16).ok()
}
/// Convert a hex string to a Unicode string
/// Handles both 2-byte (BMP) and 4-byte (supplementary) codepoints
/// Convert a ToUnicode destination hex string to Unicode.
///
/// PDF ToUnicode destinations are UTF-16BE strings. Supplementary-plane
/// characters are encoded as surrogate pairs, so treating each 4-hex chunk as
/// a scalar drops emoji like D83CDF1F.
fn hex_to_unicode_string(hex: &str) -> Option<String> {
let hex = hex.trim();
let mut result = String::new();
// Process 4 hex digits at a time
let mut i = 0;
while i + 4 <= hex.len() {
if let Ok(cp) = u32::from_str_radix(&hex[i..i + 4], 16) {
if let Some(c) = char::from_u32(cp) {
result.push(c);
}
}
i += 4;
let hex: String = hex.chars().filter(|ch| !ch.is_ascii_whitespace()).collect();
if hex.is_empty() || !hex.len().is_multiple_of(2) {
return None;
}
if result.is_empty() {
None
let bytes: Option<Vec<u8>> = (0..hex.len())
.step_by(2)
.map(|i| u8::from_str_radix(&hex[i..i + 2], 16).ok())
.collect();
let bytes = bytes?;
if bytes.len().is_multiple_of(2) {
let units: Vec<u16> = bytes
.chunks_exact(2)
.map(|chunk| u16::from_be_bytes([chunk[0], chunk[1]]))
.collect();
if let Ok(result) = String::from_utf16(&units) {
if !result.is_empty() {
return Some(normalize_tounicode_destination(result));
}
}
}
// Be permissive for non-standard one-byte destinations.
if bytes.len() == 1 {
let ch = bytes[0] as char;
if !ch.is_control() || ch == '\t' || ch == '\n' {
return Some(ch.to_string());
}
}
None
}
fn normalize_tounicode_destination(text: String) -> String {
let is_multi_char = text.chars().nth(1).is_some();
// Some malformed producer CMaps put a list of alternative whitespace or
// hyphen codepoints into one destination. Keep ordinary multi-character
// mappings intact unless that malformed signature is present.
if is_multi_char
&& text.chars().all(char::is_whitespace)
&& text.chars().any(|ch| matches!(ch, '\t' | '\n' | '\r'))
{
return if text.contains('\t') {
"\t".to_string()
} else {
" ".to_string()
};
}
if is_multi_char
&& text.contains('\u{00ad}')
&& text.chars().all(|ch| {
matches!(
ch,
'-' | '\u{00ad}' | '\u{2010}' | '\u{2011}' | '\u{2012}' | '\u{2013}' | '\u{2212}'
)
})
{
return "-".to_string();
}
text
}
fn hex_to_unicode_scalar(hex: &str) -> Option<u32> {
let text = hex_to_unicode_string(hex)?;
let mut chars = text.chars();
let ch = chars.next()?;
if chars.next().is_none() {
Some(ch as u32)
} else {
Some(result)
None
}
}
@@ -2607,6 +2661,97 @@ endbfrange
assert_eq!(cmap.lookup(0x0005), Some("C".to_string()));
}
#[test]
fn test_parse_bfchar_surrogate_pair_emoji() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
2 beginbfchar
<16> <D83CDF1F>
<9D> <D83CDFAD>
endbfchar
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.code_byte_length, 1);
assert_eq!(cmap.lookup(0x16), Some("🌟".to_string()));
assert_eq!(cmap.lookup(0x9D), Some("🎭".to_string()));
}
#[test]
fn test_parse_bfrange_surrogate_pair_base() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
1 beginbfrange
<C8> <C9> <D83CDFD8>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.code_byte_length, 1);
assert_eq!(cmap.lookup(0xC8), Some("🏘".to_string()));
assert_eq!(cmap.lookup(0xC9), Some("🏙".to_string()));
}
#[test]
fn test_parse_bfrange_preserves_single_hyphen_like_base() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
1 beginbfrange
<21> <22> <2013>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.lookup(0x21), Some("".to_string()));
assert_eq!(cmap.lookup(0x22), Some("".to_string()));
}
#[test]
fn test_parse_spaced_destination_hex_without_control_noise() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
3 beginbfchar
<21> < 0009 000d 0020 00a0 >
<22> < 002d 00ad 2010 >
<23> <00a0>
endbfchar
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.lookup(0x21), Some("\t".to_string()));
assert_eq!(cmap.lookup(0x22), Some("-".to_string()));
assert_eq!(cmap.lookup(0x23), Some("\u{00a0}".to_string()));
}
#[test]
fn test_parse_preserves_valid_multi_character_destinations() {
let cmap_content = r#"
1 begincodespacerange
<00> <FF>
endcodespacerange
4 beginbfchar
<21> <002d002d>
<22> <20132013>
<23> <002000a0>
<24> <00660069>
endbfchar
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
assert_eq!(cmap.lookup(0x21), Some("--".to_string()));
assert_eq!(cmap.lookup(0x22), Some("––".to_string()));
assert_eq!(cmap.lookup(0x23), Some(" \u{00a0}".to_string()));
assert_eq!(cmap.lookup(0x24), Some("fi".to_string()));
}
#[test]
fn test_remap_to_sequential() {
// Simulate a broken CMap where GIDs are from pre-subsetting:
+32 -7
View File
@@ -116,6 +116,10 @@ pub struct TextItem {
pub is_bold: bool,
/// Whether the font is italic
pub is_italic: bool,
/// Whether the text is underlined (drawn rule/thin rect under the
/// baseline — PDFs have no underline font flag, so this is detected
/// geometrically after extraction; see `extractor::underline`).
pub is_underline: bool,
/// Type of item (text, image, link)
pub item_type: ItemType,
/// Marked Content ID from the content stream's BDC/BMC operator.
@@ -137,12 +141,17 @@ pub struct TextLine {
impl TextLine {
pub fn text(&self) -> String {
self.text_with_formatting(false, false)
self.text_with_formatting(false, false, false)
}
/// Get text with optional bold/italic markdown formatting
pub fn text_with_formatting(&self, format_bold: bool, format_italic: bool) -> String {
if !format_bold && !format_italic {
/// Get text with optional bold/italic/underline markdown formatting
pub fn text_with_formatting(
&self,
format_bold: bool,
format_italic: bool,
format_underline: bool,
) -> String {
if !format_bold && !format_italic && !format_underline {
return self.text_plain();
}
@@ -151,6 +160,7 @@ impl TextLine {
let mut result = String::new();
let mut current_bold = false;
let mut current_italic = false;
let mut current_underline = false;
for (i, item) in self.items.iter().enumerate() {
let text = item.text.as_str();
@@ -176,9 +186,13 @@ impl TextLine {
// we push text_trimmed below (which strips it).
let has_leading_space = text.starts_with(' ');
// Check for style changes
let item_bold = format_bold && item.is_bold;
let item_italic = format_italic && item.is_italic;
// Check for style changes. Underline is exclusive: `<u>` content
// stays free of `**`/`*` markers — consumers (and the eval
// harnesses this feeds) match the tag content literally, and
// mixed `<u>**x**</u>` nesting breaks that.
let item_underline = format_underline && item.is_underline;
let item_bold = format_bold && item.is_bold && !item_underline;
let item_italic = format_italic && item.is_italic && !item_underline;
// Close previous styles if they change
if current_italic && !item_italic {
@@ -189,6 +203,10 @@ impl TextLine {
result.push_str("**");
current_bold = false;
}
if current_underline && !item_underline {
result.push_str("</u>");
current_underline = false;
}
// Add space: either from spacing logic or preserved from item text
if needs_space || (has_leading_space && !result.is_empty() && !result.ends_with(' ')) {
@@ -196,6 +214,10 @@ impl TextLine {
}
// Open new styles
if item_underline && !current_underline {
result.push_str("<u>");
current_underline = true;
}
if item_bold && !current_bold {
result.push_str("**");
current_bold = true;
@@ -215,6 +237,9 @@ impl TextLine {
if current_bold {
result.push_str("**");
}
if current_underline {
result.push_str("</u>");
}
result
}
+3
View File
@@ -104,6 +104,7 @@ fn make_text_item(text: &str, x: f32, y: f32, font_size: f32, page: u32) -> Text
page,
is_bold: false,
is_italic: false,
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -129,6 +130,7 @@ fn make_text_item_with_font(
page,
is_bold: is_bold_font(font),
is_italic: is_italic_font(font),
is_underline: false,
item_type: ItemType::Text,
mcid: None,
}
@@ -1107,6 +1109,7 @@ fn test_pages_needing_ocr_field_accessible() {
page_count: 1,
processing_time_ms: 0,
pages_needing_ocr: vec![1, 3],
ocr_reasons_by_page: Vec::new(),
title: None,
confidence: 1.0,
layout: pdf_inspector::LayoutComplexity::default(),
+2 -2
View File
@@ -188,7 +188,7 @@
|156|23/7|General renovation works to Block B at Belonie Secondary School|MOE|Belvedere Builders|SR869,505.75|
|157|30/7|Procurement of Engine Block and Crankshaft for Engine A11|PUC|Ras Tek Pvt Ltd|Euro798,650.00|
|158|30/7|procurement of Wartsila Engine spares|PUC|Wartsila Eastern Africa ltd|Euro158,424.00|
|159|30/7|Proposed walkway, Drain, rock armoring , road and Bridge widening at Anse Talbot( Ex-Golden Egg)|SLTA|G&S Enterpise|SR1,113,010.00|
|159|30/7|Proposed walkway, Drain, rock armoring, road and Bridge widening at Anse Talbot( Ex-Golden Egg)|SLTA|G&S Enterpise|SR1,113,010.00|
|160|30/7|Procurement of transfer pump control panel|PUC|CA Engineering Consultancy Pte Ltd|SGD14,600.00|
|161|30/7|Consultancy service for North to South Victoria Bye- Pass road and utilities organisation|MLUH|Sonnel Seychelles LTD|SR1,332,000.00|
|162 AUG|30/7|Procurement of the supply of sodium cardonate|PUC|HPL Chemical LTD|USD42,600.00|
@@ -237,7 +237,7 @@
|201|24/9|Procurement of vehicle x 2|SLTA|Abhaye Valabhji Pty Ltd|SR1000.000.00|
||OCT|||||
|202|1/10|Proposed new traffic lane to 5th June Avenue|SLTA|Divy Constrution|SR2,864,589.00|
|203|1/10||Proposed Walkway, Drain, rock armoring , road and Bridge widening at Anse Talbot( Ex-Golden Egg) - Variations SLTA|G & S Enterprise|SR200,448.00|
|203|1/10||Proposed Walkway, Drain, rock armoring, road and Bridge widening at Anse Talbot(Ex-Golden Egg) - Variations SLTA|G & S Enterprise|SR200,448.00|
|204|1/10|Proposed Reconstrcution of Burnt House-Au Cap|MLUH|Furui Construction|SR946,130.00|
|205|1/10|Variation on the project associated with the procurement of seven 100m3/day containerised plant|PUC|Tornado Group (UAE)|USD172,500.00|
|206|1/10|Works on the breaker system at Bel Omber desalination plant|PUC|United Concrete Products (Sey)Ltd|SR1,998,993.11|
+15 -15
View File
@@ -10,7 +10,7 @@ Department of the Treasury **Internal Revenue Service**
### This publication contains:
**Form 4070A, Employees Daily Record of** Tips **Form 4070, Employees Report of Tips to** Employer
**Form 4070A,** Employees Daily Record of Tips **Form 4070,** Employees Report of Tips to Employer
For the period
@@ -22,7 +22,7 @@ Name and address of employee
**Publication 1244 (Rev. 7-96)** Cat. No. 44472W
**Instructions** You must keep sufficient proof to show the amount of your tip income for the year. A daily record of your tip income is considered sufficient proof. Keep a daily record for each workday showing the amount of cash and credit card tips received directly from customers or other employees. Also keep a record of the amount of tips, if any, you paid to other employees through tip sharing, tip pooling or other arrangements, and the names of employees to whom you paid tips. Show the date that each entry is made. This date should be on or near the date you received the tip income. You may use Form 4070A, Employees Daily Record of Tips, or any other daily record to record your tips. **Reporting Tips to Your Employer.—If you** receive tips that total $20 or more for any month while working for one employer, you must report the tips to your employer. Tips include cash left by customers, tips customers add to credit card charges, and tips you receive from other employees. You must report your tips for any one month by the 10th day of the next month. If the 10th day falls on a Saturday, Sunday, or legal holiday, you may give the report to your employer on the next business day that is not a Saturday, Sunday, or legal holiday. You must report tips that total $20 or more every month regardless of your total wages and tips for the year. You may use Form 4070, Employees Report of Tips to Employer, to report your tips to your employer. See the instructions on the back of Form 4070. You must include all tips, including tips not reported to your employer, as wages on your income tax return. You may use the last page of this publication to total your tips for the year. Your employer must withhold income, social security, and Medicare (or railroad retirement) taxes on tips you report. Your employer usually deducts the withholding due on tips from your regular wages.
**Instructions** You must keep sufficient proof to show the amount of your tip income for the year. A daily record of your tip income is considered sufficient proof. Keep a daily record for each workday showing the amount of cash and credit card tips received directly from customers or other employees. Also keep a record of the amount of tips, if any, you paid to other employees through tip sharing, tip pooling or other arrangements, and the names of employees to whom you paid tips. Show the date that each entry is made. This date should be on or near the date you received the tip income. You may use **Form 4070A**, Employees Daily Record of Tips, or any other daily record to record your tips. **Reporting Tips to Your Employer.—**If you receive tips that total $20 or more for any month while working for one employer, you must report the tips to your employer. Tips include cash left by customers, tips customers add to credit card charges, and tips you receive from other employees. You must report your tips for any one month by the 10th day of the next month. If the 10th day falls on a Saturday, Sunday, or legal holiday, you may give the report to your employer on the next business day that is not a Saturday, Sunday, or legal holiday. You must report tips that total $20 or more every month regardless of your total wages and tips for the year. You may use **Form 4070**, Employees Report of Tips to Employer, to report your tips to your employer. See the instructions on the back of Form 4070. You must include all tips, including tips not reported to your employer, as wages on your income tax return. You may use the last page of this publication to total your tips for the year. Your employer must withhold income, social security, and Medicare (or railroad retirement) taxes on tips you report. Your employer usually deducts the withholding due on tips from your regular wages.
*(continued on inside of back cover)*
@@ -30,14 +30,14 @@ Form **4070A** Employees Daily Record of Tips (Rev. July 1996) **This is a vo
Establishment name (if different)
Date Date **a. Tips received**
Date Date **a.** Tips received
**b. Credit card tips c. Tips paid out to d. Names of employees to whom you**
**b.** Credit card tips **c.** Tips paid out to **d.** Names of employees to whom you
tips of directly from customers received other employees paid tips recd. entry and other employees 1 2 3 4 5 **Subtotals** **For Paperwork Reduction Act Notice, see Instructions on the back of Form 4070. Page 1**
Date Date **a. Tips received**
Date Date **a.** Tips received
**b. Credit card tips c. Tips paid out to d. Names of employees to whom you**
**b.** Credit card tips **c.** Tips paid out to **d.** Names of employees to whom you
tips of directly from customers received other employees paid tips recd. entry and other employees
7 8 9 10 11 12 13 14 15 **Subtotals**
@@ -50,9 +50,9 @@ tips of directly from customers received other employees paid tips recd. entr
27 28 29 30 31 **Subtotals** **from pages** **1, 2, and 3** **Totals**
**1.** Report total cash tips (col. a) on Form 4070, line 1.
**2.** Report total credit card tips (col. b) on Form 4070, line 2.
**3.** Report total tips paid out (col. c) on Form 4070, line 3. **Page 4**
**1.** Report total cash tips (col. **a**) on Form 4070, line **1.**
**2.** Report total credit card tips (col. **b**) on Form 4070, line **2.**
**3.** Report total tips paid out (col. **c**) on Form 4070, line **3.** **Page 4**
Form Employees Report (Rev. July 1996)
@@ -66,17 +66,17 @@ Employers name and address (include establishment name, if different) **1** C
**3** Tips paid out
Month or shorter period in which tips were received **4** Net tips (lines 1 + 2 - 3) from, 19, to, 19 Signature Date
Month or shorter period in which tips were received **4** Net tips (lines **1 + 2 - 3**) from, 19, to, 19 Signature Date
**Paperwork Reduction Act Notice.—We ask for the** information on these forms to carry out the Internal Revenue laws of the United States. You are required to give us the information. We need it to ensure that you are complying with these laws and to allow us to figure and collect the right amount of tax. You are not required to provide the information requested on a form that is subject to the Paperwork Reduction Act unless the form displays a valid OMB control number. Books or records relating to a form or its instructions must be retained as long as their contents may become material in the administration of any Internal Revenue law. Generally, tax returns and return information are confidential, as required by Code section 6103. The time needed to complete Forms 4070 and 4070A will vary depending on individual circumstances. The estimated average times are: Recordkeeping—Form 4070, 7 min.; Form 4070A, 3 hr. and 23 min.; Learning **about the law—each form, 2 min.; Preparing Form 4070,** 13 min.; Form 4070A, 55 min.; and Copying and **providing Form 4070, 10 min.; Form 4070A, 14 min.** If you have comments concerning the accuracy of these time estimates or suggestions for making these
**Paperwork Reduction Act Notice.—**We ask for the information on these forms to carry out the Internal Revenue laws of the United States. You are required to give us the information. We need it to ensure that you are complying with these laws and to allow us to figure and collect the right amount of tax. You are not required to provide the information requested on a form that is subject to the Paperwork Reduction Act unless the form displays a valid OMB control number. Books or records relating to a form or its instructions must be retained as long as their contents may become material in the administration of any Internal Revenue law. Generally, tax returns and return information are confidential, as required by Code section 6103. The time needed to complete Forms 4070 and 4070A will vary depending on individual circumstances. The estimated average times are: **Recordkeeping**—Form 4070, 7 min.; Form 4070A, 3 hr. and 23 min.; **Learning** **about the law**—each form, 2 min.; **Preparing** Form 4070, 13 min.; Form 4070A, 55 min.; and **Copying and** **providing** Form 4070, 10 min.; Form 4070A, 14 min. If you have comments concerning the accuracy of these time estimates or suggestions for making these
forms simpler, we would be happy to hear from you. You can write to the Tax Forms Committee, Western Area Distribution Center, Rancho Cordova, CA 95743-0001. **Purpose.—Use this form to report tips you receive to** your employer. This includes cash tips, tips you receive from other employees, and credit card tips. You must report tips every month regardless of your total wages and tips for the year. However, you do not have to report tips to your employer for any month you received less than $20 in tips while working for that employer. Report tips by the 10th day of the month following the month that you receive them. If the 10th day is a Saturday, Sunday, or legal holiday, report tips by the next day that is not a Saturday, Sunday, or legal holiday. See Pub. 531, Reporting Tip Income, for more information. You can get additional copies of Pub. 1244, Employees Daily Record of Tips and Report to Employer, which contains both Forms 4070A and 4070, by calling 1-800-TAX-FORM (1-800-829-3676).
forms simpler, we would be happy to hear from you. You can write to the Tax Forms Committee, Western Area Distribution Center, Rancho Cordova, CA 95743-0001. **Purpose.—**Use this form to report tips you receive to your employer. This includes cash tips, tips you receive from other employees, and credit card tips. You must report tips every month regardless of your total wages and tips for the year. However, you do not have to report tips to your employer for any month you received less than $20 in tips while working for that employer. Report tips by the 10th day of the month following the month that you receive them. If the 10th day is a Saturday, Sunday, or legal holiday, report tips by the next day that is not a Saturday, Sunday, or legal holiday. See **Pub. 531**, Reporting Tip Income, for more information. You can get additional copies of **Pub. 1244**, Employees Daily Record of Tips and Report to Employer, which contains both Forms 4070A and 4070, by calling 1-800-TAX-FORM (1-800-829-3676).
**Instructions (continued)**
**Instructions** *(continued)*
**Unreported Tips.—If you received tips of $20 or** more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you must use Form 1040 and Form 4137, Social Security and Medicare Tax on Unreported Tip Income, to report them. You may not use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act cannot use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—Get Pub. 531, Reporting** Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—If you do not keep a daily** record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
**Unreported Tips.—**If you received tips of $20 or more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you **must** use Form 1040 and **Form 4137,** Social Security and Medicare Tax on Unreported Tip Income, to report them. You may **not** use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act **cannot** use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—**Get **Pub. 531,** Reporting Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—**If you do not keep a daily record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
### Instructions (continued)
**Instructions** *(continued)*
Use this space to total your tips for the year
+5 -3
View File
@@ -6,7 +6,7 @@
8 4 Z E L L / L U R I E R E A L E S T A T E C E N T E R
**Table I: Cap rate correlations** **Cap Rate Correlation With:*** **BBB Corp** **10-Year Bond Yield S&P Dividend** **Treasury (10-15 yr) Yield** Multifamily 0.187 0.771 0.068 Industrial-0.221 0.748-0.307 CBD Office-0.449 0.694-0.458 Retail-0.181 0.649-02.58
**Table I:** Cap rate correlations **Cap Rate Correlation With:*** **BBB Corp** **10-Year Bond Yield S&P Dividend** **Treasury (10-15 yr) Yield** Multifamily 0.187 0.771 0.068 Industrial-0.221 0.748-0.307 CBD Office-0.449 0.694-0.458 Retail-0.181 0.649-02.58
* Based on 25 years of data for the 10-yrT & S&P DivYld; and 14 years for BBB.
**Figure 1:** NCREIF cap rates vs. 10-yearTreasury
@@ -20,7 +20,9 @@ R E V I E W 8 5
**Figure 2:** Capratespreadsover10-yearTreasury
**Basis Points -200** -400
**Basis Points** -200
-400
-600
@@ -32,7 +34,7 @@ R E V I E W 8 5
1982 1986 1990 1994 1998 2002 2006
**Table II: Correlationsofspreadsbypropertytype** **Correlation of Cap Rate Spreads Over Treasury** **Multifamily Industrial CBD Office**
**Table II:** Correlationsofspreadsbypropertytype **Correlation of Cap Rate Spreads Over Treasury** **Multifamily Industrial CBD Office**
||Multifamily|Industrial|CBD Office|
|---|---|---|---|
+16 -16
View File
@@ -1,8 +1,8 @@
(e) [Reserved]. For further guidance, see §1.1563-3T(e)(1). Par. 50. Section 1.1563-3T is added to read as follows:
§1.1563-3T Rules for determining stock ownership (temporary).
<u>§1.1563-3T Rules for determining stock ownership (temporary)</u>.
(a) through (d)(2)(iii) [Reserved]. For further guidance, see §1.1563-3(a)
through (d)(2)(iii). (iv) Statement. If the application of paragraph (d)(2)(ii) or (iii) of §1.1563-3 does not result in a corporation being treated as a component member of only one controlled group of corporations on a December 31, then such corporation will be treated as a component member of only one such group on such date. Such corporation may elect the group in which it is to be included by including on or with its income tax return a statement entitled, “STATEMENT TO ELECT CONTROLLED GROUP PURSUANT TO §1.1563-3T(d)(2)(iv).” The statement must include--
through (d)(2)(iii). (iv) <u>Statement</u>. If the application of paragraph (d)(2)(ii) or (iii) of §1.1563-3 does not result in a corporation being treated as a component member of only one controlled group of corporations on a December 31, then such corporation will be treated as a component member of only one such group on such date. Such corporation may elect the group in which it is to be included by including on or with its income tax return a statement entitled, “STATEMENT TO ELECT CONTROLLED GROUP PURSUANT TO §1.1563-3T(d)(2)(iv).” The statement must include--
(A) A description of each of the controlled groups in which the corporation
could be included. The description must include the name and employer identification number of each component member of each such group and the stock ownership of the component members of each such group; and
@@ -10,7 +10,7 @@ could be included. The description must include the name and employer identifica
(B) The following representation: [INSERT NAME AND EMPLOYER
IDENTIFICATION NUMBER OF CORPORATION] ELECTS TO BE TREATED AS A COMPONENT MEMBER OF THE [INSERT DESIGNATION OF GROUP].
(v) Election-- (A) Election filed. An election filed under paragraph (d)(2)(iv) of
(v) <u>Election</u>-- (A) <u>Election filed</u>. An election filed under paragraph (d)(2)(iv) of
this section is irrevocable and effective until paragraph (d)(2)(ii) or (iii) of §1.1563-3 applies or until a change in the stock ownership of the corporation results in
|termination of membership in the controlled group in which such corporation has||
@@ -30,47 +30,47 @@ Federal income tax return (including any amended return filed on or before the d
2006.
(2) Expiration date. The applicability of this section will expire on May 26,
2009. Par. 51. Section 1.6012-2 is amended by revising paragraph (c) and adding paragraph (k) to read as follows: §1.6012-2 Corporations required to make returns of income.
2009. Par. 51. Section 1.6012-2 is amended by revising paragraph (c) and adding paragraph (k) to read as follows: <u>§1.6012-2 Corporations required to make returns of income</u>.
* * * * *
(c) [Reserved]. For further guidance, see §1.6012-2T(c).
* * * * *
(k) [Reserved]. For further guidance, see §1.6012-2T(k)(1).
Par. 52. Section 1.6012-2T is added to read as follows: §1.6012-2T Corporations required to make returns of income (temporary).
Par. 52. Section 1.6012-2T is added to read as follows: <u>§1.6012-2T Corporations required to make returns of income (temporary)</u>.
(a) through (b) [Reserved]. For further guidance, see §1.6012-2(a) through
(b).
(c) Insurance companies-- (1) Domestic life insurance companies-- (i) In
general. A life insurance company subject to tax under section 801 shall make a return on Form 1120L. Except as provided in paragraph (c)(4) of this section, such company shall file with its return--
<u>general</u>. A life insurance company subject to tax under section 801 shall make a return on Form 1120L. Except as provided in paragraph (c)(4) of this section, such company shall file with its return--
(A) A copy of its annual statement which shows the reserves used by the
company in computing the taxable income reported on its return; and
(B) A copy of Schedule A (real estate) and of Schedule D (bonds and stocks),
or any successor thereto, of such annual statement. (ii) Mutual savings banks. Mutual savings banks conducting life insurance business and meeting the requirements of section 594 are subject to partial tax computed on Form 1120 and partial tax computed on Form 1120L. The Form 1120L is attached as a schedule to Form 1120, together with the annual statement and schedules required to be filed with Form 1120L.
or any successor thereto, of such annual statement. (ii) <u>Mutual savings banks</u>. Mutual savings banks conducting life insurance business and meeting the requirements of section 594 are subject to partial tax computed on Form 1120 and partial tax computed on Form 1120L. The Form 1120L is attached as a schedule to Form 1120, together with the annual statement and schedules required to be filed with Form 1120L.
(2) Domestic nonlife insurance companies. Every domestic insurance
(2) <u>Domestic nonlife insurance companies</u>. Every domestic insurance
company other than a life insurance company shall make a return on Form 1120PC. This includes organizations described in section 501(m)(1) that provide commercial- type insurance and organizations described in section 833. Except as provided in paragraph (c)(4) of this section, such company shall file with its return a copy of its
annual statement (or a pro forma annual statement), including the underwriting and investment exhibit for the year covered by such return.
(3) Foreign insurance companies. The provisions of paragraphs (c)(1) and
(3) <u>Foreign insurance companies</u>. The provisions of paragraphs (c)(1) and
(c)(2) of this section concerning the returns and statements of insurance companies subject to tax under section 801 or section 831 also apply to foreign insurance companies subject to tax under those sections, except that the copy of the annual statement required to be submitted with the return shall, in the case of a foreign insurance company that is not required to file an annual statement, be a copy of the pro forma annual statement relating to the United States business of such company.
(4) Exception for insurance companies filing their Federal income tax returns
electronically. If an insurance company described in paragraph (c)(1), (c)(2), or
(4) <u>Exception for insurance companies filing their Federal income tax returns</u>
<u>electronically</u>. If an insurance company described in paragraph (c)(1), (c)(2), or
(c)(3) of this section files its Federal income tax return electronically, it should not include on or with such return its annual statement (or pro forma annual statement), or any portion thereof. Such statement must be available at all times for inspection by authorized Internal Revenue Service officers or employees and retained for so long as such statements may be material in the administration of any internal revenue law. See §1.6001-1(e).
(5) Definition. For purposes of this section, the term annual statement means
(5) <u>Definition</u>. For purposes of this section, the term <u>annual statement</u> means
the annual statement, the form of which is approved by the National Association of Insurance Commissioners (NAIC), which is filed by an insurance company for the year with the insurance departments of States, Territories, and the District of
Columbia. The term annual statement also includes a pro forma annual statement if the insurance company is not required to file the NAIC annual statement.
(d) through (j) [Reserved]. For further guidance, see §1.6012-2(d) through (j).
(k) Effective date-- (1) Applicability date. This section applies to any original
(k) <u>Effective date</u>-- (1) <u>Applicability date</u>. This section applies to any original
Federal income tax return (including any amended return filed on or before the due date (including extensions) of such original return) timely filed on or after May 30,
2006.
(2) Expiration date. The applicability of this section will expire on May 26,
(2) <u>Expiration date</u>. The applicability of this section will expire on May 26,
2009.
|||Par. 53. For each entry in the “Location” column of the following table,|
@@ -165,7 +165,7 @@ section and paragraph
PART 602--OMB CONTROL NUMBERS UNDER THE PAPERWORK REDUCTION ACT Par. 54. The authority citation for part 602 continues to read as follows: Authority: 26 U.S.C. 7805. Par. 55. In §602.101, paragraph (b) is amended to read as follows:
1. The following entries to the table are removed:
§602.101 OMB Control numbers.
<u>§602.101 OMB Control numbers</u>.
* * * * *
(b) * * *
@@ -180,7 +180,7 @@ CFR part or section where Current OMB identified or described control No.
1.1081-11………………………………………………………………. 1545-2019
* * * * * **______________________________________________________________**
2. The following entries are added in numerical order to the table:
§602.101 OMB Control numbers.
<u>§602.101 OMB Control numbers</u>.
* * * * *
(b) * * *
+2 -2
View File
@@ -26,7 +26,7 @@ A.P., NIST Standard Reference in cubic meters per kilogram Database 23, NIST the
##### Physical Properties
|Chemical Formula|CCl2F2|
|Chemical Formula|CCl₂F₂|
|---|---|
|Molecular mass|120.91|
|Boiling Point At one atmosphere|-29.75°C|
@@ -45,7 +45,7 @@ l
|Temp|Pressure||Volume|||Density||Enthalpy|||Entropy|Temp|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
|°C|[kPa]|[m3 Liquid v f|/kg]|Vapour v g|Liquid d f|[kg/m3] Vapour d g|Liquid H f|[kJ/kg] Latent H fg|Vapour H g|Liquid S f|[kJ/K-kg] Vapour S g|°C|
|°C|[kPa]|[m³ Liquid v f|/kg]|Vapour v g|Liquid d f|[kg/m³] Vapour d g|Liquid H f|[kJ/kg] Latent H fg|Vapour H g|Liquid S f|[kJ/K-kg] Vapour S g|°C|
|-100|1.2|0.0006|10.0000|1679.0|0.100|113.3|192.8|306.1|0.6077|1.7210|-100|
|---|---|---|---|---|---|---|---|---|---|---|---|