Compare commits
25
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4bee4f993b | ||
|
|
1719d24871 | ||
|
|
f114e79c8b | ||
|
|
544538b99f | ||
|
|
c8ba909407 | ||
|
|
076183e2e4 | ||
|
|
ec6e54afb8 | ||
|
|
89dd20d02c | ||
|
|
3d33ff3dbd | ||
|
|
75e9b09593 | ||
|
|
f4aab3b36f | ||
|
|
f65b25906c | ||
|
|
9947485a92 | ||
|
|
7054d6aa69 | ||
|
|
a67ee03269 | ||
|
|
965dc65f1b | ||
|
|
36dd5fa426 | ||
|
|
fabec0aec3 | ||
|
|
1f28c00a13 | ||
|
|
f4b8c9e854 | ||
|
|
69039f2728 | ||
|
|
3cca6446bd | ||
|
|
493fed498e | ||
|
|
436af97038 | ||
|
|
f731e1191c |
@@ -24,6 +24,9 @@ jobs:
|
||||
- name: Cache cargo
|
||||
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Run tests
|
||||
run: cargo test --verbose
|
||||
|
||||
|
||||
@@ -31,6 +31,9 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Check if version changed
|
||||
id: check
|
||||
run: |
|
||||
|
||||
@@ -28,6 +28,9 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Check if version changed
|
||||
id: check
|
||||
run: |
|
||||
|
||||
@@ -28,6 +28,9 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Check package version
|
||||
id: check
|
||||
run: |
|
||||
|
||||
@@ -28,6 +28,9 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Check if version changed
|
||||
id: check
|
||||
run: |
|
||||
|
||||
+1
-2
@@ -31,9 +31,8 @@ Thumbs.db
|
||||
napi/index.js
|
||||
napi/index.d.ts
|
||||
|
||||
# Local samples and scripts
|
||||
# Local samples
|
||||
samples/
|
||||
scripts/
|
||||
|
||||
# Test output
|
||||
test_output/
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "pdf-inspector"
|
||||
version = "0.1.7"
|
||||
version = "1.14.2"
|
||||
edition = "2021"
|
||||
autobins = false
|
||||
authors = ["Firecrawl Team"]
|
||||
|
||||
@@ -110,7 +110,7 @@ Or add it manually:
|
||||
|
||||
```toml
|
||||
[dependencies]
|
||||
pdf-inspector = "0.1"
|
||||
pdf-inspector = "1"
|
||||
```
|
||||
|
||||
```rust
|
||||
|
||||
+4
-3
@@ -5,14 +5,15 @@
|
||||
If you believe you've found a security vulnerability in pdf-inspector, please
|
||||
report it privately so we can fix it before public disclosure.
|
||||
|
||||
**Preferred:** Email **help@firecrawl.dev** with:
|
||||
**Preferred:** Submit through Firecrawl's Bugcrowd vulnerability disclosure
|
||||
program at <https://bugcrowd.com/engagements/firecrawl-vdp-ess>. Please include:
|
||||
|
||||
- A description of the issue and its impact
|
||||
- Steps to reproduce (a minimal PDF or input that triggers the bug is ideal)
|
||||
- The version or commit hash of pdf-inspector you tested against
|
||||
|
||||
**Alternative:** Use GitHub's private vulnerability reporting under the
|
||||
[Security tab](https://github.com/firecrawl/pdf-inspector/security/advisories/new).
|
||||
**Alternative:** If you'd rather not use Bugcrowd, email
|
||||
**help@firecrawl.dev** with the same details.
|
||||
|
||||
We'll acknowledge your report in a timely manner and keep you updated on
|
||||
remediation progress. Please do not open a public GitHub issue for security
|
||||
|
||||
+43
-24
@@ -1,38 +1,57 @@
|
||||
# Publishing
|
||||
|
||||
The Rust crate is published to [crates.io](https://crates.io/crates/pdf-inspector) with trusted publishing from GitHub Actions. The first release was published manually; future releases publish from `.github/workflows/publish-crate.yml` when a `Cargo.toml` version change lands on `main`.
|
||||
Every pdf-inspector distribution uses one shared semantic version:
|
||||
|
||||
## crates.io Trusted Publisher
|
||||
- Rust crate: `pdf-inspector`
|
||||
- Python package: `pdf-inspector`
|
||||
- Node package: `@firecrawl/pdf-inspector` and its platform packages
|
||||
- Browser package: `@firecrawl/pdf-inspector-wasm`
|
||||
- Internal NAPI and WASM Rust crates
|
||||
|
||||
Configure the trusted publisher for the `pdf-inspector` crate with:
|
||||
`Cargo.toml` is the canonical version source. Update every manifest and lockfile
|
||||
with:
|
||||
|
||||
- Repository: `firecrawl/pdf-inspector`
|
||||
- Workflow: `publish-crate.yml`
|
||||
- Environment: `crates-io`
|
||||
```bash
|
||||
python3 scripts/version.py <version>
|
||||
```
|
||||
|
||||
The workflow uses `rust-lang/crates-io-auth-action@v1` to exchange GitHub's OIDC token for a short-lived crates.io token, then passes it to `cargo publish`.
|
||||
Verify that nothing has diverged with:
|
||||
|
||||
## Release Steps
|
||||
```bash
|
||||
python3 scripts/version.py --check
|
||||
```
|
||||
|
||||
1. Update `version` in `Cargo.toml`.
|
||||
2. Merge the version bump to `main`.
|
||||
3. The publish workflow compares the new `Cargo.toml` version with `HEAD~1`, runs `cargo publish --dry-run`, then publishes if that version is not already on crates.io.
|
||||
CI and every publishing workflow run this check before building or publishing.
|
||||
|
||||
If `Cargo.toml` changes without a package version bump, the workflow exits without publishing.
|
||||
## Release steps
|
||||
|
||||
## Browser WebAssembly package
|
||||
1. Choose the next shared semantic version and run `scripts/version.py`.
|
||||
2. Review the manifest and lockfile changes in the version-bump pull request.
|
||||
3. Merge the pull request to `main`.
|
||||
4. The crates.io, PyPI, Node, and WASM workflows independently build and
|
||||
publish that version from the same commit.
|
||||
5. After all registries succeed, create one `v<version>` GitHub release that
|
||||
links to each package and describes changes since the previous shared tag.
|
||||
|
||||
The browser package is published as `@firecrawl/pdf-inspector-wasm`. Its version lives in `wasm/Cargo.toml`, and `.github/workflows/publish-wasm.yml` builds the `web` target with `wasm-pack` before publishing the generated package.
|
||||
The independent workflows are intentionally idempotent. A manual dispatch from
|
||||
`main` can repair a partial release, and already-published artifacts are skipped.
|
||||
|
||||
The npm package must exist before a trusted publisher can be configured. For the first release only:
|
||||
## Trusted publishers
|
||||
|
||||
1. Build with `wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release`.
|
||||
2. Inspect with `npm pack --dry-run ./wasm/pkg`.
|
||||
3. Publish with `npm publish ./wasm/pkg --access public` from an authorized maintainer session.
|
||||
4. In the package settings on npm, configure the GitHub Actions trusted publisher:
|
||||
- Organization: `firecrawl`
|
||||
- Repository: `pdf-inspector`
|
||||
- Workflow: `publish-wasm.yml`
|
||||
- Allowed action: `npm publish`
|
||||
The repositories use GitHub Actions OIDC instead of long-lived registry tokens.
|
||||
Configure each registry's trusted publisher for `firecrawl/pdf-inspector` and
|
||||
its corresponding workflow:
|
||||
|
||||
After that one-time bootstrap, bumping the version in `wasm/Cargo.toml` and merging it to `main` publishes through OIDC. Until the package exists, the workflow exits cleanly without attempting an unauthenticated first publish. See npm's [trusted publishing documentation](https://docs.npmjs.com/trusted-publishers/) for the registry-side setup.
|
||||
- crates.io: `publish-crate.yml`, environment `crates-io`
|
||||
- PyPI: `publish-pypi.yml`, environment `pypi`
|
||||
- npm Node package: `publish.yml`
|
||||
- npm WASM package: `publish-wasm.yml`
|
||||
|
||||
The WASM package must exist before npm trusted publishing can be configured. If
|
||||
it ever needs to be bootstrapped again, build and inspect it before publishing:
|
||||
|
||||
```bash
|
||||
wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release
|
||||
npm pack --dry-run ./wasm/pkg
|
||||
npm publish ./wasm/pkg --access public
|
||||
```
|
||||
|
||||
@@ -80,6 +80,17 @@ for page in result.pages:
|
||||
|
||||
# Restrict to specific 0-indexed pages (preserves caller order)
|
||||
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
|
||||
|
||||
# Structure-tree elements from tagged PDFs (empty list when untagged).
|
||||
# Pages are 1-indexed to match TextItem.page, so (page, mcid) joins directly
|
||||
# against extract_text_with_positions — e.g. to recover real heading levels:
|
||||
elements = pdf_inspector.extract_structure_elements("tagged.pdf")
|
||||
roles = {(e.page, e.mcid): e.role for e in elements}
|
||||
headings = [
|
||||
item.text
|
||||
for item in pdf_inspector.extract_text_with_positions("tagged.pdf")
|
||||
if item.mcid is not None and roles.get((item.page, item.mcid), "").startswith("H")
|
||||
]
|
||||
```
|
||||
|
||||
## API reference
|
||||
@@ -100,6 +111,8 @@ result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
|
||||
| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes |
|
||||
| `extract_pages_markdown(path, pages=None)` | Per-page Markdown + layout metadata (all pages by default) |
|
||||
| `extract_pages_markdown_bytes(data, pages=None)` | Per-page Markdown from bytes |
|
||||
| `extract_structure_elements(path, pages=None)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
|
||||
| `extract_structure_elements_bytes(data, pages=None)` | Structure-tree elements from bytes |
|
||||
|
||||
## Types
|
||||
|
||||
@@ -144,6 +157,12 @@ class TextItem: # extract_text_with_positions
|
||||
is_underline: bool
|
||||
is_strikeout: bool
|
||||
item_type: str
|
||||
mcid: int | None # marked-content ID for tagged PDFs (None otherwise)
|
||||
|
||||
class StructureElement: # extract_structure_elements
|
||||
page: int # 1-indexed (matches TextItem.page)
|
||||
mcid: int
|
||||
role: str # "H1".."H6", "P", "Table", ... (resolved via /RoleMap)
|
||||
|
||||
class RegionText: # extract_text_in_regions
|
||||
text: str
|
||||
|
||||
+32
-1
@@ -138,6 +138,34 @@ for page in &result.pages {
|
||||
println!("Complex layout? {}", result.is_complex);
|
||||
```
|
||||
|
||||
Extract structure-tree elements from tagged PDFs, and join them against
|
||||
`extract_text_with_positions` to attach semantic roles (heading levels,
|
||||
paragraphs, table cells) to extracted text:
|
||||
|
||||
```rust
|
||||
use pdf_inspector::{extract_structure_elements, extract_text_with_positions};
|
||||
use std::collections::HashMap;
|
||||
|
||||
// One entry per marked-content reference, sorted by (page, mcid); empty for
|
||||
// untagged PDFs. Pages are 1-indexed to match `TextItem::page`, so the
|
||||
// (page, mcid) pair is a direct join key.
|
||||
let elements = extract_structure_elements("tagged.pdf", None)?;
|
||||
let roles: HashMap<(u32, i64), &str> = elements
|
||||
.iter()
|
||||
.map(|e| ((e.page, e.mcid), e.role.as_str()))
|
||||
.collect();
|
||||
|
||||
for item in extract_text_with_positions("tagged.pdf")? {
|
||||
if let Some(mcid) = item.mcid {
|
||||
if let Some(role) = roles.get(&(item.page, mcid)) {
|
||||
if role.starts_with('H') {
|
||||
println!("{}: {}", role, item.text);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Processing modes
|
||||
|
||||
| Mode | What it does | Returns |
|
||||
@@ -163,6 +191,8 @@ println!("Complex layout? {}", result.is_complex);
|
||||
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
|
||||
| `extract_pages_markdown(path, pages)` | Per-page Markdown + layout metadata (file) |
|
||||
| `extract_pages_markdown_mem(bytes, pages)` | Per-page Markdown from bytes |
|
||||
| `extract_structure_elements(path, pages)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
|
||||
| `extract_structure_elements_mem(bytes, pages)` | Structure-tree elements from bytes |
|
||||
|
||||
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
|
||||
|
||||
@@ -178,7 +208,8 @@ Low-level detection functions are also available via the `detector` module (`det
|
||||
| `DetectionConfig` | Configuration for detection: scan strategy, thresholds |
|
||||
| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` |
|
||||
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
|
||||
| `TextItem` | Text with position, font info, and page number |
|
||||
| `TextItem` | Text with position, font info, page number, and optional structure-tree `mcid` |
|
||||
| `StructureElement` | Tagged-PDF structure reference: page (1-indexed), mcid, role (`"H1"`..`"H6"`, `"P"`, …) |
|
||||
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
|
||||
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
|
||||
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
|
||||
|
||||
Generated
+2
-2
@@ -851,7 +851,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector"
|
||||
version = "0.1.7"
|
||||
version = "1.14.2"
|
||||
dependencies = [
|
||||
"env_logger",
|
||||
"include_dir",
|
||||
@@ -867,7 +867,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector-napi"
|
||||
version = "0.2.2"
|
||||
version = "1.14.2"
|
||||
dependencies = [
|
||||
"napi",
|
||||
"napi-build",
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "pdf-inspector-napi"
|
||||
version = "0.2.2"
|
||||
version = "1.14.2"
|
||||
edition = "2021"
|
||||
|
||||
[lib]
|
||||
|
||||
@@ -83,6 +83,22 @@ for (const region of result[0].regions) {
|
||||
}
|
||||
```
|
||||
|
||||
### Async variants
|
||||
|
||||
`processPdf`, `classifyPdf`, and `extractPagesMarkdown` are synchronous and parse on the calling thread — in Node, that's the event loop. For a one-off call in a script that's fine, but in a server a large document can hold the loop for tens to hundreds of milliseconds.
|
||||
|
||||
`processPdfAsync`, `classifyPdfAsync`, and `extractPagesMarkdownAsync` take the same arguments and produce the same results, but run the parse on the libuv thread pool and return a promise, keeping the event loop free. The input buffer is copied before the call returns, so it's safe to reuse or mutate immediately:
|
||||
|
||||
```typescript
|
||||
import { classifyPdfAsync, extractPagesMarkdownAsync } from '@firecrawl/pdf-inspector'
|
||||
|
||||
const classification = await classifyPdfAsync(pdf)
|
||||
if (classification.pdfType === 'TextBased') {
|
||||
const { pages } = await extractPagesMarkdownAsync(pdf)
|
||||
// ...
|
||||
}
|
||||
```
|
||||
|
||||
## Types
|
||||
|
||||
```typescript
|
||||
|
||||
+6
-6
@@ -8,12 +8,12 @@
|
||||
"@napi-rs/cli": "^3.4.1",
|
||||
},
|
||||
"optionalDependencies": {
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.2",
|
||||
},
|
||||
},
|
||||
},
|
||||
|
||||
+7
-7
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "@firecrawl/pdf-inspector",
|
||||
"version": "1.12.0",
|
||||
"version": "1.14.2",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
@@ -52,11 +52,11 @@
|
||||
"@napi-rs/cli": "^3.4.1"
|
||||
},
|
||||
"optionalDependencies": {
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.12.0",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.12.0"
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.2",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.2"
|
||||
}
|
||||
}
|
||||
|
||||
+236
-41
@@ -89,6 +89,11 @@ pub struct TextItem {
|
||||
pub item_type: ItemType,
|
||||
/// URL for link items, `None` for other types.
|
||||
pub link_url: Option<String>,
|
||||
/// Marked Content ID from the content stream's BDC/BMC operator, `None`
|
||||
/// when the text is not part of marked content. Join with the
|
||||
/// `page`/`mcid` pairs from [`extractStructureElements`] to attach
|
||||
/// structure-tree roles (headings, paragraphs, …) in tagged PDFs.
|
||||
pub mcid: Option<i64>,
|
||||
}
|
||||
|
||||
/// A page's regions for text extraction: (page_index_0based, bboxes).
|
||||
@@ -153,9 +158,7 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
|
||||
}
|
||||
}
|
||||
|
||||
fn to_napi_page_ocr_reasons(
|
||||
reasons: Vec<pdf_inspector::PageOcrReasons>,
|
||||
) -> Vec<PageOcrReasons> {
|
||||
fn to_napi_page_ocr_reasons(reasons: Vec<pdf_inspector::PageOcrReasons>) -> Vec<PageOcrReasons> {
|
||||
reasons
|
||||
.into_iter()
|
||||
.map(|reason| PageOcrReasons {
|
||||
@@ -202,6 +205,31 @@ where
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Shared implementations (single body behind sync and async entry points)
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn process_pdf_impl(bytes: &[u8], pages: Option<Vec<u32>>) -> Result<PdfResult> {
|
||||
let mut opts = pdf_inspector::PdfOptions::new();
|
||||
if let Some(p) = pages {
|
||||
opts = opts.pages(p);
|
||||
}
|
||||
let result = pdf_inspector::process_pdf_mem_with_options(bytes, opts)
|
||||
.map_err(|e| to_napi_err(e, "process_pdf"))?;
|
||||
Ok(to_napi_result(result))
|
||||
}
|
||||
|
||||
fn classify_pdf_impl(bytes: &[u8]) -> Result<PdfClassification> {
|
||||
let result =
|
||||
pdf_inspector::classify_pdf_mem(bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
|
||||
Ok(PdfClassification {
|
||||
pdf_type: convert_pdf_type(result.pdf_type),
|
||||
page_count: result.page_count,
|
||||
pages_needing_ocr: result.pages_needing_ocr,
|
||||
confidence: result.confidence as f64,
|
||||
})
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Public NAPI API
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -210,15 +238,7 @@ where
|
||||
#[napi]
|
||||
pub fn process_pdf(buffer: Buffer, pages: Option<Vec<u32>>) -> Result<PdfResult> {
|
||||
let bytes: Vec<u8> = buffer.to_vec();
|
||||
catch_panic("process_pdf", move || {
|
||||
let mut opts = pdf_inspector::PdfOptions::new();
|
||||
if let Some(p) = pages {
|
||||
opts = opts.pages(p);
|
||||
}
|
||||
let result = pdf_inspector::process_pdf_mem_with_options(&bytes, opts)
|
||||
.map_err(|e| to_napi_err(e, "process_pdf"))?;
|
||||
Ok(to_napi_result(result))
|
||||
})
|
||||
catch_panic("process_pdf", move || process_pdf_impl(&bytes, pages))
|
||||
}
|
||||
|
||||
/// Fast detection only — no text extraction or markdown.
|
||||
@@ -238,16 +258,7 @@ pub fn detect_pdf(buffer: Buffer) -> Result<PdfResult> {
|
||||
#[napi]
|
||||
pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
|
||||
let bytes: Vec<u8> = buffer.to_vec();
|
||||
catch_panic("classify_pdf", move || {
|
||||
let result =
|
||||
pdf_inspector::classify_pdf_mem(&bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
|
||||
Ok(PdfClassification {
|
||||
pdf_type: convert_pdf_type(result.pdf_type),
|
||||
page_count: result.page_count,
|
||||
pages_needing_ocr: result.pages_needing_ocr,
|
||||
confidence: result.confidence as f64,
|
||||
})
|
||||
})
|
||||
catch_panic("classify_pdf", move || classify_pdf_impl(&bytes))
|
||||
}
|
||||
|
||||
/// Extract plain text from a PDF Buffer.
|
||||
@@ -300,12 +311,61 @@ pub fn extract_text_with_positions(
|
||||
is_strikeout: item.is_strikeout,
|
||||
item_type,
|
||||
link_url,
|
||||
mcid: item.mcid,
|
||||
}
|
||||
})
|
||||
.collect())
|
||||
})
|
||||
}
|
||||
|
||||
/// One structure-tree element reference from a tagged PDF.
|
||||
#[napi(object)]
|
||||
pub struct StructureElementJs {
|
||||
/// 1-indexed page number (matches `TextItem.page`).
|
||||
pub page: u32,
|
||||
/// Marked Content ID from the page's content stream (matches
|
||||
/// `TextItem.mcid`).
|
||||
pub mcid: i64,
|
||||
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
|
||||
/// Custom tags are resolved through the document's role map; tags with
|
||||
/// no standard mapping are returned verbatim.
|
||||
pub role: String,
|
||||
}
|
||||
|
||||
/// Extract structure-tree element references from a tagged PDF.
|
||||
///
|
||||
/// Parses the document's structure tree (when present) and returns one
|
||||
/// entry per marked-content reference, resolved to its 1-indexed page,
|
||||
/// MCID, and structure type name. Returns an empty array when the PDF is
|
||||
/// not tagged.
|
||||
///
|
||||
/// Join `(page, mcid)` against the `page`/`mcid` fields from
|
||||
/// [`extractTextWithPositions`] to attach heading levels (H1..H6) and other
|
||||
/// semantic roles to extracted text.
|
||||
///
|
||||
/// Pass 1-indexed page numbers (matching `TextItem.page`) to restrict
|
||||
/// output; omit `pages` for the whole document. Entries are sorted by
|
||||
/// `(page, mcid)`.
|
||||
#[napi]
|
||||
pub fn extract_structure_elements(
|
||||
buffer: Buffer,
|
||||
pages: Option<Vec<u32>>,
|
||||
) -> Result<Vec<StructureElementJs>> {
|
||||
let bytes: Vec<u8> = buffer.to_vec();
|
||||
catch_panic("extract_structure_elements", move || {
|
||||
let elements = pdf_inspector::extract_structure_elements_mem(&bytes, pages.as_deref())
|
||||
.map_err(|e| to_napi_err(e, "extract_structure_elements"))?;
|
||||
Ok(elements
|
||||
.into_iter()
|
||||
.map(|e| StructureElementJs {
|
||||
page: e.page,
|
||||
mcid: e.mcid,
|
||||
role: e.role,
|
||||
})
|
||||
.collect())
|
||||
})
|
||||
}
|
||||
|
||||
/// Extract text within bounding-box regions from a PDF.
|
||||
///
|
||||
/// For hybrid OCR: layout model detects regions in rendered images,
|
||||
@@ -633,25 +693,32 @@ pub fn extract_pages_markdown(
|
||||
) -> Result<PagesExtractionResult> {
|
||||
let bytes: Vec<u8> = buffer.to_vec();
|
||||
catch_panic("extract_pages_markdown", move || {
|
||||
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, pages.as_deref())
|
||||
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
|
||||
Ok(PagesExtractionResult {
|
||||
pages: result
|
||||
.pages
|
||||
.into_iter()
|
||||
.map(|r| PageMarkdownResult {
|
||||
page: r.page,
|
||||
markdown: r.markdown,
|
||||
needs_ocr: r.needs_ocr,
|
||||
ocr_reason: r.ocr_reason,
|
||||
})
|
||||
.collect(),
|
||||
pages_with_tables: result.pages_with_tables,
|
||||
pages_with_columns: result.pages_with_columns,
|
||||
pages_needing_ocr: result.pages_needing_ocr,
|
||||
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
|
||||
is_complex: result.is_complex,
|
||||
})
|
||||
extract_pages_markdown_impl(&bytes, pages.as_deref())
|
||||
})
|
||||
}
|
||||
|
||||
fn extract_pages_markdown_impl(
|
||||
bytes: &[u8],
|
||||
pages: Option<&[u32]>,
|
||||
) -> Result<PagesExtractionResult> {
|
||||
let result = pdf_inspector::extract_pages_markdown_mem(bytes, pages)
|
||||
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
|
||||
Ok(PagesExtractionResult {
|
||||
pages: result
|
||||
.pages
|
||||
.into_iter()
|
||||
.map(|r| PageMarkdownResult {
|
||||
page: r.page,
|
||||
markdown: r.markdown,
|
||||
needs_ocr: r.needs_ocr,
|
||||
ocr_reason: r.ocr_reason,
|
||||
})
|
||||
.collect(),
|
||||
pages_with_tables: result.pages_with_tables,
|
||||
pages_with_columns: result.pages_with_columns,
|
||||
pages_needing_ocr: result.pages_needing_ocr,
|
||||
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
|
||||
is_complex: result.is_complex,
|
||||
})
|
||||
}
|
||||
|
||||
@@ -692,3 +759,131 @@ fn to_page_region_texts(results: Vec<pdf_inspector::PageRegionResult>) -> Vec<Pa
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Async variants (libuv thread pool via AsyncTask)
|
||||
//
|
||||
// The synchronous exports above parse on the calling thread, which in Node is
|
||||
// the event loop. These `*Async` variants run the same shared implementations
|
||||
// on the libuv thread pool and hand JavaScript a promise, so servers under
|
||||
// concurrent load keep answering requests while a document parses. The sync
|
||||
// exports keep their names, signatures, and behaviour.
|
||||
//
|
||||
// Each factory copies the input Buffer to an owned `Vec<u8>` on the calling
|
||||
// (JS) thread — deliberately. JS execution is single-threaded, so no JS code
|
||||
// can mutate the buffer while the synchronous part of the call copies it.
|
||||
// Holding the napi `Buffer` and reading it from the worker instead would be
|
||||
// zero-copy, but a caller mutating the buffer before the promise settles
|
||||
// would then race the worker's reads — undefined behavior, not a recoverable
|
||||
// error (a known napi-rs soundness hazard with cross-thread Buffer access).
|
||||
// The copy is a one-time memcpy, negligible next to the parse it unblocks.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
pub struct ProcessPdfTask {
|
||||
bytes: Vec<u8>,
|
||||
pages: Option<Vec<u32>>,
|
||||
}
|
||||
|
||||
impl Task for ProcessPdfTask {
|
||||
type Output = PdfResult;
|
||||
type JsValue = PdfResult;
|
||||
|
||||
fn compute(&mut self) -> Result<Self::Output> {
|
||||
let bytes = std::mem::take(&mut self.bytes);
|
||||
let pages = self.pages.take();
|
||||
// AssertUnwindSafe: `bytes`/`pages` are moved into the closure and
|
||||
// dropped on unwind — no shared state can be observed broken.
|
||||
catch_panic(
|
||||
"process_pdf",
|
||||
panic::AssertUnwindSafe(move || process_pdf_impl(&bytes, pages)),
|
||||
)
|
||||
}
|
||||
|
||||
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
|
||||
Ok(output)
|
||||
}
|
||||
}
|
||||
|
||||
/// Async variant of [`processPdf`]: same result, but the parse runs on the
|
||||
/// libuv thread pool instead of the event loop and the call returns a
|
||||
/// promise. The buffer is copied before the call returns, so it may be
|
||||
/// reused or mutated immediately.
|
||||
// ts_return_type is required: napi-rs emits `Promise<unknown>` for
|
||||
// `AsyncTask<T>` returns without it.
|
||||
#[napi(ts_return_type = "Promise<PdfResult>")]
|
||||
pub fn process_pdf_async(buffer: Buffer, pages: Option<Vec<u32>>) -> AsyncTask<ProcessPdfTask> {
|
||||
AsyncTask::new(ProcessPdfTask {
|
||||
bytes: buffer.to_vec(),
|
||||
pages,
|
||||
})
|
||||
}
|
||||
|
||||
pub struct ClassifyPdfTask {
|
||||
bytes: Vec<u8>,
|
||||
}
|
||||
|
||||
impl Task for ClassifyPdfTask {
|
||||
type Output = PdfClassification;
|
||||
type JsValue = PdfClassification;
|
||||
|
||||
fn compute(&mut self) -> Result<Self::Output> {
|
||||
let bytes = std::mem::take(&mut self.bytes);
|
||||
catch_panic(
|
||||
"classify_pdf",
|
||||
panic::AssertUnwindSafe(move || classify_pdf_impl(&bytes)),
|
||||
)
|
||||
}
|
||||
|
||||
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
|
||||
Ok(output)
|
||||
}
|
||||
}
|
||||
|
||||
/// Async variant of [`classifyPdf`]: same result, but the classification runs
|
||||
/// on the libuv thread pool instead of the event loop and the call returns a
|
||||
/// promise. The buffer is copied before the call returns, so it may be
|
||||
/// reused or mutated immediately.
|
||||
#[napi(ts_return_type = "Promise<PdfClassification>")]
|
||||
pub fn classify_pdf_async(buffer: Buffer) -> AsyncTask<ClassifyPdfTask> {
|
||||
AsyncTask::new(ClassifyPdfTask {
|
||||
bytes: buffer.to_vec(),
|
||||
})
|
||||
}
|
||||
|
||||
pub struct ExtractPagesMarkdownTask {
|
||||
bytes: Vec<u8>,
|
||||
pages: Option<Vec<u32>>,
|
||||
}
|
||||
|
||||
impl Task for ExtractPagesMarkdownTask {
|
||||
type Output = PagesExtractionResult;
|
||||
type JsValue = PagesExtractionResult;
|
||||
|
||||
fn compute(&mut self) -> Result<Self::Output> {
|
||||
let bytes = std::mem::take(&mut self.bytes);
|
||||
let pages = self.pages.take();
|
||||
catch_panic(
|
||||
"extract_pages_markdown",
|
||||
panic::AssertUnwindSafe(move || extract_pages_markdown_impl(&bytes, pages.as_deref())),
|
||||
)
|
||||
}
|
||||
|
||||
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
|
||||
Ok(output)
|
||||
}
|
||||
}
|
||||
|
||||
/// Async variant of [`extractPagesMarkdown`]: same result, but the extraction
|
||||
/// runs on the libuv thread pool instead of the event loop and the call
|
||||
/// returns a promise. The buffer is copied before the call returns, so it
|
||||
/// may be reused or mutated immediately.
|
||||
#[napi(ts_return_type = "Promise<PagesExtractionResult>")]
|
||||
pub fn extract_pages_markdown_async(
|
||||
buffer: Buffer,
|
||||
pages: Option<Vec<u32>>,
|
||||
) -> AsyncTask<ExtractPagesMarkdownTask> {
|
||||
AsyncTask::new(ExtractPagesMarkdownTask {
|
||||
bytes: buffer.to_vec(),
|
||||
pages,
|
||||
})
|
||||
}
|
||||
|
||||
+110
@@ -2,16 +2,21 @@ import { readFileSync } from 'fs';
|
||||
import { strict as assert } from 'assert';
|
||||
import {
|
||||
processPdf,
|
||||
processPdfAsync,
|
||||
detectPdf,
|
||||
classifyPdf,
|
||||
classifyPdfAsync,
|
||||
extractText,
|
||||
extractTextWithPositions,
|
||||
extractStructureElements,
|
||||
extractTextInRegions,
|
||||
detectVectorGridInRegion,
|
||||
extractPagesMarkdown,
|
||||
extractPagesMarkdownAsync,
|
||||
} from './index.js';
|
||||
|
||||
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
|
||||
const taggedFixture = readFileSync('../tests/fixtures/firecrawl_docs_tagged.pdf');
|
||||
|
||||
// --- processPdf ---
|
||||
console.log('Testing processPdf...');
|
||||
@@ -79,6 +84,46 @@ assert.ok(page1Items.length > 0);
|
||||
assert.ok(page1Items.every(i => i.page === 1));
|
||||
console.log(' extractTextWithPositions with pages: OK');
|
||||
|
||||
// mcid: undefined on untagged PDFs, numeric on tagged marked content
|
||||
assert.ok(items.every(i => i.mcid === undefined || typeof i.mcid === 'number'));
|
||||
const taggedItems = extractTextWithPositions(taggedFixture);
|
||||
assert.ok(
|
||||
taggedItems.some(i => typeof i.mcid === 'number'),
|
||||
'tagged PDF text items should carry Marked Content IDs',
|
||||
);
|
||||
console.log(' extractTextWithPositions mcid: OK');
|
||||
|
||||
// --- extractStructureElements ---
|
||||
console.log('Testing extractStructureElements...');
|
||||
const structureElements = extractStructureElements(taggedFixture);
|
||||
assert.ok(structureElements.length > 0);
|
||||
assert.ok(structureElements.every(e => typeof e.page === 'number'));
|
||||
assert.ok(structureElements.every(e => typeof e.mcid === 'number'));
|
||||
assert.ok(structureElements.every(e => typeof e.role === 'string' && e.role.length > 0));
|
||||
assert.ok(
|
||||
structureElements.some(e => e.role === 'H1'),
|
||||
'tagged fixture should surface H1 heading roles',
|
||||
);
|
||||
|
||||
// (page, mcid) joins against extractTextWithPositions to recover heading text
|
||||
const h1Refs = new Set(
|
||||
structureElements.filter(e => e.role === 'H1').map(e => `${e.page}:${e.mcid}`),
|
||||
);
|
||||
const h1Text = taggedItems
|
||||
.filter(i => typeof i.mcid === 'number' && h1Refs.has(`${i.page}:${i.mcid}`))
|
||||
.map(i => i.text)
|
||||
.join('');
|
||||
assert.ok(h1Text.trim().length > 0, 'H1 join should recover heading text');
|
||||
|
||||
// pages filter is 1-indexed, matching TextItem.page
|
||||
const page1Elements = extractStructureElements(taggedFixture, [1]);
|
||||
assert.ok(page1Elements.length > 0);
|
||||
assert.ok(page1Elements.every(e => e.page === 1));
|
||||
|
||||
// untagged PDFs yield an empty array
|
||||
assert.deepEqual(extractStructureElements(fixture), []);
|
||||
console.log(' extractStructureElements: OK');
|
||||
|
||||
// --- extractTextInRegions ---
|
||||
console.log('Testing extractTextInRegions...');
|
||||
const regionResults = extractTextInRegions(fixture, [
|
||||
@@ -124,10 +169,75 @@ assert.equal(picked.pages[0].page, 2);
|
||||
assert.equal(picked.pages[1].page, 0);
|
||||
console.log(' extractPagesMarkdown with pages: OK');
|
||||
|
||||
// --- Async variants ---
|
||||
console.log('Testing async variants...');
|
||||
|
||||
// processPdfAsync returns a promise and matches the sync result
|
||||
const asyncResultPromise = processPdfAsync(fixture);
|
||||
assert.ok(asyncResultPromise instanceof Promise);
|
||||
const asyncResult = await asyncResultPromise;
|
||||
assert.equal(asyncResult.pdfType, result.pdfType);
|
||||
assert.equal(asyncResult.pageCount, result.pageCount);
|
||||
assert.equal(asyncResult.markdown, result.markdown);
|
||||
console.log(' processPdfAsync: OK');
|
||||
|
||||
// processPdfAsync with pages
|
||||
const asyncResult2 = await processPdfAsync(fixture, [1]);
|
||||
assert.equal(asyncResult2.markdown, result2.markdown);
|
||||
console.log(' processPdfAsync with pages: OK');
|
||||
|
||||
// classifyPdfAsync matches the sync result
|
||||
const asyncClassified = await classifyPdfAsync(fixture);
|
||||
assert.equal(asyncClassified.pdfType, classified.pdfType);
|
||||
assert.equal(asyncClassified.pageCount, classified.pageCount);
|
||||
assert.equal(asyncClassified.confidence, classified.confidence);
|
||||
assert.deepEqual(asyncClassified.pagesNeedingOcr, classified.pagesNeedingOcr);
|
||||
console.log(' classifyPdfAsync: OK');
|
||||
|
||||
// extractPagesMarkdownAsync matches the sync result
|
||||
const asyncAllPages = await extractPagesMarkdownAsync(fixture);
|
||||
assert.equal(asyncAllPages.pages.length, allPages.pages.length);
|
||||
assert.deepEqual(
|
||||
asyncAllPages.pages.map(p => p.markdown),
|
||||
allPages.pages.map(p => p.markdown),
|
||||
);
|
||||
assert.equal(asyncAllPages.isComplex, allPages.isComplex);
|
||||
console.log(' extractPagesMarkdownAsync: OK');
|
||||
|
||||
// selected pages preserve caller order
|
||||
const asyncPicked = await extractPagesMarkdownAsync(fixture, [2, 0]);
|
||||
assert.equal(asyncPicked.pages.length, 2);
|
||||
assert.equal(asyncPicked.pages[0].page, 2);
|
||||
assert.equal(asyncPicked.pages[1].page, 0);
|
||||
console.log(' extractPagesMarkdownAsync with pages: OK');
|
||||
|
||||
// input buffer is copied at call time: mutating it immediately after the
|
||||
// call must not affect the in-flight parse
|
||||
const scratch = Buffer.from(fixture);
|
||||
const inFlight = processPdfAsync(scratch);
|
||||
scratch.fill(0);
|
||||
const fromMutated = await inFlight;
|
||||
assert.equal(fromMutated.markdown, result.markdown);
|
||||
console.log(' processPdfAsync input copied at call time: OK');
|
||||
|
||||
// concurrent async calls all settle
|
||||
const [c1, c2, c3] = await Promise.all([
|
||||
processPdfAsync(fixture),
|
||||
classifyPdfAsync(fixture),
|
||||
extractPagesMarkdownAsync(fixture),
|
||||
]);
|
||||
assert.equal(c1.pdfType, 'TextBased');
|
||||
assert.equal(c2.pdfType, 'TextBased');
|
||||
assert.equal(c3.pages.length, 3);
|
||||
console.log(' concurrent async calls: OK');
|
||||
|
||||
// --- Error handling ---
|
||||
console.log('Testing error handling...');
|
||||
assert.throws(() => processPdf(Buffer.from('not a pdf')), /process_pdf/);
|
||||
assert.throws(() => classifyPdf(Buffer.from('')), /classify_pdf/);
|
||||
await assert.rejects(processPdfAsync(Buffer.from('not a pdf')), /process_pdf/);
|
||||
await assert.rejects(classifyPdfAsync(Buffer.from('')), /classify_pdf/);
|
||||
await assert.rejects(extractPagesMarkdownAsync(Buffer.from('')), /extract_pages_markdown/);
|
||||
console.log(' error handling: OK');
|
||||
|
||||
console.log('\nAll NAPI tests passed!');
|
||||
|
||||
@@ -51,6 +51,20 @@ class TextItem:
|
||||
is_underline: bool
|
||||
is_strikeout: bool
|
||||
item_type: str
|
||||
mcid: Optional[int]
|
||||
"""Marked Content ID from the content stream's BDC/BMC operator, None when
|
||||
the text is not part of marked content. Join with the (page, mcid) pairs
|
||||
from extract_structure_elements to attach structure-tree roles in tagged
|
||||
PDFs."""
|
||||
|
||||
class StructureElement:
|
||||
"""One structure-tree element reference from a tagged PDF."""
|
||||
page: int
|
||||
"""1-indexed page number (matches TextItem.page)."""
|
||||
mcid: int
|
||||
"""Marked Content ID from the page's content stream (matches TextItem.mcid)."""
|
||||
role: str
|
||||
"""Standard structure type name ("H1".."H6", "P", "Table", "TD", ...)."""
|
||||
|
||||
class RegionText:
|
||||
"""Extracted text for a single region."""
|
||||
@@ -132,6 +146,27 @@ def extract_text_with_positions_bytes(data: bytes, pages: Optional[list[int]] =
|
||||
"""Extract text with position information from bytes."""
|
||||
...
|
||||
|
||||
def extract_structure_elements(path: str, pages: Optional[list[int]] = None) -> list[StructureElement]:
|
||||
"""Extract structure-tree element references from a tagged PDF file.
|
||||
|
||||
Returns one entry per marked-content reference, resolved to its 1-indexed
|
||||
page, MCID, and structure type name ("H1".."H6", "P", "Table", ...), sorted
|
||||
by (page, mcid). Returns an empty list when the PDF is not tagged.
|
||||
|
||||
Args:
|
||||
path: Path to the PDF file.
|
||||
pages: Optional list of 1-indexed pages (matching ``TextItem.page``).
|
||||
When ``None`` (default), the whole document is returned.
|
||||
"""
|
||||
...
|
||||
|
||||
def extract_structure_elements_bytes(data: bytes, pages: Optional[list[int]] = None) -> list[StructureElement]:
|
||||
"""Extract structure-tree element references from tagged PDF bytes.
|
||||
|
||||
See :func:`extract_structure_elements` for details.
|
||||
"""
|
||||
...
|
||||
|
||||
def extract_text_in_regions(
|
||||
path: str,
|
||||
page_regions: list[tuple[int, list[list[float]]]],
|
||||
|
||||
+3
-3
@@ -4,9 +4,9 @@ build-backend = "maturin"
|
||||
|
||||
[project]
|
||||
name = "pdf-inspector"
|
||||
# Bump this to publish to PyPI — CI publishes automatically when the version
|
||||
# changes on main (same flow as napi/package.json for npm).
|
||||
version = "0.2.6"
|
||||
# Keep package versions in sync with `python3 scripts/version.py <version>`.
|
||||
# CI publishes automatically when the synchronized change lands on main.
|
||||
version = "1.14.2"
|
||||
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
|
||||
readme = "docs/python.md"
|
||||
license = { text = "MIT" }
|
||||
|
||||
@@ -0,0 +1,111 @@
|
||||
import json
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
|
||||
from version import PLATFORM_PACKAGES, check_versions, set_versions
|
||||
|
||||
|
||||
class VersionTests(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.temporary = tempfile.TemporaryDirectory()
|
||||
self.root = Path(self.temporary.name)
|
||||
(self.root / "napi").mkdir()
|
||||
(self.root / "site").mkdir()
|
||||
(self.root / "wasm").mkdir()
|
||||
|
||||
self._write_manifest("Cargo.toml", "package", "0.1.0")
|
||||
self._write_manifest("pyproject.toml", "project", "0.1.0")
|
||||
self._write_manifest("napi/Cargo.toml", "package", "0.1.0")
|
||||
self._write_manifest("wasm/Cargo.toml", "package", "0.1.0")
|
||||
|
||||
package = {
|
||||
"name": "@firecrawl/pdf-inspector",
|
||||
"version": "0.1.0",
|
||||
"optionalDependencies": {
|
||||
dependency: "0.1.0" for dependency in PLATFORM_PACKAGES
|
||||
},
|
||||
}
|
||||
(self.root / "napi/package.json").write_text(
|
||||
json.dumps(package), encoding="utf-8"
|
||||
)
|
||||
(self.root / "napi/bun.lock").write_text(
|
||||
"\n".join(
|
||||
f' "{dependency}": "0.1.0",'
|
||||
for dependency in PLATFORM_PACKAGES
|
||||
)
|
||||
+ "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
(self.root / "site/index.html").write_text(
|
||||
'https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.0/'
|
||||
'pdf_inspector_wasm.js\n',
|
||||
encoding="utf-8",
|
||||
)
|
||||
self._write_lock(
|
||||
"napi/Cargo.lock", ("pdf-inspector", "pdf-inspector-napi")
|
||||
)
|
||||
self._write_lock(
|
||||
"wasm/Cargo.lock", ("pdf-inspector", "pdf-inspector-wasm")
|
||||
)
|
||||
|
||||
def tearDown(self):
|
||||
self.temporary.cleanup()
|
||||
|
||||
def _write_manifest(self, relative, section, version):
|
||||
(self.root / relative).write_text(
|
||||
f'[{section}]\nname = "fixture"\nversion = "{version}"\n',
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
def _write_lock(self, relative, packages):
|
||||
content = "\n".join(
|
||||
f'[[package]]\nname = "{package}"\nversion = "0.1.0"\n'
|
||||
for package in packages
|
||||
)
|
||||
(self.root / relative).write_text(content, encoding="utf-8")
|
||||
|
||||
def test_updates_every_version_location(self):
|
||||
set_versions("1.14.0", self.root)
|
||||
|
||||
self.assertEqual(check_versions(self.root), "1.14.0")
|
||||
|
||||
def test_reports_a_divergent_package(self):
|
||||
self._write_manifest("wasm/Cargo.toml", "package", "0.2.0")
|
||||
|
||||
with self.assertRaisesRegex(ValueError, "WASM package: 0.2.0"):
|
||||
check_versions(self.root)
|
||||
|
||||
def test_rejects_an_invalid_version(self):
|
||||
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
|
||||
set_versions("next", self.root)
|
||||
|
||||
def test_rejects_numeric_prerelease_with_leading_zero(self):
|
||||
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
|
||||
|
||||
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
|
||||
set_versions("1.2.3-01", self.root)
|
||||
|
||||
self.assertEqual(
|
||||
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
|
||||
)
|
||||
|
||||
def test_preflight_failure_does_not_partially_update(self):
|
||||
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
|
||||
(self.root / "site/index.html").write_text(
|
||||
"missing module URL\n", encoding="utf-8"
|
||||
)
|
||||
|
||||
with self.assertRaisesRegex(ValueError, "Missing pinned WASM package URL"):
|
||||
set_versions("1.14.0", self.root)
|
||||
|
||||
self.assertEqual(
|
||||
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,262 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Keep every pdf-inspector package on one release version."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
PRERELEASE_IDENTIFIER = (
|
||||
r"(?:0|[1-9]\d*|[0-9A-Za-z-]*[A-Za-z-][0-9A-Za-z-]*)"
|
||||
)
|
||||
SEMVER = re.compile(
|
||||
r"^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)"
|
||||
rf"(?:-{PRERELEASE_IDENTIFIER}(?:\.{PRERELEASE_IDENTIFIER})*)?"
|
||||
r"(?:\+[0-9A-Za-z-]+(?:\.[0-9A-Za-z-]+)*)?$"
|
||||
)
|
||||
VERSION_LINE = re.compile(r'^(\s*version\s*=\s*")[^"]+(".*)$')
|
||||
SECTION_LINE = re.compile(r"^\s*\[([^]]+)]\s*$")
|
||||
PLATFORM_PACKAGES = (
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc",
|
||||
)
|
||||
TOML_VERSIONS = (
|
||||
("Rust crate", Path("Cargo.toml"), "package"),
|
||||
("Python package", Path("pyproject.toml"), "project"),
|
||||
("NAPI crate", Path("napi/Cargo.toml"), "package"),
|
||||
("WASM package", Path("wasm/Cargo.toml"), "package"),
|
||||
)
|
||||
LOCK_VERSIONS = (
|
||||
("NAPI lock: core", Path("napi/Cargo.lock"), "pdf-inspector"),
|
||||
("NAPI lock: binding", Path("napi/Cargo.lock"), "pdf-inspector-napi"),
|
||||
("WASM lock: core", Path("wasm/Cargo.lock"), "pdf-inspector"),
|
||||
("WASM lock: binding", Path("wasm/Cargo.lock"), "pdf-inspector-wasm"),
|
||||
)
|
||||
SITE_WASM_VERSION = re.compile(
|
||||
r"(@firecrawl/pdf-inspector-wasm@)([^/\"]+)(/pdf_inspector_wasm\.js)"
|
||||
)
|
||||
|
||||
|
||||
def _read_section_version(path: Path, section: str) -> str:
|
||||
active = False
|
||||
for line in path.read_text(encoding="utf-8").splitlines():
|
||||
section_match = SECTION_LINE.match(line)
|
||||
if section_match:
|
||||
active = section_match.group(1) == section
|
||||
elif active:
|
||||
version_match = VERSION_LINE.match(line)
|
||||
if version_match:
|
||||
return line.split('"', 2)[1]
|
||||
raise ValueError(f"No version found in [{section}] of {path}")
|
||||
|
||||
|
||||
def _write_section_version(path: Path, section: str, version: str) -> None:
|
||||
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
|
||||
active = False
|
||||
for index, line in enumerate(lines):
|
||||
section_match = SECTION_LINE.match(line)
|
||||
if section_match:
|
||||
active = section_match.group(1) == section
|
||||
elif active:
|
||||
version_match = VERSION_LINE.match(line)
|
||||
if version_match:
|
||||
newline = "\n" if line.endswith("\n") else ""
|
||||
replacement = (
|
||||
f"{version_match.group(1)}{version}"
|
||||
f"{version_match.group(2).rstrip()}"
|
||||
)
|
||||
lines[index] = (
|
||||
f"{replacement}{newline}"
|
||||
)
|
||||
path.write_text("".join(lines), encoding="utf-8")
|
||||
return
|
||||
raise ValueError(f"No version found in [{section}] of {path}")
|
||||
|
||||
|
||||
def _package_block(lines: list[str], package: str) -> tuple[int, int]:
|
||||
for start, line in enumerate(lines):
|
||||
if line.strip() != "[[package]]":
|
||||
continue
|
||||
end = next(
|
||||
(
|
||||
index
|
||||
for index in range(start + 1, len(lines))
|
||||
if lines[index].strip() == "[[package]]"
|
||||
),
|
||||
len(lines),
|
||||
)
|
||||
if any(line.strip() == f'name = "{package}"' for line in lines[start:end]):
|
||||
return start, end
|
||||
raise ValueError(f"No lockfile entry found for {package}")
|
||||
|
||||
|
||||
def _read_lock_version(path: Path, package: str) -> str:
|
||||
lines = path.read_text(encoding="utf-8").splitlines()
|
||||
start, end = _package_block(lines, package)
|
||||
for line in lines[start:end]:
|
||||
version_match = VERSION_LINE.match(line)
|
||||
if version_match:
|
||||
return line.split('"', 2)[1]
|
||||
raise ValueError(f"No version found for {package} in {path}")
|
||||
|
||||
|
||||
def _write_lock_version(path: Path, package: str, version: str) -> None:
|
||||
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
|
||||
start, end = _package_block(lines, package)
|
||||
for index in range(start, end):
|
||||
version_match = VERSION_LINE.match(lines[index])
|
||||
if version_match:
|
||||
newline = "\n" if lines[index].endswith("\n") else ""
|
||||
lines[index] = (
|
||||
f'{version_match.group(1)}{version}{version_match.group(2).rstrip()}'
|
||||
f"{newline}"
|
||||
)
|
||||
path.write_text("".join(lines), encoding="utf-8")
|
||||
return
|
||||
raise ValueError(f"No version found for {package} in {path}")
|
||||
|
||||
|
||||
def _node_versions(root: Path) -> dict[str, str]:
|
||||
package = json.loads((root / "napi/package.json").read_text(encoding="utf-8"))
|
||||
versions = {"Node package": package["version"]}
|
||||
optional = package.get("optionalDependencies", {})
|
||||
for dependency in PLATFORM_PACKAGES:
|
||||
if dependency not in optional:
|
||||
raise ValueError(f"Missing Node optional dependency: {dependency}")
|
||||
versions[f"Node optional dependency: {dependency}"] = optional[dependency]
|
||||
return versions
|
||||
|
||||
|
||||
def _bun_versions(root: Path) -> dict[str, str]:
|
||||
text = (root / "napi/bun.lock").read_text(encoding="utf-8")
|
||||
versions = {}
|
||||
for dependency in PLATFORM_PACKAGES:
|
||||
match = re.search(
|
||||
rf'"{re.escape(dependency)}": "([^"]+)"[,]', text
|
||||
)
|
||||
if not match:
|
||||
raise ValueError(f"Missing Bun lock dependency: {dependency}")
|
||||
versions[f"Bun lock: {dependency}"] = match.group(1)
|
||||
return versions
|
||||
|
||||
|
||||
def _site_wasm_version(root: Path) -> str:
|
||||
text = (root / "site/index.html").read_text(encoding="utf-8")
|
||||
match = SITE_WASM_VERSION.search(text)
|
||||
if not match:
|
||||
raise ValueError("Missing pinned WASM package URL in site/index.html")
|
||||
return match.group(2)
|
||||
|
||||
|
||||
def package_versions(root: Path = ROOT) -> dict[str, str]:
|
||||
versions = {
|
||||
label: _read_section_version(root / relative, section)
|
||||
for label, relative, section in TOML_VERSIONS
|
||||
}
|
||||
versions.update(_node_versions(root))
|
||||
versions.update(_bun_versions(root))
|
||||
versions["Website WASM module"] = _site_wasm_version(root)
|
||||
versions.update(
|
||||
{
|
||||
label: _read_lock_version(root / relative, package)
|
||||
for label, relative, package in LOCK_VERSIONS
|
||||
}
|
||||
)
|
||||
return versions
|
||||
|
||||
|
||||
def check_versions(root: Path = ROOT) -> str:
|
||||
versions = package_versions(root)
|
||||
expected = versions["Rust crate"]
|
||||
if not SEMVER.fullmatch(expected):
|
||||
raise ValueError(f"Rust crate has an invalid semantic version: {expected}")
|
||||
mismatches = {
|
||||
label: version for label, version in versions.items() if version != expected
|
||||
}
|
||||
if mismatches:
|
||||
details = "\n".join(
|
||||
f" - {label}: {version}" for label, version in mismatches.items()
|
||||
)
|
||||
raise ValueError(f"Expected every package to use {expected}:\n{details}")
|
||||
return expected
|
||||
|
||||
|
||||
def set_versions(version: str, root: Path = ROOT) -> None:
|
||||
if not SEMVER.fullmatch(version):
|
||||
raise ValueError(f"Invalid semantic version: {version}")
|
||||
|
||||
# Validate every expected location before writing the first file. This
|
||||
# prevents a stale manifest or generated file from leaving a partial bump.
|
||||
package_versions(root)
|
||||
|
||||
for _, relative, section in TOML_VERSIONS:
|
||||
_write_section_version(root / relative, section, version)
|
||||
|
||||
package_path = root / "napi/package.json"
|
||||
package = json.loads(package_path.read_text(encoding="utf-8"))
|
||||
package["version"] = version
|
||||
optional = package.get("optionalDependencies", {})
|
||||
for dependency in PLATFORM_PACKAGES:
|
||||
if dependency not in optional:
|
||||
raise ValueError(f"Missing Node optional dependency: {dependency}")
|
||||
optional[dependency] = version
|
||||
package_path.write_text(json.dumps(package, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
bun_path = root / "napi/bun.lock"
|
||||
bun_text = bun_path.read_text(encoding="utf-8")
|
||||
for dependency in PLATFORM_PACKAGES:
|
||||
pattern = rf'("{re.escape(dependency)}": ")[^"]+("[,])'
|
||||
bun_text, count = re.subn(
|
||||
pattern, rf"\g<1>{version}\g<2>", bun_text, count=1
|
||||
)
|
||||
if count != 1:
|
||||
raise ValueError(f"Missing Bun lock dependency: {dependency}")
|
||||
bun_path.write_text(bun_text, encoding="utf-8")
|
||||
|
||||
site_path = root / "site/index.html"
|
||||
site_text = site_path.read_text(encoding="utf-8")
|
||||
site_text, count = SITE_WASM_VERSION.subn(
|
||||
rf"\g<1>{version}\g<3>", site_text, count=1
|
||||
)
|
||||
if count != 1:
|
||||
raise ValueError("Missing pinned WASM package URL in site/index.html")
|
||||
site_path.write_text(site_text, encoding="utf-8")
|
||||
|
||||
for _, relative, package_name in LOCK_VERSIONS:
|
||||
_write_lock_version(root / relative, package_name, version)
|
||||
|
||||
check_versions(root)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("version", nargs="?", help="new shared semantic version")
|
||||
parser.add_argument(
|
||||
"--check", action="store_true", help="fail if package versions have diverged"
|
||||
)
|
||||
arguments = parser.parse_args()
|
||||
if arguments.check == bool(arguments.version):
|
||||
parser.error("provide either a version or --check")
|
||||
|
||||
try:
|
||||
if arguments.check:
|
||||
version = check_versions()
|
||||
print(f"All packages use {version}")
|
||||
else:
|
||||
set_versions(arguments.version)
|
||||
print(f"Updated all packages to {arguments.version}")
|
||||
except ValueError as error:
|
||||
parser.exit(1, f"{error}\n")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
+1
-1
@@ -975,7 +975,7 @@ result = pdf_inspector.<span class="fn">process_pdf</span>(<span class="str">"do
|
||||
<script>
|
||||
(() => {
|
||||
const MAX_FILE_SIZE = 25 * 1024 * 1024;
|
||||
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.1/pdf_inspector_wasm.js";
|
||||
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@1.14.2/pdf_inspector_wasm.js";
|
||||
const input = document.querySelector("#pdf-input");
|
||||
const dropZone = document.querySelector("#drop-zone");
|
||||
const filePanel = document.querySelector("#demo-file");
|
||||
|
||||
+98
-24
@@ -1382,7 +1382,13 @@ fn scan_content_for_text_operators(
|
||||
let is_word_end =
|
||||
|pos: usize| -> bool { pos + 1 >= content.len() || content[pos + 1].is_ascii_whitespace() };
|
||||
|
||||
// Simple state machine to find operators
|
||||
// Simple state machine to find operators.
|
||||
// Each Tj/TJ/Tf lookback stops at the previous text/font operator so a
|
||||
// malformed `] TJ` (no `[`) cannot rescan the entire prefix — that was
|
||||
// quadratic in the number of operators.
|
||||
// `Tj`/`TJ` are only counted when the preceding token closes a string or
|
||||
// array (')', '>', ']'), so `Tj` inside `(Hello Tj World)` cannot pin the floor.
|
||||
let mut operand_floor = 0usize;
|
||||
let mut i = 0;
|
||||
while i < content.len() {
|
||||
let b = content[i];
|
||||
@@ -1392,14 +1398,15 @@ fn scan_content_for_text_operators(
|
||||
let next = content[i + 1];
|
||||
if next == b'j' || next == b'J' {
|
||||
// Verify it's an operator (followed by whitespace or newline)
|
||||
if i + 2 >= content.len()
|
||||
if (i + 2 >= content.len()
|
||||
|| content[i + 2].is_ascii_whitespace()
|
||||
|| content[i + 2] == b'\n'
|
||||
|| content[i + 2] == b'\r'
|
||||
|| content[i + 2] == b'\r')
|
||||
&& preceding_operand_closer(content, i, operand_floor)
|
||||
{
|
||||
text_ops += 1;
|
||||
// Scan backward for text string operand to collect unique chars
|
||||
collect_text_chars_before(content, i, unique_chars);
|
||||
collect_text_chars_before(content, i, unique_chars, operand_floor);
|
||||
operand_floor = i;
|
||||
}
|
||||
} else if next == b'f' {
|
||||
// Tf = set font operator
|
||||
@@ -1415,12 +1422,10 @@ fn scan_content_for_text_operators(
|
||||
|| content[i + 2] == b'<'
|
||||
|| content[i + 2] == b'/'
|
||||
{
|
||||
font_changes += 1;
|
||||
// Extract the font name operand preceding the size + Tf.
|
||||
// Pattern: /FontName <size> Tf
|
||||
// Scan backward past the size number and whitespace to find /Name.
|
||||
if let Some(name) = extract_font_name_before_tf(content, i) {
|
||||
if let Some(name) = extract_font_name_before_tf(content, i, operand_floor) {
|
||||
used_font_names.insert(name);
|
||||
font_changes += 1;
|
||||
operand_floor = i;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1466,6 +1471,20 @@ fn scan_content_for_text_operators(
|
||||
(text_ops, image_count, path_ops, font_changes)
|
||||
}
|
||||
|
||||
/// True when the token before `op_pos` (skipping whitespace, not crossing
|
||||
/// `floor`) is a string/array closer. Used so `Tj` inside `(Hello Tj World)`
|
||||
/// is not treated as an operator.
|
||||
fn preceding_operand_closer(content: &[u8], op_pos: usize, floor: usize) -> bool {
|
||||
let mut j = op_pos;
|
||||
while j > floor {
|
||||
j -= 1;
|
||||
if !content[j].is_ascii_whitespace() {
|
||||
return matches!(content[j], b')' | b'>' | b']');
|
||||
}
|
||||
}
|
||||
false
|
||||
}
|
||||
|
||||
/// Extract the font name operand from content stream bytes preceding a Tf operator.
|
||||
///
|
||||
/// The Tf operator syntax is: `/FontName size Tf`
|
||||
@@ -1473,25 +1492,27 @@ fn scan_content_for_text_operators(
|
||||
/// whitespace to find the `/Name` token.
|
||||
///
|
||||
/// Returns the font name bytes (without the leading `/`), e.g. `b"F1"` for `/F1`.
|
||||
fn extract_font_name_before_tf(content: &[u8], tf_pos: usize) -> Option<Vec<u8>> {
|
||||
/// `floor` is the start of the previous text/font operator (or 0); lookback
|
||||
/// must not cross it.
|
||||
fn extract_font_name_before_tf(content: &[u8], tf_pos: usize, floor: usize) -> Option<Vec<u8>> {
|
||||
// Scan backward past whitespace before "Tf"
|
||||
let mut j = tf_pos;
|
||||
while j > 0 && content[j - 1].is_ascii_whitespace() {
|
||||
while j > floor && content[j - 1].is_ascii_whitespace() {
|
||||
j -= 1;
|
||||
}
|
||||
// Scan backward past the size number (digits, '.', '-')
|
||||
while j > 0
|
||||
while j > floor
|
||||
&& (content[j - 1].is_ascii_digit() || content[j - 1] == b'.' || content[j - 1] == b'-')
|
||||
{
|
||||
j -= 1;
|
||||
}
|
||||
// Scan backward past whitespace between font name and size
|
||||
while j > 0 && content[j - 1].is_ascii_whitespace() {
|
||||
while j > floor && content[j - 1].is_ascii_whitespace() {
|
||||
j -= 1;
|
||||
}
|
||||
// Now j should point just after the font name. Scan backward to find '/'.
|
||||
let name_end = j;
|
||||
while j > 0 && content[j - 1] != b'/' {
|
||||
while j > floor && content[j - 1] != b'/' {
|
||||
// Font names consist of regular characters (not whitespace, not delimiters)
|
||||
if content[j - 1].is_ascii_whitespace() || content[j - 1] == b'(' || content[j - 1] == b')'
|
||||
{
|
||||
@@ -1499,7 +1520,7 @@ fn extract_font_name_before_tf(content: &[u8], tf_pos: usize) -> Option<Vec<u8>>
|
||||
}
|
||||
j -= 1;
|
||||
}
|
||||
if j == 0 || content[j - 1] != b'/' {
|
||||
if j <= floor || content[j - 1] != b'/' {
|
||||
return None;
|
||||
}
|
||||
// j-1 is the '/', font name is content[j..name_end]
|
||||
@@ -1514,16 +1535,24 @@ fn extract_font_name_before_tf(content: &[u8], tf_pos: usize) -> Option<Vec<u8>>
|
||||
/// and collect unique non-whitespace bytes from it.
|
||||
///
|
||||
/// Handles both literal strings `(...)` and hex strings `<...>`.
|
||||
fn collect_text_chars_before(content: &[u8], op_pos: usize, unique_chars: &mut HashSet<u8>) {
|
||||
/// `floor` is the start of the previous text/font operator (or 0); lookback
|
||||
/// must not cross it, or a missing `[` before `TJ` rescans the whole prefix.
|
||||
fn collect_text_chars_before(
|
||||
content: &[u8],
|
||||
op_pos: usize,
|
||||
unique_chars: &mut HashSet<u8>,
|
||||
floor: usize,
|
||||
) {
|
||||
// Walk backward past whitespace to find the closing delimiter
|
||||
let mut j = op_pos;
|
||||
while j > 0 {
|
||||
while j > floor {
|
||||
j -= 1;
|
||||
if !content[j].is_ascii_whitespace() {
|
||||
break;
|
||||
}
|
||||
}
|
||||
if j == 0 {
|
||||
// All whitespace, or we landed on the previous operator token.
|
||||
if j == floor {
|
||||
return;
|
||||
}
|
||||
|
||||
@@ -1533,7 +1562,7 @@ fn collect_text_chars_before(content: &[u8], op_pos: usize, unique_chars: &mut H
|
||||
// Literal string: scan backward for matching '('
|
||||
let mut depth = 1i32;
|
||||
let mut k = j;
|
||||
while k > 0 && depth > 0 {
|
||||
while k > floor && depth > 0 {
|
||||
k -= 1;
|
||||
match content[k] {
|
||||
b')' if k == 0 || content[k - 1] != b'\\' => depth += 1,
|
||||
@@ -1552,7 +1581,7 @@ fn collect_text_chars_before(content: &[u8], op_pos: usize, unique_chars: &mut H
|
||||
} else if closing == b'>' {
|
||||
// Hex string: scan backward for '<'
|
||||
let mut k = j;
|
||||
while k > 0 {
|
||||
while k > floor {
|
||||
k -= 1;
|
||||
if content[k] == b'<' {
|
||||
break;
|
||||
@@ -1582,7 +1611,7 @@ fn collect_text_chars_before(content: &[u8], op_pos: usize, unique_chars: &mut H
|
||||
} else if closing == b']' {
|
||||
// TJ array: scan backward for '[' and collect from all strings inside
|
||||
let mut k = j;
|
||||
while k > 0 {
|
||||
while k > floor {
|
||||
k -= 1;
|
||||
if content[k] == b'[' {
|
||||
break;
|
||||
@@ -2014,6 +2043,51 @@ mod tests {
|
||||
assert_eq!(imgs3, 0);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_scan_content_successive_tj_collects_each_operand() {
|
||||
// Lookback is floored at the previous Tj/TJ/Tf so later operators must
|
||||
// still see their own operands.
|
||||
let content = b"[(Hello)] TJ [(World)] TJ (More) Tj";
|
||||
let mut uchars = HashSet::new();
|
||||
let (ops, _, _, _) =
|
||||
scan_content_for_text_operators(content, &mut uchars, &mut HashSet::new());
|
||||
assert_eq!(ops, 3);
|
||||
for &ch in b"HeloWrdM" {
|
||||
assert!(uchars.contains(&ch), "missing char {}", ch as char);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_scan_content_tj_inside_literal_is_not_an_operator() {
|
||||
// `Tj` followed by space inside a literal must not count as an operator
|
||||
// or pin the lookback floor; the real `Tj` still collects the string.
|
||||
let content = b"BT (Hello Tj World) Tj ET";
|
||||
let mut uchars = HashSet::new();
|
||||
let (ops, _, _, _) =
|
||||
scan_content_for_text_operators(content, &mut uchars, &mut HashSet::new());
|
||||
assert_eq!(ops, 1);
|
||||
for &ch in b"HeloTjWrd" {
|
||||
assert!(uchars.contains(&ch), "missing char {}", ch as char);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_scan_content_malformed_tj_lookback_stays_linear() {
|
||||
// `] TJ` with no `[` used to walk the entire prefix for every operator
|
||||
// (quadratic). 30k repeats is enough that a prefix rescan would dominate
|
||||
// the test runtime; with the floor it is a single linear pass.
|
||||
let n = 30_000usize;
|
||||
let mut content = Vec::with_capacity(n * 5);
|
||||
for _ in 0..n {
|
||||
content.extend_from_slice(b"] TJ\n");
|
||||
}
|
||||
let mut uchars = HashSet::new();
|
||||
let (ops, _, _, _) =
|
||||
scan_content_for_text_operators(&content, &mut uchars, &mut HashSet::new());
|
||||
assert_eq!(ops, n as u32);
|
||||
assert!(uchars.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_image_dominated_detection() {
|
||||
// Do operators are no longer counted as images by scan_content_for_text_operators.
|
||||
@@ -2772,14 +2846,14 @@ mod tests {
|
||||
fn test_extract_font_name_basic() {
|
||||
// Standard pattern: /F1 12 Tf
|
||||
let content = b"/F1 12 Tf";
|
||||
let name = extract_font_name_before_tf(content, 6); // 'T' is at index 6
|
||||
let name = extract_font_name_before_tf(content, 6, 0); // 'T' is at index 6
|
||||
assert_eq!(name, Some(b"F1".to_vec()));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_font_name_long_name() {
|
||||
let content = b"/ArialMT-Bold 9.5 Tf";
|
||||
let name = extract_font_name_before_tf(content, 18);
|
||||
let name = extract_font_name_before_tf(content, 18, 0);
|
||||
assert_eq!(name, Some(b"ArialMT-Bold".to_vec()));
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,328 @@
|
||||
//! Bounded content-stream decoding.
|
||||
//!
|
||||
//! `lopdf::content::Content::decode` materializes every operator before any
|
||||
//! caller can apply a limit. A compact page of `q Q` pairs can therefore
|
||||
//! allocate hundreds of megabytes and abort. Count operators first (without
|
||||
//! allocating `Operation` objects) and skip decode when the cap is exceeded.
|
||||
|
||||
use crate::PdfError;
|
||||
use lopdf::content::Content;
|
||||
|
||||
/// Maximum content-stream operators decoded for a page or a single Form
|
||||
/// XObject. Matches the previous post-decode skip threshold.
|
||||
pub(crate) const MAX_PAGE_OPERATIONS: usize = 1_000_000;
|
||||
|
||||
/// Decode `data` unless it contains more than `max_operations` operators.
|
||||
///
|
||||
/// Returns `Ok(None)` when the stream exceeds the cap, so callers can skip
|
||||
/// extraction without first allocating the operation vector.
|
||||
pub(crate) fn decode_content_bounded(
|
||||
data: &[u8],
|
||||
max_operations: usize,
|
||||
) -> Result<Option<Content>, PdfError> {
|
||||
if content_exceeds_operation_limit(data, max_operations) {
|
||||
return Ok(None);
|
||||
}
|
||||
Content::decode(data)
|
||||
.map(Some)
|
||||
.map_err(|e| PdfError::Parse(e.to_string()))
|
||||
}
|
||||
|
||||
fn content_exceeds_operation_limit(data: &[u8], max_operations: usize) -> bool {
|
||||
count_content_operators(data, max_operations.saturating_add(1)) > max_operations
|
||||
}
|
||||
|
||||
/// Count operators using the same token rules as lopdf's content parser,
|
||||
/// stopping at `limit`. Does not allocate `Operation` / `Object` values.
|
||||
fn count_content_operators(data: &[u8], limit: usize) -> usize {
|
||||
let mut i = 0;
|
||||
let mut count = 0;
|
||||
while i < data.len() && count < limit {
|
||||
skip_content_space(data, &mut i);
|
||||
if i >= data.len() {
|
||||
break;
|
||||
}
|
||||
if data[i] == b'%' {
|
||||
skip_comment(data, &mut i);
|
||||
continue;
|
||||
}
|
||||
match data[i] {
|
||||
b'(' => i = skip_literal_string(data, i),
|
||||
b'<' => {
|
||||
if data.get(i + 1) == Some(&b'<') {
|
||||
i += 2;
|
||||
} else {
|
||||
i = skip_hex_string(data, i);
|
||||
}
|
||||
}
|
||||
b'>' => {
|
||||
i += 1;
|
||||
if data.get(i) == Some(&b'>') {
|
||||
i += 1;
|
||||
}
|
||||
}
|
||||
b'[' | b']' => i += 1,
|
||||
b'/' => skip_name(data, &mut i),
|
||||
b'+' | b'-' | b'.' => skip_number(data, &mut i),
|
||||
b if b.is_ascii_digit() => skip_number(data, &mut i),
|
||||
b if is_operator_byte(b) => {
|
||||
let start = i;
|
||||
i += 1;
|
||||
while i < data.len() && is_operator_byte(data[i]) {
|
||||
i += 1;
|
||||
}
|
||||
let token = &data[start..i];
|
||||
if token == b"true" || token == b"false" || token == b"null" {
|
||||
continue;
|
||||
}
|
||||
count += 1;
|
||||
if token == b"BI" && (i >= data.len() || is_content_space(data[i])) {
|
||||
i = skip_inline_image_after_bi(data, i);
|
||||
}
|
||||
}
|
||||
_ => i += 1,
|
||||
}
|
||||
}
|
||||
count
|
||||
}
|
||||
|
||||
fn is_content_space(b: u8) -> bool {
|
||||
// PDF whitespace (ISO 32000): NUL, tab, LF, FF, CR, space. Names must
|
||||
// stop on these so a following operator is not absorbed into `/Name`.
|
||||
matches!(b, b'\0' | b'\t' | b'\n' | b'\x0c' | b'\r' | b' ')
|
||||
}
|
||||
|
||||
fn is_operator_byte(b: u8) -> bool {
|
||||
b.is_ascii_alphabetic() || matches!(b, b'*' | b'\'' | b'"')
|
||||
}
|
||||
|
||||
fn is_delimiter(b: u8) -> bool {
|
||||
matches!(
|
||||
b,
|
||||
b'(' | b')' | b'<' | b'>' | b'[' | b']' | b'{' | b'}' | b'/' | b'%'
|
||||
)
|
||||
}
|
||||
|
||||
fn skip_content_space(data: &[u8], i: &mut usize) {
|
||||
while *i < data.len() && is_content_space(data[*i]) {
|
||||
*i += 1;
|
||||
}
|
||||
}
|
||||
|
||||
fn skip_comment(data: &[u8], i: &mut usize) {
|
||||
while *i < data.len() && data[*i] != b'\n' && data[*i] != b'\r' {
|
||||
*i += 1;
|
||||
}
|
||||
}
|
||||
|
||||
fn skip_literal_string(data: &[u8], mut i: usize) -> usize {
|
||||
let mut depth = 1i32;
|
||||
i += 1;
|
||||
while i < data.len() && depth > 0 {
|
||||
match data[i] {
|
||||
b'\\' => {
|
||||
i += 1;
|
||||
if i < data.len() {
|
||||
i += 1;
|
||||
}
|
||||
}
|
||||
b'(' => {
|
||||
depth += 1;
|
||||
i += 1;
|
||||
}
|
||||
b')' => {
|
||||
depth -= 1;
|
||||
i += 1;
|
||||
}
|
||||
_ => i += 1,
|
||||
}
|
||||
}
|
||||
i
|
||||
}
|
||||
|
||||
fn skip_hex_string(data: &[u8], mut i: usize) -> usize {
|
||||
i += 1;
|
||||
while i < data.len() && data[i] != b'>' {
|
||||
i += 1;
|
||||
}
|
||||
if i < data.len() {
|
||||
i += 1;
|
||||
}
|
||||
i
|
||||
}
|
||||
|
||||
fn skip_name(data: &[u8], i: &mut usize) {
|
||||
*i += 1;
|
||||
while *i < data.len() && !is_content_space(data[*i]) && !is_delimiter(data[*i]) {
|
||||
*i += 1;
|
||||
}
|
||||
}
|
||||
|
||||
fn skip_number(data: &[u8], i: &mut usize) {
|
||||
if *i < data.len() && matches!(data[*i], b'+' | b'-') {
|
||||
*i += 1;
|
||||
}
|
||||
while *i < data.len() && data[*i].is_ascii_digit() {
|
||||
*i += 1;
|
||||
}
|
||||
if *i < data.len() && data[*i] == b'.' {
|
||||
*i += 1;
|
||||
while *i < data.len() && data[*i].is_ascii_digit() {
|
||||
*i += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// After a `BI` operator, skip inline-image data through `EI`.
|
||||
/// Uses the same PDF whitespace set as `is_content_space`. If `EI` is not
|
||||
/// found, leave the cursor in place so later operators are still counted
|
||||
/// (undercounting would let decode allocate the full vector).
|
||||
fn skip_inline_image_after_bi(data: &[u8], mut i: usize) -> usize {
|
||||
skip_content_space(data, &mut i);
|
||||
let rest = &data[i..];
|
||||
if let Some(pos) = rest.windows(4).position(|w| {
|
||||
is_content_space(w[0]) && w[1] == b'E' && w[2] == b'I' && is_content_space(w[3])
|
||||
}) {
|
||||
return i + pos + 3;
|
||||
}
|
||||
i
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn lopdf_op_count(data: &[u8]) -> usize {
|
||||
Content::decode(data)
|
||||
.map(|c| c.operations.len())
|
||||
.unwrap_or(0)
|
||||
}
|
||||
|
||||
/// DoS safety: never report fewer operators than lopdf would allocate.
|
||||
/// Overcount is acceptable (skip a page); undercount would re-open decode.
|
||||
fn assert_count_does_not_undercount(data: &[u8]) {
|
||||
let ours = count_content_operators(data, usize::MAX);
|
||||
match Content::decode(data) {
|
||||
Ok(content) => assert!(
|
||||
ours >= content.operations.len(),
|
||||
"undercount: ours={ours} lopdf={} for {:?}",
|
||||
content.operations.len(),
|
||||
String::from_utf8_lossy(data)
|
||||
),
|
||||
Err(_) => {}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn operator_count_matches_lopdf_for_typical_streams() {
|
||||
let samples: &[&[u8]] = &[
|
||||
b"q 1 0 0 1 0 0 cm BT /F1 12 Tf 72 720 Td (Hello) Tj ET Q",
|
||||
b"q Q q Q",
|
||||
b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (first) Tj (struck) ' ET",
|
||||
b"1 0 0 rg 0 0 10 10 re f",
|
||||
b"true false null q",
|
||||
b"% comment\nq Q\n",
|
||||
b"[ (a) 1 (b) ] TJ",
|
||||
b"1 0 0 1 0 0 cm /Im0 Do",
|
||||
];
|
||||
for data in samples {
|
||||
assert_eq!(
|
||||
count_content_operators(data, usize::MAX),
|
||||
lopdf_op_count(data),
|
||||
"count mismatch for {}",
|
||||
String::from_utf8_lossy(data)
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn strings_and_comments_are_not_operators() {
|
||||
let data = b"(q Q Tj) Tj % q Q\nET";
|
||||
assert_eq!(
|
||||
count_content_operators(data, usize::MAX),
|
||||
lopdf_op_count(data)
|
||||
);
|
||||
assert_eq!(count_content_operators(data, usize::MAX), 2); // Tj, ET
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn inline_image_counts_as_one_operator() {
|
||||
let data = b"BI /W 2 /H 2 /CS /RGB /BPC 8 ID \x00\x01\x02\x03 EI q";
|
||||
assert_eq!(
|
||||
count_content_operators(data, usize::MAX),
|
||||
lopdf_op_count(data)
|
||||
);
|
||||
assert_eq!(count_content_operators(data, usize::MAX), 2); // BI, q
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn inline_image_ei_accepts_pdf_whitespace() {
|
||||
let tab = b"BI /W 1 /H 1 ID \xff\tEI\t q Q";
|
||||
let nul = b"BI /W 1 /H 1 ID \xff\x00EI\x00 q Q";
|
||||
let ff = b"BI /W 1 /H 1 ID \xff\x0cEI\x0c q Q";
|
||||
for data in [tab.as_slice(), nul.as_slice(), ff.as_slice()] {
|
||||
assert_count_does_not_undercount(data);
|
||||
assert!(
|
||||
count_content_operators(data, usize::MAX) >= 3,
|
||||
"BI plus following q Q must remain visible after EI, got {} for {:?}",
|
||||
count_content_operators(data, usize::MAX),
|
||||
String::from_utf8_lossy(data)
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn decode_is_skipped_when_operator_cap_is_exceeded() {
|
||||
let mut data = Vec::new();
|
||||
for _ in 0..20 {
|
||||
data.extend_from_slice(b"q Q\n");
|
||||
}
|
||||
assert!(decode_content_bounded(&data, 10).unwrap().is_none());
|
||||
let decoded = decode_content_bounded(&data, 50).unwrap().unwrap();
|
||||
assert_eq!(decoded.operations.len(), 40);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn name_whitespace_does_not_swallow_following_operator() {
|
||||
// NUL / form-feed end a name (PDF whitespace). Absorbing `q` into
|
||||
// `/x` would undercount and let decode allocate the operator vector.
|
||||
let mut nul_sep = Vec::new();
|
||||
let mut ff_sep = Vec::new();
|
||||
for _ in 0..8_000 {
|
||||
nul_sep.extend_from_slice(b"/x\x00q");
|
||||
ff_sep.extend_from_slice(b"/x\x0cq");
|
||||
}
|
||||
assert_count_does_not_undercount(&nul_sep);
|
||||
assert_count_does_not_undercount(&ff_sep);
|
||||
assert!(count_content_operators(&ff_sep, usize::MAX) >= 8_000);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn edge_streams_do_not_undercount_vs_lopdf() {
|
||||
let samples: &[&[u8]] = &[
|
||||
b".5 0 0 .5 0 0 cm",
|
||||
b"+1 -2 3.0 rg",
|
||||
b"<0041> Tj",
|
||||
b"(unbalanced",
|
||||
b"BI /W 1 /H 1 ID \xff\xff no EI here q Q q Q",
|
||||
b"q\x00Q\x00q\x00Q",
|
||||
b"/F1\x0c12 Tf (Hi) Tj",
|
||||
b"{ 1 2 add } cvx",
|
||||
];
|
||||
for data in samples {
|
||||
assert_count_does_not_undercount(data);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn million_q_pairs_are_rejected_without_decode() {
|
||||
let mut data = Vec::with_capacity((MAX_PAGE_OPERATIONS + 1) * 2);
|
||||
for _ in 0..=MAX_PAGE_OPERATIONS {
|
||||
data.extend_from_slice(b"q\n");
|
||||
}
|
||||
assert!(content_exceeds_operation_limit(&data, MAX_PAGE_OPERATIONS));
|
||||
assert!(decode_content_bounded(&data, MAX_PAGE_OPERATIONS)
|
||||
.unwrap()
|
||||
.is_none());
|
||||
}
|
||||
}
|
||||
@@ -19,7 +19,7 @@ use super::fonts::{
|
||||
CMapDecisionCache, FontStyleCache,
|
||||
};
|
||||
use super::underline::UnderlineLine;
|
||||
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
|
||||
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, FormWalkBudget, XObjectType};
|
||||
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
|
||||
|
||||
/// Strip PDF comments (% to end of line) from content stream bytes.
|
||||
@@ -137,8 +137,11 @@ fn rise_adjusted(tm: &[f32; 6], rise: f32) -> [f32; 6] {
|
||||
]
|
||||
}
|
||||
|
||||
/// Returns `(page_extraction, has_gid_fonts)` where `has_gid_fonts` indicates
|
||||
/// the page uses fonts with unresolvable gid-encoded glyphs.
|
||||
/// Returns `(page_extraction, has_gid_fonts, coords_rotated, skipped_invisible)`
|
||||
/// where `has_gid_fonts` indicates the page uses fonts with unresolvable
|
||||
/// gid-encoded glyphs and `skipped_invisible` reports that invisible (Tr 3)
|
||||
/// text was present but suppressed — callers can use it to decide whether an
|
||||
/// `include_invisible` retry could recover anything at all.
|
||||
pub(crate) fn extract_page_text_items(
|
||||
doc: &Document,
|
||||
page_id: ObjectId,
|
||||
@@ -146,9 +149,8 @@ pub(crate) fn extract_page_text_items(
|
||||
font_cmaps: &FontCMaps,
|
||||
include_invisible: bool,
|
||||
style_cache: &mut FontStyleCache,
|
||||
) -> Result<(PageExtraction, bool, bool), PdfError> {
|
||||
use lopdf::content::Content;
|
||||
|
||||
form_budget: &mut FormWalkBudget,
|
||||
) -> Result<(PageExtraction, bool, bool, bool), PdfError> {
|
||||
let mut items = Vec::new();
|
||||
let mut rects: Vec<PdfRect> = Vec::new();
|
||||
let mut clip_rects: Vec<PdfRect> = Vec::new();
|
||||
@@ -252,22 +254,27 @@ pub(crate) fn extract_page_text_items(
|
||||
// Content::decode parser, causing it to skip operators like ET and Q.
|
||||
let content_data = strip_pdf_comments(&content_data);
|
||||
|
||||
let content = Content::decode(&content_data).map_err(|e| PdfError::Parse(e.to_string()))?;
|
||||
|
||||
const MAX_OPERATIONS: usize = 1_000_000;
|
||||
if content.operations.len() > MAX_OPERATIONS {
|
||||
log::warn!(
|
||||
"page {}: skipping extraction — {} operations exceeds limit ({})",
|
||||
page_num,
|
||||
content.operations.len(),
|
||||
MAX_OPERATIONS
|
||||
);
|
||||
return Ok(((Vec::new(), Vec::new(), Vec::new()), false, false));
|
||||
}
|
||||
let content = match super::content_decode::decode_content_bounded(
|
||||
&content_data,
|
||||
super::content_decode::MAX_PAGE_OPERATIONS,
|
||||
)? {
|
||||
Some(content) => content,
|
||||
None => {
|
||||
log::warn!(
|
||||
"page {}: skipping extraction — content stream exceeds {} operations",
|
||||
page_num,
|
||||
super::content_decode::MAX_PAGE_OPERATIONS
|
||||
);
|
||||
return Ok(((Vec::new(), Vec::new(), Vec::new()), false, false, false));
|
||||
}
|
||||
};
|
||||
|
||||
// Graphics state tracking
|
||||
let mut ctm = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0]; // Current Transformation Matrix
|
||||
let mut text_rendering_mode: i32 = 0; // 0=fill, 1=stroke, 2=fill+stroke, 3=invisible
|
||||
// Invisible (Tr 3) text was present but suppressed — reported to callers
|
||||
// so an include_invisible retry is attempted only when it can recover.
|
||||
let mut skipped_invisible = false;
|
||||
let mut line_width: f32 = 1.0;
|
||||
#[derive(Clone)]
|
||||
struct SavedGraphicsState {
|
||||
@@ -494,6 +501,14 @@ pub(crate) fn extract_page_text_items(
|
||||
// For Mixed/template PDFs, include_invisible=true extracts
|
||||
// the OCR text layer that sits behind scanned images.
|
||||
if text_rendering_mode == 3 && !include_invisible {
|
||||
if op
|
||||
.operands
|
||||
.first()
|
||||
.and_then(get_operand_bytes)
|
||||
.is_some_and(|raw| !raw.is_empty())
|
||||
{
|
||||
skipped_invisible = true;
|
||||
}
|
||||
if let Some(w_ts) = w_ts_opt {
|
||||
text_matrix[4] += w_ts * text_matrix[0];
|
||||
text_matrix[5] += w_ts * text_matrix[1];
|
||||
@@ -565,6 +580,16 @@ pub(crate) fn extract_page_text_items(
|
||||
if in_text_block && !op.operands.is_empty() {
|
||||
if let Ok(array) = op.operands[0].as_array() {
|
||||
let font_info = font_widths.get(¤t_font);
|
||||
// Numeric-only TJ arrays (pure kerning) show no
|
||||
// text — they must not trigger the invisible retry.
|
||||
if text_rendering_mode == 3
|
||||
&& !include_invisible
|
||||
&& array
|
||||
.iter()
|
||||
.any(|el| get_operand_bytes(el).is_some_and(|raw| !raw.is_empty()))
|
||||
{
|
||||
skipped_invisible = true;
|
||||
}
|
||||
let is_invisible = (text_rendering_mode == 3 && !include_invisible)
|
||||
|| suppress_glyph_extraction;
|
||||
// Capture first-glyph position for ActualText
|
||||
@@ -770,6 +795,16 @@ pub(crate) fn extract_page_text_items(
|
||||
)
|
||||
})
|
||||
});
|
||||
if text_rendering_mode == 3
|
||||
&& !include_invisible
|
||||
&& op
|
||||
.operands
|
||||
.first()
|
||||
.and_then(get_operand_bytes)
|
||||
.is_some_and(|raw| !raw.is_empty())
|
||||
{
|
||||
skipped_invisible = true;
|
||||
}
|
||||
if !((text_rendering_mode == 3 && !include_invisible)
|
||||
|| suppress_glyph_extraction
|
||||
|| op.operands.is_empty())
|
||||
@@ -882,6 +917,7 @@ pub(crate) fn extract_page_text_items(
|
||||
&ctm,
|
||||
&mut cmap_decisions,
|
||||
style_cache,
|
||||
form_budget,
|
||||
);
|
||||
items.extend(form_items);
|
||||
}
|
||||
@@ -1251,6 +1287,12 @@ pub(crate) fn extract_page_text_items(
|
||||
}
|
||||
}
|
||||
|
||||
if form_budget.was_truncated() {
|
||||
log::warn!(
|
||||
"page {page_num}: Form XObject expansion truncated (invocation or operation budget reached); nested form text may be incomplete"
|
||||
);
|
||||
}
|
||||
|
||||
// Underline detection reads only painted ink: `re` rects confirmed by
|
||||
// a paint operator plus filled-subpath rects — never clip-only rects,
|
||||
// which draw nothing.
|
||||
@@ -1300,7 +1342,12 @@ pub(crate) fn extract_page_text_items(
|
||||
|
||||
let items = super::merge_text_items(items);
|
||||
let items = super::merge_subscript_items(items);
|
||||
Ok(((items, rects, lines), has_gid_fonts, coords_rotated))
|
||||
Ok((
|
||||
(items, rects, lines),
|
||||
has_gid_fonts,
|
||||
coords_rotated,
|
||||
skipped_invisible,
|
||||
))
|
||||
}
|
||||
|
||||
/// Counts of text operators with horizontal vs rotated combined matrices.
|
||||
@@ -1498,13 +1545,14 @@ mod tests {
|
||||
|
||||
let (doc, page_id) = simple_doc_with_content(content);
|
||||
let font_cmaps = FontCMaps::from_doc(&doc);
|
||||
let ((items, _, _), _, _) = extract_page_text_items(
|
||||
let ((items, _, _), _, _, _) = extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
1,
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut FontStyleCache::new(),
|
||||
&mut FormWalkBudget::new(),
|
||||
)
|
||||
.unwrap();
|
||||
items
|
||||
@@ -1734,9 +1782,10 @@ BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut FontStyleCache::new(),
|
||||
&mut FormWalkBudget::new(),
|
||||
)
|
||||
.unwrap();
|
||||
let ((items, rects, lines), _has_gid, _coords_rotated) = result;
|
||||
let ((items, rects, lines), _has_gid, _coords_rotated, _skipped_invisible) = result;
|
||||
assert!(items.is_empty());
|
||||
assert!(rects.is_empty());
|
||||
assert!(lines.is_empty());
|
||||
@@ -1817,13 +1866,14 @@ BT 30 700 Tm <41> Tj ET";
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
|
||||
let font_cmaps = FontCMaps::from_doc(&doc);
|
||||
let ((items, _, _), _, _) = extract_page_text_items(
|
||||
let ((items, _, _), _, _, _) = extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
1,
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut FontStyleCache::new(),
|
||||
&mut FormWalkBudget::new(),
|
||||
)
|
||||
.unwrap();
|
||||
let text = items
|
||||
@@ -1887,4 +1937,19 @@ BT 30 700 Tm <41> Tj ET";
|
||||
let output = strip_pdf_comments(input);
|
||||
assert_eq!(output, b"(x\\\\) Tj \nET\n");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn oversized_content_stream_skips_extraction() {
|
||||
let mut content =
|
||||
Vec::with_capacity((super::super::content_decode::MAX_PAGE_OPERATIONS + 1) * 2);
|
||||
for _ in 0..=super::super::content_decode::MAX_PAGE_OPERATIONS {
|
||||
content.extend_from_slice(b"q\n");
|
||||
}
|
||||
content.extend_from_slice(b"BT /F1 12 Tf 72 720 Td (Hello) Tj ET\n");
|
||||
let items = extract_simple_items(&content);
|
||||
assert!(
|
||||
items.is_empty(),
|
||||
"pages over the operator cap must not be decoded"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
+107
-16
@@ -482,7 +482,11 @@ pub(crate) fn parse_cid_w_array(
|
||||
widths: &mut HashMap<u16, u16>,
|
||||
) {
|
||||
let mut i = 0;
|
||||
let mut assigned = 0usize;
|
||||
while i < w_array.len() {
|
||||
if assigned >= crate::tounicode::MAX_CID_W_EXPANSION {
|
||||
return;
|
||||
}
|
||||
let start_cid = match &w_array[i] {
|
||||
Object::Integer(n) => *n as u16,
|
||||
Object::Real(n) => *n as u16,
|
||||
@@ -501,12 +505,14 @@ pub(crate) fn parse_cid_w_array(
|
||||
Object::Array(arr) => {
|
||||
// [c [w1 w2 ...]] — consecutive widths starting at c
|
||||
for (j, w_obj) in arr.iter().enumerate() {
|
||||
let w = match w_obj {
|
||||
Object::Integer(n) => *n as u16,
|
||||
Object::Real(n) => *n as u16,
|
||||
_ => continue,
|
||||
};
|
||||
widths.insert(start_cid + j as u16, w);
|
||||
if !assign_cid_width(
|
||||
widths,
|
||||
start_cid.wrapping_add(j as u16),
|
||||
w_obj,
|
||||
&mut assigned,
|
||||
) {
|
||||
return;
|
||||
}
|
||||
}
|
||||
i += 1;
|
||||
}
|
||||
@@ -514,12 +520,14 @@ pub(crate) fn parse_cid_w_array(
|
||||
// Could be a reference to an array
|
||||
if let Ok(Object::Array(arr)) = doc.get_object(*r) {
|
||||
for (j, w_obj) in arr.iter().enumerate() {
|
||||
let w = match w_obj {
|
||||
Object::Integer(n) => *n as u16,
|
||||
Object::Real(n) => *n as u16,
|
||||
_ => continue,
|
||||
};
|
||||
widths.insert(start_cid + j as u16, w);
|
||||
if !assign_cid_width(
|
||||
widths,
|
||||
start_cid.wrapping_add(j as u16),
|
||||
w_obj,
|
||||
&mut assigned,
|
||||
) {
|
||||
return;
|
||||
}
|
||||
}
|
||||
i += 1;
|
||||
} else {
|
||||
@@ -542,8 +550,8 @@ pub(crate) fn parse_cid_w_array(
|
||||
continue;
|
||||
}
|
||||
};
|
||||
for cid in start_cid..=end {
|
||||
widths.insert(cid, w);
|
||||
if !assign_cid_width_range(widths, start_cid, end, w, &mut assigned) {
|
||||
return;
|
||||
}
|
||||
i += 1;
|
||||
}
|
||||
@@ -561,8 +569,8 @@ pub(crate) fn parse_cid_w_array(
|
||||
continue;
|
||||
}
|
||||
};
|
||||
for cid in start_cid..=end {
|
||||
widths.insert(cid, w);
|
||||
if !assign_cid_width_range(widths, start_cid, end, w, &mut assigned) {
|
||||
return;
|
||||
}
|
||||
i += 1;
|
||||
}
|
||||
@@ -573,6 +581,45 @@ pub(crate) fn parse_cid_w_array(
|
||||
}
|
||||
}
|
||||
|
||||
fn assign_cid_width(
|
||||
widths: &mut HashMap<u16, u16>,
|
||||
cid: u16,
|
||||
w_obj: &Object,
|
||||
assigned: &mut usize,
|
||||
) -> bool {
|
||||
let w = match w_obj {
|
||||
Object::Integer(n) => *n as u16,
|
||||
Object::Real(n) => *n as u16,
|
||||
_ => return true,
|
||||
};
|
||||
if *assigned >= crate::tounicode::MAX_CID_W_EXPANSION {
|
||||
return false;
|
||||
}
|
||||
widths.insert(cid, w);
|
||||
*assigned += 1;
|
||||
true
|
||||
}
|
||||
|
||||
fn assign_cid_width_range(
|
||||
widths: &mut HashMap<u16, u16>,
|
||||
start: u16,
|
||||
end: u16,
|
||||
w: u16,
|
||||
assigned: &mut usize,
|
||||
) -> bool {
|
||||
if start > end {
|
||||
return true;
|
||||
}
|
||||
for cid in start..=end {
|
||||
if *assigned >= crate::tounicode::MAX_CID_W_EXPANSION {
|
||||
return false;
|
||||
}
|
||||
widths.insert(cid, w);
|
||||
*assigned += 1;
|
||||
}
|
||||
true
|
||||
}
|
||||
|
||||
/// Compute the width of a string in text space units,
|
||||
/// given raw bytes and font width info.
|
||||
/// Returns width in text space units (font_units * units_scale * font_size).
|
||||
@@ -2313,4 +2360,48 @@ end",
|
||||
// invalid CMap result — so it must not clear the gid flag.
|
||||
assert!(gid_flagged(Some("<01> <FFFD>\n<02> <FFFD>")));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_cid_w_array_range_and_consecutive() {
|
||||
use super::parse_cid_w_array;
|
||||
use lopdf::{Document, Object};
|
||||
use std::collections::HashMap;
|
||||
|
||||
let doc = Document::new();
|
||||
let mut widths = HashMap::new();
|
||||
let w = vec![
|
||||
Object::Integer(10),
|
||||
Object::Integer(12),
|
||||
Object::Integer(500),
|
||||
Object::Integer(20),
|
||||
Object::Array(vec![Object::Integer(100), Object::Integer(200)]),
|
||||
];
|
||||
parse_cid_w_array(&doc, &w, &mut widths);
|
||||
assert_eq!(widths.get(&10), Some(&500));
|
||||
assert_eq!(widths.get(&11), Some(&500));
|
||||
assert_eq!(widths.get(&12), Some(&500));
|
||||
assert_eq!(widths.get(&20), Some(&100));
|
||||
assert_eq!(widths.get(&21), Some(&200));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_cid_w_array_repeated_full_ranges_stay_bounded() {
|
||||
use super::parse_cid_w_array;
|
||||
use crate::tounicode::MAX_CID_W_EXPANSION;
|
||||
use lopdf::{Document, Object};
|
||||
use std::collections::HashMap;
|
||||
|
||||
let doc = Document::new();
|
||||
let mut widths = HashMap::new();
|
||||
let mut w = Vec::new();
|
||||
for _ in 0..5_000 {
|
||||
w.push(Object::Integer(0));
|
||||
w.push(Object::Integer(65535));
|
||||
w.push(Object::Integer(500));
|
||||
}
|
||||
parse_cid_w_array(&doc, &w, &mut widths);
|
||||
assert!(widths.len() <= MAX_CID_W_EXPANSION);
|
||||
assert_eq!(widths.get(&0), Some(&500));
|
||||
assert_eq!(widths.get(&65535), Some(&500));
|
||||
}
|
||||
}
|
||||
|
||||
+327
-17
@@ -41,15 +41,110 @@ pub(crate) fn detect_columns(
|
||||
}
|
||||
debug!("page {}: detect_columns: {} items", page, page_items.len());
|
||||
|
||||
// Find page bounds
|
||||
let x_min = page_items.iter().map(|i| i.x).fold(f32::INFINITY, f32::min);
|
||||
let x_max = page_items
|
||||
.iter()
|
||||
.map(|i| i.x + effective_width(i))
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
// The width of one ordinary page, used three ways below: as the largest
|
||||
// credible width for a single text run, as the size of empty gap that marks
|
||||
// content as detached, and as the span past which those checks run at all.
|
||||
// This is a heuristic, not a format rule: PDF 2.0 sets no page-size limit,
|
||||
// and since PDF 1.6 `UserUnit` scales a page's physical size independently
|
||||
// of its coordinates. 14_400 units (200in at the default 1/72in unit) is
|
||||
// the traditional Acrobat architectural limit, which makes it a reasonable
|
||||
// "wider than any ordinary page" mark in coordinate space.
|
||||
const MAX_PAGE_EXTENT: f32 = 14_400.0;
|
||||
// A detached cluster is only dropped if it also holds a small minority of
|
||||
// the items, so a genuine two-part layout keeps its full bounds even when
|
||||
// the halves are far apart.
|
||||
const MAX_TRIM_FRACTION: f32 = 0.10;
|
||||
|
||||
// Position and width of each item, skipping only non-finite geometry.
|
||||
let finite_span = |i: &&TextItem| -> Option<(f32, f32)> {
|
||||
let (left, width) = (i.x, effective_width(i));
|
||||
(left.is_finite() && (left + width).is_finite()).then_some((left, width))
|
||||
};
|
||||
|
||||
let (min_left, max_right, total) = page_items.iter().filter_map(finite_span).fold(
|
||||
(f32::INFINITY, f32::NEG_INFINITY, 0usize),
|
||||
|(lo, hi, n), (left, width)| (lo.min(left), hi.max(left + width), n + 1),
|
||||
);
|
||||
|
||||
// No item had usable geometry, so there is no layout to report.
|
||||
if total == 0 {
|
||||
return vec![];
|
||||
}
|
||||
|
||||
// Every threshold below (gutter margins, spanning-item width, the XY-cut
|
||||
// margin) is a fraction of the page width, so a far item can set the scale
|
||||
// for the whole page and shrink the effective detection window to a
|
||||
// rounding error — real gutters then fall inside the margin band and a
|
||||
// genuine multi-column page collapses to one region.
|
||||
//
|
||||
// Anything inside one page extent is ordinary, so the common case keeps the
|
||||
// plain bounds and skips the work below entirely.
|
||||
let (x_min, x_max) = if max_right - min_left <= MAX_PAGE_EXTENT {
|
||||
(min_left, max_right)
|
||||
} else {
|
||||
// Discarding content needs positive evidence that it is not part of the
|
||||
// layout, because a count-based rule alone cannot tell a stray from a
|
||||
// sparse far sidebar. The evidence is geometric: positions are grouped
|
||||
// into clusters separated by more than a whole page of continuous
|
||||
// emptiness. Real content, however sparse, does not leave a void that
|
||||
// large; a malformed coordinate sits alone beyond one.
|
||||
let mut spans: Vec<(f32, f32)> = page_items.iter().filter_map(finite_span).collect();
|
||||
spans.sort_by(|a, b| a.0.total_cmp(&b.0));
|
||||
|
||||
let mut core: Option<std::ops::Range<usize>> = None;
|
||||
let mut start = 0usize;
|
||||
for i in 1..=spans.len() {
|
||||
if i < spans.len() && spans[i].0 - spans[i - 1].0 <= MAX_PAGE_EXTENT {
|
||||
continue;
|
||||
}
|
||||
if core.as_ref().is_none_or(|best| i - start > best.len()) {
|
||||
core = Some(start..i);
|
||||
}
|
||||
start = i;
|
||||
}
|
||||
let mut core = core.unwrap_or(0..spans.len());
|
||||
|
||||
// Only drop the detached clusters when they are a small minority, so a
|
||||
// genuine two-part layout keeps its full bounds.
|
||||
let dropped = spans.len() - core.len();
|
||||
if dropped as f32 > spans.len() as f32 * MAX_TRIM_FRACTION {
|
||||
core = 0..spans.len();
|
||||
}
|
||||
let core = &spans[core];
|
||||
|
||||
// Positions cannot be inflated by a bogus width, so the spread of the
|
||||
// content is a sound scale for judging one. A run much wider than the
|
||||
// page's own content is a malformed width — the test is relative, so a
|
||||
// genuinely large page keeps its genuinely long runs.
|
||||
let (lo, widest_left) = (core[0].0, core[core.len() - 1].0);
|
||||
let max_run_width = (widest_left - lo) + MAX_PAGE_EXTENT;
|
||||
let hi = core
|
||||
.iter()
|
||||
.filter(|&&(_, width)| width <= max_run_width)
|
||||
.map(|&(left, width)| left + width)
|
||||
.fold(widest_left, f32::max);
|
||||
|
||||
if lo != min_left || hi != max_right {
|
||||
debug!(
|
||||
"page {page}: bounds {min_left}..{max_right} exceed one page; \
|
||||
dropped {dropped}/{} detached item(s), using {lo}..{hi}",
|
||||
spans.len()
|
||||
);
|
||||
}
|
||||
(lo, hi)
|
||||
};
|
||||
|
||||
// Hard ceiling on the histogram size, independent of the trimming above:
|
||||
// the bounds are attacker-influenced, so an unclamped
|
||||
// `page_width / BIN_WIDTH` lets a crafted PDF force an arbitrarily large
|
||||
// `vec![0u32; num_bins]` allocation. 65_536 bins covers ~128k points at
|
||||
// BIN_WIDTH 2.0 — roughly 9x the largest legal page — so this never binds
|
||||
// on a real layout. Kept as a bound that does not depend on the outlier
|
||||
// heuristic staying correct.
|
||||
const MAX_BINS: usize = 65_536;
|
||||
|
||||
let page_width = x_max - x_min;
|
||||
if page_width < 200.0 {
|
||||
if !page_width.is_finite() || page_width < 200.0 {
|
||||
return vec![ColumnRegion { x_min, x_max }];
|
||||
}
|
||||
|
||||
@@ -57,13 +152,20 @@ pub(crate) fn detect_columns(
|
||||
return vec![ColumnRegion { x_min, x_max }];
|
||||
}
|
||||
|
||||
// Widen the bins rather than dropping the tail of the page. Clamping the
|
||||
// count alone would leave anything past MAX_BINS * BIN_WIDTH outside the
|
||||
// histogram, folded into the last bin, which places gutters at the wrong
|
||||
// coordinates. Scaling keeps full coverage under the same allocation
|
||||
// ceiling; only the resolution degrades, and only beyond ~131k points.
|
||||
let bin_width = BIN_WIDTH.max(page_width / MAX_BINS as f32);
|
||||
|
||||
// Build occupancy histogram.
|
||||
// Exclude items wider than 60% of page width — these are spanning items
|
||||
// (titles, full-width paragraphs) that would fill the gutter and prevent
|
||||
// detection of partial-page column layouts (e.g. two-column abstracts on
|
||||
// a page that also has single-column introduction text).
|
||||
let wide_threshold = page_width * 0.6;
|
||||
let num_bins = ((page_width / BIN_WIDTH).ceil() as usize).max(1);
|
||||
let num_bins = ((page_width / bin_width).ceil() as usize).clamp(1, MAX_BINS);
|
||||
let mut histogram = vec![0u32; num_bins];
|
||||
|
||||
for item in &page_items {
|
||||
@@ -71,8 +173,8 @@ pub(crate) fn detect_columns(
|
||||
if w > wide_threshold {
|
||||
continue;
|
||||
}
|
||||
let left = ((item.x - x_min) / BIN_WIDTH).floor() as usize;
|
||||
let right = (((item.x + w) - x_min) / BIN_WIDTH).ceil() as usize;
|
||||
let left = ((item.x - x_min) / bin_width).floor() as usize;
|
||||
let right = (((item.x + w) - x_min) / bin_width).ceil() as usize;
|
||||
let left = left.min(num_bins);
|
||||
let right = right.min(num_bins);
|
||||
for count in histogram.iter_mut().take(right).skip(left) {
|
||||
@@ -109,12 +211,12 @@ pub(crate) fn detect_columns(
|
||||
let valleys: Vec<(usize, usize)> = valleys
|
||||
.into_iter()
|
||||
.filter(|&(start, end)| {
|
||||
let width_pts = (end - start) as f32 * BIN_WIDTH;
|
||||
let width_pts = (end - start) as f32 * bin_width;
|
||||
if width_pts < MIN_GUTTER_WIDTH {
|
||||
return false;
|
||||
}
|
||||
// Valley center must not be within 5% of page edges
|
||||
let center_pts = ((start + end) as f32 / 2.0) * BIN_WIDTH;
|
||||
let center_pts = ((start + end) as f32 / 2.0) * bin_width;
|
||||
center_pts > margin_threshold && center_pts < (page_width - margin_threshold)
|
||||
})
|
||||
.collect();
|
||||
@@ -132,7 +234,7 @@ pub(crate) fn detect_columns(
|
||||
&histogram,
|
||||
num_bins,
|
||||
x_min,
|
||||
BIN_WIDTH,
|
||||
bin_width,
|
||||
page_width,
|
||||
margin_threshold,
|
||||
);
|
||||
@@ -141,7 +243,7 @@ pub(crate) fn detect_columns(
|
||||
&rel_valleys,
|
||||
&page_items,
|
||||
x_min,
|
||||
BIN_WIDTH,
|
||||
bin_width,
|
||||
x_max,
|
||||
MIN_ITEMS_PER_COLUMN,
|
||||
MIN_VERTICAL_SPAN_RATIO,
|
||||
@@ -182,7 +284,7 @@ pub(crate) fn detect_columns(
|
||||
&valleys,
|
||||
&page_items,
|
||||
x_min,
|
||||
BIN_WIDTH,
|
||||
bin_width,
|
||||
x_max,
|
||||
MIN_ITEMS_PER_COLUMN,
|
||||
MIN_VERTICAL_SPAN_RATIO,
|
||||
@@ -196,7 +298,7 @@ pub(crate) fn detect_columns(
|
||||
&valleys,
|
||||
&page_items,
|
||||
x_min,
|
||||
BIN_WIDTH,
|
||||
bin_width,
|
||||
x_max,
|
||||
MIN_ITEMS_PER_COLUMN,
|
||||
MIN_VERTICAL_SPAN_RATIO,
|
||||
@@ -1827,7 +1929,7 @@ fn split_column_stragglers(lines: Vec<TextLine>) -> (Vec<TextLine>, Vec<TextLine
|
||||
.unwrap();
|
||||
|
||||
let (cs, ce) = segments[core_seg];
|
||||
let mut core = Vec::with_capacity(ce - cs);
|
||||
let mut core = Vec::with_capacity(ce.saturating_sub(cs));
|
||||
let mut stragglers = Vec::new();
|
||||
for (i, line) in lines.into_iter().enumerate() {
|
||||
if i >= cs && i < ce {
|
||||
@@ -2533,6 +2635,214 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn extreme_far_coordinate_does_not_allocate_unboundedly() {
|
||||
// A crafted PDF can place a text run at an arbitrary coordinate via the
|
||||
// text matrix. The derived page width must not drive an unbounded
|
||||
// histogram allocation (previously `page_width / BIN_WIDTH` bins with no
|
||||
// upper bound would try to reserve terabytes and abort the process).
|
||||
let mut items = Vec::new();
|
||||
for i in 0..24 {
|
||||
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
|
||||
}
|
||||
// Item placed 1e12 points away — 5e11 bins if left unclamped.
|
||||
items.push(make_item(1, 1e12, 700.0, "Z"));
|
||||
|
||||
// Must return without aborting; content is preserved as a single region.
|
||||
let cols = detect_columns(&items, 1, false);
|
||||
assert!(!cols.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_finite_coordinates_never_leak_into_region_bounds() {
|
||||
// An inf/NaN coordinate must not escape as a column boundary: callers
|
||||
// treat these as page/column edges.
|
||||
for bad_x in [f32::INFINITY, f32::NEG_INFINITY, f32::NAN] {
|
||||
let mut items = Vec::new();
|
||||
for i in 0..24 {
|
||||
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
|
||||
}
|
||||
items.push(make_item(1, bad_x, 700.0, "Z"));
|
||||
|
||||
for col in detect_columns(&items, 1, false) {
|
||||
assert!(
|
||||
col.x_min.is_finite() && col.x_max.is_finite(),
|
||||
"bad_x {bad_x} leaked bounds {}..{}",
|
||||
col.x_min,
|
||||
col.x_max
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn all_non_finite_coordinates_yield_no_columns() {
|
||||
let items: Vec<TextItem> = (0..24)
|
||||
.map(|i| make_item(1, f32::NAN, 700.0 - i as f32 * 5.0, "A"))
|
||||
.collect();
|
||||
|
||||
assert!(detect_columns(&items, 1, false).is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn one_bad_item_does_not_disable_column_detection() {
|
||||
// A single stray item should not collapse a clean two-column page to
|
||||
// one region. Every gutter threshold is a fraction of the page width,
|
||||
// so an untrimmed outlier pushes real gutters inside the rejected
|
||||
// margin band. A malformed *width* at an ordinary position poisons the
|
||||
// bounds just as a malformed position does.
|
||||
for (label, bad_x, bad_width) in [
|
||||
("nan position", f32::NAN, 0.0),
|
||||
("inf position", f32::INFINITY, 0.0),
|
||||
("far position", 50_000.0, 0.0),
|
||||
("very far position", 1e12, 0.0),
|
||||
("huge width", 100.0, 1e12),
|
||||
("inf width", 100.0, f32::INFINITY),
|
||||
] {
|
||||
let mut items = Vec::new();
|
||||
items.extend(fill_zone(1, 30.0, 280.0, 750.0, 50.0));
|
||||
items.extend(fill_zone(1, 320.0, 570.0, 750.0, 50.0));
|
||||
let mut bad = make_item(1, bad_x, 400.0, "Z");
|
||||
bad.width = bad_width;
|
||||
items.push(bad);
|
||||
|
||||
let cols = detect_columns(&items, 1, false);
|
||||
assert_eq!(
|
||||
cols.len(),
|
||||
2,
|
||||
"{label}: expected 2 columns, got {}",
|
||||
cols.len()
|
||||
);
|
||||
for col in &cols {
|
||||
assert!(
|
||||
col.x_max - col.x_min <= MAX_PAGE_EXTENT_FOR_TEST,
|
||||
"{label}: region {}..{} exceeds one page",
|
||||
col.x_min,
|
||||
col.x_max
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Mirrors `MAX_PAGE_EXTENT` in `detect_columns`.
|
||||
const MAX_PAGE_EXTENT_FOR_TEST: f32 = 14_400.0;
|
||||
|
||||
#[test]
|
||||
fn very_wide_page_keeps_full_histogram_coverage() {
|
||||
// Beyond MAX_BINS * BIN_WIDTH (~131k points) the bins must widen rather
|
||||
// than stop covering the page. Three zones: the first gutter is inside
|
||||
// the old coverage limit, the second is past it. Because the first
|
||||
// gutter is found, the XY-cut fallback never runs, so a truncated
|
||||
// histogram silently reports two columns instead of three.
|
||||
let mut items = Vec::new();
|
||||
items.extend(fill_zone(1, 0.0, 60_000.0, 750.0, 700.0));
|
||||
items.extend(fill_zone(1, 70_000.0, 140_000.0, 750.0, 700.0));
|
||||
items.extend(fill_zone(1, 160_000.0, 200_000.0, 750.0, 700.0));
|
||||
|
||||
let cols = detect_columns(&items, 1, false);
|
||||
assert_eq!(
|
||||
cols.len(),
|
||||
3,
|
||||
"Expected 3 columns across a 200k-wide page, got {}",
|
||||
cols.len()
|
||||
);
|
||||
assert!(
|
||||
(140_000.0..=160_000.0).contains(&cols[1].x_max),
|
||||
"second gutter at {}, expected inside the real 140k..160k gap",
|
||||
cols[1].x_max
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn large_page_with_legitimately_long_runs_is_kept() {
|
||||
// On a very large page, individual runs can exceed one ordinary page's
|
||||
// width. They are real content, so they must not be judged malformed:
|
||||
// the page keeps its columns and its full right edge.
|
||||
let mut items = Vec::new();
|
||||
for row in 0..30 {
|
||||
let y = 750.0 - row as f32 * 14.0;
|
||||
let mut left = make_item(1, 0.0, y, "Left run");
|
||||
left.width = 20_000.0;
|
||||
let mut right = make_item(1, 25_000.0, y, "Right run");
|
||||
right.width = 20_000.0;
|
||||
items.extend([left, right]);
|
||||
}
|
||||
|
||||
let cols = detect_columns(&items, 1, false);
|
||||
assert!(
|
||||
!cols.is_empty(),
|
||||
"a page of long-but-valid runs must still report a layout"
|
||||
);
|
||||
let right_edge = cols
|
||||
.iter()
|
||||
.map(|c| c.x_max)
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
assert!(
|
||||
right_edge > 44_000.0,
|
||||
"long runs were treated as malformed: right edge {right_edge}, expected ~45_000"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn sparse_far_sidebar_on_a_large_page_is_kept() {
|
||||
// A large-format page with a thin, sparsely-populated sidebar far from
|
||||
// the main block. The sidebar is a small minority of the items, so an
|
||||
// item-count rule alone would discard it — but nothing about its
|
||||
// geometry says it is invalid, so its bounds must survive.
|
||||
let mut items = Vec::new();
|
||||
items.extend(fill_zone(1, 0.0, 12_000.0, 750.0, 500.0));
|
||||
for i in 0..12 {
|
||||
items.push(make_item(1, 24_000.0, 750.0 - i as f32 * 14.0, "Sidebar"));
|
||||
}
|
||||
|
||||
let cols = detect_columns(&items, 1, false);
|
||||
let right_edge = cols
|
||||
.iter()
|
||||
.map(|c| c.x_max)
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
assert!(
|
||||
right_edge > 24_000.0,
|
||||
"sidebar was trimmed away: right edge {right_edge}, expected >24_000"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn genuinely_wide_layout_keeps_its_true_bounds() {
|
||||
// A large-format page whose content really is spread beyond one
|
||||
// ordinary page must not be trimmed to the median cluster: its far
|
||||
// items are the majority, not strays.
|
||||
let mut items = Vec::new();
|
||||
items.extend(fill_zone(1, 100.0, 20_000.0, 750.0, 600.0));
|
||||
items.extend(fill_zone(1, 22_000.0, 40_000.0, 750.0, 600.0));
|
||||
|
||||
let cols = detect_columns(&items, 1, false);
|
||||
let widest = cols
|
||||
.iter()
|
||||
.map(|c| c.x_max)
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
assert!(
|
||||
widest > 35_000.0,
|
||||
"wide layout was trimmed: right edge {widest}, expected ~40_000"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn oversized_but_legal_page_is_not_trimmed() {
|
||||
// A wide-format page well inside the 14_400pt spec limit must keep its
|
||||
// real bounds — outlier trimming is only for spans beyond a legal page.
|
||||
let mut items = Vec::new();
|
||||
items.extend(fill_zone(1, 100.0, 4_000.0, 750.0, 400.0));
|
||||
items.extend(fill_zone(1, 4_400.0, 8_000.0, 750.0, 400.0));
|
||||
|
||||
let cols = detect_columns(&items, 1, false);
|
||||
assert_eq!(cols.len(), 2, "Expected 2 columns, got {}", cols.len());
|
||||
assert!(
|
||||
cols[1].x_max > 7_000.0,
|
||||
"right column should keep its true extent, got {}",
|
||||
cols[1].x_max
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn two_column_regression_guard() {
|
||||
// Standard 2-column layout with clear gutter at center
|
||||
|
||||
+311
-5
@@ -2,11 +2,51 @@
|
||||
|
||||
use crate::types::{ItemType, TextItem};
|
||||
use lopdf::{Document, Object, ObjectId};
|
||||
use std::collections::HashMap;
|
||||
use std::collections::{HashMap, HashSet};
|
||||
|
||||
use super::fonts::{resolve_array, resolve_dict};
|
||||
use super::get_number;
|
||||
|
||||
/// Upper bound on the number of form-field nodes visited during a single
|
||||
/// `extract_form_fields` pass. A crafted PDF can chain thousands of distinct
|
||||
/// `/Kids` fields to blow the stack even without an outright reference cycle,
|
||||
/// so we cap total traversal work in addition to detecting cycles.
|
||||
const MAX_FORM_FIELD_NODES: usize = 100_000;
|
||||
|
||||
/// Upper bound on `/Kids` recursion depth. Real AcroForm hierarchies are only
|
||||
/// a few levels deep (fields → child fields → widgets); a crafted PDF can chain
|
||||
/// tens of thousands of distinct fields into a linear `/Kids` list that would
|
||||
/// overflow the stack via depth-first recursion long before the node budget is
|
||||
/// reached. This depth cap bounds the stack independently of total node count.
|
||||
const MAX_FORM_FIELD_DEPTH: usize = 100;
|
||||
|
||||
/// Traversal budget for the AcroForm field walk. Bounds both the number of
|
||||
/// distinct nodes visited *and* the total number of `/Fields`/`/Kids` entries
|
||||
/// examined.
|
||||
///
|
||||
/// Counting `visited` alone is not enough: invalid entries (non-references) and
|
||||
/// duplicate references never grow `visited`, so an oversized array full of them
|
||||
/// would iterate to completion no matter how large. Charging every examined
|
||||
/// entry against the same budget makes it a real cap on traversal work.
|
||||
pub(crate) struct FieldWalkBudget {
|
||||
visited: HashSet<ObjectId>,
|
||||
examined: usize,
|
||||
}
|
||||
|
||||
impl FieldWalkBudget {
|
||||
fn new() -> Self {
|
||||
Self {
|
||||
visited: HashSet::new(),
|
||||
examined: 0,
|
||||
}
|
||||
}
|
||||
|
||||
/// True once the budget is spent; callers must stop iterating and recursing.
|
||||
fn exhausted(&self) -> bool {
|
||||
self.visited.len() >= MAX_FORM_FIELD_NODES || self.examined >= MAX_FORM_FIELD_NODES
|
||||
}
|
||||
}
|
||||
|
||||
pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> Vec<TextItem> {
|
||||
let mut links = Vec::new();
|
||||
|
||||
@@ -146,9 +186,12 @@ pub(crate) fn extract_form_fields(
|
||||
Err(_) => return items,
|
||||
};
|
||||
|
||||
// Borrow the array rather than cloning it: a crafted `/Fields` can be huge,
|
||||
// and cloning would pay an O(n) allocation/copy before the budget check
|
||||
// below can stop the work.
|
||||
let fields = match acroform.get(b"Fields") {
|
||||
Ok(obj) => match resolve_array(doc, obj) {
|
||||
Some(arr) => arr.clone(),
|
||||
Some(arr) => arr,
|
||||
None => return items,
|
||||
},
|
||||
Err(_) => return items,
|
||||
@@ -158,7 +201,19 @@ pub(crate) fn extract_form_fields(
|
||||
}
|
||||
let annotation_pages = annotation_page_map(doc, page_map);
|
||||
|
||||
for field_obj in &fields {
|
||||
// Bound the walk so a crafted PDF cannot send us into unbounded recursion
|
||||
// via a `/Kids` cycle, a deep chain, or an oversized array of invalid or
|
||||
// duplicate entries.
|
||||
let mut budget = FieldWalkBudget::new();
|
||||
|
||||
for field_obj in fields {
|
||||
// Stop once the budget is spent so a `/Fields` array wider than the
|
||||
// budget can't burn CPU iterating entries whose walk would no-op. Charge
|
||||
// every entry (including invalid ones) against the budget.
|
||||
if budget.exhausted() {
|
||||
break;
|
||||
}
|
||||
budget.examined += 1;
|
||||
if let Ok(field_ref) = field_obj.as_reference() {
|
||||
walk_form_fields(
|
||||
doc,
|
||||
@@ -168,6 +223,8 @@ pub(crate) fn extract_form_fields(
|
||||
page_map,
|
||||
&annotation_pages,
|
||||
&mut items,
|
||||
&mut budget,
|
||||
0,
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -202,6 +259,7 @@ fn annotation_page_map(
|
||||
}
|
||||
|
||||
/// Recursively walk the form field tree, extracting leaf field values.
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
pub(crate) fn walk_form_fields(
|
||||
doc: &Document,
|
||||
field_id: ObjectId,
|
||||
@@ -210,7 +268,22 @@ pub(crate) fn walk_form_fields(
|
||||
page_map: &HashMap<ObjectId, u32>,
|
||||
annotation_pages: &HashMap<ObjectId, u32>,
|
||||
items: &mut Vec<TextItem>,
|
||||
budget: &mut FieldWalkBudget,
|
||||
depth: usize,
|
||||
) {
|
||||
// Guard against `/Kids` cycles and pathologically large field trees.
|
||||
// Exceeding the depth cap means the chain is too deep to be a legitimate
|
||||
// form (and would overflow the stack); an exhausted budget means the tree is
|
||||
// too large. Both checks run *before* inserting so the visited set can never
|
||||
// grow past the budget.
|
||||
if depth > MAX_FORM_FIELD_DEPTH || budget.exhausted() {
|
||||
return;
|
||||
}
|
||||
// Revisiting an object ID means we hit a `/Kids` cycle.
|
||||
if !budget.visited.insert(field_id) {
|
||||
return;
|
||||
}
|
||||
|
||||
let field_dict = match doc.get_dictionary(field_id) {
|
||||
Ok(d) => d,
|
||||
Err(_) => return,
|
||||
@@ -241,9 +314,19 @@ pub(crate) fn walk_form_fields(
|
||||
|
||||
// Check for /Kids — if present, recurse into children
|
||||
if let Ok(kids_obj) = field_dict.get(b"Kids") {
|
||||
// Iterate the borrowed array directly — cloning a crafted, oversized
|
||||
// `/Kids` would allocate and copy every entry before the budget check
|
||||
// below could stop the work.
|
||||
if let Some(kids) = resolve_array(doc, kids_obj) {
|
||||
let kids = kids.clone();
|
||||
for kid in &kids {
|
||||
for kid in kids {
|
||||
// Stop once the budget is spent so a `/Kids` array wider than the
|
||||
// budget can't burn CPU iterating entries whose walk would no-op.
|
||||
// Charge every entry (including invalid/duplicate ones) against
|
||||
// the budget so this is a true traversal-work cap.
|
||||
if budget.exhausted() {
|
||||
break;
|
||||
}
|
||||
budget.examined += 1;
|
||||
if let Ok(kid_ref) = kid.as_reference() {
|
||||
walk_form_fields(
|
||||
doc,
|
||||
@@ -253,6 +336,8 @@ pub(crate) fn walk_form_fields(
|
||||
page_map,
|
||||
annotation_pages,
|
||||
items,
|
||||
budget,
|
||||
depth + 1,
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -411,4 +496,225 @@ mod tests {
|
||||
assert_eq!(items[0].page, 2);
|
||||
assert_eq!(items[0].text, "customer: Alice");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn kids_self_cycle_does_not_overflow_stack() {
|
||||
// A crafted AcroForm field that lists itself in `/Kids` must not send
|
||||
// the traversal into unbounded recursion.
|
||||
let mut doc = Document::new();
|
||||
let field_id = doc.new_object_id();
|
||||
doc.set_object(
|
||||
field_id,
|
||||
dictionary! {
|
||||
"FT" => "Tx",
|
||||
"T" => Object::string_literal("loop"),
|
||||
"Kids" => vec![Object::Reference(field_id)],
|
||||
},
|
||||
);
|
||||
let catalog_id = doc.add_object(dictionary! {
|
||||
"Type" => "Catalog",
|
||||
"AcroForm" => dictionary! {
|
||||
"Fields" => vec![Object::Reference(field_id)],
|
||||
},
|
||||
});
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
|
||||
let page_map = HashMap::new();
|
||||
// Completes (rather than overflowing the stack) and yields no items.
|
||||
let items = extract_form_fields(&doc, &page_map);
|
||||
assert!(items.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn kids_mutual_cycle_terminates() {
|
||||
// Two fields that reference each other via `/Kids` form a cycle that
|
||||
// must also terminate.
|
||||
let mut doc = Document::new();
|
||||
let field_a = doc.new_object_id();
|
||||
let field_b = doc.new_object_id();
|
||||
doc.set_object(
|
||||
field_a,
|
||||
dictionary! {
|
||||
"T" => Object::string_literal("a"),
|
||||
"Kids" => vec![Object::Reference(field_b)],
|
||||
},
|
||||
);
|
||||
doc.set_object(
|
||||
field_b,
|
||||
dictionary! {
|
||||
"T" => Object::string_literal("b"),
|
||||
"Kids" => vec![Object::Reference(field_a)],
|
||||
},
|
||||
);
|
||||
let catalog_id = doc.add_object(dictionary! {
|
||||
"Type" => "Catalog",
|
||||
"AcroForm" => dictionary! {
|
||||
"Fields" => vec![Object::Reference(field_a)],
|
||||
},
|
||||
});
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
|
||||
let page_map = HashMap::new();
|
||||
let items = extract_form_fields(&doc, &page_map);
|
||||
assert!(items.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn deep_acyclic_kids_chain_does_not_overflow_stack() {
|
||||
// A long chain of *distinct* fields (no cycle) must also terminate:
|
||||
// the visited set alone would still recurse to the chain length, so
|
||||
// the depth cap is what prevents a stack overflow here.
|
||||
let mut doc = Document::new();
|
||||
let n = MAX_FORM_FIELD_DEPTH * 500;
|
||||
let ids: Vec<ObjectId> = (0..=n).map(|_| doc.new_object_id()).collect();
|
||||
for i in 0..n {
|
||||
doc.set_object(
|
||||
ids[i],
|
||||
dictionary! {
|
||||
"FT" => "Tx",
|
||||
"Kids" => vec![Object::Reference(ids[i + 1])],
|
||||
},
|
||||
);
|
||||
}
|
||||
// Leaf carries a value; it sits far below the depth cap so it is never
|
||||
// reached, proving traversal stops early rather than crashing.
|
||||
doc.set_object(
|
||||
ids[n],
|
||||
dictionary! {
|
||||
"FT" => "Tx",
|
||||
"T" => Object::string_literal("leaf"),
|
||||
"V" => Object::string_literal("x"),
|
||||
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
|
||||
},
|
||||
);
|
||||
let catalog_id = doc.add_object(dictionary! {
|
||||
"Type" => "Catalog",
|
||||
"AcroForm" => dictionary! {
|
||||
"Fields" => vec![Object::Reference(ids[0])],
|
||||
},
|
||||
});
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
|
||||
let page_map = HashMap::new();
|
||||
let items = extract_form_fields(&doc, &page_map);
|
||||
assert!(items.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn wide_tree_traversal_stops_at_node_budget() {
|
||||
// A single field with a `/Kids` array wider than the node budget must
|
||||
// stop traversal at the cap rather than growing `visited` (and the work)
|
||||
// without bound. Each processed leaf emits one item, so the item count
|
||||
// is bounded by the budget and reaches right up to it (a couple of
|
||||
// slots go to the root and the boundary node charged against the cap).
|
||||
let mut doc = Document::new();
|
||||
let fanout = MAX_FORM_FIELD_NODES + 50;
|
||||
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
|
||||
for &leaf in &leaf_ids {
|
||||
doc.set_object(
|
||||
leaf,
|
||||
dictionary! {
|
||||
"FT" => "Tx",
|
||||
"V" => Object::string_literal("v"),
|
||||
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
|
||||
},
|
||||
);
|
||||
}
|
||||
let kids: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
|
||||
let root_id = doc.add_object(dictionary! {
|
||||
"T" => Object::string_literal("root"),
|
||||
"Kids" => kids,
|
||||
});
|
||||
let catalog_id = doc.add_object(dictionary! {
|
||||
"Type" => "Catalog",
|
||||
"AcroForm" => dictionary! {
|
||||
"Fields" => vec![Object::Reference(root_id)],
|
||||
},
|
||||
});
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
|
||||
let page_map = HashMap::new();
|
||||
let items = extract_form_fields(&doc, &page_map);
|
||||
// Extraction stops at the budget: bounded above by the cap, and it gets
|
||||
// right up to it (allowing a small delta for the root/boundary nodes
|
||||
// charged against the budget).
|
||||
assert!(items.len() <= MAX_FORM_FIELD_NODES);
|
||||
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn wide_top_level_fields_stop_at_node_budget() {
|
||||
// A top-level `/Fields` array wider than the budget must also stop at
|
||||
// the cap: the item count is bounded by the budget and reaches right up
|
||||
// to it.
|
||||
let mut doc = Document::new();
|
||||
let fanout = MAX_FORM_FIELD_NODES + 50;
|
||||
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
|
||||
for &leaf in &leaf_ids {
|
||||
doc.set_object(
|
||||
leaf,
|
||||
dictionary! {
|
||||
"FT" => "Tx",
|
||||
"V" => Object::string_literal("v"),
|
||||
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
|
||||
},
|
||||
);
|
||||
}
|
||||
let fields: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
|
||||
let catalog_id = doc.add_object(dictionary! {
|
||||
"Type" => "Catalog",
|
||||
"AcroForm" => dictionary! {
|
||||
"Fields" => fields,
|
||||
},
|
||||
});
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
|
||||
let page_map = HashMap::new();
|
||||
let items = extract_form_fields(&doc, &page_map);
|
||||
assert!(items.len() <= MAX_FORM_FIELD_NODES);
|
||||
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn duplicate_and_invalid_kids_entries_stop_at_budget() {
|
||||
// Duplicate references and non-reference junk never grow `visited`, so
|
||||
// without charging examined entries against the budget an oversized
|
||||
// array of them would iterate to completion. The walk must still
|
||||
// terminate and extract the single real leaf exactly once.
|
||||
let mut doc = Document::new();
|
||||
let leaf_id = doc.new_object_id();
|
||||
doc.set_object(
|
||||
leaf_id,
|
||||
dictionary! {
|
||||
"FT" => "Tx",
|
||||
"V" => Object::string_literal("v"),
|
||||
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
|
||||
},
|
||||
);
|
||||
// A `/Kids` array far wider than the budget: half duplicate references
|
||||
// to the same leaf, half invalid (null) entries.
|
||||
let mut kids: Vec<Object> = Vec::new();
|
||||
for i in 0..(MAX_FORM_FIELD_NODES * 2) {
|
||||
if i % 2 == 0 {
|
||||
kids.push(Object::Reference(leaf_id));
|
||||
} else {
|
||||
kids.push(Object::Null);
|
||||
}
|
||||
}
|
||||
let root_id = doc.add_object(dictionary! {
|
||||
"T" => Object::string_literal("root"),
|
||||
"Kids" => kids,
|
||||
});
|
||||
let catalog_id = doc.add_object(dictionary! {
|
||||
"Type" => "Catalog",
|
||||
"AcroForm" => dictionary! {
|
||||
"Fields" => vec![Object::Reference(root_id)],
|
||||
},
|
||||
});
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
|
||||
let page_map = HashMap::new();
|
||||
let items = extract_form_fields(&doc, &page_map);
|
||||
assert_eq!(items.len(), 1);
|
||||
}
|
||||
}
|
||||
|
||||
+253
-14
@@ -3,6 +3,7 @@
|
||||
//! This module extracts text with position information for structure detection.
|
||||
|
||||
mod base14;
|
||||
mod content_decode;
|
||||
pub(crate) mod content_stream;
|
||||
mod fonts;
|
||||
mod layout;
|
||||
@@ -37,6 +38,7 @@ pub(crate) use layout::group_prefiltered_items_into_lines_with_thresholds_and_re
|
||||
pub(crate) use layout::is_newspaper_layout;
|
||||
pub(crate) use layout::ColumnRegion;
|
||||
pub use layout::{group_into_lines, group_into_lines_preserving_all_text};
|
||||
pub(crate) use xobjects::FormWalkBudget;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Public API
|
||||
@@ -282,18 +284,22 @@ fn extract_positioned_text_impl(
|
||||
font_cmaps,
|
||||
include_invisible,
|
||||
&mut style_cache,
|
||||
&mut FormWalkBudget::new(),
|
||||
);
|
||||
let ((mut items, mut rects, mut lines), has_gid_fonts, coords_rotated) = match page_result {
|
||||
Ok(extraction) => extraction,
|
||||
Err(error) if required_pages.is_some_and(|required| !required.contains(page_num)) => {
|
||||
debug!(
|
||||
"page {}: skipping context-only extraction error: {}",
|
||||
page_num, error
|
||||
);
|
||||
continue;
|
||||
}
|
||||
Err(error) => return Err(error),
|
||||
};
|
||||
let ((mut items, mut rects, mut lines), has_gid_fonts, coords_rotated, _skipped_invisible) =
|
||||
match page_result {
|
||||
Ok(extraction) => extraction,
|
||||
Err(error)
|
||||
if required_pages.is_some_and(|required| !required.contains(page_num)) =>
|
||||
{
|
||||
debug!(
|
||||
"page {}: skipping context-only extraction error: {}",
|
||||
page_num, error
|
||||
);
|
||||
continue;
|
||||
}
|
||||
Err(error) => return Err(error),
|
||||
};
|
||||
// Clip to the visible page box: single-page extracts and imposed
|
||||
// spreads keep neighboring pages' content in the stream, positioned
|
||||
// outside the CropBox. Extracting it interleaves invisible text into
|
||||
@@ -884,6 +890,92 @@ fn tracked_run_space_floor(group: &[&TextItem], start: usize) -> Option<(usize,
|
||||
Some((end, floor * fs))
|
||||
}
|
||||
|
||||
/// Fractional font-size band within which `merge_text_items` treats two runs as
|
||||
/// the same size. Shared with `is_small_caps_continuation`, which exists only to
|
||||
/// rescue junctions this band would otherwise break.
|
||||
const MERGE_FONT_SIZE_BAND: f32 = 0.20;
|
||||
|
||||
/// Detect a small-caps continuation: typesetters render small caps as a
|
||||
/// full-size capital immediately followed by shrunken capitals in the same
|
||||
/// font (`(R) Tj` at 9.98pt, then `(OLANDO) Tj` at 6.74pt). Those runs are one
|
||||
/// word, but the font-size band in `merge_text_items` would split them,
|
||||
/// leaving "R" and "OLANDO" as separate items — which then read as separate
|
||||
/// table columns, since column boundaries cluster on item start positions.
|
||||
///
|
||||
/// Gated tightly so it cannot absorb the other reasons a smaller run follows a
|
||||
/// larger one:
|
||||
/// - runs the size band already accepts — excluded by requiring the junction
|
||||
/// to *cross* the band, so within-band pairs keep the normal word-spacing
|
||||
/// logic instead of having their space suppressed
|
||||
/// - superscripts / footnote markers — excluded by requiring an uppercase
|
||||
/// *letter* on both sides, so digits never qualify
|
||||
/// - drop caps — excluded because the body text that follows is mixed case
|
||||
/// - adjacent table cells or separate words — excluded by requiring the runs
|
||||
/// to be visually contiguous (essentially no gap)
|
||||
fn is_small_caps_continuation(
|
||||
text_so_far: &str,
|
||||
first: &TextItem,
|
||||
next: &TextItem,
|
||||
gap: f32,
|
||||
) -> bool {
|
||||
// Must shrink. Real small caps sit near 0.7-0.8 of the full cap height;
|
||||
// anything smaller is a superscript or a different run entirely.
|
||||
if first.font_size <= 0.0 || next.font_size >= first.font_size {
|
||||
return false;
|
||||
}
|
||||
// Only rescue junctions the size band would have broken. Within-band pairs
|
||||
// merge on their own, and suppressing their space would swallow real word
|
||||
// gaps between two similarly-sized uppercase words.
|
||||
if (next.font_size - first.font_size).abs() <= first.font_size * MERGE_FONT_SIZE_BAND {
|
||||
return false;
|
||||
}
|
||||
if next.font_size / first.font_size < 0.55 {
|
||||
return false;
|
||||
}
|
||||
// Visually contiguous: the capital and its small caps touch. A real word
|
||||
// space or a column gap disqualifies.
|
||||
if !(-first.font_size * 0.2..=first.font_size * 0.15).contains(&gap) {
|
||||
return false;
|
||||
}
|
||||
// The continuation must be all-uppercase letters (digits and lowercase
|
||||
// both disqualify), and must contain at least one letter.
|
||||
let mut saw_letter = false;
|
||||
for ch in next.text.chars() {
|
||||
if ch.is_alphabetic() {
|
||||
saw_letter = true;
|
||||
if !ch.is_uppercase() {
|
||||
return false;
|
||||
}
|
||||
} else if ch.is_numeric() {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
if !saw_letter {
|
||||
return false;
|
||||
}
|
||||
// What we are continuing must itself end in a capital. Check the actual
|
||||
// trailing character rather than skipping back to the nearest letter: after
|
||||
// "ANGELA M. MAZZARELLI1" the run to continue is the footnote marker, not
|
||||
// the "I" before it.
|
||||
let trimmed = text_so_far.trim_end();
|
||||
if trimmed.chars().last().is_some_and(|c| c.is_numeric()) {
|
||||
// One legitimate exception: an ordinal suffix set as a smaller run,
|
||||
// e.g. "JULY 4" + "TH". Only the four English suffixes qualify —
|
||||
// anything else after a digit is a footnote marker or numeric suffix.
|
||||
return matches!(trimmed_suffix(next), "TH" | "ST" | "ND" | "RD");
|
||||
}
|
||||
trimmed
|
||||
.chars()
|
||||
.rev()
|
||||
.find(|c| c.is_alphabetic())
|
||||
.is_some_and(|c| c.is_uppercase())
|
||||
}
|
||||
|
||||
/// The continuation run's text, trimmed — used to spot ordinal suffixes.
|
||||
fn trimmed_suffix(next: &TextItem) -> &str {
|
||||
next.text.trim()
|
||||
}
|
||||
|
||||
pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
|
||||
if items.is_empty() {
|
||||
return items;
|
||||
@@ -942,8 +1034,16 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
|
||||
let mut j = i + 1;
|
||||
while j < group.len() {
|
||||
let next = group[j];
|
||||
// Must be similar font size (within 20%)
|
||||
if (next.font_size - first.font_size).abs() > first.font_size * 0.20 {
|
||||
// A small-caps junction is mid-word: it both survives the
|
||||
// font-size band below and must never take a space.
|
||||
let small_caps_join =
|
||||
is_small_caps_continuation(&text, first, next, next.x - end_x);
|
||||
// Must be similar font size, except for genuine small-caps
|
||||
// runs, where the shrunken capitals are the same word as the
|
||||
// full-size initial (see helper).
|
||||
if (next.font_size - first.font_size).abs() > first.font_size * MERGE_FONT_SIZE_BAND
|
||||
&& !small_caps_join
|
||||
{
|
||||
break;
|
||||
}
|
||||
// Never merge across style boundaries: the merged item
|
||||
@@ -997,7 +1097,7 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
|
||||
Some((run_end, floor)) if j <= run_end => floor,
|
||||
_ => threshold,
|
||||
};
|
||||
if needs_bullet_space || gap > effective_threshold {
|
||||
if !small_caps_join && (needs_bullet_space || gap > effective_threshold) {
|
||||
text.push(' ');
|
||||
}
|
||||
text.push_str(&next.text);
|
||||
@@ -3001,6 +3101,145 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
/// Small caps as typesetters emit them: a full-size capital at 9.98pt
|
||||
/// immediately followed by shrunken capitals at 6.74pt, touching.
|
||||
/// Modelled on `199AD3d.pdf` p.5 ("ROLANDO T. ACOSTA, P.J.").
|
||||
#[test]
|
||||
fn small_caps_run_merges_into_one_word() {
|
||||
let items = vec![
|
||||
make_item_fs("R", 144.36, 581.84, 7.20, 9.98),
|
||||
make_item_fs("OLANDO", 151.56, 581.84, 30.56, 6.74),
|
||||
make_item_fs("T. A", 185.45, 581.84, 17.58, 9.98),
|
||||
make_item_fs("COSTA", 203.94, 581.84, 23.15, 6.74),
|
||||
make_item_fs(", P.J.", 227.09, 581.84, 22.56, 9.98),
|
||||
];
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1, "got {:?}", merged);
|
||||
assert_eq!(merged[0].text, "ROLANDO T. ACOSTA, P.J.");
|
||||
}
|
||||
|
||||
/// The full two-column row: both names must merge independently and the
|
||||
/// 72pt column gap between them must survive as an item boundary.
|
||||
#[test]
|
||||
fn small_caps_merge_does_not_swallow_a_second_column() {
|
||||
let items = vec![
|
||||
// Column 1: "ROLANDO T. ACOSTA, P.J." ending at x=249.65
|
||||
make_item_fs("R", 144.36, 581.84, 7.20, 9.98),
|
||||
make_item_fs("OLANDO", 151.56, 581.84, 30.56, 6.74),
|
||||
make_item_fs("T. A", 185.45, 581.84, 17.58, 9.98),
|
||||
make_item_fs("COSTA", 203.94, 581.84, 23.15, 6.74),
|
||||
make_item_fs(", P.J.", 227.09, 581.84, 22.56, 9.98),
|
||||
// Column 2 starts at x=321.96 — a 72pt gap.
|
||||
make_item_fs("A", 321.96, 581.84, 7.20, 9.98),
|
||||
make_item_fs("NIL", 329.17, 581.84, 12.72, 6.74),
|
||||
make_item_fs("C. S", 345.04, 581.84, 19.59, 9.98),
|
||||
make_item_fs("INGH", 364.62, 581.84, 19.08, 6.74),
|
||||
];
|
||||
let merged = merge_text_items(items);
|
||||
let texts: Vec<&str> = merged.iter().map(|i| i.text.as_str()).collect();
|
||||
assert_eq!(
|
||||
texts,
|
||||
vec!["ROLANDO T. ACOSTA, P.J.", "ANIL C. SINGH"],
|
||||
"column gap should keep the two names apart"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn small_caps_merge_keeps_word_space_between_same_size_capitals() {
|
||||
// Two uppercase words at sizes the merge band already accepts (9.98 and
|
||||
// 9.0, a 10% drop) separated by a real word gap. The small-caps path
|
||||
// must not claim this junction and swallow the space.
|
||||
let items = vec![
|
||||
make_item_fs("SEE", 100.0, 500.0, 18.0, 9.98),
|
||||
make_item_fs("ALSO", 119.2, 500.0, 24.0, 9.0),
|
||||
];
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1, "got {:?}", merged);
|
||||
assert_eq!(merged[0].text, "SEE ALSO");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn trailing_digit_is_not_a_capital_awaiting_small_caps() {
|
||||
// "...MAZZARELLI1" ends in a footnote marker; the backward search for an
|
||||
// uppercase letter must not skip the digit and glue the next run.
|
||||
assert!(!is_small_caps_continuation(
|
||||
"ANGELA M. MAZZARELLI1",
|
||||
&make_item_fs("ANGELA", 100.0, 500.0, 40.0, 9.98),
|
||||
&make_item_fs("SHULMAN", 140.0, 500.0, 30.0, 6.74),
|
||||
0.0,
|
||||
));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ordinal_suffix_after_a_digit_still_merges() {
|
||||
// "TUESDAY, JULY 4" + "TH" is one word in the source; the digit guard
|
||||
// must not block the four English ordinal suffixes.
|
||||
for suffix in ["TH", "ST", "ND", "RD"] {
|
||||
assert!(
|
||||
is_small_caps_continuation(
|
||||
"TUESDAY, JULY 4",
|
||||
&make_item_fs("JULY", 100.0, 500.0, 30.0, 12.0),
|
||||
&make_item_fs(suffix, 130.0, 500.0, 8.0, 8.0),
|
||||
0.0,
|
||||
),
|
||||
"{suffix} should merge after a digit"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn superscript_footnote_marker_is_not_a_small_caps_continuation() {
|
||||
// A digit must never qualify — otherwise footnote markers get glued on
|
||||
// without the superscript handling.
|
||||
assert!(!is_small_caps_continuation(
|
||||
"MAZZARELLI",
|
||||
&make_item_fs("MAZZARELLI", 100.0, 500.0, 50.0, 9.98),
|
||||
&make_item_fs("1", 150.0, 503.0, 3.0, 6.74),
|
||||
0.0,
|
||||
));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn drop_cap_is_not_a_small_caps_continuation() {
|
||||
// Mixed-case body text after a large initial is a drop cap, not small
|
||||
// caps.
|
||||
assert!(!is_small_caps_continuation(
|
||||
"T",
|
||||
&make_item_fs("T", 100.0, 500.0, 20.0, 30.0),
|
||||
&make_item_fs("he court held", 120.0, 500.0, 60.0, 10.0),
|
||||
0.0,
|
||||
));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn separate_word_is_not_a_small_caps_continuation() {
|
||||
// A real word space disqualifies even when both runs are uppercase.
|
||||
let first = make_item_fs("SEE", 100.0, 500.0, 20.0, 9.98);
|
||||
let next = make_item_fs("ALSO", 128.0, 500.0, 25.0, 6.74);
|
||||
assert!(!is_small_caps_continuation("SEE", &first, &next, 8.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn lowercase_continuation_is_not_small_caps() {
|
||||
assert!(!is_small_caps_continuation(
|
||||
"SMALL",
|
||||
&make_item_fs("SMALL", 100.0, 500.0, 30.0, 9.98),
|
||||
&make_item_fs("caps", 130.0, 500.0, 20.0, 6.74),
|
||||
0.0,
|
||||
));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn too_small_a_ratio_is_not_small_caps() {
|
||||
// 0.4 ratio is a superscript/sub-run, outside the small-caps band.
|
||||
assert!(!is_small_caps_continuation(
|
||||
"A",
|
||||
&make_item_fs("A", 100.0, 500.0, 7.0, 10.0),
|
||||
&make_item_fs("BC", 107.0, 500.0, 8.0, 4.0),
|
||||
0.0,
|
||||
));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_merge_subscript_items_chemical_formula() {
|
||||
// NH₃: "NH" at fs=8 followed by subscript "3" at fs=4.7
|
||||
|
||||
+600
-21
@@ -16,6 +16,76 @@ use super::{get_number, image_bbox_from_ctm, multiply_matrices};
|
||||
|
||||
const MAX_FORM_XOBJECT_DEPTH: u8 = 5;
|
||||
|
||||
/// Upper bound on Form XObject invocations during a single page extraction.
|
||||
/// Depth alone is not enough: an acyclic DAG where each form invokes the next
|
||||
/// N times expands to N^depth work before the depth cap is reached.
|
||||
const MAX_FORM_XOBJECT_INVOCATIONS: usize = 10_000;
|
||||
|
||||
/// Upper bound on content-stream operations walked across all Form XObject
|
||||
/// expansions for a page. Nested forms are decoded independently of the
|
||||
/// page-level operation cap, so this keeps total form work in the same
|
||||
/// ballpark as that page cap.
|
||||
const MAX_FORM_XOBJECT_OPERATIONS: usize = 1_000_000;
|
||||
|
||||
/// Shared budget for Form XObject expansion on a page. Bounds both nested DAG
|
||||
/// expansion and repeated sibling `/Do` invocations of the same form.
|
||||
pub(crate) struct FormWalkBudget {
|
||||
invocations: usize,
|
||||
operations: usize,
|
||||
max_invocations: usize,
|
||||
max_operations: usize,
|
||||
truncated: bool,
|
||||
}
|
||||
|
||||
impl FormWalkBudget {
|
||||
pub(crate) fn new() -> Self {
|
||||
Self::with_limits(MAX_FORM_XOBJECT_INVOCATIONS, MAX_FORM_XOBJECT_OPERATIONS)
|
||||
}
|
||||
|
||||
fn with_limits(max_invocations: usize, max_operations: usize) -> Self {
|
||||
Self {
|
||||
invocations: 0,
|
||||
operations: 0,
|
||||
max_invocations,
|
||||
max_operations,
|
||||
truncated: false,
|
||||
}
|
||||
}
|
||||
|
||||
fn exhausted(&mut self) -> bool {
|
||||
if self.invocations >= self.max_invocations || self.operations >= self.max_operations {
|
||||
self.truncated = true;
|
||||
true
|
||||
} else {
|
||||
false
|
||||
}
|
||||
}
|
||||
|
||||
fn charge_invocation(&mut self) -> bool {
|
||||
if self.exhausted() {
|
||||
return false;
|
||||
}
|
||||
self.invocations += 1;
|
||||
true
|
||||
}
|
||||
|
||||
/// Charge one walked content-stream operator. Independent of the
|
||||
/// invocation cap so a form that was already admitted can finish its
|
||||
/// stream (up to the operation cap).
|
||||
fn charge_operation(&mut self) -> bool {
|
||||
if self.operations >= self.max_operations {
|
||||
self.truncated = true;
|
||||
return false;
|
||||
}
|
||||
self.operations += 1;
|
||||
true
|
||||
}
|
||||
|
||||
pub(crate) fn was_truncated(&self) -> bool {
|
||||
self.truncated
|
||||
}
|
||||
}
|
||||
|
||||
pub(crate) enum XObjectType {
|
||||
Image,
|
||||
Form(ObjectId),
|
||||
@@ -109,6 +179,7 @@ fn collect_xobjects_from_dict(
|
||||
}
|
||||
|
||||
/// Extract text items from a Form XObject
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
pub(crate) fn extract_form_xobject_text(
|
||||
doc: &Document,
|
||||
form_id: ObjectId,
|
||||
@@ -117,6 +188,7 @@ pub(crate) fn extract_form_xobject_text(
|
||||
parent_ctm: &[f32; 6],
|
||||
cmap_decisions: &mut CMapDecisionCache,
|
||||
style_cache: &mut FontStyleCache,
|
||||
budget: &mut FormWalkBudget,
|
||||
) -> Vec<TextItem> {
|
||||
extract_form_xobject_text_inner(
|
||||
doc,
|
||||
@@ -127,6 +199,7 @@ pub(crate) fn extract_form_xobject_text(
|
||||
cmap_decisions,
|
||||
style_cache,
|
||||
0,
|
||||
budget,
|
||||
)
|
||||
}
|
||||
|
||||
@@ -140,11 +213,14 @@ fn extract_form_xobject_text_inner(
|
||||
cmap_decisions: &mut CMapDecisionCache,
|
||||
style_cache: &mut FontStyleCache,
|
||||
depth: u8,
|
||||
budget: &mut FormWalkBudget,
|
||||
) -> Vec<TextItem> {
|
||||
use lopdf::content::Content;
|
||||
|
||||
let mut items = Vec::new();
|
||||
|
||||
if !budget.charge_invocation() {
|
||||
return items;
|
||||
}
|
||||
|
||||
// Get the Form XObject stream
|
||||
let Ok(Object::Stream(stream)) = doc.get_object(form_id) else {
|
||||
return items;
|
||||
@@ -156,8 +232,12 @@ fn extract_form_xobject_text_inner(
|
||||
Err(_) => stream.content.clone(),
|
||||
};
|
||||
|
||||
// Decode the content stream
|
||||
let Ok(content) = Content::decode(&content_data) else {
|
||||
// Decode the content stream. Cap before lopdf materializes the operator
|
||||
// vector — the walk budget cannot help if decode itself allocates first.
|
||||
let Ok(Some(content)) = super::content_decode::decode_content_bounded(
|
||||
&content_data,
|
||||
super::content_decode::MAX_PAGE_OPERATIONS,
|
||||
) else {
|
||||
return items;
|
||||
};
|
||||
|
||||
@@ -246,19 +326,55 @@ fn extract_form_xobject_text_inner(
|
||||
let mut current_font = String::new();
|
||||
let mut current_font_size: f32 = 12.0;
|
||||
let mut text_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
|
||||
// Text line matrix (TLM) — Td/TD/T* move relative to the start of the
|
||||
// current line, not to the position left by the last show operator.
|
||||
let mut line_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
|
||||
let mut text_leading: f32 = 0.0; // TL parameter (text-space units)
|
||||
let mut char_spacing: f32 = 0.0; // Tc parameter
|
||||
let mut word_spacing: f32 = 0.0; // Tw parameter
|
||||
let mut in_text_block = false;
|
||||
let mut fill_is_white = false;
|
||||
let mut ctm = base_ctm;
|
||||
let mut ctm_stack: Vec<[f32; 6]> = Vec::new();
|
||||
|
||||
// Text state (Tc/Tw/TL/Tf) and the fill colour are part of the graphics
|
||||
// state and must be saved/restored by q/Q alongside the CTM.
|
||||
#[derive(Clone)]
|
||||
struct GraphicsState {
|
||||
ctm: [f32; 6],
|
||||
char_spacing: f32,
|
||||
word_spacing: f32,
|
||||
text_leading: f32,
|
||||
current_font: String,
|
||||
current_font_size: f32,
|
||||
fill_is_white: bool,
|
||||
}
|
||||
let mut ctm_stack: Vec<GraphicsState> = Vec::new();
|
||||
|
||||
for op in &content.operations {
|
||||
if !budget.charge_operation() {
|
||||
break;
|
||||
}
|
||||
match op.operator.as_str() {
|
||||
"q" => {
|
||||
ctm_stack.push(ctm);
|
||||
ctm_stack.push(GraphicsState {
|
||||
ctm,
|
||||
char_spacing,
|
||||
word_spacing,
|
||||
text_leading,
|
||||
current_font: current_font.clone(),
|
||||
current_font_size,
|
||||
fill_is_white,
|
||||
});
|
||||
}
|
||||
"Q" => {
|
||||
if let Some(saved) = ctm_stack.pop() {
|
||||
ctm = saved;
|
||||
ctm = saved.ctm;
|
||||
char_spacing = saved.char_spacing;
|
||||
word_spacing = saved.word_spacing;
|
||||
text_leading = saved.text_leading;
|
||||
current_font = saved.current_font;
|
||||
current_font_size = saved.current_font_size;
|
||||
fill_is_white = saved.fill_is_white;
|
||||
}
|
||||
}
|
||||
"cm" => {
|
||||
@@ -276,7 +392,7 @@ fn extract_form_xobject_text_inner(
|
||||
let xobj_name = String::from_utf8_lossy(name).to_string();
|
||||
match form_xobjects.get(&xobj_name) {
|
||||
Some(XObjectType::Form(nested_id)) => {
|
||||
if depth < MAX_FORM_XOBJECT_DEPTH {
|
||||
if depth < MAX_FORM_XOBJECT_DEPTH && !budget.exhausted() {
|
||||
let nested_items = extract_form_xobject_text_inner(
|
||||
doc,
|
||||
*nested_id,
|
||||
@@ -286,6 +402,7 @@ fn extract_form_xobject_text_inner(
|
||||
cmap_decisions,
|
||||
style_cache,
|
||||
depth + 1,
|
||||
budget,
|
||||
);
|
||||
items.extend(nested_items);
|
||||
}
|
||||
@@ -321,6 +438,7 @@ fn extract_form_xobject_text_inner(
|
||||
"BT" => {
|
||||
in_text_block = true;
|
||||
text_matrix = [1.0, 0.0, 0.0, 1.0, 0.0, 0.0];
|
||||
line_matrix = text_matrix;
|
||||
}
|
||||
"ET" => {
|
||||
in_text_block = false;
|
||||
@@ -333,12 +451,33 @@ fn extract_form_xobject_text_inner(
|
||||
current_font_size = get_number(&op.operands[1]).unwrap_or(12.0);
|
||||
}
|
||||
}
|
||||
"TL" => {
|
||||
// Set text leading (used by T*, ', and ")
|
||||
if let Some(tl) = op.operands.first().and_then(get_number) {
|
||||
text_leading = tl;
|
||||
}
|
||||
}
|
||||
"Tc" => {
|
||||
if let Some(tc) = op.operands.first().and_then(get_number) {
|
||||
char_spacing = tc;
|
||||
}
|
||||
}
|
||||
"Tw" => {
|
||||
if let Some(tw) = op.operands.first().and_then(get_number) {
|
||||
word_spacing = tw;
|
||||
}
|
||||
}
|
||||
"Td" | "TD" => {
|
||||
// Move text position: TLM = T(tx,ty) x TLM; Tm = TLM
|
||||
if op.operands.len() >= 2 {
|
||||
let tx = get_number(&op.operands[0]).unwrap_or(0.0);
|
||||
let ty = get_number(&op.operands[1]).unwrap_or(0.0);
|
||||
text_matrix[4] += tx * text_matrix[0] + ty * text_matrix[2];
|
||||
text_matrix[5] += tx * text_matrix[1] + ty * text_matrix[3];
|
||||
line_matrix[4] += tx * line_matrix[0] + ty * line_matrix[2];
|
||||
line_matrix[5] += tx * line_matrix[1] + ty * line_matrix[3];
|
||||
text_matrix = line_matrix;
|
||||
if op.operator == "TD" {
|
||||
text_leading = -ty;
|
||||
}
|
||||
}
|
||||
}
|
||||
"Tm" => {
|
||||
@@ -347,8 +486,20 @@ fn extract_form_xobject_text_inner(
|
||||
text_matrix[i] =
|
||||
get_number(operand).unwrap_or(if i == 0 || i == 3 { 1.0 } else { 0.0 });
|
||||
}
|
||||
line_matrix = text_matrix;
|
||||
}
|
||||
}
|
||||
"T*" => {
|
||||
// Move to start of next line: equivalent to `0 -TL Td`
|
||||
let tl = if text_leading != 0.0 {
|
||||
text_leading
|
||||
} else {
|
||||
current_font_size * 1.2
|
||||
};
|
||||
line_matrix[4] += (-tl) * line_matrix[2];
|
||||
line_matrix[5] += (-tl) * line_matrix[3];
|
||||
text_matrix = line_matrix;
|
||||
}
|
||||
"g" => {
|
||||
if let Some(gray) = op.operands.first().and_then(get_number) {
|
||||
fill_is_white = gray > 0.95;
|
||||
@@ -384,17 +535,33 @@ fn extract_form_xobject_text_inner(
|
||||
_ => fill_is_white = false,
|
||||
}
|
||||
}
|
||||
"Tj" => {
|
||||
if in_text_block && !op.operands.is_empty() {
|
||||
"Tj" | "'" | "\"" => {
|
||||
// `'` = move to next line then show; `"` = set word/char spacing,
|
||||
// move to next line, then show (string is the last operand).
|
||||
if op.operator != "Tj" {
|
||||
if op.operator == "\"" && op.operands.len() >= 3 {
|
||||
word_spacing = get_number(&op.operands[0]).unwrap_or(word_spacing);
|
||||
char_spacing = get_number(&op.operands[1]).unwrap_or(char_spacing);
|
||||
}
|
||||
let tl = if text_leading != 0.0 {
|
||||
text_leading
|
||||
} else {
|
||||
current_font_size * 1.2
|
||||
};
|
||||
line_matrix[4] += (-tl) * line_matrix[2];
|
||||
line_matrix[5] += (-tl) * line_matrix[3];
|
||||
text_matrix = line_matrix;
|
||||
}
|
||||
if let (true, Some(show_operand)) = (in_text_block, op.operands.last()) {
|
||||
if fill_is_white {
|
||||
if let Some(font_info) = font_widths.get(¤t_font) {
|
||||
if let Some(raw_bytes) = get_operand_bytes(&op.operands[0]) {
|
||||
if let Some(raw_bytes) = get_operand_bytes(show_operand) {
|
||||
let w_ts = compute_string_width_ts(
|
||||
raw_bytes,
|
||||
font_info,
|
||||
current_font_size,
|
||||
0.0,
|
||||
0.0,
|
||||
char_spacing,
|
||||
word_spacing,
|
||||
);
|
||||
text_matrix[4] += w_ts * text_matrix[0];
|
||||
text_matrix[5] += w_ts * text_matrix[1];
|
||||
@@ -403,7 +570,7 @@ fn extract_form_xobject_text_inner(
|
||||
continue;
|
||||
}
|
||||
if let Some(text) = extract_text_from_operand(
|
||||
&op.operands[0],
|
||||
show_operand,
|
||||
¤t_font,
|
||||
font_base_names.get(¤t_font).map(|s| s.as_str()),
|
||||
font_cmaps,
|
||||
@@ -419,13 +586,13 @@ fn extract_form_xobject_text_inner(
|
||||
* type3_scales.get(¤t_font).copied().unwrap_or(1.0);
|
||||
let (x, y) = (combined[4], combined[5]);
|
||||
let width = if let Some(font_info) = font_widths.get(¤t_font) {
|
||||
if let Some(raw_bytes) = get_operand_bytes(&op.operands[0]) {
|
||||
if let Some(raw_bytes) = get_operand_bytes(show_operand) {
|
||||
let w_ts = compute_string_width_ts(
|
||||
raw_bytes,
|
||||
font_info,
|
||||
current_font_size,
|
||||
0.0,
|
||||
0.0,
|
||||
char_spacing,
|
||||
word_spacing,
|
||||
);
|
||||
text_matrix[4] += w_ts * text_matrix[0];
|
||||
text_matrix[5] += w_ts * text_matrix[1];
|
||||
@@ -548,8 +715,8 @@ fn extract_form_xobject_text_inner(
|
||||
raw_bytes,
|
||||
fi,
|
||||
current_font_size,
|
||||
0.0,
|
||||
0.0,
|
||||
char_spacing,
|
||||
word_spacing,
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -683,3 +850,415 @@ pub(crate) fn get_form_fonts<'a>(
|
||||
|
||||
fonts
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::extractor::content_stream::extract_page_text_items;
|
||||
use lopdf::{dictionary, Dictionary, Stream};
|
||||
|
||||
/// Build an acyclic Form XObject DAG: `levels` form objects, each non-leaf
|
||||
/// invoking the next form `branches` times. The leaf draws a single `(X)`.
|
||||
/// Returns `(doc, root_form_id)`.
|
||||
fn form_dag(branches: usize, levels: usize) -> (Document, ObjectId) {
|
||||
assert!(levels >= 2);
|
||||
let mut doc = Document::new();
|
||||
let font_id = doc.add_object(dictionary! {
|
||||
"Type" => "Font",
|
||||
"Subtype" => "Type1",
|
||||
"BaseFont" => "Helvetica",
|
||||
});
|
||||
let ids: Vec<ObjectId> = (0..levels).map(|_| doc.new_object_id()).collect();
|
||||
for level in 0..levels {
|
||||
let stream = if level + 1 == levels {
|
||||
Stream::new(
|
||||
dictionary! {
|
||||
"Type" => "XObject",
|
||||
"Subtype" => "Form",
|
||||
"BBox" => vec![0.into(), 0.into(), 100.into(), 100.into()],
|
||||
"Resources" => dictionary! {
|
||||
"Font" => dictionary! {
|
||||
"F1" => Object::Reference(font_id),
|
||||
},
|
||||
},
|
||||
},
|
||||
b"BT /F1 10 Tf 10 10 Td (X) Tj ET\n".to_vec(),
|
||||
)
|
||||
} else {
|
||||
let next_name = format!("Fm{}", level + 1);
|
||||
let content = format!("/{next_name} Do\n").repeat(branches);
|
||||
let mut xobjects = Dictionary::new();
|
||||
xobjects.set(next_name, Object::Reference(ids[level + 1]));
|
||||
let mut resources = Dictionary::new();
|
||||
resources.set("XObject", Object::Dictionary(xobjects));
|
||||
let mut dict = dictionary! {
|
||||
"Type" => "XObject",
|
||||
"Subtype" => "Form",
|
||||
"BBox" => vec![0.into(), 0.into(), 100.into(), 100.into()],
|
||||
};
|
||||
dict.set("Resources", Object::Dictionary(resources));
|
||||
Stream::new(dict, content.into_bytes())
|
||||
};
|
||||
doc.set_object(ids[level], Object::Stream(stream));
|
||||
}
|
||||
(doc, ids[0])
|
||||
}
|
||||
|
||||
fn page_invoking_form(mut doc: Document, form_id: ObjectId) -> (Document, ObjectId) {
|
||||
let content_id = doc.add_object(Object::Stream(Stream::new(
|
||||
dictionary! {},
|
||||
b"/Fm0 Do\n".to_vec(),
|
||||
)));
|
||||
let page_id = doc.add_object(dictionary! {
|
||||
"Type" => "Page",
|
||||
"Contents" => Object::Reference(content_id),
|
||||
"Resources" => dictionary! {
|
||||
"XObject" => dictionary! {
|
||||
"Fm0" => Object::Reference(form_id),
|
||||
},
|
||||
},
|
||||
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
|
||||
});
|
||||
let pages_id = doc.add_object(dictionary! {
|
||||
"Type" => "Pages",
|
||||
"Count" => Object::Integer(1),
|
||||
"Kids" => vec![Object::Reference(page_id)],
|
||||
});
|
||||
let catalog_id = doc.add_object(dictionary! {
|
||||
"Type" => "Catalog",
|
||||
"Pages" => Object::Reference(pages_id),
|
||||
});
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
(doc, page_id)
|
||||
}
|
||||
|
||||
fn extract_form(
|
||||
doc: &Document,
|
||||
form_id: ObjectId,
|
||||
budget: &mut FormWalkBudget,
|
||||
) -> Vec<TextItem> {
|
||||
extract_form_xobject_text(
|
||||
doc,
|
||||
form_id,
|
||||
1,
|
||||
&FontCMaps::from_doc(doc),
|
||||
&[1.0, 0.0, 0.0, 1.0, 0.0, 0.0],
|
||||
&mut CMapDecisionCache::new(),
|
||||
&mut FontStyleCache::new(),
|
||||
budget,
|
||||
)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn nested_form_still_extracts_leaf_text() {
|
||||
let (doc, root) = form_dag(1, 3);
|
||||
let items = extract_form(&doc, root, &mut FormWalkBudget::new());
|
||||
assert_eq!(items.len(), 1);
|
||||
assert_eq!(items[0].text, "X");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn acyclic_form_dag_within_budget_keeps_all_leaves() {
|
||||
// 4 sibling invocations across 4 nested levels → 4^4 leaf drawings.
|
||||
// Default budgets are far above 256, so legitimate nesting is intact.
|
||||
let (doc, root) = form_dag(4, 5);
|
||||
let items = extract_form(&doc, root, &mut FormWalkBudget::new());
|
||||
assert_eq!(items.len(), 4usize.pow(4));
|
||||
assert!(items.iter().all(|item| item.text == "X"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn acyclic_form_dag_stops_at_invocation_budget() {
|
||||
// Same DAG as above would draw 256 leaves; a tiny invocation cap must
|
||||
// stop expansion rather than walking the full tree.
|
||||
let (doc, root) = form_dag(4, 5);
|
||||
let mut budget = FormWalkBudget::with_limits(20, MAX_FORM_XOBJECT_OPERATIONS);
|
||||
let items = extract_form(&doc, root, &mut budget);
|
||||
assert!(
|
||||
items.len() < 4usize.pow(4),
|
||||
"invocation budget must truncate DAG expansion; got {} items",
|
||||
items.len()
|
||||
);
|
||||
assert!(budget.was_truncated());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn form_operations_stop_at_budget() {
|
||||
let mut doc = Document::new();
|
||||
let font_id = doc.add_object(dictionary! {
|
||||
"Type" => "Font",
|
||||
"Subtype" => "Type1",
|
||||
"BaseFont" => "Helvetica",
|
||||
});
|
||||
let mut content = b"q Q\n".repeat(50);
|
||||
content.extend_from_slice(b"BT /F1 10 Tf 10 10 Td (X) Tj ET\n");
|
||||
let form_id = doc.add_object(Object::Stream(Stream::new(
|
||||
dictionary! {
|
||||
"Type" => "XObject",
|
||||
"Subtype" => "Form",
|
||||
"BBox" => vec![0.into(), 0.into(), 100.into(), 100.into()],
|
||||
"Resources" => dictionary! {
|
||||
"Font" => dictionary! {
|
||||
"F1" => Object::Reference(font_id),
|
||||
},
|
||||
},
|
||||
},
|
||||
content,
|
||||
)));
|
||||
|
||||
let mut budget = FormWalkBudget::with_limits(MAX_FORM_XOBJECT_INVOCATIONS, 10);
|
||||
let items = extract_form(&doc, form_id, &mut budget);
|
||||
assert!(
|
||||
items.is_empty(),
|
||||
"operation budget must stop before the trailing text show"
|
||||
);
|
||||
assert!(budget.was_truncated());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn page_level_form_dag_stays_within_production_budget() {
|
||||
// A page-level `/Do` of an 8-wide, 6-level Form DAG would expand to
|
||||
// 8^5 = 32_768 leaf drawings without a budget. The production
|
||||
// invocation cap must keep extraction bounded.
|
||||
let (doc, root) = form_dag(8, 6);
|
||||
let (doc, page_id) = page_invoking_form(doc, root);
|
||||
|
||||
let font_cmaps = FontCMaps::from_doc(&doc);
|
||||
let ((items, _, _), _, _, _) = extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
1,
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut FontStyleCache::new(),
|
||||
&mut FormWalkBudget::new(),
|
||||
)
|
||||
.unwrap();
|
||||
assert!(
|
||||
items.len() <= MAX_FORM_XOBJECT_INVOCATIONS,
|
||||
"page-level Form expansion must stay within the invocation cap; got {}",
|
||||
items.len()
|
||||
);
|
||||
assert!(
|
||||
!items.is_empty(),
|
||||
"budget must still allow some nested form text through"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn shared_form_budget_spans_two_extraction_passes() {
|
||||
// The invisible-layer retry calls extract_page_text_items twice for
|
||||
// the same page; both passes must share one budget.
|
||||
let (doc, root) = form_dag(1, 2);
|
||||
let (doc, page_id) = page_invoking_form(doc, root);
|
||||
let font_cmaps = FontCMaps::from_doc(&doc);
|
||||
// Root + leaf = 2 invocations on the first pass.
|
||||
let mut budget = FormWalkBudget::with_limits(2, MAX_FORM_XOBJECT_OPERATIONS);
|
||||
let ((first, _, _), _, _, _) = extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
1,
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut FontStyleCache::new(),
|
||||
&mut budget,
|
||||
)
|
||||
.unwrap();
|
||||
assert_eq!(first.iter().filter(|item| item.text == "X").count(), 1);
|
||||
assert!(!budget.was_truncated());
|
||||
|
||||
let ((second, _, _), _, _, _) = extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
1,
|
||||
&font_cmaps,
|
||||
true,
|
||||
&mut FontStyleCache::new(),
|
||||
&mut budget,
|
||||
)
|
||||
.unwrap();
|
||||
assert!(
|
||||
second.iter().all(|item| item.text != "X"),
|
||||
"second pass must not get a fresh invocation budget"
|
||||
);
|
||||
assert!(budget.was_truncated());
|
||||
}
|
||||
|
||||
/// Build a document whose page draws *all* of its content through a single
|
||||
/// Form XObject — the shape emitted by print-to-PDF producers like PDFlib,
|
||||
/// where the page stream itself is only `q /X1 Do Q`.
|
||||
fn doc_with_form_content(form_content: &[u8]) -> (Document, ObjectId) {
|
||||
let mut doc = Document::new();
|
||||
let widths: Vec<Object> = (0..=255).map(|_| 600.into()).collect();
|
||||
let font_id = doc.add_object(dictionary! {
|
||||
"Type" => "Font",
|
||||
"Subtype" => "Type1",
|
||||
"BaseFont" => "Helvetica",
|
||||
"FirstChar" => 0,
|
||||
"LastChar" => 255,
|
||||
"Widths" => Object::Array(widths),
|
||||
});
|
||||
let form_id = doc.add_object(Object::Stream(Stream::new(
|
||||
dictionary! {
|
||||
"Type" => "XObject",
|
||||
"Subtype" => "Form",
|
||||
"BBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
|
||||
"Resources" => dictionary! {
|
||||
"Font" => dictionary! { "F1" => Object::Reference(font_id) },
|
||||
},
|
||||
},
|
||||
form_content.to_vec(),
|
||||
)));
|
||||
let content_id = doc.add_object(Object::Stream(Stream::new(
|
||||
dictionary! {},
|
||||
b"q /X1 Do Q".to_vec(),
|
||||
)));
|
||||
let page_id = doc.add_object(dictionary! {
|
||||
"Type" => "Page",
|
||||
"Contents" => Object::Reference(content_id),
|
||||
"Resources" => dictionary! {
|
||||
"XObject" => dictionary! { "X1" => Object::Reference(form_id) },
|
||||
},
|
||||
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
|
||||
});
|
||||
let pages_id = doc.add_object(dictionary! {
|
||||
"Type" => "Pages",
|
||||
"Count" => Object::Integer(1),
|
||||
"Kids" => vec![Object::Reference(page_id)],
|
||||
});
|
||||
let catalog_id = doc.add_object(dictionary! {
|
||||
"Type" => "Catalog",
|
||||
"Pages" => Object::Reference(pages_id),
|
||||
});
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
(doc, page_id)
|
||||
}
|
||||
|
||||
fn form_items(form_content: &[u8]) -> Vec<TextItem> {
|
||||
let (doc, page_id) = doc_with_form_content(form_content);
|
||||
let font_cmaps = FontCMaps::from_doc(&doc);
|
||||
let ((items, _, _), _, _, _) = extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
1,
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut FontStyleCache::new(),
|
||||
&mut FormWalkBudget::new(),
|
||||
)
|
||||
.unwrap();
|
||||
items
|
||||
}
|
||||
|
||||
fn find<'a>(items: &'a [TextItem], text: &str) -> &'a TextItem {
|
||||
items
|
||||
.iter()
|
||||
.find(|item| item.text == text)
|
||||
.unwrap_or_else(|| {
|
||||
let found: Vec<&String> = items.iter().map(|i| &i.text).collect();
|
||||
panic!("no item {text:?} in {found:?}")
|
||||
})
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn t_star_inside_form_moves_to_next_line() {
|
||||
// T* was previously unhandled inside Form XObjects, so every line after
|
||||
// the first piled onto the preceding baseline and drifted right.
|
||||
let items =
|
||||
form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm (first) Tj T* (second) Tj ET");
|
||||
|
||||
let first = find(&items, "first");
|
||||
let second = find(&items, "second");
|
||||
assert!((first.y - 700.0).abs() < 0.1, "first y = {}", first.y);
|
||||
assert!((second.y - 688.0).abs() < 0.1, "second y = {}", second.y);
|
||||
assert!((second.x - 100.0).abs() < 0.1, "second x = {}", second.x);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn td_inside_form_is_relative_to_line_start_not_shown_text() {
|
||||
// Td moves relative to the text *line* matrix. Applying it to the
|
||||
// matrix already advanced by Tj marched each line off the right edge.
|
||||
let items = form_items(b"BT /F1 12 Tf 1 0 0 1 100 700 Tm (AAAAA) Tj 0 -12 Td (B) Tj ET");
|
||||
|
||||
let b = find(&items, "B");
|
||||
assert!((b.x - 100.0).abs() < 0.1, "B x = {} (expected 100)", b.x);
|
||||
assert!((b.y - 688.0).abs() < 0.1, "B y = {}", b.y);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn td_inside_form_sets_leading_for_later_t_star() {
|
||||
// `TD` sets the leading to -ty as a side effect; a following T* must
|
||||
// reuse it.
|
||||
let items = form_items(
|
||||
b"BT /F1 12 Tf 1 0 0 1 100 700 Tm (one) Tj 0 -15 TD (two) Tj T* (three) Tj ET",
|
||||
);
|
||||
|
||||
assert!((find(&items, "two").y - 685.0).abs() < 0.1);
|
||||
let three = find(&items, "three");
|
||||
assert!((three.y - 670.0).abs() < 0.1, "three y = {}", three.y);
|
||||
assert!((three.x - 100.0).abs() < 0.1, "three x = {}", three.x);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn quote_operator_inside_form_moves_to_next_line() {
|
||||
let items = form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm (first) Tj (second) ' ET");
|
||||
|
||||
let second = find(&items, "second");
|
||||
assert!((second.y - 688.0).abs() < 0.1, "second y = {}", second.y);
|
||||
assert!((second.x - 100.0).abs() < 0.1, "second x = {}", second.x);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn double_quote_operator_inside_form_sets_spacing_and_moves() {
|
||||
// `aw ac (string) "` — set word spacing and char spacing, then T* and show.
|
||||
let items =
|
||||
form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm (first) Tj 0 0 (second) \" ET");
|
||||
|
||||
let second = find(&items, "second");
|
||||
assert!((second.y - 688.0).abs() < 0.1, "second y = {}", second.y);
|
||||
assert!((second.x - 100.0).abs() < 0.1, "second x = {}", second.x);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn char_spacing_inside_form_widens_advance() {
|
||||
// Tc was hardcoded to 0 in the form parser, so advance widths drifted.
|
||||
// 2 glyphs x 600/1000 x 12pt = 14.4, plus 2 x Tc(2.0) = 18.4.
|
||||
let items = form_items(b"BT /F1 12 Tf 1 0 0 1 100 700 Tm 2 Tc (AB) Tj ET");
|
||||
|
||||
let ab = find(&items, "AB");
|
||||
assert!((ab.width - 18.4).abs() < 0.1, "AB width = {}", ab.width);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn q_restores_fill_colour_inside_form() {
|
||||
// A white fill set inside q/Q must not leak past the Q — otherwise the
|
||||
// following black text is treated as invisible and dropped entirely.
|
||||
let items = form_items(
|
||||
b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm q 1 g (hidden) Tj Q T* (visible) Tj ET",
|
||||
);
|
||||
|
||||
assert!(
|
||||
items.iter().any(|item| item.text == "visible"),
|
||||
"text after Q was dropped: {:?}",
|
||||
items.iter().map(|i| &i.text).collect::<Vec<_>>()
|
||||
);
|
||||
assert!(
|
||||
!items.iter().any(|item| item.text == "hidden"),
|
||||
"white-filled text should still be suppressed"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn q_restores_text_state_inside_form() {
|
||||
// Tc/TL live in the graphics state; `Q` must roll them back.
|
||||
let items =
|
||||
form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm q 30 TL (a) Tj Q T* (b) Tj ET");
|
||||
|
||||
let b = find(&items, "b");
|
||||
assert!(
|
||||
(b.y - 688.0).abs() < 0.1,
|
||||
"b y = {} (leading should restore to 12)",
|
||||
b.y
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
+40
-3
@@ -4566,9 +4566,13 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
|
||||
}
|
||||
}
|
||||
|
||||
// Try to parse uniXXXX format
|
||||
if name.starts_with("uni") && name.len() >= 7 {
|
||||
if let Ok(code) = u32::from_str_radix(&name[3..7], 16) {
|
||||
// Try to parse uniXXXX format.
|
||||
// Use `get` rather than a byte-length check + slice: `name` can contain
|
||||
// non-ASCII bytes (e.g. U+FFFD from lossy UTF-8 decoding of an attacker
|
||||
// controlled /Differences name), so byte index 7 may not be a char
|
||||
// boundary and `&name[3..7]` would panic.
|
||||
if let Some(hex) = name.strip_prefix("uni").and_then(|rest| rest.get(..4)) {
|
||||
if let Ok(code) = u32::from_str_radix(hex, 16) {
|
||||
// Strip PUA F000 offset: uniF0XX → U+00XX (Windows Symbol encoding convention)
|
||||
let code = if (0xF000..=0xF0FF).contains(&code) {
|
||||
code - 0xF000
|
||||
@@ -4588,3 +4592,36 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
|
||||
|
||||
None
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn uni_hex_parsing() {
|
||||
assert_eq!(glyph_to_char("uni0041"), Some('A'));
|
||||
assert_eq!(glyph_to_char("uni00e9"), Some('\u{00e9}'));
|
||||
// PUA F0xx symbol-encoding offset is stripped.
|
||||
assert_eq!(glyph_to_char("uniF041"), Some('A'));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn u_hex_parsing() {
|
||||
assert_eq!(glyph_to_char("u0041"), Some('A'));
|
||||
assert_eq!(glyph_to_char("u1F600"), Some('\u{1F600}'));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_ascii_uni_name_does_not_panic() {
|
||||
// A crafted /Differences name like `/uni#80#80#80#80` decodes via
|
||||
// from_utf8_lossy into "uni" followed by four U+FFFD replacements.
|
||||
// Byte index 7 lands mid-character, so a naive `&name[3..7]` slice
|
||||
// would panic. It must be handled gracefully instead.
|
||||
let crafted = format!("uni{0}{0}{0}{0}", '\u{FFFD}');
|
||||
assert_eq!(glyph_to_char(&crafted), None);
|
||||
|
||||
// Assorted non-ASCII bytes right after the "uni" prefix.
|
||||
assert_eq!(glyph_to_char("uni\u{FFFD}bc"), None);
|
||||
assert_eq!(glyph_to_char("uni\u{00e9}00"), None);
|
||||
}
|
||||
}
|
||||
|
||||
+173
-24
@@ -657,6 +657,80 @@ pub fn extract_pages_markdown<P: AsRef<Path>>(
|
||||
extract_pages_markdown_mem(&buffer, pages)
|
||||
}
|
||||
|
||||
// =========================================================================
|
||||
// Structure-tree element extraction (tagged PDFs)
|
||||
// =========================================================================
|
||||
|
||||
/// One structure-tree element reference from a tagged PDF, resolved to a
|
||||
/// page and Marked Content ID.
|
||||
///
|
||||
/// Join `(page, mcid)` against [`TextItem::page`] / [`TextItem::mcid`] from
|
||||
/// [`extract_text_with_positions`] to attach semantic roles (heading levels,
|
||||
/// paragraphs, table cells, …) to extracted text.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct StructureElement {
|
||||
/// 1-indexed page number (matches [`TextItem::page`]).
|
||||
pub page: u32,
|
||||
/// Marked Content ID from the page's content stream (matches
|
||||
/// [`TextItem::mcid`]).
|
||||
pub mcid: i64,
|
||||
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
|
||||
/// Custom tags are resolved through the document's `/RoleMap`; tags
|
||||
/// with no standard mapping are returned verbatim.
|
||||
pub role: String,
|
||||
}
|
||||
|
||||
/// Extract structure-tree element references from a tagged PDF in memory.
|
||||
///
|
||||
/// Parses `/StructTreeRoot` (when present) and returns one entry per
|
||||
/// marked-content reference, resolved to its 1-indexed page, MCID, and
|
||||
/// structure type name. Returns an empty list when the PDF is not tagged.
|
||||
///
|
||||
/// Pass `Some(&[...])` with 1-indexed page numbers (matching
|
||||
/// [`TextItem::page`]) to restrict output to those pages; pass `None` for
|
||||
/// the whole document. Entries are sorted by `(page, mcid)`.
|
||||
pub fn extract_structure_elements_mem(
|
||||
buffer: &[u8],
|
||||
pages: Option<&[u32]>,
|
||||
) -> Result<Vec<StructureElement>, PdfError> {
|
||||
validate_pdf_bytes(buffer)?;
|
||||
let (doc, _page_count) = load_document_from_mem(buffer)?;
|
||||
let Some(tree) = structure_tree::StructTree::from_doc(&doc) else {
|
||||
return Ok(Vec::new());
|
||||
};
|
||||
let page_ids = doc.get_pages();
|
||||
let roles = tree.mcid_to_roles(&page_ids);
|
||||
|
||||
let page_filter: Option<HashSet<u32>> = pages.map(|p| p.iter().copied().collect());
|
||||
let mut elements: Vec<StructureElement> = roles
|
||||
.into_iter()
|
||||
.filter(|(page, _)| page_filter.as_ref().is_none_or(|f| f.contains(page)))
|
||||
.flat_map(|(page, mcids)| {
|
||||
mcids.into_iter().map(move |(mcid, role)| StructureElement {
|
||||
page,
|
||||
mcid,
|
||||
role: role.name().to_string(),
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
elements.sort_unstable_by_key(|e| (e.page, e.mcid));
|
||||
Ok(elements)
|
||||
}
|
||||
|
||||
/// Path-based wrapper for [`extract_structure_elements_mem`].
|
||||
///
|
||||
/// Reads the PDF from disk and extracts structure-tree element references.
|
||||
/// Pass `None` for `pages` to return the whole document, or `Some(&[...])`
|
||||
/// to restrict to specific 1-indexed pages.
|
||||
pub fn extract_structure_elements<P: AsRef<Path>>(
|
||||
path: P,
|
||||
pages: Option<&[u32]>,
|
||||
) -> Result<Vec<StructureElement>, PdfError> {
|
||||
validate_pdf_file(&path)?;
|
||||
let buffer = std::fs::read(path.as_ref())?;
|
||||
extract_structure_elements_mem(&buffer, pages)
|
||||
}
|
||||
|
||||
// =========================================================================
|
||||
// Region-based text extraction (for hybrid OCR pipelines)
|
||||
// =========================================================================
|
||||
@@ -683,6 +757,23 @@ pub struct PageRegionResult {
|
||||
pub regions: Vec<RegionText>,
|
||||
}
|
||||
|
||||
/// Minimum alphanumeric mass an invisible (Tr 3) text layer must carry for
|
||||
/// the OCR-layer fallback in [`extract_text_in_regions_mem`] to adopt it. A
|
||||
/// real OCR layer carries far more; a stray watermark or artifact does not.
|
||||
const OCR_LAYER_MIN_ALNUM: usize = 40;
|
||||
|
||||
/// Alphanumeric mass of extracted items, ignoring raster placeholders.
|
||||
/// `[Image: ...]` items (ItemType::Image) are synthesized for image
|
||||
/// XObjects — they mark that pixels exist, not that text was read, so they
|
||||
/// must not count as coverage.
|
||||
fn non_placeholder_alnum(items: &[TextItem]) -> usize {
|
||||
items
|
||||
.iter()
|
||||
.filter(|it| !matches!(it.item_type, types::ItemType::Image))
|
||||
.map(|it| it.text.chars().filter(|c| c.is_alphanumeric()).count())
|
||||
.sum()
|
||||
}
|
||||
|
||||
/// Extract text within bounding-box regions from a PDF in memory.
|
||||
///
|
||||
/// This is designed for hybrid OCR pipelines: a layout model detects regions
|
||||
@@ -735,8 +826,11 @@ pub fn extract_text_in_regions_mem(
|
||||
let height = get_page_height(&doc, page_id).unwrap_or(792.0);
|
||||
page_heights.insert(*page_num, height);
|
||||
|
||||
// Extract text items for this page
|
||||
let ((mut items, _rects, _lines), has_gid, coords_rotated) =
|
||||
// Extract text items for this page. The Form XObject budget is shared
|
||||
// with the invisible-layer retry below so one page cannot consume two
|
||||
// full expansion budgets.
|
||||
let mut form_budget = extractor::FormWalkBudget::new();
|
||||
let ((mut items, _rects, _lines), mut has_gid, mut coords_rotated, skipped_invisible) =
|
||||
extractor::content_stream::extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
@@ -744,7 +838,54 @@ pub fn extract_text_in_regions_mem(
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut style_cache,
|
||||
&mut form_budget,
|
||||
)?;
|
||||
// OCR-layer fallback: scanned pages often carry their text as an
|
||||
// invisible (Tr 3) layer behind the page raster. The visible-only
|
||||
// pass sees nothing there but `[Image: ...]` placeholders, so every
|
||||
// region on the page reports needs_ocr even though the exact text is
|
||||
// embedded in the PDF — and this extractor then disagrees with the
|
||||
// markdown path, which already retries Mixed PDFs with the invisible
|
||||
// layer included. Retry page-scoped, and only when (a) the first
|
||||
// pass actually SKIPPED invisible text — blank pages and image-only
|
||||
// scans without an OCR layer must not pay a second content-stream
|
||||
// parse (review catch) — and (b) the page has NO visible text item
|
||||
// at all (punctuation counts, whitespace-only artifacts don't): an
|
||||
// invisible OCR layer transcribes the raster, so any visible glyph
|
||||
// has an invisible twin there and adoption would duplicate it
|
||||
// (review catches — strict gate, no fuzzy dedupe). Adopt the retry
|
||||
// only when it contributes real, non-garbage text.
|
||||
let has_visible_text = items.iter().any(|it| {
|
||||
!matches!(it.item_type, types::ItemType::Image) && !it.text.trim().is_empty()
|
||||
});
|
||||
if skipped_invisible && !has_visible_text {
|
||||
if let Ok(((inv_items, _inv_rects, _inv_lines), inv_gid, inv_rotated, _)) =
|
||||
extractor::content_stream::extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
*page_num,
|
||||
&font_cmaps,
|
||||
true,
|
||||
&mut style_cache,
|
||||
&mut form_budget,
|
||||
)
|
||||
{
|
||||
let inv_alnum = non_placeholder_alnum(&inv_items);
|
||||
// Judge the WHOLE recovered layer, not a prefix — a broken
|
||||
// OCR layer can hide its garbage past any fixed sample size
|
||||
// (review catch).
|
||||
let sample: String = inv_items
|
||||
.iter()
|
||||
.filter(|it| !matches!(it.item_type, types::ItemType::Image))
|
||||
.map(|it| it.text.as_str())
|
||||
.collect();
|
||||
if inv_alnum >= OCR_LAYER_MIN_ALNUM && !is_garbage_text(&sample) {
|
||||
items = inv_items;
|
||||
has_gid = inv_gid;
|
||||
coords_rotated = inv_rotated;
|
||||
}
|
||||
}
|
||||
}
|
||||
let threshold = text_utils::fix_letterspaced_items(&mut items);
|
||||
if threshold > 0.10 {
|
||||
page_thresholds.insert(*page_num, threshold);
|
||||
@@ -901,7 +1042,7 @@ pub fn extract_tables_in_regions_mem(
|
||||
let height = get_page_height(&doc, page_id).unwrap_or(792.0);
|
||||
page_heights.insert(*page_num, height);
|
||||
|
||||
let ((mut items, rects, lines), has_gid, coords_rotated) =
|
||||
let ((mut items, rects, lines), has_gid, coords_rotated, _skipped_invisible) =
|
||||
extractor::content_stream::extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
@@ -909,6 +1050,7 @@ pub fn extract_tables_in_regions_mem(
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut style_cache,
|
||||
&mut extractor::FormWalkBudget::new(),
|
||||
)?;
|
||||
let threshold = text_utils::fix_letterspaced_items(&mut items);
|
||||
if threshold > 0.10 {
|
||||
@@ -1212,7 +1354,7 @@ pub fn detect_vector_grid_in_region_mem(
|
||||
let needed_pages = HashSet::from([page_1idx]);
|
||||
let font_cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed_pages));
|
||||
let page_h = get_page_height(&doc, page_id).unwrap_or(792.0);
|
||||
let ((mut items, rects, lines), _has_gid, coords_rotated) =
|
||||
let ((mut items, rects, lines), _has_gid, coords_rotated, _skipped_invisible) =
|
||||
extractor::content_stream::extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
@@ -1220,6 +1362,7 @@ pub fn detect_vector_grid_in_region_mem(
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut extractor::FontStyleCache::new(),
|
||||
&mut extractor::FormWalkBudget::new(),
|
||||
)?;
|
||||
text_utils::fix_letterspaced_items(&mut items);
|
||||
|
||||
@@ -1406,15 +1549,17 @@ mod vector_grid_tests {
|
||||
let &page_id = pages.get(&1).unwrap();
|
||||
let needed: HashSet<u32> = HashSet::from([1]);
|
||||
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
|
||||
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
1,
|
||||
&cmaps,
|
||||
false,
|
||||
&mut crate::extractor::FontStyleCache::new(),
|
||||
)
|
||||
.unwrap();
|
||||
let ((items, rects, _lines), _has_gid, _rotated, _skipped_invisible) =
|
||||
extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
1,
|
||||
&cmaps,
|
||||
false,
|
||||
&mut crate::extractor::FontStyleCache::new(),
|
||||
&mut crate::extractor::FormWalkBudget::new(),
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, 1);
|
||||
assert_eq!(rect_tables.len(), 1, "expected one rect-detected table");
|
||||
@@ -1448,15 +1593,17 @@ mod vector_grid_tests {
|
||||
let &page_id = pages.get(&page_num).unwrap();
|
||||
let needed: HashSet<u32> = HashSet::from([page_num]);
|
||||
let cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
|
||||
let ((items, rects, _lines), _has_gid, _rotated) = extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
page_num,
|
||||
&cmaps,
|
||||
false,
|
||||
&mut crate::extractor::FontStyleCache::new(),
|
||||
)
|
||||
.unwrap();
|
||||
let ((items, rects, _lines), _has_gid, _rotated, _skipped_invisible) =
|
||||
extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
page_num,
|
||||
&cmaps,
|
||||
false,
|
||||
&mut crate::extractor::FontStyleCache::new(),
|
||||
&mut crate::extractor::FormWalkBudget::new(),
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, page_num);
|
||||
rect_tables
|
||||
@@ -2184,7 +2331,7 @@ pub fn extract_tables_with_structure_cells_mem(
|
||||
let height = get_page_height(&doc, page_id).unwrap_or(792.0);
|
||||
page_heights.insert(*page_num, height);
|
||||
|
||||
let ((mut items, _rects, _lines), _has_gid, coords_rotated) =
|
||||
let ((mut items, _rects, _lines), _has_gid, coords_rotated, _skipped_invisible) =
|
||||
extractor::content_stream::extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
@@ -2192,6 +2339,7 @@ pub fn extract_tables_with_structure_cells_mem(
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut style_cache,
|
||||
&mut extractor::FormWalkBudget::new(),
|
||||
)?;
|
||||
let threshold = text_utils::fix_letterspaced_items(&mut items);
|
||||
if threshold > 0.10 {
|
||||
@@ -2986,7 +3134,7 @@ fn detect_tsr_quality_issue(
|
||||
let mut needed: HashSet<u32> = HashSet::new();
|
||||
needed.insert(page_1idx);
|
||||
let font_cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed));
|
||||
let ((mut items, _rects, _lines), _has_gid, coords_rotated) =
|
||||
let ((mut items, _rects, _lines), _has_gid, coords_rotated, _skipped_invisible) =
|
||||
extractor::content_stream::extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
@@ -2994,6 +3142,7 @@ fn detect_tsr_quality_issue(
|
||||
&font_cmaps,
|
||||
false,
|
||||
&mut extractor::FontStyleCache::new(),
|
||||
&mut extractor::FormWalkBudget::new(),
|
||||
)?;
|
||||
let adaptive_threshold = text_utils::fix_letterspaced_items(&mut items);
|
||||
let coords = if coords_rotated {
|
||||
|
||||
@@ -171,6 +171,74 @@ pub(crate) fn is_toc_marker_heading(text: &str) -> bool {
|
||||
/// equation and absent from name-plus-number headings. A bare trailing colon
|
||||
/// is NOT a fragment signal either: real headings frequently end with colons
|
||||
/// ("Procedure:", "Steps for Using the Microscope:").
|
||||
/// True when the line opens with a section number ("3.", "2.1.4", "IV)").
|
||||
///
|
||||
/// Mirrors the acceptance of `heading::parse_numbering` rather than the
|
||||
/// stricter `convert::starts_with_section_number`, which deliberately
|
||||
/// requires two components because it bypasses isolation checks. Here a
|
||||
/// single "1." counts: numbering is independent evidence of a heading, and
|
||||
/// `heading.rs` applies its numbered-prefix allowance *after* consulting
|
||||
/// `is_heading_fragment`, so without this exemption a numbered
|
||||
/// sentence-case heading would be vetoed before that allowance can run.
|
||||
fn starts_with_numbering_prefix(t: &str) -> bool {
|
||||
let Some(first) = t.split_whitespace().next() else {
|
||||
return false;
|
||||
};
|
||||
let has_delimiter = first.ends_with(['.', ')', ':']);
|
||||
let token = first.trim_end_matches(['.', ')', ':']);
|
||||
if token.is_empty() {
|
||||
return false;
|
||||
}
|
||||
let parts: Vec<&str> = token.split('.').collect();
|
||||
let decimal = parts
|
||||
.iter()
|
||||
.all(|p| !p.is_empty() && p.len() <= 3 && p.chars().all(|c| c.is_ascii_digit()));
|
||||
if decimal {
|
||||
// "1." / "2.1." carry a delimiter; "2.3 Title" is written without
|
||||
// one, so a multi-component number is accepted bare. A bare single
|
||||
// number ("3 apples") is not — that is ordinary prose.
|
||||
return has_delimiter || parts.len() >= 2;
|
||||
}
|
||||
// Roman numerals go through the heading parser's own grammar so the two
|
||||
// agree: uppercase I/V/X/L/C only, at most 8 characters. A looser rule
|
||||
// here would exempt markers the parser rejects — "iv)" or "d)" from an
|
||||
// alphabetical list — letting an ordinary list item bypass the veto and
|
||||
// reach heading promotion.
|
||||
//
|
||||
// A delimiter is also required: a bare leading "I" is the pronoun far
|
||||
// more often than a section number.
|
||||
has_delimiter && crate::markdown::heading::roman_value(token).is_some()
|
||||
}
|
||||
|
||||
/// True when the line reads as a title rather than a sentence: every
|
||||
/// content word (ignoring minor words) starts uppercase. Used to spare real
|
||||
/// headings from the dangling-verb veto — "Bond Yields" is a section title,
|
||||
/// "the method yields" is a stranded clause, and only the casing tells them
|
||||
/// apart.
|
||||
fn looks_title_case(t: &str) -> bool {
|
||||
const MINOR: &[&str] = &[
|
||||
"a", "an", "the", "of", "and", "or", "for", "to", "in", "on", "at", "by", "with", "from",
|
||||
"as", "is", "are", "that", "than", "into",
|
||||
];
|
||||
let mut content = 0usize;
|
||||
let mut capitalized = 0usize;
|
||||
for w in t.split_whitespace() {
|
||||
let cleaned: String = w.chars().filter(|c| c.is_alphabetic()).collect();
|
||||
if cleaned.is_empty() {
|
||||
continue;
|
||||
}
|
||||
if MINOR.contains(&cleaned.to_lowercase().as_str()) {
|
||||
continue;
|
||||
}
|
||||
content += 1;
|
||||
if cleaned.chars().next().is_some_and(char::is_uppercase) {
|
||||
capitalized += 1;
|
||||
}
|
||||
}
|
||||
// A single content word ("Yields") is a title by default.
|
||||
content == 0 || capitalized == content
|
||||
}
|
||||
|
||||
pub(crate) fn is_heading_fragment(text: &str) -> bool {
|
||||
let t = text.trim_end();
|
||||
|
||||
@@ -244,9 +312,133 @@ pub(crate) fn is_heading_fragment(text: &str) -> bool {
|
||||
if t.ends_with(':') && t.split_whitespace().any(is_equation_number) {
|
||||
return true;
|
||||
}
|
||||
|
||||
// Dangling clause: a stranded sentence lead-in ends on a relational
|
||||
// verb with no terminal punctuation — "Note that the exact error equals"
|
||||
// left ahead of its formula when a phantom table dissolved.
|
||||
//
|
||||
// Gated on the line reading as prose rather than a title. Case is the
|
||||
// discriminator the trailing word alone cannot provide: a heading is
|
||||
// title case ("Bond Yields", "The Method Yields") while a stranded
|
||||
// lead-in is sentence case ("the method yields"). Without this gate the
|
||||
// veto eats real headings — "Bond Yields", "Crop Yields" and any wrapped
|
||||
// title-case heading the preprocessor failed to merge.
|
||||
if !t.ends_with(['.', '!', '?', ':', ';', ')', ']'])
|
||||
&& !looks_title_case(t)
|
||||
&& !starts_with_numbering_prefix(t)
|
||||
{
|
||||
if let Some(last) = t.split_whitespace().next_back() {
|
||||
let word: String = last
|
||||
.trim_matches(|c: char| !c.is_alphanumeric())
|
||||
.to_lowercase();
|
||||
// Relational verbs only, and only those with no common noun
|
||||
// sense. "yields" was dropped for exactly that reason: "Bond
|
||||
// Yields" is a real section title. Function words, copulas and
|
||||
// auxiliaries were measured and rejected outright — a heading
|
||||
// that wraps across lines ends on those, and suppressing them
|
||||
// destroyed real IRS Publication 17 headings.
|
||||
const DANGLING_TAIL: &[&str] =
|
||||
&["equals", "denotes", "implies", "satisfies", "signifies"];
|
||||
if DANGLING_TAIL.contains(&word.as_str()) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
false
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod fragment_heading_tests {
|
||||
use super::is_heading_fragment;
|
||||
|
||||
#[test]
|
||||
fn dangling_tail_marks_stranded_clause() {
|
||||
// opendataloader 01030000000144: left behind when a phantom table
|
||||
// dissolved, ahead of its formula on the next line.
|
||||
assert!(is_heading_fragment("Note that the exact error equals"));
|
||||
assert!(is_heading_fragment("The remainder term satisfies"));
|
||||
assert!(is_heading_fragment("we conclude that the sum equals"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn real_headings_survive() {
|
||||
assert!(!is_heading_fragment("Introduction"));
|
||||
assert!(!is_heading_fragment("Error Analysis"));
|
||||
assert!(!is_heading_fragment("Materials and Methods"));
|
||||
assert!(!is_heading_fragment("Results"));
|
||||
assert!(!is_heading_fragment("3.2 Richardson Extrapolation"));
|
||||
assert!(!is_heading_fragment("Discussion and Conclusions"));
|
||||
// Terminal punctuation means the clause is complete.
|
||||
assert!(!is_heading_fragment("What is a Derivative?"));
|
||||
assert!(!is_heading_fragment("Procedure:"));
|
||||
assert!(!is_heading_fragment("Note that this is important."));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn title_case_headings_ending_in_a_verb_survive() {
|
||||
// "yields" is also a plural noun; these are real section titles.
|
||||
assert!(!is_heading_fragment("Bond Yields"));
|
||||
assert!(!is_heading_fragment("Crop Yields"));
|
||||
assert!(!is_heading_fragment("Dividend Yields"));
|
||||
assert!(!is_heading_fragment("Yields"));
|
||||
// A wrapped title-case heading whose first line ends on a listed
|
||||
// verb must survive even if the preprocessor failed to merge it.
|
||||
assert!(!is_heading_fragment("The Theorem Implies"));
|
||||
assert!(!is_heading_fragment("What This Denotes"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn numbered_sentence_case_headings_survive() {
|
||||
// heading.rs consults is_heading_fragment BEFORE applying its
|
||||
// numbered-prefix allowance, so the veto must not pre-empt it.
|
||||
assert!(!is_heading_fragment("1. What the model implies"));
|
||||
assert!(!is_heading_fragment("2.3 How the estimator satisfies"));
|
||||
assert!(!is_heading_fragment("IV) What this denotes"));
|
||||
// Without numbering the same wording is still a stranded clause.
|
||||
assert!(is_heading_fragment("What the model implies"));
|
||||
// A bare leading number or pronoun is prose, not numbering.
|
||||
assert!(is_heading_fragment("3 apples and what that implies"));
|
||||
assert!(is_heading_fragment("I think the model implies"));
|
||||
// Markers heading::parse_numbering rejects must not be exempted
|
||||
// either, or an ordinary list item bypasses the veto: lowercase
|
||||
// roman, alphabetical markers, and over-long tokens.
|
||||
assert!(is_heading_fragment("iv) the estimator satisfies"));
|
||||
assert!(is_heading_fragment("d) the value implies"));
|
||||
// Unsupported character (M is outside the parser's I/V/X/L/C set).
|
||||
assert!(is_heading_fragment("MMMM. the value implies"));
|
||||
// Over-long token: nine valid characters, so this exercises the
|
||||
// 8-character bound rather than the character set.
|
||||
assert!(is_heading_fragment("IIIIIIIII. the value implies"));
|
||||
// Eight is still within the bound and stays exempt.
|
||||
assert!(!is_heading_fragment("IIIIIIII. What this implies"));
|
||||
// Uppercase roman within the parser's grammar is still exempt.
|
||||
assert!(!is_heading_fragment("IV. What this denotes"));
|
||||
assert!(!is_heading_fragment("XII) What this implies"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn wrapped_headings_are_not_fragments() {
|
||||
// A heading that wraps across lines ends on a function word. These
|
||||
// are real headings from IRS Publication 17 and must survive.
|
||||
assert!(!is_heading_fragment("Casualty and"));
|
||||
assert!(!is_heading_fragment("Rule 10. You Must Be at"));
|
||||
assert!(!is_heading_fragment("Higher Standard Deduction for"));
|
||||
assert!(!is_heading_fragment("Qualifying Child of"));
|
||||
assert!(!is_heading_fragment("When Can I Withdraw or"));
|
||||
// Copulas and auxiliaries also end real wrapped headings.
|
||||
assert!(!is_heading_fragment("Rule 15. Your AGI Must Be"));
|
||||
assert!(!is_heading_fragment("What Medical Expenses Are"));
|
||||
assert!(!is_heading_fragment("Rule 13. You Must Have"));
|
||||
assert!(!is_heading_fragment("When Can a Roth IRA Be"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dangling_check_is_case_insensitive() {
|
||||
// All-caps is not sentence case, so the veto must not fire there.
|
||||
assert!(!is_heading_fragment("THE REMAINDER EQUALS"));
|
||||
}
|
||||
}
|
||||
|
||||
/// Compute the Y-gap threshold for paragraph break detection.
|
||||
///
|
||||
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
|
||||
|
||||
@@ -127,7 +127,9 @@ fn visual_style(line: &TextLine) -> Option<VisualStyle> {
|
||||
})
|
||||
}
|
||||
|
||||
fn roman_value(token: &str) -> Option<u32> {
|
||||
/// Shared with `analysis::starts_with_numbering_prefix` so the veto
|
||||
/// exemption and the heading parser agree on what a roman numeral is.
|
||||
pub(super) fn roman_value(token: &str) -> Option<u32> {
|
||||
if token.is_empty() || token.len() > 8 {
|
||||
return None;
|
||||
}
|
||||
|
||||
@@ -271,6 +271,12 @@ pub struct PyTextItem {
|
||||
pub is_strikeout: bool,
|
||||
#[pyo3(get)]
|
||||
pub item_type: String,
|
||||
/// Marked Content ID from the content stream's BDC/BMC operator, None
|
||||
/// when the text is not part of marked content. Join with the
|
||||
/// (page, mcid) pairs from extract_structure_elements to attach
|
||||
/// structure-tree roles (headings, paragraphs, ...) in tagged PDFs.
|
||||
#[pyo3(get)]
|
||||
pub mcid: Option<i64>,
|
||||
}
|
||||
|
||||
#[pymethods]
|
||||
@@ -286,6 +292,32 @@ impl PyTextItem {
|
||||
}
|
||||
}
|
||||
|
||||
/// One structure-tree element reference from a tagged PDF.
|
||||
#[pyclass(name = "StructureElement")]
|
||||
#[derive(Clone)]
|
||||
pub struct PyStructureElement {
|
||||
/// 1-indexed page number (matches TextItem.page).
|
||||
#[pyo3(get)]
|
||||
pub page: u32,
|
||||
/// Marked Content ID from the page's content stream (matches
|
||||
/// TextItem.mcid).
|
||||
#[pyo3(get)]
|
||||
pub mcid: i64,
|
||||
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", ...).
|
||||
#[pyo3(get)]
|
||||
pub role: String,
|
||||
}
|
||||
|
||||
#[pymethods]
|
||||
impl PyStructureElement {
|
||||
fn __repr__(&self) -> String {
|
||||
format!(
|
||||
"StructureElement(page={}, mcid={}, role='{}')",
|
||||
self.page, self.mcid, self.role
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Helpers
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -356,6 +388,18 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
|
||||
is_underline: item.is_underline,
|
||||
is_strikeout: item.is_strikeout,
|
||||
item_type: item_type_str(&item.item_type),
|
||||
mcid: item.mcid,
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn convert_structure_elements(elements: Vec<crate::StructureElement>) -> Vec<PyStructureElement> {
|
||||
elements
|
||||
.into_iter()
|
||||
.map(|e| PyStructureElement {
|
||||
page: e.page,
|
||||
mcid: e.mcid,
|
||||
role: e.role,
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
@@ -613,6 +657,48 @@ fn extract_pages_markdown_bytes(
|
||||
Ok(to_py_pages_result(result))
|
||||
}
|
||||
|
||||
/// Extract structure-tree element references from a tagged PDF file.
|
||||
///
|
||||
/// Parses the document's structure tree (when present) and returns one
|
||||
/// entry per marked-content reference, resolved to its 1-indexed page,
|
||||
/// MCID, and structure type name ("H1".."H6", "P", "Table", ...). Returns
|
||||
/// an empty list when the PDF is not tagged.
|
||||
///
|
||||
/// Join (page, mcid) against the page/mcid attributes from
|
||||
/// [`extract_text_with_positions`] to attach heading levels and other
|
||||
/// semantic roles to extracted text.
|
||||
///
|
||||
/// Args:
|
||||
/// path: Path to the PDF file.
|
||||
/// pages: Optional list of 1-indexed pages (matching TextItem.page).
|
||||
/// When None (default), the whole document is returned.
|
||||
///
|
||||
/// Returns:
|
||||
/// List of StructureElement sorted by (page, mcid).
|
||||
#[pyfunction]
|
||||
#[pyo3(signature = (path, pages=None))]
|
||||
fn extract_structure_elements(
|
||||
path: &str,
|
||||
pages: Option<Vec<u32>>,
|
||||
) -> PyResult<Vec<PyStructureElement>> {
|
||||
let elements = crate::extract_structure_elements(path, pages.as_deref()).map_err(to_py_err)?;
|
||||
Ok(convert_structure_elements(elements))
|
||||
}
|
||||
|
||||
/// Extract structure-tree element references from tagged PDF bytes.
|
||||
///
|
||||
/// See [`extract_structure_elements`] for details.
|
||||
#[pyfunction]
|
||||
#[pyo3(signature = (data, pages=None))]
|
||||
fn extract_structure_elements_bytes(
|
||||
data: &[u8],
|
||||
pages: Option<Vec<u32>>,
|
||||
) -> PyResult<Vec<PyStructureElement>> {
|
||||
let elements =
|
||||
crate::extract_structure_elements_mem(data, pages.as_deref()).map_err(to_py_err)?;
|
||||
Ok(convert_structure_elements(elements))
|
||||
}
|
||||
|
||||
/// Python module definition.
|
||||
#[pymodule]
|
||||
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
||||
@@ -620,6 +706,7 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
||||
m.add_class::<PyPageOcrReasons>()?;
|
||||
m.add_class::<PyPdfClassification>()?;
|
||||
m.add_class::<PyTextItem>()?;
|
||||
m.add_class::<PyStructureElement>()?;
|
||||
m.add_class::<PyRegionText>()?;
|
||||
m.add_class::<PyPageRegionTexts>()?;
|
||||
m.add_class::<PyPageMarkdown>()?;
|
||||
@@ -634,6 +721,8 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
||||
m.add_function(wrap_pyfunction!(extract_text_bytes, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_text_with_positions, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_text_with_positions_bytes, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_structure_elements, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_structure_elements_bytes, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_text_in_regions, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_text_in_regions_bytes, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_pages_markdown, m)?)?;
|
||||
|
||||
+819
-38
File diff suppressed because it is too large
Load Diff
+308
-10
@@ -435,6 +435,72 @@ fn revised_table_cell_indices(
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Index of candidate "body" items (larger-font attachment targets) sorted by
|
||||
/// Y, so script-attachment checks scan a narrow Y window instead of the whole
|
||||
/// page per candidate.
|
||||
struct ScriptBodyIndex<'a> {
|
||||
/// (y, item), sorted ascending by y
|
||||
by_y: Vec<(f32, &'a TextItem)>,
|
||||
/// widest vertical attachment window any body item can produce
|
||||
max_window: f32,
|
||||
}
|
||||
|
||||
impl<'a> ScriptBodyIndex<'a> {
|
||||
fn new(items: &'a [TextItem]) -> Self {
|
||||
// Smallest table-candidate font is 6pt, so any possible attachment
|
||||
// target is at least 6 x 1.2 pt.
|
||||
let mut by_y: Vec<(f32, &TextItem)> = items
|
||||
.iter()
|
||||
.filter(|i| i.font_size >= 6.0 * 1.2)
|
||||
.map(|i| (i.y, i))
|
||||
.collect();
|
||||
by_y.sort_by(|a, b| a.0.total_cmp(&b.0));
|
||||
let max_window = by_y
|
||||
.iter()
|
||||
.map(|(_, i)| i.font_size * 0.8)
|
||||
.fold(0.0f32, f32::max);
|
||||
Self { by_y, max_window }
|
||||
}
|
||||
|
||||
/// True when a small-font item is horizontally attached to a larger-font
|
||||
/// item at a script baseline offset — a sub/superscript in running text
|
||||
/// or math (equation subscripts, footnote markers). Script attachments
|
||||
/// are not table cells; without this filter, display equations with
|
||||
/// sub/superscripts form phantom small-font table regions (e.g. TeX
|
||||
/// papers where log subscripts cluster with footnote lines into a fake
|
||||
/// 3-column table). A genuine baseline offset is required so same-line
|
||||
/// table neighbours (a small cell beside a larger label cell) are never
|
||||
/// classified as scripts.
|
||||
///
|
||||
/// `min_anchor_size` additionally constrains what counts as an
|
||||
/// attachment target: the small-font pass accepts any sufficiently
|
||||
/// larger item (0.0), while the body-font pass requires a heading-sized
|
||||
/// anchor so a body-size table cell beside a slightly larger label with
|
||||
/// baseline jitter is never treated as a script.
|
||||
fn is_script_attachment(&self, small: &TextItem, min_anchor_size: f32) -> bool {
|
||||
let attach_gap = small.font_size.max(4.0) * 0.6;
|
||||
let lo = self
|
||||
.by_y
|
||||
.partition_point(|(y, _)| *y < small.y - self.max_window);
|
||||
self.by_y[lo..]
|
||||
.iter()
|
||||
.take_while(|(y, _)| *y <= small.y + self.max_window)
|
||||
.any(|(_, body)| {
|
||||
let dy = (small.y - body.y).abs();
|
||||
body.font_size >= small.font_size * 1.2
|
||||
&& body.font_size >= min_anchor_size
|
||||
&& dy > body.font_size * 0.05
|
||||
&& dy <= body.font_size * 0.8
|
||||
&& {
|
||||
let gap_after_body = small.x - (body.x + body.width);
|
||||
let gap_before_body = body.x - (small.x + small.width);
|
||||
(-attach_gap..=attach_gap).contains(&gap_after_body)
|
||||
|| (-attach_gap..=attach_gap).contains(&gap_before_body)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// Detect tables in a set of text items from a single page
|
||||
pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bool) -> Vec<Table> {
|
||||
detect_tables_with_page_width(items, base_font_size, skip_body_font, content_width(items))
|
||||
@@ -483,6 +549,27 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
// === Pass 1: Small-font tables (existing behavior) ===
|
||||
let table_font_threshold = base_font_size * 0.90;
|
||||
|
||||
// Mark sub/superscript attachments once per pass. They stay candidates —
|
||||
// the masks only remove them from region qualification and column/row
|
||||
// geometry.
|
||||
//
|
||||
// The two passes need different anchor thresholds. In the small-font pass
|
||||
// any sufficiently larger neighbour is a plausible base for a script. In
|
||||
// the body-font pass the candidates are themselves body-sized
|
||||
// (0.85..1.05x), so a merely "slightly larger" neighbour is usually a bold
|
||||
// label or an adjacent column header, not the base of a superscript —
|
||||
// treating it as one would strip real cells out of the geometry and lose
|
||||
// the table. Requiring a heading-sized anchor (>= 1.15x base) keeps the
|
||||
// body pass to genuine scripts hanging off headings.
|
||||
let script_index = ScriptBodyIndex::new(items);
|
||||
let script_flags: Vec<bool> = items
|
||||
.iter()
|
||||
.map(|item| script_index.is_script_attachment(item, 0.0))
|
||||
.collect();
|
||||
let body_script_flags: Vec<bool> = items
|
||||
.iter()
|
||||
.map(|item| script_index.is_script_attachment(item, base_font_size * 1.15))
|
||||
.collect();
|
||||
let table_candidates: Vec<(usize, &TextItem)> = items
|
||||
.iter()
|
||||
.enumerate()
|
||||
@@ -494,7 +581,14 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
.collect();
|
||||
|
||||
if table_candidates.len() >= 6 {
|
||||
let regions = find_table_regions(&table_candidates);
|
||||
// Qualify regions from non-script items: a cluster of sub/superscripts
|
||||
// must not, on its own, mark out a table region.
|
||||
let region_evidence: Vec<(usize, &TextItem)> = table_candidates
|
||||
.iter()
|
||||
.filter(|(idx, _)| !script_flags[*idx])
|
||||
.cloned()
|
||||
.collect();
|
||||
let regions = find_table_regions(®ion_evidence);
|
||||
|
||||
for (y_min, y_max) in regions {
|
||||
let region_items: Vec<(usize, &TextItem)> = table_candidates
|
||||
@@ -508,7 +602,9 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
}
|
||||
|
||||
if let Some(mut table) =
|
||||
detect_table_in_region(®ion_items, TableDetectionMode::SmallFont)
|
||||
detect_table_in_region(®ion_items, TableDetectionMode::SmallFont, &|i| {
|
||||
script_flags[i]
|
||||
})
|
||||
{
|
||||
// Try to recover body-font header row above the small-font table
|
||||
recover_header_row(&mut table, items, table_font_threshold);
|
||||
@@ -553,8 +649,20 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
body_font_low,
|
||||
body_font_high,
|
||||
);
|
||||
// Scripts are NOT filtered out of the candidate set here, mirroring
|
||||
// the small-font pass: they must stay eligible for cell assignment so
|
||||
// a sub/superscript that belongs inside a table cell keeps its text.
|
||||
// The heading-anchored `body_script_flags` mask removes them from
|
||||
// geometry only.
|
||||
if body_candidates.len() >= 6 {
|
||||
let regions = find_table_regions_strict(&body_candidates);
|
||||
// Same reasoning as the small-font pass: scripts do not qualify
|
||||
// regions, but remain available for cell assignment within one.
|
||||
let region_evidence: Vec<(usize, &TextItem)> = body_candidates
|
||||
.iter()
|
||||
.filter(|(idx, _)| !body_script_flags[*idx])
|
||||
.cloned()
|
||||
.collect();
|
||||
let regions = find_table_regions_strict(®ion_evidence);
|
||||
log::debug!("body-font: {} strict regions found", regions.len());
|
||||
|
||||
for (y_min, y_max, _x_min, _x_max) in ®ions {
|
||||
@@ -580,7 +688,9 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
}
|
||||
|
||||
if let Some(table) =
|
||||
detect_table_in_region(®ion_items, TableDetectionMode::BodyFont)
|
||||
detect_table_in_region(®ion_items, TableDetectionMode::BodyFont, &|i| {
|
||||
body_script_flags[i]
|
||||
})
|
||||
{
|
||||
tables.push(table);
|
||||
}
|
||||
@@ -808,10 +918,30 @@ fn find_table_regions_strict(items: &[(usize, &TextItem)]) -> Vec<(f32, f32, f32
|
||||
regions
|
||||
}
|
||||
|
||||
/// Detect a table within a specific region
|
||||
fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode) -> Option<Table> {
|
||||
// Find column boundaries
|
||||
let columns = find_column_boundaries(items, mode);
|
||||
/// Detect a table within a specific region.
|
||||
///
|
||||
/// `is_script` marks items that are sub/superscript attachments. Those are
|
||||
/// excluded from the *geometry* — they must not be able to create a column,
|
||||
/// which is how equation subscript clusters used to fabricate phantom grids —
|
||||
/// but they remain eligible for cell assignment, so legitimate cell content
|
||||
/// (exponents in an engineering-notation table, footnote markers) stays in
|
||||
/// the cell it belongs to instead of leaking out into the reading order.
|
||||
fn detect_table_in_region(
|
||||
items: &[(usize, &TextItem)],
|
||||
mode: TableDetectionMode,
|
||||
is_script: &dyn Fn(usize) -> bool,
|
||||
) -> Option<Table> {
|
||||
// Column geometry from non-script items only.
|
||||
let geometry_items: Vec<(usize, &TextItem)> = items
|
||||
.iter()
|
||||
.filter(|(idx, _)| !is_script(*idx))
|
||||
.cloned()
|
||||
.collect();
|
||||
// A region that is *entirely* scripts has no table structure at all.
|
||||
if geometry_items.is_empty() {
|
||||
return None;
|
||||
}
|
||||
let columns = find_column_boundaries(&geometry_items, mode);
|
||||
let min_cols = 2;
|
||||
if columns.len() < min_cols || columns.len() > 25 {
|
||||
log::debug!(
|
||||
@@ -822,8 +952,8 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
|
||||
return None;
|
||||
}
|
||||
|
||||
// Find row boundaries
|
||||
let rows = find_row_boundaries(items);
|
||||
// Find row boundaries (geometry items only, same reasoning)
|
||||
let rows = find_row_boundaries(&geometry_items);
|
||||
let min_rows = 2;
|
||||
if rows.len() < min_rows {
|
||||
log::debug!(
|
||||
@@ -842,6 +972,11 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
|
||||
);
|
||||
|
||||
// Verify this looks like a table: multiple items should align to columns
|
||||
// Validate against ALL items, including scripts. Columns are derived from
|
||||
// non-script geometry so scripts cannot *create* a column, but excluding
|
||||
// them from validation too would let a region manufacture alignment: drop
|
||||
// the awkward items and whatever remains looks like a tidy grid. Block
|
||||
// diagrams did exactly that. Everything in the region must fit.
|
||||
let col_alignment = check_column_alignment(items, &columns, mode);
|
||||
let min_alignment = match mode {
|
||||
TableDetectionMode::SmallFont => 0.5,
|
||||
@@ -912,6 +1047,29 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
|
||||
cells.push(row_cells);
|
||||
}
|
||||
|
||||
// Validation 0 (small-font pass only): reject tiny all-numeric
|
||||
// fragments. A <=2-row grid whose every cell is a bare 1-2 digit number
|
||||
// carries no tabular information — in practice these are
|
||||
// exponent/subscript clusters from display math that happen to align.
|
||||
// Body-font tables are not subject to this veto: their cells cannot be
|
||||
// script glyphs.
|
||||
if matches!(mode, TableDetectionMode::SmallFont) {
|
||||
let nonempty_cells: Vec<&String> =
|
||||
cells.iter().flatten().filter(|c| !c.is_empty()).collect();
|
||||
if rows.len() <= 2
|
||||
&& !nonempty_cells.is_empty()
|
||||
&& nonempty_cells
|
||||
.iter()
|
||||
.all(|c| c.len() <= 2 && c.chars().all(|ch| ch.is_ascii_digit()))
|
||||
{
|
||||
log::debug!(
|
||||
" validation 0 fail: tiny all-numeric fragment ({} cells)",
|
||||
nonempty_cells.len()
|
||||
);
|
||||
return None;
|
||||
}
|
||||
}
|
||||
|
||||
// Validation 1: some rows should have content in first column.
|
||||
// Use a lower threshold (25%) for tables with wrapped cells where
|
||||
// continuation lines leave the first column empty.
|
||||
@@ -1977,6 +2135,146 @@ fn try_add_label_column(
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
|
||||
fn make_item(text: &str, x: f32, y: f32, font_size: f32, width: f32) -> TextItem {
|
||||
TextItem {
|
||||
text: text.to_string(),
|
||||
x,
|
||||
y,
|
||||
width,
|
||||
height: font_size,
|
||||
font: "TestFont".to_string(),
|
||||
font_size,
|
||||
page: 1,
|
||||
is_bold: false,
|
||||
is_italic: false,
|
||||
is_underline: false,
|
||||
is_strikeout: false,
|
||||
item_type: ItemType::Text,
|
||||
mcid: None,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_detects_subscript_after_body_text() {
|
||||
let body = make_item("log", 100.0, 500.0, 10.0, 15.0);
|
||||
let sub = make_item("10", 115.5, 497.0, 7.0, 7.0);
|
||||
let items = vec![body, sub.clone()];
|
||||
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sub, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_detects_superscript_footnote_marker() {
|
||||
let body = make_item("Hartley", 200.0, 500.0, 10.0, 35.0);
|
||||
let sup = make_item("2", 235.8, 504.0, 6.6, 3.5);
|
||||
let items = vec![body, sup.clone()];
|
||||
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sup, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_ignores_small_cell_far_from_body_text() {
|
||||
let body = make_item("Revenue", 100.0, 500.0, 10.0, 40.0);
|
||||
let cell = make_item("1,234", 180.0, 500.0, 7.0, 20.0);
|
||||
let items = vec![body, cell.clone()];
|
||||
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn body_pass_anchor_spares_cells_beside_slightly_larger_labels() {
|
||||
// A body-font table cell (10pt) sitting beside a slightly larger,
|
||||
// NON-heading label (12.5pt) with a little baseline jitter. The
|
||||
// small-font pass treats any larger neighbour as a possible script
|
||||
// base, but the body pass must not: at body sizes a slightly larger
|
||||
// neighbour is a bold label or column header, and flagging the cell
|
||||
// would strip it out of the table geometry and lose the table.
|
||||
// Cell at the low end of the body band (0.85x base) beside a 10.5pt
|
||||
// label. 10.5 clears the inherent 1.2x-of-cell rule (10.2) but falls
|
||||
// below the body pass's heading anchor (11.5), which is exactly the
|
||||
// band where the two masks must disagree.
|
||||
let label = make_item("Revenue", 100.0, 500.0, 10.5, 40.0);
|
||||
let cell = make_item("1,234", 141.0, 496.5, 8.5, 22.0);
|
||||
let items = vec![label, cell.clone()];
|
||||
let index = ScriptBodyIndex::new(&items);
|
||||
let base = 10.0;
|
||||
assert!(
|
||||
index.is_script_attachment(&cell, 0.0),
|
||||
"small-font pass anchor should still see this as an attachment"
|
||||
);
|
||||
assert!(
|
||||
!index.is_script_attachment(&cell, base * 1.15),
|
||||
"body pass must not treat a cell beside a slightly larger label \
|
||||
as a script — that removes real cells from the geometry"
|
||||
);
|
||||
// A genuine heading-sized anchor still qualifies in the body pass.
|
||||
let heading = make_item("Section", 100.0, 500.0, 20.0, 60.0);
|
||||
let sup = make_item("3", 161.0, 508.0, 10.0, 5.0);
|
||||
let h_items = vec![heading, sup.clone()];
|
||||
assert!(
|
||||
ScriptBodyIndex::new(&h_items).is_script_attachment(&sup, base * 1.15),
|
||||
"script hanging off a heading must still be excluded in the body pass"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_ignores_same_baseline_neighbor_cell() {
|
||||
// A small cell beside a larger label on the SAME baseline is a table
|
||||
// layout, not a subscript — a genuine baseline offset is required.
|
||||
let label = make_item("Total", 100.0, 500.0, 10.0, 25.0);
|
||||
let cell = make_item("42", 127.0, 500.0, 7.5, 9.0);
|
||||
let items = vec![label, cell.clone()];
|
||||
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_ignores_neighbor_on_different_line() {
|
||||
let body = make_item("Header", 100.0, 500.0, 10.0, 30.0);
|
||||
let cell = make_item("42", 131.0, 486.0, 7.0, 10.0);
|
||||
let items = vec![body, cell.clone()];
|
||||
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||
}
|
||||
|
||||
/// Equation-subscript + footnote layout from Shannon entropy.pdf page 1,
|
||||
/// with real coordinates. Without the larger-font anchors the small items
|
||||
/// alone DO form a phantom table — proving the layout reaches detection —
|
||||
/// and adding the anchors must suppress it.
|
||||
fn shannon_page1_small_items() -> Vec<TextItem> {
|
||||
vec![
|
||||
make_item("2", 267.4, 133.9, 7.4, 3.7),
|
||||
make_item("10", 306.2, 133.9, 7.4, 7.4),
|
||||
make_item("10", 342.7, 133.9, 7.4, 7.4),
|
||||
make_item("10", 325.0, 118.9, 7.4, 7.4),
|
||||
make_item("Bell System Technical Journal,", 295.7, 101.9, 8.0, 95.0),
|
||||
make_item(
|
||||
"April 1924, p. 324; Certain Topics in",
|
||||
396.7,
|
||||
101.9,
|
||||
8.0,
|
||||
130.0,
|
||||
),
|
||||
make_item("v. 47, April 1928, p. 617.", 250.9, 92.5, 8.0, 90.0),
|
||||
make_item("Bell System Technical Journal,", 264.2, 82.6, 8.0, 95.0),
|
||||
make_item("July 1928, p. 535.", 364.3, 82.6, 8.0, 65.0),
|
||||
]
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn equation_scripts_do_not_form_phantom_table() {
|
||||
let bare = shannon_page1_small_items();
|
||||
assert!(
|
||||
!detect_tables(&bare, 10.0, false).is_empty(),
|
||||
"test layout must form a phantom table when the filter cannot fire"
|
||||
);
|
||||
let mut items = shannon_page1_small_items();
|
||||
items.push(make_item("log", 253.0, 137.0, 10.0, 13.5));
|
||||
items.push(make_item("log", 291.5, 137.0, 10.0, 13.5));
|
||||
items.push(make_item("log", 328.0, 137.0, 10.0, 13.5));
|
||||
items.push(make_item("log", 310.3, 122.0, 10.0, 13.5));
|
||||
let tables = detect_tables(&items, 10.0, false);
|
||||
assert!(
|
||||
tables.is_empty(),
|
||||
"equation scripts + footnotes must not become a table: {tables:?}"
|
||||
);
|
||||
}
|
||||
use super::*;
|
||||
use crate::types::ItemType;
|
||||
|
||||
|
||||
+335
-15
@@ -1,6 +1,6 @@
|
||||
//! Rectangle-based table detection using union-find clustering.
|
||||
|
||||
use std::collections::HashMap;
|
||||
use std::collections::{BTreeMap, HashMap, HashSet};
|
||||
|
||||
use log::debug;
|
||||
|
||||
@@ -78,19 +78,111 @@ pub(crate) fn rects_overlap(a: &(f32, f32, f32, f32), b: &(f32, f32, f32, f32),
|
||||
!(a_right < b_left || b_right < a_left || a_top < b_bottom || b_top < a_bottom)
|
||||
}
|
||||
|
||||
fn grid_coord(value: f32, cell: f32) -> i32 {
|
||||
(value / cell).floor().clamp(-1_000_000.0, 1_000_000.0) as i32
|
||||
}
|
||||
|
||||
/// Inclusive grid range. `None` if the rect covers more cells than we will
|
||||
/// materialize — those rects are clustered via a bounded fallback.
|
||||
fn grid_span(lo: f32, hi: f32, cell: f32) -> Option<std::ops::RangeInclusive<i32>> {
|
||||
let a = grid_coord(lo.min(hi), cell);
|
||||
let b = grid_coord(lo.max(hi), cell);
|
||||
let span = b.saturating_sub(a);
|
||||
if span > 64 {
|
||||
return None;
|
||||
}
|
||||
Some(a..=b)
|
||||
}
|
||||
|
||||
fn union_bucket_pairs(
|
||||
uf: &mut UnionFind,
|
||||
rects: &[(f32, f32, f32, f32)],
|
||||
bucket: &[usize],
|
||||
tolerance: f32,
|
||||
) {
|
||||
let m = bucket.len();
|
||||
let mut pairs = 0usize;
|
||||
'cell: for a in 0..m {
|
||||
let i = bucket[a];
|
||||
if uf.component_size(i) >= MAX_CLUSTER_RECTS {
|
||||
continue;
|
||||
}
|
||||
for &j in &bucket[a + 1..] {
|
||||
if pairs >= MAX_CLUSTER_PAIRS_PER_CELL {
|
||||
break 'cell;
|
||||
}
|
||||
if uf.component_size(j) >= MAX_CLUSTER_RECTS {
|
||||
continue;
|
||||
}
|
||||
pairs += 1;
|
||||
if rects_overlap(&rects[i], &rects[j], tolerance) {
|
||||
uf.union(i, j);
|
||||
if uf.component_size(i) >= MAX_CLUSTER_RECTS {
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn union_rect_against_bands(
|
||||
uf: &mut UnionFind,
|
||||
rects: &[(f32, f32, f32, f32)],
|
||||
i: usize,
|
||||
bands: &BTreeMap<i32, Vec<usize>>,
|
||||
lo: i32,
|
||||
hi: i32,
|
||||
tolerance: f32,
|
||||
) {
|
||||
if uf.component_size(i) >= MAX_CLUSTER_RECTS {
|
||||
return;
|
||||
}
|
||||
let mut pairs = 0usize;
|
||||
let mut seen = HashSet::new();
|
||||
for (_, bucket) in bands.range(lo..=hi) {
|
||||
for &j in bucket {
|
||||
if !seen.insert(j) {
|
||||
continue;
|
||||
}
|
||||
if pairs >= MAX_CLUSTER_PAIRS_PER_CELL {
|
||||
return;
|
||||
}
|
||||
if i == j || uf.component_size(j) >= MAX_CLUSTER_RECTS {
|
||||
continue;
|
||||
}
|
||||
pairs += 1;
|
||||
if rects_overlap(&rects[i], &rects[j], tolerance) {
|
||||
uf.union(i, j);
|
||||
if uf.component_size(i) >= MAX_CLUSTER_RECTS {
|
||||
return;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Maximum component size for rect clustering. No real table has thousands
|
||||
/// of cell rects — once a component exceeds this, it is a vector drawing or
|
||||
/// page-spanning clipping path. We skip overlap checks for rects already in
|
||||
/// an oversized component, keeping the original O(n²) loop but making it
|
||||
/// effectively O(n) for pathological pages.
|
||||
/// an oversized component.
|
||||
const MAX_CLUSTER_RECTS: usize = 2000;
|
||||
|
||||
/// Pairwise-disjoint rects never merge, so a component-size cap does not
|
||||
/// stop an all-pairs loop. Rects are hashed into this many points of grid
|
||||
/// and compared only against others in the same cell.
|
||||
const CLUSTER_GRID_CELL: f32 = 64.0;
|
||||
|
||||
/// All-pairs AABB tests allowed inside one grid cell. A real table cell is
|
||||
/// tens of points wide, so a 64-pt cell holds a handful of neighbors — not
|
||||
/// thousands of stacked drawings.
|
||||
const MAX_CLUSTER_PAIRS_PER_CELL: usize = 16_384;
|
||||
|
||||
/// Cluster rects by spatial overlap using union-find.
|
||||
/// Returns groups of rect indices; only groups with ≥ `min_size` rects are returned.
|
||||
///
|
||||
/// Skips overlap checks for rects whose component has already exceeded
|
||||
/// [`MAX_CLUSTER_RECTS`], so pages with tens of thousands of vector-drawing
|
||||
/// rects complete in milliseconds instead of minutes.
|
||||
/// Overlap tests run inside a uniform grid so far-apart rects are never
|
||||
/// compared, and each cell is pair-capped so a dense stack cannot go
|
||||
/// quadratic or starve an independent table in another cell.
|
||||
pub(crate) fn cluster_rects(
|
||||
rects: &[(f32, f32, f32, f32)],
|
||||
tolerance: f32,
|
||||
@@ -98,23 +190,144 @@ pub(crate) fn cluster_rects(
|
||||
) -> Vec<Vec<usize>> {
|
||||
let n = rects.len();
|
||||
let mut uf = UnionFind::new(n);
|
||||
let cell = CLUSTER_GRID_CELL.max(tolerance * 4.0);
|
||||
|
||||
for i in 0..n {
|
||||
// If rect i is already in an oversized component, no point comparing
|
||||
// it against further rects — the component won't be used for table
|
||||
// detection anyway.
|
||||
let mut grid: HashMap<(i32, i32), Vec<usize>> = HashMap::new();
|
||||
let mut large: Vec<usize> = Vec::new();
|
||||
for (idx, &(x, y, w, h)) in rects.iter().enumerate() {
|
||||
match (
|
||||
grid_span(x - tolerance, x + w + tolerance, cell),
|
||||
grid_span(y - tolerance, y + h + tolerance, cell),
|
||||
) {
|
||||
(Some(xs), Some(ys)) => {
|
||||
for gx in xs {
|
||||
for gy in ys.clone() {
|
||||
grid.entry((gx, gy)).or_default().push(idx);
|
||||
}
|
||||
}
|
||||
}
|
||||
_ => large.push(idx),
|
||||
}
|
||||
}
|
||||
|
||||
let mut keys: Vec<_> = grid.keys().copied().collect();
|
||||
keys.sort_unstable();
|
||||
let mut keys_by_y: BTreeMap<i32, Vec<i32>> = BTreeMap::new();
|
||||
for &key in &keys {
|
||||
union_bucket_pairs(&mut uf, rects, &grid[&key], tolerance);
|
||||
keys_by_y.entry(key.1).or_default().push(key.0);
|
||||
}
|
||||
|
||||
// Oversized spans skip insert. Range-query occupied cells they cover so
|
||||
// later X-ranges are not starved and we do not scan unrelated rows.
|
||||
for &i in &large {
|
||||
if uf.component_size(i) >= MAX_CLUSTER_RECTS {
|
||||
continue;
|
||||
}
|
||||
for j in (i + 1)..n {
|
||||
if rects_overlap(&rects[i], &rects[j], tolerance) {
|
||||
uf.union(i, j);
|
||||
// Check if the merged component just exceeded the cap —
|
||||
// if so, no need to test more pairs for rect i.
|
||||
let (x, y, w, h) = rects[i];
|
||||
let x_lo = grid_coord(x - tolerance, cell);
|
||||
let x_hi = grid_coord(x + w + tolerance, cell);
|
||||
let y_lo = grid_coord(y - tolerance, cell);
|
||||
let y_hi = grid_coord(y + h + tolerance, cell);
|
||||
for (&gy, gxs) in keys_by_y.range(y_lo..=y_hi) {
|
||||
let start = gxs.partition_point(|&gx| gx < x_lo);
|
||||
for &gx in &gxs[start..] {
|
||||
if gx > x_hi {
|
||||
break;
|
||||
}
|
||||
let bucket = &grid[&(gx, gy)];
|
||||
let mut pairs = 0usize;
|
||||
for &j in bucket {
|
||||
if pairs >= MAX_CLUSTER_PAIRS_PER_CELL {
|
||||
break;
|
||||
}
|
||||
if uf.component_size(j) >= MAX_CLUSTER_RECTS {
|
||||
continue;
|
||||
}
|
||||
pairs += 1;
|
||||
if rects_overlap(&rects[i], &rects[j], tolerance) {
|
||||
uf.union(i, j);
|
||||
if uf.component_size(i) >= MAX_CLUSTER_RECTS {
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
if uf.component_size(i) >= MAX_CLUSTER_RECTS {
|
||||
break;
|
||||
}
|
||||
}
|
||||
if uf.component_size(i) >= MAX_CLUSTER_RECTS {
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Oversized-vs-oversized: band on the short axis so stacked or side-by-side
|
||||
// page-spanning rules stay linear. Wide vs tall pairs are matched by
|
||||
// querying the tall X-index; dual-oversized rects occupy every coarse-Y
|
||||
// cell they span.
|
||||
let mut large_x: BTreeMap<i32, Vec<usize>> = BTreeMap::new();
|
||||
let mut large_y: BTreeMap<i32, Vec<usize>> = BTreeMap::new();
|
||||
let mut large_coarse_y: BTreeMap<i32, Vec<usize>> = BTreeMap::new();
|
||||
let mut wide: Vec<usize> = Vec::new();
|
||||
let mut dual: Vec<usize> = Vec::new();
|
||||
for &i in &large {
|
||||
let (x, y, w, h) = rects[i];
|
||||
let xs = grid_span(x - tolerance, x + w + tolerance, cell);
|
||||
let ys = grid_span(y - tolerance, y + h + tolerance, cell);
|
||||
match (xs, ys) {
|
||||
(Some(xs), _) => {
|
||||
for gx in xs {
|
||||
large_x.entry(gx).or_default().push(i);
|
||||
}
|
||||
}
|
||||
(_, Some(ys)) => {
|
||||
wide.push(i);
|
||||
for gy in ys {
|
||||
large_y.entry(gy).or_default().push(i);
|
||||
}
|
||||
}
|
||||
_ => {
|
||||
dual.push(i);
|
||||
let coarse = cell * 64.0;
|
||||
match grid_span(y - tolerance, y + h + tolerance, coarse) {
|
||||
Some(ys) => {
|
||||
for gy in ys {
|
||||
large_coarse_y.entry(gy).or_default().push(i);
|
||||
}
|
||||
}
|
||||
None => {
|
||||
large_coarse_y.entry(i32::MIN).or_default().push(i);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
for bands in [&large_x, &large_y, &large_coarse_y] {
|
||||
for bucket in bands.values() {
|
||||
union_bucket_pairs(&mut uf, rects, bucket, tolerance);
|
||||
}
|
||||
}
|
||||
// Cross-orientation is |wide|×|tall| if every wide rule spans the page.
|
||||
// Skip that pass when the product cannot be a table (a few rules).
|
||||
let tall_n = large
|
||||
.len()
|
||||
.saturating_sub(wide.len())
|
||||
.saturating_sub(dual.len());
|
||||
let cross_n =
|
||||
(wide.len() + dual.len()).saturating_mul(tall_n) + dual.len().saturating_mul(wide.len());
|
||||
if cross_n > 0 && cross_n <= MAX_CLUSTER_PAIRS_PER_CELL {
|
||||
for &i in wide.iter().chain(&dual) {
|
||||
let (x, _, w, _) = rects[i];
|
||||
let x_lo = grid_coord(x - tolerance, cell);
|
||||
let x_hi = grid_coord(x + w + tolerance, cell);
|
||||
union_rect_against_bands(&mut uf, rects, i, &large_x, x_lo, x_hi, tolerance);
|
||||
}
|
||||
for &i in &dual {
|
||||
let (_, y, _, h) = rects[i];
|
||||
let y_lo = grid_coord(y - tolerance, cell);
|
||||
let y_hi = grid_coord(y + h + tolerance, cell);
|
||||
union_rect_against_bands(&mut uf, rects, i, &large_y, y_lo, y_hi, tolerance);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3810,6 +4023,113 @@ mod tests {
|
||||
assert_eq!(groups[0].len(), 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_cluster_rects_overlapping_grid_still_clusters() {
|
||||
// Neighboring cells overlap; the grid must still union the whole table.
|
||||
let mut rects = Vec::new();
|
||||
for row in 0..4 {
|
||||
for col in 0..4 {
|
||||
rects.push((col as f32 * 9.0, row as f32 * 9.0, 10.0, 10.0));
|
||||
}
|
||||
}
|
||||
let groups = cluster_rects(&rects, 0.0, 1);
|
||||
assert_eq!(groups.len(), 1);
|
||||
assert_eq!(groups[0].len(), 16);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_cluster_rects_many_disjoint_stays_subquadratic() {
|
||||
// Pairwise-disjoint rects never merge, so a component-size cap does
|
||||
// not stop all-pairs overlap tests. Spread in X so they land in
|
||||
// different grid cells; 8k is enough that n² tests would dominate.
|
||||
let n = 8_000usize;
|
||||
let rects: Vec<(f32, f32, f32, f32)> =
|
||||
(0..n).map(|i| (i as f32 * 20.0, 0.0, 10.0, 10.0)).collect();
|
||||
let groups = cluster_rects(&rects, 0.0, 2);
|
||||
assert!(groups.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_cluster_rects_stacked_disjoint_does_not_starve_later_table() {
|
||||
// Same X, spread in Y: a spatial grid must still union an overlapping
|
||||
// pair in another region of the page.
|
||||
let n = 8_000usize;
|
||||
let mut rects: Vec<(f32, f32, f32, f32)> =
|
||||
(0..n).map(|i| (0.0, i as f32 * 20.0, 10.0, 10.0)).collect();
|
||||
rects.push((500.0, 0.0, 10.0, 10.0));
|
||||
rects.push((508.0, 0.0, 10.0, 10.0));
|
||||
let groups = cluster_rects(&rects, 0.0, 2);
|
||||
assert_eq!(groups.len(), 1);
|
||||
assert_eq!(groups[0].len(), 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_cluster_rects_oversized_span_still_unions() {
|
||||
// Wider than 64 grid cells; must still union the small overlapping rect.
|
||||
let rects = vec![(0.0, 0.0, 5000.0, 10.0), (4900.0, 0.0, 10.0, 10.0)];
|
||||
let groups = cluster_rects(&rects, 0.0, 1);
|
||||
assert_eq!(groups.len(), 1);
|
||||
assert_eq!(groups[0].len(), 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_cluster_rects_many_oversized_spans_all_get_a_pass() {
|
||||
// More than 32 huge rects: the last one must still union its overlap.
|
||||
let mut rects: Vec<(f32, f32, f32, f32)> = (0..40)
|
||||
.map(|i| (0.0, i as f32 * 20.0, 5000.0, 10.0))
|
||||
.collect();
|
||||
rects.push((4900.0, 39.0 * 20.0, 10.0, 10.0));
|
||||
let groups = cluster_rects(&rects, 0.0, 2);
|
||||
assert_eq!(groups.len(), 1);
|
||||
assert_eq!(groups[0].len(), 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_cluster_rects_oversized_not_starved_by_earlier_disjoint() {
|
||||
// 9k earlier disjoint drawings would exhaust an index-order cap of
|
||||
// 8,192 before the overlapping cell is visited.
|
||||
let mut rects: Vec<(f32, f32, f32, f32)> = (0..9_000)
|
||||
.map(|i| (10_000.0, i as f32 * 20.0, 10.0, 10.0))
|
||||
.collect();
|
||||
let wide = rects.len();
|
||||
rects.push((0.0, 0.0, 5000.0, 10.0));
|
||||
let target = rects.len();
|
||||
rects.push((4900.0, 0.0, 10.0, 10.0));
|
||||
let groups = cluster_rects(&rects, 0.0, 2);
|
||||
assert!(
|
||||
groups
|
||||
.iter()
|
||||
.any(|g| g.contains(&wide) && g.contains(&target)),
|
||||
"wide rule and far-end cell must share a cluster"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_cluster_rects_wide_and_tall_oversized_union() {
|
||||
let rects = vec![(0.0, 0.0, 5000.0, 10.0), (0.0, 0.0, 10.0, 5000.0)];
|
||||
let groups = cluster_rects(&rects, 0.0, 2);
|
||||
assert_eq!(groups.len(), 1);
|
||||
assert_eq!(groups[0].len(), 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_cluster_rects_dual_oversized_spans_coarse_y() {
|
||||
let rects = vec![(0.0, 0.0, 5000.0, 5000.0), (0.0, 4500.0, 5000.0, 5000.0)];
|
||||
let groups = cluster_rects(&rects, 0.0, 2);
|
||||
assert_eq!(groups.len(), 1);
|
||||
assert_eq!(groups[0].len(), 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_cluster_rects_many_wide_and_tall_stays_subquadratic() {
|
||||
let mut rects = Vec::with_capacity(4_000);
|
||||
for i in 0..2_000 {
|
||||
rects.push((0.0, i as f32 * 20.0, 5000.0, 10.0));
|
||||
rects.push((i as f32 * 20.0, 0.0, 10.0, 5000.0));
|
||||
}
|
||||
let _groups = cluster_rects(&rects, 0.0, 2);
|
||||
}
|
||||
|
||||
// --- snap_edges ---
|
||||
|
||||
#[test]
|
||||
|
||||
+230
-30
@@ -540,18 +540,14 @@ impl ToUnicodeCMap {
|
||||
|
||||
/// Remap a CMap that references pre-subsetting GIDs to sequential post-subsetting GIDs.
|
||||
/// Collects all source CIDs, sorts them, and reassigns to 1, 2, 3, ...
|
||||
///
|
||||
/// Range expansion stops after `MAX_CID_W_EXPANSION` CID visits, counting
|
||||
/// overwrites, so repeated full-width `bfrange`s cannot re-expand the
|
||||
/// 16-bit domain. Later overlapping ranges that would have introduced new
|
||||
/// CIDs after that many visits are truncated.
|
||||
pub fn remap_to_sequential(&self) -> ToUnicodeCMap {
|
||||
let mut cid_to_unicode: HashMap<u16, String> = HashMap::new();
|
||||
|
||||
// Expand ranges first
|
||||
for &(start, end, base) in &self.ranges {
|
||||
for cid in start..=end {
|
||||
let unicode_cp = base + (cid - start) as u32;
|
||||
if let Some(ch) = char::from_u32(unicode_cp) {
|
||||
cid_to_unicode.insert(cid, ch.to_string());
|
||||
}
|
||||
}
|
||||
}
|
||||
expand_bfranges_for_remap(&self.ranges, &mut cid_to_unicode, MAX_CID_W_EXPANSION);
|
||||
|
||||
// char_map entries override range entries
|
||||
for (&cid, unicode) in &self.char_map {
|
||||
@@ -576,6 +572,33 @@ impl ToUnicodeCMap {
|
||||
}
|
||||
}
|
||||
|
||||
/// Expand `bfrange` entries into individual CID→Unicode inserts.
|
||||
/// Returns how many CIDs were visited. Counts overwrites so a repeated
|
||||
/// full-width range cannot keep working after `max_assignments`.
|
||||
fn expand_bfranges_for_remap(
|
||||
ranges: &[(u16, u16, u32)],
|
||||
cid_to_unicode: &mut HashMap<u16, String>,
|
||||
max_assignments: usize,
|
||||
) -> usize {
|
||||
let mut assigned = 0usize;
|
||||
'ranges: for &(start, end, base) in ranges {
|
||||
if start > end {
|
||||
continue;
|
||||
}
|
||||
for cid in start..=end {
|
||||
if assigned >= max_assignments {
|
||||
break 'ranges;
|
||||
}
|
||||
assigned += 1;
|
||||
let unicode_cp = base + (cid - start) as u32;
|
||||
if let Some(ch) = char::from_u32(unicode_cp) {
|
||||
cid_to_unicode.insert(cid, ch.to_string());
|
||||
}
|
||||
}
|
||||
}
|
||||
assigned
|
||||
}
|
||||
|
||||
/// Parse a hex string to u16
|
||||
fn parse_hex_u16(hex: &str) -> Option<u16> {
|
||||
u16::from_str_radix(hex.trim(), 16).ok()
|
||||
@@ -594,7 +617,7 @@ fn hex_to_unicode_string(hex: &str) -> Option<String> {
|
||||
|
||||
let bytes: Option<Vec<u8>> = (0..hex.len())
|
||||
.step_by(2)
|
||||
.map(|i| u8::from_str_radix(&hex[i..i + 2], 16).ok())
|
||||
.map(|i| u8::from_str_radix(hex.get(i..i + 2)?, 16).ok())
|
||||
.collect();
|
||||
let bytes = bytes?;
|
||||
|
||||
@@ -1561,23 +1584,31 @@ fn parse_encoding_cmap_stream(data: &[u8]) -> Option<EncodingCMap> {
|
||||
}
|
||||
|
||||
let mut map = HashMap::new();
|
||||
let mut assigned = 0usize;
|
||||
let mut pos = 0;
|
||||
while let Some(start) = text[pos..].find("begincidchar") {
|
||||
let section_start = pos + start + "begincidchar".len();
|
||||
if let Some(end) = text[section_start..].find("endcidchar") {
|
||||
let section = &text[section_start..section_start + end];
|
||||
parse_cidchar_section(section, &mut map, &mut src_hex_lengths);
|
||||
if !parse_cidchar_section(section, &mut map, &mut src_hex_lengths, &mut assigned) {
|
||||
break;
|
||||
}
|
||||
pos = section_start + end;
|
||||
} else {
|
||||
break;
|
||||
}
|
||||
}
|
||||
pos = 0;
|
||||
while let Some(start) = text[pos..].find("begincidrange") {
|
||||
while assigned < MAX_CID_W_EXPANSION {
|
||||
let Some(start) = text[pos..].find("begincidrange") else {
|
||||
break;
|
||||
};
|
||||
let section_start = pos + start + "begincidrange".len();
|
||||
if let Some(end) = text[section_start..].find("endcidrange") {
|
||||
let section = &text[section_start..section_start + end];
|
||||
parse_cidrange_section(section, &mut map, &mut src_hex_lengths);
|
||||
if !parse_cidrange_section(section, &mut map, &mut src_hex_lengths, &mut assigned) {
|
||||
break;
|
||||
}
|
||||
pos = section_start + end;
|
||||
} else {
|
||||
break;
|
||||
@@ -1612,7 +1643,8 @@ fn parse_cidchar_section(
|
||||
section: &str,
|
||||
map: &mut HashMap<u16, u16>,
|
||||
src_hex_lengths: &mut Vec<usize>,
|
||||
) {
|
||||
assigned: &mut usize,
|
||||
) -> bool {
|
||||
let mut chars = section.chars().peekable();
|
||||
loop {
|
||||
while chars.peek().is_some_and(|c| c.is_whitespace()) {
|
||||
@@ -1643,16 +1675,20 @@ fn parse_cidchar_section(
|
||||
}
|
||||
}
|
||||
if let (Some(code), Ok(cid)) = (parse_hex_u16(&src_hex), cid_str.parse::<u16>()) {
|
||||
map.insert(code, cid);
|
||||
if !assign_encoding_cid(map, code, cid, assigned) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
}
|
||||
true
|
||||
}
|
||||
|
||||
fn parse_cidrange_section(
|
||||
section: &str,
|
||||
map: &mut HashMap<u16, u16>,
|
||||
src_hex_lengths: &mut Vec<usize>,
|
||||
) {
|
||||
assigned: &mut usize,
|
||||
) -> bool {
|
||||
let mut chars = section.chars().peekable();
|
||||
loop {
|
||||
while chars.peek().is_some_and(|c| c.is_whitespace()) {
|
||||
@@ -1703,12 +1739,34 @@ fn parse_cidrange_section(
|
||||
) else {
|
||||
continue;
|
||||
};
|
||||
if start > end {
|
||||
continue;
|
||||
}
|
||||
let mut cid = start_cid;
|
||||
for code in start..=end {
|
||||
map.insert(code, cid);
|
||||
if !assign_encoding_cid(map, code, cid, assigned) {
|
||||
return false;
|
||||
}
|
||||
cid = cid.saturating_add(1);
|
||||
}
|
||||
}
|
||||
true
|
||||
}
|
||||
|
||||
fn assign_encoding_cid(
|
||||
map: &mut HashMap<u16, u16>,
|
||||
code: u16,
|
||||
cid: u16,
|
||||
assigned: &mut usize,
|
||||
) -> bool {
|
||||
// Count overwrites: unique-key coverage alone would not stop a repeated
|
||||
// full-width range from re-inserting all 65,536 codes.
|
||||
if *assigned >= MAX_CID_W_EXPANSION {
|
||||
return false;
|
||||
}
|
||||
map.insert(code, cid);
|
||||
*assigned += 1;
|
||||
true
|
||||
}
|
||||
|
||||
fn parse_binary_cmap_encoding(data: &[u8]) -> Result<EncodingCMap, String> {
|
||||
@@ -1832,6 +1890,13 @@ fn merge_cmaps(mut base: ToUnicodeCMap, overlay: ToUnicodeCMap) -> ToUnicodeCMap
|
||||
base
|
||||
}
|
||||
|
||||
/// Shared 16-bit CID expansion cap (65,536).
|
||||
/// Encoding `begincidrange`, `/W` width assignment, and ToUnicode sequential
|
||||
/// remap count every insert, including overwrites, so a repeated full-width
|
||||
/// range cannot keep working after the domain is filled. The `/W` unicode
|
||||
/// heuristic caps unique CIDs with the same number.
|
||||
pub(crate) const MAX_CID_W_EXPANSION: usize = 65_536;
|
||||
|
||||
/// Check if a CIDFont's /W (widths) array contains CID values that look like
|
||||
/// Unicode codepoints rather than low-value GIDs.
|
||||
///
|
||||
@@ -1843,20 +1908,23 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
|
||||
_ => return false,
|
||||
};
|
||||
|
||||
// The /W array format: [cid [w1 w2 ...]] or [cid_start cid_end w]
|
||||
// We extract all CID values (the first element of each group).
|
||||
let mut cids: Vec<u16> = Vec::new();
|
||||
// The /W array format: [cid [w1 w2 ...]] or [cid_start cid_end w].
|
||||
// Collect unique CIDs only: repeating a full-width range must not grow a
|
||||
// temporary vector (or the sort) with the range length on every copy.
|
||||
let mut seen = HashSet::new();
|
||||
let mut i = 0;
|
||||
while i < w_arr.len() {
|
||||
while i < w_arr.len() && seen.len() < MAX_CID_W_EXPANSION {
|
||||
if let Ok(cid) = w_arr[i].as_i64() {
|
||||
cids.push(cid as u16);
|
||||
// Skip the width data
|
||||
let start = cid as u16;
|
||||
if i + 1 < w_arr.len() {
|
||||
match &w_arr[i + 1] {
|
||||
Object::Array(widths) => {
|
||||
// [cid [w1 w2 ...]] — CIDs are cid, cid+1, ..., cid+len-1
|
||||
for j in 1..widths.len() {
|
||||
cids.push((cid as u16).wrapping_add(j as u16));
|
||||
for j in 0..widths.len() {
|
||||
if seen.len() >= MAX_CID_W_EXPANSION {
|
||||
break;
|
||||
}
|
||||
seen.insert(start.wrapping_add(j as u16));
|
||||
}
|
||||
i += 2;
|
||||
}
|
||||
@@ -1864,9 +1932,7 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
|
||||
// [cid_start cid_end w] — range of CIDs
|
||||
if i + 2 < w_arr.len() {
|
||||
if let Ok(cid_end) = w_arr[i + 1].as_i64() {
|
||||
for c in (cid as u16)..=(cid_end as u16) {
|
||||
cids.push(c);
|
||||
}
|
||||
record_unique_cid_range(start, cid_end as u16, &mut seen);
|
||||
}
|
||||
i += 3;
|
||||
} else {
|
||||
@@ -1875,6 +1941,7 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
|
||||
}
|
||||
}
|
||||
} else {
|
||||
seen.insert(start);
|
||||
i += 1;
|
||||
}
|
||||
} else {
|
||||
@@ -1882,10 +1949,11 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
|
||||
}
|
||||
}
|
||||
|
||||
if cids.is_empty() {
|
||||
if seen.is_empty() {
|
||||
return false;
|
||||
}
|
||||
|
||||
let mut cids: Vec<u16> = seen.into_iter().collect();
|
||||
cids.sort_unstable();
|
||||
let median = cids[cids.len() / 2];
|
||||
// Unicode text CIDs are typically >= 0x20 (space) with letters at 0x41+.
|
||||
@@ -1894,6 +1962,18 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
|
||||
median >= 0x41
|
||||
}
|
||||
|
||||
fn record_unique_cid_range(start: u16, end: u16, seen: &mut HashSet<u16>) {
|
||||
if start > end {
|
||||
return;
|
||||
}
|
||||
for cid in start..=end {
|
||||
if seen.len() >= MAX_CID_W_EXPANSION {
|
||||
return;
|
||||
}
|
||||
seen.insert(cid);
|
||||
}
|
||||
}
|
||||
|
||||
/// Build a ToUnicodeCMap from predefined CID→Unicode mapping based on CIDSystemInfo.
|
||||
///
|
||||
/// Supports Adobe-Korea1 (Korean) character collection. Can be extended for
|
||||
@@ -2606,6 +2686,24 @@ endcmap
|
||||
assert_eq!(cmap.lookup(0x0025), Some("B".to_string()));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_hex_to_unicode_non_ascii_no_panic() {
|
||||
// A destination containing a multi-byte char makes the byte length even
|
||||
// while a byte offset can land inside a char. Slicing must not panic;
|
||||
// it should be rejected gracefully.
|
||||
assert_eq!(hex_to_unicode_string("XéY"), None);
|
||||
assert_eq!(hex_to_unicode_string("\u{fffd}0"), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_parse_bfchar_non_ascii_destination_no_panic() {
|
||||
// Crafted /ToUnicode CMap: a non-hex, non-ASCII destination previously
|
||||
// triggered a char-boundary panic in hex_to_unicode_string.
|
||||
let cmap_content = "beginbfchar <0041> <XéY> endbfchar";
|
||||
// Must not panic; the malformed entry is simply skipped.
|
||||
let _ = ToUnicodeCMap::parse(cmap_content.as_bytes());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_parse_bfchar_1byte() {
|
||||
// This is the pattern that caused the CJK bug: codespace is <0000><FFFF>
|
||||
@@ -2849,6 +2947,33 @@ endbfrange
|
||||
assert!(remapped.ranges.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn remap_to_sequential_repeated_full_bfranges_stay_bounded() {
|
||||
// 5,000 copies of `<0003> <ffff>` must stop after 65,536 CID visits,
|
||||
// not 5,000 × ~65,533 expansions.
|
||||
let ranges = vec![(3u16, 65535u16, 0x41u32); 5_000];
|
||||
let mut map = std::collections::HashMap::new();
|
||||
let assigned = expand_bfranges_for_remap(&ranges, &mut map, MAX_CID_W_EXPANSION);
|
||||
assert_eq!(assigned, MAX_CID_W_EXPANSION);
|
||||
assert!(map.len() <= MAX_CID_W_EXPANSION);
|
||||
|
||||
let mut body = String::new();
|
||||
let mut remaining = 5_000usize;
|
||||
while remaining > 0 {
|
||||
let n = remaining.min(100);
|
||||
body.push_str(&format!("{n} beginbfrange\n"));
|
||||
for _ in 0..n {
|
||||
body.push_str("<0003> <ffff> <0041>\n");
|
||||
}
|
||||
body.push_str("endbfrange\n");
|
||||
remaining -= n;
|
||||
}
|
||||
let data = format!("1 begincodespacerange\n<0000> <ffff>\nendcodespacerange\n{body}");
|
||||
let cmap = ToUnicodeCMap::parse(data.as_bytes()).unwrap();
|
||||
let remapped = cmap.remap_to_sequential();
|
||||
assert_eq!(remapped.lookup(1), Some("A".to_string()));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_min_source_cid() {
|
||||
let cmap_content = r#"
|
||||
@@ -3280,4 +3405,79 @@ endbfrange
|
||||
"An indirect /Subtype naming CIDFontType2 must still reach the remap"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cid_values_look_like_unicode_letter_range() {
|
||||
let mut dict = lopdf::Dictionary::new();
|
||||
dict.set(
|
||||
"W",
|
||||
Object::Array(vec![
|
||||
Object::Integer(0x41),
|
||||
Object::Integer(0x5A),
|
||||
Object::Integer(500),
|
||||
]),
|
||||
);
|
||||
assert!(cid_values_look_like_unicode(&dict));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cid_values_look_like_unicode_low_gids() {
|
||||
let mut dict = lopdf::Dictionary::new();
|
||||
dict.set(
|
||||
"W",
|
||||
Object::Array(vec![
|
||||
Object::Integer(0),
|
||||
Object::Array(vec![Object::Integer(500); 10]),
|
||||
]),
|
||||
);
|
||||
assert!(!cid_values_look_like_unicode(&dict));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cid_values_look_like_unicode_repeated_full_ranges_stay_bounded() {
|
||||
// Repeating `[0 65535 w]` must not materialize 65,536 CIDs per copy.
|
||||
let mut w = Vec::new();
|
||||
for _ in 0..5_000 {
|
||||
w.push(Object::Integer(0));
|
||||
w.push(Object::Integer(65535));
|
||||
w.push(Object::Integer(500));
|
||||
}
|
||||
let mut dict = lopdf::Dictionary::new();
|
||||
dict.set("W", Object::Array(w));
|
||||
assert!(cid_values_look_like_unicode(&dict));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn encoding_cidrange_maps_a_normal_range() {
|
||||
let data = b"1 begincodespacerange\n<0000> <FFFF>\nendcodespacerange\n\
|
||||
1 begincidrange\n<0041> <0043> 65\nendcidrange\n";
|
||||
let enc = parse_encoding_cmap_stream(data).unwrap();
|
||||
assert_eq!(enc.map.get(&0x41), Some(&65));
|
||||
assert_eq!(enc.map.get(&0x42), Some(&66));
|
||||
assert_eq!(enc.map.get(&0x43), Some(&67));
|
||||
assert_eq!(enc.map.len(), 3);
|
||||
assert_eq!(enc.code_byte_length, 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn encoding_cidrange_repeated_full_ranges_stay_bounded() {
|
||||
// 5,000 copies of `<0000> <ffff> 0` must not re-expand the 16-bit
|
||||
// domain on every declaration.
|
||||
let mut body = String::new();
|
||||
let mut remaining = 5_000usize;
|
||||
while remaining > 0 {
|
||||
let n = remaining.min(100);
|
||||
body.push_str(&format!("{n} begincidrange\n"));
|
||||
for _ in 0..n {
|
||||
body.push_str("<0000> <ffff> 0\n");
|
||||
}
|
||||
body.push_str("endcidrange\n");
|
||||
remaining -= n;
|
||||
}
|
||||
let data = format!("1 begincodespacerange\n<0000> <FFFF>\nendcodespacerange\n{body}");
|
||||
let enc = parse_encoding_cmap_stream(data.as_bytes()).unwrap();
|
||||
assert!(enc.map.len() <= MAX_CID_W_EXPANSION);
|
||||
assert_eq!(enc.map.get(&0), Some(&0));
|
||||
assert_eq!(enc.map.get(&65535), Some(&65535));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1429,6 +1429,77 @@ fn test_firecrawl_tagged_pdf_struct_tree() {
|
||||
assert_eq!(fence_count % 2, 0, "Code fences should be balanced");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_tagged_pdf_text_items_carry_mcid() {
|
||||
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
|
||||
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
|
||||
assert!(
|
||||
items.iter().any(|i| i.mcid.is_some()),
|
||||
"Tagged PDF text items should carry Marked Content IDs"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_structure_elements_tagged_pdf() {
|
||||
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
|
||||
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
|
||||
assert!(!elements.is_empty(), "Tagged PDF should yield elements");
|
||||
assert!(
|
||||
elements.iter().any(|e| e.role == "H1"),
|
||||
"Should surface H1 heading roles"
|
||||
);
|
||||
assert!(
|
||||
elements.iter().all(|e| !e.role.is_empty()),
|
||||
"Every element should carry a role name"
|
||||
);
|
||||
|
||||
// Sorted by (page, mcid) for deterministic output
|
||||
assert!(
|
||||
elements
|
||||
.windows(2)
|
||||
.all(|w| (w[0].page, w[0].mcid) <= (w[1].page, w[1].mcid)),
|
||||
"Elements should be sorted by (page, mcid)"
|
||||
);
|
||||
|
||||
// The advertised join: (page, mcid) pairs must line up with the
|
||||
// mcid-carrying TextItems from positioned extraction, and joining the
|
||||
// H1 entries must recover non-empty heading text.
|
||||
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
|
||||
let h1_refs: std::collections::HashSet<(u32, i64)> = elements
|
||||
.iter()
|
||||
.filter(|e| e.role == "H1")
|
||||
.map(|e| (e.page, e.mcid))
|
||||
.collect();
|
||||
let h1_text: String = items
|
||||
.iter()
|
||||
.filter(|i| i.mcid.is_some_and(|mcid| h1_refs.contains(&(i.page, mcid))))
|
||||
.map(|i| i.text.as_str())
|
||||
.collect();
|
||||
assert!(
|
||||
!h1_text.trim().is_empty(),
|
||||
"Joining H1 structure elements to text items should recover heading text"
|
||||
);
|
||||
|
||||
// Page filter is 1-indexed (matching TextItem.page) and equals the
|
||||
// corresponding subset of the full document result.
|
||||
let page1 = pdf_inspector::extract_structure_elements_mem(&buf, Some(&[1])).unwrap();
|
||||
assert!(!page1.is_empty(), "Page 1 should have elements");
|
||||
assert!(page1.iter().all(|e| e.page == 1));
|
||||
let full_page1_count = elements.iter().filter(|e| e.page == 1).count();
|
||||
assert_eq!(page1.len(), full_page1_count);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_structure_elements_untagged_pdf_empty() {
|
||||
let buf = std::fs::read("tests/fixtures/thermo-freon12.pdf").unwrap();
|
||||
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
|
||||
assert!(
|
||||
elements.is_empty(),
|
||||
"Untagged PDF should yield no structure elements, got {:?}",
|
||||
elements
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_identity_h_no_tounicode_suppresses_garbage() {
|
||||
// shinagawa_identity_h.pdf uses YuGothic with Identity-H encoding and no
|
||||
@@ -1583,6 +1654,282 @@ fn test_extract_regions_mem_basic_text_pdf() {
|
||||
assert_eq!(regions[0].page, 0);
|
||||
}
|
||||
|
||||
/// Build a synthetic "scanned page" PDF: a full-page image XObject with a
|
||||
/// text layer drawn in the given render mode (3 = invisible OCR overlay,
|
||||
/// 0 = normal visible fill). `visible_extra` optionally adds a normally
|
||||
/// rendered line so double-layer behavior can be tested; `layer_lines`
|
||||
/// overrides the layer content (default: three pangram lines);
|
||||
/// `quote_ops` shows every layer line via the `'` operator instead of Tj
|
||||
/// (both are standard show-text encodings for OCR layers).
|
||||
fn make_pdf_with_custom_text_layer(
|
||||
text_render_mode: i32,
|
||||
visible_extra: Option<&str>,
|
||||
layer_lines: Option<&[&str]>,
|
||||
quote_ops: bool,
|
||||
) -> Vec<u8> {
|
||||
let mut pdf = b"%PDF-1.4\n".to_vec();
|
||||
let mut offsets = vec![0usize];
|
||||
|
||||
fn add_object(pdf: &mut Vec<u8>, offsets: &mut Vec<usize>, id: usize, body: &str) {
|
||||
offsets.push(pdf.len());
|
||||
pdf.extend_from_slice(format!("{id} 0 obj\n").as_bytes());
|
||||
pdf.extend_from_slice(body.as_bytes());
|
||||
pdf.extend_from_slice(b"\nendobj\n");
|
||||
}
|
||||
fn add_stream_object(
|
||||
pdf: &mut Vec<u8>,
|
||||
offsets: &mut Vec<usize>,
|
||||
id: usize,
|
||||
dict: &str,
|
||||
stream_bytes: &[u8],
|
||||
) {
|
||||
offsets.push(pdf.len());
|
||||
pdf.extend_from_slice(format!("{id} 0 obj\n").as_bytes());
|
||||
pdf.extend_from_slice(
|
||||
format!("<< {} /Length {} >>\nstream\n", dict, stream_bytes.len()).as_bytes(),
|
||||
);
|
||||
pdf.extend_from_slice(stream_bytes);
|
||||
pdf.extend_from_slice(b"\nendstream\nendobj\n");
|
||||
}
|
||||
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
1,
|
||||
"<< /Type /Catalog /Pages 2 0 R >>",
|
||||
);
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
2,
|
||||
"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
|
||||
);
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
3,
|
||||
"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] \
|
||||
/Resources << /Font << /F1 5 0 R >> /XObject << /Im0 6 0 R >> >> \
|
||||
/Contents 4 0 R >>",
|
||||
);
|
||||
// Full-page raster, then the text layer in the requested render mode —
|
||||
// several lines so the OCR-layer gate's alnum floor (40) is well cleared.
|
||||
let mut content = String::from("q 612 0 0 792 0 0 cm /Im0 Do Q\n");
|
||||
let default_layer = [
|
||||
"The quick brown fox jumps over the lazy dog",
|
||||
"Pack my box with five dozen liquor jugs tonight",
|
||||
"Sphinx of black quartz judge my vow carefully",
|
||||
];
|
||||
let layer: &[&str] = layer_lines.unwrap_or(&default_layer);
|
||||
if quote_ops {
|
||||
// Every line shown via `'` (move-to-next-line + show) — nothing on
|
||||
// this layer goes through Tj, pinning the `'` suppression path.
|
||||
content.push_str(&format!(
|
||||
"BT /F1 12 Tf {text_render_mode} Tr 16 TL 72 716 Td "
|
||||
));
|
||||
for line in layer {
|
||||
content.push_str(&format!("({line}) ' "));
|
||||
}
|
||||
} else {
|
||||
content.push_str(&format!("BT /F1 12 Tf {text_render_mode} Tr 72 700 Td "));
|
||||
for (i, line) in layer.iter().enumerate() {
|
||||
if i > 0 {
|
||||
content.push_str("0 -16 Td ");
|
||||
}
|
||||
content.push_str(&format!("({line}) Tj "));
|
||||
}
|
||||
}
|
||||
content.push_str("ET\n");
|
||||
if let Some(extra) = visible_extra {
|
||||
content.push_str(&format!("BT /F1 12 Tf 0 Tr 72 500 Td ({extra}) Tj ET\n"));
|
||||
}
|
||||
add_stream_object(&mut pdf, &mut offsets, 4, "", content.as_bytes());
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
5,
|
||||
"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica >>",
|
||||
);
|
||||
let image_pixel = [128u8];
|
||||
add_stream_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
6,
|
||||
"/Type /XObject /Subtype /Image /Width 1 /Height 1 \
|
||||
/ColorSpace /DeviceGray /BitsPerComponent 8",
|
||||
&image_pixel,
|
||||
);
|
||||
|
||||
let xref_start = pdf.len();
|
||||
pdf.extend_from_slice(format!("xref\n0 {}\n", offsets.len()).as_bytes());
|
||||
pdf.extend_from_slice(b"0000000000 65535 f \n");
|
||||
for offset in offsets.iter().skip(1) {
|
||||
pdf.extend_from_slice(format!("{offset:010} 00000 n \n").as_bytes());
|
||||
}
|
||||
pdf.extend_from_slice(
|
||||
format!(
|
||||
"trailer\n<< /Size {} /Root 1 0 R >>\nstartxref\n{}\n%%EOF",
|
||||
offsets.len(),
|
||||
xref_start
|
||||
)
|
||||
.as_bytes(),
|
||||
);
|
||||
pdf
|
||||
}
|
||||
|
||||
fn make_pdf_with_text_layer(text_render_mode: i32, visible_extra: Option<&str>) -> Vec<u8> {
|
||||
make_pdf_with_custom_text_layer(text_render_mode, visible_extra, None, false)
|
||||
}
|
||||
|
||||
/// A scanned page whose only text is an invisible (Tr 3) OCR layer behind
|
||||
/// the raster must serve that layer from the region extractor instead of
|
||||
/// reporting the region as needs_ocr — the exact text is already in the PDF.
|
||||
#[test]
|
||||
fn test_extract_regions_mem_recovers_invisible_ocr_layer() {
|
||||
let buf = make_pdf_with_text_layer(3, None);
|
||||
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
|
||||
assert_eq!(regions.len(), 1);
|
||||
let region = ®ions[0].regions[0];
|
||||
assert!(
|
||||
region.text.contains("quick brown fox"),
|
||||
"invisible OCR layer should be served as region text, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
assert!(
|
||||
!region.needs_ocr,
|
||||
"recovered OCR layer must not fall back to GPU OCR"
|
||||
);
|
||||
}
|
||||
|
||||
/// ANY visible text on the page — even a single short line — must block the
|
||||
/// invisible-layer adoption entirely: the invisible pass returns visible
|
||||
/// items too, so adopting it alongside visible text would duplicate the
|
||||
/// visible words. Strict zero-visible gate, no fuzzy dedupe.
|
||||
#[test]
|
||||
fn test_extract_regions_mem_visible_text_blocks_invisible_layer() {
|
||||
let buf = make_pdf_with_text_layer(3, Some("Folio 142"));
|
||||
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
|
||||
let region = ®ions[0].regions[0];
|
||||
assert!(
|
||||
region.text.contains("Folio 142"),
|
||||
"visible text should be extracted, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
assert!(
|
||||
!region.text.contains("quick brown fox"),
|
||||
"invisible layer must not be adopted when any visible text exists, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
assert_eq!(
|
||||
region.text.matches("Folio 142").count(),
|
||||
1,
|
||||
"visible text must appear exactly once, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
}
|
||||
|
||||
/// An invisible OCR layer shown entirely via the `'` show-text operator
|
||||
/// (move-to-next-line + show) must also be recovered — the skipped_invisible
|
||||
/// signal has to fire on every show-text path, not just Tj/TJ.
|
||||
#[test]
|
||||
fn test_extract_regions_mem_recovers_quote_operator_layer() {
|
||||
let buf = make_pdf_with_custom_text_layer(3, None, None, true);
|
||||
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
|
||||
let region = ®ions[0].regions[0];
|
||||
assert!(
|
||||
region.text.contains("quick brown fox"),
|
||||
"'-operator OCR layer should be recovered, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
assert!(!region.needs_ocr);
|
||||
}
|
||||
|
||||
/// An invisible layer below the 40-alnum floor (a stray watermark line)
|
||||
/// must NOT be adopted — the region keeps its needs_ocr fallback.
|
||||
#[test]
|
||||
fn test_extract_regions_mem_tiny_invisible_layer_not_adopted() {
|
||||
let buf = make_pdf_with_custom_text_layer(3, None, Some(&["Scanned by ACME"]), false);
|
||||
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
|
||||
let region = ®ions[0].regions[0];
|
||||
assert!(
|
||||
!region.text.contains("Scanned by ACME"),
|
||||
"below-floor invisible layer must not be adopted, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
// Only the raster placeholder remains — needs_ocr stays whatever main
|
||||
// reports for placeholder-only regions (false today; downstream
|
||||
// pipelines route placeholder-only text to OCR themselves, and this PR
|
||||
// deliberately does not change that contract).
|
||||
assert!(
|
||||
region.text.trim().starts_with("[Image:"),
|
||||
"region should hold only the raster placeholder, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
}
|
||||
|
||||
/// An invisible layer that clears the alnum floor but is mostly symbol
|
||||
/// garbage (a broken OCR run) must be rejected by the garbage gate.
|
||||
#[test]
|
||||
fn test_extract_regions_mem_garbage_invisible_layer_not_adopted() {
|
||||
// Each line: 5 alphanumerics among 15 symbol chars. Ten lines clear the
|
||||
// 40-alnum floor (50 alnum) while staying well under the half-alnum
|
||||
// ratio is_garbage_text requires.
|
||||
let garbage_lines: Vec<&str> = vec!["a@@b%%c&&d==e~~"; 10];
|
||||
let buf = make_pdf_with_custom_text_layer(3, None, Some(&garbage_lines), false);
|
||||
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
|
||||
let region = ®ions[0].regions[0];
|
||||
assert!(
|
||||
!region.text.contains("a@@b"),
|
||||
"garbage invisible layer must not be adopted, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
assert!(
|
||||
region.text.trim().starts_with("[Image:"),
|
||||
"region should hold only the raster placeholder, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
}
|
||||
|
||||
/// Punctuation-only visible text (zero alphanumerics) must ALSO block
|
||||
/// adoption — the gate is item-presence, not alphanumeric mass. (Real-world
|
||||
/// rationale: an invisible OCR layer transcribes the raster, so visible
|
||||
/// glyphs typically have invisible twins there; this fixture's layers are
|
||||
/// disjoint, so it pins the gate itself, not the duplication scenario.)
|
||||
#[test]
|
||||
fn test_extract_regions_mem_punctuation_visible_blocks_invisible_layer() {
|
||||
let buf = make_pdf_with_text_layer(3, Some("... --- ..."));
|
||||
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
|
||||
let region = ®ions[0].regions[0];
|
||||
assert!(
|
||||
!region.text.contains("quick brown fox"),
|
||||
"invisible layer must not be adopted over punctuation-only visible text, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
assert_eq!(
|
||||
region.text.matches("... --- ...").count(),
|
||||
1,
|
||||
"visible punctuation must be preserved exactly once, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
}
|
||||
|
||||
/// Regression guard: a normal visible-text page (render mode 0) is served
|
||||
/// once and only once — if the fallback ever mis-fired here and merged a
|
||||
/// second pass, the phrase would duplicate.
|
||||
#[test]
|
||||
fn test_extract_regions_mem_visible_layer_unchanged() {
|
||||
let buf = make_pdf_with_text_layer(0, None);
|
||||
let regions = extract_text_in_regions_mem(&buf, &full_page_regions(1)).unwrap();
|
||||
let region = ®ions[0].regions[0];
|
||||
assert_eq!(
|
||||
region.text.matches("quick brown fox").count(),
|
||||
1,
|
||||
"visible text must appear exactly once, got: {:?}",
|
||||
region.text
|
||||
);
|
||||
assert!(!region.needs_ocr);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_regions_mem_identity_h_needs_ocr() {
|
||||
let buf = std::fs::read("tests/fixtures/shinagawa_identity_h.pdf").unwrap();
|
||||
|
||||
@@ -203,6 +203,79 @@ class TestExtractTextWithPositions:
|
||||
assert len(items) > 0
|
||||
assert all(item.page == 1 for item in items)
|
||||
|
||||
def test_mcid(self):
|
||||
# Untagged fixture: mcid is None or int, never anything else
|
||||
items = pdf_inspector.extract_text_with_positions(
|
||||
fixture_path("thermo-freon12.pdf")
|
||||
)
|
||||
assert all(item.mcid is None or isinstance(item.mcid, int) for item in items)
|
||||
# Tagged fixture: marked content carries MCIDs
|
||||
tagged = pdf_inspector.extract_text_with_positions(
|
||||
fixture_path("firecrawl_docs_tagged.pdf")
|
||||
)
|
||||
assert any(item.mcid is not None for item in tagged)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# extract_structure_elements / extract_structure_elements_bytes
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
class TestExtractStructureElements:
|
||||
def test_tagged_file(self):
|
||||
elements = pdf_inspector.extract_structure_elements(
|
||||
fixture_path("firecrawl_docs_tagged.pdf")
|
||||
)
|
||||
assert len(elements) > 0
|
||||
assert all(isinstance(e.page, int) for e in elements)
|
||||
assert all(isinstance(e.mcid, int) for e in elements)
|
||||
assert all(isinstance(e.role, str) and len(e.role) > 0 for e in elements)
|
||||
assert any(e.role == "H1" for e in elements)
|
||||
|
||||
def test_join_with_text_items(self):
|
||||
# (page, mcid) joins against extract_text_with_positions to recover
|
||||
# heading text
|
||||
path = fixture_path("firecrawl_docs_tagged.pdf")
|
||||
elements = pdf_inspector.extract_structure_elements(path)
|
||||
items = pdf_inspector.extract_text_with_positions(path)
|
||||
h1_refs = {(e.page, e.mcid) for e in elements if e.role == "H1"}
|
||||
h1_text = "".join(
|
||||
item.text
|
||||
for item in items
|
||||
if item.mcid is not None and (item.page, item.mcid) in h1_refs
|
||||
)
|
||||
assert len(h1_text.strip()) > 0
|
||||
|
||||
def test_with_pages(self):
|
||||
# pages filter is 1-indexed, matching TextItem.page
|
||||
elements = pdf_inspector.extract_structure_elements(
|
||||
fixture_path("firecrawl_docs_tagged.pdf"), pages=[1]
|
||||
)
|
||||
assert len(elements) > 0
|
||||
assert all(e.page == 1 for e in elements)
|
||||
|
||||
def test_bytes(self):
|
||||
data = fixture_bytes("firecrawl_docs_tagged.pdf")
|
||||
elements = pdf_inspector.extract_structure_elements_bytes(data)
|
||||
assert len(elements) > 0
|
||||
assert any(e.role == "H1" for e in elements)
|
||||
|
||||
def test_untagged_returns_empty(self):
|
||||
elements = pdf_inspector.extract_structure_elements(
|
||||
fixture_path("thermo-freon12.pdf")
|
||||
)
|
||||
assert elements == []
|
||||
|
||||
def test_repr(self):
|
||||
elements = pdf_inspector.extract_structure_elements(
|
||||
fixture_path("firecrawl_docs_tagged.pdf")
|
||||
)
|
||||
assert "StructureElement" in repr(elements[0])
|
||||
|
||||
def test_not_a_pdf(self):
|
||||
with pytest.raises(ValueError):
|
||||
pdf_inspector.extract_structure_elements_bytes(b"not a pdf")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# extract_text_in_regions / extract_text_in_regions_bytes
|
||||
|
||||
Generated
+2
-2
@@ -724,7 +724,7 @@ checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector"
|
||||
version = "0.1.7"
|
||||
version = "1.14.2"
|
||||
dependencies = [
|
||||
"env_logger",
|
||||
"include_dir",
|
||||
@@ -740,7 +740,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector-wasm"
|
||||
version = "0.1.3"
|
||||
version = "1.14.2"
|
||||
dependencies = [
|
||||
"console_error_panic_hook",
|
||||
"js-sys",
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "pdf-inspector-wasm"
|
||||
version = "0.1.3"
|
||||
version = "1.14.2"
|
||||
edition = "2021"
|
||||
authors = ["Firecrawl Team"]
|
||||
description = "Browser WebAssembly bindings for pdf-inspector"
|
||||
|
||||
Reference in New Issue
Block a user