Compare commits
21
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
0b0ca64ffb | ||
|
|
7054d6aa69 | ||
|
|
a67ee03269 | ||
|
|
965dc65f1b | ||
|
|
36dd5fa426 | ||
|
|
fabec0aec3 | ||
|
|
1f28c00a13 | ||
|
|
f4b8c9e854 | ||
|
|
69039f2728 | ||
|
|
3cca6446bd | ||
|
|
493fed498e | ||
|
|
436af97038 | ||
|
|
f731e1191c | ||
|
|
54a1e9ab74 | ||
|
|
fabbb63521 | ||
|
|
585d36e6a6 | ||
|
|
371de80b14 | ||
|
|
ede48099c0 | ||
|
|
12e9a655e3 | ||
|
|
1d134e26aa | ||
|
|
bfd6c3eabb |
@@ -24,6 +24,9 @@ jobs:
|
|||||||
- name: Cache cargo
|
- name: Cache cargo
|
||||||
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
|
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
|
||||||
|
|
||||||
|
- name: Check package version sync
|
||||||
|
run: python3 scripts/version.py --check
|
||||||
|
|
||||||
- name: Run tests
|
- name: Run tests
|
||||||
run: cargo test --verbose
|
run: cargo test --verbose
|
||||||
|
|
||||||
|
|||||||
@@ -31,6 +31,9 @@ jobs:
|
|||||||
with:
|
with:
|
||||||
fetch-depth: 2
|
fetch-depth: 2
|
||||||
|
|
||||||
|
- name: Check package version sync
|
||||||
|
run: python3 scripts/version.py --check
|
||||||
|
|
||||||
- name: Check if version changed
|
- name: Check if version changed
|
||||||
id: check
|
id: check
|
||||||
run: |
|
run: |
|
||||||
|
|||||||
@@ -28,6 +28,9 @@ jobs:
|
|||||||
with:
|
with:
|
||||||
fetch-depth: 2
|
fetch-depth: 2
|
||||||
|
|
||||||
|
- name: Check package version sync
|
||||||
|
run: python3 scripts/version.py --check
|
||||||
|
|
||||||
- name: Check if version changed
|
- name: Check if version changed
|
||||||
id: check
|
id: check
|
||||||
run: |
|
run: |
|
||||||
|
|||||||
@@ -28,6 +28,9 @@ jobs:
|
|||||||
with:
|
with:
|
||||||
fetch-depth: 2
|
fetch-depth: 2
|
||||||
|
|
||||||
|
- name: Check package version sync
|
||||||
|
run: python3 scripts/version.py --check
|
||||||
|
|
||||||
- name: Check package version
|
- name: Check package version
|
||||||
id: check
|
id: check
|
||||||
run: |
|
run: |
|
||||||
|
|||||||
@@ -28,6 +28,9 @@ jobs:
|
|||||||
with:
|
with:
|
||||||
fetch-depth: 2
|
fetch-depth: 2
|
||||||
|
|
||||||
|
- name: Check package version sync
|
||||||
|
run: python3 scripts/version.py --check
|
||||||
|
|
||||||
- name: Check if version changed
|
- name: Check if version changed
|
||||||
id: check
|
id: check
|
||||||
run: |
|
run: |
|
||||||
|
|||||||
+2
-2
@@ -31,12 +31,12 @@ Thumbs.db
|
|||||||
napi/index.js
|
napi/index.js
|
||||||
napi/index.d.ts
|
napi/index.d.ts
|
||||||
|
|
||||||
# Local samples and scripts
|
# Local samples
|
||||||
samples/
|
samples/
|
||||||
scripts/
|
|
||||||
|
|
||||||
# Test output
|
# Test output
|
||||||
test_output/
|
test_output/
|
||||||
|
.firecrawl/
|
||||||
|
|
||||||
# Python
|
# Python
|
||||||
__pycache__/
|
__pycache__/
|
||||||
|
|||||||
@@ -61,7 +61,8 @@ src/
|
|||||||
|
|
||||||
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
|
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
|
||||||
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
|
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
|
||||||
- **Regression suite**: sibling repo `pdf-evals` with 179+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
|
- **Regression suite**: sibling repo `pdf-evals` with ~200 snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing. While iterating, prefer a subset run (`bench.py test -q` for the quick set, or `-s <name>` for a named test set) and save the full `bench.py test` for the final pre-commit check.
|
||||||
|
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
|
||||||
|
|
||||||
## Debugging
|
## Debugging
|
||||||
|
|
||||||
|
|||||||
@@ -61,7 +61,7 @@ src/
|
|||||||
|
|
||||||
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
|
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
|
||||||
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
|
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
|
||||||
- **Regression suite**: sibling repo `pdf-evals` with 187+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
|
- **Regression suite**: sibling repo `pdf-evals` with ~200 snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing. While iterating, prefer a subset run (`bench.py test -q` for the quick set, or `-s <name>` for a named test set) and save the full `bench.py test` for the final pre-commit check.
|
||||||
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
|
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
|
||||||
|
|
||||||
## Debugging
|
## Debugging
|
||||||
|
|||||||
+1
-1
@@ -1,6 +1,6 @@
|
|||||||
[package]
|
[package]
|
||||||
name = "pdf-inspector"
|
name = "pdf-inspector"
|
||||||
version = "0.1.7"
|
version = "1.14.0"
|
||||||
edition = "2021"
|
edition = "2021"
|
||||||
autobins = false
|
autobins = false
|
||||||
authors = ["Firecrawl Team"]
|
authors = ["Firecrawl Team"]
|
||||||
|
|||||||
@@ -110,7 +110,7 @@ Or add it manually:
|
|||||||
|
|
||||||
```toml
|
```toml
|
||||||
[dependencies]
|
[dependencies]
|
||||||
pdf-inspector = "0.1"
|
pdf-inspector = "1"
|
||||||
```
|
```
|
||||||
|
|
||||||
```rust
|
```rust
|
||||||
|
|||||||
+4
-3
@@ -5,14 +5,15 @@
|
|||||||
If you believe you've found a security vulnerability in pdf-inspector, please
|
If you believe you've found a security vulnerability in pdf-inspector, please
|
||||||
report it privately so we can fix it before public disclosure.
|
report it privately so we can fix it before public disclosure.
|
||||||
|
|
||||||
**Preferred:** Email **help@firecrawl.dev** with:
|
**Preferred:** Submit through Firecrawl's Bugcrowd vulnerability disclosure
|
||||||
|
program at <https://bugcrowd.com/engagements/firecrawl-vdp-ess>. Please include:
|
||||||
|
|
||||||
- A description of the issue and its impact
|
- A description of the issue and its impact
|
||||||
- Steps to reproduce (a minimal PDF or input that triggers the bug is ideal)
|
- Steps to reproduce (a minimal PDF or input that triggers the bug is ideal)
|
||||||
- The version or commit hash of pdf-inspector you tested against
|
- The version or commit hash of pdf-inspector you tested against
|
||||||
|
|
||||||
**Alternative:** Use GitHub's private vulnerability reporting under the
|
**Alternative:** If you'd rather not use Bugcrowd, email
|
||||||
[Security tab](https://github.com/firecrawl/pdf-inspector/security/advisories/new).
|
**help@firecrawl.dev** with the same details.
|
||||||
|
|
||||||
We'll acknowledge your report in a timely manner and keep you updated on
|
We'll acknowledge your report in a timely manner and keep you updated on
|
||||||
remediation progress. Please do not open a public GitHub issue for security
|
remediation progress. Please do not open a public GitHub issue for security
|
||||||
|
|||||||
+43
-24
@@ -1,38 +1,57 @@
|
|||||||
# Publishing
|
# Publishing
|
||||||
|
|
||||||
The Rust crate is published to [crates.io](https://crates.io/crates/pdf-inspector) with trusted publishing from GitHub Actions. The first release was published manually; future releases publish from `.github/workflows/publish-crate.yml` when a `Cargo.toml` version change lands on `main`.
|
Every pdf-inspector distribution uses one shared semantic version:
|
||||||
|
|
||||||
## crates.io Trusted Publisher
|
- Rust crate: `pdf-inspector`
|
||||||
|
- Python package: `pdf-inspector`
|
||||||
|
- Node package: `@firecrawl/pdf-inspector` and its platform packages
|
||||||
|
- Browser package: `@firecrawl/pdf-inspector-wasm`
|
||||||
|
- Internal NAPI and WASM Rust crates
|
||||||
|
|
||||||
Configure the trusted publisher for the `pdf-inspector` crate with:
|
`Cargo.toml` is the canonical version source. Update every manifest and lockfile
|
||||||
|
with:
|
||||||
|
|
||||||
- Repository: `firecrawl/pdf-inspector`
|
```bash
|
||||||
- Workflow: `publish-crate.yml`
|
python3 scripts/version.py <version>
|
||||||
- Environment: `crates-io`
|
```
|
||||||
|
|
||||||
The workflow uses `rust-lang/crates-io-auth-action@v1` to exchange GitHub's OIDC token for a short-lived crates.io token, then passes it to `cargo publish`.
|
Verify that nothing has diverged with:
|
||||||
|
|
||||||
## Release Steps
|
```bash
|
||||||
|
python3 scripts/version.py --check
|
||||||
|
```
|
||||||
|
|
||||||
1. Update `version` in `Cargo.toml`.
|
CI and every publishing workflow run this check before building or publishing.
|
||||||
2. Merge the version bump to `main`.
|
|
||||||
3. The publish workflow compares the new `Cargo.toml` version with `HEAD~1`, runs `cargo publish --dry-run`, then publishes if that version is not already on crates.io.
|
|
||||||
|
|
||||||
If `Cargo.toml` changes without a package version bump, the workflow exits without publishing.
|
## Release steps
|
||||||
|
|
||||||
## Browser WebAssembly package
|
1. Choose the next shared semantic version and run `scripts/version.py`.
|
||||||
|
2. Review the manifest and lockfile changes in the version-bump pull request.
|
||||||
|
3. Merge the pull request to `main`.
|
||||||
|
4. The crates.io, PyPI, Node, and WASM workflows independently build and
|
||||||
|
publish that version from the same commit.
|
||||||
|
5. After all registries succeed, create one `v<version>` GitHub release that
|
||||||
|
links to each package and describes changes since the previous shared tag.
|
||||||
|
|
||||||
The browser package is published as `@firecrawl/pdf-inspector-wasm`. Its version lives in `wasm/Cargo.toml`, and `.github/workflows/publish-wasm.yml` builds the `web` target with `wasm-pack` before publishing the generated package.
|
The independent workflows are intentionally idempotent. A manual dispatch from
|
||||||
|
`main` can repair a partial release, and already-published artifacts are skipped.
|
||||||
|
|
||||||
The npm package must exist before a trusted publisher can be configured. For the first release only:
|
## Trusted publishers
|
||||||
|
|
||||||
1. Build with `wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release`.
|
The repositories use GitHub Actions OIDC instead of long-lived registry tokens.
|
||||||
2. Inspect with `npm pack --dry-run ./wasm/pkg`.
|
Configure each registry's trusted publisher for `firecrawl/pdf-inspector` and
|
||||||
3. Publish with `npm publish ./wasm/pkg --access public` from an authorized maintainer session.
|
its corresponding workflow:
|
||||||
4. In the package settings on npm, configure the GitHub Actions trusted publisher:
|
|
||||||
- Organization: `firecrawl`
|
|
||||||
- Repository: `pdf-inspector`
|
|
||||||
- Workflow: `publish-wasm.yml`
|
|
||||||
- Allowed action: `npm publish`
|
|
||||||
|
|
||||||
After that one-time bootstrap, bumping the version in `wasm/Cargo.toml` and merging it to `main` publishes through OIDC. Until the package exists, the workflow exits cleanly without attempting an unauthenticated first publish. See npm's [trusted publishing documentation](https://docs.npmjs.com/trusted-publishers/) for the registry-side setup.
|
- crates.io: `publish-crate.yml`, environment `crates-io`
|
||||||
|
- PyPI: `publish-pypi.yml`, environment `pypi`
|
||||||
|
- npm Node package: `publish.yml`
|
||||||
|
- npm WASM package: `publish-wasm.yml`
|
||||||
|
|
||||||
|
The WASM package must exist before npm trusted publishing can be configured. If
|
||||||
|
it ever needs to be bootstrapped again, build and inspect it before publishing:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release
|
||||||
|
npm pack --dry-run ./wasm/pkg
|
||||||
|
npm publish ./wasm/pkg --access public
|
||||||
|
```
|
||||||
|
|||||||
@@ -80,6 +80,17 @@ for page in result.pages:
|
|||||||
|
|
||||||
# Restrict to specific 0-indexed pages (preserves caller order)
|
# Restrict to specific 0-indexed pages (preserves caller order)
|
||||||
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
|
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
|
||||||
|
|
||||||
|
# Structure-tree elements from tagged PDFs (empty list when untagged).
|
||||||
|
# Pages are 1-indexed to match TextItem.page, so (page, mcid) joins directly
|
||||||
|
# against extract_text_with_positions — e.g. to recover real heading levels:
|
||||||
|
elements = pdf_inspector.extract_structure_elements("tagged.pdf")
|
||||||
|
roles = {(e.page, e.mcid): e.role for e in elements}
|
||||||
|
headings = [
|
||||||
|
item.text
|
||||||
|
for item in pdf_inspector.extract_text_with_positions("tagged.pdf")
|
||||||
|
if item.mcid is not None and roles.get((item.page, item.mcid), "").startswith("H")
|
||||||
|
]
|
||||||
```
|
```
|
||||||
|
|
||||||
## API reference
|
## API reference
|
||||||
@@ -100,6 +111,8 @@ result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
|
|||||||
| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes |
|
| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes |
|
||||||
| `extract_pages_markdown(path, pages=None)` | Per-page Markdown + layout metadata (all pages by default) |
|
| `extract_pages_markdown(path, pages=None)` | Per-page Markdown + layout metadata (all pages by default) |
|
||||||
| `extract_pages_markdown_bytes(data, pages=None)` | Per-page Markdown from bytes |
|
| `extract_pages_markdown_bytes(data, pages=None)` | Per-page Markdown from bytes |
|
||||||
|
| `extract_structure_elements(path, pages=None)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
|
||||||
|
| `extract_structure_elements_bytes(data, pages=None)` | Structure-tree elements from bytes |
|
||||||
|
|
||||||
## Types
|
## Types
|
||||||
|
|
||||||
@@ -144,6 +157,12 @@ class TextItem: # extract_text_with_positions
|
|||||||
is_underline: bool
|
is_underline: bool
|
||||||
is_strikeout: bool
|
is_strikeout: bool
|
||||||
item_type: str
|
item_type: str
|
||||||
|
mcid: int | None # marked-content ID for tagged PDFs (None otherwise)
|
||||||
|
|
||||||
|
class StructureElement: # extract_structure_elements
|
||||||
|
page: int # 1-indexed (matches TextItem.page)
|
||||||
|
mcid: int
|
||||||
|
role: str # "H1".."H6", "P", "Table", ... (resolved via /RoleMap)
|
||||||
|
|
||||||
class RegionText: # extract_text_in_regions
|
class RegionText: # extract_text_in_regions
|
||||||
text: str
|
text: str
|
||||||
|
|||||||
+32
-1
@@ -138,6 +138,34 @@ for page in &result.pages {
|
|||||||
println!("Complex layout? {}", result.is_complex);
|
println!("Complex layout? {}", result.is_complex);
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Extract structure-tree elements from tagged PDFs, and join them against
|
||||||
|
`extract_text_with_positions` to attach semantic roles (heading levels,
|
||||||
|
paragraphs, table cells) to extracted text:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
use pdf_inspector::{extract_structure_elements, extract_text_with_positions};
|
||||||
|
use std::collections::HashMap;
|
||||||
|
|
||||||
|
// One entry per marked-content reference, sorted by (page, mcid); empty for
|
||||||
|
// untagged PDFs. Pages are 1-indexed to match `TextItem::page`, so the
|
||||||
|
// (page, mcid) pair is a direct join key.
|
||||||
|
let elements = extract_structure_elements("tagged.pdf", None)?;
|
||||||
|
let roles: HashMap<(u32, i64), &str> = elements
|
||||||
|
.iter()
|
||||||
|
.map(|e| ((e.page, e.mcid), e.role.as_str()))
|
||||||
|
.collect();
|
||||||
|
|
||||||
|
for item in extract_text_with_positions("tagged.pdf")? {
|
||||||
|
if let Some(mcid) = item.mcid {
|
||||||
|
if let Some(role) = roles.get(&(item.page, mcid)) {
|
||||||
|
if role.starts_with('H') {
|
||||||
|
println!("{}: {}", role, item.text);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
## Processing modes
|
## Processing modes
|
||||||
|
|
||||||
| Mode | What it does | Returns |
|
| Mode | What it does | Returns |
|
||||||
@@ -163,6 +191,8 @@ println!("Complex layout? {}", result.is_complex);
|
|||||||
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
|
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
|
||||||
| `extract_pages_markdown(path, pages)` | Per-page Markdown + layout metadata (file) |
|
| `extract_pages_markdown(path, pages)` | Per-page Markdown + layout metadata (file) |
|
||||||
| `extract_pages_markdown_mem(bytes, pages)` | Per-page Markdown from bytes |
|
| `extract_pages_markdown_mem(bytes, pages)` | Per-page Markdown from bytes |
|
||||||
|
| `extract_structure_elements(path, pages)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
|
||||||
|
| `extract_structure_elements_mem(bytes, pages)` | Structure-tree elements from bytes |
|
||||||
|
|
||||||
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
|
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
|
||||||
|
|
||||||
@@ -178,7 +208,8 @@ Low-level detection functions are also available via the `detector` module (`det
|
|||||||
| `DetectionConfig` | Configuration for detection: scan strategy, thresholds |
|
| `DetectionConfig` | Configuration for detection: scan strategy, thresholds |
|
||||||
| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` |
|
| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` |
|
||||||
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
|
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
|
||||||
| `TextItem` | Text with position, font info, and page number |
|
| `TextItem` | Text with position, font info, page number, and optional structure-tree `mcid` |
|
||||||
|
| `StructureElement` | Tagged-PDF structure reference: page (1-indexed), mcid, role (`"H1"`..`"H6"`, `"P"`, …) |
|
||||||
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
|
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
|
||||||
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
|
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
|
||||||
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
|
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
|
||||||
|
|||||||
Generated
+2
-2
@@ -851,7 +851,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
|
|||||||
|
|
||||||
[[package]]
|
[[package]]
|
||||||
name = "pdf-inspector"
|
name = "pdf-inspector"
|
||||||
version = "0.1.7"
|
version = "1.14.0"
|
||||||
dependencies = [
|
dependencies = [
|
||||||
"env_logger",
|
"env_logger",
|
||||||
"include_dir",
|
"include_dir",
|
||||||
@@ -867,7 +867,7 @@ dependencies = [
|
|||||||
|
|
||||||
[[package]]
|
[[package]]
|
||||||
name = "pdf-inspector-napi"
|
name = "pdf-inspector-napi"
|
||||||
version = "0.2.2"
|
version = "1.14.0"
|
||||||
dependencies = [
|
dependencies = [
|
||||||
"napi",
|
"napi",
|
||||||
"napi-build",
|
"napi-build",
|
||||||
|
|||||||
+1
-1
@@ -1,6 +1,6 @@
|
|||||||
[package]
|
[package]
|
||||||
name = "pdf-inspector-napi"
|
name = "pdf-inspector-napi"
|
||||||
version = "0.2.2"
|
version = "1.14.0"
|
||||||
edition = "2021"
|
edition = "2021"
|
||||||
|
|
||||||
[lib]
|
[lib]
|
||||||
|
|||||||
@@ -83,6 +83,22 @@ for (const region of result[0].regions) {
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Async variants
|
||||||
|
|
||||||
|
`processPdf`, `classifyPdf`, and `extractPagesMarkdown` are synchronous and parse on the calling thread — in Node, that's the event loop. For a one-off call in a script that's fine, but in a server a large document can hold the loop for tens to hundreds of milliseconds.
|
||||||
|
|
||||||
|
`processPdfAsync`, `classifyPdfAsync`, and `extractPagesMarkdownAsync` take the same arguments and produce the same results, but run the parse on the libuv thread pool and return a promise, keeping the event loop free. The input buffer is copied before the call returns, so it's safe to reuse or mutate immediately:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { classifyPdfAsync, extractPagesMarkdownAsync } from '@firecrawl/pdf-inspector'
|
||||||
|
|
||||||
|
const classification = await classifyPdfAsync(pdf)
|
||||||
|
if (classification.pdfType === 'TextBased') {
|
||||||
|
const { pages } = await extractPagesMarkdownAsync(pdf)
|
||||||
|
// ...
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
## Types
|
## Types
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
|
|||||||
+6
-6
@@ -8,12 +8,12 @@
|
|||||||
"@napi-rs/cli": "^3.4.1",
|
"@napi-rs/cli": "^3.4.1",
|
||||||
},
|
},
|
||||||
"optionalDependencies": {
|
"optionalDependencies": {
|
||||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.12.0",
|
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.12.0",
|
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.12.0",
|
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.12.0",
|
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.12.0",
|
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.12.0",
|
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0",
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
|
|||||||
+7
-7
@@ -1,6 +1,6 @@
|
|||||||
{
|
{
|
||||||
"name": "@firecrawl/pdf-inspector",
|
"name": "@firecrawl/pdf-inspector",
|
||||||
"version": "1.12.0",
|
"version": "1.14.0",
|
||||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||||
"main": "index.js",
|
"main": "index.js",
|
||||||
"types": "index.d.ts",
|
"types": "index.d.ts",
|
||||||
@@ -52,11 +52,11 @@
|
|||||||
"@napi-rs/cli": "^3.4.1"
|
"@napi-rs/cli": "^3.4.1"
|
||||||
},
|
},
|
||||||
"optionalDependencies": {
|
"optionalDependencies": {
|
||||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.12.0",
|
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.12.0",
|
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.12.0",
|
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.12.0",
|
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.12.0",
|
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
|
||||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.12.0"
|
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0"
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
+236
-41
@@ -89,6 +89,11 @@ pub struct TextItem {
|
|||||||
pub item_type: ItemType,
|
pub item_type: ItemType,
|
||||||
/// URL for link items, `None` for other types.
|
/// URL for link items, `None` for other types.
|
||||||
pub link_url: Option<String>,
|
pub link_url: Option<String>,
|
||||||
|
/// Marked Content ID from the content stream's BDC/BMC operator, `None`
|
||||||
|
/// when the text is not part of marked content. Join with the
|
||||||
|
/// `page`/`mcid` pairs from [`extractStructureElements`] to attach
|
||||||
|
/// structure-tree roles (headings, paragraphs, …) in tagged PDFs.
|
||||||
|
pub mcid: Option<i64>,
|
||||||
}
|
}
|
||||||
|
|
||||||
/// A page's regions for text extraction: (page_index_0based, bboxes).
|
/// A page's regions for text extraction: (page_index_0based, bboxes).
|
||||||
@@ -153,9 +158,7 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn to_napi_page_ocr_reasons(
|
fn to_napi_page_ocr_reasons(reasons: Vec<pdf_inspector::PageOcrReasons>) -> Vec<PageOcrReasons> {
|
||||||
reasons: Vec<pdf_inspector::PageOcrReasons>,
|
|
||||||
) -> Vec<PageOcrReasons> {
|
|
||||||
reasons
|
reasons
|
||||||
.into_iter()
|
.into_iter()
|
||||||
.map(|reason| PageOcrReasons {
|
.map(|reason| PageOcrReasons {
|
||||||
@@ -202,6 +205,31 @@ where
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ---------------------------------------------------------------------------
|
||||||
|
// Shared implementations (single body behind sync and async entry points)
|
||||||
|
// ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
fn process_pdf_impl(bytes: &[u8], pages: Option<Vec<u32>>) -> Result<PdfResult> {
|
||||||
|
let mut opts = pdf_inspector::PdfOptions::new();
|
||||||
|
if let Some(p) = pages {
|
||||||
|
opts = opts.pages(p);
|
||||||
|
}
|
||||||
|
let result = pdf_inspector::process_pdf_mem_with_options(bytes, opts)
|
||||||
|
.map_err(|e| to_napi_err(e, "process_pdf"))?;
|
||||||
|
Ok(to_napi_result(result))
|
||||||
|
}
|
||||||
|
|
||||||
|
fn classify_pdf_impl(bytes: &[u8]) -> Result<PdfClassification> {
|
||||||
|
let result =
|
||||||
|
pdf_inspector::classify_pdf_mem(bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
|
||||||
|
Ok(PdfClassification {
|
||||||
|
pdf_type: convert_pdf_type(result.pdf_type),
|
||||||
|
page_count: result.page_count,
|
||||||
|
pages_needing_ocr: result.pages_needing_ocr,
|
||||||
|
confidence: result.confidence as f64,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
// ---------------------------------------------------------------------------
|
// ---------------------------------------------------------------------------
|
||||||
// Public NAPI API
|
// Public NAPI API
|
||||||
// ---------------------------------------------------------------------------
|
// ---------------------------------------------------------------------------
|
||||||
@@ -210,15 +238,7 @@ where
|
|||||||
#[napi]
|
#[napi]
|
||||||
pub fn process_pdf(buffer: Buffer, pages: Option<Vec<u32>>) -> Result<PdfResult> {
|
pub fn process_pdf(buffer: Buffer, pages: Option<Vec<u32>>) -> Result<PdfResult> {
|
||||||
let bytes: Vec<u8> = buffer.to_vec();
|
let bytes: Vec<u8> = buffer.to_vec();
|
||||||
catch_panic("process_pdf", move || {
|
catch_panic("process_pdf", move || process_pdf_impl(&bytes, pages))
|
||||||
let mut opts = pdf_inspector::PdfOptions::new();
|
|
||||||
if let Some(p) = pages {
|
|
||||||
opts = opts.pages(p);
|
|
||||||
}
|
|
||||||
let result = pdf_inspector::process_pdf_mem_with_options(&bytes, opts)
|
|
||||||
.map_err(|e| to_napi_err(e, "process_pdf"))?;
|
|
||||||
Ok(to_napi_result(result))
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Fast detection only — no text extraction or markdown.
|
/// Fast detection only — no text extraction or markdown.
|
||||||
@@ -238,16 +258,7 @@ pub fn detect_pdf(buffer: Buffer) -> Result<PdfResult> {
|
|||||||
#[napi]
|
#[napi]
|
||||||
pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
|
pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
|
||||||
let bytes: Vec<u8> = buffer.to_vec();
|
let bytes: Vec<u8> = buffer.to_vec();
|
||||||
catch_panic("classify_pdf", move || {
|
catch_panic("classify_pdf", move || classify_pdf_impl(&bytes))
|
||||||
let result =
|
|
||||||
pdf_inspector::classify_pdf_mem(&bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
|
|
||||||
Ok(PdfClassification {
|
|
||||||
pdf_type: convert_pdf_type(result.pdf_type),
|
|
||||||
page_count: result.page_count,
|
|
||||||
pages_needing_ocr: result.pages_needing_ocr,
|
|
||||||
confidence: result.confidence as f64,
|
|
||||||
})
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Extract plain text from a PDF Buffer.
|
/// Extract plain text from a PDF Buffer.
|
||||||
@@ -300,12 +311,61 @@ pub fn extract_text_with_positions(
|
|||||||
is_strikeout: item.is_strikeout,
|
is_strikeout: item.is_strikeout,
|
||||||
item_type,
|
item_type,
|
||||||
link_url,
|
link_url,
|
||||||
|
mcid: item.mcid,
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
.collect())
|
.collect())
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// One structure-tree element reference from a tagged PDF.
|
||||||
|
#[napi(object)]
|
||||||
|
pub struct StructureElementJs {
|
||||||
|
/// 1-indexed page number (matches `TextItem.page`).
|
||||||
|
pub page: u32,
|
||||||
|
/// Marked Content ID from the page's content stream (matches
|
||||||
|
/// `TextItem.mcid`).
|
||||||
|
pub mcid: i64,
|
||||||
|
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
|
||||||
|
/// Custom tags are resolved through the document's role map; tags with
|
||||||
|
/// no standard mapping are returned verbatim.
|
||||||
|
pub role: String,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Extract structure-tree element references from a tagged PDF.
|
||||||
|
///
|
||||||
|
/// Parses the document's structure tree (when present) and returns one
|
||||||
|
/// entry per marked-content reference, resolved to its 1-indexed page,
|
||||||
|
/// MCID, and structure type name. Returns an empty array when the PDF is
|
||||||
|
/// not tagged.
|
||||||
|
///
|
||||||
|
/// Join `(page, mcid)` against the `page`/`mcid` fields from
|
||||||
|
/// [`extractTextWithPositions`] to attach heading levels (H1..H6) and other
|
||||||
|
/// semantic roles to extracted text.
|
||||||
|
///
|
||||||
|
/// Pass 1-indexed page numbers (matching `TextItem.page`) to restrict
|
||||||
|
/// output; omit `pages` for the whole document. Entries are sorted by
|
||||||
|
/// `(page, mcid)`.
|
||||||
|
#[napi]
|
||||||
|
pub fn extract_structure_elements(
|
||||||
|
buffer: Buffer,
|
||||||
|
pages: Option<Vec<u32>>,
|
||||||
|
) -> Result<Vec<StructureElementJs>> {
|
||||||
|
let bytes: Vec<u8> = buffer.to_vec();
|
||||||
|
catch_panic("extract_structure_elements", move || {
|
||||||
|
let elements = pdf_inspector::extract_structure_elements_mem(&bytes, pages.as_deref())
|
||||||
|
.map_err(|e| to_napi_err(e, "extract_structure_elements"))?;
|
||||||
|
Ok(elements
|
||||||
|
.into_iter()
|
||||||
|
.map(|e| StructureElementJs {
|
||||||
|
page: e.page,
|
||||||
|
mcid: e.mcid,
|
||||||
|
role: e.role,
|
||||||
|
})
|
||||||
|
.collect())
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
/// Extract text within bounding-box regions from a PDF.
|
/// Extract text within bounding-box regions from a PDF.
|
||||||
///
|
///
|
||||||
/// For hybrid OCR: layout model detects regions in rendered images,
|
/// For hybrid OCR: layout model detects regions in rendered images,
|
||||||
@@ -633,25 +693,32 @@ pub fn extract_pages_markdown(
|
|||||||
) -> Result<PagesExtractionResult> {
|
) -> Result<PagesExtractionResult> {
|
||||||
let bytes: Vec<u8> = buffer.to_vec();
|
let bytes: Vec<u8> = buffer.to_vec();
|
||||||
catch_panic("extract_pages_markdown", move || {
|
catch_panic("extract_pages_markdown", move || {
|
||||||
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, pages.as_deref())
|
extract_pages_markdown_impl(&bytes, pages.as_deref())
|
||||||
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
|
})
|
||||||
Ok(PagesExtractionResult {
|
}
|
||||||
pages: result
|
|
||||||
.pages
|
fn extract_pages_markdown_impl(
|
||||||
.into_iter()
|
bytes: &[u8],
|
||||||
.map(|r| PageMarkdownResult {
|
pages: Option<&[u32]>,
|
||||||
page: r.page,
|
) -> Result<PagesExtractionResult> {
|
||||||
markdown: r.markdown,
|
let result = pdf_inspector::extract_pages_markdown_mem(bytes, pages)
|
||||||
needs_ocr: r.needs_ocr,
|
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
|
||||||
ocr_reason: r.ocr_reason,
|
Ok(PagesExtractionResult {
|
||||||
})
|
pages: result
|
||||||
.collect(),
|
.pages
|
||||||
pages_with_tables: result.pages_with_tables,
|
.into_iter()
|
||||||
pages_with_columns: result.pages_with_columns,
|
.map(|r| PageMarkdownResult {
|
||||||
pages_needing_ocr: result.pages_needing_ocr,
|
page: r.page,
|
||||||
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
|
markdown: r.markdown,
|
||||||
is_complex: result.is_complex,
|
needs_ocr: r.needs_ocr,
|
||||||
})
|
ocr_reason: r.ocr_reason,
|
||||||
|
})
|
||||||
|
.collect(),
|
||||||
|
pages_with_tables: result.pages_with_tables,
|
||||||
|
pages_with_columns: result.pages_with_columns,
|
||||||
|
pages_needing_ocr: result.pages_needing_ocr,
|
||||||
|
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
|
||||||
|
is_complex: result.is_complex,
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -692,3 +759,131 @@ fn to_page_region_texts(results: Vec<pdf_inspector::PageRegionResult>) -> Vec<Pa
|
|||||||
})
|
})
|
||||||
.collect()
|
.collect()
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ---------------------------------------------------------------------------
|
||||||
|
// Async variants (libuv thread pool via AsyncTask)
|
||||||
|
//
|
||||||
|
// The synchronous exports above parse on the calling thread, which in Node is
|
||||||
|
// the event loop. These `*Async` variants run the same shared implementations
|
||||||
|
// on the libuv thread pool and hand JavaScript a promise, so servers under
|
||||||
|
// concurrent load keep answering requests while a document parses. The sync
|
||||||
|
// exports keep their names, signatures, and behaviour.
|
||||||
|
//
|
||||||
|
// Each factory copies the input Buffer to an owned `Vec<u8>` on the calling
|
||||||
|
// (JS) thread — deliberately. JS execution is single-threaded, so no JS code
|
||||||
|
// can mutate the buffer while the synchronous part of the call copies it.
|
||||||
|
// Holding the napi `Buffer` and reading it from the worker instead would be
|
||||||
|
// zero-copy, but a caller mutating the buffer before the promise settles
|
||||||
|
// would then race the worker's reads — undefined behavior, not a recoverable
|
||||||
|
// error (a known napi-rs soundness hazard with cross-thread Buffer access).
|
||||||
|
// The copy is a one-time memcpy, negligible next to the parse it unblocks.
|
||||||
|
// ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
pub struct ProcessPdfTask {
|
||||||
|
bytes: Vec<u8>,
|
||||||
|
pages: Option<Vec<u32>>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Task for ProcessPdfTask {
|
||||||
|
type Output = PdfResult;
|
||||||
|
type JsValue = PdfResult;
|
||||||
|
|
||||||
|
fn compute(&mut self) -> Result<Self::Output> {
|
||||||
|
let bytes = std::mem::take(&mut self.bytes);
|
||||||
|
let pages = self.pages.take();
|
||||||
|
// AssertUnwindSafe: `bytes`/`pages` are moved into the closure and
|
||||||
|
// dropped on unwind — no shared state can be observed broken.
|
||||||
|
catch_panic(
|
||||||
|
"process_pdf",
|
||||||
|
panic::AssertUnwindSafe(move || process_pdf_impl(&bytes, pages)),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
|
||||||
|
Ok(output)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Async variant of [`processPdf`]: same result, but the parse runs on the
|
||||||
|
/// libuv thread pool instead of the event loop and the call returns a
|
||||||
|
/// promise. The buffer is copied before the call returns, so it may be
|
||||||
|
/// reused or mutated immediately.
|
||||||
|
// ts_return_type is required: napi-rs emits `Promise<unknown>` for
|
||||||
|
// `AsyncTask<T>` returns without it.
|
||||||
|
#[napi(ts_return_type = "Promise<PdfResult>")]
|
||||||
|
pub fn process_pdf_async(buffer: Buffer, pages: Option<Vec<u32>>) -> AsyncTask<ProcessPdfTask> {
|
||||||
|
AsyncTask::new(ProcessPdfTask {
|
||||||
|
bytes: buffer.to_vec(),
|
||||||
|
pages,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
pub struct ClassifyPdfTask {
|
||||||
|
bytes: Vec<u8>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Task for ClassifyPdfTask {
|
||||||
|
type Output = PdfClassification;
|
||||||
|
type JsValue = PdfClassification;
|
||||||
|
|
||||||
|
fn compute(&mut self) -> Result<Self::Output> {
|
||||||
|
let bytes = std::mem::take(&mut self.bytes);
|
||||||
|
catch_panic(
|
||||||
|
"classify_pdf",
|
||||||
|
panic::AssertUnwindSafe(move || classify_pdf_impl(&bytes)),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
|
||||||
|
Ok(output)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Async variant of [`classifyPdf`]: same result, but the classification runs
|
||||||
|
/// on the libuv thread pool instead of the event loop and the call returns a
|
||||||
|
/// promise. The buffer is copied before the call returns, so it may be
|
||||||
|
/// reused or mutated immediately.
|
||||||
|
#[napi(ts_return_type = "Promise<PdfClassification>")]
|
||||||
|
pub fn classify_pdf_async(buffer: Buffer) -> AsyncTask<ClassifyPdfTask> {
|
||||||
|
AsyncTask::new(ClassifyPdfTask {
|
||||||
|
bytes: buffer.to_vec(),
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
pub struct ExtractPagesMarkdownTask {
|
||||||
|
bytes: Vec<u8>,
|
||||||
|
pages: Option<Vec<u32>>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Task for ExtractPagesMarkdownTask {
|
||||||
|
type Output = PagesExtractionResult;
|
||||||
|
type JsValue = PagesExtractionResult;
|
||||||
|
|
||||||
|
fn compute(&mut self) -> Result<Self::Output> {
|
||||||
|
let bytes = std::mem::take(&mut self.bytes);
|
||||||
|
let pages = self.pages.take();
|
||||||
|
catch_panic(
|
||||||
|
"extract_pages_markdown",
|
||||||
|
panic::AssertUnwindSafe(move || extract_pages_markdown_impl(&bytes, pages.as_deref())),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
|
||||||
|
Ok(output)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Async variant of [`extractPagesMarkdown`]: same result, but the extraction
|
||||||
|
/// runs on the libuv thread pool instead of the event loop and the call
|
||||||
|
/// returns a promise. The buffer is copied before the call returns, so it
|
||||||
|
/// may be reused or mutated immediately.
|
||||||
|
#[napi(ts_return_type = "Promise<PagesExtractionResult>")]
|
||||||
|
pub fn extract_pages_markdown_async(
|
||||||
|
buffer: Buffer,
|
||||||
|
pages: Option<Vec<u32>>,
|
||||||
|
) -> AsyncTask<ExtractPagesMarkdownTask> {
|
||||||
|
AsyncTask::new(ExtractPagesMarkdownTask {
|
||||||
|
bytes: buffer.to_vec(),
|
||||||
|
pages,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|||||||
+110
@@ -2,16 +2,21 @@ import { readFileSync } from 'fs';
|
|||||||
import { strict as assert } from 'assert';
|
import { strict as assert } from 'assert';
|
||||||
import {
|
import {
|
||||||
processPdf,
|
processPdf,
|
||||||
|
processPdfAsync,
|
||||||
detectPdf,
|
detectPdf,
|
||||||
classifyPdf,
|
classifyPdf,
|
||||||
|
classifyPdfAsync,
|
||||||
extractText,
|
extractText,
|
||||||
extractTextWithPositions,
|
extractTextWithPositions,
|
||||||
|
extractStructureElements,
|
||||||
extractTextInRegions,
|
extractTextInRegions,
|
||||||
detectVectorGridInRegion,
|
detectVectorGridInRegion,
|
||||||
extractPagesMarkdown,
|
extractPagesMarkdown,
|
||||||
|
extractPagesMarkdownAsync,
|
||||||
} from './index.js';
|
} from './index.js';
|
||||||
|
|
||||||
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
|
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
|
||||||
|
const taggedFixture = readFileSync('../tests/fixtures/firecrawl_docs_tagged.pdf');
|
||||||
|
|
||||||
// --- processPdf ---
|
// --- processPdf ---
|
||||||
console.log('Testing processPdf...');
|
console.log('Testing processPdf...');
|
||||||
@@ -79,6 +84,46 @@ assert.ok(page1Items.length > 0);
|
|||||||
assert.ok(page1Items.every(i => i.page === 1));
|
assert.ok(page1Items.every(i => i.page === 1));
|
||||||
console.log(' extractTextWithPositions with pages: OK');
|
console.log(' extractTextWithPositions with pages: OK');
|
||||||
|
|
||||||
|
// mcid: undefined on untagged PDFs, numeric on tagged marked content
|
||||||
|
assert.ok(items.every(i => i.mcid === undefined || typeof i.mcid === 'number'));
|
||||||
|
const taggedItems = extractTextWithPositions(taggedFixture);
|
||||||
|
assert.ok(
|
||||||
|
taggedItems.some(i => typeof i.mcid === 'number'),
|
||||||
|
'tagged PDF text items should carry Marked Content IDs',
|
||||||
|
);
|
||||||
|
console.log(' extractTextWithPositions mcid: OK');
|
||||||
|
|
||||||
|
// --- extractStructureElements ---
|
||||||
|
console.log('Testing extractStructureElements...');
|
||||||
|
const structureElements = extractStructureElements(taggedFixture);
|
||||||
|
assert.ok(structureElements.length > 0);
|
||||||
|
assert.ok(structureElements.every(e => typeof e.page === 'number'));
|
||||||
|
assert.ok(structureElements.every(e => typeof e.mcid === 'number'));
|
||||||
|
assert.ok(structureElements.every(e => typeof e.role === 'string' && e.role.length > 0));
|
||||||
|
assert.ok(
|
||||||
|
structureElements.some(e => e.role === 'H1'),
|
||||||
|
'tagged fixture should surface H1 heading roles',
|
||||||
|
);
|
||||||
|
|
||||||
|
// (page, mcid) joins against extractTextWithPositions to recover heading text
|
||||||
|
const h1Refs = new Set(
|
||||||
|
structureElements.filter(e => e.role === 'H1').map(e => `${e.page}:${e.mcid}`),
|
||||||
|
);
|
||||||
|
const h1Text = taggedItems
|
||||||
|
.filter(i => typeof i.mcid === 'number' && h1Refs.has(`${i.page}:${i.mcid}`))
|
||||||
|
.map(i => i.text)
|
||||||
|
.join('');
|
||||||
|
assert.ok(h1Text.trim().length > 0, 'H1 join should recover heading text');
|
||||||
|
|
||||||
|
// pages filter is 1-indexed, matching TextItem.page
|
||||||
|
const page1Elements = extractStructureElements(taggedFixture, [1]);
|
||||||
|
assert.ok(page1Elements.length > 0);
|
||||||
|
assert.ok(page1Elements.every(e => e.page === 1));
|
||||||
|
|
||||||
|
// untagged PDFs yield an empty array
|
||||||
|
assert.deepEqual(extractStructureElements(fixture), []);
|
||||||
|
console.log(' extractStructureElements: OK');
|
||||||
|
|
||||||
// --- extractTextInRegions ---
|
// --- extractTextInRegions ---
|
||||||
console.log('Testing extractTextInRegions...');
|
console.log('Testing extractTextInRegions...');
|
||||||
const regionResults = extractTextInRegions(fixture, [
|
const regionResults = extractTextInRegions(fixture, [
|
||||||
@@ -124,10 +169,75 @@ assert.equal(picked.pages[0].page, 2);
|
|||||||
assert.equal(picked.pages[1].page, 0);
|
assert.equal(picked.pages[1].page, 0);
|
||||||
console.log(' extractPagesMarkdown with pages: OK');
|
console.log(' extractPagesMarkdown with pages: OK');
|
||||||
|
|
||||||
|
// --- Async variants ---
|
||||||
|
console.log('Testing async variants...');
|
||||||
|
|
||||||
|
// processPdfAsync returns a promise and matches the sync result
|
||||||
|
const asyncResultPromise = processPdfAsync(fixture);
|
||||||
|
assert.ok(asyncResultPromise instanceof Promise);
|
||||||
|
const asyncResult = await asyncResultPromise;
|
||||||
|
assert.equal(asyncResult.pdfType, result.pdfType);
|
||||||
|
assert.equal(asyncResult.pageCount, result.pageCount);
|
||||||
|
assert.equal(asyncResult.markdown, result.markdown);
|
||||||
|
console.log(' processPdfAsync: OK');
|
||||||
|
|
||||||
|
// processPdfAsync with pages
|
||||||
|
const asyncResult2 = await processPdfAsync(fixture, [1]);
|
||||||
|
assert.equal(asyncResult2.markdown, result2.markdown);
|
||||||
|
console.log(' processPdfAsync with pages: OK');
|
||||||
|
|
||||||
|
// classifyPdfAsync matches the sync result
|
||||||
|
const asyncClassified = await classifyPdfAsync(fixture);
|
||||||
|
assert.equal(asyncClassified.pdfType, classified.pdfType);
|
||||||
|
assert.equal(asyncClassified.pageCount, classified.pageCount);
|
||||||
|
assert.equal(asyncClassified.confidence, classified.confidence);
|
||||||
|
assert.deepEqual(asyncClassified.pagesNeedingOcr, classified.pagesNeedingOcr);
|
||||||
|
console.log(' classifyPdfAsync: OK');
|
||||||
|
|
||||||
|
// extractPagesMarkdownAsync matches the sync result
|
||||||
|
const asyncAllPages = await extractPagesMarkdownAsync(fixture);
|
||||||
|
assert.equal(asyncAllPages.pages.length, allPages.pages.length);
|
||||||
|
assert.deepEqual(
|
||||||
|
asyncAllPages.pages.map(p => p.markdown),
|
||||||
|
allPages.pages.map(p => p.markdown),
|
||||||
|
);
|
||||||
|
assert.equal(asyncAllPages.isComplex, allPages.isComplex);
|
||||||
|
console.log(' extractPagesMarkdownAsync: OK');
|
||||||
|
|
||||||
|
// selected pages preserve caller order
|
||||||
|
const asyncPicked = await extractPagesMarkdownAsync(fixture, [2, 0]);
|
||||||
|
assert.equal(asyncPicked.pages.length, 2);
|
||||||
|
assert.equal(asyncPicked.pages[0].page, 2);
|
||||||
|
assert.equal(asyncPicked.pages[1].page, 0);
|
||||||
|
console.log(' extractPagesMarkdownAsync with pages: OK');
|
||||||
|
|
||||||
|
// input buffer is copied at call time: mutating it immediately after the
|
||||||
|
// call must not affect the in-flight parse
|
||||||
|
const scratch = Buffer.from(fixture);
|
||||||
|
const inFlight = processPdfAsync(scratch);
|
||||||
|
scratch.fill(0);
|
||||||
|
const fromMutated = await inFlight;
|
||||||
|
assert.equal(fromMutated.markdown, result.markdown);
|
||||||
|
console.log(' processPdfAsync input copied at call time: OK');
|
||||||
|
|
||||||
|
// concurrent async calls all settle
|
||||||
|
const [c1, c2, c3] = await Promise.all([
|
||||||
|
processPdfAsync(fixture),
|
||||||
|
classifyPdfAsync(fixture),
|
||||||
|
extractPagesMarkdownAsync(fixture),
|
||||||
|
]);
|
||||||
|
assert.equal(c1.pdfType, 'TextBased');
|
||||||
|
assert.equal(c2.pdfType, 'TextBased');
|
||||||
|
assert.equal(c3.pages.length, 3);
|
||||||
|
console.log(' concurrent async calls: OK');
|
||||||
|
|
||||||
// --- Error handling ---
|
// --- Error handling ---
|
||||||
console.log('Testing error handling...');
|
console.log('Testing error handling...');
|
||||||
assert.throws(() => processPdf(Buffer.from('not a pdf')), /process_pdf/);
|
assert.throws(() => processPdf(Buffer.from('not a pdf')), /process_pdf/);
|
||||||
assert.throws(() => classifyPdf(Buffer.from('')), /classify_pdf/);
|
assert.throws(() => classifyPdf(Buffer.from('')), /classify_pdf/);
|
||||||
|
await assert.rejects(processPdfAsync(Buffer.from('not a pdf')), /process_pdf/);
|
||||||
|
await assert.rejects(classifyPdfAsync(Buffer.from('')), /classify_pdf/);
|
||||||
|
await assert.rejects(extractPagesMarkdownAsync(Buffer.from('')), /extract_pages_markdown/);
|
||||||
console.log(' error handling: OK');
|
console.log(' error handling: OK');
|
||||||
|
|
||||||
console.log('\nAll NAPI tests passed!');
|
console.log('\nAll NAPI tests passed!');
|
||||||
|
|||||||
@@ -51,6 +51,20 @@ class TextItem:
|
|||||||
is_underline: bool
|
is_underline: bool
|
||||||
is_strikeout: bool
|
is_strikeout: bool
|
||||||
item_type: str
|
item_type: str
|
||||||
|
mcid: Optional[int]
|
||||||
|
"""Marked Content ID from the content stream's BDC/BMC operator, None when
|
||||||
|
the text is not part of marked content. Join with the (page, mcid) pairs
|
||||||
|
from extract_structure_elements to attach structure-tree roles in tagged
|
||||||
|
PDFs."""
|
||||||
|
|
||||||
|
class StructureElement:
|
||||||
|
"""One structure-tree element reference from a tagged PDF."""
|
||||||
|
page: int
|
||||||
|
"""1-indexed page number (matches TextItem.page)."""
|
||||||
|
mcid: int
|
||||||
|
"""Marked Content ID from the page's content stream (matches TextItem.mcid)."""
|
||||||
|
role: str
|
||||||
|
"""Standard structure type name ("H1".."H6", "P", "Table", "TD", ...)."""
|
||||||
|
|
||||||
class RegionText:
|
class RegionText:
|
||||||
"""Extracted text for a single region."""
|
"""Extracted text for a single region."""
|
||||||
@@ -132,6 +146,27 @@ def extract_text_with_positions_bytes(data: bytes, pages: Optional[list[int]] =
|
|||||||
"""Extract text with position information from bytes."""
|
"""Extract text with position information from bytes."""
|
||||||
...
|
...
|
||||||
|
|
||||||
|
def extract_structure_elements(path: str, pages: Optional[list[int]] = None) -> list[StructureElement]:
|
||||||
|
"""Extract structure-tree element references from a tagged PDF file.
|
||||||
|
|
||||||
|
Returns one entry per marked-content reference, resolved to its 1-indexed
|
||||||
|
page, MCID, and structure type name ("H1".."H6", "P", "Table", ...), sorted
|
||||||
|
by (page, mcid). Returns an empty list when the PDF is not tagged.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
path: Path to the PDF file.
|
||||||
|
pages: Optional list of 1-indexed pages (matching ``TextItem.page``).
|
||||||
|
When ``None`` (default), the whole document is returned.
|
||||||
|
"""
|
||||||
|
...
|
||||||
|
|
||||||
|
def extract_structure_elements_bytes(data: bytes, pages: Optional[list[int]] = None) -> list[StructureElement]:
|
||||||
|
"""Extract structure-tree element references from tagged PDF bytes.
|
||||||
|
|
||||||
|
See :func:`extract_structure_elements` for details.
|
||||||
|
"""
|
||||||
|
...
|
||||||
|
|
||||||
def extract_text_in_regions(
|
def extract_text_in_regions(
|
||||||
path: str,
|
path: str,
|
||||||
page_regions: list[tuple[int, list[list[float]]]],
|
page_regions: list[tuple[int, list[list[float]]]],
|
||||||
|
|||||||
+3
-3
@@ -4,9 +4,9 @@ build-backend = "maturin"
|
|||||||
|
|
||||||
[project]
|
[project]
|
||||||
name = "pdf-inspector"
|
name = "pdf-inspector"
|
||||||
# Bump this to publish to PyPI — CI publishes automatically when the version
|
# Keep package versions in sync with `python3 scripts/version.py <version>`.
|
||||||
# changes on main (same flow as napi/package.json for npm).
|
# CI publishes automatically when the synchronized change lands on main.
|
||||||
version = "0.2.6"
|
version = "1.14.0"
|
||||||
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
|
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
|
||||||
readme = "docs/python.md"
|
readme = "docs/python.md"
|
||||||
license = { text = "MIT" }
|
license = { text = "MIT" }
|
||||||
|
|||||||
@@ -0,0 +1,111 @@
|
|||||||
|
import json
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||||
|
|
||||||
|
from version import PLATFORM_PACKAGES, check_versions, set_versions
|
||||||
|
|
||||||
|
|
||||||
|
class VersionTests(unittest.TestCase):
|
||||||
|
def setUp(self):
|
||||||
|
self.temporary = tempfile.TemporaryDirectory()
|
||||||
|
self.root = Path(self.temporary.name)
|
||||||
|
(self.root / "napi").mkdir()
|
||||||
|
(self.root / "site").mkdir()
|
||||||
|
(self.root / "wasm").mkdir()
|
||||||
|
|
||||||
|
self._write_manifest("Cargo.toml", "package", "0.1.0")
|
||||||
|
self._write_manifest("pyproject.toml", "project", "0.1.0")
|
||||||
|
self._write_manifest("napi/Cargo.toml", "package", "0.1.0")
|
||||||
|
self._write_manifest("wasm/Cargo.toml", "package", "0.1.0")
|
||||||
|
|
||||||
|
package = {
|
||||||
|
"name": "@firecrawl/pdf-inspector",
|
||||||
|
"version": "0.1.0",
|
||||||
|
"optionalDependencies": {
|
||||||
|
dependency: "0.1.0" for dependency in PLATFORM_PACKAGES
|
||||||
|
},
|
||||||
|
}
|
||||||
|
(self.root / "napi/package.json").write_text(
|
||||||
|
json.dumps(package), encoding="utf-8"
|
||||||
|
)
|
||||||
|
(self.root / "napi/bun.lock").write_text(
|
||||||
|
"\n".join(
|
||||||
|
f' "{dependency}": "0.1.0",'
|
||||||
|
for dependency in PLATFORM_PACKAGES
|
||||||
|
)
|
||||||
|
+ "\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
(self.root / "site/index.html").write_text(
|
||||||
|
'https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.0/'
|
||||||
|
'pdf_inspector_wasm.js\n',
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
self._write_lock(
|
||||||
|
"napi/Cargo.lock", ("pdf-inspector", "pdf-inspector-napi")
|
||||||
|
)
|
||||||
|
self._write_lock(
|
||||||
|
"wasm/Cargo.lock", ("pdf-inspector", "pdf-inspector-wasm")
|
||||||
|
)
|
||||||
|
|
||||||
|
def tearDown(self):
|
||||||
|
self.temporary.cleanup()
|
||||||
|
|
||||||
|
def _write_manifest(self, relative, section, version):
|
||||||
|
(self.root / relative).write_text(
|
||||||
|
f'[{section}]\nname = "fixture"\nversion = "{version}"\n',
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
|
||||||
|
def _write_lock(self, relative, packages):
|
||||||
|
content = "\n".join(
|
||||||
|
f'[[package]]\nname = "{package}"\nversion = "0.1.0"\n'
|
||||||
|
for package in packages
|
||||||
|
)
|
||||||
|
(self.root / relative).write_text(content, encoding="utf-8")
|
||||||
|
|
||||||
|
def test_updates_every_version_location(self):
|
||||||
|
set_versions("1.14.0", self.root)
|
||||||
|
|
||||||
|
self.assertEqual(check_versions(self.root), "1.14.0")
|
||||||
|
|
||||||
|
def test_reports_a_divergent_package(self):
|
||||||
|
self._write_manifest("wasm/Cargo.toml", "package", "0.2.0")
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(ValueError, "WASM package: 0.2.0"):
|
||||||
|
check_versions(self.root)
|
||||||
|
|
||||||
|
def test_rejects_an_invalid_version(self):
|
||||||
|
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
|
||||||
|
set_versions("next", self.root)
|
||||||
|
|
||||||
|
def test_rejects_numeric_prerelease_with_leading_zero(self):
|
||||||
|
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
|
||||||
|
set_versions("1.2.3-01", self.root)
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_preflight_failure_does_not_partially_update(self):
|
||||||
|
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
|
||||||
|
(self.root / "site/index.html").write_text(
|
||||||
|
"missing module URL\n", encoding="utf-8"
|
||||||
|
)
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(ValueError, "Missing pinned WASM package URL"):
|
||||||
|
set_versions("1.14.0", self.root)
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
@@ -0,0 +1,262 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Keep every pdf-inspector package on one release version."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
PRERELEASE_IDENTIFIER = (
|
||||||
|
r"(?:0|[1-9]\d*|[0-9A-Za-z-]*[A-Za-z-][0-9A-Za-z-]*)"
|
||||||
|
)
|
||||||
|
SEMVER = re.compile(
|
||||||
|
r"^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)"
|
||||||
|
rf"(?:-{PRERELEASE_IDENTIFIER}(?:\.{PRERELEASE_IDENTIFIER})*)?"
|
||||||
|
r"(?:\+[0-9A-Za-z-]+(?:\.[0-9A-Za-z-]+)*)?$"
|
||||||
|
)
|
||||||
|
VERSION_LINE = re.compile(r'^(\s*version\s*=\s*")[^"]+(".*)$')
|
||||||
|
SECTION_LINE = re.compile(r"^\s*\[([^]]+)]\s*$")
|
||||||
|
PLATFORM_PACKAGES = (
|
||||||
|
"@firecrawl/pdf-inspector-linux-x64-gnu",
|
||||||
|
"@firecrawl/pdf-inspector-linux-x64-musl",
|
||||||
|
"@firecrawl/pdf-inspector-linux-arm64-gnu",
|
||||||
|
"@firecrawl/pdf-inspector-linux-arm64-musl",
|
||||||
|
"@firecrawl/pdf-inspector-darwin-arm64",
|
||||||
|
"@firecrawl/pdf-inspector-win32-x64-msvc",
|
||||||
|
)
|
||||||
|
TOML_VERSIONS = (
|
||||||
|
("Rust crate", Path("Cargo.toml"), "package"),
|
||||||
|
("Python package", Path("pyproject.toml"), "project"),
|
||||||
|
("NAPI crate", Path("napi/Cargo.toml"), "package"),
|
||||||
|
("WASM package", Path("wasm/Cargo.toml"), "package"),
|
||||||
|
)
|
||||||
|
LOCK_VERSIONS = (
|
||||||
|
("NAPI lock: core", Path("napi/Cargo.lock"), "pdf-inspector"),
|
||||||
|
("NAPI lock: binding", Path("napi/Cargo.lock"), "pdf-inspector-napi"),
|
||||||
|
("WASM lock: core", Path("wasm/Cargo.lock"), "pdf-inspector"),
|
||||||
|
("WASM lock: binding", Path("wasm/Cargo.lock"), "pdf-inspector-wasm"),
|
||||||
|
)
|
||||||
|
SITE_WASM_VERSION = re.compile(
|
||||||
|
r"(@firecrawl/pdf-inspector-wasm@)([^/\"]+)(/pdf_inspector_wasm\.js)"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _read_section_version(path: Path, section: str) -> str:
|
||||||
|
active = False
|
||||||
|
for line in path.read_text(encoding="utf-8").splitlines():
|
||||||
|
section_match = SECTION_LINE.match(line)
|
||||||
|
if section_match:
|
||||||
|
active = section_match.group(1) == section
|
||||||
|
elif active:
|
||||||
|
version_match = VERSION_LINE.match(line)
|
||||||
|
if version_match:
|
||||||
|
return line.split('"', 2)[1]
|
||||||
|
raise ValueError(f"No version found in [{section}] of {path}")
|
||||||
|
|
||||||
|
|
||||||
|
def _write_section_version(path: Path, section: str, version: str) -> None:
|
||||||
|
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
|
||||||
|
active = False
|
||||||
|
for index, line in enumerate(lines):
|
||||||
|
section_match = SECTION_LINE.match(line)
|
||||||
|
if section_match:
|
||||||
|
active = section_match.group(1) == section
|
||||||
|
elif active:
|
||||||
|
version_match = VERSION_LINE.match(line)
|
||||||
|
if version_match:
|
||||||
|
newline = "\n" if line.endswith("\n") else ""
|
||||||
|
replacement = (
|
||||||
|
f"{version_match.group(1)}{version}"
|
||||||
|
f"{version_match.group(2).rstrip()}"
|
||||||
|
)
|
||||||
|
lines[index] = (
|
||||||
|
f"{replacement}{newline}"
|
||||||
|
)
|
||||||
|
path.write_text("".join(lines), encoding="utf-8")
|
||||||
|
return
|
||||||
|
raise ValueError(f"No version found in [{section}] of {path}")
|
||||||
|
|
||||||
|
|
||||||
|
def _package_block(lines: list[str], package: str) -> tuple[int, int]:
|
||||||
|
for start, line in enumerate(lines):
|
||||||
|
if line.strip() != "[[package]]":
|
||||||
|
continue
|
||||||
|
end = next(
|
||||||
|
(
|
||||||
|
index
|
||||||
|
for index in range(start + 1, len(lines))
|
||||||
|
if lines[index].strip() == "[[package]]"
|
||||||
|
),
|
||||||
|
len(lines),
|
||||||
|
)
|
||||||
|
if any(line.strip() == f'name = "{package}"' for line in lines[start:end]):
|
||||||
|
return start, end
|
||||||
|
raise ValueError(f"No lockfile entry found for {package}")
|
||||||
|
|
||||||
|
|
||||||
|
def _read_lock_version(path: Path, package: str) -> str:
|
||||||
|
lines = path.read_text(encoding="utf-8").splitlines()
|
||||||
|
start, end = _package_block(lines, package)
|
||||||
|
for line in lines[start:end]:
|
||||||
|
version_match = VERSION_LINE.match(line)
|
||||||
|
if version_match:
|
||||||
|
return line.split('"', 2)[1]
|
||||||
|
raise ValueError(f"No version found for {package} in {path}")
|
||||||
|
|
||||||
|
|
||||||
|
def _write_lock_version(path: Path, package: str, version: str) -> None:
|
||||||
|
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
|
||||||
|
start, end = _package_block(lines, package)
|
||||||
|
for index in range(start, end):
|
||||||
|
version_match = VERSION_LINE.match(lines[index])
|
||||||
|
if version_match:
|
||||||
|
newline = "\n" if lines[index].endswith("\n") else ""
|
||||||
|
lines[index] = (
|
||||||
|
f'{version_match.group(1)}{version}{version_match.group(2).rstrip()}'
|
||||||
|
f"{newline}"
|
||||||
|
)
|
||||||
|
path.write_text("".join(lines), encoding="utf-8")
|
||||||
|
return
|
||||||
|
raise ValueError(f"No version found for {package} in {path}")
|
||||||
|
|
||||||
|
|
||||||
|
def _node_versions(root: Path) -> dict[str, str]:
|
||||||
|
package = json.loads((root / "napi/package.json").read_text(encoding="utf-8"))
|
||||||
|
versions = {"Node package": package["version"]}
|
||||||
|
optional = package.get("optionalDependencies", {})
|
||||||
|
for dependency in PLATFORM_PACKAGES:
|
||||||
|
if dependency not in optional:
|
||||||
|
raise ValueError(f"Missing Node optional dependency: {dependency}")
|
||||||
|
versions[f"Node optional dependency: {dependency}"] = optional[dependency]
|
||||||
|
return versions
|
||||||
|
|
||||||
|
|
||||||
|
def _bun_versions(root: Path) -> dict[str, str]:
|
||||||
|
text = (root / "napi/bun.lock").read_text(encoding="utf-8")
|
||||||
|
versions = {}
|
||||||
|
for dependency in PLATFORM_PACKAGES:
|
||||||
|
match = re.search(
|
||||||
|
rf'"{re.escape(dependency)}": "([^"]+)"[,]', text
|
||||||
|
)
|
||||||
|
if not match:
|
||||||
|
raise ValueError(f"Missing Bun lock dependency: {dependency}")
|
||||||
|
versions[f"Bun lock: {dependency}"] = match.group(1)
|
||||||
|
return versions
|
||||||
|
|
||||||
|
|
||||||
|
def _site_wasm_version(root: Path) -> str:
|
||||||
|
text = (root / "site/index.html").read_text(encoding="utf-8")
|
||||||
|
match = SITE_WASM_VERSION.search(text)
|
||||||
|
if not match:
|
||||||
|
raise ValueError("Missing pinned WASM package URL in site/index.html")
|
||||||
|
return match.group(2)
|
||||||
|
|
||||||
|
|
||||||
|
def package_versions(root: Path = ROOT) -> dict[str, str]:
|
||||||
|
versions = {
|
||||||
|
label: _read_section_version(root / relative, section)
|
||||||
|
for label, relative, section in TOML_VERSIONS
|
||||||
|
}
|
||||||
|
versions.update(_node_versions(root))
|
||||||
|
versions.update(_bun_versions(root))
|
||||||
|
versions["Website WASM module"] = _site_wasm_version(root)
|
||||||
|
versions.update(
|
||||||
|
{
|
||||||
|
label: _read_lock_version(root / relative, package)
|
||||||
|
for label, relative, package in LOCK_VERSIONS
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return versions
|
||||||
|
|
||||||
|
|
||||||
|
def check_versions(root: Path = ROOT) -> str:
|
||||||
|
versions = package_versions(root)
|
||||||
|
expected = versions["Rust crate"]
|
||||||
|
if not SEMVER.fullmatch(expected):
|
||||||
|
raise ValueError(f"Rust crate has an invalid semantic version: {expected}")
|
||||||
|
mismatches = {
|
||||||
|
label: version for label, version in versions.items() if version != expected
|
||||||
|
}
|
||||||
|
if mismatches:
|
||||||
|
details = "\n".join(
|
||||||
|
f" - {label}: {version}" for label, version in mismatches.items()
|
||||||
|
)
|
||||||
|
raise ValueError(f"Expected every package to use {expected}:\n{details}")
|
||||||
|
return expected
|
||||||
|
|
||||||
|
|
||||||
|
def set_versions(version: str, root: Path = ROOT) -> None:
|
||||||
|
if not SEMVER.fullmatch(version):
|
||||||
|
raise ValueError(f"Invalid semantic version: {version}")
|
||||||
|
|
||||||
|
# Validate every expected location before writing the first file. This
|
||||||
|
# prevents a stale manifest or generated file from leaving a partial bump.
|
||||||
|
package_versions(root)
|
||||||
|
|
||||||
|
for _, relative, section in TOML_VERSIONS:
|
||||||
|
_write_section_version(root / relative, section, version)
|
||||||
|
|
||||||
|
package_path = root / "napi/package.json"
|
||||||
|
package = json.loads(package_path.read_text(encoding="utf-8"))
|
||||||
|
package["version"] = version
|
||||||
|
optional = package.get("optionalDependencies", {})
|
||||||
|
for dependency in PLATFORM_PACKAGES:
|
||||||
|
if dependency not in optional:
|
||||||
|
raise ValueError(f"Missing Node optional dependency: {dependency}")
|
||||||
|
optional[dependency] = version
|
||||||
|
package_path.write_text(json.dumps(package, indent=2) + "\n", encoding="utf-8")
|
||||||
|
|
||||||
|
bun_path = root / "napi/bun.lock"
|
||||||
|
bun_text = bun_path.read_text(encoding="utf-8")
|
||||||
|
for dependency in PLATFORM_PACKAGES:
|
||||||
|
pattern = rf'("{re.escape(dependency)}": ")[^"]+("[,])'
|
||||||
|
bun_text, count = re.subn(
|
||||||
|
pattern, rf"\g<1>{version}\g<2>", bun_text, count=1
|
||||||
|
)
|
||||||
|
if count != 1:
|
||||||
|
raise ValueError(f"Missing Bun lock dependency: {dependency}")
|
||||||
|
bun_path.write_text(bun_text, encoding="utf-8")
|
||||||
|
|
||||||
|
site_path = root / "site/index.html"
|
||||||
|
site_text = site_path.read_text(encoding="utf-8")
|
||||||
|
site_text, count = SITE_WASM_VERSION.subn(
|
||||||
|
rf"\g<1>{version}\g<3>", site_text, count=1
|
||||||
|
)
|
||||||
|
if count != 1:
|
||||||
|
raise ValueError("Missing pinned WASM package URL in site/index.html")
|
||||||
|
site_path.write_text(site_text, encoding="utf-8")
|
||||||
|
|
||||||
|
for _, relative, package_name in LOCK_VERSIONS:
|
||||||
|
_write_lock_version(root / relative, package_name, version)
|
||||||
|
|
||||||
|
check_versions(root)
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument("version", nargs="?", help="new shared semantic version")
|
||||||
|
parser.add_argument(
|
||||||
|
"--check", action="store_true", help="fail if package versions have diverged"
|
||||||
|
)
|
||||||
|
arguments = parser.parse_args()
|
||||||
|
if arguments.check == bool(arguments.version):
|
||||||
|
parser.error("provide either a version or --check")
|
||||||
|
|
||||||
|
try:
|
||||||
|
if arguments.check:
|
||||||
|
version = check_versions()
|
||||||
|
print(f"All packages use {version}")
|
||||||
|
else:
|
||||||
|
set_versions(arguments.version)
|
||||||
|
print(f"Updated all packages to {arguments.version}")
|
||||||
|
except ValueError as error:
|
||||||
|
parser.exit(1, f"{error}\n")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
+8
-8
@@ -858,22 +858,22 @@
|
|||||||
<p>Evaluated on the <a class="text-link" href="https://github.com/opendataloader-project/opendataloader-bench">opendataloader-bench</a> corpus of 200 PDFs. This comparison covers local engines without model-based PDF parsing, with OCR disabled. Higher scores are better.</p>
|
<p>Evaluated on the <a class="text-link" href="https://github.com/opendataloader-project/opendataloader-bench">opendataloader-bench</a> corpus of 200 PDFs. This comparison covers local engines without model-based PDF parsing, with OCR disabled. Higher scores are better.</p>
|
||||||
</div>
|
</div>
|
||||||
<div class="benchmark-card">
|
<div class="benchmark-card">
|
||||||
<div class="benchmark-top"><span><strong>200 PDFs</strong> · OpenDataLoader benchmark</span><span>Apple M4 Pro · median of 3 runs</span></div>
|
<div class="benchmark-top"><span><strong>200 PDFs</strong> · OpenDataLoader benchmark</span><span>Apple M4 Pro · median of 5 runs</span></div>
|
||||||
<div class="table-scroll">
|
<div class="table-scroll">
|
||||||
<table aria-label="PDF extraction benchmark results">
|
<table aria-label="PDF extraction benchmark results">
|
||||||
<thead>
|
<thead>
|
||||||
<tr><th>Engine</th><th>Overall</th><th>Reading order</th><th>Tables</th><th>Headings</th><th>Complete run</th></tr>
|
<tr><th>Engine</th><th>Overall</th><th>Reading order</th><th>Tables</th><th>Headings</th><th>Complete run</th></tr>
|
||||||
</thead>
|
</thead>
|
||||||
<tbody>
|
<tbody>
|
||||||
<tr class="highlight"><td>pdf-inspector</td><td>0.875</td><td>0.915</td><td>0.814</td><td>0.788</td><td>2.8s</td></tr>
|
<tr class="highlight"><td>pdf-inspector</td><td>0.875</td><td>0.915</td><td>0.814</td><td>0.788</td><td>0.470s</td></tr>
|
||||||
<tr><td>LiteParse</td><td>0.870</td><td>0.908</td><td>0.693</td><td>0.811</td><td>13.9s</td></tr>
|
<tr><td>LiteParse</td><td>0.873</td><td>0.913</td><td>0.693</td><td>0.811</td><td>0.750s</td></tr>
|
||||||
<tr><td>OpenDataLoader</td><td>0.843</td><td>0.912</td><td>0.489</td><td>0.760</td><td>9.8s</td></tr>
|
<tr><td>OpenDataLoader</td><td>0.831</td><td>0.902</td><td>0.489</td><td>0.739</td><td>2.569s</td></tr>
|
||||||
<tr><td>PyMuPDF4LLM</td><td>0.735</td><td>0.886</td><td>0.401</td><td>0.424</td><td>15.5s</td></tr>
|
<tr><td>PyMuPDF4LLM</td><td>0.735</td><td>0.886</td><td>0.401</td><td>0.424</td><td>17.117s</td></tr>
|
||||||
<tr><td>MarkItDown</td><td>0.583</td><td>0.879</td><td>0.000</td><td>0.000</td><td>6.7s</td></tr>
|
<tr><td>MarkItDown</td><td>0.589</td><td>0.844</td><td>0.273</td><td>0.000</td><td>16.165s</td></tr>
|
||||||
</tbody>
|
</tbody>
|
||||||
</table>
|
</table>
|
||||||
</div>
|
</div>
|
||||||
<div class="benchmark-note">Refreshed July 16, 2026. Scores use the benchmark’s NID, TEDS, and MHS evaluators.</div>
|
<div class="benchmark-note">Refreshed July 31, 2026. Scores use the benchmark’s NID, TEDS, and MHS evaluators; speed is the median of five alternating or rotating complete corpus runs after an excluded warm-up. <a class="text-link" href="https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results">Versions and raw artifacts</a>.</div>
|
||||||
</div>
|
</div>
|
||||||
<div class="best-fit">
|
<div class="best-fit">
|
||||||
<strong>Best fit</strong>
|
<strong>Best fit</strong>
|
||||||
@@ -975,7 +975,7 @@ result = pdf_inspector.<span class="fn">process_pdf</span>(<span class="str">"do
|
|||||||
<script>
|
<script>
|
||||||
(() => {
|
(() => {
|
||||||
const MAX_FILE_SIZE = 25 * 1024 * 1024;
|
const MAX_FILE_SIZE = 25 * 1024 * 1024;
|
||||||
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.1/pdf_inspector_wasm.js";
|
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@1.14.0/pdf_inspector_wasm.js";
|
||||||
const input = document.querySelector("#pdf-input");
|
const input = document.querySelector("#pdf-input");
|
||||||
const dropZone = document.querySelector("#drop-zone");
|
const dropZone = document.querySelector("#drop-zone");
|
||||||
const filePanel = document.querySelector("#demo-file");
|
const filePanel = document.querySelector("#demo-file");
|
||||||
|
|||||||
+65
-1
@@ -1659,7 +1659,15 @@ fn hex_val(b: u8) -> Option<u8> {
|
|||||||
/// Standard page: 612x792 points (US Letter) = ~485,000 sq points
|
/// Standard page: 612x792 points (US Letter) = ~485,000 sq points
|
||||||
/// At 2x resolution that's ~1.9M pixels, so we use 250K pixels as threshold
|
/// At 2x resolution that's ~1.9M pixels, so we use 250K pixels as threshold
|
||||||
/// (accounting for varying DPI and page sizes)
|
/// (accounting for varying DPI and page sizes)
|
||||||
fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
|
/// Returns `(has_images, total_image_area, has_template_image)` for a page.
|
||||||
|
/// `has_template_image` means a single large (>50% page coverage)
|
||||||
|
/// background image — the signal `classify_pdf`/`detect_pdf_type` uses to
|
||||||
|
/// route a page to OCR regardless of any incidental native text drawn over
|
||||||
|
/// it. Exposed at crate visibility so extraction-side per-page `needs_ocr`
|
||||||
|
/// computation (`extract_pages_markdown_mem`) can consult the same signal
|
||||||
|
/// instead of maintaining its own, independent notion of "needs OCR" that
|
||||||
|
/// can silently disagree with detection — see #227.
|
||||||
|
pub(crate) fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
|
||||||
// Threshold: image covering roughly half a page at 150+ DPI
|
// Threshold: image covering roughly half a page at 150+ DPI
|
||||||
// 612 * 792 / 2 * (150/72)^2 ≈ 1M pixels, but we'll be conservative
|
// 612 * 792 / 2 * (150/72)^2 ≈ 1M pixels, but we'll be conservative
|
||||||
const TEMPLATE_IMAGE_THRESHOLD: u64 = 500_000; // 500K pixels
|
const TEMPLATE_IMAGE_THRESHOLD: u64 = 500_000; // 500K pixels
|
||||||
@@ -1741,6 +1749,62 @@ fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
|
|||||||
(has_images, total_area, has_template_image)
|
(has_images, total_area, has_template_image)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Computes both `(needs_ocr_for_template_image, has_vector_text)` for a
|
||||||
|
/// page from a single shared `analyze_page_content` pass — that call
|
||||||
|
/// decompresses and scans every content stream (page + XObjects) plus
|
||||||
|
/// image coverage, so `extract_pages_markdown_mem` must not invoke it
|
||||||
|
/// twice per page (once per signal) the way `detect_from_document` avoids
|
||||||
|
/// by caching its per-page `PageAnalysis`.
|
||||||
|
///
|
||||||
|
/// `needs_ocr_for_template_image` is true when a page's template image
|
||||||
|
/// should be treated as a scan needing OCR — a single full-page background
|
||||||
|
/// image with little/no real text — rather than a text page that happens
|
||||||
|
/// to carry a watermark, letterhead, or figure. Mirrors the two distinct
|
||||||
|
/// signals classification uses to route a template-image page to OCR:
|
||||||
|
///
|
||||||
|
/// 1. `looks_like_scan`: image_count <= 1, few text operators (<50), and
|
||||||
|
/// low alphanumeric diversity in raw string operands (unless decodable
|
||||||
|
/// CID/ToUnicode fonts explain that away) — the gate used for
|
||||||
|
/// `pages_with_template_images` and Mixed-type per-page routing.
|
||||||
|
/// 2. Insufficient real text volume, using `DetectionConfig::default()`'s
|
||||||
|
/// `min_text_ops_per_page` (3) — the same threshold Mixed-type per-page
|
||||||
|
/// routing applies via `text_operator_count < config.min_text_ops_per_page
|
||||||
|
/// && has_images` (simplified here since a template image implies
|
||||||
|
/// `has_images`). Deliberately *not* the higher `effective_min_ops`
|
||||||
|
/// floor (`min_text_ops_per_page.max(10)`) that whole-document
|
||||||
|
/// `PdfType::ImageBased`/`Scanned` classification uses for
|
||||||
|
/// `pages_with_text` — that's a cross-page aggregate decision this
|
||||||
|
/// per-page function has no way to replicate exactly, and the lower
|
||||||
|
/// per-page threshold is the one a single page's own signals can
|
||||||
|
/// actually agree with.
|
||||||
|
///
|
||||||
|
/// `has_vector_text` is true when a page has vector-outlined text (glyphs
|
||||||
|
/// drawn as paths rather than shown via text-showing operators) —
|
||||||
|
/// `detect_from_document`'s Mixed-type per-page routing always sends
|
||||||
|
/// these pages to OCR, independent of any template-image check, since
|
||||||
|
/// outlined glyphs can't be extracted as text at all.
|
||||||
|
///
|
||||||
|
/// Exposed at crate visibility so `extract_pages_markdown_mem` can apply
|
||||||
|
/// the same gates classification needs elsewhere instead of treating the
|
||||||
|
/// raw signals alone as sufficient — see #227/#231.
|
||||||
|
pub(crate) fn page_ocr_signals(doc: &Document, page_id: ObjectId) -> (bool, bool) {
|
||||||
|
let analysis = analyze_page_content(doc, page_id);
|
||||||
|
|
||||||
|
let needs_ocr_for_template_image = if !analysis.has_template_image {
|
||||||
|
false
|
||||||
|
} else {
|
||||||
|
let alphanum_low = analysis.unique_alphanum_chars < 10
|
||||||
|
&& !(analysis.has_decodable_text_fonts && analysis.text_operator_count >= 10);
|
||||||
|
let looks_like_scan =
|
||||||
|
analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_low;
|
||||||
|
let insufficient_text =
|
||||||
|
analysis.text_operator_count < DetectionConfig::default().min_text_ops_per_page;
|
||||||
|
looks_like_scan || insufficient_text
|
||||||
|
};
|
||||||
|
|
||||||
|
(needs_ocr_for_template_image, analysis.has_vector_text)
|
||||||
|
}
|
||||||
|
|
||||||
/// Recursively collect image dimensions from XObject resources,
|
/// Recursively collect image dimensions from XObject resources,
|
||||||
/// including images nested inside Form XObjects.
|
/// including images nested inside Form XObjects.
|
||||||
fn collect_images_from_resources(
|
fn collect_images_from_resources(
|
||||||
|
|||||||
@@ -41,6 +41,17 @@ fn strip_pdf_comments(data: &[u8]) -> Vec<u8> {
|
|||||||
while i < data.len() {
|
while i < data.len() {
|
||||||
let b = data[i];
|
let b = data[i];
|
||||||
match b {
|
match b {
|
||||||
|
// Inside a string literal, a backslash escapes the next byte —
|
||||||
|
// `\(`, `\)`, and `\\` must not touch the nesting depth, or a
|
||||||
|
// later `%` glyph inside a string gets stripped as a comment,
|
||||||
|
// corrupting the stream.
|
||||||
|
b'\\' if in_string > 0 => {
|
||||||
|
result.push(b);
|
||||||
|
if let Some(&next) = data.get(i + 1) {
|
||||||
|
result.push(next);
|
||||||
|
i += 1;
|
||||||
|
}
|
||||||
|
}
|
||||||
b'(' if !in_hex_string => {
|
b'(' if !in_hex_string => {
|
||||||
in_string += 1;
|
in_string += 1;
|
||||||
result.push(b);
|
result.push(b);
|
||||||
@@ -1854,4 +1865,26 @@ BT 30 700 Tm <41> Tj ET";
|
|||||||
"ET should be preserved after comment stripping"
|
"ET should be preserved after comment stripping"
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_strip_pdf_comments_escaped_parens() {
|
||||||
|
// An escaped `\)` must not close the string: the `%` after it is
|
||||||
|
// still string content, not a comment (subset fonts routinely map
|
||||||
|
// glyphs to `%` and to escaped parens in the same TJ array).
|
||||||
|
let input = b"[ (a\\)b) 1 (%) 1 (c) ] TJ\n";
|
||||||
|
let output = strip_pdf_comments(input);
|
||||||
|
assert_eq!(output, input.to_vec());
|
||||||
|
|
||||||
|
// Same for an escaped `\(` — must not open a phantom string that
|
||||||
|
// shields a real comment.
|
||||||
|
let input = b"(x\\(y) Tj % real comment\nET\n";
|
||||||
|
let output = strip_pdf_comments(input);
|
||||||
|
assert_eq!(output, b"(x\\(y) Tj \nET\n");
|
||||||
|
|
||||||
|
// Escaped backslash before a real close-paren: `\\` ends the escape,
|
||||||
|
// the `)` does close the string, and the comment is stripped.
|
||||||
|
let input = b"(x\\\\) Tj % comment\nET\n";
|
||||||
|
let output = strip_pdf_comments(input);
|
||||||
|
assert_eq!(output, b"(x\\\\) Tj \nET\n");
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
+327
-17
@@ -41,15 +41,110 @@ pub(crate) fn detect_columns(
|
|||||||
}
|
}
|
||||||
debug!("page {}: detect_columns: {} items", page, page_items.len());
|
debug!("page {}: detect_columns: {} items", page, page_items.len());
|
||||||
|
|
||||||
// Find page bounds
|
// The width of one ordinary page, used three ways below: as the largest
|
||||||
let x_min = page_items.iter().map(|i| i.x).fold(f32::INFINITY, f32::min);
|
// credible width for a single text run, as the size of empty gap that marks
|
||||||
let x_max = page_items
|
// content as detached, and as the span past which those checks run at all.
|
||||||
.iter()
|
// This is a heuristic, not a format rule: PDF 2.0 sets no page-size limit,
|
||||||
.map(|i| i.x + effective_width(i))
|
// and since PDF 1.6 `UserUnit` scales a page's physical size independently
|
||||||
.fold(f32::NEG_INFINITY, f32::max);
|
// of its coordinates. 14_400 units (200in at the default 1/72in unit) is
|
||||||
|
// the traditional Acrobat architectural limit, which makes it a reasonable
|
||||||
|
// "wider than any ordinary page" mark in coordinate space.
|
||||||
|
const MAX_PAGE_EXTENT: f32 = 14_400.0;
|
||||||
|
// A detached cluster is only dropped if it also holds a small minority of
|
||||||
|
// the items, so a genuine two-part layout keeps its full bounds even when
|
||||||
|
// the halves are far apart.
|
||||||
|
const MAX_TRIM_FRACTION: f32 = 0.10;
|
||||||
|
|
||||||
|
// Position and width of each item, skipping only non-finite geometry.
|
||||||
|
let finite_span = |i: &&TextItem| -> Option<(f32, f32)> {
|
||||||
|
let (left, width) = (i.x, effective_width(i));
|
||||||
|
(left.is_finite() && (left + width).is_finite()).then_some((left, width))
|
||||||
|
};
|
||||||
|
|
||||||
|
let (min_left, max_right, total) = page_items.iter().filter_map(finite_span).fold(
|
||||||
|
(f32::INFINITY, f32::NEG_INFINITY, 0usize),
|
||||||
|
|(lo, hi, n), (left, width)| (lo.min(left), hi.max(left + width), n + 1),
|
||||||
|
);
|
||||||
|
|
||||||
|
// No item had usable geometry, so there is no layout to report.
|
||||||
|
if total == 0 {
|
||||||
|
return vec![];
|
||||||
|
}
|
||||||
|
|
||||||
|
// Every threshold below (gutter margins, spanning-item width, the XY-cut
|
||||||
|
// margin) is a fraction of the page width, so a far item can set the scale
|
||||||
|
// for the whole page and shrink the effective detection window to a
|
||||||
|
// rounding error — real gutters then fall inside the margin band and a
|
||||||
|
// genuine multi-column page collapses to one region.
|
||||||
|
//
|
||||||
|
// Anything inside one page extent is ordinary, so the common case keeps the
|
||||||
|
// plain bounds and skips the work below entirely.
|
||||||
|
let (x_min, x_max) = if max_right - min_left <= MAX_PAGE_EXTENT {
|
||||||
|
(min_left, max_right)
|
||||||
|
} else {
|
||||||
|
// Discarding content needs positive evidence that it is not part of the
|
||||||
|
// layout, because a count-based rule alone cannot tell a stray from a
|
||||||
|
// sparse far sidebar. The evidence is geometric: positions are grouped
|
||||||
|
// into clusters separated by more than a whole page of continuous
|
||||||
|
// emptiness. Real content, however sparse, does not leave a void that
|
||||||
|
// large; a malformed coordinate sits alone beyond one.
|
||||||
|
let mut spans: Vec<(f32, f32)> = page_items.iter().filter_map(finite_span).collect();
|
||||||
|
spans.sort_by(|a, b| a.0.total_cmp(&b.0));
|
||||||
|
|
||||||
|
let mut core: Option<std::ops::Range<usize>> = None;
|
||||||
|
let mut start = 0usize;
|
||||||
|
for i in 1..=spans.len() {
|
||||||
|
if i < spans.len() && spans[i].0 - spans[i - 1].0 <= MAX_PAGE_EXTENT {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if core.as_ref().is_none_or(|best| i - start > best.len()) {
|
||||||
|
core = Some(start..i);
|
||||||
|
}
|
||||||
|
start = i;
|
||||||
|
}
|
||||||
|
let mut core = core.unwrap_or(0..spans.len());
|
||||||
|
|
||||||
|
// Only drop the detached clusters when they are a small minority, so a
|
||||||
|
// genuine two-part layout keeps its full bounds.
|
||||||
|
let dropped = spans.len() - core.len();
|
||||||
|
if dropped as f32 > spans.len() as f32 * MAX_TRIM_FRACTION {
|
||||||
|
core = 0..spans.len();
|
||||||
|
}
|
||||||
|
let core = &spans[core];
|
||||||
|
|
||||||
|
// Positions cannot be inflated by a bogus width, so the spread of the
|
||||||
|
// content is a sound scale for judging one. A run much wider than the
|
||||||
|
// page's own content is a malformed width — the test is relative, so a
|
||||||
|
// genuinely large page keeps its genuinely long runs.
|
||||||
|
let (lo, widest_left) = (core[0].0, core[core.len() - 1].0);
|
||||||
|
let max_run_width = (widest_left - lo) + MAX_PAGE_EXTENT;
|
||||||
|
let hi = core
|
||||||
|
.iter()
|
||||||
|
.filter(|&&(_, width)| width <= max_run_width)
|
||||||
|
.map(|&(left, width)| left + width)
|
||||||
|
.fold(widest_left, f32::max);
|
||||||
|
|
||||||
|
if lo != min_left || hi != max_right {
|
||||||
|
debug!(
|
||||||
|
"page {page}: bounds {min_left}..{max_right} exceed one page; \
|
||||||
|
dropped {dropped}/{} detached item(s), using {lo}..{hi}",
|
||||||
|
spans.len()
|
||||||
|
);
|
||||||
|
}
|
||||||
|
(lo, hi)
|
||||||
|
};
|
||||||
|
|
||||||
|
// Hard ceiling on the histogram size, independent of the trimming above:
|
||||||
|
// the bounds are attacker-influenced, so an unclamped
|
||||||
|
// `page_width / BIN_WIDTH` lets a crafted PDF force an arbitrarily large
|
||||||
|
// `vec![0u32; num_bins]` allocation. 65_536 bins covers ~128k points at
|
||||||
|
// BIN_WIDTH 2.0 — roughly 9x the largest legal page — so this never binds
|
||||||
|
// on a real layout. Kept as a bound that does not depend on the outlier
|
||||||
|
// heuristic staying correct.
|
||||||
|
const MAX_BINS: usize = 65_536;
|
||||||
|
|
||||||
let page_width = x_max - x_min;
|
let page_width = x_max - x_min;
|
||||||
if page_width < 200.0 {
|
if !page_width.is_finite() || page_width < 200.0 {
|
||||||
return vec![ColumnRegion { x_min, x_max }];
|
return vec![ColumnRegion { x_min, x_max }];
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -57,13 +152,20 @@ pub(crate) fn detect_columns(
|
|||||||
return vec![ColumnRegion { x_min, x_max }];
|
return vec![ColumnRegion { x_min, x_max }];
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Widen the bins rather than dropping the tail of the page. Clamping the
|
||||||
|
// count alone would leave anything past MAX_BINS * BIN_WIDTH outside the
|
||||||
|
// histogram, folded into the last bin, which places gutters at the wrong
|
||||||
|
// coordinates. Scaling keeps full coverage under the same allocation
|
||||||
|
// ceiling; only the resolution degrades, and only beyond ~131k points.
|
||||||
|
let bin_width = BIN_WIDTH.max(page_width / MAX_BINS as f32);
|
||||||
|
|
||||||
// Build occupancy histogram.
|
// Build occupancy histogram.
|
||||||
// Exclude items wider than 60% of page width — these are spanning items
|
// Exclude items wider than 60% of page width — these are spanning items
|
||||||
// (titles, full-width paragraphs) that would fill the gutter and prevent
|
// (titles, full-width paragraphs) that would fill the gutter and prevent
|
||||||
// detection of partial-page column layouts (e.g. two-column abstracts on
|
// detection of partial-page column layouts (e.g. two-column abstracts on
|
||||||
// a page that also has single-column introduction text).
|
// a page that also has single-column introduction text).
|
||||||
let wide_threshold = page_width * 0.6;
|
let wide_threshold = page_width * 0.6;
|
||||||
let num_bins = ((page_width / BIN_WIDTH).ceil() as usize).max(1);
|
let num_bins = ((page_width / bin_width).ceil() as usize).clamp(1, MAX_BINS);
|
||||||
let mut histogram = vec![0u32; num_bins];
|
let mut histogram = vec![0u32; num_bins];
|
||||||
|
|
||||||
for item in &page_items {
|
for item in &page_items {
|
||||||
@@ -71,8 +173,8 @@ pub(crate) fn detect_columns(
|
|||||||
if w > wide_threshold {
|
if w > wide_threshold {
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
let left = ((item.x - x_min) / BIN_WIDTH).floor() as usize;
|
let left = ((item.x - x_min) / bin_width).floor() as usize;
|
||||||
let right = (((item.x + w) - x_min) / BIN_WIDTH).ceil() as usize;
|
let right = (((item.x + w) - x_min) / bin_width).ceil() as usize;
|
||||||
let left = left.min(num_bins);
|
let left = left.min(num_bins);
|
||||||
let right = right.min(num_bins);
|
let right = right.min(num_bins);
|
||||||
for count in histogram.iter_mut().take(right).skip(left) {
|
for count in histogram.iter_mut().take(right).skip(left) {
|
||||||
@@ -109,12 +211,12 @@ pub(crate) fn detect_columns(
|
|||||||
let valleys: Vec<(usize, usize)> = valleys
|
let valleys: Vec<(usize, usize)> = valleys
|
||||||
.into_iter()
|
.into_iter()
|
||||||
.filter(|&(start, end)| {
|
.filter(|&(start, end)| {
|
||||||
let width_pts = (end - start) as f32 * BIN_WIDTH;
|
let width_pts = (end - start) as f32 * bin_width;
|
||||||
if width_pts < MIN_GUTTER_WIDTH {
|
if width_pts < MIN_GUTTER_WIDTH {
|
||||||
return false;
|
return false;
|
||||||
}
|
}
|
||||||
// Valley center must not be within 5% of page edges
|
// Valley center must not be within 5% of page edges
|
||||||
let center_pts = ((start + end) as f32 / 2.0) * BIN_WIDTH;
|
let center_pts = ((start + end) as f32 / 2.0) * bin_width;
|
||||||
center_pts > margin_threshold && center_pts < (page_width - margin_threshold)
|
center_pts > margin_threshold && center_pts < (page_width - margin_threshold)
|
||||||
})
|
})
|
||||||
.collect();
|
.collect();
|
||||||
@@ -132,7 +234,7 @@ pub(crate) fn detect_columns(
|
|||||||
&histogram,
|
&histogram,
|
||||||
num_bins,
|
num_bins,
|
||||||
x_min,
|
x_min,
|
||||||
BIN_WIDTH,
|
bin_width,
|
||||||
page_width,
|
page_width,
|
||||||
margin_threshold,
|
margin_threshold,
|
||||||
);
|
);
|
||||||
@@ -141,7 +243,7 @@ pub(crate) fn detect_columns(
|
|||||||
&rel_valleys,
|
&rel_valleys,
|
||||||
&page_items,
|
&page_items,
|
||||||
x_min,
|
x_min,
|
||||||
BIN_WIDTH,
|
bin_width,
|
||||||
x_max,
|
x_max,
|
||||||
MIN_ITEMS_PER_COLUMN,
|
MIN_ITEMS_PER_COLUMN,
|
||||||
MIN_VERTICAL_SPAN_RATIO,
|
MIN_VERTICAL_SPAN_RATIO,
|
||||||
@@ -182,7 +284,7 @@ pub(crate) fn detect_columns(
|
|||||||
&valleys,
|
&valleys,
|
||||||
&page_items,
|
&page_items,
|
||||||
x_min,
|
x_min,
|
||||||
BIN_WIDTH,
|
bin_width,
|
||||||
x_max,
|
x_max,
|
||||||
MIN_ITEMS_PER_COLUMN,
|
MIN_ITEMS_PER_COLUMN,
|
||||||
MIN_VERTICAL_SPAN_RATIO,
|
MIN_VERTICAL_SPAN_RATIO,
|
||||||
@@ -196,7 +298,7 @@ pub(crate) fn detect_columns(
|
|||||||
&valleys,
|
&valleys,
|
||||||
&page_items,
|
&page_items,
|
||||||
x_min,
|
x_min,
|
||||||
BIN_WIDTH,
|
bin_width,
|
||||||
x_max,
|
x_max,
|
||||||
MIN_ITEMS_PER_COLUMN,
|
MIN_ITEMS_PER_COLUMN,
|
||||||
MIN_VERTICAL_SPAN_RATIO,
|
MIN_VERTICAL_SPAN_RATIO,
|
||||||
@@ -1827,7 +1929,7 @@ fn split_column_stragglers(lines: Vec<TextLine>) -> (Vec<TextLine>, Vec<TextLine
|
|||||||
.unwrap();
|
.unwrap();
|
||||||
|
|
||||||
let (cs, ce) = segments[core_seg];
|
let (cs, ce) = segments[core_seg];
|
||||||
let mut core = Vec::with_capacity(ce - cs);
|
let mut core = Vec::with_capacity(ce.saturating_sub(cs));
|
||||||
let mut stragglers = Vec::new();
|
let mut stragglers = Vec::new();
|
||||||
for (i, line) in lines.into_iter().enumerate() {
|
for (i, line) in lines.into_iter().enumerate() {
|
||||||
if i >= cs && i < ce {
|
if i >= cs && i < ce {
|
||||||
@@ -2533,6 +2635,214 @@ mod tests {
|
|||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn extreme_far_coordinate_does_not_allocate_unboundedly() {
|
||||||
|
// A crafted PDF can place a text run at an arbitrary coordinate via the
|
||||||
|
// text matrix. The derived page width must not drive an unbounded
|
||||||
|
// histogram allocation (previously `page_width / BIN_WIDTH` bins with no
|
||||||
|
// upper bound would try to reserve terabytes and abort the process).
|
||||||
|
let mut items = Vec::new();
|
||||||
|
for i in 0..24 {
|
||||||
|
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
|
||||||
|
}
|
||||||
|
// Item placed 1e12 points away — 5e11 bins if left unclamped.
|
||||||
|
items.push(make_item(1, 1e12, 700.0, "Z"));
|
||||||
|
|
||||||
|
// Must return without aborting; content is preserved as a single region.
|
||||||
|
let cols = detect_columns(&items, 1, false);
|
||||||
|
assert!(!cols.is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn non_finite_coordinates_never_leak_into_region_bounds() {
|
||||||
|
// An inf/NaN coordinate must not escape as a column boundary: callers
|
||||||
|
// treat these as page/column edges.
|
||||||
|
for bad_x in [f32::INFINITY, f32::NEG_INFINITY, f32::NAN] {
|
||||||
|
let mut items = Vec::new();
|
||||||
|
for i in 0..24 {
|
||||||
|
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
|
||||||
|
}
|
||||||
|
items.push(make_item(1, bad_x, 700.0, "Z"));
|
||||||
|
|
||||||
|
for col in detect_columns(&items, 1, false) {
|
||||||
|
assert!(
|
||||||
|
col.x_min.is_finite() && col.x_max.is_finite(),
|
||||||
|
"bad_x {bad_x} leaked bounds {}..{}",
|
||||||
|
col.x_min,
|
||||||
|
col.x_max
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn all_non_finite_coordinates_yield_no_columns() {
|
||||||
|
let items: Vec<TextItem> = (0..24)
|
||||||
|
.map(|i| make_item(1, f32::NAN, 700.0 - i as f32 * 5.0, "A"))
|
||||||
|
.collect();
|
||||||
|
|
||||||
|
assert!(detect_columns(&items, 1, false).is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn one_bad_item_does_not_disable_column_detection() {
|
||||||
|
// A single stray item should not collapse a clean two-column page to
|
||||||
|
// one region. Every gutter threshold is a fraction of the page width,
|
||||||
|
// so an untrimmed outlier pushes real gutters inside the rejected
|
||||||
|
// margin band. A malformed *width* at an ordinary position poisons the
|
||||||
|
// bounds just as a malformed position does.
|
||||||
|
for (label, bad_x, bad_width) in [
|
||||||
|
("nan position", f32::NAN, 0.0),
|
||||||
|
("inf position", f32::INFINITY, 0.0),
|
||||||
|
("far position", 50_000.0, 0.0),
|
||||||
|
("very far position", 1e12, 0.0),
|
||||||
|
("huge width", 100.0, 1e12),
|
||||||
|
("inf width", 100.0, f32::INFINITY),
|
||||||
|
] {
|
||||||
|
let mut items = Vec::new();
|
||||||
|
items.extend(fill_zone(1, 30.0, 280.0, 750.0, 50.0));
|
||||||
|
items.extend(fill_zone(1, 320.0, 570.0, 750.0, 50.0));
|
||||||
|
let mut bad = make_item(1, bad_x, 400.0, "Z");
|
||||||
|
bad.width = bad_width;
|
||||||
|
items.push(bad);
|
||||||
|
|
||||||
|
let cols = detect_columns(&items, 1, false);
|
||||||
|
assert_eq!(
|
||||||
|
cols.len(),
|
||||||
|
2,
|
||||||
|
"{label}: expected 2 columns, got {}",
|
||||||
|
cols.len()
|
||||||
|
);
|
||||||
|
for col in &cols {
|
||||||
|
assert!(
|
||||||
|
col.x_max - col.x_min <= MAX_PAGE_EXTENT_FOR_TEST,
|
||||||
|
"{label}: region {}..{} exceeds one page",
|
||||||
|
col.x_min,
|
||||||
|
col.x_max
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Mirrors `MAX_PAGE_EXTENT` in `detect_columns`.
|
||||||
|
const MAX_PAGE_EXTENT_FOR_TEST: f32 = 14_400.0;
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn very_wide_page_keeps_full_histogram_coverage() {
|
||||||
|
// Beyond MAX_BINS * BIN_WIDTH (~131k points) the bins must widen rather
|
||||||
|
// than stop covering the page. Three zones: the first gutter is inside
|
||||||
|
// the old coverage limit, the second is past it. Because the first
|
||||||
|
// gutter is found, the XY-cut fallback never runs, so a truncated
|
||||||
|
// histogram silently reports two columns instead of three.
|
||||||
|
let mut items = Vec::new();
|
||||||
|
items.extend(fill_zone(1, 0.0, 60_000.0, 750.0, 700.0));
|
||||||
|
items.extend(fill_zone(1, 70_000.0, 140_000.0, 750.0, 700.0));
|
||||||
|
items.extend(fill_zone(1, 160_000.0, 200_000.0, 750.0, 700.0));
|
||||||
|
|
||||||
|
let cols = detect_columns(&items, 1, false);
|
||||||
|
assert_eq!(
|
||||||
|
cols.len(),
|
||||||
|
3,
|
||||||
|
"Expected 3 columns across a 200k-wide page, got {}",
|
||||||
|
cols.len()
|
||||||
|
);
|
||||||
|
assert!(
|
||||||
|
(140_000.0..=160_000.0).contains(&cols[1].x_max),
|
||||||
|
"second gutter at {}, expected inside the real 140k..160k gap",
|
||||||
|
cols[1].x_max
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn large_page_with_legitimately_long_runs_is_kept() {
|
||||||
|
// On a very large page, individual runs can exceed one ordinary page's
|
||||||
|
// width. They are real content, so they must not be judged malformed:
|
||||||
|
// the page keeps its columns and its full right edge.
|
||||||
|
let mut items = Vec::new();
|
||||||
|
for row in 0..30 {
|
||||||
|
let y = 750.0 - row as f32 * 14.0;
|
||||||
|
let mut left = make_item(1, 0.0, y, "Left run");
|
||||||
|
left.width = 20_000.0;
|
||||||
|
let mut right = make_item(1, 25_000.0, y, "Right run");
|
||||||
|
right.width = 20_000.0;
|
||||||
|
items.extend([left, right]);
|
||||||
|
}
|
||||||
|
|
||||||
|
let cols = detect_columns(&items, 1, false);
|
||||||
|
assert!(
|
||||||
|
!cols.is_empty(),
|
||||||
|
"a page of long-but-valid runs must still report a layout"
|
||||||
|
);
|
||||||
|
let right_edge = cols
|
||||||
|
.iter()
|
||||||
|
.map(|c| c.x_max)
|
||||||
|
.fold(f32::NEG_INFINITY, f32::max);
|
||||||
|
assert!(
|
||||||
|
right_edge > 44_000.0,
|
||||||
|
"long runs were treated as malformed: right edge {right_edge}, expected ~45_000"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn sparse_far_sidebar_on_a_large_page_is_kept() {
|
||||||
|
// A large-format page with a thin, sparsely-populated sidebar far from
|
||||||
|
// the main block. The sidebar is a small minority of the items, so an
|
||||||
|
// item-count rule alone would discard it — but nothing about its
|
||||||
|
// geometry says it is invalid, so its bounds must survive.
|
||||||
|
let mut items = Vec::new();
|
||||||
|
items.extend(fill_zone(1, 0.0, 12_000.0, 750.0, 500.0));
|
||||||
|
for i in 0..12 {
|
||||||
|
items.push(make_item(1, 24_000.0, 750.0 - i as f32 * 14.0, "Sidebar"));
|
||||||
|
}
|
||||||
|
|
||||||
|
let cols = detect_columns(&items, 1, false);
|
||||||
|
let right_edge = cols
|
||||||
|
.iter()
|
||||||
|
.map(|c| c.x_max)
|
||||||
|
.fold(f32::NEG_INFINITY, f32::max);
|
||||||
|
assert!(
|
||||||
|
right_edge > 24_000.0,
|
||||||
|
"sidebar was trimmed away: right edge {right_edge}, expected >24_000"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn genuinely_wide_layout_keeps_its_true_bounds() {
|
||||||
|
// A large-format page whose content really is spread beyond one
|
||||||
|
// ordinary page must not be trimmed to the median cluster: its far
|
||||||
|
// items are the majority, not strays.
|
||||||
|
let mut items = Vec::new();
|
||||||
|
items.extend(fill_zone(1, 100.0, 20_000.0, 750.0, 600.0));
|
||||||
|
items.extend(fill_zone(1, 22_000.0, 40_000.0, 750.0, 600.0));
|
||||||
|
|
||||||
|
let cols = detect_columns(&items, 1, false);
|
||||||
|
let widest = cols
|
||||||
|
.iter()
|
||||||
|
.map(|c| c.x_max)
|
||||||
|
.fold(f32::NEG_INFINITY, f32::max);
|
||||||
|
assert!(
|
||||||
|
widest > 35_000.0,
|
||||||
|
"wide layout was trimmed: right edge {widest}, expected ~40_000"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn oversized_but_legal_page_is_not_trimmed() {
|
||||||
|
// A wide-format page well inside the 14_400pt spec limit must keep its
|
||||||
|
// real bounds — outlier trimming is only for spans beyond a legal page.
|
||||||
|
let mut items = Vec::new();
|
||||||
|
items.extend(fill_zone(1, 100.0, 4_000.0, 750.0, 400.0));
|
||||||
|
items.extend(fill_zone(1, 4_400.0, 8_000.0, 750.0, 400.0));
|
||||||
|
|
||||||
|
let cols = detect_columns(&items, 1, false);
|
||||||
|
assert_eq!(cols.len(), 2, "Expected 2 columns, got {}", cols.len());
|
||||||
|
assert!(
|
||||||
|
cols[1].x_max > 7_000.0,
|
||||||
|
"right column should keep its true extent, got {}",
|
||||||
|
cols[1].x_max
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn two_column_regression_guard() {
|
fn two_column_regression_guard() {
|
||||||
// Standard 2-column layout with clear gutter at center
|
// Standard 2-column layout with clear gutter at center
|
||||||
|
|||||||
+311
-5
@@ -2,11 +2,51 @@
|
|||||||
|
|
||||||
use crate::types::{ItemType, TextItem};
|
use crate::types::{ItemType, TextItem};
|
||||||
use lopdf::{Document, Object, ObjectId};
|
use lopdf::{Document, Object, ObjectId};
|
||||||
use std::collections::HashMap;
|
use std::collections::{HashMap, HashSet};
|
||||||
|
|
||||||
use super::fonts::{resolve_array, resolve_dict};
|
use super::fonts::{resolve_array, resolve_dict};
|
||||||
use super::get_number;
|
use super::get_number;
|
||||||
|
|
||||||
|
/// Upper bound on the number of form-field nodes visited during a single
|
||||||
|
/// `extract_form_fields` pass. A crafted PDF can chain thousands of distinct
|
||||||
|
/// `/Kids` fields to blow the stack even without an outright reference cycle,
|
||||||
|
/// so we cap total traversal work in addition to detecting cycles.
|
||||||
|
const MAX_FORM_FIELD_NODES: usize = 100_000;
|
||||||
|
|
||||||
|
/// Upper bound on `/Kids` recursion depth. Real AcroForm hierarchies are only
|
||||||
|
/// a few levels deep (fields → child fields → widgets); a crafted PDF can chain
|
||||||
|
/// tens of thousands of distinct fields into a linear `/Kids` list that would
|
||||||
|
/// overflow the stack via depth-first recursion long before the node budget is
|
||||||
|
/// reached. This depth cap bounds the stack independently of total node count.
|
||||||
|
const MAX_FORM_FIELD_DEPTH: usize = 100;
|
||||||
|
|
||||||
|
/// Traversal budget for the AcroForm field walk. Bounds both the number of
|
||||||
|
/// distinct nodes visited *and* the total number of `/Fields`/`/Kids` entries
|
||||||
|
/// examined.
|
||||||
|
///
|
||||||
|
/// Counting `visited` alone is not enough: invalid entries (non-references) and
|
||||||
|
/// duplicate references never grow `visited`, so an oversized array full of them
|
||||||
|
/// would iterate to completion no matter how large. Charging every examined
|
||||||
|
/// entry against the same budget makes it a real cap on traversal work.
|
||||||
|
pub(crate) struct FieldWalkBudget {
|
||||||
|
visited: HashSet<ObjectId>,
|
||||||
|
examined: usize,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl FieldWalkBudget {
|
||||||
|
fn new() -> Self {
|
||||||
|
Self {
|
||||||
|
visited: HashSet::new(),
|
||||||
|
examined: 0,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// True once the budget is spent; callers must stop iterating and recursing.
|
||||||
|
fn exhausted(&self) -> bool {
|
||||||
|
self.visited.len() >= MAX_FORM_FIELD_NODES || self.examined >= MAX_FORM_FIELD_NODES
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> Vec<TextItem> {
|
pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> Vec<TextItem> {
|
||||||
let mut links = Vec::new();
|
let mut links = Vec::new();
|
||||||
|
|
||||||
@@ -146,9 +186,12 @@ pub(crate) fn extract_form_fields(
|
|||||||
Err(_) => return items,
|
Err(_) => return items,
|
||||||
};
|
};
|
||||||
|
|
||||||
|
// Borrow the array rather than cloning it: a crafted `/Fields` can be huge,
|
||||||
|
// and cloning would pay an O(n) allocation/copy before the budget check
|
||||||
|
// below can stop the work.
|
||||||
let fields = match acroform.get(b"Fields") {
|
let fields = match acroform.get(b"Fields") {
|
||||||
Ok(obj) => match resolve_array(doc, obj) {
|
Ok(obj) => match resolve_array(doc, obj) {
|
||||||
Some(arr) => arr.clone(),
|
Some(arr) => arr,
|
||||||
None => return items,
|
None => return items,
|
||||||
},
|
},
|
||||||
Err(_) => return items,
|
Err(_) => return items,
|
||||||
@@ -158,7 +201,19 @@ pub(crate) fn extract_form_fields(
|
|||||||
}
|
}
|
||||||
let annotation_pages = annotation_page_map(doc, page_map);
|
let annotation_pages = annotation_page_map(doc, page_map);
|
||||||
|
|
||||||
for field_obj in &fields {
|
// Bound the walk so a crafted PDF cannot send us into unbounded recursion
|
||||||
|
// via a `/Kids` cycle, a deep chain, or an oversized array of invalid or
|
||||||
|
// duplicate entries.
|
||||||
|
let mut budget = FieldWalkBudget::new();
|
||||||
|
|
||||||
|
for field_obj in fields {
|
||||||
|
// Stop once the budget is spent so a `/Fields` array wider than the
|
||||||
|
// budget can't burn CPU iterating entries whose walk would no-op. Charge
|
||||||
|
// every entry (including invalid ones) against the budget.
|
||||||
|
if budget.exhausted() {
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
budget.examined += 1;
|
||||||
if let Ok(field_ref) = field_obj.as_reference() {
|
if let Ok(field_ref) = field_obj.as_reference() {
|
||||||
walk_form_fields(
|
walk_form_fields(
|
||||||
doc,
|
doc,
|
||||||
@@ -168,6 +223,8 @@ pub(crate) fn extract_form_fields(
|
|||||||
page_map,
|
page_map,
|
||||||
&annotation_pages,
|
&annotation_pages,
|
||||||
&mut items,
|
&mut items,
|
||||||
|
&mut budget,
|
||||||
|
0,
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -202,6 +259,7 @@ fn annotation_page_map(
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// Recursively walk the form field tree, extracting leaf field values.
|
/// Recursively walk the form field tree, extracting leaf field values.
|
||||||
|
#[allow(clippy::too_many_arguments)]
|
||||||
pub(crate) fn walk_form_fields(
|
pub(crate) fn walk_form_fields(
|
||||||
doc: &Document,
|
doc: &Document,
|
||||||
field_id: ObjectId,
|
field_id: ObjectId,
|
||||||
@@ -210,7 +268,22 @@ pub(crate) fn walk_form_fields(
|
|||||||
page_map: &HashMap<ObjectId, u32>,
|
page_map: &HashMap<ObjectId, u32>,
|
||||||
annotation_pages: &HashMap<ObjectId, u32>,
|
annotation_pages: &HashMap<ObjectId, u32>,
|
||||||
items: &mut Vec<TextItem>,
|
items: &mut Vec<TextItem>,
|
||||||
|
budget: &mut FieldWalkBudget,
|
||||||
|
depth: usize,
|
||||||
) {
|
) {
|
||||||
|
// Guard against `/Kids` cycles and pathologically large field trees.
|
||||||
|
// Exceeding the depth cap means the chain is too deep to be a legitimate
|
||||||
|
// form (and would overflow the stack); an exhausted budget means the tree is
|
||||||
|
// too large. Both checks run *before* inserting so the visited set can never
|
||||||
|
// grow past the budget.
|
||||||
|
if depth > MAX_FORM_FIELD_DEPTH || budget.exhausted() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
// Revisiting an object ID means we hit a `/Kids` cycle.
|
||||||
|
if !budget.visited.insert(field_id) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
let field_dict = match doc.get_dictionary(field_id) {
|
let field_dict = match doc.get_dictionary(field_id) {
|
||||||
Ok(d) => d,
|
Ok(d) => d,
|
||||||
Err(_) => return,
|
Err(_) => return,
|
||||||
@@ -241,9 +314,19 @@ pub(crate) fn walk_form_fields(
|
|||||||
|
|
||||||
// Check for /Kids — if present, recurse into children
|
// Check for /Kids — if present, recurse into children
|
||||||
if let Ok(kids_obj) = field_dict.get(b"Kids") {
|
if let Ok(kids_obj) = field_dict.get(b"Kids") {
|
||||||
|
// Iterate the borrowed array directly — cloning a crafted, oversized
|
||||||
|
// `/Kids` would allocate and copy every entry before the budget check
|
||||||
|
// below could stop the work.
|
||||||
if let Some(kids) = resolve_array(doc, kids_obj) {
|
if let Some(kids) = resolve_array(doc, kids_obj) {
|
||||||
let kids = kids.clone();
|
for kid in kids {
|
||||||
for kid in &kids {
|
// Stop once the budget is spent so a `/Kids` array wider than the
|
||||||
|
// budget can't burn CPU iterating entries whose walk would no-op.
|
||||||
|
// Charge every entry (including invalid/duplicate ones) against
|
||||||
|
// the budget so this is a true traversal-work cap.
|
||||||
|
if budget.exhausted() {
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
budget.examined += 1;
|
||||||
if let Ok(kid_ref) = kid.as_reference() {
|
if let Ok(kid_ref) = kid.as_reference() {
|
||||||
walk_form_fields(
|
walk_form_fields(
|
||||||
doc,
|
doc,
|
||||||
@@ -253,6 +336,8 @@ pub(crate) fn walk_form_fields(
|
|||||||
page_map,
|
page_map,
|
||||||
annotation_pages,
|
annotation_pages,
|
||||||
items,
|
items,
|
||||||
|
budget,
|
||||||
|
depth + 1,
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -411,4 +496,225 @@ mod tests {
|
|||||||
assert_eq!(items[0].page, 2);
|
assert_eq!(items[0].page, 2);
|
||||||
assert_eq!(items[0].text, "customer: Alice");
|
assert_eq!(items[0].text, "customer: Alice");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn kids_self_cycle_does_not_overflow_stack() {
|
||||||
|
// A crafted AcroForm field that lists itself in `/Kids` must not send
|
||||||
|
// the traversal into unbounded recursion.
|
||||||
|
let mut doc = Document::new();
|
||||||
|
let field_id = doc.new_object_id();
|
||||||
|
doc.set_object(
|
||||||
|
field_id,
|
||||||
|
dictionary! {
|
||||||
|
"FT" => "Tx",
|
||||||
|
"T" => Object::string_literal("loop"),
|
||||||
|
"Kids" => vec![Object::Reference(field_id)],
|
||||||
|
},
|
||||||
|
);
|
||||||
|
let catalog_id = doc.add_object(dictionary! {
|
||||||
|
"Type" => "Catalog",
|
||||||
|
"AcroForm" => dictionary! {
|
||||||
|
"Fields" => vec![Object::Reference(field_id)],
|
||||||
|
},
|
||||||
|
});
|
||||||
|
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||||
|
|
||||||
|
let page_map = HashMap::new();
|
||||||
|
// Completes (rather than overflowing the stack) and yields no items.
|
||||||
|
let items = extract_form_fields(&doc, &page_map);
|
||||||
|
assert!(items.is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn kids_mutual_cycle_terminates() {
|
||||||
|
// Two fields that reference each other via `/Kids` form a cycle that
|
||||||
|
// must also terminate.
|
||||||
|
let mut doc = Document::new();
|
||||||
|
let field_a = doc.new_object_id();
|
||||||
|
let field_b = doc.new_object_id();
|
||||||
|
doc.set_object(
|
||||||
|
field_a,
|
||||||
|
dictionary! {
|
||||||
|
"T" => Object::string_literal("a"),
|
||||||
|
"Kids" => vec![Object::Reference(field_b)],
|
||||||
|
},
|
||||||
|
);
|
||||||
|
doc.set_object(
|
||||||
|
field_b,
|
||||||
|
dictionary! {
|
||||||
|
"T" => Object::string_literal("b"),
|
||||||
|
"Kids" => vec![Object::Reference(field_a)],
|
||||||
|
},
|
||||||
|
);
|
||||||
|
let catalog_id = doc.add_object(dictionary! {
|
||||||
|
"Type" => "Catalog",
|
||||||
|
"AcroForm" => dictionary! {
|
||||||
|
"Fields" => vec![Object::Reference(field_a)],
|
||||||
|
},
|
||||||
|
});
|
||||||
|
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||||
|
|
||||||
|
let page_map = HashMap::new();
|
||||||
|
let items = extract_form_fields(&doc, &page_map);
|
||||||
|
assert!(items.is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn deep_acyclic_kids_chain_does_not_overflow_stack() {
|
||||||
|
// A long chain of *distinct* fields (no cycle) must also terminate:
|
||||||
|
// the visited set alone would still recurse to the chain length, so
|
||||||
|
// the depth cap is what prevents a stack overflow here.
|
||||||
|
let mut doc = Document::new();
|
||||||
|
let n = MAX_FORM_FIELD_DEPTH * 500;
|
||||||
|
let ids: Vec<ObjectId> = (0..=n).map(|_| doc.new_object_id()).collect();
|
||||||
|
for i in 0..n {
|
||||||
|
doc.set_object(
|
||||||
|
ids[i],
|
||||||
|
dictionary! {
|
||||||
|
"FT" => "Tx",
|
||||||
|
"Kids" => vec![Object::Reference(ids[i + 1])],
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}
|
||||||
|
// Leaf carries a value; it sits far below the depth cap so it is never
|
||||||
|
// reached, proving traversal stops early rather than crashing.
|
||||||
|
doc.set_object(
|
||||||
|
ids[n],
|
||||||
|
dictionary! {
|
||||||
|
"FT" => "Tx",
|
||||||
|
"T" => Object::string_literal("leaf"),
|
||||||
|
"V" => Object::string_literal("x"),
|
||||||
|
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
|
||||||
|
},
|
||||||
|
);
|
||||||
|
let catalog_id = doc.add_object(dictionary! {
|
||||||
|
"Type" => "Catalog",
|
||||||
|
"AcroForm" => dictionary! {
|
||||||
|
"Fields" => vec![Object::Reference(ids[0])],
|
||||||
|
},
|
||||||
|
});
|
||||||
|
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||||
|
|
||||||
|
let page_map = HashMap::new();
|
||||||
|
let items = extract_form_fields(&doc, &page_map);
|
||||||
|
assert!(items.is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn wide_tree_traversal_stops_at_node_budget() {
|
||||||
|
// A single field with a `/Kids` array wider than the node budget must
|
||||||
|
// stop traversal at the cap rather than growing `visited` (and the work)
|
||||||
|
// without bound. Each processed leaf emits one item, so the item count
|
||||||
|
// is bounded by the budget and reaches right up to it (a couple of
|
||||||
|
// slots go to the root and the boundary node charged against the cap).
|
||||||
|
let mut doc = Document::new();
|
||||||
|
let fanout = MAX_FORM_FIELD_NODES + 50;
|
||||||
|
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
|
||||||
|
for &leaf in &leaf_ids {
|
||||||
|
doc.set_object(
|
||||||
|
leaf,
|
||||||
|
dictionary! {
|
||||||
|
"FT" => "Tx",
|
||||||
|
"V" => Object::string_literal("v"),
|
||||||
|
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}
|
||||||
|
let kids: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
|
||||||
|
let root_id = doc.add_object(dictionary! {
|
||||||
|
"T" => Object::string_literal("root"),
|
||||||
|
"Kids" => kids,
|
||||||
|
});
|
||||||
|
let catalog_id = doc.add_object(dictionary! {
|
||||||
|
"Type" => "Catalog",
|
||||||
|
"AcroForm" => dictionary! {
|
||||||
|
"Fields" => vec![Object::Reference(root_id)],
|
||||||
|
},
|
||||||
|
});
|
||||||
|
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||||
|
|
||||||
|
let page_map = HashMap::new();
|
||||||
|
let items = extract_form_fields(&doc, &page_map);
|
||||||
|
// Extraction stops at the budget: bounded above by the cap, and it gets
|
||||||
|
// right up to it (allowing a small delta for the root/boundary nodes
|
||||||
|
// charged against the budget).
|
||||||
|
assert!(items.len() <= MAX_FORM_FIELD_NODES);
|
||||||
|
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn wide_top_level_fields_stop_at_node_budget() {
|
||||||
|
// A top-level `/Fields` array wider than the budget must also stop at
|
||||||
|
// the cap: the item count is bounded by the budget and reaches right up
|
||||||
|
// to it.
|
||||||
|
let mut doc = Document::new();
|
||||||
|
let fanout = MAX_FORM_FIELD_NODES + 50;
|
||||||
|
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
|
||||||
|
for &leaf in &leaf_ids {
|
||||||
|
doc.set_object(
|
||||||
|
leaf,
|
||||||
|
dictionary! {
|
||||||
|
"FT" => "Tx",
|
||||||
|
"V" => Object::string_literal("v"),
|
||||||
|
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}
|
||||||
|
let fields: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
|
||||||
|
let catalog_id = doc.add_object(dictionary! {
|
||||||
|
"Type" => "Catalog",
|
||||||
|
"AcroForm" => dictionary! {
|
||||||
|
"Fields" => fields,
|
||||||
|
},
|
||||||
|
});
|
||||||
|
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||||
|
|
||||||
|
let page_map = HashMap::new();
|
||||||
|
let items = extract_form_fields(&doc, &page_map);
|
||||||
|
assert!(items.len() <= MAX_FORM_FIELD_NODES);
|
||||||
|
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn duplicate_and_invalid_kids_entries_stop_at_budget() {
|
||||||
|
// Duplicate references and non-reference junk never grow `visited`, so
|
||||||
|
// without charging examined entries against the budget an oversized
|
||||||
|
// array of them would iterate to completion. The walk must still
|
||||||
|
// terminate and extract the single real leaf exactly once.
|
||||||
|
let mut doc = Document::new();
|
||||||
|
let leaf_id = doc.new_object_id();
|
||||||
|
doc.set_object(
|
||||||
|
leaf_id,
|
||||||
|
dictionary! {
|
||||||
|
"FT" => "Tx",
|
||||||
|
"V" => Object::string_literal("v"),
|
||||||
|
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
|
||||||
|
},
|
||||||
|
);
|
||||||
|
// A `/Kids` array far wider than the budget: half duplicate references
|
||||||
|
// to the same leaf, half invalid (null) entries.
|
||||||
|
let mut kids: Vec<Object> = Vec::new();
|
||||||
|
for i in 0..(MAX_FORM_FIELD_NODES * 2) {
|
||||||
|
if i % 2 == 0 {
|
||||||
|
kids.push(Object::Reference(leaf_id));
|
||||||
|
} else {
|
||||||
|
kids.push(Object::Null);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let root_id = doc.add_object(dictionary! {
|
||||||
|
"T" => Object::string_literal("root"),
|
||||||
|
"Kids" => kids,
|
||||||
|
});
|
||||||
|
let catalog_id = doc.add_object(dictionary! {
|
||||||
|
"Type" => "Catalog",
|
||||||
|
"AcroForm" => dictionary! {
|
||||||
|
"Fields" => vec![Object::Reference(root_id)],
|
||||||
|
},
|
||||||
|
});
|
||||||
|
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||||
|
|
||||||
|
let page_map = HashMap::new();
|
||||||
|
let items = extract_form_fields(&doc, &page_map);
|
||||||
|
assert_eq!(items.len(), 1);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
+40
-3
@@ -4566,9 +4566,13 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// Try to parse uniXXXX format
|
// Try to parse uniXXXX format.
|
||||||
if name.starts_with("uni") && name.len() >= 7 {
|
// Use `get` rather than a byte-length check + slice: `name` can contain
|
||||||
if let Ok(code) = u32::from_str_radix(&name[3..7], 16) {
|
// non-ASCII bytes (e.g. U+FFFD from lossy UTF-8 decoding of an attacker
|
||||||
|
// controlled /Differences name), so byte index 7 may not be a char
|
||||||
|
// boundary and `&name[3..7]` would panic.
|
||||||
|
if let Some(hex) = name.strip_prefix("uni").and_then(|rest| rest.get(..4)) {
|
||||||
|
if let Ok(code) = u32::from_str_radix(hex, 16) {
|
||||||
// Strip PUA F000 offset: uniF0XX → U+00XX (Windows Symbol encoding convention)
|
// Strip PUA F000 offset: uniF0XX → U+00XX (Windows Symbol encoding convention)
|
||||||
let code = if (0xF000..=0xF0FF).contains(&code) {
|
let code = if (0xF000..=0xF0FF).contains(&code) {
|
||||||
code - 0xF000
|
code - 0xF000
|
||||||
@@ -4588,3 +4592,36 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
|
|||||||
|
|
||||||
None
|
None
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[cfg(test)]
|
||||||
|
mod tests {
|
||||||
|
use super::*;
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn uni_hex_parsing() {
|
||||||
|
assert_eq!(glyph_to_char("uni0041"), Some('A'));
|
||||||
|
assert_eq!(glyph_to_char("uni00e9"), Some('\u{00e9}'));
|
||||||
|
// PUA F0xx symbol-encoding offset is stripped.
|
||||||
|
assert_eq!(glyph_to_char("uniF041"), Some('A'));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn u_hex_parsing() {
|
||||||
|
assert_eq!(glyph_to_char("u0041"), Some('A'));
|
||||||
|
assert_eq!(glyph_to_char("u1F600"), Some('\u{1F600}'));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn non_ascii_uni_name_does_not_panic() {
|
||||||
|
// A crafted /Differences name like `/uni#80#80#80#80` decodes via
|
||||||
|
// from_utf8_lossy into "uni" followed by four U+FFFD replacements.
|
||||||
|
// Byte index 7 lands mid-character, so a naive `&name[3..7]` slice
|
||||||
|
// would panic. It must be handled gracefully instead.
|
||||||
|
let crafted = format!("uni{0}{0}{0}{0}", '\u{FFFD}');
|
||||||
|
assert_eq!(glyph_to_char(&crafted), None);
|
||||||
|
|
||||||
|
// Assorted non-ASCII bytes right after the "uni" prefix.
|
||||||
|
assert_eq!(glyph_to_char("uni\u{FFFD}bc"), None);
|
||||||
|
assert_eq!(glyph_to_char("uni\u{00e9}00"), None);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
+385
-7
@@ -491,7 +491,14 @@ pub fn extract_pages_markdown_mem(
|
|||||||
|
|
||||||
// Tables need the original numeric cells; columns use folio-cleaned
|
// Tables need the original numeric cells; columns use folio-cleaned
|
||||||
// evidence so removed page numbers cannot create false layout metadata.
|
// evidence so removed page numbers cannot create false layout metadata.
|
||||||
let complexity = compute_layout_complexity(&all_items, &filtered_items, &all_rects, &all_lines);
|
let chart_regions = markdown::chart_regions_by_page(&all_items, &all_rects, &all_lines);
|
||||||
|
let complexity = compute_layout_complexity_with_chart_regions(
|
||||||
|
&all_items,
|
||||||
|
&filtered_items,
|
||||||
|
&all_rects,
|
||||||
|
&all_lines,
|
||||||
|
&chart_regions,
|
||||||
|
);
|
||||||
|
|
||||||
// Compute font stats from full document (cross-page consistency).
|
// Compute font stats from full document (cross-page consistency).
|
||||||
let font_stats = markdown::analysis::calculate_font_stats_from_items(&filtered_items);
|
let font_stats = markdown::analysis::calculate_font_stats_from_items(&filtered_items);
|
||||||
@@ -509,6 +516,7 @@ pub fn extract_pages_markdown_mem(
|
|||||||
let mut results = Vec::with_capacity(pages_slice.len());
|
let mut results = Vec::with_capacity(pages_slice.len());
|
||||||
let mut pages_needing_ocr = Vec::new();
|
let mut pages_needing_ocr = Vec::new();
|
||||||
let mut ocr_reasons_by_page = BTreeMap::new();
|
let mut ocr_reasons_by_page = BTreeMap::new();
|
||||||
|
let lopdf_pages = doc.get_pages();
|
||||||
|
|
||||||
for &page_0idx in pages_slice {
|
for &page_0idx in pages_slice {
|
||||||
// Out-of-range pages → empty + needs_ocr
|
// Out-of-range pages → empty + needs_ocr
|
||||||
@@ -542,6 +550,25 @@ pub fn extract_pages_markdown_mem(
|
|||||||
let has_gid = gid_pages.contains(&page_1idx);
|
let has_gid = gid_pages.contains(&page_1idx);
|
||||||
let has_text_quality_issue = text_quality.pages_needing_ocr.contains(&page_1idx);
|
let has_text_quality_issue = text_quality.pages_needing_ocr.contains(&page_1idx);
|
||||||
|
|
||||||
|
// A page can extract cleanly (no decoding issues, non-empty text)
|
||||||
|
// while still being fundamentally a scan: a full-page raster with
|
||||||
|
// a little genuine native text drawn over it (a header, a stamp, a
|
||||||
|
// cover-sheet annotation). Text-quality signals alone can't see
|
||||||
|
// that — consult the same "large background image" signal
|
||||||
|
// classify_pdf/detect_pdf_type already uses, so the two APIs can't
|
||||||
|
// silently disagree on whether a page needs OCR. See #227.
|
||||||
|
// Also covers vector-outlined text (glyphs drawn as paths, not
|
||||||
|
// shown via a text-showing operator): a hybrid page with real
|
||||||
|
// embedded-font body text elsewhere would otherwise still extract
|
||||||
|
// non-empty, non-garbled markdown and miss OCR routing entirely.
|
||||||
|
// detect_from_document's Mixed-type per-page routing always sends
|
||||||
|
// these pages to OCR; mirror that here too. Both signals share one
|
||||||
|
// analyze_page_content pass — see page_ocr_signals's doc comment.
|
||||||
|
let (has_template_image, has_vector_text) = lopdf_pages
|
||||||
|
.get(&page_1idx)
|
||||||
|
.map(|&page_id| detector::page_ocr_signals(&doc, page_id))
|
||||||
|
.unwrap_or((false, false));
|
||||||
|
|
||||||
// Build markdown with document-wide font stats
|
// Build markdown with document-wide font stats
|
||||||
let options = MarkdownOptions {
|
let options = MarkdownOptions {
|
||||||
base_font_size: Some(font_stats.most_common_size),
|
base_font_size: Some(font_stats.most_common_size),
|
||||||
@@ -565,6 +592,7 @@ pub fn extract_pages_markdown_mem(
|
|||||||
page_count,
|
page_count,
|
||||||
prefiltered_page_number_pages: Some(&removed_page_number_pages),
|
prefiltered_page_number_pages: Some(&removed_page_number_pages),
|
||||||
prefiltered_page_number_mask: Some(&page_number_removal_mask),
|
prefiltered_page_number_mask: Some(&page_number_removal_mask),
|
||||||
|
precomputed_chart_regions: Some(&chart_regions),
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
};
|
};
|
||||||
@@ -578,10 +606,20 @@ pub fn extract_pages_markdown_mem(
|
|||||||
OCR_REASON_SUSPECTED_GARBLED_TEXT,
|
OCR_REASON_SUSPECTED_GARBLED_TEXT,
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
if has_template_image {
|
||||||
|
add_ocr_reason(&mut ocr_reasons_by_page, page_1idx, OCR_REASON_SCANNED);
|
||||||
|
}
|
||||||
|
if has_vector_text {
|
||||||
|
add_ocr_reason(&mut ocr_reasons_by_page, page_1idx, OCR_REASON_VECTOR_TEXT);
|
||||||
|
}
|
||||||
let ocr_reason = page_ocr_reason(&ocr_reasons_by_page, page_1idx);
|
let ocr_reason = page_ocr_reason(&ocr_reasons_by_page, page_1idx);
|
||||||
|
|
||||||
let needs_ocr =
|
let needs_ocr = ocr_reason.is_some()
|
||||||
ocr_reason.is_some() || md.trim().is_empty() || has_gid || is_garbage_text(&md);
|
|| md.trim().is_empty()
|
||||||
|
|| has_gid
|
||||||
|
|| is_garbage_text(&md)
|
||||||
|
|| has_template_image
|
||||||
|
|| has_vector_text;
|
||||||
|
|
||||||
if needs_ocr {
|
if needs_ocr {
|
||||||
pages_needing_ocr.push(page_1idx);
|
pages_needing_ocr.push(page_1idx);
|
||||||
@@ -619,6 +657,80 @@ pub fn extract_pages_markdown<P: AsRef<Path>>(
|
|||||||
extract_pages_markdown_mem(&buffer, pages)
|
extract_pages_markdown_mem(&buffer, pages)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// =========================================================================
|
||||||
|
// Structure-tree element extraction (tagged PDFs)
|
||||||
|
// =========================================================================
|
||||||
|
|
||||||
|
/// One structure-tree element reference from a tagged PDF, resolved to a
|
||||||
|
/// page and Marked Content ID.
|
||||||
|
///
|
||||||
|
/// Join `(page, mcid)` against [`TextItem::page`] / [`TextItem::mcid`] from
|
||||||
|
/// [`extract_text_with_positions`] to attach semantic roles (heading levels,
|
||||||
|
/// paragraphs, table cells, …) to extracted text.
|
||||||
|
#[derive(Debug, Clone)]
|
||||||
|
pub struct StructureElement {
|
||||||
|
/// 1-indexed page number (matches [`TextItem::page`]).
|
||||||
|
pub page: u32,
|
||||||
|
/// Marked Content ID from the page's content stream (matches
|
||||||
|
/// [`TextItem::mcid`]).
|
||||||
|
pub mcid: i64,
|
||||||
|
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
|
||||||
|
/// Custom tags are resolved through the document's `/RoleMap`; tags
|
||||||
|
/// with no standard mapping are returned verbatim.
|
||||||
|
pub role: String,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Extract structure-tree element references from a tagged PDF in memory.
|
||||||
|
///
|
||||||
|
/// Parses `/StructTreeRoot` (when present) and returns one entry per
|
||||||
|
/// marked-content reference, resolved to its 1-indexed page, MCID, and
|
||||||
|
/// structure type name. Returns an empty list when the PDF is not tagged.
|
||||||
|
///
|
||||||
|
/// Pass `Some(&[...])` with 1-indexed page numbers (matching
|
||||||
|
/// [`TextItem::page`]) to restrict output to those pages; pass `None` for
|
||||||
|
/// the whole document. Entries are sorted by `(page, mcid)`.
|
||||||
|
pub fn extract_structure_elements_mem(
|
||||||
|
buffer: &[u8],
|
||||||
|
pages: Option<&[u32]>,
|
||||||
|
) -> Result<Vec<StructureElement>, PdfError> {
|
||||||
|
validate_pdf_bytes(buffer)?;
|
||||||
|
let (doc, _page_count) = load_document_from_mem(buffer)?;
|
||||||
|
let Some(tree) = structure_tree::StructTree::from_doc(&doc) else {
|
||||||
|
return Ok(Vec::new());
|
||||||
|
};
|
||||||
|
let page_ids = doc.get_pages();
|
||||||
|
let roles = tree.mcid_to_roles(&page_ids);
|
||||||
|
|
||||||
|
let page_filter: Option<HashSet<u32>> = pages.map(|p| p.iter().copied().collect());
|
||||||
|
let mut elements: Vec<StructureElement> = roles
|
||||||
|
.into_iter()
|
||||||
|
.filter(|(page, _)| page_filter.as_ref().is_none_or(|f| f.contains(page)))
|
||||||
|
.flat_map(|(page, mcids)| {
|
||||||
|
mcids.into_iter().map(move |(mcid, role)| StructureElement {
|
||||||
|
page,
|
||||||
|
mcid,
|
||||||
|
role: role.name().to_string(),
|
||||||
|
})
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
elements.sort_unstable_by_key(|e| (e.page, e.mcid));
|
||||||
|
Ok(elements)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Path-based wrapper for [`extract_structure_elements_mem`].
|
||||||
|
///
|
||||||
|
/// Reads the PDF from disk and extracts structure-tree element references.
|
||||||
|
/// Pass `None` for `pages` to return the whole document, or `Some(&[...])`
|
||||||
|
/// to restrict to specific 1-indexed pages.
|
||||||
|
pub fn extract_structure_elements<P: AsRef<Path>>(
|
||||||
|
path: P,
|
||||||
|
pages: Option<&[u32]>,
|
||||||
|
) -> Result<Vec<StructureElement>, PdfError> {
|
||||||
|
validate_pdf_file(&path)?;
|
||||||
|
let buffer = std::fs::read(path.as_ref())?;
|
||||||
|
extract_structure_elements_mem(&buffer, pages)
|
||||||
|
}
|
||||||
|
|
||||||
// =========================================================================
|
// =========================================================================
|
||||||
// Region-based text extraction (for hybrid OCR pipelines)
|
// Region-based text extraction (for hybrid OCR pipelines)
|
||||||
// =========================================================================
|
// =========================================================================
|
||||||
@@ -3515,6 +3627,7 @@ fn repair_pdf_container_candidates(buf: &[u8]) -> Vec<Vec<u8>> {
|
|||||||
let mut candidates = Vec::new();
|
let mut candidates = Vec::new();
|
||||||
|
|
||||||
add_repair_candidate(&mut candidates, append_missing_eof_marker(buf), buf);
|
add_repair_candidate(&mut candidates, append_missing_eof_marker(buf), buf);
|
||||||
|
add_repair_candidate(&mut candidates, recover_startxref_pointer(buf), buf);
|
||||||
|
|
||||||
let stripped = strip_leading_pdf_container_bytes(buf);
|
let stripped = strip_leading_pdf_container_bytes(buf);
|
||||||
if let Some(stripped_buf) = stripped.as_deref() {
|
if let Some(stripped_buf) = stripped.as_deref() {
|
||||||
@@ -3524,11 +3637,112 @@ fn repair_pdf_container_candidates(buf: &[u8]) -> Vec<Vec<u8>> {
|
|||||||
append_missing_eof_marker(stripped_buf),
|
append_missing_eof_marker(stripped_buf),
|
||||||
buf,
|
buf,
|
||||||
);
|
);
|
||||||
|
add_repair_candidate(
|
||||||
|
&mut candidates,
|
||||||
|
recover_startxref_pointer(stripped_buf),
|
||||||
|
buf,
|
||||||
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
candidates
|
candidates
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Some PDF writers emit a `startxref` pointer that doesn't actually point
|
||||||
|
/// at the cross-reference table — a single corrupted byte in the offset is
|
||||||
|
/// enough. lopdf trusts that pointer outright and fails to load rather than
|
||||||
|
/// searching for the real table, unlike pypdf/pdfium which both recover by
|
||||||
|
/// locating it directly. This finds the real (classic, non-stream) `xref`
|
||||||
|
/// table by scanning for the keyword — validating that a plausible
|
||||||
|
/// subsection header follows, not just any standalone "xref" token, since
|
||||||
|
/// this crate processes untrusted input and a coincidental match inside
|
||||||
|
/// unrelated stream/string content must not get "repaired" against a bogus
|
||||||
|
/// offset (lopdf would then load successfully against garbage instead of
|
||||||
|
/// returning a clean error) — and appends a corrected trailing
|
||||||
|
/// `startxref`/`%%EOF` block. lopdf's own `get_xref_start` always uses the
|
||||||
|
/// *last* `%%EOF` in the final 512 bytes of the buffer, so ours
|
||||||
|
/// transparently supersedes the broken one without needing to touch
|
||||||
|
/// anything already in the file.
|
||||||
|
///
|
||||||
|
/// Doesn't cover cross-reference *streams* (`N 0 obj << /Type /XRef ...`,
|
||||||
|
/// used by some PDF 1.5+ writers instead of a classic table) — recovering
|
||||||
|
/// those needs the containing object's number, not just a byte offset.
|
||||||
|
fn recover_startxref_pointer(buf: &[u8]) -> Option<Vec<u8>> {
|
||||||
|
let xref_pos = find_last_valid_xref_table_start(buf)?;
|
||||||
|
|
||||||
|
let mut repaired = Vec::with_capacity(buf.len() + 32);
|
||||||
|
repaired.extend_from_slice(buf);
|
||||||
|
if !repaired.ends_with(b"\n") {
|
||||||
|
repaired.push(b'\n');
|
||||||
|
}
|
||||||
|
repaired.extend_from_slice(format!("startxref\n{xref_pos}\n%%EOF\n").as_bytes());
|
||||||
|
Some(repaired)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Finds the last standalone `xref` token in `buf` that is immediately
|
||||||
|
/// followed by a plausible classic cross-reference subsection header
|
||||||
|
/// (`<start-id> <count>`, e.g. "0 6") — the shape every real classic xref
|
||||||
|
/// table starts with. A single reverse byte scan: O(n) even on a
|
||||||
|
/// pathological buffer with many non-matching or non-standalone "xref"
|
||||||
|
/// occurrences, unlike repeatedly re-searching a shrinking prefix.
|
||||||
|
fn find_last_valid_xref_table_start(buf: &[u8]) -> Option<usize> {
|
||||||
|
const KEYWORD: &[u8] = b"xref";
|
||||||
|
if buf.len() < KEYWORD.len() {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
let mut pos = buf.len() - KEYWORD.len();
|
||||||
|
loop {
|
||||||
|
if &buf[pos..pos + KEYWORD.len()] == KEYWORD {
|
||||||
|
let before_ok = pos == 0 || buf[pos - 1].is_ascii_whitespace();
|
||||||
|
let after_ok = buf
|
||||||
|
.get(pos + KEYWORD.len())
|
||||||
|
.is_none_or(|c| c.is_ascii_whitespace());
|
||||||
|
if before_ok && after_ok && looks_like_xref_subsection_header(buf, pos + KEYWORD.len())
|
||||||
|
{
|
||||||
|
return Some(pos);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if pos == 0 {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
pos -= 1;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Checks that `buf[pos..]` starts (after whitespace) with two
|
||||||
|
/// whitespace-separated runs of ASCII digits — `<start-id> <count>`, the
|
||||||
|
/// first subsection header of a classic PDF cross-reference table.
|
||||||
|
fn looks_like_xref_subsection_header(buf: &[u8], pos: usize) -> bool {
|
||||||
|
fn skip_ws(buf: &[u8], mut pos: usize) -> usize {
|
||||||
|
while buf.get(pos).is_some_and(u8::is_ascii_whitespace) {
|
||||||
|
pos += 1;
|
||||||
|
}
|
||||||
|
pos
|
||||||
|
}
|
||||||
|
fn skip_digits(buf: &[u8], mut pos: usize) -> usize {
|
||||||
|
while buf.get(pos).is_some_and(u8::is_ascii_digit) {
|
||||||
|
pos += 1;
|
||||||
|
}
|
||||||
|
pos
|
||||||
|
}
|
||||||
|
|
||||||
|
let pos = skip_ws(buf, pos);
|
||||||
|
let after_first_digits = skip_digits(buf, pos);
|
||||||
|
if after_first_digits == pos {
|
||||||
|
return false; // no start-id
|
||||||
|
}
|
||||||
|
let sep = skip_ws(buf, after_first_digits);
|
||||||
|
if sep == after_first_digits {
|
||||||
|
return false; // start-id and count must be whitespace-separated
|
||||||
|
}
|
||||||
|
let after_count = skip_digits(buf, sep);
|
||||||
|
if after_count == sep {
|
||||||
|
return false; // no count
|
||||||
|
}
|
||||||
|
// The count run must end at whitespace/buffer-end, not run into trailing
|
||||||
|
// garbage (e.g. a coincidental "xref\n0 6garbage" in stream content).
|
||||||
|
buf.get(after_count).is_none_or(u8::is_ascii_whitespace)
|
||||||
|
}
|
||||||
|
|
||||||
fn add_repair_candidate(
|
fn add_repair_candidate(
|
||||||
candidates: &mut Vec<Vec<u8>>,
|
candidates: &mut Vec<Vec<u8>>,
|
||||||
candidate: Option<Vec<u8>>,
|
candidate: Option<Vec<u8>>,
|
||||||
@@ -3816,7 +4030,14 @@ fn process_document(
|
|||||||
|
|
||||||
let text_quality = analyze_text_quality(&items);
|
let text_quality = analyze_text_quality(&items);
|
||||||
merge_ocr_reasons(&mut ocr_reasons_by_page, text_quality.reasons_by_page);
|
merge_ocr_reasons(&mut ocr_reasons_by_page, text_quality.reasons_by_page);
|
||||||
let layout = compute_layout_complexity(&items, &layout_items, &rects, &lines);
|
let chart_regions = markdown::chart_regions_by_page(&items, &rects, &lines);
|
||||||
|
let layout = compute_layout_complexity_with_chart_regions(
|
||||||
|
&items,
|
||||||
|
&layout_items,
|
||||||
|
&rects,
|
||||||
|
&lines,
|
||||||
|
&chart_regions,
|
||||||
|
);
|
||||||
|
|
||||||
let md = if options.mode == ProcessMode::Analyze {
|
let md = if options.mode == ProcessMode::Analyze {
|
||||||
None
|
None
|
||||||
@@ -3833,6 +4054,7 @@ fn process_document(
|
|||||||
page_count,
|
page_count,
|
||||||
prefiltered_page_number_pages: Some(&removed_pages),
|
prefiltered_page_number_pages: Some(&removed_pages),
|
||||||
prefiltered_page_number_mask: Some(removal_mask.as_slice()),
|
prefiltered_page_number_mask: Some(removal_mask.as_slice()),
|
||||||
|
precomputed_chart_regions: Some(&chart_regions),
|
||||||
},
|
},
|
||||||
))
|
))
|
||||||
};
|
};
|
||||||
@@ -5633,11 +5855,29 @@ fn select_items_with_document_folio_context(
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// Analyse extracted items and rects for layout complexity.
|
/// Analyse extracted items and rects for layout complexity.
|
||||||
|
#[cfg(test)]
|
||||||
fn compute_layout_complexity(
|
fn compute_layout_complexity(
|
||||||
items: &[types::TextItem],
|
items: &[types::TextItem],
|
||||||
column_items: &[types::TextItem],
|
column_items: &[types::TextItem],
|
||||||
rects: &[types::PdfRect],
|
rects: &[types::PdfRect],
|
||||||
lines: &[types::PdfLine],
|
lines: &[types::PdfLine],
|
||||||
|
) -> LayoutComplexity {
|
||||||
|
let page_chart_regions = markdown::chart_regions_by_page(items, rects, lines);
|
||||||
|
compute_layout_complexity_with_chart_regions(
|
||||||
|
items,
|
||||||
|
column_items,
|
||||||
|
rects,
|
||||||
|
lines,
|
||||||
|
&page_chart_regions,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn compute_layout_complexity_with_chart_regions(
|
||||||
|
items: &[types::TextItem],
|
||||||
|
column_items: &[types::TextItem],
|
||||||
|
rects: &[types::PdfRect],
|
||||||
|
lines: &[types::PdfLine],
|
||||||
|
page_chart_regions: &markdown::PageChartRegions,
|
||||||
) -> LayoutComplexity {
|
) -> LayoutComplexity {
|
||||||
use markdown::analysis::calculate_font_stats_from_items;
|
use markdown::analysis::calculate_font_stats_from_items;
|
||||||
|
|
||||||
@@ -5659,6 +5899,10 @@ fn compute_layout_complexity(
|
|||||||
let owned_items: Vec<types::TextItem> = page_items.iter().map(|i| (*i).clone()).collect();
|
let owned_items: Vec<types::TextItem> = page_items.iter().map(|i| (*i).clone()).collect();
|
||||||
let page_content_width = tables::content_width(&owned_items);
|
let page_content_width = tables::content_width(&owned_items);
|
||||||
let bands = markdown::split_side_by_side(&owned_items);
|
let bands = markdown::split_side_by_side(&owned_items);
|
||||||
|
let chart_regions = page_chart_regions
|
||||||
|
.get(&page)
|
||||||
|
.map(Vec::as_slice)
|
||||||
|
.unwrap_or_default();
|
||||||
|
|
||||||
let band_ranges: Vec<(f32, f32)> = if bands.is_empty() {
|
let band_ranges: Vec<(f32, f32)> = if bands.is_empty() {
|
||||||
// Single region — use sentinel range that includes everything
|
// Single region — use sentinel range that includes everything
|
||||||
@@ -5673,7 +5917,8 @@ fn compute_layout_complexity(
|
|||||||
let band_items: Vec<types::TextItem> = owned_items
|
let band_items: Vec<types::TextItem> = owned_items
|
||||||
.iter()
|
.iter()
|
||||||
.filter(|item| {
|
.filter(|item| {
|
||||||
x_lo == f32::MIN || (item.x >= x_lo - margin && item.x < x_hi + margin)
|
(x_lo == f32::MIN || (item.x >= x_lo - margin && item.x < x_hi + margin))
|
||||||
|
&& !markdown::item_is_in_chart_region(item, chart_regions)
|
||||||
})
|
})
|
||||||
.cloned()
|
.cloned()
|
||||||
.collect();
|
.collect();
|
||||||
@@ -5725,8 +5970,20 @@ fn compute_layout_complexity(
|
|||||||
}
|
}
|
||||||
|
|
||||||
let mut pages_with_columns: Vec<u32> = Vec::new();
|
let mut pages_with_columns: Vec<u32> = Vec::new();
|
||||||
for page in seen_pages {
|
for &page in &seen_pages {
|
||||||
let cols = extractor::detect_columns(column_items, page, pages_with_tables.contains(&page));
|
let chart_regions = page_chart_regions
|
||||||
|
.get(&page)
|
||||||
|
.map(Vec::as_slice)
|
||||||
|
.unwrap_or_default();
|
||||||
|
let page_column_items: Vec<types::TextItem> = column_items
|
||||||
|
.iter()
|
||||||
|
.filter(|item| {
|
||||||
|
item.page == page && !markdown::item_is_in_chart_region(item, chart_regions)
|
||||||
|
})
|
||||||
|
.cloned()
|
||||||
|
.collect();
|
||||||
|
let cols =
|
||||||
|
extractor::detect_columns(&page_column_items, page, pages_with_tables.contains(&page));
|
||||||
if cols.len() >= 2 {
|
if cols.len() >= 2 {
|
||||||
pages_with_columns.push(page);
|
pages_with_columns.push(page);
|
||||||
}
|
}
|
||||||
@@ -5955,6 +6212,66 @@ mod tests {
|
|||||||
assert!(filtered.pages_with_columns.is_empty());
|
assert!(filtered.pages_with_columns.is_empty());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn dense_chart_panel_is_not_reported_as_a_table() {
|
||||||
|
let mut items: Vec<TextItem> = (0..8)
|
||||||
|
.flat_map(|row| {
|
||||||
|
(0..6).map(move |column| {
|
||||||
|
test_item(
|
||||||
|
&format!("{}", row * 10 + column),
|
||||||
|
105.0 + column as f32 * 35.0,
|
||||||
|
525.0 - row as f32 * 15.0,
|
||||||
|
24.0,
|
||||||
|
10.0,
|
||||||
|
)
|
||||||
|
})
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
for row in 0..6 {
|
||||||
|
items.push(test_item(
|
||||||
|
"Left column prose continues here",
|
||||||
|
80.0,
|
||||||
|
320.0 - row as f32 * 15.0,
|
||||||
|
160.0,
|
||||||
|
10.0,
|
||||||
|
));
|
||||||
|
items.push(test_item(
|
||||||
|
"Right column prose continues here",
|
||||||
|
300.0,
|
||||||
|
320.0 - row as f32 * 15.0,
|
||||||
|
160.0,
|
||||||
|
10.0,
|
||||||
|
));
|
||||||
|
}
|
||||||
|
let mut lines: Vec<PdfLine> = (0..30)
|
||||||
|
.map(|column| PdfLine {
|
||||||
|
x1: 100.0 + column as f32 * 8.0,
|
||||||
|
y1: 400.0,
|
||||||
|
x2: 100.0 + column as f32 * 8.0,
|
||||||
|
y2: 550.0,
|
||||||
|
page: 1,
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
lines.extend((0..6).map(|row| PdfLine {
|
||||||
|
x1: 100.0,
|
||||||
|
y1: 400.0 + row as f32 * 30.0,
|
||||||
|
x2: 332.0,
|
||||||
|
y2: 400.0 + row as f32 * 30.0,
|
||||||
|
page: 1,
|
||||||
|
}));
|
||||||
|
let rects = vec![PdfRect {
|
||||||
|
x: 80.0,
|
||||||
|
y: 350.0,
|
||||||
|
width: 280.0,
|
||||||
|
height: 240.0,
|
||||||
|
page: 1,
|
||||||
|
}];
|
||||||
|
|
||||||
|
let complexity = compute_layout_complexity(&items, &items, &rects, &lines);
|
||||||
|
|
||||||
|
assert!(complexity.pages_with_tables.is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn page_selection_keeps_document_wide_folio_layout_decisions() {
|
fn page_selection_keeps_document_wide_folio_layout_decisions() {
|
||||||
let mut items = Vec::new();
|
let mut items = Vec::new();
|
||||||
@@ -6958,4 +7275,65 @@ mod tests {
|
|||||||
// Pre-filled cell was not touched.
|
// Pre-filled cell was not touched.
|
||||||
assert_eq!(cells[1].text, "Pre-filled");
|
assert_eq!(cells[1].text, "Pre-filled");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// -- recover_startxref_pointer / find_last_valid_xref_table_start ------
|
||||||
|
//
|
||||||
|
// Direct unit tests on the byte-level scan, addressing review feedback
|
||||||
|
// on #230: a coincidental standalone "xref" token that isn't actually
|
||||||
|
// followed by a subsection header (start-id + count) must not be
|
||||||
|
// treated as a real table — accepting it would let lopdf "succeed"
|
||||||
|
// against a bogus offset and silently return garbled/empty content
|
||||||
|
// instead of a clean error.
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn find_xref_rejects_standalone_token_without_subsection_header() {
|
||||||
|
// "xref" appears as a real standalone word, but nothing that looks
|
||||||
|
// like "<start-id> <count>" follows it.
|
||||||
|
let buf = b"Please refer to the xref appendix for details.";
|
||||||
|
assert_eq!(find_last_valid_xref_table_start(buf), None);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn find_xref_accepts_real_classic_table_header() {
|
||||||
|
let buf = b"garbage\nxref\n0 6\n0000000000 65535 f \n%%EOF";
|
||||||
|
let pos = find_last_valid_xref_table_start(buf).expect("should find the real table");
|
||||||
|
assert_eq!(&buf[pos..pos + 4], b"xref");
|
||||||
|
assert_eq!(&buf[pos..], b"xref\n0 6\n0000000000 65535 f \n%%EOF");
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn find_xref_skips_coincidental_match_and_finds_real_table_before_it() {
|
||||||
|
// A coincidental "xref" (no subsection header) appears *after* the
|
||||||
|
// real table in the buffer — the scan must not stop at the first
|
||||||
|
// (rightmost) standalone token it finds; it must keep looking
|
||||||
|
// backward until one actually validates.
|
||||||
|
let buf = b"xref\n0 3\n0000000000 65535 f \ntrailer\nsee the xref\n";
|
||||||
|
let pos = find_last_valid_xref_table_start(buf).expect("should find the real table");
|
||||||
|
assert_eq!(pos, 0);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn find_xref_rejects_substring_of_startxref() {
|
||||||
|
// "xref" is a substring of "startxref" but isn't a standalone
|
||||||
|
// token there (not preceded by whitespace) — must not match, even
|
||||||
|
// though a number immediately follows it.
|
||||||
|
let buf = b"startxref\n1234\n%%EOF";
|
||||||
|
assert_eq!(find_last_valid_xref_table_start(buf), None);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn find_xref_rejects_count_run_with_trailing_garbage() {
|
||||||
|
// "xref\n0 6garbage" has the right shape (digits, whitespace,
|
||||||
|
// digits) but the count run doesn't end at whitespace/EOF — it
|
||||||
|
// runs straight into non-digit garbage, so this must not be
|
||||||
|
// accepted as a real subsection header.
|
||||||
|
let buf = b"xref\n0 6garbage\n%%EOF";
|
||||||
|
assert_eq!(find_last_valid_xref_table_start(buf), None);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn recover_startxref_pointer_returns_none_without_a_valid_table() {
|
||||||
|
let buf = b"Please refer to the xref appendix for details.";
|
||||||
|
assert!(recover_startxref_pointer(buf).is_none());
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -171,6 +171,74 @@ pub(crate) fn is_toc_marker_heading(text: &str) -> bool {
|
|||||||
/// equation and absent from name-plus-number headings. A bare trailing colon
|
/// equation and absent from name-plus-number headings. A bare trailing colon
|
||||||
/// is NOT a fragment signal either: real headings frequently end with colons
|
/// is NOT a fragment signal either: real headings frequently end with colons
|
||||||
/// ("Procedure:", "Steps for Using the Microscope:").
|
/// ("Procedure:", "Steps for Using the Microscope:").
|
||||||
|
/// True when the line opens with a section number ("3.", "2.1.4", "IV)").
|
||||||
|
///
|
||||||
|
/// Mirrors the acceptance of `heading::parse_numbering` rather than the
|
||||||
|
/// stricter `convert::starts_with_section_number`, which deliberately
|
||||||
|
/// requires two components because it bypasses isolation checks. Here a
|
||||||
|
/// single "1." counts: numbering is independent evidence of a heading, and
|
||||||
|
/// `heading.rs` applies its numbered-prefix allowance *after* consulting
|
||||||
|
/// `is_heading_fragment`, so without this exemption a numbered
|
||||||
|
/// sentence-case heading would be vetoed before that allowance can run.
|
||||||
|
fn starts_with_numbering_prefix(t: &str) -> bool {
|
||||||
|
let Some(first) = t.split_whitespace().next() else {
|
||||||
|
return false;
|
||||||
|
};
|
||||||
|
let has_delimiter = first.ends_with(['.', ')', ':']);
|
||||||
|
let token = first.trim_end_matches(['.', ')', ':']);
|
||||||
|
if token.is_empty() {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
let parts: Vec<&str> = token.split('.').collect();
|
||||||
|
let decimal = parts
|
||||||
|
.iter()
|
||||||
|
.all(|p| !p.is_empty() && p.len() <= 3 && p.chars().all(|c| c.is_ascii_digit()));
|
||||||
|
if decimal {
|
||||||
|
// "1." / "2.1." carry a delimiter; "2.3 Title" is written without
|
||||||
|
// one, so a multi-component number is accepted bare. A bare single
|
||||||
|
// number ("3 apples") is not — that is ordinary prose.
|
||||||
|
return has_delimiter || parts.len() >= 2;
|
||||||
|
}
|
||||||
|
// Roman numerals go through the heading parser's own grammar so the two
|
||||||
|
// agree: uppercase I/V/X/L/C only, at most 8 characters. A looser rule
|
||||||
|
// here would exempt markers the parser rejects — "iv)" or "d)" from an
|
||||||
|
// alphabetical list — letting an ordinary list item bypass the veto and
|
||||||
|
// reach heading promotion.
|
||||||
|
//
|
||||||
|
// A delimiter is also required: a bare leading "I" is the pronoun far
|
||||||
|
// more often than a section number.
|
||||||
|
has_delimiter && crate::markdown::heading::roman_value(token).is_some()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// True when the line reads as a title rather than a sentence: every
|
||||||
|
/// content word (ignoring minor words) starts uppercase. Used to spare real
|
||||||
|
/// headings from the dangling-verb veto — "Bond Yields" is a section title,
|
||||||
|
/// "the method yields" is a stranded clause, and only the casing tells them
|
||||||
|
/// apart.
|
||||||
|
fn looks_title_case(t: &str) -> bool {
|
||||||
|
const MINOR: &[&str] = &[
|
||||||
|
"a", "an", "the", "of", "and", "or", "for", "to", "in", "on", "at", "by", "with", "from",
|
||||||
|
"as", "is", "are", "that", "than", "into",
|
||||||
|
];
|
||||||
|
let mut content = 0usize;
|
||||||
|
let mut capitalized = 0usize;
|
||||||
|
for w in t.split_whitespace() {
|
||||||
|
let cleaned: String = w.chars().filter(|c| c.is_alphabetic()).collect();
|
||||||
|
if cleaned.is_empty() {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if MINOR.contains(&cleaned.to_lowercase().as_str()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
content += 1;
|
||||||
|
if cleaned.chars().next().is_some_and(char::is_uppercase) {
|
||||||
|
capitalized += 1;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// A single content word ("Yields") is a title by default.
|
||||||
|
content == 0 || capitalized == content
|
||||||
|
}
|
||||||
|
|
||||||
pub(crate) fn is_heading_fragment(text: &str) -> bool {
|
pub(crate) fn is_heading_fragment(text: &str) -> bool {
|
||||||
let t = text.trim_end();
|
let t = text.trim_end();
|
||||||
|
|
||||||
@@ -244,9 +312,133 @@ pub(crate) fn is_heading_fragment(text: &str) -> bool {
|
|||||||
if t.ends_with(':') && t.split_whitespace().any(is_equation_number) {
|
if t.ends_with(':') && t.split_whitespace().any(is_equation_number) {
|
||||||
return true;
|
return true;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Dangling clause: a stranded sentence lead-in ends on a relational
|
||||||
|
// verb with no terminal punctuation — "Note that the exact error equals"
|
||||||
|
// left ahead of its formula when a phantom table dissolved.
|
||||||
|
//
|
||||||
|
// Gated on the line reading as prose rather than a title. Case is the
|
||||||
|
// discriminator the trailing word alone cannot provide: a heading is
|
||||||
|
// title case ("Bond Yields", "The Method Yields") while a stranded
|
||||||
|
// lead-in is sentence case ("the method yields"). Without this gate the
|
||||||
|
// veto eats real headings — "Bond Yields", "Crop Yields" and any wrapped
|
||||||
|
// title-case heading the preprocessor failed to merge.
|
||||||
|
if !t.ends_with(['.', '!', '?', ':', ';', ')', ']'])
|
||||||
|
&& !looks_title_case(t)
|
||||||
|
&& !starts_with_numbering_prefix(t)
|
||||||
|
{
|
||||||
|
if let Some(last) = t.split_whitespace().next_back() {
|
||||||
|
let word: String = last
|
||||||
|
.trim_matches(|c: char| !c.is_alphanumeric())
|
||||||
|
.to_lowercase();
|
||||||
|
// Relational verbs only, and only those with no common noun
|
||||||
|
// sense. "yields" was dropped for exactly that reason: "Bond
|
||||||
|
// Yields" is a real section title. Function words, copulas and
|
||||||
|
// auxiliaries were measured and rejected outright — a heading
|
||||||
|
// that wraps across lines ends on those, and suppressing them
|
||||||
|
// destroyed real IRS Publication 17 headings.
|
||||||
|
const DANGLING_TAIL: &[&str] =
|
||||||
|
&["equals", "denotes", "implies", "satisfies", "signifies"];
|
||||||
|
if DANGLING_TAIL.contains(&word.as_str()) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
false
|
false
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[cfg(test)]
|
||||||
|
mod fragment_heading_tests {
|
||||||
|
use super::is_heading_fragment;
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn dangling_tail_marks_stranded_clause() {
|
||||||
|
// opendataloader 01030000000144: left behind when a phantom table
|
||||||
|
// dissolved, ahead of its formula on the next line.
|
||||||
|
assert!(is_heading_fragment("Note that the exact error equals"));
|
||||||
|
assert!(is_heading_fragment("The remainder term satisfies"));
|
||||||
|
assert!(is_heading_fragment("we conclude that the sum equals"));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn real_headings_survive() {
|
||||||
|
assert!(!is_heading_fragment("Introduction"));
|
||||||
|
assert!(!is_heading_fragment("Error Analysis"));
|
||||||
|
assert!(!is_heading_fragment("Materials and Methods"));
|
||||||
|
assert!(!is_heading_fragment("Results"));
|
||||||
|
assert!(!is_heading_fragment("3.2 Richardson Extrapolation"));
|
||||||
|
assert!(!is_heading_fragment("Discussion and Conclusions"));
|
||||||
|
// Terminal punctuation means the clause is complete.
|
||||||
|
assert!(!is_heading_fragment("What is a Derivative?"));
|
||||||
|
assert!(!is_heading_fragment("Procedure:"));
|
||||||
|
assert!(!is_heading_fragment("Note that this is important."));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn title_case_headings_ending_in_a_verb_survive() {
|
||||||
|
// "yields" is also a plural noun; these are real section titles.
|
||||||
|
assert!(!is_heading_fragment("Bond Yields"));
|
||||||
|
assert!(!is_heading_fragment("Crop Yields"));
|
||||||
|
assert!(!is_heading_fragment("Dividend Yields"));
|
||||||
|
assert!(!is_heading_fragment("Yields"));
|
||||||
|
// A wrapped title-case heading whose first line ends on a listed
|
||||||
|
// verb must survive even if the preprocessor failed to merge it.
|
||||||
|
assert!(!is_heading_fragment("The Theorem Implies"));
|
||||||
|
assert!(!is_heading_fragment("What This Denotes"));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn numbered_sentence_case_headings_survive() {
|
||||||
|
// heading.rs consults is_heading_fragment BEFORE applying its
|
||||||
|
// numbered-prefix allowance, so the veto must not pre-empt it.
|
||||||
|
assert!(!is_heading_fragment("1. What the model implies"));
|
||||||
|
assert!(!is_heading_fragment("2.3 How the estimator satisfies"));
|
||||||
|
assert!(!is_heading_fragment("IV) What this denotes"));
|
||||||
|
// Without numbering the same wording is still a stranded clause.
|
||||||
|
assert!(is_heading_fragment("What the model implies"));
|
||||||
|
// A bare leading number or pronoun is prose, not numbering.
|
||||||
|
assert!(is_heading_fragment("3 apples and what that implies"));
|
||||||
|
assert!(is_heading_fragment("I think the model implies"));
|
||||||
|
// Markers heading::parse_numbering rejects must not be exempted
|
||||||
|
// either, or an ordinary list item bypasses the veto: lowercase
|
||||||
|
// roman, alphabetical markers, and over-long tokens.
|
||||||
|
assert!(is_heading_fragment("iv) the estimator satisfies"));
|
||||||
|
assert!(is_heading_fragment("d) the value implies"));
|
||||||
|
// Unsupported character (M is outside the parser's I/V/X/L/C set).
|
||||||
|
assert!(is_heading_fragment("MMMM. the value implies"));
|
||||||
|
// Over-long token: nine valid characters, so this exercises the
|
||||||
|
// 8-character bound rather than the character set.
|
||||||
|
assert!(is_heading_fragment("IIIIIIIII. the value implies"));
|
||||||
|
// Eight is still within the bound and stays exempt.
|
||||||
|
assert!(!is_heading_fragment("IIIIIIII. What this implies"));
|
||||||
|
// Uppercase roman within the parser's grammar is still exempt.
|
||||||
|
assert!(!is_heading_fragment("IV. What this denotes"));
|
||||||
|
assert!(!is_heading_fragment("XII) What this implies"));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn wrapped_headings_are_not_fragments() {
|
||||||
|
// A heading that wraps across lines ends on a function word. These
|
||||||
|
// are real headings from IRS Publication 17 and must survive.
|
||||||
|
assert!(!is_heading_fragment("Casualty and"));
|
||||||
|
assert!(!is_heading_fragment("Rule 10. You Must Be at"));
|
||||||
|
assert!(!is_heading_fragment("Higher Standard Deduction for"));
|
||||||
|
assert!(!is_heading_fragment("Qualifying Child of"));
|
||||||
|
assert!(!is_heading_fragment("When Can I Withdraw or"));
|
||||||
|
// Copulas and auxiliaries also end real wrapped headings.
|
||||||
|
assert!(!is_heading_fragment("Rule 15. Your AGI Must Be"));
|
||||||
|
assert!(!is_heading_fragment("What Medical Expenses Are"));
|
||||||
|
assert!(!is_heading_fragment("Rule 13. You Must Have"));
|
||||||
|
assert!(!is_heading_fragment("When Can a Roth IRA Be"));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn dangling_check_is_case_insensitive() {
|
||||||
|
// All-caps is not sentence case, so the veto must not fire there.
|
||||||
|
assert!(!is_heading_fragment("THE REMAINDER EQUALS"));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Compute the Y-gap threshold for paragraph break detection.
|
/// Compute the Y-gap threshold for paragraph break detection.
|
||||||
///
|
///
|
||||||
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
|
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
|
||||||
|
|||||||
@@ -127,7 +127,9 @@ fn visual_style(line: &TextLine) -> Option<VisualStyle> {
|
|||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
fn roman_value(token: &str) -> Option<u32> {
|
/// Shared with `analysis::starts_with_numbering_prefix` so the veto
|
||||||
|
/// exemption and the heading parser agree on what a roman numeral is.
|
||||||
|
pub(super) fn roman_value(token: &str) -> Option<u32> {
|
||||||
if token.is_empty() || token.len() > 8 {
|
if token.is_empty() || token.len() > 8 {
|
||||||
return None;
|
return None;
|
||||||
}
|
}
|
||||||
|
|||||||
+85
-12
@@ -69,7 +69,7 @@ fn is_chart_adjacent_label(item: &TextItem, region: (f32, f32, f32, f32)) -> boo
|
|||||||
|| (mostly_inside_chart_width && close_to_chart_edge && category_sized))
|
|| (mostly_inside_chart_width && close_to_chart_edge && category_sized))
|
||||||
}
|
}
|
||||||
|
|
||||||
fn item_is_in_chart_region(item: &TextItem, regions: &[(f32, f32, f32, f32)]) -> bool {
|
pub(crate) fn item_is_in_chart_region(item: &TextItem, regions: &[(f32, f32, f32, f32)]) -> bool {
|
||||||
regions.iter().any(|&(x0, y0, x1, y1)| {
|
regions.iter().any(|&(x0, y0, x1, y1)| {
|
||||||
let cx = item.x + item.width / 2.0;
|
let cx = item.x + item.width / 2.0;
|
||||||
let within_padded_x = cx >= x0 - CHART_REGION_PAD && cx <= x1 + CHART_REGION_PAD;
|
let within_padded_x = cx >= x0 - CHART_REGION_PAD && cx <= x1 + CHART_REGION_PAD;
|
||||||
@@ -92,6 +92,72 @@ fn items_outside_chart_regions(
|
|||||||
.collect()
|
.collect()
|
||||||
}
|
}
|
||||||
|
|
||||||
|
pub(crate) fn merge_chart_regions(
|
||||||
|
regions: impl IntoIterator<Item = (f32, f32, f32, f32)>,
|
||||||
|
) -> Vec<(f32, f32, f32, f32)> {
|
||||||
|
const MERGE_TOLERANCE: f32 = 3.0;
|
||||||
|
|
||||||
|
let mut merged: Vec<(f32, f32, f32, f32)> = Vec::new();
|
||||||
|
for (x0, y0, x1, y1) in regions {
|
||||||
|
let mut current = (x0.min(x1), y0.min(y1), x0.max(x1), y0.max(y1));
|
||||||
|
let mut index = 0;
|
||||||
|
while index < merged.len() {
|
||||||
|
let candidate = merged[index];
|
||||||
|
let overlaps = current.2 + MERGE_TOLERANCE >= candidate.0
|
||||||
|
&& candidate.2 + MERGE_TOLERANCE >= current.0
|
||||||
|
&& current.3 + MERGE_TOLERANCE >= candidate.1
|
||||||
|
&& candidate.3 + MERGE_TOLERANCE >= current.1;
|
||||||
|
if overlaps {
|
||||||
|
current = (
|
||||||
|
current.0.min(candidate.0),
|
||||||
|
current.1.min(candidate.1),
|
||||||
|
current.2.max(candidate.2),
|
||||||
|
current.3.max(candidate.3),
|
||||||
|
);
|
||||||
|
merged.swap_remove(index);
|
||||||
|
} else {
|
||||||
|
index += 1;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
merged.push(current);
|
||||||
|
}
|
||||||
|
merged
|
||||||
|
}
|
||||||
|
|
||||||
|
pub(crate) type PageChartRegions = HashMap<u32, Vec<(f32, f32, f32, f32)>>;
|
||||||
|
|
||||||
|
/// Compute the chart masks used by both layout analysis and Markdown output.
|
||||||
|
///
|
||||||
|
/// Keeping the rect-backed and dense-line heuristics behind one entry point
|
||||||
|
/// ensures metadata and extraction cannot drift when either detector changes.
|
||||||
|
pub(crate) fn chart_regions_by_page(
|
||||||
|
items: &[TextItem],
|
||||||
|
rects: &[PdfRect],
|
||||||
|
lines: &[PdfLine],
|
||||||
|
) -> PageChartRegions {
|
||||||
|
let mut page_items: HashMap<u32, Vec<TextItem>> = HashMap::new();
|
||||||
|
for item in items.iter().filter(|item| {
|
||||||
|
matches!(
|
||||||
|
&item.item_type,
|
||||||
|
crate::types::ItemType::Text | crate::types::ItemType::FormField
|
||||||
|
)
|
||||||
|
}) {
|
||||||
|
page_items.entry(item.page).or_default().push(item.clone());
|
||||||
|
}
|
||||||
|
|
||||||
|
page_items
|
||||||
|
.into_iter()
|
||||||
|
.filter_map(|(page, items)| {
|
||||||
|
let rect_regions = crate::tables::detect_chart_regions(&items, rects, page);
|
||||||
|
let line_regions = crate::tables::detect_dense_line_chart_regions(lines, rects, page)
|
||||||
|
.into_iter()
|
||||||
|
.filter(|®ion| chart_region_separates_prose_columns(&items, region));
|
||||||
|
let regions = merge_chart_regions(rect_regions.into_iter().chain(line_regions));
|
||||||
|
(!regions.is_empty()).then_some((page, regions))
|
||||||
|
})
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
/// Detect side-by-side table layout by finding a significant X-position gap.
|
/// Detect side-by-side table layout by finding a significant X-position gap.
|
||||||
///
|
///
|
||||||
/// Returns X-band boundaries `[(x_min, split_x), (split_x, x_max)]` when a
|
/// Returns X-band boundaries `[(x_min, split_x), (split_x, x_max)]` when a
|
||||||
@@ -329,6 +395,15 @@ fn chart_spans_prose_split(region: (f32, f32, f32, f32), split_x: f32) -> bool {
|
|||||||
split_x - left >= MIN_CHART_WIDTH_PER_SIDE && right - split_x >= MIN_CHART_WIDTH_PER_SIDE
|
split_x - left >= MIN_CHART_WIDTH_PER_SIDE && right - split_x >= MIN_CHART_WIDTH_PER_SIDE
|
||||||
}
|
}
|
||||||
|
|
||||||
|
pub(crate) fn chart_region_separates_prose_columns(
|
||||||
|
items: &[TextItem],
|
||||||
|
region: (f32, f32, f32, f32),
|
||||||
|
) -> bool {
|
||||||
|
let outside = items_outside_chart_regions(items, &[region]);
|
||||||
|
chart_page_prose_column_split(&outside)
|
||||||
|
.is_some_and(|split_x| chart_spans_prose_split(region, split_x))
|
||||||
|
}
|
||||||
|
|
||||||
/// True when adjacent physical rows form an unterminated, lowercase prose
|
/// True when adjacent physical rows form an unterminated, lowercase prose
|
||||||
/// continuation in the same projected column.
|
/// continuation in the same projected column.
|
||||||
fn is_cross_row_prose_continuation(previous: &str, current: &str) -> bool {
|
fn is_cross_row_prose_continuation(previous: &str, current: &str) -> bool {
|
||||||
@@ -1004,6 +1079,7 @@ pub fn to_markdown_from_items_with_rects_and_page_count(
|
|||||||
page_count: document_page_count,
|
page_count: document_page_count,
|
||||||
prefiltered_page_number_pages: None,
|
prefiltered_page_number_pages: None,
|
||||||
prefiltered_page_number_mask: None,
|
prefiltered_page_number_mask: None,
|
||||||
|
precomputed_chart_regions: None,
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
@@ -1021,6 +1097,9 @@ pub(crate) struct MarkdownDocumentContext<'a> {
|
|||||||
/// Table detection consumes the original items; the mask is applied only
|
/// Table detection consumes the original items; the mask is applied only
|
||||||
/// after table claims have been established.
|
/// after table claims have been established.
|
||||||
pub(crate) prefiltered_page_number_mask: Option<&'a [bool]>,
|
pub(crate) prefiltered_page_number_mask: Option<&'a [bool]>,
|
||||||
|
/// Optional chart masks shared with layout analysis so the geometry is
|
||||||
|
/// detected once and interpreted identically by both pipelines.
|
||||||
|
pub(crate) precomputed_chart_regions: Option<&'a PageChartRegions>,
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Convert positioned text items to markdown, using rectangles and line segments for table detection.
|
/// Convert positioned text items to markdown, using rectangles and line segments for table detection.
|
||||||
@@ -1047,6 +1126,7 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
|
|||||||
page_count: document_page_count,
|
page_count: document_page_count,
|
||||||
prefiltered_page_number_pages,
|
prefiltered_page_number_pages,
|
||||||
prefiltered_page_number_mask,
|
prefiltered_page_number_mask,
|
||||||
|
precomputed_chart_regions,
|
||||||
} = context;
|
} = context;
|
||||||
|
|
||||||
if items.is_empty() {
|
if items.is_empty() {
|
||||||
@@ -1119,17 +1199,9 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
|
|||||||
|
|
||||||
// Chart regions per page: their text must not steer column detection
|
// Chart regions per page: their text must not steer column detection
|
||||||
// during line grouping (it fills the gutter and fuses two-column lines).
|
// during line grouping (it fills the gutter and fuses two-column lines).
|
||||||
let mut page_chart_map: HashMap<u32, Vec<(f32, f32, f32, f32)>> = HashMap::new();
|
let page_chart_map = precomputed_chart_regions
|
||||||
for &page in page_groups.keys() {
|
.cloned()
|
||||||
let page_items_ref: Vec<TextItem> = page_groups[&page]
|
.unwrap_or_else(|| chart_regions_by_page(&text_items, rects, pdf_lines));
|
||||||
.iter()
|
|
||||||
.map(|(_, item)| (*item).clone())
|
|
||||||
.collect();
|
|
||||||
let regions = crate::tables::detect_chart_regions(&page_items_ref, rects, page);
|
|
||||||
if !regions.is_empty() {
|
|
||||||
page_chart_map.insert(page, regions);
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
let mut pages: Vec<u32> = page_groups.keys().copied().collect();
|
let mut pages: Vec<u32> = page_groups.keys().copied().collect();
|
||||||
pages.sort();
|
pages.sort();
|
||||||
@@ -2051,6 +2123,7 @@ mod tests {
|
|||||||
page_count: 1,
|
page_count: 1,
|
||||||
prefiltered_page_number_pages: Some(&removed_pages),
|
prefiltered_page_number_pages: Some(&removed_pages),
|
||||||
prefiltered_page_number_mask: Some(&removal_mask),
|
prefiltered_page_number_mask: Some(&removal_mask),
|
||||||
|
precomputed_chart_regions: None,
|
||||||
},
|
},
|
||||||
);
|
);
|
||||||
|
|
||||||
|
|||||||
@@ -464,6 +464,7 @@ mod tests {
|
|||||||
assert!(!is_page_number_line("Hello World"));
|
assert!(!is_page_number_line("Hello World"));
|
||||||
assert!(!is_page_number_line("Chapter 1"));
|
assert!(!is_page_number_line("Chapter 1"));
|
||||||
assert!(!is_page_number_line("Total: 500"));
|
assert!(!is_page_number_line("Total: 500"));
|
||||||
|
assert!(!is_page_number_line("PAGE0-PARA2-END-MARKER-0"));
|
||||||
}
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
@@ -507,6 +508,14 @@ mod tests {
|
|||||||
assert!(result.contains("End"));
|
assert!(result.contains("End"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_remove_page_numbers_preserves_page_prefixed_content() {
|
||||||
|
let input = "PAGE0-PARA2-START substantive report text PAGE0-PARA2-END-MARKER-0";
|
||||||
|
let result = remove_page_numbers(input);
|
||||||
|
|
||||||
|
assert_eq!(result, input);
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_remove_page_numbers_multiple_patterns() {
|
fn test_remove_page_numbers_multiple_patterns() {
|
||||||
let input = "\n1\n\nContent\n\n2\n\n---\nMore\n\n3\n";
|
let input = "\n1\n\nContent\n\n2\n\n---\nMore\n\n3\n";
|
||||||
|
|||||||
@@ -271,6 +271,12 @@ pub struct PyTextItem {
|
|||||||
pub is_strikeout: bool,
|
pub is_strikeout: bool,
|
||||||
#[pyo3(get)]
|
#[pyo3(get)]
|
||||||
pub item_type: String,
|
pub item_type: String,
|
||||||
|
/// Marked Content ID from the content stream's BDC/BMC operator, None
|
||||||
|
/// when the text is not part of marked content. Join with the
|
||||||
|
/// (page, mcid) pairs from extract_structure_elements to attach
|
||||||
|
/// structure-tree roles (headings, paragraphs, ...) in tagged PDFs.
|
||||||
|
#[pyo3(get)]
|
||||||
|
pub mcid: Option<i64>,
|
||||||
}
|
}
|
||||||
|
|
||||||
#[pymethods]
|
#[pymethods]
|
||||||
@@ -286,6 +292,32 @@ impl PyTextItem {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// One structure-tree element reference from a tagged PDF.
|
||||||
|
#[pyclass(name = "StructureElement")]
|
||||||
|
#[derive(Clone)]
|
||||||
|
pub struct PyStructureElement {
|
||||||
|
/// 1-indexed page number (matches TextItem.page).
|
||||||
|
#[pyo3(get)]
|
||||||
|
pub page: u32,
|
||||||
|
/// Marked Content ID from the page's content stream (matches
|
||||||
|
/// TextItem.mcid).
|
||||||
|
#[pyo3(get)]
|
||||||
|
pub mcid: i64,
|
||||||
|
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", ...).
|
||||||
|
#[pyo3(get)]
|
||||||
|
pub role: String,
|
||||||
|
}
|
||||||
|
|
||||||
|
#[pymethods]
|
||||||
|
impl PyStructureElement {
|
||||||
|
fn __repr__(&self) -> String {
|
||||||
|
format!(
|
||||||
|
"StructureElement(page={}, mcid={}, role='{}')",
|
||||||
|
self.page, self.mcid, self.role
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// ---------------------------------------------------------------------------
|
// ---------------------------------------------------------------------------
|
||||||
// Helpers
|
// Helpers
|
||||||
// ---------------------------------------------------------------------------
|
// ---------------------------------------------------------------------------
|
||||||
@@ -356,6 +388,18 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
|
|||||||
is_underline: item.is_underline,
|
is_underline: item.is_underline,
|
||||||
is_strikeout: item.is_strikeout,
|
is_strikeout: item.is_strikeout,
|
||||||
item_type: item_type_str(&item.item_type),
|
item_type: item_type_str(&item.item_type),
|
||||||
|
mcid: item.mcid,
|
||||||
|
})
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
fn convert_structure_elements(elements: Vec<crate::StructureElement>) -> Vec<PyStructureElement> {
|
||||||
|
elements
|
||||||
|
.into_iter()
|
||||||
|
.map(|e| PyStructureElement {
|
||||||
|
page: e.page,
|
||||||
|
mcid: e.mcid,
|
||||||
|
role: e.role,
|
||||||
})
|
})
|
||||||
.collect()
|
.collect()
|
||||||
}
|
}
|
||||||
@@ -613,6 +657,48 @@ fn extract_pages_markdown_bytes(
|
|||||||
Ok(to_py_pages_result(result))
|
Ok(to_py_pages_result(result))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Extract structure-tree element references from a tagged PDF file.
|
||||||
|
///
|
||||||
|
/// Parses the document's structure tree (when present) and returns one
|
||||||
|
/// entry per marked-content reference, resolved to its 1-indexed page,
|
||||||
|
/// MCID, and structure type name ("H1".."H6", "P", "Table", ...). Returns
|
||||||
|
/// an empty list when the PDF is not tagged.
|
||||||
|
///
|
||||||
|
/// Join (page, mcid) against the page/mcid attributes from
|
||||||
|
/// [`extract_text_with_positions`] to attach heading levels and other
|
||||||
|
/// semantic roles to extracted text.
|
||||||
|
///
|
||||||
|
/// Args:
|
||||||
|
/// path: Path to the PDF file.
|
||||||
|
/// pages: Optional list of 1-indexed pages (matching TextItem.page).
|
||||||
|
/// When None (default), the whole document is returned.
|
||||||
|
///
|
||||||
|
/// Returns:
|
||||||
|
/// List of StructureElement sorted by (page, mcid).
|
||||||
|
#[pyfunction]
|
||||||
|
#[pyo3(signature = (path, pages=None))]
|
||||||
|
fn extract_structure_elements(
|
||||||
|
path: &str,
|
||||||
|
pages: Option<Vec<u32>>,
|
||||||
|
) -> PyResult<Vec<PyStructureElement>> {
|
||||||
|
let elements = crate::extract_structure_elements(path, pages.as_deref()).map_err(to_py_err)?;
|
||||||
|
Ok(convert_structure_elements(elements))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Extract structure-tree element references from tagged PDF bytes.
|
||||||
|
///
|
||||||
|
/// See [`extract_structure_elements`] for details.
|
||||||
|
#[pyfunction]
|
||||||
|
#[pyo3(signature = (data, pages=None))]
|
||||||
|
fn extract_structure_elements_bytes(
|
||||||
|
data: &[u8],
|
||||||
|
pages: Option<Vec<u32>>,
|
||||||
|
) -> PyResult<Vec<PyStructureElement>> {
|
||||||
|
let elements =
|
||||||
|
crate::extract_structure_elements_mem(data, pages.as_deref()).map_err(to_py_err)?;
|
||||||
|
Ok(convert_structure_elements(elements))
|
||||||
|
}
|
||||||
|
|
||||||
/// Python module definition.
|
/// Python module definition.
|
||||||
#[pymodule]
|
#[pymodule]
|
||||||
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
||||||
@@ -620,6 +706,7 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
|||||||
m.add_class::<PyPageOcrReasons>()?;
|
m.add_class::<PyPageOcrReasons>()?;
|
||||||
m.add_class::<PyPdfClassification>()?;
|
m.add_class::<PyPdfClassification>()?;
|
||||||
m.add_class::<PyTextItem>()?;
|
m.add_class::<PyTextItem>()?;
|
||||||
|
m.add_class::<PyStructureElement>()?;
|
||||||
m.add_class::<PyRegionText>()?;
|
m.add_class::<PyRegionText>()?;
|
||||||
m.add_class::<PyPageRegionTexts>()?;
|
m.add_class::<PyPageRegionTexts>()?;
|
||||||
m.add_class::<PyPageMarkdown>()?;
|
m.add_class::<PyPageMarkdown>()?;
|
||||||
@@ -634,6 +721,8 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
|||||||
m.add_function(wrap_pyfunction!(extract_text_bytes, m)?)?;
|
m.add_function(wrap_pyfunction!(extract_text_bytes, m)?)?;
|
||||||
m.add_function(wrap_pyfunction!(extract_text_with_positions, m)?)?;
|
m.add_function(wrap_pyfunction!(extract_text_with_positions, m)?)?;
|
||||||
m.add_function(wrap_pyfunction!(extract_text_with_positions_bytes, m)?)?;
|
m.add_function(wrap_pyfunction!(extract_text_with_positions_bytes, m)?)?;
|
||||||
|
m.add_function(wrap_pyfunction!(extract_structure_elements, m)?)?;
|
||||||
|
m.add_function(wrap_pyfunction!(extract_structure_elements_bytes, m)?)?;
|
||||||
m.add_function(wrap_pyfunction!(extract_text_in_regions, m)?)?;
|
m.add_function(wrap_pyfunction!(extract_text_in_regions, m)?)?;
|
||||||
m.add_function(wrap_pyfunction!(extract_text_in_regions_bytes, m)?)?;
|
m.add_function(wrap_pyfunction!(extract_text_in_regions_bytes, m)?)?;
|
||||||
m.add_function(wrap_pyfunction!(extract_pages_markdown, m)?)?;
|
m.add_function(wrap_pyfunction!(extract_pages_markdown, m)?)?;
|
||||||
|
|||||||
+819
-38
File diff suppressed because it is too large
Load Diff
+308
-10
@@ -435,6 +435,72 @@ fn revised_table_cell_indices(
|
|||||||
.collect()
|
.collect()
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Index of candidate "body" items (larger-font attachment targets) sorted by
|
||||||
|
/// Y, so script-attachment checks scan a narrow Y window instead of the whole
|
||||||
|
/// page per candidate.
|
||||||
|
struct ScriptBodyIndex<'a> {
|
||||||
|
/// (y, item), sorted ascending by y
|
||||||
|
by_y: Vec<(f32, &'a TextItem)>,
|
||||||
|
/// widest vertical attachment window any body item can produce
|
||||||
|
max_window: f32,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl<'a> ScriptBodyIndex<'a> {
|
||||||
|
fn new(items: &'a [TextItem]) -> Self {
|
||||||
|
// Smallest table-candidate font is 6pt, so any possible attachment
|
||||||
|
// target is at least 6 x 1.2 pt.
|
||||||
|
let mut by_y: Vec<(f32, &TextItem)> = items
|
||||||
|
.iter()
|
||||||
|
.filter(|i| i.font_size >= 6.0 * 1.2)
|
||||||
|
.map(|i| (i.y, i))
|
||||||
|
.collect();
|
||||||
|
by_y.sort_by(|a, b| a.0.total_cmp(&b.0));
|
||||||
|
let max_window = by_y
|
||||||
|
.iter()
|
||||||
|
.map(|(_, i)| i.font_size * 0.8)
|
||||||
|
.fold(0.0f32, f32::max);
|
||||||
|
Self { by_y, max_window }
|
||||||
|
}
|
||||||
|
|
||||||
|
/// True when a small-font item is horizontally attached to a larger-font
|
||||||
|
/// item at a script baseline offset — a sub/superscript in running text
|
||||||
|
/// or math (equation subscripts, footnote markers). Script attachments
|
||||||
|
/// are not table cells; without this filter, display equations with
|
||||||
|
/// sub/superscripts form phantom small-font table regions (e.g. TeX
|
||||||
|
/// papers where log subscripts cluster with footnote lines into a fake
|
||||||
|
/// 3-column table). A genuine baseline offset is required so same-line
|
||||||
|
/// table neighbours (a small cell beside a larger label cell) are never
|
||||||
|
/// classified as scripts.
|
||||||
|
///
|
||||||
|
/// `min_anchor_size` additionally constrains what counts as an
|
||||||
|
/// attachment target: the small-font pass accepts any sufficiently
|
||||||
|
/// larger item (0.0), while the body-font pass requires a heading-sized
|
||||||
|
/// anchor so a body-size table cell beside a slightly larger label with
|
||||||
|
/// baseline jitter is never treated as a script.
|
||||||
|
fn is_script_attachment(&self, small: &TextItem, min_anchor_size: f32) -> bool {
|
||||||
|
let attach_gap = small.font_size.max(4.0) * 0.6;
|
||||||
|
let lo = self
|
||||||
|
.by_y
|
||||||
|
.partition_point(|(y, _)| *y < small.y - self.max_window);
|
||||||
|
self.by_y[lo..]
|
||||||
|
.iter()
|
||||||
|
.take_while(|(y, _)| *y <= small.y + self.max_window)
|
||||||
|
.any(|(_, body)| {
|
||||||
|
let dy = (small.y - body.y).abs();
|
||||||
|
body.font_size >= small.font_size * 1.2
|
||||||
|
&& body.font_size >= min_anchor_size
|
||||||
|
&& dy > body.font_size * 0.05
|
||||||
|
&& dy <= body.font_size * 0.8
|
||||||
|
&& {
|
||||||
|
let gap_after_body = small.x - (body.x + body.width);
|
||||||
|
let gap_before_body = body.x - (small.x + small.width);
|
||||||
|
(-attach_gap..=attach_gap).contains(&gap_after_body)
|
||||||
|
|| (-attach_gap..=attach_gap).contains(&gap_before_body)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Detect tables in a set of text items from a single page
|
/// Detect tables in a set of text items from a single page
|
||||||
pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bool) -> Vec<Table> {
|
pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bool) -> Vec<Table> {
|
||||||
detect_tables_with_page_width(items, base_font_size, skip_body_font, content_width(items))
|
detect_tables_with_page_width(items, base_font_size, skip_body_font, content_width(items))
|
||||||
@@ -483,6 +549,27 @@ pub(crate) fn detect_tables_with_page_width(
|
|||||||
// === Pass 1: Small-font tables (existing behavior) ===
|
// === Pass 1: Small-font tables (existing behavior) ===
|
||||||
let table_font_threshold = base_font_size * 0.90;
|
let table_font_threshold = base_font_size * 0.90;
|
||||||
|
|
||||||
|
// Mark sub/superscript attachments once per pass. They stay candidates —
|
||||||
|
// the masks only remove them from region qualification and column/row
|
||||||
|
// geometry.
|
||||||
|
//
|
||||||
|
// The two passes need different anchor thresholds. In the small-font pass
|
||||||
|
// any sufficiently larger neighbour is a plausible base for a script. In
|
||||||
|
// the body-font pass the candidates are themselves body-sized
|
||||||
|
// (0.85..1.05x), so a merely "slightly larger" neighbour is usually a bold
|
||||||
|
// label or an adjacent column header, not the base of a superscript —
|
||||||
|
// treating it as one would strip real cells out of the geometry and lose
|
||||||
|
// the table. Requiring a heading-sized anchor (>= 1.15x base) keeps the
|
||||||
|
// body pass to genuine scripts hanging off headings.
|
||||||
|
let script_index = ScriptBodyIndex::new(items);
|
||||||
|
let script_flags: Vec<bool> = items
|
||||||
|
.iter()
|
||||||
|
.map(|item| script_index.is_script_attachment(item, 0.0))
|
||||||
|
.collect();
|
||||||
|
let body_script_flags: Vec<bool> = items
|
||||||
|
.iter()
|
||||||
|
.map(|item| script_index.is_script_attachment(item, base_font_size * 1.15))
|
||||||
|
.collect();
|
||||||
let table_candidates: Vec<(usize, &TextItem)> = items
|
let table_candidates: Vec<(usize, &TextItem)> = items
|
||||||
.iter()
|
.iter()
|
||||||
.enumerate()
|
.enumerate()
|
||||||
@@ -494,7 +581,14 @@ pub(crate) fn detect_tables_with_page_width(
|
|||||||
.collect();
|
.collect();
|
||||||
|
|
||||||
if table_candidates.len() >= 6 {
|
if table_candidates.len() >= 6 {
|
||||||
let regions = find_table_regions(&table_candidates);
|
// Qualify regions from non-script items: a cluster of sub/superscripts
|
||||||
|
// must not, on its own, mark out a table region.
|
||||||
|
let region_evidence: Vec<(usize, &TextItem)> = table_candidates
|
||||||
|
.iter()
|
||||||
|
.filter(|(idx, _)| !script_flags[*idx])
|
||||||
|
.cloned()
|
||||||
|
.collect();
|
||||||
|
let regions = find_table_regions(®ion_evidence);
|
||||||
|
|
||||||
for (y_min, y_max) in regions {
|
for (y_min, y_max) in regions {
|
||||||
let region_items: Vec<(usize, &TextItem)> = table_candidates
|
let region_items: Vec<(usize, &TextItem)> = table_candidates
|
||||||
@@ -508,7 +602,9 @@ pub(crate) fn detect_tables_with_page_width(
|
|||||||
}
|
}
|
||||||
|
|
||||||
if let Some(mut table) =
|
if let Some(mut table) =
|
||||||
detect_table_in_region(®ion_items, TableDetectionMode::SmallFont)
|
detect_table_in_region(®ion_items, TableDetectionMode::SmallFont, &|i| {
|
||||||
|
script_flags[i]
|
||||||
|
})
|
||||||
{
|
{
|
||||||
// Try to recover body-font header row above the small-font table
|
// Try to recover body-font header row above the small-font table
|
||||||
recover_header_row(&mut table, items, table_font_threshold);
|
recover_header_row(&mut table, items, table_font_threshold);
|
||||||
@@ -553,8 +649,20 @@ pub(crate) fn detect_tables_with_page_width(
|
|||||||
body_font_low,
|
body_font_low,
|
||||||
body_font_high,
|
body_font_high,
|
||||||
);
|
);
|
||||||
|
// Scripts are NOT filtered out of the candidate set here, mirroring
|
||||||
|
// the small-font pass: they must stay eligible for cell assignment so
|
||||||
|
// a sub/superscript that belongs inside a table cell keeps its text.
|
||||||
|
// The heading-anchored `body_script_flags` mask removes them from
|
||||||
|
// geometry only.
|
||||||
if body_candidates.len() >= 6 {
|
if body_candidates.len() >= 6 {
|
||||||
let regions = find_table_regions_strict(&body_candidates);
|
// Same reasoning as the small-font pass: scripts do not qualify
|
||||||
|
// regions, but remain available for cell assignment within one.
|
||||||
|
let region_evidence: Vec<(usize, &TextItem)> = body_candidates
|
||||||
|
.iter()
|
||||||
|
.filter(|(idx, _)| !body_script_flags[*idx])
|
||||||
|
.cloned()
|
||||||
|
.collect();
|
||||||
|
let regions = find_table_regions_strict(®ion_evidence);
|
||||||
log::debug!("body-font: {} strict regions found", regions.len());
|
log::debug!("body-font: {} strict regions found", regions.len());
|
||||||
|
|
||||||
for (y_min, y_max, _x_min, _x_max) in ®ions {
|
for (y_min, y_max, _x_min, _x_max) in ®ions {
|
||||||
@@ -580,7 +688,9 @@ pub(crate) fn detect_tables_with_page_width(
|
|||||||
}
|
}
|
||||||
|
|
||||||
if let Some(table) =
|
if let Some(table) =
|
||||||
detect_table_in_region(®ion_items, TableDetectionMode::BodyFont)
|
detect_table_in_region(®ion_items, TableDetectionMode::BodyFont, &|i| {
|
||||||
|
body_script_flags[i]
|
||||||
|
})
|
||||||
{
|
{
|
||||||
tables.push(table);
|
tables.push(table);
|
||||||
}
|
}
|
||||||
@@ -808,10 +918,30 @@ fn find_table_regions_strict(items: &[(usize, &TextItem)]) -> Vec<(f32, f32, f32
|
|||||||
regions
|
regions
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Detect a table within a specific region
|
/// Detect a table within a specific region.
|
||||||
fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode) -> Option<Table> {
|
///
|
||||||
// Find column boundaries
|
/// `is_script` marks items that are sub/superscript attachments. Those are
|
||||||
let columns = find_column_boundaries(items, mode);
|
/// excluded from the *geometry* — they must not be able to create a column,
|
||||||
|
/// which is how equation subscript clusters used to fabricate phantom grids —
|
||||||
|
/// but they remain eligible for cell assignment, so legitimate cell content
|
||||||
|
/// (exponents in an engineering-notation table, footnote markers) stays in
|
||||||
|
/// the cell it belongs to instead of leaking out into the reading order.
|
||||||
|
fn detect_table_in_region(
|
||||||
|
items: &[(usize, &TextItem)],
|
||||||
|
mode: TableDetectionMode,
|
||||||
|
is_script: &dyn Fn(usize) -> bool,
|
||||||
|
) -> Option<Table> {
|
||||||
|
// Column geometry from non-script items only.
|
||||||
|
let geometry_items: Vec<(usize, &TextItem)> = items
|
||||||
|
.iter()
|
||||||
|
.filter(|(idx, _)| !is_script(*idx))
|
||||||
|
.cloned()
|
||||||
|
.collect();
|
||||||
|
// A region that is *entirely* scripts has no table structure at all.
|
||||||
|
if geometry_items.is_empty() {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
let columns = find_column_boundaries(&geometry_items, mode);
|
||||||
let min_cols = 2;
|
let min_cols = 2;
|
||||||
if columns.len() < min_cols || columns.len() > 25 {
|
if columns.len() < min_cols || columns.len() > 25 {
|
||||||
log::debug!(
|
log::debug!(
|
||||||
@@ -822,8 +952,8 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
|
|||||||
return None;
|
return None;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Find row boundaries
|
// Find row boundaries (geometry items only, same reasoning)
|
||||||
let rows = find_row_boundaries(items);
|
let rows = find_row_boundaries(&geometry_items);
|
||||||
let min_rows = 2;
|
let min_rows = 2;
|
||||||
if rows.len() < min_rows {
|
if rows.len() < min_rows {
|
||||||
log::debug!(
|
log::debug!(
|
||||||
@@ -842,6 +972,11 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
|
|||||||
);
|
);
|
||||||
|
|
||||||
// Verify this looks like a table: multiple items should align to columns
|
// Verify this looks like a table: multiple items should align to columns
|
||||||
|
// Validate against ALL items, including scripts. Columns are derived from
|
||||||
|
// non-script geometry so scripts cannot *create* a column, but excluding
|
||||||
|
// them from validation too would let a region manufacture alignment: drop
|
||||||
|
// the awkward items and whatever remains looks like a tidy grid. Block
|
||||||
|
// diagrams did exactly that. Everything in the region must fit.
|
||||||
let col_alignment = check_column_alignment(items, &columns, mode);
|
let col_alignment = check_column_alignment(items, &columns, mode);
|
||||||
let min_alignment = match mode {
|
let min_alignment = match mode {
|
||||||
TableDetectionMode::SmallFont => 0.5,
|
TableDetectionMode::SmallFont => 0.5,
|
||||||
@@ -912,6 +1047,29 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
|
|||||||
cells.push(row_cells);
|
cells.push(row_cells);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Validation 0 (small-font pass only): reject tiny all-numeric
|
||||||
|
// fragments. A <=2-row grid whose every cell is a bare 1-2 digit number
|
||||||
|
// carries no tabular information — in practice these are
|
||||||
|
// exponent/subscript clusters from display math that happen to align.
|
||||||
|
// Body-font tables are not subject to this veto: their cells cannot be
|
||||||
|
// script glyphs.
|
||||||
|
if matches!(mode, TableDetectionMode::SmallFont) {
|
||||||
|
let nonempty_cells: Vec<&String> =
|
||||||
|
cells.iter().flatten().filter(|c| !c.is_empty()).collect();
|
||||||
|
if rows.len() <= 2
|
||||||
|
&& !nonempty_cells.is_empty()
|
||||||
|
&& nonempty_cells
|
||||||
|
.iter()
|
||||||
|
.all(|c| c.len() <= 2 && c.chars().all(|ch| ch.is_ascii_digit()))
|
||||||
|
{
|
||||||
|
log::debug!(
|
||||||
|
" validation 0 fail: tiny all-numeric fragment ({} cells)",
|
||||||
|
nonempty_cells.len()
|
||||||
|
);
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// Validation 1: some rows should have content in first column.
|
// Validation 1: some rows should have content in first column.
|
||||||
// Use a lower threshold (25%) for tables with wrapped cells where
|
// Use a lower threshold (25%) for tables with wrapped cells where
|
||||||
// continuation lines leave the first column empty.
|
// continuation lines leave the first column empty.
|
||||||
@@ -1977,6 +2135,146 @@ fn try_add_label_column(
|
|||||||
|
|
||||||
#[cfg(test)]
|
#[cfg(test)]
|
||||||
mod tests {
|
mod tests {
|
||||||
|
|
||||||
|
fn make_item(text: &str, x: f32, y: f32, font_size: f32, width: f32) -> TextItem {
|
||||||
|
TextItem {
|
||||||
|
text: text.to_string(),
|
||||||
|
x,
|
||||||
|
y,
|
||||||
|
width,
|
||||||
|
height: font_size,
|
||||||
|
font: "TestFont".to_string(),
|
||||||
|
font_size,
|
||||||
|
page: 1,
|
||||||
|
is_bold: false,
|
||||||
|
is_italic: false,
|
||||||
|
is_underline: false,
|
||||||
|
is_strikeout: false,
|
||||||
|
item_type: ItemType::Text,
|
||||||
|
mcid: None,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn script_attachment_detects_subscript_after_body_text() {
|
||||||
|
let body = make_item("log", 100.0, 500.0, 10.0, 15.0);
|
||||||
|
let sub = make_item("10", 115.5, 497.0, 7.0, 7.0);
|
||||||
|
let items = vec![body, sub.clone()];
|
||||||
|
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sub, 0.0));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn script_attachment_detects_superscript_footnote_marker() {
|
||||||
|
let body = make_item("Hartley", 200.0, 500.0, 10.0, 35.0);
|
||||||
|
let sup = make_item("2", 235.8, 504.0, 6.6, 3.5);
|
||||||
|
let items = vec![body, sup.clone()];
|
||||||
|
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sup, 0.0));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn script_attachment_ignores_small_cell_far_from_body_text() {
|
||||||
|
let body = make_item("Revenue", 100.0, 500.0, 10.0, 40.0);
|
||||||
|
let cell = make_item("1,234", 180.0, 500.0, 7.0, 20.0);
|
||||||
|
let items = vec![body, cell.clone()];
|
||||||
|
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn body_pass_anchor_spares_cells_beside_slightly_larger_labels() {
|
||||||
|
// A body-font table cell (10pt) sitting beside a slightly larger,
|
||||||
|
// NON-heading label (12.5pt) with a little baseline jitter. The
|
||||||
|
// small-font pass treats any larger neighbour as a possible script
|
||||||
|
// base, but the body pass must not: at body sizes a slightly larger
|
||||||
|
// neighbour is a bold label or column header, and flagging the cell
|
||||||
|
// would strip it out of the table geometry and lose the table.
|
||||||
|
// Cell at the low end of the body band (0.85x base) beside a 10.5pt
|
||||||
|
// label. 10.5 clears the inherent 1.2x-of-cell rule (10.2) but falls
|
||||||
|
// below the body pass's heading anchor (11.5), which is exactly the
|
||||||
|
// band where the two masks must disagree.
|
||||||
|
let label = make_item("Revenue", 100.0, 500.0, 10.5, 40.0);
|
||||||
|
let cell = make_item("1,234", 141.0, 496.5, 8.5, 22.0);
|
||||||
|
let items = vec![label, cell.clone()];
|
||||||
|
let index = ScriptBodyIndex::new(&items);
|
||||||
|
let base = 10.0;
|
||||||
|
assert!(
|
||||||
|
index.is_script_attachment(&cell, 0.0),
|
||||||
|
"small-font pass anchor should still see this as an attachment"
|
||||||
|
);
|
||||||
|
assert!(
|
||||||
|
!index.is_script_attachment(&cell, base * 1.15),
|
||||||
|
"body pass must not treat a cell beside a slightly larger label \
|
||||||
|
as a script — that removes real cells from the geometry"
|
||||||
|
);
|
||||||
|
// A genuine heading-sized anchor still qualifies in the body pass.
|
||||||
|
let heading = make_item("Section", 100.0, 500.0, 20.0, 60.0);
|
||||||
|
let sup = make_item("3", 161.0, 508.0, 10.0, 5.0);
|
||||||
|
let h_items = vec![heading, sup.clone()];
|
||||||
|
assert!(
|
||||||
|
ScriptBodyIndex::new(&h_items).is_script_attachment(&sup, base * 1.15),
|
||||||
|
"script hanging off a heading must still be excluded in the body pass"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn script_attachment_ignores_same_baseline_neighbor_cell() {
|
||||||
|
// A small cell beside a larger label on the SAME baseline is a table
|
||||||
|
// layout, not a subscript — a genuine baseline offset is required.
|
||||||
|
let label = make_item("Total", 100.0, 500.0, 10.0, 25.0);
|
||||||
|
let cell = make_item("42", 127.0, 500.0, 7.5, 9.0);
|
||||||
|
let items = vec![label, cell.clone()];
|
||||||
|
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn script_attachment_ignores_neighbor_on_different_line() {
|
||||||
|
let body = make_item("Header", 100.0, 500.0, 10.0, 30.0);
|
||||||
|
let cell = make_item("42", 131.0, 486.0, 7.0, 10.0);
|
||||||
|
let items = vec![body, cell.clone()];
|
||||||
|
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Equation-subscript + footnote layout from Shannon entropy.pdf page 1,
|
||||||
|
/// with real coordinates. Without the larger-font anchors the small items
|
||||||
|
/// alone DO form a phantom table — proving the layout reaches detection —
|
||||||
|
/// and adding the anchors must suppress it.
|
||||||
|
fn shannon_page1_small_items() -> Vec<TextItem> {
|
||||||
|
vec![
|
||||||
|
make_item("2", 267.4, 133.9, 7.4, 3.7),
|
||||||
|
make_item("10", 306.2, 133.9, 7.4, 7.4),
|
||||||
|
make_item("10", 342.7, 133.9, 7.4, 7.4),
|
||||||
|
make_item("10", 325.0, 118.9, 7.4, 7.4),
|
||||||
|
make_item("Bell System Technical Journal,", 295.7, 101.9, 8.0, 95.0),
|
||||||
|
make_item(
|
||||||
|
"April 1924, p. 324; Certain Topics in",
|
||||||
|
396.7,
|
||||||
|
101.9,
|
||||||
|
8.0,
|
||||||
|
130.0,
|
||||||
|
),
|
||||||
|
make_item("v. 47, April 1928, p. 617.", 250.9, 92.5, 8.0, 90.0),
|
||||||
|
make_item("Bell System Technical Journal,", 264.2, 82.6, 8.0, 95.0),
|
||||||
|
make_item("July 1928, p. 535.", 364.3, 82.6, 8.0, 65.0),
|
||||||
|
]
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn equation_scripts_do_not_form_phantom_table() {
|
||||||
|
let bare = shannon_page1_small_items();
|
||||||
|
assert!(
|
||||||
|
!detect_tables(&bare, 10.0, false).is_empty(),
|
||||||
|
"test layout must form a phantom table when the filter cannot fire"
|
||||||
|
);
|
||||||
|
let mut items = shannon_page1_small_items();
|
||||||
|
items.push(make_item("log", 253.0, 137.0, 10.0, 13.5));
|
||||||
|
items.push(make_item("log", 291.5, 137.0, 10.0, 13.5));
|
||||||
|
items.push(make_item("log", 328.0, 137.0, 10.0, 13.5));
|
||||||
|
items.push(make_item("log", 310.3, 122.0, 10.0, 13.5));
|
||||||
|
let tables = detect_tables(&items, 10.0, false);
|
||||||
|
assert!(
|
||||||
|
tables.is_empty(),
|
||||||
|
"equation scripts + footnotes must not become a table: {tables:?}"
|
||||||
|
);
|
||||||
|
}
|
||||||
use super::*;
|
use super::*;
|
||||||
use crate::types::ItemType;
|
use crate::types::ItemType;
|
||||||
|
|
||||||
|
|||||||
+450
-2
@@ -4,10 +4,10 @@
|
|||||||
//! gridlines. Many IRS forms and government PDFs use these instead of
|
//! gridlines. Many IRS forms and government PDFs use these instead of
|
||||||
//! `re` (rectangle) operators.
|
//! `re` (rectangle) operators.
|
||||||
|
|
||||||
use std::collections::HashSet;
|
use std::collections::{HashMap, HashSet};
|
||||||
|
|
||||||
use crate::tables::Table;
|
use crate::tables::Table;
|
||||||
use crate::types::{PdfLine, TextItem};
|
use crate::types::{PdfLine, PdfRect, TextItem};
|
||||||
|
|
||||||
use super::detect_rects::{assign_items_to_grid, snap_edges};
|
use super::detect_rects::{assign_items_to_grid, snap_edges};
|
||||||
|
|
||||||
@@ -15,11 +15,33 @@ const RULE_Y_TOLERANCE: f32 = 2.0;
|
|||||||
const RULE_JOIN_GAP: f32 = 6.0;
|
const RULE_JOIN_GAP: f32 = 6.0;
|
||||||
const RULE_SPAN_TOLERANCE: f32 = 8.0;
|
const RULE_SPAN_TOLERANCE: f32 = 8.0;
|
||||||
const TEXT_ROW_TOLERANCE: f32 = 2.5;
|
const TEXT_ROW_TOLERANCE: f32 = 2.5;
|
||||||
|
const DENSE_CHART_MIN_VERTICAL_EDGES: usize = 27;
|
||||||
|
const DENSE_CHART_LABEL_PAD: f32 = 20.0;
|
||||||
|
const DENSE_CHART_MAX_SHARED_PANEL_GRIDS: usize = 4;
|
||||||
|
|
||||||
type HorizontalRule = (f32, f32, f32); // (y, x_min, x_max)
|
type HorizontalRule = (f32, f32, f32); // (y, x_min, x_max)
|
||||||
type VerticalRule = (f32, f32, f32); // (x, y_min, y_max)
|
type VerticalRule = (f32, f32, f32); // (x, y_min, y_max)
|
||||||
type AnchoredRow<'a> = (f32, Vec<(usize, &'a TextItem)>);
|
type AnchoredRow<'a> = (f32, Vec<(usize, &'a TextItem)>);
|
||||||
|
|
||||||
|
fn dense_chart_grids_are_co_located(
|
||||||
|
left: (f32, f32, f32, f32),
|
||||||
|
right: (f32, f32, f32, f32),
|
||||||
|
) -> bool {
|
||||||
|
let left_width = left.2 - left.0;
|
||||||
|
let right_width = right.2 - right.0;
|
||||||
|
let left_height = left.3 - left.1;
|
||||||
|
let right_height = right.3 - right.1;
|
||||||
|
let horizontal_overlap = (left.2.min(right.2) - left.0.max(right.0)).max(0.0);
|
||||||
|
let vertical_overlap = (left.3.min(right.3) - left.1.max(right.1)).max(0.0);
|
||||||
|
let horizontal_gap = (left.0.max(right.0) - left.2.min(right.2)).max(0.0);
|
||||||
|
let vertical_gap = (left.1.max(right.1) - left.3.min(right.3)).max(0.0);
|
||||||
|
|
||||||
|
(vertical_overlap >= left_height.min(right_height) * 0.5
|
||||||
|
&& horizontal_gap <= left_width.min(right_width) * 0.5)
|
||||||
|
|| (horizontal_overlap >= left_width.min(right_width) * 0.5
|
||||||
|
&& vertical_gap <= left_height.min(right_height) * 0.5)
|
||||||
|
}
|
||||||
|
|
||||||
#[derive(Debug)]
|
#[derive(Debug)]
|
||||||
struct TextAnchorTable {
|
struct TextAnchorTable {
|
||||||
table: Table,
|
table: Table,
|
||||||
@@ -1187,6 +1209,279 @@ pub fn detect_tables_from_lines(items: &[TextItem], lines: &[PdfLine], page: u32
|
|||||||
detect_tables_from_lines_inner(items, lines, page, true, true)
|
detect_tables_from_lines_inner(items, lines, page, true, true)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Bounding boxes of chart panels backed by a very dense vector grid.
|
||||||
|
///
|
||||||
|
/// Tables support at most 25 columns, so a panel with at least 27 distinct,
|
||||||
|
/// long vertical coordinates plus repeated horizontal rules is treated as
|
||||||
|
/// chart geometry. When the grid is enclosed by a painted panel rectangle,
|
||||||
|
/// the region expands to that rectangle so axis labels, legends, and source
|
||||||
|
/// notes remain part of the figure instead of forming a heuristic table.
|
||||||
|
pub(crate) fn detect_dense_line_chart_regions(
|
||||||
|
lines: &[PdfLine],
|
||||||
|
rects: &[PdfRect],
|
||||||
|
page: u32,
|
||||||
|
) -> Vec<(f32, f32, f32, f32)> {
|
||||||
|
const ANGLE_TOLERANCE: f32 = 0.035;
|
||||||
|
const MIN_GRID_LINE_LENGTH: f32 = 40.0;
|
||||||
|
const EXTENT_TOLERANCE: f32 = 6.0;
|
||||||
|
|
||||||
|
let mut verticals = Vec::new();
|
||||||
|
let mut horizontals = Vec::new();
|
||||||
|
for line in lines.iter().filter(|line| line.page == page) {
|
||||||
|
let dx = (line.x2 - line.x1).abs();
|
||||||
|
let dy = (line.y2 - line.y1).abs();
|
||||||
|
let length = dx.hypot(dy);
|
||||||
|
if length < MIN_GRID_LINE_LENGTH {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if dy > 0.01 && dx / dy <= ANGLE_TOLERANCE {
|
||||||
|
verticals.push((
|
||||||
|
(line.x1 + line.x2) / 2.0,
|
||||||
|
line.y1.min(line.y2),
|
||||||
|
line.y1.max(line.y2),
|
||||||
|
));
|
||||||
|
} else if dx > 0.01 && dy / dx <= ANGLE_TOLERANCE {
|
||||||
|
horizontals.push((
|
||||||
|
(line.y1 + line.y2) / 2.0,
|
||||||
|
line.x1.min(line.x2),
|
||||||
|
line.x1.max(line.x2),
|
||||||
|
));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if verticals.len() < DENSE_CHART_MIN_VERTICAL_EDGES || horizontals.len() < 3 {
|
||||||
|
return Vec::new();
|
||||||
|
}
|
||||||
|
|
||||||
|
// Group similar vertical extents once. Neighboring buckets are consulted
|
||||||
|
// below so coordinates that straddle a bucket boundary still form one
|
||||||
|
// family, while each line participates in only a constant number of
|
||||||
|
// candidates instead of being re-scanned for every vertical anchor.
|
||||||
|
let extent_key = |value: f32| (value / EXTENT_TOLERANCE).round() as i32;
|
||||||
|
let mut extent_buckets: HashMap<(i32, i32), Vec<VerticalRule>> = HashMap::new();
|
||||||
|
for vertical in verticals {
|
||||||
|
extent_buckets
|
||||||
|
.entry((extent_key(vertical.1), extent_key(vertical.2)))
|
||||||
|
.or_default()
|
||||||
|
.push(vertical);
|
||||||
|
}
|
||||||
|
|
||||||
|
let mut grid_regions = Vec::new();
|
||||||
|
let extent_keys: Vec<(i32, i32)> = extent_buckets.keys().copied().collect();
|
||||||
|
for key in extent_keys {
|
||||||
|
let anchor_family = &extent_buckets[&key];
|
||||||
|
let anchor_bottom = anchor_family.iter().map(|vertical| vertical.1).sum::<f32>()
|
||||||
|
/ anchor_family.len() as f32;
|
||||||
|
let anchor_top = anchor_family.iter().map(|vertical| vertical.2).sum::<f32>()
|
||||||
|
/ anchor_family.len() as f32;
|
||||||
|
|
||||||
|
let mut family = Vec::new();
|
||||||
|
for bottom_offset in -1..=1 {
|
||||||
|
for top_offset in -1..=1 {
|
||||||
|
if let Some(bucket) =
|
||||||
|
extent_buckets.get(&(key.0 + bottom_offset, key.1 + top_offset))
|
||||||
|
{
|
||||||
|
family.extend(bucket.iter().copied().filter(|vertical| {
|
||||||
|
(vertical.1 - anchor_bottom).abs() <= EXTENT_TOLERANCE
|
||||||
|
&& (vertical.2 - anchor_top).abs() <= EXTENT_TOLERANCE
|
||||||
|
}));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
let xs = snap_edges(&family.iter().map(|&(x, _, _)| x).collect::<Vec<_>>(), 3.0);
|
||||||
|
if xs.len() < DENSE_CHART_MIN_VERTICAL_EDGES {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
let grid_bottom =
|
||||||
|
family.iter().map(|vertical| vertical.1).sum::<f32>() / family.len() as f32;
|
||||||
|
let grid_top = family.iter().map(|vertical| vertical.2).sum::<f32>() / family.len() as f32;
|
||||||
|
if grid_top - grid_bottom < 60.0 {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
// A horizontal rule must support the same contiguous dense run of
|
||||||
|
// vertical coordinates. Splitting at sparse X gaps prevents a shared
|
||||||
|
// rule from joining a chart to a neighboring ruled table. Keying by
|
||||||
|
// the covered X-index range also lets multiple chart panels sharing
|
||||||
|
// the same Y extents produce independent regions.
|
||||||
|
let mut supported_spans: HashMap<(usize, usize), Vec<f32>> = HashMap::new();
|
||||||
|
for &(y, line_left, line_right) in &horizontals {
|
||||||
|
if y < grid_bottom - EXTENT_TOLERANCE || y > grid_top + EXTENT_TOLERANCE {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let start = xs.partition_point(|&x| x < line_left - EXTENT_TOLERANCE);
|
||||||
|
let end = xs.partition_point(|&x| x <= line_right + EXTENT_TOLERANCE);
|
||||||
|
if end - start < DENSE_CHART_MIN_VERTICAL_EDGES {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
let mut gaps: Vec<f32> = xs[start..end]
|
||||||
|
.windows(2)
|
||||||
|
.map(|pair| pair[1] - pair[0])
|
||||||
|
.collect();
|
||||||
|
gaps.sort_by(f32::total_cmp);
|
||||||
|
let dense_gap = gaps[gaps.len() / 4];
|
||||||
|
let run_break = (dense_gap * 3.0).max(12.0);
|
||||||
|
let locally_dense_gap_limit = (dense_gap * 1.5).max(6.0);
|
||||||
|
let locally_dense_gaps = gaps
|
||||||
|
.iter()
|
||||||
|
.filter(|&&gap| gap <= locally_dense_gap_limit)
|
||||||
|
.count();
|
||||||
|
|
||||||
|
let mut run_start = start;
|
||||||
|
let mut retained_dense_run = false;
|
||||||
|
for index in start..end - 1 {
|
||||||
|
if xs[index + 1] - xs[index] <= run_break {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let run_end = index + 1;
|
||||||
|
if run_end - run_start >= DENSE_CHART_MIN_VERTICAL_EDGES
|
||||||
|
&& xs[run_end - 1] - xs[run_start] >= 120.0
|
||||||
|
{
|
||||||
|
supported_spans
|
||||||
|
.entry((run_start, run_end))
|
||||||
|
.or_default()
|
||||||
|
.push(y);
|
||||||
|
retained_dense_run = true;
|
||||||
|
}
|
||||||
|
run_start = run_end;
|
||||||
|
}
|
||||||
|
if end - run_start >= DENSE_CHART_MIN_VERTICAL_EDGES
|
||||||
|
&& xs[end - 1] - xs[run_start] >= 120.0
|
||||||
|
{
|
||||||
|
supported_spans.entry((run_start, end)).or_default().push(y);
|
||||||
|
retained_dense_run = true;
|
||||||
|
}
|
||||||
|
|
||||||
|
// One or two wider category gaps may split an otherwise dense
|
||||||
|
// chart into sub-threshold runs. Keep the full family only when
|
||||||
|
// its total width remains close to the expected dense spacing;
|
||||||
|
// a neighboring sparse table makes this ratio much larger.
|
||||||
|
let span_width = xs[end - 1] - xs[start];
|
||||||
|
let expected_dense_width = dense_gap * (end - start - 1) as f32;
|
||||||
|
if !retained_dense_run
|
||||||
|
&& span_width >= 120.0
|
||||||
|
&& span_width <= expected_dense_width * 1.35
|
||||||
|
&& gaps.len().saturating_sub(locally_dense_gaps) <= 2
|
||||||
|
{
|
||||||
|
supported_spans.entry((start, end)).or_default().push(y);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
for ((start, end), ys) in supported_spans {
|
||||||
|
if snap_edges(&ys, 3.0).len() >= 3 {
|
||||||
|
grid_regions.push((xs[start], grid_bottom, xs[end - 1], grid_top));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Prefer the smallest qualifying region when a broad rule happens to
|
||||||
|
// cover a denser nested panel, and retain every non-overlapping panel.
|
||||||
|
grid_regions.sort_by(|left, right| {
|
||||||
|
let left_area = (left.2 - left.0) * (left.3 - left.1);
|
||||||
|
let right_area = (right.2 - right.0) * (right.3 - right.1);
|
||||||
|
left_area.total_cmp(&right_area)
|
||||||
|
});
|
||||||
|
let mut selected_regions: Vec<(f32, f32, f32, f32)> = Vec::new();
|
||||||
|
for region in grid_regions {
|
||||||
|
let area = (region.2 - region.0) * (region.3 - region.1);
|
||||||
|
let duplicates_existing = selected_regions.iter().any(|existing| {
|
||||||
|
let overlap_width = (region.2.min(existing.2) - region.0.max(existing.0)).max(0.0);
|
||||||
|
let overlap_height = (region.3.min(existing.3) - region.1.max(existing.1)).max(0.0);
|
||||||
|
let overlap_area = overlap_width * overlap_height;
|
||||||
|
let existing_area = (existing.2 - existing.0) * (existing.3 - existing.1);
|
||||||
|
overlap_area >= area.min(existing_area) * 0.8
|
||||||
|
});
|
||||||
|
if !duplicates_existing {
|
||||||
|
selected_regions.push(region);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
let all_grid_regions = selected_regions.clone();
|
||||||
|
let mut regions: Vec<_> = selected_regions
|
||||||
|
.into_iter()
|
||||||
|
.map(|grid_region| {
|
||||||
|
let (grid_left, grid_bottom, grid_right, grid_top) = grid_region;
|
||||||
|
let enclosing_panel = rects
|
||||||
|
.iter()
|
||||||
|
.filter(|rect| rect.page == page)
|
||||||
|
.filter_map(|rect| {
|
||||||
|
let (left, width) = if rect.width < 0.0 {
|
||||||
|
(rect.x + rect.width, -rect.width)
|
||||||
|
} else {
|
||||||
|
(rect.x, rect.width)
|
||||||
|
};
|
||||||
|
let (bottom, height) = if rect.height < 0.0 {
|
||||||
|
(rect.y + rect.height, -rect.height)
|
||||||
|
} else {
|
||||||
|
(rect.y, rect.height)
|
||||||
|
};
|
||||||
|
let right = left + width;
|
||||||
|
let top = bottom + height;
|
||||||
|
let enclosed_grids: Vec<_> = all_grid_regions
|
||||||
|
.iter()
|
||||||
|
.filter(|&&(other_left, other_bottom, other_right, other_top)| {
|
||||||
|
left <= other_left + EXTENT_TOLERANCE
|
||||||
|
&& right >= other_right - EXTENT_TOLERANCE
|
||||||
|
&& bottom <= other_bottom + EXTENT_TOLERANCE
|
||||||
|
&& top >= other_top - EXTENT_TOLERANCE
|
||||||
|
})
|
||||||
|
.copied()
|
||||||
|
.collect();
|
||||||
|
if enclosed_grids.len() > DENSE_CHART_MAX_SHARED_PANEL_GRIDS
|
||||||
|
|| enclosed_grids.iter().any(|&other| {
|
||||||
|
other != grid_region
|
||||||
|
&& !dense_chart_grids_are_co_located(grid_region, other)
|
||||||
|
})
|
||||||
|
{
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
let enclosed_grid_bounds =
|
||||||
|
enclosed_grids.into_iter().reduce(|bounds, other| {
|
||||||
|
(
|
||||||
|
bounds.0.min(other.0),
|
||||||
|
bounds.1.min(other.1),
|
||||||
|
bounds.2.max(other.2),
|
||||||
|
bounds.3.max(other.3),
|
||||||
|
)
|
||||||
|
})?;
|
||||||
|
let enclosed_width = enclosed_grid_bounds.2 - enclosed_grid_bounds.0;
|
||||||
|
let enclosed_height = enclosed_grid_bounds.3 - enclosed_grid_bounds.1;
|
||||||
|
(left <= grid_left + EXTENT_TOLERANCE
|
||||||
|
&& right >= grid_right - EXTENT_TOLERANCE
|
||||||
|
&& bottom <= grid_bottom + EXTENT_TOLERANCE
|
||||||
|
&& top >= grid_top - EXTENT_TOLERANCE
|
||||||
|
&& width <= enclosed_width * 2.0
|
||||||
|
&& height <= enclosed_height * 4.0
|
||||||
|
&& !(left < 5.0 && bottom < 5.0))
|
||||||
|
.then_some(((left, bottom, right, top), width * height))
|
||||||
|
})
|
||||||
|
.min_by(|left, right| left.1.total_cmp(&right.1))
|
||||||
|
.map(|(region, _)| region);
|
||||||
|
|
||||||
|
enclosing_panel.unwrap_or((
|
||||||
|
grid_left - DENSE_CHART_LABEL_PAD,
|
||||||
|
grid_bottom - DENSE_CHART_LABEL_PAD,
|
||||||
|
grid_right + DENSE_CHART_LABEL_PAD,
|
||||||
|
grid_top + DENSE_CHART_LABEL_PAD,
|
||||||
|
))
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
regions.sort_by(|left, right| {
|
||||||
|
left.0
|
||||||
|
.total_cmp(&right.0)
|
||||||
|
.then_with(|| left.1.total_cmp(&right.1))
|
||||||
|
});
|
||||||
|
regions.dedup_by(|left, right| {
|
||||||
|
(left.0 - right.0).abs() <= EXTENT_TOLERANCE
|
||||||
|
&& (left.1 - right.1).abs() <= EXTENT_TOLERANCE
|
||||||
|
&& (left.2 - right.2).abs() <= EXTENT_TOLERANCE
|
||||||
|
&& (left.3 - right.3).abs() <= EXTENT_TOLERANCE
|
||||||
|
});
|
||||||
|
regions
|
||||||
|
}
|
||||||
|
|
||||||
/// Detect only tables whose cell grid is backed by explicit vector geometry.
|
/// Detect only tables whose cell grid is backed by explicit vector geometry.
|
||||||
///
|
///
|
||||||
/// Region-level TSR callers need physical cell boundaries for crop bboxes, so
|
/// Region-level TSR callers need physical cell boundaries for crop bboxes, so
|
||||||
@@ -1621,6 +1916,159 @@ mod tests {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn dense_vector_grid_expands_to_enclosing_chart_panel() {
|
||||||
|
let mut lines: Vec<PdfLine> = (0..30)
|
||||||
|
.map(|column| make_vline(100.0 + column as f32 * 8.0, 400.0, 550.0, 1))
|
||||||
|
.collect();
|
||||||
|
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 332.0, 1)));
|
||||||
|
let rects = vec![PdfRect {
|
||||||
|
x: 80.0,
|
||||||
|
y: 350.0,
|
||||||
|
width: 280.0,
|
||||||
|
height: 240.0,
|
||||||
|
page: 1,
|
||||||
|
}];
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
detect_dense_line_chart_regions(&lines, &rects, 1),
|
||||||
|
vec![(80.0, 350.0, 360.0, 590.0)]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn frameless_dense_vector_grid_includes_label_padding() {
|
||||||
|
let mut lines: Vec<PdfLine> = (0..30)
|
||||||
|
.map(|column| make_vline(100.0 + column as f32 * 8.0, 400.0, 550.0, 1))
|
||||||
|
.collect();
|
||||||
|
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 332.0, 1)));
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
detect_dense_line_chart_regions(&lines, &[], 1),
|
||||||
|
vec![(80.0, 380.0, 352.0, 570.0)]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn multiple_dense_vector_panels_are_retained() {
|
||||||
|
let mut lines = Vec::new();
|
||||||
|
for panel_left in [60.0, 380.0] {
|
||||||
|
lines.extend(
|
||||||
|
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
|
||||||
|
);
|
||||||
|
lines.extend((0..6).map(|row| {
|
||||||
|
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
|
||||||
|
}));
|
||||||
|
}
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
detect_dense_line_chart_regions(&lines, &[], 1),
|
||||||
|
vec![(40.0, 380.0, 312.0, 570.0), (360.0, 380.0, 632.0, 570.0),]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn multiple_dense_vector_panels_use_shared_enclosing_panel() {
|
||||||
|
let mut lines = Vec::new();
|
||||||
|
for panel_left in [60.0, 380.0] {
|
||||||
|
lines.extend(
|
||||||
|
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
|
||||||
|
);
|
||||||
|
lines.extend((0..6).map(|row| {
|
||||||
|
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
|
||||||
|
}));
|
||||||
|
}
|
||||||
|
let rects = vec![PdfRect {
|
||||||
|
x: 40.0,
|
||||||
|
y: 350.0,
|
||||||
|
width: 592.0,
|
||||||
|
height: 240.0,
|
||||||
|
page: 1,
|
||||||
|
}];
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
detect_dense_line_chart_regions(&lines, &rects, 1),
|
||||||
|
vec![(40.0, 350.0, 632.0, 590.0)]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn shared_rules_do_not_join_dense_chart_to_adjacent_table() {
|
||||||
|
let mut lines: Vec<PdfLine> = (0..30)
|
||||||
|
.map(|column| make_vline(60.0 + column as f32 * 8.0, 400.0, 550.0, 1))
|
||||||
|
.collect();
|
||||||
|
|
||||||
|
lines.extend(
|
||||||
|
[330.0, 390.0, 450.0, 510.0, 570.0, 630.0]
|
||||||
|
.into_iter()
|
||||||
|
.map(|x| make_vline(x, 400.0, 550.0, 1)),
|
||||||
|
);
|
||||||
|
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 630.0, 1)));
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
detect_dense_line_chart_regions(&lines, &[], 1),
|
||||||
|
vec![(40.0, 380.0, 312.0, 570.0)]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn uneven_dense_spacing_keeps_the_complete_chart_region() {
|
||||||
|
let mut xs: Vec<f32> = (0..15).map(|column| 60.0 + column as f32 * 8.0).collect();
|
||||||
|
xs.extend((0..15).map(|column| 212.0 + column as f32 * 8.0));
|
||||||
|
let mut lines: Vec<PdfLine> = xs.iter().map(|&x| make_vline(x, 400.0, 550.0, 1)).collect();
|
||||||
|
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 324.0, 1)));
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
detect_dense_line_chart_regions(&lines, &[], 1),
|
||||||
|
vec![(40.0, 380.0, 344.0, 570.0)]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn subthreshold_dense_run_does_not_absorb_adjacent_sparse_grid() {
|
||||||
|
let mut xs: Vec<f32> = (0..21).map(|column| 60.0 + column as f32 * 8.0).collect();
|
||||||
|
xs.extend((0..6).map(|column| 248.0 + column as f32 * 18.0));
|
||||||
|
let mut lines: Vec<PdfLine> = xs.iter().map(|&x| make_vline(x, 400.0, 550.0, 1)).collect();
|
||||||
|
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 338.0, 1)));
|
||||||
|
|
||||||
|
assert!(detect_dense_line_chart_regions(&lines, &[], 1).is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn broad_frame_does_not_merge_distant_dense_grids() {
|
||||||
|
let mut lines = Vec::new();
|
||||||
|
for panel_left in [60.0, 700.0] {
|
||||||
|
lines.extend(
|
||||||
|
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
|
||||||
|
);
|
||||||
|
lines.extend((0..6).map(|row| {
|
||||||
|
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
|
||||||
|
}));
|
||||||
|
}
|
||||||
|
let rects = vec![PdfRect {
|
||||||
|
x: 40.0,
|
||||||
|
y: 350.0,
|
||||||
|
width: 912.0,
|
||||||
|
height: 240.0,
|
||||||
|
page: 1,
|
||||||
|
}];
|
||||||
|
|
||||||
|
assert_eq!(
|
||||||
|
detect_dense_line_chart_regions(&lines, &rects, 1),
|
||||||
|
vec![(40.0, 380.0, 312.0, 570.0), (680.0, 380.0, 952.0, 570.0)]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn supported_width_vector_table_is_not_a_dense_chart() {
|
||||||
|
let mut lines: Vec<PdfLine> = (0..26)
|
||||||
|
.map(|column| make_vline(100.0 + column as f32 * 10.0, 400.0, 550.0, 1))
|
||||||
|
.collect();
|
||||||
|
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 350.0, 1)));
|
||||||
|
|
||||||
|
assert!(detect_dense_line_chart_regions(&lines, &[], 1).is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_basic_grid_detection() {
|
fn test_basic_grid_detection() {
|
||||||
// 3x2 grid with horizontal lines at y=500, 480, 460 and vertical at x=100, 200, 300
|
// 3x2 grid with horizontal lines at y=500, 480, 460 and vertical at x=100, 200, 300
|
||||||
|
|||||||
+454
-10
@@ -1996,17 +1996,249 @@ fn without_dominant_page_backgrounds(rects: &[(f32, f32, f32, f32)]) -> Vec<(f32
|
|||||||
.collect()
|
.collect()
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Detect a table from cell-background rects that failed grid detection.
|
/// Repeated rows of touching cell rectangles are stronger table evidence
|
||||||
|
/// than the bar-length variation used by the chart detector.
|
||||||
///
|
///
|
||||||
/// Uses rect Y-edges for row boundaries and text X-position clustering for
|
/// Ruled tables with wrapped labels naturally have variable row heights, and
|
||||||
/// columns. Handles tables with cell backgrounds that don't form a clean
|
/// numeric-heavy cells can otherwise resemble horizontal or vertical bars.
|
||||||
/// X-edge grid (variable column widths, decorative fills).
|
/// Require several rows to repeat a shared edge schema before overriding the
|
||||||
/// Chart-bar signature: ≥3 rects sharing an aligned bottom edge (the axis),
|
/// chart hypothesis so sparse plots and independent bars remain unaffected.
|
||||||
/// with similar widths (bars) but strongly varying heights (data-driven),
|
fn is_repeated_cell_grid(group_rects: &[(f32, f32, f32, f32)]) -> bool {
|
||||||
/// holding at most a single numeric data label each. Bar charts drawn as
|
type RowGroup = (f32, f32, Vec<(f32, f32)>);
|
||||||
/// filled rects otherwise read as cell rects and grid their axis labels
|
|
||||||
/// into a phantom table. The mirrored check catches horizontal bar charts.
|
const ROW_EDGE_TOLERANCE: f32 = 3.0;
|
||||||
fn is_chart_bar_cluster(
|
const MIN_GRID_ROWS: usize = 4;
|
||||||
|
const MIN_CELLS_PER_ROW: usize = 3;
|
||||||
|
|
||||||
|
if group_rects.len() < MIN_GRID_ROWS * MIN_CELLS_PER_ROW {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
let mut row_groups: Vec<RowGroup> = Vec::new();
|
||||||
|
for &(x, y, width, height) in group_rects {
|
||||||
|
if width < 5.0 || height < 5.0 {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let top = y + height;
|
||||||
|
if let Some((_, _, cells)) = row_groups.iter_mut().find(|(bottom, row_top, _)| {
|
||||||
|
(y - *bottom).abs() <= ROW_EDGE_TOLERANCE
|
||||||
|
&& (top - *row_top).abs() <= ROW_EDGE_TOLERANCE
|
||||||
|
}) {
|
||||||
|
cells.push((x, x + width));
|
||||||
|
} else {
|
||||||
|
row_groups.push((y, top, vec![(x, x + width)]));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
let mut row_schemas = Vec::new();
|
||||||
|
for (_, _, mut cells) in row_groups {
|
||||||
|
if cells.len() < MIN_CELLS_PER_ROW {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let mut widths: Vec<f32> = cells.iter().map(|&(left, right)| right - left).collect();
|
||||||
|
widths.sort_by(f32::total_cmp);
|
||||||
|
let median_width = widths[widths.len() / 2];
|
||||||
|
cells.retain(|&(left, right)| right - left <= median_width * 2.5);
|
||||||
|
cells.sort_by(|left, right| {
|
||||||
|
left.0
|
||||||
|
.total_cmp(&right.0)
|
||||||
|
.then_with(|| left.1.total_cmp(&right.1))
|
||||||
|
});
|
||||||
|
cells.dedup_by(|left, right| {
|
||||||
|
(left.0 - right.0).abs() <= ROW_EDGE_TOLERANCE
|
||||||
|
&& (left.1 - right.1).abs() <= ROW_EDGE_TOLERANCE
|
||||||
|
});
|
||||||
|
if cells.len() < MIN_CELLS_PER_ROW
|
||||||
|
|| cells
|
||||||
|
.windows(2)
|
||||||
|
.any(|pair| pair[1].0 > pair[0].1 + ROW_EDGE_TOLERANCE)
|
||||||
|
{
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let edges: Vec<f32> = cells
|
||||||
|
.iter()
|
||||||
|
.flat_map(|&(left, right)| [left, right])
|
||||||
|
.collect();
|
||||||
|
let schema = snap_edges(&edges, ROW_EDGE_TOLERANCE);
|
||||||
|
if schema.len() > MIN_CELLS_PER_ROW {
|
||||||
|
row_schemas.push(schema);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if row_schemas.len() < MIN_GRID_ROWS {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
let reference = row_schemas
|
||||||
|
.iter()
|
||||||
|
.max_by_key(|schema| schema.len())
|
||||||
|
.expect("grid rows are non-empty");
|
||||||
|
row_schemas
|
||||||
|
.iter()
|
||||||
|
.filter(|schema| {
|
||||||
|
let comparable_edges = reference.len().min(schema.len());
|
||||||
|
let matched_edges = schema
|
||||||
|
.iter()
|
||||||
|
.filter(|edge| {
|
||||||
|
reference
|
||||||
|
.iter()
|
||||||
|
.any(|reference_edge| (*edge - *reference_edge).abs() <= ROW_EDGE_TOLERANCE)
|
||||||
|
})
|
||||||
|
.count();
|
||||||
|
matched_edges > MIN_CELLS_PER_ROW && matched_edges * 4 >= comparable_edges * 3
|
||||||
|
})
|
||||||
|
.count()
|
||||||
|
>= MIN_GRID_ROWS
|
||||||
|
}
|
||||||
|
|
||||||
|
fn repeated_cell_grid_overrides_bar_hypothesis(group_rects: &[(f32, f32, f32, f32)]) -> bool {
|
||||||
|
is_repeated_cell_grid(group_rects)
|
||||||
|
&& without_dominant_page_backgrounds(group_rects).len() == group_rects.len()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Detect horizontal segmented stacks from aligned rows of touching rects.
|
||||||
|
///
|
||||||
|
/// Category rows must have visible gutters and data-varying internal segment
|
||||||
|
/// boundaries, unlike the stable boundaries of a ruled table.
|
||||||
|
struct SegmentedBarGeometry {
|
||||||
|
bounds: (f32, f32, f32, f32),
|
||||||
|
row_bands: Vec<(f32, f32)>,
|
||||||
|
}
|
||||||
|
|
||||||
|
fn segmented_stacked_bar_geometry(
|
||||||
|
group_rects: &[(f32, f32, f32, f32)],
|
||||||
|
) -> Option<SegmentedBarGeometry> {
|
||||||
|
type BarRow = (f32, f32, Vec<(f32, f32)>);
|
||||||
|
|
||||||
|
const EDGE_TOLERANCE: f32 = 3.0;
|
||||||
|
const MIN_ROWS: usize = 4;
|
||||||
|
const MIN_SEGMENTS: usize = 3;
|
||||||
|
|
||||||
|
let mut rows: Vec<BarRow> = Vec::new();
|
||||||
|
for &(x, y, width, height) in group_rects {
|
||||||
|
if width < 5.0 || height < 5.0 {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let top = y + height;
|
||||||
|
if let Some((_, _, segments)) = rows.iter_mut().find(|(bottom, row_top, _)| {
|
||||||
|
(y - *bottom).abs() <= EDGE_TOLERANCE && (top - *row_top).abs() <= EDGE_TOLERANCE
|
||||||
|
}) {
|
||||||
|
segments.push((x, x + width));
|
||||||
|
} else {
|
||||||
|
rows.push((y, top, vec![(x, x + width)]));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
rows.retain_mut(|(_, _, segments)| {
|
||||||
|
segments.sort_by(|left, right| left.0.total_cmp(&right.0));
|
||||||
|
segments.len() >= MIN_SEGMENTS
|
||||||
|
&& segments
|
||||||
|
.windows(2)
|
||||||
|
.all(|pair| (pair[1].0 - pair[0].1).abs() <= EDGE_TOLERANCE)
|
||||||
|
});
|
||||||
|
if rows.len() < MIN_ROWS {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
rows.sort_by(|left, right| left.0.total_cmp(&right.0));
|
||||||
|
|
||||||
|
// Table rows normally share borders. Horizontal stacked bars instead
|
||||||
|
// leave a visible gutter between category rows.
|
||||||
|
if rows.windows(2).any(|pair| {
|
||||||
|
let shorter_height = (pair[0].1 - pair[0].0).min(pair[1].1 - pair[1].0);
|
||||||
|
pair[1].0 - pair[0].1 < (shorter_height * 0.25).max(2.0)
|
||||||
|
}) {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
|
||||||
|
// At least two rows must move an internal segment boundary. Stable
|
||||||
|
// boundaries across every row are stronger evidence for a ruled table.
|
||||||
|
let reference_edges: Vec<f32> = rows[0]
|
||||||
|
.2
|
||||||
|
.iter()
|
||||||
|
.take(rows[0].2.len() - 1)
|
||||||
|
.map(|segment| segment.1)
|
||||||
|
.collect();
|
||||||
|
let drifting_rows = rows
|
||||||
|
.iter()
|
||||||
|
.skip(1)
|
||||||
|
.filter(|(_, _, segments)| {
|
||||||
|
let edges: Vec<f32> = segments
|
||||||
|
.iter()
|
||||||
|
.take(segments.len() - 1)
|
||||||
|
.map(|segment| segment.1)
|
||||||
|
.collect();
|
||||||
|
edges.len() == reference_edges.len()
|
||||||
|
&& edges
|
||||||
|
.iter()
|
||||||
|
.zip(&reference_edges)
|
||||||
|
.any(|(edge, reference)| (edge - reference).abs() > EDGE_TOLERANCE)
|
||||||
|
})
|
||||||
|
.count();
|
||||||
|
if drifting_rows < 2 {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
|
||||||
|
let left = rows
|
||||||
|
.iter()
|
||||||
|
.flat_map(|row| &row.2)
|
||||||
|
.map(|segment| segment.0)
|
||||||
|
.reduce(f32::min)?;
|
||||||
|
let right = rows
|
||||||
|
.iter()
|
||||||
|
.flat_map(|row| &row.2)
|
||||||
|
.map(|segment| segment.1)
|
||||||
|
.reduce(f32::max)?;
|
||||||
|
let bottom = rows.iter().map(|row| row.0).reduce(f32::min)?;
|
||||||
|
let top = rows.iter().map(|row| row.1).reduce(f32::max)?;
|
||||||
|
let row_bands = rows.iter().map(|row| (row.0, row.1)).collect();
|
||||||
|
Some(SegmentedBarGeometry {
|
||||||
|
bounds: (left, bottom, right, top),
|
||||||
|
row_bands,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Category labels beside multiple bar rows are independent chart evidence:
|
||||||
|
/// numeric table text stays inside its cells, regardless of whether the table
|
||||||
|
/// has an outer border or extra padding.
|
||||||
|
fn has_external_segmented_bar_labels(
|
||||||
|
items: &[TextItem],
|
||||||
|
page: u32,
|
||||||
|
geometry: &SegmentedBarGeometry,
|
||||||
|
) -> bool {
|
||||||
|
const LABEL_EDGE_TOLERANCE: f32 = 3.0;
|
||||||
|
const LABEL_CLAIM_PAD: f32 = 20.0;
|
||||||
|
|
||||||
|
let (content_left, _, content_right, _) = geometry.bounds;
|
||||||
|
let labeled_rows = geometry
|
||||||
|
.row_bands
|
||||||
|
.iter()
|
||||||
|
.filter(|&&(row_bottom, row_top)| {
|
||||||
|
items.iter().any(|item| {
|
||||||
|
if item.page != page || item.text.trim().is_empty() {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
let item_left = item.x.min(item.x + item.width);
|
||||||
|
let item_right = item.x.max(item.x + item.width);
|
||||||
|
let item_center_x = (item_left + item_right) / 2.0;
|
||||||
|
let item_center_y = item.y + item.height / 2.0;
|
||||||
|
let beside_stack = (item_center_x <= content_left + LABEL_EDGE_TOLERANCE
|
||||||
|
&& item_center_x >= content_left - LABEL_CLAIM_PAD
|
||||||
|
&& item_left < content_left)
|
||||||
|
|| (item_center_x >= content_right - LABEL_EDGE_TOLERANCE
|
||||||
|
&& item_center_x <= content_right + LABEL_CLAIM_PAD
|
||||||
|
&& item_right > content_right);
|
||||||
|
beside_stack
|
||||||
|
&& item_center_y >= row_bottom - LABEL_EDGE_TOLERANCE
|
||||||
|
&& item_center_y <= row_top + LABEL_EDGE_TOLERANCE
|
||||||
|
})
|
||||||
|
})
|
||||||
|
.count();
|
||||||
|
|
||||||
|
labeled_rows >= 2 && labeled_rows * 2 >= geometry.row_bands.len()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Recognize filled vertical or horizontal bars whose geometry and labels are
|
||||||
|
/// data-driven rather than uniform table cells.
|
||||||
|
fn has_chart_bar_signature(
|
||||||
items: &[TextItem],
|
items: &[TextItem],
|
||||||
group_rects: &[(f32, f32, f32, f32)],
|
group_rects: &[(f32, f32, f32, f32)],
|
||||||
page: u32,
|
page: u32,
|
||||||
@@ -2113,6 +2345,30 @@ fn is_chart_bar_cluster(
|
|||||||
|| bar_family(|r| r.1, |r| r.3, |r| r.2, |r| r.0)
|
|| bar_family(|r| r.1, |r| r.3, |r| r.2, |r| r.0)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn is_chart_bar_cluster(
|
||||||
|
items: &[TextItem],
|
||||||
|
group_rects: &[(f32, f32, f32, f32)],
|
||||||
|
page: u32,
|
||||||
|
) -> bool {
|
||||||
|
let has_bar_signature = has_chart_bar_signature(items, group_rects, page);
|
||||||
|
|
||||||
|
// A segmented horizontal chart can share most of its edges across rows.
|
||||||
|
// Row-aligned category labels outside the stack distinguish it from a
|
||||||
|
// numeric table without depending on whether either shape has a frame.
|
||||||
|
if has_bar_signature {
|
||||||
|
if let Some(geometry) = segmented_stacked_bar_geometry(group_rects) {
|
||||||
|
if has_external_segmented_bar_labels(items, page, &geometry) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if repeated_cell_grid_overrides_bar_hypothesis(group_rects) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
has_bar_signature
|
||||||
|
}
|
||||||
|
|
||||||
fn detect_row_stripe_table_from_cell_rects(
|
fn detect_row_stripe_table_from_cell_rects(
|
||||||
items: &[TextItem],
|
items: &[TextItem],
|
||||||
group_rects: &[(f32, f32, f32, f32)],
|
group_rects: &[(f32, f32, f32, f32)],
|
||||||
@@ -3083,6 +3339,194 @@ mod tests {
|
|||||||
assert!(detect_chart_regions(&items, &rects, 1).is_empty());
|
assert!(detect_chart_regions(&items, &rects, 1).is_empty());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn variable_height_ruled_grid_overrides_bar_hypothesis() {
|
||||||
|
let edge_sets = [
|
||||||
|
[80.0, 140.0, 200.0, 260.0, 320.0, 380.0, 440.0, 500.0, 560.0],
|
||||||
|
[80.0, 140.0, 210.0, 260.0, 320.0, 380.0, 450.0, 500.0, 560.0],
|
||||||
|
];
|
||||||
|
let heights = [20.0, 34.0, 26.0, 42.0, 20.0, 34.0];
|
||||||
|
let edge_variants = [0, 0, 0, 0, 1, 1];
|
||||||
|
let mut rects = Vec::new();
|
||||||
|
let mut y = 650.0;
|
||||||
|
for (row, height) in heights.into_iter().enumerate() {
|
||||||
|
let edges = edge_sets[edge_variants[row]];
|
||||||
|
rects.extend(
|
||||||
|
edges
|
||||||
|
.windows(2)
|
||||||
|
.map(|edge| (edge[0], y, edge[1] - edge[0], height)),
|
||||||
|
);
|
||||||
|
y -= height;
|
||||||
|
}
|
||||||
|
|
||||||
|
assert!(is_repeated_cell_grid(&rects));
|
||||||
|
assert!(has_chart_bar_signature(&[], &rects, 1));
|
||||||
|
assert!(repeated_cell_grid_overrides_bar_hypothesis(&rects));
|
||||||
|
assert!(segmented_stacked_bar_geometry(&rects).is_none());
|
||||||
|
assert!(!is_chart_bar_cluster(&[], &rects, 1));
|
||||||
|
|
||||||
|
let mut with_page_fills =
|
||||||
|
vec![(0.0, 0.0, 600.0, 800.0); DOMINANT_PAGE_BACKGROUND_MIN_REPETITIONS];
|
||||||
|
with_page_fills.extend(rects);
|
||||||
|
assert!(!repeated_cell_grid_overrides_bar_hypothesis(
|
||||||
|
&with_page_fills
|
||||||
|
));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn touching_segments_with_spaced_rows_remain_a_chart() {
|
||||||
|
let row_edges = [
|
||||||
|
[100.0, 140.0, 180.0, 220.0, 260.0],
|
||||||
|
[100.0, 140.0, 180.0, 228.0, 260.0],
|
||||||
|
[100.0, 140.0, 180.0, 214.0, 260.0],
|
||||||
|
[100.0, 140.0, 180.0, 232.0, 260.0],
|
||||||
|
];
|
||||||
|
let mut raw_rects = vec![(90.0, 530.0, 190.0, 100.0)];
|
||||||
|
for (row, edges) in row_edges.into_iter().enumerate() {
|
||||||
|
let y = 540.0 + row as f32 * 20.0;
|
||||||
|
raw_rects.extend(
|
||||||
|
edges
|
||||||
|
.windows(2)
|
||||||
|
.map(|edge| (edge[0], y, edge[1] - edge[0], 12.0)),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
let items: Vec<TextItem> = (0..4)
|
||||||
|
.map(|row| make_item("Category", 62.0, 541.0 + row as f32 * 20.0, 9.0))
|
||||||
|
.collect();
|
||||||
|
|
||||||
|
assert!(is_repeated_cell_grid(&raw_rects));
|
||||||
|
assert!(has_chart_bar_signature(&items, &raw_rects, 1));
|
||||||
|
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented stack");
|
||||||
|
assert!(has_external_segmented_bar_labels(&items, 1, &geometry));
|
||||||
|
assert!(is_chart_bar_cluster(&items, &raw_rects, 1));
|
||||||
|
|
||||||
|
let numeric_items: Vec<TextItem> = (0..4)
|
||||||
|
.map(|row| make_item("2024", 80.0, 541.0 + row as f32 * 18.0, 9.0))
|
||||||
|
.collect();
|
||||||
|
assert!(has_external_segmented_bar_labels(
|
||||||
|
&numeric_items,
|
||||||
|
1,
|
||||||
|
&geometry
|
||||||
|
));
|
||||||
|
assert!(is_chart_bar_cluster(&numeric_items, &raw_rects, 1));
|
||||||
|
|
||||||
|
let edge_adjacent_items: Vec<TextItem> = (0..4)
|
||||||
|
.map(|row| make_item("2024", 92.0, 541.0 + row as f32 * 18.0, 9.0))
|
||||||
|
.collect();
|
||||||
|
assert!(has_external_segmented_bar_labels(
|
||||||
|
&edge_adjacent_items,
|
||||||
|
1,
|
||||||
|
&geometry
|
||||||
|
));
|
||||||
|
assert!(is_chart_bar_cluster(&edge_adjacent_items, &raw_rects, 1));
|
||||||
|
|
||||||
|
let far_items: Vec<TextItem> = (0..4)
|
||||||
|
.map(|row| make_item("Category", 20.0, 541.0 + row as f32 * 18.0, 9.0))
|
||||||
|
.collect();
|
||||||
|
assert!(!has_external_segmented_bar_labels(&far_items, 1, &geometry));
|
||||||
|
assert!(!is_chart_bar_cluster(&far_items, &raw_rects, 1));
|
||||||
|
|
||||||
|
let rects: Vec<PdfRect> = raw_rects
|
||||||
|
.into_iter()
|
||||||
|
.map(|(x, y, width, height)| PdfRect {
|
||||||
|
x,
|
||||||
|
y,
|
||||||
|
width,
|
||||||
|
height,
|
||||||
|
page: 1,
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
assert_eq!(detect_chart_regions(&items, &rects, 1).len(), 1);
|
||||||
|
let (tables, hints) = detect_tables_from_rects(&items, &rects, 1);
|
||||||
|
assert!(tables.is_empty());
|
||||||
|
assert!(hints.is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn padded_numeric_grid_frame_remains_a_table() {
|
||||||
|
let row_edges = [
|
||||||
|
[100.0, 140.0, 180.0, 220.0, 260.0],
|
||||||
|
[100.0, 140.0, 180.0, 228.0, 260.0],
|
||||||
|
[100.0, 140.0, 180.0, 214.0, 260.0],
|
||||||
|
[100.0, 140.0, 180.0, 232.0, 260.0],
|
||||||
|
];
|
||||||
|
let mut raw_rects = vec![(96.0, 536.0, 168.0, 80.0)];
|
||||||
|
let mut items = Vec::new();
|
||||||
|
for (row, edges) in row_edges.into_iter().enumerate() {
|
||||||
|
let y = 540.0 + row as f32 * 20.0;
|
||||||
|
for edge in edges.windows(2) {
|
||||||
|
raw_rects.push((edge[0], y, edge[1] - edge[0], 12.0));
|
||||||
|
items.push(make_item("42", edge[0] + 8.0, y + 1.0, 9.0));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
assert!(is_repeated_cell_grid(&raw_rects));
|
||||||
|
assert!(has_chart_bar_signature(&items, &raw_rects, 1));
|
||||||
|
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented rows");
|
||||||
|
assert!(!has_external_segmented_bar_labels(&items, 1, &geometry));
|
||||||
|
assert!(!is_chart_bar_cluster(&items, &raw_rects, 1));
|
||||||
|
|
||||||
|
let flush_items: Vec<TextItem> = (0..4)
|
||||||
|
.map(|row| make_item("1", 100.0, 541.0 + row as f32 * 20.0, 9.0))
|
||||||
|
.collect();
|
||||||
|
assert!(!has_external_segmented_bar_labels(
|
||||||
|
&flush_items,
|
||||||
|
1,
|
||||||
|
&geometry
|
||||||
|
));
|
||||||
|
assert!(!is_chart_bar_cluster(&flush_items, &raw_rects, 1));
|
||||||
|
|
||||||
|
let rects: Vec<PdfRect> = raw_rects
|
||||||
|
.into_iter()
|
||||||
|
.map(|(x, y, width, height)| PdfRect {
|
||||||
|
x,
|
||||||
|
y,
|
||||||
|
width,
|
||||||
|
height,
|
||||||
|
page: 1,
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
assert!(detect_chart_regions(&items, &rects, 1).is_empty());
|
||||||
|
assert!(!detect_tables_from_rects(&items, &rects, 1).0.is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn frameless_segmented_chart_with_category_labels_remains_a_chart() {
|
||||||
|
let row_edges = [
|
||||||
|
[100.0, 140.0, 180.0, 220.0, 260.0],
|
||||||
|
[100.0, 140.0, 180.0, 228.0, 260.0],
|
||||||
|
[100.0, 140.0, 180.0, 214.0, 260.0],
|
||||||
|
[100.0, 140.0, 180.0, 232.0, 260.0],
|
||||||
|
];
|
||||||
|
let mut raw_rects = Vec::new();
|
||||||
|
let mut items = Vec::new();
|
||||||
|
for (row, edges) in row_edges.into_iter().enumerate() {
|
||||||
|
let y = 540.0 + row as f32 * 18.0;
|
||||||
|
raw_rects.extend(
|
||||||
|
edges
|
||||||
|
.windows(2)
|
||||||
|
.map(|edge| (edge[0], y, edge[1] - edge[0], 12.0)),
|
||||||
|
);
|
||||||
|
items.push(make_item("Category", 62.0, y + 1.0, 9.0));
|
||||||
|
}
|
||||||
|
|
||||||
|
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented stack");
|
||||||
|
assert!(has_external_segmented_bar_labels(&items, 1, &geometry));
|
||||||
|
assert!(is_chart_bar_cluster(&items, &raw_rects, 1));
|
||||||
|
|
||||||
|
let rects: Vec<PdfRect> = raw_rects
|
||||||
|
.into_iter()
|
||||||
|
.map(|(x, y, width, height)| PdfRect {
|
||||||
|
x,
|
||||||
|
y,
|
||||||
|
width,
|
||||||
|
height,
|
||||||
|
page: 1,
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
assert_eq!(detect_chart_regions(&items, &rects, 1).len(), 1);
|
||||||
|
}
|
||||||
|
|
||||||
// --- detect_stacked_box_table ---
|
// --- detect_stacked_box_table ---
|
||||||
|
|
||||||
/// N stacked boxes at x=100, w=300, h=22, top-to-bottom from y=600.
|
/// N stacked boxes at x=100, w=300, h=22, top-to-bottom from y=600.
|
||||||
|
|||||||
+3
-1
@@ -16,7 +16,9 @@ pub(crate) use detect_heuristic::{
|
|||||||
content_width, detect_tables_with_page_width, is_table_of_contents,
|
content_width, detect_tables_with_page_width, is_table_of_contents,
|
||||||
};
|
};
|
||||||
pub use detect_lines::detect_tables_from_lines;
|
pub use detect_lines::detect_tables_from_lines;
|
||||||
pub(crate) use detect_lines::detect_vector_grid_tables_from_lines;
|
pub(crate) use detect_lines::{
|
||||||
|
detect_dense_line_chart_regions, detect_vector_grid_tables_from_lines,
|
||||||
|
};
|
||||||
pub(crate) use detect_rects::cluster_rects;
|
pub(crate) use detect_rects::cluster_rects;
|
||||||
pub use detect_rects::{detect_chart_regions, detect_tables_from_rects, RectHintRegion};
|
pub use detect_rects::{detect_chart_regions, detect_tables_from_rects, RectHintRegion};
|
||||||
pub use detect_struct::detect_tables_from_struct_tree;
|
pub use detect_struct::detect_tables_from_struct_tree;
|
||||||
|
|||||||
+10
-3
@@ -70,10 +70,17 @@ pub(crate) fn is_page_number_line(text: &str) -> bool {
|
|||||||
|
|
||||||
let lowercase = text.trim().to_ascii_lowercase();
|
let lowercase = text.trim().to_ascii_lowercase();
|
||||||
lowercase.strip_prefix("page").is_some_and(|rest| {
|
lowercase.strip_prefix("page").is_some_and(|rest| {
|
||||||
rest.trim_start()
|
let mut characters = rest.trim_start().chars().peekable();
|
||||||
.chars()
|
let mut has_page_number = false;
|
||||||
.next()
|
while characters
|
||||||
|
.peek()
|
||||||
.is_some_and(|character| character.is_ascii_digit())
|
.is_some_and(|character| character.is_ascii_digit())
|
||||||
|
{
|
||||||
|
has_page_number = true;
|
||||||
|
characters.next();
|
||||||
|
}
|
||||||
|
|
||||||
|
has_page_number && characters.next().is_none_or(char::is_whitespace)
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
+140
-1
@@ -594,7 +594,7 @@ fn hex_to_unicode_string(hex: &str) -> Option<String> {
|
|||||||
|
|
||||||
let bytes: Option<Vec<u8>> = (0..hex.len())
|
let bytes: Option<Vec<u8>> = (0..hex.len())
|
||||||
.step_by(2)
|
.step_by(2)
|
||||||
.map(|i| u8::from_str_radix(&hex[i..i + 2], 16).ok())
|
.map(|i| u8::from_str_radix(hex.get(i..i + 2)?, 16).ok())
|
||||||
.collect();
|
.collect();
|
||||||
let bytes = bytes?;
|
let bytes = bytes?;
|
||||||
|
|
||||||
@@ -880,6 +880,25 @@ fn try_remap_subset_cmap(
|
|||||||
None => return (cmap, None),
|
None => return (cmap, None),
|
||||||
};
|
};
|
||||||
|
|
||||||
|
// Both repair paths below assume CIDs are glyph indices that a subsetter can
|
||||||
|
// renumber, which is only true for CIDFontType2 (TrueType). For CIDFontType0
|
||||||
|
// (CFF), CIDs are resolved through the CFF charset, so a valid CMap stays valid
|
||||||
|
// after subsetting and renumbering it corrupts otherwise-correct text.
|
||||||
|
// CIDToGIDMap is likewise CIDFontType2-only (PDF 32000-1:2008, 9.7.4.2), so this
|
||||||
|
// also ignores a CIDToGIDMap that a malformed producer attached to a CFF font.
|
||||||
|
// /Subtype may be an indirect reference, so resolve it through the document.
|
||||||
|
// Only bail out when the descendant is *explicitly* something other than
|
||||||
|
// CIDFontType2: a missing or unresolvable /Subtype keeps the previous
|
||||||
|
// behaviour rather than silently disabling the repair.
|
||||||
|
let subtype = cid_font_dict.get(b"Subtype").ok().and_then(|o| match o {
|
||||||
|
Object::Reference(r) => doc.get_object(*r).ok().and_then(|o| o.as_name().ok()),
|
||||||
|
other => other.as_name().ok(),
|
||||||
|
});
|
||||||
|
if subtype.is_some_and(|name| name != b"CIDFontType2") {
|
||||||
|
debug!("Subset remap skipped for obj={obj_num}: descendant is not CIDFontType2");
|
||||||
|
return (cmap, None);
|
||||||
|
}
|
||||||
|
|
||||||
// If there's an explicit CIDToGIDMap, build a repaired CMap using it.
|
// If there's an explicit CIDToGIDMap, build a repaired CMap using it.
|
||||||
if let Some(cid_to_gid) = get_cid_to_gid_map(cid_font_dict, doc) {
|
if let Some(cid_to_gid) = get_cid_to_gid_map(cid_font_dict, doc) {
|
||||||
if let Some(repaired) = build_cmap_with_cid_to_gid_map(&cmap, &cid_to_gid) {
|
if let Some(repaired) = build_cmap_with_cid_to_gid_map(&cmap, &cid_to_gid) {
|
||||||
@@ -2587,6 +2606,24 @@ endcmap
|
|||||||
assert_eq!(cmap.lookup(0x0025), Some("B".to_string()));
|
assert_eq!(cmap.lookup(0x0025), Some("B".to_string()));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_hex_to_unicode_non_ascii_no_panic() {
|
||||||
|
// A destination containing a multi-byte char makes the byte length even
|
||||||
|
// while a byte offset can land inside a char. Slicing must not panic;
|
||||||
|
// it should be rejected gracefully.
|
||||||
|
assert_eq!(hex_to_unicode_string("XéY"), None);
|
||||||
|
assert_eq!(hex_to_unicode_string("\u{fffd}0"), None);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_parse_bfchar_non_ascii_destination_no_panic() {
|
||||||
|
// Crafted /ToUnicode CMap: a non-hex, non-ASCII destination previously
|
||||||
|
// triggered a char-boundary panic in hex_to_unicode_string.
|
||||||
|
let cmap_content = "beginbfchar <0041> <XéY> endbfchar";
|
||||||
|
// Must not panic; the malformed entry is simply skipped.
|
||||||
|
let _ = ToUnicodeCMap::parse(cmap_content.as_bytes());
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_parse_bfchar_1byte() {
|
fn test_parse_bfchar_1byte() {
|
||||||
// This is the pattern that caused the CJK bug: codespace is <0000><FFFF>
|
// This is the pattern that caused the CJK bug: codespace is <0000><FFFF>
|
||||||
@@ -3159,4 +3196,106 @@ endbfrange
|
|||||||
"Remap must fire when CMap's CIDs are outside W array coverage"
|
"Remap must fire when CMap's CIDs are outside W array coverage"
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_try_remap_skipped_for_cid_font_type0() {
|
||||||
|
// Same W/CMap mismatch as the CIDFontType2 case above, but the descendant is
|
||||||
|
// CIDFontType0 (CFF). There CIDs are resolved through the CFF charset, so the
|
||||||
|
// ToUnicode CIDs stay valid after subsetting and must not be renumbered.
|
||||||
|
// Real-world case: Japanese Adobe-Japan1 PDFs (e.g. National Diet Library
|
||||||
|
// minutes) where remapping turned correct text into unrelated glyphs.
|
||||||
|
let cmap_content = r#"
|
||||||
|
1 begincodespacerange
|
||||||
|
<0000><FFFF>
|
||||||
|
endcodespacerange
|
||||||
|
1 beginbfrange
|
||||||
|
<0200> <0220> <0410>
|
||||||
|
endbfrange
|
||||||
|
"#;
|
||||||
|
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
|
||||||
|
|
||||||
|
let mut doc = Document::new();
|
||||||
|
|
||||||
|
// CIDToGIDMap is CIDFontType2-only, but a malformed producer can still emit
|
||||||
|
// one on a CFF font. Use a real stream (not /Identity, which is treated as
|
||||||
|
// "no map") so this also fails if the guard is moved back below the
|
||||||
|
// CIDToGIDMap branch: cid 1 -> gid 0x0200, which the CMap resolves.
|
||||||
|
let mut cid_to_gid = vec![0u8; 68];
|
||||||
|
cid_to_gid[2] = 0x02;
|
||||||
|
cid_to_gid[3] = 0x00;
|
||||||
|
let cid_to_gid_id =
|
||||||
|
doc.add_object(lopdf::Stream::new(lopdf::Dictionary::new(), cid_to_gid));
|
||||||
|
|
||||||
|
let mut cid_font = lopdf::Dictionary::new();
|
||||||
|
cid_font.set("Subtype", lopdf::Object::Name(b"CIDFontType0".to_vec()));
|
||||||
|
cid_font.set("CIDToGIDMap", lopdf::Object::Reference(cid_to_gid_id));
|
||||||
|
cid_font.set(
|
||||||
|
"W",
|
||||||
|
lopdf::Object::Array(vec![
|
||||||
|
lopdf::Object::Integer(1),
|
||||||
|
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
|
||||||
|
]),
|
||||||
|
);
|
||||||
|
let cid_font_id = doc.add_object(cid_font);
|
||||||
|
|
||||||
|
let mut font_dict = lopdf::Dictionary::new();
|
||||||
|
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
|
||||||
|
font_dict.set(
|
||||||
|
"DescendantFonts",
|
||||||
|
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
|
||||||
|
);
|
||||||
|
|
||||||
|
let (primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 789);
|
||||||
|
assert!(
|
||||||
|
remapped.is_none(),
|
||||||
|
"Remap must be skipped for CIDFontType0 (CFF) descendants, including a \
|
||||||
|
CIDToGIDMap a malformed producer attached to one"
|
||||||
|
);
|
||||||
|
// The original CMap must still resolve its own CIDs.
|
||||||
|
assert_eq!(primary.lookup(0x0200), Some("\u{0410}".to_string()));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_try_remap_resolves_indirect_subtype() {
|
||||||
|
// /Subtype may be stored as an indirect reference. A genuine CIDFontType2
|
||||||
|
// font must still get the repair, so the guard has to dereference it rather
|
||||||
|
// than treat the unresolved value as "not CIDFontType2".
|
||||||
|
let cmap_content = r#"
|
||||||
|
1 begincodespacerange
|
||||||
|
<0000><FFFF>
|
||||||
|
endcodespacerange
|
||||||
|
1 beginbfrange
|
||||||
|
<0200> <0220> <0410>
|
||||||
|
endbfrange
|
||||||
|
"#;
|
||||||
|
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
|
||||||
|
|
||||||
|
let mut doc = Document::new();
|
||||||
|
let subtype_id = doc.add_object(lopdf::Object::Name(b"CIDFontType2".to_vec()));
|
||||||
|
|
||||||
|
let mut cid_font = lopdf::Dictionary::new();
|
||||||
|
cid_font.set("Subtype", lopdf::Object::Reference(subtype_id));
|
||||||
|
cid_font.set("CIDToGIDMap", lopdf::Object::Name(b"Identity".to_vec()));
|
||||||
|
cid_font.set(
|
||||||
|
"W",
|
||||||
|
lopdf::Object::Array(vec![
|
||||||
|
lopdf::Object::Integer(0),
|
||||||
|
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
|
||||||
|
]),
|
||||||
|
);
|
||||||
|
let cid_font_id = doc.add_object(cid_font);
|
||||||
|
|
||||||
|
let mut font_dict = lopdf::Dictionary::new();
|
||||||
|
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
|
||||||
|
font_dict.set(
|
||||||
|
"DescendantFonts",
|
||||||
|
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
|
||||||
|
);
|
||||||
|
|
||||||
|
let (_primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 790);
|
||||||
|
assert!(
|
||||||
|
remapped.is_some(),
|
||||||
|
"An indirect /Subtype naming CIDFontType2 must still reach the remap"
|
||||||
|
);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
+68
@@ -0,0 +1,68 @@
|
|||||||
|
%PDF-1.3
|
||||||
|
%“Œ‹ž ReportLab Generated PDF document (opensource)
|
||||||
|
1 0 obj
|
||||||
|
<<
|
||||||
|
/F1 2 0 R
|
||||||
|
>>
|
||||||
|
endobj
|
||||||
|
2 0 obj
|
||||||
|
<<
|
||||||
|
/BaseFont /Helvetica /Encoding /WinAnsiEncoding /Name /F1 /Subtype /Type1 /Type /Font
|
||||||
|
>>
|
||||||
|
endobj
|
||||||
|
3 0 obj
|
||||||
|
<<
|
||||||
|
/Contents 7 0 R /MediaBox [ 0 0 612 792 ] /Parent 6 0 R /Resources <<
|
||||||
|
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
|
||||||
|
>> /Rotate 0 /Trans <<
|
||||||
|
|
||||||
|
>>
|
||||||
|
/Type /Page
|
||||||
|
>>
|
||||||
|
endobj
|
||||||
|
4 0 obj
|
||||||
|
<<
|
||||||
|
/PageMode /UseNone /Pages 6 0 R /Type /Catalog
|
||||||
|
>>
|
||||||
|
endobj
|
||||||
|
5 0 obj
|
||||||
|
<<
|
||||||
|
/Author (anonymous) /CreationDate (D:20260803112923+00'00') /Creator (anonymous) /Keywords () /ModDate (D:20260803112923+00'00') /Producer (ReportLab PDF Library - \(opensource\))
|
||||||
|
/Subject (unspecified) /Title (untitled) /Trapped /False
|
||||||
|
>>
|
||||||
|
endobj
|
||||||
|
6 0 obj
|
||||||
|
<<
|
||||||
|
/Count 1 /Kids [ 3 0 R ] /Type /Pages
|
||||||
|
>>
|
||||||
|
endobj
|
||||||
|
7 0 obj
|
||||||
|
<<
|
||||||
|
/Filter [ /ASCII85Decode /FlateDecode ] /Length 202
|
||||||
|
>>
|
||||||
|
stream
|
||||||
|
GarW05mr9@&;9NOME,dW.,B;'jAjYq0S4Z`*D9aMA;]$5J)A/$3lESen1?F)ZJsa4$4&N%-%cs)#qW5EVhhbPiRDrAV>MC%.spto@CU"ZdipR'TtFiMR_%m*Hm$N%qL7a"ckkp9T/s[N2"Og377mP*M^akb2XQZ@'l*qT(9bVtDb5+)S&Q.#%E)<]Ao`TSk2AE'/E\fn~>endstream
|
||||||
|
endobj
|
||||||
|
xref
|
||||||
|
0 8
|
||||||
|
0000000000 65535 f
|
||||||
|
0000000061 00000 n
|
||||||
|
0000000092 00000 n
|
||||||
|
0000000199 00000 n
|
||||||
|
0000000392 00000 n
|
||||||
|
0000000460 00000 n
|
||||||
|
0000000721 00000 n
|
||||||
|
0000000780 00000 n
|
||||||
|
trailer
|
||||||
|
<<
|
||||||
|
/ID
|
||||||
|
[<6d7ea1213c5974c78613d5d2a08423b5><6d7ea1213c5974c78613d5d2a08423b5>]
|
||||||
|
% ReportLab generated PDF document -- digest (opensource)
|
||||||
|
|
||||||
|
/Info 5 0 R
|
||||||
|
/Root 4 0 R
|
||||||
|
/Size 8
|
||||||
|
>>
|
||||||
|
startxref
|
||||||
|
9072
|
||||||
|
%%EOF
|
||||||
+79
File diff suppressed because one or more lines are too long
Binary file not shown.
+1239
File diff suppressed because it is too large
Load Diff
@@ -1429,6 +1429,77 @@ fn test_firecrawl_tagged_pdf_struct_tree() {
|
|||||||
assert_eq!(fence_count % 2, 0, "Code fences should be balanced");
|
assert_eq!(fence_count % 2, 0, "Code fences should be balanced");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_tagged_pdf_text_items_carry_mcid() {
|
||||||
|
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
|
||||||
|
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
|
||||||
|
assert!(
|
||||||
|
items.iter().any(|i| i.mcid.is_some()),
|
||||||
|
"Tagged PDF text items should carry Marked Content IDs"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_extract_structure_elements_tagged_pdf() {
|
||||||
|
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
|
||||||
|
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
|
||||||
|
assert!(!elements.is_empty(), "Tagged PDF should yield elements");
|
||||||
|
assert!(
|
||||||
|
elements.iter().any(|e| e.role == "H1"),
|
||||||
|
"Should surface H1 heading roles"
|
||||||
|
);
|
||||||
|
assert!(
|
||||||
|
elements.iter().all(|e| !e.role.is_empty()),
|
||||||
|
"Every element should carry a role name"
|
||||||
|
);
|
||||||
|
|
||||||
|
// Sorted by (page, mcid) for deterministic output
|
||||||
|
assert!(
|
||||||
|
elements
|
||||||
|
.windows(2)
|
||||||
|
.all(|w| (w[0].page, w[0].mcid) <= (w[1].page, w[1].mcid)),
|
||||||
|
"Elements should be sorted by (page, mcid)"
|
||||||
|
);
|
||||||
|
|
||||||
|
// The advertised join: (page, mcid) pairs must line up with the
|
||||||
|
// mcid-carrying TextItems from positioned extraction, and joining the
|
||||||
|
// H1 entries must recover non-empty heading text.
|
||||||
|
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
|
||||||
|
let h1_refs: std::collections::HashSet<(u32, i64)> = elements
|
||||||
|
.iter()
|
||||||
|
.filter(|e| e.role == "H1")
|
||||||
|
.map(|e| (e.page, e.mcid))
|
||||||
|
.collect();
|
||||||
|
let h1_text: String = items
|
||||||
|
.iter()
|
||||||
|
.filter(|i| i.mcid.is_some_and(|mcid| h1_refs.contains(&(i.page, mcid))))
|
||||||
|
.map(|i| i.text.as_str())
|
||||||
|
.collect();
|
||||||
|
assert!(
|
||||||
|
!h1_text.trim().is_empty(),
|
||||||
|
"Joining H1 structure elements to text items should recover heading text"
|
||||||
|
);
|
||||||
|
|
||||||
|
// Page filter is 1-indexed (matching TextItem.page) and equals the
|
||||||
|
// corresponding subset of the full document result.
|
||||||
|
let page1 = pdf_inspector::extract_structure_elements_mem(&buf, Some(&[1])).unwrap();
|
||||||
|
assert!(!page1.is_empty(), "Page 1 should have elements");
|
||||||
|
assert!(page1.iter().all(|e| e.page == 1));
|
||||||
|
let full_page1_count = elements.iter().filter(|e| e.page == 1).count();
|
||||||
|
assert_eq!(page1.len(), full_page1_count);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn test_extract_structure_elements_untagged_pdf_empty() {
|
||||||
|
let buf = std::fs::read("tests/fixtures/thermo-freon12.pdf").unwrap();
|
||||||
|
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
|
||||||
|
assert!(
|
||||||
|
elements.is_empty(),
|
||||||
|
"Untagged PDF should yield no structure elements, got {:?}",
|
||||||
|
elements
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_identity_h_no_tounicode_suppresses_garbage() {
|
fn test_identity_h_no_tounicode_suppresses_garbage() {
|
||||||
// shinagawa_identity_h.pdf uses YuGothic with Identity-H encoding and no
|
// shinagawa_identity_h.pdf uses YuGothic with Identity-H encoding and no
|
||||||
@@ -3926,6 +3997,65 @@ fn encrypted_pdf_decrypts_with_correct_password() {
|
|||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Regression for the #231 review finding: `extract_pages_markdown`'s
|
||||||
|
/// `has_template_image` check must be gated the same way
|
||||||
|
/// `classify_pdf`/`detect_pdf_type` gates it (image_count <= 1, few text
|
||||||
|
/// ops, low alphanumeric diversity) — not treated as sufficient on its
|
||||||
|
/// own. The fixture is a real text page with substantial, richly varied
|
||||||
|
/// body text (>=50 Tj ops) drawn over a full-bleed background image
|
||||||
|
/// (e.g. letterhead/watermark). Before the fix, has_template_image alone
|
||||||
|
/// forced needs_ocr=true and discarded the page's clean markdown; now the
|
||||||
|
/// page must extract normally.
|
||||||
|
#[test]
|
||||||
|
fn test_extract_pages_markdown_does_not_ocr_text_page_with_watermark_image() {
|
||||||
|
let buf = std::fs::read("tests/fixtures/text_page_with_watermark_image.pdf").unwrap();
|
||||||
|
|
||||||
|
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
|
||||||
|
let page = &ext.pages[0];
|
||||||
|
assert!(
|
||||||
|
!page.needs_ocr,
|
||||||
|
"a text page with substantial real text should not be routed to OCR \
|
||||||
|
just because it has a background image"
|
||||||
|
);
|
||||||
|
assert!(
|
||||||
|
page.markdown.contains("watermark"),
|
||||||
|
"expected the page's real body text to be preserved, got: {:?}",
|
||||||
|
page.markdown
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Regression for the #231 review finding: `extract_pages_markdown` never
|
||||||
|
/// checked `has_vector_text` at all, even though `detect_from_document`'s
|
||||||
|
/// Mixed-type per-page routing always sends vector-outlined-text pages to
|
||||||
|
/// OCR (outlined glyphs can't be extracted as text). A page with massive
|
||||||
|
/// path ops (outlined decorative text) plus a short genuine caption would
|
||||||
|
/// extract that caption cleanly — non-empty, non-garbled — so the
|
||||||
|
/// existing empty/garbage-text checks alone couldn't catch it.
|
||||||
|
#[test]
|
||||||
|
fn test_extract_pages_markdown_ocrs_page_with_vector_outlined_text() {
|
||||||
|
let buf = std::fs::read("tests/fixtures/vector_outlined_text_with_caption.pdf").unwrap();
|
||||||
|
|
||||||
|
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
|
||||||
|
assert!(
|
||||||
|
cls.pages_needing_ocr.contains(&1),
|
||||||
|
"classify_pdf should flag page 1 as needing OCR (vector-outlined text), got: {:?}",
|
||||||
|
cls.pages_needing_ocr
|
||||||
|
);
|
||||||
|
|
||||||
|
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
|
||||||
|
let page = &ext.pages[0];
|
||||||
|
assert!(
|
||||||
|
page.needs_ocr,
|
||||||
|
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
|
||||||
|
);
|
||||||
|
assert!(
|
||||||
|
page.markdown.is_empty(),
|
||||||
|
"a page flagged needs_ocr must not return markdown as if extraction were \
|
||||||
|
trustworthy, got: {:?}",
|
||||||
|
page.markdown
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn pdf_options_debug_redacts_password() {
|
fn pdf_options_debug_redacts_password() {
|
||||||
let opts = PdfOptions::new().password("secret123");
|
let opts = PdfOptions::new().password("secret123");
|
||||||
@@ -3936,3 +4066,60 @@ fn pdf_options_debug_redacts_password() {
|
|||||||
);
|
);
|
||||||
assert!(dbg.contains("REDACTED"), "expected redaction marker: {dbg}");
|
assert!(dbg.contains("REDACTED"), "expected redaction marker: {dbg}");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Regression for #228: a `startxref` pointer corrupted to point at the
|
||||||
|
/// wrong byte offset (a single flipped digit — a real, common writer bug)
|
||||||
|
/// must not make the whole file unprocessable. The real classic xref table
|
||||||
|
/// is still present and findable by scanning for the `xref` keyword; both
|
||||||
|
/// pypdf and pdfium recover the same way. Before this fix, every entry
|
||||||
|
/// point raised "Invalid PDF structure" on a file whose object data was
|
||||||
|
/// otherwise completely intact.
|
||||||
|
#[test]
|
||||||
|
fn test_process_pdf_recovers_corrupted_startxref_pointer() {
|
||||||
|
let result = process_pdf_with_options(
|
||||||
|
"tests/fixtures/broken_startxref_pointer.pdf",
|
||||||
|
PdfOptions::new(),
|
||||||
|
)
|
||||||
|
.expect("a corrupted startxref pointer should be recoverable, like pypdf/pdfium");
|
||||||
|
|
||||||
|
assert_eq!(result.page_count, 1);
|
||||||
|
let md = result.markdown.unwrap_or_default();
|
||||||
|
assert!(
|
||||||
|
md.contains("Order Detail Report by Account") && md.contains("WIDGET ASSEMBLY"),
|
||||||
|
"recovered document should extract its real text, got: {md:?}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Regression for #227: `extract_pages_markdown`'s per-page `needs_ocr`
|
||||||
|
/// must agree with `classify_pdf`/`detect_pdf_type` on the same page. The
|
||||||
|
/// fixture is a full-page raster "scan" with a single line of genuine
|
||||||
|
/// native text drawn over it (a header) — the native text extracts
|
||||||
|
/// perfectly cleanly (no decoding issues, non-empty), so a needs_ocr
|
||||||
|
/// computation based on text-quality signals alone says `false`, while
|
||||||
|
/// detection correctly sees a dominant background image and says the page
|
||||||
|
/// needs OCR. Both must now agree it needs OCR, and the markdown must not
|
||||||
|
/// be returned as if the extraction were trustworthy.
|
||||||
|
#[test]
|
||||||
|
fn test_extract_pages_markdown_agrees_with_classify_on_scan_with_native_header() {
|
||||||
|
let buf = std::fs::read("tests/fixtures/scan_with_native_header_text.pdf").unwrap();
|
||||||
|
|
||||||
|
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
|
||||||
|
assert!(
|
||||||
|
cls.pages_needing_ocr.contains(&1),
|
||||||
|
"classify_pdf should flag page 1 as needing OCR (image-dominated), got: {:?}",
|
||||||
|
cls.pages_needing_ocr
|
||||||
|
);
|
||||||
|
|
||||||
|
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
|
||||||
|
let page = &ext.pages[0];
|
||||||
|
assert!(
|
||||||
|
page.needs_ocr,
|
||||||
|
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
|
||||||
|
);
|
||||||
|
assert!(
|
||||||
|
page.markdown.is_empty(),
|
||||||
|
"a page flagged needs_ocr must not return markdown as if extraction were \
|
||||||
|
trustworthy, got: {:?}",
|
||||||
|
page.markdown
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|||||||
@@ -203,6 +203,79 @@ class TestExtractTextWithPositions:
|
|||||||
assert len(items) > 0
|
assert len(items) > 0
|
||||||
assert all(item.page == 1 for item in items)
|
assert all(item.page == 1 for item in items)
|
||||||
|
|
||||||
|
def test_mcid(self):
|
||||||
|
# Untagged fixture: mcid is None or int, never anything else
|
||||||
|
items = pdf_inspector.extract_text_with_positions(
|
||||||
|
fixture_path("thermo-freon12.pdf")
|
||||||
|
)
|
||||||
|
assert all(item.mcid is None or isinstance(item.mcid, int) for item in items)
|
||||||
|
# Tagged fixture: marked content carries MCIDs
|
||||||
|
tagged = pdf_inspector.extract_text_with_positions(
|
||||||
|
fixture_path("firecrawl_docs_tagged.pdf")
|
||||||
|
)
|
||||||
|
assert any(item.mcid is not None for item in tagged)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# extract_structure_elements / extract_structure_elements_bytes
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
class TestExtractStructureElements:
|
||||||
|
def test_tagged_file(self):
|
||||||
|
elements = pdf_inspector.extract_structure_elements(
|
||||||
|
fixture_path("firecrawl_docs_tagged.pdf")
|
||||||
|
)
|
||||||
|
assert len(elements) > 0
|
||||||
|
assert all(isinstance(e.page, int) for e in elements)
|
||||||
|
assert all(isinstance(e.mcid, int) for e in elements)
|
||||||
|
assert all(isinstance(e.role, str) and len(e.role) > 0 for e in elements)
|
||||||
|
assert any(e.role == "H1" for e in elements)
|
||||||
|
|
||||||
|
def test_join_with_text_items(self):
|
||||||
|
# (page, mcid) joins against extract_text_with_positions to recover
|
||||||
|
# heading text
|
||||||
|
path = fixture_path("firecrawl_docs_tagged.pdf")
|
||||||
|
elements = pdf_inspector.extract_structure_elements(path)
|
||||||
|
items = pdf_inspector.extract_text_with_positions(path)
|
||||||
|
h1_refs = {(e.page, e.mcid) for e in elements if e.role == "H1"}
|
||||||
|
h1_text = "".join(
|
||||||
|
item.text
|
||||||
|
for item in items
|
||||||
|
if item.mcid is not None and (item.page, item.mcid) in h1_refs
|
||||||
|
)
|
||||||
|
assert len(h1_text.strip()) > 0
|
||||||
|
|
||||||
|
def test_with_pages(self):
|
||||||
|
# pages filter is 1-indexed, matching TextItem.page
|
||||||
|
elements = pdf_inspector.extract_structure_elements(
|
||||||
|
fixture_path("firecrawl_docs_tagged.pdf"), pages=[1]
|
||||||
|
)
|
||||||
|
assert len(elements) > 0
|
||||||
|
assert all(e.page == 1 for e in elements)
|
||||||
|
|
||||||
|
def test_bytes(self):
|
||||||
|
data = fixture_bytes("firecrawl_docs_tagged.pdf")
|
||||||
|
elements = pdf_inspector.extract_structure_elements_bytes(data)
|
||||||
|
assert len(elements) > 0
|
||||||
|
assert any(e.role == "H1" for e in elements)
|
||||||
|
|
||||||
|
def test_untagged_returns_empty(self):
|
||||||
|
elements = pdf_inspector.extract_structure_elements(
|
||||||
|
fixture_path("thermo-freon12.pdf")
|
||||||
|
)
|
||||||
|
assert elements == []
|
||||||
|
|
||||||
|
def test_repr(self):
|
||||||
|
elements = pdf_inspector.extract_structure_elements(
|
||||||
|
fixture_path("firecrawl_docs_tagged.pdf")
|
||||||
|
)
|
||||||
|
assert "StructureElement" in repr(elements[0])
|
||||||
|
|
||||||
|
def test_not_a_pdf(self):
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
pdf_inspector.extract_structure_elements_bytes(b"not a pdf")
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# extract_text_in_regions / extract_text_in_regions_bytes
|
# extract_text_in_regions / extract_text_in_regions_bytes
|
||||||
|
|||||||
Generated
+2
-2
@@ -724,7 +724,7 @@ checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
|
|||||||
|
|
||||||
[[package]]
|
[[package]]
|
||||||
name = "pdf-inspector"
|
name = "pdf-inspector"
|
||||||
version = "0.1.7"
|
version = "1.14.0"
|
||||||
dependencies = [
|
dependencies = [
|
||||||
"env_logger",
|
"env_logger",
|
||||||
"include_dir",
|
"include_dir",
|
||||||
@@ -740,7 +740,7 @@ dependencies = [
|
|||||||
|
|
||||||
[[package]]
|
[[package]]
|
||||||
name = "pdf-inspector-wasm"
|
name = "pdf-inspector-wasm"
|
||||||
version = "0.1.3"
|
version = "1.14.0"
|
||||||
dependencies = [
|
dependencies = [
|
||||||
"console_error_panic_hook",
|
"console_error_panic_hook",
|
||||||
"js-sys",
|
"js-sys",
|
||||||
|
|||||||
+1
-1
@@ -1,6 +1,6 @@
|
|||||||
[package]
|
[package]
|
||||||
name = "pdf-inspector-wasm"
|
name = "pdf-inspector-wasm"
|
||||||
version = "0.1.3"
|
version = "1.14.0"
|
||||||
edition = "2021"
|
edition = "2021"
|
||||||
authors = ["Firecrawl Team"]
|
authors = ["Firecrawl Team"]
|
||||||
description = "Browser WebAssembly bindings for pdf-inspector"
|
description = "Browser WebAssembly bindings for pdf-inspector"
|
||||||
|
|||||||
Reference in New Issue
Block a user