Compare commits
3
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
0b0ca64ffb | ||
|
|
7054d6aa69 | ||
|
|
a67ee03269 |
@@ -24,6 +24,9 @@ jobs:
|
||||
- name: Cache cargo
|
||||
uses: Swatinem/rust-cache@c19371144df3bb44fab255c43d04cbc2ab54d1c4 # v2.9.1
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Run tests
|
||||
run: cargo test --verbose
|
||||
|
||||
|
||||
@@ -31,6 +31,9 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Check if version changed
|
||||
id: check
|
||||
run: |
|
||||
|
||||
@@ -28,6 +28,9 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Check if version changed
|
||||
id: check
|
||||
run: |
|
||||
|
||||
@@ -28,6 +28,9 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Check package version
|
||||
id: check
|
||||
run: |
|
||||
|
||||
@@ -28,6 +28,9 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check package version sync
|
||||
run: python3 scripts/version.py --check
|
||||
|
||||
- name: Check if version changed
|
||||
id: check
|
||||
run: |
|
||||
|
||||
+1
-2
@@ -31,9 +31,8 @@ Thumbs.db
|
||||
napi/index.js
|
||||
napi/index.d.ts
|
||||
|
||||
# Local samples and scripts
|
||||
# Local samples
|
||||
samples/
|
||||
scripts/
|
||||
|
||||
# Test output
|
||||
test_output/
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "pdf-inspector"
|
||||
version = "0.1.8"
|
||||
version = "1.14.0"
|
||||
edition = "2021"
|
||||
autobins = false
|
||||
authors = ["Firecrawl Team"]
|
||||
|
||||
@@ -110,7 +110,7 @@ Or add it manually:
|
||||
|
||||
```toml
|
||||
[dependencies]
|
||||
pdf-inspector = "0.1"
|
||||
pdf-inspector = "1"
|
||||
```
|
||||
|
||||
```rust
|
||||
|
||||
+43
-24
@@ -1,38 +1,57 @@
|
||||
# Publishing
|
||||
|
||||
The Rust crate is published to [crates.io](https://crates.io/crates/pdf-inspector) with trusted publishing from GitHub Actions. The first release was published manually; future releases publish from `.github/workflows/publish-crate.yml` when a `Cargo.toml` version change lands on `main`.
|
||||
Every pdf-inspector distribution uses one shared semantic version:
|
||||
|
||||
## crates.io Trusted Publisher
|
||||
- Rust crate: `pdf-inspector`
|
||||
- Python package: `pdf-inspector`
|
||||
- Node package: `@firecrawl/pdf-inspector` and its platform packages
|
||||
- Browser package: `@firecrawl/pdf-inspector-wasm`
|
||||
- Internal NAPI and WASM Rust crates
|
||||
|
||||
Configure the trusted publisher for the `pdf-inspector` crate with:
|
||||
`Cargo.toml` is the canonical version source. Update every manifest and lockfile
|
||||
with:
|
||||
|
||||
- Repository: `firecrawl/pdf-inspector`
|
||||
- Workflow: `publish-crate.yml`
|
||||
- Environment: `crates-io`
|
||||
```bash
|
||||
python3 scripts/version.py <version>
|
||||
```
|
||||
|
||||
The workflow uses `rust-lang/crates-io-auth-action@v1` to exchange GitHub's OIDC token for a short-lived crates.io token, then passes it to `cargo publish`.
|
||||
Verify that nothing has diverged with:
|
||||
|
||||
## Release Steps
|
||||
```bash
|
||||
python3 scripts/version.py --check
|
||||
```
|
||||
|
||||
1. Update `version` in `Cargo.toml`.
|
||||
2. Merge the version bump to `main`.
|
||||
3. The publish workflow compares the new `Cargo.toml` version with `HEAD~1`, runs `cargo publish --dry-run`, then publishes if that version is not already on crates.io.
|
||||
CI and every publishing workflow run this check before building or publishing.
|
||||
|
||||
If `Cargo.toml` changes without a package version bump, the workflow exits without publishing.
|
||||
## Release steps
|
||||
|
||||
## Browser WebAssembly package
|
||||
1. Choose the next shared semantic version and run `scripts/version.py`.
|
||||
2. Review the manifest and lockfile changes in the version-bump pull request.
|
||||
3. Merge the pull request to `main`.
|
||||
4. The crates.io, PyPI, Node, and WASM workflows independently build and
|
||||
publish that version from the same commit.
|
||||
5. After all registries succeed, create one `v<version>` GitHub release that
|
||||
links to each package and describes changes since the previous shared tag.
|
||||
|
||||
The browser package is published as `@firecrawl/pdf-inspector-wasm`. Its version lives in `wasm/Cargo.toml`, and `.github/workflows/publish-wasm.yml` builds the `web` target with `wasm-pack` before publishing the generated package.
|
||||
The independent workflows are intentionally idempotent. A manual dispatch from
|
||||
`main` can repair a partial release, and already-published artifacts are skipped.
|
||||
|
||||
The npm package must exist before a trusted publisher can be configured. For the first release only:
|
||||
## Trusted publishers
|
||||
|
||||
1. Build with `wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release`.
|
||||
2. Inspect with `npm pack --dry-run ./wasm/pkg`.
|
||||
3. Publish with `npm publish ./wasm/pkg --access public` from an authorized maintainer session.
|
||||
4. In the package settings on npm, configure the GitHub Actions trusted publisher:
|
||||
- Organization: `firecrawl`
|
||||
- Repository: `pdf-inspector`
|
||||
- Workflow: `publish-wasm.yml`
|
||||
- Allowed action: `npm publish`
|
||||
The repositories use GitHub Actions OIDC instead of long-lived registry tokens.
|
||||
Configure each registry's trusted publisher for `firecrawl/pdf-inspector` and
|
||||
its corresponding workflow:
|
||||
|
||||
After that one-time bootstrap, bumping the version in `wasm/Cargo.toml` and merging it to `main` publishes through OIDC. Until the package exists, the workflow exits cleanly without attempting an unauthenticated first publish. See npm's [trusted publishing documentation](https://docs.npmjs.com/trusted-publishers/) for the registry-side setup.
|
||||
- crates.io: `publish-crate.yml`, environment `crates-io`
|
||||
- PyPI: `publish-pypi.yml`, environment `pypi`
|
||||
- npm Node package: `publish.yml`
|
||||
- npm WASM package: `publish-wasm.yml`
|
||||
|
||||
The WASM package must exist before npm trusted publishing can be configured. If
|
||||
it ever needs to be bootstrapped again, build and inspect it before publishing:
|
||||
|
||||
```bash
|
||||
wasm-pack build wasm --target web --scope firecrawl --out-dir pkg --release
|
||||
npm pack --dry-run ./wasm/pkg
|
||||
npm publish ./wasm/pkg --access public
|
||||
```
|
||||
|
||||
@@ -80,6 +80,17 @@ for page in result.pages:
|
||||
|
||||
# Restrict to specific 0-indexed pages (preserves caller order)
|
||||
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
|
||||
|
||||
# Structure-tree elements from tagged PDFs (empty list when untagged).
|
||||
# Pages are 1-indexed to match TextItem.page, so (page, mcid) joins directly
|
||||
# against extract_text_with_positions — e.g. to recover real heading levels:
|
||||
elements = pdf_inspector.extract_structure_elements("tagged.pdf")
|
||||
roles = {(e.page, e.mcid): e.role for e in elements}
|
||||
headings = [
|
||||
item.text
|
||||
for item in pdf_inspector.extract_text_with_positions("tagged.pdf")
|
||||
if item.mcid is not None and roles.get((item.page, item.mcid), "").startswith("H")
|
||||
]
|
||||
```
|
||||
|
||||
## API reference
|
||||
@@ -100,6 +111,8 @@ result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
|
||||
| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes |
|
||||
| `extract_pages_markdown(path, pages=None)` | Per-page Markdown + layout metadata (all pages by default) |
|
||||
| `extract_pages_markdown_bytes(data, pages=None)` | Per-page Markdown from bytes |
|
||||
| `extract_structure_elements(path, pages=None)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
|
||||
| `extract_structure_elements_bytes(data, pages=None)` | Structure-tree elements from bytes |
|
||||
|
||||
## Types
|
||||
|
||||
@@ -144,6 +157,12 @@ class TextItem: # extract_text_with_positions
|
||||
is_underline: bool
|
||||
is_strikeout: bool
|
||||
item_type: str
|
||||
mcid: int | None # marked-content ID for tagged PDFs (None otherwise)
|
||||
|
||||
class StructureElement: # extract_structure_elements
|
||||
page: int # 1-indexed (matches TextItem.page)
|
||||
mcid: int
|
||||
role: str # "H1".."H6", "P", "Table", ... (resolved via /RoleMap)
|
||||
|
||||
class RegionText: # extract_text_in_regions
|
||||
text: str
|
||||
|
||||
+32
-1
@@ -138,6 +138,34 @@ for page in &result.pages {
|
||||
println!("Complex layout? {}", result.is_complex);
|
||||
```
|
||||
|
||||
Extract structure-tree elements from tagged PDFs, and join them against
|
||||
`extract_text_with_positions` to attach semantic roles (heading levels,
|
||||
paragraphs, table cells) to extracted text:
|
||||
|
||||
```rust
|
||||
use pdf_inspector::{extract_structure_elements, extract_text_with_positions};
|
||||
use std::collections::HashMap;
|
||||
|
||||
// One entry per marked-content reference, sorted by (page, mcid); empty for
|
||||
// untagged PDFs. Pages are 1-indexed to match `TextItem::page`, so the
|
||||
// (page, mcid) pair is a direct join key.
|
||||
let elements = extract_structure_elements("tagged.pdf", None)?;
|
||||
let roles: HashMap<(u32, i64), &str> = elements
|
||||
.iter()
|
||||
.map(|e| ((e.page, e.mcid), e.role.as_str()))
|
||||
.collect();
|
||||
|
||||
for item in extract_text_with_positions("tagged.pdf")? {
|
||||
if let Some(mcid) = item.mcid {
|
||||
if let Some(role) = roles.get(&(item.page, mcid)) {
|
||||
if role.starts_with('H') {
|
||||
println!("{}: {}", role, item.text);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Processing modes
|
||||
|
||||
| Mode | What it does | Returns |
|
||||
@@ -163,6 +191,8 @@ println!("Complex layout? {}", result.is_complex);
|
||||
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
|
||||
| `extract_pages_markdown(path, pages)` | Per-page Markdown + layout metadata (file) |
|
||||
| `extract_pages_markdown_mem(bytes, pages)` | Per-page Markdown from bytes |
|
||||
| `extract_structure_elements(path, pages)` | Structure-tree elements from tagged PDFs (page, mcid, role) |
|
||||
| `extract_structure_elements_mem(bytes, pages)` | Structure-tree elements from bytes |
|
||||
|
||||
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
|
||||
|
||||
@@ -178,7 +208,8 @@ Low-level detection functions are also available via the `detector` module (`det
|
||||
| `DetectionConfig` | Configuration for detection: scan strategy, thresholds |
|
||||
| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` |
|
||||
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
|
||||
| `TextItem` | Text with position, font info, and page number |
|
||||
| `TextItem` | Text with position, font info, page number, and optional structure-tree `mcid` |
|
||||
| `StructureElement` | Tagged-PDF structure reference: page (1-indexed), mcid, role (`"H1"`..`"H6"`, `"P"`, …) |
|
||||
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
|
||||
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
|
||||
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
|
||||
|
||||
Generated
+2
-2
@@ -851,7 +851,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector"
|
||||
version = "0.1.8"
|
||||
version = "1.14.0"
|
||||
dependencies = [
|
||||
"env_logger",
|
||||
"include_dir",
|
||||
@@ -867,7 +867,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector-napi"
|
||||
version = "0.2.3"
|
||||
version = "1.14.0"
|
||||
dependencies = [
|
||||
"napi",
|
||||
"napi-build",
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "pdf-inspector-napi"
|
||||
version = "0.2.3"
|
||||
version = "1.14.0"
|
||||
edition = "2021"
|
||||
|
||||
[lib]
|
||||
|
||||
+6
-6
@@ -8,12 +8,12 @@
|
||||
"@napi-rs/cli": "^3.4.1",
|
||||
},
|
||||
"optionalDependencies": {
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0",
|
||||
},
|
||||
},
|
||||
},
|
||||
|
||||
+7
-7
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "@firecrawl/pdf-inspector",
|
||||
"version": "1.13.0",
|
||||
"version": "1.14.0",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
@@ -52,11 +52,11 @@
|
||||
"@napi-rs/cli": "^3.4.1"
|
||||
},
|
||||
"optionalDependencies": {
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.13.0",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.13.0"
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0"
|
||||
}
|
||||
}
|
||||
|
||||
@@ -89,6 +89,11 @@ pub struct TextItem {
|
||||
pub item_type: ItemType,
|
||||
/// URL for link items, `None` for other types.
|
||||
pub link_url: Option<String>,
|
||||
/// Marked Content ID from the content stream's BDC/BMC operator, `None`
|
||||
/// when the text is not part of marked content. Join with the
|
||||
/// `page`/`mcid` pairs from [`extractStructureElements`] to attach
|
||||
/// structure-tree roles (headings, paragraphs, …) in tagged PDFs.
|
||||
pub mcid: Option<i64>,
|
||||
}
|
||||
|
||||
/// A page's regions for text extraction: (page_index_0based, bboxes).
|
||||
@@ -306,12 +311,61 @@ pub fn extract_text_with_positions(
|
||||
is_strikeout: item.is_strikeout,
|
||||
item_type,
|
||||
link_url,
|
||||
mcid: item.mcid,
|
||||
}
|
||||
})
|
||||
.collect())
|
||||
})
|
||||
}
|
||||
|
||||
/// One structure-tree element reference from a tagged PDF.
|
||||
#[napi(object)]
|
||||
pub struct StructureElementJs {
|
||||
/// 1-indexed page number (matches `TextItem.page`).
|
||||
pub page: u32,
|
||||
/// Marked Content ID from the page's content stream (matches
|
||||
/// `TextItem.mcid`).
|
||||
pub mcid: i64,
|
||||
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
|
||||
/// Custom tags are resolved through the document's role map; tags with
|
||||
/// no standard mapping are returned verbatim.
|
||||
pub role: String,
|
||||
}
|
||||
|
||||
/// Extract structure-tree element references from a tagged PDF.
|
||||
///
|
||||
/// Parses the document's structure tree (when present) and returns one
|
||||
/// entry per marked-content reference, resolved to its 1-indexed page,
|
||||
/// MCID, and structure type name. Returns an empty array when the PDF is
|
||||
/// not tagged.
|
||||
///
|
||||
/// Join `(page, mcid)` against the `page`/`mcid` fields from
|
||||
/// [`extractTextWithPositions`] to attach heading levels (H1..H6) and other
|
||||
/// semantic roles to extracted text.
|
||||
///
|
||||
/// Pass 1-indexed page numbers (matching `TextItem.page`) to restrict
|
||||
/// output; omit `pages` for the whole document. Entries are sorted by
|
||||
/// `(page, mcid)`.
|
||||
#[napi]
|
||||
pub fn extract_structure_elements(
|
||||
buffer: Buffer,
|
||||
pages: Option<Vec<u32>>,
|
||||
) -> Result<Vec<StructureElementJs>> {
|
||||
let bytes: Vec<u8> = buffer.to_vec();
|
||||
catch_panic("extract_structure_elements", move || {
|
||||
let elements = pdf_inspector::extract_structure_elements_mem(&bytes, pages.as_deref())
|
||||
.map_err(|e| to_napi_err(e, "extract_structure_elements"))?;
|
||||
Ok(elements
|
||||
.into_iter()
|
||||
.map(|e| StructureElementJs {
|
||||
page: e.page,
|
||||
mcid: e.mcid,
|
||||
role: e.role,
|
||||
})
|
||||
.collect())
|
||||
})
|
||||
}
|
||||
|
||||
/// Extract text within bounding-box regions from a PDF.
|
||||
///
|
||||
/// For hybrid OCR: layout model detects regions in rendered images,
|
||||
|
||||
@@ -8,6 +8,7 @@ import {
|
||||
classifyPdfAsync,
|
||||
extractText,
|
||||
extractTextWithPositions,
|
||||
extractStructureElements,
|
||||
extractTextInRegions,
|
||||
detectVectorGridInRegion,
|
||||
extractPagesMarkdown,
|
||||
@@ -15,6 +16,7 @@ import {
|
||||
} from './index.js';
|
||||
|
||||
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
|
||||
const taggedFixture = readFileSync('../tests/fixtures/firecrawl_docs_tagged.pdf');
|
||||
|
||||
// --- processPdf ---
|
||||
console.log('Testing processPdf...');
|
||||
@@ -82,6 +84,46 @@ assert.ok(page1Items.length > 0);
|
||||
assert.ok(page1Items.every(i => i.page === 1));
|
||||
console.log(' extractTextWithPositions with pages: OK');
|
||||
|
||||
// mcid: undefined on untagged PDFs, numeric on tagged marked content
|
||||
assert.ok(items.every(i => i.mcid === undefined || typeof i.mcid === 'number'));
|
||||
const taggedItems = extractTextWithPositions(taggedFixture);
|
||||
assert.ok(
|
||||
taggedItems.some(i => typeof i.mcid === 'number'),
|
||||
'tagged PDF text items should carry Marked Content IDs',
|
||||
);
|
||||
console.log(' extractTextWithPositions mcid: OK');
|
||||
|
||||
// --- extractStructureElements ---
|
||||
console.log('Testing extractStructureElements...');
|
||||
const structureElements = extractStructureElements(taggedFixture);
|
||||
assert.ok(structureElements.length > 0);
|
||||
assert.ok(structureElements.every(e => typeof e.page === 'number'));
|
||||
assert.ok(structureElements.every(e => typeof e.mcid === 'number'));
|
||||
assert.ok(structureElements.every(e => typeof e.role === 'string' && e.role.length > 0));
|
||||
assert.ok(
|
||||
structureElements.some(e => e.role === 'H1'),
|
||||
'tagged fixture should surface H1 heading roles',
|
||||
);
|
||||
|
||||
// (page, mcid) joins against extractTextWithPositions to recover heading text
|
||||
const h1Refs = new Set(
|
||||
structureElements.filter(e => e.role === 'H1').map(e => `${e.page}:${e.mcid}`),
|
||||
);
|
||||
const h1Text = taggedItems
|
||||
.filter(i => typeof i.mcid === 'number' && h1Refs.has(`${i.page}:${i.mcid}`))
|
||||
.map(i => i.text)
|
||||
.join('');
|
||||
assert.ok(h1Text.trim().length > 0, 'H1 join should recover heading text');
|
||||
|
||||
// pages filter is 1-indexed, matching TextItem.page
|
||||
const page1Elements = extractStructureElements(taggedFixture, [1]);
|
||||
assert.ok(page1Elements.length > 0);
|
||||
assert.ok(page1Elements.every(e => e.page === 1));
|
||||
|
||||
// untagged PDFs yield an empty array
|
||||
assert.deepEqual(extractStructureElements(fixture), []);
|
||||
console.log(' extractStructureElements: OK');
|
||||
|
||||
// --- extractTextInRegions ---
|
||||
console.log('Testing extractTextInRegions...');
|
||||
const regionResults = extractTextInRegions(fixture, [
|
||||
|
||||
@@ -51,6 +51,20 @@ class TextItem:
|
||||
is_underline: bool
|
||||
is_strikeout: bool
|
||||
item_type: str
|
||||
mcid: Optional[int]
|
||||
"""Marked Content ID from the content stream's BDC/BMC operator, None when
|
||||
the text is not part of marked content. Join with the (page, mcid) pairs
|
||||
from extract_structure_elements to attach structure-tree roles in tagged
|
||||
PDFs."""
|
||||
|
||||
class StructureElement:
|
||||
"""One structure-tree element reference from a tagged PDF."""
|
||||
page: int
|
||||
"""1-indexed page number (matches TextItem.page)."""
|
||||
mcid: int
|
||||
"""Marked Content ID from the page's content stream (matches TextItem.mcid)."""
|
||||
role: str
|
||||
"""Standard structure type name ("H1".."H6", "P", "Table", "TD", ...)."""
|
||||
|
||||
class RegionText:
|
||||
"""Extracted text for a single region."""
|
||||
@@ -132,6 +146,27 @@ def extract_text_with_positions_bytes(data: bytes, pages: Optional[list[int]] =
|
||||
"""Extract text with position information from bytes."""
|
||||
...
|
||||
|
||||
def extract_structure_elements(path: str, pages: Optional[list[int]] = None) -> list[StructureElement]:
|
||||
"""Extract structure-tree element references from a tagged PDF file.
|
||||
|
||||
Returns one entry per marked-content reference, resolved to its 1-indexed
|
||||
page, MCID, and structure type name ("H1".."H6", "P", "Table", ...), sorted
|
||||
by (page, mcid). Returns an empty list when the PDF is not tagged.
|
||||
|
||||
Args:
|
||||
path: Path to the PDF file.
|
||||
pages: Optional list of 1-indexed pages (matching ``TextItem.page``).
|
||||
When ``None`` (default), the whole document is returned.
|
||||
"""
|
||||
...
|
||||
|
||||
def extract_structure_elements_bytes(data: bytes, pages: Optional[list[int]] = None) -> list[StructureElement]:
|
||||
"""Extract structure-tree element references from tagged PDF bytes.
|
||||
|
||||
See :func:`extract_structure_elements` for details.
|
||||
"""
|
||||
...
|
||||
|
||||
def extract_text_in_regions(
|
||||
path: str,
|
||||
page_regions: list[tuple[int, list[list[float]]]],
|
||||
|
||||
+3
-3
@@ -4,9 +4,9 @@ build-backend = "maturin"
|
||||
|
||||
[project]
|
||||
name = "pdf-inspector"
|
||||
# Bump this to publish to PyPI — CI publishes automatically when the version
|
||||
# changes on main (same flow as napi/package.json for npm).
|
||||
version = "0.2.7"
|
||||
# Keep package versions in sync with `python3 scripts/version.py <version>`.
|
||||
# CI publishes automatically when the synchronized change lands on main.
|
||||
version = "1.14.0"
|
||||
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
|
||||
readme = "docs/python.md"
|
||||
license = { text = "MIT" }
|
||||
|
||||
@@ -0,0 +1,111 @@
|
||||
import json
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
|
||||
from version import PLATFORM_PACKAGES, check_versions, set_versions
|
||||
|
||||
|
||||
class VersionTests(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.temporary = tempfile.TemporaryDirectory()
|
||||
self.root = Path(self.temporary.name)
|
||||
(self.root / "napi").mkdir()
|
||||
(self.root / "site").mkdir()
|
||||
(self.root / "wasm").mkdir()
|
||||
|
||||
self._write_manifest("Cargo.toml", "package", "0.1.0")
|
||||
self._write_manifest("pyproject.toml", "project", "0.1.0")
|
||||
self._write_manifest("napi/Cargo.toml", "package", "0.1.0")
|
||||
self._write_manifest("wasm/Cargo.toml", "package", "0.1.0")
|
||||
|
||||
package = {
|
||||
"name": "@firecrawl/pdf-inspector",
|
||||
"version": "0.1.0",
|
||||
"optionalDependencies": {
|
||||
dependency: "0.1.0" for dependency in PLATFORM_PACKAGES
|
||||
},
|
||||
}
|
||||
(self.root / "napi/package.json").write_text(
|
||||
json.dumps(package), encoding="utf-8"
|
||||
)
|
||||
(self.root / "napi/bun.lock").write_text(
|
||||
"\n".join(
|
||||
f' "{dependency}": "0.1.0",'
|
||||
for dependency in PLATFORM_PACKAGES
|
||||
)
|
||||
+ "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
(self.root / "site/index.html").write_text(
|
||||
'https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.0/'
|
||||
'pdf_inspector_wasm.js\n',
|
||||
encoding="utf-8",
|
||||
)
|
||||
self._write_lock(
|
||||
"napi/Cargo.lock", ("pdf-inspector", "pdf-inspector-napi")
|
||||
)
|
||||
self._write_lock(
|
||||
"wasm/Cargo.lock", ("pdf-inspector", "pdf-inspector-wasm")
|
||||
)
|
||||
|
||||
def tearDown(self):
|
||||
self.temporary.cleanup()
|
||||
|
||||
def _write_manifest(self, relative, section, version):
|
||||
(self.root / relative).write_text(
|
||||
f'[{section}]\nname = "fixture"\nversion = "{version}"\n',
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
def _write_lock(self, relative, packages):
|
||||
content = "\n".join(
|
||||
f'[[package]]\nname = "{package}"\nversion = "0.1.0"\n'
|
||||
for package in packages
|
||||
)
|
||||
(self.root / relative).write_text(content, encoding="utf-8")
|
||||
|
||||
def test_updates_every_version_location(self):
|
||||
set_versions("1.14.0", self.root)
|
||||
|
||||
self.assertEqual(check_versions(self.root), "1.14.0")
|
||||
|
||||
def test_reports_a_divergent_package(self):
|
||||
self._write_manifest("wasm/Cargo.toml", "package", "0.2.0")
|
||||
|
||||
with self.assertRaisesRegex(ValueError, "WASM package: 0.2.0"):
|
||||
check_versions(self.root)
|
||||
|
||||
def test_rejects_an_invalid_version(self):
|
||||
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
|
||||
set_versions("next", self.root)
|
||||
|
||||
def test_rejects_numeric_prerelease_with_leading_zero(self):
|
||||
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
|
||||
|
||||
with self.assertRaisesRegex(ValueError, "Invalid semantic version"):
|
||||
set_versions("1.2.3-01", self.root)
|
||||
|
||||
self.assertEqual(
|
||||
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
|
||||
)
|
||||
|
||||
def test_preflight_failure_does_not_partially_update(self):
|
||||
before = (self.root / "Cargo.toml").read_text(encoding="utf-8")
|
||||
(self.root / "site/index.html").write_text(
|
||||
"missing module URL\n", encoding="utf-8"
|
||||
)
|
||||
|
||||
with self.assertRaisesRegex(ValueError, "Missing pinned WASM package URL"):
|
||||
set_versions("1.14.0", self.root)
|
||||
|
||||
self.assertEqual(
|
||||
(self.root / "Cargo.toml").read_text(encoding="utf-8"), before
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,262 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Keep every pdf-inspector package on one release version."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
PRERELEASE_IDENTIFIER = (
|
||||
r"(?:0|[1-9]\d*|[0-9A-Za-z-]*[A-Za-z-][0-9A-Za-z-]*)"
|
||||
)
|
||||
SEMVER = re.compile(
|
||||
r"^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)"
|
||||
rf"(?:-{PRERELEASE_IDENTIFIER}(?:\.{PRERELEASE_IDENTIFIER})*)?"
|
||||
r"(?:\+[0-9A-Za-z-]+(?:\.[0-9A-Za-z-]+)*)?$"
|
||||
)
|
||||
VERSION_LINE = re.compile(r'^(\s*version\s*=\s*")[^"]+(".*)$')
|
||||
SECTION_LINE = re.compile(r"^\s*\[([^]]+)]\s*$")
|
||||
PLATFORM_PACKAGES = (
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu",
|
||||
"@firecrawl/pdf-inspector-linux-x64-musl",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-gnu",
|
||||
"@firecrawl/pdf-inspector-linux-arm64-musl",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc",
|
||||
)
|
||||
TOML_VERSIONS = (
|
||||
("Rust crate", Path("Cargo.toml"), "package"),
|
||||
("Python package", Path("pyproject.toml"), "project"),
|
||||
("NAPI crate", Path("napi/Cargo.toml"), "package"),
|
||||
("WASM package", Path("wasm/Cargo.toml"), "package"),
|
||||
)
|
||||
LOCK_VERSIONS = (
|
||||
("NAPI lock: core", Path("napi/Cargo.lock"), "pdf-inspector"),
|
||||
("NAPI lock: binding", Path("napi/Cargo.lock"), "pdf-inspector-napi"),
|
||||
("WASM lock: core", Path("wasm/Cargo.lock"), "pdf-inspector"),
|
||||
("WASM lock: binding", Path("wasm/Cargo.lock"), "pdf-inspector-wasm"),
|
||||
)
|
||||
SITE_WASM_VERSION = re.compile(
|
||||
r"(@firecrawl/pdf-inspector-wasm@)([^/\"]+)(/pdf_inspector_wasm\.js)"
|
||||
)
|
||||
|
||||
|
||||
def _read_section_version(path: Path, section: str) -> str:
|
||||
active = False
|
||||
for line in path.read_text(encoding="utf-8").splitlines():
|
||||
section_match = SECTION_LINE.match(line)
|
||||
if section_match:
|
||||
active = section_match.group(1) == section
|
||||
elif active:
|
||||
version_match = VERSION_LINE.match(line)
|
||||
if version_match:
|
||||
return line.split('"', 2)[1]
|
||||
raise ValueError(f"No version found in [{section}] of {path}")
|
||||
|
||||
|
||||
def _write_section_version(path: Path, section: str, version: str) -> None:
|
||||
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
|
||||
active = False
|
||||
for index, line in enumerate(lines):
|
||||
section_match = SECTION_LINE.match(line)
|
||||
if section_match:
|
||||
active = section_match.group(1) == section
|
||||
elif active:
|
||||
version_match = VERSION_LINE.match(line)
|
||||
if version_match:
|
||||
newline = "\n" if line.endswith("\n") else ""
|
||||
replacement = (
|
||||
f"{version_match.group(1)}{version}"
|
||||
f"{version_match.group(2).rstrip()}"
|
||||
)
|
||||
lines[index] = (
|
||||
f"{replacement}{newline}"
|
||||
)
|
||||
path.write_text("".join(lines), encoding="utf-8")
|
||||
return
|
||||
raise ValueError(f"No version found in [{section}] of {path}")
|
||||
|
||||
|
||||
def _package_block(lines: list[str], package: str) -> tuple[int, int]:
|
||||
for start, line in enumerate(lines):
|
||||
if line.strip() != "[[package]]":
|
||||
continue
|
||||
end = next(
|
||||
(
|
||||
index
|
||||
for index in range(start + 1, len(lines))
|
||||
if lines[index].strip() == "[[package]]"
|
||||
),
|
||||
len(lines),
|
||||
)
|
||||
if any(line.strip() == f'name = "{package}"' for line in lines[start:end]):
|
||||
return start, end
|
||||
raise ValueError(f"No lockfile entry found for {package}")
|
||||
|
||||
|
||||
def _read_lock_version(path: Path, package: str) -> str:
|
||||
lines = path.read_text(encoding="utf-8").splitlines()
|
||||
start, end = _package_block(lines, package)
|
||||
for line in lines[start:end]:
|
||||
version_match = VERSION_LINE.match(line)
|
||||
if version_match:
|
||||
return line.split('"', 2)[1]
|
||||
raise ValueError(f"No version found for {package} in {path}")
|
||||
|
||||
|
||||
def _write_lock_version(path: Path, package: str, version: str) -> None:
|
||||
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
|
||||
start, end = _package_block(lines, package)
|
||||
for index in range(start, end):
|
||||
version_match = VERSION_LINE.match(lines[index])
|
||||
if version_match:
|
||||
newline = "\n" if lines[index].endswith("\n") else ""
|
||||
lines[index] = (
|
||||
f'{version_match.group(1)}{version}{version_match.group(2).rstrip()}'
|
||||
f"{newline}"
|
||||
)
|
||||
path.write_text("".join(lines), encoding="utf-8")
|
||||
return
|
||||
raise ValueError(f"No version found for {package} in {path}")
|
||||
|
||||
|
||||
def _node_versions(root: Path) -> dict[str, str]:
|
||||
package = json.loads((root / "napi/package.json").read_text(encoding="utf-8"))
|
||||
versions = {"Node package": package["version"]}
|
||||
optional = package.get("optionalDependencies", {})
|
||||
for dependency in PLATFORM_PACKAGES:
|
||||
if dependency not in optional:
|
||||
raise ValueError(f"Missing Node optional dependency: {dependency}")
|
||||
versions[f"Node optional dependency: {dependency}"] = optional[dependency]
|
||||
return versions
|
||||
|
||||
|
||||
def _bun_versions(root: Path) -> dict[str, str]:
|
||||
text = (root / "napi/bun.lock").read_text(encoding="utf-8")
|
||||
versions = {}
|
||||
for dependency in PLATFORM_PACKAGES:
|
||||
match = re.search(
|
||||
rf'"{re.escape(dependency)}": "([^"]+)"[,]', text
|
||||
)
|
||||
if not match:
|
||||
raise ValueError(f"Missing Bun lock dependency: {dependency}")
|
||||
versions[f"Bun lock: {dependency}"] = match.group(1)
|
||||
return versions
|
||||
|
||||
|
||||
def _site_wasm_version(root: Path) -> str:
|
||||
text = (root / "site/index.html").read_text(encoding="utf-8")
|
||||
match = SITE_WASM_VERSION.search(text)
|
||||
if not match:
|
||||
raise ValueError("Missing pinned WASM package URL in site/index.html")
|
||||
return match.group(2)
|
||||
|
||||
|
||||
def package_versions(root: Path = ROOT) -> dict[str, str]:
|
||||
versions = {
|
||||
label: _read_section_version(root / relative, section)
|
||||
for label, relative, section in TOML_VERSIONS
|
||||
}
|
||||
versions.update(_node_versions(root))
|
||||
versions.update(_bun_versions(root))
|
||||
versions["Website WASM module"] = _site_wasm_version(root)
|
||||
versions.update(
|
||||
{
|
||||
label: _read_lock_version(root / relative, package)
|
||||
for label, relative, package in LOCK_VERSIONS
|
||||
}
|
||||
)
|
||||
return versions
|
||||
|
||||
|
||||
def check_versions(root: Path = ROOT) -> str:
|
||||
versions = package_versions(root)
|
||||
expected = versions["Rust crate"]
|
||||
if not SEMVER.fullmatch(expected):
|
||||
raise ValueError(f"Rust crate has an invalid semantic version: {expected}")
|
||||
mismatches = {
|
||||
label: version for label, version in versions.items() if version != expected
|
||||
}
|
||||
if mismatches:
|
||||
details = "\n".join(
|
||||
f" - {label}: {version}" for label, version in mismatches.items()
|
||||
)
|
||||
raise ValueError(f"Expected every package to use {expected}:\n{details}")
|
||||
return expected
|
||||
|
||||
|
||||
def set_versions(version: str, root: Path = ROOT) -> None:
|
||||
if not SEMVER.fullmatch(version):
|
||||
raise ValueError(f"Invalid semantic version: {version}")
|
||||
|
||||
# Validate every expected location before writing the first file. This
|
||||
# prevents a stale manifest or generated file from leaving a partial bump.
|
||||
package_versions(root)
|
||||
|
||||
for _, relative, section in TOML_VERSIONS:
|
||||
_write_section_version(root / relative, section, version)
|
||||
|
||||
package_path = root / "napi/package.json"
|
||||
package = json.loads(package_path.read_text(encoding="utf-8"))
|
||||
package["version"] = version
|
||||
optional = package.get("optionalDependencies", {})
|
||||
for dependency in PLATFORM_PACKAGES:
|
||||
if dependency not in optional:
|
||||
raise ValueError(f"Missing Node optional dependency: {dependency}")
|
||||
optional[dependency] = version
|
||||
package_path.write_text(json.dumps(package, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
bun_path = root / "napi/bun.lock"
|
||||
bun_text = bun_path.read_text(encoding="utf-8")
|
||||
for dependency in PLATFORM_PACKAGES:
|
||||
pattern = rf'("{re.escape(dependency)}": ")[^"]+("[,])'
|
||||
bun_text, count = re.subn(
|
||||
pattern, rf"\g<1>{version}\g<2>", bun_text, count=1
|
||||
)
|
||||
if count != 1:
|
||||
raise ValueError(f"Missing Bun lock dependency: {dependency}")
|
||||
bun_path.write_text(bun_text, encoding="utf-8")
|
||||
|
||||
site_path = root / "site/index.html"
|
||||
site_text = site_path.read_text(encoding="utf-8")
|
||||
site_text, count = SITE_WASM_VERSION.subn(
|
||||
rf"\g<1>{version}\g<3>", site_text, count=1
|
||||
)
|
||||
if count != 1:
|
||||
raise ValueError("Missing pinned WASM package URL in site/index.html")
|
||||
site_path.write_text(site_text, encoding="utf-8")
|
||||
|
||||
for _, relative, package_name in LOCK_VERSIONS:
|
||||
_write_lock_version(root / relative, package_name, version)
|
||||
|
||||
check_versions(root)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("version", nargs="?", help="new shared semantic version")
|
||||
parser.add_argument(
|
||||
"--check", action="store_true", help="fail if package versions have diverged"
|
||||
)
|
||||
arguments = parser.parse_args()
|
||||
if arguments.check == bool(arguments.version):
|
||||
parser.error("provide either a version or --check")
|
||||
|
||||
try:
|
||||
if arguments.check:
|
||||
version = check_versions()
|
||||
print(f"All packages use {version}")
|
||||
else:
|
||||
set_versions(arguments.version)
|
||||
print(f"Updated all packages to {arguments.version}")
|
||||
except ValueError as error:
|
||||
parser.exit(1, f"{error}\n")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
+1
-1
@@ -975,7 +975,7 @@ result = pdf_inspector.<span class="fn">process_pdf</span>(<span class="str">"do
|
||||
<script>
|
||||
(() => {
|
||||
const MAX_FILE_SIZE = 25 * 1024 * 1024;
|
||||
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@0.1.1/pdf_inspector_wasm.js";
|
||||
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@1.14.0/pdf_inspector_wasm.js";
|
||||
const input = document.querySelector("#pdf-input");
|
||||
const dropZone = document.querySelector("#drop-zone");
|
||||
const filePanel = document.querySelector("#demo-file");
|
||||
|
||||
+74
@@ -657,6 +657,80 @@ pub fn extract_pages_markdown<P: AsRef<Path>>(
|
||||
extract_pages_markdown_mem(&buffer, pages)
|
||||
}
|
||||
|
||||
// =========================================================================
|
||||
// Structure-tree element extraction (tagged PDFs)
|
||||
// =========================================================================
|
||||
|
||||
/// One structure-tree element reference from a tagged PDF, resolved to a
|
||||
/// page and Marked Content ID.
|
||||
///
|
||||
/// Join `(page, mcid)` against [`TextItem::page`] / [`TextItem::mcid`] from
|
||||
/// [`extract_text_with_positions`] to attach semantic roles (heading levels,
|
||||
/// paragraphs, table cells, …) to extracted text.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct StructureElement {
|
||||
/// 1-indexed page number (matches [`TextItem::page`]).
|
||||
pub page: u32,
|
||||
/// Marked Content ID from the page's content stream (matches
|
||||
/// [`TextItem::mcid`]).
|
||||
pub mcid: i64,
|
||||
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", …).
|
||||
/// Custom tags are resolved through the document's `/RoleMap`; tags
|
||||
/// with no standard mapping are returned verbatim.
|
||||
pub role: String,
|
||||
}
|
||||
|
||||
/// Extract structure-tree element references from a tagged PDF in memory.
|
||||
///
|
||||
/// Parses `/StructTreeRoot` (when present) and returns one entry per
|
||||
/// marked-content reference, resolved to its 1-indexed page, MCID, and
|
||||
/// structure type name. Returns an empty list when the PDF is not tagged.
|
||||
///
|
||||
/// Pass `Some(&[...])` with 1-indexed page numbers (matching
|
||||
/// [`TextItem::page`]) to restrict output to those pages; pass `None` for
|
||||
/// the whole document. Entries are sorted by `(page, mcid)`.
|
||||
pub fn extract_structure_elements_mem(
|
||||
buffer: &[u8],
|
||||
pages: Option<&[u32]>,
|
||||
) -> Result<Vec<StructureElement>, PdfError> {
|
||||
validate_pdf_bytes(buffer)?;
|
||||
let (doc, _page_count) = load_document_from_mem(buffer)?;
|
||||
let Some(tree) = structure_tree::StructTree::from_doc(&doc) else {
|
||||
return Ok(Vec::new());
|
||||
};
|
||||
let page_ids = doc.get_pages();
|
||||
let roles = tree.mcid_to_roles(&page_ids);
|
||||
|
||||
let page_filter: Option<HashSet<u32>> = pages.map(|p| p.iter().copied().collect());
|
||||
let mut elements: Vec<StructureElement> = roles
|
||||
.into_iter()
|
||||
.filter(|(page, _)| page_filter.as_ref().is_none_or(|f| f.contains(page)))
|
||||
.flat_map(|(page, mcids)| {
|
||||
mcids.into_iter().map(move |(mcid, role)| StructureElement {
|
||||
page,
|
||||
mcid,
|
||||
role: role.name().to_string(),
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
elements.sort_unstable_by_key(|e| (e.page, e.mcid));
|
||||
Ok(elements)
|
||||
}
|
||||
|
||||
/// Path-based wrapper for [`extract_structure_elements_mem`].
|
||||
///
|
||||
/// Reads the PDF from disk and extracts structure-tree element references.
|
||||
/// Pass `None` for `pages` to return the whole document, or `Some(&[...])`
|
||||
/// to restrict to specific 1-indexed pages.
|
||||
pub fn extract_structure_elements<P: AsRef<Path>>(
|
||||
path: P,
|
||||
pages: Option<&[u32]>,
|
||||
) -> Result<Vec<StructureElement>, PdfError> {
|
||||
validate_pdf_file(&path)?;
|
||||
let buffer = std::fs::read(path.as_ref())?;
|
||||
extract_structure_elements_mem(&buffer, pages)
|
||||
}
|
||||
|
||||
// =========================================================================
|
||||
// Region-based text extraction (for hybrid OCR pipelines)
|
||||
// =========================================================================
|
||||
|
||||
@@ -271,6 +271,12 @@ pub struct PyTextItem {
|
||||
pub is_strikeout: bool,
|
||||
#[pyo3(get)]
|
||||
pub item_type: String,
|
||||
/// Marked Content ID from the content stream's BDC/BMC operator, None
|
||||
/// when the text is not part of marked content. Join with the
|
||||
/// (page, mcid) pairs from extract_structure_elements to attach
|
||||
/// structure-tree roles (headings, paragraphs, ...) in tagged PDFs.
|
||||
#[pyo3(get)]
|
||||
pub mcid: Option<i64>,
|
||||
}
|
||||
|
||||
#[pymethods]
|
||||
@@ -286,6 +292,32 @@ impl PyTextItem {
|
||||
}
|
||||
}
|
||||
|
||||
/// One structure-tree element reference from a tagged PDF.
|
||||
#[pyclass(name = "StructureElement")]
|
||||
#[derive(Clone)]
|
||||
pub struct PyStructureElement {
|
||||
/// 1-indexed page number (matches TextItem.page).
|
||||
#[pyo3(get)]
|
||||
pub page: u32,
|
||||
/// Marked Content ID from the page's content stream (matches
|
||||
/// TextItem.mcid).
|
||||
#[pyo3(get)]
|
||||
pub mcid: i64,
|
||||
/// Standard structure type name ("H1".."H6", "P", "Table", "TD", ...).
|
||||
#[pyo3(get)]
|
||||
pub role: String,
|
||||
}
|
||||
|
||||
#[pymethods]
|
||||
impl PyStructureElement {
|
||||
fn __repr__(&self) -> String {
|
||||
format!(
|
||||
"StructureElement(page={}, mcid={}, role='{}')",
|
||||
self.page, self.mcid, self.role
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Helpers
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -356,6 +388,18 @@ fn convert_text_items(items: Vec<crate::TextItem>) -> Vec<PyTextItem> {
|
||||
is_underline: item.is_underline,
|
||||
is_strikeout: item.is_strikeout,
|
||||
item_type: item_type_str(&item.item_type),
|
||||
mcid: item.mcid,
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn convert_structure_elements(elements: Vec<crate::StructureElement>) -> Vec<PyStructureElement> {
|
||||
elements
|
||||
.into_iter()
|
||||
.map(|e| PyStructureElement {
|
||||
page: e.page,
|
||||
mcid: e.mcid,
|
||||
role: e.role,
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
@@ -613,6 +657,48 @@ fn extract_pages_markdown_bytes(
|
||||
Ok(to_py_pages_result(result))
|
||||
}
|
||||
|
||||
/// Extract structure-tree element references from a tagged PDF file.
|
||||
///
|
||||
/// Parses the document's structure tree (when present) and returns one
|
||||
/// entry per marked-content reference, resolved to its 1-indexed page,
|
||||
/// MCID, and structure type name ("H1".."H6", "P", "Table", ...). Returns
|
||||
/// an empty list when the PDF is not tagged.
|
||||
///
|
||||
/// Join (page, mcid) against the page/mcid attributes from
|
||||
/// [`extract_text_with_positions`] to attach heading levels and other
|
||||
/// semantic roles to extracted text.
|
||||
///
|
||||
/// Args:
|
||||
/// path: Path to the PDF file.
|
||||
/// pages: Optional list of 1-indexed pages (matching TextItem.page).
|
||||
/// When None (default), the whole document is returned.
|
||||
///
|
||||
/// Returns:
|
||||
/// List of StructureElement sorted by (page, mcid).
|
||||
#[pyfunction]
|
||||
#[pyo3(signature = (path, pages=None))]
|
||||
fn extract_structure_elements(
|
||||
path: &str,
|
||||
pages: Option<Vec<u32>>,
|
||||
) -> PyResult<Vec<PyStructureElement>> {
|
||||
let elements = crate::extract_structure_elements(path, pages.as_deref()).map_err(to_py_err)?;
|
||||
Ok(convert_structure_elements(elements))
|
||||
}
|
||||
|
||||
/// Extract structure-tree element references from tagged PDF bytes.
|
||||
///
|
||||
/// See [`extract_structure_elements`] for details.
|
||||
#[pyfunction]
|
||||
#[pyo3(signature = (data, pages=None))]
|
||||
fn extract_structure_elements_bytes(
|
||||
data: &[u8],
|
||||
pages: Option<Vec<u32>>,
|
||||
) -> PyResult<Vec<PyStructureElement>> {
|
||||
let elements =
|
||||
crate::extract_structure_elements_mem(data, pages.as_deref()).map_err(to_py_err)?;
|
||||
Ok(convert_structure_elements(elements))
|
||||
}
|
||||
|
||||
/// Python module definition.
|
||||
#[pymodule]
|
||||
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
||||
@@ -620,6 +706,7 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
||||
m.add_class::<PyPageOcrReasons>()?;
|
||||
m.add_class::<PyPdfClassification>()?;
|
||||
m.add_class::<PyTextItem>()?;
|
||||
m.add_class::<PyStructureElement>()?;
|
||||
m.add_class::<PyRegionText>()?;
|
||||
m.add_class::<PyPageRegionTexts>()?;
|
||||
m.add_class::<PyPageMarkdown>()?;
|
||||
@@ -634,6 +721,8 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
|
||||
m.add_function(wrap_pyfunction!(extract_text_bytes, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_text_with_positions, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_text_with_positions_bytes, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_structure_elements, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_structure_elements_bytes, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_text_in_regions, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_text_in_regions_bytes, m)?)?;
|
||||
m.add_function(wrap_pyfunction!(extract_pages_markdown, m)?)?;
|
||||
|
||||
@@ -122,6 +122,65 @@ impl StructRole {
|
||||
)
|
||||
}
|
||||
|
||||
/// The standard structure type name for this role ("H1", "P", "Table", …).
|
||||
///
|
||||
/// Inverse of [`StructRole::from_name`]: for [`StructRole::Other`] the
|
||||
/// custom tag name is returned verbatim.
|
||||
pub fn name(&self) -> &str {
|
||||
match self {
|
||||
Self::Document => "Document",
|
||||
Self::Part => "Part",
|
||||
Self::Art => "Art",
|
||||
Self::Sect => "Sect",
|
||||
Self::Div => "Div",
|
||||
Self::BlockQuote => "BlockQuote",
|
||||
Self::Caption => "Caption",
|
||||
Self::TOC => "TOC",
|
||||
Self::TOCI => "TOCI",
|
||||
Self::Index => "Index",
|
||||
Self::NonStruct => "NonStruct",
|
||||
Self::Private => "Private",
|
||||
Self::H => "H",
|
||||
Self::H1 => "H1",
|
||||
Self::H2 => "H2",
|
||||
Self::H3 => "H3",
|
||||
Self::H4 => "H4",
|
||||
Self::H5 => "H5",
|
||||
Self::H6 => "H6",
|
||||
Self::P => "P",
|
||||
Self::L => "L",
|
||||
Self::LI => "LI",
|
||||
Self::Lbl => "Lbl",
|
||||
Self::LBody => "LBody",
|
||||
Self::Table => "Table",
|
||||
Self::TR => "TR",
|
||||
Self::TH => "TH",
|
||||
Self::TD => "TD",
|
||||
Self::THead => "THead",
|
||||
Self::TBody => "TBody",
|
||||
Self::TFoot => "TFoot",
|
||||
Self::Span => "Span",
|
||||
Self::Quote => "Quote",
|
||||
Self::Note => "Note",
|
||||
Self::Reference => "Reference",
|
||||
Self::BibEntry => "BibEntry",
|
||||
Self::Code => "Code",
|
||||
Self::Link => "Link",
|
||||
Self::Annot => "Annot",
|
||||
Self::Figure => "Figure",
|
||||
Self::Formula => "Formula",
|
||||
Self::Form => "Form",
|
||||
Self::Ruby => "Ruby",
|
||||
Self::RB => "RB",
|
||||
Self::RT => "RT",
|
||||
Self::RP => "RP",
|
||||
Self::Warichu => "Warichu",
|
||||
Self::WT => "WT",
|
||||
Self::WP => "WP",
|
||||
Self::Other(name) => name,
|
||||
}
|
||||
}
|
||||
|
||||
fn from_name(name: &str) -> Self {
|
||||
match name {
|
||||
"Document" => Self::Document,
|
||||
@@ -1233,6 +1292,66 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_struct_role_name_roundtrip() {
|
||||
// `name()` is the inverse of `from_name` for every standard type.
|
||||
for name in [
|
||||
"Document",
|
||||
"Part",
|
||||
"Art",
|
||||
"Sect",
|
||||
"Div",
|
||||
"BlockQuote",
|
||||
"Caption",
|
||||
"TOC",
|
||||
"TOCI",
|
||||
"Index",
|
||||
"NonStruct",
|
||||
"Private",
|
||||
"H",
|
||||
"H1",
|
||||
"H2",
|
||||
"H3",
|
||||
"H4",
|
||||
"H5",
|
||||
"H6",
|
||||
"P",
|
||||
"L",
|
||||
"LI",
|
||||
"Lbl",
|
||||
"LBody",
|
||||
"Table",
|
||||
"TR",
|
||||
"TH",
|
||||
"TD",
|
||||
"THead",
|
||||
"TBody",
|
||||
"TFoot",
|
||||
"Span",
|
||||
"Quote",
|
||||
"Note",
|
||||
"Reference",
|
||||
"BibEntry",
|
||||
"Code",
|
||||
"Link",
|
||||
"Annot",
|
||||
"Figure",
|
||||
"Formula",
|
||||
"Form",
|
||||
"Ruby",
|
||||
"RB",
|
||||
"RT",
|
||||
"RP",
|
||||
"Warichu",
|
||||
"WT",
|
||||
"WP",
|
||||
] {
|
||||
assert_eq!(StructRole::from_name(name).name(), name);
|
||||
}
|
||||
// Custom tags pass through verbatim.
|
||||
assert_eq!(StructRole::from_name("CustomTag").name(), "CustomTag");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_struct_role_with_role_map() {
|
||||
let mut role_map = HashMap::new();
|
||||
|
||||
@@ -1429,6 +1429,77 @@ fn test_firecrawl_tagged_pdf_struct_tree() {
|
||||
assert_eq!(fence_count % 2, 0, "Code fences should be balanced");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_tagged_pdf_text_items_carry_mcid() {
|
||||
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
|
||||
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
|
||||
assert!(
|
||||
items.iter().any(|i| i.mcid.is_some()),
|
||||
"Tagged PDF text items should carry Marked Content IDs"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_structure_elements_tagged_pdf() {
|
||||
let buf = std::fs::read("tests/fixtures/firecrawl_docs_tagged.pdf").unwrap();
|
||||
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
|
||||
assert!(!elements.is_empty(), "Tagged PDF should yield elements");
|
||||
assert!(
|
||||
elements.iter().any(|e| e.role == "H1"),
|
||||
"Should surface H1 heading roles"
|
||||
);
|
||||
assert!(
|
||||
elements.iter().all(|e| !e.role.is_empty()),
|
||||
"Every element should carry a role name"
|
||||
);
|
||||
|
||||
// Sorted by (page, mcid) for deterministic output
|
||||
assert!(
|
||||
elements
|
||||
.windows(2)
|
||||
.all(|w| (w[0].page, w[0].mcid) <= (w[1].page, w[1].mcid)),
|
||||
"Elements should be sorted by (page, mcid)"
|
||||
);
|
||||
|
||||
// The advertised join: (page, mcid) pairs must line up with the
|
||||
// mcid-carrying TextItems from positioned extraction, and joining the
|
||||
// H1 entries must recover non-empty heading text.
|
||||
let items = pdf_inspector::extractor::extract_text_with_positions_mem(&buf).unwrap();
|
||||
let h1_refs: std::collections::HashSet<(u32, i64)> = elements
|
||||
.iter()
|
||||
.filter(|e| e.role == "H1")
|
||||
.map(|e| (e.page, e.mcid))
|
||||
.collect();
|
||||
let h1_text: String = items
|
||||
.iter()
|
||||
.filter(|i| i.mcid.is_some_and(|mcid| h1_refs.contains(&(i.page, mcid))))
|
||||
.map(|i| i.text.as_str())
|
||||
.collect();
|
||||
assert!(
|
||||
!h1_text.trim().is_empty(),
|
||||
"Joining H1 structure elements to text items should recover heading text"
|
||||
);
|
||||
|
||||
// Page filter is 1-indexed (matching TextItem.page) and equals the
|
||||
// corresponding subset of the full document result.
|
||||
let page1 = pdf_inspector::extract_structure_elements_mem(&buf, Some(&[1])).unwrap();
|
||||
assert!(!page1.is_empty(), "Page 1 should have elements");
|
||||
assert!(page1.iter().all(|e| e.page == 1));
|
||||
let full_page1_count = elements.iter().filter(|e| e.page == 1).count();
|
||||
assert_eq!(page1.len(), full_page1_count);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_structure_elements_untagged_pdf_empty() {
|
||||
let buf = std::fs::read("tests/fixtures/thermo-freon12.pdf").unwrap();
|
||||
let elements = pdf_inspector::extract_structure_elements_mem(&buf, None).unwrap();
|
||||
assert!(
|
||||
elements.is_empty(),
|
||||
"Untagged PDF should yield no structure elements, got {:?}",
|
||||
elements
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_identity_h_no_tounicode_suppresses_garbage() {
|
||||
// shinagawa_identity_h.pdf uses YuGothic with Identity-H encoding and no
|
||||
|
||||
@@ -203,6 +203,79 @@ class TestExtractTextWithPositions:
|
||||
assert len(items) > 0
|
||||
assert all(item.page == 1 for item in items)
|
||||
|
||||
def test_mcid(self):
|
||||
# Untagged fixture: mcid is None or int, never anything else
|
||||
items = pdf_inspector.extract_text_with_positions(
|
||||
fixture_path("thermo-freon12.pdf")
|
||||
)
|
||||
assert all(item.mcid is None or isinstance(item.mcid, int) for item in items)
|
||||
# Tagged fixture: marked content carries MCIDs
|
||||
tagged = pdf_inspector.extract_text_with_positions(
|
||||
fixture_path("firecrawl_docs_tagged.pdf")
|
||||
)
|
||||
assert any(item.mcid is not None for item in tagged)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# extract_structure_elements / extract_structure_elements_bytes
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
class TestExtractStructureElements:
|
||||
def test_tagged_file(self):
|
||||
elements = pdf_inspector.extract_structure_elements(
|
||||
fixture_path("firecrawl_docs_tagged.pdf")
|
||||
)
|
||||
assert len(elements) > 0
|
||||
assert all(isinstance(e.page, int) for e in elements)
|
||||
assert all(isinstance(e.mcid, int) for e in elements)
|
||||
assert all(isinstance(e.role, str) and len(e.role) > 0 for e in elements)
|
||||
assert any(e.role == "H1" for e in elements)
|
||||
|
||||
def test_join_with_text_items(self):
|
||||
# (page, mcid) joins against extract_text_with_positions to recover
|
||||
# heading text
|
||||
path = fixture_path("firecrawl_docs_tagged.pdf")
|
||||
elements = pdf_inspector.extract_structure_elements(path)
|
||||
items = pdf_inspector.extract_text_with_positions(path)
|
||||
h1_refs = {(e.page, e.mcid) for e in elements if e.role == "H1"}
|
||||
h1_text = "".join(
|
||||
item.text
|
||||
for item in items
|
||||
if item.mcid is not None and (item.page, item.mcid) in h1_refs
|
||||
)
|
||||
assert len(h1_text.strip()) > 0
|
||||
|
||||
def test_with_pages(self):
|
||||
# pages filter is 1-indexed, matching TextItem.page
|
||||
elements = pdf_inspector.extract_structure_elements(
|
||||
fixture_path("firecrawl_docs_tagged.pdf"), pages=[1]
|
||||
)
|
||||
assert len(elements) > 0
|
||||
assert all(e.page == 1 for e in elements)
|
||||
|
||||
def test_bytes(self):
|
||||
data = fixture_bytes("firecrawl_docs_tagged.pdf")
|
||||
elements = pdf_inspector.extract_structure_elements_bytes(data)
|
||||
assert len(elements) > 0
|
||||
assert any(e.role == "H1" for e in elements)
|
||||
|
||||
def test_untagged_returns_empty(self):
|
||||
elements = pdf_inspector.extract_structure_elements(
|
||||
fixture_path("thermo-freon12.pdf")
|
||||
)
|
||||
assert elements == []
|
||||
|
||||
def test_repr(self):
|
||||
elements = pdf_inspector.extract_structure_elements(
|
||||
fixture_path("firecrawl_docs_tagged.pdf")
|
||||
)
|
||||
assert "StructureElement" in repr(elements[0])
|
||||
|
||||
def test_not_a_pdf(self):
|
||||
with pytest.raises(ValueError):
|
||||
pdf_inspector.extract_structure_elements_bytes(b"not a pdf")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# extract_text_in_regions / extract_text_in_regions_bytes
|
||||
|
||||
Generated
+2
-2
@@ -724,7 +724,7 @@ checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector"
|
||||
version = "0.1.8"
|
||||
version = "1.14.0"
|
||||
dependencies = [
|
||||
"env_logger",
|
||||
"include_dir",
|
||||
@@ -740,7 +740,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector-wasm"
|
||||
version = "0.1.4"
|
||||
version = "1.14.0"
|
||||
dependencies = [
|
||||
"console_error_panic_hook",
|
||||
"js-sys",
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "pdf-inspector-wasm"
|
||||
version = "0.1.4"
|
||||
version = "1.14.0"
|
||||
edition = "2021"
|
||||
authors = ["Firecrawl Team"]
|
||||
description = "Browser WebAssembly bindings for pdf-inspector"
|
||||
|
||||
Reference in New Issue
Block a user