Compare commits

..
13 changed files with 1702 additions and 288 deletions
+16 -2
View File
@@ -5,7 +5,7 @@
[![PyPI](https://img.shields.io/pypi/v/pdf-inspector.svg)](https://pypi.org/project/pdf-inspector/)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
Fast Rust library for PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Includes bindings for [Python](docs/python.md), [Node.js](napi/README.md), and [browser WebAssembly](wasm/README.md).
Fast Rust library for PDF classification and text extraction. By default it detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown without OCR. Native Rust and CLI consumers can opt into selective OCR. Includes bindings for [Python](docs/python.md), [Node.js](napi/README.md), and [browser WebAssembly](wasm/README.md).
Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.
@@ -18,9 +18,10 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
- **CID font support** — ToUnicode CMap decoding for Type0/Identity-H fonts, UTF-16BE, UTF-8, and Latin-1 encodings.
- **Multi-column layout** — Automatic detection of newspaper-style columns, sequential reading order, and RTL text support.
- **Encoding issue detection** — Automatically flags broken font encodings so callers can fall back to OCR.
- **Optional OCR** — An opt-in Rust and CLI feature selectively renders only pages that need OCR, runs PP-OCRv6 Small locally, and preserves per-page provenance and hosted-fallback recommendations.
- **Single document load** — The document is parsed once and shared between detection and extraction, avoiding redundant I/O.
- **Browser WebAssembly** — Run the same Rust parser locally in browsers and Web Workers, with embedded CMaps and no server round trip.
- **Lightweight** — Pure Rust, no ML models, no external services. Single dependency on `lopdf` for PDF parsing.
- **Lightweight by default** — The default build is pure Rust with no ML models or external services. PDFium, ONNX Runtime, and OCR models are added only when the native `ocr` feature is selected and remain external runtime artifacts.
## Benchmark
@@ -160,6 +161,19 @@ detect-pdf document.pdf --json
detect-pdf document.pdf --analyze --json
```
OCR is a separate native CLI build and does not change the default package:
```bash
cargo install pdf-inspector --features ocr --bin pdf2md
PDFIUM_LIB_PATH=/path/to/libpdfium ORT_DYLIB_PATH=/path/to/libonnxruntime \
pdf2md scan.pdf --ocr auto --json
```
The OCR JSON envelope is versioned and reports routed pages, per-page source
and confidence, warnings, and pages recommended for the hosted document
pipeline. See the [Rust API guide](docs/rust-api.md#complete-ocr-api) for model
cache and offline configuration.
From a source checkout, use `cargo run --bin pdf2md -- document.pdf` or `cargo run --bin detect-pdf -- document.pdf` instead.
## Architecture
+49 -24
View File
@@ -1,6 +1,6 @@
# pdf-inspector
Fast PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. The default build is pure Rust, has no ML models or external services, and uses [lopdf](https://crates.io/crates/lopdf) for PDF parsing. Also available for [Python](https://pypi.org/project/pdf-inspector/) and [Node.js](https://www.npmjs.com/package/@firecrawl/pdf-inspector/).
Fast PDF classification and text extraction. The default build detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown without OCR. It is pure Rust, has no ML models or external services, and uses [lopdf](https://crates.io/crates/lopdf) for PDF parsing. Native Rust and CLI consumers can opt into selective OCR. Also available for [Python](https://pypi.org/project/pdf-inspector/) and [Node.js](https://www.npmjs.com/package/@firecrawl/pdf-inspector/).
Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.
@@ -123,10 +123,10 @@ The native-only `vision` feature exposes the stable seam used by OCR
integrations without selecting or embedding an inference runtime. The
separate `model-cache` feature adds pinned artifact management:
- `PageRenderer`, `OcrEngine`, and `LayoutEngine` traits;
- `PageRenderer` and `OcrEngine` traits;
- renderer-neutral owned page buffers and affine pixel↔PDF transforms;
- `OcrOptions` and opt-in `Off`/`Auto`/`Force` routing modes;
- positioned OCR/layout results and per-page provenance types; and
- positioned OCR results and per-page provenance types; and
- a versioned PP-OCRv6 Small manifest with checksum-verified, locked, atomic
model-cache installation and explicit offline-directory overrides.
@@ -135,9 +135,9 @@ separate `model-cache` feature adds pinned artifact management:
pdf-inspector = { version = "1", features = ["vision", "model-cache"] }
```
The OCR contracts preserve existing behavior by default: OCR is `Off`, learned
layout is disabled, and model resolution is never reached. `ModelStore` itself
does not access the network. The optional `model-download` feature provides an
The OCR contracts preserve existing behavior by default: OCR is `Off` and
model resolution is never reached. `ModelStore` itself does not access the
network. The optional `model-download` feature provides an
HTTPS downloader that streams pinned artifacts into the checksum-verified
cache only after routing has selected OCR work. Offline consumers set an
explicit model directory and `ModelDownloadPolicy::Offline`. Renderer-only
@@ -171,9 +171,10 @@ and `PdfiumRenderer` implements the renderer-neutral `PageRenderer` trait.
pdf-inspector = { version = "1", features = ["render-pdfium"] }
```
PDFium is loaded at runtime. Set `PDFIUM_LIB_PATH`, place its shared library
next to the executable, or use another discovery route supported by
`firecrawl-pdfium`.
PDFium is loaded at runtime and is not bundled into the crate. Set
`PDFIUM_LIB_PATH` to the platform shared library, place that library next to
the executable, or use another discovery route supported by
`firecrawl-pdfium`. A load failure reports this prerequisite directly.
```rust
use pdf_inspector::vision::{PdfiumRenderer, RenderOptions};
@@ -205,9 +206,11 @@ The native-only `ocr-oar` feature adds a CPU PP-OCRv6 Small implementation of
not enable model auto-download, ONNX Runtime download, or PDF rendering. Model
files remain external, must match the pinned manifest, and are opened only
after `ModelStore` verifies their exact size and SHA-256 digest. Install an
ONNX Runtime shared library separately and set `ORT_DYLIB_PATH` when it is not
available through the platform library search path. The feature currently
requires Rust 1.95 or newer, matching OAR 0.9.1's MSRV.
ONNX Runtime shared library separately and set `ORT_DYLIB_PATH` to its full
path when it is not available through the platform library search path. The
runtime is resolved only when an OCR engine is first constructed; clean
`Auto` requests do not require it. The feature currently requires Rust 1.95
or newer, matching OAR 0.9.1's MSRV.
```toml
[dependencies]
@@ -329,7 +332,10 @@ let fused = fuse_ocr_pages(
for page in &fused.pages {
println!("{}", page.markdown);
if page.provenance.hosted_recommended {
eprintln!("page {} needs the hosted document pipeline", page.page + 1);
eprintln!(
"page {} needs the hosted document pipeline",
page.page_number,
);
}
}
```
@@ -354,15 +360,11 @@ pdf-inspector = { version = "1", features = ["ocr"] }
```
```rust
use pdf_inspector::vision::{
process_pdf_with_ocr, OcrMode, OcrPdfOptions,
};
use pdf_inspector::vision::{process_pdf_with_ocr, OcrPdfOptions};
let result = process_pdf_with_ocr(
"document.pdf",
OcrPdfOptions::new()
.mode(OcrMode::Auto)
.pages([1, 2, 3]),
OcrPdfOptions::auto().page_numbers([1, 2, 3]),
)?;
println!("{}", result.markdown);
@@ -377,8 +379,28 @@ Native extraction always runs first. In `Auto`, a clean PDF returns before
PDFium loading, model-cache access, HTTP, or OAR initialization. Model files
remain external and the default crate feature set remains unchanged. `Off`
provides the same native-only behavior through the OCR result/provenance
shape; `Force` renders every selected page. Learned layout intentionally
returns an explicit unsupported error in this lightweight pipeline.
shape; `Force` renders every selected page. OCR uses the existing deterministic
table, column, reading-order, and Markdown assembly path; no learned layout
model is included.
For ambiguous mixed pages, `Auto` privately retains clean native fragments
instead of discarding them when OCR is selected. After recognition it compares
script-agnostic text quality, OCR confidence, character overlap, and material
new coverage. Exact native text wins over a duplicate or weak OCR hypothesis;
complementary image-backed text is fused; and pages where both candidates are
weak recommend the hosted document pipeline. A page routed because native
coverage appeared incomplete also recommends hosted processing when confident
OCR only duplicates the retained fragment: the agreement preserves trustworthy
text, but neither hypothesis proves full-page coverage. Public native-only
extraction continues to suppress pages marked unreliable, and clean text
documents pay no renderer or model-initialization cost.
In `Auto`, pages routed only for suspicious font encoding or vectorized text
first get a bounded positioned-text probe through PDFium. A credible recovered
text layer with sufficient geometric page coverage skips rasterization and
model loading for that page; garbled, partial, or insubstantial recovery
continues through OCR. Recovered tables are reflected in the same document
metadata as tables found by the primary extractor.
The one-call API keeps the most recently used verified OCR engine in process.
Long-lived workers therefore verify the pinned artifacts and build the ONNX
@@ -391,11 +413,13 @@ not hot-reload a running process; restart the process when intentionally
replacing files at the same paths. CPU inference uses at most four intra-op
threads per ONNX session so a single small page does not oversubscribe larger
hosts, and recognizes variable-width line crops individually to avoid
padding-heavy CPU batches.
padding-heavy CPU batches. The high-level pipeline renders and fuses at most
four routed pages at a time, bounding bitmap memory on long documents.
Build the CLI with the same opt-in feature:
```bash
cargo install pdf-inspector --features ocr --bin pdf2md
cargo build --release --features ocr --bin pdf2md
pdf2md document.pdf --ocr auto --raw
pdf2md document.pdf --ocr auto --json
@@ -404,10 +428,11 @@ pdf2md document.pdf --ocr auto --ocr-offline --ocr-model-dir /opt/models/pp-ocrv
CLI controls include `--ocr-dpi`, `--ocr-min-confidence`,
`--ocr-hosted-threshold`, `--select-pages`, and the existing encrypted-PDF
`--password` option. JSON output includes per-page Markdown, source/model
`--password` option. JSON output has `schema_version: 1` and includes per-page Markdown, source/model
provenance, confidence, timings, warnings, routed pages, and hosted-fallback
recommendations. Page numbers in `OcrPdfResult` and its per-page provenance
are 1-indexed, matching the PDF page numbers accepted by `OcrPdfOptions::pages`.
are 1-indexed, matching the PDF page numbers accepted by
`OcrPdfOptions::page_numbers`.
Extract per-page Markdown (one string per page, plus document-wide layout
metadata):
+61 -29
View File
@@ -165,8 +165,8 @@ fn format_ocr_json(result: &OcrPdfResult) -> String {
.collect::<Vec<_>>()
.join(",");
format!(
r#"{{"page":{},"source":"{}","markdown":"{}","ocr_model":{},"render_dpi":{},"ocr_confidence":{},"hosted_recommended":{},"timings":{{"render_ms":{},"ocr_ms":{},"layout_ms":{},"assembly_ms":{}}},"warnings":[{}]}}"#,
provenance.page,
r#"{{"page":{},"source":"{}","markdown":"{}","ocr_model":{},"render_dpi":{},"ocr_confidence":{},"hosted_recommended":{},"timings":{{"render_ms":{},"ocr_ms":{},"assembly_ms":{}}},"warnings":[{}]}}"#,
provenance.page_number,
source,
json_escape(&page.markdown),
model,
@@ -175,7 +175,6 @@ fn format_ocr_json(result: &OcrPdfResult) -> String {
provenance.hosted_recommended,
provenance.timings.render_ms,
provenance.timings.ocr_ms,
provenance.timings.layout_ms,
provenance.timings.assembly_ms,
warnings,
)
@@ -196,7 +195,7 @@ fn format_ocr_json(result: &OcrPdfResult) -> String {
.join(",");
let ocr_reasons = format_ocr_reasons_by_page(&result.ocr_reasons_by_page);
format!(
r#"{{"page_count":{},"processing_time_ms":{},"render_time_ms":{},"ocr_time_ms":{},"pages_recommended_for_ocr":[{}],"pages_routed_to_ocr":[{}],"pages_recommending_hosted":[{}],"ocr_reasons_by_page":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"pages":[{}],"markdown":"{}"}}"#,
r#"{{"schema_version":1,"page_count":{},"processing_time_ms":{},"render_time_ms":{},"ocr_time_ms":{},"pages_recommended_for_ocr":[{}],"pages_routed_to_ocr":[{}],"pages_recommending_hosted":[{}],"ocr_reasons_by_page":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"pages":[{}],"markdown":"{}"}}"#,
result.page_count,
result.processing_time_ms,
result.render_time_ms,
@@ -224,6 +223,19 @@ fn argument_value<'a>(args: &'a [String], name: &str) -> Result<Option<&'a str>,
.transpose()
}
fn format_ocr_error_json(error: &str) -> String {
format!(r#"{{"schema_version":1,"error":"{}"}}"#, json_escape(error))
}
fn exit_ocr_error(error: &str, json_output: bool) -> ! {
if json_output {
println!("{}", format_ocr_error_json(error));
} else {
eprintln!("Error: {error}");
}
process::exit(1);
}
#[cfg(all(feature = "ocr", not(target_arch = "wasm32")))]
fn float_argument(args: &[String], name: &str, default: f32) -> Result<f32, String> {
argument_value(args, name)?
@@ -247,7 +259,9 @@ fn extract_items_json(
#[cfg(test)]
mod tests {
use super::{extract_items_json, format_items_json};
use super::{extract_items_json, format_items_json, format_ocr_error_json};
#[cfg(all(feature = "ocr", not(target_arch = "wasm32")))]
use super::{format_ocr_json, process_pdf_with_ocr, OcrPdfOptions};
use pdf_inspector::extractor::ItemType;
use pdf_inspector::TextItem;
@@ -297,6 +311,27 @@ mod tests {
"decrypted item JSON should contain fixture text, got {json}"
);
}
#[cfg(all(feature = "ocr", not(target_arch = "wasm32")))]
#[test]
fn ocr_json_has_a_versioned_stable_envelope() {
let result =
process_pdf_with_ocr("tests/fixtures/thermo-freon12.pdf", OcrPdfOptions::new())
.unwrap();
let json = format_ocr_json(&result);
assert!(json.starts_with(r#"{"schema_version":1,"page_count":3,"#));
assert!(json.contains(r#""page":1,"source":"native""#));
assert!(!json.contains("layout_ms"));
}
#[test]
fn ocr_json_errors_use_the_same_versioned_envelope() {
assert_eq!(
format_ocr_error_json("bad \"value\""),
r#"{"schema_version":1,"error":"bad \"value\""}"#
);
}
}
/// Parse a page specification like "1,3,5-10,20" into a HashSet of page numbers.
@@ -441,23 +476,27 @@ fn main() {
.iter()
.any(|option| args.iter().any(|argument| argument == option));
if ocr_mode_argument.is_none() && has_ocr_only_option {
eprintln!("Error: OCR options require --ocr off, --ocr auto, or --ocr force");
process::exit(1);
exit_ocr_error(
"OCR options require --ocr off, --ocr auto, or --ocr force",
json_output,
);
}
if let Some(mode) = ocr_mode_argument {
if items_json_output || detect_only || analyze {
eprintln!(
"Error: --ocr cannot be combined with --items-json, --detect-only, or --analyze"
exit_ocr_error(
"--ocr cannot be combined with --items-json, --detect-only, or --analyze",
json_output,
);
process::exit(1);
}
#[cfg(not(all(feature = "ocr", not(target_arch = "wasm32"))))]
{
let _ = mode;
eprintln!("Error: this pdf2md build does not include OCR; rebuild with --features ocr");
process::exit(1);
exit_ocr_error(
"this pdf2md build does not include OCR; rebuild with --features ocr",
json_output,
);
}
#[cfg(all(feature = "ocr", not(target_arch = "wasm32")))]
@@ -467,28 +506,26 @@ fn main() {
"auto" => OcrMode::Auto,
"force" => OcrMode::Force,
value => {
eprintln!("Error: invalid --ocr mode {value:?}; expected off, auto, or force");
process::exit(1);
exit_ocr_error(
&format!("invalid --ocr mode {value:?}; expected off, auto, or force"),
json_output,
);
}
};
let dpi = float_argument(&args, "--ocr-dpi", 150.0).unwrap_or_else(|error| {
eprintln!("Error: {error}");
process::exit(1);
exit_ocr_error(&error, json_output);
});
let minimum_confidence = float_argument(&args, "--ocr-min-confidence", 0.0)
.unwrap_or_else(|error| {
eprintln!("Error: {error}");
process::exit(1);
exit_ocr_error(&error, json_output);
});
let hosted_threshold = float_argument(&args, "--ocr-hosted-threshold", 0.5)
.unwrap_or_else(|error| {
eprintln!("Error: {error}");
process::exit(1);
exit_ocr_error(&error, json_output);
});
let model_directory =
argument_value(&args, "--ocr-model-dir").unwrap_or_else(|error| {
eprintln!("Error: {error}");
process::exit(1);
exit_ocr_error(&error, json_output);
});
let mut ocr = OcrOptions::new()
@@ -511,7 +548,7 @@ fn main() {
.markdown(markdown)
.hosted_recommendation_confidence(hosted_threshold);
if let Some(pages) = page_filter.clone() {
pdf_options = pdf_options.pages(pages);
pdf_options = pdf_options.page_numbers(pages);
}
if let Some(password) = password.clone() {
pdf_options = pdf_options.password(password);
@@ -549,12 +586,7 @@ fn main() {
}
}
Err(error) => {
if json_output {
println!(r#"{{"error":"{}"}}"#, json_escape(&error.to_string()));
} else {
eprintln!("Error: {error}");
}
process::exit(1);
exit_ocr_error(&error.to_string(), json_output);
}
}
return;
+20 -3
View File
@@ -459,8 +459,15 @@ pub fn extract_pages_markdown_mem(
buffer: &[u8],
pages: Option<&[u32]>,
) -> Result<PagesExtractionResult, PdfError> {
extract_pages_markdown_mem_impl(buffer, pages, None, &MarkdownOptions::default(), false)
.map(|(result, _)| result)
extract_pages_markdown_mem_impl(
buffer,
pages,
None,
&MarkdownOptions::default(),
false,
false,
)
.map(|(result, _)| result)
}
#[cfg(all(feature = "ocr", not(target_arch = "wasm32")))]
@@ -476,6 +483,7 @@ pub(crate) fn extract_pages_markdown_mem_for_ocr(
password,
markdown_options,
markdown_options.strip_headers_footers,
true,
)
}
@@ -485,6 +493,7 @@ fn extract_pages_markdown_mem_impl(
password: Option<&str>,
markdown_options: &MarkdownOptions,
strip_repeated_headers_footers: bool,
preserve_ocr_candidates: bool,
) -> Result<(PagesExtractionResult, u32), PdfError> {
validate_pdf_bytes(buffer)?;
let (doc, page_count) = load_document_from_mem_with_password(buffer, password)?;
@@ -663,7 +672,15 @@ fn extract_pages_markdown_mem_impl(
results.push(PageMarkdown {
page: page_0idx,
markdown: if needs_ocr { String::new() } else { md },
// The public native extractor continues to suppress unreliable
// text. The OCR orchestrator retains clean partial text
// internally so it can compare/fuse it with OCR before deciding
// what is safe to return.
markdown: if needs_ocr && !preserve_ocr_candidates {
String::new()
} else {
md
},
needs_ocr,
ocr_reason,
});
+3 -152
View File
@@ -1,4 +1,4 @@
//! Public contracts between rendering, OCR, layout, and orchestration.
//! Public contracts between rendering, OCR, and orchestration.
use std::error::Error;
use std::path::PathBuf;
@@ -18,19 +18,6 @@ pub enum OcrMode {
Force,
}
/// Resource/quality profile for the OCR engine.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
#[non_exhaustive]
pub enum OcrProfile {
/// Lowest latency and memory footprint.
Edge,
/// OCR-oriented balance of quality and CPU cost.
#[default]
Balanced,
/// Highest quality within the lightweight model family.
Quality,
}
/// Controls whether missing model artifacts may be fetched.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
#[non_exhaustive]
@@ -47,12 +34,8 @@ pub enum ModelDownloadPolicy {
pub struct OcrOptions {
/// Page-level routing behavior.
pub mode: OcrMode,
/// Local quality/resource profile.
pub profile: OcrProfile,
/// Drop recognition spans below this confidence threshold.
pub minimum_confidence: f32,
/// Optional language hints understood by the selected engine.
pub languages: Vec<String>,
/// Optional directory containing an offline model set.
pub model_directory: Option<PathBuf>,
/// Whether a missing pinned artifact may be downloaded.
@@ -63,9 +46,7 @@ impl Default for OcrOptions {
fn default() -> Self {
Self {
mode: OcrMode::Off,
profile: OcrProfile::Balanced,
minimum_confidence: 0.0,
languages: Vec::new(),
model_directory: None,
model_downloads: ModelDownloadPolicy::IfMissing,
}
@@ -84,24 +65,12 @@ impl OcrOptions {
self
}
/// Sets the local resource/quality profile.
pub fn profile(mut self, profile: OcrProfile) -> Self {
self.profile = profile;
self
}
/// Sets the minimum accepted recognition confidence.
pub fn minimum_confidence(mut self, minimum_confidence: f32) -> Self {
self.minimum_confidence = minimum_confidence;
self
}
/// Replaces the language hints passed to the OCR engine.
pub fn languages(mut self, languages: impl IntoIterator<Item = impl Into<String>>) -> Self {
self.languages = languages.into_iter().map(Into::into).collect();
self
}
/// Uses an explicit model directory, suitable for offline packaging.
pub fn model_directory(mut self, directory: impl Into<PathBuf>) -> Self {
self.model_directory = Some(directory.into());
@@ -115,55 +84,6 @@ impl OcrOptions {
}
}
/// Configuration for an optional learned layout engine.
///
/// Layout inference is disabled by default. Existing deterministic layout,
/// table, and Markdown logic remains the assembly path when this is disabled.
#[derive(Debug, Clone, PartialEq)]
pub struct LayoutOptions {
/// Whether the learned layout extension may run.
pub enabled: bool,
/// Drop layout regions below this confidence threshold.
pub minimum_confidence: f32,
/// Optional directory containing an offline layout model set.
pub model_directory: Option<PathBuf>,
}
impl Default for LayoutOptions {
fn default() -> Self {
Self {
enabled: false,
minimum_confidence: 0.0,
model_directory: None,
}
}
}
impl LayoutOptions {
/// Creates layout options with learned layout disabled.
pub fn new() -> Self {
Self::default()
}
/// Enables or disables learned layout inference.
pub fn enabled(mut self, enabled: bool) -> Self {
self.enabled = enabled;
self
}
/// Sets the minimum accepted region confidence.
pub fn minimum_confidence(mut self, minimum_confidence: f32) -> Self {
self.minimum_confidence = minimum_confidence;
self
}
/// Uses an explicit layout model directory.
pub fn model_directory(mut self, directory: impl Into<PathBuf>) -> Self {
self.model_directory = Some(directory.into());
self
}
}
/// A point in bitmap space, measured from the top-left in pixels.
#[derive(Debug, Clone, Copy, Default, PartialEq)]
pub struct ImagePoint {
@@ -230,7 +150,7 @@ pub struct OcrSpan {
#[derive(Debug, Clone, PartialEq)]
pub struct OcrPage {
/// 1-indexed PDF page number.
pub page: u32,
pub page_number: u32,
/// Positioned recognition spans.
pub spans: Vec<OcrSpan>,
/// Mean confidence across accepted spans, when available.
@@ -243,54 +163,6 @@ pub struct OcrPage {
pub warnings: Vec<String>,
}
/// Normalized semantic class emitted by a learned layout engine.
#[derive(Debug, Clone, PartialEq, Eq)]
#[non_exhaustive]
pub enum LayoutRegionKind {
/// Body or other prose text.
Text,
/// Document heading or title.
Heading,
/// Table region.
Table,
/// Figure/image region.
Figure,
/// Figure or table caption.
Caption,
/// Header/footer/page furniture.
Furniture,
/// Model-specific class retained without changing the common taxonomy.
Other(String),
}
/// One learned layout region in bitmap coordinates.
#[derive(Debug, Clone, PartialEq)]
pub struct LayoutRegion {
/// Normalized semantic class.
pub kind: LayoutRegionKind,
/// Region polygon in the original rendered page's pixel space.
pub polygon: ImageQuad,
/// Model confidence in the inclusive range 01.
pub confidence: f32,
/// Optional model-provided reading-order position.
pub reading_order: Option<u32>,
}
/// Learned layout output for one 1-indexed page.
#[derive(Debug, Clone, PartialEq)]
pub struct LayoutPage {
/// 1-indexed PDF page number.
pub page: u32,
/// Semantic regions.
pub regions: Vec<LayoutRegion>,
/// Exact model identity used for this result.
pub model: ModelIdentity,
/// Layout inference wall time for this page.
pub processing_time_ms: u64,
/// Non-fatal engine warnings.
pub warnings: Vec<String>,
}
/// How final page content was sourced.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
#[non_exhaustive]
@@ -310,8 +182,6 @@ pub struct VisionTimings {
pub render_ms: u64,
/// OCR wall time.
pub ocr_ms: u64,
/// Optional learned layout wall time.
pub layout_ms: u64,
/// Native/OCR fusion and assembly wall time.
pub assembly_ms: u64,
}
@@ -320,13 +190,11 @@ pub struct VisionTimings {
#[derive(Debug, Clone, PartialEq)]
pub struct PageProvenance {
/// 1-indexed PDF page number.
pub page: u32,
pub page_number: u32,
/// Final page-content source.
pub source: PageContentSource,
/// OCR model, when OCR ran.
pub ocr_model: Option<ModelIdentity>,
/// Learned layout model, when layout inference ran.
pub layout_model: Option<ModelIdentity>,
/// Render resolution used for local vision.
pub render_dpi: Option<f32>,
/// Mean accepted OCR confidence, when available.
@@ -371,23 +239,6 @@ pub trait OcrEngine: Send + Sync {
) -> Result<Vec<OcrPage>, Self::Error>;
}
/// Optional learned semantic layout extension.
pub trait LayoutEngine: Send + Sync {
/// Engine-specific failure type.
type Error: Error + Send + Sync + 'static;
/// Exact model identity used by this engine instance.
fn model(&self) -> &ModelIdentity;
/// Analyzes rendered pages, optionally using their OCR spans.
fn analyze(
&self,
pages: &[RenderedPage],
ocr: &[OcrPage],
options: &LayoutOptions,
) -> Result<Vec<LayoutPage>, Self::Error>;
}
#[cfg(test)]
mod tests {
use super::*;
+416 -14
View File
@@ -6,6 +6,7 @@ use std::time::Instant;
use thiserror::Error;
use crate::markdown::{to_markdown_from_items_with_rects_and_page_count, MarkdownOptions};
use crate::text_quality::{detect_encoding_issues, is_cid_garbage, is_garbage_text};
use crate::types::{ItemType, TextItem};
use crate::PageMarkdown;
@@ -73,7 +74,7 @@ impl OcrFusionOptions {
#[derive(Debug, Clone, PartialEq)]
pub struct FusedPageMarkdown {
/// 1-indexed document page number, matching OCR and provenance fields.
pub page: u32,
pub page_number: u32,
/// Final page Markdown.
pub markdown: String,
/// Native/OCR source, model, timing, and fallback metadata.
@@ -91,6 +92,73 @@ pub struct FusedPages {
pub ocr_time_ms: u64,
}
/// Origin of a trustworthy native-text candidate retained for adaptive OCR.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub(crate) enum NativeCandidateOrigin {
/// Text produced by pdf-inspector's normal native extractor.
Extractor,
/// Positioned text independently recovered through PDFium.
Pdfium,
}
impl NativeCandidateOrigin {
fn description(self) -> &'static str {
match self {
Self::Extractor => "native extraction",
Self::Pdfium => "PDFium native recovery",
}
}
}
#[derive(Debug, Clone, Copy)]
struct TextCandidateQuality {
alphanumeric_chars: usize,
score: f32,
}
/// Clean native text retained while an ambiguous page is compared with OCR.
#[derive(Debug, Clone)]
pub(crate) struct NativeFallbackCandidate {
markdown: String,
quality: TextCandidateQuality,
origin: NativeCandidateOrigin,
}
impl NativeFallbackCandidate {
/// True when an independent native recovery is substantial enough to
/// cancel OCR for recoverable font/vector routing reasons.
pub(crate) fn is_complete_recovery(&self) -> bool {
self.quality.alphanumeric_chars >= 40 && self.quality.score >= 0.68
}
pub(crate) fn markdown(&self) -> &str {
&self.markdown
}
pub(crate) fn is_stronger_than(&self, other: &Self) -> bool {
self.quality.alphanumeric_chars > other.quality.alphanumeric_chars
|| (self.quality.alphanumeric_chars == other.quality.alphanumeric_chars
&& self.quality.score > other.quality.score)
}
}
/// Validates and scores native Markdown for possible post-OCR comparison.
///
/// This intentionally uses script-agnostic evidence. A native candidate only
/// needs to be trustworthy, not necessarily complete: a clean native header
/// can still be fused with an image-backed OCR body.
pub(crate) fn assess_native_candidate(
markdown: String,
origin: NativeCandidateOrigin,
) -> Option<NativeFallbackCandidate> {
let quality = assess_text_candidate(&markdown)?;
(quality.alphanumeric_chars >= 8).then_some(NativeFallbackCandidate {
markdown,
quality,
origin,
})
}
/// Converts positioned OCR spans to Markdown through pdf-inspector's existing
/// deterministic geometry, reading-order, table, and Markdown pipeline.
///
@@ -116,13 +184,47 @@ pub fn ocr_page_to_markdown(
/// OCR replaces pages whose native extraction was already rejected. On clean
/// native pages (for example in `Force` mode), normalized duplicate OCR blocks
/// are removed and only genuinely additional blocks are appended. Pages that
/// needed OCR but still have no credible local result recommend the hosted
/// document pipeline instead of silently presenting an empty result as final.
/// needed OCR but still have no credible OCR result, or whose confident OCR
/// only repeats an incomplete native fragment, recommend the hosted document
/// pipeline instead of silently presenting partial content as final.
pub fn fuse_ocr_pages(
native_pages: &[PageMarkdown],
ocr_run: &OcrRun,
document_page_count: u32,
options: &OcrFusionOptions,
) -> Result<FusedPages, OcrFusionError> {
fuse_ocr_pages_impl(
native_pages,
ocr_run,
document_page_count,
options,
&BTreeMap::new(),
)
}
/// OCR-pipeline fusion with trustworthy partial native candidates.
pub(crate) fn fuse_ocr_pages_adaptive(
native_pages: &[PageMarkdown],
ocr_run: &OcrRun,
document_page_count: u32,
options: &OcrFusionOptions,
native_candidates: &BTreeMap<u32, NativeFallbackCandidate>,
) -> Result<FusedPages, OcrFusionError> {
fuse_ocr_pages_impl(
native_pages,
ocr_run,
document_page_count,
options,
native_candidates,
)
}
fn fuse_ocr_pages_impl(
native_pages: &[PageMarkdown],
ocr_run: &OcrRun,
document_page_count: u32,
options: &OcrFusionOptions,
native_candidates: &BTreeMap<u32, NativeFallbackCandidate>,
) -> Result<FusedPages, OcrFusionError> {
validate_options(options)?;
@@ -178,16 +280,31 @@ pub fn fuse_ocr_pages(
document_page_count,
);
let ocr_markdown = preserve_ocr_line_breaks(&ocr_markdown, local);
let (markdown, source) = if native.markdown.trim().is_empty() || native.needs_ocr {
(ocr_markdown, PageContentSource::Ocr)
let (markdown, source, adaptive_recommends_hosted) = if let Some(candidate) =
native_candidates
.get(&page_number)
.filter(|_| native.needs_ocr)
{
let choice = choose_adaptive_content(
candidate,
&ocr_markdown,
local.ocr.mean_confidence,
options.hosted_recommendation_confidence,
);
warnings.push(choice.warning);
(choice.markdown, choice.source, choice.recommend_hosted)
} else if native.markdown.trim().is_empty() || native.needs_ocr {
(ocr_markdown, PageContentSource::Ocr, false)
} else {
merge_native_and_ocr(&native.markdown, &ocr_markdown)
let (markdown, source) = merge_native_and_ocr(&native.markdown, &ocr_markdown);
(markdown, source, false)
};
let weak_ocr = local
.ocr
.mean_confidence
.is_none_or(|confidence| confidence < options.hosted_recommendation_confidence);
let recommend_hosted = native.needs_ocr && (markdown.trim().is_empty() || weak_ocr);
let recommend_hosted = native.needs_ocr
&& (markdown.trim().is_empty() || weak_ocr || adaptive_recommends_hosted);
if native.needs_ocr && markdown.trim().is_empty() {
warnings.push("OCR produced no usable text".to_string());
}
@@ -214,13 +331,12 @@ pub fn fuse_ocr_pages(
};
pages.push(FusedPageMarkdown {
page: page_number,
page_number,
markdown,
provenance: PageProvenance {
page: page_number,
page_number,
source,
ocr_model,
layout_model: None,
render_dpi: ocr_by_page
.contains_key(&page_number)
.then_some(options.render_dpi),
@@ -228,7 +344,6 @@ pub fn fuse_ocr_pages(
timings: VisionTimings {
render_ms: render_by_page.get(&page_number).copied().unwrap_or(0),
ocr_ms,
layout_ms: 0,
assembly_ms: elapsed_ms(assembly_started),
},
warnings,
@@ -244,6 +359,159 @@ pub fn fuse_ocr_pages(
})
}
struct AdaptiveContentChoice {
markdown: String,
source: PageContentSource,
warning: String,
recommend_hosted: bool,
}
fn choose_adaptive_content(
native: &NativeFallbackCandidate,
ocr: &str,
ocr_confidence: Option<f32>,
weak_ocr_threshold: f32,
) -> AdaptiveContentChoice {
let ocr_quality = assess_text_candidate(ocr);
let weak_ocr = ocr_confidence.is_none_or(|confidence| confidence < weak_ocr_threshold);
if weak_ocr || ocr_quality.is_none() {
return AdaptiveContentChoice {
markdown: ensure_trailing_newline(native.markdown()),
source: PageContentSource::Native,
warning: format!(
"kept trustworthy {} because OCR was weak or unusable",
native.origin.description()
),
recommend_hosted: true,
};
}
let ocr_quality = ocr_quality.expect("checked above");
let overlap = content_overlap(native.markdown(), ocr);
let ocr_novel_chars = ocr_quality
.alphanumeric_chars
.saturating_sub(overlap.shared_chars);
let material_novelty =
ocr_novel_chars >= 2 && ocr_novel_chars * 8 >= ocr_quality.alphanumeric_chars.max(1);
let native_substantially_covered =
overlap.shared_chars * 4 >= native.quality.alphanumeric_chars.max(1) * 3;
if native_substantially_covered && !material_novelty {
return AdaptiveContentChoice {
markdown: ensure_trailing_newline(native.markdown()),
source: PageContentSource::Native,
warning: format!(
"kept trustworthy {} because OCR added no material coverage",
native.origin.description()
),
// The candidate exists only because this page was routed with
// incomplete native coverage. Agreement between two partial
// hypotheses preserves trustworthy text, but does not prove that
// the rest of the page was recovered.
recommend_hosted: true,
};
}
// A shorter, materially lower-quality OCR hypothesis should not displace
// exact native text even when its mean confidence happens to be high.
if ocr_quality.score + 0.12 < native.quality.score
&& ocr_quality.alphanumeric_chars * 10
<= native.quality.alphanumeric_chars.saturating_mul(11)
{
return AdaptiveContentChoice {
markdown: ensure_trailing_newline(native.markdown()),
source: PageContentSource::Native,
warning: format!(
"kept higher-quality {} after comparing OCR",
native.origin.description()
),
recommend_hosted: false,
};
}
let (markdown, source) = merge_native_and_ocr(native.markdown(), ocr);
let warning = match source {
PageContentSource::Native => format!(
"kept trustworthy {} because OCR duplicated its content",
native.origin.description()
),
PageContentSource::Fused => format!(
"fused trustworthy {} with complementary OCR",
native.origin.description()
),
PageContentSource::Ocr => unreachable!("merge never returns OCR-only content"),
};
AdaptiveContentChoice {
markdown,
source,
warning,
recommend_hosted: false,
}
}
#[derive(Debug, Clone, Copy)]
struct ContentOverlap {
shared_chars: usize,
}
fn content_overlap(first: &str, second: &str) -> ContentOverlap {
let mut first_counts = BTreeMap::<char, usize>::new();
for character in normalized_content_chars(first) {
*first_counts.entry(character).or_insert(0) += 1;
}
let mut second_counts = BTreeMap::<char, usize>::new();
for character in normalized_content_chars(second) {
*second_counts.entry(character).or_insert(0) += 1;
}
let shared_chars = first_counts
.iter()
.map(|(character, count)| (*count).min(second_counts.get(character).copied().unwrap_or(0)))
.sum();
ContentOverlap { shared_chars }
}
fn normalized_content_chars(text: &str) -> impl Iterator<Item = char> + '_ {
text.chars()
.flat_map(char::to_lowercase)
.filter(|character| character.is_alphanumeric())
}
fn assess_text_candidate(markdown: &str) -> Option<TextCandidateQuality> {
if markdown.trim().is_empty()
|| is_garbage_text(markdown)
|| is_cid_garbage(markdown)
|| detect_encoding_issues(markdown)
{
return None;
}
let alphanumeric_chars = markdown
.chars()
.filter(|character| character.is_alphanumeric())
.count();
if alphanumeric_chars == 0 {
return None;
}
let visible_chars = markdown
.chars()
.filter(|character| !character.is_whitespace())
.count()
.max(1);
let density = alphanumeric_chars as f32 / visible_chars as f32;
let length_score = (alphanumeric_chars as f32 / 160.0).min(1.0);
let nonempty_lines = markdown
.lines()
.filter(|line| !line.trim().is_empty())
.count()
.max(1);
let line_score = (alphanumeric_chars as f32 / nonempty_lines as f32 / 12.0).min(1.0);
let score = (0.45 + length_score * 0.25 + density * 0.20 + line_score * 0.10).min(1.0);
Some(TextCandidateQuality {
alphanumeric_chars,
score,
})
}
/// Converts recognized line polygons to ordinary PDF-space text items.
fn ocr_text_items(page: &RoutedOcrPage) -> (Vec<TextItem>, usize) {
let mut discarded = 0usize;
@@ -746,7 +1014,7 @@ mod tests {
RoutedOcrPage {
rendered: rendered_page(page),
ocr: OcrPage {
page,
page_number: page,
spans,
mean_confidence: confidence,
model: ModelIdentity::new("test-ocr", "v1"),
@@ -778,6 +1046,10 @@ mod tests {
}
}
fn native_candidate(markdown: &str) -> NativeFallbackCandidate {
assess_native_candidate(markdown.to_string(), NativeCandidateOrigin::Extractor).unwrap()
}
#[test]
fn scanned_page_uses_geometry_ordered_ocr_and_provenance() {
let native = [native(0, "", true)];
@@ -797,8 +1069,11 @@ mod tests {
< result.pages[0].markdown.find("Second").unwrap()
);
assert_eq!(result.pages[0].provenance.source, PageContentSource::Ocr);
assert_eq!(result.pages[0].page, 1);
assert_eq!(result.pages[0].page, result.pages[0].provenance.page);
assert_eq!(result.pages[0].page_number, 1);
assert_eq!(
result.pages[0].page_number,
result.pages[0].provenance.page_number
);
assert_eq!(
result.pages[0].provenance.ocr_model.as_ref().unwrap().name,
"test-ocr"
@@ -930,6 +1205,133 @@ mod tests {
assert_eq!(result.pages[0].provenance.source, PageContentSource::Native);
}
#[test]
fn adaptive_fallback_keeps_native_text_when_ocr_is_weak() {
let native = [native(0, "", true)];
let run = run(vec![routed_page(
1,
vec![span("Inv0ice total uncertain", 10.0, 0.3)],
Some(0.3),
)]);
let candidates = BTreeMap::from([(
1,
native_candidate("Invoice total: $420.00\nPayment received\n"),
)]);
let result =
fuse_ocr_pages_adaptive(&native, &run, 1, &OcrFusionOptions::new(), &candidates)
.unwrap();
assert_eq!(result.pages[0].provenance.source, PageContentSource::Native);
assert_eq!(
result.pages[0].markdown,
"Invoice total: $420.00\nPayment received\n"
);
assert!(result.pages[0].provenance.hosted_recommended);
assert!(result.pages[0].provenance.warnings[0].contains("OCR was weak"));
}
#[test]
fn adaptive_fallback_prefers_exact_native_text_over_duplicate_ocr() {
let native = [native(0, "", true)];
let run = run(vec![routed_page(
1,
vec![span("Invoice total 420.00 Payment received", 10.0, 0.98)],
Some(0.98),
)]);
let candidates = BTreeMap::from([(
1,
native_candidate("Invoice total: $420.00\nPayment received\n"),
)]);
let result =
fuse_ocr_pages_adaptive(&native, &run, 1, &OcrFusionOptions::new(), &candidates)
.unwrap();
assert_eq!(result.pages[0].provenance.source, PageContentSource::Native);
assert!(result.pages[0].provenance.hosted_recommended);
assert_eq!(result.pages[0].markdown.matches("Invoice").count(), 1);
assert!(result.pages[0].provenance.warnings[0].contains("no material coverage"));
}
#[test]
fn adaptive_fallback_fuses_short_novel_ocr_content() {
let native = [native(0, "", true)];
let run = run(vec![routed_page(
1,
vec![span("Status ready 42", 10.0, 0.98)],
Some(0.98),
)]);
let candidates = BTreeMap::from([(1, native_candidate("Status ready\n"))]);
let result =
fuse_ocr_pages_adaptive(&native, &run, 1, &OcrFusionOptions::new(), &candidates)
.unwrap();
assert_eq!(result.pages[0].provenance.source, PageContentSource::Fused);
assert!(result.pages[0].markdown.contains("42"));
assert!(!result.pages[0].provenance.hosted_recommended);
}
#[test]
fn adaptive_fallback_recommends_hosted_for_high_confidence_garbage_ocr() {
let native = [native(0, "", true)];
let garbage = "a@@b%%c&&d==e~~".repeat(12);
let run = run(vec![routed_page(
1,
vec![span(&garbage, 10.0, 0.99)],
Some(0.99),
)]);
let candidates = BTreeMap::from([(1, native_candidate("Invoice total 420\n"))]);
let result =
fuse_ocr_pages_adaptive(&native, &run, 1, &OcrFusionOptions::new(), &candidates)
.unwrap();
assert_eq!(result.pages[0].provenance.source, PageContentSource::Native);
assert!(result.pages[0].provenance.hosted_recommended);
assert!(!result.pages[0].markdown.contains("@@"));
}
#[test]
fn adaptive_fallback_fuses_native_header_with_scanned_body() {
let native = [native(0, "", true)];
let run = run(vec![routed_page(
1,
vec![
span("Quarterly account report", 10.0, 0.96),
span("March revenue 420 units", 30.0, 0.96),
span("April revenue 510 units", 50.0, 0.96),
],
Some(0.96),
)]);
let candidates = BTreeMap::from([(1, native_candidate("# Quarterly account report\n"))]);
let result =
fuse_ocr_pages_adaptive(&native, &run, 1, &OcrFusionOptions::new(), &candidates)
.unwrap();
assert_eq!(result.pages[0].provenance.source, PageContentSource::Fused);
assert_eq!(result.pages[0].markdown.matches("Quarterly").count(), 1);
assert!(result.pages[0].markdown.contains("March revenue 420 units"));
assert!(result.pages[0].markdown.contains("April revenue 510 units"));
assert!(result.pages[0].provenance.warnings[0].contains("complementary"));
}
#[test]
fn native_candidate_scoring_is_multilingual_and_rejects_garbage() {
assert!(assess_native_candidate(
"請求書 合計金額 4200円\n支払済み\n".to_string(),
NativeCandidateOrigin::Extractor,
)
.is_some());
assert!(assess_native_candidate(
"a@@b%%c&&d==e~~".repeat(12),
NativeCandidateOrigin::Extractor,
)
.is_none());
}
#[test]
fn force_mode_appends_only_additional_ocr_blocks() {
let native = [native(0, "Native title\n", false)];
+2 -3
View File
@@ -30,9 +30,8 @@ mod pdfium;
#[cfg(all(feature = "vision", not(target_arch = "wasm32")))]
pub use contracts::{
ImagePoint, ImageQuad, LayoutEngine, LayoutOptions, LayoutPage, LayoutRegion, LayoutRegionKind,
ModelDownloadPolicy, ModelIdentity, OcrEngine, OcrMode, OcrOptions, OcrPage, OcrProfile,
OcrSpan, PageContentSource, PageProvenance, PageRenderer, VisionTimings,
ImagePoint, ImageQuad, ModelDownloadPolicy, ModelIdentity, OcrEngine, OcrMode, OcrOptions,
OcrPage, OcrSpan, PageContentSource, PageProvenance, PageRenderer, VisionTimings,
};
#[cfg(all(feature = "model-download", not(target_arch = "wasm32")))]
pub use download::{HttpModelDownloadError, HttpModelDownloader, DEFAULT_MODEL_DOWNLOAD_TIMEOUT};
+4 -6
View File
@@ -49,7 +49,9 @@ pub enum OarOcrError {
page: u32,
},
/// The external ONNX Runtime shared library could not be loaded.
#[error("failed to load ONNX Runtime from {path}: {source}")]
#[error(
"failed to load ONNX Runtime from {path}; install a compatible ONNX Runtime shared library or set ORT_DYLIB_PATH to its path: {source}"
)]
OnnxRuntimeLoad {
/// Requested shared-library path or platform library name.
path: PathBuf,
@@ -143,10 +145,6 @@ impl OarOcrEngine {
}
let mut warnings = Vec::new();
if !options.languages.is_empty() {
warnings
.push("language hints are not used by the PP-OCRv6 Small OAR backend".to_string());
}
if missing_recognition > 0 {
warnings.push(format!(
"discarded {missing_recognition} regions without usable recognition output"
@@ -166,7 +164,7 @@ impl OarOcrEngine {
let processing_time_ms = u64::try_from(started.elapsed().as_millis()).unwrap_or(u64::MAX);
Ok(OcrPage {
page: page.page(),
page_number: page.page(),
spans,
mean_confidence,
model: self.model.clone(),
+215 -3
View File
@@ -2,9 +2,11 @@
use std::path::Path;
use firecrawl_pdfium::{Pdfium, PixelFormat, PixelPoint, RenderConfig};
use firecrawl_pdfium::{PageChar, Pdfium, PixelFormat, PixelPoint, RenderConfig};
use thiserror::Error;
use crate::types::{ItemType, TextItem};
use super::{
PageRenderer, PageTransform, RenderBufferError, RenderOptions, RenderPixelFormat, RenderedPage,
};
@@ -47,6 +49,15 @@ pub enum RenderError {
/// Number of pages in the document.
page_count: usize,
},
/// The PDFium shared library could not be discovered or loaded.
#[error(
"failed to load PDFium; install a compatible PDFium shared library or set PDFIUM_LIB_PATH to its path"
)]
PdfiumLoad {
/// Dynamic loading failure.
#[source]
source: firecrawl_pdfium::Error,
},
/// PDFium loading, document parsing, form setup, or rendering failed.
#[error(transparent)]
Pdfium(#[from] firecrawl_pdfium::Error),
@@ -65,18 +76,28 @@ pub struct PdfiumRenderer {
pdfium: Pdfium,
}
/// Positioned native text recovered from one selected PDF page.
#[derive(Debug)]
pub(crate) struct PdfiumTextPage {
pub(crate) page: u32,
pub(crate) page_width: f32,
pub(crate) page_height: f32,
pub(crate) items: Vec<TextItem>,
}
impl PdfiumRenderer {
/// Loads PDFium using `firecrawl-pdfium`'s documented discovery chain.
pub fn load() -> Result<Self, RenderError> {
Ok(Self {
pdfium: Pdfium::load()?,
pdfium: Pdfium::load().map_err(|source| RenderError::PdfiumLoad { source })?,
})
}
/// Loads PDFium from an explicit native library path.
pub fn load_from_path(path: impl AsRef<Path>) -> Result<Self, RenderError> {
Ok(Self {
pdfium: Pdfium::load_from_path(path)?,
pdfium: Pdfium::load_from_path(path)
.map_err(|source| RenderError::PdfiumLoad { source })?,
})
}
@@ -100,6 +121,56 @@ impl PdfiumRenderer {
self.render_pages_impl(pdf_bytes, pages, password, options)
}
/// Extracts positioned native text from selected 1-indexed pages.
///
/// This is deliberately separate from rendering: callers can probe a
/// suspicious embedded text layer before paying for rasterization and
/// OCR. A page-level text failure is treated as an unavailable recovery
/// candidate so the caller can continue to its normal OCR fallback.
pub(crate) fn extract_text_pages(
&self,
pdf_bytes: &[u8],
pages: &[u32],
password: Option<&str>,
) -> Result<Vec<PdfiumTextPage>, RenderError> {
const MAX_TEXT_CHARS_PER_PAGE: usize = 250_000;
if pages.is_empty() {
return Ok(Vec::new());
}
if pages.contains(&0) {
return Err(RenderError::InvalidPageNumber);
}
let document = self.pdfium.load_document(pdf_bytes.to_vec(), password)?;
let page_count = document.page_count();
if let Some(&page) = pages.iter().find(|&&page| page as usize > page_count) {
return Err(RenderError::PageOutOfBounds { page, page_count });
}
let mut recovered = Vec::with_capacity(pages.len());
for &page_number in pages {
let page = document.page(page_number as usize - 1)?;
let page_size = page.size();
let text = match page.text_with_limit(MAX_TEXT_CHARS_PER_PAGE) {
Ok(text) => text,
Err(error) => {
log::debug!(
"page {page_number}: positioned native text recovery unavailable: {error}"
);
continue;
}
};
recovered.push(PdfiumTextPage {
page: page_number,
page_width: page_size.width,
page_height: page_size.height,
items: text_chars_to_items(text.chars(), page_number),
});
}
Ok(recovered)
}
fn render_pages_impl(
&self,
pdf_bytes: &[u8],
@@ -145,6 +216,106 @@ impl PdfiumRenderer {
}
}
fn text_chars_to_items(chars: &[PageChar], page: u32) -> Vec<TextItem> {
#[derive(Debug, Clone, Copy)]
struct Bounds {
left: f64,
bottom: f64,
right: f64,
top: f64,
}
fn flush(items: &mut Vec<TextItem>, text: &mut String, bounds: &mut Option<Bounds>, page: u32) {
let Some(bounds) = bounds.take() else {
text.clear();
return;
};
if text.is_empty() {
return;
}
let width = (bounds.right - bounds.left) as f32;
let height = (bounds.top - bounds.bottom) as f32;
let x = bounds.left as f32;
let y = bounds.bottom as f32;
if !x.is_finite()
|| !y.is_finite()
|| !width.is_finite()
|| !height.is_finite()
|| width <= 0.0
|| height <= 0.0
{
text.clear();
return;
}
items.push(TextItem {
text: std::mem::take(text),
x,
y,
width,
height,
font: "PDFium native text".to_string(),
font_size: height.max(1.0),
page,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
});
}
let mut items = Vec::new();
let mut text = String::new();
let mut bounds: Option<Bounds> = None;
for character in chars {
let Some(value) = character.unicode else {
flush(&mut items, &mut text, &mut bounds, page);
continue;
};
if value.is_whitespace() {
flush(&mut items, &mut text, &mut bounds, page);
continue;
}
let rect = character.loose_bounds.normalized();
if !rect.left.is_finite()
|| !rect.bottom.is_finite()
|| !rect.right.is_finite()
|| !rect.top.is_finite()
|| rect.width() <= 0.0
|| rect.height() <= 0.0
{
flush(&mut items, &mut text, &mut bounds, page);
continue;
}
text.push(value);
bounds = Some(match bounds {
Some(bounds) => Bounds {
left: bounds.left.min(rect.left),
bottom: bounds.bottom.min(rect.bottom),
right: bounds.right.max(rect.right),
top: bounds.top.max(rect.top),
},
None => Bounds {
left: rect.left,
bottom: rect.bottom,
right: rect.right,
top: rect.top,
},
});
}
flush(&mut items, &mut text, &mut bounds, page);
items.sort_by(|first, second| {
first
.page
.cmp(&second.page)
.then(second.y.total_cmp(&first.y))
.then(first.x.total_cmp(&second.x))
});
items
}
impl PageRenderer for PdfiumRenderer {
type Error = RenderError;
@@ -237,6 +408,17 @@ fn bgr_to_rgb_in_place(
#[cfg(test)]
mod tests {
use super::*;
use firecrawl_pdfium::{PagePoint, PageRect};
fn page_char(value: char, bounds: PageRect) -> PageChar {
PageChar {
unicode: Some(value),
code: value as u32,
bounds,
loose_bounds: bounds,
origin: PagePoint::new(bounds.left, bounds.bottom),
}
}
#[test]
fn bgr_pixels_are_converted_to_rgb_in_place() {
@@ -263,4 +445,34 @@ mod tests {
Err(RenderBufferError::InvalidBufferLength { .. })
));
}
#[test]
fn invalid_character_geometry_splits_text_runs() {
let chars = [
page_char('A', PageRect::new(0.0, 0.0, 8.0, 10.0)),
page_char('X', PageRect::new(10.0, 0.0, 10.0, 10.0)),
page_char('B', PageRect::new(20.0, 0.0, 28.0, 10.0)),
];
let items = text_chars_to_items(&chars, 1);
assert_eq!(
items
.iter()
.map(|item| item.text.as_str())
.collect::<Vec<_>>(),
["A", "B"]
);
}
#[test]
fn coordinates_that_overflow_f32_are_discarded() {
let left = f64::from(f32::MAX) * 2.0;
let chars = [page_char(
'A',
PageRect::new(left, 0.0, left + 1.0e30, 10.0),
)];
assert!(text_chars_to_items(&chars, 1).is_empty());
}
}
+851 -41
View File
File diff suppressed because it is too large Load Diff
+6 -2
View File
@@ -106,7 +106,11 @@ where
source: Box::new(source),
})?;
let ocr_time_ms = elapsed_ms(ocr_started);
validate_page_order("OCR engine", pages, recognized.iter().map(|page| page.page))?;
validate_page_order(
"OCR engine",
pages,
recognized.iter().map(|page| page.page_number),
)?;
Ok(OcrRun {
pages: rendered
@@ -268,7 +272,7 @@ mod tests {
Ok(pages
.iter()
.map(|page| OcrPage {
page: page.page(),
page_number: page.page(),
spans: vec![OcrSpan {
text: format!("page {}", page.page()),
polygon: ImageQuad::new([
+1 -3
View File
@@ -5,9 +5,7 @@ use pdf_inspector::vision::{PdfiumRenderer, RenderError, RenderOptions, RenderPi
fn load_renderer() -> Option<PdfiumRenderer> {
match PdfiumRenderer::load() {
Ok(renderer) => Some(renderer),
Err(RenderError::Pdfium(firecrawl_pdfium::Error::Load(
firecrawl_pdfium::LoadError::LibraryNotFound { .. },
))) => {
Err(RenderError::PdfiumLoad { .. }) => {
eprintln!("skipping PDFium runtime test because no native library is installed");
None
}
+58 -6
View File
@@ -2,7 +2,8 @@
#[cfg(feature = "ocr")]
use pdf_inspector::vision::{
process_pdf_with_ocr_mem, ModelDownloadPolicy, OcrPdfOptions, PageContentSource,
process_pdf_with_ocr_mem, ModelDownloadPolicy, OcrPdfOptions, OcrPipelineError,
PageContentSource,
};
use pdf_inspector::vision::{
ModelStore, OarOcrEngine, OcrEngine, OcrMode, OcrOptions, PageTransform, RenderPixelFormat,
@@ -19,9 +20,7 @@ const EXPECTED_TEXT_ENV: &str = "PDF_INSPECTOR_OCR_TEST_EXPECTED";
fn load_renderer() -> Option<PdfiumRenderer> {
match PdfiumRenderer::load() {
Ok(renderer) => Some(renderer),
Err(RenderError::Pdfium(firecrawl_pdfium::Error::Load(
firecrawl_pdfium::LoadError::LibraryNotFound { .. },
))) => {
Err(RenderError::PdfiumLoad { .. }) => {
eprintln!("skipping OCR runtime test because no native PDFium library is installed");
None
}
@@ -127,7 +126,16 @@ fn complete_ocr_pipeline_routes_and_assembles_a_scanned_fixture() {
assert_eq!(result.pages_routed_to_ocr, vec![1]);
assert!(!result.markdown.trim().is_empty());
assert_eq!(result.pages[0].provenance.source, PageContentSource::Ocr);
assert!(result
.markdown
.contains("Order Date Item Code Description Status Unit Cost\n\n03/14/2024"));
assert!(result.markdown.contains("$482,110.40\n\n05/02/2024"));
assert_eq!(result.pages[0].provenance.source, PageContentSource::Fused);
assert!(result.pages[0]
.provenance
.warnings
.iter()
.any(|warning| warning.contains("complementary OCR")));
assert_eq!(
result.pages[0].provenance.ocr_model.as_ref().unwrap().name,
PP_OCR_V6_SMALL.id
@@ -135,6 +143,50 @@ fn complete_ocr_pipeline_routes_and_assembles_a_scanned_fixture() {
assert_eq!(repeated.markdown, result.markdown);
}
#[cfg(all(feature = "ocr", feature = "render-pdfium"))]
#[test]
fn auto_recovers_credible_native_text_before_loading_ocr_models() {
let Some(_renderer) = load_renderer() else {
return;
};
let bytes = std::fs::read("tests/fixtures/shinagawa_identity_h.pdf").unwrap();
let ocr = OcrOptions::new()
.mode(OcrMode::Auto)
.model_directory("/models/must-not-be-read")
.model_downloads(ModelDownloadPolicy::Offline);
let result = process_pdf_with_ocr_mem(&bytes, OcrPdfOptions::new().ocr(ocr)).unwrap();
assert_eq!(result.pages_recommended_for_ocr, vec![1]);
assert!(result.pages_routed_to_ocr.is_empty());
assert!(result.markdown.contains("羽田空港新飛行経路"));
assert!(result.markdown.contains("|4月30日|有|81.0|"));
assert!(result.markdown.contains("※1 最大騒音レベル"));
assert!(result.pages_with_tables.contains(&1));
assert_eq!(result.pages[0].provenance.source, PageContentSource::Native);
assert!(result.pages[0].provenance.ocr_model.is_none());
}
#[cfg(all(feature = "ocr", feature = "render-pdfium"))]
#[test]
fn auto_rejects_garbled_native_recovery_and_continues_to_ocr() {
let Some(_renderer) = load_renderer() else {
return;
};
let bytes = std::fs::read("tests/fixtures/shifted_cipher_tounicode.pdf").unwrap();
let ocr = OcrOptions::new()
.mode(OcrMode::Auto)
.model_directory("/models/must-not-be-read")
.model_downloads(ModelDownloadPolicy::Offline);
let error = process_pdf_with_ocr_mem(&bytes, OcrPdfOptions::new().ocr(ocr)).unwrap_err();
assert!(matches!(
error,
OcrPipelineError::ModelAcquire(_) | OcrPipelineError::ModelStore(_)
));
}
fn recognize(
model_directory: &std::ffi::OsStr,
pages: &[RenderedPage],
@@ -154,7 +206,7 @@ fn recognize(
fn assert_usable_result(results: &[pdf_inspector::vision::OcrPage]) {
assert_eq!(results.len(), 1);
assert_eq!(results[0].page, 1);
assert_eq!(results[0].page_number, 1);
assert_eq!(results[0].model.name, PP_OCR_V6_SMALL.id);
assert_eq!(results[0].model.revision, PP_OCR_V6_SMALL.revision);
assert!(!results[0].spans.is_empty());