* feat(vision): add OAR OCR engine * fix(vision): harden OAR runtime loading * refactor(vision): use OCR engine terminology
14 KiB
pdf-inspector
Fast PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. The default build is pure Rust, has no ML models or external services, and uses lopdf for PDF parsing. Also available for Python and Node.js.
Built by Firecrawl to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.
Features
- Smart classification — TextBased / Scanned / ImageBased / Mixed in ~10–50ms, with a confidence score and per-page OCR routing.
- Markdown conversion — headings, lists, code blocks, bold/italic, URL linking, and dual-mode table detection (PDF drawing ops + text-alignment heuristics).
- Layout-aware extraction — multi-column reading order, position and font info per text item, RTL support.
- Robust text decoding — CID/Type0 fonts via ToUnicode CMaps, plus automatic flagging of broken encodings so callers can fall back to OCR.
- Lightweight — pure Rust, no ML models, no external services; single PDF dependency (lopdf).
Benchmark
opendataloader-bench corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 0–1, higher is better:
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|---|---|---|---|---|---|
| pdf-inspector | 0.875 | 0.915 | 0.814 | 0.788 | 0.470s |
| liteparse | 0.873 | 0.913 | 0.693 | 0.811 | 0.750s |
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the repo README, with raw timings and artifacts in the results branch.
Install
cargo add pdf-inspector
For the latest unreleased changes, use the git dependency instead:
[dependencies]
pdf-inspector = { git = "https://github.com/firecrawl/pdf-inspector" }
The crate also ships CLI binaries — pdf2md (PDF → Markdown, with --json, --pages, --select-pages, and the opt-in token-saving --compact profile) and detect-pdf (classification, with --analyze --json):
cargo install pdf-inspector
Usage
Detect and extract in one call:
use pdf_inspector::process_pdf;
let result = process_pdf("document.pdf")?;
println!("Type: {:?}", result.pdf_type); // TextBased, Scanned, ImageBased, Mixed
println!("Confidence: {:.0}%", result.confidence * 100.0);
println!("Pages: {}", result.page_count);
if let Some(markdown) = &result.markdown {
println!("{}", markdown);
}
Fast metadata-only detection (no text extraction or markdown generation):
use pdf_inspector::detect_pdf;
let info = detect_pdf("document.pdf")?;
match info.pdf_type {
pdf_inspector::PdfType::TextBased => {
// Extract locally — fast and free
}
_ => {
// Route to OCR service
// info.pages_needing_ocr tells you exactly which pages
}
}
Customize processing with PdfOptions:
use pdf_inspector::{process_pdf_with_options, PdfOptions, ProcessMode, DetectionConfig, ScanStrategy};
// Analyze layout without generating markdown
let result = process_pdf_with_options(
"document.pdf",
PdfOptions::new().mode(ProcessMode::Analyze),
)?;
// Full extraction with custom detection strategy
let result = process_pdf_with_options(
"large.pdf",
PdfOptions::new().detection(DetectionConfig {
strategy: ScanStrategy::Sample(5),
..Default::default()
}),
)?;
// Process only specific pages
let result = process_pdf_with_options(
"document.pdf",
PdfOptions::new().pages([1, 3, 5]),
)?;
Process from a byte buffer (no filesystem needed):
use pdf_inspector::process_pdf_mem;
let bytes = std::fs::read("document.pdf")?;
let result = process_pdf_mem(&bytes)?;
Vision extension contracts
The native-only vision feature exposes the stable seam used by OCR
integrations without selecting or embedding an inference runtime. The
separate model-cache feature adds pinned artifact management:
PageRenderer,OcrEngine, andLayoutEnginetraits;- renderer-neutral owned page buffers and affine pixel↔PDF transforms;
OcrOptionsand opt-inOff/Auto/Forcerouting modes;- positioned OCR/layout results and per-page provenance types; and
- a versioned PP-OCRv6 Small manifest with checksum-verified, locked, atomic model-cache installation and explicit offline-directory overrides.
[dependencies]
pdf-inspector = { version = "1", features = ["vision", "model-cache"] }
The OCR contracts preserve existing behavior by default: OCR is Off, learned
layout is disabled, and model resolution is never reached. ModelStore itself
does not access the network; a runtime integration can fetch a manifest's
canonical URL only when allowed and pass the stream to ModelStore::install.
Offline consumers set an explicit model directory and ModelDownloadPolicy::Offline.
Renderer-only consumers do not enable model-cache and therefore do not compile
its filesystem, locking, or hashing dependencies.
use pdf_inspector::vision::{
ModelDownloadPolicy, ModelStore, OcrMode, OcrOptions, PP_OCR_V6_SMALL,
};
let ocr = OcrOptions::new()
.mode(OcrMode::Auto)
.model_directory("/opt/firecrawl/models/pp-ocrv6-small")
.model_downloads(ModelDownloadPolicy::Offline);
// Verifies exact sizes and SHA-256 digests before an engine opens the files.
let models = ModelStore::from_options(&ocr)?.resolve(&PP_OCR_V6_SMALL)?;
println!("using {} at {}", models.manifest_id(), models.revision());
Optional native page rendering
The render-pdfium feature adds a native-only page renderer backed by
firecrawl-pdfium. It is the
rendering boundary for OCR pipelines; enabling it does not include an OCR
model or change the existing extraction functions. It implies vision,
and PdfiumRenderer implements the renderer-neutral PageRenderer trait.
[dependencies]
pdf-inspector = { version = "1", features = ["render-pdfium"] }
PDFium is loaded at runtime. Set PDFIUM_LIB_PATH, place its shared library
next to the executable, or use another discovery route supported by
firecrawl-pdfium.
use pdf_inspector::vision::{PdfiumRenderer, RenderOptions};
let renderer = PdfiumRenderer::load()?;
let bytes = std::fs::read("document.pdf")?;
let pages = renderer.render_pages(
&bytes,
&[1, 3], // 1-indexed, matching pages_needing_ocr
None, // optional PDF password
&RenderOptions::new().dpi(150.0),
)?;
for page in pages {
// Owned RGB pixels can leave the PDFium critical section and be sent to
// an OCR worker. OCR pixel boxes can be mapped back to PDF coordinates.
let rect = page.pixel_rect_to_pdf_rect(20.0, 30.0, 100.0, 24.0);
println!("page {}: {}x{}, rect={rect:?}", page.page(), page.width(), page.height());
}
Browser WASM remains on the default text-only path and does not expose native PDFium rendering.
Optional OCR engine
The native-only ocr-oar feature adds a CPU PP-OCRv6 Small implementation of
OcrEngine backed by OAR and ONNX Runtime. It implies model-cache, but does
not enable model auto-download, ONNX Runtime download, or PDF rendering. Model
files remain external, must match the pinned manifest, and are opened only
after ModelStore verifies their exact size and SHA-256 digest. Install an
ONNX Runtime shared library separately and set ORT_DYLIB_PATH when it is not
available through the platform library search path. The feature currently
requires Rust 1.95 or newer, matching OAR 0.9.1's MSRV.
[dependencies]
pdf-inspector = { version = "1", features = ["ocr-oar", "render-pdfium"] }
Direct engine invocation is intentionally separate from extraction routing and native/OCR fusion:
use pdf_inspector::vision::{
ModelDownloadPolicy, ModelStore, OarOcrEngine, OcrEngine, OcrMode,
OcrOptions, PdfiumRenderer, RenderOptions, PP_OCR_V6_SMALL,
};
let options = OcrOptions::new()
.mode(OcrMode::Force)
.minimum_confidence(0.45)
.model_directory("/opt/firecrawl/models/pp-ocrv6-small")
.model_downloads(ModelDownloadPolicy::Offline);
let models = ModelStore::from_options(&options)?.resolve(&PP_OCR_V6_SMALL)?;
let engine = OarOcrEngine::from_models(&models)?;
let renderer = PdfiumRenderer::load()?;
let bytes = std::fs::read("scan.pdf")?;
let pages = renderer.render_pages(&bytes, &[1], None, &RenderOptions::new())?;
let ocr_pages = engine.recognize(&pages, &options)?;
for span in &ocr_pages[0].spans {
println!("{:.3}: {}", span.confidence, span.text);
}
The engine accepts renderer-neutral RGB, RGBA, and grayscale pages, preserves
OAR's positioned quadrilaterals in bitmap coordinates, filters spans using
minimum_confidence, and records the pinned model revision in every OcrPage.
OcrMode::Off is rejected at the engine boundary so default options cannot run
inference accidentally. Selective routing and OCR/native-text fusion are added
by higher stack layers.
Extract per-page Markdown (one string per page, plus document-wide layout metadata):
use pdf_inspector::extract_pages_markdown;
// Pass `None` for every page in document order, or a slice of 0-indexed
// pages to restrict the output (caller-supplied order is preserved).
let result = extract_pages_markdown("document.pdf", None)?;
for page in &result.pages {
if page.needs_ocr {
// Route this page to OCR
} else {
println!("Page {}: {}", page.page, page.markdown);
}
}
println!("Complex layout? {}", result.is_complex);
Extract structure-tree elements from tagged PDFs, and join them against
extract_text_with_positions to attach semantic roles (heading levels,
paragraphs, table cells) to extracted text:
use pdf_inspector::{extract_structure_elements, extract_text_with_positions};
use std::collections::HashMap;
// One entry per marked-content reference, sorted by (page, mcid); empty for
// untagged PDFs. Pages are 1-indexed to match `TextItem::page`, so the
// (page, mcid) pair is a direct join key.
let elements = extract_structure_elements("tagged.pdf", None)?;
let roles: HashMap<(u32, i64), &str> = elements
.iter()
.map(|e| ((e.page, e.mcid), e.role.as_str()))
.collect();
for item in extract_text_with_positions("tagged.pdf")? {
if let Some(mcid) = item.mcid {
if let Some(role) = roles.get(&(item.page, mcid)) {
if role.starts_with('H') {
println!("{}: {}", role, item.text);
}
}
}
}
Processing modes
| Mode | What it does | Returns |
|---|---|---|
ProcessMode::Full (default) |
Detect + extract + convert to Markdown | Everything populated |
ProcessMode::Analyze |
Detect + extract + layout analysis (no Markdown) | markdown is None, layout is populated |
ProcessMode::DetectOnly |
Classification only (fastest) | markdown is None, layout is default |
Functions
| Function | Description |
|---|---|
process_pdf(path) |
Full processing with defaults |
detect_pdf(path) |
Fast metadata-only detection (no extraction) |
process_pdf_with_options(path, options) |
Process with custom PdfOptions |
process_pdf_mem(bytes) |
Full processing from a byte buffer |
detect_pdf_mem(bytes) |
Fast detection from a byte buffer |
process_pdf_mem_with_options(bytes, options) |
Process from bytes with custom options |
extract_text(path) |
Plain text extraction |
extract_text_with_positions(path) |
Text with X/Y coordinates and font info |
to_markdown(text, options) |
Convert plain text to Markdown |
to_markdown_from_items(items, options) |
Markdown from pre-extracted TextItems |
to_markdown_from_items_with_rects(items, options, rects) |
Markdown with rectangle-based table detection |
extract_pages_markdown(path, pages) |
Per-page Markdown + layout metadata (file) |
extract_pages_markdown_mem(bytes, pages) |
Per-page Markdown from bytes |
extract_structure_elements(path, pages) |
Structure-tree elements from tagged PDFs (page, mcid, role) |
extract_structure_elements_mem(bytes, pages) |
Structure-tree elements from bytes |
Low-level detection functions are also available via the detector module (detect_pdf_type, detect_pdf_type_with_config, etc.) for callers who need PdfTypeResult instead of PdfProcessResult.
Types
| Type | Description |
|---|---|
PdfOptions |
Builder for processing configuration (mode, detection, markdown, page filter) |
ProcessMode |
DetectOnly, Analyze, Full |
PdfType |
TextBased, Scanned, ImageBased, Mixed |
PdfProcessResult |
Full result: pdf_type, markdown, page_count, confidence, layout, has_encoding_issues, timing |
PdfTypeResult |
Low-level detection result: type, confidence, page count, pages needing OCR |
DetectionConfig |
Configuration for detection: scan strategy, thresholds |
ScanStrategy |
EarlyExit, Full, Sample(n), Pages(vec) |
LayoutComplexity |
Layout analysis: is_complex, pages_with_tables, pages_with_columns |
TextItem |
Text with position, font info, page number, and optional structure-tree mcid |
StructureElement |
Tagged-PDF structure reference: page (1-indexed), mcid, role ("H1".."H6", "P", …) |
MarkdownOptions |
Configuration for Markdown formatting (page numbers, etc.) |
PageMarkdown |
Per-page result: page (0-indexed), markdown, needs_ocr |
PagesExtractionResult |
Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
PdfError |
Io, Parse, Encrypted, InvalidStructure, NotAPdf |