From a4b1c714e8252ffa0bab2f58393cefe9b99a14b4 Mon Sep 17 00:00:00 2001 From: Abimael Martell <1450169+abimaelmartell@users.noreply.github.com> Date: Mon, 17 Aug 2026 10:30:29 -0700 Subject: [PATCH] chore(ocr): polish launch readiness (#409) Polish the selective OCR runtime documentation, packaging guidance, and cross-language launch smoke coverage. --- .github/workflows/ci.yml | 19 ++++++++ README.md | 8 ++-- docs/ocr-runtime.md | 97 ++++++++++++++++++++++++++++++++++++++++ docs/python.md | 4 +- docs/rust-api.md | 4 ++ napi/README.md | 4 +- 6 files changed, 130 insertions(+), 6 deletions(-) create mode 100644 docs/ocr-runtime.md diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index b258161..3ead3f5 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -180,18 +180,31 @@ jobs: test -n "$ort_path" echo "ORT_DYLIB_PATH=$ort_path" >> "$GITHUB_ENV" + - name: Configure isolated model cache + shell: bash + run: echo "PDF_INSPECTOR_MODEL_CACHE=$RUNNER_TEMP/pdf-inspector-models" >> "$GITHUB_ENV" + - name: Build OCR CLI run: cargo build --features ocr --bin pdf2md - name: Test PDFium runtime run: cargo test --features ocr --test local_render_tests + - name: Provision OCR model cache + shell: bash + run: | + target/debug/pdf2md \ + tests/fixtures/scan_with_native_header_text.pdf \ + --ocr force \ + --json > /dev/null + - name: Run OCR CLI shell: bash run: | target/debug/pdf2md \ tests/fixtures/scan_with_native_header_text.pdf \ --ocr auto \ + --ocr-offline \ --json > "$RUNNER_TEMP/ocr-result.json" - name: Validate OCR JSON contract @@ -213,6 +226,12 @@ jobs: assert "layout_ms" not in result["pages"][0]["timings"] PY + - name: Run OCR launch smoke set + shell: bash + run: | + export PDF_INSPECTOR_OCR_TEST_MODELS="$PDF_INSPECTOR_MODEL_CACHE/pp-ocrv6-small/oar-ocr-v0.7.0" + cargo test --features ocr --test ocr_tests -- --nocapture + - name: Build Node binding working-directory: napi run: | diff --git a/README.md b/README.md index 30a8574..8be9758 100644 --- a/README.md +++ b/README.md @@ -48,8 +48,7 @@ Use the [paired benchmark harness](docs/benchmarking.md) to compare two local bu ### Python ```bash -pip install maturin -maturin develop --release +pip install pdf-inspector ``` ```python @@ -182,8 +181,9 @@ and confidence, warnings, and pages recommended for the hosted document pipeline. Native Python and Node packages expose the same pipeline without a source-build feature. All native entry points still require separately installed PDFium and ONNX Runtime libraries only when OCR is routed. See the -[Rust API guide](docs/rust-api.md#complete-ocr-api) for model cache and offline -configuration. +[OCR runtime setup guide](docs/ocr-runtime.md) for pinned downloads, platform +support, model-cache behavior, and hosted-fallback integration. See the +[Rust API guide](docs/rust-api.md#complete-ocr-api) for lower-level controls. From a source checkout, use `cargo run --bin pdf2md -- document.pdf` or `cargo run --bin detect-pdf -- document.pdf` instead. diff --git a/docs/ocr-runtime.md b/docs/ocr-runtime.md new file mode 100644 index 0000000..28763d4 --- /dev/null +++ b/docs/ocr-runtime.md @@ -0,0 +1,97 @@ +# OCR runtime setup + +Selective OCR is available from the Rust library and CLI, Python, and Node.js. +Clean native-text documents do not load an OCR dependency or download a model. +When `auto` routes at least one page, the process needs PDFium, ONNX Runtime, +and the pinned PP-OCRv6 Small model set. + +## Validated versions + +The reproducible runtime path uses these builds: + +- [Firecrawl PDFium `native-v7988`](https://github.com/firecrawl/pdfium-rs/releases/tag/native-v7988), + containing PDFium `153.0.7988.0` +- [ONNX Runtime `1.27.0`](https://github.com/microsoft/onnxruntime/releases/tag/v1.27.0) +- PP-OCRv6 Small artifact revision `oar-ocr-v0.7.0` + +Use these versions for the reproducible path. Other compatible shared-library +builds may work, but are not part of the release smoke test. + +## Install the shared libraries + +Download and extract the matching archives: + +| Platform | PDFium asset | ONNX Runtime asset | +|---|---|---| +| Linux x64 | `firecrawl-pdfium-linux-x64.tgz` | `onnxruntime-linux-x64-1.27.0.tgz` | +| Linux ARM64 | `firecrawl-pdfium-linux-arm64.tgz` | `onnxruntime-linux-aarch64-1.27.0.tgz` | +| macOS Apple Silicon | `firecrawl-pdfium-mac-arm64.tgz` | `onnxruntime-osx-arm64-1.27.0.tgz` | +| Windows x64 | `firecrawl-pdfium-win-x64.tgz` | `onnxruntime-win-x64-1.27.0.zip` | + +The PDFium release publishes `SHA256SUMS`, build provenance, license files, +and an SPDX document for every platform archive. GitHub publishes a SHA-256 +digest with each ONNX Runtime asset. + +Point pdf-inspector at the extracted shared libraries when they are not on the +platform library search path: + +```bash +export PDFIUM_LIB_PATH=/absolute/path/to/libpdfium.so +export ORT_DYLIB_PATH=/absolute/path/to/libonnxruntime.so +pdf2md scan.pdf --ocr auto --json +``` + +On macOS the filenames end in `.dylib`. On Windows, use PowerShell and point +the variables at `pdfium.dll` and `onnxruntime.dll`: + +```powershell +$env:PDFIUM_LIB_PATH = "C:\absolute\path\to\pdfium.dll" +$env:ORT_DYLIB_PATH = "C:\absolute\path\to\onnxruntime.dll" +pdf2md scan.pdf --ocr auto --json +``` + +The native extraction packages also support platforms without these exact +runtime assets. In particular, the Python package has an Intel macOS wheel, +but ONNX Runtime 1.27.0 does not publish an Intel macOS archive; local OCR on +that target requires a compatible custom ONNX Runtime build. + +The full OCR path is exercised end to end on Linux x64 in CI. macOS and +Windows compile and run the feature's platform-independent tests, while their +external-runtime paths should be treated as preview until equivalent smoke +jobs are added. + +## Model cache and offline mode + +The first routed page downloads and SHA-256-verifies three pinned artifacts: +the detection model, recognition model, and character dictionary. Together +they are about 31 MB. They are stored below the platform cache directory. +Set `PDF_INSPECTOR_MODEL_CACHE` to choose a managed cache root. + +For hermetic deployments, populate the model directory ahead of time and use +the language-specific offline option: + +- CLI: `--ocr-offline --ocr-model-dir /models/pp-ocrv6-small` +- Rust: `ModelDownloadPolicy::Offline` with `OcrOptions::model_directory` +- Python: `offline=True, model_directory="/models/pp-ocrv6-small"` +- Node.js: `offline: true, modelDirectory: "/models/pp-ocrv6-small"` + +The model artifacts come from +[`GreatV/oar-ocr`](https://github.com/GreatV/oar-ocr/releases/tag/v0.7.0), +whose OCR implementation and upstream PaddleOCR project use Apache-2.0 +licensing. Models are downloaded at runtime and are not embedded in any +pdf-inspector package. + +## Hosted fallback boundary + +`pages_recommending_hosted` is available after the local pipeline completes. +It marks pages whose completed OCR result is empty, low-confidence, or still +appears incomplete. + +Setup and execution failures happen before that result exists. A missing or +incompatible PDFium/ONNX Runtime library, failed model acquisition, or OCR +execution error is returned as an error. A downstream integration that has a +hosted parser should catch that error and route the document to the hosted +path. This keeps deployment problems distinct from page-quality judgments. + +In `auto`, documents with no routed pages return successfully without touching +PDFium, ONNX Runtime, the model cache, or the network. diff --git a/docs/python.md b/docs/python.md index 423c874..242c23f 100644 --- a/docs/python.md +++ b/docs/python.md @@ -44,7 +44,9 @@ OCR calls that route work require compatible PDFium and ONNX Runtime shared libraries. Set `PDFIUM_LIB_PATH` and `ORT_DYLIB_PATH` when they are not on the platform library search path. The pinned OCR model set is downloaded and checksum-verified on the first routed page; use `offline=True` with a warm -cache or `model_directory` to prohibit network access. +cache or `model_directory` to prohibit network access. See the +[OCR runtime setup guide](https://github.com/firecrawl/pdf-inspector/blob/main/docs/ocr-runtime.md) +for pinned downloads, supported platforms, and hosted-fallback behavior. ## Usage diff --git a/docs/rust-api.md b/docs/rust-api.md index 0150610..13ed38a 100644 --- a/docs/rust-api.md +++ b/docs/rust-api.md @@ -383,6 +383,10 @@ shape; `Force` renders every selected page. OCR uses the existing deterministic table, column, reading-order, and Markdown assembly path; no learned layout model is included. +The [OCR runtime setup guide](https://github.com/firecrawl/pdf-inspector/blob/main/docs/ocr-runtime.md) +lists the pinned PDFium and ONNX Runtime builds, environment variables, model +cache behavior, and the error boundary downstream hosted fallbacks should use. + For ambiguous mixed pages, `Auto` privately retains clean native fragments instead of discarding them when OCR is selected. After recognition it compares script-agnostic text quality, OCR confidence, character overlap, and material diff --git a/napi/README.md b/napi/README.md index a89dac6..bbaacb3 100644 --- a/napi/README.md +++ b/napi/README.md @@ -41,7 +41,9 @@ OCR calls that route work require compatible PDFium and ONNX Runtime shared libraries. Set `PDFIUM_LIB_PATH` and `ORT_DYLIB_PATH` when they are not on the platform library search path. The pinned OCR model set is downloaded and checksum-verified on the first routed page; use `offline: true` with a warm -cache or `modelDirectory` to prohibit network access. +cache or `modelDirectory` to prohibit network access. See the +[OCR runtime setup guide](https://github.com/firecrawl/pdf-inspector/blob/main/docs/ocr-runtime.md) +for pinned downloads, supported platforms, and hosted-fallback behavior. ## API