From c66819c1654e50d0d0f64c53db8b07865560956c Mon Sep 17 00:00:00 2001 From: Abimael Martell Date: Thu, 2 Apr 2026 12:28:17 -0700 Subject: [PATCH] Reorganize README: move API references to docs/ Move detailed Python, Rust, and debugging docs into docs/ to keep the main README focused on overview and quick start examples. Co-Authored-By: Claude Opus 4.6 --- README.md | 236 +++------------------------------------------- docs/debugging.md | 29 ++++++ docs/python.md | 74 +++++++++++++++ docs/rust-api.md | 122 ++++++++++++++++++++++++ 4 files changed, 240 insertions(+), 221 deletions(-) create mode 100644 docs/debugging.md create mode 100644 docs/python.md create mode 100644 docs/rust-api.md diff --git a/README.md b/README.md index d4a6fbd..b0eb716 100644 --- a/README.md +++ b/README.md @@ -12,85 +12,30 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in - **Table detection** — Dual-mode: rectangle-based detection from PDF drawing ops, plus heuristic detection from text alignment. Handles financial tables, footnotes, and continuation tables across pages. - **CID font support** — ToUnicode CMap decoding for Type0/Identity-H fonts, UTF-16BE, UTF-8, and Latin-1 encodings. - **Multi-column layout** — Automatic detection of newspaper-style columns, sequential reading order, and RTL text support. -- **Encoding issue detection** — Automatically flags broken font encodings (garbled text, replacement characters) so callers can fall back to OCR. +- **Encoding issue detection** — Automatically flags broken font encodings so callers can fall back to OCR. - **Single document load** — The document is parsed once and shared between detection and extraction, avoiding redundant I/O. - **Lightweight** — Pure Rust, no ML models, no external services. Single dependency on `lopdf` for PDF parsing. -- **Python bindings** — Use from Python via PyO3. Install with `pip install pdf-inspector` or build from source with `maturin`. ## Quick start ### Python -Install from source (requires Rust toolchain): - ```bash pip install maturin maturin develop --release ``` -Use it: - ```python import pdf_inspector -# Full processing: detect + extract + convert to Markdown result = pdf_inspector.process_pdf("document.pdf") -print(result.pdf_type) # "text_based", "scanned", "image_based", "mixed" -print(result.confidence) # 0.0 - 1.0 -print(result.page_count) # number of pages -print(result.markdown) # Markdown string or None - -# Process specific pages only -result = pdf_inspector.process_pdf("document.pdf", pages=[1, 3, 5]) - -# Process from bytes (no filesystem needed) -with open("document.pdf", "rb") as f: - result = pdf_inspector.process_pdf_bytes(f.read()) - -# Fast detection only (no text extraction) -result = pdf_inspector.detect_pdf("document.pdf") -if result.pdf_type == "text_based": - print("Can extract locally!") -else: - print(f"Pages needing OCR: {result.pages_needing_ocr}") - -# Plain text extraction -text = pdf_inspector.extract_text("document.pdf") - -# Positioned text items with font info -items = pdf_inspector.extract_text_with_positions("document.pdf") -for item in items[:5]: - print(f"'{item.text}' at ({item.x:.0f}, {item.y:.0f}) size={item.font_size}") +print(result.pdf_type) # "text_based", "scanned", "image_based", "mixed" +print(result.markdown) # Markdown string or None ``` -#### Python API reference +> Full API reference: [docs/python.md](docs/python.md) -| Function | Description | -|---|---| -| `process_pdf(path, pages=None)` | Full processing (detect + extract + markdown) | -| `process_pdf_bytes(data, pages=None)` | Full processing from bytes | -| `detect_pdf(path)` | Fast detection only (returns PdfResult) | -| `detect_pdf_bytes(data)` | Fast detection from bytes | -| `classify_pdf(path)` | Lightweight classification (returns PdfClassification) | -| `classify_pdf_bytes(data)` | Lightweight classification from bytes | -| `extract_text(path)` | Plain text extraction | -| `extract_text_bytes(data)` | Plain text extraction from bytes | -| `extract_text_with_positions(path, pages=None)` | Text with X/Y coords and font info | -| `extract_text_with_positions_bytes(data, pages=None)` | Text with positions from bytes | -| `extract_text_in_regions(path, page_regions)` | Extract text in bounding-box regions | -| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes | - -**`PdfResult` fields:** `pdf_type`, `markdown`, `page_count`, `processing_time_ms`, `pages_needing_ocr`, `title`, `confidence`, `is_complex_layout`, `pages_with_tables`, `pages_with_columns`, `has_encoding_issues` - -**`PdfClassification` fields:** `pdf_type`, `page_count`, `pages_needing_ocr` (0-indexed), `confidence` - -**`TextItem` fields:** `text`, `x`, `y`, `width`, `height`, `font`, `font_size`, `page`, `is_bold`, `is_italic`, `item_type` - -**`RegionText` fields:** `text`, `needs_ocr` - -**`PageRegionTexts` fields:** `page` (0-indexed), `regions` (list of RegionText) - -### Node.js (NAPI) +### Node.js ```bash npm install @firecrawl/pdf-inspector-js @@ -98,114 +43,33 @@ npm install @firecrawl/pdf-inspector-js ```javascript import { readFileSync } from 'fs'; -import { processPdf, classifyPdf, extractTextInRegions } from '@firecrawl/pdf-inspector-js'; +import { processPdf, classifyPdf } from '@firecrawl/pdf-inspector-js'; -const buffer = readFileSync('document.pdf'); - -// Full processing -const result = processPdf(buffer); +const result = processPdf(readFileSync('document.pdf')); console.log(result.pdfType); // "TextBased", "Scanned", "ImageBased", "Mixed" console.log(result.markdown); // Markdown string or null - -// Lightweight classification -const cls = classifyPdf(buffer); -console.log(cls.pdfType, cls.pagesNeedingOcr); - -// Region-based extraction (for hybrid OCR pipelines) -const regions = extractTextInRegions(buffer, [ - { page: 0, regions: [[0, 0, 600, 100]] } -]); ``` -#### Node.js API reference - -| Function | Description | -|---|---| -| `processPdf(buffer, pages?)` | Full processing (detect + extract + markdown) | -| `detectPdf(buffer)` | Fast detection only (returns PdfResult) | -| `classifyPdf(buffer)` | Lightweight classification (returns PdfClassification) | -| `extractText(buffer)` | Plain text extraction | -| `extractTextWithPositions(buffer, pages?)` | Text with X/Y coords and font info | -| `extractTextInRegions(buffer, pageRegions)` | Extract text in bounding-box regions | +> Full API reference: [napi/README.md](napi/README.md) ### Rust -Add to your `Cargo.toml`: - ```toml [dependencies] pdf-inspector = { git = "https://github.com/firecrawl/pdf-inspector" } ``` -Detect and extract in one call: - ```rust use pdf_inspector::process_pdf; let result = process_pdf("document.pdf")?; - -println!("Type: {:?}", result.pdf_type); // TextBased, Scanned, ImageBased, Mixed -println!("Confidence: {:.0}%", result.confidence * 100.0); -println!("Pages: {}", result.page_count); - +println!("Type: {:?}", result.pdf_type); if let Some(markdown) = &result.markdown { println!("{}", markdown); } ``` -Fast metadata-only detection (no text extraction or markdown generation): - -```rust -use pdf_inspector::detect_pdf; - -let info = detect_pdf("document.pdf")?; - -match info.pdf_type { - pdf_inspector::PdfType::TextBased => { - // Extract locally — fast and free - } - _ => { - // Route to OCR service - // info.pages_needing_ocr tells you exactly which pages - } -} -``` - -Customize processing with `PdfOptions`: - -```rust -use pdf_inspector::{process_pdf_with_options, PdfOptions, ProcessMode, DetectionConfig, ScanStrategy}; - -// Analyze layout without generating markdown -let result = process_pdf_with_options( - "document.pdf", - PdfOptions::new().mode(ProcessMode::Analyze), -)?; - -// Full extraction with custom detection strategy -let result = process_pdf_with_options( - "large.pdf", - PdfOptions::new().detection(DetectionConfig { - strategy: ScanStrategy::Sample(5), - ..Default::default() - }), -)?; - -// Process only specific pages -let result = process_pdf_with_options( - "document.pdf", - PdfOptions::new().pages([1, 3, 5]), -)?; -``` - -Process from a byte buffer (no filesystem needed): - -```rust -use pdf_inspector::process_pdf_mem; - -let bytes = std::fs::read("document.pdf")?; -let result = process_pdf_mem(&bytes)?; -``` +> Full API reference: [docs/rust-api.md](docs/rust-api.md) ### CLI @@ -230,7 +94,6 @@ cargo run --bin detect-pdf -- document.pdf cargo run --bin detect-pdf -- document.pdf --json # Detection + layout analysis (tables, columns) -cargo run --bin detect-pdf -- document.pdf --analyze cargo run --bin detect-pdf -- document.pdf --analyze --json ``` @@ -280,6 +143,7 @@ src/ tables/ — Table detection and formatting markdown/ — Markdown conversion and structure detection bin/ — CLI tools (pdf2md, detect_pdf) +napi/ — Node.js/Bun bindings (napi-rs) ``` ## How classification works @@ -300,50 +164,6 @@ This detects 300+ page PDFs in milliseconds. The result includes `pages_needing_ | `Sample(n)` | Sample `n` evenly distributed pages (first, last, middle) | Very large PDFs where speed matters more than precision | | `Pages(vec)` | Only scan specific 1-indexed page numbers | When the caller knows which pages to check | -## Rust API - -### Processing modes - -| Mode | What it does | Returns | -|---|---|---| -| `ProcessMode::Full` (default) | Detect + extract + convert to Markdown | Everything populated | -| `ProcessMode::Analyze` | Detect + extract + layout analysis (no Markdown) | `markdown` is `None`, `layout` is populated | -| `ProcessMode::DetectOnly` | Classification only (fastest) | `markdown` is `None`, `layout` is default | - -### Functions - -| Function | Description | -|---|---| -| `process_pdf(path)` | Full processing with defaults | -| `detect_pdf(path)` | Fast metadata-only detection (no extraction) | -| `process_pdf_with_options(path, options)` | Process with custom `PdfOptions` | -| `process_pdf_mem(bytes)` | Full processing from a byte buffer | -| `detect_pdf_mem(bytes)` | Fast detection from a byte buffer | -| `process_pdf_mem_with_options(bytes, options)` | Process from bytes with custom options | -| `extract_text(path)` | Plain text extraction | -| `extract_text_with_positions(path)` | Text with X/Y coordinates and font info | -| `to_markdown(text, options)` | Convert plain text to Markdown | -| `to_markdown_from_items(items, options)` | Markdown from pre-extracted `TextItem`s | -| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection | - -Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`. - -### Types - -| Type | Description | -|---|---| -| `PdfOptions` | Builder for processing configuration (mode, detection, markdown, page filter) | -| `ProcessMode` | `DetectOnly`, `Analyze`, `Full` | -| `PdfType` | `TextBased`, `Scanned`, `ImageBased`, `Mixed` | -| `PdfProcessResult` | Full result: pdf_type, markdown, page_count, confidence, layout, has_encoding_issues, timing | -| `PdfTypeResult` | Low-level detection result: type, confidence, page count, pages needing OCR | -| `DetectionConfig` | Configuration for detection: scan strategy, thresholds | -| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` | -| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns | -| `TextItem` | Text with position, font info, and page number | -| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) | -| `PdfError` | `Io`, `Parse`, `Encrypted`, `InvalidStructure`, `NotAPdf` | - ## Markdown output The converter handles: @@ -366,36 +186,6 @@ The converter handles: | Drop caps | Large initial letters merged with following text | | Dot leaders | TOC-style dots collapsed to " ... " | -## Debugging with RUST_LOG - -Structured logging via `RUST_LOG` replaces the former debug binaries. Set the environment variable to control which sections emit debug output on stderr: - -```bash -# Raw PDF content stream operators (replaces dump_ops) -RUST_LOG=pdf_inspector::extractor::content_stream=trace cargo run --bin pdf2md -- file.pdf > /dev/null - -# Font metadata, encodings, ligatures (replaces debug_fonts / debug_ligatures) -RUST_LOG=pdf_inspector::extractor::fonts=debug cargo run --bin pdf2md -- file.pdf > /dev/null - -# ToUnicode CMap parsing -RUST_LOG=pdf_inspector::tounicode=debug cargo run --bin pdf2md -- file.pdf > /dev/null - -# Text items per page with x/y/width (replaces debug_spaces / debug_pages) -RUST_LOG=pdf_inspector::extractor=debug cargo run --bin pdf2md -- file.pdf > /dev/null - -# Column detection and reading order (replaces debug_order) -RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- file.pdf > /dev/null - -# Y-gap analysis and paragraph thresholds (replaces debug_ygaps) -RUST_LOG=pdf_inspector::markdown::analysis=debug cargo run --bin pdf2md -- file.pdf > /dev/null - -# Table detection -RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- file.pdf > /dev/null - -# Everything -RUST_LOG=pdf_inspector=debug cargo run --bin pdf2md -- file.pdf > /dev/null -``` - ## Use case: smart PDF routing pdf-inspector was built for pipelines that process PDFs at scale. Instead of sending every PDF through OCR: @@ -410,6 +200,10 @@ PDF arrives This saves cost and latency for the majority of PDFs that are already text-based (reports, papers, invoices, legal docs). +## Debugging + +See [docs/debugging.md](docs/debugging.md) for `RUST_LOG` environment variable usage. + ## License MIT diff --git a/docs/debugging.md b/docs/debugging.md new file mode 100644 index 0000000..963eb24 --- /dev/null +++ b/docs/debugging.md @@ -0,0 +1,29 @@ +# Debugging with RUST_LOG + +Structured logging via `RUST_LOG` replaces the former debug binaries. Set the environment variable to control which sections emit debug output on stderr: + +```bash +# Raw PDF content stream operators (replaces dump_ops) +RUST_LOG=pdf_inspector::extractor::content_stream=trace cargo run --bin pdf2md -- file.pdf > /dev/null + +# Font metadata, encodings, ligatures (replaces debug_fonts / debug_ligatures) +RUST_LOG=pdf_inspector::extractor::fonts=debug cargo run --bin pdf2md -- file.pdf > /dev/null + +# ToUnicode CMap parsing +RUST_LOG=pdf_inspector::tounicode=debug cargo run --bin pdf2md -- file.pdf > /dev/null + +# Text items per page with x/y/width (replaces debug_spaces / debug_pages) +RUST_LOG=pdf_inspector::extractor=debug cargo run --bin pdf2md -- file.pdf > /dev/null + +# Column detection and reading order (replaces debug_order) +RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- file.pdf > /dev/null + +# Y-gap analysis and paragraph thresholds (replaces debug_ygaps) +RUST_LOG=pdf_inspector::markdown::analysis=debug cargo run --bin pdf2md -- file.pdf > /dev/null + +# Table detection +RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- file.pdf > /dev/null + +# Everything +RUST_LOG=pdf_inspector=debug cargo run --bin pdf2md -- file.pdf > /dev/null +``` diff --git a/docs/python.md b/docs/python.md new file mode 100644 index 0000000..8416460 --- /dev/null +++ b/docs/python.md @@ -0,0 +1,74 @@ +# Python API + +Python bindings via [PyO3](https://pyo3.rs). Requires Rust toolchain for building from source. + +## Install + +```bash +pip install maturin +maturin develop --release +``` + +## Usage + +```python +import pdf_inspector + +# Full processing: detect + extract + convert to Markdown +result = pdf_inspector.process_pdf("document.pdf") +print(result.pdf_type) # "text_based", "scanned", "image_based", "mixed" +print(result.confidence) # 0.0 - 1.0 +print(result.page_count) # number of pages +print(result.markdown) # Markdown string or None + +# Process specific pages only +result = pdf_inspector.process_pdf("document.pdf", pages=[1, 3, 5]) + +# Process from bytes (no filesystem needed) +with open("document.pdf", "rb") as f: + result = pdf_inspector.process_pdf_bytes(f.read()) + +# Fast detection only (no text extraction) +result = pdf_inspector.detect_pdf("document.pdf") +if result.pdf_type == "text_based": + print("Can extract locally!") +else: + print(f"Pages needing OCR: {result.pages_needing_ocr}") + +# Plain text extraction +text = pdf_inspector.extract_text("document.pdf") + +# Positioned text items with font info +items = pdf_inspector.extract_text_with_positions("document.pdf") +for item in items[:5]: + print(f"'{item.text}' at ({item.x:.0f}, {item.y:.0f}) size={item.font_size}") +``` + +## API reference + +| Function | Description | +|---|---| +| `process_pdf(path, pages=None)` | Full processing (detect + extract + markdown) | +| `process_pdf_bytes(data, pages=None)` | Full processing from bytes | +| `detect_pdf(path)` | Fast detection only (returns PdfResult) | +| `detect_pdf_bytes(data)` | Fast detection from bytes | +| `classify_pdf(path)` | Lightweight classification (returns PdfClassification) | +| `classify_pdf_bytes(data)` | Lightweight classification from bytes | +| `extract_text(path)` | Plain text extraction | +| `extract_text_bytes(data)` | Plain text extraction from bytes | +| `extract_text_with_positions(path, pages=None)` | Text with X/Y coords and font info | +| `extract_text_with_positions_bytes(data, pages=None)` | Text with positions from bytes | +| `extract_text_in_regions(path, page_regions)` | Extract text in bounding-box regions | +| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes | + +## Types + +**`PdfResult` fields:** `pdf_type`, `markdown`, `page_count`, `processing_time_ms`, `pages_needing_ocr`, `title`, `confidence`, `is_complex_layout`, `pages_with_tables`, `pages_with_columns`, `has_encoding_issues` + +**`PdfClassification` fields:** `pdf_type`, `page_count`, `pages_needing_ocr` (0-indexed), `confidence` + +**`TextItem` fields:** `text`, `x`, `y`, `width`, `height`, `font`, `font_size`, `page`, `is_bold`, `is_italic`, `item_type` + +**`RegionText` fields:** `text`, `needs_ocr` + +**`PageRegionTexts` fields:** `page` (0-indexed), `regions` (list of RegionText) diff --git a/docs/rust-api.md b/docs/rust-api.md new file mode 100644 index 0000000..02eec00 --- /dev/null +++ b/docs/rust-api.md @@ -0,0 +1,122 @@ +# Rust API + +Add to your `Cargo.toml`: + +```toml +[dependencies] +pdf-inspector = { git = "https://github.com/firecrawl/pdf-inspector" } +``` + +## Usage + +Detect and extract in one call: + +```rust +use pdf_inspector::process_pdf; + +let result = process_pdf("document.pdf")?; + +println!("Type: {:?}", result.pdf_type); // TextBased, Scanned, ImageBased, Mixed +println!("Confidence: {:.0}%", result.confidence * 100.0); +println!("Pages: {}", result.page_count); + +if let Some(markdown) = &result.markdown { + println!("{}", markdown); +} +``` + +Fast metadata-only detection (no text extraction or markdown generation): + +```rust +use pdf_inspector::detect_pdf; + +let info = detect_pdf("document.pdf")?; + +match info.pdf_type { + pdf_inspector::PdfType::TextBased => { + // Extract locally — fast and free + } + _ => { + // Route to OCR service + // info.pages_needing_ocr tells you exactly which pages + } +} +``` + +Customize processing with `PdfOptions`: + +```rust +use pdf_inspector::{process_pdf_with_options, PdfOptions, ProcessMode, DetectionConfig, ScanStrategy}; + +// Analyze layout without generating markdown +let result = process_pdf_with_options( + "document.pdf", + PdfOptions::new().mode(ProcessMode::Analyze), +)?; + +// Full extraction with custom detection strategy +let result = process_pdf_with_options( + "large.pdf", + PdfOptions::new().detection(DetectionConfig { + strategy: ScanStrategy::Sample(5), + ..Default::default() + }), +)?; + +// Process only specific pages +let result = process_pdf_with_options( + "document.pdf", + PdfOptions::new().pages([1, 3, 5]), +)?; +``` + +Process from a byte buffer (no filesystem needed): + +```rust +use pdf_inspector::process_pdf_mem; + +let bytes = std::fs::read("document.pdf")?; +let result = process_pdf_mem(&bytes)?; +``` + +## Processing modes + +| Mode | What it does | Returns | +|---|---|---| +| `ProcessMode::Full` (default) | Detect + extract + convert to Markdown | Everything populated | +| `ProcessMode::Analyze` | Detect + extract + layout analysis (no Markdown) | `markdown` is `None`, `layout` is populated | +| `ProcessMode::DetectOnly` | Classification only (fastest) | `markdown` is `None`, `layout` is default | + +## Functions + +| Function | Description | +|---|---| +| `process_pdf(path)` | Full processing with defaults | +| `detect_pdf(path)` | Fast metadata-only detection (no extraction) | +| `process_pdf_with_options(path, options)` | Process with custom `PdfOptions` | +| `process_pdf_mem(bytes)` | Full processing from a byte buffer | +| `detect_pdf_mem(bytes)` | Fast detection from a byte buffer | +| `process_pdf_mem_with_options(bytes, options)` | Process from bytes with custom options | +| `extract_text(path)` | Plain text extraction | +| `extract_text_with_positions(path)` | Text with X/Y coordinates and font info | +| `to_markdown(text, options)` | Convert plain text to Markdown | +| `to_markdown_from_items(items, options)` | Markdown from pre-extracted `TextItem`s | +| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection | + +Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`. + +## Types + +| Type | Description | +|---|---| +| `PdfOptions` | Builder for processing configuration (mode, detection, markdown, page filter) | +| `ProcessMode` | `DetectOnly`, `Analyze`, `Full` | +| `PdfType` | `TextBased`, `Scanned`, `ImageBased`, `Mixed` | +| `PdfProcessResult` | Full result: pdf_type, markdown, page_count, confidence, layout, has_encoding_issues, timing | +| `PdfTypeResult` | Low-level detection result: type, confidence, page count, pages needing OCR | +| `DetectionConfig` | Configuration for detection: scan strategy, thresholds | +| `ScanStrategy` | `EarlyExit`, `Full`, `Sample(n)`, `Pages(vec)` | +| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns | +| `TextItem` | Text with position, font info, and page number | +| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) | +| `PdfError` | `Io`, `Parse`, `Encrypted`, `InvalidStructure`, `NotAPdf` |