From 6466e5927107a51d9f12243c433400092f38735f Mon Sep 17 00:00:00 2001 From: Abimael Martell Date: Fri, 17 Apr 2026 08:33:37 -0700 Subject: [PATCH] update node pacakge paths in docs (add agents file too) --- AGENTS.md | 79 ++++++++++++++++++++++++++++++++++++++++++++++++++ README.md | 4 +-- napi/README.md | 10 +++---- 3 files changed, 86 insertions(+), 7 deletions(-) create mode 100644 AGENTS.md diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..05ccaa0 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,79 @@ +# pdf-inspector + +Fast PDF text extraction to structured Markdown. CLI binary: `pdf2md`. Detection binary: `detect-pdf`. + +## Build & Test + +```bash +cargo fmt # format +cargo clippy -- -D warnings # lint (enforced, zero warnings) +cargo test # unit + integration tests (267+ unit, 73+ integration) +cargo build --release # release binary for benchmarks +``` + +All three must pass before committing. + +## Binaries + +- `pdf2md` — extract PDF → Markdown. Supports `--json` for structured output. +- `detect-pdf` — classify PDF type (TextBased/Scanned/Mixed/ImageBased). Supports `--analyze --json`. + +## Architecture + +``` +src/ + lib.rs – public API, process_pdf_with_options, encoding issue detection + detector.rs – PDF type classification, tiled-scan detection, page sampling + types.rs – TextItem, TextLine, PdfRect, PdfLine + tounicode.rs – CMap/ToUnicode parsing, CID decoding + text_utils.rs – CJK/RTL handling, Otsu threshold, ligature expansion, NFKC + extractor/ + mod.rs – top-level extraction orchestrator + content_stream.rs – PDF operator state machine (Tj/TJ/Td/Tm/q/Q) + fonts.rs – font width/encoding, CMapDecisionCache, TrueType cmap fallback + layout.rs – column detection (histogram), newspaper/tabular classification, + spanning-line pre-masking, sidebar detection + tables/ + detect_rects.rs – rect-based table detection (union-find clustering) + detect_heuristic.rs – heuristic table detection (gap-histogram, body-font tables) + detect_lines.rs – line-based table detection (H/V line grids) + grid.rs – column/row boundaries, cell assignment + format.rs – table→Markdown formatting, continuation row merging + markdown/ + convert.rs – core line→Markdown loop, struct-tree role support + analysis.rs – font stats, heading tiers, paragraph thresholds + classify.rs – line classification (header, list, code, caption) + preprocess.rs – drop cap merging, heading line merging + postprocess.rs – dot leaders, hyphenation, page numbers, URL formatting +``` + +## Key design decisions + +- **Primary audience is AI agents.** Output optimized for token efficiency and semantic quality, not visual formatting. No cosmetic padding. +- **Three table detection strategies** run in priority order: rect-based → line-based → heuristic. First valid result wins. +- **Column detection** uses horizontal projection histograms with valley detection. Multi-item spanning lines (titles, headers) are pre-masked using column-aware thresholds before column assignment. +- **Newspaper vs tabular** classification determines reading order: newspaper reads columns sequentially, tabular Y-interleaves them. +- **Tiled-scan detection** catches scanned PDFs with JBIG2/strip images where no single tile exceeds the template threshold but aggregate area does (≥2M pixels). +- **Garbage text upgrade** reclassifies Mixed PDFs as Scanned when extracted text is <50% alphanumeric. +- **Tagged PDF support** uses structure tree roles (H1-H6, P, L, Code, BlockQuote) when available, falling back to font-size heuristics. + +## Testing + +- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data. +- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`. +- **Regression suite**: sibling repo `pdf-evals` with 179+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing. + +## Debugging + +```bash +RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- file.pdf +RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- file.pdf +RUST_LOG=pdf_inspector::detector=debug cargo run --release --bin detect-pdf -- file.pdf +``` + +## Conventions + +- Clippy: use `is_some_and(...)` not `map_or(false, ...)` +- lopdf quirk: `ParseError` is private — match by string for `InvalidFileHeader` +- Column limit for tables: 25 (wide statistical tables) +- `propagate_merged_cells` skipped for >10 columns (spanning rects = background fills) diff --git a/README.md b/README.md index ed7e2ba..ebac7ec 100644 --- a/README.md +++ b/README.md @@ -55,12 +55,12 @@ print(result.markdown) # Markdown string or None ### Node.js ```bash -npm install firecrawl-pdf-inspector +npm install @firecrawl/pdf-inspector ``` ```javascript import { readFileSync } from 'fs'; -import { processPdf, classifyPdf } from 'firecrawl-pdf-inspector'; +import { processPdf, classifyPdf } from '@firecrawl/pdf-inspector'; const result = processPdf(readFileSync('document.pdf')); console.log(result.pdfType); // "TextBased", "Scanned", "ImageBased", "Mixed" diff --git a/napi/README.md b/napi/README.md index e1c987a..c8a7c27 100644 --- a/napi/README.md +++ b/napi/README.md @@ -1,4 +1,4 @@ -# firecrawl-pdf-inspector +# PDF Inspector Fast PDF classification and region-based text extraction for Node.js/Bun. Native Rust performance via [napi-rs](https://napi.rs). @@ -7,9 +7,9 @@ Built by [Firecrawl](https://firecrawl.dev) for hybrid OCR pipelines — extract ## Install ```bash -npm install firecrawl-pdf-inspector +npm install @firecrawl/pdf-inspector # or -bun add firecrawl-pdf-inspector +bun add @firecrawl/pdf-inspector ``` Prebuilt binaries included for **linux-x64** and **macOS ARM64**. No Rust toolchain needed. @@ -21,7 +21,7 @@ Prebuilt binaries included for **linux-x64** and **macOS ARM64**. No Rust toolch Classify a PDF as TextBased, Scanned, Mixed, or ImageBased (~10-50ms). Returns which pages need OCR. ```typescript -import { classifyPdf } from 'firecrawl-pdf-inspector' +import { classifyPdf } from '@firecrawl/pdf-inspector' import { readFileSync } from 'fs' const pdf = readFileSync('document.pdf') @@ -40,7 +40,7 @@ Extract text within bounding-box regions from a PDF. Designed for hybrid OCR pip Each region result includes a `needsOcr` flag that signals unreliable extraction (empty text, GID-encoded fonts, garbage text, encoding issues). ```typescript -import { extractTextInRegions } from 'firecrawl-pdf-inspector' +import { extractTextInRegions } from '@firecrawl/pdf-inspector' const result = extractTextInRegions(pdf, [ {