6.9 KiB
AGENTS.md — pdf-inspector Codebase Guide
Project Overview
Rust crate (pdf-inspector) that extracts text from PDFs and converts it to structured Markdown. Ships a CLI binary pdf2md.
- Crate name:
pdf-inspector - Binary:
pdf2md(src/bin/pdf2md.rs) - PDF parsing:
lopdfcrate (v0.39.0, git dependency) - Test PDFs:
/Users/abimaelmartell/Code/pdf-evals/pdfs/
Module Map
src/
lib.rs — Public API, re-exports
types.rs — Shared types: TextItem, TextLine, PdfRect, ItemType
text_utils.rs — Character/text helpers: CJK, RTL, ligatures, bold/italic
detector.rs — Fast PDF type detection (text vs scanned) without full load
glyph_names.rs — Adobe Glyph List → Unicode mapping
tounicode.rs — ToUnicode CMap parsing for CID-encoded text
extractor/
mod.rs — Public API: extract_text, extract_text_with_positions
fonts.rs — Font width parsing, encoding, text decoding
content_stream.rs — PDF operator state machine (Tm, Td, Tj, TJ, etc.)
xobjects.rs — Form XObject and image XObject extraction
links.rs — Hyperlink and AcroForm field extraction
layout.rs — Column detection, line grouping, reading order
tables/
mod.rs — Table struct, TableDetectionMode, re-exports
detect_rects.rs — Rectangle-based table detection (union-find clustering)
detect_heuristic.rs — Heuristic table detection + validation
financial.rs — Financial token splitting for consolidated values
grid.rs — Column/row boundaries, cell assignment
format.rs — Table → Markdown formatting, footnotes
markdown/
mod.rs — MarkdownOptions, public API (to_markdown, to_markdown_from_items)
convert.rs — Core line-to-markdown loop, table/image interleaving
analysis.rs — Font statistics, heading tiers, paragraph thresholds
classify.rs — Caption, list, code detection
preprocess.rs — Heading merging, drop cap handling
postprocess.rs — Cleanup: dot leaders, hyphenation, page numbers, URLs
bin/
pdf2md.rs — CLI: PDF → Markdown
detect_pdf.rs — CLI: detect PDF type
Data Flow
PDF bytes
│
├─► detector.rs → PdfType (TextBased / Scanned / ImageBased)
│
└─► extractor/
├─ fonts.rs → font widths, encodings
├─ content_stream.rs → walk operators → Vec<TextItem> + Vec<PdfRect>
├─ xobjects.rs → Form XObject text, image placeholders
├─ links.rs → hyperlinks, AcroForm fields
└─ layout.rs → column detection → group_into_lines → Vec<TextLine>
│
├─► tables/
│ ├─ detect_rects.rs → rect-based tables (PdfRect clusters)
│ ├─ detect_heuristic.rs → heuristic tables (font-size + alignment)
│ ├─ grid.rs → column/row assignment → cells
│ └─ format.rs → Table → Markdown string
│
└─► markdown/
├─ analysis.rs → font stats, heading tiers
├─ preprocess.rs → merge headings, drop caps
├─ convert.rs → line loop + table/image insertion
├─ classify.rs → captions, lists, code
└─ postprocess.rs → cleanup → final Markdown string
Critical Implementation Details
Text Matrix Math
PDF text positioning uses two matrices: text_matrix (Tm) and line_matrix. The Td/TD operators provide offsets in text space, which must be scaled by line_matrix:
e += tx * a + ty * c
f += tx * b + ty * d
When Tm has scaling (e.g., [12,0,0,12,x,y]), failing to apply this scaling produces incorrect positions. The T* and ' operators are equivalent to 0 -TL Td and need the same treatment.
Font Size
Font size can come from the Tf operand or the Tm matrix scaling. Use effective_font_size() from text_utils.rs to get the correct value.
White-Fill Text
Text drawn with white fill (1 g before text ops) should be skipped during extraction but the text matrix must still advance to keep positions correct.
CID Fonts
Fonts named C2_* or C0_* are CID fonts that emit one word per Tj operator. Spaces must be inserted between consecutive Tj items.
lopdf Quirks
lopdf::error::ParseErroris private — match by string forInvalidFileHeader- Clippy enforces
-D warnings— useis_some_and(...)instead ofmap_or(false, ...)
Testing
cargo test # Run all 66 unit tests
cargo clippy -- -D warnings # Lint (enforced in CI)
cargo fmt --check # Format check
cargo run --release --bin pdf2md -- <file.pdf> # Smoke test
Debugging with RUST_LOG
All debug output uses structured logging via the log crate. Set RUST_LOG to control output:
RUST_LOG=pdf_inspector::extractor::content_stream=trace # raw PDF operators
RUST_LOG=pdf_inspector::extractor::fonts=debug # font metadata + encodings
RUST_LOG=pdf_inspector::tounicode=debug # CMap parsing
RUST_LOG=pdf_inspector::extractor=debug # text items per page
RUST_LOG=pdf_inspector::extractor::layout=debug # columns, reading order
RUST_LOG=pdf_inspector::markdown::analysis=debug # Y-gaps, paragraph threshold
RUST_LOG=pdf_inspector::tables=debug # table detection
RUST_LOG=pdf_inspector=debug # everything
Example: RUST_LOG=pdf_inspector::extractor::fonts=debug cargo run --bin pdf2md -- file.pdf > /dev/null
Common Tasks
| Task | Where to Edit |
|---|---|
| Fix text positioning bugs | extractor/content_stream.rs |
| Add font encoding support | extractor/fonts.rs, tounicode.rs |
| Fix column/reading order | extractor/layout.rs |
| Improve table detection | tables/detect_heuristic.rs |
| Fix table formatting | tables/format.rs, tables/grid.rs |
| Add rectangle-based tables | tables/detect_rects.rs |
| Change heading detection | markdown/analysis.rs |
| Fix list/code detection | markdown/classify.rs |
| Fix paragraph breaks | markdown/convert.rs, markdown/analysis.rs |
| Fix URL/hyphenation cleanup | markdown/postprocess.rs |
| Add new PDF type detection | detector.rs |
| Add new text item type | types.rs, then update consumers |