Commit Graph
42 Commits
Author SHA1 Message Date
Abimael MartellandClaude Opus 4.6 30244078d8 fix: sort panic and missing CID garbage check in extract_text_in_regions_mem
Two bugs in collect_text_in_region / extract_text_in_regions_mem:

1. The threshold-based sort comparator in collect_text_in_region was not
   transitive, causing Rust's sort to panic on certain PDFs. Replaced with
   strict total_cmp ordering — the line-grouping phase already handles
   fuzzy Y matching via threshold.

2. The needs_ocr check was missing is_cid_garbage, so Identity-H fonts
   with CID garbage (C1 control chars, high Latin mojibake) could pass
   all quality checks and be served as real text with needs_ocr=false.

Also adds 7 integration tests for extract_text_in_regions_mem (previously
had zero coverage): basic extraction, Identity-H needs_ocr, multiple
regions, nonexistent page, empty region, invalid input, and a fast-vs-normal
comparison test across all text-based fixtures.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-02 18:13:11 -07:00
Abimael MartellandClaude Opus 4.6 57673ebb69 perf: skip expensive TrueType fallback in extract_text_in_regions
FontCMaps::from_doc can spend 4.5+ seconds decompressing and parsing
large embedded TrueType fonts for CID fonts with sparse ToUnicode
CMaps. For extract_text_in_regions (hybrid OCR pipeline), this is
unnecessary — fonts that can't be decoded cheaply will produce
empty/garbage text, triggering needs_ocr=true and GPU OCR fallback.

Changes:
- Add FontCMaps::from_doc_pages_fast() that skips TrueType font
  fallback parsing (build_fallback_cmap_for_type0) and Identity-H/V
  second pass entirely
- Add FontCMaps::from_doc_pages() for filtered page sets
- extract_text_in_regions_mem uses fast mode
- Restructure fallback chain: try cheap fallbacks first, only attempt
  expensive TrueType parsing when needed and not in fast mode

Benchmark on nihms-1771367.pdf (19-page chemistry paper):
- FontCMaps fast:  201µs
- FontCMaps slow:  4.47s
- 22,000x speedup on font parsing

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 18:13:11 -07:00
Abimael MartellandClaude Opus 4.6 fba0a644ef fix: replace partial_cmp with total_cmp to prevent NaN sort panics (#20)
* fix sort panics on NaN values from bogus PDF font metrics

Replace all `partial_cmp(...).unwrap_or(Ordering::Equal)` and bare
`partial_cmp(...).unwrap()` with `total_cmp()` across the codebase.

`partial_cmp` returns `None` for NaN, and mapping that to `Equal`
violates total ordering: `a == NaN` and `NaN == b` but `a != b`.
Rust 1.81+ detects this and panics in sort_by. `total_cmp` handles
NaN deterministically (sorts to end) and guarantees total ordering.

The critical crash was in `extract_text_in_regions` (lib.rs:478)
where PDFs with bogus font ascent/descent values produced NaN in
text item coordinates, causing process abort via NAPI.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix missed partial_cmp in layout.rs and restore napi exports

- Convert two remaining b.y.partial_cmp(&a.y) calls to total_cmp
  in group_single_column and column layout sorting
- Restore missing napi exports: detectPdf, extractText,
  extractTextWithPositions, processPdf

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 14:39:42 -07:00
Abimael Martell 5f829f5258 Merge branch 'main' into feat/python-bindings 2026-04-02 11:35:18 -07:00
Abimael MartellandClaude Opus 4.6 7d0e0b295b fix clippy: remove unnecessary f32 cast in obj_to_f32
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 10:52:29 -07:00
Abimael MartellandClaude Opus 4.6 af6ecebaa6 add per-region quality checks to extract_text_in_regions_mem
Return RegionText with a needs_ocr flag per region, set when:
- extracted text is empty (image region with no PDF text)
- page uses GID-encoded fonts (unreliable CID mapping)
- text fails garbage detection (mostly non-alphanumeric)
- text has encoding issues (U+FFFD, dollar-as-space patterns)

This lets callers skip GPU OCR only when text quality is reliable,
falling back for any region where extraction is suspect.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 10:46:38 -07:00
Abimael MartellandClaude Opus 4.6 263ed0aa42 add region-based text extraction for hybrid OCR pipelines
Add `extract_text_in_regions_mem` that takes layout-detected bounding
boxes (top-left origin, PDF points) and returns text within each region
in reading order. Designed for pipelines where a layout model detects
regions and text-based pages can skip GPU OCR by extracting text from
the PDF structure directly.

Also add `classify_pdf_mem` for lightweight PDF type classification
returning 0-indexed pages_needing_ocr.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 02:35:05 -07:00
Abimael MartellandClaude Opus 4.6 4d52d7af52 fix(extractor): content stream comment parsing + CJK mojibake detection (#18)
* fix(extractor): strip PDF comments that break lopdf content stream parsing

Some PDF generators (notably PD4ML used by school districts) embed
comments (% to end of line) in content streams. lopdf's Content::decode
parser fails to parse operators that follow comments, silently dropping
ET (end text) and Q (restore graphics state) operators. This caused
entire pages to produce 0 text items despite having valid text.

Fix: pre-process content streams to strip comments before parsing.
Comments inside string literals (parentheses) and hex strings are
preserved. The comment is replaced with a space to maintain token
separation.

Impact: fixes 13+ school district PDFs and similar PD4ML-generated
documents that were producing near-empty output (454 → 31,955 chars
for a 22-page school improvement plan).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(detect): flag sparse-extraction pages as needing OCR

When a TEXT-BASED PDF produces <50 chars/page average with <500 total
chars, flag all pages as needing OCR. This catches PDFs where the
extractable text is minimal (form templates, image-heavy layouts)
and the bulk of content requires OCR to access.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(detect): improve CID mojibake detection for Japanese/CJK PDFs

Extend is_cid_garbage to detect CID-as-Latin-1 mojibake: when ≥40%
of characters are high Latin-1 (U+00A0-00FF) and <33% are ASCII
letters, the text is likely CID values misinterpreted as Latin-1
characters (common in Japanese/CJK PDFs with broken ToUnicode CMaps).

Also add sparse-extraction OCR flagging: TEXT-BASED PDFs with
<50 chars/page and <500 total chars get all pages flagged for OCR.

Impact: Softbank Japanese PDFs now produce empty output with
pages_needing_ocr=all instead of mojibake garbage. Korean PDFs
with valid extraction (nexo-price-en) remain unaffected.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(detect): sparse extraction check only when markdown is generated

The sparse-extraction OCR check was triggering in Analyze mode where
markdown is not generated (md_len=0), causing false OCR flags on
every PDF processed via detect-pdf --analyze.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 22:58:56 -07:00
Abimael MartellandClaude Opus 4.6 b6764e7ca9 feat(tables): extract tables from tagged PDF structure tree (#17)
When a PDF has a well-formed structure tree with /Table > /TR > /TD|TH
elements linked to MCIDs, build tables directly from the semantic
hierarchy. Runs as highest-priority detection (step 0) before rect-based,
line-based, and heuristic strategies.

- Add StructTree::extract_tables() to walk the tree and collect table
  descriptors with row/cell/MCID info
- Add detect_tables_from_struct_tree() to match MCIDs to TextItems
- Reject tables with <30% MCID cell coverage (stale structure trees)
- Update 2013-app2 snapshot (struct-tree gives valid but different
  column ordering)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 12:33:39 -07:00
Abimael MartellandClaude Opus 4.6 95aac6a7cd feat(layout): relative valley column detection for justified text (#15)
* feat(layout): relative valley column detection for justified text

Add fallback column detection using relative valley analysis for PDFs with
justified text where item widths extend past gutter boundaries. The absolute
valley detector fails on these layouts because gutter bins are at ~40% of
peak (well above the 15% noise threshold).

The relative valley detector smooths the histogram with a 5-bin moving
average, finds local minima where contrast < 0.60 of surrounding peaks,
and validates with peak balance >= 0.40. Limited to single best valley
(max 2 columns) and requires >= 100 items per page.

Tested on IRS Publication 17 (2002), a 289-page 2-column justified text
document: column detection went from ~40 pages to 165 pages.

190 passed, 0 regressions across 191 eval PDFs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(layout): tighten relative valley thresholds to reduce false positives

Reduce PEAK_WINDOW from 40 to 25 bins (50pt) so valleys are only validated
against nearby peaks, not distant ones. Add MIN_PEAK_HEIGHT of 20 (smoothed)
to reject sparse pages where histogram peaks are too low to indicate dense
two-column text.

Previous thresholds caused 13 regressions across the eval suite by splitting
tables, TOCs, checklists, and forms. Now: 188 passed, 0 regressions (2 minor
metadata-only diffs on IRS P17 and 9978293).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(layout): skip relative valley detection on pages with tables

Table column gaps in the histogram look identical to text column gutters
but the table pipeline already handles reading order for those pages.
Pass page_has_table flag through detect_columns to suppress the relative
valley fallback on pages where tables were detected.

This eliminates all remaining regressions from relative valley detection:
190 passed, 0 regressions across 191 eval PDFs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(layout): prose density validation for relative valley detection

Add columns_have_prose() to validate relative valley column splits.
Checks that both sides of a proposed split contain paragraph-like
content (fill ratio >= 40%, avg items/line <= 3.5) before committing
to a column split. Combined with the table-page guard, this prevents
false column splits on financial statements, forms, and tabular
layouts where long labels or dot leaders fill the column width.

Also tightens find_relative_valleys() thresholds (PEAK_WINDOW 40->25,
MIN_PEAK_HEIGHT 5->20) to reduce false positive valley candidates.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-24 13:11:58 -07:00
Abimael MartellandClaude Opus 4.6 ad10f670bf chore: migrate from deprecated load_mem_with_password to load_mem_with_options
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 17:40:58 -07:00
Abimael MartellandClaude Opus 4.6 aa387503c3 fix(fonts): suppress garbage CID text on OCR-flagged pages (#9)
When Identity-H or Type3 fonts lack a ToUnicode CMap and the
CID-as-Unicode passthrough doesn't produce valid text, the raw CID
byte values appear as mojibake (random Latin Extended characters mixed
with C1 control codes).

The detector already flags these pages in pages_needing_ocr, but the
markdown pipeline still emitted the garbage. Now, for TextBased PDFs,
we check each OCR-flagged page's extracted text for CID garbage
(C1 control characters U+0080–U+009F at ≥5% density) and strip items
from pages that fail the check.

This is scoped to TextBased PDFs only — Mixed PDFs flag pages for OCR
due to template images, not font encoding issues.

Adds test fixture (shinagawa_identity_h.pdf) and integration test.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-23 10:35:45 -07:00
Abimael MartellandClaude Opus 4.6 967b70a788 fix: suppress markdown when all pages have gid-encoded fonts
When every page uses fonts with unresolvable gid-encoded glyphs,
the extracted text is unreliable garbage. Suppress the markdown
output so consumers know to use OCR instead.

Only triggers for gid-encoded fonts specifically, not for other
OCR signals (Identity-H without ToUnicode, template images) where
the text layer may still be partially useful.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-19 18:37:06 -07:00
Abimael MartellandClaude Opus 4.6 587c4bed95 feat(detect): flag Identity-H fonts without ToUnicode for OCR (#7)
* feat(detect): flag Identity-H fonts without ToUnicode for OCR

Cyrillic (and other non-Latin) PDFs with Type0/Identity-H encoded fonts
and no ToUnicode CMap produce garbage text from direct extraction. Two
fixes:

1. Detector: new `page_has_identity_h_no_tounicode` check adds affected
   pages to `pages_needing_ocr` regardless of PDF classification.
2. Extraction: extend garbage-text safety net to TextBased PDFs — when
   extracted text is <50% alphanumeric, drop the markdown, set
   `has_encoding_issues`, and flag all pages for OCR.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: replace integration tests with synthetic unit tests

Remove PDF fixture dependencies from detector and lib tests. Use
in-memory lopdf documents to test Identity-H/ToUnicode detection logic.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-19 15:40:48 -07:00
Abimael MartellandClaude Opus 4.6 74d416e8ce feat(detect): flag pages with gid-encoded fonts for OCR
Fonts using raw glyph ID names (gidNNNNN) in their Differences
encoding cannot be decoded to Unicode without the original font's
cmap table. Detect this pattern during font parsing and add
affected pages to pages_needing_ocr so downstream consumers
know to use OCR instead.

Fixes text extraction on PDFs like Tezukuri_Food-Menu.pdf where
the main body font (AcuminVariableConcept) uses gid-encoded
glyphs — even PyMuPDF and ODL fail on these.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-19 08:36:16 -07:00
Abimael MartellandClaude Opus 4.6 b8211db4b0 feat: extract invisible OCR text layer from Mixed/template PDFs
Scanned PDFs with OCR text layers use rendering mode 3 (invisible text)
positioned behind page images. Previously we skipped all Tr=3 text.

Now for Mixed/template PDFs, if normal extraction produces garbage or
empty output, we retry with invisible text included. This unlocks text
from OCR-generated PDFs without requiring external OCR.

Also adds Windows Unicode BMP (3,1) subtable support to the TrueType
cmap fallback, and allows TrueType fonts with explicit encoding to
extract their embedded cmap (OCR fonts often lie about encoding).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-18 11:39:38 -07:00
Abimael MartellandClaude Opus 4.6 3ff40987d6 fix(detect): classify tiled scans with garbage OCR as Scanned
Scanned PDFs with JBIG2/tiled image strips were misclassified as
TextBased because no individual image tile exceeded the template
threshold. Now checks aggregate image area per page (≥2M pixels).

Also adds is_garbage_text() check: if a Mixed/template PDF's extracted
text is predominantly non-alphanumeric (<50%), upgrade to Scanned so
callers use proper OCR instead of the garbage text layer.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-18 10:40:30 -07:00
Abimael MartellandClaude Opus 4.6 d9c2143c32 feat: tagged PDF structure tree support (#4)
* feat: tagged PDF structure tree support for semantic markdown generation

Parse /StructTreeRoot from tagged PDFs and use semantic roles (H1-H6, P,
LI, BlockQuote, Code, Caption) to improve markdown output. Structure tree
headings add to font-size heuristics without suppressing them. Coverage
threshold (≥50%) ensures only properly tagged PDFs activate this path.

Phase 1: Parse structure tree with role maps, MCID collection, flattening
Phase 2: Capture MCIDs from BMC/BDC operators, tag TextItems
Phase 3: Structure-aware markdown generation in convert loop

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: accumulate consecutive code lines into single fenced block

Per-line code fencing produced broken markdown for multi-line code
blocks (separate ``` open/close per line). Unify struct-tree Code
role and font-based monospace detection into a single is_code_line
check with in_code_block state for proper accumulation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test: add tagged PDF fixture with Firecrawl docs content

Synthetic 7-page PDF with rich structure tree exercising H1, H2, H3,
P, Code, LI, Caption, TH, TD roles. Generated via fpdf2 script.
Integration test verifies struct tree parsing and code fence output.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: remove python PDF generator script from repo

Keep the generated fixture PDF but don't track the generator script.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: handle malformed bare-name struct types in tagged PDFs

Some PDF generators (e.g. fpdf2) write /S Code instead of /S /Code
in structure elements. lopdf silently drops these objects since bare
tokens are invalid PDF syntax. Add a pre-processor that scans for
known bare struct type names and prepends / before loading.

Unifies path and memory loading through the same fix pipeline.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: update lopdf dependency to main branch

The firecrawl/zlib-checksum-encrypted branch was merged and deleted.
Point to main which includes all previously merged fixes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: switch lopdf to upstream repo pinned at 845cd3d

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-17 20:50:24 -07:00
Abimael MartellandClaude Opus 4.6 152de8b56d feat(text): adaptive join threshold for Canva-style letter-spaced PDFs
Canva-generated PDFs render text character-by-character with CSS-style
letter-spacing (~0.5-0.9× font_size). The hardcoded 0.10 threshold
caused every character to get a space inserted ("K a r i b i b").

Detect Canva pages via fix_letterspaced_items (≥50% items match "a b c"
pattern), compute an IQR-based threshold (median × 1.55) on the gap
distribution BEFORE space removal, then propagate per-page thresholds
through PageThresholds → group_into_lines_with_thresholds → TextLine
→ should_join_items.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 17:54:49 -07:00
Jacques DumoraandClaude Opus 4.6 0bf5463a1e feat: add Python bindings via PyO3
Expose the pdf-inspector Rust library as a Python package using PyO3 + maturin.
Python users can now `pip install` and use `import pdf_inspector` for PDF
classification, text extraction, and markdown conversion with native Rust speed.

Adds:
- src/python.rs: PyO3 bindings (process_pdf, detect_pdf, extract_text, etc.)
- pyproject.toml: maturin build configuration
- pdf_inspector.pyi: type stubs for IDE support
- tests/test_python.py: 21 pytest tests covering all Python API functions
- examples/basic_usage.py: example script demonstrating all features
- Updated README with Python quick start and API reference

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 17:38:53 +01:00
Abimael MartellandClaude Opus 4.6 0dba1e66ed feat: detect side-by-side tables via X-gap pre-splitting
Split pages with two independent table regions (e.g. D'Addario tension
chart) into separate bands before running rect/line/heuristic detection.
The split function finds X-position gaps ≥40pt, validates that <2% of
items cross the split point, and requires ≥30% balance on each side.

Also adds heuristic table detection fallback in compute_layout_complexity.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 17:01:18 -08:00
Abimael MartellandClaude Opus 4.6 2b5fcd7bbb feat: detect tables from PDF line path operators (m/l/S)
Many IRS forms and government PDFs draw table gridlines using path
operators (m/l/S) instead of rectangle (re) operators. This adds
line-based table detection to capture these tables.

- Add PdfLine type for line segments from path operators
- Capture m/l/h/S/s/B/b/f/n path operators in content_stream.rs
- Thread Vec<PdfLine> through extraction pipeline
- New detect_lines.rs: classify lines, snap to grid, validate and
  assign items with extensive false-positive filters
- Integrate in markdown pipeline: rects first, then lines as fallback

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 12:58:59 -08:00
Abimael MartellandClaude Opus 4.6 f37c911e8e refactor: redesign public API for better usability
- Remove `process_mode` from `MarkdownOptions` (it controlled the
  pipeline, not markdown formatting)
- Remove dead `text` field from `PdfProcessResult`
- Add `PdfOptions` builder consolidating mode, detection, markdown,
  and page filter configuration
- Add convenience functions: `detect_pdf()`, `detect_pdf_mem()`,
  `process_pdf_with_options()`, `process_pdf_mem_with_options()`
- Eliminate double document parsing: load once, share between
  detection and extraction via internal `pub(crate)` helpers
- Deprecate old `process_pdf_with_config*` functions (kept as shims)
- Update binaries to use new API

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 13:00:21 -08:00
Abimael MartellandClaude Opus 4.6 433132d195 Handle encrypted PDFs with empty user password retry
Map lopdf's Unimplemented("encrypted...") error to PdfError::Encrypted
instead of falling through to PdfError::Parse. Retry all Document::load
calls with an empty password when encryption is detected, so
owner-password-only PDFs can be opened. If the retry also fails, the
user now sees a clear "PDF is encrypted" message.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:25:42 -08:00
Abimael MartellandClaude Opus 4.6 43c9a81388 Lower dollar-as-space threshold to catch chest-wall PDF pattern
The chest-wall PDF uses $ as separator but also has trailing $ after
spaces, giving only 13.7% letter-dollar-letter ratio (594 of 4332).
Add absolute count threshold (>20) alongside the ratio check so both
concentrated and dispersed substitution patterns are caught.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 08:59:19 -08:00
Abimael MartellandClaude Opus 4.6 8725e2b487 Detect broken font encodings and flag for OCR fallback
Add has_encoding_issues field to PdfProcessResult that detects garbled
text from broken ToUnicode CMaps (U+FFFD replacement characters or
systematic dollar-as-space substitution). Surfaced in JSON output so
clients can fall back to OCR for affected PDFs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 22:06:29 -08:00
Abimael MartellandClaude Opus 4.6 7a5af1e9c3 feat: Add configurable ProcessMode (detect-only / analyze / full)
Introduces a ProcessMode enum that controls how far the PDF pipeline
runs, enabling fast document triage without paying extraction or
markdown conversion costs. Exposed via --detect-only and --analyze
flags in both pdf2md and detect-pdf CLIs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-19 10:06:30 -08:00
Abimael MartellandClaude Opus 4.6 e62444fee0 feat: Add layout complexity detection (tables and multi-column)
Add LayoutComplexity struct to PdfProcessResult so callers can detect
when a PDF has complex layout (tables or multi-column text) and decide
whether to use the extracted markdown or fall back to OCR.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 20:38:50 -08:00
Abimael MartellandClaude Opus 4.6 421690c7c4 fix(fonts): Fix decoding for custom encodings and CID Korean fonts
Merge CMap and Differences encoding at the byte level for single-byte
fonts, fixing garbled output when partial ToUnicode CMaps caused Latin-1
fallback to block the Differences path. Add Adobe-Korea1 CID-to-Unicode
predefined mapping for Identity-H CID fonts without ToUnicode streams,
and support TrueType cmap extraction via ttf-parser for embedded fonts.
Also handle suffixed glyph names (e.g. zero.tf, a.ss01) per Adobe spec.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 15:32:11 -08:00
Abimael Martell 7328af91bd chore(refactor): Better organization of the codebase, split modules, update docs, add AGENTS.md 2026-02-18 10:16:11 -08:00
Abimael MartellandClaude Opus 4.6 b547e685de fix(extractor): Include CTM scaling in font_size to fix two-column merge
effective_font_size() was computed from the text matrix alone, ignoring
CTM scaling. In PDFs with a small CTM scale (e.g. 0.24×) and large Tm
values (e.g. 58), font_size was inflated (58pt instead of ~14pt), causing
the merge threshold to bridge inter-column gaps and merge two-column
text into single lines. Now compute the combined matrix (text_matrix × CTM)
before calling effective_font_size() at all 6 call sites.

Also adds PdfRect extraction from `re` operators for future table-grid
detection, and related plumbing changes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 12:59:09 -08:00
Abimael MartellandClaude Opus 4.6 350bbc0086 feat(cli): Add --pages and --select-pages flags, fix CMap matching and whitespace tracking
- Add --pages flag to insert <!-- Page N --> markers between pages
- Add --select-pages flag to process only specific pages (e.g. 1,3,5-10)
- Wire MarkdownOptions and page filter through process_pdf_with_config
- Fix fuzzy CMap matching that caused Cyrillic substitution on Latin text
- Fix whitespace Tj items not advancing text matrix (broke gap detection)
- Tighten single-char fragment join threshold from 0.25 to 0.20
- Gitignore debug/diagnostic binaries and remove tracked ones

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 08:58:58 -08:00
Abimael MartellandClaude Opus 4.6 2b0e6b4153 feat(detector): Add configurable ScanStrategy for page selection
Replace hardcoded max_pages_to_sample with a ScanStrategy enum that
supports EarlyExit (default), Full, Sample(n), and Pages(vec) modes.
Add process_pdf_with_config and process_pdf_mem_with_config to the
public API. Update README with usage examples and strategy docs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-16 16:50:38 -08:00
Abimael MartellandClaude Opus 4.6 f54a296c44 feat(detector): Scan all pages with early-exit for reliable classification
Change max_pages_to_sample from 5 to u32::MAX so every page is analyzed.
This prevents misclassifying Mixed PDFs as TextBased when scanned pages
fall outside the old 5-page sample window.

Add early-exit: stop scanning as soon as a non-text page is found, since
the PDF can't be purely TextBased. A 492-page mixed PDF exits after 2
pages instead of scanning all 492.

Also add title and confidence fields to PdfProcessResult for downstream
consumers (NAPI wrapper, feature-flag gating).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-16 14:45:07 -08:00
Abimael MartellandClaude Opus 4.6 c815de9844 feat(detector): Add per-page OCR routing with pages_needing_ocr field
For mixed PDFs, callers can now see exactly which pages need OCR instead
of re-analyzing the document. Phase 2 scan iterates all pages for Mixed
PDFs (caching sampled results), while TextBased gets empty and
Scanned/ImageBased gets all pages.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-16 10:13:52 -08:00
Abimael MartellandClaude Opus 4.6 fb4d168882 feat(errors): Add NotAPdf error variant for graceful non-PDF file handling
Validate files against %PDF- magic before parsing, returning a
machine-readable NotAPdf error with a hint about the actual file type
(HTML, XML, JSON, PNG, JPEG, ZIP, plain text). Improves From<lopdf::Error>
with structured matching for IO, encryption, and structural errors.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-15 21:06:51 -08:00
Abimael Martell 68643e0c37 fix encoding, and spacing on custom fonts 2026-02-11 12:17:42 -08:00
Abimael Martell ec96311a65 implement ToUnicode CMap support for proper text extraction from PDFs with custom font encodings 2026-02-09 09:46:07 -08:00
Abimael Martell 15981b7a15 add table detection 2026-02-07 19:39:42 -08:00
Abimael MartellandClaude Opus 4.5 c14e26495d Improve text extraction with visual reading order and header detection
Major changes:
- Switch from lopdf.extract_text() to position-aware extraction
- Text is now sorted by visual reading order (top→bottom, left→right)
- Fixed font size calculation to account for text matrix scaling
- Headers are now detected based on font size ratios
- Skip very short text (≤3 chars) for header detection to avoid drop caps

This significantly improves output quality for PDFs with complex layouts.
Before: Title appeared at end, no structure detected
After: Correct reading order, headers marked with # syntax

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-02-07 13:34:37 -08:00
Abimael Martell 0ef7b2adfc add ci, plus new package 2026-02-06 21:39:21 -08:00
Abimael MartellandClaude Opus 4.5 135ce518c1 Initial commit: Rust PDF-to-Markdown library
- Smart PDF type detection (text vs scanned) without full document load
- Text extraction using lopdf directly
- Markdown conversion with header/list/code detection
- CLI tools: detect-pdf, pdf2md

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-02-06 11:51:41 -08:00