Commit Graph
122 Commits
Author SHA1 Message Date
Abimael MartellandClaude Opus 4.6 0aa4e0a9bc Point lopdf at firecrawl/lopdf firecrawl-fixes branch
Includes %%EOF boundary fix and inline image skip fix.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 17:29:54 -08:00
Abimael MartellandClaude Opus 4.6 3a1ada1c6b fix(reader): point lopdf at fork with %%EOF boundary fix (abimaelmartell/lopdf@88f7d9e)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 17:07:06 -08:00
Abimael Martell a2eab803cf update readme 2026-02-25 13:54:42 -08:00
Abimael MartellandClaude Opus 4.6 f37c911e8e refactor: redesign public API for better usability
- Remove `process_mode` from `MarkdownOptions` (it controlled the
  pipeline, not markdown formatting)
- Remove dead `text` field from `PdfProcessResult`
- Add `PdfOptions` builder consolidating mode, detection, markdown,
  and page filter configuration
- Add convenience functions: `detect_pdf()`, `detect_pdf_mem()`,
  `process_pdf_with_options()`, `process_pdf_mem_with_options()`
- Eliminate double document parsing: load once, share between
  detection and extraction via internal `pub(crate)` helpers
- Deprecate old `process_pdf_with_config*` functions (kept as shims)
- Update binaries to use new API

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 13:00:21 -08:00
Abimael MartellandClaude Opus 4.6 4ed73f7c3b Point lopdf at upstream repo (J-F-Liu/lopdf@ae7c584)
The R6 encryption fix has been merged upstream, so switch from the
fork back to the canonical repository.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 11:05:38 -08:00
Abimael MartellandClaude Opus 4.6 e72152ce63 Normalize typographic spaces (U+2000–U+200A) to ASCII space
EM SPACE from PDF ActualText entries was lost by text.trim(), breaking
spacing after bullets and numbered list markers. Normalizing to ASCII
space lets should_join_items detect word boundaries naturally. NBSP
(U+00A0) is excluded as it's handled by coordinate-based spacing.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-24 12:08:24 -08:00
Abimael MartellandClaude Opus 4.6 1afc1fdaa2 Strip invisible Unicode chars and expand ligatures on ActualText path
Soft hyphens, zero-width spaces, BOM, ZWNJ/ZWJ, and word joiners now get
stripped in expand_ligatures(). Also call expand_ligatures() on the ActualText
code path which was the only TextItem creation site missing it.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 16:00:08 -08:00
Abimael MartellandClaude Opus 4.6 93c5c21384 Point lopdf at fork with R6 padded /O /U fix
Use abimaelmartell/lopdf firecrawl/fix-r6-encryption branch which
accepts /O and /U values longer than 48 bytes in V5/R6 encrypted PDFs.
This unblocks 4 encrypted PDFs in the bench suite that were previously
rejected with InvalidHashLength.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:28:02 -08:00
Abimael MartellandClaude Opus 4.6 433132d195 Handle encrypted PDFs with empty user password retry
Map lopdf's Unimplemented("encrypted...") error to PdfError::Encrypted
instead of falling through to PdfError::Parse. Retry all Document::load
calls with an empty password when encryption is detected, so
owner-password-only PDFs can be opened. If the retry also fails, the
user now sees a clear "PDF is encrypted" message.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:25:42 -08:00
Abimael MartellandClaude Opus 4.6 1bacf34bd0 Add tests for row-stripe table detection
Unit test verifying short sub-header rows (e.g. month names) are not
merged as continuation rows in table formatting. Snapshot test for
2013_app2.pdf pinning the full row-stripe table output.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:43:23 -08:00
Abimael MartellandClaude Opus 4.6 dcbea3f652 Add row-stripe rect fallback for table detection
When PDF tables use full-width alternating row shading (row-stripe rects),
the normal rect grid detection fails because all rects share the same X/width,
collapsing to ~1 column. Add a fallback that uses rect Y-edges for rows and
text X-position clustering for columns, with a lower 15pt threshold to
separate narrow columns like row numbers and dates.

Also fix continuation-row merging in table formatting to not merge short
single-cell rows (≤5 chars) that are section sub-headers (e.g. month names).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:39:25 -08:00
Abimael MartellandClaude Opus 4.6 43c9a81388 Lower dollar-as-space threshold to catch chest-wall PDF pattern
The chest-wall PDF uses $ as separator but also has trailing $ after
spaces, giving only 13.7% letter-dollar-letter ratio (594 of 4332).
Add absolute count threshold (>20) alongside the ratio check so both
concentrated and dispersed substitution patterns are caught.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 08:59:19 -08:00
Abimael MartellandClaude Opus 4.6 8725e2b487 Detect broken font encodings and flag for OCR fallback
Add has_encoding_issues field to PdfProcessResult that detects garbled
text from broken ToUnicode CMaps (U+FFFD replacement characters or
systematic dollar-as-space substitution). Surfaced in JSON output so
clients can fall back to OCR for affected PDFs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 22:06:29 -08:00
Abimael MartellandClaude Opus 4.6 a7dbb16e86 Strip PUA F000-F0FF characters from decoded text
Symbol/Wingdings fonts often map codes to PUA (U+F000-F0FF) via ToUnicode
CMaps, producing invisible/tofu characters in markdown output. Add
clean_symbol_pua() post-processing to extract_text_from_operand() that
converts these to standard Unicode: bullets to U+2022, checkmark to U+2713,
and ASCII/Latin-1 range by stripping the F000 offset.

Eliminates 1,405 PUA characters across 25+ eval PDFs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-21 22:42:34 -08:00
Abimael MartellandClaude Opus 4.6 a4a2e1f267 Fix PUA glyph decoding, Symbol font cmap mapping, and UTF-8 detection
- Strip PUA F000 offset for "uniF0XX" glyph names (e.g. uniF072 → 'r')
- Use Mac Roman/Symbol cmap subtables for proper code→GID→Unicode mapping
  instead of assuming GID equals character code in subsetted fonts
- Skip fallback CMap building for fonts with explicit encoding
- Move UTF-8 detection before lopdf single-byte encoding decoder

Fixes systematic letter substitutions (CITY→CITQ), missing French letters,
and UTF-8 mojibake (José→José). Eval: missing_text -32%, encoding_issue 7→4.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-21 13:54:12 -08:00
Abimael Martell 97910e86bf fix specific decoding issues 2026-02-20 21:16:12 -08:00
Abimael Martell b23d073553 Update real-estate-pricing snapshot 2026-02-20 19:45:09 -08:00
Abimael Martell 7d6dae66df Escape invisible glyph literals 2026-02-20 19:15:52 -08:00
Abimael Martell cbad087ec7 Fix control character literals in glyph map 2026-02-20 19:14:09 -08:00
Abimael Martell 18a01d366a Inline full glyph map and refine encoding fallbacks 2026-02-20 19:13:38 -08:00
Abimael Martell 7775d35207 replace range check with includes 2026-02-20 18:57:32 -08:00
Abimael Martell 7b7bff60ea Fix encoding fallbacks for ASCII and symbol fonts 2026-02-20 18:54:53 -08:00
Abimael MartellandClaude Opus 4.6 65e305d544 feat: Improve CMap handling with binary CMap support, inline ToUnicode, and fallback decoding
- Add pdf.js binary CMap (bcmap) parser for Japan1/GB1/CNS1 CID fonts
- Support inline ToUnicode streams (not just object references)
- Add CMapDecisionCache for heuristic primary vs remapped CMap selection
- Build fallback CMaps from embedded font data and CIDSystemInfo
- Handle simple fonts without ToUnicode via embedded font cmap
- Add Symbol/Wingdings/ZapfDingbats font decoding fallback
- Support usecmap chaining in both text and binary CMap formats
- Scan XObject Forms for text operators in detector
- Change default detection strategy from EarlyExit to Sample(8)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-20 18:47:10 -08:00
Abimael MartellandClaude Opus 4.6 66c88a4fc9 fix(tounicode): Remap broken ToUnicode CMaps from subset fonts with GID mismatch
Some PDF generators subset-embed fonts by renumbering GIDs sequentially
but fail to update the ToUnicode CMap, which still references original
GID values. Detect this mismatch (Identity-H, no CIDToGIDMap, W array
starts at CID ≤ 2 but CMap min CID > 2) and remap to sequential positions.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-20 14:12:30 -08:00
Abimael MartellandClaude Opus 4.6 23056dc5ba fix: Properly escape JSON string output in CLI binaries
The hand-rolled JSON escaping only handled \, ", and \n but missed
tabs, carriage returns, and other control characters (U+0000..U+001F),
producing invalid JSON that Python's json.loads would reject. Add a
proper json_escape() function covering all required escapes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-19 10:52:31 -08:00
Abimael MartellandClaude Opus 4.6 7a5af1e9c3 feat: Add configurable ProcessMode (detect-only / analyze / full)
Introduces a ProcessMode enum that controls how far the PDF pipeline
runs, enabling fast document triage without paying extraction or
markdown conversion costs. Exposed via --detect-only and --analyze
flags in both pdf2md and detect-pdf CLIs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-19 10:06:30 -08:00
Abimael MartellandClaude Opus 4.6 e62444fee0 feat: Add layout complexity detection (tables and multi-column)
Add LayoutComplexity struct to PdfProcessResult so callers can detect
when a PDF has complex layout (tables or multi-column text) and decide
whether to use the extracted markdown or fall back to OCR.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 20:38:50 -08:00
Abimael MartellandClaude Opus 4.6 0cf80025f6 fix(tables): Prevent graph labels from merging into adjacent tables
When a PDF page has a graph directly above a table (e.g. Figure 2 above
Table II in real-estate-pricing.pdf), the heuristic table detector would
merge graph axis labels and legend items into the table, producing a
corrupted result.

Add rect hint regions: when a small cluster of cell-border rects (4-6)
fails full grid validation (e.g. only row borders, no column dividers),
extract their Y bounding box as a "hint region". The heuristic detector
then runs separately on items inside vs outside hint regions, preventing
unrelated content from being merged into tables.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 18:03:25 -08:00
Abimael MartellandClaude Opus 4.6 b6828fff8c fix(tables): Handle vertically-merged cells and protect header from continuation merge
Rects spanning multiple grid rows (e.g. classification labels) now have
their text consolidated into the first sub-row via propagate_merged_cells().
Also prevent continuation-row merge from folding body data into the header.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 16:40:57 -08:00
Abimael MartellandClaude Opus 4.6 28313e1f2d test: Add snapshot regression tests with PDF fixtures
Add 5 stripped public-domain PDF fixtures and golden markdown snapshots
for CI regression testing. Fix non-deterministic output caused by
HashMap iteration order in font stats, table heuristics, and rect
clustering by adding deterministic tie-breaking.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 16:13:55 -08:00
Abimael MartellandClaude Opus 4.6 ceb0cf5913 fix(tables): Increase rect edge snap tolerance to fix table detection
The snap_edges tolerance of 2.0 was too tight for PDFs where header and
body cell boundaries differ by ~3pt due to cell padding. This created
phantom empty columns that caused valid tables to be rejected. Increase
snap and cell coverage tolerances from 2.0/3.0 to 6.0. Also add debug
logging to rect-based table detection pipeline.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 15:49:19 -08:00
Abimael MartellandClaude Opus 4.6 421690c7c4 fix(fonts): Fix decoding for custom encodings and CID Korean fonts
Merge CMap and Differences encoding at the byte level for single-byte
fonts, fixing garbled output when partial ToUnicode CMaps caused Latin-1
fallback to block the Differences path. Add Adobe-Korea1 CID-to-Unicode
predefined mapping for Identity-H CID fonts without ToUnicode streams,
and support TrueType cmap extraction via ttf-parser for embedded fonts.
Also handle suffixed glyph names (e.g. zero.tf, a.ss01) per Adobe spec.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 15:32:11 -08:00
Abimael MartellandClaude Opus 4.6 42c1d639f2 chore: Pin lopdf to commit 0387137b
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 13:35:03 -08:00
Abimael MartellandClaude Opus 4.6 a00ce46ab0 refactor(tounicode): Replace raw byte scanning with lopdf document model for CMap extraction
Eliminates double-parsing of PDFs by using the lopdf document model exclusively
for ToUnicode CMap extraction. The old raw byte scanner parsed PDFs separately
from lopdf and only handled FlateDecode, while lopdf handles FlateDecode + LZW +
ASCII85. The new approach walks page fonts and Form XObject fonts via the document
API, yielding ~7x speedup on text-heavy PDFs and ~1.2x overall.

- Add FontCMaps::from_doc() with recursive Form XObject font walking
- Remove ~270 lines of raw byte scanning code (from_pdf_bytes, extract_stream_from_raw_pdf, etc.)
- Remove flate2 dependency (lopdf handles decompression internally)
- Remove dead CMap lookup fallback branches (by_name, get_with_obj, base_font_name)
- Remove unused font_base_names parameter from extract_text_from_operand
- Skip U+FFFD replacement characters in CMap decode (PDF notdef glyph markers)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 13:33:32 -08:00
Abimael Martell 64d2438e2e chore(debug): Add debugging binary for position issues 2026-02-18 10:38:34 -08:00
Abimael Martell 5401354f2e add logging 2026-02-18 10:34:32 -08:00
Abimael Martell 7328af91bd chore(refactor): Better organization of the codebase, split modules, update docs, add AGENTS.md 2026-02-18 10:16:11 -08:00
Abimael MartellandClaude Opus 4.6 e9c0737bd8 feat(extractor): Sequential newspaper-style column reading order
Detect newspaper-style two-column layouts (dense independent text
flows) and emit columns sequentially instead of Y-interleaving them.
This fixes academic papers where column text was being merged across
columns at the same Y position.

Key changes:
- Add is_newspaper_layout() heuristic (≥15 lines/col, >50% Y-collision)
- Add split_column_stragglers() to separate core column clusters from
  header remnants and per-word items that landed in column buckets
- Filter wide items (>60% page width) from column detection histogram
  so full-width text doesn't fill the gutter and prevent detection
- Emit: above spanning → core columns sequentially → column stragglers
  sequentially → below spanning items
- Fix markdown paragraph break to detect backward Y jumps (y_gap.abs())
  so sequential columns on the same page get proper paragraph breaks

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 21:20:45 -08:00
Abimael MartellandClaude Opus 4.6 dca267ac44 feat(extractor): Add RTL text support, vertical WMode, and Korean Hangul
Detect RTL scripts (Hebrew, Arabic, Syriac, etc.) by Unicode ranges and
sort line items by X descending when a line is predominantly RTL. Fix
gap calculation in should_join_items() to be direction-neutral. Extract
WMode from Type0 font dictionaries for vertical text foundation. Add
missing Korean Hangul ranges to is_cjk_char().

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 16:42:04 -08:00
Abimael MartellandClaude Opus 4.6 0f53de0387 feat(tables): Cluster rects with union-find for multi-table detection
PDFs like privacy notices have 6-7 distinct visual table sections per
page, each with its own set of cell rects. Previously all rects were
merged into a single giant sparse grid that failed validation. Now
spatially connected rects are clustered via union-find before running
grid detection independently on each cluster.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 16:11:00 -08:00
Abimael MartellandClaude Opus 4.6 e7455a0a3c fix(markdown): Add st ligatures and preserve TOC line breaks
- Add U+FB05/FB06 (ſt/st→st) to ligature table matching PyMuPDF's full set
- Refactor expand_ligatures to single-pass match loop (clippy fix)
- Detect dot-leader lines (TOC entries) and prevent list continuation
  merging, which was joining sub-entries onto parent numbered items
- Force newline breaks between consecutive dot-leader lines in paragraphs

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 15:32:47 -08:00
Abimael MartellandClaude Opus 4.6 e403acd7f7 fix(extractor): Y-interleave multi-column merge for correct line breaks
Replace section-based column merge (which emitted all left-column lines
then all right-column lines) with Y-interleaved merge that combines
lines at the same Y position from different columns into a single line.

This fixes tabular layouts (like court dockets) where rows span across
columns — previously hearing descriptions were detached from their
entries and line breaks between rows were lost.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 13:26:50 -08:00
Abimael MartellandClaude Opus 4.6 342eb86d06 feat(extractor): Extract AcroForm field values from fillable PDFs
Fillable PDF forms store typed values in the AcroForm dictionary, not in
the page content stream. This adds extraction of form field names and
values, emitting them as TextItems positioned at the field's Rect so they
flow naturally into the markdown pipeline.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 13:10:50 -08:00
Abimael MartellandClaude Opus 4.6 b547e685de fix(extractor): Include CTM scaling in font_size to fix two-column merge
effective_font_size() was computed from the text matrix alone, ignoring
CTM scaling. In PDFs with a small CTM scale (e.g. 0.24×) and large Tm
values (e.g. 58), font_size was inflated (58pt instead of ~14pt), causing
the merge threshold to bridge inter-column gaps and merge two-column
text into single lines. Now compute the combined matrix (text_matrix × CTM)
before calling effective_font_size() at all 6 call sites.

Also adds PdfRect extraction from `re` operators for future table-grid
detection, and related plumbing changes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 12:59:09 -08:00
Abimael MartellandClaude Opus 4.6 d35adf22d9 fix(extractor): Defer ActualText emission to EMC with correct glyph width
Emit ActualText items at EMC (end of marked content) instead of at the
first Tj/TJ inside the BDC region. This computes the width from the
text matrix delta across the entire BDC..EMC span, so ligature
replacements like "fi" get proper widths and merge correctly with
adjacent text items (e.g., "Defi nitions" → "Definitions").

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 11:32:46 -08:00
Abimael MartellandClaude Opus 4.6 3cf6970136 feat(extractor): Add Tr operator, ActualText/BDC/EMC, char-word merge, skip images
- Handle text rendering mode (Tr=3) to skip invisible OCR overlay text
- Support BDC/EMC marked content with ActualText for tagged PDFs
- Merge adjacent single-char TextItems into words at extraction layer
- Skip image XObjects instead of emitting [Image: ...] placeholders

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 11:15:25 -08:00
Abimael Martell facdf11a68 better tables 2026-02-17 10:58:00 -08:00
Abimael MartellandClaude Opus 4.6 08bc993ebd fix(extractor): Tighten per-glyph word boundary threshold for character-by-character PDFs
PDFs like SEC filings render each glyph as a separate text operator with
intra-word gaps ≈ 0 and word gaps ≈ 0.15× font_size. The previous default
threshold (0.15) was exactly at the boundary, causing words to merge
(e.g. "ofOperations", "endedJune", "furnishedas").

Reduce single-char-to-single-char alphabetic threshold from 0.15 to 0.10,
cleanly separating word boundaries from intra-word kerning.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 10:04:51 -08:00
Abimael MartellandClaude Opus 4.6 9ab3d23350 fix: Resolve clippy field_reassign_with_default warning
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 09:19:44 -08:00
Abimael MartellandClaude Opus 4.6 350bbc0086 feat(cli): Add --pages and --select-pages flags, fix CMap matching and whitespace tracking
- Add --pages flag to insert <!-- Page N --> markers between pages
- Add --select-pages flag to process only specific pages (e.g. 1,3,5-10)
- Wire MarkdownOptions and page filter through process_pdf_with_config
- Fix fuzzy CMap matching that caused Cyrillic substitution on Latin text
- Fix whitespace Tj items not advancing text matrix (broke gap detection)
- Tighten single-char fragment join threshold from 0.25 to 0.20
- Gitignore debug/diagnostic binaries and remove tracked ones

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 08:58:58 -08:00