Commit Graph
431 Commits
Author SHA1 Message Date
Abimael MartellandClaude Opus 4.6 d5e03b8915 fix: track sc/scn color operators to avoid skipping visible text
PDFs that set fill color via `sc`/`scn` (named color space operators)
instead of `g`/`rg`/`k` were incorrectly keeping fill_is_white=true
from earlier page background fills, causing all text to be skipped.

Infer color model from operand count: 3=RGB, 4=CMYK, other=non-white.

Fixes extraction of tension_chart PDF (D'Addario string specifications)
which previously output only glyph names.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 15:04:18 -08:00
Abimael Martell ccd833dd59 remove agents 2026-03-04 14:05:38 -08:00
Abimael MartellandClaude Opus 4.6 7d73c6a1ce refactor: add PageExtraction type alias for extraction return tuples
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 11:27:31 -08:00
Abimael MartellandClaude Opus 4.6 f7afc0c439 feat: improve rect-based table detection for wide statistical tables
- Add width-based outlier filter to remove page-spanning clipping paths
  without losing row-stripe background rects
- Deduplicate sub-rects (cell-internal decorations) to prevent spurious
  Y-edge splits, with height constraint to preserve row-stripe patterns
- Raise column limit from 12 to 25 for statistical lookup tables (MWU,
  chi-square)
- Skip propagate_merged_cells for wide tables (>10 cols) where spanning
  rects are background fills, not true merged cells
- Add numeric cell check to continuation-row heuristic so short text
  labels (e.g. "Liquid", "Vapour") are still merged while numeric data
  rows are kept separate
- Raise row-stripe content density threshold from 25% to 40% to reject
  false-positive tables from alternating-shade prose sections

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 11:05:58 -08:00
Abimael MartellandClaude Opus 4.6 2b5fcd7bbb feat: detect tables from PDF line path operators (m/l/S)
Many IRS forms and government PDFs draw table gridlines using path
operators (m/l/S) instead of rectangle (re) operators. This adds
line-based table detection to capture these tables.

- Add PdfLine type for line segments from path operators
- Capture m/l/h/S/s/B/b/f/n path operators in content_stream.rs
- Thread Vec<PdfLine> through extraction pipeline
- New detect_lines.rs: classify lines, snap to grid, validate and
  assign items with extensive false-positive filters
- Integrate in markdown pipeline: rects first, then lines as fallback

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 12:58:59 -08:00
Abimael MartellandClaude Opus 4.6 8adfacd3e3 feat: enhance header/footer stripping with Y-band coalescing and page number normalization
Strip leading/trailing digit sequences (page numbers) before frequency
comparison so headers like "Page 5" and "Page 6" are treated as identical.
Group TextLines at the same Y position into Y-bands and propagate removal
to all siblings when any member is stripped. Increase EDGE_LINE_COUNT from
4 to 5 to cover 5-row form column headers (e.g., IRS p1244).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 09:46:51 -08:00
Abimael MartellandClaude Opus 4.6 eca2db119c fix: increase edge line count to catch form column headers as repeated content
Bumps EDGE_LINE_COUNT from 3 to 4 to strip repeated form column headers
that sit just inside the page margin (4th-from-edge Y position).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 00:14:47 -08:00
Abimael MartellandClaude Opus 4.6 7b6bb69feb feat: strip repeated headers/footers from markdown output
Add strip_repeated_lines() preprocessing step that removes running
headers and footers before markdown conversion. Uses multiple guard
rails to avoid false positives:

- Edge-line detection: only considers lines among the first/last 3
  distinct Y positions on each page (not percentage-based margins)
- Y-position consistency: requires low variance across pages (stddev
  < 5% of page span) to distinguish headers from scattered table content
- Position-aware removal: only strips instances at page edges, preserving
  body content that happens to match a header/footer text
- Skips structural lines (headings, list items), short lines (<10 chars),
  and decorative separators (repeated single characters)
- Frequency threshold: >= max(3, page_count * 30%) distinct pages

New strip_headers_footers option on MarkdownOptions (default: true).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 00:04:17 -08:00
Abimael Martell 8f652b81a9 update lopdf 2026-03-02 10:12:59 -08:00
Abimael MartellandClaude Opus 4.6 0aa4e0a9bc Point lopdf at firecrawl/lopdf firecrawl-fixes branch
Includes %%EOF boundary fix and inline image skip fix.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 17:29:54 -08:00
Abimael MartellandClaude Opus 4.6 3a1ada1c6b fix(reader): point lopdf at fork with %%EOF boundary fix (abimaelmartell/lopdf@88f7d9e)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 17:07:06 -08:00
Abimael Martell a2eab803cf update readme 2026-02-25 13:54:42 -08:00
Abimael MartellandClaude Opus 4.6 f37c911e8e refactor: redesign public API for better usability
- Remove `process_mode` from `MarkdownOptions` (it controlled the
  pipeline, not markdown formatting)
- Remove dead `text` field from `PdfProcessResult`
- Add `PdfOptions` builder consolidating mode, detection, markdown,
  and page filter configuration
- Add convenience functions: `detect_pdf()`, `detect_pdf_mem()`,
  `process_pdf_with_options()`, `process_pdf_mem_with_options()`
- Eliminate double document parsing: load once, share between
  detection and extraction via internal `pub(crate)` helpers
- Deprecate old `process_pdf_with_config*` functions (kept as shims)
- Update binaries to use new API

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 13:00:21 -08:00
Abimael MartellandClaude Opus 4.6 4ed73f7c3b Point lopdf at upstream repo (J-F-Liu/lopdf@ae7c584)
The R6 encryption fix has been merged upstream, so switch from the
fork back to the canonical repository.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 11:05:38 -08:00
Abimael MartellandClaude Opus 4.6 e72152ce63 Normalize typographic spaces (U+2000–U+200A) to ASCII space
EM SPACE from PDF ActualText entries was lost by text.trim(), breaking
spacing after bullets and numbered list markers. Normalizing to ASCII
space lets should_join_items detect word boundaries naturally. NBSP
(U+00A0) is excluded as it's handled by coordinate-based spacing.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-24 12:08:24 -08:00
Abimael MartellandClaude Opus 4.6 1afc1fdaa2 Strip invisible Unicode chars and expand ligatures on ActualText path
Soft hyphens, zero-width spaces, BOM, ZWNJ/ZWJ, and word joiners now get
stripped in expand_ligatures(). Also call expand_ligatures() on the ActualText
code path which was the only TextItem creation site missing it.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 16:00:08 -08:00
Abimael MartellandClaude Opus 4.6 93c5c21384 Point lopdf at fork with R6 padded /O /U fix
Use abimaelmartell/lopdf firecrawl/fix-r6-encryption branch which
accepts /O and /U values longer than 48 bytes in V5/R6 encrypted PDFs.
This unblocks 4 encrypted PDFs in the bench suite that were previously
rejected with InvalidHashLength.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 14:28:02 -08:00
Abimael MartellandClaude Opus 4.6 433132d195 Handle encrypted PDFs with empty user password retry
Map lopdf's Unimplemented("encrypted...") error to PdfError::Encrypted
instead of falling through to PdfError::Parse. Retry all Document::load
calls with an empty password when encryption is detected, so
owner-password-only PDFs can be opened. If the retry also fails, the
user now sees a clear "PDF is encrypted" message.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 10:25:42 -08:00
Abimael MartellandClaude Opus 4.6 1bacf34bd0 Add tests for row-stripe table detection
Unit test verifying short sub-header rows (e.g. month names) are not
merged as continuation rows in table formatting. Snapshot test for
2013_app2.pdf pinning the full row-stripe table output.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:43:23 -08:00
Abimael MartellandClaude Opus 4.6 dcbea3f652 Add row-stripe rect fallback for table detection
When PDF tables use full-width alternating row shading (row-stripe rects),
the normal rect grid detection fails because all rects share the same X/width,
collapsing to ~1 column. Add a fallback that uses rect Y-edges for rows and
text X-position clustering for columns, with a lower 15pt threshold to
separate narrow columns like row numbers and dates.

Also fix continuation-row merging in table formatting to not merge short
single-cell rows (≤5 chars) that are section sub-headers (e.g. month names).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 09:39:25 -08:00
Abimael MartellandClaude Opus 4.6 43c9a81388 Lower dollar-as-space threshold to catch chest-wall PDF pattern
The chest-wall PDF uses $ as separator but also has trailing $ after
spaces, giving only 13.7% letter-dollar-letter ratio (594 of 4332).
Add absolute count threshold (>20) alongside the ratio check so both
concentrated and dispersed substitution patterns are caught.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 08:59:19 -08:00
Abimael MartellandClaude Opus 4.6 8725e2b487 Detect broken font encodings and flag for OCR fallback
Add has_encoding_issues field to PdfProcessResult that detects garbled
text from broken ToUnicode CMaps (U+FFFD replacement characters or
systematic dollar-as-space substitution). Surfaced in JSON output so
clients can fall back to OCR for affected PDFs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 22:06:29 -08:00
Abimael MartellandClaude Opus 4.6 a7dbb16e86 Strip PUA F000-F0FF characters from decoded text
Symbol/Wingdings fonts often map codes to PUA (U+F000-F0FF) via ToUnicode
CMaps, producing invisible/tofu characters in markdown output. Add
clean_symbol_pua() post-processing to extract_text_from_operand() that
converts these to standard Unicode: bullets to U+2022, checkmark to U+2713,
and ASCII/Latin-1 range by stripping the F000 offset.

Eliminates 1,405 PUA characters across 25+ eval PDFs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-21 22:42:34 -08:00
Abimael MartellandClaude Opus 4.6 a4a2e1f267 Fix PUA glyph decoding, Symbol font cmap mapping, and UTF-8 detection
- Strip PUA F000 offset for "uniF0XX" glyph names (e.g. uniF072 → 'r')
- Use Mac Roman/Symbol cmap subtables for proper code→GID→Unicode mapping
  instead of assuming GID equals character code in subsetted fonts
- Skip fallback CMap building for fonts with explicit encoding
- Move UTF-8 detection before lopdf single-byte encoding decoder

Fixes systematic letter substitutions (CITY→CITQ), missing French letters,
and UTF-8 mojibake (José→José). Eval: missing_text -32%, encoding_issue 7→4.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-21 13:54:12 -08:00
Abimael Martell 97910e86bf fix specific decoding issues 2026-02-20 21:16:12 -08:00
Abimael Martell b23d073553 Update real-estate-pricing snapshot 2026-02-20 19:45:09 -08:00
Abimael Martell 7d6dae66df Escape invisible glyph literals 2026-02-20 19:15:52 -08:00
Abimael Martell cbad087ec7 Fix control character literals in glyph map 2026-02-20 19:14:09 -08:00
Abimael Martell 18a01d366a Inline full glyph map and refine encoding fallbacks 2026-02-20 19:13:38 -08:00
Abimael Martell 7775d35207 replace range check with includes 2026-02-20 18:57:32 -08:00
Abimael Martell 7b7bff60ea Fix encoding fallbacks for ASCII and symbol fonts 2026-02-20 18:54:53 -08:00
Abimael MartellandClaude Opus 4.6 65e305d544 feat: Improve CMap handling with binary CMap support, inline ToUnicode, and fallback decoding
- Add pdf.js binary CMap (bcmap) parser for Japan1/GB1/CNS1 CID fonts
- Support inline ToUnicode streams (not just object references)
- Add CMapDecisionCache for heuristic primary vs remapped CMap selection
- Build fallback CMaps from embedded font data and CIDSystemInfo
- Handle simple fonts without ToUnicode via embedded font cmap
- Add Symbol/Wingdings/ZapfDingbats font decoding fallback
- Support usecmap chaining in both text and binary CMap formats
- Scan XObject Forms for text operators in detector
- Change default detection strategy from EarlyExit to Sample(8)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-20 18:47:10 -08:00
Abimael MartellandClaude Opus 4.6 66c88a4fc9 fix(tounicode): Remap broken ToUnicode CMaps from subset fonts with GID mismatch
Some PDF generators subset-embed fonts by renumbering GIDs sequentially
but fail to update the ToUnicode CMap, which still references original
GID values. Detect this mismatch (Identity-H, no CIDToGIDMap, W array
starts at CID ≤ 2 but CMap min CID > 2) and remap to sequential positions.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-20 14:12:30 -08:00
Abimael MartellandClaude Opus 4.6 23056dc5ba fix: Properly escape JSON string output in CLI binaries
The hand-rolled JSON escaping only handled \, ", and \n but missed
tabs, carriage returns, and other control characters (U+0000..U+001F),
producing invalid JSON that Python's json.loads would reject. Add a
proper json_escape() function covering all required escapes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-19 10:52:31 -08:00
Abimael MartellandClaude Opus 4.6 7a5af1e9c3 feat: Add configurable ProcessMode (detect-only / analyze / full)
Introduces a ProcessMode enum that controls how far the PDF pipeline
runs, enabling fast document triage without paying extraction or
markdown conversion costs. Exposed via --detect-only and --analyze
flags in both pdf2md and detect-pdf CLIs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-19 10:06:30 -08:00
Abimael MartellandClaude Opus 4.6 e62444fee0 feat: Add layout complexity detection (tables and multi-column)
Add LayoutComplexity struct to PdfProcessResult so callers can detect
when a PDF has complex layout (tables or multi-column text) and decide
whether to use the extracted markdown or fall back to OCR.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 20:38:50 -08:00
Abimael MartellandClaude Opus 4.6 0cf80025f6 fix(tables): Prevent graph labels from merging into adjacent tables
When a PDF page has a graph directly above a table (e.g. Figure 2 above
Table II in real-estate-pricing.pdf), the heuristic table detector would
merge graph axis labels and legend items into the table, producing a
corrupted result.

Add rect hint regions: when a small cluster of cell-border rects (4-6)
fails full grid validation (e.g. only row borders, no column dividers),
extract their Y bounding box as a "hint region". The heuristic detector
then runs separately on items inside vs outside hint regions, preventing
unrelated content from being merged into tables.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 18:03:25 -08:00
Abimael MartellandClaude Opus 4.6 b6828fff8c fix(tables): Handle vertically-merged cells and protect header from continuation merge
Rects spanning multiple grid rows (e.g. classification labels) now have
their text consolidated into the first sub-row via propagate_merged_cells().
Also prevent continuation-row merge from folding body data into the header.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 16:40:57 -08:00
Abimael MartellandClaude Opus 4.6 28313e1f2d test: Add snapshot regression tests with PDF fixtures
Add 5 stripped public-domain PDF fixtures and golden markdown snapshots
for CI regression testing. Fix non-deterministic output caused by
HashMap iteration order in font stats, table heuristics, and rect
clustering by adding deterministic tie-breaking.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 16:13:55 -08:00
Abimael MartellandClaude Opus 4.6 ceb0cf5913 fix(tables): Increase rect edge snap tolerance to fix table detection
The snap_edges tolerance of 2.0 was too tight for PDFs where header and
body cell boundaries differ by ~3pt due to cell padding. This created
phantom empty columns that caused valid tables to be rejected. Increase
snap and cell coverage tolerances from 2.0/3.0 to 6.0. Also add debug
logging to rect-based table detection pipeline.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 15:49:19 -08:00
Abimael MartellandClaude Opus 4.6 421690c7c4 fix(fonts): Fix decoding for custom encodings and CID Korean fonts
Merge CMap and Differences encoding at the byte level for single-byte
fonts, fixing garbled output when partial ToUnicode CMaps caused Latin-1
fallback to block the Differences path. Add Adobe-Korea1 CID-to-Unicode
predefined mapping for Identity-H CID fonts without ToUnicode streams,
and support TrueType cmap extraction via ttf-parser for embedded fonts.
Also handle suffixed glyph names (e.g. zero.tf, a.ss01) per Adobe spec.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 15:32:11 -08:00
Abimael MartellandClaude Opus 4.6 42c1d639f2 chore: Pin lopdf to commit 0387137b
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 13:35:03 -08:00
Abimael MartellandClaude Opus 4.6 a00ce46ab0 refactor(tounicode): Replace raw byte scanning with lopdf document model for CMap extraction
Eliminates double-parsing of PDFs by using the lopdf document model exclusively
for ToUnicode CMap extraction. The old raw byte scanner parsed PDFs separately
from lopdf and only handled FlateDecode, while lopdf handles FlateDecode + LZW +
ASCII85. The new approach walks page fonts and Form XObject fonts via the document
API, yielding ~7x speedup on text-heavy PDFs and ~1.2x overall.

- Add FontCMaps::from_doc() with recursive Form XObject font walking
- Remove ~270 lines of raw byte scanning code (from_pdf_bytes, extract_stream_from_raw_pdf, etc.)
- Remove flate2 dependency (lopdf handles decompression internally)
- Remove dead CMap lookup fallback branches (by_name, get_with_obj, base_font_name)
- Remove unused font_base_names parameter from extract_text_from_operand
- Skip U+FFFD replacement characters in CMap decode (PDF notdef glyph markers)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 13:33:32 -08:00
Abimael Martell 64d2438e2e chore(debug): Add debugging binary for position issues 2026-02-18 10:38:34 -08:00
Abimael Martell 5401354f2e add logging 2026-02-18 10:34:32 -08:00
Abimael Martell 7328af91bd chore(refactor): Better organization of the codebase, split modules, update docs, add AGENTS.md 2026-02-18 10:16:11 -08:00
Abimael MartellandClaude Opus 4.6 e9c0737bd8 feat(extractor): Sequential newspaper-style column reading order
Detect newspaper-style two-column layouts (dense independent text
flows) and emit columns sequentially instead of Y-interleaving them.
This fixes academic papers where column text was being merged across
columns at the same Y position.

Key changes:
- Add is_newspaper_layout() heuristic (≥15 lines/col, >50% Y-collision)
- Add split_column_stragglers() to separate core column clusters from
  header remnants and per-word items that landed in column buckets
- Filter wide items (>60% page width) from column detection histogram
  so full-width text doesn't fill the gutter and prevent detection
- Emit: above spanning → core columns sequentially → column stragglers
  sequentially → below spanning items
- Fix markdown paragraph break to detect backward Y jumps (y_gap.abs())
  so sequential columns on the same page get proper paragraph breaks

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 21:20:45 -08:00
Abimael MartellandClaude Opus 4.6 dca267ac44 feat(extractor): Add RTL text support, vertical WMode, and Korean Hangul
Detect RTL scripts (Hebrew, Arabic, Syriac, etc.) by Unicode ranges and
sort line items by X descending when a line is predominantly RTL. Fix
gap calculation in should_join_items() to be direction-neutral. Extract
WMode from Type0 font dictionaries for vertical text foundation. Add
missing Korean Hangul ranges to is_cjk_char().

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 16:42:04 -08:00
Abimael MartellandClaude Opus 4.6 0f53de0387 feat(tables): Cluster rects with union-find for multi-table detection
PDFs like privacy notices have 6-7 distinct visual table sections per
page, each with its own set of cell rects. Previously all rects were
merged into a single giant sparse grid that failed validation. Now
spatially connected rects are clustered via union-find before running
grid detection independently on each cluster.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 16:11:00 -08:00
Abimael MartellandClaude Opus 4.6 e7455a0a3c fix(markdown): Add st ligatures and preserve TOC line breaks
- Add U+FB05/FB06 (ſt/st→st) to ligature table matching PyMuPDF's full set
- Refactor expand_ligatures to single-pass match loop (clippy fix)
- Detect dot-leader lines (TOC entries) and prevent list continuation
  merging, which was joining sub-entries onto parent numbered items
- Force newline breaks between consecutive dot-leader lines in paragraphs

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 15:32:47 -08:00