Commit Graph
20 Commits
Author SHA1 Message Date
Abimael Martell 64a0930f9f feat(layout): order image-anchored regions (#174)
* feat(layout): order image-anchored regions

* fix(layout): preserve region flow boundaries

* fix(layout): gate image-backed column flows
2026-07-16 12:37:54 -07:00
Abimael MartellandClaude Fable 5 4aa2c0c208 fix(headings): rescue wrapped bold headings on interleaved column pages (#164)
* fix(headings): rescue wrapped bold headings on interleaved column pages

Report pages whose columns can't be detected (6pt gutters) interleave
both columns' lines, which breaks every whitespace signal the bold
heading heuristic relies on: para_threshold inflates to ~3x line
height and a wrapped heading's own internal line gap defeats isolation.
'9.5. Adapting to the New Normal: Changing / Business Models' merged
into the following paragraph.

Four changes:
- merge_wrapped_bold_heading_groups: 2-3 consecutive all-bold
  body-size lines merge into one line when the group is isolated
  (column-locally, judged by x-overlapping lines only) or starts with
  a section number.
- Section-numbered all-bold lines ('9.5. ...') classify as headings
  without the standalone/isolation score gate.
- Line unfusing extends to uppercase-start continuations, gated on a
  bold-style mismatch between the runs (a bold heading beside regular
  body text) — same-style label rows stay joined.
- The unfuse line-side wordiness requirement drops to 2 words so a
  wrapped heading's short last line ('Business Models') still splits
  from the neighboring column.

opendataloader-bench: overall 0.8554 -> 0.8576, MHS 0.769 -> 0.777;
docs 037 +0.161, 111 +0.157, 039 +0.091, 198 +0.028, none down.
pdf-evals: 63 snapshots, composite 0.5952 -> 0.5964, sole >0.02 mover
positive. thermo-freon12 snapshot regenerated (cosmetic churn on an
already-scrambled 3-column legend).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(headings): review follow-ups — multi-component section numbers, wholly-bold line gate

Single '1. ' prefixes are ordered list items and no longer bypass
isolation; the uppercase unfuse requires the whole line bold (a
heading), not merely its last run, so mixed bold-label/value rows
stay joined.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:00:54 -07:00
Abimael MartellandClaude Fable 5 31918ff62f fix(layout): unfuse independent column runs sharing a baseline (#163)
* fix(layout): unfuse independent column runs sharing a baseline

Two-column report pages with charts fused headings into the adjacent
column's body text: the columns' ~6pt gutter is below what histogram
valley detection can safely use, so the page grouped single-column and
same-baseline items from both columns joined into one line ('6.2.
Expectations for Re-Hiring Employees' + mid-sentence text, killing
MHS and NID on the whole survey-report doc family).

Three changes:
- Line grouping splits same-baseline runs separated by a wide void
  (>3x font size, >=30pt) when the incoming run starts lowercase
  (mid-sentence continuation from another column) and both sides are
  multi-word prose. TOC page numbers, dot leaders, and table cells
  (numbered/capitalized) stay joined.
- Column detection is blind to chart-region text (tight 2pt bounds —
  wider padding ate rows adjacent to charts), via a chart-aware line
  grouping variant wired from the markdown pipeline.
- validate_and_build_columns computes its vertical span from
  histogram-eligible items only, so full-width captions no longer sink
  the overlap ratio for partial-page column regions.

opendataloader-bench: overall 0.8532 -> 0.8554, MHS 0.761 -> 0.769;
doc 038 +0.434, no regressions. pdf-evals: 34 snapshots change,
semantic composite wash (0.5749 -> 0.5748), no per-doc mover >0.015.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(layout): review follow-ups — chart-aware band-split grouping, single chart scan per page

Band-split pages now route through the chart-aware grouping too, and
the band loop reuses the precomputed page_chart_map instead of
re-scanning the rect list per page.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 11:58:14 -07:00
Abimael MartellandClaude Fable 5 57335f8bcf feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection (#125)
* feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection

Two style-recall gaps, both invisible to the existing name-based
heuristics:

1. Subset fonts with opaque BaseFont names ("Tc1", "AAAAAB+Amplitude")
   defeat is_italic_font/is_bold_font. New descriptor_style_flags reads
   the FontDescriptor (ItalicAngle beyond 4 degrees, Flags bit 7 Italic,
   bit 19 ForceBold) and, when the descriptor claims upright, falls back
   to the embedded font file: ttf-parser's OS/2 fsSelection + post
   italicAngle for sfnt fonts, and the CFF Name INDEX PostScript name
   for bare-CFF FontFile3 (descriptor rewritten to ItalicAngle 0 while
   embedding "Amplitude-LightItalic" was observed in the wild).
   ORed into is_bold/is_italic at item creation (content streams and
   form XObjects).

2. No strikeout signal existed. New is_strikeout on TextItem, detected
   in the same pass as underline: same rules pipeline (stroked lines /
   thin filled rects, table-ruling suppression), different vertical
   window — a rule crossing the glyphs at 12-55% of the em above the
   baseline instead of sitting at it. Exposed through napi and python
   bindings and pdf2md --items-json.

Verified on public ParseBench corpus docs: previously-missed italic
council titles and bold CJK itinerary headings now flagged (render-
checked); 24/508 docs gain flags, none lose any; 35 strikeout items
detected corpus-wide, disjoint from underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): quote-op advance width, Ts text rise, doc-level font style cache (PR #125 review)

Address three valid findings from review:

- The ' (move-to-next-line-and-show-text) operator emitted zero-width
  items and never advanced the text matrix, so geometric underline/
  strikeout detection (which requires width > 0) could never mark its
  text, and following show ops overlapped it. Reuse Tj's advance-width
  computation and matrix advance.

- Ts (text rise) was dropped entirely: raised/lowered runs kept the
  unshifted baseline, so rules drawn at the risen glyph position missed
  the strike/underline windows. Track rise in the text state (saved and
  restored with q/Q) and shift the rendering position through the text
  matrix's y column; advances stay on the unshifted matrix per spec.

- descriptor_style_flags re-decompressed and re-parsed the same embedded
  font program on every page whenever the descriptor left a style flag
  unset (the common case). Add a document-scoped FontStyleCache keyed by
  the FontFile2/FontFile3 object id, threaded through page and form
  extraction alongside the existing CMapDecisionCache.

The fourth finding (Form XObject rules never reach geometric detection)
is real but pre-existing for underline and needs the form walker to grow
path/paint tracking plus a new return type; deferred as a follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): ActualText items render at their glyphs' text rise (PR #125 review)

The EMC-built ActualText item used the captured text matrix without the
rise adjustment the ordinary Tj/TJ/' emission sites apply, so a tagged
run shown with Ts landed on the unshifted baseline — off the strikeout/
underline windows and inconsistent with untagged runs. The rise is
captured together with the first-glyph matrix (and at BDC for the
entry-position fallback): the item must render at the rise of its
GLYPHS, not whatever rise is set by EMC time.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): capture ActualText glyph position after the quote op's line move (PR #125 review)

The `'` handler skipped the entire suppressed-extraction block, so a
tagged span whose show op is `'` never captured its glyph matrix/rise —
the EMC item fell back to the BDC-entry matrix, which sits on the
PREVIOUS line (the `'` line move happens after BDC) with no rise. The
capture now happens right after the line move, matching the Tj/TJ
paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

* fix(review): style-boundary gate on subscript merge + strikeout suppression coverage (PR #125 review)

merge_subscript_items absorbed a script digit into its parent
regardless of underline/strikeout flags — dropping the digit's own mark
or widening the parent's over it. The merged item carries one flag, so
differing marks now break the merge, mirroring merge_text_items'
style-boundary rule (pre-existing for underline as well).

Also extends the table-suppression test to assert is_strikeout is
cleared alongside is_underline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 17:30:47 -07:00
Abimael MartellandCursor 422a2ff118 feat(extractor): geometric underline detection on TextItem (#116)
* feat(extractor): geometric underline detection on TextItem (ENG-5015)

PDFs carry no underline font flag — underlines are stroked horizontal
lines or thin filled rects drawn under the baseline. Correlate those
graphics (already parsed from the content stream) with text items in a
post-pass: a rule within ~0.35em below the baseline covering >=60% of
an item's width marks is_underline.

Exposed through the napi and python bindings. Verified on real docs:
4/4 underlined sentences flagged on a Japanese report, links/headings
flagged on 8 of 10 underline-bearing eval docs, zero flags on docs
without underlines. Known FP source (table cell borders) documented —
downstream applies inline styling only to plain-text regions.

napi 1.9.8 -> 1.9.9.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): underline rules only from painted rects, normalized extents (review)

Two review fixes: (1) normalize rect extents before the thickness/width
checks — `re` operands pass through the CTM so width/height can be
negative, which missed negative-width rules and let negative-height
bands pass as thin; (2) only feed painted rects to underline detection —
`re` rects now wait in a pending list until a paint operator (S/s, f/F/
f*, B/B*/b/b*) confirms them, and `re W n` clip-only paths are discarded
at `n`, so invisible clip boundaries no longer underline nearby text.
Marking moved into content_stream where paint state lives (pre-rotation,
consistent device space).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): harden underline detection

* feat(cli): export positioned text item json

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-08 10:20:18 -07:00
Abimael Martell 1b2e2c76d6 fix(extractor): make trace previews unicode-safe (#113) 2026-06-24 01:27:27 -06:00
Abimael Martell 8b63ceb084 emit ItemType::Image bboxes for Image XObjects (was: silently dropped) (#94)
* emit ItemType::Image bboxes for Image XObjects (was: silently dropped)

Background. ItemType::Image, MarkdownOptions::include_images, and the
markdown emitter's image-collection path have all been in the tree
for a while, but no producer ever populated them — content_stream.rs
explicitly `// Skip images — text extraction only` at the Do
operator, and the nested Form-XObject walker in xobjects.rs only
matched XObjectType::Form, silently dropping Image entries. The
declared types were dead code.

This PR lights them up. At every Do that resolves to an Image
XObject (both top-level and nested inside Form XObjects), we now
compute the page-space bbox from the current CTM via a new
`image_bbox_from_ctm` helper — handling both axis-aligned and
rotated/sheared placements via 4-corner AABB — and emit a TextItem
with `item_type: ItemType::Image` and the legacy `[Image: <name>]`
text payload that the markdown emitter already knows how to render.

Callers can now find raster figures via `extract_text_with_positions`
(and the `_mem` variant, newly re-exported at the crate root) without
needing to re-parse the PDF or run a vision/layout model. The intended
consumer is layout-aware text pipelines that want to crop figures and
caption them out-of-band.

Two backstops to avoid silent breakage for existing callers:

  1. `MarkdownOptions::include_images` default flipped `true → false`.
     If it stayed at `true`, every existing user of
     `extract_pages_markdown` would suddenly see `![Image: Im0](image)`
     placeholders inserted throughout their output the moment they
     upgraded. Image data is still available structurally via
     `extract_text_with_positions`; rendering it into markdown is now
     an opt-in. New regression test asserts `extract_pages_markdown`
     output is unchanged for the image-bearing fixture.

  2. Image items now also skip the layout heuristics
     (`detect_columns`, `detect_tables_from_rects`) via a new
     `is_text_layout_item` predicate. Without this filter, an image's
     left edge would land in the column-projection profile and skew
     table column detection — surfaced by
     `vector_grid_tests::upstage_key_functions_four_cols` going from 4
     detected columns to 5 in CI before the filter was added.

Re-exporting `extract_text_with_positions_mem` at the crate root —
strictly additive; mirrors how `extract_pages_markdown_mem` is already
available there.

Tests:

  - test_extract_text_with_positions_emits_image_bboxes — minimal PDF
    with one 200×100 image at (50, 600); asserts one Image item with
    correct bbox + page + text.
  - test_image_xobject_bbox_handles_rotated_ctm — 90° rotated image
    via shear-component CTM; asserts AABB is correct (handles non-
    axis-aligned placements via 4-corner clamp).
  - test_image_emission_does_not_change_default_markdown — asserts no
    `Image:` token leaks into default markdown output, regression
    guard for the include_images flip.
  - test_markdown_options_default_has_include_images_false — explicit
    sentinel so anyone flipping it back catches it in CI.

* Bump version from 1.8.15 to 1.9.0
2026-05-20 13:43:05 -07:00
Abimael MartellandClaude Opus 4.7 f44509ac3a fix: don't treat bullet-marker columns as real text columns (#55)
On PDFs where every list item starts with ● at the left margin and content
at a fixed offset, histogram column detection sees the gap between marker
and content as a gutter and splits each line across two phantom "columns,"
scrambling the reading order (Anthropic's Mythos system card p.73–74).

- layout: reject gutter candidates where the smaller side is ≥80%
  standalone bullet-marker glyphs (•, ●, ○, ◦, ▪, ▫, ◆, ◇, ■, □)
- markdown/classify: add starts_with_bullet_marker helper (narrower than
  is_list_item — excludes numbered/lettered patterns like 1. and a) so
  numbered section headings stay as headings)
- markdown/convert: skip heuristic heading detection on lines that start
  with a bullet marker
- markdown/classify: strip a leading bullet wrapped in a bold/italic run
  (e.g. "**● Label:**" → "- **Label:**") — some PDFs put the marker inside
  the same bold run as the label

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 22:47:34 -07:00
Abimael MartellandClaude Opus 4.6 10dd7e2881 feat: XY-cut fallback for column detection on asymmetric layouts
When the histogram-based column detector finds no valleys (common with
sidebar/asymmetric layouts), fall back to a simplified XY-cut: find the
largest horizontal gap between item edges and split there if both sides
have enough items with vertical overlap.

Inspired by opendataloader's XY-Cut++ algorithm but implemented as a
single-level fallback rather than full recursive segmentation.

Doc 156: NID 0.545→0.966, Doc 157: NID 0.564→0.962.
NID-S +0.007, TEDS-S +0.066 across 200 docs. No regressions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 12:11:45 -07:00
Abimael MartellandClaude Opus 4.6 9c463b9c95 fix: prevent multi-column text from being misdetected as tables
On pages where column detection finds 2+ columns, skip body-font
heuristic table detection in the merged-band retry path. This prevents
sidebar/two-column prose from being formatted as markdown tables.

The fix is targeted: per-band heuristic detection still runs (bands
are scoped to single columns), so real tables within columns are
still detected. Only the merged-band retry (which sees all items
across columns) is gated.

Also relaxes column validation to accept asymmetric layouts (sidebars)
where one side has fewer items, and tries center-based item assignment
before edge-based to improve column splitting for asymmetric layouts.

Benchmark: NID 0.865→0.869, NID-S 0.798→0.805, overall +0.002.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 02:01:24 -07:00
Abimael MartellandClaude Opus 4.6 fba0a644ef fix: replace partial_cmp with total_cmp to prevent NaN sort panics (#20)
* fix sort panics on NaN values from bogus PDF font metrics

Replace all `partial_cmp(...).unwrap_or(Ordering::Equal)` and bare
`partial_cmp(...).unwrap()` with `total_cmp()` across the codebase.

`partial_cmp` returns `None` for NaN, and mapping that to `Equal`
violates total ordering: `a == NaN` and `NaN == b` but `a != b`.
Rust 1.81+ detects this and panics in sort_by. `total_cmp` handles
NaN deterministically (sorts to end) and guarantees total ordering.

The critical crash was in `extract_text_in_regions` (lib.rs:478)
where PDFs with bogus font ascent/descent values produced NaN in
text item coordinates, causing process abort via NAPI.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix missed partial_cmp in layout.rs and restore napi exports

- Convert two remaining b.y.partial_cmp(&a.y) calls to total_cmp
  in group_single_column and column layout sorting
- Restore missing napi exports: detectPdf, extractText,
  extractTextWithPositions, processPdf

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 14:39:42 -07:00
Abimael MartellandClaude Opus 4.6 95aac6a7cd feat(layout): relative valley column detection for justified text (#15)
* feat(layout): relative valley column detection for justified text

Add fallback column detection using relative valley analysis for PDFs with
justified text where item widths extend past gutter boundaries. The absolute
valley detector fails on these layouts because gutter bins are at ~40% of
peak (well above the 15% noise threshold).

The relative valley detector smooths the histogram with a 5-bin moving
average, finds local minima where contrast < 0.60 of surrounding peaks,
and validates with peak balance >= 0.40. Limited to single best valley
(max 2 columns) and requires >= 100 items per page.

Tested on IRS Publication 17 (2002), a 289-page 2-column justified text
document: column detection went from ~40 pages to 165 pages.

190 passed, 0 regressions across 191 eval PDFs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(layout): tighten relative valley thresholds to reduce false positives

Reduce PEAK_WINDOW from 40 to 25 bins (50pt) so valleys are only validated
against nearby peaks, not distant ones. Add MIN_PEAK_HEIGHT of 20 (smoothed)
to reject sparse pages where histogram peaks are too low to indicate dense
two-column text.

Previous thresholds caused 13 regressions across the eval suite by splitting
tables, TOCs, checklists, and forms. Now: 188 passed, 0 regressions (2 minor
metadata-only diffs on IRS P17 and 9978293).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(layout): skip relative valley detection on pages with tables

Table column gaps in the histogram look identical to text column gutters
but the table pipeline already handles reading order for those pages.
Pass page_has_table flag through detect_columns to suppress the relative
valley fallback on pages where tables were detected.

This eliminates all remaining regressions from relative valley detection:
190 passed, 0 regressions across 191 eval PDFs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(layout): prose density validation for relative valley detection

Add columns_have_prose() to validate relative valley column splits.
Checks that both sides of a proposed split contain paragraph-like
content (fill ratio >= 40%, avg items/line <= 3.5) before committing
to a column split. Combined with the table-page guard, this prevents
false column splits on financial statements, forms, and tabular
layouts where long labels or dot leaders fill the column width.

Also tightens find_relative_valleys() thresholds (PEAK_WINDOW 40->25,
MIN_PEAK_HEIGHT 5->20) to reduce false positive valley candidates.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-24 13:11:58 -07:00
Abimael MartellandClaude Opus 4.6 442a169ddf feat(layout): pre-mask spanning lines for multi-column pages
Multi-item lines (titles, section headers) that span across detected
columns were being split into individual column buckets, corrupting
newspaper detection and reading order. Add identify_spanning_lines()
that groups items by Y-proximity and marks lines wider than 1.3x the
widest column that have no gap at a detected gutter boundary.

Uses column-aware thresholds and gutter-interval gap detection for
precision — only fires on pages with detected columns, preserves
original item order via boolean mask.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 22:22:19 -07:00
Abimael MartellandClaude Opus 4.6 d9c2143c32 feat: tagged PDF structure tree support (#4)
* feat: tagged PDF structure tree support for semantic markdown generation

Parse /StructTreeRoot from tagged PDFs and use semantic roles (H1-H6, P,
LI, BlockQuote, Code, Caption) to improve markdown output. Structure tree
headings add to font-size heuristics without suppressing them. Coverage
threshold (≥50%) ensures only properly tagged PDFs activate this path.

Phase 1: Parse structure tree with role maps, MCID collection, flattening
Phase 2: Capture MCIDs from BMC/BDC operators, tag TextItems
Phase 3: Structure-aware markdown generation in convert loop

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: accumulate consecutive code lines into single fenced block

Per-line code fencing produced broken markdown for multi-line code
blocks (separate ``` open/close per line). Unify struct-tree Code
role and font-based monospace detection into a single is_code_line
check with in_code_block state for proper accumulation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test: add tagged PDF fixture with Firecrawl docs content

Synthetic 7-page PDF with rich structure tree exercising H1, H2, H3,
P, Code, LI, Caption, TH, TD roles. Generated via fpdf2 script.
Integration test verifies struct tree parsing and code fence output.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: remove python PDF generator script from repo

Keep the generated fixture PDF but don't track the generator script.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: handle malformed bare-name struct types in tagged PDFs

Some PDF generators (e.g. fpdf2) write /S Code instead of /S /Code
in structure elements. lopdf silently drops these objects since bare
tokens are invalid PDF syntax. Add a pre-processor that scans for
known bare struct type names and prepends / before loading.

Unifies path and memory loading through the same fix pipeline.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: update lopdf dependency to main branch

The firecrawl/zlib-checksum-encrypted branch was merged and deleted.
Point to main which includes all previously merged fixes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: switch lopdf to upstream repo pinned at 845cd3d

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-17 20:50:24 -07:00
Abimael MartellandClaude Opus 4.6 76ea52680b feat(layout): detect sidebar annotations as newspaper columns
Narrow annotation columns beside wide body columns were being
Y-interleaved with body text, producing garbled reading order.
Add sidebar detection with width ratio, line balance, and sparse
density guards to correctly trigger sequential column reading.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 17:38:00 -07:00
Abimael MartellandClaude Opus 4.6 152de8b56d feat(text): adaptive join threshold for Canva-style letter-spaced PDFs
Canva-generated PDFs render text character-by-character with CSS-style
letter-spacing (~0.5-0.9× font_size). The hardcoded 0.10 threshold
caused every character to get a space inserted ("K a r i b i b").

Detect Canva pages via fix_letterspaced_items (≥50% items match "a b c"
pattern), compute an IQR-based threshold (median × 1.55) on the gap
distribution BEFORE space removal, then propagate per-page thresholds
through PageThresholds → group_into_lines_with_thresholds → TextLine
→ should_join_items.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 17:54:49 -07:00
Abimael MartellandClaude Opus 4.6 e3990dc065 feat(tables): add rect-guided calendar table builder and extractor improvements
Rect-guided tables:
- Add try_build_rect_guided_table() to build tables from rect cluster
  X positions as column boundaries, bypassing heuristic detection
- Split merged multi-number TextItems (e.g. "10 11 12...31") into
  individual day-column cells using column-advancing boundary assignment
- Interpolate missing column boundaries for holiday/non-work days
  that lack colored rects
- Strip tilde-leader noise from cells (legend text bleeding)
- Filter legend text beyond table area via max-X threshold
- Pass cluster_rects through RectHintRegion for downstream use

Extractor improvements:
- Deduplicate clip-path and fill-path rects before using as table hints
- Add font width fallback for missing glyph metrics via average width
- Improve column detection scoring to prefer balanced gutters
- Handle side-by-side layouts with hint-region derived split points

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 15:09:17 -07:00
Abimael MartellandClaude Opus 4.6 7a5af1e9c3 feat: Add configurable ProcessMode (detect-only / analyze / full)
Introduces a ProcessMode enum that controls how far the PDF pipeline
runs, enabling fast document triage without paying extraction or
markdown conversion costs. Exposed via --detect-only and --analyze
flags in both pdf2md and detect-pdf CLIs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-19 10:06:30 -08:00
Abimael Martell 5401354f2e add logging 2026-02-18 10:34:32 -08:00
Abimael Martell 7328af91bd chore(refactor): Better organization of the codebase, split modules, update docs, add AGENTS.md 2026-02-18 10:16:11 -08:00