8950a36218bc00af5140efd81f229ad310c16f30
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
39c31a8404 |
fix(underline): rescue snug-owned underlines from the table-ruling filters (#143)
* fix(underline): rescue snug-owned underlines from the table-ruling filters Documents that underline many full-width lines (dense CJK business docs, legal redlines, 10-K section links) produce span-similar rules at 3+ y-levels — exactly what the repeated-ruling filter treats as table rulings, so every semantic underline on such pages was discarded. Three changes fix detection without re-marking real tables: 1. Snug-owner rescue: a rule survives the repeated-ruling filter when the union of touching text runs on its baseline row owns it (rule contained within the union's span +0.75em, runs cover >=60% of the rule, no column-sized gaps between runs). Table row separators fail ownership: they overshoot their cells' text or match gapped items. Same-row segmented rules (column-header separators) always stay discarded, and a rule enclosed by a drawn cell-sized box (rect-grid tables) is never rescued. 2. Vertical window widened 0.35em -> 0.72em below the baseline: CJK layouts draw underlines under the full em box, measured at ~0.67em. 3. Prose-table guard in the positions-path suppressor: a detected 'table' whose cells hold flowing prose (>=30% of cells over 100 chars) is a detection artifact of boxed callouts + stacked rules, not a real table — suppressing there erased every underline on the page. Also fixes cluster_x_positions fabricating phantom table columns from style-split continuation runs (touching items, gap <2pt, now feed one column start) — the fix that keeps rect-grid table shapes stable while underlined links inside cells are correctly marked. Snapshot updates are underline gains on regulation/form fixtures and one empty spacer-column change in a subscripted header. Corpus (508-doc public bench sweep): text output byte-identical on all docs; underlined items +224/-0; strikeout now fires on redline docs. Item-level GT coverage: is_underline 86->151/405, is_strikeout 0->10/44. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB * fix(underline): fraction-bar guard + subscript merge across underline marks Found by a 202-doc real-world corpus diff (pdf-evals) that exercises the full markdown pipeline, which the bench-corpus item sweep does not: 1. Math fraction bars and lattice grid lines are underline geometry — short horizontal rules under digits. Guard: a narrow rule (<=60pt) with bar-sized text hanging just below it (denominator) never marks. The below-text width bound matters: tightly-leaded REAL underlines have a full-width next line below, which must not trip the guard. 2. merge_subscript_items refused to merge when the parent was underlined but the tiny digit was not (the drawn rule easily misses the digit's own overlap window) — losing the merge broke subscript tokens inside table cells (b+2 no longer became b₂). Strikeout boundaries still block the merge in both directions; only parent-underlined/digit-bare merges, absorbing with the parent's flags. Corpus after refinement: underlined items +220/-2 (the 2 are fraction bars the old code wrongly marked), GT rule-text coverage 149/405 underline + 10/44 strikeout, text output identical on all 508 bench docs and word-count-identical on the 202 pdf-evals docs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB * fix(underline): address review — grid-evidence veto, strikeout-safe fraction guard, bounded gaps - Cell-box veto now requires GRID EVIDENCE (a vertically abutting neighbor rect with x-overlap) instead of a height window: multiline table cells taller than the old 90pt ceiling veto again, and isolated filled callout panels (which legitimately contain underlines) no longer veto at all. - The fraction guard gates only UNDERLINE marking; rule_strikes_item still evaluates, so short strikeouts near lower text survive. - Fraction hug distance tightened to 0.3em so a short last-line at normal leading is not mistaken for a denominator. - Continuation-run suppression bounds the negative gap (-4pt): text overhanging from an adjacent cell keeps its own column start. Corpus after review fixes: underlined items +222/-2, GT coverage 150/405 underline + 10/44 strikeout, bench text output identical on all 508 docs, pdf-evals word loss bounded at equation-reflow noise. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB * chore: appease clippy (redundant closure) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b375d6f102 |
feat(markdown): underline emission, Unicode scripts, style-preserving merges (#117)
* feat(markdown): underline emission, Unicode scripts, style-preserving merges (ENG-5015 2b)
Three formatting losses in the direct-extraction markdown path:
1. text_with_formatting gains <u> run emission (detect_underline option,
default on) using the geometric is_underline flag from 1.9.9.
Underline runs stay free of nested bold/italic markers — consumers
match tag content literally. Heading lines keep plain text for
bold/italic but preserve <u>: the tag carries meaning `#` doesn't.
2. merge_subscript_items now maps absorbed digit scripts to Unicode
sub/superscript forms with direction from the baseline offset
("H"+"2" -> "H₂", "word"+raised "2" -> "word²", "m"+"3" -> "m³").
NFKC/NFKD folds these back to plain digits so text matching
downstream is unaffected; renderers keep the script semantics.
3. merge_text_items no longer merges across bold/italic boundaries —
absorbing a styled run into a plain neighbor erased the styling
before markdown emission ever saw it. On eval docs this recovers
20-82 italic runs per document that previously emitted as plain.
Snapshots regenerated (diffs are the features: CCl₂F₂, m³, underlined
legal section headings, finer bold runs). pdf-evals regression suite:
202/202 real PDFs pass. napi 1.9.9 -> 1.9.10.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(extractor): break merges at underline boundaries too (review)
OR-merging underline stretched the eventual <u> span over neighboring
plain fragments. Merge runs now break on any style-flag change, the
redundant accumulator is gone, and format_list_item learned to move
bullet markers outside <u> wrappers so fully-underlined bullet lines
still render as markdown lists. td9264 snapshot regenerated — spans are
tighter (trailing periods correctly outside the tag).
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(markdown): strip stray spaces before sentence punctuation (review)
Style-boundary item splits can strand a trailing period in its own
fragment, and multiple assembly paths join fragments with spaces,
yielding "word ." artifacts. Rather than chasing every join site, a
postprocess pass removes a space before `.`/`,`/`;` when the mark ends
its token (whitespace, cell boundary `|`, or end of text follows).
Dot leaders/ellipses and mid-token periods are untouched.
Fixes the td9264 "companies ." artifacts and two pre-existing
"armoring ," artifacts in the 2013-app2 snapshot. pdf-evals: zero
markdown diffs across all 203 corpus PDFs vs committed baselines.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(tables): trim spaces inside parenthetical cell fragments
* fix(tables): reject sparse prose row-stripe tables
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
|
||
|
|
59b17f372a |
tables: tighten prose-in-frame rejection (#77)
* tables: lift detection on shaded-header + alt-row tables (#wired-grids) Production telemetry on `wired_high_confidence`-classified table regions showed `detect_vector_grid_in_region_mem` returning a usable grid only ~27% of the time, with the rest falling through to GLM-OCR. Three surgical fixes target the dominant production shapes: * Path-fill cell backgrounds: when the page has no `re` rects but draws cell backgrounds via `m`/`l`/`h`/`f*` sequences, prefer the fill-derived rects over the few section-level `W*` clip paths that previously won the priority gate. Activated when fill rects outnumber clip rects ≥3×. * Dedup-induced cluster splits: page-background rects could pose as containers in the sub-rect dedup and evict a slightly smaller table-frame rect, breaking adjacency between column-cell groups so each column became its own cluster. Origin-anchored containers are now disqualified from sub-rect dedup. A separate exact-duplicate pass collapses the cell-padding/text-bg/cell-border triple emissions some PDFs produce, preserving original order to avoid reshuffling table output on multi-table pages. * Prose-words rejection: the `cell-rect` fallback's whole-grid prose threshold also rejected real tables that include a description column. Now relaxed when content is well-distributed (≥75% of cols filled), while keeping the original strictness for prose-in-a-frame layouts. Two regression fixtures from the opendataloader-bench corpus, covering the dominant production failure categories: * `greencomp_competence.pdf` — 2-col shaded-header + plain-body glossary. Mirrors production crops #1 (Contractions glossary) and #6 (BIO 350 course header). * `upstage_key_functions.pdf` — 4-col shaded-header + alt-row backgrounds + merged left column. Mirrors production crops #2 (Parameter/Value alt-row), #7 (Spanish XML schema), and #8 (Córdoba multi-row header). Existing fixtures stay green (doc 51 wrapped-label, doc 128 forecast six-cols, td9264 snapshot). 133 unit + integration tests pass; clippy clean. Bumps napi/package.json 1.8.4 → 1.8.5. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * tables: tighten prose-in-frame rejection — fixes pdf-evals #30 regression PR #76's shaded-header detection lift surfaced a regression on accessory_building_permit_application_1 (TEDS 0.10 → 0.05): a paragraph of legal text laid out in a 2-column justified block was being admitted as a 10×2 fake table where every cell holds a sentence fragment ("I agree to comply...", "I", "It is the property owner's responsibility..."). Per pdf-evals PR #30 review, this is the kind of regression production users will notice — the markdown is structurally and semantically misleading. Root cause: PR #76's prose-rejection only fires for `num_cols >= 4`, so the 2-col prose-in-a-frame case slipped past it entirely. The new fill-priority + dedup changes started producing rects for this layout that 1.8.4 correctly ignored. Fix: tighten the prose-in-frame check. - Lower the column-count guard from `>= 4` to `>= 2`. - Add a content-length signal as the primary discriminator: when the prose-words trigger fires AND mean non-empty cell length exceeds 65 chars, reject regardless of column distribution. The 65-char threshold cleanly separates observed cases: accessory_building (prose-in-frame): mean 74 chars → REJECT upstage_key_functions (real 4-col table): mean 53 → admit greencomp_competence (real 2-col glossary): mean 20 → admit accessory_building (real 5×3 form data): mean 10 → admit The well-distributed-cols relaxation that PR #76 added stays — "label / value / description / benefit" tables (#7, #8 from the production crops) still pass, but only when their mean cell length stays below the prose threshold. New regression test `accessory_building_rejects_prose_in_frame` asserts both that the real 5×3 form data table survives AND the 10×2 prose block is rejected. Snapshot test `test_snapshot_td9264` updated to match new output — old snapshot captured the same prose-in-frame bug on regulatory text (paragraphs emitted as 3-col `||text||` fake-table rows). New snapshot emits clean prose paragraphs, which is correct. Verification: - cargo test --all: 424 lib + 133 integration + 2 doc tests pass - cargo fmt --check clean - cargo clippy -- -D warnings clean (lib-level; pre-existing test-level clippy issues on the wired-grids branch unaffected) - Existing fixtures stay green: forecast_table_chart_six_cols (PR #72), bits_pilani_* (PR #73), greencomp_competence_two_cols and upstage_key_functions_four_cols (PR #76). This branch is based on abimaelmartell/wired-grids so it includes PR #76's commits plus this fix on top. Suggest merging this and closing #76, OR rebasing #76 to incorporate this fix. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * tables: drop early-dedup atom — caused broad TOC + matrix corruption Bisected PR #76's 4 atoms against the SEC 10-K 0001104659-25-093871_183e44ac.pdf which appeared as a TEDS regression in pdf-evals PR #30. Result: atom | TOC | perf-graph | qualifications ---------------------------|-----|------------|--------------- fill-priority | ✓ | ✓ | ✓ early-dedup | ✗ | ✗ | ✗ page-bg disqualification | ✓ | ✓ | ✓ prose-relaxation | ✓ | ✓ | ✓ Early-dedup was the SOLE source of all three regressions on this doc. Tried a more conservative variant (≥3 copies only — pair-duplicates appear in legit multi-section layouts like 10-K dividers above + below section headers); didn't fix the regression. The triplet+ duplicates on this doc are real, intentional rects, not the cell-border + inner- fill + text-bg pattern PR #76 was targeting. Drop early-dedup. Mark `greencomp_competence_two_cols` as #[ignore] since that wired-grid lift only worked WITH early-dedup; a more surgical lift in `try_build_grid` / `snap_edges` for the cell-border + inner-fill + text-bg triplet pattern is the right follow-up. The other PR #76 wins (upstage_key_functions / production crops #2, #7, #8) still hold; greencomp / production crops #1, #6 revert to GLM until the surgical fix. Validation on the regression doc: 0001104659 TOC PART II markers: 4 (matches main, was 2 with PR#76) 0001104659 perf-graph data row: 2 (matches main, was 1) 0001104659 qualifications rows: 9 (matches main, was 6) Validation on the prose-frame doc: accessory_building fake-table: 0 (matches main, was 1 with PR#76) accessory_building prose intact: 1 (matches main) cargo test --all clean, cargo fmt --check clean, cargo clippy --lib -- -D warnings clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
6819852541 |
fix: reject cell-rect "tables" that are actually prose in a framed box (#58)
The rect-based cell fallback in detect_row_stripe_table_from_cell_rects derives columns purely from text X-position clustering. When prose wraps inside a bounding-box rect (chat transcripts, stylized figures), the word-boundary gaps cluster into many spurious columns, producing a multi-column "table" that is just fragmented prose. Count cells containing common English function words (articles, prepositions, pronouns, common verbs) and reject the fallback when 20%+ of non-empty cells contain any such word. Real tabular data — labels, units, numbers, short identifiers — rarely contains these. Update the td9264 snapshot: the government document section that previously rendered as a malformed table now renders as cleaner prose + a proper CFR list. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
2876fa4b3e |
fix: tighten rect-row span check in propagate_merged_cells (#57)
* fix: don't reclassify wrapped bold list leads as headings When a numbered/bulleted list item's bold lead phrase wraps onto a second visual line, that line is all_bold + standalone, which scored above the rarity heading threshold and was emitted as #### in the middle of the item. That reset in_list, so the body continuation below picked up a stray `- ` bullet via the struct-tree LI path, shattering a single item into heading + stray bullets. Guard the font heuristic: when already inside a list, skip heading classification for lines at the list continuation indent with a Y gap within para_threshold. Structure-tree headings still win. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: tighten rect-row span check in propagate_merged_cells propagate_merged_cells used an overlap-based predicate with ±tol slop that returned true at shared row boundaries — a rect whose top exactly equals row N's bottom lies entirely below the row, yet the predicate considered it to span row N. When multiple background rects aligned on a shared Y edge (e.g. consecutive row-stripe shading), each adjacent rect would over-reach by one row, cascading labels and data from unrelated rows into a single merged cell. Replace the overlap predicate with a containment check: rect bottom at or below row bottom, rect top at or above row top (each within tol). Genuine merged-cell rects fully contain the rows they span; tangent rects do not. Update two snapshots that were encoding the old buggy output. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
dbc2de3e9b |
revert: undo table splitting changes that caused TEDS regression
Reverts commits |
||
|
|
1dcb0c6feb |
fix: improve table row splitting for stacked sub-tables
Three fixes for better table extraction from spreadsheet-exported PDFs:
1. Convert thin filled rects (< 2pt) to PdfLine objects before line-based
table detection. Many PDFs draw table borders as narrow filled rectangles
instead of stroked paths — these were invisible to the line detector.
2. Relax uniform row spacing rejection (CV 0.05 → 0.02). Spreadsheet
exports have very even row heights that were being rejected as "chart
grids".
3. Fix continuation row merging: don't merge rows where the only non-first
cell content is a long label (section headers like "Category No. 03").
Don't merge first-cell-only rows with long text ("Note: ...").
Also adds multi-Y row splitting in line-based detection and column-aware
table detection skipping for multi-column pages.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
44092bcc9e |
feat(tables): borderless table detection improvements (#17)
* feat(tables): improve heuristic detection for borderless wrapped-cell tables Three changes to the body-font heuristic detector: 1. Adaptive Y-gap in find_table_regions_strict: use median qualifying-row spacing × 3 instead of fixed 25pt. Tables with wrapped cells have larger gaps between qualifying rows (those with 3+ X-clusters). 2. Y-only region filtering: use full X range when collecting region items. The strict X bounds from qualifying rows excluded continuation lines in wrapped cells, starving find_column_boundaries of items. 3. Merged-band retry: when split_side_by_side splits a page into bands but no band produces a table, retry heuristic detection on all items merged as a single band. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(tables): gap-histogram column detection for small tables + lower avg_cells Two changes to fix PDF 045 (borderless table with narrow "No." column): 1. Extend gap-histogram column threshold to small tables: when the gap between within-column jitter and between-column spacing is >10pt (unambiguous bimodal signal), use the detected threshold even with fewer than 500 items. Previously only triggered for dense tables. 2. Lower BodyFont avg_cells_per_row minimum from 2.5 to 2.0 to handle tables with wrapped multi-line cells where continuation lines have only 1 filled cell. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tables): trim empty outer columns + relax partial H-line validation - Rect detection: trim empty first/last columns instead of rejecting the whole table. Rect edges often extend beyond text boundaries. - Line detection: accept tables with 6+ partial horizontal lines (>15% width) when <3 full-spanning lines exist. Handles tables with column-level separators instead of full-width rules. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(tables): cell-rect fallback for tables with variable-width backgrounds When rect clustering produces a grid that fails validation (empty interior columns from variable-width cell backgrounds), fall through to a new strategy: use rect Y-edges for row boundaries and text X-position clustering for columns. This handles tables like the opendataloader-bench 088-090 comparison tables where each cell has its own background rect at different widths. Also widen failed-cluster hint width cap for large clusters (≥30 rects) to allow page-spanning table regions. TEDS score on opendataloader-bench: 0.300 → 0.353. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tables): relax vertical line spanning validation for partial borders Accept tables with 4+ partial vertical lines (>10% table height) when fewer than 2 span >30%. Handles tables like opendataloader-bench 053 with column-level vertical separators that don't extend the full height. TEDS: 0.353 → 0.377 on opendataloader-bench. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tables): lower cell-rect density threshold + add validation logging - Lower cell-rect density minimum from 25% to 15% to accept sparser tables with decorative backgrounds (fixes 147). - Add debug logging to all heuristic validation paths for diagnosability. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(tables): relax body-font validations for text-only and 2-column tables Four fixes closing 71% of the TEDS gap vs opendataloader: 1. Validation 7 (table-like content): bypass numeric content requirement for tables with 3+ columns that passed all structural checks. Text-only tables (category lists, program descriptions) are legitimate. 2. Qualifying row threshold: lower from 3+ to 2+ X-clusters per row. Enables 2-column body-font table detection (fixes 166). 3. Row-stripe max cell length: raise from 500 to 2000 for 3+ column tables. Tables with paragraph descriptions in one column are valid (fixes 121). 4. Row-stripe empty-column trimming: apply the same outer-column trim as grid detection (fixes 121 column-0 rejection). TEDS: 0.377 → 0.438 on opendataloader-bench (gap: -0.056 vs odl). TEDS=0 docs: 14 → 9. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(tables): enable 2-column body-font tables + lower all minimums - Lower BodyFont minimum columns from 3 to 2 in detect_table_in_region - Lower BodyFont minimum rows from 3 to 2 - Lower avg_cells_per_row minimum from 2.0 to 1.5 (handles wrapped cells in 2-column tables) - Apply empty-outer-column trimming to row-stripe detection (not just grid) TEDS: 0.438 → 0.468 on opendataloader-bench (gap: -0.027 vs odl). TEDS=0 docs: 9 → 8. 86% of original gap closed. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(tables): text-based row fallback + fix 120 flow-chart and 188 leaderboard Three changes that push TEDS past opendataloader: 1. Cell-rect Y-edge fallback: when rects have too few Y-edges for row structure, derive rows from text Y-position clustering within the rect bounding box. Fixes flow-chart tables (120) and column-header- only rects (188). 2. Lower cell-rect minimum from 20 to 6 rects to catch smaller tables. 3. Relax validation 1 (first-column presence) from 50% to 25% of rows. Tables with wrapped model names have continuation lines without first column content. TEDS: 0.468 → 0.508 on opendataloader-bench. Now BEATS opendataloader (0.508 vs 0.494, gap=+0.014). TEDS=0 docs: 8 → 5. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(tables): wrapped-cell continuation row merging Merge rows that have fewer filled cells than the header row into the previous row. Handles wrapped multi-line cells where text overflow creates extra rows (e.g., "Direct" + "communications" → "Direct communications"). Conditions: fewer filled cells than header, more than previous row had, not a data row (numeric), not a short subheader label. TEDS: 0.508 → 0.522 on opendataloader-bench (now +0.028 vs odl). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tables): tighten cell-rect validation + fix continuation-row merging Cell-rect false positives: - Raise density threshold back to 25% (from 15%) - Add max cell length check (500 chars) to reject paragraph content - Reject disproportionate grids (>20 rows, <4 cols) Continuation-row merging: - Wide tables (5+ cols): only merge rows with ≤50% header cells - Narrow tables (2-4 cols): merge rows with fewer cells than header - Prevents merging normal data rows in large tables (6_KE_Chart) while keeping wrapped-cell merging for narrow tables (178) TEDS: 0.498 on opendataloader-bench (still +0.004 vs odl). pdf-evals: 191/192 passed, 0 regressions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
8c4181434f |
fix(text): handle Tc/Tw character and word spacing (#11)
* fix(text): handle Tc/Tw character and word spacing in text width computation PDFs using Tc (character spacing) and Tw (word spacing) operators for text justification had words incorrectly split across TextItems. The computed advance width didn't account for these spacing parameters, causing spurious spaces mid-word (e.g. "deve lopers" instead of "developers"). - Add Tc/Tw operator handling and graphics state save/restore - Incorporate char_spacing and word_spacing into compute_string_width_ts - Add adaptive merge threshold: tighter for lowercase→lowercase junctions, wider before joining punctuation - Add unit tests for Tc/Tw width computation and merge behavior Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * test(fonts): add large Tc width computation test Verifies that large character spacing values are applied in full without any artificial cap, matching PDF spec behavior. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(text): guard against Tc/Tw-inflated widths in merge and join paths Two targeted fixes to prevent character-spacing (Tc) and word-spacing (Tw) inflation from causing data quality regressions: 1. should_join_items: reject large negative gaps (< -font_size) that arise when Tc/Tw inflate item widths past adjacent items. Fixes FY_2015 merged numbers (e.g. "239.696.0" → "239.69 6.0"). 2. merge_text_items: cap effective width for gap computation when Tw inflates space-containing items beyond 0.85× font_size per char. Prevents column-level gaps from collapsing into merge range, recovering table detection for Baldwin-Edwards and similar PDFs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
258210a821 |
fix(postprocess): collapse consecutive spaces in extracted text
OCR text layers and some PDF producers emit trailing spaces on each
text item, which combine with gap-based joining to produce double
spaces ("Vice President"). Now collapses runs of 2+ spaces to single
space within lines, preserving leading indentation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
80a3b81ff9 |
perf(tables): remove cosmetic padding from markdown tables
Drop column-width padding from table output. Saves ~37% bytes on table-heavy PDFs. Markdown renders identically — padding was purely cosmetic. Optimizes token efficiency for AI agent consumers. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
28313e1f2d |
test: Add snapshot regression tests with PDF fixtures
Add 5 stripped public-domain PDF fixtures and golden markdown snapshots for CI regression testing. Fix non-deterministic output caused by HashMap iteration order in font stats, table heuristics, and rect clustering by adding deterministic tie-breaking. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |