72842a0734389389db79ac888505bea1f40f9173
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
20bb22aa3c |
fix(tables): flatten page-number tables of contents instead of gridding (#142)
* fix(tables): flatten page-number tables of contents instead of gridding
A title-based contents page ("About the Publisher vii", "Experiment #1
… 3") with no dot leaders and no section numbers was detected as a
2-column data table and rendered as a markdown grid, scrambling the
linear reading order (a top cause of NID loss on affected docs) and
scoring 0 on table structure.
Add is_page_number_toc: a narrow (2-3 col) list whose last column is
mostly page numbers (short integers or roman numerals) that are mostly
non-decreasing, with a text-title first column and NO header row (a
TOC's first row is already an entry). Such tables now route through the
existing flat-list TOC renderer.
The no-header + narrow-width + monotonic guards keep real data tables
intact — e.g. a 4-column regional table, or a 2-column "Mineral | CEC"
table with a header row and ascending values.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: bump reading-order (NID) benchmark to 0.89
Reflects the phantom-TOC fix in this PR: NID 0.88 -> 0.89 on the
200-doc benchmark. Other cells are unchanged at 2-decimal precision.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: canonical roman validation, real-first-row header check, roman page cells
- page_number_value now requires a *canonical* roman numeral (re-encode
and compare), so words like "civil"/"mix"/"ill" are no longer parsed
as page numbers.
- The no-header guard checks the actual first row's last cell instead of
the first non-empty one, so a blank header cell ("Category | ") still
rejects the TOC heuristic.
- format::is_page_number_cell recognizes canonical roman numerals, so
roman front-matter pages (vii, ix) get proper title/page separation in
the flat TOC list.
- Fix the non-monotonic test to use 5 rows so it exercises the
monotonicity guard rather than the row-count early return; add
roman-lookalike and blank-header rejection tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: share roman helper, widen length to 8, require page-span for TOC
- Extract canonical_roman_value + to_roman_lower into tables/mod.rs and
use them from both the TOC detector and the formatter, removing the
duplicated mapping/loop and keeping them in sync. The shared helper
accepts ≤8 chars, so longer front-matter numerals (xxxviii) flatten
consistently on both sides.
- Add a page-span guard to is_page_number_toc: real page numbers skip
through the document (range >> entry count), so a dense consecutive
ordinal/rank/ID column (1,2,3,…) is rejected — monotonicity alone did
not separate those data tables from contents.
Costs ~0.001 aggregate on the benchmark (NID 0.888->0.887) for the added
precision; still a clear win over baseline (NID 0.883, TEDS unchanged).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: recover consecutive-page TOCs via a title signal
The strict page-span rule rejected legitimate one-page-per-entry TOCs
(range ~= entry count). Relax it: accept any page sequence with a gap
(nearly all real contents). Only a *perfectly dense* consecutive run —
which rank/ID/ordinal columns produce, but a chapter-per-page TOC can
too — falls back to a title signal: flatten when the first-column
entries average multi-word headings, keep as a table when they are the
short single-word labels typical of leaderboards/ID lists.
Recovers the ~0.001 the range-only rule cost (NID back to 0.888) while
still rejecting dense ordinal data tables.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
42f0e52987 |
feat(site): GitHub Pages landing page (#135)
* feat(site): add GitHub Pages landing page Self-contained landing page (single index.html, no build step) plus a Pages deploy workflow that publishes site/ on push to main. Covers the pitch, install commands for all three registries, feature grid, benchmark, and tabbed quick-start for Rust/Python/Node/CLI. Benchmark numbers mirror the README's current published table; both should be refreshed together in a follow-up. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(site): add Firecrawl Parse upsell to closing section Replace the single CTA with a two-path split: run pdf-inspector locally (OSS) vs. hand scanned/OCR/at-scale documents to Firecrawl Parse (hosted). Links the OSS library back to the paid product for the cases local parsing can't cover. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(site): add Firecrawl branding (flame mark + wordmark) Firecrawl flame mark anchors the hosted-parse card; charcoal wordmark in the footer credit. Brand SVGs referenced as-is (exact colors). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(site): refresh benchmark with current main numbers pdf-inspector row updated from a fresh run on latest main: overall 0.78→0.83, tables 0.59→0.66, headings 0.57→0.74. Now within 0.01 of opendataloader overall, best tables of the group, headings on par. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |