Compare commits

...
Author SHA1 Message Date
Abimael Martell e8a7529e27 fix TSR cell text assignment for overlapping bboxes
Made-with: Cursor
2026-04-26 16:13:34 -07:00
Abimael MartellandClaude Opus 4.7 5ade93440b TSR follow-ups: header-aware separator, cells API, v1.6.0
- cells_to_markdown emits the separator after the LAST row that contains
  is_header=true cells, falling back to "after row 0" when no header is
  flagged. Multi-row theads now render correctly. Three new unit tests
  cover: multi-row header, header not on row 0, no headers (fallback).
- New public extract_tables_with_structure_cells_mem returning
  Vec<Vec<StructuredCell>> so callers can drive their own rendering or
  debug overlays without re-doing the parse + extraction. The markdown
  variant now wraps it. The previously-unused page_pt_bbox field is
  surfaced through this API.
- New napi binding extractTablesWithStructureCells + StructuredCellJs.
- Bump @firecrawl/pdf-inspector to 1.6.0.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 00:53:09 -07:00
Abimael MartellandClaude Opus 4.7 cbd45e96d9 feat: TSR-aware table extraction (extract_tables_with_structure_mem)
New public function that consumes raw structure-recovery output (HTML
structure tokens + per-cell bboxes from a model like SLANet) and assembles
markdown tables by pulling cell text from the native PDF — no OCR, no
geometry inference.

Why: the existing extract_tables_in_regions_mem infers grid geometry from
text positions only and can't distinguish merged cells from multiple narrow
columns. Pairing structure recovery from a layout/TSR model with native
PDF text gets perfect text quality with proper row/col/span structure.

- New module src/tables/structured.rs: token state machine, polygon→AABB,
  crop-px→page-pt, rowspan/colspan-aware cell layout, markdown emitter.
  Accepts both 4-element rects and 8-element 4-corner polygons.
- New public extract_tables_with_structure_mem in src/lib.rs that reuses
  extract_page_text_items, region_overlaps_item, and the shared region
  text-collection helper. No existing public function modified.
- napi binding extractTablesWithStructure mirroring the existing
  extractTablesInRegions shape (f64 in JS → f32 internally).
- 14 unit tests + 5 integration tests, including a real-PDF gold-standard
  match against bits_pilani_feedback.pdf.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 00:41:42 -07:00
Abimael Martell 5c4c6e8d33 Table improvements (v1.5.0) 2026-04-22 09:52:15 -07:00
Abimael Martell d7fb493a61 fix tagged table header recovery (#60) 2026-04-22 09:51:15 -07:00
Abimael MartellandClaude Opus 4.7 6819852541 fix: reject cell-rect "tables" that are actually prose in a framed box (#58)
The rect-based cell fallback in detect_row_stripe_table_from_cell_rects
derives columns purely from text X-position clustering. When prose
wraps inside a bounding-box rect (chat transcripts, stylized figures),
the word-boundary gaps cluster into many spurious columns, producing
a multi-column "table" that is just fragmented prose.

Count cells containing common English function words (articles,
prepositions, pronouns, common verbs) and reject the fallback when
20%+ of non-empty cells contain any such word. Real tabular data —
labels, units, numbers, short identifiers — rarely contains these.

Update the td9264 snapshot: the government document section that
previously rendered as a malformed table now renders as cleaner
prose + a proper CFR list.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 10:23:28 -07:00
Abimael MartellandClaude Opus 4.7 2876fa4b3e fix: tighten rect-row span check in propagate_merged_cells (#57)
* fix: don't reclassify wrapped bold list leads as headings

When a numbered/bulleted list item's bold lead phrase wraps onto a
second visual line, that line is all_bold + standalone, which scored
above the rarity heading threshold and was emitted as #### in the
middle of the item. That reset in_list, so the body continuation
below picked up a stray `- ` bullet via the struct-tree LI path,
shattering a single item into heading + stray bullets.

Guard the font heuristic: when already inside a list, skip heading
classification for lines at the list continuation indent with a Y
gap within para_threshold. Structure-tree headings still win.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: tighten rect-row span check in propagate_merged_cells

propagate_merged_cells used an overlap-based predicate with ±tol
slop that returned true at shared row boundaries — a rect whose
top exactly equals row N's bottom lies entirely below the row, yet
the predicate considered it to span row N. When multiple background
rects aligned on a shared Y edge (e.g. consecutive row-stripe
shading), each adjacent rect would over-reach by one row, cascading
labels and data from unrelated rows into a single merged cell.

Replace the overlap predicate with a containment check: rect bottom
at or below row bottom, rect top at or above row top (each within
tol). Genuine merged-cell rects fully contain the rows they span;
tangent rects do not.

Update two snapshots that were encoding the old buggy output.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 09:33:43 -07:00
Abimael MartellandClaude Opus 4.7 5859ed651b fix: don't reclassify wrapped bold list leads as headings (#56)
When a numbered/bulleted list item's bold lead phrase wraps onto a
second visual line, that line is all_bold + standalone, which scored
above the rarity heading threshold and was emitted as #### in the
middle of the item. That reset in_list, so the body continuation
below picked up a stray `- ` bullet via the struct-tree LI path,
shattering a single item into heading + stray bullets.

Guard the font heuristic: when already inside a list, skip heading
classification for lines at the list continuation indent with a Y
gap within para_threshold. Structure-tree headings still win.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 08:47:47 -07:00
Abimael Martell df023474f1 update eval instructions 2026-04-20 22:50:57 -07:00
Abimael MartellandClaude Opus 4.7 f44509ac3a fix: don't treat bullet-marker columns as real text columns (#55)
On PDFs where every list item starts with ● at the left margin and content
at a fixed offset, histogram column detection sees the gap between marker
and content as a gutter and splits each line across two phantom "columns,"
scrambling the reading order (Anthropic's Mythos system card p.73–74).

- layout: reject gutter candidates where the smaller side is ≥80%
  standalone bullet-marker glyphs (•, ●, ○, ◦, ▪, ▫, ◆, ◇, ■, □)
- markdown/classify: add starts_with_bullet_marker helper (narrower than
  is_list_item — excludes numbered/lettered patterns like 1. and a) so
  numbered section headings stay as headings)
- markdown/convert: skip heuristic heading detection on lines that start
  with a bullet marker
- markdown/classify: strip a leading bullet wrapped in a bold/italic run
  (e.g. "**● Label:**" → "- **Label:**") — some PDFs put the marker inside
  the same bold run as the label

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 22:47:34 -07:00
Abimael MartellandClaude Opus 4.7 2e933ba8c1 fix: don't split tagged-LI continuation lines into separate bullets (#54)
Some tagged PDFs use a flat style where each wrapped visual line of a list
item gets its own MCID tagged directly under /LI. The LI branch was
unconditionally prefixing every such line with "- ", turning continuation
lines into their own bullets. Only emit a new bullet when we're not already
inside a list; otherwise fall through to the existing continuation logic so
the text is appended to the previous item.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 21:59:53 -07:00
Abimael MartellandClaude Opus 4.7 4b5ae91f54 feat: expose per-page markdown extraction to Python and Node (#53)
* feat: expose per-page markdown extraction to Python and Node (#49)

Implements the feature requested in issue #49: a list-of-pages markdown output
from the Python API. Matching the existing project pattern, the feature lives
in the Rust core and is surfaced through every binding.

- Rust core: `extract_pages_markdown` (path) and `extract_pages_markdown_mem`
  (bytes) now take `Option<&[u32]>` — `None` returns every page in document
  order; a slice restricts and preserves caller order.
- Python: new `extract_pages_markdown(path, pages=None)` and
  `extract_pages_markdown_bytes(data, pages=None)` functions plus
  `PageMarkdown` / `PagesExtractionResult` classes; stub file updated.
- Node: `extractPagesMarkdown(buffer, pages?)` — `pages` is now optional.
- Tests: 2 new Rust integration tests, 9 new Python tests, 2 new Node
  assertions. All 372 unit + 107 integration + 53 Python tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump version from 1.3.0 to 1.4.0

Minor bump for the new per-page markdown extraction API exposed through
the Python and Node bindings.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 17:19:03 -07:00
Abimael MartellandClaude Opus 4.7 b4c34ba4b7 refactor: classify Table as Data vs Toc once at construction (#52)
The TOC/data-table distinction was being recomputed at every consumer:
- format.rs::table_to_markdown ran is_table_of_contents to decide between
  flat-list and markdown-table rendering.
- compute_layout_complexity ran it again to filter TOCs out of
  pages_with_tables.
- detect_heuristic validations used it to decide whether to relax val 1/9.

Each caller had to remember tables can be either kind, which leaks the
TOC concept across the codebase.

Add `TableKind { Data, Toc }` and a `Table::new` constructor that classifies
once from the cells. All five detectors (heuristic, rect, line, struct,
columns) now go through `Table::new`. Consumers match on `kind` instead of
re-running classification.

Pure refactor — no behavior change. Verified: pdf-evals output is byte-for-
byte identical (0 changed snapshots).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 14:04:50 -07:00
Abimael MartellandClaude Opus 4.7 462a7fdfde fix: exclude TOC pages from pages_with_tables metadata (#51)
* fix: detect hierarchical-indent TOCs as tables

Validation 1 (≥25% of rows have first col) and validation 9 (paragraph
content) were rejecting TOC pages where top-level chapters indent at the
leftmost X but subsections cascade further right. The leftmost column ends
up sparse (only chapter rows land there) and most cells are empty, but the
structure is still an unambiguous TOC.

Skip both validations when cells form a TOC pattern AND the table is narrow
(≤5 cols). The width cap preserves existing handling of wide multi-column
TOCs (e.g. back-of-book 2-up indices) where format_toc_as_list would mash
adjacent visual entries together.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: exclude TOC pages from pages_with_tables metadata

TOC detection routes through the table pipeline (detected as a table, then
formatted as a flat list with tab-aligned page numbers via
format_toc_as_list). That made TOC pages appear in LayoutComplexity's
pages_with_tables, which:

- Misleads downstream consumers that read this field as "this page has a
  data table".
- Trips the table-page guard in column detection, which switches to a
  different valley-detection threshold for table pages.

Add an is_table_of_contents check in compute_layout_complexity so TOC-shaped
detections don't count toward the table flag. The TOC still renders correctly
as a flat list — only the metadata classification changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 13:48:36 -07:00
Abimael MartellandClaude Opus 4.7 9abcbbb359 fix: detect hierarchical-indent TOCs as tables (#50)
Validation 1 (≥25% of rows have first col) and validation 9 (paragraph
content) were rejecting TOC pages where top-level chapters indent at the
leftmost X but subsections cascade further right. The leftmost column ends
up sparse (only chapter rows land there) and most cells are empty, but the
structure is still an unambiguous TOC.

Skip both validations when cells form a TOC pattern AND the table is narrow
(≤5 cols). The width cap preserves existing handling of wide multi-column
TOCs (e.g. back-of-book 2-up indices) where format_toc_as_list would mash
adjacent visual entries together.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 13:41:59 -07:00
Abimael Martell 0b3ba8784c Bump version from 1.2.0 to 1.3.0 2026-04-20 11:09:16 -07:00
25 changed files with 3822 additions and 211 deletions
+5
View File
@@ -35,3 +35,8 @@ scripts/
# Test output
test_output/
# Python
__pycache__/
*.pyc
.pytest_cache/
+2 -1
View File
@@ -61,7 +61,8 @@ src/
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
- **Regression suite**: sibling repo `pdf-evals` with 179+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
- **Regression suite**: sibling repo `pdf-evals` with 187+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
## Debugging
+14
View File
@@ -42,6 +42,14 @@ text = pdf_inspector.extract_text("document.pdf")
items = pdf_inspector.extract_text_with_positions("document.pdf")
for item in items[:5]:
print(f"'{item.text}' at ({item.x:.0f}, {item.y:.0f}) size={item.font_size}")
# Per-page markdown (one Markdown string per page, plus layout metadata)
result = pdf_inspector.extract_pages_markdown("document.pdf")
for page in result.pages:
print(f"Page {page.page}: {len(page.markdown)} chars, needs_ocr={page.needs_ocr}")
# Restrict to specific 0-indexed pages (preserves caller order)
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
```
## API reference
@@ -60,6 +68,8 @@ for item in items[:5]:
| `extract_text_with_positions_bytes(data, pages=None)` | Text with positions from bytes |
| `extract_text_in_regions(path, page_regions)` | Extract text in bounding-box regions |
| `extract_text_in_regions_bytes(data, page_regions)` | Region extraction from bytes |
| `extract_pages_markdown(path, pages=None)` | Per-page Markdown + layout metadata (all pages by default) |
| `extract_pages_markdown_bytes(data, pages=None)` | Per-page Markdown from bytes |
## Types
@@ -72,3 +82,7 @@ for item in items[:5]:
**`RegionText` fields:** `text`, `needs_ocr`
**`PageRegionTexts` fields:** `page` (0-indexed), `regions` (list of RegionText)
**`PageMarkdown` fields:** `page` (0-indexed), `markdown`, `needs_ocr`
**`PagesExtractionResult` fields:** `pages` (list of PageMarkdown), `pages_with_tables` (1-indexed), `pages_with_columns` (1-indexed), `pages_needing_ocr` (1-indexed), `is_complex`
+25
View File
@@ -79,6 +79,27 @@ let bytes = std::fs::read("document.pdf")?;
let result = process_pdf_mem(&bytes)?;
```
Extract per-page Markdown (one string per page, plus document-wide layout
metadata):
```rust
use pdf_inspector::extract_pages_markdown;
// Pass `None` for every page in document order, or a slice of 0-indexed
// pages to restrict the output (caller-supplied order is preserved).
let result = extract_pages_markdown("document.pdf", None)?;
for page in &result.pages {
if page.needs_ocr {
// Route this page to OCR
} else {
println!("Page {}: {}", page.page, page.markdown);
}
}
println!("Complex layout? {}", result.is_complex);
```
## Processing modes
| Mode | What it does | Returns |
@@ -102,6 +123,8 @@ let result = process_pdf_mem(&bytes)?;
| `to_markdown(text, options)` | Convert plain text to Markdown |
| `to_markdown_from_items(items, options)` | Markdown from pre-extracted `TextItem`s |
| `to_markdown_from_items_with_rects(items, options, rects)` | Markdown with rectangle-based table detection |
| `extract_pages_markdown(path, pages)` | Per-page Markdown + layout metadata (file) |
| `extract_pages_markdown_mem(bytes, pages)` | Per-page Markdown from bytes |
Low-level detection functions are also available via the `detector` module (`detect_pdf_type`, `detect_pdf_type_with_config`, etc.) for callers who need `PdfTypeResult` instead of `PdfProcessResult`.
@@ -119,4 +142,6 @@ Low-level detection functions are also available via the `detector` module (`det
| `LayoutComplexity` | Layout analysis: is_complex, pages_with_tables, pages_with_columns |
| `TextItem` | Text with position, font info, and page number |
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
| `PdfError` | `Io`, `Parse`, `Encrypted`, `InvalidStructure`, `NotAPdf` |
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.2.0",
"version": "1.6.0",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
+143 -4
View File
@@ -317,6 +317,141 @@ pub fn extract_tables_in_regions(
})
}
/// One cropped table region plus its raw structure-recovery output, for
/// `extractTablesWithStructure`.
///
/// `structureTokens` and `cellBboxes` are typically produced by an external
/// table-structure recognition model (e.g. SLANet on PaddleOCR) running on
/// a rendered crop of the page. pdf-inspector uses the structure to lay out
/// the cells and pulls the cell text from the native PDF — no OCR involved.
#[napi(object)]
pub struct TsrTableInputJs {
/// 0-indexed page number where the crop was taken from.
pub page: u32,
/// Crop bbox on the page, `[x1, y1, x2, y2]` in PDF points with
/// top-left origin.
pub crop_pdf_pt_bbox: Vec<f64>,
/// DPI the crop image was rendered at (e.g. `200.0`).
pub render_dpi: f64,
/// Raw structure tokens emitted by the TSR model, in document order.
pub structure_tokens: Vec<String>,
/// One bbox per cell (in document order). May be 4-element
/// `[x1,y1,x2,y2]` or 8-element 4-corner polygon, in crop image-pixel
/// space.
pub cell_bboxes: Vec<Vec<f64>>,
}
/// Extract markdown tables using externally-supplied structure recovery.
///
/// For each input, pairs structure tokens with cell bboxes (rowspan/colspan
/// aware), converts each cell bbox from crop image-pixels into page PDF
/// points, pulls the cell's text from the native PDF, and emits a markdown
/// pipe-table.
///
/// Returns one markdown string per input, in input order.
#[napi]
pub fn extract_tables_with_structure(
buffer: Buffer,
inputs: Vec<TsrTableInputJs>,
) -> Result<Vec<String>> {
let bytes: Vec<u8> = buffer.to_vec();
let parsed = parse_tsr_inputs(&inputs);
catch_panic("extract_tables_with_structure", move || {
pdf_inspector::extract_tables_with_structure_mem(&bytes, &parsed)
.map_err(|e| to_napi_err(e, "extract_tables_with_structure"))
})
}
/// One resolved cell from `extractTablesWithStructureCells`.
#[napi(object)]
pub struct StructuredCellJs {
/// 0-indexed grid row.
pub row: u32,
/// 0-indexed grid column.
pub col: u32,
/// 1 for a normal cell.
pub rowspan: u32,
/// 1 for a normal cell.
pub colspan: u32,
/// `true` when the cell is a `<th>` or sits inside `<thead>`.
pub is_header: bool,
/// Text extracted from the native PDF for this cell (may be empty).
pub text: String,
/// Axis-aligned bbox `[x1, y1, x2, y2]` in page PDF-points, top-left
/// origin. Useful for debug overlays or per-cell post-processing.
pub page_pt_bbox: Vec<f64>,
}
/// Extract structured cells using externally-supplied structure recovery.
///
/// Lower-level sibling of [`extractTablesWithStructure`]: instead of
/// rendering markdown, returns the resolved cells (row, col, rowspan,
/// colspan, isHeader, text, pagePtBbox) so callers can drive their own
/// rendering, debug overlays, or per-cell post-processing.
///
/// Returns one `Array<StructuredCellJs>` per input, in input order.
#[napi]
pub fn extract_tables_with_structure_cells(
buffer: Buffer,
inputs: Vec<TsrTableInputJs>,
) -> Result<Vec<Vec<StructuredCellJs>>> {
let bytes: Vec<u8> = buffer.to_vec();
let parsed = parse_tsr_inputs(&inputs);
catch_panic("extract_tables_with_structure_cells", move || {
let result = pdf_inspector::extract_tables_with_structure_cells_mem(&bytes, &parsed)
.map_err(|e| to_napi_err(e, "extract_tables_with_structure_cells"))?;
Ok(result
.into_iter()
.map(|cells| {
cells
.into_iter()
.map(|c| StructuredCellJs {
row: c.row as u32,
col: c.col as u32,
rowspan: c.rowspan as u32,
colspan: c.colspan as u32,
is_header: c.is_header,
text: c.text,
page_pt_bbox: c.page_pt_bbox.iter().map(|v| *v as f64).collect(),
})
.collect()
})
.collect())
})
}
fn parse_tsr_inputs(inputs: &[TsrTableInputJs]) -> Vec<pdf_inspector::TsrTableInput> {
inputs
.iter()
.map(|i| {
let crop = if i.crop_pdf_pt_bbox.len() == 4 {
[
i.crop_pdf_pt_bbox[0] as f32,
i.crop_pdf_pt_bbox[1] as f32,
i.crop_pdf_pt_bbox[2] as f32,
i.crop_pdf_pt_bbox[3] as f32,
]
} else {
[0.0, 0.0, 0.0, 0.0]
};
let cell_bboxes: Vec<Vec<f32>> = i
.cell_bboxes
.iter()
.map(|bb| bb.iter().map(|v| *v as f32).collect())
.collect();
pdf_inspector::TsrTableInput {
page: i.page,
crop_pdf_pt_bbox: crop,
render_dpi: i.render_dpi as f32,
structure_tokens: i.structure_tokens.clone(),
cell_bboxes,
}
})
.collect()
}
/// Per-page markdown extraction result.
#[napi(object)]
pub struct PageMarkdownResult {
@@ -343,20 +478,24 @@ pub struct PagesExtractionResult {
pub is_complex: bool,
}
/// Extract formatted markdown for specific pages of a PDF, with layout
/// classification metadata.
/// Extract formatted markdown for pages of a PDF, with layout classification
/// metadata.
///
/// Returns per-page markdown and classification data (tables, columns,
/// OCR needs) from a single parse. Font statistics are computed from the
/// full document so header detection is consistent across pages.
///
/// Omit `pages` (or pass `undefined`) to return every page in document
/// order. Pass an array of 0-indexed page numbers to restrict output to
/// those pages, in caller-supplied order.
#[napi]
pub fn extract_pages_markdown(
buffer: Buffer,
pages: Vec<u32>,
pages: Option<Vec<u32>>,
) -> Result<PagesExtractionResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_pages_markdown", move || {
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, &pages)
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, pages.as_deref())
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
+23
View File
@@ -7,6 +7,7 @@ import {
extractText,
extractTextWithPositions,
extractTextInRegions,
extractPagesMarkdown,
} from './index.js';
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
@@ -89,6 +90,28 @@ assert.equal(typeof regionResults[0].regions[0].text, 'string');
assert.equal(typeof regionResults[0].regions[0].needsOcr, 'boolean');
console.log(' extractTextInRegions: OK');
// --- extractPagesMarkdown ---
console.log('Testing extractPagesMarkdown...');
// omit pages → every page in document order
const allPages = extractPagesMarkdown(fixture);
assert.equal(allPages.pages.length, 3);
assert.deepEqual(allPages.pages.map(p => p.page), [0, 1, 2]);
assert.ok(typeof allPages.pages[0].markdown === 'string');
assert.equal(typeof allPages.pages[0].needsOcr, 'boolean');
assert.ok(Array.isArray(allPages.pagesWithTables));
assert.ok(Array.isArray(allPages.pagesWithColumns));
assert.ok(Array.isArray(allPages.pagesNeedingOcr));
assert.equal(typeof allPages.isComplex, 'boolean');
console.log(' extractPagesMarkdown (no pages arg): OK');
// selected pages preserve caller order
const picked = extractPagesMarkdown(fixture, [2, 0]);
assert.equal(picked.pages.length, 2);
assert.equal(picked.pages[0].page, 2);
assert.equal(picked.pages[1].page, 0);
console.log(' extractPagesMarkdown with pages: OK');
// --- Error handling ---
console.log('Testing error handling...');
assert.throws(() => processPdf(Buffer.from('not a pdf')), /process_pdf/);
+50
View File
@@ -52,6 +52,28 @@ class PageRegionTexts:
"""0-indexed page number."""
regions: list[RegionText]
class PageMarkdown:
"""Per-page markdown extraction result."""
page: int
"""0-indexed page number."""
markdown: str
"""Formatted markdown for this page (empty string when needs_ocr is True)."""
needs_ocr: bool
"""True when text on this page is unreliable and OCR should be used instead."""
class PagesExtractionResult:
"""Per-page markdown output with document-wide layout classification."""
pages: list[PageMarkdown]
"""Per-page markdown results, in the order requested."""
pages_with_tables: list[int]
"""1-indexed pages where tables were detected."""
pages_with_columns: list[int]
"""1-indexed pages where multi-column layout was detected."""
pages_needing_ocr: list[int]
"""1-indexed pages that need OCR."""
is_complex: bool
"""True if any page has tables or multi-column layout."""
def process_pdf(path: str, pages: Optional[list[int]] = None) -> PdfResult:
"""Process a PDF: detect type, extract text, convert to Markdown."""
...
@@ -115,3 +137,31 @@ def extract_text_in_regions_bytes(
page_regions: List of (page_0indexed, [[x1, y1, x2, y2], ...]) tuples.
"""
...
def extract_pages_markdown(
path: str,
pages: Optional[list[int]] = None,
) -> PagesExtractionResult:
"""Extract formatted markdown for pages of a PDF, with layout classification.
Args:
path: Path to the PDF file.
pages: Optional list of 0-indexed pages. When ``None`` (default), every
page is returned in document order. Otherwise, output matches the
caller-supplied order.
Returns:
PagesExtractionResult with per-page markdown and document-wide layout
classification (tables, columns, OCR needs).
"""
...
def extract_pages_markdown_bytes(
data: bytes,
pages: Optional[list[int]] = None,
) -> PagesExtractionResult:
"""Extract formatted markdown for pages of a PDF from bytes.
See :func:`extract_pages_markdown` for details.
"""
...
+97
View File
@@ -626,6 +626,33 @@ fn find_relative_valleys(
valleys
}
/// Detect whether a side of a gutter consists predominantly of list-marker
/// glyphs (•, ●, ○, ◦, ▪, ▫, ◆, ◇). A column of bullets on the left margin
/// creates a spurious histogram valley between the bullet and the content.
/// Treating it as a real column splits each list item's text across two
/// "columns," so we reject these candidates.
fn is_list_marker_column(items: &[&&TextItem]) -> bool {
const LIST_MARKERS: &[char] = &['•', '●', '○', '◦', '▪', '▫', '◆', '◇', '■', '□'];
if items.is_empty() {
return false;
}
let marker_count = items
.iter()
.filter(|i| {
let t = i.text.trim();
let mut chars = t.chars();
match (chars.next(), chars.next()) {
(Some(c), None) => LIST_MARKERS.contains(&c),
_ => false,
}
})
.count();
// Require ≥80% of items on this side to be standalone markers. A handful
// of non-marker items (stray page numbers, footnote refs) shouldn't
// defeat the check.
marker_count as f32 / items.len() as f32 >= 0.8
}
/// Validate valley candidates with vertical consistency checks and build column regions.
///
/// When `center_assign` is true, items are assigned to columns based on their
@@ -694,6 +721,19 @@ fn validate_and_build_columns(
continue;
}
// Reject valleys where the smaller side is just a column of list
// markers (bullets aligned at the left margin). This is a common
// pattern in PDFs where ● starts each list item: histogram detection
// sees the gap between bullet and content as a gutter.
let smaller_items: &[&&TextItem] = if left_items.len() <= right_items.len() {
&left_items
} else {
&right_items
};
if is_list_marker_column(smaller_items) {
continue;
}
// Check vertical overlap
if y_range > 0.0 {
let left_y_min = left_items.iter().map(|i| i.y).fold(f32::INFINITY, f32::min);
@@ -1780,6 +1820,63 @@ mod tests {
);
}
#[test]
fn bullet_marker_column_not_detected_as_column() {
// Pattern: every line is `● <content>`, with ● at x=90 and content
// starting at x=104. Histogram detection sees a gutter between them
// and would split the page into a "bullet column" and "content column",
// scrambling every list item.
let mut items = Vec::new();
for i in 0..15 {
let y = 750.0 - i as f32 * 30.0;
items.push(make_item(1, 90.0, y, ""));
items.push(make_item(
1,
104.0,
y,
"FullContentLineTextHere________________",
));
}
// Pad with content to satisfy min item count for column detection.
for i in 0..15 {
let y = 300.0 - i as f32 * 14.0;
items.push(make_item(1, 72.0, y, "FootnoteText_____________________"));
}
let cols = detect_columns(&items, 1, false);
assert_eq!(
cols.len(),
1,
"Bullet markers aligned at left margin should not be treated as their own column"
);
}
#[test]
fn is_list_marker_column_detects_bullets() {
let items = vec![
make_item(1, 90.0, 100.0, ""),
make_item(1, 90.0, 114.0, ""),
make_item(1, 90.0, 128.0, ""),
make_item(1, 90.0, 142.0, ""),
];
let refs: Vec<&TextItem> = items.iter().collect();
let wrapped: Vec<&&TextItem> = refs.iter().collect();
assert!(is_list_marker_column(&wrapped));
}
#[test]
fn is_list_marker_column_rejects_prose() {
let items = vec![
make_item(1, 30.0, 100.0, "Regular prose line"),
make_item(1, 30.0, 114.0, "Another sentence"),
make_item(1, 30.0, 128.0, "Third line"),
make_item(1, 30.0, 142.0, "Fourth line"),
];
let refs: Vec<&TextItem> = items.iter().collect();
let wrapped: Vec<&&TextItem> = refs.iter().collect();
assert!(!is_list_marker_column(&wrapped));
}
#[test]
fn premask_narrow_line_not_masked() {
// Items that form a line spanning only ~40% of column width → not masked
+386 -7
View File
@@ -336,13 +336,17 @@ pub struct PagesExtractionResult {
pub is_complex: bool,
}
/// Extract formatted markdown for specific pages of a PDF, with layout
/// Extract formatted markdown for pages of a PDF, with layout
/// classification metadata.
///
/// Unlike [`process_pdf_mem`] which returns one concatenated markdown string,
/// this returns per-page markdown so callers can mix direct extraction
/// (for simple text pages) with GPU OCR (for complex/scanned pages).
///
/// When `pages` is `None`, every page (0-indexed, in document order) is
/// returned. When `Some(&[...])`, only the listed 0-indexed pages are
/// returned, in the caller's order.
///
/// Font statistics are computed from the full document so header
/// detection thresholds are consistent regardless of which pages are
/// requested. Per-page `needs_ocr` is set when the page has GID-encoded
@@ -352,7 +356,7 @@ pub struct PagesExtractionResult {
/// at near-zero cost since the items/rects/lines are already in memory.
pub fn extract_pages_markdown_mem(
buffer: &[u8],
pages: &[u32],
pages: Option<&[u32]>,
) -> Result<PagesExtractionResult, PdfError> {
validate_pdf_bytes(buffer)?;
let (doc, page_count) = load_document_from_mem(buffer)?;
@@ -368,10 +372,20 @@ pub fn extract_pages_markdown_mem(
// Compute font stats from full document (cross-page consistency).
let font_stats = markdown::analysis::calculate_font_stats_from_items(&all_items);
let mut results = Vec::with_capacity(pages.len());
// When caller doesn't specify pages, return every page in document order.
let all_pages: Vec<u32>;
let pages_slice: &[u32] = match pages {
Some(p) => p,
None => {
all_pages = (0..page_count).collect();
&all_pages
}
};
let mut results = Vec::with_capacity(pages_slice.len());
let mut pages_needing_ocr = Vec::new();
for &page_0idx in pages {
for &page_0idx in pages_slice {
// Out-of-range pages → empty + needs_ocr
if page_0idx >= page_count {
pages_needing_ocr.push(page_0idx + 1);
@@ -444,6 +458,20 @@ pub fn extract_pages_markdown_mem(
})
}
/// Path-based wrapper for [`extract_pages_markdown_mem`].
///
/// Reads the PDF from disk and extracts per-page markdown. Pass `None` for
/// `pages` to return every page in document order, or `Some(&[...])` to
/// restrict to specific 0-indexed pages (in caller-supplied order).
pub fn extract_pages_markdown<P: AsRef<Path>>(
path: P,
pages: Option<&[u32]>,
) -> Result<PagesExtractionResult, PdfError> {
validate_pdf_file(&path)?;
let buffer = std::fs::read(path.as_ref())?;
extract_pages_markdown_mem(&buffer, pages)
}
// =========================================================================
// Region-based text extraction (for hybrid OCR pipelines)
// =========================================================================
@@ -756,6 +784,191 @@ pub fn extract_tables_in_regions_mem(
Ok(results)
}
// =========================================================================
// Region-based table extraction with external structure recovery (TSR)
// =========================================================================
/// Input for [`extract_tables_with_structure_mem`]: one cropped table region
/// plus the raw structure-recovery output for it.
///
/// The structure tokens and bboxes are typically produced by an external
/// table-structure recognition model (e.g. SLANet on PaddleOCR) running on
/// a rendered crop of the page. pdf-inspector uses the structure to lay out
/// the cells and pulls the cell text from the native PDF — no OCR involved.
#[derive(Debug, Clone)]
pub struct TsrTableInput {
/// 0-indexed page number where the crop was taken from.
pub page: u32,
/// Crop bbox on the page, `[x1, y1, x2, y2]` in PDF points with
/// **top-left origin** (matches the layout model's coordinate space).
pub crop_pdf_pt_bbox: [f32; 4],
/// DPI the crop image was rendered at (e.g. `200.0`). Used to convert
/// cell bboxes from image-pixels back to PDF points.
pub render_dpi: f32,
/// Raw structure tokens emitted by the TSR model, in document order.
/// See [`tables::structured::parse_structure`] for the accepted grammar.
pub structure_tokens: Vec<String>,
/// One bbox per cell (in document order, parallel to the cell open-tags
/// in `structure_tokens`). May be 4-element `[x1,y1,x2,y2]` or
/// 8-element 4-corner polygon, in **crop image-pixel space**.
pub cell_bboxes: Vec<Vec<f32>>,
}
/// Extract structured cells using externally-supplied structure recovery.
///
/// For each input, this:
/// 1. Pairs each cell open-tag in `structure_tokens` with the next bbox in
/// `cell_bboxes` (document order), tracking row/col with rowspan/colspan
/// awareness.
/// 2. Converts each cell bbox from crop image-pixels into page PDF-points.
/// 3. Pulls the cell's text by overlap-testing PDF text items inside that
/// bbox — same primitives used by [`extract_text_in_regions_mem`].
///
/// Returns one `Vec<StructuredCell>` per input, in input order. Each cell
/// carries its (row, col, rowspan, colspan, is_header) metadata, the
/// extracted text, and its page-PDF-pt bbox so callers can do their own
/// rendering, debug overlays, or per-cell post-processing.
///
/// Inputs whose page is out of range or whose tokens parse to zero cells
/// produce an empty `Vec`.
///
/// See [`extract_tables_with_structure_mem`] if you just want the rendered
/// markdown.
pub fn extract_tables_with_structure_cells_mem(
buffer: &[u8],
inputs: &[TsrTableInput],
) -> Result<Vec<Vec<tables::StructuredCell>>, PdfError> {
use tables::structured::{
cell_px_to_page_pt, normalize_cell_bands, parse_structure, polygon_to_aabb, StructuredCell,
};
validate_pdf_bytes(buffer)?;
let (doc, _page_count) = load_document_from_mem(buffer)?;
let pages = doc.get_pages();
let needed_pages: HashSet<u32> = inputs.iter().map(|t| t.page + 1).collect();
let font_cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed_pages));
let mut items_by_page: HashMap<u32, Vec<TextItem>> = HashMap::new();
let mut page_heights: HashMap<u32, f32> = HashMap::new();
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
let mut rotated_pages: HashSet<u32> = HashSet::new();
for (page_num, &page_id) in pages.iter() {
if !needed_pages.contains(page_num) {
continue;
}
let height = get_page_height(&doc, page_id).unwrap_or(792.0);
page_heights.insert(*page_num, height);
let ((mut items, _rects, _lines), _has_gid, coords_rotated) =
extractor::content_stream::extract_page_text_items(
&doc,
page_id,
*page_num,
&font_cmaps,
false,
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
page_thresholds.insert(*page_num, threshold);
}
if coords_rotated {
rotated_pages.insert(*page_num);
}
items_by_page.insert(*page_num, items);
}
let mut results: Vec<Vec<StructuredCell>> = Vec::with_capacity(inputs.len());
for input in inputs {
let page_1idx = input.page + 1;
let Some(items) = items_by_page.get(&page_1idx) else {
// Out-of-range page or page with no extractable text — emit empty.
results.push(Vec::new());
continue;
};
let page_h = page_heights.get(&page_1idx).copied().unwrap_or(792.0);
let adaptive_threshold = page_thresholds.get(&page_1idx).copied().unwrap_or(0.10);
let coords = if rotated_pages.contains(&page_1idx) {
RegionCoordSpace::Rotated90Ccw
} else {
RegionCoordSpace::Standard
};
let crop_origin = [input.crop_pdf_pt_bbox[0], input.crop_pdf_pt_bbox[1]];
let slots = parse_structure(&input.structure_tokens);
if slots.is_empty() {
results.push(Vec::new());
continue;
}
let mut cells: Vec<StructuredCell> = Vec::with_capacity(slots.len());
for slot in &slots {
let page_pt_bbox;
if let Some(coords_arr) = input.cell_bboxes.get(slot.bbox_idx) {
if let Some(aabb_px) = polygon_to_aabb(coords_arr) {
page_pt_bbox = cell_px_to_page_pt(aabb_px, input.render_dpi, crop_origin);
} else {
page_pt_bbox = [0.0, 0.0, 0.0, 0.0];
}
} else {
page_pt_bbox = [0.0, 0.0, 0.0, 0.0];
}
cells.push(StructuredCell {
row: slot.row,
col: slot.col,
rowspan: slot.rowspan,
colspan: slot.colspan,
is_header: slot.is_header,
text: String::new(),
page_pt_bbox,
});
}
normalize_cell_bands(&mut cells);
for cell in &mut cells {
let [x1, y1, x2, y2] = cell.page_pt_bbox;
let raw =
collect_text_in_tsr_cell(items, x1, y1, x2, y2, page_h, coords, adaptive_threshold);
// Markdown cells must be one line — collapse line breaks produced
// by the line-grouping pass.
cell.text = raw.replace(['\n', '\r'], " ");
}
results.push(cells);
}
Ok(results)
}
/// Extract markdown tables using externally-supplied structure recovery.
///
/// Convenience wrapper around [`extract_tables_with_structure_cells_mem`]
/// that renders each cell list to markdown via
/// [`tables::cells_to_markdown`]. Returns one markdown string per input,
/// in input order. Inputs whose page is out of range or whose tokens parse
/// to zero cells produce an empty string.
pub fn extract_tables_with_structure_mem(
buffer: &[u8],
inputs: &[TsrTableInput],
) -> Result<Vec<String>, PdfError> {
let cells_lists = extract_tables_with_structure_cells_mem(buffer, inputs)?;
Ok(cells_lists
.into_iter()
.map(|cells| {
if cells.is_empty() {
String::new()
} else {
tables::cells_to_markdown(&cells)
}
})
.collect())
}
/// Get page height in points from MediaBox.
fn get_page_height(doc: &Document, page_id: lopdf::ObjectId) -> Option<f32> {
let page_dict = doc.get_dictionary(page_id).ok()?;
@@ -842,6 +1055,30 @@ fn collect_text_in_region_with_options(
.filter(|item| region_overlaps_item(item, bounds))
.cloned()
.collect();
collect_text_from_matched_items(matched, adaptive_threshold)
}
#[allow(clippy::too_many_arguments)]
fn collect_text_in_tsr_cell(
items: &[TextItem],
rx1: f32,
ry1: f32,
rx2: f32,
ry2: f32,
page_height: f32,
coord_space: RegionCoordSpace,
adaptive_threshold: f32,
) -> String {
let bounds = region_bounds(rx1, ry1, rx2, ry2, page_height, coord_space);
let matched: Vec<TextItem> = items
.iter()
.filter(|item| tsr_region_contains_item(item, bounds))
.cloned()
.collect();
collect_text_from_matched_items(matched, adaptive_threshold)
}
fn collect_text_from_matched_items(matched: Vec<TextItem>, adaptive_threshold: f32) -> String {
if matched.is_empty() {
return String::new();
}
@@ -943,6 +1180,30 @@ fn region_overlaps_item(item: &TextItem, bounds: RegionBounds) -> bool {
x_overlap > 0.0 && y_overlap > 0.0
}
fn tsr_region_contains_item(item: &TextItem, bounds: RegionBounds) -> bool {
let item_x_min = item.x;
let item_x_max = item.x + text_utils::effective_width(item);
let item_y_min = item.y;
let item_y_max = item.y + item.height;
let center_x = (item_x_min + item_x_max) * 0.5;
let center_y = (item_y_min + item_y_max) * 0.5;
if center_x >= bounds.x_min
&& center_x <= bounds.x_max
&& center_y >= bounds.y_min
&& center_y <= bounds.y_max
{
return true;
}
let x_overlap = (item_x_max.min(bounds.x_max) - item_x_min.max(bounds.x_min)).max(0.0);
let y_overlap = (item_y_max.min(bounds.y_max) - item_y_min.max(bounds.y_min)).max(0.0);
let item_width = (item_x_max - item_x_min).max(0.1);
let item_height = (item_y_max - item_y_min).max(0.1);
x_overlap / item_width >= 0.6 && y_overlap / item_height >= 0.6
}
// =========================================================================
// Internal: single-load document pipeline
// =========================================================================
@@ -1782,19 +2043,26 @@ fn compute_layout_complexity(
markdown::filter_lines_to_band(lines, page, x_lo, x_hi)
};
// TOC pages route through the table detector but render as flat
// lists. They aren't tables in any user-facing sense, so don't
// count them toward LayoutComplexity (would also trip the
// table-page guard in column detection below).
let has_data_table =
|tables: &[tables::Table]| tables.iter().any(|t| t.kind == tables::TableKind::Data);
let (rect_tables, _) = tables::detect_tables_from_rects(&band_items, &band_rects, page);
if !rect_tables.is_empty() {
if has_data_table(&rect_tables) {
found_table = true;
break;
}
let line_tables = tables::detect_tables_from_lines(&band_items, &band_lines, page);
if !line_tables.is_empty() {
if has_data_table(&line_tables) {
found_table = true;
break;
}
// Heuristic fallback for borderless tables
let heuristic_tables = tables::detect_tables(&band_items, base_size, false);
if !heuristic_tables.is_empty() {
if has_data_table(&heuristic_tables) {
found_table = true;
break;
}
@@ -1990,6 +2258,24 @@ pub(crate) fn validate_pdf_file<P: AsRef<Path>>(path: P) -> Result<(), PdfError>
#[cfg(test)]
mod tests {
use super::*;
use crate::types::ItemType;
fn test_item(text: &str, x: f32, y: f32, width: f32, height: f32) -> TextItem {
TextItem {
text: text.to_string(),
x,
y,
width,
height,
font: "Helvetica".to_string(),
font_size: height,
page: 1,
is_bold: false,
is_italic: false,
item_type: ItemType::Text,
mcid: None,
}
}
#[test]
fn test_detect_encoding_issues_fffd() {
@@ -2070,4 +2356,97 @@ mod tests {
"Valid Japanese text should not be flagged as garbage"
);
}
#[test]
fn tsr_text_fill_does_not_pull_neighboring_overlapping_rows() {
use crate::tables::structured::normalize_cell_bands;
use crate::tables::StructuredCell;
let items = vec![
test_item("Branch Name", 12.0, 88.0, 55.0, 8.0),
test_item("Deposits", 112.0, 88.0, 36.0, 8.0),
test_item("Oak Street", 12.0, 72.0, 48.0, 8.0),
test_item("100", 112.0, 72.0, 18.0, 8.0),
test_item("Boardwalk", 12.0, 55.2, 46.0, 8.0),
test_item("200", 112.0, 55.2, 18.0, 8.0),
];
let mut cells = vec![
StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 1,
is_header: true,
text: String::new(),
page_pt_bbox: [10.0, 100.0, 100.0, 125.0],
},
StructuredCell {
row: 0,
col: 1,
rowspan: 1,
colspan: 1,
is_header: true,
text: String::new(),
page_pt_bbox: [100.0, 100.0, 170.0, 125.0],
},
StructuredCell {
row: 1,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [10.0, 116.0, 100.0, 141.0],
},
StructuredCell {
row: 1,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [100.0, 116.0, 170.0, 141.0],
},
StructuredCell {
row: 2,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [10.0, 132.8, 100.0, 157.8],
},
StructuredCell {
row: 2,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [100.0, 132.8, 170.0, 157.8],
},
];
normalize_cell_bands(&mut cells);
for cell in &mut cells {
let [x1, y1, x2, y2] = cell.page_pt_bbox;
cell.text = collect_text_in_tsr_cell(
&items,
x1,
y1,
x2,
y2,
200.0,
RegionCoordSpace::Standard,
0.10,
);
}
assert_eq!(cells[0].text, "Branch Name");
assert_eq!(cells[2].text, "Oak Street");
assert_eq!(cells[4].text, "Boardwalk");
assert!(!cells[0].text.contains("Oak Street"));
assert!(!cells[2].text.contains("Branch Name"));
assert!(!cells[2].text.contains("Boardwalk"));
}
}
+65
View File
@@ -64,6 +64,22 @@ pub(crate) fn is_caption_line(text: &str) -> bool {
false
}
/// Check if text starts with an unambiguous bullet marker (●, •, ○, ◦).
///
/// Narrower than [`is_list_item`]: it excludes numbered/lettered patterns
/// like `1.` or `a)`, which legitimately appear as section headings in many
/// documents. Used by the heading classifier to reject bullet lines without
/// also demoting numbered headings.
pub(crate) fn starts_with_bullet_marker(text: &str) -> bool {
let trimmed = text.trim_start();
trimmed.starts_with("")
|| trimmed.starts_with("")
|| trimmed.starts_with("")
|| trimmed.starts_with("")
|| trimmed.starts_with("- ")
|| trimmed.starts_with("* ")
}
/// Check if text looks like a list item
pub(crate) fn is_list_item(text: &str) -> bool {
let trimmed = text.trim_start();
@@ -115,6 +131,16 @@ pub(crate) fn format_list_item(text: &str) -> String {
if let Some(rest) = trimmed.strip_prefix(*bullet) {
return format!("- {}", rest.trim_start());
}
// Bullet inside a leading bold/italic run (e.g. "**● Label:** rest").
// The run wraps both the marker and the following label because both
// use a bold font in the PDF.
for wrapper in ["**", "*"] {
if let Some(after_open) = trimmed.strip_prefix(wrapper) {
if let Some(rest) = after_open.strip_prefix(*bullet) {
return format!("- {}{}", wrapper, rest.trim_start());
}
}
}
}
if trimmed.starts_with("- ") || trimmed.starts_with("* ") {
@@ -198,3 +224,42 @@ pub(crate) fn is_monospace_font(font_name: &str) -> bool {
patterns.iter().any(|p| lower.contains(p))
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn format_list_item_plain_bullet() {
assert_eq!(format_list_item("● Item"), "- Item");
assert_eq!(format_list_item("• Item"), "- Item");
}
#[test]
fn format_list_item_bullet_inside_bold() {
// PDF that uses bold font for both the marker and the label produces
// a single bold run like "**● Label:** rest"; the bullet must still
// be stripped and the bold wrapper preserved on the label.
assert_eq!(
format_list_item("**● Fraud: Willing cooperation;**"),
"- **Fraud: Willing cooperation;**"
);
assert_eq!(
format_list_item("**● Label:** rest of line"),
"- **Label:** rest of line"
);
assert_eq!(format_list_item("*● Italic:* rest"), "- *Italic:* rest");
}
#[test]
fn format_list_item_already_dash() {
assert_eq!(format_list_item("- existing"), "- existing");
}
#[test]
fn is_list_item_with_bullet_space() {
assert!(is_list_item("● Item"));
assert!(is_list_item("• Item"));
assert!(is_list_item("- Item"));
}
}
+149 -2
View File
@@ -9,7 +9,9 @@ use super::analysis::{
bold_heading_level, calculate_font_stats, compute_heading_tiers, compute_paragraph_threshold,
detect_header_level, font_size_rarity, has_dot_leaders,
};
use super::classify::{format_list_item, is_caption_line, is_list_item, is_monospace_font};
use super::classify::{
format_list_item, is_caption_line, is_list_item, is_monospace_font, starts_with_bullet_marker,
};
use super::postprocess::clean_markdown;
use super::preprocess::{merge_drop_caps, merge_heading_lines};
use super::MarkdownOptions;
@@ -584,9 +586,30 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
.as_ref()
.and_then(struct_role_heading_level)
.filter(|level| !overused_heading_levels.contains(level));
// Protect wrapped list items: when inside a list, a visually-continuing
// line (same indent, line-wrap spacing) must not be reclassified as a
// heading by the font heuristic — PDFs often bold the lead phrase of a
// list item across multiple wrap lines, and an all-bold middle line
// would otherwise split one item into a heading + stray body text.
// We gate on the document's paragraph threshold so genuine section
// headings that follow a numbered paragraph (y_gap > para_threshold)
// remain detectable.
let looks_like_list_continuation = in_list
&& match (last_list_x, line.items.first().map(|i| i.x)) {
(Some(list_x), Some(curr_x)) => {
let x_ok = curr_x >= list_x - 5.0 && curr_x <= list_x + 50.0;
let y_ok = y_gap >= 0.0 && y_gap <= para_threshold;
x_ok && y_ok && !is_list_item(plain_trimmed)
}
_ => false,
};
let heuristic_heading = if options.detect_headers
&& !looks_like_list_continuation
&& plain_trimmed.len() > 3
&& plain_trimmed.split_whitespace().count() <= 15
&& !starts_with_bullet_marker(plain_trimmed)
{
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
detect_header_level(line_font_size, base_size, &heading_tiers).or_else(|| {
@@ -641,11 +664,16 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
continue;
}
// Structure-tree list item (LI only — LBody is a continuation, not a new item)
// Structure-tree list item (LI only — LBody is a continuation, not a new item).
// Some tagged PDFs use a "flat" style where every wrapped line in a list item
// gets its own MCID tagged directly under LI. When we're already inside a list
// and the line has no visible bullet marker, treat it as a continuation (falls
// through to the continuation logic below) rather than a new list item.
if struct_role
.as_ref()
.is_some_and(|r| matches!(r, StructRole::LI))
&& !is_list_item(plain_trimmed)
&& !in_list
{
if in_paragraph {
output.push_str("\n\n");
@@ -1090,6 +1118,59 @@ mod tests {
);
}
#[test]
fn test_struct_role_li_flat_continuation_lines_merge() {
// Regression: some tagged PDFs put each wrapped visual line of a list
// item under its own MCID, all tagged directly as LI. Continuation
// lines (no bullet marker) must merge into the bulleted parent item,
// not each become their own list item.
let make = |text: &str, mcid: i64, x: f32, y: f32| {
let mut item = make_item(text, 1, Some(mcid));
item.x = x;
item.y = y;
item
};
let lines = vec![
make_line(vec![make("● First item that wraps onto", 0, 90.0, 322.0)]),
make_line(vec![make("a continuation line.", 1, 108.0, 306.0)]),
make_line(vec![make("● Second bullet also wraps", 2, 90.0, 290.0)]),
make_line(vec![make("to a second line here.", 3, 108.0, 274.0)]),
];
let mut page_roles = HashMap::new();
for mcid in 0..4 {
page_roles.insert(mcid, StructRole::LI);
}
let mut roles = HashMap::new();
roles.insert(1u32, page_roles);
let md = to_markdown_from_lines_with_tables_and_images(
lines,
MarkdownOptions::default(),
HashMap::new(),
HashMap::new(),
&std::collections::HashSet::new(),
Some(&roles),
);
assert!(
md.contains("- First item that wraps onto a continuation line."),
"continuation should merge into first bullet: {md}"
);
assert!(
md.contains("- Second bullet also wraps to a second line here."),
"continuation should merge into second bullet: {md}"
);
assert!(
!md.contains("- a continuation line."),
"continuation line should not get its own bullet: {md}"
);
assert!(
!md.contains("- to a second line here."),
"continuation line should not get its own bullet: {md}"
);
}
#[test]
fn test_struct_role_blockquote() {
let lines = vec![make_line(vec![make_item("Quoted text", 1, Some(0))])];
@@ -1392,4 +1473,70 @@ mod tests {
overused
);
}
#[test]
fn test_wrapped_bold_lead_in_list_item_not_heading() {
// Regression: numbered-list items whose bold "lead" phrase wraps onto
// a second line (e.g. definitions in system cards) must not have the
// wrapped line reclassified as a heading. The middle line is
// all_bold + standalone (in_paragraph=false while in_list), which
// previously tripped the rarity heuristic and emitted #### in the
// middle of the item, splitting the body into stray bullets.
let make = |text: &str, x: f32, y: f32, bold: bool| {
let mut item = make_item(text, 1, None);
item.x = x;
item.y = y;
item.is_bold = bold;
item
};
let lines = vec![
// "1. **bold lead phrase start**"
make_line(vec![
make("1. ", 72.0, 700.0, false),
make(
"Chemical and biological weapons threat model 1 (CB-1): Non-novel",
90.0,
700.0,
true,
),
]),
// wrapped continuation of the bold lead — all_bold, same indent
make_line(vec![make(
"chemical/biological weapons production capabilities: A model has CB-1",
90.0,
686.0,
true,
)]),
// body text of the same list item
make_line(vec![make(
"capabilities if it has the ability to significantly help individuals.",
90.0,
672.0,
false,
)]),
];
let md = to_markdown_from_lines_with_tables_and_images(
lines,
MarkdownOptions::default(),
HashMap::new(),
HashMap::new(),
&std::collections::HashSet::new(),
None,
);
assert!(
!md.contains("#### "),
"wrapped bold lead must not become a heading: {md}"
);
assert!(
md.lines().filter(|l| l.starts_with("- ")).count() == 0,
"continuation body must not become a stray bullet: {md}"
);
assert!(
md.contains("1. ") && md.contains("A model has CB-1"),
"numbered list item should remain intact: {md}"
);
}
}
+121
View File
@@ -146,6 +146,67 @@ impl PyPageRegionTexts {
// Text item wrapper
// ---------------------------------------------------------------------------
/// Per-page markdown extraction result.
#[pyclass(name = "PageMarkdown")]
#[derive(Clone)]
pub struct PyPageMarkdown {
/// 0-indexed page number.
#[pyo3(get)]
pub page: u32,
/// Formatted markdown for this page.
#[pyo3(get)]
pub markdown: String,
/// True when text on this page is unreliable (GID-encoded fonts,
/// encoding issues, garbage text, or empty extraction).
#[pyo3(get)]
pub needs_ocr: bool,
}
#[pymethods]
impl PyPageMarkdown {
fn __repr__(&self) -> String {
format!(
"PageMarkdown(page={}, markdown='{}', needs_ocr={})",
self.page,
self.markdown.chars().take(40).collect::<String>(),
self.needs_ocr
)
}
}
/// Combined per-page markdown extraction and layout classification result.
#[pyclass(name = "PagesExtractionResult")]
#[derive(Clone)]
pub struct PyPagesExtractionResult {
/// Per-page markdown results, in the order requested.
#[pyo3(get)]
pub pages: Vec<PyPageMarkdown>,
/// 1-indexed pages where tables were detected.
#[pyo3(get)]
pub pages_with_tables: Vec<u32>,
/// 1-indexed pages where multi-column layout was detected.
#[pyo3(get)]
pub pages_with_columns: Vec<u32>,
/// 1-indexed pages that need OCR (scanned/image-based or unreliable text).
#[pyo3(get)]
pub pages_needing_ocr: Vec<u32>,
/// True if any page has tables or columns.
#[pyo3(get)]
pub is_complex: bool,
}
#[pymethods]
impl PyPagesExtractionResult {
fn __repr__(&self) -> String {
format!(
"PagesExtractionResult(pages={}, pages_with_tables={:?}, is_complex={})",
self.pages.len(),
self.pages_with_tables,
self.is_complex
)
}
}
/// A positioned text item extracted from a PDF.
#[pyclass(name = "TextItem")]
#[derive(Clone)]
@@ -280,6 +341,24 @@ fn parse_page_regions(
.collect()
}
fn to_py_pages_result(r: crate::PagesExtractionResult) -> PyPagesExtractionResult {
PyPagesExtractionResult {
pages: r
.pages
.into_iter()
.map(|p| PyPageMarkdown {
page: p.page,
markdown: p.markdown,
needs_ocr: p.needs_ocr,
})
.collect(),
pages_with_tables: r.pages_with_tables,
pages_with_columns: r.pages_with_columns,
pages_needing_ocr: r.pages_needing_ocr,
is_complex: r.is_complex,
}
}
fn convert_region_results(results: Vec<crate::PageRegionResult>) -> Vec<PyPageRegionTexts> {
results
.into_iter()
@@ -442,6 +521,44 @@ fn extract_text_in_regions_bytes(
Ok(convert_region_results(results))
}
/// Extract formatted markdown for pages of a PDF file, with layout
/// classification metadata.
///
/// Returns per-page markdown and classification data (tables, columns,
/// OCR needs) from a single parse. Font statistics are computed from the
/// full document so header detection is consistent across pages.
///
/// Args:
/// path: Path to the PDF file.
/// pages: Optional list of 0-indexed pages. When None (default), every
/// page is returned in document order. When provided, output
/// matches the caller-supplied order.
///
/// Returns:
/// PagesExtractionResult with per-page markdown and classification data.
#[pyfunction]
#[pyo3(signature = (path, pages=None))]
fn extract_pages_markdown(
path: &str,
pages: Option<Vec<u32>>,
) -> PyResult<PyPagesExtractionResult> {
let result = crate::extract_pages_markdown(path, pages.as_deref()).map_err(to_py_err)?;
Ok(to_py_pages_result(result))
}
/// Extract formatted markdown for pages of a PDF from bytes.
///
/// See [`extract_pages_markdown`] for details.
#[pyfunction]
#[pyo3(signature = (data, pages=None))]
fn extract_pages_markdown_bytes(
data: &[u8],
pages: Option<Vec<u32>>,
) -> PyResult<PyPagesExtractionResult> {
let result = crate::extract_pages_markdown_mem(data, pages.as_deref()).map_err(to_py_err)?;
Ok(to_py_pages_result(result))
}
/// Python module definition.
#[pymodule]
fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
@@ -450,6 +567,8 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_class::<PyTextItem>()?;
m.add_class::<PyRegionText>()?;
m.add_class::<PyPageRegionTexts>()?;
m.add_class::<PyPageMarkdown>()?;
m.add_class::<PyPagesExtractionResult>()?;
m.add_function(wrap_pyfunction!(process_pdf, m)?)?;
m.add_function(wrap_pyfunction!(process_pdf_bytes, m)?)?;
m.add_function(wrap_pyfunction!(detect_pdf, m)?)?;
@@ -462,5 +581,7 @@ fn pdf_inspector(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_function(wrap_pyfunction!(extract_text_with_positions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions, m)?)?;
m.add_function(wrap_pyfunction!(extract_text_in_regions_bytes, m)?)?;
m.add_function(wrap_pyfunction!(extract_pages_markdown, m)?)?;
m.add_function(wrap_pyfunction!(extract_pages_markdown_bytes, m)?)?;
Ok(())
}
+71 -10
View File
@@ -581,8 +581,14 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
// Validation 1: some rows should have content in first column.
// Use a lower threshold (25%) for tables with wrapped cells where
// continuation lines leave the first column empty.
// Skip when cells form a narrow TOC pattern: hierarchical entries indented
// across multiple X levels leave the leftmost column sparse (only top-level
// chapters land there) but the structure is still a valid TOC. Narrow only
// (<=5 cols) — wide multi-column TOCs (e.g. 2-up indices) would render
// poorly through format_toc_as_list, which assumes one entry per row.
let rows_with_first_col = cells.iter().filter(|row| !row[0].is_empty()).count();
if rows_with_first_col < rows.len() / 4 {
let is_narrow_toc = columns.len() <= 5 && is_table_of_contents(&cells);
if rows_with_first_col < rows.len() / 4 && !is_narrow_toc {
log::debug!(
" validation 1 fail: {}/{} rows have first col",
rows_with_first_col,
@@ -653,8 +659,12 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
return None;
}
// Validation 8: Reject paragraph-like content falsely detected as tables
if is_paragraph_content(&cells) {
// Validation 8: Reject paragraph-like content falsely detected as tables.
// TOC pages with deep indentation (top-level chapters in col 0, subsections
// in cols 1-3, page numbers in last col) leave most cells empty and trip
// the paragraph heuristic; TOC shape is a safer signal here. Narrow only
// — see narrow-TOC rationale at validation 1.
if is_paragraph_content(&cells) && !is_narrow_toc {
log::debug!(" validation 9 fail: paragraph content");
return None;
}
@@ -677,12 +687,7 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
item_indices.len()
);
Some(Table {
columns,
rows,
cells,
item_indices,
})
Some(Table::new(columns, rows, cells, item_indices))
}
/// Check if this looks like a key-value pair layout rather than a table
@@ -906,7 +911,7 @@ fn looks_like_number(s: &str) -> bool {
/// Check if this looks like a Table of Contents (either style).
///
/// Used by format.rs to render TOCs as flat lists instead of markdown tables.
pub(super) fn is_table_of_contents(cells: &[Vec<String>]) -> bool {
pub fn is_table_of_contents(cells: &[Vec<String>]) -> bool {
is_dot_leader_toc(cells) || is_tabular_toc(cells)
}
@@ -1601,6 +1606,62 @@ mod tests {
);
}
#[test]
fn is_table_of_contents_accepts_hierarchical_indented_toc() {
// Mythos system card pages 4-5: top-level chapters indent at col 0,
// subsections at cols 1-2, leaving col 0 mostly empty (only ~10% of
// rows). Validation 1 was rejecting these even though the structure
// is unambiguously a TOC.
let cells = vec![
vec!["Abstract".to_string(), String::new(), "3".to_string()],
vec![
"1 Introduction".to_string(),
String::new(),
"10".to_string(),
],
vec![
String::new(),
"1.1 Model training".to_string(),
"11".to_string(),
],
vec![
String::new(),
"1.1.1 Training data".to_string(),
"11".to_string(),
],
vec![
String::new(),
"1.1.2 Crowd workers".to_string(),
"12".to_string(),
],
vec![
String::new(),
"1.2 Release decision".to_string(),
"13".to_string(),
],
vec![
"2 RSP evaluations".to_string(),
String::new(),
"16".to_string(),
],
vec![
String::new(),
"2.1 RSP risk assessment".to_string(),
"16".to_string(),
],
vec![String::new(), "2.1.1 Context".to_string(), "16".to_string()],
vec![
String::new(),
"2.2 CB evaluations".to_string(),
"20".to_string(),
],
];
assert!(
is_table_of_contents(&cells),
"hierarchical TOC with sparse col 0 should still be detected"
);
}
#[test]
fn is_table_of_contents_rejects_dotless_toc() {
// Tabular TOC without leader dots: first column starts with dotted
+4 -4
View File
@@ -265,12 +265,12 @@ pub fn detect_tables_from_lines(items: &[TextItem], lines: &[PdfLine], page: u32
page, num_rows, num_cols, item_indices.len(), page_item_count, non_empty_rows, cols_with_content
);
vec![Table {
columns: col_edges,
rows: row_edges_desc[..num_rows].to_vec(),
vec![Table::new(
col_edges,
row_edges_desc[..num_rows].to_vec(),
cells,
item_indices,
}]
)]
}
#[cfg(test)]
+85 -29
View File
@@ -1037,12 +1037,7 @@ fn try_build_grid(
(columns, cells)
};
GridResult::Ok(Table {
columns,
rows,
cells,
item_indices,
})
GridResult::Ok(Table::new(columns, rows, cells, item_indices))
}
/// Deduplicate nearby edge values within a tolerance, returning sorted unique edges.
@@ -1161,11 +1156,23 @@ fn propagate_merged_cells(
continue;
}
// Find first and last grid rows that the rect spans
let first_row = (0..num_rows)
.find(|&r| ry <= row_edges[r] + tol && (ry + rh) >= row_edges[r + 1] - tol);
let last_row = (0..num_rows)
.rfind(|&r| ry <= row_edges[r] + tol && (ry + rh) >= row_edges[r + 1] - tol);
// Find first and last grid rows that the rect spans.
//
// Require a rect to actually overlap the row by more than `tol`
// to count as a span. A "rect bottom ≤ row top + tol AND rect
// top ≥ row bottom tol" check gives false positives at shared
// row boundaries — a rect whose top equals row N's bottom lies
// entirely below the row but still passes the tolerance-slack
// check, cascading body text from unrelated rows into one
// merged cell.
let spans = |r: usize| {
let row_top = row_edges[r];
let row_bot = row_edges[r + 1];
let overlap = (row_top.min(ry + rh) - row_bot.max(ry)).max(0.0);
overlap > tol
};
let first_row = (0..num_rows).find(|&r| spans(r));
let last_row = (0..num_rows).rfind(|&r| spans(r));
let (first, last) = match (first_row, last_row) {
(Some(f), Some(l)) if l > f => (f, l),
@@ -1442,12 +1449,7 @@ fn detect_row_stripe_table(
content_ratio * 100.0
);
Some(Table {
columns: column_centers,
rows: row_centers,
cells,
item_indices,
})
Some(Table::new(column_centers, row_centers, cells, item_indices))
}
/// Detect a table from cell-background rects that failed grid detection.
@@ -1677,6 +1679,49 @@ fn detect_row_stripe_table_from_cell_rects(
return None;
}
// Reject "tables" that are actually prose in a framed region.
// Columns here come from text X-position clustering; when prose wraps
// inside a bounding-box rect (e.g. chat-transcript figures) the
// word-boundary gaps cluster into many spurious columns, and the
// resulting cells hold sentence fragments riddled with common English
// function words. Count cells with any such word and reject when
// 20%+ of non-empty cells match — real tabular data (labels, units,
// numbers) rarely contains these words.
if num_cols >= 4 {
const PROSE_WORDS: &[&str] = &[
"a", "an", "the", "of", "to", "is", "was", "are", "were", "be", "been", "in", "on",
"at", "with", "for", "by", "as", "and", "or", "but", "this", "that", "these", "those",
"from", "into", "has", "have", "had", "not", "don't", "doesn't", "it's", "its", "it",
"i", "me", "my", "we", "our", "us", "you", "your", "they", "them", "their", "he",
"she", "his", "her",
];
let mut prose_cells = 0usize;
let mut counted = 0usize;
for row in &cells {
for cell in row {
let t = cell.trim();
if t.is_empty() {
continue;
}
counted += 1;
let lower = t.to_ascii_lowercase();
let has_prose_word = lower
.split(|c: char| !c.is_ascii_alphabetic() && c != '\'')
.any(|w| PROSE_WORDS.contains(&w));
if has_prose_word {
prose_cells += 1;
}
}
}
if counted > 0 && prose_cells * 5 >= counted {
debug!(
" cell-rect rejected: {}/{} cells contain prose function words — likely prose",
prose_cells, counted
);
return None;
}
}
let column_centers: Vec<f32> = (0..num_cols)
.map(|c| (col_edges[c] + col_edges[c + 1]) / 2.0)
.collect();
@@ -1691,12 +1736,7 @@ fn detect_row_stripe_table_from_cell_rects(
non_empty_cells as f32 / total_cells * 100.0
);
Some(Table {
columns: column_centers,
rows: row_centers,
cells,
item_indices,
})
Some(Table::new(column_centers, row_centers, cells, item_indices))
}
/// Detect a table by merging all cluster rects into one group.
@@ -1875,12 +1915,7 @@ fn detect_merged_cluster_table(
content_ratio * 100.0
);
Some(Table {
columns: column_centers,
rows: row_centers,
cells,
item_indices,
})
Some(Table::new(column_centers, row_centers, cells, item_indices))
}
/// Cluster text item X positions into column centers with a given minimum threshold.
@@ -2370,6 +2405,27 @@ mod tests {
assert_eq!(cells[1][0], "B");
}
#[test]
fn test_propagate_merged_cells_rect_tangent_to_row_boundary() {
// Regression: a rect whose top exactly equals a row's bottom lies
// entirely outside that row, so it must not be considered to span
// it. With the old overlap-based predicate this cascaded into body
// text from unrelated rows being merged into a single header cell
// (mythos system card CB task-based evaluations table).
//
// Layout: two rows 0..80 and 80..160 (bottom → top in PDF coords),
// rect occupies only the lower row (y=0..80). Its top equals the
// upper row's bottom; it must not span the upper row.
let col_edges = vec![0.0, 50.0];
let row_edges = vec![160.0, 80.0, 0.0]; // top → bot
let mut cells = vec![vec!["Upper".to_string()], vec!["Lower".to_string()]];
let group_rects = vec![(0.0, 0.0, 50.0, 80.0)]; // rect at y=0..80
let skip = vec![false];
propagate_merged_cells(&mut cells, &col_edges, &row_edges, &group_rects, &skip);
assert_eq!(cells[0][0], "Upper", "upper row must not be merged");
assert_eq!(cells[1][0], "Lower", "lower row must not be touched");
}
#[test]
fn test_propagate_merged_cells_empty_cells_preserved() {
let col_edges = vec![0.0, 50.0];
+846 -62
View File
@@ -4,15 +4,385 @@
//! elements linked to MCIDs, this module builds `Table` structs directly from
//! the semantic hierarchy — no geometry heuristics needed.
use std::collections::HashMap;
use std::collections::{HashMap, HashSet};
use log::debug;
use crate::structure_tree::StructTable;
use crate::structure_tree::{StructTable, StructTableRow};
use crate::types::TextItem;
use super::Table;
#[derive(Debug, Clone)]
struct MatchedCell {
text: String,
item_indices: Vec<usize>,
x: Option<f32>,
y: Option<f32>,
}
fn legacy_column_positions(
page_rows: &[&StructTableRow],
mcid_to_items: &HashMap<i64, Vec<usize>>,
items: &[TextItem],
page: u32,
num_cols: usize,
) -> Vec<f32> {
let mut col_positions: Vec<f32> = vec![0.0; num_cols];
for (col, col_pos) in col_positions.iter_mut().enumerate() {
for row in page_rows {
if col < row.cells.len() {
if let Some(x) = row.cells[col]
.mcids
.iter()
.filter(|(_, p)| *p == page)
.filter_map(|(mcid, _)| mcid_to_items.get(mcid))
.flatten()
.map(|&idx| items[idx].x)
.reduce(f32::min)
{
*col_pos = x;
break;
}
}
}
}
col_positions
}
fn infer_column_positions(
raw_rows: &[Vec<MatchedCell>],
fallback_positions: &[f32],
num_cols: usize,
) -> Vec<f32> {
const SAME_COLUMN_TOLERANCE: f32 = 18.0;
let mut anchors = raw_rows
.iter()
.max_by_key(|row| row.iter().filter(|cell| cell.x.is_some()).count())
.map(|row| row.iter().filter_map(|cell| cell.x).collect::<Vec<_>>())
.unwrap_or_default();
if anchors.len() > num_cols {
anchors.truncate(num_cols);
}
let mut additional_positions: Vec<f32> = raw_rows
.iter()
.flat_map(|row| row.iter().filter_map(|cell| cell.x))
.collect();
additional_positions.sort_by(|a, b| a.total_cmp(b));
for x in additional_positions {
if anchors.len() >= num_cols {
break;
}
if anchors
.iter()
.all(|existing| (x - *existing).abs() > SAME_COLUMN_TOLERANCE)
{
anchors.push(x);
anchors.sort_by(|a, b| a.total_cmp(b));
}
}
if anchors.len() < num_cols {
for &x in fallback_positions {
if anchors.len() >= num_cols {
break;
}
if anchors
.iter()
.all(|existing| (x - *existing).abs() > SAME_COLUMN_TOLERANCE)
{
anchors.push(x);
anchors.sort_by(|a, b| a.total_cmp(b));
}
}
}
if anchors.is_empty() {
return fallback_positions.to_vec();
}
while anchors.len() < num_cols {
anchors.push(*anchors.last().unwrap());
}
anchors
}
fn align_positions_to_columns(cell_xs: &[f32], columns: &[f32]) -> Vec<usize> {
if cell_xs.is_empty() || columns.is_empty() {
return Vec::new();
}
if cell_xs.len() >= columns.len() {
return (0..cell_xs.len().min(columns.len())).collect();
}
let mut dp = vec![vec![f32::INFINITY; columns.len() + 1]; cell_xs.len() + 1];
let mut take = vec![vec![false; columns.len() + 1]; cell_xs.len() + 1];
for value in &mut dp[0] {
*value = 0.0;
}
for i in 1..=cell_xs.len() {
for j in 1..=columns.len() {
let skip_cost = dp[i][j - 1];
let take_cost = dp[i - 1][j - 1] + (cell_xs[i - 1] - columns[j - 1]).abs();
if take_cost <= skip_cost {
dp[i][j] = take_cost;
take[i][j] = true;
} else {
dp[i][j] = skip_cost;
}
}
}
let mut assignments_rev = Vec::with_capacity(cell_xs.len());
let mut i = cell_xs.len();
let mut j = columns.len();
while i > 0 && j > 0 {
if take[i][j] {
assignments_rev.push(j - 1);
i -= 1;
j -= 1;
} else {
j -= 1;
}
}
assignments_rev.reverse();
assignments_rev
}
fn align_struct_rows(
raw_rows: &[Vec<MatchedCell>],
col_positions: &[f32],
) -> (Vec<Vec<String>>, Vec<f32>, Vec<usize>) {
let mut cells: Vec<Vec<String>> = Vec::with_capacity(raw_rows.len());
let mut row_positions: Vec<f32> = Vec::with_capacity(raw_rows.len());
let mut all_item_indices: Vec<usize> = Vec::new();
for row in raw_rows {
let present_cells: Vec<&MatchedCell> = row
.iter()
.filter(|cell| {
!cell.item_indices.is_empty() || !cell.text.is_empty() || cell.x.is_some()
})
.collect();
let cell_xs: Vec<f32> = present_cells.iter().filter_map(|cell| cell.x).collect();
let assignments = if cell_xs.len() == present_cells.len() {
align_positions_to_columns(&cell_xs, col_positions)
} else {
(0..present_cells.len().min(col_positions.len())).collect()
};
let mut row_cells = vec![String::new(); col_positions.len()];
for (cell, &col_idx) in present_cells.iter().zip(assignments.iter()) {
if !cell.text.is_empty() {
if !row_cells[col_idx].is_empty() {
row_cells[col_idx].push(' ');
}
row_cells[col_idx].push_str(&cell.text);
}
all_item_indices.extend(cell.item_indices.iter().copied());
}
let row_y = row
.iter()
.filter_map(|cell| cell.y)
.reduce(f32::max)
.unwrap_or(0.0);
cells.push(row_cells);
row_positions.push(row_y);
}
(cells, row_positions, all_item_indices)
}
fn left_align_struct_rows(
raw_rows: &[Vec<MatchedCell>],
num_cols: usize,
) -> (Vec<Vec<String>>, Vec<f32>, Vec<usize>) {
let mut cells: Vec<Vec<String>> = Vec::with_capacity(raw_rows.len());
let mut row_positions: Vec<f32> = Vec::with_capacity(raw_rows.len());
let mut all_item_indices: Vec<usize> = Vec::new();
for row in raw_rows {
let mut row_cells: Vec<String> = row.iter().map(|cell| cell.text.clone()).collect();
row_cells.truncate(num_cols);
while row_cells.len() < num_cols {
row_cells.push(String::new());
}
cells.push(row_cells);
all_item_indices.extend(
row.iter()
.flat_map(|cell| cell.item_indices.iter().copied()),
);
row_positions.push(
row.iter()
.filter_map(|cell| cell.y)
.reduce(f32::max)
.unwrap_or(0.0),
);
}
(cells, row_positions, all_item_indices)
}
fn recover_unclaimed_header_row(table: &mut Table, items: &[TextItem], has_ragged_rows: bool) {
if !has_ragged_rows || table.rows.is_empty() || table.columns.len() < 3 {
return;
}
const MAX_HEADER_DISTANCE: f32 = 90.0;
const MAX_GAP_TO_TABLE: f32 = 35.0;
const MAX_INTER_HEADER_GAP: f32 = 25.0;
const MAX_HEADER_ROWS: usize = 3;
const Y_TOLERANCE: f32 = 5.0;
let top_row_y = table.rows[0];
let x_min = table.columns.first().copied().unwrap_or(0.0) - 25.0;
let x_max = table.columns.last().copied().unwrap_or(0.0) + 120.0;
let claimed: HashSet<usize> = table.item_indices.iter().copied().collect();
let mut candidate_rows: Vec<(f32, Vec<(usize, &TextItem)>)> = Vec::new();
for (idx, item) in items.iter().enumerate() {
if claimed.contains(&idx)
|| item.text.trim().is_empty()
|| item.y <= top_row_y
|| item.y - top_row_y > MAX_HEADER_DISTANCE
|| item.x < x_min
|| item.x > x_max
{
continue;
}
if let Some((_, row_items)) = candidate_rows
.iter_mut()
.find(|(row_y, _)| (item.y - *row_y).abs() < Y_TOLERANCE)
{
row_items.push((idx, item));
} else {
candidate_rows.push((item.y, vec![(idx, item)]));
}
}
if candidate_rows.is_empty() {
return;
}
for (_, row_items) in &mut candidate_rows {
row_items.sort_by(|a, b| a.1.x.total_cmp(&b.1.x));
}
candidate_rows.sort_by(|a, b| a.0.total_cmp(&b.0));
if candidate_rows[0].0 - top_row_y > MAX_GAP_TO_TABLE {
return;
}
let mut candidate_iter = candidate_rows.into_iter();
let Some(first_row) = candidate_iter.next() else {
return;
};
let mut selected_rows: Vec<(f32, Vec<(usize, &TextItem)>)> = vec![first_row];
let mut prev_y = selected_rows[0].0;
for (row_y, row_items) in candidate_iter {
if selected_rows.len() >= MAX_HEADER_ROWS {
break;
}
if row_y - prev_y > MAX_INTER_HEADER_GAP {
break;
}
prev_y = row_y;
selected_rows.push((row_y, row_items));
}
if selected_rows.is_empty() {
return;
}
let mut assigned_rows: Vec<(f32, Vec<String>, Vec<usize>)> = Vec::new();
let mut closest_row_populated = 0usize;
let mut combined_cols: HashSet<usize> = HashSet::new();
for (row_idx, (row_y, row_items)) in selected_rows.iter().enumerate() {
if row_items.len() > table.columns.len() {
return;
}
let row_xs: Vec<f32> = row_items.iter().map(|(_, item)| item.x).collect();
let assignments = align_positions_to_columns(&row_xs, &table.columns);
if assignments.len() != row_items.len() {
return;
}
let mut row_cells = vec![String::new(); table.columns.len()];
let mut row_indices = Vec::with_capacity(row_items.len());
let mut populated_cols: HashSet<usize> = HashSet::new();
for ((idx, item), &col_idx) in row_items.iter().zip(assignments.iter()) {
let text = item.text.trim();
if text.is_empty() {
continue;
}
if !row_cells[col_idx].is_empty() {
row_cells[col_idx].push(' ');
}
row_cells[col_idx].push_str(text);
row_indices.push(*idx);
populated_cols.insert(col_idx);
}
if row_idx == 0 {
closest_row_populated = populated_cols.len();
}
combined_cols.extend(populated_cols.iter().copied());
assigned_rows.push((*row_y, row_cells, row_indices));
}
let required_cols = if table.columns.len() <= 4 {
table.columns.len()
} else {
table.columns.len() - 1
};
if closest_row_populated < 2 || combined_cols.len() < required_cols {
return;
}
let mut header_cells = vec![String::new(); table.columns.len()];
let mut header_indices = Vec::new();
for (_, row_cells, row_indices) in assigned_rows.iter().rev() {
for (col_idx, cell_text) in row_cells.iter().enumerate() {
if cell_text.is_empty() {
continue;
}
if !header_cells[col_idx].is_empty() {
header_cells[col_idx].push(' ');
}
header_cells[col_idx].push_str(cell_text);
}
header_indices.extend(row_indices.iter().copied());
}
table.rows.insert(
0,
assigned_rows
.iter()
.map(|(row_y, _, _)| *row_y)
.reduce(f32::max)
.unwrap_or(top_row_y),
);
table.cells.insert(0, header_cells);
table.item_indices.extend(header_indices);
table.item_indices.sort_unstable();
table.item_indices.dedup();
}
/// Build tables from structure-tree table descriptors by matching MCIDs to TextItems.
///
/// Returns tables for the given page. Tables where fewer than 50% of cells
@@ -67,18 +437,14 @@ pub fn detect_tables_from_struct_tree(
continue;
}
// Build cell text and collect item indices
let mut cells: Vec<Vec<String>> = Vec::new();
let mut all_item_indices: Vec<usize> = Vec::new();
// Build cell text and geometry for alignment and header recovery.
let mut raw_rows: Vec<Vec<MatchedCell>> = Vec::new();
let mut total_cells = 0u32;
let mut matched_cells = 0u32;
for row in &page_rows {
let mut row_cells = Vec::with_capacity(num_cols);
for (col_idx, cell) in row.cells.iter().enumerate() {
if col_idx >= num_cols {
break;
}
let mut row_cells = Vec::with_capacity(row.cells.len());
for cell in &row.cells {
total_cells += 1;
// Collect all items for this cell's MCIDs
@@ -115,18 +481,18 @@ pub fn detect_tables_from_struct_tree(
.collect::<Vec<_>>()
.join(" ");
for (idx, _) in &cell_items {
all_item_indices.push(*idx);
}
let item_indices = cell_items.iter().map(|(idx, _)| *idx).collect::<Vec<_>>();
let x = cell_items.iter().map(|(_, item)| item.x).reduce(f32::min);
let y = cell_items.iter().map(|(_, item)| item.y).reduce(f32::max);
row_cells.push(text);
row_cells.push(MatchedCell {
text,
item_indices,
x,
y,
});
}
// Pad to num_cols
while row_cells.len() < num_cols {
row_cells.push(String::new());
}
cells.push(row_cells);
raw_rows.push(row_cells);
}
// Reject if too few cells matched (stale structure tree)
@@ -148,51 +514,54 @@ pub fn detect_tables_from_struct_tree(
continue;
}
// Derive row/column positions from item geometry
let mut row_positions: Vec<f32> = Vec::new();
for row in &page_rows {
let y = row
.cells
.iter()
.flat_map(|c| c.mcids.iter())
.filter(|(_, p)| *p == page)
.filter_map(|(mcid, _)| mcid_to_items.get(mcid))
.flatten()
.map(|&idx| items[idx].y)
.reduce(f32::max)
.unwrap_or(0.0);
row_positions.push(y);
}
let has_ragged_rows = raw_rows
.iter()
.any(|row| row.iter().filter(|cell| cell.x.is_some()).count() < num_cols);
let first_row_has_tagged_header = page_rows.first().is_some_and(|row| {
let header_cells = row.cells.iter().filter(|cell| cell.is_header).count();
header_cells * 2 >= row.cells.len()
});
let fallback_col_positions =
legacy_column_positions(&page_rows, &mcid_to_items, items, page, num_cols);
let (legacy_cells, legacy_row_positions, mut legacy_item_indices) =
left_align_struct_rows(&raw_rows, num_cols);
legacy_item_indices.sort_unstable();
legacy_item_indices.dedup();
let legacy_table = Table::new(
fallback_col_positions.clone(),
legacy_row_positions,
legacy_cells,
legacy_item_indices,
);
// Column positions: use X positions of first non-empty cell in each column
let mut col_positions: Vec<f32> = vec![0.0; num_cols];
for (col, col_pos) in col_positions.iter_mut().enumerate() {
for row in &page_rows {
if col < row.cells.len() {
if let Some(x) = row.cells[col]
.mcids
.iter()
.filter(|(_, p)| *p == page)
.filter_map(|(mcid, _)| mcid_to_items.get(mcid))
.flatten()
.map(|&idx| items[idx].x)
.reduce(f32::min)
{
*col_pos = x;
break;
}
}
}
}
let col_positions = infer_column_positions(&raw_rows, &fallback_col_positions, num_cols);
let (aligned_cells, aligned_row_positions, mut aligned_item_indices) =
align_struct_rows(&raw_rows, &col_positions);
aligned_item_indices.sort_unstable();
aligned_item_indices.dedup();
all_item_indices.sort_unstable();
all_item_indices.dedup();
let mut aligned_table = Table::new(
col_positions,
aligned_row_positions,
aligned_cells,
aligned_item_indices,
);
let item_count_before_header = aligned_table.item_indices.len();
let row_count_before_header = aligned_table.cells.len();
recover_unclaimed_header_row(
&mut aligned_table,
items,
has_ragged_rows && !first_row_has_tagged_header,
);
tables.push(Table {
columns: col_positions,
rows: row_positions,
cells,
item_indices: all_item_indices,
let recovered_header = aligned_table.item_indices.len() > item_count_before_header
|| aligned_table.cells.len() > row_count_before_header;
let prefer_aligned = recovered_header;
tables.push(if prefer_aligned {
aligned_table
} else {
legacy_table
});
}
@@ -374,4 +743,419 @@ mod tests {
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 2);
assert_eq!(tables.len(), 1);
}
#[test]
fn realigns_ragged_rows_and_recovers_untagged_header() {
let items = vec![
make_item("Category", 50.0, 120.0, 1, None),
make_item("Potentially", 150.0, 120.0, 1, None),
make_item("Summary", 250.0, 120.0, 1, None),
make_item("Most commonly", 350.0, 120.0, 1, None),
make_item("concerning aspect", 150.0, 110.0, 1, None),
make_item("suggested", 350.0, 110.0, 1, None),
make_item("of circumstances", 150.0, 100.0, 1, None),
make_item("intervention", 350.0, 100.0, 1, None),
make_item("Existence of red-teaming", 150.0, 80.0, 1, Some(10)),
make_item("Important for safety", 250.0, 80.0, 1, Some(11)),
make_item("Ensure welfare interviews", 350.0, 80.0, 1, Some(12)),
make_item("Identity & self-knowledge", 50.0, 60.0, 1, Some(20)),
make_item("Lack of knowledge", 150.0, 60.0, 1, Some(21)),
make_item("Overall negative", 250.0, 60.0, 1, Some(22)),
make_item("Describe training process", 350.0, 60.0, 1, Some(23)),
make_item("Uncertainty around other copies", 150.0, 40.0, 1, Some(30)),
make_item("High uncertainty", 250.0, 40.0, 1, Some(31)),
make_item("No intervention suggested", 350.0, 40.0, 1, Some(32)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(10, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(11, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(12, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(22, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(23, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(30, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(31, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(32, 1)],
},
],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells.len(), 4);
assert_eq!(
table.cells[0],
vec![
"Category",
"Potentially concerning aspect of circumstances",
"Summary",
"Most commonly suggested intervention",
]
);
assert_eq!(table.cells[1][0], "");
assert_eq!(table.cells[1][1], "Existence of red-teaming");
assert_eq!(table.cells[2][0], "Identity & self-knowledge");
assert_eq!(table.cells[3][0], "");
assert_eq!(table.columns.len(), 4);
assert!(table.columns.windows(2).all(|w| w[0] < w[1]));
assert_eq!(table.item_indices.len(), items.len());
}
#[test]
fn does_not_absorb_caption_above_ragged_struct_table() {
let items = vec![
make_item("Table 5-7: Summary of responses", 50.0, 120.0, 1, None),
make_item("Aspect one", 150.0, 80.0, 1, Some(10)),
make_item("Summary one", 250.0, 80.0, 1, Some(11)),
make_item("Category", 50.0, 60.0, 1, Some(20)),
make_item("Aspect two", 150.0, 60.0, 1, Some(21)),
make_item("Summary two", 250.0, 60.0, 1, Some(22)),
make_item("Aspect three", 150.0, 40.0, 1, Some(30)),
make_item("Summary three", 250.0, 40.0, 1, Some(31)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(10, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(11, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(22, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(30, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(31, 1)],
},
],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells.len(), 3);
assert!(
table
.cells
.iter()
.flatten()
.all(|cell| !cell.contains("Table 5-7")),
"caption must stay outside the table"
);
assert!(!table.item_indices.contains(&0));
}
#[test]
fn keeps_existing_tagged_header_without_absorbing_intro_or_caption() {
let items = vec![
make_item(
"Eighteen people left other comments regarding the Project.",
50.0,
130.0,
1,
None,
),
make_item("Table 5-1:", 220.0, 130.0, 1, None),
make_item("Other Comments", 350.0, 130.0, 1, None),
make_item("Theme", 50.0, 110.0, 1, Some(10)),
make_item("Specific Concern/Inquiry", 200.0, 110.0, 1, Some(11)),
make_item("Response", 420.0, 110.0, 1, Some(12)),
make_item("Traffic", 50.0, 90.0, 1, Some(20)),
make_item("Road conditions", 200.0, 90.0, 1, Some(21)),
make_item("Maintenance response", 420.0, 90.0, 1, Some(22)),
make_item("Noise", 50.0, 70.0, 1, Some(30)),
make_item("Dust concerns", 200.0, 70.0, 1, Some(31)),
make_item("Mitigation response", 420.0, 70.0, 1, Some(32)),
make_item("Resource Use", 50.0, 50.0, 1, Some(40)),
make_item("Snowmobile trails", 200.0, 50.0, 1, Some(41)),
make_item("Access response", 420.0, 50.0, 1, Some(42)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![
StructTableCell {
is_header: true,
mcids: vec![(10, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(11, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(12, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(22, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(30, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(31, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(32, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(40, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(41, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(42, 1)],
},
],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells.len(), 4);
assert_eq!(
table.cells[0],
vec!["Theme", "Specific Concern/Inquiry", "Response"]
);
assert!(!table.item_indices.contains(&0));
assert!(!table.item_indices.contains(&1));
assert!(!table.item_indices.contains(&2));
}
#[test]
fn does_not_recover_header_for_narrow_two_column_table() {
let items = vec![
make_item("Alpha", 50.0, 120.0, 1, None),
make_item("Beta", 200.0, 120.0, 1, None),
make_item("First value", 200.0, 80.0, 1, Some(10)),
make_item("Only labeled row", 50.0, 60.0, 1, Some(20)),
make_item("Second value", 200.0, 60.0, 1, Some(21)),
make_item("Third value", 200.0, 40.0, 1, Some(30)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![StructTableCell {
is_header: false,
mcids: vec![(10, 1)],
}],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
],
},
StructTableRow {
cells: vec![StructTableCell {
is_header: false,
mcids: vec![(30, 1)],
}],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells.len(), 3);
assert_eq!(table.cells[0], vec!["First value", ""]);
assert_eq!(table.cells[1], vec!["Only labeled row", "Second value"]);
assert_eq!(table.cells[2], vec!["Third value", ""]);
assert!(!table.item_indices.contains(&0));
assert!(!table.item_indices.contains(&1));
}
#[test]
fn ragged_rows_without_recovered_header_keep_legacy_alignment() {
let items = vec![
make_item("Date", 150.0, 120.0, 1, Some(10)),
make_item("Title", 250.0, 120.0, 1, Some(11)),
make_item("PE", 350.0, 120.0, 1, Some(12)),
make_item("Bidder", 450.0, 120.0, 1, Some(13)),
make_item("Amount", 550.0, 120.0, 1, Some(14)),
make_item("1", 50.0, 100.0, 1, Some(20)),
make_item("8/1", 150.0, 100.0, 1, Some(21)),
make_item("Procurement", 250.0, 100.0, 1, Some(22)),
make_item("PUC", 350.0, 100.0, 1, Some(23)),
make_item("Vendor", 450.0, 100.0, 1, Some(24)),
make_item("SR1", 550.0, 100.0, 1, Some(25)),
];
let struct_tables = vec![StructTable {
rows: vec![
StructTableRow {
cells: vec![
StructTableCell {
is_header: true,
mcids: vec![(10, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(11, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(12, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(13, 1)],
},
StructTableCell {
is_header: true,
mcids: vec![(14, 1)],
},
],
},
StructTableRow {
cells: vec![
StructTableCell {
is_header: false,
mcids: vec![(20, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(21, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(22, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(23, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(24, 1)],
},
StructTableCell {
is_header: false,
mcids: vec![(25, 1)],
},
],
},
],
}];
let tables = detect_tables_from_struct_tree(&items, &struct_tables, 1);
assert_eq!(tables.len(), 1);
let table = &tables[0];
assert_eq!(table.cells[0][0], "Date");
assert_eq!(table.cells[0][4], "Amount");
assert_eq!(table.cells[0][5], "");
}
}
+96 -21
View File
@@ -1,26 +1,19 @@
//! Table-to-markdown formatting and cell cleanup.
use super::detect_heuristic::is_table_of_contents;
use super::Table;
use super::{Table, TableKind};
pub fn table_to_markdown(table: &Table) -> String {
if table.cells.is_empty() || table.cells[0].is_empty() {
return String::new();
}
// Detect TOC on the raw cells: clean_table_cells merges rows in ways
// that can make genuine data tables superficially resemble a TOC
// (short numeric cells, few columns) — but the raw detection here
// preserves the original multi-column structure and only matches the
// true TOC pattern.
//
// Tables of contents render poorly as markdown tables — emit a flat
// per-row text list instead so the page numbers stay aligned with
// their section titles rather than drifting to a separate column.
// Format from raw cells: continuation-row merging collapses separate
// TOC entries (e.g. "6.2 Contamination" + "6.2.1 SWE-bench") into a
// single line because sub-entries leave column 0 empty.
if is_table_of_contents(&table.cells) {
// TOCs render poorly as markdown tables — emit a flat per-row text list
// instead so the page numbers stay aligned with their section titles
// rather than drifting to a separate column. Format from raw cells
// because continuation-row merging in clean_table_cells collapses
// separate TOC entries (e.g. "6.2 Contamination" + "6.2.1 SWE-bench")
// into one line where sub-entries leave column 0 empty.
if table.kind == TableKind::Toc {
return format_toc_as_list(&table.cells, &[]);
}
@@ -161,6 +154,12 @@ fn is_dots_only(cell: &str) -> bool {
dots >= 3 && t.chars().all(|c| c == '.' || c.is_whitespace())
}
fn starts_with_uppercase_word(cell: &str) -> bool {
cell.chars()
.find(|c| c.is_alphanumeric())
.is_some_and(|c| c.is_uppercase())
}
/// Clean up table cells: merge continuation rows, extract footnotes, remove empty rows
fn clean_table_cells(cells: &[Vec<String>]) -> (Vec<Vec<String>>, Vec<String>) {
let mut cleaned: Vec<Vec<String>> = Vec::new();
@@ -219,11 +218,20 @@ fn clean_table_cells(cells: &[Vec<String>]) -> (Vec<Vec<String>>, Vec<String>) {
let looks_like_data_row = non_first_cells.len() >= 2
&& avg_cell_len <= 10.0
&& numeric_cells > non_first_cells.len() / 2;
let uppercase_leading_cells = non_first_cells
.iter()
.filter(|cell| starts_with_uppercase_word(cell))
.count();
let looks_like_spanning_first_column_row = first_cell.is_empty()
&& row.len() >= 4
&& non_first_cells.len() == row.len().saturating_sub(1)
&& uppercase_leading_cells >= non_first_cells.len().saturating_sub(1);
// Classic continuation: first cell empty, content in other cells
let is_classic_continuation = first_cell.is_empty()
&& !non_first_cells.is_empty()
&& !is_short_subheader
&& !looks_like_data_row
&& !looks_like_spanning_first_column_row
&& cleaned.len() > 1;
// Wrapped-cell continuation: row has fewer filled cells than the header
@@ -253,6 +261,7 @@ fn clean_table_cells(cells: &[Vec<String>]) -> (Vec<Vec<String>>, Vec<String>) {
&& filled_cells <= max_filled_for_merge
&& prev_filled > filled_cells
&& !looks_like_data_row
&& !looks_like_spanning_first_column_row
&& !is_short_subheader;
let is_continuation = is_classic_continuation || is_wrapped_continuation;
@@ -422,6 +431,65 @@ mod tests {
assert_eq!(cleaned.len(), 3);
}
#[test]
fn test_clean_table_cells_spanning_first_column_row_not_merged() {
let cells = vec![
vec![
"Category".into(),
"Potentially concerning aspect".into(),
"Summary".into(),
"Intervention".into(),
],
vec![
"Identity & self-knowledge".into(),
"Lack of knowledge".into(),
"Overall negative".into(),
"Describe training".into(),
],
vec![
"".into(),
"Uncertainty around other copies".into(),
"High uncertainty".into(),
"No intervention suggested".into(),
],
];
let (cleaned, _) = clean_table_cells(&cells);
assert_eq!(cleaned.len(), 3);
assert_eq!(cleaned[2][0], "");
assert_eq!(cleaned[2][1], "Uncertainty around other copies");
}
#[test]
fn test_clean_table_cells_full_width_continuation_row_still_merges_when_lowercase() {
let cells = vec![
vec![
"Classification".into(),
"Before tax".into(),
"After tax".into(),
"Standard equipment".into(),
"Options".into(),
],
vec![
"Exclusive Special".into(),
"83,500,000".into(),
"79,275,000".into(),
"Standard equipment".into(),
"Option A".into(),
],
vec![
"".into(),
"with 3.5% individual consumption tax applied".into(),
"with 3.5% individual consumption tax applied".into(),
"lighting(crash pad)".into(),
"sound system".into(),
],
];
let (cleaned, _) = clean_table_cells(&cells);
assert_eq!(cleaned.len(), 2);
assert!(cleaned[1][1].contains("83,500,000"));
assert!(cleaned[1][1].contains("with 3.5%"));
}
#[test]
fn test_clean_table_cells_header_row_not_merged() {
// Continuation requires cleaned.len() > 1 (don't merge into header)
@@ -470,6 +538,7 @@ mod tests {
vec!["Bob".into(), "25".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
assert!(md.contains("|Name|"));
@@ -485,6 +554,7 @@ mod tests {
rows: vec![500.0],
cells: vec![vec!["Only".into(), "Row".into()]],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
assert!(md.contains("|Only|"));
@@ -498,6 +568,7 @@ mod tests {
rows: vec![],
cells: vec![],
item_indices: vec![],
kind: TableKind::Data,
};
assert_eq!(table_to_markdown(&table), "");
}
@@ -513,6 +584,7 @@ mod tests {
vec!["(1)".into(), "Footnote text".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
assert!(md.contains("(1) Footnote text"));
@@ -528,6 +600,7 @@ mod tests {
vec!["太郎".into(), "25".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
assert!(md.contains("名前"));
@@ -541,6 +614,7 @@ mod tests {
rows: vec![500.0],
cells: vec![vec![]],
item_indices: vec![],
kind: TableKind::Data,
};
assert_eq!(table_to_markdown(&table), "");
}
@@ -550,10 +624,10 @@ mod tests {
// A TOC-shaped table with section numbers in col 0 and page numbers
// in the last column should render as a flat list, not a markdown
// table, so the page numbers stay on the same line as their titles.
let table = Table {
columns: vec![50.0, 80.0, 300.0],
rows: vec![500.0; 5],
cells: vec![
let table = Table::new(
vec![50.0, 80.0, 300.0],
vec![500.0; 5],
vec![
vec![
"4.3".into(),
"Case studies and targeted evaluations".into(),
@@ -572,8 +646,9 @@ mod tests {
vec!["4.4".into(), "Capability evaluations".into(), "101".into()],
vec!["4.5".into(), "White-box analyses".into(), "113".into()],
],
item_indices: vec![],
};
vec![],
);
assert_eq!(table.kind, TableKind::Toc);
let md = table_to_markdown(&table);
assert!(
!md.contains("|---|"),
+6
View File
@@ -499,6 +499,7 @@ pub(crate) fn recover_header_row(
#[cfg(test)]
mod tests {
use super::*;
use crate::tables::TableKind;
use crate::types::ItemType;
fn make_item(text: &str, x: f32, y: f32, font_size: f32) -> TextItem {
@@ -756,6 +757,7 @@ mod tests {
rows: vec![500.0, 480.0],
cells: vec![vec!["A".into(), "B".into()], vec!["C".into(), "D".into()]],
item_indices: vec![2, 3],
kind: TableKind::Data,
};
recover_header_row(&mut table, &all_items, 9.0);
@@ -774,6 +776,7 @@ mod tests {
rows: vec![500.0],
cells: vec![vec!["A".into(), "B".into()]],
item_indices: vec![0, 1],
kind: TableKind::Data,
};
let rows_before = table.rows.len();
@@ -794,6 +797,7 @@ mod tests {
rows: vec![500.0, 480.0],
cells: vec![vec!["A".into(), "B".into()], vec!["C".into(), "D".into()]],
item_indices: vec![2, 3],
kind: TableKind::Data,
};
let rows_before = table.rows.len();
@@ -814,6 +818,7 @@ mod tests {
rows: vec![500.0],
cells: vec![vec!["A".into(), "B".into()]],
item_indices: vec![1, 2],
kind: TableKind::Data,
};
let rows_before = table.rows.len();
@@ -829,6 +834,7 @@ mod tests {
rows: vec![],
cells: vec![],
item_indices: vec![],
kind: TableKind::Data,
};
recover_header_row(&mut table, &all_items, 9.0);
+51 -11
View File
@@ -9,13 +9,16 @@ mod detect_struct;
mod financial;
mod format;
mod grid;
pub mod structured;
pub use detect_heuristic::detect_tables;
pub(crate) use detect_heuristic::is_table_of_contents;
pub use detect_lines::detect_tables_from_lines;
pub(crate) use detect_rects::cluster_rects;
pub use detect_rects::{detect_tables_from_rects, RectHintRegion};
pub use detect_struct::detect_tables_from_struct_tree;
pub use format::table_to_markdown;
pub use structured::{cells_to_markdown, StructuredCell};
use crate::types::TextItem;
@@ -166,12 +169,12 @@ pub(crate) fn try_build_rect_guided_table(
used_indices.sort_unstable();
used_indices.dedup();
Some(Table {
columns: col_boundaries,
rows: row_boundaries,
Some(Table::new(
col_boundaries,
row_boundaries,
cells,
item_indices: used_indices,
})
used_indices,
))
}
/// Split a TextItem whose text contains multiple whitespace-separated tokens
@@ -541,12 +544,22 @@ pub(crate) fn try_build_table_from_columns(items: &[TextItem], page: u32) -> Opt
multi_col_rows
);
Some(Table {
columns: col_xs,
rows: row_ys,
cells,
item_indices,
})
Some(Table::new(col_xs, row_ys, cells, item_indices))
}
/// What kind of structure a detected `Table` represents. Classification is
/// computed once at construction so consumers don't have to re-analyze the
/// cells (and stay consistent across detection backends).
#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
pub enum TableKind {
/// A real data table — renders as markdown table syntax.
#[default]
Data,
/// A table of contents — renders as a flat list with tab-aligned page
/// numbers via `format_toc_as_list`. Detected through the table pipeline
/// because TOCs share row/column structure with tables, but they are not
/// data tables and shouldn't appear in `pages_with_tables` etc.
Toc,
}
/// A detected table.
@@ -560,6 +573,31 @@ pub struct Table {
pub cells: Vec<Vec<String>>,
/// Items that belong to this table
pub item_indices: Vec<usize>,
/// Data table vs TOC. Set by `Table::new` from `cells`.
pub kind: TableKind,
}
impl Table {
/// Build a table and classify it (data vs TOC) from its cells.
pub fn new(
columns: Vec<f32>,
rows: Vec<f32>,
cells: Vec<Vec<String>>,
item_indices: Vec<usize>,
) -> Self {
let kind = if is_table_of_contents(&cells) {
TableKind::Toc
} else {
TableKind::Data
};
Self {
columns,
rows,
cells,
item_indices,
kind,
}
}
}
#[cfg(test)]
@@ -642,6 +680,7 @@ mod tests {
vec!["Cell 1".into(), "Cell 2".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
@@ -1002,6 +1041,7 @@ mod tests {
vec!["3".into(), "5/2".into(), "Item C".into(), "300".into()],
],
item_indices: vec![],
kind: TableKind::Data,
};
let md = table_to_markdown(&table);
+972
View File
@@ -0,0 +1,972 @@
//! Structure-recovery-aware (TSR) table assembly.
//!
//! Consumes the raw output of an external table-structure recognition model
//! (e.g. SLANet on PaddleOCR): a flat list of HTML structure tokens plus a
//! parallel list of per-cell bboxes. Pairs each cell open-tag with its bbox
//! in document order, tracks row/column position with rowspan/colspan
//! awareness, and emits a markdown pipe-table.
//!
//! No real HTML parser is needed — the token grammar is restricted (see
//! [`parse_structure`]), so a small state machine is enough.
//!
//! Cell text is supplied separately by the caller (typically by overlap-
//! testing PDF text items against each cell's page-PDF-pt bbox).
use std::collections::{HashMap, HashSet};
/// A single resolved cell, with both structural metadata and its bbox in
/// page PDF-points (top-left origin).
#[derive(Debug, Clone)]
pub struct StructuredCell {
/// 0-indexed grid row.
pub row: usize,
/// 0-indexed grid column.
pub col: usize,
/// 1 for a normal cell.
pub rowspan: usize,
/// 1 for a normal cell.
pub colspan: usize,
/// `true` when the cell is a `<th>` or sits inside `<thead>`.
pub is_header: bool,
/// Cell text (filled in by the caller after overlap-testing PDF items).
pub text: String,
/// Axis-aligned bbox `[x1, y1, x2, y2]` in page PDF-points, top-left origin.
pub page_pt_bbox: [f32; 4],
}
/// Intermediate parse result before the caller fills in text + page coords.
#[derive(Debug, Clone)]
pub(crate) struct CellSlot {
pub row: usize,
pub col: usize,
pub rowspan: usize,
pub colspan: usize,
pub is_header: bool,
/// Index into the parallel `cell_bboxes` array.
pub bbox_idx: usize,
}
/// Parse a sequence of SLANet structure tokens into ordered cell slots.
///
/// Token grammar (no real HTML parsing required):
/// - Section markers: `<thead>`, `</thead>`, `<tbody>`, `</tbody>` and
/// wrapper tokens (`<html>`, `<body>`, `<table>`, plus closing variants)
/// are tracked or skipped.
/// - Row markers: `<tr>` opens a new row, `</tr>` is informational.
/// - Empty cell, single token: `<td></td>` or `<th></th>`.
/// - Cell with attributes, multi-token sequence: `<td` (or `<th`), then
/// attribute fragments like ` colspan="4"`, then `>`, then later `</td>`
/// (or `</th>`). Cells get paired with the next bbox in document order.
///
/// Cells inside `<thead>` and any `<th>` cells are flagged as headers.
/// rowspan/colspan attributes are honoured and prior-row rowspans push
/// later-row cells to the right.
pub(crate) fn parse_structure(tokens: &[String]) -> Vec<CellSlot> {
let mut slots: Vec<CellSlot> = Vec::new();
let mut occupied: HashSet<(usize, usize)> = HashSet::new();
let mut row: usize = 0;
let mut col: usize = 0;
let mut bbox_idx: usize = 0;
let mut in_thead = false;
let mut started_first_row = false;
let mut i = 0;
while i < tokens.len() {
let tok = tokens[i].trim();
match tok {
"<thead>" => {
in_thead = true;
}
"</thead>" => {
in_thead = false;
}
"<tr>" => {
if started_first_row {
row += 1;
}
col = 0;
started_first_row = true;
}
"<td></td>" | "<th></th>" => {
let is_th = tok == "<th></th>";
while occupied.contains(&(row, col)) {
col += 1;
}
slots.push(CellSlot {
row,
col,
rowspan: 1,
colspan: 1,
is_header: in_thead || is_th,
bbox_idx,
});
bbox_idx += 1;
col += 1;
}
"<td" | "<th" => {
let is_th = tok == "<th";
let mut rowspan: usize = 1;
let mut colspan: usize = 1;
// Consume attribute fragments until we hit ">".
i += 1;
while i < tokens.len() && tokens[i].trim() != ">" {
let attr = tokens[i].as_str();
if let Some(v) = parse_int_attr(attr, "rowspan") {
rowspan = v.max(1);
} else if let Some(v) = parse_int_attr(attr, "colspan") {
colspan = v.max(1);
}
i += 1;
}
// i now points at the `>` token (or off the end if malformed).
while occupied.contains(&(row, col)) {
col += 1;
}
slots.push(CellSlot {
row,
col,
rowspan,
colspan,
is_header: in_thead || is_th,
bbox_idx,
});
for r in row..row + rowspan {
for c in col..col + colspan {
occupied.insert((r, c));
}
}
bbox_idx += 1;
col += colspan;
}
// Wrapper / informational tokens — no-op.
_ => {}
}
i += 1;
}
slots
}
/// Parse an attribute fragment like ` colspan="4"` or `rowspan='2'`.
///
/// Tolerates leading whitespace and either single or double quotes.
fn parse_int_attr(s: &str, name: &str) -> Option<usize> {
let trimmed = s.trim();
if !trimmed.starts_with(name) {
return None;
}
let rest = trimmed[name.len()..].trim_start();
let rest = rest.strip_prefix('=')?.trim_start();
let value = rest
.trim_start_matches(['"', '\''])
.trim_end_matches(['"', '\'']);
value.parse().ok()
}
/// Convert a SLANet polygon (4 or 8 elements) into an axis-aligned
/// `[x1, y1, x2, y2]` rect.
///
/// 8-element form: `[x1,y1, x2,y1, x2,y2, x1,y2]` (4 corners). We ignore the
/// implicit corner order and just take min/max so rotated polygons collapse
/// to a sane bounding box.
///
/// 4-element form: `[x1, y1, x2, y2]` (axis-aligned, older SLANet variants).
pub(crate) fn polygon_to_aabb(coords: &[f32]) -> Option<[f32; 4]> {
match coords.len() {
4 => {
let x1 = coords[0].min(coords[2]);
let y1 = coords[1].min(coords[3]);
let x2 = coords[0].max(coords[2]);
let y2 = coords[1].max(coords[3]);
Some([x1, y1, x2, y2])
}
8 => {
let xs = [coords[0], coords[2], coords[4], coords[6]];
let ys = [coords[1], coords[3], coords[5], coords[7]];
let x1 = xs.iter().copied().fold(f32::INFINITY, f32::min);
let y1 = ys.iter().copied().fold(f32::INFINITY, f32::min);
let x2 = xs.iter().copied().fold(f32::NEG_INFINITY, f32::max);
let y2 = ys.iter().copied().fold(f32::NEG_INFINITY, f32::max);
if x1.is_finite() && y1.is_finite() && x2.is_finite() && y2.is_finite() {
Some([x1, y1, x2, y2])
} else {
None
}
}
_ => None,
}
}
/// Convert a cell rect from crop image-pixel space to page PDF-points
/// (top-left origin), given the crop's PDF-point offset on the page and the
/// DPI the crop image was rendered at.
pub(crate) fn cell_px_to_page_pt(
cell_px: [f32; 4],
render_dpi: f32,
crop_origin_pt: [f32; 2],
) -> [f32; 4] {
let pt_per_px = if render_dpi > 0.0 {
72.0 / render_dpi
} else {
1.0
};
let [x_off, y_off] = crop_origin_pt;
[
cell_px[0] * pt_per_px + x_off,
cell_px[1] * pt_per_px + y_off,
cell_px[2] * pt_per_px + x_off,
cell_px[3] * pt_per_px + y_off,
]
}
/// Refine TSR cell bboxes into non-overlapping row/column bands.
///
/// SLANet-style bboxes are often plausible but too tall on dense borderless
/// tables. Native PDF text assignment is more reliable when each parsed row
/// owns the band between neighboring row centers instead of the full model box.
pub(crate) fn normalize_cell_bands(cells: &mut [StructuredCell]) {
if cells.len() < 2 {
return;
}
let row_bands = derive_axis_bands(cells, Axis::Y);
let col_bands = derive_axis_bands(cells, Axis::X);
for cell in cells {
let row_end = cell.row + cell.rowspan.max(1).saturating_sub(1);
if let (Some(&(y1, _)), Some(&(_, y2))) =
(row_bands.get(&cell.row), row_bands.get(&row_end))
{
let clamped_y1 = cell.page_pt_bbox[1].max(y1);
let clamped_y2 = cell.page_pt_bbox[3].min(y2);
if clamped_y1 < clamped_y2 {
cell.page_pt_bbox[1] = clamped_y1;
cell.page_pt_bbox[3] = clamped_y2;
}
}
let col_end = cell.col + cell.colspan.max(1).saturating_sub(1);
if let (Some(&(x1, _)), Some(&(_, x2))) =
(col_bands.get(&cell.col), col_bands.get(&col_end))
{
let clamped_x1 = cell.page_pt_bbox[0].max(x1);
let clamped_x2 = cell.page_pt_bbox[2].min(x2);
if clamped_x1 < clamped_x2 {
cell.page_pt_bbox[0] = clamped_x1;
cell.page_pt_bbox[2] = clamped_x2;
}
}
}
}
#[derive(Clone, Copy)]
enum Axis {
X,
Y,
}
fn derive_axis_bands(cells: &[StructuredCell], axis: Axis) -> HashMap<usize, (f32, f32)> {
let mut by_index: HashMap<usize, Vec<(f32, f32)>> = HashMap::new();
// Prefer non-spanning cells so colspan/rowspan boxes do not skew a single
// column/row center. If an axis has no non-spanning examples for an index,
// fall back to anchored cells below.
for cell in cells {
let span = match axis {
Axis::X => cell.colspan.max(1),
Axis::Y => cell.rowspan.max(1),
};
if span == 1 {
let idx = match axis {
Axis::X => cell.col,
Axis::Y => cell.row,
};
by_index
.entry(idx)
.or_default()
.push(axis_bounds(cell.page_pt_bbox, axis));
}
}
for cell in cells {
let idx = match axis {
Axis::X => cell.col,
Axis::Y => cell.row,
};
if !by_index.contains_key(&idx) {
by_index
.entry(idx)
.or_default()
.push(axis_bounds(cell.page_pt_bbox, axis));
}
}
let mut rows: Vec<(usize, f32, f32, f32)> = by_index
.into_iter()
.filter_map(|(idx, bounds)| {
let mut min_edge = f32::INFINITY;
let mut max_edge = f32::NEG_INFINITY;
let mut center_sum = 0.0;
let mut count = 0usize;
for (lo, hi) in bounds {
if lo.is_finite() && hi.is_finite() && lo < hi {
min_edge = min_edge.min(lo);
max_edge = max_edge.max(hi);
center_sum += (lo + hi) * 0.5;
count += 1;
}
}
(count > 0).then_some((idx, center_sum / count as f32, min_edge, max_edge))
})
.collect();
if rows.len() < 2 {
return rows
.into_iter()
.map(|(idx, _center, lo, hi)| (idx, (lo, hi)))
.collect();
}
rows.sort_by_key(|(idx, _, _, _)| *idx);
let mut bands = HashMap::new();
for i in 0..rows.len() {
let (idx, _center, min_edge, max_edge) = rows[i];
let lo = if i == 0 {
min_edge
} else {
(rows[i - 1].1 + rows[i].1) * 0.5
};
let hi = if i + 1 == rows.len() {
max_edge
} else {
(rows[i].1 + rows[i + 1].1) * 0.5
};
if lo.is_finite() && hi.is_finite() && lo < hi {
bands.insert(idx, (lo, hi));
}
}
bands
}
fn axis_bounds(bbox: [f32; 4], axis: Axis) -> (f32, f32) {
match axis {
Axis::X => (bbox[0].min(bbox[2]), bbox[0].max(bbox[2])),
Axis::Y => (bbox[1].min(bbox[3]), bbox[1].max(bbox[3])),
}
}
/// Sanitize cell text for inclusion in a markdown pipe-table cell:
/// collapse whitespace runs, drop newlines/tabs (cells must be one line),
/// and escape pipes that would otherwise break the table.
fn sanitize_cell(text: &str) -> String {
let mut s = String::with_capacity(text.len());
let mut prev_space = false;
for c in text.chars() {
match c {
'|' => {
s.push_str("\\|");
prev_space = false;
}
'\n' | '\r' | '\t' | ' ' => {
if !prev_space {
s.push(' ');
}
prev_space = true;
}
other => {
s.push(other);
prev_space = false;
}
}
}
s.trim().to_string()
}
/// Render a list of explicitly-positioned cells as a markdown pipe-table.
///
/// Grid dimensions are inferred from the cells' (row, col, rowspan, colspan)
/// extents. A cell with colspan/rowspan > 1 is rendered in its top-left
/// position; the absorbed grid positions are emitted as empty cells so the
/// markdown stays a valid rectangular grid that downstream readers can
/// column-count correctly.
///
/// The separator row (`|---|...|`) is emitted after the **last** row that
/// contains a header cell (`is_header == true`). When no cells are flagged
/// as headers — e.g. the upstream TSR model didn't emit `<thead>`/`<th>` —
/// the separator falls back to "after row 0" so the output is still a
/// valid pipe-table.
pub fn cells_to_markdown(cells: &[StructuredCell]) -> String {
if cells.is_empty() {
return String::new();
}
let num_rows = cells
.iter()
.map(|c| c.row + c.rowspan.max(1))
.max()
.unwrap_or(0);
let num_cols = cells
.iter()
.map(|c| c.col + c.colspan.max(1))
.max()
.unwrap_or(0);
if num_rows == 0 || num_cols == 0 {
return String::new();
}
// Separator goes after the last header row, falling back to row 0 when
// no header cells exist. Clamped into range so a malformed cell with
// row >= num_rows can't push it past the table.
let separator_after_row = cells
.iter()
.filter(|c| c.is_header)
.map(|c| c.row)
.max()
.unwrap_or(0)
.min(num_rows.saturating_sub(1));
let mut grid: Vec<Vec<String>> = vec![vec![String::new(); num_cols]; num_rows];
for cell in cells {
if cell.row < num_rows && cell.col < num_cols {
grid[cell.row][cell.col] = sanitize_cell(&cell.text);
}
}
let mut output = String::new();
for (row_idx, row) in grid.iter().enumerate() {
output.push('|');
for cell in row {
output.push_str(cell);
output.push('|');
}
output.push('\n');
if row_idx == separator_after_row {
output.push('|');
for _ in 0..num_cols {
output.push_str("---|");
}
output.push('\n');
}
}
output
}
#[cfg(test)]
mod tests {
use super::*;
fn t(s: &str) -> String {
s.to_string()
}
/// Tokens for the synthetic 3×3 grid example (one colspan-4 row + two
/// data rows of 4 cells each = 9 cells total, 3 rows × 4 cols).
fn synthetic_3x3_tokens() -> Vec<String> {
vec![
"<html>",
"<body>",
"<table>",
"<tbody>",
"<tr>",
"<td",
" colspan=\"4\"",
">",
"</td>",
"</tr>",
"<tr>",
"<td></td>",
"<td></td>",
"<td></td>",
"<td></td>",
"</tr>",
"<tr>",
"<td></td>",
"<td></td>",
"<td></td>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
"</body>",
"</html>",
]
.into_iter()
.map(t)
.collect()
}
/// Bboxes for the synthetic 3×3 grid (8-element polygon form), all
/// within a 400×120 px crop.
fn synthetic_3x3_bboxes() -> Vec<Vec<f32>> {
vec![
vec![3.0, 2.0, 395.0, 2.0, 396.0, 59.0, 3.0, 59.0],
vec![26.0, 62.0, 140.0, 62.0, 141.0, 120.0, 26.0, 120.0],
vec![149.0, 64.0, 248.0, 64.0, 248.0, 119.0, 149.0, 119.0],
vec![257.0, 64.0, 350.0, 64.0, 350.0, 119.0, 257.0, 119.0],
vec![359.0, 64.0, 395.0, 64.0, 395.0, 119.0, 359.0, 119.0],
vec![26.0, 122.0, 140.0, 122.0, 140.0, 178.0, 26.0, 178.0],
vec![149.0, 124.0, 248.0, 124.0, 248.0, 179.0, 149.0, 179.0],
vec![257.0, 124.0, 350.0, 124.0, 350.0, 179.0, 257.0, 179.0],
vec![359.0, 124.0, 395.0, 124.0, 395.0, 179.0, 359.0, 179.0],
]
}
#[test]
fn parse_structure_synthetic_3x3() {
let tokens = synthetic_3x3_tokens();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 9, "should parse 9 cells");
// Cell 0: row 0 col 0, colspan 4
assert_eq!(slots[0].row, 0);
assert_eq!(slots[0].col, 0);
assert_eq!(slots[0].colspan, 4);
assert_eq!(slots[0].rowspan, 1);
// Cells 1..5: row 1, cols 0..3
for (i, slot) in slots.iter().enumerate().skip(1).take(4) {
assert_eq!(slot.row, 1, "cell {i}: row should be 1");
assert_eq!(slot.col, i - 1, "cell {i}: col should be {}", i - 1);
assert_eq!(slot.colspan, 1);
assert_eq!(slot.rowspan, 1);
}
// Cells 5..9: row 2, cols 0..3
for (i, slot) in slots.iter().enumerate().skip(5).take(4) {
assert_eq!(slot.row, 2, "cell {i}: row should be 2");
assert_eq!(slot.col, i - 5);
assert_eq!(slot.colspan, 1);
}
}
#[test]
fn polygon_to_aabb_8elt() {
// Synthetic cell bbox 0
let coords = vec![3.0, 2.0, 395.0, 2.0, 396.0, 59.0, 3.0, 59.0];
let aabb = polygon_to_aabb(&coords).unwrap();
assert_eq!(aabb, [3.0, 2.0, 396.0, 59.0]);
}
#[test]
fn polygon_to_aabb_4elt() {
let coords = vec![5.0, 10.0, 50.0, 60.0];
let aabb = polygon_to_aabb(&coords).unwrap();
assert_eq!(aabb, [5.0, 10.0, 50.0, 60.0]);
}
#[test]
fn polygon_to_aabb_4elt_unordered() {
// Caller may pass corners in any order; min/max should normalise.
let coords = vec![50.0, 60.0, 5.0, 10.0];
let aabb = polygon_to_aabb(&coords).unwrap();
assert_eq!(aabb, [5.0, 10.0, 50.0, 60.0]);
}
#[test]
fn polygon_to_aabb_invalid_len() {
assert!(polygon_to_aabb(&[1.0, 2.0, 3.0]).is_none());
assert!(polygon_to_aabb(&[1.0; 6]).is_none());
assert!(polygon_to_aabb(&[]).is_none());
}
#[test]
fn synthetic_3x3_aabbs_inside_crop() {
// All 9 bboxes should produce valid (x1<x2, y1<y2) rects within the
// crop bounds (400 wide, ~180 tall by inspection of the fixture).
let bboxes = synthetic_3x3_bboxes();
assert_eq!(bboxes.len(), 9);
for (i, bb) in bboxes.iter().enumerate() {
let aabb = polygon_to_aabb(bb).unwrap_or_else(|| panic!("bbox {i} invalid"));
assert!(aabb[0] < aabb[2], "bbox {i}: x1 < x2");
assert!(aabb[1] < aabb[3], "bbox {i}: y1 < y2");
assert!(aabb[0] >= 0.0 && aabb[2] <= 500.0, "bbox {i}: within crop");
assert!(aabb[1] >= 0.0 && aabb[3] <= 200.0, "bbox {i}: within crop");
}
}
#[test]
fn normalize_cell_bands_splits_overlapping_slanet_rows() {
let mut cells = vec![
StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 1,
is_header: true,
text: String::new(),
page_pt_bbox: [10.0, 100.0, 90.0, 120.0],
},
StructuredCell {
row: 0,
col: 1,
rowspan: 1,
colspan: 1,
is_header: true,
text: String::new(),
page_pt_bbox: [90.0, 100.0, 170.0, 120.0],
},
StructuredCell {
row: 1,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [10.0, 116.0, 90.0, 136.0],
},
StructuredCell {
row: 1,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [90.0, 116.0, 170.0, 136.0],
},
];
normalize_cell_bands(&mut cells);
assert_eq!(cells[0].page_pt_bbox[3], cells[2].page_pt_bbox[1]);
assert_eq!(cells[1].page_pt_bbox[3], cells[3].page_pt_bbox[1]);
assert!(
(cells[0].page_pt_bbox[3] - 118.0).abs() < 0.01,
"row separator should be midpoint between row centers: {:?}",
cells
);
}
#[test]
fn normalize_cell_bands_preserves_colspan_extent() {
let mut cells = vec![
StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 2,
is_header: true,
text: String::new(),
page_pt_bbox: [8.0, 80.0, 172.0, 98.0],
},
StructuredCell {
row: 1,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [10.0, 96.0, 90.0, 114.0],
},
StructuredCell {
row: 1,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: String::new(),
page_pt_bbox: [88.0, 96.0, 170.0, 114.0],
},
];
normalize_cell_bands(&mut cells);
assert!(
cells[0].page_pt_bbox[0] <= cells[1].page_pt_bbox[0],
"spanning cell should retain the first column's left edge"
);
assert!(
cells[0].page_pt_bbox[2] >= cells[2].page_pt_bbox[2],
"spanning cell should retain the last column's right edge"
);
}
#[test]
fn parse_int_attr_basic() {
assert_eq!(parse_int_attr(" colspan=\"4\"", "colspan"), Some(4));
assert_eq!(parse_int_attr(" rowspan=\"2\"", "rowspan"), Some(2));
assert_eq!(parse_int_attr("colspan='3'", "colspan"), Some(3));
assert_eq!(parse_int_attr(" colspan=\"4\"", "rowspan"), None);
assert_eq!(parse_int_attr(" class=\"foo\"", "colspan"), None);
}
#[test]
fn parse_structure_rowspan_pushes_next_row_right() {
// <tr><td rowspan="2">A</td><td>B</td></tr><tr><td>C</td></tr>
// Expected: A at (0,0), B at (0,1), C at (1,1) — col 0 of row 1
// is occupied by A's rowspan.
let tokens: Vec<String> = vec![
"<table>",
"<tbody>",
"<tr>",
"<td",
" rowspan=\"2\"",
">",
"</td>",
"<td></td>",
"</tr>",
"<tr>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(t)
.collect();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 3);
assert_eq!((slots[0].row, slots[0].col), (0, 0));
assert_eq!(slots[0].rowspan, 2);
assert_eq!((slots[1].row, slots[1].col), (0, 1));
// C should be at (1, 1) because (1, 0) is occupied by A's rowspan.
assert_eq!((slots[2].row, slots[2].col), (1, 1));
}
#[test]
fn parse_structure_thead_marks_headers() {
// <thead><tr><th>H1</th><th>H2</th></tr></thead>
// <tbody><tr><td>D1</td><td>D2</td></tr></tbody>
let tokens: Vec<String> = vec![
"<table>",
"<thead>",
"<tr>",
"<th></th>",
"<th></th>",
"</tr>",
"</thead>",
"<tbody>",
"<tr>",
"<td></td>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(t)
.collect();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 4);
assert!(slots[0].is_header && slots[1].is_header);
assert!(!slots[2].is_header && !slots[3].is_header);
}
#[test]
fn parse_structure_th_outside_thead_still_header() {
// A row-header style: leading <th> in tbody.
let tokens: Vec<String> = vec![
"<table>",
"<tbody>",
"<tr>",
"<th></th>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(t)
.collect();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 2);
assert!(slots[0].is_header);
assert!(!slots[1].is_header);
}
#[test]
fn parse_structure_th_with_attrs() {
let tokens: Vec<String> = vec![
"<table>",
"<thead>",
"<tr>",
"<th",
" colspan=\"2\"",
">",
"</th>",
"</tr>",
"</thead>",
"</table>",
]
.into_iter()
.map(t)
.collect();
let slots = parse_structure(&tokens);
assert_eq!(slots.len(), 1);
assert_eq!(slots[0].colspan, 2);
assert!(slots[0].is_header);
}
#[test]
fn cells_to_markdown_synthetic_3x3() {
// Build the cells the parser would produce for the synthetic grid,
// and provide some sample text so we can sanity-check output.
let cells = vec![
StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 4,
is_header: false,
text: "Title".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 1,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: "a".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 1,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: "b".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 1,
col: 2,
rowspan: 1,
colspan: 1,
is_header: false,
text: "c".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 1,
col: 3,
rowspan: 1,
colspan: 1,
is_header: false,
text: "d".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
];
let md = cells_to_markdown(&cells);
// Header row contains the spanning cell text in col 0 and pads to 4 cols.
// Absorbed-by-colspan positions render as empty cells (no padding).
assert!(md.starts_with("|Title||||\n"), "got: {md}");
assert!(md.contains("|---|---|---|---|\n"));
assert!(md.contains("|a|b|c|d|\n"));
}
#[test]
fn cells_to_markdown_escapes_pipes() {
let cells = vec![
StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: "a|b".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
StructuredCell {
row: 0,
col: 1,
rowspan: 1,
colspan: 1,
is_header: false,
text: "x".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
},
];
let md = cells_to_markdown(&cells);
assert!(md.contains("|a\\|b|x|"));
}
#[test]
fn cells_to_markdown_collapses_whitespace_and_newlines() {
let cells = vec![StructuredCell {
row: 0,
col: 0,
rowspan: 1,
colspan: 1,
is_header: false,
text: "foo \n bar\tbaz".into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
}];
let md = cells_to_markdown(&cells);
assert!(md.contains("|foo bar baz|"));
}
fn cell(row: usize, col: usize, is_header: bool, text: &str) -> StructuredCell {
StructuredCell {
row,
col,
rowspan: 1,
colspan: 1,
is_header,
text: text.into(),
page_pt_bbox: [0.0, 0.0, 0.0, 0.0],
}
}
#[test]
fn cells_to_markdown_separator_after_last_header_row() {
// Two-row header (a multi-row thead), then two body rows. Separator
// should land after row 1 (the LAST header row), not after row 0.
let cells = vec![
cell(0, 0, true, "H0a"),
cell(0, 1, true, "H0b"),
cell(1, 0, true, "H1a"),
cell(1, 1, true, "H1b"),
cell(2, 0, false, "d0a"),
cell(2, 1, false, "d0b"),
cell(3, 0, false, "d1a"),
cell(3, 1, false, "d1b"),
];
let md = cells_to_markdown(&cells);
let expected = "|H0a|H0b|\n|H1a|H1b|\n|---|---|\n|d0a|d0b|\n|d1a|d1b|\n";
assert_eq!(md, expected, "got: {md}");
}
#[test]
fn cells_to_markdown_separator_when_row_0_not_header() {
// Row 0 is not flagged as a header but row 1 is. Separator should
// follow row 1 (the header), demonstrating that we don't blindly
// emit after row 0.
let cells = vec![
cell(0, 0, false, "x0a"),
cell(0, 1, false, "x0b"),
cell(1, 0, true, "Hdr1"),
cell(1, 1, true, "Hdr2"),
cell(2, 0, false, "data1"),
cell(2, 1, false, "data2"),
];
let md = cells_to_markdown(&cells);
// Confirm the separator is NOT after row 0.
assert!(!md.starts_with("|x0a|x0b|\n|---|"), "got: {md}");
// Confirm it IS after row 1.
assert!(
md.contains("|Hdr1|Hdr2|\n|---|---|\n|data1|data2|"),
"got: {md}"
);
}
#[test]
fn cells_to_markdown_no_headers_falls_back_to_row_0() {
// No header cells at all — fallback: separator after row 0 so the
// output is still a valid markdown pipe-table.
let cells = vec![
cell(0, 0, false, "a"),
cell(0, 1, false, "b"),
cell(1, 0, false, "c"),
cell(1, 1, false, "d"),
];
let md = cells_to_markdown(&cells);
assert_eq!(md, "|a|b|\n|---|---|\n|c|d|\n");
}
}
+487 -15
View File
@@ -4,10 +4,10 @@ use pdf_inspector::detector::{DetectionConfig, ScanStrategy};
use pdf_inspector::extractor::group_into_lines;
use pdf_inspector::types::TextLine;
use pdf_inspector::{
detect_pdf_type, extract_pages_markdown_mem, extract_tables_in_regions_mem, extract_text,
extract_text_in_regions_mem, extract_text_with_positions, process_pdf_mem,
process_pdf_with_options, to_markdown, MarkdownOptions, PdfError, PdfOptions, PdfType,
TextItem,
detect_pdf_type, extract_pages_markdown, extract_pages_markdown_mem,
extract_tables_in_regions_mem, extract_text, extract_text_in_regions_mem,
extract_text_with_positions, process_pdf_mem, process_pdf_with_options, to_markdown,
MarkdownOptions, PdfError, PdfOptions, PdfType, TextItem,
};
use std::collections::HashSet;
@@ -1474,6 +1474,440 @@ fn test_bits_pilani_page8_table_detection() {
assert!(!region.needs_ocr, "Page 8 table should still be detected");
}
// =========================================================================
// extract_tables_with_structure_mem tests (TSR-aware path)
// =========================================================================
/// Build an 8-element 4-corner polygon `[x1,y1, x2,y1, x2,y2, x1,y2]` from
/// an axis-aligned rect — matches the format SLANet emits for cell bboxes.
fn poly(x1: f32, y1: f32, x2: f32, y2: f32) -> Vec<f32> {
vec![x1, y1, x2, y1, x2, y2, x1, y2]
}
fn synthetic_dense_table_pdf() -> Vec<u8> {
use lopdf::content::{Content, Operation};
use lopdf::{dictionary, Document, Object, Stream};
let mut doc = Document::with_version("1.5");
let pages_id = doc.new_object_id();
let page_id = doc.new_object_id();
let font_id = doc.new_object_id();
let content_id = doc.new_object_id();
doc.objects.insert(
font_id,
dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
}
.into(),
);
let operations = vec![
Operation::new("BT", vec![]),
Operation::new("Tf", vec!["F1".into(), 10.into()]),
Operation::new("Td", vec![20.into(), 700.into()]),
Operation::new("Tj", vec![Object::string_literal("Branch Name")]),
Operation::new("Td", vec![100.into(), 0.into()]),
Operation::new("Tj", vec![Object::string_literal("Deposits")]),
Operation::new("Td", vec![Object::Integer(-100), Object::Real(-16.8)]),
Operation::new("Tj", vec![Object::string_literal("Oak Street")]),
Operation::new("Td", vec![100.into(), 0.into()]),
Operation::new("Tj", vec![Object::string_literal("100")]),
Operation::new("Td", vec![Object::Integer(-100), Object::Real(-16.8)]),
Operation::new("Tj", vec![Object::string_literal("Boardwalk")]),
Operation::new("Td", vec![100.into(), 0.into()]),
Operation::new("Tj", vec![Object::string_literal("200")]),
Operation::new("ET", vec![]),
];
let content = Content { operations }.encode().unwrap();
doc.objects
.insert(content_id, Stream::new(dictionary! {}, content).into());
doc.objects.insert(
page_id,
dictionary! {
"Type" => "Page",
"Parent" => pages_id,
"MediaBox" => vec![0.into(), 0.into(), 200.into(), 800.into()],
"Resources" => dictionary! {
"Font" => dictionary! {
"F1" => font_id,
},
},
"Contents" => content_id,
}
.into(),
);
doc.objects.insert(
pages_id,
dictionary! {
"Type" => "Pages",
"Kids" => vec![page_id.into()],
"Count" => 1,
}
.into(),
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => pages_id,
});
doc.trailer.set("Root", catalog_id);
let mut bytes = Vec::new();
doc.save_to(&mut bytes).unwrap();
bytes
}
#[test]
fn test_extract_tables_with_structure_real_pdf_bits_pilani() {
use pdf_inspector::{extract_tables_with_structure_mem, TsrTableInput};
// Hand-crafted TSR fixture targeting page 4 (0-indexed=3) of
// bits_pilani_feedback.pdf, which contains a clean tabular layout.
//
// We construct a 2×2 table:
// row 0 (header): "Department" "Core Courses"
// row 1 (data): "BIO" "8.23"
//
// The PDF page is US Letter (792pt tall). We render at 72 dpi so
// image-px maps 1:1 to PDF-pt — that lets us write cell bboxes in
// the same units as our hand-measured page-pt coordinates.
let buf = std::fs::read("tests/fixtures/bits_pilani_feedback.pdf").unwrap();
// The PDF page is A4 in points (≈595.44 × 841.68). The table sits in
// the upper part of the page; we crop a window large enough to enclose
// both rows we care about.
//
// Crop bounds in PDF points (top-left origin):
// x: 80..280, y: 170..240
let crop = [80.0_f32, 170.0, 280.0, 240.0];
let dpi = 72.0_f32;
// Cell bboxes in CROP image-pixel space (= crop-relative PDF-pt at
// 72 dpi). The y ranges are tightened against neighbouring rows
// ("First Degree" above the header at native y=666.7, "Feedback Score"
// between the header and data rows at native y=640.9, "CE" below the
// BIO row at native y=591.1) so each cell only overlaps its target
// text item.
let cell_bboxes = vec![
// Header row: y crop-relative (7, 18) → page-pt y (177, 188)
poly(10.0, 7.0, 100.0, 18.0), // "Department" (item at page-pt x=107.1)
poly(110.0, 7.0, 200.0, 18.0), // "Core Courses" (item at page-pt x=199.0)
// Data row: y crop-relative (35, 60) → page-pt y (205, 230)
poly(10.0, 35.0, 100.0, 60.0), // "BIO" (item at page-pt x=104.1)
poly(110.0, 35.0, 200.0, 60.0), // "8.23" (item at page-pt x=221.2)
];
// Minimal SLANet-style token stream: a 2-row table with a thead and tbody.
let tokens: Vec<String> = [
"<table>",
"<thead>",
"<tr>",
"<th></th>",
"<th></th>",
"</tr>",
"</thead>",
"<tbody>",
"<tr>",
"<td></td>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(String::from)
.collect();
let inputs = vec![TsrTableInput {
page: 3,
crop_pdf_pt_bbox: crop,
render_dpi: dpi,
structure_tokens: tokens,
cell_bboxes,
}];
let mds = extract_tables_with_structure_mem(&buf, &inputs).unwrap();
assert_eq!(mds.len(), 1);
let md = &mds[0];
// Hand-written gold standard for the rendered markdown.
let expected = "|Department|Core Courses|\n|---|---|\n|BIO|8.23|\n";
assert_eq!(
md, expected,
"structured-table markdown should match the gold standard exactly\nactual: {md}"
);
}
#[test]
fn test_extract_tables_with_structure_dense_overlapping_slanet_boxes() {
use pdf_inspector::{extract_tables_with_structure_mem, TsrTableInput};
let buf = synthetic_dense_table_pdf();
let tokens: Vec<String> = [
"<table>",
"<thead>",
"<tr>",
"<th></th>",
"<th></th>",
"</tr>",
"</thead>",
"<tbody>",
"<tr>",
"<td></td>",
"<td></td>",
"</tr>",
"<tr>",
"<td></td>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(String::from)
.collect();
// Rows are spaced 16.8pt apart, while the SLANet-style boxes are 40pt
// tall and overlap adjacent rows. Text must still land in only one row.
let cell_bboxes = vec![
poly(10.0, 72.0, 100.0, 112.0),
poly(90.0, 72.0, 180.0, 112.0),
poly(10.0, 88.8, 100.0, 128.8),
poly(90.0, 88.8, 180.0, 128.8),
poly(10.0, 105.6, 100.0, 145.6),
poly(90.0, 105.6, 180.0, 145.6),
];
let mds = extract_tables_with_structure_mem(
&buf,
&[TsrTableInput {
page: 0,
crop_pdf_pt_bbox: [0.0, 0.0, 200.0, 800.0],
render_dpi: 72.0,
structure_tokens: tokens,
cell_bboxes,
}],
)
.unwrap();
let expected = "|Branch Name|Deposits|\n|---|---|\n|Oak Street|100|\n|Boardwalk|200|\n";
assert_eq!(mds[0], expected);
assert!(!mds[0].contains("Branch Name Oak Street"));
assert!(!mds[0].contains("Oak Street Boardwalk"));
}
#[test]
fn test_extract_tables_with_structure_input_order_preserved() {
use pdf_inspector::{extract_tables_with_structure_mem, TsrTableInput};
let buf = std::fs::read("tests/fixtures/bits_pilani_feedback.pdf").unwrap();
// Two inputs; both target the same page but with different shapes.
// We just need to confirm we get 2 outputs in the same order.
let make_input = |toks: Vec<&str>, cells: Vec<Vec<f32>>| TsrTableInput {
page: 3,
crop_pdf_pt_bbox: [80.0, 170.0, 280.0, 240.0],
render_dpi: 72.0,
structure_tokens: toks.into_iter().map(String::from).collect(),
cell_bboxes: cells,
};
let inputs = vec![
make_input(
vec!["<table>", "<tr>", "<td></td>", "</tr>", "</table>"],
vec![poly(10.0, 35.0, 100.0, 60.0)],
),
make_input(
vec!["<table>", "<tr>", "<td></td>", "</tr>", "</table>"],
vec![poly(110.0, 35.0, 200.0, 60.0)],
),
];
let mds = extract_tables_with_structure_mem(&buf, &inputs).unwrap();
assert_eq!(mds.len(), 2);
assert!(
mds[0].contains("BIO"),
"input 0 should pull 'BIO': {}",
mds[0]
);
assert!(
mds[1].contains("8.23"),
"input 1 should pull '8.23': {}",
mds[1]
);
}
#[test]
fn test_extract_tables_with_structure_out_of_range_page() {
use pdf_inspector::{extract_tables_with_structure_mem, TsrTableInput};
let buf = std::fs::read("tests/fixtures/bits_pilani_feedback.pdf").unwrap();
let inputs = vec![TsrTableInput {
page: 9999,
crop_pdf_pt_bbox: [0.0, 0.0, 100.0, 100.0],
render_dpi: 72.0,
structure_tokens: vec![
"<table>".into(),
"<tr>".into(),
"<td></td>".into(),
"</tr>".into(),
"</table>".into(),
],
cell_bboxes: vec![poly(0.0, 0.0, 50.0, 50.0)],
}];
let mds = extract_tables_with_structure_mem(&buf, &inputs).unwrap();
assert_eq!(mds.len(), 1);
assert!(
mds[0].is_empty(),
"out-of-range page should yield empty string"
);
}
#[test]
fn test_extract_tables_with_structure_not_a_pdf() {
use pdf_inspector::extract_tables_with_structure_mem;
let result = extract_tables_with_structure_mem(b"not a pdf", &[]);
assert!(result.is_err());
}
#[test]
fn test_extract_tables_with_structure_empty_inputs() {
use pdf_inspector::extract_tables_with_structure_mem;
let buf = std::fs::read("tests/fixtures/bits_pilani_feedback.pdf").unwrap();
let mds = extract_tables_with_structure_mem(&buf, &[]).unwrap();
assert!(mds.is_empty());
}
#[test]
fn test_extract_tables_with_structure_cells_real_pdf_bits_pilani() {
use pdf_inspector::{extract_tables_with_structure_cells_mem, TsrTableInput};
// Same fixture as test_extract_tables_with_structure_real_pdf_bits_pilani
// but exercising the cell-level API. Verifies that callers receive
// structured per-cell metadata (row/col/spans/is_header/page_pt_bbox)
// alongside the extracted text.
let buf = std::fs::read("tests/fixtures/bits_pilani_feedback.pdf").unwrap();
let crop = [80.0_f32, 170.0, 280.0, 240.0];
let dpi = 72.0_f32;
let cell_bboxes = vec![
poly(10.0, 7.0, 100.0, 18.0),
poly(110.0, 7.0, 200.0, 18.0),
poly(10.0, 35.0, 100.0, 60.0),
poly(110.0, 35.0, 200.0, 60.0),
];
let tokens: Vec<String> = [
"<table>",
"<thead>",
"<tr>",
"<th></th>",
"<th></th>",
"</tr>",
"</thead>",
"<tbody>",
"<tr>",
"<td></td>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(String::from)
.collect();
let inputs = vec![TsrTableInput {
page: 3,
crop_pdf_pt_bbox: crop,
render_dpi: dpi,
structure_tokens: tokens,
cell_bboxes,
}];
let cells_lists = extract_tables_with_structure_cells_mem(&buf, &inputs).unwrap();
assert_eq!(cells_lists.len(), 1);
let cells = &cells_lists[0];
assert_eq!(cells.len(), 4);
// Header row: both cells flagged as headers (they were in <thead>/<th>).
assert!(cells[0].is_header);
assert!(cells[1].is_header);
assert_eq!((cells[0].row, cells[0].col), (0, 0));
assert_eq!((cells[1].row, cells[1].col), (0, 1));
assert_eq!(cells[0].text, "Department");
assert_eq!(cells[1].text, "Core Courses");
// Data row: not flagged as header.
assert!(!cells[2].is_header);
assert!(!cells[3].is_header);
assert_eq!((cells[2].row, cells[2].col), (1, 0));
assert_eq!((cells[3].row, cells[3].col), (1, 1));
assert_eq!(cells[2].text, "BIO");
assert_eq!(cells[3].text, "8.23");
// Every cell carries a non-degenerate page-pt bbox.
for c in cells {
let [x1, y1, x2, y2] = c.page_pt_bbox;
assert!(
x1 < x2 && y1 < y2,
"cell bbox should be non-empty: {:?}",
c.page_pt_bbox
);
}
}
#[test]
fn test_extract_tables_with_structure_separator_after_thead() {
use pdf_inspector::{extract_tables_with_structure_mem, TsrTableInput};
// Re-run the same 2x2 fixture but assert exact markdown output: with
// <thead> + <th> headers, the separator should land after the header
// row (which is also row 0 here, so the gold-standard hasn't changed).
let buf = std::fs::read("tests/fixtures/bits_pilani_feedback.pdf").unwrap();
let crop = [80.0_f32, 170.0, 280.0, 240.0];
let dpi = 72.0_f32;
let cell_bboxes = vec![
poly(10.0, 7.0, 100.0, 18.0),
poly(110.0, 7.0, 200.0, 18.0),
poly(10.0, 35.0, 100.0, 60.0),
poly(110.0, 35.0, 200.0, 60.0),
];
let tokens: Vec<String> = [
"<table>",
"<thead>",
"<tr>",
"<th></th>",
"<th></th>",
"</tr>",
"</thead>",
"<tbody>",
"<tr>",
"<td></td>",
"<td></td>",
"</tr>",
"</tbody>",
"</table>",
]
.into_iter()
.map(String::from)
.collect();
let mds = extract_tables_with_structure_mem(
&buf,
&[TsrTableInput {
page: 3,
crop_pdf_pt_bbox: crop,
render_dpi: dpi,
structure_tokens: tokens,
cell_bboxes,
}],
)
.unwrap();
assert_eq!(mds.len(), 1);
assert_eq!(mds[0], "|Department|Core Courses|\n|---|---|\n|BIO|8.23|\n");
}
// =========================================================================
// extract_pages_markdown_mem tests
// =========================================================================
@@ -1481,7 +1915,7 @@ fn test_bits_pilani_page8_table_detection() {
#[test]
fn test_extract_pages_markdown_basic() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, &[0, 1]).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&[0, 1])).unwrap();
assert_eq!(result.pages.len(), 2);
assert_eq!(result.pages[0].page, 0);
@@ -1495,7 +1929,7 @@ fn test_extract_pages_markdown_basic() {
fn test_extract_pages_markdown_page_ordering() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
// Request pages in non-sequential order
let result = extract_pages_markdown_mem(&buf, &[1, 0]).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&[1, 0])).unwrap();
assert_eq!(result.pages.len(), 2);
// Results should match input order, not document order
@@ -1506,7 +1940,7 @@ fn test_extract_pages_markdown_page_ordering() {
#[test]
fn test_extract_pages_markdown_out_of_range() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, &[9999]).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&[9999])).unwrap();
assert_eq!(result.pages.len(), 1);
assert_eq!(result.pages[0].page, 9999);
@@ -1518,14 +1952,14 @@ fn test_extract_pages_markdown_out_of_range() {
#[test]
fn test_extract_pages_markdown_empty_pages_list() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, &[]).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&[])).unwrap();
assert!(result.pages.is_empty());
}
#[test]
fn test_extract_pages_markdown_single_page() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, &[0]).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&[0])).unwrap();
assert_eq!(result.pages.len(), 1);
assert_eq!(result.pages[0].page, 0);
@@ -1535,7 +1969,7 @@ fn test_extract_pages_markdown_single_page() {
#[test]
fn test_extract_pages_markdown_invalid_buffer() {
let result = extract_pages_markdown_mem(b"not a pdf", &[0]);
let result = extract_pages_markdown_mem(b"not a pdf", Some(&[0]));
assert!(result.is_err());
}
@@ -1543,7 +1977,7 @@ fn test_extract_pages_markdown_invalid_buffer() {
fn test_extract_pages_markdown_gid_pages_need_ocr() {
// shinagawa_identity_h.pdf has GID-encoded fonts
let buf = std::fs::read("tests/fixtures/shinagawa_identity_h.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, &[0]).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&[0])).unwrap();
assert_eq!(result.pages.len(), 1);
assert!(result.pages[0].needs_ocr);
@@ -1556,7 +1990,7 @@ fn test_extract_pages_markdown_classification_with_tables() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
let page_count = process_pdf_mem(&buf).unwrap().page_count;
let page_indices: Vec<u32> = (0..page_count).collect();
let result = extract_pages_markdown_mem(&buf, &page_indices).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&page_indices)).unwrap();
assert!(
!result.pages_with_tables.is_empty(),
@@ -1569,7 +2003,7 @@ fn test_extract_pages_markdown_classification_with_tables() {
fn test_extract_pages_markdown_simple_pdf_no_complexity() {
// bare_name_struct.pdf is a simple document with a heading and code block
let buf = std::fs::read("tests/fixtures/bare_name_struct.pdf").unwrap();
let result = extract_pages_markdown_mem(&buf, &[0]).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&[0])).unwrap();
assert!(result.pages_with_tables.is_empty());
assert!(result.pages_with_columns.is_empty());
@@ -1582,7 +2016,7 @@ fn test_extract_pages_markdown_classification_matches_process_pdf() {
let full = process_pdf_mem(&buf).unwrap();
let page_count = full.page_count;
let page_indices: Vec<u32> = (0..page_count).collect();
let result = extract_pages_markdown_mem(&buf, &page_indices).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&page_indices)).unwrap();
assert_eq!(
result.pages_with_tables, full.layout.pages_with_tables,
@@ -1605,7 +2039,7 @@ fn test_extract_pages_markdown_consistency_with_process_pdf() {
// Get per-page output for all pages
let page_count = full.page_count;
let page_indices: Vec<u32> = (0..page_count).collect();
let result = extract_pages_markdown_mem(&buf, &page_indices).unwrap();
let result = extract_pages_markdown_mem(&buf, Some(&page_indices)).unwrap();
// Concatenated per-page markdown should contain substantial overlap with
// the full output (exact match not expected due to header/footer stripping
@@ -1630,3 +2064,41 @@ fn test_extract_pages_markdown_consistency_with_process_pdf() {
full_md.len()
);
}
#[test]
fn test_extract_pages_markdown_none_returns_all_pages() {
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
let page_count = process_pdf_mem(&buf).unwrap().page_count;
let result = extract_pages_markdown_mem(&buf, None).unwrap();
assert_eq!(result.pages.len() as u32, page_count);
for (i, page) in result.pages.iter().enumerate() {
assert_eq!(page.page, i as u32, "pages should be in document order");
}
}
#[test]
fn test_extract_pages_markdown_path_api() {
let path = "tests/fixtures/nexo-price-en.pdf";
let buf = std::fs::read(path).unwrap();
let via_path = extract_pages_markdown(path, Some(&[0])).unwrap();
let via_mem = extract_pages_markdown_mem(&buf, Some(&[0])).unwrap();
assert_eq!(via_path.pages.len(), via_mem.pages.len());
assert_eq!(via_path.pages[0].markdown, via_mem.pages[0].markdown);
assert_eq!(via_path.pages[0].needs_ocr, via_mem.pages[0].needs_ocr);
assert_eq!(via_path.is_complex, via_mem.is_complex);
}
#[test]
fn test_extract_pages_markdown_path_none_returns_all_pages() {
let path = "tests/fixtures/nexo-price-en.pdf";
let page_count = process_pdf_mem(&std::fs::read(path).unwrap())
.unwrap()
.page_count;
let result = extract_pages_markdown(path, None).unwrap();
assert_eq!(result.pages.len() as u32, page_count);
}
+4 -4
View File
@@ -1,9 +1,9 @@
본 가격표는 국내 거주 중인 외국인을 위한 한국어 가격표의 비공식 번역본입니다. ※ The post-tax benefit sales price is provided for your reference only, reflecting the current tax benefits and eco-friendly vehicle individual consumption tax reductions. 본 가격표와 한국어 가격표의 내용이 상이한 경우 한국어 가격표의 내용이 우선하므로, 반드시 한국어 가격표의 내용을 확인하십시오. The final sales price may vary depending on the addition of optional items and whether the eco-friendly vehicle criteria are met, so please be sure to check the quotation. This price list is an unofficial translation of the Korean price list for the convenience of foreign residents in South Korea. ※ Please check the Korean price list for information on colors, details, and fuel consumption for each model. If the price list differs from the Korean price list, please check the contents of the Korean price list first. ※ All optional item prices are listed based on pre-tax reduction amounts. The actual sales price, which reflects the total individual consumption tax reduction including optional items, may differ depending on applicable tax benefits. ※ The items (specifications, colors, etc.) and prices listed in this pricing table are subject to change without prior notice depending on the holding of new car launch events, improvements made in automobile performance, introduction of related laws and regulations, and changes in company circumstances. The all-new NEXO Release Date: June 10, 2025 / (Unit: KRW)
|Classification Exclusive Exclusive|Selling price before tax benefit Supply value(surtax) 80,509,000 73,190,000(7,319,000)|Selling price after tax benefit 76,435,000|Standard equipment • Powertrain/Performance: Fuel cell system(150kW drive motor, lithium-ion battery, and reducer), Regenerative braking system, Column-Type Shift By Wire(vibration warning), Drive mode select • Safety: 9 airbag system(1st-row advanced/center side airbags, 1st/2nd-row side airbags, and rollover-resistant curtain airbags), Multi-Collision Brake System, Active hood system(for pedestrian protection), Safety unlock function, Artificial engine sound(for pedestrian protection), Child seat fastening system (2 in 2nd-row), Fire extinguisher for vehicles, Pedal Misapplication Safety Assist • Smart Safety Technology: Forward Collision-avoidance Assist(vehicles/ pedestrians/two-wheeled vehicles/junction turning/front oncoming), Smart Cruise Control with Stop & Go, Lane Keeping Assist, Lane Following Assist 2, Blind-spot Collision Warning(driving), Blind-spot Collision-avoidance Assist(forward exit), Rear Cross-traffic Collision-avoidance Assist, Safety Exit Assist, Driver Attention Warning, High Beam Assist, Advanced Rear Occupant Alert, Intelligent Speed Limit Assist, Hands-On Detection, Highway Driving Assist, Navigation-based Smart Cruise Control(safety speed zone/curve control), Vibration warning steering wheel • Exterior: Full LED headlamps(projection type), LED turn signal lamps (front and rear), LED Daytime Running Lights, LED positioning lights, LED rear combination lamps, LED third brake lights, 18-inch alloy wheels & tires, Solar glass(windshield), Double-glazed soundproof glass(windshield, and 1st/ 2nd-row doors), Outside mirror(heating, power-folding, power adjustment,|Options (before tax benefit) ▶ Hi-pass(e hi-pass) [200,000]|
|Classification|Selling price before tax benefit Supply value(surtax)|Selling price after tax benefit|Standard equipment|Options (before tax benefit)|
|---|---|---|---|---|
|Special|with 3.5% individual consumption tax applied 79,287,000 83,500,000 75,909,091(7,590,909) with 3.5% individual consumption tax applied 82,232,000|with 3.5% individual consumption tax applied 76,435,000 79,275,000 with 3.5% individual consumption tax applied 79,275,000|and LED turn signal lamps), Auto flush door handles, Black door garnish • Interior: Panoramic curved display, 12.3-inch color LCD cluster, Leather- upholstered steering wheel(with heating, two-tone color, and Interactive Pixel Lights), LED interior lamp (map lamp, personal lamp, sun visor lamp, and luggage lamp), Metallic door scuff plate • Seat: Synthetic leather seats, 1st-row manual seats, Heated 1st-row seats, 2nd-row 60/40-split folding seats(reclining) • Convenience: Proximity key with push-button start, Smart key remote start, Electronic Parking Brake(with automatic vehicle hold), Paddle shift(regenerative control), Dual-zone full automatic air conditioning(with high-performance antibacterial combination filter, auto defog system, fine dust sensor, air cleaning mode, and after-blow function), 2nd-row seat air vent, Auto light control system, USB Type-C Ports(1×27W switchable charging/data port in 1st-row, and 2×100W charging ports in both 1st and 2nd-row), ECM room mirror(frameless), Rain sensor, Power windows with pinch protection(1st/2nd-row), Power outlet (1 in 1st-row), Parking Distance Warning-Forward/Reverse, Rear View Monitor, Wireless phone charger(single), Walk-away lock, Route planner, Hyundai AI Assistant • Infotainment: 12.3-inch navigation(Bluelink, phone projection, Bluetooth hands-free, and In-car Payment), Audio system(6 speakers), Over-The-Air navigation updates Standard equipment of Exclusive plus • Smart Safety Technology: Forward Collision-avoidance Assist(intersection crossing/changing lanes in oncoming traffic/approaching from either side/ evasive steering assist), Highway Driving Assist 2, Navigation-based Smart Cruise Control(access road) • Exterior: Roof rack • Interior: Metallic pedal, Driving mode-dependent ambient mood lighting(crash pad, 1st/2nd-row door trim) • Seat: Synthetic leather seats(patch applied), Power-adjustable driver's|▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [950,000] Parking Assist ▶ [1,150,000] Audio by BANG & OLUFSEN|
|Prestige|87,893,000 79,902,727(7,990,273) with 3.5% individual consumption tax applied 86,559,000|83,445,000 with 3.5% individual consumption tax applied 83,445,000|seat(8-way, lumbar support, and Integrated Memory System(driver's seat and outside mirror connected)), Power-adjustable front passengers seat(8-way), Ventilated 1st-row seats, Heated 2nd-row seats • Convenience: Hi-pass(e hi-pass), In-car fingerprint authentication system(personalization, startup, payment, and etc.), Smart power tailgate ▶ Standard equipment of Exclusive Special plus • Smart Safety Technology: Remote Smart Parking Assist 2, Parking Collison- avoidance Assist(front/side/rear) • Exterior: Intelligent Front-Lighting System(IFS), Dynamic welcome/escort lighting(1 type), Sequential turn signals(front and rear), Ambient lighting auto flush door handles, Two-tone door garnish, Glossy black rear diffuserInterior: Recycled PET suede interior materials(headlining/sunvisor), Fabric upholstered crash pad • Seat: BIO-processed natural leather seats(metal patch applied, embossed design punching), Passenger's seat walk-in device, 1st-row relaxation comfort seats(leg rest included), Ventilated 2nd-row seats • Convenience: Parking Distance Warning-Side, Head-Up Display, Digital key 2, Wireless phone charger(dual), Surround View Monitor, Blind-spot View Monitor, LED reverse light guide • Infotainment: Audio by BANG & OLUFSEN sound system(14 speakers, including external amp), Active road noise control, Active Sound Design|sound system ▶ [250,000] 19-inch alloy wheels & tires ▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [900,000] Vision roof ▶ [1,380,000] Digital side mirror ▶ [750,000] Camera package ▶ [250,000] 19-inch alloy wheels & tires|
|Exclusive|80,509,000 73,190,000(7,319,000) with 3.5% individual consumption tax applied 79,287,000|76,435,000 with 3.5% individual consumption tax applied 76,435,000|• Powertrain/Performance: Fuel cell system(150kW drive motor, lithium-ion battery, and reducer), Regenerative braking system, Column-Type Shift By Wire(vibration warning), Drive mode select • Safety: 9 airbag system(1st-row advanced/center side airbags, 1st/2nd-row side airbags, and rollover-resistant curtain airbags), Multi-Collision Brake System, Active hood system(for pedestrian protection), Safety unlock function, Artificial engine sound(for pedestrian protection), Child seat fastening system (2 in 2nd-row), Fire extinguisher for vehicles, Pedal Misapplication Safety Assist • Smart Safety Technology: Forward Collision-avoidance Assist(vehicles/ pedestrians/two-wheeled vehicles/junction turning/front oncoming), Smart Cruise Control with Stop & Go, Lane Keeping Assist, Lane Following Assist 2, Blind-spot Collision Warning(driving), Blind-spot Collision-avoidance Assist(forward exit), Rear Cross-traffic Collision-avoidance Assist, Safety Exit Assist, Driver Attention Warning, High Beam Assist, Advanced Rear Occupant Alert, Intelligent Speed Limit Assist, Hands-On Detection, Highway Driving Assist, Navigation-based Smart Cruise Control(safety speed zone/curve control), Vibration warning steering wheel • Exterior: Full LED headlamps(projection type), LED turn signal lamps (front and rear), LED Daytime Running Lights, LED positioning lights, LED rear combination lamps, LED third brake lights, 18-inch alloy wheels & tires, Solar glass(windshield), Double-glazed soundproof glass(windshield, and 1st/ 2nd-row doors), Outside mirror(heating, power-folding, power adjustment, and LED turn signal lamps), Auto flush door handles, Black door garnish • Interior: Panoramic curved display, 12.3-inch color LCD cluster, Leather- upholstered steering wheel(with heating, two-tone color, and Interactive Pixel Lights), LED interior lamp (map lamp, personal lamp, sun visor lamp, and luggage lamp), Metallic door scuff plate • Seat: Synthetic leather seats, 1st-row manual seats, Heated 1st-row seats, 2nd-row 60/40-split folding seats(reclining) • Convenience: Proximity key with push-button start, Smart key remote start, Electronic Parking Brake(with automatic vehicle hold), Paddle shift(regenerative control), Dual-zone full automatic air conditioning(with high-performance antibacterial combination filter, auto defog system, fine dust sensor, air cleaning mode, and after-blow function), 2nd-row seat air vent, Auto light control system, USB Type-C Ports(1×27W switchable charging/data port in 1st-row, and 2×100W charging ports in both 1st and 2nd-row), ECM room mirror(frameless), Rain sensor, Power windows with pinch protection(1st/2nd-row), Power outlet (1 in 1st-row), Parking Distance Warning-Forward/Reverse, Rear View Monitor, Wireless phone charger(single), Walk-away lock, Route planner, Hyundai AI Assistant • Infotainment: 12.3-inch navigation(Bluelink, phone projection, Bluetooth hands-free, and In-car Payment), Audio system(6 speakers), Over-The-Air navigation updates|Hi-pass(e hi-pass) [200,000]|
|Exclusive Special|83,500,000 75,909,091(7,590,909) with 3.5% individual consumption tax applied 82,232,000|79,275,000 with 3.5% individual consumption tax applied 79,275,000|▶ Standard equipment of Exclusive plus • Smart Safety Technology: Forward Collision-avoidance Assist(intersection crossing/changing lanes in oncoming traffic/approaching from either side/ evasive steering assist), Highway Driving Assist 2, Navigation-based Smart Cruise Control(access road)Exterior: Roof rack • Interior: Metallic pedal, Driving mode-dependent ambient mood lighting(crash pad, 1st/2nd-row door trim) • Seat: Synthetic leather seats(patch applied), Power-adjustable driver's seat(8-way, lumbar support, and Integrated Memory System(driver's seat and outside mirror connected)), Power-adjustable front passengers seat(8-way), Ventilated 1st-row seats, Heated 2nd-row seats • Convenience: Hi-pass(e hi-pass), In-car fingerprint authentication system(personalization, startup, payment, and etc.), Smart power tailgate|▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [950,000] Parking Assist ▶ [1,150,000] Audio by BANG & OLUFSEN sound system ▶ [250,000] 19-inch alloy wheels & tires|
|Prestige|87,893,000 79,902,727(7,990,273) with 3.5% individual consumption tax applied 86,559,000|83,445,000 with 3.5% individual consumption tax applied 83,445,000|▶ Standard equipment of Exclusive Special plus • Smart Safety Technology: Remote Smart Parking Assist 2, Parking Collison- avoidance Assist(front/side/rear) • Exterior: Intelligent Front-Lighting System(IFS), Dynamic welcome/escort lighting(1 type), Sequential turn signals(front and rear), Ambient lighting auto flush door handles, Two-tone door garnish, Glossy black rear diffuser • Interior: Recycled PET suede interior materials(headlining/sunvisor), Fabric upholstered crash pad • Seat: BIO-processed natural leather seats(metal patch applied, embossed design punching), Passenger's seat walk-in device, 1st-row relaxation comfort seats(leg rest included), Ventilated 2nd-row seats • Convenience: Parking Distance Warning-Side, Head-Up Display, Digital key 2, Wireless phone charger(dual), Surround View Monitor, Blind-spot View Monitor, LED reverse light guide • Infotainment: Audio by BANG & OLUFSEN sound system(14 speakers, including external amp), Active road noise control, Active Sound Design|▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [900,000] Vision roof ▶ [1,380,000] Digital side mirror ▶ [750,000] Camera package ▶ [250,000] 19-inch alloy wheels & tires|
**Classification Details** **Indoor/outdoor V2L** Indoor V2L, Outdoor V2L(connectorless type) **Parking Assist** Surround View Monitor, Blind-spot View Monitor, Parking Distance Warning-Side, Parking Collison-avoidance Assist-Rear **Audio by BANG & OLUFSEN** Audio by BANG & OLUFSEN sound system(14 speakers, including external amp.), Active road noise control, Active Sound Design **sound system** **Camera package** Digital center mirror(with camera sensor cleaning system), Driver monitoring system THE ALL-NEW NEXO /// ECO-FRIENDLY CAR
+47 -40
View File
@@ -13,8 +13,10 @@ IDENTIFICATION NUMBER OF CORPORATION] ELECTS TO BE TREATED AS A COMPONENT MEMBER
(v) Election-- (A) Election filed. An election filed under paragraph (d)(2)(iv) of
this section is irrevocable and effective until paragraph (d)(2)(ii) or (iii) of §1.1563-3 applies or until a change in the stock ownership of the corporation results in
|termination of membership in the controlled group in which such corporation has been included. (B) Election not filed.|In the event no election is filed in accordance with the|
|termination of membership in the controlled group in which such corporation has||
|---|---|
|been included.||
|(B) Election not filed.|In the event no election is filed in accordance with the|
|provisions of paragraph (d)(2)(iv) of this section, then the Internal Revenue Service||
|will determine the group in which such corporation is to be included. Such||
|determination will be binding for all subsequent years unless the corporation files a||
@@ -101,45 +103,32 @@ section and paragraph
(c)(4), and (c)(5) of this section, and paragraph
(c)(2) of §1.382-8T
||§1.382-8(g), Example|
|---|---|
||The first sentence of §1.382-8(g), Example §1.382-8(g), Example §1.382-8(g), Example|
(2)(c)
(2)(e)
(3)(b)
(3)(c)(1)(B)
||The second sentence of|
|---|---|
||§1.382-8(g), Example The second sentence of §1.382-8(g), Example The first sentence of §1.1502-32(b)(4)(v)(A) The first sentence of §1.1502-32(b)(4)(v)(B)|
(4)(c)
(5)(c)
|The fifth sentence of|paragraph (c) of this|paragraphs (c)(1), (c)(3),|
|---|---|---|
|§1.382-8(f)|section|(c)(4), and (c)(5) of this section, and paragraph (c)(2) of §1.382-8T|
|§1.382-8(g), Example|paragraph (c) of this|paragraphs (c)(1), (c)(3), section, and paragraph (c)(2) of §1.382-8T|
|The second sentence of|paragraph (c) of this|paragraphs (c)(1), (c)(3),|
|§1.382-8(g), Example|section|(c)(4), and (c)(5) of this|
|§1.382-8(g), Example|section paragraph (c)(2) of this section paragraph (c)(2) of this section paragraph (c)(2) of this section paragraphs (c)(1) and (2) of this section paragraph (c)(2) of this section paragraph (c)(2) of this section paragraph (b)(4)(iv) of this section paragraph (b)(4)(iv) of this section|(c)(4), and (c)(5) of this (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (c)(1) of this section and paragraph (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (c)(2) of §1.382-8T paragraph (b)(4)(iv) of §1.1502-32T paragraph (b)(4)(iv) of §1.1502-32T|
(1)(b)(2) section (c)(4), and (c)(5) of this
(1)(c) section, and paragraph
(c)(2) of §1.382-8T
|§1.382-8(g), Example|paragraph (c)(2) of this|paragraph (c)(2) of|
|---|---|---|
|The first sentence of|paragraph (c)(2) of this|paragraph (c)(2) of|
|§1.382-8(g), Example|section|§1.382-8T|
|§1.382-8(g), Example|paragraph (c)(2) of this|paragraph (c)(2) of|
|§1.382-8(g), Example|paragraphs (c)(1) and (2)|paragraph (c)(1) of this|
(2)(c) section §1.382-8T
(2)(e)
(3)(b) section §1.382-8T
(3)(c)(1)(B) of this section section and paragraph
(c)(2) of §1.382-8T
|The second sentence of|paragraph (c)(2) of this|paragraph (c)(2) of|
|---|---|---|
|§1.382-8(g), Example|section|§1.382-8T|
|The second sentence of|paragraph (c)(2) of this|paragraph (c)(2) of|
|§1.382-8(g), Example|section|§1.382-8T|
|The first sentence of|paragraph (b)(4)(iv) of|paragraph (b)(4)(iv) of|
|§1.1502-32(b)(4)(v)(A)|this section|§1.1502-32T|
|The first sentence of|paragraph (b)(4)(iv) of|paragraph (b)(4)(iv) of|
|§1.1502-32(b)(4)(v)(B)|this section|§1.1502-32T|
(4)(c)
(5)(c)
|§1.1502-35(c)(4)(ii)(B)|§1.1502-76(b)(2)(ii)(D)|§1.1502-76T(b)(2)(ii)(D)|
|---|---|---|
|§1.1502-76(b)(2)(ii)(A)(2)|paragraph (b)(2)(ii)(D) of this section|paragraph (b)(2)(ii)(D) of §1.1502-76T|
@@ -168,14 +157,33 @@ section and paragraph
|§1.6043-2(a)|or 1.1081-11|3T(a), or §1.1081-11T|
|The first sentence of §301.6011-5T(a) (twice)|§1.6012-2|paragraphs (a), (b) and (d) through (j) of §1.6012- 2, and paragraph (c) of §1.6012-2T|
|||PART 602--OMB CONTROL NUMBERS UNDER THE PAPERWORK||
|---|---|---|---|
||REDUCTION ACT Authority: 26 U.S.C. 7805. 1. The following entries to the table are removed: §602.101 OMB Control numbers.|Par. 54. The authority citation for part 602 continues to read as follows: Par. 55. In §602.101, paragraph (b) is amended to read as follows:||
|* * * * *|(b) * * * CFR part or section where identified or described||Current OMB control No.|
|* * * * *|1.332-6………………………………………………………………….|1.382-11……………………………………………………………….. 1545-2019 1.351-3…………………………………………………………………. 1545-2019 1.355-5…………………………………………………………………. 1545-2019 1.368-3…………………………………………………………………. 1545-2019 1.1081-11………………………………………………………………. 1545-2019|1545-2019|
|* * * * *|§602.101 OMB Control numbers.|______________________________________________________________ 2. The following entries are added in numerical order to the table:||
|* * * * *|(b) * * * CFR part or section where identified or described||Current OMB control No.|
|* * * * *|1.302-2T………………………………………………………………… 1545 1.302-4T………………………………………………………………… 1545||-2019 -2019|
PART 602--OMB CONTROL NUMBERS UNDER THE PAPERWORK REDUCTION ACT Par. 54. The authority citation for part 602 continues to read as follows: Authority: 26 U.S.C. 7805. Par. 55. In §602.101, paragraph (b) is amended to read as follows:
1. The following entries to the table are removed:
§602.101 OMB Control numbers.
* * * * *
(b) * * *
CFR part or section where Current OMB identified or described control No.
* * * * *
1.332-6…………………………………………………………………. 1545-2019
1.382-11……………………………………………………………….. 1545-2019
1.351-3…………………………………………………………………. 1545-2019
1.355-5…………………………………………………………………. 1545-2019
1.368-3…………………………………………………………………. 1545-2019
1.1081-11………………………………………………………………. 1545-2019
* * * * * **______________________________________________________________**
2. The following entries are added in numerical order to the table:
§602.101 OMB Control numbers.
* * * * *
(b) * * *
CFR part or section where Current OMB identified or described control No.
* * * * *
1.302-2T………………………………………………………………… 1545-2019
1.302-4T………………………………………………………………… 1545-2019
|1.331-1T………………………………………………………………… 1545|-2019|
|---|---|
@@ -193,4 +201,3 @@ section and paragraph
Deputy Commissioner for Services and Enforcement.
Approved: May 19, 2006 Eric Solomon Acting Deputy Assistant Secretary of the Treasury (Tax Policy).
+72
View File
@@ -270,6 +270,78 @@ class TestExtractTextInRegions:
)
# ---------------------------------------------------------------------------
# extract_pages_markdown / extract_pages_markdown_bytes
# ---------------------------------------------------------------------------
class TestExtractPagesMarkdown:
def test_default_returns_all_pages(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf")
)
assert len(result.pages) == 3
assert [p.page for p in result.pages] == [0, 1, 2]
assert all(isinstance(p.markdown, str) for p in result.pages)
def test_bytes_default_returns_all_pages(self):
data = fixture_bytes("thermo-freon12.pdf")
result = pdf_inspector.extract_pages_markdown_bytes(data)
assert len(result.pages) == 3
def test_selected_pages_preserve_order(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf"), pages=[2, 0]
)
assert [p.page for p in result.pages] == [2, 0]
def test_bytes_selected_pages_preserve_order(self):
data = fixture_bytes("thermo-freon12.pdf")
result = pdf_inspector.extract_pages_markdown_bytes(data, pages=[1])
assert len(result.pages) == 1
assert result.pages[0].page == 1
def test_page_fields(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf"), pages=[0]
)
page = result.pages[0]
assert isinstance(page.page, int)
assert isinstance(page.markdown, str)
assert isinstance(page.needs_ocr, bool)
assert not page.needs_ocr # text-based fixture
assert len(page.markdown) > 0
def test_result_fields(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf")
)
assert isinstance(result.pages, list)
assert isinstance(result.pages_with_tables, list)
assert isinstance(result.pages_with_columns, list)
assert isinstance(result.pages_needing_ocr, list)
assert isinstance(result.is_complex, bool)
def test_out_of_range_page_marks_needs_ocr(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf"), pages=[9999]
)
assert len(result.pages) == 1
assert result.pages[0].needs_ocr
assert result.pages[0].markdown == ""
def test_repr(self):
result = pdf_inspector.extract_pages_markdown(
fixture_path("thermo-freon12.pdf"), pages=[0]
)
assert "PagesExtractionResult" in repr(result)
assert "PageMarkdown" in repr(result.pages[0])
def test_not_a_pdf(self):
with pytest.raises(ValueError):
pdf_inspector.extract_pages_markdown_bytes(b"not a pdf")
# ---------------------------------------------------------------------------
# Error handling
# ---------------------------------------------------------------------------