Compare commits
34
Commits
@@ -2,14 +2,41 @@ name: Publish npm package
|
||||
|
||||
on:
|
||||
push:
|
||||
tags: ['v*']
|
||||
branches: [main]
|
||||
paths: ['napi/package.json']
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
id-token: write
|
||||
|
||||
jobs:
|
||||
check-version:
|
||||
name: Check version change
|
||||
runs-on: ubuntu-latest
|
||||
outputs:
|
||||
changed: ${{ steps.check.outputs.changed }}
|
||||
version: ${{ steps.check.outputs.version }}
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check if version changed
|
||||
id: check
|
||||
run: |
|
||||
NEW_VERSION=$(node -p "require('./napi/package.json').version")
|
||||
OLD_VERSION=$(git show HEAD~1:napi/package.json | node -p "JSON.parse(require('fs').readFileSync('/dev/stdin','utf8')).version")
|
||||
echo "old=$OLD_VERSION new=$NEW_VERSION"
|
||||
if [ "$NEW_VERSION" != "$OLD_VERSION" ]; then
|
||||
echo "changed=true" >> "$GITHUB_OUTPUT"
|
||||
echo "version=$NEW_VERSION" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "changed=false" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
build:
|
||||
needs: check-version
|
||||
if: needs.check-version.outputs.changed == 'true'
|
||||
name: Build ${{ matrix.target }}
|
||||
runs-on: ${{ matrix.os }}
|
||||
strategy:
|
||||
@@ -19,6 +46,8 @@ jobs:
|
||||
target: x86_64-unknown-linux-gnu
|
||||
- os: macos-14
|
||||
target: aarch64-apple-darwin
|
||||
- os: windows-latest
|
||||
target: x86_64-pc-windows-msvc
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
@@ -68,7 +97,7 @@ jobs:
|
||||
|
||||
publish:
|
||||
name: Publish to npm
|
||||
needs: build
|
||||
needs: [check-version, build]
|
||||
runs-on: ubuntu-latest
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
@@ -0,0 +1,79 @@
|
||||
# pdf-inspector
|
||||
|
||||
Fast PDF text extraction to structured Markdown. CLI binary: `pdf2md`. Detection binary: `detect-pdf`.
|
||||
|
||||
## Build & Test
|
||||
|
||||
```bash
|
||||
cargo fmt # format
|
||||
cargo clippy -- -D warnings # lint (enforced, zero warnings)
|
||||
cargo test # unit + integration tests (267+ unit, 73+ integration)
|
||||
cargo build --release # release binary for benchmarks
|
||||
```
|
||||
|
||||
All three must pass before committing.
|
||||
|
||||
## Binaries
|
||||
|
||||
- `pdf2md` — extract PDF → Markdown. Supports `--json` for structured output.
|
||||
- `detect-pdf` — classify PDF type (TextBased/Scanned/Mixed/ImageBased). Supports `--analyze --json`.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
src/
|
||||
lib.rs – public API, process_pdf_with_options, encoding issue detection
|
||||
detector.rs – PDF type classification, tiled-scan detection, page sampling
|
||||
types.rs – TextItem, TextLine, PdfRect, PdfLine
|
||||
tounicode.rs – CMap/ToUnicode parsing, CID decoding
|
||||
text_utils.rs – CJK/RTL handling, Otsu threshold, ligature expansion, NFKC
|
||||
extractor/
|
||||
mod.rs – top-level extraction orchestrator
|
||||
content_stream.rs – PDF operator state machine (Tj/TJ/Td/Tm/q/Q)
|
||||
fonts.rs – font width/encoding, CMapDecisionCache, TrueType cmap fallback
|
||||
layout.rs – column detection (histogram), newspaper/tabular classification,
|
||||
spanning-line pre-masking, sidebar detection
|
||||
tables/
|
||||
detect_rects.rs – rect-based table detection (union-find clustering)
|
||||
detect_heuristic.rs – heuristic table detection (gap-histogram, body-font tables)
|
||||
detect_lines.rs – line-based table detection (H/V line grids)
|
||||
grid.rs – column/row boundaries, cell assignment
|
||||
format.rs – table→Markdown formatting, continuation row merging
|
||||
markdown/
|
||||
convert.rs – core line→Markdown loop, struct-tree role support
|
||||
analysis.rs – font stats, heading tiers, paragraph thresholds
|
||||
classify.rs – line classification (header, list, code, caption)
|
||||
preprocess.rs – drop cap merging, heading line merging
|
||||
postprocess.rs – dot leaders, hyphenation, page numbers, URL formatting
|
||||
```
|
||||
|
||||
## Key design decisions
|
||||
|
||||
- **Primary audience is AI agents.** Output optimized for token efficiency and semantic quality, not visual formatting. No cosmetic padding.
|
||||
- **Three table detection strategies** run in priority order: rect-based → line-based → heuristic. First valid result wins.
|
||||
- **Column detection** uses horizontal projection histograms with valley detection. Multi-item spanning lines (titles, headers) are pre-masked using column-aware thresholds before column assignment.
|
||||
- **Newspaper vs tabular** classification determines reading order: newspaper reads columns sequentially, tabular Y-interleaves them.
|
||||
- **Tiled-scan detection** catches scanned PDFs with JBIG2/strip images where no single tile exceeds the template threshold but aggregate area does (≥2M pixels).
|
||||
- **Garbage text upgrade** reclassifies Mixed PDFs as Scanned when extracted text is <50% alphanumeric.
|
||||
- **Tagged PDF support** uses structure tree roles (H1-H6, P, L, Code, BlockQuote) when available, falling back to font-size heuristics.
|
||||
|
||||
## Testing
|
||||
|
||||
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
|
||||
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
|
||||
- **Regression suite**: sibling repo `pdf-evals` with 179+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
|
||||
|
||||
## Debugging
|
||||
|
||||
```bash
|
||||
RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- file.pdf
|
||||
RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- file.pdf
|
||||
RUST_LOG=pdf_inspector::detector=debug cargo run --release --bin detect-pdf -- file.pdf
|
||||
```
|
||||
|
||||
## Conventions
|
||||
|
||||
- Clippy: use `is_some_and(...)` not `map_or(false, ...)`
|
||||
- lopdf quirk: `ParseError` is private — match by string for `InvalidFileHeader`
|
||||
- Column limit for tables: 25 (wide statistical tables)
|
||||
- `propagate_merged_cells` skipped for >10 columns (spanning rects = background fills)
|
||||
@@ -22,7 +22,7 @@ Evaluated on the [opendataloader-bench](https://github.com/opendataloader-projec
|
||||
|
||||
| Engine | Overall | Reading Order (NID) | Tables (TEDS) | Headings (MHS) | Speed (200 docs) |
|
||||
|---|---|---|---|---|---|
|
||||
| pdf-inspector | 0.77 | 0.87 | 0.52 | 0.58 | 4s |
|
||||
| pdf-inspector | 0.78 | 0.87 | 0.59 | 0.57 | 4s |
|
||||
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
|
||||
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
|
||||
| markitdown | 0.58 | 0.88 | 0.00 | 0.00 | 8s |
|
||||
@@ -55,12 +55,12 @@ print(result.markdown) # Markdown string or None
|
||||
### Node.js
|
||||
|
||||
```bash
|
||||
npm install @firecrawl/pdf-inspector-js
|
||||
npm install @firecrawl/pdf-inspector
|
||||
```
|
||||
|
||||
```javascript
|
||||
import { readFileSync } from 'fs';
|
||||
import { processPdf, classifyPdf } from '@firecrawl/pdf-inspector-js';
|
||||
import { processPdf, classifyPdf } from '@firecrawl/pdf-inspector';
|
||||
|
||||
const result = processPdf(readFileSync('document.pdf'));
|
||||
console.log(result.pdfType); // "TextBased", "Scanned", "ImageBased", "Mixed"
|
||||
|
||||
Generated
+1
-19
@@ -129,12 +129,6 @@ version = "3.20.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "5d20789868f4b01b2f2caec9f5c4e0213b41e3e5702a50157d699ae31ced2fcb"
|
||||
|
||||
[[package]]
|
||||
name = "bytecount"
|
||||
version = "0.6.9"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "175812e0be2bccb6abe50bb8d566126198344f707e304f45c648fd8f2cc0365e"
|
||||
|
||||
[[package]]
|
||||
name = "cbc"
|
||||
version = "0.1.2"
|
||||
@@ -679,7 +673,7 @@ checksum = "5e5032e24019045c762d3c0f28f5b6b8bbf38563a65908389bf7978758920897"
|
||||
[[package]]
|
||||
name = "lopdf"
|
||||
version = "0.40.0"
|
||||
source = "git+https://github.com/J-F-Liu/lopdf?rev=052674053814a9f4897af94f0b8e46a545c9b329#052674053814a9f4897af94f0b8e46a545c9b329"
|
||||
source = "git+https://github.com/J-F-Liu/lopdf?rev=7a05512d831415b1f2b1ce522391d6beab8a1284#7a05512d831415b1f2b1ce522391d6beab8a1284"
|
||||
dependencies = [
|
||||
"aes",
|
||||
"bitflags",
|
||||
@@ -695,7 +689,6 @@ dependencies = [
|
||||
"log",
|
||||
"md-5",
|
||||
"nom",
|
||||
"nom_locate",
|
||||
"rand",
|
||||
"rangemap",
|
||||
"rayon",
|
||||
@@ -807,17 +800,6 @@ dependencies = [
|
||||
"memchr",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "nom_locate"
|
||||
version = "5.0.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "0b577e2d69827c4740cba2b52efaad1c4cc7c73042860b199710b3575c68438d"
|
||||
dependencies = [
|
||||
"bytecount",
|
||||
"memchr",
|
||||
"nom",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "num-conv"
|
||||
version = "0.2.1"
|
||||
|
||||
+5
-5
@@ -1,4 +1,4 @@
|
||||
# firecrawl-pdf-inspector
|
||||
# PDF Inspector
|
||||
|
||||
Fast PDF classification and region-based text extraction for Node.js/Bun. Native Rust performance via [napi-rs](https://napi.rs).
|
||||
|
||||
@@ -7,9 +7,9 @@ Built by [Firecrawl](https://firecrawl.dev) for hybrid OCR pipelines — extract
|
||||
## Install
|
||||
|
||||
```bash
|
||||
npm install firecrawl-pdf-inspector
|
||||
npm install @firecrawl/pdf-inspector
|
||||
# or
|
||||
bun add firecrawl-pdf-inspector
|
||||
bun add @firecrawl/pdf-inspector
|
||||
```
|
||||
|
||||
Prebuilt binaries included for **linux-x64** and **macOS ARM64**. No Rust toolchain needed.
|
||||
@@ -21,7 +21,7 @@ Prebuilt binaries included for **linux-x64** and **macOS ARM64**. No Rust toolch
|
||||
Classify a PDF as TextBased, Scanned, Mixed, or ImageBased (~10-50ms). Returns which pages need OCR.
|
||||
|
||||
```typescript
|
||||
import { classifyPdf } from 'firecrawl-pdf-inspector'
|
||||
import { classifyPdf } from '@firecrawl/pdf-inspector'
|
||||
import { readFileSync } from 'fs'
|
||||
|
||||
const pdf = readFileSync('document.pdf')
|
||||
@@ -40,7 +40,7 @@ Extract text within bounding-box regions from a PDF. Designed for hybrid OCR pip
|
||||
Each region result includes a `needsOcr` flag that signals unreliable extraction (empty text, GID-encoded fonts, garbage text, encoding issues).
|
||||
|
||||
```typescript
|
||||
import { extractTextInRegions } from 'firecrawl-pdf-inspector'
|
||||
import { extractTextInRegions } from '@firecrawl/pdf-inspector'
|
||||
|
||||
const result = extractTextInRegions(pdf, [
|
||||
{
|
||||
|
||||
Executable
+131
@@ -0,0 +1,131 @@
|
||||
#!/usr/bin/env node
|
||||
|
||||
import { readFileSync, writeFileSync } from "fs";
|
||||
import { createRequire } from "module";
|
||||
|
||||
const require = createRequire(import.meta.url);
|
||||
const { version } = require("../package.json");
|
||||
|
||||
const HELP = `pdf-inspector v${version} — Fast PDF text extraction to Markdown
|
||||
|
||||
Usage:
|
||||
pdf-inspector <file> Extract markdown (default)
|
||||
pdf-inspector detect <file> Classify PDF type
|
||||
|
||||
Options:
|
||||
--json Output as JSON
|
||||
--pages <pages> Comma-separated page numbers (e.g. 1,3,5)
|
||||
-o, --output <file> Write output to file instead of stdout
|
||||
-h, --help Show this help
|
||||
-v, --version Show version
|
||||
|
||||
Examples:
|
||||
pdf-inspector document.pdf
|
||||
pdf-inspector document.pdf --json
|
||||
pdf-inspector document.pdf --pages 1,2,3
|
||||
pdf-inspector detect document.pdf --json
|
||||
cat document.pdf | pdf-inspector -`;
|
||||
|
||||
function die(msg) {
|
||||
process.stderr.write(`error: ${msg}\n`);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
function parseArgs(argv) {
|
||||
const opts = { json: false, pages: null, output: null, file: null, command: "extract" };
|
||||
let i = 0;
|
||||
|
||||
// Check for subcommand
|
||||
if (argv[0] === "detect") {
|
||||
opts.command = "detect";
|
||||
i = 1;
|
||||
}
|
||||
|
||||
while (i < argv.length) {
|
||||
const arg = argv[i];
|
||||
if (arg === "-h" || arg === "--help") {
|
||||
process.stdout.write(HELP + "\n");
|
||||
process.exit(0);
|
||||
} else if (arg === "-v" || arg === "--version") {
|
||||
process.stdout.write(`${version}\n`);
|
||||
process.exit(0);
|
||||
} else if (arg === "--json") {
|
||||
opts.json = true;
|
||||
} else if (arg === "--pages") {
|
||||
i++;
|
||||
if (!argv[i]) die("--pages requires a value (e.g. 1,3,5)");
|
||||
opts.pages = argv[i].split(",").map((p) => {
|
||||
const n = parseInt(p.trim(), 10);
|
||||
if (Number.isNaN(n) || n < 1) die(`invalid page number: ${p}`);
|
||||
return n;
|
||||
});
|
||||
} else if (arg === "-o" || arg === "--output") {
|
||||
i++;
|
||||
if (!argv[i]) die("-o requires a filename");
|
||||
opts.output = argv[i];
|
||||
} else if (arg === "-" || !arg.startsWith("-")) {
|
||||
if (opts.file) die(`unexpected argument: ${arg}`);
|
||||
opts.file = arg;
|
||||
} else {
|
||||
die(`unknown option: ${arg}`);
|
||||
}
|
||||
i++;
|
||||
}
|
||||
|
||||
return opts;
|
||||
}
|
||||
|
||||
function readInput(file) {
|
||||
if (file === "-") {
|
||||
return readFileSync(0); // stdin fd
|
||||
}
|
||||
try {
|
||||
return readFileSync(file);
|
||||
} catch (err) {
|
||||
if (err.code === "ENOENT") die(`file not found: ${file}`);
|
||||
die(err.message);
|
||||
}
|
||||
}
|
||||
|
||||
function output(text, outputPath) {
|
||||
if (outputPath) {
|
||||
writeFileSync(outputPath, text);
|
||||
} else {
|
||||
process.stdout.write(text);
|
||||
}
|
||||
}
|
||||
|
||||
// ---- main ----
|
||||
|
||||
const opts = parseArgs(process.argv.slice(2));
|
||||
|
||||
if (!opts.file) {
|
||||
// Check if stdin is piped
|
||||
if (process.stdin.isTTY !== false) {
|
||||
process.stderr.write(HELP + "\n");
|
||||
process.exit(1);
|
||||
}
|
||||
opts.file = "-";
|
||||
}
|
||||
|
||||
const { processPdf, classifyPdf } = await import("../index.js");
|
||||
const buffer = readInput(opts.file);
|
||||
|
||||
if (opts.command === "detect") {
|
||||
const result = classifyPdf(buffer);
|
||||
if (opts.json) {
|
||||
output(JSON.stringify(result, null, 2) + "\n", opts.output);
|
||||
} else {
|
||||
const ocr = result.pagesNeedingOcr.length > 0
|
||||
? `, ${result.pagesNeedingOcr.length} pages need OCR`
|
||||
: "";
|
||||
output(`${result.pdfType} (${result.pageCount} pages, confidence: ${result.confidence.toFixed(2)}${ocr})\n`, opts.output);
|
||||
}
|
||||
} else {
|
||||
const result = processPdf(buffer, opts.pages ?? undefined);
|
||||
if (opts.json) {
|
||||
output(JSON.stringify(result, null, 2) + "\n", opts.output);
|
||||
} else {
|
||||
output((result.markdown ?? "") + "\n", opts.output);
|
||||
}
|
||||
}
|
||||
+8
-3
@@ -1,9 +1,12 @@
|
||||
{
|
||||
"name": "firecrawl-pdf-inspector",
|
||||
"version": "0.3.3",
|
||||
"name": "@firecrawl/pdf-inspector",
|
||||
"version": "1.2.0",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
"bin": {
|
||||
"pdf-inspector": "bin/pdf-inspector.mjs"
|
||||
},
|
||||
"license": "MIT",
|
||||
"keywords": [
|
||||
"pdf",
|
||||
@@ -20,6 +23,7 @@
|
||||
"index.js",
|
||||
"index.d.ts",
|
||||
"*.node",
|
||||
"bin/",
|
||||
"README.md"
|
||||
],
|
||||
"repository": {
|
||||
@@ -34,7 +38,8 @@
|
||||
"binaryName": "pdf-inspector",
|
||||
"targets": [
|
||||
"x86_64-unknown-linux-gnu",
|
||||
"aarch64-apple-darwin"
|
||||
"aarch64-apple-darwin",
|
||||
"x86_64-pc-windows-msvc"
|
||||
],
|
||||
"package": {
|
||||
"name": "@firecrawl/pdf-inspector-js"
|
||||
|
||||
+167
-48
@@ -5,6 +5,28 @@ use napi_derive::napi;
|
||||
use std::collections::HashSet;
|
||||
use std::panic;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Enums
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// PDF document type classification.
|
||||
#[napi(string_enum)]
|
||||
pub enum PdfType {
|
||||
TextBased,
|
||||
Scanned,
|
||||
ImageBased,
|
||||
Mixed,
|
||||
}
|
||||
|
||||
/// Type of a positioned text item.
|
||||
#[napi(string_enum)]
|
||||
pub enum ItemType {
|
||||
Text,
|
||||
Image,
|
||||
Link,
|
||||
FormField,
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Result types
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -12,7 +34,7 @@ use std::panic;
|
||||
/// Full PDF processing result with markdown and metadata.
|
||||
#[napi(object)]
|
||||
pub struct PdfResult {
|
||||
pub pdf_type: String,
|
||||
pub pdf_type: PdfType,
|
||||
pub markdown: Option<String>,
|
||||
pub page_count: u32,
|
||||
pub processing_time_ms: u32,
|
||||
@@ -29,7 +51,7 @@ pub struct PdfResult {
|
||||
/// Lightweight PDF classification result.
|
||||
#[napi(object)]
|
||||
pub struct PdfClassification {
|
||||
pub pdf_type: String,
|
||||
pub pdf_type: PdfType,
|
||||
pub page_count: u32,
|
||||
/// 0-indexed page numbers that need OCR.
|
||||
pub pages_needing_ocr: Vec<u32>,
|
||||
@@ -49,7 +71,9 @@ pub struct TextItem {
|
||||
pub page: u32,
|
||||
pub is_bold: bool,
|
||||
pub is_italic: bool,
|
||||
pub item_type: String,
|
||||
pub item_type: ItemType,
|
||||
/// URL for link items, `None` for other types.
|
||||
pub link_url: Option<String>,
|
||||
}
|
||||
|
||||
/// A page's regions for text extraction: (page_index_0based, bboxes).
|
||||
@@ -79,18 +103,18 @@ pub struct PageRegionTexts {
|
||||
// Helpers
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn pdf_type_string(t: pdf_inspector::PdfType) -> String {
|
||||
fn convert_pdf_type(t: pdf_inspector::PdfType) -> PdfType {
|
||||
match t {
|
||||
pdf_inspector::PdfType::TextBased => "TextBased".to_string(),
|
||||
pdf_inspector::PdfType::Scanned => "Scanned".to_string(),
|
||||
pdf_inspector::PdfType::ImageBased => "ImageBased".to_string(),
|
||||
pdf_inspector::PdfType::Mixed => "Mixed".to_string(),
|
||||
pdf_inspector::PdfType::TextBased => PdfType::TextBased,
|
||||
pdf_inspector::PdfType::Scanned => PdfType::Scanned,
|
||||
pdf_inspector::PdfType::ImageBased => PdfType::ImageBased,
|
||||
pdf_inspector::PdfType::Mixed => PdfType::Mixed,
|
||||
}
|
||||
}
|
||||
|
||||
fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
|
||||
PdfResult {
|
||||
pdf_type: pdf_type_string(r.pdf_type),
|
||||
pdf_type: convert_pdf_type(r.pdf_type),
|
||||
markdown: r.markdown,
|
||||
page_count: r.page_count,
|
||||
processing_time_ms: r.processing_time_ms as u32,
|
||||
@@ -104,12 +128,12 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
|
||||
}
|
||||
}
|
||||
|
||||
fn item_type_string(t: &pdf_inspector::types::ItemType) -> String {
|
||||
fn convert_item_type(t: &pdf_inspector::types::ItemType) -> (ItemType, Option<String>) {
|
||||
match t {
|
||||
pdf_inspector::types::ItemType::Text => "text".into(),
|
||||
pdf_inspector::types::ItemType::Image => "image".into(),
|
||||
pdf_inspector::types::ItemType::Link(url) => format!("link:{url}"),
|
||||
pdf_inspector::types::ItemType::FormField => "form_field".into(),
|
||||
pdf_inspector::types::ItemType::Text => (ItemType::Text, None),
|
||||
pdf_inspector::types::ItemType::Image => (ItemType::Image, None),
|
||||
pdf_inspector::types::ItemType::Link(url) => (ItemType::Link, Some(url.clone())),
|
||||
pdf_inspector::types::ItemType::FormField => (ItemType::FormField, None),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -181,7 +205,7 @@ pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
|
||||
let result =
|
||||
pdf_inspector::classify_pdf_mem(&bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
|
||||
Ok(PdfClassification {
|
||||
pdf_type: pdf_type_string(result.pdf_type),
|
||||
pdf_type: convert_pdf_type(result.pdf_type),
|
||||
page_count: result.page_count,
|
||||
pages_needing_ocr: result.pages_needing_ocr,
|
||||
confidence: result.confidence as f64,
|
||||
@@ -222,18 +246,22 @@ pub fn extract_text_with_positions(
|
||||
|
||||
Ok(items
|
||||
.into_iter()
|
||||
.map(|item| TextItem {
|
||||
text: item.text,
|
||||
x: item.x as f64,
|
||||
y: item.y as f64,
|
||||
width: item.width as f64,
|
||||
height: item.height as f64,
|
||||
font: item.font,
|
||||
font_size: item.font_size as f64,
|
||||
page: item.page,
|
||||
is_bold: item.is_bold,
|
||||
is_italic: item.is_italic,
|
||||
item_type: item_type_string(&item.item_type),
|
||||
.map(|item| {
|
||||
let (item_type, link_url) = convert_item_type(&item.item_type);
|
||||
TextItem {
|
||||
text: item.text,
|
||||
x: item.x as f64,
|
||||
y: item.y as f64,
|
||||
width: item.width as f64,
|
||||
height: item.height as f64,
|
||||
font: item.font,
|
||||
font_size: item.font_size as f64,
|
||||
page: item.page,
|
||||
is_bold: item.is_bold,
|
||||
is_italic: item.is_italic,
|
||||
item_type,
|
||||
link_url,
|
||||
}
|
||||
})
|
||||
.collect())
|
||||
})
|
||||
@@ -255,7 +283,101 @@ pub fn extract_text_in_regions(
|
||||
page_regions: Vec<PageRegions>,
|
||||
) -> Result<Vec<PageRegionTexts>> {
|
||||
let bytes: Vec<u8> = buffer.to_vec();
|
||||
let regions: Vec<(u32, Vec<[f32; 4]>)> = page_regions
|
||||
let regions = parse_page_regions(&page_regions);
|
||||
|
||||
catch_panic("extract_text_in_regions", move || {
|
||||
let results = pdf_inspector::extract_text_in_regions_mem(&bytes, ®ions)
|
||||
.map_err(|e| to_napi_err(e, "extract_text_in_regions"))?;
|
||||
Ok(to_page_region_texts(results))
|
||||
})
|
||||
}
|
||||
|
||||
/// Extract markdown tables within bounding-box regions from a PDF.
|
||||
///
|
||||
/// Like `extractTextInRegions` but runs table detection on items within each
|
||||
/// region and returns markdown pipe-tables instead of flat text.
|
||||
///
|
||||
/// When table structure is detected, `text` contains a markdown pipe-table and
|
||||
/// `needsOcr` is `false`. When no table is found, `text` is empty and
|
||||
/// `needsOcr` is `true` so the caller can fall back to GPU OCR.
|
||||
///
|
||||
/// Coordinates are PDF points with top-left origin.
|
||||
#[napi]
|
||||
pub fn extract_tables_in_regions(
|
||||
buffer: Buffer,
|
||||
page_regions: Vec<PageRegions>,
|
||||
) -> Result<Vec<PageRegionTexts>> {
|
||||
let bytes: Vec<u8> = buffer.to_vec();
|
||||
let regions = parse_page_regions(&page_regions);
|
||||
|
||||
catch_panic("extract_tables_in_regions", move || {
|
||||
let results = pdf_inspector::extract_tables_in_regions_mem(&bytes, ®ions)
|
||||
.map_err(|e| to_napi_err(e, "extract_tables_in_regions"))?;
|
||||
Ok(to_page_region_texts(results))
|
||||
})
|
||||
}
|
||||
|
||||
/// Per-page markdown extraction result.
|
||||
#[napi(object)]
|
||||
pub struct PageMarkdownResult {
|
||||
/// 0-indexed page number.
|
||||
pub page: u32,
|
||||
/// Formatted markdown for this page.
|
||||
pub markdown: String,
|
||||
/// `true` when text on this page is unreliable.
|
||||
pub needs_ocr: bool,
|
||||
}
|
||||
|
||||
/// Combined per-page markdown extraction and layout classification result.
|
||||
#[napi(object)]
|
||||
pub struct PagesExtractionResult {
|
||||
/// Per-page markdown results.
|
||||
pub pages: Vec<PageMarkdownResult>,
|
||||
/// 1-indexed pages where tables were detected.
|
||||
pub pages_with_tables: Vec<u32>,
|
||||
/// 1-indexed pages where multi-column layout was detected.
|
||||
pub pages_with_columns: Vec<u32>,
|
||||
/// 1-indexed pages that need OCR (scanned/image-based).
|
||||
pub pages_needing_ocr: Vec<u32>,
|
||||
/// True if any page has tables or columns.
|
||||
pub is_complex: bool,
|
||||
}
|
||||
|
||||
/// Extract formatted markdown for specific pages of a PDF, with layout
|
||||
/// classification metadata.
|
||||
///
|
||||
/// Returns per-page markdown and classification data (tables, columns,
|
||||
/// OCR needs) from a single parse. Font statistics are computed from the
|
||||
/// full document so header detection is consistent across pages.
|
||||
#[napi]
|
||||
pub fn extract_pages_markdown(
|
||||
buffer: Buffer,
|
||||
pages: Vec<u32>,
|
||||
) -> Result<PagesExtractionResult> {
|
||||
let bytes: Vec<u8> = buffer.to_vec();
|
||||
catch_panic("extract_pages_markdown", move || {
|
||||
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, &pages)
|
||||
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
|
||||
Ok(PagesExtractionResult {
|
||||
pages: result
|
||||
.pages
|
||||
.into_iter()
|
||||
.map(|r| PageMarkdownResult {
|
||||
page: r.page,
|
||||
markdown: r.markdown,
|
||||
needs_ocr: r.needs_ocr,
|
||||
})
|
||||
.collect(),
|
||||
pages_with_tables: result.pages_with_tables,
|
||||
pages_with_columns: result.pages_with_columns,
|
||||
pages_needing_ocr: result.pages_needing_ocr,
|
||||
is_complex: result.is_complex,
|
||||
})
|
||||
})
|
||||
}
|
||||
|
||||
fn parse_page_regions(page_regions: &[PageRegions]) -> Vec<(u32, Vec<[f32; 4]>)> {
|
||||
page_regions
|
||||
.iter()
|
||||
.map(|pr| {
|
||||
let bboxes: Vec<[f32; 4]> = pr
|
||||
@@ -271,25 +393,22 @@ pub fn extract_text_in_regions(
|
||||
.collect();
|
||||
(pr.page, bboxes)
|
||||
})
|
||||
.collect();
|
||||
.collect()
|
||||
}
|
||||
|
||||
catch_panic("extract_text_in_regions", move || {
|
||||
let results = pdf_inspector::extract_text_in_regions_mem(&bytes, ®ions)
|
||||
.map_err(|e| to_napi_err(e, "extract_text_in_regions"))?;
|
||||
|
||||
Ok(results
|
||||
.into_iter()
|
||||
.map(|page_result| PageRegionTexts {
|
||||
page: page_result.page,
|
||||
regions: page_result
|
||||
.regions
|
||||
.into_iter()
|
||||
.map(|r| RegionText {
|
||||
text: r.text,
|
||||
needs_ocr: r.needs_ocr,
|
||||
})
|
||||
.collect(),
|
||||
})
|
||||
.collect())
|
||||
})
|
||||
fn to_page_region_texts(results: Vec<pdf_inspector::PageRegionResult>) -> Vec<PageRegionTexts> {
|
||||
results
|
||||
.into_iter()
|
||||
.map(|page_result| PageRegionTexts {
|
||||
page: page_result.page,
|
||||
regions: page_result
|
||||
.regions
|
||||
.into_iter()
|
||||
.map(|r| RegionText {
|
||||
text: r.text,
|
||||
needs_ocr: r.needs_ocr,
|
||||
})
|
||||
.collect(),
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
+2180
-97
File diff suppressed because it is too large
Load Diff
@@ -216,6 +216,7 @@ pub(crate) fn extract_page_text_items(
|
||||
let mut marked_content_stack: Vec<MarkedContentEntry> = Vec::new();
|
||||
let mut suppress_glyph_extraction = false;
|
||||
let mut actual_text_start_tm: Option<[f32; 6]> = None; // text matrix at BDC entry
|
||||
let mut actual_text_glyph_tm: Option<[f32; 6]> = None; // text matrix at first glyph inside BDC
|
||||
/// Get the innermost MCID from the marked content stack.
|
||||
fn current_mcid(stack: &[MarkedContentEntry]) -> Option<i64> {
|
||||
stack.iter().rev().find_map(|e| e.mcid)
|
||||
@@ -349,8 +350,15 @@ pub(crate) fn extract_page_text_items(
|
||||
)
|
||||
})
|
||||
});
|
||||
// ActualText: suppress glyph extraction, just advance text matrix
|
||||
// ActualText: suppress glyph extraction, just advance text matrix.
|
||||
// Capture the FIRST glyph's text matrix as the rendering position
|
||||
// for the ActualText item. Td ops between BDC and the first Tj
|
||||
// may have moved the position to the correct line — the BDC-entry
|
||||
// position (actual_text_start_tm) can be on the previous line.
|
||||
if suppress_glyph_extraction {
|
||||
if actual_text_glyph_tm.is_none() {
|
||||
actual_text_glyph_tm = Some(text_matrix);
|
||||
}
|
||||
if let Some(w_ts) = w_ts_opt {
|
||||
text_matrix[4] += w_ts * text_matrix[0];
|
||||
text_matrix[5] += w_ts * text_matrix[1];
|
||||
@@ -425,6 +433,10 @@ pub(crate) fn extract_page_text_items(
|
||||
let font_info = font_widths.get(¤t_font);
|
||||
let is_invisible = (text_rendering_mode == 3 && !include_invisible)
|
||||
|| suppress_glyph_extraction;
|
||||
// Capture first-glyph position for ActualText
|
||||
if suppress_glyph_extraction && actual_text_glyph_tm.is_none() {
|
||||
actual_text_glyph_tm = Some(text_matrix);
|
||||
}
|
||||
|
||||
// Compute space threshold based on font metrics when available
|
||||
let space_threshold = if let Some(font_info) = font_info {
|
||||
@@ -700,6 +712,7 @@ pub(crate) fn extract_page_text_items(
|
||||
if actual_text.is_some() {
|
||||
suppress_glyph_extraction = true;
|
||||
actual_text_start_tm = Some(text_matrix);
|
||||
actual_text_glyph_tm = None; // reset — will be captured at first Tj/TJ
|
||||
}
|
||||
marked_content_stack.push(MarkedContentEntry { actual_text, mcid });
|
||||
}
|
||||
@@ -707,8 +720,13 @@ pub(crate) fn extract_page_text_items(
|
||||
// End Marked Content — emit ActualText item with correct width
|
||||
if let Some(entry) = marked_content_stack.pop() {
|
||||
if let Some(at) = entry.actual_text {
|
||||
// Compute width from text matrix advancement during BDC..EMC
|
||||
if let Some(start_tm) = actual_text_start_tm.take() {
|
||||
// Use the first-glyph position (if available) instead of the
|
||||
// BDC-entry position. Td operators between BDC and the first
|
||||
// Tj may have moved the text position to the correct line —
|
||||
// the BDC-entry position can be on the previous line.
|
||||
let glyph_tm = actual_text_glyph_tm.take();
|
||||
let entry_tm = actual_text_start_tm.take();
|
||||
if let Some(start_tm) = glyph_tm.or(entry_tm) {
|
||||
let combined = multiply_matrices(&start_tm, &ctm);
|
||||
if combined[0].abs() >= combined[1].abs() {
|
||||
rotation_votes.horizontal += 1;
|
||||
|
||||
+160
-1
@@ -164,6 +164,10 @@ pub(crate) fn detect_columns(
|
||||
}
|
||||
}
|
||||
}
|
||||
// Try XY-cut fallback before giving up
|
||||
if let Some(columns) = try_xy_cut_split(&page_items, x_min, x_max, page) {
|
||||
return columns;
|
||||
}
|
||||
return vec![ColumnRegion { x_min, x_max }];
|
||||
}
|
||||
|
||||
@@ -184,7 +188,7 @@ pub(crate) fn detect_columns(
|
||||
if result.len() > 1 {
|
||||
return result;
|
||||
}
|
||||
return validate_and_build_columns(
|
||||
let result = validate_and_build_columns(
|
||||
&valleys,
|
||||
&page_items,
|
||||
x_min,
|
||||
@@ -195,6 +199,161 @@ pub(crate) fn detect_columns(
|
||||
page,
|
||||
false, // edge-based fallback
|
||||
);
|
||||
if result.len() > 1 {
|
||||
return result;
|
||||
}
|
||||
|
||||
// Fallback: XY-cut style gap detection. When the histogram finds no
|
||||
// clear valleys (common with asymmetric/sidebar layouts), look for the
|
||||
// largest horizontal gap between item edges. This is a simplified
|
||||
// single-level XY-cut inspired by opendataloader's XY-Cut++ algorithm.
|
||||
if page_items.len() >= 20 && !page_has_table {
|
||||
if let Some(columns) = try_xy_cut_split(&page_items, x_min, x_max, page) {
|
||||
return columns;
|
||||
}
|
||||
}
|
||||
|
||||
vec![ColumnRegion { x_min, x_max }]
|
||||
}
|
||||
|
||||
/// Simplified single-level XY-cut: find the largest horizontal gap between
|
||||
/// item right-edges and left-edges. If the gap is wide enough and both sides
|
||||
/// have sufficient items with vertical overlap, split into two columns.
|
||||
///
|
||||
/// Inspired by opendataloader's XY-Cut++ algorithm but without full recursion.
|
||||
/// Handles asymmetric layouts (sidebars) that the histogram misses because
|
||||
/// the narrow column has too few items to register in the occupancy profile.
|
||||
fn try_xy_cut_split(
|
||||
page_items: &[&TextItem],
|
||||
page_x_min: f32,
|
||||
page_x_max: f32,
|
||||
page: u32,
|
||||
) -> Option<Vec<ColumnRegion>> {
|
||||
const MIN_GAP: f32 = 15.0; // minimum gap to consider a split
|
||||
const MIN_ITEMS_MAJOR: usize = 10; // major column must have ≥10 items
|
||||
const MIN_ITEMS_MINOR: usize = 3; // minor column (sidebar) must have ≥3
|
||||
|
||||
let page_width = page_x_max - page_x_min;
|
||||
if page_width < 200.0 {
|
||||
return None;
|
||||
}
|
||||
|
||||
// Collect all item edges: (right_edge, left_edge) pairs sorted by right_edge
|
||||
// The gap between one item's right edge and the next item's left edge
|
||||
// reveals column gutters.
|
||||
let mut edges: Vec<(f32, f32)> = page_items
|
||||
.iter()
|
||||
.map(|i| (i.x, i.x + effective_width(i)))
|
||||
.collect();
|
||||
edges.sort_by(|a, b| a.0.total_cmp(&b.0));
|
||||
|
||||
// Find the largest gap between consecutive items (by left edge).
|
||||
// Use a sweep: sort left edges, find max gap between sorted right edges
|
||||
// of items to the left and left edges of items to the right.
|
||||
let mut left_edges: Vec<f32> = page_items.iter().map(|i| i.x).collect();
|
||||
left_edges.sort_by(|a, b| a.total_cmp(b));
|
||||
|
||||
// Build prefix max of right edges (for items sorted by left edge)
|
||||
let mut sorted_by_left: Vec<(f32, f32)> = page_items
|
||||
.iter()
|
||||
.map(|i| (i.x, i.x + effective_width(i)))
|
||||
.collect();
|
||||
sorted_by_left.sort_by(|a, b| a.0.total_cmp(&b.0));
|
||||
|
||||
let mut best_gap = 0.0f32;
|
||||
let mut best_split = 0.0f32;
|
||||
let mut max_right_so_far = f32::NEG_INFINITY;
|
||||
|
||||
for i in 0..sorted_by_left.len() - 1 {
|
||||
let (_, right) = sorted_by_left[i];
|
||||
max_right_so_far = max_right_so_far.max(right);
|
||||
|
||||
let (next_left, _) = sorted_by_left[i + 1];
|
||||
let gap = next_left - max_right_so_far;
|
||||
if gap > best_gap {
|
||||
best_gap = gap;
|
||||
best_split = (max_right_so_far + next_left) / 2.0;
|
||||
}
|
||||
}
|
||||
|
||||
if best_gap < MIN_GAP {
|
||||
return None;
|
||||
}
|
||||
|
||||
// Don't split at page margins (within 10% of edges)
|
||||
let margin = page_width * 0.10;
|
||||
if best_split - page_x_min < margin || page_x_max - best_split < margin {
|
||||
return None;
|
||||
}
|
||||
|
||||
// Count items on each side
|
||||
let left_count = page_items
|
||||
.iter()
|
||||
.filter(|i| i.x + effective_width(i) / 2.0 <= best_split)
|
||||
.count();
|
||||
let right_count = page_items
|
||||
.iter()
|
||||
.filter(|i| i.x + effective_width(i) / 2.0 > best_split)
|
||||
.count();
|
||||
|
||||
let (minor, major) = if left_count <= right_count {
|
||||
(left_count, right_count)
|
||||
} else {
|
||||
(right_count, left_count)
|
||||
};
|
||||
|
||||
if major < MIN_ITEMS_MAJOR || minor < MIN_ITEMS_MINOR {
|
||||
return None;
|
||||
}
|
||||
|
||||
// Check vertical overlap — both sides should span a meaningful Y range
|
||||
let left_items: Vec<&&TextItem> = page_items
|
||||
.iter()
|
||||
.filter(|i| i.x + effective_width(i) / 2.0 <= best_split)
|
||||
.collect();
|
||||
let right_items: Vec<&&TextItem> = page_items
|
||||
.iter()
|
||||
.filter(|i| i.x + effective_width(i) / 2.0 > best_split)
|
||||
.collect();
|
||||
|
||||
let l_y_min = left_items.iter().map(|i| i.y).fold(f32::INFINITY, f32::min);
|
||||
let l_y_max = left_items
|
||||
.iter()
|
||||
.map(|i| i.y)
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
let r_y_min = right_items
|
||||
.iter()
|
||||
.map(|i| i.y)
|
||||
.fold(f32::INFINITY, f32::min);
|
||||
let r_y_max = right_items
|
||||
.iter()
|
||||
.map(|i| i.y)
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
|
||||
let overlap_min = l_y_min.max(r_y_min);
|
||||
let overlap_max = l_y_max.min(r_y_max);
|
||||
let overlap = (overlap_max - overlap_min).max(0.0);
|
||||
let y_range = (l_y_max.max(r_y_max) - l_y_min.min(r_y_min)).max(1.0);
|
||||
|
||||
if overlap / y_range < 0.20 {
|
||||
return None;
|
||||
}
|
||||
|
||||
debug!(
|
||||
"page {}: XY-cut split at x={:.1} (gap={:.1}pt, left={}, right={})",
|
||||
page, best_split, best_gap, left_count, right_count
|
||||
);
|
||||
|
||||
Some(vec![
|
||||
ColumnRegion {
|
||||
x_min: page_x_min,
|
||||
x_max: best_split,
|
||||
},
|
||||
ColumnRegion {
|
||||
x_min: best_split,
|
||||
x_max: page_x_max,
|
||||
},
|
||||
])
|
||||
}
|
||||
|
||||
/// Check whether each proposed column contains paragraph-like content.
|
||||
|
||||
+173
-9
@@ -315,21 +315,79 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
|
||||
return items;
|
||||
}
|
||||
|
||||
// Group items by (page, Y position) with 5pt tolerance
|
||||
// Group items by (page, Y position) with 5pt tolerance, walking in
|
||||
// stream order. A new group also starts when the incoming item's X
|
||||
// would be a significant backtrack — signature of a new PDF text
|
||||
// block (`BT ... Tm`) starting at the left of the same Y band.
|
||||
// Without this split, slide-deck PDFs that stack several copy-paste
|
||||
// layers at the same tiny-text Y (e.g. a color-palette row and a
|
||||
// disclaimer body) later sort-by-X into character-interleaved
|
||||
// gibberish.
|
||||
let y_tolerance = 5.0;
|
||||
let mut line_groups: Vec<(u32, f32, Vec<&TextItem>)> = Vec::new();
|
||||
struct LineGroup<'a> {
|
||||
page: u32,
|
||||
y: f32,
|
||||
end_x: f32,
|
||||
font_size: f32,
|
||||
items: Vec<&'a TextItem>,
|
||||
}
|
||||
let mut line_groups: Vec<LineGroup> = Vec::new();
|
||||
|
||||
for item in &items {
|
||||
let found = line_groups
|
||||
.iter_mut()
|
||||
.find(|(pg, y, _)| *pg == item.page && (item.y - *y).abs() < y_tolerance);
|
||||
if let Some((_, _, group)) = found {
|
||||
group.push(item);
|
||||
} else {
|
||||
line_groups.push((item.page, item.y, vec![item]));
|
||||
// Find the most recently touched matching (page, y) bucket that
|
||||
// the item can legitimately continue (no sharp X backtrack).
|
||||
// Scanning from the end means the last-updated bucket wins, so a
|
||||
// new back-tracked block reliably opens a fresh group instead of
|
||||
// re-joining the old one.
|
||||
let item_right_limit = item.x + effective_merge_width(item);
|
||||
let mut target: Option<usize> = None;
|
||||
for (i, g) in line_groups.iter().enumerate().rev() {
|
||||
if g.page != item.page {
|
||||
continue;
|
||||
}
|
||||
if (g.y - item.y).abs() >= y_tolerance {
|
||||
continue;
|
||||
}
|
||||
// Detect a sharp X backtrack that can only be a new PDF text
|
||||
// block (`BT ... Tm`) starting at the same Y band. Only
|
||||
// trigger for very small fonts: in slide-deck PDFs, multiple
|
||||
// text blocks get stacked at the same Y in a tiny copy-paste
|
||||
// layer, and sort-by-X later interleaves them. Normal-sized
|
||||
// fonts (≥ 3pt) stay on the original behaviour — splitting
|
||||
// there tends to break tables and other valid layouts.
|
||||
if item.font_size < 3.0 && g.font_size < 3.0 {
|
||||
let backtrack = g.end_x - item.x;
|
||||
let min_backtrack = (item.font_size * 5.0).max(5.0);
|
||||
if backtrack > min_backtrack {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
target = Some(i);
|
||||
break;
|
||||
}
|
||||
match target {
|
||||
Some(i) => {
|
||||
line_groups[i].end_x = line_groups[i].end_x.max(item_right_limit);
|
||||
line_groups[i].items.push(item);
|
||||
}
|
||||
None => {
|
||||
line_groups.push(LineGroup {
|
||||
page: item.page,
|
||||
y: item.y,
|
||||
end_x: item_right_limit,
|
||||
font_size: item.font_size,
|
||||
items: vec![item],
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Re-shape into the tuple layout the rest of this function expects.
|
||||
let mut line_groups: Vec<(u32, f32, Vec<&TextItem>)> = line_groups
|
||||
.into_iter()
|
||||
.map(|g| (g.page, g.y, g.items))
|
||||
.collect();
|
||||
|
||||
// Sort each group by X position (direction-aware)
|
||||
for (_, _, group) in &mut line_groups {
|
||||
let rtl = is_rtl_text(group.iter().map(|i| &i.text));
|
||||
@@ -573,6 +631,112 @@ mod tests {
|
||||
assert_eq!(merged[0].text, "hello world");
|
||||
}
|
||||
|
||||
fn make_tiny_item(text: &str, x: f32, width: f32) -> TextItem {
|
||||
TextItem {
|
||||
text: text.into(),
|
||||
x,
|
||||
y: 700.0,
|
||||
width,
|
||||
height: 1.3,
|
||||
font: "F1".into(),
|
||||
font_size: 1.3, // tiny copy-paste / accessibility layer font
|
||||
page: 1,
|
||||
is_bold: false,
|
||||
is_italic: false,
|
||||
item_type: ItemType::Text,
|
||||
mcid: None,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn merge_items_splits_tiny_interleaved_text_blocks_at_same_y() {
|
||||
// Slide-deck copy-paste layers render several `BT` blocks into a
|
||||
// tiny (<3pt) text layer all at the same Y. Before the fix,
|
||||
// grouping by Y alone would fuse them into one bucket, and the
|
||||
// subsequent sort-by-X would character-interleave them:
|
||||
// "A O B S G T..." instead of "Alpha...Omega...".
|
||||
//
|
||||
// The sharp X backtrack between blocks must open a new group so
|
||||
// block-A items stay before block-B items in the output, even
|
||||
// after their shared Y bucket would be X-sorted.
|
||||
let items = vec![
|
||||
make_tiny_item("Alpha", 10.0, 2.0),
|
||||
make_tiny_item("Gamma", 40.0, 2.0),
|
||||
// Block B: Tm resets X back to the left. With the fix, this
|
||||
// opens a new group; without, X-sort would put "Omega"
|
||||
// between "Alpha" and "Gamma" (interleaving).
|
||||
make_tiny_item("Omega", 20.0, 2.0),
|
||||
make_tiny_item("Theta", 50.0, 2.0),
|
||||
];
|
||||
let merged = merge_text_items(items);
|
||||
let texts: Vec<String> = merged.iter().map(|m| m.text.clone()).collect();
|
||||
// The critical property is that "Alpha" + "Gamma" appear
|
||||
// together in stream order before "Omega" + "Theta" — not
|
||||
// interleaved. Individual merges within each block depend on
|
||||
// gap thresholds that don't matter for this invariant.
|
||||
let alpha_idx = texts.iter().position(|t| t.contains("Alpha")).unwrap();
|
||||
let gamma_idx = texts.iter().position(|t| t.contains("Gamma")).unwrap();
|
||||
let omega_idx = texts.iter().position(|t| t.contains("Omega")).unwrap();
|
||||
let theta_idx = texts.iter().position(|t| t.contains("Theta")).unwrap();
|
||||
assert!(
|
||||
alpha_idx < omega_idx && gamma_idx < omega_idx,
|
||||
"block-A items (Alpha, Gamma) must precede block-B items (Omega, Theta); got {texts:?}"
|
||||
);
|
||||
assert!(
|
||||
omega_idx < theta_idx,
|
||||
"Omega must come before Theta within block B; got {texts:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn merge_items_does_not_split_normal_font_table_rows() {
|
||||
// Normal body/table text at 12pt must not be split by the
|
||||
// X-backtrack rule — that rule applies only to tiny (<3pt)
|
||||
// copy-paste layers. A single 12pt group here means the rule
|
||||
// didn't fire: the items all end up X-sorted together, so
|
||||
// "extra" at X=50 sits between "Row1-ColA" (10) and "Row1-ColB"
|
||||
// (100) rather than staying at the end.
|
||||
let items = vec![
|
||||
make_merge_item("Row1-ColA", 10.0, 50.0),
|
||||
make_merge_item("Row1-ColB", 100.0, 50.0),
|
||||
make_merge_item("Row1-ColC", 200.0, 50.0),
|
||||
make_merge_item("extra", 50.0, 30.0),
|
||||
];
|
||||
let merged = merge_text_items(items);
|
||||
let texts: Vec<String> = merged.iter().map(|m| m.text.clone()).collect();
|
||||
// Find positions of the non-merged anchors.
|
||||
let a_idx = texts
|
||||
.iter()
|
||||
.position(|t| t.contains("Row1-ColA"))
|
||||
.expect("ColA in output");
|
||||
let extra_idx = texts
|
||||
.iter()
|
||||
.position(|t| t.contains("extra"))
|
||||
.expect("extra in output");
|
||||
let b_idx = texts
|
||||
.iter()
|
||||
.position(|t| t.contains("Row1-ColB"))
|
||||
.expect("ColB in output");
|
||||
assert!(
|
||||
a_idx < extra_idx && extra_idx < b_idx,
|
||||
"12pt items must be X-sorted together (Row1-ColA, extra, Row1-ColB); got {texts:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn merge_items_tolerates_small_kerning_overlap() {
|
||||
// Small negative gaps (ligature/kerning) must NOT split a word.
|
||||
// A ~1pt overlap between "Wo" and "rld" is well within normal
|
||||
// kerning and should stay in one group.
|
||||
let items = vec![
|
||||
make_merge_item("Wo", 100.0, 20.0), // end = 120
|
||||
make_merge_item("rld", 119.0, 25.0), // 1pt overlap: kerning, not a reset
|
||||
];
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1);
|
||||
assert_eq!(merged[0].text, "World");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_group_into_lines() {
|
||||
let items = vec![
|
||||
|
||||
+719
-9
@@ -1,3 +1,9 @@
|
||||
// Rust 1.95 introduced collapsible_match for `if` inside match arms.
|
||||
// The content-stream parsers use this pattern extensively (match on operator
|
||||
// name, then check `in_text_block && !op.operands.is_empty()`). Collapsing
|
||||
// these into match guards would hurt readability. Allow crate-wide.
|
||||
#![allow(clippy::collapsible_match)]
|
||||
|
||||
//! Smart PDF detection and text extraction using lopdf
|
||||
//!
|
||||
//! # Quick start
|
||||
@@ -299,6 +305,149 @@ pub fn classify_pdf_mem(buffer: &[u8]) -> Result<PdfClassification, PdfError> {
|
||||
})
|
||||
}
|
||||
|
||||
// =========================================================================
|
||||
// Per-page markdown extraction
|
||||
// =========================================================================
|
||||
|
||||
/// Per-page markdown extraction result.
|
||||
#[derive(Debug)]
|
||||
pub struct PageMarkdown {
|
||||
/// 0-indexed page number.
|
||||
pub page: u32,
|
||||
/// Formatted markdown for this page.
|
||||
pub markdown: String,
|
||||
/// `true` when text on this page is unreliable (GID-encoded fonts,
|
||||
/// encoding issues, garbage text, or empty extraction).
|
||||
pub needs_ocr: bool,
|
||||
}
|
||||
|
||||
/// Combined per-page markdown extraction and layout classification result.
|
||||
#[derive(Debug)]
|
||||
pub struct PagesExtractionResult {
|
||||
/// Per-page markdown results.
|
||||
pub pages: Vec<PageMarkdown>,
|
||||
/// 1-indexed pages where tables were detected.
|
||||
pub pages_with_tables: Vec<u32>,
|
||||
/// 1-indexed pages where multi-column layout was detected.
|
||||
pub pages_with_columns: Vec<u32>,
|
||||
/// 1-indexed pages that need OCR (scanned/image-based).
|
||||
pub pages_needing_ocr: Vec<u32>,
|
||||
/// True if any page has tables or columns.
|
||||
pub is_complex: bool,
|
||||
}
|
||||
|
||||
/// Extract formatted markdown for specific pages of a PDF, with layout
|
||||
/// classification metadata.
|
||||
///
|
||||
/// Unlike [`process_pdf_mem`] which returns one concatenated markdown string,
|
||||
/// this returns per-page markdown so callers can mix direct extraction
|
||||
/// (for simple text pages) with GPU OCR (for complex/scanned pages).
|
||||
///
|
||||
/// Font statistics are computed from the full document so header
|
||||
/// detection thresholds are consistent regardless of which pages are
|
||||
/// requested. Per-page `needs_ocr` is set when the page has GID-encoded
|
||||
/// fonts, encoding issues, or garbage text.
|
||||
///
|
||||
/// Layout complexity (tables, columns) is computed from the full document
|
||||
/// at near-zero cost since the items/rects/lines are already in memory.
|
||||
pub fn extract_pages_markdown_mem(
|
||||
buffer: &[u8],
|
||||
pages: &[u32],
|
||||
) -> Result<PagesExtractionResult, PdfError> {
|
||||
validate_pdf_bytes(buffer)?;
|
||||
let (doc, page_count) = load_document_from_mem(buffer)?;
|
||||
let font_cmaps = FontCMaps::from_doc(&doc);
|
||||
|
||||
// Extract ALL pages to get accurate, document-wide font stats.
|
||||
let ((all_items, all_rects, all_lines), page_thresholds, gid_pages) =
|
||||
extractor::extract_positioned_text_from_doc(&doc, &font_cmaps, None)?;
|
||||
|
||||
// Compute layout complexity from full document (near-zero cost).
|
||||
let complexity = compute_layout_complexity(&all_items, &all_rects, &all_lines);
|
||||
|
||||
// Compute font stats from full document (cross-page consistency).
|
||||
let font_stats = markdown::analysis::calculate_font_stats_from_items(&all_items);
|
||||
|
||||
let mut results = Vec::with_capacity(pages.len());
|
||||
let mut pages_needing_ocr = Vec::new();
|
||||
|
||||
for &page_0idx in pages {
|
||||
// Out-of-range pages → empty + needs_ocr
|
||||
if page_0idx >= page_count {
|
||||
pages_needing_ocr.push(page_0idx + 1);
|
||||
results.push(PageMarkdown {
|
||||
page: page_0idx,
|
||||
markdown: String::new(),
|
||||
needs_ocr: true,
|
||||
});
|
||||
continue;
|
||||
}
|
||||
|
||||
let page_1idx = page_0idx + 1;
|
||||
|
||||
// Filter items/rects for this page only
|
||||
let page_items: Vec<TextItem> = all_items
|
||||
.iter()
|
||||
.filter(|i| i.page == page_1idx)
|
||||
.cloned()
|
||||
.collect();
|
||||
|
||||
let page_rects: Vec<PdfRect> = all_rects
|
||||
.iter()
|
||||
.filter(|r| r.page == page_1idx)
|
||||
.cloned()
|
||||
.collect();
|
||||
|
||||
let has_gid = gid_pages.contains(&page_1idx);
|
||||
|
||||
// Build markdown with document-wide font stats
|
||||
let options = MarkdownOptions {
|
||||
base_font_size: Some(font_stats.most_common_size),
|
||||
include_page_numbers: false,
|
||||
strip_headers_footers: false,
|
||||
..MarkdownOptions::default()
|
||||
};
|
||||
|
||||
let md = markdown::to_markdown_from_items_with_rects_and_lines(
|
||||
page_items,
|
||||
options,
|
||||
&page_rects,
|
||||
&[],
|
||||
&page_thresholds,
|
||||
None,
|
||||
&[],
|
||||
);
|
||||
|
||||
let needs_ocr = md.trim().is_empty()
|
||||
|| has_gid
|
||||
|| is_garbage_text(&md)
|
||||
|| is_cid_garbage(&md)
|
||||
|| detect_encoding_issues(&md);
|
||||
|
||||
if needs_ocr {
|
||||
pages_needing_ocr.push(page_1idx);
|
||||
}
|
||||
|
||||
results.push(PageMarkdown {
|
||||
page: page_0idx,
|
||||
markdown: if needs_ocr { String::new() } else { md },
|
||||
needs_ocr,
|
||||
});
|
||||
}
|
||||
|
||||
Ok(PagesExtractionResult {
|
||||
pages: results,
|
||||
pages_with_tables: complexity.pages_with_tables,
|
||||
pages_with_columns: complexity.pages_with_columns,
|
||||
pages_needing_ocr,
|
||||
is_complex: complexity.is_complex,
|
||||
})
|
||||
}
|
||||
|
||||
// =========================================================================
|
||||
// Region-based text extraction (for hybrid OCR pipelines)
|
||||
// =========================================================================
|
||||
|
||||
/// Result for a single region's text extraction.
|
||||
#[derive(Debug)]
|
||||
pub struct RegionText {
|
||||
@@ -399,7 +548,7 @@ pub fn extract_text_in_regions_mem(
|
||||
let page_1idx = page_0idx + 1;
|
||||
let items = items_by_page.get(&page_1idx);
|
||||
let page_h = page_heights.get(&page_1idx).copied().unwrap_or(792.0);
|
||||
let page_has_gid = gid_pages.contains(&page_1idx);
|
||||
let _page_has_gid = gid_pages.contains(&page_1idx);
|
||||
let adaptive_threshold = page_thresholds.get(&page_1idx).copied().unwrap_or(0.10);
|
||||
let coords = if rotated_pages.contains(&page_1idx) {
|
||||
RegionCoordSpace::Rotated90Ccw
|
||||
@@ -426,8 +575,10 @@ pub fn extract_text_in_regions_mem(
|
||||
None => String::new(),
|
||||
};
|
||||
|
||||
// Check per-region text quality instead of blanket page-level
|
||||
// GID rejection. A GID font in a logo elsewhere on the page
|
||||
// shouldn't force GPU OCR for clean text regions.
|
||||
let needs_ocr = text.trim().is_empty()
|
||||
|| page_has_gid
|
||||
|| is_garbage_text(&text)
|
||||
|| is_cid_garbage(&text)
|
||||
|| detect_encoding_issues(&text);
|
||||
@@ -444,6 +595,167 @@ pub fn extract_text_in_regions_mem(
|
||||
Ok(results)
|
||||
}
|
||||
|
||||
/// Extract tables within bounding-box regions from a PDF in memory.
|
||||
///
|
||||
/// Similar to [`extract_text_in_regions_mem`] but runs table detection on items
|
||||
/// within each region and returns markdown pipe-tables instead of flat text.
|
||||
///
|
||||
/// When table structure is detected, `text` contains a markdown pipe-table and
|
||||
/// `needs_ocr` is `false`. When no table is found (too few items, poor alignment,
|
||||
/// GID fonts, etc.), `text` is empty and `needs_ocr` is `true` so the caller can
|
||||
/// fall back to GPU OCR.
|
||||
pub fn extract_tables_in_regions_mem(
|
||||
buffer: &[u8],
|
||||
page_regions: &[(u32, Vec<[f32; 4]>)],
|
||||
) -> Result<Vec<PageRegionResult>, PdfError> {
|
||||
validate_pdf_bytes(buffer)?;
|
||||
let (doc, _page_count) = load_document_from_mem(buffer)?;
|
||||
let pages = doc.get_pages();
|
||||
|
||||
let needed_pages: HashSet<u32> = page_regions.iter().map(|(p, _)| p + 1).collect();
|
||||
let font_cmaps = FontCMaps::from_doc_pages_fast(&doc, Some(&needed_pages));
|
||||
|
||||
let mut items_by_page: HashMap<u32, Vec<TextItem>> = HashMap::new();
|
||||
let mut page_heights: HashMap<u32, f32> = HashMap::new();
|
||||
let mut gid_pages: HashSet<u32> = HashSet::new();
|
||||
let mut page_thresholds: HashMap<u32, f32> = HashMap::new();
|
||||
let mut rotated_pages: HashSet<u32> = HashSet::new();
|
||||
|
||||
for (page_num, &page_id) in pages.iter() {
|
||||
if !needed_pages.contains(page_num) {
|
||||
continue;
|
||||
}
|
||||
let height = get_page_height(&doc, page_id).unwrap_or(792.0);
|
||||
page_heights.insert(*page_num, height);
|
||||
|
||||
let ((mut items, _rects, _lines), has_gid, coords_rotated) =
|
||||
extractor::content_stream::extract_page_text_items(
|
||||
&doc,
|
||||
page_id,
|
||||
*page_num,
|
||||
&font_cmaps,
|
||||
false,
|
||||
)?;
|
||||
let threshold = text_utils::fix_letterspaced_items(&mut items);
|
||||
if threshold > 0.10 {
|
||||
page_thresholds.insert(*page_num, threshold);
|
||||
}
|
||||
if has_gid {
|
||||
gid_pages.insert(*page_num);
|
||||
}
|
||||
if coords_rotated {
|
||||
rotated_pages.insert(*page_num);
|
||||
}
|
||||
items_by_page.insert(*page_num, items);
|
||||
}
|
||||
|
||||
let mut results = Vec::with_capacity(page_regions.len());
|
||||
|
||||
for (page_0idx, regions) in page_regions {
|
||||
let page_1idx = page_0idx + 1;
|
||||
let items = items_by_page.get(&page_1idx);
|
||||
let page_h = page_heights.get(&page_1idx).copied().unwrap_or(792.0);
|
||||
let _page_has_gid = gid_pages.contains(&page_1idx);
|
||||
let coords = if rotated_pages.contains(&page_1idx) {
|
||||
RegionCoordSpace::Rotated90Ccw
|
||||
} else {
|
||||
RegionCoordSpace::Standard
|
||||
};
|
||||
|
||||
let mut page_results = Vec::with_capacity(regions.len());
|
||||
|
||||
for rect in regions {
|
||||
let [rx1, ry1, rx2, ry2] = *rect;
|
||||
|
||||
// Note: we intentionally DO NOT bail on page_has_gid here.
|
||||
// The GID flag means some font on the page uses unresolvable
|
||||
// glyph IDs, but that font may only appear in a logo or
|
||||
// header — not in the table region. Instead we let the
|
||||
// per-region text quality checks (is_garbage_text, is_cid_garbage,
|
||||
// detect_encoding_issues) reject based on the actual extracted
|
||||
// content. This avoids rejecting clean tables just because an
|
||||
// unrelated decorative font on the same page is GID-encoded.
|
||||
|
||||
let matched: Vec<TextItem> = match items {
|
||||
Some(items) => {
|
||||
let bounds = region_bounds(rx1, ry1, rx2, ry2, page_h, coords);
|
||||
items
|
||||
.iter()
|
||||
.filter(|item| region_overlaps_item(item, bounds))
|
||||
.cloned()
|
||||
.collect()
|
||||
}
|
||||
None => Vec::new(),
|
||||
};
|
||||
|
||||
if matched.is_empty() {
|
||||
page_results.push(RegionText {
|
||||
text: String::new(),
|
||||
needs_ocr: true,
|
||||
});
|
||||
continue;
|
||||
}
|
||||
|
||||
// Compute base_font_size as most common font size in the region
|
||||
let base_font_size = {
|
||||
let mut freq: HashMap<i32, usize> = HashMap::new();
|
||||
for item in &matched {
|
||||
*freq.entry((item.font_size * 10.0) as i32).or_default() += 1;
|
||||
}
|
||||
freq.into_iter()
|
||||
.max_by_key(|(_, count)| *count)
|
||||
.map(|(size, _)| size as f32 / 10.0)
|
||||
.unwrap_or(12.0)
|
||||
};
|
||||
|
||||
// Run heuristic table detection; skip_body_font = false since
|
||||
// the layout model already identified this region as a table.
|
||||
let detected = tables::detect_tables(&matched, base_font_size, false);
|
||||
|
||||
if let Some(table) = detected.into_iter().next() {
|
||||
let md = tables::table_to_markdown(&table);
|
||||
if md.trim().is_empty() {
|
||||
page_results.push(RegionText {
|
||||
text: String::new(),
|
||||
needs_ocr: true,
|
||||
});
|
||||
} else {
|
||||
// needs_ocr fires on any of:
|
||||
// - garbage text (non-alphanumeric heavy)
|
||||
// - CID/Latin-1 mojibake
|
||||
// - encoding issues (U+FFFD, dollar-as-space)
|
||||
// - structural giveaways that the table is partial /
|
||||
// mis-detected (numeric "header", empty header cells,
|
||||
// duplicate header cells). Caught GLM-OCR-as-baseline
|
||||
// scoring 0 TEDS on real prod tables in eval.
|
||||
// Layout model already identified this region as a table,
|
||||
// so use relaxed partial-table checks (layout_assisted=true).
|
||||
let needs_ocr = is_garbage_text(&md)
|
||||
|| is_cid_garbage(&md)
|
||||
|| detect_encoding_issues(&md)
|
||||
|| looks_like_partial_table_ex(&md, true);
|
||||
page_results.push(RegionText {
|
||||
text: if needs_ocr { String::new() } else { md },
|
||||
needs_ocr,
|
||||
});
|
||||
}
|
||||
} else {
|
||||
page_results.push(RegionText {
|
||||
text: String::new(),
|
||||
needs_ocr: true,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
results.push(PageRegionResult {
|
||||
page: *page_0idx,
|
||||
regions: page_results,
|
||||
});
|
||||
}
|
||||
|
||||
Ok(results)
|
||||
}
|
||||
|
||||
/// Get page height in points from MediaBox.
|
||||
fn get_page_height(doc: &Document, page_id: lopdf::ObjectId) -> Option<f32> {
|
||||
let page_dict = doc.get_dictionary(page_id).ok()?;
|
||||
@@ -525,9 +837,6 @@ fn collect_text_in_region_with_options(
|
||||
adaptive_threshold: f32,
|
||||
) -> String {
|
||||
let bounds = region_bounds(rx1, ry1, rx2, ry2, page_height, coord_space);
|
||||
let Some(page) = items.first().map(|item| item.page) else {
|
||||
return String::new();
|
||||
};
|
||||
let matched: Vec<TextItem> = items
|
||||
.iter()
|
||||
.filter(|item| region_overlaps_item(item, bounds))
|
||||
@@ -536,11 +845,39 @@ fn collect_text_in_region_with_options(
|
||||
if matched.is_empty() {
|
||||
return String::new();
|
||||
}
|
||||
let mut thresholds = HashMap::new();
|
||||
if adaptive_threshold > 0.10 {
|
||||
thresholds.insert(page, adaptive_threshold);
|
||||
|
||||
// Simple extraction: the caller (fire-pdf) already handles reading order
|
||||
// and column splitting via the layout model. We just need to sort items
|
||||
// top-to-bottom, left-to-right and group into lines.
|
||||
let mut sorted = matched;
|
||||
sorted.sort_by(|a, b| b.y.total_cmp(&a.y).then(a.x.total_cmp(&b.x)));
|
||||
|
||||
let y_tolerance = 3.0;
|
||||
let mut lines: Vec<extractor::TextLine> = Vec::new();
|
||||
|
||||
for item in sorted {
|
||||
let should_merge = lines.last().is_some_and(|last_line: &extractor::TextLine| {
|
||||
last_line.page == item.page && (last_line.y - item.y).abs() < y_tolerance
|
||||
});
|
||||
if should_merge {
|
||||
lines.last_mut().unwrap().items.push(item);
|
||||
} else {
|
||||
let y = item.y;
|
||||
let page = item.page;
|
||||
lines.push(extractor::TextLine {
|
||||
items: vec![item],
|
||||
y,
|
||||
page,
|
||||
adaptive_threshold,
|
||||
});
|
||||
}
|
||||
}
|
||||
let lines = extractor::group_into_lines_with_thresholds(matched, &thresholds, &HashSet::new());
|
||||
|
||||
// Sort items within each line by X position
|
||||
for line in &mut lines {
|
||||
text_utils::sort_line_items(&mut line.items);
|
||||
}
|
||||
|
||||
lines
|
||||
.into_iter()
|
||||
.map(|line| line.text())
|
||||
@@ -1016,6 +1353,379 @@ fn is_cid_garbage(text: &str) -> bool {
|
||||
high_latin * 5 >= total * 2 && ascii_letters * 3 < total
|
||||
}
|
||||
|
||||
/// Detect markdown tables with suspicious structure that suggest the heuristic
|
||||
/// missed/mangled rows or columns. Returns true when the caller should treat
|
||||
/// the result as `needs_ocr` and fall back to GPU OCR.
|
||||
///
|
||||
/// Catches three failure modes observed in production:
|
||||
///
|
||||
/// 1. **Header row looks like a data row** — first row starts with a numeric
|
||||
/// value (e.g. `|2|...`), suggesting we missed the actual header above it.
|
||||
/// Real headers almost never start with a bare number.
|
||||
///
|
||||
/// 2. **Header has empty cells in a multi-column table** — e.g.
|
||||
/// `|Position||Administration|Administration|` (3+ cols, ≥1 empty cell).
|
||||
/// Indicates poor column boundary detection.
|
||||
///
|
||||
/// 3. **Header has duplicate non-empty cells** in a multi-column table —
|
||||
/// e.g. `Administration|Administration` appearing as adjacent cells means
|
||||
/// we collapsed multi-line headers wrong.
|
||||
///
|
||||
/// Conservative by design: a few false positives (perfectly fine tables flagged)
|
||||
/// just mean we run GPU OCR which is the existing safe path.
|
||||
/// When `layout_assisted` is true (the layout model identified this region
|
||||
/// as a table), we relax boundary-detection heuristics (numeric header,
|
||||
/// empty header cells, sparse first data row) because the layout model
|
||||
/// already gave us the table bbox — we're not guessing "is this a table?"
|
||||
/// anymore, only "can we extract it correctly?". Paragraph and duplicate-
|
||||
/// header checks stay, since those indicate genuine extraction quality
|
||||
/// issues regardless of how the region was identified.
|
||||
fn looks_like_partial_table_ex(markdown: &str, layout_assisted: bool) -> bool {
|
||||
let lines: Vec<&str> = markdown.lines().filter(|l| l.starts_with('|')).collect();
|
||||
if lines.len() < 2 {
|
||||
return false;
|
||||
}
|
||||
// Header is the first pipe-line; separator is the second
|
||||
let header_line = lines[0];
|
||||
let separator_line = lines.get(1).copied().unwrap_or("");
|
||||
let is_separator = |l: &str| l.chars().all(|c| matches!(c, '|' | '-' | ' '));
|
||||
if !is_separator(separator_line) {
|
||||
// No separator after the first line — not a well-formed pipe-table.
|
||||
// table_to_markdown always emits one when it returns content, so this
|
||||
// shouldn't happen in practice. If it does, fall through to OCR.
|
||||
return true;
|
||||
}
|
||||
|
||||
// Parse header cells: split on '|', drop the leading/trailing empty pieces
|
||||
let cells: Vec<&str> = header_line.split('|').map(|s| s.trim()).collect::<Vec<_>>();
|
||||
// The first and last items are always empty (string starts and ends with '|')
|
||||
if cells.len() < 3 {
|
||||
return false;
|
||||
}
|
||||
let header_cells: Vec<&str> = cells[1..cells.len() - 1].to_vec();
|
||||
let n_cols = header_cells.len();
|
||||
if n_cols < 2 {
|
||||
// Single-column tables are usually lists/keys, not tables. Keep them
|
||||
// (caller can decide), but multi-column header checks below don't
|
||||
// apply.
|
||||
return false;
|
||||
}
|
||||
|
||||
// Failure mode 1: header starts with a bare number (likely we missed
|
||||
// the real header row above). Skip when layout-assisted — the layout
|
||||
// model's bbox includes the real header; a numeric first cell (e.g.,
|
||||
// a year "2024") is legitimate.
|
||||
if !layout_assisted {
|
||||
if let Some(first) = header_cells.first() {
|
||||
let trimmed = first.trim();
|
||||
if !trimmed.is_empty() && trimmed.chars().all(|c| c.is_ascii_digit()) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Failure mode 2: header has empty cells in a multi-column table.
|
||||
// When layout-assisted, allow up to 1 empty header cell (common in
|
||||
// tables with merged/spanning header cells that we can't represent).
|
||||
let empty_count = header_cells.iter().filter(|c| c.is_empty()).count();
|
||||
if layout_assisted {
|
||||
// Reject only if >1 empty header cell (2+ means serious boundary issue)
|
||||
if n_cols >= 3 && empty_count >= 2 {
|
||||
return true;
|
||||
}
|
||||
} else if n_cols >= 3 && empty_count >= 1 {
|
||||
return true;
|
||||
}
|
||||
|
||||
// Failure mode 3: header has duplicate non-empty cells
|
||||
let mut seen: std::collections::HashSet<&str> = std::collections::HashSet::new();
|
||||
for cell in &header_cells {
|
||||
if cell.is_empty() {
|
||||
continue;
|
||||
}
|
||||
if !seen.insert(cell) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
// Failure mode 4: first data row has many empty cells in a multi-column
|
||||
// table. Real tables rarely have a leading row with most cells blank;
|
||||
// when this happens it usually means the heuristic split a multi-row
|
||||
// header (e.g. "Position\nAdministration (1986-1992) | Administration
|
||||
// (1992-1998)") into a single-row header + a sparse data row.
|
||||
if let Some(first_data_line) = lines.get(2) {
|
||||
let data_cells: Vec<&str> = first_data_line
|
||||
.split('|')
|
||||
.map(|s| s.trim())
|
||||
.collect::<Vec<_>>();
|
||||
if data_cells.len() >= 3 {
|
||||
let data_inner = &data_cells[1..data_cells.len() - 1];
|
||||
let empty_data = data_inner.iter().filter(|c| c.is_empty()).count();
|
||||
// ≥3 cols, and significant portion of cells in the first data
|
||||
// row are empty → likely we mis-split a multi-row header.
|
||||
// When layout-assisted, relax from 33% to 50% — the bbox is
|
||||
// more reliable, and real tables with one sparse first row
|
||||
// (totals, subtotals) are common.
|
||||
let threshold = if layout_assisted { 2 } else { 3 };
|
||||
if n_cols >= 3 && empty_data * threshold >= n_cols {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Failure mode 5: cells flow as continuation paragraph (text wrapping
|
||||
// mistaken for column structure). When a paragraph of prose gets mis-
|
||||
// detected as a multi-column table, cells in the same column tend to
|
||||
// start with lowercase letters or punctuation (continuation), not
|
||||
// capital letters / digits (new entries). Real tables almost never
|
||||
// have most data cells starting lowercase.
|
||||
//
|
||||
// Signal: ≥2 cols, ≥4 data rows, and ≥60% of non-empty data cells
|
||||
// start with a lowercase letter or continuation punctuation.
|
||||
let data_rows: Vec<Vec<&str>> = lines
|
||||
.iter()
|
||||
.skip(2) // header + separator
|
||||
.map(|l| {
|
||||
let parts: Vec<&str> = l.split('|').map(|s| s.trim()).collect();
|
||||
if parts.len() >= 3 {
|
||||
parts[1..parts.len() - 1].to_vec()
|
||||
} else {
|
||||
Vec::new()
|
||||
}
|
||||
})
|
||||
.filter(|cells| !cells.is_empty())
|
||||
.collect();
|
||||
|
||||
if n_cols >= 2 && data_rows.len() >= 4 {
|
||||
let mut continuation = 0;
|
||||
let mut total = 0;
|
||||
for row in &data_rows {
|
||||
for cell in row {
|
||||
let trimmed = cell.trim();
|
||||
if trimmed.is_empty() {
|
||||
continue;
|
||||
}
|
||||
total += 1;
|
||||
let first = trimmed.chars().next().unwrap();
|
||||
// Continuation indicators: lowercase letter, common
|
||||
// mid-sentence punctuation, closing quote
|
||||
if first.is_lowercase()
|
||||
|| matches!(first, ',' | '.' | ';' | ')' | '"' | '\'' | '”' | '’')
|
||||
{
|
||||
continuation += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
if total > 0 && continuation * 5 >= total * 3 {
|
||||
// ≥60% of cells look like sentence continuations → paragraph
|
||||
// misread as table.
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
false
|
||||
}
|
||||
|
||||
/// Original strict validation (no layout assistance). Used by tests and
|
||||
/// full-page extraction paths that don't have layout model assistance.
|
||||
#[cfg(test)]
|
||||
fn looks_like_partial_table(markdown: &str) -> bool {
|
||||
looks_like_partial_table_ex(markdown, false)
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod looks_like_partial_table_tests {
|
||||
use super::{looks_like_partial_table, looks_like_partial_table_ex};
|
||||
|
||||
#[test]
|
||||
fn good_table_passes() {
|
||||
let md = "|Name|Year|Country|\n|---|---|---|\n|Alice|2020|US|\n|Bob|2021|UK|";
|
||||
assert!(
|
||||
!looks_like_partial_table(md),
|
||||
"should not flag well-formed table"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn header_starting_with_number_is_partial() {
|
||||
// Heuristic missed the actual header row above
|
||||
let md = "|2|Cambodian Women for Peace|9,835|\n|---|---|---|\n|3|Association|711|";
|
||||
assert!(looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn header_with_empty_cells_in_3col_is_partial() {
|
||||
// Empty cell in 3+ column header → bad column detection
|
||||
let md =
|
||||
"|Position||Administration|Administration|\n|---|---|---|---|\n|Senate|24|8.3|16.7|";
|
||||
assert!(looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn header_with_duplicate_cells_is_partial() {
|
||||
// Duplicate "Administration" → collapsed multi-line header wrong
|
||||
let md =
|
||||
"|Position|Administration|Administration|Notes|\n|---|---|---|---|\n|Senate|24|16|x|";
|
||||
assert!(looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn two_column_with_one_empty_cell_passes() {
|
||||
// Many real two-column tables have key-only rows; don't penalise.
|
||||
let md = "|Key||\n|---|---|\n|Alice|123|\n|Bob|456|";
|
||||
// Header "Key|" has one empty cell but only 2 cols total — keep it.
|
||||
assert!(!looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn single_column_table_is_kept() {
|
||||
// Single-column "tables" are common (lists). Caller can decide; we
|
||||
// don't second-guess based on column count alone.
|
||||
let md = "|Item|\n|---|\n|First|\n|Second|";
|
||||
assert!(!looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn no_table_at_all_returns_true() {
|
||||
// table_to_markdown should never produce this, but defensive — if
|
||||
// there's no separator, treat as not-a-table.
|
||||
let md = "Just some text\nWith multiple lines";
|
||||
// No lines start with '|' so we return false (no header to inspect).
|
||||
assert!(!looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn first_data_row_with_many_empty_cells_is_partial() {
|
||||
// Multi-row header collapsed to single-row → first "data row" has
|
||||
// most cells empty (the actual sub-header values).
|
||||
let md = "|Government|No. of Seats|Aquino|Ramos|\n|---|---|---|---|\n|Position|||(1986-1992)|\n|Senate|24|8.3|16.7|";
|
||||
assert!(looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn first_data_row_with_one_empty_cell_in_4col_passes() {
|
||||
// Real data rows can have one empty cell (e.g. missing value);
|
||||
// only flag when ≥1/3 of cells are empty.
|
||||
let md = "|A|B|C|D|\n|---|---|---|---|\n|x|y||z|\n|p|q|r|s|";
|
||||
assert!(!looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn paragraph_misread_as_two_column_table_is_partial() {
|
||||
// Real production failure: text-wrapped paragraph mis-detected as
|
||||
// 2-col table. Each cell continues the previous one as prose.
|
||||
let md = "|Approval is needed from the|Acquisitions of|\n\
|
||||
|---|---|\n\
|
||||
|Treasurer if the acquisition|residential and|\n\
|
||||
|constitutes a \"significant|agricultural|\n\
|
||||
|action,\" including acquiring an|land by foreign|\n\
|
||||
|interest in different types of|persons must be|\n\
|
||||
|land where the monetary|reported to the|";
|
||||
assert!(looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn real_multi_word_table_is_kept() {
|
||||
// Real table with multi-word entries — cells start with capital
|
||||
// letters / proper nouns, NOT lowercase continuations.
|
||||
let md = "|Country|Capital|Notes|\n\
|
||||
|---|---|---|\n\
|
||||
|United States|Washington DC|Federal capital|\n\
|
||||
|United Kingdom|London|City of London is a separate|\n\
|
||||
|France|Paris|Île-de-France region|\n\
|
||||
|Germany|Berlin|Reunified 1990|\n\
|
||||
|Spain|Madrid|Largest city in Spain|";
|
||||
assert!(!looks_like_partial_table(md));
|
||||
}
|
||||
|
||||
// --- layout_assisted relaxation tests ---
|
||||
|
||||
#[test]
|
||||
fn numeric_header_accepted_when_layout_assisted() {
|
||||
// Year as first header cell is valid when layout model gave us the bbox.
|
||||
let md = "|2024|Revenue|Growth|\n|---|---|---|\n|Q1|1.2M|5%|\n|Q2|1.4M|8%|";
|
||||
assert!(
|
||||
looks_like_partial_table(md),
|
||||
"strict mode rejects numeric header"
|
||||
);
|
||||
assert!(
|
||||
!looks_like_partial_table_ex(md, true),
|
||||
"layout-assisted should accept"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn one_empty_header_accepted_when_layout_assisted() {
|
||||
// Common in merged-header tables: one spanning cell leaves a gap.
|
||||
let md = "|Position||Senate|House|\n|---|---|---|---|\n|Chair|1|2|3|\n|Vice|4|5|6|";
|
||||
assert!(
|
||||
looks_like_partial_table(md),
|
||||
"strict rejects 1 empty header"
|
||||
);
|
||||
assert!(
|
||||
!looks_like_partial_table_ex(md, true),
|
||||
"layout-assisted allows 1 empty"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn two_empty_headers_still_rejected_when_layout_assisted() {
|
||||
// 2+ empty headers is still bad even with layout assistance.
|
||||
let md = "|A|||D|\n|---|---|---|---|\n|x|y|z|w|";
|
||||
assert!(
|
||||
looks_like_partial_table_ex(md, true),
|
||||
"2 empty headers rejected even layout-assisted"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn sparse_first_row_relaxed_when_layout_assisted() {
|
||||
// 1/4 empty = 25%, below strict 33% threshold but accepted by layout-assisted 50%.
|
||||
let md = "|A|B|C|D|\n|---|---|---|---|\n|x||y|z|\n|p|q|r|s|";
|
||||
assert!(!looks_like_partial_table(md), "strict: 25% empty is OK");
|
||||
// 2/4 = 50%, strict would flag (2*3>=4), relaxed threshold (2*2>=4) would also flag.
|
||||
let md2 = "|A|B|C|D|\n|---|---|---|---|\n|||y|z|\n|p|q|r|s|";
|
||||
assert!(looks_like_partial_table(md2), "strict: 50% empty flagged");
|
||||
assert!(
|
||||
looks_like_partial_table_ex(md2, true),
|
||||
"layout-assisted: 50% also flagged"
|
||||
);
|
||||
// 2/6 = 33%, strict flags (2*3>=6), relaxed does not (2*2<6)
|
||||
let md3 = "|A|B|C|D|E|F|\n|---|---|---|---|---|---|\n|x|||y|z|w|\n|a|b|c|d|e|f|";
|
||||
assert!(looks_like_partial_table(md3), "strict: 33% flagged");
|
||||
assert!(
|
||||
!looks_like_partial_table_ex(md3, true),
|
||||
"layout-assisted: 33% accepted"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn paragraph_still_rejected_when_layout_assisted() {
|
||||
// Paragraph detection is not relaxed — it's a genuine extraction issue.
|
||||
let md = "|Approval is needed from the|Acquisitions of|\n\
|
||||
|---|---|\n\
|
||||
|Treasurer if the acquisition|residential and|\n\
|
||||
|constitutes a \"significant|agricultural|\n\
|
||||
|action,\" including acquiring an|land by foreign|\n\
|
||||
|interest in different types of|persons must be|\n\
|
||||
|land where the monetary|reported to the|";
|
||||
assert!(
|
||||
looks_like_partial_table_ex(md, true),
|
||||
"paragraph rejection stays strict"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn duplicate_headers_still_rejected_when_layout_assisted() {
|
||||
let md =
|
||||
"|Position|Administration|Administration|Notes|\n|---|---|---|---|\n|Senate|24|16|x|";
|
||||
assert!(
|
||||
looks_like_partial_table_ex(md, true),
|
||||
"duplicate headers rejected even layout-assisted"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/// Analyse extracted items and rects for layout complexity.
|
||||
fn compute_layout_complexity(
|
||||
items: &[types::TextItem],
|
||||
|
||||
+302
-12
@@ -1,6 +1,6 @@
|
||||
//! Core line-to-markdown conversion loop with table/image interleaving.
|
||||
|
||||
use std::collections::HashSet;
|
||||
use std::collections::{HashMap, HashSet};
|
||||
|
||||
use crate::structure_tree::StructRole;
|
||||
use crate::types::TextLine;
|
||||
@@ -14,6 +14,139 @@ use super::postprocess::clean_markdown;
|
||||
use super::preprocess::{merge_drop_caps, merge_heading_lines};
|
||||
use super::MarkdownOptions;
|
||||
|
||||
/// Pre-scan struct heading tags to find levels that are overused — i.e., tagged on
|
||||
/// so many lines that they clearly represent body text, not real headings.
|
||||
/// Returns the set of heading levels (1–6) that should be suppressed.
|
||||
///
|
||||
/// Some PDFs (e.g. British Academy grant guidance) tag every numbered paragraph
|
||||
/// line as H2, producing hundreds of false headings. We detect this by checking
|
||||
/// if any heading level accounts for >25% of tagged lines.
|
||||
fn detect_overused_struct_heading_levels(
|
||||
lines: &[TextLine],
|
||||
struct_roles: Option<
|
||||
&std::collections::HashMap<u32, std::collections::HashMap<i64, StructRole>>,
|
||||
>,
|
||||
) -> HashSet<usize> {
|
||||
let mut overused = HashSet::new();
|
||||
let Some(roles) = struct_roles else {
|
||||
return overused;
|
||||
};
|
||||
|
||||
let mut level_counts: HashMap<usize, usize> = HashMap::new();
|
||||
let mut total = 0usize;
|
||||
|
||||
for line in lines {
|
||||
if let Some(role) = resolve_line_struct_role(line, roles) {
|
||||
total += 1;
|
||||
if let Some(level) = struct_role_heading_level(&role) {
|
||||
*level_counts.entry(level).or_insert(0) += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if total < 20 {
|
||||
return overused;
|
||||
}
|
||||
|
||||
for (&level, &count) in &level_counts {
|
||||
let ratio = count as f32 / total as f32;
|
||||
if ratio > 0.15 {
|
||||
log::debug!(
|
||||
"struct heading H{} overused: {}/{} lines ({:.0}%), suppressing",
|
||||
level,
|
||||
count,
|
||||
total,
|
||||
ratio * 100.0
|
||||
);
|
||||
overused.insert(level);
|
||||
}
|
||||
}
|
||||
|
||||
overused
|
||||
}
|
||||
|
||||
/// Pre-scan lines to find "isolated" ones: short lines with paragraph breaks both
|
||||
/// before and after. These are heading candidates even at body font size — common
|
||||
/// in academic papers ("Acknowledgements", "B.3 Prompt Engineering").
|
||||
fn find_isolated_lines(lines: &[TextLine], base_size: f32, para_threshold: f32) -> HashSet<usize> {
|
||||
let mut set = HashSet::new();
|
||||
for i in 0..lines.len() {
|
||||
let line = &lines[i];
|
||||
let plain = line.text();
|
||||
let trimmed = plain.trim();
|
||||
let word_count = trimmed.split_whitespace().count();
|
||||
if !(1..=6).contains(&word_count) || trimmed.len() <= 3 {
|
||||
continue;
|
||||
}
|
||||
let font_size = line.items.first().map(|it| it.font_size).unwrap_or(0.0);
|
||||
if font_size < base_size * 0.95 {
|
||||
continue;
|
||||
}
|
||||
if is_list_item(trimmed) || is_caption_line(trimmed) {
|
||||
continue;
|
||||
}
|
||||
|
||||
// Reject lines that look like wrapped paragraph text:
|
||||
// ends with hyphen, comma, preposition, or lowercase continuation
|
||||
let last_char = trimmed.chars().last().unwrap_or(' ');
|
||||
if last_char == '-' || last_char == ',' || last_char == ';' {
|
||||
continue;
|
||||
}
|
||||
// Last word is a common continuation word → wrapped paragraph
|
||||
let last_word = trimmed.split_whitespace().last().unwrap_or("");
|
||||
let continuation_words = [
|
||||
"the", "a", "an", "and", "or", "of", "in", "to", "for", "with", "by", "on", "at",
|
||||
"from", "as", "is", "are", "was", "were", "be", "that", "this", "their", "its", "our",
|
||||
"your", "has", "have", "had", "not",
|
||||
];
|
||||
if continuation_words.contains(&last_word.to_lowercase().as_str()) {
|
||||
continue;
|
||||
}
|
||||
|
||||
// Paragraph break BEFORE
|
||||
let break_before = if i == 0 {
|
||||
true
|
||||
} else {
|
||||
let prev = &lines[i - 1];
|
||||
prev.page != line.page || (prev.y - line.y).abs() > para_threshold
|
||||
};
|
||||
|
||||
// Paragraph break AFTER
|
||||
let break_after = if i + 1 >= lines.len() {
|
||||
true
|
||||
} else {
|
||||
let next = &lines[i + 1];
|
||||
next.page != line.page || (line.y - next.y).abs() > para_threshold
|
||||
};
|
||||
|
||||
if !break_before || !break_after {
|
||||
continue;
|
||||
}
|
||||
|
||||
set.insert(i);
|
||||
}
|
||||
|
||||
// Density guard: if too many lines on a page are "isolated", they're
|
||||
// all paragraph lines in a multi-column layout, not headings. Real
|
||||
// headings are rare — at most ~20% of lines on a page.
|
||||
let mut page_line_counts: HashMap<u32, (usize, usize)> = HashMap::new(); // (total, isolated)
|
||||
for (i, line) in lines.iter().enumerate() {
|
||||
let entry = page_line_counts.entry(line.page).or_insert((0, 0));
|
||||
entry.0 += 1;
|
||||
if set.contains(&i) {
|
||||
entry.1 += 1;
|
||||
}
|
||||
}
|
||||
for (&page, &(total, isolated)) in &page_line_counts {
|
||||
if total > 0 && isolated as f32 / total as f32 > 0.25 {
|
||||
// Too many isolated lines on this page — remove them all
|
||||
set.retain(|&i| lines[i].page != page);
|
||||
}
|
||||
}
|
||||
|
||||
set
|
||||
}
|
||||
|
||||
/// Resolve the dominant structure role for a text line by looking up its items' MCIDs.
|
||||
///
|
||||
/// Returns the first non-container role found (skipping Document/Part/Sect/Div/NonStruct/Span).
|
||||
@@ -256,6 +389,16 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
// threshold and cause every line to be treated as a paragraph break.
|
||||
let para_threshold = compute_paragraph_threshold(&lines, base_size);
|
||||
|
||||
// Pre-scan: identify isolated lines (paragraph break before AND after).
|
||||
// These are heading candidates even without bold/large font — common in
|
||||
// academic papers where section titles like "Acknowledgements" sit alone
|
||||
// between paragraphs at body font size. Inspired by opendataloader's
|
||||
// lookahead in HeadingProcessor (prevNode/nextNode context).
|
||||
let isolated_lines = find_isolated_lines(&lines, base_size, para_threshold);
|
||||
|
||||
// Detect struct heading levels that are overused (body text mistagged as headings)
|
||||
let overused_heading_levels = detect_overused_struct_heading_levels(&lines, struct_roles);
|
||||
|
||||
let mut output = String::new();
|
||||
let mut current_page = 0u32;
|
||||
let mut prev_y = f32::MAX;
|
||||
@@ -277,7 +420,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
all_content_pages.sort();
|
||||
all_content_pages.dedup();
|
||||
|
||||
for line in lines {
|
||||
for (line_idx, line) in lines.iter().enumerate() {
|
||||
// Page break
|
||||
if line.page != current_page {
|
||||
// Flush current page's remaining tables and images
|
||||
@@ -405,7 +548,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
|
||||
// Detect figure/table captions and source citations
|
||||
// These should be on their own line followed by a paragraph break
|
||||
let struct_role = struct_roles.and_then(|roles| resolve_line_struct_role(&line, roles));
|
||||
let struct_role = struct_roles.and_then(|roles| resolve_line_struct_role(line, roles));
|
||||
|
||||
// Determine if this line is code (struct-tree or font-based) for block accumulation
|
||||
let is_code_line = struct_role
|
||||
@@ -437,7 +580,10 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
// Structure roles ADD headings (e.g. same-size text tagged H2) but do NOT
|
||||
// suppress headings that the font heuristic would detect (some tagged PDFs
|
||||
// mark obvious headings as P or Span).
|
||||
let struct_heading = struct_role.as_ref().and_then(struct_role_heading_level);
|
||||
let struct_heading = struct_role
|
||||
.as_ref()
|
||||
.and_then(struct_role_heading_level)
|
||||
.filter(|level| !overused_heading_levels.contains(level));
|
||||
let heuristic_heading = if options.detect_headers
|
||||
&& plain_trimmed.len() > 3
|
||||
&& plain_trimmed.split_whitespace().count() <= 15
|
||||
@@ -445,8 +591,9 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
|
||||
detect_header_level(line_font_size, base_size, &heading_tiers).or_else(|| {
|
||||
// Rarity-based heading detection (inspired by opendataloader).
|
||||
// Score = font_rarity * 0.5 + bold * 0.3 + standalone * 0.2
|
||||
// Lines scoring above threshold are promoted to headings.
|
||||
// Heading probability scoring with lookahead context.
|
||||
// Score = rarity * 0.5 + bold * 0.3 + standalone * 0.2
|
||||
// + isolated * 0.3 (paragraph break before AND after)
|
||||
// Only consider lines at or above body font size.
|
||||
if line_font_size < base_size * 0.95 {
|
||||
return None;
|
||||
@@ -458,13 +605,21 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
let rarity = font_size_rarity(line_font_size, &font_stats);
|
||||
let all_bold = !line.items.is_empty() && line.items.iter().all(|i| i.is_bold);
|
||||
let standalone = !in_paragraph;
|
||||
let isolated = isolated_lines.contains(&line_idx);
|
||||
|
||||
let score = rarity * 0.5
|
||||
+ if all_bold { 0.3 } else { 0.0 }
|
||||
+ if standalone { 0.2 } else { 0.0 };
|
||||
+ if standalone { 0.2 } else { 0.0 }
|
||||
+ if isolated { 0.3 } else { 0.0 };
|
||||
|
||||
// Require standalone + at least one other signal
|
||||
if score >= 0.5 && standalone && word_count >= 3 {
|
||||
// Require standalone + at least one strong signal.
|
||||
// Non-bold, non-isolated lines need very high rarity (≥0.97)
|
||||
// to avoid classifying ordinary body text as headings in
|
||||
// multi-column layouts where column switches break
|
||||
// paragraph continuity and minor font-size variation
|
||||
// inflates rarity scores.
|
||||
let has_strong_signal = all_bold || isolated || (rarity >= 0.97 && word_count <= 8);
|
||||
if score >= 0.5 && standalone && word_count >= 2 && has_strong_signal {
|
||||
Some(bold_heading_level(&heading_tiers))
|
||||
} else {
|
||||
None
|
||||
@@ -652,6 +807,8 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
|
||||
// Compute the typical line spacing for paragraph break detection
|
||||
let para_threshold = compute_paragraph_threshold(&lines, base_size);
|
||||
|
||||
let isolated_lines = find_isolated_lines(&lines, base_size, para_threshold);
|
||||
|
||||
let mut output = String::new();
|
||||
let mut current_page = 0u32;
|
||||
let mut prev_y = f32::MAX;
|
||||
@@ -660,7 +817,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
|
||||
let mut last_list_x: Option<f32> = None;
|
||||
let mut prev_had_dot_leaders = false;
|
||||
|
||||
for line in lines {
|
||||
for (line_idx, line) in lines.iter().enumerate() {
|
||||
// Page break
|
||||
if line.page != current_page {
|
||||
if current_page > 0 {
|
||||
@@ -736,10 +893,12 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
|
||||
let rarity = font_size_rarity(line_font_size, &font_stats);
|
||||
let all_bold = !line.items.is_empty() && line.items.iter().all(|i| i.is_bold);
|
||||
let standalone = !in_paragraph;
|
||||
let isolated = isolated_lines.contains(&line_idx);
|
||||
let score = rarity * 0.5
|
||||
+ if all_bold { 0.3 } else { 0.0 }
|
||||
+ if standalone { 0.2 } else { 0.0 };
|
||||
if score >= 0.5 && standalone && word_count >= 3 {
|
||||
+ if standalone { 0.2 } else { 0.0 }
|
||||
+ if isolated { 0.3 } else { 0.0 };
|
||||
if score >= 0.5 && standalone && word_count >= 2 {
|
||||
return Some(bold_heading_level(&heading_tiers));
|
||||
}
|
||||
None
|
||||
@@ -1060,6 +1219,62 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_rarity_heading_requires_strong_signal() {
|
||||
// Simulate a two-column academic paper where body text lines become
|
||||
// "standalone" due to column switches. Body text at the same font
|
||||
// size as most of the document should NOT be classified as headings
|
||||
// just because of moderate rarity + standalone.
|
||||
//
|
||||
// Regression: previously, lines with rarity ~0.62 and standalone=true
|
||||
// scored 0.51 (>=0.5 threshold), producing hundreds of false ## headings.
|
||||
|
||||
// Create many body-text lines at font_size=10.9 (most common)
|
||||
let mut lines = Vec::new();
|
||||
for i in 0..20 {
|
||||
let mut item = make_item("This is ordinary body text in a paragraph.", 1, None);
|
||||
item.font_size = 10.9;
|
||||
item.y = 700.0 - i as f32 * 14.0;
|
||||
lines.push(make_line(vec![item]));
|
||||
}
|
||||
// A few lines at a slightly different size (simulating column B text)
|
||||
for i in 0..10 {
|
||||
let mut item = make_item("Another body text line from the second column.", 1, None);
|
||||
item.font_size = 11.0; // slightly different → non-zero rarity
|
||||
item.y = 700.0 - i as f32 * 14.0;
|
||||
item.x = 320.0; // right column
|
||||
lines.push(make_line(vec![item]));
|
||||
}
|
||||
// One genuine bold heading
|
||||
let mut heading_item = make_item("3 Philosophical Perspectives", 1, None);
|
||||
heading_item.font_size = 10.9;
|
||||
heading_item.is_bold = true;
|
||||
heading_item.y = 200.0;
|
||||
lines.push(make_line(vec![heading_item]));
|
||||
|
||||
let md = to_markdown_from_lines_with_tables_and_images(
|
||||
lines,
|
||||
MarkdownOptions::default(),
|
||||
HashMap::new(),
|
||||
HashMap::new(),
|
||||
&std::collections::HashSet::new(),
|
||||
None,
|
||||
);
|
||||
|
||||
// The bold heading should be detected
|
||||
assert!(
|
||||
md.contains("## 3 Philosophical Perspectives"),
|
||||
"Bold heading should be detected: {md}"
|
||||
);
|
||||
|
||||
// Body text lines should NOT be headings
|
||||
let heading_count = md.lines().filter(|l| l.starts_with("##")).count();
|
||||
assert!(
|
||||
heading_count <= 2,
|
||||
"Expected at most 2 headings but found {heading_count} in:\n{md}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_struct_role_code_multiline_accumulation() {
|
||||
let mut line1 = make_item("fn main() {", 1, Some(0));
|
||||
@@ -1102,4 +1317,79 @@ mod tests {
|
||||
"Should not have adjacent close/open fences: {md}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_overused_struct_heading_suppressed() {
|
||||
// Simulate a PDF where H2 is mistagged on body text lines.
|
||||
// 30 lines total: 5 tagged H1 (real headings), 20 tagged H2 (mistagged body),
|
||||
// 5 tagged P.
|
||||
let mut lines = Vec::new();
|
||||
let mut page_roles = HashMap::new();
|
||||
let mut mcid = 0i64;
|
||||
|
||||
for i in 0..30 {
|
||||
let mut item = make_item(&format!("Line {i}"), 1, Some(mcid));
|
||||
item.y = 700.0 - (i as f32 * 15.0);
|
||||
lines.push(make_line(vec![item]));
|
||||
|
||||
let role = if i < 5 {
|
||||
StructRole::H1
|
||||
} else if i < 25 {
|
||||
StructRole::H2
|
||||
} else {
|
||||
StructRole::P
|
||||
};
|
||||
page_roles.insert(mcid, role);
|
||||
mcid += 1;
|
||||
}
|
||||
|
||||
let mut roles = HashMap::new();
|
||||
roles.insert(1u32, page_roles);
|
||||
|
||||
let overused = detect_overused_struct_heading_levels(&lines, Some(&roles));
|
||||
// H2 is on 20/30 = 67% of lines — should be suppressed
|
||||
assert!(
|
||||
overused.contains(&2),
|
||||
"H2 should be detected as overused: {:?}",
|
||||
overused
|
||||
);
|
||||
// H1 is on 5/30 = 17% — should also be suppressed at >15% threshold
|
||||
assert!(
|
||||
overused.contains(&1),
|
||||
"H1 at 17% should also be suppressed: {:?}",
|
||||
overused
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_normal_struct_headings_not_suppressed() {
|
||||
// Normal document: a few headings, mostly body text
|
||||
let mut lines = Vec::new();
|
||||
let mut page_roles = HashMap::new();
|
||||
let mut mcid = 0i64;
|
||||
|
||||
for i in 0..50 {
|
||||
let mut item = make_item(&format!("Line {i}"), 1, Some(mcid));
|
||||
item.y = 700.0 - (i as f32 * 14.0);
|
||||
lines.push(make_line(vec![item]));
|
||||
|
||||
let role = if i % 10 == 0 {
|
||||
StructRole::H1 // 5 headings out of 50 = 10%
|
||||
} else {
|
||||
StructRole::P
|
||||
};
|
||||
page_roles.insert(mcid, role);
|
||||
mcid += 1;
|
||||
}
|
||||
|
||||
let mut roles = HashMap::new();
|
||||
roles.insert(1u32, page_roles);
|
||||
|
||||
let overused = detect_overused_struct_heading_levels(&lines, Some(&roles));
|
||||
assert!(
|
||||
overused.is_empty(),
|
||||
"No heading level should be overused: {:?}",
|
||||
overused
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
+482
-57
@@ -653,15 +653,20 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
|
||||
return None;
|
||||
}
|
||||
|
||||
// Validation 8: Check for Table of Contents pattern
|
||||
if is_table_of_contents(&cells) {
|
||||
log::debug!(" validation 8 fail: table of contents");
|
||||
// Validation 8: Reject paragraph-like content falsely detected as tables
|
||||
if is_paragraph_content(&cells) {
|
||||
log::debug!(" validation 9 fail: paragraph content");
|
||||
return None;
|
||||
}
|
||||
|
||||
// Validation 9: Reject paragraph-like content falsely detected as tables
|
||||
if is_paragraph_content(&cells) {
|
||||
log::debug!(" validation 9 fail: paragraph content");
|
||||
// Validation 9: Reject wide "index" layouts where every cell carries a
|
||||
// full "label ... page" fragment (back-of-book IRS-style indices).
|
||||
// These render poorly in any structured form; text flow is the best
|
||||
// fallback. Narrow dot-leader TOCs (2-3 cols) are kept so format.rs
|
||||
// can emit them as a per-row flat list with titles tab-joined to page
|
||||
// numbers.
|
||||
if is_inline_leader_index(&cells) {
|
||||
log::debug!(" validation 9 fail: inline-leader index");
|
||||
return None;
|
||||
}
|
||||
|
||||
@@ -898,77 +903,247 @@ fn looks_like_number(s: &str) -> bool {
|
||||
&& s.chars().any(|c| c.is_ascii_digit())
|
||||
}
|
||||
|
||||
/// Check if this looks like a Table of Contents
|
||||
/// TOCs have characteristic patterns: leader dots, page numbers, section names
|
||||
fn is_table_of_contents(cells: &[Vec<String>]) -> bool {
|
||||
/// Check if this looks like a Table of Contents (either style).
|
||||
///
|
||||
/// Used by format.rs to render TOCs as flat lists instead of markdown tables.
|
||||
pub(super) fn is_table_of_contents(cells: &[Vec<String>]) -> bool {
|
||||
is_dot_leader_toc(cells) || is_tabular_toc(cells)
|
||||
}
|
||||
|
||||
/// Dot-leader TOC: any "Chapter 1 ........ 42" style with explicit leader
|
||||
/// dots. Covers both narrow 2-3 col TOCs (where the leader is a dedicated
|
||||
/// cell) and wide indices (where each cell encodes a full "label ... page"
|
||||
/// fragment). Used by format.rs to render as a flat list.
|
||||
pub(super) fn is_dot_leader_toc(cells: &[Vec<String>]) -> bool {
|
||||
has_structural_dot_leader(cells) || is_inline_leader_index(cells)
|
||||
}
|
||||
|
||||
/// Rows with a dedicated dots-only cell flanked by label + number (2-3 col
|
||||
/// TOC layout). Format.rs handles these well via per-row flat-list
|
||||
/// rendering; they should NOT be rejected at detect time.
|
||||
fn has_structural_dot_leader(cells: &[Vec<String>]) -> bool {
|
||||
if cells.is_empty() {
|
||||
return false;
|
||||
}
|
||||
let structural_rows = cells.iter().filter(|row| row_has_dot_leader(row)).count();
|
||||
structural_rows as f32 / cells.len() as f32 >= 0.3
|
||||
}
|
||||
|
||||
let num_cols = cells[0].len();
|
||||
let mut dot_cells = 0;
|
||||
let mut page_number_cells = 0;
|
||||
let mut total_cells = 0;
|
||||
// Track which columns contain dots vs numbers to distinguish
|
||||
// TOC (dots span middle, page number at end) from data tables
|
||||
// (dots only in label column, many number columns).
|
||||
let mut dot_cols = vec![0u32; num_cols];
|
||||
let mut numeric_cols = vec![0u32; num_cols];
|
||||
|
||||
/// Wide index layout: each cell holds a full "label ... page" fragment
|
||||
/// because the column detector kept multi-column indices as single cells.
|
||||
/// These render poorly both as markdown tables (column boundaries are
|
||||
/// arbitrary) and as flat lists (each row holds 3+ separate index
|
||||
/// entries). Reject these at detect time so they fall back to the page's
|
||||
/// normal text flow.
|
||||
pub(super) fn is_inline_leader_index(cells: &[Vec<String>]) -> bool {
|
||||
let mut inline_cells = 0;
|
||||
let mut total_nonempty = 0;
|
||||
for row in cells {
|
||||
for (ci, cell) in row.iter().enumerate() {
|
||||
for cell in row {
|
||||
let trimmed = cell.trim();
|
||||
if trimmed.is_empty() {
|
||||
continue;
|
||||
}
|
||||
total_cells += 1;
|
||||
|
||||
// Check for leader dots (sequences of periods)
|
||||
// TOCs often have "........" or ". . . ." patterns
|
||||
let dot_count = trimmed.chars().filter(|&c| c == '.').count();
|
||||
let is_mostly_dots = dot_count > trimmed.len() / 2 && dot_count >= 3;
|
||||
if is_mostly_dots {
|
||||
dot_cells += 1;
|
||||
if ci < num_cols {
|
||||
dot_cols[ci] += 1;
|
||||
}
|
||||
}
|
||||
|
||||
// Check for standalone page numbers (1-4 digits, possibly with spaces)
|
||||
let digits_only: String = trimmed.chars().filter(|c| !c.is_whitespace()).collect();
|
||||
if digits_only.len() <= 4
|
||||
&& !digits_only.is_empty()
|
||||
&& digits_only.chars().all(|c| c.is_ascii_digit())
|
||||
{
|
||||
page_number_cells += 1;
|
||||
if ci < num_cols {
|
||||
numeric_cols[ci] += 1;
|
||||
}
|
||||
total_nonempty += 1;
|
||||
if cell_is_inline_leader(trimmed) {
|
||||
inline_cells += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
total_nonempty >= 4 && inline_cells as f32 / total_nonempty as f32 >= 0.25
|
||||
}
|
||||
|
||||
if total_cells == 0 {
|
||||
/// A row with a dot-leader. Accepts two layouts:
|
||||
/// 1. A dedicated dots-only cell ("....") with a text label somewhere
|
||||
/// to its left and a page number somewhere to its right.
|
||||
/// 2. A "title ... " cell (trailing leader dots glued to the title)
|
||||
/// with a page number elsewhere in the same row.
|
||||
fn row_has_dot_leader(row: &[String]) -> bool {
|
||||
let has_page_number = row.iter().any(|c| row_cell_is_page_number(c));
|
||||
|
||||
for (ci, cell) in row.iter().enumerate() {
|
||||
let trimmed = cell.trim();
|
||||
|
||||
// Pattern 1: dedicated dots-only cell.
|
||||
let dot_count = trimmed.chars().filter(|&c| c == '.').count();
|
||||
let is_mostly_dots = dot_count >= 3
|
||||
&& dot_count > trimmed.len() / 2
|
||||
&& trimmed.chars().all(|c| c == '.' || c.is_whitespace());
|
||||
if is_mostly_dots {
|
||||
let has_label_left = row[..ci].iter().any(|c| {
|
||||
let t = c.trim();
|
||||
!t.is_empty() && t.chars().any(|ch| ch.is_alphabetic())
|
||||
});
|
||||
if has_label_left && has_page_number {
|
||||
return true;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Pattern 2: cell ends with a trailing " ... " run after a label.
|
||||
if has_page_number && cell_has_trailing_leader(trimmed) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
false
|
||||
}
|
||||
|
||||
/// Cell ends with a run of ≥3 dots preceded by alphabetic text and a
|
||||
/// space — the "Title ... " layout where the leader is glued to the name.
|
||||
/// Alphabetic (not alphanumeric) so that data-table row labels like
|
||||
/// "1973 ... " do not register as titles.
|
||||
fn cell_has_trailing_leader(cell: &str) -> bool {
|
||||
let trimmed = cell.trim_end();
|
||||
if !trimmed.ends_with('.') {
|
||||
return false;
|
||||
}
|
||||
let without_dots = trimmed.trim_end_matches('.');
|
||||
let dot_run = trimmed.len() - without_dots.len();
|
||||
if dot_run < 3 {
|
||||
return false;
|
||||
}
|
||||
// Require a space before the dot run (rules out "etc..." / "Mr...") and
|
||||
// at least one alphabetic char (rules out "1973 ... " data-row labels).
|
||||
without_dots.ends_with(' ') && without_dots.trim().chars().any(|c| c.is_alphabetic())
|
||||
}
|
||||
|
||||
/// Page-number shape: single ≤4-digit integer, a ", "-separated list of
|
||||
/// ≤4-digit integers ("18, 36, 107"), or a dashed section-page ID
|
||||
/// ("A-1", "5-21"). Rejects decimal cells ("4. 0"), thousands-separated
|
||||
/// values ("189,164"), and other long numeric data that appears in
|
||||
/// statistical tables.
|
||||
fn row_cell_is_page_number(cell: &str) -> bool {
|
||||
let t = cell.trim();
|
||||
if t.is_empty() {
|
||||
return false;
|
||||
}
|
||||
if looks_like_section_page_id(t) {
|
||||
return true;
|
||||
}
|
||||
// Page list: ", " separator (with space) distinguishes real page lists
|
||||
// from thousands-separated numbers like "189,164".
|
||||
let parts: Vec<&str> = t.split(", ").collect();
|
||||
parts
|
||||
.iter()
|
||||
.all(|p| !p.is_empty() && p.len() <= 4 && p.chars().all(|c| c.is_ascii_digit()))
|
||||
}
|
||||
|
||||
/// A cell shaped like an index leader fragment. Accepts two forms:
|
||||
/// - "text ... number" — label + dots + page number in one cell
|
||||
/// - "... number" — bare leader + number (row where the label
|
||||
/// landed in a separate column)
|
||||
///
|
||||
/// Both only count if followed by pure numeric content (optionally
|
||||
/// comma-separated page lists like "127, 213").
|
||||
fn cell_is_inline_leader(cell: &str) -> bool {
|
||||
let cell = cell.trim();
|
||||
|
||||
// Find the first "..." run. Surrounding-whitespace checks below
|
||||
// reject intra-word ellipses ("etc...").
|
||||
let idx = match cell.match_indices("...").next() {
|
||||
Some((i, _)) => i,
|
||||
None => return false,
|
||||
};
|
||||
|
||||
let before = &cell[..idx];
|
||||
let after_dots = &cell[idx + 3..];
|
||||
// Allow extra dots (e.g. "....") by skipping any additional '.'
|
||||
let after = after_dots.trim_start_matches('.');
|
||||
|
||||
// Require space (or start-of-cell) before the dots and space/digit
|
||||
// after — blocks intra-word ellipses.
|
||||
let before_ok = before.is_empty() || before.ends_with(' ');
|
||||
let after_ok = after.starts_with(' ') || after.is_empty();
|
||||
if !before_ok || !after_ok {
|
||||
return false;
|
||||
}
|
||||
|
||||
// Data tables with dot leaders (e.g. "1973....") have dots concentrated
|
||||
// in one column (the label column) while many other columns contain numbers.
|
||||
// True TOCs have dots spanning the middle and one page-number column at the end.
|
||||
// If dots are confined to ≤1 column AND there are ≥3 columns with numbers,
|
||||
// this is a data table, not a TOC.
|
||||
let cols_with_dots = dot_cols.iter().filter(|&&c| c >= 2).count();
|
||||
let cols_with_numbers = numeric_cols.iter().filter(|&&c| c >= 2).count();
|
||||
if cols_with_dots <= 1 && cols_with_numbers >= 3 {
|
||||
let after_trim = after.trim();
|
||||
if after_trim.is_empty() {
|
||||
return false;
|
||||
}
|
||||
// Tail must be purely numeric/page-list content.
|
||||
let tail_numeric = after_trim
|
||||
.chars()
|
||||
.all(|c| c.is_ascii_digit() || matches!(c, ',' | ' ' | '.' | '-' | '$'))
|
||||
&& after_trim.chars().any(|c| c.is_ascii_digit());
|
||||
if !tail_numeric {
|
||||
return false;
|
||||
}
|
||||
|
||||
// If a significant portion of cells are dots or page numbers, it's likely a TOC
|
||||
let dot_ratio = dot_cells as f32 / total_cells as f32;
|
||||
let page_num_ratio = page_number_cells as f32 / total_cells as f32;
|
||||
// Either we have a label before, or the leader is bare (starts the cell)
|
||||
// — both are legitimate index fragments.
|
||||
before.chars().any(|c| c.is_alphabetic()) || before.trim().is_empty()
|
||||
}
|
||||
|
||||
// TOC typically has >15% dot cells and >10% page number cells
|
||||
dot_ratio > 0.15 || (dot_ratio > 0.05 && page_num_ratio > 0.15)
|
||||
/// Dot-less tabular TOC: tagged PDFs emit entries as rows where the first
|
||||
/// column starts with a dotted section number (e.g. "4.3.1 Something") and
|
||||
/// the last column is one or more page numbers. These have no leader dots
|
||||
/// and benefit from flat-list formatting (page numbers aligned to titles).
|
||||
pub(super) fn is_tabular_toc(cells: &[Vec<String>]) -> bool {
|
||||
if cells.is_empty() {
|
||||
return false;
|
||||
}
|
||||
let num_cols = cells[0].len();
|
||||
if num_cols < 2 || cells.len() < 4 {
|
||||
return false;
|
||||
}
|
||||
|
||||
let section_rows = cells
|
||||
.iter()
|
||||
.filter(|row| {
|
||||
row.iter()
|
||||
.find(|c| !c.trim().is_empty())
|
||||
.is_some_and(|c| starts_with_section_number(c.trim()))
|
||||
})
|
||||
.count();
|
||||
|
||||
let last_col = num_cols - 1;
|
||||
let (last_filled, last_page_num) = cells.iter().fold((0u32, 0u32), |(f, n), row| {
|
||||
let cell = row.get(last_col).map(|s| s.trim()).unwrap_or("");
|
||||
if cell.is_empty() {
|
||||
return (f, n);
|
||||
}
|
||||
let is_page_nums = cell
|
||||
.split_whitespace()
|
||||
.all(|tok| !tok.is_empty() && tok.chars().all(|c| c.is_ascii_digit()));
|
||||
(f + 1, n + if is_page_nums { 1 } else { 0 })
|
||||
});
|
||||
|
||||
let section_ratio = section_rows as f32 / cells.len() as f32;
|
||||
let page_num_last_ratio = if last_filled > 0 {
|
||||
last_page_num as f32 / last_filled as f32
|
||||
} else {
|
||||
0.0
|
||||
};
|
||||
|
||||
section_ratio >= 0.6 && last_filled >= 3 && page_num_last_ratio >= 0.7
|
||||
}
|
||||
|
||||
/// Matches dashed section-page identifiers used in technical manuals:
|
||||
/// "5-21", "A-1", "B--3", "TC-2". At least one ASCII digit is required.
|
||||
fn looks_like_section_page_id(s: &str) -> bool {
|
||||
let ok = s
|
||||
.chars()
|
||||
.all(|c| c.is_ascii_digit() || c.is_ascii_uppercase() || c == '-');
|
||||
ok && s.chars().any(|c| c.is_ascii_digit())
|
||||
}
|
||||
|
||||
/// Returns true when the leading token looks like a dotted section number:
|
||||
/// "1", "1.2", "1.2.3", "4.3.1.2" — integer components joined by dots,
|
||||
/// with at least one dot (single-number prefixes are too ambiguous).
|
||||
fn starts_with_section_number(s: &str) -> bool {
|
||||
let Some(first) = s.split_whitespace().next() else {
|
||||
return false;
|
||||
};
|
||||
let first = first.trim_end_matches('.');
|
||||
let parts: Vec<&str> = first.split('.').collect();
|
||||
if parts.len() < 2 || parts.len() > 6 {
|
||||
return false;
|
||||
}
|
||||
parts
|
||||
.iter()
|
||||
.all(|p| !p.is_empty() && p.len() <= 3 && p.chars().all(|c| c.is_ascii_digit()))
|
||||
}
|
||||
|
||||
/// Check if detected "table" cells are actually paragraph text fragments.
|
||||
@@ -1144,6 +1319,33 @@ pub(crate) fn find_first_table_row(
|
||||
continue;
|
||||
}
|
||||
|
||||
// Skip rows that have duplicate non-empty cells. These are spanning
|
||||
// super-headers (e.g., "First Degree | First Degree | Higher Degree")
|
||||
// that sit above the real column header row. Using them as the markdown
|
||||
// header produces duplicate column names that downstream validation
|
||||
// rejects. Only skip if a subsequent row looks like a better header
|
||||
// (denser fill or has data).
|
||||
if filled_count >= 2 && !has_data {
|
||||
let mut text_counts: std::collections::HashMap<&str, usize> =
|
||||
std::collections::HashMap::new();
|
||||
for cell in &filled_cells {
|
||||
*text_counts.entry(cell.trim()).or_insert(0) += 1;
|
||||
}
|
||||
let has_duplicates = text_counts.values().any(|&count| count >= 2);
|
||||
if has_duplicates {
|
||||
// Check if a later row is a better header candidate
|
||||
let has_better_below = cells.iter().skip(row_idx + 1).take(3).any(|r| {
|
||||
let next_filled = r.iter().filter(|c| !c.trim().is_empty()).count();
|
||||
let next_fill = next_filled as f32 / total_cols as f32;
|
||||
let next_numeric = r.iter().filter(|c| looks_like_number(c.trim())).count();
|
||||
next_fill >= 0.4 || next_numeric >= 2
|
||||
});
|
||||
if has_better_below {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Data rows are definitely table content
|
||||
if has_data {
|
||||
first_table_row = row_idx;
|
||||
@@ -1398,4 +1600,227 @@ mod tests {
|
||||
"data table with dot-leader labels should not be rejected as TOC"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn is_table_of_contents_rejects_dotless_toc() {
|
||||
// Tabular TOC without leader dots: first column starts with dotted
|
||||
// section numbers, last column is page numbers. Pattern from
|
||||
// Mythos system card pages 6-8.
|
||||
let cells = vec![
|
||||
vec![
|
||||
"4.3 Case studies and targeted evaluations".to_string(),
|
||||
String::new(),
|
||||
"86".to_string(),
|
||||
],
|
||||
vec![
|
||||
"4.3.1 Destructive or reckless actions".to_string(),
|
||||
"4.3.1.1 Synthetic-backend evaluation".to_string(),
|
||||
"86 86".to_string(),
|
||||
],
|
||||
vec![
|
||||
"4.3.2 Adherence to constitution".to_string(),
|
||||
"4.3.2.1 Overview".to_string(),
|
||||
"89 89".to_string(),
|
||||
],
|
||||
vec![
|
||||
"4.3.3 Honesty and hallucinations".to_string(),
|
||||
"4.3.3.1 Factual hallucinations".to_string(),
|
||||
"93 94".to_string(),
|
||||
],
|
||||
vec![
|
||||
"4.4 Capability evaluations".to_string(),
|
||||
String::new(),
|
||||
"101".to_string(),
|
||||
],
|
||||
];
|
||||
assert!(
|
||||
is_table_of_contents(&cells),
|
||||
"dot-less TOC with section numbers + page numbers should be rejected"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dot_leader_toc_accepts_short_inline_leaders() {
|
||||
// Index-style cells where the full "label ... number" pattern is
|
||||
// preserved in a single cell (IRS Publication 17 back-of-book index).
|
||||
let cells = vec![
|
||||
vec!["Child tax credit ... 235".to_string(), String::new()],
|
||||
vec!["Church employee ... 252".to_string(), String::new()],
|
||||
vec!["Citizens outside the U.S ... 6".to_string(), String::new()],
|
||||
vec![
|
||||
"Claim for refund ... 18, 36, 107".to_string(),
|
||||
String::new(),
|
||||
],
|
||||
vec!["Clergy ... 7, 52".to_string(), String::new()],
|
||||
];
|
||||
assert!(is_dot_leader_toc(&cells));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dot_leader_toc_allows_ellipsis_data_table() {
|
||||
// Data tables using "..." as a row-omission marker must not be
|
||||
// mistaken for dot-leader TOCs. Based on MCF5235RM QSPI RAM layout.
|
||||
let cells = vec![
|
||||
vec![
|
||||
"0x00".to_string(),
|
||||
"QTR0".to_string(),
|
||||
"Transmit RAM".to_string(),
|
||||
],
|
||||
vec!["0x01".to_string(), "QTR1".to_string(), String::new()],
|
||||
vec![
|
||||
"...".to_string(),
|
||||
"...".to_string(),
|
||||
"16 bits wide".to_string(),
|
||||
],
|
||||
vec!["0x0F".to_string(), "QTR15".to_string(), String::new()],
|
||||
vec![
|
||||
"0x10".to_string(),
|
||||
"QRR0".to_string(),
|
||||
"Receive RAM".to_string(),
|
||||
],
|
||||
vec!["0x11".to_string(), "QRR1".to_string(), String::new()],
|
||||
vec![
|
||||
"...".to_string(),
|
||||
"...".to_string(),
|
||||
"16 bits wide".to_string(),
|
||||
],
|
||||
vec!["0x1F".to_string(), "QRR15".to_string(), String::new()],
|
||||
];
|
||||
assert!(
|
||||
!is_dot_leader_toc(&cells),
|
||||
"ellipsis markers in a data table should not match TOC detection"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dot_leader_toc_rejects_year_row_data_table() {
|
||||
// ERP-2025 economic data tables: year labels with trailing " ... ",
|
||||
// a final " ... " column, and decimal-looking numeric cells. The
|
||||
// detection previously classified these as dot-leader TOCs and
|
||||
// routed them through flat-list formatting, destroying the grid.
|
||||
let cells = vec![
|
||||
vec![
|
||||
"1973 ... ".to_string(),
|
||||
"4. 0".to_string(),
|
||||
"1. 8".to_string(),
|
||||
"0. 4".to_string(),
|
||||
"3. 2".to_string(),
|
||||
" ... ".to_string(),
|
||||
],
|
||||
vec![
|
||||
"1974 ... ".to_string(),
|
||||
"–1. 9".to_string(),
|
||||
"–1. 6".to_string(),
|
||||
"–5. 6".to_string(),
|
||||
"2. 4".to_string(),
|
||||
" ... ".to_string(),
|
||||
],
|
||||
vec![
|
||||
"1975 ... ".to_string(),
|
||||
"2. 6".to_string(),
|
||||
"5. 1".to_string(),
|
||||
"6. 1".to_string(),
|
||||
"4. 1".to_string(),
|
||||
" ... ".to_string(),
|
||||
],
|
||||
vec![
|
||||
"1976 ... ".to_string(),
|
||||
"4. 3".to_string(),
|
||||
"5. 4".to_string(),
|
||||
"6. 4".to_string(),
|
||||
"4. 5".to_string(),
|
||||
" ... ".to_string(),
|
||||
],
|
||||
];
|
||||
assert!(
|
||||
!is_dot_leader_toc(&cells),
|
||||
"year-indexed data tables with decimal cells must not match TOC detection"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dot_leader_toc_rejects_monthly_data_table() {
|
||||
// ERP-2025 Table B-22: monthly labor-force rows with "Jan ... ",
|
||||
// "Feb ... " labels and thousands-separated cells ("189,164").
|
||||
// Previously matched TOC detection because "Jan ..." has alphabetic
|
||||
// text and "189,164" passed the page-number shape check.
|
||||
let cells = vec![
|
||||
vec![
|
||||
"2023: Jan ... ".to_string(),
|
||||
"265,962".to_string(),
|
||||
"165,871".to_string(),
|
||||
"160,152".to_string(),
|
||||
"62. 4".to_string(),
|
||||
],
|
||||
vec![
|
||||
"Feb ... ".to_string(),
|
||||
"266,112".to_string(),
|
||||
"166,263".to_string(),
|
||||
"160,301".to_string(),
|
||||
"62. 5".to_string(),
|
||||
],
|
||||
vec![
|
||||
"Mar ... ".to_string(),
|
||||
"266,272".to_string(),
|
||||
"166,690".to_string(),
|
||||
"160,824".to_string(),
|
||||
"62. 6".to_string(),
|
||||
],
|
||||
vec![
|
||||
"Apr ... ".to_string(),
|
||||
"266,443".to_string(),
|
||||
"166,678".to_string(),
|
||||
"160,962".to_string(),
|
||||
"62. 6".to_string(),
|
||||
],
|
||||
];
|
||||
assert!(
|
||||
!is_dot_leader_toc(&cells),
|
||||
"monthly labor-force rows with thousands-separated data must not match TOC detection"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn tabular_toc_requires_section_numbers_and_pages() {
|
||||
// Dot-less tabular TOC matches is_tabular_toc but not dot-leader.
|
||||
let cells = vec![
|
||||
vec![
|
||||
"4.3 Case studies".to_string(),
|
||||
String::new(),
|
||||
"86".to_string(),
|
||||
],
|
||||
vec![
|
||||
"4.3.1 Destructive actions".to_string(),
|
||||
String::new(),
|
||||
"86".to_string(),
|
||||
],
|
||||
vec![
|
||||
"4.3.2 Adherence".to_string(),
|
||||
String::new(),
|
||||
"89".to_string(),
|
||||
],
|
||||
vec!["4.3.3 Honesty".to_string(), String::new(), "93".to_string()],
|
||||
];
|
||||
assert!(is_tabular_toc(&cells));
|
||||
assert!(!is_dot_leader_toc(&cells));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn starts_with_section_number_matches_dotted() {
|
||||
assert!(starts_with_section_number("1.2"));
|
||||
assert!(starts_with_section_number("4.3.1"));
|
||||
assert!(starts_with_section_number("4.3.1.2"));
|
||||
assert!(starts_with_section_number("4.3 Case studies"));
|
||||
assert!(starts_with_section_number("2.2.5.1 Expert red teaming"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn starts_with_section_number_rejects_non_sections() {
|
||||
assert!(!starts_with_section_number("Chapter 1"));
|
||||
assert!(!starts_with_section_number("1973"));
|
||||
assert!(!starts_with_section_number("1.5M"));
|
||||
assert!(!starts_with_section_number("10.0%"));
|
||||
assert!(!starts_with_section_number(""));
|
||||
assert!(!starts_with_section_number("Hello world"));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
//! Table-to-markdown formatting and cell cleanup.
|
||||
|
||||
use super::detect_heuristic::is_table_of_contents;
|
||||
use super::Table;
|
||||
|
||||
pub fn table_to_markdown(table: &Table) -> String {
|
||||
@@ -7,6 +8,22 @@ pub fn table_to_markdown(table: &Table) -> String {
|
||||
return String::new();
|
||||
}
|
||||
|
||||
// Detect TOC on the raw cells: clean_table_cells merges rows in ways
|
||||
// that can make genuine data tables superficially resemble a TOC
|
||||
// (short numeric cells, few columns) — but the raw detection here
|
||||
// preserves the original multi-column structure and only matches the
|
||||
// true TOC pattern.
|
||||
//
|
||||
// Tables of contents render poorly as markdown tables — emit a flat
|
||||
// per-row text list instead so the page numbers stay aligned with
|
||||
// their section titles rather than drifting to a separate column.
|
||||
// Format from raw cells: continuation-row merging collapses separate
|
||||
// TOC entries (e.g. "6.2 Contamination" + "6.2.1 SWE-bench") into a
|
||||
// single line because sub-entries leave column 0 empty.
|
||||
if is_table_of_contents(&table.cells) {
|
||||
return format_toc_as_list(&table.cells, &[]);
|
||||
}
|
||||
|
||||
// Clean up the table: merge continuation rows, extract footnotes, remove empty rows
|
||||
let (cleaned_cells, footnotes) = clean_table_cells(&table.cells);
|
||||
|
||||
@@ -49,6 +66,101 @@ pub fn table_to_markdown(table: &Table) -> String {
|
||||
output
|
||||
}
|
||||
|
||||
/// Render a table-of-contents as a flat per-row text block.
|
||||
///
|
||||
/// Each row becomes one line: non-empty cells joined with spaces, and the
|
||||
/// last cell (typically a page number) is separated by a tab so the page
|
||||
/// numbers stay aligned with their titles instead of being pulled into a
|
||||
/// separate column by the column-aware reader.
|
||||
fn format_toc_as_list(cells: &[Vec<String>], footnotes: &[String]) -> String {
|
||||
let mut output = String::new();
|
||||
|
||||
for row in cells {
|
||||
let trimmed: Vec<&str> = row.iter().map(|c| c.trim()).collect();
|
||||
let last_idx = trimmed.iter().rposition(|c| !c.is_empty());
|
||||
let Some(last_idx) = last_idx else {
|
||||
continue;
|
||||
};
|
||||
|
||||
let last_cell = trimmed[last_idx];
|
||||
let last_is_page = is_page_number_cell(last_cell);
|
||||
|
||||
let (title_cells, trailing) = if last_is_page && last_idx > 0 {
|
||||
(&trimmed[..last_idx], Some(last_cell))
|
||||
} else {
|
||||
(&trimmed[..=last_idx], None)
|
||||
};
|
||||
|
||||
// Skip dots-only cells when joining the title — in a detected TOC
|
||||
// layout, a "...." cell is a leader separator, not part of the
|
||||
// entry name.
|
||||
let title = title_cells
|
||||
.iter()
|
||||
.filter(|c| !c.is_empty() && !is_dots_only(c))
|
||||
.copied()
|
||||
.collect::<Vec<_>>()
|
||||
.join(" ");
|
||||
|
||||
if title.is_empty() && trailing.is_none() {
|
||||
continue;
|
||||
}
|
||||
|
||||
if !title.is_empty() {
|
||||
output.push_str(&title);
|
||||
}
|
||||
if let Some(page) = trailing {
|
||||
if !title.is_empty() {
|
||||
output.push('\t');
|
||||
}
|
||||
output.push_str(page);
|
||||
}
|
||||
output.push('\n');
|
||||
}
|
||||
|
||||
if !footnotes.is_empty() {
|
||||
output.push('\n');
|
||||
for footnote in footnotes {
|
||||
output.push_str(footnote);
|
||||
output.push('\n');
|
||||
}
|
||||
}
|
||||
|
||||
output
|
||||
}
|
||||
|
||||
/// True when the cell looks like a page number. Accepts:
|
||||
/// - plain digit tokens: "42", "86 86"
|
||||
/// - dashed section-page IDs: "5-21", "A-1", "B--3", "TC-2" (common in
|
||||
/// technical manuals)
|
||||
fn is_page_number_cell(cell: &str) -> bool {
|
||||
let tokens: Vec<&str> = cell.split_whitespace().collect();
|
||||
if tokens.is_empty() {
|
||||
return false;
|
||||
}
|
||||
tokens.iter().all(|t| {
|
||||
if t.is_empty() || t.len() > 8 {
|
||||
return false;
|
||||
}
|
||||
let all_digits = t.chars().all(|c| c.is_ascii_digit());
|
||||
if all_digits {
|
||||
return t.len() <= 4;
|
||||
}
|
||||
// Section-page form: uppercase letters, digits, dashes; at least
|
||||
// one digit present.
|
||||
t.chars()
|
||||
.all(|c| c.is_ascii_digit() || c.is_ascii_uppercase() || c == '-')
|
||||
&& t.chars().any(|c| c.is_ascii_digit())
|
||||
})
|
||||
}
|
||||
|
||||
/// True when the cell is purely leader dots (any length ≥ 3) with optional
|
||||
/// whitespace.
|
||||
fn is_dots_only(cell: &str) -> bool {
|
||||
let t = cell.trim();
|
||||
let dots = t.chars().filter(|&c| c == '.').count();
|
||||
dots >= 3 && t.chars().all(|c| c == '.' || c.is_whitespace())
|
||||
}
|
||||
|
||||
/// Clean up table cells: merge continuation rows, extract footnotes, remove empty rows
|
||||
fn clean_table_cells(cells: &[Vec<String>]) -> (Vec<Vec<String>>, Vec<String>) {
|
||||
let mut cleaned: Vec<Vec<String>> = Vec::new();
|
||||
@@ -432,4 +544,42 @@ mod tests {
|
||||
};
|
||||
assert_eq!(table_to_markdown(&table), "");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_table_to_markdown_toc_renders_as_flat_list() {
|
||||
// A TOC-shaped table with section numbers in col 0 and page numbers
|
||||
// in the last column should render as a flat list, not a markdown
|
||||
// table, so the page numbers stay on the same line as their titles.
|
||||
let table = Table {
|
||||
columns: vec![50.0, 80.0, 300.0],
|
||||
rows: vec![500.0; 5],
|
||||
cells: vec![
|
||||
vec![
|
||||
"4.3".into(),
|
||||
"Case studies and targeted evaluations".into(),
|
||||
"86".into(),
|
||||
],
|
||||
vec![
|
||||
"4.3.1".into(),
|
||||
"Destructive or reckless actions".into(),
|
||||
"86".into(),
|
||||
],
|
||||
vec![
|
||||
"4.3.2".into(),
|
||||
"Adherence to its constitution".into(),
|
||||
"89".into(),
|
||||
],
|
||||
vec!["4.4".into(), "Capability evaluations".into(), "101".into()],
|
||||
vec!["4.5".into(), "White-box analyses".into(), "113".into()],
|
||||
],
|
||||
item_indices: vec![],
|
||||
};
|
||||
let md = table_to_markdown(&table);
|
||||
assert!(
|
||||
!md.contains("|---|"),
|
||||
"TOC should not render as a markdown table: {md}"
|
||||
);
|
||||
assert!(md.contains("4.3 Case studies and targeted evaluations\t86"));
|
||||
assert!(md.contains("4.5 White-box analyses\t113"));
|
||||
}
|
||||
}
|
||||
|
||||
+132
-12
@@ -82,33 +82,42 @@ pub(crate) fn find_column_boundaries(
|
||||
}
|
||||
}
|
||||
|
||||
let mut columns = Vec::new();
|
||||
let mut cluster_items: Vec<f32> = vec![x_positions[0]];
|
||||
// Track cluster membership: for each cluster, store the list of x positions
|
||||
let mut cluster_xs: Vec<Vec<f32>> = vec![vec![x_positions[0]]];
|
||||
|
||||
for &x in &x_positions[1..] {
|
||||
let last_cluster = cluster_xs.last().unwrap();
|
||||
// For dense columns (gap-histogram triggered), use edge-based clustering:
|
||||
// compare with the last item to avoid center-drift that merges adjacent
|
||||
// narrow columns. For normal tables, use center-based (original behavior).
|
||||
let reference = if use_edge_clustering {
|
||||
*cluster_items.last().unwrap()
|
||||
*last_cluster.last().unwrap()
|
||||
} else {
|
||||
cluster_items.iter().sum::<f32>() / cluster_items.len() as f32
|
||||
last_cluster.iter().sum::<f32>() / last_cluster.len() as f32
|
||||
};
|
||||
|
||||
if x - reference > cluster_threshold {
|
||||
let cluster_center = cluster_items.iter().sum::<f32>() / cluster_items.len() as f32;
|
||||
columns.push(cluster_center);
|
||||
cluster_items = vec![x];
|
||||
cluster_xs.push(vec![x]);
|
||||
} else {
|
||||
cluster_items.push(x);
|
||||
cluster_xs.last_mut().unwrap().push(x);
|
||||
}
|
||||
}
|
||||
|
||||
// Don't forget last cluster
|
||||
if !cluster_items.is_empty() {
|
||||
columns.push(cluster_items.iter().sum::<f32>() / cluster_items.len() as f32);
|
||||
// Numeric column merge pass: when a sparse cluster (few items, typically
|
||||
// header text) is adjacent to a dense numeric cluster and within 1.5×
|
||||
// threshold, merge them. This fixes tables where multi-line wrapped
|
||||
// headers have slightly different X positions than the data columns,
|
||||
// causing the header and data to split into separate clusters.
|
||||
let columns_before_merge = cluster_xs.len();
|
||||
if columns_before_merge >= 3 {
|
||||
cluster_xs = merge_numeric_adjacent_clusters(cluster_xs, items, cluster_threshold);
|
||||
}
|
||||
|
||||
let columns: Vec<f32> = cluster_xs
|
||||
.iter()
|
||||
.map(|xs| xs.iter().sum::<f32>() / xs.len() as f32)
|
||||
.collect();
|
||||
|
||||
// Filter columns - each should have multiple items
|
||||
let min_items_per_col = (items.len() / columns.len().max(1) / 4).max(2);
|
||||
let columns: Vec<f32> = columns
|
||||
@@ -123,8 +132,9 @@ pub(crate) fn find_column_boundaries(
|
||||
.collect();
|
||||
|
||||
log::debug!(
|
||||
" find_column_boundaries: {} columns before filter, threshold={:.1}, {} items",
|
||||
" find_column_boundaries: {} columns (merged from {}), threshold={:.1}, {} items",
|
||||
columns.len(),
|
||||
columns_before_merge,
|
||||
cluster_threshold,
|
||||
items.len()
|
||||
);
|
||||
@@ -148,6 +158,116 @@ pub(crate) fn find_column_boundaries(
|
||||
columns
|
||||
}
|
||||
|
||||
/// Check if a text string looks like a number (digits, decimals, sign, comma).
|
||||
fn is_numeric_text(s: &str) -> bool {
|
||||
let s = s.trim();
|
||||
if s.is_empty() {
|
||||
return false;
|
||||
}
|
||||
// Match patterns like: 8.23, -1.05, 9.99, 7.12, 100, 3,456.78, +5%, ---
|
||||
// But NOT: BIO, Department, Core Courses
|
||||
s.chars()
|
||||
.all(|c| c.is_ascii_digit() || c == '.' || c == ',' || c == '-' || c == '+' || c == '%')
|
||||
&& s.chars().any(|c| c.is_ascii_digit())
|
||||
}
|
||||
|
||||
/// Merge adjacent X-position clusters when one is a sparse header cluster
|
||||
/// and the other is a dense numeric data cluster. This prevents multi-line
|
||||
/// wrapped headers from splitting a logical column into two clusters.
|
||||
fn merge_numeric_adjacent_clusters(
|
||||
mut clusters: Vec<Vec<f32>>,
|
||||
items: &[(usize, &TextItem)],
|
||||
threshold: f32,
|
||||
) -> Vec<Vec<f32>> {
|
||||
// For each cluster, compute: center, item count, numeric fraction
|
||||
struct ClusterInfo {
|
||||
center: f32,
|
||||
count: usize,
|
||||
numeric_frac: f32,
|
||||
}
|
||||
|
||||
let compute_info = |xs: &[f32]| -> ClusterInfo {
|
||||
let center = xs.iter().sum::<f32>() / xs.len() as f32;
|
||||
// Count items and numeric fraction for items near this cluster center
|
||||
let mut total = 0;
|
||||
let mut numeric = 0;
|
||||
for (_, item) in items {
|
||||
if (item.x - center).abs() < threshold {
|
||||
total += 1;
|
||||
if is_numeric_text(&item.text) {
|
||||
numeric += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
ClusterInfo {
|
||||
center,
|
||||
count: total,
|
||||
numeric_frac: if total > 0 {
|
||||
numeric as f32 / total as f32
|
||||
} else {
|
||||
0.0
|
||||
},
|
||||
}
|
||||
};
|
||||
|
||||
// Merge distance: allow merging clusters that are slightly beyond the
|
||||
// original threshold. Use 1.5× threshold to catch header-vs-data splits.
|
||||
let merge_dist = threshold * 1.5;
|
||||
|
||||
// Iterate and merge adjacent pairs. Use a simple left-to-right scan.
|
||||
let mut merged = true;
|
||||
while merged {
|
||||
merged = false;
|
||||
let mut i = 0;
|
||||
while i + 1 < clusters.len() {
|
||||
let info_a = compute_info(&clusters[i]);
|
||||
let info_b = compute_info(&clusters[i + 1]);
|
||||
let dist = (info_b.center - info_a.center).abs();
|
||||
|
||||
if dist > merge_dist {
|
||||
i += 1;
|
||||
continue;
|
||||
}
|
||||
|
||||
// Determine if one cluster is sparse (header) and the other
|
||||
// is dense and numeric (data). A cluster is "sparse" if it has
|
||||
// significantly fewer items than the other.
|
||||
let (sparse, dense) = if info_a.count < info_b.count {
|
||||
(&info_a, &info_b)
|
||||
} else {
|
||||
(&info_b, &info_a)
|
||||
};
|
||||
|
||||
// Merge if the dense cluster is predominantly numeric (>50%)
|
||||
// and the sparse cluster has at most 1/3 the items of the dense one.
|
||||
let should_merge =
|
||||
dense.numeric_frac > 0.50 && sparse.count <= dense.count / 2 && sparse.count <= 5;
|
||||
|
||||
if should_merge {
|
||||
log::debug!(
|
||||
" merging column clusters: center {:.1} ({} items, {:.0}% numeric) + {:.1} ({} items, {:.0}% numeric), dist={:.1}",
|
||||
info_a.center,
|
||||
info_a.count,
|
||||
info_a.numeric_frac * 100.0,
|
||||
info_b.center,
|
||||
info_b.count,
|
||||
info_b.numeric_frac * 100.0,
|
||||
dist,
|
||||
);
|
||||
// Merge cluster i+1 into cluster i
|
||||
let next = clusters.remove(i + 1);
|
||||
clusters[i].extend(next);
|
||||
merged = true;
|
||||
// Don't increment i — check if the merged cluster can merge further
|
||||
} else {
|
||||
i += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
clusters
|
||||
}
|
||||
|
||||
/// Find row boundaries by clustering Y positions
|
||||
pub(crate) fn find_row_boundaries(items: &[(usize, &TextItem)]) -> Vec<f32> {
|
||||
let mut y_positions: Vec<f32> = items.iter().map(|(_, i)| i.y).collect();
|
||||
|
||||
+285
-1
@@ -520,6 +520,18 @@ impl ToUnicodeCMap {
|
||||
}
|
||||
}
|
||||
|
||||
/// Get the maximum source CID across all mappings (char_map + ranges).
|
||||
fn max_source_cid(&self) -> Option<u16> {
|
||||
let char_max = self.char_map.keys().copied().max();
|
||||
let range_max = self.ranges.iter().map(|&(_, end, _)| end).max();
|
||||
match (char_max, range_max) {
|
||||
(Some(a), Some(b)) => Some(a.max(b)),
|
||||
(a @ Some(_), None) => a,
|
||||
(None, b @ Some(_)) => b,
|
||||
(None, None) => None,
|
||||
}
|
||||
}
|
||||
|
||||
/// Remap a CMap that references pre-subsetting GIDs to sequential post-subsetting GIDs.
|
||||
/// Collects all source CIDs, sorts them, and reassigns to 1, 2, 3, ...
|
||||
pub fn remap_to_sequential(&self) -> ToUnicodeCMap {
|
||||
@@ -657,6 +669,81 @@ fn get_w_array_start_cid(cid_font_dict: &lopdf::Dictionary, doc: &Document) -> O
|
||||
}
|
||||
}
|
||||
|
||||
/// Return true if the CIDFont's W (widths) array explicitly covers the given CID.
|
||||
///
|
||||
/// The W array uses two formats (PDF 32000-1:2008, §9.7.4.3):
|
||||
/// 1. `c [w1 w2 ... wn]` — widths for CIDs c, c+1, ..., c+n-1
|
||||
/// 2. `c_first c_last w` — CIDs c_first..c_last all have width w
|
||||
fn w_array_covers_cid(cid_font_dict: &lopdf::Dictionary, doc: &Document, target: u16) -> bool {
|
||||
let Ok(w_obj) = cid_font_dict.get(b"W") else {
|
||||
return false;
|
||||
};
|
||||
let arr = match w_obj {
|
||||
Object::Array(arr) => arr,
|
||||
Object::Reference(r) => match doc.get_object(*r) {
|
||||
Ok(Object::Array(arr)) => arr,
|
||||
_ => return false,
|
||||
},
|
||||
_ => return false,
|
||||
};
|
||||
|
||||
let resolve_int = |o: &Object| -> Option<i64> {
|
||||
match o {
|
||||
Object::Integer(n) => Some(*n),
|
||||
Object::Reference(r) => match doc.get_object(*r) {
|
||||
Ok(Object::Integer(n)) => Some(*n),
|
||||
_ => None,
|
||||
},
|
||||
_ => None,
|
||||
}
|
||||
};
|
||||
|
||||
let resolve_arr = |o: &Object| -> Option<Vec<Object>> {
|
||||
match o {
|
||||
Object::Array(a) => Some(a.clone()),
|
||||
Object::Reference(r) => match doc.get_object(*r) {
|
||||
Ok(Object::Array(a)) => Some(a.clone()),
|
||||
_ => None,
|
||||
},
|
||||
_ => None,
|
||||
}
|
||||
};
|
||||
|
||||
let target = target as i64;
|
||||
let mut i = 0usize;
|
||||
while i < arr.len() {
|
||||
let Some(first) = resolve_int(&arr[i]) else {
|
||||
break;
|
||||
};
|
||||
i += 1;
|
||||
if i >= arr.len() {
|
||||
break;
|
||||
}
|
||||
// Peek at arr[i] to decide format.
|
||||
if let Some(widths) = resolve_arr(&arr[i]) {
|
||||
// Format 1: c [w1 ... wn]
|
||||
let last = first + widths.len() as i64 - 1;
|
||||
if target >= first && target <= last {
|
||||
return true;
|
||||
}
|
||||
i += 1;
|
||||
} else if let Some(last) = resolve_int(&arr[i]) {
|
||||
// Format 2: c_first c_last w
|
||||
i += 1;
|
||||
if i < arr.len() {
|
||||
i += 1; // skip the width value
|
||||
}
|
||||
if target >= first && target <= last {
|
||||
return true;
|
||||
}
|
||||
} else {
|
||||
// Unknown token — abort parsing safely
|
||||
break;
|
||||
}
|
||||
}
|
||||
false
|
||||
}
|
||||
|
||||
/// Extract CIDToGIDMap as a vector of GIDs (u16) indexed by CID.
|
||||
fn get_cid_to_gid_map(cid_font_dict: &lopdf::Dictionary, doc: &Document) -> Option<Vec<u16>> {
|
||||
let obj = cid_font_dict.get(b"CIDToGIDMap").ok()?;
|
||||
@@ -752,6 +839,20 @@ fn try_remap_subset_cmap(
|
||||
_ => return (cmap, None),
|
||||
};
|
||||
|
||||
// If the W array actually covers the CMap's max source CID, the CMap is
|
||||
// aligned with the font — no sequential renumbering happened. A sparse W
|
||||
// array starting at CID 0 (for .notdef) with additional high-CID entries
|
||||
// matching the CMap is the normal subset layout, not a mismatch.
|
||||
if let Some(max_cid) = cmap.max_source_cid() {
|
||||
if w_array_covers_cid(cid_font_dict, doc, max_cid) {
|
||||
debug!(
|
||||
"Subset remap skipped for obj={}: W array covers CMap max CID {}",
|
||||
obj_num, max_cid
|
||||
);
|
||||
return (cmap, None);
|
||||
}
|
||||
}
|
||||
|
||||
debug!(
|
||||
"Subset GID mismatch detected for obj={}: W starts at CID {}, CMap min CID {}. Remapping to sequential.",
|
||||
obj_num, w_start, min_cid
|
||||
@@ -1650,7 +1751,7 @@ fn merge_cmaps(mut base: ToUnicodeCMap, overlay: ToUnicodeCMap) -> ToUnicodeCMap
|
||||
///
|
||||
/// Returns true if the median CID is >= 0x41 (letter 'A'), indicating
|
||||
/// the PDF generator likely used Unicode codepoints as CIDs.
|
||||
fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) -> bool {
|
||||
pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) -> bool {
|
||||
let w_arr = match cid_font_dict.get(b"W").ok() {
|
||||
Some(Object::Array(arr)) => arr,
|
||||
_ => return false,
|
||||
@@ -2717,4 +2818,187 @@ endbfchar
|
||||
assert_eq!(remapped.unwrap().char_map.len(), 50);
|
||||
assert_eq!(fallback.unwrap().char_map.len(), 10);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_max_source_cid() {
|
||||
let cmap_content = r#"
|
||||
1 begincodespacerange
|
||||
<0000><FFFF>
|
||||
endcodespacerange
|
||||
2 beginbfchar
|
||||
<0003> <0020>
|
||||
<0031> <004E>
|
||||
endbfchar
|
||||
1 beginbfrange
|
||||
<0208> <0227> <0430>
|
||||
endbfrange
|
||||
"#;
|
||||
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
|
||||
assert_eq!(cmap.min_source_cid(), Some(0x0003));
|
||||
assert_eq!(cmap.max_source_cid(), Some(0x0227));
|
||||
}
|
||||
|
||||
/// Helper: build a minimal CIDFont dict with a W array and check coverage.
|
||||
fn cid_font_dict_with_w(w_items: Vec<lopdf::Object>) -> lopdf::Dictionary {
|
||||
let mut d = lopdf::Dictionary::new();
|
||||
d.set("W", lopdf::Object::Array(w_items));
|
||||
d
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_w_array_covers_cid_format1() {
|
||||
// Format 1: `c [w1 w2 ... wn]` — widths for CIDs c..c+n-1.
|
||||
// Mimics the 16.pdf Tahoma W array: 0[1000] 3[313] 5[401] 11[383 383] 16[363 303 382]
|
||||
let doc = Document::new();
|
||||
let d = cid_font_dict_with_w(vec![
|
||||
lopdf::Object::Integer(0),
|
||||
lopdf::Object::Array(vec![lopdf::Object::Integer(1000)]),
|
||||
lopdf::Object::Integer(3),
|
||||
lopdf::Object::Array(vec![lopdf::Object::Integer(313)]),
|
||||
lopdf::Object::Integer(5),
|
||||
lopdf::Object::Array(vec![lopdf::Object::Integer(401)]),
|
||||
lopdf::Object::Integer(11),
|
||||
lopdf::Object::Array(vec![
|
||||
lopdf::Object::Integer(383),
|
||||
lopdf::Object::Integer(383),
|
||||
]),
|
||||
lopdf::Object::Integer(16),
|
||||
lopdf::Object::Array(vec![
|
||||
lopdf::Object::Integer(363),
|
||||
lopdf::Object::Integer(303),
|
||||
lopdf::Object::Integer(382),
|
||||
]),
|
||||
lopdf::Object::Integer(570),
|
||||
lopdf::Object::Array(vec![lopdf::Object::Integer(667); 26]),
|
||||
]);
|
||||
|
||||
assert!(w_array_covers_cid(&d, &doc, 0));
|
||||
assert!(w_array_covers_cid(&d, &doc, 3));
|
||||
assert!(w_array_covers_cid(&d, &doc, 5));
|
||||
assert!(w_array_covers_cid(&d, &doc, 11));
|
||||
assert!(w_array_covers_cid(&d, &doc, 12));
|
||||
assert!(w_array_covers_cid(&d, &doc, 16));
|
||||
assert!(w_array_covers_cid(&d, &doc, 18));
|
||||
assert!(w_array_covers_cid(&d, &doc, 570));
|
||||
assert!(w_array_covers_cid(&d, &doc, 595));
|
||||
// Gaps are NOT covered
|
||||
assert!(!w_array_covers_cid(&d, &doc, 1));
|
||||
assert!(!w_array_covers_cid(&d, &doc, 4));
|
||||
assert!(!w_array_covers_cid(&d, &doc, 19));
|
||||
assert!(!w_array_covers_cid(&d, &doc, 596));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_w_array_covers_cid_format2() {
|
||||
// Format 2: `c_first c_last w` — CIDs c_first..c_last all have width w.
|
||||
let doc = Document::new();
|
||||
let d = cid_font_dict_with_w(vec![
|
||||
lopdf::Object::Integer(100),
|
||||
lopdf::Object::Integer(120),
|
||||
lopdf::Object::Integer(500),
|
||||
]);
|
||||
|
||||
assert!(w_array_covers_cid(&d, &doc, 100));
|
||||
assert!(w_array_covers_cid(&d, &doc, 110));
|
||||
assert!(w_array_covers_cid(&d, &doc, 120));
|
||||
assert!(!w_array_covers_cid(&d, &doc, 99));
|
||||
assert!(!w_array_covers_cid(&d, &doc, 121));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_w_array_covers_cid_missing_w() {
|
||||
let doc = Document::new();
|
||||
let d = lopdf::Dictionary::new();
|
||||
assert!(!w_array_covers_cid(&d, &doc, 3));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_try_remap_skipped_when_w_covers_cmap() {
|
||||
// Simulates 16.pdf: CMap's max source CID (0x0279 = 633) is explicitly
|
||||
// in the W array, so no subset-renumbering happened — remap must NOT fire.
|
||||
let cmap_content = r#"
|
||||
1 begincodespacerange
|
||||
<0000><FFFF>
|
||||
endcodespacerange
|
||||
2 beginbfchar
|
||||
<0003> <0020>
|
||||
<0031> <004E>
|
||||
endbfchar
|
||||
2 beginbfrange
|
||||
<023A> <0253> <0410>
|
||||
<0255> <0279> <042B>
|
||||
endbfrange
|
||||
"#;
|
||||
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
|
||||
|
||||
let mut doc = Document::new();
|
||||
// Build a CIDFont dict with Identity CIDToGIDMap and a W array that
|
||||
// covers CID 633 via `597 [widths...]`.
|
||||
let mut cid_font = lopdf::Dictionary::new();
|
||||
cid_font.set("CIDToGIDMap", lopdf::Object::Name(b"Identity".to_vec()));
|
||||
cid_font.set(
|
||||
"W",
|
||||
lopdf::Object::Array(vec![
|
||||
lopdf::Object::Integer(0),
|
||||
lopdf::Object::Array(vec![lopdf::Object::Integer(750)]),
|
||||
lopdf::Object::Integer(597),
|
||||
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 37]), // 597..633
|
||||
]),
|
||||
);
|
||||
let cid_font_id = doc.add_object(cid_font);
|
||||
|
||||
// Build the Type0 font dict with Identity-H + DescendantFonts ref.
|
||||
let mut font_dict = lopdf::Dictionary::new();
|
||||
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
|
||||
font_dict.set(
|
||||
"DescendantFonts",
|
||||
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
|
||||
);
|
||||
|
||||
let (primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 123);
|
||||
assert!(
|
||||
remapped.is_none(),
|
||||
"Remap must be skipped when W covers CMap max CID (this is 16.pdf)"
|
||||
);
|
||||
assert_eq!(primary.lookup(0x0003), Some(" ".to_string()));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_try_remap_fires_for_true_subset_mismatch() {
|
||||
// True mismatch: CMap has high CIDs (512-544) but W only lists low sequential CIDs.
|
||||
let cmap_content = r#"
|
||||
1 begincodespacerange
|
||||
<0000><FFFF>
|
||||
endcodespacerange
|
||||
1 beginbfrange
|
||||
<0200> <0220> <0410>
|
||||
endbfrange
|
||||
"#;
|
||||
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
|
||||
|
||||
let mut doc = Document::new();
|
||||
let mut cid_font = lopdf::Dictionary::new();
|
||||
cid_font.set("CIDToGIDMap", lopdf::Object::Name(b"Identity".to_vec()));
|
||||
cid_font.set(
|
||||
"W",
|
||||
lopdf::Object::Array(vec![
|
||||
lopdf::Object::Integer(0),
|
||||
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]), // 0..33
|
||||
]),
|
||||
);
|
||||
let cid_font_id = doc.add_object(cid_font);
|
||||
|
||||
let mut font_dict = lopdf::Dictionary::new();
|
||||
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
|
||||
font_dict.set(
|
||||
"DescendantFonts",
|
||||
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
|
||||
);
|
||||
|
||||
let (_primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 456);
|
||||
assert!(
|
||||
remapped.is_some(),
|
||||
"Remap must fire when CMap's CIDs are outside W array coverage"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
BIN
Binary file not shown.
+289
-3
@@ -4,9 +4,10 @@ use pdf_inspector::detector::{DetectionConfig, ScanStrategy};
|
||||
use pdf_inspector::extractor::group_into_lines;
|
||||
use pdf_inspector::types::TextLine;
|
||||
use pdf_inspector::{
|
||||
detect_pdf_type, extract_text, extract_text_in_regions_mem, extract_text_with_positions,
|
||||
process_pdf_mem, process_pdf_with_options, to_markdown, MarkdownOptions, PdfError, PdfOptions,
|
||||
PdfType, TextItem,
|
||||
detect_pdf_type, extract_pages_markdown_mem, extract_tables_in_regions_mem, extract_text,
|
||||
extract_text_in_regions_mem, extract_text_with_positions, process_pdf_mem,
|
||||
process_pdf_with_options, to_markdown, MarkdownOptions, PdfError, PdfOptions, PdfType,
|
||||
TextItem,
|
||||
};
|
||||
use std::collections::HashSet;
|
||||
|
||||
@@ -1344,3 +1345,288 @@ fn test_extract_regions_fast_vs_normal_comparison() {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// =========================================================================
|
||||
// extract_tables_in_regions_mem tests
|
||||
// =========================================================================
|
||||
|
||||
#[test]
|
||||
fn test_extract_tables_in_regions_table_pdf() {
|
||||
// tnagriculture has a clear table with district names and spice columns
|
||||
let buf = std::fs::read("tests/fixtures/tnagriculture_06_12.pdf").unwrap();
|
||||
let results =
|
||||
extract_tables_in_regions_mem(&buf, &[(0, vec![[0.0, 0.0, 1200.0, 1200.0]])]).unwrap();
|
||||
|
||||
assert_eq!(results.len(), 1);
|
||||
assert_eq!(results[0].regions.len(), 1);
|
||||
|
||||
let region = &results[0].regions[0];
|
||||
// Should detect a table with pipe-delimited markdown
|
||||
if !region.needs_ocr {
|
||||
assert!(
|
||||
region.text.contains('|'),
|
||||
"Table output should contain pipe delimiters"
|
||||
);
|
||||
// Should have separator row
|
||||
assert!(
|
||||
region.text.lines().any(|l| l.contains("---")),
|
||||
"Table output should contain separator row"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_tables_in_regions_non_table_region() {
|
||||
// Use a small region that likely won't contain enough items for a table
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
let results =
|
||||
extract_tables_in_regions_mem(&buf, &[(0, vec![[0.0, 0.0, 50.0, 50.0]])]).unwrap();
|
||||
|
||||
assert_eq!(results.len(), 1);
|
||||
assert_eq!(results[0].regions.len(), 1);
|
||||
|
||||
let region = &results[0].regions[0];
|
||||
// Small region with few items should fall back to needs_ocr
|
||||
assert!(
|
||||
region.needs_ocr,
|
||||
"Non-table region should set needs_ocr = true"
|
||||
);
|
||||
assert!(
|
||||
region.text.is_empty(),
|
||||
"Non-table region should have empty text"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_tables_in_regions_empty_region() {
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
let results = extract_tables_in_regions_mem(&buf, &[(0, vec![[0.0, 0.0, 0.0, 0.0]])]).unwrap();
|
||||
|
||||
assert_eq!(results.len(), 1);
|
||||
let region = &results[0].regions[0];
|
||||
assert!(region.needs_ocr);
|
||||
assert!(region.text.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_tables_in_regions_identity_h_needs_ocr() {
|
||||
let buf = std::fs::read("tests/fixtures/shinagawa_identity_h.pdf").unwrap();
|
||||
let results =
|
||||
extract_tables_in_regions_mem(&buf, &[(0, vec![[0.0, 0.0, 1200.0, 1200.0]])]).unwrap();
|
||||
|
||||
assert_eq!(results.len(), 1);
|
||||
let region = &results[0].regions[0];
|
||||
assert!(region.needs_ocr, "Identity-H font should trigger needs_ocr");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_tables_in_regions_not_a_pdf() {
|
||||
let result =
|
||||
extract_tables_in_regions_mem(b"not a pdf", &[(0, vec![[0.0, 0.0, 100.0, 100.0]])]);
|
||||
assert!(result.is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_tables_in_regions_nonexistent_page() {
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
let results =
|
||||
extract_tables_in_regions_mem(&buf, &[(9999, vec![[0.0, 0.0, 1200.0, 1200.0]])]).unwrap();
|
||||
|
||||
assert_eq!(results.len(), 1);
|
||||
let region = &results[0].regions[0];
|
||||
assert!(region.needs_ocr);
|
||||
assert!(region.text.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_bits_pilani_page4_table_detection() {
|
||||
// Page 4 (0-indexed 3) has a table with multi-line wrapped headers and
|
||||
// numeric data columns. The heuristic detector previously failed because:
|
||||
// 1. Header items at different X positions than data created extra column
|
||||
// clusters (6 cols instead of 4)
|
||||
// 2. Spanning super-header row ("First Degree | First Degree") produced
|
||||
// duplicate header cells that looks_like_partial_table_ex rejected
|
||||
let buf = std::fs::read("tests/fixtures/bits_pilani_feedback.pdf").unwrap();
|
||||
let results =
|
||||
extract_tables_in_regions_mem(&buf, &[(3, vec![[0.0, 0.0, 612.0, 792.0]])]).unwrap();
|
||||
assert_eq!(results.len(), 1);
|
||||
let region = &results[0].regions[0];
|
||||
assert!(
|
||||
!region.needs_ocr,
|
||||
"Page 4 table should be detected, got needs_ocr=true"
|
||||
);
|
||||
assert!(
|
||||
region.text.contains("BIO"),
|
||||
"Should contain department name BIO"
|
||||
);
|
||||
assert!(region.text.contains("8.23"), "Should contain numeric data");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_bits_pilani_page8_table_detection() {
|
||||
// Page 8 (0-indexed 7) has a numbered-row table that already worked.
|
||||
// Verify it still works after changes.
|
||||
let buf = std::fs::read("tests/fixtures/bits_pilani_feedback.pdf").unwrap();
|
||||
let results =
|
||||
extract_tables_in_regions_mem(&buf, &[(7, vec![[0.0, 0.0, 612.0, 792.0]])]).unwrap();
|
||||
assert_eq!(results.len(), 1);
|
||||
let region = &results[0].regions[0];
|
||||
assert!(!region.needs_ocr, "Page 8 table should still be detected");
|
||||
}
|
||||
|
||||
// =========================================================================
|
||||
// extract_pages_markdown_mem tests
|
||||
// =========================================================================
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_basic() {
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
let result = extract_pages_markdown_mem(&buf, &[0, 1]).unwrap();
|
||||
|
||||
assert_eq!(result.pages.len(), 2);
|
||||
assert_eq!(result.pages[0].page, 0);
|
||||
assert_eq!(result.pages[1].page, 1);
|
||||
// Text-based PDF should produce non-empty markdown
|
||||
assert!(!result.pages[0].markdown.is_empty());
|
||||
assert!(!result.pages[0].needs_ocr);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_page_ordering() {
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
// Request pages in non-sequential order
|
||||
let result = extract_pages_markdown_mem(&buf, &[1, 0]).unwrap();
|
||||
|
||||
assert_eq!(result.pages.len(), 2);
|
||||
// Results should match input order, not document order
|
||||
assert_eq!(result.pages[0].page, 1);
|
||||
assert_eq!(result.pages[1].page, 0);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_out_of_range() {
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
let result = extract_pages_markdown_mem(&buf, &[9999]).unwrap();
|
||||
|
||||
assert_eq!(result.pages.len(), 1);
|
||||
assert_eq!(result.pages[0].page, 9999);
|
||||
assert!(result.pages[0].markdown.is_empty());
|
||||
assert!(result.pages[0].needs_ocr);
|
||||
assert!(result.pages_needing_ocr.contains(&10000)); // 1-indexed
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_empty_pages_list() {
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
let result = extract_pages_markdown_mem(&buf, &[]).unwrap();
|
||||
assert!(result.pages.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_single_page() {
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
let result = extract_pages_markdown_mem(&buf, &[0]).unwrap();
|
||||
|
||||
assert_eq!(result.pages.len(), 1);
|
||||
assert_eq!(result.pages[0].page, 0);
|
||||
assert!(!result.pages[0].markdown.is_empty());
|
||||
assert!(!result.pages[0].needs_ocr);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_invalid_buffer() {
|
||||
let result = extract_pages_markdown_mem(b"not a pdf", &[0]);
|
||||
assert!(result.is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_gid_pages_need_ocr() {
|
||||
// shinagawa_identity_h.pdf has GID-encoded fonts
|
||||
let buf = std::fs::read("tests/fixtures/shinagawa_identity_h.pdf").unwrap();
|
||||
let result = extract_pages_markdown_mem(&buf, &[0]).unwrap();
|
||||
|
||||
assert_eq!(result.pages.len(), 1);
|
||||
assert!(result.pages[0].needs_ocr);
|
||||
assert!(result.pages_needing_ocr.contains(&1)); // 1-indexed
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_classification_with_tables() {
|
||||
// nexo-price-en.pdf is known to have tables
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
let page_count = process_pdf_mem(&buf).unwrap().page_count;
|
||||
let page_indices: Vec<u32> = (0..page_count).collect();
|
||||
let result = extract_pages_markdown_mem(&buf, &page_indices).unwrap();
|
||||
|
||||
assert!(
|
||||
!result.pages_with_tables.is_empty(),
|
||||
"nexo-price-en.pdf should have pages with tables"
|
||||
);
|
||||
assert!(result.is_complex);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_simple_pdf_no_complexity() {
|
||||
// bare_name_struct.pdf is a simple document with a heading and code block
|
||||
let buf = std::fs::read("tests/fixtures/bare_name_struct.pdf").unwrap();
|
||||
let result = extract_pages_markdown_mem(&buf, &[0]).unwrap();
|
||||
|
||||
assert!(result.pages_with_tables.is_empty());
|
||||
assert!(result.pages_with_columns.is_empty());
|
||||
assert!(!result.is_complex);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_classification_matches_process_pdf() {
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
let full = process_pdf_mem(&buf).unwrap();
|
||||
let page_count = full.page_count;
|
||||
let page_indices: Vec<u32> = (0..page_count).collect();
|
||||
let result = extract_pages_markdown_mem(&buf, &page_indices).unwrap();
|
||||
|
||||
assert_eq!(
|
||||
result.pages_with_tables, full.layout.pages_with_tables,
|
||||
"pages_with_tables should match process_pdf"
|
||||
);
|
||||
assert_eq!(
|
||||
result.pages_with_columns, full.layout.pages_with_columns,
|
||||
"pages_with_columns should match process_pdf"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_consistency_with_process_pdf() {
|
||||
let buf = std::fs::read("tests/fixtures/nexo-price-en.pdf").unwrap();
|
||||
|
||||
// Get full process_pdf output
|
||||
let full = process_pdf_mem(&buf).unwrap();
|
||||
let full_md = full.markdown.unwrap_or_default();
|
||||
|
||||
// Get per-page output for all pages
|
||||
let page_count = full.page_count;
|
||||
let page_indices: Vec<u32> = (0..page_count).collect();
|
||||
let result = extract_pages_markdown_mem(&buf, &page_indices).unwrap();
|
||||
|
||||
// Concatenated per-page markdown should contain substantial overlap with
|
||||
// the full output (exact match not expected due to header/footer stripping
|
||||
// and cross-page paragraph merging differences)
|
||||
let concat: String = result
|
||||
.pages
|
||||
.iter()
|
||||
.map(|p| p.markdown.as_str())
|
||||
.collect::<Vec<_>>()
|
||||
.join("\n");
|
||||
|
||||
// Both should be non-empty for a text-based PDF
|
||||
assert!(!full_md.is_empty());
|
||||
assert!(!concat.is_empty());
|
||||
|
||||
// The per-page version should contain at least 50% of the full content's
|
||||
// length (accounting for header/footer stripping differences)
|
||||
assert!(
|
||||
concat.len() * 2 >= full_md.len(),
|
||||
"per-page concat ({} chars) is too short vs full ({} chars)",
|
||||
concat.len(),
|
||||
full_md.len()
|
||||
);
|
||||
}
|
||||
|
||||
@@ -76,7 +76,7 @@ forms simpler, we would be happy to hear from you. You can write to the Tax Form
|
||||
|
||||
**Unreported Tips.—If you received tips of $20 or** more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you must use Form 1040 and Form 4137, Social Security and Medicare Tax on Unreported Tip Income, to report them. You may not use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act cannot use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—Get Pub. 531, Reporting** Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—If you do not keep a daily** record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
|
||||
|
||||
**Instructions (continued)**
|
||||
### Instructions (continued)
|
||||
|
||||
Use this space to total your tips for the year
|
||||
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
**Technical Information**
|
||||
##### Technical Information
|
||||
|
||||
## l T-12 SI
|
||||
|
||||
DuPont Fluorochemicals
|
||||
##### DuPont Fluorochemicals
|
||||
|
||||
#### Thermodynamic Properties
|
||||
|
||||
@@ -20,11 +20,11 @@ Tables of the thermodynamic **Units** properties of R-12 have been developed and
|
||||
|
||||
S.A., Lemmon, E.W., and Peskin, Vf = Fluid (liquid) specific volume
|
||||
A.P., NIST Standard Reference in cubic meters per kilogram Database 23, NIST thermodynamic and transport properties of Vg = Vapour (gas) specific volume refrigerants and refrigerant in cubic meters per kilogram mixtures – REFPROP version 6.01, Standard Reference Data Program, df and dg = Fluid and Vapour National Institute of Standards and (respectively) densities in Technology, 1998). kilograms per cubic meter
|
||||
H = Enthalpy (kJ/kg)
|
||||
##### H = Enthalpy (kJ/kg)
|
||||
|
||||
S = Entropy (kJ/kg.K)
|
||||
##### S = Entropy (kJ/kg.K)
|
||||
|
||||
**Physical Properties**
|
||||
##### Physical Properties
|
||||
|
||||
|Chemical Formula|CCl2F2|
|
||||
|---|---|
|
||||
|
||||
Reference in New Issue
Block a user