Compare commits

..
Author SHA1 Message Date
Abimael Martell b67b45cedc fix(tables): stop dropping body-band scripts from the candidate set
Valid: the body-font pass filtered scripts out of body_candidates
itself, so body_script_flags and its two downstream uses were dead. The
mask filters region_evidence and feeds detect_table_in_region's is_script
closure, but neither ever saw a script item because the candidate set no
longer contained any.

Consequences: a body-band sub/superscript attached to a heading-sized
anchor was dropped from the table outright rather than assigned to a
cell, so its text was lost — the opposite of what both the
body_script_flags comment ('they stay candidates') and the
detect_table_in_region docstring ('they remain eligible for cell
assignment') describe, and inconsistent with the small-font pass.

Root cause: the geometry rework removed the candidate-level filter from
the small-font pass, but the body one had been reflowed onto a single
line by rustfmt so the same edit missed it. Adding body_script_flags in
a later review then wired a mask that the surviving filter made
unreachable.

No measured change on either benchmark — pdf-evals still 20 documents
and -367 table rows, opendataloader still overall +0.0003 / mhs +0.0003
with one document changed — because the combination it affects (a
body-sized script attached to a heading-sized anchor) does not occur in
either corpus. The fix is for correctness and consistency between the
two passes, not for a score.

964 tests pass, clippy clean.
2026-08-06 16:24:14 -07:00
Abimael Martell efd875a944 test: cover the roman length bound with a nine-character token
Valid P3. The 'MMMM.' case fails on the unsupported M, not on length, so
the 8-character bound in roman_value had no coverage and could regress
silently. Added a nine-'I' token, which is rejected only by the bound,
plus an eight-'I' token that must stay exempt to pin the boundary from
both sides.
2026-08-06 13:34:09 -07:00
Abimael Martell f68cc39972 review: share roman_value so the veto exemption matches the parser
Valid. My numbering predicate accepted tokens heading::parse_numbering
rejects — lowercase 'iv)', alphabetical 'd)', over-long 'MMMM.' — because
it case-folded and allowed D and M. Anything the parser rejects is not
numbering, so exempting it let ordinary list items bypass the
dangling-verb veto and reach font-based heading promotion.

Rather than restate the grammar, roman_value is now pub(super) and the
exemption calls it, so the two cannot drift. Its rules apply as written:
uppercase I/V/X/L/C only, at most 8 characters, positive total.

Decimal numbering keeps its slightly broader acceptance (bare '2.3' with
no trailing delimiter), which is deliberate and documented — that form is
common in real headings and being permissive in a veto exemption cannot
manufacture a heading, only decline to suppress one. The roman case is
different because single letters collide with alphabetical list markers.

No measured change: opendataloader overall +0.0003 / mhs +0.0003, doc
01030000000144 still 0.732 -> 0.785. 964 tests pass, clippy clean.
2026-08-06 13:18:38 -07:00
Abimael Martell dafcff5ea3 review: exempt section-numbered lines from the dangling-verb veto
Valid ordering bug. heading.rs consults is_heading_fragment at line 282
and only applies its numbered-prefix allowance at line 288, so the veto
pre-empted it: '1. What the model implies' is sentence case and ends on
a listed verb, so it was discarded before numbering could vouch for it.

Numbering is independent evidence of a heading, so the veto now skips
any line opening with a section number.

Acceptance is deliberately a little broader than heading::parse_numbering
(which requires a trailing delimiter) because '2.3 Section Title' is
written without one, and being permissive in a veto exemption can only
avoid suppressing headings. Two guards keep it from swallowing prose:

- a bare single number needs a delimiter ('1.' yes, '3 apples' no)
- roman numerals always need one, since a leading 'I' is the pronoun far
  more often than a section number

Not reused from convert::starts_with_section_number, which deliberately
demands two components because it bypasses isolation checks — that would
reject the reviewer's single-'1.' case.

No measured change: opendataloader still overall +0.0003 / mhs +0.0003
with doc 01030000000144 at 0.732 -> 0.785, pdf-evals still 20 documents,
target case still suppressed. 964 tests pass, clippy clean.
2026-08-06 13:04:21 -07:00
Abimael Martell d883b44e1d review: gate the dangling-verb veto on sentence case, drop 'yields'
Cubic review of a5a6e8f — both findings valid.

1. 'yields' is also a plural noun. 'Bond Yields', 'Crop Yields' and
   'Dividend Yields' are real section titles in financial documents,
   which this corpus contains. Removed from the list; my claim that
   these verbs 'never end a heading in any register' was wrong for it.

2. A wrapped title-case heading whose first line ends on one of these
   verbs would be suppressed if the heading preprocessor failed to
   merge it.

Both are fixed by the same gate, which is the discriminator I was
missing: case. A heading is title case ('Bond Yields', 'The Theorem
Implies'); a stranded lead-in is sentence case ('Note that the exact
error equals', 'the method yields'). The veto now applies only when
every content word is NOT capitalized, so titles are spared regardless
of their final word.

This is also why the earlier function-word and copula variants failed:
they had no way to tell 'Rule 15. Your AGI Must Be' from 'the tax
burden should be'. Case separates those two as well.

No measured cost. opendataloader is unchanged from the previous
revision — overall +0.0003, mhs +0.0003, doc 01030000000144 still
0.732 -> 0.785 — and pdf-evals still 20 documents. 963 tests pass,
clippy clean.
2026-08-06 12:41:09 -07:00
Abimael Martell a5a6e8fbea fix(markdown): reject headings that end on a relational verb
A heading candidate ending in 'equals', 'denotes', 'implies' and the
like is the first half of a sentence, not a title. This shows up when a
block dissolves and strands its lead-in ahead of the formula it
introduced — opendataloader 01030000000144 produced

    ## Note that the exact error equals
    M - Q(h) = e - 2.7525... = -0.0342....

Deliberately a very short list. Broader variants were tried and
measured, then rejected:

- Function words (of/and/for/the): a heading that WRAPS across lines
  ends on exactly those. Destroyed real IRS Publication 17 headings —
  'Casualty and' -> 'Casualty and Theft Losses', 'Rule 10. You Must Be
  at' -> '... At Least Age 25'. 52 documents affected, -619 headings.
- Copulas and auxiliaries (is/are/be/have): same failure. 'Rule 15.
  Your AGI Must Be', 'What Medical Expenses Are' and 'When Can a Roth
  IRA Be' are real wrapped headings, while 'the tax burden should be'
  is a genuine fragment. The trailing word cannot separate them; that
  needs the next line's context, which this text-only predicate lacks.

The verbs kept never end a heading in any register, so they are safe
without context. Standalone the guard is a no-op on both benchmarks
(0 documents on opendataloader, 4 on pdf-evals with no net heading
change) — its value is as a companion to the table filter in this PR,
which is what strands these lead-ins.

Combined effect on opendataloader (200 docs, vs a control build of
main), where the table filter alone regressed:

                 table filter    + this guard
    overall        -0.0003          +0.0003
    mhs            -0.0019          +0.0003
    doc ...144     -0.063           +0.053
    doc ...144 mhs -0.203           +0.028
2026-08-06 12:20:53 -07:00
Abimael Martell 85a5150062 fix(tables): use a heading-anchored script mask in the body-font pass
Cubic review of #264: a single script mask computed with a 0.0 anchor
was applied to both passes, including body-font region qualification and
geometry. The body pass is supposed to require a heading-sized anchor —
that distinction existed before the geometry rework and was lost in it.

Why it matters: body-pass candidates are themselves body-sized
(0.85..1.05x base). A cell at the low end of that band, say 8.5pt,
sitting beside a 10.5pt label clears the inherent 'anchor >= 1.2x cell'
rule (10.2) and so was flagged as a script attachment. At body sizes a
slightly larger neighbour is a bold label or column header, not the base
of a superscript, so flagging it stripped real cells out of the region
evidence and column geometry and could lose the table entirely.

Two masks now: the small-font pass keeps the 0.0 anchor, the body pass
requires >= 1.15x base. Note the threshold only bites below base size —
for a cell at base, 1.2x-of-cell already exceeds 1.15x-of-base — which
is exactly the 0.85..1.0x band cubic identified.

Corpus: 20 documents, net -367 table rows (was -359 with the single
mask), so the body pass now keeps 8 rows of real table it had been
discarding. 958 tests pass, clippy clean.
2026-08-04 20:27:09 -07:00
Abimael Martell beba62d8e5 fix(tables): suppress script column evidence instead of dropping candidates
Reworked after reviewing the corpus diffs: the first approach removed
sub/superscripts from table candidates entirely, which had two failure
modes beyond the intended fix.

- Legitimate cell content was displaced. citizen-sr-282 is a calculator
  manual whose engineering-notation table lists M = 10^6, k = 10^3.
  Those exponents are superscripts, so they were dropped from the table
  and resurfaced elsewhere in the reading order ('9 mega 6 kilo = 10 3
  milli').
- Removing items changed the candidate geometry, so different spurious
  structure could form from what remained.

Scripts are now kept as candidates and excluded only from the geometry:
they cannot create a column (find_column_boundaries), cannot qualify a
region on their own (find_table_regions / _strict), but are still
assigned to cells. Column alignment is validated against ALL items
including scripts — validating only the non-script subset would let a
region manufacture alignment by ignoring its awkward items, which is
what block-diagram pages did.

Corpus: 20 of 186 documents, net -359 table rows, no content lost.
Token-level comparison shows the only text changes are merges in the
right direction: 'X' + '10' becomes 'X10', 'L' + 'g' becomes 'Lg' —
subscripts joining their base instead of floating free.

Remaining artifact: MCF5235RM (and its _nxp duplicate) gains a small
spurious table from a block-diagram label line, and M2019_mordeste
gains 11 rows from a 3-column grid re-detected as 2-column. Both are
borderline regions where the previous output was also wrong; documented
rather than tuned away.

805 unit + 148 integration tests pass, clippy clean.
2026-08-04 18:46:30 -07:00
Abimael Martell dbd56681b5 fix(tables): exclude script attachments and tiny numeric fragments from detection
Split out of #242 (draft) so it can be reviewed on its own evidence.

Display equations with sub/superscripts form phantom small-font table
regions: the subscripts cluster with nearby small text (footnotes, axis
labels) into fake multi-column grids. Two guards:

- Script attachment: a small-font item horizontally adjacent to a
  larger-font item at a genuine baseline offset is a sub/superscript,
  not a table cell, and is excluded from candidates. A real baseline
  offset is required so a small cell beside a larger same-baseline
  label is never filtered. Attachment targets are indexed by Y and
  scanned through a bounded window rather than a full-page sweep.
  The body-font pass applies the same exclusion but only for
  heading-sized anchors (>= 1.15x base), so body-size table cells
  beside slightly larger labels are untouched.
- Tiny numeric fragments: a <=2-row grid whose every cell is a bare
  1-2 digit number carries no tabular information. Restricted to the
  small-font pass, where the pattern is overwhelmingly exponent
  clusters; body-font numeric grids are unaffected.

Corpus impact: 20 of 186 documents, measured against a control build of
main so main's own drift is excluded. Table rows fall in 17 of 18
inspected documents and no content is lost — 2103_07786 drops all 21
rows, every one a math fragment ('|X 42 43|1|||'); Stijn_SB_doc drops
173 rows of footnote text that had been shredded into cells, with word
count slightly UP and footnote markers intact. M2019_mordeste gains 11
rows from a 3-column grid re-detected as 2-column, neither clearly
better nor worse.

Note the 20 documents is far more than the 3 that per-change ablation
suggested: that figure measured sole-cause attribution inside the
original combined PR, where other heuristics changed the same files and
masked this one. Reach and sole-cause are different measurements.

805 unit + 148 integration tests pass, clippy clean, in-repo snapshots
unchanged.
2026-08-04 18:07:34 -07:00
22 changed files with 96 additions and 3658 deletions
-1
View File
@@ -37,7 +37,6 @@ scripts/
# Test output
test_output/
.firecrawl/
# Python
__pycache__/
+3 -4
View File
@@ -5,15 +5,14 @@
If you believe you've found a security vulnerability in pdf-inspector, please
report it privately so we can fix it before public disclosure.
**Preferred:** Submit through Firecrawl's Bugcrowd vulnerability disclosure
program at <https://bugcrowd.com/engagements/firecrawl-vdp-ess>. Please include:
**Preferred:** Email **help@firecrawl.dev** with:
- A description of the issue and its impact
- Steps to reproduce (a minimal PDF or input that triggers the bug is ideal)
- The version or commit hash of pdf-inspector you tested against
**Alternative:** If you'd rather not use Bugcrowd, email
**help@firecrawl.dev** with the same details.
**Alternative:** Use GitHub's private vulnerability reporting under the
[Security tab](https://github.com/firecrawl/pdf-inspector/security/advisories/new).
We'll acknowledge your report in a timely manner and keep you updated on
remediation progress. Please do not open a public GitHub issue for security
-16
View File
@@ -83,22 +83,6 @@ for (const region of result[0].regions) {
}
```
### Async variants
`processPdf`, `classifyPdf`, and `extractPagesMarkdown` are synchronous and parse on the calling thread — in Node, that's the event loop. For a one-off call in a script that's fine, but in a server a large document can hold the loop for tens to hundreds of milliseconds.
`processPdfAsync`, `classifyPdfAsync`, and `extractPagesMarkdownAsync` take the same arguments and produce the same results, but run the parse on the libuv thread pool and return a promise, keeping the event loop free. The input buffer is copied before the call returns, so it's safe to reuse or mutate immediately:
```typescript
import { classifyPdfAsync, extractPagesMarkdownAsync } from '@firecrawl/pdf-inspector'
const classification = await classifyPdfAsync(pdf)
if (classification.pdfType === 'TextBased') {
const { pages } = await extractPagesMarkdownAsync(pdf)
// ...
}
```
## Types
```typescript
+41 -182
View File
@@ -153,7 +153,9 @@ fn to_napi_result(r: pdf_inspector::PdfProcessResult) -> PdfResult {
}
}
fn to_napi_page_ocr_reasons(reasons: Vec<pdf_inspector::PageOcrReasons>) -> Vec<PageOcrReasons> {
fn to_napi_page_ocr_reasons(
reasons: Vec<pdf_inspector::PageOcrReasons>,
) -> Vec<PageOcrReasons> {
reasons
.into_iter()
.map(|reason| PageOcrReasons {
@@ -200,31 +202,6 @@ where
}
}
// ---------------------------------------------------------------------------
// Shared implementations (single body behind sync and async entry points)
// ---------------------------------------------------------------------------
fn process_pdf_impl(bytes: &[u8], pages: Option<Vec<u32>>) -> Result<PdfResult> {
let mut opts = pdf_inspector::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = pdf_inspector::process_pdf_mem_with_options(bytes, opts)
.map_err(|e| to_napi_err(e, "process_pdf"))?;
Ok(to_napi_result(result))
}
fn classify_pdf_impl(bytes: &[u8]) -> Result<PdfClassification> {
let result =
pdf_inspector::classify_pdf_mem(bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
Ok(PdfClassification {
pdf_type: convert_pdf_type(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
}
// ---------------------------------------------------------------------------
// Public NAPI API
// ---------------------------------------------------------------------------
@@ -233,7 +210,15 @@ fn classify_pdf_impl(bytes: &[u8]) -> Result<PdfClassification> {
#[napi]
pub fn process_pdf(buffer: Buffer, pages: Option<Vec<u32>>) -> Result<PdfResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("process_pdf", move || process_pdf_impl(&bytes, pages))
catch_panic("process_pdf", move || {
let mut opts = pdf_inspector::PdfOptions::new();
if let Some(p) = pages {
opts = opts.pages(p);
}
let result = pdf_inspector::process_pdf_mem_with_options(&bytes, opts)
.map_err(|e| to_napi_err(e, "process_pdf"))?;
Ok(to_napi_result(result))
})
}
/// Fast detection only — no text extraction or markdown.
@@ -253,7 +238,16 @@ pub fn detect_pdf(buffer: Buffer) -> Result<PdfResult> {
#[napi]
pub fn classify_pdf(buffer: Buffer) -> Result<PdfClassification> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("classify_pdf", move || classify_pdf_impl(&bytes))
catch_panic("classify_pdf", move || {
let result =
pdf_inspector::classify_pdf_mem(&bytes).map_err(|e| to_napi_err(e, "classify_pdf"))?;
Ok(PdfClassification {
pdf_type: convert_pdf_type(result.pdf_type),
page_count: result.page_count,
pages_needing_ocr: result.pages_needing_ocr,
confidence: result.confidence as f64,
})
})
}
/// Extract plain text from a PDF Buffer.
@@ -639,32 +633,25 @@ pub fn extract_pages_markdown(
) -> Result<PagesExtractionResult> {
let bytes: Vec<u8> = buffer.to_vec();
catch_panic("extract_pages_markdown", move || {
extract_pages_markdown_impl(&bytes, pages.as_deref())
})
}
fn extract_pages_markdown_impl(
bytes: &[u8],
pages: Option<&[u32]>,
) -> Result<PagesExtractionResult> {
let result = pdf_inspector::extract_pages_markdown_mem(bytes, pages)
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
.pages
.into_iter()
.map(|r| PageMarkdownResult {
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
is_complex: result.is_complex,
let result = pdf_inspector::extract_pages_markdown_mem(&bytes, pages.as_deref())
.map_err(|e| to_napi_err(e, "extract_pages_markdown"))?;
Ok(PagesExtractionResult {
pages: result
.pages
.into_iter()
.map(|r| PageMarkdownResult {
page: r.page,
markdown: r.markdown,
needs_ocr: r.needs_ocr,
ocr_reason: r.ocr_reason,
})
.collect(),
pages_with_tables: result.pages_with_tables,
pages_with_columns: result.pages_with_columns,
pages_needing_ocr: result.pages_needing_ocr,
ocr_reasons_by_page: to_napi_page_ocr_reasons(result.ocr_reasons_by_page),
is_complex: result.is_complex,
})
})
}
@@ -705,131 +692,3 @@ fn to_page_region_texts(results: Vec<pdf_inspector::PageRegionResult>) -> Vec<Pa
})
.collect()
}
// ---------------------------------------------------------------------------
// Async variants (libuv thread pool via AsyncTask)
//
// The synchronous exports above parse on the calling thread, which in Node is
// the event loop. These `*Async` variants run the same shared implementations
// on the libuv thread pool and hand JavaScript a promise, so servers under
// concurrent load keep answering requests while a document parses. The sync
// exports keep their names, signatures, and behaviour.
//
// Each factory copies the input Buffer to an owned `Vec<u8>` on the calling
// (JS) thread — deliberately. JS execution is single-threaded, so no JS code
// can mutate the buffer while the synchronous part of the call copies it.
// Holding the napi `Buffer` and reading it from the worker instead would be
// zero-copy, but a caller mutating the buffer before the promise settles
// would then race the worker's reads — undefined behavior, not a recoverable
// error (a known napi-rs soundness hazard with cross-thread Buffer access).
// The copy is a one-time memcpy, negligible next to the parse it unblocks.
// ---------------------------------------------------------------------------
pub struct ProcessPdfTask {
bytes: Vec<u8>,
pages: Option<Vec<u32>>,
}
impl Task for ProcessPdfTask {
type Output = PdfResult;
type JsValue = PdfResult;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
let pages = self.pages.take();
// AssertUnwindSafe: `bytes`/`pages` are moved into the closure and
// dropped on unwind — no shared state can be observed broken.
catch_panic(
"process_pdf",
panic::AssertUnwindSafe(move || process_pdf_impl(&bytes, pages)),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`processPdf`]: same result, but the parse runs on the
/// libuv thread pool instead of the event loop and the call returns a
/// promise. The buffer is copied before the call returns, so it may be
/// reused or mutated immediately.
// ts_return_type is required: napi-rs emits `Promise<unknown>` for
// `AsyncTask<T>` returns without it.
#[napi(ts_return_type = "Promise<PdfResult>")]
pub fn process_pdf_async(buffer: Buffer, pages: Option<Vec<u32>>) -> AsyncTask<ProcessPdfTask> {
AsyncTask::new(ProcessPdfTask {
bytes: buffer.to_vec(),
pages,
})
}
pub struct ClassifyPdfTask {
bytes: Vec<u8>,
}
impl Task for ClassifyPdfTask {
type Output = PdfClassification;
type JsValue = PdfClassification;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
catch_panic(
"classify_pdf",
panic::AssertUnwindSafe(move || classify_pdf_impl(&bytes)),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`classifyPdf`]: same result, but the classification runs
/// on the libuv thread pool instead of the event loop and the call returns a
/// promise. The buffer is copied before the call returns, so it may be
/// reused or mutated immediately.
#[napi(ts_return_type = "Promise<PdfClassification>")]
pub fn classify_pdf_async(buffer: Buffer) -> AsyncTask<ClassifyPdfTask> {
AsyncTask::new(ClassifyPdfTask {
bytes: buffer.to_vec(),
})
}
pub struct ExtractPagesMarkdownTask {
bytes: Vec<u8>,
pages: Option<Vec<u32>>,
}
impl Task for ExtractPagesMarkdownTask {
type Output = PagesExtractionResult;
type JsValue = PagesExtractionResult;
fn compute(&mut self) -> Result<Self::Output> {
let bytes = std::mem::take(&mut self.bytes);
let pages = self.pages.take();
catch_panic(
"extract_pages_markdown",
panic::AssertUnwindSafe(move || extract_pages_markdown_impl(&bytes, pages.as_deref())),
)
}
fn resolve(&mut self, _env: Env, output: Self::Output) -> Result<Self::JsValue> {
Ok(output)
}
}
/// Async variant of [`extractPagesMarkdown`]: same result, but the extraction
/// runs on the libuv thread pool instead of the event loop and the call
/// returns a promise. The buffer is copied before the call returns, so it
/// may be reused or mutated immediately.
#[napi(ts_return_type = "Promise<PagesExtractionResult>")]
pub fn extract_pages_markdown_async(
buffer: Buffer,
pages: Option<Vec<u32>>,
) -> AsyncTask<ExtractPagesMarkdownTask> {
AsyncTask::new(ExtractPagesMarkdownTask {
bytes: buffer.to_vec(),
pages,
})
}
-68
View File
@@ -2,16 +2,13 @@ import { readFileSync } from 'fs';
import { strict as assert } from 'assert';
import {
processPdf,
processPdfAsync,
detectPdf,
classifyPdf,
classifyPdfAsync,
extractText,
extractTextWithPositions,
extractTextInRegions,
detectVectorGridInRegion,
extractPagesMarkdown,
extractPagesMarkdownAsync,
} from './index.js';
const fixture = readFileSync('../tests/fixtures/thermo-freon12.pdf');
@@ -127,75 +124,10 @@ assert.equal(picked.pages[0].page, 2);
assert.equal(picked.pages[1].page, 0);
console.log(' extractPagesMarkdown with pages: OK');
// --- Async variants ---
console.log('Testing async variants...');
// processPdfAsync returns a promise and matches the sync result
const asyncResultPromise = processPdfAsync(fixture);
assert.ok(asyncResultPromise instanceof Promise);
const asyncResult = await asyncResultPromise;
assert.equal(asyncResult.pdfType, result.pdfType);
assert.equal(asyncResult.pageCount, result.pageCount);
assert.equal(asyncResult.markdown, result.markdown);
console.log(' processPdfAsync: OK');
// processPdfAsync with pages
const asyncResult2 = await processPdfAsync(fixture, [1]);
assert.equal(asyncResult2.markdown, result2.markdown);
console.log(' processPdfAsync with pages: OK');
// classifyPdfAsync matches the sync result
const asyncClassified = await classifyPdfAsync(fixture);
assert.equal(asyncClassified.pdfType, classified.pdfType);
assert.equal(asyncClassified.pageCount, classified.pageCount);
assert.equal(asyncClassified.confidence, classified.confidence);
assert.deepEqual(asyncClassified.pagesNeedingOcr, classified.pagesNeedingOcr);
console.log(' classifyPdfAsync: OK');
// extractPagesMarkdownAsync matches the sync result
const asyncAllPages = await extractPagesMarkdownAsync(fixture);
assert.equal(asyncAllPages.pages.length, allPages.pages.length);
assert.deepEqual(
asyncAllPages.pages.map(p => p.markdown),
allPages.pages.map(p => p.markdown),
);
assert.equal(asyncAllPages.isComplex, allPages.isComplex);
console.log(' extractPagesMarkdownAsync: OK');
// selected pages preserve caller order
const asyncPicked = await extractPagesMarkdownAsync(fixture, [2, 0]);
assert.equal(asyncPicked.pages.length, 2);
assert.equal(asyncPicked.pages[0].page, 2);
assert.equal(asyncPicked.pages[1].page, 0);
console.log(' extractPagesMarkdownAsync with pages: OK');
// input buffer is copied at call time: mutating it immediately after the
// call must not affect the in-flight parse
const scratch = Buffer.from(fixture);
const inFlight = processPdfAsync(scratch);
scratch.fill(0);
const fromMutated = await inFlight;
assert.equal(fromMutated.markdown, result.markdown);
console.log(' processPdfAsync input copied at call time: OK');
// concurrent async calls all settle
const [c1, c2, c3] = await Promise.all([
processPdfAsync(fixture),
classifyPdfAsync(fixture),
extractPagesMarkdownAsync(fixture),
]);
assert.equal(c1.pdfType, 'TextBased');
assert.equal(c2.pdfType, 'TextBased');
assert.equal(c3.pages.length, 3);
console.log(' concurrent async calls: OK');
// --- Error handling ---
console.log('Testing error handling...');
assert.throws(() => processPdf(Buffer.from('not a pdf')), /process_pdf/);
assert.throws(() => classifyPdf(Buffer.from('')), /classify_pdf/);
await assert.rejects(processPdfAsync(Buffer.from('not a pdf')), /process_pdf/);
await assert.rejects(classifyPdfAsync(Buffer.from('')), /classify_pdf/);
await assert.rejects(extractPagesMarkdownAsync(Buffer.from('')), /extract_pages_markdown/);
console.log(' error handling: OK');
console.log('\nAll NAPI tests passed!');
+7 -7
View File
@@ -858,22 +858,22 @@
<p>Evaluated on the <a class="text-link" href="https://github.com/opendataloader-project/opendataloader-bench">opendataloader-bench</a> corpus of 200 PDFs. This comparison covers local engines without model-based PDF parsing, with OCR disabled. Higher scores are better.</p>
</div>
<div class="benchmark-card">
<div class="benchmark-top"><span><strong>200 PDFs</strong> · OpenDataLoader benchmark</span><span>Apple M4 Pro · median of 5 runs</span></div>
<div class="benchmark-top"><span><strong>200 PDFs</strong> · OpenDataLoader benchmark</span><span>Apple M4 Pro · median of 3 runs</span></div>
<div class="table-scroll">
<table aria-label="PDF extraction benchmark results">
<thead>
<tr><th>Engine</th><th>Overall</th><th>Reading order</th><th>Tables</th><th>Headings</th><th>Complete run</th></tr>
</thead>
<tbody>
<tr class="highlight"><td>pdf-inspector</td><td>0.875</td><td>0.915</td><td>0.814</td><td>0.788</td><td>0.470s</td></tr>
<tr><td>LiteParse</td><td>0.873</td><td>0.913</td><td>0.693</td><td>0.811</td><td>0.750s</td></tr>
<tr><td>OpenDataLoader</td><td>0.831</td><td>0.902</td><td>0.489</td><td>0.739</td><td>2.569s</td></tr>
<tr><td>PyMuPDF4LLM</td><td>0.735</td><td>0.886</td><td>0.401</td><td>0.424</td><td>17.117s</td></tr>
<tr><td>MarkItDown</td><td>0.589</td><td>0.844</td><td>0.273</td><td>0.000</td><td>16.165s</td></tr>
<tr class="highlight"><td>pdf-inspector</td><td>0.875</td><td>0.915</td><td>0.814</td><td>0.788</td><td>2.8s</td></tr>
<tr><td>LiteParse</td><td>0.870</td><td>0.908</td><td>0.693</td><td>0.811</td><td>13.9s</td></tr>
<tr><td>OpenDataLoader</td><td>0.843</td><td>0.912</td><td>0.489</td><td>0.760</td><td>9.8s</td></tr>
<tr><td>PyMuPDF4LLM</td><td>0.735</td><td>0.886</td><td>0.401</td><td>0.424</td><td>15.5s</td></tr>
<tr><td>MarkItDown</td><td>0.583</td><td>0.879</td><td>0.000</td><td>0.000</td><td>6.7s</td></tr>
</tbody>
</table>
</div>
<div class="benchmark-note">Refreshed July 31, 2026. Scores use the benchmarks NID, TEDS, and MHS evaluators; speed is the median of five alternating or rotating complete corpus runs after an excluded warm-up. <a class="text-link" href="https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results">Versions and raw artifacts</a>.</div>
<div class="benchmark-note">Refreshed July 16, 2026. Scores use the benchmarks NID, TEDS, and MHS evaluators.</div>
</div>
<div class="best-fit">
<strong>Best fit</strong>
+1 -65
View File
@@ -1659,15 +1659,7 @@ fn hex_val(b: u8) -> Option<u8> {
/// Standard page: 612x792 points (US Letter) = ~485,000 sq points
/// At 2x resolution that's ~1.9M pixels, so we use 250K pixels as threshold
/// (accounting for varying DPI and page sizes)
/// Returns `(has_images, total_image_area, has_template_image)` for a page.
/// `has_template_image` means a single large (>50% page coverage)
/// background image — the signal `classify_pdf`/`detect_pdf_type` uses to
/// route a page to OCR regardless of any incidental native text drawn over
/// it. Exposed at crate visibility so extraction-side per-page `needs_ocr`
/// computation (`extract_pages_markdown_mem`) can consult the same signal
/// instead of maintaining its own, independent notion of "needs OCR" that
/// can silently disagree with detection — see #227.
pub(crate) fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
// Threshold: image covering roughly half a page at 150+ DPI
// 612 * 792 / 2 * (150/72)^2 ≈ 1M pixels, but we'll be conservative
const TEMPLATE_IMAGE_THRESHOLD: u64 = 500_000; // 500K pixels
@@ -1749,62 +1741,6 @@ pub(crate) fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u
(has_images, total_area, has_template_image)
}
/// Computes both `(needs_ocr_for_template_image, has_vector_text)` for a
/// page from a single shared `analyze_page_content` pass — that call
/// decompresses and scans every content stream (page + XObjects) plus
/// image coverage, so `extract_pages_markdown_mem` must not invoke it
/// twice per page (once per signal) the way `detect_from_document` avoids
/// by caching its per-page `PageAnalysis`.
///
/// `needs_ocr_for_template_image` is true when a page's template image
/// should be treated as a scan needing OCR — a single full-page background
/// image with little/no real text — rather than a text page that happens
/// to carry a watermark, letterhead, or figure. Mirrors the two distinct
/// signals classification uses to route a template-image page to OCR:
///
/// 1. `looks_like_scan`: image_count <= 1, few text operators (<50), and
/// low alphanumeric diversity in raw string operands (unless decodable
/// CID/ToUnicode fonts explain that away) — the gate used for
/// `pages_with_template_images` and Mixed-type per-page routing.
/// 2. Insufficient real text volume, using `DetectionConfig::default()`'s
/// `min_text_ops_per_page` (3) — the same threshold Mixed-type per-page
/// routing applies via `text_operator_count < config.min_text_ops_per_page
/// && has_images` (simplified here since a template image implies
/// `has_images`). Deliberately *not* the higher `effective_min_ops`
/// floor (`min_text_ops_per_page.max(10)`) that whole-document
/// `PdfType::ImageBased`/`Scanned` classification uses for
/// `pages_with_text` — that's a cross-page aggregate decision this
/// per-page function has no way to replicate exactly, and the lower
/// per-page threshold is the one a single page's own signals can
/// actually agree with.
///
/// `has_vector_text` is true when a page has vector-outlined text (glyphs
/// drawn as paths rather than shown via text-showing operators) —
/// `detect_from_document`'s Mixed-type per-page routing always sends
/// these pages to OCR, independent of any template-image check, since
/// outlined glyphs can't be extracted as text at all.
///
/// Exposed at crate visibility so `extract_pages_markdown_mem` can apply
/// the same gates classification needs elsewhere instead of treating the
/// raw signals alone as sufficient — see #227/#231.
pub(crate) fn page_ocr_signals(doc: &Document, page_id: ObjectId) -> (bool, bool) {
let analysis = analyze_page_content(doc, page_id);
let needs_ocr_for_template_image = if !analysis.has_template_image {
false
} else {
let alphanum_low = analysis.unique_alphanum_chars < 10
&& !(analysis.has_decodable_text_fonts && analysis.text_operator_count >= 10);
let looks_like_scan =
analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_low;
let insufficient_text =
analysis.text_operator_count < DetectionConfig::default().min_text_ops_per_page;
looks_like_scan || insufficient_text
};
(needs_ocr_for_template_image, analysis.has_vector_text)
}
/// Recursively collect image dimensions from XObject resources,
/// including images nested inside Form XObjects.
fn collect_images_from_resources(
+5 -311
View File
@@ -2,51 +2,11 @@
use crate::types::{ItemType, TextItem};
use lopdf::{Document, Object, ObjectId};
use std::collections::{HashMap, HashSet};
use std::collections::HashMap;
use super::fonts::{resolve_array, resolve_dict};
use super::get_number;
/// Upper bound on the number of form-field nodes visited during a single
/// `extract_form_fields` pass. A crafted PDF can chain thousands of distinct
/// `/Kids` fields to blow the stack even without an outright reference cycle,
/// so we cap total traversal work in addition to detecting cycles.
const MAX_FORM_FIELD_NODES: usize = 100_000;
/// Upper bound on `/Kids` recursion depth. Real AcroForm hierarchies are only
/// a few levels deep (fields → child fields → widgets); a crafted PDF can chain
/// tens of thousands of distinct fields into a linear `/Kids` list that would
/// overflow the stack via depth-first recursion long before the node budget is
/// reached. This depth cap bounds the stack independently of total node count.
const MAX_FORM_FIELD_DEPTH: usize = 100;
/// Traversal budget for the AcroForm field walk. Bounds both the number of
/// distinct nodes visited *and* the total number of `/Fields`/`/Kids` entries
/// examined.
///
/// Counting `visited` alone is not enough: invalid entries (non-references) and
/// duplicate references never grow `visited`, so an oversized array full of them
/// would iterate to completion no matter how large. Charging every examined
/// entry against the same budget makes it a real cap on traversal work.
pub(crate) struct FieldWalkBudget {
visited: HashSet<ObjectId>,
examined: usize,
}
impl FieldWalkBudget {
fn new() -> Self {
Self {
visited: HashSet::new(),
examined: 0,
}
}
/// True once the budget is spent; callers must stop iterating and recursing.
fn exhausted(&self) -> bool {
self.visited.len() >= MAX_FORM_FIELD_NODES || self.examined >= MAX_FORM_FIELD_NODES
}
}
pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> Vec<TextItem> {
let mut links = Vec::new();
@@ -186,12 +146,9 @@ pub(crate) fn extract_form_fields(
Err(_) => return items,
};
// Borrow the array rather than cloning it: a crafted `/Fields` can be huge,
// and cloning would pay an O(n) allocation/copy before the budget check
// below can stop the work.
let fields = match acroform.get(b"Fields") {
Ok(obj) => match resolve_array(doc, obj) {
Some(arr) => arr,
Some(arr) => arr.clone(),
None => return items,
},
Err(_) => return items,
@@ -201,19 +158,7 @@ pub(crate) fn extract_form_fields(
}
let annotation_pages = annotation_page_map(doc, page_map);
// Bound the walk so a crafted PDF cannot send us into unbounded recursion
// via a `/Kids` cycle, a deep chain, or an oversized array of invalid or
// duplicate entries.
let mut budget = FieldWalkBudget::new();
for field_obj in fields {
// Stop once the budget is spent so a `/Fields` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op. Charge
// every entry (including invalid ones) against the budget.
if budget.exhausted() {
break;
}
budget.examined += 1;
for field_obj in &fields {
if let Ok(field_ref) = field_obj.as_reference() {
walk_form_fields(
doc,
@@ -223,8 +168,6 @@ pub(crate) fn extract_form_fields(
page_map,
&annotation_pages,
&mut items,
&mut budget,
0,
);
}
}
@@ -259,7 +202,6 @@ fn annotation_page_map(
}
/// Recursively walk the form field tree, extracting leaf field values.
#[allow(clippy::too_many_arguments)]
pub(crate) fn walk_form_fields(
doc: &Document,
field_id: ObjectId,
@@ -268,22 +210,7 @@ pub(crate) fn walk_form_fields(
page_map: &HashMap<ObjectId, u32>,
annotation_pages: &HashMap<ObjectId, u32>,
items: &mut Vec<TextItem>,
budget: &mut FieldWalkBudget,
depth: usize,
) {
// Guard against `/Kids` cycles and pathologically large field trees.
// Exceeding the depth cap means the chain is too deep to be a legitimate
// form (and would overflow the stack); an exhausted budget means the tree is
// too large. Both checks run *before* inserting so the visited set can never
// grow past the budget.
if depth > MAX_FORM_FIELD_DEPTH || budget.exhausted() {
return;
}
// Revisiting an object ID means we hit a `/Kids` cycle.
if !budget.visited.insert(field_id) {
return;
}
let field_dict = match doc.get_dictionary(field_id) {
Ok(d) => d,
Err(_) => return,
@@ -314,19 +241,9 @@ pub(crate) fn walk_form_fields(
// Check for /Kids — if present, recurse into children
if let Ok(kids_obj) = field_dict.get(b"Kids") {
// Iterate the borrowed array directly — cloning a crafted, oversized
// `/Kids` would allocate and copy every entry before the budget check
// below could stop the work.
if let Some(kids) = resolve_array(doc, kids_obj) {
for kid in kids {
// Stop once the budget is spent so a `/Kids` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op.
// Charge every entry (including invalid/duplicate ones) against
// the budget so this is a true traversal-work cap.
if budget.exhausted() {
break;
}
budget.examined += 1;
let kids = kids.clone();
for kid in &kids {
if let Ok(kid_ref) = kid.as_reference() {
walk_form_fields(
doc,
@@ -336,8 +253,6 @@ pub(crate) fn walk_form_fields(
page_map,
annotation_pages,
items,
budget,
depth + 1,
);
}
}
@@ -496,225 +411,4 @@ mod tests {
assert_eq!(items[0].page, 2);
assert_eq!(items[0].text, "customer: Alice");
}
#[test]
fn kids_self_cycle_does_not_overflow_stack() {
// A crafted AcroForm field that lists itself in `/Kids` must not send
// the traversal into unbounded recursion.
let mut doc = Document::new();
let field_id = doc.new_object_id();
doc.set_object(
field_id,
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("loop"),
"Kids" => vec![Object::Reference(field_id)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
// Completes (rather than overflowing the stack) and yields no items.
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn kids_mutual_cycle_terminates() {
// Two fields that reference each other via `/Kids` form a cycle that
// must also terminate.
let mut doc = Document::new();
let field_a = doc.new_object_id();
let field_b = doc.new_object_id();
doc.set_object(
field_a,
dictionary! {
"T" => Object::string_literal("a"),
"Kids" => vec![Object::Reference(field_b)],
},
);
doc.set_object(
field_b,
dictionary! {
"T" => Object::string_literal("b"),
"Kids" => vec![Object::Reference(field_a)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_a)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn deep_acyclic_kids_chain_does_not_overflow_stack() {
// A long chain of *distinct* fields (no cycle) must also terminate:
// the visited set alone would still recurse to the chain length, so
// the depth cap is what prevents a stack overflow here.
let mut doc = Document::new();
let n = MAX_FORM_FIELD_DEPTH * 500;
let ids: Vec<ObjectId> = (0..=n).map(|_| doc.new_object_id()).collect();
for i in 0..n {
doc.set_object(
ids[i],
dictionary! {
"FT" => "Tx",
"Kids" => vec![Object::Reference(ids[i + 1])],
},
);
}
// Leaf carries a value; it sits far below the depth cap so it is never
// reached, proving traversal stops early rather than crashing.
doc.set_object(
ids[n],
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("leaf"),
"V" => Object::string_literal("x"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(ids[0])],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn wide_tree_traversal_stops_at_node_budget() {
// A single field with a `/Kids` array wider than the node budget must
// stop traversal at the cap rather than growing `visited` (and the work)
// without bound. Each processed leaf emits one item, so the item count
// is bounded by the budget and reaches right up to it (a couple of
// slots go to the root and the boundary node charged against the cap).
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let kids: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
// Extraction stops at the budget: bounded above by the cap, and it gets
// right up to it (allowing a small delta for the root/boundary nodes
// charged against the budget).
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn wide_top_level_fields_stop_at_node_budget() {
// A top-level `/Fields` array wider than the budget must also stop at
// the cap: the item count is bounded by the budget and reaches right up
// to it.
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let fields: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => fields,
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn duplicate_and_invalid_kids_entries_stop_at_budget() {
// Duplicate references and non-reference junk never grow `visited`, so
// without charging examined entries against the budget an oversized
// array of them would iterate to completion. The walk must still
// terminate and extract the single real leaf exactly once.
let mut doc = Document::new();
let leaf_id = doc.new_object_id();
doc.set_object(
leaf_id,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
// A `/Kids` array far wider than the budget: half duplicate references
// to the same leaf, half invalid (null) entries.
let mut kids: Vec<Object> = Vec::new();
for i in 0..(MAX_FORM_FIELD_NODES * 2) {
if i % 2 == 0 {
kids.push(Object::Reference(leaf_id));
} else {
kids.push(Object::Null);
}
}
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert_eq!(items.len(), 1);
}
}
+3 -40
View File
@@ -4566,13 +4566,9 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
}
}
// Try to parse uniXXXX format.
// Use `get` rather than a byte-length check + slice: `name` can contain
// non-ASCII bytes (e.g. U+FFFD from lossy UTF-8 decoding of an attacker
// controlled /Differences name), so byte index 7 may not be a char
// boundary and `&name[3..7]` would panic.
if let Some(hex) = name.strip_prefix("uni").and_then(|rest| rest.get(..4)) {
if let Ok(code) = u32::from_str_radix(hex, 16) {
// Try to parse uniXXXX format
if name.starts_with("uni") && name.len() >= 7 {
if let Ok(code) = u32::from_str_radix(&name[3..7], 16) {
// Strip PUA F000 offset: uniF0XX → U+00XX (Windows Symbol encoding convention)
let code = if (0xF000..=0xF0FF).contains(&code) {
code - 0xF000
@@ -4592,36 +4588,3 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
None
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn uni_hex_parsing() {
assert_eq!(glyph_to_char("uni0041"), Some('A'));
assert_eq!(glyph_to_char("uni00e9"), Some('\u{00e9}'));
// PUA F0xx symbol-encoding offset is stripped.
assert_eq!(glyph_to_char("uniF041"), Some('A'));
}
#[test]
fn u_hex_parsing() {
assert_eq!(glyph_to_char("u0041"), Some('A'));
assert_eq!(glyph_to_char("u1F600"), Some('\u{1F600}'));
}
#[test]
fn non_ascii_uni_name_does_not_panic() {
// A crafted /Differences name like `/uni#80#80#80#80` decodes via
// from_utf8_lossy into "uni" followed by four U+FFFD replacements.
// Byte index 7 lands mid-character, so a naive `&name[3..7]` slice
// would panic. It must be handled gracefully instead.
let crafted = format!("uni{0}{0}{0}{0}", '\u{FFFD}');
assert_eq!(glyph_to_char(&crafted), None);
// Assorted non-ASCII bytes right after the "uni" prefix.
assert_eq!(glyph_to_char("uni\u{FFFD}bc"), None);
assert_eq!(glyph_to_char("uni\u{00e9}00"), None);
}
}
+7 -311
View File
@@ -491,14 +491,7 @@ pub fn extract_pages_markdown_mem(
// Tables need the original numeric cells; columns use folio-cleaned
// evidence so removed page numbers cannot create false layout metadata.
let chart_regions = markdown::chart_regions_by_page(&all_items, &all_rects, &all_lines);
let complexity = compute_layout_complexity_with_chart_regions(
&all_items,
&filtered_items,
&all_rects,
&all_lines,
&chart_regions,
);
let complexity = compute_layout_complexity(&all_items, &filtered_items, &all_rects, &all_lines);
// Compute font stats from full document (cross-page consistency).
let font_stats = markdown::analysis::calculate_font_stats_from_items(&filtered_items);
@@ -516,7 +509,6 @@ pub fn extract_pages_markdown_mem(
let mut results = Vec::with_capacity(pages_slice.len());
let mut pages_needing_ocr = Vec::new();
let mut ocr_reasons_by_page = BTreeMap::new();
let lopdf_pages = doc.get_pages();
for &page_0idx in pages_slice {
// Out-of-range pages → empty + needs_ocr
@@ -550,25 +542,6 @@ pub fn extract_pages_markdown_mem(
let has_gid = gid_pages.contains(&page_1idx);
let has_text_quality_issue = text_quality.pages_needing_ocr.contains(&page_1idx);
// A page can extract cleanly (no decoding issues, non-empty text)
// while still being fundamentally a scan: a full-page raster with
// a little genuine native text drawn over it (a header, a stamp, a
// cover-sheet annotation). Text-quality signals alone can't see
// that — consult the same "large background image" signal
// classify_pdf/detect_pdf_type already uses, so the two APIs can't
// silently disagree on whether a page needs OCR. See #227.
// Also covers vector-outlined text (glyphs drawn as paths, not
// shown via a text-showing operator): a hybrid page with real
// embedded-font body text elsewhere would otherwise still extract
// non-empty, non-garbled markdown and miss OCR routing entirely.
// detect_from_document's Mixed-type per-page routing always sends
// these pages to OCR; mirror that here too. Both signals share one
// analyze_page_content pass — see page_ocr_signals's doc comment.
let (has_template_image, has_vector_text) = lopdf_pages
.get(&page_1idx)
.map(|&page_id| detector::page_ocr_signals(&doc, page_id))
.unwrap_or((false, false));
// Build markdown with document-wide font stats
let options = MarkdownOptions {
base_font_size: Some(font_stats.most_common_size),
@@ -592,7 +565,6 @@ pub fn extract_pages_markdown_mem(
page_count,
prefiltered_page_number_pages: Some(&removed_page_number_pages),
prefiltered_page_number_mask: Some(&page_number_removal_mask),
precomputed_chart_regions: Some(&chart_regions),
},
)
};
@@ -606,20 +578,10 @@ pub fn extract_pages_markdown_mem(
OCR_REASON_SUSPECTED_GARBLED_TEXT,
);
}
if has_template_image {
add_ocr_reason(&mut ocr_reasons_by_page, page_1idx, OCR_REASON_SCANNED);
}
if has_vector_text {
add_ocr_reason(&mut ocr_reasons_by_page, page_1idx, OCR_REASON_VECTOR_TEXT);
}
let ocr_reason = page_ocr_reason(&ocr_reasons_by_page, page_1idx);
let needs_ocr = ocr_reason.is_some()
|| md.trim().is_empty()
|| has_gid
|| is_garbage_text(&md)
|| has_template_image
|| has_vector_text;
let needs_ocr =
ocr_reason.is_some() || md.trim().is_empty() || has_gid || is_garbage_text(&md);
if needs_ocr {
pages_needing_ocr.push(page_1idx);
@@ -3553,7 +3515,6 @@ fn repair_pdf_container_candidates(buf: &[u8]) -> Vec<Vec<u8>> {
let mut candidates = Vec::new();
add_repair_candidate(&mut candidates, append_missing_eof_marker(buf), buf);
add_repair_candidate(&mut candidates, recover_startxref_pointer(buf), buf);
let stripped = strip_leading_pdf_container_bytes(buf);
if let Some(stripped_buf) = stripped.as_deref() {
@@ -3563,112 +3524,11 @@ fn repair_pdf_container_candidates(buf: &[u8]) -> Vec<Vec<u8>> {
append_missing_eof_marker(stripped_buf),
buf,
);
add_repair_candidate(
&mut candidates,
recover_startxref_pointer(stripped_buf),
buf,
);
}
candidates
}
/// Some PDF writers emit a `startxref` pointer that doesn't actually point
/// at the cross-reference table — a single corrupted byte in the offset is
/// enough. lopdf trusts that pointer outright and fails to load rather than
/// searching for the real table, unlike pypdf/pdfium which both recover by
/// locating it directly. This finds the real (classic, non-stream) `xref`
/// table by scanning for the keyword — validating that a plausible
/// subsection header follows, not just any standalone "xref" token, since
/// this crate processes untrusted input and a coincidental match inside
/// unrelated stream/string content must not get "repaired" against a bogus
/// offset (lopdf would then load successfully against garbage instead of
/// returning a clean error) — and appends a corrected trailing
/// `startxref`/`%%EOF` block. lopdf's own `get_xref_start` always uses the
/// *last* `%%EOF` in the final 512 bytes of the buffer, so ours
/// transparently supersedes the broken one without needing to touch
/// anything already in the file.
///
/// Doesn't cover cross-reference *streams* (`N 0 obj << /Type /XRef ...`,
/// used by some PDF 1.5+ writers instead of a classic table) — recovering
/// those needs the containing object's number, not just a byte offset.
fn recover_startxref_pointer(buf: &[u8]) -> Option<Vec<u8>> {
let xref_pos = find_last_valid_xref_table_start(buf)?;
let mut repaired = Vec::with_capacity(buf.len() + 32);
repaired.extend_from_slice(buf);
if !repaired.ends_with(b"\n") {
repaired.push(b'\n');
}
repaired.extend_from_slice(format!("startxref\n{xref_pos}\n%%EOF\n").as_bytes());
Some(repaired)
}
/// Finds the last standalone `xref` token in `buf` that is immediately
/// followed by a plausible classic cross-reference subsection header
/// (`<start-id> <count>`, e.g. "0 6") — the shape every real classic xref
/// table starts with. A single reverse byte scan: O(n) even on a
/// pathological buffer with many non-matching or non-standalone "xref"
/// occurrences, unlike repeatedly re-searching a shrinking prefix.
fn find_last_valid_xref_table_start(buf: &[u8]) -> Option<usize> {
const KEYWORD: &[u8] = b"xref";
if buf.len() < KEYWORD.len() {
return None;
}
let mut pos = buf.len() - KEYWORD.len();
loop {
if &buf[pos..pos + KEYWORD.len()] == KEYWORD {
let before_ok = pos == 0 || buf[pos - 1].is_ascii_whitespace();
let after_ok = buf
.get(pos + KEYWORD.len())
.is_none_or(|c| c.is_ascii_whitespace());
if before_ok && after_ok && looks_like_xref_subsection_header(buf, pos + KEYWORD.len())
{
return Some(pos);
}
}
if pos == 0 {
return None;
}
pos -= 1;
}
}
/// Checks that `buf[pos..]` starts (after whitespace) with two
/// whitespace-separated runs of ASCII digits — `<start-id> <count>`, the
/// first subsection header of a classic PDF cross-reference table.
fn looks_like_xref_subsection_header(buf: &[u8], pos: usize) -> bool {
fn skip_ws(buf: &[u8], mut pos: usize) -> usize {
while buf.get(pos).is_some_and(u8::is_ascii_whitespace) {
pos += 1;
}
pos
}
fn skip_digits(buf: &[u8], mut pos: usize) -> usize {
while buf.get(pos).is_some_and(u8::is_ascii_digit) {
pos += 1;
}
pos
}
let pos = skip_ws(buf, pos);
let after_first_digits = skip_digits(buf, pos);
if after_first_digits == pos {
return false; // no start-id
}
let sep = skip_ws(buf, after_first_digits);
if sep == after_first_digits {
return false; // start-id and count must be whitespace-separated
}
let after_count = skip_digits(buf, sep);
if after_count == sep {
return false; // no count
}
// The count run must end at whitespace/buffer-end, not run into trailing
// garbage (e.g. a coincidental "xref\n0 6garbage" in stream content).
buf.get(after_count).is_none_or(u8::is_ascii_whitespace)
}
fn add_repair_candidate(
candidates: &mut Vec<Vec<u8>>,
candidate: Option<Vec<u8>>,
@@ -3956,14 +3816,7 @@ fn process_document(
let text_quality = analyze_text_quality(&items);
merge_ocr_reasons(&mut ocr_reasons_by_page, text_quality.reasons_by_page);
let chart_regions = markdown::chart_regions_by_page(&items, &rects, &lines);
let layout = compute_layout_complexity_with_chart_regions(
&items,
&layout_items,
&rects,
&lines,
&chart_regions,
);
let layout = compute_layout_complexity(&items, &layout_items, &rects, &lines);
let md = if options.mode == ProcessMode::Analyze {
None
@@ -3980,7 +3833,6 @@ fn process_document(
page_count,
prefiltered_page_number_pages: Some(&removed_pages),
prefiltered_page_number_mask: Some(removal_mask.as_slice()),
precomputed_chart_regions: Some(&chart_regions),
},
))
};
@@ -5781,29 +5633,11 @@ fn select_items_with_document_folio_context(
}
/// Analyse extracted items and rects for layout complexity.
#[cfg(test)]
fn compute_layout_complexity(
items: &[types::TextItem],
column_items: &[types::TextItem],
rects: &[types::PdfRect],
lines: &[types::PdfLine],
) -> LayoutComplexity {
let page_chart_regions = markdown::chart_regions_by_page(items, rects, lines);
compute_layout_complexity_with_chart_regions(
items,
column_items,
rects,
lines,
&page_chart_regions,
)
}
fn compute_layout_complexity_with_chart_regions(
items: &[types::TextItem],
column_items: &[types::TextItem],
rects: &[types::PdfRect],
lines: &[types::PdfLine],
page_chart_regions: &markdown::PageChartRegions,
) -> LayoutComplexity {
use markdown::analysis::calculate_font_stats_from_items;
@@ -5825,10 +5659,6 @@ fn compute_layout_complexity_with_chart_regions(
let owned_items: Vec<types::TextItem> = page_items.iter().map(|i| (*i).clone()).collect();
let page_content_width = tables::content_width(&owned_items);
let bands = markdown::split_side_by_side(&owned_items);
let chart_regions = page_chart_regions
.get(&page)
.map(Vec::as_slice)
.unwrap_or_default();
let band_ranges: Vec<(f32, f32)> = if bands.is_empty() {
// Single region — use sentinel range that includes everything
@@ -5843,8 +5673,7 @@ fn compute_layout_complexity_with_chart_regions(
let band_items: Vec<types::TextItem> = owned_items
.iter()
.filter(|item| {
(x_lo == f32::MIN || (item.x >= x_lo - margin && item.x < x_hi + margin))
&& !markdown::item_is_in_chart_region(item, chart_regions)
x_lo == f32::MIN || (item.x >= x_lo - margin && item.x < x_hi + margin)
})
.cloned()
.collect();
@@ -5896,20 +5725,8 @@ fn compute_layout_complexity_with_chart_regions(
}
let mut pages_with_columns: Vec<u32> = Vec::new();
for &page in &seen_pages {
let chart_regions = page_chart_regions
.get(&page)
.map(Vec::as_slice)
.unwrap_or_default();
let page_column_items: Vec<types::TextItem> = column_items
.iter()
.filter(|item| {
item.page == page && !markdown::item_is_in_chart_region(item, chart_regions)
})
.cloned()
.collect();
let cols =
extractor::detect_columns(&page_column_items, page, pages_with_tables.contains(&page));
for page in seen_pages {
let cols = extractor::detect_columns(column_items, page, pages_with_tables.contains(&page));
if cols.len() >= 2 {
pages_with_columns.push(page);
}
@@ -6138,66 +5955,6 @@ mod tests {
assert!(filtered.pages_with_columns.is_empty());
}
#[test]
fn dense_chart_panel_is_not_reported_as_a_table() {
let mut items: Vec<TextItem> = (0..8)
.flat_map(|row| {
(0..6).map(move |column| {
test_item(
&format!("{}", row * 10 + column),
105.0 + column as f32 * 35.0,
525.0 - row as f32 * 15.0,
24.0,
10.0,
)
})
})
.collect();
for row in 0..6 {
items.push(test_item(
"Left column prose continues here",
80.0,
320.0 - row as f32 * 15.0,
160.0,
10.0,
));
items.push(test_item(
"Right column prose continues here",
300.0,
320.0 - row as f32 * 15.0,
160.0,
10.0,
));
}
let mut lines: Vec<PdfLine> = (0..30)
.map(|column| PdfLine {
x1: 100.0 + column as f32 * 8.0,
y1: 400.0,
x2: 100.0 + column as f32 * 8.0,
y2: 550.0,
page: 1,
})
.collect();
lines.extend((0..6).map(|row| PdfLine {
x1: 100.0,
y1: 400.0 + row as f32 * 30.0,
x2: 332.0,
y2: 400.0 + row as f32 * 30.0,
page: 1,
}));
let rects = vec![PdfRect {
x: 80.0,
y: 350.0,
width: 280.0,
height: 240.0,
page: 1,
}];
let complexity = compute_layout_complexity(&items, &items, &rects, &lines);
assert!(complexity.pages_with_tables.is_empty());
}
#[test]
fn page_selection_keeps_document_wide_folio_layout_decisions() {
let mut items = Vec::new();
@@ -7201,65 +6958,4 @@ mod tests {
// Pre-filled cell was not touched.
assert_eq!(cells[1].text, "Pre-filled");
}
// -- recover_startxref_pointer / find_last_valid_xref_table_start ------
//
// Direct unit tests on the byte-level scan, addressing review feedback
// on #230: a coincidental standalone "xref" token that isn't actually
// followed by a subsection header (start-id + count) must not be
// treated as a real table — accepting it would let lopdf "succeed"
// against a bogus offset and silently return garbled/empty content
// instead of a clean error.
#[test]
fn find_xref_rejects_standalone_token_without_subsection_header() {
// "xref" appears as a real standalone word, but nothing that looks
// like "<start-id> <count>" follows it.
let buf = b"Please refer to the xref appendix for details.";
assert_eq!(find_last_valid_xref_table_start(buf), None);
}
#[test]
fn find_xref_accepts_real_classic_table_header() {
let buf = b"garbage\nxref\n0 6\n0000000000 65535 f \n%%EOF";
let pos = find_last_valid_xref_table_start(buf).expect("should find the real table");
assert_eq!(&buf[pos..pos + 4], b"xref");
assert_eq!(&buf[pos..], b"xref\n0 6\n0000000000 65535 f \n%%EOF");
}
#[test]
fn find_xref_skips_coincidental_match_and_finds_real_table_before_it() {
// A coincidental "xref" (no subsection header) appears *after* the
// real table in the buffer — the scan must not stop at the first
// (rightmost) standalone token it finds; it must keep looking
// backward until one actually validates.
let buf = b"xref\n0 3\n0000000000 65535 f \ntrailer\nsee the xref\n";
let pos = find_last_valid_xref_table_start(buf).expect("should find the real table");
assert_eq!(pos, 0);
}
#[test]
fn find_xref_rejects_substring_of_startxref() {
// "xref" is a substring of "startxref" but isn't a standalone
// token there (not preceded by whitespace) — must not match, even
// though a number immediately follows it.
let buf = b"startxref\n1234\n%%EOF";
assert_eq!(find_last_valid_xref_table_start(buf), None);
}
#[test]
fn find_xref_rejects_count_run_with_trailing_garbage() {
// "xref\n0 6garbage" has the right shape (digits, whitespace,
// digits) but the count run doesn't end at whitespace/EOF — it
// runs straight into non-digit garbage, so this must not be
// accepted as a real subsection header.
let buf = b"xref\n0 6garbage\n%%EOF";
assert_eq!(find_last_valid_xref_table_start(buf), None);
}
#[test]
fn recover_startxref_pointer_returns_none_without_a_valid_table() {
let buf = b"Please refer to the xref appendix for details.";
assert!(recover_startxref_pointer(buf).is_none());
}
}
+12 -85
View File
@@ -69,7 +69,7 @@ fn is_chart_adjacent_label(item: &TextItem, region: (f32, f32, f32, f32)) -> boo
|| (mostly_inside_chart_width && close_to_chart_edge && category_sized))
}
pub(crate) fn item_is_in_chart_region(item: &TextItem, regions: &[(f32, f32, f32, f32)]) -> bool {
fn item_is_in_chart_region(item: &TextItem, regions: &[(f32, f32, f32, f32)]) -> bool {
regions.iter().any(|&(x0, y0, x1, y1)| {
let cx = item.x + item.width / 2.0;
let within_padded_x = cx >= x0 - CHART_REGION_PAD && cx <= x1 + CHART_REGION_PAD;
@@ -92,72 +92,6 @@ fn items_outside_chart_regions(
.collect()
}
pub(crate) fn merge_chart_regions(
regions: impl IntoIterator<Item = (f32, f32, f32, f32)>,
) -> Vec<(f32, f32, f32, f32)> {
const MERGE_TOLERANCE: f32 = 3.0;
let mut merged: Vec<(f32, f32, f32, f32)> = Vec::new();
for (x0, y0, x1, y1) in regions {
let mut current = (x0.min(x1), y0.min(y1), x0.max(x1), y0.max(y1));
let mut index = 0;
while index < merged.len() {
let candidate = merged[index];
let overlaps = current.2 + MERGE_TOLERANCE >= candidate.0
&& candidate.2 + MERGE_TOLERANCE >= current.0
&& current.3 + MERGE_TOLERANCE >= candidate.1
&& candidate.3 + MERGE_TOLERANCE >= current.1;
if overlaps {
current = (
current.0.min(candidate.0),
current.1.min(candidate.1),
current.2.max(candidate.2),
current.3.max(candidate.3),
);
merged.swap_remove(index);
} else {
index += 1;
}
}
merged.push(current);
}
merged
}
pub(crate) type PageChartRegions = HashMap<u32, Vec<(f32, f32, f32, f32)>>;
/// Compute the chart masks used by both layout analysis and Markdown output.
///
/// Keeping the rect-backed and dense-line heuristics behind one entry point
/// ensures metadata and extraction cannot drift when either detector changes.
pub(crate) fn chart_regions_by_page(
items: &[TextItem],
rects: &[PdfRect],
lines: &[PdfLine],
) -> PageChartRegions {
let mut page_items: HashMap<u32, Vec<TextItem>> = HashMap::new();
for item in items.iter().filter(|item| {
matches!(
&item.item_type,
crate::types::ItemType::Text | crate::types::ItemType::FormField
)
}) {
page_items.entry(item.page).or_default().push(item.clone());
}
page_items
.into_iter()
.filter_map(|(page, items)| {
let rect_regions = crate::tables::detect_chart_regions(&items, rects, page);
let line_regions = crate::tables::detect_dense_line_chart_regions(lines, rects, page)
.into_iter()
.filter(|&region| chart_region_separates_prose_columns(&items, region));
let regions = merge_chart_regions(rect_regions.into_iter().chain(line_regions));
(!regions.is_empty()).then_some((page, regions))
})
.collect()
}
/// Detect side-by-side table layout by finding a significant X-position gap.
///
/// Returns X-band boundaries `[(x_min, split_x), (split_x, x_max)]` when a
@@ -395,15 +329,6 @@ fn chart_spans_prose_split(region: (f32, f32, f32, f32), split_x: f32) -> bool {
split_x - left >= MIN_CHART_WIDTH_PER_SIDE && right - split_x >= MIN_CHART_WIDTH_PER_SIDE
}
pub(crate) fn chart_region_separates_prose_columns(
items: &[TextItem],
region: (f32, f32, f32, f32),
) -> bool {
let outside = items_outside_chart_regions(items, &[region]);
chart_page_prose_column_split(&outside)
.is_some_and(|split_x| chart_spans_prose_split(region, split_x))
}
/// True when adjacent physical rows form an unterminated, lowercase prose
/// continuation in the same projected column.
fn is_cross_row_prose_continuation(previous: &str, current: &str) -> bool {
@@ -1079,7 +1004,6 @@ pub fn to_markdown_from_items_with_rects_and_page_count(
page_count: document_page_count,
prefiltered_page_number_pages: None,
prefiltered_page_number_mask: None,
precomputed_chart_regions: None,
},
)
}
@@ -1097,9 +1021,6 @@ pub(crate) struct MarkdownDocumentContext<'a> {
/// Table detection consumes the original items; the mask is applied only
/// after table claims have been established.
pub(crate) prefiltered_page_number_mask: Option<&'a [bool]>,
/// Optional chart masks shared with layout analysis so the geometry is
/// detected once and interpreted identically by both pipelines.
pub(crate) precomputed_chart_regions: Option<&'a PageChartRegions>,
}
/// Convert positioned text items to markdown, using rectangles and line segments for table detection.
@@ -1126,7 +1047,6 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
page_count: document_page_count,
prefiltered_page_number_pages,
prefiltered_page_number_mask,
precomputed_chart_regions,
} = context;
if items.is_empty() {
@@ -1199,9 +1119,17 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
// Chart regions per page: their text must not steer column detection
// during line grouping (it fills the gutter and fuses two-column lines).
let page_chart_map = precomputed_chart_regions
.cloned()
.unwrap_or_else(|| chart_regions_by_page(&text_items, rects, pdf_lines));
let mut page_chart_map: HashMap<u32, Vec<(f32, f32, f32, f32)>> = HashMap::new();
for &page in page_groups.keys() {
let page_items_ref: Vec<TextItem> = page_groups[&page]
.iter()
.map(|(_, item)| (*item).clone())
.collect();
let regions = crate::tables::detect_chart_regions(&page_items_ref, rects, page);
if !regions.is_empty() {
page_chart_map.insert(page, regions);
}
}
let mut pages: Vec<u32> = page_groups.keys().copied().collect();
pages.sort();
@@ -2123,7 +2051,6 @@ mod tests {
page_count: 1,
prefiltered_page_number_pages: Some(&removed_pages),
prefiltered_page_number_mask: Some(&removal_mask),
precomputed_chart_regions: None,
},
);
-9
View File
@@ -464,7 +464,6 @@ mod tests {
assert!(!is_page_number_line("Hello World"));
assert!(!is_page_number_line("Chapter 1"));
assert!(!is_page_number_line("Total: 500"));
assert!(!is_page_number_line("PAGE0-PARA2-END-MARKER-0"));
}
#[test]
@@ -508,14 +507,6 @@ mod tests {
assert!(result.contains("End"));
}
#[test]
fn test_remove_page_numbers_preserves_page_prefixed_content() {
let input = "PAGE0-PARA2-START substantive report text PAGE0-PARA2-END-MARKER-0";
let result = remove_page_numbers(input);
assert_eq!(result, input);
}
#[test]
fn test_remove_page_numbers_multiple_patterns() {
let input = "\n1\n\nContent\n\n2\n\n---\nMore\n\n3\n";
+2 -450
View File
@@ -4,10 +4,10 @@
//! gridlines. Many IRS forms and government PDFs use these instead of
//! `re` (rectangle) operators.
use std::collections::{HashMap, HashSet};
use std::collections::HashSet;
use crate::tables::Table;
use crate::types::{PdfLine, PdfRect, TextItem};
use crate::types::{PdfLine, TextItem};
use super::detect_rects::{assign_items_to_grid, snap_edges};
@@ -15,33 +15,11 @@ const RULE_Y_TOLERANCE: f32 = 2.0;
const RULE_JOIN_GAP: f32 = 6.0;
const RULE_SPAN_TOLERANCE: f32 = 8.0;
const TEXT_ROW_TOLERANCE: f32 = 2.5;
const DENSE_CHART_MIN_VERTICAL_EDGES: usize = 27;
const DENSE_CHART_LABEL_PAD: f32 = 20.0;
const DENSE_CHART_MAX_SHARED_PANEL_GRIDS: usize = 4;
type HorizontalRule = (f32, f32, f32); // (y, x_min, x_max)
type VerticalRule = (f32, f32, f32); // (x, y_min, y_max)
type AnchoredRow<'a> = (f32, Vec<(usize, &'a TextItem)>);
fn dense_chart_grids_are_co_located(
left: (f32, f32, f32, f32),
right: (f32, f32, f32, f32),
) -> bool {
let left_width = left.2 - left.0;
let right_width = right.2 - right.0;
let left_height = left.3 - left.1;
let right_height = right.3 - right.1;
let horizontal_overlap = (left.2.min(right.2) - left.0.max(right.0)).max(0.0);
let vertical_overlap = (left.3.min(right.3) - left.1.max(right.1)).max(0.0);
let horizontal_gap = (left.0.max(right.0) - left.2.min(right.2)).max(0.0);
let vertical_gap = (left.1.max(right.1) - left.3.min(right.3)).max(0.0);
(vertical_overlap >= left_height.min(right_height) * 0.5
&& horizontal_gap <= left_width.min(right_width) * 0.5)
|| (horizontal_overlap >= left_width.min(right_width) * 0.5
&& vertical_gap <= left_height.min(right_height) * 0.5)
}
#[derive(Debug)]
struct TextAnchorTable {
table: Table,
@@ -1209,279 +1187,6 @@ pub fn detect_tables_from_lines(items: &[TextItem], lines: &[PdfLine], page: u32
detect_tables_from_lines_inner(items, lines, page, true, true)
}
/// Bounding boxes of chart panels backed by a very dense vector grid.
///
/// Tables support at most 25 columns, so a panel with at least 27 distinct,
/// long vertical coordinates plus repeated horizontal rules is treated as
/// chart geometry. When the grid is enclosed by a painted panel rectangle,
/// the region expands to that rectangle so axis labels, legends, and source
/// notes remain part of the figure instead of forming a heuristic table.
pub(crate) fn detect_dense_line_chart_regions(
lines: &[PdfLine],
rects: &[PdfRect],
page: u32,
) -> Vec<(f32, f32, f32, f32)> {
const ANGLE_TOLERANCE: f32 = 0.035;
const MIN_GRID_LINE_LENGTH: f32 = 40.0;
const EXTENT_TOLERANCE: f32 = 6.0;
let mut verticals = Vec::new();
let mut horizontals = Vec::new();
for line in lines.iter().filter(|line| line.page == page) {
let dx = (line.x2 - line.x1).abs();
let dy = (line.y2 - line.y1).abs();
let length = dx.hypot(dy);
if length < MIN_GRID_LINE_LENGTH {
continue;
}
if dy > 0.01 && dx / dy <= ANGLE_TOLERANCE {
verticals.push((
(line.x1 + line.x2) / 2.0,
line.y1.min(line.y2),
line.y1.max(line.y2),
));
} else if dx > 0.01 && dy / dx <= ANGLE_TOLERANCE {
horizontals.push((
(line.y1 + line.y2) / 2.0,
line.x1.min(line.x2),
line.x1.max(line.x2),
));
}
}
if verticals.len() < DENSE_CHART_MIN_VERTICAL_EDGES || horizontals.len() < 3 {
return Vec::new();
}
// Group similar vertical extents once. Neighboring buckets are consulted
// below so coordinates that straddle a bucket boundary still form one
// family, while each line participates in only a constant number of
// candidates instead of being re-scanned for every vertical anchor.
let extent_key = |value: f32| (value / EXTENT_TOLERANCE).round() as i32;
let mut extent_buckets: HashMap<(i32, i32), Vec<VerticalRule>> = HashMap::new();
for vertical in verticals {
extent_buckets
.entry((extent_key(vertical.1), extent_key(vertical.2)))
.or_default()
.push(vertical);
}
let mut grid_regions = Vec::new();
let extent_keys: Vec<(i32, i32)> = extent_buckets.keys().copied().collect();
for key in extent_keys {
let anchor_family = &extent_buckets[&key];
let anchor_bottom = anchor_family.iter().map(|vertical| vertical.1).sum::<f32>()
/ anchor_family.len() as f32;
let anchor_top = anchor_family.iter().map(|vertical| vertical.2).sum::<f32>()
/ anchor_family.len() as f32;
let mut family = Vec::new();
for bottom_offset in -1..=1 {
for top_offset in -1..=1 {
if let Some(bucket) =
extent_buckets.get(&(key.0 + bottom_offset, key.1 + top_offset))
{
family.extend(bucket.iter().copied().filter(|vertical| {
(vertical.1 - anchor_bottom).abs() <= EXTENT_TOLERANCE
&& (vertical.2 - anchor_top).abs() <= EXTENT_TOLERANCE
}));
}
}
}
let xs = snap_edges(&family.iter().map(|&(x, _, _)| x).collect::<Vec<_>>(), 3.0);
if xs.len() < DENSE_CHART_MIN_VERTICAL_EDGES {
continue;
}
let grid_bottom =
family.iter().map(|vertical| vertical.1).sum::<f32>() / family.len() as f32;
let grid_top = family.iter().map(|vertical| vertical.2).sum::<f32>() / family.len() as f32;
if grid_top - grid_bottom < 60.0 {
continue;
}
// A horizontal rule must support the same contiguous dense run of
// vertical coordinates. Splitting at sparse X gaps prevents a shared
// rule from joining a chart to a neighboring ruled table. Keying by
// the covered X-index range also lets multiple chart panels sharing
// the same Y extents produce independent regions.
let mut supported_spans: HashMap<(usize, usize), Vec<f32>> = HashMap::new();
for &(y, line_left, line_right) in &horizontals {
if y < grid_bottom - EXTENT_TOLERANCE || y > grid_top + EXTENT_TOLERANCE {
continue;
}
let start = xs.partition_point(|&x| x < line_left - EXTENT_TOLERANCE);
let end = xs.partition_point(|&x| x <= line_right + EXTENT_TOLERANCE);
if end - start < DENSE_CHART_MIN_VERTICAL_EDGES {
continue;
}
let mut gaps: Vec<f32> = xs[start..end]
.windows(2)
.map(|pair| pair[1] - pair[0])
.collect();
gaps.sort_by(f32::total_cmp);
let dense_gap = gaps[gaps.len() / 4];
let run_break = (dense_gap * 3.0).max(12.0);
let locally_dense_gap_limit = (dense_gap * 1.5).max(6.0);
let locally_dense_gaps = gaps
.iter()
.filter(|&&gap| gap <= locally_dense_gap_limit)
.count();
let mut run_start = start;
let mut retained_dense_run = false;
for index in start..end - 1 {
if xs[index + 1] - xs[index] <= run_break {
continue;
}
let run_end = index + 1;
if run_end - run_start >= DENSE_CHART_MIN_VERTICAL_EDGES
&& xs[run_end - 1] - xs[run_start] >= 120.0
{
supported_spans
.entry((run_start, run_end))
.or_default()
.push(y);
retained_dense_run = true;
}
run_start = run_end;
}
if end - run_start >= DENSE_CHART_MIN_VERTICAL_EDGES
&& xs[end - 1] - xs[run_start] >= 120.0
{
supported_spans.entry((run_start, end)).or_default().push(y);
retained_dense_run = true;
}
// One or two wider category gaps may split an otherwise dense
// chart into sub-threshold runs. Keep the full family only when
// its total width remains close to the expected dense spacing;
// a neighboring sparse table makes this ratio much larger.
let span_width = xs[end - 1] - xs[start];
let expected_dense_width = dense_gap * (end - start - 1) as f32;
if !retained_dense_run
&& span_width >= 120.0
&& span_width <= expected_dense_width * 1.35
&& gaps.len().saturating_sub(locally_dense_gaps) <= 2
{
supported_spans.entry((start, end)).or_default().push(y);
}
}
for ((start, end), ys) in supported_spans {
if snap_edges(&ys, 3.0).len() >= 3 {
grid_regions.push((xs[start], grid_bottom, xs[end - 1], grid_top));
}
}
}
// Prefer the smallest qualifying region when a broad rule happens to
// cover a denser nested panel, and retain every non-overlapping panel.
grid_regions.sort_by(|left, right| {
let left_area = (left.2 - left.0) * (left.3 - left.1);
let right_area = (right.2 - right.0) * (right.3 - right.1);
left_area.total_cmp(&right_area)
});
let mut selected_regions: Vec<(f32, f32, f32, f32)> = Vec::new();
for region in grid_regions {
let area = (region.2 - region.0) * (region.3 - region.1);
let duplicates_existing = selected_regions.iter().any(|existing| {
let overlap_width = (region.2.min(existing.2) - region.0.max(existing.0)).max(0.0);
let overlap_height = (region.3.min(existing.3) - region.1.max(existing.1)).max(0.0);
let overlap_area = overlap_width * overlap_height;
let existing_area = (existing.2 - existing.0) * (existing.3 - existing.1);
overlap_area >= area.min(existing_area) * 0.8
});
if !duplicates_existing {
selected_regions.push(region);
}
}
let all_grid_regions = selected_regions.clone();
let mut regions: Vec<_> = selected_regions
.into_iter()
.map(|grid_region| {
let (grid_left, grid_bottom, grid_right, grid_top) = grid_region;
let enclosing_panel = rects
.iter()
.filter(|rect| rect.page == page)
.filter_map(|rect| {
let (left, width) = if rect.width < 0.0 {
(rect.x + rect.width, -rect.width)
} else {
(rect.x, rect.width)
};
let (bottom, height) = if rect.height < 0.0 {
(rect.y + rect.height, -rect.height)
} else {
(rect.y, rect.height)
};
let right = left + width;
let top = bottom + height;
let enclosed_grids: Vec<_> = all_grid_regions
.iter()
.filter(|&&(other_left, other_bottom, other_right, other_top)| {
left <= other_left + EXTENT_TOLERANCE
&& right >= other_right - EXTENT_TOLERANCE
&& bottom <= other_bottom + EXTENT_TOLERANCE
&& top >= other_top - EXTENT_TOLERANCE
})
.copied()
.collect();
if enclosed_grids.len() > DENSE_CHART_MAX_SHARED_PANEL_GRIDS
|| enclosed_grids.iter().any(|&other| {
other != grid_region
&& !dense_chart_grids_are_co_located(grid_region, other)
})
{
return None;
}
let enclosed_grid_bounds =
enclosed_grids.into_iter().reduce(|bounds, other| {
(
bounds.0.min(other.0),
bounds.1.min(other.1),
bounds.2.max(other.2),
bounds.3.max(other.3),
)
})?;
let enclosed_width = enclosed_grid_bounds.2 - enclosed_grid_bounds.0;
let enclosed_height = enclosed_grid_bounds.3 - enclosed_grid_bounds.1;
(left <= grid_left + EXTENT_TOLERANCE
&& right >= grid_right - EXTENT_TOLERANCE
&& bottom <= grid_bottom + EXTENT_TOLERANCE
&& top >= grid_top - EXTENT_TOLERANCE
&& width <= enclosed_width * 2.0
&& height <= enclosed_height * 4.0
&& !(left < 5.0 && bottom < 5.0))
.then_some(((left, bottom, right, top), width * height))
})
.min_by(|left, right| left.1.total_cmp(&right.1))
.map(|(region, _)| region);
enclosing_panel.unwrap_or((
grid_left - DENSE_CHART_LABEL_PAD,
grid_bottom - DENSE_CHART_LABEL_PAD,
grid_right + DENSE_CHART_LABEL_PAD,
grid_top + DENSE_CHART_LABEL_PAD,
))
})
.collect();
regions.sort_by(|left, right| {
left.0
.total_cmp(&right.0)
.then_with(|| left.1.total_cmp(&right.1))
});
regions.dedup_by(|left, right| {
(left.0 - right.0).abs() <= EXTENT_TOLERANCE
&& (left.1 - right.1).abs() <= EXTENT_TOLERANCE
&& (left.2 - right.2).abs() <= EXTENT_TOLERANCE
&& (left.3 - right.3).abs() <= EXTENT_TOLERANCE
});
regions
}
/// Detect only tables whose cell grid is backed by explicit vector geometry.
///
/// Region-level TSR callers need physical cell boundaries for crop bboxes, so
@@ -1916,159 +1621,6 @@ mod tests {
}
}
#[test]
fn dense_vector_grid_expands_to_enclosing_chart_panel() {
let mut lines: Vec<PdfLine> = (0..30)
.map(|column| make_vline(100.0 + column as f32 * 8.0, 400.0, 550.0, 1))
.collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 332.0, 1)));
let rects = vec![PdfRect {
x: 80.0,
y: 350.0,
width: 280.0,
height: 240.0,
page: 1,
}];
assert_eq!(
detect_dense_line_chart_regions(&lines, &rects, 1),
vec![(80.0, 350.0, 360.0, 590.0)]
);
}
#[test]
fn frameless_dense_vector_grid_includes_label_padding() {
let mut lines: Vec<PdfLine> = (0..30)
.map(|column| make_vline(100.0 + column as f32 * 8.0, 400.0, 550.0, 1))
.collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 332.0, 1)));
assert_eq!(
detect_dense_line_chart_regions(&lines, &[], 1),
vec![(80.0, 380.0, 352.0, 570.0)]
);
}
#[test]
fn multiple_dense_vector_panels_are_retained() {
let mut lines = Vec::new();
for panel_left in [60.0, 380.0] {
lines.extend(
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
);
lines.extend((0..6).map(|row| {
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
}));
}
assert_eq!(
detect_dense_line_chart_regions(&lines, &[], 1),
vec![(40.0, 380.0, 312.0, 570.0), (360.0, 380.0, 632.0, 570.0),]
);
}
#[test]
fn multiple_dense_vector_panels_use_shared_enclosing_panel() {
let mut lines = Vec::new();
for panel_left in [60.0, 380.0] {
lines.extend(
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
);
lines.extend((0..6).map(|row| {
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
}));
}
let rects = vec![PdfRect {
x: 40.0,
y: 350.0,
width: 592.0,
height: 240.0,
page: 1,
}];
assert_eq!(
detect_dense_line_chart_regions(&lines, &rects, 1),
vec![(40.0, 350.0, 632.0, 590.0)]
);
}
#[test]
fn shared_rules_do_not_join_dense_chart_to_adjacent_table() {
let mut lines: Vec<PdfLine> = (0..30)
.map(|column| make_vline(60.0 + column as f32 * 8.0, 400.0, 550.0, 1))
.collect();
lines.extend(
[330.0, 390.0, 450.0, 510.0, 570.0, 630.0]
.into_iter()
.map(|x| make_vline(x, 400.0, 550.0, 1)),
);
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 630.0, 1)));
assert_eq!(
detect_dense_line_chart_regions(&lines, &[], 1),
vec![(40.0, 380.0, 312.0, 570.0)]
);
}
#[test]
fn uneven_dense_spacing_keeps_the_complete_chart_region() {
let mut xs: Vec<f32> = (0..15).map(|column| 60.0 + column as f32 * 8.0).collect();
xs.extend((0..15).map(|column| 212.0 + column as f32 * 8.0));
let mut lines: Vec<PdfLine> = xs.iter().map(|&x| make_vline(x, 400.0, 550.0, 1)).collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 324.0, 1)));
assert_eq!(
detect_dense_line_chart_regions(&lines, &[], 1),
vec![(40.0, 380.0, 344.0, 570.0)]
);
}
#[test]
fn subthreshold_dense_run_does_not_absorb_adjacent_sparse_grid() {
let mut xs: Vec<f32> = (0..21).map(|column| 60.0 + column as f32 * 8.0).collect();
xs.extend((0..6).map(|column| 248.0 + column as f32 * 18.0));
let mut lines: Vec<PdfLine> = xs.iter().map(|&x| make_vline(x, 400.0, 550.0, 1)).collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 338.0, 1)));
assert!(detect_dense_line_chart_regions(&lines, &[], 1).is_empty());
}
#[test]
fn broad_frame_does_not_merge_distant_dense_grids() {
let mut lines = Vec::new();
for panel_left in [60.0, 700.0] {
lines.extend(
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
);
lines.extend((0..6).map(|row| {
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
}));
}
let rects = vec![PdfRect {
x: 40.0,
y: 350.0,
width: 912.0,
height: 240.0,
page: 1,
}];
assert_eq!(
detect_dense_line_chart_regions(&lines, &rects, 1),
vec![(40.0, 380.0, 312.0, 570.0), (680.0, 380.0, 952.0, 570.0)]
);
}
#[test]
fn supported_width_vector_table_is_not_a_dense_chart() {
let mut lines: Vec<PdfLine> = (0..26)
.map(|column| make_vline(100.0 + column as f32 * 10.0, 400.0, 550.0, 1))
.collect();
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 350.0, 1)));
assert!(detect_dense_line_chart_regions(&lines, &[], 1).is_empty());
}
#[test]
fn test_basic_grid_detection() {
// 3x2 grid with horizontal lines at y=500, 480, 460 and vertical at x=100, 200, 300
+10 -454
View File
@@ -1996,249 +1996,17 @@ fn without_dominant_page_backgrounds(rects: &[(f32, f32, f32, f32)]) -> Vec<(f32
.collect()
}
/// Repeated rows of touching cell rectangles are stronger table evidence
/// than the bar-length variation used by the chart detector.
/// Detect a table from cell-background rects that failed grid detection.
///
/// Ruled tables with wrapped labels naturally have variable row heights, and
/// numeric-heavy cells can otherwise resemble horizontal or vertical bars.
/// Require several rows to repeat a shared edge schema before overriding the
/// chart hypothesis so sparse plots and independent bars remain unaffected.
fn is_repeated_cell_grid(group_rects: &[(f32, f32, f32, f32)]) -> bool {
type RowGroup = (f32, f32, Vec<(f32, f32)>);
const ROW_EDGE_TOLERANCE: f32 = 3.0;
const MIN_GRID_ROWS: usize = 4;
const MIN_CELLS_PER_ROW: usize = 3;
if group_rects.len() < MIN_GRID_ROWS * MIN_CELLS_PER_ROW {
return false;
}
let mut row_groups: Vec<RowGroup> = Vec::new();
for &(x, y, width, height) in group_rects {
if width < 5.0 || height < 5.0 {
continue;
}
let top = y + height;
if let Some((_, _, cells)) = row_groups.iter_mut().find(|(bottom, row_top, _)| {
(y - *bottom).abs() <= ROW_EDGE_TOLERANCE
&& (top - *row_top).abs() <= ROW_EDGE_TOLERANCE
}) {
cells.push((x, x + width));
} else {
row_groups.push((y, top, vec![(x, x + width)]));
}
}
let mut row_schemas = Vec::new();
for (_, _, mut cells) in row_groups {
if cells.len() < MIN_CELLS_PER_ROW {
continue;
}
let mut widths: Vec<f32> = cells.iter().map(|&(left, right)| right - left).collect();
widths.sort_by(f32::total_cmp);
let median_width = widths[widths.len() / 2];
cells.retain(|&(left, right)| right - left <= median_width * 2.5);
cells.sort_by(|left, right| {
left.0
.total_cmp(&right.0)
.then_with(|| left.1.total_cmp(&right.1))
});
cells.dedup_by(|left, right| {
(left.0 - right.0).abs() <= ROW_EDGE_TOLERANCE
&& (left.1 - right.1).abs() <= ROW_EDGE_TOLERANCE
});
if cells.len() < MIN_CELLS_PER_ROW
|| cells
.windows(2)
.any(|pair| pair[1].0 > pair[0].1 + ROW_EDGE_TOLERANCE)
{
continue;
}
let edges: Vec<f32> = cells
.iter()
.flat_map(|&(left, right)| [left, right])
.collect();
let schema = snap_edges(&edges, ROW_EDGE_TOLERANCE);
if schema.len() > MIN_CELLS_PER_ROW {
row_schemas.push(schema);
}
}
if row_schemas.len() < MIN_GRID_ROWS {
return false;
}
let reference = row_schemas
.iter()
.max_by_key(|schema| schema.len())
.expect("grid rows are non-empty");
row_schemas
.iter()
.filter(|schema| {
let comparable_edges = reference.len().min(schema.len());
let matched_edges = schema
.iter()
.filter(|edge| {
reference
.iter()
.any(|reference_edge| (*edge - *reference_edge).abs() <= ROW_EDGE_TOLERANCE)
})
.count();
matched_edges > MIN_CELLS_PER_ROW && matched_edges * 4 >= comparable_edges * 3
})
.count()
>= MIN_GRID_ROWS
}
fn repeated_cell_grid_overrides_bar_hypothesis(group_rects: &[(f32, f32, f32, f32)]) -> bool {
is_repeated_cell_grid(group_rects)
&& without_dominant_page_backgrounds(group_rects).len() == group_rects.len()
}
/// Detect horizontal segmented stacks from aligned rows of touching rects.
///
/// Category rows must have visible gutters and data-varying internal segment
/// boundaries, unlike the stable boundaries of a ruled table.
struct SegmentedBarGeometry {
bounds: (f32, f32, f32, f32),
row_bands: Vec<(f32, f32)>,
}
fn segmented_stacked_bar_geometry(
group_rects: &[(f32, f32, f32, f32)],
) -> Option<SegmentedBarGeometry> {
type BarRow = (f32, f32, Vec<(f32, f32)>);
const EDGE_TOLERANCE: f32 = 3.0;
const MIN_ROWS: usize = 4;
const MIN_SEGMENTS: usize = 3;
let mut rows: Vec<BarRow> = Vec::new();
for &(x, y, width, height) in group_rects {
if width < 5.0 || height < 5.0 {
continue;
}
let top = y + height;
if let Some((_, _, segments)) = rows.iter_mut().find(|(bottom, row_top, _)| {
(y - *bottom).abs() <= EDGE_TOLERANCE && (top - *row_top).abs() <= EDGE_TOLERANCE
}) {
segments.push((x, x + width));
} else {
rows.push((y, top, vec![(x, x + width)]));
}
}
rows.retain_mut(|(_, _, segments)| {
segments.sort_by(|left, right| left.0.total_cmp(&right.0));
segments.len() >= MIN_SEGMENTS
&& segments
.windows(2)
.all(|pair| (pair[1].0 - pair[0].1).abs() <= EDGE_TOLERANCE)
});
if rows.len() < MIN_ROWS {
return None;
}
rows.sort_by(|left, right| left.0.total_cmp(&right.0));
// Table rows normally share borders. Horizontal stacked bars instead
// leave a visible gutter between category rows.
if rows.windows(2).any(|pair| {
let shorter_height = (pair[0].1 - pair[0].0).min(pair[1].1 - pair[1].0);
pair[1].0 - pair[0].1 < (shorter_height * 0.25).max(2.0)
}) {
return None;
}
// At least two rows must move an internal segment boundary. Stable
// boundaries across every row are stronger evidence for a ruled table.
let reference_edges: Vec<f32> = rows[0]
.2
.iter()
.take(rows[0].2.len() - 1)
.map(|segment| segment.1)
.collect();
let drifting_rows = rows
.iter()
.skip(1)
.filter(|(_, _, segments)| {
let edges: Vec<f32> = segments
.iter()
.take(segments.len() - 1)
.map(|segment| segment.1)
.collect();
edges.len() == reference_edges.len()
&& edges
.iter()
.zip(&reference_edges)
.any(|(edge, reference)| (edge - reference).abs() > EDGE_TOLERANCE)
})
.count();
if drifting_rows < 2 {
return None;
}
let left = rows
.iter()
.flat_map(|row| &row.2)
.map(|segment| segment.0)
.reduce(f32::min)?;
let right = rows
.iter()
.flat_map(|row| &row.2)
.map(|segment| segment.1)
.reduce(f32::max)?;
let bottom = rows.iter().map(|row| row.0).reduce(f32::min)?;
let top = rows.iter().map(|row| row.1).reduce(f32::max)?;
let row_bands = rows.iter().map(|row| (row.0, row.1)).collect();
Some(SegmentedBarGeometry {
bounds: (left, bottom, right, top),
row_bands,
})
}
/// Category labels beside multiple bar rows are independent chart evidence:
/// numeric table text stays inside its cells, regardless of whether the table
/// has an outer border or extra padding.
fn has_external_segmented_bar_labels(
items: &[TextItem],
page: u32,
geometry: &SegmentedBarGeometry,
) -> bool {
const LABEL_EDGE_TOLERANCE: f32 = 3.0;
const LABEL_CLAIM_PAD: f32 = 20.0;
let (content_left, _, content_right, _) = geometry.bounds;
let labeled_rows = geometry
.row_bands
.iter()
.filter(|&&(row_bottom, row_top)| {
items.iter().any(|item| {
if item.page != page || item.text.trim().is_empty() {
return false;
}
let item_left = item.x.min(item.x + item.width);
let item_right = item.x.max(item.x + item.width);
let item_center_x = (item_left + item_right) / 2.0;
let item_center_y = item.y + item.height / 2.0;
let beside_stack = (item_center_x <= content_left + LABEL_EDGE_TOLERANCE
&& item_center_x >= content_left - LABEL_CLAIM_PAD
&& item_left < content_left)
|| (item_center_x >= content_right - LABEL_EDGE_TOLERANCE
&& item_center_x <= content_right + LABEL_CLAIM_PAD
&& item_right > content_right);
beside_stack
&& item_center_y >= row_bottom - LABEL_EDGE_TOLERANCE
&& item_center_y <= row_top + LABEL_EDGE_TOLERANCE
})
})
.count();
labeled_rows >= 2 && labeled_rows * 2 >= geometry.row_bands.len()
}
/// Recognize filled vertical or horizontal bars whose geometry and labels are
/// data-driven rather than uniform table cells.
fn has_chart_bar_signature(
/// Uses rect Y-edges for row boundaries and text X-position clustering for
/// columns. Handles tables with cell backgrounds that don't form a clean
/// X-edge grid (variable column widths, decorative fills).
/// Chart-bar signature: ≥3 rects sharing an aligned bottom edge (the axis),
/// with similar widths (bars) but strongly varying heights (data-driven),
/// holding at most a single numeric data label each. Bar charts drawn as
/// filled rects otherwise read as cell rects and grid their axis labels
/// into a phantom table. The mirrored check catches horizontal bar charts.
fn is_chart_bar_cluster(
items: &[TextItem],
group_rects: &[(f32, f32, f32, f32)],
page: u32,
@@ -2345,30 +2113,6 @@ fn has_chart_bar_signature(
|| bar_family(|r| r.1, |r| r.3, |r| r.2, |r| r.0)
}
fn is_chart_bar_cluster(
items: &[TextItem],
group_rects: &[(f32, f32, f32, f32)],
page: u32,
) -> bool {
let has_bar_signature = has_chart_bar_signature(items, group_rects, page);
// A segmented horizontal chart can share most of its edges across rows.
// Row-aligned category labels outside the stack distinguish it from a
// numeric table without depending on whether either shape has a frame.
if has_bar_signature {
if let Some(geometry) = segmented_stacked_bar_geometry(group_rects) {
if has_external_segmented_bar_labels(items, page, &geometry) {
return true;
}
}
}
if repeated_cell_grid_overrides_bar_hypothesis(group_rects) {
return false;
}
has_bar_signature
}
fn detect_row_stripe_table_from_cell_rects(
items: &[TextItem],
group_rects: &[(f32, f32, f32, f32)],
@@ -3339,194 +3083,6 @@ mod tests {
assert!(detect_chart_regions(&items, &rects, 1).is_empty());
}
#[test]
fn variable_height_ruled_grid_overrides_bar_hypothesis() {
let edge_sets = [
[80.0, 140.0, 200.0, 260.0, 320.0, 380.0, 440.0, 500.0, 560.0],
[80.0, 140.0, 210.0, 260.0, 320.0, 380.0, 450.0, 500.0, 560.0],
];
let heights = [20.0, 34.0, 26.0, 42.0, 20.0, 34.0];
let edge_variants = [0, 0, 0, 0, 1, 1];
let mut rects = Vec::new();
let mut y = 650.0;
for (row, height) in heights.into_iter().enumerate() {
let edges = edge_sets[edge_variants[row]];
rects.extend(
edges
.windows(2)
.map(|edge| (edge[0], y, edge[1] - edge[0], height)),
);
y -= height;
}
assert!(is_repeated_cell_grid(&rects));
assert!(has_chart_bar_signature(&[], &rects, 1));
assert!(repeated_cell_grid_overrides_bar_hypothesis(&rects));
assert!(segmented_stacked_bar_geometry(&rects).is_none());
assert!(!is_chart_bar_cluster(&[], &rects, 1));
let mut with_page_fills =
vec![(0.0, 0.0, 600.0, 800.0); DOMINANT_PAGE_BACKGROUND_MIN_REPETITIONS];
with_page_fills.extend(rects);
assert!(!repeated_cell_grid_overrides_bar_hypothesis(
&with_page_fills
));
}
#[test]
fn touching_segments_with_spaced_rows_remain_a_chart() {
let row_edges = [
[100.0, 140.0, 180.0, 220.0, 260.0],
[100.0, 140.0, 180.0, 228.0, 260.0],
[100.0, 140.0, 180.0, 214.0, 260.0],
[100.0, 140.0, 180.0, 232.0, 260.0],
];
let mut raw_rects = vec![(90.0, 530.0, 190.0, 100.0)];
for (row, edges) in row_edges.into_iter().enumerate() {
let y = 540.0 + row as f32 * 20.0;
raw_rects.extend(
edges
.windows(2)
.map(|edge| (edge[0], y, edge[1] - edge[0], 12.0)),
);
}
let items: Vec<TextItem> = (0..4)
.map(|row| make_item("Category", 62.0, 541.0 + row as f32 * 20.0, 9.0))
.collect();
assert!(is_repeated_cell_grid(&raw_rects));
assert!(has_chart_bar_signature(&items, &raw_rects, 1));
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented stack");
assert!(has_external_segmented_bar_labels(&items, 1, &geometry));
assert!(is_chart_bar_cluster(&items, &raw_rects, 1));
let numeric_items: Vec<TextItem> = (0..4)
.map(|row| make_item("2024", 80.0, 541.0 + row as f32 * 18.0, 9.0))
.collect();
assert!(has_external_segmented_bar_labels(
&numeric_items,
1,
&geometry
));
assert!(is_chart_bar_cluster(&numeric_items, &raw_rects, 1));
let edge_adjacent_items: Vec<TextItem> = (0..4)
.map(|row| make_item("2024", 92.0, 541.0 + row as f32 * 18.0, 9.0))
.collect();
assert!(has_external_segmented_bar_labels(
&edge_adjacent_items,
1,
&geometry
));
assert!(is_chart_bar_cluster(&edge_adjacent_items, &raw_rects, 1));
let far_items: Vec<TextItem> = (0..4)
.map(|row| make_item("Category", 20.0, 541.0 + row as f32 * 18.0, 9.0))
.collect();
assert!(!has_external_segmented_bar_labels(&far_items, 1, &geometry));
assert!(!is_chart_bar_cluster(&far_items, &raw_rects, 1));
let rects: Vec<PdfRect> = raw_rects
.into_iter()
.map(|(x, y, width, height)| PdfRect {
x,
y,
width,
height,
page: 1,
})
.collect();
assert_eq!(detect_chart_regions(&items, &rects, 1).len(), 1);
let (tables, hints) = detect_tables_from_rects(&items, &rects, 1);
assert!(tables.is_empty());
assert!(hints.is_empty());
}
#[test]
fn padded_numeric_grid_frame_remains_a_table() {
let row_edges = [
[100.0, 140.0, 180.0, 220.0, 260.0],
[100.0, 140.0, 180.0, 228.0, 260.0],
[100.0, 140.0, 180.0, 214.0, 260.0],
[100.0, 140.0, 180.0, 232.0, 260.0],
];
let mut raw_rects = vec![(96.0, 536.0, 168.0, 80.0)];
let mut items = Vec::new();
for (row, edges) in row_edges.into_iter().enumerate() {
let y = 540.0 + row as f32 * 20.0;
for edge in edges.windows(2) {
raw_rects.push((edge[0], y, edge[1] - edge[0], 12.0));
items.push(make_item("42", edge[0] + 8.0, y + 1.0, 9.0));
}
}
assert!(is_repeated_cell_grid(&raw_rects));
assert!(has_chart_bar_signature(&items, &raw_rects, 1));
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented rows");
assert!(!has_external_segmented_bar_labels(&items, 1, &geometry));
assert!(!is_chart_bar_cluster(&items, &raw_rects, 1));
let flush_items: Vec<TextItem> = (0..4)
.map(|row| make_item("1", 100.0, 541.0 + row as f32 * 20.0, 9.0))
.collect();
assert!(!has_external_segmented_bar_labels(
&flush_items,
1,
&geometry
));
assert!(!is_chart_bar_cluster(&flush_items, &raw_rects, 1));
let rects: Vec<PdfRect> = raw_rects
.into_iter()
.map(|(x, y, width, height)| PdfRect {
x,
y,
width,
height,
page: 1,
})
.collect();
assert!(detect_chart_regions(&items, &rects, 1).is_empty());
assert!(!detect_tables_from_rects(&items, &rects, 1).0.is_empty());
}
#[test]
fn frameless_segmented_chart_with_category_labels_remains_a_chart() {
let row_edges = [
[100.0, 140.0, 180.0, 220.0, 260.0],
[100.0, 140.0, 180.0, 228.0, 260.0],
[100.0, 140.0, 180.0, 214.0, 260.0],
[100.0, 140.0, 180.0, 232.0, 260.0],
];
let mut raw_rects = Vec::new();
let mut items = Vec::new();
for (row, edges) in row_edges.into_iter().enumerate() {
let y = 540.0 + row as f32 * 18.0;
raw_rects.extend(
edges
.windows(2)
.map(|edge| (edge[0], y, edge[1] - edge[0], 12.0)),
);
items.push(make_item("Category", 62.0, y + 1.0, 9.0));
}
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented stack");
assert!(has_external_segmented_bar_labels(&items, 1, &geometry));
assert!(is_chart_bar_cluster(&items, &raw_rects, 1));
let rects: Vec<PdfRect> = raw_rects
.into_iter()
.map(|(x, y, width, height)| PdfRect {
x,
y,
width,
height,
page: 1,
})
.collect();
assert_eq!(detect_chart_regions(&items, &rects, 1).len(), 1);
}
// --- detect_stacked_box_table ---
/// N stacked boxes at x=100, w=300, h=22, top-to-bottom from y=600.
+1 -3
View File
@@ -16,9 +16,7 @@ pub(crate) use detect_heuristic::{
content_width, detect_tables_with_page_width, is_table_of_contents,
};
pub use detect_lines::detect_tables_from_lines;
pub(crate) use detect_lines::{
detect_dense_line_chart_regions, detect_vector_grid_tables_from_lines,
};
pub(crate) use detect_lines::detect_vector_grid_tables_from_lines;
pub(crate) use detect_rects::cluster_rects;
pub use detect_rects::{detect_chart_regions, detect_tables_from_rects, RectHintRegion};
pub use detect_struct::detect_tables_from_struct_tree;
+3 -10
View File
@@ -70,17 +70,10 @@ pub(crate) fn is_page_number_line(text: &str) -> bool {
let lowercase = text.trim().to_ascii_lowercase();
lowercase.strip_prefix("page").is_some_and(|rest| {
let mut characters = rest.trim_start().chars().peekable();
let mut has_page_number = false;
while characters
.peek()
rest.trim_start()
.chars()
.next()
.is_some_and(|character| character.is_ascii_digit())
{
has_page_number = true;
characters.next();
}
has_page_number && characters.next().is_none_or(char::is_whitespace)
})
}
+1 -140
View File
@@ -594,7 +594,7 @@ fn hex_to_unicode_string(hex: &str) -> Option<String> {
let bytes: Option<Vec<u8>> = (0..hex.len())
.step_by(2)
.map(|i| u8::from_str_radix(hex.get(i..i + 2)?, 16).ok())
.map(|i| u8::from_str_radix(&hex[i..i + 2], 16).ok())
.collect();
let bytes = bytes?;
@@ -880,25 +880,6 @@ fn try_remap_subset_cmap(
None => return (cmap, None),
};
// Both repair paths below assume CIDs are glyph indices that a subsetter can
// renumber, which is only true for CIDFontType2 (TrueType). For CIDFontType0
// (CFF), CIDs are resolved through the CFF charset, so a valid CMap stays valid
// after subsetting and renumbering it corrupts otherwise-correct text.
// CIDToGIDMap is likewise CIDFontType2-only (PDF 32000-1:2008, 9.7.4.2), so this
// also ignores a CIDToGIDMap that a malformed producer attached to a CFF font.
// /Subtype may be an indirect reference, so resolve it through the document.
// Only bail out when the descendant is *explicitly* something other than
// CIDFontType2: a missing or unresolvable /Subtype keeps the previous
// behaviour rather than silently disabling the repair.
let subtype = cid_font_dict.get(b"Subtype").ok().and_then(|o| match o {
Object::Reference(r) => doc.get_object(*r).ok().and_then(|o| o.as_name().ok()),
other => other.as_name().ok(),
});
if subtype.is_some_and(|name| name != b"CIDFontType2") {
debug!("Subset remap skipped for obj={obj_num}: descendant is not CIDFontType2");
return (cmap, None);
}
// If there's an explicit CIDToGIDMap, build a repaired CMap using it.
if let Some(cid_to_gid) = get_cid_to_gid_map(cid_font_dict, doc) {
if let Some(repaired) = build_cmap_with_cid_to_gid_map(&cmap, &cid_to_gid) {
@@ -2606,24 +2587,6 @@ endcmap
assert_eq!(cmap.lookup(0x0025), Some("B".to_string()));
}
#[test]
fn test_hex_to_unicode_non_ascii_no_panic() {
// A destination containing a multi-byte char makes the byte length even
// while a byte offset can land inside a char. Slicing must not panic;
// it should be rejected gracefully.
assert_eq!(hex_to_unicode_string("XéY"), None);
assert_eq!(hex_to_unicode_string("\u{fffd}0"), None);
}
#[test]
fn test_parse_bfchar_non_ascii_destination_no_panic() {
// Crafted /ToUnicode CMap: a non-hex, non-ASCII destination previously
// triggered a char-boundary panic in hex_to_unicode_string.
let cmap_content = "beginbfchar <0041> <XéY> endbfchar";
// Must not panic; the malformed entry is simply skipped.
let _ = ToUnicodeCMap::parse(cmap_content.as_bytes());
}
#[test]
fn test_parse_bfchar_1byte() {
// This is the pattern that caused the CJK bug: codespace is <0000><FFFF>
@@ -3196,106 +3159,4 @@ endbfrange
"Remap must fire when CMap's CIDs are outside W array coverage"
);
}
#[test]
fn test_try_remap_skipped_for_cid_font_type0() {
// Same W/CMap mismatch as the CIDFontType2 case above, but the descendant is
// CIDFontType0 (CFF). There CIDs are resolved through the CFF charset, so the
// ToUnicode CIDs stay valid after subsetting and must not be renumbered.
// Real-world case: Japanese Adobe-Japan1 PDFs (e.g. National Diet Library
// minutes) where remapping turned correct text into unrelated glyphs.
let cmap_content = r#"
1 begincodespacerange
<0000><FFFF>
endcodespacerange
1 beginbfrange
<0200> <0220> <0410>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
let mut doc = Document::new();
// CIDToGIDMap is CIDFontType2-only, but a malformed producer can still emit
// one on a CFF font. Use a real stream (not /Identity, which is treated as
// "no map") so this also fails if the guard is moved back below the
// CIDToGIDMap branch: cid 1 -> gid 0x0200, which the CMap resolves.
let mut cid_to_gid = vec![0u8; 68];
cid_to_gid[2] = 0x02;
cid_to_gid[3] = 0x00;
let cid_to_gid_id =
doc.add_object(lopdf::Stream::new(lopdf::Dictionary::new(), cid_to_gid));
let mut cid_font = lopdf::Dictionary::new();
cid_font.set("Subtype", lopdf::Object::Name(b"CIDFontType0".to_vec()));
cid_font.set("CIDToGIDMap", lopdf::Object::Reference(cid_to_gid_id));
cid_font.set(
"W",
lopdf::Object::Array(vec![
lopdf::Object::Integer(1),
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
]),
);
let cid_font_id = doc.add_object(cid_font);
let mut font_dict = lopdf::Dictionary::new();
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
font_dict.set(
"DescendantFonts",
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
);
let (primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 789);
assert!(
remapped.is_none(),
"Remap must be skipped for CIDFontType0 (CFF) descendants, including a \
CIDToGIDMap a malformed producer attached to one"
);
// The original CMap must still resolve its own CIDs.
assert_eq!(primary.lookup(0x0200), Some("\u{0410}".to_string()));
}
#[test]
fn test_try_remap_resolves_indirect_subtype() {
// /Subtype may be stored as an indirect reference. A genuine CIDFontType2
// font must still get the repair, so the guard has to dereference it rather
// than treat the unresolved value as "not CIDFontType2".
let cmap_content = r#"
1 begincodespacerange
<0000><FFFF>
endcodespacerange
1 beginbfrange
<0200> <0220> <0410>
endbfrange
"#;
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
let mut doc = Document::new();
let subtype_id = doc.add_object(lopdf::Object::Name(b"CIDFontType2".to_vec()));
let mut cid_font = lopdf::Dictionary::new();
cid_font.set("Subtype", lopdf::Object::Reference(subtype_id));
cid_font.set("CIDToGIDMap", lopdf::Object::Name(b"Identity".to_vec()));
cid_font.set(
"W",
lopdf::Object::Array(vec![
lopdf::Object::Integer(0),
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
]),
);
let cid_font_id = doc.add_object(cid_font);
let mut font_dict = lopdf::Dictionary::new();
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
font_dict.set(
"DescendantFonts",
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
);
let (_primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 790);
assert!(
remapped.is_some(),
"An indirect /Subtype naming CIDFontType2 must still reach the remap"
);
}
}
-68
View File
@@ -1,68 +0,0 @@
%PDF-1.3
%“Œ‹ž ReportLab Generated PDF document (opensource)
1 0 obj
<<
/F1 2 0 R
>>
endobj
2 0 obj
<<
/BaseFont /Helvetica /Encoding /WinAnsiEncoding /Name /F1 /Subtype /Type1 /Type /Font
>>
endobj
3 0 obj
<<
/Contents 7 0 R /MediaBox [ 0 0 612 792 ] /Parent 6 0 R /Resources <<
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>> /Rotate 0 /Trans <<
>>
/Type /Page
>>
endobj
4 0 obj
<<
/PageMode /UseNone /Pages 6 0 R /Type /Catalog
>>
endobj
5 0 obj
<<
/Author (anonymous) /CreationDate (D:20260803112923+00'00') /Creator (anonymous) /Keywords () /ModDate (D:20260803112923+00'00') /Producer (ReportLab PDF Library - \(opensource\))
/Subject (unspecified) /Title (untitled) /Trapped /False
>>
endobj
6 0 obj
<<
/Count 1 /Kids [ 3 0 R ] /Type /Pages
>>
endobj
7 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ] /Length 202
>>
stream
GarW05mr9@&;9NOME,dW.,B;'jAjYq0S4Z`*D9aMA;]$5J)A/$3lESen1?F)ZJsa4$4&N%-%cs)#qW5EVhhbPiRDrAV>MC%.spto@CU"ZdipR'TtFiMR_%m*Hm$N%qL7a"ckkp9T/s[N2"Og377mP*M^akb2XQZ@'l*qT(9bVtDb5+)S&Q.#%E)<]Ao`TSk2AE'/E\fn~>endstream
endobj
xref
0 8
0000000000 65535 f
0000000061 00000 n
0000000092 00000 n
0000000199 00000 n
0000000392 00000 n
0000000460 00000 n
0000000721 00000 n
0000000780 00000 n
trailer
<<
/ID
[<6d7ea1213c5974c78613d5d2a08423b5><6d7ea1213c5974c78613d5d2a08423b5>]
% ReportLab generated PDF document -- digest (opensource)
/Info 5 0 R
/Root 4 0 R
/Size 8
>>
startxref
9072
%%EOF
File diff suppressed because one or more lines are too long
Binary file not shown.
File diff suppressed because it is too large Load Diff
-116
View File
@@ -3926,65 +3926,6 @@ fn encrypted_pdf_decrypts_with_correct_password() {
);
}
/// Regression for the #231 review finding: `extract_pages_markdown`'s
/// `has_template_image` check must be gated the same way
/// `classify_pdf`/`detect_pdf_type` gates it (image_count <= 1, few text
/// ops, low alphanumeric diversity) — not treated as sufficient on its
/// own. The fixture is a real text page with substantial, richly varied
/// body text (>=50 Tj ops) drawn over a full-bleed background image
/// (e.g. letterhead/watermark). Before the fix, has_template_image alone
/// forced needs_ocr=true and discarded the page's clean markdown; now the
/// page must extract normally.
#[test]
fn test_extract_pages_markdown_does_not_ocr_text_page_with_watermark_image() {
let buf = std::fs::read("tests/fixtures/text_page_with_watermark_image.pdf").unwrap();
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
let page = &ext.pages[0];
assert!(
!page.needs_ocr,
"a text page with substantial real text should not be routed to OCR \
just because it has a background image"
);
assert!(
page.markdown.contains("watermark"),
"expected the page's real body text to be preserved, got: {:?}",
page.markdown
);
}
/// Regression for the #231 review finding: `extract_pages_markdown` never
/// checked `has_vector_text` at all, even though `detect_from_document`'s
/// Mixed-type per-page routing always sends vector-outlined-text pages to
/// OCR (outlined glyphs can't be extracted as text). A page with massive
/// path ops (outlined decorative text) plus a short genuine caption would
/// extract that caption cleanly — non-empty, non-garbled — so the
/// existing empty/garbage-text checks alone couldn't catch it.
#[test]
fn test_extract_pages_markdown_ocrs_page_with_vector_outlined_text() {
let buf = std::fs::read("tests/fixtures/vector_outlined_text_with_caption.pdf").unwrap();
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
assert!(
cls.pages_needing_ocr.contains(&1),
"classify_pdf should flag page 1 as needing OCR (vector-outlined text), got: {:?}",
cls.pages_needing_ocr
);
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
let page = &ext.pages[0];
assert!(
page.needs_ocr,
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
);
assert!(
page.markdown.is_empty(),
"a page flagged needs_ocr must not return markdown as if extraction were \
trustworthy, got: {:?}",
page.markdown
);
}
#[test]
fn pdf_options_debug_redacts_password() {
let opts = PdfOptions::new().password("secret123");
@@ -3995,60 +3936,3 @@ fn pdf_options_debug_redacts_password() {
);
assert!(dbg.contains("REDACTED"), "expected redaction marker: {dbg}");
}
/// Regression for #228: a `startxref` pointer corrupted to point at the
/// wrong byte offset (a single flipped digit — a real, common writer bug)
/// must not make the whole file unprocessable. The real classic xref table
/// is still present and findable by scanning for the `xref` keyword; both
/// pypdf and pdfium recover the same way. Before this fix, every entry
/// point raised "Invalid PDF structure" on a file whose object data was
/// otherwise completely intact.
#[test]
fn test_process_pdf_recovers_corrupted_startxref_pointer() {
let result = process_pdf_with_options(
"tests/fixtures/broken_startxref_pointer.pdf",
PdfOptions::new(),
)
.expect("a corrupted startxref pointer should be recoverable, like pypdf/pdfium");
assert_eq!(result.page_count, 1);
let md = result.markdown.unwrap_or_default();
assert!(
md.contains("Order Detail Report by Account") && md.contains("WIDGET ASSEMBLY"),
"recovered document should extract its real text, got: {md:?}"
);
}
/// Regression for #227: `extract_pages_markdown`'s per-page `needs_ocr`
/// must agree with `classify_pdf`/`detect_pdf_type` on the same page. The
/// fixture is a full-page raster "scan" with a single line of genuine
/// native text drawn over it (a header) — the native text extracts
/// perfectly cleanly (no decoding issues, non-empty), so a needs_ocr
/// computation based on text-quality signals alone says `false`, while
/// detection correctly sees a dominant background image and says the page
/// needs OCR. Both must now agree it needs OCR, and the markdown must not
/// be returned as if the extraction were trustworthy.
#[test]
fn test_extract_pages_markdown_agrees_with_classify_on_scan_with_native_header() {
let buf = std::fs::read("tests/fixtures/scan_with_native_header_text.pdf").unwrap();
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
assert!(
cls.pages_needing_ocr.contains(&1),
"classify_pdf should flag page 1 as needing OCR (image-dominated), got: {:?}",
cls.pages_needing_ocr
);
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
let page = &ext.pages[0];
assert!(
page.needs_ocr,
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
);
assert!(
page.markdown.is_empty(),
"a page flagged needs_ocr must not return markdown as if extraction were \
trustworthy, got: {:?}",
page.markdown
);
}