Compare commits
4
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a813de5561 | ||
|
|
217e1745fb | ||
|
|
2a7ad5891d | ||
|
|
1228a2c2ca |
@@ -37,7 +37,6 @@ scripts/
|
||||
|
||||
# Test output
|
||||
test_output/
|
||||
.firecrawl/
|
||||
|
||||
# Python
|
||||
__pycache__/
|
||||
|
||||
@@ -61,8 +61,7 @@ src/
|
||||
|
||||
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
|
||||
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
|
||||
- **Regression suite**: sibling repo `pdf-evals` with ~200 snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing. While iterating, prefer a subset run (`bench.py test -q` for the quick set, or `-s <name>` for a named test set) and save the full `bench.py test` for the final pre-commit check.
|
||||
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
|
||||
- **Regression suite**: sibling repo `pdf-evals` with 179+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
|
||||
|
||||
## Debugging
|
||||
|
||||
|
||||
@@ -61,7 +61,7 @@ src/
|
||||
|
||||
- **Unit tests**: inline `#[cfg(test)] mod tests` in each module with synthetic data.
|
||||
- **Integration tests**: `tests/integration_tests.rs` with fixture PDFs in `tests/fixtures/`.
|
||||
- **Regression suite**: sibling repo `pdf-evals` with ~200 snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing. While iterating, prefer a subset run (`bench.py test -q` for the quick set, or `-s <name>` for a named test set) and save the full `bench.py test` for the final pre-commit check.
|
||||
- **Regression suite**: sibling repo `pdf-evals` with 187+ snapshot PDFs. Run `cargo build --release` then `bench.py test` in that repo before committing.
|
||||
- **Semantic quality**: run `bench.py score` in `pdf-evals` for the semantic verdict (TEDS + MHS + reading order + char/word + list preservation, composited). Character-level diff alone misclassifies structural improvements (e.g., column-detection rewrites) as regressions — `score` is the tie-breaker. See `pdf-evals/CLAUDE.md` "Semantic scoring".
|
||||
|
||||
## Debugging
|
||||
|
||||
+7
-7
@@ -858,22 +858,22 @@
|
||||
<p>Evaluated on the <a class="text-link" href="https://github.com/opendataloader-project/opendataloader-bench">opendataloader-bench</a> corpus of 200 PDFs. This comparison covers local engines without model-based PDF parsing, with OCR disabled. Higher scores are better.</p>
|
||||
</div>
|
||||
<div class="benchmark-card">
|
||||
<div class="benchmark-top"><span><strong>200 PDFs</strong> · OpenDataLoader benchmark</span><span>Apple M4 Pro · median of 5 runs</span></div>
|
||||
<div class="benchmark-top"><span><strong>200 PDFs</strong> · OpenDataLoader benchmark</span><span>Apple M4 Pro · median of 3 runs</span></div>
|
||||
<div class="table-scroll">
|
||||
<table aria-label="PDF extraction benchmark results">
|
||||
<thead>
|
||||
<tr><th>Engine</th><th>Overall</th><th>Reading order</th><th>Tables</th><th>Headings</th><th>Complete run</th></tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr class="highlight"><td>pdf-inspector</td><td>0.875</td><td>0.915</td><td>0.814</td><td>0.788</td><td>0.470s</td></tr>
|
||||
<tr><td>LiteParse</td><td>0.873</td><td>0.913</td><td>0.693</td><td>0.811</td><td>0.750s</td></tr>
|
||||
<tr><td>OpenDataLoader</td><td>0.831</td><td>0.902</td><td>0.489</td><td>0.739</td><td>2.569s</td></tr>
|
||||
<tr><td>PyMuPDF4LLM</td><td>0.735</td><td>0.886</td><td>0.401</td><td>0.424</td><td>17.117s</td></tr>
|
||||
<tr><td>MarkItDown</td><td>0.589</td><td>0.844</td><td>0.273</td><td>0.000</td><td>16.165s</td></tr>
|
||||
<tr class="highlight"><td>pdf-inspector</td><td>0.875</td><td>0.915</td><td>0.814</td><td>0.788</td><td>2.8s</td></tr>
|
||||
<tr><td>LiteParse</td><td>0.870</td><td>0.908</td><td>0.693</td><td>0.811</td><td>13.9s</td></tr>
|
||||
<tr><td>OpenDataLoader</td><td>0.843</td><td>0.912</td><td>0.489</td><td>0.760</td><td>9.8s</td></tr>
|
||||
<tr><td>PyMuPDF4LLM</td><td>0.735</td><td>0.886</td><td>0.401</td><td>0.424</td><td>15.5s</td></tr>
|
||||
<tr><td>MarkItDown</td><td>0.583</td><td>0.879</td><td>0.000</td><td>0.000</td><td>6.7s</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
<div class="benchmark-note">Refreshed July 31, 2026. Scores use the benchmark’s NID, TEDS, and MHS evaluators; speed is the median of five alternating or rotating complete corpus runs after an excluded warm-up. <a class="text-link" href="https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results">Versions and raw artifacts</a>.</div>
|
||||
<div class="benchmark-note">Refreshed July 16, 2026. Scores use the benchmark’s NID, TEDS, and MHS evaluators.</div>
|
||||
</div>
|
||||
<div class="best-fit">
|
||||
<strong>Best fit</strong>
|
||||
|
||||
+1
-65
@@ -1659,15 +1659,7 @@ fn hex_val(b: u8) -> Option<u8> {
|
||||
/// Standard page: 612x792 points (US Letter) = ~485,000 sq points
|
||||
/// At 2x resolution that's ~1.9M pixels, so we use 250K pixels as threshold
|
||||
/// (accounting for varying DPI and page sizes)
|
||||
/// Returns `(has_images, total_image_area, has_template_image)` for a page.
|
||||
/// `has_template_image` means a single large (>50% page coverage)
|
||||
/// background image — the signal `classify_pdf`/`detect_pdf_type` uses to
|
||||
/// route a page to OCR regardless of any incidental native text drawn over
|
||||
/// it. Exposed at crate visibility so extraction-side per-page `needs_ocr`
|
||||
/// computation (`extract_pages_markdown_mem`) can consult the same signal
|
||||
/// instead of maintaining its own, independent notion of "needs OCR" that
|
||||
/// can silently disagree with detection — see #227.
|
||||
pub(crate) fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
|
||||
fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u64, bool) {
|
||||
// Threshold: image covering roughly half a page at 150+ DPI
|
||||
// 612 * 792 / 2 * (150/72)^2 ≈ 1M pixels, but we'll be conservative
|
||||
const TEMPLATE_IMAGE_THRESHOLD: u64 = 500_000; // 500K pixels
|
||||
@@ -1749,62 +1741,6 @@ pub(crate) fn analyze_page_images(doc: &Document, page_id: ObjectId) -> (bool, u
|
||||
(has_images, total_area, has_template_image)
|
||||
}
|
||||
|
||||
/// Computes both `(needs_ocr_for_template_image, has_vector_text)` for a
|
||||
/// page from a single shared `analyze_page_content` pass — that call
|
||||
/// decompresses and scans every content stream (page + XObjects) plus
|
||||
/// image coverage, so `extract_pages_markdown_mem` must not invoke it
|
||||
/// twice per page (once per signal) the way `detect_from_document` avoids
|
||||
/// by caching its per-page `PageAnalysis`.
|
||||
///
|
||||
/// `needs_ocr_for_template_image` is true when a page's template image
|
||||
/// should be treated as a scan needing OCR — a single full-page background
|
||||
/// image with little/no real text — rather than a text page that happens
|
||||
/// to carry a watermark, letterhead, or figure. Mirrors the two distinct
|
||||
/// signals classification uses to route a template-image page to OCR:
|
||||
///
|
||||
/// 1. `looks_like_scan`: image_count <= 1, few text operators (<50), and
|
||||
/// low alphanumeric diversity in raw string operands (unless decodable
|
||||
/// CID/ToUnicode fonts explain that away) — the gate used for
|
||||
/// `pages_with_template_images` and Mixed-type per-page routing.
|
||||
/// 2. Insufficient real text volume, using `DetectionConfig::default()`'s
|
||||
/// `min_text_ops_per_page` (3) — the same threshold Mixed-type per-page
|
||||
/// routing applies via `text_operator_count < config.min_text_ops_per_page
|
||||
/// && has_images` (simplified here since a template image implies
|
||||
/// `has_images`). Deliberately *not* the higher `effective_min_ops`
|
||||
/// floor (`min_text_ops_per_page.max(10)`) that whole-document
|
||||
/// `PdfType::ImageBased`/`Scanned` classification uses for
|
||||
/// `pages_with_text` — that's a cross-page aggregate decision this
|
||||
/// per-page function has no way to replicate exactly, and the lower
|
||||
/// per-page threshold is the one a single page's own signals can
|
||||
/// actually agree with.
|
||||
///
|
||||
/// `has_vector_text` is true when a page has vector-outlined text (glyphs
|
||||
/// drawn as paths rather than shown via text-showing operators) —
|
||||
/// `detect_from_document`'s Mixed-type per-page routing always sends
|
||||
/// these pages to OCR, independent of any template-image check, since
|
||||
/// outlined glyphs can't be extracted as text at all.
|
||||
///
|
||||
/// Exposed at crate visibility so `extract_pages_markdown_mem` can apply
|
||||
/// the same gates classification needs elsewhere instead of treating the
|
||||
/// raw signals alone as sufficient — see #227/#231.
|
||||
pub(crate) fn page_ocr_signals(doc: &Document, page_id: ObjectId) -> (bool, bool) {
|
||||
let analysis = analyze_page_content(doc, page_id);
|
||||
|
||||
let needs_ocr_for_template_image = if !analysis.has_template_image {
|
||||
false
|
||||
} else {
|
||||
let alphanum_low = analysis.unique_alphanum_chars < 10
|
||||
&& !(analysis.has_decodable_text_fonts && analysis.text_operator_count >= 10);
|
||||
let looks_like_scan =
|
||||
analysis.image_count <= 1 && analysis.text_operator_count < 50 && alphanum_low;
|
||||
let insufficient_text =
|
||||
analysis.text_operator_count < DetectionConfig::default().min_text_ops_per_page;
|
||||
looks_like_scan || insufficient_text
|
||||
};
|
||||
|
||||
(needs_ocr_for_template_image, analysis.has_vector_text)
|
||||
}
|
||||
|
||||
/// Recursively collect image dimensions from XObject resources,
|
||||
/// including images nested inside Form XObjects.
|
||||
fn collect_images_from_resources(
|
||||
|
||||
@@ -41,17 +41,6 @@ fn strip_pdf_comments(data: &[u8]) -> Vec<u8> {
|
||||
while i < data.len() {
|
||||
let b = data[i];
|
||||
match b {
|
||||
// Inside a string literal, a backslash escapes the next byte —
|
||||
// `\(`, `\)`, and `\\` must not touch the nesting depth, or a
|
||||
// later `%` glyph inside a string gets stripped as a comment,
|
||||
// corrupting the stream.
|
||||
b'\\' if in_string > 0 => {
|
||||
result.push(b);
|
||||
if let Some(&next) = data.get(i + 1) {
|
||||
result.push(next);
|
||||
i += 1;
|
||||
}
|
||||
}
|
||||
b'(' if !in_hex_string => {
|
||||
in_string += 1;
|
||||
result.push(b);
|
||||
@@ -1865,26 +1854,4 @@ BT 30 700 Tm <41> Tj ET";
|
||||
"ET should be preserved after comment stripping"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_strip_pdf_comments_escaped_parens() {
|
||||
// An escaped `\)` must not close the string: the `%` after it is
|
||||
// still string content, not a comment (subset fonts routinely map
|
||||
// glyphs to `%` and to escaped parens in the same TJ array).
|
||||
let input = b"[ (a\\)b) 1 (%) 1 (c) ] TJ\n";
|
||||
let output = strip_pdf_comments(input);
|
||||
assert_eq!(output, input.to_vec());
|
||||
|
||||
// Same for an escaped `\(` — must not open a phantom string that
|
||||
// shields a real comment.
|
||||
let input = b"(x\\(y) Tj % real comment\nET\n";
|
||||
let output = strip_pdf_comments(input);
|
||||
assert_eq!(output, b"(x\\(y) Tj \nET\n");
|
||||
|
||||
// Escaped backslash before a real close-paren: `\\` ends the escape,
|
||||
// the `)` does close the string, and the comment is stripped.
|
||||
let input = b"(x\\\\) Tj % comment\nET\n";
|
||||
let output = strip_pdf_comments(input);
|
||||
assert_eq!(output, b"(x\\\\) Tj \nET\n");
|
||||
}
|
||||
}
|
||||
|
||||
+7
-311
@@ -491,14 +491,7 @@ pub fn extract_pages_markdown_mem(
|
||||
|
||||
// Tables need the original numeric cells; columns use folio-cleaned
|
||||
// evidence so removed page numbers cannot create false layout metadata.
|
||||
let chart_regions = markdown::chart_regions_by_page(&all_items, &all_rects, &all_lines);
|
||||
let complexity = compute_layout_complexity_with_chart_regions(
|
||||
&all_items,
|
||||
&filtered_items,
|
||||
&all_rects,
|
||||
&all_lines,
|
||||
&chart_regions,
|
||||
);
|
||||
let complexity = compute_layout_complexity(&all_items, &filtered_items, &all_rects, &all_lines);
|
||||
|
||||
// Compute font stats from full document (cross-page consistency).
|
||||
let font_stats = markdown::analysis::calculate_font_stats_from_items(&filtered_items);
|
||||
@@ -516,7 +509,6 @@ pub fn extract_pages_markdown_mem(
|
||||
let mut results = Vec::with_capacity(pages_slice.len());
|
||||
let mut pages_needing_ocr = Vec::new();
|
||||
let mut ocr_reasons_by_page = BTreeMap::new();
|
||||
let lopdf_pages = doc.get_pages();
|
||||
|
||||
for &page_0idx in pages_slice {
|
||||
// Out-of-range pages → empty + needs_ocr
|
||||
@@ -550,25 +542,6 @@ pub fn extract_pages_markdown_mem(
|
||||
let has_gid = gid_pages.contains(&page_1idx);
|
||||
let has_text_quality_issue = text_quality.pages_needing_ocr.contains(&page_1idx);
|
||||
|
||||
// A page can extract cleanly (no decoding issues, non-empty text)
|
||||
// while still being fundamentally a scan: a full-page raster with
|
||||
// a little genuine native text drawn over it (a header, a stamp, a
|
||||
// cover-sheet annotation). Text-quality signals alone can't see
|
||||
// that — consult the same "large background image" signal
|
||||
// classify_pdf/detect_pdf_type already uses, so the two APIs can't
|
||||
// silently disagree on whether a page needs OCR. See #227.
|
||||
// Also covers vector-outlined text (glyphs drawn as paths, not
|
||||
// shown via a text-showing operator): a hybrid page with real
|
||||
// embedded-font body text elsewhere would otherwise still extract
|
||||
// non-empty, non-garbled markdown and miss OCR routing entirely.
|
||||
// detect_from_document's Mixed-type per-page routing always sends
|
||||
// these pages to OCR; mirror that here too. Both signals share one
|
||||
// analyze_page_content pass — see page_ocr_signals's doc comment.
|
||||
let (has_template_image, has_vector_text) = lopdf_pages
|
||||
.get(&page_1idx)
|
||||
.map(|&page_id| detector::page_ocr_signals(&doc, page_id))
|
||||
.unwrap_or((false, false));
|
||||
|
||||
// Build markdown with document-wide font stats
|
||||
let options = MarkdownOptions {
|
||||
base_font_size: Some(font_stats.most_common_size),
|
||||
@@ -592,7 +565,6 @@ pub fn extract_pages_markdown_mem(
|
||||
page_count,
|
||||
prefiltered_page_number_pages: Some(&removed_page_number_pages),
|
||||
prefiltered_page_number_mask: Some(&page_number_removal_mask),
|
||||
precomputed_chart_regions: Some(&chart_regions),
|
||||
},
|
||||
)
|
||||
};
|
||||
@@ -606,20 +578,10 @@ pub fn extract_pages_markdown_mem(
|
||||
OCR_REASON_SUSPECTED_GARBLED_TEXT,
|
||||
);
|
||||
}
|
||||
if has_template_image {
|
||||
add_ocr_reason(&mut ocr_reasons_by_page, page_1idx, OCR_REASON_SCANNED);
|
||||
}
|
||||
if has_vector_text {
|
||||
add_ocr_reason(&mut ocr_reasons_by_page, page_1idx, OCR_REASON_VECTOR_TEXT);
|
||||
}
|
||||
let ocr_reason = page_ocr_reason(&ocr_reasons_by_page, page_1idx);
|
||||
|
||||
let needs_ocr = ocr_reason.is_some()
|
||||
|| md.trim().is_empty()
|
||||
|| has_gid
|
||||
|| is_garbage_text(&md)
|
||||
|| has_template_image
|
||||
|| has_vector_text;
|
||||
let needs_ocr =
|
||||
ocr_reason.is_some() || md.trim().is_empty() || has_gid || is_garbage_text(&md);
|
||||
|
||||
if needs_ocr {
|
||||
pages_needing_ocr.push(page_1idx);
|
||||
@@ -3553,7 +3515,6 @@ fn repair_pdf_container_candidates(buf: &[u8]) -> Vec<Vec<u8>> {
|
||||
let mut candidates = Vec::new();
|
||||
|
||||
add_repair_candidate(&mut candidates, append_missing_eof_marker(buf), buf);
|
||||
add_repair_candidate(&mut candidates, recover_startxref_pointer(buf), buf);
|
||||
|
||||
let stripped = strip_leading_pdf_container_bytes(buf);
|
||||
if let Some(stripped_buf) = stripped.as_deref() {
|
||||
@@ -3563,112 +3524,11 @@ fn repair_pdf_container_candidates(buf: &[u8]) -> Vec<Vec<u8>> {
|
||||
append_missing_eof_marker(stripped_buf),
|
||||
buf,
|
||||
);
|
||||
add_repair_candidate(
|
||||
&mut candidates,
|
||||
recover_startxref_pointer(stripped_buf),
|
||||
buf,
|
||||
);
|
||||
}
|
||||
|
||||
candidates
|
||||
}
|
||||
|
||||
/// Some PDF writers emit a `startxref` pointer that doesn't actually point
|
||||
/// at the cross-reference table — a single corrupted byte in the offset is
|
||||
/// enough. lopdf trusts that pointer outright and fails to load rather than
|
||||
/// searching for the real table, unlike pypdf/pdfium which both recover by
|
||||
/// locating it directly. This finds the real (classic, non-stream) `xref`
|
||||
/// table by scanning for the keyword — validating that a plausible
|
||||
/// subsection header follows, not just any standalone "xref" token, since
|
||||
/// this crate processes untrusted input and a coincidental match inside
|
||||
/// unrelated stream/string content must not get "repaired" against a bogus
|
||||
/// offset (lopdf would then load successfully against garbage instead of
|
||||
/// returning a clean error) — and appends a corrected trailing
|
||||
/// `startxref`/`%%EOF` block. lopdf's own `get_xref_start` always uses the
|
||||
/// *last* `%%EOF` in the final 512 bytes of the buffer, so ours
|
||||
/// transparently supersedes the broken one without needing to touch
|
||||
/// anything already in the file.
|
||||
///
|
||||
/// Doesn't cover cross-reference *streams* (`N 0 obj << /Type /XRef ...`,
|
||||
/// used by some PDF 1.5+ writers instead of a classic table) — recovering
|
||||
/// those needs the containing object's number, not just a byte offset.
|
||||
fn recover_startxref_pointer(buf: &[u8]) -> Option<Vec<u8>> {
|
||||
let xref_pos = find_last_valid_xref_table_start(buf)?;
|
||||
|
||||
let mut repaired = Vec::with_capacity(buf.len() + 32);
|
||||
repaired.extend_from_slice(buf);
|
||||
if !repaired.ends_with(b"\n") {
|
||||
repaired.push(b'\n');
|
||||
}
|
||||
repaired.extend_from_slice(format!("startxref\n{xref_pos}\n%%EOF\n").as_bytes());
|
||||
Some(repaired)
|
||||
}
|
||||
|
||||
/// Finds the last standalone `xref` token in `buf` that is immediately
|
||||
/// followed by a plausible classic cross-reference subsection header
|
||||
/// (`<start-id> <count>`, e.g. "0 6") — the shape every real classic xref
|
||||
/// table starts with. A single reverse byte scan: O(n) even on a
|
||||
/// pathological buffer with many non-matching or non-standalone "xref"
|
||||
/// occurrences, unlike repeatedly re-searching a shrinking prefix.
|
||||
fn find_last_valid_xref_table_start(buf: &[u8]) -> Option<usize> {
|
||||
const KEYWORD: &[u8] = b"xref";
|
||||
if buf.len() < KEYWORD.len() {
|
||||
return None;
|
||||
}
|
||||
let mut pos = buf.len() - KEYWORD.len();
|
||||
loop {
|
||||
if &buf[pos..pos + KEYWORD.len()] == KEYWORD {
|
||||
let before_ok = pos == 0 || buf[pos - 1].is_ascii_whitespace();
|
||||
let after_ok = buf
|
||||
.get(pos + KEYWORD.len())
|
||||
.is_none_or(|c| c.is_ascii_whitespace());
|
||||
if before_ok && after_ok && looks_like_xref_subsection_header(buf, pos + KEYWORD.len())
|
||||
{
|
||||
return Some(pos);
|
||||
}
|
||||
}
|
||||
if pos == 0 {
|
||||
return None;
|
||||
}
|
||||
pos -= 1;
|
||||
}
|
||||
}
|
||||
|
||||
/// Checks that `buf[pos..]` starts (after whitespace) with two
|
||||
/// whitespace-separated runs of ASCII digits — `<start-id> <count>`, the
|
||||
/// first subsection header of a classic PDF cross-reference table.
|
||||
fn looks_like_xref_subsection_header(buf: &[u8], pos: usize) -> bool {
|
||||
fn skip_ws(buf: &[u8], mut pos: usize) -> usize {
|
||||
while buf.get(pos).is_some_and(u8::is_ascii_whitespace) {
|
||||
pos += 1;
|
||||
}
|
||||
pos
|
||||
}
|
||||
fn skip_digits(buf: &[u8], mut pos: usize) -> usize {
|
||||
while buf.get(pos).is_some_and(u8::is_ascii_digit) {
|
||||
pos += 1;
|
||||
}
|
||||
pos
|
||||
}
|
||||
|
||||
let pos = skip_ws(buf, pos);
|
||||
let after_first_digits = skip_digits(buf, pos);
|
||||
if after_first_digits == pos {
|
||||
return false; // no start-id
|
||||
}
|
||||
let sep = skip_ws(buf, after_first_digits);
|
||||
if sep == after_first_digits {
|
||||
return false; // start-id and count must be whitespace-separated
|
||||
}
|
||||
let after_count = skip_digits(buf, sep);
|
||||
if after_count == sep {
|
||||
return false; // no count
|
||||
}
|
||||
// The count run must end at whitespace/buffer-end, not run into trailing
|
||||
// garbage (e.g. a coincidental "xref\n0 6garbage" in stream content).
|
||||
buf.get(after_count).is_none_or(u8::is_ascii_whitespace)
|
||||
}
|
||||
|
||||
fn add_repair_candidate(
|
||||
candidates: &mut Vec<Vec<u8>>,
|
||||
candidate: Option<Vec<u8>>,
|
||||
@@ -3956,14 +3816,7 @@ fn process_document(
|
||||
|
||||
let text_quality = analyze_text_quality(&items);
|
||||
merge_ocr_reasons(&mut ocr_reasons_by_page, text_quality.reasons_by_page);
|
||||
let chart_regions = markdown::chart_regions_by_page(&items, &rects, &lines);
|
||||
let layout = compute_layout_complexity_with_chart_regions(
|
||||
&items,
|
||||
&layout_items,
|
||||
&rects,
|
||||
&lines,
|
||||
&chart_regions,
|
||||
);
|
||||
let layout = compute_layout_complexity(&items, &layout_items, &rects, &lines);
|
||||
|
||||
let md = if options.mode == ProcessMode::Analyze {
|
||||
None
|
||||
@@ -3980,7 +3833,6 @@ fn process_document(
|
||||
page_count,
|
||||
prefiltered_page_number_pages: Some(&removed_pages),
|
||||
prefiltered_page_number_mask: Some(removal_mask.as_slice()),
|
||||
precomputed_chart_regions: Some(&chart_regions),
|
||||
},
|
||||
))
|
||||
};
|
||||
@@ -5781,29 +5633,11 @@ fn select_items_with_document_folio_context(
|
||||
}
|
||||
|
||||
/// Analyse extracted items and rects for layout complexity.
|
||||
#[cfg(test)]
|
||||
fn compute_layout_complexity(
|
||||
items: &[types::TextItem],
|
||||
column_items: &[types::TextItem],
|
||||
rects: &[types::PdfRect],
|
||||
lines: &[types::PdfLine],
|
||||
) -> LayoutComplexity {
|
||||
let page_chart_regions = markdown::chart_regions_by_page(items, rects, lines);
|
||||
compute_layout_complexity_with_chart_regions(
|
||||
items,
|
||||
column_items,
|
||||
rects,
|
||||
lines,
|
||||
&page_chart_regions,
|
||||
)
|
||||
}
|
||||
|
||||
fn compute_layout_complexity_with_chart_regions(
|
||||
items: &[types::TextItem],
|
||||
column_items: &[types::TextItem],
|
||||
rects: &[types::PdfRect],
|
||||
lines: &[types::PdfLine],
|
||||
page_chart_regions: &markdown::PageChartRegions,
|
||||
) -> LayoutComplexity {
|
||||
use markdown::analysis::calculate_font_stats_from_items;
|
||||
|
||||
@@ -5825,10 +5659,6 @@ fn compute_layout_complexity_with_chart_regions(
|
||||
let owned_items: Vec<types::TextItem> = page_items.iter().map(|i| (*i).clone()).collect();
|
||||
let page_content_width = tables::content_width(&owned_items);
|
||||
let bands = markdown::split_side_by_side(&owned_items);
|
||||
let chart_regions = page_chart_regions
|
||||
.get(&page)
|
||||
.map(Vec::as_slice)
|
||||
.unwrap_or_default();
|
||||
|
||||
let band_ranges: Vec<(f32, f32)> = if bands.is_empty() {
|
||||
// Single region — use sentinel range that includes everything
|
||||
@@ -5843,8 +5673,7 @@ fn compute_layout_complexity_with_chart_regions(
|
||||
let band_items: Vec<types::TextItem> = owned_items
|
||||
.iter()
|
||||
.filter(|item| {
|
||||
(x_lo == f32::MIN || (item.x >= x_lo - margin && item.x < x_hi + margin))
|
||||
&& !markdown::item_is_in_chart_region(item, chart_regions)
|
||||
x_lo == f32::MIN || (item.x >= x_lo - margin && item.x < x_hi + margin)
|
||||
})
|
||||
.cloned()
|
||||
.collect();
|
||||
@@ -5896,20 +5725,8 @@ fn compute_layout_complexity_with_chart_regions(
|
||||
}
|
||||
|
||||
let mut pages_with_columns: Vec<u32> = Vec::new();
|
||||
for &page in &seen_pages {
|
||||
let chart_regions = page_chart_regions
|
||||
.get(&page)
|
||||
.map(Vec::as_slice)
|
||||
.unwrap_or_default();
|
||||
let page_column_items: Vec<types::TextItem> = column_items
|
||||
.iter()
|
||||
.filter(|item| {
|
||||
item.page == page && !markdown::item_is_in_chart_region(item, chart_regions)
|
||||
})
|
||||
.cloned()
|
||||
.collect();
|
||||
let cols =
|
||||
extractor::detect_columns(&page_column_items, page, pages_with_tables.contains(&page));
|
||||
for page in seen_pages {
|
||||
let cols = extractor::detect_columns(column_items, page, pages_with_tables.contains(&page));
|
||||
if cols.len() >= 2 {
|
||||
pages_with_columns.push(page);
|
||||
}
|
||||
@@ -6138,66 +5955,6 @@ mod tests {
|
||||
assert!(filtered.pages_with_columns.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dense_chart_panel_is_not_reported_as_a_table() {
|
||||
let mut items: Vec<TextItem> = (0..8)
|
||||
.flat_map(|row| {
|
||||
(0..6).map(move |column| {
|
||||
test_item(
|
||||
&format!("{}", row * 10 + column),
|
||||
105.0 + column as f32 * 35.0,
|
||||
525.0 - row as f32 * 15.0,
|
||||
24.0,
|
||||
10.0,
|
||||
)
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
for row in 0..6 {
|
||||
items.push(test_item(
|
||||
"Left column prose continues here",
|
||||
80.0,
|
||||
320.0 - row as f32 * 15.0,
|
||||
160.0,
|
||||
10.0,
|
||||
));
|
||||
items.push(test_item(
|
||||
"Right column prose continues here",
|
||||
300.0,
|
||||
320.0 - row as f32 * 15.0,
|
||||
160.0,
|
||||
10.0,
|
||||
));
|
||||
}
|
||||
let mut lines: Vec<PdfLine> = (0..30)
|
||||
.map(|column| PdfLine {
|
||||
x1: 100.0 + column as f32 * 8.0,
|
||||
y1: 400.0,
|
||||
x2: 100.0 + column as f32 * 8.0,
|
||||
y2: 550.0,
|
||||
page: 1,
|
||||
})
|
||||
.collect();
|
||||
lines.extend((0..6).map(|row| PdfLine {
|
||||
x1: 100.0,
|
||||
y1: 400.0 + row as f32 * 30.0,
|
||||
x2: 332.0,
|
||||
y2: 400.0 + row as f32 * 30.0,
|
||||
page: 1,
|
||||
}));
|
||||
let rects = vec![PdfRect {
|
||||
x: 80.0,
|
||||
y: 350.0,
|
||||
width: 280.0,
|
||||
height: 240.0,
|
||||
page: 1,
|
||||
}];
|
||||
|
||||
let complexity = compute_layout_complexity(&items, &items, &rects, &lines);
|
||||
|
||||
assert!(complexity.pages_with_tables.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn page_selection_keeps_document_wide_folio_layout_decisions() {
|
||||
let mut items = Vec::new();
|
||||
@@ -7201,65 +6958,4 @@ mod tests {
|
||||
// Pre-filled cell was not touched.
|
||||
assert_eq!(cells[1].text, "Pre-filled");
|
||||
}
|
||||
|
||||
// -- recover_startxref_pointer / find_last_valid_xref_table_start ------
|
||||
//
|
||||
// Direct unit tests on the byte-level scan, addressing review feedback
|
||||
// on #230: a coincidental standalone "xref" token that isn't actually
|
||||
// followed by a subsection header (start-id + count) must not be
|
||||
// treated as a real table — accepting it would let lopdf "succeed"
|
||||
// against a bogus offset and silently return garbled/empty content
|
||||
// instead of a clean error.
|
||||
|
||||
#[test]
|
||||
fn find_xref_rejects_standalone_token_without_subsection_header() {
|
||||
// "xref" appears as a real standalone word, but nothing that looks
|
||||
// like "<start-id> <count>" follows it.
|
||||
let buf = b"Please refer to the xref appendix for details.";
|
||||
assert_eq!(find_last_valid_xref_table_start(buf), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn find_xref_accepts_real_classic_table_header() {
|
||||
let buf = b"garbage\nxref\n0 6\n0000000000 65535 f \n%%EOF";
|
||||
let pos = find_last_valid_xref_table_start(buf).expect("should find the real table");
|
||||
assert_eq!(&buf[pos..pos + 4], b"xref");
|
||||
assert_eq!(&buf[pos..], b"xref\n0 6\n0000000000 65535 f \n%%EOF");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn find_xref_skips_coincidental_match_and_finds_real_table_before_it() {
|
||||
// A coincidental "xref" (no subsection header) appears *after* the
|
||||
// real table in the buffer — the scan must not stop at the first
|
||||
// (rightmost) standalone token it finds; it must keep looking
|
||||
// backward until one actually validates.
|
||||
let buf = b"xref\n0 3\n0000000000 65535 f \ntrailer\nsee the xref\n";
|
||||
let pos = find_last_valid_xref_table_start(buf).expect("should find the real table");
|
||||
assert_eq!(pos, 0);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn find_xref_rejects_substring_of_startxref() {
|
||||
// "xref" is a substring of "startxref" but isn't a standalone
|
||||
// token there (not preceded by whitespace) — must not match, even
|
||||
// though a number immediately follows it.
|
||||
let buf = b"startxref\n1234\n%%EOF";
|
||||
assert_eq!(find_last_valid_xref_table_start(buf), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn find_xref_rejects_count_run_with_trailing_garbage() {
|
||||
// "xref\n0 6garbage" has the right shape (digits, whitespace,
|
||||
// digits) but the count run doesn't end at whitespace/EOF — it
|
||||
// runs straight into non-digit garbage, so this must not be
|
||||
// accepted as a real subsection header.
|
||||
let buf = b"xref\n0 6garbage\n%%EOF";
|
||||
assert_eq!(find_last_valid_xref_table_start(buf), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn recover_startxref_pointer_returns_none_without_a_valid_table() {
|
||||
let buf = b"Please refer to the xref appendix for details.";
|
||||
assert!(recover_startxref_pointer(buf).is_none());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -171,74 +171,6 @@ pub(crate) fn is_toc_marker_heading(text: &str) -> bool {
|
||||
/// equation and absent from name-plus-number headings. A bare trailing colon
|
||||
/// is NOT a fragment signal either: real headings frequently end with colons
|
||||
/// ("Procedure:", "Steps for Using the Microscope:").
|
||||
/// True when the line opens with a section number ("3.", "2.1.4", "IV)").
|
||||
///
|
||||
/// Mirrors the acceptance of `heading::parse_numbering` rather than the
|
||||
/// stricter `convert::starts_with_section_number`, which deliberately
|
||||
/// requires two components because it bypasses isolation checks. Here a
|
||||
/// single "1." counts: numbering is independent evidence of a heading, and
|
||||
/// `heading.rs` applies its numbered-prefix allowance *after* consulting
|
||||
/// `is_heading_fragment`, so without this exemption a numbered
|
||||
/// sentence-case heading would be vetoed before that allowance can run.
|
||||
fn starts_with_numbering_prefix(t: &str) -> bool {
|
||||
let Some(first) = t.split_whitespace().next() else {
|
||||
return false;
|
||||
};
|
||||
let has_delimiter = first.ends_with(['.', ')', ':']);
|
||||
let token = first.trim_end_matches(['.', ')', ':']);
|
||||
if token.is_empty() {
|
||||
return false;
|
||||
}
|
||||
let parts: Vec<&str> = token.split('.').collect();
|
||||
let decimal = parts
|
||||
.iter()
|
||||
.all(|p| !p.is_empty() && p.len() <= 3 && p.chars().all(|c| c.is_ascii_digit()));
|
||||
if decimal {
|
||||
// "1." / "2.1." carry a delimiter; "2.3 Title" is written without
|
||||
// one, so a multi-component number is accepted bare. A bare single
|
||||
// number ("3 apples") is not — that is ordinary prose.
|
||||
return has_delimiter || parts.len() >= 2;
|
||||
}
|
||||
// Roman numerals go through the heading parser's own grammar so the two
|
||||
// agree: uppercase I/V/X/L/C only, at most 8 characters. A looser rule
|
||||
// here would exempt markers the parser rejects — "iv)" or "d)" from an
|
||||
// alphabetical list — letting an ordinary list item bypass the veto and
|
||||
// reach heading promotion.
|
||||
//
|
||||
// A delimiter is also required: a bare leading "I" is the pronoun far
|
||||
// more often than a section number.
|
||||
has_delimiter && crate::markdown::heading::roman_value(token).is_some()
|
||||
}
|
||||
|
||||
/// True when the line reads as a title rather than a sentence: every
|
||||
/// content word (ignoring minor words) starts uppercase. Used to spare real
|
||||
/// headings from the dangling-verb veto — "Bond Yields" is a section title,
|
||||
/// "the method yields" is a stranded clause, and only the casing tells them
|
||||
/// apart.
|
||||
fn looks_title_case(t: &str) -> bool {
|
||||
const MINOR: &[&str] = &[
|
||||
"a", "an", "the", "of", "and", "or", "for", "to", "in", "on", "at", "by", "with", "from",
|
||||
"as", "is", "are", "that", "than", "into",
|
||||
];
|
||||
let mut content = 0usize;
|
||||
let mut capitalized = 0usize;
|
||||
for w in t.split_whitespace() {
|
||||
let cleaned: String = w.chars().filter(|c| c.is_alphabetic()).collect();
|
||||
if cleaned.is_empty() {
|
||||
continue;
|
||||
}
|
||||
if MINOR.contains(&cleaned.to_lowercase().as_str()) {
|
||||
continue;
|
||||
}
|
||||
content += 1;
|
||||
if cleaned.chars().next().is_some_and(char::is_uppercase) {
|
||||
capitalized += 1;
|
||||
}
|
||||
}
|
||||
// A single content word ("Yields") is a title by default.
|
||||
content == 0 || capitalized == content
|
||||
}
|
||||
|
||||
pub(crate) fn is_heading_fragment(text: &str) -> bool {
|
||||
let t = text.trim_end();
|
||||
|
||||
@@ -312,133 +244,9 @@ pub(crate) fn is_heading_fragment(text: &str) -> bool {
|
||||
if t.ends_with(':') && t.split_whitespace().any(is_equation_number) {
|
||||
return true;
|
||||
}
|
||||
|
||||
// Dangling clause: a stranded sentence lead-in ends on a relational
|
||||
// verb with no terminal punctuation — "Note that the exact error equals"
|
||||
// left ahead of its formula when a phantom table dissolved.
|
||||
//
|
||||
// Gated on the line reading as prose rather than a title. Case is the
|
||||
// discriminator the trailing word alone cannot provide: a heading is
|
||||
// title case ("Bond Yields", "The Method Yields") while a stranded
|
||||
// lead-in is sentence case ("the method yields"). Without this gate the
|
||||
// veto eats real headings — "Bond Yields", "Crop Yields" and any wrapped
|
||||
// title-case heading the preprocessor failed to merge.
|
||||
if !t.ends_with(['.', '!', '?', ':', ';', ')', ']'])
|
||||
&& !looks_title_case(t)
|
||||
&& !starts_with_numbering_prefix(t)
|
||||
{
|
||||
if let Some(last) = t.split_whitespace().next_back() {
|
||||
let word: String = last
|
||||
.trim_matches(|c: char| !c.is_alphanumeric())
|
||||
.to_lowercase();
|
||||
// Relational verbs only, and only those with no common noun
|
||||
// sense. "yields" was dropped for exactly that reason: "Bond
|
||||
// Yields" is a real section title. Function words, copulas and
|
||||
// auxiliaries were measured and rejected outright — a heading
|
||||
// that wraps across lines ends on those, and suppressing them
|
||||
// destroyed real IRS Publication 17 headings.
|
||||
const DANGLING_TAIL: &[&str] =
|
||||
&["equals", "denotes", "implies", "satisfies", "signifies"];
|
||||
if DANGLING_TAIL.contains(&word.as_str()) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
false
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod fragment_heading_tests {
|
||||
use super::is_heading_fragment;
|
||||
|
||||
#[test]
|
||||
fn dangling_tail_marks_stranded_clause() {
|
||||
// opendataloader 01030000000144: left behind when a phantom table
|
||||
// dissolved, ahead of its formula on the next line.
|
||||
assert!(is_heading_fragment("Note that the exact error equals"));
|
||||
assert!(is_heading_fragment("The remainder term satisfies"));
|
||||
assert!(is_heading_fragment("we conclude that the sum equals"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn real_headings_survive() {
|
||||
assert!(!is_heading_fragment("Introduction"));
|
||||
assert!(!is_heading_fragment("Error Analysis"));
|
||||
assert!(!is_heading_fragment("Materials and Methods"));
|
||||
assert!(!is_heading_fragment("Results"));
|
||||
assert!(!is_heading_fragment("3.2 Richardson Extrapolation"));
|
||||
assert!(!is_heading_fragment("Discussion and Conclusions"));
|
||||
// Terminal punctuation means the clause is complete.
|
||||
assert!(!is_heading_fragment("What is a Derivative?"));
|
||||
assert!(!is_heading_fragment("Procedure:"));
|
||||
assert!(!is_heading_fragment("Note that this is important."));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn title_case_headings_ending_in_a_verb_survive() {
|
||||
// "yields" is also a plural noun; these are real section titles.
|
||||
assert!(!is_heading_fragment("Bond Yields"));
|
||||
assert!(!is_heading_fragment("Crop Yields"));
|
||||
assert!(!is_heading_fragment("Dividend Yields"));
|
||||
assert!(!is_heading_fragment("Yields"));
|
||||
// A wrapped title-case heading whose first line ends on a listed
|
||||
// verb must survive even if the preprocessor failed to merge it.
|
||||
assert!(!is_heading_fragment("The Theorem Implies"));
|
||||
assert!(!is_heading_fragment("What This Denotes"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn numbered_sentence_case_headings_survive() {
|
||||
// heading.rs consults is_heading_fragment BEFORE applying its
|
||||
// numbered-prefix allowance, so the veto must not pre-empt it.
|
||||
assert!(!is_heading_fragment("1. What the model implies"));
|
||||
assert!(!is_heading_fragment("2.3 How the estimator satisfies"));
|
||||
assert!(!is_heading_fragment("IV) What this denotes"));
|
||||
// Without numbering the same wording is still a stranded clause.
|
||||
assert!(is_heading_fragment("What the model implies"));
|
||||
// A bare leading number or pronoun is prose, not numbering.
|
||||
assert!(is_heading_fragment("3 apples and what that implies"));
|
||||
assert!(is_heading_fragment("I think the model implies"));
|
||||
// Markers heading::parse_numbering rejects must not be exempted
|
||||
// either, or an ordinary list item bypasses the veto: lowercase
|
||||
// roman, alphabetical markers, and over-long tokens.
|
||||
assert!(is_heading_fragment("iv) the estimator satisfies"));
|
||||
assert!(is_heading_fragment("d) the value implies"));
|
||||
// Unsupported character (M is outside the parser's I/V/X/L/C set).
|
||||
assert!(is_heading_fragment("MMMM. the value implies"));
|
||||
// Over-long token: nine valid characters, so this exercises the
|
||||
// 8-character bound rather than the character set.
|
||||
assert!(is_heading_fragment("IIIIIIIII. the value implies"));
|
||||
// Eight is still within the bound and stays exempt.
|
||||
assert!(!is_heading_fragment("IIIIIIII. What this implies"));
|
||||
// Uppercase roman within the parser's grammar is still exempt.
|
||||
assert!(!is_heading_fragment("IV. What this denotes"));
|
||||
assert!(!is_heading_fragment("XII) What this implies"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn wrapped_headings_are_not_fragments() {
|
||||
// A heading that wraps across lines ends on a function word. These
|
||||
// are real headings from IRS Publication 17 and must survive.
|
||||
assert!(!is_heading_fragment("Casualty and"));
|
||||
assert!(!is_heading_fragment("Rule 10. You Must Be at"));
|
||||
assert!(!is_heading_fragment("Higher Standard Deduction for"));
|
||||
assert!(!is_heading_fragment("Qualifying Child of"));
|
||||
assert!(!is_heading_fragment("When Can I Withdraw or"));
|
||||
// Copulas and auxiliaries also end real wrapped headings.
|
||||
assert!(!is_heading_fragment("Rule 15. Your AGI Must Be"));
|
||||
assert!(!is_heading_fragment("What Medical Expenses Are"));
|
||||
assert!(!is_heading_fragment("Rule 13. You Must Have"));
|
||||
assert!(!is_heading_fragment("When Can a Roth IRA Be"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dangling_check_is_case_insensitive() {
|
||||
// All-caps is not sentence case, so the veto must not fire there.
|
||||
assert!(!is_heading_fragment("THE REMAINDER EQUALS"));
|
||||
}
|
||||
}
|
||||
|
||||
/// Compute the Y-gap threshold for paragraph break detection.
|
||||
///
|
||||
/// Instead of using a fixed multiple of base_size (which fails for double-spaced
|
||||
|
||||
@@ -127,9 +127,7 @@ fn visual_style(line: &TextLine) -> Option<VisualStyle> {
|
||||
})
|
||||
}
|
||||
|
||||
/// Shared with `analysis::starts_with_numbering_prefix` so the veto
|
||||
/// exemption and the heading parser agree on what a roman numeral is.
|
||||
pub(super) fn roman_value(token: &str) -> Option<u32> {
|
||||
fn roman_value(token: &str) -> Option<u32> {
|
||||
if token.is_empty() || token.len() > 8 {
|
||||
return None;
|
||||
}
|
||||
|
||||
+12
-85
@@ -69,7 +69,7 @@ fn is_chart_adjacent_label(item: &TextItem, region: (f32, f32, f32, f32)) -> boo
|
||||
|| (mostly_inside_chart_width && close_to_chart_edge && category_sized))
|
||||
}
|
||||
|
||||
pub(crate) fn item_is_in_chart_region(item: &TextItem, regions: &[(f32, f32, f32, f32)]) -> bool {
|
||||
fn item_is_in_chart_region(item: &TextItem, regions: &[(f32, f32, f32, f32)]) -> bool {
|
||||
regions.iter().any(|&(x0, y0, x1, y1)| {
|
||||
let cx = item.x + item.width / 2.0;
|
||||
let within_padded_x = cx >= x0 - CHART_REGION_PAD && cx <= x1 + CHART_REGION_PAD;
|
||||
@@ -92,72 +92,6 @@ fn items_outside_chart_regions(
|
||||
.collect()
|
||||
}
|
||||
|
||||
pub(crate) fn merge_chart_regions(
|
||||
regions: impl IntoIterator<Item = (f32, f32, f32, f32)>,
|
||||
) -> Vec<(f32, f32, f32, f32)> {
|
||||
const MERGE_TOLERANCE: f32 = 3.0;
|
||||
|
||||
let mut merged: Vec<(f32, f32, f32, f32)> = Vec::new();
|
||||
for (x0, y0, x1, y1) in regions {
|
||||
let mut current = (x0.min(x1), y0.min(y1), x0.max(x1), y0.max(y1));
|
||||
let mut index = 0;
|
||||
while index < merged.len() {
|
||||
let candidate = merged[index];
|
||||
let overlaps = current.2 + MERGE_TOLERANCE >= candidate.0
|
||||
&& candidate.2 + MERGE_TOLERANCE >= current.0
|
||||
&& current.3 + MERGE_TOLERANCE >= candidate.1
|
||||
&& candidate.3 + MERGE_TOLERANCE >= current.1;
|
||||
if overlaps {
|
||||
current = (
|
||||
current.0.min(candidate.0),
|
||||
current.1.min(candidate.1),
|
||||
current.2.max(candidate.2),
|
||||
current.3.max(candidate.3),
|
||||
);
|
||||
merged.swap_remove(index);
|
||||
} else {
|
||||
index += 1;
|
||||
}
|
||||
}
|
||||
merged.push(current);
|
||||
}
|
||||
merged
|
||||
}
|
||||
|
||||
pub(crate) type PageChartRegions = HashMap<u32, Vec<(f32, f32, f32, f32)>>;
|
||||
|
||||
/// Compute the chart masks used by both layout analysis and Markdown output.
|
||||
///
|
||||
/// Keeping the rect-backed and dense-line heuristics behind one entry point
|
||||
/// ensures metadata and extraction cannot drift when either detector changes.
|
||||
pub(crate) fn chart_regions_by_page(
|
||||
items: &[TextItem],
|
||||
rects: &[PdfRect],
|
||||
lines: &[PdfLine],
|
||||
) -> PageChartRegions {
|
||||
let mut page_items: HashMap<u32, Vec<TextItem>> = HashMap::new();
|
||||
for item in items.iter().filter(|item| {
|
||||
matches!(
|
||||
&item.item_type,
|
||||
crate::types::ItemType::Text | crate::types::ItemType::FormField
|
||||
)
|
||||
}) {
|
||||
page_items.entry(item.page).or_default().push(item.clone());
|
||||
}
|
||||
|
||||
page_items
|
||||
.into_iter()
|
||||
.filter_map(|(page, items)| {
|
||||
let rect_regions = crate::tables::detect_chart_regions(&items, rects, page);
|
||||
let line_regions = crate::tables::detect_dense_line_chart_regions(lines, rects, page)
|
||||
.into_iter()
|
||||
.filter(|®ion| chart_region_separates_prose_columns(&items, region));
|
||||
let regions = merge_chart_regions(rect_regions.into_iter().chain(line_regions));
|
||||
(!regions.is_empty()).then_some((page, regions))
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Detect side-by-side table layout by finding a significant X-position gap.
|
||||
///
|
||||
/// Returns X-band boundaries `[(x_min, split_x), (split_x, x_max)]` when a
|
||||
@@ -395,15 +329,6 @@ fn chart_spans_prose_split(region: (f32, f32, f32, f32), split_x: f32) -> bool {
|
||||
split_x - left >= MIN_CHART_WIDTH_PER_SIDE && right - split_x >= MIN_CHART_WIDTH_PER_SIDE
|
||||
}
|
||||
|
||||
pub(crate) fn chart_region_separates_prose_columns(
|
||||
items: &[TextItem],
|
||||
region: (f32, f32, f32, f32),
|
||||
) -> bool {
|
||||
let outside = items_outside_chart_regions(items, &[region]);
|
||||
chart_page_prose_column_split(&outside)
|
||||
.is_some_and(|split_x| chart_spans_prose_split(region, split_x))
|
||||
}
|
||||
|
||||
/// True when adjacent physical rows form an unterminated, lowercase prose
|
||||
/// continuation in the same projected column.
|
||||
fn is_cross_row_prose_continuation(previous: &str, current: &str) -> bool {
|
||||
@@ -1079,7 +1004,6 @@ pub fn to_markdown_from_items_with_rects_and_page_count(
|
||||
page_count: document_page_count,
|
||||
prefiltered_page_number_pages: None,
|
||||
prefiltered_page_number_mask: None,
|
||||
precomputed_chart_regions: None,
|
||||
},
|
||||
)
|
||||
}
|
||||
@@ -1097,9 +1021,6 @@ pub(crate) struct MarkdownDocumentContext<'a> {
|
||||
/// Table detection consumes the original items; the mask is applied only
|
||||
/// after table claims have been established.
|
||||
pub(crate) prefiltered_page_number_mask: Option<&'a [bool]>,
|
||||
/// Optional chart masks shared with layout analysis so the geometry is
|
||||
/// detected once and interpreted identically by both pipelines.
|
||||
pub(crate) precomputed_chart_regions: Option<&'a PageChartRegions>,
|
||||
}
|
||||
|
||||
/// Convert positioned text items to markdown, using rectangles and line segments for table detection.
|
||||
@@ -1126,7 +1047,6 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
|
||||
page_count: document_page_count,
|
||||
prefiltered_page_number_pages,
|
||||
prefiltered_page_number_mask,
|
||||
precomputed_chart_regions,
|
||||
} = context;
|
||||
|
||||
if items.is_empty() {
|
||||
@@ -1199,9 +1119,17 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
|
||||
|
||||
// Chart regions per page: their text must not steer column detection
|
||||
// during line grouping (it fills the gutter and fuses two-column lines).
|
||||
let page_chart_map = precomputed_chart_regions
|
||||
.cloned()
|
||||
.unwrap_or_else(|| chart_regions_by_page(&text_items, rects, pdf_lines));
|
||||
let mut page_chart_map: HashMap<u32, Vec<(f32, f32, f32, f32)>> = HashMap::new();
|
||||
for &page in page_groups.keys() {
|
||||
let page_items_ref: Vec<TextItem> = page_groups[&page]
|
||||
.iter()
|
||||
.map(|(_, item)| (*item).clone())
|
||||
.collect();
|
||||
let regions = crate::tables::detect_chart_regions(&page_items_ref, rects, page);
|
||||
if !regions.is_empty() {
|
||||
page_chart_map.insert(page, regions);
|
||||
}
|
||||
}
|
||||
|
||||
let mut pages: Vec<u32> = page_groups.keys().copied().collect();
|
||||
pages.sort();
|
||||
@@ -2123,7 +2051,6 @@ mod tests {
|
||||
page_count: 1,
|
||||
prefiltered_page_number_pages: Some(&removed_pages),
|
||||
prefiltered_page_number_mask: Some(&removal_mask),
|
||||
precomputed_chart_regions: None,
|
||||
},
|
||||
);
|
||||
|
||||
|
||||
@@ -464,7 +464,6 @@ mod tests {
|
||||
assert!(!is_page_number_line("Hello World"));
|
||||
assert!(!is_page_number_line("Chapter 1"));
|
||||
assert!(!is_page_number_line("Total: 500"));
|
||||
assert!(!is_page_number_line("PAGE0-PARA2-END-MARKER-0"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -508,14 +507,6 @@ mod tests {
|
||||
assert!(result.contains("End"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_remove_page_numbers_preserves_page_prefixed_content() {
|
||||
let input = "PAGE0-PARA2-START substantive report text PAGE0-PARA2-END-MARKER-0";
|
||||
let result = remove_page_numbers(input);
|
||||
|
||||
assert_eq!(result, input);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_remove_page_numbers_multiple_patterns() {
|
||||
let input = "\n1\n\nContent\n\n2\n\n---\nMore\n\n3\n";
|
||||
|
||||
@@ -150,61 +150,6 @@ pub(crate) fn merge_heading_lines(
|
||||
/// Merge drop caps with the appropriate line.
|
||||
/// A drop cap is a single large letter at the start of a paragraph.
|
||||
/// Due to PDF coordinate sorting, the drop cap may appear AFTER the line it belongs to.
|
||||
/// True when the text ends a sentence, as opposed to merely ending in a
|
||||
/// period. An abbreviation or list marker ("e.g.", "Fig.", "Mr.", "1.")
|
||||
/// closes with a period mid-sentence, so treating those as paragraph
|
||||
/// boundaries would let a drop cap be prepended to a continuation.
|
||||
fn ends_sentence(text: &str) -> bool {
|
||||
let t = text.trim_end();
|
||||
if t.ends_with(['!', '?']) {
|
||||
return true;
|
||||
}
|
||||
let Some(stripped) = t.strip_suffix('.') else {
|
||||
return false;
|
||||
};
|
||||
let last = stripped.split_whitespace().next_back().unwrap_or("");
|
||||
if last.is_empty() {
|
||||
return false;
|
||||
}
|
||||
|
||||
// Abbreviations carry an internal period between very short segments
|
||||
// ("e.g.", "i.e.", "U.S."). Domains and decimals have the same shape but
|
||||
// longer or numeric segments ("example.com.", "3.14."), and those end
|
||||
// sentences perfectly well, so require every segment to be short and
|
||||
// alphabetic before reading the internal period as an abbreviation.
|
||||
if last.contains('.')
|
||||
&& last
|
||||
.split('.')
|
||||
.filter(|seg| !seg.is_empty())
|
||||
// Characters, not bytes: a two-letter non-ASCII abbreviation
|
||||
// ("т.е.", "ú.d.") measures four or more bytes and would
|
||||
// otherwise be read as a completed sentence.
|
||||
.all(|seg| seg.chars().count() <= 2 && seg.chars().all(char::is_alphabetic))
|
||||
{
|
||||
return false;
|
||||
}
|
||||
|
||||
// Enumerators stand alone on their line ("1.", "ii.", "IV."). A number
|
||||
// or numeral in the tail of a sentence does not — "published in 2020.",
|
||||
// "He scored 5." and "after World War II." all end sentences, and
|
||||
// treating them as markers would block a legitimate drop-cap merge.
|
||||
if stripped.split_whitespace().count() == 1 {
|
||||
let is_numeric = last.chars().all(|c| c.is_ascii_digit());
|
||||
let is_roman = last
|
||||
.chars()
|
||||
.all(|c| matches!(c.to_ascii_uppercase(), 'I' | 'V' | 'X' | 'L' | 'C'));
|
||||
if is_numeric || is_roman {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
const ABBREVIATIONS: &[&str] = &[
|
||||
"Fig", "No", "Mr", "Mrs", "Ms", "Dr", "St", "vs", "etc", "al", "Ed", "Eq", "Ch", "pp",
|
||||
"Vol", "cf", "Prof", "Inc", "Ltd", "Jr", "Sr",
|
||||
];
|
||||
!ABBREVIATIONS.iter().any(|a| a.eq_ignore_ascii_case(last))
|
||||
}
|
||||
|
||||
pub(crate) fn merge_drop_caps(lines: Vec<TextLine>, base_size: f32) -> Vec<TextLine> {
|
||||
let mut result: Vec<TextLine> = Vec::with_capacity(lines.len());
|
||||
|
||||
@@ -224,169 +169,6 @@ pub(crate) fn merge_drop_caps(lines: Vec<TextLine>, base_size: f32) -> Vec<TextL
|
||||
.map(|c| c.is_uppercase())
|
||||
.unwrap_or(false);
|
||||
|
||||
// Embedded drop cap: a two-line cap's baseline aligns with the
|
||||
// paragraph's SECOND line, so Y-grouping puts the glyph at the start
|
||||
// of that line rather than on a line of its own. Left there it
|
||||
// surfaces mid-sentence once the paragraph is joined — Shannon's
|
||||
// "A Mathematical Theory of Communication" reads "...which exchange
|
||||
// T bandwidth for signal-to-noise ratio...". Detect it, prepend the
|
||||
// character to the paragraph's first line, and drop it from this one.
|
||||
//
|
||||
// The size gate is 1.8x rather than 2.5x because bitmap (Type3) caps
|
||||
// report their glyph bbox rather than the em box, so a two-line cap
|
||||
// can measure as little as ~1.9x the body size.
|
||||
if line.items.len() > 1 {
|
||||
let first = &line.items[0];
|
||||
// The remainder must be a substantive body run: a lone label or
|
||||
// math fragment beside a large glyph is not a drop-cap paragraph.
|
||||
let rest_letters: usize = line.items[1..]
|
||||
.iter()
|
||||
.map(|i| i.text.chars().filter(|c| c.is_alphabetic()).count())
|
||||
.sum();
|
||||
let is_embedded_cap = first.font_size >= base_size * 1.8
|
||||
&& first.text.trim().chars().count() == 1
|
||||
&& first
|
||||
.text
|
||||
.trim()
|
||||
.chars()
|
||||
.next()
|
||||
.is_some_and(char::is_uppercase)
|
||||
&& line.items[1..]
|
||||
.iter()
|
||||
.all(|i| i.font_size < base_size * 1.5)
|
||||
&& line.items[1..].iter().any(|i| i.x > first.x)
|
||||
&& rest_letters >= 8;
|
||||
if is_embedded_cap {
|
||||
let drop_char = first.text.trim().chars().next().unwrap();
|
||||
let cap_x = first.x;
|
||||
let line_y = line.y;
|
||||
// Text on the cap's own line, pushed right to clear the glyph.
|
||||
let rest_x = line.items[1].x;
|
||||
|
||||
// Walk up the run of lines the cap has indented. A drop cap
|
||||
// pushes every line it covers to the right of the glyph, so
|
||||
// the paragraph's first line is the TOPMOST line sharing that
|
||||
// indent — however many lines the cap spans. Using the indent
|
||||
// rather than the cap's font size is what makes this work for
|
||||
// three- and four-line initials as well as two-line ones;
|
||||
// deriving a line count from the em size does not survive
|
||||
// contact with real documents, where 36-47pt initials sit
|
||||
// over 11-14pt leading.
|
||||
//
|
||||
// A cap in a different column has no such run (its neighbours
|
||||
// sit at an unrelated x), so it is left alone — which is
|
||||
// correct when the cap's own line already carries the rest of
|
||||
// the word.
|
||||
const INDENT_TOLERANCE: f32 = 2.0;
|
||||
const MAX_CAP_LINES: usize = 8;
|
||||
let max_step = base_size * 2.5;
|
||||
let mut target_idx = result.len();
|
||||
let mut expected_y = line_y;
|
||||
while target_idx > 0 && result.len() - target_idx < MAX_CAP_LINES {
|
||||
let cand = &result[target_idx - 1];
|
||||
let step = cand.y - expected_y;
|
||||
let shares_indent = cand
|
||||
.items
|
||||
.first()
|
||||
.is_some_and(|i| (i.x - rest_x).abs() <= INDENT_TOLERANCE);
|
||||
if cand.page != line.page || step <= 0.0 || step > max_step || !shares_indent {
|
||||
break;
|
||||
}
|
||||
expected_y = cand.y;
|
||||
target_idx -= 1;
|
||||
}
|
||||
|
||||
// The topmost line of the run is the paragraph's first line.
|
||||
// The line above THAT tells us whether it starts a paragraph.
|
||||
let before_target = target_idx
|
||||
.checked_sub(1)
|
||||
.and_then(|i| result.get(i))
|
||||
.filter(|l| l.page == line.page)
|
||||
.map(|l| (l.text().trim_end().to_string(), l.y));
|
||||
// Leading within the run: the step from the target down to the
|
||||
// next line of the paragraph, which is the cap's own line when
|
||||
// the run is a single line.
|
||||
let run_step = result
|
||||
.get(target_idx)
|
||||
.map(|t| {
|
||||
let below_y = result.get(target_idx + 1).map_or(line_y, |b| b.y);
|
||||
t.y - below_y
|
||||
})
|
||||
.unwrap_or(0.0);
|
||||
let step_for_gap = if run_step > 0.0 {
|
||||
run_step
|
||||
} else {
|
||||
base_size * 1.2
|
||||
};
|
||||
|
||||
let target = (target_idx < result.len())
|
||||
.then(|| &mut result[target_idx])
|
||||
.filter(|prev| {
|
||||
let prev_text = prev.text();
|
||||
let prev_trimmed = prev_text.trim();
|
||||
// A hyphen on the line above means the target resumes
|
||||
// a split word, so it continues a paragraph rather
|
||||
// than starting one (polkuja_ylakoulu: "ylakou-" +
|
||||
// "lulaisten").
|
||||
//
|
||||
// Case cannot serve as a continuation signal here: the
|
||||
// target legitimately starts lowercase, because the
|
||||
// cap removes the word's first letter and leaves
|
||||
// "ver the course..." for "Over".
|
||||
let continues_previous = before_target
|
||||
.as_ref()
|
||||
.is_some_and(|(b, _)| b.ends_with('-'));
|
||||
// The target must START a paragraph: extra leading
|
||||
// above it, a completed sentence on the line above, or
|
||||
// nothing above it at all.
|
||||
let starts_paragraph = match before_target.as_ref() {
|
||||
None => true,
|
||||
Some((text, y)) => {
|
||||
y - prev.y > step_for_gap * 1.15 || ends_sentence(text)
|
||||
}
|
||||
};
|
||||
!continues_previous
|
||||
&& starts_paragraph
|
||||
&& prev.page == line.page
|
||||
&& prev.y > line_y
|
||||
// Indented past the cap glyph, not merely to its
|
||||
// right by an arbitrary amount.
|
||||
&& prev
|
||||
.items
|
||||
.first()
|
||||
.is_some_and(|i| i.x > cap_x && i.x - cap_x <= first.font_size * 2.0)
|
||||
// Body text, so headings, labels and table
|
||||
// fragments are never rewritten.
|
||||
&& prev_trimmed
|
||||
.chars()
|
||||
.next()
|
||||
.is_some_and(char::is_alphabetic)
|
||||
&& prev_trimmed.chars().filter(|c| c.is_alphabetic()).count() >= 8
|
||||
});
|
||||
if let Some(prev_line) = target {
|
||||
if let Some(first_item) = prev_line.items.first_mut() {
|
||||
// A mid-word cap ("T" + "HE recent") joins directly.
|
||||
// Leading whitespace only marks a word boundary when
|
||||
// the cap is itself a single-letter word, since the
|
||||
// paragraph's indent can also arrive as whitespace.
|
||||
const SINGLE_LETTER_WORDS: &[char] = &['A', 'I', 'O', 'U', 'Y', 'E'];
|
||||
let had_leading_ws = first_item.text.starts_with(char::is_whitespace)
|
||||
&& SINGLE_LETTER_WORDS.contains(&drop_char);
|
||||
let rest = first_item.text.trim_start().to_string();
|
||||
first_item.text = if had_leading_ws {
|
||||
format!("{} {}", drop_char, rest)
|
||||
} else {
|
||||
format!("{}{}", drop_char, rest)
|
||||
};
|
||||
}
|
||||
let mut line = line.clone();
|
||||
line.items.remove(0);
|
||||
result.push(line);
|
||||
continue;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if is_drop_cap {
|
||||
let drop_char = trimmed.chars().next().unwrap();
|
||||
|
||||
@@ -816,314 +598,6 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
fn make_item_at(text: &str, font_size: f32, x: f32) -> TextItem {
|
||||
let mut item = make_item(text, font_size, None);
|
||||
item.x = x;
|
||||
item.width = text.len() as f32 * font_size * 0.5;
|
||||
item
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn embedded_drop_cap_moves_to_paragraph_start() {
|
||||
// A two-line cap baseline-aligns with the paragraph's SECOND line,
|
||||
// so it lands as that line's first item (Shannon entropy.pdf p.1).
|
||||
let first_line = TextLine {
|
||||
items: vec![make_item_at(
|
||||
"HE recent development which exchange",
|
||||
10.0,
|
||||
90.0,
|
||||
)],
|
||||
// 16pt baseline step under a 25pt cap: a genuine two-line cap.
|
||||
y: 716.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let second_line = TextLine {
|
||||
items: vec![
|
||||
make_item_at("T", 25.0, 72.0),
|
||||
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
|
||||
],
|
||||
y: 700.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let result = merge_drop_caps(vec![first_line, second_line], 10.0);
|
||||
assert_eq!(result.len(), 2);
|
||||
assert!(
|
||||
result[0].text().starts_with("THE recent"),
|
||||
"cap should prepend to the paragraph start: {}",
|
||||
result[0].text()
|
||||
);
|
||||
assert!(
|
||||
result[1].text().starts_with("bandwidth"),
|
||||
"cap must be removed from the second line: {}",
|
||||
result[1].text()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn embedded_drop_cap_walks_a_multi_line_initial_to_the_paragraph_start() {
|
||||
// A 47pt initial over 13pt leading covers four lines, so the
|
||||
// paragraph's first line is three lines above the cap rather than
|
||||
// immediately above it (polkuja_ylakoulu). The indented run, not the
|
||||
// cap's em size, is what locates it.
|
||||
let mut lines = vec![TextLine {
|
||||
items: vec![make_item_at("Previous paragraph ends here.", 10.0, 72.0)],
|
||||
y: 766.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
}];
|
||||
for (i, text) in [
|
||||
"rilaiset mediasisallot ovat tarkea osa",
|
||||
"useimpien ylakoululaisten elamaa ja",
|
||||
"muuta tekstia jatkuu tassa viela",
|
||||
]
|
||||
.iter()
|
||||
.enumerate()
|
||||
{
|
||||
lines.push(TextLine {
|
||||
items: vec![make_item_at(text, 10.0, 90.0)],
|
||||
y: 753.0 - 13.0 * i as f32,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
});
|
||||
}
|
||||
lines.push(TextLine {
|
||||
items: vec![
|
||||
make_item_at("E", 47.0, 72.0),
|
||||
make_item_at("loppuosa tekstista tassa", 10.0, 90.0),
|
||||
],
|
||||
y: 714.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
});
|
||||
|
||||
let result = merge_drop_caps(lines, 10.0);
|
||||
assert!(
|
||||
result[1].text().starts_with("Erilaiset"),
|
||||
"cap belongs on the topmost line of the indented run: {}",
|
||||
result[1].text()
|
||||
);
|
||||
assert!(
|
||||
result[2].text().starts_with("useimpien"),
|
||||
"intervening run lines must be untouched: {}",
|
||||
result[2].text()
|
||||
);
|
||||
assert!(
|
||||
result[4].text().starts_with("loppuosa"),
|
||||
"cap must be removed from its own line: {}",
|
||||
result[4].text()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn embedded_drop_cap_ignores_non_paragraph_neighbours() {
|
||||
// Same geometry, but the preceding line is a short label rather than
|
||||
// body text, so it must not be rewritten.
|
||||
let label = TextLine {
|
||||
items: vec![make_item_at("Fig. 2", 10.0, 90.0)],
|
||||
y: 716.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let second_line = TextLine {
|
||||
items: vec![
|
||||
make_item_at("T", 25.0, 72.0),
|
||||
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
|
||||
],
|
||||
y: 700.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let result = merge_drop_caps(vec![label, second_line], 10.0);
|
||||
assert_eq!(result[0].text().trim(), "Fig. 2");
|
||||
assert!(
|
||||
result[1].text().starts_with('T'),
|
||||
"cap stays put: {}",
|
||||
result[1].text()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn embedded_drop_cap_keeps_a_space_for_standalone_word_caps() {
|
||||
// Leading whitespace on the paragraph's first item marks the cap as
|
||||
// a word of its own rather than the first letter of one.
|
||||
let mut lead = make_item_at("long time ago in a galaxy far away", 10.0, 90.0);
|
||||
lead.text = " long time ago in a galaxy far away".to_string();
|
||||
let first_line = TextLine {
|
||||
items: vec![lead],
|
||||
y: 716.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let second_line = TextLine {
|
||||
items: vec![
|
||||
make_item_at("A", 25.0, 72.0),
|
||||
make_item_at("continued here with more body text", 10.0, 90.0),
|
||||
],
|
||||
y: 700.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let result = merge_drop_caps(vec![first_line, second_line], 10.0);
|
||||
assert!(
|
||||
result[0].text().starts_with("A long time ago"),
|
||||
"standalone-word cap keeps one space: {}",
|
||||
result[0].text()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn embedded_drop_cap_skips_hyphenation_continuation_targets() {
|
||||
// The line above the RUN ends on a hyphen, so the run's topmost line
|
||||
// resumes a split word rather than starting a paragraph. It sits at
|
||||
// the paragraph margin (x=72), outside the cap's indent, so it is not
|
||||
// part of the run itself.
|
||||
let split_word = TextLine {
|
||||
items: vec![make_item_at(
|
||||
"mediasisallot ovat osa useimpien ylakou-",
|
||||
10.0,
|
||||
72.0,
|
||||
)],
|
||||
y: 728.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let run_top = TextLine {
|
||||
items: vec![make_item_at(
|
||||
"lulaisten elamaa ja muuta tekstia",
|
||||
10.0,
|
||||
90.0,
|
||||
)],
|
||||
y: 714.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let cap_line = TextLine {
|
||||
items: vec![
|
||||
make_item_at("E", 25.0, 72.0),
|
||||
make_item_at("jatkuu tassa lisaa leipatekstia", 10.0, 90.0),
|
||||
],
|
||||
y: 700.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let result = merge_drop_caps(vec![split_word, run_top, cap_line], 10.0);
|
||||
assert!(
|
||||
result[1].text().starts_with("lulaisten"),
|
||||
"a run resuming a split word must not receive the cap: {}",
|
||||
result[1].text()
|
||||
);
|
||||
assert!(
|
||||
result[2].text().starts_with('E'),
|
||||
"cap stays put when no valid target exists: {}",
|
||||
result[2].text()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn embedded_drop_cap_indent_is_not_a_word_boundary() {
|
||||
// The paragraph's first line is indented to clear the cap, and that
|
||||
// indent can arrive as leading whitespace. A mid-word cap must still
|
||||
// join directly — "T HE recent" would be the defect this fixes.
|
||||
let mut lead = make_item_at("HE recent development and more body text", 10.0, 90.0);
|
||||
lead.text = " HE recent development and more body text".to_string();
|
||||
let first_line = TextLine {
|
||||
items: vec![lead],
|
||||
y: 716.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let cap_line = TextLine {
|
||||
items: vec![
|
||||
make_item_at("T", 25.0, 72.0),
|
||||
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
|
||||
],
|
||||
y: 700.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let result = merge_drop_caps(vec![first_line, cap_line], 10.0);
|
||||
assert!(
|
||||
result[0].text().starts_with("THE recent"),
|
||||
"indent must not be read as a word boundary: {}",
|
||||
result[0].text()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ends_sentence_rejects_abbreviations_and_markers() {
|
||||
use super::ends_sentence;
|
||||
assert!(ends_sentence("This completes the thought."));
|
||||
assert!(ends_sentence("Is that so?"));
|
||||
assert!(ends_sentence("Stop!"));
|
||||
// Periods that do not end a sentence.
|
||||
assert!(!ends_sentence("as shown in Fig."));
|
||||
assert!(!ends_sentence("see e.g."));
|
||||
// Non-ASCII abbreviations. The two-CHARACTER segment is the case
|
||||
// that distinguishes a character count from a byte count: "пр" is
|
||||
// 2 chars but 4 bytes, so a byte-based bound would reject it and
|
||||
// read the line as a completed sentence.
|
||||
assert!(!ends_sentence("и т.пр."));
|
||||
assert!(!ends_sentence("см. т.е."));
|
||||
assert!(!ends_sentence("napr. ú.d."));
|
||||
assert!(!ends_sentence("reviewed by Dr."));
|
||||
// Standalone enumerators, any case.
|
||||
assert!(!ends_sentence("1."));
|
||||
assert!(!ends_sentence("IV."));
|
||||
assert!(!ends_sentence("ii."));
|
||||
assert!(!ends_sentence("xii."));
|
||||
// Numbers and numerals that genuinely end a sentence must count,
|
||||
// or a legitimate drop-cap merge is blocked.
|
||||
assert!(ends_sentence("The paper was published in 2020."));
|
||||
assert!(ends_sentence("He scored 5."));
|
||||
assert!(ends_sentence("after World War II."));
|
||||
assert!(ends_sentence("the constant equals 3.14."));
|
||||
assert!(ends_sentence("documented at example.com."));
|
||||
assert!(!ends_sentence("a trailing clause with no period"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn embedded_drop_cap_allows_first_paragraph_on_a_new_page() {
|
||||
// The line two back is on the previous page, so its y is unrelated
|
||||
// and must not be used as leading evidence.
|
||||
let prev_page_tail = TextLine {
|
||||
items: vec![make_item_at(
|
||||
"tail of the previous page body text",
|
||||
10.0,
|
||||
90.0,
|
||||
)],
|
||||
y: 90.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let first_line = TextLine {
|
||||
items: vec![make_item_at(
|
||||
"HE recent development which exchange",
|
||||
10.0,
|
||||
90.0,
|
||||
)],
|
||||
y: 716.0,
|
||||
page: 2,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let cap_line = TextLine {
|
||||
items: vec![
|
||||
make_item_at("T", 25.0, 72.0),
|
||||
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
|
||||
],
|
||||
y: 700.0,
|
||||
page: 2,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let result = merge_drop_caps(vec![prev_page_tail, first_line, cap_line], 10.0);
|
||||
assert!(
|
||||
result[1].text().starts_with("THE recent"),
|
||||
"a page break must not suppress the merge: {}",
|
||||
result[1].text()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_merge_struct_tree_headings() {
|
||||
// Two consecutive lines tagged as H2 via struct tree, same font size as body
|
||||
|
||||
+10
-308
@@ -435,72 +435,6 @@ fn revised_table_cell_indices(
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Index of candidate "body" items (larger-font attachment targets) sorted by
|
||||
/// Y, so script-attachment checks scan a narrow Y window instead of the whole
|
||||
/// page per candidate.
|
||||
struct ScriptBodyIndex<'a> {
|
||||
/// (y, item), sorted ascending by y
|
||||
by_y: Vec<(f32, &'a TextItem)>,
|
||||
/// widest vertical attachment window any body item can produce
|
||||
max_window: f32,
|
||||
}
|
||||
|
||||
impl<'a> ScriptBodyIndex<'a> {
|
||||
fn new(items: &'a [TextItem]) -> Self {
|
||||
// Smallest table-candidate font is 6pt, so any possible attachment
|
||||
// target is at least 6 x 1.2 pt.
|
||||
let mut by_y: Vec<(f32, &TextItem)> = items
|
||||
.iter()
|
||||
.filter(|i| i.font_size >= 6.0 * 1.2)
|
||||
.map(|i| (i.y, i))
|
||||
.collect();
|
||||
by_y.sort_by(|a, b| a.0.total_cmp(&b.0));
|
||||
let max_window = by_y
|
||||
.iter()
|
||||
.map(|(_, i)| i.font_size * 0.8)
|
||||
.fold(0.0f32, f32::max);
|
||||
Self { by_y, max_window }
|
||||
}
|
||||
|
||||
/// True when a small-font item is horizontally attached to a larger-font
|
||||
/// item at a script baseline offset — a sub/superscript in running text
|
||||
/// or math (equation subscripts, footnote markers). Script attachments
|
||||
/// are not table cells; without this filter, display equations with
|
||||
/// sub/superscripts form phantom small-font table regions (e.g. TeX
|
||||
/// papers where log subscripts cluster with footnote lines into a fake
|
||||
/// 3-column table). A genuine baseline offset is required so same-line
|
||||
/// table neighbours (a small cell beside a larger label cell) are never
|
||||
/// classified as scripts.
|
||||
///
|
||||
/// `min_anchor_size` additionally constrains what counts as an
|
||||
/// attachment target: the small-font pass accepts any sufficiently
|
||||
/// larger item (0.0), while the body-font pass requires a heading-sized
|
||||
/// anchor so a body-size table cell beside a slightly larger label with
|
||||
/// baseline jitter is never treated as a script.
|
||||
fn is_script_attachment(&self, small: &TextItem, min_anchor_size: f32) -> bool {
|
||||
let attach_gap = small.font_size.max(4.0) * 0.6;
|
||||
let lo = self
|
||||
.by_y
|
||||
.partition_point(|(y, _)| *y < small.y - self.max_window);
|
||||
self.by_y[lo..]
|
||||
.iter()
|
||||
.take_while(|(y, _)| *y <= small.y + self.max_window)
|
||||
.any(|(_, body)| {
|
||||
let dy = (small.y - body.y).abs();
|
||||
body.font_size >= small.font_size * 1.2
|
||||
&& body.font_size >= min_anchor_size
|
||||
&& dy > body.font_size * 0.05
|
||||
&& dy <= body.font_size * 0.8
|
||||
&& {
|
||||
let gap_after_body = small.x - (body.x + body.width);
|
||||
let gap_before_body = body.x - (small.x + small.width);
|
||||
(-attach_gap..=attach_gap).contains(&gap_after_body)
|
||||
|| (-attach_gap..=attach_gap).contains(&gap_before_body)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// Detect tables in a set of text items from a single page
|
||||
pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bool) -> Vec<Table> {
|
||||
detect_tables_with_page_width(items, base_font_size, skip_body_font, content_width(items))
|
||||
@@ -549,27 +483,6 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
// === Pass 1: Small-font tables (existing behavior) ===
|
||||
let table_font_threshold = base_font_size * 0.90;
|
||||
|
||||
// Mark sub/superscript attachments once per pass. They stay candidates —
|
||||
// the masks only remove them from region qualification and column/row
|
||||
// geometry.
|
||||
//
|
||||
// The two passes need different anchor thresholds. In the small-font pass
|
||||
// any sufficiently larger neighbour is a plausible base for a script. In
|
||||
// the body-font pass the candidates are themselves body-sized
|
||||
// (0.85..1.05x), so a merely "slightly larger" neighbour is usually a bold
|
||||
// label or an adjacent column header, not the base of a superscript —
|
||||
// treating it as one would strip real cells out of the geometry and lose
|
||||
// the table. Requiring a heading-sized anchor (>= 1.15x base) keeps the
|
||||
// body pass to genuine scripts hanging off headings.
|
||||
let script_index = ScriptBodyIndex::new(items);
|
||||
let script_flags: Vec<bool> = items
|
||||
.iter()
|
||||
.map(|item| script_index.is_script_attachment(item, 0.0))
|
||||
.collect();
|
||||
let body_script_flags: Vec<bool> = items
|
||||
.iter()
|
||||
.map(|item| script_index.is_script_attachment(item, base_font_size * 1.15))
|
||||
.collect();
|
||||
let table_candidates: Vec<(usize, &TextItem)> = items
|
||||
.iter()
|
||||
.enumerate()
|
||||
@@ -581,14 +494,7 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
.collect();
|
||||
|
||||
if table_candidates.len() >= 6 {
|
||||
// Qualify regions from non-script items: a cluster of sub/superscripts
|
||||
// must not, on its own, mark out a table region.
|
||||
let region_evidence: Vec<(usize, &TextItem)> = table_candidates
|
||||
.iter()
|
||||
.filter(|(idx, _)| !script_flags[*idx])
|
||||
.cloned()
|
||||
.collect();
|
||||
let regions = find_table_regions(®ion_evidence);
|
||||
let regions = find_table_regions(&table_candidates);
|
||||
|
||||
for (y_min, y_max) in regions {
|
||||
let region_items: Vec<(usize, &TextItem)> = table_candidates
|
||||
@@ -602,9 +508,7 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
}
|
||||
|
||||
if let Some(mut table) =
|
||||
detect_table_in_region(®ion_items, TableDetectionMode::SmallFont, &|i| {
|
||||
script_flags[i]
|
||||
})
|
||||
detect_table_in_region(®ion_items, TableDetectionMode::SmallFont)
|
||||
{
|
||||
// Try to recover body-font header row above the small-font table
|
||||
recover_header_row(&mut table, items, table_font_threshold);
|
||||
@@ -649,20 +553,8 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
body_font_low,
|
||||
body_font_high,
|
||||
);
|
||||
// Scripts are NOT filtered out of the candidate set here, mirroring
|
||||
// the small-font pass: they must stay eligible for cell assignment so
|
||||
// a sub/superscript that belongs inside a table cell keeps its text.
|
||||
// The heading-anchored `body_script_flags` mask removes them from
|
||||
// geometry only.
|
||||
if body_candidates.len() >= 6 {
|
||||
// Same reasoning as the small-font pass: scripts do not qualify
|
||||
// regions, but remain available for cell assignment within one.
|
||||
let region_evidence: Vec<(usize, &TextItem)> = body_candidates
|
||||
.iter()
|
||||
.filter(|(idx, _)| !body_script_flags[*idx])
|
||||
.cloned()
|
||||
.collect();
|
||||
let regions = find_table_regions_strict(®ion_evidence);
|
||||
let regions = find_table_regions_strict(&body_candidates);
|
||||
log::debug!("body-font: {} strict regions found", regions.len());
|
||||
|
||||
for (y_min, y_max, _x_min, _x_max) in ®ions {
|
||||
@@ -688,9 +580,7 @@ pub(crate) fn detect_tables_with_page_width(
|
||||
}
|
||||
|
||||
if let Some(table) =
|
||||
detect_table_in_region(®ion_items, TableDetectionMode::BodyFont, &|i| {
|
||||
body_script_flags[i]
|
||||
})
|
||||
detect_table_in_region(®ion_items, TableDetectionMode::BodyFont)
|
||||
{
|
||||
tables.push(table);
|
||||
}
|
||||
@@ -918,30 +808,10 @@ fn find_table_regions_strict(items: &[(usize, &TextItem)]) -> Vec<(f32, f32, f32
|
||||
regions
|
||||
}
|
||||
|
||||
/// Detect a table within a specific region.
|
||||
///
|
||||
/// `is_script` marks items that are sub/superscript attachments. Those are
|
||||
/// excluded from the *geometry* — they must not be able to create a column,
|
||||
/// which is how equation subscript clusters used to fabricate phantom grids —
|
||||
/// but they remain eligible for cell assignment, so legitimate cell content
|
||||
/// (exponents in an engineering-notation table, footnote markers) stays in
|
||||
/// the cell it belongs to instead of leaking out into the reading order.
|
||||
fn detect_table_in_region(
|
||||
items: &[(usize, &TextItem)],
|
||||
mode: TableDetectionMode,
|
||||
is_script: &dyn Fn(usize) -> bool,
|
||||
) -> Option<Table> {
|
||||
// Column geometry from non-script items only.
|
||||
let geometry_items: Vec<(usize, &TextItem)> = items
|
||||
.iter()
|
||||
.filter(|(idx, _)| !is_script(*idx))
|
||||
.cloned()
|
||||
.collect();
|
||||
// A region that is *entirely* scripts has no table structure at all.
|
||||
if geometry_items.is_empty() {
|
||||
return None;
|
||||
}
|
||||
let columns = find_column_boundaries(&geometry_items, mode);
|
||||
/// Detect a table within a specific region
|
||||
fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode) -> Option<Table> {
|
||||
// Find column boundaries
|
||||
let columns = find_column_boundaries(items, mode);
|
||||
let min_cols = 2;
|
||||
if columns.len() < min_cols || columns.len() > 25 {
|
||||
log::debug!(
|
||||
@@ -952,8 +822,8 @@ fn detect_table_in_region(
|
||||
return None;
|
||||
}
|
||||
|
||||
// Find row boundaries (geometry items only, same reasoning)
|
||||
let rows = find_row_boundaries(&geometry_items);
|
||||
// Find row boundaries
|
||||
let rows = find_row_boundaries(items);
|
||||
let min_rows = 2;
|
||||
if rows.len() < min_rows {
|
||||
log::debug!(
|
||||
@@ -972,11 +842,6 @@ fn detect_table_in_region(
|
||||
);
|
||||
|
||||
// Verify this looks like a table: multiple items should align to columns
|
||||
// Validate against ALL items, including scripts. Columns are derived from
|
||||
// non-script geometry so scripts cannot *create* a column, but excluding
|
||||
// them from validation too would let a region manufacture alignment: drop
|
||||
// the awkward items and whatever remains looks like a tidy grid. Block
|
||||
// diagrams did exactly that. Everything in the region must fit.
|
||||
let col_alignment = check_column_alignment(items, &columns, mode);
|
||||
let min_alignment = match mode {
|
||||
TableDetectionMode::SmallFont => 0.5,
|
||||
@@ -1047,29 +912,6 @@ fn detect_table_in_region(
|
||||
cells.push(row_cells);
|
||||
}
|
||||
|
||||
// Validation 0 (small-font pass only): reject tiny all-numeric
|
||||
// fragments. A <=2-row grid whose every cell is a bare 1-2 digit number
|
||||
// carries no tabular information — in practice these are
|
||||
// exponent/subscript clusters from display math that happen to align.
|
||||
// Body-font tables are not subject to this veto: their cells cannot be
|
||||
// script glyphs.
|
||||
if matches!(mode, TableDetectionMode::SmallFont) {
|
||||
let nonempty_cells: Vec<&String> =
|
||||
cells.iter().flatten().filter(|c| !c.is_empty()).collect();
|
||||
if rows.len() <= 2
|
||||
&& !nonempty_cells.is_empty()
|
||||
&& nonempty_cells
|
||||
.iter()
|
||||
.all(|c| c.len() <= 2 && c.chars().all(|ch| ch.is_ascii_digit()))
|
||||
{
|
||||
log::debug!(
|
||||
" validation 0 fail: tiny all-numeric fragment ({} cells)",
|
||||
nonempty_cells.len()
|
||||
);
|
||||
return None;
|
||||
}
|
||||
}
|
||||
|
||||
// Validation 1: some rows should have content in first column.
|
||||
// Use a lower threshold (25%) for tables with wrapped cells where
|
||||
// continuation lines leave the first column empty.
|
||||
@@ -2135,146 +1977,6 @@ fn try_add_label_column(
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
|
||||
fn make_item(text: &str, x: f32, y: f32, font_size: f32, width: f32) -> TextItem {
|
||||
TextItem {
|
||||
text: text.to_string(),
|
||||
x,
|
||||
y,
|
||||
width,
|
||||
height: font_size,
|
||||
font: "TestFont".to_string(),
|
||||
font_size,
|
||||
page: 1,
|
||||
is_bold: false,
|
||||
is_italic: false,
|
||||
is_underline: false,
|
||||
is_strikeout: false,
|
||||
item_type: ItemType::Text,
|
||||
mcid: None,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_detects_subscript_after_body_text() {
|
||||
let body = make_item("log", 100.0, 500.0, 10.0, 15.0);
|
||||
let sub = make_item("10", 115.5, 497.0, 7.0, 7.0);
|
||||
let items = vec![body, sub.clone()];
|
||||
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sub, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_detects_superscript_footnote_marker() {
|
||||
let body = make_item("Hartley", 200.0, 500.0, 10.0, 35.0);
|
||||
let sup = make_item("2", 235.8, 504.0, 6.6, 3.5);
|
||||
let items = vec![body, sup.clone()];
|
||||
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sup, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_ignores_small_cell_far_from_body_text() {
|
||||
let body = make_item("Revenue", 100.0, 500.0, 10.0, 40.0);
|
||||
let cell = make_item("1,234", 180.0, 500.0, 7.0, 20.0);
|
||||
let items = vec![body, cell.clone()];
|
||||
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn body_pass_anchor_spares_cells_beside_slightly_larger_labels() {
|
||||
// A body-font table cell (10pt) sitting beside a slightly larger,
|
||||
// NON-heading label (12.5pt) with a little baseline jitter. The
|
||||
// small-font pass treats any larger neighbour as a possible script
|
||||
// base, but the body pass must not: at body sizes a slightly larger
|
||||
// neighbour is a bold label or column header, and flagging the cell
|
||||
// would strip it out of the table geometry and lose the table.
|
||||
// Cell at the low end of the body band (0.85x base) beside a 10.5pt
|
||||
// label. 10.5 clears the inherent 1.2x-of-cell rule (10.2) but falls
|
||||
// below the body pass's heading anchor (11.5), which is exactly the
|
||||
// band where the two masks must disagree.
|
||||
let label = make_item("Revenue", 100.0, 500.0, 10.5, 40.0);
|
||||
let cell = make_item("1,234", 141.0, 496.5, 8.5, 22.0);
|
||||
let items = vec![label, cell.clone()];
|
||||
let index = ScriptBodyIndex::new(&items);
|
||||
let base = 10.0;
|
||||
assert!(
|
||||
index.is_script_attachment(&cell, 0.0),
|
||||
"small-font pass anchor should still see this as an attachment"
|
||||
);
|
||||
assert!(
|
||||
!index.is_script_attachment(&cell, base * 1.15),
|
||||
"body pass must not treat a cell beside a slightly larger label \
|
||||
as a script — that removes real cells from the geometry"
|
||||
);
|
||||
// A genuine heading-sized anchor still qualifies in the body pass.
|
||||
let heading = make_item("Section", 100.0, 500.0, 20.0, 60.0);
|
||||
let sup = make_item("3", 161.0, 508.0, 10.0, 5.0);
|
||||
let h_items = vec![heading, sup.clone()];
|
||||
assert!(
|
||||
ScriptBodyIndex::new(&h_items).is_script_attachment(&sup, base * 1.15),
|
||||
"script hanging off a heading must still be excluded in the body pass"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_ignores_same_baseline_neighbor_cell() {
|
||||
// A small cell beside a larger label on the SAME baseline is a table
|
||||
// layout, not a subscript — a genuine baseline offset is required.
|
||||
let label = make_item("Total", 100.0, 500.0, 10.0, 25.0);
|
||||
let cell = make_item("42", 127.0, 500.0, 7.5, 9.0);
|
||||
let items = vec![label, cell.clone()];
|
||||
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_ignores_neighbor_on_different_line() {
|
||||
let body = make_item("Header", 100.0, 500.0, 10.0, 30.0);
|
||||
let cell = make_item("42", 131.0, 486.0, 7.0, 10.0);
|
||||
let items = vec![body, cell.clone()];
|
||||
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||
}
|
||||
|
||||
/// Equation-subscript + footnote layout from Shannon entropy.pdf page 1,
|
||||
/// with real coordinates. Without the larger-font anchors the small items
|
||||
/// alone DO form a phantom table — proving the layout reaches detection —
|
||||
/// and adding the anchors must suppress it.
|
||||
fn shannon_page1_small_items() -> Vec<TextItem> {
|
||||
vec![
|
||||
make_item("2", 267.4, 133.9, 7.4, 3.7),
|
||||
make_item("10", 306.2, 133.9, 7.4, 7.4),
|
||||
make_item("10", 342.7, 133.9, 7.4, 7.4),
|
||||
make_item("10", 325.0, 118.9, 7.4, 7.4),
|
||||
make_item("Bell System Technical Journal,", 295.7, 101.9, 8.0, 95.0),
|
||||
make_item(
|
||||
"April 1924, p. 324; Certain Topics in",
|
||||
396.7,
|
||||
101.9,
|
||||
8.0,
|
||||
130.0,
|
||||
),
|
||||
make_item("v. 47, April 1928, p. 617.", 250.9, 92.5, 8.0, 90.0),
|
||||
make_item("Bell System Technical Journal,", 264.2, 82.6, 8.0, 95.0),
|
||||
make_item("July 1928, p. 535.", 364.3, 82.6, 8.0, 65.0),
|
||||
]
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn equation_scripts_do_not_form_phantom_table() {
|
||||
let bare = shannon_page1_small_items();
|
||||
assert!(
|
||||
!detect_tables(&bare, 10.0, false).is_empty(),
|
||||
"test layout must form a phantom table when the filter cannot fire"
|
||||
);
|
||||
let mut items = shannon_page1_small_items();
|
||||
items.push(make_item("log", 253.0, 137.0, 10.0, 13.5));
|
||||
items.push(make_item("log", 291.5, 137.0, 10.0, 13.5));
|
||||
items.push(make_item("log", 328.0, 137.0, 10.0, 13.5));
|
||||
items.push(make_item("log", 310.3, 122.0, 10.0, 13.5));
|
||||
let tables = detect_tables(&items, 10.0, false);
|
||||
assert!(
|
||||
tables.is_empty(),
|
||||
"equation scripts + footnotes must not become a table: {tables:?}"
|
||||
);
|
||||
}
|
||||
use super::*;
|
||||
use crate::types::ItemType;
|
||||
|
||||
|
||||
+2
-450
@@ -4,10 +4,10 @@
|
||||
//! gridlines. Many IRS forms and government PDFs use these instead of
|
||||
//! `re` (rectangle) operators.
|
||||
|
||||
use std::collections::{HashMap, HashSet};
|
||||
use std::collections::HashSet;
|
||||
|
||||
use crate::tables::Table;
|
||||
use crate::types::{PdfLine, PdfRect, TextItem};
|
||||
use crate::types::{PdfLine, TextItem};
|
||||
|
||||
use super::detect_rects::{assign_items_to_grid, snap_edges};
|
||||
|
||||
@@ -15,33 +15,11 @@ const RULE_Y_TOLERANCE: f32 = 2.0;
|
||||
const RULE_JOIN_GAP: f32 = 6.0;
|
||||
const RULE_SPAN_TOLERANCE: f32 = 8.0;
|
||||
const TEXT_ROW_TOLERANCE: f32 = 2.5;
|
||||
const DENSE_CHART_MIN_VERTICAL_EDGES: usize = 27;
|
||||
const DENSE_CHART_LABEL_PAD: f32 = 20.0;
|
||||
const DENSE_CHART_MAX_SHARED_PANEL_GRIDS: usize = 4;
|
||||
|
||||
type HorizontalRule = (f32, f32, f32); // (y, x_min, x_max)
|
||||
type VerticalRule = (f32, f32, f32); // (x, y_min, y_max)
|
||||
type AnchoredRow<'a> = (f32, Vec<(usize, &'a TextItem)>);
|
||||
|
||||
fn dense_chart_grids_are_co_located(
|
||||
left: (f32, f32, f32, f32),
|
||||
right: (f32, f32, f32, f32),
|
||||
) -> bool {
|
||||
let left_width = left.2 - left.0;
|
||||
let right_width = right.2 - right.0;
|
||||
let left_height = left.3 - left.1;
|
||||
let right_height = right.3 - right.1;
|
||||
let horizontal_overlap = (left.2.min(right.2) - left.0.max(right.0)).max(0.0);
|
||||
let vertical_overlap = (left.3.min(right.3) - left.1.max(right.1)).max(0.0);
|
||||
let horizontal_gap = (left.0.max(right.0) - left.2.min(right.2)).max(0.0);
|
||||
let vertical_gap = (left.1.max(right.1) - left.3.min(right.3)).max(0.0);
|
||||
|
||||
(vertical_overlap >= left_height.min(right_height) * 0.5
|
||||
&& horizontal_gap <= left_width.min(right_width) * 0.5)
|
||||
|| (horizontal_overlap >= left_width.min(right_width) * 0.5
|
||||
&& vertical_gap <= left_height.min(right_height) * 0.5)
|
||||
}
|
||||
|
||||
#[derive(Debug)]
|
||||
struct TextAnchorTable {
|
||||
table: Table,
|
||||
@@ -1209,279 +1187,6 @@ pub fn detect_tables_from_lines(items: &[TextItem], lines: &[PdfLine], page: u32
|
||||
detect_tables_from_lines_inner(items, lines, page, true, true)
|
||||
}
|
||||
|
||||
/// Bounding boxes of chart panels backed by a very dense vector grid.
|
||||
///
|
||||
/// Tables support at most 25 columns, so a panel with at least 27 distinct,
|
||||
/// long vertical coordinates plus repeated horizontal rules is treated as
|
||||
/// chart geometry. When the grid is enclosed by a painted panel rectangle,
|
||||
/// the region expands to that rectangle so axis labels, legends, and source
|
||||
/// notes remain part of the figure instead of forming a heuristic table.
|
||||
pub(crate) fn detect_dense_line_chart_regions(
|
||||
lines: &[PdfLine],
|
||||
rects: &[PdfRect],
|
||||
page: u32,
|
||||
) -> Vec<(f32, f32, f32, f32)> {
|
||||
const ANGLE_TOLERANCE: f32 = 0.035;
|
||||
const MIN_GRID_LINE_LENGTH: f32 = 40.0;
|
||||
const EXTENT_TOLERANCE: f32 = 6.0;
|
||||
|
||||
let mut verticals = Vec::new();
|
||||
let mut horizontals = Vec::new();
|
||||
for line in lines.iter().filter(|line| line.page == page) {
|
||||
let dx = (line.x2 - line.x1).abs();
|
||||
let dy = (line.y2 - line.y1).abs();
|
||||
let length = dx.hypot(dy);
|
||||
if length < MIN_GRID_LINE_LENGTH {
|
||||
continue;
|
||||
}
|
||||
if dy > 0.01 && dx / dy <= ANGLE_TOLERANCE {
|
||||
verticals.push((
|
||||
(line.x1 + line.x2) / 2.0,
|
||||
line.y1.min(line.y2),
|
||||
line.y1.max(line.y2),
|
||||
));
|
||||
} else if dx > 0.01 && dy / dx <= ANGLE_TOLERANCE {
|
||||
horizontals.push((
|
||||
(line.y1 + line.y2) / 2.0,
|
||||
line.x1.min(line.x2),
|
||||
line.x1.max(line.x2),
|
||||
));
|
||||
}
|
||||
}
|
||||
if verticals.len() < DENSE_CHART_MIN_VERTICAL_EDGES || horizontals.len() < 3 {
|
||||
return Vec::new();
|
||||
}
|
||||
|
||||
// Group similar vertical extents once. Neighboring buckets are consulted
|
||||
// below so coordinates that straddle a bucket boundary still form one
|
||||
// family, while each line participates in only a constant number of
|
||||
// candidates instead of being re-scanned for every vertical anchor.
|
||||
let extent_key = |value: f32| (value / EXTENT_TOLERANCE).round() as i32;
|
||||
let mut extent_buckets: HashMap<(i32, i32), Vec<VerticalRule>> = HashMap::new();
|
||||
for vertical in verticals {
|
||||
extent_buckets
|
||||
.entry((extent_key(vertical.1), extent_key(vertical.2)))
|
||||
.or_default()
|
||||
.push(vertical);
|
||||
}
|
||||
|
||||
let mut grid_regions = Vec::new();
|
||||
let extent_keys: Vec<(i32, i32)> = extent_buckets.keys().copied().collect();
|
||||
for key in extent_keys {
|
||||
let anchor_family = &extent_buckets[&key];
|
||||
let anchor_bottom = anchor_family.iter().map(|vertical| vertical.1).sum::<f32>()
|
||||
/ anchor_family.len() as f32;
|
||||
let anchor_top = anchor_family.iter().map(|vertical| vertical.2).sum::<f32>()
|
||||
/ anchor_family.len() as f32;
|
||||
|
||||
let mut family = Vec::new();
|
||||
for bottom_offset in -1..=1 {
|
||||
for top_offset in -1..=1 {
|
||||
if let Some(bucket) =
|
||||
extent_buckets.get(&(key.0 + bottom_offset, key.1 + top_offset))
|
||||
{
|
||||
family.extend(bucket.iter().copied().filter(|vertical| {
|
||||
(vertical.1 - anchor_bottom).abs() <= EXTENT_TOLERANCE
|
||||
&& (vertical.2 - anchor_top).abs() <= EXTENT_TOLERANCE
|
||||
}));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
let xs = snap_edges(&family.iter().map(|&(x, _, _)| x).collect::<Vec<_>>(), 3.0);
|
||||
if xs.len() < DENSE_CHART_MIN_VERTICAL_EDGES {
|
||||
continue;
|
||||
}
|
||||
|
||||
let grid_bottom =
|
||||
family.iter().map(|vertical| vertical.1).sum::<f32>() / family.len() as f32;
|
||||
let grid_top = family.iter().map(|vertical| vertical.2).sum::<f32>() / family.len() as f32;
|
||||
if grid_top - grid_bottom < 60.0 {
|
||||
continue;
|
||||
}
|
||||
|
||||
// A horizontal rule must support the same contiguous dense run of
|
||||
// vertical coordinates. Splitting at sparse X gaps prevents a shared
|
||||
// rule from joining a chart to a neighboring ruled table. Keying by
|
||||
// the covered X-index range also lets multiple chart panels sharing
|
||||
// the same Y extents produce independent regions.
|
||||
let mut supported_spans: HashMap<(usize, usize), Vec<f32>> = HashMap::new();
|
||||
for &(y, line_left, line_right) in &horizontals {
|
||||
if y < grid_bottom - EXTENT_TOLERANCE || y > grid_top + EXTENT_TOLERANCE {
|
||||
continue;
|
||||
}
|
||||
let start = xs.partition_point(|&x| x < line_left - EXTENT_TOLERANCE);
|
||||
let end = xs.partition_point(|&x| x <= line_right + EXTENT_TOLERANCE);
|
||||
if end - start < DENSE_CHART_MIN_VERTICAL_EDGES {
|
||||
continue;
|
||||
}
|
||||
|
||||
let mut gaps: Vec<f32> = xs[start..end]
|
||||
.windows(2)
|
||||
.map(|pair| pair[1] - pair[0])
|
||||
.collect();
|
||||
gaps.sort_by(f32::total_cmp);
|
||||
let dense_gap = gaps[gaps.len() / 4];
|
||||
let run_break = (dense_gap * 3.0).max(12.0);
|
||||
let locally_dense_gap_limit = (dense_gap * 1.5).max(6.0);
|
||||
let locally_dense_gaps = gaps
|
||||
.iter()
|
||||
.filter(|&&gap| gap <= locally_dense_gap_limit)
|
||||
.count();
|
||||
|
||||
let mut run_start = start;
|
||||
let mut retained_dense_run = false;
|
||||
for index in start..end - 1 {
|
||||
if xs[index + 1] - xs[index] <= run_break {
|
||||
continue;
|
||||
}
|
||||
let run_end = index + 1;
|
||||
if run_end - run_start >= DENSE_CHART_MIN_VERTICAL_EDGES
|
||||
&& xs[run_end - 1] - xs[run_start] >= 120.0
|
||||
{
|
||||
supported_spans
|
||||
.entry((run_start, run_end))
|
||||
.or_default()
|
||||
.push(y);
|
||||
retained_dense_run = true;
|
||||
}
|
||||
run_start = run_end;
|
||||
}
|
||||
if end - run_start >= DENSE_CHART_MIN_VERTICAL_EDGES
|
||||
&& xs[end - 1] - xs[run_start] >= 120.0
|
||||
{
|
||||
supported_spans.entry((run_start, end)).or_default().push(y);
|
||||
retained_dense_run = true;
|
||||
}
|
||||
|
||||
// One or two wider category gaps may split an otherwise dense
|
||||
// chart into sub-threshold runs. Keep the full family only when
|
||||
// its total width remains close to the expected dense spacing;
|
||||
// a neighboring sparse table makes this ratio much larger.
|
||||
let span_width = xs[end - 1] - xs[start];
|
||||
let expected_dense_width = dense_gap * (end - start - 1) as f32;
|
||||
if !retained_dense_run
|
||||
&& span_width >= 120.0
|
||||
&& span_width <= expected_dense_width * 1.35
|
||||
&& gaps.len().saturating_sub(locally_dense_gaps) <= 2
|
||||
{
|
||||
supported_spans.entry((start, end)).or_default().push(y);
|
||||
}
|
||||
}
|
||||
|
||||
for ((start, end), ys) in supported_spans {
|
||||
if snap_edges(&ys, 3.0).len() >= 3 {
|
||||
grid_regions.push((xs[start], grid_bottom, xs[end - 1], grid_top));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Prefer the smallest qualifying region when a broad rule happens to
|
||||
// cover a denser nested panel, and retain every non-overlapping panel.
|
||||
grid_regions.sort_by(|left, right| {
|
||||
let left_area = (left.2 - left.0) * (left.3 - left.1);
|
||||
let right_area = (right.2 - right.0) * (right.3 - right.1);
|
||||
left_area.total_cmp(&right_area)
|
||||
});
|
||||
let mut selected_regions: Vec<(f32, f32, f32, f32)> = Vec::new();
|
||||
for region in grid_regions {
|
||||
let area = (region.2 - region.0) * (region.3 - region.1);
|
||||
let duplicates_existing = selected_regions.iter().any(|existing| {
|
||||
let overlap_width = (region.2.min(existing.2) - region.0.max(existing.0)).max(0.0);
|
||||
let overlap_height = (region.3.min(existing.3) - region.1.max(existing.1)).max(0.0);
|
||||
let overlap_area = overlap_width * overlap_height;
|
||||
let existing_area = (existing.2 - existing.0) * (existing.3 - existing.1);
|
||||
overlap_area >= area.min(existing_area) * 0.8
|
||||
});
|
||||
if !duplicates_existing {
|
||||
selected_regions.push(region);
|
||||
}
|
||||
}
|
||||
|
||||
let all_grid_regions = selected_regions.clone();
|
||||
let mut regions: Vec<_> = selected_regions
|
||||
.into_iter()
|
||||
.map(|grid_region| {
|
||||
let (grid_left, grid_bottom, grid_right, grid_top) = grid_region;
|
||||
let enclosing_panel = rects
|
||||
.iter()
|
||||
.filter(|rect| rect.page == page)
|
||||
.filter_map(|rect| {
|
||||
let (left, width) = if rect.width < 0.0 {
|
||||
(rect.x + rect.width, -rect.width)
|
||||
} else {
|
||||
(rect.x, rect.width)
|
||||
};
|
||||
let (bottom, height) = if rect.height < 0.0 {
|
||||
(rect.y + rect.height, -rect.height)
|
||||
} else {
|
||||
(rect.y, rect.height)
|
||||
};
|
||||
let right = left + width;
|
||||
let top = bottom + height;
|
||||
let enclosed_grids: Vec<_> = all_grid_regions
|
||||
.iter()
|
||||
.filter(|&&(other_left, other_bottom, other_right, other_top)| {
|
||||
left <= other_left + EXTENT_TOLERANCE
|
||||
&& right >= other_right - EXTENT_TOLERANCE
|
||||
&& bottom <= other_bottom + EXTENT_TOLERANCE
|
||||
&& top >= other_top - EXTENT_TOLERANCE
|
||||
})
|
||||
.copied()
|
||||
.collect();
|
||||
if enclosed_grids.len() > DENSE_CHART_MAX_SHARED_PANEL_GRIDS
|
||||
|| enclosed_grids.iter().any(|&other| {
|
||||
other != grid_region
|
||||
&& !dense_chart_grids_are_co_located(grid_region, other)
|
||||
})
|
||||
{
|
||||
return None;
|
||||
}
|
||||
let enclosed_grid_bounds =
|
||||
enclosed_grids.into_iter().reduce(|bounds, other| {
|
||||
(
|
||||
bounds.0.min(other.0),
|
||||
bounds.1.min(other.1),
|
||||
bounds.2.max(other.2),
|
||||
bounds.3.max(other.3),
|
||||
)
|
||||
})?;
|
||||
let enclosed_width = enclosed_grid_bounds.2 - enclosed_grid_bounds.0;
|
||||
let enclosed_height = enclosed_grid_bounds.3 - enclosed_grid_bounds.1;
|
||||
(left <= grid_left + EXTENT_TOLERANCE
|
||||
&& right >= grid_right - EXTENT_TOLERANCE
|
||||
&& bottom <= grid_bottom + EXTENT_TOLERANCE
|
||||
&& top >= grid_top - EXTENT_TOLERANCE
|
||||
&& width <= enclosed_width * 2.0
|
||||
&& height <= enclosed_height * 4.0
|
||||
&& !(left < 5.0 && bottom < 5.0))
|
||||
.then_some(((left, bottom, right, top), width * height))
|
||||
})
|
||||
.min_by(|left, right| left.1.total_cmp(&right.1))
|
||||
.map(|(region, _)| region);
|
||||
|
||||
enclosing_panel.unwrap_or((
|
||||
grid_left - DENSE_CHART_LABEL_PAD,
|
||||
grid_bottom - DENSE_CHART_LABEL_PAD,
|
||||
grid_right + DENSE_CHART_LABEL_PAD,
|
||||
grid_top + DENSE_CHART_LABEL_PAD,
|
||||
))
|
||||
})
|
||||
.collect();
|
||||
regions.sort_by(|left, right| {
|
||||
left.0
|
||||
.total_cmp(&right.0)
|
||||
.then_with(|| left.1.total_cmp(&right.1))
|
||||
});
|
||||
regions.dedup_by(|left, right| {
|
||||
(left.0 - right.0).abs() <= EXTENT_TOLERANCE
|
||||
&& (left.1 - right.1).abs() <= EXTENT_TOLERANCE
|
||||
&& (left.2 - right.2).abs() <= EXTENT_TOLERANCE
|
||||
&& (left.3 - right.3).abs() <= EXTENT_TOLERANCE
|
||||
});
|
||||
regions
|
||||
}
|
||||
|
||||
/// Detect only tables whose cell grid is backed by explicit vector geometry.
|
||||
///
|
||||
/// Region-level TSR callers need physical cell boundaries for crop bboxes, so
|
||||
@@ -1916,159 +1621,6 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dense_vector_grid_expands_to_enclosing_chart_panel() {
|
||||
let mut lines: Vec<PdfLine> = (0..30)
|
||||
.map(|column| make_vline(100.0 + column as f32 * 8.0, 400.0, 550.0, 1))
|
||||
.collect();
|
||||
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 332.0, 1)));
|
||||
let rects = vec![PdfRect {
|
||||
x: 80.0,
|
||||
y: 350.0,
|
||||
width: 280.0,
|
||||
height: 240.0,
|
||||
page: 1,
|
||||
}];
|
||||
|
||||
assert_eq!(
|
||||
detect_dense_line_chart_regions(&lines, &rects, 1),
|
||||
vec![(80.0, 350.0, 360.0, 590.0)]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn frameless_dense_vector_grid_includes_label_padding() {
|
||||
let mut lines: Vec<PdfLine> = (0..30)
|
||||
.map(|column| make_vline(100.0 + column as f32 * 8.0, 400.0, 550.0, 1))
|
||||
.collect();
|
||||
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 332.0, 1)));
|
||||
|
||||
assert_eq!(
|
||||
detect_dense_line_chart_regions(&lines, &[], 1),
|
||||
vec![(80.0, 380.0, 352.0, 570.0)]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn multiple_dense_vector_panels_are_retained() {
|
||||
let mut lines = Vec::new();
|
||||
for panel_left in [60.0, 380.0] {
|
||||
lines.extend(
|
||||
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
|
||||
);
|
||||
lines.extend((0..6).map(|row| {
|
||||
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
|
||||
}));
|
||||
}
|
||||
|
||||
assert_eq!(
|
||||
detect_dense_line_chart_regions(&lines, &[], 1),
|
||||
vec![(40.0, 380.0, 312.0, 570.0), (360.0, 380.0, 632.0, 570.0),]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn multiple_dense_vector_panels_use_shared_enclosing_panel() {
|
||||
let mut lines = Vec::new();
|
||||
for panel_left in [60.0, 380.0] {
|
||||
lines.extend(
|
||||
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
|
||||
);
|
||||
lines.extend((0..6).map(|row| {
|
||||
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
|
||||
}));
|
||||
}
|
||||
let rects = vec![PdfRect {
|
||||
x: 40.0,
|
||||
y: 350.0,
|
||||
width: 592.0,
|
||||
height: 240.0,
|
||||
page: 1,
|
||||
}];
|
||||
|
||||
assert_eq!(
|
||||
detect_dense_line_chart_regions(&lines, &rects, 1),
|
||||
vec![(40.0, 350.0, 632.0, 590.0)]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn shared_rules_do_not_join_dense_chart_to_adjacent_table() {
|
||||
let mut lines: Vec<PdfLine> = (0..30)
|
||||
.map(|column| make_vline(60.0 + column as f32 * 8.0, 400.0, 550.0, 1))
|
||||
.collect();
|
||||
|
||||
lines.extend(
|
||||
[330.0, 390.0, 450.0, 510.0, 570.0, 630.0]
|
||||
.into_iter()
|
||||
.map(|x| make_vline(x, 400.0, 550.0, 1)),
|
||||
);
|
||||
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 630.0, 1)));
|
||||
|
||||
assert_eq!(
|
||||
detect_dense_line_chart_regions(&lines, &[], 1),
|
||||
vec![(40.0, 380.0, 312.0, 570.0)]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn uneven_dense_spacing_keeps_the_complete_chart_region() {
|
||||
let mut xs: Vec<f32> = (0..15).map(|column| 60.0 + column as f32 * 8.0).collect();
|
||||
xs.extend((0..15).map(|column| 212.0 + column as f32 * 8.0));
|
||||
let mut lines: Vec<PdfLine> = xs.iter().map(|&x| make_vline(x, 400.0, 550.0, 1)).collect();
|
||||
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 324.0, 1)));
|
||||
|
||||
assert_eq!(
|
||||
detect_dense_line_chart_regions(&lines, &[], 1),
|
||||
vec![(40.0, 380.0, 344.0, 570.0)]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn subthreshold_dense_run_does_not_absorb_adjacent_sparse_grid() {
|
||||
let mut xs: Vec<f32> = (0..21).map(|column| 60.0 + column as f32 * 8.0).collect();
|
||||
xs.extend((0..6).map(|column| 248.0 + column as f32 * 18.0));
|
||||
let mut lines: Vec<PdfLine> = xs.iter().map(|&x| make_vline(x, 400.0, 550.0, 1)).collect();
|
||||
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 60.0, 338.0, 1)));
|
||||
|
||||
assert!(detect_dense_line_chart_regions(&lines, &[], 1).is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn broad_frame_does_not_merge_distant_dense_grids() {
|
||||
let mut lines = Vec::new();
|
||||
for panel_left in [60.0, 700.0] {
|
||||
lines.extend(
|
||||
(0..30).map(|column| make_vline(panel_left + column as f32 * 8.0, 400.0, 550.0, 1)),
|
||||
);
|
||||
lines.extend((0..6).map(|row| {
|
||||
make_hline(400.0 + row as f32 * 30.0, panel_left, panel_left + 232.0, 1)
|
||||
}));
|
||||
}
|
||||
let rects = vec![PdfRect {
|
||||
x: 40.0,
|
||||
y: 350.0,
|
||||
width: 912.0,
|
||||
height: 240.0,
|
||||
page: 1,
|
||||
}];
|
||||
|
||||
assert_eq!(
|
||||
detect_dense_line_chart_regions(&lines, &rects, 1),
|
||||
vec![(40.0, 380.0, 312.0, 570.0), (680.0, 380.0, 952.0, 570.0)]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn supported_width_vector_table_is_not_a_dense_chart() {
|
||||
let mut lines: Vec<PdfLine> = (0..26)
|
||||
.map(|column| make_vline(100.0 + column as f32 * 10.0, 400.0, 550.0, 1))
|
||||
.collect();
|
||||
lines.extend((0..6).map(|row| make_hline(400.0 + row as f32 * 30.0, 100.0, 350.0, 1)));
|
||||
|
||||
assert!(detect_dense_line_chart_regions(&lines, &[], 1).is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_basic_grid_detection() {
|
||||
// 3x2 grid with horizontal lines at y=500, 480, 460 and vertical at x=100, 200, 300
|
||||
|
||||
+10
-454
@@ -1996,249 +1996,17 @@ fn without_dominant_page_backgrounds(rects: &[(f32, f32, f32, f32)]) -> Vec<(f32
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Repeated rows of touching cell rectangles are stronger table evidence
|
||||
/// than the bar-length variation used by the chart detector.
|
||||
/// Detect a table from cell-background rects that failed grid detection.
|
||||
///
|
||||
/// Ruled tables with wrapped labels naturally have variable row heights, and
|
||||
/// numeric-heavy cells can otherwise resemble horizontal or vertical bars.
|
||||
/// Require several rows to repeat a shared edge schema before overriding the
|
||||
/// chart hypothesis so sparse plots and independent bars remain unaffected.
|
||||
fn is_repeated_cell_grid(group_rects: &[(f32, f32, f32, f32)]) -> bool {
|
||||
type RowGroup = (f32, f32, Vec<(f32, f32)>);
|
||||
|
||||
const ROW_EDGE_TOLERANCE: f32 = 3.0;
|
||||
const MIN_GRID_ROWS: usize = 4;
|
||||
const MIN_CELLS_PER_ROW: usize = 3;
|
||||
|
||||
if group_rects.len() < MIN_GRID_ROWS * MIN_CELLS_PER_ROW {
|
||||
return false;
|
||||
}
|
||||
|
||||
let mut row_groups: Vec<RowGroup> = Vec::new();
|
||||
for &(x, y, width, height) in group_rects {
|
||||
if width < 5.0 || height < 5.0 {
|
||||
continue;
|
||||
}
|
||||
let top = y + height;
|
||||
if let Some((_, _, cells)) = row_groups.iter_mut().find(|(bottom, row_top, _)| {
|
||||
(y - *bottom).abs() <= ROW_EDGE_TOLERANCE
|
||||
&& (top - *row_top).abs() <= ROW_EDGE_TOLERANCE
|
||||
}) {
|
||||
cells.push((x, x + width));
|
||||
} else {
|
||||
row_groups.push((y, top, vec![(x, x + width)]));
|
||||
}
|
||||
}
|
||||
|
||||
let mut row_schemas = Vec::new();
|
||||
for (_, _, mut cells) in row_groups {
|
||||
if cells.len() < MIN_CELLS_PER_ROW {
|
||||
continue;
|
||||
}
|
||||
let mut widths: Vec<f32> = cells.iter().map(|&(left, right)| right - left).collect();
|
||||
widths.sort_by(f32::total_cmp);
|
||||
let median_width = widths[widths.len() / 2];
|
||||
cells.retain(|&(left, right)| right - left <= median_width * 2.5);
|
||||
cells.sort_by(|left, right| {
|
||||
left.0
|
||||
.total_cmp(&right.0)
|
||||
.then_with(|| left.1.total_cmp(&right.1))
|
||||
});
|
||||
cells.dedup_by(|left, right| {
|
||||
(left.0 - right.0).abs() <= ROW_EDGE_TOLERANCE
|
||||
&& (left.1 - right.1).abs() <= ROW_EDGE_TOLERANCE
|
||||
});
|
||||
if cells.len() < MIN_CELLS_PER_ROW
|
||||
|| cells
|
||||
.windows(2)
|
||||
.any(|pair| pair[1].0 > pair[0].1 + ROW_EDGE_TOLERANCE)
|
||||
{
|
||||
continue;
|
||||
}
|
||||
let edges: Vec<f32> = cells
|
||||
.iter()
|
||||
.flat_map(|&(left, right)| [left, right])
|
||||
.collect();
|
||||
let schema = snap_edges(&edges, ROW_EDGE_TOLERANCE);
|
||||
if schema.len() > MIN_CELLS_PER_ROW {
|
||||
row_schemas.push(schema);
|
||||
}
|
||||
}
|
||||
if row_schemas.len() < MIN_GRID_ROWS {
|
||||
return false;
|
||||
}
|
||||
|
||||
let reference = row_schemas
|
||||
.iter()
|
||||
.max_by_key(|schema| schema.len())
|
||||
.expect("grid rows are non-empty");
|
||||
row_schemas
|
||||
.iter()
|
||||
.filter(|schema| {
|
||||
let comparable_edges = reference.len().min(schema.len());
|
||||
let matched_edges = schema
|
||||
.iter()
|
||||
.filter(|edge| {
|
||||
reference
|
||||
.iter()
|
||||
.any(|reference_edge| (*edge - *reference_edge).abs() <= ROW_EDGE_TOLERANCE)
|
||||
})
|
||||
.count();
|
||||
matched_edges > MIN_CELLS_PER_ROW && matched_edges * 4 >= comparable_edges * 3
|
||||
})
|
||||
.count()
|
||||
>= MIN_GRID_ROWS
|
||||
}
|
||||
|
||||
fn repeated_cell_grid_overrides_bar_hypothesis(group_rects: &[(f32, f32, f32, f32)]) -> bool {
|
||||
is_repeated_cell_grid(group_rects)
|
||||
&& without_dominant_page_backgrounds(group_rects).len() == group_rects.len()
|
||||
}
|
||||
|
||||
/// Detect horizontal segmented stacks from aligned rows of touching rects.
|
||||
///
|
||||
/// Category rows must have visible gutters and data-varying internal segment
|
||||
/// boundaries, unlike the stable boundaries of a ruled table.
|
||||
struct SegmentedBarGeometry {
|
||||
bounds: (f32, f32, f32, f32),
|
||||
row_bands: Vec<(f32, f32)>,
|
||||
}
|
||||
|
||||
fn segmented_stacked_bar_geometry(
|
||||
group_rects: &[(f32, f32, f32, f32)],
|
||||
) -> Option<SegmentedBarGeometry> {
|
||||
type BarRow = (f32, f32, Vec<(f32, f32)>);
|
||||
|
||||
const EDGE_TOLERANCE: f32 = 3.0;
|
||||
const MIN_ROWS: usize = 4;
|
||||
const MIN_SEGMENTS: usize = 3;
|
||||
|
||||
let mut rows: Vec<BarRow> = Vec::new();
|
||||
for &(x, y, width, height) in group_rects {
|
||||
if width < 5.0 || height < 5.0 {
|
||||
continue;
|
||||
}
|
||||
let top = y + height;
|
||||
if let Some((_, _, segments)) = rows.iter_mut().find(|(bottom, row_top, _)| {
|
||||
(y - *bottom).abs() <= EDGE_TOLERANCE && (top - *row_top).abs() <= EDGE_TOLERANCE
|
||||
}) {
|
||||
segments.push((x, x + width));
|
||||
} else {
|
||||
rows.push((y, top, vec![(x, x + width)]));
|
||||
}
|
||||
}
|
||||
|
||||
rows.retain_mut(|(_, _, segments)| {
|
||||
segments.sort_by(|left, right| left.0.total_cmp(&right.0));
|
||||
segments.len() >= MIN_SEGMENTS
|
||||
&& segments
|
||||
.windows(2)
|
||||
.all(|pair| (pair[1].0 - pair[0].1).abs() <= EDGE_TOLERANCE)
|
||||
});
|
||||
if rows.len() < MIN_ROWS {
|
||||
return None;
|
||||
}
|
||||
rows.sort_by(|left, right| left.0.total_cmp(&right.0));
|
||||
|
||||
// Table rows normally share borders. Horizontal stacked bars instead
|
||||
// leave a visible gutter between category rows.
|
||||
if rows.windows(2).any(|pair| {
|
||||
let shorter_height = (pair[0].1 - pair[0].0).min(pair[1].1 - pair[1].0);
|
||||
pair[1].0 - pair[0].1 < (shorter_height * 0.25).max(2.0)
|
||||
}) {
|
||||
return None;
|
||||
}
|
||||
|
||||
// At least two rows must move an internal segment boundary. Stable
|
||||
// boundaries across every row are stronger evidence for a ruled table.
|
||||
let reference_edges: Vec<f32> = rows[0]
|
||||
.2
|
||||
.iter()
|
||||
.take(rows[0].2.len() - 1)
|
||||
.map(|segment| segment.1)
|
||||
.collect();
|
||||
let drifting_rows = rows
|
||||
.iter()
|
||||
.skip(1)
|
||||
.filter(|(_, _, segments)| {
|
||||
let edges: Vec<f32> = segments
|
||||
.iter()
|
||||
.take(segments.len() - 1)
|
||||
.map(|segment| segment.1)
|
||||
.collect();
|
||||
edges.len() == reference_edges.len()
|
||||
&& edges
|
||||
.iter()
|
||||
.zip(&reference_edges)
|
||||
.any(|(edge, reference)| (edge - reference).abs() > EDGE_TOLERANCE)
|
||||
})
|
||||
.count();
|
||||
if drifting_rows < 2 {
|
||||
return None;
|
||||
}
|
||||
|
||||
let left = rows
|
||||
.iter()
|
||||
.flat_map(|row| &row.2)
|
||||
.map(|segment| segment.0)
|
||||
.reduce(f32::min)?;
|
||||
let right = rows
|
||||
.iter()
|
||||
.flat_map(|row| &row.2)
|
||||
.map(|segment| segment.1)
|
||||
.reduce(f32::max)?;
|
||||
let bottom = rows.iter().map(|row| row.0).reduce(f32::min)?;
|
||||
let top = rows.iter().map(|row| row.1).reduce(f32::max)?;
|
||||
let row_bands = rows.iter().map(|row| (row.0, row.1)).collect();
|
||||
Some(SegmentedBarGeometry {
|
||||
bounds: (left, bottom, right, top),
|
||||
row_bands,
|
||||
})
|
||||
}
|
||||
|
||||
/// Category labels beside multiple bar rows are independent chart evidence:
|
||||
/// numeric table text stays inside its cells, regardless of whether the table
|
||||
/// has an outer border or extra padding.
|
||||
fn has_external_segmented_bar_labels(
|
||||
items: &[TextItem],
|
||||
page: u32,
|
||||
geometry: &SegmentedBarGeometry,
|
||||
) -> bool {
|
||||
const LABEL_EDGE_TOLERANCE: f32 = 3.0;
|
||||
const LABEL_CLAIM_PAD: f32 = 20.0;
|
||||
|
||||
let (content_left, _, content_right, _) = geometry.bounds;
|
||||
let labeled_rows = geometry
|
||||
.row_bands
|
||||
.iter()
|
||||
.filter(|&&(row_bottom, row_top)| {
|
||||
items.iter().any(|item| {
|
||||
if item.page != page || item.text.trim().is_empty() {
|
||||
return false;
|
||||
}
|
||||
let item_left = item.x.min(item.x + item.width);
|
||||
let item_right = item.x.max(item.x + item.width);
|
||||
let item_center_x = (item_left + item_right) / 2.0;
|
||||
let item_center_y = item.y + item.height / 2.0;
|
||||
let beside_stack = (item_center_x <= content_left + LABEL_EDGE_TOLERANCE
|
||||
&& item_center_x >= content_left - LABEL_CLAIM_PAD
|
||||
&& item_left < content_left)
|
||||
|| (item_center_x >= content_right - LABEL_EDGE_TOLERANCE
|
||||
&& item_center_x <= content_right + LABEL_CLAIM_PAD
|
||||
&& item_right > content_right);
|
||||
beside_stack
|
||||
&& item_center_y >= row_bottom - LABEL_EDGE_TOLERANCE
|
||||
&& item_center_y <= row_top + LABEL_EDGE_TOLERANCE
|
||||
})
|
||||
})
|
||||
.count();
|
||||
|
||||
labeled_rows >= 2 && labeled_rows * 2 >= geometry.row_bands.len()
|
||||
}
|
||||
|
||||
/// Recognize filled vertical or horizontal bars whose geometry and labels are
|
||||
/// data-driven rather than uniform table cells.
|
||||
fn has_chart_bar_signature(
|
||||
/// Uses rect Y-edges for row boundaries and text X-position clustering for
|
||||
/// columns. Handles tables with cell backgrounds that don't form a clean
|
||||
/// X-edge grid (variable column widths, decorative fills).
|
||||
/// Chart-bar signature: ≥3 rects sharing an aligned bottom edge (the axis),
|
||||
/// with similar widths (bars) but strongly varying heights (data-driven),
|
||||
/// holding at most a single numeric data label each. Bar charts drawn as
|
||||
/// filled rects otherwise read as cell rects and grid their axis labels
|
||||
/// into a phantom table. The mirrored check catches horizontal bar charts.
|
||||
fn is_chart_bar_cluster(
|
||||
items: &[TextItem],
|
||||
group_rects: &[(f32, f32, f32, f32)],
|
||||
page: u32,
|
||||
@@ -2345,30 +2113,6 @@ fn has_chart_bar_signature(
|
||||
|| bar_family(|r| r.1, |r| r.3, |r| r.2, |r| r.0)
|
||||
}
|
||||
|
||||
fn is_chart_bar_cluster(
|
||||
items: &[TextItem],
|
||||
group_rects: &[(f32, f32, f32, f32)],
|
||||
page: u32,
|
||||
) -> bool {
|
||||
let has_bar_signature = has_chart_bar_signature(items, group_rects, page);
|
||||
|
||||
// A segmented horizontal chart can share most of its edges across rows.
|
||||
// Row-aligned category labels outside the stack distinguish it from a
|
||||
// numeric table without depending on whether either shape has a frame.
|
||||
if has_bar_signature {
|
||||
if let Some(geometry) = segmented_stacked_bar_geometry(group_rects) {
|
||||
if has_external_segmented_bar_labels(items, page, &geometry) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
if repeated_cell_grid_overrides_bar_hypothesis(group_rects) {
|
||||
return false;
|
||||
}
|
||||
|
||||
has_bar_signature
|
||||
}
|
||||
|
||||
fn detect_row_stripe_table_from_cell_rects(
|
||||
items: &[TextItem],
|
||||
group_rects: &[(f32, f32, f32, f32)],
|
||||
@@ -3339,194 +3083,6 @@ mod tests {
|
||||
assert!(detect_chart_regions(&items, &rects, 1).is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn variable_height_ruled_grid_overrides_bar_hypothesis() {
|
||||
let edge_sets = [
|
||||
[80.0, 140.0, 200.0, 260.0, 320.0, 380.0, 440.0, 500.0, 560.0],
|
||||
[80.0, 140.0, 210.0, 260.0, 320.0, 380.0, 450.0, 500.0, 560.0],
|
||||
];
|
||||
let heights = [20.0, 34.0, 26.0, 42.0, 20.0, 34.0];
|
||||
let edge_variants = [0, 0, 0, 0, 1, 1];
|
||||
let mut rects = Vec::new();
|
||||
let mut y = 650.0;
|
||||
for (row, height) in heights.into_iter().enumerate() {
|
||||
let edges = edge_sets[edge_variants[row]];
|
||||
rects.extend(
|
||||
edges
|
||||
.windows(2)
|
||||
.map(|edge| (edge[0], y, edge[1] - edge[0], height)),
|
||||
);
|
||||
y -= height;
|
||||
}
|
||||
|
||||
assert!(is_repeated_cell_grid(&rects));
|
||||
assert!(has_chart_bar_signature(&[], &rects, 1));
|
||||
assert!(repeated_cell_grid_overrides_bar_hypothesis(&rects));
|
||||
assert!(segmented_stacked_bar_geometry(&rects).is_none());
|
||||
assert!(!is_chart_bar_cluster(&[], &rects, 1));
|
||||
|
||||
let mut with_page_fills =
|
||||
vec![(0.0, 0.0, 600.0, 800.0); DOMINANT_PAGE_BACKGROUND_MIN_REPETITIONS];
|
||||
with_page_fills.extend(rects);
|
||||
assert!(!repeated_cell_grid_overrides_bar_hypothesis(
|
||||
&with_page_fills
|
||||
));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn touching_segments_with_spaced_rows_remain_a_chart() {
|
||||
let row_edges = [
|
||||
[100.0, 140.0, 180.0, 220.0, 260.0],
|
||||
[100.0, 140.0, 180.0, 228.0, 260.0],
|
||||
[100.0, 140.0, 180.0, 214.0, 260.0],
|
||||
[100.0, 140.0, 180.0, 232.0, 260.0],
|
||||
];
|
||||
let mut raw_rects = vec![(90.0, 530.0, 190.0, 100.0)];
|
||||
for (row, edges) in row_edges.into_iter().enumerate() {
|
||||
let y = 540.0 + row as f32 * 20.0;
|
||||
raw_rects.extend(
|
||||
edges
|
||||
.windows(2)
|
||||
.map(|edge| (edge[0], y, edge[1] - edge[0], 12.0)),
|
||||
);
|
||||
}
|
||||
let items: Vec<TextItem> = (0..4)
|
||||
.map(|row| make_item("Category", 62.0, 541.0 + row as f32 * 20.0, 9.0))
|
||||
.collect();
|
||||
|
||||
assert!(is_repeated_cell_grid(&raw_rects));
|
||||
assert!(has_chart_bar_signature(&items, &raw_rects, 1));
|
||||
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented stack");
|
||||
assert!(has_external_segmented_bar_labels(&items, 1, &geometry));
|
||||
assert!(is_chart_bar_cluster(&items, &raw_rects, 1));
|
||||
|
||||
let numeric_items: Vec<TextItem> = (0..4)
|
||||
.map(|row| make_item("2024", 80.0, 541.0 + row as f32 * 18.0, 9.0))
|
||||
.collect();
|
||||
assert!(has_external_segmented_bar_labels(
|
||||
&numeric_items,
|
||||
1,
|
||||
&geometry
|
||||
));
|
||||
assert!(is_chart_bar_cluster(&numeric_items, &raw_rects, 1));
|
||||
|
||||
let edge_adjacent_items: Vec<TextItem> = (0..4)
|
||||
.map(|row| make_item("2024", 92.0, 541.0 + row as f32 * 18.0, 9.0))
|
||||
.collect();
|
||||
assert!(has_external_segmented_bar_labels(
|
||||
&edge_adjacent_items,
|
||||
1,
|
||||
&geometry
|
||||
));
|
||||
assert!(is_chart_bar_cluster(&edge_adjacent_items, &raw_rects, 1));
|
||||
|
||||
let far_items: Vec<TextItem> = (0..4)
|
||||
.map(|row| make_item("Category", 20.0, 541.0 + row as f32 * 18.0, 9.0))
|
||||
.collect();
|
||||
assert!(!has_external_segmented_bar_labels(&far_items, 1, &geometry));
|
||||
assert!(!is_chart_bar_cluster(&far_items, &raw_rects, 1));
|
||||
|
||||
let rects: Vec<PdfRect> = raw_rects
|
||||
.into_iter()
|
||||
.map(|(x, y, width, height)| PdfRect {
|
||||
x,
|
||||
y,
|
||||
width,
|
||||
height,
|
||||
page: 1,
|
||||
})
|
||||
.collect();
|
||||
assert_eq!(detect_chart_regions(&items, &rects, 1).len(), 1);
|
||||
let (tables, hints) = detect_tables_from_rects(&items, &rects, 1);
|
||||
assert!(tables.is_empty());
|
||||
assert!(hints.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn padded_numeric_grid_frame_remains_a_table() {
|
||||
let row_edges = [
|
||||
[100.0, 140.0, 180.0, 220.0, 260.0],
|
||||
[100.0, 140.0, 180.0, 228.0, 260.0],
|
||||
[100.0, 140.0, 180.0, 214.0, 260.0],
|
||||
[100.0, 140.0, 180.0, 232.0, 260.0],
|
||||
];
|
||||
let mut raw_rects = vec![(96.0, 536.0, 168.0, 80.0)];
|
||||
let mut items = Vec::new();
|
||||
for (row, edges) in row_edges.into_iter().enumerate() {
|
||||
let y = 540.0 + row as f32 * 20.0;
|
||||
for edge in edges.windows(2) {
|
||||
raw_rects.push((edge[0], y, edge[1] - edge[0], 12.0));
|
||||
items.push(make_item("42", edge[0] + 8.0, y + 1.0, 9.0));
|
||||
}
|
||||
}
|
||||
|
||||
assert!(is_repeated_cell_grid(&raw_rects));
|
||||
assert!(has_chart_bar_signature(&items, &raw_rects, 1));
|
||||
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented rows");
|
||||
assert!(!has_external_segmented_bar_labels(&items, 1, &geometry));
|
||||
assert!(!is_chart_bar_cluster(&items, &raw_rects, 1));
|
||||
|
||||
let flush_items: Vec<TextItem> = (0..4)
|
||||
.map(|row| make_item("1", 100.0, 541.0 + row as f32 * 20.0, 9.0))
|
||||
.collect();
|
||||
assert!(!has_external_segmented_bar_labels(
|
||||
&flush_items,
|
||||
1,
|
||||
&geometry
|
||||
));
|
||||
assert!(!is_chart_bar_cluster(&flush_items, &raw_rects, 1));
|
||||
|
||||
let rects: Vec<PdfRect> = raw_rects
|
||||
.into_iter()
|
||||
.map(|(x, y, width, height)| PdfRect {
|
||||
x,
|
||||
y,
|
||||
width,
|
||||
height,
|
||||
page: 1,
|
||||
})
|
||||
.collect();
|
||||
assert!(detect_chart_regions(&items, &rects, 1).is_empty());
|
||||
assert!(!detect_tables_from_rects(&items, &rects, 1).0.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn frameless_segmented_chart_with_category_labels_remains_a_chart() {
|
||||
let row_edges = [
|
||||
[100.0, 140.0, 180.0, 220.0, 260.0],
|
||||
[100.0, 140.0, 180.0, 228.0, 260.0],
|
||||
[100.0, 140.0, 180.0, 214.0, 260.0],
|
||||
[100.0, 140.0, 180.0, 232.0, 260.0],
|
||||
];
|
||||
let mut raw_rects = Vec::new();
|
||||
let mut items = Vec::new();
|
||||
for (row, edges) in row_edges.into_iter().enumerate() {
|
||||
let y = 540.0 + row as f32 * 18.0;
|
||||
raw_rects.extend(
|
||||
edges
|
||||
.windows(2)
|
||||
.map(|edge| (edge[0], y, edge[1] - edge[0], 12.0)),
|
||||
);
|
||||
items.push(make_item("Category", 62.0, y + 1.0, 9.0));
|
||||
}
|
||||
|
||||
let geometry = segmented_stacked_bar_geometry(&raw_rects).expect("segmented stack");
|
||||
assert!(has_external_segmented_bar_labels(&items, 1, &geometry));
|
||||
assert!(is_chart_bar_cluster(&items, &raw_rects, 1));
|
||||
|
||||
let rects: Vec<PdfRect> = raw_rects
|
||||
.into_iter()
|
||||
.map(|(x, y, width, height)| PdfRect {
|
||||
x,
|
||||
y,
|
||||
width,
|
||||
height,
|
||||
page: 1,
|
||||
})
|
||||
.collect();
|
||||
assert_eq!(detect_chart_regions(&items, &rects, 1).len(), 1);
|
||||
}
|
||||
|
||||
// --- detect_stacked_box_table ---
|
||||
|
||||
/// N stacked boxes at x=100, w=300, h=22, top-to-bottom from y=600.
|
||||
|
||||
+1
-3
@@ -16,9 +16,7 @@ pub(crate) use detect_heuristic::{
|
||||
content_width, detect_tables_with_page_width, is_table_of_contents,
|
||||
};
|
||||
pub use detect_lines::detect_tables_from_lines;
|
||||
pub(crate) use detect_lines::{
|
||||
detect_dense_line_chart_regions, detect_vector_grid_tables_from_lines,
|
||||
};
|
||||
pub(crate) use detect_lines::detect_vector_grid_tables_from_lines;
|
||||
pub(crate) use detect_rects::cluster_rects;
|
||||
pub use detect_rects::{detect_chart_regions, detect_tables_from_rects, RectHintRegion};
|
||||
pub use detect_struct::detect_tables_from_struct_tree;
|
||||
|
||||
+3
-10
@@ -70,17 +70,10 @@ pub(crate) fn is_page_number_line(text: &str) -> bool {
|
||||
|
||||
let lowercase = text.trim().to_ascii_lowercase();
|
||||
lowercase.strip_prefix("page").is_some_and(|rest| {
|
||||
let mut characters = rest.trim_start().chars().peekable();
|
||||
let mut has_page_number = false;
|
||||
while characters
|
||||
.peek()
|
||||
rest.trim_start()
|
||||
.chars()
|
||||
.next()
|
||||
.is_some_and(|character| character.is_ascii_digit())
|
||||
{
|
||||
has_page_number = true;
|
||||
characters.next();
|
||||
}
|
||||
|
||||
has_page_number && characters.next().is_none_or(char::is_whitespace)
|
||||
})
|
||||
}
|
||||
|
||||
|
||||
@@ -880,25 +880,6 @@ fn try_remap_subset_cmap(
|
||||
None => return (cmap, None),
|
||||
};
|
||||
|
||||
// Both repair paths below assume CIDs are glyph indices that a subsetter can
|
||||
// renumber, which is only true for CIDFontType2 (TrueType). For CIDFontType0
|
||||
// (CFF), CIDs are resolved through the CFF charset, so a valid CMap stays valid
|
||||
// after subsetting and renumbering it corrupts otherwise-correct text.
|
||||
// CIDToGIDMap is likewise CIDFontType2-only (PDF 32000-1:2008, 9.7.4.2), so this
|
||||
// also ignores a CIDToGIDMap that a malformed producer attached to a CFF font.
|
||||
// /Subtype may be an indirect reference, so resolve it through the document.
|
||||
// Only bail out when the descendant is *explicitly* something other than
|
||||
// CIDFontType2: a missing or unresolvable /Subtype keeps the previous
|
||||
// behaviour rather than silently disabling the repair.
|
||||
let subtype = cid_font_dict.get(b"Subtype").ok().and_then(|o| match o {
|
||||
Object::Reference(r) => doc.get_object(*r).ok().and_then(|o| o.as_name().ok()),
|
||||
other => other.as_name().ok(),
|
||||
});
|
||||
if subtype.is_some_and(|name| name != b"CIDFontType2") {
|
||||
debug!("Subset remap skipped for obj={obj_num}: descendant is not CIDFontType2");
|
||||
return (cmap, None);
|
||||
}
|
||||
|
||||
// If there's an explicit CIDToGIDMap, build a repaired CMap using it.
|
||||
if let Some(cid_to_gid) = get_cid_to_gid_map(cid_font_dict, doc) {
|
||||
if let Some(repaired) = build_cmap_with_cid_to_gid_map(&cmap, &cid_to_gid) {
|
||||
@@ -3178,106 +3159,4 @@ endbfrange
|
||||
"Remap must fire when CMap's CIDs are outside W array coverage"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_try_remap_skipped_for_cid_font_type0() {
|
||||
// Same W/CMap mismatch as the CIDFontType2 case above, but the descendant is
|
||||
// CIDFontType0 (CFF). There CIDs are resolved through the CFF charset, so the
|
||||
// ToUnicode CIDs stay valid after subsetting and must not be renumbered.
|
||||
// Real-world case: Japanese Adobe-Japan1 PDFs (e.g. National Diet Library
|
||||
// minutes) where remapping turned correct text into unrelated glyphs.
|
||||
let cmap_content = r#"
|
||||
1 begincodespacerange
|
||||
<0000><FFFF>
|
||||
endcodespacerange
|
||||
1 beginbfrange
|
||||
<0200> <0220> <0410>
|
||||
endbfrange
|
||||
"#;
|
||||
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
|
||||
|
||||
let mut doc = Document::new();
|
||||
|
||||
// CIDToGIDMap is CIDFontType2-only, but a malformed producer can still emit
|
||||
// one on a CFF font. Use a real stream (not /Identity, which is treated as
|
||||
// "no map") so this also fails if the guard is moved back below the
|
||||
// CIDToGIDMap branch: cid 1 -> gid 0x0200, which the CMap resolves.
|
||||
let mut cid_to_gid = vec![0u8; 68];
|
||||
cid_to_gid[2] = 0x02;
|
||||
cid_to_gid[3] = 0x00;
|
||||
let cid_to_gid_id =
|
||||
doc.add_object(lopdf::Stream::new(lopdf::Dictionary::new(), cid_to_gid));
|
||||
|
||||
let mut cid_font = lopdf::Dictionary::new();
|
||||
cid_font.set("Subtype", lopdf::Object::Name(b"CIDFontType0".to_vec()));
|
||||
cid_font.set("CIDToGIDMap", lopdf::Object::Reference(cid_to_gid_id));
|
||||
cid_font.set(
|
||||
"W",
|
||||
lopdf::Object::Array(vec![
|
||||
lopdf::Object::Integer(1),
|
||||
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
|
||||
]),
|
||||
);
|
||||
let cid_font_id = doc.add_object(cid_font);
|
||||
|
||||
let mut font_dict = lopdf::Dictionary::new();
|
||||
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
|
||||
font_dict.set(
|
||||
"DescendantFonts",
|
||||
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
|
||||
);
|
||||
|
||||
let (primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 789);
|
||||
assert!(
|
||||
remapped.is_none(),
|
||||
"Remap must be skipped for CIDFontType0 (CFF) descendants, including a \
|
||||
CIDToGIDMap a malformed producer attached to one"
|
||||
);
|
||||
// The original CMap must still resolve its own CIDs.
|
||||
assert_eq!(primary.lookup(0x0200), Some("\u{0410}".to_string()));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_try_remap_resolves_indirect_subtype() {
|
||||
// /Subtype may be stored as an indirect reference. A genuine CIDFontType2
|
||||
// font must still get the repair, so the guard has to dereference it rather
|
||||
// than treat the unresolved value as "not CIDFontType2".
|
||||
let cmap_content = r#"
|
||||
1 begincodespacerange
|
||||
<0000><FFFF>
|
||||
endcodespacerange
|
||||
1 beginbfrange
|
||||
<0200> <0220> <0410>
|
||||
endbfrange
|
||||
"#;
|
||||
let cmap = ToUnicodeCMap::parse(cmap_content.as_bytes()).unwrap();
|
||||
|
||||
let mut doc = Document::new();
|
||||
let subtype_id = doc.add_object(lopdf::Object::Name(b"CIDFontType2".to_vec()));
|
||||
|
||||
let mut cid_font = lopdf::Dictionary::new();
|
||||
cid_font.set("Subtype", lopdf::Object::Reference(subtype_id));
|
||||
cid_font.set("CIDToGIDMap", lopdf::Object::Name(b"Identity".to_vec()));
|
||||
cid_font.set(
|
||||
"W",
|
||||
lopdf::Object::Array(vec![
|
||||
lopdf::Object::Integer(0),
|
||||
lopdf::Object::Array(vec![lopdf::Object::Integer(500); 34]),
|
||||
]),
|
||||
);
|
||||
let cid_font_id = doc.add_object(cid_font);
|
||||
|
||||
let mut font_dict = lopdf::Dictionary::new();
|
||||
font_dict.set("Encoding", lopdf::Object::Name(b"Identity-H".to_vec()));
|
||||
font_dict.set(
|
||||
"DescendantFonts",
|
||||
lopdf::Object::Array(vec![lopdf::Object::Reference(cid_font_id)]),
|
||||
);
|
||||
|
||||
let (_primary, remapped) = try_remap_subset_cmap(cmap, &font_dict, &doc, 790);
|
||||
assert!(
|
||||
remapped.is_some(),
|
||||
"An indirect /Subtype naming CIDFontType2 must still reach the remap"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
-68
@@ -1,68 +0,0 @@
|
||||
%PDF-1.3
|
||||
%“Œ‹ž ReportLab Generated PDF document (opensource)
|
||||
1 0 obj
|
||||
<<
|
||||
/F1 2 0 R
|
||||
>>
|
||||
endobj
|
||||
2 0 obj
|
||||
<<
|
||||
/BaseFont /Helvetica /Encoding /WinAnsiEncoding /Name /F1 /Subtype /Type1 /Type /Font
|
||||
>>
|
||||
endobj
|
||||
3 0 obj
|
||||
<<
|
||||
/Contents 7 0 R /MediaBox [ 0 0 612 792 ] /Parent 6 0 R /Resources <<
|
||||
/Font 1 0 R /ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
|
||||
>> /Rotate 0 /Trans <<
|
||||
|
||||
>>
|
||||
/Type /Page
|
||||
>>
|
||||
endobj
|
||||
4 0 obj
|
||||
<<
|
||||
/PageMode /UseNone /Pages 6 0 R /Type /Catalog
|
||||
>>
|
||||
endobj
|
||||
5 0 obj
|
||||
<<
|
||||
/Author (anonymous) /CreationDate (D:20260803112923+00'00') /Creator (anonymous) /Keywords () /ModDate (D:20260803112923+00'00') /Producer (ReportLab PDF Library - \(opensource\))
|
||||
/Subject (unspecified) /Title (untitled) /Trapped /False
|
||||
>>
|
||||
endobj
|
||||
6 0 obj
|
||||
<<
|
||||
/Count 1 /Kids [ 3 0 R ] /Type /Pages
|
||||
>>
|
||||
endobj
|
||||
7 0 obj
|
||||
<<
|
||||
/Filter [ /ASCII85Decode /FlateDecode ] /Length 202
|
||||
>>
|
||||
stream
|
||||
GarW05mr9@&;9NOME,dW.,B;'jAjYq0S4Z`*D9aMA;]$5J)A/$3lESen1?F)ZJsa4$4&N%-%cs)#qW5EVhhbPiRDrAV>MC%.spto@CU"ZdipR'TtFiMR_%m*Hm$N%qL7a"ckkp9T/s[N2"Og377mP*M^akb2XQZ@'l*qT(9bVtDb5+)S&Q.#%E)<]Ao`TSk2AE'/E\fn~>endstream
|
||||
endobj
|
||||
xref
|
||||
0 8
|
||||
0000000000 65535 f
|
||||
0000000061 00000 n
|
||||
0000000092 00000 n
|
||||
0000000199 00000 n
|
||||
0000000392 00000 n
|
||||
0000000460 00000 n
|
||||
0000000721 00000 n
|
||||
0000000780 00000 n
|
||||
trailer
|
||||
<<
|
||||
/ID
|
||||
[<6d7ea1213c5974c78613d5d2a08423b5><6d7ea1213c5974c78613d5d2a08423b5>]
|
||||
% ReportLab generated PDF document -- digest (opensource)
|
||||
|
||||
/Info 5 0 R
|
||||
/Root 4 0 R
|
||||
/Size 8
|
||||
>>
|
||||
startxref
|
||||
9072
|
||||
%%EOF
|
||||
-79
File diff suppressed because one or more lines are too long
Binary file not shown.
-1239
File diff suppressed because it is too large
Load Diff
@@ -3926,65 +3926,6 @@ fn encrypted_pdf_decrypts_with_correct_password() {
|
||||
);
|
||||
}
|
||||
|
||||
/// Regression for the #231 review finding: `extract_pages_markdown`'s
|
||||
/// `has_template_image` check must be gated the same way
|
||||
/// `classify_pdf`/`detect_pdf_type` gates it (image_count <= 1, few text
|
||||
/// ops, low alphanumeric diversity) — not treated as sufficient on its
|
||||
/// own. The fixture is a real text page with substantial, richly varied
|
||||
/// body text (>=50 Tj ops) drawn over a full-bleed background image
|
||||
/// (e.g. letterhead/watermark). Before the fix, has_template_image alone
|
||||
/// forced needs_ocr=true and discarded the page's clean markdown; now the
|
||||
/// page must extract normally.
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_does_not_ocr_text_page_with_watermark_image() {
|
||||
let buf = std::fs::read("tests/fixtures/text_page_with_watermark_image.pdf").unwrap();
|
||||
|
||||
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
|
||||
let page = &ext.pages[0];
|
||||
assert!(
|
||||
!page.needs_ocr,
|
||||
"a text page with substantial real text should not be routed to OCR \
|
||||
just because it has a background image"
|
||||
);
|
||||
assert!(
|
||||
page.markdown.contains("watermark"),
|
||||
"expected the page's real body text to be preserved, got: {:?}",
|
||||
page.markdown
|
||||
);
|
||||
}
|
||||
|
||||
/// Regression for the #231 review finding: `extract_pages_markdown` never
|
||||
/// checked `has_vector_text` at all, even though `detect_from_document`'s
|
||||
/// Mixed-type per-page routing always sends vector-outlined-text pages to
|
||||
/// OCR (outlined glyphs can't be extracted as text). A page with massive
|
||||
/// path ops (outlined decorative text) plus a short genuine caption would
|
||||
/// extract that caption cleanly — non-empty, non-garbled — so the
|
||||
/// existing empty/garbage-text checks alone couldn't catch it.
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_ocrs_page_with_vector_outlined_text() {
|
||||
let buf = std::fs::read("tests/fixtures/vector_outlined_text_with_caption.pdf").unwrap();
|
||||
|
||||
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
|
||||
assert!(
|
||||
cls.pages_needing_ocr.contains(&1),
|
||||
"classify_pdf should flag page 1 as needing OCR (vector-outlined text), got: {:?}",
|
||||
cls.pages_needing_ocr
|
||||
);
|
||||
|
||||
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
|
||||
let page = &ext.pages[0];
|
||||
assert!(
|
||||
page.needs_ocr,
|
||||
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
|
||||
);
|
||||
assert!(
|
||||
page.markdown.is_empty(),
|
||||
"a page flagged needs_ocr must not return markdown as if extraction were \
|
||||
trustworthy, got: {:?}",
|
||||
page.markdown
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn pdf_options_debug_redacts_password() {
|
||||
let opts = PdfOptions::new().password("secret123");
|
||||
@@ -3995,60 +3936,3 @@ fn pdf_options_debug_redacts_password() {
|
||||
);
|
||||
assert!(dbg.contains("REDACTED"), "expected redaction marker: {dbg}");
|
||||
}
|
||||
|
||||
/// Regression for #228: a `startxref` pointer corrupted to point at the
|
||||
/// wrong byte offset (a single flipped digit — a real, common writer bug)
|
||||
/// must not make the whole file unprocessable. The real classic xref table
|
||||
/// is still present and findable by scanning for the `xref` keyword; both
|
||||
/// pypdf and pdfium recover the same way. Before this fix, every entry
|
||||
/// point raised "Invalid PDF structure" on a file whose object data was
|
||||
/// otherwise completely intact.
|
||||
#[test]
|
||||
fn test_process_pdf_recovers_corrupted_startxref_pointer() {
|
||||
let result = process_pdf_with_options(
|
||||
"tests/fixtures/broken_startxref_pointer.pdf",
|
||||
PdfOptions::new(),
|
||||
)
|
||||
.expect("a corrupted startxref pointer should be recoverable, like pypdf/pdfium");
|
||||
|
||||
assert_eq!(result.page_count, 1);
|
||||
let md = result.markdown.unwrap_or_default();
|
||||
assert!(
|
||||
md.contains("Order Detail Report by Account") && md.contains("WIDGET ASSEMBLY"),
|
||||
"recovered document should extract its real text, got: {md:?}"
|
||||
);
|
||||
}
|
||||
|
||||
/// Regression for #227: `extract_pages_markdown`'s per-page `needs_ocr`
|
||||
/// must agree with `classify_pdf`/`detect_pdf_type` on the same page. The
|
||||
/// fixture is a full-page raster "scan" with a single line of genuine
|
||||
/// native text drawn over it (a header) — the native text extracts
|
||||
/// perfectly cleanly (no decoding issues, non-empty), so a needs_ocr
|
||||
/// computation based on text-quality signals alone says `false`, while
|
||||
/// detection correctly sees a dominant background image and says the page
|
||||
/// needs OCR. Both must now agree it needs OCR, and the markdown must not
|
||||
/// be returned as if the extraction were trustworthy.
|
||||
#[test]
|
||||
fn test_extract_pages_markdown_agrees_with_classify_on_scan_with_native_header() {
|
||||
let buf = std::fs::read("tests/fixtures/scan_with_native_header_text.pdf").unwrap();
|
||||
|
||||
let cls = pdf_inspector::detector::detect_pdf_type_mem(&buf).expect("fixture should classify");
|
||||
assert!(
|
||||
cls.pages_needing_ocr.contains(&1),
|
||||
"classify_pdf should flag page 1 as needing OCR (image-dominated), got: {:?}",
|
||||
cls.pages_needing_ocr
|
||||
);
|
||||
|
||||
let ext = extract_pages_markdown_mem(&buf, None).expect("fixture should extract");
|
||||
let page = &ext.pages[0];
|
||||
assert!(
|
||||
page.needs_ocr,
|
||||
"extract_pages_markdown must agree with classify_pdf that this page needs OCR"
|
||||
);
|
||||
assert!(
|
||||
page.markdown.is_empty(),
|
||||
"a page flagged needs_ocr must not return markdown as if extraction were \
|
||||
trustworthy, got: {:?}",
|
||||
page.markdown
|
||||
);
|
||||
}
|
||||
|
||||
@@ -1,12 +1,16 @@
|
||||
Reprinted with corrections from *The Bell System Technical Journal,* Vol. 27, pp. 379–423, 623–656, July, October, 1948.
|
||||
|
||||
# A Mathematical Theory of Communication
|
||||
## A Mathematical Theory of Communication
|
||||
|
||||
## By C. E. SHANNON
|
||||
### By C. E. SHANNON
|
||||
|
||||
INTRODUCTION
|
||||
|
||||
THE recent development of various methods of modulation such as PCM and PPM which exchange bandwidth for signal-to-noise ratio has intensified the interest in a general theory of communication. A basis for such a theory is contained in the important papers of Nyquist¹ and Hartley² on this subject. In the present paper we will extend the theory to include a number of new factors, in particular the effect of noise in the channel, and the savings possible due to the statistical structure of the original message and due to the nature of the final destination of the information. The fundamental problem of communication is that of reproducing at one point either exactly or ap- proximately a message selected at another point. Frequently the messages have *meaning*; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one *selected from a set* of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design. If the number of messages in the set is finite then this number or any monotonic function of this number can be regarded as a measure of the information produced when one message is chosen from the set, all choices being equally likely. As was pointed out by Hartley the most natural choice is the logarithmic function. Although this definition must be generalized considerably when we consider the influence of the statistics of the message and when we have a continuous range of messages, we will in all cases use an essentially logarithmic measure. The logarithmic measure is more convenient for various reasons:
|
||||
HE recent development of various methods of modulation such as PCM and PPM which exchange
|
||||
|
||||
# Tbandwidth for signal-to-noise ratio has intensified the interest in a general theory of communication. A
|
||||
|
||||
basis for such a theory is contained in the important papers of Nyquist¹ and Hartley² on this subject. In the present paper we will extend the theory to include a number of new factors, in particular the effect of noise in the channel, and the savings possible due to the statistical structure of the original message and due to the nature of the final destination of the information. The fundamental problem of communication is that of reproducing at one point either exactly or ap- proximately a message selected at another point. Frequently the messages have *meaning*; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one *selected from a set* of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design. If the number of messages in the set is finite then this number or any monotonic function of this number can be regarded as a measure of the information produced when one message is chosen from the set, all choices being equally likely. As was pointed out by Hartley the most natural choice is the logarithmic function. Although this definition must be generalized considerably when we consider the influence of the statistics of the message and when we have a continuous range of messages, we will in all cases use an essentially logarithmic measure. The logarithmic measure is more convenient for various reasons:
|
||||
|
||||
1. It is practically more useful. Parameters of engineering importance such as time, bandwidth, number of relays, etc., tend to vary linearly with the logarithm of the number of possibilities. For example, adding one relay to a group doubles the number of possible states of the relays. It adds 1 to the base 2 logarithm of this number. Doubling the time roughly squares the number of possible messages, or doubles the logarithm, etc.
|
||||
2. It is nearer to our intuitive feeling as to the proper measure. This is closely related to (1) since we in- tuitively measures entities by linear comparison with common standards. One feels, for example, that two punched cards should have twice the capacity of one for information storage, and two identical channels twice the capacity of one for transmitting information.
|
||||
|
||||
Reference in New Issue
Block a user