Compare commits

..
Author SHA1 Message Date
Cursor AgentandAbimael Martell 83d378ebe8 Judge run width against page content, not a fixed extent
Treating any run wider than 14_400 units as malformed penalised valid
large pages: one made entirely of such runs reported no columns at all,
and a mixed page lost the right edge of every long run.

Judge width relative to the page's own content instead. Positions cannot
be inflated by a bogus width, so the spread of the core cluster is a
sound scale: a run wider than that spread plus one page is malformed.
A genuinely large page keeps its genuinely long runs, while a 1e12-wide
run beside ordinary text is still rejected.

Cluster on positions rather than filled intervals, so a bogus width can
no longer merge everything into one cluster, and keep ordinary pages on
an O(n) fast path that skips the sort entirely.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 19:11:40 +00:00
Cursor AgentandAbimael Martell d628a6f7d8 Require geometric evidence before detaching far content
The median-window trim narrowed any content more than one page from the
centre, so a valid large page with a sparse far sidebar lost the sidebar
from its bounds and its text fell into column 0. An item-count minority
rule cannot tell that layout from malformed coordinates.

Group content into clusters separated by more than a whole page of
continuous emptiness, and only drop a cluster that is both detached by
such a void and a small minority of items. Real content does not leave a
gap that large; a stray coordinate sits alone beyond one.

A single run wider than one page is treated as a malformed width, which
also covers the huge-width case the cluster sweep cannot see (such an
item spans everything and leaves no gap).

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 17:42:00 +00:00
Cursor AgentandAbimael Martell 3895b469ed Scale bin width so the histogram spans the whole page
Clamping the bin count alone left anything past MAX_BINS * BIN_WIDTH
(~131k points) outside the histogram, folded into the final bin. A page
wide enough to hit that lost real gutters: with a visible gutter inside
the covered range the XY-cut fallback never runs, so a three-column
layout silently reported two. Derive bin_width from page_width instead,
keeping the same allocation ceiling and degrading only resolution.

Also anchor the trimming median on the same finite left/right items that
bounds() accepts, so a malformed width cannot shift which items count as
strays.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 08:34:40 +00:00
Cursor AgentandAbimael Martell c8b617895f Harden bounds trimming against widths and wide layouts
Check both item edges when trimming: a malformed width at an ordinary
position poisoned x_max just as a malformed position poisoned x_min, so
a huge width still collapsed a two-column page to one region.

Only trim when the far items are a small minority (<=10%). A genuinely
large-format page has content spread across its full width, so it now
keeps its true bounds instead of being reduced to the median cluster.

Correct the MAX_PAGE_EXTENT comment: 14_400 units is the traditional
Acrobat architectural limit, not a format cap. PDF 2.0 sets no page-size
limit and UserUnit scales physical size, so this is a heuristic.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 03:46:28 +00:00
Cursor AgentandAbimael Martell d79bd936ea Trim far-outlier coordinates from page bounds
Gutter margins, spanning-item width and the XY-cut margin are all
fractions of page_width, so a single far-but-finite item (x=50_000 is
enough) set the scale for the whole page: real gutters fell inside the
rejected margin band and a genuine two-column page collapsed to one
region. When the span exceeds one legal page (14_400 units), re-derive
the bounds from items clustered around the median x. Outliers keep their
text because column assignment buckets by nearest overlap.

The MAX_BINS ceiling stays as an allocation bound that does not depend
on this heuristic.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-10 00:15:18 +00:00
Cursor AgentandAbimael Martell 6fa36be0a0 Exclude non-finite coordinates from page bounds
Items at NaN/inf positions are now skipped when folding the page bounds,
so a malformed coordinate can no longer escape as a ColumnRegion
boundary, and an all-non-finite page returns no columns. Bad items are
dropped individually rather than failing the page, so one stray glyph
does not disable column detection.

Addresses review feedback on the finite-width guard.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 23:43:51 +00:00
Cursor AgentandAbimael Martell 0eda2eb778 Bound column-detection histogram size
Derive the projection histogram from a clamped bin count and skip
non-finite page widths. Extreme or malformed text-item coordinates
(from the content-stream text matrix) could otherwise drive a very
large allocation. 65,536 bins is ~9x the largest legal page, so real
layouts are unaffected. Adds regression tests.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 17:53:06 +00:00
69039f2728 Fix char-boundary panic in hex_to_unicode_string (#320)
Use hex.get(i..i+2) instead of &hex[i..i+2] so a non-hex, non-ASCII
destination in a /ToUnicode CMap can no longer trigger a UTF-8
char-boundary panic. An even byte length does not guarantee the byte
offset falls on a char boundary; get() returns None on a non-boundary
or out-of-range index, folding cleanly into the existing flow.

Add regression tests covering a multi-byte destination char and a
replacement-char byte.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 08:21:31 -07:00
3cca6446bd fix(glyph_names): handle non-ASCII input in uniXXXX glyph name parsing (#321)
Use str::get instead of a byte-length check plus slice when parsing the
uniXXXX glyph-name form. The byte-length guard only proved the index was
in bounds, not on a UTF-8 char boundary, so a glyph name containing
non-ASCII bytes could cause a slice on a non-boundary index. Switch to a
checked slice that folds into the existing Option flow, and add tests.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-09 08:21:22 -07:00
493fed498e fix(links): prevent stack-overflow DoS from AcroForm /Kids self-cycle (#314)
* fix(links): guard AcroForm /Kids traversal against cycles and huge trees

A crafted PDF whose AcroForm field lists itself (or another ancestor) in
/Kids caused walk_form_fields to recurse indefinitely, overflowing the
stack and aborting pdf2md (exit 134) — an application-level DoS from a
~730-byte input.

Track visited field object IDs to break /Kids cycles, and cap total
field-node traversal at 100k nodes to bound pathologically large trees.

Adds regression tests for self-cycle and mutual-cycle field graphs.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): cap AcroForm /Kids recursion depth to stop deep-chain overflow

The visited-set guard stops cyclic /Kids graphs, but a long *acyclic*
chain of distinct fields still recurses to the chain length and overflows
the stack (a ~1.6MB PDF with 20k linked fields aborts pdf2md, exit 134)
before the 100k node budget is reached.

Add an explicit recursion depth cap (100 levels — far above any legitimate
form hierarchy) so stack usage is bounded independently of node count.

Adds a deep-acyclic-chain regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): enforce form-field node budget before insertion

The node-budget guard inserted each field ID into the visited set before
checking the budget, so the check triggered an early return but never
actually capped the set. A field with a huge /Kids array kept inserting
post-budget IDs, letting visited (memory and work) grow with the crafted
input rather than stopping at MAX_FORM_FIELD_NODES.

Check depth and budget before inserting, so visited can never exceed the
cap. Adds a wide-tree regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): stop /Fields and /Kids iteration once node budget is spent

Checking the budget before insertion capped the visited set, but callers
still iterated every remaining entry of a wide /Fields or /Kids array
after the budget was exhausted — each walk returned immediately, yet the
O(N) sibling iteration let a single multi-million-entry array burn
extraction CPU unbounded. Break out of both the top-level and recursive
loops once visited reaches the cap, making the budget a true
traversal-work cap. Adds a top-level wide-/Fields regression test.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): charge examined entries against the field-node budget

The budget counted only distinct visited nodes, so /Fields or /Kids
arrays full of invalid (non-reference) or duplicate entries never grew
visited and ran to completion regardless of size — the node budget did
not actually cap traversal work.

Introduce FieldWalkBudget tracking both visited nodes and total entries
examined; charge every array entry (valid, invalid, or duplicate) and
stop once either hits MAX_FORM_FIELD_NODES. Adds a regression test with a
huge /Kids array of duplicate + null entries.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* fix(links): iterate /Fields and /Kids arrays by borrow, not clone

Both arrays were cloned in full before the budget check, so a crafted
oversized /Fields or /Kids array forced an O(n) allocation and copy
regardless of the cap. resolve_array already returns a borrow tied to the
document and the walker only needs a shared &Document, so iterate the
borrowed arrays directly — the early break now bounds how many entries
are even touched, before any per-array allocation.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

* docs(links): correct wide-array test comments to match range assertions

The two wide-array tests assert item counts within a range near the
budget, not an exact value (charging entries in the entry guard shifts
the boundary by one or two). Fix the stale comments that claimed exact
counts.

Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Abimael Martell <abimaelmartell@users.noreply.github.com>
2026-08-08 23:43:33 -07:00
6 changed files with 704 additions and 555 deletions
+327 -17
View File
@@ -41,15 +41,110 @@ pub(crate) fn detect_columns(
}
debug!("page {}: detect_columns: {} items", page, page_items.len());
// Find page bounds
let x_min = page_items.iter().map(|i| i.x).fold(f32::INFINITY, f32::min);
let x_max = page_items
.iter()
.map(|i| i.x + effective_width(i))
.fold(f32::NEG_INFINITY, f32::max);
// The width of one ordinary page, used three ways below: as the largest
// credible width for a single text run, as the size of empty gap that marks
// content as detached, and as the span past which those checks run at all.
// This is a heuristic, not a format rule: PDF 2.0 sets no page-size limit,
// and since PDF 1.6 `UserUnit` scales a page's physical size independently
// of its coordinates. 14_400 units (200in at the default 1/72in unit) is
// the traditional Acrobat architectural limit, which makes it a reasonable
// "wider than any ordinary page" mark in coordinate space.
const MAX_PAGE_EXTENT: f32 = 14_400.0;
// A detached cluster is only dropped if it also holds a small minority of
// the items, so a genuine two-part layout keeps its full bounds even when
// the halves are far apart.
const MAX_TRIM_FRACTION: f32 = 0.10;
// Position and width of each item, skipping only non-finite geometry.
let finite_span = |i: &&TextItem| -> Option<(f32, f32)> {
let (left, width) = (i.x, effective_width(i));
(left.is_finite() && (left + width).is_finite()).then_some((left, width))
};
let (min_left, max_right, total) = page_items.iter().filter_map(finite_span).fold(
(f32::INFINITY, f32::NEG_INFINITY, 0usize),
|(lo, hi, n), (left, width)| (lo.min(left), hi.max(left + width), n + 1),
);
// No item had usable geometry, so there is no layout to report.
if total == 0 {
return vec![];
}
// Every threshold below (gutter margins, spanning-item width, the XY-cut
// margin) is a fraction of the page width, so a far item can set the scale
// for the whole page and shrink the effective detection window to a
// rounding error — real gutters then fall inside the margin band and a
// genuine multi-column page collapses to one region.
//
// Anything inside one page extent is ordinary, so the common case keeps the
// plain bounds and skips the work below entirely.
let (x_min, x_max) = if max_right - min_left <= MAX_PAGE_EXTENT {
(min_left, max_right)
} else {
// Discarding content needs positive evidence that it is not part of the
// layout, because a count-based rule alone cannot tell a stray from a
// sparse far sidebar. The evidence is geometric: positions are grouped
// into clusters separated by more than a whole page of continuous
// emptiness. Real content, however sparse, does not leave a void that
// large; a malformed coordinate sits alone beyond one.
let mut spans: Vec<(f32, f32)> = page_items.iter().filter_map(finite_span).collect();
spans.sort_by(|a, b| a.0.total_cmp(&b.0));
let mut core: Option<std::ops::Range<usize>> = None;
let mut start = 0usize;
for i in 1..=spans.len() {
if i < spans.len() && spans[i].0 - spans[i - 1].0 <= MAX_PAGE_EXTENT {
continue;
}
if core.as_ref().is_none_or(|best| i - start > best.len()) {
core = Some(start..i);
}
start = i;
}
let mut core = core.unwrap_or(0..spans.len());
// Only drop the detached clusters when they are a small minority, so a
// genuine two-part layout keeps its full bounds.
let dropped = spans.len() - core.len();
if dropped as f32 > spans.len() as f32 * MAX_TRIM_FRACTION {
core = 0..spans.len();
}
let core = &spans[core];
// Positions cannot be inflated by a bogus width, so the spread of the
// content is a sound scale for judging one. A run much wider than the
// page's own content is a malformed width — the test is relative, so a
// genuinely large page keeps its genuinely long runs.
let (lo, widest_left) = (core[0].0, core[core.len() - 1].0);
let max_run_width = (widest_left - lo) + MAX_PAGE_EXTENT;
let hi = core
.iter()
.filter(|&&(_, width)| width <= max_run_width)
.map(|&(left, width)| left + width)
.fold(widest_left, f32::max);
if lo != min_left || hi != max_right {
debug!(
"page {page}: bounds {min_left}..{max_right} exceed one page; \
dropped {dropped}/{} detached item(s), using {lo}..{hi}",
spans.len()
);
}
(lo, hi)
};
// Hard ceiling on the histogram size, independent of the trimming above:
// the bounds are attacker-influenced, so an unclamped
// `page_width / BIN_WIDTH` lets a crafted PDF force an arbitrarily large
// `vec![0u32; num_bins]` allocation. 65_536 bins covers ~128k points at
// BIN_WIDTH 2.0 — roughly 9x the largest legal page — so this never binds
// on a real layout. Kept as a bound that does not depend on the outlier
// heuristic staying correct.
const MAX_BINS: usize = 65_536;
let page_width = x_max - x_min;
if page_width < 200.0 {
if !page_width.is_finite() || page_width < 200.0 {
return vec![ColumnRegion { x_min, x_max }];
}
@@ -57,13 +152,20 @@ pub(crate) fn detect_columns(
return vec![ColumnRegion { x_min, x_max }];
}
// Widen the bins rather than dropping the tail of the page. Clamping the
// count alone would leave anything past MAX_BINS * BIN_WIDTH outside the
// histogram, folded into the last bin, which places gutters at the wrong
// coordinates. Scaling keeps full coverage under the same allocation
// ceiling; only the resolution degrades, and only beyond ~131k points.
let bin_width = BIN_WIDTH.max(page_width / MAX_BINS as f32);
// Build occupancy histogram.
// Exclude items wider than 60% of page width — these are spanning items
// (titles, full-width paragraphs) that would fill the gutter and prevent
// detection of partial-page column layouts (e.g. two-column abstracts on
// a page that also has single-column introduction text).
let wide_threshold = page_width * 0.6;
let num_bins = ((page_width / BIN_WIDTH).ceil() as usize).max(1);
let num_bins = ((page_width / bin_width).ceil() as usize).clamp(1, MAX_BINS);
let mut histogram = vec![0u32; num_bins];
for item in &page_items {
@@ -71,8 +173,8 @@ pub(crate) fn detect_columns(
if w > wide_threshold {
continue;
}
let left = ((item.x - x_min) / BIN_WIDTH).floor() as usize;
let right = (((item.x + w) - x_min) / BIN_WIDTH).ceil() as usize;
let left = ((item.x - x_min) / bin_width).floor() as usize;
let right = (((item.x + w) - x_min) / bin_width).ceil() as usize;
let left = left.min(num_bins);
let right = right.min(num_bins);
for count in histogram.iter_mut().take(right).skip(left) {
@@ -109,12 +211,12 @@ pub(crate) fn detect_columns(
let valleys: Vec<(usize, usize)> = valleys
.into_iter()
.filter(|&(start, end)| {
let width_pts = (end - start) as f32 * BIN_WIDTH;
let width_pts = (end - start) as f32 * bin_width;
if width_pts < MIN_GUTTER_WIDTH {
return false;
}
// Valley center must not be within 5% of page edges
let center_pts = ((start + end) as f32 / 2.0) * BIN_WIDTH;
let center_pts = ((start + end) as f32 / 2.0) * bin_width;
center_pts > margin_threshold && center_pts < (page_width - margin_threshold)
})
.collect();
@@ -132,7 +234,7 @@ pub(crate) fn detect_columns(
&histogram,
num_bins,
x_min,
BIN_WIDTH,
bin_width,
page_width,
margin_threshold,
);
@@ -141,7 +243,7 @@ pub(crate) fn detect_columns(
&rel_valleys,
&page_items,
x_min,
BIN_WIDTH,
bin_width,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
@@ -182,7 +284,7 @@ pub(crate) fn detect_columns(
&valleys,
&page_items,
x_min,
BIN_WIDTH,
bin_width,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
@@ -196,7 +298,7 @@ pub(crate) fn detect_columns(
&valleys,
&page_items,
x_min,
BIN_WIDTH,
bin_width,
x_max,
MIN_ITEMS_PER_COLUMN,
MIN_VERTICAL_SPAN_RATIO,
@@ -1827,7 +1929,7 @@ fn split_column_stragglers(lines: Vec<TextLine>) -> (Vec<TextLine>, Vec<TextLine
.unwrap();
let (cs, ce) = segments[core_seg];
let mut core = Vec::with_capacity(ce - cs);
let mut core = Vec::with_capacity(ce.saturating_sub(cs));
let mut stragglers = Vec::new();
for (i, line) in lines.into_iter().enumerate() {
if i >= cs && i < ce {
@@ -2533,6 +2635,214 @@ mod tests {
);
}
#[test]
fn extreme_far_coordinate_does_not_allocate_unboundedly() {
// A crafted PDF can place a text run at an arbitrary coordinate via the
// text matrix. The derived page width must not drive an unbounded
// histogram allocation (previously `page_width / BIN_WIDTH` bins with no
// upper bound would try to reserve terabytes and abort the process).
let mut items = Vec::new();
for i in 0..24 {
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
}
// Item placed 1e12 points away — 5e11 bins if left unclamped.
items.push(make_item(1, 1e12, 700.0, "Z"));
// Must return without aborting; content is preserved as a single region.
let cols = detect_columns(&items, 1, false);
assert!(!cols.is_empty());
}
#[test]
fn non_finite_coordinates_never_leak_into_region_bounds() {
// An inf/NaN coordinate must not escape as a column boundary: callers
// treat these as page/column edges.
for bad_x in [f32::INFINITY, f32::NEG_INFINITY, f32::NAN] {
let mut items = Vec::new();
for i in 0..24 {
items.push(make_item(1, i as f32 * 10.0, 700.0 - i as f32 * 5.0, "A"));
}
items.push(make_item(1, bad_x, 700.0, "Z"));
for col in detect_columns(&items, 1, false) {
assert!(
col.x_min.is_finite() && col.x_max.is_finite(),
"bad_x {bad_x} leaked bounds {}..{}",
col.x_min,
col.x_max
);
}
}
}
#[test]
fn all_non_finite_coordinates_yield_no_columns() {
let items: Vec<TextItem> = (0..24)
.map(|i| make_item(1, f32::NAN, 700.0 - i as f32 * 5.0, "A"))
.collect();
assert!(detect_columns(&items, 1, false).is_empty());
}
#[test]
fn one_bad_item_does_not_disable_column_detection() {
// A single stray item should not collapse a clean two-column page to
// one region. Every gutter threshold is a fraction of the page width,
// so an untrimmed outlier pushes real gutters inside the rejected
// margin band. A malformed *width* at an ordinary position poisons the
// bounds just as a malformed position does.
for (label, bad_x, bad_width) in [
("nan position", f32::NAN, 0.0),
("inf position", f32::INFINITY, 0.0),
("far position", 50_000.0, 0.0),
("very far position", 1e12, 0.0),
("huge width", 100.0, 1e12),
("inf width", 100.0, f32::INFINITY),
] {
let mut items = Vec::new();
items.extend(fill_zone(1, 30.0, 280.0, 750.0, 50.0));
items.extend(fill_zone(1, 320.0, 570.0, 750.0, 50.0));
let mut bad = make_item(1, bad_x, 400.0, "Z");
bad.width = bad_width;
items.push(bad);
let cols = detect_columns(&items, 1, false);
assert_eq!(
cols.len(),
2,
"{label}: expected 2 columns, got {}",
cols.len()
);
for col in &cols {
assert!(
col.x_max - col.x_min <= MAX_PAGE_EXTENT_FOR_TEST,
"{label}: region {}..{} exceeds one page",
col.x_min,
col.x_max
);
}
}
}
/// Mirrors `MAX_PAGE_EXTENT` in `detect_columns`.
const MAX_PAGE_EXTENT_FOR_TEST: f32 = 14_400.0;
#[test]
fn very_wide_page_keeps_full_histogram_coverage() {
// Beyond MAX_BINS * BIN_WIDTH (~131k points) the bins must widen rather
// than stop covering the page. Three zones: the first gutter is inside
// the old coverage limit, the second is past it. Because the first
// gutter is found, the XY-cut fallback never runs, so a truncated
// histogram silently reports two columns instead of three.
let mut items = Vec::new();
items.extend(fill_zone(1, 0.0, 60_000.0, 750.0, 700.0));
items.extend(fill_zone(1, 70_000.0, 140_000.0, 750.0, 700.0));
items.extend(fill_zone(1, 160_000.0, 200_000.0, 750.0, 700.0));
let cols = detect_columns(&items, 1, false);
assert_eq!(
cols.len(),
3,
"Expected 3 columns across a 200k-wide page, got {}",
cols.len()
);
assert!(
(140_000.0..=160_000.0).contains(&cols[1].x_max),
"second gutter at {}, expected inside the real 140k..160k gap",
cols[1].x_max
);
}
#[test]
fn large_page_with_legitimately_long_runs_is_kept() {
// On a very large page, individual runs can exceed one ordinary page's
// width. They are real content, so they must not be judged malformed:
// the page keeps its columns and its full right edge.
let mut items = Vec::new();
for row in 0..30 {
let y = 750.0 - row as f32 * 14.0;
let mut left = make_item(1, 0.0, y, "Left run");
left.width = 20_000.0;
let mut right = make_item(1, 25_000.0, y, "Right run");
right.width = 20_000.0;
items.extend([left, right]);
}
let cols = detect_columns(&items, 1, false);
assert!(
!cols.is_empty(),
"a page of long-but-valid runs must still report a layout"
);
let right_edge = cols
.iter()
.map(|c| c.x_max)
.fold(f32::NEG_INFINITY, f32::max);
assert!(
right_edge > 44_000.0,
"long runs were treated as malformed: right edge {right_edge}, expected ~45_000"
);
}
#[test]
fn sparse_far_sidebar_on_a_large_page_is_kept() {
// A large-format page with a thin, sparsely-populated sidebar far from
// the main block. The sidebar is a small minority of the items, so an
// item-count rule alone would discard it — but nothing about its
// geometry says it is invalid, so its bounds must survive.
let mut items = Vec::new();
items.extend(fill_zone(1, 0.0, 12_000.0, 750.0, 500.0));
for i in 0..12 {
items.push(make_item(1, 24_000.0, 750.0 - i as f32 * 14.0, "Sidebar"));
}
let cols = detect_columns(&items, 1, false);
let right_edge = cols
.iter()
.map(|c| c.x_max)
.fold(f32::NEG_INFINITY, f32::max);
assert!(
right_edge > 24_000.0,
"sidebar was trimmed away: right edge {right_edge}, expected >24_000"
);
}
#[test]
fn genuinely_wide_layout_keeps_its_true_bounds() {
// A large-format page whose content really is spread beyond one
// ordinary page must not be trimmed to the median cluster: its far
// items are the majority, not strays.
let mut items = Vec::new();
items.extend(fill_zone(1, 100.0, 20_000.0, 750.0, 600.0));
items.extend(fill_zone(1, 22_000.0, 40_000.0, 750.0, 600.0));
let cols = detect_columns(&items, 1, false);
let widest = cols
.iter()
.map(|c| c.x_max)
.fold(f32::NEG_INFINITY, f32::max);
assert!(
widest > 35_000.0,
"wide layout was trimmed: right edge {widest}, expected ~40_000"
);
}
#[test]
fn oversized_but_legal_page_is_not_trimmed() {
// A wide-format page well inside the 14_400pt spec limit must keep its
// real bounds — outlier trimming is only for spans beyond a legal page.
let mut items = Vec::new();
items.extend(fill_zone(1, 100.0, 4_000.0, 750.0, 400.0));
items.extend(fill_zone(1, 4_400.0, 8_000.0, 750.0, 400.0));
let cols = detect_columns(&items, 1, false);
assert_eq!(cols.len(), 2, "Expected 2 columns, got {}", cols.len());
assert!(
cols[1].x_max > 7_000.0,
"right column should keep its true extent, got {}",
cols[1].x_max
);
}
#[test]
fn two_column_regression_guard() {
// Standard 2-column layout with clear gutter at center
+311 -5
View File
@@ -2,11 +2,51 @@
use crate::types::{ItemType, TextItem};
use lopdf::{Document, Object, ObjectId};
use std::collections::HashMap;
use std::collections::{HashMap, HashSet};
use super::fonts::{resolve_array, resolve_dict};
use super::get_number;
/// Upper bound on the number of form-field nodes visited during a single
/// `extract_form_fields` pass. A crafted PDF can chain thousands of distinct
/// `/Kids` fields to blow the stack even without an outright reference cycle,
/// so we cap total traversal work in addition to detecting cycles.
const MAX_FORM_FIELD_NODES: usize = 100_000;
/// Upper bound on `/Kids` recursion depth. Real AcroForm hierarchies are only
/// a few levels deep (fields → child fields → widgets); a crafted PDF can chain
/// tens of thousands of distinct fields into a linear `/Kids` list that would
/// overflow the stack via depth-first recursion long before the node budget is
/// reached. This depth cap bounds the stack independently of total node count.
const MAX_FORM_FIELD_DEPTH: usize = 100;
/// Traversal budget for the AcroForm field walk. Bounds both the number of
/// distinct nodes visited *and* the total number of `/Fields`/`/Kids` entries
/// examined.
///
/// Counting `visited` alone is not enough: invalid entries (non-references) and
/// duplicate references never grow `visited`, so an oversized array full of them
/// would iterate to completion no matter how large. Charging every examined
/// entry against the same budget makes it a real cap on traversal work.
pub(crate) struct FieldWalkBudget {
visited: HashSet<ObjectId>,
examined: usize,
}
impl FieldWalkBudget {
fn new() -> Self {
Self {
visited: HashSet::new(),
examined: 0,
}
}
/// True once the budget is spent; callers must stop iterating and recursing.
fn exhausted(&self) -> bool {
self.visited.len() >= MAX_FORM_FIELD_NODES || self.examined >= MAX_FORM_FIELD_NODES
}
}
pub fn extract_page_links(doc: &Document, page_id: ObjectId, page_num: u32) -> Vec<TextItem> {
let mut links = Vec::new();
@@ -146,9 +186,12 @@ pub(crate) fn extract_form_fields(
Err(_) => return items,
};
// Borrow the array rather than cloning it: a crafted `/Fields` can be huge,
// and cloning would pay an O(n) allocation/copy before the budget check
// below can stop the work.
let fields = match acroform.get(b"Fields") {
Ok(obj) => match resolve_array(doc, obj) {
Some(arr) => arr.clone(),
Some(arr) => arr,
None => return items,
},
Err(_) => return items,
@@ -158,7 +201,19 @@ pub(crate) fn extract_form_fields(
}
let annotation_pages = annotation_page_map(doc, page_map);
for field_obj in &fields {
// Bound the walk so a crafted PDF cannot send us into unbounded recursion
// via a `/Kids` cycle, a deep chain, or an oversized array of invalid or
// duplicate entries.
let mut budget = FieldWalkBudget::new();
for field_obj in fields {
// Stop once the budget is spent so a `/Fields` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op. Charge
// every entry (including invalid ones) against the budget.
if budget.exhausted() {
break;
}
budget.examined += 1;
if let Ok(field_ref) = field_obj.as_reference() {
walk_form_fields(
doc,
@@ -168,6 +223,8 @@ pub(crate) fn extract_form_fields(
page_map,
&annotation_pages,
&mut items,
&mut budget,
0,
);
}
}
@@ -202,6 +259,7 @@ fn annotation_page_map(
}
/// Recursively walk the form field tree, extracting leaf field values.
#[allow(clippy::too_many_arguments)]
pub(crate) fn walk_form_fields(
doc: &Document,
field_id: ObjectId,
@@ -210,7 +268,22 @@ pub(crate) fn walk_form_fields(
page_map: &HashMap<ObjectId, u32>,
annotation_pages: &HashMap<ObjectId, u32>,
items: &mut Vec<TextItem>,
budget: &mut FieldWalkBudget,
depth: usize,
) {
// Guard against `/Kids` cycles and pathologically large field trees.
// Exceeding the depth cap means the chain is too deep to be a legitimate
// form (and would overflow the stack); an exhausted budget means the tree is
// too large. Both checks run *before* inserting so the visited set can never
// grow past the budget.
if depth > MAX_FORM_FIELD_DEPTH || budget.exhausted() {
return;
}
// Revisiting an object ID means we hit a `/Kids` cycle.
if !budget.visited.insert(field_id) {
return;
}
let field_dict = match doc.get_dictionary(field_id) {
Ok(d) => d,
Err(_) => return,
@@ -241,9 +314,19 @@ pub(crate) fn walk_form_fields(
// Check for /Kids — if present, recurse into children
if let Ok(kids_obj) = field_dict.get(b"Kids") {
// Iterate the borrowed array directly — cloning a crafted, oversized
// `/Kids` would allocate and copy every entry before the budget check
// below could stop the work.
if let Some(kids) = resolve_array(doc, kids_obj) {
let kids = kids.clone();
for kid in &kids {
for kid in kids {
// Stop once the budget is spent so a `/Kids` array wider than the
// budget can't burn CPU iterating entries whose walk would no-op.
// Charge every entry (including invalid/duplicate ones) against
// the budget so this is a true traversal-work cap.
if budget.exhausted() {
break;
}
budget.examined += 1;
if let Ok(kid_ref) = kid.as_reference() {
walk_form_fields(
doc,
@@ -253,6 +336,8 @@ pub(crate) fn walk_form_fields(
page_map,
annotation_pages,
items,
budget,
depth + 1,
);
}
}
@@ -411,4 +496,225 @@ mod tests {
assert_eq!(items[0].page, 2);
assert_eq!(items[0].text, "customer: Alice");
}
#[test]
fn kids_self_cycle_does_not_overflow_stack() {
// A crafted AcroForm field that lists itself in `/Kids` must not send
// the traversal into unbounded recursion.
let mut doc = Document::new();
let field_id = doc.new_object_id();
doc.set_object(
field_id,
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("loop"),
"Kids" => vec![Object::Reference(field_id)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
// Completes (rather than overflowing the stack) and yields no items.
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn kids_mutual_cycle_terminates() {
// Two fields that reference each other via `/Kids` form a cycle that
// must also terminate.
let mut doc = Document::new();
let field_a = doc.new_object_id();
let field_b = doc.new_object_id();
doc.set_object(
field_a,
dictionary! {
"T" => Object::string_literal("a"),
"Kids" => vec![Object::Reference(field_b)],
},
);
doc.set_object(
field_b,
dictionary! {
"T" => Object::string_literal("b"),
"Kids" => vec![Object::Reference(field_a)],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(field_a)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn deep_acyclic_kids_chain_does_not_overflow_stack() {
// A long chain of *distinct* fields (no cycle) must also terminate:
// the visited set alone would still recurse to the chain length, so
// the depth cap is what prevents a stack overflow here.
let mut doc = Document::new();
let n = MAX_FORM_FIELD_DEPTH * 500;
let ids: Vec<ObjectId> = (0..=n).map(|_| doc.new_object_id()).collect();
for i in 0..n {
doc.set_object(
ids[i],
dictionary! {
"FT" => "Tx",
"Kids" => vec![Object::Reference(ids[i + 1])],
},
);
}
// Leaf carries a value; it sits far below the depth cap so it is never
// reached, proving traversal stops early rather than crashing.
doc.set_object(
ids[n],
dictionary! {
"FT" => "Tx",
"T" => Object::string_literal("leaf"),
"V" => Object::string_literal("x"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(ids[0])],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.is_empty());
}
#[test]
fn wide_tree_traversal_stops_at_node_budget() {
// A single field with a `/Kids` array wider than the node budget must
// stop traversal at the cap rather than growing `visited` (and the work)
// without bound. Each processed leaf emits one item, so the item count
// is bounded by the budget and reaches right up to it (a couple of
// slots go to the root and the boundary node charged against the cap).
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let kids: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
// Extraction stops at the budget: bounded above by the cap, and it gets
// right up to it (allowing a small delta for the root/boundary nodes
// charged against the budget).
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn wide_top_level_fields_stop_at_node_budget() {
// A top-level `/Fields` array wider than the budget must also stop at
// the cap: the item count is bounded by the budget and reaches right up
// to it.
let mut doc = Document::new();
let fanout = MAX_FORM_FIELD_NODES + 50;
let leaf_ids: Vec<ObjectId> = (0..fanout).map(|_| doc.new_object_id()).collect();
for &leaf in &leaf_ids {
doc.set_object(
leaf,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
}
let fields: Vec<Object> = leaf_ids.iter().map(|&id| Object::Reference(id)).collect();
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => fields,
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert!(items.len() <= MAX_FORM_FIELD_NODES);
assert!(items.len() >= MAX_FORM_FIELD_NODES - 3);
}
#[test]
fn duplicate_and_invalid_kids_entries_stop_at_budget() {
// Duplicate references and non-reference junk never grow `visited`, so
// without charging examined entries against the budget an oversized
// array of them would iterate to completion. The walk must still
// terminate and extract the single real leaf exactly once.
let mut doc = Document::new();
let leaf_id = doc.new_object_id();
doc.set_object(
leaf_id,
dictionary! {
"FT" => "Tx",
"V" => Object::string_literal("v"),
"Rect" => vec![10.into(), 20.into(), 110.into(), 40.into()],
},
);
// A `/Kids` array far wider than the budget: half duplicate references
// to the same leaf, half invalid (null) entries.
let mut kids: Vec<Object> = Vec::new();
for i in 0..(MAX_FORM_FIELD_NODES * 2) {
if i % 2 == 0 {
kids.push(Object::Reference(leaf_id));
} else {
kids.push(Object::Null);
}
}
let root_id = doc.add_object(dictionary! {
"T" => Object::string_literal("root"),
"Kids" => kids,
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"AcroForm" => dictionary! {
"Fields" => vec![Object::Reference(root_id)],
},
});
doc.trailer.set("Root", Object::Reference(catalog_id));
let page_map = HashMap::new();
let items = extract_form_fields(&doc, &page_map);
assert_eq!(items.len(), 1);
}
}
+40 -3
View File
@@ -4566,9 +4566,13 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
}
}
// Try to parse uniXXXX format
if name.starts_with("uni") && name.len() >= 7 {
if let Ok(code) = u32::from_str_radix(&name[3..7], 16) {
// Try to parse uniXXXX format.
// Use `get` rather than a byte-length check + slice: `name` can contain
// non-ASCII bytes (e.g. U+FFFD from lossy UTF-8 decoding of an attacker
// controlled /Differences name), so byte index 7 may not be a char
// boundary and `&name[3..7]` would panic.
if let Some(hex) = name.strip_prefix("uni").and_then(|rest| rest.get(..4)) {
if let Ok(code) = u32::from_str_radix(hex, 16) {
// Strip PUA F000 offset: uniF0XX → U+00XX (Windows Symbol encoding convention)
let code = if (0xF000..=0xF0FF).contains(&code) {
code - 0xF000
@@ -4588,3 +4592,36 @@ pub fn glyph_to_char(name: &str) -> Option<char> {
None
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn uni_hex_parsing() {
assert_eq!(glyph_to_char("uni0041"), Some('A'));
assert_eq!(glyph_to_char("uni00e9"), Some('\u{00e9}'));
// PUA F0xx symbol-encoding offset is stripped.
assert_eq!(glyph_to_char("uniF041"), Some('A'));
}
#[test]
fn u_hex_parsing() {
assert_eq!(glyph_to_char("u0041"), Some('A'));
assert_eq!(glyph_to_char("u1F600"), Some('\u{1F600}'));
}
#[test]
fn non_ascii_uni_name_does_not_panic() {
// A crafted /Differences name like `/uni#80#80#80#80` decodes via
// from_utf8_lossy into "uni" followed by four U+FFFD replacements.
// Byte index 7 lands mid-character, so a naive `&name[3..7]` slice
// would panic. It must be handled gracefully instead.
let crafted = format!("uni{0}{0}{0}{0}", '\u{FFFD}');
assert_eq!(glyph_to_char(&crafted), None);
// Assorted non-ASCII bytes right after the "uni" prefix.
assert_eq!(glyph_to_char("uni\u{FFFD}bc"), None);
assert_eq!(glyph_to_char("uni\u{00e9}00"), None);
}
}
-526
View File
@@ -150,61 +150,6 @@ pub(crate) fn merge_heading_lines(
/// Merge drop caps with the appropriate line.
/// A drop cap is a single large letter at the start of a paragraph.
/// Due to PDF coordinate sorting, the drop cap may appear AFTER the line it belongs to.
/// True when the text ends a sentence, as opposed to merely ending in a
/// period. An abbreviation or list marker ("e.g.", "Fig.", "Mr.", "1.")
/// closes with a period mid-sentence, so treating those as paragraph
/// boundaries would let a drop cap be prepended to a continuation.
fn ends_sentence(text: &str) -> bool {
let t = text.trim_end();
if t.ends_with(['!', '?']) {
return true;
}
let Some(stripped) = t.strip_suffix('.') else {
return false;
};
let last = stripped.split_whitespace().next_back().unwrap_or("");
if last.is_empty() {
return false;
}
// Abbreviations carry an internal period between very short segments
// ("e.g.", "i.e.", "U.S."). Domains and decimals have the same shape but
// longer or numeric segments ("example.com.", "3.14."), and those end
// sentences perfectly well, so require every segment to be short and
// alphabetic before reading the internal period as an abbreviation.
if last.contains('.')
&& last
.split('.')
.filter(|seg| !seg.is_empty())
// Characters, not bytes: a two-letter non-ASCII abbreviation
// ("т.е.", "ú.d.") measures four or more bytes and would
// otherwise be read as a completed sentence.
.all(|seg| seg.chars().count() <= 2 && seg.chars().all(char::is_alphabetic))
{
return false;
}
// Enumerators stand alone on their line ("1.", "ii.", "IV."). A number
// or numeral in the tail of a sentence does not — "published in 2020.",
// "He scored 5." and "after World War II." all end sentences, and
// treating them as markers would block a legitimate drop-cap merge.
if stripped.split_whitespace().count() == 1 {
let is_numeric = last.chars().all(|c| c.is_ascii_digit());
let is_roman = last
.chars()
.all(|c| matches!(c.to_ascii_uppercase(), 'I' | 'V' | 'X' | 'L' | 'C'));
if is_numeric || is_roman {
return false;
}
}
const ABBREVIATIONS: &[&str] = &[
"Fig", "No", "Mr", "Mrs", "Ms", "Dr", "St", "vs", "etc", "al", "Ed", "Eq", "Ch", "pp",
"Vol", "cf", "Prof", "Inc", "Ltd", "Jr", "Sr",
];
!ABBREVIATIONS.iter().any(|a| a.eq_ignore_ascii_case(last))
}
pub(crate) fn merge_drop_caps(lines: Vec<TextLine>, base_size: f32) -> Vec<TextLine> {
let mut result: Vec<TextLine> = Vec::with_capacity(lines.len());
@@ -224,169 +169,6 @@ pub(crate) fn merge_drop_caps(lines: Vec<TextLine>, base_size: f32) -> Vec<TextL
.map(|c| c.is_uppercase())
.unwrap_or(false);
// Embedded drop cap: a two-line cap's baseline aligns with the
// paragraph's SECOND line, so Y-grouping puts the glyph at the start
// of that line rather than on a line of its own. Left there it
// surfaces mid-sentence once the paragraph is joined — Shannon's
// "A Mathematical Theory of Communication" reads "...which exchange
// T bandwidth for signal-to-noise ratio...". Detect it, prepend the
// character to the paragraph's first line, and drop it from this one.
//
// The size gate is 1.8x rather than 2.5x because bitmap (Type3) caps
// report their glyph bbox rather than the em box, so a two-line cap
// can measure as little as ~1.9x the body size.
if line.items.len() > 1 {
let first = &line.items[0];
// The remainder must be a substantive body run: a lone label or
// math fragment beside a large glyph is not a drop-cap paragraph.
let rest_letters: usize = line.items[1..]
.iter()
.map(|i| i.text.chars().filter(|c| c.is_alphabetic()).count())
.sum();
let is_embedded_cap = first.font_size >= base_size * 1.8
&& first.text.trim().chars().count() == 1
&& first
.text
.trim()
.chars()
.next()
.is_some_and(char::is_uppercase)
&& line.items[1..]
.iter()
.all(|i| i.font_size < base_size * 1.5)
&& line.items[1..].iter().any(|i| i.x > first.x)
&& rest_letters >= 8;
if is_embedded_cap {
let drop_char = first.text.trim().chars().next().unwrap();
let cap_x = first.x;
let line_y = line.y;
// Text on the cap's own line, pushed right to clear the glyph.
let rest_x = line.items[1].x;
// Walk up the run of lines the cap has indented. A drop cap
// pushes every line it covers to the right of the glyph, so
// the paragraph's first line is the TOPMOST line sharing that
// indent — however many lines the cap spans. Using the indent
// rather than the cap's font size is what makes this work for
// three- and four-line initials as well as two-line ones;
// deriving a line count from the em size does not survive
// contact with real documents, where 36-47pt initials sit
// over 11-14pt leading.
//
// A cap in a different column has no such run (its neighbours
// sit at an unrelated x), so it is left alone — which is
// correct when the cap's own line already carries the rest of
// the word.
const INDENT_TOLERANCE: f32 = 2.0;
const MAX_CAP_LINES: usize = 8;
let max_step = base_size * 2.5;
let mut target_idx = result.len();
let mut expected_y = line_y;
while target_idx > 0 && result.len() - target_idx < MAX_CAP_LINES {
let cand = &result[target_idx - 1];
let step = cand.y - expected_y;
let shares_indent = cand
.items
.first()
.is_some_and(|i| (i.x - rest_x).abs() <= INDENT_TOLERANCE);
if cand.page != line.page || step <= 0.0 || step > max_step || !shares_indent {
break;
}
expected_y = cand.y;
target_idx -= 1;
}
// The topmost line of the run is the paragraph's first line.
// The line above THAT tells us whether it starts a paragraph.
let before_target = target_idx
.checked_sub(1)
.and_then(|i| result.get(i))
.filter(|l| l.page == line.page)
.map(|l| (l.text().trim_end().to_string(), l.y));
// Leading within the run: the step from the target down to the
// next line of the paragraph, which is the cap's own line when
// the run is a single line.
let run_step = result
.get(target_idx)
.map(|t| {
let below_y = result.get(target_idx + 1).map_or(line_y, |b| b.y);
t.y - below_y
})
.unwrap_or(0.0);
let step_for_gap = if run_step > 0.0 {
run_step
} else {
base_size * 1.2
};
let target = (target_idx < result.len())
.then(|| &mut result[target_idx])
.filter(|prev| {
let prev_text = prev.text();
let prev_trimmed = prev_text.trim();
// A hyphen on the line above means the target resumes
// a split word, so it continues a paragraph rather
// than starting one (polkuja_ylakoulu: "ylakou-" +
// "lulaisten").
//
// Case cannot serve as a continuation signal here: the
// target legitimately starts lowercase, because the
// cap removes the word's first letter and leaves
// "ver the course..." for "Over".
let continues_previous = before_target
.as_ref()
.is_some_and(|(b, _)| b.ends_with('-'));
// The target must START a paragraph: extra leading
// above it, a completed sentence on the line above, or
// nothing above it at all.
let starts_paragraph = match before_target.as_ref() {
None => true,
Some((text, y)) => {
y - prev.y > step_for_gap * 1.15 || ends_sentence(text)
}
};
!continues_previous
&& starts_paragraph
&& prev.page == line.page
&& prev.y > line_y
// Indented past the cap glyph, not merely to its
// right by an arbitrary amount.
&& prev
.items
.first()
.is_some_and(|i| i.x > cap_x && i.x - cap_x <= first.font_size * 2.0)
// Body text, so headings, labels and table
// fragments are never rewritten.
&& prev_trimmed
.chars()
.next()
.is_some_and(char::is_alphabetic)
&& prev_trimmed.chars().filter(|c| c.is_alphabetic()).count() >= 8
});
if let Some(prev_line) = target {
if let Some(first_item) = prev_line.items.first_mut() {
// A mid-word cap ("T" + "HE recent") joins directly.
// Leading whitespace only marks a word boundary when
// the cap is itself a single-letter word, since the
// paragraph's indent can also arrive as whitespace.
const SINGLE_LETTER_WORDS: &[char] = &['A', 'I', 'O', 'U', 'Y', 'E'];
let had_leading_ws = first_item.text.starts_with(char::is_whitespace)
&& SINGLE_LETTER_WORDS.contains(&drop_char);
let rest = first_item.text.trim_start().to_string();
first_item.text = if had_leading_ws {
format!("{} {}", drop_char, rest)
} else {
format!("{}{}", drop_char, rest)
};
}
let mut line = line.clone();
line.items.remove(0);
result.push(line);
continue;
}
}
}
if is_drop_cap {
let drop_char = trimmed.chars().next().unwrap();
@@ -816,314 +598,6 @@ mod tests {
}
}
fn make_item_at(text: &str, font_size: f32, x: f32) -> TextItem {
let mut item = make_item(text, font_size, None);
item.x = x;
item.width = text.len() as f32 * font_size * 0.5;
item
}
#[test]
fn embedded_drop_cap_moves_to_paragraph_start() {
// A two-line cap baseline-aligns with the paragraph's SECOND line,
// so it lands as that line's first item (Shannon entropy.pdf p.1).
let first_line = TextLine {
items: vec![make_item_at(
"HE recent development which exchange",
10.0,
90.0,
)],
// 16pt baseline step under a 25pt cap: a genuine two-line cap.
y: 716.0,
page: 1,
adaptive_threshold: 0.10,
};
let second_line = TextLine {
items: vec![
make_item_at("T", 25.0, 72.0),
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![first_line, second_line], 10.0);
assert_eq!(result.len(), 2);
assert!(
result[0].text().starts_with("THE recent"),
"cap should prepend to the paragraph start: {}",
result[0].text()
);
assert!(
result[1].text().starts_with("bandwidth"),
"cap must be removed from the second line: {}",
result[1].text()
);
}
#[test]
fn embedded_drop_cap_walks_a_multi_line_initial_to_the_paragraph_start() {
// A 47pt initial over 13pt leading covers four lines, so the
// paragraph's first line is three lines above the cap rather than
// immediately above it (polkuja_ylakoulu). The indented run, not the
// cap's em size, is what locates it.
let mut lines = vec![TextLine {
items: vec![make_item_at("Previous paragraph ends here.", 10.0, 72.0)],
y: 766.0,
page: 1,
adaptive_threshold: 0.10,
}];
for (i, text) in [
"rilaiset mediasisallot ovat tarkea osa",
"useimpien ylakoululaisten elamaa ja",
"muuta tekstia jatkuu tassa viela",
]
.iter()
.enumerate()
{
lines.push(TextLine {
items: vec![make_item_at(text, 10.0, 90.0)],
y: 753.0 - 13.0 * i as f32,
page: 1,
adaptive_threshold: 0.10,
});
}
lines.push(TextLine {
items: vec![
make_item_at("E", 47.0, 72.0),
make_item_at("loppuosa tekstista tassa", 10.0, 90.0),
],
y: 714.0,
page: 1,
adaptive_threshold: 0.10,
});
let result = merge_drop_caps(lines, 10.0);
assert!(
result[1].text().starts_with("Erilaiset"),
"cap belongs on the topmost line of the indented run: {}",
result[1].text()
);
assert!(
result[2].text().starts_with("useimpien"),
"intervening run lines must be untouched: {}",
result[2].text()
);
assert!(
result[4].text().starts_with("loppuosa"),
"cap must be removed from its own line: {}",
result[4].text()
);
}
#[test]
fn embedded_drop_cap_ignores_non_paragraph_neighbours() {
// Same geometry, but the preceding line is a short label rather than
// body text, so it must not be rewritten.
let label = TextLine {
items: vec![make_item_at("Fig. 2", 10.0, 90.0)],
y: 716.0,
page: 1,
adaptive_threshold: 0.10,
};
let second_line = TextLine {
items: vec![
make_item_at("T", 25.0, 72.0),
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![label, second_line], 10.0);
assert_eq!(result[0].text().trim(), "Fig. 2");
assert!(
result[1].text().starts_with('T'),
"cap stays put: {}",
result[1].text()
);
}
#[test]
fn embedded_drop_cap_keeps_a_space_for_standalone_word_caps() {
// Leading whitespace on the paragraph's first item marks the cap as
// a word of its own rather than the first letter of one.
let mut lead = make_item_at("long time ago in a galaxy far away", 10.0, 90.0);
lead.text = " long time ago in a galaxy far away".to_string();
let first_line = TextLine {
items: vec![lead],
y: 716.0,
page: 1,
adaptive_threshold: 0.10,
};
let second_line = TextLine {
items: vec![
make_item_at("A", 25.0, 72.0),
make_item_at("continued here with more body text", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![first_line, second_line], 10.0);
assert!(
result[0].text().starts_with("A long time ago"),
"standalone-word cap keeps one space: {}",
result[0].text()
);
}
#[test]
fn embedded_drop_cap_skips_hyphenation_continuation_targets() {
// The line above the RUN ends on a hyphen, so the run's topmost line
// resumes a split word rather than starting a paragraph. It sits at
// the paragraph margin (x=72), outside the cap's indent, so it is not
// part of the run itself.
let split_word = TextLine {
items: vec![make_item_at(
"mediasisallot ovat osa useimpien ylakou-",
10.0,
72.0,
)],
y: 728.0,
page: 1,
adaptive_threshold: 0.10,
};
let run_top = TextLine {
items: vec![make_item_at(
"lulaisten elamaa ja muuta tekstia",
10.0,
90.0,
)],
y: 714.0,
page: 1,
adaptive_threshold: 0.10,
};
let cap_line = TextLine {
items: vec![
make_item_at("E", 25.0, 72.0),
make_item_at("jatkuu tassa lisaa leipatekstia", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![split_word, run_top, cap_line], 10.0);
assert!(
result[1].text().starts_with("lulaisten"),
"a run resuming a split word must not receive the cap: {}",
result[1].text()
);
assert!(
result[2].text().starts_with('E'),
"cap stays put when no valid target exists: {}",
result[2].text()
);
}
#[test]
fn embedded_drop_cap_indent_is_not_a_word_boundary() {
// The paragraph's first line is indented to clear the cap, and that
// indent can arrive as leading whitespace. A mid-word cap must still
// join directly — "T HE recent" would be the defect this fixes.
let mut lead = make_item_at("HE recent development and more body text", 10.0, 90.0);
lead.text = " HE recent development and more body text".to_string();
let first_line = TextLine {
items: vec![lead],
y: 716.0,
page: 1,
adaptive_threshold: 0.10,
};
let cap_line = TextLine {
items: vec![
make_item_at("T", 25.0, 72.0),
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![first_line, cap_line], 10.0);
assert!(
result[0].text().starts_with("THE recent"),
"indent must not be read as a word boundary: {}",
result[0].text()
);
}
#[test]
fn ends_sentence_rejects_abbreviations_and_markers() {
use super::ends_sentence;
assert!(ends_sentence("This completes the thought."));
assert!(ends_sentence("Is that so?"));
assert!(ends_sentence("Stop!"));
// Periods that do not end a sentence.
assert!(!ends_sentence("as shown in Fig."));
assert!(!ends_sentence("see e.g."));
// Non-ASCII abbreviations. The two-CHARACTER segment is the case
// that distinguishes a character count from a byte count: "пр" is
// 2 chars but 4 bytes, so a byte-based bound would reject it and
// read the line as a completed sentence.
assert!(!ends_sentence("и т.пр."));
assert!(!ends_sentence("см. т.е."));
assert!(!ends_sentence("napr. ú.d."));
assert!(!ends_sentence("reviewed by Dr."));
// Standalone enumerators, any case.
assert!(!ends_sentence("1."));
assert!(!ends_sentence("IV."));
assert!(!ends_sentence("ii."));
assert!(!ends_sentence("xii."));
// Numbers and numerals that genuinely end a sentence must count,
// or a legitimate drop-cap merge is blocked.
assert!(ends_sentence("The paper was published in 2020."));
assert!(ends_sentence("He scored 5."));
assert!(ends_sentence("after World War II."));
assert!(ends_sentence("the constant equals 3.14."));
assert!(ends_sentence("documented at example.com."));
assert!(!ends_sentence("a trailing clause with no period"));
}
#[test]
fn embedded_drop_cap_allows_first_paragraph_on_a_new_page() {
// The line two back is on the previous page, so its y is unrelated
// and must not be used as leading evidence.
let prev_page_tail = TextLine {
items: vec![make_item_at(
"tail of the previous page body text",
10.0,
90.0,
)],
y: 90.0,
page: 1,
adaptive_threshold: 0.10,
};
let first_line = TextLine {
items: vec![make_item_at(
"HE recent development which exchange",
10.0,
90.0,
)],
y: 716.0,
page: 2,
adaptive_threshold: 0.10,
};
let cap_line = TextLine {
items: vec![
make_item_at("T", 25.0, 72.0),
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
],
y: 700.0,
page: 2,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![prev_page_tail, first_line, cap_line], 10.0);
assert!(
result[1].text().starts_with("THE recent"),
"a page break must not suppress the merge: {}",
result[1].text()
);
}
#[test]
fn test_merge_struct_tree_headings() {
// Two consecutive lines tagged as H2 via struct tree, same font size as body
+19 -1
View File
@@ -594,7 +594,7 @@ fn hex_to_unicode_string(hex: &str) -> Option<String> {
let bytes: Option<Vec<u8>> = (0..hex.len())
.step_by(2)
.map(|i| u8::from_str_radix(&hex[i..i + 2], 16).ok())
.map(|i| u8::from_str_radix(hex.get(i..i + 2)?, 16).ok())
.collect();
let bytes = bytes?;
@@ -2606,6 +2606,24 @@ endcmap
assert_eq!(cmap.lookup(0x0025), Some("B".to_string()));
}
#[test]
fn test_hex_to_unicode_non_ascii_no_panic() {
// A destination containing a multi-byte char makes the byte length even
// while a byte offset can land inside a char. Slicing must not panic;
// it should be rejected gracefully.
assert_eq!(hex_to_unicode_string("XéY"), None);
assert_eq!(hex_to_unicode_string("\u{fffd}0"), None);
}
#[test]
fn test_parse_bfchar_non_ascii_destination_no_panic() {
// Crafted /ToUnicode CMap: a non-hex, non-ASCII destination previously
// triggered a char-boundary panic in hex_to_unicode_string.
let cmap_content = "beginbfchar <0041> <XéY> endbfchar";
// Must not panic; the malformed entry is simply skipped.
let _ = ToUnicodeCMap::parse(cmap_content.as_bytes());
}
#[test]
fn test_parse_bfchar_1byte() {
// This is the pattern that caused the CJK bug: codespace is <0000><FFFF>
+7 -3
View File
@@ -1,12 +1,16 @@
Reprinted with corrections from *The Bell System Technical Journal,* Vol. 27, pp. 379423, 623656, July, October, 1948.
# A Mathematical Theory of Communication
## A Mathematical Theory of Communication
## By C. E. SHANNON
### By C. E. SHANNON
INTRODUCTION
THE recent development of various methods of modulation such as PCM and PPM which exchange bandwidth for signal-to-noise ratio has intensified the interest in a general theory of communication. A basis for such a theory is contained in the important papers of Nyquist¹ and Hartley² on this subject. In the present paper we will extend the theory to include a number of new factors, in particular the effect of noise in the channel, and the savings possible due to the statistical structure of the original message and due to the nature of the final destination of the information. The fundamental problem of communication is that of reproducing at one point either exactly or ap- proximately a message selected at another point. Frequently the messages have *meaning*; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one *selected from a set* of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design. If the number of messages in the set is finite then this number or any monotonic function of this number can be regarded as a measure of the information produced when one message is chosen from the set, all choices being equally likely. As was pointed out by Hartley the most natural choice is the logarithmic function. Although this definition must be generalized considerably when we consider the influence of the statistics of the message and when we have a continuous range of messages, we will in all cases use an essentially logarithmic measure. The logarithmic measure is more convenient for various reasons:
HE recent development of various methods of modulation such as PCM and PPM which exchange
# Tbandwidth for signal-to-noise ratio has intensified the interest in a general theory of communication. A
basis for such a theory is contained in the important papers of Nyquist¹ and Hartley² on this subject. In the present paper we will extend the theory to include a number of new factors, in particular the effect of noise in the channel, and the savings possible due to the statistical structure of the original message and due to the nature of the final destination of the information. The fundamental problem of communication is that of reproducing at one point either exactly or ap- proximately a message selected at another point. Frequently the messages have *meaning*; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one *selected from a set* of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design. If the number of messages in the set is finite then this number or any monotonic function of this number can be regarded as a measure of the information produced when one message is chosen from the set, all choices being equally likely. As was pointed out by Hartley the most natural choice is the logarithmic function. Although this definition must be generalized considerably when we consider the influence of the statistics of the message and when we have a continuous range of messages, we will in all cases use an essentially logarithmic measure. The logarithmic measure is more convenient for various reasons:
1. It is practically more useful. Parameters of engineering importance such as time, bandwidth, number of relays, etc., tend to vary linearly with the logarithm of the number of possibilities. For example, adding one relay to a group doubles the number of possible states of the relays. It adds 1 to the base 2 logarithm of this number. Doubling the time roughly squares the number of possible messages, or doubles the logarithm, etc.
2. It is nearer to our intuitive feeling as to the proper measure. This is closely related to (1) since we in- tuitively measures entities by linear comparison with common standards. One feels, for example, that two punched cards should have twice the capacity of one for information storage, and two identical channels twice the capacity of one for transmitting information.