Compare commits

...
Author SHA1 Message Date
Abimael Martell afca0c5ac8 fix(markdown): locate drop-cap paragraphs by indent run, not cap size
The previous rule only handled two-line caps: it took the line
immediately above the cap and required the cap's em to be under 2.5
line-steps. That fired on exactly one document in 386 — Shannon — and
rejected every other real initial, because production initials are
36-47pt over 11-14pt leading and cover three or four lines.

A drop cap pushes every line it covers to the right of the glyph, so
the paragraph's first line is the topmost line sharing that indent,
whatever the line count. Walk that run instead of deriving a count from
the font size, and evaluate the paragraph-start and continuation guards
against the run's top rather than the cap's immediate neighbour.

A cap in a different column has no such run, so it is left alone.

Across pdf-evals (192) and opendataloader-bench (200) this now fixes
three documents instead of one, each losing a truncated first word and
a junk heading built from the orphaned cap ("# Omany patients"). The
opendataloader score moves 0.87836 -> 0.87912 overall and 0.7878 ->
0.7906 MHS with no metric regressing; 275_41c50d3a's pdf-evals
composite moves 0.5772 -> 0.5862.
2026-08-08 12:41:02 -07:00
Abimael Martell 764ec6e081 test: cover the byte-vs-character bound with a real two-character segment
Valid, and the third time this review has caught a test that did not
exercise its own fix. My cases used single-character segments ('т', 'е',
'ú'), which are 2 bytes each and so already satisfied the old
seg.len() <= 2 bound — a revert to len() would have passed them.

Added 'и т.пр.', whose final segment 'пр' is 2 characters but 4 bytes,
which is precisely the case the two bounds disagree on.

This time I verified the discrimination directly rather than assuming
it: with seg.len() the new assertion fails, with seg.chars().count() it
passes. 998 tests, clippy clean.
2026-08-07 20:14:48 -07:00
Abimael Martell 631f3ff192 review: measure abbreviation segments in characters, not bytes
Valid. The segment-length bound used seg.len(), which counts bytes, so a
two-letter non-ASCII abbreviation ('т.е.', 'ú.d.') measured four or more
bytes, failed the check, and was read as a completed sentence — which
would let a drop cap merge into a continuation.

Counts characters now, with both examples pinned in the test.

998 tests pass, clippy clean; Shannon and the corpus are unchanged.
2026-08-07 19:47:00 -07:00
Abimael Martell ba1807dce5 review: make enumerator and abbreviation tests precise in ends_sentence
Two findings, pulling in opposite directions, which is what made the
original rule wrong in both.

conf 10 — lowercase roman markers ('ii.') were not recognised, because
the numeral check compared uppercase characters only. Now
case-insensitive.

conf 7 — the same check was too broad in the other direction: any last
token of digits or roman letters counted as a marker, so 'published in
2020.', 'He scored 5.' and 'after World War II.' were not sentence ends
and a legitimate drop cap would be left in place mid-sentence.

Resolved by using position rather than shape: an enumerator stands alone
on its line ('1.', 'ii.', 'IV.'), whereas a number or numeral closing a
sentence has words before it. The marker test now applies only when the
line is a single token.

The abbreviation test had the same flaw. 'e.g.' and 'example.com.' both
carry an internal period with letters, so the old rule classified the
domain as an abbreviation. Abbreviation segments are short and
alphabetic ('e.g.', 'U.S.'), so every dot-separated segment must now be
at most two alphabetic characters; domains and decimals ('3.14.') fall
through as sentence ends.

Unchanged: Shannon reads 'THE recent development ...', 0 of 186
pdf-evals documents affected. 998 tests pass, clippy clean.
2026-08-07 19:41:45 -07:00
Abimael Martell 597e125731 review: drop stray diagnostic, scope to page, require a real sentence end
Three findings from the third review, all valid.

P2 (conf 10) — a GEO log::warn! diagnostic I used while measuring the
cap geometry was left in the committed code, so every embedded-cap
candidate emitted a warning during ordinary conversion, including
candidates deliberately left untouched. Removed. It should never have
been committed; verified with RUST_LOG=warn that conversion is now
silent apart from a pre-existing lopdf encoding warning.

P2 (conf 9) — a cap on the first paragraph of a new page could be
suppressed because before_target could come from the previous page,
whose y is unrelated, making the leading comparison meaningless.
before_target is now scoped to the cap's own page.

P3 (conf 4) — starts_paragraph accepted any line ending in a period, so
'as shown in Fig.', 'see e.g.', 'item 1.' and 'reviewed by Dr.' read as
paragraph boundaries and could let a continuation take the cap. Replaced
with ends_sentence(), which rejects internal-period abbreviations,
numeric and roman list markers, and a short abbreviation list.

Unchanged: Shannon reads 'THE recent development ...', 0 of 186
pdf-evals documents affected. 998 tests pass, clippy clean.
2026-08-07 19:33:21 -07:00
Abimael Martell b0b42c4949 review: gate on cap line-count and require a paragraph start
Three findings from the second review, all valid.

P1 — ordinary continuations could still take the cap. Rejecting only a
preceding hyphen treated any non-hyphenated context as a paragraph
start, so a continuation sharing the left edge became 'Tcontinuation'.
The target must now START a paragraph: either extra leading above it
(> 1.15x the local step) or a completed sentence on the line above.
Shannon's target has 18.6pt of leading above against a 12.1pt step.

P2 — the vertical-step bound was expressed as a fraction of the cap em,
but the step is the body leading and is independent of cap size. A 25pt
cap over 12pt leading is a genuine two-line cap yet failed 'step >= em/2'.
Recast as a line count: the cap must be at most 2.5 line-steps tall.
Shannon measures 1.53; the four-line initials this cannot place measure
3.5 and 4.0. Tall caps over tight leading now pass, as they should.

P2 — the hyphenation regression test did not exercise its own guard: at
a 12pt step under a 25pt cap the continuation was already rejected by
the vertical gate, so deleting the hyphen check would not have failed
it. Raised the step to 14pt so the hyphen guard is what does the work.

Unchanged: Shannon reads 'THE recent development ...', 0 of 186
pdf-evals documents affected, 996 tests pass, clippy clean.
2026-08-07 16:32:41 -07:00
Abimael Martell dd68cc221f review: scope drop-cap merging to genuine two-line caps
Addresses both cubic findings, and measuring the geometry changed what
this PR claims.

Finding 1 (target may be a continuation): real. Two guards now.

- The line above the target must not end on a hyphen, or the target
  resumes a split word (polkuja_ylakoulu: 'ylakou-' + 'lulaisten').
- The target and the cap's own line must share a left edge, since both
  clear the glyph; a continuation sits at the paragraph margin instead.

A third guard I tried and removed: rejecting targets that start
lowercase. That is backwards here — the cap takes the word's first
letter, so the paragraph's first line legitimately starts lowercase
('ver the course...' for 'Over').

Finding 2 (indent misread as a word boundary): real. Leading whitespace
now only marks a standalone-word cap when the cap is itself a
single-letter word, so an indented mid-word cap cannot yield 'T HE'.

The substantive change is scope. Measuring the two corpus documents
showed both use caps of 44-47pt over an 11-13pt leading — four-line
initials, not two-line. For those the paragraph's first line is several
lines up, and taking the line immediately above is a coin flip: right
for 275_41c50d3a, wrong for polkuja_ylakoulu, where it produced
'Ntulla'. Requiring the baseline step to be at least half the cap height
restricts this to genuine two-line caps, where the immediately preceding
line is the paragraph start by construction.

Consequences, stated plainly:

- Shannon still reads 'THE recent development ...' — the reported case.
- Corpus impact is now zero: 0 of 186 pdf-evals documents and no change
  on opendataloader beyond what #264 already contributes.
- The earlier +0.0033 mhs came from the unscoped version merging
  multi-line caps, which also introduced the polkuja corruption. That
  gain was not real and is withdrawn.
- Multi-line initials remain unhandled; locating their paragraph start
  needs a different approach than looking one line up.

1004 tests pass, clippy clean.
2026-08-07 16:16:43 -07:00
Abimael Martell 20c5a0a82a fix(markdown): place two-line drop caps at the paragraph start
A two-line drop cap's baseline aligns with the paragraph's SECOND line,
so Y-grouping puts the glyph at the start of that line rather than on a
line of its own. The existing merge_drop_caps only handled caps that
occupy their own line, so the embedded form survived into the output and
surfaced mid-sentence once the paragraph was joined.

Shannon's 'A Mathematical Theory of Communication' — the PDF that
prompted this work — opened with

    HE recent development of various methods of modulation such as PCM
    and PPM which exchange T bandwidth for signal-to-noise ratio ...

and now reads 'THE recent development ... which exchange bandwidth'.

Detection requires the cap to be a single uppercase glyph >= 1.8x body
size (bitmap Type3 caps report their glyph bbox rather than the em box,
so a two-line cap can measure as little as ~1.9x), followed on the same
line by a substantive body run, with the paragraph's first line directly
above in the same column and indented past the cap. The target must
itself read as body text, so headings, labels and table fragments that
merely fall in the geometric window are never rewritten.

Removing the cap from the line also stops it inflating heading tiers:
Shannon's title moves from h2 to h1, and 275_41c50d3a's heading levels
all shift up one, both correct.

Corpus impact is narrow — 2 of 186 pdf-evals documents:

- 275_41c50d3a: heading levels shift up one, as above.
- polkuja_ylakoulu: a wash. The cap is misplaced both before and after
  (main prefixes it to 'pelaaminen' inside a spurious heading, this
  version prefixes it to 'lulaisten'), but the spurious heading is gone
  and the paragraph joins correctly. Noted rather than tuned away — the
  target selection takes the immediately preceding line, which can be
  the wrong column on multi-column pages.

opendataloader-bench (200 docs, ground truth): overall +0.0012, mhs
+0.0036 against pre-#264 main, of which #264 contributed +0.0003 /
+0.0003 — so this change accounts for roughly +0.0009 overall and
+0.0033 mhs, the largest heading-metric gain measured in this series.

994 tests pass, clippy clean.
2026-08-07 15:51:27 -07:00
2 changed files with 529 additions and 7 deletions
+526
View File
@@ -150,6 +150,61 @@ pub(crate) fn merge_heading_lines(
/// Merge drop caps with the appropriate line.
/// A drop cap is a single large letter at the start of a paragraph.
/// Due to PDF coordinate sorting, the drop cap may appear AFTER the line it belongs to.
/// True when the text ends a sentence, as opposed to merely ending in a
/// period. An abbreviation or list marker ("e.g.", "Fig.", "Mr.", "1.")
/// closes with a period mid-sentence, so treating those as paragraph
/// boundaries would let a drop cap be prepended to a continuation.
fn ends_sentence(text: &str) -> bool {
let t = text.trim_end();
if t.ends_with(['!', '?']) {
return true;
}
let Some(stripped) = t.strip_suffix('.') else {
return false;
};
let last = stripped.split_whitespace().next_back().unwrap_or("");
if last.is_empty() {
return false;
}
// Abbreviations carry an internal period between very short segments
// ("e.g.", "i.e.", "U.S."). Domains and decimals have the same shape but
// longer or numeric segments ("example.com.", "3.14."), and those end
// sentences perfectly well, so require every segment to be short and
// alphabetic before reading the internal period as an abbreviation.
if last.contains('.')
&& last
.split('.')
.filter(|seg| !seg.is_empty())
// Characters, not bytes: a two-letter non-ASCII abbreviation
// ("т.е.", "ú.d.") measures four or more bytes and would
// otherwise be read as a completed sentence.
.all(|seg| seg.chars().count() <= 2 && seg.chars().all(char::is_alphabetic))
{
return false;
}
// Enumerators stand alone on their line ("1.", "ii.", "IV."). A number
// or numeral in the tail of a sentence does not — "published in 2020.",
// "He scored 5." and "after World War II." all end sentences, and
// treating them as markers would block a legitimate drop-cap merge.
if stripped.split_whitespace().count() == 1 {
let is_numeric = last.chars().all(|c| c.is_ascii_digit());
let is_roman = last
.chars()
.all(|c| matches!(c.to_ascii_uppercase(), 'I' | 'V' | 'X' | 'L' | 'C'));
if is_numeric || is_roman {
return false;
}
}
const ABBREVIATIONS: &[&str] = &[
"Fig", "No", "Mr", "Mrs", "Ms", "Dr", "St", "vs", "etc", "al", "Ed", "Eq", "Ch", "pp",
"Vol", "cf", "Prof", "Inc", "Ltd", "Jr", "Sr",
];
!ABBREVIATIONS.iter().any(|a| a.eq_ignore_ascii_case(last))
}
pub(crate) fn merge_drop_caps(lines: Vec<TextLine>, base_size: f32) -> Vec<TextLine> {
let mut result: Vec<TextLine> = Vec::with_capacity(lines.len());
@@ -169,6 +224,169 @@ pub(crate) fn merge_drop_caps(lines: Vec<TextLine>, base_size: f32) -> Vec<TextL
.map(|c| c.is_uppercase())
.unwrap_or(false);
// Embedded drop cap: a two-line cap's baseline aligns with the
// paragraph's SECOND line, so Y-grouping puts the glyph at the start
// of that line rather than on a line of its own. Left there it
// surfaces mid-sentence once the paragraph is joined — Shannon's
// "A Mathematical Theory of Communication" reads "...which exchange
// T bandwidth for signal-to-noise ratio...". Detect it, prepend the
// character to the paragraph's first line, and drop it from this one.
//
// The size gate is 1.8x rather than 2.5x because bitmap (Type3) caps
// report their glyph bbox rather than the em box, so a two-line cap
// can measure as little as ~1.9x the body size.
if line.items.len() > 1 {
let first = &line.items[0];
// The remainder must be a substantive body run: a lone label or
// math fragment beside a large glyph is not a drop-cap paragraph.
let rest_letters: usize = line.items[1..]
.iter()
.map(|i| i.text.chars().filter(|c| c.is_alphabetic()).count())
.sum();
let is_embedded_cap = first.font_size >= base_size * 1.8
&& first.text.trim().chars().count() == 1
&& first
.text
.trim()
.chars()
.next()
.is_some_and(char::is_uppercase)
&& line.items[1..]
.iter()
.all(|i| i.font_size < base_size * 1.5)
&& line.items[1..].iter().any(|i| i.x > first.x)
&& rest_letters >= 8;
if is_embedded_cap {
let drop_char = first.text.trim().chars().next().unwrap();
let cap_x = first.x;
let line_y = line.y;
// Text on the cap's own line, pushed right to clear the glyph.
let rest_x = line.items[1].x;
// Walk up the run of lines the cap has indented. A drop cap
// pushes every line it covers to the right of the glyph, so
// the paragraph's first line is the TOPMOST line sharing that
// indent — however many lines the cap spans. Using the indent
// rather than the cap's font size is what makes this work for
// three- and four-line initials as well as two-line ones;
// deriving a line count from the em size does not survive
// contact with real documents, where 36-47pt initials sit
// over 11-14pt leading.
//
// A cap in a different column has no such run (its neighbours
// sit at an unrelated x), so it is left alone — which is
// correct when the cap's own line already carries the rest of
// the word.
const INDENT_TOLERANCE: f32 = 2.0;
const MAX_CAP_LINES: usize = 8;
let max_step = base_size * 2.5;
let mut target_idx = result.len();
let mut expected_y = line_y;
while target_idx > 0 && result.len() - target_idx < MAX_CAP_LINES {
let cand = &result[target_idx - 1];
let step = cand.y - expected_y;
let shares_indent = cand
.items
.first()
.is_some_and(|i| (i.x - rest_x).abs() <= INDENT_TOLERANCE);
if cand.page != line.page || step <= 0.0 || step > max_step || !shares_indent {
break;
}
expected_y = cand.y;
target_idx -= 1;
}
// The topmost line of the run is the paragraph's first line.
// The line above THAT tells us whether it starts a paragraph.
let before_target = target_idx
.checked_sub(1)
.and_then(|i| result.get(i))
.filter(|l| l.page == line.page)
.map(|l| (l.text().trim_end().to_string(), l.y));
// Leading within the run: the step from the target down to the
// next line of the paragraph, which is the cap's own line when
// the run is a single line.
let run_step = result
.get(target_idx)
.map(|t| {
let below_y = result.get(target_idx + 1).map_or(line_y, |b| b.y);
t.y - below_y
})
.unwrap_or(0.0);
let step_for_gap = if run_step > 0.0 {
run_step
} else {
base_size * 1.2
};
let target = (target_idx < result.len())
.then(|| &mut result[target_idx])
.filter(|prev| {
let prev_text = prev.text();
let prev_trimmed = prev_text.trim();
// A hyphen on the line above means the target resumes
// a split word, so it continues a paragraph rather
// than starting one (polkuja_ylakoulu: "ylakou-" +
// "lulaisten").
//
// Case cannot serve as a continuation signal here: the
// target legitimately starts lowercase, because the
// cap removes the word's first letter and leaves
// "ver the course..." for "Over".
let continues_previous = before_target
.as_ref()
.is_some_and(|(b, _)| b.ends_with('-'));
// The target must START a paragraph: extra leading
// above it, a completed sentence on the line above, or
// nothing above it at all.
let starts_paragraph = match before_target.as_ref() {
None => true,
Some((text, y)) => {
y - prev.y > step_for_gap * 1.15 || ends_sentence(text)
}
};
!continues_previous
&& starts_paragraph
&& prev.page == line.page
&& prev.y > line_y
// Indented past the cap glyph, not merely to its
// right by an arbitrary amount.
&& prev
.items
.first()
.is_some_and(|i| i.x > cap_x && i.x - cap_x <= first.font_size * 2.0)
// Body text, so headings, labels and table
// fragments are never rewritten.
&& prev_trimmed
.chars()
.next()
.is_some_and(char::is_alphabetic)
&& prev_trimmed.chars().filter(|c| c.is_alphabetic()).count() >= 8
});
if let Some(prev_line) = target {
if let Some(first_item) = prev_line.items.first_mut() {
// A mid-word cap ("T" + "HE recent") joins directly.
// Leading whitespace only marks a word boundary when
// the cap is itself a single-letter word, since the
// paragraph's indent can also arrive as whitespace.
const SINGLE_LETTER_WORDS: &[char] = &['A', 'I', 'O', 'U', 'Y', 'E'];
let had_leading_ws = first_item.text.starts_with(char::is_whitespace)
&& SINGLE_LETTER_WORDS.contains(&drop_char);
let rest = first_item.text.trim_start().to_string();
first_item.text = if had_leading_ws {
format!("{} {}", drop_char, rest)
} else {
format!("{}{}", drop_char, rest)
};
}
let mut line = line.clone();
line.items.remove(0);
result.push(line);
continue;
}
}
}
if is_drop_cap {
let drop_char = trimmed.chars().next().unwrap();
@@ -598,6 +816,314 @@ mod tests {
}
}
fn make_item_at(text: &str, font_size: f32, x: f32) -> TextItem {
let mut item = make_item(text, font_size, None);
item.x = x;
item.width = text.len() as f32 * font_size * 0.5;
item
}
#[test]
fn embedded_drop_cap_moves_to_paragraph_start() {
// A two-line cap baseline-aligns with the paragraph's SECOND line,
// so it lands as that line's first item (Shannon entropy.pdf p.1).
let first_line = TextLine {
items: vec![make_item_at(
"HE recent development which exchange",
10.0,
90.0,
)],
// 16pt baseline step under a 25pt cap: a genuine two-line cap.
y: 716.0,
page: 1,
adaptive_threshold: 0.10,
};
let second_line = TextLine {
items: vec![
make_item_at("T", 25.0, 72.0),
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![first_line, second_line], 10.0);
assert_eq!(result.len(), 2);
assert!(
result[0].text().starts_with("THE recent"),
"cap should prepend to the paragraph start: {}",
result[0].text()
);
assert!(
result[1].text().starts_with("bandwidth"),
"cap must be removed from the second line: {}",
result[1].text()
);
}
#[test]
fn embedded_drop_cap_walks_a_multi_line_initial_to_the_paragraph_start() {
// A 47pt initial over 13pt leading covers four lines, so the
// paragraph's first line is three lines above the cap rather than
// immediately above it (polkuja_ylakoulu). The indented run, not the
// cap's em size, is what locates it.
let mut lines = vec![TextLine {
items: vec![make_item_at("Previous paragraph ends here.", 10.0, 72.0)],
y: 766.0,
page: 1,
adaptive_threshold: 0.10,
}];
for (i, text) in [
"rilaiset mediasisallot ovat tarkea osa",
"useimpien ylakoululaisten elamaa ja",
"muuta tekstia jatkuu tassa viela",
]
.iter()
.enumerate()
{
lines.push(TextLine {
items: vec![make_item_at(text, 10.0, 90.0)],
y: 753.0 - 13.0 * i as f32,
page: 1,
adaptive_threshold: 0.10,
});
}
lines.push(TextLine {
items: vec![
make_item_at("E", 47.0, 72.0),
make_item_at("loppuosa tekstista tassa", 10.0, 90.0),
],
y: 714.0,
page: 1,
adaptive_threshold: 0.10,
});
let result = merge_drop_caps(lines, 10.0);
assert!(
result[1].text().starts_with("Erilaiset"),
"cap belongs on the topmost line of the indented run: {}",
result[1].text()
);
assert!(
result[2].text().starts_with("useimpien"),
"intervening run lines must be untouched: {}",
result[2].text()
);
assert!(
result[4].text().starts_with("loppuosa"),
"cap must be removed from its own line: {}",
result[4].text()
);
}
#[test]
fn embedded_drop_cap_ignores_non_paragraph_neighbours() {
// Same geometry, but the preceding line is a short label rather than
// body text, so it must not be rewritten.
let label = TextLine {
items: vec![make_item_at("Fig. 2", 10.0, 90.0)],
y: 716.0,
page: 1,
adaptive_threshold: 0.10,
};
let second_line = TextLine {
items: vec![
make_item_at("T", 25.0, 72.0),
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![label, second_line], 10.0);
assert_eq!(result[0].text().trim(), "Fig. 2");
assert!(
result[1].text().starts_with('T'),
"cap stays put: {}",
result[1].text()
);
}
#[test]
fn embedded_drop_cap_keeps_a_space_for_standalone_word_caps() {
// Leading whitespace on the paragraph's first item marks the cap as
// a word of its own rather than the first letter of one.
let mut lead = make_item_at("long time ago in a galaxy far away", 10.0, 90.0);
lead.text = " long time ago in a galaxy far away".to_string();
let first_line = TextLine {
items: vec![lead],
y: 716.0,
page: 1,
adaptive_threshold: 0.10,
};
let second_line = TextLine {
items: vec![
make_item_at("A", 25.0, 72.0),
make_item_at("continued here with more body text", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![first_line, second_line], 10.0);
assert!(
result[0].text().starts_with("A long time ago"),
"standalone-word cap keeps one space: {}",
result[0].text()
);
}
#[test]
fn embedded_drop_cap_skips_hyphenation_continuation_targets() {
// The line above the RUN ends on a hyphen, so the run's topmost line
// resumes a split word rather than starting a paragraph. It sits at
// the paragraph margin (x=72), outside the cap's indent, so it is not
// part of the run itself.
let split_word = TextLine {
items: vec![make_item_at(
"mediasisallot ovat osa useimpien ylakou-",
10.0,
72.0,
)],
y: 728.0,
page: 1,
adaptive_threshold: 0.10,
};
let run_top = TextLine {
items: vec![make_item_at(
"lulaisten elamaa ja muuta tekstia",
10.0,
90.0,
)],
y: 714.0,
page: 1,
adaptive_threshold: 0.10,
};
let cap_line = TextLine {
items: vec![
make_item_at("E", 25.0, 72.0),
make_item_at("jatkuu tassa lisaa leipatekstia", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![split_word, run_top, cap_line], 10.0);
assert!(
result[1].text().starts_with("lulaisten"),
"a run resuming a split word must not receive the cap: {}",
result[1].text()
);
assert!(
result[2].text().starts_with('E'),
"cap stays put when no valid target exists: {}",
result[2].text()
);
}
#[test]
fn embedded_drop_cap_indent_is_not_a_word_boundary() {
// The paragraph's first line is indented to clear the cap, and that
// indent can arrive as leading whitespace. A mid-word cap must still
// join directly — "T HE recent" would be the defect this fixes.
let mut lead = make_item_at("HE recent development and more body text", 10.0, 90.0);
lead.text = " HE recent development and more body text".to_string();
let first_line = TextLine {
items: vec![lead],
y: 716.0,
page: 1,
adaptive_threshold: 0.10,
};
let cap_line = TextLine {
items: vec![
make_item_at("T", 25.0, 72.0),
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
],
y: 700.0,
page: 1,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![first_line, cap_line], 10.0);
assert!(
result[0].text().starts_with("THE recent"),
"indent must not be read as a word boundary: {}",
result[0].text()
);
}
#[test]
fn ends_sentence_rejects_abbreviations_and_markers() {
use super::ends_sentence;
assert!(ends_sentence("This completes the thought."));
assert!(ends_sentence("Is that so?"));
assert!(ends_sentence("Stop!"));
// Periods that do not end a sentence.
assert!(!ends_sentence("as shown in Fig."));
assert!(!ends_sentence("see e.g."));
// Non-ASCII abbreviations. The two-CHARACTER segment is the case
// that distinguishes a character count from a byte count: "пр" is
// 2 chars but 4 bytes, so a byte-based bound would reject it and
// read the line as a completed sentence.
assert!(!ends_sentence("и т.пр."));
assert!(!ends_sentence("см. т.е."));
assert!(!ends_sentence("napr. ú.d."));
assert!(!ends_sentence("reviewed by Dr."));
// Standalone enumerators, any case.
assert!(!ends_sentence("1."));
assert!(!ends_sentence("IV."));
assert!(!ends_sentence("ii."));
assert!(!ends_sentence("xii."));
// Numbers and numerals that genuinely end a sentence must count,
// or a legitimate drop-cap merge is blocked.
assert!(ends_sentence("The paper was published in 2020."));
assert!(ends_sentence("He scored 5."));
assert!(ends_sentence("after World War II."));
assert!(ends_sentence("the constant equals 3.14."));
assert!(ends_sentence("documented at example.com."));
assert!(!ends_sentence("a trailing clause with no period"));
}
#[test]
fn embedded_drop_cap_allows_first_paragraph_on_a_new_page() {
// The line two back is on the previous page, so its y is unrelated
// and must not be used as leading evidence.
let prev_page_tail = TextLine {
items: vec![make_item_at(
"tail of the previous page body text",
10.0,
90.0,
)],
y: 90.0,
page: 1,
adaptive_threshold: 0.10,
};
let first_line = TextLine {
items: vec![make_item_at(
"HE recent development which exchange",
10.0,
90.0,
)],
y: 716.0,
page: 2,
adaptive_threshold: 0.10,
};
let cap_line = TextLine {
items: vec![
make_item_at("T", 25.0, 72.0),
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
],
y: 700.0,
page: 2,
adaptive_threshold: 0.10,
};
let result = merge_drop_caps(vec![prev_page_tail, first_line, cap_line], 10.0);
assert!(
result[1].text().starts_with("THE recent"),
"a page break must not suppress the merge: {}",
result[1].text()
);
}
#[test]
fn test_merge_struct_tree_headings() {
// Two consecutive lines tagged as H2 via struct tree, same font size as body
+3 -7
View File
@@ -1,16 +1,12 @@
Reprinted with corrections from *The Bell System Technical Journal,* Vol. 27, pp. 379423, 623656, July, October, 1948.
## A Mathematical Theory of Communication
# A Mathematical Theory of Communication
### By C. E. SHANNON
## By C. E. SHANNON
INTRODUCTION
HE recent development of various methods of modulation such as PCM and PPM which exchange
# Tbandwidth for signal-to-noise ratio has intensified the interest in a general theory of communication. A
basis for such a theory is contained in the important papers of Nyquist¹ and Hartley² on this subject. In the present paper we will extend the theory to include a number of new factors, in particular the effect of noise in the channel, and the savings possible due to the statistical structure of the original message and due to the nature of the final destination of the information. The fundamental problem of communication is that of reproducing at one point either exactly or ap- proximately a message selected at another point. Frequently the messages have *meaning*; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one *selected from a set* of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design. If the number of messages in the set is finite then this number or any monotonic function of this number can be regarded as a measure of the information produced when one message is chosen from the set, all choices being equally likely. As was pointed out by Hartley the most natural choice is the logarithmic function. Although this definition must be generalized considerably when we consider the influence of the statistics of the message and when we have a continuous range of messages, we will in all cases use an essentially logarithmic measure. The logarithmic measure is more convenient for various reasons:
THE recent development of various methods of modulation such as PCM and PPM which exchange bandwidth for signal-to-noise ratio has intensified the interest in a general theory of communication. A basis for such a theory is contained in the important papers of Nyquist¹ and Hartley² on this subject. In the present paper we will extend the theory to include a number of new factors, in particular the effect of noise in the channel, and the savings possible due to the statistical structure of the original message and due to the nature of the final destination of the information. The fundamental problem of communication is that of reproducing at one point either exactly or ap- proximately a message selected at another point. Frequently the messages have *meaning*; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one *selected from a set* of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design. If the number of messages in the set is finite then this number or any monotonic function of this number can be regarded as a measure of the information produced when one message is chosen from the set, all choices being equally likely. As was pointed out by Hartley the most natural choice is the logarithmic function. Although this definition must be generalized considerably when we consider the influence of the statistics of the message and when we have a continuous range of messages, we will in all cases use an essentially logarithmic measure. The logarithmic measure is more convenient for various reasons:
1. It is practically more useful. Parameters of engineering importance such as time, bandwidth, number of relays, etc., tend to vary linearly with the logarithm of the number of possibilities. For example, adding one relay to a group doubles the number of possible states of the relays. It adds 1 to the base 2 logarithm of this number. Doubling the time roughly squares the number of possible messages, or doubles the logarithm, etc.
2. It is nearer to our intuitive feeling as to the proper measure. This is closely related to (1) since we in- tuitively measures entities by linear comparison with common standards. One feels, for example, that two punched cards should have twice the capacity of one for information storage, and two identical channels twice the capacity of one for transmitting information.