feat(extractor): stamp items with the font family name, not the resource tag (#415)
* feat(extractor): stamp items with the font family name, not the resource tag
TextItem::font carried the page's font resource name ("F2", "T22") —
an arbitrary per-page tag — even though both content-stream parsers
already resolve the /BaseFont family name for bold/italic detection at
every item-creation site. Stamp that resolved family name instead
("ABCDEF+CMMI10", "Courier"), from a single item_font_name helper so
the two parsers cannot drift.
One deliberate carve-out, documented on the helper: resource names
using Distiller's CID convention (C2_0, C0_1) are kept as-is, because
text_utils::is_cid_font keys on that prefix for micro-gap joining and
the family name carries no CID marker to replace it.
Consumers that match on font names start working against real names:
- Code detection (is_monospace_font) previously never fired against
opaque resource tags. It now does — so line classification also moves
from any-item matching to a majority-by-characters rule
(line_is_monospace): code lines are wholly monospace, while a lone
URL or identifier styled in a mono face inside a prose line must not
fence the surrounding sentence.
- Heading/body font grouping now merges resource aliases of the same
family instead of treating them as distinct fonts.
- Positioned-item output (--items-json and the bindings) reports real
face names.
Regression corpus: code-heavy manuals improve substantially (assembly
and C snippets previously emitted as prose now fence with line
structure preserved); remaining churn reviewed as improvements.
* fix(markdown): address review of font-name consumers
- Monotype is a foundry prefix on proportional faces (Monotype Corsiva,
Monotype Garamond); it must not satisfy is_monospace_font's generic
"mono" token. Regression tests pin both directions.
- Flush the pending code block before inserting a positioned table or
image, so a block that falls between two code lines cannot be emitted
ahead of code that precedes it in reading order; a code line after
the block reopens a new fence naturally.
* fix(markdown): emit sub-3-char mono fragments as plain text, not fences
A lone registered-trademark glyph or stray bullet set in a mono face is
not code; a fenced block containing one character reads as noise.
* fix(markdown): font-based code blocks open only at paragraph boundaries
HTML-to-PDF producers smear an inline code literal's mono style across
whole wrapped lines, so a prose paragraph can alternate body and mono
fonts line by line. Fencing those lines cut sentences in three: prose
head, fenced middle, prose tail. A mono-set line that continues an open
prose paragraph now stays prose; font-based blocks open at paragraph
boundaries (or continue an open block), and struct-tree Code roles are
honored unconditionally.
* refactor(markdown): drop paragraph-flush branch made unreachable by the boundary gate
The enclosing guard proves in_paragraph is false, so the nested flush
could never run; the guard and mono check collapse into one condition.
This commit is contained in:
@@ -561,7 +561,11 @@ pub(crate) fn extract_page_text_items(
|
||||
y,
|
||||
width,
|
||||
height: rendered_size,
|
||||
font: current_font.clone(),
|
||||
font: crate::extractor::fonts::item_font_name(
|
||||
¤t_font,
|
||||
base_font,
|
||||
)
|
||||
.to_string(),
|
||||
font_size: rendered_size,
|
||||
page: page_num,
|
||||
is_bold: is_bold_font(base_font) || desc_bold,
|
||||
@@ -745,7 +749,11 @@ pub(crate) fn extract_page_text_items(
|
||||
y,
|
||||
width,
|
||||
height: rendered_size,
|
||||
font: current_font.clone(),
|
||||
font: crate::extractor::fonts::item_font_name(
|
||||
¤t_font,
|
||||
base_font,
|
||||
)
|
||||
.to_string(),
|
||||
font_size: rendered_size,
|
||||
page: page_num,
|
||||
is_bold: is_bold_font(base_font) || desc_bold,
|
||||
@@ -852,7 +860,11 @@ pub(crate) fn extract_page_text_items(
|
||||
y,
|
||||
width,
|
||||
height: rendered_size,
|
||||
font: current_font.clone(),
|
||||
font: crate::extractor::fonts::item_font_name(
|
||||
¤t_font,
|
||||
base_font,
|
||||
)
|
||||
.to_string(),
|
||||
font_size: rendered_size,
|
||||
page: page_num,
|
||||
is_bold: is_bold_font(base_font) || desc_bold,
|
||||
@@ -1005,7 +1017,11 @@ pub(crate) fn extract_page_text_items(
|
||||
y,
|
||||
width,
|
||||
height: rendered_size,
|
||||
font: current_font.clone(),
|
||||
font: crate::extractor::fonts::item_font_name(
|
||||
¤t_font,
|
||||
base_font,
|
||||
)
|
||||
.to_string(),
|
||||
font_size: rendered_size,
|
||||
page: page_num,
|
||||
is_bold: is_bold_font(base_font) || desc_bold,
|
||||
|
||||
@@ -225,6 +225,26 @@ pub(crate) fn build_type3_scales(
|
||||
scales
|
||||
}
|
||||
|
||||
/// The name a `TextItem` carries for its font: the `/BaseFont` family name
|
||||
/// ("ABCDEF+CMMI10"), which identifies the actual face, rather than the
|
||||
/// arbitrary per-page resource tag ("F2").
|
||||
///
|
||||
/// Exception: resource names using Distiller's CID convention (`C2_0`,
|
||||
/// `C0_1`) are kept as-is — `text_utils::is_cid_font` keys on that prefix
|
||||
/// for micro-gap joining, and the family name carries no CID marker to
|
||||
/// replace it. This is a known, deliberate wart: `TextItem::font` is the
|
||||
/// face name except for this one producer convention. The clean fix is an
|
||||
/// explicit CID flag on `TextItem`, which touches its ~29 construction
|
||||
/// sites; do that migration when `TextItem` next changes shape, and delete
|
||||
/// this carve-out with it.
|
||||
pub(crate) fn item_font_name<'a>(resource_name: &'a str, base_font: &'a str) -> &'a str {
|
||||
if crate::text_utils::is_cid_font(resource_name) {
|
||||
resource_name
|
||||
} else {
|
||||
base_font
|
||||
}
|
||||
}
|
||||
|
||||
/// Parse font widths from a font dictionary, dispatching by Subtype
|
||||
pub(crate) fn parse_font_widths(
|
||||
doc: &Document,
|
||||
@@ -1664,6 +1684,17 @@ fn score_text(text: &str) -> i32 {
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
|
||||
#[test]
|
||||
fn item_font_name_prefers_family_over_resource_tag() {
|
||||
use super::item_font_name;
|
||||
assert_eq!(item_font_name("F2", "ABCDEF+CMMI10"), "ABCDEF+CMMI10");
|
||||
assert_eq!(item_font_name("T22", "Times-Roman"), "Times-Roman");
|
||||
// Distiller CID-convention resources keep the resource name:
|
||||
// is_cid_font keys on the C2_/C0_ prefix for micro-gap joining.
|
||||
assert_eq!(item_font_name("C2_0", "ABCDEE+SimSun"), "C2_0");
|
||||
assert_eq!(item_font_name("C0_1", "ABCDEE+MSMincho"), "C0_1");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn type3_scale_resolves_indirect_matrix_and_bbox_numbers() {
|
||||
use lopdf::{dictionary, Document, Object};
|
||||
|
||||
@@ -620,7 +620,11 @@ fn extract_form_xobject_text_inner(
|
||||
y,
|
||||
width,
|
||||
height: rendered_size,
|
||||
font: current_font.clone(),
|
||||
font: crate::extractor::fonts::item_font_name(
|
||||
¤t_font,
|
||||
base_font,
|
||||
)
|
||||
.to_string(),
|
||||
font_size: rendered_size,
|
||||
page: page_num,
|
||||
is_bold: is_bold_font(base_font) || desc_bold,
|
||||
@@ -775,7 +779,11 @@ fn extract_form_xobject_text_inner(
|
||||
y,
|
||||
width,
|
||||
height: rendered_size,
|
||||
font: current_font.clone(),
|
||||
font: crate::extractor::fonts::item_font_name(
|
||||
¤t_font,
|
||||
base_font,
|
||||
)
|
||||
.to_string(),
|
||||
font_size: rendered_size,
|
||||
page: page_num,
|
||||
is_bold: is_bold_font(base_font) || desc_bold,
|
||||
|
||||
Reference in New Issue
Block a user