Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
150dde604b | ||
|
|
fe65ec72e3 |
+1
-1
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "pdf-inspector"
|
||||
version = "0.1.7"
|
||||
version = "0.1.6"
|
||||
edition = "2021"
|
||||
autobins = false
|
||||
authors = ["Firecrawl Team"]
|
||||
|
||||
@@ -28,17 +28,17 @@ Evaluated on the [opendataloader-bench](https://github.com/opendataloader-projec
|
||||
|
||||
| Engine | Overall | Reading Order (NID) | Tables (TEDS) | Headings (MHS) | Speed (200 docs) |
|
||||
|---|---|---|---|---|---|
|
||||
| pdf-inspector | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
|
||||
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
|
||||
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
|
||||
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
|
||||
| pdf-inspector | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
|
||||
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
|
||||
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
|
||||
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
|
||||
|
||||
Results were refreshed on July 31, 2026, on an Apple M4 Pro. Engine versions were pdf-inspector 0.2.6, LiteParse 2.10.1, OpenDataLoader 2.2.1, PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.5. Speed is the median of five alternating or rotating complete corpus runs after an excluded warm-up run, with each parser processing documents sequentially in a single process.
|
||||
Results were refreshed on July 16, 2026, on an Apple M4 Pro. Engine versions were pdf-inspector 0.1.6, LiteParse 2.6.0, OpenDataLoader 2.1.1, PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.4. Speed is the median of three complete corpus runs.
|
||||
|
||||
The complete parser configuration, per-document predictions, evaluator output, and generated charts are available in the [reproducible results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
|
||||
For context, engines that use OCR or model-based document parsing (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus — pdf-inspector reaches the top of that range without either, in 2.8 seconds.
|
||||
|
||||
**Best fit:** Native-text PDFs where speed, reading order, and table structure matter. In this comparison, pdf-inspector delivered the higher overall, reading-order, and table scores, along with the fastest complete run. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.
|
||||
**Best fit:** Native-text PDFs where speed, reading order, and table structure matter. pdf-inspector delivered the highest overall, reading-order, and table scores, along with the fastest complete run in this benchmark. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.
|
||||
|
||||
Use the [paired benchmark harness](docs/benchmarking.md) to compare two local builds against the exact same corpus and evaluator revision.
|
||||
|
||||
@@ -215,7 +215,7 @@ wasm/ — Browser bindings (wasm-bindgen)
|
||||
## How classification works
|
||||
|
||||
1. Parse the xref table and page tree (no full object load)
|
||||
2. Select pages based on `ScanStrategy` (default: all pages with early exit)
|
||||
2. Select pages based on `ScanStrategy` (default: sample up to 8 evenly distributed pages)
|
||||
3. Look for `Tj`/`TJ` (text operators) and `Do` (image operators) in content streams
|
||||
4. Classify based on text operator presence across sampled pages
|
||||
|
||||
@@ -225,9 +225,9 @@ This detects 300+ page PDFs in milliseconds. The result includes `pages_needing_
|
||||
|
||||
| Strategy | Behavior | Best for |
|
||||
|---|---|---|
|
||||
| `EarlyExit` (default) | Scan all pages, stop on first non-text page | Pipelines routing TextBased PDFs to fast extraction |
|
||||
| `EarlyExit` | Scan all pages, stop on first non-text page | Pipelines routing TextBased PDFs to fast extraction |
|
||||
| `Full` | Scan all pages, no early exit | Accurate Mixed vs Scanned classification |
|
||||
| `Sample(n)` | Sample `n` evenly distributed pages (first, last, middle) | Very large PDFs where speed matters more than precision |
|
||||
| `Sample(n)` (default: `n = 8`) | Sample `n` evenly distributed pages (first, last, middle) | Very large PDFs where speed matters more than precision |
|
||||
| `Pages(vec)` | Only scan specific 1-indexed page numbers | When the caller knows which pages to check |
|
||||
|
||||
## Markdown output
|
||||
|
||||
@@ -30,14 +30,11 @@ overwrite one another.
|
||||
|
||||
## Published comparison protocol
|
||||
|
||||
The public benchmark table was refreshed on July 31, 2026, on an Apple M4 Pro
|
||||
using pdf-inspector 0.2.6, LiteParse 2.10.1, OpenDataLoader 2.2.1,
|
||||
PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.5. Every engine processed the same 200
|
||||
PDFs sequentially in a single process with OCR disabled. Reported speed is the
|
||||
median of five alternating or rotating complete corpus runs after an excluded
|
||||
warm-up run; quality scores come from the benchmark evaluator over all 200
|
||||
outputs. Raw timings, predictions, evaluations, and charts are available in the
|
||||
[results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
|
||||
The public benchmark table was refreshed on July 16, 2026, on an Apple M4 Pro
|
||||
using pdf-inspector 0.1.6, LiteParse 2.6.0, OpenDataLoader 2.1.1,
|
||||
PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.4. Every engine processed the same 200
|
||||
PDFs with OCR disabled. Reported speed is the median of three complete corpus
|
||||
runs; quality scores come from the benchmark evaluator over all 200 outputs.
|
||||
|
||||
## Optional backend evidence probe
|
||||
|
||||
|
||||
+6
-6
@@ -18,13 +18,13 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
|
||||
|
||||
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|
||||
|---|---|---|---|---|---|
|
||||
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
|
||||
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
|
||||
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
|
||||
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
|
||||
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
|
||||
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
|
||||
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
|
||||
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
|
||||
|
||||
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark), with raw timings and artifacts in the [results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
|
||||
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
|
||||
|
||||
## Install
|
||||
|
||||
|
||||
+7
-7
@@ -18,13 +18,13 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
|
||||
|
||||
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|
||||
|---|---|---|---|---|---|
|
||||
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
|
||||
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
|
||||
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
|
||||
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
|
||||
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
|
||||
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
|
||||
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
|
||||
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
|
||||
|
||||
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark), with raw timings and artifacts in the [results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
|
||||
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
|
||||
|
||||
## Install
|
||||
|
||||
@@ -182,4 +182,4 @@ Low-level detection functions are also available via the `detector` module (`det
|
||||
| `MarkdownOptions` | Configuration for Markdown formatting (page numbers, etc.) |
|
||||
| `PageMarkdown` | Per-page result: page (0-indexed), markdown, needs_ocr |
|
||||
| `PagesExtractionResult` | Per-page output + 1-indexed pages_with_tables / pages_with_columns / pages_needing_ocr, is_complex |
|
||||
| `PdfError` | `Io`, `Parse`, `Encrypted`, `InvalidStructure`, `NotAPdf` |
|
||||
| `PdfError` | `Io`, `Parse`, `Encrypted`, `InvalidOptions`, `InvalidStructure`, `NotAPdf` |
|
||||
|
||||
Generated
+1
-1
@@ -830,7 +830,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector"
|
||||
version = "0.1.7"
|
||||
version = "0.1.6"
|
||||
dependencies = [
|
||||
"env_logger",
|
||||
"log",
|
||||
|
||||
+6
-6
@@ -18,13 +18,13 @@ Built by [Firecrawl](https://firecrawl.dev) for hybrid OCR pipelines — extract
|
||||
|
||||
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|
||||
|---|---|---|---|---|---|
|
||||
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
|
||||
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
|
||||
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
|
||||
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
|
||||
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
|
||||
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
|
||||
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
|
||||
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
|
||||
|
||||
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark), with raw timings and artifacts in the [results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
|
||||
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
|
||||
|
||||
## Install
|
||||
|
||||
|
||||
+4
-4
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "@firecrawl/pdf-inspector",
|
||||
"version": "1.11.2",
|
||||
"version": "1.11.1",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
@@ -49,8 +49,8 @@
|
||||
"@napi-rs/cli": "^3.4.1"
|
||||
},
|
||||
"optionalDependencies": {
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.11.2",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.11.2",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.11.2"
|
||||
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.11.1",
|
||||
"@firecrawl/pdf-inspector-darwin-arm64": "1.11.1",
|
||||
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.11.1"
|
||||
}
|
||||
}
|
||||
|
||||
+1
-1
@@ -6,7 +6,7 @@ build-backend = "maturin"
|
||||
name = "pdf-inspector"
|
||||
# Bump this to publish to PyPI — CI publishes automatically when the version
|
||||
# changes on main (same flow as napi/package.json for npm).
|
||||
version = "0.2.6"
|
||||
version = "0.2.5"
|
||||
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
|
||||
readme = "docs/python.md"
|
||||
license = { text = "MIT" }
|
||||
|
||||
+6
-1
@@ -25,7 +25,7 @@ pub enum PdfType {
|
||||
/// Strategy for which pages to scan during detection
|
||||
#[derive(Debug, Clone)]
|
||||
pub enum ScanStrategy {
|
||||
/// Scan all pages, stop on first non-text page (current default).
|
||||
/// Scan all pages, stop on first non-text page.
|
||||
/// Best for pipelines that route TextBased PDFs to fast extraction.
|
||||
EarlyExit,
|
||||
/// Scan all pages, no early exit.
|
||||
@@ -205,6 +205,11 @@ pub(crate) fn detect_from_document(
|
||||
.collect();
|
||||
valid.sort();
|
||||
valid.dedup();
|
||||
if valid.is_empty() {
|
||||
return Err(PdfError::InvalidOptions(format!(
|
||||
"ScanStrategy::Pages contains no in-range page numbers for a {total_pages}-page PDF"
|
||||
)));
|
||||
}
|
||||
(valid, false)
|
||||
}
|
||||
};
|
||||
|
||||
@@ -1,837 +0,0 @@
|
||||
//! Built-in glyph metrics for the 14 standard PDF fonts.
|
||||
//!
|
||||
//! PDFs may omit `/Widths` for non-embedded base-14 fonts (Times, Helvetica,
|
||||
//! Courier, Symbol, ZapfDingbats); per the PDF spec the reader must supply
|
||||
//! the metrics. Without them every text item gets width 0, which breaks
|
||||
//! space synthesis, sub/superscript detection, and table column detection
|
||||
//! (common in 1990s dvips/Distiller output).
|
||||
//!
|
||||
//! Tables are generated from the Adobe Core 14 AFM files (via reportlab's
|
||||
//! `_fontdata`), keyed by Unicode char, sorted for binary search.
|
||||
//! Generator: scratchpad/gen_base14.py (session tooling, not checked in).
|
||||
|
||||
/// Width in 1000ths of an em for `c` in the given base-14 font, or `None`
|
||||
/// if the font is not one of the base 14 (after name normalization) or the
|
||||
/// char has no glyph in its AFM.
|
||||
pub(crate) fn base14_char_width(base_font: &str, c: char) -> Option<u16> {
|
||||
let table = base14_table(base_font)?;
|
||||
// AFM tables key visible glyphs only; alias the invisible variants the
|
||||
// cp1252 fallback can produce so they get the metric of their visible
|
||||
// counterpart instead of the generic default.
|
||||
let c = match c {
|
||||
'\u{00A0}' => ' ', // no-break space -> space
|
||||
'\u{00AD}' => '-', // soft hyphen -> hyphen
|
||||
_ => c,
|
||||
};
|
||||
table
|
||||
.binary_search_by_key(&c, |&(ch, _)| ch)
|
||||
.ok()
|
||||
.map(|i| table[i].1)
|
||||
}
|
||||
|
||||
/// True when the base font name normalizes to one of the standard 14 fonts.
|
||||
pub(crate) fn is_base14_font(base_font: &str) -> bool {
|
||||
base14_table(base_font).is_some()
|
||||
}
|
||||
|
||||
/// Code → Unicode through the font's BUILT-IN encoding, for the base-14
|
||||
/// fonts whose repertoire is not Latin (Symbol, ZapfDingbats). Their glyphs
|
||||
/// live at byte positions that have nothing to do with cp1252 (Symbol 0x61
|
||||
/// renders α, Zapf 0x21 renders ✁), so advance widths must be resolved
|
||||
/// through this mapping — the renderer draws these glyphs regardless of how
|
||||
/// the text decoder transliterates them. Returns `None` for the Latin text
|
||||
/// fonts, which follow standard single-byte encodings.
|
||||
pub(crate) fn builtin_encoding_char(base_font: &str, code: u8) -> Option<char> {
|
||||
let table = base14_table(base_font)?;
|
||||
let enc: &[(u8, char)] = if std::ptr::eq(table, SYMBOL) {
|
||||
SYMBOL_ENCODING
|
||||
} else if std::ptr::eq(table, ZAPFDINGBATS) {
|
||||
ZAPFDINGBATS_ENCODING
|
||||
} else {
|
||||
return None;
|
||||
};
|
||||
enc.binary_search_by_key(&code, |&(b, _)| b)
|
||||
.ok()
|
||||
.map(|i| enc[i].1)
|
||||
}
|
||||
|
||||
/// Map a BaseFont name (possibly subset-prefixed, e.g. "ABCDEF+Times-Bold",
|
||||
/// or a common alias like "Arial" / "TimesNewRomanPSMT") to its width table.
|
||||
fn base14_table(base_font: &str) -> Option<&'static [(char, u16)]> {
|
||||
// Strip subset prefix "ABCDEF+"
|
||||
let name = match base_font.split_once('+') {
|
||||
Some((prefix, rest))
|
||||
if prefix.len() == 6 && prefix.chars().all(|c| c.is_ascii_uppercase()) =>
|
||||
{
|
||||
rest
|
||||
}
|
||||
_ => base_font,
|
||||
};
|
||||
let lower = name.to_ascii_lowercase();
|
||||
let bold = lower.contains("bold");
|
||||
let italic = lower.contains("italic") || lower.contains("oblique");
|
||||
if lower.contains("courier") {
|
||||
return Some(match (bold, italic) {
|
||||
(false, false) => COURIER,
|
||||
(true, false) => COURIER_BOLD,
|
||||
(false, true) => COURIER_OBLIQUE,
|
||||
(true, true) => COURIER_BOLDOBLIQUE,
|
||||
});
|
||||
}
|
||||
if lower.contains("helvetica") || lower.contains("arial") {
|
||||
return Some(match (bold, italic) {
|
||||
(false, false) => HELVETICA,
|
||||
(true, false) => HELVETICA_BOLD,
|
||||
(false, true) => HELVETICA_OBLIQUE,
|
||||
(true, true) => HELVETICA_BOLDOBLIQUE,
|
||||
});
|
||||
}
|
||||
if lower.contains("times") {
|
||||
return Some(match (bold, italic) {
|
||||
(false, false) => TIMES_ROMAN,
|
||||
(true, false) => TIMES_BOLD,
|
||||
(false, true) => TIMES_ITALIC,
|
||||
(true, true) => TIMES_BOLDITALIC,
|
||||
});
|
||||
}
|
||||
// Symbol and ZapfDingbats have unique glyph repertoires, so only exact
|
||||
// names (plus the common MT/ITC aliases) qualify — a custom font that
|
||||
// merely mentions "Symbol" in its name must not get these metrics.
|
||||
match lower.as_str() {
|
||||
"zapfdingbats" | "dingbats" | "itczapfdingbats" | "zapfdingbatsitc" => {
|
||||
return Some(ZAPFDINGBATS)
|
||||
}
|
||||
"symbol" | "symbolmt" | "symbolitc" => return Some(SYMBOL),
|
||||
_ => {}
|
||||
}
|
||||
None
|
||||
}
|
||||
|
||||
#[rustfmt::skip]
|
||||
static COURIER: &[(char, u16)] = &[
|
||||
(' ', 600), ('!', 600), ('"', 600), ('#', 600), ('$', 600), ('%', 600),
|
||||
('&', 600), ('\'', 600), ('(', 600), (')', 600), ('*', 600), ('+', 600),
|
||||
(',', 600), ('-', 600), ('.', 600), ('/', 600), ('0', 600), ('1', 600),
|
||||
('2', 600), ('3', 600), ('4', 600), ('5', 600), ('6', 600), ('7', 600),
|
||||
('8', 600), ('9', 600), (':', 600), (';', 600), ('<', 600), ('=', 600),
|
||||
('>', 600), ('?', 600), ('@', 600), ('A', 600), ('B', 600), ('C', 600),
|
||||
('D', 600), ('E', 600), ('F', 600), ('G', 600), ('H', 600), ('I', 600),
|
||||
('J', 600), ('K', 600), ('L', 600), ('M', 600), ('N', 600), ('O', 600),
|
||||
('P', 600), ('Q', 600), ('R', 600), ('S', 600), ('T', 600), ('U', 600),
|
||||
('V', 600), ('W', 600), ('X', 600), ('Y', 600), ('Z', 600), ('[', 600),
|
||||
('\\', 600), (']', 600), ('^', 600), ('_', 600), ('`', 600), ('a', 600),
|
||||
('b', 600), ('c', 600), ('d', 600), ('e', 600), ('f', 600), ('g', 600),
|
||||
('h', 600), ('i', 600), ('j', 600), ('k', 600), ('l', 600), ('m', 600),
|
||||
('n', 600), ('o', 600), ('p', 600), ('q', 600), ('r', 600), ('s', 600),
|
||||
('t', 600), ('u', 600), ('v', 600), ('w', 600), ('x', 600), ('y', 600),
|
||||
('z', 600), ('{', 600), ('|', 600), ('}', 600), ('~', 600), ('\u{00A1}', 600),
|
||||
('\u{00A2}', 600), ('\u{00A3}', 600), ('\u{00A4}', 600), ('\u{00A5}', 600), ('\u{00A6}', 600), ('\u{00A7}', 600),
|
||||
('\u{00A8}', 600), ('\u{00A9}', 600), ('\u{00AA}', 600), ('\u{00AB}', 600), ('\u{00AC}', 600), ('\u{00AE}', 600),
|
||||
('\u{00AF}', 600), ('\u{00B0}', 600), ('\u{00B1}', 600), ('\u{00B2}', 600), ('\u{00B3}', 600), ('\u{00B4}', 600),
|
||||
('\u{00B5}', 600), ('\u{00B6}', 600), ('\u{00B7}', 600), ('\u{00B8}', 600), ('\u{00B9}', 600), ('\u{00BA}', 600),
|
||||
('\u{00BB}', 600), ('\u{00BC}', 600), ('\u{00BD}', 600), ('\u{00BE}', 600), ('\u{00BF}', 600), ('\u{00C0}', 600),
|
||||
('\u{00C1}', 600), ('\u{00C2}', 600), ('\u{00C3}', 600), ('\u{00C4}', 600), ('\u{00C5}', 600), ('\u{00C6}', 600),
|
||||
('\u{00C7}', 600), ('\u{00C8}', 600), ('\u{00C9}', 600), ('\u{00CA}', 600), ('\u{00CB}', 600), ('\u{00CC}', 600),
|
||||
('\u{00CD}', 600), ('\u{00CE}', 600), ('\u{00CF}', 600), ('\u{00D0}', 600), ('\u{00D1}', 600), ('\u{00D2}', 600),
|
||||
('\u{00D3}', 600), ('\u{00D4}', 600), ('\u{00D5}', 600), ('\u{00D6}', 600), ('\u{00D7}', 600), ('\u{00D8}', 600),
|
||||
('\u{00D9}', 600), ('\u{00DA}', 600), ('\u{00DB}', 600), ('\u{00DC}', 600), ('\u{00DD}', 600), ('\u{00DE}', 600),
|
||||
('\u{00DF}', 600), ('\u{00E0}', 600), ('\u{00E1}', 600), ('\u{00E2}', 600), ('\u{00E3}', 600), ('\u{00E4}', 600),
|
||||
('\u{00E5}', 600), ('\u{00E6}', 600), ('\u{00E7}', 600), ('\u{00E8}', 600), ('\u{00E9}', 600), ('\u{00EA}', 600),
|
||||
('\u{00EB}', 600), ('\u{00EC}', 600), ('\u{00ED}', 600), ('\u{00EE}', 600), ('\u{00EF}', 600), ('\u{00F0}', 600),
|
||||
('\u{00F1}', 600), ('\u{00F2}', 600), ('\u{00F3}', 600), ('\u{00F4}', 600), ('\u{00F5}', 600), ('\u{00F6}', 600),
|
||||
('\u{00F7}', 600), ('\u{00F8}', 600), ('\u{00F9}', 600), ('\u{00FA}', 600), ('\u{00FB}', 600), ('\u{00FC}', 600),
|
||||
('\u{00FD}', 600), ('\u{00FE}', 600), ('\u{00FF}', 600), ('\u{0131}', 600), ('\u{0141}', 600), ('\u{0142}', 600),
|
||||
('\u{0152}', 600), ('\u{0153}', 600), ('\u{0160}', 600), ('\u{0161}', 600), ('\u{0178}', 600), ('\u{017D}', 600),
|
||||
('\u{017E}', 600), ('\u{0192}', 600), ('\u{02C6}', 600), ('\u{02C7}', 600), ('\u{02D8}', 600), ('\u{02D9}', 600),
|
||||
('\u{02DA}', 600), ('\u{02DB}', 600), ('\u{02DC}', 600), ('\u{02DD}', 600), ('\u{2013}', 600), ('\u{2014}', 600),
|
||||
('\u{2018}', 600), ('\u{2019}', 600), ('\u{201A}', 600), ('\u{201C}', 600), ('\u{201D}', 600), ('\u{201E}', 600),
|
||||
('\u{2020}', 600), ('\u{2021}', 600), ('\u{2022}', 600), ('\u{2026}', 600), ('\u{2030}', 600), ('\u{2039}', 600),
|
||||
('\u{203A}', 600), ('\u{2044}', 600), ('\u{20AC}', 600), ('\u{2122}', 600), ('\u{2212}', 600), ('\u{FB01}', 600),
|
||||
('\u{FB02}', 600),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static COURIER_BOLD: &[(char, u16)] = &[
|
||||
(' ', 600), ('!', 600), ('"', 600), ('#', 600), ('$', 600), ('%', 600),
|
||||
('&', 600), ('\'', 600), ('(', 600), (')', 600), ('*', 600), ('+', 600),
|
||||
(',', 600), ('-', 600), ('.', 600), ('/', 600), ('0', 600), ('1', 600),
|
||||
('2', 600), ('3', 600), ('4', 600), ('5', 600), ('6', 600), ('7', 600),
|
||||
('8', 600), ('9', 600), (':', 600), (';', 600), ('<', 600), ('=', 600),
|
||||
('>', 600), ('?', 600), ('@', 600), ('A', 600), ('B', 600), ('C', 600),
|
||||
('D', 600), ('E', 600), ('F', 600), ('G', 600), ('H', 600), ('I', 600),
|
||||
('J', 600), ('K', 600), ('L', 600), ('M', 600), ('N', 600), ('O', 600),
|
||||
('P', 600), ('Q', 600), ('R', 600), ('S', 600), ('T', 600), ('U', 600),
|
||||
('V', 600), ('W', 600), ('X', 600), ('Y', 600), ('Z', 600), ('[', 600),
|
||||
('\\', 600), (']', 600), ('^', 600), ('_', 600), ('`', 600), ('a', 600),
|
||||
('b', 600), ('c', 600), ('d', 600), ('e', 600), ('f', 600), ('g', 600),
|
||||
('h', 600), ('i', 600), ('j', 600), ('k', 600), ('l', 600), ('m', 600),
|
||||
('n', 600), ('o', 600), ('p', 600), ('q', 600), ('r', 600), ('s', 600),
|
||||
('t', 600), ('u', 600), ('v', 600), ('w', 600), ('x', 600), ('y', 600),
|
||||
('z', 600), ('{', 600), ('|', 600), ('}', 600), ('~', 600), ('\u{00A1}', 600),
|
||||
('\u{00A2}', 600), ('\u{00A3}', 600), ('\u{00A4}', 600), ('\u{00A5}', 600), ('\u{00A6}', 600), ('\u{00A7}', 600),
|
||||
('\u{00A8}', 600), ('\u{00A9}', 600), ('\u{00AA}', 600), ('\u{00AB}', 600), ('\u{00AC}', 600), ('\u{00AE}', 600),
|
||||
('\u{00AF}', 600), ('\u{00B0}', 600), ('\u{00B1}', 600), ('\u{00B2}', 600), ('\u{00B3}', 600), ('\u{00B4}', 600),
|
||||
('\u{00B5}', 600), ('\u{00B6}', 600), ('\u{00B7}', 600), ('\u{00B8}', 600), ('\u{00B9}', 600), ('\u{00BA}', 600),
|
||||
('\u{00BB}', 600), ('\u{00BC}', 600), ('\u{00BD}', 600), ('\u{00BE}', 600), ('\u{00BF}', 600), ('\u{00C0}', 600),
|
||||
('\u{00C1}', 600), ('\u{00C2}', 600), ('\u{00C3}', 600), ('\u{00C4}', 600), ('\u{00C5}', 600), ('\u{00C6}', 600),
|
||||
('\u{00C7}', 600), ('\u{00C8}', 600), ('\u{00C9}', 600), ('\u{00CA}', 600), ('\u{00CB}', 600), ('\u{00CC}', 600),
|
||||
('\u{00CD}', 600), ('\u{00CE}', 600), ('\u{00CF}', 600), ('\u{00D0}', 600), ('\u{00D1}', 600), ('\u{00D2}', 600),
|
||||
('\u{00D3}', 600), ('\u{00D4}', 600), ('\u{00D5}', 600), ('\u{00D6}', 600), ('\u{00D7}', 600), ('\u{00D8}', 600),
|
||||
('\u{00D9}', 600), ('\u{00DA}', 600), ('\u{00DB}', 600), ('\u{00DC}', 600), ('\u{00DD}', 600), ('\u{00DE}', 600),
|
||||
('\u{00DF}', 600), ('\u{00E0}', 600), ('\u{00E1}', 600), ('\u{00E2}', 600), ('\u{00E3}', 600), ('\u{00E4}', 600),
|
||||
('\u{00E5}', 600), ('\u{00E6}', 600), ('\u{00E7}', 600), ('\u{00E8}', 600), ('\u{00E9}', 600), ('\u{00EA}', 600),
|
||||
('\u{00EB}', 600), ('\u{00EC}', 600), ('\u{00ED}', 600), ('\u{00EE}', 600), ('\u{00EF}', 600), ('\u{00F0}', 600),
|
||||
('\u{00F1}', 600), ('\u{00F2}', 600), ('\u{00F3}', 600), ('\u{00F4}', 600), ('\u{00F5}', 600), ('\u{00F6}', 600),
|
||||
('\u{00F7}', 600), ('\u{00F8}', 600), ('\u{00F9}', 600), ('\u{00FA}', 600), ('\u{00FB}', 600), ('\u{00FC}', 600),
|
||||
('\u{00FD}', 600), ('\u{00FE}', 600), ('\u{00FF}', 600), ('\u{0131}', 600), ('\u{0141}', 600), ('\u{0142}', 600),
|
||||
('\u{0152}', 600), ('\u{0153}', 600), ('\u{0160}', 600), ('\u{0161}', 600), ('\u{0178}', 600), ('\u{017D}', 600),
|
||||
('\u{017E}', 600), ('\u{0192}', 600), ('\u{02C6}', 600), ('\u{02C7}', 600), ('\u{02D8}', 600), ('\u{02D9}', 600),
|
||||
('\u{02DA}', 600), ('\u{02DB}', 600), ('\u{02DC}', 600), ('\u{02DD}', 600), ('\u{2013}', 600), ('\u{2014}', 600),
|
||||
('\u{2018}', 600), ('\u{2019}', 600), ('\u{201A}', 600), ('\u{201C}', 600), ('\u{201D}', 600), ('\u{201E}', 600),
|
||||
('\u{2020}', 600), ('\u{2021}', 600), ('\u{2022}', 600), ('\u{2026}', 600), ('\u{2030}', 600), ('\u{2039}', 600),
|
||||
('\u{203A}', 600), ('\u{2044}', 600), ('\u{20AC}', 600), ('\u{2122}', 600), ('\u{2212}', 600), ('\u{FB01}', 600),
|
||||
('\u{FB02}', 600),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static COURIER_OBLIQUE: &[(char, u16)] = &[
|
||||
(' ', 600), ('!', 600), ('"', 600), ('#', 600), ('$', 600), ('%', 600),
|
||||
('&', 600), ('\'', 600), ('(', 600), (')', 600), ('*', 600), ('+', 600),
|
||||
(',', 600), ('-', 600), ('.', 600), ('/', 600), ('0', 600), ('1', 600),
|
||||
('2', 600), ('3', 600), ('4', 600), ('5', 600), ('6', 600), ('7', 600),
|
||||
('8', 600), ('9', 600), (':', 600), (';', 600), ('<', 600), ('=', 600),
|
||||
('>', 600), ('?', 600), ('@', 600), ('A', 600), ('B', 600), ('C', 600),
|
||||
('D', 600), ('E', 600), ('F', 600), ('G', 600), ('H', 600), ('I', 600),
|
||||
('J', 600), ('K', 600), ('L', 600), ('M', 600), ('N', 600), ('O', 600),
|
||||
('P', 600), ('Q', 600), ('R', 600), ('S', 600), ('T', 600), ('U', 600),
|
||||
('V', 600), ('W', 600), ('X', 600), ('Y', 600), ('Z', 600), ('[', 600),
|
||||
('\\', 600), (']', 600), ('^', 600), ('_', 600), ('`', 600), ('a', 600),
|
||||
('b', 600), ('c', 600), ('d', 600), ('e', 600), ('f', 600), ('g', 600),
|
||||
('h', 600), ('i', 600), ('j', 600), ('k', 600), ('l', 600), ('m', 600),
|
||||
('n', 600), ('o', 600), ('p', 600), ('q', 600), ('r', 600), ('s', 600),
|
||||
('t', 600), ('u', 600), ('v', 600), ('w', 600), ('x', 600), ('y', 600),
|
||||
('z', 600), ('{', 600), ('|', 600), ('}', 600), ('~', 600), ('\u{00A1}', 600),
|
||||
('\u{00A2}', 600), ('\u{00A3}', 600), ('\u{00A4}', 600), ('\u{00A5}', 600), ('\u{00A6}', 600), ('\u{00A7}', 600),
|
||||
('\u{00A8}', 600), ('\u{00A9}', 600), ('\u{00AA}', 600), ('\u{00AB}', 600), ('\u{00AC}', 600), ('\u{00AE}', 600),
|
||||
('\u{00AF}', 600), ('\u{00B0}', 600), ('\u{00B1}', 600), ('\u{00B2}', 600), ('\u{00B3}', 600), ('\u{00B4}', 600),
|
||||
('\u{00B5}', 600), ('\u{00B6}', 600), ('\u{00B7}', 600), ('\u{00B8}', 600), ('\u{00B9}', 600), ('\u{00BA}', 600),
|
||||
('\u{00BB}', 600), ('\u{00BC}', 600), ('\u{00BD}', 600), ('\u{00BE}', 600), ('\u{00BF}', 600), ('\u{00C0}', 600),
|
||||
('\u{00C1}', 600), ('\u{00C2}', 600), ('\u{00C3}', 600), ('\u{00C4}', 600), ('\u{00C5}', 600), ('\u{00C6}', 600),
|
||||
('\u{00C7}', 600), ('\u{00C8}', 600), ('\u{00C9}', 600), ('\u{00CA}', 600), ('\u{00CB}', 600), ('\u{00CC}', 600),
|
||||
('\u{00CD}', 600), ('\u{00CE}', 600), ('\u{00CF}', 600), ('\u{00D0}', 600), ('\u{00D1}', 600), ('\u{00D2}', 600),
|
||||
('\u{00D3}', 600), ('\u{00D4}', 600), ('\u{00D5}', 600), ('\u{00D6}', 600), ('\u{00D7}', 600), ('\u{00D8}', 600),
|
||||
('\u{00D9}', 600), ('\u{00DA}', 600), ('\u{00DB}', 600), ('\u{00DC}', 600), ('\u{00DD}', 600), ('\u{00DE}', 600),
|
||||
('\u{00DF}', 600), ('\u{00E0}', 600), ('\u{00E1}', 600), ('\u{00E2}', 600), ('\u{00E3}', 600), ('\u{00E4}', 600),
|
||||
('\u{00E5}', 600), ('\u{00E6}', 600), ('\u{00E7}', 600), ('\u{00E8}', 600), ('\u{00E9}', 600), ('\u{00EA}', 600),
|
||||
('\u{00EB}', 600), ('\u{00EC}', 600), ('\u{00ED}', 600), ('\u{00EE}', 600), ('\u{00EF}', 600), ('\u{00F0}', 600),
|
||||
('\u{00F1}', 600), ('\u{00F2}', 600), ('\u{00F3}', 600), ('\u{00F4}', 600), ('\u{00F5}', 600), ('\u{00F6}', 600),
|
||||
('\u{00F7}', 600), ('\u{00F8}', 600), ('\u{00F9}', 600), ('\u{00FA}', 600), ('\u{00FB}', 600), ('\u{00FC}', 600),
|
||||
('\u{00FD}', 600), ('\u{00FE}', 600), ('\u{00FF}', 600), ('\u{0131}', 600), ('\u{0141}', 600), ('\u{0142}', 600),
|
||||
('\u{0152}', 600), ('\u{0153}', 600), ('\u{0160}', 600), ('\u{0161}', 600), ('\u{0178}', 600), ('\u{017D}', 600),
|
||||
('\u{017E}', 600), ('\u{0192}', 600), ('\u{02C6}', 600), ('\u{02C7}', 600), ('\u{02D8}', 600), ('\u{02D9}', 600),
|
||||
('\u{02DA}', 600), ('\u{02DB}', 600), ('\u{02DC}', 600), ('\u{02DD}', 600), ('\u{2013}', 600), ('\u{2014}', 600),
|
||||
('\u{2018}', 600), ('\u{2019}', 600), ('\u{201A}', 600), ('\u{201C}', 600), ('\u{201D}', 600), ('\u{201E}', 600),
|
||||
('\u{2020}', 600), ('\u{2021}', 600), ('\u{2022}', 600), ('\u{2026}', 600), ('\u{2030}', 600), ('\u{2039}', 600),
|
||||
('\u{203A}', 600), ('\u{2044}', 600), ('\u{20AC}', 600), ('\u{2122}', 600), ('\u{2212}', 600), ('\u{FB01}', 600),
|
||||
('\u{FB02}', 600),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static COURIER_BOLDOBLIQUE: &[(char, u16)] = &[
|
||||
(' ', 600), ('!', 600), ('"', 600), ('#', 600), ('$', 600), ('%', 600),
|
||||
('&', 600), ('\'', 600), ('(', 600), (')', 600), ('*', 600), ('+', 600),
|
||||
(',', 600), ('-', 600), ('.', 600), ('/', 600), ('0', 600), ('1', 600),
|
||||
('2', 600), ('3', 600), ('4', 600), ('5', 600), ('6', 600), ('7', 600),
|
||||
('8', 600), ('9', 600), (':', 600), (';', 600), ('<', 600), ('=', 600),
|
||||
('>', 600), ('?', 600), ('@', 600), ('A', 600), ('B', 600), ('C', 600),
|
||||
('D', 600), ('E', 600), ('F', 600), ('G', 600), ('H', 600), ('I', 600),
|
||||
('J', 600), ('K', 600), ('L', 600), ('M', 600), ('N', 600), ('O', 600),
|
||||
('P', 600), ('Q', 600), ('R', 600), ('S', 600), ('T', 600), ('U', 600),
|
||||
('V', 600), ('W', 600), ('X', 600), ('Y', 600), ('Z', 600), ('[', 600),
|
||||
('\\', 600), (']', 600), ('^', 600), ('_', 600), ('`', 600), ('a', 600),
|
||||
('b', 600), ('c', 600), ('d', 600), ('e', 600), ('f', 600), ('g', 600),
|
||||
('h', 600), ('i', 600), ('j', 600), ('k', 600), ('l', 600), ('m', 600),
|
||||
('n', 600), ('o', 600), ('p', 600), ('q', 600), ('r', 600), ('s', 600),
|
||||
('t', 600), ('u', 600), ('v', 600), ('w', 600), ('x', 600), ('y', 600),
|
||||
('z', 600), ('{', 600), ('|', 600), ('}', 600), ('~', 600), ('\u{00A1}', 600),
|
||||
('\u{00A2}', 600), ('\u{00A3}', 600), ('\u{00A4}', 600), ('\u{00A5}', 600), ('\u{00A6}', 600), ('\u{00A7}', 600),
|
||||
('\u{00A8}', 600), ('\u{00A9}', 600), ('\u{00AA}', 600), ('\u{00AB}', 600), ('\u{00AC}', 600), ('\u{00AE}', 600),
|
||||
('\u{00AF}', 600), ('\u{00B0}', 600), ('\u{00B1}', 600), ('\u{00B2}', 600), ('\u{00B3}', 600), ('\u{00B4}', 600),
|
||||
('\u{00B5}', 600), ('\u{00B6}', 600), ('\u{00B7}', 600), ('\u{00B8}', 600), ('\u{00B9}', 600), ('\u{00BA}', 600),
|
||||
('\u{00BB}', 600), ('\u{00BC}', 600), ('\u{00BD}', 600), ('\u{00BE}', 600), ('\u{00BF}', 600), ('\u{00C0}', 600),
|
||||
('\u{00C1}', 600), ('\u{00C2}', 600), ('\u{00C3}', 600), ('\u{00C4}', 600), ('\u{00C5}', 600), ('\u{00C6}', 600),
|
||||
('\u{00C7}', 600), ('\u{00C8}', 600), ('\u{00C9}', 600), ('\u{00CA}', 600), ('\u{00CB}', 600), ('\u{00CC}', 600),
|
||||
('\u{00CD}', 600), ('\u{00CE}', 600), ('\u{00CF}', 600), ('\u{00D0}', 600), ('\u{00D1}', 600), ('\u{00D2}', 600),
|
||||
('\u{00D3}', 600), ('\u{00D4}', 600), ('\u{00D5}', 600), ('\u{00D6}', 600), ('\u{00D7}', 600), ('\u{00D8}', 600),
|
||||
('\u{00D9}', 600), ('\u{00DA}', 600), ('\u{00DB}', 600), ('\u{00DC}', 600), ('\u{00DD}', 600), ('\u{00DE}', 600),
|
||||
('\u{00DF}', 600), ('\u{00E0}', 600), ('\u{00E1}', 600), ('\u{00E2}', 600), ('\u{00E3}', 600), ('\u{00E4}', 600),
|
||||
('\u{00E5}', 600), ('\u{00E6}', 600), ('\u{00E7}', 600), ('\u{00E8}', 600), ('\u{00E9}', 600), ('\u{00EA}', 600),
|
||||
('\u{00EB}', 600), ('\u{00EC}', 600), ('\u{00ED}', 600), ('\u{00EE}', 600), ('\u{00EF}', 600), ('\u{00F0}', 600),
|
||||
('\u{00F1}', 600), ('\u{00F2}', 600), ('\u{00F3}', 600), ('\u{00F4}', 600), ('\u{00F5}', 600), ('\u{00F6}', 600),
|
||||
('\u{00F7}', 600), ('\u{00F8}', 600), ('\u{00F9}', 600), ('\u{00FA}', 600), ('\u{00FB}', 600), ('\u{00FC}', 600),
|
||||
('\u{00FD}', 600), ('\u{00FE}', 600), ('\u{00FF}', 600), ('\u{0131}', 600), ('\u{0141}', 600), ('\u{0142}', 600),
|
||||
('\u{0152}', 600), ('\u{0153}', 600), ('\u{0160}', 600), ('\u{0161}', 600), ('\u{0178}', 600), ('\u{017D}', 600),
|
||||
('\u{017E}', 600), ('\u{0192}', 600), ('\u{02C6}', 600), ('\u{02C7}', 600), ('\u{02D8}', 600), ('\u{02D9}', 600),
|
||||
('\u{02DA}', 600), ('\u{02DB}', 600), ('\u{02DC}', 600), ('\u{02DD}', 600), ('\u{2013}', 600), ('\u{2014}', 600),
|
||||
('\u{2018}', 600), ('\u{2019}', 600), ('\u{201A}', 600), ('\u{201C}', 600), ('\u{201D}', 600), ('\u{201E}', 600),
|
||||
('\u{2020}', 600), ('\u{2021}', 600), ('\u{2022}', 600), ('\u{2026}', 600), ('\u{2030}', 600), ('\u{2039}', 600),
|
||||
('\u{203A}', 600), ('\u{2044}', 600), ('\u{20AC}', 600), ('\u{2122}', 600), ('\u{2212}', 600), ('\u{FB01}', 600),
|
||||
('\u{FB02}', 600),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static HELVETICA: &[(char, u16)] = &[
|
||||
(' ', 278), ('!', 278), ('"', 355), ('#', 556), ('$', 556), ('%', 889),
|
||||
('&', 667), ('\'', 191), ('(', 333), (')', 333), ('*', 389), ('+', 584),
|
||||
(',', 278), ('-', 333), ('.', 278), ('/', 278), ('0', 556), ('1', 556),
|
||||
('2', 556), ('3', 556), ('4', 556), ('5', 556), ('6', 556), ('7', 556),
|
||||
('8', 556), ('9', 556), (':', 278), (';', 278), ('<', 584), ('=', 584),
|
||||
('>', 584), ('?', 556), ('@', 1015), ('A', 667), ('B', 667), ('C', 722),
|
||||
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 722), ('I', 278),
|
||||
('J', 500), ('K', 667), ('L', 556), ('M', 833), ('N', 722), ('O', 778),
|
||||
('P', 667), ('Q', 778), ('R', 722), ('S', 667), ('T', 611), ('U', 722),
|
||||
('V', 667), ('W', 944), ('X', 667), ('Y', 667), ('Z', 611), ('[', 278),
|
||||
('\\', 278), (']', 278), ('^', 469), ('_', 556), ('`', 333), ('a', 556),
|
||||
('b', 556), ('c', 500), ('d', 556), ('e', 556), ('f', 278), ('g', 556),
|
||||
('h', 556), ('i', 222), ('j', 222), ('k', 500), ('l', 222), ('m', 833),
|
||||
('n', 556), ('o', 556), ('p', 556), ('q', 556), ('r', 333), ('s', 500),
|
||||
('t', 278), ('u', 556), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
|
||||
('z', 500), ('{', 334), ('|', 260), ('}', 334), ('~', 584), ('\u{00A1}', 333),
|
||||
('\u{00A2}', 556), ('\u{00A3}', 556), ('\u{00A4}', 556), ('\u{00A5}', 556), ('\u{00A6}', 260), ('\u{00A7}', 556),
|
||||
('\u{00A8}', 333), ('\u{00A9}', 737), ('\u{00AA}', 370), ('\u{00AB}', 556), ('\u{00AC}', 584), ('\u{00AE}', 737),
|
||||
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 584), ('\u{00B2}', 333), ('\u{00B3}', 333), ('\u{00B4}', 333),
|
||||
('\u{00B5}', 556), ('\u{00B6}', 537), ('\u{00B7}', 278), ('\u{00B8}', 333), ('\u{00B9}', 333), ('\u{00BA}', 365),
|
||||
('\u{00BB}', 556), ('\u{00BC}', 834), ('\u{00BD}', 834), ('\u{00BE}', 834), ('\u{00BF}', 611), ('\u{00C0}', 667),
|
||||
('\u{00C1}', 667), ('\u{00C2}', 667), ('\u{00C3}', 667), ('\u{00C4}', 667), ('\u{00C5}', 667), ('\u{00C6}', 1000),
|
||||
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 278),
|
||||
('\u{00CD}', 278), ('\u{00CE}', 278), ('\u{00CF}', 278), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
|
||||
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 584), ('\u{00D8}', 778),
|
||||
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 667), ('\u{00DE}', 667),
|
||||
('\u{00DF}', 611), ('\u{00E0}', 556), ('\u{00E1}', 556), ('\u{00E2}', 556), ('\u{00E3}', 556), ('\u{00E4}', 556),
|
||||
('\u{00E5}', 556), ('\u{00E6}', 889), ('\u{00E7}', 500), ('\u{00E8}', 556), ('\u{00E9}', 556), ('\u{00EA}', 556),
|
||||
('\u{00EB}', 556), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 556),
|
||||
('\u{00F1}', 556), ('\u{00F2}', 556), ('\u{00F3}', 556), ('\u{00F4}', 556), ('\u{00F5}', 556), ('\u{00F6}', 556),
|
||||
('\u{00F7}', 584), ('\u{00F8}', 611), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
|
||||
('\u{00FD}', 500), ('\u{00FE}', 556), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 556), ('\u{0142}', 222),
|
||||
('\u{0152}', 1000), ('\u{0153}', 944), ('\u{0160}', 667), ('\u{0161}', 500), ('\u{0178}', 667), ('\u{017D}', 611),
|
||||
('\u{017E}', 500), ('\u{0192}', 556), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
|
||||
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 556), ('\u{2014}', 1000),
|
||||
('\u{2018}', 222), ('\u{2019}', 222), ('\u{201A}', 222), ('\u{201C}', 333), ('\u{201D}', 333), ('\u{201E}', 333),
|
||||
('\u{2020}', 556), ('\u{2021}', 556), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
|
||||
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 556), ('\u{2122}', 1000), ('\u{2212}', 584), ('\u{FB01}', 500),
|
||||
('\u{FB02}', 500),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static HELVETICA_BOLD: &[(char, u16)] = &[
|
||||
(' ', 278), ('!', 333), ('"', 474), ('#', 556), ('$', 556), ('%', 889),
|
||||
('&', 722), ('\'', 238), ('(', 333), (')', 333), ('*', 389), ('+', 584),
|
||||
(',', 278), ('-', 333), ('.', 278), ('/', 278), ('0', 556), ('1', 556),
|
||||
('2', 556), ('3', 556), ('4', 556), ('5', 556), ('6', 556), ('7', 556),
|
||||
('8', 556), ('9', 556), (':', 333), (';', 333), ('<', 584), ('=', 584),
|
||||
('>', 584), ('?', 611), ('@', 975), ('A', 722), ('B', 722), ('C', 722),
|
||||
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 722), ('I', 278),
|
||||
('J', 556), ('K', 722), ('L', 611), ('M', 833), ('N', 722), ('O', 778),
|
||||
('P', 667), ('Q', 778), ('R', 722), ('S', 667), ('T', 611), ('U', 722),
|
||||
('V', 667), ('W', 944), ('X', 667), ('Y', 667), ('Z', 611), ('[', 333),
|
||||
('\\', 278), (']', 333), ('^', 584), ('_', 556), ('`', 333), ('a', 556),
|
||||
('b', 611), ('c', 556), ('d', 611), ('e', 556), ('f', 333), ('g', 611),
|
||||
('h', 611), ('i', 278), ('j', 278), ('k', 556), ('l', 278), ('m', 889),
|
||||
('n', 611), ('o', 611), ('p', 611), ('q', 611), ('r', 389), ('s', 556),
|
||||
('t', 333), ('u', 611), ('v', 556), ('w', 778), ('x', 556), ('y', 556),
|
||||
('z', 500), ('{', 389), ('|', 280), ('}', 389), ('~', 584), ('\u{00A1}', 333),
|
||||
('\u{00A2}', 556), ('\u{00A3}', 556), ('\u{00A4}', 556), ('\u{00A5}', 556), ('\u{00A6}', 280), ('\u{00A7}', 556),
|
||||
('\u{00A8}', 333), ('\u{00A9}', 737), ('\u{00AA}', 370), ('\u{00AB}', 556), ('\u{00AC}', 584), ('\u{00AE}', 737),
|
||||
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 584), ('\u{00B2}', 333), ('\u{00B3}', 333), ('\u{00B4}', 333),
|
||||
('\u{00B5}', 611), ('\u{00B6}', 556), ('\u{00B7}', 278), ('\u{00B8}', 333), ('\u{00B9}', 333), ('\u{00BA}', 365),
|
||||
('\u{00BB}', 556), ('\u{00BC}', 834), ('\u{00BD}', 834), ('\u{00BE}', 834), ('\u{00BF}', 611), ('\u{00C0}', 722),
|
||||
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 1000),
|
||||
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 278),
|
||||
('\u{00CD}', 278), ('\u{00CE}', 278), ('\u{00CF}', 278), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
|
||||
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 584), ('\u{00D8}', 778),
|
||||
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 667), ('\u{00DE}', 667),
|
||||
('\u{00DF}', 611), ('\u{00E0}', 556), ('\u{00E1}', 556), ('\u{00E2}', 556), ('\u{00E3}', 556), ('\u{00E4}', 556),
|
||||
('\u{00E5}', 556), ('\u{00E6}', 889), ('\u{00E7}', 556), ('\u{00E8}', 556), ('\u{00E9}', 556), ('\u{00EA}', 556),
|
||||
('\u{00EB}', 556), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 611),
|
||||
('\u{00F1}', 611), ('\u{00F2}', 611), ('\u{00F3}', 611), ('\u{00F4}', 611), ('\u{00F5}', 611), ('\u{00F6}', 611),
|
||||
('\u{00F7}', 584), ('\u{00F8}', 611), ('\u{00F9}', 611), ('\u{00FA}', 611), ('\u{00FB}', 611), ('\u{00FC}', 611),
|
||||
('\u{00FD}', 556), ('\u{00FE}', 611), ('\u{00FF}', 556), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
|
||||
('\u{0152}', 1000), ('\u{0153}', 944), ('\u{0160}', 667), ('\u{0161}', 556), ('\u{0178}', 667), ('\u{017D}', 611),
|
||||
('\u{017E}', 500), ('\u{0192}', 556), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
|
||||
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 556), ('\u{2014}', 1000),
|
||||
('\u{2018}', 278), ('\u{2019}', 278), ('\u{201A}', 278), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
|
||||
('\u{2020}', 556), ('\u{2021}', 556), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
|
||||
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 556), ('\u{2122}', 1000), ('\u{2212}', 584), ('\u{FB01}', 611),
|
||||
('\u{FB02}', 611),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static HELVETICA_OBLIQUE: &[(char, u16)] = &[
|
||||
(' ', 278), ('!', 278), ('"', 355), ('#', 556), ('$', 556), ('%', 889),
|
||||
('&', 667), ('\'', 191), ('(', 333), (')', 333), ('*', 389), ('+', 584),
|
||||
(',', 278), ('-', 333), ('.', 278), ('/', 278), ('0', 556), ('1', 556),
|
||||
('2', 556), ('3', 556), ('4', 556), ('5', 556), ('6', 556), ('7', 556),
|
||||
('8', 556), ('9', 556), (':', 278), (';', 278), ('<', 584), ('=', 584),
|
||||
('>', 584), ('?', 556), ('@', 1015), ('A', 667), ('B', 667), ('C', 722),
|
||||
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 722), ('I', 278),
|
||||
('J', 500), ('K', 667), ('L', 556), ('M', 833), ('N', 722), ('O', 778),
|
||||
('P', 667), ('Q', 778), ('R', 722), ('S', 667), ('T', 611), ('U', 722),
|
||||
('V', 667), ('W', 944), ('X', 667), ('Y', 667), ('Z', 611), ('[', 278),
|
||||
('\\', 278), (']', 278), ('^', 469), ('_', 556), ('`', 333), ('a', 556),
|
||||
('b', 556), ('c', 500), ('d', 556), ('e', 556), ('f', 278), ('g', 556),
|
||||
('h', 556), ('i', 222), ('j', 222), ('k', 500), ('l', 222), ('m', 833),
|
||||
('n', 556), ('o', 556), ('p', 556), ('q', 556), ('r', 333), ('s', 500),
|
||||
('t', 278), ('u', 556), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
|
||||
('z', 500), ('{', 334), ('|', 260), ('}', 334), ('~', 584), ('\u{00A1}', 333),
|
||||
('\u{00A2}', 556), ('\u{00A3}', 556), ('\u{00A4}', 556), ('\u{00A5}', 556), ('\u{00A6}', 260), ('\u{00A7}', 556),
|
||||
('\u{00A8}', 333), ('\u{00A9}', 737), ('\u{00AA}', 370), ('\u{00AB}', 556), ('\u{00AC}', 584), ('\u{00AE}', 737),
|
||||
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 584), ('\u{00B2}', 333), ('\u{00B3}', 333), ('\u{00B4}', 333),
|
||||
('\u{00B5}', 556), ('\u{00B6}', 537), ('\u{00B7}', 278), ('\u{00B8}', 333), ('\u{00B9}', 333), ('\u{00BA}', 365),
|
||||
('\u{00BB}', 556), ('\u{00BC}', 834), ('\u{00BD}', 834), ('\u{00BE}', 834), ('\u{00BF}', 611), ('\u{00C0}', 667),
|
||||
('\u{00C1}', 667), ('\u{00C2}', 667), ('\u{00C3}', 667), ('\u{00C4}', 667), ('\u{00C5}', 667), ('\u{00C6}', 1000),
|
||||
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 278),
|
||||
('\u{00CD}', 278), ('\u{00CE}', 278), ('\u{00CF}', 278), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
|
||||
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 584), ('\u{00D8}', 778),
|
||||
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 667), ('\u{00DE}', 667),
|
||||
('\u{00DF}', 611), ('\u{00E0}', 556), ('\u{00E1}', 556), ('\u{00E2}', 556), ('\u{00E3}', 556), ('\u{00E4}', 556),
|
||||
('\u{00E5}', 556), ('\u{00E6}', 889), ('\u{00E7}', 500), ('\u{00E8}', 556), ('\u{00E9}', 556), ('\u{00EA}', 556),
|
||||
('\u{00EB}', 556), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 556),
|
||||
('\u{00F1}', 556), ('\u{00F2}', 556), ('\u{00F3}', 556), ('\u{00F4}', 556), ('\u{00F5}', 556), ('\u{00F6}', 556),
|
||||
('\u{00F7}', 584), ('\u{00F8}', 611), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
|
||||
('\u{00FD}', 500), ('\u{00FE}', 556), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 556), ('\u{0142}', 222),
|
||||
('\u{0152}', 1000), ('\u{0153}', 944), ('\u{0160}', 667), ('\u{0161}', 500), ('\u{0178}', 667), ('\u{017D}', 611),
|
||||
('\u{017E}', 500), ('\u{0192}', 556), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
|
||||
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 556), ('\u{2014}', 1000),
|
||||
('\u{2018}', 222), ('\u{2019}', 222), ('\u{201A}', 222), ('\u{201C}', 333), ('\u{201D}', 333), ('\u{201E}', 333),
|
||||
('\u{2020}', 556), ('\u{2021}', 556), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
|
||||
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 556), ('\u{2122}', 1000), ('\u{2212}', 584), ('\u{FB01}', 500),
|
||||
('\u{FB02}', 500),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static HELVETICA_BOLDOBLIQUE: &[(char, u16)] = &[
|
||||
(' ', 278), ('!', 333), ('"', 474), ('#', 556), ('$', 556), ('%', 889),
|
||||
('&', 722), ('\'', 238), ('(', 333), (')', 333), ('*', 389), ('+', 584),
|
||||
(',', 278), ('-', 333), ('.', 278), ('/', 278), ('0', 556), ('1', 556),
|
||||
('2', 556), ('3', 556), ('4', 556), ('5', 556), ('6', 556), ('7', 556),
|
||||
('8', 556), ('9', 556), (':', 333), (';', 333), ('<', 584), ('=', 584),
|
||||
('>', 584), ('?', 611), ('@', 975), ('A', 722), ('B', 722), ('C', 722),
|
||||
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 722), ('I', 278),
|
||||
('J', 556), ('K', 722), ('L', 611), ('M', 833), ('N', 722), ('O', 778),
|
||||
('P', 667), ('Q', 778), ('R', 722), ('S', 667), ('T', 611), ('U', 722),
|
||||
('V', 667), ('W', 944), ('X', 667), ('Y', 667), ('Z', 611), ('[', 333),
|
||||
('\\', 278), (']', 333), ('^', 584), ('_', 556), ('`', 333), ('a', 556),
|
||||
('b', 611), ('c', 556), ('d', 611), ('e', 556), ('f', 333), ('g', 611),
|
||||
('h', 611), ('i', 278), ('j', 278), ('k', 556), ('l', 278), ('m', 889),
|
||||
('n', 611), ('o', 611), ('p', 611), ('q', 611), ('r', 389), ('s', 556),
|
||||
('t', 333), ('u', 611), ('v', 556), ('w', 778), ('x', 556), ('y', 556),
|
||||
('z', 500), ('{', 389), ('|', 280), ('}', 389), ('~', 584), ('\u{00A1}', 333),
|
||||
('\u{00A2}', 556), ('\u{00A3}', 556), ('\u{00A4}', 556), ('\u{00A5}', 556), ('\u{00A6}', 280), ('\u{00A7}', 556),
|
||||
('\u{00A8}', 333), ('\u{00A9}', 737), ('\u{00AA}', 370), ('\u{00AB}', 556), ('\u{00AC}', 584), ('\u{00AE}', 737),
|
||||
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 584), ('\u{00B2}', 333), ('\u{00B3}', 333), ('\u{00B4}', 333),
|
||||
('\u{00B5}', 611), ('\u{00B6}', 556), ('\u{00B7}', 278), ('\u{00B8}', 333), ('\u{00B9}', 333), ('\u{00BA}', 365),
|
||||
('\u{00BB}', 556), ('\u{00BC}', 834), ('\u{00BD}', 834), ('\u{00BE}', 834), ('\u{00BF}', 611), ('\u{00C0}', 722),
|
||||
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 1000),
|
||||
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 278),
|
||||
('\u{00CD}', 278), ('\u{00CE}', 278), ('\u{00CF}', 278), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
|
||||
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 584), ('\u{00D8}', 778),
|
||||
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 667), ('\u{00DE}', 667),
|
||||
('\u{00DF}', 611), ('\u{00E0}', 556), ('\u{00E1}', 556), ('\u{00E2}', 556), ('\u{00E3}', 556), ('\u{00E4}', 556),
|
||||
('\u{00E5}', 556), ('\u{00E6}', 889), ('\u{00E7}', 556), ('\u{00E8}', 556), ('\u{00E9}', 556), ('\u{00EA}', 556),
|
||||
('\u{00EB}', 556), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 611),
|
||||
('\u{00F1}', 611), ('\u{00F2}', 611), ('\u{00F3}', 611), ('\u{00F4}', 611), ('\u{00F5}', 611), ('\u{00F6}', 611),
|
||||
('\u{00F7}', 584), ('\u{00F8}', 611), ('\u{00F9}', 611), ('\u{00FA}', 611), ('\u{00FB}', 611), ('\u{00FC}', 611),
|
||||
('\u{00FD}', 556), ('\u{00FE}', 611), ('\u{00FF}', 556), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
|
||||
('\u{0152}', 1000), ('\u{0153}', 944), ('\u{0160}', 667), ('\u{0161}', 556), ('\u{0178}', 667), ('\u{017D}', 611),
|
||||
('\u{017E}', 500), ('\u{0192}', 556), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
|
||||
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 556), ('\u{2014}', 1000),
|
||||
('\u{2018}', 278), ('\u{2019}', 278), ('\u{201A}', 278), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
|
||||
('\u{2020}', 556), ('\u{2021}', 556), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
|
||||
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 556), ('\u{2122}', 1000), ('\u{2212}', 584), ('\u{FB01}', 611),
|
||||
('\u{FB02}', 611),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static TIMES_ROMAN: &[(char, u16)] = &[
|
||||
(' ', 250), ('!', 333), ('"', 408), ('#', 500), ('$', 500), ('%', 833),
|
||||
('&', 778), ('\'', 180), ('(', 333), (')', 333), ('*', 500), ('+', 564),
|
||||
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
|
||||
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
|
||||
('8', 500), ('9', 500), (':', 278), (';', 278), ('<', 564), ('=', 564),
|
||||
('>', 564), ('?', 444), ('@', 921), ('A', 722), ('B', 667), ('C', 667),
|
||||
('D', 722), ('E', 611), ('F', 556), ('G', 722), ('H', 722), ('I', 333),
|
||||
('J', 389), ('K', 722), ('L', 611), ('M', 889), ('N', 722), ('O', 722),
|
||||
('P', 556), ('Q', 722), ('R', 667), ('S', 556), ('T', 611), ('U', 722),
|
||||
('V', 722), ('W', 944), ('X', 722), ('Y', 722), ('Z', 611), ('[', 333),
|
||||
('\\', 278), (']', 333), ('^', 469), ('_', 500), ('`', 333), ('a', 444),
|
||||
('b', 500), ('c', 444), ('d', 500), ('e', 444), ('f', 333), ('g', 500),
|
||||
('h', 500), ('i', 278), ('j', 278), ('k', 500), ('l', 278), ('m', 778),
|
||||
('n', 500), ('o', 500), ('p', 500), ('q', 500), ('r', 333), ('s', 389),
|
||||
('t', 278), ('u', 500), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
|
||||
('z', 444), ('{', 480), ('|', 200), ('}', 480), ('~', 541), ('\u{00A1}', 333),
|
||||
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 200), ('\u{00A7}', 500),
|
||||
('\u{00A8}', 333), ('\u{00A9}', 760), ('\u{00AA}', 276), ('\u{00AB}', 500), ('\u{00AC}', 564), ('\u{00AE}', 760),
|
||||
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 564), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
|
||||
('\u{00B5}', 500), ('\u{00B6}', 453), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 310),
|
||||
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 444), ('\u{00C0}', 722),
|
||||
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 889),
|
||||
('\u{00C7}', 667), ('\u{00C8}', 611), ('\u{00C9}', 611), ('\u{00CA}', 611), ('\u{00CB}', 611), ('\u{00CC}', 333),
|
||||
('\u{00CD}', 333), ('\u{00CE}', 333), ('\u{00CF}', 333), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 722),
|
||||
('\u{00D3}', 722), ('\u{00D4}', 722), ('\u{00D5}', 722), ('\u{00D6}', 722), ('\u{00D7}', 564), ('\u{00D8}', 722),
|
||||
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 722), ('\u{00DE}', 556),
|
||||
('\u{00DF}', 500), ('\u{00E0}', 444), ('\u{00E1}', 444), ('\u{00E2}', 444), ('\u{00E3}', 444), ('\u{00E4}', 444),
|
||||
('\u{00E5}', 444), ('\u{00E6}', 667), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
|
||||
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
|
||||
('\u{00F1}', 500), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
|
||||
('\u{00F7}', 564), ('\u{00F8}', 500), ('\u{00F9}', 500), ('\u{00FA}', 500), ('\u{00FB}', 500), ('\u{00FC}', 500),
|
||||
('\u{00FD}', 500), ('\u{00FE}', 500), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
|
||||
('\u{0152}', 889), ('\u{0153}', 722), ('\u{0160}', 556), ('\u{0161}', 389), ('\u{0178}', 722), ('\u{017D}', 611),
|
||||
('\u{017E}', 444), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
|
||||
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 1000),
|
||||
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 444), ('\u{201D}', 444), ('\u{201E}', 444),
|
||||
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
|
||||
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 980), ('\u{2212}', 564), ('\u{FB01}', 556),
|
||||
('\u{FB02}', 556),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static TIMES_BOLD: &[(char, u16)] = &[
|
||||
(' ', 250), ('!', 333), ('"', 555), ('#', 500), ('$', 500), ('%', 1000),
|
||||
('&', 833), ('\'', 278), ('(', 333), (')', 333), ('*', 500), ('+', 570),
|
||||
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
|
||||
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
|
||||
('8', 500), ('9', 500), (':', 333), (';', 333), ('<', 570), ('=', 570),
|
||||
('>', 570), ('?', 500), ('@', 930), ('A', 722), ('B', 667), ('C', 722),
|
||||
('D', 722), ('E', 667), ('F', 611), ('G', 778), ('H', 778), ('I', 389),
|
||||
('J', 500), ('K', 778), ('L', 667), ('M', 944), ('N', 722), ('O', 778),
|
||||
('P', 611), ('Q', 778), ('R', 722), ('S', 556), ('T', 667), ('U', 722),
|
||||
('V', 722), ('W', 1000), ('X', 722), ('Y', 722), ('Z', 667), ('[', 333),
|
||||
('\\', 278), (']', 333), ('^', 581), ('_', 500), ('`', 333), ('a', 500),
|
||||
('b', 556), ('c', 444), ('d', 556), ('e', 444), ('f', 333), ('g', 500),
|
||||
('h', 556), ('i', 278), ('j', 333), ('k', 556), ('l', 278), ('m', 833),
|
||||
('n', 556), ('o', 500), ('p', 556), ('q', 556), ('r', 444), ('s', 389),
|
||||
('t', 333), ('u', 556), ('v', 500), ('w', 722), ('x', 500), ('y', 500),
|
||||
('z', 444), ('{', 394), ('|', 220), ('}', 394), ('~', 520), ('\u{00A1}', 333),
|
||||
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 220), ('\u{00A7}', 500),
|
||||
('\u{00A8}', 333), ('\u{00A9}', 747), ('\u{00AA}', 300), ('\u{00AB}', 500), ('\u{00AC}', 570), ('\u{00AE}', 747),
|
||||
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 570), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
|
||||
('\u{00B5}', 556), ('\u{00B6}', 540), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 330),
|
||||
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 500), ('\u{00C0}', 722),
|
||||
('\u{00C1}', 722), ('\u{00C2}', 722), ('\u{00C3}', 722), ('\u{00C4}', 722), ('\u{00C5}', 722), ('\u{00C6}', 1000),
|
||||
('\u{00C7}', 722), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 389),
|
||||
('\u{00CD}', 389), ('\u{00CE}', 389), ('\u{00CF}', 389), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 778),
|
||||
('\u{00D3}', 778), ('\u{00D4}', 778), ('\u{00D5}', 778), ('\u{00D6}', 778), ('\u{00D7}', 570), ('\u{00D8}', 778),
|
||||
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 722), ('\u{00DE}', 611),
|
||||
('\u{00DF}', 556), ('\u{00E0}', 500), ('\u{00E1}', 500), ('\u{00E2}', 500), ('\u{00E3}', 500), ('\u{00E4}', 500),
|
||||
('\u{00E5}', 500), ('\u{00E6}', 722), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
|
||||
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
|
||||
('\u{00F1}', 556), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
|
||||
('\u{00F7}', 570), ('\u{00F8}', 500), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
|
||||
('\u{00FD}', 500), ('\u{00FE}', 556), ('\u{00FF}', 500), ('\u{0131}', 278), ('\u{0141}', 667), ('\u{0142}', 278),
|
||||
('\u{0152}', 1000), ('\u{0153}', 722), ('\u{0160}', 556), ('\u{0161}', 389), ('\u{0178}', 722), ('\u{017D}', 667),
|
||||
('\u{017E}', 444), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
|
||||
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 1000),
|
||||
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
|
||||
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
|
||||
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 1000), ('\u{2212}', 570), ('\u{FB01}', 556),
|
||||
('\u{FB02}', 556),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static TIMES_ITALIC: &[(char, u16)] = &[
|
||||
(' ', 250), ('!', 333), ('"', 420), ('#', 500), ('$', 500), ('%', 833),
|
||||
('&', 778), ('\'', 214), ('(', 333), (')', 333), ('*', 500), ('+', 675),
|
||||
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
|
||||
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
|
||||
('8', 500), ('9', 500), (':', 333), (';', 333), ('<', 675), ('=', 675),
|
||||
('>', 675), ('?', 500), ('@', 920), ('A', 611), ('B', 611), ('C', 667),
|
||||
('D', 722), ('E', 611), ('F', 611), ('G', 722), ('H', 722), ('I', 333),
|
||||
('J', 444), ('K', 667), ('L', 556), ('M', 833), ('N', 667), ('O', 722),
|
||||
('P', 611), ('Q', 722), ('R', 611), ('S', 500), ('T', 556), ('U', 722),
|
||||
('V', 611), ('W', 833), ('X', 611), ('Y', 556), ('Z', 556), ('[', 389),
|
||||
('\\', 278), (']', 389), ('^', 422), ('_', 500), ('`', 333), ('a', 500),
|
||||
('b', 500), ('c', 444), ('d', 500), ('e', 444), ('f', 278), ('g', 500),
|
||||
('h', 500), ('i', 278), ('j', 278), ('k', 444), ('l', 278), ('m', 722),
|
||||
('n', 500), ('o', 500), ('p', 500), ('q', 500), ('r', 389), ('s', 389),
|
||||
('t', 278), ('u', 500), ('v', 444), ('w', 667), ('x', 444), ('y', 444),
|
||||
('z', 389), ('{', 400), ('|', 275), ('}', 400), ('~', 541), ('\u{00A1}', 389),
|
||||
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 275), ('\u{00A7}', 500),
|
||||
('\u{00A8}', 333), ('\u{00A9}', 760), ('\u{00AA}', 276), ('\u{00AB}', 500), ('\u{00AC}', 675), ('\u{00AE}', 760),
|
||||
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 675), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
|
||||
('\u{00B5}', 500), ('\u{00B6}', 523), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 310),
|
||||
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 500), ('\u{00C0}', 611),
|
||||
('\u{00C1}', 611), ('\u{00C2}', 611), ('\u{00C3}', 611), ('\u{00C4}', 611), ('\u{00C5}', 611), ('\u{00C6}', 889),
|
||||
('\u{00C7}', 667), ('\u{00C8}', 611), ('\u{00C9}', 611), ('\u{00CA}', 611), ('\u{00CB}', 611), ('\u{00CC}', 333),
|
||||
('\u{00CD}', 333), ('\u{00CE}', 333), ('\u{00CF}', 333), ('\u{00D0}', 722), ('\u{00D1}', 667), ('\u{00D2}', 722),
|
||||
('\u{00D3}', 722), ('\u{00D4}', 722), ('\u{00D5}', 722), ('\u{00D6}', 722), ('\u{00D7}', 675), ('\u{00D8}', 722),
|
||||
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 556), ('\u{00DE}', 611),
|
||||
('\u{00DF}', 500), ('\u{00E0}', 500), ('\u{00E1}', 500), ('\u{00E2}', 500), ('\u{00E3}', 500), ('\u{00E4}', 500),
|
||||
('\u{00E5}', 500), ('\u{00E6}', 667), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
|
||||
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
|
||||
('\u{00F1}', 500), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
|
||||
('\u{00F7}', 675), ('\u{00F8}', 500), ('\u{00F9}', 500), ('\u{00FA}', 500), ('\u{00FB}', 500), ('\u{00FC}', 500),
|
||||
('\u{00FD}', 444), ('\u{00FE}', 500), ('\u{00FF}', 444), ('\u{0131}', 278), ('\u{0141}', 556), ('\u{0142}', 278),
|
||||
('\u{0152}', 944), ('\u{0153}', 667), ('\u{0160}', 500), ('\u{0161}', 389), ('\u{0178}', 556), ('\u{017D}', 556),
|
||||
('\u{017E}', 389), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
|
||||
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 889),
|
||||
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 556), ('\u{201D}', 556), ('\u{201E}', 556),
|
||||
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 889), ('\u{2030}', 1000), ('\u{2039}', 333),
|
||||
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 980), ('\u{2212}', 675), ('\u{FB01}', 500),
|
||||
('\u{FB02}', 500),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static TIMES_BOLDITALIC: &[(char, u16)] = &[
|
||||
(' ', 250), ('!', 389), ('"', 555), ('#', 500), ('$', 500), ('%', 833),
|
||||
('&', 778), ('\'', 278), ('(', 333), (')', 333), ('*', 500), ('+', 570),
|
||||
(',', 250), ('-', 333), ('.', 250), ('/', 278), ('0', 500), ('1', 500),
|
||||
('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500), ('7', 500),
|
||||
('8', 500), ('9', 500), (':', 333), (';', 333), ('<', 570), ('=', 570),
|
||||
('>', 570), ('?', 500), ('@', 832), ('A', 667), ('B', 667), ('C', 667),
|
||||
('D', 722), ('E', 667), ('F', 667), ('G', 722), ('H', 778), ('I', 389),
|
||||
('J', 500), ('K', 667), ('L', 611), ('M', 889), ('N', 722), ('O', 722),
|
||||
('P', 611), ('Q', 722), ('R', 667), ('S', 556), ('T', 611), ('U', 722),
|
||||
('V', 667), ('W', 889), ('X', 667), ('Y', 611), ('Z', 611), ('[', 333),
|
||||
('\\', 278), (']', 333), ('^', 570), ('_', 500), ('`', 333), ('a', 500),
|
||||
('b', 500), ('c', 444), ('d', 500), ('e', 444), ('f', 333), ('g', 500),
|
||||
('h', 556), ('i', 278), ('j', 278), ('k', 500), ('l', 278), ('m', 778),
|
||||
('n', 556), ('o', 500), ('p', 500), ('q', 500), ('r', 389), ('s', 389),
|
||||
('t', 278), ('u', 556), ('v', 444), ('w', 667), ('x', 500), ('y', 444),
|
||||
('z', 389), ('{', 348), ('|', 220), ('}', 348), ('~', 570), ('\u{00A1}', 389),
|
||||
('\u{00A2}', 500), ('\u{00A3}', 500), ('\u{00A4}', 500), ('\u{00A5}', 500), ('\u{00A6}', 220), ('\u{00A7}', 500),
|
||||
('\u{00A8}', 333), ('\u{00A9}', 747), ('\u{00AA}', 266), ('\u{00AB}', 500), ('\u{00AC}', 606), ('\u{00AE}', 747),
|
||||
('\u{00AF}', 333), ('\u{00B0}', 400), ('\u{00B1}', 570), ('\u{00B2}', 300), ('\u{00B3}', 300), ('\u{00B4}', 333),
|
||||
('\u{00B5}', 576), ('\u{00B6}', 500), ('\u{00B7}', 250), ('\u{00B8}', 333), ('\u{00B9}', 300), ('\u{00BA}', 300),
|
||||
('\u{00BB}', 500), ('\u{00BC}', 750), ('\u{00BD}', 750), ('\u{00BE}', 750), ('\u{00BF}', 500), ('\u{00C0}', 667),
|
||||
('\u{00C1}', 667), ('\u{00C2}', 667), ('\u{00C3}', 667), ('\u{00C4}', 667), ('\u{00C5}', 667), ('\u{00C6}', 944),
|
||||
('\u{00C7}', 667), ('\u{00C8}', 667), ('\u{00C9}', 667), ('\u{00CA}', 667), ('\u{00CB}', 667), ('\u{00CC}', 389),
|
||||
('\u{00CD}', 389), ('\u{00CE}', 389), ('\u{00CF}', 389), ('\u{00D0}', 722), ('\u{00D1}', 722), ('\u{00D2}', 722),
|
||||
('\u{00D3}', 722), ('\u{00D4}', 722), ('\u{00D5}', 722), ('\u{00D6}', 722), ('\u{00D7}', 570), ('\u{00D8}', 722),
|
||||
('\u{00D9}', 722), ('\u{00DA}', 722), ('\u{00DB}', 722), ('\u{00DC}', 722), ('\u{00DD}', 611), ('\u{00DE}', 611),
|
||||
('\u{00DF}', 500), ('\u{00E0}', 500), ('\u{00E1}', 500), ('\u{00E2}', 500), ('\u{00E3}', 500), ('\u{00E4}', 500),
|
||||
('\u{00E5}', 500), ('\u{00E6}', 722), ('\u{00E7}', 444), ('\u{00E8}', 444), ('\u{00E9}', 444), ('\u{00EA}', 444),
|
||||
('\u{00EB}', 444), ('\u{00EC}', 278), ('\u{00ED}', 278), ('\u{00EE}', 278), ('\u{00EF}', 278), ('\u{00F0}', 500),
|
||||
('\u{00F1}', 556), ('\u{00F2}', 500), ('\u{00F3}', 500), ('\u{00F4}', 500), ('\u{00F5}', 500), ('\u{00F6}', 500),
|
||||
('\u{00F7}', 570), ('\u{00F8}', 500), ('\u{00F9}', 556), ('\u{00FA}', 556), ('\u{00FB}', 556), ('\u{00FC}', 556),
|
||||
('\u{00FD}', 444), ('\u{00FE}', 500), ('\u{00FF}', 444), ('\u{0131}', 278), ('\u{0141}', 611), ('\u{0142}', 278),
|
||||
('\u{0152}', 944), ('\u{0153}', 722), ('\u{0160}', 556), ('\u{0161}', 389), ('\u{0178}', 611), ('\u{017D}', 611),
|
||||
('\u{017E}', 389), ('\u{0192}', 500), ('\u{02C6}', 333), ('\u{02C7}', 333), ('\u{02D8}', 333), ('\u{02D9}', 333),
|
||||
('\u{02DA}', 333), ('\u{02DB}', 333), ('\u{02DC}', 333), ('\u{02DD}', 333), ('\u{2013}', 500), ('\u{2014}', 1000),
|
||||
('\u{2018}', 333), ('\u{2019}', 333), ('\u{201A}', 333), ('\u{201C}', 500), ('\u{201D}', 500), ('\u{201E}', 500),
|
||||
('\u{2020}', 500), ('\u{2021}', 500), ('\u{2022}', 350), ('\u{2026}', 1000), ('\u{2030}', 1000), ('\u{2039}', 333),
|
||||
('\u{203A}', 333), ('\u{2044}', 167), ('\u{20AC}', 500), ('\u{2122}', 1000), ('\u{2212}', 606), ('\u{FB01}', 556),
|
||||
('\u{FB02}', 556),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static SYMBOL: &[(char, u16)] = &[
|
||||
(' ', 250), ('!', 333), ('#', 500), ('%', 833), ('&', 778), ('(', 333),
|
||||
(')', 333), ('+', 549), (',', 250), ('.', 250), ('/', 278), ('0', 500),
|
||||
('1', 500), ('2', 500), ('3', 500), ('4', 500), ('5', 500), ('6', 500),
|
||||
('7', 500), ('8', 500), ('9', 500), (':', 278), (';', 278), ('<', 549),
|
||||
('=', 549), ('>', 549), ('?', 444), ('[', 333), (']', 333), ('_', 500),
|
||||
('{', 480), ('|', 200), ('}', 480), ('\u{00AC}', 713), ('\u{00B0}', 400), ('\u{00B1}', 549),
|
||||
('\u{00B5}', 576), ('\u{00D7}', 549), ('\u{00F7}', 549), ('\u{0192}', 500), ('\u{0391}', 722), ('\u{0392}', 667),
|
||||
('\u{0393}', 603), ('\u{0395}', 611), ('\u{0396}', 611), ('\u{0397}', 722), ('\u{0398}', 741), ('\u{0399}', 333),
|
||||
('\u{039A}', 722), ('\u{039B}', 686), ('\u{039C}', 889), ('\u{039D}', 722), ('\u{039E}', 645), ('\u{039F}', 722),
|
||||
('\u{03A0}', 768), ('\u{03A1}', 556), ('\u{03A3}', 592), ('\u{03A4}', 611), ('\u{03A5}', 690), ('\u{03A6}', 763),
|
||||
('\u{03A7}', 722), ('\u{03A8}', 795), ('\u{03B1}', 631), ('\u{03B2}', 549), ('\u{03B3}', 411), ('\u{03B4}', 494),
|
||||
('\u{03B5}', 439), ('\u{03B6}', 494), ('\u{03B7}', 603), ('\u{03B8}', 521), ('\u{03B9}', 329), ('\u{03BA}', 549),
|
||||
('\u{03BB}', 549), ('\u{03BD}', 521), ('\u{03BE}', 493), ('\u{03BF}', 549), ('\u{03C0}', 549), ('\u{03C1}', 549),
|
||||
('\u{03C2}', 439), ('\u{03C3}', 603), ('\u{03C4}', 439), ('\u{03C5}', 576), ('\u{03C6}', 521), ('\u{03C7}', 549),
|
||||
('\u{03C8}', 686), ('\u{03C9}', 686), ('\u{03D1}', 631), ('\u{03D2}', 620), ('\u{03D5}', 603), ('\u{03D6}', 713),
|
||||
('\u{2022}', 460), ('\u{2026}', 1000), ('\u{2032}', 247), ('\u{2033}', 411), ('\u{2044}', 167), ('\u{20AC}', 750),
|
||||
('\u{2111}', 686), ('\u{2118}', 987), ('\u{211C}', 795), ('\u{2126}', 768), ('\u{2135}', 823), ('\u{2190}', 987),
|
||||
('\u{2191}', 603), ('\u{2192}', 987), ('\u{2193}', 603), ('\u{2194}', 1042), ('\u{21B5}', 658), ('\u{21D0}', 987),
|
||||
('\u{21D1}', 603), ('\u{21D2}', 987), ('\u{21D3}', 603), ('\u{21D4}', 1042), ('\u{2200}', 713), ('\u{2202}', 494),
|
||||
('\u{2203}', 549), ('\u{2205}', 823), ('\u{2206}', 612), ('\u{2207}', 713), ('\u{2208}', 713), ('\u{2209}', 713),
|
||||
('\u{220B}', 439), ('\u{220F}', 823), ('\u{2211}', 713), ('\u{2212}', 549), ('\u{2217}', 500), ('\u{221A}', 549),
|
||||
('\u{221D}', 713), ('\u{221E}', 713), ('\u{2220}', 768), ('\u{2227}', 603), ('\u{2228}', 603), ('\u{2229}', 768),
|
||||
('\u{222A}', 768), ('\u{222B}', 274), ('\u{2234}', 863), ('\u{223C}', 549), ('\u{2245}', 549), ('\u{2248}', 549),
|
||||
('\u{2260}', 549), ('\u{2261}', 549), ('\u{2264}', 549), ('\u{2265}', 549), ('\u{2282}', 713), ('\u{2283}', 713),
|
||||
('\u{2284}', 713), ('\u{2286}', 713), ('\u{2287}', 713), ('\u{2295}', 768), ('\u{2297}', 768), ('\u{22A5}', 658),
|
||||
('\u{22C5}', 250), ('\u{2320}', 686), ('\u{2321}', 686), ('\u{2329}', 329), ('\u{232A}', 329), ('\u{25CA}', 494),
|
||||
('\u{2660}', 753), ('\u{2663}', 753), ('\u{2665}', 753), ('\u{2666}', 753), ('\u{F6D9}', 790), ('\u{F6DA}', 790),
|
||||
('\u{F6DB}', 890), ('\u{F8E5}', 500), ('\u{F8E6}', 603), ('\u{F8E7}', 1000), ('\u{F8E8}', 790), ('\u{F8E9}', 790),
|
||||
('\u{F8EA}', 786), ('\u{F8EB}', 384), ('\u{F8EC}', 384), ('\u{F8ED}', 384), ('\u{F8EE}', 384), ('\u{F8EF}', 384),
|
||||
('\u{F8F0}', 384), ('\u{F8F1}', 494), ('\u{F8F2}', 494), ('\u{F8F3}', 494), ('\u{F8F4}', 494), ('\u{F8F5}', 686),
|
||||
('\u{F8F6}', 384), ('\u{F8F7}', 384), ('\u{F8F8}', 384), ('\u{F8F9}', 384), ('\u{F8FA}', 384), ('\u{F8FB}', 384),
|
||||
('\u{F8FC}', 494), ('\u{F8FD}', 494), ('\u{F8FE}', 494), ('\u{F8FF}', 790),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static ZAPFDINGBATS: &[(char, u16)] = &[
|
||||
(' ', 278), ('\u{2192}', 838), ('\u{2194}', 1016), ('\u{2195}', 458), ('\u{2460}', 788), ('\u{2461}', 788),
|
||||
('\u{2462}', 788), ('\u{2463}', 788), ('\u{2464}', 788), ('\u{2465}', 788), ('\u{2466}', 788), ('\u{2467}', 788),
|
||||
('\u{2468}', 788), ('\u{2469}', 788), ('\u{25A0}', 761), ('\u{25B2}', 892), ('\u{25BC}', 892), ('\u{25C6}', 788),
|
||||
('\u{25CF}', 791), ('\u{25D7}', 438), ('\u{2605}', 816), ('\u{260E}', 719), ('\u{261B}', 960), ('\u{261E}', 939),
|
||||
('\u{2660}', 626), ('\u{2663}', 776), ('\u{2665}', 694), ('\u{2666}', 595), ('\u{2701}', 974), ('\u{2702}', 961),
|
||||
('\u{2703}', 974), ('\u{2704}', 980), ('\u{2706}', 789), ('\u{2707}', 790), ('\u{2708}', 791), ('\u{2709}', 690),
|
||||
('\u{270C}', 549), ('\u{270D}', 855), ('\u{270E}', 911), ('\u{270F}', 933), ('\u{2710}', 911), ('\u{2711}', 945),
|
||||
('\u{2712}', 974), ('\u{2713}', 755), ('\u{2714}', 846), ('\u{2715}', 762), ('\u{2716}', 761), ('\u{2717}', 571),
|
||||
('\u{2718}', 677), ('\u{2719}', 763), ('\u{271A}', 760), ('\u{271B}', 759), ('\u{271C}', 754), ('\u{271D}', 494),
|
||||
('\u{271E}', 552), ('\u{271F}', 537), ('\u{2720}', 577), ('\u{2721}', 692), ('\u{2722}', 786), ('\u{2723}', 788),
|
||||
('\u{2724}', 788), ('\u{2725}', 790), ('\u{2726}', 793), ('\u{2727}', 794), ('\u{2729}', 823), ('\u{272A}', 789),
|
||||
('\u{272B}', 841), ('\u{272C}', 823), ('\u{272D}', 833), ('\u{272E}', 816), ('\u{272F}', 831), ('\u{2730}', 923),
|
||||
('\u{2731}', 744), ('\u{2732}', 723), ('\u{2733}', 749), ('\u{2734}', 790), ('\u{2735}', 792), ('\u{2736}', 695),
|
||||
('\u{2737}', 776), ('\u{2738}', 768), ('\u{2739}', 792), ('\u{273A}', 759), ('\u{273B}', 707), ('\u{273C}', 708),
|
||||
('\u{273D}', 682), ('\u{273E}', 701), ('\u{273F}', 826), ('\u{2740}', 815), ('\u{2741}', 789), ('\u{2742}', 789),
|
||||
('\u{2743}', 707), ('\u{2744}', 687), ('\u{2745}', 696), ('\u{2746}', 689), ('\u{2747}', 786), ('\u{2748}', 787),
|
||||
('\u{2749}', 713), ('\u{274A}', 791), ('\u{274B}', 785), ('\u{274D}', 873), ('\u{274F}', 762), ('\u{2750}', 762),
|
||||
('\u{2751}', 759), ('\u{2752}', 759), ('\u{2756}', 784), ('\u{2758}', 138), ('\u{2759}', 277), ('\u{275A}', 415),
|
||||
('\u{275B}', 392), ('\u{275C}', 392), ('\u{275D}', 668), ('\u{275E}', 668), ('\u{2761}', 732), ('\u{2762}', 544),
|
||||
('\u{2763}', 544), ('\u{2764}', 910), ('\u{2765}', 667), ('\u{2766}', 760), ('\u{2767}', 760), ('\u{2768}', 390),
|
||||
('\u{2769}', 390), ('\u{276A}', 317), ('\u{276B}', 317), ('\u{276C}', 276), ('\u{276D}', 276), ('\u{276E}', 509),
|
||||
('\u{276F}', 509), ('\u{2770}', 410), ('\u{2771}', 410), ('\u{2772}', 234), ('\u{2773}', 234), ('\u{2774}', 334),
|
||||
('\u{2775}', 334), ('\u{2776}', 788), ('\u{2777}', 788), ('\u{2778}', 788), ('\u{2779}', 788), ('\u{277A}', 788),
|
||||
('\u{277B}', 788), ('\u{277C}', 788), ('\u{277D}', 788), ('\u{277E}', 788), ('\u{277F}', 788), ('\u{2780}', 788),
|
||||
('\u{2781}', 788), ('\u{2782}', 788), ('\u{2783}', 788), ('\u{2784}', 788), ('\u{2785}', 788), ('\u{2786}', 788),
|
||||
('\u{2787}', 788), ('\u{2788}', 788), ('\u{2789}', 788), ('\u{278A}', 788), ('\u{278B}', 788), ('\u{278C}', 788),
|
||||
('\u{278D}', 788), ('\u{278E}', 788), ('\u{278F}', 788), ('\u{2790}', 788), ('\u{2791}', 788), ('\u{2792}', 788),
|
||||
('\u{2793}', 788), ('\u{2794}', 894), ('\u{2798}', 748), ('\u{2799}', 924), ('\u{279A}', 748), ('\u{279B}', 918),
|
||||
('\u{279C}', 927), ('\u{279D}', 928), ('\u{279E}', 928), ('\u{279F}', 834), ('\u{27A0}', 873), ('\u{27A1}', 828),
|
||||
('\u{27A2}', 924), ('\u{27A3}', 924), ('\u{27A4}', 917), ('\u{27A5}', 930), ('\u{27A6}', 931), ('\u{27A7}', 463),
|
||||
('\u{27A8}', 883), ('\u{27A9}', 836), ('\u{27AA}', 836), ('\u{27AB}', 867), ('\u{27AC}', 867), ('\u{27AD}', 696),
|
||||
('\u{27AE}', 696), ('\u{27AF}', 874), ('\u{27B1}', 874), ('\u{27B2}', 760), ('\u{27B3}', 946), ('\u{27B4}', 771),
|
||||
('\u{27B5}', 865), ('\u{27B6}', 771), ('\u{27B7}', 888), ('\u{27B8}', 967), ('\u{27B9}', 888), ('\u{27BA}', 831),
|
||||
('\u{27BB}', 873), ('\u{27BC}', 927), ('\u{27BD}', 970), ('\u{27BE}', 918),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static SYMBOL_ENCODING: &[(u8, char)] = &[
|
||||
(0x20, ' '), (0x21, '!'), (0x22, '\u{2200}'), (0x23, '#'), (0x24, '\u{2203}'), (0x25, '%'),
|
||||
(0x26, '&'), (0x27, '\u{220B}'), (0x28, '('), (0x29, ')'), (0x2A, '\u{2217}'), (0x2B, '+'),
|
||||
(0x2C, ','), (0x2D, '\u{2212}'), (0x2E, '.'), (0x2F, '/'), (0x30, '0'), (0x31, '1'),
|
||||
(0x32, '2'), (0x33, '3'), (0x34, '4'), (0x35, '5'), (0x36, '6'), (0x37, '7'),
|
||||
(0x38, '8'), (0x39, '9'), (0x3A, ':'), (0x3B, ';'), (0x3C, '<'), (0x3D, '='),
|
||||
(0x3E, '>'), (0x3F, '?'), (0x40, '\u{2245}'), (0x41, '\u{0391}'), (0x42, '\u{0392}'), (0x43, '\u{03A7}'),
|
||||
(0x44, '\u{2206}'), (0x45, '\u{0395}'), (0x46, '\u{03A6}'), (0x47, '\u{0393}'), (0x48, '\u{0397}'), (0x49, '\u{0399}'),
|
||||
(0x4A, '\u{03D1}'), (0x4B, '\u{039A}'), (0x4C, '\u{039B}'), (0x4D, '\u{039C}'), (0x4E, '\u{039D}'), (0x4F, '\u{039F}'),
|
||||
(0x50, '\u{03A0}'), (0x51, '\u{0398}'), (0x52, '\u{03A1}'), (0x53, '\u{03A3}'), (0x54, '\u{03A4}'), (0x55, '\u{03A5}'),
|
||||
(0x56, '\u{03C2}'), (0x57, '\u{2126}'), (0x58, '\u{039E}'), (0x59, '\u{03A8}'), (0x5A, '\u{0396}'), (0x5B, '['),
|
||||
(0x5C, '\u{2234}'), (0x5D, ']'), (0x5E, '\u{22A5}'), (0x5F, '_'), (0x60, '\u{F8E5}'), (0x61, '\u{03B1}'),
|
||||
(0x62, '\u{03B2}'), (0x63, '\u{03C7}'), (0x64, '\u{03B4}'), (0x65, '\u{03B5}'), (0x66, '\u{03C6}'), (0x67, '\u{03B3}'),
|
||||
(0x68, '\u{03B7}'), (0x69, '\u{03B9}'), (0x6A, '\u{03D5}'), (0x6B, '\u{03BA}'), (0x6C, '\u{03BB}'), (0x6D, '\u{00B5}'),
|
||||
(0x6E, '\u{03BD}'), (0x6F, '\u{03BF}'), (0x70, '\u{03C0}'), (0x71, '\u{03B8}'), (0x72, '\u{03C1}'), (0x73, '\u{03C3}'),
|
||||
(0x74, '\u{03C4}'), (0x75, '\u{03C5}'), (0x76, '\u{03D6}'), (0x77, '\u{03C9}'), (0x78, '\u{03BE}'), (0x79, '\u{03C8}'),
|
||||
(0x7A, '\u{03B6}'), (0x7B, '{'), (0x7C, '|'), (0x7D, '}'), (0x7E, '\u{223C}'), (0xA0, '\u{20AC}'),
|
||||
(0xA1, '\u{03D2}'), (0xA2, '\u{2032}'), (0xA3, '\u{2264}'), (0xA4, '\u{2044}'), (0xA5, '\u{221E}'), (0xA6, '\u{0192}'),
|
||||
(0xA7, '\u{2663}'), (0xA8, '\u{2666}'), (0xA9, '\u{2665}'), (0xAA, '\u{2660}'), (0xAB, '\u{2194}'), (0xAC, '\u{2190}'),
|
||||
(0xAD, '\u{2191}'), (0xAE, '\u{2192}'), (0xAF, '\u{2193}'), (0xB0, '\u{00B0}'), (0xB1, '\u{00B1}'), (0xB2, '\u{2033}'),
|
||||
(0xB3, '\u{2265}'), (0xB4, '\u{00D7}'), (0xB5, '\u{221D}'), (0xB6, '\u{2202}'), (0xB7, '\u{2022}'), (0xB8, '\u{00F7}'),
|
||||
(0xB9, '\u{2260}'), (0xBA, '\u{2261}'), (0xBB, '\u{2248}'), (0xBC, '\u{2026}'), (0xBD, '\u{F8E6}'), (0xBE, '\u{F8E7}'),
|
||||
(0xBF, '\u{21B5}'), (0xC0, '\u{2135}'), (0xC1, '\u{2111}'), (0xC2, '\u{211C}'), (0xC3, '\u{2118}'), (0xC4, '\u{2297}'),
|
||||
(0xC5, '\u{2295}'), (0xC6, '\u{2205}'), (0xC7, '\u{2229}'), (0xC8, '\u{222A}'), (0xC9, '\u{2283}'), (0xCA, '\u{2287}'),
|
||||
(0xCB, '\u{2284}'), (0xCC, '\u{2282}'), (0xCD, '\u{2286}'), (0xCE, '\u{2208}'), (0xCF, '\u{2209}'), (0xD0, '\u{2220}'),
|
||||
(0xD1, '\u{2207}'), (0xD2, '\u{F6DA}'), (0xD3, '\u{F6D9}'), (0xD4, '\u{F6DB}'), (0xD5, '\u{220F}'), (0xD6, '\u{221A}'),
|
||||
(0xD7, '\u{22C5}'), (0xD8, '\u{00AC}'), (0xD9, '\u{2227}'), (0xDA, '\u{2228}'), (0xDB, '\u{21D4}'), (0xDC, '\u{21D0}'),
|
||||
(0xDD, '\u{21D1}'), (0xDE, '\u{21D2}'), (0xDF, '\u{21D3}'), (0xE0, '\u{25CA}'), (0xE1, '\u{2329}'), (0xE2, '\u{F8E8}'),
|
||||
(0xE3, '\u{F8E9}'), (0xE4, '\u{F8EA}'), (0xE5, '\u{2211}'), (0xE6, '\u{F8EB}'), (0xE7, '\u{F8EC}'), (0xE8, '\u{F8ED}'),
|
||||
(0xE9, '\u{F8EE}'), (0xEA, '\u{F8EF}'), (0xEB, '\u{F8F0}'), (0xEC, '\u{F8F1}'), (0xED, '\u{F8F2}'), (0xEE, '\u{F8F3}'),
|
||||
(0xEF, '\u{F8F4}'), (0xF1, '\u{232A}'), (0xF2, '\u{222B}'), (0xF3, '\u{2320}'), (0xF4, '\u{F8F5}'), (0xF5, '\u{2321}'),
|
||||
(0xF6, '\u{F8F6}'), (0xF7, '\u{F8F7}'), (0xF8, '\u{F8F8}'), (0xF9, '\u{F8F9}'), (0xFA, '\u{F8FA}'), (0xFB, '\u{F8FB}'),
|
||||
(0xFC, '\u{F8FC}'), (0xFD, '\u{F8FD}'), (0xFE, '\u{F8FE}'),
|
||||
];
|
||||
|
||||
#[rustfmt::skip]
|
||||
static ZAPFDINGBATS_ENCODING: &[(u8, char)] = &[
|
||||
(0x20, ' '), (0x21, '\u{2701}'), (0x22, '\u{2702}'), (0x23, '\u{2703}'), (0x24, '\u{2704}'), (0x25, '\u{260E}'),
|
||||
(0x26, '\u{2706}'), (0x27, '\u{2707}'), (0x28, '\u{2708}'), (0x29, '\u{2709}'), (0x2A, '\u{261B}'), (0x2B, '\u{261E}'),
|
||||
(0x2C, '\u{270C}'), (0x2D, '\u{270D}'), (0x2E, '\u{270E}'), (0x2F, '\u{270F}'), (0x30, '\u{2710}'), (0x31, '\u{2711}'),
|
||||
(0x32, '\u{2712}'), (0x33, '\u{2713}'), (0x34, '\u{2714}'), (0x35, '\u{2715}'), (0x36, '\u{2716}'), (0x37, '\u{2717}'),
|
||||
(0x38, '\u{2718}'), (0x39, '\u{2719}'), (0x3A, '\u{271A}'), (0x3B, '\u{271B}'), (0x3C, '\u{271C}'), (0x3D, '\u{271D}'),
|
||||
(0x3E, '\u{271E}'), (0x3F, '\u{271F}'), (0x40, '\u{2720}'), (0x41, '\u{2721}'), (0x42, '\u{2722}'), (0x43, '\u{2723}'),
|
||||
(0x44, '\u{2724}'), (0x45, '\u{2725}'), (0x46, '\u{2726}'), (0x47, '\u{2727}'), (0x48, '\u{2605}'), (0x49, '\u{2729}'),
|
||||
(0x4A, '\u{272A}'), (0x4B, '\u{272B}'), (0x4C, '\u{272C}'), (0x4D, '\u{272D}'), (0x4E, '\u{272E}'), (0x4F, '\u{272F}'),
|
||||
(0x50, '\u{2730}'), (0x51, '\u{2731}'), (0x52, '\u{2732}'), (0x53, '\u{2733}'), (0x54, '\u{2734}'), (0x55, '\u{2735}'),
|
||||
(0x56, '\u{2736}'), (0x57, '\u{2737}'), (0x58, '\u{2738}'), (0x59, '\u{2739}'), (0x5A, '\u{273A}'), (0x5B, '\u{273B}'),
|
||||
(0x5C, '\u{273C}'), (0x5D, '\u{273D}'), (0x5E, '\u{273E}'), (0x5F, '\u{273F}'), (0x60, '\u{2740}'), (0x61, '\u{2741}'),
|
||||
(0x62, '\u{2742}'), (0x63, '\u{2743}'), (0x64, '\u{2744}'), (0x65, '\u{2745}'), (0x66, '\u{2746}'), (0x67, '\u{2747}'),
|
||||
(0x68, '\u{2748}'), (0x69, '\u{2749}'), (0x6A, '\u{274A}'), (0x6B, '\u{274B}'), (0x6C, '\u{25CF}'), (0x6D, '\u{274D}'),
|
||||
(0x6E, '\u{25A0}'), (0x6F, '\u{274F}'), (0x70, '\u{2750}'), (0x71, '\u{2751}'), (0x72, '\u{2752}'), (0x73, '\u{25B2}'),
|
||||
(0x74, '\u{25BC}'), (0x75, '\u{25C6}'), (0x76, '\u{2756}'), (0x77, '\u{25D7}'), (0x78, '\u{2758}'), (0x79, '\u{2759}'),
|
||||
(0x7A, '\u{275A}'), (0x7B, '\u{275B}'), (0x7C, '\u{275C}'), (0x7D, '\u{275D}'), (0x7E, '\u{275E}'), (0x80, '\u{2768}'),
|
||||
(0x81, '\u{2769}'), (0x82, '\u{276A}'), (0x83, '\u{276B}'), (0x84, '\u{276C}'), (0x85, '\u{276D}'), (0x86, '\u{276E}'),
|
||||
(0x87, '\u{276F}'), (0x88, '\u{2770}'), (0x89, '\u{2771}'), (0x8A, '\u{2772}'), (0x8B, '\u{2773}'), (0x8C, '\u{2774}'),
|
||||
(0x8D, '\u{2775}'), (0xA1, '\u{2761}'), (0xA2, '\u{2762}'), (0xA3, '\u{2763}'), (0xA4, '\u{2764}'), (0xA5, '\u{2765}'),
|
||||
(0xA6, '\u{2766}'), (0xA7, '\u{2767}'), (0xA8, '\u{2663}'), (0xA9, '\u{2666}'), (0xAA, '\u{2665}'), (0xAB, '\u{2660}'),
|
||||
(0xAC, '\u{2460}'), (0xAD, '\u{2461}'), (0xAE, '\u{2462}'), (0xAF, '\u{2463}'), (0xB0, '\u{2464}'), (0xB1, '\u{2465}'),
|
||||
(0xB2, '\u{2466}'), (0xB3, '\u{2467}'), (0xB4, '\u{2468}'), (0xB5, '\u{2469}'), (0xB6, '\u{2776}'), (0xB7, '\u{2777}'),
|
||||
(0xB8, '\u{2778}'), (0xB9, '\u{2779}'), (0xBA, '\u{277A}'), (0xBB, '\u{277B}'), (0xBC, '\u{277C}'), (0xBD, '\u{277D}'),
|
||||
(0xBE, '\u{277E}'), (0xBF, '\u{277F}'), (0xC0, '\u{2780}'), (0xC1, '\u{2781}'), (0xC2, '\u{2782}'), (0xC3, '\u{2783}'),
|
||||
(0xC4, '\u{2784}'), (0xC5, '\u{2785}'), (0xC6, '\u{2786}'), (0xC7, '\u{2787}'), (0xC8, '\u{2788}'), (0xC9, '\u{2789}'),
|
||||
(0xCA, '\u{278A}'), (0xCB, '\u{278B}'), (0xCC, '\u{278C}'), (0xCD, '\u{278D}'), (0xCE, '\u{278E}'), (0xCF, '\u{278F}'),
|
||||
(0xD0, '\u{2790}'), (0xD1, '\u{2791}'), (0xD2, '\u{2792}'), (0xD3, '\u{2793}'), (0xD4, '\u{2794}'), (0xD5, '\u{2192}'),
|
||||
(0xD6, '\u{2194}'), (0xD7, '\u{2195}'), (0xD8, '\u{2798}'), (0xD9, '\u{2799}'), (0xDA, '\u{279A}'), (0xDB, '\u{279B}'),
|
||||
(0xDC, '\u{279C}'), (0xDD, '\u{279D}'), (0xDE, '\u{279E}'), (0xDF, '\u{279F}'), (0xE0, '\u{27A0}'), (0xE1, '\u{27A1}'),
|
||||
(0xE2, '\u{27A2}'), (0xE3, '\u{27A3}'), (0xE4, '\u{27A4}'), (0xE5, '\u{27A5}'), (0xE6, '\u{27A6}'), (0xE7, '\u{27A7}'),
|
||||
(0xE8, '\u{27A8}'), (0xE9, '\u{27A9}'), (0xEA, '\u{27AA}'), (0xEB, '\u{27AB}'), (0xEC, '\u{27AC}'), (0xED, '\u{27AD}'),
|
||||
(0xEE, '\u{27AE}'), (0xEF, '\u{27AF}'), (0xF1, '\u{27B1}'), (0xF2, '\u{27B2}'), (0xF3, '\u{27B3}'), (0xF4, '\u{27B4}'),
|
||||
(0xF5, '\u{27B5}'), (0xF6, '\u{27B6}'), (0xF7, '\u{27B7}'), (0xF8, '\u{27B8}'), (0xF9, '\u{27B9}'), (0xFA, '\u{27BA}'),
|
||||
(0xFB, '\u{27BB}'), (0xFC, '\u{27BC}'), (0xFD, '\u{27BD}'), (0xFE, '\u{27BE}'),
|
||||
];
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn times_roman_ascii_widths() {
|
||||
assert_eq!(base14_char_width("Times-Roman", ' '), Some(250));
|
||||
assert_eq!(base14_char_width("Times-Roman", 'M'), Some(889));
|
||||
assert_eq!(base14_char_width("Times-Roman", 'i'), Some(278));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn subset_prefix_and_aliases_normalize() {
|
||||
assert_eq!(
|
||||
base14_char_width("ABCDEF+Times-Bold", ' '),
|
||||
base14_char_width("Times-Bold", ' ')
|
||||
);
|
||||
assert!(base14_char_width("ArialMT", 'a').is_some());
|
||||
assert!(base14_char_width("TimesNewRomanPSMT", 'a').is_some());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_base14_returns_none() {
|
||||
assert_eq!(base14_char_width("DejaVuSans", 'a'), None);
|
||||
assert!(!is_base14_font("Garamond"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn builtin_encoding_resolves_symbol_and_zapf_codes() {
|
||||
// Symbol 0x61 renders alpha; Zapf 0x21 renders U+2701.
|
||||
assert_eq!(builtin_encoding_char("Symbol", 0x61), Some('\u{03B1}'));
|
||||
assert_eq!(builtin_encoding_char("Symbol", 0xA5), Some('\u{221E}'));
|
||||
assert_eq!(
|
||||
builtin_encoding_char("ZapfDingbats", 0x21),
|
||||
Some('\u{2701}')
|
||||
);
|
||||
// Latin text fonts follow standard encodings — no builtin override.
|
||||
assert_eq!(builtin_encoding_char("Times-Roman", 0x61), None);
|
||||
// The resolved chars have real AFM widths.
|
||||
let alpha_w = base14_char_width("Symbol", '\u{03B1}');
|
||||
assert!(alpha_w.is_some() && alpha_w != Some(500));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn encoding_tables_are_sorted_for_binary_search() {
|
||||
for table in [SYMBOL_ENCODING, ZAPFDINGBATS_ENCODING] {
|
||||
assert!(table.windows(2).all(|w| w[0].0 < w[1].0));
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn tables_are_sorted_for_binary_search() {
|
||||
for table in [
|
||||
TIMES_ROMAN,
|
||||
TIMES_ITALIC,
|
||||
HELVETICA,
|
||||
COURIER,
|
||||
SYMBOL,
|
||||
ZAPFDINGBATS,
|
||||
] {
|
||||
assert!(table.windows(2).all(|w| w[0].0 < w[1].0));
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -14,9 +14,9 @@ use lopdf::{Document, Encoding, Object, ObjectId};
|
||||
use std::collections::HashMap;
|
||||
|
||||
use super::fonts::{
|
||||
build_font_encodings, build_font_widths, build_type3_scales, compute_string_width_ts,
|
||||
descriptor_style_flags, extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes,
|
||||
CMapDecisionCache, FontStyleCache,
|
||||
build_font_encodings, build_font_widths, compute_string_width_ts, descriptor_style_flags,
|
||||
extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
|
||||
FontStyleCache,
|
||||
};
|
||||
use super::underline::UnderlineLine;
|
||||
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
|
||||
@@ -162,11 +162,10 @@ pub(crate) fn extract_page_text_items(
|
||||
let fonts = doc.get_page_fonts(page_id).unwrap_or_default();
|
||||
|
||||
// Build font encoding maps from Differences arrays
|
||||
let (font_encodings, has_gid_fonts) = build_font_encodings(doc, &fonts, font_cmaps);
|
||||
let (font_encodings, has_gid_fonts) = build_font_encodings(doc, &fonts);
|
||||
|
||||
// Build font width info for accurate text positioning
|
||||
let font_widths = build_font_widths(doc, &fonts);
|
||||
let type3_scales = build_type3_scales(doc, &fonts);
|
||||
|
||||
// Build maps of font resource names to their base font names and ToUnicode object refs
|
||||
let mut font_base_names: std::collections::HashMap<String, String> =
|
||||
@@ -503,8 +502,7 @@ pub(crate) fn extract_page_text_items(
|
||||
) {
|
||||
let combined =
|
||||
multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm);
|
||||
let rendered_size = effective_font_size(current_font_size, &combined)
|
||||
* type3_scales.get(¤t_font).copied().unwrap_or(1.0);
|
||||
let rendered_size = effective_font_size(current_font_size, &combined);
|
||||
let (x, y) = (combined[4], combined[5]);
|
||||
if combined[0].abs() >= combined[1].abs() {
|
||||
rotation_votes.horizontal += 1;
|
||||
@@ -675,8 +673,7 @@ pub(crate) fn extract_page_text_items(
|
||||
} else {
|
||||
rotation_votes.rotated += 1;
|
||||
}
|
||||
let rendered_size = effective_font_size(current_font_size, &combined)
|
||||
* type3_scales.get(¤t_font).copied().unwrap_or(1.0);
|
||||
let rendered_size = effective_font_size(current_font_size, &combined);
|
||||
let base_font = font_base_names
|
||||
.get(¤t_font)
|
||||
.map(|s| s.as_str())
|
||||
@@ -783,8 +780,7 @@ pub(crate) fn extract_page_text_items(
|
||||
} else {
|
||||
rotation_votes.rotated += 1;
|
||||
}
|
||||
let rendered_size = effective_font_size(current_font_size, &combined)
|
||||
* type3_scales.get(¤t_font).copied().unwrap_or(1.0);
|
||||
let rendered_size = effective_font_size(current_font_size, &combined);
|
||||
let (x, y) = (combined[4], combined[5]);
|
||||
let width = w_ts_opt
|
||||
.map(|w_ts| {
|
||||
@@ -936,8 +932,7 @@ pub(crate) fn extract_page_text_items(
|
||||
} else {
|
||||
rotation_votes.rotated += 1;
|
||||
}
|
||||
let rendered_size = effective_font_size(current_font_size, &combined)
|
||||
* type3_scales.get(¤t_font).copied().unwrap_or(1.0);
|
||||
let rendered_size = effective_font_size(current_font_size, &combined);
|
||||
let (x, y) = (combined[4], combined[5]);
|
||||
// Width in device space from text matrix delta
|
||||
let delta_ts = text_matrix[4] - start_tm[4];
|
||||
|
||||
+13
-321
@@ -138,75 +138,6 @@ pub(crate) fn build_font_widths(
|
||||
widths
|
||||
}
|
||||
|
||||
/// Visual-size scale factors for Type3 fonts, keyed by resource name.
|
||||
///
|
||||
/// A Type3 font's glyph space maps to text space through FontMatrix, so the
|
||||
/// visual height of its glyphs is `nominal_size × |matrix_y| × FontBBox
|
||||
/// height`. For a well-behaved font (matrix 0.001, bbox ≈ 1000 units) that
|
||||
/// factor is ≈ 1.0 and the nominal size is already right. TeX PK bitmap
|
||||
/// fonts (dvips → Distiller) instead use FontMatrix [1 0 0 -1 0 0] with
|
||||
/// nominal sizes like 0.12, which makes every downstream font-size heuristic
|
||||
/// (drop caps, sub/superscripts, small-font tables, line heights) see
|
||||
/// nonsense. Fonts without a usable FontBBox are omitted (treated as 1.0).
|
||||
pub(crate) fn build_type3_scales(
|
||||
doc: &Document,
|
||||
fonts: &std::collections::BTreeMap<Vec<u8>, &lopdf::Dictionary>,
|
||||
) -> HashMap<String, f32> {
|
||||
let mut scales = HashMap::new();
|
||||
for (font_name, font_dict) in fonts {
|
||||
let is_type3 = font_dict
|
||||
.get(b"Subtype")
|
||||
.ok()
|
||||
.and_then(|o| o.as_name().ok())
|
||||
.is_some_and(|n| n == b"Type3");
|
||||
if !is_type3 {
|
||||
continue;
|
||||
}
|
||||
// Array elements may themselves be indirect references per PDF
|
||||
// syntax — resolve before reading the numeric value.
|
||||
let num = |o: &Object| {
|
||||
let resolved = match o {
|
||||
Object::Reference(r) => match doc.get_object(*r) {
|
||||
Ok(inner) => inner,
|
||||
Err(_) => return 0.0,
|
||||
},
|
||||
other => other,
|
||||
};
|
||||
match resolved {
|
||||
Object::Integer(i) => *i as f32,
|
||||
Object::Real(r) => *r,
|
||||
_ => 0.0,
|
||||
}
|
||||
};
|
||||
let Some(matrix) = font_dict
|
||||
.get(b"FontMatrix")
|
||||
.ok()
|
||||
.and_then(|o| resolve_array(doc, o))
|
||||
else {
|
||||
continue;
|
||||
};
|
||||
let Some(bbox) = font_dict
|
||||
.get(b"FontBBox")
|
||||
.ok()
|
||||
.and_then(|o| resolve_array(doc, o))
|
||||
else {
|
||||
continue;
|
||||
};
|
||||
if matrix.len() < 4 || bbox.len() < 4 {
|
||||
continue;
|
||||
}
|
||||
let scale_y = (num(&matrix[2]).powi(2) + num(&matrix[3]).powi(2)).sqrt();
|
||||
let bbox_h = (num(&bbox[3]) - num(&bbox[1])).abs();
|
||||
let scale = bbox_h * scale_y;
|
||||
// Only store meaningful scales; degenerate bboxes ([0 0 0 0] is legal)
|
||||
// and near-1.0 factors keep the nominal size.
|
||||
if scale > 0.01 && (scale - 1.0).abs() > 0.05 {
|
||||
scales.insert(String::from_utf8_lossy(font_name).to_string(), scale);
|
||||
}
|
||||
}
|
||||
scales
|
||||
}
|
||||
|
||||
/// Parse font widths from a font dictionary, dispatching by Subtype
|
||||
pub(crate) fn parse_font_widths(
|
||||
doc: &Document,
|
||||
@@ -218,71 +149,11 @@ pub(crate) fn parse_font_widths(
|
||||
|
||||
match subtype_name {
|
||||
b"Type0" => parse_type0_widths(doc, font_dict),
|
||||
b"Type1" | b"TrueType" | b"MMType1" => parse_simple_font_widths(doc, font_dict)
|
||||
.or_else(|| base14_fallback_widths(doc, font_dict)),
|
||||
b"Type3" => parse_simple_font_widths(doc, font_dict),
|
||||
b"Type1" | b"TrueType" | b"MMType1" | b"Type3" => parse_simple_font_widths(doc, font_dict),
|
||||
_ => None,
|
||||
}
|
||||
}
|
||||
|
||||
/// Fallback metrics for non-embedded base-14 fonts whose dictionary omits
|
||||
/// `/FirstChar`/`/Widths` (legal per the PDF spec — the reader must supply
|
||||
/// standard-font metrics). Without this, every glyph advances 0 and all
|
||||
/// downstream gap-based logic (space synthesis, script detection, table
|
||||
/// columns) collapses — common in 1990s dvips/Distiller PDFs.
|
||||
///
|
||||
/// Widths are resolved per code through the font's Differences encoding when
|
||||
/// present, falling back to the same single-byte decode the text extractor
|
||||
/// uses (cp1252-style smart punctuation for 0x80..=0x9F, Latin-1 elsewhere) —
|
||||
/// so the width of a code always matches the char we extract for it.
|
||||
fn base14_fallback_widths(doc: &Document, font_dict: &lopdf::Dictionary) -> Option<FontWidthInfo> {
|
||||
let base_font = font_dict
|
||||
.get(b"BaseFont")
|
||||
.ok()
|
||||
.and_then(|o| o.as_name().ok())
|
||||
.map(|n| String::from_utf8_lossy(n).to_string())?;
|
||||
if !crate::extractor::base14::is_base14_font(&base_font) {
|
||||
return None;
|
||||
}
|
||||
|
||||
let enc_map = parse_font_encoding(doc, font_dict)
|
||||
.map(|r| r.map)
|
||||
.unwrap_or_default();
|
||||
|
||||
let mut widths = HashMap::new();
|
||||
for code in 0u16..=255 {
|
||||
// Resolution order: Differences override, then the font's BUILT-IN
|
||||
// encoding (Symbol/ZapfDingbats glyphs live at positions unrelated
|
||||
// to cp1252 — the renderer draws α for Symbol 0x61 no matter how
|
||||
// the text decoder transliterates it, so the advance must be α's),
|
||||
// then the cp1252-style fallback used by the text decoder.
|
||||
let ch = enc_map
|
||||
.get(&(code as u8))
|
||||
.copied()
|
||||
.or_else(|| crate::extractor::base14::builtin_encoding_char(&base_font, code as u8))
|
||||
.unwrap_or_else(|| decode_single_byte_fallback_char(code as u8, true));
|
||||
if let Some(w) = crate::extractor::base14::base14_char_width(&base_font, ch) {
|
||||
widths.insert(code, w);
|
||||
}
|
||||
}
|
||||
let space_width = widths.get(&32).copied().unwrap_or(250);
|
||||
|
||||
debug!(
|
||||
" base14 fallback widths for {} ({} codes mapped)",
|
||||
base_font,
|
||||
widths.len()
|
||||
);
|
||||
|
||||
Some(FontWidthInfo {
|
||||
widths,
|
||||
default_width: 500,
|
||||
space_width,
|
||||
is_cid: false,
|
||||
units_scale: 0.001,
|
||||
wmode: 0,
|
||||
})
|
||||
}
|
||||
|
||||
/// Parse widths for simple fonts (Type1, TrueType, MMType1, Type3)
|
||||
/// Reads FirstChar, LastChar, and Widths array.
|
||||
/// For Type3 fonts, reads FontMatrix to determine the correct units_scale.
|
||||
@@ -626,13 +497,9 @@ pub(crate) fn get_operand_bytes(obj: &Object) -> Option<&[u8]> {
|
||||
/// Build encoding maps for all fonts on a page.
|
||||
/// Returns `(encodings, has_gid_fonts)` where `has_gid_fonts` is true when
|
||||
/// any font uses raw glyph ID names (gidNNNNN) that can't be decoded.
|
||||
/// Gid names whose codes the font's own ToUnicode CMap maps are decodable
|
||||
/// and do not set the flag (LibreOffice subsets write /gidNNNN Differences
|
||||
/// names alongside a complete ToUnicode CMap).
|
||||
pub(crate) fn build_font_encodings(
|
||||
doc: &Document,
|
||||
fonts: &std::collections::BTreeMap<Vec<u8>, &lopdf::Dictionary>,
|
||||
cmaps: &FontCMaps,
|
||||
) -> (PageFontEncodings, bool) {
|
||||
let mut encodings = PageFontEncodings::new();
|
||||
let mut has_gid_fonts = false;
|
||||
@@ -641,9 +508,7 @@ pub(crate) fn build_font_encodings(
|
||||
let resource_name = String::from_utf8_lossy(font_name).to_string();
|
||||
|
||||
if let Some(result) = parse_font_encoding(doc, font_dict) {
|
||||
if !result.gid_codes.is_empty()
|
||||
&& !tounicode_maps_codes(font_dict, cmaps, &result.gid_codes)
|
||||
{
|
||||
if result.gid_glyph_count > 0 {
|
||||
has_gid_fonts = true;
|
||||
}
|
||||
if !result.map.is_empty() {
|
||||
@@ -655,34 +520,6 @@ pub(crate) fn build_font_encodings(
|
||||
(encodings, has_gid_fonts)
|
||||
}
|
||||
|
||||
/// True when the font's ToUnicode CMap maps the gid-named character codes,
|
||||
/// so the Differences entries still decode through the CMap.
|
||||
fn tounicode_maps_codes(font_dict: &lopdf::Dictionary, cmaps: &FontCMaps, codes: &[u8]) -> bool {
|
||||
let Some(obj_ref) = font_dict
|
||||
.get(b"ToUnicode")
|
||||
.ok()
|
||||
.and_then(|o| o.as_reference().ok())
|
||||
else {
|
||||
return false;
|
||||
};
|
||||
let Some(entry) = cmaps.get_by_obj(obj_ref.0) else {
|
||||
return false;
|
||||
};
|
||||
// At least one gid code usably mapped means the CMap addresses these
|
||||
// codes; remaining unmapped codes are subset leftovers (e.g. the
|
||||
// component glyphs of an emoji ZWJ sequence mapped whole on its first
|
||||
// code). A mapping is usable only when extraction would accept it —
|
||||
// empty or U+FFFD results are rejected there as invalid. Fonts whose
|
||||
// CMap ignores the gid codes entirely stay flagged, and the downstream
|
||||
// garbage/encoding checks still catch partial damage.
|
||||
codes.iter().any(|&code| {
|
||||
entry
|
||||
.primary
|
||||
.lookup(code as u16)
|
||||
.is_some_and(|s| !s.is_empty() && !s.contains('\u{FFFD}'))
|
||||
})
|
||||
}
|
||||
|
||||
/// Parse font encoding from a font dictionary
|
||||
pub(crate) fn parse_font_encoding(
|
||||
doc: &Document,
|
||||
@@ -721,10 +558,11 @@ pub(crate) fn parse_font_encoding(
|
||||
/// Result of parsing an encoding dictionary's Differences array.
|
||||
pub(crate) struct EncodingResult {
|
||||
pub map: FontEncodingMap,
|
||||
/// Character codes whose glyph names match the `gidNNNNN` pattern (raw
|
||||
/// glyph IDs). These reference the original font's glyph table and are
|
||||
/// only decodable when the font's ToUnicode CMap maps the code.
|
||||
pub gid_codes: Vec<u8>,
|
||||
/// Number of glyph names matching the `gidNNNNN` pattern (raw glyph IDs).
|
||||
/// These indicate a font with unresolvable encoding — the glyph IDs
|
||||
/// reference the original font's glyph table, but without the original
|
||||
/// font's cmap there is no way to map them to Unicode.
|
||||
pub gid_glyph_count: u32,
|
||||
}
|
||||
|
||||
/// Parse an encoding dictionary with Differences array
|
||||
@@ -750,7 +588,7 @@ pub(crate) fn parse_encoding_dictionary(
|
||||
let mut encoding_map = FontEncodingMap::new();
|
||||
let mut current_code: u8 = 0;
|
||||
let mut ligature_count = 0u32;
|
||||
let mut gid_codes: Vec<u8> = Vec::new();
|
||||
let mut gid_glyph_count = 0u32;
|
||||
|
||||
for item in diff_array {
|
||||
match item {
|
||||
@@ -776,7 +614,7 @@ pub(crate) fn parse_encoding_dictionary(
|
||||
&& glyph_name.len() >= 4
|
||||
&& glyph_name[3..].chars().all(|c| c.is_ascii_digit())
|
||||
{
|
||||
gid_codes.push(current_code);
|
||||
gid_glyph_count += 1;
|
||||
}
|
||||
if let Some(ch) = mapped_char {
|
||||
encoding_map.insert(current_code, ch);
|
||||
@@ -800,16 +638,16 @@ pub(crate) fn parse_encoding_dictionary(
|
||||
);
|
||||
}
|
||||
|
||||
if !gid_codes.is_empty() {
|
||||
if gid_glyph_count > 0 {
|
||||
debug!(
|
||||
" Differences: {} gid-encoded glyphs (decodable only via ToUnicode)",
|
||||
gid_codes.len()
|
||||
" Differences: {} gid-encoded glyphs (unresolvable without original font)",
|
||||
gid_glyph_count
|
||||
);
|
||||
}
|
||||
|
||||
Some(EncodingResult {
|
||||
map: encoding_map,
|
||||
gid_codes,
|
||||
gid_glyph_count,
|
||||
})
|
||||
}
|
||||
|
||||
@@ -1599,43 +1437,6 @@ fn score_text(text: &str) -> i32 {
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
|
||||
#[test]
|
||||
fn type3_scale_resolves_indirect_matrix_and_bbox_numbers() {
|
||||
use lopdf::{dictionary, Document, Object};
|
||||
// FontMatrix/FontBBox elements may be indirect references per PDF
|
||||
// syntax; the scale must use their resolved values, not zero.
|
||||
let mut doc = Document::with_version("1.4");
|
||||
let matrix_d = doc.add_object(Object::Real(-1.0));
|
||||
let bbox_top = doc.add_object(Object::Integer(3));
|
||||
let font_dict = dictionary! {
|
||||
"Type" => "Font",
|
||||
"Subtype" => "Type3",
|
||||
"FontMatrix" => vec![
|
||||
Object::Integer(1),
|
||||
Object::Integer(0),
|
||||
Object::Integer(0),
|
||||
Object::Reference(matrix_d),
|
||||
Object::Integer(0),
|
||||
Object::Integer(0),
|
||||
],
|
||||
"FontBBox" => vec![
|
||||
Object::Integer(1),
|
||||
Object::Integer(-156),
|
||||
Object::Integer(37),
|
||||
Object::Reference(bbox_top),
|
||||
],
|
||||
};
|
||||
let mut fonts = std::collections::BTreeMap::new();
|
||||
fonts.insert(b"T2".to_vec(), &font_dict);
|
||||
let scales = super::build_type3_scales(&doc, &fonts);
|
||||
let scale = scales.get("T2").copied().unwrap_or(1.0);
|
||||
// bbox height 159 x |matrix_y| 1.0
|
||||
assert!(
|
||||
(scale - 159.0).abs() < 0.5,
|
||||
"scale should use resolved indirect values, got {scale}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn texcm_math_symbols_remap() {
|
||||
assert_eq!(
|
||||
@@ -2137,113 +1938,4 @@ mod tests {
|
||||
false
|
||||
));
|
||||
}
|
||||
|
||||
fn gid_font_doc(bfchar: Option<&str>) -> (Document, lopdf::ObjectId) {
|
||||
use lopdf::Stream;
|
||||
let mut doc = Document::with_version("1.4");
|
||||
let cmap = format!(
|
||||
"/CIDInit /ProcSet findresource begin
|
||||
12 dict begin
|
||||
begincmap
|
||||
1 begincodespacerange
|
||||
<00> <FF>
|
||||
endcodespacerange
|
||||
1 beginbfchar
|
||||
{}
|
||||
endbfchar
|
||||
endcmap
|
||||
CMapName currentdict /CMap defineresource pop
|
||||
end
|
||||
end",
|
||||
bfchar.unwrap_or_default()
|
||||
);
|
||||
let tounicode_id = doc.add_object(Object::Stream(Stream::new(
|
||||
dictionary! {},
|
||||
cmap.into_bytes(),
|
||||
)));
|
||||
let enc_id = doc.add_object(dictionary! {
|
||||
"Type" => "Encoding",
|
||||
"Differences" => vec![
|
||||
1.into(),
|
||||
Object::Name(b"gid1283".to_vec()),
|
||||
Object::Name(b"gid1464".to_vec()),
|
||||
],
|
||||
});
|
||||
let mut font = dictionary! {
|
||||
"Type" => "Font",
|
||||
"Subtype" => "TrueType",
|
||||
"BaseFont" => "ABCDEF+OpenSymbol",
|
||||
"Encoding" => Object::Reference(enc_id),
|
||||
};
|
||||
if bfchar.is_some() {
|
||||
font.set("ToUnicode", Object::Reference(tounicode_id));
|
||||
}
|
||||
let font_id = doc.add_object(font);
|
||||
let page_id = doc.add_object(dictionary! {
|
||||
"Type" => "Page",
|
||||
"Resources" => dictionary! {
|
||||
"Font" => dictionary! { "F1" => Object::Reference(font_id) },
|
||||
},
|
||||
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
|
||||
});
|
||||
let pages_id = doc.add_object(dictionary! {
|
||||
"Type" => "Pages",
|
||||
"Count" => Object::Integer(1),
|
||||
"Kids" => vec![Object::Reference(page_id)],
|
||||
});
|
||||
let catalog_id = doc.add_object(dictionary! {
|
||||
"Type" => "Catalog",
|
||||
"Pages" => Object::Reference(pages_id),
|
||||
});
|
||||
doc.trailer.set("Root", Object::Reference(catalog_id));
|
||||
(doc, page_id)
|
||||
}
|
||||
|
||||
fn gid_flagged(bfchar: Option<&str>) -> bool {
|
||||
let (doc, page_id) = gid_font_doc(bfchar);
|
||||
let cmaps = FontCMaps::from_doc(&doc);
|
||||
let fonts = doc.get_page_fonts(page_id).unwrap();
|
||||
let (_, has_gid_fonts) = build_font_encodings(&doc, &fonts, &cmaps);
|
||||
has_gid_fonts
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn gid_differences_with_covering_tounicode_are_not_flagged() {
|
||||
// LibreOffice subsets write /gidNNNN Differences names alongside a
|
||||
// ToUnicode CMap that decodes those codes; the page must not be
|
||||
// flagged as unresolvable (which would suppress the whole document's
|
||||
// markdown when every page carries such a font).
|
||||
assert!(!gid_flagged(Some("<01> <2022>\n<02> <25E6>")));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn gid_differences_with_partial_tounicode_are_not_flagged() {
|
||||
// An emoji ZWJ sequence maps whole on its first code; the remaining
|
||||
// component-glyph codes are subset leftovers, not damage.
|
||||
assert!(!gid_flagged(Some(
|
||||
"<01> <D83DDC68200DD83DDC69200DD83DDC67>"
|
||||
)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn gid_differences_without_tounicode_are_flagged() {
|
||||
assert!(
|
||||
gid_flagged(None),
|
||||
"gid glyphs without ToUnicode are unresolvable"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn gid_differences_with_disjoint_tounicode_are_flagged() {
|
||||
// A ToUnicode that never addresses the gid codes leaves them
|
||||
// unresolvable.
|
||||
assert!(gid_flagged(Some("<10> <0041>")));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn gid_differences_with_replacement_char_tounicode_are_flagged() {
|
||||
// A mapping to U+FFFD is not usable — extraction rejects it as an
|
||||
// invalid CMap result — so it must not clear the gid flag.
|
||||
assert!(gid_flagged(Some("<01> <FFFD>\n<02> <FFFD>")));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
//!
|
||||
//! This module extracts text with position information for structure detection.
|
||||
|
||||
mod base14;
|
||||
pub(crate) mod content_stream;
|
||||
mod fonts;
|
||||
mod layout;
|
||||
|
||||
@@ -8,9 +8,8 @@ use lopdf::{Document, Encoding, Object, ObjectId};
|
||||
use std::collections::HashMap;
|
||||
|
||||
use super::fonts::{
|
||||
build_font_encodings, build_font_widths, build_type3_scales, compute_string_width_ts,
|
||||
extract_text_from_operand, get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache,
|
||||
FontStyleCache,
|
||||
build_font_encodings, build_font_widths, compute_string_width_ts, extract_text_from_operand,
|
||||
get_font_file2_obj_num, get_operand_bytes, CMapDecisionCache, FontStyleCache,
|
||||
};
|
||||
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
|
||||
|
||||
@@ -163,11 +162,10 @@ fn extract_form_xobject_text_inner(
|
||||
|
||||
// Get fonts from the Form's Resources
|
||||
let form_fonts = get_form_fonts(doc, &stream.dict);
|
||||
let (font_encodings, _has_gid_fonts) = build_font_encodings(doc, &form_fonts, font_cmaps);
|
||||
let (font_encodings, _has_gid_fonts) = build_font_encodings(doc, &form_fonts);
|
||||
|
||||
// Build font width info for the form
|
||||
let font_widths = build_font_widths(doc, &form_fonts);
|
||||
let type3_scales = build_type3_scales(doc, &form_fonts);
|
||||
|
||||
// Build font base names and ToUnicode refs for the form
|
||||
let mut font_base_names: HashMap<String, String> = HashMap::new();
|
||||
@@ -415,8 +413,7 @@ fn extract_form_xobject_text_inner(
|
||||
&font_widths,
|
||||
) {
|
||||
let combined = multiply_matrices(&text_matrix, &ctm);
|
||||
let rendered_size = effective_font_size(current_font_size, &combined)
|
||||
* type3_scales.get(¤t_font).copied().unwrap_or(1.0);
|
||||
let rendered_size = effective_font_size(current_font_size, &combined);
|
||||
let (x, y) = (combined[4], combined[5]);
|
||||
let width = if let Some(font_info) = font_widths.get(¤t_font) {
|
||||
if let Some(raw_bytes) = get_operand_bytes(&op.operands[0]) {
|
||||
@@ -575,8 +572,7 @@ fn extract_form_xobject_text_inner(
|
||||
}
|
||||
if !sub_items.is_empty() {
|
||||
let combined = multiply_matrices(&text_matrix, &ctm);
|
||||
let rendered_size = effective_font_size(current_font_size, &combined)
|
||||
* type3_scales.get(¤t_font).copied().unwrap_or(1.0);
|
||||
let rendered_size = effective_font_size(current_font_size, &combined);
|
||||
let base_font = font_base_names
|
||||
.get(¤t_font)
|
||||
.map(|s| s.as_str())
|
||||
|
||||
@@ -5614,6 +5614,8 @@ pub enum PdfError {
|
||||
Parse(String),
|
||||
#[error("PDF is encrypted")]
|
||||
Encrypted,
|
||||
#[error("Invalid PDF options: {0}")]
|
||||
InvalidOptions(String),
|
||||
#[error("Invalid PDF structure")]
|
||||
InvalidStructure,
|
||||
#[error("Not a PDF: {0}")]
|
||||
|
||||
+12
-84
@@ -368,95 +368,23 @@ pub(crate) fn compute_paragraph_threshold(lines: &[TextLine], base_size: f32) ->
|
||||
/// Discover distinct heading font-size tiers in the document.
|
||||
/// Returns tiers sorted largest-first (tier 0 = H1, tier 1 = H2, …).
|
||||
/// Sizes within 0.5pt are clustered into the same tier. Capped at 4 tiers.
|
||||
/// Dominant font size of a line: the size cluster (0.5pt tolerance) covering
|
||||
/// the most alphanumeric characters. A line led by an oversized ornament
|
||||
/// (drop cap, Type3 display-math delimiter) must be classified by its body
|
||||
/// text, not its first glyph — and operators/delimiters don't get a vote,
|
||||
/// so a display equation whose parens and plus signs render larger than its
|
||||
/// variables still classifies at the variables' size. Falls back to counting
|
||||
/// all non-whitespace chars for lines with no alphanumerics. Returns `None`
|
||||
/// for lines with no text.
|
||||
pub(crate) fn line_dominant_font_size(line: &TextLine) -> Option<f32> {
|
||||
fn dominant(line: &TextLine, count_chars: fn(&str) -> usize) -> Option<f32> {
|
||||
let mut clusters: Vec<(f32, usize)> = Vec::new();
|
||||
for item in &line.items {
|
||||
let chars = count_chars(&item.text);
|
||||
if chars == 0 {
|
||||
continue;
|
||||
}
|
||||
if let Some((_, count)) = clusters
|
||||
.iter_mut()
|
||||
.find(|(s, _)| (*s - item.font_size).abs() < 0.5)
|
||||
{
|
||||
*count += chars;
|
||||
} else {
|
||||
clusters.push((item.font_size, chars));
|
||||
}
|
||||
}
|
||||
clusters
|
||||
.into_iter()
|
||||
.max_by_key(|&(_, count)| count)
|
||||
.map(|(size, _)| size)
|
||||
}
|
||||
dominant(line, |t| t.chars().filter(|c| c.is_alphanumeric()).count())
|
||||
.or_else(|| dominant(line, |t| t.chars().filter(|c| !c.is_whitespace()).count()))
|
||||
}
|
||||
|
||||
pub(crate) fn compute_heading_tiers(lines: &[TextLine], base_size: f32) -> Vec<f32> {
|
||||
let mut heading_sizes: Vec<f32> = Vec::new();
|
||||
|
||||
for line in lines {
|
||||
// Use the dominant (alphanumeric-weighted) size — the same notion
|
||||
// heading DETECTION matches against tiers — so a large leading
|
||||
// ornament or section marker can't register a tier at a size no
|
||||
// line will ever be classified with.
|
||||
let Some(dominant) = line_dominant_font_size(line) else {
|
||||
continue;
|
||||
};
|
||||
if dominant / base_size >= 1.2 {
|
||||
// Digit-only lines (page numbers, issue numbers) must not
|
||||
// define heading tiers: a large bold folio claims tier 0 and
|
||||
// blocks the bold-size fallback for the document's real
|
||||
// same-size headings.
|
||||
let text = line.text();
|
||||
let t = text.trim();
|
||||
if !t.is_empty() && t.chars().all(|c| !c.is_alphabetic()) {
|
||||
continue;
|
||||
if let Some(first) = line.items.first() {
|
||||
if first.font_size / base_size >= 1.2 {
|
||||
// Digit-only lines (page numbers, issue numbers) must not
|
||||
// define heading tiers: a large bold folio claims tier 0 and
|
||||
// blocks the bold-size fallback for the document's real
|
||||
// same-size headings.
|
||||
let text = line.text();
|
||||
let t = text.trim();
|
||||
if !t.is_empty() && t.chars().all(|c| !c.is_alphabetic()) {
|
||||
continue;
|
||||
}
|
||||
heading_sizes.push(first.font_size);
|
||||
}
|
||||
// Near-empty lines (oversized glyph runs — display-math
|
||||
// operators like ∫∫, ⟨⟩ or √ that decode as "ZZ"/"h i"/"p p"
|
||||
// in TeX bitmap fonts, drop caps) must not define tiers
|
||||
// either. Real headings have at least three letters, two of
|
||||
// them distinct.
|
||||
let mut letters: Vec<char> = t
|
||||
.chars()
|
||||
.filter(|c| c.is_alphabetic())
|
||||
.flat_map(|c| c.to_lowercase())
|
||||
.collect();
|
||||
let total_letters = letters.len();
|
||||
letters.sort_unstable();
|
||||
letters.dedup();
|
||||
if total_letters < 3 || letters.len() < 2 {
|
||||
continue;
|
||||
}
|
||||
// A heading line is uniformly sized: every item within ±20% of
|
||||
// the dominant size. Mixed-size lines (a drop cap or oversized
|
||||
// math glyph anywhere in the line — leading OR trailing) are
|
||||
// body content, not headings.
|
||||
let uniform = line
|
||||
.items
|
||||
.iter()
|
||||
.all(|i| (i.font_size - dominant).abs() <= dominant * 0.2);
|
||||
if !uniform {
|
||||
continue;
|
||||
}
|
||||
log::debug!(
|
||||
"heading tier candidate: fs={:.1} page={} {:?}",
|
||||
dominant,
|
||||
line.page,
|
||||
t.chars().take(60).collect::<String>()
|
||||
);
|
||||
heading_sizes.push(dominant);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
+4
-173
@@ -458,83 +458,6 @@ fn is_body_size_all_bold_line(line: &TextLine, base_size: f32) -> bool {
|
||||
.all(|item| item.is_bold && (item.font_size - first.font_size).abs() < 0.5)
|
||||
}
|
||||
|
||||
/// First-line-indent paragraph break. Justified documents (TeX/dvips output)
|
||||
/// often separate paragraphs only by indenting the first line — the vertical
|
||||
/// gap stays at normal line spacing, so the Y-gap rule never fires and whole
|
||||
/// sections merge into one wall of text.
|
||||
///
|
||||
/// Fires only when all of:
|
||||
/// - the line advance is a normal in-paragraph step (`0 < y_gap <= para_threshold`)
|
||||
/// - the left margin is established (2+ consecutive lines at the same X), so
|
||||
/// list continuations and ragged layouts don't trigger
|
||||
/// - this line starts 0.8–4 em deeper than that margin (a first-line indent,
|
||||
/// not a column jump)
|
||||
/// - the previous line ended short of the established right edge (a paragraph's
|
||||
/// last line is ragged; hanging-indent wraps like bibliography continuation
|
||||
/// lines follow a full-width line and must not split) — OR ended at a
|
||||
/// sentence boundary, which covers paragraphs whose justified last line
|
||||
/// happens to run flush (a citation's first line almost never wraps exactly
|
||||
/// at a sentence end)
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
fn is_indent_paragraph_break(
|
||||
line_x: f32,
|
||||
y_gap: f32,
|
||||
para_threshold: f32,
|
||||
base_size: f32,
|
||||
margin_x: f32,
|
||||
margin_run: u32,
|
||||
prev_right: f32,
|
||||
max_right: f32,
|
||||
prev_ends_sentence: bool,
|
||||
) -> bool {
|
||||
let indent = line_x - margin_x;
|
||||
margin_run >= 2
|
||||
&& y_gap > 0.0
|
||||
&& y_gap <= para_threshold
|
||||
&& indent >= base_size * 0.8
|
||||
&& indent <= base_size * 4.0
|
||||
&& (prev_right <= max_right - base_size || prev_ends_sentence)
|
||||
}
|
||||
|
||||
/// True when a line's text ends at a sentence boundary: terminal punctuation,
|
||||
/// optionally followed by a closing quote/bracket.
|
||||
fn line_ends_sentence(text: &str) -> bool {
|
||||
let trimmed = text.trim_end();
|
||||
let trimmed = trimmed.trim_end_matches(['"', '\u{201D}', '\u{2019}', ')', ']']);
|
||||
trimmed.ends_with(['.', '!', '?'])
|
||||
}
|
||||
|
||||
/// True when math/operator symbols rival the alphanumeric content of a line —
|
||||
/// the signature of a display equation rather than prose or a heading.
|
||||
fn is_symbol_dominated(text: &str) -> bool {
|
||||
let alnum = text.chars().filter(|c| c.is_alphanumeric()).count();
|
||||
// Grouping characters are deliberately excluded: parenthesized short
|
||||
// headings like "(R-12)" or "[DRAFT]" are common and are not math.
|
||||
let mathy = text
|
||||
.chars()
|
||||
.filter(|c| {
|
||||
matches!(
|
||||
c,
|
||||
'+' | '\u{2212}'
|
||||
| '='
|
||||
| '/'
|
||||
| '^'
|
||||
| '|'
|
||||
| '<'
|
||||
| '>'
|
||||
| '\u{00B1}'
|
||||
| '\u{00D7}'
|
||||
| '\u{00B7}'
|
||||
| '\u{2211}'
|
||||
| '\u{222B}'
|
||||
| '\u{221A}'
|
||||
| '\u{221E}'
|
||||
)
|
||||
})
|
||||
.count();
|
||||
mathy >= 2 && mathy * 2 >= alnum.max(1)
|
||||
}
|
||||
|
||||
fn is_wrapped_same_style_line(prev: &TextLine, next: &TextLine, para_threshold: f32) -> bool {
|
||||
if prev.page != next.page {
|
||||
return false;
|
||||
@@ -854,12 +777,6 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
let mut toc_suppress_page: Option<u32> = None;
|
||||
let mut inserted_tables: HashSet<(u32, usize)> = HashSet::new();
|
||||
let mut inserted_images: HashSet<(u32, usize)> = HashSet::new();
|
||||
// First-line-indent paragraph detection state
|
||||
let mut left_margin_x = f32::MAX;
|
||||
let mut left_margin_run = 0u32;
|
||||
let mut prev_right = 0.0f32;
|
||||
let mut page_max_right = 0.0f32;
|
||||
let mut prev_ends_sentence = false;
|
||||
|
||||
// Collect all pages that have tables or images (including image-only pages)
|
||||
let mut all_content_pages: Vec<u32> = page_tables
|
||||
@@ -934,11 +851,6 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
prev_y = f32::MAX;
|
||||
prev_x = 0.0;
|
||||
paragraph_in_wrapped_bold_run = false;
|
||||
left_margin_x = f32::MAX;
|
||||
left_margin_run = 0;
|
||||
prev_right = 0.0;
|
||||
page_max_right = 0.0;
|
||||
prev_ends_sentence = false;
|
||||
|
||||
if options.include_page_numbers {
|
||||
output.push_str(&format!("<!-- Page {} -->\n\n", current_page));
|
||||
@@ -995,22 +907,7 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
&& !line_all_bold
|
||||
&& y_gap > base_size * 1.2
|
||||
&& y_gap <= para_threshold;
|
||||
let is_indent_break = in_paragraph
|
||||
&& !in_list
|
||||
&& is_indent_paragraph_break(
|
||||
line_x,
|
||||
y_gap,
|
||||
para_threshold,
|
||||
base_size,
|
||||
left_margin_x,
|
||||
left_margin_run,
|
||||
prev_right,
|
||||
page_max_right,
|
||||
prev_ends_sentence,
|
||||
);
|
||||
if (is_para_break || is_band_switch || is_bold_to_regular_break || is_indent_break)
|
||||
&& in_paragraph
|
||||
{
|
||||
if (is_para_break || is_band_switch || is_bold_to_regular_break) && in_paragraph {
|
||||
output.push_str("\n\n");
|
||||
in_paragraph = false;
|
||||
paragraph_in_wrapped_bold_run = false;
|
||||
@@ -1019,21 +916,6 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
// Let the continuation check below decide if we're still in a list
|
||||
prev_y = line.y;
|
||||
prev_x = line_x;
|
||||
// Track the stable left margin and line right edges for
|
||||
// first-line-indent paragraph detection
|
||||
if (line_x - left_margin_x).abs() <= 2.0 {
|
||||
left_margin_run += 1;
|
||||
} else {
|
||||
left_margin_x = line_x;
|
||||
left_margin_run = 1;
|
||||
}
|
||||
prev_right = line
|
||||
.items
|
||||
.iter()
|
||||
.map(|i| i.x + i.width)
|
||||
.fold(0.0f32, f32::max);
|
||||
page_max_right = page_max_right.max(prev_right);
|
||||
prev_ends_sentence = line_ends_sentence(&line.text());
|
||||
|
||||
// Get text with optional bold/italic formatting
|
||||
let text = line.text_with_formatting(
|
||||
@@ -1125,15 +1007,8 @@ pub(super) fn to_markdown_from_lines_with_tables_and_images(
|
||||
&& !is_toc_entry_line(plain_trimmed)
|
||||
&& !is_heading_fragment(plain_trimmed)
|
||||
&& toc_suppress_page != Some(line.page)
|
||||
// Display equations must never become headings — through the
|
||||
// font-tier path or the rarity fallback. They are standalone and
|
||||
// isolated by construction, so without this gate they score
|
||||
// straight past both.
|
||||
&& !plain_trimmed.contains('=')
|
||||
&& !is_symbol_dominated(plain_trimmed)
|
||||
{
|
||||
let line_font_size =
|
||||
crate::markdown::analysis::line_dominant_font_size(line).unwrap_or(base_size);
|
||||
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
|
||||
detect_header_level(
|
||||
line_font_size,
|
||||
base_size,
|
||||
@@ -1414,12 +1289,6 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
|
||||
let mut prev_had_dot_leaders = false;
|
||||
let mut paragraph_in_wrapped_bold_run = false;
|
||||
let mut toc_suppress_page: Option<u32> = None;
|
||||
// First-line-indent paragraph detection state
|
||||
let mut left_margin_x = f32::MAX;
|
||||
let mut left_margin_run = 0u32;
|
||||
let mut prev_right = 0.0f32;
|
||||
let mut page_max_right = 0.0f32;
|
||||
let mut prev_ends_sentence = false;
|
||||
|
||||
for (line_idx, line) in lines.iter().enumerate() {
|
||||
// Page break
|
||||
@@ -1437,11 +1306,6 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
|
||||
last_list_x = None;
|
||||
prev_had_dot_leaders = false;
|
||||
paragraph_in_wrapped_bold_run = false;
|
||||
left_margin_x = f32::MAX;
|
||||
left_margin_run = 0;
|
||||
prev_right = 0.0;
|
||||
page_max_right = 0.0;
|
||||
prev_ends_sentence = false;
|
||||
|
||||
if options.include_page_numbers {
|
||||
output.push_str(&format!("<!-- Page {} -->\n\n", current_page));
|
||||
@@ -1451,7 +1315,6 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
|
||||
// Paragraph break: large forward Y gap (normal) or large backward jump
|
||||
// (newspaper columns emitted sequentially on the same page).
|
||||
let y_gap = prev_y - line.y;
|
||||
let line_x = line.items.first().map(|i| i.x).unwrap_or(0.0);
|
||||
let is_para_break = y_gap.abs() > para_threshold;
|
||||
let line_all_bold = !line.items.is_empty() && line.items.iter().all(|item| item.is_bold);
|
||||
let line_in_wrapped_bold_run = wrapped_bold_paragraph_lines.contains(&line_idx);
|
||||
@@ -1461,20 +1324,7 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
|
||||
&& !line_all_bold
|
||||
&& y_gap > base_size * 1.2
|
||||
&& y_gap <= para_threshold;
|
||||
let is_indent_break = in_paragraph
|
||||
&& !in_list
|
||||
&& is_indent_paragraph_break(
|
||||
line_x,
|
||||
y_gap,
|
||||
para_threshold,
|
||||
base_size,
|
||||
left_margin_x,
|
||||
left_margin_run,
|
||||
prev_right,
|
||||
page_max_right,
|
||||
prev_ends_sentence,
|
||||
);
|
||||
if (is_para_break || is_bold_to_regular_break || is_indent_break) && in_paragraph {
|
||||
if (is_para_break || is_bold_to_regular_break) && in_paragraph {
|
||||
output.push_str("\n\n");
|
||||
in_paragraph = false;
|
||||
paragraph_in_wrapped_bold_run = false;
|
||||
@@ -1482,21 +1332,6 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
|
||||
// Don't immediately end list on paragraph break
|
||||
// Let the continuation check below decide if we're still in a list
|
||||
prev_y = line.y;
|
||||
// Track the stable left margin and line right edges for
|
||||
// first-line-indent paragraph detection
|
||||
if (line_x - left_margin_x).abs() <= 2.0 {
|
||||
left_margin_run += 1;
|
||||
} else {
|
||||
left_margin_x = line_x;
|
||||
left_margin_run = 1;
|
||||
}
|
||||
prev_right = line
|
||||
.items
|
||||
.iter()
|
||||
.map(|i| i.x + i.width)
|
||||
.fold(0.0f32, f32::max);
|
||||
page_max_right = page_max_right.max(prev_right);
|
||||
prev_ends_sentence = line_ends_sentence(&line.text());
|
||||
|
||||
// Get text with optional bold/italic formatting
|
||||
let text = line.text_with_formatting(
|
||||
@@ -1536,12 +1371,8 @@ pub fn to_markdown_from_lines(lines: Vec<TextLine>, options: MarkdownOptions) ->
|
||||
&& !is_heading_fragment(plain_trimmed)
|
||||
&& toc_suppress_page != Some(line.page)
|
||||
&& !(options.detect_code && line.items.iter().any(|i| is_monospace_font(&i.font)))
|
||||
// Display equations must never become headings (see main path).
|
||||
&& !plain_trimmed.contains('=')
|
||||
&& !is_symbol_dominated(plain_trimmed)
|
||||
{
|
||||
let line_font_size =
|
||||
crate::markdown::analysis::line_dominant_font_size(line).unwrap_or(base_size);
|
||||
let line_font_size = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
|
||||
if let Some(header_level) = detect_header_level(
|
||||
line_font_size,
|
||||
base_size,
|
||||
|
||||
+3
-133
@@ -40,9 +40,8 @@ fn effective_heading_level(
|
||||
}
|
||||
}
|
||||
|
||||
// Fall back to font-size heuristic (dominant size, so a leading drop cap
|
||||
// or oversized math delimiter doesn't reclassify a body line)
|
||||
let font = crate::markdown::analysis::line_dominant_font_size(line).unwrap_or(base_size);
|
||||
// Fall back to font-size heuristic
|
||||
let font = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
|
||||
detect_header_level(
|
||||
font,
|
||||
base_size,
|
||||
@@ -73,10 +72,7 @@ pub(crate) fn merge_heading_lines(
|
||||
|
||||
for line in lines {
|
||||
let line_level = effective_heading_level(&line, base_size, heading_tiers, struct_roles);
|
||||
// Dominant size, consistent with effective_heading_level: a large
|
||||
// leading delimiter must not inflate the merge gap threshold.
|
||||
let line_font =
|
||||
crate::markdown::analysis::line_dominant_font_size(&line).unwrap_or(base_size);
|
||||
let line_font = line.items.first().map(|i| i.font_size).unwrap_or(base_size);
|
||||
|
||||
// Check if the previous line is a heading at the same level on the same page
|
||||
let should_merge = if let (Some(prev), Some(curr_level)) = (result.last(), line_level) {
|
||||
@@ -173,87 +169,6 @@ pub(crate) fn merge_drop_caps(lines: Vec<TextLine>, base_size: f32) -> Vec<TextL
|
||||
.map(|c| c.is_uppercase())
|
||||
.unwrap_or(false);
|
||||
|
||||
// Embedded drop cap: a two-line drop cap's baseline aligns with the
|
||||
// paragraph's SECOND line, so Y-grouping puts the glyph at the start
|
||||
// of that line instead of on its own line. Left in place, the cap
|
||||
// ends up mid-sentence after paragraph joining ("...which exchange T
|
||||
// bandwidth..."). Detect it, prepend the char to the previous line
|
||||
// (the paragraph's first line, typeset beside the cap), and drop it.
|
||||
// The size gate is 1.8x (not 2.5x): bitmap (Type3) caps measure their
|
||||
// glyph bbox, not the em box, so a two-line cap can be as small as
|
||||
// ~1.9x the body size.
|
||||
if line.items.len() > 1 {
|
||||
let first = &line.items[0];
|
||||
// The rest of the line must be a substantive body-text run —
|
||||
// a lone label or math fragment beside a big glyph is not a
|
||||
// drop-cap paragraph and must not be rewritten.
|
||||
let rest_letters: usize = line.items[1..]
|
||||
.iter()
|
||||
.map(|i| i.text.chars().filter(|c| c.is_alphabetic()).count())
|
||||
.sum();
|
||||
let is_embedded_cap = first.font_size >= base_size * 1.8
|
||||
&& first.text.trim().chars().count() == 1
|
||||
&& first
|
||||
.text
|
||||
.trim()
|
||||
.chars()
|
||||
.next()
|
||||
.is_some_and(|c| c.is_uppercase())
|
||||
&& line.items[1..]
|
||||
.iter()
|
||||
.all(|i| i.font_size < base_size * 1.5)
|
||||
&& line.items[1..].iter().any(|i| i.x > first.x)
|
||||
&& rest_letters >= 8;
|
||||
if is_embedded_cap {
|
||||
let drop_char = first.text.trim().chars().next().unwrap();
|
||||
let cap_x = first.x;
|
||||
let cap_span = first.font_size * 1.5;
|
||||
let line_y = line.y;
|
||||
// Only the IMMEDIATELY preceding line qualifies: the
|
||||
// paragraph's first line sits directly above, in the same
|
||||
// column, indented past the cap glyph by roughly the cap's
|
||||
// width. It must also read as a paragraph first line —
|
||||
// substantive body text starting with a letter — so
|
||||
// headings, labels, or table fragments that merely fall in
|
||||
// the geometric window are never rewritten.
|
||||
let target = result.last_mut().filter(|prev| {
|
||||
let prev_text = prev.text();
|
||||
let prev_trimmed = prev_text.trim();
|
||||
prev.page == line.page
|
||||
&& prev.y > line_y
|
||||
&& prev.y - line_y <= cap_span
|
||||
&& prev
|
||||
.items
|
||||
.first()
|
||||
.is_some_and(|i| i.x > cap_x && i.x - cap_x <= first.font_size * 2.5)
|
||||
&& prev_trimmed
|
||||
.chars()
|
||||
.next()
|
||||
.is_some_and(|c| c.is_alphabetic())
|
||||
&& prev_trimmed.chars().filter(|c| c.is_alphabetic()).count() >= 8
|
||||
});
|
||||
if let Some(prev_line) = target {
|
||||
if let Some(first_item) = prev_line.items.first_mut() {
|
||||
// A mid-word cap ("T" + "HE recent…") joins directly.
|
||||
// Leading whitespace on the remainder means the cap
|
||||
// is a standalone word ("A" + " long time ago…") —
|
||||
// preserve exactly one separating space.
|
||||
let had_leading_ws = first_item.text.starts_with(char::is_whitespace);
|
||||
let rest = first_item.text.trim_start().to_string();
|
||||
first_item.text = if had_leading_ws {
|
||||
format!("{} {}", drop_char, rest)
|
||||
} else {
|
||||
format!("{}{}", drop_char, rest)
|
||||
};
|
||||
}
|
||||
let mut line = line.clone();
|
||||
line.items.remove(0);
|
||||
result.push(line);
|
||||
continue;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if is_drop_cap {
|
||||
let drop_char = trimmed.chars().next().unwrap();
|
||||
|
||||
@@ -683,51 +598,6 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
fn make_item_at(text: &str, font_size: f32, x: f32) -> TextItem {
|
||||
let mut item = make_item(text, font_size, None);
|
||||
item.x = x;
|
||||
item.width = text.len() as f32 * font_size * 0.5;
|
||||
item
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_embedded_drop_cap_moves_to_paragraph_start() {
|
||||
// Two-line drop cap: the big "T" baseline-aligns with the paragraph's
|
||||
// second line, so it lands as that line's first item (Shannon
|
||||
// entropy.pdf page 1 pattern).
|
||||
let first_line = TextLine {
|
||||
items: vec![make_item_at(
|
||||
"HE recent development which exchange",
|
||||
10.0,
|
||||
90.0,
|
||||
)],
|
||||
y: 712.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let second_line = TextLine {
|
||||
items: vec![
|
||||
make_item_at("T", 25.0, 72.0),
|
||||
make_item_at("bandwidth for signal-to-noise ratio", 10.0, 90.0),
|
||||
],
|
||||
y: 700.0,
|
||||
page: 1,
|
||||
adaptive_threshold: 0.10,
|
||||
};
|
||||
let result = merge_drop_caps(vec![first_line, second_line], 10.0);
|
||||
assert_eq!(result.len(), 2);
|
||||
assert!(
|
||||
result[0].text().starts_with("THE recent"),
|
||||
"cap should prepend to paragraph start: {}",
|
||||
result[0].text()
|
||||
);
|
||||
assert!(
|
||||
result[1].text().starts_with("bandwidth"),
|
||||
"cap item should be removed from second line: {}",
|
||||
result[1].text()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_merge_struct_tree_headings() {
|
||||
// Two consecutive lines tagged as H2 via struct tree, same font size as body
|
||||
|
||||
@@ -137,71 +137,6 @@ fn expand_consolidated_items(items: &[TextItem]) -> (Vec<TextItem>, Vec<usize>)
|
||||
(expanded, index_map)
|
||||
}
|
||||
|
||||
/// Index of candidate "body" items (larger-font attachment targets) sorted by
|
||||
/// Y, so script-attachment checks scan a narrow Y window instead of the whole
|
||||
/// page per candidate.
|
||||
struct ScriptBodyIndex<'a> {
|
||||
/// (y, item), sorted ascending by y
|
||||
by_y: Vec<(f32, &'a TextItem)>,
|
||||
/// widest vertical attachment window any body item can produce
|
||||
max_window: f32,
|
||||
}
|
||||
|
||||
impl<'a> ScriptBodyIndex<'a> {
|
||||
fn new(items: &'a [TextItem]) -> Self {
|
||||
// Smallest table-candidate font is 6pt, so any possible attachment
|
||||
// target is at least 6 × 1.2 pt.
|
||||
let mut by_y: Vec<(f32, &TextItem)> = items
|
||||
.iter()
|
||||
.filter(|i| i.font_size >= 6.0 * 1.2)
|
||||
.map(|i| (i.y, i))
|
||||
.collect();
|
||||
by_y.sort_by(|a, b| a.0.total_cmp(&b.0));
|
||||
let max_window = by_y
|
||||
.iter()
|
||||
.map(|(_, i)| i.font_size * 0.8)
|
||||
.fold(0.0f32, f32::max);
|
||||
Self { by_y, max_window }
|
||||
}
|
||||
|
||||
/// True when a small-font item is horizontally attached to a larger-font
|
||||
/// item at a script baseline offset — a sub/superscript in running text
|
||||
/// or math (equation subscripts, footnote markers). Script attachments
|
||||
/// are not table cells; without this filter, display equations with
|
||||
/// sub/superscripts form phantom small-font table regions (e.g. TeX
|
||||
/// papers where log₁₀ subscripts cluster with footnote lines into a fake
|
||||
/// 3-column table). A genuine baseline offset is required so same-line
|
||||
/// table neighbors (a small cell beside a larger label cell) are never
|
||||
/// classified as scripts.
|
||||
/// `min_anchor_size` additionally constrains what counts as an
|
||||
/// attachment target: the small-font pass accepts any sufficiently
|
||||
/// larger item (0.0), while the body-font pass requires a heading-sized
|
||||
/// anchor so a body-size table cell beside a slightly larger label with
|
||||
/// baseline jitter is never treated as a script.
|
||||
fn is_script_attachment(&self, small: &TextItem, min_anchor_size: f32) -> bool {
|
||||
let attach_gap = small.font_size.max(4.0) * 0.6;
|
||||
let lo = self
|
||||
.by_y
|
||||
.partition_point(|(y, _)| *y < small.y - self.max_window);
|
||||
self.by_y[lo..]
|
||||
.iter()
|
||||
.take_while(|(y, _)| *y <= small.y + self.max_window)
|
||||
.any(|(_, body)| {
|
||||
let dy = (small.y - body.y).abs();
|
||||
body.font_size >= small.font_size * 1.2
|
||||
&& body.font_size >= min_anchor_size
|
||||
&& dy > body.font_size * 0.05
|
||||
&& dy <= body.font_size * 0.8
|
||||
&& {
|
||||
let gap_after_body = small.x - (body.x + body.width);
|
||||
let gap_before_body = body.x - (small.x + small.width);
|
||||
(-attach_gap..=attach_gap).contains(&gap_after_body)
|
||||
|| (-attach_gap..=attach_gap).contains(&gap_before_body)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// Detect tables in a set of text items from a single page
|
||||
pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bool) -> Vec<Table> {
|
||||
if items.len() < 6 {
|
||||
@@ -221,12 +156,10 @@ pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bo
|
||||
// === Pass 1: Small-font tables (existing behavior) ===
|
||||
let table_font_threshold = base_font_size * 0.90;
|
||||
|
||||
let script_index = ScriptBodyIndex::new(items);
|
||||
let table_candidates: Vec<(usize, &TextItem)> = items
|
||||
.iter()
|
||||
.enumerate()
|
||||
.filter(|(_, item)| item.font_size <= table_font_threshold && item.font_size >= 6.0)
|
||||
.filter(|(_, item)| !script_index.is_script_attachment(item, 0.0))
|
||||
.collect();
|
||||
|
||||
if table_candidates.len() >= 6 {
|
||||
@@ -243,20 +176,6 @@ pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bo
|
||||
continue;
|
||||
}
|
||||
|
||||
for (_, it) in ®ion_items {
|
||||
// Coordinates and metrics only — document text must not
|
||||
// leak into logs.
|
||||
log::debug!(
|
||||
" region item: {} chars x={:.1} y={:.1} w={:.1} fs={:.1} font={}",
|
||||
it.text.chars().count(),
|
||||
it.x,
|
||||
it.y,
|
||||
it.width,
|
||||
it.font_size,
|
||||
it.font
|
||||
);
|
||||
}
|
||||
|
||||
if let Some(mut table) =
|
||||
detect_table_in_region(®ion_items, TableDetectionMode::SmallFont)
|
||||
{
|
||||
@@ -284,11 +203,6 @@ pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bo
|
||||
let body_font_low = base_font_size * 0.85;
|
||||
let body_font_high = base_font_size * 1.05;
|
||||
|
||||
// The script exclusion applies here too — a sub/superscript attached
|
||||
// to a heading-size run can land in the body-font band and would
|
||||
// otherwise re-enter table detection through this pass — but only
|
||||
// for heading-sized anchors (>= 1.15x base): body-size table cells
|
||||
// beside slightly larger labels must never be filtered.
|
||||
let body_candidates: Vec<(usize, &TextItem)> = items
|
||||
.iter()
|
||||
.enumerate()
|
||||
@@ -298,7 +212,6 @@ pub fn detect_tables(items: &[TextItem], base_font_size: f32, skip_body_font: bo
|
||||
&& item.font_size <= body_font_high
|
||||
&& item.font_size >= 6.0
|
||||
})
|
||||
.filter(|(_, item)| !script_index.is_script_attachment(item, base_font_size * 1.15))
|
||||
.collect();
|
||||
|
||||
log::debug!(
|
||||
@@ -667,30 +580,6 @@ fn detect_table_in_region(items: &[(usize, &TextItem)], mode: TableDetectionMode
|
||||
cells.push(row_cells);
|
||||
}
|
||||
|
||||
// Validation 0 (small-font pass only): reject tiny all-numeric
|
||||
// fragments. A ≤2-row grid whose every cell is a bare 1-2 digit number
|
||||
// carries no tabular information — in practice these are
|
||||
// exponent/subscript clusters from display math (e.g. the W^-2, W^-4
|
||||
// determinant superscripts in TeX papers) that happen to align. Emitting
|
||||
// them as plain text lines is strictly better. Body-font tables are not
|
||||
// subject to this veto: their cells cannot be script glyphs.
|
||||
if matches!(mode, TableDetectionMode::SmallFont) {
|
||||
let nonempty_cells: Vec<&String> =
|
||||
cells.iter().flatten().filter(|c| !c.is_empty()).collect();
|
||||
if rows.len() <= 2
|
||||
&& !nonempty_cells.is_empty()
|
||||
&& nonempty_cells
|
||||
.iter()
|
||||
.all(|c| c.len() <= 2 && c.chars().all(|ch| ch.is_ascii_digit()))
|
||||
{
|
||||
log::debug!(
|
||||
" validation 0 fail: tiny all-numeric fragment ({} cells)",
|
||||
nonempty_cells.len()
|
||||
);
|
||||
return None;
|
||||
}
|
||||
}
|
||||
|
||||
// Validation 1: some rows should have content in first column.
|
||||
// Use a lower threshold (25%) for tables with wrapped cells where
|
||||
// continuation lines leave the first column empty.
|
||||
@@ -1757,126 +1646,6 @@ fn try_add_label_column(
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::types::ItemType;
|
||||
|
||||
fn make_item(text: &str, x: f32, y: f32, font_size: f32, width: f32) -> TextItem {
|
||||
TextItem {
|
||||
text: text.to_string(),
|
||||
x,
|
||||
y,
|
||||
width,
|
||||
height: font_size,
|
||||
font: "TestFont".to_string(),
|
||||
font_size,
|
||||
page: 1,
|
||||
is_bold: false,
|
||||
is_italic: false,
|
||||
is_underline: false,
|
||||
is_strikeout: false,
|
||||
item_type: ItemType::Text,
|
||||
mcid: None,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_detects_subscript_after_body_text() {
|
||||
// "log" at 10pt followed immediately by subscript "10" at 7pt,
|
||||
// slightly below the baseline (TeX display math pattern).
|
||||
let body = make_item("log", 100.0, 500.0, 10.0, 15.0);
|
||||
let sub = make_item("10", 115.5, 497.0, 7.0, 7.0);
|
||||
let items = vec![body, sub.clone()];
|
||||
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sub, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_detects_superscript_footnote_marker() {
|
||||
// "Hartley" at 10pt with superscript "2" raised above the baseline.
|
||||
let body = make_item("Hartley", 200.0, 500.0, 10.0, 35.0);
|
||||
let sup = make_item("2", 235.8, 504.0, 6.6, 3.5);
|
||||
let items = vec![body, sup.clone()];
|
||||
assert!(ScriptBodyIndex::new(&items).is_script_attachment(&sup, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_ignores_small_cell_far_from_body_text() {
|
||||
// A small-font table cell 40pt away from body text in another column
|
||||
// must NOT be treated as a script attachment.
|
||||
let body = make_item("Revenue", 100.0, 500.0, 10.0, 40.0);
|
||||
let cell = make_item("1,234", 180.0, 500.0, 7.0, 20.0);
|
||||
let items = vec![body, cell.clone()];
|
||||
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_ignores_same_baseline_neighbor_cell() {
|
||||
// A small cell immediately beside a larger label cell on the SAME
|
||||
// baseline is a table layout, not a subscript — a genuine baseline
|
||||
// offset is required.
|
||||
let label = make_item("Total", 100.0, 500.0, 10.0, 25.0);
|
||||
let cell = make_item("42", 127.0, 500.0, 7.5, 9.0);
|
||||
let items = vec![label, cell.clone()];
|
||||
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn script_attachment_ignores_neighbor_on_different_line() {
|
||||
// Small item directly below a body item (next row) is not a script.
|
||||
let body = make_item("Header", 100.0, 500.0, 10.0, 30.0);
|
||||
let cell = make_item("42", 131.0, 486.0, 7.0, 10.0);
|
||||
let items = vec![body, cell.clone()];
|
||||
assert!(!ScriptBodyIndex::new(&items).is_script_attachment(&cell, 0.0));
|
||||
}
|
||||
|
||||
/// The equation-subscript + footnote layout from Shannon entropy.pdf
|
||||
/// page 1, with real coordinates. Without body anchors ("log" items) the
|
||||
/// small-font items alone must form a phantom 3-column table — proving
|
||||
/// this layout exercises the detection path — and adding the anchors
|
||||
/// must suppress it via the script-attachment filter.
|
||||
fn shannon_page1_small_items() -> Vec<TextItem> {
|
||||
vec![
|
||||
// Equation subscripts (7.4pt): log2 M = log10 M / log10 2
|
||||
make_item("2", 267.4, 133.9, 7.4, 3.7),
|
||||
make_item("10", 306.2, 133.9, 7.4, 7.4),
|
||||
make_item("10", 342.7, 133.9, 7.4, 7.4),
|
||||
make_item("10", 325.0, 118.9, 7.4, 7.4),
|
||||
// Footnote block (8pt)
|
||||
make_item("Bell System Technical Journal,", 295.7, 101.9, 8.0, 95.0),
|
||||
make_item(
|
||||
"April 1924, p. 324; Certain Topics in",
|
||||
396.7,
|
||||
101.9,
|
||||
8.0,
|
||||
130.0,
|
||||
),
|
||||
make_item("v. 47, April 1928, p. 617.", 250.9, 92.5, 8.0, 90.0),
|
||||
make_item("Bell System Technical Journal,", 264.2, 82.6, 8.0, 95.0),
|
||||
make_item("July 1928, p. 535.", 364.3, 82.6, 8.0, 65.0),
|
||||
]
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn equation_scripts_do_not_form_phantom_table() {
|
||||
// Sanity: without the larger-font anchors the same items DO form a
|
||||
// phantom table, so this layout genuinely reaches region detection.
|
||||
let bare = shannon_page1_small_items();
|
||||
assert!(
|
||||
!detect_tables(&bare, 10.0, false).is_empty(),
|
||||
"test layout must form a phantom table when the filter cannot fire"
|
||||
);
|
||||
|
||||
// With the "log" anchors adjacent to each subscript, the script
|
||||
// filter removes the subscripts and no table survives.
|
||||
let mut items = shannon_page1_small_items();
|
||||
items.push(make_item("log", 253.0, 137.0, 10.0, 13.5));
|
||||
items.push(make_item("log", 291.5, 137.0, 10.0, 13.5));
|
||||
items.push(make_item("log", 328.0, 137.0, 10.0, 13.5));
|
||||
items.push(make_item("log", 310.3, 122.0, 10.0, 13.5));
|
||||
let tables = detect_tables(&items, 10.0, false);
|
||||
assert!(
|
||||
tables.is_empty(),
|
||||
"equation scripts + footnotes must not become a table: {tables:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn is_table_of_contents_rejects_toc() {
|
||||
|
||||
BIN
Binary file not shown.
+22
-12
@@ -162,6 +162,28 @@ fn test_detection_config_custom() {
|
||||
assert!((config.text_page_ratio_threshold - 0.8).abs() < 0.001);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_detection_pages_rejects_no_in_range_pages() {
|
||||
let buffer = std::fs::read("tests/fixtures/thermo-freon12.pdf").unwrap();
|
||||
let error = pdf_inspector::detect_pdf_type_mem_with_config(
|
||||
&buffer,
|
||||
DetectionConfig {
|
||||
strategy: ScanStrategy::Pages(vec![9999]),
|
||||
..DetectionConfig::default()
|
||||
},
|
||||
)
|
||||
.expect_err("out-of-range page selection must fail");
|
||||
|
||||
assert!(
|
||||
matches!(
|
||||
error,
|
||||
PdfError::InvalidOptions(ref message)
|
||||
if message.contains("contains no in-range page numbers")
|
||||
),
|
||||
"unexpected error: {error:?}"
|
||||
);
|
||||
}
|
||||
|
||||
// ============================================================================
|
||||
// PdfType Tests
|
||||
// ============================================================================
|
||||
@@ -1089,18 +1111,6 @@ fn test_snapshot_2013_app2() {
|
||||
assert_snapshot("2013-app2");
|
||||
}
|
||||
|
||||
/// First two pages of Shannon's "A Mathematical Theory of Communication"
|
||||
/// (1998 dvips 5.58 → Distiller 3 retypesetting). Canonical legacy-TeX PDF:
|
||||
/// non-embedded base-14 fonts with no /Widths (exercises the built-in AFM
|
||||
/// metrics fallback), Type3 PK bitmap math fonts with FontMatrix
|
||||
/// [1 0 0 -1 0 0] (exercises visual-size scaling), a two-line embedded drop
|
||||
/// cap, indent-only paragraph breaks, and display math that must not be
|
||||
/// detected as tables or headings.
|
||||
#[test]
|
||||
fn test_snapshot_shannon_entropy() {
|
||||
assert_snapshot("shannon-entropy-p1-2");
|
||||
}
|
||||
|
||||
// ============================================================================
|
||||
// Pages Needing OCR Tests
|
||||
// ============================================================================
|
||||
|
||||
@@ -6,7 +6,4 @@
|
||||
|Exclusive Special|83,500,000 75,909,091(7,590,909) with 3.5% individual consumption tax applied 82,232,000|79,275,000 with 3.5% individual consumption tax applied 79,275,000|▶ Standard equipment of Exclusive plus • Smart Safety Technology: Forward Collision-avoidance Assist(intersection crossing/changing lanes in oncoming traffic/approaching from either side/ evasive steering assist), Highway Driving Assist 2, Navigation-based Smart Cruise Control(access road) • Exterior: Roof rack • Interior: Metallic pedal, Driving mode-dependent ambient mood lighting(crash pad, 1st/2nd-row door trim) • Seat: Synthetic leather seats(patch applied), Power-adjustable driver's seat(8-way, lumbar support, and Integrated Memory System(driver's seat and outside mirror connected)), Power-adjustable front passenger’s seat(8-way), Ventilated 1st-row seats, Heated 2nd-row seats • Convenience: Hi-pass(e hi-pass), In-car fingerprint authentication system(personalization, startup, payment, and etc.), Smart power tailgate|▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [950,000] Parking Assist ▶ [1,150,000] Audio by BANG & OLUFSEN sound system ▶ [250,000] 19-inch alloy wheels & tires|
|
||||
|Prestige|87,893,000 79,902,727(7,990,273) with 3.5% individual consumption tax applied 86,559,000|83,445,000 with 3.5% individual consumption tax applied 83,445,000|▶ Standard equipment of Exclusive Special plus • Smart Safety Technology: Remote Smart Parking Assist 2, Parking Collison- avoidance Assist(front/side/rear) • Exterior: Intelligent Front-Lighting System(IFS), Dynamic welcome/escort lighting(1 type), Sequential turn signals(front and rear), Ambient lighting auto flush door handles, Two-tone door garnish, Glossy black rear diffuser • Interior: Recycled PET suede interior materials(headlining/sunvisor), Fabric upholstered crash pad • Seat: BIO-processed natural leather seats(metal patch applied, embossed design punching), Passenger's seat walk-in device, 1st-row relaxation comfort seats(leg rest included), Ventilated 2nd-row seats • Convenience: Parking Distance Warning-Side, Head-Up Display, Digital key 2, Wireless phone charger(dual), Surround View Monitor, Blind-spot View Monitor, LED reverse light guide • Infotainment: Audio by BANG & OLUFSEN sound system(14 speakers, including external amp), Active road noise control, Active Sound Design|▶ [600,000] Built-in Cam 2 Plus, Augmented reality navigation ▶ [850,000] Indoor/outdoor V2L ▶ [900,000] Vision roof ▶ [1,380,000] Digital side mirror ▶ [750,000] Camera package ▶ [250,000] 19-inch alloy wheels & tires|
|
||||
|
||||
**Classification Details** **Indoor/outdoor V2L** Indoor V2L, Outdoor V2L(connectorless type) **Parking Assist** Surround View Monitor, Blind-spot View Monitor, Parking Distance Warning-Side, Parking Collison-avoidance Assist-Rear **Audio by BANG & OLUFSEN**
|
||||
|
||||
Audio by BANG & OLUFSEN sound system(14 speakers, including external amp.), Active road noise control, Active Sound Design **sound system** **Camera package** Digital center mirror(with camera sensor cleaning system), Driver monitoring system THE ALL-NEW NEXO /// ECO-FRIENDLY CAR
|
||||
|
||||
**Classification Details** **Indoor/outdoor V2L** Indoor V2L, Outdoor V2L(connectorless type) **Parking Assist** Surround View Monitor, Blind-spot View Monitor, Parking Distance Warning-Side, Parking Collison-avoidance Assist-Rear **Audio by BANG & OLUFSEN** Audio by BANG & OLUFSEN sound system(14 speakers, including external amp.), Active road noise control, Active Sound Design **sound system** **Camera package** Digital center mirror(with camera sensor cleaning system), Driver monitoring system THE ALL-NEW NEXO /// ECO-FRIENDLY CAR
|
||||
|
||||
@@ -8,7 +8,7 @@ Department of the Treasury **Internal Revenue Service**
|
||||
|
||||
# and Report to Employer
|
||||
|
||||
## This publication contains:
|
||||
### This publication contains:
|
||||
|
||||
**Form 4070A,** Employee’s Daily Record of Tips **Form 4070,** Employee’s Report of Tips to Employer
|
||||
|
||||
@@ -22,7 +22,7 @@ Name and address of employee
|
||||
|
||||
**Publication 1244 (Rev. 7-96)** Cat. No. 44472W
|
||||
|
||||
## Instructions
|
||||
### Instructions
|
||||
|
||||
You must keep sufficient proof to show the amount of your tip income for the year. A daily record of your tip income is considered sufficient proof. Keep a daily record for each workday showing the amount of cash and credit card tips received directly from customers or other employees. Also keep a record of the amount of tips, if any, you paid to other employees through tip sharing, tip pooling or other arrangements, and the names of employees to whom you paid tips. Show the date that each entry is made. This date should be on or near the date you received the tip income. You may use **Form 4070A**, Employee’s Daily Record of Tips, or any other daily record to record your tips. **Reporting Tips to Your Employer.—**If you receive tips that total $20 or more for any month while working for one employer, you must report the tips to your employer. Tips include cash left by customers, tips customers add to credit card charges, and tips you receive from other employees. You must report your tips for any one month by the 10th day of the next month. If the 10th day falls on a Saturday, Sunday, or legal holiday, you may give the report to your employer on the next business day that is not a Saturday, Sunday, or legal holiday. You must report tips that total $20 or more every month regardless of your total wages and tips for the year. You may use **Form 4070**, Employee’s Report of Tips to Employer, to report your tips to your employer. See the instructions on the back of Form 4070. You must include all tips, including tips not reported to your employer, as wages on your income tax return. You may use the last page of this publication to total your tips for the year. Your employer must withhold income, social security, and Medicare (or railroad retirement) taxes on tips you report. Your employer usually deducts the withholding due on tips from your regular wages.
|
||||
|
||||
@@ -56,7 +56,11 @@ tips of directly from customers received other employees paid tips rec’d. entr
|
||||
**2.** Report total credit card tips (col. **b**) on Form 4070, line **2.**
|
||||
**3.** Report total tips paid out (col. **c**) on Form 4070, line **3.** **Page 4**
|
||||
|
||||
Form Employee’s Report (Rev. July 1996) of Tips to EmployerOMB No. 1545-0065 Department of the Treasury Internal Revenue Service' **For Paperwork Reduction Act Notice, see back of form.** Employee’s name and address **Social security number**
|
||||
Form Employee’s Report (Rev. July 1996)
|
||||
|
||||
## of Tips to EmployerOMB No. 1545-0065
|
||||
|
||||
Department of the Treasury Internal Revenue Service' **For Paperwork Reduction Act Notice, see back of form.** Employee’s name and address **Social security number**
|
||||
|
||||
Employer’s name and address (include establishment name, if different) **1** Cash tips received
|
||||
|
||||
@@ -74,7 +78,7 @@ forms simpler, we would be happy to hear from you. You can write to the Tax Form
|
||||
|
||||
**Unreported Tips.—**If you received tips of $20 or more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you **must** use Form 1040 and **Form 4137,** Social Security and Medicare Tax on Unreported Tip Income, to report them. You may **not** use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act **cannot** use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—**Get **Pub. 531,** Reporting Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—**If you do not keep a daily record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
|
||||
|
||||
## Instructions (continued)
|
||||
### Instructions (continued)
|
||||
|
||||
Use this space to total your tips for the year
|
||||
|
||||
|
||||
@@ -46,19 +46,9 @@ more about investing in tax losses than burst, cap rates spreads steadily com- r
|
||||
|
||||
8 6 Z E L L / L U R I E R E A L E S T A T E C E N T E R
|
||||
|
||||
ingdebtspreadswerepartoforiginalpro formamodels.Thiscapratespreadcom- pressionoffsetweakcashflowsinapost- recessionary economy from 2002 to 2005, while continued compression, combined with improved cash flows, pushed property values skyward in 2006 throughmid-2007.
|
||||
ingdebtspreadswerepartoforiginalpro formamodels.Thiscapratespreadcom- pressionoffsetweakcashflowsinapost- recessionary economy from 2002 to 2005, while continued compression, combined with improved cash flows, pushed property values skyward in 2006 throughmid-2007. Cap rate compression reduced the importance of the ability to add value. After all, if all you had to do to make moneywastoleveragetothehiltwhilecap ratesfell,whytakeontheextraworkand riskofattemptingtoaddvalue?Stateddif- ferently: Why print money if it is laying everywhereonthestreets? In Tables III and IV, we demonstrate thepowerofcapratecompressionviavery simple pro forma cash flow analyses that assume Year 1 NOI of $100; a going-in cap rate of 9 percent; an LTV of 70 per- cent; and an interest rate of 7 percent. Withineachfigure,wedisplaytwoscenar- ios, which vary based on NOI growth assumptions.ScenarioIassumesthatNOI growsby3percentperyear,whileScenario IIassumesavalue-addNOIgrowthof20 percentbetweenyearstwoandthree. The only other difference between TablesIIIandIVisinresidualcaprates, which are assumed to be 6 percent and 9 percent, respectively. Based on these assumptions, we calculate the equity IRRs. It is clear that cap rate compres- sion is a significant factor in driving
|
||||
|
||||
Cap rate compression reduced the importance of the ability to add value. After all, if all you had to do to make moneywastoleveragetothehiltwhilecap ratesfell,whytakeontheextraworkand riskofattemptingtoaddvalue?Stateddif- ferently: Why print money if it is laying everywhereonthestreets?
|
||||
|
||||
In Tables III and IV, we demonstrate thepowerofcapratecompressionviavery simple pro forma cash flow analyses that assume Year 1 NOI of $100; a going-in cap rate of 9 percent; an LTV of 70 per- cent; and an interest rate of 7 percent. Withineachfigure,wedisplaytwoscenar- ios, which vary based on NOI growth assumptions.ScenarioIassumesthatNOI growsby3percentperyear,whileScenario IIassumesavalue-addNOIgrowthof20 percentbetweenyearstwoandthree.
|
||||
|
||||
The only other difference between TablesIIIandIVisinresidualcaprates, which are assumed to be 6 percent and 9 percent, respectively. Based on these assumptions, we calculate the equity IRRs. It is clear that cap rate compres- sion is a significant factor in driving
|
||||
|
||||
returns. That is, cap rate compression from 9 percent to 6 percent increased IRR on leveraged stabilized properties by 250 percent, to a staggering 57 per- cent. Who needs to take on value add riskatthisreturnforstabilizedassets?
|
||||
|
||||
Intheearly1980s,moneywasmadein real estate by mastering the creation and syndication of tax gimmicks. In the late 1980s, one made money by mastering bank and S&L connections to over-lever- age.Intheearly1990s,onemademoneyin realestatebyhavingaccesstoequity—the morethebetter.Duringthelate1990s,one made money from real estate by realizing large spreads between cap rates and debt costs.And,overthepastfiveyears,theway to make money in real estate was to own realestateonahighlyleveragedbasisascap ratesplunged.
|
||||
|
||||
Theclassicassetpricingmodelisthe capital asset pricing model (CAPM). CAPM is a simple, yet elegant, model that relates asset pricing to the risk-free rate(F),theabilityofanassettoreduce portfolio variance (B), and the expected rate of return on the market bundle of investableassets(M).CAPMisfarfrom perfect,butprovidesacrudebenchmark for asset pricing, around which discrep- ancies and novelties arise. Specifically, CAPM states that an asset’s price is set suchthattheexpectedreturnforanasset
|
||||
returns. That is, cap rate compression from 9 percent to 6 percent increased IRR on leveraged stabilized properties by 250 percent, to a staggering 57 per- cent. Who needs to take on value add riskatthisreturnforstabilizedassets? Intheearly1980s,moneywasmadein real estate by mastering the creation and syndication of tax gimmicks. In the late 1980s, one made money by mastering bank and S&L connections to over-lever- age.Intheearly1990s,onemademoneyin realestatebyhavingaccesstoequity—the morethebetter.Duringthelate1990s,one made money from real estate by realizing large spreads between cap rates and debt costs.And,overthepastfiveyears,theway to make money in real estate was to own realestateonahighlyleveragedbasisascap ratesplunged. Theclassicassetpricingmodelisthe capital asset pricing model (CAPM). CAPM is a simple, yet elegant, model that relates asset pricing to the risk-free rate(F),theabilityofanassettoreduce portfolio variance (B), and the expected rate of return on the market bundle of investableassets(M).CAPMisfarfrom perfect,butprovidesacrudebenchmark for asset pricing, around which discrep- ancies and novelties arise. Specifically, CAPM states that an asset’s price is set suchthattheexpectedreturnforanasset
|
||||
|
||||
(R)is R=F+ β(M-F).
|
||||
R E V I E W 8 7
|
||||
|
||||
@@ -1,47 +0,0 @@
|
||||
Reprinted with corrections from *The Bell System Technical Journal,* Vol. 27, pp. 379–423, 623–656, July, October, 1948.
|
||||
|
||||
# A Mathematical Theory of Communication
|
||||
|
||||
## By C. E. SHANNON
|
||||
|
||||
INTRODUCTION
|
||||
|
||||
THE recent development of various methods of modulation such as PCM and PPM which exchange bandwidth for signal-to-noise ratio has intensified the interest in a general theory of communication. A basis for such a theory is contained in the important papers of Nyquist¹ and Hartley² on this subject. In the present paper we will extend the theory to include a number of new factors, in particular the effect of noise in the channel, and the savings possible due to the statistical structure of the original message and due to the nature of the final destination of the information.
|
||||
|
||||
The fundamental problem of communication is that of reproducing at one point either exactly or ap- proximately a message selected at another point. Frequently the messages have *meaning*; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one *selected from a set* of possible messages. The system must be designed to operate for each possible selection, not just the one which will actually be chosen since this is unknown at the time of design.
|
||||
|
||||
If the number of messages in the set is finite then this number or any monotonic function of this number can be regarded as a measure of the information produced when one message is chosen from the set, all choices being equally likely. As was pointed out by Hartley the most natural choice is the logarithmic function. Although this definition must be generalized considerably when we consider the influence of the statistics of the message and when we have a continuous range of messages, we will in all cases use an essentially logarithmic measure.
|
||||
|
||||
The logarithmic measure is more convenient for various reasons:
|
||||
|
||||
1. It is practically more useful. Parameters of engineering importance such as time, bandwidth, number of relays, etc., tend to vary linearly with the logarithm of the number of possibilities. For example, adding one relay to a group doubles the number of possible states of the relays. It adds 1 to the base 2 logarithm of this number. Doubling the time roughly squares the number of possible messages, or doubles the logarithm, etc.
|
||||
2. It is nearer to our intuitive feeling as to the proper measure. This is closely related to (1) since we in- tuitively measures entities by linear comparison with common standards. One feels, for example, that two punched cards should have twice the capacity of one for information storage, and two identical channels twice the capacity of one for transmitting information.
|
||||
3. It is mathematically more suitable. Many of the limiting operations are simple in terms of the loga- rithm but would require clumsy restatement in terms of the number of possibilities. The choice of a logarithmic base corresponds to the choice of a unit for measuring information. If the
|
||||
base 2 is used the resulting units may be called binary digits, or more briefly *bits,* a word suggested by
|
||||
|
||||
J. W. Tukey. A device with two stable positions, such as a relay or a flip-flop circuit, can store one bit of information. *N* such devices can store*N* bits, since the total number of possible states is 2
|
||||
*N* and log₂2 *N* = *N*. If the base 10 is used the units may be called decimal digits. Since
|
||||
|
||||
log₂*M* = log₁₀*M*= log₁₀2 = 3:32 log₁₀*M*;
|
||||
|
||||
1 Nyquist, H., “Certain Factors Affecting Telegraph Speed,” *Bell System Technical Journal,* April 1924, p. 324; “Certain Topics in Telegraph Transmission Theory,” *A.I.E.E. Trans.,* v. 47, April 1928, p. 617. Hartley, R. V. L., “Transmission of Information,” *Bell System Technical Journal,* July 1928, p. 535.
|
||||
|
||||
INFORMATION SOURCE TRANSMITTER RECEIVER DESTINATION
|
||||
|
||||
SIGNAL RECEIVED SIGNAL MESSAGE MESSAGE
|
||||
|
||||
NOISE SOURCE
|
||||
|
||||
Fig. 1 — Schematic diagram of a general communication system.
|
||||
|
||||
a decimal digit is about 3 13 bits. A digit wheel on a desk computing machine has ten stable positions and therefore has a storage capacity of one decimal digit. In analytical work where integration and differentiation are involved the base *e* is sometimes useful. The resulting units of information will be called natural units. Change from the base *a* to base *b* merely requires multiplication by log*ba*.
|
||||
|
||||
By a communication system we will mean a system of the type indicated schematically in Fig. 1. It consists of essentially five parts:
|
||||
|
||||
1. An *information source* which produces a message or sequence of messages to be communicated to the receiving terminal. The message may be of various types: (a) A sequence of letters as in a telegraph of teletype system; (b) A single function of time *f* (*t*) as in radio or telephony; (c) A function of time and other variables as in black and white television — here the message may be thought of as a function *f* (*x*; *y*;*t*) of two space coordinates and time, the light intensity at point (*x*; *y*) and time *t* on a pickup tube plate; (d) Two or more functions of time, say *f* (*t*), *g*(*t*), *h*(*t*) — this is the case in “three- dimensional” sound transmission or if the system is intended to service several individual channels in multiplex; (e) Several functions of several variables — in color television the message consists of three functions *f* (*x*; *y*;*t*), *g*(*x*; *y*;*t*), *h*(*x*; *y*;*t*) defined in a three-dimensional continuum — we may also think of these three functions as components of a vector field defined in the region — similarly, several black and white television sources would produce “messages” consisting of a number of functions of three variables; (f) Various combinations also occur, for example in television with an associated audio channel.
|
||||
2. A *transmitter* which operates on the message in some way to produce a signal suitable for trans- mission over the channel. In telephony this operation consists merely of changing sound pressure into a proportional electrical current. In telegraphy we have an encoding operation which produces a sequence of dots, dashes and spaces on the channel corresponding to the message. In a multiplex PCM system the different speech functions must be sampled, compressed, quantized and encoded, and finally interleaved properly to construct the signal. Vocoder systems, television and frequency modulation are other examples of complex operations applied to the message to obtain the signal.
|
||||
3. The *channel* is merely the medium used to transmit the signal from transmitter to receiver. It may be a pair of wires, a coaxial cable, a band of radio frequencies, a beam of light, etc.
|
||||
4. The *receiver* ordinarily performs the inverse operation of that done by the transmitter, reconstructing the message from the signal.
|
||||
5. The *destination* is the person (or thing) for whom the message is intended. We wish to consider certain general problems involving communication systems. To do this it is first
|
||||
necessary to represent the various elements involved as mathematical entities, suitably idealized from their
|
||||
|
||||
@@ -1,20 +1,18 @@
|
||||
### Technical Information
|
||||
##### Technical Information
|
||||
|
||||
#### l T-12 SI
|
||||
## l T-12 SI
|
||||
|
||||
#### DuPont Fluorochemicals
|
||||
##### DuPont Fluorochemicals
|
||||
|
||||
## Thermodynamic Properties
|
||||
#### Thermodynamic Properties
|
||||
|
||||
**of**
|
||||
|
||||
®
|
||||
**of** ®
|
||||
|
||||
# Freon 12
|
||||
|
||||
### (R-12)
|
||||
##### (R-12)
|
||||
|
||||
#### Technical Information Technical Information
|
||||
##### Technical Information Technical Information
|
||||
|
||||
**®** **Thermodynamic Properties of Freon 12 Refrigerant** **(R-12)** **SI Units**
|
||||
|
||||
@@ -24,11 +22,11 @@ S.A., Lemmon, E.W., and Peskin, Vf = Fluid (liquid) specific volume
|
||||
A.P., NIST Standard Reference in cubic meters per kilogram Database 23, NIST thermodynamic and transport properties of Vg = Vapour (gas) specific volume refrigerants and refrigerant in cubic meters per kilogram mixtures – REFPROP version 6.01, Standard Reference Data Program, df and dg = Fluid and Vapour National Institute of Standards and (respectively) densities in Technology, 1998).
|
||||
kilograms per cubic meter
|
||||
|
||||
H = Enthalpy (kJ/kg)
|
||||
##### H = Enthalpy (kJ/kg)
|
||||
|
||||
S = Entropy (kJ/kg.K)
|
||||
##### S = Entropy (kJ/kg.K)
|
||||
|
||||
#### Physical Properties
|
||||
##### Physical Properties
|
||||
|
||||
|Chemical Formula|CCl₂F₂|
|
||||
|---|---|
|
||||
|
||||
Generated
+1
-1
@@ -724,7 +724,7 @@ checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
|
||||
|
||||
[[package]]
|
||||
name = "pdf-inspector"
|
||||
version = "0.1.7"
|
||||
version = "0.1.6"
|
||||
dependencies = [
|
||||
"env_logger",
|
||||
"include_dir",
|
||||
|
||||
+16
-2
@@ -11,7 +11,7 @@ npm install @firecrawl/pdf-inspector-wasm
|
||||
## Usage
|
||||
|
||||
```ts
|
||||
import init, { processPdf } from "@firecrawl/pdf-inspector-wasm";
|
||||
import init, { detectPdf, processPdf } from "@firecrawl/pdf-inspector-wasm";
|
||||
|
||||
await init();
|
||||
|
||||
@@ -33,10 +33,24 @@ const result = processPdf(pdf, {
|
||||
});
|
||||
```
|
||||
|
||||
Detection can scan every page, stop early, sample a fixed number of pages, or
|
||||
inspect a caller-selected set of 1-indexed pages:
|
||||
|
||||
```ts
|
||||
const full = detectPdf(pdf, { strategy: "full" });
|
||||
const earlyExit = detectPdf(pdf, { strategy: "earlyExit" });
|
||||
const sampled = detectPdf(pdf, { strategy: { sample: 12 } });
|
||||
const selected = detectPdf(pdf, { strategy: { pages: [1, 50, 100] } });
|
||||
```
|
||||
|
||||
The same `strategy`, `minTextOpsPerPage`, and `textPageRatioThreshold` options
|
||||
are accepted by `processPdf` and `classifyPdf`. Unknown or unsupported option
|
||||
fields throw an error instead of being ignored.
|
||||
|
||||
The package also exports:
|
||||
|
||||
- `detectPdf(pdf, options?)` for detection without extraction.
|
||||
- `classifyPdf(pdf)` for the lightweight result shape shared with the native Node.js API.
|
||||
- `classifyPdf(pdf, options?)` for the lightweight result shape shared with the native Node.js API.
|
||||
- `extractText(pdf)` for plain text.
|
||||
- `version()` for the WASM package version.
|
||||
|
||||
|
||||
+491
-22
@@ -1,6 +1,6 @@
|
||||
use pdf_inspector::{
|
||||
LayoutComplexity, MarkdownProfile, PageOcrReasons, PdfOptions, PdfProcessResult, PdfType,
|
||||
ProcessMode,
|
||||
DetectionConfig, LayoutComplexity, MarkdownProfile, PageOcrReasons, PdfOptions,
|
||||
PdfProcessResult, PdfType, ProcessMode, ScanStrategy,
|
||||
};
|
||||
use serde::{Deserialize, Serialize};
|
||||
use wasm_bindgen::prelude::*;
|
||||
@@ -10,11 +10,29 @@ const TYPESCRIPT_TYPES: &str = r#"
|
||||
export type PdfType = "TextBased" | "Scanned" | "ImageBased" | "Mixed";
|
||||
export type MarkdownProfile = "fidelity" | "compact";
|
||||
|
||||
export interface ProcessOptions {
|
||||
/** Restrict extraction to these 1-indexed page numbers. */
|
||||
pages?: number[];
|
||||
export type ScanStrategy =
|
||||
| "earlyExit"
|
||||
| "full"
|
||||
| { sample: number }
|
||||
| { pages: number[] };
|
||||
|
||||
export interface DetectionOptions {
|
||||
/** Which pages detection inspects. Defaults to `{ sample: 8 }`. */
|
||||
strategy?: ScanStrategy;
|
||||
/** Minimum text operators required for a page to count as text-based. */
|
||||
minTextOpsPerPage?: number;
|
||||
/** Text-page ratio required for TextBased classification (0.0–1.0). */
|
||||
textPageRatioThreshold?: number;
|
||||
}
|
||||
|
||||
export interface DetectOptions extends DetectionOptions {
|
||||
/** Password for an encrypted PDF. */
|
||||
password?: string;
|
||||
}
|
||||
|
||||
export interface ProcessOptions extends DetectOptions {
|
||||
/** Restrict extraction to these 1-indexed page numbers. */
|
||||
pages?: number[];
|
||||
/** Source-faithful output by default, or compact output for fewer tokens. */
|
||||
profile?: MarkdownProfile;
|
||||
/** Insert `<!-- Page N -->` markers between pages. */
|
||||
@@ -60,8 +78,8 @@ export interface PdfClassification {
|
||||
}
|
||||
|
||||
export function processPdf(data: Uint8Array, options?: ProcessOptions): PdfProcessResult;
|
||||
export function detectPdf(data: Uint8Array, options?: Pick<ProcessOptions, "password">): PdfProcessResult;
|
||||
export function classifyPdf(data: Uint8Array): PdfClassification;
|
||||
export function detectPdf(data: Uint8Array, options?: DetectOptions): PdfProcessResult;
|
||||
export function classifyPdf(data: Uint8Array, options?: DetectOptions): PdfClassification;
|
||||
export function extractText(data: Uint8Array): string;
|
||||
export function version(): string;
|
||||
"#;
|
||||
@@ -69,6 +87,9 @@ export function version(): string;
|
||||
#[derive(Debug, Default, Deserialize)]
|
||||
#[serde(default, rename_all = "camelCase", deny_unknown_fields)]
|
||||
struct WasmProcessOptions {
|
||||
strategy: Option<WasmScanStrategy>,
|
||||
min_text_ops_per_page: Option<u32>,
|
||||
text_page_ratio_threshold: Option<f32>,
|
||||
pages: Option<Vec<u32>>,
|
||||
password: Option<String>,
|
||||
profile: Option<WasmMarkdownProfile>,
|
||||
@@ -76,6 +97,31 @@ struct WasmProcessOptions {
|
||||
include_images: Option<bool>,
|
||||
}
|
||||
|
||||
#[derive(Debug, Deserialize)]
|
||||
#[serde(untagged)]
|
||||
enum WasmScanStrategy {
|
||||
Named(WasmNamedScanStrategy),
|
||||
Sample(WasmSampleStrategy),
|
||||
Pages(WasmPagesStrategy),
|
||||
}
|
||||
|
||||
#[derive(Debug, Deserialize)]
|
||||
#[serde(rename_all = "camelCase")]
|
||||
enum WasmNamedScanStrategy {
|
||||
EarlyExit,
|
||||
Full,
|
||||
}
|
||||
|
||||
#[derive(Debug, Deserialize)]
|
||||
struct WasmSampleStrategy {
|
||||
sample: u32,
|
||||
}
|
||||
|
||||
#[derive(Debug, Deserialize)]
|
||||
struct WasmPagesStrategy {
|
||||
pages: Vec<u32>,
|
||||
}
|
||||
|
||||
#[derive(Debug, Deserialize)]
|
||||
#[serde(rename_all = "lowercase")]
|
||||
enum WasmMarkdownProfile {
|
||||
@@ -175,16 +221,120 @@ fn js_error(context: &str, error: impl std::fmt::Display) -> JsValue {
|
||||
js_sys::Error::new(&format!("{context}: {error}")).into()
|
||||
}
|
||||
|
||||
fn deserialize_options(value: JsValue) -> Result<WasmProcessOptions, JsValue> {
|
||||
fn validate_object_fields(
|
||||
value: &JsValue,
|
||||
allowed_fields: &[&str],
|
||||
object_name: &str,
|
||||
) -> Result<(), JsValue> {
|
||||
// serde-wasm-bindgen reads the fields named by the Rust struct but does
|
||||
// not enumerate other JavaScript object keys, so Serde's
|
||||
// deny_unknown_fields cannot catch them by itself.
|
||||
value.dyn_ref::<js_sys::Object>().ok_or_else(|| {
|
||||
js_error(
|
||||
"invalid options",
|
||||
format!("{object_name} must be an object"),
|
||||
)
|
||||
})?;
|
||||
|
||||
let keys = js_sys::Reflect::own_keys(value).map_err(|_| {
|
||||
js_error(
|
||||
"invalid options",
|
||||
format!("could not inspect {object_name} fields"),
|
||||
)
|
||||
})?;
|
||||
for key in keys.iter() {
|
||||
let Some(key) = key.as_string() else {
|
||||
return Err(js_error(
|
||||
"invalid options",
|
||||
format!("{object_name} fields must use string keys"),
|
||||
));
|
||||
};
|
||||
if !allowed_fields.contains(&key.as_str()) {
|
||||
return Err(js_error(
|
||||
"invalid options",
|
||||
format!("unknown {object_name} field `{key}`"),
|
||||
));
|
||||
}
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn validate_strategy_fields(value: &JsValue) -> Result<(), JsValue> {
|
||||
if value.is_undefined() || value.is_null() || value.as_string().is_some() {
|
||||
return Ok(());
|
||||
}
|
||||
|
||||
value.dyn_ref::<js_sys::Object>().ok_or_else(|| {
|
||||
js_error(
|
||||
"invalid options",
|
||||
"strategy must be \"earlyExit\", \"full\", { sample: number }, or { pages: number[] }",
|
||||
)
|
||||
})?;
|
||||
let keys = js_sys::Reflect::own_keys(value)
|
||||
.map_err(|_| js_error("invalid options", "could not inspect strategy fields"))?;
|
||||
if keys.length() != 1 {
|
||||
return Err(js_error(
|
||||
"invalid options",
|
||||
"strategy objects must contain exactly one of `sample` or `pages`",
|
||||
));
|
||||
}
|
||||
|
||||
let key = keys
|
||||
.get(0)
|
||||
.as_string()
|
||||
.ok_or_else(|| js_error("invalid options", "strategy fields must use string keys"))?;
|
||||
if key != "sample" && key != "pages" {
|
||||
return Err(js_error(
|
||||
"invalid options",
|
||||
format!("unknown strategy field `{key}`"),
|
||||
));
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn validate_option_fields(value: &JsValue, mode: &ProcessMode) -> Result<(), JsValue> {
|
||||
const DETECTION_FIELDS: &[&str] = &[
|
||||
"strategy",
|
||||
"minTextOpsPerPage",
|
||||
"textPageRatioThreshold",
|
||||
"password",
|
||||
];
|
||||
const PROCESS_FIELDS: &[&str] = &[
|
||||
"strategy",
|
||||
"minTextOpsPerPage",
|
||||
"textPageRatioThreshold",
|
||||
"password",
|
||||
"pages",
|
||||
"profile",
|
||||
"includePageMarkers",
|
||||
"includeImages",
|
||||
];
|
||||
|
||||
let allowed_fields = match mode {
|
||||
ProcessMode::Full => PROCESS_FIELDS,
|
||||
ProcessMode::DetectOnly => DETECTION_FIELDS,
|
||||
ProcessMode::Analyze => PROCESS_FIELDS,
|
||||
};
|
||||
validate_object_fields(value, allowed_fields, "option")?;
|
||||
|
||||
let strategy = js_sys::Reflect::get(value, &JsValue::from_str("strategy"))
|
||||
.map_err(|_| js_error("invalid options", "could not read `strategy`"))?;
|
||||
validate_strategy_fields(&strategy)
|
||||
}
|
||||
|
||||
fn deserialize_options(value: JsValue, mode: &ProcessMode) -> Result<WasmProcessOptions, JsValue> {
|
||||
if value.is_undefined() || value.is_null() {
|
||||
return Ok(WasmProcessOptions::default());
|
||||
}
|
||||
|
||||
validate_option_fields(&value, mode)?;
|
||||
serde_wasm_bindgen::from_value(value).map_err(|error| js_error("invalid options", error))
|
||||
}
|
||||
|
||||
fn build_options(value: JsValue, mode: ProcessMode) -> Result<PdfOptions, JsValue> {
|
||||
let options = deserialize_options(value)?;
|
||||
let options = deserialize_options(value, &mode)?;
|
||||
if options
|
||||
.pages
|
||||
.as_ref()
|
||||
@@ -196,7 +346,60 @@ fn build_options(value: JsValue, mode: ProcessMode) -> Result<PdfOptions, JsValu
|
||||
));
|
||||
}
|
||||
|
||||
let mut detection = DetectionConfig::default();
|
||||
if let Some(strategy) = options.strategy {
|
||||
detection.strategy = match strategy {
|
||||
WasmScanStrategy::Named(WasmNamedScanStrategy::EarlyExit) => ScanStrategy::EarlyExit,
|
||||
WasmScanStrategy::Named(WasmNamedScanStrategy::Full) => ScanStrategy::Full,
|
||||
WasmScanStrategy::Sample(WasmSampleStrategy { sample }) => {
|
||||
if sample == 0 {
|
||||
return Err(js_error(
|
||||
"invalid options",
|
||||
"strategy.sample must be at least 1",
|
||||
));
|
||||
}
|
||||
ScanStrategy::Sample(sample)
|
||||
}
|
||||
WasmScanStrategy::Pages(WasmPagesStrategy { pages }) => {
|
||||
if pages.is_empty() {
|
||||
return Err(js_error(
|
||||
"invalid options",
|
||||
"strategy.pages must not be empty",
|
||||
));
|
||||
}
|
||||
if pages.contains(&0) {
|
||||
return Err(js_error(
|
||||
"invalid options",
|
||||
"strategy.pages are 1-indexed; page 0 is invalid",
|
||||
));
|
||||
}
|
||||
ScanStrategy::Pages(pages)
|
||||
}
|
||||
};
|
||||
}
|
||||
if let Some(min_text_ops_per_page) = options.min_text_ops_per_page {
|
||||
if min_text_ops_per_page == 0 {
|
||||
return Err(js_error(
|
||||
"invalid options",
|
||||
"minTextOpsPerPage must be at least 1",
|
||||
));
|
||||
}
|
||||
detection.min_text_ops_per_page = min_text_ops_per_page;
|
||||
}
|
||||
if let Some(text_page_ratio_threshold) = options.text_page_ratio_threshold {
|
||||
if !text_page_ratio_threshold.is_finite()
|
||||
|| !(0.0..=1.0).contains(&text_page_ratio_threshold)
|
||||
{
|
||||
return Err(js_error(
|
||||
"invalid options",
|
||||
"textPageRatioThreshold must be between 0.0 and 1.0",
|
||||
));
|
||||
}
|
||||
detection.text_page_ratio_threshold = text_page_ratio_threshold;
|
||||
}
|
||||
|
||||
let mut result = PdfOptions::new().mode(mode);
|
||||
result = result.detection(detection);
|
||||
if let Some(pages) = options.pages {
|
||||
result = result.pages(pages);
|
||||
}
|
||||
@@ -252,14 +455,19 @@ pub fn detect_pdf(data: &[u8], options: JsValue) -> Result<JsValue, JsValue> {
|
||||
|
||||
/// Return the lightweight classification shape used by the native Node API.
|
||||
#[wasm_bindgen(js_name = classifyPdf, skip_typescript)]
|
||||
pub fn classify_pdf(data: &[u8]) -> Result<JsValue, JsValue> {
|
||||
pub fn classify_pdf(data: &[u8], options: JsValue) -> Result<JsValue, JsValue> {
|
||||
initialize();
|
||||
let result =
|
||||
pdf_inspector::classify_pdf_mem(data).map_err(|error| js_error("classify PDF", error))?;
|
||||
let options = build_options(options, ProcessMode::DetectOnly)?;
|
||||
let result = pdf_inspector::process_pdf_mem_with_options(data, options)
|
||||
.map_err(|error| js_error("classify PDF", error))?;
|
||||
serialize(&WasmPdfClassification {
|
||||
pdf_type: pdf_type_name(result.pdf_type),
|
||||
page_count: result.page_count,
|
||||
pages_needing_ocr: result.pages_needing_ocr,
|
||||
pages_needing_ocr: result
|
||||
.pages_needing_ocr
|
||||
.into_iter()
|
||||
.map(|page| page - 1)
|
||||
.collect(),
|
||||
confidence: result.confidence as f64,
|
||||
})
|
||||
}
|
||||
@@ -295,6 +503,33 @@ mod tests {
|
||||
const TEXT_PDF: &[u8] = include_bytes!("../../tests/fixtures/thermo-freon12.pdf");
|
||||
const ENCRYPTED_PDF: &[u8] = include_bytes!("../../tests/fixtures/encrypted-secret123.pdf");
|
||||
|
||||
fn string_property(value: &JsValue, name: &str) -> String {
|
||||
Reflect::get(value, &JsValue::from_str(name))
|
||||
.unwrap_or_else(|_| panic!("read {name}"))
|
||||
.as_string()
|
||||
.unwrap_or_else(|| panic!("{name} string"))
|
||||
}
|
||||
|
||||
fn error_message(error: &JsValue) -> String {
|
||||
Reflect::get(error, &JsValue::from_str("message"))
|
||||
.expect("read error message")
|
||||
.as_string()
|
||||
.expect("error message string")
|
||||
}
|
||||
|
||||
fn define_non_enumerable_property(object: &js_sys::Object, name: &str, value: &JsValue) {
|
||||
let descriptor = js_sys::Object::new();
|
||||
Reflect::set(&descriptor, &JsValue::from_str("value"), value)
|
||||
.expect("set descriptor value");
|
||||
Reflect::set(
|
||||
&descriptor,
|
||||
&JsValue::from_str("enumerable"),
|
||||
&JsValue::FALSE,
|
||||
)
|
||||
.expect("set descriptor enumerable");
|
||||
js_sys::Object::define_property(object, &JsValue::from_str(name), &descriptor);
|
||||
}
|
||||
|
||||
fn synthetic_korea1_pdf() -> Vec<u8> {
|
||||
let mut pdf = b"%PDF-1.4\n".to_vec();
|
||||
let mut offsets = vec![0usize];
|
||||
@@ -377,13 +612,120 @@ mod tests {
|
||||
pdf
|
||||
}
|
||||
|
||||
/// A 20-page document where the eight pages selected by the default
|
||||
/// Sample(8) strategy are text, while every other page is image-backed.
|
||||
fn synthetic_heterogeneous_pdf() -> Vec<u8> {
|
||||
const PAGE_COUNT: usize = 20;
|
||||
const TEXT_CONTENT_ID: usize = 23;
|
||||
const IMAGE_CONTENT_ID: usize = 24;
|
||||
const FONT_ID: usize = 25;
|
||||
const IMAGE_ID: usize = 26;
|
||||
|
||||
let mut pdf = b"%PDF-1.4\n".to_vec();
|
||||
let mut offsets = vec![0usize];
|
||||
|
||||
fn add_object(pdf: &mut Vec<u8>, offsets: &mut Vec<usize>, id: usize, body: &str) {
|
||||
offsets.push(pdf.len());
|
||||
pdf.extend_from_slice(format!("{id} 0 obj\n").as_bytes());
|
||||
pdf.extend_from_slice(body.as_bytes());
|
||||
pdf.extend_from_slice(b"\nendobj\n");
|
||||
}
|
||||
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
1,
|
||||
"<< /Type /Catalog /Pages 2 0 R >>",
|
||||
);
|
||||
let kids = (3..3 + PAGE_COUNT)
|
||||
.map(|id| format!("{id} 0 R"))
|
||||
.collect::<Vec<_>>()
|
||||
.join(" ");
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
2,
|
||||
&format!("<< /Type /Pages /Kids [{kids}] /Count {PAGE_COUNT} >>"),
|
||||
);
|
||||
|
||||
// distribute_pages(8, 20) selects exactly these page numbers.
|
||||
let default_sample = [1usize, 3, 5, 7, 9, 11, 13, 20];
|
||||
for page in 1..=PAGE_COUNT {
|
||||
let page_id = page + 2;
|
||||
let body = if default_sample.contains(&page) {
|
||||
format!(
|
||||
"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] \
|
||||
/Resources << /Font << /F1 {FONT_ID} 0 R >> >> \
|
||||
/Contents {TEXT_CONTENT_ID} 0 R >>"
|
||||
)
|
||||
} else {
|
||||
format!(
|
||||
"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 200 200] \
|
||||
/Resources << /XObject << /Im0 {IMAGE_ID} 0 R >> >> \
|
||||
/Contents {IMAGE_CONTENT_ID} 0 R >>"
|
||||
)
|
||||
};
|
||||
add_object(&mut pdf, &mut offsets, page_id, &body);
|
||||
}
|
||||
|
||||
let text_content =
|
||||
"BT /F1 12 Tf 10 100 Td (Hello World) Tj (More Text) Tj (Sample Page) Tj ET";
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
TEXT_CONTENT_ID,
|
||||
&format!(
|
||||
"<< /Length {} >>\nstream\n{}\nendstream",
|
||||
text_content.len(),
|
||||
text_content
|
||||
),
|
||||
);
|
||||
let image_content = "/Im0 Do";
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
IMAGE_CONTENT_ID,
|
||||
&format!(
|
||||
"<< /Length {} >>\nstream\n{}\nendstream",
|
||||
image_content.len(),
|
||||
image_content
|
||||
),
|
||||
);
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
FONT_ID,
|
||||
"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica >>",
|
||||
);
|
||||
add_object(
|
||||
&mut pdf,
|
||||
&mut offsets,
|
||||
IMAGE_ID,
|
||||
"<< /Type /XObject /Subtype /Image /Width 1000 /Height 1000 \
|
||||
/ColorSpace /DeviceGray /BitsPerComponent 8 /Length 1 >>\nstream\n0\nendstream",
|
||||
);
|
||||
|
||||
let xref_start = pdf.len();
|
||||
pdf.extend_from_slice(format!("xref\n0 {}\n", offsets.len()).as_bytes());
|
||||
pdf.extend_from_slice(b"0000000000 65535 f \n");
|
||||
for offset in offsets.iter().skip(1) {
|
||||
pdf.extend_from_slice(format!("{offset:010} 00000 n \n").as_bytes());
|
||||
}
|
||||
pdf.extend_from_slice(
|
||||
format!(
|
||||
"trailer\n<< /Size {} /Root 1 0 R >>\nstartxref\n{}\n%%EOF",
|
||||
offsets.len(),
|
||||
xref_start
|
||||
)
|
||||
.as_bytes(),
|
||||
);
|
||||
pdf
|
||||
}
|
||||
|
||||
#[wasm_bindgen_test]
|
||||
fn processes_pdf_to_markdown() {
|
||||
let result = process_pdf(TEXT_PDF, JsValue::UNDEFINED).expect("process PDF");
|
||||
let pdf_type = Reflect::get(&result, &JsValue::from_str("pdfType"))
|
||||
.expect("pdfType")
|
||||
.as_string()
|
||||
.expect("pdfType string");
|
||||
let pdf_type = string_property(&result, "pdfType");
|
||||
let markdown = Reflect::get(&result, &JsValue::from_str("markdown"))
|
||||
.expect("markdown")
|
||||
.as_string()
|
||||
@@ -400,11 +742,8 @@ mod tests {
|
||||
|
||||
#[wasm_bindgen_test]
|
||||
fn classifies_and_extracts_plain_text() {
|
||||
let classification = classify_pdf(TEXT_PDF).expect("classify PDF");
|
||||
let pdf_type = Reflect::get(&classification, &JsValue::from_str("pdfType"))
|
||||
.expect("pdfType")
|
||||
.as_string()
|
||||
.expect("pdfType string");
|
||||
let classification = classify_pdf(TEXT_PDF, JsValue::UNDEFINED).expect("classify PDF");
|
||||
let pdf_type = string_property(&classification, "pdfType");
|
||||
let text = extract_text(TEXT_PDF).expect("extract text");
|
||||
|
||||
assert_eq!(pdf_type, "TextBased");
|
||||
@@ -437,4 +776,134 @@ mod tests {
|
||||
|
||||
assert!(!markdown.is_empty());
|
||||
}
|
||||
|
||||
#[wasm_bindgen_test]
|
||||
fn full_strategy_scans_pages_missed_by_default_sample() {
|
||||
let pdf = synthetic_heterogeneous_pdf();
|
||||
let sampled = detect_pdf(&pdf, JsValue::UNDEFINED).expect("sampled detection");
|
||||
assert_eq!(string_property(&sampled, "pdfType"), "TextBased");
|
||||
|
||||
let options = js_sys::Object::new();
|
||||
Reflect::set(
|
||||
&options,
|
||||
&JsValue::from_str("strategy"),
|
||||
&JsValue::from_str("full"),
|
||||
)
|
||||
.expect("set full strategy");
|
||||
|
||||
let full = detect_pdf(&pdf, options.clone().into()).expect("full detection");
|
||||
assert_eq!(string_property(&full, "pdfType"), "Mixed");
|
||||
|
||||
let classification = classify_pdf(&pdf, options.into()).expect("full classification");
|
||||
assert_eq!(string_property(&classification, "pdfType"), "Mixed");
|
||||
let pages = js_sys::Array::from(
|
||||
&Reflect::get(&classification, &JsValue::from_str("pagesNeedingOcr"))
|
||||
.expect("pagesNeedingOcr"),
|
||||
);
|
||||
assert!(
|
||||
pages.includes(&JsValue::from_f64(1.0), 0),
|
||||
"classifyPdf should report image-backed page 2 as zero-indexed page 1"
|
||||
);
|
||||
}
|
||||
|
||||
#[wasm_bindgen_test]
|
||||
fn accepts_every_scan_strategy_variant() {
|
||||
let sample = js_sys::Object::new();
|
||||
Reflect::set(
|
||||
&sample,
|
||||
&JsValue::from_str("sample"),
|
||||
&JsValue::from_f64(1.0),
|
||||
)
|
||||
.expect("set sample count");
|
||||
|
||||
let pages = js_sys::Object::new();
|
||||
Reflect::set(
|
||||
&pages,
|
||||
&JsValue::from_str("pages"),
|
||||
&js_sys::Array::of1(&JsValue::from_f64(1.0)),
|
||||
)
|
||||
.expect("set strategy pages");
|
||||
|
||||
for strategy in [
|
||||
JsValue::from_str("earlyExit"),
|
||||
JsValue::from_str("full"),
|
||||
sample.into(),
|
||||
pages.into(),
|
||||
] {
|
||||
let options = js_sys::Object::new();
|
||||
Reflect::set(&options, &JsValue::from_str("strategy"), &strategy)
|
||||
.expect("set strategy");
|
||||
detect_pdf(TEXT_PDF, options.into()).expect("supported strategy");
|
||||
}
|
||||
}
|
||||
|
||||
#[wasm_bindgen_test]
|
||||
fn rejects_unknown_and_unsupported_detection_options() {
|
||||
let unknown = js_sys::Object::new();
|
||||
Reflect::set(&unknown, &JsValue::from_str("fullScan"), &JsValue::TRUE)
|
||||
.expect("set unknown option");
|
||||
let error = detect_pdf(TEXT_PDF, unknown.into()).expect_err("unknown option must fail");
|
||||
assert!(error_message(&error).contains("unknown option field `fullScan`"));
|
||||
|
||||
let hidden = js_sys::Object::new();
|
||||
define_non_enumerable_property(&hidden, "hiddenOption", &JsValue::TRUE);
|
||||
let error =
|
||||
detect_pdf(TEXT_PDF, hidden.into()).expect_err("non-enumerable option must fail");
|
||||
assert!(error_message(&error).contains("unknown option field `hiddenOption`"));
|
||||
|
||||
let symbol_keyed = js_sys::Object::new();
|
||||
Reflect::set(
|
||||
&symbol_keyed,
|
||||
&js_sys::Symbol::for_("unsupportedOption").into(),
|
||||
&JsValue::TRUE,
|
||||
)
|
||||
.expect("set symbol option");
|
||||
let error = detect_pdf(TEXT_PDF, symbol_keyed.into()).expect_err("symbol option must fail");
|
||||
assert!(error_message(&error).contains("option fields must use string keys"));
|
||||
|
||||
let process_only = js_sys::Object::new();
|
||||
Reflect::set(
|
||||
&process_only,
|
||||
&JsValue::from_str("profile"),
|
||||
&JsValue::from_str("compact"),
|
||||
)
|
||||
.expect("set process-only option");
|
||||
let error = detect_pdf(TEXT_PDF, process_only.into())
|
||||
.expect_err("unsupported detection option must fail");
|
||||
assert!(error_message(&error).contains("unknown option field `profile`"));
|
||||
|
||||
let malformed_strategy = js_sys::Object::new();
|
||||
Reflect::set(
|
||||
&malformed_strategy,
|
||||
&JsValue::from_str("sample"),
|
||||
&JsValue::from_f64(8.0),
|
||||
)
|
||||
.expect("set sample");
|
||||
define_non_enumerable_property(&malformed_strategy, "hiddenTypo", &JsValue::TRUE);
|
||||
let options = js_sys::Object::new();
|
||||
Reflect::set(
|
||||
&options,
|
||||
&JsValue::from_str("strategy"),
|
||||
&malformed_strategy,
|
||||
)
|
||||
.expect("set malformed strategy");
|
||||
let error = detect_pdf(TEXT_PDF, options.into()).expect_err("malformed strategy must fail");
|
||||
assert!(error_message(&error).contains("exactly one of `sample` or `pages`"));
|
||||
}
|
||||
|
||||
#[wasm_bindgen_test]
|
||||
fn rejects_strategy_pages_when_none_are_in_range() {
|
||||
let pages = js_sys::Object::new();
|
||||
Reflect::set(
|
||||
&pages,
|
||||
&JsValue::from_str("pages"),
|
||||
&js_sys::Array::of1(&JsValue::from_f64(9999.0)),
|
||||
)
|
||||
.expect("set out-of-range pages");
|
||||
let options = js_sys::Object::new();
|
||||
Reflect::set(&options, &JsValue::from_str("strategy"), &pages).expect("set pages strategy");
|
||||
|
||||
let error = detect_pdf(TEXT_PDF, options.into()).expect_err("out-of-range pages must fail");
|
||||
assert!(error_message(&error).contains("contains no in-range page numbers"));
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user