Compare commits
13
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7ab187fea7 | ||
|
|
a91db44851 | ||
|
|
62b3269301 | ||
|
|
eac8af0df8 | ||
|
|
39c31a8404 | ||
|
|
cb3906e7b8 | ||
|
|
ebfd096b78 | ||
|
|
d8eb33e390 | ||
|
|
a38efcf142 | ||
|
|
42f0e52987 | ||
|
|
0ee3cd7d10 | ||
|
|
f5be40143c | ||
|
|
fed3b90d37 |
@@ -0,0 +1,36 @@
|
||||
name: Deploy landing page
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
paths: ['site/**', '.github/workflows/pages.yml']
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
pages: write
|
||||
id-token: write
|
||||
|
||||
# Allow one concurrent deployment; don't cancel an in-progress production deploy.
|
||||
concurrency:
|
||||
group: pages
|
||||
cancel-in-progress: false
|
||||
|
||||
jobs:
|
||||
deploy:
|
||||
name: Build & deploy to GitHub Pages
|
||||
runs-on: ubuntu-latest
|
||||
environment:
|
||||
name: github-pages
|
||||
url: ${{ steps.deploy.outputs.page_url }}
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- name: Upload site artifact
|
||||
uses: actions/upload-pages-artifact@v3
|
||||
with:
|
||||
path: site
|
||||
|
||||
- name: Deploy to GitHub Pages
|
||||
id: deploy
|
||||
uses: actions/deploy-pages@v4
|
||||
@@ -26,16 +26,16 @@ Evaluated on the [opendataloader-bench](https://github.com/opendataloader-projec
|
||||
|
||||
| Engine | Overall | Reading Order (NID) | Tables (TEDS) | Headings (MHS) | Speed (200 docs) |
|
||||
|---|---|---|---|---|---|
|
||||
| pdf-inspector | 0.78 | 0.87 | 0.59 | 0.57 | 4s |
|
||||
| pdf-inspector | 0.83 | 0.88 | 0.66 | 0.74 | 4s |
|
||||
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
|
||||
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
|
||||
| markitdown | 0.58 | 0.88 | 0.00 | 0.00 | 8s |
|
||||
|
||||
For context, engines that use OCR/ML (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus.
|
||||
For context, engines that use OCR/ML (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus — pdf-inspector reaches the low end of that range without any OCR, in 4 seconds.
|
||||
|
||||
**Where we do well:** Speed (fastest of all engines), reading order, table detection vs other direct-text tools.
|
||||
**Where we do well:** Speed (fastest of all engines), the best table detection of any engine shown, and heading detection now on par with opendataloader. Overall lands within 0.01 of opendataloader at roughly 2.5× the speed.
|
||||
|
||||
**Where we lag:** Heading detection trails opendataloader — many PDFs use bold text at body font size for headings, or headings that are only slightly larger than body text. Table detection trails OCR-based engines that can see visual table structure.
|
||||
**Where we lag:** Reading order still trails opendataloader slightly, and table structure trails OCR-based engines that can see visual layout.
|
||||
|
||||
## Quick start
|
||||
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "@firecrawl/pdf-inspector",
|
||||
"version": "1.10.0",
|
||||
"version": "1.10.3",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
|
||||
@@ -0,0 +1,3 @@
|
||||
<svg width="200" height="284" viewBox="0 0 200 284" fill="none" xmlns="http://www.w3.org/2000/svg">
|
||||
<path d="M166.862 90.7716C155.812 94.0514 147.483 101.471 141.383 109.53C140.073 111.26 137.343 109.96 137.863 107.841C149.543 59.8136 134.113 19.896 86.0157 0.247269C83.5758 -0.752669 81.0359 1.43719 81.6759 3.99704C103.555 91.8416 11.5294 84.432 23.1588 184.016C23.3588 185.726 21.4389 186.896 20.039 185.896C15.6792 182.766 10.8095 176.236 7.46963 171.647C6.48968 170.297 4.36978 170.677 3.9198 172.287C1.25994 181.906 0 190.965 0 199.965C0 234.963 17.9891 265.771 45.2177 283.63C46.7777 284.65 48.7776 283.19 48.2476 281.4C46.8477 276.7 46.0577 271.74 45.9977 266.611C45.9977 263.461 46.1977 260.241 46.6877 257.241C47.8276 249.702 50.4475 242.522 54.8473 235.983C69.9365 213.334 100.185 191.455 95.3552 161.747C95.0453 159.867 97.2651 158.627 98.6651 159.917C119.974 179.386 124.194 205.575 120.694 229.063C120.394 231.103 122.954 232.193 124.244 230.593C127.504 226.513 131.483 222.933 135.813 220.244C136.893 219.574 138.333 220.084 138.743 221.284C141.153 228.293 144.733 234.873 148.113 241.452C152.152 249.362 154.302 258.391 153.962 267.951C153.792 272.6 153.022 277.1 151.732 281.38C151.182 283.19 153.162 284.7 154.752 283.66C182.001 265.801 200 234.993 200 199.975C200 187.806 197.87 175.876 193.84 164.697C185.391 141.248 163.952 123.64 169.372 93.0815C169.632 91.6216 168.282 90.3517 166.862 90.7716Z" fill="#FA5D19" style="fill:#FA5D19;fill:color(display-p3 0.9816 0.3634 0.0984);fill-opacity:1;"/>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 1.5 KiB |
@@ -0,0 +1,12 @@
|
||||
<svg width="172" height="40" viewBox="0 0 172 40" fill="none" xmlns="http://www.w3.org/2000/svg">
|
||||
<path d="M23.3606 12.8281C21.8137 13.2873 20.6476 14.3261 19.7936 15.4544C19.6102 15.6966 19.228 15.5146 19.3008 15.2178C20.936 8.49401 18.7759 2.90556 12.0422 0.154735C11.7006 0.0147436 11.345 0.321324 11.4346 0.679702C14.4977 12.9779 1.61412 11.9406 3.24224 25.8823C3.27024 26.1217 3.00145 26.2855 2.80546 26.1455C2.19509 25.7073 1.51332 24.7932 1.04575 24.1506C0.908555 23.9616 0.611769 24.0148 0.548773 24.2402C0.176391 25.5869 0 26.8553 0 28.1152C0 33.0149 2.51847 37.328 6.33048 39.8283C6.54887 39.9711 6.82886 39.7667 6.75466 39.5161C6.55867 38.8581 6.44808 38.1638 6.43968 37.4456C6.43968 37.0046 6.46768 36.5539 6.53627 36.1339C6.69587 35.0784 7.06265 34.0732 7.67862 33.1577C9.79111 29.9869 14.0259 26.9239 13.3497 22.7647C13.3063 22.5015 13.6171 22.328 13.8131 22.5085C16.7964 25.2342 17.3871 28.9005 16.8972 32.1889C16.8552 32.4745 17.2135 32.6271 17.3941 32.4031C17.8505 31.832 18.4077 31.3308 19.0138 30.9542C19.165 30.8604 19.3666 30.9318 19.424 31.0998C19.7614 32.0811 20.2626 33.0023 20.7358 33.9234C21.3013 35.0308 21.6023 36.2949 21.5547 37.6332C21.5309 38.2842 21.4231 38.9141 21.2425 39.5133C21.1655 39.7667 21.4427 39.9781 21.6653 39.8325C25.4801 37.3322 28 33.0191 28 28.1166C28 26.4129 27.7018 24.7428 27.1376 23.1777C25.9547 19.8949 22.9533 17.4297 23.712 13.1515C23.7484 12.9471 23.5594 12.7693 23.3606 12.8281Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
<path d="M41 34.0521V10.9618H55.7586V14.3264H44.7969V21.0226H53.8436V24.2882H44.7969V34.0521H41Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
<path d="M59.9569 14.7882C58.7352 14.7882 57.7777 13.8976 57.7777 12.6441C57.7777 11.3906 58.7352 10.5 59.9569 10.5C61.1785 10.5 62.136 11.3906 62.136 12.6441C62.136 13.8976 61.1785 14.7882 59.9569 14.7882ZM58.1409 34.0521V17.1632H61.7068V34.0521H58.1409Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
<path d="M73.5885 17.1632H74.3809V20.4948H72.796C69.6264 20.4948 68.6029 22.9687 68.6029 25.5747V34.0521H65.0371V17.1632H68.2067L68.6029 19.7031C69.4613 18.2847 70.815 17.1632 73.5885 17.1632Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
<path d="M83.632 34.25C78.3163 34.25 74.9816 30.8194 74.9816 25.6406C74.9816 20.4288 78.3163 16.9653 83.3019 16.9653C88.1884 16.9653 91.457 20.066 91.5561 25.0139C91.5561 25.4427 91.5231 25.9045 91.457 26.3663H78.7125V26.5972C78.8116 29.467 80.6275 31.3472 83.4339 31.3472C85.613 31.3472 87.1979 30.2587 87.6931 28.3785H91.2589C90.6646 31.7101 87.8252 34.25 83.632 34.25ZM78.8446 23.7604H87.8582C87.561 21.2535 85.8112 19.8351 83.3349 19.8351C81.0567 19.8351 79.1087 21.3524 78.8446 23.7604Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
<path d="M102.033 34.25C96.9151 34.25 93.6465 30.9184 93.6465 25.6406C93.6465 20.4288 97.0142 16.9653 102.132 16.9653C106.49 16.9653 109.197 19.3733 109.891 23.1997H106.16C105.698 21.2205 104.278 20 102.066 20C99.1933 20 97.3113 22.309 97.3113 25.6406C97.3113 28.9392 99.1933 31.2153 102.066 31.2153C104.245 31.2153 105.698 29.9618 106.127 28.0156H109.891C109.23 31.842 106.358 34.25 102.033 34.25Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
<path d="M121.006 17.1632H121.799V20.4948H120.214C117.044 20.4948 116.021 22.9687 116.021 25.5747V34.0521H112.455V17.1632H115.625L116.021 19.7031C116.879 18.2847 118.233 17.1632 121.006 17.1632Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
<path d="M130.614 16.9653C135.104 16.9653 137.679 19.1094 137.679 23.1007V34.0521H134.576L134.279 31.6441C133.123 33.1615 131.505 34.25 128.831 34.25C125.133 34.25 122.657 32.4358 122.657 29.3021C122.657 25.8385 125.166 23.8924 129.92 23.8924H134.147V22.8698C134.147 20.9896 132.793 19.8351 130.449 19.8351C128.336 19.8351 126.916 20.8247 126.652 22.309H123.152C123.515 19.0104 126.355 16.9653 130.614 16.9653ZM129.425 31.4792C132.397 31.4792 134.114 29.7309 134.147 27.125V26.5312H129.722C127.51 26.5312 126.289 27.3559 126.289 29.0712C126.289 30.4896 127.477 31.4792 129.425 31.4792Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
<path d="M144.653 34.0521L139.139 17.1632H142.903L146.766 30.0937L150.629 17.1632H153.897L157.595 30.0937L161.59 17.1632H165.222L159.609 34.0521H155.779L152.214 22.5729L148.516 34.0521H144.653Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
<path d="M166.934 34.0521V10.9618H170.5V34.0521H166.934Z" fill="#262626" style="fill:#262626;fill:color(display-p3 0.1500 0.1500 0.1500);fill-opacity:1;"/>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 4.8 KiB |
+428
@@ -0,0 +1,428 @@
|
||||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||
<title>pdf-inspector — PDF classification & text extraction, no OCR</title>
|
||||
<meta name="description" content="Fast Rust library that classifies PDFs (text-based vs scanned) and extracts clean Markdown — no OCR, no ML models. Bindings for Rust, Python, and Node.js.">
|
||||
<meta property="og:title" content="pdf-inspector">
|
||||
<meta property="og:description" content="Classify PDFs and extract clean Markdown in milliseconds. No OCR. No ML. Pure Rust.">
|
||||
<meta property="og:type" content="website">
|
||||
<link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 100 100'%3E%3Ctext y='.9em' font-size='90'%3E%F0%9F%93%84%3C/text%3E%3C/svg%3E">
|
||||
<link rel="preconnect" href="https://fonts.googleapis.com">
|
||||
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
|
||||
<link href="https://fonts.googleapis.com/css2?family=Bricolage+Grotesque:opsz,wght@12..96,400;12..96,600;12..96,800&family=Hanken+Grotesk:wght@400;500;600&family=JetBrains+Mono:wght@400;500;700&display=swap" rel="stylesheet">
|
||||
<style>
|
||||
:root {
|
||||
--paper: #f4efe4;
|
||||
--paper-2: #eee7d8;
|
||||
--ink: #1b1712;
|
||||
--ink-soft: #4a433a;
|
||||
--muted: #8b8375;
|
||||
--line: #d9cfbb;
|
||||
--accent: #dd3f22;
|
||||
--accent-deep: #b32d15;
|
||||
--card: #faf6ec;
|
||||
--display: "Bricolage Grotesque", serif;
|
||||
--body: "Hanken Grotesk", sans-serif;
|
||||
--mono: "JetBrains Mono", monospace;
|
||||
}
|
||||
* { box-sizing: border-box; margin: 0; padding: 0; }
|
||||
html { scroll-behavior: smooth; }
|
||||
body {
|
||||
background: var(--paper);
|
||||
color: var(--ink);
|
||||
font-family: var(--body);
|
||||
font-size: 17px;
|
||||
line-height: 1.6;
|
||||
-webkit-font-smoothing: antialiased;
|
||||
overflow-x: hidden;
|
||||
background-image:
|
||||
radial-gradient(circle at 1px 1px, rgba(27,23,18,0.05) 1px, transparent 0);
|
||||
background-size: 22px 22px;
|
||||
}
|
||||
::selection { background: var(--accent); color: var(--paper); }
|
||||
a { color: inherit; text-decoration: none; }
|
||||
|
||||
.wrap { max-width: 1120px; margin: 0 auto; padding: 0 28px; }
|
||||
|
||||
/* ── nav ── */
|
||||
nav {
|
||||
position: sticky; top: 0; z-index: 50;
|
||||
background: rgba(244,239,228,0.82);
|
||||
backdrop-filter: blur(10px);
|
||||
border-bottom: 1px solid var(--line);
|
||||
}
|
||||
.nav-in { display: flex; align-items: center; gap: 22px; height: 60px; }
|
||||
.brand { font-family: var(--mono); font-weight: 700; font-size: 15px; letter-spacing: -0.02em; display: flex; align-items: center; gap: 9px; }
|
||||
.brand .dot { width: 9px; height: 9px; background: var(--accent); border-radius: 50%; box-shadow: 0 0 0 3px rgba(221,63,34,0.18); }
|
||||
.nav-links { margin-left: auto; display: flex; gap: 24px; align-items: center; font-size: 14.5px; font-weight: 500; }
|
||||
.nav-links a { color: var(--ink-soft); transition: color .15s; }
|
||||
.nav-links a:hover { color: var(--accent); }
|
||||
.nav-gh { border: 1px solid var(--ink); border-radius: 999px; padding: 6px 15px; color: var(--ink) !important; transition: all .15s; }
|
||||
.nav-gh:hover { background: var(--ink); color: var(--paper) !important; }
|
||||
@media (max-width: 680px) { .nav-links .hide-sm { display: none; } }
|
||||
|
||||
/* ── hero ── */
|
||||
header { padding: 74px 0 40px; position: relative; }
|
||||
.eyebrow { font-family: var(--mono); font-size: 12.5px; letter-spacing: 0.16em; text-transform: uppercase; color: var(--accent-deep); margin-bottom: 22px; }
|
||||
h1 {
|
||||
font-family: var(--display);
|
||||
font-weight: 800;
|
||||
font-size: clamp(2.9rem, 8vw, 6.1rem);
|
||||
line-height: 0.96;
|
||||
letter-spacing: -0.035em;
|
||||
max-width: 15ch;
|
||||
}
|
||||
h1 .em { color: var(--accent); font-style: normal; position: relative; }
|
||||
h1 .strike { position: relative; white-space: nowrap; }
|
||||
h1 .strike::after { content: ""; position: absolute; left: -2%; right: -2%; top: 54%; height: 0.09em; background: var(--accent); transform: rotate(-3deg); }
|
||||
.lede { margin-top: 28px; font-size: clamp(1.05rem, 2.2vw, 1.32rem); color: var(--ink-soft); max-width: 46ch; line-height: 1.5; }
|
||||
.lede b { color: var(--ink); font-weight: 600; }
|
||||
|
||||
.hero-grid { display: grid; grid-template-columns: 1.15fr 0.85fr; gap: 48px; align-items: end; }
|
||||
@media (max-width: 880px) { .hero-grid { grid-template-columns: 1fr; gap: 40px; } }
|
||||
|
||||
/* readout card */
|
||||
.readout {
|
||||
background: var(--ink); color: var(--paper);
|
||||
border-radius: 14px; padding: 22px 22px 20px;
|
||||
font-family: var(--mono); font-size: 13px;
|
||||
box-shadow: 14px 14px 0 rgba(27,23,18,0.09);
|
||||
position: relative;
|
||||
}
|
||||
.readout .rlabel { color: #b8ad98; font-size: 11px; letter-spacing: 0.14em; text-transform: uppercase; margin-bottom: 16px; display: flex; justify-content: space-between; }
|
||||
.readout .rrow { display: flex; justify-content: space-between; align-items: center; padding: 9px 0; border-top: 1px solid rgba(255,255,255,0.09); }
|
||||
.readout .rrow:first-of-type { border-top: none; }
|
||||
.readout .k { color: #cfc6b3; }
|
||||
.readout .v { font-weight: 700; }
|
||||
.readout .v.hot { color: var(--accent); }
|
||||
.bar { height: 6px; background: rgba(255,255,255,0.1); border-radius: 3px; overflow: hidden; margin-top: 3px; width: 96px; }
|
||||
.bar > i { display: block; height: 100%; background: var(--accent); border-radius: 3px; }
|
||||
|
||||
/* ── install row ── */
|
||||
.installs { display: grid; grid-template-columns: repeat(3,1fr); gap: 14px; margin-top: 54px; }
|
||||
@media (max-width: 720px) { .installs { grid-template-columns: 1fr; } }
|
||||
.inst {
|
||||
background: var(--card); border: 1px solid var(--line); border-radius: 11px;
|
||||
padding: 15px 17px; transition: transform .16s, border-color .16s, box-shadow .16s;
|
||||
cursor: pointer;
|
||||
}
|
||||
.inst:hover { transform: translateY(-3px); border-color: var(--accent); box-shadow: 0 8px 22px rgba(27,23,18,0.07); }
|
||||
.inst .reg { font-family: var(--mono); font-size: 11px; letter-spacing: 0.12em; text-transform: uppercase; color: var(--muted); margin-bottom: 8px; display: flex; justify-content: space-between; }
|
||||
.inst code { font-family: var(--mono); font-size: 14px; color: var(--ink); font-weight: 500; }
|
||||
.inst .arrow { color: var(--accent); opacity: 0; transition: opacity .16s; }
|
||||
.inst:hover .arrow { opacity: 1; }
|
||||
|
||||
/* ── section scaffold ── */
|
||||
section { padding: 66px 0; border-top: 1px solid var(--line); }
|
||||
.sec-head { display: flex; align-items: baseline; gap: 16px; margin-bottom: 40px; flex-wrap: wrap; }
|
||||
.sec-num { font-family: var(--mono); font-size: 13px; color: var(--accent); font-weight: 700; }
|
||||
.sec-title { font-family: var(--display); font-weight: 600; font-size: clamp(1.7rem, 4vw, 2.6rem); letter-spacing: -0.025em; }
|
||||
.sec-sub { color: var(--ink-soft); max-width: 52ch; font-size: 1.02rem; }
|
||||
|
||||
/* ── features ── */
|
||||
.feat-grid { display: grid; grid-template-columns: repeat(4,1fr); gap: 1px; background: var(--line); border: 1px solid var(--line); border-radius: 14px; overflow: hidden; }
|
||||
@media (max-width: 900px) { .feat-grid { grid-template-columns: repeat(2,1fr); } }
|
||||
@media (max-width: 560px) { .feat-grid { grid-template-columns: 1fr; } }
|
||||
.feat { background: var(--card); padding: 24px 22px; transition: background .16s; }
|
||||
.feat:hover { background: #fff; }
|
||||
.feat .fn { font-family: var(--mono); font-size: 12px; color: var(--accent); font-weight: 700; }
|
||||
.feat h3 { font-family: var(--display); font-weight: 600; font-size: 1.16rem; margin: 12px 0 8px; letter-spacing: -0.01em; }
|
||||
.feat p { font-size: 14.5px; color: var(--ink-soft); line-height: 1.5; }
|
||||
|
||||
/* ── benchmark ── */
|
||||
.bench {
|
||||
border: 1px solid var(--line); border-radius: 14px; overflow: hidden;
|
||||
background: var(--card);
|
||||
}
|
||||
table { width: 100%; border-collapse: collapse; font-size: 15px; }
|
||||
thead th { font-family: var(--mono); font-size: 11px; letter-spacing: 0.08em; text-transform: uppercase; color: var(--muted); text-align: right; padding: 15px 18px; background: var(--paper-2); border-bottom: 1px solid var(--line); font-weight: 500; }
|
||||
thead th:first-child { text-align: left; }
|
||||
tbody td { padding: 14px 18px; text-align: right; font-family: var(--mono); border-bottom: 1px solid var(--line); }
|
||||
tbody td:first-child { text-align: left; font-family: var(--body); font-weight: 500; }
|
||||
tbody tr:last-child td { border-bottom: none; }
|
||||
tbody tr.us { background: rgba(221,63,34,0.06); }
|
||||
tbody tr.us td:first-child { color: var(--accent-deep); font-weight: 700; }
|
||||
tbody tr.us td:first-child::before { content: "▸ "; color: var(--accent); }
|
||||
.bench-foot { padding: 15px 18px; font-size: 13.5px; color: var(--ink-soft); background: var(--paper-2); border-top: 1px solid var(--line); }
|
||||
.bench-wrap { overflow-x: auto; }
|
||||
.callouts { display: grid; grid-template-columns: 1fr 1fr; gap: 16px; margin-top: 22px; }
|
||||
@media (max-width: 640px) { .callouts { grid-template-columns: 1fr; } }
|
||||
.callout { border-left: 3px solid var(--accent); padding: 4px 0 4px 16px; }
|
||||
.callout .ct { font-family: var(--mono); font-size: 11px; letter-spacing: 0.1em; text-transform: uppercase; color: var(--accent-deep); margin-bottom: 5px; }
|
||||
.callout p { font-size: 14.5px; color: var(--ink-soft); }
|
||||
|
||||
/* ── quickstart tabs ── */
|
||||
.tabs input { position: absolute; opacity: 0; pointer-events: none; }
|
||||
.tablist { display: flex; gap: 6px; margin-bottom: 0; }
|
||||
.tablist label {
|
||||
font-family: var(--mono); font-size: 13px; font-weight: 500;
|
||||
padding: 10px 18px; cursor: pointer; color: var(--muted);
|
||||
border: 1px solid var(--line); border-bottom: none;
|
||||
border-radius: 9px 9px 0 0; background: var(--paper-2); transition: all .15s;
|
||||
}
|
||||
.tablist label:hover { color: var(--ink); }
|
||||
.panel { display: none; }
|
||||
.code {
|
||||
background: var(--ink); border-radius: 0 12px 12px 12px;
|
||||
padding: 22px 24px; overflow-x: auto;
|
||||
font-family: var(--mono); font-size: 13.5px; line-height: 1.7;
|
||||
color: #e9e2d3;
|
||||
box-shadow: 12px 12px 0 rgba(27,23,18,0.07);
|
||||
}
|
||||
.code .cm { color: #8a8069; }
|
||||
.code .kw { color: #ff9f7a; }
|
||||
.code .st { color: #cbb78a; }
|
||||
.code .fn { color: #f4efe4; font-weight: 700; }
|
||||
#t-rust:checked ~ .tablist label[for=t-rust],
|
||||
#t-py:checked ~ .tablist label[for=t-py],
|
||||
#t-node:checked ~ .tablist label[for=t-node],
|
||||
#t-cli:checked ~ .tablist label[for=t-cli] {
|
||||
background: var(--ink); color: var(--paper); border-color: var(--ink);
|
||||
}
|
||||
#t-rust:checked ~ .panels #p-rust,
|
||||
#t-py:checked ~ .panels #p-py,
|
||||
#t-node:checked ~ .panels #p-node,
|
||||
#t-cli:checked ~ .panels #p-cli { display: block; }
|
||||
.code a.ref { color: #ff9f7a; border-bottom: 1px dotted #ff9f7a; }
|
||||
|
||||
/* ── closing split (OSS vs hosted) ── */
|
||||
.split { display: grid; grid-template-columns: 1fr 1fr; gap: 18px; }
|
||||
@media (max-width: 780px) { .split { grid-template-columns: 1fr; } }
|
||||
.path { border: 1px solid var(--line); border-radius: 16px; padding: 32px 30px; background: var(--card); display: flex; flex-direction: column; }
|
||||
.path .ptag { font-family: var(--mono); font-size: 11px; letter-spacing: 0.12em; text-transform: uppercase; color: var(--muted); margin-bottom: 15px; }
|
||||
.path h3 { font-family: var(--display); font-weight: 600; font-size: 1.5rem; letter-spacing: -0.02em; line-height: 1.05; margin-bottom: 12px; }
|
||||
.path p { color: var(--ink-soft); font-size: 15px; line-height: 1.5; flex: 1; margin-bottom: 24px; }
|
||||
.pbtns { display: flex; gap: 12px; flex-wrap: wrap; }
|
||||
.path-pro { background: var(--ink); border-color: var(--ink); box-shadow: 14px 14px 0 rgba(27,23,18,0.09); }
|
||||
.path-pro .ptag { color: #b8ad98; }
|
||||
.path-pro .ptag b { color: var(--accent); font-weight: 700; }
|
||||
.path-pro h3 { color: var(--paper); }
|
||||
.path-pro p { color: #cfc6b3; }
|
||||
.btn { font-family: var(--mono); font-size: 14px; font-weight: 500; padding: 13px 24px; border-radius: 999px; transition: all .15s; border: 1px solid var(--ink); }
|
||||
.btn-p { background: var(--accent); border-color: var(--accent); color: var(--paper); }
|
||||
.btn-p:hover { background: var(--accent-deep); border-color: var(--accent-deep); }
|
||||
.btn-s:hover { background: var(--ink); color: var(--paper); }
|
||||
.btn-pro { background: var(--accent); border-color: var(--accent); color: var(--paper); }
|
||||
.btn-pro:hover { background: var(--accent-deep); border-color: var(--accent-deep); }
|
||||
.btn-ghost { color: var(--paper); border-color: rgba(255,255,255,0.3); }
|
||||
.btn-ghost:hover { border-color: var(--paper); background: rgba(255,255,255,0.08); }
|
||||
.fc-mark { height: 34px; width: auto; display: block; margin-bottom: 20px; }
|
||||
.fc-wordmark { height: 15px; width: auto; vertical-align: -2px; transition: opacity .15s; }
|
||||
.fc-wordmark:hover { opacity: 0.65; }
|
||||
footer { border-top: 1px solid var(--line); padding: 34px 0; font-size: 14px; color: var(--muted); }
|
||||
.foot-in { display: flex; justify-content: space-between; align-items: center; gap: 18px; flex-wrap: wrap; }
|
||||
.foot-in a { color: var(--ink-soft); }
|
||||
.foot-in a:hover { color: var(--accent); }
|
||||
.foot-links { display: flex; gap: 20px; font-family: var(--mono); font-size: 13px; }
|
||||
|
||||
/* ── load animation ── */
|
||||
.reveal { opacity: 0; transform: translateY(14px); animation: rise .7s cubic-bezier(.2,.7,.3,1) forwards; }
|
||||
@keyframes rise { to { opacity: 1; transform: none; } }
|
||||
.d1 { animation-delay: .05s; } .d2 { animation-delay: .15s; } .d3 { animation-delay: .25s; }
|
||||
.d4 { animation-delay: .35s; } .d5 { animation-delay: .45s; } .d6 { animation-delay: .55s; }
|
||||
@media (prefers-reduced-motion: reduce) { .reveal { animation: none; opacity: 1; transform: none; } }
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
|
||||
<nav>
|
||||
<div class="wrap nav-in">
|
||||
<a href="#top" class="brand"><span class="dot"></span>pdf-inspector</a>
|
||||
<div class="nav-links">
|
||||
<a href="#features" class="hide-sm">Features</a>
|
||||
<a href="#benchmark" class="hide-sm">Benchmark</a>
|
||||
<a href="#start">Quick start</a>
|
||||
<a class="nav-gh" href="https://github.com/firecrawl/pdf-inspector">GitHub ↗</a>
|
||||
</div>
|
||||
</div>
|
||||
</nav>
|
||||
|
||||
<header id="top">
|
||||
<div class="wrap hero-grid">
|
||||
<div>
|
||||
<div class="eyebrow reveal d1">Rust · Python · Node · CLI</div>
|
||||
<h1 class="reveal d2">Classify PDFs. Extract Markdown. <span class="strike">No OCR.</span></h1>
|
||||
<p class="lede reveal d3">A fast Rust library that tells text-based PDFs from scanned ones, then extracts position-aware text and clean Markdown — <b>locally, in milliseconds</b>. Skip the OCR bill for the ~54% of PDFs that never needed it.</p>
|
||||
</div>
|
||||
<div class="readout reveal d4" aria-hidden="true">
|
||||
<div class="rlabel"><span>classify_pdf()</span><span>~12ms</span></div>
|
||||
<div class="rrow"><span class="k">type</span><span class="v hot">TextBased</span></div>
|
||||
<div class="rrow"><span class="k">confidence</span><span class="v">0.98</span></div>
|
||||
<div class="rrow"><span class="k">needs_ocr</span><span class="v">false</span></div>
|
||||
<div class="rrow"><span class="k">route</span><span class="v">local → md</span></div>
|
||||
<div class="rrow" style="border-top:1px solid rgba(255,255,255,.09);padding-top:13px">
|
||||
<span class="k">signal</span>
|
||||
<span class="bar"><i style="width:98%"></i></span>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="wrap">
|
||||
<div class="installs">
|
||||
<a class="inst reveal d4" href="https://crates.io/crates/pdf-inspector">
|
||||
<div class="reg"><span>crates.io</span><span class="arrow">↗</span></div>
|
||||
<code>cargo add pdf-inspector</code>
|
||||
</a>
|
||||
<a class="inst reveal d5" href="https://pypi.org/project/pdf-inspector/">
|
||||
<div class="reg"><span>PyPI</span><span class="arrow">↗</span></div>
|
||||
<code>pip install pdf-inspector</code>
|
||||
</a>
|
||||
<a class="inst reveal d6" href="https://www.npmjs.com/package/@firecrawl/pdf-inspector">
|
||||
<div class="reg"><span>npm</span><span class="arrow">↗</span></div>
|
||||
<code>npm i @firecrawl/pdf-inspector</code>
|
||||
</a>
|
||||
</div>
|
||||
</div>
|
||||
</header>
|
||||
|
||||
<section id="features">
|
||||
<div class="wrap">
|
||||
<div class="sec-head">
|
||||
<span class="sec-num">01</span>
|
||||
<h2 class="sec-title">Built for routing, not just reading</h2>
|
||||
</div>
|
||||
<div class="feat-grid">
|
||||
<div class="feat"><div class="fn">01</div><h3>Smart classification</h3><p>TextBased, Scanned, ImageBased, or Mixed in ~10–50ms by sampling content streams. Returns a confidence score and per-page OCR routing.</p></div>
|
||||
<div class="feat"><div class="fn">02</div><h3>Position-aware text</h3><p>Extraction with font info, X/Y coordinates, and automatic multi-column reading order.</p></div>
|
||||
<div class="feat"><div class="fn">03</div><h3>Markdown conversion</h3><p>Headings, bullet/numbered lists, code blocks, tables, bold/italic, URL linking, and page breaks.</p></div>
|
||||
<div class="feat"><div class="fn">04</div><h3>Table detection</h3><p>Rectangle-based detection from drawing ops plus heuristic alignment detection. Financial tables, footnotes, and cross-page continuations.</p></div>
|
||||
<div class="feat"><div class="fn">05</div><h3>CID font support</h3><p>ToUnicode CMap decoding for Type0/Identity-H fonts, with UTF-16BE, UTF-8, and Latin-1 encodings.</p></div>
|
||||
<div class="feat"><div class="fn">06</div><h3>Multi-column layout</h3><p>Newspaper-style column detection, sequential reading order, and right-to-left text support.</p></div>
|
||||
<div class="feat"><div class="fn">07</div><h3>Encoding checks</h3><p>Flags broken font encodings automatically so callers can fall back to OCR only when it's actually needed.</p></div>
|
||||
<div class="feat"><div class="fn">08</div><h3>Lightweight</h3><p>Pure Rust. No ML models, no external services. A single parse shared between detection and extraction.</p></div>
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section id="benchmark">
|
||||
<div class="wrap">
|
||||
<div class="sec-head">
|
||||
<span class="sec-num">02</span>
|
||||
<h2 class="sec-title">Fastest of the direct-text engines</h2>
|
||||
<p class="sec-sub">Evaluated on the <a href="https://github.com/opendataloader-project/opendataloader-bench" style="color:var(--accent-deep);border-bottom:1px solid var(--line)">opendataloader-bench</a> corpus (200 PDFs). Direct text-extraction engines only — no OCR, no ML. Higher is better.</p>
|
||||
</div>
|
||||
<div class="bench">
|
||||
<div class="bench-wrap">
|
||||
<table>
|
||||
<thead>
|
||||
<tr><th>Engine</th><th>Overall</th><th>Reading order</th><th>Tables</th><th>Headings</th><th>200 docs</th></tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr class="us"><td>pdf-inspector</td><td>0.83</td><td>0.88</td><td>0.66</td><td>0.74</td><td>4s</td></tr>
|
||||
<tr><td>opendataloader</td><td>0.84</td><td>0.91</td><td>0.49</td><td>0.74</td><td>11s</td></tr>
|
||||
<tr><td>pymupdf4llm</td><td>0.73</td><td>0.89</td><td>0.40</td><td>0.41</td><td>18s</td></tr>
|
||||
<tr><td>markitdown</td><td>0.58</td><td>0.88</td><td>0.00</td><td>0.00</td><td>8s</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
<div class="bench-foot">OCR/ML engines (docling, marker, mineru) score 0.83–0.88 overall — but take 2–180 minutes on the same corpus.</div>
|
||||
</div>
|
||||
<div class="callouts">
|
||||
<div class="callout"><div class="ct">Where we win</div><p>Fastest engine measured, the best table detection of any engine here, and heading quality now on par with opendataloader — at ~2.5× its speed.</p></div>
|
||||
<div class="callout"><div class="ct">Where we're working</div><p>Reading order still trails opendataloader slightly, and tables that need the visual structure only an OCR engine can see.</p></div>
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section id="start">
|
||||
<div class="wrap">
|
||||
<div class="sec-head">
|
||||
<span class="sec-num">03</span>
|
||||
<h2 class="sec-title">Three lines to Markdown</h2>
|
||||
</div>
|
||||
<div class="tabs">
|
||||
<input type="radio" name="tab" id="t-rust" checked>
|
||||
<input type="radio" name="tab" id="t-py">
|
||||
<input type="radio" name="tab" id="t-node">
|
||||
<input type="radio" name="tab" id="t-cli">
|
||||
<div class="tablist">
|
||||
<label for="t-rust">Rust</label>
|
||||
<label for="t-py">Python</label>
|
||||
<label for="t-node">Node.js</label>
|
||||
<label for="t-cli">CLI</label>
|
||||
</div>
|
||||
<div class="panels">
|
||||
<div class="panel" id="p-rust"><pre class="code"><span class="kw">use</span> pdf_inspector::process_pdf;
|
||||
|
||||
<span class="kw">let</span> result = <span class="fn">process_pdf</span>(<span class="st">"document.pdf"</span>)?;
|
||||
<span class="fn">println!</span>(<span class="st">"Type: {:?}"</span>, result.pdf_type);
|
||||
<span class="kw">if let</span> <span class="kw">Some</span>(markdown) = &result.markdown {
|
||||
<span class="fn">println!</span>(<span class="st">"{}"</span>, markdown);
|
||||
}
|
||||
<span class="cm">// full reference → </span><a class="ref" href="https://github.com/firecrawl/pdf-inspector/blob/main/docs/rust-api.md">docs/rust-api.md</a></pre></div>
|
||||
<div class="panel" id="p-py"><pre class="code"><span class="kw">import</span> pdf_inspector
|
||||
|
||||
result = pdf_inspector.<span class="fn">process_pdf</span>(<span class="st">"document.pdf"</span>)
|
||||
<span class="fn">print</span>(result.pdf_type) <span class="cm"># "text_based" | "scanned" | "image_based" | "mixed"</span>
|
||||
<span class="fn">print</span>(result.markdown) <span class="cm"># Markdown string or None</span>
|
||||
<span class="cm"># full reference → </span><a class="ref" href="https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md">docs/python.md</a></pre></div>
|
||||
<div class="panel" id="p-node"><pre class="code"><span class="kw">import</span> { readFileSync } <span class="kw">from</span> <span class="st">'fs'</span>;
|
||||
<span class="kw">import</span> { processPdf } <span class="kw">from</span> <span class="st">'@firecrawl/pdf-inspector'</span>;
|
||||
|
||||
<span class="kw">const</span> result = <span class="fn">processPdf</span>(<span class="fn">readFileSync</span>(<span class="st">'document.pdf'</span>));
|
||||
console.<span class="fn">log</span>(result.pdfType); <span class="cm">// "TextBased" | "Scanned" | ...</span>
|
||||
console.<span class="fn">log</span>(result.markdown); <span class="cm">// Markdown string or null</span>
|
||||
<span class="cm">// full reference → </span><a class="ref" href="https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md">napi/README.md</a></pre></div>
|
||||
<div class="panel" id="p-cli"><pre class="code"><span class="cm"># install the CLI tools</span>
|
||||
cargo <span class="fn">install</span> pdf-inspector
|
||||
|
||||
<span class="cm"># convert a PDF to Markdown</span>
|
||||
<span class="fn">pdf2md</span> document.pdf
|
||||
|
||||
<span class="cm"># classify only — is it scanned?</span>
|
||||
<span class="fn">detect-pdf</span> document.pdf --analyze --json
|
||||
|
||||
<span class="cm"># structured output for pipelines</span>
|
||||
<span class="fn">pdf2md</span> document.pdf --json</pre></div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="close">
|
||||
<div class="wrap">
|
||||
<div class="sec-head">
|
||||
<span class="sec-num">04</span>
|
||||
<h2 class="sec-title">Two ways to parse</h2>
|
||||
<p class="sec-sub">Run the classifier locally for text-based PDFs; hand the scanned, OCR, and at-scale work to Firecrawl.</p>
|
||||
</div>
|
||||
<div class="split">
|
||||
<div class="path">
|
||||
<div class="ptag">Open source · runs local</div>
|
||||
<h3>Use pdf-inspector yourself</h3>
|
||||
<p>Pure-Rust library and CLI. Classify and extract text-based PDFs on your own machine in milliseconds — no external calls, no OCR bill, MIT licensed.</p>
|
||||
<div class="pbtns">
|
||||
<a class="btn btn-p" href="https://github.com/firecrawl/pdf-inspector">Get started</a>
|
||||
<a class="btn btn-s" href="https://crates.io/crates/pdf-inspector">crates.io</a>
|
||||
</div>
|
||||
</div>
|
||||
<div class="path path-pro">
|
||||
<img class="fc-mark" src="assets/firecrawl-mark.svg" alt="Firecrawl" width="24" height="34">
|
||||
<div class="ptag"><b>Firecrawl Parse</b> · hosted API</div>
|
||||
<h3>Or let Firecrawl handle the hard ones</h3>
|
||||
<p>Scanned documents, OCR, DOCX / XLSX / HTML, and parsing at scale — clean, LLM-ready Markdown from one API call. The downstream route for everything local parsing can't reach.</p>
|
||||
<div class="pbtns">
|
||||
<a class="btn btn-pro" href="https://docs.firecrawl.dev/api-reference/endpoint/parse">Firecrawl Parse ↗</a>
|
||||
<a class="btn btn-ghost" href="https://firecrawl.dev">firecrawl.dev</a>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<footer>
|
||||
<div class="wrap foot-in">
|
||||
<div style="display:flex;align-items:center;gap:7px">Built by <a href="https://firecrawl.dev"><img class="fc-wordmark" src="assets/firecrawl-wordmark.svg" alt="Firecrawl"></a> · MIT licensed</div>
|
||||
<div class="foot-links">
|
||||
<a href="https://github.com/firecrawl/pdf-inspector">GitHub</a>
|
||||
<a href="https://crates.io/crates/pdf-inspector">crates.io</a>
|
||||
<a href="https://pypi.org/project/pdf-inspector/">PyPI</a>
|
||||
<a href="https://www.npmjs.com/package/@firecrawl/pdf-inspector">npm</a>
|
||||
</div>
|
||||
</div>
|
||||
</footer>
|
||||
|
||||
</body>
|
||||
</html>
|
||||
+43
-2
@@ -11,6 +11,37 @@ use std::process;
|
||||
use std::time::Instant;
|
||||
|
||||
/// Escape a string for embedding in a JSON string value.
|
||||
fn format_detector_ocr_reasons(reasons: &std::collections::BTreeMap<u32, Vec<String>>) -> String {
|
||||
reasons
|
||||
.iter()
|
||||
.map(|(page, page_reasons)| {
|
||||
let reasons_json = page_reasons
|
||||
.iter()
|
||||
.map(|reason| format!(r#""{}""#, json_escape(reason)))
|
||||
.collect::<Vec<_>>()
|
||||
.join(",");
|
||||
format!(r#"{{"page":{},"reasons":[{}]}}"#, page, reasons_json)
|
||||
})
|
||||
.collect::<Vec<_>>()
|
||||
.join(",")
|
||||
}
|
||||
|
||||
fn format_ocr_reasons_by_page(reasons: &[pdf_inspector::PageOcrReasons]) -> String {
|
||||
reasons
|
||||
.iter()
|
||||
.map(|entry| {
|
||||
let reasons_json = entry
|
||||
.reasons
|
||||
.iter()
|
||||
.map(|reason| format!(r#""{}""#, json_escape(reason)))
|
||||
.collect::<Vec<_>>()
|
||||
.join(",");
|
||||
format!(r#"{{"page":{},"reasons":[{}]}}"#, entry.page, reasons_json)
|
||||
})
|
||||
.collect::<Vec<_>>()
|
||||
.join(",")
|
||||
}
|
||||
|
||||
fn json_escape(s: &str) -> String {
|
||||
let mut out = String::with_capacity(s.len() + 16);
|
||||
for ch in s.chars() {
|
||||
@@ -117,11 +148,13 @@ fn run_analyze(pdf_path: &str, json_output: bool, start: Instant) {
|
||||
.iter()
|
||||
.map(|p| p.to_string())
|
||||
.collect();
|
||||
let ocr_reasons = format_ocr_reasons_by_page(&result.ocr_reasons_by_page);
|
||||
println!(
|
||||
r#"{{"pdf_type":"{}","page_count":{},"pages_needing_ocr":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"detection_time_ms":{}}}"#,
|
||||
r#"{{"pdf_type":"{}","page_count":{},"pages_needing_ocr":[{}],"ocr_reasons_by_page":[{}],"is_complex":{},"pages_with_tables":[{}],"pages_with_columns":[{}],"detection_time_ms":{}}}"#,
|
||||
pdf_type_str(&result.pdf_type),
|
||||
result.page_count,
|
||||
ocr_pages.join(","),
|
||||
ocr_reasons,
|
||||
result.layout.is_complex,
|
||||
table_pages.join(","),
|
||||
col_pages.join(","),
|
||||
@@ -144,6 +177,9 @@ fn run_analyze(pdf_path: &str, json_output: bool, start: Instant) {
|
||||
println!("Page count: {}", result.page_count);
|
||||
if !result.pages_needing_ocr.is_empty() {
|
||||
println!("Pages needing OCR: {:?}", result.pages_needing_ocr);
|
||||
for entry in &result.ocr_reasons_by_page {
|
||||
println!(" page {}: {}", entry.page, entry.reasons.join(", "));
|
||||
}
|
||||
}
|
||||
println!();
|
||||
if result.layout.is_complex {
|
||||
@@ -184,8 +220,9 @@ fn run_detect_only(pdf_path: &str, json_output: bool, start: Instant) {
|
||||
.iter()
|
||||
.map(|p| p.to_string())
|
||||
.collect();
|
||||
let ocr_reasons = format_detector_ocr_reasons(&result.ocr_reasons_by_page);
|
||||
println!(
|
||||
r#"{{"pdf_type":"{}","page_count":{},"pages_sampled":{},"pages_with_text":{},"confidence":{:.2},"title":{},"ocr_recommended":{},"pages_needing_ocr":[{}],"detection_time_ms":{}}}"#,
|
||||
r#"{{"pdf_type":"{}","page_count":{},"pages_sampled":{},"pages_with_text":{},"confidence":{:.2},"title":{},"ocr_recommended":{},"pages_needing_ocr":[{}],"ocr_reasons_by_page":[{}],"detection_time_ms":{}}}"#,
|
||||
pdf_type_str(&result.pdf_type),
|
||||
result.page_count,
|
||||
result.pages_sampled,
|
||||
@@ -198,6 +235,7 @@ fn run_detect_only(pdf_path: &str, json_output: bool, start: Instant) {
|
||||
.unwrap_or_else(|| "null".to_string()),
|
||||
result.ocr_recommended,
|
||||
ocr_pages.join(","),
|
||||
ocr_reasons,
|
||||
elapsed.as_millis()
|
||||
);
|
||||
} else {
|
||||
@@ -232,6 +270,9 @@ fn run_detect_only(pdf_path: &str, json_output: bool, start: Instant) {
|
||||
result.pages_needing_ocr, result.page_count
|
||||
);
|
||||
}
|
||||
for (page, reasons) in &result.ocr_reasons_by_page {
|
||||
println!(" page {}: {}", page, reasons.join(", "));
|
||||
}
|
||||
}
|
||||
if let Some(title) = &result.title {
|
||||
println!("Title: {}", title);
|
||||
|
||||
@@ -208,6 +208,7 @@ fn main() {
|
||||
eprintln!(" --raw Output only markdown (no headers)");
|
||||
eprintln!(" --pages Insert page break markers (<!-- Page N -->)");
|
||||
eprintln!(" --select-pages N Only process specified pages (e.g. 1,3,5-10)");
|
||||
eprintln!(" --password PW Password for an encrypted PDF");
|
||||
eprintln!(" --detect-only Only detect PDF type (no extraction)");
|
||||
eprintln!(" --analyze Detect + extract + layout analysis (no markdown)");
|
||||
process::exit(1);
|
||||
@@ -221,6 +222,16 @@ fn main() {
|
||||
let detect_only = args.iter().any(|a| a == "--detect-only");
|
||||
let analyze = args.iter().any(|a| a == "--analyze");
|
||||
|
||||
// Parse --password value
|
||||
let password = args.iter().position(|a| a == "--password").map(|i| {
|
||||
args.get(i + 1)
|
||||
.unwrap_or_else(|| {
|
||||
eprintln!("Error: --password requires a value");
|
||||
process::exit(1);
|
||||
})
|
||||
.clone()
|
||||
});
|
||||
|
||||
// Parse --select-pages value
|
||||
let page_filter = args
|
||||
.iter()
|
||||
@@ -269,6 +280,7 @@ fn main() {
|
||||
if let Some(pages) = page_filter {
|
||||
options.page_filter = Some(pages);
|
||||
}
|
||||
options.password = password;
|
||||
|
||||
match process_pdf_with_options(pdf_path, options) {
|
||||
Ok(result) => {
|
||||
|
||||
+111
-2
@@ -60,6 +60,10 @@ pub struct PdfTypeResult {
|
||||
/// 1-indexed page numbers that need OCR (image-only or insufficient text).
|
||||
/// Empty for TextBased. All pages for Scanned/ImageBased. Specific pages for Mixed.
|
||||
pub pages_needing_ocr: Vec<u32>,
|
||||
/// Per-page explanation for `pages_needing_ocr`: 1-indexed page → reason
|
||||
/// codes (`scanned`, `no_text`, `vector_text`, `suspected_garbled_text`).
|
||||
/// Only contains pages that need OCR.
|
||||
pub ocr_reasons_by_page: std::collections::BTreeMap<u32, Vec<String>>,
|
||||
}
|
||||
|
||||
/// Configuration for PDF type detection
|
||||
@@ -382,7 +386,12 @@ pub(crate) fn detect_from_document(
|
||||
let analysis = if let Some(cached) = analysis_cache.get(&page_num) {
|
||||
cached.clone()
|
||||
} else if let Some(&page_id) = pages.get(&page_num) {
|
||||
analyze_page_content(doc, page_id)
|
||||
// Cache the fresh analysis so the reason-classification pass
|
||||
// below sees the real signals (vector_text, etc.) instead of
|
||||
// defaulting to "scanned".
|
||||
let a = analyze_page_content(doc, page_id);
|
||||
analysis_cache.insert(page_num, a.clone());
|
||||
a
|
||||
} else {
|
||||
continue;
|
||||
};
|
||||
@@ -429,6 +438,9 @@ pub(crate) fn detect_from_document(
|
||||
let analysis = analyze_page_content(doc, page_id);
|
||||
if analysis.has_identity_h_no_tounicode || analysis.has_only_type3_fonts {
|
||||
pages_needing_ocr.push(page_num);
|
||||
// Cache so the reason pass reports suspected_garbled_text
|
||||
// rather than defaulting to "scanned".
|
||||
analysis_cache.insert(page_num, analysis);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -436,6 +448,19 @@ pub(crate) fn detect_from_document(
|
||||
pages_needing_ocr.sort();
|
||||
pages_needing_ocr.dedup();
|
||||
|
||||
// Explain each OCR-flagged page. Pages we analyzed get a signal-derived
|
||||
// reason; pages flagged only by whole-document classification (unsampled
|
||||
// pages of a Scanned/ImageBased doc) default to `scanned`.
|
||||
let mut ocr_reasons_by_page: std::collections::BTreeMap<u32, Vec<String>> =
|
||||
std::collections::BTreeMap::new();
|
||||
for &page_num in &pages_needing_ocr {
|
||||
let reasons = match analysis_cache.get(&page_num) {
|
||||
Some(analysis) => page_ocr_reasons(analysis),
|
||||
None => vec![crate::OCR_REASON_SCANNED],
|
||||
};
|
||||
ocr_reasons_by_page.insert(page_num, reasons.into_iter().map(String::from).collect());
|
||||
}
|
||||
|
||||
// Try to get title from metadata
|
||||
let title = get_document_title(doc);
|
||||
|
||||
@@ -448,6 +473,7 @@ pub(crate) fn detect_from_document(
|
||||
title,
|
||||
ocr_recommended,
|
||||
pages_needing_ocr,
|
||||
ocr_reasons_by_page,
|
||||
})
|
||||
}
|
||||
|
||||
@@ -487,7 +513,7 @@ fn distribute_pages(n: u32, total: u32) -> Vec<u32> {
|
||||
}
|
||||
|
||||
/// Page content analysis result
|
||||
#[derive(Clone)]
|
||||
#[derive(Clone, Default)]
|
||||
struct PageAnalysis {
|
||||
text_operator_count: u32,
|
||||
has_images: bool,
|
||||
@@ -523,6 +549,31 @@ struct PageAnalysis {
|
||||
has_decodable_text_fonts: bool,
|
||||
}
|
||||
|
||||
/// Explain *why* a page needs OCR, from its content analysis. Priority:
|
||||
/// undecodable fonts (`suspected_garbled_text`) and vector-outlined text
|
||||
/// (`vector_text`) come first because they persist even when a text layer is
|
||||
/// present; otherwise a page with no extractable text is `scanned` when an
|
||||
/// image backs it or `no_text` when nothing does.
|
||||
fn page_ocr_reasons(a: &PageAnalysis) -> Vec<&'static str> {
|
||||
let mut reasons = Vec::new();
|
||||
if a.has_identity_h_no_tounicode || a.has_only_type3_fonts {
|
||||
reasons.push(crate::OCR_REASON_SUSPECTED_GARBLED_TEXT);
|
||||
}
|
||||
if a.has_vector_text {
|
||||
reasons.push(crate::OCR_REASON_VECTOR_TEXT);
|
||||
}
|
||||
if reasons.is_empty() {
|
||||
let has_extractable_text = a.text_operator_count > 0 && a.unique_text_chars > 0;
|
||||
if !has_extractable_text && !a.has_images && !a.has_template_image {
|
||||
reasons.push(crate::OCR_REASON_NO_TEXT);
|
||||
} else {
|
||||
// Image-backed with no usable text, or too little text to trust.
|
||||
reasons.push(crate::OCR_REASON_SCANNED);
|
||||
}
|
||||
}
|
||||
reasons
|
||||
}
|
||||
|
||||
/// Extracted font information from a Resource dictionary entry.
|
||||
/// Stores the properties needed for decodability/identity-h checks
|
||||
/// without holding a reference to the document.
|
||||
@@ -1809,6 +1860,64 @@ fn get_document_title(doc: &Document) -> Option<String> {
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn page_ocr_reasons_classify() {
|
||||
// Scanned: no text, full-page image.
|
||||
let scanned = PageAnalysis {
|
||||
has_template_image: true,
|
||||
..Default::default()
|
||||
};
|
||||
assert_eq!(page_ocr_reasons(&scanned), vec![crate::OCR_REASON_SCANNED]);
|
||||
|
||||
// Image-only page (no template flag, but has an image).
|
||||
let image_only = PageAnalysis {
|
||||
has_images: true,
|
||||
..Default::default()
|
||||
};
|
||||
assert_eq!(
|
||||
page_ocr_reasons(&image_only),
|
||||
vec![crate::OCR_REASON_SCANNED]
|
||||
);
|
||||
|
||||
// No text, no image → no_text.
|
||||
let blank = PageAnalysis::default();
|
||||
assert_eq!(page_ocr_reasons(&blank), vec![crate::OCR_REASON_NO_TEXT]);
|
||||
|
||||
// Vector-outlined text.
|
||||
let vector = PageAnalysis {
|
||||
has_vector_text: true,
|
||||
..Default::default()
|
||||
};
|
||||
assert_eq!(
|
||||
page_ocr_reasons(&vector),
|
||||
vec![crate::OCR_REASON_VECTOR_TEXT]
|
||||
);
|
||||
|
||||
// Undecodable fonts → garbled, and it wins over the fall-through.
|
||||
let garbled = PageAnalysis {
|
||||
has_identity_h_no_tounicode: true,
|
||||
has_images: true,
|
||||
..Default::default()
|
||||
};
|
||||
assert_eq!(
|
||||
page_ocr_reasons(&garbled),
|
||||
vec![crate::OCR_REASON_SUSPECTED_GARBLED_TEXT]
|
||||
);
|
||||
|
||||
// A page with real extractable text and an image is not flagged here
|
||||
// as scanned/no_text (only reached for pages already needing OCR).
|
||||
let text_with_image = PageAnalysis {
|
||||
text_operator_count: 40,
|
||||
unique_text_chars: 120,
|
||||
has_images: true,
|
||||
..Default::default()
|
||||
};
|
||||
assert_eq!(
|
||||
page_ocr_reasons(&text_with_image),
|
||||
vec![crate::OCR_REASON_SCANNED]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_scan_content_operators() {
|
||||
let mut uchars = HashSet::new();
|
||||
|
||||
+290
-9
@@ -9,7 +9,7 @@ mod links;
|
||||
pub(crate) mod underline;
|
||||
mod xobjects;
|
||||
|
||||
use crate::text_utils::is_rtl_text;
|
||||
use crate::text_utils::{is_cjk_char, is_rtl_text};
|
||||
use crate::tounicode::FontCMaps;
|
||||
use crate::types::{PageExtraction, PdfLine, PdfRect, TextItem};
|
||||
use crate::PdfError;
|
||||
@@ -252,17 +252,47 @@ fn suppress_table_underlines(
|
||||
}
|
||||
|
||||
let mut table_item_indices: HashSet<usize> = HashSet::new();
|
||||
// A detected "table" that swallows nearly every text item on the page
|
||||
// is a detection artifact (prose pages with boxed callouts or stacked
|
||||
// underline rules read as one giant grid), not a real table — letting
|
||||
// it through here erased every legitimate underline on the page
|
||||
// (text_dense__underline: rect detection claimed 52/52 items). Real
|
||||
// ruled tables share the page with headings, captions, and body text.
|
||||
let plausible = |table: &crate::tables::Table| {
|
||||
// Content sanity gate: prose pages with boxed callouts and stacked
|
||||
// underline rules can detect as a structurally rich "table" that
|
||||
// swallows every item on the page (text_dense__underline: a 4x8
|
||||
// grid claiming 52/52 items, one "cell" holding 806 chars of body
|
||||
// text) — suppressing there erased every legitimate underline on
|
||||
// the page. Real data-table cells are short values; a cell with
|
||||
// hundreds of characters means the grid captured flowing prose.
|
||||
let lens: Vec<usize> = table
|
||||
.cells
|
||||
.iter()
|
||||
.flatten()
|
||||
.filter(|cell| !cell.trim().is_empty())
|
||||
.map(|cell| cell.chars().count())
|
||||
.collect();
|
||||
if lens.is_empty() {
|
||||
return false;
|
||||
}
|
||||
let long = lens.iter().filter(|&&n| n > 100).count();
|
||||
(long as f32) < (lens.len() as f32) * 0.3
|
||||
};
|
||||
|
||||
if !rects.is_empty() {
|
||||
let (rect_tables, _) = crate::tables::detect_tables_from_rects(items, rects, page);
|
||||
for table in rect_tables {
|
||||
table_item_indices.extend(table.item_indices);
|
||||
for table in rect_tables.iter().filter(|t| plausible(t)) {
|
||||
table_item_indices.extend(table.item_indices.iter().copied());
|
||||
}
|
||||
}
|
||||
|
||||
if !lines.is_empty() {
|
||||
for table in crate::tables::detect_tables_from_lines(items, lines, page) {
|
||||
table_item_indices.extend(table.item_indices);
|
||||
for table in crate::tables::detect_tables_from_lines(items, lines, page)
|
||||
.iter()
|
||||
.filter(|t| plausible(t))
|
||||
{
|
||||
table_item_indices.extend(table.item_indices.iter().copied());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -527,6 +557,136 @@ fn should_preserve_overlapping_stream_order(group: &[&TextItem]) -> bool {
|
||||
saw_backtrack
|
||||
}
|
||||
|
||||
/// Detect a tracked (letter-spaced) run of single-glyph items and derive its
|
||||
/// run-local space floor.
|
||||
///
|
||||
/// Display type set with tracking renders one glyph per show op; the merge
|
||||
/// loop's fixed thresholds (0.08-0.13 em) then read every letter gap as a
|
||||
/// word boundary and emit "H O W" instead of "HOW". Within such a run the
|
||||
/// gaps carry the real signal: letter gaps cluster tightly just above the
|
||||
/// fixed threshold, word gaps sit clearly higher. Returns (run_end_index,
|
||||
/// space_floor) when the run starting at `start` is tracked — spaces are
|
||||
/// then inserted only at gaps above the floor (infinity = single word).
|
||||
/// Normal text (multi-char items, or single-char runs with sub-threshold
|
||||
/// gaps) returns None and keeps the existing behavior.
|
||||
/// Han/Kana scripts write without inter-word spaces. Hangul (Korean) DOES
|
||||
/// space between words and deliberately stays out of this set — a Korean
|
||||
/// tracked run keeps normal word-boundary handling.
|
||||
fn is_spaceless_cjk(c: char) -> bool {
|
||||
matches!(c,
|
||||
'\u{3000}'..='\u{303F}' // CJK Symbols and Punctuation
|
||||
| '\u{3040}'..='\u{309F}' // Hiragana
|
||||
| '\u{30A0}'..='\u{30FF}' // Katakana
|
||||
| '\u{4E00}'..='\u{9FFF}' // CJK Unified Ideographs
|
||||
| '\u{F900}'..='\u{FAFF}' // CJK Compatibility Ideographs
|
||||
| '\u{FF00}'..='\u{FFEF}' // Halfwidth and Fullwidth Forms
|
||||
)
|
||||
}
|
||||
|
||||
fn tracked_run_space_floor(group: &[&TextItem], start: usize) -> Option<(usize, f32)> {
|
||||
const MIN_GAPS: usize = 4;
|
||||
let first = group[start];
|
||||
if first.text.trim().chars().count() != 1 {
|
||||
return None;
|
||||
}
|
||||
let fs = first.font_size;
|
||||
if fs <= 0.0 {
|
||||
return None;
|
||||
}
|
||||
|
||||
// Walk the run under the SAME break conditions as the merge loop
|
||||
// (size band, style equality, mergeable gap) so indices stay aligned.
|
||||
let mut gaps: Vec<f32> = Vec::new();
|
||||
let mut end_x = first.x + effective_merge_width(first);
|
||||
let mut end = start;
|
||||
for (offset, next) in group[start + 1..].iter().enumerate() {
|
||||
if next.text.trim().chars().count() != 1 {
|
||||
break;
|
||||
}
|
||||
if (next.font_size - fs).abs() > fs * 0.20 {
|
||||
break;
|
||||
}
|
||||
if next.is_bold != first.is_bold
|
||||
|| next.is_italic != first.is_italic
|
||||
|| next.is_underline != first.is_underline
|
||||
|| next.is_strikeout != first.is_strikeout
|
||||
{
|
||||
break;
|
||||
}
|
||||
let gap = next.x - end_x;
|
||||
if gap > fs * 0.5 || gap < -fs * 0.5 {
|
||||
break;
|
||||
}
|
||||
gaps.push(gap / fs);
|
||||
end_x = next.x + effective_merge_width(next);
|
||||
end = start + 1 + offset;
|
||||
}
|
||||
if gaps.len() < 2 {
|
||||
return None;
|
||||
}
|
||||
|
||||
// Tracked signature: the run's TYPICAL gap clears the fixed space
|
||||
// threshold (0.08) — the merge loop would break almost every letter
|
||||
// pair into "words". Short runs (2-3 gaps: "H O W") demand a stricter
|
||||
// shape — clearly wide, uniform, ALL-CAPS — because a genuine spaced
|
||||
// sequence of single letters ("x y z" variables) has the same gap
|
||||
// count; display tracking is a caps convention.
|
||||
let mut sorted = gaps.clone();
|
||||
sorted.sort_by(|a, b| a.total_cmp(b));
|
||||
let median = sorted[sorted.len() / 2];
|
||||
// Typographic convention gate, both tiers: display tracking is an
|
||||
// all-caps convention, and Han/Kana never space between glyphs. Mixed-
|
||||
// or lowercase Latin runs keep their boundaries because geometry alone
|
||||
// cannot distinguish spaced singles ("A b c d e") from a tracked
|
||||
// title-case word ("B u f f a l o").
|
||||
let run_chars = || {
|
||||
group[start..=end]
|
||||
.iter()
|
||||
.flat_map(|it| it.text.trim().chars())
|
||||
};
|
||||
let spaceless_cjk = run_chars().all(|c| is_spaceless_cjk(c) || !c.is_alphanumeric())
|
||||
&& run_chars().any(is_spaceless_cjk);
|
||||
let all_caps = run_chars().all(|c| c.is_uppercase() || is_cjk_char(c) || !c.is_alphabetic());
|
||||
if !(spaceless_cjk || all_caps) {
|
||||
return None;
|
||||
}
|
||||
|
||||
if gaps.len() >= MIN_GAPS {
|
||||
if median <= 0.075 {
|
||||
return None;
|
||||
}
|
||||
} else {
|
||||
let uniform = sorted[sorted.len() - 1] <= sorted[0].max(0.01) * 1.4;
|
||||
if median < 0.09 || !uniform {
|
||||
return None;
|
||||
}
|
||||
}
|
||||
|
||||
// Han/Kana: no inter-glyph spaces, period — a nonuniform gap
|
||||
// distribution (punctuation spacing, justification) must not
|
||||
// manufacture word boundaries.
|
||||
if spaceless_cjk {
|
||||
return Some((end, f32::INFINITY));
|
||||
}
|
||||
|
||||
// Word gaps, if present, form a second mode above the letter-gap
|
||||
// cluster: split at the largest relative jump. Unimodal → one word.
|
||||
let mut best_jump = 1.0f32;
|
||||
let mut floor = f32::INFINITY;
|
||||
for pair in sorted.windows(2) {
|
||||
let (lo, hi) = (pair[0].max(0.01), pair[1].max(0.01));
|
||||
let jump = hi / lo;
|
||||
if jump > best_jump {
|
||||
best_jump = jump;
|
||||
floor = (lo + hi) / 2.0;
|
||||
}
|
||||
}
|
||||
if best_jump < 1.4 {
|
||||
floor = f32::INFINITY;
|
||||
}
|
||||
Some((end, floor * fs))
|
||||
}
|
||||
|
||||
pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
|
||||
if items.is_empty() {
|
||||
return items;
|
||||
@@ -574,6 +734,14 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
|
||||
let mut text = first.text.clone();
|
||||
let mut end_x = first.x + effective_merge_width(first);
|
||||
|
||||
// Tracked display text: run-local space floor overrides the
|
||||
// fixed thresholds for this run's junctions (see helper).
|
||||
let tracked = if *preserve_stream_order {
|
||||
None
|
||||
} else {
|
||||
tracked_run_space_floor(group, i)
|
||||
};
|
||||
|
||||
let mut j = i + 1;
|
||||
while j < group.len() {
|
||||
let next = group[j];
|
||||
@@ -628,7 +796,11 @@ pub(crate) fn merge_text_items(items: Vec<TextItem>) -> Vec<TextItem> {
|
||||
let needs_bullet_space = *preserve_stream_order
|
||||
&& is_standalone_bullet_text(&text)
|
||||
&& !next.text.trim().is_empty();
|
||||
if needs_bullet_space || gap > threshold {
|
||||
let effective_threshold = match tracked {
|
||||
Some((run_end, floor)) if j <= run_end => floor,
|
||||
_ => threshold,
|
||||
};
|
||||
if needs_bullet_space || gap > effective_threshold {
|
||||
text.push(' ');
|
||||
}
|
||||
text.push_str(&next.text);
|
||||
@@ -731,9 +903,17 @@ pub(crate) fn merge_subscript_items(items: Vec<TextItem>) -> Vec<TextItem> {
|
||||
.chars()
|
||||
.last()
|
||||
.is_some_and(|c| c.is_alphabetic());
|
||||
let same_marks = parent.is_underline == item.is_underline
|
||||
&& parent.is_strikeout == item.is_strikeout;
|
||||
if parent.font_size >= sub_threshold && ends_with_letter && same_marks {
|
||||
// Strikeout boundaries block the merge (a struck word
|
||||
// must not extend its strike over a live footnote digit,
|
||||
// and a struck digit must not lose its own mark). An
|
||||
// underlined parent with an unmarked digit DOES merge:
|
||||
// the drawn rule easily misses the tiny digit's overlap
|
||||
// window, and refusing costs the whole subscript token
|
||||
// ("b"+"2" staying split). Visually the rule spans both.
|
||||
let marks_ok = parent.is_strikeout == item.is_strikeout
|
||||
&& (parent.is_underline == item.is_underline
|
||||
|| (parent.is_underline && !item.is_underline));
|
||||
if parent.font_size >= sub_threshold && ends_with_letter && marks_ok {
|
||||
let parent_right = parent.x + parent.width;
|
||||
let gap = item.x - parent_right;
|
||||
// Subscripts must be tightly adjacent (within ~1pt)
|
||||
@@ -794,6 +974,107 @@ mod tests {
|
||||
use crate::types::{ItemType, PdfLine, TextLine};
|
||||
use layout::{detect_columns, is_newspaper_layout, ColumnRegion};
|
||||
|
||||
/// Glyph-per-item run at `fs`=12 with the given inter-glyph gap (pt).
|
||||
fn glyph_run(chars: &str, start_x: f32, glyph_w: f32, gap: f32) -> Vec<TextItem> {
|
||||
let mut x = start_x;
|
||||
let mut out = Vec::new();
|
||||
for c in chars.chars() {
|
||||
out.push(make_merge_item(&c.to_string(), x, glyph_w));
|
||||
x += glyph_w + gap;
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn tracked_caps_run_collapses_to_word() {
|
||||
// Display tracking: every letter gap (0.19 em) clears the fixed
|
||||
// space threshold — without the run-local floor this reads "H O W".
|
||||
let items = glyph_run("HOW", 100.0, 10.0, 2.3);
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1);
|
||||
assert_eq!(merged[0].text, "HOW");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn tracked_run_keeps_word_gaps_bimodal() {
|
||||
// Letters at 0.19 em, word gaps at 0.42 em (below the 0.5 em item
|
||||
// break): the split must land between the modes. Needs >=4 gaps to
|
||||
// enter the bimodal tier — short runs use the strict uniform gate.
|
||||
let mut items = glyph_run("ITISOK", 100.0, 8.0, 2.3);
|
||||
for i in 2..6 {
|
||||
items[i].x += 2.8; // word gap at T|I
|
||||
}
|
||||
for i in 4..6 {
|
||||
items[i].x += 2.8; // word gap at S|O
|
||||
}
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1);
|
||||
assert_eq!(merged[0].text, "IT IS OK");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn lowercase_spaced_singles_stay_words() {
|
||||
// "x y z" variables: same gap shape but lowercase — the short-run
|
||||
// caps requirement keeps genuine spaced singles apart.
|
||||
let items = glyph_run("xyz", 100.0, 6.0, 2.3);
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1);
|
||||
assert_eq!(merged[0].text, "x y z");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn kerned_singles_unaffected() {
|
||||
// Tiny kerning gaps never triggered spaces before and still don't.
|
||||
let items = glyph_run("WORD", 100.0, 8.0, 0.3);
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1);
|
||||
assert_eq!(merged[0].text, "WORD");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn long_lowercase_spaced_singles_keep_boundaries() {
|
||||
// Review: a 5+ single-letter lowercase list has the tracked gap
|
||||
// shape at any length — the convention gate must protect it in
|
||||
// the >=4-gap tier too.
|
||||
let items = glyph_run("abcde", 100.0, 6.0, 2.3);
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1);
|
||||
assert_eq!(merged[0].text, "a b c d e");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn han_run_with_nonuniform_gaps_never_gains_spaces() {
|
||||
// Review: a bimodal gap distribution (justification, punctuation
|
||||
// spacing) must not manufacture word boundaries in Han text.
|
||||
let mut items = glyph_run("北京时事快报", 100.0, 12.0, 1.4);
|
||||
for item in items.iter_mut().skip(3) {
|
||||
item.x += 3.0; // wide gap after the third glyph
|
||||
}
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1);
|
||||
assert_eq!(merged[0].text, "北京时事快报");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn uppercase_leading_spaced_singles_keep_boundaries() {
|
||||
// "A b c d e" is indistinguishable from a title-case tracked word
|
||||
// without reliable tracking metadata, so preserve its boundaries.
|
||||
let items = glyph_run("Abcde", 100.0, 7.0, 2.3);
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1);
|
||||
assert_eq!(merged[0].text, "A b c d e");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cjk_glyph_run_collapses_without_spaces() {
|
||||
// CJK sets one glyph per item with loose gaps; CJK uses no spaces,
|
||||
// and the non-alphabetic run passes the caps gate.
|
||||
let items = glyph_run("北京时事", 100.0, 12.0, 1.4);
|
||||
let merged = merge_text_items(items);
|
||||
assert_eq!(merged.len(), 1);
|
||||
assert_eq!(merged[0].text, "北京时事");
|
||||
}
|
||||
|
||||
fn make_merge_item(text: &str, x: f32, width: f32) -> TextItem {
|
||||
TextItem {
|
||||
text: text.into(),
|
||||
|
||||
+288
-7
@@ -115,7 +115,13 @@ fn rules_from_graphics(rects: &[PdfRect], lines: &[UnderlineLine], page: u32) ->
|
||||
rules
|
||||
}
|
||||
|
||||
fn discard_repeated_ruling_rules(rules: Vec<Rule>) -> Vec<Rule> {
|
||||
fn discard_repeated_ruling_rules(
|
||||
rules: Vec<Rule>,
|
||||
items: &[TextItem],
|
||||
rects: &[PdfRect],
|
||||
lines: &[UnderlineLine],
|
||||
page: u32,
|
||||
) -> Vec<Rule> {
|
||||
if rules.len() < MIN_REPEATED_RULE_LEVELS {
|
||||
return rules;
|
||||
}
|
||||
@@ -123,12 +129,142 @@ fn discard_repeated_ruling_rules(rules: Vec<Rule>) -> Vec<Rule> {
|
||||
rules
|
||||
.iter()
|
||||
.filter(|rule| {
|
||||
!is_repeated_ruling_rule(rule, &rules) && !is_segmented_row_ruling_rule(rule, &rules)
|
||||
// A rule snugly owned by one text line is an underline even when
|
||||
// span-similar rules repeat down the page — documents that
|
||||
// underline many full-width lines (dense CJK business docs) look
|
||||
// exactly like table rulings to the repetition check, which used
|
||||
// to discard every one of them. Table rulings fail snugness:
|
||||
// row separators extend past their cells' text (or have no text
|
||||
// on the baseline above), and multi-column matches are still
|
||||
// culled by the tabular filter afterwards.
|
||||
// Same-row segmented rules (column-header separators) are
|
||||
// always rulings — each segment snugly owns its column label,
|
||||
// so snugness must not override that check.
|
||||
!is_segmented_row_ruling_rule(rule, &rules)
|
||||
&& ((has_snug_text_owner(rule, items)
|
||||
&& !has_flanking_verticals(rule, rects, lines, page))
|
||||
|| !is_repeated_ruling_rule(rule, &rules))
|
||||
})
|
||||
.cloned()
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// True when a single text item both matches the rule vertically (baseline
|
||||
/// window) and horizontally contains it: the rule may not extend past the
|
||||
/// item's span by more than ~0.75em on either side. Underlines are drawn to
|
||||
/// the width of the text they decorate; table/form rulings span cells or
|
||||
/// full table width and overshoot any single item.
|
||||
/// A rule flanked by vertical strokes at its ends is a table/box border
|
||||
/// row edge, not an underline — underlined text lines have no vertical
|
||||
/// rules rising from their ends. Checked against raw stroked lines: a
|
||||
/// near-vertical segment whose x sits at either end of the rule and whose
|
||||
/// y-range covers the rule's row.
|
||||
fn has_flanking_verticals(
|
||||
rule: &Rule,
|
||||
rects: &[PdfRect],
|
||||
lines: &[UnderlineLine],
|
||||
page: u32,
|
||||
) -> bool {
|
||||
// A drawn rect that CONTAINS the rule vetoes rescue only with GRID
|
||||
// EVIDENCE: another drawn rect abutting it vertically (cell rows tile).
|
||||
// Height alone can't separate a table cell from a decorative callout
|
||||
// panel — genuine underlines live inside isolated filled panels, and
|
||||
// multiline table cells can be arbitrarily tall.
|
||||
let norm = |r: &PdfRect| {
|
||||
let (x_lo, x_hi) = if r.width >= 0.0 {
|
||||
(r.x, r.x + r.width)
|
||||
} else {
|
||||
(r.x + r.width, r.x)
|
||||
};
|
||||
let (y_lo, y_hi) = if r.height >= 0.0 {
|
||||
(r.y, r.y + r.height)
|
||||
} else {
|
||||
(r.y + r.height, r.y)
|
||||
};
|
||||
(x_lo, x_hi, y_lo, y_hi)
|
||||
};
|
||||
let page_rects: Vec<(f32, f32, f32, f32)> = rects
|
||||
.iter()
|
||||
.filter(|r| r.page == page && r.height.abs() > 6.0)
|
||||
.map(norm)
|
||||
.collect();
|
||||
let rect_flank = page_rects.iter().any(|&(x_lo, x_hi, y_lo, y_hi)| {
|
||||
let contains = x_lo <= rule.x1 + 2.0
|
||||
&& x_hi >= rule.x2 - 2.0
|
||||
&& y_lo <= rule.y + 2.0
|
||||
&& y_hi >= rule.y - 2.0;
|
||||
if !contains {
|
||||
return false;
|
||||
}
|
||||
// Grid evidence: a vertically abutting neighbor box with x-overlap.
|
||||
page_rects.iter().any(|&(nx_lo, nx_hi, ny_lo, ny_hi)| {
|
||||
let x_overlap = nx_hi.min(x_hi) - nx_lo.max(x_lo);
|
||||
if x_overlap <= 10.0 {
|
||||
return false;
|
||||
}
|
||||
(ny_lo - y_hi).abs() <= 3.0 || (y_lo - ny_hi).abs() <= 3.0
|
||||
})
|
||||
});
|
||||
if rect_flank {
|
||||
return true;
|
||||
}
|
||||
lines.iter().any(|l| {
|
||||
if l.page != page || (l.x1 - l.x2).abs() > 2.0 {
|
||||
return false;
|
||||
}
|
||||
let x = (l.x1 + l.x2) / 2.0;
|
||||
let near_end = (x - rule.x1).abs() <= 6.0 || (x - rule.x2).abs() <= 6.0;
|
||||
if !near_end {
|
||||
return false;
|
||||
}
|
||||
let (y_lo, y_hi) = if l.y1 <= l.y2 {
|
||||
(l.y1, l.y2)
|
||||
} else {
|
||||
(l.y2, l.y1)
|
||||
};
|
||||
y_lo <= rule.y + 2.0 && y_hi >= rule.y - 2.0
|
||||
})
|
||||
}
|
||||
|
||||
fn has_snug_text_owner(rule: &Rule, items: &[TextItem]) -> bool {
|
||||
// Underlines are drawn to the width of the text they decorate, but the
|
||||
// text may be split into several runs (CJK lines mix scripts and font
|
||||
// switches) — so ownership is judged against the UNION of the runs on
|
||||
// the rule's baseline row. Table/form rulings overshoot their row's
|
||||
// text (row separators span cell padding and empty columns), so they
|
||||
// fail either containment or coverage.
|
||||
let matched: Vec<&TextItem> = items
|
||||
.iter()
|
||||
.filter(|item| is_underline_candidate(item) && rule_matches_item(rule, item))
|
||||
.collect();
|
||||
if matched.is_empty() {
|
||||
return false;
|
||||
}
|
||||
let x1 = matched.iter().map(|i| i.x).fold(f32::INFINITY, f32::min);
|
||||
let x2 = matched
|
||||
.iter()
|
||||
.map(|i| i.x + i.width)
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
let max_fs = matched.iter().map(|i| i.font_size).fold(0.0, f32::max);
|
||||
let pad = (max_fs * 0.75).max(4.0);
|
||||
if rule.x1 < x1 - pad || rule.x2 > x2 + pad {
|
||||
return false;
|
||||
}
|
||||
let covered: f32 = matched.iter().map(|i| i.width).sum();
|
||||
if covered < rule.width() * 0.6 {
|
||||
return false;
|
||||
}
|
||||
// A table row also unions to the rule's span — but its cells sit apart.
|
||||
// An underlined text line is contiguous runs with word-sized gaps; any
|
||||
// column-sized hole between matched runs means this is a row ruling.
|
||||
let mut sorted = matched;
|
||||
sorted.sort_by(|a, b| a.x.total_cmp(&b.x));
|
||||
sorted.windows(2).all(|pair| {
|
||||
let gap = pair[1].x - (pair[0].x + pair[0].width);
|
||||
gap <= (max_fs * 2.0).max(12.0)
|
||||
})
|
||||
}
|
||||
|
||||
fn is_repeated_ruling_rule(rule: &Rule, rules: &[Rule]) -> bool {
|
||||
let mut y_levels: Vec<f32> = rules
|
||||
.iter()
|
||||
@@ -215,9 +351,11 @@ fn is_underline_candidate(item: &TextItem) -> bool {
|
||||
|
||||
fn rule_matches_item(rule: &Rule, item: &TextItem) -> bool {
|
||||
// Vertical window: underlines sit at or slightly below the baseline.
|
||||
// Fonts draw them at roughly 5-15% of the em below; allow up to 35%
|
||||
// (min 3pt) below and 1pt above for rounding.
|
||||
let below = (item.font_size * 0.35).max(3.0);
|
||||
// Latin fonts draw them at roughly 5-15% of the em below; CJK layouts
|
||||
// put them under the full em box, measured up to ~0.67em below the
|
||||
// baseline (text_dense__underline). Allow 0.72em (min 3pt) below and
|
||||
// 1pt above for rounding.
|
||||
let below = (item.font_size * 0.72).max(3.0);
|
||||
let y_min = item.y - below;
|
||||
let y_max = item.y + 1.0;
|
||||
if rule.y < y_min || rule.y > y_max {
|
||||
@@ -260,12 +398,48 @@ pub(crate) fn mark_underlined_items(
|
||||
lines: &[UnderlineLine],
|
||||
page: u32,
|
||||
) {
|
||||
let rules = discard_repeated_ruling_rules(rules_from_graphics(rects, lines, page));
|
||||
let rules = discard_repeated_ruling_rules(
|
||||
rules_from_graphics(rects, lines, page),
|
||||
items,
|
||||
rects,
|
||||
lines,
|
||||
page,
|
||||
);
|
||||
if rules.is_empty() {
|
||||
return;
|
||||
}
|
||||
let tabular_rules = tabular_row_separator_rule_indices(&rules, items);
|
||||
|
||||
// Math fraction bars are short horizontal lines with the numerator just
|
||||
// above AND the denominator just below — underline geometry from above,
|
||||
// but no underline has text hanging directly beneath it at fraction
|
||||
// distance. Only narrow rules qualify: real underlines under short
|
||||
// labels have their next text line a full line-pitch away.
|
||||
let fraction_rules: HashSet<usize> = rules
|
||||
.iter()
|
||||
.enumerate()
|
||||
.filter(|(_, rule)| {
|
||||
rule.width() <= 60.0
|
||||
&& items.iter().any(|item| {
|
||||
if !is_underline_candidate(item) {
|
||||
return false;
|
||||
}
|
||||
// A denominator HUGS the bar (fraction typesetting
|
||||
// leaves ~0.1-0.2em) and is bar-sized. Both bounds
|
||||
// matter: a short last-line of a paragraph at normal
|
||||
// leading sits further below, and a full next text
|
||||
// line is far wider than the rule.
|
||||
let dy = rule.y - (item.y + item.height);
|
||||
let overlap = rule.x2.min(item.x + item.width) - rule.x1.max(item.x);
|
||||
dy > 0.0
|
||||
&& dy <= item.font_size * 0.3
|
||||
&& overlap > rule.width() * 0.5
|
||||
&& item.width <= rule.width() * 1.5
|
||||
})
|
||||
})
|
||||
.map(|(i, _)| i)
|
||||
.collect();
|
||||
|
||||
for item in items.iter_mut() {
|
||||
if !is_underline_candidate(item) {
|
||||
continue;
|
||||
@@ -275,7 +449,10 @@ pub(crate) fn mark_underlined_items(
|
||||
if tabular_rules.contains(&rule_idx) {
|
||||
continue;
|
||||
}
|
||||
if rule_matches_item(rule, item) {
|
||||
// The fraction guard only gates UNDERLINE marking — a rule that
|
||||
// reads as a fraction bar from below can still legitimately
|
||||
// strike through a line above it.
|
||||
if !fraction_rules.contains(&rule_idx) && rule_matches_item(rule, item) {
|
||||
item.is_underline = true;
|
||||
}
|
||||
if rule_strikes_item(rule, item) {
|
||||
@@ -323,6 +500,16 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
fn cell_rect(x: f32, y: f32, width: f32, height: f32) -> PdfRect {
|
||||
PdfRect {
|
||||
x,
|
||||
y,
|
||||
width,
|
||||
height,
|
||||
page: 1,
|
||||
}
|
||||
}
|
||||
|
||||
fn thin_rect(x: f32, y: f32, width: f32) -> PdfRect {
|
||||
PdfRect {
|
||||
x,
|
||||
@@ -543,6 +730,100 @@ mod tests {
|
||||
assert!(items.iter().all(|item| !item.is_underline));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn repeated_snug_underlines_survive_ruling_filter() {
|
||||
// Dense docs underline many full-width lines: span-similar rules at
|
||||
// 3+ y-levels used to be discarded wholesale as table rulings.
|
||||
// Each rule here snugly matches one text line, so all must mark.
|
||||
let mut items = vec![
|
||||
item("first underlined line of text", 50.0, 700.0, 300.0, 11.0),
|
||||
item("second underlined line here", 50.0, 650.0, 300.0, 11.0),
|
||||
item("third underlined line as well", 50.0, 600.0, 300.0, 11.0),
|
||||
];
|
||||
let lines = vec![
|
||||
hline(50.0, 350.0, 697.0),
|
||||
hline(50.0, 350.0, 647.0),
|
||||
hline(50.0, 350.0, 597.0),
|
||||
];
|
||||
|
||||
mark_underlined_items(&mut items, &[], &lines, 1);
|
||||
|
||||
assert!(items.iter().all(|item| item.is_underline));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn snug_rescue_spans_split_runs_on_one_line() {
|
||||
// A single underlined line is often split into several runs (script
|
||||
// or font switches). The union of touching runs owns the rule.
|
||||
let mut items = vec![
|
||||
item("run one", 50.0, 700.0, 100.0, 11.0),
|
||||
item("run two", 150.5, 700.0, 100.0, 11.0),
|
||||
item("run three", 251.0, 700.0, 99.0, 11.0),
|
||||
item("other a", 50.0, 650.0, 300.0, 11.0),
|
||||
item("other b", 50.0, 600.0, 300.0, 11.0),
|
||||
];
|
||||
let lines = vec![
|
||||
hline(50.0, 350.0, 697.0),
|
||||
hline(50.0, 350.0, 647.0),
|
||||
hline(50.0, 350.0, 597.0),
|
||||
];
|
||||
|
||||
mark_underlined_items(&mut items, &[], &lines, 1);
|
||||
|
||||
assert!(items[0].is_underline && items[1].is_underline && items[2].is_underline);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn snug_rescue_denied_for_row_with_cell_gaps() {
|
||||
// A full-width rule whose baseline row is several items separated by
|
||||
// column-sized gaps is a table row separator, not an underline —
|
||||
// even when span-similar rules repeat down the page.
|
||||
let mut items = vec![
|
||||
item("cell a", 50.0, 700.0, 60.0, 11.0),
|
||||
item("cell b", 190.0, 700.0, 60.0, 11.0),
|
||||
item("cell c", 330.0, 700.0, 70.0, 11.0),
|
||||
item("cell d", 50.0, 650.0, 60.0, 11.0),
|
||||
item("cell e", 190.0, 650.0, 60.0, 11.0),
|
||||
item("cell f", 330.0, 650.0, 70.0, 11.0),
|
||||
];
|
||||
let lines = vec![
|
||||
hline(50.0, 400.0, 697.0),
|
||||
hline(50.0, 400.0, 647.0),
|
||||
hline(50.0, 400.0, 597.0),
|
||||
];
|
||||
|
||||
mark_underlined_items(&mut items, &[], &lines, 1);
|
||||
|
||||
assert!(items.iter().all(|item| !item.is_underline));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn snug_rescue_denied_inside_cell_box() {
|
||||
// A rule snugly under one text line but enclosed by a drawn cell
|
||||
// box that TILES with vertical neighbors (grid evidence) is a row
|
||||
// ruling of a rect-grid table. Isolated boxes (callout panels) do
|
||||
// not veto — see repeated_snug_underlines_survive_ruling_filter.
|
||||
let mut items = vec![
|
||||
item("one wide cell row", 50.0, 700.0, 300.0, 11.0),
|
||||
item("second wide cell", 50.0, 650.0, 300.0, 11.0),
|
||||
item("third wide cell", 50.0, 600.0, 300.0, 11.0),
|
||||
];
|
||||
let lines = vec![
|
||||
hline(50.0, 350.0, 697.0),
|
||||
hline(50.0, 350.0, 647.0),
|
||||
hline(50.0, 350.0, 597.0),
|
||||
];
|
||||
let boxes = vec![
|
||||
cell_rect(45.0, 690.0, 320.0, 50.0),
|
||||
cell_rect(45.0, 640.0, 320.0, 50.0),
|
||||
cell_rect(45.0, 590.0, 320.0, 50.0),
|
||||
];
|
||||
|
||||
mark_underlined_items(&mut items, &boxes, &lines, 1);
|
||||
|
||||
assert!(items.iter().all(|item| !item.is_underline));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn same_row_spaced_rule_segments_do_not_mark_column_labels() {
|
||||
let mut items = vec![
|
||||
|
||||
+170
-29
@@ -71,6 +71,18 @@ use tounicode::FontCMaps;
|
||||
/// broken font decoding or mojibake.
|
||||
pub const OCR_REASON_SUSPECTED_GARBLED_TEXT: &str = "suspected_garbled_text";
|
||||
|
||||
/// OCR reason: the page is a scanned image (a full-page raster / image-only
|
||||
/// page) with no usable text layer.
|
||||
pub const OCR_REASON_SCANNED: &str = "scanned";
|
||||
|
||||
/// OCR reason: the page has no extractable text and no image to OCR — blank,
|
||||
/// or content the parser cannot reach.
|
||||
pub const OCR_REASON_NO_TEXT: &str = "no_text";
|
||||
|
||||
/// OCR reason: the page's text is drawn as vector outlines (path operators)
|
||||
/// rather than real text operators, so it cannot be extracted as characters.
|
||||
pub const OCR_REASON_VECTOR_TEXT: &str = "vector_text";
|
||||
|
||||
// =========================================================================
|
||||
// Result type
|
||||
// =========================================================================
|
||||
@@ -125,7 +137,7 @@ pub struct PdfProcessResult {
|
||||
/// .mode(ProcessMode::Analyze)
|
||||
/// .pages([1, 3, 5]);
|
||||
/// ```
|
||||
#[derive(Debug, Clone)]
|
||||
#[derive(Clone)]
|
||||
pub struct PdfOptions {
|
||||
/// How far the pipeline should run (default: [`ProcessMode::Full`]).
|
||||
pub mode: ProcessMode,
|
||||
@@ -135,6 +147,23 @@ pub struct PdfOptions {
|
||||
pub markdown: MarkdownOptions,
|
||||
/// Optional set of 1-indexed pages to process. `None` = all pages.
|
||||
pub page_filter: Option<HashSet<u32>>,
|
||||
/// Password for decrypting an encrypted PDF. `None` falls back to the
|
||||
/// empty password (owner-only encryption).
|
||||
pub password: Option<String>,
|
||||
}
|
||||
|
||||
// Manual `Debug` so the password is never leaked through debug logging or a
|
||||
// panic that formats the options; it renders as `Some("[REDACTED]")`.
|
||||
impl std::fmt::Debug for PdfOptions {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
f.debug_struct("PdfOptions")
|
||||
.field("mode", &self.mode)
|
||||
.field("detection", &self.detection)
|
||||
.field("markdown", &self.markdown)
|
||||
.field("page_filter", &self.page_filter)
|
||||
.field("password", &self.password.as_ref().map(|_| "[REDACTED]"))
|
||||
.finish()
|
||||
}
|
||||
}
|
||||
|
||||
impl Default for PdfOptions {
|
||||
@@ -144,6 +173,7 @@ impl Default for PdfOptions {
|
||||
detection: DetectionConfig::default(),
|
||||
markdown: MarkdownOptions::default(),
|
||||
page_filter: None,
|
||||
password: None,
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -185,6 +215,12 @@ impl PdfOptions {
|
||||
self.page_filter = Some(pages.into_iter().collect());
|
||||
self
|
||||
}
|
||||
|
||||
/// Set the password used to decrypt an encrypted PDF.
|
||||
pub fn password(mut self, password: impl Into<String>) -> Self {
|
||||
self.password = Some(password.into());
|
||||
self
|
||||
}
|
||||
}
|
||||
|
||||
// =========================================================================
|
||||
@@ -217,7 +253,8 @@ pub fn process_pdf_with_options<P: AsRef<Path>>(
|
||||
validate_pdf_file(&path)?;
|
||||
|
||||
// Load the document once — shared by detection AND extraction.
|
||||
let (doc, page_count) = load_document_from_path(&path)?;
|
||||
let (doc, page_count) =
|
||||
load_document_from_path_with_password(&path, options.password.as_deref())?;
|
||||
|
||||
process_document(doc, page_count, options, start)
|
||||
}
|
||||
@@ -242,7 +279,8 @@ pub fn process_pdf_mem_with_options(
|
||||
let start = std::time::Instant::now();
|
||||
validate_pdf_bytes(buffer)?;
|
||||
|
||||
let (doc, page_count) = load_document_from_mem(buffer)?;
|
||||
let (doc, page_count) =
|
||||
load_document_from_mem_with_password(buffer, options.password.as_deref())?;
|
||||
|
||||
process_document(doc, page_count, options, start)
|
||||
}
|
||||
@@ -635,18 +673,51 @@ pub fn extract_text_in_regions_mem(
|
||||
|
||||
let mut page_results = Vec::with_capacity(regions.len());
|
||||
|
||||
for rect in regions {
|
||||
let [rx1, ry1, rx2, ry2] = *rect;
|
||||
// Exclusive item->region assignment: overlapping layout regions used
|
||||
// to extract shared items into EVERY region they touched (the
|
||||
// 1.5pt inclusion margin makes borders generous), duplicating whole
|
||||
// lines in the final markdown on 21% of bench docs — and downstream
|
||||
// duplicate-handling sometimes dropped the variant holding a
|
||||
// sentence tail, turning duplication into content LOSS. Each item
|
||||
// now belongs to the single region with the largest overlap area;
|
||||
// items are partitioned, never suppressed, so no content can vanish.
|
||||
let all_bounds: Vec<RegionBounds> = regions
|
||||
.iter()
|
||||
.map(|rect| {
|
||||
let [rx1, ry1, rx2, ry2] = *rect;
|
||||
region_bounds(rx1, ry1, rx2, ry2, page_h, coords)
|
||||
})
|
||||
.collect();
|
||||
// Single pass over items: assign each to the best-overlap region and
|
||||
// bucket the clone directly (review: avoid a second O(items x
|
||||
// regions) traversal). `had_candidates` marks regions that touched
|
||||
// at least one item even if every one was assigned elsewhere.
|
||||
let mut region_items: Vec<Vec<TextItem>> = vec![Vec::new(); regions.len()];
|
||||
let mut had_candidates: Vec<bool> = vec![false; regions.len()];
|
||||
if let Some(items) = items {
|
||||
for item in items {
|
||||
let mut best: Option<usize> = None;
|
||||
let mut best_area = 0.0_f32;
|
||||
for (ri, b) in all_bounds.iter().enumerate() {
|
||||
if !region_overlaps_item(item, *b) {
|
||||
continue;
|
||||
}
|
||||
had_candidates[ri] = true;
|
||||
let area = region_item_overlap_area(item, *b);
|
||||
if area > best_area {
|
||||
best_area = area;
|
||||
best = Some(ri);
|
||||
}
|
||||
}
|
||||
if let Some(ri) = best {
|
||||
region_items[ri].push(item.clone());
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
let bounds = region_bounds(rx1, ry1, rx2, ry2, page_h, coords);
|
||||
let matched: Vec<TextItem> = match items {
|
||||
Some(items) => items
|
||||
.iter()
|
||||
.filter(|item| region_overlaps_item(item, bounds))
|
||||
.cloned()
|
||||
.collect(),
|
||||
None => Vec::new(),
|
||||
};
|
||||
for (region_idx, _rect) in regions.iter().enumerate() {
|
||||
let matched: Vec<TextItem> = std::mem::take(&mut region_items[region_idx]);
|
||||
let assigned_count = matched.len();
|
||||
let has_text_quality_issue = region_items_have_decoding_issue(&matched);
|
||||
let text = collect_text_from_matched_items(matched, adaptive_threshold);
|
||||
let has_cid_issue = is_cid_garbage(&text);
|
||||
@@ -660,8 +731,21 @@ pub fn extract_text_in_regions_mem(
|
||||
// Check per-region text quality instead of blanket page-level
|
||||
// GID rejection. A GID font in a logo elsewhere on the page
|
||||
// shouldn't force GPU OCR for clean text regions.
|
||||
let needs_ocr =
|
||||
ocr_reason.is_some() || text.trim().is_empty() || is_garbage_text(&text);
|
||||
// A region whose ONLY overlapping items were assigned to a
|
||||
// better-overlapping neighbor must not fall back to OCR: the
|
||||
// pixels it would re-read belong to that neighbor, and OCR
|
||||
// would reintroduce the duplication exclusivity removed.
|
||||
// Before exclusive assignment these regions were non-empty
|
||||
// native (no OCR), so this preserves the old OCR load too.
|
||||
// Requires ZERO items assigned HERE: a region whose own
|
||||
// assigned items materialize to empty text (whitespace-only,
|
||||
// collector-filtered) keeps its OCR fallback.
|
||||
let lost_to_neighbor = text.trim().is_empty()
|
||||
&& ocr_reason.is_none()
|
||||
&& assigned_count == 0
|
||||
&& had_candidates[region_idx];
|
||||
let needs_ocr = !lost_to_neighbor
|
||||
&& (ocr_reason.is_some() || text.trim().is_empty() || is_garbage_text(&text));
|
||||
|
||||
page_results.push(RegionText {
|
||||
text,
|
||||
@@ -3177,8 +3261,26 @@ fn region_bounds(
|
||||
}
|
||||
}
|
||||
|
||||
/// Inclusion margin shared by the region/item overlap predicates and the
|
||||
/// exclusive-assignment area score — these MUST stay in sync: an item that
|
||||
/// passes the boolean guard must always have positive overlap area.
|
||||
const REGION_MARGIN: f32 = 1.5;
|
||||
|
||||
/// Overlap area between an item and region bounds (same margin as the
|
||||
/// boolean test) — the exclusive-assignment score.
|
||||
fn region_item_overlap_area(item: &TextItem, bounds: RegionBounds) -> f32 {
|
||||
let item_x_max = item.x + text_utils::effective_width(item);
|
||||
let item_y_max = item.y + item.height;
|
||||
let x_overlap = (item_x_max.min(bounds.x_max + REGION_MARGIN)
|
||||
- item.x.max(bounds.x_min - REGION_MARGIN))
|
||||
.max(0.0);
|
||||
let y_overlap = (item_y_max.min(bounds.y_max + REGION_MARGIN)
|
||||
- item.y.max(bounds.y_min - REGION_MARGIN))
|
||||
.max(0.0);
|
||||
x_overlap * y_overlap
|
||||
}
|
||||
|
||||
fn region_overlaps_item(item: &TextItem, bounds: RegionBounds) -> bool {
|
||||
const REGION_MARGIN: f32 = 1.5;
|
||||
let item_x_min = item.x;
|
||||
let item_x_max = item.x + text_utils::effective_width(item);
|
||||
let item_y_min = item.y;
|
||||
@@ -3194,7 +3296,6 @@ fn region_overlaps_item(item: &TextItem, bounds: RegionBounds) -> bool {
|
||||
}
|
||||
|
||||
fn region_overlaps_rect(rect: &PdfRect, bounds: RegionBounds) -> bool {
|
||||
const REGION_MARGIN: f32 = 1.5;
|
||||
let (x_min, y_min, x_max, y_max) = normalized_rect_edges(rect);
|
||||
ranges_overlap(
|
||||
x_min,
|
||||
@@ -3210,7 +3311,6 @@ fn region_overlaps_rect(rect: &PdfRect, bounds: RegionBounds) -> bool {
|
||||
}
|
||||
|
||||
fn region_overlaps_line(line: &PdfLine, bounds: RegionBounds) -> bool {
|
||||
const REGION_MARGIN: f32 = 1.5;
|
||||
let x_min = line.x1.min(line.x2);
|
||||
let x_max = line.x1.max(line.x2);
|
||||
let y_min = line.y1.min(line.y2);
|
||||
@@ -3267,24 +3367,40 @@ fn tsr_region_contains_item(item: &TextItem, bounds: RegionBounds) -> bool {
|
||||
/// page count from it directly to avoid the metadata-only round-trip.
|
||||
pub(crate) fn load_document_from_path<P: AsRef<Path>>(
|
||||
path: P,
|
||||
) -> Result<(Document, u32), PdfError> {
|
||||
load_document_from_path_with_password(path, None)
|
||||
}
|
||||
|
||||
/// Load a PDF file, decrypting with `password` if the file is encrypted.
|
||||
pub(crate) fn load_document_from_path_with_password<P: AsRef<Path>>(
|
||||
path: P,
|
||||
password: Option<&str>,
|
||||
) -> Result<(Document, u32), PdfError> {
|
||||
let buffer = std::fs::read(&path)?;
|
||||
load_document_from_mem(&buffer)
|
||||
load_document_from_mem_with_password(&buffer, password)
|
||||
}
|
||||
|
||||
/// Load a PDF from a memory buffer.
|
||||
pub(crate) fn load_document_from_mem(buffer: &[u8]) -> Result<(Document, u32), PdfError> {
|
||||
load_document_from_mem_with_password(buffer, None)
|
||||
}
|
||||
|
||||
/// Load a PDF from a memory buffer, decrypting with `password` if encrypted.
|
||||
pub(crate) fn load_document_from_mem_with_password(
|
||||
buffer: &[u8],
|
||||
password: Option<&str>,
|
||||
) -> Result<(Document, u32), PdfError> {
|
||||
// Fix malformed struct element names before parsing. Some PDF generators
|
||||
// write bare names (/S Code) instead of proper PDF names (/S /Code), which
|
||||
// causes lopdf to silently drop the entire object.
|
||||
let fixed = structure_tree::fix_bare_struct_names(buffer);
|
||||
let buf = fixed.as_ref();
|
||||
|
||||
let doc = match load_document_bytes(buf) {
|
||||
let doc = match load_document_bytes(buf, password) {
|
||||
Ok(doc) => doc,
|
||||
Err(first_err) => {
|
||||
for repaired in repair_pdf_container_candidates(buf) {
|
||||
match load_document_bytes(&repaired) {
|
||||
match load_document_bytes(&repaired, password) {
|
||||
Ok(doc) => {
|
||||
log::debug!("loaded PDF after repairing malformed container bytes");
|
||||
let page_count = doc.get_pages().len() as u32;
|
||||
@@ -3304,16 +3420,34 @@ pub(crate) fn load_document_from_mem(buffer: &[u8]) -> Result<(Document, u32), P
|
||||
Ok((doc, page_count))
|
||||
}
|
||||
|
||||
fn load_document_bytes(buf: &[u8]) -> Result<Document, lopdf::Error> {
|
||||
fn load_document_bytes(buf: &[u8], password: Option<&str>) -> Result<Document, lopdf::Error> {
|
||||
match Document::load_mem(buf) {
|
||||
// Some encrypted PDFs load structurally but leave their streams
|
||||
// encrypted (`is_encrypted()` stays true); reading them yields garbage
|
||||
// until we re-load with a password. Others fail load_mem outright with
|
||||
// an encryption error. Handle both by re-loading with the password.
|
||||
Ok(doc) if doc.is_encrypted() => decrypt_document_bytes(buf, password),
|
||||
Ok(doc) => Ok(doc),
|
||||
Err(ref e) if is_encrypted_lopdf_error(e) => {
|
||||
Document::load_mem_with_options(buf, lopdf::LoadOptions::with_password(""))
|
||||
}
|
||||
Err(ref e) if is_encrypted_lopdf_error(e) => decrypt_document_bytes(buf, password),
|
||||
Err(e) => Err(e),
|
||||
}
|
||||
}
|
||||
|
||||
/// Re-load an encrypted PDF, decrypting with `password`. Falls back to the
|
||||
/// empty password (owner-only encryption, the common "protected" case) when a
|
||||
/// non-empty password was supplied but rejected.
|
||||
fn decrypt_document_bytes(buf: &[u8], password: Option<&str>) -> Result<Document, lopdf::Error> {
|
||||
let pw = password.unwrap_or("");
|
||||
match Document::load_mem_with_options(buf, lopdf::LoadOptions::with_password(pw)) {
|
||||
Ok(doc) => Ok(doc),
|
||||
Err(inner) if !pw.is_empty() => {
|
||||
Document::load_mem_with_options(buf, lopdf::LoadOptions::with_password(""))
|
||||
.map_err(|_| inner)
|
||||
}
|
||||
Err(inner) => Err(inner),
|
||||
}
|
||||
}
|
||||
|
||||
fn repair_pdf_container_candidates(buf: &[u8]) -> Vec<Vec<u8>> {
|
||||
let mut candidates = Vec::new();
|
||||
|
||||
@@ -3405,6 +3539,7 @@ fn process_document(
|
||||
let pages_needing_ocr = detection.pages_needing_ocr;
|
||||
let title = detection.title;
|
||||
let confidence = detection.confidence;
|
||||
let detection_ocr_reasons = detection.ocr_reasons_by_page;
|
||||
|
||||
// DetectOnly → return immediately
|
||||
if options.mode == ProcessMode::DetectOnly {
|
||||
@@ -3414,7 +3549,7 @@ fn process_document(
|
||||
page_count,
|
||||
processing_time_ms: start.elapsed().as_millis() as u64,
|
||||
pages_needing_ocr,
|
||||
ocr_reasons_by_page: Vec::new(),
|
||||
ocr_reasons_by_page: page_ocr_reasons_vec(detection_ocr_reasons),
|
||||
title,
|
||||
confidence,
|
||||
layout: LayoutComplexity::default(),
|
||||
@@ -3430,7 +3565,7 @@ fn process_document(
|
||||
page_count,
|
||||
processing_time_ms: start.elapsed().as_millis() as u64,
|
||||
pages_needing_ocr,
|
||||
ocr_reasons_by_page: Vec::new(),
|
||||
ocr_reasons_by_page: page_ocr_reasons_vec(detection_ocr_reasons),
|
||||
title,
|
||||
confidence,
|
||||
layout: LayoutComplexity::default(),
|
||||
@@ -3696,7 +3831,13 @@ fn process_document(
|
||||
page_count,
|
||||
processing_time_ms: start.elapsed().as_millis() as u64,
|
||||
pages_needing_ocr,
|
||||
ocr_reasons_by_page: page_ocr_reasons_vec(text_quality_reasons_by_page),
|
||||
ocr_reasons_by_page: {
|
||||
// Detector reasons (scanned / no_text / vector_text / garbled) merged
|
||||
// with the markdown-stage garbled detection, deduped per page.
|
||||
let mut merged = detection_ocr_reasons;
|
||||
merge_ocr_reasons(&mut merged, text_quality_reasons_by_page);
|
||||
page_ocr_reasons_vec(merged)
|
||||
},
|
||||
title,
|
||||
confidence,
|
||||
layout,
|
||||
|
||||
@@ -174,6 +174,21 @@ pub(crate) fn is_toc_marker_heading(text: &str) -> bool {
|
||||
pub(crate) fn is_heading_fragment(text: &str) -> bool {
|
||||
let t = text.trim_end();
|
||||
|
||||
// A lowercase-initial one-or-two-word "heading" is a mid-sentence
|
||||
// fragment beside display math ("or inversely", "and therefore") —
|
||||
// real headings that short start uppercase. Measured as spurious
|
||||
// headings on academic docs (fire-pdf ENG-5029 / opendataloader MHS).
|
||||
{
|
||||
let words: Vec<&str> = t.split_whitespace().collect();
|
||||
if words.len() <= 2 {
|
||||
if let Some(first_alpha) = t.chars().find(|c| c.is_alphabetic()) {
|
||||
if first_alpha.is_lowercase() {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn is_equation_number(s: &str) -> bool {
|
||||
s.strip_prefix('(')
|
||||
.and_then(|r| r.strip_suffix(')'))
|
||||
@@ -460,6 +475,10 @@ mod tests {
|
||||
#[test]
|
||||
fn heading_fragments() {
|
||||
// Equation lead-ins: colon ending + inline equation reference
|
||||
assert!(is_heading_fragment("or inversely"));
|
||||
assert!(is_heading_fragment("and therefore"));
|
||||
assert!(!is_heading_fragment("Introduction"));
|
||||
assert!(!is_heading_fragment("iPhone Sales Strategy Overview")); // 4 words, exempt
|
||||
assert!(is_heading_fragment("Rearranging Equation (8) gives:"));
|
||||
// Display-equation neighbours ending in an equation number
|
||||
assert!(is_heading_fragment("S = kB ln W, (2)"));
|
||||
|
||||
@@ -2367,7 +2367,29 @@ fn detect_merged_cluster_table(
|
||||
/// suitable for rect-backed tables where we already know tabular structure exists
|
||||
/// (no need for anti-paragraph safeguards).
|
||||
fn cluster_x_positions(items: &[(usize, &TextItem)], min_threshold: f32) -> Vec<f32> {
|
||||
let mut x_positions: Vec<f32> = items.iter().map(|(_, i)| i.x).collect();
|
||||
// Column edges come from where text STARTS. An item whose left edge hugs
|
||||
// the previous item's right edge on the same line is a continuation run
|
||||
// (style boundary, script change, underline split) — feeding its x-start
|
||||
// in here fabricates a phantom column mid-cell.
|
||||
let mut sorted: Vec<&TextItem> = items.iter().map(|&(_, i)| i).collect();
|
||||
sorted.sort_by(|a, b| a.y.total_cmp(&b.y).then(a.x.total_cmp(&b.x)));
|
||||
let mut x_positions: Vec<f32> = Vec::with_capacity(sorted.len());
|
||||
for (idx, item) in sorted.iter().enumerate() {
|
||||
let is_continuation = idx > 0 && {
|
||||
let prev = sorted[idx - 1];
|
||||
// Style/underline splits leave runs that TOUCH (gap ~0); real
|
||||
// cell boundaries in even the tightest tables keep a visible
|
||||
// gap. 2pt separates the two without eating dense-table columns.
|
||||
// The negative side is bounded too: text overhanging from an
|
||||
// adjacent cell overlaps by far more than italic kerning ever
|
||||
// does, and must still start its own column.
|
||||
let gap = item.x - (prev.x + prev.width);
|
||||
(prev.y - item.y).abs() <= 2.0 && gap < 2.0 && gap > -4.0 && item.x >= prev.x
|
||||
};
|
||||
if !is_continuation {
|
||||
x_positions.push(item.x);
|
||||
}
|
||||
}
|
||||
x_positions.sort_by(|a, b| a.total_cmp(b));
|
||||
|
||||
if x_positions.is_empty() {
|
||||
|
||||
BIN
Binary file not shown.
@@ -1102,6 +1102,7 @@ fn test_pages_needing_ocr_field_accessible() {
|
||||
title: None,
|
||||
ocr_recommended: false,
|
||||
pages_needing_ocr: Vec::new(),
|
||||
ocr_reasons_by_page: std::collections::BTreeMap::new(),
|
||||
};
|
||||
assert!(detection_result.pages_needing_ocr.is_empty());
|
||||
|
||||
@@ -3606,3 +3607,45 @@ fn test_markdown_options_default_has_include_images_false() {
|
||||
let opts = MarkdownOptions::default();
|
||||
assert!(!opts.include_images);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn encrypted_pdf_decrypts_with_correct_password() {
|
||||
let path = "tests/fixtures/encrypted-secret123.pdf";
|
||||
|
||||
// No password: the file is encrypted and can't be read.
|
||||
let no_pw = process_pdf_with_options(path, PdfOptions::new());
|
||||
assert!(
|
||||
matches!(no_pw, Err(PdfError::Encrypted)),
|
||||
"expected Encrypted without a password, got {no_pw:?}"
|
||||
);
|
||||
|
||||
// Wrong password: still rejected.
|
||||
let wrong = process_pdf_with_options(path, PdfOptions::new().password("wrong"));
|
||||
assert!(
|
||||
matches!(wrong, Err(PdfError::Encrypted)),
|
||||
"expected Encrypted with a wrong password, got {wrong:?}"
|
||||
);
|
||||
|
||||
// Correct password: decrypts and extracts real content.
|
||||
let ok = process_pdf_with_options(path, PdfOptions::new().password("secret123"))
|
||||
.expect("correct password should decrypt");
|
||||
let md = ok.markdown.unwrap_or_default();
|
||||
// Assert a stable fixture token so a garbled-but-long extraction (the
|
||||
// encrypted-stream regression this guards) still fails the test.
|
||||
assert!(
|
||||
md.contains("Procurement"),
|
||||
"decrypted markdown should contain the fixture's real text, got {} chars",
|
||||
md.len()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn pdf_options_debug_redacts_password() {
|
||||
let opts = PdfOptions::new().password("secret123");
|
||||
let dbg = format!("{opts:?}");
|
||||
assert!(
|
||||
!dbg.contains("secret123"),
|
||||
"password leaked in Debug: {dbg}"
|
||||
);
|
||||
assert!(dbg.contains("REDACTED"), "expected redaction marker: {dbg}");
|
||||
}
|
||||
|
||||
@@ -72,10 +72,11 @@ Month or shorter period in which tips were received **4** Net tips (lines **1 +
|
||||
|
||||
forms simpler, we would be happy to hear from you. You can write to the Tax Forms Committee, Western Area Distribution Center, Rancho Cordova, CA 95743-0001. **Purpose.—**Use this form to report tips you receive to your employer. This includes cash tips, tips you receive from other employees, and credit card tips. You must report tips every month regardless of your total wages and tips for the year. However, you do not have to report tips to your employer for any month you received less than $20 in tips while working for that employer. Report tips by the 10th day of the month following the month that you receive them. If the 10th day is a Saturday, Sunday, or legal holiday, report tips by the next day that is not a Saturday, Sunday, or legal holiday. See **Pub. 531**, Reporting Tip Income, for more information. You can get additional copies of **Pub. 1244**, Employee’s Daily Record of Tips and Report to Employer, which contains both Forms 4070A and 4070, by calling 1-800-TAX-FORM (1-800-829-3676).
|
||||
|
||||
**Instructions** *(continued)*
|
||||
<u>Instructions (continued)</u>
|
||||
|
||||
**Unreported Tips.—**If you received tips of $20 or more for any month while working for one employer but did not report them to your employer, you must figure and pay social security and Medicare taxes on the unreported tips when you file your tax return. If you have unreported tips, you **must** use Form 1040 and **Form 4137,** Social Security and Medicare Tax on Unreported Tip Income, to report them. You may **not** use Form 1040A or 1040EZ. Employees subject to the Railroad Retirement Tax Act **cannot** use Form 4137 to pay railroad retirement tax on unreported tips. To get railroad retirement credit, you must report tips to your employer. If you do not report tips to your employer as required, you may be charged a penalty of 50% of the social security and Medicare taxes (or railroad retirement tax) due on the unreported tips unless there was reasonable cause for not reporting them. **Additional Information.—**Get **Pub. 531,** Reporting Tip Income, and Form 4137 for more information on tips. If you are an employee of certain large food or beverage establishments, see Pub. 531 for tip allocation rules. **Recordkeeping.—**If you do not keep a daily record of tips, you must keep other reliable proof of the tip income you received. This proof includes copies of restaurant bills and credit card charges that show amounts customers added as tips. Keep your tip income records for as long as the information on them may be needed in the administration of any Internal Revenue law.
|
||||
|
||||
### Instructions (continued)
|
||||
|
||||
Use this space to total your tips for the year
|
||||
|
||||
|
||||
@@ -1,17 +1,8 @@
|
||||
(e) [Reserved]. For further guidance, see §1.1563-3T(e)(1). Par. 50. Section 1.1563-3T is added to read as follows:
|
||||
<u>§1.1563-3T Rules for determining stock ownership (temporary)</u>.
|
||||
|
||||
(a) through (d)(2)(iii) [Reserved]. For further guidance, see §1.1563-3(a)
|
||||
through (d)(2)(iii). (iv) <u>Statement</u>. If the application of paragraph (d)(2)(ii) or (iii) of §1.1563-3 does not result in a corporation being treated as a component member of only one controlled group of corporations on a December 31, then such corporation will be treated as a component member of only one such group on such date. Such corporation may elect the group in which it is to be included by including on or with its income tax return a statement entitled, “STATEMENT TO ELECT CONTROLLED GROUP PURSUANT TO §1.1563-3T(d)(2)(iv).” The statement must include--
|
||||
|
||||
(A) A description of each of the controlled groups in which the corporation
|
||||
could be included. The description must include the name and employer identification number of each component member of each such group and the stock ownership of the component members of each such group; and
|
||||
|
||||
(B) The following representation: [INSERT NAME AND EMPLOYER
|
||||
IDENTIFICATION NUMBER OF CORPORATION] ELECTS TO BE TREATED AS A COMPONENT MEMBER OF THE [INSERT DESIGNATION OF GROUP].
|
||||
|
||||
(v) <u>Election</u>-- (A) <u>Election filed</u>. An election filed under paragraph (d)(2)(iv) of
|
||||
this section is irrevocable and effective until paragraph (d)(2)(ii) or (iii) of §1.1563-3 applies or until a change in the stock ownership of the corporation results in
|
||||
||||(e) [Reserved]. For further guidance, see §1.1563-3T(e)(1). Par. 50. Section 1.1563-3T is added to read as follows: §1.1563-3T Rules for determining stock ownership (temporary). (a) through (d)(2)(iii) [Reserved]. For further guidance, see §1.1563-3(a)|
|
||||
|---|---|---|---|
|
||||
||through (d)(2)(iii).|||
|
||||
||(iv)|Statement|. If the application of paragraph (d)(2)(ii) or (iii) of §1.1563-3 does not result in a corporation being treated as a component member of only one controlled group of corporations on a December 31, then such corporation will be treated as a component member of only one such group on such date. Such corporation may elect the group in which it is to be included by including on or with its income tax return a statement entitled, “STATEMENT TO ELECT CONTROLLED GROUP PURSUANT TO §1.1563-3T(d)(2)(iv).” The statement must include-- (A) A description of each of the controlled groups in which the corporation could be included. The description must include the name and employer identification number of each component member of each such group and the stock ownership of the component members of each such group; and (B) The following representation: [INSERT NAME AND EMPLOYER IDENTIFICATION NUMBER OF CORPORATION] ELECTS TO BE TREATED AS A COMPONENT MEMBER OF THE [INSERT DESIGNATION OF GROUP].|
|
||||
||(v)|Election|-- (A) Election filed. An election filed under paragraph (d)(2)(iv) of this section is irrevocable and effective until paragraph (d)(2)(ii) or (iii) of §1.1563-3 applies or until a change in the stock ownership of the corporation results in|
|
||||
|
||||
|termination of membership in the controlled group in which such corporation has||
|
||||
|---|---|
|
||||
@@ -29,7 +20,7 @@ this section is irrevocable and effective until paragraph (d)(2)(ii) or (iii) of
|
||||
Federal income tax return (including any amended return filed on or before the due date (including extensions) of such original return) timely filed on or after May 30,
|
||||
|
||||
2006.
|
||||
(2) Expiration date. The applicability of this section will expire on May 26,
|
||||
(2) <u>Expiration date</u>. The applicability of this section will expire on May 26,
|
||||
2009. Par. 51. Section 1.6012-2 is amended by revising paragraph (c) and adding paragraph (k) to read as follows: <u>§1.6012-2 Corporations required to make returns of income</u>.
|
||||
* * * * *
|
||||
(c) [Reserved]. For further guidance, see §1.6012-2T(c).
|
||||
|
||||
@@ -43,9 +43,9 @@ l
|
||||
|
||||
**Freon** **®** **12 Saturation Properties-Temperature Table**
|
||||
|
||||
|Temp|Pressure||Volume|||Density||Enthalpy|||Entropy|Temp|
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
|°C|[kPa]|[m³ Liquid v f|/kg]|Vapour v g|Liquid d f|[kg/m³] Vapour d g|Liquid H f|[kJ/kg] Latent H fg|Vapour H g|Liquid S f|[kJ/K-kg] Vapour S g|°C|
|
||||
|Temp|Pressure||Volume||Density||Enthalpy|||Entropy|Temp|
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
|°C|[kPa]|[m³ Liquid v f|/kg] Vapour v g|[kg/m³ Liquid d f|] Vapour d g|Liquid H f|[kJ/kg] Latent H fg|Vapour H g|Liquid S f|[kJ/K-kg] Vapour S g|°C|
|
||||
|
||||
|-100|1.2|0.0006|10.0000|1679.0|0.100|113.3|192.8|306.1|0.6077|1.7210|-100|
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
|
||||
Reference in New Issue
Block a user