Compare commits

...
Author SHA1 Message Date
Abimael Martell 7bfa25bdc6 test(markdown): make the German suspension test exercise the length floor
Review finding: 'klein' appeared only at the break, so vocabulary scrubbing
removed it and the compound rule could never fire — the test passed
trivially. It now includes a standalone 'klein', so only the length floor
keeps 'klein-und' from fusing; verified the test fails when the floor is
lowered.
2026-08-14 15:03:59 -07:00
Abimael Martell 3696605d26 refactor(markdown): replace SUSPENSION_WORDS with a length floor
Same critique as the compound-suffix list, same answer: 'to' and 'or' were
provably dead entries (two-letter words can never enter the 3+-letter
vocabulary, so nothing downstream would join them), and 'and' guarded
exactly one flaw in the both-words rule — while German 'und', Finnish
'tai', French 'et' got no protection at all, and the both-words rule was
silently fusing their suspended hyphens ('vor- und' -> 'vor-und',
'kuva- tai' -> 'kuva-tai').

The generic replacement: the both-words rule now requires a continuation of
four or more letters. Conjunctions are one to three letters in essentially
every language, so suspended hyphens fall through to Keep with no word list
in any language.

Corpus effect (27 of 203 files): German and Finnish suspended hyphens are
fixed (they were silent wrong fusions); English three-letter compound
continuations ('factory- set', 'stand- off') revert to visible breaks when
the document offers no hyphenated evidence — the documented no-evidence
floor. Known remaining limit: four-letter conjunctions (German 'oder')
still pass the floor; that miss predates this change.

The policy now contains zero hard-coded words: vocabulary evidence,
capitalization, and two length facts.
2026-08-14 14:57:02 -07:00
Abimael Martell ed830a9e6a refactor(markdown): drop the hard-coded compound-suffix list
The list ('based', 'class', 'type', ...) was seeded from two observed
cases and grew by intuition: English-only in an otherwise script-agnostic
design, unfalsifiable membership, and able to invent hyphens
('proto- type' -> 'proto-type'). It was added to prevent corruption under
the old unconditional default, and outlived the default it patched — under
evidence-only joining its entire corpus footprint is 8 cosmetic joins
across 7 documents, all of which now revert to honest visible breaks (the
documented no-evidence floor).

Every remaining rule is document evidence or script-agnostic typography.
A real hyphenation dictionary (e.g. TeX patterns) as an opt-in feature is
the principled path if those cases ever matter.
2026-08-14 13:54:24 -07:00
Abimael Martell 7c52a613a3 fix(markdown): unicode word class, stable-loop joins, pre-collapse spaces
Three review findings on #388, all applied:

- WORD is now \p{L} instead of a hard-coded Latin-plus-accents subset, so
  evidence-backed breaks join in any script (Cyrillic, Turkish, and the
  full German alphabet included — the old subset lacked 'ß', which made the
  regex match accidental sub-fragments of German words).
- The per-line pass loops until stable instead of capping at three passes;
  every successful join removes a break, so the loop is bounded.
- collapse_consecutive_spaces runs before the hyphenation passes, so a
  double-spaced break ('de-  fendant') still rejoins. It also feeds the
  pre-existing spaced-hyphen rule cleaner input.

Fixture snapshots are byte-stable. Corpus effect measured against the
previous revision: 24 documents change by a few bytes each, dominated by
newly-joined Cyrillic and Turkish breaks; every join is evidence-backed.
2026-08-14 13:40:06 -07:00
Abimael Martell 203d1f944d refactor(markdown): make the join decision an explicit type
Review finding: join_decision returned rendered strings and the emphasis
handler reverse-inferred the decision from the string shape. It now returns
Join::{Plain,Hyphen,Keep} and each call site renders it — the emphasis path
preserves its own markers on Keep instead of reconstructing them.

Behavior-identical: fixture snapshots byte-stable, 1,109 tests pass.
2026-08-14 13:19:40 -07:00
Abimael Martell a82ca82c98 refactor(markdown): exclude code blocks from dehyphenation vocabulary
Two review findings:

- Fenced code blocks no longer contribute vocabulary evidence: a fused
  identifier like 'comreal' must not justify joining unrelated prose.
  Table rows deliberately stay — cells carry genuine document vocabulary,
  and excluding them would cost real joins to guard a contrived case.
- The join policy moved out of the closure into join_decision(), so the
  decision rules are separated from the Markdown scanning around them.

Fixture snapshots are byte-identical; 1,109 tests pass.
2026-08-14 13:15:29 -07:00
Abimael Martell 06daa34353 fix(markdown): join breaks only on evidence, never by default
Second local-review finding: the unconditional lowercase fallback fused
interleaved-column fragments ('com- real' -> 'comreal') that slipped under
the length gate. Measured on a 1,370-page justified document, the fallback
contributed only ~1% of joins (94 of ~8,000 — the vocabulary rules do the
rest) while being the sole rule able to corrupt output. The no-evidence case
now leaves the break untouched: a visible break is honest, a silent fusion
is not.

Word recall on that document is unchanged at 99.1%. Tests updated to supply
vocabulary evidence where joins are expected, plus a new no-evidence case;
fixture snapshots regenerated.
2026-08-14 13:10:23 -07:00
Abimael Martell e89b016127 fix(markdown): length-gate dehyphenation against fused column noise
Local review finding: interleaved-column reading-order noise arrives already
fused into long pseudo-words, and joining across its breaks compounded the
damage. Real syllable fragments are short — a combined 40-character cap
blocks the fused case while still admitting long German compounds
(Bundesausbildungsförderungs- gesetz joins; fused column text does not).

Regenerated the affected fixture snapshot; 1,108 tests pass.
2026-08-14 13:04:12 -07:00
Abimael Martell 5ebb690170 feat(markdown): rejoin words hyphenated at line breaks
Justified print breaks words at syllables; after paragraph lines are joined
with spaces those breaks survive as "de- fendant" — thousands of them in a
long document — and, when an emphasis span was split with the word, as
"Bap-** **tist".

Whether the hyphen itself belongs in the word cannot be decided locally
("de- fendant" is one word, "Third- Party" is a hyphenated compound), so the
document is used as its own dictionary. For each break:

  1. fragments appear joined elsewhere ("defendant")      -> join plain
  2. appear hyphenated elsewhere ("six-month")            -> keep the hyphen
  3. capitalized continuation ("Hinds- Radix")            -> keep the hyphen
  4. continuation is to/and/or ("mid- to long-term")      -> suspended
     hyphen, leave untouched
  5. both fragments are words the document uses
     ("commercial- type")                                 -> keep the hyphen
  6. default (syllable breaks dominate)                   -> join plain

The vocabulary is collected after scrubbing the break patterns themselves,
otherwise every broken word donates its own fragments ("evi", "dence") and
rule 5 misfires. Split emphasis spans rejoin inside the markers
("**Bap-** **tist**" -> "**Baptist**") when the markers match. Table rows
and fenced code blocks are left untouched.

Runs under the existing fix_hyphenation option (default on). On a 1,370-page
legal reporter this rejoins ~8,000 broken words with a single deliberate
survivor (a genuine suspended hyphen); word recall against a reference
extraction rises from 97.8% to 99.1%, and no "six-month" -> "sixmonth" class
errors are introduced (a failure mode common in extractors that strip
line-break hyphens unconditionally).

Regression-checked against a ~200-document corpus with semantic scoring
against an OCR baseline. Three in-repo fixture snapshots regenerated; each
diff inspected (syllable joins, kept compounds).

Adds 11 unit tests covering every rule, the vocabulary scrubbing, chained
breaks, mismatched emphasis markers, accented words, and the table/code
skips.
2026-08-14 13:00:51 -07:00
3 changed files with 391 additions and 7 deletions
+386 -2
View File
@@ -13,7 +13,11 @@ pub(crate) fn clean_markdown(mut text: String, options: &MarkdownOptions) -> Str
text = collapse_dot_leaders(&text);
}
// Fix hyphenation first (before other processing)
// Collapse runs of spaces first: double-spaced breaks ("de- fendant")
// must look like single-spaced ones before the hyphenation passes.
collapse_consecutive_spaces(&mut text);
// Fix hyphenation (before other processing)
if options.fix_hyphenation {
text = fix_hyphenation(&text);
}
@@ -143,7 +147,213 @@ fn fix_hyphenation(text: &str) -> String {
})
.to_string();
result
dehyphenate_line_breaks(&result)
}
/// What a line-break hyphen pair should become. Policy output only — how the
/// decision is rendered (plain text vs. inside split emphasis markers) is the
/// caller's business.
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
enum Join {
/// The fragments are one word: "de- fendant" -> "defendant".
Plain,
/// The fragments are a hyphenated compound: "six- month" -> "six-month".
Hyphen,
/// No evidence (or a suspended hyphen): leave the break as it is.
Keep,
}
/// Decide what a line-break hyphen pair becomes, given the document's
/// vocabulary evidence. The whole dehyphenation policy lives here so it can
/// be reasoned about (and tested) apart from the Markdown scanning around it;
/// see [`dehyphenate_line_breaks`] for the rule rationale.
fn join_decision(
a: &str,
b: &str,
words: &std::collections::HashSet<String>,
hyphenated: &std::collections::HashSet<String>,
) -> Join {
// Word-length sanity: syllable fragments are short, and even long German
// compounds stay under this. Fragments beyond it are already fused
// reading-order noise (interleaved columns); joining would compound the
// damage.
if a.chars().count() + b.chars().count() > 40 {
return Join::Keep;
}
let key_plain = format!("{}{}", a.to_lowercase(), b.to_lowercase());
let key_hyphen = format!("{}-{}", a.to_lowercase(), b.to_lowercase());
if words.contains(&key_plain) {
Join::Plain
} else if hyphenated.contains(&key_hyphen) || b.chars().next().is_some_and(|c| c.is_uppercase())
{
Join::Hyphen
} else if b.chars().count() >= 4
&& words.contains(&a.to_lowercase())
&& words.contains(&b.to_lowercase())
{
// No direct evidence, but both fragments are themselves words the
// document uses ("commercial- type"): hyphenated compounds are made
// of words, while syllable fragments ("evi", "judg", "mo") are not.
//
// The continuation must be at least four letters. Suspended hyphens
// ("mid- and long-term", "klein- und mittelgroß") put a conjunction
// after the hyphen, and conjunctions are near-universally one to
// three letters in any language — the length floor keeps this rule
// off them without a hard-coded conjunction list.
Join::Hyphen
} else {
// No evidence at all: leave the break as it is. An unconditional join
// here covered only ~1% more breaks on a vocabulary-rich document,
// but it was the sole rule able to corrupt output — fusing
// interleaved-column fragments ("com- real" -> "comreal") into
// unrecoverable tokens. A visible break is honest; a silent fusion
// is not.
Join::Keep
}
}
/// Rejoin words hyphenated at the original line breaks.
///
/// Justified print breaks words at syllables; after paragraph lines are
/// joined with spaces those breaks survive as "de- fendant" (and, when an
/// emphasis span was split with the word, "Bap-** **tist"). Whether the
/// hyphen itself belongs in the word cannot be decided locally — "de-
/// fendant" is one word but "Third- Party" is a hyphenated compound — so the
/// document is its own dictionary:
///
/// 1. fragments appear elsewhere joined plain ("defendant") — join plain;
/// 2. appear elsewhere hyphenated ("six-month"), or the continuation is
/// capitalized ("Hinds- Radix", "Third- Party") — keep the hyphen;
/// 3. both fragments are words the document uses and the continuation
/// has four or more letters ("commercial- type" where "commercial"
/// and "type" appear elsewhere) — a compound, keep the hyphen. The
/// length floor keeps this rule off suspended hyphens ("mid- and
/// long-term", "klein- und mittelgroß"): conjunctions are one to
/// three letters in essentially every language, so no conjunction
/// list is needed;
/// 4. no evidence — leave the break untouched. Evidence covers ~99% of
/// breaks on vocabulary-rich documents, and an unconditional join was
/// the one rule able to corrupt output (fusing interleaved-column
/// fragments into unrecoverable tokens).
///
/// Every rule is either document evidence or script-agnostic typography;
/// deliberately no hard-coded word lists beyond the three suspension
/// conjunctions (a curated suffix list was tried and removed — it was
/// English-only, its membership was unfalsifiable, and it could invent
/// hyphens: "proto- type" -> "proto-type").
///
/// Table rows and fenced code blocks are left untouched.
fn dehyphenate_line_breaks(text: &str) -> String {
use once_cell::sync::Lazy;
use std::collections::HashSet;
const WORD: &str = r"\p{L}";
// "de- fendant"
static BREAK_RE: Lazy<Regex> =
Lazy::new(|| Regex::new(&format!("({WORD}{{2,}})- ({WORD}{{2,}})")).unwrap());
// "Bap-** **tist" — an emphasis span split together with the word. The
// regex crate has no backreferences, so both markers are captured and
// compared in the replacement closure.
static BREAK_EMPH_RE: Lazy<Regex> = Lazy::new(|| {
Regex::new(&format!(
r"({WORD}{{2,}})-(\*{{1,2}}) (\*{{1,2}})({WORD}{{2,}})"
))
.unwrap()
});
static PLAIN_WORD_RE: Lazy<Regex> = Lazy::new(|| Regex::new(&format!("{WORD}{{3,}}")).unwrap());
static HYPHENATED_RE: Lazy<Regex> =
Lazy::new(|| Regex::new(&format!("({WORD}{{2,}})-({WORD}{{2,}})")).unwrap());
// The document is its own dictionary: whole words and hyphenated
// compounds as they appear away from line breaks. Two exclusions keep the
// evidence sound:
// - the break pairs themselves are scrubbed first, otherwise every
// broken word donates its own fragments ("evi", "dence") and the
// compound rule would see them as words;
// - fenced code blocks are skipped: identifiers are a different
// language, and one like `comreal` must not justify fusing prose.
// Table rows stay — cells carry genuine document vocabulary.
let mut prose = String::with_capacity(text.len());
let mut in_code = false;
for line in text.split('\n') {
if line.trim_start().starts_with("```") {
in_code = !in_code;
continue;
}
if !in_code {
prose.push_str(line);
prose.push('\n');
}
}
let scrubbed = BREAK_EMPH_RE.replace_all(&prose, " ");
let scrubbed = BREAK_RE.replace_all(&scrubbed, " ");
let mut words: HashSet<String> = HashSet::new();
let mut hyphenated: HashSet<String> = HashSet::new();
for m in PLAIN_WORD_RE.find_iter(&scrubbed) {
words.insert(m.as_str().to_lowercase());
}
for caps in HYPHENATED_RE.captures_iter(&scrubbed) {
hyphenated.insert(format!(
"{}-{}",
caps[1].to_lowercase(),
caps[2].to_lowercase()
));
}
let join = |a: &str, b: &str| join_decision(a, b, &words, &hyphenated);
let mut in_code_block = false;
let mut out = String::with_capacity(text.len());
for (i, line) in text.split('\n').enumerate() {
if i > 0 {
out.push('\n');
}
if line.trim_start().starts_with("```") {
in_code_block = !in_code_block;
}
if in_code_block || line.trim_start().starts_with('|') {
out.push_str(line);
continue;
}
// A break can chain ("unconsti- tu- tional"); each pass joins one
// junction, and every successful join removes a break, so running
// until the line is stable is bounded by the number of breaks.
let mut current = line.to_string();
loop {
let next = BREAK_EMPH_RE
.replace_all(&current, |caps: &regex::Captures| {
// Mismatched markers aren't a split span; a Keep decision
// preserves the original spacing and markers.
if caps[2] != caps[3] {
return caps[0].to_string();
}
let (a, b) = (&caps[1], &caps[4]);
match join(a, b) {
Join::Plain => format!("{a}{b}"),
Join::Hyphen => format!("{a}-{b}"),
Join::Keep => caps[0].to_string(),
}
})
.to_string();
let next = BREAK_RE
.replace_all(&next, |caps: &regex::Captures| {
let (a, b) = (&caps[1], &caps[2]);
match join(a, b) {
Join::Plain => format!("{a}{b}"),
Join::Hyphen => format!("{a}-{b}"),
Join::Keep => caps[0].to_string(),
}
})
.to_string();
if next == current {
break;
}
current = next;
}
out.push_str(&current);
}
out
}
/// Remove isolated page-number expressions from Markdown.
@@ -384,6 +594,180 @@ mod tests {
assert_eq!(t, "version 3 .14 released");
}
// --- dehyphenate_line_breaks ---
#[test]
fn line_break_joins_plain_on_vocabulary_evidence() {
// "defendant" appears whole elsewhere, so the broken form joins plain.
let text = "The defendant appeared. The de- fendant argued.";
assert_eq!(
dehyphenate_line_breaks(text),
"The defendant appeared. The defendant argued."
);
}
#[test]
fn line_break_keeps_hyphen_on_hyphenated_evidence() {
// "six-month" appears hyphenated elsewhere, so the broken form keeps it.
let text = "A six-month term. After a six- month delay.";
assert_eq!(
dehyphenate_line_breaks(text),
"A six-month term. After a six-month delay."
);
}
#[test]
fn no_evidence_leaves_the_break_untouched() {
// Without document evidence a join cannot be distinguished from
// interleaved-column noise; the visible break is kept.
let text = "The evi- dence was clear.";
assert_eq!(dehyphenate_line_breaks(text), text);
}
#[test]
fn line_break_keeps_hyphen_before_capitalized_continuation() {
// Broken compounds: "Third-Party", "Hinds-Radix".
let text = "The Third- Party complaint by Hinds- Radix.";
assert_eq!(
dehyphenate_line_breaks(text),
"The Third-Party complaint by Hinds-Radix."
);
}
#[test]
fn cyrillic_words_join_on_evidence() {
// The word class is Unicode-wide, not a hard-coded Latin subset.
let text = "Это решение важно. Это реше- ние суда.";
assert_eq!(
dehyphenate_line_breaks(text),
"Это решение важно. Это решение суда."
);
}
#[test]
fn double_spaced_breaks_join_through_clean_markdown() {
// Space collapsing runs before hyphenation, so a break that arrives
// with two spaces ("de- fendant") still rejoins.
let options = MarkdownOptions::default();
let out = clean_markdown(
"The defendant appeared. The de- fendant argued.".to_string(),
&options,
);
assert_eq!(
out.trim_end(),
"The defendant appeared. The defendant argued."
);
}
#[test]
fn code_block_identifiers_are_not_vocabulary_evidence() {
// A fused identifier in code must not justify fusing unrelated prose.
let text = "```\nlet comreal = 1;\n```\nThe com- real estate story.";
assert_eq!(dehyphenate_line_breaks(text), text);
}
#[test]
fn fused_column_noise_is_not_joined() {
// Interleaved-column garbage arrives already fused; joining across
// its breaks would compound the damage. Real syllable fragments are
// short; fragments this long are left exactly as they are.
let text = "spreadswerenegativeintheearlytomid- seriouslyflawedduetoappraisallags";
assert_eq!(dehyphenate_line_breaks(text), text);
// Long German compounds stay under the length gate and join on
// vocabulary evidence.
let german =
"Das Bundesausbildungsförderungsgesetz. Das Bundesausbildungsförderungs- gesetz gilt.";
assert_eq!(
dehyphenate_line_breaks(german),
"Das Bundesausbildungsförderungsgesetz. Das Bundesausbildungsförderungsgesetz gilt."
);
}
#[test]
fn no_evidence_compounds_stay_visibly_broken() {
// No hard-coded suffix list: without document evidence even a likely
// compound keeps its visible break. A curated list was tried and
// removed — English-only, unfalsifiable membership, and able to
// invent hyphens ("proto- type" -> "proto-type").
let text = "Their world- class support and proto- type systems.";
assert_eq!(dehyphenate_line_breaks(text), text);
}
#[test]
fn both_fragments_being_words_keeps_the_hyphen() {
// "commercial-type insurance": no evidence either way, but both
// fragments are words the document uses, so this is a compound.
let text = "Any commercial firm of this type offering commercial- type insurance.";
assert_eq!(
dehyphenate_line_breaks(text),
"Any commercial firm of this type offering commercial-type insurance."
);
}
#[test]
fn suspended_hyphen_is_preserved() {
// "mid- to long-term": joining would fuse unrelated words. No
// conjunction list is involved — conjunctions are 1-3 letters in
// essentially every language, and the compound rule requires a
// 4-letter continuation, so suspended hyphens fall through to Keep.
let text = "Planned over the mid- to long-term horizon, in- and out-of-possession.";
assert_eq!(dehyphenate_line_breaks(text), text);
// Same construction in German, which a hard-coded English list
// would have missed. "klein" appears standalone so it IS in the
// vocabulary — only the length floor (continuation "und" has three
// letters) keeps the compound rule from fusing "klein-und".
let german = "Das klein geschriebene Wort und die klein- und mittelgroßen Betriebe.";
assert_eq!(dehyphenate_line_breaks(german), german);
}
#[test]
fn split_emphasis_span_joins_inside_markers() {
// Vocabulary evidence ("Baptist", "Consolidated" elsewhere) drives
// the join; the split emphasis markers collapse with it.
let text = "The Baptist and Consolidated cases. By **Bap-** **tist** pastors and *Consoli-* *dated* Edison.";
assert_eq!(
dehyphenate_line_breaks(text),
"The Baptist and Consolidated cases. By **Baptist** pastors and *Consolidated* Edison."
);
}
#[test]
fn mismatched_emphasis_markers_are_left_alone() {
let text = "Odd **Bap-** *tist* markers.";
assert_eq!(dehyphenate_line_breaks(text), text);
}
#[test]
fn chained_breaks_join_stepwise_with_evidence() {
// A word broken twice joins across passes when each junction has
// vocabulary evidence for its intermediate form.
let text = "The word unconstitutional, and unconstitu appears too: unconsti- tu- tional.";
assert_eq!(
dehyphenate_line_breaks(text),
"The word unconstitutional, and unconstitu appears too: unconstitutional."
);
}
#[test]
fn table_rows_and_code_blocks_are_untouched() {
let text =
"The defendant.\n|de- fendant|value|\n```\nlet x = de- fendant;\n```\nThe de- fendant won.";
assert_eq!(
dehyphenate_line_breaks(text),
"The defendant.\n|de- fendant|value|\n```\nlet x = de- fendant;\n```\nThe defendant won."
);
}
#[test]
fn accented_words_join() {
// Spanish syllable break with accented continuation, evidence-backed.
let text = "Una resolución firme. La resolu- ción fue clara.";
assert_eq!(
dehyphenate_line_breaks(text),
"Una resolución firme. La resolución fue clara."
);
}
// --- fix_hyphenation ---
#[test]
+2 -2
View File
@@ -46,9 +46,9 @@ more about investing in tax losses than burst, cap rates spreads steadily com- r
8 6 Z E L L / L U R I E R E A L E S T A T E C E N T E R
ingdebtspreadswerepartoforiginalpro formamodels.Thiscapratespreadcom- pressionoffsetweakcashflowsinapost- recessionary economy from 2002 to 2005, while continued compression, combined with improved cash flows, pushed property values skyward in 2006 throughmid-2007. Cap rate compression reduced the importance of the ability to add value. After all, if all you had to do to make moneywastoleveragetothehiltwhilecap ratesfell,whytakeontheextraworkand riskofattemptingtoaddvalue?Stateddif- ferently: Why print money if it is laying everywhereonthestreets? In Tables III and IV, we demonstrate thepowerofcapratecompressionviavery simple pro forma cash flow analyses that assume Year 1 NOI of $100; a going-in cap rate of 9 percent; an LTV of 70 per- cent; and an interest rate of 7 percent. Withineachfigure,wedisplaytwoscenar- ios, which vary based on NOI growth assumptions.ScenarioIassumesthatNOI growsby3percentperyear,whileScenario IIassumesavalue-addNOIgrowthof20 percentbetweenyearstwoandthree. The only other difference between TablesIIIandIVisinresidualcaprates, which are assumed to be 6 percent and 9 percent, respectively. Based on these assumptions, we calculate the equity IRRs. It is clear that cap rate compres- sion is a significant factor in driving
ingdebtspreadswerepartoforiginalpro formamodels.Thiscapratespreadcom- pressionoffsetweakcashflowsinapost- recessionary economy from 2002 to 2005, while continued compression, combined with improved cash flows, pushed property values skyward in 2006 throughmid-2007. Cap rate compression reduced the importance of the ability to add value. After all, if all you had to do to make moneywastoleveragetothehiltwhilecap ratesfell,whytakeontheextraworkand riskofattemptingtoaddvalue?Stateddif- ferently: Why print money if it is laying everywhereonthestreets? In Tables III and IV, we demonstrate thepowerofcapratecompressionviavery simple pro forma cash flow analyses that assume Year 1 NOI of $100; a going-in cap rate of 9 percent; an LTV of 70 percent; and an interest rate of 7 percent. Withineachfigure,wedisplaytwoscenar- ios, which vary based on NOI growth assumptions.ScenarioIassumesthatNOI growsby3percentperyear,whileScenario IIassumesavalue-addNOIgrowthof20 percentbetweenyearstwoandthree. The only other difference between TablesIIIandIVisinresidualcaprates, which are assumed to be 6 percent and 9 percent, respectively. Based on these assumptions, we calculate the equity IRRs. It is clear that cap rate compression is a significant factor in driving
returns. That is, cap rate compression from 9 percent to 6 percent increased IRR on leveraged stabilized properties by 250 percent, to a staggering 57 per- cent. Who needs to take on value add riskatthisreturnforstabilizedassets? Intheearly1980s,moneywasmadein real estate by mastering the creation and syndication of tax gimmicks. In the late 1980s, one made money by mastering bank and S&L connections to over-lever- age.Intheearly1990s,onemademoneyin realestatebyhavingaccesstoequity—the morethebetter.Duringthelate1990s,one made money from real estate by realizing large spreads between cap rates and debt costs.And,overthepastfiveyears,theway to make money in real estate was to own realestateonahighlyleveragedbasisascap ratesplunged. Theclassicassetpricingmodelisthe capital asset pricing model (CAPM). CAPM is a simple, yet elegant, model that relates asset pricing to the risk-free rate(F),theabilityofanassettoreduce portfolio variance (B), and the expected rate of return on the market bundle of investableassets(M).CAPMisfarfrom perfect,butprovidesacrudebenchmark for asset pricing, around which discrep- ancies and novelties arise. Specifically, CAPM states that an assets price is set suchthattheexpectedreturnforanasset
returns. That is, cap rate compression from 9 percent to 6 percent increased IRR on leveraged stabilized properties by 250 percent, to a staggering 57 percent. Who needs to take on value add riskatthisreturnforstabilizedassets? Intheearly1980s,moneywasmadein real estate by mastering the creation and syndication of tax gimmicks. In the late 1980s, one made money by mastering bank and S&L connections to over-leverage.Intheearly1990s,onemademoneyin realestatebyhavingaccesstoequity—the morethebetter.Duringthelate1990s,one made money from real estate by realizing large spreads between cap rates and debt costs.And,overthepastfiveyears,theway to make money in real estate was to own realestateonahighlyleveragedbasisascap ratesplunged. Theclassicassetpricingmodelisthe capital asset pricing model (CAPM). CAPM is a simple, yet elegant, model that relates asset pricing to the risk-free rate(F),theabilityofanassettoreduce portfolio variance (B), and the expected rate of return on the market bundle of investableassets(M).CAPMisfarfrom perfect,butprovidesacrudebenchmark for asset pricing, around which discrep- ancies and novelties arise. Specifically, CAPM states that an assets price is set suchthattheexpectedreturnforanasset
(R)is R=F+ β(M-F).
R E V I E W 8 7
+3 -3
View File
@@ -14,7 +14,7 @@ basis for such a theory is contained in the important papers of Nyquist¹ and Ha
1. It is practically more useful. Parameters of engineering importance such as time, bandwidth, number of relays, etc., tend to vary linearly with the logarithm of the number of possibilities. For example, adding one relay to a group doubles the number of possible states of the relays. It adds 1 to the base 2 logarithm of this number. Doubling the time roughly squares the number of possible messages, or doubles the logarithm, etc.
2. It is nearer to our intuitive feeling as to the proper measure. This is closely related to (1) since we in- tuitively measures entities by linear comparison with common standards. One feels, for example, that two punched cards should have twice the capacity of one for information storage, and two identical channels twice the capacity of one for transmitting information.
3. It is mathematically more suitable. Many of the limiting operations are simple in terms of the loga- rithm but would require clumsy restatement in terms of the number of possibilities. The choice of a logarithmic base corresponds to the choice of a unit for measuring information. If the
3. It is mathematically more suitable. Many of the limiting operations are simple in terms of the logarithm but would require clumsy restatement in terms of the number of possibilities. The choice of a logarithmic base corresponds to the choice of a unit for measuring information. If the
base 2 is used the resulting units may be called binary digits, or more briefly *bits,* a word suggested by
J. W. Tukey. A device with two stable positions, such as a relay or a flip-flop circuit, can store one bit of information. *N* such devices can store*N* bits, since the total number of possible states is 2
@@ -34,8 +34,8 @@ Fig. 1 — Schematic diagram of a general communication system.
a decimal digit is about 3 13 bits. A digit wheel on a desk computing machine has ten stable positions and therefore has a storage capacity of one decimal digit. In analytical work where integration and differentiation are involved the base *e* is sometimes useful. The resulting units of information will be called natural units. Change from the base *a* to base *b* merely requires multiplication by log*ba*. By a communication system we will mean a system of the type indicated schematically in Fig. 1. It consists of essentially five parts:
1. An *information source* which produces a message or sequence of messages to be communicated to the receiving terminal. The message may be of various types: (a) A sequence of letters as in a telegraph of teletype system; (b) A single function of time *f* (*t*) as in radio or telephony; (c) A function of time and other variables as in black and white television — here the message may be thought of as a function *f* (*x*; *y*;*t*) of two space coordinates and time, the light intensity at point (*x*; *y*) and time *t* on a pickup tube plate; (d) Two or more functions of time, say *f* (*t*), *g*(*t*), *h*(*t*) — this is the case in “three- dimensional” sound transmission or if the system is intended to service several individual channels in multiplex; (e) Several functions of several variables — in color television the message consists of three functions *f* (*x*; *y*;*t*), *g*(*x*; *y*;*t*), *h*(*x*; *y*;*t*) defined in a three-dimensional continuum — we may also think of these three functions as components of a vector field defined in the region — similarly, several black and white television sources would produce “messages” consisting of a number of functions of three variables; (f) Various combinations also occur, for example in television with an associated audio channel.
2. A *transmitter* which operates on the message in some way to produce a signal suitable for trans- mission over the channel. In telephony this operation consists merely of changing sound pressure into a proportional electrical current. In telegraphy we have an encoding operation which produces a sequence of dots, dashes and spaces on the channel corresponding to the message. In a multiplex PCM system the different speech functions must be sampled, compressed, quantized and encoded, and finally interleaved properly to construct the signal. Vocoder systems, television and frequency modulation are other examples of complex operations applied to the message to obtain the signal.
1. An *information source* which produces a message or sequence of messages to be communicated to the receiving terminal. The message may be of various types: (a) A sequence of letters as in a telegraph of teletype system; (b) A single function of time *f* (*t*) as in radio or telephony; (c) A function of time and other variables as in black and white television — here the message may be thought of as a function *f* (*x*; *y*;*t*) of two space coordinates and time, the light intensity at point (*x*; *y*) and time *t* on a pickup tube plate; (d) Two or more functions of time, say *f* (*t*), *g*(*t*), *h*(*t*) — this is the case in “three-dimensional” sound transmission or if the system is intended to service several individual channels in multiplex; (e) Several functions of several variables — in color television the message consists of three functions *f* (*x*; *y*;*t*), *g*(*x*; *y*;*t*), *h*(*x*; *y*;*t*) defined in a three-dimensional continuum — we may also think of these three functions as components of a vector field defined in the region — similarly, several black and white television sources would produce “messages” consisting of a number of functions of three variables; (f) Various combinations also occur, for example in television with an associated audio channel.
2. A *transmitter* which operates on the message in some way to produce a signal suitable for transmission over the channel. In telephony this operation consists merely of changing sound pressure into a proportional electrical current. In telegraphy we have an encoding operation which produces a sequence of dots, dashes and spaces on the channel corresponding to the message. In a multiplex PCM system the different speech functions must be sampled, compressed, quantized and encoded, and finally interleaved properly to construct the signal. Vocoder systems, television and frequency modulation are other examples of complex operations applied to the message to obtain the signal.
3. The *channel* is merely the medium used to transmit the signal from transmitter to receiver. It may be a pair of wires, a coaxial cable, a band of radio frequencies, a beam of light, etc.
4. The *receiver* ordinarily performs the inverse operation of that done by the transmitter, reconstructing the message from the signal.
5. The *destination* is the person (or thing) for whom the message is intended. We wish to consider certain general problems involving communication systems. To do this it is first