../

Probability & risk

How to reason when outcomes are uncertain: probabilities as beliefs, base rates and Bayes, expected value and its limits, regression and selection effects, fat tails, ruin, position sizing, optionality, compounding and forecasting. The general toolkit is in general thinking, how to turn these numbers into choices is in decision-making, and the mental errors that distort probability judgments are in cognitive biases. Algebra and logarithms are on math fundamentals.

Probabilistic thinking basics

A probability is a number from 0 to 1 that says how strongly you expect something. For one-off events ("will this launch hit 1,000 sign-ups in a month?") it can only be a degree of belief, not a long-run frequency. That is fine: beliefs can still be scored, compared and improved.

IdeaMeaningPractical rule
degree of beliefprobability as confidence, conditioned on what you knowalways ask "given what information?"
calibrationof all the things you call 70%, about 70% happenkeep score; see forecasting
resolutionyour forecasts vary and separate what happens from what doesn'tsaying 50% on everything is calibrated but useless
ranges over pointsgive an interval with a confidence level, not a single number"4–9 weeks, 80% confident" beats "6 weeks"
complementsP(not A)=1−P(A)P(\text{not } A) = 1 - P(A)if success is 30%, failure is 70%: plan for it
conjunctionP(A and B)=P(A) P(B∣A)≤P(A)P(A \text{ and } B) = P(A)\,P(B \mid A) \le P(A)every added detail makes a story less likely
disjunctionP(A or B)=P(A)+P(B)−P(A and B)P(A \text{ or } B) = P(A) + P(B) - P(A \text{ and } B)many small risks add up to a big one
independenceP(A and B)=P(A) P(B)P(A \text{ and } B) = P(A)\,P(B) only if unrelatedshared causes (one cloud region, one founder) break it
never 0 or 1a prior of exactly 0 or 1 can never be updatedreserve them for logic and definitions

Conjunctions shrink fast. A plan with 8 independent steps, each 90% likely, succeeds with probability 0.98≈0.430.9^8 \approx 0.43. Disjunctions grow fast. Ten independent risks at 5% each give a 1−0.9510≈40%1- 0.95^{10} \approx 40\% chance that at least one bites.

Ranges and confidence intervals

  • A 90% interval should contain the true value 9 times in 10. Most people's "90%" intervals are far too narrow: in Alpert and Raiffa's classic studies the true answer fell outside people's supposedly wide intervals far more often than the stated confidence allowed. Overconfidence is one of the most robust findings in judgment research.
  • Widen until you would be genuinely surprised by a value outside the range, on either side.
  • Anchor the ends separately: "what low value would surprise me?" then "what high value would?".
  • For quantities that multiply (market sizes, timelines, returns), think in ratios: "×2 either side" rather than "± 10".

Words vs numbers

Vague words hide disagreement. Sherman Kent (CIA, 1964, "Words of Estimative Probability") proposed a mapping:

TermProbabilityGive or take
certain100%0
almost certain93%about 6%
probable75%about 12%
chances about even50%about 10%
probably not30%about 10%
almost certainly not7%about 5%
impossible0%0

Better still: say the number. Words like "likely" or "a real possibility" mean very different things to different listeners, so two people can agree on the wording and disagree widely on the odds.

Fermi estimates (decompose an unknown into factors you can bound, multiply, and check) are the fastest way to get a prior when you have no data. Method and examples are on first principles.

Base rates and Bayes' theorem

The base rate (prior) is how common something is before you look at the specific case. The base-rate fallacy is judging by how well the evidence fits the story while ignoring how rare the story is.

P(H∣E)=P(E∣H) P(H)P(E∣H) P(H)+P(E∣¬H) P(¬H)P(H \mid E) = \frac{P(E \mid H)\,P(H)}{P(E \mid H)\,P(H) + P(E \mid \neg H)\,P(\neg H)}
SymbolNameMedical test reading
P(H)P(H)prior, base rateprevalence of the disease
P(E∣H)P(E \mid H)likelihood, true positive ratesensitivity
P(E∣¬H)P(E \mid \neg H)false positive rate1−1- specificity
P(H∣E)P(H \mid E)posteriorpositive predictive value (PPV)

Worked example: a positive screening test

A disease affects 1% of people screened. The test catches 90% of real cases (sensitivity 90%) and wrongly flags 9% of healthy people (specificity 91%). You test positive. What is the chance you have it?

P(D∣+)=0.90×0.010.90×0.01+0.09×0.99=0.0090.009+0.0891≈9.2%P(D \mid +) = \frac{0.90 \times 0.01}{0.90 \times 0.01 + 0.09 \times 0.99} = \frac{0.009}{0.009 + 0.0891} \approx 9.2\%

Most people, including many doctors, answer "about 90%". The error is confusing P(+∣D)P(+ \mid D) with P(D∣+)P(D \mid +).

The same problem in natural frequencies

Gigerenzer and Hoffrage (1995) showed that counts of people are far easier to reason with than conditional probabilities. Imagine 1,000 people:

1,000 people screened
├── 10 have the disease (1%)
│   ├──  9 test positive   (90% sensitivity)
│   └──  1 tests negative
└── 990 are healthy
    ├── 89 test positive   (9% false positives)
    └── 901 test negative
 
Positives: 9 + 89 = 98. Sick among them: 9.
P(disease | positive) = 9 / 98 ≈ 9%

In Gigerenzer et al. (2007), only 21% of 160 gynaecologists picked the right answer to a mammography version of this problem when it was stated as probabilities; after training in natural frequencies, 87% did.

LessonWhy
rare conditions produce mostly false positiveshealthy people vastly outnumber sick ones
a negative result is very reassuring hereP(D∣−)=1/902≈0.1%P(D \mid -) = 1/902 \approx 0.1\%
a second, independent positive changes a lotsee the odds form below: about 50%
screen high-risk groups, not everyonea higher prior raises the PPV

The same structure governs fraud alerts, intrusion detection, "this candidate interviewed brilliantly", and flaky test failures: when the thing you are hunting is rare, most alarms are false.

Bayesian updating as a habit

Bayes is easiest in odds form. Odds are p/(1−p)p/(1-p): 1% is 1:99, 50% is 1:1, 75% is 3:1.

P(H∣E)P(¬H∣E)⏟posterior odds=P(H)P(¬H)⏟prior odds×P(E∣H)P(E∣¬H)⏟likelihood ratio\underbrace{\frac{P(H \mid E)}{P(\neg H \mid E)}}_{\text{posterior odds}} = \underbrace{\frac{P(H)}{P(\neg H)}}_{\text{prior odds}} \times \underbrace{\frac{P(E \mid H)}{P(E \mid \neg H)}}_{\text{likelihood ratio}}

The likelihood ratio (LR) asks one question: how much more likely is this evidence if the hypothesis is true than if it is false? Evidence that is equally likely either way (LR = 1) is worthless, however vivid.

Test example again. Prior odds 1:99. LR of a positive result =0.90/0.09=10= 0.90 / 0.09 = 10. Posterior odds 10:99, which is 10/109≈9.2%10/109 \approx 9.2\%. A second independent positive multiplies by 10 again: 100:99, about 50%.

Likelihood ratioStrength of evidenceEffect on 50% priorEffect on 10% prior
1none50%10%
2weak67%18%
5moderate83%36%
10strong91%53%
100very strong99%92%
0.5 / 0.1evidence against33% / 9%5% / 1%

The habit

  1. Start from a base rate

    Pick a reference class ("seed-stage B2B startups", "migrations of this size in our codebase") and write down its frequency. This is the outside view.

  2. Ask what the evidence is diagnostic of

    For each new fact, estimate how likely you would see it if you were right and if you were wrong. Ignore evidence with a likelihood ratio near 1.

  3. Update in proportion

    Multiply the odds. Move a little on weak evidence and a lot on strong evidence. Many small updates beat rare lurches.

  4. Write it down and revisit

    Record the prior, the evidence and the posterior. Without a record, hindsight rewrites what you believed.

Worked example: will a pilot convert? Of your last 20 enterprise pilots, 6 converted to paid (prior 30%, odds 3:7). The new pilot's champion has budget authority. Looking back, 5 of the 6 converters (83%) had such a champion, as did 4 of the 14 non-converters (29%). LR =0.83/0.29≈2.9= 0.83/0.29 \approx 2.9. Posterior odds 3/7×2.9≈1.243/7 \times 2.9 \approx 1.24, so about 55%. Better than the base rate, still close to a coin flip: do not book the revenue.

Failure modeWhat it looks like
base-rate neglectjudging only by how well the case fits the story
conservatismupdating too little on strong evidence (Edwards, 1968)
overreactionlurching on one vivid anecdote
double countingtreating correlated evidence (three articles quoting one source) as independent
non-falsifiable priorsa hypothesis that fits every outcome has learned nothing from any of them

Expected value, utility and variance

Expected value is the probability-weighted average outcome:

E[X]=∑ipi xiE[X] = \sum_i p_i\,x_i

It is the right yardstick when the bet is small relative to your wealth and repeated many times (pricing insurance across a portfolio, choosing between A/B-test variants, spending engineering hours on bugs). It is the wrong yardstick when one outcome can wipe you out. Decision trees built on expected value are on decision-making.

Expected utility replaces money with how much the money matters to you. Daniel Bernoulli (1738) proposed logarithmic utility, u(w)=ln⁡wu(w) = \ln w, to resolve the St Petersburg paradox: a coin game with infinite expected value that nobody would pay much to play.

Worked example: a positive-EV bet you should refuse

You have $100,000. A bet pays +$100,000 or −$80,000 on a fair coin.

MeasureCalculationResult
expected value0.5(100,000)+0.5(−80,000)0.5(100{,}000) + 0.5(-80{,}000)+$10,000: looks good
expected log utility0.5ln⁡200,000+0.5ln⁡20,0000.5\ln 200{,}000 + 0.5 \ln 20{,}000lower than ln⁡100,000\ln 100{,}000
certainty equivalent200,000×20,000\sqrt{200{,}000 \times 20{,}000}about $63,000: worse than keeping $100,000

A log-utility investor treats this +EV bet as equivalent to losing $37,000 for sure. The downside cuts deeper than the upside helps because each dollar matters more when you have fewer.

ConceptFormulaMeaning
varianceVar⁡(X)=E[(X−μ)2]\operatorname{Var}(X) = E[(X - \mu)^2]how spread out outcomes are
standard deviationσ=Var⁡(X)\sigma = \sqrt{\operatorname{Var}(X)}spread in the same units as XX
certainty equivalentu(CE)=E[u(X)]u(CE) = E[u(X)]the sure amount you would swap the gamble for
risk premiumE[X]−CEE[X] - CEwhat you would pay to remove the risk
risk aversionconcave uu: E[u(X)]<u(E[X])E[u(X)] \lt u(E[X])why insurance exists
diversificationσavg=σ/n\sigma_{\text{avg}} = \sigma / \sqrt{n} for nn independent betsthe average of many small bets is stable; one big bet is not

Why the average isn't enough

PlanOutcomesEVSpreadVerdict
A$1m for sure$1mnonesafe
B50%: $0, 50%: $2.2m$1.1mhugebetter EV, but half the time you have nothing
C99%: $1.2m, 1%: −$10m (personal liability)$1.09mfat left tailthe tail is the whole story

Always ask for the distribution, not just the mean: the median, the 10th and 90th percentiles, and the worst case. A 1% annual chance of ruin sounds small; over 30 years it is 1−0.9930≈26%1- 0.99^{30} \approx 26\%.

Small samples and regression to the mean

The law of large numbers: as the sample grows, the sample average converges on the true average. Its error shrinks like 1/n1/\sqrt{n}, so 4× the data only halves the noise.

The "law of small numbers" (Tversky and Kahneman, 1971, Psychological Bulletin) is the mistaken belief that small samples are as representative as large ones. They found that even trained research psychologists expected small studies to replicate far more reliably than the math allows.

SymptomExampleCorrection
extremes come from small samplesthe counties with the highest and lowest cancer rates are both mostly small and ruralexpect small groups at both ends of every league table
early results look decisivean A/B test "wins" on day 2 with 40 conversionsfix sample size in advance; don't peek and stop
five interviews = a market"every user we spoke to wanted it"treat as hypothesis generation, not evidence of demand
a hot month = a trendone strong sales monthcompare with the month-to-month noise first

The hospital problem (Kahneman and Tversky, 1972): a large hospital has about 45 births a day, a small one about 15. Which records more days when over 60% of babies are boys? The small one, by a wide margin, because small samples swing further. Most people say "about the same".

Regression to the mean

When a measurement is part skill and part luck, an extreme result is likely to be followed by a less extreme one, because the luck does not repeat. Francis Galton described it in 1886 ("Regression toward mediocrity in hereditary stature"): the children of very tall parents were tall, but on average less tall than their parents.

In standardized units, if two measurements correlate with coefficient rr:

z^next=r⋅znow\hat z_{\text{next}} = r \cdot z_{\text{now}}

With r=0.5r = 0.5, someone two standard deviations above average this time is expected to be one above next time. With r=0r = 0 (pure luck), expect average. With r=1r = 1 (pure skill), no regression.

Story people tellWhat is really happening
praise makes people worse, criticism makes them betterKahneman's Israeli Air Force flight instructors (recounted in Thinking, Fast and Slow, ch. 17): cadets praised after an unusually good maneuver did worse next time, those shouted at after a bad one did better. Regression, not feedback.
the Sports Illustrated cover jinxathletes make the cover after an exceptional streak; an ordinary stretch follows
the new manager turned the team roundmanagers are fired after the worst runs, which tend to end anyway
the treatment workedpatients enroll when symptoms peak; many improve without treatment. Only a control group separates the two.
last year's top fund manager is a geniustop-quartile funds mostly fail to stay top quartile
our fix cut incidents by halffixes are made after the worst weeks

Rule: whenever you select on an extreme and then measure again, you need a control group that was selected the same way and not treated.

Selection effects

A sample is only evidence about the population it was drawn from by the process that drew it. If the process filters on the outcome, the sample lies.

Survivorship bias

In the Second World War, Abraham Wald of Columbia's Statistical Research Group wrote memoranda (around 1943) on estimating how vulnerable each part of an aircraft was from the damage on planes that returned. The insight: hits on returning planes mark places a plane can be hit and survive; the planes hit elsewhere never came back to be counted. The memoranda were published by the Center for Naval Analyses in 1980 and explained by Mangel and Samaniego (JASA, 1984).

DomainSurvivor sampleWhat is missing
startup advicefounders of successful companies explaining what workedthe thousands who did the same and failed (YC advice is useful partly because YC sees the failures too)
investingfunds and stocks that still exist in the indexclosed funds and delisted companies
architecture"old buildings were built better"the badly built ones fell down
hiring"our best engineers all came from X"how many from X you hired who didn't work out
products"users love feature Y" in a survey of current userseveryone who churned because of it
tech"we ran without backups for years and it was fine"the teams whose disks died

Other selection effects

EffectMechanismExample
Berkson's paradoxselecting on A or B makes A and B look negatively correlatedBerkson (1946) on hospital patients: two diseases look anticorrelated among in-patients because either one gets you admitted. Among startups you've heard of, great tech and great distribution seem anticorrelated because either gets you noticed.
self-selectionpeople choose whether to be in the samplecustomers who answer an NPS survey are the keen ones
non-response / attritiondropouts differ from stayersa trial where the sickest patients leave
collider biasconditioning on a common effectamong admitted students, test scores and grades look unrelated
publication biassignificant results get published, null results don'tthe published literature overstates effects
the Lindy questionwhat has survived a long time tends to be robustold ideas have passed a filter; new ones haven't yet

Test: "What would the data look like if my hypothesis were false, and would I have seen it?"

Correlation, causation and Simpson's paradox

Correlation between A and B has at least five explanations. Causation is only one of them.

ExplanationPatternExample
A causes BA → Bexercise lowers resting heart rate
B causes A (reverse)B → Asuccessful companies spend more on brand, not only the other way round
confounderC → A and C → Bice-cream sales and drownings both rise with summer heat
selection / colliderconditioning on a result of A and BBerkson's paradox, above
chancemany comparisons, small samplesspurious correlations between unrelated time series

Ways to get closer to causation: a randomized experiment (A/B test) where possible; a natural experiment or discontinuity; a dose–response relationship; a plausible mechanism; consistency across settings and methods; the cause preceding the effect. These echo Bradford Hill's 1965 considerations for inferring causation in epidemiology.

Simpson's paradox

A trend in every subgroup can reverse when the groups are combined, because the groups are different sizes and the grouping variable is related to both the "treatment" and the outcome.

Kidney stones (Charig et al., BMJ, 1986):

Stone sizeTreatment A (open surgery)Treatment B (PCNL)
small93% (81/87)87% (234/270)
large73% (192/263)69% (55/80)
all78% (273/350)83% (289/350)

A is better for small stones and for large stones, yet worse overall, because A was given mostly to the harder, large-stone cases.

UC Berkeley, 1973. Graduate admissions looked biased against women in aggregate: about 44% of 8,442 men were admitted versus about 35% of 4,321 women. Bickel, Hammel and O'Connell (Science, 1975) looked department by department and found no pattern of discrimination against women by admissions committees (if anything, a small bias in their favor). Women had applied disproportionately to the more competitive departments with low admission rates for everyone.

Rule: before trusting an aggregate comparison, ask what the groups differ on, and look within strata. Then ask which is the causally right comparison: stratifying on a confounder helps; stratifying on a consequence of the treatment can mislead.

Thin tails, fat tails and power laws

MediocristanExtremistan
coined byNassim Nicholas Taleb, The Black Swan (2007)same
typical distributionnormal (Gaussian), thin-tailedpower law, fat-tailed
examplesheight, weight, reaction times, manufacturing errorwealth, book sales, city sizes, startup outcomes, casualties in wars, cyber losses
one observationcannot move the total muchcan dominate the total
averagesstable, informativeunstable; the sample mean understates the true mean
past maximumgood guide to future maximumpoor guide: records keep being broken
strategyoptimize the typical casesurvive the worst case, position for the best

A normal distribution puts about 68% of values within 1 standard deviation, 95% within 2 and 99.7% within 3. Its tails fall off so fast that a 5σ event is essentially impossible. Real markets do not behave like that: on 19 October 1987 the Dow Jones fell 22.6% in one day, a move of more than 20 standard deviations by normal daily-volatility standards.

A power law (Pareto) tail has

P(X>x)∝x−αP(X > x) \propto x^{-\alpha}
Tail exponent α\alphaConsequence
α>2\alpha > 2mean and variance exist, but tails are still much heavier than normal
1<α≤21\lt \alpha \le 2mean exists, variance is infinite: sample standard deviations are meaningless
α≤1\alpha \le 1even the mean is infinite: the sum is dominated by the largest item
α≈1.16\alpha \approx 1.16gives the classic 80/20 split
NameObservation
Pareto principleVilfredo Pareto observed that about 80% of the land in Italy was owned by about 20% of the people; Joseph Juran later generalized it as "the vital few"
Zipf's lawthe kk-th most common word appears with frequency roughly ∝1/k\propto 1/k (George Kingsley Zipf, 1930s–1949); city sizes behave similarly
venture returnsin Horsley Bridge data reported by Chris Dixon (a16z, 2015), about 6% of investments, 4.5% of dollars invested, produced about 60% of total returns
softwarea few bugs cause most crashes; a few customers produce most revenue and most support load

Black swans

Taleb's black swan is an event that (1) lies outside regular expectations, (2) has an extreme impact, and (3) is explained after the fact as if it had been predictable. You cannot forecast specific black swans. You can decide in advance how exposed you are to them.

DoDon't
cap downside exposure: limits, insurance, redundancyuse standard deviation or VaR as "the risk" in fat-tailed domains
keep slack: cash, spare capacity, timeoptimize for the average day with no buffer
keep many small positive-tail bets openassume the largest past loss is the largest possible loss
stress-test with scenarios, not just statisticstreat "it has never happened" as "it can't happen"

Ergodicity and the ruin problem

A process is ergodic when the average over many people at one moment (the ensemble average) equals the average for one person over a long time (the time average). Most risky choices you make are not ergodic: you live one path, and you cannot average across parallel versions of yourself.

Russian roulette (a favorite example of Taleb's): six people each playing once for $1m have a five in six chance each of winning. One person playing six times has a (5/6)6≈33%(5/6)^6 \approx 33\% chance of surviving, and someone who keeps playing is eventually certain to die. Same expected value per round; entirely different fate.

Peters's coin (Ole Peters, "The ergodicity problem in economics", Nature Physics 15, 2019): each round, wealth rises 50% on heads or falls 40% on tails.

AverageCalculationPer round
ensemble (expected value)0.5×1.5+0.5×0.6=1.050.5\times 1.5 + 0.5 \times 0.6 = 1.05+5%: looks attractive
time average (typical path)1.5×0.6=0.9≈0.949\sqrt{1.5 \times 0.6} = \sqrt{0.9} \approx 0.949about −5%: you go broke

After 50 rounds the expected wealth is 1.0550≈11.5×1.05^{50} \approx 11.5\times the start, yet the typical player has 0.925≈0.07×0.9^{25} \approx 0.07\times. A tiny minority of lucky paths holds nearly all the ensemble's wealth. For multiplicative bets, maximize the growth rate (the expected log), not the expected value.

Rules for ruin

RuleWhy
survive first; optimize seconda return of −100% ends the game; no later return compounds from zero
ruin includes non-financial ruinreputation, health, trust, legal standing, a key relationship
small risks repeated are large risksper-event risk × number of exposures
losses need larger gains to recovera 50% loss needs a 100% gain; see compounding
don't borrow to raise a bet you can already affordleverage turns volatility into ruin
never risk what you need for what you merely wantBuffett's verdict on Long-Term Capital Management (1998, paraphrased): they risked money they had and needed to make money they didn't have and didn't need

Position sizing: the Kelly criterion

John Kelly (Bell Labs, 1956, "A New Interpretation of Information Rate") found the fraction of your bankroll to stake on a favorable repeated bet that maximizes the long-run growth rate. For a bet that pays bb to 1, with win probability pp and loss probability q=1−pq = 1 - p:

f∗=p−qb=bp−qbf^* = p - \frac{q}{b} = \frac{bp - q}{b}

It comes from maximizing the expected log of wealth after one bet:

g(f)=pln⁡(1+bf)+qln⁡(1−f),g′(f)=0  ⇒  f∗=bp−qbg(f) = p \ln(1 + bf) + q \ln(1 - f), \qquad g'(f) = 0 \;\Rightarrow\; f^* = \frac{bp - q}{b}

If f∗≤0f^* \le 0 there is no edge: don't bet.

Worked example

A repeated even-money bet (b=1b = 1) that you win 60% of the time: f∗=0.6−0.4/1=0.2f^* = 0.6 - 0.4/1 = 0.2. Stake 20% of current wealth each round.

Fraction of KellyStakeGrowth per round g(f)g(f)Share of max growth
¼ Kelly5%0.88%44%
½ Kelly10%1.50%75%
full Kelly20%2.01%100%
1.5× Kelly30%1.47%73%
2× Kelly40%−0.25%negative: long-run ruin

Another example: a 2-to-1 payout won 40% of the time: f∗=0.4−0.6/2=0.1f^* = 0.4 - 0.6/2 = 0.1.

Why practitioners use fractional Kelly

  • Half Kelly keeps about 75% of the growth with half the volatility. The growth curve is flat near the top and falls off a cliff past it, so undershooting is cheap and overshooting is expensive.
  • Your edge is estimated, not known. If you overestimate pp, full Kelly is really over-Kelly. At 2× the true Kelly fraction, long-run growth is roughly zero or negative.
  • Full Kelly is a rough ride. Deep drawdowns are routine: in the standard continuous approximation, a full-Kelly bettor has a 50% chance of seeing their bankroll halve at some point.
  • Real bets are not independent coin flips. Correlated positions, fat tails and illiquidity all argue for smaller sizes.

For a founder, Kelly is a way of thinking rather than a formula: size each bet (a hire, a market, a personal guarantee) so that being wrong is survivable and being right still matters.

Risk, uncertainty and margin of safety

Frank Knight (Risk, Uncertainty and Profit, 1921) distinguished risk (outcomes unknown, but the probabilities are measurable, as in dice or actuarial tables) from uncertainty (the probabilities themselves are unknown). He argued that profit is the reward for bearing uncertainty that cannot be insured.

RiskUncertainty
probabilitiesknown or estimable from dataunknown; no good reference class
examplescard games, car insurance, server failure ratesa new market, a regulatory change, a pandemic's second-order effects
toolsexpected value, Kelly, statisticsscenarios, robustness, optionality, margin of safety
failure modenone, if the model is righttreating uncertainty as risk: false precision

Margin of safety

Benjamin Graham titled the last chapter of The Intelligent Investor (1949) "Margin of Safety as the Central Concept of Investment": buy only well below your estimate of value, so that errors in the estimate and bad luck still leave you whole. Engineers call it a safety factor; founders call it runway. It is the practical answer to Knightian uncertainty: when you cannot trust the probabilities, make sure the plan survives being wrong. Size the margin to the uncertainty of the estimate and the cost of being wrong. Domain-by-domain examples and where it misleads are on general thinking.

Asymmetry, optionality and convexity

An asymmetric bet has limited downside and large upside. Optionality is the right, but not the obligation, to act later when you know more. A payoff is convex when it gains more from good surprises than it loses from bad ones.

PositionBad outcomeNormalGreat outcomeShape
owning a sharelose proportionallysmall gaingain proportionallylinear
holding a call optionlose the premium onlylose the premiumlarge gainconvex (likes volatility)
selling insurance / optionslarge losscollect premiumcollect premiumconcave (hates volatility)
startup equity (for a founder)lose time and salarymodestenormousconvex
taking on leveragewiped outamplified gainamplified gainconcave near ruin
a two-week prototypelose two weekslearn somethingfind a productconvex

Jensen's inequality is the math behind it. For a convex payoff ff, E[f(X)]≥f(E[X])E[f(X)] \ge f(E[X]), so more variability in XX raises the average payoff. For a concave payoff the inequality flips and variability hurts.

How to use it:

  • Seek convex exposures: cheap experiments, small early bets on many ideas, options on future choices (an MVP before a platform, a contractor before a hire).
  • Avoid concave ones you don't get paid enough for: picking up pennies in front of a steamroller, guarantees, uncapped liabilities, single points of failure.
  • Value reversibility: a reversible decision is a cheap option; decide fast. An irreversible one deserves the slow process in decision-making.
  • The barbell (Taleb): most resources very safe, a small portion in high-upside bets, little in the middle.
  • Options have a price. Keeping every door open costs focus. Exercise options (commit) when the information has arrived.

Compounding and the rule of 72

Growth at rate rr per period for nn periods:

FV=PV(1+r)nFV = PV (1 + r)^n

Doubling time solves (1+r)t=2(1 + r)^t = 2:

t=ln⁡2ln⁡(1+r)≈0.693r−r2/2≈0.693r+0.35t = \frac{\ln 2}{\ln(1 + r)} \approx \frac{0.693}{r - r^2/2} \approx \frac{0.693}{r} + 0.35

using ln⁡(1+r)≈r−r2/2\ln(1 + r) \approx r - r^2/2 for small rr. With rr in percent, 69.3/r69.3/r is exact for continuous compounding; for annual compounding the +0.35+0.35 correction means 72/r72/r is almost exact near 8% and within a few per cent of the truth from about 4% to 12%. 72 also divides evenly by 2, 3, 4, 6, 8, 9 and 12.

Rate per periodExact doubling timeRule of 72Rule of 69.3
2%35.036.034.7
4%17.718.017.3
6%11.912.011.6
8%9.09.08.7
10%7.37.26.9
15%5.04.84.6
25%3.12.92.8

Uses: 7% a year doubles in about 10 years; a product growing 10% a week doubles about every 7 weeks; 3% annual inflation halves purchasing power in about 24 years.

Losses compound too, and asymmetrically:

LossGain needed to recover
−10%+11%
−25%+33%
−50%+100%
−75%+300%
−90%+900%

This is why volatility drags on growth: the geometric mean return is roughly the arithmetic mean minus half the variance (g≈μ−σ2/2g \approx \mu - \sigma^2/2). Money, skills, reputation, code quality and technical debt all compound; so do small daily frictions. More on investing on the investing sheet.

Common statistical traps

TrapWhat happensDefense
p-hackingtrying analyses until p<0.05p \lt 0.05; Simmons, Nelson and Simonsohn (2011) showed flexible choices can make false positives likelypre-register the analysis; report everything you tried
garden of forking pathsGelman and Loken: even without conscious fishing, data-dependent choices inflate false positivesdecide the analysis before seeing the data; replicate
multiple comparisonstest 20 metrics at p=0.05p = 0.05 and expect one "win" by chancename one primary metric; correct (Bonferroni, false discovery rate)
peekingstopping an A/B test when it first looks significantfixed horizon or a sequential method built for peeking
base-rate neglectignoring the prior; see the test examplestart from the reference class
gambler's fallacyexpecting independent events to "even out" (red is due)independent means no memory
hot-hand fallacy (contested)Gilovich, Vallone and Tversky (1985) found no hot hand in basketball shooting; Miller and Sanjurjo (Econometrica, 2018) showed their method had a streak-selection bias, and after correcting it the data show evidence for a hot handdon't cite the hot hand as a textbook fallacy; streaks can be real where skill states vary
regression to the meancrediting an intervention for natural reversioncontrol group
survivorshipanalyzing only winnersfind the denominator
Texas sharpshooterdrawing the target around the cluster after the factstate hypotheses before looking
relative vs absolute risk"halves the risk" from 2 in 1,000 to 1 in 1,000always ask for absolute numbers; number needed to treat here is 1,000
denominator neglect"500 incidents" without the number of deploymentsquote rates, not counts
ecological fallacyinferring individual behavior from group averagesgroup-level data answer group-level questions
Goodhart's lawa measure that becomes a target stops measuring wellpair metrics; watch for gaming
extrapolationtrend lines beyond the data (every S-curve looks exponential early)ask what limits the growth

Forecasting and scoring

Philip Tetlock's Expert Political Judgment (2005) tracked tens of thousands of predictions by experts over nearly two decades and found most did little better than simple baselines; generalists who drew on many ideas (his "foxes") beat those with one big theory ("hedgehogs"). In the IARPA ACE forecasting tournament (2011–2015), Tetlock and Barbara Mellers's Good Judgment Project recruited volunteers, scored every forecast, and identified the top performers as superforecasters: roughly the top 2%. GJP won the tournament; Good Judgment reports that its forecasts were over 30% more accurate than intelligence analysts with access to classified information. The story is told in Superforecasting (Tetlock and Gardner, 2015).

The Brier score

For binary events, with forecast probability ftf_t and outcome ot∈{0,1}o_t \in \{0, 1\}:

BS=1N∑t=1N(ft−ot)2BS = \frac{1}{N} \sum_{t=1}^{N} (f_t - o_t)^2

0 is perfect, 0.25 is what always saying 50% earns, 1 is confidently wrong every time. Lower is better. Brier's original 1950 version sums over both outcomes, doubling the score to a 0–2 range; that is the scale Tetlock's books use.

ForecastHappened?Score
90%yes0.01
90%no0.81
60%yes0.16
50%either0.25
10%no0.01

The squared error punishes confident misses hard, and the score is proper: you minimize your expected score only by reporting what you actually believe.

Habits of good forecasters

Paraphrased from Tetlock and Gardner's findings:

HabitIn practice
outside view firststart from a base rate for similar cases, then adjust for specifics
break the question downFermi-style: what would have to be true, and how likely is each part?
granular probabilitiesdistinguish 60% from 70%; the best forecasters' fine distinctions carried real information
update often, in small stepsmove with each piece of news, rarely lurch
actively open-mindedseek out the strongest version of the other side
keep scorea forecast you never check teaches nothing
work in teamsaggregating independent forecasts beats most individuals
post-mortem hits and misseswas it the reasoning or the luck?

Useful ways to practice: keep a forecast log with dates and probabilities, and play on public forecasting platforms such as Metaculus or Good Judgment Open.

Risk checklist

RISK CHECKLIST (run before any bet you can't cheaply undo)
 
Base rate
[ ] What is the reference class, and how often does this work?
[ ] How is my case different, and would I bet on that?
 
Downside
[ ] What is the worst realistic outcome? Is it survivable?
[ ] Could this be ruin: money, reputation, health, trust?
[ ] Does the downside compound, cascade or trigger others?
[ ] Is anything here irreversible?
 
Distribution
[ ] Thin- or fat-tailed domain? Am I using averages where
    tails dominate?
[ ] What are the median, 10th and 90th percentile outcomes?
[ ] Are my risks correlated (one customer, one cloud, one
    founder, one bank)?
 
Evidence
[ ] Is my data survivorship- or selection-biased?
[ ] Could this be regression to the mean?
[ ] Confounders? Is the aggregate hiding a Simpson reversal?
[ ] How many things did I test before this looked good?
 
Sizing
[ ] Is the position small enough that being wrong is fine?
[ ] Am I estimating the edge honestly (use half of it)?
[ ] Does leverage turn a bad outcome into ruin?
 
Shape
[ ] Is the payoff convex (capped loss, open upside)?
[ ] Can I buy an option (pilot, prototype, trial) first?
[ ] Where is my margin of safety, and how big is it?
 
Record
[ ] Written probability, range and reasoning, dated.
[ ] Named signals that would make me change my mind.
[ ] Review date set.

References