VISTA·RESEARCH
How it worksThe LibraryHeritageAboutTalk to the Desk
The Library

First we work out what kind
of question you asked.

Most consequential questions are a known kind of problem wearing industry clothing. We identify the kind, say what evidence would actually settle it, and then go and get that evidence: structured interviews with the people who have run it, the published record, and where they are possible, surveys and experiments. This page shows the thinking. The Desk applies it to your decision.

Jump to the 84 entries

Thirty-year-old mathematics, finally computable.

The founder studied formal social science in Northwestern's Mathematical Methods in the Social Sciences program from 1989 to 1993, where his game theory professor was Roger Myerson, later a Nobel laureate. Across that program, the economics department, and later the University of Chicago Booth School of Business, he studied under six Nobel laureates. Not one of them gave him an A. The models they taught, and the decision sciences tradition of Kellogg's Managerial Economics and Decision Sciences department that ran alongside, were always the right way to think about a consequential decision. They were rarely usable, because the variables they needed could not be observed at the price of a research engagement. That is what changed. The Library restates those models as research moves, and Vista does the part general AI cannot: it goes and gets the missing evidence.

The shapes a question arrives in

Will people choose it, or pay for it?
Asking them is close to worthless. We make them choose between real alternatives and read what they give up.
Did X cause Y, or would it?
Before-and-after settles nothing on its own. This needs a comparison group, a threshold, or a change in timing.
Will it happen, and when?
Timing is a different question from probability and it is usually the one that decides. It is answered from comparable cases, not opinion.
How will the other side respond?
No survey answers this. It needs people who have sat in that seat and made that call under the same pressure.
Where should scarce things go?
Allocation questions often have exact answers once the constraints are written down honestly, and writing them down is most of the work.
What should we learn first, and when do we stop?
We rank the unknowns by how much each would move the decision, and we tell you when more research will not change it.
Who decides, and how will they decide?
Committees, boards, regulators, and voters follow rules and agendas that shape the outcome as much as the merits do.
Who would actually know?
Not the highest title. The person the information passed through, or the person who solved the same problem in a neighboring industry.
How will it move over time, and how big a bet is it?
Capacity cycles, contagion, tipping points, feedback, and the arithmetic of sizing a position: the pure mathematics that decides whether a plan survives motion.

Six steps, stated in advance.

1. State the decision and the money behind it.
Not the topic. The choice, the alternatives, the deadline, and what is at stake.
2. List what has to be true.
The five to seven assumptions the thesis rests on, each as a sentence that could be false.
3. Find the crux.
The one or two assumptions that would flip the decision if they fell. Most theses have one.
4. Rank the unknowns by what they would change.
Then research the top of that list, not the top of yours.
5. Go get the evidence.
Structured interviews with two or three vetted Advisors, a sweep of the record, and where the question needs it, a designed survey or experiment. Three to ten days.
6. Say what changed, and whether to stop.
What moved, why, where the transcripts support it, and whether further research is worth doing at all.
Every Vista brief ends with one sentence most research firms cannot write: whether more research would change this decision enough to justify its cost. When the answer is no, we say so, and we say it about our own work.

The Desk draws on the Library, as your decision warrants.

When you describe a decision to the Desk on our homepage, it is matching what you tell it against a library of formal structures a decision can have: the ways a question can turn out to be a known kind of problem. When one fits, the Desk may name the model it comes from, says the move in plain English, asks the one question that matters next, and tells you what evidence would settle it and how a Research Director would go and get it. It does not compute your answer, it does not give you a probability, and it does not tell you what to do. It shows you the shape of your question and what it would take to answer it. Most questions do not need a model at all, and the Desk does not invent one. Either way the engagement is the same: the right people interviewed, the record swept, one report with the transcripts. That is the beginning of every engagement, and it is free.

The entries

84 models, mapped.

Each one gives the model as it is taught, what it was originally built to answer, what it earns its keep on now, and the mathematics in one line. Open an entry for the full treatment: where it began, when it applies, what we would ask, what evidence would settle it, where the method breaks, and the sources.

84 of 84

Deciding under uncertainty

Whether to act, wait, stage, or learn more, and how much any of it is worth.

Expected utility and risk aversion

Expected value is not the decision

We separate what the bet is worth on average from what it is worth to you, given the size of the loss you cannot absorb.

Born to answer
Bernoulli, 1738: why nobody pays much to play a game of infinite expected value.
What it does now
Sizing a bet whose worst case is a loss the fund cannot absorb.
The mathematics
EU = sum of p(s) x U(X(s))
The full entry
Where it began

Daniel Bernoulli posed it in 1738 as the St. Petersburg game: a coin is flipped until it lands heads, and the pot doubles with every tail. The expected payout is infinite, and nobody will pay more than a few coins to play. His answer, that the value of money to a person bends, was the birth of utility. Pratt's 1964 paper turned the bend into a measurable number, the coefficient of risk aversion, which is why one investor's acceptable bet is another's ruin.

Why it works now

The bend was always real; what was missing was the distribution to apply it to. A base case with a bull and a bear beside it is not a distribution. Today the interviews can be designed to elicit ranges and tails from operators, the record can be read at scale for how often the bad state actually arrived, and a fund's own constraints can be captured at scoping. The arithmetic is unchanged and still done outside the conversation. What changed is that the inputs are obtainable inside a research engagement instead of being assumed.

When it applies

Two options have similar average payoffs but very different downsides, or a fund is sizing a position whose worst case is permanent impairment.

What we would ask

What is the loss that would actually change how you operate, not the loss that would merely hurt?
Is the downside a bad quarter or a permanent impairment?

What it needs

A distribution of outcomes, not a point estimate; the client's real constraints (mandate, concentration, liquidity, career risk).

How we would get it

Operators and authorities on the outcome distribution, especially the tail; the client's own constraints captured at scoping rather than assumed.

Where it breaks

Probabilities asserted rather than derived. Utility used as a cover for whatever the client already wanted to do.

For example

Two private-credit positions carry the same expected return. One has a small chance of a total loss that would breach a concentration limit and trigger redemptions; the other does not. On average they are twins. To this fund they are not, and the research question becomes the size and shape of that tail, not the average.

The mathematics in fullEV = sum over states of p(s) x X(s); the decision uses EU = sum of p(s) x U(X(s)) with U concave, and the gap between EV and the certainty equivalent is the price of the risk. Taught in

MMSS 300; DECS-430; MECN-451; MECS 550-1

Sources

Pratt (1964). Risk Aversion in the Small and in the Large. Econometrica 32(1/2), 122. https://doi.org/10.2307/1913738
Kahneman, Tversky (1979). Prospect Theory: An Analysis of Decision under Risk. Econometrica 47(2), 263. https://doi.org/10.2307/1914185

Decision trees and backward induction

Research is a decision, so is waiting

We lay out the choices and the unknowns in the order they will actually arrive, so that doing it in stages, or finding something out first, shows up as an option with a value.

Born to answer
Raiffa, 1968: a wildcatter deciding whether to drill, or to survey first.
What it does now
Finding the tranche, pilot, or milestone the yes-or-no framing hid.
The mathematics
V = max over branches, solved from the leaves back
The full entry
Where it began

Howard Raiffa's 1968 lectures built decision analysis around a wildcatter deciding whether to drill an oil well. The tree had two branches until Raiffa added a third: pay for a seismic survey first. Drawn that way, the survey is a decision node placed before the drilling decision, and its value is visible on the page. Kellogg still teaches trees with cases of the same shape.

Why it works now

Trees were used sparingly because filling in the chance nodes required data nobody had time to gather. Now the branches an investor actually faces, tranche, pilot, milestone, wait, can be enumerated in one conversation, and the chance nodes populated from comparable cases retrieved from the record and from Advisors who have run the staged version. The Desk draws the tree in words. The Research Director fills it from the interviews. Nobody in the conversation solves it.

When it applies

A single yes-or-no decision is on the table but the client could in fact invest in tranches, run a pilot, wait for a milestone, or research first.

What we would ask

What do you learn between now and the next point at which you could change course?
Is there a version of this where you commit less now and more later?

What it needs

The real sequence of decisions and information events, with what each intermediate result would reveal.

How we would get it

Advisors who have run the staged version of this decision; the record on how comparable milestones resolved.

Where it breaks

The tree omits the branch that mattered, usually wait or stage; the chance-node probabilities are asserted.

For example

A growth investor sees invest twenty-five million or pass. Drawn as a tree, a third branch appears: fund ten now against a retention milestone in two quarters, which is worth more than either original branch precisely because the milestone resolves the crux variable before the rest of the money goes in.

The mathematics in fullAt chance nodes V = sum of p(i) x V(i); at decision nodes V = max over branches; solved from the leaves back, so an information node placed before a decision node raises the value of the tree. Taught in

DECS-433; MECN-451; MMSS 300

Sources

Raiffa, Howard (1968). Decision Analysis: Introductory Lectures on Choices under Uncertainty. Addison-Wesley. (book)
Bertsimas, Dimitris and Freund, Robert M. (2004). Data, Models, and Decisions: The Fundamentals of Management Science. Dynamic Ideas. (book)

Value of information

What is the answer worth?

Before we research anything we ask whether any result could change what you do. If no result would, the research is worth nothing, however interesting.

Born to answer
Howard, 1966: what a seismic survey is worth before you drill.
What it does now
Refusing to sell research that no result could act on.
The mathematics
EVPI = E[max U] - max E[U]
The full entry
Where it began

Ronald Howard's 1966 paper gave the idea its name: information is worth the difference between the decision you would make with it and the one you would make without it. The teaching case in that tradition is the same wildcatter, asking how much a seismic survey is worth before drilling, and the answer is sometimes zero, because no survey result would change the decision to drill.

Why it works now

The idea was easy to state and hard to use, because it requires knowing what each possible answer would do to the decision, and that took a modeling exercise nobody commissioned. Today the Desk can work a client's decision back to the questions whose answers would flip it in the first conversation, and the whole engagement can be organized around those questions instead of around a topic. This is the single most commercially important idea in the Library, and it now costs a conversation rather than a study.

When it applies

A client has a list of open questions and is about to work through them in order, or is proposing research that could not change the decision whatever it found.

What we would ask

Which result, if we came back with it, would make you not do this?
And which would make you do it larger?

What it needs

The decision and its alternatives stated explicitly; the plausible results of each proposed study and what each would trigger.

How we would get it

This is scoping, not fieldwork: the Research Director works the decision back to the questions whose answers would flip it, then designs the study around those.

Where it breaks

The decision is already made and the study is theater; the set of actions was drawn too narrow for anything to flip.

For example

An investment team has forty diligence questions on a software company. Thirty-one of them cannot change the decision at any plausible answer. Two can. The engagement becomes those two, and the report says so.

The mathematics in fullEVPI = E[max over a of U(a, s)] - max over a of E[U(a, s)]; a real study is worth EVSI = E over results of [max over a of E[U | result]] minus the no-study value, and is pursued only when EVSI exceeds its cost. Taught in

MECN-451; DECS-433; MMSS 300

Sources

Howard (1966). Information Value Theory. IEEE Transactions on Systems Science and Cybernetics 2(1), 22-26. https://doi.org/10.1109/TSSC.1966.300074
Howard (1988). Decision Analysis: Practice and Promise. Management Science 34(6), 679-695. https://doi.org/10.1287/mnsc.34.6.679
Raiffa, Howard (1968). Decision Analysis: Introductory Lectures on Choices under Uncertainty. Addison-Wesley. (book)

Sensitivity and analysis ranking

The one question to answer next

We rank the unknowns by how much each would move the decision, per dollar and per day, and start with the top of that list rather than the top of yours.

Born to answer
Howard, 1988: spend the budget on the uncertainties the decision turns on.
What it does now
Ranking forty diligence questions by what each would change per week.
The mathematics
q* = argmax (EVSI - cost) / time
The full entry
Where it began

Howard's 1988 account of decision analysis in practice describes the cycle used at large companies: identify the uncertainties, find which ones the decision is sensitive to, and spend the analysis budget on those, resolving the most valuable first. The ranking, not the analysis, was the contribution; it told the team where to stop reading and start asking.

Why it works now

The ranking used to require a quantitative model of the decision before any research began, so it was reserved for the largest capital projects. Now a client's forty open questions can be laid against the decision in one sitting, each scored by what its plausible answers would change per week of effort, and the plan can be re-ranked as each answer arrives. The Research Director does the ranking; the Desk shows the client that it will be done, which is often the first thing that distinguishes Vista from a call.

When it applies

Time or budget will not cover every open question, or the client is spending on the questions that are easiest to answer rather than the ones that matter.

What we would ask

If you could resolve only one uncertainty before the committee meets, which one would you choose, and why that one?
What does it cost, in time as well as money, to answer each of the others?

What it needs

For each candidate question: what it would cost, how long it would take, and what its plausible answers would do to the decision.

How we would get it

A ranked research plan is the first deliverable; the fieldwork follows the ranking, and the plan is re-ranked as answers arrive.

Where it breaks

The ranking is done by interest rather than decision leverage; the cost of the slow answers is ignored.

For example

Eight enterprise-customer interviews, a competitor pricing study, more market-size work, and one former executive. Ranked by what each could change per week of effort, the former executive goes first, the customers second, and the market-size work is dropped because no plausible result would move anything.

The mathematics in fullChoose q* = argmax over studies q of (EVSI(q) - cost(q)) / time(q); re-solve after each result, since the value of the remaining studies changes as beliefs move. Taught in

MECN-451; MECS 560-1

Sources

Howard (1988). Decision Analysis: Practice and Promise. Management Science 34(6), 679-695. https://doi.org/10.1287/mnsc.34.6.679
Kantorovich (1960). Mathematical Methods of Organizing and Planning Production. Management Science 6(4), 366-422. https://doi.org/10.1287/mnsc.6.4.366

Optimal stopping

When to stop paying for research

We tell you when further research is unlikely to change the decision enough to justify its cost, including research from us.

Born to answer
Wald, 1945: stop inspecting once the evidence crosses a threshold.
What it does now
Telling a client to stop paying for research, including ours.
The mathematics
Stop when the best remaining net VOI is at or below zero
The full entry
Where it began

Abraham Wald developed sequential analysis during the Second World War for inspecting munitions: instead of testing a fixed number of items, stop as soon as the accumulated evidence crosses a threshold either way. It cut inspection effort by a large fraction and was classified until 1945. Bellman's dynamic programming later gave the general form: stop when the value of continuing falls below the value of acting.

Why it works now

A research firm paid by the hour has no reason to compute a stopping rule and every reason not to. A firm on a fixed fee can state the rule in advance and mean it. What AI adds is the ability to track, across every interview and document, whether the recommendation is still moving, so that saturation is observed rather than declared. The stopping sentence in a Vista brief is Wald's rule applied to a client's money.

When it applies

The client is on the third round of diligence and each round moves the picture less; or a firm paid by the hour keeps finding one more source.

What we would ask

What did the last round of work change about the decision?
Is there any remaining question whose plausible answers would flip it?

What it needs

A record of what each round of research changed; the plausible range of the remaining unknowns.

How we would get it

Stated as a rule at scoping: the engagement ends when the recommendation is stable across the plausible results of every remaining study, and the report says so in one sentence.

Where it breaks

Self-serving in either direction: hourly firms never stop, fixed-fee firms stop early. The defense is stating the rule before the work starts.

For example

After the customer interviews and the former executive, the thesis is fragile on one variable that no obtainable evidence will resolve before the deadline. The honest sentence is that more research will not settle it, that the decision is being made under that uncertainty, and that the client should stop paying, including paying us.

The mathematics in fullStop when U(stop | state) is at least -cost(q) + E[V(state after q)] for every remaining study q; equivalently when the best remaining net value of information is at or below zero. Taught in

MECS 560-2; MECN-451

Sources

Wald (1945). Sequential Tests of Statistical Hypotheses. The Annals of Mathematical Statistics 16(2), 117-186. https://doi.org/10.1214/aoms/1177731118
Bellman (1954). The theory of dynamic programming. Bulletin of the American Mathematical Society 60(6), 503-515. https://doi.org/10.1090/S0002-9904-1954-09848-8

Real options

Waiting has a value, and sometimes a price

We work out whether the next quarter resolves a variable this decision turns on, in which case waiting is worth something, and whether the window will still be open, in which case it is not.

Born to answer
McDonald and Siegel, 1986: an irreversible investment worth waiting on.
What it does now
Whether the next quarter resolves the crux, and whether the window stays open.
The mathematics
Invest only when V exceeds a threshold V* above cost
The full entry
Where it began

McDonald and Siegel's 1986 paper asked when a firm should make an irreversible investment whose value fluctuates, and found that the usual rule, invest when value exceeds cost, is wrong. The right rule waits until value exceeds cost by a wide margin, in their calibration roughly double, because investing kills the option to invest later on better information. Dixit and Pindyck built the real-options field on it.

Why it works now

The model needs to know how fast the uncertainty resolves and whether the window stays open, which used to be guesses. Both are now researchable: operators can say how quickly the crux variable actually clears in their industry, and the record shows how often comparable windows closed. The Desk asks the two questions that decide it, what would you know in six months and who can take this away from you meanwhile, and the engagement answers them.

When it applies

An irreversible commitment under high uncertainty: a plant, an acquisition, a platform bet, a market entry; or the reverse, a client waiting for certainty that will never arrive while a competitor moves.

What we would ask

What would you know in six months that you do not know now?
Who else can take this decision away from you in the meantime?

What it needs

How much of the uncertainty resolves with time versus never; the reversibility of the commitment; the competitive clock.

How we would get it

Operators on how fast the crux variable actually resolves in this industry; the record on comparable windows that closed.

Where it breaks

The option is illusory because the seller walks or the competitor enters; or the uncertainty is permanent and waiting only delays.

For example

A buyer can close a carve-out now or wait for the target's next two quarters of results. The wait is worth a great deal if those quarters reveal whether the churn is structural, and worth nothing if a strategic bidder is already in the data room.

The mathematics in fullInvest now only when value V exceeds a threshold V* strictly above the cost I; the gap V* - I grows with the volatility of V and with the irreversibility of I, and shrinks to zero when the option can be taken away. Taught in

DECS-433; MECN-451

Sources

McDonald, Siegel (1986). The Value of Waiting to Invest. The Quarterly Journal of Economics 101(4), 707. https://doi.org/10.2307/1884175
Dixit, Avinash K. and Pindyck, Robert S. (1994). Investment under Uncertainty. Princeton University Press. (book)

Monte Carlo simulation and Jensen's inequality

The base case is not the expected case

We replace bull, base, and bear with the whole spread of outcomes, because a plan built on average inputs does not produce the average result.

Born to answer
Jensen, 1906; Ulam at Los Alamos, 1949: answers from many random trials.
What it does now
Replacing bull, base and bear with the whole spread of outcomes.
The mathematics
E[f(X)] is not f(E[X]) whenever f bends
The full entry
Where it began

The mathematics is Jensen's 1906 inequality on convex functions. The name is Sam Savage's, from his joke about the statistician who drowned crossing a river that was on average three feet deep. Monte Carlo itself was born at Los Alamos in the late 1940s, when Stanislaw Ulam, playing solitaire while convalescing, realized that many random trials could answer questions the equations could not, and Metropolis and Ulam published the method in 1949.

Why it works now

Running the simulation was never the hard part; deciding what to put in it was. Distributions invented at a desk produce a histogram of assumptions. What is now practical is eliciting ranges and co-movements from the people who have watched the variables move, reading the record for their historical spread, and feeding a model the client already owns. Vista designs the inputs. The client or a specialist runs the model. The Desk explains why the base case is not the expected case.

When it applies

The model has a base case with a bull and a bear beside it; the thesis depends on several uncertain inputs at once; or the inputs move together in bad states.

What we would ask

Which three inputs, if they came in at their worst plausible values together, would break the plan?
Do those inputs tend to go wrong at the same time?

What it needs

Ranges rather than points for the inputs that matter, and the correlations among them; which the interviews are designed to elicit.

How we would get it

Advisors asked for ranges and for what moves together, not for point forecasts; the record for the historical spread of the same variables.

Where it breaks

The input distributions are made up, so the histogram is a picture of assumptions; correlations are ignored and the tails are understated.

For example

A leveraged buyout model shows a comfortable base case. Run across the ranges the operators actually gave for growth, margin, and exit multiple, with the three falling together in a downturn, the probability of a covenant breach is not a footnote. It is the crux, and it points at the one variable to research next.

The mathematics in fullY = f(X1..Xk) with each X drawn from its range and correlated where they move together; by Jensen, E[f(X)] is not f(E[X]) whenever f bends, and leverage makes f bend. Taught in

MECN-451; DECS-433

Sources

Metropolis, Ulam (1949). The Monte Carlo Method. Journal of the American Statistical Association 44(247), 335-341. https://doi.org/10.1080/01621459.1949.10483310
Jensen (1906). Sur les fonctions convexes et les inégalités entre les valeurs moyennes. Acta Mathematica 30(0), 175-193. https://doi.org/10.1007/BF02418571
Bertsimas, Dimitris and Freund, Robert M. (2004). Data, Models, and Decisions: The Fundamentals of Management Science. Dynamic Ideas. (book)

Heavy tails

Near misses are data about the disaster

We treat the incidents that almost happened as evidence about the one that could, and fit the tail from them instead of from the calm years.

Born to answer
Mandelbrot, 1963: cotton price moves far larger than a bell curve allows.
What it does now
Fitting the disaster from the near misses nobody counted.
The mathematics
P(X > x) falls like a power of x, not exponentially
The full entry
Where it began

Benoit Mandelbrot's 1963 study of cotton prices showed that the big moves came far more often than a bell curve allows, and that the tails, not the middle, held the risk. The management version is taught at Kellogg through the Boeing 737 MAX: a rare disaster that the near-miss record had been describing for some time to anyone who counted it.

Why it works now

Near misses were always the evidence and were never collected, because nobody read every incident report. Reading every incident report is now cheap. Former plant, safety, and quality leaders can be asked for the count that never reached the board, the record can be swept for comparable events at peers, and the tail can be fitted from what almost happened rather than from the calm years. Where the fit needs real extreme-value statistics, a specialist does it; recognizing that it is needed is now a conversation.

When it applies

A rare, severe event dominates the thesis: recall, outage, safety failure, regulatory action, fraud; and the client is reasoning from the years in which it did not happen.

What we would ask

How many near misses has this business had, and who counted them?
What would the event cost, and is that a bad year or the end?

What it needs

Incident and near-miss history, including the ones not reported; comparable events at peers; the size of the loss conditional on the event.

How we would get it

Former operators, safety and quality leads, and regulators on the near-miss record; the record on comparable failures and their consequences.

Where it breaks

The tail is fitted to a handful of observations; the absence of the event is read as evidence of its impossibility.

For example

An industrial company has had no serious incident in a decade and four near misses in two years. The decade is not the evidence. The four are, and the interviews with former plant managers are where the count comes from.

The mathematics in fullFit the frequency of severe events from the full incident distribution rather than the severe ones alone; the loss is E[L | event] x P(event), and heavy tails with exponent below two make the sample average an unreliable estimate of either term. Taught in

MECN-451; MECNX-435

Sources

Mandelbrot (1963). The Variation of Certain Speculative Prices. The Journal of Business 36(4), 394. https://doi.org/10.1086/294632
Clauset, Shalizi, Newman (2009). Power-Law Distributions in Empirical Data. SIAM Review 51(4), 661-703. https://doi.org/10.1137/070710111

Ambiguity and Knightian uncertainty

Some uncertainties have no defensible number

When the evidence cannot supply a probability, we say so, and we design the decision to be robust across the range of beliefs the evidence actually permits, rather than pretending to one number.

Born to answer
Ellsberg, 1961: people prefer the urn whose mix they know.
What it does now
Saying when the evidence supports a range of beliefs and no single number.
The mathematics
max over acts of min over priors of E[U]
The full entry
Where it began

Daniel Ellsberg's 1961 experiment offered people two urns, one with a known fifty-fifty mix of colors and one with an unknown mix, and found that most preferred to bet on the known urn whichever color they were betting on. No single probability can explain that. Gilboa and Schmeidler in 1989 and Klibanoff, Marinacci, and Mukerji in 2005 built decision theories in which a careful person holds a set of beliefs rather than one, and is not being irrational when the set is wide.

Why it works now

For decades the practical consequence was nil, because nobody could say how wide the set of defensible beliefs actually was. Now the record can be searched for whether a reference class exists at all, and Advisors can be asked for the range a careful person could hold, so the width of the set becomes a finding rather than a mood. The Desk says when a question has no defensible number, and the engagement reports which decisions survive the whole range instead of pretending to a point.

When it applies

The client wants a probability for something with no comparable history: a novel regulation, a first-of-kind technology, a geopolitical break; or a report has supplied a precise number that nothing supports.

What we would ask

What class of past cases would this belong to, and does one exist?
Across the range of beliefs a careful person could hold, which decision holds up?

What it needs

Honesty about the reference class: whether one exists, and how wide the range of defensible beliefs is.

How we would get it

Authorities on whether any base rate exists; operators on what the plausible range is; the report gives a range of priors and the decision that survives all of them.

Where it breaks

Ambiguity used as an excuse not to estimate the estimable; or a single number reported where a range of priors was the truth.

For example

A fund wants the probability that a proposed rule takes effect as drafted. There is no base rate for this rule. The honest deliverable is the range a careful reader of the record would hold, and which position sizes survive the whole range.

The mathematics in fullRather than a single prior, a set of priors; the robust choice maximizes the minimum expected utility over the set, or, with smooth ambiguity, an ambiguity-averse aggregate of the expected utilities across it. Taught in

MECS 550-1; MMSS 300

Sources

Ellsberg (1961). Risk, Ambiguity, and the Savage Axioms. The Quarterly Journal of Economics 75(4), 643. https://doi.org/10.2307/1884324
Gilboa, Schmeidler (1989). Maxmin expected utility with non-unique prior. Journal of Mathematical Economics 18(2), 141-153. https://doi.org/10.1016/0304-4068(89)90018-9
Klibanoff, Marinacci, Mukerji (2005). A Smooth Model of Decision Making under Ambiguity. Econometrica 73(6), 1849-1892. https://doi.org/10.1111/j.1468-0262.2005.00640.x

The outside view

The outside view before the inside view

Before we argue about this case, we look at how the whole class of cases like it turned out, and we run the failure post-mortem before the failure.

Born to answer
Kahneman and Lovallo, 1993; Flyvbjerg on project overruns.
What it does now
Starting from how the class of comparable deals actually turned out.
The mathematics
Forecast = the class distribution, then adjust
The full entry
Where it began

Kahneman and Lovallo's 1993 paper distinguished the inside view, which reasons from the specifics of the case, from the outside view, which starts from how similar cases turned out, and showed that the inside view produces bold forecasts and timid choices. Bent Flyvbjerg turned the outside view into a procedure, reference class forecasting, first for transport projects, where the class of comparable projects predicted cost overruns the planners never did, and it was adopted into official appraisal guidance.

Why it works now

The outside view was rarely taken because assembling the reference class was a research project in itself. It is now the cheapest part of an engagement: the record can be read for how the class of comparable deals, migrations, launches, or turnarounds actually resolved, and the Advisors can be asked the pre-mortem question directly. The Desk starts from the class and asks what makes this case different, which is the reverse of how most investment memos are written.

When it applies

A forecast built from the specifics of the deal, the plan, or the team, with no reference to how similar deals, plans, and teams have done; or a projection everyone in the room finds persuasive.

What we would ask

What is the class of comparable cases, and how did they actually turn out?
If this has failed two years from now, what is the most likely reason?

What it needs

A defensible reference class and its outcome distribution; the plan's own assumptions listed so they can be checked against it.

How we would get it

The record for the reference class; operators who have seen the failures; a structured pre-mortem with the Advisors as the second half of each interview.

Where it breaks

The reference class is chosen to flatter; the bias story is applied after the fact and cannot be falsified.

For example

A platform migration plan promises eighteen months. The class of comparable migrations at companies of this size ran twice the planned time in most cases. The research starts from that distribution and asks what makes this one different, rather than starting from eighteen months and asking what could go wrong.

The mathematics in fullForecast = distribution of outcomes in the reference class, then adjust for the case at hand; the adjustment is the inside view and should be smaller than the decision maker wants it to be. Taught in

MECNX-435; MECN-451

Sources

Kahneman, Lovallo (1993). Timid Choices and Bold Forecasts: A Cognitive Perspective on Risk Taking. Management Science 39(1), 17-31. https://doi.org/10.1287/mnsc.39.1.17
Flyvbjerg (2006). From Nobel Prize to Project Management: Getting Risks Right. Project Management Journal 37(3), 5-15. https://doi.org/10.1177/875697280603700302
Kahneman, Tversky (1973). On the psychology of prediction. Psychological Review 80(4), 237-251. https://doi.org/10.1037/h0034747

The limits of subjective probability

The question has no answer worth buying

Some questions cannot be answered defensibly from any evidence obtainable in the time available. When that is true we say so before you spend, and it is often the most useful sentence in the conversation.

Born to answer
Knight, 1921; Ellsberg, 1961: uncertainty that admits no probability.
What it does now
Declining the question before the client spends on it.
The mathematics
No structure applies, and naming that is the discipline
The full entry
Where it began

Frank Knight drew the line in 1921 between risk, which can be measured, and uncertainty, which cannot, and Ellsberg's urns made the line experimentally real. Gilboa's modern treatment keeps the honest case: some questions have no probability that a careful person could defend, and the rational response is to decide in a way that does not require one.

Why it works now

The temptation to manufacture a number has grown with the tools that can produce one. The discipline that has become cheaper is the opposite: checking, quickly and thoroughly, whether anyone anywhere is positioned to know the thing, and if nobody is, saying so before the client spends. The Desk does this in the conversation. It is the entry in the Library that exists to refuse.

When it applies

The question turns on something nobody can observe in time (a private decision not yet made, a hidden preference, a future negotiation), or on a quantity with no comparable history and no one positioned to know it.

What we would ask

Who, anywhere, is actually positioned to know this today?
If nobody is, what decision would you make under that uncertainty, and is that the real question?

What it needs

Candor about what is observable; a reframing of the decision so it can be made under the uncertainty rather than after resolving it.

How we would get it

None sold. The Research Director reframes the decision to the question that can be answered, if there is one, and otherwise declines.

Where it breaks

Declining what could have been partially answered; substance first, handoff last is the rule for everything short of this case.

For example

A client wants to know whether a specific acquirer will bid for a specific target next quarter. Nobody outside that acquirer's board knows, and they will not say. The answerable question is what the target is worth to that acquirer and what a bid would need to look like, which is a different engagement.

The mathematics in fullNo structure applies; this is the boundary of the library, and naming it is the discipline. Taught in

MECS 550-1

Sources

Ellsberg (1961). Risk, Ambiguity, and the Savage Axioms. The Quarterly Journal of Economics 75(4), 643. https://doi.org/10.2307/1884324
Gilboa, Itzhak (2009). Theory of Decision under Uncertainty. Cambridge University Press. (book)

Framing and the strategy table

The option you did not list

Most decisions arrive as a choice between two things. Before analyzing either, we generate the alternatives that were never listed, and we separate what has already been decided from what is being decided now and what can be deferred.

Born to answer
Howard at Stanford: the decision hierarchy and generated alternatives.
What it does now
Finding the third option a yes-or-no memo left out.
The mathematics
A strategy is one coherent path across the dimensions
The full entry
Where it began

Howard's decision analysis at Stanford insists that the quality of a decision is decided before any analysis, by the frame and the alternatives, and it supplies two tools: the decision hierarchy, which separates what is taken as given from what is being decided and what is deferred, and the strategy table, which lays out every dimension of a strategy with its options so that whole strategies can be composed rather than two options compared. The Stanford text calls the failure to recognize alternatives one of the most expensive errors in practice.

Why it works now

Alternatives were generated in a workshop by whoever was in the room. The table can now be built in a conversation, with the dimensions and options drawn from comparable decisions in the record and from Advisors who have seen the versions that were never listed. The Desk asks what has already been decided that this decision takes as given, and what the client would do if both listed options were unavailable.

When it applies

A yes-or-no decision; a choice between two options that both feel wrong; a plan whose alternatives were set by whoever wrote the memo.

What we would ask

What has already been decided that this decision takes as given, and should it?
If both of these options were unavailable, what would you do?

What it needs

The decision hierarchy stated explicitly, and a structured search for alternatives across the dimensions of the strategy.

How we would get it

The Research Director builds the strategy table with the client and the Advisors before any evidence is gathered: each dimension of the decision, its options, and the coherent combinations.

Where it breaks

Analyzing the two listed options with great care; the answer was a third one.

For example

A board is deciding whether to sell a division or keep it. Laid out as a table, there are also a partial sale, a joint venture, a carve-out with a service agreement, and a run-off. Two of those are worth more than either original option, and the research is redirected to them.

The mathematics in fullA decision hierarchy separates the givens, the decision, and the deferred; a strategy table lists each dimension of the strategy with its options, and a strategy is one coherent path across the columns. Taught in

MS&E 252 and 352 (Stanford); MECN-451

Sources

Howard, Ronald A. and Abbas, Ali E. (2016). Foundations of Decision Analysis. Pearson. (book)
Howard (1988). Decision Analysis: Practice and Promise. Management Science 34(6), 679-695. https://doi.org/10.1287/mnsc.34.6.679

Weighing evidence

How much each source and each fact should move a belief, and how to know when it did.

Bayesian updating

The thesis as a list of things that must be true

We break the thesis into the five to seven assumptions it rests on, say how each stands on the evidence, and record what moved each one and why.

Born to answer
Bayes, 1763: inferring a ball's position from where later balls fall.
What it does now
The thesis as five to seven claims, each with what moved it.
The mathematics
Odds(H | E) = Odds(H) x likelihood ratio
The full entry
Where it began

Thomas Bayes's essay, published in 1763 after his death, imagined a ball rolled onto a table and its position inferred from where later balls landed to its left or right: with each new ball, the belief about the first one's position updates. Laplace made it a working method. The odds form, in which each piece of evidence multiplies the odds by how much likelier it is under one hypothesis than the other, is what a working analyst actually uses.

Why it works now

The ledger was never kept because writing down every assumption and what moved it was clerical work nobody did. AI makes the clerical part free: every interview and document can be tagged to the assumption it bears on and the direction it moved it. What must not change is that the movement is stated in words, more likely than not, the evidence conflicts, unresolved, and not as a number that nothing calibrates. Vista keeps the ledger in the brief and preserves it afterward, which is how a track record starts.

When it applies

A thesis stated as a narrative rather than as claims that could be false; or a client who cannot say which assumption the whole thing rests on.

What we would ask

What has to be true for this to work? Say it as a sentence that could be false.
Which of those, if it fell, takes the rest with it?

What it needs

Each assumption stated falsifiably, with the evidence for and against it kept separate and traceable to its source.

How we would get it

The report opens with the ledger: each assumption, a plain-language confidence band (more likely than not; the evidence conflicts; unresolved), what moved it, and where the transcript supports it. No invented percentages.

Where it breaks

Numbers dressed as calibrated probabilities that are in fact intuition; a ledger written after the conclusion to justify it.

For example

A thesis on a software company reduces to seven sentences. Two are well supported, three conflict, and one is unresolved and decisive: that customers past two hundred seats face real switching costs. The engagement is built around that sentence.

The mathematics in fullOdds(H | E) = Odds(H) x LR; stated in prose as which direction each piece of evidence moved each assumption and how strongly, with the likelihood ratio the reason and never a displayed number. Taught in

DECS-430; MATH 385

Sources

Bayes (1763). An Essay towards Solving a Problem in the Doctrine of Chances. Philosophical Transactions of the Royal Society of London(53), 370-418. https://doi.org/10.1098/rstl.1763.0053

Informant accuracy

Who is reliable about what

We weight each source by what they were positioned to know, not by their title. A former salesperson is a strong source on pricing and a weak one on technology, and the report says which is which.

Born to answer
Romney, Weller and Batchelder, 1986: competence with no answer key.
What it does now
Weighting a source by what they were positioned to see.
The mathematics
Weight source i by its reliability r(i, d) in domain d
The full entry
Where it began

Romney, Weller, and Batchelder's 1986 cultural consensus model came from anthropology, where informants are asked about their own culture and there is no answer key. The model estimates each informant's competence from how much they agree with the others, and recovers the consensus answers without knowing them in advance. Competence, it turned out, is specific to the domain of the questions.

Why it works now

Expert networks have always sold access and left reliability to the client's intuition. What is now feasible is to record, for every Advisor, what they could observe from where and when, to tag each claim in a transcript by the domain it belongs to, and to weight the synthesis accordingly, so that a former salesperson counts heavily on pricing and lightly on technology. Over many engagements that becomes a dataset about which kinds of sources are reliable about which kinds of claims, which no one has ever built.

When it applies

Evidence from several sources of unequal access, recency, and incentive is being averaged as if it were equal; or a senior title is being treated as expertise on everything.

What we would ask

What did this person actually see, from where, and how long ago?
What do they gain if you believe them?

What it needs

For each source: role, access, recency, incentives, conflicts, and independence from the other sources.

How we would get it

Advisors are chosen by vantage point, and each interview records what the Advisor could and could not observe; the synthesis weights claims by that record, not by seniority.

Where it breaks

A single universal credibility score per source; reliability treated as a trait rather than as domain-specific.

For example

Three sources agree the product is winning. One is a former sales lead who saw the pipeline, one a customer who saw the product, one an analyst who saw the pitch. On pricing, the first counts most; on product fit, the second; the third counts on neither.

The mathematics in fullFor source i on domain d, a reliability r(i, d) enters the likelihood ratio of that source's report; the aggregate is a weighted sum of log likelihood ratios, with weights for reliability, relevance, recency, and independence. Taught in

MMSS 1995 (informant accuracy); DECS-430

Sources

Romney, Weller, Batchelder (1986). Culture as Consensus: A Theory of Culture and Informant Accuracy. American Anthropologist 88(2), 313-338. https://doi.org/10.1525/aa.1986.88.2.02a00020

Informational cascades

Five sources, or one source heard five times

Before we count confirming sources we check whether they saw the thing themselves or heard it from each other. Five customers who share an implementation partner are closer to one observation than to five.

Born to answer
Banerjee and Bikhchandani, 1992: diners choosing the fuller restaurant.
What it does now
Counting independent clusters of evidence, not voices.
The mathematics
Log odds add only across independent evidence
The full entry
Where it began

Abhijit Banerjee's 1992 model has diners choosing between two restaurants: each sees a private signal and the choices of those ahead, and after a few people the crowd follows the crowd regardless of what anyone privately knows. Bikhchandani, Hirshleifer, and Welch published the general theory of informational cascades the same year and used it to explain fads and fashions that flip on almost nothing.

Why it works now

A research firm used to count confirming sources. It could not trace where each source got the story, so five agreeing calls looked like strong evidence even when four had heard it from the fifth. Provenance is now traceable: transcripts can be read for who saw what directly and who is repeating, and Advisors can be recruited from deliberately separate channels. The brief counts independent clusters, not voices.

When it applies

Consensus among sources is being treated as strong evidence; the sources share a channel, a consultant, a conference, or an outage; or a market view has converged quickly.

What we would ask

How did each of these people come to know this? Did any of them see it directly?
Do they talk to each other, or to the same third party?

What it needs

The provenance chain of each claim; a map of which sources share upstream sources.

How we would get it

Advisors recruited deliberately from different channels and vantage points; the report clusters evidence by independence and counts clusters, not voices.

Where it breaks

A cascade is read as convergent evidence; independence is asserted because the sources have different job titles.

For example

Six channel partners report the same competitor weakness. Five learned it from the same distributor's sales kickoff. The evidence is one distributor's claim plus one independent observation, and the report says so.

The mathematics in fullLog odds add across pieces of evidence only when they are independent; with pairwise correlation rho among n confirming sources, the effective number of observations approaches 1/rho rather than n. Taught in

MECN-451; DECS-430

Sources

Banerjee (1992). A Simple Model of Herd Behavior. The Quarterly Journal of Economics 107(3), 797-817. https://doi.org/10.2307/2118364
Bikhchandani, Hirshleifer, Welch (1992). A Theory of Fads, Fashion, Custom, and Cultural Change as Informational Cascades. Journal of Political Economy 100(5), 992-1026. https://doi.org/10.1086/261849

Signaling and cheap talk

What they said versus what it cost them

We separate claims, which cost nothing, from signals, which would have been expensive for a liar, from commitments, which are hard to reverse. The plan is weighed on the second and third.

Born to answer
Spence, 1973: education as a costly signal; Crawford and Sobel, 1982.
What it does now
Separating what management said from what it spent.
The mathematics
A signal separates when c(s, high) < c(s, low)
The full entry
Where it began

Michael Spence's 1973 model explained why education pays even if it teaches nothing: a degree is costly to obtain, more costly for the less able, so it separates the types and employers rationally reward it. Crawford and Sobel's 1982 paper showed the mirror case: when talk is free and interests differ, what gets communicated is coarse at best, and the wider the difference in interests, the less can be said.

Why it works now

Reading a management team for costly actions rather than statements has always been what good analysts did by instinct on a handful of names. It can now be done systematically across filings, hiring data, pricing behavior, insider transactions, and contract terms, and then tested in interviews with operators who know what a real commitment looks like in that industry. The Desk sorts the deck into claims, signals, and commitments before the first Advisor is recruited.

When it applies

Management says the pipeline is strong, the seller says the customer is loyal, the founder says the round is oversubscribed; or the client is reading intent into statements that cost the speaker nothing.

What we would ask

What have they done that would have been foolish if the claim were false?
What have they done that they cannot easily undo?

What it needs

The actions taken, their cost to the actor, and their reversibility; a record of what was said against what was done.

How we would get it

Operators who can say what a real commitment looks like in this industry; the record (hiring, capital spend, contracts, insider purchases, pricing discipline) checked against the claims.

Where it breaks

The supposedly costly action was cheap for the sender; the receiver reads intent into noise.

For example

A CEO describes an extremely strong enterprise pipeline. The company has also hired forty salespeople, refused to discount in the quarter, and the CEO has bought stock. The sentence is talk; the three actions are the evidence, and the interviews test whether they were costly.

The mathematics in fullA signal separates types only when its cost differs across them, c(s, honest) < c(s, bluffing); a costless message is informative only when the speaker's interests are aligned with the listener's, which in a sale they are not. Taught in

MMSS 311-1; MECN-452; MECS 465

Sources

Spence (1973). Job Market Signaling. The Quarterly Journal of Economics 87(3), 355. https://doi.org/10.2307/1882010
Crawford, Sobel (1982). Strategic Information Transmission. Econometrica 50(6), 1431. https://doi.org/10.2307/1913390

Calibration and proper scoring rules

Was the seventy percent right seventy percent of the time?

We record what we said, in the words and bands we said it, and we score it when the outcome arrives, so that our judgment has a track record rather than a reputation.

Born to answer
Brier, 1950: scoring a forecaster's stated chance of rain.
What it does now
Making a research firm's judgment checkable years later.
The mathematics
Brier = mean of (p - outcome) squared
The full entry
Where it began

Glenn Brier's 1950 paper gave weather forecasters a score for probability forecasts: the squared distance between the stated chance of rain and whether it rained, averaged over many days. Meteorology adopted it and became, as a result, one of the few professions whose seventy percent means seventy percent.

Why it works now

Investment research has never been scored, because the forecasts were buried in prose and revised in memory. A brief whose ledger is written in explicit bands and preserved unaltered can be scored when the outcomes arrive, and the outcomes can now be captured from the record as they happen. This takes years to mature and Vista says so. It is the only route to a research firm whose judgment has a track record rather than a reputation.

When it applies

A research provider, an analyst, or an internal team makes confident calls with no record of how past calls turned out; or the client wants to know whether Vista's own bands mean anything.

What we would ask

Of the last twenty calls like this, how many resolved the way they were called?
Were the confident ones more often right than the hedged ones?

What it needs

Predictions preserved unaltered at the time they were made, with resolution dates and outcomes.

How we would get it

Every Vista brief keeps its ledger as written; outcomes are captured as they resolve; the score is reported to clients when the history is long enough to mean something, and not before.

Where it breaks

Forecasts revised after the fact; bands so wide they cannot be wrong; a score reported on a handful of cases.

For example

A provider says a merger is very likely to close. It closes. That is one observation. Twenty such calls with outcomes, scored, are a track record; the difference is why Vista keeps its ledgers.

The mathematics in fullBrier = (1/N) x sum of (p(i) - y(i))^2; calibration means that among forecasts near p, the realized frequency is near p. The library of predictions must be preserved to be scorable. Taught in

DECS-430; MECN-451

Sources

Brier (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review 78(1), 1-3. https://doi.org/10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2

Regression to the mean

Last year's best performer

Whatever was measured at an extreme was partly luck, and the luck does not repeat. We separate the part of last year's result that was skill from the part that was noise before anyone is rewarded, punished, or bought.

Born to answer
Galton, 1886: tall parents, somewhat less tall children.
What it does now
Not paying a premium for last year's luck.
The mathematics
Next = mean + r x (this value - mean)
The full entry
Where it began

Galton discovered it in 1886 measuring parents and children, and it has been rediscovered by every generation since. Tversky and Kahneman's 1971 paper on belief in the law of small numbers showed that trained scientists ignore it too, and Kahneman later told the story of flight instructors who believed praise made pilots worse and criticism made them better, because a great landing was followed by a worse one and a bad landing by a better one, whatever the instructor said.

Why it works now

Rewarding last year's best and punishing last year's worst has been organizational reflex because separating skill from noise required several periods of data for the unit and its peers, and nobody assembled them. The record can now be assembled quickly and the persistence of the measure estimated. The Desk asks how much the measure swings for reasons nobody controls, and how much of the recovery would have happened anyway.

When it applies

A manager, fund, region, product, or salesperson selected because of an extreme result; a turnaround credited to an intervention that followed a bad year; a rank list used to allocate.

What we would ask

How much does this measure vary from year to year for reasons nobody controls?
If you had done nothing after the bad year, how much of the recovery would have happened anyway?

What it needs

Several periods of the measure for the unit and its peers, enough to estimate how much of the variation is persistent.

How we would get it

The record for the measure across periods and peers; operators on what drives the year-to-year swings.

Where it breaks

Reading every reversal as an effect; punishing after bad luck and crediting the punishment.

For example

A fund allocates to the three regional managers with the best results last year. Their results were the most extreme of thirty, and the year-to-year correlation of the measure is low. Most of their advantage was noise, and next year's leaders will be different names.

The mathematics in fullExpected next value = mean + r x (this value - mean), where r is the correlation between periods; the further from the mean the observation, the larger the expected fall back, and with r near zero the extreme is almost entirely noise. Taught in

MATH 385; MATH 386-1

Sources

Galton (1886). Regression Towards Mediocrity in Hereditary Stature. The Journal of the Anthropological Institute of Great Britain and Ireland 15, 246. https://doi.org/10.2307/2841583
Tversky, Kahneman (1971). Belief in the law of small numbers. Psychological Bulletin 76(2), 105-110. https://doi.org/10.1037/h0031322

Probability encoding

How to ask an expert for a number

An expert asked for a forecast gives an anchored, overconfident one. We ask in a fixed order that starts from the extremes and checks for consistency, so that the range an Advisor gives us is one we can defend.

Born to answer
Spetzler and Stael von Holstein, 1975, at Stanford.
What it does now
Asking an Advisor for a range that is worth something.
The mathematics
Elicit the 10th and 90th fractiles before the 50th
The full entry
Where it began

Carl Spetzler and Carl-Axel Stael von Holstein's 1975 paper, from the Stanford decision analysis group, set out the protocol for encoding an expert's probability: fix the quantity precisely, elicit the extremes before the center, alternate the form of the question, check for consistency, and calibrate. Tversky and Kahneman's 1974 catalogue of anchoring and overconfidence is the reason the protocol exists. Ronald Howard's Stanford course still teaches it.

Why it works now

Vista's product is asking experts questions, and asking for a quantity badly produces a confident wrong answer. The protocol was known for fifty years and used by a handful of decision analysts, because applying it took training and time. It can now be built into every structured interview, with the fractiles asked in the right order and the consistency checks run as the conversation happens. The Desk asks whether the client's own experts start from a central number or from the extremes.

When it applies

Any engagement in which Advisors will be asked for ranges, probabilities, or timing; a client whose own experts have produced narrow ranges that keep being wrong.

What we would ask

When your team estimates a range, do they start from a central number or from the extremes?
How often do outcomes fall outside the ranges your experts give?

What it needs

A structured protocol applied consistently across Advisors; calibration checks on quantities with known answers.

How we would get it

Every Vista interview that asks for a quantity uses the protocol: the surprising values first, the central estimate last, consistency checks between fixed-value and fixed-probability questions.

Where it breaks

Anchoring on the first number mentioned; an expert with a stake in the answer; ranges reported without the protocol that produced them.

For example

Three operators are asked when a competitor's plant will reach full capacity. Asked directly, all three say eighteen months. Asked first for the date they would be surprised to see it beat, then the date they would be surprised to see it miss, the honest range runs from twelve to thirty-six months, and the thesis has to survive the whole range.

The mathematics in fullElicit the tenth, ninetieth, and then fiftieth fractiles before any central estimate; alternate fixed-value and fixed-probability questions; calibration means that about ten percent of outcomes should fall below the stated tenth fractile. Taught in

MS&E 252 (Stanford); MECN-451

Sources

Spetzler, Stael von Holstein (1975). Probability Encoding in Decision Analysis. Management Science 22(3), 340-358. https://doi.org/10.1287/mnsc.22.3.340
Tversky, Kahneman (1974). Judgment under Uncertainty: Heuristics and Biases. Science 185(4157), 1124-1131. https://doi.org/10.1126/science.185.4157.1124
Howard, Ronald A. and Abbas, Ali E. (2016). Foundations of Decision Analysis. Pearson. (book)

Strategic response

What the other side does when you move, and what that does to the plan.

Nash equilibrium

Assume the competitor is not asleep

We ask what a competent competitor does after you move, and whether your plan still works once they have.

Born to answer
Cournot, 1838, two sellers of mineral water; Nash, 1951, in general.
What it does now
What a competent rival does the day after you move.
The mathematics
No player gains by changing strategy alone
The full entry
Where it began

Antoine Augustin Cournot analyzed two sellers of mineral water in 1838 and found the outcome where each produces its best quantity given the other's, which John Nash generalized in 1951 to any number of players and any game. Bulow, Geanakoplos, and Klemperer's 1985 paper added the practical taxonomy: when moves are strategic substitutes a rival meets aggression by backing off, when they are complements by matching it, and the sign decides whether a plan works.

Why it works now

Every investment model that assumed passive competitors did so because modeling the response required knowing the rival's payoffs, and nobody outside the rival knew them. Former executives of the rival do, and they can be recruited. The Desk asks the two questions, what is the best thing your strongest competitor can do after you move and does the plan survive it, and the engagement puts the people who ran that competitor on the record.

When it applies

A plan (price rise, entry, new product, acquisition, channel change) whose model assumes rivals stay where they are; or a client puzzled by a rival's move.

What we would ask

If you do this, what is the best thing your strongest competitor can do in response?
Does your plan still work if they do it?

What it needs

The rivals' real payoffs and constraints, which are usually not the ones the client assumes; who moves first and who can observe whom.

How we would get it

Former executives of the competitors, who know how those companies actually decide; channel partners and customers who saw past responses.

Where it breaks

The game is mis-specified: wrong players, wrong moves, wrong information; or several stable outcomes exist and one was picked for convenience.

For example

A company plans a price increase on the assumption that its main rival follows. The rival's former head of pricing, interviewed, explains that the rival's incentive plan rewards share, not margin, and it will not follow. The plan is a different plan.

The mathematics in fullA set of strategies is stable when no player gains by changing alone, given the others; with strategic substitutes (quantities, capacity) a rival responds to aggression by retreating, with strategic complements (prices) by matching, and the sign decides the plan. Taught in

MMSS 211-2; MECN-452; MECN-441

Sources

Nash (1951). Non-Cooperative Games. The Annals of Mathematics 54(2), 286. https://doi.org/10.2307/1969529
Bulow, Geanakoplos, Klemperer (1985). Multimarket Oligopoly: Strategic Substitutes and Complements. Journal of Political Economy 93(3), 488-511. https://doi.org/10.1086/261312
Gibbons, Robert (1992). Game Theory for Applied Economists. Princeton University Press. (book)

Entry deterrence and reputation

Will they enter, and can you stop them?

We work out whether the incumbent's threats are credible, which usually means whether it has already spent money it cannot get back, and whether a reputation for fighting is real or a story.

Born to answer
Selten's chain store; Kreps and Wilson, 1982; Dixit, 1980.
What it does now
Whether a moat is sunk commitment or a story.
The mathematics
A threat is credible only if carrying it out is optimal
The full entry
Where it began

Reinhard Selten's chain-store paradox asked why an incumbent would fight entrants when fighting costs more than accommodating, and backward induction said it never would. Kreps and Wilson, and Milgrom and Roberts, resolved it in 1982: a small chance that the incumbent is genuinely tough makes fighting early entrants worthwhile as reputation. Avinash Dixit in 1980 showed the other route, capacity built before entry that cannot be unbuilt, which turns a threat into a fact.

Why it works now

Whether a moat is real used to be argued from the incumbent's rhetoric. What can now be established is what the incumbent has actually sunk, what it did the last three times someone entered, and what each response cost it, from the record and from former executives of both the incumbent and the entrants. The Desk asks what has been committed that cannot be reversed. The interviews answer it.

When it applies

A client is deciding whether to enter a market with an aggressive incumbent, or is an incumbent deciding how to respond to entry, or is underwriting a company whose moat is an entry barrier.

What we would ask

What has the incumbent already committed that it cannot reverse?
When someone last entered, what did the incumbent actually do, and did it pay for it?

What it needs

The incumbent's sunk commitments, its history of responses to entry, and what those responses cost it.

How we would get it

Former executives of the incumbent and of past entrants; the record of past entries and the price and capacity responses that followed.

Where it breaks

The threat is not credible because carrying it out would hurt the incumbent more than accommodating; or the moat was regulatory and the regulation is changing.

For example

A venture investor is underwriting an entrant against a dominant incumbent said to crush competition. The last three entrants were bought, ignored, and out-priced respectively; the interviews establish which of those this entrant resembles and what the incumbent's capacity position makes credible now.

The mathematics in fullA threat is credible only when carrying it out is the incumbent's best response at the point of entry; irreversible investment before entry changes that best response, and a reputation for toughness is worth more the more entrants remain to be deterred. Taught in

MMSS 311-1; MECN-441; MECS 449

Sources

Dixit (1980). The Role of Investment in Entry-Deterrence. The Economic Journal 90(357), 95. https://doi.org/10.2307/2231658
Milgrom, Roberts (1982). Predation, reputation, and entry deterrence. Journal of Economic Theory 27(2), 280-312. https://doi.org/10.1016/0022-0531(82)90031-X
Besanko, Dranove, Shanley and Schaefer. Economics of Strategy, 7th ed. Wiley. (book)

Repeated games and tacit collusion

Why the truce holds, and when it breaks

We look at what keeps rivals from cutting price: how much future business they expect from each other, how well they can see each other's moves, and whether anyone is about to leave the game.

Born to answer
Green and Porter, 1984, on a nineteenth-century railroad cartel.
What it does now
What holds pricing discipline, and who is about to leave the game.
The mathematics
Cooperation holds while gain G is below the discounted future loss
The full entry
Where it began

Green and Porter's 1984 model explained price wars without anyone cheating: when rivals cannot see each other's prices and only observe a noisy market, a bad quarter looks like a secret price cut, and the punishment phase that follows is what keeps the truce credible the rest of the time. The theory was built with a nineteenth-century railroad cartel in mind, whose periodic price wars Porter had studied.

Why it works now

Underwriting pricing discipline in a concentrated industry has always meant guessing at three things: how transparent prices are among rivals, how long each expects to be in the game, and whether anyone is about to leave. All three are now questions former pricing and sales leaders across the rivals can answer, and the record shows what preceded every past price break. The Desk asks who is about to sell, retire, or exit, because that is what shortens the shadow of the future.

When it applies

An industry with stable prices that the thesis assumes will stay stable; or a price war the client did not expect; or a rival that is about to be sold, exit, or change hands.

What we would ask

How quickly would the others see a secret price cut, and what would they do?
Is anyone in this market about to leave it, sell, or retire?

What it needs

Transparency of pricing among rivals, the horizon each expects, and the events that shorten it.

How we would get it

Former pricing and sales leaders across the rivals; the record of past price breaks and what preceded them.

Where it breaks

The end of the game is nearer than assumed; punishment is not credible; deviations cannot be observed clearly enough to punish.

For example

A thesis rests on rational pricing in a four-player market. One of the four is being prepared for sale, which shortens its horizon and changes its incentives in the last year. The interviews with former executives of that player are where the thesis is tested.

The mathematics in fullCooperation holds when the one-period gain from cutting, G, is at most the discounted future loss delta x (V(cooperate) - V(punish)) / (1 - delta); anything that lowers delta or blurs observation of the cut breaks the truce. Taught in

MMSS 311-1; MECN-441; MECN-452

Sources

Green, Porter (1984). Noncooperative Collusion under Imperfect Price Information. Econometrica 52(1), 87. https://doi.org/10.2307/1911462
Fudenberg, Maskin (1986). The Folk Theorem in Repeated Games with Discounting or with Incomplete Information. Econometrica 54(3), 533. https://doi.org/10.2307/1911307
Kreps, Wilson (1982). Reputation and imperfect information. Journal of Economic Theory 27(2), 253-279. https://doi.org/10.1016/0022-0531(82)90030-8

Nash bargaining and alternating offers

The outside option decides the split

We establish what each side gets if talks fail, because that, not the argument in the room, decides most negotiations, and we research the other side's outside option as hard as yours.

Born to answer
Nash, 1950; Rubinstein, 1982: patience decides the split.
What it does now
Researching the other side's fallback, not only your own.
The mathematics
Maximize (u1 - d1) x (u2 - d2)
The full entry
Where it began

John Nash's 1950 paper solved the bargaining problem from four axioms and found that the split depends on what each side gets if talks fail. Ariel Rubinstein's 1982 model of alternating offers showed the same thing dynamically: the more patient party captures more of the surplus, and impatience is expensive. Kellogg teaches the Nash scheme in its game theory course as a tool for actual negotiations.

Why it works now

Negotiators have always known their own fallback and guessed the other side's. The other side's fallback is now researchable: former executives and advisors from that side's world can say what a failed deal actually costs a party in that position, and the record shows the alternatives and the timing pressure. The Desk asks who is more able to wait. The engagement finds out before the first meeting instead of after the last.

When it applies

A client preparing a negotiation (acquisition, partnership, supply contract, renewal, joint venture) who knows their own fallback and is guessing the other side's.

What we would ask

What happens to them if this does not close? Not what they say, what actually happens.
Who is more able to wait?

What it needs

Both sides' fallbacks, patience, and ability to commit, established from evidence rather than posture.

How we would get it

Former executives and advisors from the other side's industry on what a failed deal costs a party in that position; the record on the other side's alternatives and timing pressure.

Where it breaks

The outside options are misread (the seller does have another buyer, or does not); a commitment was not credible.

For example

A buyer believes a founder must sell. The founder's former CFO, interviewed, explains that the company just refinanced and the founder can run it for years. The buyer's leverage was imaginary, and the negotiation plan changes before the first meeting.

The mathematics in fullThe split maximizes (u1 - d1) x (u2 - d2), where d1 and d2 are the payoffs if no deal is reached; in alternating offers the more patient side captures more, so improving your own d or worsening theirs moves the outcome more than any concession. Taught in

MECN-452; MECS 540-2; MMSS 311-1

Sources

Nash (1950). The Bargaining Problem. Econometrica 18(2), 155. https://doi.org/10.2307/1907266
Rubinstein (1982). Perfect Equilibrium in a Bargaining Model. Econometrica 50(1), 97. https://doi.org/10.2307/1912531
Muthoo, Abhinay (1999). Bargaining Theory with Applications. Cambridge University Press. (book)

The Myerson-Satterthwaite theorem

Deals that should close and do not

When both sides would gain from a deal and it still is not closing, the cause is usually private information about value or a reputation for holding out, and we design the research to find which.

Born to answer
Myerson and Satterthwaite, 1983: some worthwhile trades must fail.
What it does now
Diagnosing a stalled deal that makes obvious economic sense.
The mathematics
No mechanism guarantees trade under two-sided private values
The full entry
Where it began

Myerson and Satterthwaite proved in 1983 that when a buyer and a seller each privately know their own value, no bargaining procedure can guarantee trade whenever trade is worth doing; some good deals must fail. James Fearon's 1995 paper applied the same logic to war: rational states fight because of private information they have reason to misrepresent and because they cannot commit. Abreu and Gul showed in 2000 how a reputation for stubbornness lets a party win by waiting.

Why it works now

A stalled deal used to be read as irrationality or bad faith. It can now be diagnosed: what each side privately knows, what each has committed to in public, and what waiting costs each of them are all things former participants in comparable stalemates can speak to, and the record shows the public positions. The Desk asks which side's cost of delay is real. That tells the client who folds first and when.

When it applies

A transaction that makes obvious economic sense is stalled; a negotiation has become a waiting game; or a client wants to know whether a stalled deal will close.

What we would ask

What does each side know about the value that the other does not?
Is someone holding out because folding once would cost them in every future deal?

What it needs

What each side privately knows, what each has committed to publicly, and each side's cost of continued delay.

How we would get it

Advisors who have been on the other side of comparable stalemates; the record of the parties' public positions and the costs of waiting.

Where it breaks

The efficiency of the deal is assumed to guarantee it closes; the hold-out is read as irrational when it is reputational.

For example

A strategic buyer and a target agree the combination creates value and have not closed in nine months. Each believes the other is overstating its walk-away. The interviews establish which side's cost of delay is real, which tells the client who folds first and when.

The mathematics in fullWith private values on both sides, no mechanism guarantees efficient trade whenever a deal is worth doing; delay is the cost that separates types, and a party with a reputation for holding out can win by waiting even when waiting is expensive. Taught in

MECS 465; MECS 540-2

Sources

Myerson, Satterthwaite (1983). Efficient mechanisms for bilateral trading. Journal of Economic Theory 29(2), 265-281. https://doi.org/10.1016/0022-0531(83)90048-0
Abreu, Gul (2000). Bargaining and Reputation. Econometrica 68(1), 85-117. https://doi.org/10.1111/1468-0262.00094
Fearon (1995). Rationalist explanations for war. International Organization 49(3), 379-414. https://doi.org/10.1017/S0020818300033324

Global games and bank runs

Everyone runs because everyone else might

We look at whether the outcome depends on what each participant believes the others will do, in which case small changes in confidence can flip a whole market, and we find the participants whose beliefs matter most.

Born to answer
Diamond and Dybvig, 1983; Carlsson and van Damme, 1993.
What it does now
Whose beliefs about everyone else's beliefs decide a run.
The mathematics
Each participant acts when its estimate crosses a threshold
The full entry
Where it began

Carlsson and van Damme's 1993 paper on global games showed that when players have slightly different information about the fundamentals, a coordination game with many possible outcomes collapses to one: each player acts when its own estimate crosses a threshold. Morris and Shin applied it to currency attacks, and Baliga and Sjöström's 2004 work on arms races formalized Schelling's reciprocal fear of surprise attack, where each side arms because it fears the other might.

Why it works now

Runs and panics were treated as sentiment because the decision rules of the large participants were unknowable. They are knowable: treasurers, portfolio managers, and operators at the kind of institutions that hold the deposits, the bonds, or the platform positions can be asked what they would do and what public news would change what they believe about each other. The Desk finds the participants whose beliefs matter most. The engagement asks them.

When it applies

A run, a panic, a rush for the exit, a standards war, a platform tipping point; or a thesis that assumes a stable equilibrium in a market with two.

What we would ask

If the largest three participants believed the others would stay, would they stay?
What single piece of public news would change what everyone believes everyone else believes?

What it needs

The participants' payoffs from staying versus leaving as a function of how many others leave; what information is common to all of them.

How we would get it

Operators on the actual decision rules of the large participants; the record on comparable coordination breaks and what triggered them.

Where it breaks

The model assumes participants know each other's payoffs precisely when they only estimate them; multiple equilibria are collapsed to one.

For example

A thesis on a regional lender assumes deposits are sticky. The question is not the average depositor's loyalty but whether the largest depositors believe the others will stay; the interviews go to treasurers of the kind of companies that hold those deposits.

The mathematics in fullWith a small amount of private information about fundamentals, the many equilibria of a coordination game collapse to one threshold: each participant acts when its estimate of the fundamental crosses a cutoff, so the size of the crowd that moves is a function of public news, not of mood. Taught in

MECS 540-2; MMSS 311-1

Sources

Carlsson, van Damme (1993). Global Games and Equilibrium Selection. Econometrica 61(5), 989. https://doi.org/10.2307/2951491
Baliga, Sjostrom (2004). Arms Races and Negotiations. Review of Economic Studies 71(2), 351-369. https://doi.org/10.1111/0034-6527.00287
Diamond, Dybvig (1983). Bank Runs, Deposit Insurance, and Liquidity. Journal of Political Economy 91(3), 401-419. https://doi.org/10.1086/261155

Commitment and audience costs

A threat is only as good as its cost to withdraw

We test whether a commitment (a price floor, a no-sale promise, a walk-away, a regulatory pledge) would actually be kept when the moment comes, which depends on what breaking it would cost, not on how firmly it was said.

Born to answer
Schelling, 1960: burning the bridges behind the army.
What it does now
Whether a pledge by a counterparty or a government will hold.
The mathematics
Credible when reversal costs more than honoring
The full entry
Where it began

Thomas Schelling's 1960 book turned commitment into strategy: the general who burns the bridges behind his army, the negotiator who ties his own hands, the focal point two strangers pick when told to meet in New York without communicating. James Fearon's 1994 paper added audience costs: a leader who makes a public threat pays a domestic price for backing down, which is exactly what makes the threat believable.

Why it works now

Whether a government, a regulator, or a counterparty will honor a pledge used to be judged from the firmness of its language. The cost of reversal is now the thing to research, and it can be: who the pledge was made in front of, what reversing it has cost that party before, and what it would cost now, from former officials and from the record. The Desk asks what backing down would cost them, and treats the answer, not the rhetoric, as the probability.

When it applies

A plan relies on a promise or a threat by the client, a counterparty, a regulator, or a government; or a client is trying to make its own commitment believable.

What we would ask

What would it cost them, publicly and privately, to back down?
Have they backed down from something like this before?

What it needs

The audience in front of which the commitment was made and the cost of reversal before it; the history of reversals.

How we would get it

Advisors who have watched the party in question honor or abandon commitments; the record of prior reversals and their consequences.

Where it breaks

Firmness of language is mistaken for cost of reversal; the audience that would punish a reversal is assumed rather than identified.

For example

A government has pledged a subsidy regime through the decade. The pledge was made to voters and investors before an election. The interviews establish what reversing it would cost after the election, which is the actual probability the thesis depends on.

The mathematics in fullA commitment is credible when reversing it is more costly than honoring it at the moment of decision; audience costs, sunk investment, and reputation are the three ways that cost is created, and each can be researched. Taught in

MECS 540-2; MMSS 211-3

Sources

Schelling, Thomas C. (1960). The Strategy of Conflict. Harvard University Press. (book)
Fearon (1994). Domestic Political Audiences and the Escalation of International Disputes. American Political Science Review 88(3), 577-592. https://doi.org/10.2307/2944796

Replicator dynamics and stable strategies

Which practices survive in an industry

Instead of asking what a rational competitor would do, we ask which strategies are gaining share in the population of firms and which are losing it, because an industry converges on what works whether or not anyone planned it.

Born to answer
Maynard Smith and Price, 1973: why animal contests rarely end in death.
What it does now
Which business practices gain share whether or not anyone planned it.
The mathematics
dx(i)/dt = x(i) x (payoff of i minus the average)
The full entry
Where it began

John Maynard Smith and George Price asked in 1973 why animal contests are so rarely fights to the death, and answered with the evolutionarily stable strategy: a mix of behavior that no rare alternative can invade. Taylor and Jonker wrote the dynamics down in 1978 as the replicator equation, in which a strategy's share grows when it does better than the population average. MMSS teaches it in both game theory courses because it explains convergence without assuming anyone is rational.

Why it works now

Watching which practices gain and lose share across an industry used to require a trade association's survey and a decade. The share of each strategy in a market can now be read from filings, job postings, pricing pages, and the trade press year by year, and operators can say which practices are being copied. The Desk asks which practices are gaining share among the firms in this market. The population answers the question the argument cannot.

When it applies

A thesis that assumes rivals stay irrational or old-fashioned; a practice spreading through an industry; a question about whether a business model will be copied or will fade.

What we would ask

Which practices are gaining share among the firms in this market, and which are losing it?
Could a single entrant doing something different take share from the incumbents, or would it be squeezed out?

What it needs

The mix of strategies in the industry over time, and the relative performance of each in the current mix.

How we would get it

Operators across the industry on which practices are winning and being copied; the record for share by strategy over time.

Where it breaks

Payoffs treated as fixed when they shift as the mix changes; entry and imitation ignored.

For example

A thesis assumes an incumbent's high-touch sales model will keep beating a rival's self-serve model. Across the industry, self-serve is gaining share every year and the high-touch firms are the ones being acquired or shrinking. The population, not the argument, says where this ends.

The mathematics in fulldx(i)/dt = x(i) x (f(i)(x) - fbar(x)): the share of a strategy grows when its payoff exceeds the average; a strategy that cannot be invaded by a rare alternative is where the industry settles. Taught in

MMSS 211-2; MMSS 311-1

Sources

Maynard Smith, Price (1973). The Logic of Animal Conflict. Nature 246(5427), 15-18. https://doi.org/10.1038/246015a0
Taylor, Jonker (1978). Evolutionary stable strategies and game dynamics. Mathematical Biosciences 40(1-2), 145-156. https://doi.org/10.1016/0025-5564(78)90077-9

Private information and incentives

What the seller knows, what the bidder should fear, what a contract will actually reward.

Adverse selection

Why is this for sale?

We treat the fact that something is offered as evidence about it. Owners who know their asset is bad sell readily; so the research asks what the seller knows and why they are selling now.

Born to answer
Akerlof, 1970: the used car market unravels.
What it does now
Reading the fact that something is for sale as evidence about it.
The mathematics
Quality offered = E[q | the seller accepts p]
The full entry
Where it began

George Akerlof's 1970 paper started with used cars: owners know which cars are lemons and buyers do not, so buyers pay for average quality, owners of good cars withdraw, quality falls, and the market can vanish. Rothschild and Stiglitz showed in 1976 that insurers face the same problem and respond with menus of contracts that make customers reveal their risk by what they choose.

Why it works now

Diligence on a company for sale used to be a data-room exercise, which is precisely the information the seller chose to show. The seller's private information and motive are now researchable from outside the room: former employees and advisors of the seller, prior sale attempts, the buyers who walked and what they saw. The Desk asks why now, and who passed. The engagement is built to find out what the seller knows.

When it applies

An asset, a company, a loan book, a portfolio, or a customer contract is on offer, and the seller knows more about it than the buyer can learn from the data room.

What we would ask

Why is the seller selling now, and what has changed for them that the data room does not show?
Who else looked and passed, and what did they see?

What it needs

The seller's private information and motives; the history of the asset with previous owners and previous bidders.

How we would get it

Former employees and advisors of the seller; the record of prior sale attempts, withdrawn processes, and the buyers who walked.

Where it breaks

The seller has a legitimate reason to sell that the analysis dismisses; or the buyer assumes the data room is the whole truth.

For example

A specialty lender is selling a loan book at an attractive yield. The former head of underwriting, interviewed, explains the vintage that was originated under a relaxed policy in the year before the sale. The yield is the price of that vintage.

The mathematics in fullAt any price p, the quality offered is E[q | seller accepts p], which falls as p falls; the buyer who ignores this pays for the average and receives the below-average, and the market can unravel entirely. Taught in

MMSS 311-1; MECN-452; MECS 465

Sources

Akerlof (1970). The Market for "Lemons": Quality Uncertainty and the Market Mechanism. The Quarterly Journal of Economics 84(3), 488. https://doi.org/10.2307/1879431
Rothschild, Stiglitz (1976). Equilibrium in Competitive Insurance Markets: An Essay on the Economics of Imperfect Information. The Quarterly Journal of Economics 90(4), 629. https://doi.org/10.2307/1885326

The winner's curse

Why are you the one winning?

In a competitive process, winning is itself evidence that your estimate was the most optimistic in the room. We find out who else looked, who passed, and at what price, before you congratulate yourself.

Born to answer
Capen, Clapp and Campbell, 1971: Gulf of Mexico lease auctions.
What it does now
Asking who looked and passed before celebrating a win.
The mathematics
E[V | you won] is below E[V | your estimate]
The full entry
Where it began

Three petroleum engineers at Atlantic Richfield, Capen, Clapp, and Campbell, published the winner's curse in 1971 after noticing that the companies winning Gulf of Mexico lease auctions earned poor returns on them: the winner is the bidder whose estimate was most optimistic. Richard Thaler's 1988 account includes the classroom version, an auction for a jar of coins, in which the winner reliably overpays.

Why it works now

The correction requires knowing how many sophisticated parties looked and why the ones who passed did, which used to be gossip. It is now obtainable: advisors to other bidders in the same or comparable processes can be recruited, the process record shows the withdrawals, and the reasons can be put on the record. The Desk asks why you are the one winning. The engagement finds what the two bidders who walked after diligence saw.

When it applies

A competitive auction or financing where the client is the high bidder, or is about to be; or a client puzzled that it keeps winning the deals it wants.

What we would ask

How many sophisticated parties evaluated this, and how many passed?
What did the ones who passed know?

What it needs

The number and quality of other bidders, their reasons for passing, and the dispersion of estimates.

How we would get it

Advisors who advised other bidders in the same or comparable processes; the record of the process and its withdrawals.

Where it breaks

The bidders are treated as having independent private values when the asset has one true value everyone is estimating; the client's own optimism is not modeled.

For example

A fund wins a competitive process for an industrial asset against six bidders. Two withdrew after diligence on the same environmental issue the fund's advisors called immaterial. The interviews go to people who saw what those two saw.

The mathematics in fullWith a common value V and noisy estimates, E[V | you won] is less than E[V | your estimate], and the gap grows with the number of bidders; the correction is to bid as if you have already learned that everyone else estimated lower. Taught in

MECN-452; DECS-430; MMSS 311-1

Sources

Capen, Clapp, Campbell (1971). Competitive Bidding in High-Risk Situations. Journal of Petroleum Technology 23(06), 641-653. https://doi.org/10.2118/2993-PA
Thaler (1988). Anomalies: The Winner's Curse. Journal of Economic Perspectives 2(1), 191-202. https://doi.org/10.1257/jep.2.1.191
Milgrom, Weber (1982). A Theory of Auctions and Competitive Bidding. Econometrica 50(5), 1089. https://doi.org/10.2307/1911865

Auction theory

The format decides who wins and what they pay

Whether a client is selling or bidding, we work out how the rules of the process shape the result: sealed or open, one round or several, reserve or none, and how the other side will bid under them.

Born to answer
Vickrey, 1961; Myerson, 1981: the format decides the outcome.
What it does now
A company's exposure to the rules of auctions it runs or enters.
The mathematics
Truthful bidding is optimal under a second price
The full entry
Where it began

William Vickrey's 1961 paper introduced the sealed-bid second-price auction, in which bidding your true value is the best strategy, and showed that different formats can raise the same expected revenue. Roger Myerson's 1981 paper characterized the revenue-maximizing auction and earned its share of a Nobel. Edelman, Ostrovsky, and Schwarz showed in 2007 that the search-advertising auctions run by Google and Yahoo were a generalized second-price format with their own quirks.

Why it works now

Auction theory was applied to spectrum and search advertising by the few firms that could hire the theorists. What has changed for an investor is that a company's exposure to auction rules, as a seller of inventory, a bidder for contracts, or a platform running the auction, can now be read from the record and tested with people who have run comparable processes. The Desk asks whether the bidders know their own values or are all guessing at the same one, because the two cases behave differently.

When it applies

A client is designing a sale process, entering one, bidding for spectrum or ad inventory or a contract, or underwriting a business whose revenue comes from auctions it runs or bids in.

What we would ask

Do the bidders each know their own value, or are they all estimating the same unknown?
Can the bidders talk to each other, and can the seller change the rules midway?

What it needs

Whether values are private or common; the number of bidders; the seller's ability to commit to the rules; the possibility of collusion.

How we would get it

Advisors who have run or advised in comparable processes; the record of comparable auctions and their outcomes.

Where it breaks

Bidders collude; the seller cannot commit; the format assumes rational bidders in a room with none.

For example

A company's revenue is set in online ad auctions. Its thesis assumes those auctions are competitive. The former head of an ad exchange, interviewed, explains how a change in the auction format two years earlier moved margin from the platform to the bidders, and what the next change would do.

The mathematics in fullTruthful bidding is a best response in a second-price format and not in a first-price one; under standard assumptions formats that allocate to the same bidder raise the same expected revenue, so the seller's lever is the reserve and the entry of bidders, not the format. Taught in

MECN-452; MECN-446; MECS 465

Sources

Vickrey (1961). Counterspeculation, Auctions, and Competitive Sealed Tenders. The Journal of Finance 16(1), 8-37. https://doi.org/10.1111/j.1540-6261.1961.tb02789.x
Myerson (1981). Optimal Auction Design. Mathematics of Operations Research 6(1), 58-73. https://doi.org/10.1287/moor.6.1.58
Edelman, Ostrovsky, Schwarz (2007). Internet Advertising and the Generalized Second-Price Auction: Selling Billions of Dollars Worth of Keywords. American Economic Review 97(1), 242-259. https://doi.org/10.1257/aer.97.1.242

Moral hazard and the informativeness principle

The plan rewards what it measures

We read every incentive plan, earnout, and management contract by asking what measured thing it rewards and what unmeasured thing it therefore sacrifices, because that is what the people under it will do.

Born to answer
Holmstrom, 1979: paying an agent whose effort you cannot see.
What it does now
Reading an earnout or a pay plan for what it will actually reward.
The mathematics
Include a signal only if it informs about effort
The full entry
Where it began

Bengt Holmström's 1979 paper set up the problem of paying an agent whose effort cannot be seen and proved the informativeness principle: any signal that carries information about the effort, however noisy, belongs in the contract, and any that does not should be left out. The trade-off between insurance and incentives that every compensation plan embodies is the one his model made explicit.

Why it works now

Reading an earnout, a management plan, or a supplier contract for what it will actually reward has always been possible and rarely done thoroughly, because the terms were long and the second-order effects took an expert to see. The terms can now be read completely, the measured quantity identified, and the question put to former executives who worked under comparable plans: what would a clever person do to hit this number. The Desk asks what is measured and who controls the measurement.

When it applies

An earnout, a management incentive plan, a sales compensation change, a supplier contract, or a portfolio company whose behavior the client cannot explain until they read the pay plan.

What we would ask

What exactly is measured, and who controls the measurement?
What would a clever person do to hit the number without doing the thing you wanted?

What it needs

The full contract terms, the measurement process, and the noise in the measure relative to the effort it is meant to reward.

How we would get it

Former executives who worked under comparable plans; operators on how such measures are gamed in this industry; the record of past plans and what followed them.

Where it breaks

The measure is gameable, the agent controls it, or the plan rewards the measured task at the expense of the unmeasured one.

For example

A portfolio company's management is paid on annual EBITDA. Maintenance capex has fallen for three years. The two facts are one fact, and the interviews with former plant managers establish what the deferred maintenance will cost.

The mathematics in fullPay w(y) on an observable y to maximize E[y - w(y)] subject to the agent choosing effort to maximize E[u(w(y)) - c(e)]; a noisy measure with a risk-averse agent forces weak incentives, and a signal belongs in the contract only if it adds information about effort. Taught in

MECS 465; MECN-441

Sources

Holmstrom (1979). Moral Hazard and Observability. The Bell Journal of Economics 10(1), 74. https://doi.org/10.2307/3003320
Salanie, Bernard (2005). The Economics of Contracts: A Primer, 2nd ed. MIT Press. (book)

Screening and second-degree price discrimination

Let them sort themselves

When you cannot tell the types apart, we design the offer so that each type picks the version meant for it, and we test whether the menu you have is doing that or leaking.

Born to answer
Mussa and Rosen, 1978: a menu that makes each type reveal itself.
What it does now
The tier valuable customers take because its limits never bind.
The mathematics
Distort the low type's quality to deter imitation
The full entry
Where it began

Mussa and Rosen's 1978 paper showed how a seller who cannot tell customers apart offers a menu of qualities and prices so that each type selects the version meant for it, and found the cost of doing so: the cheap version must be made worse than it needs to be, so the valuable customer will not take it. Rothschild and Stiglitz found the same structure in insurance the same decade.

Why it works now

Whether a pricing menu is sorting customers or leaking value could only be inferred from aggregate take-up. Customer-level data can now be read for who is taking which tier, and the customers themselves can be asked which limit would actually bind for them. The Desk asks which valuable customers are buying the cheap version and why they can live with it. That is a question for the customers, not the product team, and the engagement asks them.

When it applies

Pricing tiers, insurance plans, loan products, term sheets, or service levels where the seller cannot observe which customer is which and the wrong customers are picking the wrong tier.

What we would ask

Which customers are taking the cheap version who would have paid for the expensive one?
What is it about the cheap version that the valuable customers can live with?

What it needs

The distribution of customer types, what each values, and the current take-up of each tier by type.

How we would get it

Customers and former sales leaders on who buys what and why; a designed choice exercise where the current data cannot separate the types.

Where it breaks

Arbitrage undoes the segmentation; the low tier is too good, so the high type takes it; competitors offer the high type a better menu.

For example

A software company's mid tier is taken by enterprises that would have paid for the top tier, because the mid tier's limits do not bind for them. The research is which limit would bind, which is a question for the customers, not the product team.

The mathematics in fullA menu of (quality, price) pairs is designed so each type prefers its own; the low type's version is distorted downward, below what it would get alone, to keep the high type from imitating, and the high type earns an information rent. Taught in

MECN-446; MECS 465; MMSS 211-1

Sources

Mussa, Rosen (1978). Monopoly and product quality. Journal of Economic Theory 18(2), 301-317. https://doi.org/10.1016/0022-0531(78)90085-6
Rothschild, Stiglitz (1976). Equilibrium in Competitive Insurance Markets: An Essay on the Economics of Imperfect Information. The Quarterly Journal of Economics 90(4), 629. https://doi.org/10.2307/1885326

Markets, pricing, and choice

What people will choose, what they will pay, and how a market is really shaped.

Discrete choice and conjoint analysis

Do not ask what they would pay

Asking people what they would pay is close to worthless. We make them choose between realistic alternatives and read what they give up, which is where the willingness to pay actually shows.

Born to answer
McFadden's BART forecast; Green and Srinivasan, 1978.
What it does now
Reading willingness to pay from trade-offs instead of asking.
The mathematics
P(j) = exp(V(j)) / sum over k of exp(V(k))
The full entry
Where it began

Daniel McFadden built discrete choice modeling in the early 1970s and used it to forecast ridership on the Bay Area Rapid Transit system before it opened, from how people chose among the alternatives they already had; the forecast held up. Green and Srinivasan's 1978 review brought the same logic to consumer research as conjoint analysis: present realistic bundles, record the choices, and recover what each attribute is worth.

Why it works now

A proper choice study needed a specialist to design and estimate it and a panel to field it, so it was reserved for consumer products with large budgets. The design is now the cheap part: Advisors and customers define the real alternatives and attributes in the interviews, and the Desk can explain the move in one sentence. Vista designs the study and brings in a specialist to field and estimate it. It does not run a survey business, and it says so.

When it applies

A pricing, packaging, or product decision rests on a survey of stated intentions, on management's belief about value, or on a handful of customer conversations that asked the direct question.

What we would ask

Who exactly is the buyer, and what are the real alternatives in front of them, including doing nothing?
Which two or three attributes do they actually trade off against price?

What it needs

A defined buyer population, a realistic set of alternatives and attributes, and enough respondents from the actual buyer population to estimate trade-offs.

How we would get it

Advisors and customers define the attributes and alternatives; a designed choice exercise among qualified buyers produces the trade-offs; Vista designs it and brokers the fielding rather than running a survey business.

Where it breaks

Stated choices differ from purchases; the sample is not the buyer; the price range shown anchors the answers; the attributes are the product team's, not the buyer's.

For example

A medical device company believes hospitals will pay a premium for a feature. Hospital purchasing leads, presented with realistic bundles, trade that feature away for service response time every time. The premium is in the wrong place.

The mathematics in fullP(choose j) = exp(V(j)) / sum over k of exp(V(k)), with V(j) = beta x attributes(j); willingness to pay for an attribute is the ratio -beta(attribute) / beta(price), estimated from choices rather than asked. Taught in

MMSS 386-2 (qualitative-choice models); MECN-446

Sources

Green, Srinivasan (1978). Conjoint Analysis in Consumer Research: Issues and Outlook. Journal of Consumer Research 5(2), 103-123. https://doi.org/10.1086/208721
Train (2009). Discrete Choice Methods with Simulation. Cambridge University Press (book). https://doi.org/10.1017/CBO9780511805271

Price elasticity of demand

How much volume does a price change buy or lose?

Pricing power is a number, the volume lost per point of price, and it is almost never known and almost always asked about directly, which does not work. We infer it from what happened when prices actually moved.

Born to answer
Marshall, 1890, and every intermediate microeconomics course since.
What it does now
Inferring elasticity from what happened when prices actually moved.
The mathematics
(P - MC) / P = -1 / e
The full entry
Where it began

Alfred Marshall named elasticity in 1890, and every intermediate microeconomics course since, including MMSS Turbo Micro, has taught that a seller's power is the volume lost per point of price. The classic measurement failure is as old as the concept: asking customers what they would pay produces an answer, and the answer is wrong.

Why it works now

The elasticity at the current price is now inferable rather than asked: past price moves can be matched to customer-level volume responses in the client's own data, competitor responses can be read from the record, and category buyers can be asked what the next increase would trigger. The Desk asks what happened to volume the last two times prices moved, and who left. The engagement puts the largest customer's buyer on the record.

When it applies

A thesis rests on the company's ability to raise prices; a company is deciding on a price change; or a market's pricing discipline is being underwritten from the outside.

What we would ask

When did prices last move, and what happened to volume, customer by customer if possible?
Who left, and where did they go?

What it needs

Past price changes with volume responses, ideally at the customer level; competitor responses to those changes; the customers' alternatives.

How we would get it

Former pricing leaders and category buyers on real responses to past increases; the record of price moves and share shifts; a designed choice exercise where the history is silent.

Where it breaks

The elasticity is estimated from a survey; competitors respond and the estimate assumed they would not; the past increase coincided with something else.

For example

An investor underwrites a consumables business as having pricing power. The last two increases held volume, but the category buyer at the largest customer, interviewed, explains that the third will trigger a dual-sourcing program already approved. The power is real and it is nearly used up.

The mathematics in fullElasticity e = (dQ/Q) / (dP/P); a single profit-maximizing price satisfies (P - MC) / P = -1/e, so the whole question is e at the current price, which is inferred from natural experiments, discontinuities, or designed choices, never from asking. Taught in

MMSS 211-1; MECN-430; MECN-446

Sources

Frank, Robert H. Microeconomics and Behavior. McGraw-Hill. (book)
Besanko, David and Braeutigam, Ronald. Microeconomics, 5th ed. Wiley. (book)

Bundling and versioning

One price leaves money on the table

We work out which customers value what, and whether versions, bundles, or quantity terms would capture value a single price leaves behind, and whether the customers could undo it by trading among themselves.

Born to answer
Adams and Yellen, 1976: why a bundle beats two separate prices.
What it does now
Capturing the value that a single price leaves behind.
The mathematics
Bundle when component values are negatively correlated
The full entry
Where it began

Adams and Yellen's 1976 paper showed with a handful of customers and two goods why bundling raises profit when customers who value one good highly value the other little: the bundle captures value that two separate prices leave behind. Stigler had seen the same logic earlier in the block booking of films. Mussa and Rosen supplied the theory of versions.

Why it works now

The correlation of values across components, which decides whether a bundle works, was invisible because nobody could ask enough customers. It is now the object of a designed choice exercise among qualified buyers, informed by interviews that say which segments exist. The Desk asks who gets the most out of the product and how you would tell them apart at the point of sale. The engagement sizes the segments.

When it applies

A single price for a product whose customers value it very differently; a bundle whose logic nobody can state; or a competitor's tiering that is winning.

What we would ask

Which customers get the most out of this, and how would you tell them apart at the point of sale?
Could the customers who pay less resell to the ones who pay more?

What it needs

The distribution of value across customers and the correlation of values across components; the feasibility of keeping segments apart.

How we would get it

Customers and former sales leaders on value by segment; a choice exercise across candidate bundles among qualified buyers.

Where it breaks

Arbitrage; segments that cannot be identified at the point of sale; competitors who unbundle.

For example

A data vendor sells one enterprise license. Its customers split into a group that values the history and a group that values the feed, and the two groups' values run opposite. A bundle captures more than two separate prices would, and the research is the size of the two groups.

The mathematics in fullBundling profits when reservation values for the components are negatively correlated across customers; a menu of versions works when the low version is degraded just enough that the high-value customer will not take it. Taught in

MECN-446; MECN-430

Sources

Adams, Yellen (1976). Commodity Bundling and the Burden of Monopoly. The Quarterly Journal of Economics 90(3), 475. https://doi.org/10.2307/1886045
Mussa, Rosen (1978). Monopoly and product quality. Journal of Economic Theory 18(2), 301-317. https://doi.org/10.1016/0022-0531(78)90085-6

Strategic substitutes and complements

What works for a monopolist backfires in an oligopoly

In a market with a few players, the success of any strategy depends on how the others respond. We map the industry's structure first and then ask which moves soften competition and which invite it.

Born to answer
Bulow, Geanakoplos and Klemperer, 1985.
What it does now
Whether aggression here invites retreat or matching.
The mathematics
The sign of the rival's best response decides the move
The full entry
Where it began

The theory of strategic interaction among a few firms runs from Cournot and Bertrand through the game-theoretic industrial organization of the 1980s; Bulow, Geanakoplos, and Klemperer's 1985 taxonomy of strategic substitutes and complements is the working tool, and Fudenberg and Tirole gave the postures their nicknames. Kellogg's competitive strategy course teaches it through cases in which a move that would work for a monopolist backfires in an oligopoly.

Why it works now

Mapping an industry's structure and each player's incentives took a consulting engagement. The structure can now be assembled from filings and trade press quickly, and the incentives from former executives across the players, so that the history of moves and responses is on the table before the strategy is chosen. The Desk asks how many players actually set price here and what the last aggressive move provoked.

When it applies

A strategy (product proliferation, positioning, loyalty programs, capacity additions, a price cut) in a concentrated industry; or a thesis that reads margins without reading structure.

What we would ask

How many players actually set price here, and what does each one want?
What did the last aggressive move in this industry provoke?

What it needs

Industry structure, each player's cost position and incentives, and the history of moves and responses.

How we would get it

Former executives across the players; the record of past moves; the trade press and regulatory filings for structure.

Where it breaks

The industry is assumed to behave like its average firm; the response is assumed away; the structure is changing (entry, consolidation, substitutes).

For example

A company plans to fill every price point with a variant to crowd out a rival. In its three-player market the rival's former strategy head explains that the last such move triggered a capacity race that cut everyone's margin for four years.

The mathematics in fullWhether a move is strategic substitute or complement decides the response: aggression met by retreat in the first case and by matching in the second; the sign is a property of the industry, and it can be established from history before it is modeled. Taught in

MECN-441; MECS 449; MMSS 211-2

Sources

Besanko, Dranove, Shanley and Schaefer. Economics of Strategy, 7th ed. Wiley. (book)
Bulow, Geanakoplos, Klemperer (1985). Multimarket Oligopoly: Strategic Substitutes and Complements. Journal of Political Economy 93(3), 488-511. https://doi.org/10.1086/261312

Prospect theory in markets

Customers, and managers, are not the model

We look for the places where the people in the market do not behave as the textbook assumes: reference prices, loss aversion, sunk-cost pricing, anchors; and we check whether the client's own managers are doing it too.

Born to answer
Kahneman, Knetsch and Thaler, 1991: the Cornell coffee mugs.
What it does now
Why a price cut and a price rise do not mirror each other.
The mathematics
Value v(x) from a reference point, steeper in losses
The full entry
Where it began

Kahneman, Knetsch, and Thaler's 1991 review reported the coffee mug experiments at Cornell, where students given a mug demanded about twice as much to sell it as others would pay to buy it. Arkes and Blumer's 1985 theater experiment sold season tickets at different discounts and found that people who paid full price attended more, because the sunk cost felt like a reason. Al-Najjar, Baliga, and Besanko showed in 2008 that managers who fold sunk costs into prices survive in markets that the textbook says should punish them.

Why it works now

Behavioral anomalies were explanations offered after the fact. They can now be turned into predictions and tested: customers can be asked what number they compare a price to, responses to increases and decreases can be measured separately in the data, and former pricing managers can say how prices were really set. The Desk asks whether the reaction to a cut and to a raise was symmetric, and where the reference price came from.

When it applies

Pricing that customers react to asymmetrically; managers who price off fully loaded cost including sunk investment; a market where the standard model keeps being wrong in the same direction.

What we would ask

When you raised the price and when you cut it, was the reaction symmetric?
What number are your customers comparing your price to, and where did that number come from?

What it needs

Customer responses to increases and decreases separately; the reference points in the buyers' heads; the cost allocation the managers price from.

How we would get it

Customers and category buyers on reference prices and reactions; former pricing managers on how prices were actually set.

Where it breaks

Every anomaly is called behavioral after the fact; the bias story is never turned into a testable prediction.

For example

A subscription business finds cuts win few customers and increases lose many. The asymmetry is loss aversion around the reference price, and the research is what that reference price is and whether it can be moved before the next increase.

The mathematics in fullValue is v(x) relative to a reference point, steeper for losses than gains; managers who allocate sunk cost into price behave as if the reference were fully loaded cost, and a market of such managers prices differently from the textbook one. Taught in

MECN-943; MECNX-435

Sources

Kahneman, Knetsch, Thaler (1991). Anomalies: The Endowment Effect, Loss Aversion, and Status Quo Bias. Journal of Economic Perspectives 5(1), 193-206. https://doi.org/10.1257/jep.5.1.193
Al-Najjar, Baliga, Besanko (2008). Market forces meet behavioral biases: cost misallocation and irrational pricing. The RAND Journal of Economics 39(1), 214-237. https://doi.org/10.1111/j.0741-6261.2008.00011.x
Arkes, Blumer (1985). The psychology of sunk cost. Organizational Behavior and Human Decision Processes 35(1), 124-140. https://doi.org/10.1016/0749-5978(85)90049-4

The Bass diffusion model

Exponential for a while, then not

Adoption follows an S-curve driven by early adopters and then by imitation, and the two things that matter, the ceiling and the timing of the inflection, are invisible in the early data that looks exponential. We find them from comparable adoptions.

Born to answer
Bass, 1969, fitting the adoption of color television.
What it does now
Finding the ceiling and the inflection the early data cannot show.
The mathematics
Adopters = (p + q F) x (M - cumulative)
The full entry
Where it began

Frank Bass's 1969 model fitted the adoption of color televisions, refrigerators, room air conditioners, and other durables with two parameters, innovation and imitation, and a ceiling, and predicted the timing of the color television peak. It became the standard forecasting model for new products, and its lesson has not changed: the early data look exponential and are not.

Why it works now

The ceiling and the inflection, which the early data cannot identify, used to be assumed. Both can now be researched: comparable adoption curves can be retrieved from the record with their inflection points, and the people who have not adopted can be interviewed about why. The Desk asks what the ceiling is and how you know, and which past adoption this most resembles. The engagement finds the non-adopters.

When it applies

A growth thesis extrapolated from early adoption; a market whose size is inferred from its growth rate; a product whose adoption has stalled unexpectedly.

What we would ask

What is the ceiling, and how do you know? Who will never adopt this, and why?
Which past adoption does this most resemble, and where was its inflection?

What it needs

The addressable ceiling from evidence; the imitation dynamics of this category; comparable adoption curves with their inflection points.

How we would get it

Operators who lived through the comparable adoptions; customers who have not adopted on why; the record of comparable curves.

Where it breaks

The S-curve is fitted to its own first third; the ceiling is assumed; the growth is subsidy-driven and mistaken for organic imitation.

For example

A venture thesis extrapolates two years of triple-digit growth. The comparable category inflected at eleven percent penetration, and the non-adopters, interviewed, have a reason that the product does not address. The ceiling, not the growth rate, is the thesis.

The mathematics in fullAdoptions at t = (p + q x F(t)) x (M - cumulative adopters): innovation rate p, imitation rate q, ceiling M; early data identifies p and q poorly and M not at all, so M comes from research, not from the curve. Taught in

MECN-451; MMSS 1995 (exponential growth and decline)

Sources

Bass (1969). A New Product Growth Model for Consumer Durables. Management Science 15(5), 215-227. https://doi.org/10.1287/mnsc.15.5.215

Pareto and power laws

Is this market a few customers wearing a crowd?

We ask whether the market is driven by a small tail, because when it is, the average customer is nearly meaningless and the thesis lives or dies on a handful of names.

Born to answer
Pareto in the 1890s on income; Gabaix on city sizes, 1999.
What it does now
Whether a market is its average customer or a handful of names.
The mathematics
P(X > x) = (x_min / x)^alpha
The full entry
Where it began

Vilfredo Pareto found in the 1890s that income follows a distribution in which a small fraction of people hold most of the wealth, and the same shape turned up in city sizes, which Xavier Gabaix explained in 1999. Clauset, Shalizi, and Newman's 2009 paper tested twenty-four datasets and found that many claimed power laws were not, which is the caution the Library carries.

Why it works now

Concentration was a line in the risk factors. It can now be measured from the client's data and the record, customer by customer and partner by partner, and the largest accounts themselves can be interviewed. The Desk asks what share of revenue, profit, or leads the top five percent produce, and what is left if two of them leave. When the answer is most of it, the thesis is those names, and the engagement goes to them.

When it applies

A total addressable market computed as customers times average spend; a customer base described by its average; a channel where three partners produce most of the leads.

What we would ask

What share of revenue, profit, or leads comes from the top five percent?
If two of those left, what is left?

What it needs

The concentration of value across customers, products, channels, and geographies, from the data rather than from the deck.

How we would get it

The record for concentration; interviews with the operators who know which accounts actually carry the business; the largest customers themselves where reachable.

Where it breaks

A lognormal is mistaken for a power law; the tail is extrapolated from a handful of observations; the concentration is treated as a risk factor rather than as the thesis.

For example

A marketplace thesis rests on average take rate across thousands of sellers. Twelve sellers produce most of the volume and are negotiating fees as a group. The market is those twelve, and the interviews go to them.

The mathematics in fullP(X > x) = (x_min / x)^alpha; below alpha of two the variance is infinite and sample averages never settle, so any thesis stated in averages is a thesis about a number that does not exist. Taught in

MMSS 1995 (rank-size and Pareto laws); MECN-451

Sources

Gabaix (1999). Zipf's Law for Cities: An Explanation. The Quarterly Journal of Economics 114(3), 739-767. https://doi.org/10.1162/003355399556133
Clauset, Shalizi, Newman (2009). Power-Law Distributions in Empirical Data. SIAM Review 51(4), 661-703. https://doi.org/10.1137/070710111

Mixture and latent class models

The healthy average hides the cohort that is leaving

We test whether one customer base is really several, because a stable average can conceal a segment that is deteriorating fast, and the segment is where the thesis is decided.

Born to answer
Pearson, 1894, fitting Naples crabs as two populations.
What it does now
The cohort deteriorating underneath a flat average.
The mathematics
p(x) = sum of pi(k) f(x | theta(k))
The full entry
Where it began

Karl Pearson fitted the first mixture model in 1894 to measurements of crabs from Naples, suspecting one population was really two, and it was. Dempster, Laird, and Rubin's 1977 expectation-maximization algorithm made estimating hidden classes routine: guess the classes, assign the units, re-estimate, repeat.

Why it works now

Segments used to be named by the marketing department and confirmed by nobody. Customer-level data can now be split by cohort, channel, and usage in minutes, and the candidate segments proposed by the people who know the customers can be tested against it. The Desk asks whether any group looks different when the base is split by how, when, and through whom customers bought. The engagement finds the cohort that is leaving under a flat average.

When it applies

Aggregate retention or satisfaction that looks fine while something feels wrong; a customer base described as one thing; a new cohort behaving differently from the old ones.

What we would ask

If you split customers by how they bought, when they bought, and what they use, does any group look different?
Which group is growing?

What it needs

Customer-level data that can be split by cohort, channel, and usage; qualitative segments proposed by the people who know the customers.

How we would get it

Advisors and customers propose the candidate segments; the data says whether the segments are real; the report says which segment carries the risk.

Where it breaks

Segments named by the analyst before the data confirmed them; a mixture fitted to noise; the number of segments chosen for the story.

For example

A company reports flat net retention. Split by cohort, customers acquired through a partner channel in the last two years churn at three times the rate of the rest and are now a third of the base. The average is flat because the old cohort is still there.

The mathematics in fullp(x) = sum over classes k of pi(k) x f(x | theta(k)); the classes and their shares are estimated from the data, and a segmentation that changes with the number of classes assumed is a story, not a finding. Taught in

MMSS 386-2; DECS-431

Sources

Dempster, Laird, Rubin (1977). Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm. Journal of the Royal Statistical Society Series B: Statistical Methodology 39(1), 1-22. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x

Externalities and the Coase theorem

Who should own the spillover

When one party's activity imposes a cost or a benefit on another, the question is not who is at fault but who can fix it most cheaply and whether the two can bargain. We find out what the bargaining costs and who holds the right.

Born to answer
Coase, 1960: cattle straying into a neighbor's crops.
What it does now
Who should own a spillover, and whether the parties can bargain.
The mathematics
With low transaction costs the assignment does not change the outcome
The full entry
Where it began

Ronald Coase's 1960 paper used a rancher whose cattle stray onto a farmer's crops, and sparks from a railway that set fields alight, to make the point that when the parties can bargain cheaply, the efficient outcome is reached whichever of them holds the right; the assignment changes who pays, not what happens. When bargaining is costly, which is most of the time, the assignment decides the outcome.

Why it works now

The costs on each side of a spillover, and the cost of a deal among the parties, used to be assumed. They are now researchable: operators can say what abatement actually costs, former regulators can say how the right has been assigned and reassigned, and the record shows comparable bargains. The Desk asks who can fix the problem most cheaply and whether the two sides can actually bargain.

When it applies

Pollution, congestion, noise, data, shared platforms, or any business whose activity lands costs or benefits on parties outside the transaction; a regulation that assigns the right to one side.

What we would ask

Who can reduce the harm at the lowest cost, the one causing it or the one suffering it?
Can the two sides actually bargain, or are there too many of them?

What it needs

The cost of abatement on each side, the number of parties, and the transaction costs of a deal among them.

How we would get it

Operators on the real cost of abatement; former regulators on how the right has been assigned and re-assigned; the record on comparable bargains.

Where it breaks

Assuming the polluter should always pay when the sufferer could adapt more cheaply; assuming bargaining is possible among thousands.

For example

A logistics company's night deliveries impose noise on a neighborhood, and a proposed rule would ban them. The cheapest fix is quieter trucks, which the company would buy for less than the ban costs it, if the rule assigned the right and let the parties settle. The research is the cost on each side.

The mathematics in fullWith a clear assignment of the right and low transaction costs, the parties bargain to the efficient outcome whichever side holds the right; the distribution changes, the efficiency does not. With high transaction costs, the assignment decides the outcome. Taught in

MMSS 211-1; MECS 540-2

Sources

Coase (1960). The Problem of Social Cost. The Journal of Law and Economics 3, 1-44. https://doi.org/10.1086/466560

Hotelling's spatial competition

Why rivals end up next to each other

When customers pick the nearest option, competitors are pulled toward the middle of the market and end up nearly identical. We ask whether this market rewards moving to the center or being the one firm that does not.

Born to answer
Hotelling, 1929: two sellers on a beach converge on the middle.
What it does now
Whether to move toward the center of a market or away from it.
The mathematics
Location alone pulls together; price competition pushes apart
The full entry
Where it began

Harold Hotelling's 1929 paper put two sellers on a beach with customers spread along it, each customer walking to the nearer seller, and showed that both sellers end up side by side in the middle. Anthony Downs applied the same logic to political parties in 1957, which is why MMSS teaches it in the formal political models course as electoral competition. Add price competition and the sellers move apart, because being close invites undercutting.

Why it works now

Positioning has been decided from a two-by-two drawn in a workshop. The distribution of customer preferences along the dimension that matters can now be measured from choices and from interviews, and former strategy leaders can say what past moves toward or away from the center actually did. The Desk asks whether customers choose the closest option or have strong preferences for the extremes.

When it applies

A positioning decision: where to put a product, a store, a price point, or a platform relative to rivals; a market where all the competitors look the same; a political or standards contest.

What we would ask

Do customers choose the closest option to what they want, or do they have strong preferences for the extremes?
If you moved toward your rival, would you take their customers or start a price war?

What it needs

The distribution of customer preferences along the dimension that matters, and whether firms compete on position, price, or both.

How we would get it

Customers on how they actually choose; former strategy leaders on past positioning moves and their consequences.

Where it breaks

Assuming the center is always best; with price competition, firms often do better apart than together.

For example

Two regional grocers have converged on the same format in the same suburbs, and margins have compressed. The model says that is exactly what competing on location alone produces. The research question is whether a differentiated position exists that customers would pay for.

The mathematics in fullWith customers spread along a line and each choosing the nearest seller, two sellers choosing location alone converge on the center; add price competition and they separate, because proximity invites undercutting. Taught in

MMSS 211-3; MECN-441

Sources

Hotelling (1929). Stability in Competition. The Economic Journal 39(153), 41. https://doi.org/10.2307/2224214
Downs (1957). An Economic Theory of Political Action in a Democracy. Journal of Political Economy 65(2), 135-150. https://doi.org/10.1086/257897

Clustering and multidimensional scaling

Which markets behave alike

Before a thesis assumes what worked in one market transfers to another, we measure which markets, customers, or products actually behave alike, on the dimensions that matter, rather than on the map.

Born to answer
The MMSS network curriculum: which units resemble which.
What it does now
Which of twenty new markets actually resemble the three that worked.
The mathematics
Merge nearest pairs by distance on the chosen dimensions
The full entry
Where it began

Hierarchical clustering and multidimensional scaling were taught in the 1995 MMSS network unit as ways to find which units in a population resemble each other and to draw the resemblance on a page. Hotelling's 1933 principal components is the older cousin: reduce many measurements to the few directions that carry the variation.

Why it works now

Rollout decisions have grouped markets by geography because geography was the available dimension. The dimensions that actually drive behavior, density, competitive intensity, labor cost, channel mix, can now be assembled for every candidate market and the grouping done on those. The Desk asks on which dimensions these markets resemble each other and whether those are the dimensions that drove the result.

When it applies

An expansion from one market to the next; a portfolio of products or regions treated as one; a claim that a result in one segment generalizes.

What we would ask

On which dimensions do these markets resemble each other, and are those the dimensions that drove the result?
Which pair that looks alike on the map behaves most differently in the data?

What it needs

Data on the units across the dimensions that matter, and a defined notion of similarity.

How we would get it

Operators on which dimensions actually drive behavior; the data grouped on those dimensions rather than on geography.

Where it breaks

Grouping on what is easy to measure; letting the clustering method decide the number of groups.

For example

A retailer plans to roll out a format that worked in three cities to twenty. Grouped by customer density, competitive intensity, and labor cost rather than by region, only six of the twenty resemble the three, and the rollout plan is a different plan.

The mathematics in fullDefine a distance between units on the chosen dimensions; hierarchical clustering merges the nearest pairs step by step, and multidimensional scaling lays the units out so that distances on the page approximate the distances in the data. Taught in

MMSS 1995 (hierarchical clustering and multidimensional scaling)

Sources

Hotelling (1933). Analysis of a complex of statistical variables into principal components. Journal of Educational Psychology 24(6), 417-441. https://doi.org/10.1037/h0071325
Dempster, Laird, Rubin (1977). Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm. Journal of the Royal Statistical Society Series B: Statistical Methodology 39(1), 1-22. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x

Cause and effect

Whether X did cause Y, whether it would, and for whom.

Ordinary least squares

The transparent baseline any fancier model must beat

Before anything elaborate, we fit the simple, checkable relationship, and we insist that anything more complicated beat it; and we remember that what it shows is association, not cause.

Born to answer
Galton, 1886, and the transparent baseline ever since.
What it does now
The simple relationship any elaborate model has to beat.
The mathematics
beta = (X'X)^-1 X'Y
The full entry
Where it began

Francis Galton measured the heights of parents and their adult children in 1886 and found that the children of tall parents were tall, but less so, and the children of short parents short, but less so: regression toward mediocrity, from which the method took its name. Every econometrics course since, MMSS's included, has taught it as the transparent baseline and the place where causal claims quietly begin.

Why it works now

The baseline is faster to fit than it has ever been, which is not the change that matters. The change is that the omitted variables an industry insider would name in a sentence can now be asked for, in the interviews, before the regression is run, and that an elaborate model can be held to the discipline of beating the simple one. The Desk asks what else changed at the same time. That question is most of applied econometrics.

When it applies

A claim that some observable predicts an outcome (sales hiring predicts bookings, reviews predict churn, turnover predicts performance); or a complex model whose value over a simple one is unstated.

What we would ask

What is the simplest relationship in the data, and does the elaborate model actually do better than it?
What else changed at the same time that could explain the pattern?

What it needs

The data at the right grain; the variables that could drive both sides of the relationship.

How we would get it

The record and the client's own data; Advisors on which omitted variables matter in this industry.

Where it breaks

Omitted variables, reverse causation, selection into the sample, a specification chosen after seeing the results.

For example

A deck shows that customers who use a feature renew more. The feature's users are also the largest and longest-tenured customers, and once that is held constant the relationship disappears. The feature is a marker, not a cause.

The mathematics in fullY = X beta + e with beta-hat = (X'X)^-1 X'Y; each coefficient is an association holding the other regressors fixed, and an omitted driver of both X and Y appears as a coefficient on X. Taught in

MATH 386-1; DECS-431

Sources

Galton (1886). Regression Towards Mediocrity in Hereditary Stature. The Journal of the Anthropological Institute of Great Britain and Ireland 15, 246. https://doi.org/10.2307/2841583
Wooldridge, Jeffrey M. Introductory Econometrics: A Modern Approach. Cengage. (book)

Potential outcomes

Compared to what?

Every causal claim is a comparison with a world that did not happen. We state that world explicitly, and then ask what could stand in for it: a control group, a comparison period, a threshold, a natural experiment. Before-and-after on its own settles nothing.

Born to answer
Rubin, 1974: one outcome you see, one you never will.
What it does now
Making every causal claim name the world that did not happen.
The mathematics
tau = Y(1) - Y(0), only one ever observed
The full entry
Where it began

Donald Rubin's 1974 paper wrote the causal question in its modern form: every unit has an outcome under the treatment and an outcome without it, only one of which can ever be observed, and every method is a way of estimating the other. The setting was educational programs, and the framework now underlies the entire MMSS econometrics sequence.

Why it works now

Management decks are full of causal claims supported by before-and-after charts, and readers have always lacked the time to ask what else happened. The comparison group is usually available in the record, the market, the closest competitor, the regions that did not get the change, and it can now be assembled quickly. The Desk asks what would have happened to the same business over the same period without the change. The brief grades the causal evidence honestly.

When it applies

A claim that something caused something (the new CEO caused growth, the price change caused churn, the redesign lifted conversion, advertising drove bookings) supported by a before-and-after chart.

What we would ask

What would have happened to the same business over the same period without the change?
Who did not get the change, and how did they do?

What it needs

A defensible stand-in for the counterfactual: units that did not receive the treatment, a period the treatment could not have affected, or a rule that assigned it.

How we would get it

The Research Director identifies the comparison group or the natural experiment in the record; Advisors say what else changed; the report grades the causal evidence honestly.

Where it breaks

The identifying assumption fails on the facts: the comparison group differs in the way that matters, the trends were not parallel, the instrument has a second channel.

For example

A management deck credits a new sales leader with a doubling of bookings. The market doubled in the same year and the closest competitor's bookings doubled too. The comparison group was available all along; it was just not in the deck.

The mathematics in fulltau(i) = Y(i, 1) - Y(i, 0), only one of which is observed; every method is a way to estimate the unobserved term, and the assumption that licenses it (parallel trends, exclusion, no manipulation at the cutoff) is a statement about the unobserved world that the data cannot test. Taught in

MATH 386-2; DECS-431

Sources

Rubin (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology 66(5), 688-701. https://doi.org/10.1037/h0037350

Difference-in-differences

Before and after, with someone who did not get it

We compare the change in the group that received the treatment with the change in a group that did not, over the same period. It is much stronger than before-and-after alone and it still rests on one assumption we state out loud.

Born to answer
Ashenfelter and Card, 1985: the earnings dip before a training program.
What it does now
Measuring a change against the group that did not get it.
The mathematics
(T after - T before) - (C after - C before)
The full entry
Where it began

Ashenfelter and Card's 1985 study compared the earnings of people who entered a federal training program with a comparison group before and after, and found the famous dip: trainees' earnings fell just before they enrolled, which is why before-and-after alone overstated the program's effect. Card and Krueger's minimum wage comparison of New Jersey and Pennsylvania made the design famous a decade later, and Bertrand, Duflo, and Mullainathan showed in 2004 how easily it produces false precision when the errors are handled carelessly.

Why it works now

The design needs outcomes for a treated group and an untreated one before and after, which a company usually has and has never lined up. The data can now be organized that way in an afternoon, and Advisors can say what else differed between the groups. The Desk asks who got the change and who did not, and whether the two were moving together before it. The estimate is a specialist's; the design is Vista's.

When it applies

A change was introduced to some customers, regions, products, or plants and not others, and the client wants to know what it did.

What we would ask

Who got the change and who did not, and were the two groups moving in parallel before it?
What else happened to the treated group that did not happen to the others?

What it needs

Outcomes for both groups before and after; evidence that the groups trended together before the change.

How we would get it

The client's data or the record for both groups; Advisors on what else differed between them.

Where it breaks

The trends were not parallel; the control group was contaminated; the treated group was chosen because it was already changing.

For example

A company claims its new pricing lifted revenue per customer by eighteen percent. Customers on the new pricing rose eighteen percent; comparable customers still on the old pricing rose eleven. The effect is seven, and the report says so with the assumption attached.

The mathematics in fulltau = (Y(treated, after) - Y(treated, before)) - (Y(control, after) - Y(control, before)); valid under parallel trends, and the standard errors must respect that outcomes within a group are correlated over time. Taught in

MATH 386-2; DECS-431

Sources

Ashenfelter, Card (1985). Using the Longitudinal Structure of Earnings to Estimate the Effect of Training Programs. The Review of Economics and Statistics 67(4), 648. https://doi.org/10.2307/1924810
Bertrand, Duflo, Mullainathan (2004). How Much Should We Trust Differences-In-Differences Estimates?. The Quarterly Journal of Economics 119(1), 249-275. https://doi.org/10.1162/003355304772839588

Regression discontinuity

The people just above and just below the line

When a rule assigns something by a threshold (a credit score, a size cutoff, an eligibility date), the units just either side of the line are nearly identical except for the treatment, and comparing them is close to an experiment.

Born to answer
Thistlethwaite and Campbell, 1960, either side of a scholarship cutoff.
What it does now
Reading a rule's effect from the units just either side of its line.
The mathematics
Compare the limits as X approaches the cutoff
The full entry
Where it began

Thistlethwaite and Campbell's 1960 paper compared students just above and just below the test-score cutoff for a National Merit certificate, who were nearly identical in everything but the award, and read the award's effect from the gap between them. Imbens and Lemieux's 2008 guide is the modern practitioner's manual.

Why it works now

Thresholds are everywhere in business and regulation and were rarely exploited, because nobody thought of a size cutoff or an eligibility date as an experiment. The rules and the outcomes on both sides of them can now be found in the record quickly, and operators can say whether the threshold was gamed. The Desk asks whether a rule with a cutoff decided who got this. When one did, the cleanest evidence obtainable is sitting either side of the line.

When it applies

A treatment was assigned by a cutoff: a loan program, a regulatory threshold, a pricing tier, a promotion rule; and the client wants its effect.

What we would ask

Is there a rule with a threshold that decided who got this?
Could anyone have moved themselves across the line on purpose?

What it needs

The assignment rule, outcomes on both sides of it, and enough units near the threshold.

How we would get it

The record for the rule and the outcomes; operators on whether the threshold was gamed.

Where it breaks

Units manipulated their position at the cutoff; the effect near the line is not the effect far from it.

For example

A regulation applies to firms above a revenue threshold. Firms just above and just below it differ in nothing but the regulation, and the comparison of their subsequent growth is the cleanest evidence obtainable on what the regulation costs.

The mathematics in fullEffect = lim of E[Y | X just above c] - E[Y | X just below c], where c is the cutoff assigning treatment; identification requires no manipulation of X at c and continuity of everything else across it. Taught in

MATH 386-2

Sources

Thistlethwaite, Campbell (1960). Regression-discontinuity analysis: An alternative to the ex post facto experiment. Journal of Educational Psychology 51(6), 309-317. https://doi.org/10.1037/h0044319
Imbens, Lemieux (2008). Regression discontinuity designs: A guide to practice. Journal of Econometrics 142(2), 615-635. https://doi.org/10.1016/j.jeconom.2007.05.001

Instrumental variables

Something that moved the cause and nothing else

When the supposed cause was chosen by the people it affected, we look for something that shifted it from outside, and affected the outcome only through it. Finding one is the hard part, and we say plainly when there is none.

Born to answer
Angrist, Imbens and Rubin, 1996; the Vietnam draft lottery.
What it does now
When the cause was chosen by the people it affected.
The mathematics
Effect = Cov(Z, Y) / Cov(Z, D)
The full entry
Where it began

Angrist, Imbens, and Rubin's 1996 paper clarified what an instrument identifies: the effect for the people the instrument moved. The teaching example is Angrist's use of the Vietnam draft lottery, which assigned military service by birthdate for reasons unrelated to later earnings, to measure the effect of service on earnings.

Why it works now

Finding a valid instrument is still the hard part and always will be; no tool changes that. What has changed is that the accidents, rules, and shocks that moved a variable for reasons unrelated to the outcome, a regulatory change in some states, a supply disruption, a policy phased in by region, can be surfaced from the record and from Advisors far faster than before. The Desk asks what moved the cause for reasons that had nothing to do with the outcome. It says plainly when there is nothing.

When it applies

The cause and the effect choose each other: advertising and sales, sales intensity and growth, adoption and productivity, expansion and performance.

What we would ask

What moved the cause for reasons that had nothing to do with the outcome?
Could that thing have affected the outcome any other way?

What it needs

A source of variation in the cause that is unrelated to the outcome except through the cause; and the honesty to say when none exists.

How we would get it

Advisors on the accidents, rules, and shocks that moved the variable in this industry; the record for them.

Where it breaks

The instrument has a second channel; it is weak; it identifies the effect only for the units it moved.

For example

Does advertising drive bookings, or do companies advertise when bookings are strong? A regulatory change that restricted advertising in some states and not others moved advertising for reasons unrelated to demand, and the comparison across states answers the question.

The mathematics in fullZ shifts D and affects Y only through D; the effect is Cov(Z, Y) / Cov(Z, D), estimated in two stages, and it is the effect for the units Z moved, not for everyone. Taught in

MATH 386-2

Sources

Angrist, Imbens, Rubin (1996). Identification of Causal Effects Using Instrumental Variables. Journal of the American Statistical Association 91(434), 444-455. https://doi.org/10.1080/01621459.1996.10476902

Heterogeneous treatment effects

For whom, not whether

An average effect hides the customers for whom it is large and the ones for whom it is zero. We ask which segments the product, the policy, or the change actually works for, because a thesis that assumes it generalizes usually does not.

Born to answer
Athey and Imbens, 2016: an honest search for where an effect lives.
What it does now
Which customers a product works for, rather than the average.
The mathematics
CATE(x) = E[Y(1) - Y(0) | X = x]
The full entry
Where it began

Athey and Imbens's 2016 paper built causal trees, which search for the subgroups where a treatment's effect differs while keeping the estimates honest by splitting the sample. It formalized what every experienced operator knows: the average effect of a product, a policy, or a change hides the customers for whom it is large and the ones for whom it is nothing.

Why it works now

Heterogeneity used to be discovered by accident after a launch failed in a segment. It can now be looked for on purpose: customers across segments can be interviewed about what changed for them, the data split by the characteristics the Advisors name, and the market thesis tested against the segment the client actually plans to sell to. The Desk asks for which customers it worked best and what they have in common.

When it applies

A claim that a product improves something by an average amount; a thesis that a result in one segment or geography will transfer to others; a pilot that worked.

What we would ask

For which customers did it work best, and what do they have in common?
Is the segment you plan to sell to more like the ones where it worked or the ones where it did not?

What it needs

Outcomes by segment with enough units per segment; the characteristics that could predict where the effect lives.

How we would get it

Customers across segments on what changed for them; the data split by the characteristics the Advisors say matter.

Where it breaks

Segments chosen after seeing the results; too few units per segment; the effect in the pilot segment assumed for the market.

For example

A workflow product improves productivity eight percent on average. In firms with a dedicated administrator it improves twenty; in firms without one it improves nothing. The market thesis was built on the average, and the addressable market is the first kind of firm.

The mathematics in fullCATE(x) = E[Y(1) - Y(0) | X = x]; estimated by splitting on covariates chosen before the outcome is examined, with honest sample splitting so the segments are not fitted to the noise. Taught in

MATH 386-2

Sources

Athey, Imbens (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences 113(27), 7353-7360. https://doi.org/10.1073/pnas.1510489113

Randomized controlled experiments

Sometimes the answer is to try it on a few

When the decision can be tested on a slice before it is made for everyone, the cleanest evidence obtainable is an experiment, and we design it so that it is large enough to detect the effect that would matter.

Born to answer
Fisher at Rothamsted in the 1920s, testing fertilizer on plots.
What it does now
Testing a decision on a random slice before making it for everyone.
The mathematics
Sample needed scales with 1 / (effect size squared)
The full entry
Where it began

Ronald Fisher designed the first randomized field experiments at Rothamsted in the 1920s to test fertilizers on plots of land, and randomization has been the gold standard since. Kohavi and colleagues' 2009 guide brought it to the web, where Microsoft and Amazon learned that most confident product ideas fail a controlled test, and that the test must be large enough to see the effect that matters.

Why it works now

A business decision can be tested on a random slice before it is made for everyone, and most are not, because designing the test was a specialist's job and the sample-size arithmetic an afterthought. The design is now the deliverable: Advisors and the client's operators shape it, the smallest effect that matters sets the sample, and the client runs it. The Desk asks whether this could be tried on a random slice first. When it can, that is the cleanest evidence there is.

When it applies

A pricing change, a product change, a channel, or a policy that could be rolled out to a random subset first; or a pilot that was run without a control or without enough units.

What we would ask

Could you try this on a random slice of customers, regions, or products before you decide?
How small an effect would still matter, and how many units does it take to see it?

What it needs

Random assignment, a control group, an outcome measured the same way in both, and a sample large enough for the smallest effect that matters.

How we would get it

Vista designs the experiment with the Advisors and the client's operators; the client runs it; the report reads it, including the null.

Where it breaks

The pilot was not random; the sample was too small and the null is meaningless; the outcome was measured differently in the two arms; the experiment changed behavior by being observed.

For example

A company wants to know whether a lower price increases lifetime value. It can offer the lower price to a random half of new customers in two regions for a quarter. The design is the deliverable; the quarter is the evidence.

The mathematics in fullWith random assignment, E[Y | treated] - E[Y | control] is the average effect; the sample needed to detect an effect of size d at conventional thresholds scales with 1/d^2, so halving the smallest effect that matters quadruples the sample. Taught in

MATH 386-2; DECS-430

Sources

Kohavi, Longbotham, Sommerfield (2009). Controlled experiments on the web: survey and practical guide. Data Mining and Knowledge Discovery 18(1), 140-181. https://doi.org/10.1007/s10618-008-0114-1
Cohen (1992). A power primer. Psychological Bulletin 112(1), 155-159. https://doi.org/10.1037/0033-2909.112.1.155

Selection and nonresponse bias

The survivors, the responders, and the eager

Who you heard from decides what you learned. Customers who answer differ from customers who do not; surviving customers cannot tell you about churn; experts who volunteer differ from experts who know. We design the sample before we trust the finding.

Born to answer
The 1936 Literary Digest poll; Heckman, 1979.
What it does now
Recruiting the churned customers, not the seller's reference list.
The mathematics
E[Y | observed] is not E[Y]
The full entry
Where it began

The Literary Digest polled millions of its readers in 1936 and predicted a landslide for Landon; Roosevelt won in one, because the readers were not the electorate. James Heckman's 1979 paper gave selection its statistics with the case of women's wages, observed only for the women who chose to work. Groves's 2006 review showed that a low response rate does not by itself bias a survey; the difference between those who answer and those who do not does.

Why it works now

Reference calls chosen by the seller and surveys answered by the happy have always been the sample a buyer got. The missing population can now be found and recruited deliberately: churned customers, investors who passed, employees who left. The Desk asks who was in a position to be in this sample and who was not. The engagement interviews the ones who were not.

When it applies

A finding from a customer survey, a set of reference calls chosen by the seller, a panel of volunteers, or any sample whose members chose to be in it.

What we would ask

Who was in a position to be in this sample, and who was not?
Did the ones who declined differ from the ones who answered in the way that matters?

What it needs

The mechanism that put each unit in the sample; some evidence about the ones who are missing.

How we would get it

The Research Director recruits from the missing population deliberately: churned customers, passed investors, former employees who left; the report weights accordingly.

Where it breaks

The sample selected itself and the finding is the selection; the missing units are assumed to resemble the present ones.

For example

A seller provides ten reference customers, all happy. The research recruits five customers who left in the last two years. The reference calls describe the product; the churned customers describe the decision.

The mathematics in fullE[Y | observed] differs from E[Y] whenever the probability of being observed depends on Y; the bias in a mean is roughly the nonresponse rate times the gap between respondents and nonrespondents, so the gap has to be estimated, not assumed away. Taught in

MATH 386-2; DECS-430

Sources

Heckman (1979). Sample Selection Bias as a Specification Error. Econometrica 47(1), 153. https://doi.org/10.2307/1912352
Groves (2006). Nonresponse Rates and Nonresponse Bias in Household Surveys. Public Opinion Quarterly 70(5), 646-675. https://doi.org/10.1093/poq/nfl033

Panel data and fixed effects

Compare the company to itself

When companies, regions, or customers differ in ways nobody can measure, we compare each one to itself over time rather than to each other, which removes everything about the unit that does not change.

Born to answer
Nickell, 1981, on the bias in short dynamic panels.
What it does now
Comparing a plant to its own past rather than to other plants.
The mathematics
y(i,t) = a(i) + b x(i,t) + e(i,t)
The full entry
Where it began

Panel data, the same units observed repeatedly over time, let an analyst compare each unit to itself and remove everything about it that does not change. Stephen Nickell's 1981 paper is the standard warning about the method: with a lagged outcome and few periods, the fixed-effects estimate is biased in a known direction. The 1996 MMSS curriculum taught panel studies and lagged regression as a unit.

Why it works now

Companies have always had panel data on their own plants, regions, and customers and rarely analyzed it as such, because organizing it took a project. It can now be organized in an afternoon, and Advisors can say which permanent differences between units would otherwise contaminate a cross-section comparison. The Desk asks whether the same units are observed before and after, or different ones.

When it applies

A cross-section comparison where the units differ in unmeasured ways; repeated observations of the same customers, plants, or firms; a claim that a change worked because the changed units are now better than the unchanged ones.

What we would ask

Do you observe the same units before and after, or different units?
What about each unit never changes and could explain the difference?

What it needs

Repeated observations of the same units over time, with the change occurring at different times for different units.

How we would get it

The client's own panel data; Advisors on what fixed differences between units matter.

Where it breaks

Fixed effects remove permanent differences, not changing ones; with short panels and lagged outcomes the estimates are biased.

For example

Plants that adopted a new maintenance system have less downtime than plants that did not. The adopting plants were also newer. Comparing each plant to its own downtime before and after adoption, the effect is a third of the cross-section estimate.

The mathematics in fully(i,t) = a(i) + b x(i,t) + e(i,t): the unit effect a(i) absorbs everything permanent about unit i, so b is estimated only from changes within units; with a lagged outcome on the right and few periods, b is biased. Taught in

MATH 386-1; MMSS 1996 (panel studies)

Sources

Nickell (1981). Biases in Dynamic Models with Fixed Effects. Econometrica 49(6), 1417. https://doi.org/10.2307/1911408
Wooldridge, Jeffrey M. Introductory Econometrics: A Modern Approach. Cengage. (book)

The two cultures of statistical modeling

Predict, or explain

A model that forecasts well can be a black box, and a model you will act on must identify a cause. We ask which of the two the decision needs, because the same data answers them differently and the wrong one is confidently wrong.

Born to answer
Breiman, 2001: fitting a mechanism versus predicting well.
What it does now
Whether a churn model is a forecast or a lever.
The mathematics
Out-of-sample error, or an identifying assumption
The full entry
Where it began

Leo Breiman's 2001 paper on the two cultures of statistical modeling drew the line between models that fit a mechanism and models that predict without one, and argued that prediction on held-out data was the honest test of the second. Stanford's data science core teaches both cultures, and Athey and Imbens's causal trees are the modern attempt to keep the predictive machinery while identifying a cause.

Why it works now

Every company now has a model that predicts something, and every executive wants to act on its top feature, which is the mistake the distinction exists to prevent. What has changed is that both questions can be answered, separately and quickly, once someone asks which one the decision needs. The Desk asks whether the client needs to know which customers will leave or what would make them stay.

When it applies

A machine-learning model presented as a reason to intervene; a churn or default score used to decide what to change; an executive asking why when the model only says what.

What we would ask

Do you need to know which customers will leave, or what would make them stay?
If you act on this model's most important feature, do you expect the outcome to change?

What it needs

Clarity about the decision; for prediction, held-out accuracy; for explanation, an identifying assumption.

How we would get it

The Research Director states which question is being answered; a prediction is validated out of sample, an explanation is designed with a counterfactual, and the report does not confuse them.

Where it breaks

Using a predictor to choose an intervention; demanding a causal story from a forecasting tool.

For example

A lender's model predicts default from three hundred features, and the top feature is the number of times a customer checked their balance. The model is a good predictor and a useless lever: stopping customers from checking their balance would change nothing. What would change default is a different question, with a different design.

The mathematics in fullA predictor minimizes out-of-sample error and is validated by holding data back; an explanation identifies the effect of changing X on Y and is validated by the credibility of its identifying assumption; a feature that predicts an outcome is not thereby a lever on it. Taught in

Stanford data science core; MATH 386-2

Sources

Breiman (2001). Statistical Modeling: The Two Cultures. Statistical Science 16(3). https://doi.org/10.1214/ss/1009213726
Athey, Imbens (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences 113(27), 7353-7360. https://doi.org/10.1073/pnas.1510489113

When, not whether

Timing questions, which are usually the ones that decide the outcome.

Survival analysis and hazard models

Not whether, but when

Whether something will happen and when it will happen are different questions, and the second one usually decides the outcome. We answer it from the history of comparable cases, not from opinion.

Born to answer
Kaplan and Meier, 1958; Cox, 1972, on clinical survival times.
What it does now
When a borrower refinances or a customer leaves, not whether.
The mathematics
h(t | X) = h0(t) exp(X beta)
The full entry
Where it began

Kaplan and Meier's 1958 paper showed how to estimate a survival curve when many of the patients were still alive at the end of the study, and David Cox's 1972 model let the characteristics of each patient shift the hazard, the chance the event happens next given it has not happened yet. Both were built for clinical survival times, and they transferred whole to customers, loans, and companies.

Why it works now

Whether a borrower refinances, a customer churns, or a chief executive leaves was answered with a probability when the decision turned on a date. Comparable durations can now be assembled from the record, including the cases still running, and Advisors who have watched the clock in that industry can say what shortens it. The Desk asks whether the question is whether or when, and which answer changes the decision. Timing questions get timing machinery.

When it applies

A question about timing: when a customer churns, a company needs capital, a CEO leaves, a competitor enters, a regulator acts, a contract renews, a covenant is breached; or a probability question that is really a timing question in disguise.

What we would ask

Is the question whether this happens, or when? And which answer changes the decision?
In comparable cases, how long did it take, and what shortened it?

What it needs

A set of comparable cases with their durations, including the ones in which the event has not yet happened; the characteristics that shift the timing.

How we would get it

The record for comparable durations; Advisors who have seen the clock run in this industry on what accelerates it.

Where it breaks

Cases still alive at the end of the data are dropped or treated as if the event never comes; the covariates were measured after the process began.

For example

A credit investor asks whether a borrower will need to refinance. The borrower will; every borrower does. The question is whether it happens inside the fund's horizon, and the comparable cases put the hazard sharply higher in the eighteen months around a maturity wall the deck does not mention.

The mathematics in fullThe hazard h(t) is the chance the event occurs in the next interval given it has not yet; survival S(t) = exp(-H(t)); h(t | X) = h0(t) x exp(X beta) lets the characteristics shift the timing, and the shape of h0 (rising, falling, spiking at dates) is the finding. Taught in

MMSS 1996 (event history and proportional hazards); MATH 386-2

Sources

Cox (1972). Regression Models and Life-Tables. Journal of the Royal Statistical Society Series B: Statistical Methodology 34(2), 187-202. https://doi.org/10.1111/j.2517-6161.1972.tb00899.x
Kaplan, Meier (1958). Nonparametric Estimation from Incomplete Observations. Journal of the American Statistical Association 53(282), 457-481. https://doi.org/10.1080/01621459.1958.10501452

Markov chains

The retention number is stable; the transitions are not

We describe the customer base, the loan book, or the pipeline as states and the rates of moving between them, because aggregate figures can look flat while the movements underneath have already turned.

Born to answer
Markov, 1906, counting vowels and consonants in Pushkin.
What it does now
The transitions turning underneath a stable retention number.
The mathematics
p(t + 1) = p(t) P
The full entry
Where it began

Andrey Markov's chains date to 1906, and his first published application was to the sequence of vowels and consonants in Pushkin's verse. The idea that a system's next state depends on its current state and a table of transition probabilities became the workhorse of everything from credit ratings to customer lifecycles.

Why it works now

Aggregate retention has been the number everyone reported because unit-level state histories were tedious to build. They can now be built from a company's own data in hours, with the states defined by the operators who know what at risk really looks like. The Desk asks how many customers who were healthy a year ago are at risk now, and which transition changed most. The number under the number becomes visible.

When it applies

A stable aggregate (net retention, delinquency rate, pipeline conversion) that the client suspects is hiding something; a base that is aging; a pipeline that is filling at the top and emptying in the middle.

What we would ask

Of the customers who were healthy a year ago, how many are at risk now?
Which transition has changed most, and when did it change?

What it needs

Unit-level state histories over several periods; a state definition that the operators agree is real.

How we would get it

The client's data, with the states defined by the operators who know what at risk actually looks like; former operators on which transitions predict the next.

Where it breaks

The process has memory the model ignores; the transitions are estimated from too few moves; the states are defined by the analyst rather than the business.

For example

A software company reports net retention above one hundred and ten percent for six quarters. The transition from expanded to at-risk has doubled in the last three, and the aggregate is held up by expansion in the oldest cohort. The clock is visible in the transitions and invisible in the number.

The mathematics in fullp(t + 1) = p(t) x P with P(i, j) the probability of moving from state i to state j; the long-run mix solves pi = pi x P, and a small rise in the transition out of the healthy state compounds into a large change in the mix a year out. Taught in

MMSS 1995 (matrix models and stochastic processes); DECS-430

Poisson processes

Is six in a quarter a signal or a coincidence?

Complaints, incidents, departures, regulatory inquiries, and outages arrive at a natural rate. When a cluster arrives well above it, we quantify how unusual that is before anyone calls it noise or calls it a crisis.

Born to answer
Bortkiewicz, 1898, on Prussian cavalrymen killed by horse kicks.
What it does now
Whether four executive departures in a quarter is a signal.
The mathematics
P(N = k) = exp(-lt) (lt)^k / k!
The full entry
Where it began

Ladislaus Bortkiewicz applied the Poisson distribution in 1898 to deaths of Prussian cavalrymen from horse kicks, corps by corps and year by year, and showed that rare events arriving at a steady rate produce clusters by chance more often than intuition allows. The same arithmetic says when a cluster is too large to be chance.

Why it works now

A run of executive departures or product incidents was read by gut. The base rate for a company of that kind can now be found from the record, and the events themselves examined for a shared cause, which collapses several events into one. The Desk asks what the normal rate is and where that number comes from, and whether the events shared a cause or only a quarter. Four departures under one new executive is one event with a name.

When it applies

A run of adverse events (executive departures, product failures, customer incidents, regulatory inquiries, security breaches) that the client cannot tell from bad luck.

What we would ask

What is the normal rate of this event for a company of this kind, and where does that number come from?
Did the events share a cause, or only a quarter?

What it needs

The historical rate of the event for this company and its peers; the timing and causes of the recent events.

How we would get it

The record for the base rate; operators and former employees on whether the events are one event or several.

Where it breaks

The window is chosen after the cluster; the events are not independent and the count is inflated; the base rate is assumed rather than found.

For example

Four vice presidents leave a company in one quarter. At the company's historical rate, four in a quarter is very unlikely by chance; the interviews establish that three of the four report to the same new executive, which makes it one event with a name.

The mathematics in fullWith events arriving at rate lambda, P(N(t) = k) = exp(-lambda t) (lambda t)^k / k!, so k events where lambda t were expected has a computable improbability; if the events share a cause they are not independent and the count must be reduced to the number of causes. Taught in

MMSS 1995 (stochastic processes); DECS-430

Dynamic programming

Decide the order before the answer

When a decision recurs over time, the right move now depends on the moves it leaves open later. We work backward from the last decision to the first, which often reverses what looked obvious.

Born to answer
Bellman, 1954: solve the last stage first and work backward.
What it does now
Choosing the first move for the options it leaves open.
The mathematics
V(s, t) = max [ payoff + E V(s', t + 1) ]
The full entry
Where it began

Richard Bellman's dynamic programming, published in 1954, solves a multistage decision by starting from the last stage and working backward, so that each early choice is valued for the choices it leaves open. He later admitted he chose the name partly because his sponsor disliked the word research.

Why it works now

Multiyear programs, capacity plans, and roadmaps have been solved forward, greedily, because working backward required a model of what would be learned at each stage. The stages and what each reveals can now be laid out in a conversation, and operators who ran comparable programs can say where the sequence went wrong. The Desk asks what the last decision in the sequence is and which early choice forecloses the most later ones. The answer often reverses the obvious first step.

When it applies

A multi-year program of investments, a product roadmap, a capacity plan, a phased acquisition; or a client deciding the first step without having thought about the last.

What we would ask

What is the last decision in this sequence, and what does it need to be true when you get there?
Which early choice forecloses the most later ones?

What it needs

The sequence of decisions, what each one forecloses, and what is learned between them.

How we would get it

Operators who have run comparable multi-stage programs on where the sequence went wrong; the record on comparable programs.

Where it breaks

The problem is solved forward, greedily; the value of keeping options open is ignored; the horizon is cut short.

For example

A manufacturer plans three capacity additions over six years. Solved forward, the first is the largest. Solved backward from the demand uncertainty that resolves after the first, the first is the smallest, and the program is cheaper in every scenario but the best one.

The mathematics in fullV(state, t) = max over actions of [ payoff(action, state) + E[V(next state, t + 1)] ], solved from the final period backward; the early action is chosen for the states it leaves reachable, not for its own payoff. Taught in

MECS 560-2; MMSS 300

Sources

Bellman (1954). The theory of dynamic programming. Bulletin of the American Mathematical Society 60(6), 503-515. https://doi.org/10.1090/S0002-9904-1954-09848-8

Committees and institutions

How groups, boards, regulators, and voters actually decide.

The Condorcet jury theorem

A committee is not a single mind

A group of imperfect decision makers can be far better or far worse than its members, depending on whether their errors are independent and on the rule they vote by. We look at how the committee that will decide actually decides.

Born to answer
Condorcet, 1785; Galton's ox at the 1907 country fair.
What it does now
Whether an investment committee is many minds or one.
The mathematics
Majority accuracy rises in n only if errors are independent
The full entry
Where it began

Condorcet proved in 1785 that a majority of independent members who are each more likely right than wrong becomes more reliable as the group grows, and Francis Galton found the same thing at a Plymouth livestock fair in 1907, where the median of nearly eight hundred guesses of an ox's weight came within one percent of the truth. Austen-Smith and Banks showed in 1996 that the result fails when members vote strategically, and later work showed it fails when their errors are correlated.

Why it works now

Presenting to an investment committee has always meant guessing how it decides. How it decides is now researchable: former members of comparable committees can say whether views form independently or after the same senior voice, and the record shows what the committee has done before. The Desk asks whether the members form their views independently and what rule decides. A committee that hears the same expert before voting is one mind with several votes.

When it applies

An investment committee, a board, a credit committee, a panel of buyers, or any group whose decision is the outcome; or a client preparing to present to one.

What we would ask

Do the members form their views independently, or after hearing the same person?
What rule decides: majority, unanimity, or one voice?

What it needs

The committee's rule, its information flow, and the independence of its members' judgments.

How we would get it

Former members of comparable committees on how the decision is really made; the record of the committee's past decisions.

Where it breaks

A committee that hears the same expert before voting is one mind with several votes; unanimity rules are assumed to be conservative when they can be the opposite.

For example

A client will present to an investment committee that requires unanimity. Its members read the same memo and defer to the same senior partner. The research is not the memo; it is what that partner believes, and what would change it.

The mathematics in fulln independent members each right with probability p above one half are right by majority with probability rising toward one in n; with correlated errors the gain vanishes, and under unanimity a strategic voter may vote against private evidence, so the rule shapes the outcome. Taught in

MECS 540-3; MMSS 211-3

Sources

Austen-Smith, Banks (1996). Information Aggregation, Rationality, and the Condorcet Jury Theorem. American Political Science Review 90(1), 34-45. https://doi.org/10.2307/2082796
Galton (1907). Vox Populi. Nature 75(1949), 450-451. https://doi.org/10.1038/075450a0

Median voter and agenda setting

Whoever sets the agenda sets the answer

In any group that votes, the order in which options come up and the option that stands if nothing passes shape the result as much as the votes do. We find out who controls the agenda and what the default is.

Born to answer
Black, 1948; Romer and Rosenthal, 1978, on school budget votes.
What it does now
Why the obvious option keeps losing a board vote.
The mathematics
Majority rule selects the median ideal point
The full entry
Where it began

Duncan Black showed in 1948 that under majority rule with single-peaked preferences the median voter's preference wins, and Anthony Downs built a theory of party competition on it in 1957. Romer and Rosenthal's 1978 agenda-setter model, which they went on to test against Oregon school budget referenda, showed the other half: whoever sets the agenda, and whatever happens if the proposal fails, can move the outcome a long way from the median. Kenneth Arrow had proved in 1950 that no voting rule escapes all such trouble.

Why it works now

Why the obvious option keeps losing in a board, a rulemaking, or a standards body used to be a mystery attributed to politics. The agenda setter, the default, and the members' rankings can now be established from former participants and from the record of past votes. The Desk asks what happens if nothing passes and who benefits, and who decides what comes to a vote and in what order.

When it applies

A board vote, a shareholder proposal, a regulatory rulemaking, a standards committee, a legislature; or a client who cannot understand why the obvious option keeps losing.

What we would ask

What happens if nothing passes, and who benefits from that?
Who decides which options are put to a vote, and in what order?

What it needs

The decision rule, the status quo, the agenda setter, and the members' rankings of the options.

How we would get it

Former participants in the body on how agendas are really set; the record of past votes and the defaults that prevailed.

Where it breaks

The aggregation rule is chosen by whoever wants the result; preferences are assumed single-peaked when they are not.

For example

A shareholder proposal that most holders support keeps failing. The board controls the agenda and pairs it each year with an alternative that splits its supporters. The research is the pairing, not the support.

The mathematics in fullUnder majority rule with single-peaked preferences the outcome is the median member's ideal point; an agenda setter who controls the alternative to the status quo can move the outcome a long way from the median toward its own. Taught in

MMSS 211-3; MECS 540-3

Sources

Black (1948). On the Rationale of Group Decision-making. Journal of Political Economy 56(1), 23-34. https://doi.org/10.1086/256633
Romer, Rosenthal (1978). Political resource allocation, controlled agendas, and the status quo. Public Choice 33(4), 27-43. https://doi.org/10.1007/bf03187594
Arrow (1950). A Difficulty in the Concept of Social Welfare. Journal of Political Economy 58(4), 328-346. https://doi.org/10.1086/256963
Downs (1957). An Economic Theory of Political Action in a Democracy. Journal of Political Economy 65(2), 135-150. https://doi.org/10.1086/257897

Political economy of institutions

Regulators and governments are players with payoffs

We treat the regulator, the ministry, or the legislature as an actor with its own incentives, audiences, and commitment problems, and we research those rather than reading the press release.

Born to answer
Fearon, 1994: leaders, audiences, and the cost of backing down.
What it does now
Pricing a regulator as an actor with incentives, not as weather.
The mathematics
The institution best-responds to the audience that punishes it
The full entry
Where it began

James Fearon's 1994 paper on audience costs and his 1995 paper on rationalist explanations for war treated states and leaders as actors with incentives, audiences, and commitment problems rather than as weather. MMSS teaches the same formal political models to undergraduates; Kellogg's political economy sequence teaches them at the doctoral level.

Why it works now

Political and regulatory risk has been priced from headlines because the decision makers inside the institutions were unreachable. Former regulators, former officials, and the advisors who have worked the institution can now be recruited like any other Advisor, and the record shows which provisions survive every budget cycle and which are revised in each. The Desk asks who inside the institution decides and what they answer to.

When it applies

A thesis exposed to a regulatory decision, a subsidy, a tariff, a license, an election, or a policy reversal; or a client trying to price political risk from headlines.

What we would ask

Who inside the institution decides this, and what do they answer to?
What would it cost them to reverse course, and have they paid that cost before?

What it needs

The decision maker's incentives, the audiences in front of which commitments were made, and the history of reversals.

How we would get it

Former regulators, former officials, and the advisors who have worked the institution; the record of comparable decisions and their timing.

Where it breaks

The institution is treated as a monolith; the reversal cost is assumed from rhetoric; the timing is guessed from the calendar.

For example

A renewable developer's returns depend on a tariff regime. The former director of the agency that administers it, interviewed, explains which provisions are politically untouchable and which are revised every budget cycle. The thesis is exposed to the second kind.

The mathematics in fullThe institution's choice is a best response given its payoffs and the audiences that punish reversal; commitment problems and asymmetric information explain policy conflicts that would otherwise be settled by bargaining, and both can be researched. Taught in

MMSS 211-3; MECS 540-2; MECS 540-3

Sources

Fearon (1994). Domestic Political Audiences and the Escalation of International Disputes. American Political Science Review 88(3), 577-592. https://doi.org/10.2307/2944796
Fearon (1995). Rationalist explanations for war. International Organization 49(3), 379-414. https://doi.org/10.1017/S0020818300033324

Social choice and strategic voting

Governance is a voting game

Shareholder votes, board contests, activist campaigns, and creditor committees are decided by rules, coalitions, and the strategic misreporting of preferences. We model who votes, under what rule, and who can be brought across.

Born to answer
Arrow, 1950; Gibbard, 1973: every rule can be manipulated.
What it does now
Finding the swing holder's real price in a contested vote.
The mathematics
No non-dictatorial rule is immune to manipulation
The full entry
Where it began

Kenneth Arrow's 1950 impossibility theorem showed that no rule for combining individual rankings into a group ranking satisfies a short list of reasonable conditions at once, and Allan Gibbard proved in 1973 that any non-dictatorial rule over three or more options can be manipulated by strategic voting. Shareholder votes, board contests, and creditor committees are the corporate versions.

Why it works now

Contested votes used to be counted by stated positions. The swing holder's real incentives, and what it took to move that holder in comparable situations before, can now be found from former activists, proxy advisors, and governance specialists, and from the record of past votes. The Desk asks which coalition wins under this rule and who the swing is. The research is that holder's price, not the headcount.

When it applies

An activist situation, a contested merger vote, a creditor committee, a founder-control structure, or a governance change that the thesis depends on.

What we would ask

Under this rule, which coalition wins, and who is the swing?
Who might vote against their stated position, and why?

What it needs

The rule, the holders and their preferences, the coalitions available, and the incentives to misreport.

How we would get it

Former activists, proxy advisors, and governance specialists on how comparable votes were won; the record of the holders and their past votes.

Where it breaks

Stated positions are taken as votes; the swing holder's incentives are not researched; the rule is misread.

For example

A merger requires a majority of the minority. The largest minority holder has stated opposition and has also been in a comparable situation twice, both times settling for a small price bump before the vote. The research is that holder's price, not the vote count.

The mathematics in fullAny non-dictatorial rule over three or more options can be manipulated by strategic voting; the outcome is a coalition outcome under the rule, and the swing member's incentives, not the headcount, decide it. Taught in

MECS 540-3; MMSS 211-3

Sources

Gibbard (1973). Manipulation of Voting Schemes: A General Result. Econometrica 41(4), 587. https://doi.org/10.2307/1914083
Arrow (1950). A Difficulty in the Concept of Social Welfare. Journal of Political Economy 58(4), 328-346. https://doi.org/10.1086/256963

Public goods and the commons

Everyone benefits, nobody pays

When a group shares an interest and each member gains most by letting the others do the work, the thing does not get done. We ask who bears the cost, who captures the benefit, and what would make contributing individually worthwhile.

Born to answer
Samuelson, 1954; Olson, 1965; Hardin, 1968.
What it does now
Why the consortium everyone wants goes unfunded.
The mathematics
Contribute only when the share B/n exceeds the cost c
The full entry
Where it began

Paul Samuelson wrote down the mathematics of a good that everyone consumes and nobody can be excluded from in 1954, and showed why the market underprovides it. Mancur Olson's 1965 book explained why large groups with a shared interest fail to act on it while small groups with concentrated stakes do, and Garrett Hardin's 1968 essay on the commons made the overuse of a shared resource the most cited example in the field.

Why it works now

Whether a consortium, a standard, or a shared piece of infrastructure will actually get funded has always been a matter of reading personalities. The distribution of stakes can now be mapped from the record, and former participants in comparable efforts can say why they held or failed. The Desk asks who pays for this and what they get that the free riders do not, and whether a small group has enough at stake to carry it alone.

When it applies

A consortium, standards body, industry association, shared infrastructure, open-source dependency, or joint venture that everyone wants and nobody funds; a common resource being run down.

What we would ask

Who pays for this, and what do they get that the free riders do not?
Is there a small group whose stake is large enough to carry it alone?

What it needs

The distribution of stakes across members, the cost of contributing, and whether contribution can be observed and rewarded.

How we would get it

Former participants in comparable consortia and shared efforts on why they held or failed; the record of who funded what.

Where it breaks

Assuming shared interest produces shared action; ignoring that small groups with concentrated stakes act and large groups with diffuse ones do not.

For example

A thesis depends on an industry standard being adopted. Every firm wants the standard and every firm gains most by waiting for the others to bear the switching cost. The three largest have enough at stake to move alone, and the research is whether they will.

The mathematics in fullEach member contributes c and the group gains B shared by n; a member contributes only when its share B/n exceeds c, so large groups with diffuse benefits underprovide, and a common resource is overused because each user bears 1/n of the damage it causes. Taught in

MMSS 300; MMSS 211-1; MMSS 211-2

Sources

Hardin (1968). The Tragedy of the Commons. Science 162(3859), 1243-1248. https://doi.org/10.1126/science.162.3859.1243
Samuelson (1954). The Pure Theory of Public Expenditure. The Review of Economics and Statistics 36(4), 387. https://doi.org/10.2307/1925895
Olson, Mancur (1965). The Logic of Collective Action. Harvard University Press. (book)

Finding the right people

Who knows, who is central, and who occupies the same position in a neighboring world.

Network centrality

The person the information passed through

Instead of asking who had the highest title, we ask who occupied the position through which the relevant information passed, and we put that person on the panel.

Born to answer
Freeman, 1978; Granovetter on weak ties, 1973.
What it does now
Interviewing the analyst everyone called, not the executive who signed.
The mathematics
Ax = lambda x: central when tied to the central
The full entry
Where it began

Linton Freeman's 1978 paper sorted out what centrality means, distinguishing the person with the most ties from the person who sits on the most paths between others. Mark Granovetter had found in 1973 that Boston job seekers heard about their jobs more often from acquaintances than from close friends, because weak ties reach into other circles. Network analysis was a named part of the MMSS curriculum in the 1990s.

Why it works now

Expert search has been keyword search on titles, which finds the people who signed off and misses the person who ran the model and briefed all of them. The map of who talked to whom about a decision can now be built from the first two interviews and the record, and the panel recruited from its center. The Desk asks whose desk the decision crossed and who everyone called.

When it applies

An expert search that has produced impressive titles and thin knowledge; a question about how a decision was really made inside an organization or an industry.

What we would ask

When this was decided, whose desk did it cross, and who did everyone call?
Who is connected to the most people who would know?

What it needs

A map of who talked to whom about the question, built from the first interviews and from the record.

How we would get it

The Research Director builds the map from the first two interviews and the record, then recruits from its center rather than from its org chart.

Where it breaks

The network is drawn from one starting point and finds only that point's contacts; centrality is confused with seniority.

For example

A client wants to understand a pricing decision at a supplier. The senior executives who signed it off know the outcome. The mid-level analyst who ran the model and briefed all of them knows the reasoning, and is the interview.

The mathematics in fullDegree centrality counts a person's ties; eigenvector centrality solves A x = lambda x so a person is central when tied to central people; the boundary of the network drawn decides whose centrality is being measured. Taught in

MMSS 1995 (network analysis and centrality)

Sources

Freeman (1978). Centrality in social networks conceptual clarification. Social Networks 1(3), 215-239. https://doi.org/10.1016/0378-8733(78)90021-7
Bonacich (1987). Power and Centrality: A Family of Measures. American Journal of Sociology 92(5), 1170-1182. https://doi.org/10.1086/228631
Granovetter (1973). The Strength of Weak Ties. American Journal of Sociology 78(6), 1360-1380. https://doi.org/10.1086/225469

Structural equivalence

The expert on a market that has no experts

When a market is too new or too small to have experts, we look for people who occupy the same structural position in a neighboring industry and have solved the same problem in different clothes.

Born to answer
Lorrain and White, 1971: same position, no direct tie.
What it does now
Finding an expert for a market that has none.
The mathematics
Distance d(i, j) is small when tie patterns match
The full entry
Where it began

Lorrain and White's 1971 paper defined structural equivalence: two people occupy the same position in a network when they have the same pattern of ties to everyone else, whether or not they know each other. The idea let sociologists find roles rather than names, and it is the formal version of noticing that an exchange operator and a cloud capacity allocator are solving the same problem.

Why it works now

A first-of-kind market has no experts, and the search used to stop there. The question can now be restated as a structure, a liquidity manager, an allocator of a scarce shared resource, a two-sided market operator, and the industries where that structure is mature can be searched for the people who have solved it. The Desk asks what problem the business actually solves, stripped of its industry vocabulary, and who has solved that exact problem somewhere else.

When it applies

A first-of-kind market, technology, or business model where the usual expert search returns nobody with direct experience.

What we would ask

What problem does this business actually solve, stripped of its industry vocabulary?
Who has solved that exact problem somewhere else?

What it needs

A structural description of the role or problem, independent of the industry, and a search across industries for the same structure.

How we would get it

The Research Director restates the question as a structure (a liquidity manager, an allocator of a scarce shared resource, a two-sided market operator) and recruits from the industries where that structure is mature.

Where it breaks

The analogy is superficial; the structural match is asserted rather than checked against the mechanics.

For example

A client is underwriting a marketplace for a resource that has never been traded. There are no experts on that marketplace. There are exchange operators, transportation network planners, and cloud capacity allocators who have run the same coordination problem, and two of them are the panel.

The mathematics in fullTwo actors are structurally equivalent when their patterns of ties match, d(i, j) = sqrt of sum over k of (A(i, k) - A(j, k))^2 small, whether or not they are connected to each other; the search is for small d across industries, not for shared keywords. Taught in

MMSS 1995 (structural equivalence and network roles)

Sources

Lorrain, White (1971). Structural equivalence of individuals in social networks. The Journal of Mathematical Sociology 1(1), 49-80. https://doi.org/10.1080/0022250X.1971.9989788

Cultural consensus analysis

Several partial accounts, one reconstructed fact

When nobody saw the whole thing, we interview several people who each saw part, and we recover the fact from the pattern of their agreement rather than from any one account.

Born to answer
Romney, Weller and Batchelder, 1986, in anthropology.
What it does now
Rebuilding what happened from several partial accounts.
The mathematics
Agreement between sources estimates each one's competence
The full entry
Where it began

The cultural consensus model of Romney, Weller, and Batchelder in 1986 solved the anthropologist's problem of informants who each know part of the culture and none of whom can be checked against an answer key: the pattern of agreement among them estimates each one's competence and recovers the consensus. It is, in effect, a method for reconstructing a fact from several partial accounts.

Why it works now

Finding out what really happened inside a company or a deal has always meant interviewing several people who each saw a piece, and then trusting the most vivid one. Parallel structured interviews that ask each about the same facts can now be run and compared systematically, so the overlap scores the accounts and the contested facts are weighted by it. The Desk asks who saw which part and where the accounts overlap.

When it applies

A question about what really happened inside a company, a deal, or a negotiation, where every available source saw only a piece and some have reasons to shade it.

What we would ask

Who saw which part of this, and where do their accounts overlap?
On the parts that overlap, do they agree?

What it needs

Several sources with different vantage points and, ideally, no coordination; a structured interview that asks each about the same facts.

How we would get it

Parallel structured interviews; the synthesis scores each source by agreement with the others on the overlapping facts and weights the contested facts accordingly.

Where it breaks

The sources have coordinated; the overlap is too small to estimate agreement; one source dominates the synthesis by being vivid.

For example

A client wants to know why a competitor abandoned a product line. The former product head, the former CFO, and a former channel partner each saw part of it. On the facts all three saw, two agree and one does not, and the reconstruction weights the contested facts toward the two.

The mathematics in fullWith independent informants, the agreement between each pair estimates each informant's competence, and the consensus answer to each question is recovered by weighting each informant by that competence, without knowing the answers in advance. Taught in

MMSS 1995 (informant accuracy)

Sources

Romney, Weller, Batchelder (1986). Culture as Consensus: A Theory of Culture and Informant Accuracy. American Anthropologist 88(2), 313-338. https://doi.org/10.1525/aa.1986.88.2.02a00020

The gravity model

Who will trade with whom

Interaction between two parties scales with their size and falls with the distance between them, and distance can be commercial, institutional, or cultural as well as geographic. We use it to predict partnerships, encroachment, and where a customer will migrate.

Born to answer
Tinbergen, 1962, predicting trade from size and distance.
What it does now
Which distance actually binds: regulation, systems, or miles.
The mathematics
F = G x M(i)^a x M(j)^b / D^c
The full entry
Where it began

Jan Tinbergen borrowed Newton's gravity in 1962 to predict trade between countries from their sizes and the distance between them, and it fit remarkably well. Anderson and van Wincoop's 2003 paper resolved the border puzzle, why Canadian provinces traded so much more with each other than with equally close American states, by showing that relative trade costs, not miles, are the distance that matters. Gravity models were in the MMSS curriculum of the mid-1990s.

Why it works now

The insight for an investor is that distance can be regulatory, cultural, or institutional, and those distances were never measured because nobody could. Operators can now say which dimension of distance actually binds in a market, and the record of past interactions can be fitted to find out. The Desk asks what the real distance is between two parties, and who is large and close that the plan ignored.

When it applies

A question about which buyers a seller will reach, which partners will form, which competitor will encroach on which territory, or where an expansion will succeed.

What we would ask

Between these two, what is the real distance: geography, regulation, relationships, language, or systems?
Who is large and close, and therefore likely, that the plan has ignored?

What it needs

Sizes of the parties and a defensible measure of the distances between them on the dimensions that matter here.

How we would get it

Operators on which dimensions of distance actually bind in this market; the record of past interactions for the fit.

Where it breaks

Distance is measured on the dimension that is easy to measure rather than the one that binds; the fit is extrapolated outside the data.

For example

A distributor plans expansion into an adjacent region on the strength of proximity. The binding distance is not miles but a certification regime that its current customers do not require and the new ones do; the model, once the right distance is used, points at a different region.

The mathematics in fullF(i, j) = G x M(i)^a x M(j)^b / D(i, j)^c, with D a weighted combination of geographic, commercial, institutional, and network distance; the weights are the research question. Taught in

MMSS 1995 (gravity models)

Sources

Anderson, van Wincoop (2003). Gravity with Gravitas: A Solution to the Border Puzzle. American Economic Review 93(1), 170-192. https://doi.org/10.1257/000282803321455214

Allocating scarce things

Where budget, capacity, slots, and attention should go, often with an exact answer.

Linear programming

The research budget is itself an allocation

Given a budget, a deadline, and a cap on interviews, we choose the set of studies that buys the most decision value inside the constraints, and we tell you which constraint is costing you the most.

Born to answer
Kantorovich, 1939, planning a Leningrad plywood trust.
What it does now
Choosing the studies that buy the most decision inside a deadline.
The mathematics
max sum of EVSI(q) x(q) subject to cost and time
The full entry
Where it began

Leonid Kantorovich invented linear programming in 1939 to plan production at a Leningrad plywood trust, and his 1960 paper in Management Science brought the work to the West; George Dantzig's simplex method made it computable for the United States Air Force after the war. The shadow price, what one more unit of a constraint is worth, was there from the start.

Why it works now

A research engagement is an allocation of a budget and a deadline across candidate studies, and it has never been solved as one. It can be: each study's value to the decision, cost, and time can be laid out and the plan solved, with the shadow price of the deadline reported. The Desk asks what the budget and the deadline really are, and what one more week would be spent on. The report names the workstreams left out and what they would have cost.

When it applies

An engagement, or a client's own diligence, with more candidate questions than time or money; a client asking whether to enlarge the budget.

What we would ask

What is the deadline and what is the budget, really?
If we could add one more week or one more interview, what would you want it spent on?

What it needs

For each candidate study: its value to the decision, its cost, and its time; the binding constraints.

How we would get it

The research plan is solved as an allocation; the shadow price of the deadline and the budget is reported, which is the argument for or against enlarging the engagement.

Where it breaks

The objective is information rather than decision value; a constraint is invented to force the plan.

For example

Six workstreams, a fixed fee, and seventy-two hours. Four fit. The two left out are named in the report with what they would have cost and what they might have changed, so the client can buy them if the decision warrants.

The mathematics in fullChoose x(q) in {0, 1} to maximize the sum of EVSI(q) x x(q) subject to the sum of cost(q) x x(q) at most the budget and the sum of time(q) x x(q) at most the deadline; the shadow price of each constraint is the gain from relaxing it by one unit. Taught in

MECS 560-1; MECN-451

Sources

Kantorovich (1960). Mathematical Methods of Organizing and Planning Production. Management Science 6(4), 366-422. https://doi.org/10.1287/mnsc.6.4.366

Constrained optimization

Write down the constraints and the answer often falls out

Where should the capital, the capacity, the inventory, or the people go? Many such questions have exact answers once the objective and the constraints are written down honestly, and writing them down is most of the work.

Born to answer
Kantorovich and Dantzig; the shadow price of a binding constraint.
What it does now
Where capital and capacity go once the constraints are honest.
The mathematics
At the optimum, gradients of the binding constraints combine
The full entry
Where it began

The same Kantorovich and Dantzig work made allocation a solved problem once the objective and constraints were written down, and MECS teaches the Karush-Kuhn-Tucker conditions that generalize it. The lesson that survived seventy years is that writing the constraints down honestly, including the political ones, is most of the work.

Why it works now

Capital has been allocated across divisions by last year's share because the true returns and the binding constraints were never assembled in one place. They can be assembled now, from the client's data and from operators who know which constraints are real and which are habits, and the formal solution can be run as a benchmark for the judgment. The Desk asks what is being maximized, which constraint binds today, and what one more unit of it would be worth.

When it applies

A capital allocation across plants, products, regions, or funds; a capacity decision; an inventory or network design question dressed up as a judgment call.

What we would ask

What are you maximizing, and what are the constraints you cannot relax?
Which constraint is binding today, and what would one more unit of it be worth?

What it needs

The objective stated honestly; the constraints, including the political ones; the data for the coefficients.

How we would get it

Operators on the real constraints and on which are negotiable; the record for the coefficients; the formal solution as a benchmark for the judgment.

Where it breaks

The objective is mis-stated; a constraint is political and unwritten; the solution is applied without the judgment it was meant to inform.

For example

A company allocates capital to its divisions by last year's share. Written as an allocation with the true returns and the covenant constraints, one division should receive nothing and another twice its share; the shadow price on the covenant is the argument for refinancing.

The mathematics in fullMaximize f(x) subject to g(i)(x) at most zero and h(j)(x) equal to zero; at the optimum the gradient of f is a nonnegative combination of the gradients of the binding constraints, and each multiplier is the value of relaxing its constraint by one unit. Taught in

MECS 560-1; MATH 285

Sources

Kantorovich (1960). Mathematical Methods of Organizing and Planning Production. Management Science 6(4), 366-422. https://doi.org/10.1287/mnsc.6.4.366

Deferred acceptance

Some markets match rather than price

Where the good cannot be priced, who gets which slot, seat, organ, school, or partner is decided by matching rules, and the rules decide who can game the market. We read the rules the way a market designer would.

Born to answer
Gale and Shapley, 1962; Roth's redesign of the physician match.
What it does now
Reading a marketplace's matching rules for who they favor.
The mathematics
Stable when no pair would both rather have each other
The full entry
Where it began

Gale and Shapley's 1962 paper on college admissions and stable marriage gave the deferred acceptance algorithm, which always finds a matching no pair would both want to abandon, and showed that the side that proposes gets its best stable outcome. Alvin Roth and Elliott Peranson used it in 1999 to redesign the match that assigns American medical graduates to residencies, and the design has run every year since.

Why it works now

Marketplace and staffing businesses run matching rules that determine who is happy and who churns, and investors have read the churn as a marketing problem. The rules can now be read as a market designer would, the outcomes measured from the platform's data, and operators of comparable markets asked where the rules were gamed. The Desk asks which side proposes and which side has to accept what it is offered.

When it applies

A marketplace, staffing, or platform business whose economics depend on how it matches the two sides; a hiring or admissions process; an allocation with no price.

What we would ask

Which side proposes, and which side has to accept what it is offered?
Can the participants improve their outcome by misreporting what they want?

What it needs

The matching rule, the preferences on both sides, and the complementarities that break stability.

How we would get it

Operators of comparable matching markets on where the rules were gamed; the record of the platform's matching outcomes.

Where it breaks

Participants misreport strategically; couples and complementarities make no stable match exist; the design favors one side without saying so.

For example

A staffing platform's thesis assumes both sides are happy with its matches. Its rule lets clients propose and forces workers to accept the best offer so far. That is the client-optimal stable match, and the workers' churn, which the thesis calls a marketing problem, is the design.

The mathematics in fullA matching is stable when no pair would both prefer each other to their assigned partners; deferred acceptance produces a stable match that is optimal for the proposing side and worst for the other, so the choice of proposer is a distributional decision. Taught in

MECS 560-1; MECS 470

Sources

Gale, Shapley (1962). College Admissions and the Stability of Marriage. The American Mathematical Monthly 69(1), 9-15. https://doi.org/10.2307/2312726
Roth, Peranson (1999). The Redesign of the Matching Market for American Physicians: Some Engineering Aspects of Economic Design. American Economic Review 89(4), 748-780. https://doi.org/10.1257/aer.89.4.748

Mean-variance portfolio theory

Attractive on its own, wrong for this portfolio

An investment is not evaluated alone. We ask what it adds to what the client already holds, which depends on how it moves with the rest, especially in the bad states.

Born to answer
Markowitz, 1952: risk is covariance, not standalone variance.
What it does now
What a position adds to the book that already exists.
The mathematics
max w'mu - (lambda / 2) w'Sigma w
The full entry
Where it began

Harry Markowitz's 1952 paper showed that what an asset adds to a portfolio depends on how it moves with the rest, not on its own risk, and that a set of positions can be chosen to give the most return for a given variance. It was the beginning of modern portfolio theory and of his Nobel.

Why it works now

The theory has always been applied to listed securities with long price histories and almost never to a private position, a lender, or a portfolio company, whose co-movements with the rest of a book are not in any database. They can now be researched: Advisors can say which exposures co-move in that sector and how comparable positions behaved in past stresses, and the client's book can be captured at scoping. The Desk asks what the client already holds that goes wrong in the same states this does.

When it applies

A position that looks attractive standalone but adds sector, rate, regulatory, or liquidity concentration; a client evaluating the deal rather than the portfolio it enters.

What we would ask

What do you already hold that goes wrong in the same states this does?
In the last stress, how did the closest thing you own to this behave?

What it needs

The client's existing exposures and how this position covaries with them, in stressed periods as well as calm ones.

How we would get it

Advisors on the exposures that co-move in this sector; the record on how comparable positions behaved in past stresses; the client's own book captured at scoping.

Where it breaks

Correlations estimated in calm periods and assumed in stressed ones; the position judged on its own merits alone.

For example

A fund likes a specialty lender. Its book already holds two positions that draw on the same wholesale funding market. On its own the lender is fine; in the portfolio it triples an exposure that no one had named, and the research question becomes that funding market.

The mathematics in fullThe contribution of an asset to portfolio risk depends on its covariance with the portfolio, not on its own variance; maximize w'mu - (lambda / 2) x w'Sigma w with the weights summing to one, and re-estimate Sigma in stressed periods. Taught in

MECN-451; DECS-430

Sources

Markowitz (1952). Portfolio Selection. The Journal of Finance 7(1), 77-91. https://doi.org/10.1111/j.1540-6261.1952.tb01525.x

Dynamics and pure mathematics

How things move over time, how much of a bet to take, and the pure mathematics that decides whether a plan survives the movement.

The cobweb model

Why capacity arrives just as demand leaves

When supply responds to price with a lag, because a ship, a fab, or a plant takes years to build, the market cycles on its own. We ask how long the lag is and where in the cycle the decision sits.

Born to answer
Ezekiel, 1938: farmers planting on last season's price.
What it does now
Shipping, fabs and chemicals, where capacity arrives after the peak.
The mathematics
q(t) = S(p(t - L)), supply on a lagged price
The full entry
Where it began

Mordecai Ezekiel's 1938 paper named the cobweb theorem after the shape the price and quantity trace on a diagram when supply responds to last season's price: farmers plant when prices are high, the crop arrives, prices fall, they plant less, prices rise. The same mathematics governs any market where capacity takes years to build, and difference equations were part of the MMSS curriculum in the 1990s.

Why it works now

The order book of capacity under construction, the single number that tells you where the cycle is, has always existed and rarely been read against demand. It can now be assembled from filings, trade press, and operators quickly, and the history of past cycles retrieved for the turning points. The Desk asks how long it takes for capacity to arrive and how much is on order right now.

When it applies

Shipping, semiconductors, chemicals, agriculture, real estate, mining, or any business where capacity takes years to add and cannot be removed quickly; a thesis that extrapolates today's price.

What we would ask

How long from the decision to add capacity until it arrives, and how much is on order right now?
Where are we in the cycle, and what did the last peak look like?

What it needs

The order book of capacity under construction, the lag, and the history of past cycles.

How we would get it

Operators on the real lead times and the capacity in the pipeline; the record of past cycles and their turning points.

Where it breaks

Treating today's price as a level rather than a point on a cycle; ignoring the capacity already ordered.

For example

Freight rates are at a decade high and a thesis buys shipping equity. The order book for new vessels is at a decade high too, and they arrive in three years. The model says the rates fall as the ships arrive, as they did last time, and the question is only how far.

The mathematics in fullSupply at t responds to the price at t minus the lag: q(t) = S(p(t - L)), while demand clears at the current price; with a lag the system spirals around the equilibrium, converging when supply is less responsive than demand and diverging when it is more. Taught in

MMSS 1995 (difference equations); MMSS 211-1

Sources

Ezekiel (1938). The Cobweb Theorem. The Quarterly Journal of Economics 52(2), 255. https://doi.org/10.2307/1881734

Leontief input-output analysis

What breaks first in the supply chain

Every sector buys from others to make what it sells, so a shock to one sector travels through everyone downstream. We trace the shock through the purchase matrix to find where it lands hardest and first.

Born to answer
Leontief, 1936, mapping what every sector buys from every other.
What it does now
Tracing a tariff or an outage to where it lands hardest.
The mathematics
x = (I - A)^-1 d
The full entry
Where it began

Wassily Leontief built the first input-output table of the American economy in 1936, a matrix of what every sector buys from every other, and showed that a change in demand for one product could be traced through the purchases to its total effect on every sector. He won a Nobel for it, and the matrix models in the 1995 MMSS curriculum descend from it.

Why it works now

The Desk was asked, in a real conversation, what breaks first in a supply chain under a sustained tariff shock, and the honest answer is that the question has a structure. The purchase relationships can now be assembled for the relevant sectors and the shock traced through them, and procurement leads can say where substitutes exist and where they do not. The Desk asks which input comes through a sector that feeds many others and has no substitute.

When it applies

A tariff, sanction, outage, price spike, or bankruptcy in one part of a supply chain; a thesis about which industries a shock helps or hurts; a company whose inputs come through a sector under stress.

What we would ask

Which of your inputs comes through a sector that feeds many others and has few substitutes?
When the shock hits that sector, how many steps away are you, and how fast does it reach you?

What it needs

The purchase relationships among the relevant sectors, with substitution possibilities where they exist.

How we would get it

Operators and procurement leads on the actual sourcing chain; the record for input-output relationships; the analysis traced through the matrix.

Where it breaks

Fixed coefficients when firms substitute; national tables for a global chain; the second-order effects ignored.

For example

A tariff is imposed on one category of industrial input. Traced through the purchase matrix, the sectors hit hardest are not the direct importers but two steps down, in a small sector with no substitute that feeds four larger ones. That sector is the crux, and its operators are the panel.

The mathematics in fullOutput x solves x = A x + d, so x = (I - A)^-1 d; a shock to sector j propagates down column j of the inverse, and the sectors with large entries in that column absorb the shock first and hardest. Taught in

MMSS 1995 (matrix models); MATH 285

Sources

Leontief (1936). Quantitative Input and Output Relations in the Economic Systems of the United States. The Review of Economics and Statistics 18(3), 105. https://doi.org/10.2307/1927837

The Kelly criterion

How much of the bet to take

Given an edge, there is a size that maximizes long-run growth, and betting beyond it reduces growth while raising the chance of ruin. We work out which side of that line the proposed size sits on, and how sure the edge really is.

Born to answer
Kelly at Bell Labs, 1956; Thorp at the blackjack table.
What it does now
The bet size beyond which growth falls and ruin rises.
The mathematics
f* = p - (1 - p) / b
The full entry
Where it began

John Kelly at Bell Labs worked out in 1956 how a gambler with inside information on a noisy channel should size bets to maximize the growth of capital, and Edward Thorp used the result at the blackjack table and then on Wall Street; his 1969 paper is the practitioner's account. Robert Merton's 1969 continuous-time portfolio problem gives the same fraction, excess return over variance, for an investor.

Why it works now

Position sizing has been set by conviction, and conviction runs highest right after a winning year, which is when the edge is most overestimated. The arithmetic of a track record, its edge, variance, and correlation with the rest of the book, can now be assembled quickly, and the growth-optimal fraction compared with the proposed size. The Desk asks how the confidence in the edge was measured and how much of the capital is gone if it loses.

When it applies

Position sizing, allocation to a strategy, a capital commitment sized by conviction, a founder betting the company; any decision where the size of the bet is being set by enthusiasm.

What we would ask

How confident are you in the edge, and how was that confidence measured?
If this loses, how much of the capital is gone, and is that recoverable?

What it needs

An honest estimate of the edge and its uncertainty, the variance of the outcome, and the correlation with other bets.

How we would get it

The record and the Advisors on the true distribution of outcomes for bets of this kind; the client's own book at scoping for the correlations.

Where it breaks

The edge estimated on too few cases; correlated bets treated as independent; a horizon too short for the long-run argument.

For example

A fund with a documented edge in a strategy proposes to double its allocation after a strong year. Sized by the arithmetic of its own track record and variance, the current allocation is already near the growth-optimal fraction, and doubling it lowers expected growth while doubling the chance of a drawdown the fund cannot survive.

The mathematics in fullf* = p - (1 - p) / b for a bet at odds b, or (mu - r) / sigma^2 in the continuous case; growth g(f) is maximized at f* and turns negative beyond about twice it, and because the edge is estimated with error, fractional Kelly is the working rule. Taught in

MATH 385; MECN-451

Sources

Kelly (1956). A New Interpretation of Information Rate. Bell System Technical Journal 35(4), 917-926. https://doi.org/10.1002/j.1538-7305.1956.tb03809.x
Thorp (1969). Optimal Gambling Systems for Favorable Games. Revue de l'Institut International de Statistique / Review of the International Statistical Institute 37(3), 273. https://doi.org/10.2307/1402118
Merton (1969). Lifetime Portfolio Selection under Uncertainty: The Continuous-Time Case. The Review of Economics and Statistics 51(3), 247. https://doi.org/10.2307/1926560

The SIR epidemic model

Will it spread, or fizzle

Anything that spreads by contact spreads only if each carrier passes it to more than one other on average. We estimate that number and the size of the pool it can spread through, because below the threshold the early numbers mean nothing.

Born to answer
Kermack and McKendrick, 1927.
What it does now
Whether adoption or default actually spreads, or fizzles.
The mathematics
R0 = b / g; spread needs R0 x S above one
The full entry
Where it began

Kermack and McKendrick's 1927 model divided a population into the susceptible, the infected, and the recovered and derived the threshold: an epidemic takes off only when each case produces more than one new case. The same equations describe the spread of products, practices, and defaults, and Frank Bass's 1969 diffusion model is a cousin with the imitation term made explicit.

Why it works now

The pass-along rate, the one number that decides whether a thing spreads, used to be a marketing department's hope. It can now be measured from a company's own data, outside the enthusiast cohort, and compared with the threshold; the susceptible pool can be sized from the record. The Desk asks how many others each adopter brings in, measured rather than hoped.

When it applies

A product, practice, default, or rumor spreading through a population; a viral growth thesis; a credit book where one default triggers others; an adoption that has stalled.

What we would ask

For each customer who adopts, how many others do they bring in, measured rather than hoped?
How much of the reachable population is still susceptible?

What it needs

The measured pass-along rate, the recovery or churn rate, and the size and structure of the susceptible pool.

How we would get it

The client's own data for the pass-along rate; operators on comparable spreads; the record for how comparable contagions ended.

Where it breaks

Uniform mixing assumed when the network has hubs; the susceptible pool overestimated; the pass-along rate measured in the enthusiast segment.

For example

A consumer app reports that each user invites two others, and a thesis extrapolates. Measured outside the launch cohort, each user brings in nought point eight others, below the threshold, and the growth is paid acquisition wearing a viral costume.

The mathematics in fulldI/dt = b S I - g I with R0 = b / g; spread takes off only when R0 times the susceptible share exceeds one, and the final size is a fixed point strictly less than everyone. Taught in

MS&E 221 (Stanford); MMSS 1995 (exponential growth); MECN-451

Sources

Kermack, McKendrick (1927). A contribution to the mathematical theory of epidemics. Proceedings of the Royal Society of London. Series A, Containing Papers of a Mathematical and Physical Character 115(772), 700-721. https://doi.org/10.1098/rspa.1927.0118
Bass (1969). A New Product Growth Model for Consumer Durables. Management Science 15(5), 215-227. https://doi.org/10.1287/mnsc.15.5.215

Threshold models and increasing returns

Critical mass and lock-in

When each person's choice depends on how many others have chosen, a market can sit still for years and then flip in months, and once flipped it may not flip back. We find the threshold and how close the market is to it.

Born to answer
Schelling, 1971; Granovetter, 1978; Arthur, 1989.
What it does now
How close a platform market is to tipping, and whether it tips back.
The mathematics
Equilibrium is a fixed point of the threshold distribution
The full entry
Where it began

Thomas Schelling's 1971 model of segregation showed that mild individual preferences about neighbors can tip a whole neighborhood, and Mark Granovetter's 1978 threshold model generalized it: each person acts when enough others have, and the outcome depends on the whole distribution of thresholds. Brian Arthur's 1989 paper on competing technologies added increasing returns and showed that history, not merit, can decide which technology locks in.

Why it works now

Platform and standards contests have been called from the quality of the products, which the models say is often not what decides them. The adoption thresholds of the participants can now be researched from the customers themselves, and the current adoption level measured against them, so that the distance to the tip is an estimate rather than a feeling. The Desk asks how many others need to have adopted before the marginal customer will.

When it applies

A platform, standard, technology, or network good where value rises with adoption; a market that seems stuck; a thesis about a challenger displacing an entrenched incumbent.

What we would ask

How many others need to have adopted before the marginal customer will?
Once the market has tipped, what would it take to tip it back?

What it needs

The distribution of adoption thresholds across the population, the current adoption level, and the strength of the increasing returns.

How we would get it

Customers on the adoption level at which they would switch; operators on comparable tips; the record for the adoption curve and its inflection.

Where it breaks

Assuming a superior product wins when the incumbent has already locked in; assuming the tip is reversible.

For example

A challenger payment network has a better product and a fifth of the merchants. Merchants adopt when most of their customers use it and customers when most merchants take it. The threshold model says the market tips only if the challenger reaches about forty percent of one side, and the research is whether any path gets it there.

The mathematics in fullEach actor adopts when the adopting share exceeds its personal threshold; the equilibrium share is a fixed point of the cumulative threshold distribution, and with increasing returns two stable equilibria exist, so the outcome depends on history, not on which technology is better. Taught in

MMSS 311-2; MECN-451

Sources

Schelling (1971). Dynamic models of segregation. The Journal of Mathematical Sociology 1(2), 143-186. https://doi.org/10.1080/0022250X.1971.9989794
Granovetter (1978). Threshold Models of Collective Behavior. American Journal of Sociology 83(6), 1420-1443. https://doi.org/10.1086/226707
Arthur (1989). Competing Technologies, Increasing Returns, and Lock-In by Historical Events. The Economic Journal 99(394), 116. https://doi.org/10.2307/2234208

System dynamics

Why inventories, hiring, and capacity oscillate

Decisions are made on a reading of the system that reflects decisions made a delay ago, and people running such systems overshoot in a predictable pattern. We measure the delay and ask whether the plan accounts for it.

Born to answer
Forrester, 1958; Sterman's beer distribution game, 1989.
What it does now
Why inventories, hiring and capacity oscillate on their own.
The mathematics
d(Stock)/dt = inflow - outflow, set on a delayed reading
The full entry
Where it began

Jay Forrester's 1958 article founded system dynamics with an industrial supply chain in which a small change in retail demand produced wild swings in factory output, entirely because of delays in the ordering rules. John Sterman's 1989 beer distribution game showed that intelligent people playing that supply chain reproduce the oscillation every time, and Lee, Padmanabhan, and Whang named the amplification the bullwhip effect in 1997.

Why it works now

The delay between a decision and its visible effect is the number that explains most oscillation in inventories, hiring, and capacity, and it has rarely been measured because nobody thought of the rule of thumb as a model. Operators can now be asked for the delays and the rules they manage by, and the client's own order and inventory data reveal the cycle the rule produces. The Desk asks how long it takes for a decision's effect to show up in the numbers the client manages by.

When it applies

Inventory swings, hiring that alternates between freezes and sprees, capacity that arrives after demand has turned, a supply chain whose orders swing more than its sales.

What we would ask

How long between a decision here and the point when its effect is visible in the numbers you manage by?
When the numbers last turned, did the plan respond to the level or to the change?

What it needs

The structure of the decision rule, the delay, and the history of the stock and the flows.

How we would get it

Operators on the actual delays and the rules of thumb they manage by; the client's own data on orders, inventory, and demand.

Where it breaks

Treating a delayed system as immediate; blaming demand for oscillation the decision rule created.

For example

A distributor's orders to its suppliers swing three times as much as its sales, and the thesis blames volatile customers. The customers are steady. The distributor reorders on a delayed reading of stock, and the swing is the arithmetic of that rule. The fix is the rule, not the customers.

The mathematics in fulld(Stock)/dt = inflow(t) - outflow(t), with the inflow set on a reading of the stock that is D periods old; the delay turns a stabilizing rule into an oscillating one, and the variance of orders amplifies at each stage upstream. Taught in

MMSS 1995 (difference equations); MECN-451

Sources

Sterman (1989). Modeling Managerial Behavior: Misperceptions of Feedback in a Dynamic Decision Making Experiment. Management Science 35(3), 321-339. https://doi.org/10.1287/mnsc.35.3.321
Lee, Padmanabhan, Whang (1997). Information Distortion in a Supply Chain: The Bullwhip Effect. Management Science 43(4), 546-558. https://doi.org/10.1287/mnsc.43.4.546
Forrester, Jay W. (1958). Industrial Dynamics: A Major Breakthrough for Decision Makers. Harvard Business Review, 36(4), 37-66. (book)
Sterman, John D. (2000). Business Dynamics: Systems Thinking and Modeling for a Complex World. McGraw-Hill. (book)

Queueing theory and Little's law

Capacity, waiting, and the cliff at ninety percent

Whatever is waiting obeys one identity, the number in the system equals the arrival rate times the time each spends there, and near full capacity waiting does not rise gently but explodes. We check the utilization the plan assumes.

Born to answer
Little, 1961: an identity true of any stable queue.
What it does now
The waiting-time cliff a high-utilization plan walks into.
The mathematics
L = lambda x W; wait grows like rho / (1 - rho)
The full entry
Where it began

John Little proved in 1961 that for any stable queue, whatever its internal rules, the average number waiting equals the arrival rate times the average time each spends in the system. The identity is exact and needs no assumptions about the process, which is why it is the one piece of queueing theory every operations course teaches, and why the cliff near full utilization surprises people who have not seen the curve.

Why it works now

Service operations have been praised for high utilization by analysts who never saw the wait-time curve. Arrival and service rates and their variability can now be pulled from a company's own systems, and operators can say how the arrivals actually cluster. The Desk asks at what utilization the operation runs and what happens to waiting when demand rises ten percent.

When it applies

A service business, a hospital, a call center, a fulfillment operation, a software team, or any plan that runs an operation near full capacity; a thesis that credits high utilization as efficiency.

What we would ask

At what utilization does this operation run, and what happens to waiting time when demand rises ten percent?
How variable are the arrivals?

What it needs

Arrival rates, service rates, their variability, and the utilization at which the operation runs.

How we would get it

Operators on the actual utilization and its variability; the client's own throughput and wait data.

Where it breaks

Steady arrivals assumed when they cluster; a plan that runs at a utilization where the mathematics guarantees queues.

For example

A thesis praises a clinic chain for running its rooms at ninety-five percent utilization. At that utilization the wait for a room is about twenty times the visit length whenever arrivals bunch, which they do every Monday. The utilization is not efficiency; it is the reason patients leave.

The mathematics in fullL = lambda x W for any stable system; with utilization rho = lambda / mu, the expected wait in the simplest queue grows like rho / (1 - rho), so from ninety to ninety-five percent utilization the wait roughly doubles. Taught in

MS&E 221 (Stanford); MECS 470

Sources

Little (1961). A Proof for the Queuing Formula: L = lambda W. Operations Research 9(3), 383-387. https://doi.org/10.1287/opre.9.3.383

Lotka-Volterra and Lanchester's square law

Share wars and the square law

When two rivals grind against each other, the outcome depends on relative strength in a way that is far from linear, and the dynamics can oscillate rather than settle. We work out which regime this contest is in before the client commits to a fight.

Born to answer
Volterra, 1926, on Adriatic fish; Lanchester, 1916, on force.
What it does now
Whether outspending a rival buys share in proportion, or its square.
The mathematics
Coupled dx/dt and dy/dt; strength counts as its square
The full entry
Where it began

Vito Volterra wrote the predator-prey equations in 1926 to explain why fish populations in the Adriatic oscillated, and Alfred Lotka had derived the same system a year earlier in a book on physical biology. Frederick Lanchester, an aeronautical engineer, published the square law of concentrated force in 1916: when both sides fire at each other, an advantage in numbers counts as its square. The law was later applied to market share contests.

Why it works now

Share wars have been fought on the belief that spend buys share linearly. The relative strengths of the rivals on the dimension that decides share, and the history of past contests in the market, can now be assembled quickly, and former executives of both sides can say how those contests actually resolved. The Desk asks whether the larger side gains proportionally more than its size when both spend, and whether the contest has oscillated before.

When it applies

A market-share contest between two large rivals, a price war, a platform fighting a challenger, or a plan to win by spending more; a competitive relationship that cycles.

What we would ask

When you and your rival both spend to take share, does the larger side gain proportionally more than its size, or only its share?
Has this contest oscillated before, and what set the period?

What it needs

The relative strength of the rivals on the dimension that decides share, and the history of past contests.

How we would get it

Former executives of both rivals on how past contests actually resolved; the record for share over time.

Where it breaks

Assuming a linear relationship between spend and share; ignoring that the losing side can withdraw rather than fight.

For example

A challenger with a third of the incumbent's marketing budget plans to win share by outspending in one region. Under the square law of concentrated force, the incumbent's advantage in a head-on contest is nine to one, not three to one; the challenger's only winning move is a region where the incumbent is not concentrated.

The mathematics in fullCoupled equations dx/dt = f(x, y), dy/dt = g(x, y): predator-prey systems produce cycles rather than a settled share, and under the square law losses are proportional to the opponent's total strength, so an advantage in numbers counts as its square. Taught in

MMSS 1995 (difference equations); MECN-441

Sources

Volterra (1926). Fluctuations in the Abundance of a Species considered Mathematically1. Nature 118(2972), 558-560. https://doi.org/10.1038/118558a0
Lanchester, F. W. (1916). Aircraft in Warfare: The Dawn of the Fourth Arm. Constable. (book)
Lotka, Alfred J. (1925). Elements of Physical Biology. Williams and Wilkins. (book)

Ornstein-Uhlenbeck and the random walk

Is this trend a trend

Some quantities wander with no memory and some are pulled back toward a level. We test which this one is, because a margin, a valuation, or a spread that reverts is an opportunity with a known half-life, and one that wanders is not.

Born to answer
Uhlenbeck and Ornstein, 1930; Black and Scholes, 1973.
What it does now
Whether a margin above its history comes back, and how fast.
The mathematics
dx = theta (mu - x) dt + sigma dW
The full entry
Where it began

Uhlenbeck and Ornstein's 1930 paper on Brownian motion gave the equation for a quantity that wanders but is pulled back toward a level, with a half-life for the return. The random walk with no pull is the model underneath Black and Scholes's 1973 option pricing, and the difference between the two is the question every claim of a trend has to answer first.

Why it works now

Whether a margin or a multiple is reverting or has moved for good has been decided by rhetoric, because estimating the pull needs a long history assembled carefully. The history can now be assembled and the half-life estimated, and operators can say whether the structural driver of the level has actually changed. The Desk asks how long it took to come back the last time the quantity moved this far, and whether it did.

When it applies

A margin, multiple, spread, or market share above or below its history; a thesis that extrapolates a run; a claim that something has permanently changed.

What we would ask

When this quantity has moved this far from its history before, how long did it take to come back, and did it?
What would have to be true for the old level to no longer be the level?

What it needs

A long enough history to estimate the pull toward the mean and its half-life, and a view on whether the regime has changed.

How we would get it

The record for the series; operators on whether the structural drivers of the level have changed; the estimate of the half-life.

Where it breaks

The pull estimated on a window too short to tell it from zero; a regime change that moved the level.

For example

A company's margins are five points above their twenty-year average and a thesis capitalizes them. Over that history, departures of this size reverted with a half-life of about two years, and every claim that this time was different accompanied one of them. The research is whether the driver has actually changed.

The mathematics in fullRandom walk: x(t + 1) = x(t) + e, no level is home; mean reversion: dx = theta (mu - x) dt + sigma dW with half-life ln 2 / theta; the question is whether theta is distinguishable from zero on the data available. Taught in

STATS 217 (Stanford); MATH 385; MATH 386-1

Sources

Uhlenbeck, Ornstein (1930). On the Theory of the Brownian Motion. Physical Review 36(5), 823-841. https://doi.org/10.1103/PhysRev.36.823
Black, Scholes (1973). The Pricing of Options and Corporate Liabilities. Journal of Political Economy 81(3), 637-654. https://doi.org/10.1086/260062

Sensitive dependence on initial conditions

How far ahead anyone can see

In a system where small differences grow, forecasts have a horizon beyond which they are worthless whatever the model, and the honest product is a spread of futures rather than one. We ask where the horizon is for this question.

Born to answer
Lorenz, 1963: a weather model diverging from a rounding error.
What it does now
How far ahead a forecast in this market is worth anything.
The mathematics
Separation grows like d x exp(lambda t)
The full entry
Where it began

Edward Lorenz found in 1963 that a simple model of the atmosphere run twice from starting points that differed in the third decimal place produced completely different weather within weeks. Sensitive dependence on initial conditions gave every forecast a horizon beyond which it is worthless, and weather forecasting responded by running ensembles of perturbed starts and reporting the spread.

Why it works now

Long-range business forecasts are still delivered as single paths to two decimal places, decades after meteorology stopped doing that. The track record of past forecasts by horizon in any market can now be assembled from the record, which shows where the horizon actually sits, and a model can be rerun from perturbed starts in minutes. The Desk asks how far ahead anyone in this market has ever forecast accurately.

When it applies

A long-range forecast presented with precision; a plan that depends on a point estimate five years out; an argument between two models about a distant future.

What we would ask

How far out has anyone in this market forecast accurately, ever?
If you ran the same model from slightly different starting assumptions, how quickly would the paths diverge?

What it needs

The track record of forecasts in this domain at each horizon, and the sensitivity of the outcome to initial conditions.

How we would get it

The record of past forecasts against outcomes by horizon; Advisors on what actually decided past outcomes; the report gives a spread, not a point, beyond the horizon.

Where it breaks

Precision beyond the horizon; a single path where an ensemble was honest.

For example

A ten-year demand forecast for a commodity is presented to two decimal places. Past ten-year forecasts in the same market were wrong by a factor of two in either direction, and the errors were not the forecasters' fault. The decision is redesigned to be robust across the spread rather than optimized for the point.

The mathematics in fullIn a chaotic system two trajectories that start a distance d apart separate like d x exp(lambda t); the forecast horizon is roughly the time for that separation to reach the size of the quantity itself, and an ensemble of runs from perturbed starts is the honest forecast. Taught in

MATH 285; MECN-451

Sources

Lorenz (1963). Deterministic Nonperiodic Flow. Journal of the Atmospheric Sciences 20(2), 130-141. https://doi.org/10.1175/1520-0469(1963)020<0130:DNF>2.0.CO;2

Extreme value theory

The distribution of the worst case

The largest loss in a hundred periods has its own distribution, and it is not the one the average describes. We fit the tail from the few largest observations and report the level exceeded once in a horizon, with the honest width of that estimate.

Born to answer
Fisher and Tippett, 1928: the maximum has a law of its own.
What it does now
The drawdown to size limits against, not the worst one yet seen.
The mathematics
Maxima converge to Gumbel, Frechet, or Weibull
The full entry
Where it began

Fisher and Tippett proved in 1928 that the largest of many observations converges to one of only three distributions, whatever the observations came from, which is the theorem that lets the hundred-year flood be estimated from thirty years of data. The Embrechts, Kluppelberg, and Mikosch textbook brought the theory to insurance and finance, and the 1987 paper on self-organized criticality showed why avalanche sizes in many systems follow power-law tails.

Why it works now

Risk limits have been set to the worst observed loss, which is not the worst possible loss and is not even an estimate of it. The tail can now be fitted from the extremes in the client's history and comparable series, and reported with the honest width of the estimate, and operators can say what preceded past extremes. The Desk asks what the worst has ever been and how many observations that rests on.

When it applies

Maximum drawdown, the hundred-year event, a covenant or margin call triggered by an extreme, a plan sized to survive a bad year but not the worst year.

What we would ask

What is the worst this has ever been, and how many observations is that based on?
Do the bad periods cluster?

What it needs

Enough history to observe several extremes, and a view on whether the process generating them has changed.

How we would get it

The record for the extremes in this and comparable series; operators on what preceded past extremes; the fitted tail with its interval.

Where it breaks

Extrapolating far beyond the data; treating clustered extremes as independent; a process that changed.

For example

A fund sizes leverage to survive the worst monthly loss in its fifteen-year history. Fitted to the tail, the once-in-fifty-years loss is nearly twice the worst observed, and the interval around it is wide. The leverage is set to the interval, not the history.

The mathematics in fullThe maximum of many draws converges to one of three families indexed by a tail parameter; the return level is the value exceeded once in T periods, read from the fitted tail, and its confidence interval is wide because it rests on the handful of largest observations. Taught in

MATH 385; MECN-451

Sources

Fisher, Tippett (1928). Limiting forms of the frequency distribution of the largest or smallest member of a sample. Mathematical Proceedings of the Cambridge Philosophical Society 24(2), 180-190. https://doi.org/10.1017/S0305004100015681
Embrechts, Klüppelberg, Mikosch (1997). Modelling Extremal Events. Springer (book). https://doi.org/10.1007/978-3-642-33483-2
Bak, Tang, Wiesenfeld (1987). Self-organized criticality: an explanation of the 1/f noise. Physical Review Letters 59(4), 381-384. https://doi.org/10.1103/PhysRevLett.59.381

Principal components

What really drives the portfolio

A portfolio, a product line, or a P&L that seems to have twenty moving parts often has two or three that explain most of the movement. We find them, because risk that looked diversified across twenty names is often concentrated in one factor.

Born to answer
Hotelling, 1933: the few directions carrying most of the variation.
What it does now
Forty positions that turn out to be one bet.
The mathematics
Sigma v = lambda v
The full entry
Where it began

Harold Hotelling's 1933 paper on principal components found the few directions in which a set of measurements varies most, and showed how much of the total variation each explains. Harry Markowitz's portfolio theory made the same point from the other side: what an asset adds to a portfolio depends on how it moves with the rest. Between them, a portfolio of forty positions can be one bet.

Why it works now

Family offices and funds have believed themselves diversified because they held many names, and the decomposition that would have told them otherwise was a quant's job. It can now be run on any book in minutes, and Advisors can say in economic terms what the leading factor actually is. The Desk asks what moved together in the client's worst month.

When it applies

A portfolio or business with many exposures; a diversification claim; a P&L whose swings nobody can attribute; a set of regions or products that all move together.

What we would ask

When your worst month happened, what moved together?
If one factor explained most of your variance, what would you guess it is?

What it needs

A history of the components long enough to estimate their co-movement.

How we would get it

The client's own data decomposed; Advisors on what the leading factor actually is in economic terms.

Where it breaks

Reading components as causes; a decomposition that changes with rescaling; correlations estimated in calm periods.

For example

A family office holds forty positions across sectors and believes it is diversified. Two components explain most of the variance, and both load on the same interest-rate exposure. The forty positions are one bet, and the research is what that bet actually is.

The mathematics in fullSigma v = lambda v; the first component is the direction of greatest variance and lambda1 / sum of lambdas its share; loadings show what moves together, and a few components explaining most of the variance means the exposures are fewer than they look. Taught in

MATH 285; MMSS 1995 (multidimensional scaling)

Sources

Hotelling (1933). Analysis of a complex of statistical variables into principal components. Journal of Educational Psychology 24(6), 417-441. https://doi.org/10.1037/h0071325
Markowitz (1952). Portfolio Selection. The Journal of Finance 7(1), 77-91. https://doi.org/10.1111/j.1540-6261.1952.tb01525.x

Benford's law

Numbers that were typed rather than measured

Quantities that arise from real processes have leading digits that follow a known distribution, with one appearing about thirty percent of the time. Numbers invented by people do not. We use the pattern as a screen for figures that deserve a closer look.

Born to answer
Newcomb, 1881, on worn logarithm tables; Benford, 1938.
What it does now
A screen for figures typed by a person rather than produced.
The mathematics
P(first digit = d) = log10(1 + 1/d)
The full entry
Where it began

Simon Newcomb noticed in 1881 that the early pages of logarithm tables were dirtier than the later ones, and Frank Benford documented the law of leading digits across twenty thousand numbers in 1938. Theodore Hill's 1995 paper gave it a proper derivation, and accountants adopted it as a screen because figures invented by people do not follow the pattern.

Why it works now

A data room now arrives with more figures than any team can check by hand. The digit screen can be run across all of them in seconds, and the categories that deviate become the place the Advisors are asked to look first. It is a screen, not a conclusion, and the Desk says so. The Desk asks which figures were produced by a process and which were entered by a person.

When it applies

A data room, a set of financial statements, expense reports, or operating figures whose provenance is uncertain; a diligence process with more numbers than time to check them.

What we would ask

Which of these figures were produced by a process and which were entered by a person?
Which categories, if any, deviate from the expected digit pattern?

What it needs

A large enough set of figures spanning several orders of magnitude, from a process that should produce the pattern.

How we would get it

The client's data screened for the pattern; the deviations, if any, handed to the Advisors as the place to start asking questions.

Where it breaks

Applying the screen to figures that have no reason to follow the pattern, such as assigned numbers or figures with a floor and a cap; treating a deviation as proof rather than as a place to look.

For example

A buyer receives five years of monthly operating figures from a target. The leading digits follow the expected pattern in every category but one, where the pattern breaks in the two years before the sale. That category is where the operators are asked to look first.

The mathematics in fullP(first digit = d) = log10(1 + 1/d), so one leads about thirty percent of the time and nine under five percent; the law holds for quantities that are scale-invariant or arise from multiplicative processes, and departures are a screen, not a conclusion. Taught in

MATH 385

Sources

Hill (1995). A Statistical Derivation of the Significant-Digit Law. Statistical Science 10(4). https://doi.org/10.1214/ss/1177009869

The Library, written down.

We keep our method in the open. The Decision Library lists every structure, each with where it began, why it works now, when it applies, what we would ask, what evidence would settle it, how we would get it, where the method breaks, a worked example, the mathematics underneath in symbols, and the papers it rests on. No equation is evaluated there and no answer is given. The mathematics underneath is set out at How the models work (45 structures), and every source is listed on the reading list (122 papers, 28 books).

Bring us the decision