Most consequential questions are a known kind of problem wearing industry clothing. We identify the kind, say what evidence would actually settle it, and then go and get that evidence: structured interviews with the people who have run it, the published record, and where they are possible, surveys and experiments. This page shows the thinking. The Desk applies it to your decision.
The founder studied formal social science in Northwestern's Mathematical Methods in the Social Sciences program from 1989 to 1993, where his game theory professor was Roger Myerson, later a Nobel laureate. Across that program, the economics department, and later the University of Chicago Booth School of Business, he studied under six Nobel laureates. Not one of them gave him an A. The models they taught, and the decision sciences tradition of Kellogg's Managerial Economics and Decision Sciences department that ran alongside, were always the right way to think about a consequential decision. They were rarely usable, because the variables they needed could not be observed at the price of a research engagement. That is what changed. The Library restates those models as research moves, and Vista does the part general AI cannot: it goes and gets the missing evidence.
When you describe a decision to the Desk on our homepage, it is matching what you tell it against a library of formal structures a decision can have: the ways a question can turn out to be a known kind of problem. When one fits, the Desk may name the model it comes from, says the move in plain English, asks the one question that matters next, and tells you what evidence would settle it and how a Research Director would go and get it. It does not compute your answer, it does not give you a probability, and it does not tell you what to do. It shows you the shape of your question and what it would take to answer it. Most questions do not need a model at all, and the Desk does not invent one. Either way the engagement is the same: the right people interviewed, the record swept, one report with the transcripts. That is the beginning of every engagement, and it is free.
Each one gives the model as it is taught, what it was originally built to answer, what it earns its keep on now, and the mathematics in one line. Open an entry for the full treatment: where it began, when it applies, what we would ask, what evidence would settle it, where the method breaks, and the sources.
Whether to act, wait, stage, or learn more, and how much any of it is worth.
We separate what the bet is worth on average from what it is worth to you, given the size of the loss you cannot absorb.
EU = sum of p(s) x U(X(s))Daniel Bernoulli posed it in 1738 as the St. Petersburg game: a coin is flipped until it lands heads, and the pot doubles with every tail. The expected payout is infinite, and nobody will pay more than a few coins to play. His answer, that the value of money to a person bends, was the birth of utility. Pratt's 1964 paper turned the bend into a measurable number, the coefficient of risk aversion, which is why one investor's acceptable bet is another's ruin.
Why it works nowThe bend was always real; what was missing was the distribution to apply it to. A base case with a bull and a bear beside it is not a distribution. Today the interviews can be designed to elicit ranges and tails from operators, the record can be read at scale for how often the bad state actually arrived, and a fund's own constraints can be captured at scoping. The arithmetic is unchanged and still done outside the conversation. What changed is that the inputs are obtainable inside a research engagement instead of being assumed.
When it appliesTwo options have similar average payoffs but very different downsides, or a fund is sizing a position whose worst case is permanent impairment.
What we would askWhat is the loss that would actually change how you operate, not the loss that would merely hurt?
Is the downside a bad quarter or a permanent impairment?
A distribution of outcomes, not a point estimate; the client's real constraints (mandate, concentration, liquidity, career risk).
How we would get itOperators and authorities on the outcome distribution, especially the tail; the client's own constraints captured at scoping rather than assumed.
Where it breaksProbabilities asserted rather than derived. Utility used as a cover for whatever the client already wanted to do.
For exampleTwo private-credit positions carry the same expected return. One has a small chance of a total loss that would breach a concentration limit and trigger redemptions; the other does not. On average they are twins. To this fund they are not, and the research question becomes the size and shape of that tail, not the average.
The mathematics in fullEV = sum over states of p(s) x X(s); the decision uses EU = sum of p(s) x U(X(s)) with U concave, and the gap between EV and the certainty equivalent is the price of the risk.
Taught inMMSS 300; DECS-430; MECN-451; MECS 550-1
SourcesPratt (1964). Risk Aversion in the Small and in the Large. Econometrica 32(1/2), 122. https://doi.org/10.2307/1913738
Kahneman, Tversky (1979). Prospect Theory: An Analysis of Decision under Risk. Econometrica 47(2), 263. https://doi.org/10.2307/1914185
We lay out the choices and the unknowns in the order they will actually arrive, so that doing it in stages, or finding something out first, shows up as an option with a value.
V = max over branches, solved from the leaves backHoward Raiffa's 1968 lectures built decision analysis around a wildcatter deciding whether to drill an oil well. The tree had two branches until Raiffa added a third: pay for a seismic survey first. Drawn that way, the survey is a decision node placed before the drilling decision, and its value is visible on the page. Kellogg still teaches trees with cases of the same shape.
Why it works nowTrees were used sparingly because filling in the chance nodes required data nobody had time to gather. Now the branches an investor actually faces, tranche, pilot, milestone, wait, can be enumerated in one conversation, and the chance nodes populated from comparable cases retrieved from the record and from Advisors who have run the staged version. The Desk draws the tree in words. The Research Director fills it from the interviews. Nobody in the conversation solves it.
When it appliesA single yes-or-no decision is on the table but the client could in fact invest in tranches, run a pilot, wait for a milestone, or research first.
What we would askWhat do you learn between now and the next point at which you could change course?
Is there a version of this where you commit less now and more later?
The real sequence of decisions and information events, with what each intermediate result would reveal.
How we would get itAdvisors who have run the staged version of this decision; the record on how comparable milestones resolved.
Where it breaksThe tree omits the branch that mattered, usually wait or stage; the chance-node probabilities are asserted.
For exampleA growth investor sees invest twenty-five million or pass. Drawn as a tree, a third branch appears: fund ten now against a retention milestone in two quarters, which is worth more than either original branch precisely because the milestone resolves the crux variable before the rest of the money goes in.
The mathematics in fullAt chance nodes V = sum of p(i) x V(i); at decision nodes V = max over branches; solved from the leaves back, so an information node placed before a decision node raises the value of the tree.
Taught inDECS-433; MECN-451; MMSS 300
SourcesRaiffa, Howard (1968). Decision Analysis: Introductory Lectures on Choices under Uncertainty. Addison-Wesley. (book)
Bertsimas, Dimitris and Freund, Robert M. (2004). Data, Models, and Decisions: The Fundamentals of Management Science. Dynamic Ideas. (book)
Before we research anything we ask whether any result could change what you do. If no result would, the research is worth nothing, however interesting.
EVPI = E[max U] - max E[U]Ronald Howard's 1966 paper gave the idea its name: information is worth the difference between the decision you would make with it and the one you would make without it. The teaching case in that tradition is the same wildcatter, asking how much a seismic survey is worth before drilling, and the answer is sometimes zero, because no survey result would change the decision to drill.
Why it works nowThe idea was easy to state and hard to use, because it requires knowing what each possible answer would do to the decision, and that took a modeling exercise nobody commissioned. Today the Desk can work a client's decision back to the questions whose answers would flip it in the first conversation, and the whole engagement can be organized around those questions instead of around a topic. This is the single most commercially important idea in the Library, and it now costs a conversation rather than a study.
When it appliesA client has a list of open questions and is about to work through them in order, or is proposing research that could not change the decision whatever it found.
What we would askWhich result, if we came back with it, would make you not do this?
And which would make you do it larger?
The decision and its alternatives stated explicitly; the plausible results of each proposed study and what each would trigger.
How we would get itThis is scoping, not fieldwork: the Research Director works the decision back to the questions whose answers would flip it, then designs the study around those.
Where it breaksThe decision is already made and the study is theater; the set of actions was drawn too narrow for anything to flip.
For exampleAn investment team has forty diligence questions on a software company. Thirty-one of them cannot change the decision at any plausible answer. Two can. The engagement becomes those two, and the report says so.
The mathematics in fullEVPI = E[max over a of U(a, s)] - max over a of E[U(a, s)]; a real study is worth EVSI = E over results of [max over a of E[U | result]] minus the no-study value, and is pursued only when EVSI exceeds its cost.
Taught inMECN-451; DECS-433; MMSS 300
SourcesHoward (1966). Information Value Theory. IEEE Transactions on Systems Science and Cybernetics 2(1), 22-26. https://doi.org/10.1109/TSSC.1966.300074
Howard (1988). Decision Analysis: Practice and Promise. Management Science 34(6), 679-695. https://doi.org/10.1287/mnsc.34.6.679
Raiffa, Howard (1968). Decision Analysis: Introductory Lectures on Choices under Uncertainty. Addison-Wesley. (book)
We rank the unknowns by how much each would move the decision, per dollar and per day, and start with the top of that list rather than the top of yours.
q* = argmax (EVSI - cost) / timeHoward's 1988 account of decision analysis in practice describes the cycle used at large companies: identify the uncertainties, find which ones the decision is sensitive to, and spend the analysis budget on those, resolving the most valuable first. The ranking, not the analysis, was the contribution; it told the team where to stop reading and start asking.
Why it works nowThe ranking used to require a quantitative model of the decision before any research began, so it was reserved for the largest capital projects. Now a client's forty open questions can be laid against the decision in one sitting, each scored by what its plausible answers would change per week of effort, and the plan can be re-ranked as each answer arrives. The Research Director does the ranking; the Desk shows the client that it will be done, which is often the first thing that distinguishes Vista from a call.
When it appliesTime or budget will not cover every open question, or the client is spending on the questions that are easiest to answer rather than the ones that matter.
What we would askIf you could resolve only one uncertainty before the committee meets, which one would you choose, and why that one?
What does it cost, in time as well as money, to answer each of the others?
For each candidate question: what it would cost, how long it would take, and what its plausible answers would do to the decision.
How we would get itA ranked research plan is the first deliverable; the fieldwork follows the ranking, and the plan is re-ranked as answers arrive.
Where it breaksThe ranking is done by interest rather than decision leverage; the cost of the slow answers is ignored.
For exampleEight enterprise-customer interviews, a competitor pricing study, more market-size work, and one former executive. Ranked by what each could change per week of effort, the former executive goes first, the customers second, and the market-size work is dropped because no plausible result would move anything.
The mathematics in fullChoose q* = argmax over studies q of (EVSI(q) - cost(q)) / time(q); re-solve after each result, since the value of the remaining studies changes as beliefs move.
Taught inMECN-451; MECS 560-1
SourcesHoward (1988). Decision Analysis: Practice and Promise. Management Science 34(6), 679-695. https://doi.org/10.1287/mnsc.34.6.679
Kantorovich (1960). Mathematical Methods of Organizing and Planning Production. Management Science 6(4), 366-422. https://doi.org/10.1287/mnsc.6.4.366
We tell you when further research is unlikely to change the decision enough to justify its cost, including research from us.
Stop when the best remaining net VOI is at or below zeroAbraham Wald developed sequential analysis during the Second World War for inspecting munitions: instead of testing a fixed number of items, stop as soon as the accumulated evidence crosses a threshold either way. It cut inspection effort by a large fraction and was classified until 1945. Bellman's dynamic programming later gave the general form: stop when the value of continuing falls below the value of acting.
Why it works nowA research firm paid by the hour has no reason to compute a stopping rule and every reason not to. A firm on a fixed fee can state the rule in advance and mean it. What AI adds is the ability to track, across every interview and document, whether the recommendation is still moving, so that saturation is observed rather than declared. The stopping sentence in a Vista brief is Wald's rule applied to a client's money.
When it appliesThe client is on the third round of diligence and each round moves the picture less; or a firm paid by the hour keeps finding one more source.
What we would askWhat did the last round of work change about the decision?
Is there any remaining question whose plausible answers would flip it?
A record of what each round of research changed; the plausible range of the remaining unknowns.
How we would get itStated as a rule at scoping: the engagement ends when the recommendation is stable across the plausible results of every remaining study, and the report says so in one sentence.
Where it breaksSelf-serving in either direction: hourly firms never stop, fixed-fee firms stop early. The defense is stating the rule before the work starts.
For exampleAfter the customer interviews and the former executive, the thesis is fragile on one variable that no obtainable evidence will resolve before the deadline. The honest sentence is that more research will not settle it, that the decision is being made under that uncertainty, and that the client should stop paying, including paying us.
The mathematics in fullStop when U(stop | state) is at least -cost(q) + E[V(state after q)] for every remaining study q; equivalently when the best remaining net value of information is at or below zero.
Taught inMECS 560-2; MECN-451
SourcesWald (1945). Sequential Tests of Statistical Hypotheses. The Annals of Mathematical Statistics 16(2), 117-186. https://doi.org/10.1214/aoms/1177731118
Bellman (1954). The theory of dynamic programming. Bulletin of the American Mathematical Society 60(6), 503-515. https://doi.org/10.1090/S0002-9904-1954-09848-8
We work out whether the next quarter resolves a variable this decision turns on, in which case waiting is worth something, and whether the window will still be open, in which case it is not.
Invest only when V exceeds a threshold V* above costMcDonald and Siegel's 1986 paper asked when a firm should make an irreversible investment whose value fluctuates, and found that the usual rule, invest when value exceeds cost, is wrong. The right rule waits until value exceeds cost by a wide margin, in their calibration roughly double, because investing kills the option to invest later on better information. Dixit and Pindyck built the real-options field on it.
Why it works nowThe model needs to know how fast the uncertainty resolves and whether the window stays open, which used to be guesses. Both are now researchable: operators can say how quickly the crux variable actually clears in their industry, and the record shows how often comparable windows closed. The Desk asks the two questions that decide it, what would you know in six months and who can take this away from you meanwhile, and the engagement answers them.
When it appliesAn irreversible commitment under high uncertainty: a plant, an acquisition, a platform bet, a market entry; or the reverse, a client waiting for certainty that will never arrive while a competitor moves.
What we would askWhat would you know in six months that you do not know now?
Who else can take this decision away from you in the meantime?
How much of the uncertainty resolves with time versus never; the reversibility of the commitment; the competitive clock.
How we would get itOperators on how fast the crux variable actually resolves in this industry; the record on comparable windows that closed.
Where it breaksThe option is illusory because the seller walks or the competitor enters; or the uncertainty is permanent and waiting only delays.
For exampleA buyer can close a carve-out now or wait for the target's next two quarters of results. The wait is worth a great deal if those quarters reveal whether the churn is structural, and worth nothing if a strategic bidder is already in the data room.
The mathematics in fullInvest now only when value V exceeds a threshold V* strictly above the cost I; the gap V* - I grows with the volatility of V and with the irreversibility of I, and shrinks to zero when the option can be taken away.
Taught inDECS-433; MECN-451
SourcesMcDonald, Siegel (1986). The Value of Waiting to Invest. The Quarterly Journal of Economics 101(4), 707. https://doi.org/10.2307/1884175
Dixit, Avinash K. and Pindyck, Robert S. (1994). Investment under Uncertainty. Princeton University Press. (book)
We replace bull, base, and bear with the whole spread of outcomes, because a plan built on average inputs does not produce the average result.
E[f(X)] is not f(E[X]) whenever f bendsThe mathematics is Jensen's 1906 inequality on convex functions. The name is Sam Savage's, from his joke about the statistician who drowned crossing a river that was on average three feet deep. Monte Carlo itself was born at Los Alamos in the late 1940s, when Stanislaw Ulam, playing solitaire while convalescing, realized that many random trials could answer questions the equations could not, and Metropolis and Ulam published the method in 1949.
Why it works nowRunning the simulation was never the hard part; deciding what to put in it was. Distributions invented at a desk produce a histogram of assumptions. What is now practical is eliciting ranges and co-movements from the people who have watched the variables move, reading the record for their historical spread, and feeding a model the client already owns. Vista designs the inputs. The client or a specialist runs the model. The Desk explains why the base case is not the expected case.
When it appliesThe model has a base case with a bull and a bear beside it; the thesis depends on several uncertain inputs at once; or the inputs move together in bad states.
What we would askWhich three inputs, if they came in at their worst plausible values together, would break the plan?
Do those inputs tend to go wrong at the same time?
Ranges rather than points for the inputs that matter, and the correlations among them; which the interviews are designed to elicit.
How we would get itAdvisors asked for ranges and for what moves together, not for point forecasts; the record for the historical spread of the same variables.
Where it breaksThe input distributions are made up, so the histogram is a picture of assumptions; correlations are ignored and the tails are understated.
For exampleA leveraged buyout model shows a comfortable base case. Run across the ranges the operators actually gave for growth, margin, and exit multiple, with the three falling together in a downturn, the probability of a covenant breach is not a footnote. It is the crux, and it points at the one variable to research next.
The mathematics in fullY = f(X1..Xk) with each X drawn from its range and correlated where they move together; by Jensen, E[f(X)] is not f(E[X]) whenever f bends, and leverage makes f bend.
Taught inMECN-451; DECS-433
SourcesMetropolis, Ulam (1949). The Monte Carlo Method. Journal of the American Statistical Association 44(247), 335-341. https://doi.org/10.1080/01621459.1949.10483310
Jensen (1906). Sur les fonctions convexes et les inégalités entre les valeurs moyennes. Acta Mathematica 30(0), 175-193. https://doi.org/10.1007/BF02418571
Bertsimas, Dimitris and Freund, Robert M. (2004). Data, Models, and Decisions: The Fundamentals of Management Science. Dynamic Ideas. (book)
We treat the incidents that almost happened as evidence about the one that could, and fit the tail from them instead of from the calm years.
P(X > x) falls like a power of x, not exponentiallyBenoit Mandelbrot's 1963 study of cotton prices showed that the big moves came far more often than a bell curve allows, and that the tails, not the middle, held the risk. The management version is taught at Kellogg through the Boeing 737 MAX: a rare disaster that the near-miss record had been describing for some time to anyone who counted it.
Why it works nowNear misses were always the evidence and were never collected, because nobody read every incident report. Reading every incident report is now cheap. Former plant, safety, and quality leaders can be asked for the count that never reached the board, the record can be swept for comparable events at peers, and the tail can be fitted from what almost happened rather than from the calm years. Where the fit needs real extreme-value statistics, a specialist does it; recognizing that it is needed is now a conversation.
When it appliesA rare, severe event dominates the thesis: recall, outage, safety failure, regulatory action, fraud; and the client is reasoning from the years in which it did not happen.
What we would askHow many near misses has this business had, and who counted them?
What would the event cost, and is that a bad year or the end?
Incident and near-miss history, including the ones not reported; comparable events at peers; the size of the loss conditional on the event.
How we would get itFormer operators, safety and quality leads, and regulators on the near-miss record; the record on comparable failures and their consequences.
Where it breaksThe tail is fitted to a handful of observations; the absence of the event is read as evidence of its impossibility.
For exampleAn industrial company has had no serious incident in a decade and four near misses in two years. The decade is not the evidence. The four are, and the interviews with former plant managers are where the count comes from.
The mathematics in fullFit the frequency of severe events from the full incident distribution rather than the severe ones alone; the loss is E[L | event] x P(event), and heavy tails with exponent below two make the sample average an unreliable estimate of either term.
Taught inMECN-451; MECNX-435
SourcesMandelbrot (1963). The Variation of Certain Speculative Prices. The Journal of Business 36(4), 394. https://doi.org/10.1086/294632
Clauset, Shalizi, Newman (2009). Power-Law Distributions in Empirical Data. SIAM Review 51(4), 661-703. https://doi.org/10.1137/070710111
When the evidence cannot supply a probability, we say so, and we design the decision to be robust across the range of beliefs the evidence actually permits, rather than pretending to one number.
max over acts of min over priors of E[U]Daniel Ellsberg's 1961 experiment offered people two urns, one with a known fifty-fifty mix of colors and one with an unknown mix, and found that most preferred to bet on the known urn whichever color they were betting on. No single probability can explain that. Gilboa and Schmeidler in 1989 and Klibanoff, Marinacci, and Mukerji in 2005 built decision theories in which a careful person holds a set of beliefs rather than one, and is not being irrational when the set is wide.
Why it works nowFor decades the practical consequence was nil, because nobody could say how wide the set of defensible beliefs actually was. Now the record can be searched for whether a reference class exists at all, and Advisors can be asked for the range a careful person could hold, so the width of the set becomes a finding rather than a mood. The Desk says when a question has no defensible number, and the engagement reports which decisions survive the whole range instead of pretending to a point.
When it appliesThe client wants a probability for something with no comparable history: a novel regulation, a first-of-kind technology, a geopolitical break; or a report has supplied a precise number that nothing supports.
What we would askWhat class of past cases would this belong to, and does one exist?
Across the range of beliefs a careful person could hold, which decision holds up?
Honesty about the reference class: whether one exists, and how wide the range of defensible beliefs is.
How we would get itAuthorities on whether any base rate exists; operators on what the plausible range is; the report gives a range of priors and the decision that survives all of them.
Where it breaksAmbiguity used as an excuse not to estimate the estimable; or a single number reported where a range of priors was the truth.
For exampleA fund wants the probability that a proposed rule takes effect as drafted. There is no base rate for this rule. The honest deliverable is the range a careful reader of the record would hold, and which position sizes survive the whole range.
The mathematics in fullRather than a single prior, a set of priors; the robust choice maximizes the minimum expected utility over the set, or, with smooth ambiguity, an ambiguity-averse aggregate of the expected utilities across it.
Taught inMECS 550-1; MMSS 300
SourcesEllsberg (1961). Risk, Ambiguity, and the Savage Axioms. The Quarterly Journal of Economics 75(4), 643. https://doi.org/10.2307/1884324
Gilboa, Schmeidler (1989). Maxmin expected utility with non-unique prior. Journal of Mathematical Economics 18(2), 141-153. https://doi.org/10.1016/0304-4068(89)90018-9
Klibanoff, Marinacci, Mukerji (2005). A Smooth Model of Decision Making under Ambiguity. Econometrica 73(6), 1849-1892. https://doi.org/10.1111/j.1468-0262.2005.00640.x
Before we argue about this case, we look at how the whole class of cases like it turned out, and we run the failure post-mortem before the failure.
Forecast = the class distribution, then adjustKahneman and Lovallo's 1993 paper distinguished the inside view, which reasons from the specifics of the case, from the outside view, which starts from how similar cases turned out, and showed that the inside view produces bold forecasts and timid choices. Bent Flyvbjerg turned the outside view into a procedure, reference class forecasting, first for transport projects, where the class of comparable projects predicted cost overruns the planners never did, and it was adopted into official appraisal guidance.
Why it works nowThe outside view was rarely taken because assembling the reference class was a research project in itself. It is now the cheapest part of an engagement: the record can be read for how the class of comparable deals, migrations, launches, or turnarounds actually resolved, and the Advisors can be asked the pre-mortem question directly. The Desk starts from the class and asks what makes this case different, which is the reverse of how most investment memos are written.
When it appliesA forecast built from the specifics of the deal, the plan, or the team, with no reference to how similar deals, plans, and teams have done; or a projection everyone in the room finds persuasive.
What we would askWhat is the class of comparable cases, and how did they actually turn out?
If this has failed two years from now, what is the most likely reason?
A defensible reference class and its outcome distribution; the plan's own assumptions listed so they can be checked against it.
How we would get itThe record for the reference class; operators who have seen the failures; a structured pre-mortem with the Advisors as the second half of each interview.
Where it breaksThe reference class is chosen to flatter; the bias story is applied after the fact and cannot be falsified.
For exampleA platform migration plan promises eighteen months. The class of comparable migrations at companies of this size ran twice the planned time in most cases. The research starts from that distribution and asks what makes this one different, rather than starting from eighteen months and asking what could go wrong.
The mathematics in fullForecast = distribution of outcomes in the reference class, then adjust for the case at hand; the adjustment is the inside view and should be smaller than the decision maker wants it to be.
Taught inMECNX-435; MECN-451
SourcesKahneman, Lovallo (1993). Timid Choices and Bold Forecasts: A Cognitive Perspective on Risk Taking. Management Science 39(1), 17-31. https://doi.org/10.1287/mnsc.39.1.17
Flyvbjerg (2006). From Nobel Prize to Project Management: Getting Risks Right. Project Management Journal 37(3), 5-15. https://doi.org/10.1177/875697280603700302
Kahneman, Tversky (1973). On the psychology of prediction. Psychological Review 80(4), 237-251. https://doi.org/10.1037/h0034747
Some questions cannot be answered defensibly from any evidence obtainable in the time available. When that is true we say so before you spend, and it is often the most useful sentence in the conversation.
No structure applies, and naming that is the disciplineFrank Knight drew the line in 1921 between risk, which can be measured, and uncertainty, which cannot, and Ellsberg's urns made the line experimentally real. Gilboa's modern treatment keeps the honest case: some questions have no probability that a careful person could defend, and the rational response is to decide in a way that does not require one.
Why it works nowThe temptation to manufacture a number has grown with the tools that can produce one. The discipline that has become cheaper is the opposite: checking, quickly and thoroughly, whether anyone anywhere is positioned to know the thing, and if nobody is, saying so before the client spends. The Desk does this in the conversation. It is the entry in the Library that exists to refuse.
When it appliesThe question turns on something nobody can observe in time (a private decision not yet made, a hidden preference, a future negotiation), or on a quantity with no comparable history and no one positioned to know it.
What we would askWho, anywhere, is actually positioned to know this today?
If nobody is, what decision would you make under that uncertainty, and is that the real question?
Candor about what is observable; a reframing of the decision so it can be made under the uncertainty rather than after resolving it.
How we would get itNone sold. The Research Director reframes the decision to the question that can be answered, if there is one, and otherwise declines.
Where it breaksDeclining what could have been partially answered; substance first, handoff last is the rule for everything short of this case.
For exampleA client wants to know whether a specific acquirer will bid for a specific target next quarter. Nobody outside that acquirer's board knows, and they will not say. The answerable question is what the target is worth to that acquirer and what a bid would need to look like, which is a different engagement.
The mathematics in fullNo structure applies; this is the boundary of the library, and naming it is the discipline.
Taught inMECS 550-1
SourcesEllsberg (1961). Risk, Ambiguity, and the Savage Axioms. The Quarterly Journal of Economics 75(4), 643. https://doi.org/10.2307/1884324
Gilboa, Itzhak (2009). Theory of Decision under Uncertainty. Cambridge University Press. (book)
Most decisions arrive as a choice between two things. Before analyzing either, we generate the alternatives that were never listed, and we separate what has already been decided from what is being decided now and what can be deferred.
A strategy is one coherent path across the dimensionsHoward's decision analysis at Stanford insists that the quality of a decision is decided before any analysis, by the frame and the alternatives, and it supplies two tools: the decision hierarchy, which separates what is taken as given from what is being decided and what is deferred, and the strategy table, which lays out every dimension of a strategy with its options so that whole strategies can be composed rather than two options compared. The Stanford text calls the failure to recognize alternatives one of the most expensive errors in practice.
Why it works nowAlternatives were generated in a workshop by whoever was in the room. The table can now be built in a conversation, with the dimensions and options drawn from comparable decisions in the record and from Advisors who have seen the versions that were never listed. The Desk asks what has already been decided that this decision takes as given, and what the client would do if both listed options were unavailable.
When it appliesA yes-or-no decision; a choice between two options that both feel wrong; a plan whose alternatives were set by whoever wrote the memo.
What we would askWhat has already been decided that this decision takes as given, and should it?
If both of these options were unavailable, what would you do?
The decision hierarchy stated explicitly, and a structured search for alternatives across the dimensions of the strategy.
How we would get itThe Research Director builds the strategy table with the client and the Advisors before any evidence is gathered: each dimension of the decision, its options, and the coherent combinations.
Where it breaksAnalyzing the two listed options with great care; the answer was a third one.
For exampleA board is deciding whether to sell a division or keep it. Laid out as a table, there are also a partial sale, a joint venture, a carve-out with a service agreement, and a run-off. Two of those are worth more than either original option, and the research is redirected to them.
The mathematics in fullA decision hierarchy separates the givens, the decision, and the deferred; a strategy table lists each dimension of the strategy with its options, and a strategy is one coherent path across the columns.
Taught inMS&E 252 and 352 (Stanford); MECN-451
SourcesHoward, Ronald A. and Abbas, Ali E. (2016). Foundations of Decision Analysis. Pearson. (book)
Howard (1988). Decision Analysis: Practice and Promise. Management Science 34(6), 679-695. https://doi.org/10.1287/mnsc.34.6.679
How much each source and each fact should move a belief, and how to know when it did.
We break the thesis into the five to seven assumptions it rests on, say how each stands on the evidence, and record what moved each one and why.
Odds(H | E) = Odds(H) x likelihood ratioThomas Bayes's essay, published in 1763 after his death, imagined a ball rolled onto a table and its position inferred from where later balls landed to its left or right: with each new ball, the belief about the first one's position updates. Laplace made it a working method. The odds form, in which each piece of evidence multiplies the odds by how much likelier it is under one hypothesis than the other, is what a working analyst actually uses.
Why it works nowThe ledger was never kept because writing down every assumption and what moved it was clerical work nobody did. AI makes the clerical part free: every interview and document can be tagged to the assumption it bears on and the direction it moved it. What must not change is that the movement is stated in words, more likely than not, the evidence conflicts, unresolved, and not as a number that nothing calibrates. Vista keeps the ledger in the brief and preserves it afterward, which is how a track record starts.
When it appliesA thesis stated as a narrative rather than as claims that could be false; or a client who cannot say which assumption the whole thing rests on.
What we would askWhat has to be true for this to work? Say it as a sentence that could be false.
Which of those, if it fell, takes the rest with it?
Each assumption stated falsifiably, with the evidence for and against it kept separate and traceable to its source.
How we would get itThe report opens with the ledger: each assumption, a plain-language confidence band (more likely than not; the evidence conflicts; unresolved), what moved it, and where the transcript supports it. No invented percentages.
Where it breaksNumbers dressed as calibrated probabilities that are in fact intuition; a ledger written after the conclusion to justify it.
For exampleA thesis on a software company reduces to seven sentences. Two are well supported, three conflict, and one is unresolved and decisive: that customers past two hundred seats face real switching costs. The engagement is built around that sentence.
The mathematics in fullOdds(H | E) = Odds(H) x LR; stated in prose as which direction each piece of evidence moved each assumption and how strongly, with the likelihood ratio the reason and never a displayed number.
Taught inDECS-430; MATH 385
SourcesBayes (1763). An Essay towards Solving a Problem in the Doctrine of Chances. Philosophical Transactions of the Royal Society of London(53), 370-418. https://doi.org/10.1098/rstl.1763.0053
We weight each source by what they were positioned to know, not by their title. A former salesperson is a strong source on pricing and a weak one on technology, and the report says which is which.
Weight source i by its reliability r(i, d) in domain dRomney, Weller, and Batchelder's 1986 cultural consensus model came from anthropology, where informants are asked about their own culture and there is no answer key. The model estimates each informant's competence from how much they agree with the others, and recovers the consensus answers without knowing them in advance. Competence, it turned out, is specific to the domain of the questions.
Why it works nowExpert networks have always sold access and left reliability to the client's intuition. What is now feasible is to record, for every Advisor, what they could observe from where and when, to tag each claim in a transcript by the domain it belongs to, and to weight the synthesis accordingly, so that a former salesperson counts heavily on pricing and lightly on technology. Over many engagements that becomes a dataset about which kinds of sources are reliable about which kinds of claims, which no one has ever built.
When it appliesEvidence from several sources of unequal access, recency, and incentive is being averaged as if it were equal; or a senior title is being treated as expertise on everything.
What we would askWhat did this person actually see, from where, and how long ago?
What do they gain if you believe them?
For each source: role, access, recency, incentives, conflicts, and independence from the other sources.
How we would get itAdvisors are chosen by vantage point, and each interview records what the Advisor could and could not observe; the synthesis weights claims by that record, not by seniority.
Where it breaksA single universal credibility score per source; reliability treated as a trait rather than as domain-specific.
For exampleThree sources agree the product is winning. One is a former sales lead who saw the pipeline, one a customer who saw the product, one an analyst who saw the pitch. On pricing, the first counts most; on product fit, the second; the third counts on neither.
The mathematics in fullFor source i on domain d, a reliability r(i, d) enters the likelihood ratio of that source's report; the aggregate is a weighted sum of log likelihood ratios, with weights for reliability, relevance, recency, and independence.
Taught inMMSS 1995 (informant accuracy); DECS-430
SourcesRomney, Weller, Batchelder (1986). Culture as Consensus: A Theory of Culture and Informant Accuracy. American Anthropologist 88(2), 313-338. https://doi.org/10.1525/aa.1986.88.2.02a00020
Before we count confirming sources we check whether they saw the thing themselves or heard it from each other. Five customers who share an implementation partner are closer to one observation than to five.
Log odds add only across independent evidenceAbhijit Banerjee's 1992 model has diners choosing between two restaurants: each sees a private signal and the choices of those ahead, and after a few people the crowd follows the crowd regardless of what anyone privately knows. Bikhchandani, Hirshleifer, and Welch published the general theory of informational cascades the same year and used it to explain fads and fashions that flip on almost nothing.
Why it works nowA research firm used to count confirming sources. It could not trace where each source got the story, so five agreeing calls looked like strong evidence even when four had heard it from the fifth. Provenance is now traceable: transcripts can be read for who saw what directly and who is repeating, and Advisors can be recruited from deliberately separate channels. The brief counts independent clusters, not voices.
When it appliesConsensus among sources is being treated as strong evidence; the sources share a channel, a consultant, a conference, or an outage; or a market view has converged quickly.
What we would askHow did each of these people come to know this? Did any of them see it directly?
Do they talk to each other, or to the same third party?
The provenance chain of each claim; a map of which sources share upstream sources.
How we would get itAdvisors recruited deliberately from different channels and vantage points; the report clusters evidence by independence and counts clusters, not voices.
Where it breaksA cascade is read as convergent evidence; independence is asserted because the sources have different job titles.
For exampleSix channel partners report the same competitor weakness. Five learned it from the same distributor's sales kickoff. The evidence is one distributor's claim plus one independent observation, and the report says so.
The mathematics in fullLog odds add across pieces of evidence only when they are independent; with pairwise correlation rho among n confirming sources, the effective number of observations approaches 1/rho rather than n.
Taught inMECN-451; DECS-430
SourcesBanerjee (1992). A Simple Model of Herd Behavior. The Quarterly Journal of Economics 107(3), 797-817. https://doi.org/10.2307/2118364
Bikhchandani, Hirshleifer, Welch (1992). A Theory of Fads, Fashion, Custom, and Cultural Change as Informational Cascades. Journal of Political Economy 100(5), 992-1026. https://doi.org/10.1086/261849
We separate claims, which cost nothing, from signals, which would have been expensive for a liar, from commitments, which are hard to reverse. The plan is weighed on the second and third.
A signal separates when c(s, high) < c(s, low)Michael Spence's 1973 model explained why education pays even if it teaches nothing: a degree is costly to obtain, more costly for the less able, so it separates the types and employers rationally reward it. Crawford and Sobel's 1982 paper showed the mirror case: when talk is free and interests differ, what gets communicated is coarse at best, and the wider the difference in interests, the less can be said.
Why it works nowReading a management team for costly actions rather than statements has always been what good analysts did by instinct on a handful of names. It can now be done systematically across filings, hiring data, pricing behavior, insider transactions, and contract terms, and then tested in interviews with operators who know what a real commitment looks like in that industry. The Desk sorts the deck into claims, signals, and commitments before the first Advisor is recruited.
When it appliesManagement says the pipeline is strong, the seller says the customer is loyal, the founder says the round is oversubscribed; or the client is reading intent into statements that cost the speaker nothing.
What we would askWhat have they done that would have been foolish if the claim were false?
What have they done that they cannot easily undo?
The actions taken, their cost to the actor, and their reversibility; a record of what was said against what was done.
How we would get itOperators who can say what a real commitment looks like in this industry; the record (hiring, capital spend, contracts, insider purchases, pricing discipline) checked against the claims.
Where it breaksThe supposedly costly action was cheap for the sender; the receiver reads intent into noise.
For exampleA CEO describes an extremely strong enterprise pipeline. The company has also hired forty salespeople, refused to discount in the quarter, and the CEO has bought stock. The sentence is talk; the three actions are the evidence, and the interviews test whether they were costly.
The mathematics in fullA signal separates types only when its cost differs across them, c(s, honest) < c(s, bluffing); a costless message is informative only when the speaker's interests are aligned with the listener's, which in a sale they are not.
Taught inMMSS 311-1; MECN-452; MECS 465
SourcesSpence (1973). Job Market Signaling. The Quarterly Journal of Economics 87(3), 355. https://doi.org/10.2307/1882010
Crawford, Sobel (1982). Strategic Information Transmission. Econometrica 50(6), 1431. https://doi.org/10.2307/1913390
We record what we said, in the words and bands we said it, and we score it when the outcome arrives, so that our judgment has a track record rather than a reputation.
Brier = mean of (p - outcome) squaredGlenn Brier's 1950 paper gave weather forecasters a score for probability forecasts: the squared distance between the stated chance of rain and whether it rained, averaged over many days. Meteorology adopted it and became, as a result, one of the few professions whose seventy percent means seventy percent.
Why it works nowInvestment research has never been scored, because the forecasts were buried in prose and revised in memory. A brief whose ledger is written in explicit bands and preserved unaltered can be scored when the outcomes arrive, and the outcomes can now be captured from the record as they happen. This takes years to mature and Vista says so. It is the only route to a research firm whose judgment has a track record rather than a reputation.
When it appliesA research provider, an analyst, or an internal team makes confident calls with no record of how past calls turned out; or the client wants to know whether Vista's own bands mean anything.
What we would askOf the last twenty calls like this, how many resolved the way they were called?
Were the confident ones more often right than the hedged ones?
Predictions preserved unaltered at the time they were made, with resolution dates and outcomes.
How we would get itEvery Vista brief keeps its ledger as written; outcomes are captured as they resolve; the score is reported to clients when the history is long enough to mean something, and not before.
Where it breaksForecasts revised after the fact; bands so wide they cannot be wrong; a score reported on a handful of cases.
For exampleA provider says a merger is very likely to close. It closes. That is one observation. Twenty such calls with outcomes, scored, are a track record; the difference is why Vista keeps its ledgers.
The mathematics in fullBrier = (1/N) x sum of (p(i) - y(i))^2; calibration means that among forecasts near p, the realized frequency is near p. The library of predictions must be preserved to be scorable.
Taught inDECS-430; MECN-451
SourcesBrier (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review 78(1), 1-3. https://doi.org/10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2
Whatever was measured at an extreme was partly luck, and the luck does not repeat. We separate the part of last year's result that was skill from the part that was noise before anyone is rewarded, punished, or bought.
Next = mean + r x (this value - mean)Galton discovered it in 1886 measuring parents and children, and it has been rediscovered by every generation since. Tversky and Kahneman's 1971 paper on belief in the law of small numbers showed that trained scientists ignore it too, and Kahneman later told the story of flight instructors who believed praise made pilots worse and criticism made them better, because a great landing was followed by a worse one and a bad landing by a better one, whatever the instructor said.
Why it works nowRewarding last year's best and punishing last year's worst has been organizational reflex because separating skill from noise required several periods of data for the unit and its peers, and nobody assembled them. The record can now be assembled quickly and the persistence of the measure estimated. The Desk asks how much the measure swings for reasons nobody controls, and how much of the recovery would have happened anyway.
When it appliesA manager, fund, region, product, or salesperson selected because of an extreme result; a turnaround credited to an intervention that followed a bad year; a rank list used to allocate.
What we would askHow much does this measure vary from year to year for reasons nobody controls?
If you had done nothing after the bad year, how much of the recovery would have happened anyway?
Several periods of the measure for the unit and its peers, enough to estimate how much of the variation is persistent.
How we would get itThe record for the measure across periods and peers; operators on what drives the year-to-year swings.
Where it breaksReading every reversal as an effect; punishing after bad luck and crediting the punishment.
For exampleA fund allocates to the three regional managers with the best results last year. Their results were the most extreme of thirty, and the year-to-year correlation of the measure is low. Most of their advantage was noise, and next year's leaders will be different names.
The mathematics in fullExpected next value = mean + r x (this value - mean), where r is the correlation between periods; the further from the mean the observation, the larger the expected fall back, and with r near zero the extreme is almost entirely noise.
Taught inMATH 385; MATH 386-1
SourcesGalton (1886). Regression Towards Mediocrity in Hereditary Stature. The Journal of the Anthropological Institute of Great Britain and Ireland 15, 246. https://doi.org/10.2307/2841583
Tversky, Kahneman (1971). Belief in the law of small numbers. Psychological Bulletin 76(2), 105-110. https://doi.org/10.1037/h0031322
An expert asked for a forecast gives an anchored, overconfident one. We ask in a fixed order that starts from the extremes and checks for consistency, so that the range an Advisor gives us is one we can defend.
Elicit the 10th and 90th fractiles before the 50thCarl Spetzler and Carl-Axel Stael von Holstein's 1975 paper, from the Stanford decision analysis group, set out the protocol for encoding an expert's probability: fix the quantity precisely, elicit the extremes before the center, alternate the form of the question, check for consistency, and calibrate. Tversky and Kahneman's 1974 catalogue of anchoring and overconfidence is the reason the protocol exists. Ronald Howard's Stanford course still teaches it.
Why it works nowVista's product is asking experts questions, and asking for a quantity badly produces a confident wrong answer. The protocol was known for fifty years and used by a handful of decision analysts, because applying it took training and time. It can now be built into every structured interview, with the fractiles asked in the right order and the consistency checks run as the conversation happens. The Desk asks whether the client's own experts start from a central number or from the extremes.
When it appliesAny engagement in which Advisors will be asked for ranges, probabilities, or timing; a client whose own experts have produced narrow ranges that keep being wrong.
What we would askWhen your team estimates a range, do they start from a central number or from the extremes?
How often do outcomes fall outside the ranges your experts give?
A structured protocol applied consistently across Advisors; calibration checks on quantities with known answers.
How we would get itEvery Vista interview that asks for a quantity uses the protocol: the surprising values first, the central estimate last, consistency checks between fixed-value and fixed-probability questions.
Where it breaksAnchoring on the first number mentioned; an expert with a stake in the answer; ranges reported without the protocol that produced them.
For exampleThree operators are asked when a competitor's plant will reach full capacity. Asked directly, all three say eighteen months. Asked first for the date they would be surprised to see it beat, then the date they would be surprised to see it miss, the honest range runs from twelve to thirty-six months, and the thesis has to survive the whole range.
The mathematics in fullElicit the tenth, ninetieth, and then fiftieth fractiles before any central estimate; alternate fixed-value and fixed-probability questions; calibration means that about ten percent of outcomes should fall below the stated tenth fractile.
Taught inMS&E 252 (Stanford); MECN-451
SourcesSpetzler, Stael von Holstein (1975). Probability Encoding in Decision Analysis. Management Science 22(3), 340-358. https://doi.org/10.1287/mnsc.22.3.340
Tversky, Kahneman (1974). Judgment under Uncertainty: Heuristics and Biases. Science 185(4157), 1124-1131. https://doi.org/10.1126/science.185.4157.1124
Howard, Ronald A. and Abbas, Ali E. (2016). Foundations of Decision Analysis. Pearson. (book)
What the other side does when you move, and what that does to the plan.
We ask what a competent competitor does after you move, and whether your plan still works once they have.
No player gains by changing strategy aloneAntoine Augustin Cournot analyzed two sellers of mineral water in 1838 and found the outcome where each produces its best quantity given the other's, which John Nash generalized in 1951 to any number of players and any game. Bulow, Geanakoplos, and Klemperer's 1985 paper added the practical taxonomy: when moves are strategic substitutes a rival meets aggression by backing off, when they are complements by matching it, and the sign decides whether a plan works.
Why it works nowEvery investment model that assumed passive competitors did so because modeling the response required knowing the rival's payoffs, and nobody outside the rival knew them. Former executives of the rival do, and they can be recruited. The Desk asks the two questions, what is the best thing your strongest competitor can do after you move and does the plan survive it, and the engagement puts the people who ran that competitor on the record.
When it appliesA plan (price rise, entry, new product, acquisition, channel change) whose model assumes rivals stay where they are; or a client puzzled by a rival's move.
What we would askIf you do this, what is the best thing your strongest competitor can do in response?
Does your plan still work if they do it?
The rivals' real payoffs and constraints, which are usually not the ones the client assumes; who moves first and who can observe whom.
How we would get itFormer executives of the competitors, who know how those companies actually decide; channel partners and customers who saw past responses.
Where it breaksThe game is mis-specified: wrong players, wrong moves, wrong information; or several stable outcomes exist and one was picked for convenience.
For exampleA company plans a price increase on the assumption that its main rival follows. The rival's former head of pricing, interviewed, explains that the rival's incentive plan rewards share, not margin, and it will not follow. The plan is a different plan.
The mathematics in fullA set of strategies is stable when no player gains by changing alone, given the others; with strategic substitutes (quantities, capacity) a rival responds to aggression by retreating, with strategic complements (prices) by matching, and the sign decides the plan.
Taught inMMSS 211-2; MECN-452; MECN-441
SourcesNash (1951). Non-Cooperative Games. The Annals of Mathematics 54(2), 286. https://doi.org/10.2307/1969529
Bulow, Geanakoplos, Klemperer (1985). Multimarket Oligopoly: Strategic Substitutes and Complements. Journal of Political Economy 93(3), 488-511. https://doi.org/10.1086/261312
Gibbons, Robert (1992). Game Theory for Applied Economists. Princeton University Press. (book)
We work out whether the incumbent's threats are credible, which usually means whether it has already spent money it cannot get back, and whether a reputation for fighting is real or a story.
A threat is credible only if carrying it out is optimalReinhard Selten's chain-store paradox asked why an incumbent would fight entrants when fighting costs more than accommodating, and backward induction said it never would. Kreps and Wilson, and Milgrom and Roberts, resolved it in 1982: a small chance that the incumbent is genuinely tough makes fighting early entrants worthwhile as reputation. Avinash Dixit in 1980 showed the other route, capacity built before entry that cannot be unbuilt, which turns a threat into a fact.
Why it works nowWhether a moat is real used to be argued from the incumbent's rhetoric. What can now be established is what the incumbent has actually sunk, what it did the last three times someone entered, and what each response cost it, from the record and from former executives of both the incumbent and the entrants. The Desk asks what has been committed that cannot be reversed. The interviews answer it.
When it appliesA client is deciding whether to enter a market with an aggressive incumbent, or is an incumbent deciding how to respond to entry, or is underwriting a company whose moat is an entry barrier.
What we would askWhat has the incumbent already committed that it cannot reverse?
When someone last entered, what did the incumbent actually do, and did it pay for it?
The incumbent's sunk commitments, its history of responses to entry, and what those responses cost it.
How we would get itFormer executives of the incumbent and of past entrants; the record of past entries and the price and capacity responses that followed.
Where it breaksThe threat is not credible because carrying it out would hurt the incumbent more than accommodating; or the moat was regulatory and the regulation is changing.
For exampleA venture investor is underwriting an entrant against a dominant incumbent said to crush competition. The last three entrants were bought, ignored, and out-priced respectively; the interviews establish which of those this entrant resembles and what the incumbent's capacity position makes credible now.
The mathematics in fullA threat is credible only when carrying it out is the incumbent's best response at the point of entry; irreversible investment before entry changes that best response, and a reputation for toughness is worth more the more entrants remain to be deterred.
Taught inMMSS 311-1; MECN-441; MECS 449
SourcesDixit (1980). The Role of Investment in Entry-Deterrence. The Economic Journal 90(357), 95. https://doi.org/10.2307/2231658
Milgrom, Roberts (1982). Predation, reputation, and entry deterrence. Journal of Economic Theory 27(2), 280-312. https://doi.org/10.1016/0022-0531(82)90031-X
Besanko, Dranove, Shanley and Schaefer. Economics of Strategy, 7th ed. Wiley. (book)
We look at what keeps rivals from cutting price: how much future business they expect from each other, how well they can see each other's moves, and whether anyone is about to leave the game.
Cooperation holds while gain G is below the discounted future lossGreen and Porter's 1984 model explained price wars without anyone cheating: when rivals cannot see each other's prices and only observe a noisy market, a bad quarter looks like a secret price cut, and the punishment phase that follows is what keeps the truce credible the rest of the time. The theory was built with a nineteenth-century railroad cartel in mind, whose periodic price wars Porter had studied.
Why it works nowUnderwriting pricing discipline in a concentrated industry has always meant guessing at three things: how transparent prices are among rivals, how long each expects to be in the game, and whether anyone is about to leave. All three are now questions former pricing and sales leaders across the rivals can answer, and the record shows what preceded every past price break. The Desk asks who is about to sell, retire, or exit, because that is what shortens the shadow of the future.
When it appliesAn industry with stable prices that the thesis assumes will stay stable; or a price war the client did not expect; or a rival that is about to be sold, exit, or change hands.
What we would askHow quickly would the others see a secret price cut, and what would they do?
Is anyone in this market about to leave it, sell, or retire?
Transparency of pricing among rivals, the horizon each expects, and the events that shorten it.
How we would get itFormer pricing and sales leaders across the rivals; the record of past price breaks and what preceded them.
Where it breaksThe end of the game is nearer than assumed; punishment is not credible; deviations cannot be observed clearly enough to punish.
For exampleA thesis rests on rational pricing in a four-player market. One of the four is being prepared for sale, which shortens its horizon and changes its incentives in the last year. The interviews with former executives of that player are where the thesis is tested.
The mathematics in fullCooperation holds when the one-period gain from cutting, G, is at most the discounted future loss delta x (V(cooperate) - V(punish)) / (1 - delta); anything that lowers delta or blurs observation of the cut breaks the truce.
Taught inMMSS 311-1; MECN-441; MECN-452
SourcesGreen, Porter (1984). Noncooperative Collusion under Imperfect Price Information. Econometrica 52(1), 87. https://doi.org/10.2307/1911462
Fudenberg, Maskin (1986). The Folk Theorem in Repeated Games with Discounting or with Incomplete Information. Econometrica 54(3), 533. https://doi.org/10.2307/1911307
Kreps, Wilson (1982). Reputation and imperfect information. Journal of Economic Theory 27(2), 253-279. https://doi.org/10.1016/0022-0531(82)90030-8
We establish what each side gets if talks fail, because that, not the argument in the room, decides most negotiations, and we research the other side's outside option as hard as yours.
Maximize (u1 - d1) x (u2 - d2)John Nash's 1950 paper solved the bargaining problem from four axioms and found that the split depends on what each side gets if talks fail. Ariel Rubinstein's 1982 model of alternating offers showed the same thing dynamically: the more patient party captures more of the surplus, and impatience is expensive. Kellogg teaches the Nash scheme in its game theory course as a tool for actual negotiations.
Why it works nowNegotiators have always known their own fallback and guessed the other side's. The other side's fallback is now researchable: former executives and advisors from that side's world can say what a failed deal actually costs a party in that position, and the record shows the alternatives and the timing pressure. The Desk asks who is more able to wait. The engagement finds out before the first meeting instead of after the last.
When it appliesA client preparing a negotiation (acquisition, partnership, supply contract, renewal, joint venture) who knows their own fallback and is guessing the other side's.
What we would askWhat happens to them if this does not close? Not what they say, what actually happens.
Who is more able to wait?
Both sides' fallbacks, patience, and ability to commit, established from evidence rather than posture.
How we would get itFormer executives and advisors from the other side's industry on what a failed deal costs a party in that position; the record on the other side's alternatives and timing pressure.
Where it breaksThe outside options are misread (the seller does have another buyer, or does not); a commitment was not credible.
For exampleA buyer believes a founder must sell. The founder's former CFO, interviewed, explains that the company just refinanced and the founder can run it for years. The buyer's leverage was imaginary, and the negotiation plan changes before the first meeting.
The mathematics in fullThe split maximizes (u1 - d1) x (u2 - d2), where d1 and d2 are the payoffs if no deal is reached; in alternating offers the more patient side captures more, so improving your own d or worsening theirs moves the outcome more than any concession.
Taught inMECN-452; MECS 540-2; MMSS 311-1
SourcesNash (1950). The Bargaining Problem. Econometrica 18(2), 155. https://doi.org/10.2307/1907266
Rubinstein (1982). Perfect Equilibrium in a Bargaining Model. Econometrica 50(1), 97. https://doi.org/10.2307/1912531
Muthoo, Abhinay (1999). Bargaining Theory with Applications. Cambridge University Press. (book)
When both sides would gain from a deal and it still is not closing, the cause is usually private information about value or a reputation for holding out, and we design the research to find which.
No mechanism guarantees trade under two-sided private valuesMyerson and Satterthwaite proved in 1983 that when a buyer and a seller each privately know their own value, no bargaining procedure can guarantee trade whenever trade is worth doing; some good deals must fail. James Fearon's 1995 paper applied the same logic to war: rational states fight because of private information they have reason to misrepresent and because they cannot commit. Abreu and Gul showed in 2000 how a reputation for stubbornness lets a party win by waiting.
Why it works nowA stalled deal used to be read as irrationality or bad faith. It can now be diagnosed: what each side privately knows, what each has committed to in public, and what waiting costs each of them are all things former participants in comparable stalemates can speak to, and the record shows the public positions. The Desk asks which side's cost of delay is real. That tells the client who folds first and when.
When it appliesA transaction that makes obvious economic sense is stalled; a negotiation has become a waiting game; or a client wants to know whether a stalled deal will close.
What we would askWhat does each side know about the value that the other does not?
Is someone holding out because folding once would cost them in every future deal?
What each side privately knows, what each has committed to publicly, and each side's cost of continued delay.
How we would get itAdvisors who have been on the other side of comparable stalemates; the record of the parties' public positions and the costs of waiting.
Where it breaksThe efficiency of the deal is assumed to guarantee it closes; the hold-out is read as irrational when it is reputational.
For exampleA strategic buyer and a target agree the combination creates value and have not closed in nine months. Each believes the other is overstating its walk-away. The interviews establish which side's cost of delay is real, which tells the client who folds first and when.
The mathematics in fullWith private values on both sides, no mechanism guarantees efficient trade whenever a deal is worth doing; delay is the cost that separates types, and a party with a reputation for holding out can win by waiting even when waiting is expensive.
Taught inMECS 465; MECS 540-2
SourcesMyerson, Satterthwaite (1983). Efficient mechanisms for bilateral trading. Journal of Economic Theory 29(2), 265-281. https://doi.org/10.1016/0022-0531(83)90048-0
Abreu, Gul (2000). Bargaining and Reputation. Econometrica 68(1), 85-117. https://doi.org/10.1111/1468-0262.00094
Fearon (1995). Rationalist explanations for war. International Organization 49(3), 379-414. https://doi.org/10.1017/S0020818300033324
We look at whether the outcome depends on what each participant believes the others will do, in which case small changes in confidence can flip a whole market, and we find the participants whose beliefs matter most.
Each participant acts when its estimate crosses a thresholdCarlsson and van Damme's 1993 paper on global games showed that when players have slightly different information about the fundamentals, a coordination game with many possible outcomes collapses to one: each player acts when its own estimate crosses a threshold. Morris and Shin applied it to currency attacks, and Baliga and Sjöström's 2004 work on arms races formalized Schelling's reciprocal fear of surprise attack, where each side arms because it fears the other might.
Why it works nowRuns and panics were treated as sentiment because the decision rules of the large participants were unknowable. They are knowable: treasurers, portfolio managers, and operators at the kind of institutions that hold the deposits, the bonds, or the platform positions can be asked what they would do and what public news would change what they believe about each other. The Desk finds the participants whose beliefs matter most. The engagement asks them.
When it appliesA run, a panic, a rush for the exit, a standards war, a platform tipping point; or a thesis that assumes a stable equilibrium in a market with two.
What we would askIf the largest three participants believed the others would stay, would they stay?
What single piece of public news would change what everyone believes everyone else believes?
The participants' payoffs from staying versus leaving as a function of how many others leave; what information is common to all of them.
How we would get itOperators on the actual decision rules of the large participants; the record on comparable coordination breaks and what triggered them.
Where it breaksThe model assumes participants know each other's payoffs precisely when they only estimate them; multiple equilibria are collapsed to one.
For exampleA thesis on a regional lender assumes deposits are sticky. The question is not the average depositor's loyalty but whether the largest depositors believe the others will stay; the interviews go to treasurers of the kind of companies that hold those deposits.
The mathematics in fullWith a small amount of private information about fundamentals, the many equilibria of a coordination game collapse to one threshold: each participant acts when its estimate of the fundamental crosses a cutoff, so the size of the crowd that moves is a function of public news, not of mood.
Taught inMECS 540-2; MMSS 311-1
SourcesCarlsson, van Damme (1993). Global Games and Equilibrium Selection. Econometrica 61(5), 989. https://doi.org/10.2307/2951491
Baliga, Sjostrom (2004). Arms Races and Negotiations. Review of Economic Studies 71(2), 351-369. https://doi.org/10.1111/0034-6527.00287
Diamond, Dybvig (1983). Bank Runs, Deposit Insurance, and Liquidity. Journal of Political Economy 91(3), 401-419. https://doi.org/10.1086/261155
We test whether a commitment (a price floor, a no-sale promise, a walk-away, a regulatory pledge) would actually be kept when the moment comes, which depends on what breaking it would cost, not on how firmly it was said.
Credible when reversal costs more than honoringThomas Schelling's 1960 book turned commitment into strategy: the general who burns the bridges behind his army, the negotiator who ties his own hands, the focal point two strangers pick when told to meet in New York without communicating. James Fearon's 1994 paper added audience costs: a leader who makes a public threat pays a domestic price for backing down, which is exactly what makes the threat believable.
Why it works nowWhether a government, a regulator, or a counterparty will honor a pledge used to be judged from the firmness of its language. The cost of reversal is now the thing to research, and it can be: who the pledge was made in front of, what reversing it has cost that party before, and what it would cost now, from former officials and from the record. The Desk asks what backing down would cost them, and treats the answer, not the rhetoric, as the probability.
When it appliesA plan relies on a promise or a threat by the client, a counterparty, a regulator, or a government; or a client is trying to make its own commitment believable.
What we would askWhat would it cost them, publicly and privately, to back down?
Have they backed down from something like this before?
The audience in front of which the commitment was made and the cost of reversal before it; the history of reversals.
How we would get itAdvisors who have watched the party in question honor or abandon commitments; the record of prior reversals and their consequences.
Where it breaksFirmness of language is mistaken for cost of reversal; the audience that would punish a reversal is assumed rather than identified.
For exampleA government has pledged a subsidy regime through the decade. The pledge was made to voters and investors before an election. The interviews establish what reversing it would cost after the election, which is the actual probability the thesis depends on.
The mathematics in fullA commitment is credible when reversing it is more costly than honoring it at the moment of decision; audience costs, sunk investment, and reputation are the three ways that cost is created, and each can be researched.
Taught inMECS 540-2; MMSS 211-3
SourcesSchelling, Thomas C. (1960). The Strategy of Conflict. Harvard University Press. (book)
Fearon (1994). Domestic Political Audiences and the Escalation of International Disputes. American Political Science Review 88(3), 577-592. https://doi.org/10.2307/2944796
Instead of asking what a rational competitor would do, we ask which strategies are gaining share in the population of firms and which are losing it, because an industry converges on what works whether or not anyone planned it.
dx(i)/dt = x(i) x (payoff of i minus the average)John Maynard Smith and George Price asked in 1973 why animal contests are so rarely fights to the death, and answered with the evolutionarily stable strategy: a mix of behavior that no rare alternative can invade. Taylor and Jonker wrote the dynamics down in 1978 as the replicator equation, in which a strategy's share grows when it does better than the population average. MMSS teaches it in both game theory courses because it explains convergence without assuming anyone is rational.
Why it works nowWatching which practices gain and lose share across an industry used to require a trade association's survey and a decade. The share of each strategy in a market can now be read from filings, job postings, pricing pages, and the trade press year by year, and operators can say which practices are being copied. The Desk asks which practices are gaining share among the firms in this market. The population answers the question the argument cannot.
When it appliesA thesis that assumes rivals stay irrational or old-fashioned; a practice spreading through an industry; a question about whether a business model will be copied or will fade.
What we would askWhich practices are gaining share among the firms in this market, and which are losing it?
Could a single entrant doing something different take share from the incumbents, or would it be squeezed out?
The mix of strategies in the industry over time, and the relative performance of each in the current mix.
How we would get itOperators across the industry on which practices are winning and being copied; the record for share by strategy over time.
Where it breaksPayoffs treated as fixed when they shift as the mix changes; entry and imitation ignored.
For exampleA thesis assumes an incumbent's high-touch sales model will keep beating a rival's self-serve model. Across the industry, self-serve is gaining share every year and the high-touch firms are the ones being acquired or shrinking. The population, not the argument, says where this ends.
The mathematics in fulldx(i)/dt = x(i) x (f(i)(x) - fbar(x)): the share of a strategy grows when its payoff exceeds the average; a strategy that cannot be invaded by a rare alternative is where the industry settles.
Taught inMMSS 211-2; MMSS 311-1
SourcesMaynard Smith, Price (1973). The Logic of Animal Conflict. Nature 246(5427), 15-18. https://doi.org/10.1038/246015a0
Taylor, Jonker (1978). Evolutionary stable strategies and game dynamics. Mathematical Biosciences 40(1-2), 145-156. https://doi.org/10.1016/0025-5564(78)90077-9
What the seller knows, what the bidder should fear, what a contract will actually reward.
We treat the fact that something is offered as evidence about it. Owners who know their asset is bad sell readily; so the research asks what the seller knows and why they are selling now.
Quality offered = E[q | the seller accepts p]George Akerlof's 1970 paper started with used cars: owners know which cars are lemons and buyers do not, so buyers pay for average quality, owners of good cars withdraw, quality falls, and the market can vanish. Rothschild and Stiglitz showed in 1976 that insurers face the same problem and respond with menus of contracts that make customers reveal their risk by what they choose.
Why it works nowDiligence on a company for sale used to be a data-room exercise, which is precisely the information the seller chose to show. The seller's private information and motive are now researchable from outside the room: former employees and advisors of the seller, prior sale attempts, the buyers who walked and what they saw. The Desk asks why now, and who passed. The engagement is built to find out what the seller knows.
When it appliesAn asset, a company, a loan book, a portfolio, or a customer contract is on offer, and the seller knows more about it than the buyer can learn from the data room.
What we would askWhy is the seller selling now, and what has changed for them that the data room does not show?
Who else looked and passed, and what did they see?
The seller's private information and motives; the history of the asset with previous owners and previous bidders.
How we would get itFormer employees and advisors of the seller; the record of prior sale attempts, withdrawn processes, and the buyers who walked.
Where it breaksThe seller has a legitimate reason to sell that the analysis dismisses; or the buyer assumes the data room is the whole truth.
For exampleA specialty lender is selling a loan book at an attractive yield. The former head of underwriting, interviewed, explains the vintage that was originated under a relaxed policy in the year before the sale. The yield is the price of that vintage.
The mathematics in fullAt any price p, the quality offered is E[q | seller accepts p], which falls as p falls; the buyer who ignores this pays for the average and receives the below-average, and the market can unravel entirely.
Taught inMMSS 311-1; MECN-452; MECS 465
SourcesAkerlof (1970). The Market for "Lemons": Quality Uncertainty and the Market Mechanism. The Quarterly Journal of Economics 84(3), 488. https://doi.org/10.2307/1879431
Rothschild, Stiglitz (1976). Equilibrium in Competitive Insurance Markets: An Essay on the Economics of Imperfect Information. The Quarterly Journal of Economics 90(4), 629. https://doi.org/10.2307/1885326
In a competitive process, winning is itself evidence that your estimate was the most optimistic in the room. We find out who else looked, who passed, and at what price, before you congratulate yourself.
E[V | you won] is below E[V | your estimate]Three petroleum engineers at Atlantic Richfield, Capen, Clapp, and Campbell, published the winner's curse in 1971 after noticing that the companies winning Gulf of Mexico lease auctions earned poor returns on them: the winner is the bidder whose estimate was most optimistic. Richard Thaler's 1988 account includes the classroom version, an auction for a jar of coins, in which the winner reliably overpays.
Why it works nowThe correction requires knowing how many sophisticated parties looked and why the ones who passed did, which used to be gossip. It is now obtainable: advisors to other bidders in the same or comparable processes can be recruited, the process record shows the withdrawals, and the reasons can be put on the record. The Desk asks why you are the one winning. The engagement finds what the two bidders who walked after diligence saw.
When it appliesA competitive auction or financing where the client is the high bidder, or is about to be; or a client puzzled that it keeps winning the deals it wants.
What we would askHow many sophisticated parties evaluated this, and how many passed?
What did the ones who passed know?
The number and quality of other bidders, their reasons for passing, and the dispersion of estimates.
How we would get itAdvisors who advised other bidders in the same or comparable processes; the record of the process and its withdrawals.
Where it breaksThe bidders are treated as having independent private values when the asset has one true value everyone is estimating; the client's own optimism is not modeled.
For exampleA fund wins a competitive process for an industrial asset against six bidders. Two withdrew after diligence on the same environmental issue the fund's advisors called immaterial. The interviews go to people who saw what those two saw.
The mathematics in fullWith a common value V and noisy estimates, E[V | you won] is less than E[V | your estimate], and the gap grows with the number of bidders; the correction is to bid as if you have already learned that everyone else estimated lower.
Taught inMECN-452; DECS-430; MMSS 311-1
SourcesCapen, Clapp, Campbell (1971). Competitive Bidding in High-Risk Situations. Journal of Petroleum Technology 23(06), 641-653. https://doi.org/10.2118/2993-PA
Thaler (1988). Anomalies: The Winner's Curse. Journal of Economic Perspectives 2(1), 191-202. https://doi.org/10.1257/jep.2.1.191
Milgrom, Weber (1982). A Theory of Auctions and Competitive Bidding. Econometrica 50(5), 1089. https://doi.org/10.2307/1911865
Whether a client is selling or bidding, we work out how the rules of the process shape the result: sealed or open, one round or several, reserve or none, and how the other side will bid under them.
Truthful bidding is optimal under a second priceWilliam Vickrey's 1961 paper introduced the sealed-bid second-price auction, in which bidding your true value is the best strategy, and showed that different formats can raise the same expected revenue. Roger Myerson's 1981 paper characterized the revenue-maximizing auction and earned its share of a Nobel. Edelman, Ostrovsky, and Schwarz showed in 2007 that the search-advertising auctions run by Google and Yahoo were a generalized second-price format with their own quirks.
Why it works nowAuction theory was applied to spectrum and search advertising by the few firms that could hire the theorists. What has changed for an investor is that a company's exposure to auction rules, as a seller of inventory, a bidder for contracts, or a platform running the auction, can now be read from the record and tested with people who have run comparable processes. The Desk asks whether the bidders know their own values or are all guessing at the same one, because the two cases behave differently.
When it appliesA client is designing a sale process, entering one, bidding for spectrum or ad inventory or a contract, or underwriting a business whose revenue comes from auctions it runs or bids in.
What we would askDo the bidders each know their own value, or are they all estimating the same unknown?
Can the bidders talk to each other, and can the seller change the rules midway?
Whether values are private or common; the number of bidders; the seller's ability to commit to the rules; the possibility of collusion.
How we would get itAdvisors who have run or advised in comparable processes; the record of comparable auctions and their outcomes.
Where it breaksBidders collude; the seller cannot commit; the format assumes rational bidders in a room with none.
For exampleA company's revenue is set in online ad auctions. Its thesis assumes those auctions are competitive. The former head of an ad exchange, interviewed, explains how a change in the auction format two years earlier moved margin from the platform to the bidders, and what the next change would do.
The mathematics in fullTruthful bidding is a best response in a second-price format and not in a first-price one; under standard assumptions formats that allocate to the same bidder raise the same expected revenue, so the seller's lever is the reserve and the entry of bidders, not the format.
Taught inMECN-452; MECN-446; MECS 465
SourcesVickrey (1961). Counterspeculation, Auctions, and Competitive Sealed Tenders. The Journal of Finance 16(1), 8-37. https://doi.org/10.1111/j.1540-6261.1961.tb02789.x
Myerson (1981). Optimal Auction Design. Mathematics of Operations Research 6(1), 58-73. https://doi.org/10.1287/moor.6.1.58
Edelman, Ostrovsky, Schwarz (2007). Internet Advertising and the Generalized Second-Price Auction: Selling Billions of Dollars Worth of Keywords. American Economic Review 97(1), 242-259. https://doi.org/10.1257/aer.97.1.242
We read every incentive plan, earnout, and management contract by asking what measured thing it rewards and what unmeasured thing it therefore sacrifices, because that is what the people under it will do.
Include a signal only if it informs about effortBengt Holmström's 1979 paper set up the problem of paying an agent whose effort cannot be seen and proved the informativeness principle: any signal that carries information about the effort, however noisy, belongs in the contract, and any that does not should be left out. The trade-off between insurance and incentives that every compensation plan embodies is the one his model made explicit.
Why it works nowReading an earnout, a management plan, or a supplier contract for what it will actually reward has always been possible and rarely done thoroughly, because the terms were long and the second-order effects took an expert to see. The terms can now be read completely, the measured quantity identified, and the question put to former executives who worked under comparable plans: what would a clever person do to hit this number. The Desk asks what is measured and who controls the measurement.
When it appliesAn earnout, a management incentive plan, a sales compensation change, a supplier contract, or a portfolio company whose behavior the client cannot explain until they read the pay plan.
What we would askWhat exactly is measured, and who controls the measurement?
What would a clever person do to hit the number without doing the thing you wanted?
The full contract terms, the measurement process, and the noise in the measure relative to the effort it is meant to reward.
How we would get itFormer executives who worked under comparable plans; operators on how such measures are gamed in this industry; the record of past plans and what followed them.
Where it breaksThe measure is gameable, the agent controls it, or the plan rewards the measured task at the expense of the unmeasured one.
For exampleA portfolio company's management is paid on annual EBITDA. Maintenance capex has fallen for three years. The two facts are one fact, and the interviews with former plant managers establish what the deferred maintenance will cost.
The mathematics in fullPay w(y) on an observable y to maximize E[y - w(y)] subject to the agent choosing effort to maximize E[u(w(y)) - c(e)]; a noisy measure with a risk-averse agent forces weak incentives, and a signal belongs in the contract only if it adds information about effort.
Taught inMECS 465; MECN-441
SourcesHolmstrom (1979). Moral Hazard and Observability. The Bell Journal of Economics 10(1), 74. https://doi.org/10.2307/3003320
Salanie, Bernard (2005). The Economics of Contracts: A Primer, 2nd ed. MIT Press. (book)
What people will choose, what they will pay, and how a market is really shaped.
Asking people what they would pay is close to worthless. We make them choose between realistic alternatives and read what they give up, which is where the willingness to pay actually shows.
P(j) = exp(V(j)) / sum over k of exp(V(k))Daniel McFadden built discrete choice modeling in the early 1970s and used it to forecast ridership on the Bay Area Rapid Transit system before it opened, from how people chose among the alternatives they already had; the forecast held up. Green and Srinivasan's 1978 review brought the same logic to consumer research as conjoint analysis: present realistic bundles, record the choices, and recover what each attribute is worth.
Why it works nowA proper choice study needed a specialist to design and estimate it and a panel to field it, so it was reserved for consumer products with large budgets. The design is now the cheap part: Advisors and customers define the real alternatives and attributes in the interviews, and the Desk can explain the move in one sentence. Vista designs the study and brings in a specialist to field and estimate it. It does not run a survey business, and it says so.
When it appliesA pricing, packaging, or product decision rests on a survey of stated intentions, on management's belief about value, or on a handful of customer conversations that asked the direct question.
What we would askWho exactly is the buyer, and what are the real alternatives in front of them, including doing nothing?
Which two or three attributes do they actually trade off against price?
A defined buyer population, a realistic set of alternatives and attributes, and enough respondents from the actual buyer population to estimate trade-offs.
How we would get itAdvisors and customers define the attributes and alternatives; a designed choice exercise among qualified buyers produces the trade-offs; Vista designs it and brokers the fielding rather than running a survey business.
Where it breaksStated choices differ from purchases; the sample is not the buyer; the price range shown anchors the answers; the attributes are the product team's, not the buyer's.
For exampleA medical device company believes hospitals will pay a premium for a feature. Hospital purchasing leads, presented with realistic bundles, trade that feature away for service response time every time. The premium is in the wrong place.
The mathematics in fullP(choose j) = exp(V(j)) / sum over k of exp(V(k)), with V(j) = beta x attributes(j); willingness to pay for an attribute is the ratio -beta(attribute) / beta(price), estimated from choices rather than asked.
Taught inMMSS 386-2 (qualitative-choice models); MECN-446
SourcesGreen, Srinivasan (1978). Conjoint Analysis in Consumer Research: Issues and Outlook. Journal of Consumer Research 5(2), 103-123. https://doi.org/10.1086/208721
Train (2009). Discrete Choice Methods with Simulation. Cambridge University Press (book). https://doi.org/10.1017/CBO9780511805271
Pricing power is a number, the volume lost per point of price, and it is almost never known and almost always asked about directly, which does not work. We infer it from what happened when prices actually moved.
(P - MC) / P = -1 / eAlfred Marshall named elasticity in 1890, and every intermediate microeconomics course since, including MMSS Turbo Micro, has taught that a seller's power is the volume lost per point of price. The classic measurement failure is as old as the concept: asking customers what they would pay produces an answer, and the answer is wrong.
Why it works nowThe elasticity at the current price is now inferable rather than asked: past price moves can be matched to customer-level volume responses in the client's own data, competitor responses can be read from the record, and category buyers can be asked what the next increase would trigger. The Desk asks what happened to volume the last two times prices moved, and who left. The engagement puts the largest customer's buyer on the record.
When it appliesA thesis rests on the company's ability to raise prices; a company is deciding on a price change; or a market's pricing discipline is being underwritten from the outside.
What we would askWhen did prices last move, and what happened to volume, customer by customer if possible?
Who left, and where did they go?
Past price changes with volume responses, ideally at the customer level; competitor responses to those changes; the customers' alternatives.
How we would get itFormer pricing leaders and category buyers on real responses to past increases; the record of price moves and share shifts; a designed choice exercise where the history is silent.
Where it breaksThe elasticity is estimated from a survey; competitors respond and the estimate assumed they would not; the past increase coincided with something else.
For exampleAn investor underwrites a consumables business as having pricing power. The last two increases held volume, but the category buyer at the largest customer, interviewed, explains that the third will trigger a dual-sourcing program already approved. The power is real and it is nearly used up.
The mathematics in fullElasticity e = (dQ/Q) / (dP/P); a single profit-maximizing price satisfies (P - MC) / P = -1/e, so the whole question is e at the current price, which is inferred from natural experiments, discontinuities, or designed choices, never from asking.
Taught inMMSS 211-1; MECN-430; MECN-446
SourcesFrank, Robert H. Microeconomics and Behavior. McGraw-Hill. (book)
Besanko, David and Braeutigam, Ronald. Microeconomics, 5th ed. Wiley. (book)
We work out which customers value what, and whether versions, bundles, or quantity terms would capture value a single price leaves behind, and whether the customers could undo it by trading among themselves.
Bundle when component values are negatively correlatedAdams and Yellen's 1976 paper showed with a handful of customers and two goods why bundling raises profit when customers who value one good highly value the other little: the bundle captures value that two separate prices leave behind. Stigler had seen the same logic earlier in the block booking of films. Mussa and Rosen supplied the theory of versions.
Why it works nowThe correlation of values across components, which decides whether a bundle works, was invisible because nobody could ask enough customers. It is now the object of a designed choice exercise among qualified buyers, informed by interviews that say which segments exist. The Desk asks who gets the most out of the product and how you would tell them apart at the point of sale. The engagement sizes the segments.
When it appliesA single price for a product whose customers value it very differently; a bundle whose logic nobody can state; or a competitor's tiering that is winning.
What we would askWhich customers get the most out of this, and how would you tell them apart at the point of sale?
Could the customers who pay less resell to the ones who pay more?
The distribution of value across customers and the correlation of values across components; the feasibility of keeping segments apart.
How we would get itCustomers and former sales leaders on value by segment; a choice exercise across candidate bundles among qualified buyers.
Where it breaksArbitrage; segments that cannot be identified at the point of sale; competitors who unbundle.
For exampleA data vendor sells one enterprise license. Its customers split into a group that values the history and a group that values the feed, and the two groups' values run opposite. A bundle captures more than two separate prices would, and the research is the size of the two groups.
The mathematics in fullBundling profits when reservation values for the components are negatively correlated across customers; a menu of versions works when the low version is degraded just enough that the high-value customer will not take it.
Taught inMECN-446; MECN-430
SourcesAdams, Yellen (1976). Commodity Bundling and the Burden of Monopoly. The Quarterly Journal of Economics 90(3), 475. https://doi.org/10.2307/1886045
Mussa, Rosen (1978). Monopoly and product quality. Journal of Economic Theory 18(2), 301-317. https://doi.org/10.1016/0022-0531(78)90085-6
In a market with a few players, the success of any strategy depends on how the others respond. We map the industry's structure first and then ask which moves soften competition and which invite it.
The sign of the rival's best response decides the moveThe theory of strategic interaction among a few firms runs from Cournot and Bertrand through the game-theoretic industrial organization of the 1980s; Bulow, Geanakoplos, and Klemperer's 1985 taxonomy of strategic substitutes and complements is the working tool, and Fudenberg and Tirole gave the postures their nicknames. Kellogg's competitive strategy course teaches it through cases in which a move that would work for a monopolist backfires in an oligopoly.
Why it works nowMapping an industry's structure and each player's incentives took a consulting engagement. The structure can now be assembled from filings and trade press quickly, and the incentives from former executives across the players, so that the history of moves and responses is on the table before the strategy is chosen. The Desk asks how many players actually set price here and what the last aggressive move provoked.
When it appliesA strategy (product proliferation, positioning, loyalty programs, capacity additions, a price cut) in a concentrated industry; or a thesis that reads margins without reading structure.
What we would askHow many players actually set price here, and what does each one want?
What did the last aggressive move in this industry provoke?
Industry structure, each player's cost position and incentives, and the history of moves and responses.
How we would get itFormer executives across the players; the record of past moves; the trade press and regulatory filings for structure.
Where it breaksThe industry is assumed to behave like its average firm; the response is assumed away; the structure is changing (entry, consolidation, substitutes).
For exampleA company plans to fill every price point with a variant to crowd out a rival. In its three-player market the rival's former strategy head explains that the last such move triggered a capacity race that cut everyone's margin for four years.
The mathematics in fullWhether a move is strategic substitute or complement decides the response: aggression met by retreat in the first case and by matching in the second; the sign is a property of the industry, and it can be established from history before it is modeled.
Taught inMECN-441; MECS 449; MMSS 211-2
SourcesBesanko, Dranove, Shanley and Schaefer. Economics of Strategy, 7th ed. Wiley. (book)
Bulow, Geanakoplos, Klemperer (1985). Multimarket Oligopoly: Strategic Substitutes and Complements. Journal of Political Economy 93(3), 488-511. https://doi.org/10.1086/261312
We look for the places where the people in the market do not behave as the textbook assumes: reference prices, loss aversion, sunk-cost pricing, anchors; and we check whether the client's own managers are doing it too.
Value v(x) from a reference point, steeper in lossesKahneman, Knetsch, and Thaler's 1991 review reported the coffee mug experiments at Cornell, where students given a mug demanded about twice as much to sell it as others would pay to buy it. Arkes and Blumer's 1985 theater experiment sold season tickets at different discounts and found that people who paid full price attended more, because the sunk cost felt like a reason. Al-Najjar, Baliga, and Besanko showed in 2008 that managers who fold sunk costs into prices survive in markets that the textbook says should punish them.
Why it works nowBehavioral anomalies were explanations offered after the fact. They can now be turned into predictions and tested: customers can be asked what number they compare a price to, responses to increases and decreases can be measured separately in the data, and former pricing managers can say how prices were really set. The Desk asks whether the reaction to a cut and to a raise was symmetric, and where the reference price came from.
When it appliesPricing that customers react to asymmetrically; managers who price off fully loaded cost including sunk investment; a market where the standard model keeps being wrong in the same direction.
What we would askWhen you raised the price and when you cut it, was the reaction symmetric?
What number are your customers comparing your price to, and where did that number come from?
Customer responses to increases and decreases separately; the reference points in the buyers' heads; the cost allocation the managers price from.
How we would get itCustomers and category buyers on reference prices and reactions; former pricing managers on how prices were actually set.
Where it breaksEvery anomaly is called behavioral after the fact; the bias story is never turned into a testable prediction.
For exampleA subscription business finds cuts win few customers and increases lose many. The asymmetry is loss aversion around the reference price, and the research is what that reference price is and whether it can be moved before the next increase.
The mathematics in fullValue is v(x) relative to a reference point, steeper for losses than gains; managers who allocate sunk cost into price behave as if the reference were fully loaded cost, and a market of such managers prices differently from the textbook one.
Taught inMECN-943; MECNX-435
SourcesKahneman, Knetsch, Thaler (1991). Anomalies: The Endowment Effect, Loss Aversion, and Status Quo Bias. Journal of Economic Perspectives 5(1), 193-206. https://doi.org/10.1257/jep.5.1.193
Al-Najjar, Baliga, Besanko (2008). Market forces meet behavioral biases: cost misallocation and irrational pricing. The RAND Journal of Economics 39(1), 214-237. https://doi.org/10.1111/j.0741-6261.2008.00011.x
Arkes, Blumer (1985). The psychology of sunk cost. Organizational Behavior and Human Decision Processes 35(1), 124-140. https://doi.org/10.1016/0749-5978(85)90049-4
Adoption follows an S-curve driven by early adopters and then by imitation, and the two things that matter, the ceiling and the timing of the inflection, are invisible in the early data that looks exponential. We find them from comparable adoptions.
Adopters = (p + q F) x (M - cumulative)Frank Bass's 1969 model fitted the adoption of color televisions, refrigerators, room air conditioners, and other durables with two parameters, innovation and imitation, and a ceiling, and predicted the timing of the color television peak. It became the standard forecasting model for new products, and its lesson has not changed: the early data look exponential and are not.
Why it works nowThe ceiling and the inflection, which the early data cannot identify, used to be assumed. Both can now be researched: comparable adoption curves can be retrieved from the record with their inflection points, and the people who have not adopted can be interviewed about why. The Desk asks what the ceiling is and how you know, and which past adoption this most resembles. The engagement finds the non-adopters.
When it appliesA growth thesis extrapolated from early adoption; a market whose size is inferred from its growth rate; a product whose adoption has stalled unexpectedly.
What we would askWhat is the ceiling, and how do you know? Who will never adopt this, and why?
Which past adoption does this most resemble, and where was its inflection?
The addressable ceiling from evidence; the imitation dynamics of this category; comparable adoption curves with their inflection points.
How we would get itOperators who lived through the comparable adoptions; customers who have not adopted on why; the record of comparable curves.
Where it breaksThe S-curve is fitted to its own first third; the ceiling is assumed; the growth is subsidy-driven and mistaken for organic imitation.
For exampleA venture thesis extrapolates two years of triple-digit growth. The comparable category inflected at eleven percent penetration, and the non-adopters, interviewed, have a reason that the product does not address. The ceiling, not the growth rate, is the thesis.
The mathematics in fullAdoptions at t = (p + q x F(t)) x (M - cumulative adopters): innovation rate p, imitation rate q, ceiling M; early data identifies p and q poorly and M not at all, so M comes from research, not from the curve.
Taught inMECN-451; MMSS 1995 (exponential growth and decline)
SourcesBass (1969). A New Product Growth Model for Consumer Durables. Management Science 15(5), 215-227. https://doi.org/10.1287/mnsc.15.5.215
We ask whether the market is driven by a small tail, because when it is, the average customer is nearly meaningless and the thesis lives or dies on a handful of names.
P(X > x) = (x_min / x)^alphaVilfredo Pareto found in the 1890s that income follows a distribution in which a small fraction of people hold most of the wealth, and the same shape turned up in city sizes, which Xavier Gabaix explained in 1999. Clauset, Shalizi, and Newman's 2009 paper tested twenty-four datasets and found that many claimed power laws were not, which is the caution the Library carries.
Why it works nowConcentration was a line in the risk factors. It can now be measured from the client's data and the record, customer by customer and partner by partner, and the largest accounts themselves can be interviewed. The Desk asks what share of revenue, profit, or leads the top five percent produce, and what is left if two of them leave. When the answer is most of it, the thesis is those names, and the engagement goes to them.
When it appliesA total addressable market computed as customers times average spend; a customer base described by its average; a channel where three partners produce most of the leads.
What we would askWhat share of revenue, profit, or leads comes from the top five percent?
If two of those left, what is left?
The concentration of value across customers, products, channels, and geographies, from the data rather than from the deck.
How we would get itThe record for concentration; interviews with the operators who know which accounts actually carry the business; the largest customers themselves where reachable.
Where it breaksA lognormal is mistaken for a power law; the tail is extrapolated from a handful of observations; the concentration is treated as a risk factor rather than as the thesis.
For exampleA marketplace thesis rests on average take rate across thousands of sellers. Twelve sellers produce most of the volume and are negotiating fees as a group. The market is those twelve, and the interviews go to them.
The mathematics in fullP(X > x) = (x_min / x)^alpha; below alpha of two the variance is infinite and sample averages never settle, so any thesis stated in averages is a thesis about a number that does not exist.
Taught inMMSS 1995 (rank-size and Pareto laws); MECN-451
SourcesGabaix (1999). Zipf's Law for Cities: An Explanation. The Quarterly Journal of Economics 114(3), 739-767. https://doi.org/10.1162/003355399556133
Clauset, Shalizi, Newman (2009). Power-Law Distributions in Empirical Data. SIAM Review 51(4), 661-703. https://doi.org/10.1137/070710111
When one party's activity imposes a cost or a benefit on another, the question is not who is at fault but who can fix it most cheaply and whether the two can bargain. We find out what the bargaining costs and who holds the right.
With low transaction costs the assignment does not change the outcomeRonald Coase's 1960 paper used a rancher whose cattle stray onto a farmer's crops, and sparks from a railway that set fields alight, to make the point that when the parties can bargain cheaply, the efficient outcome is reached whichever of them holds the right; the assignment changes who pays, not what happens. When bargaining is costly, which is most of the time, the assignment decides the outcome.
Why it works nowThe costs on each side of a spillover, and the cost of a deal among the parties, used to be assumed. They are now researchable: operators can say what abatement actually costs, former regulators can say how the right has been assigned and reassigned, and the record shows comparable bargains. The Desk asks who can fix the problem most cheaply and whether the two sides can actually bargain.
When it appliesPollution, congestion, noise, data, shared platforms, or any business whose activity lands costs or benefits on parties outside the transaction; a regulation that assigns the right to one side.
What we would askWho can reduce the harm at the lowest cost, the one causing it or the one suffering it?
Can the two sides actually bargain, or are there too many of them?
The cost of abatement on each side, the number of parties, and the transaction costs of a deal among them.
How we would get itOperators on the real cost of abatement; former regulators on how the right has been assigned and re-assigned; the record on comparable bargains.
Where it breaksAssuming the polluter should always pay when the sufferer could adapt more cheaply; assuming bargaining is possible among thousands.
For exampleA logistics company's night deliveries impose noise on a neighborhood, and a proposed rule would ban them. The cheapest fix is quieter trucks, which the company would buy for less than the ban costs it, if the rule assigned the right and let the parties settle. The research is the cost on each side.
The mathematics in fullWith a clear assignment of the right and low transaction costs, the parties bargain to the efficient outcome whichever side holds the right; the distribution changes, the efficiency does not. With high transaction costs, the assignment decides the outcome.
Taught inMMSS 211-1; MECS 540-2
SourcesCoase (1960). The Problem of Social Cost. The Journal of Law and Economics 3, 1-44. https://doi.org/10.1086/466560
When customers pick the nearest option, competitors are pulled toward the middle of the market and end up nearly identical. We ask whether this market rewards moving to the center or being the one firm that does not.
Location alone pulls together; price competition pushes apartHarold Hotelling's 1929 paper put two sellers on a beach with customers spread along it, each customer walking to the nearer seller, and showed that both sellers end up side by side in the middle. Anthony Downs applied the same logic to political parties in 1957, which is why MMSS teaches it in the formal political models course as electoral competition. Add price competition and the sellers move apart, because being close invites undercutting.
Why it works nowPositioning has been decided from a two-by-two drawn in a workshop. The distribution of customer preferences along the dimension that matters can now be measured from choices and from interviews, and former strategy leaders can say what past moves toward or away from the center actually did. The Desk asks whether customers choose the closest option or have strong preferences for the extremes.
When it appliesA positioning decision: where to put a product, a store, a price point, or a platform relative to rivals; a market where all the competitors look the same; a political or standards contest.
What we would askDo customers choose the closest option to what they want, or do they have strong preferences for the extremes?
If you moved toward your rival, would you take their customers or start a price war?
The distribution of customer preferences along the dimension that matters, and whether firms compete on position, price, or both.
How we would get itCustomers on how they actually choose; former strategy leaders on past positioning moves and their consequences.
Where it breaksAssuming the center is always best; with price competition, firms often do better apart than together.
For exampleTwo regional grocers have converged on the same format in the same suburbs, and margins have compressed. The model says that is exactly what competing on location alone produces. The research question is whether a differentiated position exists that customers would pay for.
The mathematics in fullWith customers spread along a line and each choosing the nearest seller, two sellers choosing location alone converge on the center; add price competition and they separate, because proximity invites undercutting.
Taught inMMSS 211-3; MECN-441
SourcesHotelling (1929). Stability in Competition. The Economic Journal 39(153), 41. https://doi.org/10.2307/2224214
Downs (1957). An Economic Theory of Political Action in a Democracy. Journal of Political Economy 65(2), 135-150. https://doi.org/10.1086/257897
Before a thesis assumes what worked in one market transfers to another, we measure which markets, customers, or products actually behave alike, on the dimensions that matter, rather than on the map.
Merge nearest pairs by distance on the chosen dimensionsHierarchical clustering and multidimensional scaling were taught in the 1995 MMSS network unit as ways to find which units in a population resemble each other and to draw the resemblance on a page. Hotelling's 1933 principal components is the older cousin: reduce many measurements to the few directions that carry the variation.
Why it works nowRollout decisions have grouped markets by geography because geography was the available dimension. The dimensions that actually drive behavior, density, competitive intensity, labor cost, channel mix, can now be assembled for every candidate market and the grouping done on those. The Desk asks on which dimensions these markets resemble each other and whether those are the dimensions that drove the result.
When it appliesAn expansion from one market to the next; a portfolio of products or regions treated as one; a claim that a result in one segment generalizes.
What we would askOn which dimensions do these markets resemble each other, and are those the dimensions that drove the result?
Which pair that looks alike on the map behaves most differently in the data?
Data on the units across the dimensions that matter, and a defined notion of similarity.
How we would get itOperators on which dimensions actually drive behavior; the data grouped on those dimensions rather than on geography.
Where it breaksGrouping on what is easy to measure; letting the clustering method decide the number of groups.
For exampleA retailer plans to roll out a format that worked in three cities to twenty. Grouped by customer density, competitive intensity, and labor cost rather than by region, only six of the twenty resemble the three, and the rollout plan is a different plan.
The mathematics in fullDefine a distance between units on the chosen dimensions; hierarchical clustering merges the nearest pairs step by step, and multidimensional scaling lays the units out so that distances on the page approximate the distances in the data.
Taught inMMSS 1995 (hierarchical clustering and multidimensional scaling)
SourcesHotelling (1933). Analysis of a complex of statistical variables into principal components. Journal of Educational Psychology 24(6), 417-441. https://doi.org/10.1037/h0071325
Dempster, Laird, Rubin (1977). Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm. Journal of the Royal Statistical Society Series B: Statistical Methodology 39(1), 1-22. https://doi.org/10.1111/j.2517-6161.1977.tb01600.x
Whether X did cause Y, whether it would, and for whom.
Before anything elaborate, we fit the simple, checkable relationship, and we insist that anything more complicated beat it; and we remember that what it shows is association, not cause.
beta = (X'X)^-1 X'YFrancis Galton measured the heights of parents and their adult children in 1886 and found that the children of tall parents were tall, but less so, and the children of short parents short, but less so: regression toward mediocrity, from which the method took its name. Every econometrics course since, MMSS's included, has taught it as the transparent baseline and the place where causal claims quietly begin.
Why it works nowThe baseline is faster to fit than it has ever been, which is not the change that matters. The change is that the omitted variables an industry insider would name in a sentence can now be asked for, in the interviews, before the regression is run, and that an elaborate model can be held to the discipline of beating the simple one. The Desk asks what else changed at the same time. That question is most of applied econometrics.
When it appliesA claim that some observable predicts an outcome (sales hiring predicts bookings, reviews predict churn, turnover predicts performance); or a complex model whose value over a simple one is unstated.
What we would askWhat is the simplest relationship in the data, and does the elaborate model actually do better than it?
What else changed at the same time that could explain the pattern?
The data at the right grain; the variables that could drive both sides of the relationship.
How we would get itThe record and the client's own data; Advisors on which omitted variables matter in this industry.
Where it breaksOmitted variables, reverse causation, selection into the sample, a specification chosen after seeing the results.
For exampleA deck shows that customers who use a feature renew more. The feature's users are also the largest and longest-tenured customers, and once that is held constant the relationship disappears. The feature is a marker, not a cause.
The mathematics in fullY = X beta + e with beta-hat = (X'X)^-1 X'Y; each coefficient is an association holding the other regressors fixed, and an omitted driver of both X and Y appears as a coefficient on X.
Taught inMATH 386-1; DECS-431
SourcesGalton (1886). Regression Towards Mediocrity in Hereditary Stature. The Journal of the Anthropological Institute of Great Britain and Ireland 15, 246. https://doi.org/10.2307/2841583
Wooldridge, Jeffrey M. Introductory Econometrics: A Modern Approach. Cengage. (book)
Every causal claim is a comparison with a world that did not happen. We state that world explicitly, and then ask what could stand in for it: a control group, a comparison period, a threshold, a natural experiment. Before-and-after on its own settles nothing.
tau = Y(1) - Y(0), only one ever observedDonald Rubin's 1974 paper wrote the causal question in its modern form: every unit has an outcome under the treatment and an outcome without it, only one of which can ever be observed, and every method is a way of estimating the other. The setting was educational programs, and the framework now underlies the entire MMSS econometrics sequence.
Why it works nowManagement decks are full of causal claims supported by before-and-after charts, and readers have always lacked the time to ask what else happened. The comparison group is usually available in the record, the market, the closest competitor, the regions that did not get the change, and it can now be assembled quickly. The Desk asks what would have happened to the same business over the same period without the change. The brief grades the causal evidence honestly.
When it appliesA claim that something caused something (the new CEO caused growth, the price change caused churn, the redesign lifted conversion, advertising drove bookings) supported by a before-and-after chart.
What we would askWhat would have happened to the same business over the same period without the change?
Who did not get the change, and how did they do?
A defensible stand-in for the counterfactual: units that did not receive the treatment, a period the treatment could not have affected, or a rule that assigned it.
How we would get itThe Research Director identifies the comparison group or the natural experiment in the record; Advisors say what else changed; the report grades the causal evidence honestly.
Where it breaksThe identifying assumption fails on the facts: the comparison group differs in the way that matters, the trends were not parallel, the instrument has a second channel.
For exampleA management deck credits a new sales leader with a doubling of bookings. The market doubled in the same year and the closest competitor's bookings doubled too. The comparison group was available all along; it was just not in the deck.
The mathematics in fulltau(i) = Y(i, 1) - Y(i, 0), only one of which is observed; every method is a way to estimate the unobserved term, and the assumption that licenses it (parallel trends, exclusion, no manipulation at the cutoff) is a statement about the unobserved world that the data cannot test.
Taught inMATH 386-2; DECS-431
SourcesRubin (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology 66(5), 688-701. https://doi.org/10.1037/h0037350
We compare the change in the group that received the treatment with the change in a group that did not, over the same period. It is much stronger than before-and-after alone and it still rests on one assumption we state out loud.
(T after - T before) - (C after - C before)Ashenfelter and Card's 1985 study compared the earnings of people who entered a federal training program with a comparison group before and after, and found the famous dip: trainees' earnings fell just before they enrolled, which is why before-and-after alone overstated the program's effect. Card and Krueger's minimum wage comparison of New Jersey and Pennsylvania made the design famous a decade later, and Bertrand, Duflo, and Mullainathan showed in 2004 how easily it produces false precision when the errors are handled carelessly.
Why it works nowThe design needs outcomes for a treated group and an untreated one before and after, which a company usually has and has never lined up. The data can now be organized that way in an afternoon, and Advisors can say what else differed between the groups. The Desk asks who got the change and who did not, and whether the two were moving together before it. The estimate is a specialist's; the design is Vista's.
When it appliesA change was introduced to some customers, regions, products, or plants and not others, and the client wants to know what it did.
What we would askWho got the change and who did not, and were the two groups moving in parallel before it?
What else happened to the treated group that did not happen to the others?
Outcomes for both groups before and after; evidence that the groups trended together before the change.
How we would get itThe client's data or the record for both groups; Advisors on what else differed between them.
Where it breaksThe trends were not parallel; the control group was contaminated; the treated group was chosen because it was already changing.
For exampleA company claims its new pricing lifted revenue per customer by eighteen percent. Customers on the new pricing rose eighteen percent; comparable customers still on the old pricing rose eleven. The effect is seven, and the report says so with the assumption attached.
The mathematics in fulltau = (Y(treated, after) - Y(treated, before)) - (Y(control, after) - Y(control, before)); valid under parallel trends, and the standard errors must respect that outcomes within a group are correlated over time.
Taught inMATH 386-2; DECS-431
SourcesAshenfelter, Card (1985). Using the Longitudinal Structure of Earnings to Estimate the Effect of Training Programs. The Review of Economics and Statistics 67(4), 648. https://doi.org/10.2307/1924810
Bertrand, Duflo, Mullainathan (2004). How Much Should We Trust Differences-In-Differences Estimates?. The Quarterly Journal of Economics 119(1), 249-275. https://doi.org/10.1162/003355304772839588
When a rule assigns something by a threshold (a credit score, a size cutoff, an eligibility date), the units just either side of the line are nearly identical except for the treatment, and comparing them is close to an experiment.
Compare the limits as X approaches the cutoffThistlethwaite and Campbell's 1960 paper compared students just above and just below the test-score cutoff for a National Merit certificate, who were nearly identical in everything but the award, and read the award's effect from the gap between them. Imbens and Lemieux's 2008 guide is the modern practitioner's manual.
Why it works nowThresholds are everywhere in business and regulation and were rarely exploited, because nobody thought of a size cutoff or an eligibility date as an experiment. The rules and the outcomes on both sides of them can now be found in the record quickly, and operators can say whether the threshold was gamed. The Desk asks whether a rule with a cutoff decided who got this. When one did, the cleanest evidence obtainable is sitting either side of the line.
When it appliesA treatment was assigned by a cutoff: a loan program, a regulatory threshold, a pricing tier, a promotion rule; and the client wants its effect.
What we would askIs there a rule with a threshold that decided who got this?
Could anyone have moved themselves across the line on purpose?
The assignment rule, outcomes on both sides of it, and enough units near the threshold.
How we would get itThe record for the rule and the outcomes; operators on whether the threshold was gamed.
Where it breaksUnits manipulated their position at the cutoff; the effect near the line is not the effect far from it.
For exampleA regulation applies to firms above a revenue threshold. Firms just above and just below it differ in nothing but the regulation, and the comparison of their subsequent growth is the cleanest evidence obtainable on what the regulation costs.
The mathematics in fullEffect = lim of E[Y | X just above c] - E[Y | X just below c], where c is the cutoff assigning treatment; identification requires no manipulation of X at c and continuity of everything else across it.
Taught inMATH 386-2
SourcesThistlethwaite, Campbell (1960). Regression-discontinuity analysis: An alternative to the ex post facto experiment. Journal of Educational Psychology 51(6), 309-317. https://doi.org/10.1037/h0044319
Imbens, Lemieux (2008). Regression discontinuity designs: A guide to practice. Journal of Econometrics 142(2), 615-635. https://doi.org/10.1016/j.jeconom.2007.05.001
When the supposed cause was chosen by the people it affected, we look for something that shifted it from outside, and affected the outcome only through it. Finding one is the hard part, and we say plainly when there is none.
Effect = Cov(Z, Y) / Cov(Z, D)Angrist, Imbens, and Rubin's 1996 paper clarified what an instrument identifies: the effect for the people the instrument moved. The teaching example is Angrist's use of the Vietnam draft lottery, which assigned military service by birthdate for reasons unrelated to later earnings, to measure the effect of service on earnings.
Why it works nowFinding a valid instrument is still the hard part and always will be; no tool changes that. What has changed is that the accidents, rules, and shocks that moved a variable for reasons unrelated to the outcome, a regulatory change in some states, a supply disruption, a policy phased in by region, can be surfaced from the record and from Advisors far faster than before. The Desk asks what moved the cause for reasons that had nothing to do with the outcome. It says plainly when there is nothing.
When it appliesThe cause and the effect choose each other: advertising and sales, sales intensity and growth, adoption and productivity, expansion and performance.
What we would askWhat moved the cause for reasons that had nothing to do with the outcome?
Could that thing have affected the outcome any other way?
A source of variation in the cause that is unrelated to the outcome except through the cause; and the honesty to say when none exists.
How we would get itAdvisors on the accidents, rules, and shocks that moved the variable in this industry; the record for them.
Where it breaksThe instrument has a second channel; it is weak; it identifies the effect only for the units it moved.
For exampleDoes advertising drive bookings, or do companies advertise when bookings are strong? A regulatory change that restricted advertising in some states and not others moved advertising for reasons unrelated to demand, and the comparison across states answers the question.
The mathematics in fullZ shifts D and affects Y only through D; the effect is Cov(Z, Y) / Cov(Z, D), estimated in two stages, and it is the effect for the units Z moved, not for everyone.
Taught inMATH 386-2
SourcesAngrist, Imbens, Rubin (1996). Identification of Causal Effects Using Instrumental Variables. Journal of the American Statistical Association 91(434), 444-455. https://doi.org/10.1080/01621459.1996.10476902
An average effect hides the customers for whom it is large and the ones for whom it is zero. We ask which segments the product, the policy, or the change actually works for, because a thesis that assumes it generalizes usually does not.
CATE(x) = E[Y(1) - Y(0) | X = x]Athey and Imbens's 2016 paper built causal trees, which search for the subgroups where a treatment's effect differs while keeping the estimates honest by splitting the sample. It formalized what every experienced operator knows: the average effect of a product, a policy, or a change hides the customers for whom it is large and the ones for whom it is nothing.
Why it works nowHeterogeneity used to be discovered by accident after a launch failed in a segment. It can now be looked for on purpose: customers across segments can be interviewed about what changed for them, the data split by the characteristics the Advisors name, and the market thesis tested against the segment the client actually plans to sell to. The Desk asks for which customers it worked best and what they have in common.
When it appliesA claim that a product improves something by an average amount; a thesis that a result in one segment or geography will transfer to others; a pilot that worked.
What we would askFor which customers did it work best, and what do they have in common?
Is the segment you plan to sell to more like the ones where it worked or the ones where it did not?
Outcomes by segment with enough units per segment; the characteristics that could predict where the effect lives.
How we would get itCustomers across segments on what changed for them; the data split by the characteristics the Advisors say matter.
Where it breaksSegments chosen after seeing the results; too few units per segment; the effect in the pilot segment assumed for the market.
For exampleA workflow product improves productivity eight percent on average. In firms with a dedicated administrator it improves twenty; in firms without one it improves nothing. The market thesis was built on the average, and the addressable market is the first kind of firm.
The mathematics in fullCATE(x) = E[Y(1) - Y(0) | X = x]; estimated by splitting on covariates chosen before the outcome is examined, with honest sample splitting so the segments are not fitted to the noise.
Taught inMATH 386-2
SourcesAthey, Imbens (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences 113(27), 7353-7360. https://doi.org/10.1073/pnas.1510489113
When the decision can be tested on a slice before it is made for everyone, the cleanest evidence obtainable is an experiment, and we design it so that it is large enough to detect the effect that would matter.
Sample needed scales with 1 / (effect size squared)Ronald Fisher designed the first randomized field experiments at Rothamsted in the 1920s to test fertilizers on plots of land, and randomization has been the gold standard since. Kohavi and colleagues' 2009 guide brought it to the web, where Microsoft and Amazon learned that most confident product ideas fail a controlled test, and that the test must be large enough to see the effect that matters.
Why it works nowA business decision can be tested on a random slice before it is made for everyone, and most are not, because designing the test was a specialist's job and the sample-size arithmetic an afterthought. The design is now the deliverable: Advisors and the client's operators shape it, the smallest effect that matters sets the sample, and the client runs it. The Desk asks whether this could be tried on a random slice first. When it can, that is the cleanest evidence there is.
When it appliesA pricing change, a product change, a channel, or a policy that could be rolled out to a random subset first; or a pilot that was run without a control or without enough units.
What we would askCould you try this on a random slice of customers, regions, or products before you decide?
How small an effect would still matter, and how many units does it take to see it?
Random assignment, a control group, an outcome measured the same way in both, and a sample large enough for the smallest effect that matters.
How we would get itVista designs the experiment with the Advisors and the client's operators; the client runs it; the report reads it, including the null.
Where it breaksThe pilot was not random; the sample was too small and the null is meaningless; the outcome was measured differently in the two arms; the experiment changed behavior by being observed.
For exampleA company wants to know whether a lower price increases lifetime value. It can offer the lower price to a random half of new customers in two regions for a quarter. The design is the deliverable; the quarter is the evidence.
The mathematics in fullWith random assignment, E[Y | treated] - E[Y | control] is the average effect; the sample needed to detect an effect of size d at conventional thresholds scales with 1/d^2, so halving the smallest effect that matters quadruples the sample.
Taught inMATH 386-2; DECS-430
SourcesKohavi, Longbotham, Sommerfield (2009). Controlled experiments on the web: survey and practical guide. Data Mining and Knowledge Discovery 18(1), 140-181. https://doi.org/10.1007/s10618-008-0114-1
Cohen (1992). A power primer. Psychological Bulletin 112(1), 155-159. https://doi.org/10.1037/0033-2909.112.1.155
Who you heard from decides what you learned. Customers who answer differ from customers who do not; surviving customers cannot tell you about churn; experts who volunteer differ from experts who know. We design the sample before we trust the finding.
E[Y | observed] is not E[Y]The Literary Digest polled millions of its readers in 1936 and predicted a landslide for Landon; Roosevelt won in one, because the readers were not the electorate. James Heckman's 1979 paper gave selection its statistics with the case of women's wages, observed only for the women who chose to work. Groves's 2006 review showed that a low response rate does not by itself bias a survey; the difference between those who answer and those who do not does.
Why it works nowReference calls chosen by the seller and surveys answered by the happy have always been the sample a buyer got. The missing population can now be found and recruited deliberately: churned customers, investors who passed, employees who left. The Desk asks who was in a position to be in this sample and who was not. The engagement interviews the ones who were not.
When it appliesA finding from a customer survey, a set of reference calls chosen by the seller, a panel of volunteers, or any sample whose members chose to be in it.
What we would askWho was in a position to be in this sample, and who was not?
Did the ones who declined differ from the ones who answered in the way that matters?
The mechanism that put each unit in the sample; some evidence about the ones who are missing.
How we would get itThe Research Director recruits from the missing population deliberately: churned customers, passed investors, former employees who left; the report weights accordingly.
Where it breaksThe sample selected itself and the finding is the selection; the missing units are assumed to resemble the present ones.
For exampleA seller provides ten reference customers, all happy. The research recruits five customers who left in the last two years. The reference calls describe the product; the churned customers describe the decision.
The mathematics in fullE[Y | observed] differs from E[Y] whenever the probability of being observed depends on Y; the bias in a mean is roughly the nonresponse rate times the gap between respondents and nonrespondents, so the gap has to be estimated, not assumed away.
Taught inMATH 386-2; DECS-430
SourcesHeckman (1979). Sample Selection Bias as a Specification Error. Econometrica 47(1), 153. https://doi.org/10.2307/1912352
Groves (2006). Nonresponse Rates and Nonresponse Bias in Household Surveys. Public Opinion Quarterly 70(5), 646-675. https://doi.org/10.1093/poq/nfl033
A chart of one rising series against another proves nothing, because two series that both trend will always appear related. We difference the series, or ask whether the relationship survives once the trend is removed.
Difference the series before believing the coefficientClive Granger and Paul Newbold showed in 1974 that regressing one random walk on another produces a statistically significant relationship most of the time, with respectable-looking fit and nothing behind it. The paper changed how time series are handled in econometrics, and MMSS's econometrics text devotes a section of chapters to the discipline it created.
Why it works nowDecks full of charts with two rising lines have never been easier to produce. The test that defeats them, redoing the analysis on changes rather than levels, or checking whether the two series share a common trend, now takes minutes with the client's own data, and Advisors can name what else rose over the period. The Desk asks whether the relationship survives once the trend is removed.
When it appliesA thesis supported by a chart of two trending series; a claim that A drives B because both rose over the decade; any regression on time-series data.
What we would askIf you look at year-to-year changes rather than levels, does the relationship survive?
What else rose over the same period?
The series themselves, at the right frequency, and enough history to distinguish trend from relationship.
How we would get itThe record for the series; Advisors on what else moved; the analysis redone on changes rather than levels.
Where it breaksNearly every regression of one trending series on another.
For exampleA deck shows a company's revenue rising with the number of its app downloads and concludes downloads drive revenue. Both rose with the market. On year-to-year changes the relationship disappears, and the growth belongs to the category.
The mathematics in fully(t) = a + b x(t) + e(t) yields a significant b for any two trending series; the test is whether b survives differencing or whether the two share a common stochastic trend.
Taught inMATH 386-1; MMSS 1996 (lagged regression)
SourcesGranger, Newbold (1974). Spurious regressions in econometrics. Journal of Econometrics 2(2), 111-120. https://doi.org/10.1016/0304-4076(74)90034-7
When companies, regions, or customers differ in ways nobody can measure, we compare each one to itself over time rather than to each other, which removes everything about the unit that does not change.
y(i,t) = a(i) + b x(i,t) + e(i,t)Panel data, the same units observed repeatedly over time, let an analyst compare each unit to itself and remove everything about it that does not change. Stephen Nickell's 1981 paper is the standard warning about the method: with a lagged outcome and few periods, the fixed-effects estimate is biased in a known direction. The 1996 MMSS curriculum taught panel studies and lagged regression as a unit.
Why it works nowCompanies have always had panel data on their own plants, regions, and customers and rarely analyzed it as such, because organizing it took a project. It can now be organized in an afternoon, and Advisors can say which permanent differences between units would otherwise contaminate a cross-section comparison. The Desk asks whether the same units are observed before and after, or different ones.
When it appliesA cross-section comparison where the units differ in unmeasured ways; repeated observations of the same customers, plants, or firms; a claim that a change worked because the changed units are now better than the unchanged ones.
What we would askDo you observe the same units before and after, or different units?
What about each unit never changes and could explain the difference?
Repeated observations of the same units over time, with the change occurring at different times for different units.
How we would get itThe client's own panel data; Advisors on what fixed differences between units matter.
Where it breaksFixed effects remove permanent differences, not changing ones; with short panels and lagged outcomes the estimates are biased.
For examplePlants that adopted a new maintenance system have less downtime than plants that did not. The adopting plants were also newer. Comparing each plant to its own downtime before and after adoption, the effect is a third of the cross-section estimate.
The mathematics in fully(i,t) = a(i) + b x(i,t) + e(i,t): the unit effect a(i) absorbs everything permanent about unit i, so b is estimated only from changes within units; with a lagged outcome on the right and few periods, b is biased.
Taught inMATH 386-1; MMSS 1996 (panel studies)
SourcesNickell (1981). Biases in Dynamic Models with Fixed Effects. Econometrica 49(6), 1417. https://doi.org/10.2307/1911408
Wooldridge, Jeffrey M. Introductory Econometrics: A Modern Approach. Cengage. (book)
A model that forecasts well can be a black box, and a model you will act on must identify a cause. We ask which of the two the decision needs, because the same data answers them differently and the wrong one is confidently wrong.
Out-of-sample error, or an identifying assumptionLeo Breiman's 2001 paper on the two cultures of statistical modeling drew the line between models that fit a mechanism and models that predict without one, and argued that prediction on held-out data was the honest test of the second. Stanford's data science core teaches both cultures, and Athey and Imbens's causal trees are the modern attempt to keep the predictive machinery while identifying a cause.
Why it works nowEvery company now has a model that predicts something, and every executive wants to act on its top feature, which is the mistake the distinction exists to prevent. What has changed is that both questions can be answered, separately and quickly, once someone asks which one the decision needs. The Desk asks whether the client needs to know which customers will leave or what would make them stay.
When it appliesA machine-learning model presented as a reason to intervene; a churn or default score used to decide what to change; an executive asking why when the model only says what.
What we would askDo you need to know which customers will leave, or what would make them stay?
If you act on this model's most important feature, do you expect the outcome to change?
Clarity about the decision; for prediction, held-out accuracy; for explanation, an identifying assumption.
How we would get itThe Research Director states which question is being answered; a prediction is validated out of sample, an explanation is designed with a counterfactual, and the report does not confuse them.
Where it breaksUsing a predictor to choose an intervention; demanding a causal story from a forecasting tool.
For exampleA lender's model predicts default from three hundred features, and the top feature is the number of times a customer checked their balance. The model is a good predictor and a useless lever: stopping customers from checking their balance would change nothing. What would change default is a different question, with a different design.
The mathematics in fullA predictor minimizes out-of-sample error and is validated by holding data back; an explanation identifies the effect of changing X on Y and is validated by the credibility of its identifying assumption; a feature that predicts an outcome is not thereby a lever on it.
Taught inStanford data science core; MATH 386-2
SourcesBreiman (2001). Statistical Modeling: The Two Cultures. Statistical Science 16(3). https://doi.org/10.1214/ss/1009213726
Athey, Imbens (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences 113(27), 7353-7360. https://doi.org/10.1073/pnas.1510489113
Timing questions, which are usually the ones that decide the outcome.
Whether something will happen and when it will happen are different questions, and the second one usually decides the outcome. We answer it from the history of comparable cases, not from opinion.
h(t | X) = h0(t) exp(X beta)Kaplan and Meier's 1958 paper showed how to estimate a survival curve when many of the patients were still alive at the end of the study, and David Cox's 1972 model let the characteristics of each patient shift the hazard, the chance the event happens next given it has not happened yet. Both were built for clinical survival times, and they transferred whole to customers, loans, and companies.
Why it works nowWhether a borrower refinances, a customer churns, or a chief executive leaves was answered with a probability when the decision turned on a date. Comparable durations can now be assembled from the record, including the cases still running, and Advisors who have watched the clock in that industry can say what shortens it. The Desk asks whether the question is whether or when, and which answer changes the decision. Timing questions get timing machinery.
When it appliesA question about timing: when a customer churns, a company needs capital, a CEO leaves, a competitor enters, a regulator acts, a contract renews, a covenant is breached; or a probability question that is really a timing question in disguise.
What we would askIs the question whether this happens, or when? And which answer changes the decision?
In comparable cases, how long did it take, and what shortened it?
A set of comparable cases with their durations, including the ones in which the event has not yet happened; the characteristics that shift the timing.
How we would get itThe record for comparable durations; Advisors who have seen the clock run in this industry on what accelerates it.
Where it breaksCases still alive at the end of the data are dropped or treated as if the event never comes; the covariates were measured after the process began.
For exampleA credit investor asks whether a borrower will need to refinance. The borrower will; every borrower does. The question is whether it happens inside the fund's horizon, and the comparable cases put the hazard sharply higher in the eighteen months around a maturity wall the deck does not mention.
The mathematics in fullThe hazard h(t) is the chance the event occurs in the next interval given it has not yet; survival S(t) = exp(-H(t)); h(t | X) = h0(t) x exp(X beta) lets the characteristics shift the timing, and the shape of h0 (rising, falling, spiking at dates) is the finding.
Taught inMMSS 1996 (event history and proportional hazards); MATH 386-2
SourcesCox (1972). Regression Models and Life-Tables. Journal of the Royal Statistical Society Series B: Statistical Methodology 34(2), 187-202. https://doi.org/10.1111/j.2517-6161.1972.tb00899.x
Kaplan, Meier (1958). Nonparametric Estimation from Incomplete Observations. Journal of the American Statistical Association 53(282), 457-481. https://doi.org/10.1080/01621459.1958.10501452
We describe the customer base, the loan book, or the pipeline as states and the rates of moving between them, because aggregate figures can look flat while the movements underneath have already turned.
p(t + 1) = p(t) PAndrey Markov's chains date to 1906, and his first published application was to the sequence of vowels and consonants in Pushkin's verse. The idea that a system's next state depends on its current state and a table of transition probabilities became the workhorse of everything from credit ratings to customer lifecycles.
Why it works nowAggregate retention has been the number everyone reported because unit-level state histories were tedious to build. They can now be built from a company's own data in hours, with the states defined by the operators who know what at risk really looks like. The Desk asks how many customers who were healthy a year ago are at risk now, and which transition changed most. The number under the number becomes visible.
When it appliesA stable aggregate (net retention, delinquency rate, pipeline conversion) that the client suspects is hiding something; a base that is aging; a pipeline that is filling at the top and emptying in the middle.
What we would askOf the customers who were healthy a year ago, how many are at risk now?
Which transition has changed most, and when did it change?
Unit-level state histories over several periods; a state definition that the operators agree is real.
How we would get itThe client's data, with the states defined by the operators who know what at risk actually looks like; former operators on which transitions predict the next.
Where it breaksThe process has memory the model ignores; the transitions are estimated from too few moves; the states are defined by the analyst rather than the business.
For exampleA software company reports net retention above one hundred and ten percent for six quarters. The transition from expanded to at-risk has doubled in the last three, and the aggregate is held up by expansion in the oldest cohort. The clock is visible in the transitions and invisible in the number.
The mathematics in fullp(t + 1) = p(t) x P with P(i, j) the probability of moving from state i to state j; the long-run mix solves pi = pi x P, and a small rise in the transition out of the healthy state compounds into a large change in the mix a year out.
Taught inMMSS 1995 (matrix models and stochastic processes); DECS-430
Complaints, incidents, departures, regulatory inquiries, and outages arrive at a natural rate. When a cluster arrives well above it, we quantify how unusual that is before anyone calls it noise or calls it a crisis.
P(N = k) = exp(-lt) (lt)^k / k!Ladislaus Bortkiewicz applied the Poisson distribution in 1898 to deaths of Prussian cavalrymen from horse kicks, corps by corps and year by year, and showed that rare events arriving at a steady rate produce clusters by chance more often than intuition allows. The same arithmetic says when a cluster is too large to be chance.
Why it works nowA run of executive departures or product incidents was read by gut. The base rate for a company of that kind can now be found from the record, and the events themselves examined for a shared cause, which collapses several events into one. The Desk asks what the normal rate is and where that number comes from, and whether the events shared a cause or only a quarter. Four departures under one new executive is one event with a name.
When it appliesA run of adverse events (executive departures, product failures, customer incidents, regulatory inquiries, security breaches) that the client cannot tell from bad luck.
What we would askWhat is the normal rate of this event for a company of this kind, and where does that number come from?
Did the events share a cause, or only a quarter?
The historical rate of the event for this company and its peers; the timing and causes of the recent events.
How we would get itThe record for the base rate; operators and former employees on whether the events are one event or several.
Where it breaksThe window is chosen after the cluster; the events are not independent and the count is inflated; the base rate is assumed rather than found.
For exampleFour vice presidents leave a company in one quarter. At the company's historical rate, four in a quarter is very unlikely by chance; the interviews establish that three of the four report to the same new executive, which makes it one event with a name.
The mathematics in fullWith events arriving at rate lambda, P(N(t) = k) = exp(-lambda t) (lambda t)^k / k!, so k events where lambda t were expected has a computable improbability; if the events share a cause they are not independent and the count must be reduced to the number of causes.
Taught inMMSS 1995 (stochastic processes); DECS-430
When a decision recurs over time, the right move now depends on the moves it leaves open later. We work backward from the last decision to the first, which often reverses what looked obvious.
V(s, t) = max [ payoff + E V(s', t + 1) ]Richard Bellman's dynamic programming, published in 1954, solves a multistage decision by starting from the last stage and working backward, so that each early choice is valued for the choices it leaves open. He later admitted he chose the name partly because his sponsor disliked the word research.
Why it works nowMultiyear programs, capacity plans, and roadmaps have been solved forward, greedily, because working backward required a model of what would be learned at each stage. The stages and what each reveals can now be laid out in a conversation, and operators who ran comparable programs can say where the sequence went wrong. The Desk asks what the last decision in the sequence is and which early choice forecloses the most later ones. The answer often reverses the obvious first step.
When it appliesA multi-year program of investments, a product roadmap, a capacity plan, a phased acquisition; or a client deciding the first step without having thought about the last.
What we would askWhat is the last decision in this sequence, and what does it need to be true when you get there?
Which early choice forecloses the most later ones?
The sequence of decisions, what each one forecloses, and what is learned between them.
How we would get itOperators who have run comparable multi-stage programs on where the sequence went wrong; the record on comparable programs.
Where it breaksThe problem is solved forward, greedily; the value of keeping options open is ignored; the horizon is cut short.
For exampleA manufacturer plans three capacity additions over six years. Solved forward, the first is the largest. Solved backward from the demand uncertainty that resolves after the first, the first is the smallest, and the program is cheaper in every scenario but the best one.
The mathematics in fullV(state, t) = max over actions of [ payoff(action, state) + E[V(next state, t + 1)] ], solved from the final period backward; the early action is chosen for the states it leaves reachable, not for its own payoff.
Taught inMECS 560-2; MMSS 300
SourcesBellman (1954). The theory of dynamic programming. Bulletin of the American Mathematical Society 60(6), 503-515. https://doi.org/10.1090/S0002-9904-1954-09848-8
How groups, boards, regulators, and voters actually decide.
A group of imperfect decision makers can be far better or far worse than its members, depending on whether their errors are independent and on the rule they vote by. We look at how the committee that will decide actually decides.
Majority accuracy rises in n only if errors are independentCondorcet proved in 1785 that a majority of independent members who are each more likely right than wrong becomes more reliable as the group grows, and Francis Galton found the same thing at a Plymouth livestock fair in 1907, where the median of nearly eight hundred guesses of an ox's weight came within one percent of the truth. Austen-Smith and Banks showed in 1996 that the result fails when members vote strategically, and later work showed it fails when their errors are correlated.
Why it works nowPresenting to an investment committee has always meant guessing how it decides. How it decides is now researchable: former members of comparable committees can say whether views form independently or after the same senior voice, and the record shows what the committee has done before. The Desk asks whether the members form their views independently and what rule decides. A committee that hears the same expert before voting is one mind with several votes.
When it appliesAn investment committee, a board, a credit committee, a panel of buyers, or any group whose decision is the outcome; or a client preparing to present to one.
What we would askDo the members form their views independently, or after hearing the same person?
What rule decides: majority, unanimity, or one voice?
The committee's rule, its information flow, and the independence of its members' judgments.
How we would get itFormer members of comparable committees on how the decision is really made; the record of the committee's past decisions.
Where it breaksA committee that hears the same expert before voting is one mind with several votes; unanimity rules are assumed to be conservative when they can be the opposite.
For exampleA client will present to an investment committee that requires unanimity. Its members read the same memo and defer to the same senior partner. The research is not the memo; it is what that partner believes, and what would change it.
The mathematics in fulln independent members each right with probability p above one half are right by majority with probability rising toward one in n; with correlated errors the gain vanishes, and under unanimity a strategic voter may vote against private evidence, so the rule shapes the outcome.
Taught inMECS 540-3; MMSS 211-3
SourcesAusten-Smith, Banks (1996). Information Aggregation, Rationality, and the Condorcet Jury Theorem. American Political Science Review 90(1), 34-45. https://doi.org/10.2307/2082796
Galton (1907). Vox Populi. Nature 75(1949), 450-451. https://doi.org/10.1038/075450a0
In any group that votes, the order in which options come up and the option that stands if nothing passes shape the result as much as the votes do. We find out who controls the agenda and what the default is.
Majority rule selects the median ideal pointDuncan Black showed in 1948 that under majority rule with single-peaked preferences the median voter's preference wins, and Anthony Downs built a theory of party competition on it in 1957. Romer and Rosenthal's 1978 agenda-setter model, which they went on to test against Oregon school budget referenda, showed the other half: whoever sets the agenda, and whatever happens if the proposal fails, can move the outcome a long way from the median. Kenneth Arrow had proved in 1950 that no voting rule escapes all such trouble.
Why it works nowWhy the obvious option keeps losing in a board, a rulemaking, or a standards body used to be a mystery attributed to politics. The agenda setter, the default, and the members' rankings can now be established from former participants and from the record of past votes. The Desk asks what happens if nothing passes and who benefits, and who decides what comes to a vote and in what order.
When it appliesA board vote, a shareholder proposal, a regulatory rulemaking, a standards committee, a legislature; or a client who cannot understand why the obvious option keeps losing.
What we would askWhat happens if nothing passes, and who benefits from that?
Who decides which options are put to a vote, and in what order?
The decision rule, the status quo, the agenda setter, and the members' rankings of the options.
How we would get itFormer participants in the body on how agendas are really set; the record of past votes and the defaults that prevailed.
Where it breaksThe aggregation rule is chosen by whoever wants the result; preferences are assumed single-peaked when they are not.
For exampleA shareholder proposal that most holders support keeps failing. The board controls the agenda and pairs it each year with an alternative that splits its supporters. The research is the pairing, not the support.
The mathematics in fullUnder majority rule with single-peaked preferences the outcome is the median member's ideal point; an agenda setter who controls the alternative to the status quo can move the outcome a long way from the median toward its own.
Taught inMMSS 211-3; MECS 540-3
SourcesBlack (1948). On the Rationale of Group Decision-making. Journal of Political Economy 56(1), 23-34. https://doi.org/10.1086/256633
Romer, Rosenthal (1978). Political resource allocation, controlled agendas, and the status quo. Public Choice 33(4), 27-43. https://doi.org/10.1007/bf03187594
Arrow (1950). A Difficulty in the Concept of Social Welfare. Journal of Political Economy 58(4), 328-346. https://doi.org/10.1086/256963
Downs (1957). An Economic Theory of Political Action in a Democracy. Journal of Political Economy 65(2), 135-150. https://doi.org/10.1086/257897
We treat the regulator, the ministry, or the legislature as an actor with its own incentives, audiences, and commitment problems, and we research those rather than reading the press release.
The institution best-responds to the audience that punishes itJames Fearon's 1994 paper on audience costs and his 1995 paper on rationalist explanations for war treated states and leaders as actors with incentives, audiences, and commitment problems rather than as weather. MMSS teaches the same formal political models to undergraduates; Kellogg's political economy sequence teaches them at the doctoral level.
Why it works nowPolitical and regulatory risk has been priced from headlines because the decision makers inside the institutions were unreachable. Former regulators, former officials, and the advisors who have worked the institution can now be recruited like any other Advisor, and the record shows which provisions survive every budget cycle and which are revised in each. The Desk asks who inside the institution decides and what they answer to.
When it appliesA thesis exposed to a regulatory decision, a subsidy, a tariff, a license, an election, or a policy reversal; or a client trying to price political risk from headlines.
What we would askWho inside the institution decides this, and what do they answer to?
What would it cost them to reverse course, and have they paid that cost before?
The decision maker's incentives, the audiences in front of which commitments were made, and the history of reversals.
How we would get itFormer regulators, former officials, and the advisors who have worked the institution; the record of comparable decisions and their timing.
Where it breaksThe institution is treated as a monolith; the reversal cost is assumed from rhetoric; the timing is guessed from the calendar.
For exampleA renewable developer's returns depend on a tariff regime. The former director of the agency that administers it, interviewed, explains which provisions are politically untouchable and which are revised every budget cycle. The thesis is exposed to the second kind.
The mathematics in fullThe institution's choice is a best response given its payoffs and the audiences that punish reversal; commitment problems and asymmetric information explain policy conflicts that would otherwise be settled by bargaining, and both can be researched.
Taught inMMSS 211-3; MECS 540-2; MECS 540-3
SourcesFearon (1994). Domestic Political Audiences and the Escalation of International Disputes. American Political Science Review 88(3), 577-592. https://doi.org/10.2307/2944796
Fearon (1995). Rationalist explanations for war. International Organization 49(3), 379-414. https://doi.org/10.1017/S0020818300033324
Shareholder votes, board contests, activist campaigns, and creditor committees are decided by rules, coalitions, and the strategic misreporting of preferences. We model who votes, under what rule, and who can be brought across.
No non-dictatorial rule is immune to manipulationKenneth Arrow's 1950 impossibility theorem showed that no rule for combining individual rankings into a group ranking satisfies a short list of reasonable conditions at once, and Allan Gibbard proved in 1973 that any non-dictatorial rule over three or more options can be manipulated by strategic voting. Shareholder votes, board contests, and creditor committees are the corporate versions.
Why it works nowContested votes used to be counted by stated positions. The swing holder's real incentives, and what it took to move that holder in comparable situations before, can now be found from former activists, proxy advisors, and governance specialists, and from the record of past votes. The Desk asks which coalition wins under this rule and who the swing is. The research is that holder's price, not the headcount.
When it appliesAn activist situation, a contested merger vote, a creditor committee, a founder-control structure, or a governance change that the thesis depends on.
What we would askUnder this rule, which coalition wins, and who is the swing?
Who might vote against their stated position, and why?
The rule, the holders and their preferences, the coalitions available, and the incentives to misreport.
How we would get itFormer activists, proxy advisors, and governance specialists on how comparable votes were won; the record of the holders and their past votes.
Where it breaksStated positions are taken as votes; the swing holder's incentives are not researched; the rule is misread.
For exampleA merger requires a majority of the minority. The largest minority holder has stated opposition and has also been in a comparable situation twice, both times settling for a small price bump before the vote. The research is that holder's price, not the vote count.
The mathematics in fullAny non-dictatorial rule over three or more options can be manipulated by strategic voting; the outcome is a coalition outcome under the rule, and the swing member's incentives, not the headcount, decide it.
Taught inMECS 540-3; MMSS 211-3
SourcesGibbard (1973). Manipulation of Voting Schemes: A General Result. Econometrica 41(4), 587. https://doi.org/10.2307/1914083
Arrow (1950). A Difficulty in the Concept of Social Welfare. Journal of Political Economy 58(4), 328-346. https://doi.org/10.1086/256963
When a group shares an interest and each member gains most by letting the others do the work, the thing does not get done. We ask who bears the cost, who captures the benefit, and what would make contributing individually worthwhile.
Contribute only when the share B/n exceeds the cost cPaul Samuelson wrote down the mathematics of a good that everyone consumes and nobody can be excluded from in 1954, and showed why the market underprovides it. Mancur Olson's 1965 book explained why large groups with a shared interest fail to act on it while small groups with concentrated stakes do, and Garrett Hardin's 1968 essay on the commons made the overuse of a shared resource the most cited example in the field.
Why it works nowWhether a consortium, a standard, or a shared piece of infrastructure will actually get funded has always been a matter of reading personalities. The distribution of stakes can now be mapped from the record, and former participants in comparable efforts can say why they held or failed. The Desk asks who pays for this and what they get that the free riders do not, and whether a small group has enough at stake to carry it alone.
When it appliesA consortium, standards body, industry association, shared infrastructure, open-source dependency, or joint venture that everyone wants and nobody funds; a common resource being run down.
What we would askWho pays for this, and what do they get that the free riders do not?
Is there a small group whose stake is large enough to carry it alone?
The distribution of stakes across members, the cost of contributing, and whether contribution can be observed and rewarded.
How we would get itFormer participants in comparable consortia and shared efforts on why they held or failed; the record of who funded what.
Where it breaksAssuming shared interest produces shared action; ignoring that small groups with concentrated stakes act and large groups with diffuse ones do not.
For exampleA thesis depends on an industry standard being adopted. Every firm wants the standard and every firm gains most by waiting for the others to bear the switching cost. The three largest have enough at stake to move alone, and the research is whether they will.
The mathematics in fullEach member contributes c and the group gains B shared by n; a member contributes only when its share B/n exceeds c, so large groups with diffuse benefits underprovide, and a common resource is overused because each user bears 1/n of the damage it causes.
Taught inMMSS 300; MMSS 211-1; MMSS 211-2
SourcesHardin (1968). The Tragedy of the Commons. Science 162(3859), 1243-1248. https://doi.org/10.1126/science.162.3859.1243
Samuelson (1954). The Pure Theory of Public Expenditure. The Review of Economics and Statistics 36(4), 387. https://doi.org/10.2307/1925895
Olson, Mancur (1965). The Logic of Collective Action. Harvard University Press. (book)
Who knows, who is central, and who occupies the same position in a neighboring world.
Instead of asking who had the highest title, we ask who occupied the position through which the relevant information passed, and we put that person on the panel.
Ax = lambda x: central when tied to the centralLinton Freeman's 1978 paper sorted out what centrality means, distinguishing the person with the most ties from the person who sits on the most paths between others. Mark Granovetter had found in 1973 that Boston job seekers heard about their jobs more often from acquaintances than from close friends, because weak ties reach into other circles. Network analysis was a named part of the MMSS curriculum in the 1990s.
Why it works nowExpert search has been keyword search on titles, which finds the people who signed off and misses the person who ran the model and briefed all of them. The map of who talked to whom about a decision can now be built from the first two interviews and the record, and the panel recruited from its center. The Desk asks whose desk the decision crossed and who everyone called.
When it appliesAn expert search that has produced impressive titles and thin knowledge; a question about how a decision was really made inside an organization or an industry.
What we would askWhen this was decided, whose desk did it cross, and who did everyone call?
Who is connected to the most people who would know?
A map of who talked to whom about the question, built from the first interviews and from the record.
How we would get itThe Research Director builds the map from the first two interviews and the record, then recruits from its center rather than from its org chart.
Where it breaksThe network is drawn from one starting point and finds only that point's contacts; centrality is confused with seniority.
For exampleA client wants to understand a pricing decision at a supplier. The senior executives who signed it off know the outcome. The mid-level analyst who ran the model and briefed all of them knows the reasoning, and is the interview.
The mathematics in fullDegree centrality counts a person's ties; eigenvector centrality solves A x = lambda x so a person is central when tied to central people; the boundary of the network drawn decides whose centrality is being measured.
Taught inMMSS 1995 (network analysis and centrality)
SourcesFreeman (1978). Centrality in social networks conceptual clarification. Social Networks 1(3), 215-239. https://doi.org/10.1016/0378-8733(78)90021-7
Bonacich (1987). Power and Centrality: A Family of Measures. American Journal of Sociology 92(5), 1170-1182. https://doi.org/10.1086/228631
Granovetter (1973). The Strength of Weak Ties. American Journal of Sociology 78(6), 1360-1380. https://doi.org/10.1086/225469
When a market is too new or too small to have experts, we look for people who occupy the same structural position in a neighboring industry and have solved the same problem in different clothes.
Distance d(i, j) is small when tie patterns matchLorrain and White's 1971 paper defined structural equivalence: two people occupy the same position in a network when they have the same pattern of ties to everyone else, whether or not they know each other. The idea let sociologists find roles rather than names, and it is the formal version of noticing that an exchange operator and a cloud capacity allocator are solving the same problem.
Why it works nowA first-of-kind market has no experts, and the search used to stop there. The question can now be restated as a structure, a liquidity manager, an allocator of a scarce shared resource, a two-sided market operator, and the industries where that structure is mature can be searched for the people who have solved it. The Desk asks what problem the business actually solves, stripped of its industry vocabulary, and who has solved that exact problem somewhere else.
When it appliesA first-of-kind market, technology, or business model where the usual expert search returns nobody with direct experience.
What we would askWhat problem does this business actually solve, stripped of its industry vocabulary?
Who has solved that exact problem somewhere else?
A structural description of the role or problem, independent of the industry, and a search across industries for the same structure.
How we would get itThe Research Director restates the question as a structure (a liquidity manager, an allocator of a scarce shared resource, a two-sided market operator) and recruits from the industries where that structure is mature.
Where it breaksThe analogy is superficial; the structural match is asserted rather than checked against the mechanics.
For exampleA client is underwriting a marketplace for a resource that has never been traded. There are no experts on that marketplace. There are exchange operators, transportation network planners, and cloud capacity allocators who have run the same coordination problem, and two of them are the panel.
The mathematics in fullTwo actors are structurally equivalent when their patterns of ties match, d(i, j) = sqrt of sum over k of (A(i, k) - A(j, k))^2 small, whether or not they are connected to each other; the search is for small d across industries, not for shared keywords.
Taught inMMSS 1995 (structural equivalence and network roles)
SourcesLorrain, White (1971). Structural equivalence of individuals in social networks. The Journal of Mathematical Sociology 1(1), 49-80. https://doi.org/10.1080/0022250X.1971.9989788
When nobody saw the whole thing, we interview several people who each saw part, and we recover the fact from the pattern of their agreement rather than from any one account.
Agreement between sources estimates each one's competenceThe cultural consensus model of Romney, Weller, and Batchelder in 1986 solved the anthropologist's problem of informants who each know part of the culture and none of whom can be checked against an answer key: the pattern of agreement among them estimates each one's competence and recovers the consensus. It is, in effect, a method for reconstructing a fact from several partial accounts.
Why it works nowFinding out what really happened inside a company or a deal has always meant interviewing several people who each saw a piece, and then trusting the most vivid one. Parallel structured interviews that ask each about the same facts can now be run and compared systematically, so the overlap scores the accounts and the contested facts are weighted by it. The Desk asks who saw which part and where the accounts overlap.
When it appliesA question about what really happened inside a company, a deal, or a negotiation, where every available source saw only a piece and some have reasons to shade it.
What we would askWho saw which part of this, and where do their accounts overlap?
On the parts that overlap, do they agree?
Several sources with different vantage points and, ideally, no coordination; a structured interview that asks each about the same facts.
How we would get itParallel structured interviews; the synthesis scores each source by agreement with the others on the overlapping facts and weights the contested facts accordingly.
Where it breaksThe sources have coordinated; the overlap is too small to estimate agreement; one source dominates the synthesis by being vivid.
For exampleA client wants to know why a competitor abandoned a product line. The former product head, the former CFO, and a former channel partner each saw part of it. On the facts all three saw, two agree and one does not, and the reconstruction weights the contested facts toward the two.
The mathematics in fullWith independent informants, the agreement between each pair estimates each informant's competence, and the consensus answer to each question is recovered by weighting each informant by that competence, without knowing the answers in advance.
Taught inMMSS 1995 (informant accuracy)
SourcesRomney, Weller, Batchelder (1986). Culture as Consensus: A Theory of Culture and Informant Accuracy. American Anthropologist 88(2), 313-338. https://doi.org/10.1525/aa.1986.88.2.02a00020
Interaction between two parties scales with their size and falls with the distance between them, and distance can be commercial, institutional, or cultural as well as geographic. We use it to predict partnerships, encroachment, and where a customer will migrate.
F = G x M(i)^a x M(j)^b / D^cJan Tinbergen borrowed Newton's gravity in 1962 to predict trade between countries from their sizes and the distance between them, and it fit remarkably well. Anderson and van Wincoop's 2003 paper resolved the border puzzle, why Canadian provinces traded so much more with each other than with equally close American states, by showing that relative trade costs, not miles, are the distance that matters. Gravity models were in the MMSS curriculum of the mid-1990s.
Why it works nowThe insight for an investor is that distance can be regulatory, cultural, or institutional, and those distances were never measured because nobody could. Operators can now say which dimension of distance actually binds in a market, and the record of past interactions can be fitted to find out. The Desk asks what the real distance is between two parties, and who is large and close that the plan ignored.
When it appliesA question about which buyers a seller will reach, which partners will form, which competitor will encroach on which territory, or where an expansion will succeed.
What we would askBetween these two, what is the real distance: geography, regulation, relationships, language, or systems?
Who is large and close, and therefore likely, that the plan has ignored?
Sizes of the parties and a defensible measure of the distances between them on the dimensions that matter here.
How we would get itOperators on which dimensions of distance actually bind in this market; the record of past interactions for the fit.
Where it breaksDistance is measured on the dimension that is easy to measure rather than the one that binds; the fit is extrapolated outside the data.
For exampleA distributor plans expansion into an adjacent region on the strength of proximity. The binding distance is not miles but a certification regime that its current customers do not require and the new ones do; the model, once the right distance is used, points at a different region.
The mathematics in fullF(i, j) = G x M(i)^a x M(j)^b / D(i, j)^c, with D a weighted combination of geographic, commercial, institutional, and network distance; the weights are the research question.
Taught inMMSS 1995 (gravity models)
SourcesAnderson, van Wincoop (2003). Gravity with Gravitas: A Solution to the Border Puzzle. American Economic Review 93(1), 170-192. https://doi.org/10.1257/000282803321455214
Where budget, capacity, slots, and attention should go, often with an exact answer.
Given a budget, a deadline, and a cap on interviews, we choose the set of studies that buys the most decision value inside the constraints, and we tell you which constraint is costing you the most.
max sum of EVSI(q) x(q) subject to cost and timeLeonid Kantorovich invented linear programming in 1939 to plan production at a Leningrad plywood trust, and his 1960 paper in Management Science brought the work to the West; George Dantzig's simplex method made it computable for the United States Air Force after the war. The shadow price, what one more unit of a constraint is worth, was there from the start.
Why it works nowA research engagement is an allocation of a budget and a deadline across candidate studies, and it has never been solved as one. It can be: each study's value to the decision, cost, and time can be laid out and the plan solved, with the shadow price of the deadline reported. The Desk asks what the budget and the deadline really are, and what one more week would be spent on. The report names the workstreams left out and what they would have cost.
When it appliesAn engagement, or a client's own diligence, with more candidate questions than time or money; a client asking whether to enlarge the budget.
What we would askWhat is the deadline and what is the budget, really?
If we could add one more week or one more interview, what would you want it spent on?
For each candidate study: its value to the decision, its cost, and its time; the binding constraints.
How we would get itThe research plan is solved as an allocation; the shadow price of the deadline and the budget is reported, which is the argument for or against enlarging the engagement.
Where it breaksThe objective is information rather than decision value; a constraint is invented to force the plan.
For exampleSix workstreams, a fixed fee, and seventy-two hours. Four fit. The two left out are named in the report with what they would have cost and what they might have changed, so the client can buy them if the decision warrants.
The mathematics in fullChoose x(q) in {0, 1} to maximize the sum of EVSI(q) x x(q) subject to the sum of cost(q) x x(q) at most the budget and the sum of time(q) x x(q) at most the deadline; the shadow price of each constraint is the gain from relaxing it by one unit.
Taught inMECS 560-1; MECN-451
SourcesKantorovich (1960). Mathematical Methods of Organizing and Planning Production. Management Science 6(4), 366-422. https://doi.org/10.1287/mnsc.6.4.366
Where should the capital, the capacity, the inventory, or the people go? Many such questions have exact answers once the objective and the constraints are written down honestly, and writing them down is most of the work.
At the optimum, gradients of the binding constraints combineThe same Kantorovich and Dantzig work made allocation a solved problem once the objective and constraints were written down, and MECS teaches the Karush-Kuhn-Tucker conditions that generalize it. The lesson that survived seventy years is that writing the constraints down honestly, including the political ones, is most of the work.
Why it works nowCapital has been allocated across divisions by last year's share because the true returns and the binding constraints were never assembled in one place. They can be assembled now, from the client's data and from operators who know which constraints are real and which are habits, and the formal solution can be run as a benchmark for the judgment. The Desk asks what is being maximized, which constraint binds today, and what one more unit of it would be worth.
When it appliesA capital allocation across plants, products, regions, or funds; a capacity decision; an inventory or network design question dressed up as a judgment call.
What we would askWhat are you maximizing, and what are the constraints you cannot relax?
Which constraint is binding today, and what would one more unit of it be worth?
The objective stated honestly; the constraints, including the political ones; the data for the coefficients.
How we would get itOperators on the real constraints and on which are negotiable; the record for the coefficients; the formal solution as a benchmark for the judgment.
Where it breaksThe objective is mis-stated; a constraint is political and unwritten; the solution is applied without the judgment it was meant to inform.
For exampleA company allocates capital to its divisions by last year's share. Written as an allocation with the true returns and the covenant constraints, one division should receive nothing and another twice its share; the shadow price on the covenant is the argument for refinancing.
The mathematics in fullMaximize f(x) subject to g(i)(x) at most zero and h(j)(x) equal to zero; at the optimum the gradient of f is a nonnegative combination of the gradients of the binding constraints, and each multiplier is the value of relaxing its constraint by one unit.
Taught inMECS 560-1; MATH 285
SourcesKantorovich (1960). Mathematical Methods of Organizing and Planning Production. Management Science 6(4), 366-422. https://doi.org/10.1287/mnsc.6.4.366
Where the good cannot be priced, who gets which slot, seat, organ, school, or partner is decided by matching rules, and the rules decide who can game the market. We read the rules the way a market designer would.
Stable when no pair would both rather have each otherGale and Shapley's 1962 paper on college admissions and stable marriage gave the deferred acceptance algorithm, which always finds a matching no pair would both want to abandon, and showed that the side that proposes gets its best stable outcome. Alvin Roth and Elliott Peranson used it in 1999 to redesign the match that assigns American medical graduates to residencies, and the design has run every year since.
Why it works nowMarketplace and staffing businesses run matching rules that determine who is happy and who churns, and investors have read the churn as a marketing problem. The rules can now be read as a market designer would, the outcomes measured from the platform's data, and operators of comparable markets asked where the rules were gamed. The Desk asks which side proposes and which side has to accept what it is offered.
When it appliesA marketplace, staffing, or platform business whose economics depend on how it matches the two sides; a hiring or admissions process; an allocation with no price.
What we would askWhich side proposes, and which side has to accept what it is offered?
Can the participants improve their outcome by misreporting what they want?
The matching rule, the preferences on both sides, and the complementarities that break stability.
How we would get itOperators of comparable matching markets on where the rules were gamed; the record of the platform's matching outcomes.
Where it breaksParticipants misreport strategically; couples and complementarities make no stable match exist; the design favors one side without saying so.
For exampleA staffing platform's thesis assumes both sides are happy with its matches. Its rule lets clients propose and forces workers to accept the best offer so far. That is the client-optimal stable match, and the workers' churn, which the thesis calls a marketing problem, is the design.
The mathematics in fullA matching is stable when no pair would both prefer each other to their assigned partners; deferred acceptance produces a stable match that is optimal for the proposing side and worst for the other, so the choice of proposer is a distributional decision.
Taught inMECS 560-1; MECS 470
SourcesGale, Shapley (1962). College Admissions and the Stability of Marriage. The American Mathematical Monthly 69(1), 9-15. https://doi.org/10.2307/2312726
Roth, Peranson (1999). The Redesign of the Matching Market for American Physicians: Some Engineering Aspects of Economic Design. American Economic Review 89(4), 748-780. https://doi.org/10.1257/aer.89.4.748
An investment is not evaluated alone. We ask what it adds to what the client already holds, which depends on how it moves with the rest, especially in the bad states.
max w'mu - (lambda / 2) w'Sigma wHarry Markowitz's 1952 paper showed that what an asset adds to a portfolio depends on how it moves with the rest, not on its own risk, and that a set of positions can be chosen to give the most return for a given variance. It was the beginning of modern portfolio theory and of his Nobel.
Why it works nowThe theory has always been applied to listed securities with long price histories and almost never to a private position, a lender, or a portfolio company, whose co-movements with the rest of a book are not in any database. They can now be researched: Advisors can say which exposures co-move in that sector and how comparable positions behaved in past stresses, and the client's book can be captured at scoping. The Desk asks what the client already holds that goes wrong in the same states this does.
When it appliesA position that looks attractive standalone but adds sector, rate, regulatory, or liquidity concentration; a client evaluating the deal rather than the portfolio it enters.
What we would askWhat do you already hold that goes wrong in the same states this does?
In the last stress, how did the closest thing you own to this behave?
The client's existing exposures and how this position covaries with them, in stressed periods as well as calm ones.
How we would get itAdvisors on the exposures that co-move in this sector; the record on how comparable positions behaved in past stresses; the client's own book captured at scoping.
Where it breaksCorrelations estimated in calm periods and assumed in stressed ones; the position judged on its own merits alone.
For exampleA fund likes a specialty lender. Its book already holds two positions that draw on the same wholesale funding market. On its own the lender is fine; in the portfolio it triples an exposure that no one had named, and the research question becomes that funding market.
The mathematics in fullThe contribution of an asset to portfolio risk depends on its covariance with the portfolio, not on its own variance; maximize w'mu - (lambda / 2) x w'Sigma w with the weights summing to one, and re-estimate Sigma in stressed periods.
Taught inMECN-451; DECS-430
SourcesMarkowitz (1952). Portfolio Selection. The Journal of Finance 7(1), 77-91. https://doi.org/10.1111/j.1540-6261.1952.tb01525.x
How things move over time, how much of a bet to take, and the pure mathematics that decides whether a plan survives the movement.
When supply responds to price with a lag, because a ship, a fab, or a plant takes years to build, the market cycles on its own. We ask how long the lag is and where in the cycle the decision sits.
q(t) = S(p(t - L)), supply on a lagged priceMordecai Ezekiel's 1938 paper named the cobweb theorem after the shape the price and quantity trace on a diagram when supply responds to last season's price: farmers plant when prices are high, the crop arrives, prices fall, they plant less, prices rise. The same mathematics governs any market where capacity takes years to build, and difference equations were part of the MMSS curriculum in the 1990s.
Why it works nowThe order book of capacity under construction, the single number that tells you where the cycle is, has always existed and rarely been read against demand. It can now be assembled from filings, trade press, and operators quickly, and the history of past cycles retrieved for the turning points. The Desk asks how long it takes for capacity to arrive and how much is on order right now.
When it appliesShipping, semiconductors, chemicals, agriculture, real estate, mining, or any business where capacity takes years to add and cannot be removed quickly; a thesis that extrapolates today's price.
What we would askHow long from the decision to add capacity until it arrives, and how much is on order right now?
Where are we in the cycle, and what did the last peak look like?
The order book of capacity under construction, the lag, and the history of past cycles.
How we would get itOperators on the real lead times and the capacity in the pipeline; the record of past cycles and their turning points.
Where it breaksTreating today's price as a level rather than a point on a cycle; ignoring the capacity already ordered.
For exampleFreight rates are at a decade high and a thesis buys shipping equity. The order book for new vessels is at a decade high too, and they arrive in three years. The model says the rates fall as the ships arrive, as they did last time, and the question is only how far.
The mathematics in fullSupply at t responds to the price at t minus the lag: q(t) = S(p(t - L)), while demand clears at the current price; with a lag the system spirals around the equilibrium, converging when supply is less responsive than demand and diverging when it is more.
Taught inMMSS 1995 (difference equations); MMSS 211-1
SourcesEzekiel (1938). The Cobweb Theorem. The Quarterly Journal of Economics 52(2), 255. https://doi.org/10.2307/1881734
Every sector buys from others to make what it sells, so a shock to one sector travels through everyone downstream. We trace the shock through the purchase matrix to find where it lands hardest and first.
x = (I - A)^-1 dWassily Leontief built the first input-output table of the American economy in 1936, a matrix of what every sector buys from every other, and showed that a change in demand for one product could be traced through the purchases to its total effect on every sector. He won a Nobel for it, and the matrix models in the 1995 MMSS curriculum descend from it.
Why it works nowThe Desk was asked, in a real conversation, what breaks first in a supply chain under a sustained tariff shock, and the honest answer is that the question has a structure. The purchase relationships can now be assembled for the relevant sectors and the shock traced through them, and procurement leads can say where substitutes exist and where they do not. The Desk asks which input comes through a sector that feeds many others and has no substitute.
When it appliesA tariff, sanction, outage, price spike, or bankruptcy in one part of a supply chain; a thesis about which industries a shock helps or hurts; a company whose inputs come through a sector under stress.
What we would askWhich of your inputs comes through a sector that feeds many others and has few substitutes?
When the shock hits that sector, how many steps away are you, and how fast does it reach you?
The purchase relationships among the relevant sectors, with substitution possibilities where they exist.
How we would get itOperators and procurement leads on the actual sourcing chain; the record for input-output relationships; the analysis traced through the matrix.
Where it breaksFixed coefficients when firms substitute; national tables for a global chain; the second-order effects ignored.
For exampleA tariff is imposed on one category of industrial input. Traced through the purchase matrix, the sectors hit hardest are not the direct importers but two steps down, in a small sector with no substitute that feeds four larger ones. That sector is the crux, and its operators are the panel.
The mathematics in fullOutput x solves x = A x + d, so x = (I - A)^-1 d; a shock to sector j propagates down column j of the inverse, and the sectors with large entries in that column absorb the shock first and hardest.
Taught inMMSS 1995 (matrix models); MATH 285
SourcesLeontief (1936). Quantitative Input and Output Relations in the Economic Systems of the United States. The Review of Economics and Statistics 18(3), 105. https://doi.org/10.2307/1927837
Given an edge, there is a size that maximizes long-run growth, and betting beyond it reduces growth while raising the chance of ruin. We work out which side of that line the proposed size sits on, and how sure the edge really is.
f* = p - (1 - p) / bJohn Kelly at Bell Labs worked out in 1956 how a gambler with inside information on a noisy channel should size bets to maximize the growth of capital, and Edward Thorp used the result at the blackjack table and then on Wall Street; his 1969 paper is the practitioner's account. Robert Merton's 1969 continuous-time portfolio problem gives the same fraction, excess return over variance, for an investor.
Why it works nowPosition sizing has been set by conviction, and conviction runs highest right after a winning year, which is when the edge is most overestimated. The arithmetic of a track record, its edge, variance, and correlation with the rest of the book, can now be assembled quickly, and the growth-optimal fraction compared with the proposed size. The Desk asks how the confidence in the edge was measured and how much of the capital is gone if it loses.
When it appliesPosition sizing, allocation to a strategy, a capital commitment sized by conviction, a founder betting the company; any decision where the size of the bet is being set by enthusiasm.
What we would askHow confident are you in the edge, and how was that confidence measured?
If this loses, how much of the capital is gone, and is that recoverable?
An honest estimate of the edge and its uncertainty, the variance of the outcome, and the correlation with other bets.
How we would get itThe record and the Advisors on the true distribution of outcomes for bets of this kind; the client's own book at scoping for the correlations.
Where it breaksThe edge estimated on too few cases; correlated bets treated as independent; a horizon too short for the long-run argument.
For exampleA fund with a documented edge in a strategy proposes to double its allocation after a strong year. Sized by the arithmetic of its own track record and variance, the current allocation is already near the growth-optimal fraction, and doubling it lowers expected growth while doubling the chance of a drawdown the fund cannot survive.
The mathematics in fullf* = p - (1 - p) / b for a bet at odds b, or (mu - r) / sigma^2 in the continuous case; growth g(f) is maximized at f* and turns negative beyond about twice it, and because the edge is estimated with error, fractional Kelly is the working rule.
Taught inMATH 385; MECN-451
SourcesKelly (1956). A New Interpretation of Information Rate. Bell System Technical Journal 35(4), 917-926. https://doi.org/10.1002/j.1538-7305.1956.tb03809.x
Thorp (1969). Optimal Gambling Systems for Favorable Games. Revue de l'Institut International de Statistique / Review of the International Statistical Institute 37(3), 273. https://doi.org/10.2307/1402118
Merton (1969). Lifetime Portfolio Selection under Uncertainty: The Continuous-Time Case. The Review of Economics and Statistics 51(3), 247. https://doi.org/10.2307/1926560
Anything that spreads by contact spreads only if each carrier passes it to more than one other on average. We estimate that number and the size of the pool it can spread through, because below the threshold the early numbers mean nothing.
R0 = b / g; spread needs R0 x S above oneKermack and McKendrick's 1927 model divided a population into the susceptible, the infected, and the recovered and derived the threshold: an epidemic takes off only when each case produces more than one new case. The same equations describe the spread of products, practices, and defaults, and Frank Bass's 1969 diffusion model is a cousin with the imitation term made explicit.
Why it works nowThe pass-along rate, the one number that decides whether a thing spreads, used to be a marketing department's hope. It can now be measured from a company's own data, outside the enthusiast cohort, and compared with the threshold; the susceptible pool can be sized from the record. The Desk asks how many others each adopter brings in, measured rather than hoped.
When it appliesA product, practice, default, or rumor spreading through a population; a viral growth thesis; a credit book where one default triggers others; an adoption that has stalled.
What we would askFor each customer who adopts, how many others do they bring in, measured rather than hoped?
How much of the reachable population is still susceptible?
The measured pass-along rate, the recovery or churn rate, and the size and structure of the susceptible pool.
How we would get itThe client's own data for the pass-along rate; operators on comparable spreads; the record for how comparable contagions ended.
Where it breaksUniform mixing assumed when the network has hubs; the susceptible pool overestimated; the pass-along rate measured in the enthusiast segment.
For exampleA consumer app reports that each user invites two others, and a thesis extrapolates. Measured outside the launch cohort, each user brings in nought point eight others, below the threshold, and the growth is paid acquisition wearing a viral costume.
The mathematics in fulldI/dt = b S I - g I with R0 = b / g; spread takes off only when R0 times the susceptible share exceeds one, and the final size is a fixed point strictly less than everyone.
Taught inMS&E 221 (Stanford); MMSS 1995 (exponential growth); MECN-451
SourcesKermack, McKendrick (1927). A contribution to the mathematical theory of epidemics. Proceedings of the Royal Society of London. Series A, Containing Papers of a Mathematical and Physical Character 115(772), 700-721. https://doi.org/10.1098/rspa.1927.0118
Bass (1969). A New Product Growth Model for Consumer Durables. Management Science 15(5), 215-227. https://doi.org/10.1287/mnsc.15.5.215
When each person's choice depends on how many others have chosen, a market can sit still for years and then flip in months, and once flipped it may not flip back. We find the threshold and how close the market is to it.
Equilibrium is a fixed point of the threshold distributionThomas Schelling's 1971 model of segregation showed that mild individual preferences about neighbors can tip a whole neighborhood, and Mark Granovetter's 1978 threshold model generalized it: each person acts when enough others have, and the outcome depends on the whole distribution of thresholds. Brian Arthur's 1989 paper on competing technologies added increasing returns and showed that history, not merit, can decide which technology locks in.
Why it works nowPlatform and standards contests have been called from the quality of the products, which the models say is often not what decides them. The adoption thresholds of the participants can now be researched from the customers themselves, and the current adoption level measured against them, so that the distance to the tip is an estimate rather than a feeling. The Desk asks how many others need to have adopted before the marginal customer will.
When it appliesA platform, standard, technology, or network good where value rises with adoption; a market that seems stuck; a thesis about a challenger displacing an entrenched incumbent.
What we would askHow many others need to have adopted before the marginal customer will?
Once the market has tipped, what would it take to tip it back?
The distribution of adoption thresholds across the population, the current adoption level, and the strength of the increasing returns.
How we would get itCustomers on the adoption level at which they would switch; operators on comparable tips; the record for the adoption curve and its inflection.
Where it breaksAssuming a superior product wins when the incumbent has already locked in; assuming the tip is reversible.
For exampleA challenger payment network has a better product and a fifth of the merchants. Merchants adopt when most of their customers use it and customers when most merchants take it. The threshold model says the market tips only if the challenger reaches about forty percent of one side, and the research is whether any path gets it there.
The mathematics in fullEach actor adopts when the adopting share exceeds its personal threshold; the equilibrium share is a fixed point of the cumulative threshold distribution, and with increasing returns two stable equilibria exist, so the outcome depends on history, not on which technology is better.
Taught inMMSS 311-2; MECN-451
SourcesSchelling (1971). Dynamic models of segregation. The Journal of Mathematical Sociology 1(2), 143-186. https://doi.org/10.1080/0022250X.1971.9989794
Granovetter (1978). Threshold Models of Collective Behavior. American Journal of Sociology 83(6), 1420-1443. https://doi.org/10.1086/226707
Arthur (1989). Competing Technologies, Increasing Returns, and Lock-In by Historical Events. The Economic Journal 99(394), 116. https://doi.org/10.2307/2234208
Decisions are made on a reading of the system that reflects decisions made a delay ago, and people running such systems overshoot in a predictable pattern. We measure the delay and ask whether the plan accounts for it.
d(Stock)/dt = inflow - outflow, set on a delayed readingJay Forrester's 1958 article founded system dynamics with an industrial supply chain in which a small change in retail demand produced wild swings in factory output, entirely because of delays in the ordering rules. John Sterman's 1989 beer distribution game showed that intelligent people playing that supply chain reproduce the oscillation every time, and Lee, Padmanabhan, and Whang named the amplification the bullwhip effect in 1997.
Why it works nowThe delay between a decision and its visible effect is the number that explains most oscillation in inventories, hiring, and capacity, and it has rarely been measured because nobody thought of the rule of thumb as a model. Operators can now be asked for the delays and the rules they manage by, and the client's own order and inventory data reveal the cycle the rule produces. The Desk asks how long it takes for a decision's effect to show up in the numbers the client manages by.
When it appliesInventory swings, hiring that alternates between freezes and sprees, capacity that arrives after demand has turned, a supply chain whose orders swing more than its sales.
What we would askHow long between a decision here and the point when its effect is visible in the numbers you manage by?
When the numbers last turned, did the plan respond to the level or to the change?
The structure of the decision rule, the delay, and the history of the stock and the flows.
How we would get itOperators on the actual delays and the rules of thumb they manage by; the client's own data on orders, inventory, and demand.
Where it breaksTreating a delayed system as immediate; blaming demand for oscillation the decision rule created.
For exampleA distributor's orders to its suppliers swing three times as much as its sales, and the thesis blames volatile customers. The customers are steady. The distributor reorders on a delayed reading of stock, and the swing is the arithmetic of that rule. The fix is the rule, not the customers.
The mathematics in fulld(Stock)/dt = inflow(t) - outflow(t), with the inflow set on a reading of the stock that is D periods old; the delay turns a stabilizing rule into an oscillating one, and the variance of orders amplifies at each stage upstream.
Taught inMMSS 1995 (difference equations); MECN-451
SourcesSterman (1989). Modeling Managerial Behavior: Misperceptions of Feedback in a Dynamic Decision Making Experiment. Management Science 35(3), 321-339. https://doi.org/10.1287/mnsc.35.3.321
Lee, Padmanabhan, Whang (1997). Information Distortion in a Supply Chain: The Bullwhip Effect. Management Science 43(4), 546-558. https://doi.org/10.1287/mnsc.43.4.546
Forrester, Jay W. (1958). Industrial Dynamics: A Major Breakthrough for Decision Makers. Harvard Business Review, 36(4), 37-66. (book)
Sterman, John D. (2000). Business Dynamics: Systems Thinking and Modeling for a Complex World. McGraw-Hill. (book)
Whatever is waiting obeys one identity, the number in the system equals the arrival rate times the time each spends there, and near full capacity waiting does not rise gently but explodes. We check the utilization the plan assumes.
L = lambda x W; wait grows like rho / (1 - rho)John Little proved in 1961 that for any stable queue, whatever its internal rules, the average number waiting equals the arrival rate times the average time each spends in the system. The identity is exact and needs no assumptions about the process, which is why it is the one piece of queueing theory every operations course teaches, and why the cliff near full utilization surprises people who have not seen the curve.
Why it works nowService operations have been praised for high utilization by analysts who never saw the wait-time curve. Arrival and service rates and their variability can now be pulled from a company's own systems, and operators can say how the arrivals actually cluster. The Desk asks at what utilization the operation runs and what happens to waiting when demand rises ten percent.
When it appliesA service business, a hospital, a call center, a fulfillment operation, a software team, or any plan that runs an operation near full capacity; a thesis that credits high utilization as efficiency.
What we would askAt what utilization does this operation run, and what happens to waiting time when demand rises ten percent?
How variable are the arrivals?
Arrival rates, service rates, their variability, and the utilization at which the operation runs.
How we would get itOperators on the actual utilization and its variability; the client's own throughput and wait data.
Where it breaksSteady arrivals assumed when they cluster; a plan that runs at a utilization where the mathematics guarantees queues.
For exampleA thesis praises a clinic chain for running its rooms at ninety-five percent utilization. At that utilization the wait for a room is about twenty times the visit length whenever arrivals bunch, which they do every Monday. The utilization is not efficiency; it is the reason patients leave.
The mathematics in fullL = lambda x W for any stable system; with utilization rho = lambda / mu, the expected wait in the simplest queue grows like rho / (1 - rho), so from ninety to ninety-five percent utilization the wait roughly doubles.
Taught inMS&E 221 (Stanford); MECS 470
SourcesLittle (1961). A Proof for the Queuing Formula: L = lambda W. Operations Research 9(3), 383-387. https://doi.org/10.1287/opre.9.3.383
Some quantities wander with no memory and some are pulled back toward a level. We test which this one is, because a margin, a valuation, or a spread that reverts is an opportunity with a known half-life, and one that wanders is not.
dx = theta (mu - x) dt + sigma dWUhlenbeck and Ornstein's 1930 paper on Brownian motion gave the equation for a quantity that wanders but is pulled back toward a level, with a half-life for the return. The random walk with no pull is the model underneath Black and Scholes's 1973 option pricing, and the difference between the two is the question every claim of a trend has to answer first.
Why it works nowWhether a margin or a multiple is reverting or has moved for good has been decided by rhetoric, because estimating the pull needs a long history assembled carefully. The history can now be assembled and the half-life estimated, and operators can say whether the structural driver of the level has actually changed. The Desk asks how long it took to come back the last time the quantity moved this far, and whether it did.
When it appliesA margin, multiple, spread, or market share above or below its history; a thesis that extrapolates a run; a claim that something has permanently changed.
What we would askWhen this quantity has moved this far from its history before, how long did it take to come back, and did it?
What would have to be true for the old level to no longer be the level?
A long enough history to estimate the pull toward the mean and its half-life, and a view on whether the regime has changed.
How we would get itThe record for the series; operators on whether the structural drivers of the level have changed; the estimate of the half-life.
Where it breaksThe pull estimated on a window too short to tell it from zero; a regime change that moved the level.
For exampleA company's margins are five points above their twenty-year average and a thesis capitalizes them. Over that history, departures of this size reverted with a half-life of about two years, and every claim that this time was different accompanied one of them. The research is whether the driver has actually changed.
The mathematics in fullRandom walk: x(t + 1) = x(t) + e, no level is home; mean reversion: dx = theta (mu - x) dt + sigma dW with half-life ln 2 / theta; the question is whether theta is distinguishable from zero on the data available.
Taught inSTATS 217 (Stanford); MATH 385; MATH 386-1
SourcesUhlenbeck, Ornstein (1930). On the Theory of the Brownian Motion. Physical Review 36(5), 823-841. https://doi.org/10.1103/PhysRev.36.823
Black, Scholes (1973). The Pricing of Options and Corporate Liabilities. Journal of Political Economy 81(3), 637-654. https://doi.org/10.1086/260062
In a system where small differences grow, forecasts have a horizon beyond which they are worthless whatever the model, and the honest product is a spread of futures rather than one. We ask where the horizon is for this question.
Separation grows like d x exp(lambda t)Edward Lorenz found in 1963 that a simple model of the atmosphere run twice from starting points that differed in the third decimal place produced completely different weather within weeks. Sensitive dependence on initial conditions gave every forecast a horizon beyond which it is worthless, and weather forecasting responded by running ensembles of perturbed starts and reporting the spread.
Why it works nowLong-range business forecasts are still delivered as single paths to two decimal places, decades after meteorology stopped doing that. The track record of past forecasts by horizon in any market can now be assembled from the record, which shows where the horizon actually sits, and a model can be rerun from perturbed starts in minutes. The Desk asks how far ahead anyone in this market has ever forecast accurately.
When it appliesA long-range forecast presented with precision; a plan that depends on a point estimate five years out; an argument between two models about a distant future.
What we would askHow far out has anyone in this market forecast accurately, ever?
If you ran the same model from slightly different starting assumptions, how quickly would the paths diverge?
The track record of forecasts in this domain at each horizon, and the sensitivity of the outcome to initial conditions.
How we would get itThe record of past forecasts against outcomes by horizon; Advisors on what actually decided past outcomes; the report gives a spread, not a point, beyond the horizon.
Where it breaksPrecision beyond the horizon; a single path where an ensemble was honest.
For exampleA ten-year demand forecast for a commodity is presented to two decimal places. Past ten-year forecasts in the same market were wrong by a factor of two in either direction, and the errors were not the forecasters' fault. The decision is redesigned to be robust across the spread rather than optimized for the point.
The mathematics in fullIn a chaotic system two trajectories that start a distance d apart separate like d x exp(lambda t); the forecast horizon is roughly the time for that separation to reach the size of the quantity itself, and an ensemble of runs from perturbed starts is the honest forecast.
Taught inMATH 285; MECN-451
SourcesLorenz (1963). Deterministic Nonperiodic Flow. Journal of the Atmospheric Sciences 20(2), 130-141. https://doi.org/10.1175/1520-0469(1963)020<0130:DNF>2.0.CO;2
The largest loss in a hundred periods has its own distribution, and it is not the one the average describes. We fit the tail from the few largest observations and report the level exceeded once in a horizon, with the honest width of that estimate.
Maxima converge to Gumbel, Frechet, or WeibullFisher and Tippett proved in 1928 that the largest of many observations converges to one of only three distributions, whatever the observations came from, which is the theorem that lets the hundred-year flood be estimated from thirty years of data. The Embrechts, Kluppelberg, and Mikosch textbook brought the theory to insurance and finance, and the 1987 paper on self-organized criticality showed why avalanche sizes in many systems follow power-law tails.
Why it works nowRisk limits have been set to the worst observed loss, which is not the worst possible loss and is not even an estimate of it. The tail can now be fitted from the extremes in the client's history and comparable series, and reported with the honest width of the estimate, and operators can say what preceded past extremes. The Desk asks what the worst has ever been and how many observations that rests on.
When it appliesMaximum drawdown, the hundred-year event, a covenant or margin call triggered by an extreme, a plan sized to survive a bad year but not the worst year.
What we would askWhat is the worst this has ever been, and how many observations is that based on?
Do the bad periods cluster?
Enough history to observe several extremes, and a view on whether the process generating them has changed.
How we would get itThe record for the extremes in this and comparable series; operators on what preceded past extremes; the fitted tail with its interval.
Where it breaksExtrapolating far beyond the data; treating clustered extremes as independent; a process that changed.
For exampleA fund sizes leverage to survive the worst monthly loss in its fifteen-year history. Fitted to the tail, the once-in-fifty-years loss is nearly twice the worst observed, and the interval around it is wide. The leverage is set to the interval, not the history.
The mathematics in fullThe maximum of many draws converges to one of three families indexed by a tail parameter; the return level is the value exceeded once in T periods, read from the fitted tail, and its confidence interval is wide because it rests on the handful of largest observations.
Taught inMATH 385; MECN-451
SourcesFisher, Tippett (1928). Limiting forms of the frequency distribution of the largest or smallest member of a sample. Mathematical Proceedings of the Cambridge Philosophical Society 24(2), 180-190. https://doi.org/10.1017/S0305004100015681
Embrechts, Klüppelberg, Mikosch (1997). Modelling Extremal Events. Springer (book). https://doi.org/10.1007/978-3-642-33483-2
Bak, Tang, Wiesenfeld (1987). Self-organized criticality: an explanation of the 1/f noise. Physical Review Letters 59(4), 381-384. https://doi.org/10.1103/PhysRevLett.59.381
Quantities that arise from real processes have leading digits that follow a known distribution, with one appearing about thirty percent of the time. Numbers invented by people do not. We use the pattern as a screen for figures that deserve a closer look.
P(first digit = d) = log10(1 + 1/d)Simon Newcomb noticed in 1881 that the early pages of logarithm tables were dirtier than the later ones, and Frank Benford documented the law of leading digits across twenty thousand numbers in 1938. Theodore Hill's 1995 paper gave it a proper derivation, and accountants adopted it as a screen because figures invented by people do not follow the pattern.
Why it works nowA data room now arrives with more figures than any team can check by hand. The digit screen can be run across all of them in seconds, and the categories that deviate become the place the Advisors are asked to look first. It is a screen, not a conclusion, and the Desk says so. The Desk asks which figures were produced by a process and which were entered by a person.
When it appliesA data room, a set of financial statements, expense reports, or operating figures whose provenance is uncertain; a diligence process with more numbers than time to check them.
What we would askWhich of these figures were produced by a process and which were entered by a person?
Which categories, if any, deviate from the expected digit pattern?
A large enough set of figures spanning several orders of magnitude, from a process that should produce the pattern.
How we would get itThe client's data screened for the pattern; the deviations, if any, handed to the Advisors as the place to start asking questions.
Where it breaksApplying the screen to figures that have no reason to follow the pattern, such as assigned numbers or figures with a floor and a cap; treating a deviation as proof rather than as a place to look.
For exampleA buyer receives five years of monthly operating figures from a target. The leading digits follow the expected pattern in every category but one, where the pattern breaks in the two years before the sale. That category is where the operators are asked to look first.
The mathematics in fullP(first digit = d) = log10(1 + 1/d), so one leads about thirty percent of the time and nine under five percent; the law holds for quantities that are scale-invariant or arise from multiplicative processes, and departures are a screen, not a conclusion.
Taught inMATH 385
SourcesHill (1995). A Statistical Derivation of the Significant-Digit Law. Statistical Science 10(4). https://doi.org/10.1214/ss/1177009869
We keep our method in the open. The Decision Library lists every structure, each with where it began, why it works now, when it applies, what we would ask, what evidence would settle it, how we would get it, where the method breaks, a worked example, the mathematics underneath in symbols, and the papers it rests on. No equation is evaluated there and no answer is given. The mathematics underneath is set out at How the models work (45 structures), and every source is listed on the reading list (122 papers, 28 books).