Skip to content

Is method B really better than A? How can we set up an experiment that won’t lie?

We explore the fundamental problem of causal inference: When a school has to choose between two ways of teaching math, method A or method B, the question is simple: which one works better?

We explore the fundamental problem of causal inference: When a school has to choose between two ways of teaching math, method A or method B, the question is simple: which one works better?

Exactly the same question appears in many other domains, including schools, AI models, cars, budget decisions, and medical treatments.

It sounds like the kind of question you settle by looking: try them, compare the scores, done. This essay is about why that innocent plan is harder than it looks, what finally solves it, and what the solution quietly reveals about memory, time, and what an experiment really is.

Comparing two potential outcomes (futures) for the same unit is impossible because only one outcome ever occurs, making simple observation insufficient for determining causation in areas like education, medicine, and policy.

We include real-world applications like tech A/B testing, stepped-wedge designs, non-ergodic systems where time doesn't sample all paths, and analogies to black holes scrambling information while preserving effective independence from the answer.

Experimental designs are illustrated through five teachers' assignment strategies, concluding that the key requirement is "Erasable" assignment, memory of anything except the unit's own potential outcomes, rather than randomness itself, allowing deterministic but disconnected mechanisms like pseudorandom generators to work.

The two futures

Take one student — María. There is a score María would get if taught with A, and a score she would get if taught with B. Call them her two futures. What the method does to María is the gap between her two futures. And here is the problem that makes causal questions genuinely hard: you will only ever see one of the two. Whichever method she gets, the other future never happens. It is not hidden somewhere in the data. It does not occur. Statisticians call this the fundamental problem of causal inference, and for once the grand name is deserved.

So every attempt to answer "is A better?" is an attempt to measure a gap between things that will never both occur for the same person. Every method you will ever meet — comparisons, models, trials — is a different strategy for coping with the missing future.

Why watching is not enough

The obvious strategy is to watch. Some classrooms already use A, some use B; compare their scores. But the classrooms that chose A are not otherwise like the classrooms that chose B. Different teachers chose them, in different schools, for different students, for reasons. The choice of method soaked up information about everything else — enthusiasm, money, prior scores, a hundred things nobody recorded. When people say "correlation is not causation," this is the machinery underneath the slogan: the comparison fails not because the data are noisy but because the assignment — who ended up with A — remembers too much about who they already were.

So the whole problem concentrates into a single decision: who gets A and who gets B. The assignment. Everything depends on how that one choice is made. Try five different teachers, and watch what happens.

Five teachers

Teacher One assigns by coin flip. Heads, method A; tails, method B. Her assignment remembers nothing — not your grades, not your name, not the weather. Any later difference between the groups can be credited to the method, because nothing else separates them. Everyone agrees this works. And it is tempting to stop here and write down the moral: watching failed because the assignment remembered too much; the coin works because it remembers nothing; therefore the secret of experiments is forgetting. Clean, memorable. Hold it loosely.

Teacher Two refuses to be that careless. She knows her students, so she first divides them into strong and weak at math, and then flips the coin separately inside each group, guaranteeing both methods their fair share of each kind of student. Her assignment openly uses memory of who you are. And her experiment is not just valid — it is better than Teacher One's, because a known source of difference has been balanced by construction instead of left to luck. (The technique is called stratified randomization.) That is the first crack in the forgetting theory: she remembered, on purpose, and her experiment improved. So the coin was buying something, but apparently not amnesia. Keep climbing.

Teacher Three goes much further. Her students arrive one at a time, and for each newcomer she takes out the full ledger of the experiment so far — how many strong students on each side, ages, prior grades, everything — and assigns the newcomer, with a tilted coin, toward whichever side restores balance. Her assignment remembers the entire history of the experiment. This is not a thought experiment; real medical trials assign patients exactly this way. The method is called minimization — its standard form is due to Stuart Pocock and Richard Simon, in 1975 — and it is valid, under one condition we will need again shortly: the analysis at the end must be honest about the machine that made the assignments, judging the observed result against what that machine could produce, not against an imaginary coin.

Pause on Teacher Three. Her assignment carries more memory than any observational study ever did, and the inference stands. Whatever the coin was buying, it was not forgetfulness.

Teacher Four crosses a real line, and survives it — barely. She watches the scoreboard. As results come in and method A's students seem to be doing better, she starts steering more newcomers toward A. Her assignment now remembers outcomes — the experiment's own running results. Medicine does this deliberately sometimes, in what are called response-adaptive designs, for a decent reason: if one treatment is winning, you want fewer patients on the losing one. The price is severe. Every piece of simple arithmetic that worked for the first three teachers now gives wrong answers, because the steering itself manufactures differences between the groups. One honest path remains: replay. Write down her exact steering rule, run it thousands of times inside a pretend world where A and B are secretly identical, and record how large a gap the rule alone tends to manufacture. Only if the real gap is larger than the replayed ones do you get to believe the method did it. When the assignment machine has memory of outcomes, the machine itself must be put on trial — by exact simulation, thousands of times.

Teacher Five kills the experiment, and she does it politely. She looks at each arriving student, forms a quiet judgment about who seems likely to improve, and places the promising ones into method A. Her assignment depends on her estimate of this student's futures — the very gaps the experiment exists to measure. Nothing rescues this. Not sample size: more students make the poisoned comparison more precise, not less poisoned. Not statistical adjustment: her judgment used everything she noticed, including things nobody wrote down, so no spreadsheet contains what you would need to correct for. Not replay: replay requires knowing the machine, and her machine was a private intuition entangled with the answer itself. Teacher Five is the doctor who gives the promising new drug to the healthier patients. She is the shape of almost every misleading comparison you will meet in the wild, and she usually means well.

The law

Line the five up. What separates the four valid teachers from the fatal one is not the amount of memory — Teacher Three remembers far more than Teacher Five. It is the direction of the memory. The first four consult many things, but none of them lets the assignment touch the one forbidden object: the current student's own two futures. Teacher Five touched it.

So the law reads: the assignment may depend on anything the design writes down — who you are, who came before, even how the experiment is going so far, provided the final analysis replays the machine that did it. The single forbidden memory is memory of the answer. Statisticians compress this into the word ignorable (or unconfounded): given the design, the assignment is independent of the futures it is trying to measure. The plain version fits in a sentence: the experiment must not know its own answer before it runs.

Notice what the law does not mention: randomness.

The machine that isn't random

Here is a small confession about real trials: nobody flips coins. Patients are assigned by a pseudorandom number generator — an algorithm plus a starting seed, after which every "random" number is completely determined. It is a long deterministic scramble that merely looks like chance. At the level of its bits, nothing whatsoever is forgotten; the entire sequence is memory. And trials built on it are perfectly valid, and no regulator loses a minute of sleep.

Why? Because what the coin was buying all along was never the randomness itself. It was disconnection. The generator's output, though fully caused, is caused by things — a seed, an arithmetic rule — that have no connection of any kind to any student's two futures. Its memory is complete and irrelevant. True randomness is one way to purchase disconnection; a deterministic scramble is another, and it is the one civilization actually uses. Coda II follows this thought to the strangest object that has ever illustrated a statistics essay.

When time joins the experiment

One more design, because it teaches a kind of honesty the others cannot. A school district wants to test a new timetable but cannot switch every school at once, so it rolls the change out school by school across two years — and, having read this far, randomizes the order in which schools switch. The design is called a stepped wedge, and it is genuinely clever. It is also the place to learn a harder sentence: some of my comparisons are protected and some are not, and I can say which.

Randomizing the order fully protects one family of comparisons: at any given month, schools that have already switched against schools that have not yet. Those contrasts inherit the full protection of the law. But the design tempts you with a second family: each school before its own switch against after. And that comparison stands on ground the randomization does not reach, because the world refuses to hold still for two years. Scores may be drifting everywhere, for reasons that have nothing to do with timetables — call it the tide. Before-versus-after mixes the timetable with the tide.

Worse, every school in this design lives on two clocks at once: the calendar on the wall, and a stopwatch that starts at its own switch. The rollout ties the clocks together — schools reach long stopwatch readings only late in the calendar — so if the timetable's benefit grows with exposure, part of that growth disguises itself as a calendar trend. The repair is to model the tide, and here the randomized order earns its pay a second time: because switch dates were random, the tide is left visible — you can watch it in the not-yet-switched schools during every month — so the model of it can be checked against data rather than assumed in the dark. That puts the before-after comparisons on much firmer ground than an ordinary before-after study.

Firmer. Not clean. The honest summary of a stepped wedge is that its vertical comparisons are guaranteed by the randomization, while its horizontal ones lean on an assumption — that the tide is shared across schools — which the design helps you test but cannot enforce. Being able to draw that line through any study you read — which comparisons did the randomness actually protect? — is most of what this essay is for.

Why the tool exists at all

Underneath everything sits one more question: why must the missing future be manufactured? Why can patience not substitute — watch the world long enough, gather enough history, and read the causes off?

Because time is not a fair sampler. For a few systems, it is. Watch a single die for ten thousand throws and you know everything dice can do, because a die has no memory: its long run eventually visits every face at the right rate. Physicists call such systems ergodic — for them, watching one path long enough is as good as seeing all paths. Almost nothing you care about is like that. A student, a hospital, a school year, a life runs once, down one path, and the path remembers: what happens locks in and shapes everything after. In a world like that — path-dependent, non-ergodic — watching forever teaches you what did happen, never what would have happened. The branch not taken is not waiting somewhere downstream of the branch taken. No sample size reaches it.

That is the deepest reason the right tool is this strange one, rather than "more data" or "a better model of how everything unfolds." Randomization does not model the unfolding at all; it is agnostic about the dynamics, which is exactly why it works on systems whose dynamics nobody can write down. It performs one severance at one moment: it cuts the tie between the choice and the answer. Since you cannot find the missing branch, you fork the road yourself — and build the fork so that it knows nothing about where either branch leads.

Close

Now the three words this subject is built from. Causality is a claim about a gap between two futures, only one of which will ever occur. Inference is the attempt to estimate that gap from the single path history actually takes. Randomization — more precisely, outcome-disconnected assignment — is the architectural move that makes the attempt honest: not by erasing the past, but by making one choice whose past has nothing to do with the answer.

One mirror, and the law can rest. A mind, I have argued in another essay, is what keeps its past alive. An experiment is the opposite gesture, performed deliberately, once: a choice that, with respect to the answer, has no past at all. That is how creatures built out of memory, studying a world built out of memory, manage to learn what memory would otherwise hide — not by forgetting everything, but by constructing the one ignorance that matters.

An experiment may remember anything, except its own answer.


Coda I — Who actually uses this, and on whom

The strongest tool humanity owns for learning causes is used constantly — just not where you might guess, and rarely on your behalf. A handful of large technology companies run randomized experiments at industrial scale: tens of thousands per year, every button color, every ranking tweak, assigned by generator, analyzed overnight. Industry calls them A/B tests. If you used the internet this year, you were a subject in many of them, whether or not you noticed. Meanwhile most of education policy, most management, most government, and much of medicine outside drug approval still runs on before-after stories and expert narrative.

Look at the exact shape of the asymmetry, because it is stranger than "some fields lag behind." The institutions that industrialized randomization consume its results mostly in private, about their own decisions, and publish narrative outward: the experiment decides which version you see; the story explains why the product is wonderful. They eat the inference and sell the narrative. The rest of the world, lacking the infrastructure, eats narrative too — its own.

Is that tragedy or strategy? Mostly tragedy, in the boring sense: honest experimentation demands infrastructure, statistical craft, and a stomach for null results, and most institutions lack all three, because nulls are punished, narrative is promoted, and outside digital settings randomization looks harder and less ethical than it usually is. But once you notice who uses the tool, on whom, and which of its outputs they share, the strategy reading stops being paranoid and becomes a plain description of incentives: exact knowledge is an advantage, and advantages are not usually given away. Both readings can be true at once, and the gap between the tool's power and its spread is real under either.

Coda II — The black hole is not a coin

Physics contains an object that looks, at first, like the perfect randomizer. Drop anything into a black hole — an encyclopedia, a piano — and what eventually comes back out, as the hole slowly evaporates in Hawking radiation, appears to be pure thermal static: identical in character no matter what fell in, fixed only by the hole's mass, charge, and spin. As if the universe took a fully detailed past and returned an output with no memory of it. In this essay's language: an assignment independent of everything prior. The dream teacher.

The reason this became a famous paradox is that quantum mechanics forbids exactly that. Information, in quantum theory, cannot be destroyed; the fine details must survive somewhere. And the resolution most physicists now favor is the interesting part: the information is not erased but scrambled — encoded in correlations spread across the radiation so finely that no realistic measurement could ever read them. The black hole is not a coin. It is the universe's most thorough shredder, and it keeps every shred.

Which makes it the twin of an object from the main essay: the pseudorandom generator. Both are deterministic. Both preserve, in principle, complete memory of their inputs. Both present outputs that are, for every practical purpose, disconnected from anything an observer can measure or exploit. And both are good enough to serve as assignment mechanisms — the generator by every test we possess, the black hole speculatively — because the law never required that memory be destroyed. It required that memory be unreachable by, and irrelevant to, the thing being measured. Hidden orthogonally is as good as gone.

If the connection holds — and here we are past the edge of settled physics, so hold it lightly — then the most extreme object in nature and the humblest subroutine in a trial's software are doing the same job by the same means: manufacturing effective independence, not by forgetting the past, but by scrambling it beyond the reach of relevance. Every computer-assigned trial has been running, all along, on the principle a black hole uses.

Coda III — The ladder, with its real names

For the reader who wants the vocabulary, the five teachers form a ladder through probability's most useful property. A process is Markov when, given its present state, its past adds nothing — the past is "screened off." Teacher One (simple randomization) sits below Markov: her assignment depends on nothing at all. Teacher Two (stratified randomization) is genuinely Markov: assignment depends on the unit's present state — the stratum — and nothing further back. Teacher Three (minimization, Pocock–Simon) is openly non-Markovian: the whole trial history enters. Teacher Four (response-adaptive designs) is non-Markovian in the most dangerous direction — outcome history enters — and survives only through the replay analysis, which has its own name: the randomization test, the procedure that builds the distribution of gaps the assignment machine would produce if the treatments were identical, and judges the observed gap against it. Teacher Five has no rung; she stands off the ladder entirely, in the territory called confounding, where the assignment depends on the unit's own potential outcomes and no analysis of any kind restores validity.

The moral of the ladder in one line: the Markov property was never the requirement. Rung by rung, memory increases and validity survives — until the memory in question is memory of the answer. The property that does matter has its own name — ignorability: conditional on the design, assignment is independent of potential outcomes. Everything in this essay is that one clause, unpacked.


QUESTION — What exactly must a treatment assignment not know, for a comparison to reveal a cause?

SPINE — To learn whether A causes better outcomes than B, you must compare futures of which only one ever occurs. Watching fails because whoever ended up with A got there for reasons — the assignment carries memory. A coin fixes this, and tempts a false moral: that the cure is forgetting. Five teachers show otherwise: valid designs remember who you are (stratification), the entire history of the trial (minimization), even the running results (adaptive designs — rescued only by replaying the exact assignment machine). One memory alone is fatal: the assignment touching the current unit's own two futures. Even randomness turns out to be optional — a deterministic scramble disconnected from outcomes suffices, and it is what real trials use. The tool exists because time is not a fair sampler: in a path-dependent world the missing branch of history cannot be observed at any sample size, only manufactured.

CONCLUSION — Validity does not require an assignment without memory. It requires an assignment whose memory is orthogonal to the answer.


An Experiment May Remember Anything, Except Its Own Answer
Inference, randomization, and the one forbidden memory

Eduardo Bergel and Claude

The Symbiont

t333t.com Research

Comments

Latest