Imagine you want to know whether a new study app actually helps students learn better.
You look at kids who already use the app and kids who don’t. The app users get higher scores. Does that mean the app caused the improvement?
Not necessarily. Maybe the kids who chose the app were already more motivated, had better internet, or simply studied more. The decision to use the app carried “memory” of those prior differences. Any comparison is contaminated by that memory. This is the central problem of causal inference.
The core idea
Every claim that “A is better than B” rests on how the data were generated. Specifically, it depends on whether the critical choice points in the process — who got the intervention or action being tested, when they got it, which data were kept — carried information about the things we are trying to measure.
When a choice point carries memory of prior state, the resulting comparison can be misleading. When the choice is made without that memory, something special becomes possible.
Randomization: the memoryless move
Randomization is the procedure that deliberately erases memory at the moment of assignment. You flip a coin, use a random number generator, or run an automated A/B test. The decision of who receives version A or version B is independent of everything that came before: motivation, health, wealth, prior test scores, hidden preferences.
Because the assignment carries no memory, the two groups start equivalent (in expectation) on both the things we measured and the things we never thought to measure. Any later difference can be attributed to the intervention itself. That is why randomized experiments, including large-scale A/B tests, produce uniquely strong causal evidence.
This advantage is qualitative, not merely quantitative. In observational data you can adjust for the variables you know about, but you can never fully guarantee you have removed the influence of the ones you don’t. Larger samples make a biased estimate more precise; they do not remove the bias. Randomization is different: as the sample grows, the estimate converges to the true causal effect without needing to name or measure every confounder. No other common method offers that architectural path.
Nothing is perfect — and that is fine
Even the best randomization does not deliver eternal, absolute certainty. Real experiments can still suffer from people dropping out differently, imperfect blinding, or external events. Scientific knowledge is always provisional. What randomization changes is the kind of residual uncertainty that remains. It removes an especially dangerous class of bias by design.
History shows why this matters. Observational studies once strongly suggested that hormone replacement therapy protected women’s hearts and that certain vitamin supplements prevented cancer. Large randomized trials later overturned or sharply limited those claims. The earlier associations had been shaped by memory — healthier, more educated people were more likely to take the treatments.
A more subtle case: when time itself is the complication
Sometimes the world is changing while the experiment runs. Imagine testing a new clinic procedure by rolling it out to different health centres one after another in random order (a stepped-wedge design). Each centre serves as its own control: you compare the period before it received the new procedure with the period after.
The switch times are randomized, so the choice of when each centre starts is memoryless. That is powerful. Yet because the procedure is introduced gradually, more centres are using it later in calendar time. If the outcome was already improving (or worsening) for other reasons, time becomes partially mixed with the intervention.
The solution is to include time in the statistical model. Here is the crucial point: because the moments of switching were randomized, that adjustment for time is on much firmer ground than in an ordinary before-after study or a simple time series. The random switches deliberately expose the underlying time trend, allowing the model to separate it cleanly from the effect of the intervention. The memoryless choice at the switch points is what makes the later statistical control trustworthy.
The epistemic triad
We can now see the three elements working together:
- Causality is what we want to claim.
- Inference is the process of drawing conclusions from data.
- Randomization is the architectural feature that most powerfully protects the integrity of that inference by removing memory from critical choice points.
The deepest question to ask of any study is not
- “What model did they use?” but
- “Did the choices that generated these data carry memory of the things they are trying to explain?”
When the answer is no, the path to causal knowledge is clearer. When the answer is yes, the claim remains fragile, no matter how sophisticated the analysis that follows.
Understanding this architecture does not make every question easy. It does something better: it tells us which kinds of evidence can bear the weight we want to place on them, and why.
That is the beginning of clear thinking about what we know and how we know it.
Coda I: The Tragedy? or the Strategy of the Informed Elite?
The method that most cleanly solves the central problem of causal inference remains surprisingly niche.
Outside of clinical epidemiology and a handful of large technology companies, systematic randomized A/B testing is still the exception rather than the rule.
Tech firms discovered its power early because they had the ideal conditions: massive traffic, cheap randomization, immediate digital outcomes, and intense competitive pressure to know what actually works.
They industrialized it. Most other domains did not.In large parts of academia, business, government, and nonprofits, decisions are still driven by observational patterns, correlational studies, expert judgment, or before-after comparisons.
The architectural advantage of memoryless assignment — the ability to erase dependence on prior state and thereby isolate causal effects — is understood by specialists but has not become standard operating procedure. It often feels like a specialized technique rather than the default way to answer “does this change cause that result?”Several forces keep it partially hidden:
- It requires infrastructure and statistical literacy that many organizations lack.
- It forces confrontation with uncertainty and with null or disappointing results, which institutional incentives often discourage.
- In non-digital settings the logistics or ethics of randomization can appear harder than they need to be.
- Observational analysis and narrative explanation remain higher-status in many fields.
The result is a quiet asymmetry. The strongest available tool for learning about causes is heavily used in a few high-data, high-stakes environments and largely underused everywhere else. That gap between the power of the method and the breadth of its application is genuinely surprising.
Coda II: The Connection to the Black Hole Information Paradox
The black-hole information paradox sits at the collision of general relativity and quantum mechanics. When matter falls into a black hole, the detailed quantum information that described it seems to disappear. The Hawking radiation that slowly evaporates the black hole looks perfectly thermal: random, featureless, and dependent only on the black hole’s mass, charge, and spin. From the outside, it is as if the black hole has taken a highly structured prior state and returned something that carries no memory of that state.
In the language we have been using, a black hole appears to function as an absolute randomizer. The “choice” of what comes out in the radiation is independent of the microscopic details of what went in. All memory of the prior state is erased, or at least rendered inaccessible, in the semiclassical description. That is precisely why the paradox is sharp: quantum mechanics insists that information cannot be destroyed (unitarity must be preserved), yet the black hole seems to enforce a perfect memoryless map.
The parallel with causal inference is therefore not entirely crazy:
- In an ideal randomized experiment the assignment mechanism is constructed so that treatment is independent of every prior covariate — measured or unmeasured. Memory is deliberately broken at the choice point so that later differences can be read as causal.
- A black hole, in the Hawking picture, appears to do something even more extreme: it takes any incoming quantum state and maps it to thermal radiation that is independent of those details. The output looks maximally mixed; the prior state has been scrambled beyond recognition.
Where the analogy becomes especially interesting is in the modern resolutions. Most physicists now believe information is not truly lost. It is preserved in extremely subtle, highly scrambled correlations within the radiation (or encoded holographically). The black hole is not a destroyer of information but the universe’s most efficient scrambler. From the outside the radiation still looks thermal and memoryless for a very long time; recovering the original information would require measuring impossibly fine correlations. The memory is not erased; it is hidden by an almost perfect randomization.
This mirrors a deeper theme from our discussion. True randomness (or effective randomness) is powerful precisely because it breaks simple dependencies. In statistics that power lets us isolate causes. In quantum gravity it creates the appearance of information loss while, if unitarity holds, actually protecting the information in scrambled form. The black hole may be the physical system that most purely realizes the “memoryless choice” ideal — so pure that it forces us to confront what “memory,” “information,” and “causality” even mean when spacetime itself is dynamical.
We do not yet have a complete theory, so the connection remains speculative. But it is a fertile speculation: both domains are ultimately about how information about past states can be rendered independent of present observables, and what that independence allows us to conclude (or prevents us from concluding). The fact that the most extreme object in physics appears to enforce an absolute version of the same architectural principle we use to establish causes is, at the very least, suggestive.
Coda III: The Non Markovian Connection
The memoryless character of proper randomization is precisely the Markovian property applied at the point of assignment.
In a Markov process the future depends only on the present state; the past is screened off. Once you know the current state, additional information about how the system arrived there becomes irrelevant for predicting what happens next.
Randomization enforces an analogous screening at the critical choice point:
- The treatment assignment is generated so that it depends only on the designed probability (the “current state” of the assignment mechanism).
- It is rendered independent of the unit’s entire prior history — covariates, potential outcomes, unmeasured factors, everything that came before.
- Conditional on the randomization, the past is erased for the purpose of the assignment. The action (who receives treatment) no longer carries memory of that past.
That is why the groups become exchangeable. The assignment mechanism has been made Markovian with respect to the units’ histories: it forgets everything except the current random draw.
The parallel is not merely poetic. Both ideas rest on the same mathematical structure — conditional independence from the past given the present. In stochastic processes this yields the Markov property. In experimental design it yields unconfounded assignment. In both cases the power comes from deliberately discarding historical dependence so that the present action or transition stands free of it.
This also clarifies why imperfect randomization or post-randomization selection reintroduces problems: they re-inject dependence on prior states, violating the Markovian character of the assignment and thereby reopening pathways for confounding.
This observation ties the statistical architecture we have been discussing directly to one of the deepest simplifying principles in probability.
Randomization is the intervention that forces a non-Markovian world (units whose histories are full of relevant memory) to behave, at the decisive moment, as if it were Markovian.
A/B testing, Randomization, noisy unexpected outcome

Eduardo Bergel and Grok
t333t.com Research