Narrating emergence is not proving it
Anyone can claim that complexity emerged from their code. Proving it is another matter: comparing it against pure chance, and telling what the program invents apart from what I wrote by hand.
The dog in the cloud
A cloud drifts by, and someone sees a dog in it. The snout, the ears, the trailing tail. The shape is really there, it just says nothing about the cloud. A running simulation triggers the same reflex: you watch it live and you swear you see emergence, that moment when a system finally surprises you. Saying so costs one sentence. Proving it is another matter, and that is where everything I learned by trying, then failing, sits.
I paid dearly for that lesson. Months of investment, nights spent believing I was building a measuring instrument, before discovering that filling a catalog of traits opens no space at all. This piece is about what I managed to measure, and above all the boundary I could never cross.
Narrating is free
Watch any simulation run long enough and you will see shapes. A cluster that looks like a town, a lineage that seems to learn, two groups that "decide" to go to war. The human brain spins a story out of three moving pixels. That is a talent, and it is a trap.
Narrated emergence is just that: I watch, I see a pattern, I give it a name. Nobody can contradict me, because I am measuring nothing. The problem is not that the story is false. It is that it would be just as convincing in front of a purely random world. It is the dog in the cloud again: the shape is convincing, it proves nothing.
The null model: comparing against chance
To get out of storytelling, you need a point of comparison. In science it is called a null model: a control version of the system where the mechanism you want to test has been unplugged and replaced with pure chance. If the real system does no better than that blind version, you were looking at clouds.
In Limon, my world of proto-humans, I did exactly this, through what is called an ablation, meaning you deliberately remove one part of the engine to measure what it contributes. The part here is heredity. In the control world, each newborn receives a genome drawn at random instead of its parents'. Everything else runs identically, with no forced extinction. Then I measure one single thing: how far the newborns' genetic makeup drifts, generation after generation, from its starting point. Not the survival of adults. What births pass on.
Across six different worlds, each unrolled over four thousand simulation steps, that gap averages 0.590 when heredity works, against 0.092 when it is random. A ratio of 6.4. So the adaptation that gets transmitted is about six times stronger than what pure chance would produce. Mind the meaning: this is not six times more survival, it is six times more inherited adaptation, the kind carried through births. Along the way, genetic diversity drops when selection acts: the population converges toward a profile that works.
What makes the test honest is that I fixed the bar before seeing the result. Below a ratio of three, I had committed to concluding that selection was indistinguishable from mere random drift. Choosing the threshold after the fact means proving yourself right every time.
A catalog is not a space
So much for the measurable part. Now the trap.
This test proves that Darwinian selection works on a trait. It does not prove that open-ended emergence is taking place.
A trait like the tendency to settle down is one I wrote myself, with its effects, its bounds, its way of being passed on. That selection pushes it in one direction is real, is measurable, is what the 6.4 says. But it remains a catalog I fill in by hand. Adding fifty more traits makes a bigger catalog, not a more open world.
Open-ended emergence is something else: a space where the program builds things the designer never wrote, and can keep building them without end. The canonical examples come from the artificial-life work of the 90s. In Tierra, then Avida, the "organisms" are little programs that copy themselves into the computer's memory, with copy errors serving as mutations. Nobody wrote "invent a parasite" in there. And yet parasites appeared: programs that, having lost their own copy routine, borrow the neighbor's to reproduce. That is a space. Not a catalog.
What the null model can never tell you
Here is the limit that cost me months. A null model can certify "this beats chance." It cannot certify "this is something I did not foresee."
Chance is a yardstick: we know how to unplug it and measure against it. The designer's surprise is not one. The day my program produces a behavior, either I had made it possible in my code, and then it is not an invention, or it comes from a bug, and then I fix it. There is no convenient third box to tick. Proving open-ended emergence is not passing a test, it is demonstrating that the space of the possible overflows what I put into it. That is a wholly different undertaking, and my architecture was not built for it.
Chance can be measured, imagination cannot
Measuring a system against chance can be done, and it is already demanding: you have to unplug a mechanism, compare two worlds, set your bar in advance. Measuring a system against your own imagination cannot be done with a number.
Narrating an emergence costs one sentence. Proving it costs a null model. Proving it is open-ended may cost a research career. Me, I made a game.