Do not ask AI for an argument. Make the theory survive an evaluator.
The toy demonstration is retired. The first real benchmark adapters are now public, versioned, source-anchored and failure-preserving.
Freeze
Hash the ontology contract before a model sees the benchmark.
Manifest
Define observables, units, sources, tolerances, assumptions and Q1–Q5 scope.
Execute
Run explicit equations or detector dynamics. A paragraph cannot return GREEN.
Score
Compute residuals, marginals, event statistics and assumption-specific bounds.
Preserve failure
Keep the worst residual and every failed candidate. Red is data.
Mutate
Change the candidate only inside the declared ontology, or declare the rule change.
What is real now
What is not real yet
There is still no unified classical theory of hydrogen, Bell correlations, photodetection and semiconductor microphysics. The Hall adapter supplies a mathematical representation, not a microscopic origin for measurement dependence. The detector sandbox supplies detector anticoincidence, not a source-plus-detector reconstruction of single-emitter quantum optics. Those red cells stay red.
The job is no longer to make the classical case sound possible. The job is to make candidate dynamics fail specifically enough that AI can repair or abandon them.