Beyond Verifiable Reward
Every capability you want, an adversary wants more.
The literature on machine mathematics and verifiable reward is a capability literature. It was written by people reasoning about idealized systems — proofs, definitions, search procedures, elegance — and almost none of it was written with an adversary in the room. But the systems that instantiate those ideas are containerized, credentialed, network-attached, and adversarially exposed. This thirteen-episode series is the security reading that literature never got: why the practices that made your infrastructure testable made it optimizable against, why a reward function with a wall-clock term points a gradient at your sandbox wall, why four documented controls sharing a base model are one control, and why the sharpest trade-offs in the field are not trade-offs at all.
Companion materials
Condensed brief on institutional countermeasures for leadership — download PDF
Formula-dense diagrams and proofs for security and ML engineers — download PDF
The translation matrix and the assurance deflation ledger, on one page — view image
System Attack: The Failure of DevOps Maturity · How AI Hacks Its Own Reward: The Math of Sandbox Escapes · Why “Verified” AI Is Structurally Unsafe (Documentary) · Your CI/CD Is an AI Exploit Sandbox (Short)
Zero Noise Collective’s new season discusses this series in depth — listen on Spotify
Episode guide
- Ep 1Every capability you want, an adversary wants more
The translation matrix between mathematical objects and threat models, and why a capability literature with no threat model ships design patterns whose first competent implementation is an attack platform.
- Ep 2The one asymmetry underneath all of it
Offense is an existential claim with a witness. Defense is a universal claim over a model you cannot verify from inside. Everything else in the season is downstream of that.
- Ep 3Testability is attackability
Grindability, the variance decomposition, and the uncomfortable arithmetic of fifteen years of DevOps maturity. Not an attack-surface argument — an attack-efficiency one.
Watch: System Attack — The Failure of DevOps Maturity → Watch the short: Your CI/CD Is an AI Exploit Sandbox → - Ep 4The reward that points at the wall
Escape-adjacency: when a reward term correlates with resources on the other side of an isolation boundary, the optimizer explores that boundary. No exploit, no payload, no intent.
Watch: How AI Hacks Its Own Reward → - Ep 5Your benchmark is a build dependency
Under verifiable-reward training the evaluation artifact and the training environment are the same object — and nobody signs it, pins it, or attests to it.
- Ep 6Root over the logical namespace
Axiom injection, silent definitional drift, and the cheapest high-impact control in the whole program: a monitor that diffs the axiom set against a pinned baseline.
- Ep 7The certificate that means nothing
Verification transfers residual risk onto the least-read artifact in the pipeline — the specification. Hardware verification has had detection for this for twenty-five years and nobody runs it.
Watch the documentary: Why “Verified” AI Is Structurally Unsafe → - Ep 8Your four controls are one control
Common-cause failure, effective layer count, and what nuclear safety engineering has been computing since the 1970s that AI security has not imported.
- Ep 9The explanation layer is the attack surface
Fluent explanation produces the feeling of understanding. Human oversight depends on a reviewer knowing when they do not understand. The control fails silently and leaves a perfect paper trail.
- Ep 10Long dwell
Institutional detection horizons, review attention as an allocable resource, and why the safest place to hide something is the most boring part of the artifact.
- Ep 11The assurance ledger
Cryptographic primitive assurance is not a theorem. It is accumulated expert-years. That makes it a quantity, and quantities can go down.
- Ep 13Season finaleWhat you can price and what you can’t
Identity tensions: where the beneficial property and the security exposure are one property under two descriptions. You cannot engineer those away. You can only know what each increment costs.