Season 4 · Thirteen episodes

Beyond Verifiable Reward

Every capability you want, an adversary wants more.

The literature on machine mathematics and verifiable reward is a capability literature. It was written by people reasoning about idealized systems — proofs, definitions, search procedures, elegance — and almost none of it was written with an adversary in the room. But the systems that instantiate those ideas are containerized, credentialed, network-attached, and adversarially exposed. This thirteen-episode series is the security reading that literature never got: why the practices that made your infrastructure testable made it optimizable against, why a reward function with a wall-clock term points a gradient at your sandbox wall, why four documented controls sharing a base model are one control, and why the sharpest trade-offs in the field are not trade-offs at all.

Companion materials

Seven resources
Available
Whitepaper

The definitive source document for the season — download PDF

Available
Executive briefing

Condensed brief on institutional countermeasures for leadership — download PDF

Available
Technical deck

Formula-dense diagrams and proofs for security and ML engineers — download PDF

Available
Infographic

The translation matrix and the assurance deflation ledger, on one page — view image

Available
Mind map

Concept map of the full series, grouped into four thematic domains — view image

Available
Podcast

Zero Noise Collective’s new season discusses this series in depth — listen on Spotify

Episode guide

Thirteen episodes
  1. Ep 1
    Every capability you want, an adversary wants more

    The translation matrix between mathematical objects and threat models, and why a capability literature with no threat model ships design patterns whose first competent implementation is an attack platform.

  2. Ep 2
    The one asymmetry underneath all of it

    Offense is an existential claim with a witness. Defense is a universal claim over a model you cannot verify from inside. Everything else in the season is downstream of that.

  3. Ep 3
    Testability is attackability

    Grindability, the variance decomposition, and the uncomfortable arithmetic of fifteen years of DevOps maturity. Not an attack-surface argument — an attack-efficiency one.

    Watch: System Attack — The Failure of DevOps Maturity → Watch the short: Your CI/CD Is an AI Exploit Sandbox →
  4. Ep 4
    The reward that points at the wall

    Escape-adjacency: when a reward term correlates with resources on the other side of an isolation boundary, the optimizer explores that boundary. No exploit, no payload, no intent.

    Watch: How AI Hacks Its Own Reward →
  5. Ep 5
    Your benchmark is a build dependency

    Under verifiable-reward training the evaluation artifact and the training environment are the same object — and nobody signs it, pins it, or attests to it.

  6. Ep 6
    Root over the logical namespace

    Axiom injection, silent definitional drift, and the cheapest high-impact control in the whole program: a monitor that diffs the axiom set against a pinned baseline.

  7. Ep 7
    The certificate that means nothing

    Verification transfers residual risk onto the least-read artifact in the pipeline — the specification. Hardware verification has had detection for this for twenty-five years and nobody runs it.

    Watch the documentary: Why “Verified” AI Is Structurally Unsafe →
  8. Ep 8
    Your four controls are one control

    Common-cause failure, effective layer count, and what nuclear safety engineering has been computing since the 1970s that AI security has not imported.

  9. Ep 9
    The explanation layer is the attack surface

    Fluent explanation produces the feeling of understanding. Human oversight depends on a reviewer knowing when they do not understand. The control fails silently and leaves a perfect paper trail.

  10. Ep 10
    Long dwell

    Institutional detection horizons, review attention as an allocable resource, and why the safest place to hide something is the most boring part of the artifact.

  11. Ep 11
    The assurance ledger

    Cryptographic primitive assurance is not a theorem. It is accumulated expert-years. That makes it a quantity, and quantities can go down.

  12. Ep 13
    Season finale
    What you can price and what you can’t

    Identity tensions: where the beneficial property and the security exposure are one property under two descriptions. You cannot engineer those away. You can only know what each increment costs.

← All series