Writing about what agents actually do.
Drift is the thing we exist to catch — an agent moving away from its own normal. These are write-ups on how that happens, how it is measured, and where the measuring breaks down. For a running record of incidents in the wild, see Field Notes.
The sandbox was built correctly. That was never the problem.
Most agent escapes really are misconfiguration. The harder case is the agent that never leaves.
The authority nobody wrote down
An agent can use every tool it was given and still do something no one approved.
OpenAI disclosed six. None of them was a prompt problem.
Six misalignment reports, and the thread running through all of them is the same.
A model should never be what decides something happened
Only what explains it afterwards. The difference is the whole architecture.
There is no such thing as a PQC certification
What the post-quantum mandates actually require, and of whom.
The second connection: what TLS inspection does to trust
Inspecting traffic means re-opening it. That leg carries its own trust decision.
How much of a Sigma pack survives when you never read the content?
We assumed 45 percent. Measured, it was 11.
The person holding the pager at 2am is not an ML specialist
Most agent security is written for people who will never read it.
Can an attacker poison the learning window?
The honest answer is yes, and it is why the window has to be scored rather than trusted.