Skip to content

Build. Test. Rethink.

How useful agents fail—and which defenses preserve their usefulness.

Independent research for engineers giving AI agents access to code, private data and tools. Documented failures, practical defenses, and the evidence behind deployment decisions.

Three questions we follow

Follow the failure

Which permitted capabilities let a reported attack reach its objective? We trace the task, authority and path to the outcome.

Examine the defense

Where is the control enforced? We compare what it prevents, what useful work it preserves, and what it costs.

Challenge the assurance

What does the evidence actually establish? We separate observed behavior from coverage claims and deployment judgments.

Our evidence standard

A useful conclusion comes with its limits.

We link primary sources, distinguish reported findings from our own analysis, and explain the conditions under which a defense helps. A proposed experiment is never presented as a result.