Safe by Default: Building Agents That Can't Wreck Your Data
Read/write enforcement, principal scoping, approval gates, sandboxes and budgets: five safety properties, every one enforced in code, not in a prompt.
Agent Harness03
Subject · 4 posts
Read/write enforcement, principal scoping, approval gates, sandboxes and budgets: five safety properties, every one enforced in code, not in a prompt.
Agent Harness03
The model's output came from your trusted system, so it feels safe. It is not: untrusted input shaped it, and one rendered image can carry your data to the attacker.
LLM and Agent Security05
The signature attack on language models. How instructions hide in the text a model reads, why you cannot filter your way out, and what containment looks like.
LLM and Agent Security02
Every security bug in the last series was data mistaken for code. A language model makes that its whole way of working: to it, all text is instructions.
LLM and Agent Security01