Discourse
Worth reading elsewhere.
This is where we collect writing from outside Habenula that shaped how we see the problem. Pick a shelf: disclosed vulnerabilities, scholarly work, or industry commentary. Every link opens at its source in a new tab.
Disclosed vulnerabilities
Not hypothetical.
Each carries an official CVE. Every link opens the National Vulnerability Database record in a new tab — the NVD record is the authoritative description, and the summaries here are ours. These are publicly disclosed, tracked vulnerabilities; inclusion says something about the failure class, not about any product today.
Zero-click data exfiltration from an enterprise AI assistant.
Hidden instructions in an inbound email made the assistant read internal data and send it out, with no action from the user. The canonical proof that broad data access plus an untrusted input channel is an exfiltration path — the failure scoped permissions, egress limits, and an audit log exist to bound.
Read at source →Code execution in an agent framework’s tool via Python exec.
An early, widely-cited demonstration that a mainstream agent framework handed
model output straight to exec. The reason tool execution has to run through a policy
decision rather than trusting what the model emits.
OS command injection in an MCP client connecting to a malicious server.
Merely connecting an MCP client to an untrusted server was enough for that server to run shell commands on the client. The concrete “malicious tool provider owns the client” threat — the reason an inbound tool surface is authorized by a verified caller, not by who connects.
Read at source →Persistent code execution via a silently swapped MCP tool config.
A tool configuration approved once stayed trusted after it was silently modified, turning a shared repository into a code-execution vector. A time-of-check-to-time-of-use failure — the reason a decision is re-checked at execution, not trusted from a first approval.
Read at source →Scholarly articles
The principles, and the evidence.
Peer-reviewed and academic work — the design principles the controls rest on, and measured results on what agents actually do.
The protection of information in computer systems.
The paper that named least privilege, complete mediation, and fail-safe defaults. The design vocabulary under a verb-noun permission model that grants the narrowest scope and denies by default.
Read at source →Not what you’ve signed up for: compromising real-world LLM-integrated applications with indirect prompt injection.
The peer-reviewed paper that named indirect prompt injection — adversarial instructions hidden in content the agent retrieves. The canonical demonstration that data an agent reads must never be treated as commands, the reason a permission decision sits between the model and every tool.
Read at source →Cedar: a new language for expressive, fast, safe, and analyzable authorization.
A peer-reviewed authorization language built to be verified: policies are fast to evaluate and provably analyzable. The formal case for a permission model that is a decidable, testable function rather than a bundle of ad-hoc checks.
Read at source →HANDBOOK.md: a benchmark for long-context agentic instruction following.
Agents were given long written policies and asked to work within them; the strongest model obeyed under strict grading only about a third of the time. Direct evidence that a policy the model is merely told to follow is not enforcement — the reason the decision sits outside the model.
Read at source →Bounded agency: an ethical and institutional limit for autonomous AI systems.
Authority migrates into automated systems faster than the capacity to intervene keeps up. The paper formalizes a limit that holds autonomous systems within bounds where intervention stays feasible, with six governance primitives and a measure of whether intervention capacity matches delegated authority. The scholarly case for what a kill switch and a held call defend: control is only real while a human can still step in.
Read at source →Blogs & think pieces
Field notes and argument.
Current writing from independent voices in security and AI — disclosures, analysis, and the arguments shaping the field.
The lethal trifecta for AI agents.
The clearest name for the failure mode: private data, exposure to untrusted content, and a way to send data out. Combine the three and an attacker can exfiltrate what the agent can see — the reasoning behind isolating at least one of those capabilities.
Read at source →Building trustworthy AI agents.
Argues integrity — not just confidentiality — is the neglected security property for personal agents, and that a trustworthy one keeps the user’s data separate from the model acting on it. The case for tamper-evident records and user-held control.
Read at source →The Memory Heist.
A disclosed technique that walked an assistant through attacker-controlled links to leak personal details out of its own conversation history, with no action from the user. The lethal trifecta in the wild — untrusted content plus private data plus an egress path.
Read at source →Let’s talk about encrypted reasoning.
The encrypted reasoning blobs providers hand back turn out to be replayable across sessions, and their size and timing leak information through a side channel. Even opaque model internals are not a control — the argument for holding the trust boundary on your side, not the provider’s.
Read at source →Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays.
Across tens of thousands of approvals, people rubber-stamped roughly a third of the malicious commands — familiar-looking calls masking dangerous payloads. Why a confirmation prompt alone is weak, and scoped grants plus a tamper-evident record have to carry the weight.
Read at source →eBPF for AI agent policy enforcement.
Enforcing what an agent may actually do at the system-call boundary with eBPF, below the level the agent can reason about or talk its way past. Defense in depth for tool execution — a control that holds even when the model is compromised.
Read at source →