Discourse

Worth reading elsewhere.

This is where we collect writing from outside Habenula that shaped how we see the problem. Pick a shelf: disclosed vulnerabilities, scholarly work, or industry commentary. Every link opens at its source in a new tab.

Disclosed vulnerabilities

Not hypothetical.

Each carries an official CVE. Every link opens the National Vulnerability Database record in a new tab — the NVD record is the authoritative description, and the summaries here are ours. These are publicly disclosed, tracked vulnerabilities; inclusion says something about the failure class, not about any product today.

Zero-click data exfiltration from an enterprise AI assistant.

Hidden instructions in an inbound email made the assistant read internal data and send it out, with no action from the user. The canonical proof that broad data access plus an untrusted input channel is an exfiltration path — the failure scoped permissions, egress limits, and an audit log exist to bound.

Read at source →

Code execution in an agent framework’s tool via Python exec.

An early, widely-cited demonstration that a mainstream agent framework handed model output straight to exec. The reason tool execution has to run through a policy decision rather than trusting what the model emits.

Read at source →

OS command injection in an MCP client connecting to a malicious server.

Merely connecting an MCP client to an untrusted server was enough for that server to run shell commands on the client. The concrete “malicious tool provider owns the client” threat — the reason an inbound tool surface is authorized by a verified caller, not by who connects.

Read at source →

Persistent code execution via a silently swapped MCP tool config.

A tool configuration approved once stayed trusted after it was silently modified, turning a shared repository into a code-execution vector. A time-of-check-to-time-of-use failure — the reason a decision is re-checked at execution, not trusted from a first approval.

Read at source →