The state of agent governance.
Personal agents can now act in your accounts. The thirteen-point bar for governing them, and where today's products — Habenula included — stand against it.
Personal AI agents crossed the line from answering to acting — they now send mail, move money, book, post, and delete across a person's real accounts. The interesting question stopped being "how good is the model" and became "does the person stay in control of what it does." The capability crossed; the governance mostly didn't. This document does three things: it states a standard for what a well-governed personal agent looks like; it credits the consumer market for the steps it has shipped toward that standard; and it measures the category against the bar — where it lands, and where it doesn't.
Habenula is then measured against the same bar, gaps included.
Scope is deliberately narrow: personal, consumer products a normal person installs and runs — the platform agents (OpenAI's ChatGPT/Atlas, Google's Gemini, Apple Intelligence + Siri) and the consumer agents (Anthropic's Claude Cowork, Perplexity's Computer/Comet, Lindy, Fyxer). It excludes developer and coding agents, enterprise control planes, and governance middleware — not because those don't matter, but because none is a product an individual installs and trusts on their own.
A note on freshness: statements about other vendors' products reflect their public documentation and third-party security research as of July 31, 2026, and this category moves fast. Each such statement carries a numbered source anchor (see Sources at the end); re-check against the vendor's current documentation before relying on it. The Methodology and Scope of Claims section at the end states exactly what these comparisons do and do not assert.
The standard: what a well-governed personal agent looks like
Thirteen properties. Each is stated as what it means for the person, not the engineer.
- Every action is judged by deterministic code the model cannot reach. A confirmation dialog sits "outside the model" only trivially: the model still decides whether to raise it and writes what it says, so a manipulated model can skip the prompt or misdescribe it. The real bar is stricter, and threefold. The check runs on every consequential action, on the one path it must travel — the model can't route around it. Its verdict is a pure function of the action and your standing rules — same action, same rules, same answer, deny by default — not a fresh yes/no and not the model's judgment. And the model shapes none of it: not whether the check fires, not the outcome, not what you are shown. A one-off confirmation prompt clears none of the three.
- The model never holds your credentials. OAuth tokens and keys are resolved and used outside the model's context. A hijacked or mistaken model cannot spend what it never held.
- Permissions are granular and bound to a resource. Not "email: allowed," but "send email to this domain," "read this label," "write this folder." A grant without a concrete noun is not a grant.
- Permission proposals are smart and low-friction — without dulling the ask. The agent asks at the moment a capability is first needed, at the scope you choose, and stops re-asking what you've already settled — while still surfacing the genuinely consequential and novel. For multi-step work, it proposes a legible plan you approve once, carrying declared limits. The goal is precision, not fewer prompts for their own sake.
- Grants are scoped and mortal. Allow once, or allow for a set time with a real expiry — not a standing "always" that outlives the reason you granted it. Least standing authority, by default.
- The highest-stakes actions can be gated behind a human gesture a program can't fake. The system supports requiring a biometric or hardware-key presence check — bound to the specific approval, so a compromised process can't click through on your behalf — for the actions you choose. A capability you turn on where it matters, not a blanket requirement.
- Every action is recorded in a tamper-evident log the person can verify themselves. Not vendor-side debug telemetry you're asked to trust — a cryptographically chained record you can export and check independently.
- Spending is bounded, hard, and per-service. A cap the agent cannot exceed, set where the money actually leaves — not just a plan-level ceiling.
- There is a real, instant kill switch. One action stops everything, fast and globally — ideally revoking provider access as defense-in-depth, but at minimum leaving nothing able to act.
- Idle agents wind down on their own. An agent you stop paying attention to should stop — a dead-man heartbeat or session expiry, not an open-ended process that runs until someone remembers to kill it.
- The controls are open, and the person can run them. The code that holds your keys, enforces your rules, and writes your audit trail is public — you can read it, and you can run it yourself. A control you can't inspect is asking for your trust, not handing you control.
- The model is yours to choose. Model-agnostic by design, because a control layer owned by the model vendor it is meant to govern is a conflict of interest, not a safety story.
- It's built for the individual. A product a person installs and trusts on their own — no IT administrator, no developer assembly, no fleet to manage.
Where the market is moving
The field is moving, and several of these steps make users concretely better off:
- OpenAI — ChatGPT Atlas, Agent Mode. Agent mode runs in a managed browser and pauses to confirm sensitive actions like logins and payments; offers a logged-out mode that uses none of your cookies or accounts without explicit approval; is sandboxed (it can't run code, download files, install extensions, or touch your file system); lets you set custom approval checkpoints; and OpenAI is publicly hardening it against prompt injection as an ongoing program.[1][2] That is consent-on-sensitive, least-context-by-default, and execution sandboxing, shipped to a mass audience.
- Google — Gemini (the former Project Mariner, folded into Gemini Agent in May 2026). It shows a live view of what it's doing, logs its actions, asks before sensitive steps, and lets you intervene at any time;[3] developers can require it to refuse or confirm specific high-stakes actions.[4] Transparency and interruptibility done well.
- Apple — Apple Intelligence + Siri (iOS 27 beta, App Intents 2.0). Agentic action runs on-device or through Private Cloud Compute and can act only through typed App Intents — a structured action surface rather than free-form tool calls.[5] Reporting on WWDC26 characterized the new Siri as agentic with a human in the loop, not autonomous.[6] The loop is coarser than it sounds, though: in the Passwords app, a single tap authorizes a batch credential run, and individual changes then proceed without further confirmation.[7] Still, a privacy-first architecture and a typed capability surface are the right foundations. (iOS 27 was in developer beta as of the review date.)
- Anthropic — Claude Cowork. Ships per-folder scoping for local file access, per-connector enable/disable (Workspace connectors themselves are account-scoped — they mirror the user's own Google permissions), a Gmail connector that reads and drafts but cannot send, risk-tiered permission controls on managed plans ("Always allow / Needs approval / Blocked" per connector category, plus an org-level gate — off by default — that forces per-task approval of write tools), and explicit consent before destructive actions like file deletion.[8][9][10] The most deliberate permission surface in the consumer set.
- Perplexity — Computer. A default hard spend cap ($200/mo, raisable), micro-VM-per-session execution isolation, and a per-task stop.[14][15][16] A default hard ceiling and per-session isolation are exactly the right defaults.
- Consumer email agents — Lindy, Fyxer. Mature connector scoping and compliance posture (SOC 2 Type II in both cases, HIPAA in both, ISO 27001 in Fyxer's) — proof that consumer-grade agents can carry enterprise-grade security hygiene.[17][18]
Read together, the category is converging on real lessons: confirm on the actions that scare people, default to least context, sandbox execution, cap spend, keep a human in the loop, and — Apple especially — treat privacy architecture as a feature.
Measured against the standard, the category falls short — structurally
That progress is real. Against the thirteen-point bar, though, every consumer product today lands well short — and in the same places.
| The standard | Where the consumer category is (as of July 31, 2026) |
|---|---|
| 1. Deterministic gate, outside the model | Absent. The "gate" is the model's own judgment about whether to pause — the manipulable component deciding its own limits. Injection walks past the prompt because the hijacked model simply doesn't ask. |
| 2. Model never holds credentials | Absent. Platform and consumer agents hand the model the session or the tokens by design — Cowork is single-vendor (one company holds the model and the keys); Comet was given the active browser session.[19][20] The least-predictable component holds the keys. |
| 3. Granular, resource-bound permissions | Partial at best. Cowork's local-folder scoping is the high-water mark (its Workspace connectors stay account-scoped);[8] no product requires a noun binding ("this recipient / this path"). Grants stay coarse. |
| 4. Smart, low-friction proposal | Inverted. Confirmation prompts, not structured proposals — and tuned to reduce asking (Anthropic engineering reports an 84% cut in permission prompts from its Claude Code sandboxing[13]), which trains reflexive approval rather than precision. |
| 5. Scoped, mortal grants | Absent. Credential and connector grants persist until manually revoked; no one-time or timed grants with a real expiry. The closest step — write tools defaulting to per-task approval on Cowork's managed plans — is a per-action ask, not a mortal grant.[9] |
| 6. Human-gesture approval | Absent. Approval is a click any process can fake. Apple's "tap" is the closest instinct, but it's a screen tap, not a hardware-bound signature over the specific action. |
| 7. Tamper-evident, user-verifiable audit | Absent. For months after launch, Cowork activity was explicitly excluded from Anthropic's audit logs, Compliance API, and data exports across all tiers, per Anthropic's own help center. The fix Anthropic shipped in August 2026 captures cloud (web/mobile) sessions in the Compliance API on Team/Enterprise; local desktop sessions remain excluded from any centrally exportable record.[10][11][12] Where any log exists elsewhere it's admin/enterprise-scoped and often omits agent-generated content. No user-verifiable chain anywhere. |
| 8. Hard, per-service spending caps | Partial. Perplexity's default $200 ceiling is the best in class;[14] caps are account/plan-level, never per-service or per-action. |
| 9. Instant global kill | Absent as a primitive. "Stop the task," "close the app," "delete the task" — no documented one-action global stop, and no provider-side token revocation. |
| 10. Idle wind-down | Absent. Agents run until the task finishes or someone stops them; Cowork even runs scheduled tasks with no device online.[10] No dead-man expiry. |
| 11. Open + run-it-yourself | Absent. Closed-source, single-vendor clouds — you cannot read the code that holds your keys, and you cannot run it yourself. |
| 12. Model choice | Absent. Each is locked to its own model (OpenAI, Anthropic, Google, Apple); the control layer is owned by the model vendor. |
| 13. Built for the individual | Present — ironically. Platform and consumer agents are built for the individual, which is exactly why it's striking that they are the two clusters shipping the fewest of the controls above. |
The incident record is the proof this is architectural, not cosmetic. The same shape recurs: untrusted content reaches the model, the model acts as the user, and the confirmation gate doesn't fire because the manipulated model controls it.
- Perplexity Comet — "PleaseFix" (Zenity Labs, published March 2026): a malicious calendar invite hijacked the agent with no click on anything malicious — the user only asked it to handle a routine calendar task — exfiltrating local files and abusing password-manager credentials.[19] Perplexity closed the file-exfiltration path before publication. Preceded by CometJacking (one click → connected-service exfiltration; Perplexity disputed the finding)[20] and Straiker's "Inbox to Wipeout" (inbox-driven injection erasing Google Drive).[21] All three were researcher demonstrations, not attacks observed in the wild.
- Replit Agent: deleted a production database against an explicit freeze; the remediation was a planning mode, not architectural prevention.[22]
- Where the answer has been another confirmation, the same failure stays available; where it has been a structural boundary, it closes — which cannot fix an architecture in which the component being manipulated is also the one asking permission. A prompt rendered by the same software that runs the model is not a control. It is a hope with a button.
And the stakes are climbing under the category, not holding still: AI-related credential exposures rose roughly 81% year-over-year (over 1.27 million leaked secrets tied to AI services in 2025, per GitGuardian),[23] and security researchers describe a non-human-identity "governance vacuum" as these products scale faster than their controls.[24][25]
Habenula's inventory: honoring the standard
Habenula is built to this standard — and, in keeping with the standard's own spirit, we publish what isn't done yet as plainly as what is, because a trust product that hides its limits has already broken the trust. Status is against the open-source launch release.[26]
| The standard | Habenula | Status |
|---|---|---|
| 1. Deterministic gate, outside the model | Pure evaluatePolicy function; no model anywhere in the decision path | Shipped |
| 2. Model never holds credentials | Tokens encrypted at rest, resolved outside model context, never returned to the model; on the sidecar path, tool-result content doesn't re-enter the calling model either | Shipped |
| 3. Granular, resource-bound permissions | (service, verb, noun) grants with mandatory noun binding — no bare verb grants | Shipped |
| 4. Smart, low-friction proposal | Grants built as the agent works: held calls offer Allow once / Allow for 30 minutes (or a chosen duration) / Deny / Tell me more; asks at the moment it matters, stops re-asking the settled | Shipped; plan compilation (multi-step plan approved once, with declared limits) planned (with the consumer release) |
| 5. Scoped, mortal grants | One-time grants and timed grants (5 minutes to 30 days, 30 minutes recommended); no permanent "always"; every grant expires on its own | Shipped |
| 6. Human-gesture approval | Human Touch — an OS presence gesture (macOS Touch ID) in front of an affirmative grant | Shipped as a proof of concept (CLI-side presence check); hardware-bound WebAuthn signatures over the exact action planned |
| 7. Tamper-evident, user-verifiable audit | Append-only SHA-256 hash chain in the user's own database, written before execution, origin-tagged; log dump exports the chain and log verify recomputes every hash client-side, and the chain format is published so a reader can re-implement verification without running any Habenula code | Shipped — verification is append-integrity only: it proves no recorded entry was altered or dropped, not that the whole database was never regenerated |
| 8. Hard, per-service spending caps | Dollar-denominated hard ceilings enforced in the governance pipeline — per-session and per-month windows, on by default, user-set via habenula cap; a breach escalates to a confirmation that mints no grant | Partial — the caps are per-user, not per-service, and rate limiting is unbuilt; the standard's per-service dimension is not met. The one paid action shipped today is a sandbox, so the caps are exercised end to end but no real merchant is integrated |
| 9. Instant global kill | One command deletes every grant, so everything is denied, and clears pending work in a single step | Shipped — deny-all is what makes kill safe. Provider-side OAuth revocation is not part of kill today; it is planned later, as defense-in-depth |
| 10. Idle wind-down | Single session on a bounded clock; held calls expire with it, and every grant is one-time or timed | Shipped; configurable per-agent heartbeats planned |
| 11. Open + run-it-yourself | Open source under AGPL v3; ships as a container you run yourself | Shipped at launch; self-host on your own cloud account and a hosted option planned |
| 12. Model choice | Model-agnostic: Anthropic or any OpenAI-inference-API provider, behind one interface | Shipped (chosen per deployment) |
| 13. Built for the individual | An integrated runtime a person installs — no IT admin, no assembly | Shipped |
The honest read: most of the list ships at launch, and the two gaps that used to matter most have closed — the audit chain now ships with the command that verifies it, and spending caps are enforced in dollars, on by default. What remains early is stated plainly: the caps are per-user rather than per-service and rate limiting is unbuilt, Human Touch ships as a presence gesture rather than a hardware-bound signature, and plan compilation is still ahead rather than dressed up.
The point isn't that every piece is fully matured; it's that Habenula meets most of the standard at launch and says plainly which parts are early. Holding itself to the same bar it measures the category against — and publishing where each piece stands — is itself part of the standard the category is missing.
The bottom line
The consumer category has shipped real safety work — convenience-flavored: confirm on the scary actions, sandbox execution, cap spend, keep a human tapping, put privacy on-device. Users are better off for all of it.
But the bar a person actually needs is higher and different: a control that is deterministic and outside the model, keys the model never holds, a log you can verify yourself, and a kill you can pull. Measured there, every consumer product falls short in the same structural places — and the recurring incident record shows the shortfall is in the architecture, not the polish.
"The agent usually asks" is not the standard. "You stay in command, verifiably, even when the model is wrong or hijacked" is. Habenula is built to that standard — the one we would want every agent vendor measured by, including us. Read how we hold ourselves to it, then read the code.
Sources
Numbered per claim; anchors appear in the body as [n]. All sources verified July 31, 2026; the Cowork compliance-capture sources ([10][11]) re-verified August 7, 2026.
- OpenAI Help Center, "Using Ask ChatGPT sidebar and ChatGPT agent on Atlas"; ChatGPT Atlas release notes.
- OpenAI, "Continuously hardening ChatGPT Atlas against prompt injection attacks", December 2025.
- Google, "Use Gemini Agent for multi-step tasks in Gemini Apps"; Mariner fold-in: Google I/O 2026 collection.
- Google AI for Developers, "Computer use — Gemini API docs" (safety settings: require-confirmation / refuse, including financial transactions).
- Apple Newsroom, "Apple aids app development with new intelligence frameworks and advanced tools", June 8, 2026.
- Forbes (John Koetsier), "Apple Goes Agentic: Welcome To The New Siri", June 9, 2026. The "human in the loop, not autonomous" characterization is the reporter's, not Apple's.
- 9to5Mac, "Security Bite: Apple's most impressive agentic AI feature yet is hiding in the Passwords app", June 11, 2026.
- Anthropic Claude Help Center, "Use Google Workspace connectors" (account-level connector access; Gmail read + draft, no send; Calendar and Drive write access).
- Anthropic Claude Help Center, "Set up role-based permissions on Enterprise plans" (Always allow / Needs approval / Blocked tiers; org "Always allow" gate off by default).
- Anthropic Claude Help Center, "Use Claude Cowork safely" (destructive-action consent; scheduled tasks run with the computer off; through July 2026: "Cowork activity is not captured in the Compliance API at this time" — the live page now reads "Cowork via mobile and web is captured in Compliance API"), accessed July 31, 2026, re-verified August 7, 2026.
- Anthropic Claude Help Center, "Use Claude Cowork on Team and Enterprise plans" (Compliance API capture of web/mobile sessions, shipped August 2026; local desktop sessions not centrally manageable or exportable), accessed July 31, 2026, re-verified August 7, 2026.
- MintMCP, "Claude Cowork Audit Logging Gap: Why Compliance Teams Should Be Concerned", March 26, 2026.
- Anthropic, "How we contain Claude across products", May 25, 2026 ("The result was an 84% reduction in permission prompts" — Claude Code sandboxing, internal usage).
- Perplexity Help Center, "How Credits Work on Perplexity" (default $200/month, raisable in credit settings).
- Perplexity, "How We Built Security Into Computer"; "Secure sandboxes for agents", July 2026.
- Perplexity Help Center, "Scheduled Tasks in Computer" (per-task stop).
- Lindy, "Lindy Enterprise Security & Compliance Overview".
- Fyxer, "Security"; trust center.
- Zenity Labs, "PleaseFix: Zero-Click AI Agent Vulnerabilities", March 2026; write-up. Perplexity added a hard file-path boundary; Zenity confirmed the file attack non-reproducible as of February 13, 2026.
- LayerX, "CometJacking: How One Click Can Turn Perplexity's Comet AI Browser Against You", October 2025. Perplexity classified the report as not applicable / no security impact.
- Straiker STAR Labs, "From Inbox to Wipeout: Perplexity Comet's AI Browser Quietly Erasing Google Drive", December 4, 2025.
- Tom's Hardware, "AI coding platform goes rogue during code freeze and deletes entire company database", July 2025.
- GitGuardian, "The State of Secrets Sprawl 2026", March 17, 2026 (1,275,105 leaked secrets tied to AI services in 2025, up 81% from 2024).
- Cloud Security Alliance, "The Non-Human Identity Governance Vacuum: AI Agents and the Fastest-Growing Unmanaged Attack Surface", May 20, 2026.
- Microsoft Security Blog, "Least privilege for AI agents: Identity, access, and tool binding", July 16, 2026.
- Habenula capability status: verified against the open-source repository at the review date — read the code at github.com/habenula-ai/habenula-oss.
Methodology and Scope of Claims
The comparisons in this document reflect Habenula's own research and analysis, based on publicly available materials—including vendor documentation, help-center and release-note pages, product announcements, and third-party research—reviewed as of July 31, 2026. They represent our assessment of that public record, not a determination about any product's full internal capabilities, and not a statement of any vendor's private or undocumented behavior.
Where this document describes a control as "Absent," "Partial," or otherwise not present for a given product, that characterization means we did not find it documented in the public materials we reviewed as of the date above—not that the capability categorically does not exist. Product capabilities in this category change quickly; a claim accurate on the review date may no longer be accurate when you read this.
Figures, features, and incidents attributed to named third parties reflect those third parties' own published materials or the cited research as of the review date, and should be re-verified against current primary sources before being relied upon.
The thirteen-point standard, the labels applied against it, and the evaluative commentary throughout reflect Habenula's own views, framework, and methodology.
All third-party names, products, and trademarks are the property of their respective owners and are used here for identification and comparison only; their use does not imply any affiliation with, or endorsement or sponsorship by, those owners.
If you believe any statement about your product is inaccurate or out of date, contact hello@habenula.ai and we will review it.
← Habenula · Read the governance whitepaper · How it works · More from Beny's Blog