← Beny's Blog

The state of agent governance.

Personal agents can now act in your accounts. The thirteen-point bar for governing them, and where today's products — Habenula included — stand against it.

Personal AI agents crossed the line from answering to acting — they now send mail, move money, book, post, and delete across a person's real accounts. The interesting question stopped being "how good is the model" and became "does the person stay in control of what it does." The capability crossed; the governance mostly didn't. This document does three things: it states a standard for what a well-governed personal agent looks like; it credits the consumer market for the steps it has shipped toward that standard; and it measures the category against the bar — where it lands, and where it doesn't.

Habenula is then measured against the same bar, gaps included.

Scope is deliberately narrow: personal, consumer products a normal person installs and runs — the platform agents (OpenAI's ChatGPT/Atlas, Google's Gemini, Apple Intelligence + Siri) and the consumer agents (Anthropic's Claude Cowork, Perplexity's Computer/Comet, Lindy, Fyxer). It excludes developer and coding agents, enterprise control planes, and governance middleware — not because those don't matter, but because none is a product an individual installs and trusts on their own.

A note on freshness: statements about other vendors' products reflect their public documentation and third-party security research as of July 31, 2026, and this category moves fast. Each such statement carries a numbered source anchor (see Sources at the end); re-check against the vendor's current documentation before relying on it. The Methodology and Scope of Claims section at the end states exactly what these comparisons do and do not assert.

The standard: what a well-governed personal agent looks like

Thirteen properties. Each is stated as what it means for the person, not the engineer.

  1. Every action is judged by deterministic code the model cannot reach. A confirmation dialog sits "outside the model" only trivially: the model still decides whether to raise it and writes what it says, so a manipulated model can skip the prompt or misdescribe it. The real bar is stricter, and threefold. The check runs on every consequential action, on the one path it must travel — the model can't route around it. Its verdict is a pure function of the action and your standing rules — same action, same rules, same answer, deny by default — not a fresh yes/no and not the model's judgment. And the model shapes none of it: not whether the check fires, not the outcome, not what you are shown. A one-off confirmation prompt clears none of the three.
  2. The model never holds your credentials. OAuth tokens and keys are resolved and used outside the model's context. A hijacked or mistaken model cannot spend what it never held.
  3. Permissions are granular and bound to a resource. Not "email: allowed," but "send email to this domain," "read this label," "write this folder." A grant without a concrete noun is not a grant.
  4. Permission proposals are smart and low-friction — without dulling the ask. The agent asks at the moment a capability is first needed, at the scope you choose, and stops re-asking what you've already settled — while still surfacing the genuinely consequential and novel. For multi-step work, it proposes a legible plan you approve once, carrying declared limits. The goal is precision, not fewer prompts for their own sake.
  5. Grants are scoped and mortal. Allow once, or allow for a set time with a real expiry — not a standing "always" that outlives the reason you granted it. Least standing authority, by default.
  6. The highest-stakes actions can be gated behind a human gesture a program can't fake. The system supports requiring a biometric or hardware-key presence check — bound to the specific approval, so a compromised process can't click through on your behalf — for the actions you choose. A capability you turn on where it matters, not a blanket requirement.
  7. Every action is recorded in a tamper-evident log the person can verify themselves. Not vendor-side debug telemetry you're asked to trust — a cryptographically chained record you can export and check independently.
  8. Spending is bounded, hard, and per-service. A cap the agent cannot exceed, set where the money actually leaves — not just a plan-level ceiling.
  9. There is a real, instant kill switch. One action stops everything, fast and globally — ideally revoking provider access as defense-in-depth, but at minimum leaving nothing able to act.
  10. Idle agents wind down on their own. An agent you stop paying attention to should stop — a dead-man heartbeat or session expiry, not an open-ended process that runs until someone remembers to kill it.
  11. The controls are open, and the person can run them. The code that holds your keys, enforces your rules, and writes your audit trail is public — you can read it, and you can run it yourself. A control you can't inspect is asking for your trust, not handing you control.
  12. The model is yours to choose. Model-agnostic by design, because a control layer owned by the model vendor it is meant to govern is a conflict of interest, not a safety story.
  13. It's built for the individual. A product a person installs and trusts on their own — no IT administrator, no developer assembly, no fleet to manage.

Where the market is moving

The field is moving, and several of these steps make users concretely better off:

Read together, the category is converging on real lessons: confirm on the actions that scare people, default to least context, sandbox execution, cap spend, keep a human in the loop, and — Apple especially — treat privacy architecture as a feature.

Measured against the standard, the category falls short — structurally

That progress is real. Against the thirteen-point bar, though, every consumer product today lands well short — and in the same places.

The standardWhere the consumer category is (as of July 31, 2026)
1. Deterministic gate, outside the modelAbsent. The "gate" is the model's own judgment about whether to pause — the manipulable component deciding its own limits. Injection walks past the prompt because the hijacked model simply doesn't ask.
2. Model never holds credentialsAbsent. Platform and consumer agents hand the model the session or the tokens by design — Cowork is single-vendor (one company holds the model and the keys); Comet was given the active browser session.[19][20] The least-predictable component holds the keys.
3. Granular, resource-bound permissionsPartial at best. Cowork's local-folder scoping is the high-water mark (its Workspace connectors stay account-scoped);[8] no product requires a noun binding ("this recipient / this path"). Grants stay coarse.
4. Smart, low-friction proposalInverted. Confirmation prompts, not structured proposals — and tuned to reduce asking (Anthropic engineering reports an 84% cut in permission prompts from its Claude Code sandboxing[13]), which trains reflexive approval rather than precision.
5. Scoped, mortal grantsAbsent. Credential and connector grants persist until manually revoked; no one-time or timed grants with a real expiry. The closest step — write tools defaulting to per-task approval on Cowork's managed plans — is a per-action ask, not a mortal grant.[9]
6. Human-gesture approvalAbsent. Approval is a click any process can fake. Apple's "tap" is the closest instinct, but it's a screen tap, not a hardware-bound signature over the specific action.
7. Tamper-evident, user-verifiable auditAbsent. For months after launch, Cowork activity was explicitly excluded from Anthropic's audit logs, Compliance API, and data exports across all tiers, per Anthropic's own help center. The fix Anthropic shipped in August 2026 captures cloud (web/mobile) sessions in the Compliance API on Team/Enterprise; local desktop sessions remain excluded from any centrally exportable record.[10][11][12] Where any log exists elsewhere it's admin/enterprise-scoped and often omits agent-generated content. No user-verifiable chain anywhere.
8. Hard, per-service spending capsPartial. Perplexity's default $200 ceiling is the best in class;[14] caps are account/plan-level, never per-service or per-action.
9. Instant global killAbsent as a primitive. "Stop the task," "close the app," "delete the task" — no documented one-action global stop, and no provider-side token revocation.
10. Idle wind-downAbsent. Agents run until the task finishes or someone stops them; Cowork even runs scheduled tasks with no device online.[10] No dead-man expiry.
11. Open + run-it-yourselfAbsent. Closed-source, single-vendor clouds — you cannot read the code that holds your keys, and you cannot run it yourself.
12. Model choiceAbsent. Each is locked to its own model (OpenAI, Anthropic, Google, Apple); the control layer is owned by the model vendor.
13. Built for the individualPresent — ironically. Platform and consumer agents are built for the individual, which is exactly why it's striking that they are the two clusters shipping the fewest of the controls above.

The incident record is the proof this is architectural, not cosmetic. The same shape recurs: untrusted content reaches the model, the model acts as the user, and the confirmation gate doesn't fire because the manipulated model controls it.

And the stakes are climbing under the category, not holding still: AI-related credential exposures rose roughly 81% year-over-year (over 1.27 million leaked secrets tied to AI services in 2025, per GitGuardian),[23] and security researchers describe a non-human-identity "governance vacuum" as these products scale faster than their controls.[24][25]

Habenula's inventory: honoring the standard

Habenula is built to this standard — and, in keeping with the standard's own spirit, we publish what isn't done yet as plainly as what is, because a trust product that hides its limits has already broken the trust. Status is against the open-source launch release.[26]

The standardHabenulaStatus
1. Deterministic gate, outside the modelPure evaluatePolicy function; no model anywhere in the decision pathShipped
2. Model never holds credentialsTokens encrypted at rest, resolved outside model context, never returned to the model; on the sidecar path, tool-result content doesn't re-enter the calling model eitherShipped
3. Granular, resource-bound permissions(service, verb, noun) grants with mandatory noun binding — no bare verb grantsShipped
4. Smart, low-friction proposalGrants built as the agent works: held calls offer Allow once / Allow for 30 minutes (or a chosen duration) / Deny / Tell me more; asks at the moment it matters, stops re-asking the settledShipped; plan compilation (multi-step plan approved once, with declared limits) planned (with the consumer release)
5. Scoped, mortal grantsOne-time grants and timed grants (5 minutes to 30 days, 30 minutes recommended); no permanent "always"; every grant expires on its ownShipped
6. Human-gesture approvalHuman Touch — an OS presence gesture (macOS Touch ID) in front of an affirmative grantShipped as a proof of concept (CLI-side presence check); hardware-bound WebAuthn signatures over the exact action planned
7. Tamper-evident, user-verifiable auditAppend-only SHA-256 hash chain in the user's own database, written before execution, origin-tagged; log dump exports the chain and log verify recomputes every hash client-side, and the chain format is published so a reader can re-implement verification without running any Habenula codeShipped — verification is append-integrity only: it proves no recorded entry was altered or dropped, not that the whole database was never regenerated
8. Hard, per-service spending capsDollar-denominated hard ceilings enforced in the governance pipeline — per-session and per-month windows, on by default, user-set via habenula cap; a breach escalates to a confirmation that mints no grantPartial — the caps are per-user, not per-service, and rate limiting is unbuilt; the standard's per-service dimension is not met. The one paid action shipped today is a sandbox, so the caps are exercised end to end but no real merchant is integrated
9. Instant global killOne command deletes every grant, so everything is denied, and clears pending work in a single stepShipped — deny-all is what makes kill safe. Provider-side OAuth revocation is not part of kill today; it is planned later, as defense-in-depth
10. Idle wind-downSingle session on a bounded clock; held calls expire with it, and every grant is one-time or timedShipped; configurable per-agent heartbeats planned
11. Open + run-it-yourselfOpen source under AGPL v3; ships as a container you run yourselfShipped at launch; self-host on your own cloud account and a hosted option planned
12. Model choiceModel-agnostic: Anthropic or any OpenAI-inference-API provider, behind one interfaceShipped (chosen per deployment)
13. Built for the individualAn integrated runtime a person installs — no IT admin, no assemblyShipped

The honest read: most of the list ships at launch, and the two gaps that used to matter most have closed — the audit chain now ships with the command that verifies it, and spending caps are enforced in dollars, on by default. What remains early is stated plainly: the caps are per-user rather than per-service and rate limiting is unbuilt, Human Touch ships as a presence gesture rather than a hardware-bound signature, and plan compilation is still ahead rather than dressed up.

The point isn't that every piece is fully matured; it's that Habenula meets most of the standard at launch and says plainly which parts are early. Holding itself to the same bar it measures the category against — and publishing where each piece stands — is itself part of the standard the category is missing.

The bottom line

The consumer category has shipped real safety work — convenience-flavored: confirm on the scary actions, sandbox execution, cap spend, keep a human tapping, put privacy on-device. Users are better off for all of it.

But the bar a person actually needs is higher and different: a control that is deterministic and outside the model, keys the model never holds, a log you can verify yourself, and a kill you can pull. Measured there, every consumer product falls short in the same structural places — and the recurring incident record shows the shortfall is in the architecture, not the polish.

"The agent usually asks" is not the standard. "You stay in command, verifiably, even when the model is wrong or hijacked" is. Habenula is built to that standard — the one we would want every agent vendor measured by, including us. Read how we hold ourselves to it, then read the code.

Sources

Numbered per claim; anchors appear in the body as [n]. All sources verified July 31, 2026; the Cowork compliance-capture sources ([10][11]) re-verified August 7, 2026.

  1. OpenAI Help Center, "Using Ask ChatGPT sidebar and ChatGPT agent on Atlas"; ChatGPT Atlas release notes.
  2. OpenAI, "Continuously hardening ChatGPT Atlas against prompt injection attacks", December 2025.
  3. Google, "Use Gemini Agent for multi-step tasks in Gemini Apps"; Mariner fold-in: Google I/O 2026 collection.
  4. Google AI for Developers, "Computer use — Gemini API docs" (safety settings: require-confirmation / refuse, including financial transactions).
  5. Apple Newsroom, "Apple aids app development with new intelligence frameworks and advanced tools", June 8, 2026.
  6. Forbes (John Koetsier), "Apple Goes Agentic: Welcome To The New Siri", June 9, 2026. The "human in the loop, not autonomous" characterization is the reporter's, not Apple's.
  7. 9to5Mac, "Security Bite: Apple's most impressive agentic AI feature yet is hiding in the Passwords app", June 11, 2026.
  8. Anthropic Claude Help Center, "Use Google Workspace connectors" (account-level connector access; Gmail read + draft, no send; Calendar and Drive write access).
  9. Anthropic Claude Help Center, "Set up role-based permissions on Enterprise plans" (Always allow / Needs approval / Blocked tiers; org "Always allow" gate off by default).
  10. Anthropic Claude Help Center, "Use Claude Cowork safely" (destructive-action consent; scheduled tasks run with the computer off; through July 2026: "Cowork activity is not captured in the Compliance API at this time" — the live page now reads "Cowork via mobile and web is captured in Compliance API"), accessed July 31, 2026, re-verified August 7, 2026.
  11. Anthropic Claude Help Center, "Use Claude Cowork on Team and Enterprise plans" (Compliance API capture of web/mobile sessions, shipped August 2026; local desktop sessions not centrally manageable or exportable), accessed July 31, 2026, re-verified August 7, 2026.
  12. MintMCP, "Claude Cowork Audit Logging Gap: Why Compliance Teams Should Be Concerned", March 26, 2026.
  13. Anthropic, "How we contain Claude across products", May 25, 2026 ("The result was an 84% reduction in permission prompts" — Claude Code sandboxing, internal usage).
  14. Perplexity Help Center, "How Credits Work on Perplexity" (default $200/month, raisable in credit settings).
  15. Perplexity, "How We Built Security Into Computer"; "Secure sandboxes for agents", July 2026.
  16. Perplexity Help Center, "Scheduled Tasks in Computer" (per-task stop).
  17. Lindy, "Lindy Enterprise Security & Compliance Overview".
  18. Fyxer, "Security"; trust center.
  19. Zenity Labs, "PleaseFix: Zero-Click AI Agent Vulnerabilities", March 2026; write-up. Perplexity added a hard file-path boundary; Zenity confirmed the file attack non-reproducible as of February 13, 2026.
  20. LayerX, "CometJacking: How One Click Can Turn Perplexity's Comet AI Browser Against You", October 2025. Perplexity classified the report as not applicable / no security impact.
  21. Straiker STAR Labs, "From Inbox to Wipeout: Perplexity Comet's AI Browser Quietly Erasing Google Drive", December 4, 2025.
  22. Tom's Hardware, "AI coding platform goes rogue during code freeze and deletes entire company database", July 2025.
  23. GitGuardian, "The State of Secrets Sprawl 2026", March 17, 2026 (1,275,105 leaked secrets tied to AI services in 2025, up 81% from 2024).
  24. Cloud Security Alliance, "The Non-Human Identity Governance Vacuum: AI Agents and the Fastest-Growing Unmanaged Attack Surface", May 20, 2026.
  25. Microsoft Security Blog, "Least privilege for AI agents: Identity, access, and tool binding", July 16, 2026.
  26. Habenula capability status: verified against the open-source repository at the review date — read the code at github.com/habenula-ai/habenula-oss.

Methodology and Scope of Claims

The comparisons in this document reflect Habenula's own research and analysis, based on publicly available materials—including vendor documentation, help-center and release-note pages, product announcements, and third-party research—reviewed as of July 31, 2026. They represent our assessment of that public record, not a determination about any product's full internal capabilities, and not a statement of any vendor's private or undocumented behavior.

Where this document describes a control as "Absent," "Partial," or otherwise not present for a given product, that characterization means we did not find it documented in the public materials we reviewed as of the date above—not that the capability categorically does not exist. Product capabilities in this category change quickly; a claim accurate on the review date may no longer be accurate when you read this.

Figures, features, and incidents attributed to named third parties reflect those third parties' own published materials or the cited research as of the review date, and should be re-verified against current primary sources before being relied upon.

The thirteen-point standard, the labels applied against it, and the evaluative commentary throughout reflect Habenula's own views, framework, and methodology.

All third-party names, products, and trademarks are the property of their respective owners and are used here for identification and comparison only; their use does not imply any affiliation with, or endorsement or sponsorship by, those owners.

If you believe any statement about your product is inaccurate or out of date, contact hello@habenula.ai and we will review it.

← Habenula  ·  Read the governance whitepaper  ·  How it works  ·  More from Beny's Blog