What everyone means by “agent.”
A plain-language guide to what an AI agent actually is, how it works, and why the difference between talking and acting matters to you.
- A chatbot talks; an agent acts. It doesn't just tell you how to get the refund — it opens your email, fills the form, and files it.
- Underneath, it is a loop: the AI proposes one small step, a tool carries that step out in your real accounts, the result comes back, and it goes around again until the job is done.
- It is the first software you instruct rather than operate — you brief it in plain language and it improvises the how. That is the power, and it is why it can act on a wrong guess with complete confidence.
- Because it both acts and guesses, it needs controls: a clear view of what it is doing, a say over the steps that carry weight, and a way to stop it at once. Building those is what Habenula does.
If you have used AI at all, you have probably used a chatbot. You type a question, it types an answer. Maybe it drafted an email for you, or explained something your doctor said, or helped with homework at ten at night. Whatever it did, it did it in one place: a text box. It talked. You read.
Lately a different word has been showing up — in product launches, in headlines, in the settings of apps you already use. Agent. Every AI company is suddenly announcing one, and the word is doing a lot of work without much explanation attached.
Here is the explanation. An agent is AI that doesn't just tell you things. It does things — on your behalf, in your actual accounts, out in the world. That one difference changes almost everything about how you should think about it, and this post is a plain-language tour of why.
A chatbot talks — an agent acts
Ask a chatbot to help you get a refund from an airline, and it will do something genuinely useful: it will explain the airline's policy, draft the complaint email, tell you which form to look for. Then it stops, because that is the edge of what it can reach. The words sit in the chat window, and every remaining step is yours. You open your email. You paste the draft. You find the form. You press send.
A chatbot is a very good advisor trapped behind glass. Nothing it says leaves the window unless you carry it out by hand.
An agent is what happens when the glass is removed. Ask an agent for that refund and it opens your email to find the booking confirmation, pulls the flight details out of it, fills in the airline's form, writes the complaint, and submits it — then tells you what it did. The thing it produces is not a paragraph for you to act on. It is the acted-on thing itself: the sent message, the filed request, the booked ticket, the moved meeting.
That is the whole distinction, and it is worth sitting with for a second, because it is easy to hear “agent” as “a better chatbot.” It isn't a better chatbot. It is the same kind of intelligence given hands.
Under the hood, it is a loop, not a mind
“Given hands” sounds mysterious, so let's take the mystery out of it, because the mechanics are surprisingly ordinary.
The AI at the center of an agent is a language model — software that is remarkably good at reading and producing text, and that is all it can do. It cannot click anything. It cannot open your inbox. Left alone, it is a chatbot forever.
What turns it into an agent is a set of tools: small, boring pieces of ordinary software that can each do one real thing — search the web, read an email, add a calendar event, fill a form, make a payment. The model cannot press any buttons, but it can write a note saying which button should be pressed. The tools press it.
Put those together and you get a loop, and the loop is the agent:
You state a goal. “Find my flight confirmation and get me checked in.”
The model proposes a step. Not the whole job — one step. Search the inbox for the airline's name.
A tool carries it out. The search actually runs, in your real inbox.
The result comes back to the model. Here are the three emails that matched.
And around again. Read the newest one. Extract the confirmation code. Open the check-in page. Fill the code in. Repeat — until the goal is done or the agent gets stuck and asks you.
Each pass through the loop is small and understandable. No single step is magic. The impressive thing, the thing that feels like watching a person work, is the accumulation: dozens of little propose-act-observe passes, chained together, adding up to “it handled the whole thing while I made coffee.”
When you strip the marketing away, that loop is the definition. An agent is a language model, a set of tools, and a loop that runs until the job is done.
It is the first software you instruct rather than operate
Here is what makes agents genuinely new — not just faster or smarter, but a different kind of thing than any software before them.
Every program you have ever used was operated. The spreadsheet, the banking app, the booking site: each does exactly what it was built to do, and you drive it, click by click. It never surprises you, because it can't. Everything it will ever do was decided by the people who wrote it.
An agent is instructed. You tell it what you want in plain language, the way you'd brief a person, and it improvises the how. Nobody wrote a program called “get you a refund from the airline.” The agent composes that program on the fly, out of loop steps, differently every time. Ask twice and you may get two different paths to the same result.
That is the power, and it is also the strangeness, and you should hold both. The power: software that can handle the endless, messy, one-off tasks no developer would ever build a dedicated app for — which is to say, most of life's actual chores. The strangeness: for the first time, the software's next move is a judgment call rather than a lookup.
The model is making its best guess at what you meant and how to get there. It is very often right. It can also misread you, or misread the world, and be wrong with complete confidence — not because it is broken, but because guessing is what it is.
Old software failed by stopping. Agents can fail by proceeding — doing something, sincerely, that is not what you wanted.
The same loop that books your flight can touch everything else
Now connect the two halves of this post. An agent acts, and an agent guesses. Those two facts together are why agents deserve a different level of attention than any app you have installed before.
To be useful, an agent needs reach into your real life: the inbox, the calendar, the accounts, sometimes the card on file. And the loop that lets it send your check-in also lets it send anything else. Read one email, read the rest. Spend eleven dollars, spend eleven hundred. The verbs an agent uses on your behalf are exactly the verbs you would not hand a stranger: send, spend, post, delete.
With a chatbot, a wrong answer costs you nothing but the reading of it. It is text in a box; you shrug and close the window. When an agent guesses wrong, there is no window to close. The email went out under your name. The meeting was canceled for real. The money moved. A mistake stops being a bad answer and becomes an event — one that happened to actual people, on your actual accounts, with your actual signature on it.
None of this is a reason to avoid agents, any more than the risk of a crash is a reason to avoid cars. It is a reason to insist on the three things that make that kind of power safe to hand over: a clear view of what the agent is doing, a say over the actions that carry real weight before they happen, and a way to stop it the instant something looks wrong.
The measure of a good agent setup is not only how much it can do for you. It is how clearly you can see what it is doing, and how surely you can stop it.
Watch one work
There is a difference between knowing the definition of swimming and watching someone swim. A governed loop starts with a plan, not a step (coming with the consumer release). Ask it to book a flight and the model proposes the best plan, in order: check the flights, buy the right one within your limits, then file the ticket and put the flight on your calendar. It cannot buy the right flight before it has checked them. You approve the plan once; each step then runs in turn, its results feeding the next, and stops for a quick yes only where a step needs one — here, the purchase.
We build Habenula, the control layer in that second picture: the part that keeps you in command of the loop rather than inside it. The 10-second loop on our homepage is a guided walkthrough in a safe sandbox, where you grant each step, read the record it leaves, and hold the stop yourself.
← Habenula · Try the loop · How it works · More from Beny's Blog