ArticlesBuyer's map
Chatbot, copilot, agent, operating system: a buyer's map
Chatbot, copilot, agent and operating system are four different purchases, not four words for the same thing. This map places each one on six things you can check before you sign: what you install, who runs the model, what it remembers, what it may do on its own, what it costs to run, and the way it breaks. Use it on the product in the tab next to this one.
Copilot vs agent vs chatbot: what the words actually name
The four words do not sit on a ladder from small to big. They name four different deals with four different bills and four different ways of going wrong. A vendor may use any of them for any product, so the word on the page is a weak signal and you should treat it as one.
Here is the short form. A chatbot is a box you type into and it types back. A copilot lives inside a tool you already use and it makes the next move in that tool cheap. An agent takes a goal, runs a loop of its own, and stops when it thinks it is done. An operating system is a set of named jobs, a shared memory, and a rule about what has to stay true, on top of a model that someone else runs.
None of that is a ranking. A chatbot is the right buy for a lot of work. The point of the map is that you can tell which one you are being sold in about twenty minutes, and then you can ask the right questions instead of the general ones.
The six checks below are the map. They are dull on purpose. Each one has an answer that a vendor can give you in a sentence, and a way for you to see the answer for yourself if they will not.
- What you install: a tab, a plugin, a service, or a config on a host you already have.
- Who runs the model: them, you, or your host.
- What it remembers: nothing, the thread, or a store you can read.
- What it may do alone: draft, act on a short list, or act on anything.
- What it costs to run: a seat, your tokens, or both.
- How it fails: the wrong answer, the wrong move, or the wrong move at scale.
Chatbot: a box that answers
What you install is a tab, or a widget on your own site. There is no state on your side and often no admin to speak of. This is the whole appeal: you can have one by lunch.
Who runs the model is the vendor, almost always. That matters more than it sounds. If they run it, they set the model, they set the limits, and they can change both without asking you. Your prompt and your text go to their service to be answered.
What it remembers is the thread, and only while the thread is open. Some keep a longer history per user. Ask the plain question: if I close this and come back on Monday, what does it still know about me? A vendor who cannot answer that in one sentence has not built it.
What it may do alone is answer. That is the safety of the category. It cannot send your invoice to the wrong client, because it cannot send anything.
What it costs is a seat or a bundle of messages, and the number on the invoice is close to the number you budgeted. That is a real advantage of the deal and you should count it.
How it fails is the confident wrong answer, given straight to a customer with your name at the top. The damage is not technical. It is that you now own a sentence you did not write. The fix people reach for is more grounding and more rules, and there is a point where you have paid for a worse copilot.
Copilot: help inside the tool you already use
What you install is a plugin, an add-on, or a feature switch inside a tool you already pay for. The tool is the point. A copilot with no host is a chatbot with a different price.
Who runs the model is the vendor of that tool, and you rarely get a choice. Read where the text goes. Your draft, your sheet, your code and your mail are the input, and the input goes wherever their service is.
What it remembers is what is on the screen and what the tool already knew. This is the whole trick of a copilot and it is why the good ones feel so much better than the chatbot in the other tab: the context is free, because it is already open.
What it may do alone is edit the thing in front of you, with an undo. It suggests, you keep or drop it. Some now write to other systems too, and the moment they do, you are buying an agent and the word on the box has not caught up.
What it costs is a seat, added to a seat you already pay. That is why a copilot is the easiest sale in this list and the hardest to measure. You will not see its cost as a line of its own.
How it fails is drift. It is right most of the time, so you stop reading closely, and the one in twenty that is wrong goes in unread. Nobody notices for a month. The tradeoff is real and it is the price of the category: a tool that is useful enough to trust is a tool you will stop checking.
Agent: it runs its own loop
What you install is a service with keys. It needs access to the systems it acts on, and that access is the whole install. If setup is only a sign-in, it is a copilot with a new name.
Who runs the model is either the vendor, on their key, or you, on yours. Find out, because the answer changes the bill by an order of magnitude in either direction, and it also decides whose limits you hit when the loop gets long.
What it remembers is the run. A goal, the steps it took, and the results it got back. Most keep that for the length of the run and drop it. If you want it next week, ask where it is stored and whether you can read it without the vendor.
What it may do alone is the only question that matters here, and the honest answer is: as much as its tools let it. An agent is a loop with a list of tools. Its reach is the sum of that list, not the sum of what it is meant to do. The OWASP Top 10 for LLM applications names this one directly as excessive agency, and splits it into three causes: too much function, too many permissions, and too much autonomy.
What it costs is tokens, and the number moves with the work, not with your headcount. One loop can be twenty model calls or two hundred. Any agent you buy should be able to tell you what one run cost. If nobody can answer that, nobody is watching it.
How it fails is the wrong move made quickly, and then made again. The loop does not know it went wrong, so it keeps going. Wrong twice is a bug and wrong four hundred times is an incident, and the gap between them is a few minutes.
Operating system: named jobs, shared memory, and a gate
What you install is the odd one. There is no new app. You add a config to a host you already use, and a set of named jobs shows up inside it. Nothing new on the dock, nothing new to log into.
Who runs the model is the host you already pay. That is the defining trait of the category and it is easy to check: if the seller can keep serving you while your own model bill is zero, they are running it, and they are not this.
What it remembers is meant to outlast the session. A line of work has things that stay true: who the reader is, what the offer is, what you refuse to say. An OS that forgets the constraint between two sessions is an agent with a folder.
What it may do alone is draft and propose, and then stop at a line you drew. The gate is the product. Work that publishes, sends, spends or wipes should wait for a named person to say yes, and the record of that yes should exist afterward.
What it costs is a flat price for the system and your own tokens for the work. You pay twice, in two places, and you should know that going in. The upside of that shape is that the seller has no reason to make the model chatty.
How it fails is quieter than the others and worse to unpick. The system holds a constraint that was true in March and is not true now, and every piece of work it produces is neatly, consistently wrong. The failure of an OS is stale truth, not a stray action.
Who runs the model, and why it shows up on your bill
This is the cheapest question on the list and the one people skip. There are three answers and they are easy to tell apart. The vendor runs it on their key. You run it on your key, given to them. Or your host runs it, and nobody hands a key to anyone.
If the vendor runs it, your price is flat and their margin moves with your use. That means they have a reason to keep answers short and a reason to move you to a cheaper model when costs bite. Neither is wrong. It is the shape of the deal and you should know it is there.
If you hand over a key, you get the choice of model and you also get the bill for every retry you never saw. Ask what happens on a loop that does not end. Ask whether there is a cap, and whether the cap is yours to set.
If your host runs it, the model is whatever your host runs today, and the seller is a config on the other side of a protocol. The MCP documentation is blunt that the protocol handles context exchange and says nothing about how the application uses a model, which is why this arrangement can exist at all. The host keeps the conversation, the server offers tools, and the tokens are spent where they were always spent.
Ask it out loud, in these words: if I stopped paying my model provider tomorrow, does your product still work? The answer places the product on this axis with no room to wriggle.
What it remembers: session, run, or store
Memory is where the four categories separate most clearly, and where marketing pages are vaguest. There are three honest shapes and you can name the one you are being sold.
Session memory lasts as long as the window is open. It is not a feature, it is the absence of one, and there is nothing wrong with that if the work is one question long.
Run memory lasts as long as one job. The agent knows what it did four steps ago because the steps are in the prompt. It ends when the run ends. This is enough for a loop and not enough for a line of work.
Store memory is the only one that survives you closing the laptop. It has a shape, it has a place, and someone can read it. The questions to ask are: where is it, can I read it without you, can I delete one thing from it, and what happens to it if I cancel.
A quick test in a chat window: tell it a fact that matters, close everything, come back the next day and ask about it. If it knows, you have store memory. If it apologises, you have session memory and a marketing page that said otherwise.
- Session: dies with the window. Fine for a single question.
- Run: lives for one job. Fine for a loop with an end.
- Store: survives. Ask where it is and who can read it.
- The test: tell it something on Monday, ask on Tuesday.
What it is allowed to do alone
Every product in this map has a line between what it does on its own and what it asks you first. Find the line. It is the single most useful thing you can learn in a demo, and it is almost never on the pricing page.
The OWASP Top 10 for LLM applications is the plainest public writing on this, and it is free to read. Its entry on excessive agency names three separate causes, which is useful because they have three separate fixes. Too much function: the tool can do more than the job needs. Too many permissions: the account behind the tool can reach more than the job needs. Too much autonomy: nobody has to say yes before something that matters happens.
That split is worth carrying into a call, because a vendor will often answer one and leave the others. A read-only integration with an admin key is still a bad day waiting. A tightly scoped key attached to a tool that can do anything is the same bad day from the other side.
The approval question is the third one and it is the one you can check yourself. Ask for a demo of a rejected action. Not an approved one, a rejected one. You want to see what the product does when the person says no: does it stop, does it retry, does it half finish, and is there a record of who said no and when.
Public guidance points the same way. The NIST AI Risk Management Framework, which is voluntary and free, is organised around four functions: govern, map, measure and manage. None of them is a prompt. That is a slow way of saying the interesting part of an AI purchase is the part where a person is still in the loop on the things that matter, and where someone can tell you afterwards what happened.
What it costs to run, beyond the sticker
There are three bills and a product can have any two of them. The seat or the licence. The tokens. The people who now have to check the output.
The third one is the one nobody quotes and it is often the biggest. A copilot that writes a draft you must read line by line has moved the work, not removed it. That can still be a good trade. It is not the trade on the slide.
Token cost is the one that surprises finance, because it moves with volume and not with headcount. Ask for the shape of it: per message, per run, per document, per month. Then ask for the worst case, which is the number that shows up in the month you actually rely on the thing.
The honest thing to say here is that we cannot tell you what any of this costs at your volume, and nor can anyone else before you run it. What you can do is insist that the product can show you its own usage, broken down by the unit of work you care about. A product that cannot show you that is a product nobody has had to justify yet.
How to place a product in twenty minutes
Run this on the thing you are looking at right now. Each step has something you see, and the thing you see is the answer.
One: open the setup page and count what you install. A URL, a browser add-on, a service account with keys, or a config file on a tool you already run. Four different answers, four different categories.
Two: look for a field that asks for your model key. If there is one, you run the model. If there is not, and the product still answers, they do. If the setup is a config on a host you already pay for, the host does.
Three: tell it one durable fact. Not a task, a fact. Close the session. Come back tomorrow and ask. You now know whether memory is session, run, or store, and you know it better than the docs do.
Four: ask for the list of things it can do without you. If the answer is a paragraph rather than a list, that is the answer. A product with a real gate has a list, because someone had to write the gate.
Five: ask to see a rejection. Watch what happens after the no.
Six: ask what one unit of work costs and what the worst unit last month cost. Silence here is information too.
At the end of those six you will have placed it, and you will have six answers to compare against the next one. That is worth more than the category name, which was always the least reliable part of the page.
- Say this: if I stopped paying my model provider tomorrow, does your product still work?
- Say this: show me a run where the person clicked no, and show me the record of it.
- Say this: what did one unit of work cost last month, and what did the worst one cost?
- Say this: list what it can do without asking me. Not describe. List.
- Say this: if I cancel, what happens to what it has learned about my company?
Where this map breaks
Products move between categories without changing their name. A copilot that gains the ability to write to a second system has become an agent in every way that matters to you, and the release note will call it an improvement. Re-run the six checks after a big update, not just before you sign.
The map also says nothing about quality. A well built chatbot beats a badly built operating system on almost any day, on almost any work. The categories tell you what kind of risk and what kind of bill you are taking on. They do not tell you whether the thing is any good, and nothing in a category name ever will.
And the boundaries are genuinely soft in one place: between an agent and an operating system. A single agent with a durable store and a hard approval gate is doing most of what an OS does. The difference is whether there is more than one named job and something that holds the constraint across them. If you only have one job, you may not need the other thing, and paying for it is paying for structure you will not use.
Last limit, and it is ours: these six checks are the ones we found useful. They are not a standard, nobody audits them, and a determined vendor can answer all six well and still ship something you do not want. They narrow the field. They do not pick for you.
Where Agentik sits on its own map
Agentik {OS} is in the fourth column, and it is fair to hold us to the same six checks. What you install is a config on a host you already pay: Claude, Claude Code, Cursor, ChatGPT, Codex or Hermes. Who runs the model is that host. We never buy tokens, which also means we cannot promise you a bill we do not control.
What it remembers is per project, so one OS can hold a separate store for each client. What it may do alone is draft and propose, and work that publishes, sends, spends or resets waits for approve. What it costs is $19.99 once for a single official OS, or $19 a month for three, $99 for ten and $199 for unlimited, plus your own model bill. The official systems today are Content OS, Growth OS, Librarian OS and Builder OS.
The failure we hit is the one named above: a constraint that was true when it was written and is not true now. There is no clever fix for it in the product. Someone has to reread what the system believes. If that sounds like work, it is, and it is the part of this category nobody puts on a slide. /docs/mcp has the install side and /glossary has the 647 terms if a word in this piece was new.
| Check | Chatbot | Copilot | Agent | Operating system |
|---|---|---|---|---|
| What you install | A tab or a widget | A plugin in a tool you pay for | A service with keys to your systems | A config on a host you already run |
| Who runs the model | The vendor | The tool's vendor | The vendor, or you on your key | Your host |
| What it remembers | The open thread | What is on screen | The run | A store per project, across sessions |
| What it may do alone | Answer | Edit what is in front of you, with undo | Whatever its tool list allows | Draft and propose, then wait for a yes |
| What it costs to run | A seat or a message bundle | A seat on top of a seat | Tokens that move with the work | A flat price plus your own tokens |
| The failure you will hit | A confident wrong answer to a customer | Drift: right enough that you stop reading | The wrong move, repeated fast | A stale constraint, applied consistently |
Sources
Questions
What is the difference between a copilot and an agent?
A copilot acts inside the tool you have open and waits for you to keep or drop each suggestion. An agent takes a goal, runs its own loop across systems, and stops when it decides it is finished.
Is a chatbot just a worse agent?
No, it is a different purchase with a different risk. A chatbot cannot act, so its worst case is a wrong answer, while an agent's worst case is a wrong action repeated before anyone looks.
How do I tell which category a product is in?
Count what you install, then find out who holds the model key. Those two answers place almost every product, and the setup page usually gives you both.
What does an AI operating system add over an agent?
More than one named job, a store that survives the session, and an approval line for anything that leaves the building. If you only have one job to automate, you may not need it.
Who pays for the tokens?
Whoever holds the key, and that is worth settling before you sign. If the vendor holds it your price is flat and their limits are yours, and if your host holds it the bill moves with your own use.
Does the category tell me if the product is good?
No. It tells you what kind of bill and what kind of failure you are taking on, and a well built product in a simpler category beats a badly built one in a grander category most days.