Chatbot
answers- Understands the question
- Writes a friendly reply
- Asks for the item number and address
- ···
A chatbot hands you a text. An agent hands you a result: read, checked, booked, sent, logged. Here is how one is built, which tasks actually hold up today and what you need to settle before you start.
An example run. The last step is the important one: an agent that cannot hand over is not an agent, it is a liability.
Same email, two systems. The difference is not the quality of the prose, it is the question of who did the work in the end.
“Can you send me a quote for 500 units?”
Agents rarely fail at the language model. They fail at one of these seven parts. Click through them and you will see early where your own plan will get stuck.
What counts as done?
An agent needs a result you can tick off, not a topic. “Handle customer enquiries” is not a goal. “Answer every delivery-time enquiry within 15 minutes with the confirmed date” is one.
Nobody on the team can say in one sentence when the agent is finished.
Almost every failed agent project started at the wrong level. Drag the dial and watch what moves with it: the work, the risk and what you need to have settled first.
The agent finishes it, you release it.
The agent completes the task in full, but the result waits for your click. This is where almost every good project starts.
We almost always start at L1 and only move up once four weeks of operating numbers justify it.
Not futurology. These are the cuts that run reliably in small and mid-sized companies, because they have a clear goal, digital tools and a handover point.
Read enquiries, classify them, answer or pass them on with context to the right team.
Understand the request, calculate the line items, build the document, send it, open the case.
Read documents, match them against order and delivery note, post them, report discrepancies.
Check availability, propose slots, confirm, send the invitation and the prep material.
Reconcile information across systems, find duplicates, report gaps.
Follow sources, filter what matters, summarise it briefly, route it to the right person.
Six questions, answered honestly. The scoring runs in your browser, we never see any of it.
Four patterns we see again and again. None of them is about the language model.
“The agent runs our sales” fails reliably. The cut is so wide that nobody can say when the agent has done its job well.
One task, one trigger, one result. You widen it once the first number holds.
Without a clean handover point the agent keeps guessing when in doubt. Automation turns into quiet chaos that only surfaces at the customer.
Define the stop conditions first, not last. And put a named owner behind them.
If nobody defines what “done” means, neither quality nor savings can be evidenced. The project ends in opinions.
Count from day one: cases closed, escalation rate, rework, hours saved.
An agent with full write access to the ERP is not a productivity project, it is a security incident on a delay.
Smallest possible permission per tool, hard limits in code and a kill switch everyone on the team knows.
An agent is not automatically high-risk AI. But it triggers obligations faster than a chatbot, because it acts and reaches the outside world. This is what affects most companies.
Anyone deploying AI must make sure the people working with it understand it. For agents that means something concrete: the team knows what the agent may do, how an error looks and how to stop it.
Where your agent communicates with people, they must be able to tell that AI is involved. This deadline was expressly not postponed.
A transition period for machine-readable marking of AI-generated content runs until 2 December 2026. Relevant as soon as your agent produces text, images or audio that gets published.
If your agent has a hand in decisions about job applications, creditworthiness or access to services, the full obligations apply: risk management, data quality, documentation, human oversight under Art. 14 and logging under Art. 12.
Independent of the AI Act: decisions with legal effect must not be made by automation alone. That is exactly what part 06 is for, not as a convenience feature but as evidence.
This is an orientation, not legal advice. The official text of the regulation governs.
Not a pilot that fizzles out. Four steps, and after each one you keep something, even if we stop there.
We watch the process once, together with the people who run it. Afterwards it is on paper: which step takes how long, where decisions are made and what the agent must not touch.
A process map, an honest feasibility call and a cost range.
One agent, one task, level L1. Tools with the smallest possible permissions, guardrails in code, logging from the start.
A running agent in your environment, delivering drafts for release.
Four weeks of real operation on real cases. We count closed cases, escalations, rework and hours saved, instead of debating impressions.
Numbers that turn the decision about the next level into more than a hunch.
Only once the numbers hold does it continue: raise the autonomy level or add the next task. Never both at once.
A system that grows with your experience, not with your nerve.
Send us a process that costs your team hours every week. We will tell you whether an agent is worth it, which level we would start at and what range it sits in. If doing it by hand is cheaper, we say that too.
Describe your task →