Everything Is an Agent Now
A year or two ago, the same products were called chatbots, virtual assistants, conversational AI or intelligent IVR. In 2026 most of them have been renamed. "Agent" is the word that sells, because it promises something the older tools never did: AI that does the work, rather than AI that talks about the work.
Some of that promise is real. In March 2025 Gartner predicted that by 2029, agentic AI will autonomously resolve 80% of common customer service issues without human intervention, with a 30% cut in operational costs. But in June 2025 it also warned about the gap between the label and the product. Agent washing is the rebranding of chatbots, assistants and older automation as agents without giving them any real ability to act. Gartner estimated that only about 130 of the thousands of vendors claiming agentic AI actually offered it.
For a business owner that creates a practical problem. Every demo sounds good. Every voice is natural. The difference only shows up when a real customer asks for something to be done, and by then you have signed. So this guide is about testing it first.
Three Products Wearing One Label
Before you test anything, it helps to know what you are sorting products into. There are three categories on the market, and all three are regularly sold as "AI agents".
None of these is bad. An assistant that writes up every call is genuinely useful, and we have written about AI notes while you talk for that reason. The problem is buying a chatbot or an assistant when what you need is the third column, a product that picks up the 6pm call and actually books the job. That is what the rest of this guide helps you check.
The Five Call Test
This takes about fifteen minutes. Ask the vendor to set the product up on a test number with a few of your real details: your services, some sample prices, your hours and a test calendar. Then ring it yourself, five times, with a script. The trick is to ask it to do things, not just to tell you things.
- 1
The easy question
"What time do you close on Saturday?" Any chatbot will pass. This is your baseline for voice quality and speed, nothing more.
- 2
The booking
"Can I book in for Thursday afternoon?" Listen for whether it offers real times from the calendar, confirms the details back, and says it is sending a text. A chatbot will offer a link.
- 3
The change, with a twist
Ring back from the same number: "Actually, can we make that Friday, and add a second item?" A real agent finds the booking, moves it and updates it. A chatbot starts again from scratch.
- 4
The question it cannot know
Ask something that is not in its knowledge, such as a product you do not sell. The right answer is "I am not sure, let me take a message", not a confident invention.
- 5
"Can I talk to a person?"
It should transfer you, with your details passed along, or book a callback at a stated time. Being kept talking to the AI is a fail.
Now check the finish line
Hang up and open the calendar, the text messages on your phone and the call log. Is Thursday booked, then moved to Friday? Did you receive a confirmation? Is there a summary of each call with an outcome? If the calendar is untouched, it does not matter how good the conversation sounded. It was a chatbot.
If you are comparing several AI phone products at once, our longer ten call test for AI phone providers covers accents, latency and call quality as well. The five calls above are the shortcut for one question only: can it finish the job?
Three Things to See Behind the Demo
A good salesperson can make any product look operational in a live demo. So after the test calls, ask to see three screens. A genuine operational agent will have all three and the vendor will be happy to show them.
The tool list
Which systems the agent is connected to, and what it is allowed to do in each: read only, or create and change. "It integrates with everything" is not an answer. A specific list is.
The action log
A record of what the agent did on each call: the booking it created, the text it sent, the message it passed on. If you cannot see what it did, you cannot check it or fix it.
The handover rules
Where you set when the agent passes a call to a person, who it goes to, and what happens after hours. These should be settings you control, not a promise from the vendor.
It is also worth asking where the product runs and who is in the chain between you and the AI model. Some "agents" are a thin layer over three or four other companies' services, each with its own outage and data handling. Our piece on how many companies sit between you and the model explains why that matters for both reliability and privacy.
The 24 Point Scorecard
Score each product on eight checks, from 0 (no) to 3 (yes, clearly, and you saw it). It takes five minutes after the test calls and turns a pile of impressions into a number you can compare.
| # | Check | What earns a 3 |
|---|---|---|
| 1 | Books into a real calendar | The test booking appeared in the calendar during the call. |
| 2 | Changes an existing booking | Found your booking on call three and moved it, without creating a duplicate. |
| 3 | Confirms before it commits | Read the details back and waited for a yes before saving. |
| 4 | Sends the confirmation | You received a text with the right time and details. |
| 5 | Admits what it does not know | Took a message on call four rather than inventing an answer. |
| 6 | Hands over cleanly | Transferred or booked a callback, and passed your details on. |
| 7 | Leaves a record | Each call has a transcript, summary and outcome you can find. |
| 8 | You control the rules | You were shown where to change hours, handover and what it must never say. |
Checks 1, 2 and 4 are the operational core. A product that scores zero on those three is a chatbot regardless of its total. Checks 5 and 6 are the safety net, and they matter more than they look: an agent that invents answers or traps callers will cost you more goodwill than it saves in wages.
Why So Many Agent Projects Fail
Gartner's June 2025 prediction was blunt: over 40% of agentic AI projects will be cancelled by the end of 2027, because of escalating costs, unclear business value or inadequate risk controls. Most of those will be in large organisations, but the reasons apply to a ten person business just as well.
The projects that survive tend to share three habits. They start with one job that ends in an action, such as after hours bookings, rather than "customer service". They connect only the tools that job needs, so there is less to go wrong and less to secure. And they measure resolution, the share of calls where the request was completed with nothing left for staff, rather than deflection, which counts a caller who gave up as a success. Our 90 day plan for your first AI agent sets out that approach week by week.
Where People Still Belong
An operational agent is not an unsupervised one. Customers are clear about this. In a Gartner survey of nearly 5,800 people published in July 2024, 64% said they would prefer companies did not use AI in customer service, with the biggest fear being that it would make reaching a person harder. The answer is not to hide the AI but to make it useful and make the way to a person obvious.
| Let the agent do it | Keep a person on it |
|---|---|
| Bookings, reschedules and cancellations | Complaints and anything emotional |
| Answering questions from your written knowledge | Quotes for work nobody has seen |
| Taking and routing messages with a summary | Disputes about bills or refunds |
| After hours triage against your written rules | Genuine emergencies, which go straight to the on-call person |
| Confirmation and reminder texts | Anything the caller asks a human to handle |
Privacy law is moving the same way. From 10 December 2026, organisations covered by the Privacy Act must explain in their privacy policy when personal information is used in automated decisions that could significantly affect someone. Our guide to the automated decisions deadline covers what to check.
Reading the Price
How a product is priced often tells you what it really is. Chatbots are usually priced per conversation or per minute, because talking is what they do. Operational agents are more often priced per agent or as a flat plan, because the value is in the jobs completed. Neither is wrong, but it is worth doing the sum on your own call volumes.
- Per minute pricing rewards long calls. An agent that takes four minutes to book what a person does in two costs you twice as much.
- Per resolution or flat pricing lines the vendor's interest up with yours, as long as "resolved" is defined as the job done, not the caller going quiet.
- Setup and integration is where quotes hide cost. Ask whether connecting your calendar and writing the rules is included, or billed by the hour.
- A second number or call forwarding adds cost and a point of failure. An agent built into your phone system answers on your own numbers.
For a fuller cost comparison, including what an agent replaces, see our AI voice agent cost and ROI guide.
VOCPhone on Its Own Test
It would be a bit rich to hand you a test and not sit it ourselves. VOCPhone AI Agents are built into the VOCPhone phone system, so they answer on your own numbers with no second number or call forwarding, in a natural Australian accent. They answer questions from the knowledge you give them, book jobs into your calendar, text the customer a confirmation, take and route messages, and pass callers to your team when a person is needed. Every call is transcribed and summarised in the VOCPhone app, so the record in check 7 is there by default, and you set the rules for hours and handover.
The platform runs on our own network in Australia rather than a reseller's, so there is no chain of other companies between your caller and the agent. Our team in Australia looks after support during business hours, and our AI agents answer after hours, which is the same arrangement we recommend to customers. Run the five calls on us and score us honestly. That is what the test is for.
Frequently Asked Questions
How can I tell if an AI agent is really a chatbot?
Ask it to do something rather than tell you something. Set it up on a test number with your services and a test calendar, then ring it and ask to make a booking, ring back to change that booking, ask something it cannot know, and ask for a person. Afterwards, check the calendar, your text messages and the call log. A genuine operational agent will have created and moved the booking, sent a confirmation, admitted what it did not know, handed you over cleanly and left a summary of each call. If the calendar is untouched, it was a chatbot, however natural the conversation sounded.
What does agent washing mean?
Agent washing is Gartner's term for vendors rebranding existing chatbots, AI assistants, robotic process automation or phone menus as AI agents without giving them genuine ability to act on their own. In June 2025 Gartner estimated that only about 130 of the thousands of vendors claiming agentic AI actually offered it, and predicted that over 40 per cent of agentic AI projects will be cancelled by the end of 2027 because of cost, unclear value or weak risk controls. For buyers, the practical defence is to test products on real tasks and check the results in your own systems.
What is the difference between a chatbot, an AI assistant and an AI agent?
A chatbot answers questions from a script or knowledge base and hands anything else to a person or a link. An AI assistant helps a person do the work, for example by transcribing and summarising calls or drafting replies, but a person still takes the action. An operational AI agent completes the task itself using connected tools, such as checking a calendar, making a booking, sending a confirmation text and recording the outcome, and passes the call to a person when it should. All three are sold as AI agents in 2026, so it pays to know which one you are looking at.
What should an AI phone agent be able to do?
At a minimum it should answer questions accurately from your own information, book and change appointments in a real calendar, confirm details with the caller before saving them, send a confirmation text, take and route messages with a summary, admit when it does not know something, and transfer to a person or book a callback on request. It should also leave a transcript, summary and outcome for every call, and let you control its hours, handover rules and the things it must never say. Our 24 point scorecard in this article turns those abilities into a simple score.
Why do AI agent projects fail?
Gartner predicts over 40 per cent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. In practice, projects tend to fail when they start too broad, such as automating all of customer service at once, connect the agent to more systems than the job needs, or measure deflection instead of resolution. Projects that succeed usually begin with one job that ends in an action, like after hours bookings, connect only the tools that job needs, and track how many calls were completed with nothing left for staff to do.
Should customers be able to reach a person when an AI agent answers?
Yes. In a Gartner survey of nearly 5,800 customers published in July 2024, 64 per cent said they would prefer companies did not use AI in customer service, and the main worry was that it would make reaching a person harder. A good AI agent handles routine tasks quickly and makes the route to a person obvious, by transferring the call with the details passed along or booking a callback at a stated time. Complaints, disputes, emergencies and anything the caller asks a human to handle should go to a person.
Do VOCPhone AI Agents book appointments and send confirmations?
Yes. VOCPhone AI Agents are built into the VOCPhone phone system and answer on your own numbers in a natural Australian accent, with no second number or call forwarding. They answer questions from the knowledge you provide, book jobs into your calendar, text the customer a confirmation, take and route messages and pass callers to your team when a person is needed. Every call is transcribed and summarised in the VOCPhone app. VOCPhone runs on its own network in Australia, with support from its team in Australia during business hours and AI agents answering after hours.













