What Is an AI Agent? Four Questions for Your Vendor

The demonstration will be good. That is the first thing to understand about buying AI phone answering in 2026: the part everybody watches has been solved for about two years, so every vendor can put a natural sounding voice in front of you that handles an interruption, copes with a broad Australian accent, and answers a question about your opening hours without sounding like a recording from 1998. If you judge on the demonstration you will judge them all identically, because on that measure they very nearly are identical. The differences sit entirely in the part nobody demonstrates: what the system is allowed to do, what it does when it does not know, where the call goes when it cannot help, and which country your customers' conversations end up sitting in. Those four things decide whether the product removes work from your business or simply moves the work to a quieter part of the day while adding a monthly bill. They are not technical questions and you do not need to understand how any of this is built to ask them. You need four sentences, asked in a particular order, and the discipline to watch the screen rather than listen to the voice. Here they are, with the answers that mean yes, the answers that mean no, and the answers that mean the person in front of you has not thought about it.

Buying AI · Vendor Questions · 2026

Four Questions, And You Will Know in Ten Minutes

In 2026 every phone provider in Australia sells an AI agent, and the demonstrations are all excellent, because the conversational part of this has been solved for two years. What separates the products is everything that happens after the caller stops talking. Four questions, asked in this order, will tell you whether you are buying something that finishes work or something that has a very pleasant conversation and then puts the job back on your desk.

📅 ⏱ 16 min read 🇦🇺 Australian owned · Australian network · Australian support
TL;DR

Question one: can it write, or only read? Ask to see a call where the outcome lands in the calendar or the CRM with nobody typing it in afterwards. Watch the screen, not the transcript. "It emails your team who then enter it" means the work is still yours. Question two: what does it do when it does not know? The right answer is to say so and offer a callback. A system that answers your pricing question from general knowledge instead of your records will eventually invent a price. Question three: what happens when it cannot help? A person must always be reachable, the context must travel with the call, and after hours escalation has to mean a booked callback with a time you keep. Question four: where does the conversation live? Which country, for how long, and who can see it. Then choose two call types, not eight, start after hours where the alternative is a voicemail nobody returns, and measure resolution rather than containment, because a caller who gave up counts as contained.

Automate hours, bookings, status and overflow. Never automate complaints, distressed callers, or the first call from a large new customer.

Why the Demonstration Tells You Nothing

Sit through four vendor demonstrations in a week and something odd happens: they blur. The voices are equally natural. Each one handles being interrupted. Each one gets the trick question about the public holiday right. You come away with an impression rather than a comparison, and impressions are decided by whoever presented last.

That is not the vendors being slippery. It is that the conversational layer is now a commodity. The models that produce natural speech and understand ordinary sentences are broadly available, and no provider in Australia has a meaningful advantage there. What is not a commodity is the connection to your business: what the system can see, what it is permitted to change, how it behaves at its own limits, and where the resulting data sits. None of that is visible in a scripted demonstration, and all of it decides whether the thing is worth the monthly fee.

The reframe that makes this easy

Stop asking how good the AI is. Ask what it is allowed to do. The first question has no answer you can verify in a meeting. The second has an answer you can watch happen on a screen in ninety seconds.

Question One: Can It Write, or Only Read?

Say this: show me a call where the outcome appears in the calendar or the CRM, and nobody types it in afterwards.

This is the whole ballgame. Systems that read from your business can answer questions accurately and pleasantly. Systems that write to your business can finish the job. The gap between those two is not a matter of degree, it is a different product with different requirements, and it is the difference between a call that ends and a call that has been deferred.

What you will hearWhat it meansVerdict
"It sends your team an email with the details."Somebody still has to read it and enter it. The work has moved to 8am tomorrow.Reads only
"It creates a task in our portal for someone to action."Same, with a queue. Also a second place your staff now have to look.Reads only
"It logs the enquiry and notifies you."Message taking with better handwriting. Genuinely useful, not what you are being quoted.Reads only
"The booking is in your calendar, the customer has the SMS, and the note is on the contact."The job is finished. Nobody touched it.Writes

Reading is not worthless. For a lot of businesses a system that answers the twelve questions your phone gets all day, accurately, at 9pm, is a fine purchase and considerably cheaper than one that acts. The problem is only that the two are sold under the same word at very different prices. Know which you are buying.

Then ask the follow up, because it is the one that reveals the engineering

When it writes something wrong, who finds out, and how fast can it be undone? A serious product has an answer involving a confirmation step before commit, a visible log of what the agent changed, and a way to reverse it. A product that has not been built for real use will treat this as a hypothetical. It is not hypothetical. It happens in the first fortnight, usually with a customer who gave two dates in one sentence.

Question Two: What Does It Do When It Does Not Know?

Say this: ask it something about my business that is not in anything you have loaded, and let me hear what it says.

There are three possible behaviours and only one of them is safe.

It says it does not know

Plainly, without apologising for four sentences, and offers to have somebody call back or to take the question down. Every reasonable customer accepts this. It is what a good new employee does in week one.

⚠️

It deflects into a menu

"I can help with bookings, hours or accounts." Acceptable if it happens once. If it is the answer to everything unfamiliar, you have bought a menu with a nicer voice and it will frustrate people faster than the old one did.

🚫

It answers anyway

Confidently, plausibly, and from general knowledge rather than your records. This is the failure that ends up on social media. A system that will guess your opening hours will eventually guess your price.

The fix is architectural rather than a matter of the model being better. The agent should be restricted to your published material and your own records for anything factual about your business, with not knowing as an explicitly designed outcome rather than an accident. Ask the vendor how they enforce that. If the answer is that the model is very accurate, that is not an answer, it is a hope.

Worth testing with a nasty one

Ask it something adjacent and slightly wrong, of the kind customers actually ask. "Do you do the same job for units as for houses, and is it the same price?" A grounded system will answer the part it knows and separate the part it does not. An ungrounded one will produce a confident, tidy, invented answer, and it will sound better than the correct one. That is exactly why it is dangerous.

Question Three: What Happens When It Cannot Help?

Say this: show me the handover, including at seven in the evening when nobody is there.

Customers do not form their opinion of automated answering on the calls that go well. They form it at the moment the machine reaches its limit, because that is when they discover whether your business has thought about them or has simply installed something to keep them away. Five things have to be true.

RequirementWhat good looks like
A person is always reachable"Can I talk to someone" works as the first sentence of the call, with no negotiation and no attempt to talk the caller out of it. Hiding the exit is the most resented pattern in phone automation and it long predates AI.
Context travelsWhoever picks up already sees the number, the request, what the agent did, and what it could not do. If your customer has to explain it twice, the automation has made their experience worse, not better.
Frustration triggers escalationRepetition, interruption, a raised voice, a second attempt at the same thing. The system should give up before the customer does, even when it believes it can help.
Some subjects never enter the agentA fixed list, not a judgement: complaints, anything legal, anything with a safety dimension, anything involving a vulnerable person. This is set by you and enforced by the system.
After hours means something realWhen there is genuinely no one to transfer to, escalation is a booked callback with a stated time, and the discipline to keep it. An unanswered transfer at 7pm is worse than the voicemail it replaced.

Ask to see the escalation before you look at anything else, and be suspicious of a vendor who wants to show you the clever part first. Our note on replacing phone menus covers why the exit route has always mattered more than the greeting.

Question Four: Where Does the Conversation Live?

Say this: which country are the recordings and transcripts stored in, for how long, and who can see them?

Every one of these calls produces a transcript, and the transcripts contain things people say on the phone: names, addresses, medical details, financial circumstances, the reason they need a locksmith at 11pm. That material now sits somewhere, under somebody's law, accessible to some list of people. Most businesses discover where only when a customer asks them.

AskWhy it matters
Which country holds the audio and the transcripts?For health, legal, government, aged care and NDIS work this is frequently a hard requirement rather than a preference, and it is much cheaper to get right at the start.
Is my call content used to train anyone's model?You want a clear no, in the contract, not a paragraph on a website that can be edited.
How long is it kept, and can I set that?Retention you cannot control is retention somebody else decided. Some industries need years, most need months.
Who inside the provider can listen?There is usually a legitimate support answer. There should also be a log.
How many companies are in the chain?Plenty of AI phone products are a thin layer over two or three other providers. Each hop adds a place your data rests and a party who can change their terms.

That last one also explains something you can hear. When the audio has to travel offshore to be understood and then come back, the pause before each reply stretches, and a pause of half a second is the difference between a conversation and an interrogation. Latency is not only an engineering concern, it is the main reason some of these systems feel uncomfortable to talk to. Our note on who actually runs the infrastructure goes through the layers.

The Fifth Question, If You Have Time

What stops someone talking your agent into doing something it should not?

If the system reads incoming messages, emails, web forms or attached documents as part of its work, then whoever wrote that content can attempt to give it instructions. It is the same shape of problem as a malicious email link and it has a boring, effective answer: content that arrives from outside is treated as data and never as instructions, the actions available to the agent are kept narrow, and anything that moves money or changes an account requires confirmation. You are not asking this because you expect an attack next Tuesday. You are asking because a vendor who has never considered the question has told you a great deal about how the product was built.

Picking the Calls to Start With

Once you have a vendor, the next mistake is scope. Businesses that automate eight things at once end up with eight mediocre experiences, and mediocre is what customers remember and repeat. Two, done properly, changes the week.

Start hereBe careful hereLeave alone
Opening hours, address, parking, what you do and do not coverQuoting anything variableComplaints of any kind
Booking, rescheduling and cancellingTaking paymentDistressed or vulnerable callers
Order, job or delivery statusChanging account detailsAnything with a safety or medical dimension
After hours capture with a real outcomeAnything that creates an obligationA cancellation you would fight to keep
Overflow when everyone is already on a callAnything with a deposit ruleThe first call from a large new customer
The honest reason to do this

It is rarely that a machine answers better than your best person. It is that at 8:40 on a Monday your best person is on the first call and cannot take the fourth, and the people who ring during that window mostly do not leave a message and do not ring back. They ring the next business on the list, and you never find out they existed. That invisible loss is what automation actually addresses, which is also why the number to watch is total answered enquiries rather than anything about call quality. Our piece on the five numbers your phone system should tell you covers how to see it.

The Part of the Project Nobody Quotes For

An agent knows what it can see. Getting it to see the right things is most of the work, and it is almost never in the quote.

What it needsWhere it comes fromThe actual effort
Services, inclusions, exclusions, pricesYour website, plus the rules that live in people's headsThe hard part. Writing down what everybody knows and nobody wrote.
Hours, holidays, coverage areaThe phone system and your calendarAn hour, then the habit of keeping it current.
Availability and booking rulesCalendar or job managementConnecting it, and deciding what it may book without asking.
Customer historyCRMMatching on the calling number, and deciding what may be read out before identity is confirmed.
Who handles what, and whenCurrently nowhereAn afternoon. Worth doing even if you never buy the agent.

A business whose systems already talk to each other can be live in days. A business with a CRM nobody updates, a calendar in one person's phone, and pricing that depends on who answers will spend the project fixing those, and will get most of the benefit from having fixed them. That is not a wasted project, but it should be planned as what it is. Our guide to open APIs and keeping your software in control covers the connection work.

The Numbers, and the One That Lies

Containment rate is the figure every provider puts on the dashboard, and it is the figure most capable of flattering a failure. A caller who hung up in frustration is contained. A caller told to ring back tomorrow is contained. Watch it next to five honest ones.

Resolved
Calls where the need was actually met, checked against the record that got written. This is the real number.
48h
Repeat contact within two days. Rising repeats alongside rising containment is the classic false positive.
Total
Answered enquiries overall, against the same month last year. The number that pays for the system.

Add escalation quality, meaning whether handovers arrive with context and how long the human then spends, and abandonment inside the agent, which is almost always one broken conversation flow rather than a fault of the model. Then read twenty randomly chosen transcripts a week for the first two months. Not the ones the system flagged. Random ones. It takes half an hour and it will tell you more about your customers than any dashboard will.

What Australian Law Expects of You

ObligationIn practice
Be honest that it is automatedA short plain statement in the greeting. Not currently a specific telco rule, increasingly expected, and pretending to be a person is a trust problem that outlasts the call.
Notify recordingRecording law is state based and inconsistent, so one organisation wide standard applied to every path including the automated one is simpler and errs the right way.
Disclose automated decisionsPrivacy Act reforms require disclosure in your privacy policy where automated systems make or substantially influence decisions about individuals. If the agent decides priority or eligibility, that belongs in the policy. See our note on the privacy policy deadline.
Keep 000 out of itAn emergency call must never enter an automated flow. Verify it rather than assume it. Our guide to Triple Zero on a cloud phone system covers the surrounding obligations.
Outbound rules apply unchangedAn agent making calls is bound by the Do Not Call Register, consent, calling hours and identification exactly as a person is. See the outbound compliance mistakes.
Know that voice is no longer identificationSynthetic voice has made a familiar voice on the phone worthless as proof of who is calling. Approvals should not depend on recognising one. Our note on voice cloning and vishing covers the defence.

Reading the Price

Three pricing shapes are common. Per minute is predictable and quietly rewards a system that rushes people. Per resolved interaction lines the incentives up nicely and requires you to agree, in writing and before signing, what counts as resolved. Included in the platform with a fair use limit is simplest, and worth testing against your busiest week rather than your average one.

The comparison to avoid is a receptionist's salary, because most businesses asking this question were never going to hire one, so the saving is imaginary. Compare instead against the enquiries currently being lost. Take your weekly enquiry count, your honest estimate of how many go unanswered, your conversion rate and your average job value, and the sum takes two minutes. Our AI voice agent cost and ROI guide sets it out with the numbers filled in.

Two costs sit outside every quote. The setup described above, measured in days. And curation: an hour a week for two months, then an hour a month, spent reading transcripts and correcting the three things it keeps getting wrong. Deployments do not usually fail loudly. They decay, at about six weeks, in exactly this spot.

Starting Small on Purpose

Begin after hours. It is the safest trial available anywhere in the business, because the thing it is replacing is a voicemail most people never leave and nobody returns until Tuesday. The floor is low, the learning is real, and no customer is worse off than they were.

Then add overflow during the day: calls that arrive when every person is already on a call. Again you are competing with an unanswered ring rather than with a human being. Only after both of those are behaving well should the agent take a call that a person could have answered, and by then you will have read enough transcripts to know exactly which calls those should be.

Tell your team before it goes live rather than after. Staff who understand that this is taking the fourth simultaneous call will help you improve it, and they are the ones who know the unwritten rules you need in week two. Staff who find out by hearing a strange voice answer their phone will assume the worst, and they will be the reason it gets switched off.

How We Answer These Four

It would be poor form to hand you four questions and dodge them. VOCPhone writes as well as reads, into the calendar, the CRM and the ticket, with a confirmation step before anything commits and a log of every change the agent made. It answers from your material and your records, and not knowing is a designed outcome rather than an accident. Escalation is built before the conversation is, a person is always reachable, context travels with the transfer, and after hours the fallback is a booked callback rather than a ring into an empty office.

On the fourth question: we own and operate our own network and platform in Australia, the audio does not leave the country to be understood, your call content is not used to train anybody's model, and there is no chain of resellers between you and the thing answering your phone. That is also why it does not have the pause. Ask the other four vendors the same four questions, in that order, and compare the answers rather than the voices.

Ask us the four questions

Bring your calendar, your CRM and the two questions your phone gets most. We will show you a call that finishes the job, on your data, and you can watch the record appear.

Talk to us Or call 1300 663 222

Frequently Asked Questions

What is an AI agent on a phone system?
An AI agent is software that takes a goal, works out the steps itself, carries them out using tools it has been given, and reports what it did. On a business call the goal arrives as an ordinary sentence from a customer, the tools are your calendar, CRM, job management system and messaging, and the report is a spoken confirmation to the caller plus a record left behind for your team. The word doing the work in that definition is tools. Software with no tools can only talk, and it can talk very well. The commercial distinction is short: an assistant answers questions, an agent finishes jobs. If the caller still has to ring back, fill in a form, or wait for somebody to read a message before anything happens, you have an assistant. That is a legitimate product and it costs much less. The confusion in 2026 is that both are sold under the word agent at very different prices. Agentic, used properly rather than as marketing, describes the second kind: it decides the sequence of steps itself rather than following a script written in advance, it can use more than one tool inside a single call, and it recognises when it has reached the edge of what it should be doing and hands over.
What questions should I ask an AI phone provider?
Four, in this order. One: can it write, or only read? Ask to see a call where the outcome appears in the calendar or CRM without anybody typing it in afterwards, and watch the screen rather than the transcript. Two: what does it do when it does not know? The only safe behaviour is to say so plainly and offer a callback. A system that will guess your opening hours will eventually guess your price. Three: what happens when it cannot help? Show me the handover, including at seven in the evening when nobody is there. A person must always be reachable, context has to travel with the call, frustration should trigger escalation before the customer gives up, some subjects must never enter the agent at all, and after hours escalation has to mean a booked callback with a time you keep. Four: where does the conversation live? Which country holds the audio and transcripts, for how long, who can listen, whether your call content trains anybody's model, and how many companies sit in the chain. If you have time, add a fifth: what stops somebody talking the agent into doing something it should not, which matters as soon as it reads incoming messages, emails or documents.
Why do all the AI phone demonstrations look the same?
Because the part being demonstrated is now a commodity. The models that produce natural speech, understand ordinary sentences, handle interruptions and cope with an Australian accent are broadly available, and no provider has a meaningful advantage there. Sit through four demonstrations in a week and they blur, which is why the decision usually goes to whoever presented last. What is not a commodity is everything the demonstration does not show: what the system can see of your business, what it is permitted to change, how it behaves when it reaches its own limits, and where the resulting recordings and transcripts sit. Those four things decide whether the product removes work or simply moves it to a quieter part of the day while adding a monthly bill. The useful reframe is to stop asking how good the AI is, which has no answer you can verify in a meeting, and start asking what it is allowed to do, which has an answer you can watch happen on a screen in about ninety seconds. Insist that any trial runs on your real data with your real exceptions, because every one of these systems performs beautifully on a fictional business with no awkward cases.
Will an AI agent make up answers about my business?
It can, and preventing it is a design decision rather than a matter of waiting for better models. There are three possible behaviours when a system meets a question it cannot answer from what it holds. It can say so plainly and offer a callback, which is what a good new employee does in their first week and what every reasonable customer accepts. It can deflect into a list of things it does handle, which is tolerable once and infuriating as a standard response, because that is a menu with a nicer voice. Or it can answer anyway, confidently and plausibly, from general knowledge rather than your records, which is the failure that ends up screenshotted. The protection is architectural: restrict the agent to your published material and your own records for anything factual about your business, and make not knowing an explicitly designed outcome. Ask the vendor how they enforce it, and treat the model is very accurate as a hope rather than an answer. Test it with a question that is adjacent and slightly wrong, of the sort customers actually ask, because an ungrounded system will produce a tidy invented answer that sounds better than the correct one.
What should happen when the AI cannot help the caller?
Five things have to be true. A person is always reachable, so can I talk to someone works as the first sentence of the call with no negotiation and no attempt to talk the caller out of it; hiding the exit is the most resented pattern in phone automation and it long predates AI. Context travels, so whoever picks up already sees the number, the request, what the agent did and what it could not do, because a customer who has to explain it twice has had a worse experience than if you had never automated anything. Frustration triggers escalation, meaning repetition, interruption, a raised voice or a second attempt at the same request cause the system to give up before the customer does, even when it believes it can help. Some subjects never enter the agent at all, as a fixed list rather than a judgement call: complaints, anything legal, anything with a safety dimension, anything involving a vulnerable person. And after hours escalation means something real, because when there is genuinely nobody to transfer to, the answer is a booked callback with a stated time and the discipline to keep it. An unanswered transfer at 7pm is worse than the voicemail it replaced.
Where is my call data stored with an AI phone system?
That depends entirely on the provider, and most businesses find out only when a customer asks them. Every call produces a transcript, and transcripts contain what people actually say on the phone: names, addresses, medical details, financial circumstances, the reason somebody needs a locksmith at 11pm. Ask five things. Which country holds the audio and the transcripts, which is frequently a hard requirement rather than a preference for health, legal, government, aged care and NDIS work. Whether your call content is used to train anybody's model, where you want a clear no in the contract rather than a paragraph on a website that can be edited. How long it is kept and whether you can set that, because retention you cannot control is retention somebody else decided. Who inside the provider can listen, where there is usually a legitimate support answer and there should also be a log. And how many companies are in the chain, since many AI phone products are a thin layer over two or three other providers, each hop adding a place your data rests and a party who can change their terms. That last question also explains the pause some systems have, because audio travelling offshore to be understood and back again is what stretches the gap before each reply.
Which calls should I automate first?
Start with two, not eight. Businesses that automate everything at once end up with eight mediocre experiences, and mediocre is what customers remember and describe to other people. Good first candidates are repetitive, structured, high in volume and low in emotion: opening hours, address, parking and what you do or do not cover; booking, rescheduling and cancelling; order, job or delivery status; after hours capture that produces a real outcome instead of a voicemail; and overflow when everybody is already on a call. Be careful with anything variable that has to be quoted, taking payment, changing account details, and anything that creates an obligation, all of which work but need tighter confirmation and narrower permissions. Leave alone complaints, distressed or vulnerable callers, anything with a safety or medical dimension, a cancellation you would fight to keep, and the first call from a large prospective customer, because in each of those the caller is deciding whether you take them seriously. Begin after hours, where the thing you are replacing is a voicemail nobody returns until Tuesday, then add daytime overflow, where you are competing with an unanswered ring rather than with a person.
How do I know if the AI agent is actually working?
Do not judge it on containment rate, the share of calls that never reached a human. Every provider puts it on the dashboard and it is the figure most capable of flattering a failure, because a caller who hung up in frustration is contained and a caller told to ring back tomorrow is contained. Watch it beside five honest measures. Resolution rate, meaning calls where the need was actually met, verified against the record that got written; that is the real number. Repeat contact within 48 hours, which shows whether the call finished from the customer's point of view, and where rising repeats alongside rising containment is the classic false positive. Total answered enquiries against the same month last year, which is the number that pays for the system. Escalation quality, meaning whether handovers arrive with context and how long the human then spends, since a handover slower than no automation is a net loss. And abandonment inside the agent, which is almost always one broken conversation flow rather than a fault of the model. Then read twenty randomly chosen transcripts a week for the first two months, not the ones the system flagged, which takes half an hour and teaches you more than the dashboard.

What to Read Next

Your next reads

VOCPhone — the Australian-owned cloud phone platform that owns and operates its own network. vocphone.com | 1300 663 222

Related Articles