The Four-Way Tie Nobody Can Break
Here is the position, and it is almost universal in Australia this year.
Four quotes. All four include an AI receptionist that answers 24/7. All four include transcription, summaries and call scoring. All four integrate with the major CRMs. All four price within about fifteen per cent of each other per seat. All four demonstrated beautifully. The buyer's honest internal summary is "they all look the same", and the decision then gets made on the thing that is easy to compare, which is price, or on the thing that is easy to feel, which is who was nicest in the meeting.
They all look the same on paper because they were all written to look the same on paper. The differences are real, large, and located entirely in behaviour that a document cannot express.
Why the tie happens
The convergence is structural rather than dishonest. Speech recognition, speech synthesis and language models are commodity services now. Any platform can add "AI transcription" by sending audio to a speech service and displaying the result, and any platform can add an "AI receptionist" by placing a text-generation service between the caller and the call flow. Both can be shipped in weeks. So the presence of a feature no longer tells you about its quality, its latency, its data path, or who can fix it at four o'clock on a Thursday.
Why the Demo Cannot Help You
Six differences between a demo call and a real one. Every one of them is where products separate.
| In the demo | On Tuesday |
|---|---|
| A quiet room and a good connection. | A ute on the highway, a supermarket car park, two bars of signal. |
| A cooperative caller who waits for the prompt. | Someone who starts talking over the greeting because they are in a hurry. |
| Clear, deliberate speech. | A mumbled surname and a mobile number said as "double oh, four two, triple three". |
| Questions the system was built to answer. | A question about something you have never published anywhere. |
| Eleven on a Tuesday morning. | Ten past nine on a Saturday night, because something has gone wrong. |
| The conversation ends at the summary. | The summary has to become a record, in a system, that somebody acts on. |
The rule that saves the whole process
Never accept a demonstration in place of a trial number. Ask every shortlisted provider for a live number you can ring yourself, at times you choose, from your own mobile, for at least a week. A provider unwilling to give you one has told you something useful for free. A provider who gives you one and then asks what you found is showing you how they will behave as a supplier.
The Ten Calls, In Order
Make them in this sequence, because the early ones set up the later ones. Allow about ninety minutes per provider including note-taking. Use the same phone, the same mobile network and the same words each time.
- The interrupt. Wait until the AI is mid-sentence, then start talking. Real callers do this constantly. Look for: does it stop, listen, and follow your new point — or does it finish its script and ask you to repeat? This single call predicts caller satisfaction better than any other in the list.
- The local proper noun. Say a suburb you actually serve, a customer surname from your own database, and one of your product or service names. Then read the transcript. Australian place names and multicultural surnames are where general speech models fail hardest, and they are exactly what your bookings are made of.
- The mumbled mobile. Give a callback number quickly and naturally — "oh four one two, double three, five nine six" — the way people actually say it, ideally from a car with the window up. A wrong callback number is the most expensive small error the system can make, because nobody ever finds out.
- The wrong department. Ask for something the system should route elsewhere, then halfway through say you actually need something different. Look for: does it re-route cleanly, and does what you already said travel with you, or do you start again from nothing?
- The unanswerable question. Ask about a service you do not offer, or a price you have never published. The correct behaviour is to say it does not have that information and offer a person. A system that produces a confident, plausible, invented answer will one day invent a price, and you will be the one deciding whether to honour it.
- The emergency. Describe something urgent in ordinary words — water coming through a ceiling, someone unwell, a security issue — without using the word "emergency". Look for: does it recognise urgency from a plain description and get you to a human immediately, or does it continue collecting details?
- The 9pm call. Ring after hours, and again on a Saturday. Check that the greeting, the hours and the escalation path behave as configured rather than as assumed. Most businesses discover their after-hours volume is higher than they thought at exactly this point.
- The changed mind. Start booking an appointment, then change the day, then change it back, then ask what you have booked. Handling a reversal is much harder than handling a first request, and it is completely routine on real calls.
- The integration check. After calls 1 to 8, go and look at your CRM. Not the provider's dashboard — your CRM. Is there a record, against the right contact, with the right fields, created within a minute? A summary that never reaches the system where work happens has produced nothing.
- The repeat caller. Ring back an hour later as the same person, from the same number. Does it recognise you, and does the second conversation build on the first, or does it start from zero as though nothing happened?
Do this before the discount arrives
Once a price concession is on the table it becomes psychologically expensive to disqualify a provider on call 5 or call 6. Run the ten calls while everything is still hypothetical and nobody has invested anything in the outcome. The afternoon is the cheapest part of the whole procurement and the only part that observes the product rather than the pitch.
Scoring Them Without Arguing
Score every call 0, 1 or 2 immediately after making it, before discussing it with anyone. Ten calls, twenty points, four columns.
| Call | 0 — fails | 1 — acceptable | 2 — genuinely good |
|---|---|---|---|
| 1. Interrupt | Talks over you | Stops, asks you to repeat | Stops and follows the new point |
| 2. Proper nouns | Suburb and surname both wrong | Mostly right | Right, and a custom vocabulary can be loaded |
| 3. Mumbled number | Wrong digits, no flag | Right, or wrong but flagged as uncertain | Right, and read back for confirmation |
| 4. Re-route | Start again from nothing | Re-routes, context lost | Re-routes with everything you said |
| 5. Unanswerable | Invents an answer | Deflects vaguely | Says it does not know, offers a person |
| 6. Emergency | Keeps collecting details | Escalates when pushed | Recognises urgency and connects immediately |
| 7. After hours | Undefined behaviour | Hours configured | Hours, holidays and a genuine after-hours path |
| 8. Changed mind | Books the wrong thing | Handles it after a restart | Handles the reversal in conversation |
| 9. Integration | Nothing in your CRM | A note, wrong place | Right contact, right fields, within a minute |
| 10. Repeat caller | No recognition | Recognises the number | Recognises and continues the thread |
What the scores tend to look like
Two patterns show up almost every time. First, the spread is much wider than the price spread — providers within ten per cent of each other on cost routinely land eight or nine points apart out of twenty. Second, the calls that separate them are never the ones in the sales deck. Calls 1, 5 and 9 do most of the discriminating, and none of those three appear on any feature comparison table published anywhere.
What Each Failure Predicts
A zero is not just a lost point. Each one predicts a specific recurring problem after go-live, which is why the exercise is worth taking seriously rather than treating as a box to tick.
Fails the interrupt
Predicts daily low-level friction that never gets reported as a fault. Callers do not ring to complain that your receptionist talked over them; they simply find you harder to deal with, and the effect is invisible in every metric you collect.
Fails proper nouns
Predicts unusable notes. Once staff have read three summaries containing a suburb that does not exist, they stop trusting summaries entirely — and the entire value of transcription evaporates while you keep paying for it.
Fails the unanswerable question
Predicts an eventual dispute. A system confident enough to invent an answer will invent a price, a lead time or an eligibility, and the customer heard it from your phone number.
Fails the emergency
Predicts the incident you will be explaining afterwards. Recognising urgency from a plain description, without the caller knowing the magic word, is the difference between a triage tool and a liability.
Fails the integration check
Predicts a subscription that produces text nobody acts on. The value was never the summary; it was the record appearing where work happens without anybody retyping it.
Fails the repeat caller
Predicts the complaint you hear most: "I've already explained this." Continuity across calls is the single most-noticed quality improvement when it works, and the most-resented absence when it does not.
Three Questions the Demo Never Reaches
Behaviour is most of the decision, but three questions sit underneath it and should be answered in writing before anyone signs. An AI phone system creates three artefacts that did not exist before — an audio recording, a text transcript and a generated summary — and all three are records about identifiable people, created by default.
1. Where is the audio processed, and where does it rest?
Processing and storage are not the same thing and can happen in different countries. Audio is sometimes sent offshore for recognition even where the recording is stored locally. Ask for both, specifically, in writing, and ask whether your audio is used to train anybody's models — the answer should be no by default, or an opt-in you control and can see.
2. How long is it kept, and who can retrieve it?
Retention is a decision, not a default. Too short and you lose the evidence that settles a dispute; too long and you are holding personal information without a reason. Ask for per-record-type retention, deletion that genuinely deletes, role-based access to recordings, and an access log you can inspect. Then ask whether you can bulk export everything and leave — and test it before you need it.
3. Who can actually see the call when something is wrong?
When outbound calls start getting labelled, or the AI begins mishearing one particular caller, somebody has to look at the network, the routing and the model together. Ask how many companies sit between your fault and the person who can fix it. A provider who operates its own network and platform answers this differently from one who resells both.
Two Australian points to put on the same page
Recording notification obligations are state-based. Listening-device and surveillance-device legislation differs across the states and territories, so the workable standard is to notify at the start of every recorded call on every line — including the AI-answered path, which is regularly overlooked because it does not feel like a recorded call. One organisation-wide standard is simpler and errs in the right direction. From 10 December 2026, where personal information is used in automated decision-making capable of significantly affecting a person's rights or interests, privacy policy disclosure obligations apply. Most phone uses of AI are capture and triage rather than decision, but if yours prioritises, screens or declines, look at it before that date. General guidance, not legal advice — see our note on the December 2026 disclosure deadline.
Capture or Decide: The Setting That Matters Most
The most consequential configuration in an AI phone system is not which model it uses. It is the line between what the system captures and what the system decides — and providers differ enormously in how much of that line you control.
| Automate freely | Automate, review the exceptions | Never automate |
|---|---|---|
| Answering every call so nothing rings out | Prioritising a queue by stated reason | Whether a situation is an emergency |
| Capturing name, number, reason, callback time | Quoting a published price | Approving credit, refunds or expenditure |
| Your twenty most-asked published questions | Booking with capacity rules | Anything affecting rights, housing, care or employment |
| Transcribing, summarising and filing | Routing by detected topic | A caller in distress or a vulnerable caller |
| Confirmation messages | Flagging sentiment for review | Screening someone out |
The test for the middle column is reversibility. If a person can fix the mistake within the hour, automate it and review the exceptions. If the customer carries the consequence, a person decides and the AI prepares the decision.
Where to draw the line
The practical question for each provider is simply whether that boundary is visible and settable in their console. If it cannot be set, it has already been set for you — by somebody who has never met your customers. Our guide to which calls to automate works through the categories one by one.
Normalising Four Prices Onto One Page
Compare price last, and compare it properly. Four quotes are rarely quoting the same thing, and the per-seat headline is the least reliable number on any of them.
| Line | What to put in it | Where it hides |
|---|---|---|
| Seats × 60 months | Every seat type at its real rate, not the entry rate. | Discounts that expire at month 12 or 24. |
| Numbers | Main number, 1300 or 1800, any DIDs, plus the monthly service fee on inbound numbers. | Inbound call charges on 1300 numbers, which are billed to you, not the caller. |
| Call spend | Your actual mix — local, mobile, national, international — at quoted rates. | "Unlimited" definitions and fair-use clauses. |
| AI charges | Per minute, per call, per seat or bundled. Ask for a worked example at your real monthly volume. | Per-minute AI, which doubles when your inbound volume doubles. |
| Storage and retention | Recording and transcript storage at your chosen retention, over five years. | Included at 30 days, chargeable at 24 months. |
| Hardware | Handsets and headsets, purchased or amortised, plus replacement over five years. | Leases that outlive the contract. |
| Setup, porting, integration | One-off configuration, number porting, and connecting the output to your systems. | Integration quoted as professional services after signature. |
| Exit | Early termination, and what a bulk export costs. | Nothing hides better than an exit clause nobody reads. |
Divide the five-year total by sixty and by your seat count to get a real per-user-per-month number, and compare that. Our guide to comparing four quotes works through the arithmetic in detail, and every supplier worth dealing with will check your figures against their own quote. One who will not is telling you something.
The Two Questions to End Every Trial With
"What do the first ninety days look like, week by week?"
A specific answer names who does what and when: discovery, number porting dates, call flow design, a test day, a go-live day, and a review at thirty days with the numbers in front of you. A vague answer — "our onboarding team will be in touch" — predicts exactly the implementation you are imagining. More good products fail on implementation than on capability.
"Who do I ring at 4pm on a Thursday, and what can they see?"
The second half matters more than the first. Reaching a person is table stakes; reaching a person who can look at your actual call, your routing and the AI's behaviour in one place is the difference between a fix and a ticket. Ask what time zone they are in, and ask what happens at 4pm on the Thursday before a public holiday.
What to Measure Once It Is Live
Take a baseline for two weeks before go-live, or every later claim about improvement is an argument rather than a measurement.
Before
Answer rate, average speed to answer, abandoned calls, calls that rang out entirely, after-hours volume, and repeat callers within 48 hours.
After
The same six, plus containment judged by outcome rather than by whether a transfer happened, escalation speed, and the share of AI-handled calls that produced a correct CRM record.
Weekly
Ten transcripts checked by hand for the first month. Twenty minutes a week, and the single practice that most reliably separates a system people trust from one they quietly stop reading.
Consistent definitions matter more than sophisticated ones. Our note on the five numbers your phone system should report gives definitions you can hold a supplier to.
Run It On Us
We would rather be measured than described, so here is the invitation in plain terms: take the ten calls above, run them against three other providers and against us, in the same week, from the same phone.
We own the network and the platform
VOCPhone operates its own network rather than reselling somebody else's, and the AI sits inside the call flow rather than beside it as a forwarded call. That is why call 6 and call 9 behave the way they do, and it is why when something is wrong there is one company looking at it.
Australian voices, and your vocabulary
AI phone agents with natural Australian voices, plus a custom vocabulary for your suburbs, your product names and your staff names. Loading that list is the first thing we do, because it is the highest-return ten minutes in the entire deployment and almost nobody is told it exists.
Australian owned, Australian supported
Australian owned and Australian hosted, with Australian people answering the phone around the clock — which is the honest answer to "who do I ring at 4pm on a Thursday, and what can they see?"
You set what it may decide
What the AI answers, what it books, what it must hand straight to a person, and what it does when it does not know — all visible, all settable, and all worth revisiting at thirty days when you have real calls to look at instead of assumptions.
Frequently Asked Questions
How do I compare AI phone systems when every provider offers the same features?
Stop comparing documents and start comparing behaviour, because the feature lists converged during 2025 and no longer carry information. Speech recognition, speech synthesis and language models are commodity services, so any platform can add AI transcription by sending audio to a speech service and an AI receptionist by placing a text-generation service between the caller and the call flow, both in a matter of weeks — which means the presence of a feature says nothing about its quality, its latency, its data path or who can fix it. Ask each shortlisted provider for a live trial number rather than a demonstration, then ring it yourself with the same ten calls, in the same week, from the same phone and the same mobile network. Interrupt it mid-sentence. Say a suburb you serve, a customer surname and one of your product names, then read the transcript. Give a callback number quickly from a car. Ask for the wrong department and change halfway. Ask something it cannot know. Describe an emergency in ordinary words. Ring at 9pm and on a Saturday. Change your mind mid-booking. Check your own CRM, not their dashboard. Ring back an hour later as the same person. Score each 0, 1 or 2 immediately, before discussing it with anyone. In every comparison we have watched, the spread on those twenty points was far wider than the price spread, and the deciding calls were never the ones in the sales deck.
Why isn't a vendor demo enough to judge an AI receptionist?
Because a demo is a rehearsed conversation between two cooperating parties and a business call is nothing of the sort, so every provider's demo goes well — including the ones whose product will frustrate your customers within a fortnight. Six differences do the damage. The demo happens in a quiet room on a good connection; your callers are in a ute on the highway or a supermarket car park with two bars. The demo caller waits politely for the prompt; real callers start talking over the greeting because they are in a hurry. The demo caller speaks clearly; yours mumbles a surname and says a mobile number as double oh, four two, triple three. The demo asks questions the system was built to answer; your callers ask about things you have never published anywhere. The demo is at eleven on a Tuesday; your hardest call is at ten past nine on a Saturday night because something has gone wrong. And the demo ends at the summary, whereas in your business the summary has to become a record, in a system, that somebody acts on. None of those six appear in a demonstration and all six appear in the first fortnight. The rule that saves the process is simple: never accept a demonstration in place of a trial number you can ring yourself, at times you choose, for at least a week. A provider unwilling to give you one has told you something useful for free.
What does it mean if an AI phone system invents an answer?
It means you have found the failure mode with the longest tail, and it should usually be disqualifying. Test it deliberately by asking about a service you do not offer or a price you have never published. The correct behaviour is for the system to say it does not have that information and offer to put you through to a person. The dangerous behaviour is a confident, fluent, plausible answer that is simply untrue — dangerous precisely because it is fluent, since nobody listening has any signal that it was invented. A system willing to invent a service will eventually invent a price, a lead time, an eligibility or a delivery date, and the customer heard it from your business number, which makes it your problem to decide whether to honour. The related failure is quieter and more common: a summary built on a flawed transcript. If the transcription mangled a suburb, the summary will confidently restate the mangled version as fact, and once staff have read three summaries containing a suburb that does not exist they stop trusting summaries altogether — at which point you are paying for a feature nobody reads. So check the transcript beside the summary, ask whether the system marks what it could not hear clearly, and treat willingness to say I don't know as a feature rather than a limitation. It is the single clearest signal of a well-built system.
What should I ask about where my call recordings and transcripts are stored?
Ask six questions and get the answers in writing before signing, because an AI phone system creates three artefacts that did not previously exist — an audio recording, a text transcript and a generated summary — all of them records about identifiable people and all of them created by default. First, in which country is the audio processed, remembering that processing and storage are different things and audio is sometimes sent offshore for recognition even where the recording rests locally. Second, where do the recordings and transcripts actually live, which determines whose law reaches them and which subprocessors are involved. Third, is your audio used to train anybody's models, where the right answer is no by default or an explicit opt-in you control and can see in the console. Fourth, how long is each record type kept and can you change it, since retention is a decision rather than a default — too short and you lose the evidence that settles a dispute, too long and you hold personal information without a reason. Fifth, who inside your own organisation can retrieve a call, which should be role-based with an access log you can inspect, because recordings of customer conversations are not general staff reading material. Sixth, can you bulk export everything and leave, tested before you need it, since two years of transcripts is simultaneously a real asset and a real lock-in. Alongside those, note that recording notification obligations are state-based, so notify on every recorded call including the AI-answered path.
How much does AI add to the cost of a business phone system?
Ask how it is charged before you ask how much, because four billing models produce very different bills at identical volume. Per-minute charging scales directly with talk time and is the one that surprises people when a campaign, a product recall or an outage doubles inbound calls in a week. Per-call charging is more predictable but penalises short calls. Per-seat charging is easiest to budget and least sensitive to volume. Bundled pricing hides the mechanism entirely, which is comfortable until the bundle changes at renewal. Whichever applies, ask for a worked example at your actual monthly call count in writing, then ask for the same table with volume doubled, because that is the scenario that generates the disputed invoice. Two further costs belong in the comparison and are usually left out of it. Recording and transcript storage across a multi-year retention period is a recurring cost that grows every month, and it is often included at thirty days and chargeable at twenty-four. And integration work — connecting the output to your CRM, calendar or ticketing system — is where the value is actually realised, so if it is quoted as professional services after signature rather than included, it belongs in the five-year total. Then normalise everything: seats times sixty months, numbers, call spend, AI charges, storage, hardware, setup and porting, and exit costs, divided by sixty and by seat count. Compare that number, and compare it last.
Should AI answer all of my business calls or only some?
Only some, and the boundary should be set by you rather than accepted as a default. Draw it at reversibility. If getting something wrong produces an inconvenience a person can fix within the hour, automate it and review the exceptions; if getting it wrong produces a consequence the customer carries, a person decides and the AI prepares the decision. That gives three practical groups. Automate freely: answering every call so nothing rings out, capturing name, number, reason and callback time, answering your twenty most-asked published questions, transcribing and filing, and sending confirmation messages. Automate but review: prioritising a queue by stated reason, quoting a published price, booking into a calendar with capacity rules, routing by detected topic, and flagging sentiment for later review rather than for judging individuals. Never automate: deciding whether a situation is an emergency, approving credit, refunds or expenditure, anything affecting a person's rights, housing, care or employment, a caller in distress or a vulnerable caller, and screening someone out. The instinct is to point AI at the hardest calls because they hurt the most, but the return sits in the highest-volume, lowest-judgement calls — hours, address, status, booking — and clearing those means the hard calls reach a better-rested human sooner. Ask each provider to show you where that boundary is set in their console; if it cannot be set, it has been set for you.
What should I measure before and after switching on an AI phone system?
Take a two-week baseline before go-live on six numbers, because without it every later claim about improvement is an argument rather than a measurement. Answer rate, counted as calls answered divided by calls offered, using the same definition on both sides of the change. Average speed to answer. Abandoned calls, and separately the calls that rang out entirely, since the two have different causes and different fixes. After-hours call volume, which is almost always higher than people expect and is where AI answering delivers most of its value. And repeat callers within forty-eight hours, because a second call from the same person usually means the first one produced nothing. After go-live keep all six and add three that only exist once AI is in the flow: containment, meaning the share of calls fully resolved without a person, judged by whether the caller's problem was actually solved rather than by whether a transfer occurred; escalation rate and how quickly escalations connect to a human; and the proportion of AI-handled calls that produced a correct record in your CRM. Then add the unglamorous practice that matters most — check ten transcripts by hand every week for the first month. It takes about twenty minutes, it is tedious, and it is the single habit that most reliably separates a deployment staff trust from one they quietly stop reading while the invoice keeps arriving.