The Anatomy: Four Calls, One Loss
This is a composite of a pattern rather than one business's story, and the shape is what matters. Read it as a sequence, because that is how it was built.
- Tuesday, 10:14am — the spelling check. A pleasant caller confirms the correct email address for accounts. Thirty seconds. No request, nothing asked for, nothing given away that is not on the website. The purpose of this call is not information. It is to establish that a call from this number is normal.
- Thursday, 2:40pm — the process question. Someone asks who approves changes to supplier details. They mention having dealt with the business before. They are given a first name, because it is a reasonable question and the answer is not secret.
- Monday, 11:05am — the credibility call. A caller references the accounts email address and the approver's first name. Both facts are correct, so the caller is treated as legitimate — which is the entire mechanism. Knowledge is being mistaken for authorisation. A small, plausible query is resolved. Trust is now established.
- Wednesday, 4:50pm — the ask. A voice indistinguishable from the supplier's accounts manager explains their bank has changed, apologises for the timing, notes the details are in the email already sent, and asks that the payment due Friday go to the new account. It is ten minutes before close on the day before the payment run.
Every element of that sequence was designed to be individually forgettable. The forgetting is the attack.
— why sampling does not work
Notice the timing of the final call. Ten to five, the day before the run, with a reference to an email that supposedly already exists. Three separate pressures — time, precedent, and a document you have not read yet — deployed on a person who has been softened by three prior contacts they will not consciously recall.
Why Each Call Passed Review
Suppose this business recorded every call and had somebody review two per cent of them, chosen at random. What are the odds that reviewer sees the pattern?
| Call | If reviewed in isolation, what would a reviewer note? |
|---|---|
| 1. Spelling check | Nothing. A courteous thirty-second call, handled well. Possibly a positive example. |
| 2. Process question | Nothing. A staff member answered a reasonable question helpfully. Arguably good service. |
| 3. Credibility call | Nothing. The caller knew internal details, so of course they were treated as known. |
| 4. The ask | This one is catchable — but only after the money is gone, and only if this specific call happens to fall in the two per cent sample. |
The structural problem is that human review looks for a bad call, and the threat is a series of good ones. No amount of diligence fixes that, because the information required to spot it does not exist inside any single call. It exists in the relationship between four calls spread across nine days — different staff, different departments, no shared memory. That is not a human-attention problem. It is a data problem, and it has a data answer.
What It Costs When It Works
The published national figures, once, for context rather than alarm.
$97,000
Average incident cost, medium business — up 55%
$56,600
Average incident cost, small business — up 14%
$202,700
Average, organisations over 200 staff — up 219%
6 min
Roughly how often a cybercrime is reported in Australia
More than 84,700 cybercrime reports in a year. Over 1,200 incidents responded to, up eleven per cent. Average cost per incident up fifty per cent to $80,850. And an explicit assessment that AI almost certainly enables malicious actors to execute attacks on a larger scale and at a faster rate, through AI-generated phishing, cloned voices and cloned websites.
The line worth reading twice
The steepest increase — 219% — belongs to the largest organisations, the ones with security teams, budgets and mature controls. That tells you the attack side improved faster than the defence side at the top of the market. The inference for a smaller business is not comfort. It is that the attempts arriving at a fifteen-person company are now built with the same tooling as those aimed at organisations employing security professionals, and a proportionate response has to account for that.
The Honest Matrix: Helps, Partly Helps, Does Nothing
Every AI security article should contain this table and almost none do. Three columns, no hedging.
| Threat | Does AI on the voice channel help? | What actually addresses it |
|---|---|---|
| Multi-call social engineering sequence | Materially | Cross-call pattern analysis. This is the case AI is genuinely built for. |
| Toll fraud / system compromise for call traffic | Materially | Real-time traffic anomaly detection with somebody empowered to block. |
| Inconsistent identity verification | Materially | Automated challenge sequences that do not respond to urgency or flattery. |
| Cannot prove what was said or authorised | Materially | Searchable, timestamped transcripts with a defined retention rule. |
| Policy drift under pressure | Materially | Reviewing every call rather than a sample, so breaches surface as rates. |
| Business comms on personal channels | Materially, indirectly | Answering calls reliably so staff stop inventing workarounds. |
| Synthetic / cloned voice | Partly | Not detection — a callback policy on the request type, applied without exception. |
| Invoice fraud by email with a voice follow-up | Partly | Voice-side flagging helps; the control is dual authorisation on payment changes. |
| Compromised credentials | No | Multi-factor authentication, password manager, conditional access. |
| Unpatched systems | No | Patching. Still the highest-return activity in security. |
| Malicious insider with legitimate access | Barely | Least privilege, separation of duties, dual authorisation. |
| No process to escalate to | No — makes it worse | Naming, in writing, who is called and what they may do. |
Where AI Materially Helps
Six items from that first block, with the mechanism rather than the marketing.
Patterns across the history
Repeated calls from a number that never transacts. Callers asking about internal process rather than products. The same voice across departments. Requests clustering before close of business. Individually nothing; together a profile worth handing to a person.
Traffic anomalies
Destinations never called before, volume outside your own established pattern, concurrent call counts that are arithmetically impossible for your headcount, extensions registering from unexpected networks. Overnight, when nobody is watching.
Verification that does not get tired
The same challenge sequence at 4:55pm on a Friday as at 10am on a Tuesday. An automated first stage is not embarrassed to ask a third question, does not respond to urgency and cannot be flattered — which is exactly what the attacker was counting on.
Transcripts you can search
When a fraud pattern surfaces on Thursday, you can ask whether anyone took a matching call this month or last quarter. Without transcripts that question has no answer at all.
All the calls, not two per cent
Policy breaches become a measured rate instead of an anecdote, coaching targets what people actually do under pressure, and slow drift becomes visible. A two per cent sample is close to useless as a security control.
Coverage that removes the excuse
Calls answered after hours and at peak means nobody needs a personal mobile workaround. This is the largest and least-credited security benefit on the list.
The Synthetic Voice Problem, and the Only Durable Answer
The fourth call in the anatomy used a voice indistinguishable from a known contact. That is now cheap and it is the part businesses find hardest to accept.
Do not buy a detector
Automated detection of synthetic audio is improving and it is not reliable enough to be a control. Generation is currently outrunning detection. A confident vendor claim to detect deepfake voices is a reason for scepticism, not comfort — and a control you cannot trust is worse than no control, because it produces the confidence without the protection.
Stop authenticating the voice. Start authenticating the request. Voice as an identity proof is finished, permanently, and no technology restores it. What survives is process: a defined list of high-risk actions, each requiring out-of-band verification, applied without exception and regardless of who appears to be asking. AI's genuine contribution is that it applies that rule identically at ten to five on a Wednesday — which is precisely when the request will arrive.
The Callback Policy, Written Out
This is the single highest-value control in the article and it costs nothing but discipline. Adopt it verbatim if you like.
High-risk actions requiring callback verification
The following are never actioned on the strength of an inbound call, an email, or a message — regardless of who appears to be asking, how urgent it sounds, or how senior they claim to be:
1. Changing bank or payment details for any supplier, customer or employee.
2. Making an unscheduled or out-of-cycle payment.
3. Granting, resetting or changing account access.
4. Changing a delivery address on an existing order.
5. Releasing customer or staff personal information.
6. Changing contact details on an account.
Verification method: hang up and call back on a number already held in our own records — never a number supplied during the request, and never a number in the email that arrived with it. If the person cannot be reached on the held number, the action waits. No staff member will ever be criticised for delaying one of these six actions. That last sentence is not padding; it is the part that makes the policy survive contact with a persuasive caller.
| Why each element is there | The failure it prevents |
|---|---|
| A number from our own records | Attackers supply their own callback number and answer it professionally. This is the most common way callback policies fail. |
| Regardless of seniority claimed | Impersonating an executive is the highest-yield version of the attack, and juniors will not challenge a boss without explicit permission. |
| The action waits if unreachable | Removes the loophole an urgent caller will immediately reach for. |
| Nobody gets criticised for delaying | Staff break policy because they fear being unhelpful more than being defrauded. Fix the incentive and the policy holds. |
The Indirect Win: Shadow Channels
Worth its own section because businesses never file it under security.
When calls go unanswered, staff improvise. A personal mobile number for a good customer. A messaging group with a supplier. A personal email because the shared inbox is chaos. Each is a rational response to an operational failure. Collectively they are business communication on channels the business does not control, cannot log, cannot search, and cannot revoke when somebody leaves.
The offboarding hole
Disabling a platform account removes access. It does not remove a customer's habit of texting a former employee's personal number, and no policy fixes that retrospectively.
The evidence hole
If the commitment was made in a message on somebody's personal phone, you cannot produce it. In a dispute, that is the same as it never having been said.
The detection hole
The pattern analysis in this article only sees traffic on the platform. Every conversation that leaves the platform leaves the detection surface as well.
The fix is boring
Answer the calls. AI agents covering after hours and peak periods, in natural Australian accents, remove the reason the workarounds exist. Solve the operational problem and the security problem goes with it.
Where AI Does Nothing and Your Money Should Go Elsewhere
Four categories, and we would rather say this than sell you the wrong thing.
Do these first, or the rest is decoration
Multi-factor authentication on everything that supports it, and patching on a schedule somebody owns. Neither is interesting and both outperform anything in this article. A business with sophisticated call analytics and shared passwords has bought the fascinating control and skipped the effective one.
| Threat | Why AI on voice cannot help | What does |
|---|---|---|
| Stolen credentials | The attacker is already inside and authenticated. Nothing on the phone layer sees it. | Multi-factor authentication, password manager, conditional access. |
| Unpatched vulnerabilities | The exploit path does not touch the voice channel. | Patching, asset inventory, end-of-support tracking. |
| Malicious insider | The activity looks authorised because it is authorised. | Least privilege, separation of duties, dual authorisation on payments. |
| No escalation process | AI produces alerts. Without someone empowered to act, alerts accumulate and manufacture false confidence. | A named person, a defined authority, and a documented out-of-hours path. |
The Compliance Clock
Three dates that intersect with anything you deploy on the voice channel.
| When | What | Relevance to voice and messaging |
|---|---|---|
| Live since 1 July 2026 | SMS Sender ID Register — unregistered branded sender IDs are labelled Unverified | If you send branded SMS and have not registered, your messages now carry a warning label to your own customers. |
| From 1 September 2026 | Scams Prevention Framework rules commence, with broader obligations following from 31 March 2027; penalties up to $50 million; external dispute resolution membership requirements | Banks, telcos and digital platforms carry legal duties to prevent, detect, disrupt and respond to scams. Businesses sit downstream of those duties and will feel them as process changes. |
| From 10 December 2026 | Automated decision-making transparency obligation under the Privacy Act | Covered in the next section. This is the one most likely to catch a business unaware. |
Disclosure: 10 December 2026
The obligation, plainly
From 10 December 2026, entities that use personal information in automated decision-making capable of affecting a person's rights or interests must disclose in their privacy policy the kinds of personal information used and the kinds of decisions made. The drafting is broad — it captures rule-based tools and automated assessment technologies, not only systems anybody calls AI — the regulator has signalled a wide reading, and penalties for serious breaches are very large.
Why it reaches phone systems: an AI agent that qualifies a lead, prioritises a queue, decides who reaches a human quickly, or scores an interaction is plausibly making a decision affecting rights or interests using personal information. The disclosure itself is manageable. The risk is entirely that nobody connects the AI switched on by operations with the privacy policy last reviewed in 2023.
- Inventory every automated decision your systems make — including the ones that predate anybody using the word AI. Call routing rules count.
- Get that inventory to whoever maintains the privacy policy, well before December rather than in it.
- Ask each vendor what their disclosure position is. A vendor who has not considered this will not help you with it, and that is worth knowing now.
A Thirty-Day Plan, Ordered by Return
Deliberately ordered by value rather than by novelty, which means the technology comes last.
- Days 1–2. The callback policy. Adopt the six high-risk actions above, circulate it, and say the last sentence out loud: nobody gets criticised for delaying one of these. Free, immediate, and the single highest-return item here.
- Days 3–7. Multi-factor authentication and patching. If either is incomplete, stop reading about AI and finish them. Nothing below substitutes.
- Days 8–10. Name the escalation path. Who is called when something looks wrong, what they are authorised to do, and what happens at 2am. Write it down or it does not exist.
- Days 11–14. Ask who watches your call traffic. Put the question to your provider: what triggers a toll fraud block, what are the limits, and who is on the other end of the alert. The answer is either reassuring or extremely informative.
- Days 15–21. Turn on transcripts and recording, properly. With a retention period, an access rule, a stated purpose and staff consultation done first. Doing this badly costs more than not doing it.
- Days 22–25. Move from sampled review to complete review. Frame it as policy compliance, not performance monitoring, and say what will and will not be used for.
- Days 26–28. Close the shadow channels. Work out why staff use personal numbers and fix that reason — usually it is unanswered calls, which is an operational fix rather than a policy one.
- Days 29–30. The ADM inventory. List every automated decision and hand it to whoever owns the privacy policy.
What to Ask Us, and Everyone Else
| Question | What a good answer sounds like |
|---|---|
| Where is the processing done? | A named jurisdiction. Ours is Australia, on infrastructure we operate. |
| Is my audio or transcript retained, for how long, and can I set it? | A specific period, configurable, documented. |
| Is anything used to train a model? | No by default, in the contract rather than in a blog post. |
| Who are the subprocessors? | A list. You inherit their posture whether you asked or not. |
| Can a human always override? | Yes. If no, do not use it for a security decision. |
| What triggers a block, and who acts on it? | Named thresholds and a staffed response. An unstaffed alert is theatre. |
| Do you detect synthetic voice? | An honest "not reliably, here is what we do instead". An emphatic yes is a warning. |
| What does this not protect me from? | A real answer. A vendor who cannot name their product's limits has not examined it. |
The summary
Serious voice fraud is a sequence of unremarkable calls, which is why sampled human review cannot find it and cross-call pattern analysis can. AI materially helps with patterns, traffic anomalies, verification consistency, transcript evidence, complete review and closing shadow channels. It partially helps with synthetic voice — so authenticate the request, not the voice, using a callback policy on a defined list of high-risk actions. It does nothing for stolen credentials, unpatched systems, malicious insiders, or an alert nobody answers. Do the callback policy, multi-factor authentication and patching first; they are free or cheap and they outperform everything else. Then the technology, with disclosure sorted before 10 December 2026.
Related reading: voice cloning and vishing in depth, what a breach looks like from outside, the December 2026 disclosure deadline in detail, reviewing every call rather than a sample, and the SMS Sender ID Register.
Frequently Asked Questions
How does a social engineering attack on the phone actually work?
As a sequence of unremarkable calls rather than one suspicious one, which is the whole reason it succeeds. A representative pattern runs like this. On a Tuesday a pleasant caller confirms the correct email address for accounts — thirty seconds, no request, nothing given away that is not on the website; the purpose is simply to make a call from that number seem normal. On Thursday someone asks who approves changes to supplier details, mentions having dealt with the business before, and is told a first name, because it is a reasonable question and the answer is not secret. The following Monday a caller references both the accounts email and the approver's first name, and is therefore treated as legitimate — knowledge being mistaken for authorisation, which is the core mechanism. Then on Wednesday at ten to five, the day before the payment run, a voice indistinguishable from a supplier's accounts manager explains their bank has changed, apologises for the timing, notes the detail is in an email already sent, and asks that Friday's payment go to the new account. Three pressures are deployed at once — time, established precedent, and a document you have not read — on somebody softened by three earlier contacts they will not consciously remember. Not one of the first three calls would trouble any reviewer.
Why can't reviewing recorded calls catch this kind of fraud?
Because human review looks for a bad call and the threat is a series of good ones. Take each call in a typical sequence and ask what a reviewer would note if they heard it in isolation. The spelling check: nothing, a courteous thirty-second call handled well, possibly a positive example. The process question: nothing, a staff member answered a reasonable question helpfully, arguably good service. The credibility call: nothing, the caller knew internal details so of course they were treated as known. Only the final request is catchable, and that is after the money has gone — and only if that specific call happens to fall in whatever sample somebody reviews, which for most businesses is around two per cent chosen by whoever had time. No amount of diligence fixes this, because the information needed to spot the attack does not exist inside any single call. It exists in the relationship between four calls spread across nine days, taken by different staff in different departments with no shared memory of each other's conversations. That is not a human attention problem, it is a data problem: the signal is in the aggregate. Pattern analysis across an entire call history is structurally able to see it, and sampled review is structurally unable to.
Can AI detect a cloned or deepfake voice on a business call?
Not reliably, and a confident vendor claim here should increase your scepticism rather than your comfort. Automated detection of synthetic audio is improving but generation is currently outrunning detection, and a control you cannot trust is worse than no control at all because it delivers the confidence without the protection. The durable answer is to change what you are authenticating. Stop trying to authenticate the voice and start authenticating the request: voice as an identity proof is finished permanently, and no technology restores it. Define a list of high-risk actions and require out-of-band verification for every one of them, without exception and regardless of who appears to be asking. The list should cover changing bank or payment details for any supplier, customer or employee; making an unscheduled or out-of-cycle payment; granting, resetting or changing account access; changing a delivery address on an existing order; releasing customer or staff personal information; and changing contact details on an account. Verification means hanging up and calling back on a number already held in your own records — never one supplied during the request or contained in the accompanying email. AI's genuine contribution is applying that rule identically at ten to five on a Wednesday, which is exactly when the request arrives.
What should a callback verification policy actually say?
It needs four elements and the fourth is the one businesses omit and then wonder why the policy failed. First, a specific list of high-risk actions never actioned on the strength of an inbound call, email or message: changing bank or payment details for any supplier, customer or employee; unscheduled or out-of-cycle payments; granting, resetting or changing account access; changing a delivery address on an existing order; releasing customer or staff personal information; and changing contact details on an account. Second, a verification method that specifies calling back on a number already held in your own records, never a number supplied during the request and never one contained in the email that arrived with it — because attackers supply their own callback number and answer it professionally, and this is the most common way callback policies fail in practice. Third, an explicit statement that the rule applies regardless of seniority claimed, since impersonating an executive is the highest-yield version of the attack and junior staff will not challenge a boss without written permission to do so. Fourth, and critically: no staff member will ever be criticised for delaying one of these actions. Staff break security policy because they fear being unhelpful more than being defrauded, so fixing that incentive is what makes the policy survive contact with a persuasive caller.
What are shadow channels and why are they a security problem?
Shadow channels are the workarounds staff invent when the official channels do not work — a personal mobile number given to a good customer, a messaging group with a supplier, a personal email address used because the shared inbox is unmanageable. Each is a rational individual response to an operational failure, usually the failure to answer calls reliably, and collectively they represent business communication happening on channels the business does not control, cannot log, cannot search and cannot revoke. That creates three specific holes. The offboarding hole: disabling a platform account removes access, but it does not remove a customer's habit of texting a departed employee's personal number, and no policy fixes that retrospectively. The evidence hole: if a commitment was made in a message on somebody's personal phone, you cannot produce it, which in a dispute is the same as it never having been said. And the detection hole: any pattern analysis only sees traffic on the platform, so every conversation that leaves the platform also leaves the detection surface. The fix is unglamorous and it is operational rather than disciplinary. Answer the calls — including after hours and at peak, which is what AI phone agents are genuinely good at — and the reason for the workarounds disappears. Solve the operational problem and the security problem leaves with it.
What should I do first if my budget is limited?
The three highest-return items are free or cheap and none of them is AI. First, the callback policy: adopt the list of high-risk actions, circulate it, and say out loud that nobody will be criticised for delaying one of them. That takes two days, costs nothing, and directly addresses the attack that ends in a payment going to the wrong account. Second, multi-factor authentication on everything that supports it, plus patching on a schedule somebody actually owns. Neither is interesting and both outperform anything on the AI list — a business with sophisticated call analytics and shared passwords has bought the fascinating control and skipped the effective one. Third, name the escalation path in writing: who gets called when something looks wrong, what they are authorised to do, and what happens at two in the morning. An alert with nobody behind it is worse than no alert, because it manufactures false confidence. Only after those three should you spend on capability. Then the order is: ask your provider what watches your call traffic overnight and who acts on it; turn on transcripts with a retention period, access rule, stated purpose and staff consultation done beforehand; move from sampled to complete call review framed as policy compliance rather than performance monitoring; and close the shadow channels by fixing whatever operational failure created them.
Which compliance dates should I have in my calendar?
Three, and they all intersect with anything you deploy on the voice or messaging channel. The SMS Sender ID Register has been live since 1 July 2026, and unregistered branded sender IDs are now labelled Unverified — so if you send branded SMS and have not registered, your messages currently carry a warning label to your own customers, which is a marketing problem as much as a compliance one. From 1 September 2026 the Scams Prevention Framework rules commence, with broader obligations following from 31 March 2027, penalties up to $50 million, and external dispute resolution membership requirements; banks, telcos and digital platforms carry legal duties to prevent, detect, disrupt and respond to scams, and ordinary businesses sit downstream of those duties and will experience them as process changes rather than direct obligations. From 10 December 2026, the automated decision-making transparency obligation under the Privacy Act requires entities using personal information in automated decisions capable of affecting rights or interests to disclose in their privacy policy the kinds of information used and decisions made. That last one is the most likely to catch a business unaware, because it reaches AI agents that qualify leads, prioritise queues or score interactions, and the failure mode is simply that nobody connects what operations switched on with a privacy policy last reviewed years ago.