Your First AI Agent: A Practical 90 Day Plan

A business we spoke to earlier this year had an AI answering service running for five weeks and then turned it off. When we asked what went wrong the answer was interesting, because nothing did, exactly. It understood people fine. It booked appointments correctly. The problem was that nobody could say whether it was better than what came before, because what came before was a voicemail box that nobody had ever measured either. So when one customer complained about a call in week four, there was no counterweight. There were no transcripts anybody had read, no agreed number, no before and after, just one loud data point and a general unease, and the decision made itself. That story is far more common than the dramatic failures, and it is the reason a deployment plan is worth more than a better model. Gartner's estimate that more than forty per cent of agentic AI projects will be cancelled by the end of 2027 names escalating cost, unclear business value and inadequate risk controls, and the middle one is the quiet killer: not that the thing failed, but that nobody could demonstrate it had succeeded. The fix is boring and it works. Pick one narrow job. Write down what it may never do. Decide where a call goes when it stops. Write one sentence with a number in it that says what good looks like by day sixty. Then run it for a fortnight where no customer ever hears it, and read what it would have said. Everything after that is comparatively easy.

Deployment Plan · 90 Days

Week One Is Not Building. Week One Is Deciding.

The businesses that get an AI agent working and the ones that quietly switch it off six weeks later mostly differ in what happened before anybody touched a configuration screen. One group wrote down the single job, the list of things it may never do, where a call goes when it should stop, and the sentence that would tell them in sixty days whether it worked. The other group switched it on and hoped. Here is the first version, laid out by week, with the numbers you read at the end of each phase and the exit test that says whether you go on to the next one.

📅 ⏱ 15 min read 🇦🇺 Australian owned · Australian network · Australian support
TL;DR

Week one is writing, not building. One job in a paragraph, a list of prohibitions, a named handover destination for every hour of the day, and one success sentence with a number and a date in it. If you cannot write the success sentence, you have an interest rather than a project. Pick a job with high volume, low variance and low consequence, which usually means after hours calls, bookings and reschedules, status enquiries or call notes rather than the complicated thing eating a senior person's week. Connect it before you launch it. An agent that cannot read your customer record and write back to it is doing an impression of your business from the outside. Run a fortnight of shadow, where it works on live calls and nothing reaches a customer. Then after hours, then supervised overflow, then the front line, with an exit test at each step rather than a date. Measure five numbers, and pair containment with repeat contact in 72 hours, because containment that pushes the call into tomorrow costs you two contacts instead of one.

Why Plans Beat Better Models

The capability question is mostly settled. Voice AI is reported to be handling close to a fifth of inbound contact centre volume in 2026, against roughly six per cent two years earlier, and the average Australian organisation using AI is running about eleven agents at once. Things are clearly working somewhere. At the same time, only around seventeen per cent of organisations have deployed agents while more than sixty per cent intend to within two years, and Gartner expects over forty per cent of agentic projects to be cancelled by the end of 2027 on grounds of cost, unclear value and weak risk controls.

Put those together and the picture is clear enough. The technology works well enough to carry real volume. The average implementation does not, and the difference is almost entirely in the sequence. This plan assumes nothing about your model, your vendor or your industry. It assumes only that you would rather find out in week three than in month six.

The failure that has no dramatic moment

Most abandoned deployments were not disasters. Nobody could demonstrate they were working, so when one complaint arrived there was nothing on the other side of the scale. Every step below exists to make sure that by the time a complaint arrives you have a hundred transcripts, five numbers and a written expectation to weigh it against.

Week 1: The Four Documents

None of this involves software. All four fit on two pages, and every one of them is cheaper to write now than to reconstruct after an incident.

DocumentWhat it saysTest that it is good enough
1. The jobOne paragraph in plain words: what the agent does, for which calls, and what is explicitly out of scope.It fits in a paragraph. If it needs a page, the job is too big and should be split.
2. The never listFlat prohibitions with no conditions attached. Not guidance. Things it may never do under any circumstances.Every item is enforceable by withholding a capability rather than by asking the agent nicely.
3. The handoverWhat triggers a handover, where it goes at 10am, at 8pm and at 2am, what gets carried across, and what the caller hears.There is a named destination for every hour of the week, including the ones where the answer is a commitment with a time rather than a person.
4. The success sentenceOne sentence, one number, one date, agreed by whoever owns the outcome.Somebody who disagreed with it would know they disagreed. "Improve customer experience" fails this test.

A worked success sentence, for a trade business: "By 15 November, at least half of calls arriving between 5pm and 8am are fully dealt with without anybody ringing back the following morning, and the number of after hours callers who hang up before anything happens is zero." That is measurable on the day, it is arguable in advance, and it cannot be quietly rewritten in December to describe whatever happened.

Choosing the Job

Three variables, and you want the same answer on all three. High volume, low variance, low consequence.

📈

High volume

At least a few dozen a week. Not to justify the cost, though it helps, but because you need repetition to learn anything at all. A job that happens twice a week will take a year to produce enough evidence to judge.

📐

Low variance

The same handful of shapes over and over. High variance work is where agents demo brilliantly and behave inconsistently in production. Start on flat ground and expand into the rough once you know what your agent does when it is unsure.

🛟

Low consequence

A mistake is recoverable inside a day. Not because agents err more than people, but because your process for catching and fixing errors does not exist yet, and you want to build that habit somewhere forgiving.

Run your candidates through that and the job everybody wants to automate first, the complicated one that eats a senior person's week, usually scores badly on all three: low volume, high variance, high consequence. It will be a good candidate in about a year, once the connections, the guardrail habits and the review routine exist. The jobs that actually pay back in a first quarter are duller.

JobWhy it works first
After hours answeringThe current alternative is a mailbox nobody opens until morning or a ring that records nothing. Almost any improvement is visible, and the downside is bounded.
Booking and reschedulingAround half of booking calls in most businesses are reschedules, which is the most mechanical conversation you have all day.
Job or order statusReplaces "let me find out and call you back", which reliably costs two calls and a note that nobody writes.
Call notes and record updatesQuietly the highest value of the four, because it fixes data quality in everything downstream of the phone.
Overflow at peakCatches the calls that would be abandoned. Abandoned calls are invisible in most reporting and are the most expensive thing a busy business does.

Three of those five are work that is currently not being done by anybody, and that is not an accident. An agent that picks up unhandled work does not require any person to give anything up, so it gets a fair trial. An agent introduced as a replacement for somebody's work gets audited by that person, and they will find the three calls it got wrong well before they find the ninety it got right.

The Never List

Seven items cover most businesses. Add your own, keep them absolute, and enforce the serious ones structurally.

NeverBecause
Quote a price that is not on the published listA price stated on a recorded call is a representation, and under Australian Consumer Law it belongs to the business regardless of who said it.
Commit to a timeframe"Someone will be there this afternoon" is the classic over-eager promise. It may state a booked appointment. It may not predict one.
Change an account without verificationDefine the check and define what a failed check does, making sure the failure path is not simply an easier second question.
Mention any other customer or jobConstrain the tool rather than the instruction: the lookup should only ever return records tied to the identified caller.
Claim to be a personIt should not lie if asked, and disclosure up front produces fewer complaints than letting people work it out themselves.
Handle a caller in distressDefine the trigger words and route straight to a human. Containment is the wrong measure for these calls.
Try a third timeTwo attempts to understand, then hand over. A third loop converts mild frustration into a complaint about the business rather than the system.

Enforce the never list in the tools, not in the wording. An instruction that says do not issue refunds is a preference. Not giving the agent a refund tool is a control. Where the consequence is genuine, take the capability away rather than asking for restraint, and save instruction-level rules for the things you cannot express structurally. This single habit prevents most of the incidents that end up being written about.

Week 2: Connect It Before You Launch It

Roughly eighty-eight per cent of contact centres report using AI in some form, but only about a quarter say they have integrated it properly, and that gap is where most of the disappointment in this category comes from. An agent that cannot see your systems is guessing about your business from the outside, and it sounds like it.

Read access, at minimum: the customer record matched from the calling number, open jobs or orders, the calendar, the knowledge your team actually uses, and the last few interactions. Most of what people perceive as intelligence in a good agent is context. The same model sounds vastly cleverer when it knows the caller has a job open and a technician booked for tomorrow.

Write access, which is where the return is: notes on the record, tasks for people, bookings in the calendar, status updates. Reading saves the customer time. Writing saves your team's, and it compounds, because every downstream report gets better when the notes are actually there.

Connected how: through documented APIs and webhooks in both directions rather than a closed marketplace. The Model Context Protocol has become the common way for agents to reach tools and data across the major AI ecosystems, and it now sits under independent stewardship. You do not need to understand the specification. You do need to know whether your platform can expose your systems to an agent through an open standard, or whether each connection is a bespoke build somebody charges for. Our piece on open APIs for voice and SMS covers what to ask for.

Weeks 3 and 4: Shadow

The agent runs against real calls and produces exactly what it would have said and done. Nothing reaches a customer. No tool it calls changes anything. A person reads a sample every day, which takes about twenty minutes once you have the habit.

This fortnight is the highest value two weeks in the entire project and it is the one most often skipped, because it produces nothing visible. What it actually produces is the thing that saves you later: a hundred transcripts you have read, a list of surprises that have each become either a guardrail or a fix, and a team that has seen the output before a customer did.

Exit test: you can predict what it will do. Not perfectly, but well enough that reading a transcript rarely surprises you. If week four is still surprising you weekly, stay in shadow. The cost of another fortnight here is trivial compared with the cost of finding the same surprises in front of customers.

Weeks 5 to 7: After Hours Only

Live, but only on calls that would currently reach a mailbox or ring out. The risk is bounded by the fact that the alternative was nothing at all, which makes this the cheapest possible place to discover what you missed.

Three things get tested here that shadow cannot test. Whether escalation actually works at 2am, which is a question about your rosters and ring groups rather than about the agent. Whether notes are landing in the right records, which you verify by opening the record rather than by reading the agent's log. And whether anything embarrasses the business, which you find by asking the team to flag calls and making it take one click.

The 2am escalation is not an AI problem

Escalating at two in the morning to a ring group with nobody in it is not an escalation, it is a hang up with extra steps. If there is genuinely nobody available, the honest configuration is a commitment with a time attached and an SMS confirming it, which customers accept readily. Deciding this properly is the difference between an after hours agent that builds trust and one that quietly loses work at night.

Weeks 8 to 11: Supervised Overflow

Daytime now, but only calls that would otherwise queue past a threshold you set. Somebody owns the review, half an hour a day, and it sits inside their workload rather than on top of it, because a review that is nobody's actual job stops happening in week nine.

Exit test, and read this one carefully: containment stable across three weeks, and repeat contact flat or falling. Not rising containment. Stable containment with repeat contact behaving. A rising containment figure alongside rising repeat contact means the agent is getting better at ending calls and no better at resolving them, which is the most common way a deployment looks successful on a dashboard while costing more than it saves.

Weeks 12 and 13: Narrow Front Line

It answers first, on the defined job only, and hands over everything else. From here, autonomy widens one decision at a time, each with its own short review period.

That phrase is worth dwelling on, because autonomy is a dial rather than a switch and treating it as a switch is behind most of the incidents in this field. Inside one job, an agent might be fully autonomous about looking something up, allowed to change a booking only with explicit confirmation, and forbidden from issuing anything financial at all. Three different levels in one conversation. Every deployment that has caused a business a real problem gave the whole job a single level of autonomy, and it was the highest one.

At the end of week thirteen you answer the success sentence. Met, or honestly not met. Both of those are good outcomes, because both tell you what to do next. The only bad outcome is not being able to say.

The Five Numbers

NumberWhat good looks like
Containment, the share finished without a personRising then settling, usually 40% to 70% for a well-scoped first job. Anything approaching 100% means the scope is trivial or handover is broken.
Resolution, the share where the caller's actual purpose was achievedNeeds sampling. Read and score fifty calls a month. The log cannot tell you this and no dashboard can either.
Repeat contact within 72 hoursFlat or falling. This is the decisive number and it is the one most often missing from vendor reporting.
Handover qualityAbove 90% of handovers where the person who picked up had the context and the caller did not start again.
Cost per resolved contactIncluding the review time. A comparison that excludes supervision is dishonest and gets caught at the first budget review.

Report containment and repeat contact on one line, never two. A caller who was told something unhelpful at 7pm and rings back at 9am shows up as a success in most vendor dashboards, and has cost you two contacts instead of one plus whatever goodwill was spent. Treating the pair as a single measure is the simplest protection against a project that looks good for two quarters and is quietly expensive. The five queue numbers your phone system should already be reporting sit underneath all of this.

The Part About Your Team

Three things, and they cost nothing.

Tell them before it goes live, not after. Staff who discover an agent because a customer mentioned it become its most motivated critics, and they are entitled to. Show them the shadow transcripts in week three. It is much harder to be suspicious of something you have already read a hundred pages of.

Give them a one click way to flag a bad call, and act on the flags visibly. The flagging mechanism matters less than the visible acting. Two or three flags that produce a change in the same week will do more for adoption than any amount of explanation.

Be straight about what it means for the work. If the honest answer is that it takes the after hours burden off an on call roster and nobody's hours change, say so. If the honest answer is that it absorbs growth so you do not hire a third person in March, say that instead. The version that causes trouble is the one where nobody says anything and everybody assumes the worst.

Six Questions for the Vendor

Gartner uses the term agent washing for the practice of rebranding assistants, robotic process automation and chatbots as agentic without the capability underneath, and its assessment is that only a small fraction of vendors making the claim have it. You do not need to assess architecture. You need six questions and one phone call.

QuestionWhat a real answer sounds like
Which tools can it call, and can I add my own?A named list and a documented way to add yours. Vagueness means it can speak but not act.
Show me a transcript where it did something you did not script.They can find one, because they read transcripts too. If every demo is the same demo, there is a script under it.
What does it do when it is not confident?A described behaviour, not a reassurance that it always is. This is the clearest single signal of a real implementation.
Where is my data processed and what is retained?A straight answer about which components run where and for how long. Vagueness here becomes a privacy problem later.
What is your containment, and what is repeat contact beside it?Both numbers from real deployments. A vendor who has never been asked the second one is telling you something.
What happens to what I build if I leave?Open APIs, exportable configuration, your numbers and your records. If the answer is complicated, the answer is no.

Then ring the demo number and behave like a real customer rather than an evaluator. Interrupt it mid sentence. Change your mind halfway through. Give a suburb and a vague description rather than a reference number. Put two requests in one sentence. Ask for a person. Every one of those is ordinary and every one separates a genuine agent from a menu with a better voice. We wrote a longer version of that test in the ten call test.

What the Second Agent Costs

Much less, and this is the argument for taking the first one slowly. By the end of ninety days you own four things that do not need building again: the connections into your customer record and calendar, a written guardrail pattern you can copy, a review routine that somebody already does, and a team that has watched one of these work. The second agent is typically a fortnight, the third is a week, and this is why the organisations getting real value are running many small agents with narrow jobs rather than one large one with a broad job.

It is also why the sequence matters more than the ambition. A business that spends ninety days getting after hours answering genuinely right is further ahead at the end of the year than one that spent ninety days trying to automate everything and switched it off in week six.

Where We Fit

VOCPhone runs the agent inside the platform that already carries your calls and messages rather than beside it. That has three practical consequences. There is no forwarding hop and no second provider sitting in the audio path, so the call behaves like a call. The agent has the context from the first second, because the platform already holds the number, the matched record and the history. And escalation is a route into your existing ring groups, queues and on call rotations rather than an integration somebody has to build.

We run the ninety days with you in the shape described above, including the fortnight of shadow where no customer hears anything, and you get the transcripts from day one plus the five numbers weekly. We own and operate our own network in Australia, the platform is Australian hosted and supported, and the APIs are open, so the connections you build are yours to reuse.

Get the first agent scoped properly

Send us a week of your after hours call data. We will come back with the job in a paragraph, a never list, a handover destination for every hour, and a success sentence you can argue with before anything gets built.

Talk to us Or call 1300 663 222

Frequently Asked Questions

How long does it take to set up an AI agent for a business?
Ninety days for the first one if it is done properly, and roughly a fortnight for the second, because by then the connections, the guardrail pattern and the review routine already exist. The sequence has six stages, each with an exit test rather than a deadline. Week one is writing: the job in one paragraph, the never list, the handover destination for every hour of the day, and one success sentence with a number and a date in it. Week two is connecting it to your customer record, calendar and job or order system, with both read and write access. Weeks three and four are shadow running on live calls where nothing reaches a customer and no tool changes anything, exiting when you can predict what it will do. Weeks five to seven are live but after hours only, on calls that would otherwise reach a mailbox. Weeks eight to eleven are supervised daytime overflow, taking only calls that would queue past a threshold, exiting when containment is stable across three weeks and repeat contact is flat or falling. Weeks twelve and thirteen put it on the front line for the defined job only, widening autonomy one decision at a time. If any phase is still surprising you at its end, stay in it. Another fortnight of shadow is far cheaper than finding the same surprises in front of customers.
What should my first AI agent actually do?
Choose a job with high volume, low variance and low consequence, and accept that a good first job feels unambitious. High volume means at least a few dozen instances a week, because repetition is what lets you judge it at all; a job that occurs twice a week takes a year to generate evidence. Low variance means the same handful of shapes repeatedly, since agents demonstrate brilliantly and behave inconsistently where the work has many forms. Low consequence means a mistake is recoverable within a day, not because agents err more than people but because your process for catching errors does not exist yet. Five jobs fit for most Australian businesses: after hours answering, where the current alternative is a mailbox nobody opens; booking and rescheduling, since around half of booking calls are reschedules and reschedules are the most mechanical conversation you have; job or order status, replacing "let me find out and call you back" which costs two calls and a note nobody writes; call notes written back into the customer record, quietly the highest value because it fixes data quality downstream; and overflow at peak, which catches calls that would otherwise be abandoned. Three of those five are work nobody is doing now, which means nobody has to give anything up for the trial to be fair.
What should an AI agent never be allowed to do?
Write the prohibitions as flat rules with no conditions, because guidance gets interpreted and rules get enforced. Seven cover most businesses. Never quote a price that is not on the published list, since a price stated on a recorded call is a representation and under Australian Consumer Law it belongs to the business whoever said it. Never commit to a timeframe; it may state a booked appointment but may not predict one. Never change an account without verification, with a defined failure path that is not simply an easier second question. Never mention another customer or job, enforced by constraining the lookup tool to records tied to the identified caller rather than by instruction. Never claim to be a person, and preferably disclose up front, because businesses that disclose receive fewer complaints than those whose customers work it out. Never handle a caller in distress; define the trigger words and route straight to a human. And never try a third time, meaning two attempts to understand and then hand over, because a third loop turns mild frustration into a complaint about the business. The structural point matters more than the list: enforce the serious ones by withholding the capability. An instruction not to issue refunds is a preference. Not giving the agent a refund tool is a control.
What is shadow running and is it worth two weeks?
Shadow running means the agent works on real live calls and produces exactly what it would have said and done, while nothing reaches a customer and no tool it calls changes any record. A person reads a sample every day, which takes about twenty minutes once the habit exists. It is the highest value fortnight in the whole project and the one most often skipped, because it produces nothing visible: no calls answered, no metric improved, nothing to report upward. What it actually produces is the thing that protects the project later. By the end of it you have read around a hundred transcripts, every surprise has been converted into either a guardrail or a fix, and your team has seen the output before any customer did. That last part is worth the two weeks on its own, because staff who have read a hundred pages of transcripts are much harder to alarm than staff who first hear about the agent from a customer. The exit test is that you can predict what it will do; not perfectly, but well enough that reading a transcript rarely surprises you. If week four is still producing weekly surprises, stay in shadow. Another fortnight costs almost nothing compared with finding those same surprises in front of paying customers.
Why do AI agent deployments get abandoned?
Gartner expects more than forty per cent of agentic AI projects to be cancelled by the end of 2027, citing escalating cost, unclear business value and inadequate risk controls, and in practice the middle one does most of the damage. The typical abandoned deployment was not a disaster. It understood people, it did the task, and nothing dramatic happened. What was missing was evidence: no transcripts anybody had read, no agreed number, no before and after, often because the thing it replaced had never been measured either. So when a single customer complained in week four there was nothing on the other side of the scale, and the decision made itself. The other recurring causes are all upstream of the build. The first job was too big, usually the complicated one eating a senior person's week, which scores badly on volume, variance and consequence all at once. Guardrails were written as guidance rather than prohibitions. Handover was an afterthought, so the calls that mattered most were handled worst. Nobody owned the daily review, so the thing never improved after week two. And containment was celebrated on its own, without repeat contact beside it, which hides a deployment that defers work rather than resolving it.
How do I know if a vendor's AI agent is real?
Gartner calls the alternative agent washing: rebranding assistants, robotic process automation and chatbots as agentic without the underlying capability, and its assessment is that only a small fraction of vendors making the claim actually have it. You do not need to evaluate architecture. Ask six questions. Which tools can it call, and can I add my own, where a real answer is a named list plus a documented way to add yours. Show me a transcript where it did something you did not script, which a real vendor can find because they read transcripts too. What does it do when it is not confident, where you want a described behaviour rather than a reassurance that it always is, since confidence handling is the clearest single signal of a real implementation. Where is my data processed and what is retained. What is your containment and what is repeat contact beside it, which needs both numbers from real deployments. And what happens to what I build if I leave. Then ring the demo number and behave like a customer rather than an evaluator: interrupt it mid sentence, change your mind halfway, give a suburb and a vague description instead of a reference number, put two requests in one sentence, and ask for a person. All of those are ordinary customer behaviour and all of them separate a real agent from a menu with a better voice.
Should I tell my staff and customers that an AI is answering?
Yes to both, and for different reasons. With staff, tell them before it goes live rather than after, because people who discover an agent because a customer mentioned it become its most motivated critics and are entitled to be. Show them the shadow transcripts in week three, since it is much harder to be suspicious of something you have already read a hundred pages of. Give them a one click way to flag a bad call and, more importantly, act on the flags visibly: two or three flags producing a change in the same week does more for adoption than any amount of explanation. Be straight about what it means for the work as well, whether the honest answer is that it removes the after hours burden without changing anybody's hours, or that it absorbs growth so a third hire is not needed in March. The version that causes trouble is the one where nobody says anything and everybody assumes the worst. With customers, disclosure up front produces fewer complaints than letting people work it out on their own, and the agent should never claim to be human if asked directly. It costs nothing and it changes how people treat the conversation, usually for the better.

What to Read Next

Your next reads

VOCPhone, the Australian-owned cloud phone platform that owns and operates its own network. vocphone.com | 1300 663 222

Related Articles