Leet Force
Back to Blog
DevelopmentUpdated

AI Phone Agents: What They Can Actually Handle, What a Minute Costs, and the 200ms Problem

Human turn-taking runs at 0–200 ms; the best agents manage 500. A well-built minute costs three to seven cents before margin. Gartner says cost per resolution passes $3 by 2030. What to let it answer, what never to, and how we build one.

Hammad Tariq
By Hammad TariqCTO & Co-Founder, Leet Force
AI Phone Agents: What They Can Actually Handle, What a Minute Costs, and the 200ms Problem

The phone in a dental practice rings most at 12:40, when the receptionist is checking someone out and the hygienist is on the other line. The call goes to voicemail. The 42% of people who leave a message get a callback at three; the rest ring the next practice on the map. Everyone in the building knows this happens and nobody can fix it, because the fix is a second receptionist for forty minutes a day.

That forty minutes is what AI phone agents are actually for. Not "replacing the front desk" — the demos that claim that are selling the wrong thing — but answering the calls that currently go nowhere, doing the four boring things well, and handing the rest to a person with the context attached.

We've now built enough of these to have opinions, and the opinions are mostly about three numbers: 200 milliseconds, ten cents, and 91 percent. Here's what each one means for whether you should do this, and how.

First, the missed-call statistic everyone quotes

You've seen "62% of calls to small businesses go unanswered". It's on every vendor's homepage. It comes from a 2016 mystery-shopper exercise across 85 businesses, and I'd not build a business case on it. What I trust more is your own phone system's report, which every VoIP provider produces and almost nobody reads: calls received, calls answered, calls abandoned, average time to answer, by hour of day. Pull last month's. If abandoned plus voicemail is above 20% in any two-hour window, you have the problem this article is about. If it isn't, save your money and read something else.

The 200 millisecond problem

Human conversation runs on a clock most people never notice until a machine breaks it. In 2009, Stivers and colleagues published in PNAS a study of turn-taking across ten languages on five continents. The gap between one person finishing and the next starting had a mode of 0 to 200 milliseconds in every language measured; the cross-linguistic median was +100 ms. As the authors point out, that's faster than people can name an object out loud. We start planning our reply before the other person finishes, and any system that waits for silence, thinks, and then speaks is already late.

This is the design constraint that decides whether an AI agent feels like a competent receptionist or a phone tree with better vocabulary. A cascaded agent — the standard architecture — has to do five things in sequence before you hear anything:

Latency budget of a cascaded voice agent, stage by stage, against the human 200 millisecond turn gap. Hover a segment. Typical cascaded pipeline, well tuned (milliseconds after the caller stops speaking) Endpointing — waiting long enough to be sure the caller has finished. 300–500 ms. The single largest and least visible cost; too short and the agent interrupts.Endpointing300–500 ms Speech-to-text finalisation — 100–200 ms with a streaming model that has been transcribing all along.STT100–200 ms Language model, time to first token — 300–800 ms depending on model size, prompt length and whether a tool call is needed first.LLM first token300–800 ms Text-to-speech, first audio byte — 100–300 ms with a streaming voice.TTS100–300 ms Telephony and network — 100–200 ms round trip through the carrier and media server.Net100–200 ms ≈ 0.9–2.0 s Human turn gap: mode 0–200 ms, median +100 ms (Stivers et al., PNAS 2009) Speech-to-speech model (one model hears and speaks; no STT/TTS hop) Speech-to-speech — roughly 500–800 ms end to end in our measurements, but you give up the ability to inspect and edit the text between hearing and speaking.≈ 500–800 msfewer hops, less control
Our working latency budget for inbound agents. Hover a stage for what it does and why it costs what it costs.

The sum is somewhere between 0.9 and 2 seconds on a good day, and it's the first item — endpointing, the wait to be sure the caller has actually stopped — that people underestimate. Set it short and the agent talks over "and also—". Set it long and every reply starts with a beat of dead air. Speech-to-speech models like OpenAI's gpt-realtime family collapse the STT and TTS hops and get closer to 500–800 ms, at the cost of giving up the text in between, which is where you'd normally validate what the model is about to say.

You will not reach 200 ms. Nobody does. What we've found works is not chasing the number but designing around it: a short acknowledgement ("let me check that") that streams while the calendar lookup runs; endpointing tuned by call type, because someone reading out a policy number pauses differently from someone asking a question; and never, ever a synthesised "um" to fill the gap, which callers clock instantly and hate.

What a minute actually costs

Vendors quote per-minute prices, but the underlying models bill per token, and the conversion is public. On OpenAI's current pricing, audio is metered at one token per 100 ms of caller speech and one per 50 ms of agent speech — 600 and 1,200 tokens per minute respectively. That gives you the arithmetic:

ComponentBasisCost per minute of call
gpt-realtime-2.1, caller speaking600 tokens × $32 / 1M$0.019
gpt-realtime-2.1, agent speaking1,200 tokens × $64 / 1M$0.077
Blended, roughly half each way—≈ $0.05
Same, on the mini model ($10 / $20)—≈ $0.015
Cascaded: transcriptiongpt-4o-transcribe $0.006/min; Deepgram Nova-3 ≈ $0.005/min$0.005–0.006
Cascaded: text model + TTSSmall text model; tts-1 at $15 / 1M characters, ≈ 900 characters per spoken minute$0.02–0.03
Telephony (US inbound)Carrier per-minute rate≈ $0.01

So the raw cost of a well-built inbound agent is somewhere between three and seven cents a minute, before anyone's margin. Practitioners who've measured real sessions report higher figures — roughly $0.06–0.11 a minute on the flagship model once prompt caching is working, and two to four times that when it isn't, because every turn re-sends the conversation so far. Bundled platforms that include telephony land at $0.10–0.18 a minute in the vendor roundups we've checked, and the "bring your own keys" platforms end up at $0.13–0.31 once you add the model, voice, transcription and carrier they don't include. The headline "$0.05 a minute" on a pricing page is almost never the number on the invoice.

All-in cost per minute of call, four ways of building the same agent. Hover a bar. ALL-IN COST PER CALL MINUTE (RANGE) Custom, mini S2S model$0.03–0.05 — mini speech-to-speech plus carrier. Fine for booking and FAQ; weaker on accents and interruptions.$0.03–0.05 Custom, cascaded$0.04–0.07 — streaming STT, small text model, streaming TTS, carrier. Most control; most engineering.$0.04–0.07 Custom, flagship S2S$0.06–0.11 — flagship speech-to-speech with caching working; two to four times that without.$0.06–0.11 Bundled platform$0.10–0.18 — telephony included; fastest to launch; per-minute margin is the platform's business model.$0.10–0.18 BYOK platform, all-in$0.13–0.31 — platform fee plus the model, voice, transcription and carrier it doesn't include.$0.13–0.31
Ranges from OpenAI and Deepgram list prices, plus published practitioner measurements and vendor roundups as of September 2026. Hover for what's inside each.

Two warnings about that chart. First, the token price is the part that's been falling; it's the carrier, the platform fee and the engineering that don't. Second, Gartner's read of the direction is the opposite of the vendors'. In January 2026 they predicted the cost per resolution for generative AI in customer service will exceed $3 by 2030 — higher than many offshore human agents — as vendors move from subsidised growth to profitability and use cases get more complex and token-hungry. Their conclusion: "full automation will be prohibitively expensive for most organizations". I don't fully agree with the forecast, but I agree with the implication. Build the agent to do a narrow set of things at a low cost per call, not to be a general-purpose employee at any cost.

The 91 percent, and why it's a warning

Gartner surveyed 321 customer service leaders in October 2025 and found 91% were under pressure from executives to implement AI this year. By August, AI spending in those functions was up 38% against total budgets up 2%. That is a lot of money moving on the basis of pressure rather than a measured problem, and it is exactly how you end up with an agent that answers every call, resolves none, and is quietly switched off in month four.

The projects that have worked for us started from the phone system report, not from the board. They picked the narrow band of calls that are (a) high-volume, (b) fully answerable from systems the agent can read, and (c) not the moment the relationship is decided.

If you want the wider adoption picture, GoodFirms maintains a roundup of AI adoption statistics compiled from published surveys. It is useful background, and it is also the trap this section is about: industry-wide adoption rates tell you what everyone else is doing, not whether your own phone is ringing at a volume that justifies any of it.

What to let it handle, and what never to

We use the same five-stage framework for phone that we use for every other intake channel — Capture, Understand, Act, Escalate, Report — and the phone version looks like this.

Let it handle

  • Booking, rescheduling and cancelling against the live calendar, with the agent reading real availability rather than guessing. This alone is most of the 12:40 problem for clinics.
  • "Where's my order / where's my driver / what's my claim status" — a lookup with a number the caller already has.
  • Hours, location, parking, insurance accepted, what to bring — the questions that make up a third of a front desk's day and none of its value.
  • After-hours and overflow intake: take the name, number, reason for calling and urgency, create the record, and tell the caller exactly when they'll hear back. This is the highest-return use, because it converts calls that currently vanish.
  • Outbound reminders with one-word replies — confirm, reschedule, cancel — where the caller has consented to be contacted.

Never let it handle

  • Anything that sounds like a symptom, a chest pain, a medication question. Route to a human on the first sentence, not after a triage attempt.
  • A complaint. The AI-handled complaint is the one that ends up on a review site.
  • Pricing negotiation, discounts, refunds beyond a fixed policy.
  • A caller who has asked for a person. Once. Not after "I understand you'd like to speak to someone, but first…".
  • Anything with legal consequence in the insurance world — coverage confirmations, claim decisions, a statement about what a policy does or doesn't cover.

Escalate means a warm transfer with the transcript, the caller's record and a one-line summary already on the human's screen. If the person picking up has to say "sorry, can you start again?", you've built a slower voicemail.

Compliance, in the order it will bite you

Tell people it's an AI. The EU AI Act's transparency duties under Article 50 apply from August 2026: anyone interacting with an AI system must be told, unless it's obvious. Several US states have bot-disclosure laws of their own. Practically, the agent's first sentence says what it is. Callers don't mind; they mind being fooled.

Outbound is regulated; inbound mostly isn't. In February 2024 the FCC ruled that AI-generated voices count as "artificial" under the TCPA, which means an outbound call using one needs the caller's prior express consent, and the penalties per call are real. A caller phoning you is not a robocall. A reminder you place to someone who ticked a box is fine. A cold call from a synthetic voice is a lawsuit. We cover the current state of TCPA on our insurance intake page, because that's where it matters most.

Recording consent. Around a dozen US states, and most of Europe, require all parties to consent to a recording. The agent announces it; the transcript still exists, so treat it as a recording either way.

Healthcare. If a caller can say anything about their health — and they will — every vendor in the chain that touches the audio or the transcript needs a signed Business Associate Agreement: carrier, transcription, model provider, storage. Not all of them offer one. Check before you build, not after. Our healthcare page goes through the BAA question in detail.

How we build one

For readers who want the architecture rather than the pitch, this is the shape that has held up.

  1. Telephony via a programmable carrier (Twilio or a SIP trunk into a media server), streaming audio both ways over a WebSocket. Keep your existing number; forward on no-answer after two rings, or overflow, or out of hours.
  2. Understanding — either a streaming STT feeding a text model, or a speech-to-speech model. We default to cascaded for anything that touches money or health, because the text in the middle is where validation lives.
  3. Tools, not prompts, for facts. The agent never states an appointment time, an order status or a price it didn't just get from a function call against the system of record. The model chooses which tool; the tool returns the truth.
  4. Confirmation before commitment. Read back the slot, the name, the number. Then write. This adds two seconds and removes the worst class of failure.
  5. Escalation with context: warm transfer where staff are available, structured callback record where they aren't, with urgency scored by rules rather than vibes.
  6. Reporting that you'll actually read: calls handled, calls escalated, why, bookings made, callers who asked for a human, and — the one everyone forgets — calls where the agent said something that later turned out to be wrong. Sample transcripts weekly. Sample the escalations, not the successes.

Buy the platform or build it?

Buy if your calls are generic, volume is under a few hundred a month, and you have no systems to integrate — a bundled platform at ten to eighteen cents a minute is the cheapest way to find out whether callers accept it. Build (or have built) when the calls are worth money, the agent has to read and write your calendar, CRM or claims system, compliance requires you to know exactly where the audio goes, or the per-minute margin at your volume costs more than the engineering did. The crossover in our experience sits around two to three thousand call minutes a month, and it moves down every time a token price does.

If you want to see what the numbers look like for your own volume, the patient no-show calculator and the dispatcher time calculator both run on the model above; put your call counts in and it'll show you the breakeven rather than a promise.

Answer the calls that go nowhere. Do four things well. Hand over early, with the context attached. Everything beyond that is the part Gartner's warning about.

Frequently asked questions

What can an AI phone agent reliably handle for a small business?

The narrow, high-volume calls that can be fully answered from systems the agent can read: booking, rescheduling and cancelling against a live calendar; order, driver or claim status lookups; hours, location, insurance accepted and what to bring; after-hours and overflow intake that captures name, number, reason and urgency and promises a callback time; and consented reminders with one-word replies. It should not handle anything that sounds like a symptom, a complaint, price negotiation, coverage or claim decisions, or a caller who has asked once for a person. Those go to a human immediately with the transcript attached.

How much does an AI voice agent cost per minute?

The underlying models bill per token, and OpenAI meters audio at 600 tokens per minute of caller speech and 1,200 per minute of agent speech. On gpt-realtime-2.1 at 32 and 64 dollars per million tokens that is about 5 cents per blended minute, and about 1.5 cents on the mini model. A cascaded build using streaming transcription at around half a cent a minute, a small text model and text-to-speech lands at 4 to 7 cents including carrier charges. Practitioners measuring real sessions report 6 to 11 cents on the flagship model with prompt caching working. Bundled platforms that include telephony charge roughly 10 to 18 cents a minute, and bring-your-own-keys platforms end up at 13 to 31 cents once the model, voice, transcription and carrier are added.

Why do AI voice agents feel slow, and can it be fixed?

Human conversation turns over with a gap of 0 to 200 milliseconds in every language studied, according to Stivers and colleagues in PNAS, with a cross-linguistic median of 100 milliseconds. A cascaded agent has to detect that the caller has finished, finalise the transcription, get a first token from the language model, synthesise speech and send it through the carrier, which adds up to roughly 0.9 to 2 seconds. Speech-to-speech models cut that to around 500 to 800 milliseconds by removing two hops, at the cost of the text in between where validation normally happens. No system reaches 200 milliseconds; good ones design around it with streamed acknowledgements while lookups run and endpointing tuned per call type.

Are AI voice agents legal for outbound calls in the United States?

Only with consent. In February 2024 the FCC ruled that AI-generated voices are artificial voices under the Telephone Consumer Protection Act, so any outbound call using one requires the recipient's prior express consent, with statutory damages per call. Inbound calls that the customer places to you are not robocalls. Reminders to people who have opted in are fine; cold calls with a synthetic voice are not. Separately, the EU AI Act's Article 50 transparency duties apply from August 2026 and several US states have bot-disclosure laws, so the agent should say what it is in its first sentence, and around a dozen US states plus most of Europe require all-party consent for recording.

Should a business buy a voice agent platform or have one built?

Buy when calls are generic, volume is under a few hundred a month and there is nothing to integrate; a bundled platform at 10 to 18 cents a minute is the cheapest way to learn whether callers accept it. Build when the calls are worth money, the agent must read and write a calendar, CRM or claims system, compliance requires knowing exactly where the audio goes, or the per-minute platform margin at your volume exceeds the engineering cost. In our experience the crossover sits around two to three thousand call minutes a month and moves lower each time token prices fall.

#ai#voice-agents#ai-receptionist#automation#customer-support