AI Calling Agents
How AI calling agents work
An AI calling agent sits on top of a phone connection, and that connection is the part that decides whether the rest works. The telephony layer is what puts the agent on the call: either a SIP trunk from a carrier with its own numbers, or the calling platform the company already runs, in which case the agent inherits the existing numbering plan and directory. Above that layer, every turn of the conversation runs the same loop: speech to text, a model that decides the next action and can call an external system, then text to speech back onto the line.
Two properties of that stack are worth understanding before sizing a deployment.
Concurrency is the number of calls the agent can hold at the same time. An IVR is sized by ports and each port does almost no work. An AI calling agent runs a full transcription, model and speech pipeline per call, so twenty simultaneous callers means twenty pipelines. Concurrency is what breaks first on a peak, and it breaks silently: the twenty-first caller hears whatever the fallback does, not an error.
Cost per minute is metered, not fixed. Speech recognition, the model, speech synthesis and the telephony minutes are each billed by use, so a four-minute call in which the agent looks up three records costs more than a forty-second call that answers opening hours. This is the opposite of the IVR cost model, where the cost is the licence and a longer call is free.
Who starts the call: inbound and outbound
AI calling agents split into two groups, and the dividing question is who originates the call.
Inbound AI calling agents answer a number that already exists. The work is not the conversation, it is everything around the number: which calls reach the agent and which bypass it, what the agent may resolve alone, and where the call lands when it cannot. An inbound deployment is a telephony project with a conversation on top.
- Covering the main line outside office hours and during peaks
- Answering the questions that repeat, in parallel, with no hold
- Booking and rescheduling against a live calendar
- Taking a message and passing it on already structured
Outbound AI calling agents originate the call, and that adds three requirements an inbound agent never has. A reason to call each person on the list. A consent record proving that this person agreed to this kind of contact, retrievable when someone complains. And a disclosure at the start of the call stating that the voice is automated, because several jurisdictions apply their rules on automated calls to synthetic voices as well.
- Appointment and payment reminders
- Calling back a web lead while the form is still fresh
- Confirming or rescheduling a delivery slot
- Reactivating dormant customers
The two groups also measure differently. An inbound agent is judged on what it resolved. An outbound agent is judged first on connect rate, the share of calls that reach a person at all, because a perfect conversation with a voicemail box is a cost with no outcome.
How do AI calling agents differ from traditional IVR?
A traditional IVR and an AI calling agent are two different machines with two different cost models, and only one of them can dial.
An IVR is a routing appliance. It answers on a fixed number of ports, plays a recorded prompt, collects a keypress, and hands the call to a destination. Its cost is a licence, its capacity is engineered once, and it has no notion of placing a call: an IVR can only react to a call that already exists.
An AI calling agent is a conversational resource billed by use. It scales with concurrency rather than ports, costs more the longer it talks, and can originate a call as easily as answer one.
The practical consequence shows up on a missed call. When forty calls arrive at once, an IVR answers all forty and parks thirty-five of them, and the five who hang up are gone: there is no mechanism in an IVR to reach them again. An AI calling agent can hold as many conversations as it has capacity for, and for the ones it could not take it can place the callback itself, from the same flow, without anyone building a list. Recovering the abandoned call is not a better version of what an IVR does. It is something an IVR structurally cannot do.
What to check before you deploy one
Five questions about running the thing, not demoing it.
How many calls can the agent hold at the same time, and what does caller number forty-one hear?
A real answer names a concurrency limit and the behaviour past it. If the answer is “it scales automatically”, ask what it scaled to during the last peak and who was told.
Price me a four-minute call in which the agent checks two records.
A real answer breaks the number down by component and names what is metered. An answer expressed only per user or per month means the usage cost sits somewhere else, and it will surface after go-live.
What number does the person see when the agent calls out, and what happens when they call it back?
Outbound from a number nobody answers turns every successful call into an inbound problem. The return path has to land somewhere on purpose.
When the AI platform is unreachable, what does the caller hear?
This is a telephony question, not an intent question. A real answer names what takes over on the line: a queue, a night service, a human. Uptime figures are not an answer.
For an outbound campaign, show me the consent record for one specific contact.
It has to be per person and per purpose, and it has to be retrievable in seconds. A clause in a contract or a tick in a CRM field with no timestamp is not evidence.