Introduction
Phone calls still drive revenue and service in healthcare, real estate, home services, legal, financial services, and most appointment-based businesses. An AI voice agent can answer those calls in natural conversation, take action in your business systems, and hand complex cases to your team.
The decision most companies face in 2026 is not whether to use one, but how to get one: build it in-house or buy a production-ready platform. This guide compares both paths on ownership, cost, latency, integrations, and compliance so you can choose with confidence.
The short answer
Buy if your goal is a business outcome: answering more calls, booking appointments, qualifying leads, recovering missed calls, or running outbound follow-ups. A production-ready platform gets you there faster, with less risk.
Build if voice AI is your product, or a core competitive advantage, and you can staff real-time AI, telephony, and reliability engineering for the long haul.
Most US businesses fall into the first group. The rest of this guide helps you confirm which one you're in.
What you'd actually own: the eight layers of a voice agent
An AI voice agent is software that holds a spoken conversation and takes business actions during or after it. Under the hood, that means eight layers working in real time. The table shows who carries each one.
Layer | What it does | If you build | If you buy |
|---|---|---|---|
Telephony | Connects inbound and outbound calls, transfers, numbers | Your team | Usually included |
Speech-to-text | Turns the caller's words into text as they speak | Your team selects, tunes, and monitors | Platform-managed |
Turn detection | Decides when the caller has finished talking | Your team | Platform-managed |
Reasoning (LLM) | Understands intent and decides what to say or do | Your team, including prompts and fallbacks | Configured by you, run by the vendor |
Guardrails and logic | Limits what the agent may say and do | Your team designs and enforces | Built-in controls you configure |
Integrations | Reads and writes CRMs, calendars, helpdesks | Custom code for each system | Connectors plus APIs |
Text-to-speech | Speaks the response in a natural voice | Your team | Platform-managed |
Analytics and observability | Transcripts, outcomes, failure tracing | Your team builds dashboards and alerts | Usually included |
Building gives you maximum control over every row. It also makes every row your pager's problem at 2 a.m. Buying doesn't remove your responsibility for how the agent is used, especially for security and compliance, but it moves most of the infrastructure burden to the vendor.
When building earns its cost
There are good reasons to build. They just tend to apply to a narrow set of companies. Use this as a checklist; if you can't tick most of it, building will likely cost more than it returns.
Voice AI is the product, or a capability customers choose you for.
You already employ real-time AI and telephony engineers, not just developers who could learn.
Your workflows genuinely don't fit configurable platforms, even with APIs and custom tools.
You need infrastructure-level control, such as specific models, private networking, or unusual data residency.
Your volume is very large and predictable, enough that owning parts of the stack changes the economics.
You can fund continuous QA, observability, and security work, indefinitely, not just for launch.
Notice what's missing from the list: wanting an AI receptionist, a support agent, an appointment booker, or an outbound caller. Those are business goals, and they rarely justify owning the stack.
Total cost of ownership, on both sides
Building: the costs that never make the prototype budget
The API bill is the smallest line. The real cost sits in the failure cases your team has to design for, one at a time:
The CRM is down mid-call.
The model times out, or text-to-speech fails.
The caller interrupts, or the agent mishears a name.
The calendar returns conflicting availability.
An outbound call reaches voicemail.
Two systems return contradictory data.
Each needs handling, testing, and monitoring. Then comes ongoing QA: new accents, new customer behaviors, knowledge updates, and model changes can all break a flow that worked last month. Add security and compliance work (recording consent, retention, access control, vendor agreements), and in healthcare, a Business Associate Agreement with any cloud vendor that touches protected health information.
The largest cost is usually invisible: every engineer tuning turn detection is an engineer not working on your actual product.
Buying: read past the headline price
Platforms have costs too, and they vary in structure more than in size. Before comparing vendors, list every line that applies at your volume: platform fees, included minutes, overage rates, telephony, phone numbers, connector or integration charges, concurrency limits, implementation, premium support, and compliance add-ons. Two vendors with similar per-minute rates can differ widely once these are included.
A simple ROI model you can actually fill in
Skip the vendor's savings claim and build your own number. Three inputs are enough:
Recovered revenue = additional qualified bookings or sales × average contribution margin
Released capacity = hours of repetitive phone work avoided × loaded hourly cost
Total cost = platform (or build) cost + implementation + ongoing management time
Net annual benefit = recovered revenue + released capacity − total cost
Take a home-services company receiving 4,000 calls a month. If an agent covers after-hours calls, routine scheduling, FAQs, and first-pass qualification, you can pull every input from data the company already has: how many after-hours calls go to voicemail today, what share of those would have booked, what a booked job is worth, and how many staff hours go to rescheduling. The result is specific to that business, which is exactly what makes it credible to a finance team.
Notice that value shows up on both sides: money that was being lost, and time that was being spent. Most ROI cases that stall only count one.
Latency: measure what the caller hears
On a phone call, silence feels longer than it is. After the caller stops speaking, the system has to hear, transcribe, understand, decide, generate, and speak, and every step adds delay.
Twilio has published starting benchmarks for cascaded voice AI systems (speech-to-text, then LLM, then text-to-speech):
Component | Target | Upper starting benchmark |
|---|---|---|
Speech-to-text | ~350 ms | ~500 ms |
LLM time to first token | ~375 ms | ~750 ms |
TTS time to first audio | ~100 ms | ~250 ms |
Platform turn gap | ~885 ms | ~1,100 ms |
Mouth-to-ear turn gap | ~1,115 ms | ~1,400 ms |
Twilio frames these as starting points, not best-case performance. Deepgram, for its part, has shown a controlled deployment with a median end-to-end latency of about 660 ms and a 90th percentile under one second.
Those numbers are useful context, but a vendor's headline figure tells you little on its own. Ask instead:
Is this server-side latency, or what callers actually hear?
Does it include endpoint detection and telephony?
Is it the median, P90, or P95?
What happens during a tool call, such as a calendar lookup?
How does it hold up at peak traffic, and for callers far from the servers?
Then test it yourself with real calls from the regions you serve. LuMay publishes how it approaches this on its voice agent latency page.
Voice agent or traditional IVR?
They solve different problems, and plenty of businesses will run both.
Traditional IVR | AI voice agent | |
|---|---|---|
How callers interact | Keypad or fixed phrases | Natural speech |
How calls move | Menu tree | By intent |
Unexpected questions | Handled poorly | Handled within configured boundaries |
Business actions | Preconfigured | Tool calls to live systems |
Cost profile | Mature and predictable | Depends on usage and platform |
Best at | Simple, deterministic routing | Conversations that end in an action |
If a call always follows the same three presses, an IVR is fine. Once callers describe what they need in their own words, or the call should end with a booking or a record update, a voice agent earns its place.
How to evaluate an AI voice agent: a working scorecard
Voice quality is the first thing people notice in a demo and the least likely reason a deployment fails. Score platforms on these instead:
Area | What "good" looks like |
|---|---|
Conversation | Natural turn-taking, clean interruption handling, low latency as heard by callers |
Coverage | Inbound and outbound calling from one platform |
Actions | CRM, calendar, helpdesk, and knowledge-base access during the call |
Escalation | Transfers to a person with the context already passed along |
Control | Guardrails, permission limits, and configurable call flows |
Visibility | Transcripts, outcomes, and analytics tied to business results |
Extensibility | APIs and webhooks for anything without a connector |
Regulated use | Access controls, encryption, retention settings, audit logs, and BAAs or DPAs where required |
Integration rules that separate a useful agent from a talking one
An agent that can act is worth far more than one that can only chat. A few rules keep that power safe:
Least privilege. Give the agent only the CRM and system access each workflow needs.
Check before you offer. For bookings, the sequence is check availability, offer, confirm, book, then send confirmation.
Store outcomes, not just transcripts. Write structured fields: reason, result, qualification status, transfer outcome, appointment status, campaign source.
Own the knowledge. Assign an owner and review cadence to every document the agent reads, or an outdated page becomes an outdated spoken answer.
Map calls to pipeline steps. For example: qualified, opportunity created, owner assigned, follow-up task set.
Tighten rules around money and sensitive data. Payments, account changes, and health or financial details need stronger authentication.
Measure outcomes. "Calls handled" is activity. Bookings, resolved requests, and qualified leads are results.
You can see how LuMay handles these areas, including post-call summaries and its analytics dashboard, on the voice agent features page.
Where voice agents pay off first
The strongest early deployments share a pattern: the call is frequent, the conversation is predictable, and the outcome is a clear action. The scenarios below are illustrative, not LuMay customer case studies.
HVAC company, calls missed during installs. The agent answers every call, separates emergencies from routine service, checks the service area, books available slots, and sends urgent jobs straight to dispatch. Technicians stop taking calls from rooftops.
Dental practice, a front desk buried in rescheduling. The agent checks the calendar, confirms a new time, books it, and texts a confirmation. Staff get their attention back for patients in the chair.
SaaS support team, the same "how do I" questions daily. The agent answers approved Tier-1 questions from the knowledge base and opens a ticket with context when it can't. Specialists spend their time on harder problems.
Real estate team, leads arriving after hours. The agent asks about location, budget, timing, and financing, then writes structured answers into the pipeline. Agents start the morning with qualified leads instead of voicemails.
The same pattern fits local service businesses, legal intake, missed-call recovery, and approved outbound reminders or follow-ups. Healthcare front desks qualify too, as long as the implementation meets privacy, security, and consent requirements. For how inbound calls flow through an agent end to end, see LuMay's inbound voice agent overview.
The platform landscape, and where LuMay fits
If you decide to buy, the next choice is what kind of platform. They differ less in what they can do than in who they're built for and how they charge.
Platform | Built for | Published pricing |
|---|---|---|
LuMay Voice Agent | Businesses wanting voice automation tied to workflows, integrations, and configurable call flows | Basic $249/mo, Starter $499/mo, Growth $999/mo; Enterprise by quote |
Retell AI | Development teams wanting flexible voice infrastructure | $0.07 to $0.31/min pay as you go, 20 concurrent calls included (pricing) |
Bland AI | Teams wanting API-driven voice with bundled AI costs | $0.14/min (Start); $0.12/min plus $299/mo (Build); telephony may be separate |
Synthflow | Larger deployments needing implementation and enterprise support | Enterprise pricing from $30,000/year |
ElevenLabs Agents | Teams prioritizing voice quality and conversational experience | Usage-based agent calls, with LLM costs passed through |
No single headline number decides this. A per-minute rate can look cheaper than a monthly plan until you add telephony, LLM pass-through, and connector costs. Model each option at your expected volume and integration pattern.
Where LuMay is a good fit
LuMay Voice Agent is built for the "buy" side of this decision: companies that want to automate phone work without assembling the stack. It combines inbound and outbound calling, configurable conversation flows, integrations and APIs, guardrails around agent behavior, human handoff, and analytics in one platform. Its package pricing (tiers vary by included minutes, agents, connectors, analytics, and API access) suits teams that prefer a predictable monthly budget over component billing. Older LuMay articles quoted an approximate per-minute figure; use the current pricing page when modeling costs.
That doesn't make it the right choice for everyone. A developer team that wants to swap every model may prefer an infrastructure-first platform. Whichever you shortlist, pilot it on your own calls, integrations, accents, and compliance requirements before committing.
Outbound calling in the US: compliance comes first
A platform that can place calls doesn't make any particular campaign legal. The business running the campaign stays responsible for how it calls people, whether it built the agent or bought it.
The FTC's Telemarketing Sales Rule restricts prerecorded telemarketing calls and sets consent and opt-out requirements in applicable situations. Before launching any AI outbound calling, review with counsel:
TCPA requirements and consent records
Telemarketing Sales Rule obligations
Federal and state Do Not Call rules
Call-recording consent laws, which vary by state
Disclosure, identification, and opt-out requirements
Industry-specific privacy rules, such as HIPAA for healthcare
Inbound automation carries lighter obligations, but recording, retention, and data access still need clear policies.
FAQs
Q1: What is the best AI voice agent for a small business in 2026?
A: The one that fits your call volume, systems, and budget. Small businesses should prioritize fast setup, appointment booking, CRM connections, human transfer, predictable pricing, and responsive support over headline latency or price claims.
Q2: How do I build an AI voice agent for customer service?
A: You'd need telephony, streaming speech-to-text, turn detection, an LLM, text-to-speech, business logic, knowledge retrieval, integrations, guardrails, analytics, and escalation handling. You can assemble these yourself or use a platform that provides most of them.
Q3: How much does an AI voice agent cost?
A: Pricing mixes platform fees, per-minute usage, telephony, connectors, concurrency, and implementation. Compare total cost at your expected volume rather than a single rate.
Q4: What is good latency for an AI voice agent?
A: There's no universal threshold, but callers notice long pauses quickly. Twilio suggests starting around 1.1 seconds mouth-to-ear for cascaded systems, and optimized deployments can go below one second. Test with real calls.
Q5: Can an AI voice agent make outbound sales calls legally?
A: Technically it can place the calls. Whether a campaign is legal depends on consent, Do Not Call rules, disclosures, opt-outs, and state law, so review those before launch.
Q6: Can an AI voice agent integrate with a CRM?
A: Yes. Through connectors, APIs, or webhooks, it can pull customer context during a call and write back outcomes, appointments, tickets, and follow-up tasks.
Your next move: a 30-day pilot
Don't decide build versus buy in a meeting. Decide it with evidence from a small, controlled test.
Week 1: pick one call type. Choose a frequent, predictable conversation that ends in an action, such as rescheduling or after-hours intake. Record today's cost, volume, and what's lost when it goes unanswered.
Week 2: connect the essentials. Link only the systems that call needs, with least-privilege access, and define exactly when the agent hands off to a person.
Week 3: go live on a slice of traffic. Route a share of real calls, such as after-hours only, and keep your current process as the fallback.
Week 4: measure and decide. Compare latency as callers hear it, containment and transfer rates, bookings or qualified leads, and customer outcomes against your baseline. Scale what works.
If the pilot shows you need infrastructure no platform can provide, you'll have a precise spec for building. If it doesn't, you'll have results in weeks instead of quarters.
Explore LuMay Voice Agent or book a demo and bring the call type you'd pilot first. We'll walk through how it would run on your systems.





