Back to blog

48 AI Phone Agent Pricing Models: Full Cost Breakdown 2026

AI phone agent pricing hides three cost layers enterprise buyers never see. Avoid budget surprises with our full 2026 vendor breakdown.

Updated October 5, 202637 min read

Your vendor's per-minute rate is one layer of four. Here is what the other three layers cost, why they never appear on the pricing page, and how to read an invoice before you sign one.

Most enterprise buyers in regulated industries think: "If I collect per-minute rates from each vendor's pricing page, I'll know my real monthly cost before signing." It's a reasonable assumption. It's also structurally wrong, and the gap between that estimate and the actual invoice is where real budget surprises live. The architecture underneath most AI voice agents makes accurate pricing-page estimates nearly impossible, not because vendors are being deceptive, but because the pricing page only describes one layer of a four-layer cost stack.

Most AI phone agent platforms are assembled, not built. A typical deployment stitches together a telephony relay, a third-party large language model for reasoning, a separate speech-to-text provider for transcription, and a text-to-speech provider for voice output. Each vendor in that chain meters usage independently and bills the platform, which either absorbs those costs or passes them through. The pricing page reflects what the platform charges you for the telephony layer. The other three vendors rarely appear on it.

Image: Four-layer AI voice agent cost stack showing where pricing pages stop and hidden costs begin

This is an architecture problem, not a transparency problem. A platform that routes calls through external inference APIs cannot publish a single honest per-minute rate, because that rate changes with every model update, every token-pricing revision, and every call's unique token mix. The telephony relay is the only layer most pricing pages quote.

A single GPT-4o Realtime API session meters audio input tokens, audio output tokens, text input tokens, and text output tokens simultaneously at distinct per-token rates. 22/min or higher on a three-minute support call with a verbose response. STT providers bill per second of audio processed; TTS providers bill per character synthesized.

Key takeaways#

  • Most AI phone agent pricing pages quote a base per-minute rate that excludes the token charges, third-party STT/TTS fees, and latency penalties that compound on top of it, enterprise buyers consistently report invoices running 40-60% above their initial projection.
  • The real cost of an AI phone agent emerges from three compounding layers: the platform's base rate, the pass-through fees from every outsourced stack component, and the per-call overhead from network latency, none of which appear on the same pricing page.
  • Vendors that outsource speech recognition, voice synthesis, or LLM inference to third parties cannot quote a stable all-in rate, because their costs shift with upstream provider pricing and usage spikes they don't control.
  • Billing during hold time, silence, and IVR prompts varies by vendor and is almost never disclosed upfront, it's one of the most common sources of invoice shock for high-volume call operations.
  • Collecting per-minute rates from 48 pricing pages and dividing by call volume produces a number that is structurally wrong before the first call is placed, the formula ignores concurrency ceilings, compliance tier surcharges, and transfer fees.
  • Owning the full inference stack end-to-end is the only architectural condition under which a vendor's quoted rate can reliably match the invoice, every handoff point between outsourced components is both a latency source and an uncontrolled cost variable.
  • Bland's fine-tuned voice models run on dedicated infrastructure with co-located GPUs, and its Voice Delivery Network routes every call to the nearest server, that stack design is what eliminates the hidden pass-through fees and the processing latency that inflate per-minute costs on platforms that don't own their own infrastructure.

The 3 AI Phone Agent Pricing Models - Per-Minute Cost Breakdown for Each Deployment Type#

Per-minute rates are the first thing you see on an AI phone agent pricing page, and in most cases the last thing you should rely on when forecasting your actual monthly invoice.

Bland AI's noise cancellation can be configured at three levels: the agent (Persona) level, per individual call, or per inbound phone number.

The reason comes down to stack architecture. Three distinct deployment models exist in this market, and they produce structurally different invoice totals for identical call volumes, because each model bills a different set of metered units. Understanding which model a vendor uses tells you more about your real cost than the headline rate ever will. For operations teams whose goal is to reduce cost-per-contact and improve customer service ROI by replacing or augmenting human agents with AI phone agents, choosing the wrong model is where budget overruns begin.

1. Infrastructure/DIY API Stack - $0.05-$0.15/min Advertised, $0.20-$0.35/min Effective#

AI Phone Agent Pricing - infrastructure diy api stack

Consumption-based pricing on infrastructure platforms like Twilio covers telephony only. Buyers who assemble a DIY stack must sum every component independently. STT, TTS, and agent orchestration each bill separately across most infrastructure providers, with no single bundled rate surfaced on the pricing page, though exceptions exist, as Deepgram explicitly offers a single, unified Voice Agent API that bundles STT, TTS, and LLM orchestration together.

Add LLM token charges, which scale with conversation length rather than call minutes, and a stack showing $0.08/min in base costs routinely lands at $0.20-$0.35/min when the invoice arrives. This model suits engineering teams with the capacity to own infrastructure, monitor vendor rate changes, and absorb billing complexity. It is not a realistic path for ops teams without dedicated AI engineering support, and it works against the goal of cutting call center costs by 50%+ when hidden charges erode every efficiency gain.

2. Managed All-in-One SaaS Platform - $0.25-$1.00/min Bundled, Consumption or Subscription Billing#

AI Phone Agent Pricing - managed all in one

Managed SaaS platforms bundle STT, TTS, and LLM inference into a single per-minute rate, removing most of the arithmetic burden. The tradeoff is that "bundled" does not always mean "all-in." Some platforms pass through upstream LLM token costs at variable rates, so the bundled rate is a floor with no guaranteed ceiling.

Platforms such as Retell AI and Vapi offer bundled per-minute pricing that includes real-time transcription, premium voices and clones, and LLM inference with no separate token charges. What you see on the pricing page is what appears on the invoice. For teams running continuous outbound campaigns, sales, follow-ups, reminders, or 24/7 inbound call handling without scaling headcount, that billing certainty matters more than the headline rate alone.

Both plans include:

  • Conversational pathways, automations, and integrations
  • Version lock
  • Up to 15 knowledge bases (Build: 50, Scale: 100)
  • Voice clones (Build: 5, Scale: 15)

These capabilities deflect repetitive inquiries to AI voice agents without requiring a separate vendor for each layer.

For regulated industries or high-volume organizations where compliance documentation, dedicated infrastructure, and enterprise controls are non-negotiable, bland.ai's Enterprise plan adds a Business Associate Agreement (BAA), SSO, data residency, on-prem/VPC deployment, JWT signatures, a dedicated orchestration server, and unlimited concurrent calls sized to volume. Compliance documentation is available under NDA.

For operations teams already running on Amazon Connect, bland.ai's Amazon Connect Integration allows AI voice agents to substitute for or augment human agents within existing inbound and outbound call flows, without migrating to a new platform.

Platforms built on self-hosted infrastructure that own their inference layer can quote a rate where the advertised number and the invoiced number are the same figure. That architecture makes cutting cost-per-contact a predictable outcome rather than a projection.

3. Custom Agency Build - $1,500-$6,000+ Upfront + $300-$1,500/Month Retainer#

Custom agency builds price on scope rather than per-minute consumption. Simple implementations run $1,500-$2,000 upfront plus $300-$500/month. Complex builds carry higher retainers and longer delivery timelines, costs that compound before a single call is made. For teams that need 24/7 phone coverage most beneficial at high call volumes, an agency build delays the point at which AI agents begin deflecting repetitive inquiries and reducing cost-per-contact. That delay has a real dollar value that rarely appears in the agency's proposal.

48 AI Phone Agent Pricing Models, Plans, and Cost Structures - Every Model Ranked and Explained#

AI Phone Agent Pricing Models, Plans, and Cost Structures#

Three invoices sitting next to three pricing pages rarely match. That gap is not a billing error. It is the predictable result of a cost structure that no single pricing page fully discloses, because the real price of an AI phone agent emerges from three compounding multiplier layers that interact with your specific call behavior.

The base rate on the vendor's page is only the first layer. The second layer is how the vendor defines a billable minute. The third is what happens when calls fail, transfer, or spike past your concurrent limit.

Each layer inflates the invoice independently, and they compound. A call that triggers all three simultaneously can cost three to four times the headline rate. Enterprise buyers who audit only the base rate before signing are evaluating roughly one-third of their actual cost structure.

"AI phone agents that call repeatedly to verify reservations clog up business phone lines, blocking real customers from getting through, a direct operational cost concern for businesses evaluating AI phone agent plans."

— what we hear from restaurant and hospitality businesses

Enterprise buyers who audit only the base rate before signing are evaluating roughly one-third of their actual cost structure.

Bland Speech is free to start, with new accounts receiving 133,000 characters, roughly two hours of speech, before pay-as-you-go pricing applies.

As of May 2026, AI voice agent cost structures involve multiple overlapping billing dimensions, including per-minute consumption rates, platform fees, concurrent call limits, and feature-gate add-ons, meaning that dividing a per-minute rate by call volume alone does not capture the true monthly cost. The variables do not add. They multiply.

The 44 models, structures, and billing mechanics below map every lever that determines what you actually pay. Read them as a checklist before you sign any contract.

1. Bland.ai Start Plan - $0.14/min, $0 Platform Fee, 10 Concurrent Calls#

The Start plan carries no platform fee and requires no card to begin, making it the lowest-friction entry point for developers building and testing AI phone agents before committing to volume. At $0.14/min, STT, LLM inference, TTS, and real-time transcription are all included in that single rate. The honest limitation: 10 concurrent calls and a 100-call daily cap make it unsuitable for any production workload above light pilot scale.

2. Starter/Basic SaaS Tier - Limited Minutes, Basic Call Routing, Automated SMS Follow-Ups ($50-$150/Month)#

SaaS Starter tiers from platforms like Aircall, JustCall, and CloudTalk typically land between $50 and $150 per month and include limited bundled minutes, basic inbound call routing, and automated SMS follow-ups. Bundles at this tier often cost $30 to $200 per seat per month. The tradeoff is real: minute caps are low enough that a single busy day can push you into overage territory, and CRM sync is usually gated behind the next tier up.

3. Professional/Business SaaS Tier - Custom Voice Cloning, Multi-Agent Workflows, Advanced Integrations ($300-$600/Month)#

Professional tiers step up to $300-$600 per month and unlock custom voice cloning, multi-agent workflow orchestration, and native integrations with Salesforce, HubSpot, Pipedrive, and Zendesk. The limitation buyers miss: "advanced integrations" often means API pass-through rather than native bidirectional sync, so CRM data quality depends on your engineering team's maintenance bandwidth.

4. Enterprise SaaS Tier - HIPAA Compliance, Custom Database Lookups, Dedicated Infrastructure ($1,000-$3,000+/Month)#

Enterprise SaaS tiers reach $1,000 to $3,000 or more per month and add HIPAA-compliant call handling, custom database lookups mid-call, dedicated infrastructure, and audit log access. The cost catch: compliance features at this tier are almost never disclosed on the public pricing page. They surface in the sales call, after you have already built your cost model around the published number.

5. Pure Pay-As-You-Go Per-Minute Model - No Platform Fee, Billed Only on Talk Time#

Pure PAYG models charge only for minutes consumed, with no monthly floor. Per-minute rates across provider types range from $0.05 to $1.00 per minute depending on features and volume commitments. PAYG is the right structure for unpredictable or seasonal call volumes where a fixed monthly commitment would create waste. The risk: at high volume, the per-minute rate is almost always higher than what a subscription tier would cost per effective minute, and there is no rate lock.

6. Subscription-First Model with Included Minute Bundles - Predictable Monthly Cost Floor#

Subscription-first models charge a fixed monthly fee that includes a set minute bundle, with per-minute overages kicking in above the threshold. The trap is the bundle sizing: vendors set minute allocations based on average usage, and contact centers with seasonal spikes will routinely exceed their bundle in peak months, triggering overage rates that are often 30 to 50 percent higher than the base per-minute equivalent.

7. Hybrid Billing Structure - Monthly Platform Fee Plus Per-Minute Consumption Above a Base Allowance#

Hybrid structures combine a monthly platform fee with a per-minute consumption rate. Bland.ai's Build plan ($299/month platform fee, $0.12/min) is a clean example. The advantage for mid-volume buyers is that the platform fee buys down the per-minute rate compared to pure PAYG. The decision variable is whether your monthly minute volume is high enough to amortize the platform fee into a lower effective rate than the no-fee alternative.

8. Token Surcharge Model - LLM Token Costs Billed Separately on Top of Per-Minute Rate#

Some platforms quote a low base per-minute rate and then add LLM token charges on top, billed per 1,000 output tokens. Token surcharges are one of the hidden variables that do not appear in a simple per-minute rate comparison. A 4-minute call generating 800 output tokens on a frontier model can add $0.24 to $1.20 in token fees on top of the $0.32 base charge. At 10,000 minutes per month, that delta is not rounding error.

9. All-In Per-Minute Model - Single Rate Covering STT, LLM, TTS, and Telephony#

All-in per-minute pricing bundles speech-to-text, LLM inference, text-to-speech, and telephony into one rate, so the number on the pricing page equals the number on the invoice. The condition under which it delivers full value: the vendor must own the full stack. A platform that calls itself "all-in" but routes inference through a third-party API can still pass through upstream price changes without warning.

10. Transfer-Minute Billing - Billing Continues Through Human Handoff and Hold Time#

Transfer-minute billing is a hidden cost variable: some vendors continue the billing clock after a call is transferred to a human agent. A contact center running 5,000 warm transfers per month at an average 3-minute post-transfer duration, billed at even $0.08/min, adds $1,200 per month to the invoice for time the AI is no longer doing anything. Bland.ai's Scale plan lists a separate transfer-minute rate of $0.03/min, which at least makes the cost visible and auditable.

11. Concurrent Call Limit Pricing - Scalability Gated by Simultaneous Call Ceiling#

Concurrent call limits determine how many calls your account can handle at the same moment. Exceeding your limit does not just cost more; it means calls queue or drop. Bland.ai's Build plan supports 50 concurrent calls; Scale supports that same figure. For outbound campaigns with burst dialing patterns, the concurrent ceiling is often the binding constraint that forces an upgrade, not the per-minute rate.

12. Talk-Time-Only Billing - Meter Runs Only During Active Speech, Not Total Session Time#

Talk-time-only billing starts the meter when a caller or agent speaks and pauses it during silence, hold music, or processing latency. If 20 percent of a typical call is non-speech time, talk-time billing can reduce the effective billed duration by a meaningful fraction compared to total session billing. Always confirm the definition in the vendor's billing policy, not the sales deck.

13. Total Session Time Billing - Meter Runs from Call Connect to Disconnect Regardless of Silence#

Total session time billing starts the clock the moment the call connects and runs until disconnection, regardless of silence, hold music, or processing pauses. Buyers who assume they are paying for conversation time and are actually paying for session time will systematically underestimate their costs. Silence and hold billing vary by vendor and are not visible on standard pricing pages.

14. Silence Billing Exposure - How Dead Air During AI Processing Inflates Effective Cost#

Every conversational turn in an AI phone call includes a processing window: STT transcribes the caller's utterance, the LLM generates a response, and TTS renders audio. On platforms routing inference through third-party APIs, this latency window can run 800 milliseconds to over 2 seconds per turn. On a total-session billing model, that dead air is billed time. A 6-minute call with 20 turns and 1.5 seconds of processing latency per turn adds 30 seconds of billed silence. Across thousands of calls, the cumulative cost is real.

15. Hold Time Billing - Meter Continues While Caller Waits on Hold During AI-Initiated Pauses#

Hold time billing covers the period when the AI agent explicitly places a caller on hold to look up data, initiate a transfer, or wait for a backend API response. If the meter runs during this window, every slow CRM lookup or delayed API response becomes a billable event. The risk is highest for workflows with real-time database lookups mid-call, which are common in healthcare scheduling, financial account verification, and insurance intake flows.

16. Infrastructure DIY Stack - Self-Hosted STT, LLM, TTS, and Telephony at Component Cost#

Building your own stack from component APIs gives maximum control and, at face value, the lowest per-component cost. The engineering reality is more complicated: latency management across three or four separate API calls per conversational turn, version drift when any upstream provider updates a model, and the absence of a single support path when something breaks in production. The headline per-minute rate looks low until you price in the engineering hours.

17. Agency Build Model - White-Label AI Phone Agent Deployment with Markup Margin#

Image: AI Phone Agent Pricing - agency build model white

Agency-built AI phone agents follow a two-layer cost structure: the agency pays platform costs and marks up to the end client. Simple deployments typically run that same figure to $2,000 upfront plus $300 to $500 per month in retainer. The risk is dependency: if the agency's underlying platform changes pricing, the client's invoice changes too.

18. Voice Cloning Feature Gate - Custom Voice Locked Behind Professional Tier or Add-On Fee#

Custom voice cloning, training the AI to speak in a brand-specific or individual voice, is consistently gated behind Professional/Business tiers or sold as a standalone add-on at $50-$200/month per voice. For brands where voice consistency is a core customer experience element, this feature gate forces an upgrade even if no other Professional-tier features are needed. Buyers should evaluate whether a vendor's stock voice library is sufficient before paying the cloning premium.

19. HIPAA Compliance Feature Gate - BAA and Encrypted Call Handling Locked Behind Enterprise Tier#

HIPAA-compliant call handling, including a signed Business Associate Agreement and encrypted call storage, is locked behind Enterprise tiers across virtually every major platform. A healthcare buyer who builds a cost model on Starter or Professional tier pricing and then discovers the BAA requires Enterprise pricing has wasted the entire evaluation period. Confirm BAA availability and pricing before any other feature comparison.

20. CRM Integration Feature Gate - Native Salesforce, HubSpot, and Zoho Sync Behind Mid-Tier Plans#

Native CRM sync is typically gated behind mid-tier or higher plans. Aircall's platform supports 200-plus integrations including Salesforce, HubSpot, Intercom, Zendesk, Pipedrive, and MS Dynamics, but availability depends on plan tier. Call Agent AI's Zoho CRM integration is available as a $20/month add-on for Starter and Growth plans but is included in Pro and Enterprise. Map your required integrations to specific tiers before pricing.

21. Multi-Agent Workflow Feature Gate - Parallel Agent Orchestration Locked Behind Business Tier#

Multi-agent workflows are gated behind Business or Enterprise tiers on most platforms. For high-volume outbound campaigns or complex inbound routing trees, this is a functional requirement, not a nice-to-have. Buyers who pilot on a lower tier and then discover multi-agent orchestration requires an upgrade face both a cost surprise and a re-architecture of their call flows.

22. Flat-Rate All-In Monthly Model - Unlimited Minutes Within a Fixed Monthly Fee#

Flat-rate unlimited-minute models charge a single monthly fee regardless of call volume. Vendors who offer unlimited minutes must price the flat rate to cover their worst-case usage scenario, which means low-volume buyers subsidize high-volume ones. This model is worth evaluating only if your monthly minute consumption is high enough that the flat rate is cheaper than a per-minute equivalent.

23. Per-Outcome Billing Model - Charged Per Successful Appointment, Lead, or Resolution#

Per-outcome models charge only when the AI agent achieves a defined result: a booked appointment, a qualified lead, or a resolved support ticket. The practical challenge is defining "success" contractually in a way that is auditable and resistant to gaming. For regulated industries, outcome-based billing also raises the question of what happens when a call achieves a business outcome but fails a compliance check.

24. Seat-Based Pricing Model - Fixed Cost Per Agent Seat Regardless of Call Volume#

Seat-based pricing charges a fixed monthly fee per AI agent 'seat', typically $50-$200/seat/month, regardless of how many minutes that agent handles. This model is familiar to buyers coming from traditional CCaaS platforms and simplifies budgeting. The limitation is that it penalizes efficient deployments: a single AI agent handling 10,000 minutes costs the same as one handling 500 minutes, making it economically inferior to per-minute models at high utilization rates.

25. Volume Discount Tier Pricing - Per-Minute Rate Decreases at Committed Monthly Minute Thresholds#

Volume discount structures reduce the per-minute rate as buyers commit to higher monthly minute thresholds. The negotiation lever is the commitment term: vendors will often offer a lower rate in exchange for a 6 or 12-month volume commitment. The risk is over-committing: if actual volume falls short of the committed threshold, the effective per-minute rate rises, sometimes above the no-commitment rate.

26. Annual Commitment Discount - 15-25% Rate Reduction for 12-Month Contract Prepayment#

Annual prepayment discounts of 15 to 25 percent are standard across SaaS platforms. A $299/month platform fee paid annually at a 20 percent discount saves roughly $717 over the year. The risk: if your call volume drops, your use case changes, or the vendor's product quality degrades, you have limited recourse on a prepaid annual contract. Confirm cancellation and credit-back terms before prepaying.

27. Telephony Pass-Through Add-On - PSTN Call Costs Billed Separately from AI Processing Fees#

Some platforms bill AI processing at one rate and pass through PSTN telephony costs as a separate line item. Twilio's standard voice pricing, for example, is billed per minute of call time and sits entirely outside any AI processing fee. Buyers building cost models from a single per-minute rate who miss the telephony pass-through will underestimate their invoice by a consistent margin on every single call.

28. Phone Number Provisioning Add-On - Monthly Fee Per DID Number Beyond the Included Allocation#

Direct inward dial number provisioning is frequently excluded from base plan pricing. Platforms typically include one or two numbers in the base plan and charge $1 to $5 per additional number per month. For multi-location deployments or campaigns requiring local presence numbers across dozens of area codes, DID provisioning fees can add hundreds of dollars per month to an invoice that the per-minute rate comparison never surfaced.

29. Call Recording and Transcription Add-On - Storage and Processing Fees Beyond Base Plan#

Call recording and real-time transcription are often included at base tiers but subject to storage caps or retention limits. For regulated industries where call recording is a compliance requirement, the storage cost is non-negotiable and must be modeled as a fixed cost component. Confirm retention period, storage pricing, and export rights before treating transcription as "included."

30. Real-Time Analytics Dashboard Add-On - Advanced Reporting Gated Behind Premium Tier#

Real-time call analytics, sentiment scoring, intent classification, conversion tracking, and agent performance dashboards, are gated behind premium tiers or sold as add-ons at $50-$150/month. For operations teams using AI phone agents to replace human agents, these analytics are essential for quality assurance and ROI measurement. Buyers who treat analytics as optional during vendor selection often find themselves paying the add-on fee within 60 days of deployment.

31. Outbound Dialing Add-On - Separate Per-Minute Rate for AI-Initiated Outbound Calls#

Many platforms price inbound and outbound calls at different per-minute rates, with outbound AI-initiated calls often carrying a premium. For buyers running AI-powered outbound calling campaigns, the outbound rate is the rate that matters, and it is frequently buried in the pricing page's fine print rather than the headline rate. Confirm the outbound rate explicitly before modeling campaign costs.

32. Retell AI Pricing - Per-Minute Consumption with Included LLM and TTS Components#

Retell AI uses a per-minute consumption model. The platform's cost structure includes LLM and TTS components but also surfaces token surcharges and feature-gate fees as hidden variables that do not appear in a simple per-minute comparison. Buyers evaluating Retell alongside all-in alternatives should audit whether the per-minute rate is truly all-in or whether token and TTS costs layer on top under specific usage conditions.

33. VAPI Pricing - Developer-First Per-Minute Model with Bring-Your-Own LLM Keys#

VAPI's developer-first model allows buyers to bring their own LLM API keys, meaning the LLM cost is entirely separate from the platform fee and billed directly by the LLM provider. This gives sophisticated engineering teams maximum control over model selection and cost optimization.

34. Synthflow Pricing - Subscription Tiers with Included Minutes and No-Code Agent Builder#

Synthflow offers subscription tiers starting around $29-$500/month with included minute bundles and a no-code agent builder targeting non-technical buyers. The no-code interface reduces deployment time but limits customization depth compared to API-first platforms. Buyers who outgrow the included minutes face overage rates that can make the effective cost competitive with higher-tier plans, making volume forecasting critical before selecting a Synthflow tier.

35. Twilio Voice Intelligence Pricing - Usage-Based AI Layer on Top of Existing Twilio Telephony#

Twilio's AI voice intelligence layer adds per-minute AI processing costs on top of existing Twilio telephony rates, making it a natural fit for businesses already using Twilio infrastructure. The combined cost, Twilio telephony ($0.013/min) plus AI processing ($0.05-$0.10/min), is competitive for high-volume deployments. The tradeoff is that building a full AI phone agent on Twilio requires significant developer investment compared to turnkey platforms.

36. Deepgram STT Component Pricing - $0.0043/min Speech-to-Text as a Stack Building Block#

Deepgram's Nova-2 model prices real-time speech-to-text at approximately $0.0043/minute, making it the most cost-efficient STT option for DIY stack builders. At 10,000 minutes/month, Deepgram STT costs $43 versus $150-$300 for the STT component bundled inside managed platforms. The limitation is integration complexity: buyers must handle streaming audio, WebSocket connections, and error handling themselves rather than relying on a managed platform's abstraction layer.

37. ElevenLabs TTS Add-On Pricing - Premium Voice Quality at $0.006-$0.018/min Billed by Character#

ElevenLabs text-to-speech is billed by character rather than by minute, with costs translating to approximately $0.006-$0.018/min of generated audio depending on the voice tier selected. It delivers the highest naturalness scores among TTS providers but costs 3-5x more than alternatives like Google TTS or Amazon Polly. For brands where voice quality is a primary differentiator, the premium is justified; for high-volume commodity call flows, cheaper TTS options significantly reduce stack cost.

38. Cost-Per-Resolution Pricing Model - Billed Per Successfully Resolved Customer Issue#

Emerging cost-per-resolution models charge only when the AI phone agent fully resolves a customer issue without human escalation, typically $1.50-$8 per resolved interaction. This model aligns perfectly with contact center ROI metrics and eliminates payment for failed or escalated calls. The limitation is measurement complexity: 'resolution' must be contractually defined, and vendors may define it more broadly than buyers expect, leading to disputes over what constitutes a billable successful outcome.

39. Professional Services and Setup Fee - One-Time Deployment Cost Beyond Subscription Pricing#

Most Enterprise and many Professional-tier AI phone agent deployments include one-time professional services fees for agent configuration, integration development, and voice training, typically $2,000-$25,000 depending on complexity. These fees are rarely included in published pricing and can represent 3-12 months of subscription cost. Buyers should request a full deployment cost estimate including professional services before comparing total cost of ownership across vendors.

40. Pay-As-You-Go vs. Subscription Break-Even Analysis - The Volume Crossover Point#

The break-even between PAYG and subscription models depends on the platform fee and rate differential. For a platform charging $299/month for a $0.02/min rate discount (e.g., $0.14 PAYG vs. $0.12 subscription), break-even occurs at 14,950 minutes/month. Below that threshold, PAYG is cheaper; above it, subscription wins. Buyers should calculate their 3-month average call volume before selecting a billing model, accounting for seasonal variance that could push them above or below the crossover point.

41. Token Cost Delta at 10,000+ Minutes/Month - All-In vs. Token-Surcharge Model Comparison#

At 10,000 minutes per month, the cost delta between all-in and token-surcharge models becomes material. An all-in model at $0.18/min costs $1,800. A token-surcharge model at $0.12/min base plus $0.04-$0.08/min in LLM tokens costs $1,600-$2,000 depending on conversation complexity. For simple, short-turn call flows, token-surcharge models are cheaper; for complex, multi-turn conversations, all-in models provide cost certainty and often lower total cost.

42. Minimum Monthly Spend Commitment - Floor Charge Regardless of Actual Usage#

Some vendors impose a minimum monthly spend, typically $100-$500, regardless of actual minute consumption. This protects vendor revenue but penalizes low-volume buyers who pay for capacity they don't use. Minimum spend commitments are common in Enterprise contracts and occasionally appear in mid-tier plans. Buyers with variable or seasonal call volume should negotiate minimum spend waivers or select PAYG models without floor charges.

43. Overage Rate Penalty - Per-Minute Rate Spike When Exceeding Bundled Minute Allowance#

Subscription plans with bundled minutes typically charge overage rates of $0.15-$0.35/min when usage exceeds the included allowance, often 20-50% higher than the effective per-minute rate within the bundle. For businesses with unpredictable call spikes, overage charges can double a monthly invoice. Buyers should negotiate overage rate caps or select plans with included minute buffers that accommodate their 90th-percentile usage month, not just their average.

44. Multi-Tenant Agency Pricing - Per-Client Workspace Fee with Shared Minute Pool#

Agency-focused pricing tiers charge a per-client workspace fee ($20-$75/month per client account) plus access to a shared minute pool billed at wholesale rates. This model allows agencies to manage multiple client deployments under one account while maintaining billing separation. The tradeoff is that shared minute pools create cross-client cost allocation complexity, and a single high-volume client can consume disproportionate pool capacity, affecting other clients' effective costs.

AI Phone Agent Cost Estimator - How to Calculate Your Real Monthly Bill Before You Sign#

Per-client minute caps protect the agency's margin structure, but they do nothing to close the gap that appears between a buyer's initial projection and the invoice that arrives after the first full billing cycle. That gap, which buyers consistently report running 40-60% above their initial projection, is almost never caused by a math error. It is caused by a structural blind spot: the formula most buyers use treats the pricing page as the whole picture, when it is actually just the top layer of a four-layer cost stack.

One of the clearest pain points we see among teams evaluating AI phone agents for the first time is that there is no obvious, pre-commitment way to estimate what the platform will actually cost month-to-month, which makes it nearly impossible to build a credible ROI case against the hours currently lost to manual calling. The cost model below is designed to eliminate that blind spot before you sign anything.

Four-card grid showing the four hidden cost layers in an AI phone agent monthly bill

The Four-Layer Stack Audit - What You're Actually Multiplying Before You Multiply Anything#

Before any multiplication makes sense, you need to know what you are multiplying. Pricing analysis, per-minute rates listed on vendor pricing pages are only one component of the true monthly bill. Buyers must also account for platform fees, token surcharges from third-party LLMs, and speech-to-text or text-to-speech overages not included in the advertised base rate. On platforms that route through external inference providers, Aircall's research confirms that AI voice agent pricing ranges from $0.05 to $1+ per minute once those layers are added back in.

The four layers are: (1) platform or seat fee, (2) per-minute talk-time rate, (3) LLM token charges billed per that same figure tokens, and (4) STT/TTS metering billed per character or per second by the upstream vendor.

Bland.ai's pricing structure collapses layers 3 and 4 into layer 2. Across all self-serve tiers, Start at $0.14/min, Build at $0.12/min, and Scale at $0.11/min, LLM token charges, real-time transcription (STT), and premium voices including voice clones (TTS) are included in the per-minute rate. There are no token surcharges billed on top. The effective rate on the pricing page is the rate you multiply, which is a meaningful structural advantage when building a cost model that survives contact with the first invoice.

Step-by-Step Cost Estimation Formula - Minutes x Rate Is Only the Starting Point#

Start with your projected monthly call volume and multiply by your expected average call duration to get total minutes. Industry data for enterprise deployments points to average call durations in the 3-to-5-minute range for qualification and scheduling calls, and 6-to-10 minutes for complex support or intake calls. Then apply the per-minute talk rate, add the platform fee, and add any token or STT/TTS surcharge expressed as a per-minute equivalent.

The honest formula: (total minutes × all-in effective rate) + platform fee + integration costs + compliance add-ons = projected monthly bill. Skipping any term produces a number that will not survive contact with the first invoice.

Ai's published plan facts: a team on the Build plan running that same figure calls per month at an average of 4 minutes per call generates 4,000 talk minutes. 12/min with no token or STT/TTS surcharge layered on top, that is $480 in usage plus the $299/month platform fee, a total of $779 before any transfer minutes. 04/transfer min on Build.

That is a complete, auditable number. 11/min, $499/month platform fee) at the same volume produces $440 in usage plus $499, totaling $939, a higher absolute figure but one that becomes favorable as volume climbs past the break-even threshold. Running that comparison before you commit is the entire point of the four-layer audit.

One health insurance team using Bland.ai for open enrollment outreach described the operational impact directly: the AI handles the first touch, qualifies the lead, and transfers to the human team, enabling the organization to scale outreach volume without hiring additional agents. That model, AI for high-frequency first-touch calls, humans for high-value closes, is exactly where the cost-per-outcome math becomes favorable, and where maintaining complete control and observability over AI agent behavior matters most.

How Concurrent Call Limits Force Tier Jumps That Dwarf Per-Minute Savings#

As industry research documents, concurrent call capacity determines plan tier eligibility. A buyer optimizing for the lowest per-minute rate, without checking the concurrent call ceiling, can find that a single busy Monday morning forces a tier jump that costs more per month than an entire year of per-minute savings would have recovered.

Ai, concurrent call ceilings are explicit and published: Start supports 10 concurrent calls with a that same figure-call daily cap, Build supports 50 concurrent calls with a 2,000-call daily cap, and Scale supports that same figure concurrent calls with a 5,000-call daily cap. If your outbound campaign or open enrollment window generates call spikes that exceed those ceilings at your current tier, the tier jump is predictable and priceable in advance. Enterprise concurrency is sized to your contracted volume with no published ceiling, which is the appropriate structure for organizations running continuous inbound and outbound call flows at scale.

The ability to identify call volume trends in real time, including sentiment patterns across all customer calls, gives operations teams the observability needed to anticipate when a concurrent ceiling is about to become a constraint, before it forces a reactive tier upgrade.

---

Pre-Signature Vendor Pricing Checklist#

Use this checklist before signing any AI phone agent contract. Each unchecked item is a potential invoice surprise.

  • Budget fit → What to ask: Does the advertised per-minute rate include STT, LLM, and TTS, or are those billed separately? → Why it matters: Separates all-in from pass-through models.
  • Billing start point → What to ask: Does the billing clock start at call connection or at first spoken word? → Why it matters: Session-time vs. talk-time billing can differ by 15–25%.
  • Hold and silence billing → What to ask: Does billing pause during hold music and silence? → Why it matters: Silence billing inflates the effective rate on processing-heavy workflows.
  • Transfer minutes → What to ask: Is a separate rate charged for transfer minutes after human handoff? → Why it matters: Post-handoff billing doubles the cost of every escalation.
  • LLM token charges → What to ask: Are LLM token charges billed on top of the per-minute rate? → Why it matters: Token surcharges are invisible on the pricing page.
  • Concurrent call ceiling → What to ask: What is the concurrent call ceiling at my tier, and what does an upgrade cost? → Why it matters: Concurrent limits often force tier jumps that dwarf per-minute savings.
  • HIPAA/BAA compliance → What to ask: Is HIPAA/BAA compliance included at the quoted tier or gated behind Enterprise? → Why it matters: Compliance gates invalidate any cost model built on a lower tier.
  • Included call features → What to ask: Are phone number provisioning, call recording, and outbound dialing included? → Why it matters: These add-ons are routinely excluded from headline pricing.
  • Billing increment → What to ask: What is the billing increment: per second, per 6 seconds, or per full minute? → Why it matters: Rounding alone can inflate invoices by 2–4× across high call volumes.
  • Overage rates → What to ask: What are the overage rates if I exceed my included minute bundle or concurrent cap? → Why it matters: Overage rates are typically 30–50% above the base effective rate.

Hidden Fees and Pricing Pitfalls That Make Your AI Phone Agent Invoice Unrecognizable#

The surprise on your invoice is a design problem. The common assumption is that if you collect per-minute rates from each vendor's pricing page and divide by call volume, you will know your real monthly cost before signing.

That assumption is wrong. Every unexpected line item traces back to a platform that outsourced a layer of its stack, speech recognition, voice synthesis, LLM inference, and then passed that vendor's metering model straight through to you. Reading the pricing page more carefully will not save you, because the cost structure is baked into how the platform was built, not hidden in the footnotes.

One fear that surfaces early for teams deploying AI voice at scale is that the agent will quote costs or capabilities that were never sanctioned, creating billing or customer-expectation mismatches that are genuinely hard to walk back. That risk compounds when the pricing model itself is opaque. A platform whose rate structure is knowable before the first call goes out removes that risk. For businesses handling high call volumes or needing 24/7 phone coverage without scaling headcount, the math has to be deterministic from day one.

1. LLM Token Surcharges Stacked on Top of Per-Minute AI Phone Agent Pricing#

Platforms that route calls through frontier LLM providers do not bill you a flat rate for intelligence. They bill you for every token the model processes. Industry data confirms that audio input and output tokens are metered independently of any platform or telephony fee, and a five-minute conversation can consume thousands of tokens with that meter running in parallel with your per-minute rate.

The result is a cost multiplication effect that never appears in the headline number. Platforms built on self-hosted inference eliminate this layer entirely. When the LLM runs on the vendor's own infrastructure, the per-minute rate is the inference rate.

Key takeaway: A five-minute conversation can consume thousands of LLM tokens billed in parallel with your per-minute rate. Platforms that self-host inference eliminate this layer entirely, making the per-minute rate the inference rate.

Bland.ai's per-minute rates are all-in by design: LLM inference carries no token charges and is included in the per-minute rate across every plan tier. On the Start plan for developers, that means $0.14/min with no platform fee and no card required to begin. Teams graduating to Build pay $0.12/min plus a $299/month platform fee and gain 50 concurrent calls and 2,000 daily call capacity. High-volume operations on Scale pay $0.11/min plus a $499/month platform fee with 100 concurrent calls, 1,000 hourly cap, and 5,000 daily calls, with the lowest per-minute rate reflecting the volume those operations actually run. In every case, the per-minute number is the number; there is no parallel token meter running underneath it.

2. STT/TTS Metering That Layers Character and Second-Based Fees onto Every Call#

Speech-to-text and text-to-speech components each carry their own metering logic. Across the market, when STT and TTS are not self-hosted, vendors pass through character- or second-based charges that stack invisibly on the headline per-minute price. A platform advertising $0.07 per minute can land closer to $0.11 to $0.14 per minute on a typical 300-word-per-minute conversation. Two meters run simultaneously on every call, one for what the caller says, one for every character the agent speaks back.

Bland.ai includes real-time transcription and premium voices plus voice clones inside the same per-minute rate on every plan. Here is how voice clone allotments break down across tiers, all covered under the flat per-minute rate, not metered separately by character or second:

  • Start: 1 voice clone, access to 15 voices
  • Build: 5 voice clones
  • Scale: 15 voice clones

For teams running outbound campaigns such as sales calls, follow-ups, and reminders alongside inbound call handling for customer support and intake, this matters most at scale: at that same figure of concurrent calls running on Scale, a hidden STT/TTS meter would compound across every simultaneous session. Because both layers are folded into $0.11/min, the per-minute figure you model before launch is the per-minute figure that appears on the invoice.

3. Transfer-Minute Billing - The Clock That Keeps Running After Human Handoff#

When escalation happens on platforms that continue billing after handoff, the cost of a single transferred call can dwarf the cost of the AI leg that preceded it. Bland.ai meters transfer minutes at a separate, lower rate that is visible before you commit. That stepped structure means the more volume you run, the lower your transfer-minute exposure per call, consistent with the same volume logic that governs the talk-time rate. The figure is published, not discovered post-invoice.

  • Start → Talk-time rate: $0.14/min → Transfer-minute rate: $0.05/transfer min.
  • Build → Talk-time rate: $0.12/min → Transfer-minute rate: $0.04/transfer min.
  • Scale → Talk-time rate: $0.11/min → Transfer-minute rate: $0.03/transfer min.

For regulated or complex call flows where human handoff is a compliance requirement rather than an edge case, the kind of calls that most teams running at scale identify as a persistent cost-modeling blind spot, Bland.ai's Enterprise tier sizes concurrency and billing to contracted volume, with dedicated infrastructure, a 99.9% uptime SLA, and compliance documentation available under NDA. The forward-deployed engineering team operates on a defined deployment framework, scope, build, gray/red/green-team test, and go live, with Enterprise deployments live in production in 30 days. For organizations already running on Amazon Connect, Bland.ai's integrations platform allows AI voice agents to be added to existing inbound and outbound call flows without migrating to a new platform, keeping the cost model legible within infrastructure the team already operates and understands.

4. Silence and Hold-Time Billing - Paying Full Rate for Dead Air#

Some AI phone agent platforms meter total wall-clock session time rather than active talk time, meaning every second a caller spends on hold, navigating IVR prompts, or sitting in silence while the agent processes a query is billed at the full per-minute rate. In contact center environments where average silence ratios run 15-25% of call duration, this billing method inflates costs significantly. Always request explicit contract language distinguishing active talk time from total session time before committing.

5. Concurrency Overage Charges and Call Cap Risks Hidden in AI Phone Agent Plans#

Most AI phone agent pricing tiers impose hard limits on simultaneous concurrent calls, and exceeding those limits triggers steep overage fees, often 2-3x the base per-minute rate for burst capacity. Worse, some platforms enforce daily or hourly call caps that cause calls to fail outright rather than queue, creating operational risk alongside financial risk. Businesses with unpredictable inbound volume must negotiate explicit concurrency guarantees and overage rate caps before deployment, not after the first billing cycle.

AI Phone Agent Pricing FAQ - Billing, Minutes, Transfers, and Compliance Costs Answered#

A pricing page tells you the rate. The invoice tells you the truth. For enterprise buyers running thousands of calls a month in regulated environments, the distance between those two numbers is where budget surprises live, and where due diligence either pays off or fails you.

Does Billing Run During Hold Time, Silence, and IVR Prompts, or Only During Active Speech?#

Do and don't checklist for confirming AI phone agent billing terms before signing

Billing practices vary significantly by vendor: some platforms bill for the full session duration including silence and hold time, while others bill only for active talk time. For workflows with long IVR sequences or frequent hold periods, that distinction can materially inflate your effective per-minute rate without changing a single line on the pricing page.

Before signing, confirm the following with any vendor in writing:

  • Does the billing clock start at call connection or at first spoken word?
  • Does billing pause during hold music?

Are Transfer Minutes Billed at the Same Rate as Talk Minutes, and Who Pays After the Handoff?#

Some vendors continue charging their AI per-minute rate even after a call transfers to a human agent, billing twice for the same minutes. For any workflow with a meaningful human-handoff rate, that continuation policy can turn a seemingly competitive rate into a significant cost problem at scale.

Transfer-minute billing deserves its own contract line. Bland.ai, for instance, publishes a separate $0.03/transfer minute rate on its Scale plan, compared to the $0.11/min talk rate, so the distinction is explicit rather than buried. Before committing, ask vendors:

  • Does a separate, lower rate apply post-handoff?
  • Does billing stop entirely at transfer?
  • Are warm transfers metered differently than cold ones?

How AI Minutes Are Actually Calculated - Talk Time vs. Session Time vs. Connection Time#

Three vendors can use three different definitions of a "minute," and all three can be technically accurate. Session time starts at connection and runs until disconnect. Talk time excludes silence. Some platforms round up to the nearest 6-second increment; others bill full minutes. Two vendors advertising identical per-minute rates can produce invoices that differ by 2-4x for the same call set, because billing increment rounding and session-time definitions compound across thousands of calls.

  • Demand contractual disclosure of the billing increment and the session-time definition.

Why the Lowest Latency Platform Delivers the Lowest Real Cost - and How to Validate It Before You Buy#

Owning the full inference stack is what allows a vendor to quote a pricing page rate that matches the invoice. That same architectural choice, co-located STT, LLM, and TTS on vendor-controlled hardware, is also what eliminates the network hops responsible for processing latency. When a vendor operates its own infrastructure end-to-end, it removes the handoff points where both metered costs and response-time delays accumulate. Every third-party hop a packet must cross is a billing event for the vendor and a latency tax for the caller; removing those hops to control pricing also, as a structural side effect, produces the shortest possible inference path. The contract clause you audit to confirm pricing transparency is therefore auditing the same architectural fact that determines call quality.

Bland Speech v3 was trained on over 100 million real human conversations, teaching the model conversational speech patterns rather than polished studio delivery.

Side-by-side comparison of stitched third-party stack versus Bland's fully owned inference stack

Buyers who have watched their AI phone agent invoices balloon past the quoted per-minute rate already sense something is structurally wrong. Latency and cost are jointly determined by a single architectural choice: whether the platform owns its inference stack end-to-end or stitches together third-party STT, LLM, and TTS providers. Fixing one fixes the other, because they share the same root cause.

Key takeaway: Latency and cost share the same root cause; every third-party API hop is simultaneously a billing event and a latency tax. A platform that owns its inference stack end-to-end eliminates both at once.

The Architectural Reason Latency and Cost Move Together, Not Separately#

Every call an AI phone agent handles moves through a sequence: speech goes in, gets converted to text, a language model reasons over it, text comes back, and a voice synthesizes the response. On a platform that owns its inference stack end-to-end, that sequence runs on co-located hardware. On a platform that stitches together third-party providers, each step crosses a network boundary. Those boundary crossings are where latency and cost are jointly created. Buyers who score latency on the performance tab and cost on the finance tab are evaluating the same variable twice without realizing it.

How Third-Party API Hops Silently Inflate Both Response Time and Your Invoice#

Builders who have run cascade pipelines through external STT, LLM, and TTS providers consistently report 3 to 5 seconds of end-to-end latency, a threshold that independent benchmarks and practitioner reports describe as making conversations feel broken rather than natural. (Latency accumulation across multi-hop inference pipelines is a well-documented structural property of stitched architectures.) Each hop adds up to 400 milliseconds of network round-trip time, a range consistent with standard TCP/IP round-trip benchmarks for cross-datacenter API calls documented in network performance literature and corroborated by what most teams operating these pipelines report.

It also adds a metered charge. As broader market patterns confirm, platforms routing through third-party providers pass those per-token or per-character costs on.

Next steps#

If your actual invoices keep arriving 40-60% above the number you projected from the pricing page, the path forward starts with understanding that per-minute rates are entry tickets, not total costs. The real price is determined by which billing layers stack underneath: session-time definitions, transfer-minute continuation, LLM token surcharges, and STT/TTS metering, none of which appear on a pricing page but all of which compound on the invoice. Start with our best AI phone agent platform for enterprises.

Two findings from this breakdown make the next step obvious. First, two vendors advertising identical per-minute rates can produce invoices that differ by 2-4x for the same call set, because billing increment rounding, session-time definitions, and transfer policies compound independently across thousands of calls. Second, latency and cost are jointly determined by the same architectural choice: platforms that own their inference stack end-to-end eliminate both the network hops that create processing delays and the upstream metering events that inflate invoices. Evaluating them separately means missing that fixing one fixes the other. Together, those two findings point to auditing architecture before auditing rates.

Start with bland.ai to see a pricing model where STT, TTS, and LLM inference are folded into a single per-minute rate with no token surcharges layered underneath. From there, you can run the four-layer stack audit from the cost estimator section against your actual call volume, concurrency ceiling, and compliance tier, and arrive at a projected monthly bill that survives contact with the first invoice.

Frequently Asked Questions#

Why is my actual invoice so much higher than the per-minute rate I saw on the pricing page?#

Most AI phone agent platforms are assembled from four separate layers, telephony, a large language model, speech-to-text, and text-to-speech, each of which meters and bills usage independently. The pricing page typically quotes only the telephony layer, so LLM token charges, STT per-second fees, and TTS per-character fees all land on the invoice without ever appearing on the pricing page. A platform quoting $0.09/min that also passes through LLM token costs can reach $0.22/min or higher on a single three-minute support call.

What am I actually paying for when I see a per-minute rate?#

A per-minute rate on most pricing pages covers only the telephony relay, not the LLM inference, speech-to-text transcription, or text-to-speech synthesis that also run during every call. Platforms that own their full stack and bundle all four layers into one rate are the exception; with those platforms, like bland.ai's paid plans, the advertised rate covers STT, LLM inference, TTS, and real-time transcription with no separate token charges added on top.

Do I keep getting charged after my AI agent transfers a call to a human?#

It depends on the vendor. Some platforms continue the billing clock after a call is transferred to a human agent, meaning you are paying per-minute for time the AI is no longer doing anything. Bland.ai's Scale plan lists a separate transfer-minute rate of $0.03/min, which at least makes that cost visible and auditable rather than burying it in the base rate.

How do concurrent call limits actually affect what I pay?#

Concurrent call limits cap how many calls your account can handle at the exact same moment, exceeding the limit means calls queue or drop, not just cost more. Bland.ai's Build plan supports 50 concurrent calls and the Scale plan supports 100, and for outbound campaigns with burst dialing patterns, hitting that ceiling is often what forces a plan upgrade, not the per-minute rate itself.

Am I billed during silence and hold time, or only when someone is speaking?#

That depends on whether the vendor uses talk-time-only billing or total session billing. Talk-time-only billing pauses the meter during silence, hold music, or processing latency, and if 20 percent of a typical call is non-speech time, that difference can meaningfully reduce your billed duration compared to a vendor who meters the full session. Always confirm the definition in the vendor's billing terms before signing.

See Bland on your actual call volume.

10 to 15 minutes with the team that ships your first agent. We come prepared with answers, not a pitch deck.

Book a call