Skip to content

The runtime under every enterprise voice agent.

Hundreds of teams are building voice agents. Very few build what those agents run on. Voicing is the runtime underneath: speech in, decision, action, speech out, inside your firewall.

  • ~10,000voice agents live
  • 99.98%agent uptime
  • 30+languages
  • 700 msend to end, 80th percentile, real telephony
The wall01 / 08

Agents hold in the demo. They break in production.

Past simple questions, every extra step adds tokens, tools and latency. The failure is architectural, not a model choice, and it shows in four places.

See what it looks like in production

  1. Basic questions
  2. Multi-turn handling
  3. Intent and routing
  4. Personalised support
  5. Proactive resolution
  6. Complex problem solving
  7. Negotiation and resolution
  • Tokens consumed per flow
  • Tasks and their complexity
  • The complexity ceiling
  • General-purpose models, rented part by part
  • A specialised, cost-aware runtime

In the demoIn production

  1. End-to-end latency

    About 700 ms1,400 to 1,700 ms

  2. Entity capturePhone, card, date of birth

    Feels accurate18 to 34% error

  3. Function calling

    90%+ on textUnder 70% on 8+ tools

  4. Cost per minute

    $0.08 to $0.1440 to 50% higher

See what it looks like in production

The category02 / 08

Hundreds build the agent. Few build what it runs on.

An application can be swapped in a quarter. The layer underneath takes a security review, a telephony integration and a model validation to replace. Choose the runtime first.

The application layerVoice bots, IVR replacement, collections, claims, agent assist
  • 300+vendors
  • Per-seat pricing
  • One quarterto replace
runs on
The voice runtimeSpeech in, decision, action, speech out, inside your firewall
  • Under 5providers
  • Per-minute pricing
  • Multi-yearto displace

Four products ship on it. Or build your own on the SDK.

The runtime03 / 08

Seven layers. One owner.

Every layer is researched, trained and served by Voicing. Wrappers fail exactly where enterprise value lives: in the hops between rented parts.

  1. React, Python, Node and REST, with pre-built agents130+agent templates
  2. Stateful flows, guardrails, redaction, consent and audit98%process adherence
  3. Action-first fine-tunes, self-hosted, with vendor fallback97%function calling
  4. A tuned ASR ensemble that agrees on names, numbers, dates150 msto a final transcript
  5. Fine-tuned voices, prosody control, low-latency streaming38 msto first audio
  6. Semantic endpointing, barge-in and overlap, no silence timer60 to 120 mstrigger
  7. SIP, WebRTC and PSTN, with memory that persists across turnsUnder 5 msturn state

Rented, part by part

A hop between every part, and latency in every hop.

Illustrative sequence

  1. Application SDK: React, Python, Node, REST, WebSocket, Runtime, Verification, Payments, Handoff. Your agents attach here. Pre-built parts, your logic
  2. Orchestration and policy: Consent, Audit, Greet, Verify, Redact, Act, Retry, Card ending 4417. Redaction and consent in the path. Every step stamped for audit
  3. Agentic LLM router: Voicing-Talk 3B, Good morning. Voicing-Act 8B, Move my payment date. Voicing-Orchestrate 32B, Two accounts, one dispute. Bring your own model. The smallest model for the job. Self-hosted, with a fallback
  4. STT and entity capture: Name, Date of birth, Card, Transcript, My name is Dana Harlow, Marlow, Dana Harlow, 12 Mar 1986, ending 4417. Three recognisers, one answer. Entities validated before use
  5. TTS engine: Prosody, Text, Audio, Your payment now moves to Friday. Speaks from the first words. Prosody tuned to the brand
  6. Turn-taking and V-VAD: Caller, Agent, A silence timer fires here, Endpoint, Barge-in, Your payment now moves to Friday. Waits through the pause. Yields when the caller cuts in. F1 0.92 endpoint detection. About 45% fewer false barge-ins
  7. Transport and memory: SIP, WebRTC, PSTN, Memory, Caller, Account, Reason, Promise. One line in, any carrier. Memory outlives the turn

The caller stops speaking.

Endpoint detected60 to 120 ms

Final transcript150 ms

Routed to the action model

Read through, nothing copied out

First audio38 ms

A silence timer

A rented stack, part by part

System of record

700 ms

80th percentile, banking, on-premises, real telephony

1,500 ms

The rented stack answers

Three times under the industry average.

Illustrative turn. Marked figures are measured.

0 ms400 ms800 ms1,200 ms1,600 msHumanNaturalNoticeableBrokenHuman to 400 ms, natural to 800 ms, noticeable to 1,200 ms, broken beyond
Where it wins05 / 08

Talking is the demo. Completing the transaction is the product.

A small model trained through the voice gap, on production audio and graded tool calls, beats larger general models at the moment of value.

Read the research

  • It completes the call

    97%Voicing action model

    One hundred calls each. Hollow is a failed call.

    Voicing action model

    97%

    Claude 4.5 Sonnet

    91%

    GPT-4o

    84%

    Llama 3.3 70B

    82%
    • 8Bparameters in the model that beat all three
    • 2,800voice-transcribed calls, graded schema-strict
    • 8 to 15 pointslost by general models on transcribed speech

    One in elevencalls fails at the moment of value

  • It hears it right

    5.8%Open ASR average

    • Voicing consensus router
    • Parakeet-TDT 0.6B v3
    • Whisper large-v3
    • Whisper large-v3-turbo

    Word error rate. Lower is better.

    LibriSpeech clean

    1.7%1.93%2.7%3.0%

    LibriSpeech other

    3.1%3.59%5.2%6.1%

    Open ASR average

    5.8%6.34%~7.4%~7.7%
    • 150 msto a final transcript
    • 12M+production minutes behind the measurement

    Three recognisers, one answer, lower error than any of them alone

  • It serves it cheaply

    $0.013a minuteagainst $0.14for a stack of APIs

    Each block is one second of audio, decoded.

    0.02Xon GPU

    Fifty seconds of audio, decoded every second

    0.05Xon CPU

    Quantised, on-premises, with no GPU required

    • $0.14for a stack of APIs

    Accuracy that only makes commercial sense when you own the models

The harness06 / 08

If you cannot measure the turn, you cannot ship the runtime.

Every turn is recorded, stage by stage. Every regression is caught overnight, before a caller meets it.

See AI Observability

  1. Entity curves

    Per slot, language and line

    32slots
  2. Stage latency

    Traced at p50, p95 and p99

    15dimensions
  3. Schema audits

    Every tool call is graded

    100%of tool calls
  4. Data flywheel

    Corrections feed the retrain

    Nightly
  5. Call quality

    Loss, echo and silence, live

    Realtime
  6. Semantic outcomes

    Completion, escalation, retry

    Per journey
One recorded turn

See AI Observability

The boundary07 / 08

It clears the review that stops cloud-only stacks.

In banking and healthcare the security review is the sales cycle. Policy is enforced in the runtime: one code path for consent, redaction, residency and audit.

On-premises, your cloud or Voicing Cloud.

Read the deployment options

Request the security package

  • SOC 2 Type 2
  • HIPAA
  • PCI DSS
  • ISO 27001:2022
  • ISO 42001:2023
  • GDPR and CCPA
It compounds08 / 08

One integration. Every expansion after it.

A runtime contract grows inside the account. Expansion was the customer’s decision in each of these, taken on the figures from the first weeks.

Read the case studies

  • ~10,000voice agents live
  • 99.98%agent uptime
  • 30+languages
  • Eightindustry-tuned LLMs
  • ~4 weeksto go-live
  • 60%lower handle time
  • >87%first-call resolution
  • >97%task completion
  • 4.3CSAT on AI calls
  • <0.5%hallucination rate

A tenth of inbound volume to a quarter, in six weeks. At go-live, After expansion, Intents live.

Equifax

2.5Xdaily calls handled

A tenth of inbound volume to a quarter, in six weeks

  • 10% to 25%of inbound volume
  • 7 to 12intents live
  • 6 weeksfrom go-live to expansion

No re-integration, no change window

An IVR replacement that now runs the outbound book. Inbound, Outbound. 6.9K, 32.9K, 48.2K, 44.2K, 78.5K, 125.9K Feb, Mar, Apr, May, Jun, Jul

UTI Mutual Fund

441Kminutes served

An IVR replacement that now runs the outbound book

  • 2% to 43%of minutes outbound
  • 125.9Kin the latest month
  • English and Hindi, around the clock

Outbound on the same runtime, no new integration

Interpreter queues to real-time multilingual service. Pilot, Expansion.

Great American Insurance Group

4 to 14languages

Interpreter queues to real-time multilingual service

  • 300+agents live
  • 4.2 of 5CSAT, post-call
  • 3.5xthe language coverage

Interpreter queue removed from the call path

A Spanish-first pilot to ten times the capacity. Pilot, Scaling, Production.

Telmex

10Xcapacity growth

A Spanish-first pilot to ten times the capacity

  • 60%lower handling time
  • 83%average resolution rate
  • +30%AI agents, month on month

English and Spanish live, same runtime

Read the case studies

Put your hardest call on the runtime.

A working session with an architect: your telephony, your systems of record, your security review. You leave with a deployment outline.

Book a Demo

Read the architecture

The runtime

Seven owned layers, running inside your perimeter