Architecture
Pilots are easy. Production is an architecture test.
Concurrency, longer context and real integrations are what break a voice agent between the pilot and the floor. Voicing answers each of them inside one runtime you can point at.
- One runtime, not four vendors
- Your perimeter
- Nine layers, scheduled
One turn, one process, nothing in between.
A stitched stack pays a network hop and a queue at every step of a turn, and the budget is gone before the model answers. Here a whole turn happens inside one process on one machine.
Channel
The trunk ends on your own PBX, so no call leaves your network.
Orchestration
Policy runs before the model, so consent and redaction cannot be skipped.
Application
Agents read the systems of record in place, so nothing is copied out.
Model
Hearing, reasoning and speech are one process, so a turn costs one hop.
Infrastructure
Ordinary CPU servers, so capacity is racks your team already buys.
Where the audio sits has an exact deployment boundary.
Every review asks the same question: where does the audio live, and who can reach it. The deployment boundary answers it, and the same image runs on all three.
Where the audio sits has an exact deployment boundary.
Every review asks the same question: where does the audio live, and who can reach it. The deployment boundary answers it, and the same image runs on all three.
On-premises
Air-gapped where the review requires it, CPU only, zero egress.
Telephony and mediaOrchestration and policyAgents and retrievalSpeech and languageCPU computeConversation storeTrunkNo egress portBuilt perimeter. Plan of the runtime standing inside the On-premises perimeter, with six rooms off one corridor and the trunk entering through the perimeter and the building’s own door. Nothing. The trunk terminates on your own PBX and stays there. - Leaves the wall
- Nothing. The trunk terminates on your own PBX and stays there.
- Stays inside
- Audio, transcripts, weights, traces and the conversation store.
Your cloud
Inside your own subscription, behind your own IAM boundary.
Telephony and mediaOrchestration and policyAgents and retrievalSpeech and languageCPU computeConversation storeTrunkYour control planeYour account boundary, one port. Plan of the runtime standing inside the Your cloud perimeter, with six rooms off one corridor and the trunk entering through the perimeter and the building’s own door. Your own control plane, billed and logged on your own account. - Leaves the wall
- Your own control plane, billed and logged on your own account.
- Stays inside
- Audio, transcripts, weights, traces and the conversation store.
Voicing Cloud
Operated by us, on the same image, in a region you pin.
Telephony and mediaOrchestration and policyAgents and retrievalSpeech and languageCPU computeConversation storeTrunkOperations, in regionNamed region, one port. Plan of the runtime standing inside the Voicing Cloud perimeter, with six rooms off one corridor and the trunk entering through the perimeter and the building’s own door. Operations traffic, to the region written into the agreement. - Leaves the wall
- Operations traffic, to the region written into the agreement.
- Stays inside
- Processing, which does not leave the region you named.
Four ways an agent fails, each with a control.
The four places an agent programme actually fails, with the consequence of each failure made clear.
- Decode budgetCaller speechSpeech modelSpoken replyDecode budgetCaller speechSpoken replySpeech model
Speech as spoken
Trains on disfluent, interrupted and overlapping speech, then answers each turn inside a fixed decode budget.
A caller who changes their mind is understood.
- Filter firstEntitled sourcesPermission filterCited answerFilter firstEntitled sourcesCited answerPermission filter
Permissions first
Retrieval checks this caller’s entitlements before generation, then returns each answer with its source.
The agent cannot say what this caller may not hear.
- Schema-strictTool callSchema gateCore recordSchema-strictSchema gateTool callCore record
Schema before action
Each tool call must match the core system’s schema before it is sent. Failed arguments are retried or handed on.
A wrong field never reaches your core system.
- Riser, all floorsFloor tracesTrace riserTrace storeRiser, all floorsFloor tracesTrace riserTrace store
One trace, every floor
Turns, tools, retrievals, handoffs and model versions reach a store in real time, so one trace holds each call.
A bad call is traced to its own cause, not guessed at.
A model changes. Your call flow should not.
Each layer carries two lines, with Voicing first and a stitched stack second. Read them together and the difference is what can change without you.
A stitched stack is the usual alternative: a hosted voice API for the calls, a speech vendor for the transcript, a frontier model provider for the reasoning and a second synthesis vendor for the reply.
Solid, with Voicing.
Dashed, with a stitched stack.
One runtime.
Nine layers end here, on infrastructure you control.
01Telephony ingress
With Voicing: YouEnds on your own PBX
With a stitched stack: The vendorA rented number, not yours
02Speech recognition
With Voicing: VoicingTuned models, on your disk
With a stitched stack: The vendorEvery second is metered
03The language model
With Voicing: VoicingSelf-hosted, pinned by you
With a stitched stack: The vendorShared, no pin, no budget
04Voice synthesis
With Voicing: VoicingOur voices, same machine
With a stitched stack: The vendorSecond hop, second invoice
05Orchestration and tools
With Voicing: SharedYour policy, our runtime
With a stitched stack: SharedTwo vendors and your glue
06Knowledge and retrieval
With Voicing: YouIndexed inside the wall
With a stitched stack: SharedCopied to an outside index
07Conversation data
With Voicing: YouKept on your retention
With a stitched stack: The vendorKept on their retention
08Evaluation and traces
With Voicing: YouTraced into your own store
With a stitched stack: The vendorA sampled view per vendor
09Release cadence
With Voicing: SharedYou approve the window
With a stitched stack: The vendorChanged on their schedule
Your perimeter
Built for the floor at peak, not for the demo.
Concurrency, context and integration are where an agent programme loses its value. These are the figures a capacity review asks for.
What scale costs an agent programme
Up to 58%
The fall in Agent Value Multiple as a programme moves from pilot to full scale and adds concurrency, complexity and context, per Gartner.
700 ms
Measured across BFSI, on‑premises, real telephony, 80th percentile
97%
Function calls on 2,800 calls, schema‑strict grading, eight‑tool suite
Under 0.5%
Hallucination rate on in‑scope intents, measured in production
$0.017
Per minute on CPU, on‑premises, at 10 m minutes per month
The capacity model, CPU footprint and concurrency curve come to the architecture review on request.
Bring your architect. Bring the hard questions.

The runtime, on your infrastructure
A working session with an engineer who has deployed inside a bank’s perimeter. We map your telephony, data boundary and failure modes on this same topology, and tell you what we would not automate.