Skip to content

Platform · Runtime and models

The models are ours.
The compute is yours.

Speech, language and voice models ship with the runtime and run on CPU inside your boundary. No GPU, and no third-party endpoint in the call path.

  1. The runtimeOne deployment, four capabilities
  2. ModelsSpeech, language, voice
  3. InferenceCPU-only, your hardware
  4. IntegrationTelephony, CRM, records
The runtime01 / 05

One runtime, every surface.

Voice AI Agents, Real-Time Translation, Knowledge Mesh and Agent Assist ship together on one deployment, and Voice Identity Transformation and AI Observability run on the same runtime.

  • Voice AI Agents

    Inbound and outbound calls, long context, and a handoff to a person mid-sentence.

    Ships together
  • Real-Time Translation

    Simultaneous translation on a live call, with accent normalisation on both sides.

    Ships together
  • Knowledge Mesh

    Retrieval and governance: comprehensive in, salient out, permission-filtered and cited.

    Ships together
  • Agent Assist

    The same cited answer on the human agent's desktop while the call is still live.

    Ships together
Voicing runtime
Your infrastructure
Models02 / 05

The models ship with the runtime.

Speech, language and voice models are built and versioned here, deployed with the runtime, and moved on your change schedule rather than on a vendor's.

  • Speech

    Recognition and synthesis on the call itself, with no audio sent anywhere else for either.

    In the deployment
  • Language

    The reasoning that reads the caller's intent, chooses the tool and writes the reply.

    In the deployment
  • Voice

    The spoken output, and the identity work that holds every agent to one brand standard.

    In the deployment

Models, at work

One turn of a call, and the three models it passes through.

  1. The caller speaks, and Speech hears it
  2. Language reads intent and picks the tool
  3. Voice speaks the reply in one brand voice
  4. The audio leaves on the trunk it came in on
One turn of a call through three modelsThe caller stands outside your deployment on your trunk. The turn enters the deployment, passes through the speech model, the language model and the voice model in that order, and leaves on the same trunk it arrived on. Nothing in the loop leaves the deployment.CallerYour trunkYour deploymentSpeechLanguageVoice1234

What the drawing marks

  • Your trunk: The carrier your calls use today
  • On the call itself: No audio goes anywhere else
  • Versioned with the runtime: Moves on your change schedule
  • Your deployment: Every model loads from inside it
Speech

Voicing consensus router 1.7% LS-clean, 3.1% LS-other, 5.8% Open ASR avg; 150 ms to a final transcript; real-time factor 0.05X on CPUEnglish word error rate, quantised on-premise, no GPU

Inference03 / 05

Inference runs on hardware you own.

The runtime is optimised for CPU, so a deployment is sized on servers an enterprise already runs instead of on a GPU estate it has to buy first.

  • CPU-optimised

    Sized on ordinary server cores, so a deployment does not wait on specialised hardware.

    No GPU
  • Inside the boundary

    Every model in the call path is loaded from your own deployment, never from an endpoint.

    No egress
  • On your schedule

    Model versions move when your change board says they move, and roll back the same way.

    Versioned

Inference, under the hood

Five techniques that take the GPU out of the call path.

Real-time factor

0.05X

on CPU, quantised, no GPU required

  • Quantised weightsA smooth curve of trained weights runs into a quantiser that holds a few levels, and comes out on the right as the same curve stepped onto those levels. An accuracy check reads the stepped weights and reports back to the quantiser.Trained weightsQuantiserServed weightsAccuracy check1234
    Quantised weights
    • Trained weights
    • Quantiser
    • Served weights
    • Accuracy check

    Quantised weights

    The speech and action models are quantised, and accuracy is verified after every pass.

  • Custom CPU kernelsA streaming decoder takes audio in frames, a tuned kernel packs it into lanes, and sixteen cores run it with eight of them at work. The cores write a cache and the decoder reads the cache back instead of computing the same thing twice.Streaming decoderTuned kernelCPU coresCache1234
    Custom CPU kernels
    • Streaming decoder
    • Tuned kernel
    • CPU cores
    • Cache

    Custom CPU kernels

    Streaming decoders and cache reuse, on kernels tuned for the instruction set you run.

  • A distilled action modelA general model stands above the path as a dashed stack of nine layers with nothing joined to it. The small action model it was distilled into stands on the path in three layers, takes the transcript, and writes a tool call with every slot filled.General modelTranscriptAction modelTool call1234
    A distilled action model
    • General model
    • Transcript
    • Action model
    • Tool call

    A distilled action model

    A small fine-tune fills the slots and makes the tool calls a general model is oversized for.

  • A deterministic schedulerAudio fills a ring buffer, the scheduler beats at one interval, and four pinned cores run their work in slots that all start on the same ticks. Every block in every lane begins on a tick and none of them runs into the next one.Audio ringSchedulerPinned cores123
    A deterministic scheduler
    • Audio ring
    • Scheduler
    • Pinned cores

    A deterministic scheduler

    Pinned tenancy and a zero-copy audio ring hold the jitter inside the budget of the turn.

  • No GPU in the pathA rack of empty slots runs across the top with nothing joined to it. The whole of the work is in the band below: the call arrives, twenty cells of your own fleet run it with a column left spare, and the reply leaves.GPUCall inYour CPU fleetReply out1234
    No GPU in the path
    • GPU
    • Call in
    • Your CPU fleet
    • Reply out

    No GPU in the path

    Capacity is the CPU fleet you already own, so growth is a formality and not a purchase.

Integration04 / 05

It joins the systems the call already touches.

Telephony, CRM and the systems of record stay where they are. The runtime reaches them from inside your network, with credentials your own team issues and revokes.

  • Telephony

    The numbers, carriers and contact-centre platform your calls already route through.

    Your carrier
  • Systems of record

    CRM, core banking, policy and scheduling systems, read and written under your permissions.

    Your credentials
  • People

    A handoff to a person mid-call, with the transcript and the state of the call attached.

    Your team

Integration, by name

The systems a call already touches, reached from inside your network.

Voicing runtime, inside your network

Telephony

  • Genesys

  • NICE CXone

  • Avaya

Your carriers and the contact-centre platform you run.

Systems of record

  • Salesforce

  • Epic EHR

  • Oracle BRM

Read and written under permissions your team controls.

People

  • Five9

  • WhatsApp Business

  • ServiceNow ITSM

A mid-call handoff, with transcript and state attached.

And it does not end here.

System marks belong to their owners.

  • A documented API

    Every action the agent takes, as a call.

  • An MCP server

    Your tools, over the Model Context Protocol.

  • Webhooks

    Events pushed to whatever listens.

Closing05 / 05

Bring one call type. Leave with an architecture.

The runtime’s layers stacked as one object, turning slowly.
Platform

Six surfaces on one runtime

A working session with an engineer who has deployed inside a bank’s perimeter. We map your telephony, data boundary and handoff rules, and tell you what we would not automate.