Disaggregated Inference: Should Avatars Split Prefill and Decode?
Disaggregated inference is worth considering when long or unpredictable prompts are disturbing the token cadence of live AI-avatar conversations. It is not a default architecture. Splitting LLM prefill and decode onto separate GPU pools can isolate two different kinds of work and let each scale independently. It also creates a state-transfer path, another scheduler and more failure modes. A private deployment should adopt the split only when measured gains across the complete conversation exceed those costs.
For banks, governments and regulated enterprises, the question is therefore not “Does disaggregation benchmark well?” It is: “Does this workload need separate prompt-processing and token-generation capacity, and can we govern the hand-off?” This guide provides a practical architecture decision and test plan.
What disaggregated inference changes
An autoregressive large language model serves a request in two main phases:
- Prefill processes the input tokens—the system instruction, conversation history, retrieved evidence, tool definitions and current user turn—and produces the initial key–value (KV) cache. This phase is typically compute-intensive.
- Decode uses that cache to generate output tokens one at a time. It repeatedly reads model weights and growing KV state, so memory bandwidth and active-session capacity matter.
In a colocated design, one worker pool performs both phases. Long prefills may interrupt or compete with ongoing decoding. In a disaggregated design, a router sends the request to a prefill worker, the resulting KV cache moves to a decode worker, and generation continues there.
NVIDIA Dynamo describes this three-step flow and the different resource shapes of the phases. vLLM’s current documentation emphasises a narrower goal: independently tuning time to first token and inter-token latency, while warning that its disaggregated-prefill feature is experimental and does not itself improve throughput. Those qualifications are important. The pattern is an option to test, not a universal performance upgrade.
Why an avatar workload is different
A text endpoint can expose time to first token (TTFT) and inter-token latency (ITL). A real-time avatar has a longer critical path:
- detect that the user has finished speaking;
- finalise speech recognition;
- assemble authorised instructions, history and evidence;
- prefill the LLM and produce its first useful tokens;
- synthesise playable speech;
- render and transmit a synchronised response; and
- continue without audible gaps, visual freezes or lip-sync drift.
A change that improves TTFT but adds transfer jitter may still worsen time to first audio. A design that protects ITL may be valuable even if aggregate throughput stays flat, because voice synthesis needs a steady token stream. Conversely, a short-prompt kiosk may gain nothing from the added hop.
Measure the whole path using the workload contract in Yepic’s guide to load testing real-time AI avatars. Do not optimise the LLM in isolation and assume the user will feel the same improvement.
Choose among three execution patterns
1. Colocated prefill and decode
One pool runs both phases, using continuous batching or chunked prefill where supported. This is the simplest baseline: no inter-worker KV transfer, fewer compatibility checks and a smaller operational surface. It is often the right answer for smaller models, short prompts, modest concurrency or sites without a fast transfer fabric.
Its weakness is interference. A large retrieved context or long conversation history can delay token production for sessions already speaking. The team must tune batching and scheduling so new prompt work does not create unacceptable tail latency.
2. Statically disaggregated pools
Dedicated prefill and decode workers have fixed capacities. This makes ownership, testing and failure containment comparatively clear. Each phase can use a different parallelism strategy or GPU shape where the runtime supports it. However, fixed allocation can strand resources when the prompt-to-output mix changes.
3. Dynamically balanced disaggregation
A planner changes the number or role of workers as demand changes. This can improve utilisation in large, variable fleets, but scaling is not instantaneous. Model load time, cache warming, role-transition rules and minimum reserve all enter the readiness contract. A planner must never remove decode capacity that admitted live conversations still require.
Start with the colocated baseline and a static split. Dynamic rebalancing should follow only after the traffic model and safe transition behaviour are proven.
Apply five eligibility tests before splitting
1. Workload asymmetry
Disaggregation is most plausible when prefill and decode pressures differ materially: long RAG contexts with concise answers, long conversation histories, bursty tool schemas, or high concurrent generation with predictable prompt sizes. Segment real traffic by input length, output length, language, model route, retrieval volume and session type. Averages conceal the long contexts that cause interference.
2. Demonstrated interference
Show that new prefills are actually damaging active decode. Compare TTFT, ITL and tail ITL under mixed load. Correlate spikes with prompt arrivals, batch composition and cache pressure. Yepic’s guide to GPU scheduling for live conversations explains why priority and admission control should be tested before adding topology.
3. A viable KV-transfer path
The initial cache can be large. Its size depends on model architecture, precision and input length. Transfer time, queueing, topology and destination allocation can erase the benefit of separating the phases. Test the actual GPU-to-GPU or node-to-node path, including congestion and failover—not a theoretical link rate.
NVIDIA’s TensorRT-LLM documentation explicitly notes this transfer overhead and says benefits are strongest where long inputs and moderate outputs create substantial interference. The current Dynamo guidance is equally candid: short prompts, low concurrency, small models or clusters without fast KV transfer often favour the simpler aggregated layout.
4. Independent capacity pressure
The prefill pool should scale against input-token work and TTFT objectives. The decode pool should scale against active sequences, output length, KV memory and token-cadence objectives. If both pools rise and fall together, fixed disaggregation may add little. Use the broader envelope in Yepic’s GPU capacity-planning guide, then reserve capacity for speech, voice and rendering if those workloads share the same estate.
5. Operational maturity
The team needs per-phase health, routing, back-pressure, version control and rollback. It must recognise a healthy prefill worker paired with an unhealthy decode worker, reject incompatible pairs and preserve live sessions during maintenance. If those controls are not yet reliable, improve the colocated service first.
Define a ten-field split-point contract
Record the following for every permitted prefill/decode route:
- Workload class: avatar, tenant, language, context band, response band and criticality.
- Model identity: weights, tokenizer, adapter, precision and runtime build.
- KV format: layout, block size, datatype and any compression.
- Worker roles: eligible prefill and decode pools plus their trust boundaries.
- Transfer: protocol, network path, bandwidth reserve, encryption and timeout.
- Admission: destination capacity reserved before work begins and the allowed queue.
- Routing: selection policy, affinity, cache locality and retry ownership.
- Failure: cancellation, orphaned-state cleanup and user-facing recovery.
- Fallback: when a request returns to a compatible colocated route.
- Evidence: metrics, versions and test results proving the route is safe to operate.
Compatibility must be mechanical. Matching model names are not enough if the tokenizer, precision, block size or KV layout differs. Current Dynamo deployment guidance requires matching model and relevant cache configuration between workers and warns that incompatible state can produce transfer errors or corrupted output. Pin and attest both sides as one release unit.
Design the hand-off as a governed data flow
KV state is not readable conversation text, but it is derived from instructions, retrieved documents and user input. Treat it as transient sensitive state. Map where it is created, buffered, transmitted and released. Bind every transfer to an authenticated request and destination. Prevent a client from selecting worker or cache identifiers directly.
The design should answer:
- Can prefill and decode workers from different tenants ever share a process, pool or transfer buffer?
- What request identifier prevents state from being attached to the wrong decode session?
- When is destination memory reserved, and how is it released after cancellation?
- Can state traverse another region, supplier or security zone?
- Which logs prove the hand-off without recording prompts or raw cache contents?
- How are partially transferred, timed-out and orphaned blocks cleared?
This extends, rather than replaces, the isolation and lifecycle controls in Yepic’s LLM KV-cache governance guide.
Test six failure conditions
- Long-context burst: flood prefill while established avatar sessions decode; confirm token cadence and first audio remain within the approved envelope.
- Transfer congestion: constrain the link and measure queue growth, timeout behaviour and whether the router admits more work than the path can carry.
- Decode exhaustion: fill destination KV capacity and confirm prefill does not waste compute or hold orphaned reservations indefinitely.
- Version mismatch: pair incompatible workers and prove the request is rejected before state is consumed.
- Worker loss: terminate each phase during transfer and during generation; verify cancellation, cleanup and the allowed retry boundary.
- Fallback: disable the split route and prove the approved colocated path can accept new sessions without inheriting stale state.
For every test, capture TTFT, median and tail ITL, time to first audio, time to first rendered response, speech gaps, transfer bytes and duration, queue time, retries, cancellations and GPU utilisation by phase. Test response quality too: a fast recovery that silently loses instructions or retrieved evidence is a failure.
Compare deployment models honestly
A customer-hosted implementation can keep workers, KV transfers and operating evidence within the customer’s environment and on customer-controlled GPUs. That can support residency and control requirements, subject to a properly scoped design. It also makes fabric design, driver compatibility, observability, capacity reserve, patching and fault recovery direct operational responsibilities.
Private or sovereign cloud may provide controlled tenancy with a managed high-speed fabric. Public cloud can make specialised GPU pools and elasticity easier to obtain. None of these labels guarantees low transfer latency or correct isolation. Benchmark the actual model, runtime, topology and traffic distribution.
The Abu Dhabi Aviation and Oracle enterprise-avatar integration included separate development and production environments, APIs and iframe integration, real-time streaming, captions, microphone behaviour, WebRTC testing, browser remediation, cybersecurity support and ongoing maintenance. It is evidence that the complete enterprise path must be tested; it is not presented as a customer-hosted or disaggregated-inference deployment.
Twelve questions for architecture review
- Which measured interference problem is the split intended to solve?
- What do the real input and output token distributions look like?
- How does the colocated baseline perform under the same workload?
- What are the TTFT, ITL, first-audio and first-render targets?
- How large is the KV transfer across credible context lengths?
- Can the approved fabric carry peak transfers with headroom?
- Which configuration fields must match across both phases?
- How are destination capacity and admission coordinated?
- What are the tenant, residency and encryption boundaries?
- What happens when either worker or the transfer path fails?
- Can the service fall back without mixing state or duplicating a turn?
- Which team owns routing, capacity, patching and evidence for the split?
Split only after the workload earns it
Begin with a representative colocated deployment. Prove that prefill interrupts live decode, build the split-point contract, test a static disaggregated route, and compare the complete avatar experience under equal load. Adopt the pattern only where lower tail latency or better resource separation outweighs KV-transfer and operational overhead.
Yepic can scope real-time avatar architectures across cloud, private-cloud, sovereign and customer-hosted environments, including commercial-GPU optimisation and whole-path performance validation. The objective is not the most distributed design. It is the simplest governed design that protects a natural, reliable conversation.