Table of contents

Continuous Batching for LLMs: Protect Avatar Latency

2026-10-09T00:00:00.000Z
October 9, 2026
Video Agents
An enterprise AI avatar beside continuously refreshed GPU batching lanes carrying several conversation streams.

Continuous batching can increase LLM GPU utilisation without making every live AI-avatar conversation wait for the longest response in a fixed batch. But it should be governed as a latency and fairness policy, not enabled as a generic throughput switch. The runtime must decide, at every model iteration, which requests enter, continue, pause or leave the batch. Those decisions affect queue time, time to first token, speech cadence, KV-cache pressure and whether one tenant can degrade another.

For banks, governments and regulated enterprises, the right question is not “How large can the batch be?” It is: “Which live conversations may share an iteration, under what limits, and what proves the scheduler protects the user experience?” This guide provides a practical contract for answering that question across customer-hosted, private-cloud, sovereign-cloud and public-cloud deployments.

What continuous batching means

Traditional static batching groups several requests and runs them together until every request finishes. Because generated responses have different lengths, short requests can leave GPU capacity idle while the batch waits for the longest sequence.

Continuous batching—also called in-flight or iteration-level batching—rebuilds the active batch as generation progresses. Finished requests leave immediately; eligible waiting requests can join while other sequences continue decoding. The batch therefore changes from one model iteration to the next.

This is not the same as an application’s bulk-video mode, where many independent videos are submitted as a job. It is also different from autoscaling, which changes the number of workers, and from request routing, which chooses a worker or model before scheduling occurs.

NVIDIA’s inference guidance explains how in-flight batching evicts completed sequences and admits new work instead of holding a fixed batch open. Current TensorRT-LLM documentation describes the same pattern as a way to interleave context and generation work while using packed inputs rather than padding every request to the longest sequence.

Why avatar conversations make batching harder

A text service may optimise tokens per second across the fleet. A real-time avatar also needs a predictable speaking rhythm. Its user-visible path includes speech recognition, context assembly, LLM prefill, token generation, voice synthesis, avatar rendering and media delivery.

Continuous batching can help by keeping the GPU productive, but three forms of interference matter:

  • Prefill interference: a long system prompt, conversation history or retrieved context can take compute away from sessions already decoding.
  • Memory interference: admitting more sequences consumes KV-cache capacity and may trigger pre-emption, eviction or recomputation.
  • queue interference: throughput-oriented packing can allow high-volume or long-context traffic to delay a time-sensitive conversation.

A higher aggregate throughput figure is not a success if a speaking avatar develops audible gaps. Measure time to first audio and time between playable speech chunks alongside LLM metrics. Yepic’s guide to end-to-end avatar load testing explains how to test the complete conversation rather than one model endpoint.

Separate four scheduling decisions

1. Admission

Admission decides whether a request may enter the worker at all. It should account for prompt length, expected generation, model route, active sequences, KV-cache headroom and service class. An unbounded queue converts overload into hidden waiting time. A bounded queue, explicit retry response or approved fallback is easier to operate.

Current vLLM scheduler documentation exposes active-sequence, queued-request and queued-token limits. Its queued-token limit is explicitly framed as a time-to-first-token quality-of-service mechanism. These settings are implementation examples, not universal Yepic defaults.

2. Per-iteration selection

The scheduler chooses which active sequences receive work in the next iteration. A throughput-first policy may pack as much work as possible. A latency-first policy may protect decode work before using spare token budget for new prefills. The appropriate ordering depends on the conversational objective and the cost of pausing an admitted session.

3. Prefill treatment

A prompt can be processed as one block or divided into chunks. Chunked prefill can mix portions of a long prompt with ongoing decode, reducing the size of each interruption. Smaller chunks may protect inter-token latency but extend time to first token; larger chunks can do the opposite. There is no universal token budget that fits every model, GPU and workload.

4. Completion and cancellation

A completed, timed-out or disconnected request should leave the batch and release its memory promptly. Cancellation must propagate through the gateway, scheduler and model worker. Otherwise an abandoned browser session continues consuming tokens and KV state while a genuine user waits.

Choose an explicit service policy

One worker configuration should not silently serve every workload. Define named policies with measurable intent:

Conversation-first

Prioritise decode for active, user-facing sessions; constrain long prefills; reserve KV capacity; and keep queues short. This sacrifices some peak throughput to protect speech continuity. It suits assisted service, role-play and public-facing interactions where gaps are immediately visible.

Balanced

Mix decode with chunked prefill inside a bounded token budget. This may suit general enterprise assistants with moderate latency objectives and varied prompt lengths. It requires careful testing because one setting controls two competing clocks.

Throughput-first

Pack more active work and accept wider latency variance. This can be sensible for offline preparation, summaries or non-interactive tasks, but should not inherit the service objective of a live avatar merely because both use the same model.

Reserved critical lane

Keep a separate worker pool or protected share for critical conversations rather than relying solely on request priority inside one saturated batch. This costs capacity but creates a clearer isolation boundary for high-consequence services.

This extends Yepic’s GPU-scheduling framework. GPU scheduling decides which workloads receive accelerator time; the continuous-batching policy defines how LLM sequences share each iteration inside that allocation.

Build a ten-field batching contract

For every model route and service class, record:

  1. Scheduling unit: requests, prompt tokens, generated tokens or weighted work.
  2. Admission limit: maximum active requests, queued requests and queued prompt tokens.
  3. Iteration budget: maximum tokens or sequences processed together.
  4. Prefill policy: whole, chunked or separately served, including the allowed chunk range.
  5. Priority: how service classes are ordered and how starvation is prevented.
  6. Memory reserve: KV-cache capacity protected for admitted conversations.
  7. Pre-emption: whether work may pause, what state remains and the recovery cost.
  8. Cancellation: the deadline for stopping generation and releasing state.
  9. Isolation: which tenants, data classes and workloads may share a worker.
  10. Evidence: metrics and tests that justify the production setting.

Version the contract with the model, tokenizer, runtime, precision, GPU type and parallelism. A batch limit validated for one combination is not automatically valid after quantisation, a longer context window or a driver change.

Protect fairness and tenant boundaries

Continuous batching mixes computation, not authority. Every request must retain its authenticated tenant, user, policy and session identity. Shared execution must not create shared conversation state, prompt visibility or reusable cache rights.

Fairness also needs a definition. Equal request counts can be unfair when one request contains 20 times more prompt tokens. Equal token budgets can be unfair when a short critical response sits behind many low-priority turns. Choose the unit that reflects service impact and test noisy-neighbour behaviour deliberately.

Use separate pools where the threat model or service objective cannot tolerate shared process, memory or timing effects. Yepic’s guide to LLM KV-cache isolation and budgeting covers the state that continuous batching places under pressure.

Control overload before the batch collapses

More requests do not always produce more useful work. As active sequences rise, KV memory, iteration time and queue delay can reach a capacity knee. Past that point, admitting another request may make every conversation slower.

Define an ordered overload response:

  1. stop admitting optional or low-priority work;
  2. return a bounded retry signal rather than growing an invisible queue;
  3. route eligible work to another compatible pool;
  4. apply approved context reduction or a smaller model route;
  5. degrade the avatar interaction clearly; and
  6. preserve already admitted critical conversations.

Do not wait for an out-of-memory event. Yepic’s rate-limiting guide for real-time avatars connects external admission controls to measured internal capacity.

Measure the scheduler, not just the model

Collect metrics by model route, service class, tenant group, prompt-length band and output-length band:

  • queued, active and rejected requests;
  • queued prompt tokens and queue duration;
  • scheduled prefill and decode tokens per iteration;
  • active batch size and unused token budget;
  • time to first token and inter-token latency distributions;
  • time to first audio and speech-chunk gaps;
  • KV-cache utilisation, pre-emptions and recomputed tokens;
  • cancellation-to-release time; and
  • useful responses per GPU-hour.

Current vLLM metrics distinguish queue, prefill, decode and inter-token intervals. TensorRT-LLM exposes scheduled, context, generation and paused-request counts for in-flight batching. The names vary by runtime; the operational questions do not.

Run six production tests

  1. Mixed lengths: combine short and long prompts with short and long answers; verify one quadrant does not dominate the others.
  2. Prefill burst: introduce many retrieval-heavy turns while established sessions are speaking; measure tail inter-token and first-audio latency.
  3. Noisy tenant: saturate one tenant or service class; prove protected workloads maintain their envelope.
  4. Memory pressure: approach the KV limit; verify admission, pre-emption and recovery match the contract.
  5. Cancellation storm: disconnect many clients; confirm work and memory are released within the approved deadline.
  6. Configuration change: alter the iteration budget or chunk size through a controlled release; compare quality, latency and throughput before promotion.

Compare continuous batching against a static or lower-concurrency baseline under the same traffic. If aggregate throughput rises but conversational tail latency breaches its objective, the setting has not passed.

Deployment location changes the operating model

A customer-hosted deployment can keep model execution, scheduler telemetry and active state inside the customer’s environment and on customer-controlled GPUs, subject to a properly scoped implementation. It also makes the customer and implementation team responsible for queue policy, memory reserve, runtime upgrades, capacity expansion and incident response.

Private or sovereign cloud may offer controlled residency with managed scaling and high-speed infrastructure. Public cloud can make elastic pools easier to obtain. None of these options automatically chooses a fair batching policy. The best design depends on workload distribution, service criticality, hardware topology and the operating team.

At GAIN Saudi Arabia, Yepic operated a five-metre real-time agent targeting sub-one-second responses alongside the holographic Farrah across three summit days. The GAIN deployment demonstrates why live-avatar performance must be validated as an audience experience; it is not presented as a continuous-batching or customer-hosted implementation.

Twelve questions for architecture approval

  1. Which throughput or latency problem is continuous batching intended to solve?
  2. What are the real prompt and response length distributions?
  3. Which service classes may share one worker?
  4. What bounds active requests, queued requests and queued tokens?
  5. How are decode and prefill work prioritised?
  6. Is prefill chunked, and how was the chunk budget validated?
  7. How much KV capacity is reserved for admitted sessions?
  8. Can a request be paused or evicted, and what does the user experience?
  9. How is starvation prevented across tenants and request sizes?
  10. How quickly do cancellation and timeout release compute and state?
  11. Which metrics reveal the capacity knee before failure?
  12. Did the final test measure speech and rendering, not only tokens?

Optimise the conversation, not the batch

Start with representative traffic and a simple baseline. Define service classes, bound admission, reserve active-session memory, choose a prefill policy and measure the complete avatar response. Increase batch density only while tail latency, fairness and cancellation remain inside the approved envelope.

Yepic can scope real-time avatar deployments across cloud, private-cloud, sovereign and customer-hosted environments, including commercial-GPU optimisation and whole-path performance testing. The goal is not the fullest possible GPU. It is the highest safe utilisation that still feels like a responsive, reliable conversation.