GPU Scheduling for AI Inference: Protect Real-Time Conversations
GPU scheduling for AI inference should protect the live conversation before it tries to maximise accelerator utilisation. A real-time avatar cannot treat every workload as equal. Speech recognition, language reasoning, voice generation and rendering sit on the user’s critical path; indexing, evaluation, video production and maintenance usually do not. The scheduler must reserve the right capacity, place each workload on compatible hardware and limit how long an interactive request may wait to join a batch.
This is the practical difference between owning GPUs and operating them as a dependable service. A bank or government department can have enough aggregate compute while users still experience long pauses because a background job occupied memory, a large prompt delayed smaller turns or a shared tenant consumed the queue.
The aim is not permanently idle hardware. It is a measured operating policy that trades utilisation against tail latency, isolation, fairness and recovery. That policy applies whether the estate is customer-hosted, private cloud, sovereign cloud or public cloud.
GPU scheduling is three decisions, not one
Architecture discussions often mix three separate control layers.
- Cluster placement: which node, accelerator type and failure zone should run a model service? Kubernetes, for example, exposes GPUs through device plugins and can place workloads using resource requests, labels and affinity rules. Its current GPU scheduling guidance is an infrastructure mechanism, not an end-to-end latency policy.
- GPU allocation: does a workload receive a whole device, a hardware-backed partition, a virtual GPU or time-shared access? This determines memory availability, interference, isolation and operational flexibility.
- Inference scheduling: once a model is resident, which request runs next, which requests are batched and how long may each wait? This is where first-response latency, streaming cadence and fairness are won or lost.
A fourth layer—session admission—decides whether new work should enter the system at all. Yepic’s guide to rate limiting for real-time avatars covers that boundary. Scheduling begins after work has been admitted, but it must receive the same trusted workload class and deadline.
Classify the avatar workload before sharing GPUs
Do not schedule an “avatar session” as one undifferentiated job. Divide the pipeline into service classes with different deadlines and state requirements.
- Perception: voice activity detection, speech recognition, language identification and optional vision. These tasks are continuous or frequent and influence when the system understands that a turn has ended.
- Reasoning and retrieval: embedding, reranking, language generation, policy checks and selected tool calls. Prompt length and output length can make individual jobs highly variable.
- Presentation: text-to-speech, facial animation, rendering and encoding. Smooth streamed output matters as much as the time to start.
- Operational probes: health checks and controlled synthetic conversations. These require enough capacity to detect a genuine failure without displacing real users.
- Background work: knowledge indexing, evaluation batches, model conversion, asynchronous video generation and analytics. These may be valuable, but their deadlines are usually measured in minutes or hours rather than milliseconds.
- Recovery reserve: capacity retained for node failure, maintenance, traffic movement or a degraded operating mode.
For each class, record its latency objective, maximum queue time, memory profile, statefulness, pre-emption behaviour, batching tolerance, permitted hardware and failure response. That becomes the scheduling contract.
Choose a GPU-sharing model deliberately
Exclusive device
One service instance receives a complete GPU. This is operationally simple and can reduce interference, but low traffic may leave expensive capacity unused. Exclusive allocation is useful when the model needs most of the memory, latency variation is unacceptable or the isolation case cannot be made with sharing.
Hardware-backed partition
Supported accelerators can be divided into isolated resource slices. NVIDIA’s current vGPU inference decision guide distinguishes time-sliced, MIG-backed and combined modes, each with different isolation and compatibility properties. These are examples rather than a Yepic requirement. Partition support, profile sizes, memory limits, licensing and observability must be verified for the proposed hardware and virtualisation stack.
Time sharing
Multiple workloads take turns on a GPU. This can improve utilisation when jobs are bursty, but it may introduce scheduling jitter and weaker performance isolation. A low-priority batch should not be allowed to create an unpredictable pause in live speech or rendering.
Pooled full GPUs
Replicas draw from a shared pool of whole devices. The platform routes each request to a compatible, ready worker. This preserves device-level simplicity while improving fleet utilisation, but model loading, cache locality, network paths and unequal queue depth become routing concerns.
Dedicated is not always safer, and sharing is not always cheaper. Benchmark the complete journey on the proposed partition, driver, runtime and model set. The existing GPU capacity-planning guide explains how to size the estate; scheduling determines how that capacity is allocated moment by moment.
Batch only where the latency budget permits
Batching combines work so the accelerator can process it more efficiently. The benefit is higher throughput; the cost can be additional waiting and interference between requests.
Four patterns matter:
- No batching: suitable for an especially sensitive path or a model that gains little from a batch. It may sacrifice utilisation for predictable response.
- Static batching: a fixed group runs together. It is straightforward for offline work but poorly matched to conversational arrivals and variable output lengths.
- Dynamic batching: the server holds requests briefly to form a useful batch. NVIDIA Triton’s current batcher documentation exposes maximum delay, queue size, priority and timeout controls, and recommends dynamic batching for stateless models.
- Continuous batching: common in language-model serving, the scheduler replaces completed sequences as generation proceeds. Hugging Face’s continuous-batching guidance describes it as rescheduling the batch at each generation step to improve GPU utilisation.
None of these is universally best. A wait that is invisible in offline generation may be obvious at the end of a spoken turn. Speech, language generation, voice and rendering may therefore need different batch policies even when they share an estate.
Sequence-aware workloads also need care. Requests belonging to one streaming session may require ordering and state affinity. Do not assume a stateless model queue can safely schedule a media or conversational sequence.
Make priority a trusted policy decision
A priority scheduler is only as defensible as the data assigning priority. The avatar should not infer urgency from accent, apparent emotion, disability, age or how forcefully a person speaks.
Use trusted context such as:
- authenticated service and transaction stage;
- declared channel, site or operational role;
- an accessibility mode selected by the user;
- a reserved service class defined in contract or policy;
- an incident mode activated by an authorised operator;
- the request’s remaining deadline and whether work has already started.
Priority must not become starvation. Set a maximum wait, reserve some progress for lower classes and measure fairness by tenant and service. An emergency queue that remains permanently active is simply an ungoverned fast lane.
Control model residency and cold starts
A scheduler cannot place work on a GPU if the required model, voice or renderer does not fit. Repeatedly evicting and loading weights can turn an efficient fleet into an unusable conversation service.
Define which components remain warm, which can load on demand and which combinations may coexist. Consider the language mix, avatar identities, voice assets, quantised variants, driver compatibility and recovery reserve. Yepic’s guide to qualifying model quantisation across the avatar stack explains why a smaller memory footprint still needs quality and compatibility testing.
Keep background conversion, indexing and evaluation away from the interactive pool unless pre-emption and memory recovery have been proved. “Low GPU utilisation” does not mean the device is safely available if loading another model will evict a critical resident set.
Measure the scheduler, not just the GPU
GPU utilisation alone cannot explain whether scheduling is working. Capture at least:
- placement wait before a worker starts;
- queue time by model, class, tenant and priority;
- batch size and time spent waiting to form it;
- time to first meaningful audio and rendered frame;
- streaming cadence, stalls and dropped media;
- GPU memory pressure, model loads, evictions and cache state;
- deadline misses, rejections and pre-emptions;
- throughput and successful conversation outcomes;
- fairness and starvation indicators across workload classes.
Triton, for example, documents an inference queue-time metric in its current metrics reference. Whatever serving stack is used, correlate queue evidence with the end-to-end measures in Yepic’s real-time avatar load-testing method. A scheduler can improve aggregate throughput while making the slowest conversations materially worse.
Test six scheduling failures before production
- Mixed-length burst: combine short questions, long prompts and variable spoken answers. Check that long generations do not monopolise the service.
- Background interference: start indexing or evaluation work, then introduce live sessions. Verify that interactive deadlines remain protected.
- Tenant contention: let one approved tenant generate a burst. Confirm quotas, reserved capacity and fairness for others.
- Model miss: request a language, voice or model that is not resident. Measure load time, routing and the user-facing fallback.
- Worker loss: remove a GPU worker during active streams. Confirm admission, recovery and capacity reserve without duplicating actions.
- Priority abuse: submit missing, invalid and forged priority values. Ensure the server derives authority from trusted policy rather than client input.
Run these scenarios at representative concurrency and record p50, p95 and p99 outcomes by workload class. There is no universal batch delay, queue depth, partition size or utilisation target.
Cloud and customer-hosted trade-offs
Public cloud can provide elastic pools and managed scheduling components, reducing the work required to operate drivers, clusters and capacity. It may introduce data-location, network-path, supplier-dependency or cost constraints.
Private cloud can combine dedicated tenancy with cloud operating patterns, but responsibility still depends on the contract. Customer-hosted deployment provides direct control over hardware, partitions, model placement and telemetry; it also makes the customer and implementation team responsible for spare capacity, upgrades, scheduler configuration and recovery.
Yepic can support private, customer-hosted and cloud architectures through a properly scoped implementation. The correct choice depends on data boundaries, latency, workload shape, internal operating capability and the evidence the organisation must retain—not a blanket assumption that one hosting model is superior.
Twelve questions for architecture and procurement
- Which avatar components require GPUs, and which share a critical latency path?
- Are cluster placement, GPU allocation and request scheduling documented separately?
- Which workload classes receive reserved capacity?
- How is priority assigned, authenticated and prevented from causing starvation?
- Which models are stateless, sequence-aware or session-affine?
- Where is batching enabled, and what maximum wait does it introduce?
- What isolation is provided by whole devices, partitions, virtualisation or time sharing?
- Which models, voices and renderers stay resident?
- What happens when the requested model is not warm?
- Can background work be pre-empted without corrupting state or retaining memory?
- How are scheduler fairness, queue time and end-to-end conversation quality tested?
- What capacity remains after a GPU, node or site failure?
Yepic’s GAIN Saudi Arabia deployment demonstrates the operational importance of real-time performance: the five-metre agent targeted sub-one-second responses and ran alongside a holographic avatar across three summit days. It is relevant evidence of demanding live-avatar delivery, not a claim of a completed customer-hosted GPU-scheduling implementation.
Schedule for the conversation, then optimise utilisation
A private GPU estate should not be judged by how constantly busy it looks. It should be judged by whether approved users receive responsive, fair and safe conversations during normal demand, bursts, maintenance and failure.
The practical design is layered: place services on compatible nodes, select an explicit sharing and isolation mode, keep the right models resident, schedule inference within a measured queue budget and reserve headroom for recovery. Once those controls protect the user journey, batching and sharing can improve economics without turning efficiency into delay.