LLM Cold Starts: Build an Avatar Readiness Contract
An LLM cold start is not solved when a container becomes healthy. A real-time AI avatar is ready only when the complete conversation path—speech recognition, reasoning, knowledge, voice, rendering, media and authorised tools—can serve the first user within its agreed latency and quality envelope.
The practical answer is a readiness contract. Measure each start-up phase separately, keep traffic away until every mandatory dependency passes, run a synthetic first turn, and choose a deliberate warm-capacity policy. Extending a timeout merely gives a slow or incomplete system longer to fail.
This guide turns cold starts into an architecture and operations decision that bank, government and regulated-enterprise teams can review before approving a private, on-premise or cloud avatar service.
Cold start is four different clocks
Teams often report one “start-up time”, even though four events are involved:
- Infrastructure start: a node, virtual machine, container or process becomes available.
- Inference start: model artefacts are verified, weights enter memory, runtimes initialise and required compilation or graph capture completes.
- Session start: identity, policy, knowledge, voice, rendering and media dependencies attach to the conversation.
- First-turn start: the user’s first representative utterance produces the first valid audio and rendered response.
A service can pass the first clock and fail the fourth. The process may be alive while weights are still loading. The LLM may answer while text-to-speech is unavailable. The avatar renderer may publish a frame while private retrieval or an identity service is still unreachable.
Kubernetes makes a useful distinction among startup, liveness and readiness probes: starting successfully, remaining alive and being eligible for traffic are different decisions. An avatar readiness contract applies that distinction to the complete user journey rather than one container.
Why the first session is different
Warm-turn latency measures a running service. A cold session may additionally require:
- GPU allocation and runtime initialisation;
- container-image and model-weight transfer;
- CPU-to-GPU loading, memory allocation and kernel preparation;
- speech, embedding, LLM, voice and avatar-model initialisation;
- persona, policy, vocabulary and retrieval-index loading;
- browser permissions, WebRTC signalling and media negotiation;
- private identity, knowledge and action-service connections; and
- the first request’s tokenisation, prefill, audio synthesis and frame generation.
The critical path depends on architecture. NVIDIA describes model loading as storage-to-CPU and CPU-to-GPU movement, while its Run:ai Model Streamer guidance shows why storage location, format and transfer parallelism matter. A September 2026 AWS EKS study similarly found materially different loading times across configurations. These are implementation examples, not universal benchmarks or Yepic requirements.
For an avatar, faster weight transfer can still leave speech, voice, rendering or media cold. Optimise only after tracing the whole path.
Build a seven-gate readiness contract
1. Compute and runtime gate
Confirm the approved node, GPU, driver, container image and runtime versions. Verify that the expected accelerator is visible and that memory allocation has not silently fallen back to an unintended device. “Process running” is insufficient if the service cannot reserve the resources required for its supported workload.
2. Artefact and integrity gate
Verify every required model, tokenizer, voice, avatar asset, plugin and configuration against the approved release manifest. Record its source, version, hash, licence and target device. A missing optional voice may justify a degraded state; a missing policy model or primary renderer may make the instance unavailable.
3. Model-warmth gate
Load the required weights and execute representative warm-up inputs. A model can be resident yet still pay first-request costs for memory allocation, kernel selection, compilation or graph capture. NVIDIA Triton exposes model warm-up requests before marking a model ready. The specific mechanism will differ by runtime, but the principle is portable: readiness should follow a realistic inference path, not a successful file read.
4. Conversation-chain gate
Prove the linked speech-to-avatar path: audio ingestion, speech recognition, retrieval or context assembly, model generation, text-to-speech, lip synchronisation, rendering and stream publication. Capture time to first transcript, token, audio and rendered frame. The slowest mandatory stage determines whether the conversation is ready.
This extends Yepic’s GPU-scheduling guidance for real-time conversations. Placement and model residency decide where work can run; a readiness contract proves that the selected path can actually serve a user.
5. Knowledge and policy gate
Load the approved persona, system policy, language resources and active knowledge generation. Check that permissions, source versions and effective dates are available. If a private retrieval index is rebuilding, define whether the avatar remains unavailable, starts with a limited public corpus or routes to a human.
6. Identity and action gate
Confirm that workload identities, secrets and authorisation services are valid and that required private APIs are reachable. A citizen-service or banking avatar must not accept a session as transaction-capable merely because it can speak. Publish its actual mode—informational, authenticated, transactional or degraded—so the interface and policy layer can enforce it.
7. Media and synthetic-turn gate
Establish the real media path and run a harmless synthetic exchange through it. Validate captions, microphone state, first audio, first frame and expected disclosure. A server-side model probe cannot reveal a browser codec, TURN route or rendering failure.
Use Yepic’s synthetic conversation-probe method for this final outside-in check. Keep the probe identity controlled and prevent it from executing consequential production actions.
Publish readiness as a state, not a checkbox
A useful state model has at least six outcomes:
- Unavailable: the instance cannot begin its approved workload.
- Starting: infrastructure and artefacts are being prepared.
- Warming: models are resident but representative inference or media checks are incomplete.
- Ready: all mandatory gates passed for a named capability set.
- Degraded: a documented subset is available, such as text and captions without voice.
- Draining: no new sessions are admitted while existing work completes or moves.
Readiness should identify the capability, language, model set, avatar, policy generation and action mode it covers. One green status cannot prove every supported combination. Arabic speech, a specialist voice or a larger reasoning model may have different loading and memory requirements from the default path.
Record the start trigger, phase timings, artefact versions, cache state, GPU identity, probe inputs, outcome and expiry. This creates evidence for capacity reviews and incidents without requiring routine storage of customer conversations.
Choose the right warm-capacity policy
There are four common operating patterns. None is universally best.
Pinned capacity
Keep the complete model set and conversation chain resident. This gives the most predictable first session but consumes GPU memory and power while idle. It suits continuous or high-consequence services where start-up delay is unacceptable.
Warm pool
Maintain a bounded number of ready workers and replenish the pool as sessions begin. Size it from arrival bursts, start-up time and the minimum acceptable reserve—not average utilisation alone. A shared pool must preserve tenant, language and persona isolation.
Predictive pre-warming
Warm the required path before a branch opens, an appointment starts or an expected event peak arrives. This can reduce idle cost, but forecast error becomes an operational risk. Test what happens when demand arrives early or the warm-up job fails.
On-demand or scale-to-zero
Start capacity only when requested. This can be economical for infrequent workloads, but the first user carries the complete start-up path unless the interface provides a truthful waiting or fallback experience. It is unsuitable where the worst-case start time breaches the service objective.
Model the decision alongside energy per successful avatar conversation. Keeping everything warm may improve latency while increasing idle power; aggressive scale-to-zero may lower idle use while increasing abandonment and failed starts.
Test cold conditions deliberately
A “cold-start test” performed twice on the same node may only measure a warm page cache. Define the starting condition precisely:
- container image absent or present;
- model weights remote, on local disk or in page cache;
- GPU context absent or retained;
- compiled kernels and graphs absent or reusable;
- retrieval, voice and renderer assets unloaded;
- new node, restarted service or evicted model;
- new media route versus an established connection; and
- single start versus a fleet warming simultaneously.
Measure p50, p95 and p99 for each phase and for click-to-ready, first valid audio and first rendered response. Track readiness false positives, failed warm-ups, pool exhaustion, model evictions, idle GPU time and recovery after a node loss. Yepic’s end-to-end avatar load-testing framework explains why this must be repeated under representative concurrency rather than on an otherwise idle system.
Deployment location changes the trade-off
A customer-hosted deployment can keep model artefacts, warm pools, policies and readiness evidence inside the customer environment and on customer-controlled GPUs, subject to a properly scoped implementation. Local or high-throughput storage can make the loading path more controllable. The customer and implementation team also inherit capacity reservation, storage design, driver compatibility, power, patching and failed-node recovery.
Private cloud and sovereign cloud may combine controlled residency with managed orchestration. Public cloud can offer elastic capacity and mature managed services, but scale-to-zero and remote artefact paths still need tail-latency tests. On-premise is not automatically warmer, faster or more reliable; predictability depends on the whole architecture and operating model.
The Abu Dhabi Aviation and Oracle integration involved development and production environments, APIs, iframe delivery, real-time streaming, captions, microphone behaviour, WebRTC testing, browser remediation, cybersecurity support and ongoing maintenance. It is relevant evidence for testing a complete enterprise path, not a claim of a completed customer-hosted cold-start programme.
Twelve questions for production approval
- Which four start-up clocks are measured separately?
- What exact capability set does “ready” certify?
- Which models, voices, languages and renderers must be resident?
- Which integrity and licence checks occur before loading?
- Does warm-up execute representative inference on the target GPU?
- How are knowledge, policy, identity and action dependencies tested?
- Does an outside-in probe verify the browser and media path?
- What degraded modes are permitted, disclosed and blocked from actions?
- How is warm-pool reserve sized for bursts and node failure?
- Which cold conditions and cache states are included in acceptance tests?
- What prevents unready instances receiving new sessions?
- Which change to models, drivers, artefacts or infrastructure triggers requalification?
Make the first user part of the service objective
Start with one avatar, one language, one model route and one integration path. Reset it to defined cold states, measure every phase, publish a capability-specific readiness result and prove the first session under load. Then decide which capacity must remain pinned, pooled, predicted or on demand.
Yepic can scope real-time avatar deployments across cloud, private-cloud, sovereign and customer-hosted environments, including commercial-GPU optimisation and complete conversation-path testing. The goal is not to claim that a service starts instantly. It is to ensure that traffic begins only when the experience users and operators were promised is genuinely ready.