When Does On-Premise AI Cost Less? A TCO Model for Real-Time Avatars
On-premise AI becomes economically rational when demand is sufficiently sustained and predictable to keep private capacity productive, and when the organisation values control, data boundaries or latency enough to justify owning the operating responsibility. It is not automatically cheaper because there is no cloud invoice. A customer-hosted avatar replaces some variable service costs with hardware, software, facilities, engineering, support, resilience and refresh costs.
The useful comparison is therefore not “GPU purchase price versus API price”. It is the fully loaded cost per successful conversation, at the service level the organisation actually needs, across a realistic planning period.
For a bank, government or regulated enterprise, that calculation should include the complete real-time chain: speech recognition, retrieval, the language model, text-to-speech, avatar rendering, video encoding, identity, networks, monitoring, security and human support. It should also price the capacity held in reserve for peaks and failure, even when that capacity is idle.
Start with the deployment decision, not the hardware quote
There are at least three credible operating models:
- Public cloud or managed API: low initial commitment, rapid experimentation and elastic capacity, with usage, concurrency, data-boundary and supplier terms set by the service.
- Private or sovereign cloud: dedicated or controlled cloud infrastructure with stronger isolation and contractual controls, but still a consumption-based or reserved-capacity model.
- Customer-hosted: the selected avatar components run inside the customer’s environment and on approved GPUs, with infrastructure and operational responsibilities divided between the customer and supplier.
Yepic’s deployment-model comparison covers the architectural differences. This article goes one level deeper: how to construct a cost model that finance, procurement and engineering can all challenge.
Define the required service first. The cheapest architecture at 99% availability may not be comparable with one designed for round-the-clock citizen services, two-site resilience or a restricted network. Cost comparisons are only meaningful when the options meet the same requirements for latency, languages, concurrency, recovery, privacy and support.
The seven cost layers of a private AI avatar
1. Compute and physical infrastructure
Include GPU servers, CPU, memory, local storage, networking, racks, power distribution and any hardware needed for media transport or resilience. Use an annualised cost across the organisation’s approved depreciation or lease period rather than charging the whole purchase to the first year.
Do not divide this figure by theoretical maximum sessions. Capacity planning must account for the tested conversational pipeline, target latency, peak concurrency, model residency, failover and headroom. The GPU capacity-planning guide explains why a universal “sessions per GPU” claim is not credible.
2. Facilities, energy and environment
Account for electricity consumed by servers and the facility overhead required to cool and operate them. Add data-centre space, connectivity, hardware monitoring, physical access controls and any charges allocated by an internal hosting team. Existing data-centre capacity is not free merely because another department already pays for it.
3. Software, models and licences
The runtime may include operating systems, orchestration, GPU software, speech components, language models, retrieval tools, databases, voice licences and security products. Some are perpetual, some are subscriptions and some remain usage-based even in an otherwise private architecture.
Map licence conditions to the intended environment. A model that can technically run locally may not be licensed for the proposed commercial deployment, language, voice or volume.
4. Engineering and integration
Private deployment usually requires more than installing an image. Include architecture, model optimisation, enterprise identity, knowledge integration, API work, browser and WebRTC testing, network remediation, accessibility, security review, acceptance testing and deployment automation.
Separate one-off implementation work from repeatable operational work. This prevents a pilot’s build cost being mistaken for its annual run cost, while ensuring integration effort does not disappear from the business case.
5. Operations, security and support
Price the people and processes needed for monitoring, incident response, vulnerability management, updates, certificates, secrets, backups, audits and user support. Include supplier support and internal teams, then remove overlaps to avoid double-counting.
The recent guide to on-premise AI model updates and rollback shows why the service needs a controlled release process across models, prompts, data, voices, rendering and infrastructure. Those activities are part of TCO, not exceptional work that can be ignored.
6. Resilience and unused reserve
A production service rarely operates every GPU at its safe maximum. Capacity may be held for traffic peaks, node failure, maintenance, model loading or disaster recovery. If the design needs two sites, a warm standby or spare hardware, include it.
This unused reserve is economically important. It is also operationally valuable. The mistake is not carrying reserve; it is pretending reserve has no cost.
7. Refresh, migration and exit
Hardware and software support windows do not necessarily align. Budget for component upgrades, model migration, revalidation, secure disposal and the possibility that a new runtime needs different memory or GPU capabilities.
NVIDIA’s current AI Enterprise lifecycle documentation illustrates the issue: application and infrastructure layers have separate branches and support periods. A defensible TCO model includes the cost of staying on a supported stack, not merely keeping yesterday’s installation switched on.
Cloud costs need the same discipline
A fair model must fully load the managed alternative too. Include:
- platform, seat, support or committed-spend charges;
- conversation, rendering, speech, language-model and tool-call usage;
- concurrency tiers and minimum commitments;
- storage, recordings, logging, network transfer and data egress;
- private connectivity, identity, security and compliance work;
- application engineering, observability and incident management;
- unused prepaid capacity or commitments;
- migration and exit costs.
Managed services reduce some operational work, but they do not remove the need to integrate, govern and support the customer experience. Conversely, a cloud service may bundle speech, rendering, upgrades and global capacity that would otherwise require several private components. Compare like with like.
Competitor pricing pages commonly express conversational avatars as included minutes, overage rates and concurrency tiers. Those are useful inputs for a hosted scenario, but they do not answer a regulated buyer’s TCO question. Enterprise support, integration, data controls and high-volume terms are frequently custom.
Use three calculations, not one headline number
1. Planning-period TCO
Choose a period long enough to capture implementation and refresh, then calculate:
Private TCO = implementation + annualised infrastructure + facilities + licences + operations + support + resilience + refresh and exit.
Managed TCO = implementation + platform commitments + variable usage + connectivity + operations + support + migration and exit.
Microsoft’s current Cloud Adoption Framework guidance on TCO estimation makes a useful general point: accurate estimates begin with a defined architecture, its dependencies and operational requirements. A price calculator without that architecture is precision theatre.
2. Unit cost
Divide fully loaded cost by a unit that represents delivered value. For a real-time avatar, useful technical and business units can include:
- cost per connected conversation minute;
- cost per completed conversation;
- cost per successful answer or human handover;
- cost per resolved case, completed application or training outcome;
- cost per language or service channel supported.
The FinOps Foundation’s unit-economics framework recommends connecting technology spend to the value it produces, not stopping at resource metrics such as cost per token. For an avatar, “cost per GPU hour” is an engineering signal. “Cost per successfully served customer” is closer to an investment decision.
3. Break-even range
A simple first-pass formula is:
Break-even usage = (private fixed cost − managed fixed cost) ÷ (managed variable cost per unit − private variable cost per unit).
This is a range, not a prophecy. Run it under low, expected and high demand, and vary utilisation, electricity, staffing, support, model size, hardware life and failure reserve. If the denominator is zero or negative, additional usage does not make the private option cheaper under those assumptions.
Then test the result against peak concurrency. A service can have modest annual minutes but a severe ten-minute peak. The private estate must be sized for the approved peak unless queues, degraded service or cloud bursting are acceptable.
Demand shape usually decides the economics
Steady, high utilisation favours owned capacity. A bank running thousands of predictable customer or employee interactions may be able to amortise hardware and operations across a large volume.
Bursty or uncertain demand favours elasticity. Events, campaigns, pilots and occasional training may leave private GPUs idle for long periods. Public or private cloud can convert that uncertainty into variable cost.
Sharp peaks favour a hybrid design. A stable baseline can run privately while approved overflow uses additional private-cloud capacity, provided the data boundary and component compatibility permit it.
Strict isolation changes the decision. An air-gapped or restricted environment may rule out public overflow. The business case must then include enough local reserve and the controlled movement of models, patches and support evidence.
Six factors that can justify on-premise before break-even
Lowest nominal cost is not the only legitimate objective. A private deployment may be selected because it:
- keeps sensitive speech, retrieved data or operational context inside an approved boundary;
- meets sovereignty or restricted-network requirements unavailable in a hosted service;
- reduces dependency on internet paths for latency-critical interactions;
- allows the customer to control maintenance windows, model versions and change approval;
- reuses an existing, governed GPU platform and operations team;
- reduces exposure to usage-price changes for a predictable, high-volume workload.
These benefits should be stated and valued explicitly. Do not hide them inside an artificially low cost estimate. Finance can then see whether the organisation is paying a justified control premium, achieving a genuine unit-cost advantage or both.
Evidence from a real enterprise integration
Yepic’s work with Abu Dhabi Aviation and Oracle demonstrates why avatar economics extend beyond inference. The delivery included dedicated development and production environments, API and iframe integration, streaming, captions, microphone behaviour, WebRTC, browsers, corporate-network testing, cybersecurity support and ongoing production support. Read the enterprise integration case study.
That project is evidence of the complete service work that a cost model should recognise. It is not being presented as a completed on-premise deployment, a published ROI result or proof that one deployment model is always cheaper.
A twelve-line TCO worksheet
Bring these inputs to the architecture and procurement review:
- Required languages, capabilities, operating hours and service locations.
- Monthly conversation minutes and completed conversations.
- Normal, peak and failure-mode concurrency.
- Measured GPU capacity and latency from a representative load test.
- Implementation, integration and acceptance effort.
- Hardware, facility, power, network and storage costs.
- Software, model, voice and enterprise-support licences.
- Internal operations, security, governance and help-desk effort.
- Resilience, spare capacity and disaster-recovery requirements.
- Refresh period, support windows and migration allowance.
- Equivalent managed-service commitments and variable charges.
- Cost per successful outcome under low, expected and high demand.
Document the source and owner of every assumption. Replace estimates with pilot measurements as they become available. Review the unit economics after launch because changes in demand, model efficiency and supplier pricing can move the answer.
The honest answer
On-premise AI is most likely to cost less when demand is large, stable and measurable; the required stack performs efficiently on commercial GPUs; the customer can reuse mature infrastructure and operations; and the deployment remains in service long enough to amortise its fixed cost.
Cloud is often more economical for experiments, uncertain demand, small volumes, rapidly changing components and services that need global elasticity without a customer-operated platform. Private cloud can sit between those positions.
Yepic can scope cloud, private-cloud and customer-hosted avatar architectures around the organisation’s data boundary, GPUs and operating model. The next step is not to request a universal per-minute promise. It is to provide a measured workload, required service level and approved deployment boundary, then compare the fully loaded options on the same basis.
