RAG Data Freshness: Stop Avatars Using Stale Knowledge
RAG data freshness is not the age of a vector index. It is the maximum time between an authoritative source changing and an AI avatar reliably using—or refusing to use—that change. For a bank, government service or regulated enterprise, that interval must cover new documents, amendments, withdrawals, permission changes and future-dated rules. A successful ingestion job is not enough.
The practical design is to give each knowledge class a freshness objective, preserve its effective dates and lineage, publish validated index versions atomically, invalidate downstream caches and test the answer path. If the system cannot prove that a material change has propagated, the avatar should use a safer source, narrow its answer or abstain.
This guide turns “keep the knowledge base current” into an architecture and operating model that a risk, platform or procurement team can review.
Freshness is a contract, not a timestamp
A document can be recently indexed and still be wrong. It may have been superseded at source, approved for a future date, withdrawn without a replacement, transformed incorrectly or retrieved from an old cache. Conversely, an old but still-effective regulation may be exactly the right evidence.
Define freshness in relation to authority and use. A workable contract records five moments:
- Source time: when the authoritative item was created, changed or withdrawn.
- Effective time: when its content becomes valid for the business decision.
- Observed time: when the ingestion service detected the event.
- Ready time: when the validated representation became queryable.
- Use time: when retrieval selected it for an avatar turn.
These clocks answer different questions. Source-to-observed lag tests the connector. Observed-to-ready lag tests transformation and indexing. Ready-to-use tests retrieval, cache and routing. Effective time decides whether a document is applicable at all.
The NIST AI RMF Manage playbook asks organisations to document data provenance and how obsolescence will be communicated. The W3C PROV ontology provides useful concepts for generation and invalidation. Neither dictates an avatar architecture, but both support a basic rule: knowledge needs a traceable lifecycle, not merely a file name and embedding.
Set objectives by consequence and volatility
One global “refresh every night” schedule is easy to operate and hard to defend. Classify sources by how quickly staleness becomes harmful.
- Live operational state: flight status, service availability, balances and case state usually belong behind a live API or tool, not in a periodically rebuilt vector index.
- Fast-changing guidance: incident notices, product restrictions and emergency instructions need event-driven or short-interval propagation plus an explicit failure mode.
- Controlled policy: approved procedures and rates need version, effective-from, effective-to and withdrawal controls. Publication may be scheduled, but activation must be exact.
- Stable reference: manuals and public information can tolerate a longer objective if ownership and review dates remain clear.
For each class, state the maximum propagation delay, the deletion or withdrawal delay, the required availability during synchronisation failure and who may authorise an exception. The number should come from business harm and source behaviour—not a generic RAG benchmark.
This extends Yepic’s permission-aware on-premise RAG architecture. Access control answers “may this user see it?” Freshness control answers “is this the currently applicable version?” A regulated service needs both decisions before evidence enters the model context.
Build a seven-stage freshness pipeline
1. Register the authoritative source
Inventory the system of record, content owner, stable item identifier, jurisdiction, classification, permission model and change mechanism. A shared folder labelled “latest” is not a source contract. Where multiple repositories repeat the same guidance, define precedence before retrieval is asked to resolve the conflict.
2. Detect additions, changes and removals
Use the strongest signal the source offers: change feed, webhook, sequence number, version, ETag, modified time or content hash. Polling can be appropriate, especially in restricted networks, but its worst-case interval must fit the freshness objective.
Deletion deserves its own design. Microsoft’s Azure AI Search documentation notes that change detection can be automatic while deletion detection may require a soft-delete strategy. Amazon Bedrock’s knowledge-base sync documentation describes incremental handling of added, modified and deleted documents. These are implementation examples, not Yepic requirements; the general lesson is that disappearance at source must become an explicit event downstream.
3. Stage an immutable candidate version
Do not overwrite the active corpus while parsing is incomplete. Stage the new source object, extracted text, metadata, chunks and embeddings under one candidate version. Keep stable identifiers so an amendment replaces the correct logical item and a withdrawal can remove every derivative chunk.
4. Preserve temporal and permission metadata
Every retrievable unit should retain source ID, source version, content hash, effective-from and effective-to times, ingestion time, classification, entitlement reference and invalidation state. Future-dated content may be indexed early, but query-time filters must keep it inactive. Expired content may need retention for audit without remaining eligible for answers.
5. Validate before activation
Run structural checks, malware and content controls required by policy, parsing comparisons, permission tests and retrieval probes. Confirm that expected questions find the candidate, that superseded material loses priority and that unrelated queries do not change unexpectedly. A green pipeline status is not proof that the answer changed correctly.
6. Activate one coherent knowledge generation
Move traffic through an alias, manifest or equivalent atomic control so a conversation does not mix half of the old corpus with half of the new one. Record the active generation in the governed context assembled for each avatar turn. That makes it possible to explain which knowledge state the model actually saw.
7. Invalidate every derivative
A new index does not automatically clear retrieval-result caches, semantic answer caches, session summaries or pre-built prompts. Publish an invalidation event carrying the affected source IDs and versions. Yepic’s guide to safe LLM semantic caching explains why time-to-live alone is insufficient: permission, policy and knowledge-version changes can require immediate bypass or purge.
Separate recency from validity and authority
“Prefer the newest document” is not a complete retrieval policy. A recently uploaded draft may not yet be approved. A current national rule may outrank a newer local note. A notice published today may only take effect next month.
Resolve candidates in this order:
- Filter to the user’s current permissions and approved use-case scope.
- Remove withdrawn, expired and not-yet-effective material.
- Apply source authority and jurisdiction precedence.
- Rank for query relevance within the surviving set.
- Use recency as a signal only where newer material is normally preferable.
Azure AI Search now documents a freshness-aware retrieval option that boosts recent material while preserving other ranking signals. It also cautions that freshness is not a replacement for date filtering. The feature is a product-specific example; the distinction between boosting and eligibility applies to any retrieval design.
When two authoritative sources conflict, do not ask the language model to improvise precedence. Route the case to a deterministic rule, an owner or a human handover. A polished spoken answer can make unresolved conflict appear more certain than it is.
Measure propagation, not job completion
A useful freshness dashboard reports distributions and breaches by source class:
- source-to-observed, observed-to-ready and end-to-end propagation time;
- age of the oldest unprocessed change and backlog by connector;
- withdrawal and permission-revocation propagation time;
- candidate validation failures and rollback frequency;
- queries served from a superseded generation or stale cache;
- retrieval of future-dated, expired or withdrawn content;
- freshness-objective compliance by knowledge class; and
- time from a breach to safe degradation and recovery.
Measure with canary documents and synthetic questions, not production conversations alone. A canary can be introduced, amended and withdrawn on a known schedule; probes verify that each state reaches retrieval and the spoken response. Yepic’s synthetic monitoring framework shows how controlled conversation probes can test the entire path before users discover the failure.
Choose the deployment boundary honestly
Customer-hosted ingestion, embeddings and retrieval can keep source content and operational evidence within an approved environment. It can also support local change feeds and direct control over activation. But the customer then owns connector reliability, indexing capacity, high availability, patching, backups and on-call response.
A managed cloud service may provide mature connectors and incremental synchronisation with less operational work. It still requires review of regions, retention, permissions, supported schedules, deletion behaviour and external data flows. Private cloud can divide those responsibilities. Hybrid retrieval may be best when stable material is indexed locally but live facts are queried from systems of record.
On-premise is therefore a control and operating choice, not a freshness guarantee. Yepic can scope cloud, private-cloud and customer-hosted real-time avatar deployments, including integration with enterprise sources and customer GPUs. The correct architecture depends on source volatility, network constraints, data classification and who will operate the freshness pipeline.
What Yepic’s enterprise work demonstrates
In the Abu Dhabi Aviation and Oracle enterprise avatar integration, Yepic helped connect a real-time avatar to enterprise information and supported development and production environments, APIs, iframe integration, streaming, captions, microphone behaviour, WebRTC and network testing, browser remediation, cybersecurity work and ongoing maintenance.
That is relevant evidence that the knowledge path sits inside a wider production service: browser, network, media and maintenance decisions affect whether a user receives a dependable answer. It is not presented as proof of a completed customer-hosted RAG-freshness implementation.
Twelve tests for an architecture review
- Add a source item and measure when the avatar can cite it.
- Amend one paragraph and confirm the old chunk no longer appears.
- Withdraw a document and prove its chunks, caches and session derivatives are ineligible.
- Publish a future-dated policy and verify it activates at the intended time, not ingestion time.
- Expire a rule without a replacement and confirm the avatar narrows or declines the answer.
- Change a permission and measure revocation through index, cache and active session.
- Introduce conflicting sources and verify deterministic authority handling.
- Fail the connector and check that the service exposes the breach and degrades safely.
- Fail parsing for one document and ensure the previous valid generation remains coherent.
- Update during concurrent conversations and prove no turn mixes knowledge generations.
- Ask equivalent questions in each supported language and compare effective evidence.
- Reconstruct the source version, effective time, index generation and cache decision used for one material answer.
Start with one freshness-critical answer
Select a question whose answer changes in production: an airport disruption notice, a banking product restriction, a citizen-service deadline or a maintenance procedure. Map the authoritative source, every derivative copy and the maximum safe propagation delay. Then exercise addition, amendment, withdrawal and failure before expanding the corpus.
The goal is not a knowledge base that looks current in a dashboard. It is an avatar that can show which effective, authorised source supports the answer—and can recognise when the evidence pipeline is too stale to speak with confidence.
