AI Agent Idempotency: Prevent Duplicate Avatar Actions
AI agent idempotency means one intended business action produces no more than one business effect, even when an avatar, gateway, queue or downstream API retries the request. The safe design is to create a stable effect identity before execution, claim it atomically, pass the same idempotency key downstream, store the authoritative result and reconcile any uncertain outcome before trying again.
This matters because a real-time avatar operates across unreliable boundaries. A user may repeat a request, speech recognition may emit a duplicate, a model may call the same tool twice, a connection may drop after the transaction commits, or a queue may redeliver work. Treating every arrival as new intent can create two payments, two appointments or two case updates.
Idempotency is therefore not a prompt instruction. It is a control in the action gateway and systems of record. This guide gives bank, government and enterprise teams a practical design they can review before an avatar receives transactional authority.
Why conversational systems duplicate actions
Retries are necessary. They recover from temporary network, service and infrastructure failures. HTTP itself distinguishes idempotent methods because clients may need to repeat a request after a communication failure; the HTTP Semantics standard also warns that a client should not automatically retry a non-idempotent request unless it knows the operation is safe.
An avatar adds failure paths that ordinary forms may not have:
- the user repeats “confirm” because the avatar pauses;
- speech recognition delivers overlapping final transcripts;
- the model emits a second tool call after a timeout;
- the browser, WebSocket or WebRTC session reconnects;
- an API gateway or workflow engine retries automatically;
- a queue redelivers after an acknowledgement is lost;
- the downstream system commits but its response never arrives.
The last case is the hardest. A timeout says the caller does not know the outcome. It does not prove failure. Retrying as a new action can duplicate a transaction; reporting success can mislead the user.
Separate four identities
Many implementations reuse a session ID or tool-call ID as the deduplication key. Neither reliably describes the business effect. Keep four identities distinct:
- Conversation identity: the session and turn in which the request arose. Useful for traceability, not for deciding whether two effects are the same.
- Attempt identity: every individual tool invocation or delivery. It should change on each attempt so operators can see retries.
- Confirmation identity: the version of the action summary the user approved. It changes if material details change.
- Effect identity: the stable key for the intended business outcome. It stays the same across safe retries and reconnects.
For example, “pay £80 to supplier A from account B today” needs one effect identity. A retry uses the same key. Changing the amount, payee, account or execution date creates new intent and must not silently reuse the key.
The AWS Builders’ Library guidance on retry-safe APIs recommends caller-provided request identifiers because they express intent more clearly than trying to infer sameness from parameters. The principle applies regardless of cloud choice: the effect key should come from a deterministic control layer, not from the model inventing a fresh value on every call.
Define an action contract before execution
The existing safe action-gateway pattern for AI agents separates conversation from controlled execution. Idempotency extends that boundary with an action contract containing:
- tenant, user and delegated actor identities;
- action type and target resource;
- normalised business parameters or a manifest hash;
- confirmation version and validity window;
- effect identity and attempt identity;
- current state, timestamps and policy decision;
- downstream transaction reference and authoritative response.
Bind the effect identity to the normalised payload. If the same key arrives with different material parameters, reject it and require a new confirmation. Stripe’s idempotent-request documentation illustrates this pattern by returning the earlier result for the same key while rejecting changed parameters. Retention windows and exact behaviour are implementation choices; the important control is that one key cannot mean two different effects.
Use a seven-state execution model
A boolean “done” flag is too weak for an action whose outcome can be uncertain. Use explicit states:
- Proposed: the avatar has assembled a bounded action, but it has no authority to commit.
- Confirmed: the required user or operator approval is bound to the current manifest.
- Claimed: the gateway has atomically reserved the effect identity. Competing attempts receive the existing record.
- Executing: one authorised worker is attempting the downstream operation.
- Committed: the authoritative system confirms the effect and its reference is stored.
- Failed: a definitive failure occurred and policy determines whether a retry is allowed.
- Unknown: the request may have committed, so reconciliation—not blind replay—is required.
The claim must be atomic: insert-if-absent, compare-and-set, a unique constraint or an equivalent transactional primitive. Checking the ledger and then inserting in two separate steps creates a race in which two workers can both observe “not seen” and both execute.
Build the retry-safe path
1. Classify the action
Separate read, prepare, commit and restricted operations. Reads may be naturally repeatable. A prepare step can calculate eligibility or show a quote. A commit changes business state and needs an effect key, authority check and durable record. Some actions should remain unavailable to the avatar or require human execution.
2. Create the effect identity outside the model
The gateway should derive or issue the key when it creates the action manifest. Store it before the first side effect. Do not ask the model to decide that two natural-language turns represent the same transaction; ambiguity belongs in confirmation, not in deduplication.
3. Claim once and replay the result
On the first authorised request, atomically claim the key. Later attempts with the same key and payload should receive the stored state or result rather than start new work. Keep every attempt in the audit trail without creating a second business effect.
4. Propagate the key downstream
If the system of record supports idempotency, pass the same business key through every retry. Open Banking’s Read/Write API profile, for example, requires the same idempotency key and body not to create a new resource within its defined window. This is useful precedent, not a universal banking rule or a Yepic implementation claim.
If a downstream service cannot honour a key, place the strongest available transactional boundary around the call and record its native reference. For database-plus-message workflows, an outbox or inbox pattern can prevent local state and event publication from drifting apart. External side effects such as email or legacy host commands may still need their own deduplication controls.
5. Reconcile unknown outcomes
After a timeout or connection loss, move the record to unknown. Query the downstream system by effect key, native reference or business attributes. If it committed, store and return that result. If it definitively did not, policy may permit the same effect identity to resume. If the status remains unknowable, route it to an operator instead of guessing.
This distinction is why “exactly once” should be treated cautiously. A failed response can leave the application unsure whether a transaction committed. Durable keys and reconciliation make retries safer; they do not abolish uncertainty across every external system.
6. Speak from authoritative state
The avatar should say “I’m checking whether that completed” when the ledger is unknown—not “it failed” or “done”. A committed response should include the authoritative reference the user can verify. The banking transaction-boundary design explains why conversation can prepare and explain an action while a deterministic service owns execution and confirmation.
Idempotency is not the same as exactly-once delivery
Queues and networks commonly provide at-least-once behaviour. Azure Service Bus guidance on duplicate processing says consumers should remain idempotent even when broker duplicate detection is available. An idempotency ledger controls repeated effects only within its scope, retention period and consistency model.
Review these limits explicitly:
- Expiry: a retry after the record is purged may look new.
- Region and partition races: two claims can win if uniqueness is not strongly enforced where execution occurs.
- Payload drift: reusing a key after any material change can return the wrong historic result.
- Partial workflows: booking, payment and notification may commit independently.
- Irreversible effects: compensation may be a new audited action, not a true rollback.
- External actors: another channel may perform a similar action without sharing the same key.
For multi-step work, give the overall workflow an identity and each effect its own child identity. Record which steps committed, which can be compensated and which require manual resolution.
Place the ledger where the action boundary lives
In a customer-hosted avatar deployment, the action gateway, identity checks, effect ledger and private integrations can be scoped to run inside the customer environment. That can keep sensitive parameters and transaction evidence under direct organisational control. It also gives the customer responsibility for durable storage, replication, patching, backup and reconciliation operations.
A private-cloud or managed-cloud architecture may offer mature queues, workflow engines and globally consistent data services. Public cloud can reduce operational burden. On-premise is not automatically more idempotent, faster or more resilient; the right choice depends on data boundaries, downstream systems, failure domains and operating capability.
Yepic’s work with Abu Dhabi Aviation and Oracle demonstrates the practical integration work around development and production environments, APIs, iframe delivery, real-time streaming, captions, microphone behaviour, WebRTC testing, browser remediation and cybersecurity support. It is relevant evidence for disciplined enterprise integration, but it is not presented as a completed customer-hosted idempotency deployment.
Twelve tests before an avatar may commit
- Send the same confirmed action twice concurrently; only one effect occurs.
- Retry after a client timeout that follows a successful commit; the stored result is returned.
- Reuse the key with a changed amount or target; the gateway rejects it.
- Repeat the user’s confirmation in the same turn; no second manifest is created.
- Reconnect the browser and replay the last tool call; the effect is not repeated.
- Redeliver the queue message after worker failure; execution remains single.
- Lose the downstream response; the record enters unknown and reconciliation runs.
- Delay the retry beyond the retention window; the system follows an explicit expiry policy.
- Race requests across regions or replicas; uniqueness holds at the execution boundary.
- Fail after step one of a multi-step workflow; committed and compensating actions remain visible.
- Remove the user’s authority between confirmation and execution; the commit is denied or re-authorised.
- Confirm that the avatar’s spoken status matches the ledger under committed, failed and unknown outcomes.
Run these alongside fault-injection tests for real-time avatars. Success is not merely “the API returned 200”; it is that one authorised intent caused one traceable effect and that every ambiguous path was resolved safely.
Start with one bounded action
Choose one consequential but reversible workflow. Define its effect identity, material parameters, authority, retention period, downstream lookup method, user wording and manual-resolution path. Then test every retry and timeout boundary before expanding the avatar’s permissions.
Yepic can scope private, on-premise, sovereign, private-cloud or cloud avatar architectures around the customer’s GPU environment, identity systems, knowledge sources and action gateways. The design should make the conversational experience feel natural while keeping each business effect deterministic, inspectable and under enterprise control.
