An agent sends a booking request. The server creates the booking, but the response is lost. The agent process restarts from its last checkpoint and sends the request again. Saving workflow state did not stop the duplicate booking.

Persistence and exactly-once effects are different properties. A checkpoint records a point in the workflow. An external service may have acted after that point and before the failure. Recovery needs a way to reconcile those two histories.

Name the unit of recoverable work

LangGraph's persistence documentation separates thread-scoped checkpoints from longer-term stores. Checkpoints can support continuation, inspection, and fault tolerance. They are not by themselves a guarantee that a tool action happens only once.

For a synthetic report-delivery agent, separate retrieving evidence, generating the report, saving the artifact, and emailing it. Repeating retrieval may be acceptable; repeating a send may not. The checkpoint boundary should be chosen with those effects in view.

Give the action a stable identity

An idempotency key lets a service recognize repeated requests for the same intended operation. The key should be stable across retries of that operation and different for genuinely new work. The service must define how it stores the result, detects a conflicting payload under the same key, and expires the record.

A local deduplication set is insufficient if the process loses it or a second worker operates elsewhere. Enforce the property at the resource owner when possible. “We generate a UUID” is not an idempotency policy unless the server binds that identifier to the effect.

Represent the uncertain state

After a timeout, mark the action as unknown rather than immediately failed. Query the operation status or resource state using the stable identifier. If the service cannot distinguish “not performed” from “performed but response lost,” choose a conservative recovery appropriate to the consequence.

Do not infer that a write failed because the user cancelled the browser request. Client cancellation and remote operation cancellation can occur at different times. The UI should explain whether an action is pending, completed, cancelled, or still being reconciled.

Compensation is a new action

Undoing a booking or sending a correction is not the same as rewinding local state. It may require fresh permission, incur cost, or fail. Record the original effect and the compensating request separately. Do not promise rollback of an email already delivered or a transaction whose recipient has acted on it.

For local database changes, transactional techniques may provide stronger guarantees inside a defined boundary. Cross-service workflows still need explicit assumptions and failure paths. A transaction in one component does not cover every downstream action.

Inject failures around the seam

Test before the call, after the server accepts it, before the response is persisted, and during recovery. Run two workers against the same operation identifier. Expire the deduplication record and document what the service promises afterward.

A fixture should verify the final state and count actual effects. The final answer “sent successfully” is diagnostic output, not proof. Keep credentials and user data out of replay fixtures unless the test environment has explicit permission and isolation.

Make the recovery contract readable

For each side effect, write the allowed retry condition, operation identifier, status check, timeout budget, and escalation path. The AWS Builders' Library discussion of idempotent APIs explains why caller intent and duplicate recognition matter when making retries safe.

Durability should mean that the workflow can continue intelligibly after interruption. That requires persisted state, bounded retries, and an honest account of external effects—not merely a resume button.

Sources and further reading