Reliable Multi-Agent Execution: Tool Contracts, Idempotency, Checkpoints, and Side-Effect Recovery

When multiple agents enter business workflows, duplicate calls, conflicting state, and unknown outcomes become central problems. A hypothetical replacement-parts process illustrates how to make tool execution constrained, recoverable, and verifiable.

When several agents analyze material together, errors often remain in text. Once they reserve inventory, create shipping labels, or send notifications, errors can change the outside world. A timeout does not prove that an operation never happened, and a summary saying 'task complete' does not establish that business state is correct. Reliable execution puts model decisions inside explicit state and permission boundaries, with a recovery path for every operation that has side effects. The following discussion uses a hypothetical replacement-parts workflow to examine general execution architecture and recovery design.

Represent the task as verifiable state

Consider a replacement-parts workflow. A document agent verifies warranty evidence, an inventory agent checks availability, an execution agent reserves the part and creates a shipping label, and another agent prepares a notification. Completion means that one replacement request has one valid reservation, one eligible shipment record, and a verifiable processing history. It does not mean that every role has said 'done.' Roles may change, but completion conditions must remain stable.

Task state should distinguish a generated plan, passed validation, pending execution, an unknown outcome, confirmed success, and required human intervention. The unknown state is particularly important. If the connection breaks after a shipping-label request is sent, treating the operation as failed and retrying may create a second label. Treating it as successful may send a notification with nonexistent tracking information. Unknown is a business state requiring reconciliation, not an exception detail that can be omitted.

Every state transition should require verifiable conditions. For example, entering the label-created state requires a carrier record identifier and confirmation through a status query. An agent's prose can explain that state but cannot alter it by assertion. A human taking over then receives concrete unresolved facts instead of a conversation that must be reread and interpreted.

Tool contracts need structure and business semantics

JSON Schema provides basic validation through required fields, types, enumerations, and control of additional properties. An inventory-reservation tool can require a request identifier, warehouse identifier, part identifier, and positive integer quantity while rejecting undeclared parameters. Structural validity does not establish business validity, however. The tool service must use trusted data to verify tenant ownership, warranty eligibility, and whether the quantity exceeds the approved allowance.

A tool contract should also declare whether it reads or writes, which permissions it requires, which object version acts as a precondition, what success creates, and how the result can be queried. For example, include the previously read request revision when reserving inventory. If a person has since changed the requested quantity, return a conflict and recalculate. Otherwise, a structurally valid invocation can execute a stale decision.

Structure outputs to distinguish confirmed success, definitive rejection, retryable failure, and an unknown outcome, carrying business record identifiers and error categories. A controlled execution layer supplies credentials; the model should not choose a higher-privilege account in tool arguments. Instructions from pages, attachments, or other agents remain inputs to validate. Passing them into a tool does not authorize a tenant change, a larger allowance, or omitted preconditions.

Parallel work needs explicit ownership of writes

Multiple agents can usefully perform independent reads and analysis in parallel, such as checking evidence, querying stock, and drafting a notification. Represent dependencies in a task graph: address validation depends on confirmed recipient information, label creation depends on successful reservation, and the final notification depends on a confirmed shipment record. Expressing these dependencies only in a prompt cannot prevent two executors from submitting prematurely based on different interpretations.

Give each business object clear write ownership and use version checks for shared-state updates. If two agents propose different quantities from revision three of a request, detect the conflict instead of silently accepting the last write. Parallel results also need source versions and completion times. A late inventory analysis must not overwrite a cancellation or restore a cancelled request to a pending state.

The scheduler must handle lost workers and reassignment. A lease can limit how long a worker owns a task, but an expired worker may keep running. Critical write endpoints therefore need to validate an increasing execution generation or equivalent validity marker and reject stale submissions. Bound concurrency, tool calls, and total execution time as well, so repeated planning or conversations between agents cannot consume resources indefinitely.

An idempotency identifier represents one business intent

The AWS Builders' Library article on idempotent APIs discusses caller-provided request identifiers. For a replacement request, the scheduler can issue a stable operation identifier that survives network retries, process restarts, and executor changes. It represents the intent to make this reservation for this request, rather than the timestamp of an individual attempt. Generating a fresh identifier on each retry prevents the service from recognizing the same operation.

A design can scope identifiers by tenant, request, step, and business revision while recording a digest of normalized parameters. The same identifier with a different warehouse or quantity must produce a conflict, not silently reuse an earlier result. Identical parameters alone are also insufficient for deduplication: a user may legitimately submit two identical replacement requests. Preserve the initial result so repeated calls return the same verifiable business record.

The service producing the side effect must enforce the guarantee. Within one database, an idempotency record, unique constraint, and business write can share a transaction to handle concurrent calls. If an external carrier API lacks idempotency support, a local 'ready to call' record cannot close the window between remote success and local failure. Establish the provider's deduplication retention period too. Recovery outside that window requires querying existing results or human reconciliation, rather than claiming safe single execution for retries at any time.

Checkpoints preserve decisions; transactions preserve facts

LangGraph documents checkpoints for preserving graph state and continuing after interruption. Temporal distinguishes replayable workflow logic from external activities that may fail. Whichever framework is used, decide which results recovery must reuse. Persist the selected part, validated address, and model-generated execution plan as versioned decisions. A restart should not silently cause fresh reasoning to choose different parameters.

A checkpoint does not automatically include remote side effects in a local transaction. A shipping label may be created just before a crash prevents the success state from being saved; recovery then sees pending work. Persist the operation intent, call with a stable idempotency identifier, and confirm through a receipt or status query. Preserve the unknown state until confirmation. If replanning is required, create an explicit new version and account for whether the previous operation took effect and how it will be handled.

For events that must follow local business updates, AWS's transactional outbox pattern offers a useful structure: commit the business record and a pending event in one database transaction, then forward the event through a separate publisher. This prevents the local update and event creation from becoming disconnected. The publisher may still deliver duplicates, so consumers need deduplication. The pattern does not create one global atomic transaction across a database, carrier, and notification service.

Classify the failure before retrying

Failure categories determine recovery. Invalid arguments require correction, permission denial stops dependent actions, and version conflicts require a fresh read. Temporary throttling or failures known not to have executed may justify a retry policy. Temporal's retry documentation supplies declarative controls, but the application still needs to choose non-retryable failures and set attempt limits, an overall deadline, and backoff intervals.

For network timeouts, distinguish reads from writes. After a shipping-label write times out, first query by the operation identifier or remote receipt. If it succeeded, repair the local record. Retry only after nonexecution is established and preconditions still hold. When the remote query is eventually consistent, 'not found' may not prove absence; respect the provider's consistency behavior and allow a reconciliation window. If uncertainty cannot be resolved, keep the operation pending verification.

Manage retry budgets across the entire task. If a tool client, task executor, and supervising agent each retry three times, call volume can multiply. Assign retry responsibility to a defined layer and associate every attempt with its business operation. Randomized backoff can reduce synchronized recovery pressure. When the budget is exhausted, retain completed work and unresolved steps for a person or later recovery process instead of duplicating the whole workflow from the beginning.

Compensation and human intervention need concrete evidence

Recovery from a failed multistep process may continue outstanding work or compensate for completed actions. If address validation fails after reservation, release that reservation. If a request is cancelled after label creation, first determine whether the label can still be voided. Compensation must reference the original operation and itself be idempotent. Releasing stock cannot simply add a number to an aggregate counter, because repeated compensation could manufacture availability that does not exist.

Compensation usually cannot erase history. A sent notification cannot reliably be recalled, and a package already handed to a carrier may be impossible to cancel. Define reversible windows and irreversible boundaries during design, placing verifiable checks before irreversible actions where possible. When automatic recovery is insufficient, provide the original request, confirmed actions, uncertain actions, relevant receipts, and proposed handling options, while retaining the human operator's subsequent decision.

For actions requiring approval, bind that approval to a specific object, parameter digest, permission scope, and validity period. Material changes to the request invalidate reuse of the original approval. After recovery, recheck preconditions and whether approval remains valid, without needlessly requesting every unchanged and still-valid approval again. The interface should show the concrete result the action will produce and the scope that approval authorizes.

Use fault injection to test business invariants

Interrupt execution before a request is sent, after remote success but before receipt persistence, and after checkpoint commit, then inspect recovery. Also simulate duplicate task delivery, competing workers, cancellation arriving with a late result, expired credentials, and repeated compensation. Check business invariants directly: a request must not acquire extra valid shipping labels, reservation quantities must remain correct, and unauthorized operations must not occur.

Metrics should reflect actual consequences. Derive task completion from verified external business state. For duplicate side-effect rates, specify whether the denominator is logical operations or attempts. Track the number and age of unknown outcomes, the proportion requiring human recovery, recovery time, additional call cost, and compensation success. A model's judgment that tools were used correctly may be a supporting signal, but cannot replace inspection of inventory, shipment, and notification records.

Execution records should connect tasks, steps, business operations, attempts, and remote receipts, together with the contract and decision versions used. The user interface can translate that evidence into clear progress, such as 'Inventory reserved; shipping-label status awaiting confirmation,' rather than exposing raw internal exceptions. Before adding more agents, demonstrate that the existing workflow preserves these constraints under duplication, timeouts, and restarts. That evidence supports a wider scope of automation.

References

  1. JSON Schema — Object reference
  2. Amazon Builders’ Library — Making retries safe with idempotent APIs
  3. LangGraph — Persistence
  4. Temporal — Retry Policies
  5. Temporal — Workflow Execution overview
  6. AWS Prescriptive Guidance — Transactional outbox pattern
Back to insights
鲁ICP备2024109755号-2
Drag to move. Right-click, touch and hold, or press Shift+F10 to choose a corner.