2026年10月7日 • Engineering • 7 分鐘閱讀

Designing AI Workflows for Partial Failure

Designing AI Workflows for Partial Failure

A workflow creates a service case, reserves a replacement component, and sends a customer confirmation. The case is created successfully. The reservation request times out. The confirmation has not been sent.

The next decision is consequential. Restarting the whole workflow could create another case and another reservation. Continuing could tell the customer that a component is available when it is not. Canceling the case could remove useful work while leaving a reservation in place.

This hypothetical example shows why recovery must account for individual effects. The workflow has partial progress, and the missing response leaves one result unknown. Treating both conditions as a single failed run discards the information needed to choose a safe next step.

Define what success means for each step

Before connecting tools, write down the intended result of each operation. Creating a case means a particular case exists with the requested customer and issue. Reserving a component means a particular allocation has reached the state required by the business process. Sending a confirmation means the messaging system has accepted the intended message; customer delivery may need separate evidence.

These results belong to different systems. A workflow's own success label cannot establish all of them. Define which target record or receipt supports each conclusion, whether processing is immediate or asynchronous, and how that result can be checked.

For the example, give the overall request a stable reference, R-618. Preserve the case identifier C-618 returned by the first system and the operation reference used for the reservation. That relationship lets the team distinguish this workflow's effects from similar work created by another coordinator.

Also define the dependency between steps. The customer confirmation requires a verified reservation, so it remains blocked while the allocation is uncertain. Case creation does not need to be repeated to investigate that uncertainty.

Record progress without turning uncertainty into failure

The workflow record should preserve what was requested, what response arrived, and what effect was subsequently verified. Keep sensitive content limited to what operators need; record references can often provide a review path without copying complete customer documents.

In this example, the first step is verified complete. The second was submitted, but its result is unknown. The third has not started. These are three distinct recovery positions.

Observed positionWhat the evidence establishesAppropriate next decision
Not submittedNo request was issued by this workflowCheck current prerequisites before starting
Accepted for processingThe target accepted a requestCheck for the required final result
Verified completeA matching target result is confirmedPreserve the completed step and inspect the next one
Rejected before effectThe target contract and response establish no effectResolve the cause before any permitted retry
Submitted, result unknownThe caller cannot establish the target outcomeReconcile the operation before continuing or repeating
Partly appliedSome intended changes are confirmedEvaluate the remaining changes individually

An error response is not automatically proof that nothing changed. Its meaning depends on the target's operation contract. Likewise, receiving a request identifier may establish acceptance without establishing the business result.

Reconcile an unknown result at the target

For the reservation timeout, begin with the original operation reference. If the target provides an operation-status lookup, use it. If it provides a read path for allocations, inspect the specific case, component, quantity, and request relationship. A matching allocation is stronger evidence than a general change in available stock.

An empty search result also needs interpretation. The write might still be processing, or the read view might update later. Understand the target's documented behavior before concluding that the request had no effect. Reconciliation may need a bounded wait, another authoritative read, or an operator's investigation.

If the allocation is confirmed, record its identifier and continue from the next unmet dependency. If the target establishes that no allocation occurred, consider a retry under the operation's current conditions. If the result cannot be resolved, hold the workflow and assign a person to investigate. Uncertainty should remain visible until there is evidence to change it.

This approach also handles a workflow process that restarts. Its recovery decision should come from the preserved operation record and target evidence, not from the model remembering what it intended to do.

Retry a known operation under a verified contract

Idempotency means a repeated request has no additional effect under the relevant operation's rules. Amazon's Builders' Library discussion of idempotent APIs explains why a timeout can leave callers uncertain and why an explicit request identity helps distinguish a retry from a new operation. The protection is an API contract implemented by the receiving service. A caller inventing a request label does not establish that protection.

For R-618, inspect whether the reservation interface supports such a contract, its scope, and any time limits. If it does, repeat the same intended operation according to those documented rules. Changing the quantity or component creates a different decision; it should not be hidden inside a retry of the original request.

Limit retries by time and attempts, and distinguish temporary transport problems from validation errors or permission denials. Repeated requests cannot repair an invalid component identifier or restore revoked access. When the permitted retry path is exhausted, preserve the last known position and escalate rather than generating a new workflow to bypass the problem.

Reassess the remaining work after conditions change

While the reservation is being investigated, another coordinator may arrange a different repair, the customer may cancel, or the component may become unsuitable. Recovery is therefore more than restoring connectivity. Check whether the remaining action still serves the current request.

Approval is one checkpoint in that reassessment. If a changed quantity, recipient, or schedule falls outside the reviewed action, obtain the decision required by the workflow's policy before proceeding. A previous approval does not resolve the allocation's unknown result, and a fresh approval does not remove a duplicate effect.

In the example, a confirmed reservation should not automatically trigger the original message if the customer has since canceled. The team needs to decide what to do with the allocation and case first. Keep the new decision linked to the original request so that the eventual record explains both the partial progress and the change in direction.

Plan correction and handover as real work

Undoing a completed operation is a separate action with its own requirements. Releasing a reserved component might be possible before dispatch but inappropriate afterward. A message already sent cannot simply be recalled by marking the local workflow canceled. Define the permitted correction for each meaningful effect and identify who can authorize it.

An operator taking over R-618 should receive the intended outcome, verified case and allocation references, unresolved questions, attempted checks, and blocked next step. The handover should say whether another attempt is permitted and what evidence would allow the workflow to resume. "Failed; please retry" is insufficient when an external system may already have acted.

Accentrust's Studio positioning includes multi-step workflows and tool orchestration. The OpenPort research project describes predictable response semantics and distinguishes draft governance state from an execution outcome. These are relevant design considerations, rather than guarantees of a global transaction, automatic rollback, or idempotency in every connected system.

For R-618, recovery is complete only when the team can explain the case, the reservation, and the customer communication together. Preserving successful work, resolving unknown effects, and handing off the remaining decision makes that explanation possible.

繼續閱讀

查看全部