All articles
RELIABILITY & AUTOMATIONFZM2026-09-257 min read

Crash-Safe Automation: Durable Intent, Checkpoints, and Recovery Without Blind Retry

The most dangerous automation failure is not a clean RED. It is an unknown state where an external mutation may have succeeded but the caller never received confirmation. Durable intent and reconciliation solve this class of failure.

1. Why Unknown State Is Worse Than RED A timeout does not prove that an operation failed. The remote system may have created the resource, switched the release, or committed data while the response was lost. Blind retry can duplicate the effect.

2. Durable Intent Before Mutation Before an external action, record what is about to happen: target resource, candidate identity, operation, and expected result. The intent must outlive the process performing the work.

3. Checkpoint After Every External Effect A safe workflow avoids long chains of mutations between durable checkpoints. After each meaningful side effect, record what changed, which version is active, and which verification steps have completed.

4. Reconcile Before Retry After an interruption, read the real external state before repeating anything: 1. If the intended effect already exists, continue forward. 2. If the effect is absent, retry safely. 3. If state conflicts with the intent, stop and reconstruct truth.

The pattern applies to releases, payments, queues, infrastructure automation, and long-running AI-assisted workflows.

GET STARTED

Have a project in mind? Let’s define the right next step.

Describe the business need and we will help define the approach, project scope, and initial budget.