Braking a retry storm: the circuit brake

When repeated failures or burned budget open the circuit, retryfuse stops replaying.

Back to overview

What the brake is for

A retry storm is a call that keeps failing and keeps being retried, burning cost and load for no progress. retryfuse returns brake to open the circuit and stop replaying. The brake is a safety verdict, not an action: the engine decides; it never retries on your behalf.

The failure-streak trip

retryfuse counts the trailing run of failing attempts — consecutive failure or timeout outcomes at the end of the attempt history. When that streak reaches the policy's failure threshold, the circuit opens and the engine returns brake with evidence fail: a known-bad condition was detected. The default threshold is 5 consecutive failures; set policy.failureThreshold (a positive integer) to tune it. A single later success resets the trailing streak to zero.

The budget trip

Each attempt may carry a costUnits weight (default 1). retryfuse sums the cost across all attempts; when the total reaches the policy's budget, the circuit opens and the engine returns brake with evidence fail to stop the burn. The default budget is 100 units; set policy.budgetUnits (a positive number) to tune it.

The in-flight brake is different

The circuit brake is not the only reason retryfuse brakes. If a request is still in_flight under an idempotency key when a duplicate arrives, the engine also returns brake — but with evidence pass, because braking a concurrent duplicate is the correct, safe outcome rather than a detected fault. See idempotency keys and exactly-once replay for that path.

Correctness outranks the brake

The brake is evaluated after the correctness gates. If the last attempt left the write state uncertain — a post-write timeout, a timeout of unknown phase, or an unknown outcome — retryfuse returns reconcile_required even when the run is still under the failure threshold and under budget. A blind replay of a possibly-applied effect is unsafe regardless of how much budget is left, so reconciliation wins. Equally, a reused key with a conflicting body is reject_conflict, and a completed, confirmed effect is replay_stored, before the circuit is ever considered.

How the decisions line up

Condition (evaluated in order)DecisionEvidence
Uncertain write, even under threshold and budgetreconcile_requiredunknown
A request is still in flight under this keybrakepass
Trailing failure streak at or above the thresholdbrakefail
Total cost at or above the budgetbrakefail
Under threshold and budget, no uncertain effectreplay_safepass

The decision core is pure and versioned (contract version 1), so the same attempt history and policy always yield the same verdict.

Next

See the contract reference for the exact JSON input and the policy fields, or why reconcile beats a blind retry for the correctness gate that outranks the brake. The fixture matrix has worked brake examples (a failure streak, a budget burn, and an uncertain write that reconciles despite being under budget).