Backend · apply by default
/backend-resilience
Add appropriate timeouts, bounded retries, and failure handling
Use for bounded dependency failure behavior; ops-incident handles an active incident.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select backend-resilience from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv backend-resilience, /just-vibe backend-resilience and /jv:backend-resilience in Claude. See shortcut setup and context examples.
/just-vibe:backend-resilience Add bounded retry and timeout behavior without duplicating unsafe requests./just-vibe:backend-resilience Add retries without multiplying nested dependency attempts beyond the deadline./just-vibe:backend-resilience Design resilience from contracts without injecting faults into production.What the agent does
- Classify which operations and effects are safe to retry.
- Allocate an end-to-end deadline across attempts and dependencies, control exponential backoff and jitter within the total cap, and propagate cancellation.
- Simulate partial dependency failures.
Inputs
- dependency failure modes, latency budget, retry constraints, and fallback policy.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Timeouts, bounded retries, cancellation, circuit/fallback behavior, and useful errors.
- Writes
- Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
- Mode
- Apply; dependency failure modes, latency budget, retry constraints, and fallback policy.
- Prerequisites
- Service source, data/interface contracts, framework/runtime versions, and test environment. Default apply operations target local code and isolated tests; live infrastructure/data mutations require their own requested scope.
Expected output
- Timeout/retry/fallback matrix with the resilience changes and controlled outage and partial-effect checks.
How the work is checked
- A failing dependency cannot trigger unbounded retries; non-idempotent operations are not duplicated by blind retry.
When to stop or clarify
- Do not silently return stale/synthetic business results without an agreed fallback. Production fault injection requires separate authorization.
Handling missing context
- Infer
- Trace service callers, request contracts, authorization, transactions, retries and existing test infrastructure.
- Assume
- Use the existing persistence and framework; isolate local tests from live services.
- Ask
- Resolve ambiguous durability, duplication or consistency requirements before encoding them; absent production access does not prevent local implementation.
Technical guidance
- Evidence
- Inventory end-to-end deadline, nested retries, cancellation owners, concurrency limits and partial effects.
- Method
- Budget retries with jitter and bounded attempts, propagate owned cancellation and reconcile uncertain mutations before replay.
- Pitfall
- Retrying at every layer multiplies traffic; a timeout does not prove the remote operation failed.
- Check
- Simulate slow dependency, transient error, permanent refusal and response loss after success; verify total deadline and effect count.
Situational decisions
When the deadline expires with uncertain external mutation: Return an explicit uncertain/reconcilable state and preserve the stable operation ID.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.