Testing · apply by default
/test-flaky
Repair nondeterminism using repeated evidence
Use for nondeterministic failures; debug first identifies the relevant failing test/environment.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select test-flaky from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv test-flaky, /just-vibe test-flaky and /jv:test-flaky in Claude. See shortcut setup and context examples.
/just-vibe:test-flaky Fix the test that fails only in suite order; cap diagnosis at 30 repeats./just-vibe:test-flaky Repair a test that fails only after another test changes global state./just-vibe:test-flaky Inspect flake logs without rerunning expensive suites or adding arbitrary sleeps.What the agent does
- Record order, seed, clock and shared-resource conditions and the repeat budget. Without a stated budget, derive one from single-run duration and the observed failure rate, cap total wall time (state the cap once) and report the detection power of the chosen sample.
- Reproduce under controlled repeats, varying one factor at a time within the budget, and inspect the first divergent evidence.
- Fix isolation or synchronization, replacing timing guesses with explicit synchronization, and rerun bounded stress checks.
Inputs
- flaky test, failure history, environment, and repetition budget.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Nondeterminism from shared state, order, time, async behavior, randomness, or environment.
- Writes
- Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
- Mode
- Apply; flaky test, failure history, environment, and repetition budget.
- Prerequisites
- Defined behavior, existing test conventions/runners, isolated fixtures, and relevant dependencies. Requested bounded verification may use owned isolated fixtures without authorizing product edits or live-system tests. Never test destructive behavior against production by default; distinguish mocked behavior from real integration evidence.
Expected output
- Cause and trigger conditions, the isolation or synchronization fix, repeat counts, failure rates and residual uncertainty.
How the work is checked
- The triggering interleaving/order is exercised; no failures across a stated sample is not described as mathematical proof.
When to stop or clarify
- Do not hide failures with retries, sleeps, skipped tests, or inflated timeouts without evidence. Stop at the repeat budget.
Handling missing context
- Infer
- Read behavior contracts, existing runners and test conventions; distinguish fixture setup failure from a behavioral failure.
- Assume
- Use the smallest existing local runner and isolated synthetic fixtures that distinguish the requested behavior. When the method needs a library, runner, container runtime or load tool the project lacks, name the exact package or tool, the files it changes and any download, and add it only when the request authorizes new dev dependencies or tools; label a hand-written generator without shrinking, or a fake in place of a real dependency, as such.
- Ask
- Ask about an unresolved contract that changes the expected result, or the target/load limits before external testing; do not ask the user to choose a runner already configured.
Technical guidance
- Evidence
- Gather repeated outcomes, order, seed, clock, shared resources and cleanup evidence.
- Method
- Force the suspected race or shared-state condition deterministically before changing implementation or tests.
- Pitfall
- Increasing timeouts, retries or skips can hide nondeterminism rather than repair it.
- Check
- Run a bounded repeated sample with retained counts and the targeted interleaving; zero observed failures remains finite evidence.
Situational decisions
When no failure occurs during bounded repeats: Report the sample and uncertainty; do not declare the flake eliminated solely from absence.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.