ML evaluation · plan by default
/ml-robustness
Test missing inputs, noise, distribution changes, and boundaries
Use for bounded valid perturbation tests; ml-drift compares observed populations.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-robustness from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-robustness, /just-vibe ml-robustness and /jv:ml-robustness in Claude. See shortcut setup and context examples.
/just-vibe:ml-robustness Plan plausible missing-input and noise tests with fixed labels and bounded compute./just-vibe:ml-robustness Test missing optional fields while rejecting transformations that change the outcome./just-vibe:ml-robustness Design robustness tests without running a large synthetic inference sweep.What the agent does
- Define validity-preserving perturbations: which changes should preserve labels and expected behavior.
- Cap the sweep and run bounded tests in apply mode, comparing failure rate and input validity against a baseline.
- Identify failure envelopes and unsupported regions.
Inputs
- model, plausible perturbations, operating bounds, metrics, and evaluation budget.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Missing inputs, noise, boundary cases, and realistic distribution changes.
- Writes
- Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
- Mode
- Plan; model, plausible perturbations, operating bounds, metrics, and evaluation budget. Apply for requested perturbation test code or bounded runs within that budget.
- Prerequisites
- Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
Expected output
- Robustness protocol or results with the perturbation contract, tested envelope, failures, unsupported regions and prioritized mitigations.
How the work is checked
- Missing required inputs fail predictably; label-changing perturbations are not scored as ordinary invariance tests.
When to stop or clarify
- No claims of universal robustness from a finite suite. Large synthetic sweeps stop at the declared compute cap.
Handling missing context
- Infer
- Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
- Assume
- Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
- Ask
- Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.
Technical guidance
- Evidence
- Identify plausible missingness, noise, boundary values and deployment shifts with bounded perturbations.
- Method
- Test semantic-preserving perturbations separately from changed-label cases; measure quality, rejection and coverage.
- Pitfall
- Arbitrary corruption may not represent deployment, and invariance is wrong when the perturbation should change the answer.
- Check
- Include an unchanged control, boundary-valid input and deliberately unsupported input; report the exact tested threat/shift model.
Situational decisions
When a perturbation changes the true label or leaves the valid domain: Classify it separately from an invariance failure.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.