ML evaluation · plan by default

/ml-robustness

Test missing inputs, noise, distribution changes, and boundaries

Use for bounded valid perturbation tests; ml-drift compares observed populations.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-robustness from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-robustness, /just-vibe ml-robustness and /jv:ml-robustness in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-robustness Plan plausible missing-input and noise tests with fixed labels and bounded compute.
edge · apply
/just-vibe:ml-robustness Test missing optional fields while rejecting transformations that change the outcome.
blocked · inspect
/just-vibe:ml-robustness Design robustness tests without running a large synthetic inference sweep.

What the agent does

  1. Define validity-preserving perturbations: which changes should preserve labels and expected behavior.
  2. Cap the sweep and run bounded tests in apply mode, comparing failure rate and input validity against a baseline.
  3. Identify failure envelopes and unsupported regions.

Inputs

  • model, plausible perturbations, operating bounds, metrics, and evaluation budget.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Missing inputs, noise, boundary cases, and realistic distribution changes.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
Mode
Plan; model, plausible perturbations, operating bounds, metrics, and evaluation budget. Apply for requested perturbation test code or bounded runs within that budget.
Prerequisites
Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.

Expected output

  • Robustness protocol or results with the perturbation contract, tested envelope, failures, unsupported regions and prioritized mitigations.

How the work is checked

  • Missing required inputs fail predictably; label-changing perturbations are not scored as ordinary invariance tests.

When to stop or clarify

  • No claims of universal robustness from a finite suite. Large synthetic sweeps stop at the declared compute cap.

Handling missing context

Infer
Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
Assume
Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
Ask
Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.

Technical guidance

Evidence
Identify plausible missingness, noise, boundary values and deployment shifts with bounded perturbations.
Method
Test semantic-preserving perturbations separately from changed-label cases; measure quality, rejection and coverage.
Pitfall
Arbitrary corruption may not represent deployment, and invariance is wrong when the perturbation should change the answer.
Check
Include an unchanged control, boundary-valid input and deliberately unsupported input; report the exact tested threat/shift model.

Situational decisions

When a perturbation changes the true label or leaves the valid domain: Classify it separately from an invariance failure.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring