ML evaluation · inspect by default

/ml-error-analysis

Group failures into actionable patterns and examples

Use to inspect model mistakes; ml-slices computes cohort metrics and ml-debug-training diagnoses optimization.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-error-analysis from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-error-analysis, /just-vibe ml-error-analysis and /jv:ml-error-analysis in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:ml-error-analysis Group failures from these predictions into actionable patterns with denominators.
edge · inspect
/just-vibe:ml-error-analysis Analyze rare high-cost errors without treating vivid examples as prevalence.
blocked · inspect
/just-vibe:ml-error-analysis Analyze aggregate errors when sensitive examples cannot be inspected.

What the agent does

  1. Define the error event and its denominator according to the task.
  2. Group the model's errors by meaningful factors, and compare representative failures with matched successes and possible label problems.
  3. Propose targeted discriminating experiments.

Inputs

  • predictions, labels, task costs, and permitted redacted examples.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Actionable failure patterns; no automatic retraining or relabeling.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Inspect; predictions, labels, task costs, and permitted redacted examples.
Prerequisites
Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.

Expected output

  • Error taxonomy with cohort counts and impact, representative examples, and interventions to test as discriminating experiments.

How the work is checked

  • Common groups are judged against their population size; a few vivid cases do not imply prevalence.

When to stop or clarify

  • Protect sensitive records. Patterns discovered on test data must not become tuning targets without a fresh evaluation plan.

Handling missing context

Infer
Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
Assume
Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
Ask
Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.

Technical guidance

Evidence
Inspect representative failures, successes, uncertainty, labels and error severity.
Method
Group by plausible mechanism and estimate frequency before proposing a targeted data/model/product change.
Pitfall
Anecdotal errors or explanation scores do not establish a causal pattern across the population.
Check
Check the hypothesized group on independent examples and include correct predictions that resemble the failures.

Situational decisions

When the analysis uses held-out test outcomes: Keep findings exploratory and require fresh confirmation before tuning to those patterns.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring