ML evaluation · inspect by default
/ml-error-analysis
Group failures into actionable patterns and examples
Use to inspect model mistakes; ml-slices computes cohort metrics and ml-debug-training diagnoses optimization.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-error-analysis from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-error-analysis, /just-vibe ml-error-analysis and /jv:ml-error-analysis in Claude. See shortcut setup and context examples.
/just-vibe:ml-error-analysis Group failures from these predictions into actionable patterns with denominators./just-vibe:ml-error-analysis Analyze rare high-cost errors without treating vivid examples as prevalence./just-vibe:ml-error-analysis Analyze aggregate errors when sensitive examples cannot be inspected.What the agent does
- Define the error event and its denominator according to the task.
- Group the model's errors by meaningful factors, and compare representative failures with matched successes and possible label problems.
- Propose targeted discriminating experiments.
Inputs
- predictions, labels, task costs, and permitted redacted examples.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Actionable failure patterns; no automatic retraining or relabeling.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Inspect; predictions, labels, task costs, and permitted redacted examples.
- Prerequisites
- Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
Expected output
- Error taxonomy with cohort counts and impact, representative examples, and interventions to test as discriminating experiments.
How the work is checked
- Common groups are judged against their population size; a few vivid cases do not imply prevalence.
When to stop or clarify
- Protect sensitive records. Patterns discovered on test data must not become tuning targets without a fresh evaluation plan.
Handling missing context
- Infer
- Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
- Assume
- Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
- Ask
- Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.
Technical guidance
- Evidence
- Inspect representative failures, successes, uncertainty, labels and error severity.
- Method
- Group by plausible mechanism and estimate frequency before proposing a targeted data/model/product change.
- Pitfall
- Anecdotal errors or explanation scores do not establish a causal pattern across the population.
- Check
- Check the hypothesized group on independent examples and include correct predictions that resemble the failures.
Situational decisions
When the analysis uses held-out test outcomes: Keep findings exploratory and require fresh confirmation before tuning to those patterns.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.