ML evaluation · inspect by default
/ml-explain
Investigate behavior with appropriate explanation methods and limits
Use to interpret model behavior; explain describes code, teach explains concepts, and ml-ablation measures a feature group's contribution to held-out performance by retraining without it.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-explain from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-explain, /just-vibe ml-explain and /jv:ml-explain in Claude. See shortcut setup and context examples.
/just-vibe:ml-explain Explain these predictions and separate feature association from causation./just-vibe:ml-explain Explain correlated feature importance without implying causation./just-vibe:ml-explain Plan explanations without model artifacts or expensive inference authorization.What the agent does
- State whether the question concerns one prediction or global behavior, and choose a compatible method.
- Examine background or baseline data dependence, stability and correlated-feature sensitivity, and connect explanations to actual examples.
Inputs
- model, prediction/global behavior question, data access, and audience.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Appropriate feature/behavior explanation with method limitations; not causal attribution by default.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Inspect; model, prediction/global behavior question, data access, and audience.
- Prerequisites
- Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
Expected output
- Explanation with method/question fit, settings, example explanations, stability checks and limitations.
How the work is checked
- Correlated features are not treated as independent causal effects; unstable explanations are disclosed.
When to stop or clarify
- Expensive explanation runs need bounded execution. Do not expose proprietary or personal data through unnecessary examples.
Handling missing context
- Infer
- Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
- Assume
- Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
- Ask
- Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.
Technical guidance
- Evidence
- Resolve whether the question is global behavior, a local prediction, debugging or causal effect.
- Method
- Use a method compatible with model/data semantics and check explanation stability and plausible feature combinations.
- Pitfall
- Attributions are not causal effects; correlated inputs or impossible counterfactuals can make an explanation misleading.
- Check
- Compare nearby valid inputs or a known simple model and state approximation, background-data and stability limits.
Situational decisions
When explanations vary strongly with reasonable baselines: Report that dependence and avoid a causal or uniquely determined attribution claim.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.