ML evaluation · inspect by default

/ml-explain

Investigate behavior with appropriate explanation methods and limits

Use to interpret model behavior; explain describes code, teach explains concepts, and ml-ablation measures a feature group's contribution to held-out performance by retraining without it.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-explain from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-explain, /just-vibe ml-explain and /jv:ml-explain in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:ml-explain Explain these predictions and separate feature association from causation.
edge · inspect
/just-vibe:ml-explain Explain correlated feature importance without implying causation.
blocked · inspect
/just-vibe:ml-explain Plan explanations without model artifacts or expensive inference authorization.

What the agent does

  1. State whether the question concerns one prediction or global behavior, and choose a compatible method.
  2. Examine background or baseline data dependence, stability and correlated-feature sensitivity, and connect explanations to actual examples.

Inputs

  • model, prediction/global behavior question, data access, and audience.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Appropriate feature/behavior explanation with method limitations; not causal attribution by default.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Inspect; model, prediction/global behavior question, data access, and audience.
Prerequisites
Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.

Expected output

  • Explanation with method/question fit, settings, example explanations, stability checks and limitations.

How the work is checked

  • Correlated features are not treated as independent causal effects; unstable explanations are disclosed.

When to stop or clarify

  • Expensive explanation runs need bounded execution. Do not expose proprietary or personal data through unnecessary examples.

Handling missing context

Infer
Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
Assume
Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
Ask
Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.

Technical guidance

Evidence
Resolve whether the question is global behavior, a local prediction, debugging or causal effect.
Method
Use a method compatible with model/data semantics and check explanation stability and plausible feature combinations.
Pitfall
Attributions are not causal effects; correlated inputs or impossible counterfactuals can make an explanation misleading.
Check
Compare nearby valid inputs or a known simple model and state approximation, background-data and stability limits.

Situational decisions

When explanations vary strongly with reasonable baselines: Report that dependence and avoid a causal or uniquely determined attribution claim.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring