ML evaluation · inspect by default
/ml-slices
Compare meaningful cohorts or operating conditions
Use for cohort performance comparisons; ml-error-analysis investigates individual failure mechanisms.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-slices from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-slices, /just-vibe ml-slices and /jv:ml-slices in Claude. See shortcut setup and context examples.
/just-vibe:ml-slices Compare operating-condition cohorts and report uncertainty for small slices./just-vibe:ml-slices Compare overlapping cohorts with one tiny high-error subgroup./just-vibe:ml-slices Assess slice coverage without sensitive row-level records or causal fairness claims.What the agent does
- Predefine important slices and their overlap where possible, and account for dependent samples.
- Compute counts and metrics consistently, flag small groups, and distinguish planned from exploratory comparisons.
- For supplied run exports, import and compare them with the experiments guide as ml-experiments does.
Inputs
- predictions/labels, meaningful cohorts, minimum sample guidance, and operating context.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Performance across cohorts/time/conditions and coverage gaps.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Inspect; predictions/labels, meaningful cohorts, minimum sample guidance, and operating context.
- Prerequisites
- Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
Expected output
- Slice table with definitions, denominators, metrics, uncertainty, worst-supported conditions, coverage gaps and follow-up data needs.
How the work is checked
- A tiny cohort's extreme score is qualified; missing cohorts are shown as no evidence rather than zero performance.
When to stop or clarify
- Avoid causal or fairness guarantees from a metric table alone. Do not expose identifying small-group records.
Handling missing context
- Infer
- Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
- Assume
- Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
- Ask
- Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.
Technical guidance
- Evidence
- Define meaningful cohorts, support counts, denominators, overlap and intended decision.
- Method
- Compare performance and uncertainty within slices, including missing group attributes and intersectional cases where supported.
- Pitfall
- Tiny slices and many comparisons can produce dramatic noise; aggregate improvement can hide a harmed cohort.
- Check
- Report counts and uncertainty with each metric and verify membership logic on hand-labeled examples.
Situational decisions
When a cohort has no outcomes or very few positives: Report unavailable/unstable evidence rather than a confident zero or perfect score.
When aggregate performance improves while a slice regresses: Show denominators and both outcomes, inspect parity/temporal issues and avoid ranking an incompatible evaluation.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.