ML evaluation · inspect by default

/ml-slices

Compare meaningful cohorts or operating conditions

Use for cohort performance comparisons; ml-error-analysis investigates individual failure mechanisms.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-slices from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-slices, /just-vibe ml-slices and /jv:ml-slices in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:ml-slices Compare operating-condition cohorts and report uncertainty for small slices.
edge · inspect
/just-vibe:ml-slices Compare overlapping cohorts with one tiny high-error subgroup.
blocked · inspect
/just-vibe:ml-slices Assess slice coverage without sensitive row-level records or causal fairness claims.

What the agent does

  1. Predefine important slices and their overlap where possible, and account for dependent samples.
  2. Compute counts and metrics consistently, flag small groups, and distinguish planned from exploratory comparisons.
  3. For supplied run exports, import and compare them with the experiments guide as ml-experiments does.

Inputs

  • predictions/labels, meaningful cohorts, minimum sample guidance, and operating context.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Performance across cohorts/time/conditions and coverage gaps.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Inspect; predictions/labels, meaningful cohorts, minimum sample guidance, and operating context.
Prerequisites
Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.

Expected output

  • Slice table with definitions, denominators, metrics, uncertainty, worst-supported conditions, coverage gaps and follow-up data needs.

How the work is checked

  • A tiny cohort's extreme score is qualified; missing cohorts are shown as no evidence rather than zero performance.

When to stop or clarify

  • Avoid causal or fairness guarantees from a metric table alone. Do not expose identifying small-group records.

Handling missing context

Infer
Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
Assume
Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
Ask
Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.

Technical guidance

Evidence
Define meaningful cohorts, support counts, denominators, overlap and intended decision.
Method
Compare performance and uncertainty within slices, including missing group attributes and intersectional cases where supported.
Pitfall
Tiny slices and many comparisons can produce dramatic noise; aggregate improvement can hide a harmed cohort.
Check
Report counts and uncertainty with each metric and verify membership logic on hand-labeled examples.

Situational decisions

When a cohort has no outcomes or very few positives: Report unavailable/unstable evidence rather than a confident zero or perfect score.

When aggregate performance improves while a slice regresses: Show denominators and both outcomes, inspect parity/temporal issues and avoid ranking an incompatible evaluation.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring