ML evaluation · plan by default
/ml-evaluate
Evaluate using task-appropriate metrics and baselines
Use for fixed-model evaluation; ml-threshold and ml-calibrate require separate selection data.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-evaluate from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-evaluate, /just-vibe ml-evaluate and /jv:ml-evaluate in Claude. See shortcut setup and context examples.
/just-vibe:ml-evaluate Plan evaluating the frozen model against the declared baseline and untouched test split./just-vibe:ml-evaluate Evaluate shuffled prediction rows with missing outputs and delayed labels./just-vibe:ml-evaluate Evaluate operational behavior without ground truth; do not report accuracy.What the agent does
- Align predictions and labels by stable row identity, and freeze eligibility and metric definitions.
- Count missing, excluded and failed predictions, then compute results for the declared metrics with appropriate uncertainty against the baseline; run predictions only when authorized.
- For supplied run exports, import and compare them with the experiments guide as ml-experiments does.
Inputs
- model, dataset/split, task metrics, baseline, and inference budget. Explicit evaluation execution selects apply.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Fixed-protocol performance evaluation, not model tuning.
- Writes
- Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
- Mode
- Plan an evaluation; inspect supplied results; apply for requested evaluator implementation or scoped evaluation execution.
- Prerequisites
- Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
Expected output
- Reproducible evaluation report with model/data/split identity, metrics, denominators, baseline comparison, uncertainty method and limitations.
How the work is checked
- Prediction/label misalignment fails validation; missing predictions are counted rather than silently dropped into better scores.
When to stop or clarify
- No test-set-driven changes during evaluation. Missing ground truth restricts output to operational/descriptive checks.
Handling missing context
- Infer
- Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
- Assume
- Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
- Ask
- Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.
Technical guidance
- Evidence
- Resolve metric formula/direction, positive class, sample weights, population, split and retained predictions.
- Method
- Compare appropriate baselines and quantify sampling uncertainty respecting groups/time dependence.
- Pitfall
- Treating every correlated row as independent can make confidence intervals artificially narrow.
- Check
- Hand-check a small confusion matrix or error calculation, preserve denominators and keep test data out of model selection.
Situational decisions
When observations are dependent within entities or time blocks: Use an uncertainty method matching that dependence or explicitly leave uncertainty unestimated.
When aggregate performance improves while a slice regresses: Report both with denominators and use ml-slices for the cohort comparison before recommending the model.
When the request is for local preparation or implementation: Implement the metric and slice evaluator with known-label controls; missing production labels limit conclusions without blocking evaluator code.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.