ML evaluation · plan by default

/ml-report

Document data, results, limitations, and intended use

Use to communicate established ML evidence; ml-evaluate creates evaluation results.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-report from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-report, /just-vibe ml-report and /jv:ml-report in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-report Write a model report using these actual runs and identify unsupported uses.
edge · plan
/just-vibe:ml-report Write a model report with strong average performance but poor sparse-cohort evidence.
blocked · inspect
/just-vibe:ml-report Draft a report with missing test results; do not invent metrics or deployment approval.

What the agent does

  1. Reconcile every number with a run and denominator.
  2. Describe training and evaluation conditions, separating validation selection from independent test evidence, and summarize baseline and slice results.
  3. Document the deployment population, limitations, excluded uses and missing release evidence.

Inputs

  • task, dataset/model manifests, evaluation results, intended use, and audience.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Accurate model documentation and release assessment; no invented experiments or approval.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Plan; task, dataset/model manifests, evaluation results, intended use, and audience.
Prerequisites
Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.

Expected output

  • Model report or card with model/data/version provenance, metrics, baseline/slice evidence, operating assumptions, limitations, open risks and release gaps.

How the work is checked

  • Every numerical claim traces to an actual run; missing cohort evidence is disclosed rather than generalized away.

When to stop or clarify

  • Do not present a report as deployment authorization or claim suitability for untested populations.

Handling missing context

Infer
Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
Assume
Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
Ask
Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.

Technical guidance

Evidence
Collect intended use, dataset provenance, protocol, selected model, metrics, slices and deployment constraints.
Method
Separate measured results from anticipated value and list excluded uses plus concrete monitoring/revisit conditions.
Pitfall
Omitting failed runs or weak slices produces a misleading model story even if the best metric is correct.
Check
Trace every numerical claim to an artifact and verify data/model/version identity and unresolved limitations are retained.

Situational decisions

When evidence is missing for a key cohort or release gate: Keep the limitation visible and withhold the corresponding suitability claim.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring