ML evaluation · plan by default
/ml-report
Document data, results, limitations, and intended use
Use to communicate established ML evidence; ml-evaluate creates evaluation results.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-report from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-report, /just-vibe ml-report and /jv:ml-report in Claude. See shortcut setup and context examples.
/just-vibe:ml-report Write a model report using these actual runs and identify unsupported uses./just-vibe:ml-report Write a model report with strong average performance but poor sparse-cohort evidence./just-vibe:ml-report Draft a report with missing test results; do not invent metrics or deployment approval.What the agent does
- Reconcile every number with a run and denominator.
- Describe training and evaluation conditions, separating validation selection from independent test evidence, and summarize baseline and slice results.
- Document the deployment population, limitations, excluded uses and missing release evidence.
Inputs
- task, dataset/model manifests, evaluation results, intended use, and audience.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Accurate model documentation and release assessment; no invented experiments or approval.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Plan; task, dataset/model manifests, evaluation results, intended use, and audience.
- Prerequisites
- Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
Expected output
- Model report or card with model/data/version provenance, metrics, baseline/slice evidence, operating assumptions, limitations, open risks and release gaps.
How the work is checked
- Every numerical claim traces to an actual run; missing cohort evidence is disclosed rather than generalized away.
When to stop or clarify
- Do not present a report as deployment authorization or claim suitability for untested populations.
Handling missing context
- Infer
- Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
- Assume
- Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
- Ask
- Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.
Technical guidance
- Evidence
- Collect intended use, dataset provenance, protocol, selected model, metrics, slices and deployment constraints.
- Method
- Separate measured results from anticipated value and list excluded uses plus concrete monitoring/revisit conditions.
- Pitfall
- Omitting failed runs or weak slices produces a misleading model story even if the best metric is correct.
- Check
- Trace every numerical claim to an artifact and verify data/model/version identity and unresolved limitations are retained.
Situational decisions
When evidence is missing for a key cohort or release gate: Keep the limitation visible and withhold the corresponding suitability claim.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.