ML evaluation · inspect by default
/ml-calibrate
Assess predicted probabilities against observed outcomes
Use to assess or fit probability calibration; ml-threshold maps scores to decisions.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-calibrate from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-calibrate, /just-vibe ml-calibrate and /jv:ml-calibrate in Claude. See shortcut setup and context examples.
/just-vibe:ml-calibrate Assess probability calibration from the supplied held-out scores and outcomes./just-vibe:ml-calibrate Calibrate a model trained on oversampled positives./just-vibe:ml-calibrate Review probability outputs with too few outcomes to fit a reliable calibrator.What the agent does
- Check probability semantics, then inspect reliability by range and cohort with proper scoring measures.
- Fit any calibrator on permitted data separate from final evaluation, and compare it by cohort on untouched evaluation data.
Inputs
- predicted probabilities, labels, sampling/prevalence context, and intended use.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Probability reliability assessment; fitting a calibrator uses a separate held-out calibration protocol in apply mode.
- Writes
- Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
- Mode
- Inspect; predicted probabilities, labels, sampling/prevalence context, and intended use. Apply to fit a calibrator on a permitted split when requested.
- Prerequisites
- Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
Expected output
- Calibration report or calibrated artifact with split provenance, the calibration protocol, reliability evidence and an independent comparison.
How the work is checked
- Good ranking is not mistaken for calibrated probability; fitting and evaluating a calibrator on identical records is rejected.
When to stop or clarify
- Small samples and prevalence shift limit conclusions. Do not alter deployed probabilities without a rollout request.
Handling missing context
- Infer
- Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
- Assume
- Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
- Ask
- Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.
Technical guidance
- Evidence
- Inspect probability outputs, class definition, prevalence, selection split and calibration metric/binning.
- Method
- Fit calibration on permitted selection data and evaluate reliability on separate data; compare proper scoring rules and reliability curves.
- Pitfall
- Ranking quality does not imply probability accuracy; coarse bins or shifted prevalence can conceal miscalibration.
- Check
- Check perfect, constant and confidently wrong synthetic predictions, then evaluate held-out calibration with sample support per region.
Situational decisions
When prevalence or sampling changed since calibration: Assess transfer assumptions and recalibration evidence without claiming ranking quality proves reliable probabilities.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.