ML evaluation · plan by default
/ml-threshold
Choose thresholds against explicit costs or capacity limits
Use to choose a decision cutoff under explicit costs/capacity; ml-evaluate measures fixed behavior.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-threshold from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-threshold, /just-vibe ml-threshold and /jv:ml-threshold in Claude. See shortcut setup and context examples.
/just-vibe:ml-threshold Choose validation thresholds when reviewers can inspect 200 transactions daily./just-vibe:ml-threshold Select a daily review threshold with tied scores and a hard queue cap./just-vibe:ml-threshold Compare cutoffs with unknown error costs; do not optimize on test labels.What the agent does
- Compute validation tradeoffs with denominators and tie handling.
- Translate them into expected workload under the stated volume, prevalence and capacity, with uncertainty, and choose using validation data.
- Reserve independent confirmation.
Inputs
- scores/labels, explicit error costs or capacity, prevalence, and validation protocol.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Select an operating point or decision policy; no production activation.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Plan; scores/labels, explicit error costs or capacity, prevalence, and validation protocol.
- Prerequisites
- Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
Expected output
- Threshold policy and recommendation with confusion and workload estimates, assumptions, sensitivity and the independent evaluation requirement.
How the work is checked
- A daily review cap is respected under declared volume assumptions; a changed prevalence is shown to affect workload/precision.
When to stop or clarify
- Do not invent business costs or optimize against the held-out test set. Unresolved priorities yield a tradeoff curve rather than a forced value.
Handling missing context
- Infer
- Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
- Assume
- Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
- Ask
- Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.
Technical guidance
- Evidence
- Obtain score distribution, error costs, review capacity, protected requirements and selection data.
- Method
- Choose an operating point using explicit constraints and uncertainty; keep threshold selection separate from final test reporting.
- Pitfall
- A threshold maximizing F1 may violate a daily capacity or false-positive budget.
- Check
- Compute the confusion matrix and workload at nearby thresholds, including tied scores and changing prevalence assumptions.
Situational decisions
When no agreed cost or capacity preference distinguishes options: Present the tradeoff curve and the missing decision rather than selecting an arbitrary optimum.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.