ML evaluation · plan by default

/ml-threshold

Choose thresholds against explicit costs or capacity limits

Use to choose a decision cutoff under explicit costs/capacity; ml-evaluate measures fixed behavior.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-threshold from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-threshold, /just-vibe ml-threshold and /jv:ml-threshold in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-threshold Choose validation thresholds when reviewers can inspect 200 transactions daily.
edge · plan
/just-vibe:ml-threshold Select a daily review threshold with tied scores and a hard queue cap.
blocked · inspect
/just-vibe:ml-threshold Compare cutoffs with unknown error costs; do not optimize on test labels.

What the agent does

  1. Compute validation tradeoffs with denominators and tie handling.
  2. Translate them into expected workload under the stated volume, prevalence and capacity, with uncertainty, and choose using validation data.
  3. Reserve independent confirmation.

Inputs

  • scores/labels, explicit error costs or capacity, prevalence, and validation protocol.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Select an operating point or decision policy; no production activation.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Plan; scores/labels, explicit error costs or capacity, prevalence, and validation protocol.
Prerequisites
Frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.

Expected output

  • Threshold policy and recommendation with confusion and workload estimates, assumptions, sensitivity and the independent evaluation requirement.

How the work is checked

  • A daily review cap is respected under declared volume assumptions; a changed prevalence is shown to affect workload/precision.

When to stop or clarify

  • Do not invent business costs or optimize against the held-out test set. Unresolved priorities yield a tradeoff curve rather than a forced value.

Handling missing context

Infer
Read frozen model/data identities, metric definitions, denominators and supplied predictions; separate validation from test use.
Assume
Compute only supported metrics on permitted samples and label missing labels or subgroup coverage as unknown.
Ask
Ask when the operating cost/threshold or population changes the evaluation decision; do not fabricate labels to avoid a question.

Technical guidance

Evidence
Obtain score distribution, error costs, review capacity, protected requirements and selection data.
Method
Choose an operating point using explicit constraints and uncertainty; keep threshold selection separate from final test reporting.
Pitfall
A threshold maximizing F1 may violate a daily capacity or false-positive budget.
Check
Compute the confusion matrix and workload at nearby thresholds, including tied scores and changing prevalence assumptions.

Situational decisions

When no agreed cost or capacity preference distinguishes options: Present the tradeoff curve and the missing decision rather than selecting an arbitrary optimum.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring