ML data · plan by default

/ml-imbalance

Evaluate sampling, weighting, metrics, and thresholds for rare outcomes

Use when rare outcomes affect metrics or training; ml-threshold selects operational decisions.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-imbalance from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-imbalance, /just-vibe ml-imbalance and /jv:ml-imbalance in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-imbalance Compare weighting and metrics for rare fraud with limited review capacity.
edge · plan
/just-vibe:ml-imbalance Compare a high-accuracy all-negative baseline against a rare-event model.
blocked · inspect
/just-vibe:ml-imbalance Assess imbalance with few positives; do not invent stable confidence or business costs.

What the agent does

  1. Compute baseline prevalence, naive baselines and class counts by split.
  2. Choose task-relevant precision/recall measures, and compare resampling or weighting only within training folds.
  3. Assess how deployment prevalence changes the metrics and the threshold.

Inputs

  • prevalence, class definitions, error costs/capacity, and split protocol.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Sampling/weighting, evaluation, and operating-point options for rare outcomes.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Plan; prevalence, class definitions, error costs/capacity, and split protocol.
Prerequisites
Task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.

Expected output

  • Imbalance strategy with baselines, per-class denominators, the resampling protocol, bounded experiments, threshold considerations and uncertainty limits.

How the work is checked

  • A high-accuracy all-negative model is not accepted as useful; resampled training prevalence is not confused with deployed probability calibration.

When to stop or clarify

  • Do not synthesize across validation/test boundaries or invent business tradeoffs. Low positive counts require uncertainty disclosure.

Handling missing context

Infer
Read prediction moment, label horizon, entity/time keys, split policy and dataset provenance from the task and manifests.
Assume
Use explicit synthetic examples for design when raw data is unavailable; do not infer missing labels or fit preprocessing across held-out boundaries.
Ask
Ask when unresolved label timing, grouping or target semantics would change the split/features; do not demand a full dataset to explain the method.

Technical guidance

Evidence
Measure prevalence, minority counts per split, error costs and operational capacity.
Method
Apply resampling or weighting only within training folds; evaluate ranking, precision/recall and probability interpretation separately.
Pitfall
Accuracy can hide zero minority recall; resampling changes prevalence and can distort uncorrected probabilities.
Check
Compare with an appropriate naive baseline, retain minority denominators and verify thresholds on untouched selection/evaluation data.

Situational decisions

When prevalence differs between sampled training and deployment: Separate learned ranking from probability calibration and expected operational workload.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring