ML data · plan by default
/ml-imbalance
Evaluate sampling, weighting, metrics, and thresholds for rare outcomes
Use when rare outcomes affect metrics or training; ml-threshold selects operational decisions.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-imbalance from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-imbalance, /just-vibe ml-imbalance and /jv:ml-imbalance in Claude. See shortcut setup and context examples.
/just-vibe:ml-imbalance Compare weighting and metrics for rare fraud with limited review capacity./just-vibe:ml-imbalance Compare a high-accuracy all-negative baseline against a rare-event model./just-vibe:ml-imbalance Assess imbalance with few positives; do not invent stable confidence or business costs.What the agent does
- Compute baseline prevalence, naive baselines and class counts by split.
- Choose task-relevant precision/recall measures, and compare resampling or weighting only within training folds.
- Assess how deployment prevalence changes the metrics and the threshold.
Inputs
- prevalence, class definitions, error costs/capacity, and split protocol.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Sampling/weighting, evaluation, and operating-point options for rare outcomes.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Plan; prevalence, class definitions, error costs/capacity, and split protocol.
- Prerequisites
- Task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
Expected output
- Imbalance strategy with baselines, per-class denominators, the resampling protocol, bounded experiments, threshold considerations and uncertainty limits.
How the work is checked
- A high-accuracy all-negative model is not accepted as useful; resampled training prevalence is not confused with deployed probability calibration.
When to stop or clarify
- Do not synthesize across validation/test boundaries or invent business tradeoffs. Low positive counts require uncertainty disclosure.
Handling missing context
- Infer
- Read prediction moment, label horizon, entity/time keys, split policy and dataset provenance from the task and manifests.
- Assume
- Use explicit synthetic examples for design when raw data is unavailable; do not infer missing labels or fit preprocessing across held-out boundaries.
- Ask
- Ask when unresolved label timing, grouping or target semantics would change the split/features; do not demand a full dataset to explain the method.
Technical guidance
- Evidence
- Measure prevalence, minority counts per split, error costs and operational capacity.
- Method
- Apply resampling or weighting only within training folds; evaluate ranking, precision/recall and probability interpretation separately.
- Pitfall
- Accuracy can hide zero minority recall; resampling changes prevalence and can distort uncorrected probabilities.
- Check
- Compare with an appropriate naive baseline, retain minority denominators and verify thresholds on untouched selection/evaluation data.
Situational decisions
When prevalence differs between sampled training and deployment: Separate learned ranking from probability calibration and expected operational workload.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.