ML deployment · plan by default
/ml-drift
Design checks for input or prediction-distribution changes
Use for observed input/prediction distribution change; ml-evaluate requires outcomes to establish quality, and ml-parity or ml-leakage explain an offline-versus-production gap that is not population change; ml-monitor implements production drift and quality monitors and alert wiring.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-drift from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-drift, /just-vibe ml-drift and /jv:ml-drift in Claude. See shortcut setup and context examples.
/just-vibe:ml-drift Plan distribution checks that account for seasonality and sample size./just-vibe:ml-drift Investigate a new categorical value and a seasonal traffic shift./just-vibe:ml-drift Assess drift from aggregates without inferring model failure or silently refreshing the baseline.What the agent does
- Align schemas, sampling and seasonal windows of the baseline and current data.
- Choose meaningful per-feature and aggregate checks, and compare effect sizes and support changes in light of sample size.
- Separate data-pipeline changes from population changes, and define follow-up on signals.
Inputs
- reference/current data or prediction windows, features, seasonality, and sensitivity requirements.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Distribution-change detection and investigation; no automatic retraining.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Plan; reference/current data or prediction windows, features, seasonality, and sensitivity requirements.
- Prerequisites
- Versioned model and preprocessing artifacts, input/output schema, runtime/dependencies, operating targets, and authorized environment. Validate artifact trust before loading formats that can execute code. Packaging or writing monitoring configuration does not deploy a model or enable a hosted service.
Expected output
- Drift protocol or report with baseline/current identities, shift measures and thresholds, sample sizes, uncertainty and the next evidence needed.
How the work is checked
- A changed categorical support is detected; a statistically significant tiny shift is not automatically called model failure.
When to stop or clarify
- Drift is not proof of accuracy degradation without outcome evidence. Baseline refresh requires explicit policy, not silent adaptation.
Handling missing context
- Infer
- Read artifact format/trust, preprocessing schema, serving runtime, compatibility and existing rollout controls.
- Assume
- Prepare packaging/configuration and isolated checks without treating them as a live deployment.
- Ask
- Resolve the target, rollback compatibility and operating limits before rollout or load generation; missing production access does not block packaging.
Technical guidance
- Evidence
- Establish a reference population, feature semantics, seasonality, sample sizes and missing-data behavior.
- Method
- Monitor meaningful distribution changes with declared windows and thresholds; separate drift alerts from proven quality degradation.
- Pitfall
- A small p-value on a huge sample can flag irrelevant change, while missing labels prevent conclusions about accuracy.
- Check
- Inject a known distribution shift and a stable control, then check alert volume, cohort mix and delayed-label follow-up.
Situational decisions
When drift is statistically significant but outcomes are unavailable: Report a monitoring signal and follow-up, not confirmed accuracy degradation.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.