ML deployment · plan by default

/ml-drift

Design checks for input or prediction-distribution changes

Use for observed input/prediction distribution change; ml-evaluate requires outcomes to establish quality, and ml-parity or ml-leakage explain an offline-versus-production gap that is not population change; ml-monitor implements production drift and quality monitors and alert wiring.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-drift from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-drift, /just-vibe ml-drift and /jv:ml-drift in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-drift Plan distribution checks that account for seasonality and sample size.
edge · plan
/just-vibe:ml-drift Investigate a new categorical value and a seasonal traffic shift.
blocked · inspect
/just-vibe:ml-drift Assess drift from aggregates without inferring model failure or silently refreshing the baseline.

What the agent does

  1. Align schemas, sampling and seasonal windows of the baseline and current data.
  2. Choose meaningful per-feature and aggregate checks, and compare effect sizes and support changes in light of sample size.
  3. Separate data-pipeline changes from population changes, and define follow-up on signals.

Inputs

  • reference/current data or prediction windows, features, seasonality, and sensitivity requirements.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Distribution-change detection and investigation; no automatic retraining.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Plan; reference/current data or prediction windows, features, seasonality, and sensitivity requirements.
Prerequisites
Versioned model and preprocessing artifacts, input/output schema, runtime/dependencies, operating targets, and authorized environment. Validate artifact trust before loading formats that can execute code. Packaging or writing monitoring configuration does not deploy a model or enable a hosted service.

Expected output

  • Drift protocol or report with baseline/current identities, shift measures and thresholds, sample sizes, uncertainty and the next evidence needed.

How the work is checked

  • A changed categorical support is detected; a statistically significant tiny shift is not automatically called model failure.

When to stop or clarify

  • Drift is not proof of accuracy degradation without outcome evidence. Baseline refresh requires explicit policy, not silent adaptation.

Handling missing context

Infer
Read artifact format/trust, preprocessing schema, serving runtime, compatibility and existing rollout controls.
Assume
Prepare packaging/configuration and isolated checks without treating them as a live deployment.
Ask
Resolve the target, rollback compatibility and operating limits before rollout or load generation; missing production access does not block packaging.

Technical guidance

Evidence
Establish a reference population, feature semantics, seasonality, sample sizes and missing-data behavior.
Method
Monitor meaningful distribution changes with declared windows and thresholds; separate drift alerts from proven quality degradation.
Pitfall
A small p-value on a huge sample can flag irrelevant change, while missing labels prevent conclusions about accuracy.
Check
Inject a known distribution shift and a stable control, then check alert volume, cohort mix and delayed-label follow-up.

Situational decisions

When drift is statistically significant but outcomes are unavailable: Report a monitoring signal and follow-up, not confirmed accuracy degradation.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring