ML deployment · inspect by default

/ml-parity

Check training preprocessing against production inference

Use to compare training and serving transformations; ml-drift compares populations over time and ml-leakage audits prediction-time information.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-parity from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-parity, /just-vibe ml-parity and /jv:ml-parity in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:ml-parity Compare training and serving transformations on these exact versioned inputs.
edge · inspect
/just-vibe:ml-parity Diagnose high offline scores but poor serving due to reordered features.
blocked · inspect
/just-vibe:ml-parity Assess parity with missing training transforms; keep unobserved stages unknown.

What the agent does

  1. Feed identical raw rows with aligned versions through each pipeline.
  2. Compare schema, feature names and order, transformations and model outputs at each boundary, and localize the first divergence against declared tolerances.
  3. Propose fixes, applying a requested one in apply mode.

Inputs

  • training and serving pipelines/artifacts plus representative versioned inputs.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Preprocessing, feature order, types, defaults, model version, and numerical parity.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: edit the requested local implementation and perform relevant bounded checks while preserving unrelated work. Live data changes, remote actions and paid jobs require their resolved target and existing session authorization.
Mode
Inspect; training and serving pipelines/artifacts plus representative versioned inputs. Apply for a requested preprocessing fix or parity regression fixtures.
Prerequisites
Versioned model and preprocessing artifacts, input/output schema, runtime/dependencies, operating targets, and authorized environment. Validate artifact trust before loading formats that can execute code. Packaging or writing monitoring configuration does not deploy a model or enable a hosted service.

Expected output

  • Parity report with the first divergent boundary, aligned inputs and versions, semantic versus numerical differences and regression fixtures.

How the work is checked

  • Reordered features are detected; expected floating-point variance is separated from semantic mismatch.

When to stop or clarify

  • Missing training transforms or serving access limits coverage. Do not compare outputs from different model versions as a pure preprocessing test.

Handling missing context

Infer
Read artifact format/trust, preprocessing schema, serving runtime, compatibility and existing rollout controls.
Assume
Prepare packaging/configuration and isolated checks without treating them as a live deployment.
Ask
Resolve the target, rollback compatibility and operating limits before rollout or load generation; missing production access does not block packaging.

Technical guidance

Evidence
Compare fitted preprocessing, feature ordering, units, categorical vocabularies, missingness and runtime numerics.
Method
Send the same golden inputs through training transformation and packaged serving transformation before comparing predictions.
Pitfall
Equal tensor shape does not imply equal feature meaning; silently reordered columns can produce plausible wrong scores.
Check
Include missing, unseen, zero and boundary inputs and locate the first differing transform rather than comparing only final accuracy.

Situational decisions

When model/artifact versions differ: Resolve version identity before attributing output differences solely to preprocessing.

When the symptom is an offline-versus-production gap: Run one discriminating check per cause before concluding and report which causes each check excludes: recompute offline metrics on production-logged features for the same rows (skew, ml-parity), compare evaluation-window and production feature and prediction distributions (drift, ml-drift), and audit feature availability and split lineage (leakage, ml-leakage).

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring