ML deployment · inspect by default
/ml-parity
Check training preprocessing against production inference
Use to compare training and serving transformations; ml-drift compares populations over time and ml-leakage audits prediction-time information.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-parity from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-parity, /just-vibe ml-parity and /jv:ml-parity in Claude. See shortcut setup and context examples.
/just-vibe:ml-parity Compare training and serving transformations on these exact versioned inputs./just-vibe:ml-parity Diagnose high offline scores but poor serving due to reordered features./just-vibe:ml-parity Assess parity with missing training transforms; keep unobserved stages unknown.What the agent does
- Feed identical raw rows with aligned versions through each pipeline.
- Compare schema, feature names and order, transformations and model outputs at each boundary, and localize the first divergence against declared tolerances.
- Propose fixes, applying a requested one in apply mode.
Inputs
- training and serving pipelines/artifacts plus representative versioned inputs.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Preprocessing, feature order, types, defaults, model version, and numerical parity.
- Writes
- Inspect/plan: inspect or propose; save requested artifacts only. Apply: edit the requested local implementation and perform relevant bounded checks while preserving unrelated work. Live data changes, remote actions and paid jobs require their resolved target and existing session authorization.
- Mode
- Inspect; training and serving pipelines/artifacts plus representative versioned inputs. Apply for a requested preprocessing fix or parity regression fixtures.
- Prerequisites
- Versioned model and preprocessing artifacts, input/output schema, runtime/dependencies, operating targets, and authorized environment. Validate artifact trust before loading formats that can execute code. Packaging or writing monitoring configuration does not deploy a model or enable a hosted service.
Expected output
- Parity report with the first divergent boundary, aligned inputs and versions, semantic versus numerical differences and regression fixtures.
How the work is checked
- Reordered features are detected; expected floating-point variance is separated from semantic mismatch.
When to stop or clarify
- Missing training transforms or serving access limits coverage. Do not compare outputs from different model versions as a pure preprocessing test.
Handling missing context
- Infer
- Read artifact format/trust, preprocessing schema, serving runtime, compatibility and existing rollout controls.
- Assume
- Prepare packaging/configuration and isolated checks without treating them as a live deployment.
- Ask
- Resolve the target, rollback compatibility and operating limits before rollout or load generation; missing production access does not block packaging.
Technical guidance
- Evidence
- Compare fitted preprocessing, feature ordering, units, categorical vocabularies, missingness and runtime numerics.
- Method
- Send the same golden inputs through training transformation and packaged serving transformation before comparing predictions.
- Pitfall
- Equal tensor shape does not imply equal feature meaning; silently reordered columns can produce plausible wrong scores.
- Check
- Include missing, unseen, zero and boundary inputs and locate the first differing transform rather than comparing only final accuracy.
Situational decisions
When model/artifact versions differ: Resolve version identity before attributing output differences solely to preprocessing.
When the symptom is an offline-versus-production gap: Run one discriminating check per cause before concluding and report which causes each check excludes: recompute offline metrics on production-logged features for the same rows (skew, ml-parity), compare evaluation-window and production feature and prediction distributions (drift, ml-drift), and audit feature availability and split lineage (leakage, ml-leakage).
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.