ML deployment · plan by default
/ml-monitor
Define operational and model-quality monitoring, including delayed labels
Use to design or implement requested ML telemetry; ops-alerts designs response-worthy alert behavior.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-monitor from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-monitor, /just-vibe ml-monitor and /jv:ml-monitor in Claude. See shortcut setup and context examples.
/just-vibe:ml-monitor Design quality monitoring with delayed labels and model-version separation./just-vibe:ml-monitor Monitor a model whose outcomes arrive thirty days after prediction./just-vibe:ml-monitor Plan monitoring without activating alerts or claiming an ongoing watcher exists.What the agent does
- Separate service, feature, prediction and delayed-outcome signals, distinguishing leading signals from outcome metrics.
- Define stable joins, label-lag windows and model-version attribution, and choose thresholds and runbook actions.
- Implement only requested instrumentation or configuration.
Inputs
- serving/batch system, model objectives, telemetry, delayed-label process, and response ownership.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Operational health, data quality, model performance, and label-arrival monitoring.
- Writes
- Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
- Mode
- Plan monitoring; apply for requested instrumentation or configuration in the identified environment.
- Prerequisites
- Versioned model and preprocessing artifacts, input/output schema, runtime/dependencies, operating targets, and authorized environment. Validate artifact trust before loading formats that can execute code. Packaging or writing monitoring configuration does not deploy a model or enable a hosted service.
Expected output
- Monitoring design or code with the signal/window/join/threshold/action contract, metric definitions, privacy controls and delayed-label and alert tests.
How the work is checked
- Delayed labels do not make recent unlabeled cases look correct; model-version changes remain distinguishable.
When to stop or clarify
- Do not activate external alerts or promise ongoing observation without an actual authorized runtime/scheduler.
Handling missing context
- Infer
- Read artifact format/trust, preprocessing schema, serving runtime, compatibility and existing rollout controls.
- Assume
- Prepare packaging/configuration and isolated checks without treating them as a live deployment.
- Ask
- Resolve the target, rollback compatibility and operating limits before rollout or load generation; missing production access does not block packaging.
Technical guidance
- Evidence
- Trace prediction IDs to model/data versions, outcomes, label delay, errors and operational measurements.
- Method
- Define quality windows based on matured labels and connect each alert to a diagnosis/recovery action.
- Pitfall
- Missing outcomes can bias observed accuracy; unchanged inputs do not prove unchanged target relationships.
- Check
- Simulate delayed/missing labels, a model revision and an outage; verify denominators, routing and unknown-quality states.
Situational decisions
When recent predictions have not had time to receive labels: Exclude them from matured quality denominators and report pending follow-up separately.
When the request is for local preparation or implementation: Implement drift/quality monitors against synthetic windows; treat unlabeled proxy drift as a signal, not demonstrated accuracy loss.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.