ML deployment · apply by default

/ml-serving

Implement online inference with validation and observable errors

Use to implement model service behavior; ml-rollout plans traffic transition.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-serving from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-serving, /just-vibe ml-serving and /jv:ml-serving in Claude. See shortcut setup and context examples.

Example · apply
/just-vibe:ml-serving Implement local inference with validation, readiness, and safe error handling.
edge · apply
/just-vibe:ml-serving Serve a model with bounded batch size and a failed startup load.
blocked · inspect
/just-vibe:ml-serving Design serving code without provisioning an endpoint or logging raw sensitive inputs.

What the agent does

  1. Define readiness for the correct artifact, input bounds, batching/concurrency and deadlines.
  2. Validate shapes and types before inference, enforce resource and time limits, map errors, and preserve the model version in responses and redacted telemetry.
  3. Test concurrent valid and invalid requests.

Inputs

  • model package, request/response contract, latency/resource constraints, and target runtime.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Online inference service code and isolated validation; live hosting requires a deployment request.
Writes
Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
Mode
Apply; model package, request/response contract, latency/resource constraints, and target runtime.
Prerequisites
Versioned model and preprocessing artifacts, input/output schema, runtime/dependencies, operating targets, and authorized environment. Validate artifact trust before loading formats that can execute code. Packaging or writing monitoring configuration does not deploy a model or enable a hosted service.

Expected output

  • Service and configuration with its serving contract, lifecycle, failure handling, operational checks and performance evidence if measured.

How the work is checked

  • Invalid shapes/types fail safely; readiness does not pass before the correct model is usable.

When to stop or clarify

  • No live endpoint provisioning or model registry mutation implicitly. Do not log raw sensitive inference inputs by default.

Handling missing context

Infer
Read artifact format/trust, preprocessing schema, serving runtime, compatibility and existing rollout controls.
Assume
Prepare packaging/configuration and isolated checks without treating them as a live deployment.
Ask
Resolve the target, rollback compatibility and operating limits before rollout or load generation; missing production access does not block packaging.

Technical guidance

Evidence
Resolve request schema, batching, concurrency, model lifecycle, device memory and timeout budget.
Method
Validate inputs before inference, bound queues and separate loading failures from invalid requests; define model identity per response/trace.
Pitfall
Unbounded batches or concurrent model copies can exhaust memory; returning a default prediction hides infrastructure failure.
Check
Exercise valid/invalid inputs, queue saturation, model-load failure and cancellation; verify bounded resources and typed errors.

Situational decisions

When model loading fails or an incompatible schema arrives: Fail readiness or return a typed request error without serving an unidentified fallback.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring