ML deployment · plan by default

/ml-inference-perf

Measure latency, throughput, memory, and optimization tradeoffs

Use for latency/throughput/resource benchmarking; ml-training-cost covers training.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-inference-perf from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-inference-perf, /just-vibe ml-inference-perf and /jv:ml-inference-perf in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-inference-perf Plan a bounded benchmark for cold/warm latency and throughput at fixed quality.
edge · plan
/just-vibe:ml-inference-perf Benchmark batched inference under a latency deadline without hiding warmup cost.
blocked · inspect
/just-vibe:ml-inference-perf Plan a benchmark with no authorized hardware or production traffic budget.

What the agent does

  1. Specify hardware, precision, batch/concurrency and payload distribution as comparable benchmark conditions.
  2. Separate load and warmup from steady state, and measure bounded authorized workloads, including tail behavior within caps.
  3. Identify bottlenecks and check quality after any optimization.

Inputs

  • model/runtime, input distribution, concurrency, hardware, quality floor, and benchmark budget.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Latency, throughput, memory, batching, warmup, and optimization tradeoffs.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
Mode
Inspect supplied performance evidence or plan measurement; apply for requested benchmark code or bounded runs.
Prerequisites
Versioned model and preprocessing artifacts, input/output schema, runtime/dependencies, operating targets, and authorized environment. Validate artifact trust before loading formats that can execute code. Packaging or writing monitoring configuration does not deploy a model or enable a hosted service.

Expected output

  • Benchmark protocol or results with conditions, sample size, cold/warm/tail metrics and quality comparison, and an optimization recommendation or authorized patch.

How the work is checked

  • Tail latency and throughput are reported under stated load; quantization gains include a quality comparison.

When to stop or clarify

  • No production load or paid hardware implicitly. A single warm request cannot establish capacity.

Handling missing context

Infer
Read artifact format/trust, preprocessing schema, serving runtime, compatibility and existing rollout controls.
Assume
Prepare packaging/configuration and isolated checks without treating them as a live deployment.
Ask
Resolve the target, rollback compatibility and operating limits before rollout or load generation; missing production access does not block packaging.

Technical guidance

Evidence
Measure preprocessing, transfer, model compute, postprocessing, batching and queue time with representative inputs.
Method
Compare latency distribution, throughput, memory and quality at the actual workload/concurrency; synchronize device timing where required.
Pitfall
Timing asynchronous GPU dispatch without synchronization underreports work; throughput gains may violate tail-latency requirements.
Check
Warm up deliberately, retain cold-start evidence and verify optimized predictions against reference tolerances and quality constraints.

Situational decisions

When quantization or batching improves speed: Re-evaluate quality, memory and latency under the same workload before accepting it.

When the request is for local preparation or implementation: Write the requested benchmark with warmup, workload and correctness controls; report speed claims only for measured hardware and inputs.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring