ML experimentation · inspect by default

/ml-training-cost

Profile training time, memory, and resource bottlenecks

Use to analyze training resource use; ml-inference-perf measures deployed prediction work.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-training-cost from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-training-cost, /just-vibe ml-training-cost and /jv:ml-training-cost in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:ml-training-cost Analyze the supplied profile for loading, compute, and checkpoint bottlenecks.
edge · inspect
/just-vibe:ml-training-cost Diagnose an idle accelerator stalled by data loading.
blocked · inspect
/just-vibe:ml-training-cost Estimate training effort from sparse logs without inventing cloud prices.

What the agent does

  1. Separate startup, data loading, host-to-device transfer, compute, synchronization and checkpoint time.
  2. Relate batch and resource utilization to the same quality target and workload, identify bottlenecks, and propose measured optimizations or bounded profiling.

Inputs

  • training traces, hardware, workload, budget, and cost-rate evidence.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Time, memory, utilization, data loading, and cost bottlenecks.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Inspect; training traces, hardware, workload, budget, and cost-rate evidence.
Prerequisites
Dataset/split manifests, fixed objective/metric, environment/dependencies, baseline where applicable, and explicit compute limits. Record code revision, configuration, seeds, artifact paths, and resource use. Local smoke checks do not imply authorization for paid training. Never optimize on the held-out test set.

Expected output

  • Timing/resource breakdown with unit-cost assumptions and any verified price basis, and prioritized, bounded improvement experiments.

How the work is checked

  • Idle accelerator time caused by input loading is recognized; local elapsed time is not converted into a fabricated cloud bill.

When to stop or clarify

  • A short local profile is bounded local execution; new training runs and paid or shared compute need a resolved target and budget. Do not claim speedup without equivalent model quality and workload comparison.

Handling missing context

Infer
Read framework, training entry point, loss/metric, split manifests and checkpoint conventions from supplied source.
Assume
In apply mode, implement requested code and tiny isolated smoke checks with existing tools, and otherwise propose them; leave unmeasured model quality explicit.
Ask
Ask for unresolved objective/data semantics before encoding them, and environment/resource limits before launching training or a search; implementation alone does not need a hardware purchase decision.

Technical guidance

Evidence
Measure data loading, compute, synchronization, memory peaks, checkpointing and failed trials.
Method
Profile a representative bounded run; separate throughput from cost per completed useful result and use verified dated rates for money.
Pitfall
GPU utilization alone can hide pipeline stalls; mixed precision or larger batches can change convergence and effective optimization.
Check
Compare end-to-end runtime, peak memory and model-quality protocol under matched conditions, including warmup and failures.

Situational decisions

When throughput improves by changing effective batch or precision: Compare convergence/quality and total time-to-target before claiming a useful speedup.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring