ML experimentation · plan by default

/ml-train

Implement training with checkpoints and recorded configuration

Use for bounded training implementation/execution; ml-debug-training diagnoses a failed optimization process.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-train from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-train, /just-vibe ml-train and /jv:ml-train in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-train Plan checkpointed training on one GPU for at most six hours; do not provision compute.
edge · apply
/just-vibe:ml-train Implement checkpointed training that resumes mid-epoch without silently changing sample order.
blocked · inspect
/just-vibe:ml-train Plan training with no authorized hardware budget; do not provision or launch a run.

What the agent does

  1. Validate shapes and pipeline, run a small smoke test, record configuration/environment, train within bounds, checkpoint, and evaluate only the permitted validation protocol.
  2. For resumable training inventory model, optimizer, scheduler, scaler when used, step, RNG and sampler/data position; checkpoint atomically and compare interrupted versus uninterrupted continuation under declared tolerances.

Inputs

  • Model/task, data and objective relevant to the requested implementation. Actual training additionally needs data/splits, hardware, time/cost budget and artifact destination.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Reproducible training with checkpoint/resume and declared stopping criteria.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
Mode
Plan the training design; apply for requested pipeline implementation or an authorized bounded run.
Prerequisites
Dataset/split manifests, fixed objective/metric, environment/dependencies, baseline where applicable, and explicit compute limits. Record code revision, configuration, seeds, artifact paths, and resource use. Local smoke checks do not imply authorization for paid training. Never optimize on the held-out test set.

Expected output

  • Training code or run record with configuration, resource cap, artifact/state manifest, checkpoints, metrics and demonstrated resume conditions.

How the work is checked

  • Interrupted training can resume with identified state; NaNs or budget exhaustion stop clearly without a false completed result.

When to stop or clarify

  • No paid hardware provisioning implicitly. Do not label the best validation checkpoint as independently test-validated.

Handling missing context

Infer
Read framework, training entry point, loss/metric, split manifests and checkpoint conventions from supplied source.
Assume
In apply mode, implement requested code and tiny isolated smoke checks with existing tools, and otherwise propose them; leave unmeasured model quality explicit.
Ask
Ask for unresolved objective/data semantics before encoding them, and environment/resource limits before launching training or a search; implementation alone does not need a hardware purchase decision.

Technical guidance

Evidence
Inspect framework/version, shapes, loss semantics, device/dtype, optimizer, scheduler and data/sampler state.
Method
Run a bounded smoke batch, then checkpoint at a defined boundary including continuation state; load the training scenario for the actual framework.
Pitfall
Restoring weights alone is not exact resume; a seed alone does not guarantee deterministic kernels or data order.
Check
Compare interrupted and uninterrupted short runs under stated tolerances and verify atomic checkpoint recovery after an incomplete write.

Situational decisions

When only weights were saved: Treat loading as initialization unless all required continuation state is available; label the run a restart rather than exact resume.

When implementation is requested but a training run is not: Implement configuration, checkpoint and resume paths against local synthetic fixtures; ask for hardware, data access and resource limits only before a real training run.

When distributed training changes workers or accumulation boundaries: Verify sampler/metric aggregation and checkpoint ownership; declare approximate continuation if exact state cannot be restored.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring