ML experimentation · inspect by default

/ml-debug-training

Investigate exploding loss, unstable gradients, NaNs, or failure to learn

Use for NaNs, shape/device errors or non-learning; ml-error-analysis studies generalization failures.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-debug-training from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-debug-training, /just-vibe ml-debug-training and /jv:ml-debug-training in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:ml-debug-training Diagnose NaNs from this run's first failing batch and gradient logs.
edge · inspect
/just-vibe:ml-debug-training Debug a loss that becomes NaN only after mixed-precision updates.
blocked · inspect
/just-vibe:ml-debug-training Inspect saved training logs without retraining or guessing a learning-rate cure.

What the agent does

  1. Isolate a small batch and inspect its shapes, labels, scale, loss and gradients against expected scales, and check optimizer state.
  2. Locate the first non-finite value and compare optimizer updates with the intended objective.
  3. Propose overfit and gradient probes that test the leading cause, running them in apply mode.

Inputs

  • failing run logs/configuration, batches, model, and training symptom.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Numerical instability, shape/device issues, data/target mismatch, gradients, and failure to learn.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
Mode
Inspect; failing run logs/configuration, batches, model, and training symptom. Apply for requested bounded probes or a minimal corrective change.
Prerequisites
Dataset/split manifests, fixed objective/metric, environment/dependencies, baseline where applicable, and explicit compute limits. Record code revision, configuration, seeds, artifact paths, and resource use. Local smoke checks do not imply authorization for paid training. Never optimize on the held-out test set.

Expected output

  • Diagnosis with the first divergent tensor or step, hypothesis evidence, and a bounded corrective probe or, in apply mode, a minimal corrective change.

How the work is checked

  • NaNs are localized to their first source; an inability to overfit a tiny batch prompts pipeline investigation rather than a larger model.

When to stop or clarify

  • No unbounded retraining or random parameter changes. Preserve failed-run evidence and distinguish numerical repair from generalization improvement.

Handling missing context

Infer
Read framework, training entry point, loss/metric, split manifests and checkpoint conventions from supplied source.
Assume
In apply mode, implement requested code and tiny isolated smoke checks with existing tools, and otherwise propose them; leave unmeasured model quality explicit.
Ask
Ask for unresolved objective/data semantics before encoding them, and environment/resource limits before launching training or a search; implementation alone does not need a hardware purchase decision.

Technical guidance

Evidence
Capture the first divergent batch, activations/loss, gradient finiteness and parameter update.
Method
Check objective/target semantics and preprocessing before tuning; use a tiny-batch overfit probe to distinguish plumbing from generalization.
Pitfall
Switching architecture can conceal a detached graph, wrong target scale or optimizer that never updates parameters.
Check
Verify finite forward/backward values and actual parameter changes on the smallest failing input before a longer run.

Situational decisions

When a tiny-batch overfit probe fails: Investigate data/loss/gradient/update plumbing before larger architectures or more epochs.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring