ML experimentation · plan by default

/ml-tune

Design a bounded hyperparameter search with a fixed evaluation protocol

Use for a bounded search under a valid protocol; ml-ablation isolates component contribution.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-tune from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-tune, /just-vibe ml-tune and /jv:ml-tune in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-tune Plan at most 20 trials on validation PR-AUC; never tune on the test set.
edge · plan
/just-vibe:ml-tune Plan a search that tolerates failed trials under a strict GPU-hour cap.
blocked · inspect
/just-vibe:ml-tune Design tuning when compute is unavailable; do not fabricate winning hyperparameters.

What the agent does

  1. Freeze the search space, split, objective, trial and resource caps and the selection rule before searching.
  2. Select a search strategy with pruning and failure behavior, and log every trial, including failed and pruned ones.
  3. Compare candidates under equal evaluation conditions and choose by the predeclared validation criterion; keep the confirmation set untouched.

Inputs

  • baseline, search space, objective, fixed splits, trial/time/compute caps, and selection rule.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Bounded hyperparameter search; execution only within authorized resources.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
Mode
Plan a search; apply for requested search code or a run with explicit resource limits.
Prerequisites
Dataset/split manifests, fixed objective/metric, environment/dependencies, baseline where applicable, and explicit compute limits. Record code revision, configuration, seeds, artifact paths, and resource use. Local smoke checks do not imply authorization for paid training. Never optimize on the held-out test set.

Expected output

  • Search configuration and trial ledger with budgets consumed, the selected candidate, selection rationale, cost caveats and the untouched confirmation set.

How the work is checked

  • Trial failures remain visible; test labels never influence parameter selection.

When to stop or clarify

  • Stop at any resource cap and preserve partial results. Do not enlarge the search merely because no improvement appears.

Handling missing context

Infer
Read framework, training entry point, loss/metric, split manifests and checkpoint conventions from supplied source.
Assume
In apply mode, implement requested code and tiny isolated smoke checks with existing tools, and otherwise propose them; leave unmeasured model quality explicit.
Ask
Ask for unresolved objective/data semantics before encoding them, and environment/resource limits before launching training or a search; implementation alone does not need a hardware purchase decision.

Technical guidance

Evidence
Fix search space, metric direction, split, resource budget, pruning and selection rule.
Method
Track every trial including failures and choose using validation only; reserve the held-out test for final evaluation.
Pitfall
More trials can overfit the validation set; dropping failed runs understates cost and instability.
Check
Enforce trial/time caps and inspect selection provenance; evaluate the chosen configuration once under the reserved protocol.

Situational decisions

When tuning repeatedly consults held-out test results: Stop that selection loop and define fresh independent confirmation before reporting generalization.

When the request is for local preparation or implementation: Implement search-space validation, trial accounting and interruption handling locally; do not submit a search job without its dataset and resource limits.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring