ML experimentation · plan by default

/ml-reproduce

Reproduce a result from code, data, and configuration

Use to repeat a specified run; ml-baseline defines a new benchmark.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-reproduce from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-reproduce, /just-vibe ml-reproduce and /jv:ml-reproduce in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-reproduce Plan reproducing this result with exact artifact identities and a two-hour budget.
edge · plan
/just-vibe:ml-reproduce Reproduce a GPU run on another supported device with explicit tolerances.
blocked · inspect
/just-vibe:ml-reproduce Assess reproducibility when the original dataset snapshot is missing.

What the agent does

  1. Resolve exact data, artifact, code and dependency identities.
  2. Reconstruct preprocessing and evaluation, and declare nondeterminism tolerances before execution.
  3. Run authorized bounded work, compare outputs and metrics within the declared tolerance, and isolate deviations.

Inputs

  • claimed result, code/data/artifact identities, environment, tolerance, and budget.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Re-run a defined result with explicit reproducibility criteria.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
Mode
Inspect reproduction evidence or plan the attempt; apply for requested reproduction code or bounded execution.
Prerequisites
Dataset/split manifests, fixed objective/metric, environment/dependencies, baseline where applicable, and explicit compute limits. Record code revision, configuration, seeds, artifact paths, and resource use. Local smoke checks do not imply authorization for paid training. Never optimize on the held-out test set.

Expected output

  • Reproduction record and manifest with matched and different conditions, the measured result, deviations and discrepancy analysis against the justified tolerance.

How the work is checked

  • Data/version mismatch is discovered before claiming reproduction; nondeterministic hardware differences use declared tolerances.

When to stop or clarify

  • Missing original assets may make exact reproduction impossible. Do not silently substitute a different dataset or model and call it reproduced.

Handling missing context

Infer
Read framework, training entry point, loss/metric, split manifests and checkpoint conventions from supplied source.
Assume
In apply mode, implement requested code and tiny isolated smoke checks with existing tools, and otherwise propose them; leave unmeasured model quality explicit.
Ask
Ask for unresolved objective/data semantics before encoding them, and environment/resource limits before launching training or a search; implementation alone does not need a hardware purchase decision.

Technical guidance

Evidence
Resolve code revision, dependencies, artifacts, dataset access, hardware and claimed tolerance.
Method
Recreate the stated protocol; document each unavoidable substitution and whether it affects exact reproduction or only qualitative comparison.
Pitfall
Matching a seed or top-line metric does not establish the same data, selection process or training trajectory.
Check
Compare artifacts and outputs under declared tolerances and preserve failures or inaccessible inputs as limits on the claim.

Situational decisions

When original assets are unavailable and substitutes are necessary: Label the result a reimplementation or approximate reproduction and list each substitution.

When the request is for local preparation or implementation: Create the requested environment/configuration and synthetic smoke path; separate pipeline reproduction from matching a reported scientific result.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring