ML experimentation · plan by default
/ml-reproduce
Reproduce a result from code, data, and configuration
Use to repeat a specified run; ml-baseline defines a new benchmark.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-reproduce from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-reproduce, /just-vibe ml-reproduce and /jv:ml-reproduce in Claude. See shortcut setup and context examples.
/just-vibe:ml-reproduce Plan reproducing this result with exact artifact identities and a two-hour budget./just-vibe:ml-reproduce Reproduce a GPU run on another supported device with explicit tolerances./just-vibe:ml-reproduce Assess reproducibility when the original dataset snapshot is missing.What the agent does
- Resolve exact data, artifact, code and dependency identities.
- Reconstruct preprocessing and evaluation, and declare nondeterminism tolerances before execution.
- Run authorized bounded work, compare outputs and metrics within the declared tolerance, and isolate deviations.
Inputs
- claimed result, code/data/artifact identities, environment, tolerance, and budget.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Re-run a defined result with explicit reproducibility criteria.
- Writes
- Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
- Mode
- Inspect reproduction evidence or plan the attempt; apply for requested reproduction code or bounded execution.
- Prerequisites
- Dataset/split manifests, fixed objective/metric, environment/dependencies, baseline where applicable, and explicit compute limits. Record code revision, configuration, seeds, artifact paths, and resource use. Local smoke checks do not imply authorization for paid training. Never optimize on the held-out test set.
Expected output
- Reproduction record and manifest with matched and different conditions, the measured result, deviations and discrepancy analysis against the justified tolerance.
How the work is checked
- Data/version mismatch is discovered before claiming reproduction; nondeterministic hardware differences use declared tolerances.
When to stop or clarify
- Missing original assets may make exact reproduction impossible. Do not silently substitute a different dataset or model and call it reproduced.
Handling missing context
- Infer
- Read framework, training entry point, loss/metric, split manifests and checkpoint conventions from supplied source.
- Assume
- In apply mode, implement requested code and tiny isolated smoke checks with existing tools, and otherwise propose them; leave unmeasured model quality explicit.
- Ask
- Ask for unresolved objective/data semantics before encoding them, and environment/resource limits before launching training or a search; implementation alone does not need a hardware purchase decision.
Technical guidance
- Evidence
- Resolve code revision, dependencies, artifacts, dataset access, hardware and claimed tolerance.
- Method
- Recreate the stated protocol; document each unavoidable substitution and whether it affects exact reproduction or only qualitative comparison.
- Pitfall
- Matching a seed or top-line metric does not establish the same data, selection process or training trajectory.
- Check
- Compare artifacts and outputs under declared tolerances and preserve failures or inaccessible inputs as limits on the claim.
Situational decisions
When original assets are unavailable and substitutes are necessary: Label the result a reimplementation or approximate reproduction and list each substitution.
When the request is for local preparation or implementation: Create the requested environment/configuration and synthetic smoke path; separate pipeline reproduction from matching a reported scientific result.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.