ML experimentation · inspect by default

/ml-experiments

Compare runs and check that data and evaluation conditions match

Use to compare recorded runs; ml-train produces a run and ml-report communicates validated conclusions.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-experiments from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-experiments, /just-vibe ml-experiments and /jv:ml-experiments in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:ml-experiments Compare these runs and flag different datasets or metric definitions.
edge · inspect
/just-vibe:ml-experiments Compare experiments that used different test periods and rounded headline scores.
blocked · inspect
/just-vibe:ml-experiments Inspect incomplete run exports without inventing missing metrics or costs.

What the agent does

  1. Reconcile dataset, split, code and config identities, metric definitions and denominators, and selection history.
  2. Inspect failed and missing runs, compare quality and resources with failed or pruned runs included in total resource accounting, and separate incompatible cohorts.
  3. For supplied binary/regression run exports, follow the experiments guide to import actual provider metadata and aligned prediction rows with dataset/split, code/model identity, preprocessing, seed and feature maps. Keep provider-reported metrics separate from recomputed metrics; do not log into a provider or train merely to import.
  4. Compare only fresh compatible task/dataset/split and row/target/slice identities. Suppress deltas when incompatible or stale. Explain overall and per-slice changes together, small denominators, missing dimensions, threshold changes, feature parity mismatches and temporal check coverage.
  5. Investigate an aggregate gain with a subgroup regression before making a recommendation. Supplied metadata is attributed evidence, matching feature maps do not execute preprocessing, and observed differences do not establish cause. Reconcile unexported/missing predictions and use project tooling for uncertainty, unsupported tasks or larger data.

Inputs

  • run IDs/artifacts, metric of interest, and comparison scope.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Compare existing experiments and their compatibility; no automatic reruns.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Inspect; run IDs/artifacts, metric of interest, and comparison scope.
Prerequisites
Dataset/split manifests, fixed objective/metric, environment/dependencies, baseline where applicable, and explicit compute limits. Record code revision, configuration, seeds, artifact paths, and resource use. Local smoke checks do not imply authorization for paid training. Never optimize on the held-out test set.

Expected output

  • Experiment comparison in comparable-run groups with quality/resource evidence, the strongest supported result, comparability gaps and missing metadata.
  • For imported exports: normalized runs, actual input hashes, recomputed aggregate/slice metrics, comparability and parity/temporal findings.

How the work is checked

  • Different test populations are not ranked as directly comparable; failed trials are not silently omitted from cost accounting.

When to stop or clarify

  • Missing metadata prevents strong conclusions. Do not select a winner solely from rounded headline scores.

Handling missing context

Infer
Read framework, training entry point, loss/metric, split manifests and checkpoint conventions from supplied source.
Assume
In apply mode, implement requested code and tiny isolated smoke checks with existing tools, and otherwise propose them; leave unmeasured model quality explicit.
Ask
Ask for unresolved objective/data semantics before encoding them, and environment/resource limits before launching training or a search; implementation alone does not need a hardware purchase decision.

Technical guidance

Evidence
Inspect run manifests, data/split identity, metric definitions, code, failures and selection history.
Method
Group only comparable runs and flag differences that change the question; keep cost and failure rates beside headline scores.
Pitfall
Same metric names may hide different denominators, positive classes or evaluation populations.
Check
Recompute a small metric from retained predictions and reject or qualify comparisons with incompatible protocols.

Situational decisions

When runs used different populations or metric definitions: Group them separately and propose a matched comparison instead of a misleading leaderboard.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring