ML experimentation · plan by default

/ml-baseline

Establish simple, reproducible baselines

Use for a first comparable benchmark; ml-tune searches hyperparameters after protocol validity.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-baseline from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-baseline, /just-vibe ml-baseline and /jv:ml-baseline in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ml-baseline Plan reproducible naive and simple-model baselines under the fixed split.
edge · plan
/just-vibe:ml-baseline Establish a baseline when class imbalance makes accuracy misleading.
blocked · inspect
/just-vibe:ml-baseline Plan a baseline with unresolved label maturity; do not claim trustworthy scores.

What the agent does

  1. Fix the evaluation protocol: the same splits and leakage-safe preprocessing fit boundaries for every compared model.
  2. Include a task-appropriate constant or rule baseline and a simple model; run bounded training when authorized.
  3. Report each score with its uncertainty, recording resources alongside quality.

Inputs

  • task, training/validation split, metric, and run budget. Explicit baseline-run requests select apply.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Simple reproducible reference models, including a naive/rule-based comparator.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
Mode
Plan a baseline comparison; apply for requested baseline code or a bounded experiment.
Prerequisites
Dataset/split manifests, fixed objective/metric, environment/dependencies, baseline where applicable, and explicit compute limits. Record code revision, configuration, seeds, artifact paths, and resource use. Local smoke checks do not imply authorization for paid training. Never optimize on the held-out test set.

Expected output

  • Baseline configuration or code with data/split identity, metric definitions, the baseline/model comparison and a compute record when executed.

How the work is checked

  • A naive comparator is included; all candidates use identical permitted evaluation data.

When to stop or clarify

  • Do not begin extensive model search. Missing label/split validity blocks trustworthy benchmark claims even if training technically runs.

Handling missing context

Infer
Read framework, training entry point, loss/metric, split manifests and checkpoint conventions from supplied source.
Assume
In apply mode, implement requested code and tiny isolated smoke checks with existing tools, and otherwise propose them; leave unmeasured model quality explicit.
Ask
Ask for unresolved objective/data semantics before encoding them, and environment/resource limits before launching training or a search; implementation alone does not need a hardware purchase decision.

Technical guidance

Evidence
Resolve task, split, target availability, metric and naive prediction policy.
Method
Evaluate a constant/last-value or other task-appropriate naive model and a simple fitted model under identical preprocessing and data boundaries.
Pitfall
A strong model with a different split is not a fair baseline; test-set selection makes later comparisons optimistic.
Check
Preserve predictions, denominators and protocol identity and verify the baseline handles the same missing/rare cases as candidate models.

Situational decisions

When the baseline wins or matches within uncertainty: Keep the simpler option and identify the missing signal before expanding model search.

When the request is for local preparation or implementation: Implement the requested simple baseline and evaluation interface; use a small labeled fixture to check the pipeline without claiming predictive quality.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring