ML data · apply by default

/ml-dataset-version

Record dataset identity, transformations, and provenance

Use to identify reproducible data/splits; data-lineage explains transformations.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-dataset-version from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-dataset-version, /just-vibe ml-dataset-version and /jv:ml-dataset-version in Claude. See shortcut setup and context examples.

Example · apply
/just-vibe:ml-dataset-version Write a provenance manifest for the identified dataset without copying raw records.
edge · apply
/just-vibe:ml-dataset-version Version a dataset whose remote contents can change at the same path.
blocked · inspect
/just-vibe:ml-dataset-version Plan versioning without permission for a full expensive hash scan.

What the agent does

  1. Record source snapshot or content identity (immutable access references or hashes where feasible), transformation revision, schema, split membership and creation parameters, without storing secrets or raw private data.
  2. Verify each reference's accessibility and mark mutable boundaries.

Inputs

  • dataset snapshot, transforms, source identifiers, and approved manifest location.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Reproducible identity/provenance metadata; no automatic copying or committing of raw data.
Writes
Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
Mode
Apply; dataset snapshot, transforms, source identifiers, and approved manifest location.
Prerequisites
Task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.

Expected output

  • Version manifest with the provenance chain, reconstruction instructions, accessibility check and known reproducibility limits or mutable boundaries.

How the work is checked

  • Changed source content changes identity or is detected; a mutable URL alone is not presented as an immutable version.

When to stop or clarify

  • Avoid expensive full hashing without a budget. Do not place confidential data or credentials in Git/manifests.

Handling missing context

Infer
Read prediction moment, label horizon, entity/time keys, split policy and dataset provenance from the task and manifests.
Assume
Use explicit synthetic examples for design when raw data is unavailable; do not infer missing labels or fit preprocessing across held-out boundaries.
Ask
Ask when unresolved label timing, grouping or target semantics would change the split/features; do not demand a full dataset to explain the method.

Technical guidance

Evidence
Inventory source snapshot IDs, transforms, schema, split membership and label-version policy.
Method
Create a manifest that identifies exact inputs and processing without embedding confidential rows; state mutable external sources explicitly.
Pitfall
A filename, seed or hash of only a sample cannot identify the full dataset.
Check
Reconstruct membership from the manifest or record inaccessible components; changed upstream revisions must change or qualify dataset identity.

Situational decisions

When only a mutable source URL exists: Mark identity provisional and define the snapshot/hash mechanism needed for reproducibility.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring