ML data · apply by default
/ml-dataset-version
Record dataset identity, transformations, and provenance
Use to identify reproducible data/splits; data-lineage explains transformations.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-dataset-version from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-dataset-version, /just-vibe ml-dataset-version and /jv:ml-dataset-version in Claude. See shortcut setup and context examples.
/just-vibe:ml-dataset-version Write a provenance manifest for the identified dataset without copying raw records./just-vibe:ml-dataset-version Version a dataset whose remote contents can change at the same path./just-vibe:ml-dataset-version Plan versioning without permission for a full expensive hash scan.What the agent does
- Record source snapshot or content identity (immutable access references or hashes where feasible), transformation revision, schema, split membership and creation parameters, without storing secrets or raw private data.
- Verify each reference's accessibility and mark mutable boundaries.
Inputs
- dataset snapshot, transforms, source identifiers, and approved manifest location.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Reproducible identity/provenance metadata; no automatic copying or committing of raw data.
- Writes
- Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
- Mode
- Apply; dataset snapshot, transforms, source identifiers, and approved manifest location.
- Prerequisites
- Task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
Expected output
- Version manifest with the provenance chain, reconstruction instructions, accessibility check and known reproducibility limits or mutable boundaries.
How the work is checked
- Changed source content changes identity or is detected; a mutable URL alone is not presented as an immutable version.
When to stop or clarify
- Avoid expensive full hashing without a budget. Do not place confidential data or credentials in Git/manifests.
Handling missing context
- Infer
- Read prediction moment, label horizon, entity/time keys, split policy and dataset provenance from the task and manifests.
- Assume
- Use explicit synthetic examples for design when raw data is unavailable; do not infer missing labels or fit preprocessing across held-out boundaries.
- Ask
- Ask when unresolved label timing, grouping or target semantics would change the split/features; do not demand a full dataset to explain the method.
Technical guidance
- Evidence
- Inventory source snapshot IDs, transforms, schema, split membership and label-version policy.
- Method
- Create a manifest that identifies exact inputs and processing without embedding confidential rows; state mutable external sources explicitly.
- Pitfall
- A filename, seed or hash of only a sample cannot identify the full dataset.
- Check
- Reconstruct membership from the manifest or record inaccessible components; changed upstream revisions must change or qualify dataset identity.
Situational decisions
When only a mutable source URL exists: Mark identity provisional and define the snapshot/hash mechanism needed for reproducibility.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.