ML data · plan by default
/ml-split
Design splits respecting time, groups, entities, and dependencies
Use to design evaluation partitions matching deployment; ml-leakage audits actual contamination evidence.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-split from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-split, /just-vibe ml-split and /jv:ml-split in Claude. See shortcut setup and context examples.
/just-vibe:ml-split Design time/group splits for overlapping machine sensor windows./just-vibe:ml-split Split overlapping windows for forecasting on known machines and evaluate unseen machines separately./just-vibe:ml-split Plan a split with unknown label horizon; do not invent a universal gap duration.What the agent does
- Define deployment population, row/entity dependence, prediction times, model-fit cutoff, outcome horizon and label availability. Decide whether the question concerns future observations of known entities, unseen entities, or separate evaluations of both.
- Specify exact train/validation/test endpoints and an as-of snapshot. Compare timezone-aware instants rather than timestamp strings. Exclude observations that would not yet exist, and distinguish event time from ingestion and label-observation time.
- Determine training eligibility before fitting preprocessing: an otherwise training-period row with an immature outcome must not teach the historical model. Retain unknown held-out labels as unknown where the evaluation contract requires it; zero is a valid observed label.
- Derive group separation, purging or gaps from the stated deployment question and actual dependence/availability evidence. Record deterministic membership and exclusions with reasons; do not invent a universal embargo duration.
- Verify exact boundaries, label maturity, duplicates, relevant group overlap, empty partitions and transform fit membership. Report the number of rows and number with evaluable labels separately.
Inputs
- task, entity/group/time dependencies, deployment regime, and dataset version.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Train/validation/test membership and fitting boundaries; write manifests only when requested.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Plan; task, entity/group/time dependencies, deployment regime, and dataset version.
- Prerequisites
- Task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
Expected output
- Deployment question, snapshot and split/eligibility rules, deterministic membership/exclusion evidence, transform fit population and overlap/maturity checks.
How the work is checked
- No row, label or learned transform contains information unavailable at the simulated fit/prediction moment. Unknown outcomes stay unknown, and permitted known-entity overlap is distinguished from forbidden leakage.
When to stop or clarify
- Do not use random splitting by habit or repeatedly tune the split to improve scores. Document unsupported generalization claims.
Handling missing context
- Infer
- Read prediction moment, label horizon, entity/time keys, split policy and dataset provenance from the task and manifests.
- Assume
- Use explicit synthetic examples for design when raw data is unavailable; do not infer missing labels or fit preprocessing across held-out boundaries.
- Ask
- Ask when unresolved label timing, grouping or target semantics would change the split/features; do not demand a full dataset to explain the method.
Technical guidance
- Evidence
- Inspect deployment question, time/order, repeated entities, overlap and label maturity.
- Method
- Choose temporal/group boundaries that match known-entity versus unseen-entity deployment; derive gaps from actual information overlap.
- Pitfall
- Random splits can leak repeated entities, while universal group holdout can test a different task than the intended deployment.
- Check
- Assert disjoint required identities and training-time availability; report the exact generalization question and unresolved provenance.
Situational decisions
When the model will serve both known and unseen entities: Define separate evaluation questions instead of asserting one grouping rule answers both.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.