ML data · inspect by default
/ml-leakage
Find target leakage, temporal leakage, and split contamination
Use to audit demonstrated information leakage; ml-split designs the evaluation protocol, and ml-parity or ml-drift own train/serve skew and population change.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-leakage from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-leakage, /just-vibe ml-leakage and /jv:ml-leakage in Claude. See shortcut setup and context examples.
/just-vibe:ml-leakage Audit churn features for values unavailable 30 days before cancellation./just-vibe:ml-leakage Audit overlapping windows whose metadata does not prove shared sensor values or outcome events./just-vibe:ml-leakage Review lineage with missing preprocessing code and event IDs; leave unsupported claims unknown.What the agent does
- Identify the prediction moment, label horizon, feature availability, split membership and intended deployment population from supplied evidence. Mark missing definitions as unknown.
- Trace suspicious features to when their values could actually have been available. Distinguish future-derived values from valid point-in-time historical features.
- Inspect preprocessing fit/transform boundaries when code or fit records exist. Their absence means unverified, not proof of correct or incorrect fitting.
- For each finding record the exact observed rows or source paths, the conclusion those observations support, and any assumptions needed for a stronger conclusion. Separate confirmed defects, conditional risks and missing evidence in the report.
- Quantify entity, interval and outcome-horizon overlap. Do not infer identical raw measurements or a shared outcome event from metadata alone. Determine group separation from whether deployment targets known entities, new entities or new groups.
- Check label maturity against the simulated model-fit and prediction times. State the historical-deployment assumption when applying temporal cutoffs or an embargo; choose gaps from actual availability and overlap instead of a universal duration.
- Before delivering, check every claim labeled proven against its cited evidence. Correct unsupported absolutes, including assertions that all scores are invalid or a split is always wrong. Identify which scores would be affected under which assumptions, and require re-evaluation after confirmed leakage is corrected.
- Build a compact evidence ledger: field or row, availability time, prediction/fit time, observed violation, affected score and assumptions; keep overlap metadata separate from shared measurements or events.
Inputs
- task/prediction moment, features, preprocessing, labels, and split lineage.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Target proxies, future information, cross-split fitting, duplicates, and entity contamination.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Inspect; task/prediction moment, features, preprocessing, labels, and split lineage.
- Prerequisites
- Task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
Expected output
- Evidence-backed leakage audit and finding ledger separating confirmed defects, conditional risks and unknowns; each finding names the supporting rows or source, assumptions, the bounded affected evaluation and the correction or missing evidence.
How the work is checked
- Future-only features and immature training labels are flagged against the stated prediction/fit times; cross-split preprocessing fitting is detected only when source or fit history establishes it.
- Metadata-only overlap is quantified without asserting identical sensor values or a shared failure event. Group separation is conditional on the intended deployment population.
- Missing labels remain unknown outcomes, unavailable pipeline evidence remains unverified, and score invalidation is limited to affected evaluation assumptions rather than invented results.
When to stop or clarify
- Do not claim absence of leakage when provenance is missing. Remediation must invalidate affected scores rather than preserve misleading results.
Handling missing context
- Infer
- Read prediction moment, label horizon, entity/time keys, split policy and dataset provenance from the task and manifests.
- Assume
- Use explicit synthetic examples for design when raw data is unavailable; do not infer missing labels or fit preprocessing across held-out boundaries.
- Ask
- Ask when unresolved label timing, grouping or target semantics would change the split/features; do not demand a full dataset to explain the method.
Technical guidance
- Evidence
- Trace suspicious features, fit transforms, revisions, event time, availability time and split membership.
- Method
- Resolve the latest visible record version before applying historical windows; fit learned transforms within each training fold.
- Pitfall
- Correlation or overlapping metadata alone does not prove leakage; filtering versions before selecting the visible revision can resurrect stale data.
- Check
- Test late correction, boundary timestamps, duplicate versions and labels unavailable at fit time; state precisely which evaluation is invalidated.
Situational decisions
When no raw measurements, event IDs or fitting history establish dependence: Report conditional risk or unknown, not proven shared events, mandatory gap length or universal score invalidity.
When the symptom is an offline-versus-production gap: Run one discriminating check per cause before concluding and report which causes each check excludes: recompute offline metrics on production-logged features for the same rows (skew, ml-parity), compare evaluation-window and production feature and prediction distributions (drift, ml-drift), and audit feature availability and split lineage (leakage, ml-leakage).
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.