ML data · inspect by default
/ml-labels
Inspect label definitions, noise, disagreement, and missing outcomes
Use for label construction and annotation quality; ml-leakage checks prediction-time information flow.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ml-labels from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ml-labels, /just-vibe ml-labels and /jv:ml-labels in Claude. See shortcut setup and context examples.
/just-vibe:ml-labels Audit how missing outcome follow-up and annotation disagreement affect labels./just-vibe:ml-labels Audit labels when missing follow-up was encoded as no failure./just-vibe:ml-labels Assess annotation policy without identifiable raw examples or relabeling permission.What the agent does
- Trace each label's source, construction, event horizon and maturity.
- Compare annotations and outcomes, distinguishing true negatives, unobserved outcomes, contradictory annotations and policy ambiguity.
- Propose adjudication and quality checks.
Inputs
- label definitions, annotation/outcome sources, timing, and permitted samples.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Label consistency, noise, disagreement, censoring, and missing outcomes.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Inspect; label definitions, annotation/outcome sources, timing, and permitted samples.
- Prerequisites
- Task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
Expected output
- Label audit with the label definition, maturity/coverage checks, concrete patterns with rates and denominators, reproducible disagreement examples and corrective options.
How the work is checked
- Unobserved outcomes are not automatically negative; conflicting annotations are tracked rather than silently overwritten.
When to stop or clarify
- Relabeling requires explicit policy and scope. Avoid exposing sensitive examples or claiming a single annotator is ground truth without justification.
Handling missing context
- Infer
- Read prediction moment, label horizon, entity/time keys, split policy and dataset provenance from the task and manifests.
- Assume
- Use explicit synthetic examples for design when raw data is unavailable; do not infer missing labels or fit preprocessing across held-out boundaries.
- Ask
- Ask when unresolved label timing, grouping or target semantics would change the split/features; do not demand a full dataset to explain the method.
Technical guidance
- Evidence
- Read labeling policy, event identity, annotator agreement, outcome window and availability timestamps.
- Method
- Separate absent, unresolved and negative labels; trace a disagreement to policy or observation error before changing it.
- Pitfall
- Majority vote can erase systematic ambiguity; labels recorded after prediction may be valid outcomes but unavailable for historical fitting.
- Check
- Hand-check boundary examples, censored cases and disagreement resolution; report which historical training rows were label-eligible.
Situational decisions
When annotators disagree on an ambiguous definition: Preserve disagreement, clarify policy and adjudicate within scope rather than silently majority-voting it away.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.