ML data · inspect by default

/ml-labels

Inspect label definitions, noise, disagreement, and missing outcomes

Use for label construction and annotation quality; ml-leakage checks prediction-time information flow.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ml-labels from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ml-labels, /just-vibe ml-labels and /jv:ml-labels in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:ml-labels Audit how missing outcome follow-up and annotation disagreement affect labels.
edge · inspect
/just-vibe:ml-labels Audit labels when missing follow-up was encoded as no failure.
blocked · inspect
/just-vibe:ml-labels Assess annotation policy without identifiable raw examples or relabeling permission.

What the agent does

  1. Trace each label's source, construction, event horizon and maturity.
  2. Compare annotations and outcomes, distinguishing true negatives, unobserved outcomes, contradictory annotations and policy ambiguity.
  3. Propose adjudication and quality checks.

Inputs

  • label definitions, annotation/outcome sources, timing, and permitted samples.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Label consistency, noise, disagreement, censoring, and missing outcomes.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Inspect; label definitions, annotation/outcome sources, timing, and permitted samples.
Prerequisites
Task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.

Expected output

  • Label audit with the label definition, maturity/coverage checks, concrete patterns with rates and denominators, reproducible disagreement examples and corrective options.

How the work is checked

  • Unobserved outcomes are not automatically negative; conflicting annotations are tracked rather than silently overwritten.

When to stop or clarify

  • Relabeling requires explicit policy and scope. Avoid exposing sensitive examples or claiming a single annotator is ground truth without justification.

Handling missing context

Infer
Read prediction moment, label horizon, entity/time keys, split policy and dataset provenance from the task and manifests.
Assume
Use explicit synthetic examples for design when raw data is unavailable; do not infer missing labels or fit preprocessing across held-out boundaries.
Ask
Ask when unresolved label timing, grouping or target semantics would change the split/features; do not demand a full dataset to explain the method.

Technical guidance

Evidence
Read labeling policy, event identity, annotator agreement, outcome window and availability timestamps.
Method
Separate absent, unresolved and negative labels; trace a disagreement to policy or observation error before changing it.
Pitfall
Majority vote can erase systematic ambiguity; labels recorded after prediction may be valid outcomes but unavailable for historical fitting.
Check
Hand-check boundary examples, censored cases and disagreement resolution; report which historical training rows were label-eligible.

Situational decisions

When annotators disagree on an ambiguous definition: Preserve disagreement, clarify policy and adjudicate within scope rather than silently majority-voting it away.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring