Data engineering · inspect by default

/data-profile

Summarize distributions, missingness, duplicates, and suspicious values

Use for descriptive data inspection; data-quality evaluates declared rules.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select data-profile from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv data-profile, /just-vibe data-profile and /jv:data-profile in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:data-profile Profile missingness and duplicates in this bounded dataset sample.
edge · inspect
/just-vibe:data-profile Profile a time-partitioned dataset with sentinel zeros and missing recent partitions.
blocked · inspect
/just-vibe:data-profile Inspect metadata only when row access is unavailable; do not fabricate distributions.

What the agent does

  1. Inspect schema and volume before scanning.
  2. Select a bounded sample across relevant time and group strata, or an authorized aggregate scan, and state the selection limits.
  3. Compute summaries, distinguish nulls from sentinels and flag anomalies relative to declared semantics.

Inputs

  • dataset/snapshot, columns, sampling limit, and task context.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Distributions, missingness, uniqueness, duplicates, ranges, and suspicious values.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. data-pipeline or fix applies an accepted change.
Mode
Inspect; dataset/snapshot, columns, sampling limit, and task context.
Prerequisites
Data source/version, schema/semantics, transformation code, permitted sampling scope, and storage/compute budget. Prefer aggregates and redacted samples; never upload datasets to external services implicitly. Record time zones and snapshot identity for reproducibility.

Expected output

  • Profile with snapshot or sample identity, the sample-versus-full-scan distinction, summaries, anomaly examples, caveats and follow-up checks.

How the work is checked

  • Missing and sentinel values are separated; sample statistics are not presented as exact full-population counts.

When to stop or clarify

  • Unknown volume triggers size inspection before scanning. Do not print sensitive row-level data unnecessarily.

Handling missing context

Infer
Inspect schema, source snapshot, transformation code, grain, time zones and permitted sample scope.
Assume
Use bounded synthetic or supplied samples when full data is unavailable; keep unknown values distinct from zero.
Ask
Resolve ambiguous entity/grain/time semantics before reconciliation or backfill; obtain missing data/compute limits only for the dependent scan or execution.

Technical guidance

Evidence
Establish snapshot, row grain, sample method, units, timezones, sensitive fields and denominator.
Method
Report missingness, duplicates and distributions by meaningful group; distinguish sample observations from whole-population claims.
Pitfall
Converting numeric-looking IDs or imputing during profiling silently changes evidence.
Check
Reconcile row counts and missing-value definitions and inspect bounded anomalies without exporting private rows.

Situational decisions

When the sample is convenience-based or filtered: Label its population and avoid extrapolating exact counts or representativeness.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring