Data engineering · apply by default

/data-pipeline

Build ingestion or transformation with observable failures

Use for a transformation pipeline; data-incremental focuses on checkpoints and change processing.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select data-pipeline from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv data-pipeline, /just-vibe data-pipeline and /jv:data-pipeline in Claude. See shortcut setup and context examples.

Example · apply
/just-vibe:data-pipeline Implement isolated ingestion with observable failures and atomic partition writes.
edge · apply
/just-vibe:data-pipeline Build a pipeline interrupted between writing data and publishing its manifest.
blocked · inspect
/just-vibe:data-pipeline Design transformations with missing field semantics; block only the affected conversions.

What the agent does

  1. Establish stable source and output identity and keys, and validate inputs.
  2. Implement transformations validated with small hand-checked fixtures, and stage writes so completion markers follow durable output.
  3. Expose failures and test restart and bad-record handling.

Inputs

  • sources, transforms, destination, cadence, correctness criteria, and resource limits.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Ingestion/transformation implementation and isolated validation; activating production schedules requires that request.
Writes
Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
Mode
Apply; sources, transforms, destination, cadence, correctness criteria, and resource limits.
Prerequisites
Data source/version, schema/semantics, transformation code, permitted sampling scope, and storage/compute budget. Prefer aggregates and redacted samples; never upload datasets to external services implicitly. Record time zones and snapshot identity for reproducibility.

Expected output

  • Pipeline and configuration with its transform mapping, input/output reconciliation, completion protocol, failure accounting and operational instructions.

How the work is checked

  • A normal batch produces reconciled output; interrupted writes do not masquerade as a completed partition.

When to stop or clarify

  • No silent row dropping or unrequested data export. Missing semantics block affected transformations rather than guessed conversions.

Handling missing context

Infer
Inspect schema, source snapshot, transformation code, grain, time zones and permitted sample scope.
Assume
Use bounded synthetic or supplied samples when full data is unavailable; keep unknown values distinct from zero.
Ask
Resolve ambiguous entity/grain/time semantics before reconciliation or backfill; obtain missing data/compute limits only for the dependent scan or execution.

Technical guidance

Evidence
Trace source identity, transformations, sink transaction and checkpoint ownership.
Method
Make retries deterministic using stable keys and atomically aligned output/progress where possible; retain failed records with reasons.
Pitfall
Advancing a checkpoint before committing output silently loses records after a crash.
Check
Interrupt before/after sink commit, replay a batch and compare outputs to a hand-computed small fixture.

Situational decisions

When invalid records can be isolated without corrupting the batch: Quarantine with counts/reasons under the declared policy; never drop them silently.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring