Data engineering · apply by default
/data-pipeline
Build ingestion or transformation with observable failures
Use for a transformation pipeline; data-incremental focuses on checkpoints and change processing.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select data-pipeline from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv data-pipeline, /just-vibe data-pipeline and /jv:data-pipeline in Claude. See shortcut setup and context examples.
/just-vibe:data-pipeline Implement isolated ingestion with observable failures and atomic partition writes./just-vibe:data-pipeline Build a pipeline interrupted between writing data and publishing its manifest./just-vibe:data-pipeline Design transformations with missing field semantics; block only the affected conversions.What the agent does
- Establish stable source and output identity and keys, and validate inputs.
- Implement transformations validated with small hand-checked fixtures, and stage writes so completion markers follow durable output.
- Expose failures and test restart and bad-record handling.
Inputs
- sources, transforms, destination, cadence, correctness criteria, and resource limits.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Ingestion/transformation implementation and isolated validation; activating production schedules requires that request.
- Writes
- Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
- Mode
- Apply; sources, transforms, destination, cadence, correctness criteria, and resource limits.
- Prerequisites
- Data source/version, schema/semantics, transformation code, permitted sampling scope, and storage/compute budget. Prefer aggregates and redacted samples; never upload datasets to external services implicitly. Record time zones and snapshot identity for reproducibility.
Expected output
- Pipeline and configuration with its transform mapping, input/output reconciliation, completion protocol, failure accounting and operational instructions.
How the work is checked
- A normal batch produces reconciled output; interrupted writes do not masquerade as a completed partition.
When to stop or clarify
- No silent row dropping or unrequested data export. Missing semantics block affected transformations rather than guessed conversions.
Handling missing context
- Infer
- Inspect schema, source snapshot, transformation code, grain, time zones and permitted sample scope.
- Assume
- Use bounded synthetic or supplied samples when full data is unavailable; keep unknown values distinct from zero.
- Ask
- Resolve ambiguous entity/grain/time semantics before reconciliation or backfill; obtain missing data/compute limits only for the dependent scan or execution.
Technical guidance
- Evidence
- Trace source identity, transformations, sink transaction and checkpoint ownership.
- Method
- Make retries deterministic using stable keys and atomically aligned output/progress where possible; retain failed records with reasons.
- Pitfall
- Advancing a checkpoint before committing output silently loses records after a crash.
- Check
- Interrupt before/after sink commit, replay a batch and compare outputs to a hand-computed small fixture.
Situational decisions
When invalid records can be isolated without corrupting the batch: Quarantine with counts/reasons under the declared policy; never drop them silently.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.