Data engineering · apply by default
/data-incremental
Implement checkpoints, deduplication, and incremental processing
Use for replay-safe CDC or watermark processing; data-backfill handles bounded historical ranges.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select data-incremental from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv data-incremental, /just-vibe data-incremental and /jv:data-incremental in Claude. See shortcut setup and context examples.
/just-vibe:data-incremental Implement checkpointed updates that handle late records and deletion events./just-vibe:data-incremental Process late corrections while safely replaying the previous hour./just-vibe:data-incremental Review incremental design without stable source identity; identify the blocking semantic gap.What the agent does
- Define event versus arrival order, stable keys, deletions, overlap and late-arrival handling.
- Define the checkpoint transaction and persist checkpoints only after durable effects so replay is safe.
- Test crashes on both sides of the write/checkpoint boundary.
Inputs
- source change mechanism, stable keys, watermark/checkpoint, late-data policy, and destination semantics.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Incremental updates, deduplication, deletes, and resume behavior.
- Writes
- Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
- Mode
- Apply; source change mechanism, stable keys, watermark/checkpoint, late-data policy, and destination semantics.
- Prerequisites
- Data source/version, schema/semantics, transformation code, permitted sampling scope, and storage/compute budget. Prefer aggregates and redacted samples; never upload datasets to external services implicitly. Record time zones and snapshot identity for reproducibility.
Expected output
- Incremental processor with its watermark/key protocol and state format, and replay, late-update, deletion and crash tests.
How the work is checked
- Replaying a completed window does not duplicate effects; late updates and deletions follow documented rules.
When to stop or clarify
- Do not advance checkpoints before durable output. Missing stable identity or change semantics requires a design decision before implementation.
Handling missing context
- Infer
- Inspect schema, source snapshot, transformation code, grain, time zones and permitted sample scope.
- Assume
- Use bounded synthetic or supplied samples when full data is unavailable; keep unknown values distinct from zero.
- Ask
- Resolve ambiguous entity/grain/time semantics before reconciliation or backfill; obtain missing data/compute limits only for the dependent scan or execution.
Technical guidance
- Evidence
- Identify ordering key, watermark meaning, late arrival bound, change/delete events and deduplication identity.
- Method
- Use a stable tie-breaker and explicit overlap/reconciliation window; preserve progress only after durable output.
- Pitfall
- A maximum event timestamp alone skips equal-timestamp rows and late arrivals; updates and tombstones need semantics.
- Check
- Test tied timestamps, late corrections, deletes, duplicate deliveries and restart at a batch boundary without missing or duplicating results.
Situational decisions
When late updates fall behind the current watermark: Use an explicit overlap/reconciliation policy rather than silently skipping them.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.