Operations · apply by default
/ops-observability
Add useful logs, metrics, and traces to unclear execution paths
Use to implement requested diagnostic telemetry; ops-alerts turns signals into actionable notifications.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select ops-observability from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv ops-observability, /just-vibe ops-observability and /jv:ops-observability in Claude. See shortcut setup and context examples.
/just-vibe:ops-observability Add focused tracing without logging raw user payloads or high-cardinality labels./just-vibe:ops-observability Add tracing across a queue while keeping sensitive payloads out of logs./just-vibe:ops-observability Design observability without provisioning a vendor or changing production configuration.What the agent does
- Start with the questions operators must answer.
- Propagate correlation across boundaries, and choose stable bounded-cardinality metrics plus redacted structured events.
- Verify normal and error instrumentation locally.
Inputs
- unclear path, diagnostic goals, telemetry stack, privacy, and overhead constraints.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Useful logs, metrics, traces, correlation, and error context in the selected path.
- Writes
- Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
- Mode
- Apply; unclear path, diagnostic goals, telemetry stack, privacy, and overhead constraints.
- Prerequisites
- Exact service/environment, time window, revision/configuration identity, authorized logs/metrics, and operational constraints. Prefer observation before intervention; live restarts, traffic changes, restores, and notifications require the requested target/action. Redact sensitive telemetry.
Expected output
- Instrumentation changes with a question/signal/location table, field and metric definitions, privacy/overhead decisions and success/error telemetry checks.
How the work is checked
- A failed operation can be traced across the intended boundary; raw user payloads or unbounded IDs do not become unsafe metric labels.
When to stop or clarify
- No provider provisioning or production configuration rollout implicitly. Do not add high-volume logging without overhead consideration.
Handling missing context
- Infer
- Read service/environment, time window, revision, available telemetry and existing incident or recovery procedure.
- Assume
- Start from supplied logs and read-only observation; rank hypotheses without presenting an unexecuted intervention as recovery.
- Ask
- Resolve the precise target and missing authority before restart, restore, notification or traffic changes; continue evidence analysis while waiting.
Technical guidance
- Evidence
- Identify a specific unanswered operational question, request lifecycle and data sensitivity/cardinality.
- Method
- Instrument useful stage timing, failure categories and correlation while bounding labels, volume and overhead.
- Pitfall
- User IDs in metric labels cause unbounded cardinality; logging full payloads can expose credentials or private data.
- Check
- Exercise success/error/cancellation and inspect exported telemetry, redaction, label bounds and instrumentation overhead.
Situational decisions
When proposed labels contain user IDs or arbitrary payload values: Move needed detail to controlled logs/traces or aggregate dimensions instead of unbounded metric labels.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.