Operations · plan by default

/ops-alerts

Design actionable alerts with ownership and response guidance

Use to design response-worthy alerts; ops-observability supplies reliable signals.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select ops-alerts from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv ops-alerts, /just-vibe ops-alerts and /jv:ops-alerts in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:ops-alerts Design actionable error-rate alerts with recovery behavior; do not enable notifications.
edge · plan
/just-vibe:ops-alerts Design alerts that ignore brief spikes but detect sustained customer failures.
blocked · inspect
/just-vibe:ops-alerts Draft alerts without sending pages or inventing response ownership.
Example · apply
/just-vibe:ops-alerts Add a Prometheus alert rule for queue age to monitoring/alerts.yml with rule tests, leaving notifications disabled.

What the agent does

  1. Tie each alert to impact and a responder action.
  2. Define evaluation and recovery windows, thresholds, missing-data behavior and noise suppression.
  3. Test with historical or synthetic events.

Inputs

  • service objectives, telemetry, response ownership, notification destination, and noise tolerance.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Actionable alert rules and response guidance; enabling/sending notifications requires an explicit request.
Writes
Inspect/plan: design alerts; save requested artifacts only. Apply: write only the requested alert-rule files and rule tests. Notification destinations, enablement and test pages stay out of scope unless explicitly requested.
Mode
Plan by default; apply to write requested alert-rule files and rule tests without enabling notification routes. Requires service objectives, telemetry, response ownership, notification destination, and noise tolerance.
Prerequisites
Exact service/environment, time window, revision/configuration identity, authorized logs/metrics, and operational constraints. Prefer observation before intervention; live restarts, traffic changes, restores, and notifications require the requested target/action. Redact sensitive telemetry.

Expected output

  • Alert definitions as a contract with rationale, trigger/recovery examples, owner/destination assumptions, runbook links and false-positive/missed-event checks.

How the work is checked

  • Sustained harmful failure triggers; expected brief transients do not page indiscriminately.

When to stop or clarify

  • Do not invent owners or send test pages implicitly. Missing signal quality is a prerequisite, not something thresholds can hide.

Handling missing context

Infer
Read service/environment, time window, revision, available telemetry and existing incident or recovery procedure.
Assume
Start from supplied logs and read-only observation; rank hypotheses without presenting an unexecuted intervention as recovery.
Ask
Resolve the precise target and missing authority before restart, restore, notification or traffic changes; continue evidence analysis while waiting.

Technical guidance

Evidence
Establish user-impact signal, window, baseline, missing-data semantics and response owner.
Method
Define actionable thresholds, recovery and deduplication; distinguish page-worthy sustained failure from diagnostic events.
Pitfall
A threshold without a responder action creates noise; no samples must not silently become healthy.
Check
Replay healthy, failing, transient and missing-data sequences, checking firing, recovery and notification routing only within requested scope.

Situational decisions

When signal quality cannot distinguish failure from missing telemetry: Add an explicit missing-data state instead of silently treating absence as healthy.

When a service level objective exists: Derive page and ticket rules from error-budget burn rate over paired long and short windows (for a 30-day SLO: page at 14.4 times over 1 hour confirmed over 5 minutes, and 6 times over 6 hours confirmed over 30 minutes; ticket at 1 times over 3 days confirmed over 6 hours). Without an SLO, state the rationale for each threshold.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring