Architecture · plan by default
/arch-scale
Identify bottlenecks for a specified workload and growth scenario
Use for workload-driven capacity design; perf measures and repairs a specific bottleneck, and test-load executes an authorized workload.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select arch-scale from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv arch-scale, /just-vibe arch-scale and /jv:arch-scale in Claude. See shortcut setup and context examples.
/just-vibe:arch-scale Plan capacity experiments for a tenfold increase in checkout traffic./just-vibe:arch-scale Plan growth for a service bottlenecked by a serialized inventory update./just-vibe:arch-scale Assess scale without production metrics; avoid invented throughput estimates.What the agent does
- State workload shape, service objectives and measured constraints: arrival rate, service time distribution, concurrency, queue age and resource saturation. Separate observed production data from assumptions or synthetic samples.
- Locate the limiting serial/shared boundary before recommending replicas, caching, queues or extraction. Model steady-state and burst behavior, failure recovery and downstream limits with explicit units.
- Compare options against the bottleneck and consistency requirements. Define admission control/backpressure and degradation before adding unbounded concurrency; scaling callers can overload the shared dependency.
- Propose a bounded measurement or authorized load experiment with rejecting observations and recovery. Report the capacity range established by evidence and what remains unknown; do not invent traffic or throughput.
Inputs
- workload shape, expected growth, SLOs, current metrics, and resource limits.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Capacity constraints and targeted scaling strategy; no speculative whole-system rewrite.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Plan; workload shape, expected growth, SLOs, current metrics, and resource limits.
- Prerequisites
- Readable source, infrastructure/configuration definitions, and any supplied system documentation. Runtime telemetry is optional evidence, never assumed available. Architecture proposals remain plans until implementation is requested.
Expected output
- Bottleneck model, assumptions, capacity experiments, scaling sequence, and cost factors.
- Assumptions, limiting resource, incremental options and a bounded benchmark plan.
How the work is checked
- A serialized database operation is not solved by app replicas alone; estimates identify their workload assumptions.
When to stop or clarify
- No invented capacity numbers or automatic provisioning. Missing measurements yield an instrumentation/benchmark plan first.
Handling missing context
- Infer
- Trace current entry points, data owners, deployment units and documented constraints before proposing boundaries.
- Assume
- Prefer extending an existing owner while scale or organizational evidence is absent; mark capacity estimates as assumptions.
- Ask
- Ask for an unresolved consistency, compatibility or ownership requirement only if it changes the design; missing telemetry limits capacity claims, not source mapping.
Technical guidance
- Evidence
- Obtain workload shape, service-time distribution, concurrency limits, queue age and dependency quotas.
- Method
- Locate the first saturated shared resource; estimate concurrency from throughput and mean time only under stated steady-state assumptions, then measure tail behavior.
- Pitfall
- Adding replicas can exhaust a database connection budget or amplify retries before increasing throughput.
- Check
- Compare a bounded workload at the same mix and revision, including saturation, queue recovery and downstream limits.
Situational decisions
When a shared database or serialized step dominates: Quantify that constraint before recommending application replicas or new services.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.