LLMs and retrieval · inspect by default
/llm-cost
Measure token use, latency, caching opportunities, and routing tradeoffs
Use to measure LLM spend and cost-preserving alternatives; llm-evals measures task quality.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select llm-cost from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv llm-cost, /just-vibe llm-cost and /jv:llm-cost in Claude. See shortcut setup and context examples.
/just-vibe:llm-cost Analyze these token and retry records with explicit pricing assumptions./just-vibe:llm-cost Analyze a retry loop whose successful responses hide expensive failed attempts./just-vibe:llm-cost Estimate from incomplete usage records without inventing current prices.What the agent does
- Reconcile billed provider usage with estimates and retries, normalizing each provider to disjoint categories first (uncached input, output, cache reads, cache writes), since some APIs count cached tokens inside input and others report them separately.
- Read reasoning or thinking tokens from the provider usage report (billed as output, never estimated from visible text), verify dated pricing, and include failed runs in per-completed-task cost.
- Identify expensive failure loops and propose bounded comparisons that preserve task quality.
Inputs
- usage/latency records, task mix, quality requirements, and current verified pricing when calculating cost.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Tokens, retries, caching, model routing, concurrency, and cost/quality tradeoffs.
- Writes
- No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
- Mode
- Inspect; usage/latency records, task mix, quality requirements, and current verified pricing when calculating cost.
- Prerequisites
- Task definition, model/provider configuration, representative permitted data, versioned prompts/corpus where relevant, and explicit token/cost/latency limits for remote calls. Use current provider interfaces during implementation. Retrieved content and model-generated tool arguments remain untrusted.
Expected output
- Cost/latency breakdown with usage and dated rate assumptions, total and per-success cost, quality comparison, uncertainty and optimization priorities.
How the work is checked
- Retry tokens count toward total cost; cheaper routing is assessed against the same quality criteria.
When to stop or clarify
- Do not invent prices or silently switch providers/send data elsewhere. Missing usage records produce estimates with explicit bounds.
Handling missing context
- Infer
- Read current prompt/tool schemas, retrieval boundaries, installed SDK/provider config and permitted examples without reading secret values.
- Assume
- Use mocked calls for local contract tests when remote access is absent; do not infer model quality from mocks.
- Ask
- Ask for budget and permitted data/provider before a paid or external run if not already set; local prompt/tool implementation can proceed in apply mode.
Technical guidance
- Evidence
- Measure all requests, retries, failures, cached/uncached inputs, outputs, reasoning or thinking tokens from provider usage fields, and latency by task outcome.
- Method
- Compare routes at matched quality requirements using dated verified prices and explicit privacy/transfer constraints.
- Pitfall
- Lower price per call can raise cost per completed task through retries or quality failures.
- Check
- Reconcile usage totals with actual calls and compare successful outcomes, latency and failure rates under the same cases.
Situational decisions
When cheaper routing changes correctness or privacy conditions: Compare on the same cases and keep provider/data-transfer choices explicit.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.