LLMs and retrieval · plan by default
/llm-evals
Build representative evaluation cases and scoring criteria
Use to establish LLM task evaluation; llm-prompt optimizes against development cases.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select llm-evals from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv llm-evals, /just-vibe llm-evals and /jv:llm-evals in Claude. See shortcut setup and context examples.
/just-vibe:llm-evals Design an evaluation protocol covering valid, unsupported, and adversarial requests./just-vibe:llm-evals Evaluate tool use where a fluent answer hides an unauthorized action./just-vibe:llm-evals Design an evaluation with no inference budget; mark cases unexecuted.What the agent does
- Define the task distribution, expected behavior, unacceptable outcomes and a versioned evaluation set. Separate development examples from held-out assessment; record consent/provenance for any real user data.
- Choose independently checkable artifact or outcome assertions first. Where a model judge is necessary, blind/randomize presentation where feasible, calibrate against human or deterministic examples and document judge disagreement and failure modes.
- Freeze model/configuration including reasoning effort or thinking budget, prompts, tool availability, retrieval snapshot and budgets for a comparison. Repeat matched cases, preserve every attempt and distinguish answer correctness from tool side effects, scope adherence and unsupported claims.
- Report per-case failures and denominators alongside aggregate results, latency and actual token accounting. Missing traces or usage remain missing. Normalize usage per provider before summing or pricing: some APIs include cached tokens in the input total and report them as a detail, others report cache reads and writes separately from input. A token count is not automatically a dollar charge.
- Use observed failures for targeted revisions, then evaluate on fresh cases as well as regression examples. Do not call improved scores on the now-known development set evidence of generalization or overall superiority.
Inputs
- desired behavior, representative cases, failure costs, scoring rubric, and run budget.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Evaluation cases, harness, and reliable scoring; run only under authorized provider/data scope.
- Writes
- Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
- Mode
- Plan evaluation; apply for requested evaluator code or scoped evaluation runs.
- Prerequisites
- Task definition, model/provider configuration, representative permitted data, versioned prompts/corpus where relevant, and explicit token/cost/latency limits for remote calls. Use current provider interfaces during implementation. Retrieved content and model-generated tool arguments remain untrusted.
Expected output
- Versioned protocol and inputs, scorer-control evidence, per-case outcomes and failure analysis, aggregate denominators, measured resources and limits on generalization.
How the work is checked
- The scorer rejects known bad outputs and accepts known good controls. All attempted trials, failures and unavailable metrics remain visible; comparisons share documented conditions and held-out claims use genuinely unused cases.
When to stop or clarify
- No sensitive data upload or unlimited inference. A model judge is evidence, not unquestionable ground truth.
Handling missing context
- Infer
- Read current prompt/tool schemas, retrieval boundaries, installed SDK/provider config and permitted examples without reading secret values.
- Assume
- Use mocked calls for local contract tests when remote access is absent; do not infer model quality from mocks.
- Ask
- Ask for budget and permitted data/provider before a paid or external run if not already set; local prompt/tool implementation can proceed in apply mode.
Technical guidance
- Evidence
- Collect task distribution, failure examples, privacy constraints, scoring rubric and allowed cost.
- Method
- Separate capability, safety and usability dimensions; calibrate automated judges against human examples and retain disagreements.
- Pitfall
- A fluent answer or a judge score alone does not prove task completion; reused development cases invite overfitting.
- Check
- Include known-pass/fail controls and held-out cases, preserve raw outputs and report uncertainty and incomplete runs.
Situational decisions
When stochastic runs disagree or a judge favors style over correctness: Report variance/disagreement and inspect the rubric before declaring a winner.
When the request is for local preparation or implementation: Implement cases, scoring and positive/negative controls locally; paid model calls require a resolved model, dataset and budget.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.