LLMs and retrieval · plan by default

/llm-evals

Build representative evaluation cases and scoring criteria

Use to establish LLM task evaluation; llm-prompt optimizes against development cases.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select llm-evals from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv llm-evals, /just-vibe llm-evals and /jv:llm-evals in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:llm-evals Design an evaluation protocol covering valid, unsupported, and adversarial requests.
edge · plan
/just-vibe:llm-evals Evaluate tool use where a fluent answer hides an unauthorized action.
blocked · inspect
/just-vibe:llm-evals Design an evaluation with no inference budget; mark cases unexecuted.

What the agent does

  1. Define the task distribution, expected behavior, unacceptable outcomes and a versioned evaluation set. Separate development examples from held-out assessment; record consent/provenance for any real user data.
  2. Choose independently checkable artifact or outcome assertions first. Where a model judge is necessary, blind/randomize presentation where feasible, calibrate against human or deterministic examples and document judge disagreement and failure modes.
  3. Freeze model/configuration including reasoning effort or thinking budget, prompts, tool availability, retrieval snapshot and budgets for a comparison. Repeat matched cases, preserve every attempt and distinguish answer correctness from tool side effects, scope adherence and unsupported claims.
  4. Report per-case failures and denominators alongside aggregate results, latency and actual token accounting. Missing traces or usage remain missing. Normalize usage per provider before summing or pricing: some APIs include cached tokens in the input total and report them as a detail, others report cache reads and writes separately from input. A token count is not automatically a dollar charge.
  5. Use observed failures for targeted revisions, then evaluate on fresh cases as well as regression examples. Do not call improved scores on the now-known development set evidence of generalization or overall superiority.

Inputs

  • desired behavior, representative cases, failure costs, scoring rubric, and run budget.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Evaluation cases, harness, and reliable scoring; run only under authorized provider/data scope.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
Mode
Plan evaluation; apply for requested evaluator code or scoped evaluation runs.
Prerequisites
Task definition, model/provider configuration, representative permitted data, versioned prompts/corpus where relevant, and explicit token/cost/latency limits for remote calls. Use current provider interfaces during implementation. Retrieved content and model-generated tool arguments remain untrusted.

Expected output

  • Versioned protocol and inputs, scorer-control evidence, per-case outcomes and failure analysis, aggregate denominators, measured resources and limits on generalization.

How the work is checked

  • The scorer rejects known bad outputs and accepts known good controls. All attempted trials, failures and unavailable metrics remain visible; comparisons share documented conditions and held-out claims use genuinely unused cases.

When to stop or clarify

  • No sensitive data upload or unlimited inference. A model judge is evidence, not unquestionable ground truth.

Handling missing context

Infer
Read current prompt/tool schemas, retrieval boundaries, installed SDK/provider config and permitted examples without reading secret values.
Assume
Use mocked calls for local contract tests when remote access is absent; do not infer model quality from mocks.
Ask
Ask for budget and permitted data/provider before a paid or external run if not already set; local prompt/tool implementation can proceed in apply mode.

Technical guidance

Evidence
Collect task distribution, failure examples, privacy constraints, scoring rubric and allowed cost.
Method
Separate capability, safety and usability dimensions; calibrate automated judges against human examples and retain disagreements.
Pitfall
A fluent answer or a judge score alone does not prove task completion; reused development cases invite overfitting.
Check
Include known-pass/fail controls and held-out cases, preserve raw outputs and report uncertainty and incomplete runs.

Situational decisions

When stochastic runs disagree or a judge favors style over correctness: Report variance/disagreement and inspect the rubric before declaring a winner.

When the request is for local preparation or implementation: Implement cases, scoring and positive/negative controls locally; paid model calls require a resolved model, dataset and budget.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring