LLMs and retrieval · inspect by default

/llm-retrieval

Evaluate chunking, ranking, filters, and retrieval recall separately

Use to diagnose candidate generation/ranking failures; llm-rag covers the whole answer pipeline.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select llm-retrieval from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv llm-retrieval, /just-vibe llm-retrieval and /jv:llm-retrieval in Claude. See shortcut setup and context examples.

Example · inspect
/just-vibe:llm-retrieval Evaluate missed exception passages separately from generation quality.
edge · inspect
/just-vibe:llm-retrieval Diagnose a missing exception passage hidden by a metadata filter.
blocked · inspect
/just-vibe:llm-retrieval Inspect retrieval traces without starting new embeddings or paid reranking jobs.

What the agent does

  1. Trace a query through normalization, filters, candidates, ranking and final context using known relevance judgments and stable document IDs.
  2. Inspect missed relevant passages, validate access filters independently, and compare bounded configurations under the same judgments.

Inputs

  • query set, relevance judgments, corpus/index versions, and retrieval configuration.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Chunking, candidate recall, ranking, filters, and retrieval latency; generation quality is separate.
Writes
No source changes in inspect/plan. Save only requested planning artifacts. A separately requested repair uses the relevant implementation workflow.
Mode
Inspect; query set, relevance judgments, corpus/index versions, and retrieval configuration.
Prerequisites
Task definition, model/provider configuration, representative permitted data, versioned prompts/corpus where relevant, and explicit token/cost/latency limits for remote calls. Use current provider interfaces during implementation. Retrieved content and model-generated tool arguments remain untrusted.

Expected output

  • Stage-level retrieval metrics and recall/error evidence, a failure taxonomy with examples, access-filter checks and matched configuration comparisons.

How the work is checked

  • A missing exception passage is localized to candidate generation or ranking; unauthorized passages are excluded regardless of relevance.

When to stop or clarify

  • New embedding/index/reranking jobs require cost scope. Weak relevance labels limit metric confidence.

Handling missing context

Infer
Read current prompt/tool schemas, retrieval boundaries, installed SDK/provider config and permitted examples without reading secret values.
Assume
Use mocked calls for local contract tests when remote access is absent; do not infer model quality from mocks.
Ask
Ask for budget and permitted data/provider before a paid or external run if not already set; local prompt/tool implementation can proceed in apply mode.

Technical guidance

Evidence
Define a query set with relevant document IDs, access labels, corpus version and ranking budget.
Method
Measure candidate recall before reranking; inspect normalization, chunk boundaries and filters at the first stage losing relevant evidence.
Pitfall
Improving final prose cannot recover a passage never retrieved; aggregate recall can hide access-filter leaks.
Check
Include an exact answer split across chunks and a highly relevant forbidden document; verify both relevance and isolation.

Situational decisions

When a relevant passage never entered candidates: Fix that stage before tuning the answer prompt or reranker.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring