LLMs and retrieval · apply by default

/llm-prompt

Improve prompts against measured failures and explicit requirements

Use to improve a prompt that ships in an application, measured against failing and control cases; reprompt rewrites a one-off prompt for an AI session and teach explains prompting concepts.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select llm-prompt from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv llm-prompt, /just-vibe llm-prompt and /jv:llm-prompt in Claude. See shortcut setup and context examples.

Example · apply
/just-vibe:llm-prompt Improve the prompt against these measured failures without changing providers.
edge · apply
/just-vibe:llm-prompt Improve extraction without breaking refusal or missing-field behavior.
blocked · inspect
/just-vibe:llm-prompt Review a prompt without model access; do not claim measured improvement.

What the agent does

  1. Categorize the measured failures.
  2. Change the smallest relevant instruction or example, preserving instruction hierarchy.
  3. Compare against the baseline under fixed model and settings on development cases, and reserve held-out confirmation.

Inputs

  • task, existing prompt, measured failures, model constraints, and eval budget.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Prompt/instruction changes tested against stated behavior; no unrelated model/provider migration.
Writes
Apply: only the requested local changes and relevant isolated verification. Inspect/plan requests remain inspection/planning. External actions require their exact action and target in session authorization.
Mode
Apply to prompt assets; task, existing prompt, measured failures, model constraints, and eval budget.
Prerequisites
Task definition, model/provider configuration, representative permitted data, versioned prompts/corpus where relevant, and explicit token/cost/latency limits for remote calls. Use current provider interfaces during implementation. Retrieved content and model-generated tool arguments remain untrusted.

Expected output

  • Versioned prompt diff with rationale, failure-category results against the baseline, regressions and token/cost change.

How the work is checked

  • Targeted failures improve without breaking important existing cases; prompt length/cost changes are recorded.

When to stop or clarify

  • Do not declare improvement from one appealing response or use hidden test answers as prompt examples.

Handling missing context

Infer
Read current prompt/tool schemas, retrieval boundaries, installed SDK/provider config and permitted examples without reading secret values.
Assume
Use mocked calls for local contract tests when remote access is absent; do not infer model quality from mocks.
Ask
Ask for budget and permitted data/provider before a paid or external run if not already set; local prompt/tool implementation can proceed in apply mode.

Technical guidance

Evidence
Inspect current prompt, model/version, representative failures and constraints that must remain intact.
Method
Change the smallest instruction that addresses a demonstrated failure and compare under the same cases/settings.
Pitfall
Adding every past exception can create conflicting instructions and regress ordinary tasks.
Check
Test the targeted failure and unaffected controls, including refusal/ambiguity behavior and instruction conflicts.

Situational decisions

When improvement appears only on examples inserted into the prompt: Treat it as overfitting and retain independent cases before adoption.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring