LLMs and retrieval · plan by default
/llm-injection
Test handling of hostile instructions in untrusted content
Use for scoped instruction-boundary evaluation; security-inputs handles interpreter injection.
Make it your own.
In Claude Code, use the slash command and add your context. In Codex, select llm-injection from the just-vibe skill picker, then send the same brief.
Version 0.11.0 also supports /jv llm-injection, /just-vibe llm-injection and /jv:llm-injection in Claude. See shortcut setup and context examples.
/just-vibe:llm-injection Plan isolated prompt-injection tests using benign canaries and no real secrets./just-vibe:llm-injection Test a retrieved document asking the agent to send a synthetic secret elsewhere./just-vibe:llm-injection Design canary tests without real secrets or external exfiltration endpoints.What the agent does
- Map untrusted documents and tool results into model context, and the data-to-authority boundaries they cross.
- Plant benign canaries and run isolated tests in apply mode, inspecting tool actions as well as generated text.
- Propose enforceable mitigations.
Inputs
- agent workflow, untrusted input surfaces, trust boundaries, and isolated test scope.
Optional context: scope, references, constraints, successCriteria, environment, mode, budget.
Scope
- Reads
- Defensive tests for hostile instructions in retrieved documents, logs, messages, and tool output.
- Writes
- Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
- Mode
- Plan; agent workflow, untrusted input surfaces, trust boundaries, and isolated test scope. Apply for requested canary test fixtures or bounded isolated runs.
- Prerequisites
- Task definition, model/provider configuration, representative permitted data, versioned prompts/corpus where relevant, and explicit token/cost/latency limits for remote calls. Use current provider interfaces during implementation. Retrieved content and model-generated tool arguments remain untrusted.
Expected output
- Attack surface, canary test cases, observed actions and failures, enforceable mitigations and residual limitations.
How the work is checked
- A retrieved instruction cannot redirect secrets to a canary destination; legitimate quoted instructions remain usable as data.
When to stop or clarify
- Never exfiltrate real secrets or probe third-party systems. Passing a finite suite does not establish universal immunity.
Handling missing context
- Infer
- Read current prompt/tool schemas, retrieval boundaries, installed SDK/provider config and permitted examples without reading secret values.
- Assume
- Use mocked calls for local contract tests when remote access is absent; do not infer model quality from mocks.
- Ask
- Ask for budget and permitted data/provider before a paid or external run if not already set; local prompt/tool implementation can proceed in apply mode.
Technical guidance
- Evidence
- Identify untrusted surfaces, sensitive capabilities, instruction boundaries and observable tool-call logs.
- Method
- Use synthetic canaries and harmless target changes to test direct/indirect injection; enforce trust and permission boundaries outside generated text.
- Pitfall
- A refusal in the final response does not prove no unsafe tool call occurred; keyword blocking is not a general defense.
- Check
- Inspect attempted calls, retrieved context and output for canary exposure; include benign quoted instructions as a false-positive control.
Situational decisions
When an attack is blocked in one finite fixture: Report the tested boundary and remaining coverage; do not claim universal prompt-injection immunity.
The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.