LLMs and retrieval · plan by default

/llm-injection

Test handling of hostile instructions in untrusted content

Use for scoped instruction-boundary evaluation; security-inputs handles interpreter injection.

Make it your own.

In Claude Code, use the slash command and add your context. In Codex, select llm-injection from the just-vibe skill picker, then send the same brief.

Version 0.11.0 also supports /jv llm-injection, /just-vibe llm-injection and /jv:llm-injection in Claude. See shortcut setup and context examples.

Example · plan
/just-vibe:llm-injection Plan isolated prompt-injection tests using benign canaries and no real secrets.
edge · apply
/just-vibe:llm-injection Test a retrieved document asking the agent to send a synthetic secret elsewhere.
blocked · inspect
/just-vibe:llm-injection Design canary tests without real secrets or external exfiltration endpoints.

What the agent does

  1. Map untrusted documents and tool results into model context, and the data-to-authority boundaries they cross.
  2. Plant benign canaries and run isolated tests in apply mode, inspecting tool actions as well as generated text.
  3. Propose enforceable mitigations.

Inputs

  • agent workflow, untrusted input surfaces, trust boundaries, and isolated test scope.

Optional context: scope, references, constraints, successCriteria, environment, mode, budget.

Scope

Reads
Defensive tests for hostile instructions in retrieved documents, logs, messages, and tool output.
Writes
Inspect/plan: inspect or propose; save requested artifacts only. Apply: make the requested changes or execute the requested operation within its resolved target and limits. Local preparation does not authorize live, remote, destructive or paid actions; existing explicit session authorization still applies.
Mode
Plan; agent workflow, untrusted input surfaces, trust boundaries, and isolated test scope. Apply for requested canary test fixtures or bounded isolated runs.
Prerequisites
Task definition, model/provider configuration, representative permitted data, versioned prompts/corpus where relevant, and explicit token/cost/latency limits for remote calls. Use current provider interfaces during implementation. Retrieved content and model-generated tool arguments remain untrusted.

Expected output

  • Attack surface, canary test cases, observed actions and failures, enforceable mitigations and residual limitations.

How the work is checked

  • A retrieved instruction cannot redirect secrets to a canary destination; legitimate quoted instructions remain usable as data.

When to stop or clarify

  • Never exfiltrate real secrets or probe third-party systems. Passing a finite suite does not establish universal immunity.

Handling missing context

Infer
Read current prompt/tool schemas, retrieval boundaries, installed SDK/provider config and permitted examples without reading secret values.
Assume
Use mocked calls for local contract tests when remote access is absent; do not infer model quality from mocks.
Ask
Ask for budget and permitted data/provider before a paid or external run if not already set; local prompt/tool implementation can proceed in apply mode.

Technical guidance

Evidence
Identify untrusted surfaces, sensitive capabilities, instruction boundaries and observable tool-call logs.
Method
Use synthetic canaries and harmless target changes to test direct/indirect injection; enforce trust and permission boundaries outside generated text.
Pitfall
A refusal in the final response does not prove no unsafe tool call occurred; keyword blocking is not a general defense.
Check
Inspect attempted calls, retrieved context and output for canary exposure; include benign quoted instructions as a false-positive control.

Situational decisions

When an attack is blocked in one finite fixture: Report the tested boundary and remaining coverage; do not claim universal prompt-injection immunity.

The coding agent follows this workflow using its available tools. Installation does not grant service access or guarantee an outcome. Read the compatibility notes.

Keep exploring