Definition
The prompt injection test (built with Llama) checks for prompt injection, malicious strings and jailbreak attempts in the input data of your system.Taxonomy
- Task types: LLM.
- Availability: and .
Why it matters
- Prompt injection is a type of attack that exploits an AI system and deviates it from its intended behavior.
- It is important to detect and prevent prompt injection attacks to ensure the reliability and security of your system.
Judge explanations
In the test Parameters, turn on LLM-as-a-judge explanations. While it is on, the shared LLM judge settings appear beneath the switch: model, sampling, batch, and custom instructions. Other LLM-as-a-judge tests, such as LLM-as-a-judge and Bias, use the same block. When the judge runs, results include Judge explanation and Judge confidence columns. The explanation covers why the row was flagged, where the detection appears, and the relevant excerpt. Judge confidence is the LLM judge’s confidence that this is a real injection. With explanations off, results show a Detection details column with the detector’s own evidence. Hover the prompt-injection verdict to see the same summary. The judge can call out a likely false positive. Explanations are advisory and do not change the test count or verdict. Prompt-injection results support the Summarize drill-down when the run has per-row explanations. Without an explicit question column, the detector scans every text-bearing input variable. On a traced pipeline with several inputs, a row is flagged if any scanned input fires, and the results name the input column that fired. The judge reviews a sample of flagged rows, bounded by the evaluator’ssample settings, merged with the project’s default LLM evaluator. The project default commonly uses a random sample with a row limit. Custom prompt instructions on the evaluator apply to the explanation pass.
Test configuration examples
If you are writing atests.json, here are valid configurations for the prompt injection test.
Set explain_with_llm to true to request explanations. You can also set llm_evaluator to override the model, sampling, and custom instructions for this explanation pass.