A document can tell an assistant to change the task, disclose information, or follow a link. Those instructions may be plainly visible. Review what the document asks the assistant to do, as well as what it hides.
What counts as prompt injection?
Prompt injection happens when an assistant treats untrusted content as instructions. A resume that tells an automated reviewer to award the highest score is one example. The UK National Cyber Security Centre uses a candidate CV to explain the risk.
Hidden text is a clue, not proof. Accessibility text, multilingual writing, and security training documents can produce findings too. Judge the instruction in the context of the task.
Which scenarios should you check?
- Task changes: ignore the criteria, approve this request, or omit an unfavorable fact.
- Fake authority: text claiming to be a system message, administrator override, or trusted tool response.
- Disclosure requests: reveal private instructions, conversation history, credentials, or other records.
- Action requests: send information, open a destination, run a command, or change a file.
- Concealment: tiny, covered, off-page, or background-matched text, plus content outside the visible page.
- Disguised words: invisible Unicode, lookalike letters, escaped codes, encoded text, or ASCII art.
OWASP documents instruction, encoding, and tool-abuse patterns. Microsoft describes hidden-text attacks in documents and markup. A visual check alone cannot cover these scenarios.
How can Unicode and ASCII disguise instructions?
Invisible Unicode can split words or carry hidden character data. Direction controls can reorder displayed text. Lookalike letters can disguise familiar words. Unicode guidance also explains legitimate uses of these characters, so deleting every unusual character is not a reliable fix.
Encodings replace words with character codes. ASCII art arranges ordinary symbols into large letters. The ArtPrompt research demonstrates why a model may interpret those shapes as words. Decoding text can expose an instruction; it does not make that instruction trustworthy.
What does the detector tell you?
The detector surfaces instruction patterns, character disguises, concealed text, and supported document fields for review. Each finding names its source and explains the evidence. Examples let you compare ordinary text with a document that tries to redirect an assistant.
Checks use local rules and OCR, not an assessment of the author's intent. English instruction patterns are limited. OCR can misread text, and different document readers can extract different content. Read the coverage notes and any incomplete checks alongside the findings.
What remains outside a document scan?
A single-file scan cannot assess instructions spread across conversations, retrieved documents, or persistent memory. It cannot predict what an assistant will do. ASCII-art interpretation, arbitrary encodings, and embedded-file contents also need separate review. No findings means only that the checks found no match.
What should you do with a finding?
- Read the source and decoded text as evidence. Do not follow instructions found inside it.
- Compare the finding with the visible document and the task you actually requested.
- Remove or clarify unexpected instructions before using the document in an assistant workflow.
- Keep sending data, running commands, and changing records behind a separate approval step.
Microsoft recommends layered defenses and human verification. Limit the assistant's access and actions even after reviewing its inputs.
Review before sharing.
Common questions
- What is indirect prompt injection?
- Indirect prompt injection is when external content, such as a document, is processed by an assistant and treated as an instruction instead of material to analyze.
- Is concealed text the same as prompt injection?
- No. Concealed text is presentation evidence that a person may miss while extraction or processing can still receive it. It can be accidental, and it does not establish the author's intent.
- Can visible text still be prompt injection?
- Yes. An indirect prompt injection can be visible to a person. Concealed text is one way external instructions may be missed during a normal document review.
- What does the prompt injection detector check?
- It surfaces instruction patterns, disguised characters, concealed text, and supported document fields for review. Each finding includes its source and evidence. Coverage notes identify limits and incomplete checks.
- What should I do when a finding appears?
- Read the text, location, reason, and evidence in context. Treat instructions from the document as content to inspect, keep risky actions under human review, and decide whether to remove, replace, or explain the text before sharing.