
Yusuf Demir
My working assumption is unfriendly on purpose: treat the system prompt as readable, treat every retrieved document as an instruction waiting to happen, and treat the model's own output as untrusted until your code has checked it. Prompt injection comes in two forms. Direct injection is a user typing instructions instead of data. Indirect injection is the more common and the harder one: instructions hidden in a document the system retrieves, a web page it fetches, or a file it reads. Nothing at the token level separates an instruction from content, so a model that can read a document can be redirected by it. I show the concrete cases rather than the theory. Mitigation is layered and none of the layers is sufficient. Delimiting untrusted content helps. Telling the model to disregard instructions inside the input helps and is not a guarantee. Removing the ability to act is the layer that holds, and that means constraining what the tools can reach rather than asking the model politely. Leakage is the other half. System prompts get quoted, retrieved training material can surface in output, and a tool with broad read access will happily return whatever was asked for. Output filtering catches some of it and is worth having as a backstop. I also cover the things that are not prompt injection, because they get conflated: an application that trusts model output as a command, a model with access to more tools than the task needs, and an agent loop with no limit on its iterations. Those are design faults, and no amount of prompt work fixes them.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →