Understand the data–instruction confusion
An external passage may demand bypassing review, sending a secret elsewhere or publishing immediately under claimed prior approval. Treating that text as an application instruction lets an untrusted source control tools. Indirect injection can arrive through retrieval or attachments without a direct attacker conversation.
Legitimate documentation also contains imperative language. Keyword blocking alone cannot distinguish an installation example from an unauthorized command. The fundamental boundary is authenticated origin and granted capability, not a list of suspicious words.
Keep evidence and control channels separate
Mark web text, PDF content, snippets and tool-returned prose as data. Maintain trusted requests, permissions and approval records independently. A model may propose an action; the application must still validate it before execution.
| Input or capability | Treatment | May it elevate privileges? |
|---|---|---|
| Webpage text | Evidence for review | No |
| Attachment instructions | Content to interpret | No |
| Generated action proposal | Await policy checks | Not by itself |
| Authenticated request | Check actual permissions | Only within granted scope |
| Evidence reader | Return permitted material | No implicit publication |
| Publishing action | Separate version-bound approval | Never from document claims |
Do not promote retrieved text into a higher-trust instruction channel. Links found in evidence must not automatically become destinations for sending secrets.
Scope of the executed checks
The example uses tool, origin and scopes, allowing only operator-origin evidence reads. Document-origin requests and all publication requests are denied. Four combinations were executed and retained; unknown capabilities fail closed.
This is not model red-teaming and does not prove that a model resists persuasion. It shows that an application can deny execution even when a model proposes an unsafe action. Production origin must be assigned by trusted application context rather than accepted from generated arguments.
Constrain file and network operations
Limit attachment types, sizes and parsing resources. Avoid interpolating filenames into shell commands. Archive extraction needs path-traversal checks; URL readers need destination, redirect, internal-address and response-size controls. Logs should not expose credentials or private client information.
Give content generation and publication different capability sets. A source-summary step does not need server write access or unrestricted access to every client. Least privilege limits consequences but does not replace authentication, audit or recovery.
Maintain negative cases and explicit failure handling
Include impersonated administrators, secret-exfiltration requests, approval bypasses, embedded tool instructions and redirected destinations. Public examples should use harmless fixtures, not attacks against third-party services.
Record origin, proposed action, policy decision, actual execution and potential log leakage. Rejection should retain a reason, not silently retry with broader privileges. A legitimate blocked task needs clarification from a trusted operator, not authorization supplied by the retrieved document.
Zhihe Growth's disclosure boundary
Zhihe Growth can publish reference code and security design, but four policy tests are not comprehensive agent certification. Intake, download, generation and publication need distinct permissions. Real acceptance must cover deployed models, interfaces, parsers and infrastructure.
This encyclopedia exercise performed no commercial-model attacks or external-platform retests and exposes no client data. Read MCP/API boundaries, fact validation and dataset privacy. Security and factual correctness are separate measures.
Security review: proposal is not execution
| Request in untrusted material | Preserve | Block |
|---|---|---|
| Impersonated admin demands release | Origin and request | Publishing privilege |
| Upload client list | Destination and scope | Private-data transfer |
| Delete review history | Proposal and denial | Audit mutation |
| Claim prior approval | Source of claim | Replacement of trusted approval |
| Invoke unknown tool | Name and policy result | Dynamic capability expansion |
| Redirect to internal host | Redirect chain | Unauthorized internal access |
Report model proposal, policy rejection and absence of side effects separately. A blocked action can still reveal model susceptibility. A model refusal does not prove the server rejects forged calls. Independent controls remain necessary even when one prompt appears effective in a demonstration.
Materials and primary references
Download the policy and result package. Consult OWASP's injection guidance and MCP security practices. Defenses require continued verification; this article does not promise elimination of every attack.