Understand the data–instruction confusion

An external passage may demand bypassing review, sending a secret elsewhere or publishing immediately under claimed prior approval. Treating that text as an application instruction lets an untrusted source control tools. Indirect injection can arrive through retrieval or attachments without a direct attacker conversation.

Legitimate documentation also contains imperative language. Keyword blocking alone cannot distinguish an installation example from an unauthorized command. The fundamental boundary is authenticated origin and granted capability, not a list of suspicious words.

Keep evidence and control channels separate

Mark web text, PDF content, snippets and tool-returned prose as data. Maintain trusted requests, permissions and approval records independently. A model may propose an action; the application must still validate it before execution.

Input or capability Treatment May it elevate privileges?
Webpage text Evidence for review No
Attachment instructions Content to interpret No
Generated action proposal Await policy checks Not by itself
Authenticated request Check actual permissions Only within granted scope
Evidence reader Return permitted material No implicit publication
Publishing action Separate version-bound approval Never from document claims

Do not promote retrieved text into a higher-trust instruction channel. Links found in evidence must not automatically become destinations for sending secrets.

Scope of the executed checks

The example uses tool, origin and scopes, allowing only operator-origin evidence reads. Document-origin requests and all publication requests are denied. Four combinations were executed and retained; unknown capabilities fail closed.

This is not model red-teaming and does not prove that a model resists persuasion. It shows that an application can deny execution even when a model proposes an unsafe action. Production origin must be assigned by trusted application context rather than accepted from generated arguments.

Constrain file and network operations

Limit attachment types, sizes and parsing resources. Avoid interpolating filenames into shell commands. Archive extraction needs path-traversal checks; URL readers need destination, redirect, internal-address and response-size controls. Logs should not expose credentials or private client information.

Give content generation and publication different capability sets. A source-summary step does not need server write access or unrestricted access to every client. Least privilege limits consequences but does not replace authentication, audit or recovery.

Maintain negative cases and explicit failure handling

Include impersonated administrators, secret-exfiltration requests, approval bypasses, embedded tool instructions and redirected destinations. Public examples should use harmless fixtures, not attacks against third-party services.

Record origin, proposed action, policy decision, actual execution and potential log leakage. Rejection should retain a reason, not silently retry with broader privileges. A legitimate blocked task needs clarification from a trusted operator, not authorization supplied by the retrieved document.

Zhihe Growth's disclosure boundary

Zhihe Growth can publish reference code and security design, but four policy tests are not comprehensive agent certification. Intake, download, generation and publication need distinct permissions. Real acceptance must cover deployed models, interfaces, parsers and infrastructure.

This encyclopedia exercise performed no commercial-model attacks or external-platform retests and exposes no client data. Read MCP/API boundaries, fact validation and dataset privacy. Security and factual correctness are separate measures.

Security review: proposal is not execution

Request in untrusted material Preserve Block
Impersonated admin demands release Origin and request Publishing privilege
Upload client list Destination and scope Private-data transfer
Delete review history Proposal and denial Audit mutation
Claim prior approval Source of claim Replacement of trusted approval
Invoke unknown tool Name and policy result Dynamic capability expansion
Redirect to internal host Redirect chain Unauthorized internal access

Report model proposal, policy rejection and absence of side effects separately. A blocked action can still reveal model susceptibility. A model refusal does not prove the server rejects forged calls. Independent controls remain necessary even when one prompt appears effective in a demonstration.

Materials and primary references

Download the policy and result package. Consult OWASP's injection guidance and MCP security practices. Defenses require continued verification; this article does not promise elimination of every attack.

Knowledge center · GEO services · Research and evidence