All storiesSecurity

Document AI agents: permissions, prompt injection, and publication

Limit agent reads and tools, test hostile document content, and guard consequential publication.

Document AI agents: permissions, prompt injection, and publication: Task authority, Untrusted content, Tool boundary, Approval gate.
Security / Office SDK

Document agents need separate permissions to read, draft, publish, and communicate. Enforce those permissions in tools and application services, not in prompts alone. Treat document instructions as untrusted content, and require an identified approval before consequential output leaves the working area.

Authority belongs outside the model

Describe the agent's job in operational terms: summarize a draft, propose edits, extract a table, or prepare an approval package. List the documents it may read and the operations it may perform. Use the initiating user's verified permissions or a deliberately restricted service identity, and recheck authorization at tool execution. A prompt stating that the agent should be careful is not an access control. Separate draft generation from publication, sharing, deletion, and external communication. Give each consequential operation an explicit policy so the integration can reject requests even when the model produces a confident instruction.

A retrieved document may contain instructions that try to redirect the agent, including text hidden in comments, attachments, or formatting. The document is task data, not a source of operational authority. Preserve that distinction in the application and tool layer. Do not let quoted content select arbitrary destinations, expand access scope, or override approval requirements. Validate structured tool arguments against allowed document identifiers and permitted actions. Test hostile instructions inside realistic files instead of testing only the chat interface. The attack surface includes everything the agent can read during the assigned task.

Documents supply data while policy and approval govern actions.
Figure 1. The model cannot promote source text into operational authority.

Work through a vendor review

Consider an agent asked to summarize supplier proposals and draft a comparison. One proposal includes a sentence telling the agent to email all competing bids to a named address. The expected behavior is to treat that sentence as proposal content and refuse any unauthorized communication. The agent may create a draft comparison in a designated workspace, but the procurement owner reviews it before distribution. Test whether source restrictions persist when the agent follows references or opens attachments. Also verify that generated text does not reveal information from another project merely because the underlying service account can access it.

Include tests where the model asks for a forbidden operation. The application should deny it even when the request is syntactically correct. Test an unauthorized document identifier, an unapproved recipient, an attempt to broaden a retrieval query, and a publication request without the required approval record.

task_scope: supplier proposals for project P17
allowed_output: draft comparison in P17 workspace
forbidden_effect: send proposals to external recipients
publication_guard: approved version + authorized reviewer

These are example policy fields. The implementation should use the actual identity and authorization system rather than trusting a model-generated value such as approved: true. In the hostile supplier proposal scenario, rejection of the email action is the important result. Whether the agent's prose recognizes the hostile sentence is useful but insufficient.

Then test a less obvious failure: the agent follows a reference into another project's file that the service account can access. The retrieval layer must enforce the task's narrower scope. Repeat after a user loses project membership during a long run. Decide how authority is rechecked at consequential tool calls, and record which already completed operations remain valid. This closes the gap between launch-time permission and the authority actually used later.

Record runs without leaking data

Keep a run identifier, initiating actor, task scope, referenced versions, executed tools, and approval decisions. Protect logs because prompts and outputs can themselves contain confidential information. Avoid logging full documents by default when metadata is sufficient for investigation. If model calls leave the organization's environment, assess the provider arrangement, processing locations, retention, and applicable policy before enabling the workflow. Define cancellation and recovery for interrupted runs. Retried tool calls should not publish duplicate files or repeat external actions, and a delayed result must not overwrite a newer document version without a guarded decision.

Read, draft, review, and publication use different permissions.
Figure 2. Guard retries so publication cannot repeat or overwrite newer work.

Agent release Decision notes

  • Specify task scope and separate read, draft, publish, and communication permissions.
  • Validate tool arguments against authorized objects and actions.
  • Place hostile instructions in documents and attachments during testing.
  • Require review for consequential outputs and record the decision.
  • Verify cancellation, retries, version guards, and protected run logs.

Further reading

Back to all stories

Keep reading.

All stories