AI document data minimization: select the context before calling the model
Reduce unnecessary document exposure with task scoped selection, permission checks, and explicit handling of derived content.

Send the content needed for the requested task, after checking the user's access to it. Rewriting a selected clause rarely requires an entire contract and its appendices. Preserve necessary definitions and structure, but apply the same permissions and lifecycle rules to extracted chunks, prompts, and stored answers as to their sources.
Start with the task's smallest useful context
Distinguish tasks such as rewriting selected text, summarizing a document, answering a question across a folder, and extracting fields from a form. Each requires different scope. Use the user's selection or a permission filtered retrieval process to identify relevant material. Include necessary headings, table labels, or adjacent explanations so the selected text is not misleading in isolation. Document when additional context is fetched and why. Sending less content is useful only if the task remains meaningful and the selection process does not omit essential qualifications.
Context selection should produce an inspectable record: the task, authorized source versions, included passages, and the reason additional material was fetched. Keep this record without automatically retaining every prompt body. It gives developers a way to investigate overcollection and gives evaluators a way to explain missing context.
| Task | Likely starting scope | Reason to expand |
|---|---|---|
| Rewrite a selection | The selection and its immediate structure | A referenced definition is necessary to preserve meaning |
| Summarize a document | Allowed document content under the extraction policy | An explicitly included attachment forms part of the document |
| Answer across a folder | Permission-filtered relevant passages | Additional authorized evidence is needed to answer the question |
Separate retrieval relevance from authorization. A highly relevant chunk from a restricted file is still unavailable to this user. Enforce tenant and document filters at the retrieval boundary, and verify the selected source relationships before building the request. For stored answers, consider later permission changes: a previously generated paragraph may disclose the same protected information as its original source even if the citation link is now denied.

Enforce access at retrieval and later viewing
Check the user's current permission to the source documents before assembling model input. Search results and cached chunks must respect the same tenant and document boundaries as direct access. Recheck relevant permissions when showing stored answers or citations, especially if access can change between generation and later viewing. Treat embeddings, extracted text, and summaries as derived document data with their own storage and deletion rules. Moving content into a vector index does not remove the need to track which source document and version authorized its use.
Provider retention and local logs are separate reviews
Understand the selected model service's documented data handling for the actual account and configuration being used. Separately review your application's prompt logs, traces, caches, error reports, and support exports. A restrictive provider setting does not prevent your own telemetry from retaining full document text. Redact credentials and unnecessary personal data before submission where the workflow permits, while recognizing that automated redaction can miss context. Record the intended retention and authorized audience for both inputs and generated outputs rather than assuming transient processing leaves no copies anywhere.
Inspect the outgoing request for a selected clause
Use a synthetic contract with a selected termination clause, a definitions section, and unrelated employee information in an appendix. Provide the clause and any definitions needed to interpret it, but exclude the unrelated appendix. Show the user which source sections informed the summary and preserve the source version. Inspect the actual outgoing request and application logs during acceptance. This example tests whether context selection follows the task, rather than relying on a privacy statement while the implementation still sends the entire file by default.
- Map each AI task to its minimum useful source scope.
- Apply tenant and document authorization before retrieval and when displaying results.
- Track derived chunks and answers back to source versions.
- Review provider settings and local logging independently.
- Test deletion and permission changes across indexes, caches, and stored outputs.



