Designing Permission-Aware Retrieval for Internal AI Search

Design retrieval systems to enforce the requesting user’s document permissions before content reaches the language model. Carry those permissions through indexing, search, citations and cached answers, and deny retrieval when access cannot be verified. A prompt telling the model not to reveal confidential information is not an access-control boundary: the application must decide which evidence the model can receive.
Establish an authorization contract before indexing
Retrieval-augmented generation (RAG) supplies a language model with retrieved evidence to help answer a question. Microsoft’s RAG design guide describes documents being split into chunks—smaller searchable passages—and search results being assembled into model context. Permissions need to survive both transformations.
Start with a written authorization contract: who can read what, which system decides, and how quickly changes must take effect. Distinguish authentication, which establishes identity, from authorization, which determines permitted access.
For each indexed chunk, retain a source document identifier, tenant identifier where applicable, content version, permission reference and permission version. An access-control list (ACL) identifies users or groups allowed or denied access. Copying an ACL is useful only if its meaning remains faithful to the source, including inherited permissions, explicit denials and group membership.
Recommended rules are:
- Quarantine documents whose permissions cannot be mapped reliably; missing metadata must not mean public access.
- Preserve narrower section-level restrictions rather than assigning every chunk the document’s broadest access.
- Treat titles, previews and generated summaries as protected content too.
- For summaries combining multiple sources, require access to every contributing source unless an approved publication process creates a separately governed artifact.
A shared index can work when mandatory permission filters are enforceable. Separate indexes may simplify tenant isolation, but they do not replace document-level checks within each tenant.
Enforce access along every retrieval route
Resolve the user’s identity and memberships on the server. Do not accept a client-supplied group list or let the model choose its own access scope. The trusted retrieval service should attach mandatory tenant and permission constraints to every search.
Prefer filtering within the search operation before selecting the final results. Retrieving a global shortlist and removing unauthorized matches afterward can leave too little useful evidence, even when accessible documents exist elsewhere in the index. Apply the same constraints to keyword search, semantic search, follow-up searches and direct document fetches.
Before sending passages to a reranker—a component that reorders search matches—or the answering model, validate access against permission state that meets your freshness requirement. If indexed permissions can lag, use an authoritative check or exclude documents whose state is uncertain. Never send unrestricted passages to a model and ask it to remove them.
This also limits the damage from prompt injection: malicious instructions embedded in user input or retrieved documents. OWASP’s prevention guidance recommends checking tool calls against user permissions and session context, minimizing privileges, and treating model-based safeguards as additional layers rather than replacements for deterministic controls. A retrieved passage must never be able to broaden the next search’s permissions.
Worked example: a procurement assistant
Illustrative example: A buyer asks, “What delivery terms can we offer Supplier Cedar?” The repository contains three documents:
| Document | Access | Relevant evidence |
|---|---|---|
| Purchasing handbook | All procurement staff | Standard delivery terms |
| Cedar negotiation notes | Cedar deal team | An approved exception |
| Legal dispute memo | Legal only | Disputed obligations |
The buyer belongs to procurement and the Cedar deal team, but not Legal. The server establishes those memberships and searches only eligible documents. The handbook and negotiation notes provide evidence; the legal memo contributes neither text nor a title hint.
The answer describes the standard terms and approved exception, citing the two accessible sources. Citation links still require authorization when opened.
Now remove the buyer from the Cedar deal team. Repeating the question must not return the previous answer from a cache. Revalidate the answer’s supporting documents and regenerate from the handbook alone. Previously supplied negotiation text must also be excluded from model conversation history on subsequent turns. Revocation cannot undo information already seen, but it should prevent renewed disclosure through the assistant.
Make stale permissions an explicit exception path
Define a measurable revocation requirement with the data owner. If access changes must apply immediately, a periodically synchronized ACL alone cannot satisfy that requirement.
When identity resolution or the permission service fails, stop affected retrieval rather than falling back to a broadly privileged service account. Return a neutral message explaining that access could not be verified, without confirming restricted document names. Route operational details to an authorized support queue.
Partition caches by tenant and authorization context, track source dependencies, and revalidate before reuse. Restrict diagnostic logs containing passages or search results; record identifiers and decision metadata instead where possible. The same issue extends beyond retrieval into storage policy, covered in data retention questions before deploying generative AI.
Set release gates for security and usefulness
Build tests pairing each question with multiple identities and expected accessible sources. Recommended acceptance criteria include:
- Unauthorized exposure: zero restricted passages, titles or citations in model inputs and user-visible outputs across the test suite. Any observed leak blocks release; passing tests is not proof of universal safety.
- Permission-change behavior: measure time from revocation to effective denial, including caches and active conversations, against the agreed deadline.
- Failure safety: unavailable permission services, missing ACLs and unknown groups produce no unverified disclosure.
- Authorized retrieval quality: measure how many relevant accessible passages appear in the selected results, plus answer completeness and citation correctness.
- Operational performance: measure response time under realistic group sizes and permission checks, not just unrestricted search.
If you are planning an internal knowledge assistant, contact AIoverflow to scope its permission model, exception handling and evaluation gates before choosing the retrieval stack.
Sources & further reading
Prepared with AI assistance using the sources above and AIoverflow’s service context. Examples are illustrative; validate implementation decisions against your own requirements. Suggest a correction.