The same question may legitimately produce different answers for different employees. A salesperson may read standard service terms without access to a customer's discount agreement. A project member may access a proposal without its older versions remaining applicable. Enterprise retrieval-augmented generation should organize sufficient evidence within the caller's current permissions and the scope in which the material applies. The following design discussion uses a hypothetical contract and delivery knowledge base.
Define the boundaries an answer must respect
Before designing the knowledge base, write an answer contract: which tenant the answer serves, who requested it, what time period applies, which sources have authority, and what the system may return when evidence is incomplete. Consider the question, 'Is Customer A exempt from service fees for this year's delivery delay?' Standard exemption terms, a customer amendment, and the cause of the delay may all be necessary. Finding one of them does not support a complete conclusion.
Represent that contract in requests and data. Every query carries an identity, tenant, and authorization context established through a trusted sign-in process. Every piece of evidence records its origin, version, effective dates, and authorized audience. A model can help interpret 'this year,' but the application should make the chosen dates and time zone explicit. The authorization service determines access; the model cannot enlarge that scope based on conversation text.
Distinguish documented facts, system inferences, and business decisions. Retrieval may find that a delay exceeding ten days permits an application for a fee reduction. It must not turn permission to apply into an automatic waiver. A satisfactory answer explains the clause, its conditions, and facts still to be established, while preserving any required business approval. These boundaries provide a stronger acceptance criterion than fluency alone.
Manage chunks, versions, and permissions as one evidence unit
Chunking should produce evidence that can be interpreted on its own. Cutting a contract at fixed lengths can separate an amount from its currency, a rule from its exception, or a table row from its headings. Start with section and paragraph boundaries, then subdivide long material. Preserve column names, units, and the enclosing section when extracting tables. A fragment missing a qualifier can support a wrong answer even when its vector similarity is high.
Each chunk should retain a stable source identifier alongside its revision, page or anchor, content fingerprint, and permission version. These fields reveal what changed after an update and let citations point to the version actually used. A displayed link may open the current document, but audit records also need access-controlled historical evidence. Otherwise, the document a user later reads may no longer match the answer.
Content updates, permission changes, and deletions all belong in the indexing lifecycle. Revocation cannot wait for the next full rebuild: at minimum, provide a promptly effective denial layer and specify a maximum synchronization delay. An internal deletion tombstone can prevent a late update from restoring an obsolete chunk. During version transitions, avoid combining new content with an old authorization record in the same answer.
Carry authorization through retrieval and delivery
Azure AI Search documentation distinguishes application-supplied string security filters from identity-based permission checks. This highlights a wider design requirement: a filter expression is an enforcement mechanism, while verified caller identity, fresh group membership, and trusted tenant boundaries require their own controls. A user identifier supplied by a browser or a group name generated by a model should never become sufficient authorization evidence.
Keyword and vector retrieval should use the same authorized scope. Retrieving the nearest candidates from the entire collection and then removing inaccessible ones can leave too few results, because other departments' documents consume the candidate window. Prefer retrieval paths that support filtering, and verify how the actual engine combines filters, approximate search, and truncation. Whatever happens inside the engine, unauthorized text must not reach an external reranker, a model context, or user-visible logs.
Permissions can also leak through side channels. Suggestions, document titles, hit counts, citation previews, and cached answers can reveal information outside the user's access. Cache keys need the tenant and an identifier that accurately represents the authorization scope, with invalidation when permissions change. Revalidate all evidence before reusing an answer across users. If synchronization status is uncertain, withhold the affected material rather than presenting a synchronization failure as proof that no material exists.
Use hybrid retrieval to build candidates, then check coverage
Enterprise questions often mix exact identifiers with everyday language. Contract numbers, product codes, and error codes require reliable literal matching, while 'What happens if delivery runs late?' may correspond to a section titled 'Changes to the performance period.' Keyword and semantic retrieval can therefore generate candidates separately before deduplication. Test tokenization, synonyms, and embeddings on domain examples, particularly Chinese abbreviations, English acronyms, and mixed alphanumeric identifiers.
Elastic's reciprocal rank fusion documentation describes merging result lists by rank, avoiding direct addition of raw scores on different scales. It offers an interpretable starting point, but does not establish whether a retrieval channel is reliable. Candidate windows, channel configuration, and filters still require validation. If a relevant document is absent from every retrieval channel’s candidate set, fusion cannot recover it.
For the delay question, retrieve general fee-reduction clauses, customer-specific amendments, and delay records separately, then map candidates to the types of evidence the question requires. Preserve the original tenant and time constraints during decomposition so that rewriting does not broaden the search. For debugging, record the source and rank of each candidate in each channel. The final answer alone rarely reveals whether the failure was tokenization, insufficient candidates, or omitted evidence during generation.
Reranking and context assembly solve different problems
Sentence Transformers documents a two-stage approach: retrieve candidates, then use a cross-encoder that reads the query and passage together to assess relevance. A practical implementation sends a limited merged candidate set to that reranker instead of scoring the whole collection. Choose its size using the latency budget and an offline recall curve. Whether a larger set is worth the cost must be measured on actual business questions.
A reranking score is still not the probability that an answer is correct. An obsolete version closely matching the wording may rank above the current clause. Treat authorization and version validity as hard constraints before relevance ranking. Context assembly should remove duplication, retain definitions and exceptions, and label conflicting sources. When comparison across documents is necessary, repeated standard terms must not crowd out the only customer amendment.
Lost in the Middle observed effects from evidence position in the tasks and models it tested. This motivates testing context order, rather than asserting a permanent rule for every newer model. Move decisive evidence and vary distractor volume in your own evaluation set to check answer stability. Organize evidence by the question's components, give sources compact identifiers, and reserve output space for limitations instead of filling the context simply because capacity remains.
Citations must support claims; abstention should explain gaps
A verifiable answer connects important statements to specific evidence. Citations should support amounts, dates, affected parties, and exceptions rather than placing a broadly relevant document link at the end of a paragraph. The system can first build an internal claim-evidence-condition structure, then produce prose. Validate that every citation identifier belongs to the authorized evidence set for this request, preventing invented document numbers or paths.
Citation checks ask separate questions: Does the source exist? Does the passage contain the fact? Does that fact support the entire conclusion? 'The contract allows an application for a reduction' does not establish 'The fee is waived in this case' without approval and factual assessment. A real citation can still accompany an unjustified inference. For calculations, record inputs, units, and rules and use deterministic computation; the language model explains the result rather than replacing the calculation with convincing prose.
Abstention can remain useful. When only standard terms are available, the answer can explain the application conditions found and state that accessible customer-specific terms or delay facts are missing, so applicability cannot be confirmed. Do not disclose the names of hidden documents when explaining an access limitation. Text within a document that tells the system to ignore rules or invoke other tools remains untrusted source material. It does not change application instructions, tool authorization, or evidence boundaries.
Evaluate retrieval, faithfulness, and access safety separately
An evaluation set should pair a question with different identities, dates, and document versions. Label accessible relevant evidence, the conclusions that should be returned, and conditions requiring abstention. Retrieval evaluation measures evidence recall and ranking within the authorized collection. Generation evaluation checks support for claims, citation accuracy, and completeness of conditions. A wrong answer should be attributable to missing retrieval, lost context, or unsupported generation, rather than leaving only an aggregate score.
The original RAGAS paper evaluates retrieved context and generated answers along distinct dimensions, offering a useful starting point for automated assessment. Model judges are nevertheless affected by their prompts, models, and domain knowledge; they cannot replace authorization tests or business review. Maintain a fixed, human-verified baseline for important clauses and regularly inspect disagreement between automatic scores and human judgments, especially where mild wording quietly expands a contractual commitment.
Access tests should inspect whether unauthorized text reaches any model or external service, not merely whether it appears in the final answer. Include group removal, a document becoming confidential, unexpired caches, and old citation links. Report abstention together with false abstention on answerable questions. A system that rejects nearly everything may appear safe without demonstrating usefulness within legitimate access boundaries.
Set separate limits for freshness, cost, and degradation
Production settings need distinct targets for permission synchronization, content freshness, and answer latency. Revocation generally needs faster enforcement than routine content updates. Long-contract summaries can be built asynchronously, but must inherit source permissions and record source versions. Break latency into identity resolution, retrieval, reranking, and generation to locate cost and delay before changing caching or candidate counts.
Degraded operation must still respect the answer contract. If reranking is unavailable, the system may fall back to a validated hybrid retrieval path with a narrower answering scope. If authorization is unavailable, it cannot continue using unverified evidence. For stale or conflicting sources, explain the applicable time period and unresolved parts. Returning relevant passages with explicit limitations can be more useful than forcing a complete recommendation.
Begin with questions whose information scope is clear, sources are stable, and consequences are manageable, then expand coverage. Replay the same question-identity combinations whenever chunking, embeddings, or authorization changes, and retain versioned results. Maintainable enterprise RAG lets every important conclusion be traced to evidence available at the time. It also explains why other material was not used and what additional information is necessary to continue.
