A practical guide to enterprise RAG architecture, from document ingestion and hybrid search to permissions, citations and production evaluation.

An AI assistant can produce a fluent answer in seconds, but that does not mean it knows a company's current contracts, policies or project records. A general language model cannot reliably answer a question about a private document it has never seen. It also does not know which version is approved or which employee is allowed to read it.
Retrieval-Augmented Generation (RAG) addresses this gap by finding relevant company information before the model prepares an answer. In an enterprise system, the difficult work is usually in the retrieval layer: connecting sources, processing documents, preserving permissions, selecting the right passages and showing where an answer came from. The chat interface is only the visible part.
This article explains an enterprise RAG architecture from source systems to the final answer. It is intended for teams evaluating a production system over private company data, rather than a demonstration built around a few clean PDFs.
RAG lets an application use information outside the model's general training data. When a user asks a question, the application searches an index of authorised company content and provides selected results to the model as context. The model then uses that context to formulate a response. The original documents remain the source of evidence.
For example, an employee might ask, “What is the approval process for a new supplier?” A RAG system can search the current procurement policy, identify the relevant section and produce a concise answer with a link to the policy. If the policy is missing or contradictory, the system should make that clear rather than inventing a confident procedure.
This is the technical foundation behind an AI knowledge base built from company documents. It can also support specialised use cases, such as searching construction project documents, where source versions and project permissions are especially important.
A basic enterprise RAG flow has two connected processes. The first prepares company information for search. The second handles each user question.
Source systems → ingestion → parsing and metadata → searchable index
User question → identity and permission check → retrieval and ranking → selected context → language model → answer with sources
These steps can be implemented with managed services or with components the company operates itself. The design choice depends on existing infrastructure, data sensitivity, integration needs and the level of control required. In both cases, the essential question is the same: can the system retrieve the right evidence for the right user at the right time?
Useful company knowledge rarely lives in one repository. It may be spread across SharePoint, OneDrive, network drives, object storage, document management platforms, databases and internal applications. A source connector needs to collect content and enough metadata to make it usable, such as document ID, title, path, owner, dates, version and access rules.
The architecture should also identify which source is authoritative. If the same procedure appears in a controlled document library and an old email attachment, the system should know which one to favour. This may require repository-specific rules rather than one generic connector for every source.
Changes need a clear path into the index. New, revised and deleted documents should be reflected according to a defined schedule or event-driven process. Access changes matter too: if a person's permission is removed in the source system, the search layer must stop returning that content once the change is synchronised. The Azure AI Search documentation on document-level access describes why permission synchronisation is part of retrieval design, not only an application login concern.
A document is not automatically searchable just because it has been connected. The ingestion pipeline may need to extract text from PDFs, Word files, presentations and web pages, apply OCR to scans, identify tables, remove repeated headers and split long documents into smaller passages.
Splitting, often called chunking, allows the system to retrieve the section that answers a question rather than sending an entire manual to the model. The boundaries should preserve meaning. A contract clause should not be separated from its heading or an exception in the next paragraph. A table row should keep its column context. There is no universal chunk size that works equally well for policies, specifications and technical manuals.
Each passage should retain a reference to the source document and useful metadata. Project, department, document type, revision, approval status, language and publication date can all affect search quality. Without those fields, the system may retrieve a relevant-looking passage from the wrong project or an outdated version.
The pipeline also needs error handling. A failed OCR job, encrypted file or unsupported format should be visible to operators. Silent ingestion failures make the search experience look complete when important content is absent.
Embeddings turn text into numerical representations that help a search engine find passages with similar meaning. This is useful when a user's wording differs from the document. Someone asking about “supplier approval” may still find a passage titled “vendor onboarding.”
Exact terms remain important. Product codes, contract numbers, project IDs, names and dates may be better served by keyword search. A production design often combines keyword and vector retrieval, then uses metadata filters to narrow the result set. Microsoft's hybrid-search documentation describes this combination and why each method contributes different strengths.
An index should therefore contain both searchable text and the fields needed for semantic search, filtering, ranking and source display. A vector database alone is not the whole RAG architecture. Its usefulness depends on the quality of the extracted text and the logic around it.
When a question arrives, the application should first identify the user and the permitted scope. A question about Project A may also need filters for project, document status or date. The search engine then retrieves candidate passages using the methods appropriate to the query.
Initial search results are not always in the best order. A reranker can examine the question and candidate passages more closely and promote the most relevant ones. This is a second stage over retrieved results; it cannot repair a missing document or a passage that was never selected. Azure AI Search's semantic ranking guide makes that distinction clear.
The application should select enough context to answer the question without filling the model prompt with irrelevant text. It may also need to group related passages, remove duplicate copies and prefer an approved version. For high-risk questions, a minimum relevance threshold or an explicit “insufficient evidence” response is better than generating an answer from weak matches.
Enterprise RAG must respect the access rules of the source material. If an employee cannot open a confidential HR document, the assistant should not retrieve its text or summarise it. Filtering after the model has already received the content is too late.
There are several ways to implement this, depending on the identity system and search platform. The index may store access-control metadata and apply it at query time. A connector may bring source permissions into a managed knowledge service. In some cases, the application needs its own authorisation layer. The important property is that unauthorised passages never enter the context sent to the model.
Permission rules change. Group membership, project access and document sharing can be updated independently of the document text. The system must have a tested process for synchronising those changes, as well as a policy for what happens when permission metadata is missing or stale. Secure behaviour should be the default.
After retrieval, the application prepares a prompt containing the user's question, selected passages and instructions for how to answer. It may ask the model to stay within the evidence, identify uncertainty and cite the supporting documents. The answer should be returned with source titles, links and, where possible, the precise sections used.
In a common architecture, the entire company repository is not sent to the model. Only the selected context for a particular request is passed. This is useful for cost, latency and data minimisation, but it does not mean that no company data leaves private infrastructure. If the model is hosted by an external provider, the retrieved passages sent in the prompt cross that boundary. The provider's terms, retention controls, region and security configuration must be reviewed for the chosen service.
The application should treat document text as untrusted input. A retrieved page could contain instructions aimed at the assistant rather than facts for the user. The system needs clear separation between its own instructions and source content, and it should test how the model behaves when a document contains misleading or malicious text.
A well-written answer without evidence is hard to trust. Enterprise users need to see which documents support a statement, open the originals and check the version and context. Citations should point to the passages actually used, not merely to a vaguely related file.
The interface should also show when the system cannot find enough information. A useful response may be: “I found two relevant policies, but they disagree about the approval threshold. Please check the current owner.” That is more valuable than an unsupported single answer.
Some systems provide retrieval and generation as separate operations, while others combine them. For example, Amazon Bedrock Knowledge Bases exposes retrieval alone as well as retrieval with generated answers and source citations. Keeping retrieval inspectable is useful during development because the team can see whether a poor answer came from search, prompt construction or the model.
One option is a managed cloud service that handles ingestion, search and model access. It can reduce infrastructure work, but the company still needs to configure connectors, permissions, data handling and evaluation.
A second option keeps source documents and the search index in infrastructure controlled by the company, while using an external language model for the selected context. This gives more control over the knowledge layer and integration logic, though the selected passages still go to the model provider.
A third option runs the search system and model in private infrastructure. This may be appropriate for strict data residency or network requirements, but it adds hardware, model operations and upgrade responsibilities. The most private deployment is not automatically the best one; the architecture should fit the actual information, security and usage requirements.
A demonstration with ten clean files can hide the issues that determine whether enterprise RAG works. Real repositories contain duplicates, scans, obsolete versions, missing metadata and access rules that change. They also contain questions the source documents cannot answer.
Before rollout, the team should build a test set of real user questions with expected source documents. Measure whether the correct passages are retrieved, whether citations support the answer, whether unauthorised material is excluded and whether the system admits when evidence is insufficient. Repeat those tests after changes to connectors, parsing, indexing, ranking or models.
Operations need attention too. Monitor ingestion failures, index freshness, search latency, model cost and user feedback. Define who owns document quality and who investigates a bad answer. A RAG system is an ongoing information service, not a one-time chatbot deployment.
The strongest enterprise RAG projects begin with a specific group of users, a defined collection of documents and questions that matter to their work. From there, the team can design ingestion, permissions and retrieval around measurable needs.
The goal is not merely to connect an LLM to a vector database. It is to build a dependable path from a user's question to the right authorised evidence, then present an answer that can be checked. At Codativity, we design private AI knowledge systems and enterprise search around that complete path. If your organisation is evaluating RAG over internal data, we can help review the sources, architecture and first use case before implementation begins.

5/10/2026
701, Opal Tower Business Bay, Dubai United Arab Emirates
P.O.Box 126732
Let's Work Together
© 2016 - 2026 Codativity Software Solutions. All Rights Reserved.