Enterprise RAG on Azure and Databricks: From Prototype to Production
Retrieval-augmented generation, or RAG, is the most practical way for enterprises to put large language models to work on their own knowledge. Instead of relying on what a model memorized during training, a RAG application retrieves relevant, current, permissioned content and grounds its answer in it, with citations. A convincing demo takes a week. A system that employees and customers can trust takes deliberate engineering. This guide covers the path from prototype to production on Azure and Databricks.

Why RAG, and when not to use it
RAG fits when answers must reflect your organization's documents, policies, products, or data, when that content changes frequently, and when users need to verify sources. It is usually faster, cheaper, and easier to govern than fine-tuning a model on proprietary content.
It is not the answer to everything. If the question is really an analytical query over structured data, such as 'what were Q3 sales by region?', route it to SQL through a governed semantic layer, for example a Genie agent over curated gold tables, instead of asking a language model to read documents. Many successful enterprise assistants combine both.
The reference architecture
A production RAG system has five layers, each with its own quality bar:
- Content pipeline: ingest documents from SharePoint, file shares, wikis, and databases; parse, clean, and chunk them; and keep everything in governed Delta tables.
- Index: embed chunks and store them in a retrieval index such as Databricks AI Search (formerly Vector Search) or Azure AI Search, with hybrid keyword and vector search.
- Retrieval and orchestration: rewrite the query if needed, retrieve candidates, rerank them, enforce permissions, and assemble a grounded prompt.
- Generation: call a model such as Azure OpenAI in Microsoft Foundry Models or a Databricks-hosted model, instructed to answer only from the provided context and cite sources.
- Evaluation and operations: trace every request, measure quality continuously, collect feedback, and control cost and access through a gateway.
Data preparation decides quality
Most poor RAG answers are retrieval failures, and most retrieval failures start with how documents were prepared.
- Parse documents with layout awareness so tables, headings, and lists survive extraction.
- Chunk by structure, such as sections and headings, rather than fixed character counts, and keep modest overlap.
- Attach metadata to every chunk: source, title, section, owner, effective date, and access groups.
- Remove duplicates and superseded versions; conflicting documents produce conflicting answers.
- Build the pipeline incrementally with Lakeflow so updates and deletions in sources flow through to the index automatically.
Retrieval that actually finds the answer
- Use hybrid search. Keyword matching catches product codes, names, and acronyms that pure vector search misses.
- Add reranking to reorder candidates by true relevance before they reach the model.
- Filter by metadata, such as region, product line, or document status, to narrow the search space.
- For complex questions, use agentic retrieval that decomposes a question into subqueries and runs them in parallel; Azure AI Search provides this through knowledge bases that also power Foundry IQ.
Security and governance are non-negotiable
An assistant that reveals documents a user should not see is worse than no assistant.
- Enforce document-level permissions at retrieval time using the user's identity, not a shared service account.
- Keep source content, chunks, and indexes under Unity Catalog governance with lineage back to the original documents.
- Use private networking and managed identities for all service-to-service communication.
- Apply content safety filters, prompt injection defenses, and rate limits through an AI gateway, and log inputs and outputs for audit.
- Define data retention rules for prompts and responses that align with your privacy and compliance obligations.
Evaluate before and after launch
RAG quality must be measured, not eyeballed. Build an evaluation dataset of real questions with expected answers and source documents, reviewed by subject matter experts. Then measure retrieval quality (did we find the right sources?), groundedness (is the answer supported by them?), correctness, and safety.
MLflow 3 on Databricks provides tracing for each step of the chain, built-in and custom LLM judges, evaluation datasets, a review app for expert feedback, and production monitoring with the same scorers used in development. Microsoft Foundry offers comparable evaluation and observability for applications built on Azure. Run evaluations on every change to prompts, models, chunking, or retrieval settings, and block releases that regress.
Operate it like a product
- Track adoption, answer acceptance, escalation rates, latency, and cost per conversation.
- Route simple questions to smaller, cheaper models and reserve larger models for complex reasoning.
- Cache frequent answers and embeddings where content changes slowly.
- Give users an easy way to flag wrong answers, and turn that feedback into new evaluation cases.
- Assign a business owner for content freshness; stale sources are the fastest way to lose trust.
Databricks AI Search or Azure AI Search?
Both are strong choices, and the right one depends on where your content and applications live.
- Choose Databricks AI Search when source content is already processed in the lakehouse, when you want indexes to sync automatically from Delta tables, and when governance should stay entirely within Unity Catalog. It pairs naturally with Agent Bricks and MLflow for building and evaluating agents.
- Choose Azure AI Search when applications are built in Microsoft Foundry or the broader Azure application stack, when you need its built-in connectors, semantic ranking, and agentic retrieval through knowledge bases, or when multiple agents should share a managed, permission-aware knowledge layer through Foundry IQ.
- Many enterprises use both: the lakehouse prepares and governs content, and the index is chosen per application.
Failure modes to design against
- Confident answers from outdated documents: fix with content ownership, effective dates, and filtering on document status.
- Answers that ignore the retrieved context: fix with stronger grounding instructions, groundedness evaluation, and refusing to answer when evidence is weak.
- Missing the obvious document: fix with hybrid search, better chunking, metadata, and reranking.
- Permission leaks: fix by enforcing access at retrieval time and testing with users from different groups.
- Cost creep: fix with model routing, caching, prompt size limits, and per-application budgets.
Measuring business value
Agree on value metrics before the pilot starts. For internal assistants, measure time saved per query, reduction in tickets or emails to expert teams, and user satisfaction. For customer-facing assistants, measure containment rate, resolution time, and customer satisfaction, alongside escalation quality. Tie these to the quality metrics from evaluation so you can see how improvements in retrieval translate into business outcomes.
From pilot to production in phases
- Weeks 1 to 3: choose one well-bounded use case with clear owners, such as policy Q&A for HR or product support for service agents. Build the content pipeline and a baseline with hybrid search.
- Weeks 4 to 6: create the evaluation set, tune chunking and retrieval, add reranking and permissions, and run a pilot with a small user group.
- Weeks 7 to 10: harden security, add monitoring and feedback loops, define support processes, and roll out gradually with success metrics agreed in advance.
The takeaway
Enterprise RAG succeeds on the fundamentals: well-prepared content, strong retrieval, strict permissions, and continuous evaluation. Azure and Databricks provide every building block, from Lakeflow and Unity Catalog to Databricks AI Search, Azure AI Search, Microsoft Foundry, and MLflow. The difference between a demo and a dependable assistant is the engineering discipline you wrap around them.
Planning a generative AI assistant on your own data? Data Minds builds secure, evaluated RAG solutions on Azure and Databricks. Reach out through our contact page to explore your first use case.



