FutureGenSystems
ServicesPrivate AIMigrateProjectsProcessAbout
Book a scoping call →
FutureGenSystems
ServicesPrivate AIMigrateProjectsProcessAbout
Book a scoping call →
  1. Home
  2. /Blog
  3. /AI & automation
  4. /Building a RAG assistant on Azure OpenAI and Azure AI Search, lessons learned
AI & automation·4 Oct 2026·5 min read

Building a RAG assistant on Azure OpenAI and Azure AI Search, lessons learned

Practical lessons for building a RAG assistant on Azure OpenAI and AI Search: chunking, permission filters, citations, evaluation and keeping data in-tenant.

By FutureGen Systems

A RAG assistant on Azure OpenAI can look finished after an afternoon. You point an index at some PDFs, wire up a chat endpoint, ask three questions you already know the answers to, and it works. The gap between that demo and something people rely on every day is where most of the effort goes. We have built retrieval-augmented assistants over custom knowledge bases on Azure OpenAI and Azure AI Search, including Medi-Insights and Mufai, and these are the lessons we would pass on to anyone starting one.

None of this is exotic. It is mostly about treating retrieval as the product and the language model as the last step.

Why retrieval quality decides everything

In retrieval-augmented generation, the model only knows what you hand it in the prompt. If the search step returns the wrong passages, the model will either say it does not know or, worse, write a fluent answer from the wrong material. A large share of "the AI made it up" complaints trace back to retrieval, not to the model.

So the work splits roughly into four parts: getting documents into the index in a useful shape, finding the right chunks for a question, presenting them so the answer can be checked, and measuring whether any of it is working.

Chunking: smaller is not automatically better

Documents have to be split into chunks before they are embedded and indexed. The defaults in most tutorials, a fixed number of tokens with some overlap, are a reasonable start and a poor finish.

  • Split on structure first. Headings, numbered clauses, sections and table boundaries carry meaning. A chunk that starts halfway through clause 7.2 and ends in 7.3 is hard to retrieve and hard to cite.
  • Keep context with each chunk. Prepend the document title and the heading path ("Employee Handbook > Leave > Parental leave") to the chunk text. It costs a few tokens and makes both keyword and vector matching noticeably more reliable.
  • Handle tables deliberately. Tables flattened into a stream of cells are close to useless. Either keep small tables whole or convert rows into short sentences that carry the column names.
  • Re-index on change, not on schedule only. Stale chunks are a quiet source of wrong answers. Track source modification times and re-process documents that changed.

There is no universally correct chunk size. Try two or three settings against your evaluation set (more on that below) and pick based on results, not on what a blog post said.

Use hybrid search, then rerank

Azure AI Search supports keyword search, vector search and a combination of both in one query. For business documents, hybrid search is usually the right default. Pure vector search is good at paraphrase but can miss exact terms such as product codes, clause numbers or names, which keyword search handles well. Combining them covers both.

On top of that, the semantic ranker in Azure AI Search can reorder the top results with a more expensive model. It is worth testing against your own questions: for some collections it helps a lot, for others very little.

Metadata and permission filtering

Every chunk should carry metadata from its source: document ID, title, URL, last modified date, document type, and, critically, who is allowed to see it.

The permission part is the one people try to skip. If the assistant serves more than one group of users, access control has to be enforced in the search query, not in the prompt. A common pattern is to store the allowed group IDs on each chunk and apply a filter on every query using the signed-in user's group memberships from Entra ID. Telling the model "do not reveal confidential documents" is not a control. If a chunk was retrieved, assume it can end up in the answer.

Metadata filters are also useful well beyond security: restricting to a document type, a date range, a department or a jurisdiction often improves answers more than any prompt change.

Citations make answers checkable

An answer without sources asks the user to trust it blindly. An answer with citations lets them verify it in seconds, and that changes how people use the tool.

A pattern that works well:

  • Give each retrieved chunk a short ID in the prompt and instruct the model to cite those IDs inline.
  • Map IDs back to the source document title and link in the interface, ideally to the right page or section.
  • Check after generation that every cited ID was actually among the retrieved chunks, and drop or flag any that were not.
  • When retrieval returns nothing relevant, have the assistant say so plainly instead of answering from general knowledge.

Evaluate against real questions, not demo questions

This is the lesson we would give most weight to. A demo proves that some questions work. It says nothing about the questions users will actually ask.

Build a small evaluation set early:

  1. Collect 30 to 100 real questions from the people who will use the assistant, including awkward and ambiguous ones.
  2. For each, record which document or section should answer it and what a correct answer contains.
  3. Measure retrieval separately from generation. First check: did the right chunk appear in the top results at all? If not, no prompt will fix it.
  4. Then grade the answers for correctness, whether they stayed within the sources, and whether the citations point to the right place. A model can help with grading, but spot-check it by hand.
  5. Re-run the set every time you change chunking, search settings, prompts or the model version.

It is unglamorous work, and it is the only reliable way to know whether a change made things better or just different.

Keeping data in your tenant

For many organisations the reason to build on Azure in the first place is that the data never needs to leave their own environment. That only holds if every component is placed deliberately:

  • Deploy Azure OpenAI, Azure AI Search, storage and the application in the organisation's own subscription and chosen region.
  • Use private endpoints and disable public network access where the setup allows, so services talk over the private network.
  • Use managed identities between services rather than API keys pasted into configuration.
  • Keep prompt and response logs in the same tenant, with retention rules that match policy.
  • Review the data-handling terms for the specific Azure OpenAI deployment type you are using, since they are what you will be asked about.

The same pattern is what we install for professional firms in our Private AI Workspace.

Two mistakes worth avoiding

The first is starting the evaluation set too late. It is tempting to tune prompts while stakeholders watch and only measure properly at the end. Even twenty real questions collected on day one save rework on chunking and search settings later.

The second is underestimating document clean-up. Scanned PDFs, duplicated versions and outdated policies all end up in the index unless someone decides what belongs there.

If you are planning a retrieval-augmented assistant and want a second pair of eyes on the design, our AI and automation work covers this, and you can get in touch to talk it through.

  • RAG
  • Azure OpenAI
  • Azure AI Search
  • Retrieval
  • LLM evaluation

Written by FutureGen Systems.

ShareLinkedIn ↗X ↗Email ↗

Keep reading

More in AI & automation →
AI & automation

Why law and accounting firms need a private AI assistant, not public chatbots

A private AI assistant for law and accounting firms keeps client documents in your own cloud. What it looks like, how it is secured and what to ask first.

4 Oct 2026·6 min read→
Building products

Browser-based SSH key management without shared keys

Browser-based SSH key management with BastionSSH: per-person keys, roles, audit logs and rotation, plus the SSH hygiene habits we follow.

4 Oct 2026·5 min read→
Building products

Cookieless, self-hosted product analytics: why we built Linqry

Cookieless, self-hosted product analytics without the GDPR risk or per-event bills: why we built Linqry, how it works and what it trades off.

4 Oct 2026·5 min read→
Free scoping call — no obligation

Have a project in mind?

Tell us what you are building. We will come back with an architecture, a timeline, and a fixed scope — no obligation.

Start a conversation →Browse our services
FutureGen Systems

Engineer-led product studio building AI integrations, cloud infrastructure, and full-stack applications — and running four products of our own in production. Based in Dehradun, working with clients across India and abroad.

11 Kanwali Road, Dehradun, Uttarakhand 248001, India+91 70172 39393[email protected]
Company
AboutHow we workProjectsOpen sourceBlogContactFAQ
Services
Private AI WorkspaceCloud RepatriationCloud & DevOpsAI & ML integrationCustom ERPFull-stack webAll services →
© 2026 FutureGen Systems. All rights reserved.
Privacy policyTerms of serviceLinkedInGitHubInstagram