UK software developers since 2006 · fixed-price quotes · you own the code01623 650333 · info@dijitul.uk
Free chat

RAG Knowledge Base Development UK

dijitul developments builds retrieval-augmented generation (RAG) knowledge bases for UK businesses: AI search and question answering over your own policies, manuals, tickets and files, with every answer citing its sources and respecting who is allowed to see what. We measure retrieval quality on your real questions. Every project is a fixed-price quote after a free chat.

Updated 2026-10-10 · by the dijitul development team, Mansfield, UK

  • PostgreSQL
  • pgvector
  • Embeddings
  • Hybrid search
  • Microsoft Graph
  • Python
  • Laravel
  • Anthropic Claude API

Sound familiar?

  • Staff ask colleagues because searching the intranet never finds the right document
  • Answers live in old tickets and emails that nobody can search
  • Three versions of the same procedure exist and people use the wrong one
  • New starters take months to learn where everything is
  • A general AI tool gives fluent answers that ignore your actual policies
  • Some documents are confidential, so a single shared AI search would leak them

Key facts

  • RAG finds relevant passages from your content and gives them to a model to answer from
  • Answers cite the source document and section so users can check them
  • Hybrid search combines vector similarity with keyword search for part numbers and names
  • Usually built on PostgreSQL with pgvector, so vectors sit next to your existing data
  • Document-level permissions are enforced at retrieval time, not by the prompt
  • Retrieval and answers are measured against a test set of your real questions
  • Content syncs from SharePoint, Google Drive, help centres, databases and file shares

How retrieval-augmented generation works

Retrieval-augmented generation, or RAG, is the standard way to make a language model answer from your content rather than its general training. It has two halves:

  1. Indexing. Documents are parsed, split into passages (chunks), and each chunk is turned into an embedding: a list of numbers that represents its meaning. The embeddings are stored in a vector index along with the text and metadata such as title, date, department and access rules.
  2. Answering. When someone asks a question, it is embedded the same way, the closest chunks are retrieved, and the model is asked to answer using only those chunks and to cite them.

The model provides the reading and writing. The retrieval step decides whether the answer is right, which is why most of our effort goes there.

Getting retrieval right

Most poor RAG systems fail at retrieval, not generation. The details that make the difference:

  • Chunking that follows structure: splitting by headings and sections rather than fixed character counts, keeping tables intact and carrying the document title into each chunk.
  • Hybrid search: vector similarity finds passages with the same meaning; keyword search (PostgreSQL full-text or BM25) finds exact part numbers, product codes and names that embeddings handle poorly. We combine the two.
  • Re-ranking: a re-ranking model reorders the top candidates so the most relevant passages reach the model.
  • Metadata filters: restrict by date, department, product line or document status so superseded procedures are not used.

We measure all of this on a test set of real questions, checking whether the right passage was retrieved before we ever look at the written answer.

Why pgvector is often the right store

Dedicated vector databases exist, and some projects need one. For many businesses, though, the pgvector extension for PostgreSQL is the better fit. It stores embeddings in an ordinary table, supports approximate nearest-neighbour indexes such as HNSW, and lets one SQL query combine vector similarity, full-text search, permission joins and metadata filters. Your vectors are backed up, secured and hosted alongside the rest of your data, with one fewer system to run.

If you are on MySQL or another database, we can run a separate PostgreSQL instance for search or use a managed vector service. We decide based on volume, hosting and what your team can maintain.

Permissions and data protection

The quickest way to cause a data breach with RAG is to index everything into one pool and let anyone ask anything. We carry access rules from the source system, such as SharePoint permissions through Microsoft Graph, into each chunk's metadata and filter at retrieval time, so a user's question can only ever retrieve chunks they could open themselves. The OWASP Top 10 for LLM Applications 2025 lists vector and embedding weaknesses as LLM08, and cross-user leakage is the main one.

We also handle deletion properly: when a document is removed or a person exercises their right to erasure, the matching chunks and embeddings go too.

Keeping the index fresh

A knowledge base is only as good as its latest sync. We build connectors that watch for changes rather than re-importing everything: SharePoint and OneDrive change tracking through Microsoft Graph, webhooks or modified dates from help desks and wikis, and database triggers or scheduled queries for structured data. New and edited documents are re-chunked and re-embedded, and deleted ones are removed from the index, usually within minutes.

The admin view shows what was indexed, what failed to parse, and which questions returned weak results. Those weak results are often the most useful report of all: they tell you which procedures are missing, out of date or written so vaguely that nobody, human or model, can find the answer. If you change embedding model later, we re-embed the corpus in the background and switch over once the test set confirms the new index is at least as good.

Where it gets used

A RAG knowledge base can sit behind several front ends: an internal search page, a staff assistant in Microsoft Teams, a customer-facing chatbot, or an API that your other applications call. Typical sources are policies and procedures, product manuals and specifications, past support tickets, tender responses and contract libraries.

We often pair it with an intranet or customer portal. Every build is a fixed-price quote after a free chat.

What we deliver

  • Connectors that pull content from SharePoint, Google Drive, Confluence, help desks or databases
  • A processing pipeline for parsing, cleaning, chunking and embedding documents
  • A vector and keyword index, typically PostgreSQL with pgvector and full-text search
  • Permission filtering mapped from your source system's access rules
  • A search and answer interface with citations, or an API for your own apps
  • A retrieval evaluation set with scoring for recall and answer quality
  • Scheduled re-indexing and an admin view of stale or missing content

How it works and what it costs

Every project gets a fixed-price quote after a free initial chat and a short scoping stage. You own the code and the data.

  1. Free chat

    Tell us the problem in plain English: what you do now, what goes wrong and what "better" looks like. No charge, no obligation.

  2. Scoping

    We map the processes, systems and data involved, agree what is in and out, and write it down so there are no surprises.

  3. Fixed-price quote

    You get a fixed price for the agreed scope, or a phased plan for bigger builds, so you can start small and prove it works.

  4. Build and test

    We build in short stages you can see and try, test against real data, then go live carefully with a rollback plan.

  5. Hand over and look after

    You own the code and the data. We can host it, support it and keep improving it, or hand it to your own team.

Frequently asked questions

What is a RAG knowledge base?

A RAG knowledge base indexes your own documents so an AI model can find relevant passages and answer questions from them, citing its sources. RAG stands for retrieval-augmented generation. dijitul developments builds them over policies, manuals, tickets and file shares, with permissions enforced at retrieval time.

Why not just fine-tune a model on our documents?

Fine-tuning teaches style and patterns, but it is a poor way to store facts that change and it cannot show sources. RAG keeps knowledge in your documents, updates as soon as they change, cites where each answer came from and respects permissions. dijitul developments uses fine-tuning only where there is a clear reason.

Do we need a separate vector database?

Often not. The pgvector extension adds vector search to PostgreSQL, so embeddings sit beside your existing data and one query can combine similarity, keyword search and permission filters. dijitul developments recommends a dedicated vector service only when volume or hosting makes it the better choice.

Can it respect SharePoint permissions?

Yes. We read access rules from SharePoint and OneDrive through Microsoft Graph, store them with each chunk and filter results for the signed-in user at search time. Users can only retrieve passages from documents they could already open.

How do you measure whether the answers are good?

We build a test set of real questions with the documents that should answer them. dijitul developments first scores retrieval, checking whether the right passages were found, then scores the written answers for accuracy and correct citation. The suite reruns whenever content processing, prompts or models change.

How much does a RAG knowledge base cost?

It depends on the content sources, volume, permissions and front ends. dijitul developments gives a fixed-price quote after a free chat and a scoping stage. Ongoing costs are hosting plus embedding and model usage, which we estimate in advance and cap.

Related

Tell us what you need to build

Free chat, clear scope, fixed-price quote. You own everything we build.

Call usFree chat