← Blog

15 November 2024 · 6 min · LLM · RAG · AI

RAG pipelines for business LLM applications

How to build retrieval-augmented generation that answers from your own data, and shows where each answer came from.

What RAG is

Retrieval-augmented generation gives a language model relevant context from your own sources. Instead of relying only on what the model learned in training, it looks up the right documents first and adds them to the prompt.

Why it matters for organisations

  • ✦Accuracy — answers are based on your actual data
  • ✦Up to date — new documents are available straight away
  • ✦Control — answers stay within approved content
  • ✦Traceable — you can see which sources informed each answer

The vector store

Document embeddings are stored so they can be searched by meaning:

from langchain.vectorstores import Chroma
from langchain.embeddings import OpenAIEmbeddings

vectorstore = Chroma.from_documents(
    documents=chunks,
    embedding=OpenAIEmbeddings(),
    persist_directory="./chroma_db"
)

Retrieval

Mixing semantic and keyword matching usually beats either one alone:

retriever = vectorstore.as_retriever(
    search_type="mmr",
    search_kwargs={"k": 5, "fetch_k": 20}
)

What I've learned

  • ✦Chunk with care — balance the size of each piece against how relevant it is
  • ✦Re-rank — a cross-encoder improves precision a lot
  • ✦Measure — track retrieval quality and what users tell you

Done well, RAG turns a general chatbot into an assistant that knows your domain and can back up what it says. It's the approach behind TwinQuery.

Working on something like this?

Get in touch →

Next post

Delta Lake best practices for lakehouses

→