15 November 2024 · 6 min · LLM · RAG · AI
RAG pipelines for business LLM applications
How to build retrieval-augmented generation that answers from your own data, and shows where each answer came from.
What RAG is
Retrieval-augmented generation gives a language model relevant context from your own sources. Instead of relying only on what the model learned in training, it looks up the right documents first and adds them to the prompt.
Why it matters for organisations
- ✦Accuracy — answers are based on your actual data
- ✦Up to date — new documents are available straight away
- ✦Control — answers stay within approved content
- ✦Traceable — you can see which sources informed each answer
The vector store
Document embeddings are stored so they can be searched by meaning:
from langchain.vectorstores import Chroma
from langchain.embeddings import OpenAIEmbeddings
vectorstore = Chroma.from_documents(
documents=chunks,
embedding=OpenAIEmbeddings(),
persist_directory="./chroma_db"
)Retrieval
Mixing semantic and keyword matching usually beats either one alone:
retriever = vectorstore.as_retriever(
search_type="mmr",
search_kwargs={"k": 5, "fetch_k": 20}
)What I've learned
- ✦Chunk with care — balance the size of each piece against how relevant it is
- ✦Re-rank — a cross-encoder improves precision a lot
- ✦Measure — track retrieval quality and what users tell you
Done well, RAG turns a general chatbot into an assistant that knows your domain and can back up what it says. It's the approach behind TwinQuery.
Working on something like this?
Get in touch →