RAG (Retrieval-Augmented Generation) is a technique that gives LLMs access to external knowledge by retrieving relevant documents and injecting them into the prompt. Instead of relying solely on the model's training data, RAG searches a knowledge base, finds the most relevant chunks, and includes them as context. The result is more accurate, up-to-date, and verifiable responses.
Retrieval-Augmented Generation, universally known as RAG, is the most important architecture pattern in applied AI right now. It solves the fundamental limitation of large language models: they only know what was in their training data.
When you ask an LLM a question about your company's internal documentation, yesterday's sales figures, or a PDF you uploaded, the model has no way to answer accurately from its training data alone. RAG fixes this by adding a retrieval step before generation:
Think of it like giving the LLM an open-book exam instead of a closed-book exam. The model still does the reasoning, but now it has reference material to work with.
RAG matters because it enables the most common enterprise AI use case: question-answering over proprietary data. Customer support bots that reference your help docs, internal tools that search your company wiki, legal assistants that cite relevant case law. All RAG applications.
A RAG system has two phases: indexing (done once or periodically) and retrieval + generation (done at query time). Understanding both is essential.
Phase 1: Indexing (Offline)
During indexing, you prepare your knowledge base for fast, semantic search:
Phase 2: Retrieval + Generation (Online)
When a user asks a question:
A simplified example using Supabase (pgvector) and the OpenAI API:
// Indexing: embed and store a document chunk
const embedding = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: chunkText,
});
await supabase.from('documents').insert({
content: chunkText,
embedding: embedding.data[0].embedding,
metadata: { source: 'help-docs', page: 12 },
});
// Retrieval: find relevant chunks for a query
const queryEmbedding = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: userQuestion,
});
const { data: chunks } = await supabase.rpc('match_documents', {
query_embedding: queryEmbedding.data[0].embedding,
match_threshold: 0.7,
match_count: 5,
});
// Generation: answer using retrieved context
const context = chunks.map(c => c.content).join('\n\n');
const response = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [
{ role: 'system', content: `Answer based on this context:\n${context}` },
{ role: 'user', content: userQuestion },
],
});
Get weekly developer tips
Join 25,000+ developers. Practical guides, job tips, and new content — straight to your inbox.
No spam. Unsubscribe anytime.
Embeddings are the technology that makes RAG possible. An embedding is a list of numbers (a vector) that represents the meaning of a piece of text. Texts with similar meanings have vectors that are close together in the embedding space, even if they use completely different words.
For example, the sentences "How do I reset my password?" and "I forgot my login credentials" would have very similar embeddings, even though they share almost no words. This is what makes semantic search far more powerful than keyword search for RAG applications.
Embedding Models in 2026
The most commonly used embedding models are:
One rule you cannot break: use the same embedding model for indexing and querying. If you embed your documents with text-embedding-3-small, you must also embed user queries with text-embedding-3-small. Mixing models produces meaningless similarity scores.
Dimensionality matters. Higher-dimensional embeddings capture more nuance but require more storage and compute. For most applications, 1536 dimensions (text-embedding-3-small) is the sweet spot. If you are storing millions of documents, consider dimensionality reduction techniques.
A vector database is a specialized database optimized for storing and searching high-dimensional vectors. When your RAG system needs to find the 5 most relevant chunks out of 100,000 documents, the vector database makes this fast (milliseconds, not minutes).
Dedicated Vector Databases
PostgreSQL Extensions (Our Recommendation for the African Stack)
If you are already using PostgreSQL (and if you are using Supabase, you are), pgvector lets you add vector search to your existing database without introducing a new service. This is particularly valuable in the African market where minimizing infrastructure complexity and cost matters.
Supabase has first-class pgvector support. You can store embeddings alongside your regular relational data, use SQL to query them, and avoid the operational overhead of a separate vector database. For applications with fewer than a few million documents, pgvector on Supabase is the pragmatic choice.
-- Create a table with a vector column
create table documents (
id bigserial primary key,
content text,
embedding vector(1536), -- matches text-embedding-3-small
metadata jsonb,
created_at timestamptz default now()
);
-- Create an index for fast similarity search
create index on documents
using ivfflat (embedding vector_cosine_ops)
with (lists = 100);
-- Query: find the 5 most similar documents
create or replace function match_documents(
query_embedding vector(1536),
match_threshold float,
match_count int
) returns table (id bigint, content text, similarity float)
language sql stable as $$
select id, content, 1 - (embedding <=> query_embedding) as similarity
from documents
where 1 - (embedding <=> query_embedding) > match_threshold
order by embedding <=> query_embedding
limit match_count;
$$;
How you split your documents into chunks has a huge impact on retrieval quality, often more than the choice of embedding model or vector database. Poor chunking is the number one reason RAG systems give bad answers.
The goal of chunking is to create pieces of text that are:
Common Chunking Strategies
1. Fixed-size chunking. Split text every N tokens (e.g., 500 tokens) with some overlap (e.g., 50 tokens). This is the simplest approach and works surprisingly well for unstructured text. The overlap ensures you do not lose information at chunk boundaries.
2. Recursive character splitting. Try to split on paragraph breaks first, then sentences, then words. This preserves natural text boundaries. LangChain's RecursiveCharacterTextSplitter is the most popular implementation and the default recommendation for most use cases.
3. Semantic chunking. Use an embedding model to identify where the topic changes, and split there. This produces the highest-quality chunks but is slower and more expensive to compute. Use it when retrieval quality is critical and your documents have varied, complex structure.
4. Document-structure-aware chunking. Use the document's own structure (headings, sections, list items) to determine chunk boundaries. This works well for structured documents like documentation sites, legal contracts, or technical manuals.
Practical Guidelines
One of the most common questions in applied AI is whether to use RAG or fine-tuning. The short answer: start with RAG. It is cheaper, faster to implement, easier to update, and sufficient for the vast majority of use cases.
Use RAG when:
Use fine-tuning when:
Use both when:
The most powerful approach combines RAG and fine-tuning. Fine-tune the model to follow your desired format and style, then use RAG to provide it with up-to-date factual knowledge. This is what production systems at companies like Notion, Stripe, and Shopify do.
For most developers and startups, especially in the African market, RAG alone will take you very far. Fine-tuning becomes relevant when you have validated your product and need to optimize cost or quality at scale.
A basic RAG pipeline is straightforward. A production RAG system that reliably serves real users requires more thought. These techniques separate demos from products.
Hybrid Search
Combine vector search (semantic similarity) with keyword search (BM25). Vector search excels at understanding meaning but can miss exact matches. Keyword search is precise but misses semantic similarity. Hybrid search gives you the best of both. Most vector databases support this natively.
Query Rewriting
User queries are often vague, misspelled, or poorly phrased. Before searching, use an LLM to rewrite the query into a more search-friendly form. For example, "how do I fix that thing with the payments" might be rewritten to "troubleshoot M-Pesa payment integration errors."
Re-ranking
After retrieving the top 20 chunks via vector search, use a cross-encoder model (like Cohere Rerank) to re-score them based on actual relevance to the query. This second pass is slower but significantly more accurate than vector similarity alone. Retrieve broadly, then re-rank precisely.
Metadata Filtering
Do not search your entire knowledge base for every query. Use metadata to narrow the search space first. If the user is asking about billing, only search billing-related documents. If they are in a specific product context, filter by product. This improves both speed and relevance.
Evaluation
You cannot improve what you do not measure. Build an evaluation set of question-answer pairs and measure:
Tools like Ragas, DeepEval, and custom evaluation scripts can automate this. Run evaluations after every change to your chunking strategy, embedding model, or prompt template.
KES 120,000
The complete Full-Stack Software & AI Engineering course. From zero code to AI-powered products in 16 weeks. 16 modules, 208 lessons, with career support until placement.
Enrol in Full-Stack AI