Everyday Data Science
Latest
Agentic workflows now power a third of surveyed enterprise automationAfrica's AI startup ecosystem posts record funding yearNew benchmark results reshape the coding-agent leaderboardNigeria launches national AI strategy with major investment planRwanda's sovereign AI cloud enters public betaThe future of AI agents: from tools to teammates
Agentic AITutorial

Build a RAG System That Knows When It Doesn't Know

A practical AI engineering tutorial on building a retrieval-augmented generation system with semantic search, confidence checks, grounded answers, and a simple evaluation layer.

Most RAG tutorials stop when the chatbot produces an answer.

That is not the hard part.

The harder engineering problem is deciding whether the retrieved evidence is actually good enough to answer the question in the first place.

In this tutorial, we will build a small retrieval-augmented generation system in Python that:

  • converts documents into embeddings,

  • retrieves the most relevant evidence for a question,

  • gives that evidence to an LLM,

  • refuses to answer when retrieval confidence is too low,

  • and exposes enough information to debug the pipeline when it fails.

By the end, you will have a small but useful pattern for building grounded AI systems rather than simply wrapping an API around a prompt.

What you'll need

This tutorial uses:

  • Python 3.10+

  • openai

  • sentence-transformers

  • numpy

  • an OpenAI API key

  • a small collection of text documents

Install the dependencies:

pip install openai sentence-transformers numpy

Set your API key as an environment variable:

export OPENAI_API_KEY="your_api_key_here"

On Windows PowerShell:

$env:OPENAI_API_KEY="your_api_key_here"

The architecture is deliberately small:

User question
     ↓
Embedding model
     ↓
Similarity search
     ↓
Top matching documents
     ↓
Confidence check
     ↓
LLM + retrieved evidence
     ↓
Grounded answer

There is no vector database, orchestration framework or agent library.

That is intentional.

Before adding infrastructure, I want to be able to see exactly what retrieval is doing.

Step 1: Create a small knowledge base

We will use a fictional company's internal policy documents.

Each document contains an ID and the text we want our system to search.

documents = [
    {
        "id": "remote-work",
        "text": """
        Employees may work remotely up to three days per week.
        Tuesdays and Thursdays are designated in-office collaboration days.
        Fully remote arrangements require approval from the employee's
        department director.
        """
    },
    {
        "id": "vacation",
        "text": """
        Full-time employees receive 20 paid vacation days per calendar year.
        Vacation requests longer than five consecutive working days should
        be submitted at least three weeks in advance.
        """
    },
    {
        "id": "equipment",
        "text": """
        Employees may request a company laptop, external monitor, keyboard,
        and mouse. Home-office furniture is not covered by the standard
        equipment program.
        """
    },
    {
        "id": "training",
        "text": """
        Each employee receives an annual professional-development budget
        of $1,500. The budget may be used for courses, conferences,
        certification exams, and technical books with manager approval.
        """
    }
]

In a production RAG system these records might come from:

  • PDFs,

  • support articles,

  • database rows,

  • Google Drive documents,

  • product documentation,

  • or internal company policies.

The important idea is that retrieval works on units of evidence.

Before embedding anything, ask:

What should one retrievable unit represent?

A whole 40-page PDF is usually too large.

A sentence may be too small.

For many applications, a paragraph or short section is a reasonable starting point.

This is a data-engineering decision before it is an LLM decision.

Step 2: Convert the documents into embeddings

A language model works with tokens.

A retrieval system needs a way to compare the meaning of one piece of text with another.

That is where embeddings come in.

An embedding maps text into a vector:

"Can employees work from home?"
            ↓
[0.021, -0.314, 0.087, ..., 0.191]

Texts with similar meanings should generally occupy nearby regions of the embedding space.

We will use Sentence Transformers:

from sentence_transformers import SentenceTransformer

embedding_model = SentenceTransformer(
    "sentence-transformers/all-MiniLM-L6-v2"
)

document_texts = [doc["text"].strip() for doc in documents]

document_embeddings = embedding_model.encode_document(
    document_texts,
    normalize_embeddings=True
)

print(document_embeddings.shape)

You should see output similar to:

(4, 384)

We now have four documents represented by 384-dimensional vectors.

The interesting part is that we do this once for the documents.

For each incoming question, we only need to embed the new query.

Sentence Transformers distinguishes between document and query encodings for retrieval-oriented workloads. For question-to-passage retrieval, this is an asymmetric search problem: the query is short while the candidate document is usually longer.

Step 3: Retrieve evidence before touching the LLM

Now we can write the retrieval function.

Because our embeddings are normalized, their dot product gives us cosine similarity.

import numpy as np


def retrieve(query, top_k=2):
    query_embedding = embedding_model.encode_query(
        query,
        normalize_embeddings=True
    )

    scores = np.dot(document_embeddings, query_embedding)

    ranked_indexes = np.argsort(scores)[::-1][:top_k]

    results = []

    for index in ranked_indexes:
        results.append(
            {
                "id": documents[index]["id"],
                "text": documents[index]["text"].strip(),
                "score": float(scores[index])
            }
        )

    return results

Let's test it:

question = "How many days can I work from home?"

results = retrieve(question)

for result in results:
    print(
        result["id"],
        round(result["score"], 3)
    )

A typical result should look roughly like:

remote-work 0.66
vacation 0.25

The exact scores can vary with library and model versions.

But the ordering is what matters.

The remote-work document should clearly outrank the unrelated documents.

This is an important engineering checkpoint.

Do not call the LLM yet.

First inspect whether retrieval itself works.

If the wrong evidence enters the prompt, a better prompt cannot reliably repair the pipeline.

Step 4: Add a retrieval confidence gate

Now consider this question:

question = "Does the company provide dental insurance?"

Our knowledge base contains no information about dental insurance.

But vector search will still return something.

Nearest-neighbor retrieval always has a nearest neighbor.

That does not mean the neighbor is relevant.

This is one of the easiest mistakes to make in a RAG system.

Let's introduce a minimum similarity threshold.

MIN_SCORE = 0.40


def retrieve_with_confidence(query, top_k=2):
    results = retrieve(query, top_k)

    if not results:
        return {
            "answerable": False,
            "results": []
        }

    best_score = results[0]["score"]

    return {
        "answerable": best_score >= MIN_SCORE,
        "results": results
    }

Now test it:

result = retrieve_with_confidence(
    "Does the company provide dental insurance?"
)

print("Answerable:", result["answerable"])

for item in result["results"]:
    print(item["id"], round(item["score"], 3))

You may get something like:

Answerable: False
equipment 0.18
training 0.15

The exact scores are not universal.

That is why 0.40 should not be treated as some magical RAG constant.

It is a parameter we will eventually calibrate using evaluation data.

The principle matters more than the number:

retrieved something

is not equivalent to:

retrieved sufficient evidence

Step 5: Generate only from retrieved evidence

Now we are ready to introduce the language model.

The current OpenAI Python SDK can use the Responses API by creating an OpenAI client and calling client.responses.create(...).

from openai import OpenAI

client = OpenAI()

Next, build a function that combines retrieval and generation.

def answer_question(question):
    retrieval = retrieve_with_confidence(question)

    if not retrieval["answerable"]:
        return {
            "answer": (
                "I don't have enough information in the "
                "knowledge base to answer that."
            ),
            "sources": [],
            "retrieval": retrieval["results"]
        }

    context_parts = []

    for item in retrieval["results"]:
        context_parts.append(
            f"[SOURCE: {item['id']}]\n{item['text']}"
        )

    context = "\n\n".join(context_parts)

    prompt = f"""
You answer questions using only the supplied context.

Rules:
1. Do not use outside knowledge.
2. If the context does not support the answer, say so.
3. Cite the source IDs used in your answer.
4. Do not invent policies.

CONTEXT:

{context}

QUESTION:

{question}
"""

    response = client.responses.create(
        model="gpt-5",
        input=prompt
    )

    return {
        "answer": response.output_text,
        "sources": [
            item["id"]
            for item in retrieval["results"]
        ],
        "retrieval": retrieval["results"]
    }

Now ask:

result = answer_question(
    "How many days can employees work remotely?"
)

print(result["answer"])

The answer should say that employees may work remotely up to three days per week and should cite remote-work.

But ask:

result = answer_question(
    "What dental insurance does the company offer?"
)

print(result["answer"])

and the retrieval layer should stop the request before generation:

I don't have enough information in the knowledge base to answer that.

That refusal is a feature.

An AI system that knows when its evidence is weak is often more useful than one optimized to answer every question.

Step 6: Return retrieval metadata

For a demo, returning only the answer looks clean.

For an engineering system, it hides the most useful debugging information.

Keep the evidence.

result = answer_question(
    "Can I spend my training budget on a certification?"
)

print("ANSWER")
print(result["answer"])

print("\nRETRIEVAL")

for item in result["retrieval"]:
    print(
        item["id"],
        round(item["score"], 3)
    )

Now we can distinguish several different failure modes.

For example:

Question
   ↓
Wrong document retrieved
   ↓
Correct answer impossible

is a retrieval problem.

While:

Question
   ↓
Correct document retrieved
   ↓
Model gives unsupported answer

is a generation or instruction-following problem.

Those should not be debugged in the same way.

This distinction sounds obvious.

In practice, teams often look at a bad final answer and immediately start changing the prompt.

That can hide the real problem upstream.

Step 7: Build a tiny evaluation set

We should not decide whether the system works by asking it three questions manually.

Create expected retrieval cases.

evaluation_set = [
    {
        "question": "How many remote days do employees get?",
        "expected_source": "remote-work"
    },
    {
        "question": "Can I use company money for a certification exam?",
        "expected_source": "training"
    },
    {
        "question": "How much paid vacation do employees receive?",
        "expected_source": "vacation"
    },
    {
        "question": "Will the company buy me a monitor?",
        "expected_source": "equipment"
    }
]

Then measure whether the correct document appears first.

correct = 0

for example in evaluation_set:
    results = retrieve(
        example["question"],
        top_k=1
    )

    predicted = results[0]["id"]
    expected = example["expected_source"]

    passed = predicted == expected

    if passed:
        correct += 1

    print(
        example["question"],
        "->",
        predicted,
        "PASS" if passed else "FAIL"
    )

accuracy = correct / len(evaluation_set)

print(f"\nRetrieval accuracy: {accuracy:.2%}")

For this deliberately simple dataset, we would expect output like:

How many remote days do employees get? -> remote-work PASS
Can I use company money for a certification exam? -> training PASS
How much paid vacation do employees receive? -> vacation PASS
Will the company buy me a monitor? -> equipment PASS

Retrieval accuracy: 100.00%

Do not get excited about 100%.

We tested four easy questions against four clean documents.

This is a smoke test, not evidence that we built a production-quality retrieval system.

Step 8: Test questions that should NOT be answered

Positive tests are only half the problem.

We also need negative examples.

negative_questions = [
    "What dental insurance do employees get?",
    "Does the company match 401k contributions?",
    "Can employees work from another country?",
    "What is the parental leave policy?"
]

Test them:

for question in negative_questions:
    result = retrieve_with_confidence(question)

    top_score = result["results"][0]["score"]

    print(
        question,
        "->",
        round(top_score, 3),
        "ANSWER" if result["answerable"] else "REJECT"
    )

This test may reveal something uncomfortable.

Some unsupported questions may receive surprisingly high similarity scores.

That is valuable information.

Instead of choosing a confidence threshold because 0.40 looked reasonable, collect:

supported question → similarity score
unsupported question → similarity score

for dozens or hundreds of examples.

Then examine their distributions.

You want a threshold that separates the two groups reasonably well for your data.

That changes the problem from:

What similarity threshold should a RAG system use?

to:

What threshold gives an acceptable false-answer rate for this application's evaluation set?

That is a much better engineering question.

Where it breaks

This implementation is intentionally simple.

That makes its weaknesses easy to see.

1. Similarity is not truth

A high embedding similarity score means that two pieces of text are semantically related.

It does not prove that one answers the other.

Consider:

Question:
Can employees work remotely from another country?

Document:
Employees may work remotely up to three days per week.

Those texts are highly related.

But the document does not answer the international-work question.

A similarity threshold alone cannot solve every ambiguity.

2. Chunking can destroy context

Suppose a real policy says:

Employees may work remotely three days per week.

Exceptions:
Employees handling restricted customer data may only work
from approved company locations.

If those paragraphs are separated into different chunks, retrieval might find the first paragraph but miss the exception.

Your chunking strategy becomes part of system correctness.

3. Top-k retrieval can add noise

More context is not automatically better context.

Retrieving ten passages instead of three can introduce unrelated evidence and make generation harder.

Measure top-k rather than assuming that larger values are safer.

4. Embedding models have different retrieval behavior

A model that works well for sentence similarity is not automatically the best model for question-to-document retrieval.

Sentence Transformers explicitly distinguishes symmetric and asymmetric semantic search and provides retrieval-oriented models and query/document encoding interfaces.

5. Retrieval accuracy is not answer accuracy

Even perfect retrieval does not guarantee a correct generated answer.

A real evaluation pipeline should separately measure:

Retrieval quality
        ↓
Context sufficiency
        ↓
Answer correctness
        ↓
Answer groundedness
        ↓
Citation correctness

Combining all of those into one vague "the chatbot works" metric makes debugging much harder.

6. This approach does not scale indefinitely

We compute similarity against every document.

That is completely reasonable for a small corpus, and Sentence Transformers documents direct semantic search as suitable for relatively modest corpora.

For much larger systems, I would introduce an approximate nearest-neighbor index or vector-search infrastructure.

I would not introduce it on day one unless the corpus required it.

The easiest RAG system to debug is usually the one with the fewest moving pieces.

7. The confidence threshold will drift

Suppose the knowledge base grows from:

4 documents

to:

400,000 documents.

The score distribution can change.

So can the embedding model.

So can the query distribution.

So can the documents themselves.

A threshold calibrated on version one of the system should not automatically become a permanent constant.

What would make me wrong

The main claim of this tutorial is:

A RAG system should evaluate retrieval quality and evidence sufficiency separately from generation rather than assuming that retrieving a nearest neighbor means the question is answerable.

I would reconsider that claim for a particular application if controlled evaluation showed that removing the retrieval confidence layer produced equal or better grounded-answer accuracy without increasing unsupported answers.

That test should use a held-out dataset containing both:

  • answerable questions,

  • deliberately unanswerable questions.

For example, compare:

System A
nearest-neighbor retrieval
→ always generate

System B
nearest-neighbor retrieval
→ confidence gate
→ generate or reject

Then measure at least:

retrieval recall
unsupported-answer rate
answer correctness
citation correctness
abstention precision

If System A consistently matched or outperformed System B on unseen examples, particularly on unsupported-answer rate, then the extra confidence gate would not be earning its complexity.

That is the standard I would use rather than arguing from intuition.

Key takeaways

  • Retrieval comes before generation. If the wrong evidence reaches the model, prompt engineering is unlikely to rescue the system reliably.

  • Nearest does not mean relevant. Vector search will return something even when the knowledge base contains no answer, so production systems need a way to reason about evidence sufficiency.

  • Evaluate the pipeline in pieces. Retrieval quality, context sufficiency, answer correctness and groundedness are different problems. Measuring them separately makes AI systems much easier to improve.

The most important change in mindset is this:

Do not ask:

"Did the model give me a good answer?"

Ask:

"What evidence did we retrieve?
Was that evidence sufficient?
Did the answer stay inside that evidence?"

That is where a RAG demo starts becoming an AI engineering system.

Sources

  • OpenAI, “Developer Quickstart,” OpenAI API documentation.

  • Sentence Transformers, “Semantic Search,” Sentence Transformers documentation.

  • Sentence Transformers, “Quickstart,” Sentence Transformers documentation.

  • Sentence Transformers, “Pretrained Models,” Sentence Transformers documentation.

About the writer

Rodriquez Allen
Rodriquez Allen

AI Engineer

1 follower

An AI Engineer who work the work and avoid the talks

Share

Found this useful? Passing it on to someone who builds is the best way to help the publication grow.

Built something worth sharing? Write it up for us →