Wednesday, 7 October 2026

Part 2 — Building a Production-Style RAG Pipeline for a Telecom AI Support Agent

Building a Production-Style RAG Pipeline for a Telecom AI Support Agent — Part 2

Introduction

In Part 1 of the GlobalNet Support Agent project, I built an AI-powered telecom customer support application using:

  • Python

  • FastAPI

  • Google Cloud

  • Vertex AI

  • Gemini

  • Multiple AI agents

The application already contained three agents:

Customer
   ↓
Upsell Agent
   ↓
Plan Agent
   ↓
Offer Agent
   ↓
Customer Response

The Upsell Agent determines whether an upgrade is justified.

The Plan Agent determines which plan is suitable.

The Offer Agent creates a personalized customer offer.

However, there was an important limitation.

The Plan Agent initially contained plan information directly inside its prompt:

BASIC     → 100 Mbps
STANDARD  → 300 Mbps
PREMIUM   → 500 Mbps
ULTRA     → 1 Gbps

This approach works for a prototype, but it does not scale.

In this part, I will convert the application into an Agent + RAG architecture using:

JSON Knowledge Base
        ↓
Document Processing
        ↓
Chunking
        ↓
Sentence Transformers
        ↓
Embeddings
        ↓
ChromaDB
        ↓
Semantic Search
        ↓
Relevance Filtering
        ↓
Gemini
        ↓
Plan Recommendation

We will also build a basic RAG evaluation framework to test whether the retriever is actually finding the correct telecom plans.


1. What is RAG?

RAG stands for:

Retrieval-Augmented Generation

A normal LLM application might work like this:

User Question
      ↓
Prompt
      ↓
LLM
      ↓
Answer

The problem is that the LLM may not know our private business information.

For example, Gemini does not automatically know the current internal GlobalNet plan catalog used by this demo application.

We could put all plan information directly into the prompt, but this becomes difficult as the knowledge base grows.

RAG changes the architecture.

User Question
      ↓
Retriever
      ↓
Knowledge Base
      ↓
Relevant Information
      ↓
LLM
      ↓
Grounded Answer

Instead of expecting the LLM to know everything, we retrieve relevant information first and provide that information as context.


2. Why RAG for a Telecom Support Agent?

Imagine that a telecom company has:

100 internet plans
500 support documents
200 troubleshooting guides
50 promotion documents
1000 FAQ entries
Product documentation
Network policies
Upgrade rules
Eligibility rules

Sending all of this information to an LLM for every customer request would be inefficient.

Instead:

Customer:
"I have 8 devices and multiple 4K TVs."
                 ↓
             Retriever
                 ↓
       Search knowledge base
                 ↓
      Relevant plan information
                 ↓
              Gemini
                 ↓
        Recommended plan

This is where RAG becomes useful.


3. RAG Architecture for GlobalNet

The architecture developed in this project is:

                         Customer
                            |
                            v
                      FastAPI API
                            |
                            v
                     Upsell Agent
                            |
                            v
                         Gemini
                            |
                            v
                      Plan Agent
                            |
                            v
                       RAG Query
                            |
                            v
                 Sentence Transformer
                            |
                            v
                     Query Embedding
                            |
                            v
                        ChromaDB
                            |
                            v
                  Similarity Search
                            |
                            v
                    Relevant Chunks
                            |
                            v
                  Relevance Filtering
                            |
                            v
                     Context Prompt
                            |
                            v
                         Gemini
                            |
                            v
                 Plan Recommendation
                            |
                            v
                      Offer Agent
                            |
                            v
                         Gemini
                            |
                            v
                  Personalized Offer

4. Updated Project Directory Structure

After adding RAG, the project structure becomes:

GlobalNet-Support-Agent-GCP/
│
├── app/
│   ├── __init__.py
│   ├── main.py
│   │
│   ├── agents/
│   │   ├── __init__.py
│   │   ├── upsell_agent.py
│   │   ├── plan_agent.py
│   │   └── offer_agent.py
│   │
│   ├── services/
│   │   ├── __init__.py
│   │   ├── vertex_ai.py
│   │   ├── plan_catalog.py
│   │   ├── rag.py
│   │   ├── firestore.py
│   │   └── email.py
│   │
│   └── models/
│       └── customer.py
│
├── data/
│   ├── plans/
│   │   └── globalnet_plans.json
│   │
│   └── rag_evaluation.json
│
├── scripts/
│   ├── ingest_plans.py
│   ├── evaluate_rag.py
│   └── evaluate_rag_top1.py
│
├── chroma_db/
│
├── test_vertex.py
├── test_rag.py
│
├── requirements.txt
├── Dockerfile
├── .dockerignore
└── README.md

5. Create the GlobalNet Knowledge Base

The first step is moving plan information outside the Python source code.

Create:

data/plans/globalnet_plans.json

Example:

[
  {
    "plan_id": "BASIC",
    "plan_name": "GlobalNet Basic",
    "speed": "100 Mbps",
    "monthly_price": 499,
    "data_limit": "Unlimited",
    "recommended_for": [
      "1-3 devices",
      "normal browsing",
      "email",
      "social media",
      "HD streaming"
    ],
    "features": [
      "Unlimited data",
      "Standard customer support",
      "Wi-Fi router"
    ]
  },
  {
    "plan_id": "STANDARD",
    "plan_name": "GlobalNet Standard",
    "speed": "300 Mbps",
    "monthly_price": 799,
    "data_limit": "Unlimited",
    "recommended_for": [
      "3-6 devices",
      "4K streaming",
      "work from home",
      "video conferencing",
      "online gaming"
    ],
    "features": [
      "Unlimited data",
      "Priority customer support",
      "Wi-Fi 6 router"
    ]
  },
  {
    "plan_id": "PREMIUM",
    "plan_name": "GlobalNet Premium",
    "speed": "500 Mbps",
    "monthly_price": 1099,
    "data_limit": "Unlimited",
    "recommended_for": [
      "6-10 devices",
      "multiple 4K streams",
      "heavy work from home",
      "online gaming",
      "large file transfers"
    ],
    "features": [
      "Unlimited data",
      "Priority customer support",
      "Wi-Fi 6 router",
      "Premium support"
    ]
  },
  {
    "plan_id": "ULTRA",
    "plan_name": "GlobalNet Ultra",
    "speed": "1 Gbps",
    "monthly_price": 1499,
    "data_limit": "Unlimited",
    "recommended_for": [
      "10+ devices",
      "multiple simultaneous 4K streams",
      "professional users",
      "heavy cloud workloads",
      "large household usage"
    ],
    "features": [
      "Unlimited data",
      "24x7 priority support",
      "Wi-Fi 6 router",
      "Premium support",
      "Dedicated service assistance"
    ]
  }
]

These are sample plans created for this learning project.

They are not intended to represent actual commercial telecom offerings.


6. Create a Plan Catalog Service

Create:

app/services/plan_catalog.py

Code:

import json

from pathlib import Path


PLAN_FILE = (
    Path(__file__).resolve().parents[2]
    / "data"
    / "plans"
    / "globalnet_plans.json"
)


def load_plans():

    with open(
        PLAN_FILE,
        "r",
        encoding="utf-8"
    ) as file:

        return json.load(file)


def get_plan(plan_id: str):

    plans = load_plans()

    for plan in plans:

        if plan["plan_id"] == plan_id:

            return plan

    return None

This separates business data from application logic.

The application can now load plan information from:

globalnet_plans.json

instead of embedding the information inside Python prompts.


7. Moving from JSON Lookup to RAG

Simply reading JSON is not RAG.

This:

JSON
 ↓
Read everything
 ↓
Put everything into prompt
 ↓
Gemini

is still a normal context-injection approach.

True RAG introduces retrieval:

Documents
    ↓
Chunking
    ↓
Embeddings
    ↓
Vector Database

         +

Customer Question
    ↓
Embedding
    ↓
Vector Search
    ↓
Relevant Documents
    ↓
Gemini

8. Install the RAG Components

For this project I used:

ChromaDB
Sentence Transformers

Install:

pip install chromadb sentence-transformers

Verify:

python -c "import chromadb; import sentence_transformers; print('RAG packages OK')"

9. What is an Embedding?

An embedding converts text into a numerical representation.

For example:

"I have many devices and multiple 4K TVs."

is converted conceptually into:

[0.13, -0.44, 0.82, 0.07, ...]

The actual embedding contains many numerical dimensions.

Another sentence with similar meaning should have a vector located relatively close in embedding space.

For example:

"I need high-speed internet for several devices and 4K streaming."

has similar semantic meaning.

This enables semantic search.


10. Embedding Model

The project uses:

SentenceTransformer(
    "all-MiniLM-L6-v2"
)

This model runs locally and converts plan documents and customer questions into embeddings.


11. Why Use a Vector Database?

A traditional database can perform queries such as:

plan_id = "PREMIUM"

But a customer does not necessarily know the plan ID.

The customer might say:

"I have 8 devices and my family watches multiple 4K streams."

We need to find the plan whose meaning best matches this requirement.

A vector database stores embeddings and performs similarity searches.

For this project I used:

ChromaDB

12. Why Chunking Matters

Initially I stored one entire plan as one document.

For example:

PREMIUM
  |
  +-- Price
  +-- Speed
  +-- Recommended Usage
  +-- Features

That works with a tiny dataset.

However, larger knowledge bases require more granular retrieval.

I therefore split each plan into three chunks.

PREMIUM
   |
   +---- PREMIUM_basic
   |
   +---- PREMIUM_usage
   |
   +---- PREMIUM_features

The same happens for all four plans.

BASIC_basic
BASIC_usage
BASIC_features

STANDARD_basic
STANDARD_usage
STANDARD_features

PREMIUM_basic
PREMIUM_usage
PREMIUM_features

ULTRA_basic
ULTRA_usage
ULTRA_features

Therefore:

4 plans × 3 chunks = 12 chunks

Each chunk receives its own embedding.


13. Complete RAG Service

Create:

app/services/rag.py

Code:

import chromadb

from sentence_transformers import SentenceTransformer


CHROMA_PATH = "./chroma_db"

COLLECTION_NAME = "globalnet_plans"


# ------------------------------------------------------------
# Embedding Model
# ------------------------------------------------------------

embedding_model = SentenceTransformer(
    "all-MiniLM-L6-v2"
)


# ------------------------------------------------------------
# ChromaDB Client
# ------------------------------------------------------------

chroma_client = chromadb.PersistentClient(
    path=CHROMA_PATH
)


# ------------------------------------------------------------
# ChromaDB Collection
# ------------------------------------------------------------

collection = chroma_client.get_or_create_collection(
    name=COLLECTION_NAME
)


# ------------------------------------------------------------
# Create Plan Chunks
# ------------------------------------------------------------

def create_plan_chunks(plans):

    documents = []
    ids = []
    metadatas = []

    for plan in plans:

        plan_id = plan["plan_id"]

        # Basic information chunk

        basic_info = f"""
Plan ID: {plan_id}

Plan Name: {plan["plan_name"]}

Speed: {plan["speed"]}

Monthly Price: ₹{plan["monthly_price"]}

Data Limit: {plan["data_limit"]}
"""

        documents.append(
            basic_info
        )

        ids.append(
            f"{plan_id}_basic"
        )

        metadatas.append({
            "plan_id": plan_id,
            "plan_name": plan["plan_name"],
            "chunk_type": "basic",
            "speed": plan["speed"],
        })


        # Recommended usage chunk

        usage_info = f"""
Plan ID: {plan_id}

Plan Name: {plan["plan_name"]}

Recommended For:

{", ".join(plan["recommended_for"])}
"""

        documents.append(
            usage_info
        )

        ids.append(
            f"{plan_id}_usage"
        )

        metadatas.append({
            "plan_id": plan_id,
            "plan_name": plan["plan_name"],
            "chunk_type": "usage",
            "speed": plan["speed"],
        })


        # Features chunk

        feature_info = f"""
Plan ID: {plan_id}

Plan Name: {plan["plan_name"]}

Features:

{", ".join(plan["features"])}
"""

        documents.append(
            feature_info
        )

        ids.append(
            f"{plan_id}_features"
        )

        metadatas.append({
            "plan_id": plan_id,
            "plan_name": plan["plan_name"],
            "chunk_type": "features",
            "speed": plan["speed"],
        })

    return documents, ids, metadatas


# ------------------------------------------------------------
# Create Embeddings
# ------------------------------------------------------------

def create_embeddings(documents):

    embeddings = embedding_model.encode(
        documents
    )

    return embeddings.tolist()


# ------------------------------------------------------------
# Store Plans
# ------------------------------------------------------------

def store_plans(plans):

    documents, ids, metadatas = create_plan_chunks(
        plans
    )

    embeddings = create_embeddings(
        documents
    )

    collection.upsert(
        ids=ids,
        documents=documents,
        embeddings=embeddings,
        metadatas=metadatas,
    )

    return len(documents)


# ------------------------------------------------------------
# Search Plans
# ------------------------------------------------------------

def search_plans(
    query,
    top_k=5,
    max_distance=0.80
):

    query_embedding = embedding_model.encode(
        [query]
    ).tolist()

    results = collection.query(
        query_embeddings=query_embedding,
        n_results=top_k,
        include=[
            "documents",
            "metadatas",
            "distances"
        ],
    )

    filtered_documents = []
    filtered_metadatas = []
    filtered_distances = []

    documents = results.get(
        "documents",
        [[]]
    )[0]

    metadatas = results.get(
        "metadatas",
        [[]]
    )[0]

    distances = results.get(
        "distances",
        [[]]
    )[0]

    for document, metadata, distance in zip(
        documents,
        metadatas,
        distances,
    ):

        if distance <= max_distance:

            filtered_documents.append(
                document
            )

            filtered_metadatas.append(
                metadata
            )

            filtered_distances.append(
                distance
            )

    return {
        "documents": [
            filtered_documents
        ],
        "metadatas": [
            filtered_metadatas
        ],
        "distances": [
            filtered_distances
        ],
    }

14. What Happens During Ingestion?

The ingestion pipeline works like this:

globalnet_plans.json
        ↓
load_plans()
        ↓
4 Plans
        ↓
create_plan_chunks()
        ↓
12 Chunks
        ↓
SentenceTransformer
        ↓
12 Embeddings
        ↓
ChromaDB

This happens before customer queries are processed.


15. Create the Ingestion Script

Create:

scripts/ingest_plans.py

Code:

import sys

from pathlib import Path


PROJECT_ROOT = Path(__file__).resolve().parents[1]

sys.path.insert(
    0,
    str(PROJECT_ROOT)
)


from app.services.plan_catalog import load_plans
from app.services.rag import store_plans


def main():

    print(
        "Loading GlobalNet plans..."
    )

    plans = load_plans()

    print(
        f"Found {len(plans)} plans"
    )

    print(
        "Creating embeddings and storing in ChromaDB..."
    )

    count = store_plans(
        plans
    )

    print(
        f"Successfully stored {count} chunks in ChromaDB"
    )


if __name__ == "__main__":

    main()

Run:

python scripts/ingest_plans.py

Expected:

Loading GlobalNet plans...

Found 4 plans

Creating embeddings and storing in ChromaDB...

Successfully stored 12 chunks in ChromaDB

16. What Gets Stored in ChromaDB?

Each chunk contains three important things:

ID
Document
Metadata
Embedding

For example:

ID:
PREMIUM_usage

Document:

Plan ID: PREMIUM

Plan Name: GlobalNet Premium

Recommended For:

6-10 devices,
multiple 4K streams,
heavy work from home,
online gaming,
large file transfers

Metadata:

{
  "plan_id": "PREMIUM",
  "plan_name": "GlobalNet Premium",
  "chunk_type": "usage",
  "speed": "500 Mbps"
}

And internally:

Embedding Vector

This makes the knowledge searchable semantically.


17. Query-Time RAG Flow

Suppose the customer says:

"I have eight devices and use multiple 4K streams."

The system does not search for an exact string.

Instead:

Customer Question
       ↓
SentenceTransformer
       ↓
Query Embedding
       ↓
ChromaDB
       ↓
Vector Similarity Search
       ↓
Closest Chunks

A likely result could include:

PREMIUM_usage

because that document contains semantically related information:

6-10 devices
multiple 4K streams

18. Testing RAG Search

Create:

test_rag.py

Code:

from app.services.rag import search_plans


query = """
I have many devices at home and use several
4K streams. I need high bandwidth.
"""


results = search_plans(
    query,
    top_k=5
)


print(
    "\n===== RAG RESULTS =====\n"
)


documents = results["documents"][0]

metadatas = results["metadatas"][0]

distances = results["distances"][0]


for i in range(
    len(documents)
):

    print(
        f"Result: {i + 1}"
    )

    print(
        f"Plan: {metadatas[i]['plan_name']}"
    )

    print(
        f"Chunk Type: {metadatas[i]['chunk_type']}"
    )

    print(
        f"Distance: {distances[i]}"
    )

    print(
        "\nDocument:"
    )

    print(
        documents[i]
    )

    print(
        "-" * 70
    )

Run:

python test_rag.py

19. What is top_k?

Our query contains:

top_k=5

This means:

Retrieve up to the 5 nearest chunks.

Conceptually:

Customer Query
      ↓
Vector Search
      ↓
Result 1
Result 2
Result 3
Result 4
Result 5

The application then decides which results are relevant enough to use.


20. Understanding Distance

ChromaDB returns distance information for retrieved vectors.

For our chosen setup, smaller distances generally represent greater similarity.

For example:

Document A
Distance = 0.25

Document B
Distance = 0.70

Document A is generally a closer semantic match.

However, distance values should not be interpreted as universal percentages.

A value such as:

0.80

does not mean:

80% relevant

The values depend on the embedding model, vector space and configured distance behavior.


21. Relevance Filtering

A vector database normally returns the nearest available documents even when none are especially useful.

Suppose the user asks:

"What is the weather today?"

Our database contains only telecom plans.

Without filtering, ChromaDB may still return the closest telecom document.

We don't want Gemini to treat that as valid knowledge.

Therefore, the RAG service applies:

if distance <= max_distance:

The current learning-project threshold is:

max_distance=0.80

Results outside the threshold are rejected.

This creates:

Query
   ↓
Vector Search
   ↓
Candidate Documents
   ↓
Distance Filter
   ↓
Relevant Documents

22. Why Relevance Thresholds Need Evaluation

The value:

0.80

is not automatically correct for every RAG application.

The best threshold depends on:

Embedding model
Document type
Chunk size
Domain
Query style
Vector database configuration
Evaluation dataset

Therefore, a production system should determine the threshold using evaluation rather than guessing.


23. Connect RAG to the Plan Agent

The Plan Agent previously received every plan.

Now it retrieves only relevant plan information.

File:

app/agents/plan_agent.py

Code:

from app.services.vertex_ai import generate_text

from app.services.rag import search_plans


def recommend_plan(
    customer,
    upsell_analysis: str
) -> str:

    query = f"""
Customer currently has:

Plan:
{customer.current_plan}

Customer problem:
{customer.chat_text}

Find the GlobalNet plan information that is
most relevant to the customer's requirements.
"""

    rag_results = search_plans(
        query,
        top_k=5,
        max_distance=0.80
    )

    documents = (
        rag_results["documents"][0]
    )

    distances = (
        rag_results["distances"][0]
    )

    if not documents:

        return """
Current Plan:
Unknown

Recommended Plan:
No recommendation

Reason:
No sufficiently relevant plan information was found
in the GlobalNet knowledge base.

Next Action:
Ask a support specialist to review the customer.
"""

    relevant_context = ""

    for document, distance in zip(
        documents,
        distances
    ):

        relevant_context += f"""
Relevance Distance: {distance}

{document}

-----------------------------
"""

    prompt = f"""
You are a telecom plan recommendation AI.

Customer Information:

Customer ID:
{customer.customer_id}

Customer Name:
{customer.customer_name}

Current Plan:
{customer.current_plan}

Current Plan Description:
{customer.current_plan_desc}

Loyalty Status:
{customer.loyalty_status}

Tenure:
{customer.tenure_months} months

Customer Message:

{customer.chat_text}

Upsell Analysis:

{upsell_analysis}

Relevant GlobalNet Knowledge:

{relevant_context}

Rules:

1. Use only information contained in the retrieved
   GlobalNet knowledge.

2. Never invent a plan.

3. Never invent a price.

4. Never invent a speed.

5. Do not recommend an upgrade just to increase revenue.

6. If a technical issue is likely,
   recommend troubleshooting.

7. If an upgrade is justified,
   recommend the lowest suitable plan.

8. Consider the customer's actual usage.

9. Clearly explain the recommendation.

Return:

Current Plan:
Recommended Plan:
Speed:
Monthly Price:
Reason:
Next Action:
"""

    return generate_text(
        prompt
    )

24. Agent + RAG Architecture

The Plan Agent now works as:

Customer
   ↓
Plan Agent
   ↓
Build Retrieval Query
   ↓
Sentence Transformer
   ↓
Query Embedding
   ↓
ChromaDB
   ↓
Top-K Chunks
   ↓
Distance Filter
   ↓
Relevant Knowledge
   ↓
Gemini Prompt
   ↓
Plan Recommendation

This is significantly different from simply asking an LLM:

"Which plan should I recommend?"

The LLM now receives information retrieved from our own knowledge base.


25. Why Grounding Matters

Without RAG:

Customer
   ↓
Gemini
   ↓
Possible answer based on model knowledge

With RAG:

Customer
   ↓
GlobalNet Knowledge
   ↓
Relevant Plans
   ↓
Gemini
   ↓
Grounded recommendation

We also explicitly instruct Gemini:

Never invent a plan.
Never invent a price.
Never invent a speed.

RAG does not mathematically guarantee that hallucinations disappear, but retrieval plus explicit grounding rules can substantially improve control over the information available to the model.


26. Testing the Complete API

Start:

uvicorn app.main:app --reload

Open:

http://127.0.0.1:8000/docs

Test:

POST /support

Example:

{
  "customer_id": "1001",
  "chat_text": "My internet is very slow and I need much higher bandwidth because I have many devices and use 4K streaming.",
  "loyalty_status": "GOLD",
  "current_plan": "BASIC",
  "current_plan_desc": "100 Mbps",
  "tenure_months": 48,
  "customer_email": "your-email@example.com",
  "customer_name": "Raj",
  "call_type": "online"
}

The execution path is now:

POST /support
      ↓
Upsell Agent
      ↓
Gemini
      ↓
Plan Agent
      ↓
RAG
      ↓
ChromaDB
      ↓
Relevant Plans
      ↓
Gemini
      ↓
Plan Recommendation
      ↓
Offer Agent
      ↓
Gemini
      ↓
Final Response

27. Why RAG Evaluation is Necessary

Getting a response from RAG does not mean RAG is good.

We need to test:

Did RAG retrieve the correct information?

For example:

Question:
"I have 8 devices and multiple 4K streams."

Expected:
PREMIUM

If RAG returns only:

BASIC

our system has a retrieval problem.

This problem exists before Gemini even generates an answer.

Therefore, retrieval should be evaluated independently.


28. Create a RAG Evaluation Dataset

Create:

data/rag_evaluation.json

Example:

[
  {
    "question": "I only have two devices and mainly browse websites and use email.",
    "expected_plan": "BASIC"
  },
  {
    "question": "I have four devices and frequently work from home with video conferencing.",
    "expected_plan": "STANDARD"
  },
  {
    "question": "I have eight devices and use multiple 4K streams.",
    "expected_plan": "PREMIUM"
  },
  {
    "question": "I have more than ten devices and several simultaneous 4K streams.",
    "expected_plan": "ULTRA"
  },
  {
    "question": "I need internet for normal browsing, email and social media.",
    "expected_plan": "BASIC"
  },
  {
    "question": "My family has many devices and we regularly play online games and stream 4K video.",
    "expected_plan": "PREMIUM"
  }
]

This is not training data.

It is a test dataset.


29. RAG Evaluation Script

Create:

scripts/evaluate_rag.py

Code:

import json

from pathlib import Path

from app.services.rag import search_plans


PROJECT_ROOT = (
    Path(__file__).resolve().parents[1]
)


EVALUATION_FILE = (
    PROJECT_ROOT
    / "data"
    / "rag_evaluation.json"
)


def load_evaluation_data():

    with open(
        EVALUATION_FILE,
        "r",
        encoding="utf-8"
    ) as file:

        return json.load(file)


def extract_plan_ids(results):

    plan_ids = []

    for metadata in (
        results["metadatas"][0]
    ):

        plan_id = metadata["plan_id"]

        if plan_id not in plan_ids:

            plan_ids.append(
                plan_id
            )

    return plan_ids


def evaluate():

    test_cases = load_evaluation_data()

    total = len(
        test_cases
    )

    correct = 0

    print(
        "\n===== RAG EVALUATION =====\n"
    )

    for index, test_case in enumerate(
        test_cases,
        start=1
    ):

        question = (
            test_case["question"]
        )

        expected_plan = (
            test_case["expected_plan"]
        )

        results = search_plans(
            question,
            top_k=5,
            max_distance=0.80
        )

        retrieved_plans = extract_plan_ids(
            results
        )

        if expected_plan == "NONE":

            is_correct = (
                len(retrieved_plans) == 0
            )

        else:

            is_correct = (
                expected_plan
                in retrieved_plans
            )

        if is_correct:

            correct += 1

        print(
            f"Test {index}"
        )

        print(
            f"Question: {question}"
        )

        print(
            f"Expected: {expected_plan}"
        )

        print(
            f"Retrieved: {retrieved_plans}"
        )

        print(
            f"Result: {'PASS' if is_correct else 'FAIL'}"
        )

        print(
            "-" * 70
        )

    accuracy = (
        correct / total
    ) * 100

    print(
        "\n===== SUMMARY ====="
    )

    print(
        f"Total Tests: {total}"
    )

    print(
        f"Correct: {correct}"
    )

    print(
        f"Incorrect: {total - correct}"
    )

    print(
        f"Accuracy: {accuracy:.2f}%"
    )


if __name__ == "__main__":

    evaluate()

30. Run RAG Evaluation

Run:

python scripts/evaluate_rag.py

Example output:

===== RAG EVALUATION =====

Test 1

Question:
I only have two devices and mainly browse websites and use email.

Expected:
BASIC

Retrieved:
['BASIC', 'STANDARD']

Result:
PASS

At the end:

===== SUMMARY =====

Total Tests: 6

Correct: 5

Incorrect: 1

Accuracy: 83.33%

The exact result depends on retrieval behavior.

A failed test is not something to hide.

It tells us where retrieval needs improvement.


31. Top-K Evaluation

Suppose RAG returns:

1. PREMIUM
2. STANDARD
3. ULTRA
4. BASIC

and the expected answer is:

PREMIUM

For a Top-K retrieval test:

PASS

because PREMIUM was retrieved.

This answers:

Did the retriever find the relevant information somewhere in its result set?


32. Top-1 Evaluation

A stricter test checks only the first result.

For example:

Question:
8 devices + multiple 4K streams

Expected:
PREMIUM

Top Result:
PREMIUM

Result:

PASS

But:

Top Result:
STANDARD

would be:

FAIL

Top-1 accuracy tells us how often the retriever's highest-ranked result is correct.


33. Create Top-1 Evaluation

Create:

scripts/evaluate_rag_top1.py

Code:

import json

from pathlib import Path

from app.services.rag import search_plans


PROJECT_ROOT = (
    Path(__file__).resolve().parents[1]
)


EVALUATION_FILE = (
    PROJECT_ROOT
    / "data"
    / "rag_evaluation.json"
)


def load_data():

    with open(
        EVALUATION_FILE,
        "r",
        encoding="utf-8"
    ) as file:

        return json.load(file)


def evaluate():

    test_cases = load_data()

    correct = 0

    total = len(
        test_cases
    )

    print(
        "\n===== TOP-1 RAG EVALUATION =====\n"
    )

    for index, test_case in enumerate(
        test_cases,
        start=1
    ):

        question = (
            test_case["question"]
        )

        expected = (
            test_case["expected_plan"]
        )

        results = search_plans(
            question,
            top_k=1,
            max_distance=0.80
        )

        if results["metadatas"][0]:

            retrieved = (
                results["metadatas"][0][0]["plan_id"]
            )

            distance = (
                results["distances"][0][0]
            )

        else:

            retrieved = None
            distance = None

        if expected == "NONE":

            passed = (
                retrieved is None
            )

        else:

            passed = (
                retrieved == expected
            )

        if passed:

            correct += 1

        print(
            f"Test: {index}"
        )

        print(
            f"Expected: {expected}"
        )

        print(
            f"Retrieved: {retrieved}"
        )

        print(
            f"Distance: {distance}"
        )

        print(
            f"Result: {'PASS' if passed else 'FAIL'}"
        )

        print(
            "-" * 60
        )

    accuracy = (
        correct / total
    ) * 100

    print(
        f"\nTop-1 Accuracy: {accuracy:.2f}%"
    )


if __name__ == "__main__":

    evaluate()

Run:

python scripts/evaluate_rag_top1.py

34. Testing Abstention

An important test is an unrelated question.

Add:

{
  "question": "What is the capital city of India?",
  "expected_plan": "NONE"
}

Ideally:

Question
   ↓
RAG
   ↓
No sufficiently relevant documents
   ↓
No recommendation

This is called abstention.


35. Why Abstention Matters

A bad AI system may behave like:

Unknown Question
      ↓
Find nearest random document
      ↓
Give confident answer

A safer design is:

Unknown Question
      ↓
Search knowledge
      ↓
Insufficient evidence
      ↓
Do not make recommendation
      ↓
Clarify or escalate

This is especially important in enterprise AI systems.


36. Complete RAG Lifecycle

Our RAG system now has two major phases.

Phase 1 — Ingestion

GlobalNet Plan JSON
        ↓
Document Loader
        ↓
Chunking
        ↓
Metadata
        ↓
Embedding Model
        ↓
Embeddings
        ↓
ChromaDB

This happens when knowledge is created or updated.

Phase 2 — Retrieval

Customer Question
        ↓
Query Embedding
        ↓
ChromaDB Search
        ↓
Top-K Chunks
        ↓
Distance Filtering
        ↓
Relevant Context
        ↓
Gemini
        ↓
Recommendation

37. Where the LLM Fits in RAG

A common misunderstanding is that the vector database generates the answer.

It does not.

The responsibilities are different.

Sentence Transformer
        ↓
Creates embeddings

ChromaDB
        ↓
Finds relevant knowledge

Gemini
        ↓
Reasons over retrieved knowledge
and generates the response

This separation is fundamental to understanding RAG.


38. What Happens When Plan Data Changes?

Suppose:

PREMIUM
500 Mbps
₹1099

changes to:

PREMIUM
600 Mbps
₹1199

We update:

globalnet_plans.json

Then re-run:

python scripts/ingest_plans.py

Because we use ChromaDB upsert, records with matching IDs are updated.

The knowledge flow becomes:

Updated Business Data
        ↓
Re-ingestion
        ↓
New Embeddings
        ↓
Updated Vector Records
        ↓
Future Queries
        ↓
Updated Knowledge

This is far better than manually changing AI prompts throughout the application.


39. Important Production Considerations

This implementation is still a learning architecture.

For a production system, I would add:

Document versioning
Incremental ingestion
Delete handling
Metadata filtering
Hybrid search
Reranking
Structured LLM output
Authentication
Authorization
Secrets management
Observability
Tracing
Prompt versioning
Evaluation pipelines
Error handling
Retries
Caching
Audit logs

The evaluation dataset would also need to be much larger and based on realistic customer queries.


40. Future RAG Architecture

The next evolution can look like:

                    Customer
                       |
                       v
                FastAPI / API Gateway
                       |
                       v
                Agent Orchestrator
                       |
          +------------+------------+
          |                         |
          v                         v
    Upsell Agent                RAG Agent
                                    |
                                    v
                             Query Rewriter
                                    |
                                    v
                              Vector Search
                                    |
                                    v
                            Metadata Filter
                                    |
                                    v
                                Reranker
                                    |
                                    v
                           Relevant Context
                                    |
                                    v
                              Plan Agent
                                    |
                                    v
                                  Gemini
                                    |
                                    v
                              Offer Agent
                                    |
                                    v
                             Final Response

41. What I Learned From This RAG Implementation

This project covers several important GenAI engineering concepts.

Knowledge Management

External knowledge
JSON documents
Document loading
Knowledge updates

Embeddings

Text → vectors
Semantic similarity
Embedding models

Vector Databases

ChromaDB
Persistent collections
Vector storage
Similarity search

Chunking

Document decomposition
Chunk IDs
Chunk types
Metadata

Retrieval

Semantic search
Top-K retrieval
Distance scores
Relevance thresholds

Generation

Retrieved context
Prompt grounding
Gemini
Agent reasoning

Evaluation

Evaluation datasets
Top-K retrieval
Top-1 accuracy
Failure analysis
Abstention

42. RAG vs Fine-Tuning

For this use case, changing plan information does not require retraining Gemini.

For example:

Old price → ₹1099
New price → ₹1199

With RAG:

Update Knowledge
      ↓
Re-index
      ↓
Retrieve New Information

This makes RAG particularly useful for frequently changing business knowledge.

Fine-tuning and RAG solve different problems.

RAG is generally appropriate when the model needs access to changing or private knowledge.

Fine-tuning is more appropriate when we want to alter model behavior, style, task patterns, or specialized capabilities.

They can also be combined.


43. Final Architecture So Far

At this stage, the GlobalNet project looks like:

                         Customer
                            |
                            v
                         FastAPI
                            |
                            v
                    +---------------+
                    | Upsell Agent  |
                    +-------+-------+
                            |
                            v
                         Gemini
                            |
                            v
                    +---------------+
                    |  Plan Agent   |
                    +-------+-------+
                            |
                            v
                       RAG Layer
                            |
                            v
                Sentence Transformer
                            |
                            v
                         ChromaDB
                            |
                            v
                   Relevant Chunks
                            |
                            v
                   Relevance Filter
                            |
                            v
                         Gemini
                            |
                            v
                  Plan Recommendation
                            |
                            v
                    +---------------+
                    |  Offer Agent  |
                    +-------+-------+
                            |
                            v
                         Gemini
                            |
                            v
                  Personalized Offer

And separately:

                    Evaluation Dataset
                            |
                            v
                    RAG Evaluation
                            |
               +------------+------------+
               |            |            |
               v            v            v
             Top-K        Top-1      Abstention

44. From Simple AI App to AI Engineering

The progression of this project so far has been:

Python
   ↓
FastAPI
   ↓
Vertex AI
   ↓
Gemini
   ↓
Prompt Engineering
   ↓
Upsell Agent
   ↓
Plan Agent
   ↓
Offer Agent
   ↓
External Knowledge
   ↓
Document Processing
   ↓
Chunking
   ↓
Embeddings
   ↓
ChromaDB
   ↓
Vector Search
   ↓
RAG
   ↓
Relevance Filtering
   ↓
RAG Evaluation

This progression is important because building a useful GenAI system requires more than simply calling an LLM API.

The surrounding engineering determines how knowledge is retrieved, controlled, evaluated and delivered.


Conclusion

In this part of the GlobalNet Support Agent project, I transformed the Plan Agent from a hardcoded prompt-based implementation into a RAG-powered architecture.

We implemented:

GlobalNet Knowledge Base
        ↓
Document Loader
        ↓
Chunking
        ↓
Metadata
        ↓
Sentence Transformer
        ↓
Embeddings
        ↓
ChromaDB
        ↓
Semantic Search
        ↓
Top-K Retrieval
        ↓
Distance Filtering
        ↓
Plan Agent
        ↓
Gemini

We also added an evaluation layer:

Evaluation Questions
        ↓
Expected Plans
        ↓
RAG Retrieval
        ↓
Compare Results
        ↓
Top-K / Top-1 Accuracy

The next stage is to improve the retrieval layer further by returning structured RAG results, adding metadata filtering, improving retrieval quality, and eventually introducing techniques such as reranking and hybrid search.

The long-term objective is to evolve this project into a more complete AI support platform:

Multi-Agent AI
      +
RAG
      +
Customer Database
      +
Firestore
      +
Cloud Run
      +
Observability
      +
Automated Evaluation
      =
Production-Style AI Support Platform

No comments:

Post a Comment