Building a Production-Style RAG Pipeline for a Telecom AI Support Agent — Part 2
Introduction
In Part 1 of the GlobalNet Support Agent project, I built an AI-powered telecom customer support application using:
Python
FastAPI
Google Cloud
Vertex AI
Gemini
Multiple AI agents
The application already contained three agents:
Customer
↓
Upsell Agent
↓
Plan Agent
↓
Offer Agent
↓
Customer ResponseThe Upsell Agent determines whether an upgrade is justified.
The Plan Agent determines which plan is suitable.
The Offer Agent creates a personalized customer offer.
However, there was an important limitation.
The Plan Agent initially contained plan information directly inside its prompt:
BASIC → 100 Mbps
STANDARD → 300 Mbps
PREMIUM → 500 Mbps
ULTRA → 1 GbpsThis approach works for a prototype, but it does not scale.
In this part, I will convert the application into an Agent + RAG architecture using:
JSON Knowledge Base
↓
Document Processing
↓
Chunking
↓
Sentence Transformers
↓
Embeddings
↓
ChromaDB
↓
Semantic Search
↓
Relevance Filtering
↓
Gemini
↓
Plan RecommendationWe will also build a basic RAG evaluation framework to test whether the retriever is actually finding the correct telecom plans.
1. What is RAG?
RAG stands for:
Retrieval-Augmented Generation
A normal LLM application might work like this:
User Question
↓
Prompt
↓
LLM
↓
AnswerThe problem is that the LLM may not know our private business information.
For example, Gemini does not automatically know the current internal GlobalNet plan catalog used by this demo application.
We could put all plan information directly into the prompt, but this becomes difficult as the knowledge base grows.
RAG changes the architecture.
User Question
↓
Retriever
↓
Knowledge Base
↓
Relevant Information
↓
LLM
↓
Grounded AnswerInstead of expecting the LLM to know everything, we retrieve relevant information first and provide that information as context.
2. Why RAG for a Telecom Support Agent?
Imagine that a telecom company has:
100 internet plans
500 support documents
200 troubleshooting guides
50 promotion documents
1000 FAQ entries
Product documentation
Network policies
Upgrade rules
Eligibility rulesSending all of this information to an LLM for every customer request would be inefficient.
Instead:
Customer:
"I have 8 devices and multiple 4K TVs."
↓
Retriever
↓
Search knowledge base
↓
Relevant plan information
↓
Gemini
↓
Recommended planThis is where RAG becomes useful.
3. RAG Architecture for GlobalNet
The architecture developed in this project is:
Customer
|
v
FastAPI API
|
v
Upsell Agent
|
v
Gemini
|
v
Plan Agent
|
v
RAG Query
|
v
Sentence Transformer
|
v
Query Embedding
|
v
ChromaDB
|
v
Similarity Search
|
v
Relevant Chunks
|
v
Relevance Filtering
|
v
Context Prompt
|
v
Gemini
|
v
Plan Recommendation
|
v
Offer Agent
|
v
Gemini
|
v
Personalized Offer4. Updated Project Directory Structure
After adding RAG, the project structure becomes:
GlobalNet-Support-Agent-GCP/
│
├── app/
│ ├── __init__.py
│ ├── main.py
│ │
│ ├── agents/
│ │ ├── __init__.py
│ │ ├── upsell_agent.py
│ │ ├── plan_agent.py
│ │ └── offer_agent.py
│ │
│ ├── services/
│ │ ├── __init__.py
│ │ ├── vertex_ai.py
│ │ ├── plan_catalog.py
│ │ ├── rag.py
│ │ ├── firestore.py
│ │ └── email.py
│ │
│ └── models/
│ └── customer.py
│
├── data/
│ ├── plans/
│ │ └── globalnet_plans.json
│ │
│ └── rag_evaluation.json
│
├── scripts/
│ ├── ingest_plans.py
│ ├── evaluate_rag.py
│ └── evaluate_rag_top1.py
│
├── chroma_db/
│
├── test_vertex.py
├── test_rag.py
│
├── requirements.txt
├── Dockerfile
├── .dockerignore
└── README.md5. Create the GlobalNet Knowledge Base
The first step is moving plan information outside the Python source code.
Create:
data/plans/globalnet_plans.jsonExample:
[
{
"plan_id": "BASIC",
"plan_name": "GlobalNet Basic",
"speed": "100 Mbps",
"monthly_price": 499,
"data_limit": "Unlimited",
"recommended_for": [
"1-3 devices",
"normal browsing",
"email",
"social media",
"HD streaming"
],
"features": [
"Unlimited data",
"Standard customer support",
"Wi-Fi router"
]
},
{
"plan_id": "STANDARD",
"plan_name": "GlobalNet Standard",
"speed": "300 Mbps",
"monthly_price": 799,
"data_limit": "Unlimited",
"recommended_for": [
"3-6 devices",
"4K streaming",
"work from home",
"video conferencing",
"online gaming"
],
"features": [
"Unlimited data",
"Priority customer support",
"Wi-Fi 6 router"
]
},
{
"plan_id": "PREMIUM",
"plan_name": "GlobalNet Premium",
"speed": "500 Mbps",
"monthly_price": 1099,
"data_limit": "Unlimited",
"recommended_for": [
"6-10 devices",
"multiple 4K streams",
"heavy work from home",
"online gaming",
"large file transfers"
],
"features": [
"Unlimited data",
"Priority customer support",
"Wi-Fi 6 router",
"Premium support"
]
},
{
"plan_id": "ULTRA",
"plan_name": "GlobalNet Ultra",
"speed": "1 Gbps",
"monthly_price": 1499,
"data_limit": "Unlimited",
"recommended_for": [
"10+ devices",
"multiple simultaneous 4K streams",
"professional users",
"heavy cloud workloads",
"large household usage"
],
"features": [
"Unlimited data",
"24x7 priority support",
"Wi-Fi 6 router",
"Premium support",
"Dedicated service assistance"
]
}
]These are sample plans created for this learning project.
They are not intended to represent actual commercial telecom offerings.
6. Create a Plan Catalog Service
Create:
app/services/plan_catalog.pyCode:
import json
from pathlib import Path
PLAN_FILE = (
Path(__file__).resolve().parents[2]
/ "data"
/ "plans"
/ "globalnet_plans.json"
)
def load_plans():
with open(
PLAN_FILE,
"r",
encoding="utf-8"
) as file:
return json.load(file)
def get_plan(plan_id: str):
plans = load_plans()
for plan in plans:
if plan["plan_id"] == plan_id:
return plan
return NoneThis separates business data from application logic.
The application can now load plan information from:
globalnet_plans.jsoninstead of embedding the information inside Python prompts.
7. Moving from JSON Lookup to RAG
Simply reading JSON is not RAG.
This:
JSON
↓
Read everything
↓
Put everything into prompt
↓
Geminiis still a normal context-injection approach.
True RAG introduces retrieval:
Documents
↓
Chunking
↓
Embeddings
↓
Vector Database
+
Customer Question
↓
Embedding
↓
Vector Search
↓
Relevant Documents
↓
Gemini8. Install the RAG Components
For this project I used:
ChromaDB
Sentence TransformersInstall:
pip install chromadb sentence-transformersVerify:
python -c "import chromadb; import sentence_transformers; print('RAG packages OK')"9. What is an Embedding?
An embedding converts text into a numerical representation.
For example:
"I have many devices and multiple 4K TVs."is converted conceptually into:
[0.13, -0.44, 0.82, 0.07, ...]The actual embedding contains many numerical dimensions.
Another sentence with similar meaning should have a vector located relatively close in embedding space.
For example:
"I need high-speed internet for several devices and 4K streaming."has similar semantic meaning.
This enables semantic search.
10. Embedding Model
The project uses:
SentenceTransformer(
"all-MiniLM-L6-v2"
)This model runs locally and converts plan documents and customer questions into embeddings.
11. Why Use a Vector Database?
A traditional database can perform queries such as:
plan_id = "PREMIUM"But a customer does not necessarily know the plan ID.
The customer might say:
"I have 8 devices and my family watches multiple 4K streams."We need to find the plan whose meaning best matches this requirement.
A vector database stores embeddings and performs similarity searches.
For this project I used:
ChromaDB12. Why Chunking Matters
Initially I stored one entire plan as one document.
For example:
PREMIUM
|
+-- Price
+-- Speed
+-- Recommended Usage
+-- FeaturesThat works with a tiny dataset.
However, larger knowledge bases require more granular retrieval.
I therefore split each plan into three chunks.
PREMIUM
|
+---- PREMIUM_basic
|
+---- PREMIUM_usage
|
+---- PREMIUM_featuresThe same happens for all four plans.
BASIC_basic
BASIC_usage
BASIC_features
STANDARD_basic
STANDARD_usage
STANDARD_features
PREMIUM_basic
PREMIUM_usage
PREMIUM_features
ULTRA_basic
ULTRA_usage
ULTRA_featuresTherefore:
4 plans × 3 chunks = 12 chunksEach chunk receives its own embedding.
13. Complete RAG Service
Create:
app/services/rag.pyCode:
import chromadb
from sentence_transformers import SentenceTransformer
CHROMA_PATH = "./chroma_db"
COLLECTION_NAME = "globalnet_plans"
# ------------------------------------------------------------
# Embedding Model
# ------------------------------------------------------------
embedding_model = SentenceTransformer(
"all-MiniLM-L6-v2"
)
# ------------------------------------------------------------
# ChromaDB Client
# ------------------------------------------------------------
chroma_client = chromadb.PersistentClient(
path=CHROMA_PATH
)
# ------------------------------------------------------------
# ChromaDB Collection
# ------------------------------------------------------------
collection = chroma_client.get_or_create_collection(
name=COLLECTION_NAME
)
# ------------------------------------------------------------
# Create Plan Chunks
# ------------------------------------------------------------
def create_plan_chunks(plans):
documents = []
ids = []
metadatas = []
for plan in plans:
plan_id = plan["plan_id"]
# Basic information chunk
basic_info = f"""
Plan ID: {plan_id}
Plan Name: {plan["plan_name"]}
Speed: {plan["speed"]}
Monthly Price: ₹{plan["monthly_price"]}
Data Limit: {plan["data_limit"]}
"""
documents.append(
basic_info
)
ids.append(
f"{plan_id}_basic"
)
metadatas.append({
"plan_id": plan_id,
"plan_name": plan["plan_name"],
"chunk_type": "basic",
"speed": plan["speed"],
})
# Recommended usage chunk
usage_info = f"""
Plan ID: {plan_id}
Plan Name: {plan["plan_name"]}
Recommended For:
{", ".join(plan["recommended_for"])}
"""
documents.append(
usage_info
)
ids.append(
f"{plan_id}_usage"
)
metadatas.append({
"plan_id": plan_id,
"plan_name": plan["plan_name"],
"chunk_type": "usage",
"speed": plan["speed"],
})
# Features chunk
feature_info = f"""
Plan ID: {plan_id}
Plan Name: {plan["plan_name"]}
Features:
{", ".join(plan["features"])}
"""
documents.append(
feature_info
)
ids.append(
f"{plan_id}_features"
)
metadatas.append({
"plan_id": plan_id,
"plan_name": plan["plan_name"],
"chunk_type": "features",
"speed": plan["speed"],
})
return documents, ids, metadatas
# ------------------------------------------------------------
# Create Embeddings
# ------------------------------------------------------------
def create_embeddings(documents):
embeddings = embedding_model.encode(
documents
)
return embeddings.tolist()
# ------------------------------------------------------------
# Store Plans
# ------------------------------------------------------------
def store_plans(plans):
documents, ids, metadatas = create_plan_chunks(
plans
)
embeddings = create_embeddings(
documents
)
collection.upsert(
ids=ids,
documents=documents,
embeddings=embeddings,
metadatas=metadatas,
)
return len(documents)
# ------------------------------------------------------------
# Search Plans
# ------------------------------------------------------------
def search_plans(
query,
top_k=5,
max_distance=0.80
):
query_embedding = embedding_model.encode(
[query]
).tolist()
results = collection.query(
query_embeddings=query_embedding,
n_results=top_k,
include=[
"documents",
"metadatas",
"distances"
],
)
filtered_documents = []
filtered_metadatas = []
filtered_distances = []
documents = results.get(
"documents",
[[]]
)[0]
metadatas = results.get(
"metadatas",
[[]]
)[0]
distances = results.get(
"distances",
[[]]
)[0]
for document, metadata, distance in zip(
documents,
metadatas,
distances,
):
if distance <= max_distance:
filtered_documents.append(
document
)
filtered_metadatas.append(
metadata
)
filtered_distances.append(
distance
)
return {
"documents": [
filtered_documents
],
"metadatas": [
filtered_metadatas
],
"distances": [
filtered_distances
],
}14. What Happens During Ingestion?
The ingestion pipeline works like this:
globalnet_plans.json
↓
load_plans()
↓
4 Plans
↓
create_plan_chunks()
↓
12 Chunks
↓
SentenceTransformer
↓
12 Embeddings
↓
ChromaDBThis happens before customer queries are processed.
15. Create the Ingestion Script
Create:
scripts/ingest_plans.pyCode:
import sys
from pathlib import Path
PROJECT_ROOT = Path(__file__).resolve().parents[1]
sys.path.insert(
0,
str(PROJECT_ROOT)
)
from app.services.plan_catalog import load_plans
from app.services.rag import store_plans
def main():
print(
"Loading GlobalNet plans..."
)
plans = load_plans()
print(
f"Found {len(plans)} plans"
)
print(
"Creating embeddings and storing in ChromaDB..."
)
count = store_plans(
plans
)
print(
f"Successfully stored {count} chunks in ChromaDB"
)
if __name__ == "__main__":
main()Run:
python scripts/ingest_plans.pyExpected:
Loading GlobalNet plans...
Found 4 plans
Creating embeddings and storing in ChromaDB...
Successfully stored 12 chunks in ChromaDB16. What Gets Stored in ChromaDB?
Each chunk contains three important things:
ID
Document
Metadata
EmbeddingFor example:
ID:
PREMIUM_usageDocument:
Plan ID: PREMIUM
Plan Name: GlobalNet Premium
Recommended For:
6-10 devices,
multiple 4K streams,
heavy work from home,
online gaming,
large file transfersMetadata:
{
"plan_id": "PREMIUM",
"plan_name": "GlobalNet Premium",
"chunk_type": "usage",
"speed": "500 Mbps"
}And internally:
Embedding VectorThis makes the knowledge searchable semantically.
17. Query-Time RAG Flow
Suppose the customer says:
"I have eight devices and use multiple 4K streams."The system does not search for an exact string.
Instead:
Customer Question
↓
SentenceTransformer
↓
Query Embedding
↓
ChromaDB
↓
Vector Similarity Search
↓
Closest ChunksA likely result could include:
PREMIUM_usagebecause that document contains semantically related information:
6-10 devices
multiple 4K streams18. Testing RAG Search
Create:
test_rag.pyCode:
from app.services.rag import search_plans
query = """
I have many devices at home and use several
4K streams. I need high bandwidth.
"""
results = search_plans(
query,
top_k=5
)
print(
"\n===== RAG RESULTS =====\n"
)
documents = results["documents"][0]
metadatas = results["metadatas"][0]
distances = results["distances"][0]
for i in range(
len(documents)
):
print(
f"Result: {i + 1}"
)
print(
f"Plan: {metadatas[i]['plan_name']}"
)
print(
f"Chunk Type: {metadatas[i]['chunk_type']}"
)
print(
f"Distance: {distances[i]}"
)
print(
"\nDocument:"
)
print(
documents[i]
)
print(
"-" * 70
)Run:
python test_rag.py19. What is top_k?
Our query contains:
top_k=5This means:
Retrieve up to the 5 nearest chunks.Conceptually:
Customer Query
↓
Vector Search
↓
Result 1
Result 2
Result 3
Result 4
Result 5The application then decides which results are relevant enough to use.
20. Understanding Distance
ChromaDB returns distance information for retrieved vectors.
For our chosen setup, smaller distances generally represent greater similarity.
For example:
Document A
Distance = 0.25
Document B
Distance = 0.70Document A is generally a closer semantic match.
However, distance values should not be interpreted as universal percentages.
A value such as:
0.80does not mean:
80% relevantThe values depend on the embedding model, vector space and configured distance behavior.
21. Relevance Filtering
A vector database normally returns the nearest available documents even when none are especially useful.
Suppose the user asks:
"What is the weather today?"Our database contains only telecom plans.
Without filtering, ChromaDB may still return the closest telecom document.
We don't want Gemini to treat that as valid knowledge.
Therefore, the RAG service applies:
if distance <= max_distance:The current learning-project threshold is:
max_distance=0.80Results outside the threshold are rejected.
This creates:
Query
↓
Vector Search
↓
Candidate Documents
↓
Distance Filter
↓
Relevant Documents22. Why Relevance Thresholds Need Evaluation
The value:
0.80is not automatically correct for every RAG application.
The best threshold depends on:
Embedding model
Document type
Chunk size
Domain
Query style
Vector database configuration
Evaluation datasetTherefore, a production system should determine the threshold using evaluation rather than guessing.
23. Connect RAG to the Plan Agent
The Plan Agent previously received every plan.
Now it retrieves only relevant plan information.
File:
app/agents/plan_agent.pyCode:
from app.services.vertex_ai import generate_text
from app.services.rag import search_plans
def recommend_plan(
customer,
upsell_analysis: str
) -> str:
query = f"""
Customer currently has:
Plan:
{customer.current_plan}
Customer problem:
{customer.chat_text}
Find the GlobalNet plan information that is
most relevant to the customer's requirements.
"""
rag_results = search_plans(
query,
top_k=5,
max_distance=0.80
)
documents = (
rag_results["documents"][0]
)
distances = (
rag_results["distances"][0]
)
if not documents:
return """
Current Plan:
Unknown
Recommended Plan:
No recommendation
Reason:
No sufficiently relevant plan information was found
in the GlobalNet knowledge base.
Next Action:
Ask a support specialist to review the customer.
"""
relevant_context = ""
for document, distance in zip(
documents,
distances
):
relevant_context += f"""
Relevance Distance: {distance}
{document}
-----------------------------
"""
prompt = f"""
You are a telecom plan recommendation AI.
Customer Information:
Customer ID:
{customer.customer_id}
Customer Name:
{customer.customer_name}
Current Plan:
{customer.current_plan}
Current Plan Description:
{customer.current_plan_desc}
Loyalty Status:
{customer.loyalty_status}
Tenure:
{customer.tenure_months} months
Customer Message:
{customer.chat_text}
Upsell Analysis:
{upsell_analysis}
Relevant GlobalNet Knowledge:
{relevant_context}
Rules:
1. Use only information contained in the retrieved
GlobalNet knowledge.
2. Never invent a plan.
3. Never invent a price.
4. Never invent a speed.
5. Do not recommend an upgrade just to increase revenue.
6. If a technical issue is likely,
recommend troubleshooting.
7. If an upgrade is justified,
recommend the lowest suitable plan.
8. Consider the customer's actual usage.
9. Clearly explain the recommendation.
Return:
Current Plan:
Recommended Plan:
Speed:
Monthly Price:
Reason:
Next Action:
"""
return generate_text(
prompt
)24. Agent + RAG Architecture
The Plan Agent now works as:
Customer
↓
Plan Agent
↓
Build Retrieval Query
↓
Sentence Transformer
↓
Query Embedding
↓
ChromaDB
↓
Top-K Chunks
↓
Distance Filter
↓
Relevant Knowledge
↓
Gemini Prompt
↓
Plan RecommendationThis is significantly different from simply asking an LLM:
"Which plan should I recommend?"The LLM now receives information retrieved from our own knowledge base.
25. Why Grounding Matters
Without RAG:
Customer
↓
Gemini
↓
Possible answer based on model knowledgeWith RAG:
Customer
↓
GlobalNet Knowledge
↓
Relevant Plans
↓
Gemini
↓
Grounded recommendationWe also explicitly instruct Gemini:
Never invent a plan.
Never invent a price.
Never invent a speed.RAG does not mathematically guarantee that hallucinations disappear, but retrieval plus explicit grounding rules can substantially improve control over the information available to the model.
26. Testing the Complete API
Start:
uvicorn app.main:app --reloadOpen:
http://127.0.0.1:8000/docsTest:
POST /supportExample:
{
"customer_id": "1001",
"chat_text": "My internet is very slow and I need much higher bandwidth because I have many devices and use 4K streaming.",
"loyalty_status": "GOLD",
"current_plan": "BASIC",
"current_plan_desc": "100 Mbps",
"tenure_months": 48,
"customer_email": "your-email@example.com",
"customer_name": "Raj",
"call_type": "online"
}The execution path is now:
POST /support
↓
Upsell Agent
↓
Gemini
↓
Plan Agent
↓
RAG
↓
ChromaDB
↓
Relevant Plans
↓
Gemini
↓
Plan Recommendation
↓
Offer Agent
↓
Gemini
↓
Final Response27. Why RAG Evaluation is Necessary
Getting a response from RAG does not mean RAG is good.
We need to test:
Did RAG retrieve the correct information?For example:
Question:
"I have 8 devices and multiple 4K streams."
Expected:
PREMIUMIf RAG returns only:
BASICour system has a retrieval problem.
This problem exists before Gemini even generates an answer.
Therefore, retrieval should be evaluated independently.
28. Create a RAG Evaluation Dataset
Create:
data/rag_evaluation.jsonExample:
[
{
"question": "I only have two devices and mainly browse websites and use email.",
"expected_plan": "BASIC"
},
{
"question": "I have four devices and frequently work from home with video conferencing.",
"expected_plan": "STANDARD"
},
{
"question": "I have eight devices and use multiple 4K streams.",
"expected_plan": "PREMIUM"
},
{
"question": "I have more than ten devices and several simultaneous 4K streams.",
"expected_plan": "ULTRA"
},
{
"question": "I need internet for normal browsing, email and social media.",
"expected_plan": "BASIC"
},
{
"question": "My family has many devices and we regularly play online games and stream 4K video.",
"expected_plan": "PREMIUM"
}
]This is not training data.
It is a test dataset.
29. RAG Evaluation Script
Create:
scripts/evaluate_rag.pyCode:
import json
from pathlib import Path
from app.services.rag import search_plans
PROJECT_ROOT = (
Path(__file__).resolve().parents[1]
)
EVALUATION_FILE = (
PROJECT_ROOT
/ "data"
/ "rag_evaluation.json"
)
def load_evaluation_data():
with open(
EVALUATION_FILE,
"r",
encoding="utf-8"
) as file:
return json.load(file)
def extract_plan_ids(results):
plan_ids = []
for metadata in (
results["metadatas"][0]
):
plan_id = metadata["plan_id"]
if plan_id not in plan_ids:
plan_ids.append(
plan_id
)
return plan_ids
def evaluate():
test_cases = load_evaluation_data()
total = len(
test_cases
)
correct = 0
print(
"\n===== RAG EVALUATION =====\n"
)
for index, test_case in enumerate(
test_cases,
start=1
):
question = (
test_case["question"]
)
expected_plan = (
test_case["expected_plan"]
)
results = search_plans(
question,
top_k=5,
max_distance=0.80
)
retrieved_plans = extract_plan_ids(
results
)
if expected_plan == "NONE":
is_correct = (
len(retrieved_plans) == 0
)
else:
is_correct = (
expected_plan
in retrieved_plans
)
if is_correct:
correct += 1
print(
f"Test {index}"
)
print(
f"Question: {question}"
)
print(
f"Expected: {expected_plan}"
)
print(
f"Retrieved: {retrieved_plans}"
)
print(
f"Result: {'PASS' if is_correct else 'FAIL'}"
)
print(
"-" * 70
)
accuracy = (
correct / total
) * 100
print(
"\n===== SUMMARY ====="
)
print(
f"Total Tests: {total}"
)
print(
f"Correct: {correct}"
)
print(
f"Incorrect: {total - correct}"
)
print(
f"Accuracy: {accuracy:.2f}%"
)
if __name__ == "__main__":
evaluate()30. Run RAG Evaluation
Run:
python scripts/evaluate_rag.pyExample output:
===== RAG EVALUATION =====
Test 1
Question:
I only have two devices and mainly browse websites and use email.
Expected:
BASIC
Retrieved:
['BASIC', 'STANDARD']
Result:
PASSAt the end:
===== SUMMARY =====
Total Tests: 6
Correct: 5
Incorrect: 1
Accuracy: 83.33%The exact result depends on retrieval behavior.
A failed test is not something to hide.
It tells us where retrieval needs improvement.
31. Top-K Evaluation
Suppose RAG returns:
1. PREMIUM
2. STANDARD
3. ULTRA
4. BASICand the expected answer is:
PREMIUMFor a Top-K retrieval test:
PASSbecause PREMIUM was retrieved.
This answers:
Did the retriever find the relevant information somewhere in its result set?
32. Top-1 Evaluation
A stricter test checks only the first result.
For example:
Question:
8 devices + multiple 4K streams
Expected:
PREMIUM
Top Result:
PREMIUMResult:
PASSBut:
Top Result:
STANDARDwould be:
FAILTop-1 accuracy tells us how often the retriever's highest-ranked result is correct.
33. Create Top-1 Evaluation
Create:
scripts/evaluate_rag_top1.pyCode:
import json
from pathlib import Path
from app.services.rag import search_plans
PROJECT_ROOT = (
Path(__file__).resolve().parents[1]
)
EVALUATION_FILE = (
PROJECT_ROOT
/ "data"
/ "rag_evaluation.json"
)
def load_data():
with open(
EVALUATION_FILE,
"r",
encoding="utf-8"
) as file:
return json.load(file)
def evaluate():
test_cases = load_data()
correct = 0
total = len(
test_cases
)
print(
"\n===== TOP-1 RAG EVALUATION =====\n"
)
for index, test_case in enumerate(
test_cases,
start=1
):
question = (
test_case["question"]
)
expected = (
test_case["expected_plan"]
)
results = search_plans(
question,
top_k=1,
max_distance=0.80
)
if results["metadatas"][0]:
retrieved = (
results["metadatas"][0][0]["plan_id"]
)
distance = (
results["distances"][0][0]
)
else:
retrieved = None
distance = None
if expected == "NONE":
passed = (
retrieved is None
)
else:
passed = (
retrieved == expected
)
if passed:
correct += 1
print(
f"Test: {index}"
)
print(
f"Expected: {expected}"
)
print(
f"Retrieved: {retrieved}"
)
print(
f"Distance: {distance}"
)
print(
f"Result: {'PASS' if passed else 'FAIL'}"
)
print(
"-" * 60
)
accuracy = (
correct / total
) * 100
print(
f"\nTop-1 Accuracy: {accuracy:.2f}%"
)
if __name__ == "__main__":
evaluate()Run:
python scripts/evaluate_rag_top1.py34. Testing Abstention
An important test is an unrelated question.
Add:
{
"question": "What is the capital city of India?",
"expected_plan": "NONE"
}Ideally:
Question
↓
RAG
↓
No sufficiently relevant documents
↓
No recommendationThis is called abstention.
35. Why Abstention Matters
A bad AI system may behave like:
Unknown Question
↓
Find nearest random document
↓
Give confident answerA safer design is:
Unknown Question
↓
Search knowledge
↓
Insufficient evidence
↓
Do not make recommendation
↓
Clarify or escalateThis is especially important in enterprise AI systems.
36. Complete RAG Lifecycle
Our RAG system now has two major phases.
Phase 1 — Ingestion
GlobalNet Plan JSON
↓
Document Loader
↓
Chunking
↓
Metadata
↓
Embedding Model
↓
Embeddings
↓
ChromaDBThis happens when knowledge is created or updated.
Phase 2 — Retrieval
Customer Question
↓
Query Embedding
↓
ChromaDB Search
↓
Top-K Chunks
↓
Distance Filtering
↓
Relevant Context
↓
Gemini
↓
Recommendation37. Where the LLM Fits in RAG
A common misunderstanding is that the vector database generates the answer.
It does not.
The responsibilities are different.
Sentence Transformer
↓
Creates embeddings
ChromaDB
↓
Finds relevant knowledge
Gemini
↓
Reasons over retrieved knowledge
and generates the responseThis separation is fundamental to understanding RAG.
38. What Happens When Plan Data Changes?
Suppose:
PREMIUM
500 Mbps
₹1099changes to:
PREMIUM
600 Mbps
₹1199We update:
globalnet_plans.jsonThen re-run:
python scripts/ingest_plans.pyBecause we use ChromaDB upsert, records with matching IDs are updated.
The knowledge flow becomes:
Updated Business Data
↓
Re-ingestion
↓
New Embeddings
↓
Updated Vector Records
↓
Future Queries
↓
Updated KnowledgeThis is far better than manually changing AI prompts throughout the application.
39. Important Production Considerations
This implementation is still a learning architecture.
For a production system, I would add:
Document versioning
Incremental ingestion
Delete handling
Metadata filtering
Hybrid search
Reranking
Structured LLM output
Authentication
Authorization
Secrets management
Observability
Tracing
Prompt versioning
Evaluation pipelines
Error handling
Retries
Caching
Audit logsThe evaluation dataset would also need to be much larger and based on realistic customer queries.
40. Future RAG Architecture
The next evolution can look like:
Customer
|
v
FastAPI / API Gateway
|
v
Agent Orchestrator
|
+------------+------------+
| |
v v
Upsell Agent RAG Agent
|
v
Query Rewriter
|
v
Vector Search
|
v
Metadata Filter
|
v
Reranker
|
v
Relevant Context
|
v
Plan Agent
|
v
Gemini
|
v
Offer Agent
|
v
Final Response41. What I Learned From This RAG Implementation
This project covers several important GenAI engineering concepts.
Knowledge Management
External knowledge
JSON documents
Document loading
Knowledge updatesEmbeddings
Text → vectors
Semantic similarity
Embedding modelsVector Databases
ChromaDB
Persistent collections
Vector storage
Similarity searchChunking
Document decomposition
Chunk IDs
Chunk types
MetadataRetrieval
Semantic search
Top-K retrieval
Distance scores
Relevance thresholdsGeneration
Retrieved context
Prompt grounding
Gemini
Agent reasoningEvaluation
Evaluation datasets
Top-K retrieval
Top-1 accuracy
Failure analysis
Abstention42. RAG vs Fine-Tuning
For this use case, changing plan information does not require retraining Gemini.
For example:
Old price → ₹1099
New price → ₹1199With RAG:
Update Knowledge
↓
Re-index
↓
Retrieve New InformationThis makes RAG particularly useful for frequently changing business knowledge.
Fine-tuning and RAG solve different problems.
RAG is generally appropriate when the model needs access to changing or private knowledge.
Fine-tuning is more appropriate when we want to alter model behavior, style, task patterns, or specialized capabilities.
They can also be combined.
43. Final Architecture So Far
At this stage, the GlobalNet project looks like:
Customer
|
v
FastAPI
|
v
+---------------+
| Upsell Agent |
+-------+-------+
|
v
Gemini
|
v
+---------------+
| Plan Agent |
+-------+-------+
|
v
RAG Layer
|
v
Sentence Transformer
|
v
ChromaDB
|
v
Relevant Chunks
|
v
Relevance Filter
|
v
Gemini
|
v
Plan Recommendation
|
v
+---------------+
| Offer Agent |
+-------+-------+
|
v
Gemini
|
v
Personalized OfferAnd separately:
Evaluation Dataset
|
v
RAG Evaluation
|
+------------+------------+
| | |
v v v
Top-K Top-1 Abstention44. From Simple AI App to AI Engineering
The progression of this project so far has been:
Python
↓
FastAPI
↓
Vertex AI
↓
Gemini
↓
Prompt Engineering
↓
Upsell Agent
↓
Plan Agent
↓
Offer Agent
↓
External Knowledge
↓
Document Processing
↓
Chunking
↓
Embeddings
↓
ChromaDB
↓
Vector Search
↓
RAG
↓
Relevance Filtering
↓
RAG EvaluationThis progression is important because building a useful GenAI system requires more than simply calling an LLM API.
The surrounding engineering determines how knowledge is retrieved, controlled, evaluated and delivered.
Conclusion
In this part of the GlobalNet Support Agent project, I transformed the Plan Agent from a hardcoded prompt-based implementation into a RAG-powered architecture.
We implemented:
GlobalNet Knowledge Base
↓
Document Loader
↓
Chunking
↓
Metadata
↓
Sentence Transformer
↓
Embeddings
↓
ChromaDB
↓
Semantic Search
↓
Top-K Retrieval
↓
Distance Filtering
↓
Plan Agent
↓
GeminiWe also added an evaluation layer:
Evaluation Questions
↓
Expected Plans
↓
RAG Retrieval
↓
Compare Results
↓
Top-K / Top-1 AccuracyThe next stage is to improve the retrieval layer further by returning structured RAG results, adding metadata filtering, improving retrieval quality, and eventually introducing techniques such as reranking and hybrid search.
The long-term objective is to evolve this project into a more complete AI support platform:
Multi-Agent AI
+
RAG
+
Customer Database
+
Firestore
+
Cloud Run
+
Observability
+
Automated Evaluation
=
Production-Style AI Support Platform