Tuesday, 6 October 2026

Building a Telecom AI Support Agent with Python, FastAPI, Gemini and GCP

In this tutorial, I am building a practical AI-powered Telecom Customer Support Agent using Python, FastAPI, Google Cloud, Vertex AI and Gemini.

The goal is to build an AI system that can understand a telecom customer's problem and then:

  1. Analyze whether there is a legitimate upsell opportunity.

  2. Recommend the most appropriate telecom plan.

  3. Generate a personalized customer offer.

The project is called:

GlobalNet Support Agent




This is not just a chatbot. The objective is to build a modular AI-agent architecture that can later be extended with:

  • RAG

  • Firestore

  • Customer data

  • Plan knowledge base

  • Email notifications

  • Cloud Run

  • Cloud Scheduler

  • Monitoring

  • Multiple AI agents

  • Kubernetes integration

This article covers the implementation completed so far.


1. Project Architecture

The current architecture is:

                         Customer
                            |
                            v
                    FastAPI /support
                            |
                            v
                    +---------------+
                    |  Upsell Agent |
                    +---------------+
                            |
                            v
                       Gemini AI
                            |
                            v
                    +---------------+
                    |   Plan Agent  |
                    +---------------+
                            |
                            v
                       Gemini AI
                            |
                            v
                    +---------------+
                    |  Offer Agent  |
                    +---------------+
                            |
                            v
                    Personalized Offer

The current technology stack is:

Python
FastAPI
Pydantic
Google Gen AI SDK
Google Cloud Vertex AI
Gemini 2.5 Flash
PowerShell
VS Code

2. Final Directory Structure So Far

The project currently has the following structure:

GlobalNet-Support-Agent-GCP/
│
├── app/
│   ├── __init__.py
│   │
│   ├── main.py
│   │
│   ├── agents/
│   │   ├── __init__.py
│   │   ├── upsell_agent.py
│   │   ├── plan_agent.py
│   │   └── offer_agent.py
│   │
│   ├── services/
│   │   ├── __init__.py
│   │   ├── vertex_ai.py
│   │   ├── rag.py
│   │   ├── firestore.py
│   │   └── email.py
│   │
│   └── models/
│       └── customer.py
│
├── jobs/
│   └── scan_inactive_customers.py
│
├── data/
│   └── plans/
│
├── test_vertex.py
│
├── requirements.txt
├── Dockerfile
├── .dockerignore
└── README.md

Some directories such as rag.py, firestore.py, email.py, and scheduled jobs are part of the planned architecture and will be implemented in later parts.


3. Create the GCP Project

The Google Cloud project used for this application is:

globalnet-support-agent

Set the project using the Google Cloud CLI:

gcloud config set project globalnet-support-agent

Expected output:

Updated property [core/project].

4. Enable Required Google Cloud APIs

The following APIs are required for the project:

gcloud services enable `
run.googleapis.com `
cloudbuild.googleapis.com `
artifactregistry.googleapis.com `
aiplatform.googleapis.com `
firestore.googleapis.com `
storage.googleapis.com `
cloudscheduler.googleapis.com `
secretmanager.googleapis.com `
logging.googleapis.com

These services will eventually provide:

Vertex AI
Cloud Run
Cloud Build
Artifact Registry
Firestore
Cloud Storage
Cloud Scheduler
Secret Manager
Cloud Logging

5. Configure Google Cloud Authentication

For local development, Application Default Credentials can be configured using:

gcloud auth application-default login

The project can also be configured as the quota project:

gcloud auth application-default set-quota-project globalnet-support-agent

6. Configure Environment Variables

The application uses the following environment variables.

In PowerShell:

$env:GOOGLE_CLOUD_PROJECT="globalnet-support-agent"
$env:GOOGLE_CLOUD_LOCATION="asia-south1"

These values are used by the Vertex AI service.


7. Vertex AI Service

Instead of putting Gemini code directly inside every agent, I created a reusable service.

File:

app/services/vertex_ai.py

Code:

from google import genai
from google.genai.types import GenerateContentConfig
import os


PROJECT_ID = os.environ["GOOGLE_CLOUD_PROJECT"]
LOCATION = os.environ.get(
    "GOOGLE_CLOUD_LOCATION",
    "asia-south1"
)


client = genai.Client(
    vertexai=True,
    project=PROJECT_ID,
    location=LOCATION,
)


def generate_text(prompt: str) -> str:

    response = client.models.generate_content(
        model="gemini-2.5-flash",
        contents=prompt,
        config=GenerateContentConfig(
            temperature=0.2,
        ),
    )

    return response.text

This creates one reusable function:

generate_text()

All AI agents can use this function.


8. Test Gemini Independently

Before connecting Gemini to FastAPI, I created:

test_vertex.py

Code:

from app.services.vertex_ai import generate_text


prompt = """
You are a telecom customer support AI.

A customer says:

"My internet is very slow and I need more bandwidth."

Explain whether there could be an upsell opportunity.
"""


print("Calling Gemini...")


result = generate_text(prompt)


print("\n===== GEMINI RESPONSE =====\n")
print(result)

Run:

python test_vertex.py

If everything is configured correctly, Gemini returns a response.

This is an important development practice:

First verify the AI service independently, then integrate it with the application.


9. FastAPI Application

The main API is:

app/main.py

FastAPI is used to expose the AI agent through an HTTP API.

The application starts with:

from fastapi import FastAPI
from pydantic import BaseModel

The FastAPI application is created using:

app = FastAPI(
    title="GlobalNet Support Agent",
    description="AI-powered telecom customer support agent",
    version="1.0.0",
)

10. Customer Request Model

The /support API receives customer information.

class SupportRequest(BaseModel):

    customer_id: str

    chat_text: str

    loyalty_status: str

    current_plan: str

    current_plan_desc: str

    tenure_months: int

    customer_email: str

    customer_name: str

    call_type: str

Pydantic automatically validates this request.


11. Health Check API

The application contains a basic health endpoint:

@app.get("/health")
def health():

    return {
        "status": "healthy"
    }

Testing:

GET /health

Expected response:

{
  "status": "healthy"
}

This endpoint will become useful later for Cloud Run and production monitoring.


12. Root Endpoint

The root endpoint is:

@app.get("/")
def root():

    return {
        "status": "ok",
        "service": "GlobalNet Support Agent",
        "message": "API is running",
    }

13. Upsell Agent

The first AI agent is the:

Upsell Agent

File:

app/agents/upsell_agent.py

Its responsibility is to determine whether the customer's problem represents a genuine upgrade opportunity.

Complete code:

from app.services.vertex_ai import generate_text


def analyze_upsell(customer) -> str:

    prompt = f"""
You are a telecom customer support upsell AI.

Analyze the customer information below.

Customer Information:
- Customer ID: {customer.customer_id}
- Customer Name: {customer.customer_name}
- Loyalty Status: {customer.loyalty_status}
- Current Plan: {customer.current_plan}
- Current Plan Description: {customer.current_plan_desc}
- Tenure: {customer.tenure_months} months
- Call Type: {customer.call_type}

Customer Message:
{customer.chat_text}

Determine:

1. What is the customer's main problem?
2. Is there a legitimate upsell opportunity?
3. Why is there or isn't there an opportunity?
4. What type of plan would be appropriate?
5. What should the support agent do next?

Important:

- Do not recommend an upgrade just to increase revenue.
- First determine whether the customer's problem could be
  caused by a technical issue.
- If a technical issue is likely, recommend troubleshooting first.
- Only recommend an upsell when the customer's actual usage
  justifies a higher plan.

Provide a clear and concise answer.
"""

    return generate_text(prompt)

14. Why the Upsell Agent Checks Technical Problems

This is important in a real telecom environment.

A customer saying:

"My internet is very slow."

does not automatically mean:

Customer needs a bigger plan.

The problem could be:

Router problem
       |
Wi-Fi interference
       |
Packet loss
       |
DNS problem
       |
ISP network problem
       |
Congestion
       |
Signal issue
       |
Only then → insufficient bandwidth

Therefore, the AI should not automatically try to sell something.

The prompt explicitly tells the AI:

Do not recommend an upgrade just to increase revenue.

This makes the agent more customer-centric.


15. Plan Agent

The second agent is:

Plan Agent

File:

app/agents/plan_agent.py

Its responsibility is to select the most appropriate plan.

Complete code:

from app.services.vertex_ai import generate_text


def recommend_plan(customer, upsell_analysis: str) -> str:

    prompt = f"""
You are a telecom plan recommendation AI.

Analyze the customer information and upsell analysis below.

Customer Information:
- Customer ID: {customer.customer_id}
- Customer Name: {customer.customer_name}
- Current Plan: {customer.current_plan}
- Current Plan Description: {customer.current_plan_desc}
- Loyalty Status: {customer.loyalty_status}
- Tenure: {customer.tenure_months} months
- Call Type: {customer.call_type}

Customer Message:
{customer.chat_text}

Upsell Analysis:
{upsell_analysis}

Available Example Plans:

1. BASIC
   Speed: 100 Mbps

2. STANDARD
   Speed: 300 Mbps

3. PREMIUM
   Speed: 500 Mbps

4. ULTRA
   Speed: 1 Gbps

Rules:

1. Do not recommend an upgrade just to increase revenue.
2. If the problem appears to be a technical issue, recommend
   troubleshooting before upgrading.
3. If the customer genuinely needs more bandwidth, recommend
   the lowest plan that can satisfy the requirement.
4. Consider the customer's current plan.
5. Consider the customer's stated usage.
6. Explain why the recommended plan is appropriate.
7. Mention the current plan.
8. Mention the recommended plan.
9. Keep the response concise and suitable for a support agent.

Return:

Current Plan:
Recommended Plan:
Reason:
Next Action:
"""

    return generate_text(prompt)

16. Example Plan Catalog

Currently the plans are hardcoded for learning purposes:

BASIC
100 Mbps

STANDARD
300 Mbps

PREMIUM
500 Mbps

ULTRA
1 Gbps

Later, these values will not be hardcoded.

They will come from a real knowledge source:

Plan Agent
     |
     v
RAG
     |
     v
Plan Knowledge Base
     |
     v
GlobalNet Plans

This will be implemented in a future part.


17. Offer Agent

The third agent is:

Offer Agent

File:

app/agents/offer_agent.py

Its responsibility is to create a personalized offer.

Complete code:

from app.services.vertex_ai import generate_text


def generate_offer(customer, plan_recommendation: str) -> str:

    prompt = f"""
You are a telecom customer retention and offer AI.

Analyze the customer information below and create an appropriate
personalized telecom offer.

Customer Information:
- Customer ID: {customer.customer_id}
- Customer Name: {customer.customer_name}
- Current Plan: {customer.current_plan}
- Current Plan Description: {customer.current_plan_desc}
- Loyalty Status: {customer.loyalty_status}
- Tenure: {customer.tenure_months} months
- Call Type: {customer.call_type}

Customer Message:
{customer.chat_text}

Plan Recommendation:
{plan_recommendation}

Offer Rules:

1. Do not create an offer if there is no legitimate upgrade
   opportunity.
2. If a technical issue is likely, recommend troubleshooting
   instead of an upgrade offer.
3. GOLD or higher loyalty customers may receive a stronger
   retention benefit.
4. Long-tenure customers may receive a loyalty benefit.
5. Do not invent unrealistic discounts.
6. The offer should be simple and easy for a support agent
   to explain to the customer.
7. Explain why the customer is receiving the offer.
8. The offer should match the recommended plan.

Return the result in this format:

Offer Decision:
Recommended Plan:
Monthly Benefit:
Loyalty Benefit:
Reason:
Agent Action:
Customer Message:
"""

    return generate_text(prompt)

18. Complete main.py

The current complete main.py is:

from fastapi import FastAPI
from pydantic import BaseModel

from app.agents.upsell_agent import analyze_upsell
from app.agents.plan_agent import recommend_plan
from app.agents.offer_agent import generate_offer


app = FastAPI(
    title="GlobalNet Support Agent",
    description="AI-powered telecom customer support agent",
    version="1.0.0",
)


class SupportRequest(BaseModel):

    customer_id: str
    chat_text: str
    loyalty_status: str
    current_plan: str
    current_plan_desc: str
    tenure_months: int
    customer_email: str
    customer_name: str
    call_type: str


@app.get("/")
def root():

    return {
        "status": "ok",
        "service": "GlobalNet Support Agent",
        "message": "API is running",
    }


@app.get("/health")
def health():

    return {
        "status": "healthy"
    }


@app.post("/support")
def support(request: SupportRequest):

    # Step 1: Analyze upsell opportunity
    upsell_result = analyze_upsell(request)

    # Step 2: Recommend appropriate plan
    plan_result = recommend_plan(
        request,
        upsell_result
    )

    # Step 3: Generate personalized offer
    offer_result = generate_offer(
        request,
        plan_result
    )

    return {
        "status": "success",
        "customer_id": request.customer_id,
        "customer_name": request.customer_name,
        "upsell_analysis": upsell_result,
        "plan_recommendation": plan_result,
        "personalized_offer": offer_result,
    }

19. Start the Application

From the project root:

cd "C:\Users\Raj Kumar Gupta\Desktop\Raj\minikube\GlobalNet-Support-Agent-GCP"

Set the environment variables:

$env:GOOGLE_CLOUD_PROJECT="globalnet-support-agent"
$env:GOOGLE_CLOUD_LOCATION="asia-south1"

Start FastAPI:

uvicorn app.main:app --reload

Expected:

Uvicorn running on http://127.0.0.1:8000

20. Open Swagger

FastAPI automatically provides Swagger UI.

Open:

http://127.0.0.1:8000/docs

You should see:

GET  /
GET  /health
POST /support

21. Test /support

Select:

POST /support

Click:

Try it out

Use this request:

{
  "customer_id": "1001",
  "chat_text": "My internet is very slow and I need much higher bandwidth because I have many devices and use 4K streaming.",
  "loyalty_status": "GOLD",
  "current_plan": "BASIC",
  "current_plan_desc": "100 Mbps",
  "tenure_months": 48,
  "customer_email": "your-email@example.com",
  "customer_name": "Raj",
  "call_type": "online"
}

Click:

Execute

22. What Happens Internally?

The request enters:

POST /support

Then:

SupportRequest
       |
       v
analyze_upsell()
       |
       v
Gemini
       |
       v
upsell_result

Then:

upsell_result
       |
       v
recommend_plan()
       |
       v
Gemini
       |
       v
plan_result

Then:

plan_result
       |
       v
generate_offer()
       |
       v
Gemini
       |
       v
offer_result

Finally:

JSON Response

23. Final Response Structure

The API returns a structure similar to:

{
  "status": "success",
  "customer_id": "1001",
  "customer_name": "Raj",
  "upsell_analysis": "...",
  "plan_recommendation": "...",
  "personalized_offer": "..."
}

The exact AI text will vary because Gemini generates the response dynamically.


24. Complete Request Flow

The complete flow is now:

                         +----------------+
                         |    Customer    |
                         +-------+--------+
                                 |
                                 v
                         +---------------+
                         |    FastAPI    |
                         |   /support    |
                         +-------+-------+
                                 |
                                 v
                     +-----------------------+
                     |    Upsell Agent       |
                     |                       |
                     | Is upgrade justified? |
                     +----------+------------+
                                |
                                v
                           +---------+
                           | Gemini  |
                           +----+----+
                                |
                                v
                     +-----------------------+
                     |      Plan Agent       |
                     |                       |
                     | Which plan is best?   |
                     +----------+------------+
                                |
                                v
                           +---------+
                           | Gemini  |
                           +----+----+
                                |
                                v
                     +-----------------------+
                     |     Offer Agent       |
                     |                       |
                     | What offer should     |
                     | we provide?           |
                     +----------+------------+
                                |
                                v
                           +---------+
                           | Gemini  |
                           +----+----+
                                |
                                v
                     +-----------------------+
                     |   Final API Response  |
                     +-----------------------+

25. Why Use Multiple Agents?

Instead of creating one large prompt, the application separates responsibilities.

Upsell Agent

Answers:

Should we upgrade the customer?

Plan Agent

Answers:

Which plan is appropriate?

Offer Agent

Answers:

What personalized offer should we provide?

This separation makes the application easier to:

  • Maintain

  • Test

  • Debug

  • Extend

  • Replace individual agents

  • Add business rules

  • Add RAG

  • Add databases


26. Current Limitations

The current implementation is a learning prototype.

There are several things that need to be improved before production.

Hardcoded plans

Currently:

BASIC → 100 Mbps
STANDARD → 300 Mbps
PREMIUM → 500 Mbps
ULTRA → 1 Gbps

These should come from a database or knowledge base.

Free-form AI responses

The agents currently return strings.

A production implementation should return structured JSON.

For example:

{
  "upgrade_required": true,
  "recommended_plan": "PREMIUM",
  "reason": "Customer has multiple devices and 4K streaming requirements",
  "next_action": "Offer plan upgrade"
}

No RAG yet

The Plan Agent does not yet retrieve information from a real GlobalNet knowledge base.

No Firestore yet

Customer information is currently supplied directly through the API request.

No authentication

The API currently has no authentication or authorization.

No production monitoring

Cloud Logging, metrics and tracing still need to be integrated.


27. What We Will Build Next

The next stage will make the project significantly more realistic.

The planned architecture is:

Customer
   |
   v
FastAPI
   |
   v
Agent Orchestrator
   |
   +------------------+
   |                  |
   v                  v
Upsell Agent       RAG System
   |                  |
   |                  v
   |            Plan Knowledge
   |                  |
   +--------+---------+
            |
            v
       Plan Agent
            |
            v
       Offer Agent
            |
            v
         Firestore
            |
            v
        Email Agent
            |
            v
      Customer Notification

Future components will include:

RAG
ChromaDB / Vector Database
Embeddings
Firestore
Customer Database
Plan Knowledge Base
Email Service
Cloud Run
Artifact Registry
Cloud Scheduler
Cloud Logging

28. Learning Lessons From This Project

This project demonstrates several important Python and AI engineering concepts.

Python

We are learning:

Functions
Imports
Modules
Packages
Classes
Type validation
Virtual environments
Environment variables
Exception handling

FastAPI

We are learning:

REST API
Endpoints
POST requests
Pydantic models
Swagger
Health checks
API integration

Generative AI

We are learning:

LLM integration
Prompt engineering
Agent design
Multi-agent architecture
Context passing
AI decision making

Cloud

We are learning:

Google Cloud
Vertex AI
Cloud Run
Artifact Registry
Firestore
Cloud Scheduler
Secret Manager
Cloud Logging

29. Conclusion

We have now created the first working version of the GlobalNet Support Agent.

The application can:

Receive customer request
        ↓
Analyze upsell opportunity
        ↓
Recommend a suitable plan
        ↓
Generate a personalized offer
        ↓
Return the result through REST API

The most important design principle is that the AI should not simply try to sell a more expensive plan.

The system should first understand the customer's actual problem.

For example:

Customer:
"My internet is slow."

          ↓

AI investigates

          ↓

Could be technical issue?
          |
       YES ─────→ Troubleshoot first
          |
         NO
          |
          v
Does usage justify more bandwidth?
          |
       YES
          |
          v
Recommend appropriate plan
          |
          v
Generate personalized offer


Wednesday, 2 September 2026

Container OOM vs Node OOM in Kubernetes

Memory-related incidents are among the most common and confusing problems in Kubernetes production environments.

A Pod restarts unexpectedly, an application becomes unavailable, or a Kubernetes node suddenly becomes unstable. One of the first things an SRE may notice is:

OOMKilled

But does OOMKilled always mean that the Kubernetes node ran out of memory?

No.

Understanding the difference between Container OOM, Node Memory Pressure, Pod Eviction, and Node-level OOM is extremely important for Kubernetes SREs.


1. Container OOM vs Node OOM

The simplest way to understand the difference is:

Container OOM is a container-level memory problem. Node OOM is a node-level memory problem.

AreaContainer OOMNode OOM
ScopeOne container/processEntire Kubernetes node
Main causeContainer exceeds its memory limitNode runs critically low on memory
Typical evidenceOOMKilledMemoryPressure, kernel OOM messages
Who gets killed?Usually the offending container/processLinux kernel may kill processes/containers
Pod behaviorContainer may restartPods may be evicted, killed, or become unstable
Other workloads affectedUsually limitedPotentially many workloads
Memory limit involved?UsuallyNot necessarily
SeverityApplication/container-levelNode/cluster-level

The distinction becomes clearer with a production example.


2. What is Container OOM?

Consider a Kubernetes node with 64 GB of memory.

A Pod is running an application with the following configuration:

resources:
  requests:
    memory: "1Gi"
  limits:
    memory: "2Gi"

The container's application gradually consumes more memory:

500 MB
   ↓
1 GB
   ↓
1.5 GB
   ↓
2 GB
   ↓
2.2 GB

The container has a memory limit of 2 GiB.

If the container attempts to consume memory beyond its allowed limit, the process can be killed due to an out-of-memory condition.

Kubernetes may then report:

Reason: OOMKilled

You can investigate the Pod with:

kubectl describe pod myapp-xxx -n production

You may see:

Last State:
  Terminated:
    Reason: OOMKilled
    Exit Code: 137

This is an important clue that the container experienced a memory-related termination.


3. What does Exit Code 137 mean?

You will frequently encounter:

Exit Code: 137

The number comes from:

128 + 9 = 137

Signal 9 is:

SIGKILL

Therefore, exit code 137 generally indicates that the process was killed with SIGKILL.

However, an important SRE rule is:

Do not assume that every exit code 137 automatically means OOM.

Always verify the termination reason and supporting system evidence.

For example:

kubectl describe pod <pod> -n <namespace>

Look for:

Reason: OOMKilled

Then check the node and kernel logs if necessary.


4. What is Node OOM?

Now consider a different situation.

Suppose a Kubernetes node has 64 GB of memory:

Kubernetes Node
Memory = 64 GB

Several workloads are consuming memory:

Pod A       → 15 GB
Pod B       → 12 GB
Pod C       → 18 GB
Pod D       → 10 GB
System      →  8 GB
---------------------
Total       → 63 GB

The node is now under severe memory pressure.

If memory consumption continues and the Linux kernel cannot satisfy memory allocations, the kernel may invoke the OOM killer.

This is fundamentally different from a single container exceeding its configured memory limit.

The problem is now:

The node itself is running out of available memory.


5. How to identify Node Memory Pressure

From the Kubernetes control plane, check the node:

kubectl describe node <node-name>

Look at the Conditions section.

You may find:

Conditions:
  MemoryPressure   True

This means the kubelet has detected memory pressure on the node.

You should then investigate the actual node.

Run:

free -h

Example:

               total        used        free
Mem:             64Gi         62Gi       500Mi
Swap:               0           0           0

Other useful commands include:

top
vmstat 1
ps aux --sort=-%mem

These help identify which processes are consuming memory.


6. Check Linux Kernel Logs

For a node-level OOM investigation, kernel logs are extremely important.

Run:

dmesg -T | grep -i -E "oom|out of memory|killed process"

Or:

journalctl -k | grep -i -E "oom|out of memory|killed process"

You may find something similar to:

Out of memory: Killed process 12345 (java)

This is strong evidence that the Linux kernel invoked the OOM killer.

At this point, your investigation has moved beyond Kubernetes objects and into the Linux operating system layer.

That is an important skill for a Kubernetes SRE.


7. The SRE Mental Model

Think about a Kubernetes node like this:

                 KUBERNETES NODE
        ┌─────────────────────────────┐
        │                             │
        │  Pod A                      │
        │  ┌───────────────────────┐  │
        │  │ Container             │  │
        │  │ Memory limit: 2 GiB   │  │
        │  └───────────────────────┘  │
        │                             │
        │  Pod B                      │
        │                             │
        │  Pod C                      │
        │                             │
        │  kubelet                    │
        │  system processes           │
        │                             │
        └─────────────────────────────┘

Container OOM

Container memory usage
        ↓
Exceeds container limit
        ↓
Container/process killed
        ↓
Kubernetes reports OOMKilled
        ↓
Container may restart

Node-level OOM

Total node memory consumption
        ↓
Available memory becomes critically low
        ↓
Linux kernel cannot satisfy allocation
        ↓
Kernel OOM killer
        ↓
Process/container killed
        ↓
Potential impact to multiple workloads

8. Don't confuse OOM with Pod Eviction

This is where Kubernetes troubleshooting becomes more interesting.

There are at least three different situations an SRE should distinguish.

Situation 1 — Container OOM

Container exceeds memory limit
        ↓
Container/process killed
        ↓
OOMKilled

Situation 2 — Node Memory Pressure

Node available memory becomes low
        ↓
Kubelet detects memory pressure
        ↓
MemoryPressure = True
        ↓
Kubernetes may evict Pods

Situation 3 — Node-level Linux OOM

Node cannot satisfy memory allocation
        ↓
Linux kernel OOM killer
        ↓
Process/container killed

These situations are related, but they are not the same event.


9. Production Troubleshooting Methodology

When investigating a memory-related incident, don't immediately jump to:

"The application has an OOM."

Follow an evidence-based approach.

Step 1 — Check Pod status

kubectl get pod <pod> -n <namespace>

Example:

NAME        READY   STATUS             RESTARTS
myapp-01    0/1     CrashLoopBackOff   8

Step 2 — Describe the Pod

kubectl describe pod <pod> -n <namespace>

Look for:

Last State:
  Terminated:
    Reason: OOMKilled
    Exit Code: 137

Also check Events.


Step 3 — Check container logs

kubectl logs <pod> -n <namespace>

If the container has restarted:

kubectl logs <pod> -n <namespace> --previous

The --previous option is particularly useful because the current container may have already restarted.


10. Check Resource Requests and Limits

Inspect the Pod configuration:

kubectl get pod <pod> -n <namespace> -o yaml

Look for:

resources:
  requests:
    memory: "1Gi"
  limits:
    memory: "2Gi"

Ask:

  • Is the memory limit too low?

  • Is the application memory usage increasing?

  • Is there a memory leak?

  • Are requests and limits configured correctly?

  • Has application behavior changed recently?


11. Check the Node

Find which node is running the Pod:

kubectl get pod <pod> -n <namespace> -o wide

Then:

kubectl describe node <node-name>

Check:

MemoryPressure
Allocatable memory
Allocated resources
Conditions
Events

12. Check Node Memory from Linux

On the affected node:

free -h
top
vmstat 1
ps aux --sort=-%mem

Also check:

df -h

Although df -h primarily checks filesystem capacity rather than RAM, it is useful during broader node-health investigations because memory incidents can occur alongside other resource exhaustion problems.


13. Check Kernel Evidence

Finally:

dmesg -T | grep -i -E "oom|out of memory|killed process"

or:

journalctl -k | grep -i -E "oom|out of memory|killed process"

Now correlate:

Kubernetes evidence
        +
Container evidence
        +
Node evidence
        +
Linux kernel evidence
        ↓
Root Cause

This is the approach an experienced SRE should follow.


14. Example Production RCA

Suppose an application Pod restarted several times.

You run:

kubectl describe pod payment-01 -n production

and find:

Last State:
  Terminated:
    Reason: OOMKilled
    Exit Code: 137

The Pod configuration shows:

resources:
  requests:
    memory: "1Gi"
  limits:
    memory: "2Gi"

Monitoring shows the application memory usage increased steadily:

10:00 → 1.1 GiB
10:10 → 1.4 GiB
10:20 → 1.7 GiB
10:30 → 1.9 GiB
10:35 → 2.0 GiB
10:36 → OOMKilled

The evidence indicates:

Application memory consumption increased
             ↓
Container reached 2 GiB limit
             ↓
Container was killed
             ↓
Kubernetes reported OOMKilled
             ↓
Container restarted

This is primarily a container-level OOM, not proof of a node-level OOM.


15. Example Node-Level Incident

Now consider:

Node memory = 64 GiB

Pod A = 15 GiB
Pod B = 12 GiB
Pod C = 18 GiB
Pod D = 10 GiB
System = 8 GiB

The node becomes heavily memory constrained.

You observe:

kubectl describe node worker-01
MemoryPressure: True

Then:

journalctl -k | grep -i oom

returns:

Out of memory: Killed process 12345 (java)

Now you have evidence of a node-level memory exhaustion event.

The impact can potentially extend beyond one application.


16. Interview Question

A common senior Kubernetes SRE interview question is:

A Pod shows OOMKilled. Does that mean the Kubernetes node experienced an OOM?

Correct answer

No.

OOMKilled indicates that the container/process was killed because of an out-of-memory condition, commonly because the container exceeded its configured memory limit.

It does not by itself prove that the Kubernetes node experienced a system-wide OOM.

To determine whether the node itself experienced memory exhaustion, check:

kubectl describe node <node>

Look for:

MemoryPressure: True

Then check the node's Linux/kernel evidence:

free -h
dmesg -T | grep -i oom
journalctl -k | grep -i oom

This distinction is extremely important in production troubleshooting.


17. Senior SRE Interview Answer

If an interviewer asks:

"Explain the difference between Container OOM and Node OOM."

A strong answer would be:

Container OOM is a container-level memory exhaustion condition, typically occurring when a container exceeds its configured memory limit. Kubernetes may report the container termination as OOMKilled, and the container may subsequently restart.

Node OOM is a node-level memory exhaustion condition where the overall system runs critically low on memory. The Linux kernel may invoke the OOM killer and terminate processes, potentially affecting multiple workloads. Kubernetes may also detect node memory pressure and evict Pods before the system reaches a hard kernel OOM condition.

During troubleshooting, I would correlate Pod termination reasons, resource requests and limits, node MemoryPressure, kubelet events, node memory utilization, and Linux kernel logs to determine the actual root cause.

That demonstrates production troubleshooting rather than simple Kubernetes theory.


18. Quick SRE Cheat Sheet

                 MEMORY INCIDENT
                       │
                       ▼
               Check Pod status
                       │
                       ▼
              kubectl describe pod
                       │
              ┌────────┴────────┐
              ▼                 ▼
          OOMKilled?          Other?
              │                 │
              ▼                 ▼
       Check limits          Continue
       Check usage           troubleshooting
       Check logs
              │
              ▼
       Check affected node
              │
              ▼
     MemoryPressure = True?
              │
        ┌─────┴─────┐
        ▼           ▼
       YES          NO
        │            │
        ▼            ▼
 Check node       Investigate
 memory           container/app
        │
        ▼
 Check kernel logs
        │
        ▼
 Root Cause

Conclusion

The most important lesson is:

OOMKilled does not automatically mean "the Kubernetes node ran out of memory."

Always determine the scope of the memory problem:

Container
   ↓
Pod
   ↓
Node
   ↓
Cluster

For a Kubernetes SRE, the goal is not simply to identify that "OOM happened."

The real goal is to answer:

What ran out of memory? Why did it happen? What evidence proves it? What was the impact? How do we fix it? And how do we prevent it from happening again?


Sunday, 30 August 2026

Building a Kubernetes SRE Incident Investigation Agent with Python & Agentic AI

Short Description

Build a read-only AI-powered Kubernetes SRE agent that investigates pod incidents using status, logs, previous logs, and Kubernetes events. The agent dynamically collects evidence, forms and validates hypotheses, and generates a structured RCA with root-cause confidence—without making unsupported assumptions or performing remediation.





(.venv) PS C:\Users\Raj Kumar Gupta\Desktop\Raj\minikube\k8s-day5-sre-agent> kubectl get pod

NAME                          READY   STATUS             RESTARTS        AGE

app-crash-5cf8866df5-q6k4m    0/1     CrashLoopBackOff   8 (40s ago)     17m

broken-pod                    0/1     ImagePullBackOff   0               258d

payment-api-987c667b8-xrqhf   0/1     CrashLoopBackOff   45 (3m1s ago)   20h

(.venv) PS C:\Users\Raj Kumar Gupta\Desktop\Raj\minikube\k8s-day5-sre-agent> python -m agents.sre_agent

Enter namespace [default]: default

Enter pod name: app-crash-5cf8866df5-q6k4m


======================================================================

KUBERNETES SRE INVESTIGATION

======================================================================

Namespace: default

Pod: app-crash-5cf8866df5-q6k4m


======================================================================

AGENT ITERATION 1

======================================================================


======================================================================

TOOL CALL

======================================================================

Tool: get_pod_status

Arguments: {'namespace': 'default', 'pod_name': 'app-crash-5cf8866df5-q6k4m'}

[K8S] Using local kubeconfig


======================================================================

AGENT ITERATION 2

======================================================================


======================================================================

TOOL CALL

======================================================================

Tool: get_pod_previous_logs

Arguments: {'namespace': 'default', 'pod_name': 'app-crash-5cf8866df5-q6k4m'}

[K8S] Using local kubeconfig


======================================================================

TOOL CALL

======================================================================

Tool: get_pod_events

Arguments: {'namespace': 'default', 'pod_name': 'app-crash-5cf8866df5-q6k4m'}

[K8S] Using local kubeconfig


======================================================================

AGENT ITERATION 3

======================================================================


======================================================================

TOOL CALL

======================================================================

Tool: get_pod_status

Arguments: {'namespace': 'default', 'pod_name': 'app-crash-5cf8866df5-q6k4m'}

[K8S] Using local kubeconfig


======================================================================

AGENT ITERATION 4

======================================================================


======================================================================

FINAL SRE REPORT

======================================================================

SYMPTOM

- Pod default/app-crash-5cf8866df5-q6k4m is failing with CrashLoopBackOff; container is not ready and has multiple restarts.


EVIDENCE

- Pod status (facts):

  - phase: Running

  - container "busybox": ready=false, restart_count=8, state=waiting, reason=CrashLoopBackOff, message="back-off 5m0s restarting failed container=busybox pod=app-crash-5cf8866df5-q6k4m_default(69fb3e78-994f-42cd-98f2-981bbfa58147)"

  - node: minikube, pod_ip: 10.244.0.14

- Pod previous container logs (facts):

  - "Application started"

  - "ERROR database connection refused"

- Pod events (facts):

  - Scheduled on minikube successfully.

  - Image pulls for "busybox" succeeded repeatedly.

  - Warning BackOff: "Back-off restarting failed container busybox..." count=70 between 05:50:54 and 06:06:01 UTC.

  - No warnings/errors about image pull failures, scheduling failures, or probe failures in the events listed.


HYPOTHESIS

- The application starts, immediately attempts a database connection, receives "connection refused," exits, and the container restarts. Repeated exits lead to CrashLoopBackOff.

- The underlying reason for "connection refused" could be incorrect DB host/port, missing/incorrect service/endpoints, network policy, credentials, or the DB process not accepting connections. This is not confirmed.


VALIDATION

- Correlation:

  - Pod status shows CrashLoopBackOff with 8 restarts (FACT).

  - Events show repeated back-off restarts and successful image pulls/scheduling (FACT), indicating the failure occurs during/after container start, not during scheduling or image pulling.

  - Previous logs show the last messages before termination include "ERROR database connection refused" (FACT), consistent with the container exiting shortly after startup.

- What is not proven:

  - We do not have the container’s last termination exit code or reason (e.g., Error vs OOMKilled), so we cannot assert the exact process exit details (FACT). However, there is no evidence of OOM or node pressure in events, and the log content points to an application-level exit after a DB connection failure (INFERENCE).


ROOT CAUSE

- Immediate cause confirmed: The container repeatedly exits after encountering "ERROR database connection refused" at startup, causing CrashLoopBackOff.

- Underlying cause of the database connection refusal (misconfiguration, network, service/endpoints, DB availability, credentials) is not established with the current evidence.

- Root cause not confirmed.


CONFIDENCE

- Medium. Logs and events clearly align with application exit due to a DB connection refusal, but lack of last termination details and no verification of the database endpoint/config prevents confirming the underlying root cause.


NEXT INVESTIGATION

- Retrieve container last termination details (exitCode, reason, signal) to confirm process exit conditions.

- Verify DB connectivity configuration used by this pod:

  - Environment variables, config maps, and secrets for DB host, port, database name, and credentials.

  - Container command/entrypoint to see if it exits on first connection failure.

- Validate the Kubernetes Service backing the database:

  - Confirm the Service exists, DNS name is correct, and port matches what the app uses.

  - Check Endpoints for the Service to ensure there are ready targets.

- Network path checks:

  - Confirm no NetworkPolicy or firewall rules block traffic from this pod’s namespace to the DB.

  - If possible, run a temporary debug pod in the same namespace to test TCP connectivity to the DB host:port.

- Inspect the database side:

  - Ensure the DB process is listening on the expected interface/port and accepts connections from the pod’s network.

  - Review DB logs for connection attempts/refusals at the corresponding timestamps.

- If liveness/readiness probes are configured for this pod, confirm they aren’t killing the container; correlate probe failures if any appear in describe/events.

(.venv) PS C:\Users\Raj Kumar Gupta\Desktop\Raj\minikube\k8s-day5-sre-agent> 

================================================

As a Kubernetes SRE, troubleshooting a production incident usually involves collecting information from multiple sources:

  • Pod status

  • Container restart count

  • Exit codes

  • Application logs

  • Previous container logs

  • Kubernetes events

  • Node information

  • Scheduling information

  • Application symptoms

The challenge is not simply collecting this information.

The real challenge is deciding:

What should I investigate next based on the evidence I already have?

This is where Agentic AI becomes interesting.

Instead of creating a Python script that always executes:

get pod status
    ↓
get logs
    ↓
get events

we can build an AI agent that dynamically decides what information it needs next.


The investigation flow is:

Incident
   ↓
Observe
   ↓
Collect Evidence
   ↓
Reason
   ↓
Create Hypothesis
   ↓
Validate Hypothesis
   ↓
Collect More Evidence if Required
   ↓
Confirm Root Cause
   ↓
Generate RCA

This follows the structured investigation approach described : Incident → Evidence → Hypothesis → Validation → Root Cause → Confidence → Next Investigation.


1. What Are We Building?

We are going to build a Python-based Kubernetes SRE Agent.

The user provides:

Namespace
Pod Name

The AI agent then investigates the pod using Kubernetes tools.

The architecture is:

                    USER
                      |
                      v
              +---------------+
              |   SRE AGENT   |
              |      LLM      |
              +-------+-------+
                      |
              Dynamic Planning
                      |
        +-------------+-------------+
        |             |             |
        v             v             v
   Pod Status       Logs          Events
        |             |             |
        +-------------+-------------+
                      |
                      v
                  Evidence
                      |
                      v
                 Hypothesis
                      |
                      v
                  Validation
                      |
                      v
             Root Cause Confirmed?
                /             \
              NO               YES
              |                 |
              v                 v
        More Investigation      RCA
              |
              v
        Additional Tools

The key difference from a traditional script is:

The LLM decides which tool to call based on the current evidence.


2. What Will Our Agent Be Able to Investigate?

Our first version will support four Kubernetes investigation tools:

1. get_pod_status()
2. get_pod_logs()
3. get_pod_previous_logs()
4. get_pod_events()

These allow the agent to investigate common conditions such as:

CrashLoopBackOff
ImagePullBackOff
Pending
Container restart
Application startup failure
Scheduling failure



3. Why FACT, HYPOTHESIS and ROOT CAUSE Must Be Different

This is one of the most important concepts in an SRE AI agent.

Suppose the application log says:

ERROR database connection refused

A weak AI agent might immediately report:

Root Cause:
Database is down.

That is not necessarily correct.

The actual evidence only proves:

FACT:
Application received "connection refused".

We can infer:

INFERENCE:
Application could not establish a database connection.

We can form:

HYPOTHESIS:
There may be a database connectivity problem.

But we cannot yet claim:

ROOT CAUSE:
Database is down.

We need more evidence.

Potential causes could include:

Database server
Database Service
Service endpoints
DNS
NetworkPolicy
Firewall
Credentials
Wrong hostname
Wrong port
Application configuration

Therefore:

FACT != HYPOTHESIS

HYPOTHESIS != ROOT CAUSE



4. Project Directory Structure

Create the following project:

k8s-day5-sre-agent/
│
├── .venv/
│
├── agents/
│   ├── __init__.py
│   └── sre_agent.py
│
├── tools/
│   ├── __init__.py
│   └── agent_tools.py
│
├── models/
│   ├── __init__.py
│   └── incident.py
│
├── scenarios/
│   ├── app-pending.yaml
│   ├── create_incidents.ps1
│   └── cleanup.ps1
│
├── tests/
│   ├── __init__.py
│   └── test_tools.py
│
├── .env
├── .gitignore
├── requirements.txt
└── README.md

The important components are:

agents/
    AI agent

tools/
    Kubernetes tools

models/
    Structured incident model

scenarios/
    Test incidents

tests/
    Tool testing

5. Create the Project

On Windows PowerShell:

mkdir k8s-day5-sre-agent
cd k8s-day5-sre-agent

Create the virtual environment:

python -m venv .venv

Activate it:

.\.venv\Scripts\Activate.ps1

Verify:

python --version

You should see:

Python 3.x.x

6. Install Python Dependencies

Create:

requirements.txt

Add:

openai
kubernetes
python-dotenv
pydantic

Install:

pip install -r requirements.txt

Verify:

pip list

7. Configure the OpenAI API Key

Create:

.env

Add:

OPENAI_API_KEY=YOUR_API_KEY_HERE

For example:

OPENAI_API_KEY=sk-xxxxxxxxxxxxxxxx

Do not commit .env to Git.


8. Create .gitignore

Create:

.gitignore

Add:

.venv/
.env
__pycache__/
*.pyc

This prevents credentials and Python-generated files from being committed.


9. Kubernetes Tool Layer

The AI agent itself should not directly implement Kubernetes API calls everywhere.

Instead, we create a dedicated tool layer:

tools/agent_tools.py

This provides a clean separation:

AI Agent
   |
   v
Tool Layer
   |
   v
Kubernetes API

10. Complete tools/agent_tools.py

Create:

tools/agent_tools.py

Use the following complete code:

from kubernetes import client, config
from kubernetes.config.config_exception import ConfigException


def load_kubernetes():
    """
    Load Kubernetes configuration.

    First try in-cluster configuration.
    If that fails, use the local kubeconfig.
    """

    try:
        config.load_incluster_config()

        print("[K8S] Using in-cluster configuration")

    except ConfigException:

        config.load_kube_config()

        print("[K8S] Using local kubeconfig")


def get_pod_status(
    namespace: str,
    pod_name: str
):
    """
    Get Kubernetes pod status.

    Returns:
        - pod phase
        - node
        - pod IP
        - container state
        - restart count
        - exit code
        - termination reason
    """

    load_kubernetes()

    v1 = client.CoreV1Api()

    pod = v1.read_namespaced_pod(
        name=pod_name,
        namespace=namespace
    )

    containers = []

    if pod.status.container_statuses:

        for container in pod.status.container_statuses:

            state = {}

            if container.state:

                if container.state.waiting:

                    state = {
                        "state": "waiting",
                        "reason": container.state.waiting.reason,
                        "message": container.state.waiting.message
                    }

                elif container.state.running:

                    state = {
                        "state": "running",
                        "started_at": str(
                            container.state.running.started_at
                        )
                    }

                elif container.state.terminated:

                    state = {
                        "state": "terminated",
                        "reason": container.state.terminated.reason,
                        "exit_code": container.state.terminated.exit_code,
                        "signal": container.state.terminated.signal,
                        "message": container.state.terminated.message
                    }

            containers.append(
                {
                    "name": container.name,
                    "ready": container.ready,
                    "restart_count": container.restart_count,
                    "state": state
                }
            )

    return {
        "pod_name": pod.metadata.name,
        "namespace": pod.metadata.namespace,
        "phase": pod.status.phase,
        "node": pod.spec.node_name,
        "pod_ip": pod.status.pod_ip,
        "containers": containers
    }


def get_pod_logs(
    namespace: str,
    pod_name: str,
    container: str = None
):
    """
    Get current logs from a Kubernetes container.
    """

    load_kubernetes()

    v1 = client.CoreV1Api()

    logs = v1.read_namespaced_pod_log(
        name=pod_name,
        namespace=namespace,
        container=container,
        tail_lines=100
    )

    return {
        "pod_name": pod_name,
        "namespace": namespace,
        "log_type": "current",
        "logs": logs
    }


def get_pod_previous_logs(
    namespace: str,
    pod_name: str,
    container: str = None
):
    """
    Get logs from the previous terminated container.

    Very useful for CrashLoopBackOff.
    """

    load_kubernetes()

    v1 = client.CoreV1Api()

    logs = v1.read_namespaced_pod_log(
        name=pod_name,
        namespace=namespace,
        container=container,
        previous=True,
        tail_lines=100
    )

    return {
        "pod_name": pod_name,
        "namespace": namespace,
        "log_type": "previous",
        "logs": logs
    }


def get_pod_events(
    namespace: str,
    pod_name: str
):
    """
    Get Kubernetes events associated with a pod.
    """

    load_kubernetes()

    v1 = client.CoreV1Api()

    events = v1.list_namespaced_event(
        namespace=namespace,
        field_selector=(
            f"involvedObject.name={pod_name}"
        )
    )

    result = []

    for event in events.items:

        result.append(
            {
                "type": event.type,
                "reason": event.reason,
                "message": event.message,
                "count": event.count,
                "first_timestamp": str(
                    event.first_timestamp
                ),
                "last_timestamp": str(
                    event.last_timestamp
                )
            }
        )

    return {
        "pod_name": pod_name,
        "namespace": namespace,
        "events": result
    }

11. Important Troubleshooting Lesson

During implementation, one common mistake is accidentally creating:

tools/agent_tools.py

as an empty file.

For example:

agent_tools.py    0 bytes

Then this command:

python -c "from tools.agent_tools import get_pod_status"

will produce:

ImportError:
cannot import name 'get_pod_status'

The reason is simple:

sre_agent.py
     |
     | imports
     v
agent_tools.py
     |
     X
     |
     No functions

After adding the code above, verify:

Get-Item .\tools\agent_tools.py

The file should have a size greater than zero.


12. Test the Kubernetes Tool Import

Run:

python -c "from tools.agent_tools import get_pod_status; print('TOOLS IMPORT OK')"

Expected:

TOOLS IMPORT OK

Test all four:

python -c "from tools.agent_tools import get_pod_status, get_pod_logs, get_pod_previous_logs, get_pod_events; print('ALL 4 TOOLS IMPORT OK')"

Expected:

ALL 4 TOOLS IMPORT OK

13. Incident Model

Create:

models/incident.py

Code:

from typing import List
from pydantic import BaseModel, Field


class Evidence(BaseModel):

    fact: str

    source: str


class IncidentReport(BaseModel):

    incident: str

    symptom: str

    evidence: List[Evidence] = Field(
        default_factory=list
    )

    hypothesis: str

    validation: List[str] = Field(
        default_factory=list
    )

    root_cause: str

    confidence: str

    next_investigation: List[str] = Field(
        default_factory=list
    )

This allows us to represent the investigation in a structured format.

The intended report structure is:

SYMPTOM
EVIDENCE
HYPOTHESIS
VALIDATION
ROOT CAUSE
CONFIDENCE
NEXT INVESTIGATION

14. Building the AI SRE Agent

Now create:

agents/sre_agent.py

The agent will:

  1. Accept namespace and pod name.

  2. Ask the LLM to investigate.

  3. Allow the LLM to select tools.

  4. Execute the selected Kubernetes tool.

  5. Send the result back to the LLM.

  6. Allow the LLM to select another tool if required.

  7. Continue until sufficient evidence exists.

  8. Generate the final SRE report.


15. Complete agents/sre_agent.py

Use this complete code:

import json
import os

from dotenv import load_dotenv
from openai import OpenAI

from tools.agent_tools import (
    get_pod_status,
    get_pod_logs,
    get_pod_previous_logs,
    get_pod_events,
)


# ============================================================
# ENVIRONMENT
# ============================================================

load_dotenv()

api_key = os.getenv("OPENAI_API_KEY")

if not api_key:

    raise RuntimeError(
        "OPENAI_API_KEY is missing. "
        "Create a .env file in the project root."
    )


# ============================================================
# OPENAI CLIENT
# ============================================================

client = OpenAI(
    api_key=api_key
)


# ============================================================
# SYSTEM PROMPT
# ============================================================

SYSTEM_PROMPT = """
You are a senior Kubernetes SRE investigation agent.

Your job is to investigate Kubernetes incidents
using available READ-ONLY tools.

Investigation rules:

1. Never invent cluster state.

2. Always collect evidence before making conclusions.

3. Clearly distinguish FACT from INFERENCE.

4. Clearly distinguish HYPOTHESIS from CONFIRMED ROOT CAUSE.

5. Never claim a root cause without supporting evidence.

6. Use additional tools when evidence is insufficient.

7. Do not perform remediation.

8. Prefer the smallest number of tools necessary.

9. Correlate pod status, logs and Kubernetes events.

10. If the root cause cannot be proven, explicitly say:

    Root cause not confirmed.


Important:

Exit code 137 does NOT automatically prove OOMKilled.

Database connection refused does NOT automatically prove
that the database is down.

ImagePullBackOff should be investigated using Kubernetes
events and pod status.

Pending pods should be investigated using pod status and
Kubernetes scheduling events.


The final answer must contain:

SYMPTOM
EVIDENCE
HYPOTHESIS
VALIDATION
ROOT CAUSE
CONFIDENCE
NEXT INVESTIGATION
"""


# ============================================================
# TOOL DEFINITIONS
# ============================================================

TOOLS = [

    {
        "type": "function",

        "name": "get_pod_status",

        "description": (
            "Get Kubernetes pod status, container states, "
            "restart counts, exit codes and node information."
        ),

        "parameters": {

            "type": "object",

            "properties": {

                "namespace": {
                    "type": "string"
                },

                "pod_name": {
                    "type": "string"
                }
            },

            "required": [
                "namespace",
                "pod_name"
            ]
        }
    },


    {
        "type": "function",

        "name": "get_pod_logs",

        "description": (
            "Get current logs from a Kubernetes pod."
        ),

        "parameters": {

            "type": "object",

            "properties": {

                "namespace": {
                    "type": "string"
                },

                "pod_name": {
                    "type": "string"
                }
            },

            "required": [
                "namespace",
                "pod_name"
            ]
        }
    },


    {
        "type": "function",

        "name": "get_pod_previous_logs",

        "description": (
            "Get logs from the previous terminated "
            "container instance. Useful for CrashLoopBackOff."
        ),

        "parameters": {

            "type": "object",

            "properties": {

                "namespace": {
                    "type": "string"
                },

                "pod_name": {
                    "type": "string"
                }
            },

            "required": [
                "namespace",
                "pod_name"
            ]
        }
    },


    {
        "type": "function",

        "name": "get_pod_events",

        "description": (
            "Get Kubernetes events associated with a pod. "
            "Useful for scheduling, image pull, startup and "
            "restart failures."
        ),

        "parameters": {

            "type": "object",

            "properties": {

                "namespace": {
                    "type": "string"
                },

                "pod_name": {
                    "type": "string"
                }
            },

            "required": [
                "namespace",
                "pod_name"
            ]
        }
    }
]


# ============================================================
# TOOL EXECUTION
# ============================================================

def execute_tool(name, arguments):

    print()
    print("=" * 70)
    print("TOOL CALL")
    print("=" * 70)

    print("Tool:", name)

    print(
        "Arguments:",
        json.dumps(
            arguments,
            indent=2
        )
    )

    try:

        if name == "get_pod_status":

            return get_pod_status(
                arguments["namespace"],
                arguments["pod_name"]
            )


        elif name == "get_pod_logs":

            return get_pod_logs(
                arguments["namespace"],
                arguments["pod_name"]
            )


        elif name == "get_pod_previous_logs":

            return get_pod_previous_logs(
                arguments["namespace"],
                arguments["pod_name"]
            )


        elif name == "get_pod_events":

            return get_pod_events(
                arguments["namespace"],
                arguments["pod_name"]
            )


        else:

            return {
                "error": f"Unknown tool: {name}"
            }


    except Exception as exc:

        return {
            "error": str(exc)
        }


# ============================================================
# INVESTIGATION ENGINE
# ============================================================

def investigate(
    namespace: str,
    pod_name: str
):

    user_prompt = f"""
Investigate this Kubernetes incident.

Namespace:
{namespace}

Pod:
{pod_name}


Determine why this pod is failing.

Use evidence.

Do not make assumptions.

Do not perform remediation.

Use additional tools if the available evidence
is insufficient.

At the end produce a structured SRE incident report.
"""


    messages = [

        {
            "role": "system",

            "content": SYSTEM_PROMPT
        },

        {
            "role": "user",

            "content": user_prompt
        }
    ]


    max_iterations = 10


    for iteration in range(max_iterations):

        print()
        print("=" * 70)

        print(
            f"AGENT ITERATION {iteration + 1}"
        )

        print("=" * 70)


        response = client.responses.create(

            model="gpt-5",

            input=messages,

            tools=TOOLS
        )


        tool_outputs = []


        for item in response.output:

            if item.type == "function_call":

                name = item.name


                arguments = json.loads(
                    item.arguments
                )


                result = execute_tool(
                    name,
                    arguments
                )


                tool_outputs.append(

                    {
                        "type": "function_call_output",

                        "call_id": item.call_id,

                        "output": json.dumps(
                            result,
                            default=str
                        )
                    }

                )


        # ----------------------------------------------------
        # No more tools required
        # ----------------------------------------------------

        if not tool_outputs:

            return response.output_text


        # ----------------------------------------------------
        # Send assistant tool request back to conversation
        # ----------------------------------------------------

        messages.extend(
            response.output
        )


        # ----------------------------------------------------
        # Send tool results back to LLM
        # ----------------------------------------------------

        messages.extend(
            tool_outputs
        )


    return (
        "Investigation stopped because maximum "
        "agent iterations were reached."
    )


# ============================================================
# MAIN
# ============================================================

def main():

    namespace = input(
        "Enter namespace [default]: "
    ).strip()


    if not namespace:

        namespace = "default"


    pod_name = input(
        "Enter pod name: "
    ).strip()


    if not pod_name:

        print(
            "Pod name is required."
        )

        return


    print()

    print("=" * 70)

    print(
        "KUBERNETES SRE INVESTIGATION"
    )

    print("=" * 70)


    print(
        "Namespace:",
        namespace
    )

    print(
        "Pod:",
        pod_name
    )


    result = investigate(
        namespace,
        pod_name
    )


    print()

    print("=" * 70)

    print(
        "FINAL SRE REPORT"
    )

    print("=" * 70)


    print(result)


if __name__ == "__main__":

    main()

16. Why We Use python -m

Our project uses:

k8s-day5-sre-agent/
│
├── agents/
│   └── sre_agent.py
│
└── tools/
    └── agent_tools.py

Therefore, run:

python -m agents.sre_agent

Do not normally run:

python .\agents\sre_agent.py

The module approach allows Python to correctly find:

from tools.agent_tools import ...

17. Verify the Agent Import

Run:

python -c "import agents.sre_agent; print('AGENT IMPORT OK')"

Expected:

AGENT IMPORT OK

If you get:

ModuleNotFoundError: No module named 'tools'

you are probably executing the file directly instead of using:

python -m agents.sre_agent

18. Verify the API Key

Run:

python -c "from dotenv import load_dotenv; import os; load_dotenv(); print('API KEY STATUS:', 'SET' if os.getenv('OPENAI_API_KEY') else 'NOT SET')"

Expected:

API KEY STATUS: SET

If you get:

API KEY STATUS: NOT SET

check:

k8s-day5-sre-agent/
│
└── .env

and make sure it contains:

OPENAI_API_KEY=YOUR_API_KEY

Do not share your actual API key publicly.


19. Verify Kubernetes

Before running the AI agent:

minikube status

Then:

kubectl get nodes

Expected:

NAME       STATUS   ROLES           AGE
minikube   Ready    control-plane   ...

Then:

kubectl get pods -A

20. Create a CrashLoopBackOff Incident

We need a controlled incident for testing.

Run:

kubectl create deployment app-crash `
  --image=busybox `
  -- /bin/sh -c "echo Application started; echo ERROR database connection refused; exit 1"

Check:

kubectl get pods

After a while:

NAME                          READY   STATUS             RESTARTS
app-crash-xxxxxxxxxx          0/1     CrashLoopBackOff   3

Copy the actual pod name.


21. Manually Investigate the Incident

Before allowing AI to investigate, it is useful for an SRE to understand the raw evidence.

Run:

kubectl get pod YOUR_POD_NAME -o wide

Then:

kubectl describe pod YOUR_POD_NAME

Current logs:

kubectl logs YOUR_POD_NAME

Previous logs:

kubectl logs YOUR_POD_NAME --previous

Events:

kubectl get events --sort-by=.lastTimestamp

You may see:

Application started
ERROR database connection refused

and events such as:

BackOff

22. Understand the Evidence

The correct reasoning is:

FACT
Application logged:
database connection refused

Then:

INFERENCE
The application could not establish a database connection.

Then:

HYPOTHESIS
Database connectivity is a possible cause.

But:

ROOT CAUSE
Database is down

is not yet proven.

This is the difference between an AI chatbot that generates plausible answers and an SRE investigation agent that works from evidence.


23. Run the AI Agent

Run:

python -m agents.sre_agent

You should see:

Enter namespace [default]:

Enter:

default

Then:

Enter pod name:

Enter your actual pod:

app-crash-xxxxxxxxxx

24. What Happens Inside the Agent?

The first iteration may look conceptually like:

AGENT ITERATION 1
        |
        v
LLM asks:
"What is the pod state?"
        |
        v
get_pod_status()

The tool returns:

phase = Running
restart_count = 3
container state = terminated
exit_code = 1

The LLM then reasons:

The container is restarting.
I need logs.

It calls:

get_pod_previous_logs()

The result:

Application started
ERROR database connection refused

The LLM may then call:

get_pod_events()

The result:

BackOff

Now it has multiple pieces of evidence.


25. Dynamic Investigation

The important point is that the Python program does not explicitly contain:

if status == "CrashLoopBackOff":
    get_logs()

if logs contain "database":
    get_events()

Instead:

LLM
 |
 +--> get_pod_status()
 |
 +--> inspect result
 |
 +--> decide next action
 |
 +--> get_pod_previous_logs()
 |
 +--> inspect result
 |
 +--> decide next action
 |
 +--> get_pod_events()
 |
 +--> inspect result
 |
 +--> generate report

This is the beginning of agentic behavior.


26. Expected SRE Report

The exact wording generated by the LLM may vary, but the report should resemble:

SYMPTOM
-------
Pod app-crash-xxxx is repeatedly restarting
and is currently in CrashLoopBackOff.


EVIDENCE
--------
1. Container restart count is greater than zero.

2. Container terminated with exit code 1.

3. Previous container logs contain:
   "ERROR database connection refused"

4. Kubernetes events show BackOff.


HYPOTHESIS
----------
The application is failing because it cannot
establish database connectivity.


VALIDATION
----------
The previous container logs confirm that the
application reported a database connection failure.

However, database availability itself has not
been independently verified.


ROOT CAUSE
----------
Root cause not confirmed.

The evidence confirms an application-level
database connection failure but does not prove
that the database server itself is down.


CONFIDENCE
----------
Medium


NEXT INVESTIGATION
------------------
1. Check database Service.
2. Check Service endpoints.
3. Check DNS resolution.
4. Check NetworkPolicy.
5. Check database availability.
6. Validate application database configuration.

This is consistent with the Day-5 target report structure.


27. Incident 2 — ImagePullBackOff

Now create a bad image:

kubectl create deployment app-image `
  --image=nginx:this-image-does-not-exist

Check:

kubectl get pods

You should see something similar to:

app-image-xxxxxxxxxx    0/1    ImagePullBackOff

Now run:

python -m agents.sre_agent

Enter:

default

and the actual pod name.

The expected investigation pattern is:

Pod status
    |
    v
ImagePullBackOff
    |
    v
Kubernetes events
    |
    v
Failed to pull image
    |
    v
Diagnosis

The agent should not treat this like a CrashLoopBackOff investigation.

That is the purpose of dynamic planning.


28. Incident 3 — Pending Pod

A normal Minikube cluster may successfully schedule:

nginx

Therefore, simply creating an nginx deployment does not guarantee Pending.

Instead, deliberately create an impossible node selector.

Create:

scenarios/app-pending.yaml

Use:

apiVersion: apps/v1
kind: Deployment

metadata:
  name: app-pending

spec:
  replicas: 1

  selector:
    matchLabels:
      app: app-pending

  template:

    metadata:
      labels:
        app: app-pending

    spec:

      nodeSelector:
        sre-test-node: "does-not-exist"

      containers:

        - name: nginx
          image: nginx

Apply:

kubectl apply -f scenarios/app-pending.yaml

Check:

kubectl get pods

Expected:

app-pending-xxxxxxxxxx    0/1    Pending

29. Investigate Pending

Run:

kubectl describe pod YOUR_PENDING_POD

Near the bottom you should see scheduling events similar to:

Events:

Warning  FailedScheduling

0/1 nodes are available:
1 node(s) didn't match Pod's node affinity/selector.

The agent should follow:

Pending
   |
   v
Pod Status
   |
   v
FailedScheduling
   |
   v
Kubernetes Events
   |
   v
Node selector / scheduling constraint

30. Test Kubernetes Tools Independently

Create:

tests/test_tools.py

Use:

from tools.agent_tools import (
    get_pod_status,
    get_pod_logs,
    get_pod_previous_logs,
    get_pod_events
)


POD_NAME = "YOUR_POD_NAME"

NAMESPACE = "default"


print("=" * 70)

print("TEST 1: POD STATUS")

print("=" * 70)

result = get_pod_status(
    NAMESPACE,
    POD_NAME
)

print(result)


print()

print("=" * 70)

print("TEST 2: CURRENT LOGS")

print("=" * 70)

result = get_pod_logs(
    NAMESPACE,
    POD_NAME
)

print(result)


print()

print("=" * 70)

print("TEST 3: PREVIOUS LOGS")

print("=" * 70)

try:

    result = get_pod_previous_logs(
        NAMESPACE,
        POD_NAME
    )

    print(result)

except Exception as exc:

    print("Previous logs unavailable:")

    print(exc)


print()

print("=" * 70)

print("TEST 4: EVENTS")

print("=" * 70)

result = get_pod_events(
    NAMESPACE,
    POD_NAME
)

print(result)

Run:

python tests/test_tools.py

This lets us separate:

Kubernetes problem

from:

AI agent problem

That is an important engineering practice.


31. Cleanup Test Incidents

Create:

scenarios/cleanup.ps1

Use:

Write-Host "Cleaning SRE test incidents..."

kubectl delete deployment app-crash --ignore-not-found

kubectl delete deployment app-image --ignore-not-found

kubectl delete deployment app-pending --ignore-not-found

Write-Host ""

Write-Host "Cleanup complete."

kubectl get pods

Run:

.\scenarios\cleanup.ps1

32. Create an Automated Incident Generator

Create:

scenarios/create_incidents.ps1

Code:

Write-Host "============================================"
Write-Host "Creating Kubernetes SRE Test Incidents"
Write-Host "============================================"


Write-Host ""

Write-Host "[1] Creating CrashLoopBackOff incident"

kubectl create deployment app-crash `
  --image=busybox `
  -- /bin/sh -c "echo Application started; echo ERROR database connection refused; exit 1"


Write-Host ""

Write-Host "[2] Creating ImagePullBackOff incident"

kubectl create deployment app-image `
  --image=nginx:this-image-does-not-exist"


Write-Host ""

Write-Host "[3] Creating Pending incident"

kubectl apply -f scenarios/app-pending.yaml


Write-Host ""

Write-Host "============================================"

Write-Host "Incidents created"

Write-Host "============================================"


kubectl get pods

Run:

.\scenarios\create_incidents.ps1

33. Complete End-to-End Test

Now the complete workflow becomes:

cd k8s-day5-sre-agent

.\.venv\Scripts\Activate.ps1

minikube status

kubectl get nodes

pip install -r requirements.txt

.\scenarios\create_incidents.ps1

kubectl get pods

python -m agents.sre_agent

Investigate the CrashLoopBackOff pod.

Then:

python -m agents.sre_agent

Investigate the ImagePullBackOff pod.

Then:

python -m agents.sre_agent

Investigate the Pending pod.

Finally:

.\scenarios\cleanup.ps1

34. Complete Architecture

At this point the project looks like:

                         USER
                           |
                           v
                 +-------------------+
                 |    SRE AGENT      |
                 |      LLM          |
                 +---------+---------+
                           |
                           |
                    Dynamic Planning
                           |
          +----------------+----------------+
          |                |                |
          v                v                v
  get_pod_status()   get_pod_logs()   get_pod_events()
          |
          |
          v
get_pod_previous_logs()
          |
          +----------------+
                           |
                           v
                     EVIDENCE
                           |
                           v
                     HYPOTHESIS
                           |
                           v
                     VALIDATION
                           |
                           v
                 ROOT CAUSE CHECK
                    /          \
                   /            \
                  NO            YES
                  |              |
                  v              v
           More investigation   RCA
                  |
                  v
             More tools

35. Traditional Script vs Agentic SRE

A traditional script may look like:

get status
   ↓
get logs
   ↓
get events
   ↓
print output

The agentic approach is:

get status
   ↓
reason
   ↓
decide what evidence is missing
   ↓
get relevant tool
   ↓
reason again
   ↓
validate hypothesis
   ↓
decide whether root cause is proven
   ↓
generate RCA

This is a significant architectural difference.


36. Example: Exit Code 137

One of the important rules in our system is:

Exit code 137

must not automatically become:

OOMKilled

Why?

Because:

137 = 128 + 9

which corresponds to:

SIGKILL

Possible explanations require additional evidence.

The investigation should therefore be:

Exit code 137
       |
       v
SIGKILL
       |
       v
Possible OOM
       |
       v
Collect evidence
       |
       +----> Container state
       |
       +----> Kubernetes events
       |
       +----> Node memory
       |
       +----> Container memory limit
       |
       v
Confirm / Reject OOM hypothesis

This evidence-first reasoning is specifically emphasized in the Day-5 material.


37. Example: Database Connection Refused

Consider:

ERROR database connection refused

The agent should reason:

FACT
Application received connection refused.

        ↓

INFERENCE
Application could not establish DB connection.

        ↓

HYPOTHESIS
Database connectivity problem.

        ↓

VALIDATION
Need more evidence.

        ↓

INVESTIGATE
Service
Endpoints
DNS
NetworkPolicy
Database
Credentials
Configuration

        ↓

ROOT CAUSE
Only confirmed when evidence supports it.

This is much safer than letting an LLM invent a root cause.


38. Read-Only Safety Model

Our Day-5 agent is intentionally read-only.

The tools can:

READ pod status
READ logs
READ previous logs
READ events

They cannot:

DELETE pod
DELETE deployment
RESTART pod
SCALE deployment
PATCH deployment
CHANGE node
CHANGE network policy
CHANGE storage

This is an important production design principle.

The investigation agent should first establish:

What happened?

before eventually being allowed to answer:

What should we do?

And remediation should be a separate controlled capability.


39. Why Previous Logs Matter

For a normal running container:

kubectl logs POD

may be enough.

But for:

CrashLoopBackOff

the currently running container may have little or no useful information.

The previous container may contain:

startup logs
exception
stack trace
configuration failure
connection failure
termination reason

Therefore:

CrashLoopBackOff
      |
      v
Previous Logs
      |
      v
Failure Evidence

is an important SRE troubleshooting pattern.


40. Why Kubernetes Events Matter

Application logs tell us what the application experienced.

Kubernetes events tell us what Kubernetes experienced.

For example:

Application logs
----------------
database connection refused

versus:

Kubernetes events
-----------------
FailedScheduling
Failed to pull image
BackOff
FailedMount
Unhealthy

An SRE agent should correlate both.

Therefore:

Application Evidence
          +
Kubernetes Evidence
          =
Better Incident Diagnosis

41. Current Project Capabilities

After completing Day 5, our agent can:

✓ Connect to Kubernetes
✓ Inspect pod status
✓ Inspect container state
✓ Inspect restart counts
✓ Inspect exit codes
✓ Read current logs
✓ Read previous logs
✓ Read Kubernetes events
✓ Dynamically select tools
✓ Collect evidence
✓ Generate hypotheses
✓ Validate hypotheses
✓ Avoid unsupported root-cause claims
✓ Generate structured SRE reports

The Day-5 goal is to establish this structured evidence-driven investigation loop before moving toward broader observability and multi-agent capabilities.


42. What We Have Learned

The most important lesson from this project is:

An AI SRE agent should not be a chatbot that guesses the root cause.

It should behave more like an experienced SRE.

An experienced SRE thinks:

What do I know?
        ↓
What don't I know?
        ↓
What evidence would reduce uncertainty?
        ↓
Which command should I run?
        ↓
What did that command tell me?
        ↓
Does my hypothesis still hold?
        ↓
What should I investigate next?

Our AI agent is beginning to follow the same process.


43.  Mental Model

Remember this:

OBSERVE
   ↓
REASON
   ↓
HYPOTHESIZE
   ↓
COLLECT EVIDENCE
   ↓
VALIDATE
   ↓
REASON AGAIN
   ↓
CONFIRM ROOT CAUSE
   ↓
REPORT

This is the core mental model for building an Agentic AI SRE system.


44. Final Project Tree

The completed project should look like:

k8s-day5-sre-agent/
│
├── .venv/
│
├── agents/
│   ├── __init__.py
│   └── sre_agent.py
│
├── tools/
│   ├── __init__.py
│   └── agent_tools.py
│
├── models/
│   ├── __init__.py
│   └── incident.py
│
├── scenarios/
│   ├── app-pending.yaml
│   ├── create_incidents.ps1
│   └── cleanup.ps1
│
├── tests/
│   ├── __init__.py
│   └── test_tools.py
│
├── .env
├── .gitignore
├── requirements.txt
└── README.md

45. Final Execution Checklist

Use this checklist whenever you start the project:

# Activate environment
.\.venv\Scripts\Activate.ps1


# Check Python
python --version


# Check Kubernetes
minikube status


# Check node
kubectl get nodes


# Check API key
python -c "from dotenv import load_dotenv; import os; load_dotenv(); print('API KEY:', 'SET' if os.getenv('OPENAI_API_KEY') else 'NOT SET')"


# Check Kubernetes tools
python -c "from tools.agent_tools import get_pod_status, get_pod_logs, get_pod_previous_logs, get_pod_events; print('ALL 4 TOOLS IMPORT OK')"


# Check agent
python -c "import agents.sre_agent; print('AGENT IMPORT OK')"


# Run agent
python -m agents.sre_agent