Navigate the platform

Search sections and links. Use arrow keys and Enter to select.

PARAS BHANDERI

Engineering
reliable systems
for the AI era.

I design resilient cloud platforms, automate infrastructure, and build the systems that modern AI workloads depend on.

Platform Engineering · SRE · AI Infrastructure · Distributed Systems

Bangalore, India

SYSTEM TOPOLOGYCOMPUTE / CONTROL / SIGNAL
[EXPLORE THE SYSTEMHOVER OR FOCUS A NODE]
CONCEPTUAL SYSTEM · SIMULATED SIGNALS
RELIABILITY, BY DESIGN.SCROLL TO EXPLORE 01 / IDENTITY
02 / ENGINEERING IMPACTPRODUCTION OUTCOMES · FROM THE RÉSUMÉ
40%

Better deployment efficiency

Shell · Kubernetes scaling strategies

60%

Less manual deployment effort

Shell · Terraform on AWS & Azure

50+

Microservices made observable

Gore Mutual · Multi-cloud observability

20%

Reduction in cloud spend

Gore Mutual · Rightsizing & autoscaling

THE ENGINEERING PHILOSOPHY

Reliability isn’t added.
It’s engineered in.

I work at the intersection of reliability engineering, cloud platforms and AI infrastructure. My foundation is production SRE: Kubernetes, infrastructure as code, observability and automation across multi-cloud environments.

I’m applying that foundation to MCP-based systems, tool-calling agents and ML platforms—with model serving, GPU orchestration and LLMOps as the direction of my specialization.

paras@platform: ~
$ whoami
paras@platform:~$
Senior SRE · Platform Engineer
Kubernetes / Terraform / Python
Multi-cloud systems · Production reliability
ASK THE PLATFORM / LOCAL SIMULATION
01

Automate the operational surface

Replace recurring toil with repeatable, reviewable workflows.

02

Design for failure

Build recovery, scaling and workload identity into the platform.

03

Make infrastructure observable

Turn system signals into understanding and actionable response.

03 / PRODUCTION EXPERIENCE

Where reliability
meets reality.

From resolving complex failures
to engineering the platforms that prevent them.

01

Shell

Site Reliability Engineer

April 2026 — Present
PRODUCTION RELIABILITY

Engineering the operational foundation of cloud platforms.

  • Improved deployment efficiency by 40% through Kubernetes (EKS/AKS) scaling strategies.
  • Automated AWS and Azure provisioning with Terraform, reducing manual deployment effort by 60%.
  • Improved release frequency by 30% and reduced deployment failures by 35% with GitHub Actions and Azure DevOps.
  • Reduced MTTR by 25–40% using Prometheus, Grafana and Dynatrace; cut manual intervention by 50% with Python, Bash and PowerShell.
ENGINEERING ECOSYSTEM
AWS + AzureTerraformEKS / AKSGitHub ActionsPrometheusGrafanaDynatrace
02

Gore Mutual Insurance

Site Reliability Engineer

November 2022 — March 2026
GITOPS & MULTI-CLOUD

Turning infrastructure operations into repeatable, observable systems.

  • Cut release cycle time by 50% with GitHub Actions, Argo CD and GitOps across multi-cloud environments.
  • Reduced operational toil by 35% with Python-based self-healing workflows driven by Prometheus and Dynatrace signals.
  • Reduced cloud spend by 20% with FinOps rightsizing, autoscaling and cost observability.
  • Established observability across 50+ microservices; reduced troubleshooting time by 30%.
ENGINEERING ECOSYSTEM
GitHub ActionsArgo CDAWS + AzureTerraformAnsibleObservability
03

Kompusys Consultant Inc.

Platform Engineer / Technical Support Specialist

September 2020 — October 2022
PLATFORM FOUNDATIONS

Building consistency and resilience into cloud environments.

  • Reduced incidents and MTTR by 35% through database, authentication and performance troubleshooting.
  • Achieved 99.9% uptime with Kubernetes SQL/NoSQL workloads, backup and disaster recovery.
  • Reduced environment setup time by 50% using Terraform and Ansible; cut operational overhead by 30% with Python and Bash.
ENGINEERING ECOSYSTEM
AWS + AzureTerraformAnsibleKubernetesSQL / NoSQLPython
04

Apple Inc. (Kelly Services)

Technical Support Engineer

December 2019 — August 2020
SYSTEMS & TROUBLESHOOTING

Solving failures across the system, network and hardware layers.

  • Reduced downtime by 30% through enterprise Linux and macOS troubleshooting.
  • Supported Kubernetes SQL/NoSQL data platforms with backup, disaster recovery and 99.9% uptime.
  • Performed root cause analysis and collaborated with engineering and security teams on platform stability.
ENGINEERING ECOSYSTEM
Linux / macOSNetworkingKubernetesRoot cause analysis

04 / ENGINEERED, NOT JUST ENVISIONED

Systems with
substance.

Three projects. From declarative delivery
to AI-connected operations and FinOps ML.

01 FEATURED SYSTEM AI INFRASTRUCTURE

MCP PlatformCloud-native AI microservices on AWS EKS

An operator asks a question. A tool-calling agent reaches into the platform. Infrastructure becomes a conversation grounded in system results.

AWS EKSMCPTerraformArgo CDIRSAKustomize
10AI microservices
7Ordered sync waves
0Static AWS credentials
  • Zero-downtime deployment design
  • EKS production + microk8s development overlays
  • Gateway autoscaling: 2–10 replicas at 70% CPU
Explore the repository
MCP PLATFORM / SERVICE TOPOLOGY
Operator Intent
LLM Agent Claude · tool calling
MCP
MCP Control Plane IRSA identity
Cluster API
Payments API
Products API
AWS EKS Kustomize · HPA · OIDC
02 PLATFORM ENGINEERING

GitOps Platform Engineering

From commit to a consistent environment.

Declarative infrastructure and self-healing delivery, with reusable application patterns across development, staging and production.

View source
DELIVERY / DECLARATIVE BY DEFAULT
  1. 01Developer
  2. 02Git Commit
  3. 03GitHub
  4. 04CI
  5. 05Container Registry
  6. 06GitOps Repository
  7. 07Argo CD
  8. 08EKS
DEV STAGING PRODUCTION
Continuous reconciliation Desired state ↔ Cluster state
50%Better deployment consistency
40%Fewer manual releases
30%Faster service onboarding
35%Fewer deployment failures
03 FINOPS / MACHINE LEARNING

Cloud Cost Spike Detector

Make unexpected spend impossible to miss.

A time-series ML pipeline that surfaces abnormal cloud spend, estimates financial impact and brings service-level visibility to budget governance.

View source
CLOUD SPEND MONITORSYNTHETIC DATA / USD PER DAY
Observed spendExpected baseline
$330$230$125$20Day 1: $104Day 2: $108Day 3: $106Day 4: $112Day 5: $109Day 6: $110Day 7: $105Day 8: $112Day 9: $116Day 10: $111Day 11: $114Day 12: $110Day 13: $117Day 14: $113Day 15: $109Day 16: $116Day 17: $115Day 18: $112Day 19: $118Day 20: $114Day 21: $111Day 22: $115Day 23: $135Day 24: $175Day 25: $280Day 26: $340Day 27: $290Day 28: $250DAY 1DAY 7DAY 14DAY 21DAY 28
INSPECT DAY 26
$340
ANOMALY DETECTED

Example EC2

Expected baseline
$115/day
Peak observed
$340/day
Peak daily excess
$225

Interactive illustration only.
No live billing data or production savings claims.

Cloud Billing DataFeature PipelineIsolation ForestAnomaly EngineREST APIExecutive Dashboard
Burn-rate trackingService-level visibilityForecastingBudget governanceCI/CD integration

05 / THE NEXT LAYER

From cloud platforms
to AI infrastructure.

A specialization built on production fundamentals.
Extending the foundation, one system at a time.

01Cloud Infrastructure
02Platform Engineering
03Kubernetes
04Distributed Systems
05ML Infrastructure
06LLM Infrastructure
07Agentic Systems
SYSTEM BUILD / SCROLL TO EXECUTECONCEPTUAL WALKTHROUGH
01PROVISION02RECONCILE03CONNECT04OBSERVE

Start with a foundation.

Declare the infrastructure. Make every environment repeatable.

TerraformAWS / AzureCloud network

↓ CONTINUE SCROLLING TO BUILD THE NEXT LAYER

platform.tfREAD ONLY
resource "aws_eks_cluster" "platform" {  name = "ai-platform"  # Versioned infrastructure  # Reproducible environments}
$ terraform plan

Desired infrastructure defined

01 / PROVISION Declare the infrastructure. Make every environment repeatable.

02 / RECONCILE Git becomes the source of truth. Kubernetes reconciles the desired state.

03 / CONNECT Connect a tool-calling agent to platform services through MCP.

04 / OBSERVE Make signals actionable. Feed operational learning back into the platform.

INFRASTRUCTURE → PLATFORM → INTELLIGENCE → RELIABILITY01 / 04
PRODUCTION FOUNDATION

Operate with confidence.

Kubernetes, infrastructure as code, GitOps and observability form the operational base for dependable AI platforms.

AWS / AzureTerraformKubernetesSRE
PROJECT EXPERIENCE

Connect intelligence to systems.

MCP control planes, tool-calling agents and cloud-cost anomaly detection bring AI into concrete engineering workflows.

MCPTool callingIsolation ForestMLOps
SPECIALIZATION DIRECTION

Build the next infrastructure layer.

Deepening model serving, GPU orchestration, RAG, guardrails and evaluation. An evolving specialization grounded in cloud and platform engineering.

CUDAPyTorchTensorFlowLLMOpsLangChainAzure AI FoundryAWS Bedrock

07 / SYSTEM DESIGN LAB

Think in systems.
Follow the connections.

Four architectural lenses.
One principle: make the path explicit.

Kubernetes PlatformCONCEPTUAL DEMO
  1. 01Internet
  2. 02Load Balancer
  3. 03Ingress
  4. 04Services
  5. 05Pods
  6. 06Database

Route traffic through ingress and services to replicated pods. Keep state in a dedicated data tier.

Observe pods with Prometheus → Grafana. Route alerts to incident response.

apiVersion: apps/v1kind: Deploymentmetadata:  name: platform-api

08 / TECHNICAL TOOLKIT

Different tools.
One connected system.

Explore the layers of the stack.
Technologies listed in the resume; no proficiency scores.

PLATFORM / LAYER

Cloud

7 technologies
01AWS
02Azure
03GCP
04VPC
05DNS
06Load Balancing
07TCP/IP

VERIFIED LEARNING / CONTINUOUS PRACTICE

Foundations that compound.

AZ-104

Microsoft Certified Azure Administrator

AI-102

Microsoft Azure AI Engineer Associate

HCTA0-003

Terraform Associate

CKA

Certified Kubernetes Administrator

AZ-400

Designing and Implementing Microsoft DevOps Solutions

ML

Machine Learning Specialization

DL

Deep Learning Specialization

Education

GRADUATED 2023

Associate Degree in Computer Science and Programming

MIT (Massachusetts Institute of Technology), USA

GRADUATED 2019

Applied Science (AMM)

(Conestoga College), Canada

GRADUATED 2015

Bachelor of Mechanical Engineering

(RK University), India

09 / LET’S CONNECT

Let’s build systems
that scale.

Platform engineering. SRE. AI infrastructure.
Distributed systems and production GenAI platforms.