🖊 Principal Engineer → Senior SRE · Red Hat · OpenShift

Hi, I'm Ayush 👋

Senior Site Reliability Engineer at Red Hat, keeping OpenShift and Kubernetes clusters alive at scale. Previously a Principal Engineer on the OpenShift Networking team — years inside OVN-Kubernetes, Service Mesh, DNS, and Ingress routing.

I write hands-on guides and make YouTube videos because a good explanation saves hours of debugging.

SRECurrent Role
30+Public Repos
4Live Guides
UTC+5:30Timezone

Who I Am

An engineer who went deep into Kubernetes networking and came out the other side building reliability into the platforms that run it.

🏢 Now — Senior SRE @ Red Hat

Platform reliability, incident response, scalability, and making sure production OpenShift clusters don't fall over under real load. Working across cloud providers, bare-metal environments, and hybrid deployments.

🔙 Previously — Principal Engineer, OpenShift Networking

Years inside OVN-Kubernetes, Service Mesh 3, DNS Operator, and Ingress internals. Built detailed technical guides, debugged packet paths in production, and contributed to how OpenShift networking works today.
Career Arc
OpenShift NetworkingPrincipal Engineer
Platform ReliabilitySenior SRE
Content CreationYouTube · Blog

📍 Red Hat — Building What Runs the Cloud

OpenShift runs critical workloads for thousands of enterprises. Working here means debugging real networking issues in production, understanding the full stack from kernel to API server, and building reliability tooling that entire teams depend on daily.

What I Work On

From cloud provisioning down to the packet level — and back up through pipelines and GitOps.

☁ Cloud & Infrastructure AWS Oracle Cloud Terraform Bare Metal UPI 🧱 Platform Layer OpenShift 4 Kubernetes SRE / Reliability Observability 🔌 Networking Layer OVN-Kubernetes Service Mesh 3 DNS / Ingress TLS / Security 🔄 GitOps & CI/CD ArgoCD Tekton Konflux PaC
Scripting & Diagnostics
Shell / Bash
Python
must-gather / oc adm inspect
tcpdump / ovn-trace
Jupyter Notebooks
Docker / Podman
Currently Exploring
LLMs & AI Tooling
Agentic Workflows
Claude API

Projects & Published Guides

Hands-on repos and live guides built from real production experience — not toy examples.

Published Guides — Live on GitHub Pages

OpenShift 4 DNS Architecture

Deep dive into CoreDNS, the DNS Operator, node resolvers, and production troubleshooting. Explains what actually happens when a pod resolves a service name.
→ GitHub repo  ·  → Read the guide

OpenShift 4 Architecture

Control plane, worker nodes, networking, storage, security, and observability stack — with installation methods and architecture diagrams.
→ GitHub repo  ·  → Read the guide

OpenShift Service Mesh 3

Architecture, installation, traffic management, security, observability, and troubleshooting for OSSM 3 built on Istio Sail Operator.
→ GitHub repo  ·  → Read the guide

Tekton Field Guide

Practical guide for Tekton pipelines on OpenShift — task authoring, real-world patterns, and the gotchas you only learn the hard way.
→ GitHub repo  ·  → Read the guide  ·  → Watch playlist

Konflux Pipeline Blueprints

Hands-on Konflux CI guide on OpenShift — build, sign, test, release, SLSA provenance, and Enterprise Contract with production-ready YAML blueprints.
→ GitHub repo  ·  → Read the guide  ·  → Watch playlist

More Repos

OpenShift4-AWS-BareMetal-UPI

End-to-end procedure to deploy OCP 4 BareMetal UPI on AWS — every command, every config file, from infra to cluster.

OpenShift AWS UPI

must-gather-analysis

Shell scripts to verify OpenShift 4 cluster health from must-gather bundles. Spot issues fast without reading thousands of log lines manually.

Shell Diagnostics

Secured Gateway — Service Mesh

Deploy a secured gateway for applications using OpenShift Service Mesh — mTLS, certificate management, and route encryption end to end.

Security Service Mesh

collect-tcpdump-ocp

DaemonSet to capture tcpdump on every interface across all cluster nodes simultaneously — the right tool for production network debugging.

Diagnostics Networking

openshift-mcp-server

MCP server exposing 216 tools, 7 resources & 10 SRE runbook prompts for OpenShift 4 cluster management via any MCP-compatible LLM across 20 operational domains.

MCP AI OpenShift

🔗 All repos on GitHub

Personal: github.com/ay-garg  ·  Work (Red Hat): github.com/aygarg-rh

Art of Exploitation

A YouTube channel about understanding systems at the packet and syscall level — not just how to use tools, but why they behave the way they do.

▶ Art of Exploitation

OpenShift internals, networking deep dives, live packet captures, and hands-on labs. No fluff — real cluster behavior explained from first principles. The kind of content I wish had existed when I was debugging at 2am.

▶ Watch on YouTube

🎯 What You'll Find on the Channel

OVN-Kubernetes trace walkthroughs, OpenShift DNS deep dives, live packet captures, Service Mesh debugging, and the kind of root-cause analysis that doesn't fit in a Stack Overflow answer.

✍ Blog — blackhatinside.com

Written guides on the same topics — when a concept needs more depth than a video can give. Think of it as the technical notes that go alongside the recordings.

→ blackhatinside.com

Let's Connect

Find me wherever engineers hang out.

💡 If something in any repo saved you time

Drop a ⭐ — it takes 2 seconds and helps others find it. That's the only metric that matters here.