DevOps Engineer at Hind Social Network Pvt Ltd (2026-03 – Present)
Multi-tenant white-labeled community SaaS - 35 production communities on custom domains, on Kubernetes. Carry production on-call; work with engineering and product on release, reliability and cost.
- Built and operate GitOps delivery on GitHub Actions, SonarQube, GHCR and ArgoCD, replacing a mixed VM/Compose/Kubernetes release process; ~5 production deploys/week at 2–4 minute rolling updates, one rollback across ~130 production deploys.
- Cut monthly infrastructure spend ~70% by consolidating three cloud providers to one, trimming the server fleet on historical utilization, and correcting Kubernetes requests set 30–80x above actual usage. No availability impact.
- Delivered a zero-downtime migration of production and all white-labeled tenants to an India-hosted cloud for DPDP data residency; used Cloudflare weighted load balancer, rehearsed on stage, per-tier rollback targets.
- Diagnosed a 16-hour silent delivery outage: an immutable selector change had ArgoCD rejecting every apply while apps reported Healthy. Designed a side-by-side rename across 7 charts, added sync/health alerting.
- Built the observability stack on Prometheus (180d), Grafana, Loki, Tempo, OpenTelemetry and Alertmanager, all GitOps-managed, with SLOs from span metrics and 36 burn-rate alert rules as code. Caught tail sampling inflating measured error rate ~10x and firing false SLO pages; a sampling-baseline recording rule took false alerts to zero. 99.90% availability at p95 298ms.
- Operated a mixed GPU/non-GPU Kubernetes cluster for production AI/ML workloads (RAG, vLLM, Whisper, Kokoro TTS), using node affinity and taints to isolate GPU scheduling from general capacity.
Career Break - Independent Upskilling & Freelance DevOps at Self-directed Upskilling & Freelance DevOps (2025-08 – 2026-02)
Self-directed Kubernetes, Terraform and GitOps practice, plus freelance infrastructure work for an early-stage SaaS client. 20 services provisioned on Terraform/EKS/IAM, release cycle cut from ~4 hours to 30 minutes, Prometheus/Grafana monitoring established.
DevOps Engineer at Vantage Circle (Bargain Technologies Pvt. Ltd.) (2022-08 – 2025-07)
Three years on infrastructure and reliability for a multi-tenant B2B HRTech SaaS. 2M+ users, 700+ enterprise clients, 99.95% availability on EC2, RDS MySQL, Redis, S3 and CloudFront.
- Provisioned and owned four environments with Terraform and Ansible. UAT, stage, production, and an on-demand AI/ML environment stood up and torn down per run, replacing manual server setup.
- Owned CI/CD pipeline engineering and environment promotion the wider team shipped through; PR-gated merges and SonarQube quality gates cut release cycles from 8 hours to 40 minutes.
- Root-caused a platform-wide cascading failure from single-tenant Redis exhaustion; wrote the postmortem and drove batch-processing isolation to contain single-tenant blast radius.
- Cut API p95 latency from ~10s to under 1s via APM traces and slow-query analysis, isolating a missing MySQL index; throughput improved ~40%.
- Cut incident triage from 6–7 hours to minutes, tuning 30+ alerts to Golden Signals across Slack, email and PagerDuty, reducing overall MTTR ~40%.
- Built the performance engineering practice from scratch: Terraform-provisioned test infrastructure, a JMeter suite, and an InfluxDB + Grafana results pipeline, with a Python runner executing tests and reporting failures by email.
- Led an end-to-end EKS pilot on AWS credits: provisioned the cluster, ran production-representative workloads, and presented a cost and headcount analysis to leadership that informed the containerization decision.