Cloud and Platform Infrastructure Engineer
Send a job offer directly to this candidate
Cloud and Platform Infrastructure Engineer with 11+ years of experience designing, automating, and operating secure, scalable hybrid cloud environments across AWS, Azure, and GCP. Currently architecting production Kubernetes platforms at Topgolf spanning EKS, Rancher, and on-prem vSphere, with deep expertise in Terraform-based Infrastructure as Code, multi-environment provisioning, and reusable module design. Skilled in CI/CD automation using Jenkins and GitHub Actions, secrets management via HashiCorp Vault and IAM, and full-stack observability with Datadog, Prometheus, and Splunk.
Committed to reducing operational toil through automation and delivering reliable, cost-optimized platform foundations that enable development teams to ship faster and with greater confidence.
Architect at Topgolf | 2025-04 - Present | Austin, TX: Supported Topgolf's hybrid platform environment across AWS, EKS, Rancher-managed Kubernetes clusters, on-prem vCenter/vSphere infrastructure, Linux/Windows servers, Docker-based venue systems, and production venue platforms.: Managed production and non-production Kubernetes platforms hosting microservices, game-launcher workloads, venue services, observability agents, database services, registry components, platform automation tools, and containerized applications used across Topgolf venues. Performed advanced Kubernetes operations including pod troubleshooting, Deployment/StatefulSet/DaemonSet management, DaemonSet validation, Service/Ingress troubleshooting, namespace analysis, ConfigMap/Secret review, Persistent Volume validation, readiness/liveness probe analysis, event review, workload scaling, rolling restarts, and recovery of failed workloads.
Kubernetes node and scheduling issues including NotReady nodes, scheduling-disabled nodes, kubelet/Docker runtime failures, pod eviction behavior, terminating pods, image pull failures, node resource pressure, disk utilization, network connectivity, taints, tolerations, selectors, and workload placement problems.
Rancher-managed Kubernetes clusters by validating cluster health, troubleshooting cluster-agent connectivity, recovering unavailable clusters, reviewing stuck upgrades, validating node status, and coordinating recovery during infrastructure, power, or venue-impacting incidents.