Senior Platform Engineer / Senior Reliability Engineer - Anyscale - San Francisco, CA · Remote
(2023-10 - 2026-09)
- Designed and evolved a self-service platform for distributed AI workloads, combining Go, Kubernetes, GCP, and reliability patterns to help product and engineering teams operate production services independently.
- Built Go-based infrastructure APIs that provisioned and reconciled Kubernetes resources, applying controller-style patterns to standardize service onboarding and lifecycle management.
- Developed reusable Kubernetes deployment templates and Helm conventions for Ray-based compute services, giving application teams consistent defaults for resource sizing, health checks, and rollout behavior.
- Partnered with internal customers to assess platform requirements, guide cloud-native application designs, and turn recurring operational needs into reusable platform capabilities.
- Implemented workload orchestration patterns with Temporal and Go for long-running infrastructure actions, improving visibility and recoverability for asynchronous provisioning workflows.
- Strengthened highly available GCP service operations through Datadog dashboards, monitors, and SLO-oriented alerting, reducing mean time to diagnose recurring production issues by 30%.
- Automated infrastructure delivery through Terraform and GitOps workflows, shortening environment provisioning from hours to under 30 minutes for common platform requests.
- Improved Linux and Kubernetes operational readiness by tuning container resource policies, documenting runbooks, and participating in an on-call rotation for business-critical alerts.
- Explored AI-assisted incident triage workflows that summarized alerts, linked relevant runbooks, and accelerated investigation context for on-call engineers.
- Led architecture reviews and Scrum planning for cross-functional platform roadmap work, mentoring 4 engineers on Go services, Kubernetes APIs, and reliability-focused design decisions.
Senior DevOps Engineer / Cloud Platform Engineer - Shipwell - Austin, TX · Remote
(2022-01 - 2023-09)
- Expanded the cloud platform supporting Shipwell's logistics applications with Kubernetes, AWS, Terraform, and observability standards, enabling reliable delivery of shipment-management services.
- Created reusable CI/CD pipelines with GitHub Actions and Helm, reducing deployment time from approximately 90 minutes to 25 minutes across platform-supported services.
- Modernized Kubernetes application delivery through Argo CD and GitOps practices, giving development teams auditable, repeatable promotion paths across environments.
- Built platform automation in Go and Python for provisioning, configuration validation, and operational tasks used by engineering teams supporting carrier and transportation workflows.
- Implemented Datadog monitors, dashboards, and service-level indicators for API and worker workloads, helping maintain 99.95% availability for critical logistics services.
- Evaluated LLM-assisted documentation and incident-summary workflows to improve access to operational knowledge for on-call responders.
- Hardened cloud access through IAM least-privilege controls, Kubernetes RBAC, and centralized secrets handling, partnering with security stakeholders on remediation work.
- Facilitated Agile/Scrum ceremonies and platform roadmap discussions with product and application teams, while mentoring engineers on incident response and cloud-native delivery practices.
Infrastructure Engineer - Shipwell - Austin, TX
(2019-07 - 2022-01)
- Built and operated infrastructure for Shipwell's transportation-management platform using AWS, Kubernetes, Terraform, and Docker to support growing logistics application workloads.
- Automated AWS account, networking, and compute provisioning with Terraform modules, reducing manual environment setup effort by roughly 60%.
Infrastructure Engineer - HashiCorp - San Francisco, CA
(2016-07 - 2019-06)
- Supported cloud infrastructure and developer workflows using Terraform, Vault, Consul, and Linux systems for teams adopting HashiCorp's infrastructure automation products.
- Developed automation for repeatable infrastructure provisioning and configuration workflows, improving consistency across development and test environments.