Lead Data Platform & Reliability Engineer at Disney Streaming Services (2020-09 – 2026-06)
Built and operated large-scale infrastructure, observability, and data platforms supporting Disney+ production environments across hybrid cloud and on-prem infrastructure.
- Built and operated observability and data ingestion platforms supporting thousands of hosts across hybrid cloud and on-prem environments.
- Worked with AWS services including EC2, S3, RDS, and MWAA, using foundational Terraform skills to support infrastructure provisioning and configuration.
- Designed and operated a ClickHouse platform ingesting more than 100TB of log data daily to support real-time CDN decisioning and delivery-cost optimization.
- Built internal production analytics and telemetry infrastructure that reduced reliance on third-party systems and kept operational data within Disney infrastructure.
- Migrated large-scale logging infrastructure from Elasticsearch to Splunk by benchmarking workloads, designing the replacement environment, and deploying new infrastructure, reducing hardware requirements by 50% while doubling indexing performance.
- Built high-volume log, metrics, and event pipelines using Kafka, ClickHouse, Elasticsearch, Prometheus, and InfluxDB.
- Deployed Cribl as a data-routing and transformation layer for high-volume telemetry, enabling filtering, transformation, cost optimization, and delivery to multiple destinations.
- Automated infrastructure and operational workflows and supported engineers developing internal tooling for infrastructure automation, stream processing, and batch systems.
- Troubleshot production infrastructure, capacity, performance, and reliability issues across distributed systems and data platforms.
- Worked across engineering teams to plan and implement infrastructure changes while maintaining reliability of production streaming services.
Infrastructure & Platform Engineering Lead at Major League Baseball Advanced Media (MLBAM) (2015-05 – 2020-09)
Built and supported infrastructure, networking, monitoring, and automation systems for high-availability broadcast and live-event environments across MLB and NHL operations.
- Operated network and Wi-Fi infrastructure across 10+ MLB stadiums and NHL environments supporting broadcast connectivity and live-event operations.
- Built and maintained centralized monitoring infrastructure using Nagios XI for systems supporting all 30 MLB teams and NHL environments.
- Introduced PRTG to expand network monitoring, visibility, and performance troubleshooting across production infrastructure.
- Migrated Splunk workloads to Elasticsearch by benchmarking performance, redesigning infrastructure, and reducing hardware requirements.
- Built tooling for infrastructure automation, stream processing, and batch systems.
- Monitored and troubleshot production systems and network infrastructure during live broadcasts and high-traffic events.
- Coordinated infrastructure maintenance and production changes across teams while minimizing disruption to live services.
- Supported incident response, alerting, capacity planning, and performance troubleshooting across infrastructure and network environments.