AI Infrastructure & Reliability Engineer - Boeing Defense, Space & Security - Seattle, WA
(2026-06)
Developed and maintained data pipelines and ontology models using Palantir Foundry and AIP for AI-driven applications.
- Tracked ongoing issues to investigate bottlenecks and enhance platform stability.
- Collaborated with cross-functional teams to drive operational excellence and resolve challenges.
Lead - Site Reliability Engineer - Expedia Group - Seattle, United States
(2025-03 - 2026-01)
- Incident Management: Actively led high-priority incident resolutions in war rooms, ensuring rapid triage while mentoring newer engineers on troubleshooting. Maintained peak readiness by keeping the team calibrated to live-site performance.
- In-House Primary Monitoring Tool Development: Identified critical monitoring gaps, and orchestrated the end-to-end development of an enterprise-scale in-house primary monitoring tool. By leading cross-functional POCs with PagerDuty, BigPanda, Datadog, Dell, Splunk, FireHydrant, etc., defined requirements that replaced third-party limitations with a proprietary AIOps system capable of automated incident correlation and real-time revenue impact analysis.
- Automation & Correlation: In collaboration with the observability team, I led the optimization of our internal monitoring tool to drive enhanced incident correlation and automation. By establishing a rigorous testing framework and user feedback loop, we successfully integrated high-impact features, including AI-generated summaries, automated bridge creation, suspected root-cause identification, similar incidents, change-related incident tracking, and automated reporting features from the console
- Gap Remediation Framework: Designed and led the implementation of a global 'Gap Process' for 24/7 NOC pods in partnership with the Problem Management team, establishing accountability standards, and ensuring timely corrective actions across Service Owners, enabling data-driven identification of bottlenecks and SLA breaches.
- Financial Impact Analysis: During major site events, I performed revenue impact analysis, assisting in translating complex failures into financial insights for audits and high-loss incidents.
- Team Formation & Management: Built a high-performance NOC Tier 1 team from the ground up, overseeing the full lifecycle of hiring, onboarding, and technical training. I applied pedagogical strategies to accelerate 'time to autonomy,' transforming the unit into a robust first-line response. Designed a sustainable 24/7 global coverage model with co-pods and utilizing empathetic one-on-one mentorship, I successfully maintained high retention and ensured the team remained calibrated to live-site per
- AI/ML Literacy Program: Proactively designed and delivered a 'Fundamentals of Data Science and AI Ethics' curriculum to introduce core concepts, terminology, and internal AI tools. The program demystified AI capabilities and shifted the team's mindset from fear of replacement to a culture of augmentation, driving higher operational efficiency, and aligning the organization with company-wide AI adoption goals. Establishing guardrails and auditing standards for the integration of AI in operational
Site Reliability Engineer II - Expedia Group - Seattle, United States
(2024-04 - 2025-02)
Two-time Travel Excellence Award recipient for contributions in the Event Management space.
- Incident Management: Triage and incident management in war rooms, ensuring rapid resolution while mentoring newer NOC engineers on troubleshooting.
- Event Analysis: Maintained peak readiness by conducting booking trend monitoring and real-time event analysis for production readiness. Splunk, Datadog, Catchpoint, Grafana, Graphite, and PagerDuty, etc.
- Event Management: Lead event management strategies, improve correlation models in the primary monitoring tool to reduce alert noise, and accelerate root cause identification.
Site Reliability Engineer I - Expedia Group - Seattle, United States
(2021-03 - 2024-05)
- Data Analysis: Analyze large-scale telemetry and logs, building optimized queries to detect anomalies, trends, and early indicators of failure. Tested new releases, code checks with engineers, define and evaluate capabilities for enhanced event correlation models for the external event management tool to reduce noise, improve signal accuracy, and accelerate issue identification.
- Incident Management: Lead and support incident response by rapidly assessing impact, identifying root causes, coordinating responders, service owners, releases and driving resolution to minimize MTTR.
Operations and Traffic Analyst - IOTA - Expedia Group - Seattle, United States
(2017-08 - 2019-09)
- Monitoring: Proactively monitor and detect customer-impacting issues using observability tooling (Splunk, Catchpoint, Grafana/Graphite, PagerDuty, AWS).
- Triage: Partner with service owners, engineering teams, and incident managers to restore reliability and reduce MTTR.
- Release Management: A six-month tenure with Release Management, filling in for a personal absence of the release officer, overseeing the releases in test and production, monitoring the environment, impact analysis, and rollbacks as needed.