Data Engineer - Capital One - McLean, VA
(2024-11 - 2026-03)
- Architected and deployed highly scalable, production-grade data pipelines on AWS (S3, EMR) and Databricks, facilitating advanced analytics for 116M+ customer accounts.
- Managed the end-to-end data processing workflow, successfully handling and optimizing over 500M+ daily transaction records for large-scale financial datasets.
- Enhanced credit line decisioning by delivering targeted behavioral and transactional datasets, resulting in a 30% increase in campaign response rates.
- Implemented optimized Snowflake and Redshift data models that improved reporting access and elevated query efficiency by 20% for business teams.
- Improved pipeline stability through automated validation checks, comprehensive monitoring, and performance tuning across distributed Spark workloads.
- Collaborated closely with product, analytics, and engineering teams in Agile delivery environments to ensure timely release of production-ready data solutions.
Data Engineer (Databricks & AWS) - Cigna - Columbia, MD
(2023-06 - 2024-10)
- Transitioned legacy Hadoop environments to Databricks, which led to a 30% improvement in scalability and streamlined reporting for healthcare operations.
- Built automated workflows generating NDC11-level drug reports across 3,000+ healthcare plans and 150K+ NDC records daily.
- Optimized Spark processing for large-scale joins involving 450M+ daily data points across healthcare datasets.
- Developed ETL pipelines using AWS Glue and S3 to standardize ingestion and transformation of healthcare plan data.
- Implemented Databricks job scheduling and monitoring to improve workflow reliability and reduce manual intervention.
- Facilitated cross-functional testing sessions with business analysts and QA staff to identify discrepancies and uphold consistent data integrity in daily healthcare reports.
AWS Data Engineer - Comcast - Reston, VA
(2021-04 - 2023-06)
- Developed and maintained 100+ Airflow pipelines integrating data from APIs and distributed enterprise systems.
- Designed large-scale Spark and AWS-based data pipelines processing over 500 TB of structured and semi-structured data.
- Improved data accessibility by consolidating multiple source systems into centralized analytical datasets.
- Built automated monitoring and testing solutions that improved pipeline reliability and reduced operational issues and Reduced production incidents by 25%.
- Partnered with global engineering and analytics teams to deliver scalable data solutions for reporting and business insights.
- Engineered sentiment analysis initiatives using distributed processing and large-scale data transformation techniques.
Data Engineer - J.B. Hunt Transport Services - Lowell, AR
(2020-02 - 2021-03)
- Designed ETL pipelines supporting migration of operational data into enterprise data lakes and warehouses.
- Reduced backfill and rerun processing time by over 50% by redesigning production pipelines for idempotent execution.
- Developed SQL-based staging and transformation workflows integrating data from 10+ source systems.
- Investigated and resolved ETL failures, data quality issues, and warehouse processing anomalies.
- Supported large-scale data migration initiatives in collaboration with QA analysts and business stakeholders.
Data Engineer - Lowe's - Mooresville, NC
(2019-01 - 2019-12)
- Built ETL pipelines to process application log data and deliver datasets for recommendation and analytics systems.
- Automated generation of 140+ operational reports, reducing reporting turnaround time by 60%.
- Improved data quality and accessibility through cleansing and transformation of structured and unstructured datasets.
- Collaborated with analysts and database teams to support reporting and warehouse development initiatives.