Senior Data Engineer at DBHDS (2026-02 – Present)
Designed and implemented cloud-native data pipelines to ingest, process, and curate behavioral health and healthcare data from multiple source systems into a centralized enterprise data platform.
- Built scalable ETL/ELT workflows using AWS Glue, PySpark, Python, and SQL, enabling daily ingestion and transformation of structured and semi-structured healthcare datasets.
- Implemented Medallion Architecture (Bronze, Silver, Gold) on AWS to standardize data processing, improve data quality, and provide analytics-ready datasets for reporting and downstream applications.
- Developed automated ingestion frameworks for healthcare data exchanges, including HL7 and FHIR-based integrations, ensuring secure and reliable movement of patient and clinical information.
- Designed and optimized Iceberg/Delta Lake tables to support incremental processing, historical tracking, and high-performance analytical queries.
- Built data quality and validation frameworks using PySpark and SQL to perform schema validation, duplicate detection, reconciliation, and business rule enforcement before publishing data to curated layers.
- Leveraged AWS Lambda, EventBridge, and Step Functions to automate data ingestion workflows, orchestrate dependent jobs, and reduce manual operational effort.
- Developed secure API-based integration solutions to exchange healthcare and operational data with external agencies, partner systems, and internal applications.
- Optimized Spark workloads through partitioning, caching, broadcast joins, and query tuning, significantly improving processing performance and reducing cloud resource consumption.
- Collaborated with business analysts, healthcare stakeholders, and architects to gather requirements and translate complex business needs into scalable data engineering solutions.
- Implemented enterprise security controls using IAM, KMS, Lake Formation, and Secrets Manager, ensuring HIPAA-compliant handling of sensitive healthcare data.
- Built monitoring and alerting solutions using CloudWatch, SNS, and Glue job metrics, enabling proactive detection and resolution of production issues.
- Supported AI/ML and advanced analytics initiatives by creating curated datasets and feature-ready data structures for predictive modeling and healthcare insights.
- Participated in Agile ceremonies, code reviews, production support, and release planning while mentoring junior engineers and promoting engineering best practices.
- Created detailed technical documentation, source-to-target mappings, operational runbooks, and data lineage artifacts to support governance, audit, and knowledge transfer activities.
Senior Data Engineer at FM Global Insurance (2024-03 – 2025-12)
- Designed and implemented real-time streaming pipelines using Apache Kafka, Flink, and Spark Structured Streaming to process insurance policy, claims, and risk event data.
- Built Kafka producer/consumer services and ksqlDB/Flink transformations to ingest high-volume policy transactions and enrich underwriting data streams for downstream analytics.
- Architected Azure Databricks streaming pipelines consuming Kafka topics and persisting curated datasets in Azure Data Lake using Delta Lake (Bronze/Silver/Gold) architecture.
- Implemented CDC ingestion pipelines from SQL Server policy databases, publishing incremental changes to Kafka topics with Confluent Schema Registry for schema validation.
- Automated end-to-end workflows via Azure Data Factory and Databricks Jobs with full dependency management, retry logic, and error alerting.
- Embedded DataOps practices via Azure DevOps CI/CD pipelines automating deployment of notebooks, pipelines, and Terraform-managed infrastructure (Kafka connectors, Azure resources).
- Implemented data lineage and pipeline monitoring using Dynatrace and Azure Monitor, detecting anomalies and ensuring SLA compliance.
- Optimized streaming throughput via checkpointing, Kafka partition tuning, and consumer group scaling — reducing data latency from hours to minutes.
Senior Data Engineer at Mayo Clinic (2023-04 – 2024-02)
- Built scalable healthcare data pipelines on AWS (S3, Glue, EMR) to ingest and process large volumes of patient and clinical data from EHR systems in compliance with HIPAA standards.
- Developed PySpark transformation jobs on EMR to cleanse, normalize, and enrich patient encounter, diagnostic, and medical records data stored in Parquet format on S3.
- Implemented real-time healthcare event ingestion using Kafka and AWS Lambda, capturing hospital system events with low latency.
- Developed Apache Airflow DAGs orchestrating clinical data ingestion, transformation, and loading into Snowflake for analytics and patient outcome reporting.
- Built Snowflake data models supporting clinical reporting, patient outcomes analytics, and research data sets for medical research teams.
- Secured healthcare data with AWS IAM policies, KMS encryption, and role-based access controls meeting regulatory compliance requirements.
- Automated infrastructure provisioning using Terraform and CI/CD pipelines, reducing manual deployment effort and ensuring reproducible environments.
- Optimized Spark jobs on EMR by tuning memory allocation, partitioning strategies, and caching — significantly improving processing throughput for large clinical datasets.
Data Engineer at Express Scripts (Cigna) (2021-11 – 2023-03)
- Designed and implemented cloud-based enterprise data lake across AWS S3 and Google Cloud Storage to process large-scale healthcare and pharmacy claims datasets.
- Built scalable ETL pipelines using Apache Airflow, GCP Dataproc, and PySpark to transform structured/semi-structured healthcare data, reducing processing time by improving parallelism.
- Developed streaming ingestion pipelines using GCP Pub/Sub and Kafka to capture real-time prescription transactions and pharmacy claim events from operational systems.
- Migrated legacy Oracle and Netezza data warehouse workloads to BigQuery and Amazon Redshift, improving scalability and reducing query latency.
- Designed incremental CDC pipelines to capture transactional database changes and load them into cloud data lake environments for near-real-time analytics.
- Built Customer Data Platform (CDP) ingestion frameworks integrating transactional, behavioral, and marketing datasets supporting 360-degree customer analytics.
- Integrated Power BI dashboards with BigQuery and Redshift to enable real-time monitoring of prescription trends, pharmacy operations, and engagement metrics.
- Implemented security and governance controls using IAM, KMS encryption, and RBAC to ensure HIPAA compliance across cloud data environments.
Cloud Data Engineer at Aetna (CVS Health) (2020-10 – 2021-07)
Implemented and managed Enterprise Data Lake on AWS, supporting high-volume analytics, processing, and reporting use cases for healthcare insurance data. Developed fine-grained S3 access control security framework.