SENIOR DATA ENGINEER
Send a job offer directly to this candidate
Results-driven Data Engineer with around 9+ years of experience building scalable and automated data ecosystems across healthcare, retail, and finance. Expertise in AWS, ETL design, PySpark, and orchestration tools like Airflow. Skilled in data modeling, data governance, and performance tuning, delivering secure, optimized data pipelines and actionable business insights.
Developed and optimized PySpark-based transformations in Palantir Foundry, handling incremental loads, schema evolution, and large-scale joins across distributed datasets. Improved data reliability and performance for mission-critical reporting and decision-making. Built and maintained end-to-end data pipelines in Palantir Foundry using code repositories and pipeline orchestration, enabling automated data refreshes, strong lineage, and governed access to trusted datasets for analytics and business users.
Extensively worked on AWS Cloud services like EC2, VPC, IAM, RDS, ELB, EMR, EKS, Auto-scaling, S3, CloudFront, Glacier, Elastic Beanstalk, Lambda, Elastic Cache, Route53, OpsWorks, CloudWatch, CloudFormation, Redshift, DynamoDB, SNS, SQS, SES, Kinesis, Firehose, Cognito IAM. Hands-on experience in Unified Data Analytics with Databricks, Databricks Workspace Interface, Managing Databricks Notebooks, Delta Lake with Python, and Delta Lake with Spark SQL. Expertise in using Hadoop infrastructures such as MapReduce, Pig, Hive, Zookeeper, Sqoop, Oozie, Flume, Drill, and Spark for data storage and analysis, and programming with Scala, Python, SQL, and NoSQL databases such as HBase, MongoDB, Cassandra.
Efficient in writing Infrastructure as Code (IaC) in Terraform, Azure Resource Management (ARM), and AWS CloudFormation. Created reusable Terraform modules in both Azure and AWS Cloud platforms. Built CDC pipelines using AWS DMS and streaming frameworks, implementing watermarking, bookmarks/checkpointing, and idempotent MERGE upserts.
Modeled data warehouses in Amazon Redshift and Snowflake using star/snowflake schemas, SCD Type 1/2 dimensions, conformed facts, and surrogate keys to eliminate filter ambiguity. Tuned performance through predicate pushdown, column pruning, broadcast joins, skew mitigation, and repartition/coalesce strategies targeting 128–256 MB file sizes. Orchestrated workflows with Apache Airflow using task groups, retries/backoff, dataset-aware scheduling, backfills, and SLA-backed alerting to Slack/Email.
Enforced data quality and contracts through pandas/dbt tests, row counts, hash totals, referential integrity checks, schema drift monitors, and DLQ + replay paths. Hardened security and compliance with IAM least-privilege, KMS encryption, Secrets Manager, Lake Formation LF-tags, PII/PHI masking, and audit trail enforcement. Automated delivery pipelines via CI/CD using CodePipeline, CodeBuild, and Infrastructure as Code (Terraform/CloudFormation) with blue/green deployments.
Delivered measurable outcomes: faster refresh windows, stable SLAs, reduced computing/storage c
Senior Data Engineer - PG&E - Lehi Utah
(2026-03)
Designed and developed scalable data pipelines using Snowflake, Informatica IDMC, Airflow, and UC4 to support enterprise utility data processing and reporting needs.
Senior Data Engineer - MSC Industrial Supply - Melville, NY
(2023-08 - 2026-01)
Built and maintained end-to-end data pipelines in Palantir Foundry, transforming raw source data into reliable, analytics-ready datasets. Focused on creating seamless data flows that served as the backbone for business intelligence, analytics, and reporting teams.