Data engineer
Send a job offer directly to this candidate
Data Engineer with 4+ years of experience in designing and developing data pipelines and ETL/ELT workflows using Azure Databricks, PySpark, SQL, Python, Azure Data Factory (ADF), ADLS Gen2, and Delta Lake. Hands-on experience in building Medallion Architecture (Bronze, Silver, Gold), incremental data processing, data quality validation, and Spark/SQL performance optimization. Experienced in processing structured and semi-structured data from MySQL, APIs, and SFTP sources and implementing reliable batch data pipelines.
Familiar with Unity Catalog for data governance, Databricks Workflows for orchestration, and Azure Monitor/Log Analytics for monitoring and troubleshooting. Also experienced with AWS data engineering using AWS Glue, Amazon S3, PySpark, Amazon Redshift, and AWS Step Functions. Strong in Python, SQL, Apache Spark, ETL, data transformation and troubleshooting.
Designed and developed an Azure-based data lakehouse solution using ADLS Gen2, Azure Databricks, Delta Lake, and Azure Data Factory.
Bronze, Silver, and Gold data layers to process structured and semi-structured data from MySQL, Google APIs, and SFTP sources.
PySpark and Spark SQL transformations for cleansing, validation, joins, aggregations, and incremental processing. Implemented data quality checks and used Unity Catalog for data governance and access control. Orchestrated data workflows using Databricks Workflows and monitored pipeline execution using Azure Monitor and Log Analytics. Published curated Gold-layer Delta data through Databricks SQL Warehouse for analytical consumption.