Data Engineer · AWS Data Engineer · ETL & Data Pipeline Specialist
Send a job offer directly to this candidate
Data Engineer with 3+ years of production experience designing, building, and optimising large-scale batch and streaming data pipelines using Python, PySpark, and SQL across BFSI and Telecom domains. Led a Databricks/Delta Lake POC migrating two legacy Hive pipelines - validated a 20% query speedup and zero-data-loss ACID upserts, adopted into the team's lakehouse migration roadmap. Cut SLA-breaching Spark job runtimes by 25% on a 100 TB+ Hadoop/HDFS platform, contributing $75K+/yr in infrastructure savings.
Built a Python data-quality framework (schema validation, null-rate thresholds, row-count reconciliation, alerting) that consistently catches upstream failures before they reach BI/reporting layers. Working exposure to Microsoft Azure alongside primary AWS expertise; version control and code review discipline via Git. Comfortable across the full SDLC in Agile teams, from requirements gathering with stakeholders through development, testing, and deployment.
Complementary hands-on exploration of GenAI/RAG pipeline design and applied AI engineering. AWS Cloud Practitioner certified.
Data Engineer - Kyndryl - Pune, India
(2023-08 - 2024-09)
Data Engineer / Backend Developer - Businessnext - Mumbai, India
(2024-10)
Clients: Bank of India & HDFC Sales
Data Science Intern - TCR Innovation - Mumbai, India
(2021-10 - 2022-03)
B.E. - Information Technology - Savitribai Phule Pune University (2019 - 2023)