Data Engineer at UNIZEN TECHNOLOGIES PVT LMT (2023-05 – Present)
Project: E-Commerce Delta Analytics Platform (Wallmart)
- Developed scalable data pipelines using Python, PySpark, Azure Databricks, and ADF to process high-volume transactional data exceeding 15+ million records per day.
- Implemented a Medallion Architecture (Bronze, Silver, Gold layers) with Delta Lake to improve data reliability, consistency, and downstream reporting quality.
- Built ingestion workflows for structured and semi-structured data sources including JSON, CSV, and RDBMS using JDBC-based integrations and autoloader.
- Designed incremental and using last modified timestamp CDC-based ingestion mechanisms that significantly reduced end-to-end data processing delays.
- Applied complex business transformations, deduplication strategies, and validation checks using PySpark and SQL to ensure high data accuracy.
- Created dimensional data models with fact and dimension tables to support analytical reporting and improve query efficiency for BI workloads.
- Improved Spark job performance through partition optimization, caching techniques, and Z-order indexing, reducing execution time from nearly 3 hours to under 42 minutes.
- Collaborated with development teams using Git and GitHub for source control, enabling smooth code integration and controlled production deployments.
- Implemented monitoring, logging, and alerting mechanisms for pipeline health checks, helping reduce production failures and improve SLA compliance.
- Managed secure data governance and access control using Unity Catalog by organizing catalogs, schemas, and permission policies for business users and analysts.
Data Engineer at United health Group
Healthcare Management System — Health Insurance Analytics Platform
- Engineered ETL/ELT pipelines in PySpark on Azure Databricks to perform large-scale data transformations, aggregations, and cleansing across 10M+ records; reduced job runtime by 50% through cluster tuning and partition-aware execution.
- Architected a Bronze-Silver-Gold Medallion data lake on Azure Data Lake Storage (ADLS); ingested raw data from Oracle, SQL Server, and flat files into the Bronze layer and applied business transformation logic for analytics-ready Silver and Gold tables.
- Managed Delta Lake tables to support ACID-compliant incremental loads, schema evolution, and time-travel auditing; used OPTIMIZE and ZORDER commands to improve query performance by 40% on large partitioned datasets.
- Orchestrated end-to-end pipeline workflows using Azure Data Factory (ADF), scheduling Databricks notebook runs, monitoring pipeline SLAs, and configuring failure alerts to ensure data freshness and reliability.
- Implemented automated data quality checks and validation rules (null checks, schema enforcement, referential integrity) using PySpark and Spark SQL, ensuring 99%+ data accuracy across downstream analytics systems.
- Leveraged Spark SQL for partitioned data analysis, business metric computation (revenue KPIs, policy renewal rates), and query optimization using Partitioning, Bucketing, Caching, OPTIMIZE, ZORDER, Spark Joins.
- Applied performance optimization techniques including partitioning, bucketing, and caching across Hive and Delta tables to reduce shuffle overhead and improve pipeline throughput.
System Administrator at Sai Institute Polytechnic (2021-04 – 2023-04)
- Created and managed user accounts, reset passwords, and provided system access support for employees and students.
- Installed software, updated systems, and fixed hardware/network issues to ensure smooth day-to-day IT operations and minimal downtime.
- Monitored system performance and performed regular data backups to prevent data loss and maintain system reliability.