Data Engineer at Apetito (2025-10 – Present)
Industry: Healthcare Food Services
- Engineered enterprise analytics platforms using Medallion Architecture across Azure Databricks, Azure Synapse Analytics and Fabric Lakehouse, delivering datasets for downstream analytical reporting, with storage on ADLS Gen2, Onelake and Dedicated SQLPool.
- Orchestrated workflows in Synapse Workspace using dynamic linked services, datasets, and parameterized pipelines to extract and transform data from on-prem SQL Server, Salesforce, SAP HANA, and Oracle via SHIR, achieving 98% SLA compliance for Enterprise Data Warehouse loads.
- Designed snowflake schema-based EDW models in Fabric Power BI, implemented metadata driven ingestion pipelines with schema awareness, and orchestrated scalable PySpark transformation workflows through Databricks Unity Catalog and Lakeflow Declarative Pipelines.
- Optimized centralized SQL Pool performance through targeted statistics management, Index maintenance, advanced query tuning, and routine rebuild and reorganize operations, reducing scan overhead by 30%.
- Implemented Delta Lake optimized processing with Z Order, Optimize, Vacuum, schema evolution handling, and SCD Type patterns to deliver high performance, analytics ready Delta tables in ADLS.
- Identified and eliminated obsolete source from EDW by analyzing performance metrics CPU, DWU, and storage utilization in Azure Monitor, reducing unnecessary scans and optimizing query execution, resulting in a 35% reduction in monthly Synapse SQL pool costs.
- Delivered reliable releases by implementing CI/CD pipelines and GitHub Actions, leveraging centralized Git repositories, structured branching, and controlled code promotion workflows while applying strong management principles to safely implement code enhancements, reducing deployment errors by 80%.
- Supported AI/ML platform operations across Azure Foundry primarily partnering with Platform Support teams to troubleshoot platform issues, stabilize environments & ensure reliable data engineering workflows.
- Worked closely with business teams and stakeholders to share clear data insights and ensure decisions were aligned with validated analysis and business goals.
Data Engineer at Accenture (2021-01 – 2025-07)
Industry: Banking (Leading UK based Bank)
- Led end-to-end Hadoop migration into Azure by building ADLS Gen2 ingestion frameworks, Databricks-based transformation Lakeflow pipelines powering large-scale banking analytics and regulatory workloads.
- Engineered scalable Synapse orchestration pipelines to ingest and move raw and curated datasets across Datalake and Dedicated SQL Pools, leveraging Mapping Data Flows while applying cleansing, schema harmonization, late-arrival handling, partitioning, and storage optimizations.
- Built transformation frameworks in Azure Databricks to convert curated Delta Lake datasets into business ready layers, applying banking logic for fraud detection, credit risk scoring, and regulatory compliance.
- Designed metadata-driven Medallion Architecture workflows to manage batch ingestion and structured data movement, ensuring scalable and reliable pipeline execution across enterprise domains.
- Created and maintained SSRS reports, Power BI dashboards, and Paginated Reports, implementing row level security and applying performance optimization techniques to deliver actionable business insights.
- Optimized Databricks cluster performance by tuning autoscaling, Photon execution, shuffle behavior, and Delta Lake storage layout improving throughput and reducing compute cost across migration releases.
- Leveraged GitHub Actions and Azure DevOps CI/CD pipelines to automate deployment of Databricks notebooks, ADF/Synapse pipelines, and related artifacts, ensuring consistent and governed releases across dev, QA, and prod environments.
- Collaborated with product, analytics, engineering, and DevOps teams to design resilient data engineering architectures, resolve ingestion and transformation bottlenecks, and ensure reliable delivery of high-quality banking datasets.
Data Engineer at Accenture (2021-01 – 2025-07)
Industry: Sustainability Client: Count Us In
- Coordinated cross-functional data intake by working with partner teams to collect data files and with stakeholders to validate business rules, data requirements, and ingestion expectations, improving requirement clarity and reducing ETL defects by 30%.
- Architected end-to-end data platform components by integrating partner-sourced datasets into a centralized Amazon Redshift warehouse and designing scalable conceptual, logical, and physical data models, reducing downstream SQL workload issues by 40%.
- Developed automated ETL pipelines using AWS Glue and PySpark to process multi-layer data flows (raw → curated), improving data freshness across reporting layers and reducing daily batch processing time by 50%.
- Implemented event-driven ingestion using AWS Lambda and S3 triggers to detect new files, activate ETL pipelines, and publish SNS alerts, eliminating manual ingestion steps and reducing pipeline activation latency by 90%.
- Enabled automated metadata and schema management through Glue Crawlers and the AWS Data Catalog, ensuring accurate schema definitions and reducing manual schema maintenance by 70%.
- Delivered KPI driven Tableau dashboards by aligning SQL based data models with visualization needs, improving leadership's decision-making speed by 25% through interactive reporting.
- Supported full SDLC workflows by translating business requirements into technical designs, mapping end-to-end data movement, and ensuring alignment across design, development, and delivery, reducing cross-team friction by 35%.
Junior Data Analyst at Raja Spares Pvt Ltd (2019-05 – 2020-12)
Industry: Retail
- Supported requirement-gathering discussions with business teams by clarifying data needs, identifying required fields, and validating logic. Wrote SQL server queries to pull accurate datasets that helped stakeholders make operational and strategic decisions.
- Created SQL-based reports and summaries to highlight sales trends, fast-moving items, and inventory patterns. Structured queries using grouping, joins, and window functions to produce clear insights that guided 15–20% better inventory planning.
- Maintained detailed documentation including ER diagrams, source-to-target mappings, and workflow notes. Ensured documentation stayed updated so analysts & engineers could easily understand data structures and lineage.
- Performed data checks in SQL Server to validate sales and inventory accuracy. Used joins, counts, and conditional checks to identify missing, duplicate, improving reliability of downstream reporting.
- Responded to business requests by writing ad-hoc SQL queries for audits, reconciliations, and operational checks. Ensured quick turnaround and accurate extraction of required datasets.