Data Engineer - Airbus - Bengaluru, KA
(2023-03)
- Spearheaded the ETL platform modernization (IBM InfoSphere DataStage 11.5 → 11.7), improving system stability and scalability, and delivering €120K–€385K in cost savings through infrastructure optimization and cross-functional collaboration.
- Designed and built scalable PySpark pipelines on Databricks processing ∼1TB of enterprise data, implementing complex transformations like multi-table joins, window functions, aggregations, resulting in reduced downstream processing latency for analytics consumers.
- Performed Data Modeling for Medallion architecture (Bronze/Silver/Gold) on Databricks Delta Lake, designing dimensional schema structures, partitioning strategy, and slowly changing dimension logic, while leveraging ACID transactions for data reliability to support scalable, query-optimized analytics for downstream consumers.
- Optimized Databricks cluster configurations, including executor memory, core allocation and shuffle partitioning, improving job throughput and reducing execution time by ∼20–30% for large-scale data pipelines.
- Leveraged Delta Lake's Change Data Feed (CDF) to implement incremental data loads, reducing full-scan overhead and improving pipeline efficiency.
- Played a key role in evolving Solace-based POCs into a production-grade event-driven streaming platform, enabling real-time data integration and enhancing system stability and scalability.
- Developed and implemented an automated ETL solution for password renewal processes, reducing support incidents and improving operational efficiency.
- Resolved L3 production issues related to ETL Pipelines and the Solace Platform, ensuring high availability and minimal downtime.
Data Engineer - Danske IT and Support Service - Bengaluru, KA
(2021-06 - 2023-02)
- Contributed to the FPDF squad migration by converting legacy Mainframe jobs into IBM InfoSphere DataStage workflows, implementing new logic using scripting, and ensuring data accuracy through rigorous validation with CPS systems and business analysts.
- Developed PySpark batch pipelines to process End-of-Month balance and income datasets at scale, applying multi-stage transformations, aggregations and referential integrity checks to ensure accurate downstream financial reporting.
- Implemented GDPR-compliant data anonymization using Pyspark, applying masking and pseudonymization transformations within the pipeline layer to ensure regulatory compliance across the Data Warehouse.
- Processed high-volume transactional datasets using PySpark broadcast joins to enrich customer records with reference data (account types, product codes), avoiding shuffle-heavy joins and reducing job execution time significantly.
- Worked with Snowflake to design and optimize SQL-based data pipelines and models, supporting efficient data storage and query performance for enterprise reporting.
- Managed code deployments via GitHub Actions-based CI/CD pipelines, enabling reliable and consistent release process.
Junior ETL Developer - Sunera Technologies (Client - TracFone Wireless) - Hyderabad, TL
(2020-05 - 2021-06)
- Developed and enhanced ETL workflows using IBM InfoSphere DataStage by creating new jobs and optimizing existing pipelines based on evolving business requirements.
- Optimized ETL jobs, reducing production processing time by over 400 hours annually and improving overall system performance.
- Implemented file automation processes to enable near real-time data ingestion for vendor data.
- Migrated and re-engineered PL/SQL procedures and Unix shell scripts into IBM InfoSphere DataStage jobs, improving maintainability and standardizing ETL processes.
ETL Developer - DXC Technology (Client - General Motors Financial) - Bengaluru, KA
(2018-03 - 2020-05)
- Contributed to Project Gold Rush, developing ODS, Business Data Warehouse, and Data Mart layers to support international financial operations.
- Designed and implemented ETL pipelines using IBM InfoSphere DataStage to ingest data from web services (JSON format), performing parsing and data transformations to load into enterprise data warehouse and downstream data marts.
- Led development efforts in the CRT IDI release, processing XML data, applying complex business transformations, and loading into reporting tables for analytics consumption.
- Performed unit testing, deployment, UAT support, ensuring high-quality and reliable data delivery.
- Created detailed design, mapping, and business rule documentation to support maintainability and future enhancements.