Sr. Data Engineer - Vanguard - Malvern PA
(2023-12)
Led architecture and development of enterprise fraud modernization initiative, delivering scalable, event-driven data pipelines and real-time detection dashboards.
- Architected and implemented an end-to-end fraud detection pipeline using AWS Glue, Kinesis, S3, and CloudFormation, processing over millions of transactions daily.
- Designed serverless automation using Lambda and CloudWatch, reducing manual operational intervention by 70%.
- Led migration from Hive-based tables to Apache Iceberg format, enabling ACID compliance and improving query performance by 85%.
- Implemented AWS Glue Data Quality rules and governance framework to ensure data integrity across prime data layers.
- Built distributed PySpark applications to process multi-terabyte Parquet datasets in near real-time.
- Partnered with fraud analytics teams to develop dashboards improving detection time by 60%.
- Drove CI/CD standardization and infrastructure automation using CloudFormation templates.
- Heavily involved in leveraging design, development, testing, code review, and requirement collection from the Fraud team.
- Work together with cross-functional teams made up of developers, architects, and operations to make sure event-driven solutions are successfully implemented and run.
Data Engineer - Deloitte - Mechanicsburg PA
(2021-08 - 2023-12)
Designed and optimized enterprise-scale data platforms supporting Virginia Case Management System (VACMS).
- Developed secure, distributed Hadoop-based data platform handling structured and semi-structured data.
- Built optimized SQL and PL/SQL transformations supporting enterprise reporting and analytics.
- Designed object-oriented Java components integrating business logic into data frameworks.
- Improved query performance by 70% through indexing, partitioning, and performance tuning.
- Resolved production data integrity issues, reducing incident resolution time by 85%.
- Authored technical design documentation defining transformation algorithms and governance standards.
Data Engineer/AWS Developer - Vanguard - Malvern PA
(2019-08 - 2021-08)
Built enterprise Legal & Compliance Data Lake in AWS to support analytics and regulatory reporting.
- Designed scalable ETL pipelines using EMR, Glue, S3, and Python to ingest data from multiple enterprise sources.
- Automated file ingestion workflows from on-prem Windows systems into S3, reducing manual processing time.
- Developed PySpark-based data marts supporting compliance analytics for OGC and Financial Crime teams.
- Created reusable data validation framework in Python improving data accuracy by 95%.
- Built real-time streaming framework using Spark Streaming and Kinesis.
- Developed Splunk dashboards integrating CloudWatch logs for EMR step monitoring.
- Managed CI/CD pipelines using Bamboo, Bitbucket, and Control-M.
Big Data Developer Cloudera Distribution - Blue Cross Blue Shield - MA
(2017-11 - 2019-08)
Developed large-scale Hadoop-based data lake architecture with Raw, Curated, and Publish zones.
- Designed Hive-based transformations with partitioning and bucketing to optimize analytical workloads.
- Migrated sensitive PHI data from Netezza to HDFS/Hive using Sqoop while maintaining compliance controls.
- Implemented dynamic partitioning strategies improving processing efficiency.
- Developed metadata and lineage tracking framework for Hadoop pipelines.
- Performed complex HQL transformations exporting to JSON, AVRO, and Parquet formats.
Big Data Developer - Voya Financial - New York, NY
(2016-12 - 2017-09)
The project was to convert EBCDIC (Extended Binary Coded Decimal Interchange) files into HDFS. The main goal of our project is to migrate all EBCDIC data in Hadoop environment which is lead cost cutting for organization. My responsibility to work on Syncsort tool to ingest COBOL Copybook (Metadata) file and EBCDIC file into HDFS environment.
- Ingested COBOL Copybook and EBCDIC files into HDFS using Syncsort.
- Automated ingestion workflows using Oozie, reducing batch cycle time by 60%.
- Used Pig, Sqoop, and Hive to transform and validate ingested data.
- Contributed to enterprise-wide legacy system migration reducing operational costs.
Hadoop Developer - ANVAYA ANALYTICS PVT. LTD - Bengaluru, Karnataka
(2014-02 - 2015-12)
The ETL module laying the platform for data analytics team to get the insights of the web browsing patterns and ensuring the optimized path is laid for the customers reducing the bounce rates, Miss rates and thus improving the sales.
- Built Spark-based transformation pipelines processing large datasets.
- Designed Hive-based reporting datasets improving analytics turnaround time.
- Implemented Amazon EMR clusters for scalable Hadoop workloads.