Data Engineer - Centene Corporation - Atlanta, GA
(2022-06)
Centene Corporation is a leading healthcare enterprise committed to helping people live healthier lives. Centene offers affordable and high-quality products to nearly 1 in 15 individuals across the nation, including Medicaid and Medicare members including Medicare Prescription Drug Plans. This Project collects data from various source system using ADF to support numerous reports, these reports support operational & regulatory needs.
- Designed and implemented Azure Data Factory (ADF) pipelines to orchestrate workflows, ingesting data from diverse sources into ADLS and invoking Databricks notebooks for transformation processing.
- Configured scheduling and tumbling window triggers to automate pipeline execution and ensure timely data availability.
- Integrated data from multiple enterprise systems including SQL Server, Oracle, SharePoint, ADLS, and REST APIs.
- Partnered with DevOps teams to enable CI/CD for ADF pipelines using ARM templates, ensuring streamlined deployments.
- Applied the Databricks medallion architecture to load raw data into the Bronze layer via DataFrame Reader API and manage Delta tables.
- Developed SparkSQL and PySpark logic to populate dimension and fact tables in the Silver layer and deliver curated views in the Gold layer.
- Optimized Databricks performance through partitioning, table optimization, query plan analysis, and cost evaluation.
- Utilized Azure Key Vault, Azure Monitor, and Logic Apps to configure secure access, monitoring, and automated email notifications.
- Integrated Databricks Gold layer datasets with Power BI to support enterprise reporting and analytics.
- Contributed to architectural decisions, enhancing existing frameworks, managing costs, and defining designs for new requirements.
- Created multiple data pipelines in Azure Data Factory (ADF) and ETL workflows as per project need.
- Integrated Snowflake with Azure Data Factory (ADF) for automated data ingestion pipelines.
- Designed end-to-end ETL pipelines to load data from Azure Data Lake / Blob Storage into Snowflake.
- Used COPY INTO commands with external stages for efficient data ingestion.
- Built incremental and full load pipelines with error handling and logging mechanisms.
- Worked on source-to-target mappings, accordingly, created mappings for data transformation.
- Created ADF pipelines to perform ETL tasks which include data from various sources, including Azure Blob Storage, SQL Server, and APIs.
- Successfully implemented data transformations in Mapping Data Flows, also used, and Databricks to handle structured and unstructured data.
- Worked on incremental data load & real-time data synchronization using triggers and integration runtimes.
- Designed Bronze layer ingestion pipelines to land raw data from multiple sources (SQL Server, APIs, ADLS) into data lake using Azure Data Factory.
- Built Silver layer transformation pipelines using Azure Databricks to cleanse, deduplicate, and standardize data.
- Developed Gold layer curated datasets optimized for reporting and analytics using Delta Lake.
- Developed robust data processing scripts using Python for data transformation and automation.
- Strong experience in object-oriented programming (OOP) and modular coding practices.
- Used Python for data extraction, transformation, and loading (ETL) workflows.
- Implemented incremental data processing and data quality validations between Bronze, Silver, and Gold layers.
- Optimized data pipelines by applying partitioning, caching, and query optimization techniques in Spark.
- Wrote complex SQL queries to meet business requirements and used in data validation and identifying duplicate record as per business rules in Azure Synapse Analytics.
- Worked on error handling and understand difficult or poor performing area in ADF using Azure Monitor.
- Migrated ADF pipelines using CI/CD, Azure DevOps and Git repositories.
- Worked on performance tuning and Tuned pipeline by optimizing transformations, multiple partitioning strategies, and parallel processing as per need.
- Performed system testing and worked with business on user acceptance testing (UAT).
- Worked on Agile mode and participated in daily scrums, sprint planning, retrospectives etc.
Data Engineer - Fifth Third Bancorp - Cincinnati, Ohio
(2020-09 - 2022-06)
Fifth Third Bancorp operates as a diversified financial services company in the United States. The company's offers multiple services including a range of deposit and loan products to individuals and small businesses. This project involves building a data warehouse by consolidating data from a variety of Sources into a central data warehouse to solve multiple use cases to support downstream application
- Worked with business & gather/understand project requirements, set meetings multiple times with business.
- Used ADF for ETL purpose to move data from different sources into a target database.
- Built scalable ETL pipelines using Python with PySpark for large-scale data processing.
- Automated data ingestion from multiple sources (CSV, JSON, APIs, databases) using Python.
- Integrated Python scripts with Azure Data Factory (ADF) pipelines for orchestration.
- Implemented data validation and schema checks using Python.
- Created complex data pipelines in ADF to load data into cloud base Data Warehouses tables.
- Used various ADF activities like Copy Data, Data Flow, Lookup, Filter, Join, etc. to process and transform data.
- Created technical design documents after development tasks.
- Performed unit and integration testing to ensure data loads correctly as per business needs.
- Identified and later resolved performance issues in ADF pipelines and used SQL for faster data processing.
- Used SQL queries to debug and validate data transformations rules.
- Observed ADF pipeline runs, also reviewed logs, and troubleshot errors to fix data loads if any.
- Improved performance of data flows and optimized execution time at different stages of pipeline.
- Documented mappings, transformations and pipeline config for reference purposes.
- Used pipeline parameters, variables & triggers to create workflows also created reusable components too.
- Identified / resolved loading failures, including database issues.
- Participated in Agile env. Mainly in daily standups to track progress and improve workflows performance.
- Worked on tuning SQL statements
ETL Developer - HD Supply Holding - Atlanta, GA
(2018-09 - 2020-09)
HD Supply Holdings, Inc. is a holding company. The Company, through its subsidiaries, operates as an industrial distributor of products specializing in maintenance, repair & operations, infrastructure & power. Relevant data was integrated from various source system using ADF tools. Reports and Interactive dashboards were developed that allowed the company to track overall performance through reporting.
- Analysis, requirements gathering, functional/technical specification, development, deploying and testing.
- Extensive use of Informatica Tools like Designer, Workflow Manger and Workflow Monitor.
- Used transformations like Aggregators, Sorter, Dynamic lookups, Connected & unconnected lookups, Filters, Expression, Router, Joiner, Source Qualifier, Update Strategy, sequence Generator.
- Used Workflow Manager/Monitor for creating and monitoring workflows and worklets.
- Used mapping parameters, mapping variables and parameter files
- Worked on tuning SQL statements
- Prepared Test Cases and performed system and integration testing
- Created and Monitor the Sessions.
- Created Informatica mappings to load the data from staging to dimensions and fact tables.
- Conducting unit testing and Prepared Unit Test Specification Requirements.
- Configured mappings to handle updates to preserve the existing records using update strategy transformation.
- Used most of the transformations such as the Aggregators, Filters, Routers, Sequence Generator, Update Strategy, Rank, Expression and lookups while transforming the Data according to the business logic.
- Involved in troubleshooting the loading failure cases, including database problems.
- Created sessions, reusable worklets and workflows in Workflow Manager
- Improved performance by identifying the bottlenecks in Source, Target, Mapping and Session levels.
- Migrated on-prem ETL workflows from Informatica PowerCenter to Informatica Intelligent Cloud Services (IICS).
- Designed and developed ETL pipelines using Informatica Intelligent Cloud Services to load data from multiple sources into cloud data warehouses.
- Built mappings, mapping tasks, and taskflows to extract, transform, and load data from databases, flat files, and APIs.
- Implemented data transformations such as Joiner, Aggregator, Filter, Expression, Router, and Lookup to cleanse and transform data.
- Developed incremental load logic using watermark columns to efficiently process changed data.
- Integrated on-premises systems with cloud platforms using Secure Agent in Informatica Intelligent Cloud Services.
- Designed parameterized mappings and reusable components to improve pipeline reusability and reduce development time.
- Implemented error handling and logging mechanisms in taskflows to capture failed records and enable reprocessing.
- Scheduled and monitored ETL jobs using taskflows and schedules in Informatica Intelligent Cloud Services.
- Optimized ETL performance by tuning transformations, using pushdown optimization, and managing data partitioning.
- Worked with cloud storage systems such as Amazon S3 and Azure Data Lake Storage for data ingestion and storage.