Data Engineer with 4 years of experience building large-scale distributed data pipelines
Send a job offer directly to this candidate
Data Engineer with 4 years of experience building large-scale distributed data pipelines using PySpark, AWS EMR, and Apache Airflow. Proficient in developing Spark DataFrame and RDD-based workflows for structured and semi-structured data, including transformations, joins, aggregations, and window functions. Experienced in implementing Spark RDD-based workflows with Python.
Ability to troubleshoot common issues with Spark DataFrame, such as data processing errors, performance bottlenecks, and scalability limitations.
PySpark applications running on AWS EMR for efficient data processing. Debugged complex Spark data transformations in PySpark jobs on AWS EMR. Expertise in querying Hive tables using SQL-like syntax and performing data analysis using tools like Apache Spark.
Skilled in integrating Hive tables with other big data technologies, such as Hadoop. Familiarity with Hive megastore and its role in managing table metadata and schema evolution. Knowledge of Hive table formats, including ORC, Parquet, and Avro, and their advantages and disadvantages for different use cases.
Proficient in writing Sqoop commands to transfer data between Hadoop and various databases such as MySQL and SQL Server.
Data Engineer (PySpark) - Mavenir
(2024-08)
Data Engineer (PySpark) - Mavenir
(2022-08 - 2024-08)
BCA - Computer Applications - BRABU (Bihar) (2018 - 2021)
12th - BRABU (2017 - 2018)
10th - BSEB (2016 - 2017)