Big Data (Hadoop and PySpark) Developer
Send a job offer directly to this candidate
Presently working as Big Data (Hadoop and PySpark) Developer. Having around 4.6 years of overall IT experience in Big Data Development, Spark, AWS (s3, redshift, Glue). Having 4.3 years of exclusive experience in Hadoop Stack, HDFS, Map Reduce, Hive, Oozie, Airflow.
Hands on experience in Spark, PySpark, Spark Core, Spark SQL. Interacting directly to the business users to gather requirement and report progress. Understanding the requirements. Good understanding of Hadoop architecture and hands-on experience with Hadoop components such as Job Tracker, Task Tracker, Name Node, Secondary Name Node, Data Node, Map Reduce concepts and YARN architecture which includes Node manager, Resource manager and App Master.
Responsible to manage data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data. Involved in creating Hive tables, loading data and running Hive queries in the data.
Partitions, Buckets based on State to further process using Bucket based Hive. Develop business requirement with spark SQL and HQL. Worked on creating the RDDs, Data Frames for the required input data and performed the data operations using Spark-core. Writing the SQL queries to process the data using Spark SQL.
Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data.
Spark applications by using Python to connect the data sources. Excellent experience with application development and performance tuning. Used Hive queries in Spark-SQL for analyzing and processing the data. Hands on experience with Linux and SQL.
Experience on file formats like ORC, AVRO and PARQUET.
Experience working on GitHub and JIRA to track issues and crucible for code review.
Data Engineer - HCL Tech
(2021-11)
B.Sc. - YV University (2019)
MCA - Computer applications - SV University (2021)