Job Description :
We are seeking a highly skilled Python Spark Developer with 8 to 12 years of experience in building and optimizing large-scale data processing applications.
The ideal candidate will possess strong expertise in Python, Apache Spark, Hadoop ecosystem technologies, and data engineering best practices.
This role involves developing high-performance data solutions, leading junior developers, and delivering scalable, low-latency applications capable of handling large volumes of data.
Key Responsibilities :
- Design, develop, and maintain scalable data processing applications using Python and Apache Spark.
- Build and optimize batch and real-time data processing pipelines using Spark Core, Spark SQL, and Spark Streaming.
- Develop high-performance, low-latency, and highly available data applications.
- Work closely with Hadoop ecosystem components to support large-scale data processing requirements.
- Perform coding, testing, debugging, troubleshooting, and deployment activities throughout the software development lifecycle.
- Optimize application performance through tuning, scalability improvements, automation, and resource utilization enhancements.
- Analyze and resolve production issues, ensuring system reliability and operational excellence.
- Collaborate with data engineers, architects, business stakeholders, and cross-functional teams to deliver data-driven solutions.
- Conduct code reviews and enforce development best practices and coding standards.
- Mentor and guide junior Python developers, providing technical leadership and support.
- Participate in design discussions, estimations, sprint planning, and Agile delivery activities.
- Maintain technical documentation and ensure adherence to organizational development standards.
Required Skills & Experience :
- 8 to 12 years of experience in Python development and data engineering.
- Strong proficiency in Python programming.
- Extensive hands-on experience with Apache Spark.
- Experience implementing Spark Core, Spark SQL, and Spark Streaming solutions.
- Strong experience with Pandas for data manipulation, transformation, and analysis.
- Experience working with Hadoop ecosystem technologies and distributed data processing environments.
- Strong understanding of data structures, algorithms, and software engineering principles.
- Experience designing and developing high-performance, scalable, and fault-tolerant applications.
- Expertise in performance tuning, optimization, load balancing, and automation.
- Strong debugging, troubleshooting, and problem-solving skills.
- Experience leading technical teams and mentoring junior developers.
- Familiarity with Agile/Scrum development methodologies.
- Strong communication, collaboration, and stakeholder management skills.
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.