Job Description
Job title- Python Data Engineer (Kafka & PySpark)
Total Exp -5 to 8 yrs
Location- Pune, Chennai ,Bangalore ,Mumbai, Hyderabad
Notice Period - Only Immediate Joiners
Mode- Hybrid
Job Description The candidate should have strong experience in Python development with expertise in Kafka and PySpark. The role involves designing, developing, and maintaining scalable data processing solutions, building real-time data pipelines, and working with distributed data processing frameworks. The candidate should have good understanding of data engineering concepts, streaming architectures, and cloud/big data technologies.
Strong problem-solving, communication, and collaboration skills are required to work effectively with cross-functional teams.
Key Responsibilities
- Develop and maintain scalable data processing applications using Python
- Design, build, and support real-time data pipelines using Apache Kafka
- Develop data transformation and processing workflows using PySpark
- Work with large-scale distributed data processing frameworks
- Perform data ingestion, transformation, and integration activities
- Troubleshoot and resolve technical issues related to data pipelines and applications
- Optimize application performance and ensure scalability of solutions
- Collaborate with cross-functional teams to understand business requirements and deliver technical solutions
- Participate in design discussions, code reviews, and technical implementations
- Follow Software Development Life Cycle (SDLC) processes and coding best practices
- Prepare and maintain technical documentation as required
Core Skills Required:
- Minimum 5 years of experience in IT industry with strong Python development experience
- Hands-on experience in Python programming
- Strong experience with Apache Kafka
- Hands-on experience in PySpark
- Experience in developing data processing and ETL/ELT pipelines
- Good understanding of distributed data processing concepts
- Experience working with large datasets and performance optimization
- Strong debugging and problem-solving skills
- Good understanding of Software Development Life Cycle (SDLC)
- Strong communication and collaboration skills
- Good documentation skills
Good to Have
- Experience working in Big Data ecosystems
- Exposure to cloud platforms and data engineering solutions
- Experience working with enterprise-scale data applications
- Knowledge of real-time streaming architectures
- Experience working in Agile delivery environments
- Experience supporting production data pipelines.