Role Overview :
As a PySpark Developer, you will be at the forefront of building scalable data pipelines that transform raw information into actionable business intelligence. You will work closely with cross-functional teams, including data scientists, product managers, and infrastructure engineers, to design robust architectures that handle massive datasets with high efficiency. Your daily contributions will directly influence the companys data-driven decision-making capabilities, ensuring that our stakeholders have access to reliable, high-quality data to drive strategic growth and operational excellence across global markets.
Key Responsibilities :
- Architect and optimize complex data processing pipelines using PySpark to ensure high-performance data ingestion and transformation for downstream analytics.
- Collaborate with engineering teams to migrate legacy data systems to cloud-native environments on AWS, enhancing system scalability and reducing latency.
- Develop and maintain sophisticated SQL queries and Hive scripts to support complex reporting requirements and ad-hoc data analysis for business stakeholders.
- Implement best practices in Big Data engineering to ensure data integrity, security, and compliance across all production environments.
- Troubleshoot and resolve performance bottlenecks within Hadoop and Spark clusters to maintain optimal system uptime and resource utilization.
Required Skillset :
- Demonstrated expertise in building and maintaining large-scale data pipelines using PySpark and Python, with a deep understanding of distributed computing principles.
- Proven ability to design and manage data workflows within AWS ecosystems, leveraging cloud services to solve complex data engineering challenges.
- Strong proficiency in SQL and Hive for data modeling and complex query optimization, coupled with a solid grasp of Hadoop architecture.
- Exceptional communication skills with the ability to translate technical data concepts into clear insights for non-technical stakeholders and cross-functional partners.
- High degree of adaptability to work in hybrid or distributed team environments across Hyderabad, Chennai, Bangalore, Kolkata, or Pune, maintaining productivity and collaboration in fast-paced settings.
- A Bachelors or Masters degree in Computer Science, Information Technology, or a related quantitative field, supported by 4 to 10 years of hands-on experience in the Big Data domain.