Role Overview:
As a Senior Spark Developer at our Hyderabad office, you will serve as a technical anchor within our data engineering practice. You will be responsible for designing and implementing high-performance data pipelines that process massive datasets to support critical business intelligence and analytics initiatives.
Working closely with cross-functional teams, including data scientists, product managers, and infrastructure engineers, you will ensure that our data platforms remain resilient, scalable, and optimized for performance. Your work will directly influence how our clients derive actionable insights, ensuring data integrity and accessibility across the enterprise.
Key Responsibilities:
- Architect and develop complex ETL pipelines using PySpark and Apache Spark to process large-scale structured and unstructured datasets for downstream analytics.
- Optimize existing data workflows and SQL queries to improve processing efficiency and reduce latency in high-volume data environments.
- Design scalable data models that align with business requirements, ensuring data consistency and ease of consumption for end-users.
- Collaborate with stakeholders to translate complex business problems into technical data solutions, ensuring alignment with project goals and timelines.
- Implement best practices for code versioning, automated testing, and deployment to maintain high standards of software quality within the data engineering lifecycle.
- Mentor junior developers by conducting code reviews and sharing technical expertise to foster a culture of continuous learning and excellence.
Required Skillset:
- Demonstrated expertise in building distributed data processing systems using Python, PySpark, and Apache Spark, with a deep understanding of Spark internals and performance tuning.
- Proficiency in writing complex SQL queries and designing efficient data schemas for large-scale relational and non-relational databases.
- Strong command over Big Data ecosystems, including Hadoop and related storage frameworks, to manage and process petabyte-scale data.
- Ability to communicate complex technical concepts clearly to non-technical stakeholders, facilitating effective collaboration across global teams.
- Proven experience in data modeling techniques that support high-performance reporting and analytical applications.
- A minimum of 6 to 9 years of professional experience in data engineering roles, with a track record of delivering production-grade data solutions.
- Strong problem-solving skills and the ability to adapt to evolving project requirements in a fast-paced, hybrid work environment.