Phenom Intro :
Our purpose is to help a billion people find the right job! Phenom is an AI-Powered talent experience platform that is redefining the HR tech space. We have grown into a global organization with offices in 6 countries and over 2,000 employees. As an HR tech unicorn organization, innovation and creativity is within our DNA. Come help us make every talent moment Phenomenal!
Job Summary :
Phenom People is looking for an experienced Data Engineer to optimize and enhance our data pipelines that power insights across the organization. This role focuses on improving performance, ensuring data quality, automating pipeline development, and enabling data sharing through scalable and secure mechanisms.
The ideal candidate will have strong expertise in PySpark,SQL optimization, and experience with data sharing platforms such as DataHub. You'll collaborate closely with data engineering, platform, and analytics teams to modernize our data ecosystem built on- Airflow, Livy, EMR, Flink, Iceberg, and Snowflake.
If you're passionate about performance tuning, automation, and building resilient data systems - we'd love to hear from you.
What You'll Do :
- Optimize existing data pipelines built using PySpark for performance and scalability.
- Review and tune SQL queries to improve execution time and resource utilization.
- Design and develop new batch data pipelines that meet evolving business needs (real-time decisioning is not required; systems can tolerate up to 6 hours of delay).
- Automate pipeline generation and metadata management to streamline development.
- Implement data quality and validation frameworks to ensure reliability and trust in data assets.
- Collaborate with teams to integrate and share data via DataHub, leveraging connectors such as SFTP and Snowflake Data Share.
- Enhance observability by integrating and monitoring pipelines through Grafana- and other telemetry tools.
- Work closely with other engineers to improve coding practices and leverage AI-powered developer assistants like Cursor, Copilot, or Windsurf- for productivity.
- Participate in design reviews, documentation, and operational support for production systems.
Work Experience :
What You've Done :
- 5- 8 years of experience as a Data Engineer working with large-scale distributed data systems.
- Strong proficiency in Python (especially PySpark) and hands-on experience optimizing Spark jobs on- EMR or similar environments.
- Deep knowledge of SQL tuning and query optimization across analytical databases (preferably Snowflake).
- Experience working with Airflow for workflow orchestration and Livy for Spark job management.
- Exposure to Flink and Apache Iceberg for batch or incremental data processing.
- Strong understanding of data quality frameworks, testing, and validation techniques.
- Familiarity with DataHub or similar data catalog and sharing platforms.
- Experience setting up dashboards and alerts in Grafana- or equivalent monitoring tools.
- Exposure to AI coding assistants (Cursor, GitHub Copilot, Windsurf) for accelerating development is a plus.
- Passionate about automation, continuous improvement, and building scalable, maintainable systems.