About The Role OdoCore is looking for an experienced Data Engineer to join a large-scale data platform modernization engagement. This is a hands-on role focused on migrating a legacy Hadoop-based data warehouse to a modern lakehouse architecture. You'll be working directly on production-critical pipelines that power core business reporting and analytics, so real-world experience with the tools below — not just theoretical familiarity — is essential.
You'll be part of a team modernizing a large-scale data platform, covering: Migrating Spark 2 → Spark 3
Migrating Oozie → Apache Airflow
Offloading IBM Netezza and IBM DataStage workloads onto a Spark 3 / Iceberg Lakehouse Required Hard Skills Strong SQL, with real experience reading and re-optimizing complex, poorly-written legacy queries
Apache Spark (Spark 2 and Spark 3) — PySpark or Scala
Apache Hive and Apache Iceberg — table formats, partitioning, schema evolution
ETL / data warehousing fundamentals — medallion (bronze/silver/gold) architecture, dimensional modeling
Orchestration tools — Oozie and/or Apache Airflow
Linux/Unix comfort, Git
Cloudera CDP platform exposure (Impala, Ranger) — or a fast ability to ramp up on it Experience Level 3+ years in data engineering, with at least one prior migration, ETL modernization, or lakehouse build under your belt. Mid-to-senior individual contributors preferred.
Interested in this role?