Responsibilities Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java) Lead design of batch and streaming ETL/ELT systems handling large data volumes Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction