About the Role
Were looking for curious, high-ownership engineers who enjoy solving real systems and infrastructure problems at scale. This is not a traditional ETL-heavy Data Engineering role.
We want engineers who think beyond pipelinespeople who understand distributed systems, real-time architectures, backend engineering, infrastructure, and modern open-source data stacks.
What Youll Do
- Design and build scalable, high-performance data platforms and pipelines
- Work on distributed data systems across batch and real-time processing
- Take end-to-end ownership - from architecture to deployment and optimization
- Debug, optimize, and extend open-source data systems
- Solve problems at the infrastructure and systems level (not just tool usage)
- Collaborate across teams and drive data engineering best practices
What Were Looking For :
- 2 years of experience in Data Engineering / Data Platform roles
- Strong understanding of distributed systems, data storage, and compute layers
- Explore modern AI tooling and AI-assisted engineering workflows where relevant
- Ability to design systems from first principles, not just use tools
- Hands-on experience with open-source or self-managed architectures
- Strong programming skills in Python or Go (Golang)
- Experience with system design, performance tuning, and debugging at scale
- Comfortable working in a startup or fast-paced product environment
- Curious, self-driven, and excited to learn new technologies
Nice to Have
- Backend engineering experience (APIs, services, system design)
- Experience working on infrastructure and deployment
- Exposure to high-scale or real-time systems
Key Skills
- Query & OLAP : Trino, ClickHouse, Apache Pinot
- Batch & Stream Processing : Apache Spark (OSS), Apache Flink (OSS)
- Table Format : Apache Iceberg
- Cataloging : AWS Glue
- Cloud : AWS
- Languages : Go, Python