Data Engineer AI Java/Python/Spark
Infotree Global SolutionsOpis stanowiska
Role Overview:
We are looking for a Senior/Lead Data & GenAI Engineer to design, build and maintain scalable, production-grade data and AI platforms.
The role combines software engineering, distributed data processing, cloud-native technologies and Generative AI. You will work across teams, drive technical initiatives and build reusable libraries and frameworks that enable reliable, scalable and testable systems.
Key Responsibilities:
- Develop, test and maintain high-quality, production-ready software.
- Design and implement large-scale data pipelines and distributed processing systems.
- Build scalable cloud-native services and platforms using modern engineering practices.
- Provide technical leadership for cross-team initiatives and complex engineering projects.
- Design and develop reusable libraries, frameworks and platform components.
- Optimize distributed data processing workloads for performance, scalability and reliability.
- Work with data platforms including Databricks, Apache Spark and Snowflake.
- Develop and deploy applications using Python and/or Java.
- Build and operate containerized workloads using Kubernetes and cloud-native technologies.
- Design and implement GenAI/LLM-based applications and services.
- Work with frameworks such as LangChain and LangGraph for LLM orchestration and agentic workflows.
- Collaborate with data scientists, software engineers, architects and product teams.
- Establish engineering best practices around testing, observability, reliability and deployment.
- 5 years of professional software/data engineering experience.
- Strong hands-on experience with Python and/or Java.
- Strong experience with Apache Spark and distributed data processing.
- Experience with Databricks and/or modern lakehouse platforms.
- Experience with Snowflake or comparable cloud data warehouses.
- Practical experience with Kubernetes and cloud-native technologies.
- Experience designing and maintaining large-scale data pipelines.
- Strong understanding of distributed systems, scalability and production engineering.
- Experience developing ML/AI or GenAI applications.
- Experience with LLM-based applications, RAG, AI agents or LLM orchestration.
- Familiarity with LangChain, LangGraph or similar GenAI frameworks.
- Strong software engineering fundamentals including testing, code quality and system design.
- Experience with AWS, Azure or GCP.
- Experience with streaming technologies such as Kafka.
- Experience with Delta Lake / Lakehouse architecture.
- Experience building RAG pipelines and vector-search solutions.
- Experience with LLM evaluation, observability and productionization.
- Experience with AI agents, tool calling and multi-step workflows.
- Experience building internal developer platforms, frameworks or reusable engineering libraries.
- Experience leading cross-functional or cross-team technical initiatives.
The strongest candidate is not purely a Data Engineer and not purely an ML Engineer.
We are looking for someone who combines:
Software Engineering Data Engineering Cloud/Platform Engineering GenAI
Typical backgrounds may include
- Senior Data Engineer
- Lead Data Engineer
- Senior Software Engineer – Data
- Data Platform Engineer
- Senior Cloud Data Engineer
- AI/ML Platform Engineer
- Senior ML Engineer with strong data engineering experience
- GenAI Engineer with strong distributed-data/platform experience
- Data & AI Architect / Technical Lead
Languages: Python, Java
Data: Apache Spark, Databricks, Snowflake, Delta Lake
Cloud/Platform: Kubernetes, Docker, AWS/Azure/GCP, cloud-native technologies
GenAI/ML: LLMs, RAG, LangChain, LangGraph, AI agents, vector search
Engineering: Distributed systems, APIs, CI/CD, automated testing, observability, scalability
Interesuje Cię ta oferta?