Description :
We are looking for a Databricks Data Engineer with strong Life Sciences domain experience to join a client-facing delivery team. This role involves building scalable data products from scratch, while also enhancing and optimizing existing data pipelines in a high-impact, production environment.
The ideal candidate will combine deep technical expertise in Databricks and AWS with the ability to engage with stakeholders, understand business requirements, and translate them into robust data solutions.
Key Responsibilities :
- Design, develop, and deploy end-to-end data pipelines and data products using Databricks on AWS
- Build and optimize ETL/ELT workflows using Python, PySpark, and SQL for large-scale structured and unstructured datasets
- Work extensively with Delta Lake architecture, implementing efficient data storage, versioning, and performance optimization
- Develop and manage Databricks Workflows, Jobs, and Delta Live Tables (DLT) for reliable and scalable pipeline orchestration
- Configure and manage Databricks environments (clusters, autoscaling, DBFS, notebooks, Unity Catalog)
- Integrate data from multiple sources using Kafka, Airflow, and cloud-native ingestion frameworks
- Collaborate with business stakeholders and clients to translate requirements into scalable data solutions
- Ensure data quality, governance, and security, especially in regulated Life Sciences environments
- Continuously improve existing pipelines through performance tuning, cost optimization, and reliability enhancements
Must-Have Skills :
- Strong hands-on experience with :
a. Python / PySpark b. SQL (advanced querying and optimization)
c. Databricks (Workspace, Notebooks, Clusters, Autoscaling, DBFS)
a. Delta Lake b. Databricks Workflows, Jobs, and Delta Live Tables (DLT)
c. Unity Catalog (data governance and access control)
- Experience working on AWS cloud platform
- Hands-on experience with Kafka and Apache Airflow for data ingestion and orchestration
- Proven ability to build data pipelines and data products from scratch
- Experience working in client-facing roles, with strong communication and requirement-gathering skills
- Prior experience in Life Sciences / Pharma domain
Good-to-Have Skills :
- Experience with real-time/streaming data pipelines
- Exposure to data modeling (dimensional, lakehouse architecture)
- Knowledge of CI/CD pipelines and DevOps practices for data engineering
- Familiarity with data governance, compliance, and regulatory standards in Life Sciences
- Exposure to ML/AI data pipelines or feature engineering workflows