Skip to main content

Databricks Data Engineer

Technology
Gitforce
2 weeks agoUntil 14/11/2026
Service contractFully remote

Job description

About the job:We're hiring a Data Engineer to build and maintain governed, BI-ready data layers in Databricks across real client use cases. You'll work across Databricks, SQL, PySpark, Delta Lake, and modern data engineering infrastructure to turn functional requirements and source-system logic into scalable, reliable data models. This is a contract, fully remote role for an India-based engineer who is comfortable working hands-on with data pipelines, dimensional models, validation, and production-grade data systems.

About us:We're a new age tech & business services company. We're HQ'ed in Hyderabad with clients across the world. Find out more about us at https://www.linkedin.com/company/gitforce

What You'll DoTL;DR: Build governed, BI-ready data pipelines, models, and metric layers in Databricks using SQL, PySpark, Delta Lake, and Unity Catalog.

  • Build and test data pipelines, derived tables, dimensional models, and BI-ready data layers in Databricks
  • Design and maintain Unity Catalog structures, catalogs, schemas, permissions, and access controls
  • Translate functional requirements and PostgreSQL source logic into scalable data transformations
  • Design fact and dimension models and Gold-layer datasets optimized for reporting and analytics
  • Build Databricks metric views including dimensions, measures, joins, filters, and reusable business metrics
  • Develop data transformations using SQL, PySpark, Python, Spark, and Delta Lake
  • Implement source-to-target reconciliation and ensure transformed data accurately reflects source systems
  • Build data-quality checks, validation rules, pipeline tests, and reconciliation frameworks
  • Design data models that support multi-tenant SaaS applications and appropriate tenant-level data separation
  • Optimize Databricks pipelines, SQL queries, storage layouts, and workloads for performance and scalability
  • Follow Git-based development workflows and CI/CD practices for data engineering
  • Work closely with functional, validation, BI, and engineering teams to translate business requirements into reliable data solutions
  • Debug data issues, investigate discrepancies, and improve the reliability of production data pipelines
  • Work under the direction of a technical lead while independently owning assigned pipelines, models, and transformations
QualificationsTL;DR: Strong Databricks, SQL, PySpark, Delta Lake, dimensional modelling, and data validation experience are the most important things we're looking for.

You Are:

  • 2 to 5 years into your data engineering or software engineering career as a professional
  • Someone who enjoys working hands-on with data pipelines, transformations, and analytical data models
  • Comfortable taking functional requirements and translating them into practical technical implementations
  • An engineer who cares about data correctness, maintainability, performance, and production reliability
  • Comfortable working independently while collaborating closely with technical leads, functional teams, validation teams, and BI teams
Must-haves:
  • Strong hands-on experience with Databricks
  • Strong SQL programming and data transformation skills
  • Experience with PySpark, Python, Spark, and Delta Lake
  • Experience designing and maintaining Unity Catalog structures, schemas, permissions, and access controls
  • Strong understanding of dimensional modelling, including fact and dimension tables
  • Experience building BI-ready Gold-layer data models
  • Experience working with Databricks metric views including dimensions, measures, joins, and filters
  • Hands-on experience with data validation, source-to-target reconciliation, data-quality checks, and pipeline testing
  • Experience translating source-system logic into scalable data transformations
  • Experience designing data models for multi-tenant SaaS applications
  • Experience with Git-based development workflows and CI/CD practices
  • Good understanding of data engineering architecture, testing, performance optimization, and production engineering fundamentals
  • Strong problem-solving and communication skills
Nice-to-haves:
  • Experience working with PostgreSQL source systems
  • Experience building incremental data pipelines or CDC-based ingestion pipelines
  • Experience with Lakeflow Declarative Pipelines and data-quality expectations
  • Experience with Databricks Asset Bundles
  • Experience with dbt
  • Experience optimizing Databricks SQL queries and workloads
  • Experience designing reusable semantic or metrics layers for BI applications
  • Experience working in complex SaaS or enterprise data environments
Keywords
DatabricksSQLPySparkDelta LakeUnity CatalogDimensional ModellingPythonSparkData ValidationCI/CDGitMulti-tenant SaaSData PipelinesPostgreSQLdbtLakeflowCDCData EngineeringBI-readyGold-layer

Interested in this role?