Senior Data Engineer
Send a job offer directly to this candidate
Data Engineering professional with 8+ years of experience building reliable, production-grade data pipelines, ETL/ELT workflows, and data integration solutions across enterprise and financial data environments. Strong hands-on expertise in Python, SQL, Databricks, Apache Spark (PySpark) and modern Lakehouse patterns (Bronze/Silver/Gold). Strong hands-on experience developing scalable Python and Java frameworks for enterprise data processing and ETL applications.
Experienced in building cloud-native data engineering solutions using AWS Glue, Amazon S3, AWS EMR, AWS Lambda, Amazon Redshift, Athena, and Step Functions. Hands-on experience writing PyTest unit and integration test cases for Python and PySpark applications. Experienced in developing and consuming RESTful Web Services and integrating enterprise applications using APIs.
Strong knowledge of Unix Shell scripting for deployment automation, monitoring, and production support. Proven track record of designing end-to-end ETL/ELT pipelines using Azure Data Factory / Fabric Data Factory, Databricks Workflows, and orchestration tools. Deep understanding of Spark internals including RDD transformations/actions, caching/persistence, partitioning, shuffle behavior, and performance tuning for large workloads.
Experienced in creating Delta Lake tables with ACID compliance, schema evolution, MERGE/UPSERT, time travel, and optimized file layouts for scalable analytics. Strong SQL engineering background—writing and tuning complex queries, data validation logic, and warehouse-friendly transformations for reporting and downstream consumption. Hands-on experience building deployment automation with CI/CD (Azure DevOps / GitHub Actions) using YAML pipelines, approvals, RBAC, and environment-based releases (Dev/Test/Prod).
Skilled in building secure, production-grade data solutions using Key Vault, secrets management, service principals, and access controls (RBAC/Unity Catalog patterns). Designed ingestion frameworks pulling data from ERP systems (SAP, Oracle, Infor M3), APIs, files (CSV/JSON/XML), and databases into cloud platforms.
Experience implementing data quality and reconciliation checks (row counts, hash totals, schema checks, duplicate detection, referential checks) to ensure trusted data products. Comfortable with interactive development in Databricks notebooks / Jupyter, building reusable utilities and parameterized jobs for scalable execution. Familiar with orchestration patterns using Airflow (including Dockerized deployments) and legacy scheduling ecosystems (e.g., Oozie/Control-M concepts).
DevOps mindset—version control, branching strategies, code reviews, unit testing, job monitoring, alerting, and operational runbooks for supportability.
Experience with data governance and metadata—documenting KPI definitions, lineage, and standards for consistent reporting across teams. Collaborative communicator who partners with stakeholders to translate business requirements into performant data
Senior Data Engineer - Daikin - Waller, TX
(2024-01)