Job Summary :
We are seeking a skilled and detail-oriented Data Engineer to design, develop, and maintain scalable data pipelines and infrastructure. The ideal candidate will work closely with data scientists, analysts, and software engineers to build reliable data systems that support business intelligence, analytics, and machine learning initiatives.
Key Responsibilities :
- Design, develop, and optimize ETL/ELT data pipelines.
- Build and maintain scalable data warehouses and data lakes.
- Integrate data from multiple sources including APIs, databases, and third-party systems.
- Ensure data quality, integrity, security, and governance.
- Monitor and troubleshoot data pipeline performance.
- Collaborate with cross-functional teams to understand data requirements.
- Develop data models to support reporting and analytics.
- Implement automation for data ingestion and transformation processes.
- Optimize SQL queries and database performance.
- Document data architecture, workflows, and best practices.
Required Qualifications :
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Strong proficiency in SQL.
- Experience with Python, Java, or Scala.
- Hands-on experience with ETL tools and data integration frameworks.
- Knowledge of relational and NoSQL databases.
- Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Familiarity with data warehousing concepts and dimensional modeling.
- Experience with Apache Spark, Hadoop, or Kafka is a plus.
- Understanding of version control systems such as Git.
Preferred Skills :
- Experience with orchestration tools like Apache Airflow.
- Knowledge of containerization technologies such as Docker and Kubernetes.
- Familiarity with CI/CD pipelines.
- Experience working with big data technologies.
- Understanding of data governance and security best practices.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration abilities.
Technical Skills :
- SQL
- Python
- Apache Spark
- Apache Kafka
- Apache Airflow
- Hadoop
- Snowflake, Redshift, or BigQuery
- AWS, Azure, or GCP
- Git
- Docker & Kubernetes