Assistant Manager Data Engineer
Job Summary
We are looking for an experienced Data Engineer to design build and maintain scalable data pipelines and infrastructure that power analytics reporting and machine learning initiatives across the organization. The ideal candidate has a strong foundation in data modeling distributed systems and cloud-based data platforms with a track record of delivering reliable high-performance data solutions.
Design build and optimize scalable ETL/ELT pipelines to ingest transform and load data from diverse sources (databases APIs streaming platforms third-party systems)
Develop and maintain data warehouse/data lake architectures ensuring data quality consistency and reliability
Collaborate with data scientists analysts and product teams to understand data requirements and deliver well-structured accessible datasets
Build and maintain batch and real-time streaming data pipelines using tools like Apache Spark Kafka or Airflow
Implement data quality checks monitoring and alerting to ensure pipeline reliability and data integrity
Optimize database and query performance for large-scale datasets
Design and maintain data models (conceptual logical physical) and schemas that support analytics and application needs
Work with cloud platforms (AWS/Azure/GCP) to manage data infrastructure storage and compute resources
Implement and enforce data governance security and compliance best practices
Participate in code reviews CI/CD pipeline development and infrastructure-as-code practices
Document data pipelines architecture and processes for team knowledge sharing
Troubleshoot and resolve production data pipeline issues in a timely manner
Bachelors degree in Computer Science Engineering or a related field
46 years of hands-on experience as a Data Engineer or in a similar role
Strong programming skills in Python and/or Scala; solid SQL expertise
Experience with ETL/ELT tools and orchestration frameworks (Apache Airflow dbt Luigi or similar)
Hands-on experience with big data technologies (Apache Spark Hadoop Kafka)
Proficiency with relational databases (PostgreSQL MySQL SQL Server) and NoSQL databases (MongoDB Cassandra DynamoDB)
Experience working with cloud data platforms and services (AWS Redshift/Glue/S3 Azure Data Factory/Synapse GCP BigQuery/Dataflow)
Solid understanding of data modeling concepts entity relationships cardinality normalization and dimensional modeling (star/snowflake schemas)
Experience with data warehousing solutions (Snowflake Redshift BigQuery Databricks)
Familiarity with version control (Git) and CI/CD practices
Understanding of data governance security and privacy best practices (GDPR data masking access controls)
Strong problem-solving skills and ability to work with large complex datasets