Senior Data Engineer
Job Summary
We are looking for a highly skilled Data Quality Engineer with strong Data Engineering expertise to ensure the accuracy reliability and scalability of enterprise data platforms. The ideal candidate will possess hands-on experience with Databricks PySpark Hadoop Hive and Cloud technologies along with advanced SQL skills to validate data across large-scale data pipelines and Lakehouse architectures.
- 6 years of experience in Data Engineering Data Quality Engineering or Data Testing.
- Hands-on experience with Databricks and PySpark.
- Strong experience in Hadoop ecosystem components such as Hive HDFS Spark and related Apache technologies.
- Advanced SQL expertise for large-scale data validation and analysis.
- Experience working with Data Warehouses Data Lakes and Lakehouse architectures.
- Understanding of Star Schema Snowflake Schema and dimensional modeling.
- Experience with cloud platforms (Azure AWS or GCP).
- Perform end-to-end validation of data pipelines across ingestion transformation and consumption layers.
- Execute source-to-target reconciliation and data quality checks.
- Identify investigate and resolve data anomalies and inconsistencies.
- Define and implement data quality frameworks metrics and controls.
- Develop and validate data pipelines using PySpark and Databricks.
- Work with large-scale datasets in Hadoop Hive and Lakehouse environments.
- Support ETL/ELT workflows and ensure data integrity throughout the data lifecycle.
- Optimize data processing jobs for performance and scalability.
- Write advanced SQL queries for data profiling reconciliation and root cause analysis.
- Perform complex joins window functions CTEs aggregations and query optimization.
- Validate business rules and transformation logic against source systems.
- Validate and monitor data across Databricks Lakehouse architecture.
- Work with cloud platforms such as Azure AWS or GCP.
- Collaborate with Data Engineers Architects and Analysts to ensure reliable data delivery.
- Analyze production issues and conduct root cause analysis.
- Track and manage data defects through resolution.
- Implement proactive monitoring and automated quality checks.
- 6 years of experience in Data Engineering Data Quality Engineering or Data Testing.
- Hands-on experience with Databricks and PySpark.
- Strong experience in Hadoop ecosystem components such as Hive HDFS Spark and related Apache technologies.
- Advanced SQL expertise for large-scale data validation and analysis.
- Experience working with Data Warehouses Data Lakes and Lakehouse architectures.
- Understanding of Star Schema Snowflake Schema and dimensional modeling.
- Experience with cloud platforms (Azure AWS or GCP).
- Automated data testing frameworks.
- Data observability and monitoring tools.
- CI/CD implementation for data pipelines.
- Experience with Delta Lake Unity Catalog or similar technologies.
- Knowledge of Airflow Kafka or other Apache ecosystem tools.
Required Skills:
Role Overview We are looking for a DevOps Engineer with strong Linux and automation skills capable of supporting application troubleshooting monitoring and CI/CD processes in a fast-paced environment. Must Have Linux Shell Scripting ITIL / ITSM PL/SQL Application Troubleshooting Monitoring Tools (Splunk/Dynatrace preferred) DevOps Automation CI/CD Jenkins GitHub