Enter a job title or keyword

Data Engineer Tech Lead (Databricks, Pyspark)

EPAM Systems


Job Location:

London - UK

Monthly Salary: Not provided by the employer
Posted: 9 September 2026 (9 hours ago)
Application Deadline: 7 December 2026
Vacancies: 1 Vacancy

Job Summary

Were looking for a Senior Data Engineer Tech Lead (Databricks PySpark) to join our team in London UK in a hybrid working mode.

In this role you will lead the design development and optimization of scalable cloud-native data architectures focusing on Azure Databricks PySpark and Lakehouse principles. You will work hands-on to deliver performant data solutions for high-volume workloads ensuring governance reliability and best practices for enterprise-grade platforms.

As a technical leader you will define data strategies drive modernization initiatives and mentor engineers fostering excellence and innovation throughout the team. This position offers the opportunity to shape large-scale data ecosystems implement modern engineering practices and enable next-generation analytics and AI-driven solutions.

Responsibilities
  • Lead the architecture design and build of large-scale data platforms using Azure Databricks and modern cloud technologies
  • Implement and optimize ETL workflows and streaming pipelines with PySpark and Delta Live Tables following Lakehouse principles
  • Enhance performance manage cloud costs and ensure platform reliability for structured streaming workloads
  • Define data governance security and quality standards to maintain consistency across the platform
  • Collaborate with stakeholders to translate complex business requirements into actionable technical solutions
  • Develop integration approaches using Azure-native services such as Data Factory Synapse and Blob Storage
  • Mentor data engineers promote modern engineering practices and perform technical reviews
  • Drive adoption of CI/CD Infrastructure as Code and automated testing in data engineering environments
  • Implement observability and monitoring using tools like Databricks Workflows and related frameworks
  • Contribute to AI-driven initiatives by leveraging Databricks ML/MosaicML to integrate Generative AI and LLM-based solutions
Requirements
  • Bachelors or Masters degree in Computer Science Software Engineering or related field
  • Extensive experience designing and implementing production-grade platforms using Azure Databricks
  • Expertise in PySpark including advanced optimization data skew mitigation and query tuning
  • Strong programming skills in Python with knowledge of modern software design principles
  • Practical experience with structured streaming Delta Lake and Delta Live Tables
  • Proven experience in Lakehouse migration and modernization using open table formats such as Delta Lake or Apache Iceberg
  • Proficiency with cloud-native services on Azure and knowledge of multi-cloud environments (AWS or GCP)
  • Hands-on experience with CI/CD and Infrastructure as Code tools (Terraform GitHub Actions Jenkins)
  • Strong leadership ability to guide teams define epics/user stories and ensure delivery in agile environments
  • Excellent communication and stakeholder management skills for both technical and non-technical audiences
Nice to have
  • Experience operationalizing LLM or Generative AI workflows in Databricks pipelines
  • Familiarity with frameworks like LangChain LlamaIndex or Databricks ML/MosaicML
  • Knowledge of AI governance security practices and enterprise integration controls
  • Background in financial trading data or related domains
  • Official Databricks certifications such as Certified Data Engineer Professional or Apache Spark Developer

Required Experience:

Senior IC