Data Engineer
Job Summary
Start: ASAP
We are looking for a strong hands-on Data Engineer with deep Databricks and Apache Spark experience to support a Frankfurt-based banking client.
The ideal candidate will have extensive experience designing and implementing scalable data engineering solutions using Databricks Spark/PySpark with a strong understanding of data modeling and modern data architectures. Experience working with unstructured and semi-structured data for AI/ML and RAG use cases is highly desirable.
- Design develop and optimize data engineering pipelines and data processing solutions using Databricks and Apache Spark.
- Build scalable and reliable data pipelines using PySpark and related Spark technologies.
- Work with large and complex datasets across structured semi-structured and unstructured data sources.
- Design and implement effective data models to support analytics AI/ML and downstream data consumption.
- Process and transform unstructured and semi-structured data for AI-driven use cases including RAG (Retrieval-Augmented Generation).
- Develop data ingestion transformation cleansing and enrichment workflows.
- Optimize Spark jobs and Databricks workloads for performance scalability and reliability.
- Work closely with data scientists ML/AI engineers architects and business stakeholders to deliver production-ready data solutions.
- Apply strong engineering practices around data quality testing monitoring and operational reliability.
- Contribute to the design and evolution of modern cloud-based data platforms.
- Strong hands-on experience with Databricks in production environments.
- In-depth knowledge of Apache Spark and PySpark including performance tuning and optimization.
- Strong Data Engineering background with experience building production-grade data pipelines.
- Solid understanding of data modeling data structures and modern data architectures.
- Proven experience processing large-scale datasets.
- Experience working with unstructured and semi-structured data.
- Practical experience preparing and transforming data for AI/ML and RAG use cases.
- Strong Python skills particularly for data engineering and PySpark development.
- Experience with data ingestion transformation orchestration and pipeline automation.
- Ability to work independently in a fast-paced banking/enterprise environment.
- Candidate must be located within the EU.
- Experience with Generative AI / LLM / RAG architectures.
- Knowledge of vector search embeddings chunking and document-processing pipelines.
- Experience with Delta Lake / Delta tables and modern lakehouse architectures.
- Experience with cloud platforms such as Azure AWS or GCP.
- Experience in banking or other regulated financial-services environments.
- Knowledge of data governance security lineage and compliance requirements.
- Experience with CI/CD and DevOps practices for data platforms.