Data Engineer (Databricks)
Job Summary
About the Role
We are looking for a Data Engineer with 2 years of professional experience to build and maintain the data pipelines that power our analytics and machine learning initiatives on the Databricks Lakehouse Platform.
You will work in a distributed data-intensive environment alongside senior engineers data scientists and analysts. This is a hands-on delivery role: you will own pipelines end to end - ingestion transformation quality and performance - with the guidance and code review you need to grow into platform-level and architecture work.
Key Responsibilities
- Build and maintain scalable data processing pipelines and workflows on the Databricks Lakehouse Platform using Apache Spark and PySpark.
- Ingest and integrate datasets from multiple sources - databases file systems and APIs - into the Lakehouse using Auto Loader CDC and batch or streaming ingestion patterns.
- Model and implement transformation layers following the medallion architecture (bronze/silver/gold) with Delta Lake and Databricks SQL.
- Ensure data integrity consistency and quality through validation and monitoring using Delta Live Tables expectations and Lakehouse monitoring capabilities.
- Tune pipelines for performance and cost: Spark job optimization Delta table maintenance (OPTIMIZE Z-ORDER / liquid clustering) as well as cluster and compute configuration.
- Apply data security access control and privacy practices using Unity Catalog for governance permissions and lineage.
- Contribute to CI/CD and deployment automation for data assets (Databricks Asset Bundles Git-based workflows automated testing).
- Support/ migration initiatives from on-premise Cloudera environments (HDFS Hive Impala) to the Databricks Lakehouse working as part of a team led by senior engineers.
- Collaborate with data scientists and analysts to deliver reliable datasets for analytics and ML workflows.
- Occasionally support lightweight Python services and APIs that expose data to downstream consumers.
Requirements
- Academic background: Bachelors or Masters degree in Computer Science Information Systems Engineering or a related field.
- Experience: 2 years of professional experience in data engineering or software engineering with a strong data component.
- Programming: Solid Python skills in production environments with good engineering practices (version control testing code review).
- Big Data & processing: Hands-on experience with Spark/PySpark for large-scale batch or streaming data processing.
- Lakehouse & Cloud: Working experience with Databricks (or a comparable Lakehouse/cloud data platform) and one major cloud provider - Azure AWS or GCP.
- Databases: Strong SQL and solid data modelling fundamentals across relational and non-relational stores.
- Languages: Fluency in English (written and spoken). Portuguese is a plus.
Nice to Have
- Databricks certifications (e.g. Databricks Certified Data Engineer Associate).
- Experience with Delta Live Tables Structured Streaming or Unity Catalog.
- Familiarity with orchestration tools (Databricks Workflows Airflow) and event-driven architectures (Kafka).
- Exposure to Infrastructure as Code (Terraform Databricks Asset Bundles) and CI/CD tooling.
- Experience with Photon Databricks SQL Warehouses or serverless compute.
About Opplane
Opplane specializes in providing advanced data-focused solutions for financial services telecommunication and reg-tech to accelerate their digital transformation journey. Opplane leadership team is comprised of Silicon Valley serial entrepreneurs and experienced executives. Its expertise comes from years of specific industry experience at some of the worlds top companies such as PayPal Xerox Parc Amazon Wells Fargo SoFi in the areas of product management data technology data governance data privacy security machine learning and risk management.
Why Opplane
Global & Multicultural Diverse perspectives global collaboration (US Portugal India and Singapore offices)
Startup Energy Fast-moving impact-driven environment
Ownership Mindset Engineers own what they build
Collaborative & Friendly Open curious and supportive culture
Most organizations deploying AI to engineering cannot say what they got for it and respond by buying more of it or by arguing. You will build the evidence instead and then use it to decide where the next capability goes.
The measurement program is the first project. The scope is applied AI across the delivery lifecycle with real internal users a platform team to build on and leadership that will act on what you find.