Senior Databricks Engineer
Malvern, PA - USA
Job Summary
Job Title: Senior Databricks Engineer
Location: Malvern PA
Job Type: Full-Time
Experience: 10-15 Years
Job Summary
We are seeking a highly experienced Senior Databricks Engineer with deep expertise in Apache Spark Databricks Lakehouse Platform Delta Lake AWS and modern data engineering. The ideal candidate will design and build scalable secure and high-performance data pipelines while implementing enterprise-grade data governance CI/CD and MLOps practices. Experience with Generative AI LLMs and Vector Search is highly desirable.
Required Skills
Databricks & Apache Spark
- 10-15 years of experience in Data Engineering
- Expert in:
- Databricks Lakehouse Platform
- Apache Spark
- Spark SQL
- PySpark
- DataFrame API
- Distributed Data Processing
- Spark Performance Tuning
- Query Optimization
Delta Lake & Lakehouse Architecture
- Delta Lake
- Medallion Architecture (Bronze Silver Gold)
- Delta Live Tables (DLT)
- ACID Transactions
- Z-Ordering
- Data Optimization
- Time Travel
- Schema Evolution
Data Pipelines & Orchestration
- ETL / ELT Development
- Databricks Workflows
- Databricks Jobs
- Delta Live Tables (DLT)
- Workflow Automation
- Batch Processing
- Streaming Data Pipelines
- Error Handling
- Retry Mechanisms
Programming Languages
- Python
- PySpark
- SQL
- Scala (Preferred)
AWS Cloud & Storage
- Amazon S3
- Amazon Redshift
- AWS Glue
- Amazon Kinesis
- AWS Step Functions
- AWS IAM
- AWS KMS
- Cross-Account IAM Roles
- S3 Bucket Policies
Databricks Governance & Security
- Unity Catalog
- Data Lineage
- Row-Level Security
- Column-Level Security
- Data Governance
- Secure Data Sharing
- Compliance
Compute & Cost Optimization
- Cluster Policies
- Instance Profiles
- Spot Instances
- Compute Optimization
- Cost Management
CI/CD & DevOps
- Git
- Databricks Git Folders
- CI/CD Pipelines
- Version Control
- Deployment Automation
AI & MLOps
- MLflow
- Model Registry
- Experiment Tracking
- Feature Engineering
- Large Language Models (LLMs)
- Vector Search
- Generative AI Integrations
Key Responsibilities
- Design develop and maintain scalable ETL/ELT pipelines using Databricks Apache Spark and Delta Lake
- Build and optimize batch and streaming data pipelines processing data from multiple enterprise sources
- Develop high-performance PySpark and Spark SQL solutions optimizing distributed processing and execution plans
- Implement Medallion Architecture (Bronze Silver Gold) using Delta Lake best practices
- Design and automate resilient workflows using Databricks Workflows Jobs and Delta Live Tables (DLT)
- Configure Unity Catalog for enterprise data governance lineage and fine-grained access controls
- Integrate Databricks with AWS services including Amazon S3 Redshift Glue Kinesis and Step Functions
- Implement secure cloud architectures using IAM roles S3 bucket policies and AWS KMS customer-managed encryption keys
- Optimize Databricks compute resources through cluster policies instance profiles and Spot Instances
- Collaborate with Data Scientists ML Engineers Analysts and BI teams to support analytics machine learning and Generative AI initiatives
- Implement CI/CD pipelines Git-based development workflows and deployment automation
- Utilize MLflow for experiment tracking model management and MLOps lifecycle support
Preferred Qualifications
- Experience with real-time streaming architectures and event-driven data platforms
- Hands-on experience with Generative AI LLMs Vector Databases and Retrieval-Augmented Generation (RAG)
- Databricks Certified Data Engineer Professional or Associate
- AWS Certified Data Analytics or Solutions Architect certification
- Experience with Terraform or Infrastructure-as-Code (IaC)
- Experience supporting enterprise-scale cloud data platforms