Databricks Architect
Job Summary
Databricks Architect
69 Years
We are looking for an experienced Databricks Architect to design and implement scalable secure and high-performance data platforms using Databricks and Apache Spark. The ideal candidate should have strong expertise in data architecture cloud platforms data engineering Lakehouse architecture and enterprise-scale Databricks implementations.
Design and implement enterprise-grade Databricks Lakehouse architectures.
Define data architecture strategies standards governance and best practices.
Design scalable data ingestion transformation processing and analytics pipelines.
Develop architecture solutions using Databricks Apache Spark Delta Lake and cloud-native services.
Define data lake and lakehouse structures using Bronze Silver and Gold layers.
Design batch and real-time data processing solutions.
Provide technical leadership to Data Engineering and BI teams.
Review existing data platforms and recommend modernization strategies.
Design high-performance and cost-optimized Databricks solutions.
Establish security access control encryption and data governance standards.
Design and implement Unity Catalog for centralized data governance.
Define data lineage discovery auditing and access-management strategies.
Design CI/CD and DevOps processes for Databricks notebooks jobs workflows and code.
Integrate Databricks with enterprise data sources APIs warehouses and cloud storage.
Lead migration of legacy data platforms to Databricks where applicable.
Conduct architecture reviews proof-of-concepts and technology evaluations.
Collaborate with Data Engineers Data Scientists BI Developers Cloud Architects and business stakeholders.
Provide technical guidance mentoring and architectural documentation.
Strong hands-on experience with Databricks.
Expert-level knowledge of Apache Spark.
Strong experience with PySpark and/or Scala.
Experience with Delta Lake and Delta tables.
Strong knowledge of Databricks Workflows Jobs Clusters Notebooks and SQL Warehouses.
Experience designing and implementing Lakehouse architecture.
Knowledge of Spark performance tuning and optimization.
Strong understanding of:
Data Lake / Data Warehouse / Lakehouse
Medallion Architecture
Data Modeling
ETL/ELT
Batch and Streaming
Data Governance
Data Quality
Metadata Management
Experience designing enterprise data platforms and integration architectures.
Strong experience with at least one major cloud platform:
Microsoft Azure
Amazon Web Services (AWS)
Google Cloud Platform (GCP)
Preferred Azure technologies include:
Azure Data Lake Storage Gen2
Azure Data Factory
Azure Synapse
Azure Key Vault
Azure Event Hubs
Strong knowledge of Databricks Unity Catalog.
Design and implement catalogs schemas external locations and storage credentials.
Implement role-based access control and data security.
Establish data lineage and auditing.
Define enterprise data governance and compliance practices.
Experience implementing CI/CD for Databricks solutions.
Knowledge of Git/GitHub Azure DevOps GitHub Actions or Jenkins.
Experience with Infrastructure as Code such as Terraform.
Familiarity with automated testing deployment and release management.
Optimize Spark workloads SQL queries clusters and Delta tables.
Experience with partitioning caching file optimization and indexing strategies.
Knowledge of Delta Lake OPTIMIZE VACUUM Z-Ordering and liquid clustering.
Implement appropriate cluster sizing and autoscaling strategies.
Monitor and optimize Databricks platform costs.
Experience with Databricks SQL and BI integration.
Knowledge of MLflow and Machine Learning workloads.
Experience with streaming technologies such as Kafka or Azure Event Hubs.
Knowledge of Terraform and Infrastructure as Code.
Experience with Generative AI RAG or Vector Search.
Familiarity with Microsoft Fabric or Snowflake.
Experience with large-scale cloud migration projects.
Knowledge of enterprise security and regulatory requirements.
Databricks Apache Spark PySpark Scala Delta Lake Unity Catalog Databricks SQL Azure/AWS/GCP ADLS ADF Kafka Terraform Git CI/CD MLflow
Bachelors or Masters degree in Computer Science Information Technology Engineering Data Science or a related discipline.
The ideal candidate should have strong experience designing enterprise Databricks Lakehouse platforms and leading complex data engineering initiatives. The candidate should combine hands-on Databricks expertise with strong knowledge of cloud architecture Spark Delta Lake Unity Catalog security governance performance optimization and DevOps.
Experience leading large-scale Databricks migration and modernization programs will be highly valued.
Required Skills:
Databricks & Spark Strong hands-on experience with Databricks. Expert-level knowledge of Apache Spark. Strong experience with PySpark and/or Scala. Experience with Delta Lake and Delta tables. Strong knowledge of Databricks Workflows Jobs Clusters Notebooks and SQL Warehouses. Experience designing and implementing Lakehouse architecture. Knowledge of Spark performance tuning and optimization.