Enter a job title or keyword

Sr. Software Engineer Ingestion Core team

Databricks


Job Location:

San Francisco, CA - USA

Yearly Salary: USD 166000 - 225000
Posted: 18 September 2026 (5 hours ago)
Application Deadline: 16 December 2026
Vacancies: 1 Vacancy

Job Summary

Deeply understanding whats in the enterprise data has been a challenge that Databricks has been addressing by providing analytics and machine learning tools. From data warehousing with Databricks SQL to large-scale distributed processing with Spark and advanced ML tools for experimentation and model serving we empower our customers to gain insights and drive innovation.

To enable all of this on Databricks making data ingestion seamless is crucial. Thats the mission of the Ingestion Core Team: to make the ingestion of all datastructured and unstructuredsimple reliable and efficient. Simplifying the complex is hard and thats where you come in. This role requires building distributed platform systems to incrementally ingest high-volume petabyte-scale data from diverse sourcesincluding cloud storage (SQS ADLS GCS)databases (Oracle SQL Server MySQL Postgres) and file sources (Google Drive SharePoint)at high throughput and low cost. The data includes structured formats (JSON Parquet CSV) as well as unstructured data (text images docs PPTs and blobs) all of which land in Delta Lake with schema evolution and change data capture (CDC) capabilities.

Join us in making data ingestion effortless and be part of the team that powers the future of AI data at Databricks!

As an engineer on the team you will work on projects that:

  • Build distributed infrastructure to ingest data from diverse sources and support streaming ingestion incremental processing and replication. This isnt just about building plugin connectors.
  • Reduce end-to-end latency increase throughput and reduce costs from the time data appears in source systems to when it is available in Delta Lake.
  • Design and optimize streaming and distributed workloads for throughput cost latency reliability and scale.
  • Optimize streaming workloads by exploring and applying ML techniques.
  • Build monitoring and observability capabilities (customer-facing and internal) that provide visibility into ingestion workflows and the systems running them.
  • Collaborate with partner teams to enable use cases like RAG and AI agents.
  • Build Agentic SDKs with strong evaluations.

Ideal Engineer should have:

  • 5 years of experience writing production code in one of: Java Scala Go C or Python.
  • Experience architecting developing and deploying large-scale distributed and asynchronous systems.
  • Experience with distributed systems streaming Spark databases data processing or CDC.
  • Experience building or operating systems where scale throughput latency reliability and cost are important considerations.
  • Comfortable using AI tools and defining effective evals to shorten the development loop.

Required Experience:

Senior IC


About Company

Company Logo

The Databricks Platform is the world’s first data intelligence platform powered by generative AI. Infuse AI into every facet of your business.

View Profile View Profile