Lead Data Engineer Enterprise Data & Analytics

Mayo Clinic


Job Location:

Rochester, NH - USA

Monthly Salary: Not Disclosed
Posted on: 6 hours ago
Vacancies: 1 Vacancy

Job Summary

Description

Lead data design prototype and development of data pipeline architecture pipelines. Lead implementation of internal process improvements: automating manual processes optimizing data delivery re-designing infrastructure for greater scalability. Lead cause analysis on external and internal processes and data to identify opportunities for improvement and answer questions. Excellent analytic skills associated with working on unstructured datasets. Understand the architecture be a team player lead technical discussions and communicate the technical discussion. Be a senior Individual contributor of the Data or Software Engineering teams. Be part of Technical Review Board along with Manager and Principal Engineer. Be a technical liaison between Manager Software Engineers and Principal Engineers. Collaborate with software engineers to analyze develop and test functional requirements. Serve as a hands-on technical leader who actively designs develops reviews and optimizes production-grade data pipelines data products and platform capabilities. Maintain significant contribution to production codebases while establishing engineering standards mentoring team members and driving delivery of scalable resilient solutions. Mentor and Coach Engineers. Work with team members to investigate design approaches prototype new technology and evaluate technical feasibility. Work in an Agile/Safe/Scrum environment to deliver high quality software. Establish architectural principles select design patterns and then mentor team members on their appropriate application. Facilitate and drive communication between front-end back-end data and platform engineers. Play a formal Engineering lead role in the area of expertise. Keep up-to-date with industry trends and developments.


Key Responsibilities:
These positions are hands-on engineering this role employees are expected to actively design develop review and optimize production code and platform capabilities while providing technical leadership and mentorship to engineering teams.



Qualifications

Bachelors Degree in Computer Science/Engineering or related field with 6 years of experience OR an Associates degree in Computer Science/Engineering or related field with 8 years of experience. Knowledge of professional software engineering practices and best practices for the full software development life cycle (SDLC) including coding standards code reviews source control management build processes testing and operations. Have in-depth knowledge of data engineering and building data pipelines with a minimum of 5 years of experience in data engineering data science or analytical modeling and basic knowledge of related disciplines. Worked and lead Data Engineering teams in Continuous Integration / Continuous Delivery model. Build/Lead Data products highly resilient in nature. Build/Lead Test Automation suites Unit Testing coverage Data Quality Monitoring & Observability. A minimum experience of 5 years using relational databases and NoSQL Databases. Experience with cloud platforms such as GCP Azure AWS.
Continuous Integration using Jenkins Git Hub Actions or Azure Pipelines. Experience with cloud technologies development and deployment. Experience with tools like Jira GitHub SharePoint Azure Boards. Experience using advanced data processing solutions/capabilities such as Apache Spark Hive Airflow and Kafka GCP Dataflow. Experience using big data statistics and knowledge of data related aspects of machine learning. Experience with Google BigQuery FHIR APIs and Vertex AI. Knowledge of how workflow scheduling solutions such as Apache Airflow and Google Composer related to data systems. Knowledge of using Infrastructure as code (Kubernetes Docker) in a cloud environment.

The preferred candidate will possess:

  • Advanced proficiency in Python and SQL with demonstrated experience building and supporting production-grade solutions.
  • Advanced experience designing and implementing scalable distributed computing solutions using technologies such as Spark Flink Ray or comparable frameworks.
  • Deep understanding of cloud-agnostic architecture principles and modern data platform design.
  • Advanced experience with open data architecture technologies including Apache Iceberg Delta Lake and Apache Hudi.
  • Strong expertise with modern analytical data formats including Parquet Avro and ORC.
  • Experience designing data platforms that support analytics AI/ML and operational workloads at enterprise scale.
  • Experience implementing CI/CD automated testing Infrastructure-as-Code observability and engineering best practices.
  • Experience designing systems for scalability reliability security resiliency and long-term maintainability.



Required Experience:

Senior IC

DescriptionLead data design prototype and development of data pipeline architecture pipelines. Lead implementation of internal process improvements: automating manual processes optimizing data delivery re-designing infrastructure for greater scalability. Lead cause analysis on external and internal ...

About Company

Company Logo

Why Mayo Clinic Mayo Clinic is top-ranked in more specialties than any other care provider according to U.S. News & World Report. As we work together to put the needs of the patient first, we are also dedicated to our employees, investing in competitive compensation and comprehensive ... View more

View Profile View Profile