Data Engineer
New York City, NY - USA
Job Summary
Our client is seeking an experienced Data Engineer to join their Center for Data Analytics Innovation and Rigor team in New York City.
As part of the Center for Data Analytics Innovation and Rigor team you will report to the Rubric Engineering and Measurement Specialist. You will develop infrastructure to support large-scale AI evaluation frameworks design scalable data pipelines for generating and processing synthetic data implement secure data storage solutions and create infrastructure for real-time model evaluation and monitoring. You will use common frameworks platforms and languages such as Python SQL GitHub containerization tools (e.g. Docker Kubernetes) and cloud computing infrastructures (e.g. AWS Azure) to build robust and scalable data infrastructure that supports our AI research initiatives.
This is an exempt full-time hybrid position located in our NYC headquarters office or other relevant location. This position requires a minimum of four (4) days per week in the office on a schedule determined by your supervisor. The in-office requirement and schedule are subject to change based on the needs of the program and the organization.
Our client is dedicated to transforming the lives of children and families struggling with mental health and learning disorders by giving them the help they need. They have become the leading independent nonprofit in childrens mental health by providing gold-standard evidence-based care delivering educational resources to millions of families each year training educators in underserved communities and developing tomorrows breakthrough treatments.
Responsibilities:
- Create and maintain scalable data pipelines for efficient storage and retrieval of multimodal data with particular emphasis on clinical natural language and multi-turn response data.
- Create pipelines for data transformation preprocessing and management. Ensure data quality security and compliance with privacy regulations for handling sensitive data.
- Perform quality assurance of pipelines/processes to maintain integrity throughout the data lifecycle.
- Create interactive visualizations and dashboards to communicate data insights and pipeline performance metrics.
- Write documentation and relevant text for scientific clinical or public dissemination of knowledge.
- Perform additional job-related duties as assigned.
Qualifications:
- Masters degree in Neuroscience Psychology Engineering Computer Science or equivalent combination of education and experience is required.
- 5 years of experience in data analysis and data science fundamentals (e.g. algorithms data structures data visualization machine learning) preferably in a clinical or research setting.
- 5 years of experience in at least one scientific programming language (e.g. Python/R Matlab) and related toolboxes or frameworks (e.g. Tidyverse Scipy Sklearn Polars Pytorch) is required.
- 5 years of experience working in a Linux environment using version control systems (e.g. GitHub) and software virtualization platforms (e.g. Docker).
- 5 years of practical experience in Extract Transform Load (ETL) processes and database management languages (SQLNoSQL) and familiarity with associated cloud computing services and frameworks (AWS Azure Terraform).
Special Considerations:
The anticipated salary range for this position is $119000 - $150000 USD annually.
Our clients competitive compensation and benefits include medical insurance 401(k) paid parental leave dependent care flexible work schedules discounted tickets and entertainment perks programs.
The salary range for the position is posted. Factors such as candidates work experience education/training job-related skills internal peer equity as well as market and business considerations affect the salary offered within this addition this salary may be subject to a geographic adjustment (according to a specific city and state and depending on the role) if an authorization is granted to work outside of the location listed in this posting.
Our client is an equal opportunity employer and does not discriminate in employment based on race religion (including religious dress and grooming practices) color sex/gender (including pregnancy childbirth breastfeeding or related medical conditions) sex stereotype gender identity/gender expression/transgender (including whether or not you are transitioning or have transitioned) and sexual orientation; national origin (including language use restrictions and possession of a drivers license issued to persons unable to prove their presence in the United States is authorized under federal law Vehicle Code section 12801.9); ancestry physical or mental disability medical condition genetic information/characteristics marital status/registered domestic partner status age (40 and over) sexual orientation military or veteran status or any other basis protected by federal state or local law or ordinance or regulation.