Research Data Specialist
Job Summary
Role purpose
Lead and execute the management integration and architecture of biological data systems and computational infrastructure within the Bioinformatics group to support Seeds R&D in China. The role is responsible for structuring storing and governing diverse high-throughput biological experimental data as well as managing analytical codebases across high-performance computing (HPC) clusters and modern cloud environments. Working closely with bioinformaticians trait discovery scientists functional genomics teams and digital partners the Research Data Specialist will optimize data flows build scalable storage solutions and ensure high-performance computing systems are robustly architected to accelerate data-driven crop biotechnology research.
Accountabilities
- HPC & Cloud Platform Operations. Manage configure and optimize high-performance computing (HPC) cluster environments and cloud data warehouses to ensure seamless storage retrieval and analysis of large-scale biological datasets. Monitor system performance compute resources and data storage costs to optimize resource utilization.
- Biological Data & Code Lifecycle Management. Systematically manage curate and version biological experimental data (such as genomic RNA-seq data) along with accompanying bioinformatics software scripts and pipelines. Establish standardized code repositories version control workflows and documentation best practices for research teams. Pipeline Support & Integration. Partner closely with bioinformatics pipeline developers data scientists and experimental research teams to streamline data ingestion preprocessing and automated analysis workflows. Ensure raw data from multi-omics platforms and experimental assays are seamlessly transferred validated and formatted for analytical consumption.
- Cross-Functional Collaboration & Technical Training. Collaborate with local and global IT AI Engineering and Bioinformatics teams to align data storage and computing standards. Provide training and operational support to research scientists regarding data submission protocols cloud usage HPC Job Scheduling and script versioning. Continuously evaluate emerging cloud storage solutions distributed computing frameworks and database architectures (e.g. Snowflake AWS cloud-native services graph databases) to modernize Syngentas biological data infrastructure and maintain technological edge.
Qualifications :
Knowledge Experiences & Capabilities
- Knowledge:
- Strong knowledge of modern data storage data warehousing and cloud computing platforms (e.g. AWS S3/EC2 Snowflake or equivalent enterprise systems).
- Deep familiarity with Linux/Unix operating systems high-performance computing (HPC) cluster environments (e.g. Slurm SGE) and job scheduling mechanics.
- Understanding of database systems (SQL) and data engineering workflows for managing structured and unstructured data.
- Basic understanding of popular high-throughput biological data types. Backgrounds of plant biology biotechnology research and biological experimental workflows is preferred.
- Education and Experience
- Ph.D. or Masters degree in Bioinformatics Computational Biology Computer Science Data Engineering Information Technology or a related quantitative discipline.
- 3 years (for Ph.D.) or 5 years (for Masters) of experience in managing high-throughput data HPC clusters or enterprise cloud data architectures in academic or industry research environments.
- Proven track record of managing large-scale datasets building data pipelines and establishing code repository architectures.
- Demonstrated experience working on cloud infrastructure (Snowflake or equivalent platforms).
- Experience in agriculture biotechnology life sciences or related research-intensive industries is preferred.
- Capabilities
- Proficiency in Python Bash/Shell scripting and SQL with strong software development best practices (version control containerization documentation).
- Ability to architect migrate and optimize data workflows across HPC cluster servers and modern cloud platforms (AWS Snowflake).
- Strong problem-solving skills with a focus on data governance data integrity and reproducible computational research.
- Ability to communicate complex technical cloud and computational infrastructure concepts effectively to experimental scientists and non-technical stakeholders.
- Demonstrated ability to establish productive cross-functional collaborations across bioinformatics software engineering and laboratory teams.
- Excellent written and spoken English
Additional Information :
Note: Syngenta is an Equal Opportunity Employer and does not discriminate in recruitment hiring training promotion or any other employment practices for reasons of race color religion gender national origin age sexual orientation gender identity marital or veteran status disability or any other legally protected status.
Remote Work :
No
Employment Type :
Full-time
About Company
To help feed 10 billion people while reducing emissions and improve biodiversity. This is our mission as the global agriculture technology leader. With 59,000 employees in more than 100 countries and hundreds of thousands of agricultural partners worldwide, we are committed to transfo ... View more