Overview
In this role you will help build a robust, scalable bioinformatics data platform to power AI-driven drug discovery. You collaborate with ML Research, Computational Biology, Drug Development, and Chemistry teams to turn raw biological data into high-quality, standardized datasets for model training. You’ll harmonize public and internal data sources and contribute to a production-ready bioinformatics platform that accelerates research programs. This is a mission-driven role in a cross-disciplinary, innovation-focused culture.
Responsibilities
- Develop and operate large-scale bioinformatics pipelines from raw data to ML-ready datasets
- Ingest and harmonize complex datasets with best practices for data quality and versioning
- Harmonise public databases (Ensembl, UniProt, Reactome, Open Targets) with robust mapping to prevent data issues
- Partner with ML Research, Computational Biology, Drug Development, and Chemistry teams to promote standardized data primitives
- Contribute as a Deployed Engineer on research projects, delivering tailored solutions and feeding insights back to the core platform
- Provide documentation, guidance, and training on data resources and curation processes
Key requirements
- Proven experience with large-scale processing of raw bioinformatics data (e.g., FASTQ, BAM, mzXML)
- Track record of delivering bioinformatics outputs across modalities (genomics, proteomics, functional genomics, systems biology, single cell)
- Experience delivering bioinformatics solutions to research teams or industry projects with a focus on user enablement
- Production-grade coding in Python and building automated, scalable bioinformatics pipelines
- PhD or MSc in Bioinformatics, Computational Biology, or related field, or equivalent practical experience
- Cross-functional collaboration
- User enablement and communication to non-specialists
- Problem-solving and proactive initiative
- Python (production-grade)
- Bioinformatics pipelines
- Nextflow (Nice to have)
…
