Overview
In this role, you will design and build ML, NLP, and generative AI solutions to accelerate scientific discovery and knowledge extraction. You’ll work with large-scale, heterogeneous scientific content and collaborate across engineering, product, UX, analytics, and domain experts to deliver production-ready systems. You will translate ambiguous research challenges into measurable, data-driven outcomes that improve how researchers find and use knowledge. This position offers the chance to shape AI-enabled discovery at a trusted, collaborative organization with a focus on quality and impact.
Pay / Benefits
- flexible working hours
- wellbeing initiatives
- shared parental leave
- study assistance
- sabbaticals
- career development opportunities
Responsibilities
- Design and build ML, NLP, and generative AI systems for discovery, knowledge extraction, decision support, and content understanding
- Work with large-scale data including publications, datasets, knowledge graphs, ontologies, and metadata
- Apply a range of techniques (classification, regression, clustering, ranking, feature engineering, deep learning, embeddings, LLMs, retrieval, generative AI)
- Develop capabilities for semantic search, information retrieval, entity extraction, content classification, recommendation, ranking, summarization, QA, and evidence-grounded generation
- Build, evaluate, fine-tune, prompt, and integrate models into production systems with ongoing quality improvements
- Write clean, tested Python code and develop reusable data science components, pipelines for preprocessing, inference, experimentation, monitoring, and CI/CD
- Support deployment, monitoring, model maintenance, drift detection, automated retraining, and optimization
- Collaborate with cross-functional teams and communicate model behavior, insights, and trade-offs to technical and non-technical audiences
Key requirements
- Experience in data science, ML, AI, NLP, statistics, applied mathematics, computer science, or related quantitative area
- Experience with frontier LLMs (e.g., GPTs, Claude, Gemini), including fine-tuning LLMs/SLMs
- Strong Python skills and well-tested code
- Solid grasp of ML fundamentals (supervised/unsupervised learning, feature engineering, evaluation, selection, performance)
- Experience with structured, semi-structured, or unstructured data, especially large-scale text or content datasets
- Familiarity with Pandas, NumPy, SciPy, Scikit-learn, PyTorch, TensorFlow, or Matplotlib
- Ability to translate complex requirements into practical, data-driven solutions with strong analytical thinking and attention to quality
- Clear communication and collaborative mindset with stakeholders to deliver production-ready value
- clear communication
- collaboration
- analytical thinking
- Pandas
- NumPy
- SciPy
…
