Overview
As a Data Scientist III in Elsevier’s Platform Data Science group, you will design, build, and evaluate AI-powered capabilities for LeapSpace and the Search & AI Platform. You will work with cross-functional teams to translate cutting-edge AI research into production-ready workflows for researchers worldwide. You will tackle applied AI, NLP, and information retrieval challenges, enabling next-generation AI-assisted scientific discovery. This role offers hands-on experimentation, rigorous evaluation, and meaningful impact on how researchers access trusted content and insights.
Pay / Benefits
- Comprehensive Pension Plan
- Home, office, or commuting allowance
- Generous vacation entitlement and sabbatical leave
- Maternity, Paternity, Adoption and Family Care leave
- Flexible working hours
- Personal Choice budget
Responsibilities
- Develop and improve LLM-powered research workflows (QA, literature summarization, semantic discovery, citation-aware retrieval)
- Build agentic multi-step AI workflows using LangGraph and orchestration tools
- Apply NLP, Generative AI, embeddings, retrieval, and RAG techniques
- Evaluate emerging AI models and contribute experimentation recommendations
- Contribute to prompt engineering, grounding, context management, and hallucination mitigation
- Integrate scientific metadata, ontologies, and knowledge assets into AI workflows
- Design, develop, and optimize search and retrieval pipelines (lexical, vector, hybrid)
- Enhance RAG systems with trusted scientific content; experiment with embeddings and ranking strategies
- Collaborate with engineering teams to deploy and scale AI-powered solutions
- Develop evaluation frameworks and datasets; conduct offline and online experiments; report results
Key requirements
- Master’s degree in Computer Science, Data Science, ML, NLP, IR, or related field
- Hands-on experience with LLM-based applications, RAG pipelines, and retrieval systems
- Strong Python programming
- Experience with PyTorch, Hugging Face, LangChain, LangGraph, Haystack
- Experience with Databricks or similar distributed ML platforms
- Knowledge of IR and AI evaluation methods and statistical analysis
- Data visualization and analytical tooling proficiency
- Cross-functional collaboration
- Clear technical communication to both technical and non-technical audiences
- Independent project execution
- LLM-based applications and generative AI systems
- RAG pipelines and retrieval systems
- Search and retrieval architectures (lexical, vector, hybrid)
…
