Overview
In this role you will lead the development of document understanding capabilities that power Thomson Reuters’ legal AI platform. You’ll advance semantic chunking, document enrichment, classification, information extraction, and knowledge graph construction across vast legal, tax, and accounting content. You’ll shape production-grade AI systems and enable downstream search, retrieval, and agentic workflows. This is a high-impact leadership role that combines research excellence with scalable, real-world delivery.
Pay / Benefits
- hybrid work model
- flexible work balance policies (Flex My Way)
- growth and career development programs
- comprehensive benefits including mental health days, Headspace, retirement savings, tuition reimbursement
- employee incentive programs
- ESG initiatives and social impact opportunities
Responsibilities
- Design and deploy semantic chunking for lengthy legal/tax documents
- Build document enrichment pipelines to identify types, jurisdictions, entities, and metadata
- Develop hierarchical and multi-label document classification using standard and custom taxonomies
- Create LLM-based and traditional NLP extraction pipelines for entities, relationships, citations, and concepts
- Develop knowledge graph construction to connect and enrich entities and citations across large collections
- Handle tabular data understanding within complex documents
- Create document intelligence capabilities for search, retrieval, RAG, and agentic workflows
- Design evaluation frameworks using expert annotations, synthetic data, and production metrics
- Lead technical decisions on architectures, chunking, extraction, classification, and knowledge representation
- Partner with engineering teams to deliver scalable, reliable AI systems
- Provide technical leadership and roadmap input; mentor applied scientists and ML practitioners
Key requirements
- PhD with demonstrable post-degree industry experience developing and deploying document understanding systems at scale
- Hands-on depth in document analysis, information extraction, classification, knowledge representation, evaluation, and production deployment
- Experience moving NLP/AI capabilities from research to production and translating unstructured content into structured knowledge
- Collaborative mindset with ability to mentor and drive production success
- Experience building production document understanding systems that exceed basic OCR or parsing and deliver measurable business impact
- Understanding of transforming large document collections into structured knowledge assets and constructing knowledge graphs from real-world content
- collaborative mindset
- mentoring
- production-focused mindset
- NLP
- information extraction
- classification
…
