Overview
In this role you will design and deploy end-to-end AI solutions for complex legal document understanding, enabling scalable search, extraction, and reasoning across vast legal content. You will work across product teams to build semantic chunking, document enrichment, and knowledge-graph pipelines that underpin production systems. Your work will directly impact how legal professionals research and analyze documents, with production-ready innovations and scalable ML solutions. You will join a collaborative, research-driven environment that values cutting-edge NLP and responsible deployment.
Pay / Benefits
- hybrid work model
- flexible vacation
- Mental Health Days off
- Headspace app
- retirement savings
- tuition reimbursement
Responsibilities
- Design, build, test, and deploy end-to-end AI solutions for complex legal document understanding
- Develop models for semantic chunking of lengthy legal documents with adjustable granularity
- Build document enrichment systems to classify documents and extract metadata
- Create LLM-based knowledge graph construction pipelines linking entities, citations, and concepts
- Develop scalable synthetic data generation for model training and QA of legal queries
- Collaborate with Engineering to ensure reliable, scalable delivery in production
- Develop data/evaluation strategies and balance model performance with latency
- Lead architectural decisions for document understanding, taxonomy handling, and multi-hop reasoning
- Align with stakeholders across product lines to translate requirements into scalable solutions
- Maintain expertise and contribute to research publications and IP
- Explore knowledge distillation and efficiency improvements for deployment
Key requirements
- Hands-on experience building and deploying document understanding systems or knowledge graphs using deep learning, LLMs and NLP methods
- Ability to translate complex problems into AI applications balancing accuracy and efficiency
- Professional experience scaling projects and leading others in an applied research setting
- Strong programming skills (Python) and experience with PyTorch, Hugging Face Transformers, DeepSpeed
- Publications at ACL, EMNLP, ICLR, NeurIPS, SIGIR, KDD
- Demonstrated ability to collaborate with engineering and product teams to deliver production-ready solutions
- collaboration
- strong communication
- leadership and influence
- Deep understanding of document understanding fundamentals (layout analysis, semantic chunking, taxonomy handling)
- Entity recognition and linking, relation extraction, citation parsing, graph representations from text
- LLM-based information extraction, few-shot and multi-task learning, post-training and knowledge distillation
…
