Overview
As a Software Engineer in the AI Libraries team, you will build and maintain scalable Python libraries and tools that empower Wayve’s ML engineers and researchers. You’ll design robust abstractions for data loading, distributed training, inference, and evaluation workflows, enabling efficient large-scale ML. You’ll support training at scale on GPU clusters and cloud infra, ensuring reliable, well-documented, and observable tooling. You’ll contribute to maturing Wayve’s AI platform to advance autonomous driving technology for customers. This role combines hands-on engineering with a mission to scale robust ML infrastructure in a fast-paced, collaborative environment.
Responsibilities
- Design, build, and maintain scalable Python libraries and tools used by ML engineers and researchers
- Develop robust abstractions for data loading, distributed training, inference, checkpointing, and model evaluation workflows
- Support training at scale across large GPU clusters and cloud-based infrastructure
- Collaborate with ML teams to understand user needs and create reliable, well-documented, observable, and adoptable tools
- Improve engineering quality across ML systems via architecture, testing, monitoring, and maintainability practices
- Optimize data and training pipelines for multi-modal data (camera, radar, lidar, etc.)
- Contribute to evolving Wayve’s AI platform as autonomous driving capabilities scale
Key requirements
- Strong Python programming experience
- Proven experience designing, building, and maintaining software systems from concept through delivery
- Strong software architecture and system design skills
- Experience building tools, platforms, or libraries for internal or external users
- Strong understanding of testing, observability, maintainability, and engineering best practices
- Experience with cloud environments, ideally Azure
- Experience with concurrent, parallel, or distributed computing
- Familiarity with ML frameworks such as PyTorch, TensorFlow, or PyTorch Lightning
- Ability to work with technical stakeholders to refine requirements and deliver scalable solutions
- Collaborative communication
- Customer-focused empathy for ML teams
- Problem-solving mindset
- Azure cloud
- Distributed computing
- PyTorch/TensorFlow/PyTorch Lightning
…
