Overview
In this role you will engineer self-service performance tools that span models, hardware targets, and workflows to enable faster, cheaper AI inference and training. You’ll shape measurement standards and predict the impact of changes across the AI stack, working with cross-functional teams to turn profiling data into actionable recommendations. The position sits in Wayve’s AI Performance Tooling team, focused on data-driven performance decisions and operationalizing performance insights at scale. This is a hands-on role with real impact on efficiency, latency, and cost of AI workloads.
Responsibilities
- Design and build scalable performance tooling for multiple models and hardware targets
- Define and standardize metrics for measuring and predicting AI performance
- Model theoretical peak versus achieved performance and identify bottlenecks at layer and operation levels
- Forecast latency, memory, utilization, and compute cost for model or recipe changes
- Own monitoring and regression alerting for model builds and training runs
- Collaborate with training and runtime engineers to set performance targets and justify decisions with data
Key requirements
- Deep hands-on performance engineering in complex systems (profiling, roofline analysis, latency/throughput optimization)
- Proven track record of end-to-end tool or service ownership
- Strong Python skills and ability to profile/instrument large production codebases
- Hands-on experience with deep learning models in PyTorch
- Data analysis skills to turn noisy measurements into defendable conclusions
- Strong judgment to turn ambiguous performance questions into measurable ones
- Quantitative communication that can influence cross-team priorities
- Judgment and prioritization
- Data-driven decision making
- Cross-team collaboration
- Profiling and roofline analysis
- Latency and throughput optimization
- Python programming
…
