Overview
In this hands-on role, you will design and refine models that simulate compute, memory, interconnect, and communication behavior for large-scale ML systems. You’ll build tools to simulate workloads across distributed accelerator clusters and analyze end-to-end performance at scale. You’ll collaborate with hardware, software, networking, and ML teams to translate findings into design recommendations, shaping scalable ML infrastructure. This is a research-driven, technically deep position with real impact on architecture decisions at scale.
Responsibilities
- Build simulation models for compute, memory, interconnect, and communication behavior in large-scale ML systems
- Develop tools to simulate training and inference workloads across distributed accelerator clusters
- Model distributed execution patterns including collectives, synchronization, and communication bottlenecks
- Run experiments and benchmarks on real ML systems to calibrate and validate models
- Analyze end-to-end performance metrics: throughput, latency, scaling efficiency, and cost/performance tradeoffs
- Collaborate with hardware, software, networking, and ML teams to communicate findings through design recommendations
Key requirements
- Master’s or PhD in CS, Electrical or Computer Engineering, or related field
- Strong background in ML systems, distributed systems, performance engineering, or simulation
- Experience analyzing compute, communication, and memory behavior in large-scale ML systems
- Hands-on benchmarking, profiling, and measurement of ML systems
- Familiarity with distributed training concepts: data/tensor/pipeline parallelism, collectives, synchronization
- Proficiency in Python, C++, or Rust
- cross-functional collaboration
- strong communication of complex findings
- problem-solving and analytical thinking
- ML systems
- distributed systems
- performance engineering
…
