Find and remove bottlenecks across model representation, GPU kernels, runtimes, and serving for open-weight models.
Why this role exists
We are building an inference research team to find where current hardware and software waste time, memory, and energy when serving open-weight models.
You will measure those bottlenecks, test changes across the inference stack, and turn useful results into production systems. Your findings will also inform our silicon work.
What you will do
- Measure target models and workloads on current NVIDIA and AMD hardware.
- Design and implement kernels, numerical formats, memory layouts, and cache policies.
- Change runtimes, schedulers, distributed execution, and serving paths when they limit the system.
- Test latency, throughput, energy use, model quality, correctness, and reliability on real workloads.
- Turn research results into production services.
- Share research publicly so that others can learn about how inference research is done.
- Work with the silicon team to identify hardware bottlenecks and design opportunities.
Problems you may work on
- Kernels and fused execution paths for exact model shapes.
- Low-bit formats that preserve model quality.
- KV-cache layout, movement, compression, and reuse.
- Sparse and mixture-of-experts execution.
- Speculative and parallel decoding.
- Prefill and decode scheduling across multiple GPUs.
- Performance models that explain the gap between hardware limits and measured results.
- Runtime designs for long-running agent workloads.
What we are looking for
- You have built an ML systems, compiler, runtime, or kernel project, and you can explain your decisions and results in detail.
- You can work in Python and a systems language such as C++ or Rust. We use C++ mostly.
- You can write CUDA or HIP kernels, or you have strong low-level systems experience and can learn GPU programming.
- You understand transformer execution, including attention, memory movement, and numerical precision.
- You use profiling and controlled experiments to identify system limits.
- You are open to learning about hardware, and what is limiting the execution of the model.
- You can maintain production code, not only experimental code.
- You can explain hard technical solutions clearly
This role is not for you if
- You want to work within one layer of the inference stack.
- You want framework integration to be the main technical work.
- You want papers or notebook results to be the primary output.
- You aren’t excited by rewriting things from scratch, and taking time to find the globally optimal solution to a problem.
- You need a complete specification before you can investigate a problem.
#J-18808-Ljbffr…
