Overview
In this role you will advance NVIDIA’s inference and compiler stack by mapping neural workloads onto our platforms. You’ll work at the intersection of large-scale systems, compilers, and deep learning to optimize end-to-end inference performance. You will collaborate with hardware teams to shape future architectures and publish results at top venues. This is a hands-on position focused on practical, scalable techniques for real-world deployments.
Responsibilities
- Build and maintain high-performance runtime and compiler components for end-to-end inference optimization
- Define mappings of large-scale inference workloads onto NVIDIA systems
- Extend NVIDIA’s software ecosystem with libraries, tooling, and interfaces for cross-platform model deployment
- Benchmark, profile, and monitor performance to ensure efficient graph-to-hardware mappings
- Collaborate with hardware architects to feedback software observations and influence architectures
- Prototype and evaluate new compilation and runtime techniques for spatial processors
- Publish and present technical work at ML, compiler, and computer architecture venues
Key requirements
- MS or PhD in CS, ECE, or related field, or equivalent with 5+ years of experience
- Strong software engineering in systems programming (C/C++ and/or Rust)
- Hands-on compiler or runtime development experience (IR design, optimization passes, code generation)
- Experience with LLVM and/or MLIR (building passes, dialects, integrations)
- Familiarity with TensorFlow, PyTorch, and ONNX for graph formats
- Solid understanding of parallel and heterogeneous architectures (GPUs, spatial accelerators)
- Strong analytical and debugging skills using profiling and benchmarking tools
- Excellent communication and collaboration across hardware, systems, and software teams
- strong collaboration
- clear communication
- problem-solving mindset
- LLVM
- MLIR
- IR design
…
