Senior ML Systems Engineer – GPU & HPC Optimizations

Company: Sponsor Finder
Apply for the Senior ML Systems Engineer – GPU & HPC Optimizations
Location: Boston
Job Description:

Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for state-of-the-art models. You will profile PyTorch, write custom kernels in CUDA/Triton, and leverage compilers like torch.compile, TensorRT, and XLA to maximize throughput on large GPU clusters.

You’ll scale distributed training across 32–64 nodes on GCP, manage memory scaling for large complexes, and build benchmarks and profiling tools for the research team.

#J-18808-Ljbffr…

Posted: October 4th, 2026