AI Inference Engineer (Staff) – GPU ML Systems + Equity

Company: Perplexity
Apply for the AI Inference Engineer (Staff) – GPU ML Systems + Equity
Location: London
Job Description:

Perplexity is hiring an AI Inference Engineer to scale our inference engine behind every query. You will work on loading weights, scheduling requests, and managing KV-cache, with a stack of Rust, Python, CUDA, and CuTe DSL.

You will port CUDA kernels to CuTe DSL, optimize performance under tight latency and cost budgets, and build a robust Rust-based serving runtime to support growing traffic and model architectures.

#J-18808-Ljbffr…

Posted: September 25th, 2026