Overview
In this role you will advance large-scale model inference for HRT’s market-facing models within HAIL. You’ll collaborate with researchers to co-design and optimize inference across diverse devices and setups deployed globally, pushing the boundaries of performance and efficiency. The work combines systems engineering with cutting-edge ML, delivering tangible impact on trading outcomes. A strong hook is the chance to shape foundation-model-in-market capabilities at a fast-moving, research-driven trader.
Pay / Benefits
- discretionary performance-based bonuses
- competitive benefits package
Responsibilities
- Improve all aspects of large-scale model inference, including GPU kernel development
- Work on novel inference devices (ASICs, FPGAs) and data streaming
- Co-design with researchers to improve models and shape the research agenda
- Maintain and optimize an inference solution deployed across devices worldwide
- Tackle challenging, high-impact problems in a fast-evolving domain
Key requirements
- Two or more years of experience building deep learning systems in any domain
- Strong engineering skills in CUDA/Triton/Pallas/CuTe DSL kernel development or lower-level PyTorch/JAX/XLA development
- Experience with CUDA Graphs and FPGA/ASIC experience
- LLM experience is valuable but not required
- Finance experience not required
- CUDA/Triton/Pallas/CuTe DSL kernel development
- lower-level PyTorch/JAX/XLA development
- CUDA Graphs
- FPGA/ASIC experience
- deep learning systems engineering
…
