Overview
As an AI Research Engineer (Pre-training) on the HAIL team, you will advance large-scale model training for market-focused foundation models. You’ll collaborate with researchers to co-design, optimize training systems, and scale GPUs across busy clusters. This role blends deep engineering with research to push performance, fault tolerance, data pipelines, and parallelism in a demanding, finance-oriented environment. You’ll impact trading outcomes by improving the core infrastructure that drives state-of-the-art models.
Pay / Benefits
- discretionary performance-based bonuses
- competitive benefits package
- salary range 250000-300000 USD per year
- global offices
- equal opportunity employer
Responsibilities
- Improve large-scale model training end-to-end (kernel development, data loading, parallelism, networking, fault-tolerance)
- Collaborate with researchers to co-design and enhance models and research agenda
- Manage and optimize GPU clusters for high-performance training
- Advance systems capabilities across kernel, framework, and toolchains (CUDA, PyTorch/JAX/XLA, CUDA Graphs)
- Develop and maintain scalable training pipelines and fault-tolerant workflows
- Contribute to a culture of experimentation and rapid iteration to squeeze performance from systems
Key requirements
- Strong engineering skills in CUDA/Triton/Pallas/CuTe DSL kernel development
- Lower-level PyTorch/JAX/XLA development experience
- CUDA Graphs or FPGA/ASIC experience
- Two+ years building deep learning systems across domains
- Experience translating methods between applications; LLM experience valued but not required
- Finance experience not required
- Collaborative mindset
- Strong problem-solving and cross-disciplinary communication
- Adaptability in a fast-growing environment
- CUDA/Triton/Pallas/CuTe DSL kernel development
- Lower-level PyTorch/JAX/XLA development
- CUDA Graphs
…
