Overview
Join Kraken’s dedicated AI Compute and Infrastructure team to power next-gen AI workloads at scale. You will own GPU and accelerator clusters, shaping cost-efficient, production-grade infrastructure used by AI researchers and engineers. You’ll partner with ML researchers, platform engineers, and security teams to remove bottlenecks and improve reliability. This role offers high impact in a mission-driven, remote-friendly company building premier crypto products.
Responsibilities
- Own and operate GPU and accelerator clusters for training, inference, evaluation, and experimentation, including drivers, runtimes, kernels, device plugins, node configuration, scheduling primitives, and workload isolation
- Design infrastructure to run models locally on GPUs where strategic and economical, reducing external dependencies and compute costs
- Build and improve scheduling, orchestration, placement, quota management, and utilization systems across heterogeneous accelerators
- Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost using frameworks like vLLM, Triton Inference Server, TensorRT
- Collaborate with ML engineers and researchers to remove bottlenecks in training, evaluation, deployment, and production debugging workflows
- Build observability for GPU utilization, memory pressure, queue depth, saturation, token throughput, request latency, and spend
- Drive reliability, incident response, alerting, runbooks, and post-incident improvements for always-on AI compute infrastructure
- Evaluate and integrate new hardware, cloud instances, accelerators, runtimes, schedulers, and serving frameworks as the landscape evolves
- Build tooling that makes GPU usage visible and approachable for internal teams
- Contribute to long-term architecture decisions balancing performance, cost, scalability, and safety
- Shape architecture decisions to balance performance, cost efficiency, scalability, operational simplicity, and production safety
Key requirements
- 5+ years of infrastructure engineering with GPU compute, ML infrastructure, distributed systems, HPC, or large-scale production platforms
- Hands-on experience operating GPU clusters or accelerator-backed infrastructure in production environments
- Strong systems engineering fundamentals across Linux, networking, storage, containers, Kubernetes, distributed runtimes, and production debugging
- Experience with ML serving frameworks such as vLLM, Triton Inference Server, TensorRT, TorchServe, KServe, Ray Serve, or equivalents
- Proficiency in Python for infrastructure automation and tooling
- Understanding of performance tradeoffs across batching, concurrency, memory usage, GPU utilization, model size, latency, throughput, availability, and cost
- Track record of optimizing compute costs while maintaining performance and reliability
- Experience building observable systems with metrics, logs, traces, dashboards, alerts, and incident workflows
- Comfortable in high-stakes, always-on environments with emphasis on uptime and reliability
- Clear communicator capable of explaining infrastructure tradeoffs to researchers, product teams, and leadership
- clear communicator
- collaboration with cross-functional teams
- operational discipline
- GPU compute
- ML infrastructure
- Linux, networking, storage, containers, Kubernetes
…
