Overview
As a Senior ML Engineer, you will design and operate production-grade ML systems powering an AI product focused on long-running workflows and decision-making. You’ll own end-to-end ML pipelines, from training to inference, with emphasis on reliability and real-world performance. You’ll work with LLM-based systems and agent-style workflows, optimizing for latency, cost, and stability. This role offers ownership, fast iteration, and the chance to shape how production AI behaves under real usage. You’ll collaborate with a small, capable team in a remote-first environment to deliver robust, scalable ML solutions.
Pay / Benefits
- remote work
- flexible compensation
- equity
Responsibilities
- Build and operate training, inference, and evaluation pipelines
- Develop LLM-based systems and agent-style workflows
- Debug model behaviour using real-world signals
- Optimize performance across latency, cost, and reliability
- Implement production monitoring, logging, and system stability
- Own end-to-end pipelines from training to inference
- Conduct continuous evaluation and iterative improvement
- Ensure systems perform reliably under real usage conditions
Key requirements
- Shipped ML systems used by real users
- Owned ML systems end-to-end in production
- Experience with modern LLMs beyond simple API integration
- Exposure to model behaviour debugging, latency, throughput, and cost optimization
- Experience with monitoring, observability, and evaluation frameworks
- Experience scaling ML systems in production environments
- ownership
- pragmatic decision-making
- fast iteration
- Python
- PyTorch
- LLM ecosystem (OpenAI, Anthropic, etc.)
…
