Overview
In this role you will design and operate production-grade AI infrastructure at enterprise scale, focusing on reliable LLM serving, cloud and Kubernetes deployments, and deep observability. You’ll own end-to-end reliability, performance, and cost-efficiency of AI inference platforms, partnering across teams to improve stability and secure software delivery. You’ll solve hard production problems and drive continuous improvement in a fast-paced, risk-aware environment. This role offers the opportunity to shape AI infrastructure that modernizes traditional SRE practices.
Responsibilities
- Design, develop, and deliver secure, production-quality software for AI infrastructure
- Build backend services and APIs enabling reliable AI infrastructure in production
- Operate and scale LLM serving infrastructure including hosting, routing, batching, and caching
- Deploy and lifecycle-manage LLMs on cloud and on-prem GPUs using IaC and CD pipelines
- Implement observability across logs, metrics, and traces with dashboards and alerting
- Tune GPU capacity, autoscaling, and cost-efficiency using optimization techniques
- Lead reliability engineering for LLM endpoints including capacity planning and incident response
- Participate in on-call rotations, incident triage, and post-incident analyses
- Identify recurring issues and automate remediation to improve stability and developer experience
- Build and maintain multi-agent systems with orchestration and workflow control
- Drive adoption of AI-assisted engineering practices with standardized validation and code quality processes
Key requirements
- Hands-on system design, development, testing, and operations in production
- Advanced Python proficiency for production services
- Automation and continuous delivery experience
- Cloud infrastructure and IaC tooling experience
- Strong site reliability engineering knowledge (incident management, runbooks, reliability patterns)
- Observability experience across metrics, logs, and traces
- Kubernetes and container orchestration experience
- Experience hosting and serving LLMs on cloud and local GPU environments
- Knowledge of LLM reliability and risk considerations (latency, throughput, versioning, logging)
- Experience with enterprise AI-assisted software development tools and evaluating AI outputs for correctness and security
- Understanding of responsible AI use in engineering workflows and security expectations
- Collaborative mindset
- Strong incident management and communication
- Ownership and drive for reliability improvements
- Python production services
- Automation and CD/CI pipelines
- Cloud platforms and IaC tooling
…
