Overview
In this Senior Lead role, you shape secure, scalable platform capabilities for enterprise AI and data platforms. You will set tooling and runtime standards, focusing on production-grade LLM inference and Kubernetes deployment patterns. You’ll partner with engineering teams to improve reliability, developer experience, and operational readiness, while mentoring others. This is a hands-on, impact-focused position that blends deep systems work with cross-team collaboration and governance. You’ll contribute to a mission of enabling safe, scalable AI at scale across the organization.
Responsibilities
- Design and deliver platform standards and tooling to ease adoption (CLI, SDKs, libraries, templates, automated checks)
- Build and operate production LLM inference services using modern serving engines
- Drive Kubernetes-based deployment patterns, scaling, networking, and troubleshooting for reliable platform ops
- Optimize inference performance with GPU memory considerations and bottleneck analysis
- Assess inference-time quantization trade-offs for real-world workloads
- Develop secure, high-quality production code and automation for resiliency and observability
- Maintain architecture/design artifacts and enforce non-functional requirements through automation
- Promote enterprise AI-assisted engineering practices to improve quality, speed, and operational outcomes
- Leverage SDLC tools and enterprise AI-assisted development to enhance automation and value
Key requirements
- Hands-on experience building standards and tooling to improve platform adoption
- Deep, hands-on experience with LLM inference systems (e.g., vLLM, TensorRT-LLM, SGLang, LLM-D)
- Hands-on experience operating production services on AWS
- Ability to design, deploy, and troubleshoot AWS cloud infrastructure components for platform services
- Demonstrated Kubernetes expertise across deployments, scaling, networking, and troubleshooting
- Knowledge of GPU memory architecture, including key-value cache sizing and performance trade-offs
- Understanding of inference-time quantization and its impact on latency and throughput
- Ability to translate architecture into secure, scalable implementations
- Strong SDLC, CI/CD, resiliency, and security practice understanding
- Experience with enterprise AI-assisted software development tools and evaluating AI outputs for correctness, performance, and security
- Understanding of responsible AI use, data sensitivity, secure inputs/outputs, and resiliency/security expectations
- collaborative and cross-functional teamwork
- mentorship and coaching
- problem-solving mindset with focus on reliability and developer experience
- vLLM
- TensorRT-LLM
- SGLang
…
