Overview
In this role you will own the end-to-end production lifecycle of AI, from training and fine-tuning domain LLMs to building scalable inference pipelines. You will write clean, Python-based code and deliver reliable, low-latency AI solutions for voice and text applications. You’ll work with cross-functional partners to translate business needs into robust ML systems, embedded in containerized microservices. This is a hands-on role at scale, combining MLOps, performance engineering, and AI delivery to impact real-world use cases.
Pay / Benefits
- Competitive salary
- pension scheme
- discretionary bonus
- 25 days annual leave + 5
- hybrid working options
- Private Medical cover
Responsibilities
- Lead domain-specific LLM adaptation using LoRA, QLoRA, and PEFT to balance performance and resources
- Architect and optimize inference pipelines to reduce Time to First Token and increase throughput (quantization, caching, batching)
- Build and maintain real-time AI pipelines using WebSockets and SSE for voice (ASR/TTS) and text apps
- Deploy and orchestrate models within Docker/Kubernetes-based microservices with monitoring, security, and scalability
- Collaborate with Business Analysts and stakeholders to align commercial needs with technical execution
Key requirements
- 5+ years in AI/ML engineering with production-ready model deployment experience
- Deep proficiency in Python with clean coding, modular design, and testing
- Hands-on experience with LLM training cycles, PEFT, and prompt engineering
- Experience with high-performance inference servers (vLLM, TGI, or Triton) and GPU deployment optimization
- Linux-based environment expertise and proficiency with containerized workflows and CI/CD
- Experience building Retrieval-Augmented Generation (RAG) systems, including vector databases and semantic search
- collaboration with cross-functional teams
- effective communication with stakeholders
- problem-solving under performance constraints
- LoRA, QLoRA, PEFT for model fine-tuning
- inference optimization (TTFT, throughput)
- quantization, caching strategies, and efficient batching
…
