AI Engineer (Fluent in Mandarin & English)

Company: Chubb
Apply for the AI Engineer (Fluent in Mandarin & English)
Location: London
Job Description:

Overview

In this role you will own the end-to-end production lifecycle of AI, from training and fine-tuning domain LLMs to building scalable inference pipelines. You will write clean, Python-based code and deliver reliable, low-latency AI solutions for voice and text applications. You’ll work with cross-functional partners to translate business needs into robust ML systems, embedded in containerized microservices. This is a hands-on role at scale, combining MLOps, performance engineering, and AI delivery to impact real-world use cases.

Pay / Benefits

  • Competitive salary
  • pension scheme
  • discretionary bonus
  • 25 days annual leave + 5
  • hybrid working options
  • Private Medical cover

Responsibilities

  • Lead domain-specific LLM adaptation using LoRA, QLoRA, and PEFT to balance performance and resources
  • Architect and optimize inference pipelines to reduce Time to First Token and increase throughput (quantization, caching, batching)
  • Build and maintain real-time AI pipelines using WebSockets and SSE for voice (ASR/TTS) and text apps
  • Deploy and orchestrate models within Docker/Kubernetes-based microservices with monitoring, security, and scalability
  • Collaborate with Business Analysts and stakeholders to align commercial needs with technical execution

Key requirements

  • 5+ years in AI/ML engineering with production-ready model deployment experience
  • Deep proficiency in Python with clean coding, modular design, and testing
  • Hands-on experience with LLM training cycles, PEFT, and prompt engineering
  • Experience with high-performance inference servers (vLLM, TGI, or Triton) and GPU deployment optimization
  • Linux-based environment expertise and proficiency with containerized workflows and CI/CD
  • Experience building Retrieval-Augmented Generation (RAG) systems, including vector databases and semantic search
  • collaboration with cross-functional teams
  • effective communication with stakeholders
  • problem-solving under performance constraints
  • LoRA, QLoRA, PEFT for model fine-tuning
  • inference optimization (TTFT, throughput)
  • quantization, caching strategies, and efficient batching

…

Posted: October 1st, 2026