Overview
In this role you will own the behavior of LLM-driven agents in production, shaping prompts and evaluating outputs to ensure accuracy, on-brand tone, and cost efficiency. You will work within an established framework to design prompts, guardrails, and evaluation methods, driving improvements through feedback loops and A/B testing. You’ll partner with cross-functional teams to align AI outputs with business goals while maintaining safety standards. This is an opportunity to influence customer-facing AI at scale inside a premier analytics technology company.
Responsibilities
- Develop & tune prompts and architectures for production AI agents (system prompts, tool instructions, guardrails, output schemas)
- Build and run evaluation suites and regression tests; define quality metrics and acceptance thresholds with business owners
- Design and operate A/B testing for prompts and model variants; implement feedback loops from human review for agent optimization
- Support LLM optimization efforts focusing on model selection, context-window management, structured outputs, and token efficiency
- Maintain tone-of-voice and multi-language consistency for customer communications using professional British English as baseline
- Stress-test guardrails against adversarial inputs, edge cases, and hallucination risks within safety standards
- Document prompt patterns, test coverage, and evaluation results to enable structured business feedback
Key requirements
- 1+ year hands-on experience with LLM applications (prompting, RAG, structured outputs)
- Experience with systematic testing, data validation, or QA of AI/ML systems
- Strong prompt-writing craft; disciplined experimentation and basic Python/data-handling skills for evaluation pipelines
- Exceptional written English with professional B2B tone and nuance
- Clear documentation, attention to detail, and structured approach to user feedback
- Attention to detail
- Structured problem solving
- Collaborative cross-functional communication
- Prompt-writing craft
- Experimentation design and analysis
- Python or data-handling basics for evaluation pipelines
…
