Overview
In this role you will design and evolve a Python-based testing framework for generative AI agents and LLM-based systems. You’ll shape evaluation pipelines, observability, and data-driven testing, collaborating across AI, data engineering, and QA teams to improve reliability in production. You’ll implement synthetic/adversarial data generation and define quality KPIs to drive trustworthy AI behavior. This is a hands-on, independent role with a strong impact on how we validate AI systems at scale.
Responsibilities
- Design and develop Python-based testing frameworks for AI agents and LLM systems
- Contribute to architecture for evaluation pipelines, observability, and data-driven testing
- Build and maintain test data generation tools (synthetic/adversarial data)
- Define and implement evaluation strategies and QA KPIs (accuracy, robustness, bias, hallucinations)
- Integrate LLM evaluation tools, CI/CD, and external AI platforms
- Support continuous testing, benchmarking, and production monitoring of AI systems
- Collaborate with AI engineers, data engineers, and QA teams to improve reliability and scalability
Key requirements
- 2–5+ years of experience in Software Engineering, Test Automation, AI Engineering, or Data Engineering
- Strong Python proficiency and data/AI libraries (pandas, numpy, PyTorch or similar)
- Solid understanding of LLMs, NLP concepts, or generative AI systems
- Experience with test automation frameworks (PyTest, Playwright) and CI/CD
- Familiarity with data pipelines, dataset design, or feature engineering
- Experience with evaluation metrics for NLP/AI systems (BLEU, ROUGE, embedding-based metrics)
- Ability to design scalable systems and work autonomously on complex problems
- Strong analytical mindset and interest in AI system quality, reliability, and governance
- autonomous work style
- collaboration with cross-functional teams
- analytical mindset
- Python
- pandas
- numpy
…
