AI Evaluation Engineer

AI Evaluation Engineer
£700-£750/day Inside IR35 | Duration: 4 months | Clearance: BPSS + SC eligible | Location: London, Bristol or Manchester | Hybrid: 2 days onsite per week.
We're looking for an experienced AI Evaluation Engineer to join a specialist Public Sector Technology team. This hands-on role focuses on building evaluation frameworks, tooling and harnesses for AI systems, particularly LLM and agentic AI solutions. The role requires ownership, experimentation and rapid delivery.
AI/LLM Evaluation, RAG Evaluation, Ragas, Agentic AI Evaluation, Evaluation Harnesses, LLM Testing and Benchmarking, Prompt Engineering, Python Development, Model and Agent Integration.
AI Evaluation Engineer, Applied AI Engineer, AI Engineer, LLM Engineer, Harness Engineer, Prompt Engineer or Software Engineer specialising in AI.
Small Public Sector Technology team of approximately 2 to 6 people within a wider organisation.
Experience building evaluation tooling, evaluating LLMs, using Ragas, identifying AI failure modes, combining engineering with strategic thinking, and thriving in fast-moving environments.
Hands-on AI Evaluation Engineer role supporting government AI initiatives through evaluation, testing, benchmarking and framework development.