The LLM / Agentic Evaluation Rig Engineer will be responsible for building and maintaining the evaluation infrastructure that ensures AI outputs meet quality standards before being shipped. This includes curating datasets, developing scoring systems for various quality metrics, and integrating evaluations into continuous integration processes. The role emphasizes rigorous measurement of grounding, faithfulness, and hallucination in AI outputs, alongside collaboration with other teams to enhance model performance.
See how your résumé matches this role — and tailor it from what actually gets interviews.