In this role as a Machine Learning Research Scientist focused on evaluations, you will analyze model behavior to identify and diagnose failure modes in large language models (LLMs) and agents. You will design benchmarks and evaluation methods for both text and multimodal modalities, apply post-training techniques, and publish your research findings in top-tier AI conferences. Collaboration with researchers and engineers will be key to defining best practices in evaluation-driven AI development.
See how your résumé matches this role — and tailor it from what actually gets interviews.