HHiring Reality
← Phizenix

LLM / Agentic Evaluation Rig Engineer

Phizenix
Location
Hyderabad, India (Hybrid)
Posted
1 month ago
Department
External - Client Requirement
What they actually want (must-haves)
  • 4+ years in software / ML engineering, with hands-on work building LLM evaluation or quality tooling.
  • Real understanding of grounding, faithfulness, and hallucination — and how to measure them rigorously.
  • Strong Python and solid engineering practices (reproducibility, CI/CD).
  • Comfort designing evaluation for non-deterministic systems without producing flaky or meaningless metrics.
  • Familiarity with LLM eval frameworks and LLM-as-judge patterns.
Nice to have
  • Experience evaluating agentic / multi-step LLM systems.
  • Familiarity with RAG, structured output, and managed LLMs in-VPC.
  • FinTech / financial-services domain or other high-stakes, correctness-critical AI.
  • Background in statistics or measurement / metrics design.
What the job really is

The LLM / Agentic Evaluation Rig Engineer will be responsible for building and maintaining the evaluation infrastructure that ensures AI outputs meet quality standards before being shipped. This includes curating datasets, developing scoring systems for various quality metrics, and integrating evaluations into continuous integration processes. The role emphasizes rigorous measurement of grounding, faithfulness, and hallucination in AI outputs, alongside collaboration with other teams to enhance model performance.

Things to weigh
  • The role involves significant responsibility in defining quality standards for AI outputs, which may lead to high pressure.
  • The technical stack includes advanced AI evaluation frameworks, which may require continuous learning and adaptation.
  • No salary or specific benefits mentioned in the posting.
Job score2.8/5
Benefits1/5
Freshness4/5
Career value4/5
Role clarity5/5
Pay transparency0/5

Applying to Phizenix?

See how your résumé matches this role — and tailor it from what actually gets interviews.