HHiring Reality
← 84.51°

Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)

84.51°
Location
Cincinnati, OH
Posted
1 month ago
Department
Engineering
What they actually want (must-haves)
  • Minimum 10+ years of experience in software, platform, cloud, SRE or production engineering, including 3 or more years leading engineering teams.
  • Track record of building engineering capabilities, standards, automation or platforms that measurably improved system quality and reliability at scale.
  • Strong technical depth in distributed systems, Kubernetes, cloud infrastructure on GCP and/or Azure, networking, identity, automation, CI/CD and modern observability engineering.
  • Experience establishing SLOs, error budgets, on-call practices, production-readiness reviews, incident management and post-incident improvement programs.
  • Direct experience with ML, AI or LLM systems in production, or demonstrated ability to master failure modes unique to models and agents, including drift, non-determinism, quality regression, prompt and context issues, and hallucination.
  • Ability to influence architecture and roadmaps, lead through ambiguity, partner across a matrixed organization and drive outcomes without direct authority.
Nice to have
  • Hands-on experience with LLMOps, model observability and evaluation tooling such as LangSmith, MLflow, Arize, Fiddler or custom evaluation pipelines.
  • Familiarity with agentic systems, including orchestration, tool use, MCP-based integrations, registries and their unique reliability and quality challenges.
  • Background bridging software or SRE practices with ML, data science and Responsible AI engineering cultures, including experience establishing a new engineering discipline.
  • Experience with multi-region, high-availability services, hybrid operating models and regulated or high-risk enterprise workloads.
What the job really is

The Senior Manager, AI Reliability Engineering at Kroger Technology & Digital will lead the establishment of a new engineering discipline focused on making enterprise AI operationally trustworthy at scale. This role involves defining standards for production-grade AI, ensuring reliability and quality in AI systems, and collaborating with various teams to embed resilience and observability into AI architectures. The manager will also oversee incident management, team development, and drive efficiency in AI operations.

Things to weigh
  • Role involves building a new discipline from the ground up, which may come with challenges and uncertainties.
  • No specific salary or benefits information provided in the posting.
  • The position requires a strong blend of technical and leadership skills, which may not suit all candidates.
  • The role is heavily focused on operational reliability and may involve high-pressure incident management responsibilities.
Job score3/5
Benefits1/5
Freshness4/5
Career value5/5
Role clarity5/5
Pay transparency0/5

Applying to 84.51° ?

See how your résumé matches this role — and tailor it from what actually gets interviews.