As a Site Reliability Engineer at SpaceXAI, you will focus on ensuring the reliability of campus operations by designing monitoring systems, leading incident responses, and managing cross-functional reliability projects. Your role will involve technical leadership during incidents, maintaining playbooks, and collaborating with teams across various disciplines such as compute, network, and storage. You will also participate in on-call rotations and drive improvements based on incident feedback.
See how your résumé matches this role — and tailor it from what actually gets interviews.