What you will own

  • Reliability automation across deployment and regional infrastructure.
  • Observability standards, incident tooling, and capacity signals.
  • Resilience testing and failure-mode exercises.
  • Service-level objectives and production readiness reviews.

What we look for

  • Hands-on experience with Linux, networking, and automation.
  • Comfort debugging distributed systems from telemetry.
  • A pragmatic approach to operational risk.
  • Strong incident communication and post-incident learning habits.

First 90 days

Map the highest-risk failure paths, improve one operational control, and run a resilience exercise with the team.

Apply to reliability engineering

Include an example of an incident you helped resolve and what you changed afterward.

Apply for this role