Skip to content

Skills you need to be a site reliability engineer

9 skills a hiring manager would actually test for, each with the level this role expects and what it is used for. Not a syllabus — the shape of the job.

Build my path to this role

Upskili checks what you can already do, then sequences only what is missing. No account needed.

What the role requires

Ordered by how much the job depends on it. The bar is the proficiency expected of a competent site reliability engineer — not mastery, and not a passing acquaintance.

  • Kubernetes

    Essential

    Orchestrates and manages all containerized production workloads.

    Strong
  • Terraform

    Essential

    Provisions and manages all cloud infrastructure as code.

    Strong
  • Observability tools (e.g., Prometheus, Grafana)

    Essential

    Builds monitoring, alerting, and dashboards for system health.

    Strong
  • Incident management

    Essential

    Leads incident response, postmortems, and blameless culture.

    Strong
  • Linux systems administration

    Important

    Debugs performance issues and configures production servers.

    Strong
  • CI/CD pipelines (e.g., GitHub Actions, Jenkins)

    Important

    Automates build, test, and deployment workflows.

    Strong
  • Python or Go scripting

    Important

    Writes automation scripts and simple internal tooling.

    Working
  • Cloud platform (AWS, GCP, or Azure)

    Important

    Manages cloud resources, networking, and IAM policies.

    Strong
  • Git

    Useful

    Collaborates on infrastructure code and configuration changes.

    Strong

An order worth learning it in

A list of ten skills is the same unhelpful answer a catalogue gives, just sorted. This is where to actually start.

1

Start here

Essential to the role, and reachable from a standing start. Everything below rests on these.

  • Terraform
  • Incident management
  • Linux systems administration
  • CI/CD pipelines (e.g., GitHub Actions, Jenkins)
  • Python or Go scripting
  • Cloud platform (AWS, GCP, or Azure)
2

Then this

The rest of what the role is assessed on. Harder, and it builds on the foundation above.

  • Kubernetes
  • Observability tools (e.g., Prometheus, Grafana)
3

What sets you apart

Not what gets you hired, but what separates doing the job from being trusted with it.

  • Git

You almost certainly have some of this already.

That is the point of starting from the role rather than a course. Upskili checks what you can do, then builds a path across only the gap.

See my path to site reliability engineer