Recruiting developers into Site Reliability Engineering (SRE)

In this article, you will learn the following: Introduction Hiring in the Site Reliability Engineering (SRE) space is notoriously difficult. So it makes sense to figure out how to expand the hiring pool beyond existing SREs. One way to increase the hiring pool is to recruit developers (also known as SWEs) and gradually advance them … Read More

Analysis of SRE and platform setup at 10+ tech companies

In this article, you will see a breakdown of the platform setup and SRE practices within 12 non-FAANG technology companies. This is based on the case studies by Andrios Robert. “There is a lot of content available on how Google did [Site Reliability Engineering]; let’s uncover what happens with the rest of the world.” — … Read More

Where in team topologies does Site Reliability Engineering fit in?

We will explore the workings of the Team Topologies model and how Site Reliability Engineering (SRE) teams can fit into it. In more detail, I will share with you the following: an overview of the team topologies model the 4 team modalities it proposes, and finally… where SRE teams fit in team topologies Let’s get … Read More

Building the case for starting a software reliability team

This article aims to help engineering leaders consider issues before starting a software reliability team. Since I am an advocate for Site Reliability Engineering (SRE), we will now refer to such a team as the “SRE team”. Besides creating a new team, leaders face many responsibilities that are often invisible to individual contributors and their … Read More

How cloud infrastructure teams evolve – from start to maturity

I recently read a post by Will Lason, who started SRE at Uber. The post is called the Trunks and branches model for scaling infrastructure organizations. Several passages in the post covered how infrastructure teams can evolve from the startup phase. I felt it would be easier to comprehend the dense-and-rich advice with a visual … Read More

Cloud infrastructure success is a fine balance of budget and service quality

The visual summary below is based on a post by Will Larson, who started the SRE function at Uber. His post elaborates on a “trunks and branches” model for developing infrastructure-facing teams. It also covered an interesting perspective on the balancing act of budget and service quality. I will explain the visual summary underneath it. … Read More

Site Reliability Engineering Culture Patterns

Who should read this: Developers new to Site Reliability Engineering (SRE) who want to understand the culture Current SREs who are seeking to guide others like management on key aspects of SRE culture Technical leaders who want to create an ideal culture for effective software reliability practices Introduction Despite its now antiquated sounding name, Site … Read More

25+ Site Reliability Engineering OKRs

Incident Response OKRs Reduce MTTR for on-call engineers by 5% Develop buffers to ensure incidents remain at < 75% of the error budget Mitigate false positive system alerts to reduce on-call staff costs Speed up the resolution of critical incidents by 5% Increase the coverage of 4-point SLIs from 90% of services to 100% Reduce … Read More