Book
Site Reliability Engineering
Betsy Beyer, Chris Jones, Jennifer Petoff, Niall Richard Murphy
Summary
Google が提唱した Site Reliability Engineering の原典。SLI/SLO/SLA、 エラーバジェット、オンコール運用、リリースエンジニアリング、大規模分散システムの 運用思想を、運用チームの文化と合わせて体系化した、SRE 実践の必読書。
Target Readers
- 信頼性と俊敏性の両立を目指す運用エンジニア
- SRE チームを立ち上げたい組織リーダー
Tags
Colophon
- Publisher
- オライリー・ジャパン
- ISBN
- 978-4873117911
- Published
- Aug 2017
- List price
- ¥5,280incl. taxMay differ from the actual selling price on Amazon
Get this book
* The link above is an advertisement via Amazon Associates.Related Books
Prerequisites
- Recommended
The DevOps ハンドブック
Gene Kim, Jez Humble, Patrick Debois, John Willis
Reason: Where 'The DevOps Handbook' preaches collaboration between development and operations, Google's 'Site Reliability Engineering' shows the concrete form of solving the operations side 'with software engineering'. With error budgets and SLOs to engineer reliability, it is a natural evolution of DevOps.
- Related
Practical Monitoring
Mike Julian
Reason: After designing 'what to measure' with 'Practical Monitoring', advance to the philosophy of tying those metrics to organizational decisions. Google's 'Site Reliability Engineering' elevates monitoring data—via SLIs/SLOs/error budgets—into criteria for 'when to halt feature work and invest in reliability'.
- Related
Release It!, 2nd Edition
Design and Deploy Production-Ready Software
Michael T. Nygard
Reason: The stability patterns in 'Release It!' make individual services harder to break, but how to set targets for whole-system reliability and operate it organizationally is a separate question. Google's 'Site Reliability Engineering' integrates individual fault-tolerant design into organizational reliability management through error budgets and incident-response structures.
- Related
チームトポロジー
Matthew Skelton, Manuel Pais
Reason: After designing an organization for fast flow of value with 'Team Topologies', you arrive at the question of which team owns reliability and how. Google's 'Site Reliability Engineering' presents concrete practices—error budgets, sharing of operational responsibility—for embedding reliability as organizational culture.
Next Books
- Prerequisite
The Site Reliability Workbook
Betsy Beyer, Niall Richard Murphy, David K. Rensin, Kent Kawahara, Stephen Thorne
Reason: Where 'Site Reliability Engineering' articulates principles distilled from Google's practice, its sequel 'The Site Reliability Workbook' shows 'how to implement it at your own company' with concrete procedures and case studies. Theory first, then the implementation volume—a required progression.
- Recommended
Observability Engineering
Charity Majors, Liz Fong-Jones, George Miranda
Reason: SRE presupposes 'knowing the exact state of the system' to meet SLOs, but the SRE book itself stays at the philosophy of monitoring. 'Observability Engineering' supplements the techniques—distributed tracing, high-cardinality events—to explore unknown failures, satisfying at the implementation level the observation capability SRE demands.
- Recommended
Building Secure and Reliable Systems
Best Practices for Designing, Implementing, and Maintaining Systems
Heather Adkins, Betsy Beyer, Paul Blankinship, Piotr Lewandowski, Ana Oprea, Adam Stubblefield
Reason: Once SRE lets you handle reliability as engineering, you arrive at the question 'aren't security and reliability fundamentally the same design problem?'. Google's 'Building Secure and Reliable Systems' extends SRE to show principles for building security into system design rather than bolting it on, integrating reliability and security.
- Related
Seeking SRE
Conversations About Running Production Systems at Scale
David N. Blank-Edelman
Reason: After learning Google-originated SRE theory, you want to know how others adapt and practice it. Edited by Blank-Edelman, 'Seeking SRE' is a collection of contributions from many practitioners, offering diverse applications of SRE in non-Google contexts and broadening the scope of the principles.
Sources
- Related
LLMOps
Abi Aryan
Reason: Where Google's 'Site Reliability Engineering' frames engineering reliability for general services via SLIs/SLOs and error budgets, 'LLMOps' applies that operating philosophy to the new target of LLM applications, organizing the production concerns—including governance and cost management—specific to running them.
Sources