Knihobot

Niall Richard Murphy

    Niall Murphy je autorem či spoluautorem řady technických publikací, včetně "IPv6 Network Administration" pro O’Reilly, a několika RFC. Jeho rozsáhlé zkušenosti v internetovém průmyslu jsou doplněny studiem informatiky, matematiky a poezie, což mu dává jedinečnou perspektivu. V současné době se podílí na psaní historie internetu v Irsku a předsedá irskému peeringovému uzlu INEX. Jeho práce se vyznačuje hlubokým porozuměním technickým základům i kritickým pohledem na vývoj digitálního světa.

    The Site Reliability Workbook
    Site reliability engineering : how Google runs production systems
    • The Site Reliability Workbook

      Practical Ways to Implement SRE

      • 512 stránek
      • 18 hodin čtení

      In 2016, Google's <i>Site Reliability Engineering</i> book ignited an industry discussion on what it means to run production services today--and why reliability considerations are fundamental to service design. Now, Google engineers who worked on that bestseller introduce <i>The Site Reliability Workbook</i>, a hands-on companion that uses concrete examples to show you how to put SRE principles and practices to work in your environment. This new workbook not only combines practical examples from Google's experiences, but also provides case studies from Google's Cloud Platform customers who underwent this journey. Evernote, The Home Depot, <i>The New York Times</i>, and other companies outline hard-won experiences of what worked for them and what didn't. Dive into this workbook and learn how to flesh out your own SRE practice, no matter what size your company is. You'll learn: How to run reliable services in environments you don't completely control--like cloud Practical applications of how to create, monitor, and run your services via Service Level Objectives How to convert existing ops teams to SRE--including how to dig out of operational overload Methods for starting SRE from either greenfield or brownfield

      The Site Reliability Workbook2018
    • "The overwhelming majority of a software system's lifespan is spent in use, not in design or implementation. So, why does conventional wisdom insist that software engineers focus primarily on the design and development of large-scale computing systems? In this collection of essays and articles, key members of Google's Site Reliability Team explain how and why their commitment to the entire lifecycle has enabled the company to successfully build, deploy, monitor, and maintain some of the largest software systems in the world. You'll learn the principles and practices that enable Google engineers to make systems more scalable, reliable, and efficient - lessons directly applicable to your organization. This book is divided into four sections: Introduction - Learn what site reliability engineering is and why it differs from conventional IT industry practices; Principles - Examine the patterns, behaviors, and areas of concern that influence the work of a site reliability engineer (SRE); Practices - Understand the theory and practice of an SRE's day-to-day work: building and operating large distributed computing systems; Management - Explore Google's best practices for training, communication, and meetings that your organization can use."--Publisher's description

      Site reliability engineering : how Google runs production systems2016
      4,2