SRE

SRE

286 bookmarks
Custom sorting
Felix Geisendörfer on Twitter
Felix Geisendörfer on Twitter
🎉 Announcing fgtrace, a new profiler/tracer for #golang.It captures wallclock timeline views for each goroutine and it's really simple to use:defer fgtrace.Config{}.Start().Stop()Check it out & let me know what you think https://t.co/Ttdm5hl0Vi pic.twitter.com/4iP9SNVypD— Felix Geisendörfer (@felixge) September 19, 2022
·twitter.com·
Felix Geisendörfer on Twitter
What is eBPF? | An Introduction and Practical Tips
What is eBPF? | An Introduction and Practical Tips
Addr:https://ebpf.xyz/post/an_introduction_and_practical_tips March 23, 2022 This article introduces developers to eBPF and explains how it can be used to add security, networking, and other capabilities in the Linux kernel space. In Linux architecture, memory is separated into kernel space and user space. The kernel space is used to run the core kernel code and the device drivers. Processes running in kernel space have unrestricted access to all hardware, including CPU, memory, and disks.
·ebpf.xyz·
What is eBPF? | An Introduction and Practical Tips
Resilience Engineering and Strange Loops
Resilience Engineering and Strange Loops
My notes and takeaways from a long read on anomalies and system complexity called the STELLA Report from the SNAFUcatchers Workshop on Coping With Complexity, 2017. Via Matt. This paper is one of t…
·sensible.blog·
Resilience Engineering and Strange Loops
Who Destroyed Three Mile Island? - Nickolas Means | #LeadDevLondon 2018
Who Destroyed Three Mile Island? - Nickolas Means | #LeadDevLondon 2018
Check out the latest from The Lead Developer at theleaddeveloper.com. On March 28, 1979, at exactly 4 o’clock in the morning, control rods slammed into the reactor core of Three Mile Island Unit #2, halting the nuclear reaction because of a fault in the reactor cooling system. At 4:02, the automated emergency cooling system activated as the reactor core temperature continued to rise. At 4:04, one of the plant operators made the befuddling decision to switch off the emergency cooling system, dooming the reactor to partial meltdown. Why? When something bad happens, it’s easy to just blame someone and move on. Taking the time to find the systemic causes, though, will not only help keep the problem from repeating, it will enable you to build the psychological safety necessary for your team to truly collaborate. Let’s let the story of Three Mile Island teach us how to make our teams stronger through systems thinking and just culture.
·youtu.be·
Who Destroyed Three Mile Island? - Nickolas Means | #LeadDevLondon 2018
SLI, SLO, SLA explained in a way your kids will understand… maybe
SLI, SLO, SLA explained in a way your kids will understand… maybe
Imagine you are in a remote meeting using terms like SLI, SLO, or SLA, and your kid asks you what it means? How would you explain it to them? Or maybe you need to explain it to your boss or a colleague. In this article, I will try to put SLI, SLO, and SLA in a way even your kids would understand… maybe.
·thrownewexception.com·
SLI, SLO, SLA explained in a way your kids will understand… maybe
[PUBLIC] The Art of SLOs – Slides
[PUBLIC] The Art of SLOs – Slides
Self link: https://cre.page.link/art-of-slos-slides Participant Handbook: https://cre.page.link/art-of-slos-handbook Facilitator Handbook: https://cre.page.link/art-of-slos-howto SLO Worksheet: https://cre.page.link/art-of-slos-worksheet Errors in the content? https://cre.page.link/art-of-slos-bu...
·docs.google.com·
[PUBLIC] The Art of SLOs – Slides
Improving Incident Management through Role Assignments and Game Days
Improving Incident Management through Role Assignments and Game Days
John Arundel, principal consultant at Bitfield Consulting, shared his thoughts on how to ensure incidents are handled smoothly and quickly. He suggests assigning specific roles to each team member responding to the incident. Red team versus blue team exercises can also be leveraged to ensure the team is prepared to respond accurately and quickly.
·infoq.com·
Improving Incident Management through Role Assignments and Game Days
SRE book list
SRE book list
Increase your knowledge of site reliability engineering with these books
·docs.microsoft.com·
SRE book list
Steampipe | select * from cloud;
Steampipe | select * from cloud;
Steampipe is an open source tool to instantly query your cloud services (e.g. AWS, Azure, GCP and more) with SQL. No DB required.
·steampipe.io·
Steampipe | select * from cloud;
The Best Way to Organize & Manage Microservices
The Best Way to Organize & Manage Microservices
OpsLevel’s mission is to make it simpler for companies to ship and operate high-quality software through service ownership and developer portals.
·opslevel.com·
The Best Way to Organize & Manage Microservices
Simplicity and Reliability
Simplicity and Reliability
The code is a liability, not an asset. The debt begins for a team the moment the first line of code is written. The code is written when…
·medium.com·
Simplicity and Reliability
Introducing Domain-Oriented Microservice Architecture | Uber Blog
Introducing Domain-Oriented Microservice Architecture | Uber Blog
Recently there has been substantial discussion around the downsides of service oriented architectures and microservice architectures in particular. While only a few years ago, many people readily adopted microservice architectures due to the numerous benefits they provide such as flexibility in the form of independent deployments, clear ownership, improvements in system stability, and better separation of concerns, in recent years people have begun to decry microservices for their tendency to greatly increase complexity, sometimes making even trivial features difficult to build.
·uber.com·
Introducing Domain-Oriented Microservice Architecture | Uber Blog
Slowing Down to Speed Up - Circuit Breakers for Slack's CI/CD - Slack Engineering
Slowing Down to Speed Up - Circuit Breakers for Slack's CI/CD - Slack Engineering
What happens when your distributed service has challenges with stampeding herds of internal requests? How do you prevent cascading failures between internal services? How might you re-architect your workflows when naive horizontal or vertical scaling reaches their respective limits? These were the challenges facing Slack engineers during their day-to-day development workflows in 2020. Multiple internal …
·slack.engineering·
Slowing Down to Speed Up - Circuit Breakers for Slack's CI/CD - Slack Engineering
The point of a dashboard isn't to use a dashboard
The point of a dashboard isn't to use a dashboard
Every so often, an employer asks me to help make a dashboard. Usually, this causes technologists to roll their eyes. They have a vision of a CEO grandly staring at a giant projection screen, watchi…
·shkspr.mobi·
The point of a dashboard isn't to use a dashboard