Shaun Abram
Technology and Leadership Blog
Summary: The Observability Crisis
The Observability Crisis is an article from Jaya Gupta & Ashu Garg from Foundation Capital, a Silicon Valley based venture capital (VC) firm investing in tech startups.
TLDR:
Companies in the first wave of the observability space (such as Splunk, AppDynamics, Datadog and New Relic) focused on solving data storage and analysis problems. However, with ever-increasing data volumes, costs increase and extracting actionable insights becomes more difficult.
Companies in the second wave (such as Honeycomb, Grafana and Cribl) helped organizations cope with complexity and cost by better supporting an increase in the volume, velocity, and the cardinality of data. However, the engineering overhead (such as SREs) to manage these tools, control costs, and triage alerts remains high.
The third wave could involve a shift from volume-based to value-based pricing, where you pay for observability tools based on their ability to accurately identify root causes and even prevent them from recurring. This is likely to involve using LLMs along with ML techniques that:
- Offer predictive maintenance by recognizing patterns across heterogeneous data sources
- Automatically generate incident playbooks
- Synthesize data from post-mortem reports
So, while I do like this article, my main gripe is that the article leaves it unclear whether any such “3rd wave” tools exist yet.
Read on for a summary of that article. The original is ~1600 words. This is ~600.
Tags: apm, devops, incidents, observability, RCA, reliability, rootcauseanalysis, sitereliabilityengineering, sre, summary
Blog post summary: Blameless PostMortems post by John Allspaw
The following is a slightly summarized version of this blog post from John Allspaw that I really like: Blameless PostMortems and a Just Culture
Tags: blameless, postmortems, RCA, rootcauseanalysis, summary
Book chapter summary: Postmortem Culture, from the SRE Book
I’m really enjoying reading the excellent “SRE Book“. Chapter 15 “Postmortem Culture: Learning from Failure” in particular, really struck a chord with me. The following is a slightly summarized version of it.
TLDR: Failures are inevitable, especially in distributed systems. To learn from them, document in Postmortems, avoiding blame, and share the newly gained learnings across your org.
Tags: blameless, postmortems, RCA, rootcauseanalysis, sitereliabilityengineering, sre, summary, thesrebook
Subscribe to RSS Feed