Lorin Hochstein 
mastodon 4.7.3Student of complex systems failures, resilience engineering, cognitive systems engineering. Will talk your ear off about learning from incidents in software.
The three skills with a lot less overlap than you’d expect:
1. Ability to code.
2. Ability to perform well in a coding interview.
3. Ability to validate code.
Looking for a little less machine learning and a little more human learning
New blog post about Erik Hollnagel's Safety-II model: https://surfingcomplexity.blog/2026/04/26/the-normal-work-of-creating-reliability/
Automation can only solve your reliability problems if that automation is perfect. But if you were capable of writing perfect software, we wouldnt be talking about your reliability problems.
Automation is only part of the reliability story, you’ve got to be ready for when it breaks, too. Because I promise you, your reliability automation is going to break. Will you be ready for when it does?
The consequences of being able to generate code with AI much more quickly than we can validate that code is Nature’s way of teaching us about Erik Hollnagel’s ETTO Principle. https://www.erikhollnagel.com/ideas/etto-principle/index.html
According to the data, we shouldn’t trust the data
It's counterintuitive, but you can learn about the nature of successful work from incidents, even though the incident was a failure case. I wrote a blog post about that here:
I’ve heard of “Nimrod” primarily because of
New blog post about flipping the bozo bit: https://surfingcomplexity.blog/2026/05/09/flipping-the-bozo-bit-on-flips-the-learning-off/
What a time to be alive
Jim Calabro did an excellent write-up of a recent Bluesky outage, I learned a lot from it. I wrote my own post here: https://surfingcomplexity.blog/2026/04/12/thoughts-on-the-bluesky-public-incident-write-up/
There’s no way this reading is accurate
Cloudflare announces a huge layoff and then, the very same day, an availability zone in AWS’s us-east-1 region loses power due to a “thermal event”. Coincidence???
Well, yes
You get what the system is optimized for.
GitHub's had some availability issues lately. Their CTO wrote a blog post with some details about recent incidents. I wrote a quick post with my reaction here: https://surfingcomplexity.blog/2026/03/12/quick-thoughts-on-github-ctos-post-on-availability/
Prediction: Within the next two years, the New York Times will write a story about people who are unable to prove they are not AIs (à la CAPTCHAs).
Is there a burn ward nearby?
Seamlessness is, indeed, a worthwhile goal; systems fail at the seams.
New blog post on the surprising increases in load faced by AI companies and the operational consequences: https://surfingcomplexity.blog/2026/03/07/grow-fast-and-overload-things/
I’m very concerned about the aliens already walking among us who are going to insider trade on this market: https://kalshi.com/markets/kxaliens/aliens/kxaliens-27
Tired: old code that calls new code (extensibility!)
Wired: old code that breaks new code (latent failures!)
Managing the risk of saturation, AI edition: https://www.infoworld.com/article/4151196/anthropic-throttles-claude-subscriptions-to-meet-capacity.html
Which of these technologies will become practical in your lifetime?
If you live with a partner (or partners) and you all work from home, how do you communicate with each other during the work day?
All metrics are subjective, it’s just that you don’t see the subjectivity because it’s baked in at metric definition time. What you decide to measure and how you decide to measure it are both subjective decisions.
If you celebrate Passover, which of these do you commonly eat at a Passover seder?



