Blog - AI Reliability, AI SRE & System Reliability Insights | Dalton AI

Writing on reliability, AI, and production.

Engineering insights on AI reliability, AI SRE, system reliability, production resilience, and the future of DevOps.

01 FEATURED MANIFESTO · MAY 18, 2026

The Double Exposure: When Your AI Agents Run on AI-Generated Code
In March 2026, Amazon lost 6.3 million orders in a single afternoon. The cause wasn't an AI agent. It wasn't AI-generated code either. It was both - running on top of each other. This is the failure surface nobody's measuring yet.

Itamar Knafo· May 18, 2026

02 MORE POSTS · 6 ENTRIES

Innovation Mythos: What Anthropic's Own Reliability Team Says - and What It Means for Production AI
Anthropic's model card includes a candid reliability assessment of Claude Mythos. We break down what their own engineers found - where it's a step change, where it still falls short, and what this means for teams running AI in production.

Itamar Knafo· Apr 8, 2026

Engineering The Three Meanings of "Nothing" in AI Agent Tool-Calling
Mid-investigation, a tool call comes back empty. One agent retries endlessly. Another reports a failure. A third keeps searching across every data source it can find. Same empty response, three completely different problems.

Yanir Rot· Apr 5, 2026

Industry What is AI SRE? The History, the Market, and What Comes Next
AI SRE did not appear out of nowhere. Here's where it came from, how the market is forming, and why the next phase will be about prevention, not prettier incident response.

Itamar Knafo· Mar 29, 2026

Innovation AI Is Writing 46% of Your Code. Nobody's Watching What It Breaks.
We use AI coding tools every day. They're the future. But 46% AI-generated code, 98% more PRs, and 30% more change failures is what that future actually looks like - and nobody's talking about it honestly.

Itamar Knafo· Feb 7, 2026

AI Your AI Strategy Is a Data Strategy Wearing a Costume
Most companies think they have an AI problem. They don't. They have a data problem - messy pipelines, no context, garbage semantics - and they're throwing models at it hoping nobody notices.

Michael Zigelboim· Jan 14, 2026

Manifest Postmortem: Reactive Operations
We automated detection twenty years ago (to some extent...) - Investigation is still manual. Here's why that's a problem.

Itamar Knafo· Dec 18, 2025