Anthropic Details Claude Code Quality Problems in Postmortem

Anthropic says recent complaints about Claude Code quality were caused by three separate issues in the product’s harness, not by the underlying models. In one example, a bug made Claude repeatedly clear older thinking in idle sessions, which made it seem forgetful and repetitive, according to the company’s postmortem.

Anthropic Details Claude Code Quality Problems in Postmortem

Anthropic has said that recent complaints about Claude Code quality were not caused by the underlying Claude models, but by three separate problems in the product’s harness, the software layer that wraps the model and manages how it behaves in a session. The company’s postmortem, published after users reported that Claude Code had become more forgetful, repetitive, or generally worse at coding tasks, points to orchestration bugs rather than a sudden drop in model capability.

Claude Code is Anthropic’s coding-focused agent, designed to work in a terminal or developer workflow and carry out multi-step tasks with limited supervision. Like other agentic tools, it does not just answer one prompt at a time. It has to keep track of context, session state, and the internal reasoning process that helps it decide what to do next.

⚡ New to this?

Claude Code is Anthropic’s AI coding assistant, a tool that helps write and edit code inside a developer workflow. A harness is the software around the model that manages things like memory, session state, and when the AI should keep or drop context.

This matters because people often blame the AI model itself when a tool gets worse, but the problem can be in the wrapper around it. For non-experts, this is a good example of why AI systems are more than the model alone, and why bugs in the surrounding software can make a product look forgetful or unreliable.

🦞 OpenClaw angle

For AI automation builders, this is a reminder to inspect the harness before blaming the model. Session state, memory resets, and idle-time behavior can quietly break agent workflows and create reliability problems that look like model drift.

For self-hosters and IT teams, it’s a useful example of why observability around orchestration layers matters. If your system feels “forgetful,” the bug may be in the wrapper, not the model.

That distinction matters because users often experience the product as a single thing, even though there are several layers underneath it. If the model itself is fine but the surrounding system mishandles memory, resets context too aggressively, or interrupts the agent at the wrong point, the result can look like the AI has forgotten what it was doing.

One of the issues Anthropic described was a bug that repeatedly cleared older thinking in idle sessions. In practice, that meant Claude could lose parts of its earlier reasoning if a session sat idle long enough, which made later responses seem inconsistent or repetitive. According to the company, that behavior created the impression that the model had become forgetful even when the root cause was session handling.

This kind of failure is easy to miss from the outside because it presents as a quality problem, not an obvious crash. A user sees the agent loop, restate itself, or lose track of a plan, and the natural assumption is that the model has degraded. In reality, the failure can sit in the wrapper around the model, including the code that stores conversation state, resumes work, or decides what context to preserve.

Anthropic said the three issues were separate and all lived in the product harness. The company’s explanation is a reminder that agent systems are not just models with a chat box attached. They are a stack of model inference, prompt handling, state management, and workflow logic, and any one of those layers can affect the user’s experience.

For Claude Code, that matters especially because coding workflows tend to be long-running and stateful. Developers may leave a session open, return later, and expect the agent to continue with the same understanding of the task. If the harness mishandles idle time or context preservation, the system can appear to regress even when the core model has not changed.

The postmortem also fits a wider pattern in AI tools, where quality complaints often turn out to involve tooling around the model rather than the model weights themselves. That includes session memory, tool routing, prompt injection, and other orchestration layers that are easy to underestimate when everything is working and very visible when it is not.

Anthropic’s update does not describe a change in Claude’s underlying capabilities. It describes a set of product-layer failures that made the coding assistant feel less reliable than intended, including the idle-session bug that cleared older thinking and left users with an agent that seemed to forget its own work.

Source: Simon Willison ↗

More from OpenClaw News