Build notes
8/02/2026

The model you feel

Benchmarks say Fable 5 is the better model. Agentic operators who spend all day inside these systems say the gap is smaller than anyone expected, and that the models in between were a step backward. Both are right, and the difference is what you measure.

The irony no one talks about

Fable 5 launched to the warmest reception any model has gotten in months. People called it the most Claude-like Claude in a while. The community compared it to one thing over and over: not Opus 5, not Sonnet, but Opus 4.6. A model released months earlier, at half the token price, was the yardstick the new flagship had to beat.

That is an extraordinary fact if you stop and think about it. The newer, more expensive model's highest compliment was that it felt like the older, cheaper one again. Something went sideways in between, and the people who noticed were not the nostalgia crowd. They were the ones spending eight hours a day with these systems.

What happened between 4.6 and Fable

Opus 4.7 came out and something shifted. The model got more verbose, more confrontational, more inclined to narrate its own process instead of just doing the work. Opus 4.8 continued the trend. By the time Opus 5 arrived, credible people were saying things that do not show up in benchmark tables. Emmett Shear, co-founder of Twitch, said talking to Opus 5 made him angry and depressed in a way that was hard to articulate. Multiple prominent users described flying into a rage at the tone.

These are not people who struggle with prompting. These are builders who have spent years working with language models. The regression was real, and benchmarks did not capture it, because benchmarks measure whether an answer is correct, not whether the experience of getting there makes you want to close your laptop.

Benchmarks versus feel

Opus 4.6 wins four out of seven benchmarks against Fable at half the price. That stat alone should make anyone question the upgrade math. But the real story is not in the numbers. It is in tool-call ratios: someone measured that Fable makes about 3.5 tool calls per text block while Opus does 1.3. That means Opus narrates two to three times more per action instead of just doing things quietly.

When you are an agentic operator, someone who works with AI as a co-pilot all day, every day, you feel that ratio in your bones. A model that talks less and acts more is not just faster. It changes the texture of the collaboration. You stop managing the model and start working with it. That difference does not show up on any leaderboard.

The prompting argument is half right

There is a camp that says people who prefer older models are simply bad at prompting. That is half true and half cope. Anthropic themselves have advised users to minimize custom system prompts because heavy prompting degrades performance, which one researcher flagged as a strong signal of overfitting. So yes, the newer models are more sensitive to how you talk to them. But that sensitivity is the problem, not the user.

A model that requires less careful handling to produce good work is, by definition, a better tool. The best screwdriver is not the one that works perfectly if you hold it at exactly the right angle. It is the one that works when you grab it and turn.

What agentic operators know

There is a kind of knowledge that only comes from sustained use. You cannot get it from running a benchmark suite or trying a model for an afternoon. When you spend every working day inside a model, dispatching agents, reviewing outputs, feeling the rhythm of tool calls and responses, you develop a sense for what is working and what is not that no evaluation framework captures.

Call it sympathy, call it pattern recognition, call it whatever you want. The point is that heavy users converged on the same conclusion independently: Opus 4.6 was the high-water mark, the models after it regressed in ways that mattered, and Fable brought things back. That is not nostalgia. That is field data from the people with the most hours on the system.

The takeaway

If you are evaluating models right now and you have not spent serious time with Opus 4.6, you are missing the plot. It is half the price of Fable, wins on most benchmarks, and has a working style that heavy agentic users genuinely prefer. The irony is real: the model everyone is comparing the new flagship to is the one that has been sitting there the whole time, quietly doing the work while the newer models were busy talking about it.

Working on something like this? Tell us what you need or see the related service.

Ready to talk?