This website uses cookies

Read our Privacy policy and Terms of use for more information.

Prime Ai Solutions

Read time: 4 minutes

Leader, welcome back.

Two things happened this week that matter more together than apart.

Claude's most capable model, Fable 5, loses its place in your subscription on July 12.

And today, July 9, OpenAI answered with its own three-tier lineup.

If you only read one AI story this week, make it this one. It's about to change how much your team pays for the same work, and it means the "just use whichever model is smartest" habit is officially retired.

LATEST NEWS
Fable 5's free ride ends Saturday.

What happened: Fable 5, Anthropic's top-tier model, has been included in Claude Pro, Max, and Team plans since it came back online on July 1 (it was briefly pulled worldwide in June over a US export review).
That included access ends July 12. After that, every Fable 5 call gets billed separately: $10 per million input tokens, $50 per million output tokens.
A single long back-and-forth conversation can run you $10 to $15 in credits on top of your normal plan!

Why it matters: most people have been using Fable 5 like a chat model, asking it quick questions all day.
That's the expensive way to use it.
Fable 5 is built for the tasks that take a person days or weeks, not the ones that take five minutes.

My take: don't cancel your Fable access, but change how you use it.

Treat it like a senior partner you bring in for the plan, not the paperwork

Give it the hard, high-stakes job (a full migration, a messy decision with real money on the line, an audit of work you've already sunk weeks into) and let a cheaper model like Sonnet 5 handle the execution once Fable has set the direction.

Anthropic's own data backs this: teams running Sonnet 5 with Fable 5 checking in occasionally get roughly 92% of Fable's quality at about 63% of the cost.
Plan expensive, execute cheap.

GPT-5.6 launched publicly today, and the shape looks familiar.

What happened: OpenAI's new GPT-5.6 series, Sol (flagship), Terra (built for high-volume work), and Luna (fast and cheap for everyday tasks), cleared for public release this week after a US government review of its code-vulnerability capabilities. That's the same process Anthropic went through with Fable and Mythos.

Why it matters: this isn't a coincidence. Both labs are now running the same playbook: one frontier model gated behind a government safety review, tiered underneath by cost and speed, with the expectation that you'll route your own work between them instead of defaulting to "the smartest one" for everything.

My take: this is the new normal, not a one-off.
If you've built any habits around picking a model once and sticking with it, that era is over. Which brings me to the thing you should actually do about it.

STEAL THIS
The Model Bake-Off

A team ran the same batch of real work through Sonnet 5, Fable 5, and a rival lab's open-source model, then graded the output blind.
The humans picked the open-source model over Fable more than once. The AI graders picked Fable every single time.

That gap is everything.
Fable is genuinely the strongest model going for multi-step agent work, the kind of task where it strings several actions together like an assistant would.
But for the email, the summary, the first draft, the stuff most of your week is actually made of, the gap between models has nearly closed.
At that point taste wins, not horsepower. And now that Fable costs real money per call, running it on tasks a cheaper model handles just as well is money you didn't need to spend.

Here's the 20 minute exercise, run it once this month:

Pick one real task you'd do anyway.
A client email, a board summary, a first draft of something you're already writing this week.

Run the exact same prompt through two or three different models.
Doesn't need to be exotic: Sonnet 5 against whatever else you have access to.
Skip Fable for this test, it's the wrong tool for a task this size at its new pricing.

Judge on one question only. Which draft would you send with the fewest edits.
Not which one sounds the smartest. The one that saves you the most editing time is the one that wins.

Write down the result in one line. "Model A nailed the tone, Model B was faster." That note is worth more to you than any published benchmark, because it's about your work, not a leaderboard.

Do it again next month. The model that wins in July might not win in September and the only way to know is to keep testing and exploring new possibilities!

What This Means For You

If you're doing this on your own: the bake-off above isn't just about picking a model for one email. Use Fable / GPT5.6 for the planning, then let a cheaper model handle the actual building, and you can put together a working internal tool (a tracker, a client-facing calculator, a small dashboard) in an afternoon instead of a sprint.

If you're running a team: the same routing logic applies to your budget, not just your prompts. Decide who actually needs premium-tier access (the people doing genuinely hard, judgment-heavy work) and who's fine on the cheaper tier for day-to-day tasks. That one decision is the difference between a predictable AI line item and an Anthropic / OpenAI bill nobody signed off on.

Reply and tell me what you've built, or if you want a hand figuring out where your team's routing line should sit.

STEAL THIS TOO!
One Prompt Worth Stealing

Next time you're about to sign off on a real spend decision, run this first:

"I'm about to approve [spend decision] for [amount].
Here's my reasoning: [your case].
Before I commit, stress-test this like a skeptical board member would.
Give me the strongest argument against this spend, the most likely way it turns out to be a waste, and one question I should be able to answer but probably can't yet."

Two minutes. Cheaper than the mistake it might catch.

SIGNAL / NOISE

Signal: Claude Cowork just expanded to web and mobile. You can kick off a task on your laptop and approve it from your phone. Anthropic's own usage data says coding is under 9% of what people actually use it for, which tells you it was built for the rest of your job, not just the technical part of it.

Noise: Every "AI is coming for your job" headline isn't citing actual hiring data.
The one real study worth knowing: Ramp and Revelio Labs linked actual AI spending to actual workforce records across 21,559 US firms and found the heaviest AI adopters grew headcount 10.2% over two years, with entry-level hiring up 12%!

Firms that barely touched AI saw no real change either way.
Worth the caveat: this is correlation, not proof, and the heaviest adopters were already bigger and faster-growing before they started spending on AI.
Still, it's a real number, not a guess, and it points the opposite direction of most of the headlines.

I'm also opening something new on the training side: it builds around your actual stack and interests instead of a generic curriculum. If you're already in the program, this is free, you're already covered. If you're not, reply "training" and I'll send details.

That's the whole story this week: two frontier labs, two different pricing pressures, and the same lesson underneath both. Know someone still picking one AI tool and sticking with it out of habit? Forward this to them!

-Umar, Prime AI Solutions | primeai.solutions