# Grok 4.6 Shows How Fast Your AI Options Are Expanding — Transcript (2026-08-13)

https://aidailybrief.ai/e/2026-08-13 · Listen: https://pod.link/1680633614

---

260813 cold_EDIT: [00:00:00] A year ago, if you were talking about frontier models, pretty much you were referring to a model from one of either OpenAI, Anthropic, or Google By a couple of months ago, you were probably referring to a model just from either OpenAI or Anthropic.

Nathaniel Whittemore: Now, however, things have changed. over the past couple of months, any conversation about model performance has to include a recognition of Chinese open weight models that are pushing the frontier of both efficiency and cost. And as of this week, SpaceX AI's Grok is back in the conversation

260813 cold_EDIT: the just released Grok 4.6 is putting up benchmark numbers that put it in the category of a GPT 5.6 or a Fable 5. And doing so at a fraction of the cost. Although, of course, as we know, AI in the benchmarks tends to be very different than AI in the real world

After some initial testing, while users are not ready to declare Grok 4.6 a Fable or GPT class model yet, they are ready to argue fairly definitively that Grok and SpaceX AI are back [00:01:00] in the race 

260813 in_EDIT: The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Hyperagent, and Harbor

To get an ad free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors@aidailybrief.ai. And one other thing you should check out on aidailybrief.ai

As you know, we've recently updated the website, so now each episode has a full companion edition that includes all the key numbers, all the key quotes, all the key themes, each organized into different shareable cards that make it easy for you to find exactly the part that you wanna share with someone else

We have now added an archive as well to hopefully make it easier to find previous episodes about a particular theme It's organized on both an episode and a card basis and we'll be continuing to try to improve it as time goes on

Now with that out of the way, [00:02:00] let's get to the headlines, which are all about big money and into the change in the model landscape that's the subject of our main episode 

260813 hed_EDIT: Welcome back to the AI Daily Brief headlines edition, all the daily AI news you need in around five minutes. and the theme of today is big money. Cognition is seeking another funding round on the back of booming coding agent demand. Bloomberg reports that Cognition is in early talks with investors for new funding at a valuation of forty billion.

dollars. Cognition closed their last round just three months ago, raising a billion dollars at a twenty-six billion dollar valuation

For those doing the quick math, that means that the company's valuation would be up almost 50% in a quarter And the revenue figures seem to back it up. Sources familiar with the fundraising efforts said Cognition has doubled their revenue run rate to a billion dollars since they were last seeking funding.

One source said that Cognition is seeking a billion dollars in this round



260813 hed_EDIT: giving themselves a substantial increase in resources to address the current agent boom. The numbers also imply that the premium attached to coding agents is growing among venture investors. Cursor is one of the closest comps, and their last fundraising round in [00:03:00] March saw them seeking a $50 billion valuation on $2 billion in annualized revenue.

That round, of course, ended with SpaceX acquiring the company in a $60 billion all-stock deal And honestly, if Cognition has the ability to price their round at forty billion, The SpaceX deal could start to look like a bargain



260813 hed_EDIT: s- many think that the path that Cursor took with SpaceX feels inevitable for cognition as well

Writes Richard Wu, " "I wouldn't I wouldn't be surprised if within the next six to 12 months we see one of the hyperscalers preempt Cognition and offer to acquire them for 60 to 100 billion in stock. Given the success with SpaceX acquiring Cursor, the boards of these companies will put pressure on them to make a move."

Jeff Jeff Wu says, Google should buy Cognition for $200 billion and makeScott Wu CEO

Sandeep from Cognition responded, " "We We aren't selling. Also, 200 billion? The stock would move three times that in after-hours alone

alone." Next Next up, we have Lovable



260813 hed_EDIT: who who announced their $400 million Series C round at a $13.3 billion valuation

What's interesting is that you can clearly see how Lovable is evolving just in the [00:04:00] way that they describe themselves in their fundraising announcement

In In short

Lovable feels to me to be inching farther away from Claude Code and closer towards something like Shopify

They They write, " " Lovable is building the software creation platform that gives those closest to a problem the power to solve it, a generational opportunity that spans billions of people all over the world. For most people, turning an idea into software once required so much capital, technical fluency, and time that many ideas never came to life.



260813 hed_EDIT: Lovable's first chapter was about changing that."

Since our Series B in December 2025, we've been building features people need to reach customers, manage day-to-day operations, and run software securely For many builders, the product they create with Lovable is becoming the business itself. User survey data shows us that nearly eight in 10 are building a business or side project they hope to monetize, and more than one-third of those are already earning revenue.



260813 hed_EDIT: in in CEO Antoine Osika's post, he absolutely emphasizes the same idea Saying that Lovable will create, quote, "The most intuitive platform to build and run a business."

If you are looking for a place [00:05:00] to see the intersection of where what was once called vibe coding

meets the actual transformation of small and digital businesses, look no further than Lovable Lovable Now Now moving into public markets, business is booming for the Neo clouds as AI demand continues to rise. This week saw CoreWeave and Nebius report earnings, both vastly outstripping analyst expectations On On Tuesday night, CoreWeave reported that revenue had doubled over the past year to reach two point six billion for the quarter At the same time, cash burn also doubled, now running at five point seven billion per quarter.

Still, the big story for investors was a line out the door for compute. CoreWeave reported a one hundred and four billion dollar backlog in demand. In the footnotes, they added that the backlog had grown by twenty-five billion since they closed their books at the end of June.

The The story was the same for Nebius, who reported on Wednesday. They recorded four hundred and fifty-four percent revenue growth over the past year to reach five hundred and eighty-two million

Their cash burn is also escalating rapidly

But like CoreWeave, Nebius has endless demand With CEO Arkady Volozh telling investors, " Demand for what we are [00:06:00] building continues to be enormous. We could sell today our entire twenty twenty-seven capacity if we wanted." Supply is in fact so tight that Nebius is seeing huge profits on their available capacity.

Earnings per share beat analyst forecast by eighty-three percent And Volos told investors that their auctions for Blackwell Compute, which began in Q2, cleared at 15% above their previous record price for Hopper Compute. Markets rewarded both stocks with CoreWeave up 19% since reporting and Nebius gaining a staggering 34%.

Analysts believe that Neo clouds are some of the best indicators of marginal demand for AI as they service the overflow from the hyperscalers. And even during a quarter when token austerity came into vogue, demand is showing no signs of slowing

Meanwhile, Meanwhile, the infrastructure boom also is coming to China as Tencent has tripled their CapEx. Tencent reported that they spent seven point eight billion on AI infrastructure in the past quarter, boosting their training and inference fleet.

Now, Now, of course, that spending is still relatively modest compared to the US hyperscalers, where Meta had the slowest CapEx in their group and spent thirty-one point nine billion in Q2. [00:07:00] Still, there's a pretty clear attitude shift as the Chinese tech giants commit to scaling up their data center construction.



260813 hed_EDIT: during an earnings call on Wednesday, Chief Strategy Officer James Mitchell said, " "We We are allocating a very substantial portion of new compute to our own models and applications." The company's revenue is growing at eleven percent, but free cash flow has dipped into the negative, with incremental earnings going toward infrastructure.

Tencent President Martin Lau said that Tencent could monetize their compute by selling to outside customers if they wanted to, but for now they're prioritizing their own needs Basically, just like model training, it seems like China's AI build-out and the narratives around it are three to six months behind the US as well.



260813 hed_EDIT: It is uncanny how closely this is following the narratives from the US in Q1. Hyperscalers flipped negative free cash flow



260813 hed_EDIT: folks like Zuckerberg appeasing the market by telling them that he could sell his compute, but he doesn't want to

I'm not sure I think that US market participants have fully accounted for a Chinese CapEx boom

And what it does for the larger global investment environment

environment Meanwhile, Meanwhile, Samsung is seeing incredible efficiency gains from their use of AI in chip design. According to reports from a [00:08:00] Korean outlet, the first three months of integrating Claude Code into the software stack have been an outstanding success.

Development personnel have been able to cut down the time to complete complex tasks like system-on-chip verification from three months to two days In one example, a second-year engineer was able to complete a month-long task in a single day

Now, of Now, of course, this report doesn't claim that Claude Code produced efficiency gains throughout the entire chip design process

Nathaniel Whittemore: But 

260813 hed_EDIT: it does seem like an interesting example of the jagged frontier of AI adoption in the enterprise

Claude Code was able to make highly customized jobs more efficient and able to help a junior employee contribute way beyond their expertise



260813 hed_EDIT: Lastly Lastly today, some reported updates coming to the Trump administration's model testing framework. Last Tuesday, leading frontier labs were briefed on that framework, although the rest of us didn't get to learn all the details. It was reported that the policy would cover only state-of-the-art models, although we didn't know how exactly that was defined.

What we did hear with a fair degree of confidence was that the policy wouldn't cover open models. Open source advocates were relieved at that decision, but there was also a contingent of China hawks who believed that this would leave a gap. [00:09:00] On Wednesday, Wired reported that the administration has changed their mind.

An official said that the White House is expected to expand the policy to cover open models in the coming months. The policy, they added, is aimed at ensuring that as soon as open models reach the same capabilities as Mythos or GPT 5.6, they're added to the safety testing framework

White House officials said the administration had hoped the policy would be one and done, but the exponential development of model capabilities had forced them to iterate

When When it comes to the inclusion of open models, the thinking is that leaving them out of the framework could actually create a two-tiered system that would be negative for those open models. Specifically, officials are concerned that the framework could be viewed as a stamp of approval, leaving enterprises hesitant to use open models if they don't receive the same testing The The concern then is that leaving open models out could actually disincentivize US labs from developing those open models



260813 hed_EDIT: adding adding some evidence to the idea that the government is pro-US open models, Treasury Secretary Scott Bessent actually retweeted Mark Zuckerberg this week Saying, "We welcome Meta's release of Muse Glimmer, another win for American innovation.

Sustaining US leadership in AI means [00:10:00] advancing both open and closed weight models, ensuring the future is built on trusted foundations." Overall, it's still pretty clear that there's a lot of consternation around the administration policy. President Trump himself is reportedly insisting on keeping the framework voluntary, as he believes formal regulation will help China catch up.



260813 hed_EDIT: but by the same token, the safety-focused faction of the administration also isn't satisfied and are reportedly still pushing for a more formal arrangement

Who the heck knows how that's all going to turn out?



260813 hed_EDIT: But still this is a perfect segue to a broader discussion of the state of models. So for now, that's gonna do it for today's headlines. Next up, the main episode 

Hello everyone

Nathaniel Whittemore: One big change around AI is we've shifted our thinking from how we rank our pages to how do we become the source that AI trusts enough to answer with?

KPMG Aug_EDIT: At KPMG, they're seeing this firsthand. AI-generated results now surface answers directly, often without a single click. that's why they are increasingly focused on generative engine optimization or GEO, structuring content so AI systems can retrieve it, understand it, and cite it as trusted [00:11:00] authority this is not just an SEO evolution, but a visibility mandate. And indeed, the GEO mandate from KPMG is simple: If AI is shaping decisions, your expertise needs to show up inside the answer.

Read all about it at slash us/geo. Again, that is kpmg.com/us/geo



Blitzy's deep code base understanding unlocks the thing every roadmap owner cares about: shipping new features. Here's the truth about building inside a massive enterprise code base. Writing code was never the bottleneck. Context is Which system does this touch? Which contracts can't break? Which standards apply? Blitzy already knows because it reverse-engineered your entire code base into a dynamic knowledge graph before feature work began. With that complete picture, Blitzy builds features end to end.

blitzy July_EDIT: Architecture, APIs, UI, and tests all validated against your existing systems

One Blitzy customer built an AI native application from scratch with 100% autonomous completion, saving over 2,700 engineering hours. Features that respect your code base instead of fighting it [00:12:00] Stop letting your backlog grow faster than your team.

Accelerate your roadmap at blitzy.com. That's B-L-I-T-Z-Y.com 

Nathaniel Whittemore: This episode of the AI Daily Brief is brought to you by Hyperagent where you run fleets of agents your team can manage together New users get 1000 in inference Forget local agents and chat workflows waiting on your laptop to be prompted Hyperagent deploys alwayson agents in the cloud doing real work across the tools your team already uses Marketing's agent turns competitor moves into landing pages sales agent enriches leads drafts emails and updates the CRM ops agent chases the paperwork and tracks the budget Every agent has access to shared context and follows your rules about scope and approvals It's time you add agents that feel like teammates Hire yours at Hyperagent built by the team at Airtable Claim your 1000 in inference at hyperagent.com/aidailybrief. 

Every episode, I talk about the competition between OpenAI, Anthropic, SpaceX AI, Google, and Meta. And if you've been listening for a while, you might have a favorite. Maybe you think OpenAI and Anthropic can stay ahead, or [00:13:00] perhaps Meta's open source strategy can win out. Whatever your view, every AI lab creates a different investment opportunity. Harbor Capital Advisors AI Lab Ecosystem ETF suite lets you invest in the ecosystem behind the AI lab you believe in.

Harbor_EDIT: Search Harbor AI Lab Ecosystem ETFs wherever you invest or follow @HarborCapital on X to learn more. Visit harborcapital.com for a prospectus containing investment objectives, risks, fees, expenses, and other important information.

Read and consider it carefully before investing. Risks include principal loss and artificial intelligence related risks. Harbor ETFs are distributed by Foresight Fund Services LLC. Harbor is not affiliated with AI Daily Brief, and the funds are not affiliated with, sponsored by, or endorsed by any AI lab This is a paid advertisement and not personalized investment advice. Investing involves risk, including possible loss of principal. 

Welcome back to the AI Daily Brief. The big news that we are covering today is the release of Grok 4.6, which is getting some pretty good reviews out of the gate



260813 main_EDIT: but what's interesting to me is not just the model itself



260813 main_EDIT: But what it says about the state of the AI race and how that's [00:14:00] changing. Now, Now, the version of the AI race story that I am concerned with mostly here, of course, is the one that has to do not just with the achievement 

Nathaniel Whittemore: 

260813 main_EDIT: of some ill-defined far-flung goal like AGI or ASI

But the practical impacts on where different labs are for what we get to do with AI at home and in our companies

I I think CNBC's Deirdre Bosa

summed up the vibes when she tweeted yesterday, " What a difference a year makes. A year ago, frontier basically meant the big three US closed labs, OpenAI, Anthropic, and Google. Now a credible list includes xAI and multiple Chinese and open weight labs



260813 main_EDIT: implica- and while we'll get into the implications for the leading labs in a minute, I think Nathan Lambert also gets at another part of the sentiment when he writes, the vibe shifting from Anthropic is so far ahead to model competition back to all-time highs took like four weeks

So let's talk Grok 4.6 first short, the the release appears to put SpaceX AI squarely back in the frontier model competition Regular listeners will know that I take any release benchmarks not just with a grain of salt, but with an entire [00:15:00] bowlful



260813 main_EDIT: but but still, Groq's reported benchmarks are pretty hard to ignore

Nathaniel Whittemore: On 

260813 main_EDIT: on GDP Val,

which is the measure of how agentic AI performs on economically valuable tasks

SpaceX AI claims to have overtaken both GPT-56-Soul and Fable 5 By a small amount, yes, but taken over them nonetheless. Coding performance is improving as well

With Croc scoring right between, but in the range of Five, Six, Soul, and Fable five on cursor bench



260813 main_EDIT: being a few points behind on bothDeepSuite and Terminal Bench On On the overall artificial analysis intelligence index

Grok 4.6 jumped a full five points from Grok 4.5's 56to achieve an overall score of 61. That puts it ahead of Kimi K3 Tied with five six Seoul

And just a point or two behind Fable five and Opus five

What What that means is that if this was a new model from either Anthropic or or OpenAI, we'd probably be talking about how it's not quite state-of-the-art and didn't push the frontier forward But for SpaceX AI, who many had written out of the model race until fairly recently, this is a huge achievement, huge [00:16:00] summed up by the broad sense that you can see across AI circles that we once again have three frontier labs in the race

Also, while SpaceX AI has massively improved Grok's performance from 4.5, it seems like they're still working from the same base model as Grok 4.5. Pricing remains the same at two dollars per million input tokens and six dollars per million output tokens, Making it 60% cheaper than GPT-5 six Sole on a per token basis. 

of course, as we know, comparing tokens to tokens



Nathaniel Whittemore: 

260813 main_EDIT: is a seductive but ultimately fraught exercise given the massive differences in how many tokens different models might use to solve the same problem But But once again, artificial analysis' testing found that the model is,pretty token efficient as well

It completed the benchmark run at eighty-four cents per task, putting it in line with Kimi and making it 32% cheaper than GPT-5 6 Soul and 73% cheaper than Fable

Writes Writes investor Gavin Baker, "Absolute Pareto dominance for Groq and Cursor even after the OpenAI price cuts."

Now in Now in terms of reactions for the community

For many folks, [00:17:00] it was just gobsmacked at the achievement overall. Vittorio writes, " So they just caught up in three years? How does Elon do it?"

it?" Ben Ben Davis writes, " Grok 4.6 feels very good on first tests, very fast and capable and cheap, but time will tell as always. The cursor and SpaceX AI comeback is glorious to watch."



260813 main_EDIT: on Martin Casado from A16Z's highly technical tests

He found that it was strong Pavel Huryn writes, " Tried Grok four point six on my bug bench an hour after release. One hundred and five hidden bugs in two real repos judged blind."

His conclusion, looks like it may be my new default model. The best combination of time, value, and cost."



260813 main_EDIT: and and yet some folks did not have that same experience

Mehul Mohan writes, " Grok 4.6 is not as good as GPT 5.6 Sol in my 30 minutes of usage. It does incomplete work. Not incorrect, just incomplete. Maybe it's the Grok harness?

Justin Schroeder writes, " Early vibes on Grok 4.6 are not great. It's fast and is willing to do security work. I've already seen multiple instances where it [00:18:00] makes dangerous mistakes and later tries to cover up poor decisions. It even gets defensive.

Unfortunately, we cannot trust it."



260813 main_EDIT: entrepreneur Timmy McKegan writes, " writes, Grok 4.6 is one of the most oddly behaved models I've seen so far. It produces many times the output tokens compared to Terra or any similar intelligence model. It is cheap and fast, but takes everything extremely seriously and always investigates unclear information.

It values completeness above everything, including economics. the model seems to be designed to be economically viable, but acts differently."

Now, Now, when someone tried to clarify if this is a positive or a negative sign, Timmy kind of shrugged and said, "Probably positive?"

Benjamin DeCracker tried to sum up, " Lots of people acting like Grok 4.6 just beat Anthropic and OpenAI when really it didn't. The Grok 4.6 numbers show that xAI is not out of the race, but also not at the top. It's in the middle top-ish against models that the competition is already getting ready to update.

It shows that Grok still has a pulse, which is a good but different thing."

He He continues, "Or in sports terms, they advance past a critical [00:19:00] wildcard game into the playoffs, but are mid-rank against tough competition. They prove they can still hang, not yet winning everything."



260813 main_EDIT: and by and by the way, he clarified, " This is not a slight against Grok 4.6, which looks solid, just a read of the actual rankings and situation."

now, of course, what Benjamin is referring to is the fact that 4.6 is being compared against, GPT 5.6 and Fable 5 when both of those models are at this point several months old, and pretty much the only reason we don't have updates of them is that we're now past the threshold where the US government is going to be involved in every big new model release



260813 main_EDIT: And so state-of-the-art for us is very different from state-of-the-art at those top labs



260813 main_EDIT: however, however, it sounds like Grok 4.6 is itself just a waypoint

Elon Musk tweeted, "Grok 4.7 is significantly better than 4.6 and should be ready in three to four weeks. Initial training is complete, and now we're adding a massive amount of SpaceX company data in supplemental training. This will be something special." In another tweet, he said, " Grok 4.7 will exceed all current models.

That said, Anthropic is a great company and will probably release [00:20:00] improved models soon. However, the SpaceX training corpus is so awesome and unique that I would be shocked if any model is better at real-world engineering than 4.7."

Capturing Capturing the zeitgeist of credulity around these claims



260813 main_EDIT: Chubby shared both those posts and said, "I'm taking this seriously now. Grok 4.6 was the leap I've been hoping for. If the 10T model is still to come, then Elon's words can be taken seriously. It really could become the best model in general."

Although Although of course, Anthropic already has Fable 5.5 ready and just waiting to be released, that much is clear. Nevertheless, the next few weeks will be exciting, and xAI has shown just how much potential they possess



260813 main_EDIT: Leo Leo at Synthwave DD adds, " XAI have made an incredible comeback. From the days of Grok 4 to 4.3, where they were trailing the frontier by far, they're now arguably the third best lab in the world, behind only Anthropic and OpenAI



260813 main_EDIT: so where so where does this leave the rest of the field?



260813 main_EDIT: well first of all, there's Google, the company that many feel Anthropic has now overtaken as the definitive third place when it comes to state-of-the-art models [00:21:00] After last week's departure of Hassabis and longtime product leader Jeff Dean, many are basically counting Google completely out of the frontier AI race

The The counterpoint, however is that it appears that co-founder Sergey Brin is back in the picture to spur a comeback for Gemini. Reuters reported that Brin has become a key cheerleader for Google's AI team in recent months, encouraging AI engineers to catch up in the AI race.

He reportedly addressed a town hall after the release of Mythos, telling engineers it's time for Google to play catch up

Sergey Sergey had, of course, been out of the picture for several years after stepping down as president in 2019. However, he returned to frequent work at Google in 2023 and stepped into his involvement with the AI team in 2024, just as they were getting back on track ahead of the release of Gemini 2. During last week's news cycle, we had already heard that Google was relocating AI training out of the DeepMind office in London and back to the main campus in Mountain View.

That relocation would conveniently allow Brin to play a more active role working day-to-day with key researchers And of course, given what else we've heard about internal Google [00:22:00] politics, one of the big benefits to having Sergey fully engaged is that presumably he's one of the few people that could effortlessly cut through that bureaucracy to get things done at Google according according to the Reuters report that came out on Wednesday, that has already begun

Reuters writes, Brin has used the implicit power he holds as Google's co-founder to push resource allocation towards specific areas such as recursive self-improvement."

And to some, And to some, this is a good enough reason all on its own to not count Google out. Nick the CS guy from Google writes, " " Don't mess with Sergey, and definitely don't underestimate what he can do."



260813 main_EDIT: Still others think that Google is just temperamentally ill-suited

to this particular race. Computer science professor Pedro Domingos writes, " Hey Sundar, getting DeepMind to be an LLM lab is trying to shove a square peg into a round hole. You're destroying them, and you'll still lose the race. Let them focus on AI beyond LLMs, which is what they're good at, and create a nimble new lab to run the LLM race.



260813 main_EDIT: Now, Now, when it comes to what models we can expect next



260813 main_EDIT: I think at this point broad sentiment is that it would not be enough to recapture [00:23:00] momentum by releasing a competent Gemini 3.5 Pro at this point. 

Nathaniel Whittemore: 

260813 main_EDIT: we're already a couple months behind when we expected to get it, and just catching up I think would be seen as a failure

According to Leo and some other leakers I've seen

The reports are that teams are instead shifting to work on the scaled-up Gemini Four 

which while risky, I think does make sense in context

Now, Now, as Deirdre pointed out in thattweet at the top of this show



Nathaniel Whittemore: 

260813 main_EDIT: the top model labs question now has to necessarily include a bunch of entrants from China. And interestingly, just a few hours after Grok 4.6 launched, we got a significant leak out of China

Specifically, we got the benchmarks for the updated version of DeepSeek V4 Pro, and they appear on paper at least to be very competitive

For For example, these leaked benchmarks claim that the forthcoming model scored 87.9% on Terminal Bench 2.1, putting it just 0.1% behind Fable and 1.1% behind GPT 56 Soul. It also claims to beat Fable by 0.2% on Cyber Gym, the main cybersecurity benchmark

Now, Now, as always, there's the risk that this is just benchmark maxing and actual [00:24:00] performance will feel a little flat And unfortunately, almost as soon as these leaks started appearing



260813 main_EDIT: Other information came out Suggesting that the model was more significantly behind than the benchmarks would have it seem

seem Artificial Analysis's benchmark run was pretty disappointing, with V4 Pro scoring just 53. That's only one point ahead of V4 Flash and trails behind Kimmy K3 and Musepark 1.2 On the plus side, the model is pretty cheap even after DeepSeek delivered a substantial price increase this morning

At a buck 32 per million input and 396 per million output, it's about 1/12 the price of Fable and slightly cheaper than Muse Spark People's And people's first impressions also aren't that great. Lucky Faraday writes, " DeepSeek V4 Pro is benchmark slop. I had high hopes for this model, but it's complete trash. This was supposed to be a fable level model, and it can't even make a simple Minecraft clone. Even DeepSeek V4 Flash did a better job."

I know a Minecraft clone isn't a good test for a model, but come on, this is complete nonsense

And And before the "don't compare a less than $1 output model to frontier model" replies, [00:25:00] they are the ones comparing themselves to the frontier, not me

Still others pointed out when we're discussing models in the second half of 2026, 

it is less about raw performance alone and more about where they fit in the model stack

Dax from OpenCode says, "DeepSeek is insanely good at inference, using about two times less GPU time." And Augustin Lebrun writes, " I'm sure Kimi K3 and Grok 4.6 and DeepSeek V4 Pro are benched max more than Fable and GPT, but it doesn't matter. These models are an order of magnitude cheaper. As the frontier proceeds, fewer and fewer people need the bleeding edge and need it less often

Nowand at first glance

Ramp's latest AI index seemsto provide some evidence of that

Ramp's Ramp's lead economist, Ara Karazian writes New from Ramp AI Index, disappointing adoption of Fable 5. We've heard several reasons from businesses, mainly Fable 5 is just too expensive. a model so powerful it was briefly banned, and yet businesses don't think it's worth the price

Specifically, Ramp found that Fable 5 has made up only 6% of tokens that businesses purchased from [00:26:00] Anthropic And represented only 11.4% of dollars spent on anthropic models. For comparison, they write, " OpenAI's GPT-56 Soul comprises 25% of OpenAI tokens and 23% of spend."

In In fact, they say Fable 5 is less popular with businesses than GPT 5.6 overall



260813 main_EDIT: Ramp argues that, quote, " With Fable-05, we found a new upper bound to how much businesses are willing to spend on AI. Here, more performance is not worth the price tag. To encourage business adoption of the latest models, the labs will need to prove performance beyond even what Fable-05 is able to achieve, and simultaneously ensure that competitors aren't able to come reasonably close.

That seems increasingly out of reach, especially as open source models catch up to being only a few months behind."



260813 main_EDIT: However, I think that story is much less clear than they're letting on. First of all, as Simon Smith points out, ramp data overall suffers from selection bias, and this data suffers from it even more. This data comes from their token and spend management product, meaning users are predisposed to focus on cost control.



260813 main_EDIT: Fable simply isn't cost-effective for most [00:27:00] tasks. In other words, this is an extremely enfranchised set of users who are specifically using this in a product that is designed to manage spend and optimize spend away from models that are more powerful than you need

Rather than being a general assessment across a wide cross-section of businesses and business use cases Still Still to me, that isn't even the most damning thing, as perhaps one could argue that those companies and that type of spend management are a leading indicator of where others will get

I I think the bigger and more obvious issue is that Fable 5 still comes with a 30-day data retention policy

And most businesses aren't willing to touch that with a 39 and a half foot pole Indeed, indeed, Era actually came back to Twitter and retweeted himself



Nathaniel Whittemore: 

260813 main_EDIT: to add this incredibly important detail saying, " A lot of replies from employees who say they aren't allowed to use Fable because Anthropic is required to retain prompts for 30 days for US government safety checks."



260813 main_EDIT: look, it is absolutely the case that the more sophisticated buyers get, the less they're just gonna smash on the state-of-the-art model at the highest effort level for every single prompt

But the data retention policy [00:28:00] really makes this Not a particularly clear comparison

Now Now lurking behind everything we've discussed in today's show



260813 main_EDIT: is the fact that Anthropic and OpenAI both have more advanced modelsmore or less ready to go at this point that are that are being held back by a combination of government pressure, internal concern, or simply the fact that because nothing else is caught up, they don't really have pressure to move things forward faster

Still, even if on the one hand we are seeing a slowdown in the speed with which Anthropic and OpenAI specifically are dropping models, I think it's pretty hard to look around the model landscape right now and not feel like we have increasingly more rather than less choice

is gonna... Anyways, friends, some fun new treats to try for the weekend, but that is gonna do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace. 

​ 

[00:29:00] 

Nathaniel Whittemore's audio recording:
