Grok 4.6 Shows How Fast Your AI Options Are Expanding

2026-08-13 19:55:43 • 29:01

-

A year ago, if you were talking about frontier models, pretty much you were referring to a model from one of either open AI and

0:06

Thropic or Google. By a couple of months ago, you were probably referring to a model just from either open AI or

0:13

anthropic. Now however, things have changed. Over the past couple of months, any conversation about model performance

0:20

has to include a recognition of Chinese open weight models that are pushing the frontier of both efficiency and cost.

0:28

And as of this week, SpaceX AI's GROC is back in the conversation.

0:33

The just released GROC 4.6 is putting up benchmark numbers that put it in the category of a GPT 5.6 or a Fable 5

0:40

and doing so at a fraction of the cost. Although of course, as we know, AI in the benchmarks tends to be very different than AI in the real world.

0:48

After some initial testing, while users are not ready to declare GROC 4.6 a Fable or GPT class model yet,

0:54

they are ready to argue fairly definitively that GROC and SpaceX AI are back in the race.

1:01

The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.

1:13

Alright friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG,

1:18

Blitzy, Hyper Agent and Harbor. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief

1:24

or you can subscribe and Apple Podcasts. So learn more about sponsoring the show, send us a note at sponsors at

1:29

aiDailyBreathe.ai. And one other thing you should check out on AIDailyBreathe.ai. As you know, we've recently updated the website,

1:35

so now each episode has a full companion edition that includes all the key numbers, all the key quotes, all the key themes,

1:41

each organized into different shareable cards that make it easy for you to find exactly the part that you want to share with someone else.

1:47

We have now added an archive as well to hopefully make it easier to find previous episodes about a particular theme.

1:53

It's organized on both an episode and a card basis and we'll be continuing to try to improve it as time goes on.

1:58

Now with that out of the way, let's get to the headlines which are all about big money and into the change in the model landscape

2:04

that's the subject of our main episode. Welcome back to the AI Daily Brief headlines edition, all the daily AI news you need in around five minutes,

2:11

and the theme of today is big money. Cognition is seeking another funding round on the back of booming coding agent demand.

2:19

Bloomberg reports the cognition is in early talks with investors for new funding at evaluation of $40 billion.

2:25

Cognition closed their last round just three months ago, raising a billion dollars at a $26 billion valuation.

2:31

For those doing the quick math, that means that the company's valuation would be up almost 50% in a quarter.

2:36

And the revenue figures seem to back it up. Sources familiar with the fundraising effort said cognition has doubled their revenue run rate to a billion dollars since they were last seeking funding.

2:44

One source said that cognition is seeking a billion dollars in this round, giving themselves a substantial increase in resources to address the current agent boom.

2:51

The numbers also imply that the premium attached to coding agents is growing among venture investors.

2:57

Cursor is one of the closest comps and their last fundraising round in March saw them seeking a $50 billion valuation on two billion in annualized revenue.

3:04

That round of course ended with SpaceX acquiring the company in a $60 billion all stock deal.

3:09

And honestly if cognition has the ability to price their round at $40 billion, the SpaceX deal could start to look like a bargain.

3:16

Many think that the path that cursor took with SpaceX feels inevitable for cognition as well.

3:21

Wright's Richard Wu, I wouldn't be surprised if within the next six to 12 months we see one of the hyperscalers preempt cognition and offer to acquire them for $60 to $100 billion in stock.

3:30

Given the success with SpaceX acquiring cursor, the boards of these companies will put pressure on them to make a move.

3:35

Jeff Wu says Google should buy cognition for 200 billion and make Scott Wu CEO.

3:41

Sun deep from cognition responded, we aren't selling. Also 200 billion, the stock would move three times that in after hours alone.

3:48

Next up, we have Lovable who announced their $400 billion Series Sea round at a 13.3 billion dollar valuation.

3:56

What's interesting is that you can clearly see how Lovable is evolving just in the way that they describe themselves in their fundraising announcement.

4:02

In short, Lovable feels to me to be inching farther away from clawed code and closer towards something like Shopify.

4:09

They write, Lovable is building the software creation platform that gives those closest to a problem the power to solve it.

4:16

A generational opportunity that spans billions of people all over the world.

4:19

For most people, turning an idea into software once required so much capital, technical fluency, and time that many ideas never came to life.

4:26

Lovable's first chapter was about changing that.

4:28

Since our Series B in December 2025, we've been building features people need to reach customers, manage day to day operations, and run software securely.

4:36

For many builders, the product they create with Lovable is becoming the business itself.

4:40

User survey data shows us that nearly 8 and 10 are building a business or side project they hook to monetize, and more than one third of those are already earning revenue.

4:49

In CEO, CEO, Antoine Oseko's post, he absolutely emphasizes the same idea, saying that Lovable will create, quote, the most intuitive platform to build and run a business.

4:58

If you are looking for a place to see the intersection of where what was once called vibe coding meets the actual transformation of small and digital businesses, look no further than Lovable.

5:09

Now moving into public markets, businesses booming for the Neo Clouds as AI demand continues to rise.

5:15

This week saw CoreWeave and Nebius report earnings, both vastly outstripping analyst expectations.

5:21

On Tuesday night, CoreWeave reported that revenue had doubled over the past year to reach 2.6 billion for the quarter.

5:26

At the same time, Cashburn also doubled, now running at 5.7 billion per quarter.

5:30

Still, the big story for investors was a line out the door for compute.

5:34

CoreWeave reported a $104 billion backlog in demand.

5:38

In the footnotes, they added that the backlog had grown by 25 billion since they closed their books at the end of June.

5:43

The story was the same for Nebius who reported on Wednesday.

5:46

They recorded 454% revenue growth over the past year to reach 582 million.

5:51

Their Cashburn is also escalating rapidly.

5:54

But like CoreWeave, Nebius has endless demand.

5:56

With CEO Arcade Veloz telling investors, demand for what we are building continues to be enormous.

6:01

We could sell today our entire 2027 capacity if we wanted.

6:05

Supply is in fact so tight that Nebius is seeing huge profits on their available capacity.

6:09

Earnings per share beat analysts forecast by 83%.

6:13

Veloz told investors that their auctions for blackwell compute which began in Q2

6:17

cleared at 15% above their previous record price for hopper compute.

6:20

Markets rewarded both stocks with CoreWeave up 19% since reporting and Nebius gaining a staggering 34%.

6:27

Analysts believe that neoclads are some of the best indicators of marginal demand for AI as they

6:31

service the overflow from the hyperscalers.

6:33

And even during a quarter when token austerity came into vogue,

6:36

demand is showing no signs of slowing.

6:39

Meanwhile, the infrastructure boom also is coming to China as Tencent has tripled their cap-ex.

6:44

Tencent reported that they spent 7.8 billion on AI infrastructure in the past quarter,

6:48

boosting their training and inference fleet.

6:50

Now of course that spending is still relatively modest compared to the US hyperscalers,

6:54

where Meta had the slowest cap-ex in their group and spent 31.9 billion in Q2.

6:59

Still there's a pretty clear attitude shift as the Chinese tech giants commit to scaling up their

7:03

data center construction. During an earnings call on Wednesday, Chief Strategy Officer James Mitchell said,

7:08

We're allocating a very substantial portion of new compute to our own models and applications.

7:13

The company's revenue is growing at 11%, but free cash flow has dipped into the negative

7:17

with incremental earnings going toward infrastructure. Tencent President Martin Laos said that

7:21

Tencent could monetize their compute by selling to outside customers if they wanted to,

7:25

but for now they're prioritizing their own needs.

7:28

Basically just like model training, it seems like China's AI build-out and the narratives

7:31

around it are three to six months behind the US as well. It is uncanny how closely this is

7:37

following the narratives from the US in Q1. Hyperscalers flipped negative free cash flow,

7:42

folks like Zuckerberg appeasing the market by telling them that he could sell his compute,

7:45

but he doesn't want to. I'm not sure I think that US market participants have fully

7:49

accounted for a Chinese cap-ex boom and what it does for the larger global investment environment.

7:54

Meanwhile, Samsung is seeing incredible efficiency gains from their use of AI and chip design.

7:58

According to reports from a Korean outlet, the first three months of integrating

8:01

Quad Code into the software stack have been an outstanding success. Development personnel have

8:06

been able to cut down the time to complete complex tasks like system-on-chip verification from

8:10

three months to two days. In one example, a second year engineer was able to complete a month-long

8:14

task in a single day. Now of course this report doesn't claim that Quad Code produced efficiency

8:19

gains throughout the entire chip design process, but it does seem like an interesting example of

8:23

the jagged frontier of AI adoption in the enterprise. Quad Code was able to make highly customized

8:28

jobs more efficient, and able to help a junior employee contribute maybe on their expertise.

8:33

Lastly today, some reported updates coming to the Trump administration's model testing

8:37

framework. Last Tuesday, leading frontier labs were briefed on that framework, although the

8:41

rest of us didn't get to learn all the details. It was reported that the policy would cover only

8:45

state-of-the-art models, although we didn't know how exactly that was defined. What we did hear

8:50

with a fair degree of confidence was that the policy wouldn't cover open models.

8:54

Open source advocates were relieved at that decision, but there was also a contingent of China

8:57

hawks who believed that this would leave a gap. On Wednesday, Wired reported that the administration

9:02

has changed their mind. An official said that the White House is expected to expand the policy

9:06

to cover open models in the coming months. The policy they added is aimed at ensuring that as

9:10

soon as open models reach the same capabilities as Mythos or GPT 5.6, they're added to the safety

9:15

testing framework. White has official said the administration had hoped the policy would be one

9:19

and done, but the exponential development of model capabilities had forced them to iterate.

9:23

When it comes to the inclusion of open models, the thinking is that leaving them out of the

9:27

framework could actually create a two-tiered system that would be negative for those open models.

9:31

Specifically, officials are concerned that the framework could be viewed as a

9:34

stamp of approval, leaving enterprises hesitant to use open models if they don't receive the same

9:38

testing. The concern then is that leaving open models out could actually disincentivize US labs

9:43

from developing those open models. Adding some evidence to the idea that the government is pro-US

9:48

open models, Treasury Secretary Scott Besson actually retweeted Mark Zuckerberg this week, saying,

9:54

We welcome Metas release of Muse Glimmer, another win for American innovation.

9:57

Sustaining US leadership in AI means advancing both open and close-weight models, ensuring the future

10:02

is built on trusted foundations. Overall, it's still pretty clear that there's a lot of

10:06

consternation around the administration policy. President Trump himself is reportedly insisting

10:10

on keeping the framework voluntary as he believes formal regulation will help China catch up,

10:14

but by the same token, the safety-focused faction of the administration also isn't satisfied

10:18

and are reportedly still pushing for a more formal arrangement. Who the heck knows how that's all

10:23

going to turn out, but still this is a perfect segue to a broader discussion of the state of models.

10:27

So for now that's going to do it for today's headlines, next up the main episode.

10:35

Hello everyone, one big change around AI is we've shifted our thinking from how we rank our pages

10:41

to how do we become the source that AI trusts enough to answer with. At KPMG they're seeing this

10:46

first hand. AI generated results now surface answers directly often without a single click.

10:51

That's why they are increasingly focused on generative engine optimization or GEO,

10:55

structuring content so AI systems can retrieve it, understand it, and cite it as trusted authority.

11:00

This is not just an SEO evolution but a visibility mandate. And indeed the GEO mandate from KPMG

11:06

is simple. If AI is shaping decisions, your expertise needs to show up inside the answer.

11:12

Read all about it at kpmg.com slash us slash GEO again that is kpmg.com slash us slash GEO.

11:21

Blitzy's deep code based understanding unlocks the thing every roadmap owner cares about,

11:25

shipping new features. Here's the truth about building inside a massive enterprise codebase.

11:29

Writing code was never the bottleneck. Context is. Which system does this touch? Which contracts

11:34

can't break? Which standards apply? Blitzy already knows because it reversed engineered your

11:38

entire code base into a dynamic knowledge graph before feature work began. With that complete

11:43

picture, Blitzy builds features end to end. Architecture, APIs, UI, and tests all validated against

11:48

your existing systems. One Blitzy customer built an AI native application from scratch with 100%

11:53

autonomous completion, saving over 2700 engineering hours, features that respect your code base instead

11:59

of fighting it. Stop letting your backlog grow faster than your team. Accelerate your roadmap at

12:03

blitzy.com. That's B-L-I-T-Z-Y.com. This episode of the AI Daily Brief is brought to you by Hyper

12:10

Agent where you run fleets of agents your team can manage together. New users get $1,000 in

12:15

inference. Forget local agents and chat workflows waiting on your laptop to be prompted. Hyper Agent

12:20

deploys always-on agents in the cloud doing real work across the tools your team already uses.

12:25

Marketing's agent turns competitor moves into landing pages. Sales is agent and reaches leads,

12:29

drafts emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every

12:34

agent has access to shared context and follows your rules about scope and approvals. It's time you

12:39

add agents that feel like teammates. Higher yours at Hyper Agent built by the team at Air Table.

12:43

Claim your $1,000 in inference at hyperagent.com slash AI Daily Brief.

12:48

Every episode, I talk about the competition between OpenAI and Thropic, SpaceX AI, Google, and Meta.

12:54

And if you've been listening for a while, you might have a favorite. Maybe you think OpenAI

12:58

and Anthropic can stay ahead or perhaps Meta's open source strategy can win out. Whatever your

13:03

view, every AI lab creates a different investment opportunity. Harbor Capital Advisors AI Lab ecosystem

13:08

ETF suite lets you invest in the ecosystem behind the AI lab you believe in. Search Harbor AI Lab

13:14

ecosystem ETFs wherever you invest or follow at Harbor Capital on X to learn more. Visit

13:19

Harbor Capital.com for a prospectus containing investment objectives, risks, fees, expenses,

13:22

and other important information. Reading considerate carefully before investing.

13:25

Risks include principal loss and artificial intelligence related risks. Harbor ETFs are

13:29

distributed by four side fund services LLC. Harbor is not affiliated with AI Daily Brief and the

13:33

funds are not affiliated with sponsored by or endorsed by any AI lab. This is a paid advertisement and

13:37

not personalized investment advice. Investing involves risk including possible loss of principal.

13:46

Welcome back to the AI Daily Brief. The big news that we are covering today is the release of

13:50

GROC 4.6, which is getting some pretty good reviews out of the gate. But what's interesting to

13:55

me is not just the model itself, but what it says about the state of the AI race and how that's

13:59

changing. Now the version of the AI race story that I am concerned with mostly here of course,

14:04

is the one that has to do not just with the achievement of some ill-defined far-flung

14:08

goal like AGI or ASI, but the practical impacts on where different labs are for what we get to do with

14:14

AI at home and in our companies. I think CNBC's Deer Drabossa summed up the vibes when she tweeted

14:19

yesterday, what a difference a year makes. A year ago, Frontier basically meant the Big Three US

14:26

Closed Labs, OpenAI, Anthropic, and Google. Now a credible list includes XAI and multiple Chinese

14:33

and open-weight labs. And while we'll get into the implications for the leading labs in a minute,

14:38

I think Nathan Lambert also gets at another part of the sentiment when he writes,

14:41

the vibes shifting from Anthropic is so far ahead to model competition back to all-time highs

14:46

took like four weeks. So let's talk GROC 4.6 first. The release appears to put SpaceX AI

14:52

squarely back in the Frontier model competition. Regular listeners will know that I take any release

14:57

benchmarks not just with a grain of salt, but with an entire bullfull. But still, GROC's reported

15:02

benchmarks are pretty hard to ignore. On GDPVAL, which is the measure of how agent AI performs

15:07

on economically valuable tasks, SpaceX AI claims to have overtaken both GPT-56 sole and Fable 5

15:13

by a small amount, yes, but taken over them nonetheless. Coding performance is improving as well,

15:18

with GROC scoring right between, but in the range of 5.6 sole and Fable 5 on cursor bench,

15:23

being a few points behind on both deep-swee and terminal bench. On the overall artificial

15:28

analysis intelligence index, GROC 4.6 jumped a full five points from GROC 4.5's 56 to achieve an

15:35

overall score of 61, that puts it ahead of Kimi K3, tied with 5.6 sole, and just a pointer to

15:41

behind Fable 5 and Opus 5. What that means is that if this was a new model from either Anthropic

15:46

or OpenAI, we'd probably be talking about how it's not quite state of the art and didn't push

15:51

the Frontier forward, but for SpaceX AI, who many had written out of the model race until fairly

15:57

recently, this is a huge achievement, summed up by the broad sense that you can see across AI

16:02

circles that we once again have three Frontier labs in the race. Also, while SpaceX AI has massively

16:09

improved GROC's performance from 4.5, it seems like they're still working from the same base model

16:14

as GROC 4.5. Pricing remains the same at $2 per million input tokens and $6 per million output

16:20

tokens, making it 60% cheaper than GPT-56 sole on a per token basis. Of course, as we know,

16:26

comparing tokens to tokens is a seductive but ultimately fraud exercise, given the massive

16:31

differences in how many tokens different models might use to solve the same problem. But once again,

16:36

artificial analysis is testing found that the model is pretty token efficient as well.

16:40

It completed the benchmark run at $0.84 per task, putting it in line with Kimi K3,

16:45

and making it 32% cheaper than GPT-56 sole and 73% cheaper than Fable.

16:50

Right's investor Daven Baker, absolute Pareto dominance for GROC and Cursor even after the OpenAI

16:56

price cuts. Now, in terms of reactions for the community, for many folks, it was just gobsmacked

17:00

at the achievement overall. Vitorio writes, so they just caught up in three years? How does Elon

17:07

do it? Ben Davis writes, GROC 4.6 feels very good on first tests, very fast and capable and cheap,

17:14

but time will tell as always. The cursor and SpaceX AI come back as glorious to watch.

17:19

On Martin Kassato from A16Z's highly technical tests, he found that it was strong. Pueville Huron writes,

17:25

tried GROC 4.6 on my bug bench an hour after release, 105 hidden bugs and two real repos judged blind.

17:32

His conclusion looks like it may be my new default model, the best combination of time,

17:36

value, and cost. And yet, some folks did not have that same experience.

17:40

Neh-Hum-Lohan writes, GROC 4.6 is not as good as GPT-56 sole in my 30 minutes of usage. It does

17:46

incomplete work, not incorrect, just incomplete. Maybe it's the GROC harness? Just in Trotor writes,

17:53

early vibes on GROC 4.6 are not great. It's fast that it's willing to do security work.

17:57

I've already seen multiple instances where it makes dangerous mistakes and later tries to cover

18:01

up poor decisions. It even gets defensive. Unfortunately, we cannot trust it.

18:06

Entrepreneur Timmy McKegan writes, GROC 4.6 is one of the most oddly-behaved models I've seen so far.

18:12

It produces many times the output tokens compared to Terra or any similar intelligence model.

18:17

It is cheap and fast, but takes everything extremely seriously and always investigates

18:22

unclear information. It values completeness above everything, including economics.

18:26

The model seems to be designed to be economically viable, but acts differently.

18:30

Now, when someone tried to clarify if this is a positive or a negative sign,

18:34

Timmy kind of shrugged and said, probably positive? Benjamin DeCracker tried to sum up,

18:38

lots of people acting like GROC 4.6 just beat Anthropic in OpenAI when really it didn't.

18:43

The GROC 4.6 numbers show that XAI is not out of the race, but also not at the top.

18:48

It's in the middle top of against models that the competition is already getting ready to update.

18:52

It shows that GROC still has a pulse, which is a good but different thing. He continues,

18:57

or in sports terms, they advance past a critical wildcard game into the playoffs,

19:01

but are mid-rank against tough competition. They prove they can still hang, not yet winning everything.

19:06

And by the way, he clarified, this is not a slight against GROC 4.6 which looks solid just to

19:11

read of the actual rankings in situation. Now, of course, what Benjamin is referring to is the fact

19:16

that 4.6 is being compared against GBT 5.6 and Fable 5. When both of those models are at this point,

19:22

several months old, and pretty much the only reason we don't have updates of them is that we're

19:26

now past the threshold where the US government is going to be involved in every big new model release.

19:31

And so state of the art for us is very different from state of the art at those top labs.

19:35

However, it sounds like GROC 4.6 is itself just a waypoint. Elon Musk tweeted,

19:40

GROC 4.7 is significantly better than 4.6 and should be ready in 3-4 weeks.

19:45

Initial training is complete and now we're adding a massive amount of SpaceX company data in

19:49

supplemental training. This will be something special. In another tweet he said,

19:53

GROC 4.7 will exceed all current models. That said, and Thropic is a great company and will

19:58

probably release improved models soon. However, the SpaceX training corpus is so awesome and unique

20:03

that I would be shocked if any model is better at real world engineering than 4.7.

20:08

Capturing the zeitgeist of credulity around these claims,

20:11

Chubby shared both those posts and said, I'm taking this seriously now. GROC 4.6 was the leap

20:16

I've been hoping for. If the 10T model is still to come, then Elon's words can be taken seriously.

20:22

It really could become the best model in general. Although of course, Anthropic already has

20:26

Fable 5.5 ready and just waiting to be released that much is clear. Nevertheless, the next few weeks

20:31

will be exciting and XAI has shown just how much potential they possess. Leo at Synthwave

20:37

XAI have made an incredible comeback. From the days of GROC 4.4 to 4.3 where they were trailing

20:42

the frontier by far, they're now arguably the third best lab in the world, behind only Anthropic

20:48

and OpenAI. So where does this leave the rest of the field? Well, first of all, there's Google,

20:53

the company that many feel, Anthropic has now overtaken as the definitive third place when it

20:57

comes to state-of-the-art models. After last week's departure of DeepMind CEO Demisisabis and

21:03

longtime product leader Jeff Dean, many are basically counting Google completely out of the

21:07

frontier AI race. The counterpoint, however, is that it appears that co-founder Sergei

21:12

Brin is back in the picture to spur a comeback for Gemini. Rotter's reported that Brin has become

21:17

a key cheerleader for Google's AI team in recent months, encouraging AI engineers to catch up in

21:21

the AI race. He reportedly addressed a town hall after the release of Mythos, telling engineers

21:26

it's time for Google to play catch-up. Sergei had of course been out of the picture for several

21:31

years after stepping down as president in 2019, however, he returned to frequent work at Google

21:35

in 2023 and stepped into his involvement with the AI team in 2024, just as they were getting back

21:40

on track ahead of the release of Gemini too. During last week's news cycle, we had already heard

21:45

that Google was relocating AI training out of the DeepMind office in London and back to the

21:49

main campus in Mountain View. That relocation would conveniently allow Brin to play a more active role

21:54

working day-to-day with key researchers. And of course, given what else we've heard about internal

21:59

Google politics, one of the big benefits to having Sergei fully engaged is that presumably he's

22:05

one of the few people that could effortlessly cut through that bureaucracy to get things done

22:09

at Google. According to the Reuters report that came out on Wednesday that has already begun.

22:13

Reuters writes, Brin has used the implicit power he holds as Google's co-founder to push resource

22:18

allocation towards specific areas such as recursive self-improvement. And to some, this is a good

22:23

enough reason all on its own to not count Google out. Nick the CS guy from Google writes,

22:28

don't mess with Sergei and definitely don't underestimate what he can do.

22:32

So others think that Google is just temperamentally ill-suited to this particular race.

22:37

Computer science professor Pedro Domingo's writes,

22:39

Hey Sundar, getting DeepMind to be an LLM lab is trying to shove a square peg into a round hole.

22:44

You're destroying them and you'll still lose the race. Let them focus on AI beyond LLMs,

22:48

which is what they're good at and create an imbal new lab to run the LLM race.

22:53

Now when it comes to what models we can expect next, I think at this point,

22:56

broad sentiment is that it would not be enough to recapture momentum by releasing a competent

23:01

Gemini 3.5 pro at this point. We're already a couple months behind when we expect it to get it,

23:06

and just catching up I think would be seen as a failure. According to Leo and some other leakers

23:11

I've seen, the reports are that teams are instead shifting to work on the scaled up Gemini 4,

23:15

which while risky I think does make sense in context.

23:19

Now as Deirdre pointed out in that tweet at the top of the show, the top model lab's question now

23:23

has to necessarily include a bunch of entrance from China. And interestingly just a few hours

23:28

after GROC 4.6 launched, we got a significant leak out of China. Specifically we got the benchmarks

23:33

for the updated version of DeepSeat V4 Pro, and they appear on paper at least to be very competitive.

23:39

For example, these leaked benchmarks claim that the forthcoming model scored 87.9% on terminal

23:44

bench 2.1, putting it just 0.1% behind Fable and 1.1% behind GPT-5.6 sold. It also claims to

23:51

beat Fable by 0.2% on CyberGym, the main cybersecurity benchmark. Now as always there's the

23:57

risk that this is just benchmark maxing and actual performance will feel a little flat.

24:01

And unfortunately, almost as soon as these leaks started appearing, other information came out,

24:06

suggesting that the model was more significantly behind than the benchmarks would have it seem.

24:11

Artificial analysis is benchmark run was pretty disappointing with V4 Pro scoring just 53.

24:16

That's only 1.1 ahead of V4 Flash, and trails behind Kimi K3 and MuSpark 1.2.

24:22

On the plus side the model is pretty cheap, even after DeepSeat delivered a substantial price

24:26

increase this morning. At a buck 32 per million input and 396 per million output, it's about

24:31

1.12th the price of Fable and slightly cheaper than MuSpark. And people's first impressions

24:36

also aren't that great. Lucky Faraday writes, DeepSeat V4 Pro is Benchmark's slop. I had high hopes

24:42

for this model but it's complete trash. This was supposed to be a Fable level model and it can't

24:46

even make a simple Minecraft clone. Even DeepSeat V4 Flash did a better job. I know a Minecraft clone

24:51

isn't a good test for a model but come on, this is complete nonsense. And before the don't compare

24:56

a less than $1 output model to Frontier model replies, they are the ones comparing themselves to

25:01

the Frontier, not me. Still others pointed out that when we're discussing models in the second

25:06

half of 2026, it is less about raw performance alone and more about where they fit in the model stack.

25:12

Dax from Opencode says DeepSeat is insanely good at inference, using about two times less GPU time.

25:17

And Augustine LeBron writes, I'm sure Kimmy K3 and GROC 4.6 and DeepSeat V4 Pro are Benchmarks

25:23

more than Fable and GPT, but it doesn't matter. These models are an order of magnitude cheaper.

25:29

As the Frontier proceeds, fuel and fewer people need the bleeding edge and need it less often.

25:34

And at first glance, Rampslatus AI Index seems to provide some evidence of that. Rampslate

25:39

economist Arakerazean writes, New from Rampai Index, Disappointing adoption of Fable 5.

25:45

We've heard several reasons from businesses, mainly Fable 5 is just too expensive.

25:49

A model so powerful it was briefly banned and yet businesses don't think it's worth the price.

25:54

Specifically Ramp found that Fable 5 has made up only 6% of tokens that businesses purchased

25:59

from Anthropic and represented only 11.4% of dollar spent on Anthropic models.

26:04

For comparison, they write, OpenAI's GPT 56 sole comprises 25% of OpenAI tokens and 23% of spend.

26:12

In fact, they say Fable 5 is less popular with businesses than GPT 5.6 overall.

26:16

Ramp argues that quote,

26:18

With Fable 5, we found a new upper bound to how much businesses are willing to spend on AI.

26:23

Here, more performance is not worth the price tag.

26:25

To encourage business adoption of the latest models, the labs will need to prove performance

26:29

beyond even what Fable 5 is able to achieve and simultaneously ensure that competitors aren't

26:33

able to come reasonably close. That seems increasingly out of reach,

26:37

especially as open source models catch up to being only a few months behind.

26:40

However, I think that story is much less clear than they're letting on.

26:44

First of all, assignment Smith points out,

26:46

Ramp data overall suffers from selection bias and this data suffers from it even more.

26:51

This data comes from their token and spend management product,

26:54

meaning users are predisposed to focus on cost control.

26:58

Fable simply isn't cost effective for most tasks.

27:00

In other words, this is an extremely enfranchised set of users

27:04

who are specifically using this in a product that is designed to manage spend

27:09

and optimize spend away from models that are more powerful than you need,

27:13

rather than being a general assessment across a wide cross-section of businesses and business use cases.

27:18

Still to me, that isn't even the most damning thing,

27:20

as perhaps one could argue that those companies in that type of spend management

27:24

are a leading indicator of where others will get.

27:26

I think the bigger and more obvious issue is that Fable 5 still comes with a 30-day data

27:30

retention policy and most businesses aren't willing to touch that with a 39-and-a-half-foot pole.

27:36

Indeed, error actually came back to Twitter and retweeted himself to add this

27:39

incredibly important detail saying,

27:41

a lot of replies from employees who say they aren't allowed to use Fable because Anthropic

27:45

is required to retain prompts for 30 days for US government safety checks.

27:49

Look, it is absolutely the case that the more sophisticated buyers get,

27:53

the less they're just going to smash on the state-of-the-art model

27:55

at the highest effort level for every single prompt.

27:58

But the data retention policy really makes this not a particularly clear comparison.

28:02

Now lurking behind everything we've discussed in today's show is the fact that Anthropic and OpenAI

28:07

both have more advanced models, more or less ready to go at this point,

28:11

that are being held back by a combination of government pressure,

28:15

internal concern, or simply the fact that because nothing else is caught up,

28:19

they don't really have pressure to move things forward faster.

28:22

Still, even if on the one hand we are seeing a slowdown,

28:26

in the speed with which Anthropic and OpenAI specifically are dropping models,

28:30

I think it's pretty hard to look around the model landscape right now,

28:33

and not feel like we have increasingly more rather than less choice.

28:37

Anyways friends, some fun new treats to try for the weekend,

28:40

but that is going to do it for today's AI Daily Brief.

28:42

Appreciate you listening or watching as always, and until next time, peace!