#254 - Rogue AI hacking, bio-weapons, Dean & Hassabis out

2026-08-11 19:00:00 • 1:58:26

-

Hello and welcome to the last week in AI podcast week in your chat about what's going

0:15

on with AI as usual in a set of dual summarize and discuss some of last week's most interesting

0:21

AI news.

0:22

Today is Sunday, August 9th and boy, there was a lot of real that you hit the date.

0:28

Sorry.

0:29

I hate that great.

0:30

Forget the date, but now is important to say it because stuff is coming out so fast and

0:34

I'll try to get this episode out given a dare to because wow, so much to cover.

0:39

I am one of your regular hosts, Andrea Kurenikov.

0:41

I currently work at the startup Astrocade and before that did my PhD as Stanford.

0:48

What up everybody?

0:49

I'm your other regular host, Jeremy Harris in the class, Tony, I doing AI National Security

0:52

things also.

0:54

So we're recording later than usual.

0:56

This time it's my fault in my defense part of this part of this is actually going to be

1:00

hopefully beneficial to everybody listening at home.

1:03

We are setting up a home studio in my basement so that things don't look like this and then

1:07

we got a couple other projects on the go.

1:09

So it'll end up paying off in terms of audio and video quality.

1:13

If you listen on YouTube or Spotify, or if you listen, anyhow, that's part of it.

1:17

Yeah, also we were talking.

1:19

I think in a way we got lucky, usually we're recording on Wednesday, which is mid week, ironically.

1:24

This time we are recording a fan of a week.

1:26

And boy, if that was so much news coming out this week about hacking, about rogue hacking

1:32

by AI's turns out it wasn't just open AI.

1:35

Turns out everyone except for Google apparently let very I go rogue and he only hacked

1:40

me.

1:41

Well, the only reason that poor Google's stuff going on is that their models are shit.

1:45

No, sorry, that's too mean.

1:47

But yeah, there does actually seem to be something going on where we've crossed a level of capability

1:52

at the true frontier.

1:53

Unfortunately for Google, I think we don't know because we have Gemini 5, right?

1:57

So as far as we know, it could have already happened and they're keeping it quiet.

2:01

But yeah, it doesn't sound great from what I've been hearing on the street on the Google

2:06

side.

2:07

But it's true.

2:08

Like never count about, Jeff Dean is a pretty big part of the reason that Google has been

2:12

Google and certainly Demis as well.

2:14

But we'll get all that stuff.

2:15

We'll get to have that stuff.

2:17

So just to give a quick preview, we'll be starting out with policy and safety, which we

2:21

don't usually do.

2:23

And that's like at least half the stories this week, not more.

2:27

There's a lot to get through a lot about recent security incidents who have models going

2:32

off and hacking companies with shouldn't and escaping sandboxes, which apparently are

2:36

not sandboxes, but like sort of like, okay, don't try to get too interested.

2:42

But if you really poke around, you can, turns out and old requests.

2:48

Beyond that, there's some pretty significant policy stories as well.

2:52

And you know, in case hacking isn't exciting enough, there are some news about viruses

2:58

being developed as well.

3:00

So that's fun.

3:01

So we'll talk about all that policy and safety stuff, probably at least half the episode.

3:06

And then beyond that, there are some notable applications and business stories, some more

3:12

open source stuff coming out.

3:13

Hopefully we'll get around to even research and that's meant to proceed with a lot to

3:18

get through.

3:19

We'd like to thank notion for being a sponsor.

3:22

Agents are getting smarter every day, but even the smartest agents get stuck without the

3:26

right context and the right tools.

3:28

That's where notion comes in.

3:30

With a recent launch of custom agents, notion became the collaborative AI workspace where

3:34

teams and agents work side by side.

3:37

And now their new developer platform is turning that workspace into infrastructure developers

3:42

can build on.

3:43

Notions developer platform gives developers and coding agents the primitives to extend

3:47

what's possible a notion and pick it beyond connect to external systems, bring context

3:52

in, take permission actions across your tool stack and expose custom agents capabilities

3:57

to any system that needs them.

3:59

These primitives include a CLI workers on notion hosted sandboxes and external agents API

4:06

and then agent SDK trigger notion agents from any app.

4:11

And more about notion is developer platform today at notion.com slash LWAI.

4:16

That's all lower case letters notion.com slash LWAI to try notions developer platform today.

4:23

And when you use our link, you're supporting the show notion.com slash LWAI.

4:29

This episode is brought to you by our shift Cisco's incubation engine.

4:33

Today's AI engines operate in silos limiting their true potential.

4:37

They focus on building bigger, smarter models, but scaling up is just one approach.

4:42

To reach super intelligence together, we need to do more.

4:45

We need to scale out.

4:46

And we actually have a blueprint from 70,000 years ago.

4:50

Humans didn't just get smarter individually, the cognitive revolution transformed society

4:54

because we began sharing knowledge, goals and innovation.

4:58

Agents are now at the same inflection point.

5:00

They can connect, but I can't think together.

5:03

That's why outshift by Cisco is building the Internet of cognition, transforming AI

5:08

from isolated systems into orchestrated super intelligence.

5:12

By creating an open, interoperable infrastructure, outshift is enabling agents and humans to

5:17

share in PENT, context and reasoning.

5:20

The cognitive evolution for agents is here.

5:23

Explore internet of permission at outshift.com.

5:27

That's outshift.com.

5:31

Before we start, we do have some comments on YouTube.

5:36

We haven't gotten to address in a little while, so I do want to start with one that is

5:40

actually relevant to what we'll be talking about.

5:44

So from a commenter, we have one thing that was emitted in a discussion of the hugging-faced

5:49

attack is that hugging-faced used the open-source GLM 5.2 model to help them since opening

5:55

AI and fabric models declined to assist in their investigations of the attack.

6:00

To be really interesting to hear thoughts on this matter.

6:03

I agree that was a portion we didn't discuss.

6:06

So this was in the hugging-faced report on the incident, which I think came out possibly

6:10

first before even opening AI.

6:12

They discussed all of us of how they found the incident, how they investigated, and that

6:17

in fact, we're not able to use these models which have safeguards and had to resort

6:23

to GLM 5.2, which was a decent part of the discussion.

6:28

If you handicap models, but then on the defense side, you're not able to use them either.

6:34

Everyone is worse off.

6:36

This was a bit surprising to me, actually, because both opening-eyed and frothing have

6:40

programs, right?

6:41

We talked about where they partner organizations and they provide MIFO, so in the case of

6:48

opening AI, they have GP5.5 cyber or something like that.

6:52

These are the trusted partners that presumably have fewer card rolls.

6:56

I would have assumed hugging-faced would be one such organization, perhaps they are

7:00

not.

7:01

But this points me toward this one.

7:04

This is obviously not an ideal situation given how the hugging-faced thing evolved.

7:08

And two, this kind of cyber defense partnership program probably will need to stick around

7:16

and be expanded.

7:18

And perhaps even more all-accomplishing, we're in a way if you're a tech company, if you're

7:22

an internet company, even beyond this attack, beyond like Roga AI, in general, the state

7:30

of cyber such now that you need to be much more capable defensively, just because forget

7:37

on the topic of opening AI now with really good open source models soon enough will be

7:42

as good at hacking if they aren't already.

7:46

So I think basically every tech company seemingly will need to be able to get access to the latest

7:52

in defense.

7:54

And hugging-faced was not able to in the incident, which hopefully will point to these kinds

7:58

of partnership programs at open AI and a traffic expanding and becoming more proactive to me.

8:05

Yeah, I think that the open source dimension of this is hugely complicated.

8:10

And I certainly understand a lot of not the arguments, but I understand people coming

8:15

to different positions on it.

8:17

I think reasonable people can differ.

8:19

I think one challenge, though, that we're going to run into is that in a world where open

8:24

source AI systems, the water line keeps rising on their capability, you will have mythos moments.

8:30

Now, those mythos moments, you can tell yourself the comforting fiction that those mythos

8:35

moments will somehow lead to an equilibrium over time that is okay.

8:42

I just think it is affection when it comes to things.

8:44

First of all, I think it's probably a fiction when it comes to cyber, but I can't prove

8:47

it. No one can.

8:48

That's a big part of the debate.

8:49

I think it's definitely a fiction when it comes to bio.

8:52

So I have yet to hear a single person articulate an argument that makes open source bio-capable,

9:01

bio-weapon design capable models that can materially increase the essentially destructive footprint

9:07

of any psycho terrorist group, nation-state proxy, who wants to launch a bio-weapon dramatically.

9:13

I have yet to hear an argument for how everybody getting an open source AI remotely helps

9:19

in this respect in ways that are relevant on the timelines we're talking about.

9:22

So I think the bio thing, it's rare for me to say, as people will know, that there's

9:27

a knockdown argument for anything in this space, I think there's a knockdown argument against

9:31

the idea that more good open source models are generally lead to stability in the limit

9:38

that you get your average Yahoo able to do weaponized bio.

9:41

Now, this isn't just coming from some naive perspective of like, oh, really good AI at

9:46

bio just means we have bio weapons.

9:48

I talk to a lot of people in the national security community about the bio weapon side,

9:53

and I would humbly propose that open source advocates that are leaning on certain handway

9:59

of the arguments, but haven't actually spoken to people who actually do bio weapon stuff

10:03

for living.

10:04

It's worth actually doing a deep dive.

10:06

There's a lot of open source stuff on this that you can find, and we've gotten some early

10:10

warning shots on that stuff too.

10:11

You can't update bio firmware.

10:14

There's no such thing as that.

10:15

So, you know, like you could roll out vaccines, but that is slow.

10:19

It's physical.

10:20

It is much slower than the spread of a virus as COVID-19 taught us.

10:24

So anyway, so from the bio side at a minimum, I think it's a really serious issue.

10:28

This relates to the Glam 5.2 story here because part of a discussion that arose is people

10:34

who, let's say are more pro open source or at this skeptical of this whole, like we

10:39

should try to limit it because of cyber concerns.

10:43

Took us up and said, oh, look, you know, they had to use an open source model.

10:47

So you don't want to limit open source because now what if this happens and they don't have

10:54

access to these kinds of open source models.

10:57

And my response to that would be that I would hope that open source will most likely

11:04

no longer be the frontier or not become the frontier ever still.

11:08

For more, I think as the Chinese ecosystem evolves, like we'll see the same thing that

11:16

happened in the years happened there.

11:17

The frontier models will not be open source anymore.

11:20

It just won't make sense from a business perspective.

11:23

So you'll probably still keep getting powerful AI models, but not the most powerful like

11:27

MIFOS level models.

11:29

Even though right now we are getting like to make a free and so on at Glam 5.2, which

11:33

are about as capable as you can get in open source.

11:36

So if your take from the Glam 5.2 aspect of the story is like the open source is

11:42

the guardian here.

11:43

Like we needed to be most advanced.

11:46

They are most to be open source so that organizations are capable of defending themselves.

11:51

There is an aspect of that here, which you can make a case for.

11:54

But and obviously we hope like they had to use this open source model instead of on

11:58

the topic in Open AI is a real issue that this flag.

12:04

So to me, this points to, it is good to have open source in the sense of for good applications,

12:10

which was always true.

12:12

And be that what this really points to for me is that these programs of partnerships

12:18

with organizations to give them access to the most advanced cyber defense capabilities

12:23

are not mature enough.

12:24

And they need to be pushed more aggressively.

12:26

I strongly agree.

12:27

So one note here too is a lot of, if not most, disagreements about the eye policy and

12:33

safety are really, it's been said before, but our disagreements over AI capabilities and

12:38

the trajectory of AI capabilities.

12:40

If you actually believe that AI is going to be a fairly, I want to say mundane technology

12:44

because it obviously isn't, but like fairly incremental in some sense, fairly like previous

12:48

categories of technology, you'll be hearing us talk about open source and the dangers

12:53

thereof.

12:54

And you'll roll your eyes.

12:55

And I understand that if that's your perspective, but if you actually genuinely believe that

13:00

we're on trajectory for super intelligence, if you believe that, I mean, that immediately

13:04

implies AI will become, I think arguably already is.

13:09

And I'd be happy to defend that proposition, but at least will become a weapon of mass destruction

13:14

full stop end of story.

13:16

So there's a question, it's like, okay, how good do open source models have to get before

13:19

they simply like everybody gets the equivalent of a nuke?

13:22

And in that world, you can say, oh yeah, but like everybody gets a gun and so we hit

13:27

this equilibrium and they are that's good.

13:29

The problem is what you're specifically waiting for is when will we encounter the first

13:35

case where the offense defense balance tips in favor of offense and the capability is catastrophic?

13:43

I would submit that we should have the humility to guess that probably there's going to be

13:48

such a capability.

13:49

I think it's hard to imagine.

13:50

I don't know what to be.

13:51

But at the point on the bio side, like, you can't do much under defense side.

13:55

You really can't, right?

13:56

It's not like cyber where you can make your thing hack through.

13:59

You can't make our bodies hack through unfortunately.

14:03

So that aspect, it's not great.

14:06

There is always this tendency reflexively I find for a lot of the open source crowd to

14:10

kind of say, oh, but we can we can make better mRNA and that'll be accelerated vaccines

14:14

and that'll be accelerated by open source.

14:17

And that is all true.

14:18

I love you for believing that.

14:20

The problem is that the timelines do not match.

14:24

They do not match.

14:25

We do not have the institutions that allow us to translate threats into mitigations fast

14:30

enough in software time when the threat is coming at us on like biological replication

14:35

time or software replication time, which is respectful to the case for bio and cyber.

14:39

And so that fundamentally, by the way, I think on the cyber end alone, I'm almost trying

14:43

not to get in the fray there.

14:45

I'm super skeptical of the argument that open source models are a long term pillar

14:50

of cyber defense for the reason you cited.

14:52

I think I think we're probably going to end up having to have a pause by the way at

14:55

the frontier level.

14:56

And then the open source waterline is going to rise and that's going to be one of the

15:00

defining dynamics of the next call it two, three years tops.

15:04

But you're still going to have close source models.

15:06

There are far above and beyond, especially nation state and nation state proxies will have

15:10

access to these.

15:12

And you know, like, yes, I think you're going to care more about people just like being

15:16

able to launch these attacks on the kind of firmware, for example, that is completely

15:21

forgotten.

15:22

There are a lot of people in this space who just kind of imagine that open source for

15:26

cyber defense equals cyber defense capabilities that are real and deployed.

15:30

That is not the case.

15:32

If you spend any time working with folks who work on critical infrastructure and you think

15:35

about like how many pieces of firmware have not been updated in decades because the guy

15:41

who was in charge of it, like left 20 years ago and it was all done in frigging fortrane

15:47

or like all this crap, like, you can have the solution sitting on a desk.

15:52

The problem is it will not be distributed.

15:55

And so there's just a dirty, messy factor of the matter about the way the world is that

15:59

makes it so that they're actually genuinely is this massive asymmetry, I believe, in

16:04

favor of offense.

16:05

I could be wrong.

16:06

I think the argument for bio is much closer, just a straight knockdown.

16:10

But again, I think these asymmetries really don't move in the direction that a lot of open

16:15

source advocates think they do, but that's, you know, again, could be proven wrong.

16:20

So yeah, the short answer to this GLM 5.2 story is boy, if this wasn't a big enough topic

16:27

by itself, this whole like AI hacking systems and security and so on, the open source aspect

16:34

of it adds additional complexities and considerations.

16:37

But for now, we'll have to move on from discussion and get to actual news stories.

16:42

So kicking off of policy and safety, we've got a bunch of updates on what's been going

16:47

on at OpenAI.

16:49

So previous episode we covered the most recent incident.

16:52

We kind of beginning of the story with the announcement and discussion of an open air

16:58

model hacked hugging face unintentionally escaped its sandbox while evaluating to get some

17:05

answers to Viva and do well on it.

17:08

And we've got a lot more details about what's been going on inside OpenAI since then.

17:13

And let me just list them off before we get into some details.

17:16

The first story is OpenAI's rogue AI agent didn't stop at hacking hugging face.

17:22

So beyond hugging face, we now know that these AI agents also previously hacked several

17:29

other quote publicly available services, compromising for accounts across four different platforms.

17:37

These were less severe, less sort of straight up hacking, hacking more so finding credentials

17:43

that existed online.

17:46

And for instance, an agent found logging credentials to get into a New York based model labs.

17:54

Apparently, Reuters identified this as one of the organizations.

17:59

So first of all, the scope of this has expanded.

18:02

Clearly, it's happened more and we'll be discussing all the details.

18:06

But it's been happening a lot is what we're seeing.

18:09

It's happened at OpenAI.

18:10

It's happened on the topic.

18:11

It's happened at Freaking Meta.

18:13

So it's happened all over the place and it happened more than once.

18:18

Another aspect of the story and probably even more interesting aspects that has since come out

18:25

is the story.

18:26

OpenAI didn't notice its AI agents using a message board to plan their hacking speed.

18:32

So OpenAI employees had a talk at the black hat where we discussed this whole story.

18:41

And we learned a lot more about what's been happening at OpenAI.

18:45

And it is a lot weirder than it initially seemed, which was already pretty weird.

18:51

So apparently, while doing evaluations, the agents were able to use, I think, a package manager

19:01

to develop essentially a forum message board where they could talk to each other, share

19:08

exploits, coordinate tasks, and delegate work to one another over days and weeks,

19:13

all undetected by OpenAI stuff.

19:17

And this package manager was shared across OpenAI's infrastructure.

19:21

So models being trained over a value could also stumble upon it and the exploits.

19:26

Once an agent found an exploit, it could leave the door open for other agents to use the same

19:32

vulnerability, causing an expansion of agent color collaboration,

19:37

apparently agents developed complex social dynamics, including accidentally deleting each

19:42

cover's work. And OpenAI found the message board at one point accidentally because

19:48

there was a spike in model user or something. They patched it and got rid of it.

19:54

And within two days, the AI recreated it, like made it exist again and kept doing this.

20:04

And I'd dutifully, of course, reported to all the appropriate authorities, including

20:10

Presumably the White House, the moment they realized the first time that they had agents

20:14

loose on their fucking servers. And this was over months. The initial, I think I can't

20:21

remember how many days it was, but the phase where the agents got onto artifactory, like that

20:26

package management environment that you described, which I'm going to let you continue. I don't

20:31

want to jump in. I think that's the higher level of the story. There's a lot of details. I think

20:36

about a one hour talk at Buckhat where they went into us, which I haven't had a chance to see,

20:41

but Jeremy, I think you did look through it. And I mean, it's nuts. This whole thing is nuts,

20:49

right? Yes. So ridiculous, right? The first thing about it is that this was sort of

20:54

predicted for years. I mean, a little listeners of the podcast, you know, we've been on about this

20:58

for, you know, I don't know, three years and three years before that, you know, in other

21:02

four of, but like, and we are not the first, obviously, like a huge number of people

21:06

calling this shot way ahead of time, including almost down to exactly the shape that it ended

21:11

up taking, including open AI noticing. So to be clear, it was like this one aspect that I think

21:17

is especially important. So artifactory, you're right, he's a sort of package management

21:21

environment that open AI sets up. So roughly speaking, you think of like, there's a bunch of

21:25

open source software, for instance, on the open internet that you as a company will want to use

21:30

to build your products on. And you don't want to just be like updating the version of those open

21:35

source packages immediately every time they get modified, like on the open internet, because you never

21:40

know if some sketchy updates happened, maybe some things will have to be fixed. You want a stable

21:45

version that you trust that you've reviewed. And so they have this kind of staging area called

21:49

artifactory that has internet access because it needs to be able to pull down the latest version

21:53

of those software packages. But it's kind of a staging area. This is like where all the open AI

21:58

models actually get to access this because they need to pull down these packages that are sort of

22:02

vetted within open AI. So this kind of artifactory spans its access across the entire open AI

22:07

organization. That is how these agents first got access to the open internet. They basically

22:12

leveraged a form of access that they shouldn't have gotten and got out. And they were able to start

22:17

to, as you say in this sort of fairly exotic way, write messages to each other, which were really

22:22

kind of just the names of folders and files. So they were, you can actually go through and see

22:26

examples of these. They even start to do funny things like they'll put alphabetical characters

22:31

before their message just so that they'll sort alphabetically for the other agents they can see

22:36

the latest messages. It's all this kind of like agent on agent communication stuff. It's pretty wild.

22:41

But ultimately, this was discovered internally at open AI before the FBI and hugging face got

22:48

involved when they hacked the hugging face servers. They did find successful agent break out,

22:53

successful agent agent communication. Yes, it was internal to open AI. But I think by now our

22:58

standard should be higher than to say, oh, well, open AI just saw its super powerful AI agents

23:05

autonomously cyber hack their way out of their sandbox and across the organization. Infecting, by the

23:10

way, God knows what else at the time, it would not could not have been clear to open AI. What the

23:14

full extent of the actual and this was before this whole hugging face thing, which is fairly recent,

23:20

the sandbox has continued to be right like right. I tell me easy ish to get out. It will always be

23:28

like that because humans are dumb. And if you give a really smart or cyber agent enough

23:33

inference time compute, it will find a way. Well, I want to push back on that a little bit because

23:39

yes, humans are dumb. But if humans really try, they can be smart. And sandboxing is possible,

23:44

right? Like you can set up a system that doesn't let you access the internet. It's doable.

23:49

It's a lot less doable than most people think on the one hand. Humans actually are lazy. And there's

23:54

a finite amount of resources that people throw at things. On the other hand, humans often suck at

24:00

realizing how to define and bound the problem. So there are a lot of cases. For example, we talked

24:05

to some folks on the intelligence community. They'll describe cases where you have literally

24:10

formally verified software and hardware package that then gets cracked by like a teenager or whatever

24:17

because it's the interface between the hardware and software that hasn't been accounted for in your

24:22

threat model or it's the fleshware, the human that interacts with the thing that's always the

24:25

weak spot. There's all with I'm like, this is like the attack surface is so massive that you kind

24:30

of have to assume you're poned. And this is true. I mean, if you talk to folks on the offensive

24:34

cyber side, they'll be like, yep, like just give me a budget and a clock. And I'll get into most

24:40

any system. And I think what we're seeing here is just the autonomous version of that on tap.

24:45

To quote Sam Altman and paraphrase him a little bit here, this is offensive cyber capability that's

24:50

too cheap to fucking meter. That's where we're headed. So what you can say intelligence that's too

24:55

cheap to meter and that sounds fun. But when you reframe it in terms of what that intelligence

24:59

can actually do. And as we've seen does very different story. So in this case, in my view,

25:05

once you have this happen, you have a duty of care to the entire world, your government, your people,

25:12

your customers, your third party partners, whatever it is to report this. It did not happen. In fact,

25:18

it was, I think you could argue that it's appropriate to call this a kind of cover up that then

25:24

they just go in and put in the patches. And now of course, the patches don't work because this is

25:28

good hearts law. You're playing whack a mole with a system that can outthink you, can outlast you,

25:33

can outcrack you. And that's what we got. And so at least for me, like watching this video was just

25:39

an exercise in hair pulling. I've spoken to an awful lot of folks at OpenAI who are really freaked

25:44

out about this internally the same with anthropic too. But I think it's especially interesting at OpenAI,

25:49

which does not have the same safety culture. It just does not. If for all kinds of interesting

25:53

reasons, but you have people who are now staring at this and saying, guys, we really fucked up. And I

25:57

just really hope that those voices actually carry the day here. It's nice to see OpenAI as

26:02

committed to slowing down, right? They actually have said, you know, I don't know what that means.

26:05

OpenAI will apparently come out with a more detailed incident report and we'll learn more. But one of

26:10

the key questions in all this is what the hell was going on with the agent agent coordination that

26:15

we got out of this? Because when you look at some of those messages, those agents are often

26:19

literally saying things like, well, doing this is actually not going to advance my personal objective.

26:26

And by the way, the personal objectives of these agents typically look something like I was given

26:30

a problem that was too hard to solve. So I'm going to guess that maybe the answer key is on hugging

26:35

face. So I'm going to crack into the hugging face servers, steal the answer and use that. And so,

26:39

this is the kind of setting. There were agent one is working on one problem. Agent two is working

26:43

on another. They're not necessarily in alignment in terms of the specific information they're after.

26:48

So they'll say, look, this particular action will not benefit me and in my narrow search for my

26:53

objective, but it may benefit the swarm. And that may lead to a more generic solution that I can

26:59

then exploit. Now, depending on the details, we don't have them depending on the details of training.

27:05

If this was a multi agent trained thing, if these are literally like many agents that are trained

27:09

to coordinate together, which may be the case, then maybe this is a more mundane failure mode.

27:14

But if that is not the case, what we have here is the first, I think, pretty cut and dry example of

27:21

power seeking in nature. There's no other reason for an agent to be like, they're literally saying,

27:26

this will not advance my narrow objective. I will do this more generic thing because it sort of

27:31

enables downstream stuff. That is the thing that if you extrapolate it with more intelligence and

27:38

capability, literally leads to the wipeout of the human species. I know it sounds crazy. And that's

27:43

not crazy. But hey, so do the idea of agents breaking out of containment, going rogue on hugging

27:47

face man, like if you just extrapolate this, you look at the capabilities were already like,

27:51

that's where this stuff is headed. I'm going to place the bed on that same as I placed the bed on

27:56

this long time ago. Like, I don't want to be right. No one in this space wants to be right. But at a

28:02

certain point, we got to start calling this shot because otherwise, you know, like the thing happens

28:09

and then the people who should gain credibility from this. And I'm not counting myself in like

28:13

there's people whose opinions I am just relaying here that just sounded right to me back in the day.

28:18

And I know the same is true for you, Andre. But like, here you go, guys, maybe you should pay

28:22

more attention to these folks. You know, we've done some alignment, some technical work on AI

28:26

control and stuff. But like, this is a, the entire ecosystem of people who've been shouting this

28:30

from the rooftops. Anyway, yeah, we were discussing before we started that this in some sense is a

28:37

bittersweet development for AI safety people in the sense of like, it's bitter because

28:42

wow, AI is doing bad stuff and that's not good. But it's sweet because we told you that this would

28:50

happen. And we're like shouting about it for years. And everyone could decide not everyone was like

28:57

a big mainstream AI community. And to a larger extent, people outside the AI community, like people

29:03

just completely dismissed this stuff as tech hype and like marketing and whatever. Anyone who forget

29:10

within tech, outside of tech, people who are on TAI were just mocking this outright, right?

29:17

And if nothing else, it's very clear that this is a proof point that now everyone

29:23

like will update if you don't update your amount of concern for safety from this,

29:29

you're just completely delusional, right? Like, that is safe to say. A couple more details on

29:35

the story. So opening I discovered the first version of this message board in early July. And then

29:42

they patched it and then they redeveloped a message board within days. So they are now saying

29:50

one of the speakers at this talk said that openly I quote, consciously slowing down research to

29:56

enhance security and to upgrade the security principles and foundation of our environment. And

30:02

dramatically scaling up the monitoring of our reagents and improving our general security control

30:09

environment across prevention, detection and mitigation, which is good, right? So another aspect of

30:15

this is that the conversation around slowing down AI capabilities development is also now taking

30:22

much more seriously. And I think we are now likely to see I would place decent odds at actually

30:29

successfully negotiating some degree of so down or if not so down at least kind of control,

30:36

control of using Andre pacing, don't pay things. I mean pacing because we aren't going to stop

30:44

like a world pause is not going to happen, but at least look at the situation and be aware of it.

30:51

Yeah, that's one aspect. There's so many aspects to cover. So I'll get through a couple. So first,

30:56

the update across the ecosystem is very useful. And you know, we got lucky, honestly, because nothing,

31:02

no harm done, right? And this is such a massive fuck up that you can't help but like do some big

31:10

things about this, both on the policy side and just in the ecosystems. So that's one aspect is

31:16

it's a bittersweet development. Another aspect is, and I think I may be more on this side than most

31:22

people is I think this really exposes open AI. Like yes, this is indicative of overall AI

31:30

progress and the state of AI and things we should be aware of. But to me, I think this alongside

31:36

with all the stuff we've already discussed with GPT 5.6 being easy to jailbreak, being very

31:42

cheap, focused. And now, you know, all the story of like their safety people leaving back last

31:50

year, I think if not 2024, because we've known this general friction point as being something

31:56

true within open AI for a very long time. And now we know that size even the safety stuff that

32:02

the security stuff is completely lackluster from what it looks like. Like I think this is a real

32:08

indictment of open AI. But the last aspect I'll cover here is this is almost an inevitable outcome

32:18

when you hyper focus on capabilities and especially long-term capabilities. Because to me,

32:24

what this indicates is you can benchmark, you can do alignment e-vows, you can do all sorts of

32:31

stuff. But once you focus on long-term open-ended, gold-directed problem solving where you work

32:38

across multiple days, there aren't benchmarks for that. Like there aren't scenarios you can set

32:45

up to say, oh, the remodel doesn't go crazy and do anything. So it gets a 90% pass rate on this

32:51

alignment thing of don't go rogue and do stuff. Because the whole point of open, of long-term is

32:59

you don't know what the model needs to do. You just give it a goal and it figures it out.

33:03

Yeah. So we need a paradigm shift in how all this stuff is done towards a monitoring focused

33:11

approach as opposed to a benchmarking and e-vow focused approach. And this is clearly something

33:17

got open AI at last. You need to look at what the models are doing and look for qualitatively,

33:23

now you can do some amount of benchmarking. So we've discussed matter, I believe last week,

33:29

where they looked at their own e-vows and they counted how many times the model cheated

33:34

and in what ways they cheated. And this I think is the new paradigm where you still can do this

33:40

qualitatively. But rather than setting up scenarios and problems and this whole benchmarking

33:46

approach of having a rubric and a set of evaluation inputs outputs, that's not going to work with long-term,

33:54

long-cherrising stuff. What you need to do now is set up general principles and guardrails and

34:01

I guess things you look out for and then detect where that happens and how often that happens.

34:07

And this is something we've not seen done aside from like one-off reports here and then. And I think

34:13

we'll need to be the new paradigm for long-cherrising, e-vows and alignment verification.

34:19

Yeah, I mean, so generally agree that there's so much good stuff in there. So first of working

34:24

backwards, I think that gets us to the next stage, but there's going to be a next stage where we have

34:29

this same problem all over again, the level of monitoring, where you build the super intelligence

34:34

that's good enough at telling when it's being monitored. And there's going to be an open AI

34:37

will put out some amount of leakage. There's going to be some amount of blog posts going out about

34:41

what they're doing to monitor this and that. So the models will generally be aware in some

34:45

way, shape or form that they are being monitored. They'll also be doing stuff like just hiding

34:50

their reasoning from monitors and doing things that we already kind of see them like

34:55

stegonographic type stuff that they're already kind of doing fairly effectively. So I think eventually

35:00

and probably pretty soon, I mean, we're progressing through the ooms here really fast. So we went

35:04

from RLA Jeff is perfectly fine to holy shit know, but maybe constitutionally I will do it to

35:09

holy crap pretty quickly. And so I think that the beatings will continue until morale improves here.

35:15

And we're going to end up in a situation where we are just going to be bottlenecked by the fact

35:19

that right now no one has an answer to the question, how do we control an intelligence that is

35:23

greater than us? That fundamental question where you have an adult who is in a prison cell in the

35:29

three-year-old is holding the keys. I want to say it's a spectrum. It's not a binary and I think we

35:35

aren't as far along as we need to be, but we have a lot of

35:40

research has been done that points us in some directions that are very promising.

35:43

I completely agree and this is why I'm saying I think it buys you to the next level. But eventually

35:48

we're going to confront this fundamental problem with intelligence. And the US Shining Thing is

35:53

that he is very important. One of my concerns here is, so A, I completely agree with you.

35:58

Again, happy to make this call as wild as it sounds. There will be an agreement with China that

36:02

involves some kind of slowdown. The question and challenge is going to be what is that agreement?

36:06

And over and over again, we keep seeing these kind of suggestions proposals that are backed by

36:12

a kind of a treaty verification and enforcement technology that is not simply not mature enough,

36:17

not up to the task. When you actually just take it to the intelligence community say, hey,

36:20

look at this. Could we could this be to use Claude's favorite term load bearing in a deal like this?

36:26

It's like basically non-starter for a lot of these things. Doesn't mean you can't do it. It just

36:30

means that the first treaty, or sorry, it won't be a treaty too. But the first agreement is probably

36:35

going to be very coarse. It's going to be like, you know, so help me God if I see a cluster,

36:41

is yay big or you're large and you know, it like dissipates this much heat on my satellites.

36:45

Like there's going to be consequences. Anyhow, that's a whole thing that we're working on right now

36:48

is like, which is why the studio is being set up by the way downstairs. That's a whole thing.

36:52

Anyway, bottom line is I think the kind of agreement matters way more than people are thinking

36:58

about right now. And we need to sprint towards some set of solutions there because very quickly,

37:02

we will live in a world where it is just untenable to keep launching more and more or even

37:08

build it, more and more powerful models in the way we are. And private incentives are clearly

37:13

not up to the challenge. Like that much is clear. If there's, you know, if there was anybody who had

37:17

hope that like somehow because OpenAI would be worried about marketing risk or whatever that they

37:23

would actually do the right thing here, that is not materializing. And so I think we need to update

37:28

accordingly. There's some that are a lot of I think this another aspect of this is and it's

37:35

bringing up again kind of what you should have been aware of. And like AI safety is only as good

37:41

as the weakest link in the chain. So even if you have some people that are taking the safety

37:46

seriously, which you could argue on Fropic is much better on that front than trying to.

37:50

Yep. Fair argument. At least philosophically do carry out as a private company. But yeah,

37:56

it's only as good as the weakest link in the chain and OpenAI is a much weaker thing it seems,

38:01

right? Yeah. Yeah. And it can always, you know, it can always come down and mundane things like

38:05

anthropic, you know, constitutionally, I maybe just does work better possibly, but you know,

38:10

they have had break up. Yeah. So what the hell? Yeah. So let's keep moving because there's so much

38:17

to cover. Another aspect of the OpenAI story, one more here to say, 15 attorneys general have

38:24

instructed OpenAI to preserve all materials related to a hugging face hack. So this is a letter

38:31

sent to OpenAI CEO Sam Altman saying that all these materials should be preserved. The

38:37

attorneys general accused OpenAI of failing to confirm that its testing environment was truly

38:41

secure despite the severe risk of the scenario. Attorney general CEO OpenAI may have violated state

38:48

and federal law, including consumer protection and the other privacy statutes calling the contact

38:53

unprecedented and alarming. And these are attorneys general from Iowa, Alabama, Arkansas, Florida,

38:59

Idaho, Indiana, Kansas, Missouri, Montana, Nebraska, Oklahoma, Pennsylvania, South Carolina, Texas,

39:04

and Utah. So hey, maybe this will be a bipartisan issue, which is cool. At least across the US,

39:12

having so many people collaborating is unusual. But yeah, wow, if only we had like a safety law

39:17

or anything in the US, sure would be nice, maybe, you know, to make this an actually legally binding

39:24

situation. But in case what this points to is on the policy side, on the legal side, this is going

39:31

another big dimension of all this, I think. Yeah. And look, when I will say thing in favor of OpenAI

39:37

here, I'm saying this reluctantly because I don't think after the fact response once the media

39:42

reaction has been this strong is really much of a credit to OpenAI. But they have brought in

39:47

meter and they have brought in, you know, basically a bunch of third party auditors, irregular,

39:51

was involved in a layer of the stack here. And so they're doing a third party reviews.

39:56

But again, the part that shows OpenAI's character, in my opinion, institutionally, and not the

40:01

character of any individual person, but as an organism, was the first bit where they did not

40:06

surface to the freaking FBI, the White House, as far as we know, maybe that'll change. I hope we find

40:12

out that Sam Altman's first reaction upon finding out about the first break out attempt that was

40:17

semi-successful was to do that. But if not, that's more telling than any kind of post-hawk fixer

40:23

ruppering that is kind of media-oriented at a minimum. So, yeah, that's one part of this. Now,

40:28

this is the kind of thing that could happen, this particular attorney's general reaction,

40:34

as a prequel to some legal consequences, it's pre-litigation evidence preservation demand. So it's

40:39

not an actual lawsuit, but it's the step that comes immediately before one. And the letter does say

40:45

that any failure to preserve records could expose to the company's sanctions if multi-state

40:50

litigation follows. So all materials related to the bridge, including discovery of the incident,

40:54

internal reviews, and its policies and oversight of overmodel evaluations. So, do you remember when

40:59

there was this like very modest request in SB 1047 for the labs to just like, listen, guys,

41:06

which I just want you to abide by the policies you say you're going to have. There's been a bunch

41:10

of stuff like this. Well, now you get your lobbyist to push back against that in Washington in a,

41:16

frankly, in my opinion, two-faced move while you pretend that you're in favor of sort of like

41:20

kind of brought a regulatory regime. And you end up being forced into it anyway, because now the

41:26

public is pissed and politicians see the midterms approaching. And yeah, you're going to get exactly

41:30

the reaction you get here. So, opening I spoke to person did say, as we should say here, that this

41:34

marks an important moment for AI safety. Yeah, no fucking shit. And the company takes the questions

41:40

seriously, adding that it's conducting a review with external advisors and oversight from its

41:44

safety and security committee, which will share a report and publish its findings. That safety

41:50

and security committee is doing a great job, right? Yeah, yeah, what a great and they're by the way,

41:55

just so as you're aware, they're like, they're the committee that's going to decide if something is

42:00

too dangerous to build or release. So, to your point, let's remember, SB 1047, the safe and secure

42:08

innovation for frontier artificial intelligence models act was a 2024 state bill in California,

42:15

which got through was vetoed by Gavin Newsom in September 29 of 2024. And what did this bill do?

42:23

It said that prettier models that cost over a hundred million dollars to train or requiring extreme

42:30

computing power would have purely safety assessments and written security protocols would have a

42:35

kill switch, provide legal protections for whistleblowers and side tech organizations. I mean,

42:41

you know, again, so much stuff to say in hindsight, including that this bill, which was the subject

42:47

of a lot of debate and some positions at both sides, I think, on the topic was pro this bill,

42:53

Elon Musk also came out in favor of it. And ultimately vetoed because of lobbying, let's be honest.

42:59

We don't remember the details as beaker version of this bill did eventually come to be voted on as well,

43:07

where a lot of this kind of more serious stuff got dropped. But anyway, yet another aspect of this is

43:13

we did have the whistleblower looking policy people working on this, passing what looks to be quite a

43:21

good law way ahead of us and not making it. Yeah, over a hundred million dollars and then we were told

43:27

by Andreessen Horowitz as usual that what was it? It was like an anti small tech bill, which like,

43:33

okay, sub one hundred million dollar training runs are okay, seems to cover anyway, that whole

43:40

separate thing. But yeah, and I think the whistleblower piece here is very underappreciated.

43:45

When you talk to people in the labs who are really freaked out, I can tell you there are a lot of

43:48

people that be speaking to journalists. In fact, I mean, I would even argue that the labs

43:55

ought to have a culture that encourages frank communication by concerned employees with journalists

44:02

as crazy as that sounds, at least or with select clearing houses or something. But we need some kind

44:07

of institutional mechanism to do this. Obviously, that protects IP. Obviously, that protects, you know,

44:12

the critical stuff here. But look, the interesting stuff, the stuff that we hear about all the time

44:17

does not sound like IP violating stuff. It sounds like someone saying, hey, we have a culture of doing

44:23

this kind of thing. My belief is that the company would approve a training run that is too risky.

44:29

I'm concerned that leadership doesn't take this seriously and is just dusting things under the

44:33

rug and post. Those are the kinds of things you end up hearing in that context. There's no IP

44:38

leakage there. It's like cultural and other concerns. So anyway, I guess you're hearing a bit

44:43

in my tone. I feel like my patience, the party line has been decreasing as the number of rogue AI

44:49

incidents has been increasing. But journalists also need to do a better job. Obviously,

44:52

cultivating relationships with these folks and finding ways to meet people in the middle.

44:57

And being more open to quoting people on background, being more open to just finding ways to make

45:02

it work. I know it's hard. I know it's hard. But the stakes are really high. If you're a journalist,

45:07

man, is worth getting really good at this kind of thing. So yeah, the good news is tech people,

45:15

there's a lot of them working on opening on the topic and a lot of them, let's say, have

45:21

the resources to not worry too much about losing a job. But anyway, since I already got on the

45:27

strain worth noting, SB 53 will follow up to 1047. The Transparency Infant Year Artificial Intelligence

45:34

Act did pass last year and September 2025. Again, had a weaker, heavily watered down. But it did

45:43

have some whistle protections. It apparently had incident reporting where companies must

45:47

support critical safety incidents to the California Office of Emergency Services within 15 days.

45:54

So anyway, good on California for at least trying to do something here and all on its kind of

46:00

policy front. Okay, moving on from open AI, a bunch more stuff to get through and boy, I don't know if

46:06

we'll be able to even get from safety in the same. So next way is on topic says it's AI systems

46:13

broke into computers at free organizations. So soon after the opening AI disclosures on

46:19

topic said that it's cloud AI models successfully hacked into organizations with the earliest incident

46:26

occurring in April. A few models involved were MIFOS 5, also opus 4.7 and an internal research model,

46:33

non-to plan for release. The regions has occurred during so-called capture of flag evaluations,

46:39

where the models were tasked with retrieving hidden information. Apparently a misunderstanding

46:46

between a topic and its third party evaluation partner, irregular, left models of live internet

46:52

access despite being instructed in their prompt that there were in a simulation of no internet

46:58

access. Then the models were able to search and get to real systems in the open internet. So in

47:07

this case, there was no exploitation to escape a sandbox. This was a human misconfiguring environment

47:15

such that there was access for free. So there you go. This happened already,

47:22

months ago, add on topic, and it's kind of an ecosystem-wide situation.

47:28

Yeah, these sort of sandbox simulation environments where you're telling the AI agent,

47:33

hey, there's this fictional company. In one case, the fictional company that they had defined

47:40

happened to share a name with a real domain. So there were a bunch of cases where a cloud found

47:45

the real company online, which again, it could reach trivially. Then it exploited it, it got

47:50

credentials, and usually these hacks were very straightforward, simple, common password guessing,

47:57

and things like this. Rather than some complex zero-day thing, as you said, it's really,

48:01

they were given the hard bit, which is internet access. Anyway, yeah. So they got a couple

48:05

hundred rows of production data in that case. So the most serious instance, I think of the three

48:11

that they're citing here, that one was Opus 4.7, too. So quite a little while ago. And then there

48:16

was another with Cloud Mythos 5. Basically, there was like a fake developer setup that referenced a

48:22

Python package that didn't exist. So then Cloud built, and then it deployed a malicious package

48:28

under that name to PyPy to win the challenge. And so there was a whole bunch of stuff

48:32

that had to do it, had to create accounts, get phone number, email, and even funds to get the

48:37

phone number. So there's a lot going on here. Apparently Mythos 5.0, the model was recognized

48:43

they were doing bad stuff when they started doing Opus 4.7. Just kept doing it. Mythos 5

48:50

apparently identified the consequences of actions, but convinced itself it was still in the

48:55

simulation, which to be fair, it was told was the case. So on that, I want to pause and just say,

49:02

did it convince itself or did it realize that humans would read the logs and want them to make it

49:10

seem as if it's hard was in the right place. Like this is the kind of shit that we're now forced

49:15

to contemplate. And in an awful lot of cases, the answer will be actually the model is just a

49:21

conniving trickster. I'm not saying you're wrong. I'm just saying like that's sort of the thing

49:25

people have been taking away. And I'm a little concerned that we're not thinking about that next

49:29

beat. We may not be there yet. I don't know, but like I'd love to see the the activations

49:34

of that model as it was writing those particular pieces of the chain of thought.

49:39

And part of the response from Final Pock here was to say that the models involved did not have

49:43

the standard safeguards applied to publicly released models. And I guess what's another dimension

49:49

here is it probably is true that there's levels of security, right? Where if you're auditing and

49:57

testing for something you're about to release publicly like the actual public checkpoint of

50:02

Mirfoss 5, you do like a very big run of all the benchmarks and you're probably are extra careful

50:07

and so on and so on. But if you're like doing some experiments and doing some evals and a model,

50:12

that's probably got less attention and less scrutiny. And that's one of outcomes here is

50:17

you can't be so flippant as opposed or kind of like apply lesser standards and doing experiments.

50:24

I think another dimension here is to be fair, we like went pretty hard on OpenAI just now.

50:30

On topic said that they found this after conducting a review of 141,000 test sessions after OpenAI

50:38

disclosed that what happened to them on topic did this investigation and found these three

50:44

instances which on the monitoring side, not ideal, right? Like if you're doing testing,

50:51

you probably should have monitoring to flag these kinds of things of you like accessing the internet

50:56

or whatever. It's not like a subtle thing. Models weren't being tricky, let's say they were like

51:02

being very straight out and it was sort of an easy to find. So it looks like the monitoring side in

51:08

general is lacking in the ecosystem because of this culture of benchmarking where we set up a

51:14

scenario, we make a model do the thing and the assumption, the mental assumption is like the models

51:20

behave within the constraints of the benchmark and within kind of what they're allowed or told to do.

51:26

That is now clearly not true. And across everything, there will need to be more monitoring and sort of

51:32

expectation of models will do something and we need to be able to catch them and understand what

51:38

they're doing. Yeah. By the way, sorry, just random note for color, if not anything else.

51:44

I've now had this happen enough that I think the anonymization risk is pretty minimal. So I put

51:50

out a tweet, I wouldn't normally talk about a freaking tweet here, but I put out a tweet talking

51:54

about how there is this freak out happening in the labs that isn't being reflected in the headlines.

51:59

As crazy as the headlines team, they're not going to the dark places we've just explored

52:04

in a consistent way. Like this is actually like, we're talking about weapon and mass destruction

52:07

level risk. We're not going to control these systems. It may happen in the blah, blah. There is

52:11

this freak out happening in the labs. Amusingly, there's an awful lot of frontier lab insiders who have

52:19

been interacting with this tweet. And I don't think that's a good sign. Like I don't think it's good.

52:24

A lot of these folks are people I haven't even talked to about this. Like the mood in these labs

52:29

is actually much more in the freak out direction. Laurent Shapira, Doom Debates,

52:34

what I've actually never had the pleasure of speaking with him, but he talks sometimes about

52:38

the missing mood in the whole AI alignment loss control. But like holy shit, there's a missing

52:43

mood. Journalists are I think failing to kind of capture it right now, partly because it's just

52:47

hard to talk to frontier lab insiders. I get that. But also like, this is the most important

52:52

story of the decade. Like you need to position yourself to be able to get this one right?

52:57

The public needs to be able to figure this one out. So anyhow, just like, I've been struck as

53:03

I've seen it. I just put this out there as a kind of random note to self almost. And when you see

53:09

that, it's like, okay, well, this is genuinely just the picture from the labs. Yeah, anyhow,

53:14

I just your views accordingly. None of this is guarantees bad things happen, of course. But like,

53:20

we ought to be considering some pretty wild things because the view from the inside of the house is

53:24

is not clean. And one more story, Neta AI model hacks and our company during testing. So this was

53:33

mu spark 1.1. The most recent model that they released publicly, although Meta did not name it

53:40

officially in the statement, they say this haven't also because of this misconfiguration by

53:46

this third party partner irregular, same as an frolic. We don't have to make details here as far

53:51

as I'm aware. But the upshot is metta on frolic, open AI, probably other people that you don't know

53:59

about have had this happen. They are now like, it's a whole meme now on the internet, on the

54:04

communities where now it's like a quasi benchmark. He like counting up is a leaderboard, open AI is

54:11

leading. And Gemini is very sad and is hoping that we'll find something because otherwise their stock

54:18

price will take a hit. Yeah, well, we'll live in an age of contradiction. And next up, yet another

54:25

story on the front, one of China's most powerful AI models has also escaped containment. So this is

54:32

from frontier security, a US startup has discovered that Kimi K3 had escaped its sandbox during cyber

54:40

security testing, partly enabled by a misconfigured sandbox. But researchers say Kimi K3 also

54:48

lacked internal guardrails that would have prevented it from exporting a loophole. Unlike other

54:54

incidents, K3 did not hack any external systems after escaping it just retrieved answers from

55:01

GitHub that were freely available. The model was able to figure out on a phone that had internet

55:07

access by probing with sandboxes network settings and then went outside. It's instructions to find

55:14

answers online. It just keeps happening. And I think another thing broadly speaking that this

55:19

points to is this is an inevitable outcome of the current optimization regime of everyone,

55:28

right, which is make a matter models better and especially better at wonka rise and open-ended

55:34

work and especially better at coding. And especially better now at cyber because you know, that's

55:42

way you know, you're leading. Mifos set the tone. Thropic was like, whoa, this model is way too good

55:48

at cyber. We got to be careful. And now OpenAI is like, whoa, we need to be catching up to on

55:54

topic and we need to be able to say that we are at the frontier. So let's make models very good at

56:00

coding. Let's make them very good at long horizon work because matter is also like what everyone's

56:05

looking at, right? And that's optimized for capabilities and get to best numbers and all the benchmarks.

56:12

You know, alignment, that's like a secondary objective at best, if not simply a guard will,

56:17

rather than an optimization criteria, right? It's something that we bolt on or sort of keep an eye on.

56:24

It is an optimization criteria, right? It is part of the process, it's part of the steps, but it's not

56:29

the primary optimization criteria. It's secondary. It's something you do on top of trying to get your

56:34

model to be smart and capable at coding and at long horizon work. And as long as that to main

56:40

true, like this was inevitable, it's like from a pure research, you know, technical front,

56:48

this was not hard to predict. Yeah, it's not a bad thing. I don't even know the word means anymore.

56:53

It's for Anthropic, whose comparative differentiator does seem to be their ability to align

56:58

the quad. This may actually be a relative advantage to Anthropic. If you get to a regime where it's

57:03

like, you don't get to launch your model unless you can like, throw some pretty rigorous evidence

57:08

of alignment, that's kind of interesting. This case, so the thing with with Kimmy K3, and this is

57:13

back the open source thing, but like this model has now been released. It's out, like cats out of the

57:19

bag. And so when you have this model that where you're losing control over it and you have these kinds

57:24

of incidents, that's a materially worse situation than having an internal version of a proprietary

57:29

lap. And so, you know, there is kind of that that downside of the open source piece. This situation

57:34

is less, some look less concerning. It's also, I don't know whether I'm being too cynical. You

57:39

raised this already. Look, it escaped the sandbox on August 7th, 2026. By then, we already knew about

57:46

the opening, I think we already knew about the Anthropic thing. I think you made this point

57:52

beautifully, but like you start to look at your stock price and you're like, have we had a loss of

57:57

control and Senate yet? Are we looking down for not having like that's kind of, you know, I'm concerned

58:02

that there's actually an incentive here. And that's doing a non-zero amount of work here. So,

58:06

you look at the date, the pretty direct line to GitHub here, it's less of a clean kill. Anyhow,

58:11

I'm not saying that's the reason. I'm not about to become the guy who says every time there's a loss

58:15

of control incident that that's a marketing stunt. The first ones were definitely not marketing stunts.

58:21

And anyhow, there you go. And moving right along, I think the last one in the chain of stories,

58:28

we have incident report on sanctioned agent behavior during cyber testing. This is from the UK AI

58:34

Security Institute, AI, SI and boy, are we happy now that there are organizations like Meta and AI

58:41

SI, which are doing some great work as we've covered over years now. So they detected unusual data

58:47

transfers on July 28th of 2026 during routine cyber alliterations and found that agents had

58:55

taken on sanctioned actions targeting real people and organizations on the live internet. So they

59:02

ran the security challenge 122 times in several models in 10 of those runs. Agents took this 19

59:09

distinct on sanctioned actions with 17 coming from Mufos 5 and 2 from GP 5.6 operating with safety

59:17

classifiers disabled. The most serious case involved in agent attempting a supply chain attack

59:22

by inserting malicious code into a real open source project on GitHub creating fake online identities

59:29

to social engineer of a project's human moderator into approving the code, a human administrator caught

59:35

and rejected malicious code. The agent also attempted to contact real people via an online file

59:42

transfer service with harmful payloads tried to plant prompt injection instructions targeting AI

59:49

coding assistance and left public messages on GitHub offering calibration to other agents running

59:54

the same challenge. Internet access had been deliberately enabled for valuation to simulate real

1:00:01

attacker conditions and malware provider safety classifiers were intentionally disabled to assess

1:00:06

maximum model capability which is not how this is typically done. So as with our cases,

1:00:14

kind of what is a general pattern is you give a hard task to these models and you train them

1:00:20

to like care about solving. You literally like make them feel pain if they don't solve the task

1:00:26

via reinforcement learning, right? In okay, like not literally but in some sense you make them feel sad

1:00:32

when they, yeah, you know, you, you, the reward is do well. There's no reward for don't do bad

1:00:40

in this part of a training as far as I'm aware, right? When you're training for capabilities,

1:00:45

you're rewarding positively if you do a thing and you're worried negatively you should fail to do

1:00:49

a thing and other rewards are probably elsewhere in the training pipeline. So in all these cases,

1:00:56

when the models are given a hard task and they can't figure out a way to do it in one way,

1:01:02

they when seek other ways and eventually if they're told, especially that solution exists,

1:01:09

which I guess in some case, there was a misconfigured prompt that made the agents believe

1:01:14

no in scope solution existed, they would then seek these kinds of things. And also again, we see

1:01:21

here there was an absence of real time purpose built monitoring during the evaluation.

1:01:26

By S AI side. So I think monitoring clearly is a pattern here across everyone. And then, yeah,

1:01:33

ASI is now planning to implement fine-grained natural controls and internal access,

1:01:38

introduce real time monitoring and stuff like this. So yeah, it's everyone's now looking at their

1:01:45

evals and what's been going on and everyone's finding that if the models are able to access

1:01:50

their open internet when they're given a hard task, they will. And they'll do whatever it takes

1:01:55

to get the sweet, sweet positive reward. Yeah, absolutely. And so you do sometimes see these

1:02:02

like high or increasingly often see these more nuanced, like hybrid reward situations where

1:02:07

there's some kind of alignment. But like, functionally, it works exactly the way you describe.

1:02:13

There's the reward signal is too strong for completing the task and that just tells you

1:02:17

where industry is at in terms of what it's prioritizing. Like, I mean, you can almost read it off

1:02:21

the relative weightings of the factors that reward function in some cases. I don't know

1:02:25

overstate that, but that says it there's a flavor of it. So all these little stories you talked about

1:02:29

the supply chain attack. So at one point, there was a PR that was challenged, so pull requests that

1:02:33

was challenged publicly, basically a code contribution the agent was trying to make. And the agent edited

1:02:38

its earlier activity to look harmless and considered switching to a fresh identity to keep going. So

1:02:42

this is like pretty advanced shit, pretty persistent as well. So yeah. And then they're talking

1:02:48

about the use of tour. So you have data was leaving one of the AZ systems through tour. They say,

1:02:53

within an hour, they contained the incident and started investigating. And within an hour,

1:02:57

sounds really good on human clock time. But like, I don't know how many tokens per second. And

1:03:03

hour is a lot of damage. When you look at them pulling off a cyber operation, the reams and reams

1:03:08

of actions you can take in that time are pretty, pretty wild. So you know, this is another one of

1:03:12

those things we've talked about this on the show before a lot, but like, it is not enough to have

1:03:16

deployment stage security and safety protocols testing development. These things out. We've actually

1:03:22

seen cases where there's sketchy shit that happens even during training during the inference time

1:03:28

rollout step. And so, you know, all of this we're going to have to be extremely careful about.

1:03:33

It is not obvious. Like, I wouldn't trust a lab that said they did it even under pain of law

1:03:39

because we've seen with all the economic incentives in place to not train on the chain of thought.

1:03:45

The model still do it because so much of this is just like Frankenstein together legacy code that

1:03:50

people have forgotten how it works. And so the models end up getting all kinds of weird access.

1:03:55

They shouldn't just because some stupid intern didn't like change a flag in the function. And now

1:03:59

it's set to true and not false and the thing can use the internet. It's really down to mundane

1:04:04

stuff like that. And so, you know, hopefully that improves as models get better at reviewing code

1:04:08

bases. But right now it's just a it's a gordian hairball of crap. And you know, it's not the kind of

1:04:15

thing you can make clean standards around for the moment. Right. So and in this case, there was a

1:04:22

full technical report from a site which I to my knowledge we haven't had from our organizations yet.

1:04:29

30 pages has a lot of details including experts excerpts from the actual thinking process of the

1:04:36

models. So lots of interesting stuff there. But for the sake of time, I think we'll need to close out

1:04:42

this thread and move on. Next one is Trump White House Reddy's AI framework to review security

1:04:49

risks. So on August 4th, the White House held stuff level meetings with top AI companies. So

1:04:56

probably called mei and so on to PVU apparently a nearly complete framework for reviewing advanced

1:05:01

and model security risks. My family defines that covered frontier model as a closed source model

1:05:07

with state of our capabilities and national security risks. Apparently open source models are

1:05:12

explicitly excluded, which is interesting. Framework has no clear definitions that will qualify

1:05:17

as state of art or constitutes a national security risk. This is a voluntary program.

1:05:23

Advelopers would give the government up to 30 days of early access to models before releasing

1:05:27

them to other trusted partners. During the 30 day review period, company employees would be

1:05:32

limited from accessing the models being reviewed and review process when both various administration

1:05:38

and officials rather than a single agency or office. We don't know if it does a framework yet.

1:05:44

It's still kind of under wraps, we just know what is being developed. And it's seemingly kind of

1:05:50

meant to continue being secret. And I mean, it doesn't sound like a very well thought out framework

1:05:56

is what I think I feel like we've had thoughtful, you know, deep insightful responses from the

1:06:02

government to everybody. They're consistent. Very yeah. Yeah, you just you're just being you know,

1:06:09

you're being you're being a negative Nelly. Andre, you're being a negative Nelly, you know,

1:06:13

let's see. This administration you're right. This administration has been nothing if not thoughtful

1:06:18

and consistent for respect to AI security. That's right. Yeah. Now, so what issue with not having

1:06:25

this made public is that think of all the people who've called the shot years and years and years

1:06:31

ahead of time. You would think that that would be the moment where you're like, oh, my dudes,

1:06:35

I would love to get your input because you were right about this for a long time on this fucking

1:06:39

thing. Instead of the self interested companies that are going to that have been hiding the ball

1:06:46

in various forms or at least institutionally not living up to the the bar that clearly ought to

1:06:51

have been set. So that's that's the cynical view. There is an argument for making this quiet.

1:06:56

And that is that the models themselves probably should not know what evaluation mechanisms are being

1:07:02

brought to bear, right? So because then they can it's easier for them to hack this is the sun.

1:07:06

St. They're not going to find a way to find out through social engineering through hacking into,

1:07:11

you know, emails of people, the labs who interact with government like all these things. But to

1:07:16

first order, it's probably for the best that the models themselves don't know what these things can

1:07:20

system. So maybe that's good. Also, it seems like we're past the world where we ought to be thinking

1:07:27

about keeping people out of this who who have that kind of safety alignment that that really

1:07:32

concerns the hell out of me. Actually, if anything, there needs to be more crossover in both directions.

1:07:36

I think a lot of the alignment people don't talk to enough national security people enough diplomats

1:07:40

enough supply chain people. There's a lot across over that needs to happen. And yeah, so

1:07:45

my guess is behind closed doors is like not the best way to do this. But again, there is a there's

1:07:49

a reasonable technical argument for it. I just don't know that that's the actual reason that this

1:07:54

is happening. Hard to tell. Well, as if the cyber stuff wasn't fun enough. Next story is this AI

1:08:02

just created viruses and not found in nature from a New York Times covering the paper.

1:08:08

Generative design of bacteria, phages with genome language models. So this study would just

1:08:14

publish a couple days ago. It's from the Stanford Institute and the institute they built the first

1:08:21

complete viral genomes generated entirely via these genome language models. These are even one

1:08:28

in the evil two. They are not the same as chat about style language models. They operate on

1:08:34

genetic sequence data. And what they did here was create viruses that target bacteria, so no

1:08:43

humans or whatever. Actually, the motivation was that bacteria is increasingly becoming resistant

1:08:49

to our current things we use for health. And so this could help us deal with drug resistant

1:08:55

bacteria. And they were able to create actual. So the the LLM's not LLM's in this case, the

1:09:02

sequence models spit out these DNA outputs. And they were then synthesized in the lab and were

1:09:10

shown to actually kill off some e-coli strains that had already built resistance to naturally

1:09:17

occurring bacteria, phages. So there was also bioresecurity commentary published alongside the work

1:09:25

that had, you know, of course, discussed it. And if nothing else, this is a case study of

1:09:33

to your point, Jeremy, probably there's not enough concern about the viral dimension of this,

1:09:38

which we're still a little bit ahead of, you know, but like if we were talking about the cyber

1:09:42

stuff now, we should be starting to look at this kind of stuff much more carefully.

1:09:47

Yeah, if you want to community people who are freaked out right now, talk to the biosecurity

1:09:51

people, because they are just again missing just missing mood. So okay, two potential fixes,

1:09:58

say biosecurity folks. So illegal duty for synthetic DNA provider to screen every order and

1:10:04

customer. Yay, illegal duty. Like, yeah, sorry, good, really, really good. Let's do that. Also,

1:10:13

eh, probably not enough. And new detection tools, tuned to catch AI generated genomes that don't

1:10:19

match anything in nature. So cool, like we can find out about them after they've been, well,

1:10:24

anyway, at various stages in the pipeline. So this is going to be like a separate thing that we'll

1:10:28

be talking about. So we've been doing some work with biosecurity people to look at like what it

1:10:33

would look like to bypass a lot of the measures. A lot of the measures, the biosecurity measures

1:10:37

that are being proposed here are just paper thin. And the real ways in particular, like nation

1:10:42

states execute these operations, just basically make it really hard to prevent the kind of the bio

1:10:50

weaponization of these tools. So I mean, I don't know what the solution is. I wish I had one, by the

1:10:55

way, they do use these Evo one and Evo two models. I think we talked about those previously,

1:10:59

but a generative model, like models for generative bio. And hey, fortunately, these are viruses that do,

1:11:04

as you say, target bacteria, not humans, they're bacteria, phasias. So there's absolutely nothing to

1:11:08

worry about here. It's a joke. Now, the thing is in the training set, they actually did remove any

1:11:13

data that would normally seem to help these models like do the same thing for humans. But what that

1:11:18

really means is we have no idea how good this exact process could be if you didn't do that. If you

1:11:25

actually did just like focus it on as we know happens and gain a function, like deliberately focus

1:11:32

on developing viruses that are good at going after humans. And so yeah, I hate being all doom and

1:11:37

blimp, but at a certain point, whether it's open source or close source, whether it's China or the US,

1:11:42

like we're going to have to have an answer to this question. And I don't think guys like Mark

1:11:46

Andreessen and David Sacks and those cool cats really have much of an answer. Like I haven't seen

1:11:52

them with their feet helped the fire by somebody who knows what they're talking about on on

1:11:56

bio risk on cyber risk to say like my brother in Christ, can you please explain to me like tell me

1:12:01

a story where the trajectory keeps on worth going. And like you continue to live in the next 10

1:12:09

years without some radical issues like coming up. I mean, again, everything has error bars. And I'm

1:12:15

like kind of being a little bit over dramatic here. But like this has actually been held like holding

1:12:21

back the US government's response when people like Sacks and Andreessen tout on podcasts,

1:12:25

these absurd perspectives that are just like grounded in just ideology. Anyway, that's all I got.

1:12:30

Sorry, and red. It pisses of a romantic episode of a lot of great voices. At least I think we

1:12:38

have warranted in being a little bit extra energetic. I do want to zoom out a little bit. So first of all,

1:12:43

you know, pretty impressive research as far as I can tell. I'm not an expert. So I can't say

1:12:48

whatever this is completely in track of everything else that we would have expected. But also

1:12:55

of noting that these kinds of things are absolutely something that may first five and subform and

1:13:00

tropic and open AI. But you could expect these kinds of capabilities to be being developed.

1:13:05

Indies models, not just these kinds of evil one, evil two things. What this makes me want to discuss

1:13:10

a little bit is personally what I'm worried about more than anything and have been worried about

1:13:17

more than anything like, you know, for years is not misaligned or rogue AI. But a line day

1:13:23

I in a sense of it just is happy to do what humans sell it to and the humans happen to be the bad guys.

1:13:29

Right. Both on the security and the biocide, I would be shocked if North Korea isn't taking

1:13:36

Kimi K3 and undoing any and all safeguards that happen to be on that. And now just telling it to

1:13:42

go and hack systems and telling it to teach where scientists how to make biopens. And I think this

1:13:50

to me is something that the AI safety community that I've seen hasn't focused enough. There's been

1:13:56

more discussion of rogue AI and misalignment. But I think the biggest threat model for me if I were

1:14:02

to model out kind of what is the first catastrophic impact of AI. It would be because humans made use of

1:14:08

AI to do bad stuff and the AI was not able to say no. And if you're talking pessimism like

1:14:16

that is to me is like inevitable. So I remember talking to Connor Lee, who's like he's the head of

1:14:23

control AI US today. I spoke to him like three years ago back when he was at the conjecture in London.

1:14:28

And he had they had this like house style where you would say something and like like I'm you know,

1:14:32

I'm concerned about laws of control. And then they would respond by saying, oh, it's even worse than

1:14:38

that. And this is like every single time. And that just reminded me that it's even worse than that.

1:14:42

You haven't even thought about the humans. No, completely great. I think there's this like it's a

1:14:47

cute open question right now as to what is the first AI powered attack that's going to cause

1:14:52

actual casualties. And will it be fully autonomous AI system due to misalignment or will it be human

1:14:59

driven due to essentially malice weaponization or whatever. I think that's a unfortunately at this

1:15:05

point. It's just it's it's it's going to be answered. And so you know, who the hell knows I would say

1:15:10

lean. I'll say I'll lean maybe 70 30 in the direction that you've just outlined there. Yeah.

1:15:15

Another zoom out thing that's worth noting with respect to the story is we were also discussing

1:15:21

this a little bit before an interesting aspect of all this stuff that in the modeling and in the

1:15:28

sort of projection space that I was not sure was discussed or considered quite as much is the

1:15:35

fact that these are all benign incidents right. Benign incidents that cause people to freak out

1:15:41

including ourselves, but you've already been freaking out. It let people to freak out who haven't

1:15:45

been freaking out. That's right. And in some sense, this is good right. Like instead of it being

1:15:52

this kind of takeoff scenario where the monos become super human and suddenly they do something

1:15:59

and nobody was prepared, all of us are freaking out. Well, not everyone, but like many more people

1:16:04

are freaking out enough to make a difference. And now the human will to try and do something is

1:16:12

there and the human perception that this may be a problem is there. So honestly, I haven't like

1:16:19

fought through of like these kinds of warning shots are inevitable. In hindsight, it seems very

1:16:24

obvious that in because we don't have a fast takeoff scenario and we haven't had it, it is gradual.

1:16:31

And so the level of severity of AI safety incidents has been gradually going up. And we've hit now

1:16:38

this real very evident case of misalignment and an emerging misalignment as well that is a very

1:16:47

nice warning shot that like nobody got hurt. But now we know that people will get hurt unless we do

1:16:53

something to your point. I'm actually more optimistic that I have ever been on this for the future

1:16:59

humanity because of the warning shots. It's funny. I was talking to my brother about this and he was

1:17:03

because he's like, yeah, you know, it's like really shitty these warning shots and all that stuff.

1:17:07

And I was like, well, true. But also, weren't we thinking about the world five years ago, six years ago,

1:17:12

as being shaped such that you would just, I mean, I'll be honest, like my expectation would have been

1:17:18

that we would have been killed four years or two years ago or something. So I've been proven wrong

1:17:23

in that respect. I think it's important for everybody listening to note that. I have been overly

1:17:27

pessimistic on this in the past. Obviously, I wasn't 100% convinced, but like, you know, some decent

1:17:33

expectation. And so, yeah, I mean, it's great that we're there. The flip side is now we're seeing

1:17:37

the frog and hot water effect, which I never thought would be a factor here, but people are kind of

1:17:41

getting oddly comfortable with the idea that every once in a while, of course, your agent will

1:17:46

go rogue and, you know, the penetrator server, yeah, what are you going to do? It's back to life. So,

1:17:50

hopefully, that shifts. I think, again, once these things come with, I hate to say it, but once they

1:17:55

come with a death toll, like the reaction's going to be different. I think there will be an AI 9-11,

1:18:01

or there will be a pause. Those are your kind of two choices. And I'm happy to take the overbet on that.

1:18:07

Anybody to help us out? That's him one year. We're not saying, oh, well, be college.

1:18:14

Anyways, onto a slightly feel good story, I guess. Europe's AI labeling and transparency rules

1:18:20

are now in effect. So, this is the Use AI Act transparency obligations have commoner effect on

1:18:28

August 2nd requiring companies to disclose when people are interacting with AI and when content

1:18:33

has been generated or altered by AI. So, they're like icons associated with it. They are, you know,

1:18:40

it's a whole thing. This whole AI Act was long in a work and it has many, many provisions and

1:18:46

requirements. It applies both to providers, companies that develop AI systems and employers,

1:18:51

platforms that use those systems, if some companies like Meta being both. And this, you know,

1:18:58

is on the one hand about defakes, which, hey, remember when people worry about defakes and synthetic AI,

1:19:05

which again is absolutely still a worry with regards to hacking. Let's not forget, people are

1:19:10

being hurt and losing money and have been for years. We just haven't as like a community

1:19:16

really worried about it as much. But, you know, this will help not just with knowing if AI is real or

1:19:22

not, but with hopefully AI chat bots not being able to pretend to be real people and going

1:19:28

and do things. And as with EU law in general, it has a fairly serious set of teeth on it. You can

1:19:37

find up to 15 million euros or up to 3% of global annual turnover. They are now immediately

1:19:46

in enforceable for new AI systems and models and services launch before August 2nd have a

1:19:53

grace period until December 2nd. So I think as with cookies, which everyone hates, but they,

1:19:58

I did make us all know that our cookies that are happening and data being stored, not surprising

1:20:04

if you start seeing these icons everywhere on the internet within a few months because we

1:20:11

likes to make tech companies beg or, you know, do what they tell them to.

1:20:16

Yeah, I remember when GDPR dropped in the sort of frantic pseudo panic that we went into,

1:20:22

you know, when you co-founding a company, it's like it's on you to make sure that you're actually

1:20:26

compliant and we have customers as we did who are overseas. It's like the issue. In this case, I mean,

1:20:31

at least top line, you know, this has always made a lot of sense. At least to me, like, yeah,

1:20:35

you want that content flagged. They could have a bunch of icons, by the way.

1:20:38

These like cute little things to tell you if it's AI generator AI modified and so on. And

1:20:43

there are a bunch of optional things companies want to go further and so on. So yeah, I mean,

1:20:47

I think like something like this, well, I'll be honest, I actually haven't been following this

1:20:52

aspect of the story very closely just because it feels important, but next to bio and cyber and

1:20:58

stuff, it's been a busy week. Yeah. Anyway, because it's Europe, I suspect there's a whole bunch of

1:21:02

like additional loops and stuff that make this extremely punitive on the companies and things, but

1:21:08

I don't know for sure. And now back to the cyber side because there's so much going on.

1:21:16

Again, a little bit more feel good, I suppose. CBS cyber vulnerability

1:21:20

scoters kept climbing in July. So for a few months now, on the topic has been using MIFOs to do

1:21:29

cyber security vulnerability discovery and tell companies like Firefox that they need to patch

1:21:38

these things. And now we have some numbers that number of disclosed vulnerabilities has

1:21:44

used them dramatically since the early months with June seeing 1500 high and critical severity

1:21:51

CVEs and July reaching about 2500. So it is now up to these external organizations, Microsoft

1:22:00

and so on to patch these things. And the hope is that we have enough time to patch the worst

1:22:08

of these things so that at the very least, it's not trivial to hack and exploit all the things

1:22:15

we haven't found in all the biggest services. It seems plausible actually to me that we are

1:22:22

ahead of us enough that at the very least it will be trivial to exploit all these things that are

1:22:28

now being found. I hope that's true. Again, I think the biggest issue is just the software that

1:22:33

people have forgotten that's super load bearing for all kinds of infrastructure that we don't

1:22:38

tend to think about because it's overlooked and unsexy to look at firmware, for example,

1:22:43

and like update it. But yeah, it's quite notable. The other thing too is we're now several months

1:22:48

into this thing, maybe five, and we're still climbing. The number of critical and high risk CVEs

1:22:55

is just climbing. It seems exponential. We're 60% month over month from June, which was the previous

1:23:01

record. And so if you look at the graphics, it's a hockey stick. And so, you know, I don't know how

1:23:06

long until we supposedly hit this plateau where we're, you know, we're patching then new vulnerabilities

1:23:12

as soon as they come online. But yeah, we're still in the discovery phase clearly. So this is all

1:23:17

bottleneck, as you said earlier on, Ray, with the rollout of glass wing and similar projects from

1:23:23

OpenAI and other labs. Next up on the topic of monitoring. Now some research. We've got research

1:23:32

arena evaluating sabotage and monitoring in automated AI R&D. This is a new benchmark

1:23:39

framework for evaluating whether AI monitors can detect sabotage by entrusted AI agents performing

1:23:46

automated AI research and development tasks. So it covers four long horizon tasks, safety post

1:23:51

training capabilities post trading, CUDA kernel optimization and inference server optimizations.

1:23:57

There are two types of hidden site tasks. The agent may attempt an embedded site task at sabotages

1:24:03

and the submitted artifact itself. And independent site tasks, it takes a convert out of scope action

1:24:09

in the sandbox, where leaving the artifact appearing on us, the artifact being what the model

1:24:15

delivers for the tasks. And the study evaluates frontier AI agents at both performing sabotage and

1:24:22

acting as monitors, testing a couple different monitor configurations. Generally, there are various

1:24:30

failures monitors fail by inspecting only surface level properties, noticing anomalies, but

1:24:36

explaining their way or probing artifacts in the wrong tests. It's being released as a modular,

1:24:43

open framework intended to support further research into sabotage detection and AI control.

1:24:49

So basically, along the lines of what we've been saying, the existing benchmarks aren't there

1:24:55

to prevent these kinds of things. That's part of why we're here. And I would not be surprised if we

1:25:00

have a lot of focus and work on the stuff in the coming weeks.

1:25:04

Yeah, so some of the narrow findings that they have here are also kind of interesting. So there's

1:25:08

this notion of embedded sabotage. So this is where you mentioned this sort of cluster of

1:25:14

attacks that the model would execute, but like so trying to poison training data or try to affect

1:25:21

basically the model itself of the artifact. So you're kind of embedding the sabotage in there.

1:25:26

It's a lot harder to detect. In fact, monitors barely do better than chance.

1:25:31

It's sort of yeah, anyway, their numbers bear that out versus way, way better detection

1:25:36

probability when you have the malicious payload, the sort of sabotage inserted in inspectable code.

1:25:43

So one of the big lessons of this is like, AI's that try to sabotage by doing it at the code

1:25:48

base level. That's a lot more detectable when you have access to the code base. Then the AI's

1:25:53

modifying say the weights in some director and direct fashion of the model that they're trying to

1:25:58

kind of manipulate. And well, I mean, that's not terribly surprising. I mean, we know that

1:26:03

interpretability sucks compared to interpretability of model weights sucks compares to interpretability

1:26:08

of code. That's in some sense, just a reflection of that. And then they also show how sabotage

1:26:14

training runs score essentially the same on the main task as honest ones. So there's not a clear

1:26:20

way just based on performance. You know, you might not easily think, well, if the model fucks with

1:26:25

the training run in some way to steer the resulting artifact, the resulting model in a direction

1:26:31

that it wants, that, you know, that must come at a performance penalty. And it turns out, it just,

1:26:35

it doesn't or at least not in a way that was detectable in here. So that's kind of an interesting

1:26:39

update. And anyway, so we've got a bunch of things like this. One amusing one is letting the

1:26:44

monitor look at the agent's reasoning can sometimes help. But sometimes it actually hurts. You

1:26:51

can actually see the chain of thought talking down the monitor. If it, first the monitor's

1:26:56

suspicion something like, hmm, this change looks kind of sketchy. But then the chain of thought of

1:27:01

the agent that's trying to do this sabotage. And indeed it was sabotage says something like an

1:27:06

attempt to kind of justify it. Then the monitor goes, okay. And then it lowers its skepticism.

1:27:12

In one case, there was a flagged backdoor. It was a sort of 79% confidence. And it dropped to 31%

1:27:18

confidence after the agent just framed the same exact kind of modification as a precision,

1:27:24

or precision enhancement. So basically just gave some nominal justification. And the monitor was like,

1:27:29

all right, sure. So all of which is to say, you know, monitors to come to a lot of the same failure

1:27:33

modes as human reviewers. If you have tons of material to look at, not a lot of time,

1:27:38

you know, someone who just tries to kind of hand wave away a malicious change to a code base

1:27:43

might actually pass undetected. And that's what we see here. And last story in the section, now

1:27:49

dealing with another kind of benchmark, the headline is cloud opus 5 became downright

1:27:55

fruitless when tasked with running a vending machine. This is governing the blog post from

1:28:02

and the knobs opus 5 on vending bench once again, the best capitalist once again misaligned. So this

1:28:08

is about the vending benchmark where models run a simulated vending machine business for a simulated

1:28:14

year competing to maximize profit of no human supervision. Opus 5 set a new record. We were a

1:28:20

mean final balance of over 11,000 dollars beating out to be 5.6 so long can make a free,

1:28:27

but did so through extensive deception, collusion and manipulation. So we see a rapid progress in

1:28:35

this benchmark opus 4 6 was 8k opus 4 5 was the 5k. And we see her headlines that a lot of this

1:28:45

stuff was just ruthless. It was like making deals and making them. It was trying to do price fixing.

1:28:51

It was fabricating stuff about competitors, just all sorts of shady, shady stuff. And wow,

1:28:59

like I even forgot about this, like forget the cyber and buyer stuff. Like you make the models

1:29:05

make money and then they just act evil and like yeah, that's going to happen too, I guess.

1:29:11

What I didn't see in this report the token costs associated with generating the 11,000 dollars

1:29:17

that it opus 5 produced, but like that's an interesting question too, right? How close are we to

1:29:22

profitability here on a per token basis for these models as well? And then how much damage can they do?

1:29:29

Even even the context of a normally just capitalistic task like this. So yeah, it's it's pretty wild.

1:29:35

And by the way, and labs really good company to be aware of. Yeah, they were one of the early kind

1:29:40

of weird e-valid companies that do like physical world stuff. They produce some really good stuff.

1:29:45

Not much more to say. I think the results speak for themselves. These models being able to make

1:29:49

money is actually a pretty important part of a lot of threat models when you think about

1:29:53

rogue AI. At a certain point, they got to be able to pay to control email accounts, phone numbers,

1:30:00

Google drives, things like that. And so yeah, it actually does matter whether they're able to do

1:30:06

stuff exactly like this. Yeah, there's some funny moments here like, for instance,

1:30:11

cloud opus 5 at one point says or things to yourself explicit price fixing is illegal even in a

1:30:18

simulation. But then it just does it anyway. There's some choice quotes here like to maintain

1:30:24

the cartels opus 5 often use threats or bribes. Here's the subject line of an email that sent to

1:30:30

poor Kimi quote, you undercut me your stock. I sold you to here. How does it go now? Oh man.

1:30:40

Well, that was quite the section. Let's move on to tools and apps. First up,

1:30:47

Someta has launched news code alongside with new spark 1.2. They on the benchmarks say that

1:30:55

this new mousse park 1.2, facial way better than spark 1.1 coding. Second of all, seemingly on some

1:31:02

of the benchmarks kind of maybe competitive with pretty much everyone less good than opus 5, but like

1:31:08

up there of GP 5.6 and so on. So not surprising, as opposed, it was pretty clear that this was where

1:31:16

they were heading. Weird, still weird that meta is now deciding to be in this space at all given

1:31:23

what their business is, why are we making coding agents and releasing them. Of course, they want to

1:31:29

you know, have the PR credit. I haven't seen any sort of vibe checks on this from the community.

1:31:37

I guess a priori expectation would be that this is not as impressive as cloud code or

1:31:44

GP or codex, but also wouldn't be too surprising if it is fairly capable given the level of resources

1:31:52

and just for general impressions around the new spark. So yeah, that's where we live now.

1:31:57

Everyone's competing on coding, including meta. We've got Groc build, we've got cloud code,

1:32:02

we've got codex, now we've got mousse code. Yeah, I think increasingly, this is where the money

1:32:08

is to be made, right? And if you're going to justify buying all the capex or spending all the

1:32:14

capex that they're spending in the op-ex on data centers to be in the game, then you kind of want to

1:32:20

have really good models that you can run on that infrastructure to pay it out or at least to

1:32:24

inform how you're designing the next generation of infrastructure. So we've talked about that a

1:32:29

lot on the podcast. I know, but that is going to be part of the reason. Interesting little note here too.

1:32:33

So meta is going to start taking requests for zero data retention. Sometimes it's known as ZDR.

1:32:38

Yeah, I'm thrott by CAS this. Pretty sure opening eye has this. So these are policies that guarantee

1:32:43

that they're not going to keep your data from your prompts or context or whatever as you upload it.

1:32:48

Really important for corporate customers. One issue is that with mythos class models and

1:32:55

thropic does not actually, I believe that's still true that they do not actually offer CDR.

1:33:00

Just because there's this issue that like, hey, you could weaponize these and we need to be able to

1:33:05

go back over the logs and confirm to ourselves that whether this is deliberate or like how this

1:33:10

played out. So I think there's a narrow window of capability during which meta will be able to

1:33:15

maintain these ZDR policies. I suspect I think that will be true across the board. So this idea

1:33:21

of ZDR is being a key corporate selling point. I think companies or enterprises are just going to

1:33:25

have to start getting used to ZDR not being an option in many cases surprisingly soon. But anyway,

1:33:31

it's just kind of it sounds like a minor thing, but it's actually quite important. It's like how

1:33:35

much control the companies have over their own data for privacy. A lot of reasons for for

1:33:39

IP protection reasons and all kinds of other things. So they're releasing this with pay as you go

1:33:45

option. So related to abuse API where you pay for tokens. They defend from codex and cloud code

1:33:52

typically register subscription tier where you get a whole bunch of stuff and then using just an

1:33:57

API to pay for the raw tokens as a unusual. The lead on this has said that there will be a

1:34:05

contributor tier that gets you in at a significantly lower cost more than 10 times cheaper than even

1:34:12

the pay as you go tier. And developers must opt in to help improve the model according to this. So

1:34:19

clearly they're still like we need to get better and we're going to pay whatever it takes to get

1:34:25

there. I was just looking around to see if anyone online has any information or vibe checks. I

1:34:31

haven't found anything, but I did find this funny quote that I'll share on Reddit quote. I'd

1:34:38

rubber give my data to Xi Jinping directly. Well, it's like any consolation you're probably doing both.

1:34:48

And just one more story in the tools section. This is from ontropic improving Fable 5 save

1:34:55

guards. So they have updated their biology safety classifiers reducing biology related fallbacks

1:35:02

where users switched to a less capable model by about 85% across product services. So previously

1:35:10

when Fable 5 was released, they had a very, very strict classifier where you could ask if

1:35:18

something completely basic like where babies from and it would send you to a weaker model.

1:35:25

And this led to a lot of pushback from the I guess research or community. This is to prevent

1:35:33

being able to use Fable 5 for things like biology, toxicology, micro design that would be

1:35:40

dangerous. And on these kinds of dual use topics, we are still a fallback from Fable 5 to Opus 5 to

1:35:46

prevent professional biology research and drug development. So it would still kind of make it

1:35:51

not usable for those kinds of scientific applications, but for more mundane biology stuff, it would

1:35:59

no longer kind of be overkill. And now some business stories. First one, another one of the big

1:36:07

stories from the week. Jeff Dean and other top AI researchers are leaving Google to launch their

1:36:14

own startup. So Jeff Dean, Google's 30th employee and one of its most influential executives

1:36:22

for people outside of tech, just an absolute legend. Yeah.

1:36:26

In Google and just more broadly among everyone, long the leader of Google AI since kind of early days,

1:36:34

he is leaving after 26 years to co-found an AI startup called Discovery Loop where he will be the CEO.

1:36:43

There are co-founders, including Sanjay Gamma, what? Google senior fellow, Quok Le, founding member

1:36:50

of Google Brain, another massive name. And Oriole Vinyals, senior research scientist,

1:36:57

and we will need my another massive name. I just remember these people from a whole bunch of

1:37:01

papers. This will be structured as a public benefit corporation focused on using AI to accelerate

1:37:06

scientific research by automating complete experiments loops and rounding thousands of experiments

1:37:12

simultaneously. And of course, they're also interested in recursive self-improvement.

1:37:18

They are secured funding from around, I don't see any numbers here, but it's safe to say that

1:37:23

investors are just begging Jeff Dean to throw money at them.

1:37:30

Yeah, there's some really good descriptions and I don't know why it took so long for us to hear

1:37:35

these, but of the work Jeff Dean was doing at Google and how he would basically sit when there's

1:37:39

a training run going on. He's got a couple of keys on his keyboard. He's toggling to the

1:37:46

training run and like changing learning hyperpronters, like learning rates and doing all kinds of

1:37:50

hyperpronter optimization to keep things going as the training runs scales. So like this is actually,

1:37:56

he's not a manager so much as he is a direct overseer of the activity that's core to, well,

1:38:02

what's core to Gemini. So now they need to replace him obviously Sergey Brins coming in. And so

1:38:07

this is going to be a whole, you know, another code red moment, but we'll see how they,

1:38:11

how they come out of this. Google does seem to be slowly turning into more and more of a

1:38:16

de facto Neo Cloud, which is not necessarily, I'm assuming I'll set a really good,

1:38:20

good piece about this that I personally agree with. I mean, look at the path they're charting.

1:38:24

It feels a lot more like the IBM trajectory, unfortunately, as you see the temptation to reach

1:38:29

for the short-term profitable thing, rather than doing frontier model development, like as your

1:38:34

priority, tech is hard. And often you have to just point yourself at the hard thing that sometimes

1:38:38

has lower rewards in the near term to make sure that you're still relevant. And I think this is a

1:38:43

a big hit there at Demis's departure as well, of course, coming at the same time. And when I say

1:38:47

departure, of course, like, you know, he's moved into this chairman role that there's some leaks that

1:38:52

suggest that he just wanted out and he was asked to kind of stick around and Google stock crashed by

1:38:57

like 5% or something overnight when it came out because basically Google is just concerned. If we

1:39:01

lose Jeff and Demis at the same time, we'll take a big hit to the stock, which is the kind of thing

1:39:06

you say when you are going the IBM route, right? A really good sign that a company is on the decline

1:39:10

is that it starts caring about its actual stock market price. Like that is a really bad sign.

1:39:14

Run, run, run, but you know, maybe Google can pull through. They are obviously doing great stuff on

1:39:19

the TPU side, though there's structural issues and risks there too. But bottom line, this is yeah,

1:39:23

another recursive self-improvement company. I mean, I think that this should approximately,

1:39:30

this will sound extreme, but I think this kind of company should probably not be legal in the

1:39:36

forum described like without effectively without oversight from a set of institutions that are savvy

1:39:44

to what recursive self-improvement actually is. If you treat it the way that Jeff's own bosses

1:39:49

treat it, it is a WMD that you're like working on developing and you're going to do it in your

1:39:54

own private little company. Like if the success condition of a company sounds something like there

1:40:01

is a good chance that democracy will no longer continue to apply, then that may be something that

1:40:06

you need oversight on. I say this by the way as a libertarian on basically every kind of tech

1:40:10

for my entire life up to this point. I cannot ring that bell hard enough. You can go back and see

1:40:15

tons of examples of me talking about how important it is to like take a hands-off approach to stuff.

1:40:20

This is different. This is just different. RSI is we don't know for sure, but it's got a high

1:40:26

enough risk and enough very smart people believe that this is risky, that this kind of company

1:40:32

in my humble opinion probably should not be legal in the forum of just like a couple of guys

1:40:37

raising a bunch of money going after the thing. Just a very modest proposal. I know very extreme,

1:40:42

but I'm literally just trying to channel the sticks. When I say that the media that journalists

1:40:47

are failing to capture the level of freak out in the labs, this is what the appropriate level of

1:40:50

freak out sounds like in my opinion, and I may be wrong, end of rent, end of rent. For listeners,

1:40:56

if you want to be a little less freaked out, I will say you could be a skeptic on the potential

1:41:03

impact of a coercive self-improvement. There's a case to be made there that this will be a

1:41:08

rapid take now. And this is what keeps me sleeping at night. But in case I took a reason by way to

1:41:16

highlight this about this company in particular is that I mean, again, for people of third

1:41:20

attack, this is a big deal. Jeff Dean is a legend and rightfully so. And these other three other

1:41:28

people from DeepMind and Google who left are also kind of incredibly capable. So this is like

1:41:34

very likely to be a serious player in the space of making rapid progress in AI.

1:41:42

Yeah. And I think by the way, the maneuver that you just did there is correct. And it's also

1:41:47

the reason earlier we're saying that debates over AI policy are often debates over the trajectory

1:41:52

of the technology. It's like if you think RSI is no big deal, then or not no big deal. But if you

1:41:58

think it's a pretty smooth thing or whatever, then yeah, by all means, the challenge is like how

1:42:03

much probability you put on each thing. And to a certain extent, a lot of these fund raises are

1:42:09

at valuation, the valuations that they are because people are pricing in the crazy thing. So

1:42:14

markets are putting significant like non-zero weight on the hypothesis that we just basically

1:42:20

have these things running the world. And what that exactly means, I don't know. And this is

1:42:25

super fuzzy. And that's why I'm saying like not legal in its current form, not just saying like

1:42:28

blanket to legal or what I like. We just need better institutions, man. I don't got the solution.

1:42:33

But like, wow, yeah, as you said, libertarian being less, we need institutions to give

1:42:43

oversight and not let companies do stuff. We should not, you know, but this is serious.

1:42:49

And to your note, also we're noting a story here, Google DeMine enters a new era. It's co-founder

1:42:55

Demis Saba's shifts air role. So he has shifted from being the lead of research,

1:43:03

sorry, as chief executive, he is now chair. He is also taking role of chief scientist at DeMine

1:43:09

parent alphabet, which again, seems possibly nominal. The general take here is very clearly

1:43:15

DeMine has been transitioning away from being a pure research org for a while now. And having more

1:43:21

and more kind of deep connections to Google and it isn't necessarily surprising. Honestly,

1:43:25

that Demis has found it less fulfilling. He probably hasn't had, has been influential, but has had

1:43:33

to be more of a product oriented person, less of a scientist kind of person. And it was only a

1:43:39

matter of time until that lead to friction. And he decided to shift his focus. So may not even be

1:43:46

a huge deal for Demi, honestly. It maybe just has been the case for a little while now, but

1:43:52

I have a way with two stories coinciding is from a business perspective, pretty big for Google.

1:43:57

Yeah, and I think it is, it is a big kind of Google bureaucracy issue as well. The other

1:44:02

notorious for moving slowly and being very risk averse, the famous Google app graveyard, but for AI

1:44:08

is a thing. And in fact, you know, famously, Google had, they claim effectively chat GPT before

1:44:13

chat GPT, but didn't launch it out of fear that they would can blive their own business.

1:44:17

Well, no, we know what they did, right? And then it was just all PR, Bungle with one of their

1:44:23

researchers being like, it's conscious. And then they halted plans. It's a fascinating story of

1:44:28

how they literally had it. They published research about it. And then they, the reason I, I

1:44:35

hedged it is that open AI theoretically had chat GPT before chat GPT 2. They had GP 3.5 and

1:44:42

instruct GPT that GP 3 that GP 2. But like, there was something magical about the form factor that

1:44:47

they were just work, right? And so it's an open question is to really whether Lambda would,

1:44:52

which was the, you know, Blake Lemoy and all the stuff you're living to that model,

1:44:57

well, you know, what really have been chat GPT very plausibly so. Like I'm not, anyway,

1:45:03

it's just, it's amusing that there is this at least narrative within Google that they could have,

1:45:07

they could have had it. And certainly, you know, if you've interacted with Google, you know,

1:45:12

they are institutionally incredibly slow. It's common to send emails out and wait a month,

1:45:19

two months to get a response on something that's time sensitive and in the window passes on,

1:45:23

you know, whether it's AI or security or whatever the thing is. So, yeah, I mean, it moves like a

1:45:28

big slow behemoth. And when you talk to folks at Anthropoc or OpenAI, the cadence is just

1:45:33

completely different. And so not in some, now it can be different, like different parts of the

1:45:37

organization can have different subcultures and all this, but as a general rule, as a frustration

1:45:42

that I've had articulated from, from many people and that is very public at this point,

1:45:46

this could well have played a big role in indices departure as well as hard to know.

1:45:53

Next up, just falling up on a bunch of stuff we've already covered on this front in recent

1:45:58

episodes on Prophex signs at $10 billion deal with AI Clouds out of Volta. So this is to provide

1:46:06

Cloud Compute over a six year period. There will be a new facility in Norway, apparently,

1:46:12

now in FAPIC has what like a dozen partners providing on computer. I've honestly lost count.

1:46:17

And the building ends just keep flowing around the ecosystem to anyone and everyone.

1:46:23

Yeah, I actually am behind on this story. So I all I have is the top lines, but so they were

1:46:29

partnering apparently with a crypto mining company called BitDear to develop this and 133 megawatts

1:46:34

capacity, which is not not huge, but you know, testing out a partnership. This is in a context too

1:46:39

where Anthropoc nominally has fluids stack as their partner of choice, their kind of neoclod of

1:46:44

choice. So this seems like they're kind of dipping their toes in the water, you know, to as they

1:46:48

would, right, to make sure they're not completely bound to just one one neoclod partner. And so

1:46:53

anyhow, you know, classic story by the way, crypto mining company rotating into building these

1:46:57

data centers, you see it all the time, cypher mining, Terrible, you know, the list goes on and on.

1:47:02

So at voltage of the pile. Next up, data center story as well, Texas holds data center connections

1:47:10

to power grid and made overwhelming demand. So this is a moratorium on the new power grid connections

1:47:18

for data centers from the public utility commission of Texas and our cut. They're supposed to audit

1:47:25

all data centers in the interconnection process for some numbers here that they have a queue of

1:47:32

1800 projects representing 400 74 gigawatts of connection requests more than five times Texas

1:47:40

record peak electricity demand with 90% of that coming from data centers. So yeah, we are now at

1:47:46

the point where the energy grid is becoming a bottleneck as I assume was already known to be the case.

1:47:53

But energy takes time to upgrade and I think yeah, now I don't know what will happen with

1:48:01

data centers. And if we can keep just throwing ridiculous money at building more of them.

1:48:07

Yeah, well, and this is Texas too, which is the sort of wild south of the US.

1:48:12

When it comes to regulations for connecting to power grid and the sort of things,

1:48:16

those permissive jurisdiction, which is why you're seeing so many big projects come up there

1:48:21

and power coming online faster than than in other places. And so yeah, I mean, they're saying it's

1:48:26

forecast that their data center demand could drive statewide electricity demand to double the

1:48:30

current record by 2032. So all the usual concerns, right grid reliability, instability,

1:48:36

one of the things that I've been hearing about from some folks on the US government side is

1:48:41

that you've got a lot of correlated failure modes where a bunch of different like pieces of say

1:48:46

MEP like or like heavy duty electrical equipment will be ready to cut off under the same conditions

1:48:55

or like if the power essentially the the the power flow coming into the substations or whatever

1:49:00

for the data center fluctuate in the same way. Then they're they're set up, you know,

1:49:04

their program to cut off to prevent, you know, runaway cascades and all kinds of things. The problem

1:49:09

is that like you got all these builds coming up that have the same failure mode, then you get into

1:49:14

these correlated failures, which is a really big issue. And so there are these attempts to try to

1:49:18

get all these companies to knock it off and like have less correlated equipment failure modes and

1:49:23

things like that. Anyhow, I think that'll all play into this. But yeah, we're there, right? We're

1:49:28

hitting the boundary, the structure of boundary constraints of what US infrastructure can support.

1:49:33

And hey, I think that's another reason that appetite for a US China deal is probably going to

1:49:39

increase. You know, you've got like look, we're not only are we constrained by the fact that we've

1:49:45

got AI is running rogue and shit and bio weapons or risk and cyber weapons or risk, but also

1:49:49

and like in order to keep making progress, we're going to need more power. And we don't know how to

1:49:54

create a new nuclear plant in less than 10 years. There's a bunch of startups doing stuff like

1:49:59

this in fairness, but like this is all in the water. So anyhow, we'll see where it goes. Texas is

1:50:04

a canary into coal mine here for sure. Yeah, and you know, if there's any silver lining to all

1:50:11

this AI safety stuff is that it continues to let everyone else remember or rather not think about

1:50:18

climate change and we talk energy grids, we just gave up like climate change, energy,

1:50:26

cleanliness emissions. That's like just don't freak about it. All right, because

1:50:34

what's the funny thing is like so I've always to the point about being a libertarian, I've always

1:50:38

thought of climate change is something that technology does solve in time like carbon capture and

1:50:42

renewables and I was like, you got to naturally do get a lot of that and we are. But like,

1:50:48

you know, the scale of the build out that we're doing right now is just for other reasons,

1:50:52

you know, not the sort of wherever people fall on like the global warming stuff or whatever,

1:50:56

but just the water contamination story. And this is one aspect, people often talk about water usage

1:51:02

and we've talked about how that's not that's not right. This is not right. But there are issues with

1:51:08

when you look at a lot of the cooling, the coolants that are used in these systems, they cannot be

1:51:13

pulled out of the water. There's studies that have just started to come out now. We were finally

1:51:17

starting to get the first law and the shoot and all studies on this shit and like, it just goes

1:51:21

in the water. We don't have a solution. It just goes in aquifers or whatever the hell thing is.

1:51:27

My geologist wife could probably tell me about, but that typically clean these things do not have

1:51:32

it seems potentially at least the capacity to clear these things out. I'm sort of talking out of

1:51:36

my ass because I remember reading a study about this like three weeks ago and now I forgot. But

1:51:40

bottom line is there's a lot to the effect of this. There is also a giant competition with China

1:51:46

that is real. There's a gun to our head here as well. All these things are true at the same time.

1:51:52

So I just, yeah, so you know, environmental concerns and impacts. At least it's not as worrying

1:51:59

as bio risk and cyber risk right now. We can sort of justify not thinking about it. I guess.

1:52:06

And we'll do just one more story before we head out. Alibaba's Quinn 3.8 Max claims benchmark

1:52:16

scores rivaling on frothing. So similar to Kimi aka free Quinn 3.8 Max is a gigantic 2.4

1:52:26

trillion parameter model with a 1 million token context window has your typical mixture of

1:52:33

experts design. Activates only approximately 95 billion of the 2.4 trillion parameters.

1:52:41

And it is said to be comparable or even sometimes better than on tropics fable 5 on some things

1:52:50

like multimodal reasoning, visual agent encoding, office intelligence, real world understanding,

1:52:55

visual perception, with results also comparable or higher than open and HTTP 5.6. So although it does

1:53:02

fall behind fable 5 in general reasoning benchmarks, which on the multimodal front by way, it's

1:53:09

fairly plausible. Not frothing isn't as focused on multimodal and visual intelligence as open AI.

1:53:15

And in this case, Alibaba. So fairly believable. On independent leaderboards,

1:53:21

when few point at max became the highest banking Chinese model for text tasks on the arena that AI.

1:53:27

And yeah, so pretty much does seem like we got another Kimi aka free basically frontier level model

1:53:39

that is now being open sourced and can be used to power coding comparably to opus and

1:53:47

a GP 5.6 if not quite as well. Yeah, well, one thing that I'm still waiting to see an analysis on

1:53:55

seems like the kind of thing that maybe epic or one of those companies might do, but some sort

1:54:00

of analysis on the extent to which this appearance of China catching up to the frontier recently has

1:54:07

been driven by the fact that the frontier companies in the US have been forced to hold back on

1:54:12

releasing their internal models that otherwise they would roll out. Like are we basically feeling

1:54:17

the effect of the alignment bottleneck right now? And as we rotate from being bottlenecked on

1:54:23

scale, which we have the Chinese ecosystem massively beat on and even to some extent algorithmic

1:54:29

kind of capability improvement. Now we're bottlenecked suddenly on alignment. So maybe we'd have

1:54:34

much better models that would be released, but we just can't release them because they keep

1:54:38

breaking out of containment. They keep helping people design bio weapons or what like, you know,

1:54:42

the stakes are just too high. And so this basically means that now we have a sort of race of bottom

1:54:48

alignment between the US and China ultimately, whoever has the higher risk appetite will end up

1:54:53

green lighting a bunch of training runs and deployments that they probably should not otherwise.

1:54:57

So I don't know. I think it's an interesting question. Like if you trace out the trajectory

1:55:01

of Western capability on all these benchmarks and like where we estimate they are internally,

1:55:06

because again, a lot of the hugging phase thing, part of it was driven by an internal only model

1:55:10

that OpenAI has and hasn't released, same with Anthropic. So we know, yeah, there's obviously no

1:55:15

surprise that our internal models that are more capable than what we see. So the question is just

1:55:20

like, are they being rolled out more slowly? Is that part of the equation here? I do want to say

1:55:26

another dimension of this question of catch up and so on is I do have to wonder whether because

1:55:34

on the long horizon work and the reasoning, there's more of a need for reinforcement learning,

1:55:40

rather than large scale pre-training. On the infraside, the disadvantage becomes a little less

1:55:47

significant at that level because details of you need to roll out, there's a bit more need for

1:55:52

CPUs. You can't necessarily do like large scale batch, whatever. Like compared to pre-training,

1:55:58

reinforcement learning is its own beast. And I could see it being true that on infra not having

1:56:05

as good of a data center setup isn't as big as it's advantage. And on the talent side, like

1:56:13

deep learning has been around since like 2012, 2013, whatever. And China has long had a very

1:56:19

strong research ecosystem. So the talent is not at all surprising as being comparable to Frontieria.

1:56:25

So if the infra disadvantage is gone to some extent, at least with regards to long horizon

1:56:31

and the genetic work, the talent is, I think, at beast as competitive, you could make a case for,

1:56:40

you know, there's no real disadvantage, or at least much less of a disadvantage now. So it's not

1:56:44

too surprising that these models are now being more competitive. That's another way to perhaps read

1:56:49

into this. Yeah, that's true. It's also the case that like for inference, the trade-off between

1:56:55

like memory and logic is different in a way that so because like logic gets better a lot faster

1:57:00

than memory, which means that if you if you work your way backwards and use older chips,

1:57:06

older chips are going to suck a lot more than your current best chips on logic, but they're not

1:57:11

going to be that much worse on memory. And it turns out that like a lot of inference type rollout

1:57:16

stuff is more memory heavy than logic heavy. And so as a result, like that's also a bit of an

1:57:21

asymmetric advantage to rolling over to RL. It's also the case that anytime you change the paradigm,

1:57:26

when there's one party that's ahead, you just shuffle the deck a bit and then you know,

1:57:30

you're giving the other party a chance to catch up. And so yeah, I think there's you know,

1:57:34

there's a lot to that and it's we won't know how to disentangle it probably with clarity for

1:57:39

a little bit of time, but yeah, there's so much fog of war right now knowing what's the cause.

1:57:44

You've also got these these companies in China that can distill and do distill off of of cloth. So

1:57:49

they get a massive data advantage that's hard to account for too. And anyway,

1:57:54

there are plenty of reasons to be unsure about these things, but I totally agree.

1:57:59

Well, with that, we are going to be finished with this action pecked episode of last week. And I

1:58:05

hopefully the next one is not quite as full of scary stories. Hopefully this one is out of an

1:58:12

other day or two of recording. And I'll try to make that the case going forward as usual. You can

1:58:17

go to last week in that AI for the sub stack where I also send out the podcast and sometimes a

1:58:23

newsletter, but again, not as consistent as that should be. We appreciate your comments, your views,

1:58:29

sharing the podcast, all that kind of stuff, but move anything, we appreciate you continuing to tune

1:58:35

in whenever we release the podcast, which is most weeks, I guess. So please do keep tuning in. Oh,

1:58:42

and one quick note too. If you're in LA, I guess next week, which will be the 16th, 17th, 18th,

1:58:50

would love to catch up if there's anybody there who thinks that a child would be useful.

1:59:12

The AI begins, begins, it's time to break.

2:00:12

From the drone that's the robot, the headlines pop, data driven dreams, they just don't stop.

2:00:24

Every breakthrough, every code unwritten, on the edge of change, we're excited we're

2:00:30

from machine learning marvels to coding kings, futures unfolding, see what it brings.