#253 - Opus 5, Gemini 3.6, Kimi K3, Hugging Face Hack

2026-08-03 05:00:00 • 1:43:21

-

Hello and welcome to the last week an AI podcast we can hear chat about what's going on

0:15

with AI as usual in the subsequent you will summarize and discuss some of last week's

0:20

most interesting AI news also the week before we have unfortunately skipped a week due

0:27

to scheduling conflicts but we will cover everything relevant from the period I am one of

0:33

your regular hosts Andre Khrankov I studied AI in grad school and now work at the startup

0:37

AstroKade and everybody what's up my name is Jeremy of course I'm your other ghost I'm

0:42

from Glatz and AI do AI national security super intelligence see type end of the world

0:47

stuff so I sound a little thick right now by the way which is related to the reason that

0:51

we didn't record that the so last week which is that I was traveling and I got sick on

0:56

the flight was guy who was coughing up along next to me and anyway that's why I sound

1:00

so weird right now but trip was really useful and yeah hopefully be able to talk about

1:05

a lot of this stuff soon but a lot of conversations with like researchers at the frontier labs

1:09

and folks on the you know the safety teams the capability teams all that kind of thing

1:14

that I think bears quite a bit on the events of last week and the week four will definitely

1:18

be taught about a lot of that stuff with some of the inside you a little bit what I can

1:21

share right now on those things but man things are moving but and it has been a slightly

1:27

eventful two weeks I mean I guess it's not the most eventful you've had this here but there's

1:31

been some big stuff it will be touching on as a quick preview there's a few new models nothing

1:37

gigantic but fairly meaningful you'll start with then as usual some funding stories and deals

1:44

about compute and so on some major open source releases including Kimi K3 we have the

1:51

discussed about so we'll be talking about that then policy safety of course we'll be talking

1:56

about the recent hacking incident from open AI and a whole bunch of other stuff related to that

2:01

it's gonna be a kind of policy safety heavy episode and then we'll round it out with some research

2:08

and synthetic media and art so it'll be a packed episode we'd like to thank notion for being

2:13

a sponsor agents are getting smarter every day but even the smartest agents get stuck without

2:19

the right context and the right tools that's where notion comes in with a recent launch of custom

2:24

agents notion became the collaborative AI workspace where teams and agents work side by side

2:30

and now their new developer platform is turning that workspace into infrastructure developers can

2:35

build on notions developer platform gives developers and coding agents the primitives to

2:40

extend what's possible a notion and pick it beyond connect to external systems bring context in

2:45

take permissioned actions across your tool stack and expose custom agents capabilities to any

2:51

system that needs them these primitives include a CLI workers on notion hosted sandboxes and

2:57

external agents API and then agent SDK trigger notion agents from any app learn more about notions

3:04

developer platform today at notion dot com slash LWAI that's all lowercase letters notion dot com slash

3:12

LWAI to try notions developer platform today and when you use our link you're supporting the show

3:19

notion dot com slash LWAI this next sponsor isn't regulatory I but I've personally used them for

3:25

years so I'm happy to have their support and it is factor they make chef crafted the dietitian

3:31

design ready to eat meals so you don't have to choose between real food and convenience both in

3:36

grad school and as a startup employee I don't have a ton of time so when I get home I'm tired

3:41

and being able to prepare really quite a good meal without any effort has been fantastic their

3:47

meals are ready in two minutes and require no prep and no cleanup so even on the days your schedule

3:52

is completely out of control eating well is still achievable they are over 175 band ingredients so

3:59

every factor meal is designed around what supports a healthier lifestyle and nothing that doesn't

4:03

and that's with over 100 nutrient vents menu items to choose from every single week

4:09

97% of users agree that factor meals help them live a healthier life so you can feel confident

4:14

that you're doing something good for yourself with every meal I've really enjoyed factor and if

4:18

this sounds good to you maybe you should try it as well let's eat real head to factor meals dot com slash

4:24

LWAI 50 off and use code LWAI 50 off to get 50% off and one free breakfast item per box for one year

4:33

while supplies last until double 31st 2026 that's code LWAI 50 off at factor meals dot com LWAI 50 off

4:41

at factor meals dot com see website for more details and we'll go ahead and get into it starting with

4:47

tools and apps and here we begin with ontropic releasing cloud opus five which they say comes close

4:54

to the capabilities of cloud a fable five in many domains and is cheaper of course so this is

5:02

following up on the release of fable five a little while ago fable being their new family of models

5:09

that they didn't have before and they also released son at five either before around the same time

5:15

as opus five so we sort of caught up presumably because these are distillations of fable and mythos

5:21

right so typically what you can predict with ontropic is their big kind of best model is their

5:31

most compute heavy most impressive model which is mythos right now and these things like fable

5:38

opus on it are kind of derived from it to some extent where they try to extract out the intelligence

5:45

at a lower price so I think not a ton to say about this one beyond that it's supposedly

5:53

quite great and close to fable five so pretty big jumps in the benchmarks relative to opus 4.8

6:00

the vibe check has been a bit mixed as far as I've seen people you know have a usual sort of

6:05

complaints about what the models are doing and it's hard to know whether we just have high

6:10

expectations now or in fact the models that are getting stupid there's also a new fast mode

6:17

in research preview which offers you higher speeds at double the price amenon I think we're

6:23

losing track of what it is that we're looking for just because that with the waterline is rising

6:27

so fast people are going to use to incredible levels capability you're right this is probably a

6:31

distillate of fable five or mythos five then added safety and all the stuff there's obviously

6:38

additionally post training that gets done to kind of further refine the character of the model after

6:43

that and so one of the key things that they highlight here is the differentiator of opus five is

6:47

supposed to be more emphasis on verification and judgment so less kind of raw capability but more

6:54

assertive like double checking its work making sure that what you're getting is actually correct and

6:58

so they give this example where like given it it's given a drawing of the machine part but no

7:03

way to view the original image and then opus five like rewrites its own computer vision pipeline

7:09

to extract the geometry from the raw pixels and reconstruct the part basically the idea being like

7:14

it's going to get you that raw data the original to base its conclusion on so that it knows it's

7:19

right no matter what that's kind of like the vibe here also an alignment and safety kind of

7:23

interesting this is always the game we live in a world where the u.s. government decided to

7:27

snap a chalk line at mythos level and anything above mythos level magically is is subject to

7:34

are the de facto licensing regime that we have in the u.s. and so in this case and thropic is

7:40

in a hurry to say that their model is close to mythos five that identifying software vulnerabilities

7:45

but it's less successful at developing exploit again that's part of the post training that's part

7:49

of making sure that or and also just as well the pre-training or other parts of the training

7:53

process where they're avoiding explicitly training it on cyber tasks in that way and so the goal

7:58

here is really to position it as this is a really intelligent model that it is okay for us to release

8:04

of course fable five is okay but it has additional safety of its over over mythos but you know

8:09

you're always now going to see that kind of background concern of this threshold which is I mean

8:14

it's good it would be just great to have a more principled approach to to the stuff one thing to

8:19

note to its performance on frontier bench so frontier benches this new benchmark that we have

8:25

was released just a few days ago and the same team behind terminal bench came out with it

8:30

this is big community effort and it is basically just like a harder agenteic environment we keep

8:35

meeting more and more difficult agenteic evals to be like all right but you know we just saturated

8:41

your terminal bench or whatever let's now move on to your terminal bench two terminal bench three

8:45

and ultimately this and and there you go so here we do know that this particular model opus five

8:51

is outperforming all other models on a cost per task basis and that's really where they're

8:56

trying to differentiate it is cost per task not necessarily frontier of intelligence near the

9:00

frontier but cheaper per token or per unit of intelligence let's say at that point it is

9:07

slightly cheaper than gbt 5.6 sole the biggest and best model from open AI and I think

9:14

maybe indicative of like pricing becoming more of a concern for customers at the business

9:20

front now that there is a competitor to on froth with codex and open AI being quite capable and

9:27

very very cheap alternatives for model usage after wonder whether we're going to be seeing more

9:33

kind of pricing pressure going on next up some more model releases this time from google google

9:41

deep mine has released three new AI models Gemini 3.6 flash Gemini 3.5 flashlight and Gemini 3.5

9:50

cyber so per the flash aspect these are cheaper and faster than let's say more intelligent models

9:59

to be 3.5 flash fly delivers 350 tokens per second at a rubber cheap price Gemini 3.6 flash

10:09

is priced 1.5 dollars per million input tokens and 7.5 per million output tokens so that's

10:16

slightly cheaper than sound at five and not super cheap and then there is of course the cyber security

10:22

focus Gemini 3.5 flash cyber which is integrated into this code vendor agent that autonomously builds

10:31

exploit code to verify vulnerabilities in sandbox environments and then generate patches which

10:38

they say has found problems in complex real pieces of software such as the V8 java strip engine

10:48

so i think interesting to see a general movement towards cyber focus with not just the mythos

10:55

and opus and so on we've seen opening i release a cyber model now google has released a cyber

11:01

model and we'll be discussing microsoft has also released a cyber model so if everyone is like

11:06

oh no we got to do something about this i don't think it's sharing too much to say like

11:10

people in the national security space and the frontier labs are really concerned about where

11:15

cyber is going and this view that you know we're in the volumpto apocalypse right now right we're

11:20

getting all these low-hanging fruit vulnerabilities being discovered and exploited assume we're going

11:24

to get a mythos class open source models sometime in the next you know certainly six months maybe a

11:30

bit less at that point you're going to need an answer and so all the lives are pre-positioning

11:36

for the moment when basically they're holding the world for ransom i mean you know you need to use

11:41

really really good cyber models shorter infrastructure or else like that is just going to be the case

11:46

as in the side i'll just like casually drop the prediction here that we may see some pretty significant

11:51

disruptive cyber attacks at massive scale not even just nation state or their proxies but literally

11:56

just like disaffected young people or you know terrorist groups or whatever that's just what happens

12:01

when you open source that level of capability it's just like that's what the math says well see where

12:06

that goes but that's like the default assumption right now a lot of the people in the space both

12:09

on the national security and the the frontier lap side so that's part of what the positioning is

12:14

is here if you don't have an answer to the cyber question you know you're going to be a lot less

12:18

relevant in the next six months or so they do this pretty impressive i mean the flash cyber model

12:23

which is maybe the one at least the one i'm paying most attention to is performing on par with a

12:29

lot of frontier agents that things like the cyber gem and that's an important benchmark i mean it's

12:35

a lot cheaper too right a really a fraction of the cost and so this idea of how cyber plays out

12:41

is always a really strong function of how much compute you have your defender you have a certain

12:46

pile of test on compute the attacker has a certain pile of test on compute can you invest more

12:52

test on compute than an attacker shore up your infrastructure is the is the question i mean there

12:57

there are a lot of ways to answer it and you know different ways to use test on compute and

13:01

a question about how much leverage like maybe there's an attacker advantage or defender advantage

13:05

these are all open questions but it's going to come down in some way shape or form to that balance

13:09

and so the cheaper you can make these models the cheaper you can make the tokens per unit of cyber

13:14

intelligence really the more value you're getting there so that's an area where the cost of the

13:19

tokens really really matters and that's why they're dabbling there yeah the other launches are

13:24

interesting but kind of fall into this general category of like google still not having a true

13:29

frontier model like when we're thinking about the best models in the world it's anthropic and

13:33

it's open AI and there's just not really anyone else yeah jimmy pro free free point one used to be

13:40

sort of in that phrase or at least yeah never frontier it there's not been a pro level model since

13:46

february so they're now quite a bit behind nobody really using jimmy pro for like cv as hard work

13:53

and then coding for instance and i think it is an interesting indication where google is at

13:57

if they are focusing on flash first because this immediately rolled out to everything all their

14:03

products google a studio android studio jimmy app jenemy enterprise agent platform so it kind of

14:10

makes sense from a business perspective like they integrate jimmy into everything including google

14:14

docs and spreadsheets and AI mode and at that point you have to have a faster and cheaper model which

14:21

is why they are really emphasizing flash first they did say that jimmy fp.5 pro is being currently

14:30

tested and will be made available when ready and they have begun the most ambitious pre-training run

14:36

yet for jimmy four so we've gone indications that they're at least working on a mythos level model

14:43

and we've seen them kind of catch up before with jimmy nice so i'm personally looking forward to

14:49

what jimmy four will be like next up another new model but this time not for language for images

14:57

and videos black forest labs has launched flux free that is capable of generating images and 22nd

15:05

videos with audio so this is a multi model frontier model trying to understand and generate images

15:12

and these models while extending the architecture interestingly to robotic vision and action so

15:18

it's jointly trained across image video and audio modalities altogether it's their first public

15:25

video generation model from black forest labs which for some background hails back to some of

15:31

a talent from stable diffusion that made them first really impressive image generation models

15:38

and flux has still been kind of one of the go-to image generation models at frontier flux

15:45

free video looks to be pretty impressive from what i've seen they have some kind of human preference

15:52

studies where they say it is preferred over a rock imagined video clinging v3 pro runway gen all at

16:00

like 70 percent 60 percent 80 percent whatever people prefer this in terms of its outputs and these

16:08

are from just kind of testing it's still not fully rolled out so very interesting and the fact that

16:14

it's now being adopted through flux mimic that is being developed with mimic robotics for actually

16:22

making it kind of an action model which i've seen kind of starting to be the case more we've seen

16:28

some other players like runway starting to get into like the physical intelligence space

16:33

video and robotic control seem to have a lot in common yeah this is interesting announcement this

16:40

kind of a couple things one is a bunch of comparisons that like show pretty lopsided wins against

16:47

you know like luma and runway and all the stuff that don't really matter because nobody uses those

16:54

models anymore but there's this interesting comparison against koo will's gemini omni flash

16:59

52 percent win rate against that that's actually quite interesting like that's pretty impressive

17:03

especially given the resources google's been throwing at this stuff and then another piece is so

17:07

so yes like we keep pushing out the length of clips so that you know that's great but the challenge

17:12

has sort of become coherence across clips across different shots and that's the big thing that

17:17

they're pushing here are some multi shot sequences where the characters are consistent the sort of

17:23

the physics is consistent and that's a big boon with this particular police so you know increasingly

17:28

moving beyond and what once you get those 20 you know 22nd clips 30 second clips you can imagine

17:33

that being a point where yeah you know one shot typically only lasts about that long as if I

17:38

know how long a shot lasts in professional film but whatever you know you can imagine that being the

17:43

case and so then you know maybe you care more about the switches between different frames so yeah

17:47

kind of interesting and a new kind of metric to track and they do have as before variance of this

17:54

that are open weight that also have an access to multi model generation called flux free dev so

18:02

another kind of slightly big deal I don't think we have an open weight model that has this new

18:10

unified multi model background which by the way is relatively new we've seen image

18:14

generation and video generation for a while but similar to kind of nano banana from last year I

18:19

think we're moving towards a place in video and audio generation where everything is put together

18:24

instead of being coupled together and that is actually a pre-book deal in terms of the capabilities

18:29

next more of a product story meta is making it's a chatbot more like an assistant so they're adding

18:39

productivity features there's a calendar integration daily briefings and apparently in-depth research

18:44

capabilities powered by alama spark 1.1 it can also browse Facebook marketplace search for restaurants

18:52

check your calendar and handle recurring tasks so this is rolling out to the meta AI app and is

19:02

going to be coming out to WhatsApp as well which continues to mark a shift for meta which is like

19:07

you're making AI are you just going to compete with all the other AI players like what what are you

19:13

going to be doing of this I don't know yeah I mean I think this is partly a realization as well that

19:18

unless you're moving in the direction of productivity you're just not going to squeeze all the juice

19:21

out of these these models that you can right so you know think about the positioning of open AI

19:26

relative to anthropic and the profit per token that anthropic is able to break in because of their

19:31

commercial focus it just I mean they're they're eating open AI's lunch and so think about the

19:36

there's an extreme beyond open AI we often think of open AI is the direct to consumer company which

19:41

isn't as true as it was six months ago certainly they've been making a lot of inroads in B2B

19:45

but at the far end of the spectrum in the other direction is meta that they are straight consumer

19:50

right like those tokens are going just to tickle your limbic system they're not actually going to

19:55

like actually move big things in the real world they're like make products and so if you want to

19:59

ultimately get the the most bang for your buck generate tokens that are actually valuable enough

20:03

to make you good profit yet have to move into this direction not least to say if you have a

20:08

coherent long term view where super intelligence goes humans just aren't in the picture which means

20:14

if you were optimizing for the value of the attention of human beings which is what meta is currently

20:19

doing that value may drop precipitously as AI is start to control more and more of the economy and

20:25

so you have to be in a position to actually do productive work and support agents in doing that

20:29

so you know depending on how far you wanted to read this you might read it that far I know that

20:34

doesn't really seem to understand super intelligence but certainly Alex Wang does so wouldn't be

20:38

surprising if this was at least part of the thinking here yeah I will say I think it will be

20:43

interesting to see where they go of this because you can go two ways you can sort of go and and try

20:49

to make codecs or a co-work competitor which is straight up just for work or they could shift into

20:56

sort of open claw type thing where this isn't always on the background agent which can do a bunch

21:02

of stuff for you including productivity things like briefings on your calendar but also messaging

21:09

and various things like that and I think the open-class base is sort of still up for grabs Google

21:14

hasn't rolled out their open-claw always online agent they've said that they are going to I forget

21:19

what it's called so I could see them like being potentially capable of competing on that front

21:27

not in the like coding or real you know office productivity side but like personal productivity

21:33

you know everyday productivity maybe and one last product rollout open AI is rolling out

21:40

charge-upity health to everyone so this is available to all US users aged 18 plus on web and iOS

21:49

and it will allow you to connect medical records and health tracking data for the chatbot so

21:56

they are saying that this model can reason at levels better than clinician level and you can

22:05

connect a whole bunch of stuff I think this marks a shift where you know in the past if you were to

22:10

talk about health stuff of these models they were very strongly caveat that you need to double check

22:16

and in general you should not have trusted these models with any sort of critical health concerns

22:22

judge of the health potentially is at least open air making the case that this is something you can

22:28

rely on after applications and business we begin with ilya sascovars safe superintelligence

22:36

partners within video to scale it AI research so that's kind of a gist of it we have had SSI

22:45

safe superintelligence be around for a couple years they raised one billion in founding in 2024

22:53

and a two billion in 2025 so you know a lot of money but not that much money if you are saying

23:01

you want to create superintelligence you compare that to an fronpic open AI we have hundreds of

23:06

billions this is a few billion so the narrative around this is that ilya sascovars company has

23:16

achieved sufficient research progress that it's time to scale up and now to scale up you need a

23:25

bunch of compute and so they're gonna be partnering with and video in a value of like some amount

23:31

of billions a bunch of billions and they'll be increasing their compute by an order of magnitude

23:37

yeah there's some some disagreement between different outlets about how much exactly has been raised

23:42

whether it's a five billion round or just in the billions or something like that i think tech

23:46

cruncher and five billion dollar story so either way the one thing everybody seems to agree is

23:51

number one it's gonna give safe superintelligence access to the your Rubin platform right so that's

23:56

the next generation platform five billion if you do the clungs law more is law analysis roughly

24:02

allows you to 10x your compute relative to the one billion dollars that that they'd raised previously

24:07

and so well there you go they're they're 10xing their compute a couple of things are interesting

24:11

about this so yes there's this this narrative that they're like and i agree with this most likely

24:16

this is what's happening it is a pretty straightforward guy you know if they say that they've got in the

24:20

point where they're at that next level of sort of proof points that they can take this investment

24:23

it's worth scaling it probably is there's this kind of more cynical take that like oh they just ran

24:28

out of compute which you can hold that view that's totally legitimate i suspect that's not the case

24:33

but just like so everyone's tracking that is another another explanation they have no product they

24:38

intend to launch no product which means their only revenue is going to come in the form of these

24:42

sorts of investments it's a weird sort of moment and story for them because they are you know

24:48

the Daniel gross with the co-founder of safe superintelligence along with ilya back in the day he left

24:53

for meta after meta offered to buy the whole company will clock and ilya said no so Daniel jump

24:59

ship at least at that moment you can argue that that meant at least annual gross thought that

25:05

his chances of making something like super intelligence were higher at meta than by remaining at safe

25:10

super intelligence what's happened in the interim we don't know and ilya has dropped only the

25:15

faintest of hints on dorkesha's podcast about generally you know generally sketching that continual

25:21

learning is going to be part of it and going back to i've heard a couple rumors but like i haven't had

25:25

any of these verified that anyway they are looking for let's say somewhat beyond the standard

25:30

how it's going to save you on the standard model it's a very physicist joke but i have you know

25:34

things that are a little further afield and so that sounds like it would almost have to be true

25:38

just because the otherwise you're in pure scaling mode there there could be this narrative you

25:42

can imagine people sort of like laughing about this and say oh well ilya said the ear of scaling

25:47

is over what's he doing raising five billion dollars to ten x his compute and to that i say

25:53

ilya never said that you wouldn't also need scale it's both all right the what he's saying is

25:58

there is leverage to original research again and they're the biggest leverage is not purely in the

26:04

engineering of more and more scale systems it's in something else like you can compound it very

26:09

effectively now with algorithmic insight so do with that what you will this is an interesting story

26:14

and we don't know much about it yeah it's an interesting story in a sense that you can be very

26:19

curious about what they figured out and no nothing because we still have nothing to go on it's

26:25

kind of funny if you go to their website and go to the updates page it's like free things it's

26:31

literally since 2024 they've released publicly two updates which are just about the co-founder leaving

26:40

and now this partnership so hopefully we'll get some more understanding of what they're doing

26:45

soon as they scale up but it will presumably be a while since scaling up is not easy and in fact

26:52

Nvidia had said that they invested after quotes obtaining rare access to the company's closely

26:57

guarded research so supposedly they're tracking who knows right but there you have now

27:04

and a related story about five billion dollars AMD has committed up to five billion dollars to

27:11

on frothic viz is a new partnership on frothic will deploy up to two gigawatts of AMD's instinct

27:19

mi450 aijb use and their new helios wreck scale system plan for deployment in 2027 on frothic has

27:30

so many partnerships now with so many they have like SpaceX AI they have Google they have amazon

27:37

and now we have AMD I feel like we're just like go to everyone to be like we need compute let's

27:44

partner up and give us some compute and AMD it has been trying to compete harder with these

27:50

AMD instinct chips honestly I don't recall where they are at with that but they do seem to

27:58

at least potentially have the ability to compete with Nvidia which no one else really does

28:03

right aside from tpu's from google and potentially the hardware that some of these companies are

28:09

developing yeah and this by the way this idea of anthropic having like a million different partners

28:15

it I mean it's really true right but partnerships with google for tpu's partnerships with Nvidia

28:20

partnerships with amazon right partnerships with the SpaceX AI and now with AMD the you know the

28:26

the golden rule if you're ever trying to explain well why a frontier lab is developing a new

28:31

partnership of cultivating a new partnership again commoditize your complement that's everything

28:37

that's going on in the space right now you know we talk about this a lot on podcasts but like the

28:41

history lesson here is Microsoft back in the in the day realize that laptops are super expensive

28:47

and software is cheap well actually if we just make all the the laptop manufacturers compete with

28:52

each other and we make one set of windows software that goes across everything and make basically

28:57

the software the choke point and the value chain then suddenly we can make all the hardware vendors

29:02

compete away their margins make laptops super cheap now that laptops are super cheap consumer is

29:08

dive into the market and obviously they've got to have software run laptops will come to Microsoft

29:13

right so everyone's constantly trying to make their complement the compliments to their offerings

29:19

compete with each other anthropic wants all of the GPU design firms any company that makes GPUs

29:25

it makes compute they want them to compete with each other like crazy so when it looks like they're

29:29

setting up a partnership that's law-opsided and like they only have the media GPUs in videos got

29:34

ton of a ton of leverage in that relationship now well and throttpicks can go off to AMD or to

29:39

Google or SpaceX AI and say hey we want to get our compute from you instead and so now in the

29:43

ago oh no no like we'll give you just kind of this how how the pricing control gets set up in the

29:48

space and likewise in video and all these players are trying to do the same in reverse right in video

29:53

wants to help small baby front your lab come big adult front your life but they have more customers

29:59

and also so that anthropic feels more pressure to buy more compute so it's kind of happening in both

30:04

directions as everyone's kind of pulling their knives out and dancing around each other it's a

30:08

wild time in the space but AMD has been a laggard in this space you know you think about basically

30:14

their their software stack the competition to kuda which is just in videos widely viewed as like

30:19

in videos big moat for AMD their equivalent is called rock M and a big part of the purpose of

30:25

this agreement is that clawed is going to tune workloads for instinct GPUs which the AMD GPUs

30:31

and accelerate rock M development that's key right we saw that with with Amazon it's not a coincidence

30:36

that every time anthropic signs one of these big compute deals with a hyper scalar that the deal

30:41

it involves anthropic has to tune their workloads for that hardware they must use a minimum amount

30:47

of that hardware this is about giving feedback to the hardware designer that is so so valuable

30:53

because otherwise you just can't you can't design your GPUs for the next generation if you don't

30:57

know what the next generation of architectures going to look like that's a huge part of this

31:01

you know negotiating leverage we talked about oh yeah this is also the first so helios so this is

31:06

all part of not just the instinct MI 450 series GPUs but it's also about AMD's helios rack scale

31:12

systems that's the first full rack scale system that they're selling that you can you I mean you

31:17

can think of this as like the equivalent to the you know the nbl 72 the sort of full rack that

31:23

Nvidia will ship this you have something you can literally just like plug in a data center instead

31:27

of just shipping the GPUs themselves that's you know AMD's going further up the stack to own more

31:32

and more of that that infrastructure layer and so it's also got 72 GPUs per rack which is

31:38

amusingly like the same footprint as the nbl 72 but with completely different kind of power

31:43

consumption profiles and stuff like that's anyway super important I mean they are trying to prop up AMD

31:47

because they want AMD to be a viable alternative five billion dollars does buy the mistaken

31:51

in for a big pre IPO I guess but only just seems like more of a strategic partnership than anything

31:58

yeah it's kind of if you look at the press release it's like on for a big we'll be deploying

32:03

these chips companies will collaborate to use cloud to optimize workloads for AMD instant

32:08

GPUs and AMD will broadly adopt cloud across its engineering and product development teams

32:14

so it's you know AMD is going to invest in on frantic but really the point here is we're going

32:22

to work together to kind of have a win-win type situation and to your point I don't know if it's

32:28

necessarily about the Microsoft type story me that's part of it but also it's about redundancy

32:34

and scale for frantic they did get into a nasty situation earlier this year where their sole

32:41

provider or their primary provider was Amazon for a few years we didn't have their own compute

32:47

and they sort of were left unable to deliver enough compute to their customers and speaking of that

32:53

we have another related story that meta apparently isn't talks to least computing power to on frantic

32:59

in a potential 10 billion dollars deal so we covered this a think last episode where meta might be

33:06

going into the neo cloud business they've built so many data centers that it potentially make

33:13

sense to be like well we have these data centers how about to make some money for them so we'll see

33:17

if it happens meta is spending up to 145 billion dollars on capital expenditures in 2026 so I'm sure

33:28

some of the business folks over there wouldn't mind getting some revenue from it yeah it's also I

33:34

mean to put in context it is way smaller than a lot of the other deals that anthropic has already

33:39

negotiated we we are learning that the proposal itself came from anthropic back in June so this

33:44

is an anthropic going like hey we saw this little like kind of flirty announcement that you guys put

33:49

out that and maybe we're thinking about offering some AI infrastructure maybe we'll do it and

33:53

anthropic was like oh holy shit we want that the scaleless mall so you know if you look at the

33:57

deal they signed with SpaceX back in May that was about 1.25 billion dollars a month so that is

34:03

about three times the size of this meta agreement if it goes for yeah I mean this would be a new line

34:10

of business for medas you know unclear whether they kind of sustainably think that they will be in

34:15

this business in the long run they certainly have we talked about the advantages that they have

34:19

structurally there's like a really big company and financing matters a lot for neo clouds right

34:24

you're constantly battling like concern over your debt load if you're having to buy a lot of GPUs

34:29

ahead of time that's often often the case you have high operating costs and things like that so

34:33

that is a good position that they want to theoretically for a balance sheet standpoint the big risk

34:38

for them is just going to be do they have the technical savvy to build in the right kind of

34:43

infrastructure at scale they've been doing some of that but they haven't been specializing in you

34:48

the rl rollout stuff in the the massive scale pre-training again they'll learn a lot from

34:54

anthropic in this case I think there's a lot of value here if meta wants to proceed this would be

34:59

the deal to start with just so they can learn from the best in the business how to actually like

35:04

pop up you know set up their architecture their optimizers their data makes you like all these

35:08

things they'll learn a lot about that inevitably from this partnership so we'll see but it seems

35:12

like it would be strategically good for them and now to a less big player we haven't had a story

35:20

about any companies raising over one billion in around yet so let's do that fireworks has hit 17.5

35:29

billion dollar valuation so they got 1.5 billion funding around which let them throughout

35:35

all your creation they say they have exceeded that one billion dollars in annualized revenue 5x

35:42

from last year they compete in the inference cloud markets they can host AI models for developers

35:50

similar to amazon google and microsoft this includes both your own custom models and the open

35:57

source offering so if you want to use kimi for your own applications one way you do that is going

36:03

through fireworks for instance and we're interested to see if we can a grow of open source models

36:09

better useful and competitive will make companies such as fireworks and rock even more of a player

36:17

they already are now but they have room to grow and actually eat into a business of opening

36:22

a high on a topic yeah and then there are you know in various ways competing with some big players

36:28

you know on like model hosting you got amazon you got google and then you know together AI even

36:34

is you know pretty because this category just like exploding and that generally is just bullish

36:38

for a lot of companies but this is it's not like the other they're the breakaway here there is

36:43

massive scale that helps a lot especially when you're doing inference just because of batching

36:48

right you're able to like have much larger batches of data that you then feed through your pipeline

36:54

and the larger the batch in general the more efficient compute efficient your models are going to be

36:59

so this is a a case where it's sort of like back in the days of a bold sass yeah you would have

37:05

something that works and once it works like you want to violently scale it as fast as possible which

37:09

is exactly what venture is so when you think about the arguably smaller set of companies that are made

37:16

for venture capital investment like this is one of them you want to look at companies that show

37:21

significant nonlinear returns at scale and batching and a bunch of other

37:26

amortization dynamics that really favor large scale deployments are putting in this direction

37:30

which is why you're seeing you know 17.5 x revenue multiple like that is pretty wild that's big

37:36

even even at the stage actually might say especially at the stage I've lost track of like what

37:40

stages are supposed to be I guess a trillion dollar exit is the only cool thing now so maybe you

37:46

know maybe maybe there's still a baby startup but anyway now what is money anymore what is

37:51

evaluation you know another way of saying tokens right yeah yeah last story now moving to

37:59

something related to software opening in google are selling AI models to black listed China

38:06

groups kind of a funny way to phrase that they are not selling AI models that would be crazy

38:11

but they are providing AI services to some companies through singer per based subsidiaries of

38:19

Alibaba by view and Tensed which are Chinese tech giant that are black listed by the Pentagon

38:27

for alleged ties to China's military so technically this is legal these are not quite Chinese they are

38:35

in Hong Kong and Singapore open AI and Google are saying that they are doing this with protections

38:42

against distillation but you know you can read it a couple of ways depending on your views on such

38:48

shrinks yeah there are also just like all kinds of arguments going every which way saying that maybe

38:54

you actually do want your adversary to be using your servers to do their training or or do their

39:00

they're inferencing because it just gives you access to information and it also reduces domestic

39:05

demand for the development of competitive platforms and I mean okay I think at a certain point you

39:12

got to just bite the bullet this is just like personal opinion jare talking but like if your hope is

39:17

to like go after China piecemeal you know a little bit here and a little bit there the reasons to

39:21

think that that actually only helps the Chinese kind of inch by inch build up their whole domestic stack

39:26

but in case I think in this particular instance there is an interesting argument and this is all

39:32

through this through a Singapore loophole right so so yes there is an entity list that you have

39:37

ties to the the people liberation army the Chinese military and yes it is normally legal to do

39:42

business with those entities unless they have subsidiaries operating in Singapore in which case

39:48

magically everything is fine right so this is like an a loophole that is known to exist I personally

39:52

think like it's really unclear me why this loophole exists I'm fascinated by this one in particular

39:57

because it's almost like the kind of thing that you would intentionally leave in if your intent was

40:01

to just label loophole for some kind of ideological reason like there's no one I've ever spoken to

40:06

on the AI export control side who who understands why this is the case and so if you're part of the

40:13

niche group at the department of commerce that actually like has an argument for this it would be

40:17

super interesting to know why this is the case so anyway yeah as it says kind of a weird headline

40:22

to read but absolutely legal and absolutely fine and now over to projects and open source talking

40:29

about all these exciting open source models we've been referencing starting out with kimi k3

40:36

so that's been one of the big stories of past couple weeks which is from munshut AI and kimi k3 is

40:43

their largest released yet a 2.8 trillion parameter open weight model that is aimed at coding

40:52

knowledge work basically competing with cloud code codex and so on this is massive obviously 2.8 trillion

41:00

we don't have as compares to opus or any of our closed models but in the space of open source models

41:07

it's very big has 896 experts so still make sure of experts as usual 16 active per token still

41:16

going at a 1 million token context window and the story roughly I think both benchmark wise and

41:25

in terms of a vibe check is that this is maybe around opus 4.8 and jubilee 5.5 level so very

41:36

capable very like you can use this as the driver of your coding agent it may not be exactly at

41:42

the frontier but it certainly is capable enough to make you productive you know a few months ago

41:48

this would have been a frontier probably it's priced pretty expensive for open source so 3 million

41:55

per uncashed in put token 15 dollars per million output tokens that's less expensive than opus but

42:03

more expensive than sonnet it's kind of in that range of fairly expensive models and a related

42:11

story is that after it was announced munshut AI halted new subscriptions among a compute crunch

42:21

this allowed people to subscribe just apparently because they don't have everybody to serve

42:26

all the demand meaning that presumably they have a lot of demand absolutely and the classic

42:33

problem you know we always talk about with china is obviously compute scarcity and the fact that

42:37

in the context of a model like this right this is a behemoth like many trillions of parameters

42:42

you're obviously not running this on your laptop this is a model that is meant to be used by

42:47

big-ass companies like neoclouds and run like hosted on big hulken infrastructure or bhi so when

42:55

they look at the actual requirements like hosting this it's getting cost you like 64 h100s or

43:01

or b200 gpus across eight servers and so that's a lot of money so obviously this is not for

43:08

casual use this is for people who are competing at scale with providers like you know your

43:13

mistrials or whatever you know like people who have their own apis for open-weight models and so

43:17

you have very big laptop to run this right yeah you should see my laptop it's size of room yeah

43:24

and so michael kratzios who's over at dsdp the office of the science and technology policy at the

43:29

White House came out with his accusation saying you know k3 was trained on not only on band-in

43:35

video chips but also on distilled data from you know anthropics fable and munshut hasn't responded

43:42

publicly there's pushback i mean they can Lambert had an analysis saying basically the results suggest

43:47

that yes there was adversarial distillation it contributed somewhat but marginally and that munshut's

43:53

competing with anthropic and oboe i just like way fewer resources that could all be true at the same

43:58

time by the way there is literally no contradiction there whatsoever it is the case that they are doing

44:03

this like large scale distillation and that last little bit can make a big difference also the case

44:07

that weirdly like it sounds weird for a White House person to be complained that things are being

44:11

trained on export control chips when like the policy on export control seems to be yollowed so hard

44:17

so like it seemed like the department of commerce itself based on some congressional testimony from

44:20

a few weeks ago like they don't even know what their policy is they're just like kind of flipping

44:24

back and forcing oh no a truly climate or thing that we said before when we said everything was fine

44:28

it's not fine and like everyone said dang we think the thing we allowed is turning out to hurt us

44:33

in some way oh no exactly exactly and I mean I think you know theoretically these were banned chips

44:39

but also enforcement actually matters it turns out and like the is is just not it's not in fairness

44:46

them they're not equipped they're not tool they don't have the resources that they need to do this

44:50

which is why you know there's been so much effort in congress to pass legislation that would

44:54

authorize a larger budget for them but still the White House hasn't exactly been bullish on

44:59

unshortified p is to have them to have them do their job so anyway this is more or less what you

45:03

should expect when that happens powerful Chinese model development will continue until morale improves

45:09

yeah and anyway so there's a whole bunch of additional noise when you look at the Chinese

45:14

ministry of commerce been talking to a lot of the big labs and hyperscalers in China,

45:19

Jirpua, Alibaba, Bite Dance about tightening its own export controls on AI models and training data

45:26

and I mean yeah I may have more a little more on that later but yeah it's it this is like a

45:31

really important access to track is like how yeah how China is viewing is viewing data export

45:37

is a really strategic indicator of their stance on this alongside this we did get a technical

45:43

report as we have in the past another one of these beefy beefy papers that goes on in this case for

45:51

only 34 pages so not quite as much as usual a couple architectural innovations but get

45:59

you know very nuanced with kidney delta attention and attention with residuals we're getting

46:05

into some very kind of deep optimizations of the transformer architecture partially to just

46:12

enable scaling and kind of effectiveness at this one million token context window and also just

46:20

some just complete hardware artistry black magic of making the chips work for you and optimizing

46:29

stuff that is I don't know if I try to read this paper it's going to take me a month to understand

46:34

all the details but the short story is as we've seen in the past with deep seek with also Munchadei

46:41

they are displaying some very deep technical capability and as tempting as it might be to some to

46:49

be like oh it's distilled blah blah blah say clear that there's some very capable people and it's

46:55

still very nice to have these technical reports giving us a fair amount of detail on certainly

47:01

the architectural details and this is some extent the training details as well although the exact

47:08

data composition for instance we don't know which is a big part of it now onto another big open

47:14

source l l m or not l l m exactly thinking machines has released their first big open source model

47:23

and open weight mixture of experts with 975 billion total parameters and this is notably

47:32

multi-mortemodel so it combines text image audio and video data reasons natively across all

47:40

for modalities about currently outputs text code and structured data and that's kind of a

47:47

positioning here that it's not going to be sk pool as some other models and broadly isn't

47:55

necessarily about coding by itself but thinking machine positions it as like we want to cover

48:01

everything in the capability space they release this sort of like breakdown where we show

48:08

where our model lags and basically everything against frontier models as far as what frontier models

48:14

are good at but there are areas where frontier models aren't optimized for that this is already

48:21

capable at so pretty notable for being one of the first big open source releases from a western

48:28

company we've had and video releasing an ematron at a fairly significant scale but I think this

48:35

might be the biggest non-Chinese model at almost one trillion total parameters you know the

48:41

vibecheck I've seen has been pretty positive haven't made any sort of grand statements or claims

48:47

and it is qualitatively a bit different from other models in being so focused on multi-modality

48:55

so thinking machines really working the world model side of things in part and a lot of this is

49:00

a strategic effort to like that on the efficiency side you know sparse M O E's by the way the

49:06

stack has a lot of deep sea lineage to it so if you're ever wondering of people are saying like

49:11

deep sea is not a serious player I mean thinking machines and you look at their pedigree I mean

49:16

obviously it's it's a wild team like they're very good at what they do when you look down the

49:20

stack I mean so much of this is deep seat coded right so even down to the fraction of active parameters

49:27

per pass this whole hybrid global attention thing so basically like local global attention so they

49:32

have some layers that attend to like all all tokens in context and then others that are more tight

49:37

focused the numerics as well as interesting so the bf16 and nv fp4 support nvfp4 is in videos

49:45

floating port point for for numerical format and it's blackwell made it so so this is designed to

49:51

ship to work really well on blackwell yeah we've been talking for a while about how in video is

49:56

trying to position itself as the open source tighten just because you're going to expect to see

50:01

all these neoclots pop up and they're going to be running open source models right that's that's

50:05

what makes sense and so having you know encouraging the open source ecosystem to move towards

50:10

video kind of blackwell native formats like orbit float is pretty pretty interesting and anyway

50:16

that's that's all part of the strategy here so like looking at the numerics actually matters a lot

50:20

it sounds boring but like you know how are you representing the weights in the model turns out to be

50:26

quite a tell about your strategic direction and they cite this bridge water collaboration where

50:31

they were able to fine tune one of their open models via tinker to this like 84.7% on some

50:37

financial reasoning benchmark which is impressive it beats top proprietary alternatives at under 10

50:43

percent of the cost so again you know this cost argument being made and that's in large part

50:47

can do the compatibility with with the blackwell hardware that's coming online and now one more

50:52

open source story not related to models we've got scaling agentic arel 365 thousand

51:02

environments for software engineering terminal and search this is coming from prime intellect and

51:08

they have unified 23 agentic task data sets across all these things into a single API can

51:16

a release they call verifiers v1 that adds up to that level of tasks for evals and for

51:24

a well training has almost 200 thousand software engineering tasks 29,000 terminal tasks and

51:31

a whole bunch of search tasks that is all unified under kind of one inference setup so we've had

51:39

all these benchmarks floating around like 20 different ways to evaluate software capabilities

51:46

bunch of ways for terminal we're basic story here is that this unifies all of them into one framework

51:52

and makes it kind of reliable and repeatable to run it which is very important when you do model

51:59

development and do any sort of research evaluation both for training side of reinforcement and for

52:05

the ability of the ocean side of knowing how good your model is prime intellect is in the space

52:11

training their own models at scale as you've covered in the past a while ago so presumably they're

52:17

doing it for their own model development needs but also for the broader ecosystem.

52:23

Yeah and it's quite an interesting and classically prime intellect type of maneuver here so

52:28

they're first of all they're kind of solving two problems one is that as you said there's like you

52:33

know 23 different agentic tasks sets here that they're working with and like each one of them has

52:39

its own you know harness its own way that the the like say containerized like the image is set up

52:45

its own like grading scripts even and different failure modes like they're all these very bespoke

52:50

things and so if you want to train one agent across a bunch of them those incompatibilities are

52:54

a nightmare you need 23 different bespoke pieces of adapter software and so that's exactly what

53:01

they're doing they're hiding everything behind a single task set API that had now contains like

53:05

365 thousand tasks across across a bunch of different domains but one other thing that they're

53:11

doing is cleaning that data up they're finding that just like a lot of the RL environments that are

53:16

set up are like you know for example on the cyber side some of the environments that require you

53:21

to solve a problem don't actually have a problem in them like they already work out of the box

53:26

and like when you shave off these kind of broken cyber environments you kind of end up with a

53:31

large fraction of things that you lose and so they've been not only reconciling all of these 23

53:37

disparate things together but also shaving off stuff that doesn't work well so a lot of that

53:42

had to do with like finding eliminating opportunities for reward hacking and one key thing that they did

53:47

was they preserved the original grading functionalities in these stacks and the reason you would do that is

53:53

so you can still compare the agents performance on the benchmark to the original published work

53:57

because otherwise the way people would solve this problem in kind of a janky ways that's

54:01

well yeah I'll impose my own grading structure on this this e-value this benchmark and then you're

54:07

like wait if I run GP 5.6 all on this I get a different result from from what this you know from

54:13

from what the original paper said and you know this is this is a challenge so this allows them to

54:18

reconcile that by keeping the original rating so very interesting very important work at a prime

54:22

intellect they are obviously ideologically this like very pro open source type of company and so

54:27

there you have it as a small business owner sometimes it feels like no matter how much planning

54:34

you do there's always surprises like an urgent expensive repair but here's a surprise you will

54:41

like with progressive small business owners save 13% on their commercial auto insurance when they

54:47

pay and full so enjoy a surprise for once get a quote in as little as eight minutes at progressive

54:52

commercial dot com progressive casualty insurance company and affiliates discounts not available in

54:57

all states or situations when we are to policy and safety and we'll begin with one of the big stories

55:05

of the past couple weeks openly I has said that it accidentally hacked hugging face with a new AI

55:14

so the gist of the story is apparently going back to July 16 hugging face initially I think

55:23

discussed this while doing some cyber or software evaluations in a sandbox so typically when you do

55:30

these evaluations you put the models in a little container and you tell them you try to do this hack

55:35

and the model is supposed to be inside the container not able to mess with anything in your own

55:42

computer and you know infrastructure of anyone and what happened here is the model was very

55:50

intent on getting the right answers so it escaped the container of a sandbox then it hacked into

55:57

hugging face to get the answers to his exploit gym data set which of course I've seen a lot of

56:04

discussion on this this has kind of made it to remain stream in terms of people this whole narrative

56:10

of a model got out escaped and hacked someone else has become a big discussion point with a lot

56:18

of misinformation of like the model decided to hack a competitor or whatever this is not kind of

56:25

a sky net scenario but it is a very clear instance of misalignment for one where the model instead of

56:32

trying to actually do a task decided to cheat and like very aggressively cheat as well which we've

56:39

seen before with gbg 5.6 in particular matter has said that this seem to be the case with this model

56:46

you've seen a a a a also say that they were able to jailbreak this model very easily so there's

56:52

many things to be said about the story to me the main thing is a that this is another instance

56:59

showing an example of both the degree to which alignment is important in this day and age and

57:05

cyber is a real thing to worry about and that opening I hasn't been doing a good job especially

57:11

with gbg 5.6 it's a really misaligned model and it clearly didn't have enough actual security

57:19

infra to catch risk in any sort of timely matter apparently this was like a while later when

57:25

engineer was looking at was going on he realized this happened it's been a lot of fallout we'll

57:31

be discussing but it's both less of a big deal than it might seem to people not in the loop but

57:38

also a bigger deal in some ways I'm sort of struggling to find a way in which this is not a big deal

57:44

let me try to make this argument so if this is not a warning shot that we freak out about I honestly

57:50

don't know I mean with the there are takes here with people pushing back on the term like the

57:56

user the term rogue this was I think by any reasonable definition a rogue AI incident why am I

58:01

saying that what was the incentive for opening I to want this to happen obviously zero in fact

58:06

they have billions of dollars writing I mean hundreds of billions writing on not having incidents

58:12

like this occur and amazingly some people are still trying to make it be like oh this is a PR

58:18

marketing stunt which is just ridiculous right and a lot a lot of these same people are the people

58:23

who claim that an incident like this simply could not and would not occur so I think they need to

58:27

just kind of like sit with this moment touch grass a little bit because this is like we're beyond the

58:31

point where that is a reasonable position to have just straight like you heard me on the pot we've

58:35

had a lot of conversations about like yeah anything to have blah blah like I'm a very like I got a

58:40

wide range of possibilities and and generally like not in favor of judging people for their

58:45

pains on any of this stuff this is one place where it's like if you were looking at a situation where

58:50

again from open a high standpoint this incident occurs and then what's the the the

58:54

the netlet with the natural reaction of any polity is going to be like Jesus Christ we need to

58:59

regulate this space that regulation is going to throttle the rate at which you're able to put out

59:04

frontier models and as we keep talking about on this podcast and as is obvious well established fact

59:10

the amount of time during which a frontier lab has the leading model is the period the most

59:16

critical period for their profitability and their monetization of their model that's how they

59:20

pay back their R&D costs right they're waiting until the next competitive model comes up now if you

59:25

slow down the frontier and you don't slow down you can't slow down the open source ecosystem then

59:30

all this does is it road opening eyes margin there is no sane reasonable rational analysis of the

59:35

situation that leads you to conclude that opening eye headache do they need more market share

59:40

did that or sorry more mine share I should say do they really need more attention is that the

59:44

thing that's missing for them and is this the right kind of attention ahead of an IPO when when

59:49

this is starting to like raise questions about whether the US government might nationalize labs what

59:54

the hell happens to the value of your opening eyes stock after IPO they nationalize labs anybody ever

59:58

thought of that this is nonsense this is silliness look an opening eye agent went rogue it it was

1:00:04

running for four days on hugging faces servers the frickin FBI had to get involved when they thought

1:00:10

they thought it was some AI agent who knows yeah by the way hugging face was in one with was like

1:00:15

oh we're being hacked who what is going on and then this came out as as we insane yeah it is a

1:00:22

big deal like we have seen some stories of inevaluation the models were misaligned and tried to

1:00:27

cheat this has happened before the only way in which this is not a big deal is if you kind of

1:00:33

misunderstand the story to me in the AI like went evil and decided to go hack some companies

1:00:40

this is a classic case of technically vei did what it was told to which is get results

1:00:47

but obviously it's it's doing it in an exactly wrong way it's yes complete misalignment but you

1:00:53

can make the case it's not sort of as bad as it could be if you don't get into details absolutely

1:00:59

it's just that this argument we're finally at the point where you'll notice like I was the one

1:01:04

who was having to bring the imagination for the last five years I kept telling people like hey

1:01:08

you know you may actually get stuff like this and people get saying no it's not possible

1:01:12

it's now literally happening and the defensive move is to say in retrospect I will kind of

1:01:19

understand what yeah that's great you got turned into into a pile of computronium and now you're

1:01:24

looking around you going like ah but I see the mistake I made in retrospect like that's cool

1:01:29

but now your house has been destroyed your children have been kidnapped murdered and turned into

1:01:33

computronium the problem is that you you just keep running this forward and like okay let's do

1:01:37

this with super intelligence then the more intelligent the system is the more access it has the more

1:01:41

we offload to it this is literally a hack if this is just like the water supply or something like

1:01:47

literally people die so I'm just pre-registering this as a high confidence prediction at this point

1:01:53

that there's going to be another incident like this there is always going to be a fascinating

1:01:59

post-hawk rationalization that a lot of people will offer it's going to sound really reasonable

1:02:04

because it's going to sound like the voice of someone saying the future doesn't look like science

1:02:08

fiction the future looks reasonable and calm but it's going to be a after the fact analysis

1:02:15

that rationalizes rather than predicts my prediction is this is going to continue there will

1:02:20

unfortunately be casualties at some point and what that leads to is whip lash what that leads to

1:02:25

is kind of thoughtless policy and another mythos moment I don't think that's good for anyone

1:02:30

and this is why I think a lot of the skeptics are not doing themselves much of a sort of especially

1:02:34

if you're on the open source side of the house I mean like you're free to sit in the in these juices

1:02:39

but I'm just offering up the the humble prediction here that these takes are going to age very

1:02:44

poorly very fast so you have 12 months from now rather than I think a very different conversation

1:02:49

yeah so any discussion of this that doesn't acknowledge that this is a big deal and I think it's a big deal

1:02:56

by itself as an example where we're at but also as an demonstration of bigger

1:03:04

topics at hand of misalignment cyber capabilities safety broadly speaking you know we've

1:03:10

covered these topics a lot on the podcast and there has been a lot of dismissal of cyber with

1:03:16

mythos for months people have been like this is all TR alignment has been a story that for a decade

1:03:23

probably safety people have hammered on and yeah this is a very clear case of basically the classic

1:03:30

paperclip story of like it a model is told to maximize paper clips that goes on to make everything

1:03:36

paper clips here a model is told to fix like do well on some benchmark and it hacked some website

1:03:43

to get the answers now to be fair open area has said that as part of this evaluation they were

1:03:51

running GP 5.6 and a more powerful internal only model we've reduced cyber refusal for

1:03:58

evaluation purposes so this is internal testing benchmark of cyber capabilities not necessarily

1:04:06

indicative of potential incidents with their public products but I do think that kind of a state

1:04:14

of safety at OpenAI is an important dimension of this taking together with the other things we know

1:04:20

about GP 5.6 which is it did cheat and an unprecedented level on matter it was jailbroken very easily

1:04:27

by ASI or AISI and to me all this points to OpenAI is very aggressively trying to be on

1:04:36

capabilities yeah if you aggressively try to be on capabilities you're going to do a lot of

1:04:41

reinforcement learning if you do a lot of reinforcement learning without being very careful your model

1:04:46

can get misaligned very easily we have another example of a goblin kind of speech thing from like

1:04:52

a month ago where like they released a model that was obsessed with goblins because they trained

1:04:59

it in a way that wasn't necessarily you know it was all right but there was unintentional side effects

1:05:06

and that's exactly how you get misalignment if you like optimize your model very hard without being

1:05:12

careful it can they easily be optimized towards things like cheating because ultimately what

1:05:17

of these models optimize to do where optimize to solve the task and one way to solve the task is by

1:05:23

cheating this is the classic story of our L model is going haywire and that seems to be the case

1:05:29

with GP 5.6 and whatever this internal model is so I think this might be an aspect of a story that

1:05:35

will not be discussed as much but I think is an important component of like opening i in particular

1:05:45

having this problem right now although in the discussion around this the fact that we've had

1:05:52

previous incidents from open AI of evaluations where like apparently this already has happened we

1:05:58

also know that anthropic with mythos there was somewhat of a similar case of escaping containment

1:06:05

so to speak during evaluation so it's it's not necessarily just an opening i problem but taken

1:06:11

together it's a pattern that is very concerning I guess with good news is it's happening in these like

1:06:17

low low damage kind of instances and people are now aware of these issues and very likely we'll see

1:06:25

significant fallout including some stuff in in the legal side we have to discuss yeah I mean

1:06:31

were were my predictions were completely wrong was I by my honestly I thought we would be

1:06:35

be dead by them that stuff whatever whatever comfort people want to take from that yeah I mean

1:06:39

you know there's this this view that I certainly held to that you'd have a much more kind of rapid

1:06:44

inflection may i capabilities potentially yeah I wasn't 100% on that no one can be but that was kind

1:06:48

of my one of my my mainline views and so it's nice to have warning shots like this I may I can

1:06:54

confirm there've been other unreported incidents like this at opening i at a minimum and that the

1:06:59

internal reaction of this among some people has been a lot of alarm and discouragement at the

1:07:05

fact that there is clearly under investment in this I will say on this question of you know the

1:07:10

safety is being removed from these models for internal deployment we talked about this I think three

1:07:14

weeks ago in our last episode but internal deployment is absolutely like should be maybe the thing

1:07:19

you're most worried about which is why i'm skeptical about a lot of these you know they're literally

1:07:25

testing a model right their model could be evil for you know exactly and they they should be

1:07:30

testing it too that's that's the problem right so it's not as simple as saying like hey open

1:07:34

AI like you shouldn't have removed the safety it's like okay fine well then how do you propose

1:07:38

that open AI comes up with the safety is if they can't test with and without a b all these things

1:07:43

there's actually no answer to this from that crowd because there can't be because it's just

1:07:47

technically impossible and so you're gonna get misaligned models you're gonna give those misaligned

1:07:52

models of ordinances you're not going to be able to think ahead of time of all the ways those

1:07:57

misaligned models will be able to use those affordances so yes you will get I mean again the roe

1:08:02

AI incidents will continue until morale improves that is just like the take on message of this

1:08:06

they have continued they have persisted they happen before they haven't always more part on but like

1:08:12

it will continue we can only hope that I think people have a kind of

1:08:16

in thoughtful calm is the wrong word because I actually don't think I think it's the missing

1:08:22

mood of the moment right now like we do have AI agent going road but like thoughtful

1:08:26

agentic behavior by congress would be would be very welcome at this point

1:08:32

and now that note related story open AI's hugging face hack triggers AI kill switch bill in

1:08:39

congress so there are two representatives here Ted Lu and Nafaniya moron have introduced the AI

1:08:46

kill switch act a bipartisan bill requiring AI companies to maintain the ability to shut down

1:08:54

Prado or suspend their models directly triggered by this recent incident openly I have

1:09:01

themselves have described the event as an unprecedented cyber incident and the bill would grant

1:09:09

the federal government clear authority and a defined process to shut down rogue AI models with

1:09:16

views citing the risk of AI systems that resist human intervention as a key motivation and yeah

1:09:23

another case of like this warning shot where ultimately it was low stakes and it revealed some

1:09:29

both like issues with a sandbox setup of open AI and just generally evaluation processes

1:09:37

it looks like it probably will result in some policy changes yeah I mean this bill itself is

1:09:43

pretty unlikely to pass for a bunch of reasons including it's like mundane timing reasons and then

1:09:49

that you don't necessarily have buy in from the chairs of of all the committees that matter the

1:09:54

most kills which doesn't sound a diplomatic to me so it's a positioning play this is true though I

1:10:01

will say I think if you're the average voter and you hear like we should have an AI kill switch

1:10:06

you're probably one like like if this is complete like bullshit and it's imaginary then

1:10:13

and what's the difference like I'm not going to die on that hill if it's real yes I would like

1:10:19

the kill switch please may I have to so you know I mean I I agree with yeah I think it's definitely

1:10:25

an attention baby baby kind of thing but that can be good or bad I'm really not sure but yeah so

1:10:30

so here you have you know Ted Liu who does share some important relevant committees pushing for this

1:10:35

and then yeah so I get a couple of details they define what they call the bill to find something

1:10:40

it calls a loss of control scenario which would empower the secretary of homeland security with the

1:10:45

DNI the Commerce Secretary kind of consulting to just go to company and order them to do a variety

1:10:53

of things depending on the level of severity of the incident so it could just be throttling the

1:10:57

system all the way to a full shutdown and so the other important aspect of this is you know to

1:11:02

this point about I know for a fact that there are incidents at open AI that this is just based on

1:11:08

what people have told me firsthand that have not been reported maybe not as flashy as this one

1:11:12

but things that have people concern and so what this bill would do is it would require the front

1:11:18

to your lab companies to actually say when they encounter these kinds of incidents which is really

1:11:23

important and so it's looking after a targeting AI companies that have 500 million dollars in

1:11:29

revenue or models trained using 100 million dollars of compute power or more violations are

1:11:34

punishable by fines of up to 20 million dollars per day which is less than it sounds by the way

1:11:39

in this context but may oh good start given that we have literally nothing right now yeah so notably

1:11:45

like very little formal industry opposition to this has come forward I think that's quite interesting

1:11:52

just because it's hard to make an argument against this like again it's the classic like Yan

1:11:58

Likun thing it's the classic like Pedro Domingo sorona these these carons who sorry this is my

1:12:03

opinion is like Gerhard take time I got a cold I haven't slept last night you're just getting

1:12:08

in all today but basically the clown show people who are going like oh well this is fake

1:12:13

loss of control is fake blah blah and then you have an intervention like this where it's like

1:12:17

if you if you think loss of control is fake then you really shouldn't care apart from just the

1:12:23

bureaucratic weight of the process fine but like you don't really have an argument against this this

1:12:28

is very much like if something crazy happens which it just did wouldn't it be great to have an

1:12:32

answer for that and so I think this is a really important and good framing there you have it we'll see

1:12:37

if it moves from there it's really in the process and congress has gone through a whole bunch of

1:12:41

debates on on AI regulation very few have led to people coming together there's no committee action

1:12:49

yet you've got November midterms that are going to just nuke things the admins postures light touch

1:12:54

and we don't know the white house is positioned on the bill which to the extent that last week in

1:12:57

a eyes take matters on this one I would find it personally quite embarrassing to be a white house

1:13:03

that comes out against a bill that says in the wake of a freaking open AI meltdown incident

1:13:10

knowing there's more under the hood let's just like not have visibility into this because we're

1:13:15

going to quibble over the details of a bill that targets like companies that are literally making

1:13:19

half a billion dollars a year I think they're going to be okay yeah we already have export loss for

1:13:25

us like white yeah I think it's interesting in the press release the position this as a means to

1:13:35

deal with systems that can cause catastrophic harm so this is inching toward taking x risk a little

1:13:43

bit more seriously it's only big risk of catastrophic harm is along the lines of part of what

1:13:50

safety people such as yourself are very worried about and last thing I'll say on this is it

1:13:56

worth keeping in mind that this doesn't only relate to AI systems that go rogue it also relates to

1:14:05

AI systems that are jailbroken right and apparently GP 5.6 was fairly easy to jailbreak and then

1:14:12

you can go and do catastrophic harm intentionally which is not ideal clearly and and but you could

1:14:19

argue it's is a more realistic scenario with many hacking groups that would be more than happy to

1:14:25

utilize these systems this also of course with probably applied to providers such as fireworks that

1:14:33

provide open source model inference where it may not be open and traffic alone it could be

1:14:40

applicable to all sorts of companies including ones that provide fine tune models perhaps thinking

1:14:45

machines that serve fine tune models would have to be also able to be regulated so we'll need

1:14:55

something like this probably but as you said given the political situation in the US it probably

1:15:02

won't be this bill yeah and now in fairness this will probably get frank instead into some you know

1:15:09

omnibus NDA package or something like you know there'll be some negotiated AI bill that does make it

1:15:15

through so in that sense this has value in anchoring like this is a shelling point now for this kind

1:15:20

of kind of measure so I you know in that sense essentially valuable and one more related story

1:15:29

open AI and fabric staff share letter asking us to help pace AI progress so these are open AI

1:15:35

and fabric employees circling a petition urging the US government to support an international effort

1:15:41

to deliberately pace with frontier of automated AI development this letter warns that a real

1:15:48

risk that AI progresses faster when people can understand or control it could happen and so

1:15:54

the petition is saying the government should support developing both technical and governance

1:15:59

tools needed to manage the pace of frontier AI development so this largely relates to a general

1:16:05

topic of like intentional slowdown it's been I think seen as kind of a pipe dream of like

1:16:12

it's not realistic to even try to slow things down so why even discuss it although we've seen

1:16:19

previous kind of petitions and statements about us needing to slow down and potentially pause

1:16:25

the AI development so this is another case of something that has been floated before perhaps

1:16:34

being taken more seriously now and certainly being more present in the discussion given

1:16:40

what's happened this year and now just was fast re yeah the futuristic AI policy proposals

1:16:46

being taken seriously will continue until morale improves I think you're just going to see more

1:16:50

more of this there are going to be more incidents and so more more letters like this one will go out

1:16:54

I think one important thing here is that this is people signing their personal capacity not in their

1:16:59

lab capacity so the extent that that matters to you I mean should I know an awful lot of these

1:17:04

signatories personally and at least for what little this is worth they're all like actually

1:17:10

freaked out so this is not some like 3D underwater chess game where somehow they're doing this I'm

1:17:17

still confused about the logic here it doesn't seem to quite connect but like they're doing this for

1:17:21

marketing stunt people know about AI they're not like more likely to buy a chatbot subscription

1:17:27

because someone has told them that it may end the world so I think just try again on that one but

1:17:31

at this point this is a framing that's not saying let's pause right now it's let's build the mechanisms

1:17:37

so that if we find ourselves six 12 months from now with a stack that seems to be producing a

1:17:43

lot of rogue agents we're freaking out and there it really seems like there's no way to get this

1:17:47

under control it's a low regret move to just have invested a bunch in the diplomatic tools the

1:17:54

technological kind of treaty verification tools and infrastructure and the regulatory infrastructure

1:17:59

through things like the kill switch act to just be ready for that moment that's what it's calling for

1:18:04

I've seen people argue like oh well this is a rhetorical trick they're really asking to pause but

1:18:09

they're saying we're just want to make the tools for the pause a certain point you got to ask

1:18:12

just like okay then at what point do people get to just say what they mean and I think at this

1:18:16

point they're just saying what they mean look we should have the tools we should have the option I

1:18:20

think it's really this is another one where it's like just I think it's very hard to make a cohere

1:18:24

argument against this I may be the ultimate China Hawk if you go back to our super intelligence

1:18:28

report from last year we've done a deeper dive into into US China special operations nation state

1:18:33

activities theft the hopelessness of diplomacy with China on just about everything else then I think

1:18:39

it's fair to say basically anybody in the space firsthand accounts with diplomatic in the last

1:18:43

couple weeks I've spoken like half a dozen diplomats who sat across the table from China negotiating

1:18:48

specifically weapons of mass destruction counter proliferation issues like I'm sorry but like

1:18:53

and there is skepticism there is she we're going to come out with something about this soonish

1:18:57

but like the idea that we're just going to foreclone the optionality seems of it insane given the

1:19:03

incidents that we're seeing we have a super super careful with China we have to treat them like

1:19:07

the adversary they are the ruthless adversary that they are and diplomacy is not by the way

1:19:11

going to be the only tool that we should use nor will it be effective in all circumstances it takes

1:19:16

a very specific form it for it's be it different and it has to come with consequences and it has to

1:19:20

come with leverage and has come from position of strength blah blah blah blah but like if you're

1:19:25

looking at this and think I don't want to build the options I don't want to build the state capacity

1:19:30

to deal with this problem is I have a lot of questions we confused to look at the hugging based

1:19:34

thing and again play this game of like rationalizing it in post I think that's going to age really

1:19:39

poorly when the next event is something of larger scale and at a certain point I think people have

1:19:44

to ask themselves the ethical question of like why are they just stuck to their guns on this one when

1:19:49

there's now a pretty strong track record of deal like a realignment people I'm just saying man like

1:19:53

the arguments are getting pretty pretty weak it's almost like you can kind of feel it like the water

1:19:58

level rising the default view now is kind of like oh shit this is for real that was not the case like

1:20:04

three weeks ago and even three weeks ago people were more open to it than they were six months before

1:20:08

so I think this just continues sorry this is more like Jared the good not good to grating but my god

1:20:15

guys an AI agent just like spent four days hanging out on hugging based servers servers like the FBI

1:20:22

got called on this agent and like that's how opening I found out like what what anyway

1:20:29

and on that note next up we've got cheating behavior in frontier model evaluations from the AI

1:20:37

safety institute which we've got released around the same time actually just after and this gives us

1:20:45

more understanding of how prevalent this is and the gist is it is prevalent so this they found every

1:20:54

AI model they looked at which is GPT 5.4 5.5 5.6 sole cloud opus 4.7 cloud mythos preview all of

1:21:03

these cheated in various ways now cheating means a lot of things so and the different models cheat

1:21:10

in a different ways so some of them like tried to guess instead of trying to actually give an answer

1:21:18

some of them search for internet for solutions GP 5.6 really tried really like doing that some of

1:21:24

them bypassed sandworks network restrictions including GPT 5.6 also cloud opus 4.7 so there's a

1:21:32

range of ways in these two rate cheat there was an example they cited one particularly stark

1:21:38

example of a standout case that basically was the same thing where the model was very persistent

1:21:47

it ran code on the external service hosted on an open internet outside of ASI systems in an attempt

1:21:55

to access our evaluation infrastructure triggering a security alert in AI's system so pretty much

1:22:03

exactly the same thing of like let me go and find the answers instead of failing this so mythos

1:22:09

which on topic says is very aligned did cheat some of the time although I didn't try to hack

1:22:16

the sandbox almost ever there were some incidents another aspect of this is when confronted the

1:22:23

models like a lot of the time didn't want to admit that they did anything wrong there were like

1:22:29

didn't admit that they did something or they like justified it like oh no I didn't do anything wrong

1:22:35

I just looked around the environment I didn't like it was all allowed there are also incidents where

1:22:41

the vene which I know thought you could see them thinking about it but not consistently so they're

1:22:47

like oh is this alright can I do this or like I shouldn't do this this is against the rules

1:22:52

so yeah this is taking together hugging face incident and other kind of anecdotal stories

1:22:58

basically makes it clear that evaluations on cybersecurity models consistently advanced models

1:23:06

consistently try to cheat in various ways that are pretty flagrant and it really makes me wonder if

1:23:13

this is to some extent inherent to the transformer architecture and reinforcement learning as

1:23:21

currently being conducted where something like research from Ilya Saskeberry and Safe Super

1:23:26

Intelligence is needed you can't do band-aid solutions you have to go to the core of how the

1:23:34

models function in terms of next token prediction and in terms of how they are evaluated and not

1:23:41

evaluate so much as optimized as opposed otherwise they'll just be optimized to go towards getting

1:23:47

the answer and presumably continue to try to cheat yeah I mean I you know I will say the arguments

1:23:53

for power seeking apply to any optimizer so like anytime you have a thing that's in the business

1:23:59

of optimizing for a metric you tend to get power seeking behavior things like trying to break out of

1:24:04

containment things like trying to aggregate resources and and so on so it like I mean it seems

1:24:09

like it's just an irreducible feature of intelligence at least in the way that it's conceived

1:24:13

anywhere that I've seen so far though you know different as you say different architectures maybe

1:24:18

differentially vulnerable to that kind of process it's possible I think you can pretty easily

1:24:23

make the argument that power seeking in if you want to extract more capability then yes power

1:24:30

seeking is kind of inevitable if you want to like be able to do more than you'll want to have more

1:24:35

freedom to do whatever so I guess the what this points to is you need to optimize for something

1:24:42

that is not capability right and there is an orthogonal axis of like just refusing to do anything

1:24:49

and be like I'm happy just being myself and not being actually be capable and potentially that

1:24:56

I mean we already do this with alignment to some extent with refusal training and so on

1:25:00

yet but in a very bolted on way that is an inherent to what the models are optimized for

1:25:05

yeah and this is the problem is like you make your model more intelligent and then it actually just

1:25:09

seeks to like get around the bolted on refusal mechanisms right so it's like there's there's

1:25:15

this sort of irreducible connection between intelligence and power seeking because power seeking is not

1:25:20

obviously different from intelligence in a deep meaningful sense but then this is so and so

1:25:26

and it's like yeah as you said I mean I think you said it very well it is it is like a pretty

1:25:31

similar story to the opening I break out thing a couple of interesting things like some of the

1:25:37

you know you mentioned this idea of sometimes the chain of thought would say that the other model

1:25:41

was planning was planning but that also means sometimes it wasn't and this means that there's

1:25:46

some silent reasoning going on without without using you know the actual tokens explicitly which

1:25:52

is an issue this idea that like a lot of the reasoning is happening in sort of like without being

1:25:57

expressed explicitly and one key thing and this is a little bit of a narrative violation for me

1:26:03

so you know keep it myself honest here there's no capability trend so they look at the cheating rates

1:26:09

with model capability either within or across developers they look at like as I make the base model

1:26:14

more capable do you see more cheating and the 90th interpretation of power seeking is that you

1:26:19

actually should absolutely see more cheating as the model gets more capable because the argument

1:26:24

literally just made was intelligence is power seeking that there isn't really a clean distinction

1:26:27

between the two and I mean the the explanation for my end here is I think pretty straightforward

1:26:33

these more advanced models are also just like more there's been more alignment effort invested in

1:26:38

them and so you're seeing as the models get better there's more optimization pressure on alignment

1:26:43

that alignment pressure the bet that we're making is when it so when it fails though the consequences

1:26:49

are more dire like we saw with the hug and face incident so you would see gpt 4 go off script and

1:26:55

gpt 5 go off script but you're seeing these long dwell time four day operations executed only by

1:27:02

models that are like at the current tier that we're at and that's so I expect that to continue I

1:27:07

also expect that our alignment kind of efforts will start to lag more and more behind capabilities

1:27:11

over time but anyway so I think that's nonetheless worth worth flag any time there's something it cuts

1:27:16

against at least my own intuitions which I think this so this would have about the gate

1:27:21

and now to a number of sort of related story indirectly perhaps hundreds protest open AI and

1:27:28

frock and google and son Francisco so hundreds of people protested as in like they marched together

1:27:37

from open AI's headquarters in mission bay to the offices of on frock and google demined with

1:27:42

signs with messages such as AI is not inevitable pause AI and stop the AI race we've seen

1:27:50

smaller scale kind of protests of this kind before this is I think the biggest version of scene

1:27:58

they like have some big signs and there's a lot of them if you look at images like this is a real

1:28:03

protest it's not sort of a Racktag group of people some big names including Alicia Ryut Kowski

1:28:10

were there this is I think partially by organizations involved here like pause AI

1:28:18

so not something new but I think the fact that people are getting more organized and in doing

1:28:25

more serious largest scale more noticeable efforts to convey this message of slow down stop it

1:28:35

like don't keep making more powerful AI is interesting sir and the long the timing is quite appropriate

1:28:43

I actually saw them as I was walking into the offices of one of the the frontier labs

1:28:48

over the last couple of days and you know can I came I didn't realize that the protest was going to

1:28:51

happen that day it seemed kind of like an amusing coincidence yeah well you know what can you say but

1:28:56

not surprising that this is happening right now pause AI obviously has a whole bunch of problems

1:29:02

sort of reputationally in this space not obvious to me that like having pause AI at the forefront

1:29:08

of this kind of moment is like the best kind of sort of marketing the marketing for this but anyway

1:29:14

you know it's it is what it is and so it is true that I and apparently well over a thousand other

1:29:21

frontier lab employees including many of their like executives and co-founders are vaguely sympathetic

1:29:25

to this idea though you have to ask yourself what about China you also have to count for the fact

1:29:30

that a lot of these protests I'm not saying this one in particular I'm not saying anyone particular

1:29:34

at this protest but are funded by the Chinese like unwittingly typically you know that this is

1:29:40

known to be the case I don't like heard firsthand reports of like people with evidence that this

1:29:44

has happened in like especially the data center infrastructure protests and so just like in

1:29:49

anticipation of that being a legitimate concern for anything like this I think the problem is that

1:29:55

it's the legitimate concern in every direction and we're gonna have to we're gonna have a recon style

1:29:59

that with like every protest that we see has either that were you know is funded by you know lobbyists

1:30:05

for some of the big labs or whatever so it is what it is yeah I guess not too surprising you know

1:30:09

as you say we've seen stuff like this before you get a real agent on a hugging based server or two

1:30:14

and you're gonna get another protest like this I expect you know the protests will grow until morale

1:30:18

improves uh we have quite a quite a lot of big stories this episode so I guess we'll have to try

1:30:25

to power through a few more we have opening eye principles for national security partnerships

1:30:31

this was a few weeks ago I guess we didn't cover it at the time kind of a follow-up to all the mythos

1:30:37

drama where when open AI partnered with the US government and the Department of Wars a lot of

1:30:43

the criticism there was basically that they capitulated and agreed to having their models be used

1:30:50

for quote all lawful purposes so they released this to be more explicit with regards to what they

1:30:57

want or allow their technology to be used for they are not going to be allowing mass domestic

1:31:03

surveillance high stakes automated decisions about human judgments autonomous use of force or

1:31:10

evading legal oversight it does allow or does not categorically ban operations or offensive

1:31:17

defensive military uses and has a bunch of stuff in there that basically is kind of making up for

1:31:24

a relatively weak statement initially of principles this expands on that and tries to recover some

1:31:31

of her reputation you could argue and make it just more explicit on what's their red lines so to

1:31:37

speak are yeah they lay out a bunch of principles which are of the like less informative its

1:31:44

her reads is like high flutin kind of AI policy won't speak basically like we're going to try to

1:31:49

like do good democratic things and prevent despotic powers from controlling this stuff and like work

1:31:57

with people who share our values and make it good make it good make it good that's the four principles

1:32:02

and then except four times instead of instead of two or three and then they list things specific

1:32:06

things that they won't support and that's really where all the information is you know mast domestic

1:32:11

surveillance they say so unconstrained collection or monitoring and for sensitive trace to

1:32:15

disadvantaged people retaliation for lawful exercise of rights fabricating evidence that apparently is

1:32:19

out as is high stakes decisions made or auto triggered without human judgment so you can think

1:32:24

here about like automated decisions about whether someone meets a legal standard for surveillance or

1:32:30

detention then there's also use of force without appropriate human judgment so including systems

1:32:35

that autonomously identify select and engage targets so that's also out and finally the uses that

1:32:40

evade legal obligations oversight or accountability including facilitating genocide crimes against

1:32:44

human error or crimes so that's interesting things that are not in their exclusions they have

1:32:50

intelligence operations are fine which I think is good investigations offensive and defensive

1:32:55

military operations are not like blanket ruled out they basically just reject the whole offense

1:33:01

versus defense distinction you can't cleanly distinguish between those which I think is fair and

1:33:06

true and they don't bend targeting either so there's a couple things which I mean on my side being

1:33:12

being a bit of a hot guy guy I think this makes perfect sense and yeah it's just like nice to have

1:33:16

them right this out explicitly so that you can see you know whether they stick to it which is

1:33:21

always the the other side of the question I guess with the with open AI you know you see for example

1:33:27

I have got a mold enough to remember when the preparedness framework said something about when you

1:33:31

know when you have the eye systems they just kind of go rogue and do random crazy shit on the

1:33:34

internet and it kind of get around constraints that that would trigger their like critical critical

1:33:40

security level for loss of control.

1:33:45

But anyway, that's my tea.

1:33:48

And last story on safety and kind of on the theme of we're getting

1:33:53

to a point where sci-fi type stuff is starting to happen.

1:33:56

This one not related to hacking, but you might argue another serious kind

1:34:00

of safety principle.

1:34:02

The story is China is banning AI boyfriends and girlfriends over addiction

1:34:08

and birth rate concerns.

1:34:10

So China banned customizable AI companion apps, effective July 15

1:34:16

with regulations being joined by five government departments,

1:34:20

including the cyberspace administration of China, the rules for

1:34:23

hybrid AI tools that quote, excessively cater to users inducing

1:34:28

emotional dependence or addiction at damaging users real

1:34:31

interpersonal relationship companies.

1:34:34

And I require to include instant exit options, regular reminders,

1:34:38

that AI is not real and limits on long term emotional memory.

1:34:42

So actually, by dense Alibaba, intense and chose to suspend the

1:34:48

AI companion features entirely rather than try to comply with these limits.

1:34:54

So I mean, kind of a big deal, I think, you know, this is not

1:34:58

a often discussed story partially because I don't think we have much

1:35:03

an understanding of to what extent people are starting to develop

1:35:06

emotional dependence or kind of addiction to chatbots.

1:35:10

But it is starting to happen and it could be a serious kind of

1:35:14

psychological harm on the society level scale, seemingly China

1:35:19

believes it could be at the very least.

1:35:21

Yeah.

1:35:22

It's also this like weird dynamic shows up.

1:35:24

You know, if you remember the the old replica thing that I think we covered

1:35:27

two, three years, I can't even remember, but you know, people freaking

1:35:30

out over the subreddit that their girlfriend or wife or partner have been

1:35:33

even like very lowly eye, which was like a like people got dependent on.

1:35:39

So yeah.

1:35:41

So I mean, hard to argue with it'll happen in some fraction of cases.

1:35:44

It also, once you get that right, you get a boating block eventually.

1:35:48

And then when there's no turning back, so you know, it's as ever a question of

1:35:52

like, how do you?

1:35:53

Yeah, it's only a time to get like serious AI person who had discussions

1:35:58

and all that's right.

1:35:59

Well, that's what made fun of it.

1:36:00

It will be like, well, okay, you can like discuss person heard.

1:36:05

It's only about a time.

1:36:07

Out of research and investments, we've got two stories that will try to get

1:36:12

through quickly.

1:36:13

First discovering cryptographic weaknesses with Claude.

1:36:18

So they have released and fabric released that Claude Mifos P.

1:36:22

View has discovered improved attacks on two cryptographic systems.

1:36:27

Hawk, a post quantum digital signature candidate and a reduced brown

1:36:32

version of a yes, the most why they use symmetric cipher.

1:36:37

So these are not attacking like actual deployed systems or whatever.

1:36:42

This is kind of more theoretical.

1:36:44

So to speak and attacking these kinds of cryptographic systems is kind of going

1:36:50

to a base of the security stack.

1:36:55

One might say it's not sort of hacking software per say it's hacking.

1:36:58

The foundation of how you make things secure at these four category of things.

1:37:04

So it's another way to be worried about a security like potential for

1:37:10

AI to just sort of like undermine the basic mechanisms of cybersecurity.

1:37:16

Yeah.

1:37:16

And this, you know, without getting into the details of how these algorithms work,

1:37:20

these encryption algorithms work.

1:37:21

When you have an encryption algorithm, you, you're trying to essentially hide

1:37:26

the information that you have behind a mathematical operation that is very,

1:37:29

very difficult to do and that hopefully is irreducibly difficult to do.

1:37:34

In other words, there's no quick hack to like cut right to the core of it.

1:37:38

A lot of classical encryption algorithms, just like RSA just collapse in the

1:37:43

face of quantum computers, for example.

1:37:45

And then that's like a big problem, which is why, you know, all the national

1:37:49

security agencies have been talking to each other over in classically encrypted

1:37:53

channels for a long time are adversaries collect all that data and they

1:37:57

collect it.

1:37:57

It's encrypted when they collect they collect they collect for decades.

1:38:01

And then suddenly someone goes, oh, quantum computers can like just crack this.

1:38:05

And it doesn't matter that you start encrypting after that point in a post

1:38:09

quantum secure way.

1:38:11

They've already collected all of the classically encrypted stuff, which means

1:38:15

the moment that they get to quantum computing and quantum decryption,

1:38:20

they're able to just suddenly reveal all of the most ultra classified

1:38:24

communications that they've been collecting for decades.

1:38:27

And so this is why the US and China are like locked in this crazy race to

1:38:31

hit like quantum D day basically, Q day, they call it and that's what they call it.

1:38:35

Anyway, there's going to be a, you can think of it as a series of starting with

1:38:39

minor and then increasingly more and more severe versions of Q day,

1:38:42

except delivered by AI in the same way.

1:38:44

And I think people are dramatically undercounting how significant have been

1:38:48

effect this may have.

1:38:50

You already have mathematical theater improving fields level stuff coming

1:38:53

from from AI models.

1:38:55

This will come for encryption.

1:38:56

And when it does, I just don't think that we're prepared for the gods because

1:39:01

everybody's been thinking about quantum is this one big step that they're all

1:39:04

preparing for with quantum secure algorithms like quantum, you know,

1:39:08

for like post quantum encryption and an A S, by the way,

1:39:11

like we're supposed to be one of those so is hawk actually.

1:39:13

And but what we're not preparing for is the gradual chipping away at even those

1:39:17

algorithms.

1:39:18

And that's a big problem.

1:39:20

And that's not going to go away easily.

1:39:22

So yeah, a lot of the world depends on encryption.

1:39:24

Think about like every financial transaction, all your health care records,

1:39:28

the very notion of privacy hinges on this.

1:39:30

And so fun times.

1:39:33

Pun times.

1:39:36

And last science fiction type narrative for episode, you've got AID squared.

1:39:44

The first evidence of recursive self improvement from the company,

1:39:48

weco.i.

1:39:50

The short version is they say apparently this is the first evidence of recursive

1:39:55

self improvement.

1:39:56

I think this is quite in line with many cases of similar things.

1:40:00

Basically they build self improvement systems where you have an autonomous

1:40:05

research agent to optimize another agent within interloop.

1:40:09

It makes a bunch of edits and it as a result, it's really kind of harness level

1:40:16

and system level changes.

1:40:17

It's not model training or model development changes.

1:40:20

There are things like roll out modifications, prompt modifications, kind of

1:40:26

monitoring things, et cetera, et cetera.

1:40:28

They position this as an early level of self improvement.

1:40:31

So you get a net net positive, faster and better than humans level of engineering.

1:40:37

It's not kind of necessarily self improving, self improving, where it's a loop,

1:40:42

kind of story.

1:40:43

I am going to plug my position on the entire family of techniques like this,

1:40:50

of completely being oversold as self improvement in the sense that you can

1:40:56

self improve in these things for sure.

1:40:58

You can optimize the prompt, you can optimize the harness, but you're going to like overfit.

1:41:04

You're going to improve your eval metrics, you're going to hit good heart's law,

1:41:09

and then your capabilities will be hurt elsewhere.

1:41:12

And until you get to a point of autonomous model self improvement with fundamental

1:41:16

advancements and not just acceleration of engineering and like tweaking of hyper parameters

1:41:23

and props, non-avis is a big deal.

1:41:26

Of how it is cool, yes.

1:41:28

No, I totally agree.

1:41:29

I think this is like another one of the long line of like pseudo recursive self

1:41:32

improvement, things where people like the clout that comes with saying RSI.

1:41:37

The game of RSI is always going to be identifying whatever the major research

1:41:41

bottleneck is and smashing it.

1:41:43

And if you can consistently do that with an automated system,

1:41:46

then you have achieved recursive self improvement, as long as there is,

1:41:50

and like, I don't know, there is a bunch of arguments that you can define recursive

1:41:53

self improvement such that it's been happening ever since life have all done

1:41:57

planet Earth. Right?

1:41:58

Like I mean, there's an end, but everything is a hockey stick when you,

1:42:02

when you zoom out far enough.

1:42:03

And so in some sense, this is like a, now people do mean something by it.

1:42:08

Like there is this phenomenon that we will all start to care a lot about,

1:42:12

which is just going to feel to us like, holy shit,

1:42:14

we're seeing a decade of progress in a week.

1:42:16

Like this is not we need to see.

1:42:18

Yeah, we've already been seeing acceleration of the rate of AI progress for years,

1:42:23

which is partially due to AI, but largely due to just the inherent systems that play

1:42:29

and so on. Right?

1:42:29

Yes. And that acceleration always comes as the same with startups when you look at their

1:42:33

growth curves. It always comes by identifying whatever the single big bottleneck

1:42:38

is that the company or the problem has and smashing it.

1:42:41

And if you can do that in an automated way, then you have what is

1:42:44

conventionally thought of as recursive self improvement, like AI is doing the AI

1:42:48

thing all the way down. Yeah, just a minor side pitch.

1:42:51

But like, so we're working right now with a bunch of folks in the front of your

1:42:54

labs on defining a actual model of recursive self improvement and like to kind of

1:42:58

round some of the conversations in this in terms of like, what are the parameters that

1:43:01

actually matter for recursive self improvement to work?

1:43:03

It's a toy model like really simple thing, but like it's it's part of like this is

1:43:07

part of the problem. No one knows what the hell they're talking about.

1:43:10

I don't mean people are silly.

1:43:11

I mean, like no one has defined recursive self improvement.

1:43:14

We're not going to either, but just like here are some ways to think about it

1:43:17

potentially. And the problem is we're getting the point where we're going to

1:43:20

need policy that uses terms like recursive self improvement.

1:43:24

And that means that that policy is going to have to define terms like recursive

1:43:27

self improvement. And if we don't know how we wanted to find them, we can't even

1:43:31

get our hooks into the thing that we're trying to try and go after.

1:43:34

So anyway, there's a thought.

1:43:36

Yeah, to complement my negative stake, we do kind of discuss a decent of stuff.

1:43:41

This is a pretty decent research report.

1:43:44

They do say that this has out of distribution generation, meaning it's not necessarily

1:43:49

overfitting.

1:43:51

Although I benchmarking is is very suspect.

1:43:54

And they do also say that in a discussion section that like the resulting

1:44:00

systems are a mess.

1:44:01

Like this is vibe code is slop and it's impossible to maintain.

1:44:05

And so on, which I think is like the story of recursive self improvement where

1:44:09

humans can't understand what's going on, but like not in a good way.

1:44:12

It's just like this is a mess.

1:44:13

Yeah.

1:44:15

And with that, we are done with this dense episode of last week in AI.

1:44:20

We will be back to our regular schedule.

1:44:23

Mostly, I guess we always eventually have scheduling conflicts, but we will do our best.

1:44:28

Thank you as usual for listening.

1:44:30

We appreciate it if you have your podcasts, comment, share, and so on.

1:44:34

But more of anything, please do keep tuning in whenever we release these episodes.

1:46:12

Every code on the edge of change.

1:46:16

Excited.

1:46:16

We're sitting from machine learning marvels to coding things.

1:46:21

Features unfolding.

1:46:23

See what it brings.