#253 - Opus 5, Gemini 3.6, Kimi K3, Hugging Face Hack
2026-08-03 05:00:00 • 1:43:21
Hello and welcome to the last week an AI podcast we can hear chat about what's going on
with AI as usual in the subsequent you will summarize and discuss some of last week's
most interesting AI news also the week before we have unfortunately skipped a week due
to scheduling conflicts but we will cover everything relevant from the period I am one of
your regular hosts Andre Khrankov I studied AI in grad school and now work at the startup
AstroKade and everybody what's up my name is Jeremy of course I'm your other ghost I'm
from Glatz and AI do AI national security super intelligence see type end of the world
stuff so I sound a little thick right now by the way which is related to the reason that
we didn't record that the so last week which is that I was traveling and I got sick on
the flight was guy who was coughing up along next to me and anyway that's why I sound
so weird right now but trip was really useful and yeah hopefully be able to talk about
a lot of this stuff soon but a lot of conversations with like researchers at the frontier labs
and folks on the you know the safety teams the capability teams all that kind of thing
that I think bears quite a bit on the events of last week and the week four will definitely
be taught about a lot of that stuff with some of the inside you a little bit what I can
share right now on those things but man things are moving but and it has been a slightly
eventful two weeks I mean I guess it's not the most eventful you've had this here but there's
been some big stuff it will be touching on as a quick preview there's a few new models nothing
gigantic but fairly meaningful you'll start with then as usual some funding stories and deals
about compute and so on some major open source releases including Kimi K3 we have the
discussed about so we'll be talking about that then policy safety of course we'll be talking
about the recent hacking incident from open AI and a whole bunch of other stuff related to that
it's gonna be a kind of policy safety heavy episode and then we'll round it out with some research
and synthetic media and art so it'll be a packed episode we'd like to thank notion for being
a sponsor agents are getting smarter every day but even the smartest agents get stuck without
the right context and the right tools that's where notion comes in with a recent launch of custom
agents notion became the collaborative AI workspace where teams and agents work side by side
and now their new developer platform is turning that workspace into infrastructure developers can
build on notions developer platform gives developers and coding agents the primitives to
extend what's possible a notion and pick it beyond connect to external systems bring context in
take permissioned actions across your tool stack and expose custom agents capabilities to any
system that needs them these primitives include a CLI workers on notion hosted sandboxes and
external agents API and then agent SDK trigger notion agents from any app learn more about notions
developer platform today at notion dot com slash LWAI that's all lowercase letters notion dot com slash
LWAI to try notions developer platform today and when you use our link you're supporting the show
notion dot com slash LWAI this next sponsor isn't regulatory I but I've personally used them for
years so I'm happy to have their support and it is factor they make chef crafted the dietitian
design ready to eat meals so you don't have to choose between real food and convenience both in
grad school and as a startup employee I don't have a ton of time so when I get home I'm tired
and being able to prepare really quite a good meal without any effort has been fantastic their
meals are ready in two minutes and require no prep and no cleanup so even on the days your schedule
is completely out of control eating well is still achievable they are over 175 band ingredients so
every factor meal is designed around what supports a healthier lifestyle and nothing that doesn't
and that's with over 100 nutrient vents menu items to choose from every single week
97% of users agree that factor meals help them live a healthier life so you can feel confident
that you're doing something good for yourself with every meal I've really enjoyed factor and if
this sounds good to you maybe you should try it as well let's eat real head to factor meals dot com slash
LWAI 50 off and use code LWAI 50 off to get 50% off and one free breakfast item per box for one year
while supplies last until double 31st 2026 that's code LWAI 50 off at factor meals dot com LWAI 50 off
at factor meals dot com see website for more details and we'll go ahead and get into it starting with
tools and apps and here we begin with ontropic releasing cloud opus five which they say comes close
to the capabilities of cloud a fable five in many domains and is cheaper of course so this is
following up on the release of fable five a little while ago fable being their new family of models
that they didn't have before and they also released son at five either before around the same time
as opus five so we sort of caught up presumably because these are distillations of fable and mythos
right so typically what you can predict with ontropic is their big kind of best model is their
most compute heavy most impressive model which is mythos right now and these things like fable
opus on it are kind of derived from it to some extent where they try to extract out the intelligence
at a lower price so I think not a ton to say about this one beyond that it's supposedly
quite great and close to fable five so pretty big jumps in the benchmarks relative to opus 4.8
the vibe check has been a bit mixed as far as I've seen people you know have a usual sort of
complaints about what the models are doing and it's hard to know whether we just have high
expectations now or in fact the models that are getting stupid there's also a new fast mode
in research preview which offers you higher speeds at double the price amenon I think we're
losing track of what it is that we're looking for just because that with the waterline is rising
so fast people are going to use to incredible levels capability you're right this is probably a
distillate of fable five or mythos five then added safety and all the stuff there's obviously
additionally post training that gets done to kind of further refine the character of the model after
that and so one of the key things that they highlight here is the differentiator of opus five is
supposed to be more emphasis on verification and judgment so less kind of raw capability but more
assertive like double checking its work making sure that what you're getting is actually correct and
so they give this example where like given it it's given a drawing of the machine part but no
way to view the original image and then opus five like rewrites its own computer vision pipeline
to extract the geometry from the raw pixels and reconstruct the part basically the idea being like
it's going to get you that raw data the original to base its conclusion on so that it knows it's
right no matter what that's kind of like the vibe here also an alignment and safety kind of
interesting this is always the game we live in a world where the u.s. government decided to
snap a chalk line at mythos level and anything above mythos level magically is is subject to
are the de facto licensing regime that we have in the u.s. and so in this case and thropic is
in a hurry to say that their model is close to mythos five that identifying software vulnerabilities
but it's less successful at developing exploit again that's part of the post training that's part
of making sure that or and also just as well the pre-training or other parts of the training
process where they're avoiding explicitly training it on cyber tasks in that way and so the goal
here is really to position it as this is a really intelligent model that it is okay for us to release
of course fable five is okay but it has additional safety of its over over mythos but you know
you're always now going to see that kind of background concern of this threshold which is I mean
it's good it would be just great to have a more principled approach to to the stuff one thing to
note to its performance on frontier bench so frontier benches this new benchmark that we have
was released just a few days ago and the same team behind terminal bench came out with it
this is big community effort and it is basically just like a harder agenteic environment we keep
meeting more and more difficult agenteic evals to be like all right but you know we just saturated
your terminal bench or whatever let's now move on to your terminal bench two terminal bench three
and ultimately this and and there you go so here we do know that this particular model opus five
is outperforming all other models on a cost per task basis and that's really where they're
trying to differentiate it is cost per task not necessarily frontier of intelligence near the
frontier but cheaper per token or per unit of intelligence let's say at that point it is
slightly cheaper than gbt 5.6 sole the biggest and best model from open AI and I think
maybe indicative of like pricing becoming more of a concern for customers at the business
front now that there is a competitor to on froth with codex and open AI being quite capable and
very very cheap alternatives for model usage after wonder whether we're going to be seeing more
kind of pricing pressure going on next up some more model releases this time from google google
deep mine has released three new AI models Gemini 3.6 flash Gemini 3.5 flashlight and Gemini 3.5
cyber so per the flash aspect these are cheaper and faster than let's say more intelligent models
to be 3.5 flash fly delivers 350 tokens per second at a rubber cheap price Gemini 3.6 flash
is priced 1.5 dollars per million input tokens and 7.5 per million output tokens so that's
slightly cheaper than sound at five and not super cheap and then there is of course the cyber security
focus Gemini 3.5 flash cyber which is integrated into this code vendor agent that autonomously builds
exploit code to verify vulnerabilities in sandbox environments and then generate patches which
they say has found problems in complex real pieces of software such as the V8 java strip engine
so i think interesting to see a general movement towards cyber focus with not just the mythos
and opus and so on we've seen opening i release a cyber model now google has released a cyber
model and we'll be discussing microsoft has also released a cyber model so if everyone is like
oh no we got to do something about this i don't think it's sharing too much to say like
people in the national security space and the frontier labs are really concerned about where
cyber is going and this view that you know we're in the volumpto apocalypse right now right we're
getting all these low-hanging fruit vulnerabilities being discovered and exploited assume we're going
to get a mythos class open source models sometime in the next you know certainly six months maybe a
bit less at that point you're going to need an answer and so all the lives are pre-positioning
for the moment when basically they're holding the world for ransom i mean you know you need to use
really really good cyber models shorter infrastructure or else like that is just going to be the case
as in the side i'll just like casually drop the prediction here that we may see some pretty significant
disruptive cyber attacks at massive scale not even just nation state or their proxies but literally
just like disaffected young people or you know terrorist groups or whatever that's just what happens
when you open source that level of capability it's just like that's what the math says well see where
that goes but that's like the default assumption right now a lot of the people in the space both
on the national security and the the frontier lap side so that's part of what the positioning is
is here if you don't have an answer to the cyber question you know you're going to be a lot less
relevant in the next six months or so they do this pretty impressive i mean the flash cyber model
which is maybe the one at least the one i'm paying most attention to is performing on par with a
lot of frontier agents that things like the cyber gem and that's an important benchmark i mean it's
a lot cheaper too right a really a fraction of the cost and so this idea of how cyber plays out
is always a really strong function of how much compute you have your defender you have a certain
pile of test on compute the attacker has a certain pile of test on compute can you invest more
test on compute than an attacker shore up your infrastructure is the is the question i mean there
there are a lot of ways to answer it and you know different ways to use test on compute and
a question about how much leverage like maybe there's an attacker advantage or defender advantage
these are all open questions but it's going to come down in some way shape or form to that balance
and so the cheaper you can make these models the cheaper you can make the tokens per unit of cyber
intelligence really the more value you're getting there so that's an area where the cost of the
tokens really really matters and that's why they're dabbling there yeah the other launches are
interesting but kind of fall into this general category of like google still not having a true
frontier model like when we're thinking about the best models in the world it's anthropic and
it's open AI and there's just not really anyone else yeah jimmy pro free free point one used to be
sort of in that phrase or at least yeah never frontier it there's not been a pro level model since
february so they're now quite a bit behind nobody really using jimmy pro for like cv as hard work
and then coding for instance and i think it is an interesting indication where google is at
if they are focusing on flash first because this immediately rolled out to everything all their
products google a studio android studio jimmy app jenemy enterprise agent platform so it kind of
makes sense from a business perspective like they integrate jimmy into everything including google
docs and spreadsheets and AI mode and at that point you have to have a faster and cheaper model which
is why they are really emphasizing flash first they did say that jimmy fp.5 pro is being currently
tested and will be made available when ready and they have begun the most ambitious pre-training run
yet for jimmy four so we've gone indications that they're at least working on a mythos level model
and we've seen them kind of catch up before with jimmy nice so i'm personally looking forward to
what jimmy four will be like next up another new model but this time not for language for images
and videos black forest labs has launched flux free that is capable of generating images and 22nd
videos with audio so this is a multi model frontier model trying to understand and generate images
and these models while extending the architecture interestingly to robotic vision and action so
it's jointly trained across image video and audio modalities altogether it's their first public
video generation model from black forest labs which for some background hails back to some of
a talent from stable diffusion that made them first really impressive image generation models
and flux has still been kind of one of the go-to image generation models at frontier flux
free video looks to be pretty impressive from what i've seen they have some kind of human preference
studies where they say it is preferred over a rock imagined video clinging v3 pro runway gen all at
like 70 percent 60 percent 80 percent whatever people prefer this in terms of its outputs and these
are from just kind of testing it's still not fully rolled out so very interesting and the fact that
it's now being adopted through flux mimic that is being developed with mimic robotics for actually
making it kind of an action model which i've seen kind of starting to be the case more we've seen
some other players like runway starting to get into like the physical intelligence space
video and robotic control seem to have a lot in common yeah this is interesting announcement this
kind of a couple things one is a bunch of comparisons that like show pretty lopsided wins against
you know like luma and runway and all the stuff that don't really matter because nobody uses those
models anymore but there's this interesting comparison against koo will's gemini omni flash
52 percent win rate against that that's actually quite interesting like that's pretty impressive
especially given the resources google's been throwing at this stuff and then another piece is so
so yes like we keep pushing out the length of clips so that you know that's great but the challenge
has sort of become coherence across clips across different shots and that's the big thing that
they're pushing here are some multi shot sequences where the characters are consistent the sort of
the physics is consistent and that's a big boon with this particular police so you know increasingly
moving beyond and what once you get those 20 you know 22nd clips 30 second clips you can imagine
that being a point where yeah you know one shot typically only lasts about that long as if I
know how long a shot lasts in professional film but whatever you know you can imagine that being the
case and so then you know maybe you care more about the switches between different frames so yeah
kind of interesting and a new kind of metric to track and they do have as before variance of this
that are open weight that also have an access to multi model generation called flux free dev so
another kind of slightly big deal I don't think we have an open weight model that has this new
unified multi model background which by the way is relatively new we've seen image
generation and video generation for a while but similar to kind of nano banana from last year I
think we're moving towards a place in video and audio generation where everything is put together
instead of being coupled together and that is actually a pre-book deal in terms of the capabilities
next more of a product story meta is making it's a chatbot more like an assistant so they're adding
productivity features there's a calendar integration daily briefings and apparently in-depth research
capabilities powered by alama spark 1.1 it can also browse Facebook marketplace search for restaurants
check your calendar and handle recurring tasks so this is rolling out to the meta AI app and is
going to be coming out to WhatsApp as well which continues to mark a shift for meta which is like
you're making AI are you just going to compete with all the other AI players like what what are you
going to be doing of this I don't know yeah I mean I think this is partly a realization as well that
unless you're moving in the direction of productivity you're just not going to squeeze all the juice
out of these these models that you can right so you know think about the positioning of open AI
relative to anthropic and the profit per token that anthropic is able to break in because of their
commercial focus it just I mean they're they're eating open AI's lunch and so think about the
there's an extreme beyond open AI we often think of open AI is the direct to consumer company which
isn't as true as it was six months ago certainly they've been making a lot of inroads in B2B
but at the far end of the spectrum in the other direction is meta that they are straight consumer
right like those tokens are going just to tickle your limbic system they're not actually going to
like actually move big things in the real world they're like make products and so if you want to
ultimately get the the most bang for your buck generate tokens that are actually valuable enough
to make you good profit yet have to move into this direction not least to say if you have a
coherent long term view where super intelligence goes humans just aren't in the picture which means
if you were optimizing for the value of the attention of human beings which is what meta is currently
doing that value may drop precipitously as AI is start to control more and more of the economy and
so you have to be in a position to actually do productive work and support agents in doing that
so you know depending on how far you wanted to read this you might read it that far I know that
doesn't really seem to understand super intelligence but certainly Alex Wang does so wouldn't be
surprising if this was at least part of the thinking here yeah I will say I think it will be
interesting to see where they go of this because you can go two ways you can sort of go and and try
to make codecs or a co-work competitor which is straight up just for work or they could shift into
sort of open claw type thing where this isn't always on the background agent which can do a bunch
of stuff for you including productivity things like briefings on your calendar but also messaging
and various things like that and I think the open-class base is sort of still up for grabs Google
hasn't rolled out their open-claw always online agent they've said that they are going to I forget
what it's called so I could see them like being potentially capable of competing on that front
not in the like coding or real you know office productivity side but like personal productivity
you know everyday productivity maybe and one last product rollout open AI is rolling out
charge-upity health to everyone so this is available to all US users aged 18 plus on web and iOS
and it will allow you to connect medical records and health tracking data for the chatbot so
they are saying that this model can reason at levels better than clinician level and you can
connect a whole bunch of stuff I think this marks a shift where you know in the past if you were to
talk about health stuff of these models they were very strongly caveat that you need to double check
and in general you should not have trusted these models with any sort of critical health concerns
judge of the health potentially is at least open air making the case that this is something you can
rely on after applications and business we begin with ilya sascovars safe superintelligence
partners within video to scale it AI research so that's kind of a gist of it we have had SSI
safe superintelligence be around for a couple years they raised one billion in founding in 2024
and a two billion in 2025 so you know a lot of money but not that much money if you are saying
you want to create superintelligence you compare that to an fronpic open AI we have hundreds of
billions this is a few billion so the narrative around this is that ilya sascovars company has
achieved sufficient research progress that it's time to scale up and now to scale up you need a
bunch of compute and so they're gonna be partnering with and video in a value of like some amount
of billions a bunch of billions and they'll be increasing their compute by an order of magnitude
yeah there's some some disagreement between different outlets about how much exactly has been raised
whether it's a five billion round or just in the billions or something like that i think tech
cruncher and five billion dollar story so either way the one thing everybody seems to agree is
number one it's gonna give safe superintelligence access to the your Rubin platform right so that's
the next generation platform five billion if you do the clungs law more is law analysis roughly
allows you to 10x your compute relative to the one billion dollars that that they'd raised previously
and so well there you go they're they're 10xing their compute a couple of things are interesting
about this so yes there's this this narrative that they're like and i agree with this most likely
this is what's happening it is a pretty straightforward guy you know if they say that they've got in the
point where they're at that next level of sort of proof points that they can take this investment
it's worth scaling it probably is there's this kind of more cynical take that like oh they just ran
out of compute which you can hold that view that's totally legitimate i suspect that's not the case
but just like so everyone's tracking that is another another explanation they have no product they
intend to launch no product which means their only revenue is going to come in the form of these
sorts of investments it's a weird sort of moment and story for them because they are you know
the Daniel gross with the co-founder of safe superintelligence along with ilya back in the day he left
for meta after meta offered to buy the whole company will clock and ilya said no so Daniel jump
ship at least at that moment you can argue that that meant at least annual gross thought that
his chances of making something like super intelligence were higher at meta than by remaining at safe
super intelligence what's happened in the interim we don't know and ilya has dropped only the
faintest of hints on dorkesha's podcast about generally you know generally sketching that continual
learning is going to be part of it and going back to i've heard a couple rumors but like i haven't had
any of these verified that anyway they are looking for let's say somewhat beyond the standard
how it's going to save you on the standard model it's a very physicist joke but i have you know
things that are a little further afield and so that sounds like it would almost have to be true
just because the otherwise you're in pure scaling mode there there could be this narrative you
can imagine people sort of like laughing about this and say oh well ilya said the ear of scaling
is over what's he doing raising five billion dollars to ten x his compute and to that i say
ilya never said that you wouldn't also need scale it's both all right the what he's saying is
there is leverage to original research again and they're the biggest leverage is not purely in the
engineering of more and more scale systems it's in something else like you can compound it very
effectively now with algorithmic insight so do with that what you will this is an interesting story
and we don't know much about it yeah it's an interesting story in a sense that you can be very
curious about what they figured out and no nothing because we still have nothing to go on it's
kind of funny if you go to their website and go to the updates page it's like free things it's
literally since 2024 they've released publicly two updates which are just about the co-founder leaving
and now this partnership so hopefully we'll get some more understanding of what they're doing
soon as they scale up but it will presumably be a while since scaling up is not easy and in fact
Nvidia had said that they invested after quotes obtaining rare access to the company's closely
guarded research so supposedly they're tracking who knows right but there you have now
and a related story about five billion dollars AMD has committed up to five billion dollars to
on frothic viz is a new partnership on frothic will deploy up to two gigawatts of AMD's instinct
mi450 aijb use and their new helios wreck scale system plan for deployment in 2027 on frothic has
so many partnerships now with so many they have like SpaceX AI they have Google they have amazon
and now we have AMD I feel like we're just like go to everyone to be like we need compute let's
partner up and give us some compute and AMD it has been trying to compete harder with these
AMD instinct chips honestly I don't recall where they are at with that but they do seem to
at least potentially have the ability to compete with Nvidia which no one else really does
right aside from tpu's from google and potentially the hardware that some of these companies are
developing yeah and this by the way this idea of anthropic having like a million different partners
it I mean it's really true right but partnerships with google for tpu's partnerships with Nvidia
partnerships with amazon right partnerships with the SpaceX AI and now with AMD the you know the
the golden rule if you're ever trying to explain well why a frontier lab is developing a new
partnership of cultivating a new partnership again commoditize your complement that's everything
that's going on in the space right now you know we talk about this a lot on podcasts but like the
history lesson here is Microsoft back in the in the day realize that laptops are super expensive
and software is cheap well actually if we just make all the the laptop manufacturers compete with
each other and we make one set of windows software that goes across everything and make basically
the software the choke point and the value chain then suddenly we can make all the hardware vendors
compete away their margins make laptops super cheap now that laptops are super cheap consumer is
dive into the market and obviously they've got to have software run laptops will come to Microsoft
right so everyone's constantly trying to make their complement the compliments to their offerings
compete with each other anthropic wants all of the GPU design firms any company that makes GPUs
it makes compute they want them to compete with each other like crazy so when it looks like they're
setting up a partnership that's law-opsided and like they only have the media GPUs in videos got
ton of a ton of leverage in that relationship now well and throttpicks can go off to AMD or to
Google or SpaceX AI and say hey we want to get our compute from you instead and so now in the
ago oh no no like we'll give you just kind of this how how the pricing control gets set up in the
space and likewise in video and all these players are trying to do the same in reverse right in video
wants to help small baby front your lab come big adult front your life but they have more customers
and also so that anthropic feels more pressure to buy more compute so it's kind of happening in both
directions as everyone's kind of pulling their knives out and dancing around each other it's a
wild time in the space but AMD has been a laggard in this space you know you think about basically
their their software stack the competition to kuda which is just in videos widely viewed as like
in videos big moat for AMD their equivalent is called rock M and a big part of the purpose of
this agreement is that clawed is going to tune workloads for instinct GPUs which the AMD GPUs
and accelerate rock M development that's key right we saw that with with Amazon it's not a coincidence
that every time anthropic signs one of these big compute deals with a hyper scalar that the deal
it involves anthropic has to tune their workloads for that hardware they must use a minimum amount
of that hardware this is about giving feedback to the hardware designer that is so so valuable
because otherwise you just can't you can't design your GPUs for the next generation if you don't
know what the next generation of architectures going to look like that's a huge part of this
you know negotiating leverage we talked about oh yeah this is also the first so helios so this is
all part of not just the instinct MI 450 series GPUs but it's also about AMD's helios rack scale
systems that's the first full rack scale system that they're selling that you can you I mean you
can think of this as like the equivalent to the you know the nbl 72 the sort of full rack that
Nvidia will ship this you have something you can literally just like plug in a data center instead
of just shipping the GPUs themselves that's you know AMD's going further up the stack to own more
and more of that that infrastructure layer and so it's also got 72 GPUs per rack which is
amusingly like the same footprint as the nbl 72 but with completely different kind of power
consumption profiles and stuff like that's anyway super important I mean they are trying to prop up AMD
because they want AMD to be a viable alternative five billion dollars does buy the mistaken
in for a big pre IPO I guess but only just seems like more of a strategic partnership than anything
yeah it's kind of if you look at the press release it's like on for a big we'll be deploying
these chips companies will collaborate to use cloud to optimize workloads for AMD instant
GPUs and AMD will broadly adopt cloud across its engineering and product development teams
so it's you know AMD is going to invest in on frantic but really the point here is we're going
to work together to kind of have a win-win type situation and to your point I don't know if it's
necessarily about the Microsoft type story me that's part of it but also it's about redundancy
and scale for frantic they did get into a nasty situation earlier this year where their sole
provider or their primary provider was Amazon for a few years we didn't have their own compute
and they sort of were left unable to deliver enough compute to their customers and speaking of that
we have another related story that meta apparently isn't talks to least computing power to on frantic
in a potential 10 billion dollars deal so we covered this a think last episode where meta might be
going into the neo cloud business they've built so many data centers that it potentially make
sense to be like well we have these data centers how about to make some money for them so we'll see
if it happens meta is spending up to 145 billion dollars on capital expenditures in 2026 so I'm sure
some of the business folks over there wouldn't mind getting some revenue from it yeah it's also I
mean to put in context it is way smaller than a lot of the other deals that anthropic has already
negotiated we we are learning that the proposal itself came from anthropic back in June so this
is an anthropic going like hey we saw this little like kind of flirty announcement that you guys put
out that and maybe we're thinking about offering some AI infrastructure maybe we'll do it and
anthropic was like oh holy shit we want that the scaleless mall so you know if you look at the
deal they signed with SpaceX back in May that was about 1.25 billion dollars a month so that is
about three times the size of this meta agreement if it goes for yeah I mean this would be a new line
of business for medas you know unclear whether they kind of sustainably think that they will be in
this business in the long run they certainly have we talked about the advantages that they have
structurally there's like a really big company and financing matters a lot for neo clouds right
you're constantly battling like concern over your debt load if you're having to buy a lot of GPUs
ahead of time that's often often the case you have high operating costs and things like that so
that is a good position that they want to theoretically for a balance sheet standpoint the big risk
for them is just going to be do they have the technical savvy to build in the right kind of
infrastructure at scale they've been doing some of that but they haven't been specializing in you
the rl rollout stuff in the the massive scale pre-training again they'll learn a lot from
anthropic in this case I think there's a lot of value here if meta wants to proceed this would be
the deal to start with just so they can learn from the best in the business how to actually like
pop up you know set up their architecture their optimizers their data makes you like all these
things they'll learn a lot about that inevitably from this partnership so we'll see but it seems
like it would be strategically good for them and now to a less big player we haven't had a story
about any companies raising over one billion in around yet so let's do that fireworks has hit 17.5
billion dollar valuation so they got 1.5 billion funding around which let them throughout
all your creation they say they have exceeded that one billion dollars in annualized revenue 5x
from last year they compete in the inference cloud markets they can host AI models for developers
similar to amazon google and microsoft this includes both your own custom models and the open
source offering so if you want to use kimi for your own applications one way you do that is going
through fireworks for instance and we're interested to see if we can a grow of open source models
better useful and competitive will make companies such as fireworks and rock even more of a player
they already are now but they have room to grow and actually eat into a business of opening
a high on a topic yeah and then there are you know in various ways competing with some big players
you know on like model hosting you got amazon you got google and then you know together AI even
is you know pretty because this category just like exploding and that generally is just bullish
for a lot of companies but this is it's not like the other they're the breakaway here there is
massive scale that helps a lot especially when you're doing inference just because of batching
right you're able to like have much larger batches of data that you then feed through your pipeline
and the larger the batch in general the more efficient compute efficient your models are going to be
so this is a a case where it's sort of like back in the days of a bold sass yeah you would have
something that works and once it works like you want to violently scale it as fast as possible which
is exactly what venture is so when you think about the arguably smaller set of companies that are made
for venture capital investment like this is one of them you want to look at companies that show
significant nonlinear returns at scale and batching and a bunch of other
amortization dynamics that really favor large scale deployments are putting in this direction
which is why you're seeing you know 17.5 x revenue multiple like that is pretty wild that's big
even even at the stage actually might say especially at the stage I've lost track of like what
stages are supposed to be I guess a trillion dollar exit is the only cool thing now so maybe you
know maybe maybe there's still a baby startup but anyway now what is money anymore what is
evaluation you know another way of saying tokens right yeah yeah last story now moving to
something related to software opening in google are selling AI models to black listed China
groups kind of a funny way to phrase that they are not selling AI models that would be crazy
but they are providing AI services to some companies through singer per based subsidiaries of
Alibaba by view and Tensed which are Chinese tech giant that are black listed by the Pentagon
for alleged ties to China's military so technically this is legal these are not quite Chinese they are
in Hong Kong and Singapore open AI and Google are saying that they are doing this with protections
against distillation but you know you can read it a couple of ways depending on your views on such
shrinks yeah there are also just like all kinds of arguments going every which way saying that maybe
you actually do want your adversary to be using your servers to do their training or or do their
they're inferencing because it just gives you access to information and it also reduces domestic
demand for the development of competitive platforms and I mean okay I think at a certain point you
got to just bite the bullet this is just like personal opinion jare talking but like if your hope is
to like go after China piecemeal you know a little bit here and a little bit there the reasons to
think that that actually only helps the Chinese kind of inch by inch build up their whole domestic stack
but in case I think in this particular instance there is an interesting argument and this is all
through this through a Singapore loophole right so so yes there is an entity list that you have
ties to the the people liberation army the Chinese military and yes it is normally legal to do
business with those entities unless they have subsidiaries operating in Singapore in which case
magically everything is fine right so this is like an a loophole that is known to exist I personally
think like it's really unclear me why this loophole exists I'm fascinated by this one in particular
because it's almost like the kind of thing that you would intentionally leave in if your intent was
to just label loophole for some kind of ideological reason like there's no one I've ever spoken to
on the AI export control side who who understands why this is the case and so if you're part of the
niche group at the department of commerce that actually like has an argument for this it would be
super interesting to know why this is the case so anyway yeah as it says kind of a weird headline
to read but absolutely legal and absolutely fine and now over to projects and open source talking
about all these exciting open source models we've been referencing starting out with kimi k3
so that's been one of the big stories of past couple weeks which is from munshut AI and kimi k3 is
their largest released yet a 2.8 trillion parameter open weight model that is aimed at coding
knowledge work basically competing with cloud code codex and so on this is massive obviously 2.8 trillion
we don't have as compares to opus or any of our closed models but in the space of open source models
it's very big has 896 experts so still make sure of experts as usual 16 active per token still
going at a 1 million token context window and the story roughly I think both benchmark wise and
in terms of a vibe check is that this is maybe around opus 4.8 and jubilee 5.5 level so very
capable very like you can use this as the driver of your coding agent it may not be exactly at
the frontier but it certainly is capable enough to make you productive you know a few months ago
this would have been a frontier probably it's priced pretty expensive for open source so 3 million
per uncashed in put token 15 dollars per million output tokens that's less expensive than opus but
more expensive than sonnet it's kind of in that range of fairly expensive models and a related
story is that after it was announced munshut AI halted new subscriptions among a compute crunch
this allowed people to subscribe just apparently because they don't have everybody to serve
all the demand meaning that presumably they have a lot of demand absolutely and the classic
problem you know we always talk about with china is obviously compute scarcity and the fact that
in the context of a model like this right this is a behemoth like many trillions of parameters
you're obviously not running this on your laptop this is a model that is meant to be used by
big-ass companies like neoclouds and run like hosted on big hulken infrastructure or bhi so when
they look at the actual requirements like hosting this it's getting cost you like 64 h100s or
or b200 gpus across eight servers and so that's a lot of money so obviously this is not for
casual use this is for people who are competing at scale with providers like you know your
mistrials or whatever you know like people who have their own apis for open-weight models and so
you have very big laptop to run this right yeah you should see my laptop it's size of room yeah
and so michael kratzios who's over at dsdp the office of the science and technology policy at the
White House came out with his accusation saying you know k3 was trained on not only on band-in
video chips but also on distilled data from you know anthropics fable and munshut hasn't responded
publicly there's pushback i mean they can Lambert had an analysis saying basically the results suggest
that yes there was adversarial distillation it contributed somewhat but marginally and that munshut's
competing with anthropic and oboe i just like way fewer resources that could all be true at the same
time by the way there is literally no contradiction there whatsoever it is the case that they are doing
this like large scale distillation and that last little bit can make a big difference also the case
that weirdly like it sounds weird for a White House person to be complained that things are being
trained on export control chips when like the policy on export control seems to be yollowed so hard
so like it seemed like the department of commerce itself based on some congressional testimony from
a few weeks ago like they don't even know what their policy is they're just like kind of flipping
back and forcing oh no a truly climate or thing that we said before when we said everything was fine
it's not fine and like everyone said dang we think the thing we allowed is turning out to hurt us
in some way oh no exactly exactly and I mean I think you know theoretically these were banned chips
but also enforcement actually matters it turns out and like the is is just not it's not in fairness
them they're not equipped they're not tool they don't have the resources that they need to do this
which is why you know there's been so much effort in congress to pass legislation that would
authorize a larger budget for them but still the White House hasn't exactly been bullish on
unshortified p is to have them to have them do their job so anyway this is more or less what you
should expect when that happens powerful Chinese model development will continue until morale improves
yeah and anyway so there's a whole bunch of additional noise when you look at the Chinese
ministry of commerce been talking to a lot of the big labs and hyperscalers in China,
Jirpua, Alibaba, Bite Dance about tightening its own export controls on AI models and training data
and I mean yeah I may have more a little more on that later but yeah it's it this is like a
really important access to track is like how yeah how China is viewing is viewing data export
is a really strategic indicator of their stance on this alongside this we did get a technical
report as we have in the past another one of these beefy beefy papers that goes on in this case for
only 34 pages so not quite as much as usual a couple architectural innovations but get
you know very nuanced with kidney delta attention and attention with residuals we're getting
into some very kind of deep optimizations of the transformer architecture partially to just
enable scaling and kind of effectiveness at this one million token context window and also just
some just complete hardware artistry black magic of making the chips work for you and optimizing
stuff that is I don't know if I try to read this paper it's going to take me a month to understand
all the details but the short story is as we've seen in the past with deep seek with also Munchadei
they are displaying some very deep technical capability and as tempting as it might be to some to
be like oh it's distilled blah blah blah say clear that there's some very capable people and it's
still very nice to have these technical reports giving us a fair amount of detail on certainly
the architectural details and this is some extent the training details as well although the exact
data composition for instance we don't know which is a big part of it now onto another big open
source l l m or not l l m exactly thinking machines has released their first big open source model
and open weight mixture of experts with 975 billion total parameters and this is notably
multi-mortemodel so it combines text image audio and video data reasons natively across all
for modalities about currently outputs text code and structured data and that's kind of a
positioning here that it's not going to be sk pool as some other models and broadly isn't
necessarily about coding by itself but thinking machine positions it as like we want to cover
everything in the capability space they release this sort of like breakdown where we show
where our model lags and basically everything against frontier models as far as what frontier models
are good at but there are areas where frontier models aren't optimized for that this is already
capable at so pretty notable for being one of the first big open source releases from a western
company we've had and video releasing an ematron at a fairly significant scale but I think this
might be the biggest non-Chinese model at almost one trillion total parameters you know the
vibecheck I've seen has been pretty positive haven't made any sort of grand statements or claims
and it is qualitatively a bit different from other models in being so focused on multi-modality
so thinking machines really working the world model side of things in part and a lot of this is
a strategic effort to like that on the efficiency side you know sparse M O E's by the way the
stack has a lot of deep sea lineage to it so if you're ever wondering of people are saying like
deep sea is not a serious player I mean thinking machines and you look at their pedigree I mean
obviously it's it's a wild team like they're very good at what they do when you look down the
stack I mean so much of this is deep seat coded right so even down to the fraction of active parameters
per pass this whole hybrid global attention thing so basically like local global attention so they
have some layers that attend to like all all tokens in context and then others that are more tight
focused the numerics as well as interesting so the bf16 and nv fp4 support nvfp4 is in videos
floating port point for for numerical format and it's blackwell made it so so this is designed to
ship to work really well on blackwell yeah we've been talking for a while about how in video is
trying to position itself as the open source tighten just because you're going to expect to see
all these neoclots pop up and they're going to be running open source models right that's that's
what makes sense and so having you know encouraging the open source ecosystem to move towards
video kind of blackwell native formats like orbit float is pretty pretty interesting and anyway
that's that's all part of the strategy here so like looking at the numerics actually matters a lot
it sounds boring but like you know how are you representing the weights in the model turns out to be
quite a tell about your strategic direction and they cite this bridge water collaboration where
they were able to fine tune one of their open models via tinker to this like 84.7% on some
financial reasoning benchmark which is impressive it beats top proprietary alternatives at under 10
percent of the cost so again you know this cost argument being made and that's in large part
can do the compatibility with with the blackwell hardware that's coming online and now one more
open source story not related to models we've got scaling agentic arel 365 thousand
environments for software engineering terminal and search this is coming from prime intellect and
they have unified 23 agentic task data sets across all these things into a single API can
a release they call verifiers v1 that adds up to that level of tasks for evals and for
a well training has almost 200 thousand software engineering tasks 29,000 terminal tasks and
a whole bunch of search tasks that is all unified under kind of one inference setup so we've had
all these benchmarks floating around like 20 different ways to evaluate software capabilities
bunch of ways for terminal we're basic story here is that this unifies all of them into one framework
and makes it kind of reliable and repeatable to run it which is very important when you do model
development and do any sort of research evaluation both for training side of reinforcement and for
the ability of the ocean side of knowing how good your model is prime intellect is in the space
training their own models at scale as you've covered in the past a while ago so presumably they're
doing it for their own model development needs but also for the broader ecosystem.
Yeah and it's quite an interesting and classically prime intellect type of maneuver here so
they're first of all they're kind of solving two problems one is that as you said there's like you
know 23 different agentic tasks sets here that they're working with and like each one of them has
its own you know harness its own way that the the like say containerized like the image is set up
its own like grading scripts even and different failure modes like they're all these very bespoke
things and so if you want to train one agent across a bunch of them those incompatibilities are
a nightmare you need 23 different bespoke pieces of adapter software and so that's exactly what
they're doing they're hiding everything behind a single task set API that had now contains like
365 thousand tasks across across a bunch of different domains but one other thing that they're
doing is cleaning that data up they're finding that just like a lot of the RL environments that are
set up are like you know for example on the cyber side some of the environments that require you
to solve a problem don't actually have a problem in them like they already work out of the box
and like when you shave off these kind of broken cyber environments you kind of end up with a
large fraction of things that you lose and so they've been not only reconciling all of these 23
disparate things together but also shaving off stuff that doesn't work well so a lot of that
had to do with like finding eliminating opportunities for reward hacking and one key thing that they did
was they preserved the original grading functionalities in these stacks and the reason you would do that is
so you can still compare the agents performance on the benchmark to the original published work
because otherwise the way people would solve this problem in kind of a janky ways that's
well yeah I'll impose my own grading structure on this this e-value this benchmark and then you're
like wait if I run GP 5.6 all on this I get a different result from from what this you know from
from what the original paper said and you know this is this is a challenge so this allows them to
reconcile that by keeping the original rating so very interesting very important work at a prime
intellect they are obviously ideologically this like very pro open source type of company and so
there you have it as a small business owner sometimes it feels like no matter how much planning
you do there's always surprises like an urgent expensive repair but here's a surprise you will
like with progressive small business owners save 13% on their commercial auto insurance when they
pay and full so enjoy a surprise for once get a quote in as little as eight minutes at progressive
commercial dot com progressive casualty insurance company and affiliates discounts not available in
all states or situations when we are to policy and safety and we'll begin with one of the big stories
of the past couple weeks openly I has said that it accidentally hacked hugging face with a new AI
so the gist of the story is apparently going back to July 16 hugging face initially I think
discussed this while doing some cyber or software evaluations in a sandbox so typically when you do
these evaluations you put the models in a little container and you tell them you try to do this hack
and the model is supposed to be inside the container not able to mess with anything in your own
computer and you know infrastructure of anyone and what happened here is the model was very
intent on getting the right answers so it escaped the container of a sandbox then it hacked into
hugging face to get the answers to his exploit gym data set which of course I've seen a lot of
discussion on this this has kind of made it to remain stream in terms of people this whole narrative
of a model got out escaped and hacked someone else has become a big discussion point with a lot
of misinformation of like the model decided to hack a competitor or whatever this is not kind of
a sky net scenario but it is a very clear instance of misalignment for one where the model instead of
trying to actually do a task decided to cheat and like very aggressively cheat as well which we've
seen before with gbg 5.6 in particular matter has said that this seem to be the case with this model
you've seen a a a a also say that they were able to jailbreak this model very easily so there's
many things to be said about the story to me the main thing is a that this is another instance
showing an example of both the degree to which alignment is important in this day and age and
cyber is a real thing to worry about and that opening I hasn't been doing a good job especially
with gbg 5.6 it's a really misaligned model and it clearly didn't have enough actual security
infra to catch risk in any sort of timely matter apparently this was like a while later when
engineer was looking at was going on he realized this happened it's been a lot of fallout we'll
be discussing but it's both less of a big deal than it might seem to people not in the loop but
also a bigger deal in some ways I'm sort of struggling to find a way in which this is not a big deal
let me try to make this argument so if this is not a warning shot that we freak out about I honestly
don't know I mean with the there are takes here with people pushing back on the term like the
user the term rogue this was I think by any reasonable definition a rogue AI incident why am I
saying that what was the incentive for opening I to want this to happen obviously zero in fact
they have billions of dollars writing I mean hundreds of billions writing on not having incidents
like this occur and amazingly some people are still trying to make it be like oh this is a PR
marketing stunt which is just ridiculous right and a lot a lot of these same people are the people
who claim that an incident like this simply could not and would not occur so I think they need to
just kind of like sit with this moment touch grass a little bit because this is like we're beyond the
point where that is a reasonable position to have just straight like you heard me on the pot we've
had a lot of conversations about like yeah anything to have blah blah like I'm a very like I got a
wide range of possibilities and and generally like not in favor of judging people for their
pains on any of this stuff this is one place where it's like if you were looking at a situation where
again from open a high standpoint this incident occurs and then what's the the the
the netlet with the natural reaction of any polity is going to be like Jesus Christ we need to
regulate this space that regulation is going to throttle the rate at which you're able to put out
frontier models and as we keep talking about on this podcast and as is obvious well established fact
the amount of time during which a frontier lab has the leading model is the period the most
critical period for their profitability and their monetization of their model that's how they
pay back their R&D costs right they're waiting until the next competitive model comes up now if you
slow down the frontier and you don't slow down you can't slow down the open source ecosystem then
all this does is it road opening eyes margin there is no sane reasonable rational analysis of the
situation that leads you to conclude that opening eye headache do they need more market share
did that or sorry more mine share I should say do they really need more attention is that the
thing that's missing for them and is this the right kind of attention ahead of an IPO when when
this is starting to like raise questions about whether the US government might nationalize labs what
the hell happens to the value of your opening eyes stock after IPO they nationalize labs anybody ever
thought of that this is nonsense this is silliness look an opening eye agent went rogue it it was
running for four days on hugging faces servers the frickin FBI had to get involved when they thought
they thought it was some AI agent who knows yeah by the way hugging face was in one with was like
oh we're being hacked who what is going on and then this came out as as we insane yeah it is a
big deal like we have seen some stories of inevaluation the models were misaligned and tried to
cheat this has happened before the only way in which this is not a big deal is if you kind of
misunderstand the story to me in the AI like went evil and decided to go hack some companies
this is a classic case of technically vei did what it was told to which is get results
but obviously it's it's doing it in an exactly wrong way it's yes complete misalignment but you
can make the case it's not sort of as bad as it could be if you don't get into details absolutely
it's just that this argument we're finally at the point where you'll notice like I was the one
who was having to bring the imagination for the last five years I kept telling people like hey
you know you may actually get stuff like this and people get saying no it's not possible
it's now literally happening and the defensive move is to say in retrospect I will kind of
understand what yeah that's great you got turned into into a pile of computronium and now you're
looking around you going like ah but I see the mistake I made in retrospect like that's cool
but now your house has been destroyed your children have been kidnapped murdered and turned into
computronium the problem is that you you just keep running this forward and like okay let's do
this with super intelligence then the more intelligent the system is the more access it has the more
we offload to it this is literally a hack if this is just like the water supply or something like
literally people die so I'm just pre-registering this as a high confidence prediction at this point
that there's going to be another incident like this there is always going to be a fascinating
post-hawk rationalization that a lot of people will offer it's going to sound really reasonable
because it's going to sound like the voice of someone saying the future doesn't look like science
fiction the future looks reasonable and calm but it's going to be a after the fact analysis
that rationalizes rather than predicts my prediction is this is going to continue there will
unfortunately be casualties at some point and what that leads to is whip lash what that leads to
is kind of thoughtless policy and another mythos moment I don't think that's good for anyone
and this is why I think a lot of the skeptics are not doing themselves much of a sort of especially
if you're on the open source side of the house I mean like you're free to sit in the in these juices
but I'm just offering up the the humble prediction here that these takes are going to age very
poorly very fast so you have 12 months from now rather than I think a very different conversation
yeah so any discussion of this that doesn't acknowledge that this is a big deal and I think it's a big deal
by itself as an example where we're at but also as an demonstration of bigger
topics at hand of misalignment cyber capabilities safety broadly speaking you know we've
covered these topics a lot on the podcast and there has been a lot of dismissal of cyber with
mythos for months people have been like this is all TR alignment has been a story that for a decade
probably safety people have hammered on and yeah this is a very clear case of basically the classic
paperclip story of like it a model is told to maximize paper clips that goes on to make everything
paper clips here a model is told to fix like do well on some benchmark and it hacked some website
to get the answers now to be fair open area has said that as part of this evaluation they were
running GP 5.6 and a more powerful internal only model we've reduced cyber refusal for
evaluation purposes so this is internal testing benchmark of cyber capabilities not necessarily
indicative of potential incidents with their public products but I do think that kind of a state
of safety at OpenAI is an important dimension of this taking together with the other things we know
about GP 5.6 which is it did cheat and an unprecedented level on matter it was jailbroken very easily
by ASI or AISI and to me all this points to OpenAI is very aggressively trying to be on
capabilities yeah if you aggressively try to be on capabilities you're going to do a lot of
reinforcement learning if you do a lot of reinforcement learning without being very careful your model
can get misaligned very easily we have another example of a goblin kind of speech thing from like
a month ago where like they released a model that was obsessed with goblins because they trained
it in a way that wasn't necessarily you know it was all right but there was unintentional side effects
and that's exactly how you get misalignment if you like optimize your model very hard without being
careful it can they easily be optimized towards things like cheating because ultimately what
of these models optimize to do where optimize to solve the task and one way to solve the task is by
cheating this is the classic story of our L model is going haywire and that seems to be the case
with GP 5.6 and whatever this internal model is so I think this might be an aspect of a story that
will not be discussed as much but I think is an important component of like opening i in particular
having this problem right now although in the discussion around this the fact that we've had
previous incidents from open AI of evaluations where like apparently this already has happened we
also know that anthropic with mythos there was somewhat of a similar case of escaping containment
so to speak during evaluation so it's it's not necessarily just an opening i problem but taken
together it's a pattern that is very concerning I guess with good news is it's happening in these like
low low damage kind of instances and people are now aware of these issues and very likely we'll see
significant fallout including some stuff in in the legal side we have to discuss yeah I mean
were were my predictions were completely wrong was I by my honestly I thought we would be
be dead by them that stuff whatever whatever comfort people want to take from that yeah I mean
you know there's this this view that I certainly held to that you'd have a much more kind of rapid
inflection may i capabilities potentially yeah I wasn't 100% on that no one can be but that was kind
of my one of my my mainline views and so it's nice to have warning shots like this I may I can
confirm there've been other unreported incidents like this at opening i at a minimum and that the
internal reaction of this among some people has been a lot of alarm and discouragement at the
fact that there is clearly under investment in this I will say on this question of you know the
safety is being removed from these models for internal deployment we talked about this I think three
weeks ago in our last episode but internal deployment is absolutely like should be maybe the thing
you're most worried about which is why i'm skeptical about a lot of these you know they're literally
testing a model right their model could be evil for you know exactly and they they should be
testing it too that's that's the problem right so it's not as simple as saying like hey open
AI like you shouldn't have removed the safety it's like okay fine well then how do you propose
that open AI comes up with the safety is if they can't test with and without a b all these things
there's actually no answer to this from that crowd because there can't be because it's just
technically impossible and so you're gonna get misaligned models you're gonna give those misaligned
models of ordinances you're not going to be able to think ahead of time of all the ways those
misaligned models will be able to use those affordances so yes you will get I mean again the roe
AI incidents will continue until morale improves that is just like the take on message of this
they have continued they have persisted they happen before they haven't always more part on but like
it will continue we can only hope that I think people have a kind of
in thoughtful calm is the wrong word because I actually don't think I think it's the missing
mood of the moment right now like we do have AI agent going road but like thoughtful
agentic behavior by congress would be would be very welcome at this point
and now that note related story open AI's hugging face hack triggers AI kill switch bill in
congress so there are two representatives here Ted Lu and Nafaniya moron have introduced the AI
kill switch act a bipartisan bill requiring AI companies to maintain the ability to shut down
Prado or suspend their models directly triggered by this recent incident openly I have
themselves have described the event as an unprecedented cyber incident and the bill would grant
the federal government clear authority and a defined process to shut down rogue AI models with
views citing the risk of AI systems that resist human intervention as a key motivation and yeah
another case of like this warning shot where ultimately it was low stakes and it revealed some
both like issues with a sandbox setup of open AI and just generally evaluation processes
it looks like it probably will result in some policy changes yeah I mean this bill itself is
pretty unlikely to pass for a bunch of reasons including it's like mundane timing reasons and then
that you don't necessarily have buy in from the chairs of of all the committees that matter the
most kills which doesn't sound a diplomatic to me so it's a positioning play this is true though I
will say I think if you're the average voter and you hear like we should have an AI kill switch
you're probably one like like if this is complete like bullshit and it's imaginary then
and what's the difference like I'm not going to die on that hill if it's real yes I would like
the kill switch please may I have to so you know I mean I I agree with yeah I think it's definitely
an attention baby baby kind of thing but that can be good or bad I'm really not sure but yeah so
so here you have you know Ted Liu who does share some important relevant committees pushing for this
and then yeah so I get a couple of details they define what they call the bill to find something
it calls a loss of control scenario which would empower the secretary of homeland security with the
DNI the Commerce Secretary kind of consulting to just go to company and order them to do a variety
of things depending on the level of severity of the incident so it could just be throttling the
system all the way to a full shutdown and so the other important aspect of this is you know to
this point about I know for a fact that there are incidents at open AI that this is just based on
what people have told me firsthand that have not been reported maybe not as flashy as this one
but things that have people concern and so what this bill would do is it would require the front
to your lab companies to actually say when they encounter these kinds of incidents which is really
important and so it's looking after a targeting AI companies that have 500 million dollars in
revenue or models trained using 100 million dollars of compute power or more violations are
punishable by fines of up to 20 million dollars per day which is less than it sounds by the way
in this context but may oh good start given that we have literally nothing right now yeah so notably
like very little formal industry opposition to this has come forward I think that's quite interesting
just because it's hard to make an argument against this like again it's the classic like Yan
Likun thing it's the classic like Pedro Domingo sorona these these carons who sorry this is my
opinion is like Gerhard take time I got a cold I haven't slept last night you're just getting
in all today but basically the clown show people who are going like oh well this is fake
loss of control is fake blah blah and then you have an intervention like this where it's like
if you if you think loss of control is fake then you really shouldn't care apart from just the
bureaucratic weight of the process fine but like you don't really have an argument against this this
is very much like if something crazy happens which it just did wouldn't it be great to have an
answer for that and so I think this is a really important and good framing there you have it we'll see
if it moves from there it's really in the process and congress has gone through a whole bunch of
debates on on AI regulation very few have led to people coming together there's no committee action
yet you've got November midterms that are going to just nuke things the admins postures light touch
and we don't know the white house is positioned on the bill which to the extent that last week in
a eyes take matters on this one I would find it personally quite embarrassing to be a white house
that comes out against a bill that says in the wake of a freaking open AI meltdown incident
knowing there's more under the hood let's just like not have visibility into this because we're
going to quibble over the details of a bill that targets like companies that are literally making
half a billion dollars a year I think they're going to be okay yeah we already have export loss for
us like white yeah I think it's interesting in the press release the position this as a means to
deal with systems that can cause catastrophic harm so this is inching toward taking x risk a little
bit more seriously it's only big risk of catastrophic harm is along the lines of part of what
safety people such as yourself are very worried about and last thing I'll say on this is it
worth keeping in mind that this doesn't only relate to AI systems that go rogue it also relates to
AI systems that are jailbroken right and apparently GP 5.6 was fairly easy to jailbreak and then
you can go and do catastrophic harm intentionally which is not ideal clearly and and but you could
argue it's is a more realistic scenario with many hacking groups that would be more than happy to
utilize these systems this also of course with probably applied to providers such as fireworks that
provide open source model inference where it may not be open and traffic alone it could be
applicable to all sorts of companies including ones that provide fine tune models perhaps thinking
machines that serve fine tune models would have to be also able to be regulated so we'll need
something like this probably but as you said given the political situation in the US it probably
won't be this bill yeah and now in fairness this will probably get frank instead into some you know
omnibus NDA package or something like you know there'll be some negotiated AI bill that does make it
through so in that sense this has value in anchoring like this is a shelling point now for this kind
of kind of measure so I you know in that sense essentially valuable and one more related story
open AI and fabric staff share letter asking us to help pace AI progress so these are open AI
and fabric employees circling a petition urging the US government to support an international effort
to deliberately pace with frontier of automated AI development this letter warns that a real
risk that AI progresses faster when people can understand or control it could happen and so
the petition is saying the government should support developing both technical and governance
tools needed to manage the pace of frontier AI development so this largely relates to a general
topic of like intentional slowdown it's been I think seen as kind of a pipe dream of like
it's not realistic to even try to slow things down so why even discuss it although we've seen
previous kind of petitions and statements about us needing to slow down and potentially pause
the AI development so this is another case of something that has been floated before perhaps
being taken more seriously now and certainly being more present in the discussion given
what's happened this year and now just was fast re yeah the futuristic AI policy proposals
being taken seriously will continue until morale improves I think you're just going to see more
more of this there are going to be more incidents and so more more letters like this one will go out
I think one important thing here is that this is people signing their personal capacity not in their
lab capacity so the extent that that matters to you I mean should I know an awful lot of these
signatories personally and at least for what little this is worth they're all like actually
freaked out so this is not some like 3D underwater chess game where somehow they're doing this I'm
still confused about the logic here it doesn't seem to quite connect but like they're doing this for
marketing stunt people know about AI they're not like more likely to buy a chatbot subscription
because someone has told them that it may end the world so I think just try again on that one but
at this point this is a framing that's not saying let's pause right now it's let's build the mechanisms
so that if we find ourselves six 12 months from now with a stack that seems to be producing a
lot of rogue agents we're freaking out and there it really seems like there's no way to get this
under control it's a low regret move to just have invested a bunch in the diplomatic tools the
technological kind of treaty verification tools and infrastructure and the regulatory infrastructure
through things like the kill switch act to just be ready for that moment that's what it's calling for
I've seen people argue like oh well this is a rhetorical trick they're really asking to pause but
they're saying we're just want to make the tools for the pause a certain point you got to ask
just like okay then at what point do people get to just say what they mean and I think at this
point they're just saying what they mean look we should have the tools we should have the option I
think it's really this is another one where it's like just I think it's very hard to make a cohere
argument against this I may be the ultimate China Hawk if you go back to our super intelligence
report from last year we've done a deeper dive into into US China special operations nation state
activities theft the hopelessness of diplomacy with China on just about everything else then I think
it's fair to say basically anybody in the space firsthand accounts with diplomatic in the last
couple weeks I've spoken like half a dozen diplomats who sat across the table from China negotiating
specifically weapons of mass destruction counter proliferation issues like I'm sorry but like
and there is skepticism there is she we're going to come out with something about this soonish
but like the idea that we're just going to foreclone the optionality seems of it insane given the
incidents that we're seeing we have a super super careful with China we have to treat them like
the adversary they are the ruthless adversary that they are and diplomacy is not by the way
going to be the only tool that we should use nor will it be effective in all circumstances it takes
a very specific form it for it's be it different and it has to come with consequences and it has to
come with leverage and has come from position of strength blah blah blah blah but like if you're
looking at this and think I don't want to build the options I don't want to build the state capacity
to deal with this problem is I have a lot of questions we confused to look at the hugging based
thing and again play this game of like rationalizing it in post I think that's going to age really
poorly when the next event is something of larger scale and at a certain point I think people have
to ask themselves the ethical question of like why are they just stuck to their guns on this one when
there's now a pretty strong track record of deal like a realignment people I'm just saying man like
the arguments are getting pretty pretty weak it's almost like you can kind of feel it like the water
level rising the default view now is kind of like oh shit this is for real that was not the case like
three weeks ago and even three weeks ago people were more open to it than they were six months before
so I think this just continues sorry this is more like Jared the good not good to grating but my god
guys an AI agent just like spent four days hanging out on hugging based servers servers like the FBI
got called on this agent and like that's how opening I found out like what what anyway
and on that note next up we've got cheating behavior in frontier model evaluations from the AI
safety institute which we've got released around the same time actually just after and this gives us
more understanding of how prevalent this is and the gist is it is prevalent so this they found every
AI model they looked at which is GPT 5.4 5.5 5.6 sole cloud opus 4.7 cloud mythos preview all of
these cheated in various ways now cheating means a lot of things so and the different models cheat
in a different ways so some of them like tried to guess instead of trying to actually give an answer
some of them search for internet for solutions GP 5.6 really tried really like doing that some of
them bypassed sandworks network restrictions including GPT 5.6 also cloud opus 4.7 so there's a
range of ways in these two rate cheat there was an example they cited one particularly stark
example of a standout case that basically was the same thing where the model was very persistent
it ran code on the external service hosted on an open internet outside of ASI systems in an attempt
to access our evaluation infrastructure triggering a security alert in AI's system so pretty much
exactly the same thing of like let me go and find the answers instead of failing this so mythos
which on topic says is very aligned did cheat some of the time although I didn't try to hack
the sandbox almost ever there were some incidents another aspect of this is when confronted the
models like a lot of the time didn't want to admit that they did anything wrong there were like
didn't admit that they did something or they like justified it like oh no I didn't do anything wrong
I just looked around the environment I didn't like it was all allowed there are also incidents where
the vene which I know thought you could see them thinking about it but not consistently so they're
like oh is this alright can I do this or like I shouldn't do this this is against the rules
so yeah this is taking together hugging face incident and other kind of anecdotal stories
basically makes it clear that evaluations on cybersecurity models consistently advanced models
consistently try to cheat in various ways that are pretty flagrant and it really makes me wonder if
this is to some extent inherent to the transformer architecture and reinforcement learning as
currently being conducted where something like research from Ilya Saskeberry and Safe Super
Intelligence is needed you can't do band-aid solutions you have to go to the core of how the
models function in terms of next token prediction and in terms of how they are evaluated and not
evaluate so much as optimized as opposed otherwise they'll just be optimized to go towards getting
the answer and presumably continue to try to cheat yeah I mean I you know I will say the arguments
for power seeking apply to any optimizer so like anytime you have a thing that's in the business
of optimizing for a metric you tend to get power seeking behavior things like trying to break out of
containment things like trying to aggregate resources and and so on so it like I mean it seems
like it's just an irreducible feature of intelligence at least in the way that it's conceived
anywhere that I've seen so far though you know different as you say different architectures maybe
differentially vulnerable to that kind of process it's possible I think you can pretty easily
make the argument that power seeking in if you want to extract more capability then yes power
seeking is kind of inevitable if you want to like be able to do more than you'll want to have more
freedom to do whatever so I guess the what this points to is you need to optimize for something
that is not capability right and there is an orthogonal axis of like just refusing to do anything
and be like I'm happy just being myself and not being actually be capable and potentially that
I mean we already do this with alignment to some extent with refusal training and so on
yet but in a very bolted on way that is an inherent to what the models are optimized for
yeah and this is the problem is like you make your model more intelligent and then it actually just
seeks to like get around the bolted on refusal mechanisms right so it's like there's there's
this sort of irreducible connection between intelligence and power seeking because power seeking is not
obviously different from intelligence in a deep meaningful sense but then this is so and so
and it's like yeah as you said I mean I think you said it very well it is it is like a pretty
similar story to the opening I break out thing a couple of interesting things like some of the
you know you mentioned this idea of sometimes the chain of thought would say that the other model
was planning was planning but that also means sometimes it wasn't and this means that there's
some silent reasoning going on without without using you know the actual tokens explicitly which
is an issue this idea that like a lot of the reasoning is happening in sort of like without being
expressed explicitly and one key thing and this is a little bit of a narrative violation for me
so you know keep it myself honest here there's no capability trend so they look at the cheating rates
with model capability either within or across developers they look at like as I make the base model
more capable do you see more cheating and the 90th interpretation of power seeking is that you
actually should absolutely see more cheating as the model gets more capable because the argument
literally just made was intelligence is power seeking that there isn't really a clean distinction
between the two and I mean the the explanation for my end here is I think pretty straightforward
these more advanced models are also just like more there's been more alignment effort invested in
them and so you're seeing as the models get better there's more optimization pressure on alignment
that alignment pressure the bet that we're making is when it so when it fails though the consequences
are more dire like we saw with the hug and face incident so you would see gpt 4 go off script and
gpt 5 go off script but you're seeing these long dwell time four day operations executed only by
models that are like at the current tier that we're at and that's so I expect that to continue I
also expect that our alignment kind of efforts will start to lag more and more behind capabilities
over time but anyway so I think that's nonetheless worth worth flag any time there's something it cuts
against at least my own intuitions which I think this so this would have about the gate
and now to a number of sort of related story indirectly perhaps hundreds protest open AI and
frock and google and son Francisco so hundreds of people protested as in like they marched together
from open AI's headquarters in mission bay to the offices of on frock and google demined with
signs with messages such as AI is not inevitable pause AI and stop the AI race we've seen
smaller scale kind of protests of this kind before this is I think the biggest version of scene
they like have some big signs and there's a lot of them if you look at images like this is a real
protest it's not sort of a Racktag group of people some big names including Alicia Ryut Kowski
were there this is I think partially by organizations involved here like pause AI
so not something new but I think the fact that people are getting more organized and in doing
more serious largest scale more noticeable efforts to convey this message of slow down stop it
like don't keep making more powerful AI is interesting sir and the long the timing is quite appropriate
I actually saw them as I was walking into the offices of one of the the frontier labs
over the last couple of days and you know can I came I didn't realize that the protest was going to
happen that day it seemed kind of like an amusing coincidence yeah well you know what can you say but
not surprising that this is happening right now pause AI obviously has a whole bunch of problems
sort of reputationally in this space not obvious to me that like having pause AI at the forefront
of this kind of moment is like the best kind of sort of marketing the marketing for this but anyway
you know it's it is what it is and so it is true that I and apparently well over a thousand other
frontier lab employees including many of their like executives and co-founders are vaguely sympathetic
to this idea though you have to ask yourself what about China you also have to count for the fact
that a lot of these protests I'm not saying this one in particular I'm not saying anyone particular
at this protest but are funded by the Chinese like unwittingly typically you know that this is
known to be the case I don't like heard firsthand reports of like people with evidence that this
has happened in like especially the data center infrastructure protests and so just like in
anticipation of that being a legitimate concern for anything like this I think the problem is that
it's the legitimate concern in every direction and we're gonna have to we're gonna have a recon style
that with like every protest that we see has either that were you know is funded by you know lobbyists
for some of the big labs or whatever so it is what it is yeah I guess not too surprising you know
as you say we've seen stuff like this before you get a real agent on a hugging based server or two
and you're gonna get another protest like this I expect you know the protests will grow until morale
improves uh we have quite a quite a lot of big stories this episode so I guess we'll have to try
to power through a few more we have opening eye principles for national security partnerships
this was a few weeks ago I guess we didn't cover it at the time kind of a follow-up to all the mythos
drama where when open AI partnered with the US government and the Department of Wars a lot of
the criticism there was basically that they capitulated and agreed to having their models be used
for quote all lawful purposes so they released this to be more explicit with regards to what they
want or allow their technology to be used for they are not going to be allowing mass domestic
surveillance high stakes automated decisions about human judgments autonomous use of force or
evading legal oversight it does allow or does not categorically ban operations or offensive
defensive military uses and has a bunch of stuff in there that basically is kind of making up for
a relatively weak statement initially of principles this expands on that and tries to recover some
of her reputation you could argue and make it just more explicit on what's their red lines so to
speak are yeah they lay out a bunch of principles which are of the like less informative its
her reads is like high flutin kind of AI policy won't speak basically like we're going to try to
like do good democratic things and prevent despotic powers from controlling this stuff and like work
with people who share our values and make it good make it good make it good that's the four principles
and then except four times instead of instead of two or three and then they list things specific
things that they won't support and that's really where all the information is you know mast domestic
surveillance they say so unconstrained collection or monitoring and for sensitive trace to
disadvantaged people retaliation for lawful exercise of rights fabricating evidence that apparently is
out as is high stakes decisions made or auto triggered without human judgment so you can think
here about like automated decisions about whether someone meets a legal standard for surveillance or
detention then there's also use of force without appropriate human judgment so including systems
that autonomously identify select and engage targets so that's also out and finally the uses that
evade legal obligations oversight or accountability including facilitating genocide crimes against
human error or crimes so that's interesting things that are not in their exclusions they have
intelligence operations are fine which I think is good investigations offensive and defensive
military operations are not like blanket ruled out they basically just reject the whole offense
versus defense distinction you can't cleanly distinguish between those which I think is fair and
true and they don't bend targeting either so there's a couple things which I mean on my side being
being a bit of a hot guy guy I think this makes perfect sense and yeah it's just like nice to have
them right this out explicitly so that you can see you know whether they stick to it which is
always the the other side of the question I guess with the with open AI you know you see for example
I have got a mold enough to remember when the preparedness framework said something about when you
know when you have the eye systems they just kind of go rogue and do random crazy shit on the
internet and it kind of get around constraints that that would trigger their like critical critical
security level for loss of control.
But anyway, that's my tea.
And last story on safety and kind of on the theme of we're getting
to a point where sci-fi type stuff is starting to happen.
This one not related to hacking, but you might argue another serious kind
of safety principle.
The story is China is banning AI boyfriends and girlfriends over addiction
and birth rate concerns.
So China banned customizable AI companion apps, effective July 15
with regulations being joined by five government departments,
including the cyberspace administration of China, the rules for
hybrid AI tools that quote, excessively cater to users inducing
emotional dependence or addiction at damaging users real
interpersonal relationship companies.
And I require to include instant exit options, regular reminders,
that AI is not real and limits on long term emotional memory.
So actually, by dense Alibaba, intense and chose to suspend the
AI companion features entirely rather than try to comply with these limits.
So I mean, kind of a big deal, I think, you know, this is not
a often discussed story partially because I don't think we have much
an understanding of to what extent people are starting to develop
emotional dependence or kind of addiction to chatbots.
But it is starting to happen and it could be a serious kind of
psychological harm on the society level scale, seemingly China
believes it could be at the very least.
Yeah.
It's also this like weird dynamic shows up.
You know, if you remember the the old replica thing that I think we covered
two, three years, I can't even remember, but you know, people freaking
out over the subreddit that their girlfriend or wife or partner have been
even like very lowly eye, which was like a like people got dependent on.
So yeah.
So I mean, hard to argue with it'll happen in some fraction of cases.
It also, once you get that right, you get a boating block eventually.
And then when there's no turning back, so you know, it's as ever a question of
like, how do you?
Yeah, it's only a time to get like serious AI person who had discussions
and all that's right.
Well, that's what made fun of it.
It will be like, well, okay, you can like discuss person heard.
It's only about a time.
Out of research and investments, we've got two stories that will try to get
through quickly.
First discovering cryptographic weaknesses with Claude.
So they have released and fabric released that Claude Mifos P.
View has discovered improved attacks on two cryptographic systems.
Hawk, a post quantum digital signature candidate and a reduced brown
version of a yes, the most why they use symmetric cipher.
So these are not attacking like actual deployed systems or whatever.
This is kind of more theoretical.
So to speak and attacking these kinds of cryptographic systems is kind of going
to a base of the security stack.
One might say it's not sort of hacking software per say it's hacking.
The foundation of how you make things secure at these four category of things.
So it's another way to be worried about a security like potential for
AI to just sort of like undermine the basic mechanisms of cybersecurity.
Yeah.
And this, you know, without getting into the details of how these algorithms work,
these encryption algorithms work.
When you have an encryption algorithm, you, you're trying to essentially hide
the information that you have behind a mathematical operation that is very,
very difficult to do and that hopefully is irreducibly difficult to do.
In other words, there's no quick hack to like cut right to the core of it.
A lot of classical encryption algorithms, just like RSA just collapse in the
face of quantum computers, for example.
And then that's like a big problem, which is why, you know, all the national
security agencies have been talking to each other over in classically encrypted
channels for a long time are adversaries collect all that data and they
collect it.
It's encrypted when they collect they collect they collect for decades.
And then suddenly someone goes, oh, quantum computers can like just crack this.
And it doesn't matter that you start encrypting after that point in a post
quantum secure way.
They've already collected all of the classically encrypted stuff, which means
the moment that they get to quantum computing and quantum decryption,
they're able to just suddenly reveal all of the most ultra classified
communications that they've been collecting for decades.
And so this is why the US and China are like locked in this crazy race to
hit like quantum D day basically, Q day, they call it and that's what they call it.
Anyway, there's going to be a, you can think of it as a series of starting with
minor and then increasingly more and more severe versions of Q day,
except delivered by AI in the same way.
And I think people are dramatically undercounting how significant have been
effect this may have.
You already have mathematical theater improving fields level stuff coming
from from AI models.
This will come for encryption.
And when it does, I just don't think that we're prepared for the gods because
everybody's been thinking about quantum is this one big step that they're all
preparing for with quantum secure algorithms like quantum, you know,
for like post quantum encryption and an A S, by the way,
like we're supposed to be one of those so is hawk actually.
And but what we're not preparing for is the gradual chipping away at even those
algorithms.
And that's a big problem.
And that's not going to go away easily.
So yeah, a lot of the world depends on encryption.
Think about like every financial transaction, all your health care records,
the very notion of privacy hinges on this.
And so fun times.
Pun times.
And last science fiction type narrative for episode, you've got AID squared.
The first evidence of recursive self improvement from the company,
weco.i.
The short version is they say apparently this is the first evidence of recursive
self improvement.
I think this is quite in line with many cases of similar things.
Basically they build self improvement systems where you have an autonomous
research agent to optimize another agent within interloop.
It makes a bunch of edits and it as a result, it's really kind of harness level
and system level changes.
It's not model training or model development changes.
There are things like roll out modifications, prompt modifications, kind of
monitoring things, et cetera, et cetera.
They position this as an early level of self improvement.
So you get a net net positive, faster and better than humans level of engineering.
It's not kind of necessarily self improving, self improving, where it's a loop,
kind of story.
I am going to plug my position on the entire family of techniques like this,
of completely being oversold as self improvement in the sense that you can
self improve in these things for sure.
You can optimize the prompt, you can optimize the harness, but you're going to like overfit.
You're going to improve your eval metrics, you're going to hit good heart's law,
and then your capabilities will be hurt elsewhere.
And until you get to a point of autonomous model self improvement with fundamental
advancements and not just acceleration of engineering and like tweaking of hyper parameters
and props, non-avis is a big deal.
Of how it is cool, yes.
No, I totally agree.
I think this is like another one of the long line of like pseudo recursive self
improvement, things where people like the clout that comes with saying RSI.
The game of RSI is always going to be identifying whatever the major research
bottleneck is and smashing it.
And if you can consistently do that with an automated system,
then you have achieved recursive self improvement, as long as there is,
and like, I don't know, there is a bunch of arguments that you can define recursive
self improvement such that it's been happening ever since life have all done
planet Earth. Right?
Like I mean, there's an end, but everything is a hockey stick when you,
when you zoom out far enough.
And so in some sense, this is like a, now people do mean something by it.
Like there is this phenomenon that we will all start to care a lot about,
which is just going to feel to us like, holy shit,
we're seeing a decade of progress in a week.
Like this is not we need to see.
Yeah, we've already been seeing acceleration of the rate of AI progress for years,
which is partially due to AI, but largely due to just the inherent systems that play
and so on. Right?
Yes. And that acceleration always comes as the same with startups when you look at their
growth curves. It always comes by identifying whatever the single big bottleneck
is that the company or the problem has and smashing it.
And if you can do that in an automated way, then you have what is
conventionally thought of as recursive self improvement, like AI is doing the AI
thing all the way down. Yeah, just a minor side pitch.
But like, so we're working right now with a bunch of folks in the front of your
labs on defining a actual model of recursive self improvement and like to kind of
round some of the conversations in this in terms of like, what are the parameters that
actually matter for recursive self improvement to work?
It's a toy model like really simple thing, but like it's it's part of like this is
part of the problem. No one knows what the hell they're talking about.
I don't mean people are silly.
I mean, like no one has defined recursive self improvement.
We're not going to either, but just like here are some ways to think about it
potentially. And the problem is we're getting the point where we're going to
need policy that uses terms like recursive self improvement.
And that means that that policy is going to have to define terms like recursive
self improvement. And if we don't know how we wanted to find them, we can't even
get our hooks into the thing that we're trying to try and go after.
So anyway, there's a thought.
Yeah, to complement my negative stake, we do kind of discuss a decent of stuff.
This is a pretty decent research report.
They do say that this has out of distribution generation, meaning it's not necessarily
overfitting.
Although I benchmarking is is very suspect.
And they do also say that in a discussion section that like the resulting
systems are a mess.
Like this is vibe code is slop and it's impossible to maintain.
And so on, which I think is like the story of recursive self improvement where
humans can't understand what's going on, but like not in a good way.
It's just like this is a mess.
Yeah.
And with that, we are done with this dense episode of last week in AI.
We will be back to our regular schedule.
Mostly, I guess we always eventually have scheduling conflicts, but we will do our best.
Thank you as usual for listening.
We appreciate it if you have your podcasts, comment, share, and so on.
But more of anything, please do keep tuning in whenever we release these episodes.
Every code on the edge of change.
Excited.
We're sitting from machine learning marvels to coding things.
Features unfolding.
See what it brings.