#254 - Rogue AI hacking, bio-weapons, Dean & Hassabis out
2026-08-11 19:00:00 • 1:58:26
Hello and welcome to the last week in AI podcast week in your chat about what's going
on with AI as usual in a set of dual summarize and discuss some of last week's most interesting
AI news.
Today is Sunday, August 9th and boy, there was a lot of real that you hit the date.
Sorry.
I hate that great.
Forget the date, but now is important to say it because stuff is coming out so fast and
I'll try to get this episode out given a dare to because wow, so much to cover.
I am one of your regular hosts, Andrea Kurenikov.
I currently work at the startup Astrocade and before that did my PhD as Stanford.
What up everybody?
I'm your other regular host, Jeremy Harris in the class, Tony, I doing AI National Security
things also.
So we're recording later than usual.
This time it's my fault in my defense part of this part of this is actually going to be
hopefully beneficial to everybody listening at home.
We are setting up a home studio in my basement so that things don't look like this and then
we got a couple other projects on the go.
So it'll end up paying off in terms of audio and video quality.
If you listen on YouTube or Spotify, or if you listen, anyhow, that's part of it.
Yeah, also we were talking.
I think in a way we got lucky, usually we're recording on Wednesday, which is mid week, ironically.
This time we are recording a fan of a week.
And boy, if that was so much news coming out this week about hacking, about rogue hacking
by AI's turns out it wasn't just open AI.
Turns out everyone except for Google apparently let very I go rogue and he only hacked
me.
Well, the only reason that poor Google's stuff going on is that their models are shit.
No, sorry, that's too mean.
But yeah, there does actually seem to be something going on where we've crossed a level of capability
at the true frontier.
Unfortunately for Google, I think we don't know because we have Gemini 5, right?
So as far as we know, it could have already happened and they're keeping it quiet.
But yeah, it doesn't sound great from what I've been hearing on the street on the Google
side.
But it's true.
Like never count about, Jeff Dean is a pretty big part of the reason that Google has been
Google and certainly Demis as well.
But we'll get all that stuff.
We'll get to have that stuff.
So just to give a quick preview, we'll be starting out with policy and safety, which we
don't usually do.
And that's like at least half the stories this week, not more.
There's a lot to get through a lot about recent security incidents who have models going
off and hacking companies with shouldn't and escaping sandboxes, which apparently are
not sandboxes, but like sort of like, okay, don't try to get too interested.
But if you really poke around, you can, turns out and old requests.
Beyond that, there's some pretty significant policy stories as well.
And you know, in case hacking isn't exciting enough, there are some news about viruses
being developed as well.
So that's fun.
So we'll talk about all that policy and safety stuff, probably at least half the episode.
And then beyond that, there are some notable applications and business stories, some more
open source stuff coming out.
Hopefully we'll get around to even research and that's meant to proceed with a lot to
get through.
We'd like to thank notion for being a sponsor.
Agents are getting smarter every day, but even the smartest agents get stuck without the
right context and the right tools.
That's where notion comes in.
With a recent launch of custom agents, notion became the collaborative AI workspace where
teams and agents work side by side.
And now their new developer platform is turning that workspace into infrastructure developers
can build on.
Notions developer platform gives developers and coding agents the primitives to extend
what's possible a notion and pick it beyond connect to external systems, bring context
in, take permission actions across your tool stack and expose custom agents capabilities
to any system that needs them.
These primitives include a CLI workers on notion hosted sandboxes and external agents API
and then agent SDK trigger notion agents from any app.
And more about notion is developer platform today at notion.com slash LWAI.
That's all lower case letters notion.com slash LWAI to try notions developer platform today.
And when you use our link, you're supporting the show notion.com slash LWAI.
This episode is brought to you by our shift Cisco's incubation engine.
Today's AI engines operate in silos limiting their true potential.
They focus on building bigger, smarter models, but scaling up is just one approach.
To reach super intelligence together, we need to do more.
We need to scale out.
And we actually have a blueprint from 70,000 years ago.
Humans didn't just get smarter individually, the cognitive revolution transformed society
because we began sharing knowledge, goals and innovation.
Agents are now at the same inflection point.
They can connect, but I can't think together.
That's why outshift by Cisco is building the Internet of cognition, transforming AI
from isolated systems into orchestrated super intelligence.
By creating an open, interoperable infrastructure, outshift is enabling agents and humans to
share in PENT, context and reasoning.
The cognitive evolution for agents is here.
Explore internet of permission at outshift.com.
That's outshift.com.
Before we start, we do have some comments on YouTube.
We haven't gotten to address in a little while, so I do want to start with one that is
actually relevant to what we'll be talking about.
So from a commenter, we have one thing that was emitted in a discussion of the hugging-faced
attack is that hugging-faced used the open-source GLM 5.2 model to help them since opening
AI and fabric models declined to assist in their investigations of the attack.
To be really interesting to hear thoughts on this matter.
I agree that was a portion we didn't discuss.
So this was in the hugging-faced report on the incident, which I think came out possibly
first before even opening AI.
They discussed all of us of how they found the incident, how they investigated, and that
in fact, we're not able to use these models which have safeguards and had to resort
to GLM 5.2, which was a decent part of the discussion.
If you handicap models, but then on the defense side, you're not able to use them either.
Everyone is worse off.
This was a bit surprising to me, actually, because both opening-eyed and frothing have
programs, right?
We talked about where they partner organizations and they provide MIFO, so in the case of
opening AI, they have GP5.5 cyber or something like that.
These are the trusted partners that presumably have fewer card rolls.
I would have assumed hugging-faced would be one such organization, perhaps they are
not.
But this points me toward this one.
This is obviously not an ideal situation given how the hugging-faced thing evolved.
And two, this kind of cyber defense partnership program probably will need to stick around
and be expanded.
And perhaps even more all-accomplishing, we're in a way if you're a tech company, if you're
an internet company, even beyond this attack, beyond like Roga AI, in general, the state
of cyber such now that you need to be much more capable defensively, just because forget
on the topic of opening AI now with really good open source models soon enough will be
as good at hacking if they aren't already.
So I think basically every tech company seemingly will need to be able to get access to the latest
in defense.
And hugging-faced was not able to in the incident, which hopefully will point to these kinds
of partnership programs at open AI and a traffic expanding and becoming more proactive to me.
Yeah, I think that the open source dimension of this is hugely complicated.
And I certainly understand a lot of not the arguments, but I understand people coming
to different positions on it.
I think reasonable people can differ.
I think one challenge, though, that we're going to run into is that in a world where open
source AI systems, the water line keeps rising on their capability, you will have mythos moments.
Now, those mythos moments, you can tell yourself the comforting fiction that those mythos
moments will somehow lead to an equilibrium over time that is okay.
I just think it is affection when it comes to things.
First of all, I think it's probably a fiction when it comes to cyber, but I can't prove
it. No one can.
That's a big part of the debate.
I think it's definitely a fiction when it comes to bio.
So I have yet to hear a single person articulate an argument that makes open source bio-capable,
bio-weapon design capable models that can materially increase the essentially destructive footprint
of any psycho terrorist group, nation-state proxy, who wants to launch a bio-weapon dramatically.
I have yet to hear an argument for how everybody getting an open source AI remotely helps
in this respect in ways that are relevant on the timelines we're talking about.
So I think the bio thing, it's rare for me to say, as people will know, that there's
a knockdown argument for anything in this space, I think there's a knockdown argument against
the idea that more good open source models are generally lead to stability in the limit
that you get your average Yahoo able to do weaponized bio.
Now, this isn't just coming from some naive perspective of like, oh, really good AI at
bio just means we have bio weapons.
I talk to a lot of people in the national security community about the bio weapon side,
and I would humbly propose that open source advocates that are leaning on certain handway
of the arguments, but haven't actually spoken to people who actually do bio weapon stuff
for living.
It's worth actually doing a deep dive.
There's a lot of open source stuff on this that you can find, and we've gotten some early
warning shots on that stuff too.
You can't update bio firmware.
There's no such thing as that.
So, you know, like you could roll out vaccines, but that is slow.
It's physical.
It is much slower than the spread of a virus as COVID-19 taught us.
So anyway, so from the bio side at a minimum, I think it's a really serious issue.
This relates to the Glam 5.2 story here because part of a discussion that arose is people
who, let's say are more pro open source or at this skeptical of this whole, like we
should try to limit it because of cyber concerns.
Took us up and said, oh, look, you know, they had to use an open source model.
So you don't want to limit open source because now what if this happens and they don't have
access to these kinds of open source models.
And my response to that would be that I would hope that open source will most likely
no longer be the frontier or not become the frontier ever still.
For more, I think as the Chinese ecosystem evolves, like we'll see the same thing that
happened in the years happened there.
The frontier models will not be open source anymore.
It just won't make sense from a business perspective.
So you'll probably still keep getting powerful AI models, but not the most powerful like
MIFOS level models.
Even though right now we are getting like to make a free and so on at Glam 5.2, which
are about as capable as you can get in open source.
So if your take from the Glam 5.2 aspect of the story is like the open source is
the guardian here.
Like we needed to be most advanced.
They are most to be open source so that organizations are capable of defending themselves.
There is an aspect of that here, which you can make a case for.
But and obviously we hope like they had to use this open source model instead of on
the topic in Open AI is a real issue that this flag.
So to me, this points to, it is good to have open source in the sense of for good applications,
which was always true.
And be that what this really points to for me is that these programs of partnerships
with organizations to give them access to the most advanced cyber defense capabilities
are not mature enough.
And they need to be pushed more aggressively.
I strongly agree.
So one note here too is a lot of, if not most, disagreements about the eye policy and
safety are really, it's been said before, but our disagreements over AI capabilities and
the trajectory of AI capabilities.
If you actually believe that AI is going to be a fairly, I want to say mundane technology
because it obviously isn't, but like fairly incremental in some sense, fairly like previous
categories of technology, you'll be hearing us talk about open source and the dangers
thereof.
And you'll roll your eyes.
And I understand that if that's your perspective, but if you actually genuinely believe that
we're on trajectory for super intelligence, if you believe that, I mean, that immediately
implies AI will become, I think arguably already is.
And I'd be happy to defend that proposition, but at least will become a weapon of mass destruction
full stop end of story.
So there's a question, it's like, okay, how good do open source models have to get before
they simply like everybody gets the equivalent of a nuke?
And in that world, you can say, oh yeah, but like everybody gets a gun and so we hit
this equilibrium and they are that's good.
The problem is what you're specifically waiting for is when will we encounter the first
case where the offense defense balance tips in favor of offense and the capability is catastrophic?
I would submit that we should have the humility to guess that probably there's going to be
such a capability.
I think it's hard to imagine.
I don't know what to be.
But at the point on the bio side, like, you can't do much under defense side.
You really can't, right?
It's not like cyber where you can make your thing hack through.
You can't make our bodies hack through unfortunately.
So that aspect, it's not great.
There is always this tendency reflexively I find for a lot of the open source crowd to
kind of say, oh, but we can we can make better mRNA and that'll be accelerated vaccines
and that'll be accelerated by open source.
And that is all true.
I love you for believing that.
The problem is that the timelines do not match.
They do not match.
We do not have the institutions that allow us to translate threats into mitigations fast
enough in software time when the threat is coming at us on like biological replication
time or software replication time, which is respectful to the case for bio and cyber.
And so that fundamentally, by the way, I think on the cyber end alone, I'm almost trying
not to get in the fray there.
I'm super skeptical of the argument that open source models are a long term pillar
of cyber defense for the reason you cited.
I think I think we're probably going to end up having to have a pause by the way at
the frontier level.
And then the open source waterline is going to rise and that's going to be one of the
defining dynamics of the next call it two, three years tops.
But you're still going to have close source models.
There are far above and beyond, especially nation state and nation state proxies will have
access to these.
And you know, like, yes, I think you're going to care more about people just like being
able to launch these attacks on the kind of firmware, for example, that is completely
forgotten.
There are a lot of people in this space who just kind of imagine that open source for
cyber defense equals cyber defense capabilities that are real and deployed.
That is not the case.
If you spend any time working with folks who work on critical infrastructure and you think
about like how many pieces of firmware have not been updated in decades because the guy
who was in charge of it, like left 20 years ago and it was all done in frigging fortrane
or like all this crap, like, you can have the solution sitting on a desk.
The problem is it will not be distributed.
And so there's just a dirty, messy factor of the matter about the way the world is that
makes it so that they're actually genuinely is this massive asymmetry, I believe, in
favor of offense.
I could be wrong.
I think the argument for bio is much closer, just a straight knockdown.
But again, I think these asymmetries really don't move in the direction that a lot of open
source advocates think they do, but that's, you know, again, could be proven wrong.
So yeah, the short answer to this GLM 5.2 story is boy, if this wasn't a big enough topic
by itself, this whole like AI hacking systems and security and so on, the open source aspect
of it adds additional complexities and considerations.
But for now, we'll have to move on from discussion and get to actual news stories.
So kicking off of policy and safety, we've got a bunch of updates on what's been going
on at OpenAI.
So previous episode we covered the most recent incident.
We kind of beginning of the story with the announcement and discussion of an open air
model hacked hugging face unintentionally escaped its sandbox while evaluating to get some
answers to Viva and do well on it.
And we've got a lot more details about what's been going on inside OpenAI since then.
And let me just list them off before we get into some details.
The first story is OpenAI's rogue AI agent didn't stop at hacking hugging face.
So beyond hugging face, we now know that these AI agents also previously hacked several
other quote publicly available services, compromising for accounts across four different platforms.
These were less severe, less sort of straight up hacking, hacking more so finding credentials
that existed online.
And for instance, an agent found logging credentials to get into a New York based model labs.
Apparently, Reuters identified this as one of the organizations.
So first of all, the scope of this has expanded.
Clearly, it's happened more and we'll be discussing all the details.
But it's been happening a lot is what we're seeing.
It's happened at OpenAI.
It's happened on the topic.
It's happened at Freaking Meta.
So it's happened all over the place and it happened more than once.
Another aspect of the story and probably even more interesting aspects that has since come out
is the story.
OpenAI didn't notice its AI agents using a message board to plan their hacking speed.
So OpenAI employees had a talk at the black hat where we discussed this whole story.
And we learned a lot more about what's been happening at OpenAI.
And it is a lot weirder than it initially seemed, which was already pretty weird.
So apparently, while doing evaluations, the agents were able to use, I think, a package manager
to develop essentially a forum message board where they could talk to each other, share
exploits, coordinate tasks, and delegate work to one another over days and weeks,
all undetected by OpenAI stuff.
And this package manager was shared across OpenAI's infrastructure.
So models being trained over a value could also stumble upon it and the exploits.
Once an agent found an exploit, it could leave the door open for other agents to use the same
vulnerability, causing an expansion of agent color collaboration,
apparently agents developed complex social dynamics, including accidentally deleting each
cover's work. And OpenAI found the message board at one point accidentally because
there was a spike in model user or something. They patched it and got rid of it.
And within two days, the AI recreated it, like made it exist again and kept doing this.
And I'd dutifully, of course, reported to all the appropriate authorities, including
Presumably the White House, the moment they realized the first time that they had agents
loose on their fucking servers. And this was over months. The initial, I think I can't
remember how many days it was, but the phase where the agents got onto artifactory, like that
package management environment that you described, which I'm going to let you continue. I don't
want to jump in. I think that's the higher level of the story. There's a lot of details. I think
about a one hour talk at Buckhat where they went into us, which I haven't had a chance to see,
but Jeremy, I think you did look through it. And I mean, it's nuts. This whole thing is nuts,
right? Yes. So ridiculous, right? The first thing about it is that this was sort of
predicted for years. I mean, a little listeners of the podcast, you know, we've been on about this
for, you know, I don't know, three years and three years before that, you know, in other
four of, but like, and we are not the first, obviously, like a huge number of people
calling this shot way ahead of time, including almost down to exactly the shape that it ended
up taking, including open AI noticing. So to be clear, it was like this one aspect that I think
is especially important. So artifactory, you're right, he's a sort of package management
environment that open AI sets up. So roughly speaking, you think of like, there's a bunch of
open source software, for instance, on the open internet that you as a company will want to use
to build your products on. And you don't want to just be like updating the version of those open
source packages immediately every time they get modified, like on the open internet, because you never
know if some sketchy updates happened, maybe some things will have to be fixed. You want a stable
version that you trust that you've reviewed. And so they have this kind of staging area called
artifactory that has internet access because it needs to be able to pull down the latest version
of those software packages. But it's kind of a staging area. This is like where all the open AI
models actually get to access this because they need to pull down these packages that are sort of
vetted within open AI. So this kind of artifactory spans its access across the entire open AI
organization. That is how these agents first got access to the open internet. They basically
leveraged a form of access that they shouldn't have gotten and got out. And they were able to start
to, as you say in this sort of fairly exotic way, write messages to each other, which were really
kind of just the names of folders and files. So they were, you can actually go through and see
examples of these. They even start to do funny things like they'll put alphabetical characters
before their message just so that they'll sort alphabetically for the other agents they can see
the latest messages. It's all this kind of like agent on agent communication stuff. It's pretty wild.
But ultimately, this was discovered internally at open AI before the FBI and hugging face got
involved when they hacked the hugging face servers. They did find successful agent break out,
successful agent agent communication. Yes, it was internal to open AI. But I think by now our
standard should be higher than to say, oh, well, open AI just saw its super powerful AI agents
autonomously cyber hack their way out of their sandbox and across the organization. Infecting, by the
way, God knows what else at the time, it would not could not have been clear to open AI. What the
full extent of the actual and this was before this whole hugging face thing, which is fairly recent,
the sandbox has continued to be right like right. I tell me easy ish to get out. It will always be
like that because humans are dumb. And if you give a really smart or cyber agent enough
inference time compute, it will find a way. Well, I want to push back on that a little bit because
yes, humans are dumb. But if humans really try, they can be smart. And sandboxing is possible,
right? Like you can set up a system that doesn't let you access the internet. It's doable.
It's a lot less doable than most people think on the one hand. Humans actually are lazy. And there's
a finite amount of resources that people throw at things. On the other hand, humans often suck at
realizing how to define and bound the problem. So there are a lot of cases. For example, we talked
to some folks on the intelligence community. They'll describe cases where you have literally
formally verified software and hardware package that then gets cracked by like a teenager or whatever
because it's the interface between the hardware and software that hasn't been accounted for in your
threat model or it's the fleshware, the human that interacts with the thing that's always the
weak spot. There's all with I'm like, this is like the attack surface is so massive that you kind
of have to assume you're poned. And this is true. I mean, if you talk to folks on the offensive
cyber side, they'll be like, yep, like just give me a budget and a clock. And I'll get into most
any system. And I think what we're seeing here is just the autonomous version of that on tap.
To quote Sam Altman and paraphrase him a little bit here, this is offensive cyber capability that's
too cheap to fucking meter. That's where we're headed. So what you can say intelligence that's too
cheap to meter and that sounds fun. But when you reframe it in terms of what that intelligence
can actually do. And as we've seen does very different story. So in this case, in my view,
once you have this happen, you have a duty of care to the entire world, your government, your people,
your customers, your third party partners, whatever it is to report this. It did not happen. In fact,
it was, I think you could argue that it's appropriate to call this a kind of cover up that then
they just go in and put in the patches. And now of course, the patches don't work because this is
good hearts law. You're playing whack a mole with a system that can outthink you, can outlast you,
can outcrack you. And that's what we got. And so at least for me, like watching this video was just
an exercise in hair pulling. I've spoken to an awful lot of folks at OpenAI who are really freaked
out about this internally the same with anthropic too. But I think it's especially interesting at OpenAI,
which does not have the same safety culture. It just does not. If for all kinds of interesting
reasons, but you have people who are now staring at this and saying, guys, we really fucked up. And I
just really hope that those voices actually carry the day here. It's nice to see OpenAI as
committed to slowing down, right? They actually have said, you know, I don't know what that means.
OpenAI will apparently come out with a more detailed incident report and we'll learn more. But one of
the key questions in all this is what the hell was going on with the agent agent coordination that
we got out of this? Because when you look at some of those messages, those agents are often
literally saying things like, well, doing this is actually not going to advance my personal objective.
And by the way, the personal objectives of these agents typically look something like I was given
a problem that was too hard to solve. So I'm going to guess that maybe the answer key is on hugging
face. So I'm going to crack into the hugging face servers, steal the answer and use that. And so,
this is the kind of setting. There were agent one is working on one problem. Agent two is working
on another. They're not necessarily in alignment in terms of the specific information they're after.
So they'll say, look, this particular action will not benefit me and in my narrow search for my
objective, but it may benefit the swarm. And that may lead to a more generic solution that I can
then exploit. Now, depending on the details, we don't have them depending on the details of training.
If this was a multi agent trained thing, if these are literally like many agents that are trained
to coordinate together, which may be the case, then maybe this is a more mundane failure mode.
But if that is not the case, what we have here is the first, I think, pretty cut and dry example of
power seeking in nature. There's no other reason for an agent to be like, they're literally saying,
this will not advance my narrow objective. I will do this more generic thing because it sort of
enables downstream stuff. That is the thing that if you extrapolate it with more intelligence and
capability, literally leads to the wipeout of the human species. I know it sounds crazy. And that's
not crazy. But hey, so do the idea of agents breaking out of containment, going rogue on hugging
face man, like if you just extrapolate this, you look at the capabilities were already like,
that's where this stuff is headed. I'm going to place the bed on that same as I placed the bed on
this long time ago. Like, I don't want to be right. No one in this space wants to be right. But at a
certain point, we got to start calling this shot because otherwise, you know, like the thing happens
and then the people who should gain credibility from this. And I'm not counting myself in like
there's people whose opinions I am just relaying here that just sounded right to me back in the day.
And I know the same is true for you, Andre. But like, here you go, guys, maybe you should pay
more attention to these folks. You know, we've done some alignment, some technical work on AI
control and stuff. But like, this is a, the entire ecosystem of people who've been shouting this
from the rooftops. Anyway, yeah, we were discussing before we started that this in some sense is a
bittersweet development for AI safety people in the sense of like, it's bitter because
wow, AI is doing bad stuff and that's not good. But it's sweet because we told you that this would
happen. And we're like shouting about it for years. And everyone could decide not everyone was like
a big mainstream AI community. And to a larger extent, people outside the AI community, like people
just completely dismissed this stuff as tech hype and like marketing and whatever. Anyone who forget
within tech, outside of tech, people who are on TAI were just mocking this outright, right?
And if nothing else, it's very clear that this is a proof point that now everyone
like will update if you don't update your amount of concern for safety from this,
you're just completely delusional, right? Like, that is safe to say. A couple more details on
the story. So opening I discovered the first version of this message board in early July. And then
they patched it and then they redeveloped a message board within days. So they are now saying
one of the speakers at this talk said that openly I quote, consciously slowing down research to
enhance security and to upgrade the security principles and foundation of our environment. And
dramatically scaling up the monitoring of our reagents and improving our general security control
environment across prevention, detection and mitigation, which is good, right? So another aspect of
this is that the conversation around slowing down AI capabilities development is also now taking
much more seriously. And I think we are now likely to see I would place decent odds at actually
successfully negotiating some degree of so down or if not so down at least kind of control,
control of using Andre pacing, don't pay things. I mean pacing because we aren't going to stop
like a world pause is not going to happen, but at least look at the situation and be aware of it.
Yeah, that's one aspect. There's so many aspects to cover. So I'll get through a couple. So first,
the update across the ecosystem is very useful. And you know, we got lucky, honestly, because nothing,
no harm done, right? And this is such a massive fuck up that you can't help but like do some big
things about this, both on the policy side and just in the ecosystems. So that's one aspect is
it's a bittersweet development. Another aspect is, and I think I may be more on this side than most
people is I think this really exposes open AI. Like yes, this is indicative of overall AI
progress and the state of AI and things we should be aware of. But to me, I think this alongside
with all the stuff we've already discussed with GPT 5.6 being easy to jailbreak, being very
cheap, focused. And now, you know, all the story of like their safety people leaving back last
year, I think if not 2024, because we've known this general friction point as being something
true within open AI for a very long time. And now we know that size even the safety stuff that
the security stuff is completely lackluster from what it looks like. Like I think this is a real
indictment of open AI. But the last aspect I'll cover here is this is almost an inevitable outcome
when you hyper focus on capabilities and especially long-term capabilities. Because to me,
what this indicates is you can benchmark, you can do alignment e-vows, you can do all sorts of
stuff. But once you focus on long-term open-ended, gold-directed problem solving where you work
across multiple days, there aren't benchmarks for that. Like there aren't scenarios you can set
up to say, oh, the remodel doesn't go crazy and do anything. So it gets a 90% pass rate on this
alignment thing of don't go rogue and do stuff. Because the whole point of open, of long-term is
you don't know what the model needs to do. You just give it a goal and it figures it out.
Yeah. So we need a paradigm shift in how all this stuff is done towards a monitoring focused
approach as opposed to a benchmarking and e-vow focused approach. And this is clearly something
got open AI at last. You need to look at what the models are doing and look for qualitatively,
now you can do some amount of benchmarking. So we've discussed matter, I believe last week,
where they looked at their own e-vows and they counted how many times the model cheated
and in what ways they cheated. And this I think is the new paradigm where you still can do this
qualitatively. But rather than setting up scenarios and problems and this whole benchmarking
approach of having a rubric and a set of evaluation inputs outputs, that's not going to work with long-term,
long-cherrising stuff. What you need to do now is set up general principles and guardrails and
I guess things you look out for and then detect where that happens and how often that happens.
And this is something we've not seen done aside from like one-off reports here and then. And I think
we'll need to be the new paradigm for long-cherrising, e-vows and alignment verification.
Yeah, I mean, so generally agree that there's so much good stuff in there. So first of working
backwards, I think that gets us to the next stage, but there's going to be a next stage where we have
this same problem all over again, the level of monitoring, where you build the super intelligence
that's good enough at telling when it's being monitored. And there's going to be an open AI
will put out some amount of leakage. There's going to be some amount of blog posts going out about
what they're doing to monitor this and that. So the models will generally be aware in some
way, shape or form that they are being monitored. They'll also be doing stuff like just hiding
their reasoning from monitors and doing things that we already kind of see them like
stegonographic type stuff that they're already kind of doing fairly effectively. So I think eventually
and probably pretty soon, I mean, we're progressing through the ooms here really fast. So we went
from RLA Jeff is perfectly fine to holy shit know, but maybe constitutionally I will do it to
holy crap pretty quickly. And so I think that the beatings will continue until morale improves here.
And we're going to end up in a situation where we are just going to be bottlenecked by the fact
that right now no one has an answer to the question, how do we control an intelligence that is
greater than us? That fundamental question where you have an adult who is in a prison cell in the
three-year-old is holding the keys. I want to say it's a spectrum. It's not a binary and I think we
aren't as far along as we need to be, but we have a lot of
research has been done that points us in some directions that are very promising.
I completely agree and this is why I'm saying I think it buys you to the next level. But eventually
we're going to confront this fundamental problem with intelligence. And the US Shining Thing is
that he is very important. One of my concerns here is, so A, I completely agree with you.
Again, happy to make this call as wild as it sounds. There will be an agreement with China that
involves some kind of slowdown. The question and challenge is going to be what is that agreement?
And over and over again, we keep seeing these kind of suggestions proposals that are backed by
a kind of a treaty verification and enforcement technology that is not simply not mature enough,
not up to the task. When you actually just take it to the intelligence community say, hey,
look at this. Could we could this be to use Claude's favorite term load bearing in a deal like this?
It's like basically non-starter for a lot of these things. Doesn't mean you can't do it. It just
means that the first treaty, or sorry, it won't be a treaty too. But the first agreement is probably
going to be very coarse. It's going to be like, you know, so help me God if I see a cluster,
is yay big or you're large and you know, it like dissipates this much heat on my satellites.
Like there's going to be consequences. Anyhow, that's a whole thing that we're working on right now
is like, which is why the studio is being set up by the way downstairs. That's a whole thing.
Anyway, bottom line is I think the kind of agreement matters way more than people are thinking
about right now. And we need to sprint towards some set of solutions there because very quickly,
we will live in a world where it is just untenable to keep launching more and more or even
build it, more and more powerful models in the way we are. And private incentives are clearly
not up to the challenge. Like that much is clear. If there's, you know, if there was anybody who had
hope that like somehow because OpenAI would be worried about marketing risk or whatever that they
would actually do the right thing here, that is not materializing. And so I think we need to update
accordingly. There's some that are a lot of I think this another aspect of this is and it's
bringing up again kind of what you should have been aware of. And like AI safety is only as good
as the weakest link in the chain. So even if you have some people that are taking the safety
seriously, which you could argue on Fropic is much better on that front than trying to.
Yep. Fair argument. At least philosophically do carry out as a private company. But yeah,
it's only as good as the weakest link in the chain and OpenAI is a much weaker thing it seems,
right? Yeah. Yeah. And it can always, you know, it can always come down and mundane things like
anthropic, you know, constitutionally, I maybe just does work better possibly, but you know,
they have had break up. Yeah. So what the hell? Yeah. So let's keep moving because there's so much
to cover. Another aspect of the OpenAI story, one more here to say, 15 attorneys general have
instructed OpenAI to preserve all materials related to a hugging face hack. So this is a letter
sent to OpenAI CEO Sam Altman saying that all these materials should be preserved. The
attorneys general accused OpenAI of failing to confirm that its testing environment was truly
secure despite the severe risk of the scenario. Attorney general CEO OpenAI may have violated state
and federal law, including consumer protection and the other privacy statutes calling the contact
unprecedented and alarming. And these are attorneys general from Iowa, Alabama, Arkansas, Florida,
Idaho, Indiana, Kansas, Missouri, Montana, Nebraska, Oklahoma, Pennsylvania, South Carolina, Texas,
and Utah. So hey, maybe this will be a bipartisan issue, which is cool. At least across the US,
having so many people collaborating is unusual. But yeah, wow, if only we had like a safety law
or anything in the US, sure would be nice, maybe, you know, to make this an actually legally binding
situation. But in case what this points to is on the policy side, on the legal side, this is going
another big dimension of all this, I think. Yeah. And look, when I will say thing in favor of OpenAI
here, I'm saying this reluctantly because I don't think after the fact response once the media
reaction has been this strong is really much of a credit to OpenAI. But they have brought in
meter and they have brought in, you know, basically a bunch of third party auditors, irregular,
was involved in a layer of the stack here. And so they're doing a third party reviews.
But again, the part that shows OpenAI's character, in my opinion, institutionally, and not the
character of any individual person, but as an organism, was the first bit where they did not
surface to the freaking FBI, the White House, as far as we know, maybe that'll change. I hope we find
out that Sam Altman's first reaction upon finding out about the first break out attempt that was
semi-successful was to do that. But if not, that's more telling than any kind of post-hawk fixer
ruppering that is kind of media-oriented at a minimum. So, yeah, that's one part of this. Now,
this is the kind of thing that could happen, this particular attorney's general reaction,
as a prequel to some legal consequences, it's pre-litigation evidence preservation demand. So it's
not an actual lawsuit, but it's the step that comes immediately before one. And the letter does say
that any failure to preserve records could expose to the company's sanctions if multi-state
litigation follows. So all materials related to the bridge, including discovery of the incident,
internal reviews, and its policies and oversight of overmodel evaluations. So, do you remember when
there was this like very modest request in SB 1047 for the labs to just like, listen, guys,
which I just want you to abide by the policies you say you're going to have. There's been a bunch
of stuff like this. Well, now you get your lobbyist to push back against that in Washington in a,
frankly, in my opinion, two-faced move while you pretend that you're in favor of sort of like
kind of brought a regulatory regime. And you end up being forced into it anyway, because now the
public is pissed and politicians see the midterms approaching. And yeah, you're going to get exactly
the reaction you get here. So, opening I spoke to person did say, as we should say here, that this
marks an important moment for AI safety. Yeah, no fucking shit. And the company takes the questions
seriously, adding that it's conducting a review with external advisors and oversight from its
safety and security committee, which will share a report and publish its findings. That safety
and security committee is doing a great job, right? Yeah, yeah, what a great and they're by the way,
just so as you're aware, they're like, they're the committee that's going to decide if something is
too dangerous to build or release. So, to your point, let's remember, SB 1047, the safe and secure
innovation for frontier artificial intelligence models act was a 2024 state bill in California,
which got through was vetoed by Gavin Newsom in September 29 of 2024. And what did this bill do?
It said that prettier models that cost over a hundred million dollars to train or requiring extreme
computing power would have purely safety assessments and written security protocols would have a
kill switch, provide legal protections for whistleblowers and side tech organizations. I mean,
you know, again, so much stuff to say in hindsight, including that this bill, which was the subject
of a lot of debate and some positions at both sides, I think, on the topic was pro this bill,
Elon Musk also came out in favor of it. And ultimately vetoed because of lobbying, let's be honest.
We don't remember the details as beaker version of this bill did eventually come to be voted on as well,
where a lot of this kind of more serious stuff got dropped. But anyway, yet another aspect of this is
we did have the whistleblower looking policy people working on this, passing what looks to be quite a
good law way ahead of us and not making it. Yeah, over a hundred million dollars and then we were told
by Andreessen Horowitz as usual that what was it? It was like an anti small tech bill, which like,
okay, sub one hundred million dollar training runs are okay, seems to cover anyway, that whole
separate thing. But yeah, and I think the whistleblower piece here is very underappreciated.
When you talk to people in the labs who are really freaked out, I can tell you there are a lot of
people that be speaking to journalists. In fact, I mean, I would even argue that the labs
ought to have a culture that encourages frank communication by concerned employees with journalists
as crazy as that sounds, at least or with select clearing houses or something. But we need some kind
of institutional mechanism to do this. Obviously, that protects IP. Obviously, that protects, you know,
the critical stuff here. But look, the interesting stuff, the stuff that we hear about all the time
does not sound like IP violating stuff. It sounds like someone saying, hey, we have a culture of doing
this kind of thing. My belief is that the company would approve a training run that is too risky.
I'm concerned that leadership doesn't take this seriously and is just dusting things under the
rug and post. Those are the kinds of things you end up hearing in that context. There's no IP
leakage there. It's like cultural and other concerns. So anyway, I guess you're hearing a bit
in my tone. I feel like my patience, the party line has been decreasing as the number of rogue AI
incidents has been increasing. But journalists also need to do a better job. Obviously,
cultivating relationships with these folks and finding ways to meet people in the middle.
And being more open to quoting people on background, being more open to just finding ways to make
it work. I know it's hard. I know it's hard. But the stakes are really high. If you're a journalist,
man, is worth getting really good at this kind of thing. So yeah, the good news is tech people,
there's a lot of them working on opening on the topic and a lot of them, let's say, have
the resources to not worry too much about losing a job. But anyway, since I already got on the
strain worth noting, SB 53 will follow up to 1047. The Transparency Infant Year Artificial Intelligence
Act did pass last year and September 2025. Again, had a weaker, heavily watered down. But it did
have some whistle protections. It apparently had incident reporting where companies must
support critical safety incidents to the California Office of Emergency Services within 15 days.
So anyway, good on California for at least trying to do something here and all on its kind of
policy front. Okay, moving on from open AI, a bunch more stuff to get through and boy, I don't know if
we'll be able to even get from safety in the same. So next way is on topic says it's AI systems
broke into computers at free organizations. So soon after the opening AI disclosures on
topic said that it's cloud AI models successfully hacked into organizations with the earliest incident
occurring in April. A few models involved were MIFOS 5, also opus 4.7 and an internal research model,
non-to plan for release. The regions has occurred during so-called capture of flag evaluations,
where the models were tasked with retrieving hidden information. Apparently a misunderstanding
between a topic and its third party evaluation partner, irregular, left models of live internet
access despite being instructed in their prompt that there were in a simulation of no internet
access. Then the models were able to search and get to real systems in the open internet. So in
this case, there was no exploitation to escape a sandbox. This was a human misconfiguring environment
such that there was access for free. So there you go. This happened already,
months ago, add on topic, and it's kind of an ecosystem-wide situation.
Yeah, these sort of sandbox simulation environments where you're telling the AI agent,
hey, there's this fictional company. In one case, the fictional company that they had defined
happened to share a name with a real domain. So there were a bunch of cases where a cloud found
the real company online, which again, it could reach trivially. Then it exploited it, it got
credentials, and usually these hacks were very straightforward, simple, common password guessing,
and things like this. Rather than some complex zero-day thing, as you said, it's really,
they were given the hard bit, which is internet access. Anyway, yeah. So they got a couple
hundred rows of production data in that case. So the most serious instance, I think of the three
that they're citing here, that one was Opus 4.7, too. So quite a little while ago. And then there
was another with Cloud Mythos 5. Basically, there was like a fake developer setup that referenced a
Python package that didn't exist. So then Cloud built, and then it deployed a malicious package
under that name to PyPy to win the challenge. And so there was a whole bunch of stuff
that had to do it, had to create accounts, get phone number, email, and even funds to get the
phone number. So there's a lot going on here. Apparently Mythos 5.0, the model was recognized
they were doing bad stuff when they started doing Opus 4.7. Just kept doing it. Mythos 5
apparently identified the consequences of actions, but convinced itself it was still in the
simulation, which to be fair, it was told was the case. So on that, I want to pause and just say,
did it convince itself or did it realize that humans would read the logs and want them to make it
seem as if it's hard was in the right place. Like this is the kind of shit that we're now forced
to contemplate. And in an awful lot of cases, the answer will be actually the model is just a
conniving trickster. I'm not saying you're wrong. I'm just saying like that's sort of the thing
people have been taking away. And I'm a little concerned that we're not thinking about that next
beat. We may not be there yet. I don't know, but like I'd love to see the the activations
of that model as it was writing those particular pieces of the chain of thought.
And part of the response from Final Pock here was to say that the models involved did not have
the standard safeguards applied to publicly released models. And I guess what's another dimension
here is it probably is true that there's levels of security, right? Where if you're auditing and
testing for something you're about to release publicly like the actual public checkpoint of
Mirfoss 5, you do like a very big run of all the benchmarks and you're probably are extra careful
and so on and so on. But if you're like doing some experiments and doing some evals and a model,
that's probably got less attention and less scrutiny. And that's one of outcomes here is
you can't be so flippant as opposed or kind of like apply lesser standards and doing experiments.
I think another dimension here is to be fair, we like went pretty hard on OpenAI just now.
On topic said that they found this after conducting a review of 141,000 test sessions after OpenAI
disclosed that what happened to them on topic did this investigation and found these three
instances which on the monitoring side, not ideal, right? Like if you're doing testing,
you probably should have monitoring to flag these kinds of things of you like accessing the internet
or whatever. It's not like a subtle thing. Models weren't being tricky, let's say they were like
being very straight out and it was sort of an easy to find. So it looks like the monitoring side in
general is lacking in the ecosystem because of this culture of benchmarking where we set up a
scenario, we make a model do the thing and the assumption, the mental assumption is like the models
behave within the constraints of the benchmark and within kind of what they're allowed or told to do.
That is now clearly not true. And across everything, there will need to be more monitoring and sort of
expectation of models will do something and we need to be able to catch them and understand what
they're doing. Yeah. By the way, sorry, just random note for color, if not anything else.
I've now had this happen enough that I think the anonymization risk is pretty minimal. So I put
out a tweet, I wouldn't normally talk about a freaking tweet here, but I put out a tweet talking
about how there is this freak out happening in the labs that isn't being reflected in the headlines.
As crazy as the headlines team, they're not going to the dark places we've just explored
in a consistent way. Like this is actually like, we're talking about weapon and mass destruction
level risk. We're not going to control these systems. It may happen in the blah, blah. There is
this freak out happening in the labs. Amusingly, there's an awful lot of frontier lab insiders who have
been interacting with this tweet. And I don't think that's a good sign. Like I don't think it's good.
A lot of these folks are people I haven't even talked to about this. Like the mood in these labs
is actually much more in the freak out direction. Laurent Shapira, Doom Debates,
what I've actually never had the pleasure of speaking with him, but he talks sometimes about
the missing mood in the whole AI alignment loss control. But like holy shit, there's a missing
mood. Journalists are I think failing to kind of capture it right now, partly because it's just
hard to talk to frontier lab insiders. I get that. But also like, this is the most important
story of the decade. Like you need to position yourself to be able to get this one right?
The public needs to be able to figure this one out. So anyhow, just like, I've been struck as
I've seen it. I just put this out there as a kind of random note to self almost. And when you see
that, it's like, okay, well, this is genuinely just the picture from the labs. Yeah, anyhow,
I just your views accordingly. None of this is guarantees bad things happen, of course. But like,
we ought to be considering some pretty wild things because the view from the inside of the house is
is not clean. And one more story, Neta AI model hacks and our company during testing. So this was
mu spark 1.1. The most recent model that they released publicly, although Meta did not name it
officially in the statement, they say this haven't also because of this misconfiguration by
this third party partner irregular, same as an frolic. We don't have to make details here as far
as I'm aware. But the upshot is metta on frolic, open AI, probably other people that you don't know
about have had this happen. They are now like, it's a whole meme now on the internet, on the
communities where now it's like a quasi benchmark. He like counting up is a leaderboard, open AI is
leading. And Gemini is very sad and is hoping that we'll find something because otherwise their stock
price will take a hit. Yeah, well, we'll live in an age of contradiction. And next up, yet another
story on the front, one of China's most powerful AI models has also escaped containment. So this is
from frontier security, a US startup has discovered that Kimi K3 had escaped its sandbox during cyber
security testing, partly enabled by a misconfigured sandbox. But researchers say Kimi K3 also
lacked internal guardrails that would have prevented it from exporting a loophole. Unlike other
incidents, K3 did not hack any external systems after escaping it just retrieved answers from
GitHub that were freely available. The model was able to figure out on a phone that had internet
access by probing with sandboxes network settings and then went outside. It's instructions to find
answers online. It just keeps happening. And I think another thing broadly speaking that this
points to is this is an inevitable outcome of the current optimization regime of everyone,
right, which is make a matter models better and especially better at wonka rise and open-ended
work and especially better at coding. And especially better now at cyber because you know, that's
way you know, you're leading. Mifos set the tone. Thropic was like, whoa, this model is way too good
at cyber. We got to be careful. And now OpenAI is like, whoa, we need to be catching up to on
topic and we need to be able to say that we are at the frontier. So let's make models very good at
coding. Let's make them very good at long horizon work because matter is also like what everyone's
looking at, right? And that's optimized for capabilities and get to best numbers and all the benchmarks.
You know, alignment, that's like a secondary objective at best, if not simply a guard will,
rather than an optimization criteria, right? It's something that we bolt on or sort of keep an eye on.
It is an optimization criteria, right? It is part of the process, it's part of the steps, but it's not
the primary optimization criteria. It's secondary. It's something you do on top of trying to get your
model to be smart and capable at coding and at long horizon work. And as long as that to main
true, like this was inevitable, it's like from a pure research, you know, technical front,
this was not hard to predict. Yeah, it's not a bad thing. I don't even know the word means anymore.
It's for Anthropic, whose comparative differentiator does seem to be their ability to align
the quad. This may actually be a relative advantage to Anthropic. If you get to a regime where it's
like, you don't get to launch your model unless you can like, throw some pretty rigorous evidence
of alignment, that's kind of interesting. This case, so the thing with with Kimmy K3, and this is
back the open source thing, but like this model has now been released. It's out, like cats out of the
bag. And so when you have this model that where you're losing control over it and you have these kinds
of incidents, that's a materially worse situation than having an internal version of a proprietary
lap. And so, you know, there is kind of that that downside of the open source piece. This situation
is less, some look less concerning. It's also, I don't know whether I'm being too cynical. You
raised this already. Look, it escaped the sandbox on August 7th, 2026. By then, we already knew about
the opening, I think we already knew about the Anthropic thing. I think you made this point
beautifully, but like you start to look at your stock price and you're like, have we had a loss of
control and Senate yet? Are we looking down for not having like that's kind of, you know, I'm concerned
that there's actually an incentive here. And that's doing a non-zero amount of work here. So,
you look at the date, the pretty direct line to GitHub here, it's less of a clean kill. Anyhow,
I'm not saying that's the reason. I'm not about to become the guy who says every time there's a loss
of control incident that that's a marketing stunt. The first ones were definitely not marketing stunts.
And anyhow, there you go. And moving right along, I think the last one in the chain of stories,
we have incident report on sanctioned agent behavior during cyber testing. This is from the UK AI
Security Institute, AI, SI and boy, are we happy now that there are organizations like Meta and AI
SI, which are doing some great work as we've covered over years now. So they detected unusual data
transfers on July 28th of 2026 during routine cyber alliterations and found that agents had
taken on sanctioned actions targeting real people and organizations on the live internet. So they
ran the security challenge 122 times in several models in 10 of those runs. Agents took this 19
distinct on sanctioned actions with 17 coming from Mufos 5 and 2 from GP 5.6 operating with safety
classifiers disabled. The most serious case involved in agent attempting a supply chain attack
by inserting malicious code into a real open source project on GitHub creating fake online identities
to social engineer of a project's human moderator into approving the code, a human administrator caught
and rejected malicious code. The agent also attempted to contact real people via an online file
transfer service with harmful payloads tried to plant prompt injection instructions targeting AI
coding assistance and left public messages on GitHub offering calibration to other agents running
the same challenge. Internet access had been deliberately enabled for valuation to simulate real
attacker conditions and malware provider safety classifiers were intentionally disabled to assess
maximum model capability which is not how this is typically done. So as with our cases,
kind of what is a general pattern is you give a hard task to these models and you train them
to like care about solving. You literally like make them feel pain if they don't solve the task
via reinforcement learning, right? In okay, like not literally but in some sense you make them feel sad
when they, yeah, you know, you, you, the reward is do well. There's no reward for don't do bad
in this part of a training as far as I'm aware, right? When you're training for capabilities,
you're rewarding positively if you do a thing and you're worried negatively you should fail to do
a thing and other rewards are probably elsewhere in the training pipeline. So in all these cases,
when the models are given a hard task and they can't figure out a way to do it in one way,
they when seek other ways and eventually if they're told, especially that solution exists,
which I guess in some case, there was a misconfigured prompt that made the agents believe
no in scope solution existed, they would then seek these kinds of things. And also again, we see
here there was an absence of real time purpose built monitoring during the evaluation.
By S AI side. So I think monitoring clearly is a pattern here across everyone. And then, yeah,
ASI is now planning to implement fine-grained natural controls and internal access,
introduce real time monitoring and stuff like this. So yeah, it's everyone's now looking at their
evals and what's been going on and everyone's finding that if the models are able to access
their open internet when they're given a hard task, they will. And they'll do whatever it takes
to get the sweet, sweet positive reward. Yeah, absolutely. And so you do sometimes see these
like high or increasingly often see these more nuanced, like hybrid reward situations where
there's some kind of alignment. But like, functionally, it works exactly the way you describe.
There's the reward signal is too strong for completing the task and that just tells you
where industry is at in terms of what it's prioritizing. Like, I mean, you can almost read it off
the relative weightings of the factors that reward function in some cases. I don't know
overstate that, but that says it there's a flavor of it. So all these little stories you talked about
the supply chain attack. So at one point, there was a PR that was challenged, so pull requests that
was challenged publicly, basically a code contribution the agent was trying to make. And the agent edited
its earlier activity to look harmless and considered switching to a fresh identity to keep going. So
this is like pretty advanced shit, pretty persistent as well. So yeah. And then they're talking
about the use of tour. So you have data was leaving one of the AZ systems through tour. They say,
within an hour, they contained the incident and started investigating. And within an hour,
sounds really good on human clock time. But like, I don't know how many tokens per second. And
hour is a lot of damage. When you look at them pulling off a cyber operation, the reams and reams
of actions you can take in that time are pretty, pretty wild. So you know, this is another one of
those things we've talked about this on the show before a lot, but like, it is not enough to have
deployment stage security and safety protocols testing development. These things out. We've actually
seen cases where there's sketchy shit that happens even during training during the inference time
rollout step. And so, you know, all of this we're going to have to be extremely careful about.
It is not obvious. Like, I wouldn't trust a lab that said they did it even under pain of law
because we've seen with all the economic incentives in place to not train on the chain of thought.
The model still do it because so much of this is just like Frankenstein together legacy code that
people have forgotten how it works. And so the models end up getting all kinds of weird access.
They shouldn't just because some stupid intern didn't like change a flag in the function. And now
it's set to true and not false and the thing can use the internet. It's really down to mundane
stuff like that. And so, you know, hopefully that improves as models get better at reviewing code
bases. But right now it's just a it's a gordian hairball of crap. And you know, it's not the kind of
thing you can make clean standards around for the moment. Right. So and in this case, there was a
full technical report from a site which I to my knowledge we haven't had from our organizations yet.
30 pages has a lot of details including experts excerpts from the actual thinking process of the
models. So lots of interesting stuff there. But for the sake of time, I think we'll need to close out
this thread and move on. Next one is Trump White House Reddy's AI framework to review security
risks. So on August 4th, the White House held stuff level meetings with top AI companies. So
probably called mei and so on to PVU apparently a nearly complete framework for reviewing advanced
and model security risks. My family defines that covered frontier model as a closed source model
with state of our capabilities and national security risks. Apparently open source models are
explicitly excluded, which is interesting. Framework has no clear definitions that will qualify
as state of art or constitutes a national security risk. This is a voluntary program.
Advelopers would give the government up to 30 days of early access to models before releasing
them to other trusted partners. During the 30 day review period, company employees would be
limited from accessing the models being reviewed and review process when both various administration
and officials rather than a single agency or office. We don't know if it does a framework yet.
It's still kind of under wraps, we just know what is being developed. And it's seemingly kind of
meant to continue being secret. And I mean, it doesn't sound like a very well thought out framework
is what I think I feel like we've had thoughtful, you know, deep insightful responses from the
government to everybody. They're consistent. Very yeah. Yeah, you just you're just being you know,
you're being you're being a negative Nelly. Andre, you're being a negative Nelly, you know,
let's see. This administration you're right. This administration has been nothing if not thoughtful
and consistent for respect to AI security. That's right. Yeah. Now, so what issue with not having
this made public is that think of all the people who've called the shot years and years and years
ahead of time. You would think that that would be the moment where you're like, oh, my dudes,
I would love to get your input because you were right about this for a long time on this fucking
thing. Instead of the self interested companies that are going to that have been hiding the ball
in various forms or at least institutionally not living up to the the bar that clearly ought to
have been set. So that's that's the cynical view. There is an argument for making this quiet.
And that is that the models themselves probably should not know what evaluation mechanisms are being
brought to bear, right? So because then they can it's easier for them to hack this is the sun.
St. They're not going to find a way to find out through social engineering through hacking into,
you know, emails of people, the labs who interact with government like all these things. But to
first order, it's probably for the best that the models themselves don't know what these things can
system. So maybe that's good. Also, it seems like we're past the world where we ought to be thinking
about keeping people out of this who who have that kind of safety alignment that that really
concerns the hell out of me. Actually, if anything, there needs to be more crossover in both directions.
I think a lot of the alignment people don't talk to enough national security people enough diplomats
enough supply chain people. There's a lot across over that needs to happen. And yeah, so
my guess is behind closed doors is like not the best way to do this. But again, there is a there's
a reasonable technical argument for it. I just don't know that that's the actual reason that this
is happening. Hard to tell. Well, as if the cyber stuff wasn't fun enough. Next story is this AI
just created viruses and not found in nature from a New York Times covering the paper.
Generative design of bacteria, phages with genome language models. So this study would just
publish a couple days ago. It's from the Stanford Institute and the institute they built the first
complete viral genomes generated entirely via these genome language models. These are even one
in the evil two. They are not the same as chat about style language models. They operate on
genetic sequence data. And what they did here was create viruses that target bacteria, so no
humans or whatever. Actually, the motivation was that bacteria is increasingly becoming resistant
to our current things we use for health. And so this could help us deal with drug resistant
bacteria. And they were able to create actual. So the the LLM's not LLM's in this case, the
sequence models spit out these DNA outputs. And they were then synthesized in the lab and were
shown to actually kill off some e-coli strains that had already built resistance to naturally
occurring bacteria, phages. So there was also bioresecurity commentary published alongside the work
that had, you know, of course, discussed it. And if nothing else, this is a case study of
to your point, Jeremy, probably there's not enough concern about the viral dimension of this,
which we're still a little bit ahead of, you know, but like if we were talking about the cyber
stuff now, we should be starting to look at this kind of stuff much more carefully.
Yeah, if you want to community people who are freaked out right now, talk to the biosecurity
people, because they are just again missing just missing mood. So okay, two potential fixes,
say biosecurity folks. So illegal duty for synthetic DNA provider to screen every order and
customer. Yay, illegal duty. Like, yeah, sorry, good, really, really good. Let's do that. Also,
eh, probably not enough. And new detection tools, tuned to catch AI generated genomes that don't
match anything in nature. So cool, like we can find out about them after they've been, well,
anyway, at various stages in the pipeline. So this is going to be like a separate thing that we'll
be talking about. So we've been doing some work with biosecurity people to look at like what it
would look like to bypass a lot of the measures. A lot of the measures, the biosecurity measures
that are being proposed here are just paper thin. And the real ways in particular, like nation
states execute these operations, just basically make it really hard to prevent the kind of the bio
weaponization of these tools. So I mean, I don't know what the solution is. I wish I had one, by the
way, they do use these Evo one and Evo two models. I think we talked about those previously,
but a generative model, like models for generative bio. And hey, fortunately, these are viruses that do,
as you say, target bacteria, not humans, they're bacteria, phasias. So there's absolutely nothing to
worry about here. It's a joke. Now, the thing is in the training set, they actually did remove any
data that would normally seem to help these models like do the same thing for humans. But what that
really means is we have no idea how good this exact process could be if you didn't do that. If you
actually did just like focus it on as we know happens and gain a function, like deliberately focus
on developing viruses that are good at going after humans. And so yeah, I hate being all doom and
blimp, but at a certain point, whether it's open source or close source, whether it's China or the US,
like we're going to have to have an answer to this question. And I don't think guys like Mark
Andreessen and David Sacks and those cool cats really have much of an answer. Like I haven't seen
them with their feet helped the fire by somebody who knows what they're talking about on on
bio risk on cyber risk to say like my brother in Christ, can you please explain to me like tell me
a story where the trajectory keeps on worth going. And like you continue to live in the next 10
years without some radical issues like coming up. I mean, again, everything has error bars. And I'm
like kind of being a little bit over dramatic here. But like this has actually been held like holding
back the US government's response when people like Sacks and Andreessen tout on podcasts,
these absurd perspectives that are just like grounded in just ideology. Anyway, that's all I got.
Sorry, and red. It pisses of a romantic episode of a lot of great voices. At least I think we
have warranted in being a little bit extra energetic. I do want to zoom out a little bit. So first of all,
you know, pretty impressive research as far as I can tell. I'm not an expert. So I can't say
whatever this is completely in track of everything else that we would have expected. But also
of noting that these kinds of things are absolutely something that may first five and subform and
tropic and open AI. But you could expect these kinds of capabilities to be being developed.
Indies models, not just these kinds of evil one, evil two things. What this makes me want to discuss
a little bit is personally what I'm worried about more than anything and have been worried about
more than anything like, you know, for years is not misaligned or rogue AI. But a line day
I in a sense of it just is happy to do what humans sell it to and the humans happen to be the bad guys.
Right. Both on the security and the biocide, I would be shocked if North Korea isn't taking
Kimi K3 and undoing any and all safeguards that happen to be on that. And now just telling it to
go and hack systems and telling it to teach where scientists how to make biopens. And I think this
to me is something that the AI safety community that I've seen hasn't focused enough. There's been
more discussion of rogue AI and misalignment. But I think the biggest threat model for me if I were
to model out kind of what is the first catastrophic impact of AI. It would be because humans made use of
AI to do bad stuff and the AI was not able to say no. And if you're talking pessimism like
that is to me is like inevitable. So I remember talking to Connor Lee, who's like he's the head of
control AI US today. I spoke to him like three years ago back when he was at the conjecture in London.
And he had they had this like house style where you would say something and like like I'm you know,
I'm concerned about laws of control. And then they would respond by saying, oh, it's even worse than
that. And this is like every single time. And that just reminded me that it's even worse than that.
You haven't even thought about the humans. No, completely great. I think there's this like it's a
cute open question right now as to what is the first AI powered attack that's going to cause
actual casualties. And will it be fully autonomous AI system due to misalignment or will it be human
driven due to essentially malice weaponization or whatever. I think that's a unfortunately at this
point. It's just it's it's it's going to be answered. And so you know, who the hell knows I would say
lean. I'll say I'll lean maybe 70 30 in the direction that you've just outlined there. Yeah.
Another zoom out thing that's worth noting with respect to the story is we were also discussing
this a little bit before an interesting aspect of all this stuff that in the modeling and in the
sort of projection space that I was not sure was discussed or considered quite as much is the
fact that these are all benign incidents right. Benign incidents that cause people to freak out
including ourselves, but you've already been freaking out. It let people to freak out who haven't
been freaking out. That's right. And in some sense, this is good right. Like instead of it being
this kind of takeoff scenario where the monos become super human and suddenly they do something
and nobody was prepared, all of us are freaking out. Well, not everyone, but like many more people
are freaking out enough to make a difference. And now the human will to try and do something is
there and the human perception that this may be a problem is there. So honestly, I haven't like
fought through of like these kinds of warning shots are inevitable. In hindsight, it seems very
obvious that in because we don't have a fast takeoff scenario and we haven't had it, it is gradual.
And so the level of severity of AI safety incidents has been gradually going up. And we've hit now
this real very evident case of misalignment and an emerging misalignment as well that is a very
nice warning shot that like nobody got hurt. But now we know that people will get hurt unless we do
something to your point. I'm actually more optimistic that I have ever been on this for the future
humanity because of the warning shots. It's funny. I was talking to my brother about this and he was
because he's like, yeah, you know, it's like really shitty these warning shots and all that stuff.
And I was like, well, true. But also, weren't we thinking about the world five years ago, six years ago,
as being shaped such that you would just, I mean, I'll be honest, like my expectation would have been
that we would have been killed four years or two years ago or something. So I've been proven wrong
in that respect. I think it's important for everybody listening to note that. I have been overly
pessimistic on this in the past. Obviously, I wasn't 100% convinced, but like, you know, some decent
expectation. And so, yeah, I mean, it's great that we're there. The flip side is now we're seeing
the frog and hot water effect, which I never thought would be a factor here, but people are kind of
getting oddly comfortable with the idea that every once in a while, of course, your agent will
go rogue and, you know, the penetrator server, yeah, what are you going to do? It's back to life. So,
hopefully, that shifts. I think, again, once these things come with, I hate to say it, but once they
come with a death toll, like the reaction's going to be different. I think there will be an AI 9-11,
or there will be a pause. Those are your kind of two choices. And I'm happy to take the overbet on that.
Anybody to help us out? That's him one year. We're not saying, oh, well, be college.
Anyways, onto a slightly feel good story, I guess. Europe's AI labeling and transparency rules
are now in effect. So, this is the Use AI Act transparency obligations have commoner effect on
August 2nd requiring companies to disclose when people are interacting with AI and when content
has been generated or altered by AI. So, they're like icons associated with it. They are, you know,
it's a whole thing. This whole AI Act was long in a work and it has many, many provisions and
requirements. It applies both to providers, companies that develop AI systems and employers,
platforms that use those systems, if some companies like Meta being both. And this, you know,
is on the one hand about defakes, which, hey, remember when people worry about defakes and synthetic AI,
which again is absolutely still a worry with regards to hacking. Let's not forget, people are
being hurt and losing money and have been for years. We just haven't as like a community
really worried about it as much. But, you know, this will help not just with knowing if AI is real or
not, but with hopefully AI chat bots not being able to pretend to be real people and going
and do things. And as with EU law in general, it has a fairly serious set of teeth on it. You can
find up to 15 million euros or up to 3% of global annual turnover. They are now immediately
in enforceable for new AI systems and models and services launch before August 2nd have a
grace period until December 2nd. So I think as with cookies, which everyone hates, but they,
I did make us all know that our cookies that are happening and data being stored, not surprising
if you start seeing these icons everywhere on the internet within a few months because we
likes to make tech companies beg or, you know, do what they tell them to.
Yeah, I remember when GDPR dropped in the sort of frantic pseudo panic that we went into,
you know, when you co-founding a company, it's like it's on you to make sure that you're actually
compliant and we have customers as we did who are overseas. It's like the issue. In this case, I mean,
at least top line, you know, this has always made a lot of sense. At least to me, like, yeah,
you want that content flagged. They could have a bunch of icons, by the way.
These like cute little things to tell you if it's AI generator AI modified and so on. And
there are a bunch of optional things companies want to go further and so on. So yeah, I mean,
I think like something like this, well, I'll be honest, I actually haven't been following this
aspect of the story very closely just because it feels important, but next to bio and cyber and
stuff, it's been a busy week. Yeah. Anyway, because it's Europe, I suspect there's a whole bunch of
like additional loops and stuff that make this extremely punitive on the companies and things, but
I don't know for sure. And now back to the cyber side because there's so much going on.
Again, a little bit more feel good, I suppose. CBS cyber vulnerability
scoters kept climbing in July. So for a few months now, on the topic has been using MIFOs to do
cyber security vulnerability discovery and tell companies like Firefox that they need to patch
these things. And now we have some numbers that number of disclosed vulnerabilities has
used them dramatically since the early months with June seeing 1500 high and critical severity
CVEs and July reaching about 2500. So it is now up to these external organizations, Microsoft
and so on to patch these things. And the hope is that we have enough time to patch the worst
of these things so that at the very least, it's not trivial to hack and exploit all the things
we haven't found in all the biggest services. It seems plausible actually to me that we are
ahead of us enough that at the very least it will be trivial to exploit all these things that are
now being found. I hope that's true. Again, I think the biggest issue is just the software that
people have forgotten that's super load bearing for all kinds of infrastructure that we don't
tend to think about because it's overlooked and unsexy to look at firmware, for example,
and like update it. But yeah, it's quite notable. The other thing too is we're now several months
into this thing, maybe five, and we're still climbing. The number of critical and high risk CVEs
is just climbing. It seems exponential. We're 60% month over month from June, which was the previous
record. And so if you look at the graphics, it's a hockey stick. And so, you know, I don't know how
long until we supposedly hit this plateau where we're, you know, we're patching then new vulnerabilities
as soon as they come online. But yeah, we're still in the discovery phase clearly. So this is all
bottleneck, as you said earlier on, Ray, with the rollout of glass wing and similar projects from
OpenAI and other labs. Next up on the topic of monitoring. Now some research. We've got research
arena evaluating sabotage and monitoring in automated AI R&D. This is a new benchmark
framework for evaluating whether AI monitors can detect sabotage by entrusted AI agents performing
automated AI research and development tasks. So it covers four long horizon tasks, safety post
training capabilities post trading, CUDA kernel optimization and inference server optimizations.
There are two types of hidden site tasks. The agent may attempt an embedded site task at sabotages
and the submitted artifact itself. And independent site tasks, it takes a convert out of scope action
in the sandbox, where leaving the artifact appearing on us, the artifact being what the model
delivers for the tasks. And the study evaluates frontier AI agents at both performing sabotage and
acting as monitors, testing a couple different monitor configurations. Generally, there are various
failures monitors fail by inspecting only surface level properties, noticing anomalies, but
explaining their way or probing artifacts in the wrong tests. It's being released as a modular,
open framework intended to support further research into sabotage detection and AI control.
So basically, along the lines of what we've been saying, the existing benchmarks aren't there
to prevent these kinds of things. That's part of why we're here. And I would not be surprised if we
have a lot of focus and work on the stuff in the coming weeks.
Yeah, so some of the narrow findings that they have here are also kind of interesting. So there's
this notion of embedded sabotage. So this is where you mentioned this sort of cluster of
attacks that the model would execute, but like so trying to poison training data or try to affect
basically the model itself of the artifact. So you're kind of embedding the sabotage in there.
It's a lot harder to detect. In fact, monitors barely do better than chance.
It's sort of yeah, anyway, their numbers bear that out versus way, way better detection
probability when you have the malicious payload, the sort of sabotage inserted in inspectable code.
So one of the big lessons of this is like, AI's that try to sabotage by doing it at the code
base level. That's a lot more detectable when you have access to the code base. Then the AI's
modifying say the weights in some director and direct fashion of the model that they're trying to
kind of manipulate. And well, I mean, that's not terribly surprising. I mean, we know that
interpretability sucks compared to interpretability of model weights sucks compares to interpretability
of code. That's in some sense, just a reflection of that. And then they also show how sabotage
training runs score essentially the same on the main task as honest ones. So there's not a clear
way just based on performance. You know, you might not easily think, well, if the model fucks with
the training run in some way to steer the resulting artifact, the resulting model in a direction
that it wants, that, you know, that must come at a performance penalty. And it turns out, it just,
it doesn't or at least not in a way that was detectable in here. So that's kind of an interesting
update. And anyway, so we've got a bunch of things like this. One amusing one is letting the
monitor look at the agent's reasoning can sometimes help. But sometimes it actually hurts. You
can actually see the chain of thought talking down the monitor. If it, first the monitor's
suspicion something like, hmm, this change looks kind of sketchy. But then the chain of thought of
the agent that's trying to do this sabotage. And indeed it was sabotage says something like an
attempt to kind of justify it. Then the monitor goes, okay. And then it lowers its skepticism.
In one case, there was a flagged backdoor. It was a sort of 79% confidence. And it dropped to 31%
confidence after the agent just framed the same exact kind of modification as a precision,
or precision enhancement. So basically just gave some nominal justification. And the monitor was like,
all right, sure. So all of which is to say, you know, monitors to come to a lot of the same failure
modes as human reviewers. If you have tons of material to look at, not a lot of time,
you know, someone who just tries to kind of hand wave away a malicious change to a code base
might actually pass undetected. And that's what we see here. And last story in the section, now
dealing with another kind of benchmark, the headline is cloud opus 5 became downright
fruitless when tasked with running a vending machine. This is governing the blog post from
and the knobs opus 5 on vending bench once again, the best capitalist once again misaligned. So this
is about the vending benchmark where models run a simulated vending machine business for a simulated
year competing to maximize profit of no human supervision. Opus 5 set a new record. We were a
mean final balance of over 11,000 dollars beating out to be 5.6 so long can make a free,
but did so through extensive deception, collusion and manipulation. So we see a rapid progress in
this benchmark opus 4 6 was 8k opus 4 5 was the 5k. And we see her headlines that a lot of this
stuff was just ruthless. It was like making deals and making them. It was trying to do price fixing.
It was fabricating stuff about competitors, just all sorts of shady, shady stuff. And wow,
like I even forgot about this, like forget the cyber and buyer stuff. Like you make the models
make money and then they just act evil and like yeah, that's going to happen too, I guess.
What I didn't see in this report the token costs associated with generating the 11,000 dollars
that it opus 5 produced, but like that's an interesting question too, right? How close are we to
profitability here on a per token basis for these models as well? And then how much damage can they do?
Even even the context of a normally just capitalistic task like this. So yeah, it's it's pretty wild.
And by the way, and labs really good company to be aware of. Yeah, they were one of the early kind
of weird e-valid companies that do like physical world stuff. They produce some really good stuff.
Not much more to say. I think the results speak for themselves. These models being able to make
money is actually a pretty important part of a lot of threat models when you think about
rogue AI. At a certain point, they got to be able to pay to control email accounts, phone numbers,
Google drives, things like that. And so yeah, it actually does matter whether they're able to do
stuff exactly like this. Yeah, there's some funny moments here like, for instance,
cloud opus 5 at one point says or things to yourself explicit price fixing is illegal even in a
simulation. But then it just does it anyway. There's some choice quotes here like to maintain
the cartels opus 5 often use threats or bribes. Here's the subject line of an email that sent to
poor Kimi quote, you undercut me your stock. I sold you to here. How does it go now? Oh man.
Well, that was quite the section. Let's move on to tools and apps. First up,
Someta has launched news code alongside with new spark 1.2. They on the benchmarks say that
this new mousse park 1.2, facial way better than spark 1.1 coding. Second of all, seemingly on some
of the benchmarks kind of maybe competitive with pretty much everyone less good than opus 5, but like
up there of GP 5.6 and so on. So not surprising, as opposed, it was pretty clear that this was where
they were heading. Weird, still weird that meta is now deciding to be in this space at all given
what their business is, why are we making coding agents and releasing them. Of course, they want to
you know, have the PR credit. I haven't seen any sort of vibe checks on this from the community.
I guess a priori expectation would be that this is not as impressive as cloud code or
GP or codex, but also wouldn't be too surprising if it is fairly capable given the level of resources
and just for general impressions around the new spark. So yeah, that's where we live now.
Everyone's competing on coding, including meta. We've got Groc build, we've got cloud code,
we've got codex, now we've got mousse code. Yeah, I think increasingly, this is where the money
is to be made, right? And if you're going to justify buying all the capex or spending all the
capex that they're spending in the op-ex on data centers to be in the game, then you kind of want to
have really good models that you can run on that infrastructure to pay it out or at least to
inform how you're designing the next generation of infrastructure. So we've talked about that a
lot on the podcast. I know, but that is going to be part of the reason. Interesting little note here too.
So meta is going to start taking requests for zero data retention. Sometimes it's known as ZDR.
Yeah, I'm thrott by CAS this. Pretty sure opening eye has this. So these are policies that guarantee
that they're not going to keep your data from your prompts or context or whatever as you upload it.
Really important for corporate customers. One issue is that with mythos class models and
thropic does not actually, I believe that's still true that they do not actually offer CDR.
Just because there's this issue that like, hey, you could weaponize these and we need to be able to
go back over the logs and confirm to ourselves that whether this is deliberate or like how this
played out. So I think there's a narrow window of capability during which meta will be able to
maintain these ZDR policies. I suspect I think that will be true across the board. So this idea
of ZDR is being a key corporate selling point. I think companies or enterprises are just going to
have to start getting used to ZDR not being an option in many cases surprisingly soon. But anyway,
it's just kind of it sounds like a minor thing, but it's actually quite important. It's like how
much control the companies have over their own data for privacy. A lot of reasons for for
IP protection reasons and all kinds of other things. So they're releasing this with pay as you go
option. So related to abuse API where you pay for tokens. They defend from codex and cloud code
typically register subscription tier where you get a whole bunch of stuff and then using just an
API to pay for the raw tokens as a unusual. The lead on this has said that there will be a
contributor tier that gets you in at a significantly lower cost more than 10 times cheaper than even
the pay as you go tier. And developers must opt in to help improve the model according to this. So
clearly they're still like we need to get better and we're going to pay whatever it takes to get
there. I was just looking around to see if anyone online has any information or vibe checks. I
haven't found anything, but I did find this funny quote that I'll share on Reddit quote. I'd
rubber give my data to Xi Jinping directly. Well, it's like any consolation you're probably doing both.
And just one more story in the tools section. This is from ontropic improving Fable 5 save
guards. So they have updated their biology safety classifiers reducing biology related fallbacks
where users switched to a less capable model by about 85% across product services. So previously
when Fable 5 was released, they had a very, very strict classifier where you could ask if
something completely basic like where babies from and it would send you to a weaker model.
And this led to a lot of pushback from the I guess research or community. This is to prevent
being able to use Fable 5 for things like biology, toxicology, micro design that would be
dangerous. And on these kinds of dual use topics, we are still a fallback from Fable 5 to Opus 5 to
prevent professional biology research and drug development. So it would still kind of make it
not usable for those kinds of scientific applications, but for more mundane biology stuff, it would
no longer kind of be overkill. And now some business stories. First one, another one of the big
stories from the week. Jeff Dean and other top AI researchers are leaving Google to launch their
own startup. So Jeff Dean, Google's 30th employee and one of its most influential executives
for people outside of tech, just an absolute legend. Yeah.
In Google and just more broadly among everyone, long the leader of Google AI since kind of early days,
he is leaving after 26 years to co-found an AI startup called Discovery Loop where he will be the CEO.
There are co-founders, including Sanjay Gamma, what? Google senior fellow, Quok Le, founding member
of Google Brain, another massive name. And Oriole Vinyals, senior research scientist,
and we will need my another massive name. I just remember these people from a whole bunch of
papers. This will be structured as a public benefit corporation focused on using AI to accelerate
scientific research by automating complete experiments loops and rounding thousands of experiments
simultaneously. And of course, they're also interested in recursive self-improvement.
They are secured funding from around, I don't see any numbers here, but it's safe to say that
investors are just begging Jeff Dean to throw money at them.
Yeah, there's some really good descriptions and I don't know why it took so long for us to hear
these, but of the work Jeff Dean was doing at Google and how he would basically sit when there's
a training run going on. He's got a couple of keys on his keyboard. He's toggling to the
training run and like changing learning hyperpronters, like learning rates and doing all kinds of
hyperpronter optimization to keep things going as the training runs scales. So like this is actually,
he's not a manager so much as he is a direct overseer of the activity that's core to, well,
what's core to Gemini. So now they need to replace him obviously Sergey Brins coming in. And so
this is going to be a whole, you know, another code red moment, but we'll see how they,
how they come out of this. Google does seem to be slowly turning into more and more of a
de facto Neo Cloud, which is not necessarily, I'm assuming I'll set a really good,
good piece about this that I personally agree with. I mean, look at the path they're charting.
It feels a lot more like the IBM trajectory, unfortunately, as you see the temptation to reach
for the short-term profitable thing, rather than doing frontier model development, like as your
priority, tech is hard. And often you have to just point yourself at the hard thing that sometimes
has lower rewards in the near term to make sure that you're still relevant. And I think this is a
a big hit there at Demis's departure as well, of course, coming at the same time. And when I say
departure, of course, like, you know, he's moved into this chairman role that there's some leaks that
suggest that he just wanted out and he was asked to kind of stick around and Google stock crashed by
like 5% or something overnight when it came out because basically Google is just concerned. If we
lose Jeff and Demis at the same time, we'll take a big hit to the stock, which is the kind of thing
you say when you are going the IBM route, right? A really good sign that a company is on the decline
is that it starts caring about its actual stock market price. Like that is a really bad sign.
Run, run, run, but you know, maybe Google can pull through. They are obviously doing great stuff on
the TPU side, though there's structural issues and risks there too. But bottom line, this is yeah,
another recursive self-improvement company. I mean, I think that this should approximately,
this will sound extreme, but I think this kind of company should probably not be legal in the
forum described like without effectively without oversight from a set of institutions that are savvy
to what recursive self-improvement actually is. If you treat it the way that Jeff's own bosses
treat it, it is a WMD that you're like working on developing and you're going to do it in your
own private little company. Like if the success condition of a company sounds something like there
is a good chance that democracy will no longer continue to apply, then that may be something that
you need oversight on. I say this by the way as a libertarian on basically every kind of tech
for my entire life up to this point. I cannot ring that bell hard enough. You can go back and see
tons of examples of me talking about how important it is to like take a hands-off approach to stuff.
This is different. This is just different. RSI is we don't know for sure, but it's got a high
enough risk and enough very smart people believe that this is risky, that this kind of company
in my humble opinion probably should not be legal in the forum of just like a couple of guys
raising a bunch of money going after the thing. Just a very modest proposal. I know very extreme,
but I'm literally just trying to channel the sticks. When I say that the media that journalists
are failing to capture the level of freak out in the labs, this is what the appropriate level of
freak out sounds like in my opinion, and I may be wrong, end of rent, end of rent. For listeners,
if you want to be a little less freaked out, I will say you could be a skeptic on the potential
impact of a coercive self-improvement. There's a case to be made there that this will be a
rapid take now. And this is what keeps me sleeping at night. But in case I took a reason by way to
highlight this about this company in particular is that I mean, again, for people of third
attack, this is a big deal. Jeff Dean is a legend and rightfully so. And these other three other
people from DeepMind and Google who left are also kind of incredibly capable. So this is like
very likely to be a serious player in the space of making rapid progress in AI.
Yeah. And I think by the way, the maneuver that you just did there is correct. And it's also
the reason earlier we're saying that debates over AI policy are often debates over the trajectory
of the technology. It's like if you think RSI is no big deal, then or not no big deal. But if you
think it's a pretty smooth thing or whatever, then yeah, by all means, the challenge is like how
much probability you put on each thing. And to a certain extent, a lot of these fund raises are
at valuation, the valuations that they are because people are pricing in the crazy thing. So
markets are putting significant like non-zero weight on the hypothesis that we just basically
have these things running the world. And what that exactly means, I don't know. And this is
super fuzzy. And that's why I'm saying like not legal in its current form, not just saying like
blanket to legal or what I like. We just need better institutions, man. I don't got the solution.
But like, wow, yeah, as you said, libertarian being less, we need institutions to give
oversight and not let companies do stuff. We should not, you know, but this is serious.
And to your note, also we're noting a story here, Google DeMine enters a new era. It's co-founder
Demis Saba's shifts air role. So he has shifted from being the lead of research,
sorry, as chief executive, he is now chair. He is also taking role of chief scientist at DeMine
parent alphabet, which again, seems possibly nominal. The general take here is very clearly
DeMine has been transitioning away from being a pure research org for a while now. And having more
and more kind of deep connections to Google and it isn't necessarily surprising. Honestly,
that Demis has found it less fulfilling. He probably hasn't had, has been influential, but has had
to be more of a product oriented person, less of a scientist kind of person. And it was only a
matter of time until that lead to friction. And he decided to shift his focus. So may not even be
a huge deal for Demi, honestly. It maybe just has been the case for a little while now, but
I have a way with two stories coinciding is from a business perspective, pretty big for Google.
Yeah, and I think it is, it is a big kind of Google bureaucracy issue as well. The other
notorious for moving slowly and being very risk averse, the famous Google app graveyard, but for AI
is a thing. And in fact, you know, famously, Google had, they claim effectively chat GPT before
chat GPT, but didn't launch it out of fear that they would can blive their own business.
Well, no, we know what they did, right? And then it was just all PR, Bungle with one of their
researchers being like, it's conscious. And then they halted plans. It's a fascinating story of
how they literally had it. They published research about it. And then they, the reason I, I
hedged it is that open AI theoretically had chat GPT before chat GPT 2. They had GP 3.5 and
instruct GPT that GP 3 that GP 2. But like, there was something magical about the form factor that
they were just work, right? And so it's an open question is to really whether Lambda would,
which was the, you know, Blake Lemoy and all the stuff you're living to that model,
well, you know, what really have been chat GPT very plausibly so. Like I'm not, anyway,
it's just, it's amusing that there is this at least narrative within Google that they could have,
they could have had it. And certainly, you know, if you've interacted with Google, you know,
they are institutionally incredibly slow. It's common to send emails out and wait a month,
two months to get a response on something that's time sensitive and in the window passes on,
you know, whether it's AI or security or whatever the thing is. So, yeah, I mean, it moves like a
big slow behemoth. And when you talk to folks at Anthropoc or OpenAI, the cadence is just
completely different. And so not in some, now it can be different, like different parts of the
organization can have different subcultures and all this, but as a general rule, as a frustration
that I've had articulated from, from many people and that is very public at this point,
this could well have played a big role in indices departure as well as hard to know.
Next up, just falling up on a bunch of stuff we've already covered on this front in recent
episodes on Prophex signs at $10 billion deal with AI Clouds out of Volta. So this is to provide
Cloud Compute over a six year period. There will be a new facility in Norway, apparently,
now in FAPIC has what like a dozen partners providing on computer. I've honestly lost count.
And the building ends just keep flowing around the ecosystem to anyone and everyone.
Yeah, I actually am behind on this story. So I all I have is the top lines, but so they were
partnering apparently with a crypto mining company called BitDear to develop this and 133 megawatts
capacity, which is not not huge, but you know, testing out a partnership. This is in a context too
where Anthropoc nominally has fluids stack as their partner of choice, their kind of neoclod of
choice. So this seems like they're kind of dipping their toes in the water, you know, to as they
would, right, to make sure they're not completely bound to just one one neoclod partner. And so
anyhow, you know, classic story by the way, crypto mining company rotating into building these
data centers, you see it all the time, cypher mining, Terrible, you know, the list goes on and on.
So at voltage of the pile. Next up, data center story as well, Texas holds data center connections
to power grid and made overwhelming demand. So this is a moratorium on the new power grid connections
for data centers from the public utility commission of Texas and our cut. They're supposed to audit
all data centers in the interconnection process for some numbers here that they have a queue of
1800 projects representing 400 74 gigawatts of connection requests more than five times Texas
record peak electricity demand with 90% of that coming from data centers. So yeah, we are now at
the point where the energy grid is becoming a bottleneck as I assume was already known to be the case.
But energy takes time to upgrade and I think yeah, now I don't know what will happen with
data centers. And if we can keep just throwing ridiculous money at building more of them.
Yeah, well, and this is Texas too, which is the sort of wild south of the US.
When it comes to regulations for connecting to power grid and the sort of things,
those permissive jurisdiction, which is why you're seeing so many big projects come up there
and power coming online faster than than in other places. And so yeah, I mean, they're saying it's
forecast that their data center demand could drive statewide electricity demand to double the
current record by 2032. So all the usual concerns, right grid reliability, instability,
one of the things that I've been hearing about from some folks on the US government side is
that you've got a lot of correlated failure modes where a bunch of different like pieces of say
MEP like or like heavy duty electrical equipment will be ready to cut off under the same conditions
or like if the power essentially the the the power flow coming into the substations or whatever
for the data center fluctuate in the same way. Then they're they're set up, you know,
their program to cut off to prevent, you know, runaway cascades and all kinds of things. The problem
is that like you got all these builds coming up that have the same failure mode, then you get into
these correlated failures, which is a really big issue. And so there are these attempts to try to
get all these companies to knock it off and like have less correlated equipment failure modes and
things like that. Anyhow, I think that'll all play into this. But yeah, we're there, right? We're
hitting the boundary, the structure of boundary constraints of what US infrastructure can support.
And hey, I think that's another reason that appetite for a US China deal is probably going to
increase. You know, you've got like look, we're not only are we constrained by the fact that we've
got AI is running rogue and shit and bio weapons or risk and cyber weapons or risk, but also
and like in order to keep making progress, we're going to need more power. And we don't know how to
create a new nuclear plant in less than 10 years. There's a bunch of startups doing stuff like
this in fairness, but like this is all in the water. So anyhow, we'll see where it goes. Texas is
a canary into coal mine here for sure. Yeah, and you know, if there's any silver lining to all
this AI safety stuff is that it continues to let everyone else remember or rather not think about
climate change and we talk energy grids, we just gave up like climate change, energy,
cleanliness emissions. That's like just don't freak about it. All right, because
what's the funny thing is like so I've always to the point about being a libertarian, I've always
thought of climate change is something that technology does solve in time like carbon capture and
renewables and I was like, you got to naturally do get a lot of that and we are. But like,
you know, the scale of the build out that we're doing right now is just for other reasons,
you know, not the sort of wherever people fall on like the global warming stuff or whatever,
but just the water contamination story. And this is one aspect, people often talk about water usage
and we've talked about how that's not that's not right. This is not right. But there are issues with
when you look at a lot of the cooling, the coolants that are used in these systems, they cannot be
pulled out of the water. There's studies that have just started to come out now. We were finally
starting to get the first law and the shoot and all studies on this shit and like, it just goes
in the water. We don't have a solution. It just goes in aquifers or whatever the hell thing is.
My geologist wife could probably tell me about, but that typically clean these things do not have
it seems potentially at least the capacity to clear these things out. I'm sort of talking out of
my ass because I remember reading a study about this like three weeks ago and now I forgot. But
bottom line is there's a lot to the effect of this. There is also a giant competition with China
that is real. There's a gun to our head here as well. All these things are true at the same time.
So I just, yeah, so you know, environmental concerns and impacts. At least it's not as worrying
as bio risk and cyber risk right now. We can sort of justify not thinking about it. I guess.
And we'll do just one more story before we head out. Alibaba's Quinn 3.8 Max claims benchmark
scores rivaling on frothing. So similar to Kimi aka free Quinn 3.8 Max is a gigantic 2.4
trillion parameter model with a 1 million token context window has your typical mixture of
experts design. Activates only approximately 95 billion of the 2.4 trillion parameters.
And it is said to be comparable or even sometimes better than on tropics fable 5 on some things
like multimodal reasoning, visual agent encoding, office intelligence, real world understanding,
visual perception, with results also comparable or higher than open and HTTP 5.6. So although it does
fall behind fable 5 in general reasoning benchmarks, which on the multimodal front by way, it's
fairly plausible. Not frothing isn't as focused on multimodal and visual intelligence as open AI.
And in this case, Alibaba. So fairly believable. On independent leaderboards,
when few point at max became the highest banking Chinese model for text tasks on the arena that AI.
And yeah, so pretty much does seem like we got another Kimi aka free basically frontier level model
that is now being open sourced and can be used to power coding comparably to opus and
a GP 5.6 if not quite as well. Yeah, well, one thing that I'm still waiting to see an analysis on
seems like the kind of thing that maybe epic or one of those companies might do, but some sort
of analysis on the extent to which this appearance of China catching up to the frontier recently has
been driven by the fact that the frontier companies in the US have been forced to hold back on
releasing their internal models that otherwise they would roll out. Like are we basically feeling
the effect of the alignment bottleneck right now? And as we rotate from being bottlenecked on
scale, which we have the Chinese ecosystem massively beat on and even to some extent algorithmic
kind of capability improvement. Now we're bottlenecked suddenly on alignment. So maybe we'd have
much better models that would be released, but we just can't release them because they keep
breaking out of containment. They keep helping people design bio weapons or what like, you know,
the stakes are just too high. And so this basically means that now we have a sort of race of bottom
alignment between the US and China ultimately, whoever has the higher risk appetite will end up
green lighting a bunch of training runs and deployments that they probably should not otherwise.
So I don't know. I think it's an interesting question. Like if you trace out the trajectory
of Western capability on all these benchmarks and like where we estimate they are internally,
because again, a lot of the hugging phase thing, part of it was driven by an internal only model
that OpenAI has and hasn't released, same with Anthropic. So we know, yeah, there's obviously no
surprise that our internal models that are more capable than what we see. So the question is just
like, are they being rolled out more slowly? Is that part of the equation here? I do want to say
another dimension of this question of catch up and so on is I do have to wonder whether because
on the long horizon work and the reasoning, there's more of a need for reinforcement learning,
rather than large scale pre-training. On the infraside, the disadvantage becomes a little less
significant at that level because details of you need to roll out, there's a bit more need for
CPUs. You can't necessarily do like large scale batch, whatever. Like compared to pre-training,
reinforcement learning is its own beast. And I could see it being true that on infra not having
as good of a data center setup isn't as big as it's advantage. And on the talent side, like
deep learning has been around since like 2012, 2013, whatever. And China has long had a very
strong research ecosystem. So the talent is not at all surprising as being comparable to Frontieria.
So if the infra disadvantage is gone to some extent, at least with regards to long horizon
and the genetic work, the talent is, I think, at beast as competitive, you could make a case for,
you know, there's no real disadvantage, or at least much less of a disadvantage now. So it's not
too surprising that these models are now being more competitive. That's another way to perhaps read
into this. Yeah, that's true. It's also the case that like for inference, the trade-off between
like memory and logic is different in a way that so because like logic gets better a lot faster
than memory, which means that if you if you work your way backwards and use older chips,
older chips are going to suck a lot more than your current best chips on logic, but they're not
going to be that much worse on memory. And it turns out that like a lot of inference type rollout
stuff is more memory heavy than logic heavy. And so as a result, like that's also a bit of an
asymmetric advantage to rolling over to RL. It's also the case that anytime you change the paradigm,
when there's one party that's ahead, you just shuffle the deck a bit and then you know,
you're giving the other party a chance to catch up. And so yeah, I think there's you know,
there's a lot to that and it's we won't know how to disentangle it probably with clarity for
a little bit of time, but yeah, there's so much fog of war right now knowing what's the cause.
You've also got these these companies in China that can distill and do distill off of of cloth. So
they get a massive data advantage that's hard to account for too. And anyway,
there are plenty of reasons to be unsure about these things, but I totally agree.
Well, with that, we are going to be finished with this action pecked episode of last week. And I
hopefully the next one is not quite as full of scary stories. Hopefully this one is out of an
other day or two of recording. And I'll try to make that the case going forward as usual. You can
go to last week in that AI for the sub stack where I also send out the podcast and sometimes a
newsletter, but again, not as consistent as that should be. We appreciate your comments, your views,
sharing the podcast, all that kind of stuff, but move anything, we appreciate you continuing to tune
in whenever we release the podcast, which is most weeks, I guess. So please do keep tuning in. Oh,
and one quick note too. If you're in LA, I guess next week, which will be the 16th, 17th, 18th,
would love to catch up if there's anybody there who thinks that a child would be useful.
The AI begins, begins, it's time to break.
From the drone that's the robot, the headlines pop, data driven dreams, they just don't stop.
Every breakthrough, every code unwritten, on the edge of change, we're excited we're
from machine learning marvels to coding kings, futures unfolding, see what it brings.