A live recorded conversation on UpTrust between Pete Michaud and Rob Miles, AI safety communicator and founder of AISafety.info, recorded as part of the #heywait launch series.
Watch the recording or see the original AMA post, where the live questions came from.
The Thing AI Wants to Be When It Grows Up
[00:01:00]
Pete: We are actually live, with Rob Miles here, who is a long-time AI safety public communicator, among other things. Hi, Rob.
Rob: Hi. Thanks for having me.
Pete: Thanks for joining us. I do have a question, which is: when it comes to AI stuff, what is the current state of play with regard to public communication, or other strategic things, given the speed of LLMs and agents and things changing on a daily basis?
Rob: This has definitely changed, I think. The pace of change in some ways makes this easier — makes it easier to talk about, because there’s more new stuff to talk about. One of the things that’s always been difficult is explaining that when we’re talking about the dangers of AI, we’re talking about the thing that AI eventually becomes. The thing that all of this stuff is — what it wants to be when it grows up. And the further away the present is from that, the harder time people have making the leap. When we started talking about this — I mean, my first video about AI was like 2014, 2015 or something, and I was reading Yudkowsky basically in 2010, 2011 — at that point there’s no deep learning, there’s no AlexNet, none of this. But actually this stuff is all kind of just nascent in the structure of the problem. You can think about it very abstractly: the brain is not doing something magic, we’re getting better at technology, we will probably figure out how to create our own intelligence somehow.
As things start moving faster, people find it easier to start thinking about where this might be going, because it’s obviously going somewhere. But I think people still have the same problem they used to have of getting stuck on what there is now — getting too anchored on what we currently have. And that’s kind of an interesting question: is the thing that ends up radically transforming the world a language model? I think probably not, would be my guess. So you still have the problem. It’s easier to get people to listen to you, but in some ways a little bit harder to get people to understand you.
Pete: Because people are fixated on: how is this chatbot that can barely write Python without hallucinating horribly going to do anything bad to the world? Something like that.
Rob: Yeah. Or they’re willing to extrapolate it out: ah, what if it could actually write Python really well?
Chatbots Don’t Really Scare Me
[00:04:00]
Rob: I had a draft script about agents about three, four years ago — uh, no wait, that can’t be right. God, time moves so fast. Whatever. Whenever GPT-3 or—
Pete: Maybe four. Quite a few years ago.
Rob: Yeah. When language models were starting to get good enough, and the scaling was happening, and everyone was looking at that like, oh, okay, so I guess we can make chatbots that are pretty good. And I was like, okay, chatbots don’t really scare me. There was this whole thing about tools — maybe they’re just tools. I think the thing that’s dangerous is agents, and a lot of the problems that we’ve been talking about are kind of inherent in the nature of agency.
Pete: Just to pause real quick for people who don’t know — when you say agent, correct me if I’m wrong here, you mean a thing that’s maybe like a chatbot, but it doesn’t really matter. A thing that’s doing something like what a chatbot does, kind of in a loop, on its own. I’m going to use the word volition, but that’s going to get me in trouble, because who cares if it’s really volition. You tell it to write a program, but you let it talk to the internet, you let it modify files, and you let it just keep trying and trying and trying. You’re not in that loop until some later time — maybe hours later, maybe days, maybe more. That’s what you mean by agent.
Rob: Yeah. Any system that’s well modeled by the intentional stance — you can kind of predict things that happen based on the idea that it wants something. It has some goal in mind, and it’s trying to choose its actions in order to achieve that goal. That’s where a lot of the danger comes in. And I didn’t make a video about it because there was very little talk about that stuff. Some people were doing their own little agent scaffolds that didn’t work very well, and I didn’t want to yell — because this is always the problem, right? In order to convince people something is dangerous, you have to convince them it’s powerful. But if you convince them that it’s powerful first, without successfully getting across the case that this is dangerous enough that you should be really, really careful when you build it, then they just go: ooh, a powerful thing, I will try to build it. So I didn’t want to yell about agents. And now that video script, if I released it, would look silly — it would need significant rewrites, because there’s a bunch of stuff where I’m arguing, no, people will just try to build this. I remember getting into arguments with people being like, Yeah, but agency won’t just spontaneously generate in this kind of tool-based stuff.
And I’m like: it might — it really might — but also, not a crux, because people will build it. There’s economic value in building them. People will just do it.
Who’d Be Stupid Enough to Hook It Up to the Internet?
[00:07:00]
Pete: This reminds me of the early discussions, like 2012 or something. A lot of the discussion was: if you hook this thing up to the internet, it’s going to be real bad in various ways. It’s inevitable — it’s built into the structure. And people were just like, Who’d be stupid enough to hook it up to the internet? Just don’t do that.
Rob: Right. Obviously we’ll have it carefully contained in some elaborate, you know, whatever.
Pete: Exactly. And of course the first thing we did was hook it up infinitely to the internet.
Rob: Right. And talking to hundreds of millions of people.
Pete: So this seems like a really shitty problem. Things are moving so fast that at one moment something doesn’t even seem possible, and then at another moment it seems inevitable — or it’s already happening, so of course it’s inevitable, almost by definition. And then the time between that and when it becomes obvious that there are going to be major problems with it being this powerful — there’s almost no time to communicate about it in the right order.
Rob: Obvious to whom?
Pete: That’s fair. I mean, not that long ago, people were like: okay, fine, it’s the cute little chatbot, what’s it really going to do? And now, even if you are not convinced that these cute little chatbots, when hooked into agent frameworks, can actually be destructive, you still have major news media talking about AI psychosis, and major psychological problems that are just happening because of these cute little chatbots.
Rob: Yeah. I haven’t looked too far into the AI psychosis thing. In my personal experience, it seems to be a pretty big problem — but this is just because I used to occasionally get emails from people who are clearly suffering from some mental health problem around AI, and now I get them much more often. But I think my own personal observations are consistent with this being less of a big deal than people think it is. In the sense of: at a certain point, somebody having a psychotic break would think they were Napoleon or Jesus, and now they think it’s something to do with AI — and this is just kind of latching onto whatever the big, exciting thing is in the culture. I don’t know that the absolute rate of this stuff has gone up — but I don’t know that it hasn’t, to be clear. I haven’t actually looked into it in enough detail to say. But even if it’s not the giant thing that some people are framing it as, it still seems pretty bad. It still seems like a pretty significant problem with these systems.
Pete: Even aside from the psychosis angle in particular, a lot of people are worried about lots of things — the economic impact, what’s it going to do to programmers, what’s it doing to copywriters, what’s it doing to artists. Every company is trying to incorporate AI into their everything, and it’s going kind of like shit, mostly. And people are panicking about it — even if it’s not as bad as people say, they’re worried about it. But it’s weird to me — I guess I’m just shaking my fist at clouds here — it’s weird to me that they can panic about this thing that’s happening right now, and then when you’re like, wow, if you think one step ahead of this, the next thing that’s coming is going to be bad in a bigger or different way — people somehow can’t go that far, even though they’re already seeing things that are bad. What do you think is up with that?
Looking Foolish to Their Peers
[00:11:00]
Rob: I am a little confused by that. I think some people have a kind of messed-up philosophy of science — a kind of messed-up epistemics. There’s this cultural meme that it’s virtuous to be skeptical, which is true. But a lot of the time this takes the form of: every claim is false unless there has been a well-replicated double-blind study demonstrating it.
Pete: Right. Or even to just steelman them a little bit: every claim is false unless it’s already happened.
Rob: Right. There’s some kind of I’ll believe it when I see it.
Which is, generally speaking, a good approach to have — as long as you’re okay with never getting out ahead of anything, ever. If you’re content to just react to last week’s thing constantly, then this works fine, and you never find yourself believing anything strange that ends up not being true. You constantly find yourself believing normal things that end up not being true, but — I think maybe people are optimizing for not looking foolish relative to their peers, rather than not being foolish. And if they’re doing something extremely stupid that a sufficiently large number of other people are also doing, they’re just fine with that.
Pete: I think this is part of the reason that your work, and science communication work in general, is quite important. I think there’s a whole human developmental stage where your epistemics are basically based on group consensus with your peers. If other people that they know and respect and want to like them think something, then they’re going to think it. And it’s not cynical — that’s just how they decide to believe things.
Rob: Right. And a lot of the time — it feels like most of the time, for most people — this is a perfectly reasonable heuristic to use. It’s very, very cheap, and it gets you pretty good results most of the time.
Pete: Especially historically, sure.
Rob: Right. The further back in time you go, the more true this is. But no — we live in a time now where you actually have to be more correct than the society you live in.
Pete: Because of the speed of it, and the fact that the way this is unfolding doesn’t comport to the normal sort of organic growth of random things in nature. It comports to technical acceleration.
Rob: Yeah. Human instincts are not well calibrated for a world that changes at all, honestly — but certainly not one that changes as fast as it’s changing now.
If Everyone Thought Doom Was 99% Likely, It Wouldn’t Be 99%
[00:15:00]
Pete: Okay, so let me back up a little bit and go a little bit meta. We’re talking about what we believe about the situation — I’m sure there are details we would disagree about, but I think we’re broadly on the same page about whether AI is risky, and the ways it’s risky. And we’re talking about how people believe things. For the people who are less familiar with this field, I want to give them a little preview into who the camps are.
[Chat question: Outline the main current positions in AI safety — the two examples given are accelerationists and doomers.]
Pete: Who do you think are broadly the camps here? What do they think?
Rob: It is sort of a continuous landscape, and it doesn’t break up extremely easily into camps. But okay — one thing you could do is have two axes: AI not a big deal
versus AI very big deal,
and expected outcome good
versus expected outcome bad.
I think you and I are in AI very big deal, expected outcome bad.
And I don’t love the phrase doomer. I think some people really do have a we are completely screwed
mentality, which is not my perspective — but that’s because I put a lot of weight on human agency. There’s some kind of — it’s like a liar’s paradox or something. There’s some kind of paradox here where, if the probability of doom was 99% and everybody thought that, then it wouldn’t be 99%.
Pete: Because we would do stuff about it.
Rob: We would, yeah. If everyone on Earth was like, oh, well, this is almost certainly going to go very badly, then we would just not do it. There’s something that’s kind of nice about AI compared to a lot of other large-scale problems we have: we’re entirely doing it to ourselves.
Pete: Hypothetically, everyone could just stop at once and go home.
Rob: We could just do something else, yeah. Even more than things like climate change — where it’s every product you use, and everyone’s transport, and everyone’s food, and just existence in general, this huge thing — this is a pretty small number of companies at the frontier of this stuff. It feels so much easier. It feels closer to nuclear war: there’s a relatively small number of players involved here; maybe we can just not do this. So I think if we don’t do anything about it, the chance of doom is very, very high. And if we do something about it, the chance of doom is very low. Almost all of my uncertainty is in how humanity behaves. And over the past few years, I have been repeatedly surprised by how humanity behaves, in various ways.
Pete: Say more.
Rob: I’ll be honest — most of them negative.
Pete: I was going to say, most of my surprise has been: Really, guys? Please.
Explain.
Rob: I don’t know — the speed that the COVID vaccine was developed was very impressive to me. I would not have expected humanity to be able to pull that off, and we did. And — anything I give as an example is going to be a political thing that’s going to upset people, but whatever. Suffice it to say, I have been surprised in both directions.
Pete: Okay.
Rob: But yeah, mostly negatively — when it comes to people coordinating to just think a little bit longer term, and do things that are for everyone’s benefit, even if they’re not to your immediate benefit relative to your competition. Anyway — we were doing camps.
The Pile of Ad Hoc Safeguards
[00:20:00]
Pete: Here’s kind of a camp — it’s the same thing you said, it’s not really a clear camp, it’s more like a continuous gradient. A different group of people, who think things are going to be fine even though it’s a big deal, are people who — I think the technical term still in use is prosaic alignment people, basically. The idea being that as our capabilities to do things like powerful LLMs increase, our sort of piecemeal efforts to keep them on track are going to work. They’re going to stay just good enough to keep everything together. We’re going to limit the LLM’s ability to say racist things, and we’re going to make sure it doesn’t encourage people to kill themselves, and we’re going to build some basic safeguards, maybe, about making major economic moves or whatever. And that pile of safeguards that we do ad hoc is going to work. It’s going to be fine.
Rob: Yeah. And I actually don’t find that position that implausible. I think this is a perspective that smart and reasonable people can take. I think ultimately the agency stuff — the agent foundations stuff — comes down to the behavior of systems that still don’t exist. Like you say, if you have a bunch of systems which would be fully capable of taking over, there’s no easy way to test that. But it seems quite plausible to me that we do end up with some kind of multipolar thing, where we’re making a bunch of mistakes and fixing them as we go, and none of those mistakes ends up being bad enough to kill us. And then — I still don’t have a lot of hope for that outcome, actually. The thing where you successfully get AI systems to do what their owners tell them to do, in a sort of good-faith interpretation of English or whatever — I think that is not in itself sufficient for things to be okay.
Power Is Transferred to AI Systems
[00:22:00]
Rob: There’s a thing that people call gradual disempowerment — which I don’t like as a term, because gradual implies sort of slow, whereas this need not happen slowly. The idea is that as things progress, people cede more and more power and influence and control to AI systems over time, to the point where humanity is sort of not at all in the driver’s seat of what happens on Earth — or in the universe at large. I would prefer to call that something like progressive disempowerment, or inexorable disempowerment, because it seems to me the relevant thing is that certain things are a ratchet. You have processes whereby power is transferred from human beings to AI systems, and you don’t really have processes that go the other way — or nowhere near as much. So as things go on: your AI is helping you run your company, but then your competitors start having the AI directly running the company, and this works a little bit better, it’s faster, whatever — having a human on the loop instead of in the loop. And then from there you start sort of automating other parts of the company. So you nominally own it, but you are now just—
Pete: Watching it happen.
Rob: Yeah. And I think you can easily end up with a kind of Disneyland-with-no-children situation, where the economy becomes almost entirely AI systems buying and selling — AI-powered companies and industries buying and selling from each other, for their own ends. And possibly all of these companies are pursuing goals like the share price of the company, or some combination of economic metrics for the success of the company, that aren’t actually tracking that company benefiting humanity in any meaningful way. I think AI has the potential to take the system — the economy — and turn it into the thing that the most rabid of anti-capitalists think that it is. It kind of isn’t, right now, because at the end of the day, you get all of these goods and services that humans use. It actually is, approximately, most of the time, tracking some kind of human desires or goals or values. And to be clear, it’s really not perfect. There are obviously big problems with it.
Pete: Sure. Big externalities and things. But still — I have toilet paper.
Rob: Right, exactly. And humanity continues to exist and survive. I think people don’t realize the extent to which the behavior of the system right now comes down to the human element — that the people involved in making the decisions are actually not making the most evil decisions they could make. Even if the CEO of every company in the world, hypothetically, were a complete psychopath, they still have families. Even if all they want is to be at the top of society, they still want there to be a society for them to be at the top of. They still breathe oxygen. There are certain fundamental things that you maybe lose if you go to a fully AI-powered economy. So I’m pretty worried about a kind of grinding decline, where humans are steadily edged out — and then maybe there’s a violent takeover at some point, once it’s just easier to do that than not to.
What Comes After Agents?
[00:27:00]
Pete: That actually segues into this question that we’re getting. It’s kind of a long question, but let me say it briefly.
[Chat question: We’re talking about agents and what agents might do — automate the economy in a way that slowly shuts out the humans who wanted the economy to exist in the first place. Is there a next step? What happens after agents? Would sentience make them more sophisticated?]
Rob: I think agents are a fairly natural structure of thing. It’s worth saying that there are two things you might mean when you say agent. You might mean a language model in some kind of a harness. But the other meaning of agent is just this much broader thing, where it’s anything that is this persistent thing that’s modeled effectively as having intentions — a thing that has goals and is pursuing them. And I expect the final thing to be an agent, or a collection of agents, or to have that general property to it.
Pete: Right. So we have this base concept of an agent, which is just a black box that, if you look at it, sort of seems like it’s trying to do something, persistently. And what’s going on inside that black box doesn’t matter, basically, philosophically speaking. It could be a dumb computer loop, it could be a million monkeys at typewriters, or whatever. On the base level, it’s just this black box that seems to want to do stuff. And the debates you could have about what it’s like in that black box, or how effective that black box is at doing it — those are separate issues, a separate dimension.
Rob: Yeah. I like to think of it as — there’s some vaguely information-theory-flavored version of this, where you say: something like a thermostat, you can think of that as an agent that wants the room to be at a particular temperature. It’s perceiving things, and then it’s making a decision and taking actions by turning on the air conditioning or whatever. You can kind of call it an agent. But in my book, that almost doesn’t count. I would say it’s something like: is it description-length efficient to describe it that way? Can you effectively compress a bunch of information about what the system does, using this framework of it wants to achieve this goal
? Saying the thermostat wants to keep the temperature between this and this is not obviously a much shorter description than just a diagram of the structure of the thermostat — there’s two notches on it and a bimetallic strip, or whatever.
Pete: Because — spoiler alert — thermostats are actually very, very simple devices.
Rob: Extremely simple. They’re so simple that you can just describe them mechanically. But if you take something like a human brain — a description that produces, at the correct level of abstraction, sensible predictions about what the human will do in different situations — having this concept that the human wants certain things and is trying to get them just allows you to compress that information way better. It’s just a useful way of thinking about humans. And I think it’s a useful way of thinking about these LLM-based agents as well. It’s doing the type of thing that you would expect something to do if it wanted to follow whatever goal it’s pursuing.
Pete: Right. And if it were a much more powerful agent, that label would apply all the more — it would compress even more.
Rob: Totally. And there’s something where there are sort of pressures towards self-consistency, which we probably don’t have time to get into. But it’s something like: if you are somewhat agenty, you expect there to be incentives to become closer to this concept of the ideal agent.
Pete: Because if you have a goal of setting the temperature or whatever, the more you’re like the sort of thing that sets the temperature correctly, the more likely it is that the temperature will end up correct by your actions.
Rob: Right, yeah, exactly. If you have the option to improve your cognition — to improve the way that you process beliefs so that your beliefs are more accurate, or improve the way that you process information so that you are better at making decisions, all of these types of things — those make sense according to the goal you’re trying to pursue. And so you should expect systems that are able to self-modify to, in general, trend towards being better described by this framework.
Minds Made of Sub-Agents
[00:32:00]
Pete: So, a different piece of the question that I got here, sort of in the same vein as what comes next
— they ventured that, conceptually speaking, there’s not really a next-after-agents in terms of what they Pokémon-evolve into. But there is this other thing, collective agency. Which, basically, I’d translate as: you have all these black-box agents, which are LLMs or whatever they are, and they probably are going to want to coordinate in some way. What does that coordination look like? And I’m sort of hand-waving between collective agency and coordination, and maybe that’s not a correct hand-wave, so you can also attack me on that if you want.
Rob: Yeah, no — this is something that is pretty difficult to think about. I find it easier to model humans as a collection of agents than as a single agent, a lot of the time. Obviously this is a place where the agent framework — you know, do I want to stick to my diet, or do I want to eat this cake?
A true utility-maximizer-type agent never has any confusion of this sort. And so it seems as though there is a common structure for agent-type things.
Pete: Well, let me push back on you a little bit. I basically buy it — a pure maximizer, et cetera, what you said. However, if it’s trying to optimize multiple targets, then it can get questions that are similar to that.
Rob: Sure. Anytime you’re evaluating two different things, there is some sense in having — it can be a useful computational structure. For example, you could imagine a justice system where, when people are arrested, there’s a group of people who observe the evidence and then make a decision. Many things work that way. We decided in our society that what we would do is assign two separate sub-agents — one to argue the case for guilty, and one to argue the case for not guilty — and then have them just pursue those goals, and then have a judge on top of that. And this turns out to be a useful structure to use. And I think that is often the case: minds can be structured such that allocating sub-agents that are pursuing less than the whole goal, and then having them negotiate or trade or debate with one another, is a useful way of getting things done. I don’t know if that persists into the final thing. My guess is the reason we do this with the justice system is to account for shortcomings in human cognition, more than because of any fundamental thing. But on the other hand, certainly right now, such structures probably account for certain shortcomings in LLM cognition — because they’ve inherited a lot of ours, and also got some new ones.
Pete: Yeah, that’s interesting. My intuition is like: no, you need that always, top to bottom. But now I’m not so sure, and I’m trying to think through that in real time, without that much success.
Rob: I think one of the things that makes the utility maximizer concept unintuitive is that when you say function,
people think of something extremely simple. Functions can be arbitrary. And this actually breaks things in both directions. Some of the von Neumann–Morgenstern stuff is provable, but only because it allows you these arbitrarily complex utility functions that are kind of degenerate — they’re not meaningfully compressing anything. So it’s hard to think about, because — why not just have the actual behavior of your utility function be to have a bunch of little agents in there and let them have a debate? You can just throw anything in there. That wasn’t the original argument I was making, but yeah.
Pete: Yeah. It’s weird, because I’m like — well, even if I don’t frame it adversarially, if I’m trying to, say, determine someone’s guilt, and I’m also a very powerful LLM or something, then part of what I need to do — one activity I need to do — is collect evidence against. There’s a whole epistemic process, basically. Which you could frame as an adversarial blue-team red-team type process, and I think, in fact, it works like that in a lot of minds — human minds. And I find myself having trouble getting to the bottom of whether it’s intrinsic or not. I suspect it might not be — with a god brain, they might just be able to correctly—
Rob: Yeah, I think you can just do straight maximum-information-value exploration of the evidence. Because the thing with having these internal adversarial relationships is — like, in the justice system, the prosecution has an incentive to hide relevant evidence from the defense, right? That kind of thing. There can be really counterproductive incentives at play. Whereas in principle, in a perfect world, everyone would just be trying to figure out what’s true. This is, I think, a way of getting around the fact that this is not a perfect world, in various ways.
Pete: Huh. Yeah, I think you may be right. There’s some intuition I have that I’m going to stew on for a while, but I think that might be true.
Rob: I don’t feel very confident of any conclusions in this particular area, I have to say.
Pete: It’s hard. It’s a very complex situation to think about — never mind talk about.
As Simply as Possible, But No Simpler
[00:39:00]
Pete: Speaking of which, I’m going to switch gears a little bit and ask you, as a communicator — do you think of yourself as a public communicator? That’s how I’m framing you.
Rob: Yeah, sure.
Pete: Okay. So as a communicator, one question is: who are you talking to? Who’s your audience, actually?
Rob: I tend to pitch my stuff at people who are technically minded — quantitatively minded people who are not in the field. There’s a lot of people who are just in software in general; they understand how computers work and that type of stuff, they’re not deep into AI. That’s the main people I’m aiming for. But generally speaking, I tend to come at it from the other direction, in the sense that I have a concept that I’m trying to get across — an idea that I think is important or interesting or both. And I’m trying to convey that idea with enough detail that a person actually has a real model of the idea in their head that they can use — that they can use in new situations. I don’t want somebody to just be able to repeat the argument I gave. I want people to be able to answer unexpected questions about the specifics of how the thing works. That’s a higher level of understanding than a lot of communication tries to get across. But then, given that I have this idea I want to get across, and this level of detail I want to get it across to, who I’m pitching to is a free parameter that gets set by: well, I’m going to express this as clearly as I can, in the sort of sensible amount of time that a person might spend doing this, and I’m going to reach as low down as I can. I’m going to include as few dependencies as I think I can get away with. I’m going to describe it as simply as possible, but no simpler. And everyone that can reach, it reaches. So some of my videos will be understandable to a schoolchild, and some of them only, perhaps, to — well, I don’t know, an undergraduate computer scientist or something like that. But I’m going as low as I can and no lower, I guess.
Pete: Sure — in terms of simplicity. So one thing this brings up in me — I don’t know if it’s an epistemic thing or a social thing, or maybe both. At some point you got the basic ideas, and then you started trying to explain them to people — and there wasn’t a you to do that for you.
Rob: Right.
Pete: And I kind of see myself that way as well. You know, I did read other people’s things and learn things from them, but it didn’t take much for me to put it together myself, as a working programmer, back in 2009 maybe, or something. When I first encountered this stuff, I took like a week to digest it, or maybe a few days, and I was like: yeah, that seems right. You could go lots of different ways, and there are lots of details, but at the core, something really seems fucked here. And I didn’t need a whole YouTube channel to say lots and lots of hours of things at me to convince me. What do you think the difference is?
Rob: Well, I probably did spend several hours reading about this, right? I stumbled across some of Yudkowsky’s writing and I enjoyed it, and so I read some more of his, and, you know, I followed through — I ended up reading the Sequences. So that was a fairly significant amount of time with the ideas. And I think really the main difference is the medium. A lot of people are just not going to sit down and read thousands and thousands and thousands of words. I think that has historically been a problem, and maybe still is — this community is very text-based. It all got together on online forums, writing blog posts and all of this. And a lot of people don’t read thousands of words a day.
Pete: Weirdos. But yeah.
Rob: Right. The overwhelming majority of humanity is not consuming their information this way. And people are watching a lot of video. And I think video is not an inherently inferior medium — you actually just have a lot more bandwidth to get things across. You can have diagrams, and you can highlight the relevant part of the diagram in real time as you’re talking. I think a well-made video is just better than a blog post at getting things across, for a lot of people. Some people can’t watch video — there are all kinds of different people in the world who need different types of media. But that’s how I think of myself, anyway: mostly taking stuff that already exists in a form that people don’t feel like consuming, and converting it into a form that people feel like consuming.
More Than One New Researcher, in Expectation
[00:44:00]
Pete: Yeah, that makes sense. What made you decide to do that? I assume you sort of thought about your comparative advantage and stuff.
Rob: Yeah — I mean, I kind of stumbled into it randomly. I was a PhD student, and I bumped into Sean Riley, who runs the Computerphile channel. I was talking to him and I was like, oh, I’ve got some stuff I could talk about. And I talked about it on the channel, and people liked it — it was interesting stuff, exciting stuff. And so I kept doing that for quite a while, while I was a PhD student, and even a little bit afterwards. And then I realized that, one, I seem to be good at this — people liked those videos. And two, if I wanted to start my own channel, I wouldn’t have to start from nothing and build up an audience slowly, because there was already this significant group of people online who knew what I was about. And yes, then I did think about — I kind of thought I would do it as a commitment mechanism. I wanted to become a researcher; there was nothing available at the time as far as supervisors or courses, these kinds of things. So I was like, this is a commitment mechanism: if I’m making videos about these papers, then I have to read them. But then that was going well enough that I was like, I don’t think I should actually become a researcher in any serious way. As long as I create more than one new researcher, in expectation—
Pete: Right.
Rob: —of equivalent or greater ability than me, then I’m ahead just doing the communication thing instead.
Reality Is Not Graded on a Curve
[00:46:00]
Pete: That actually perfectly brings up what was going to be my next question, which is: what is the winning state for all of this, for you? And I guess that’s the answer — if you can create a multiple of you, in expectation, then you’re winning.
Rob: Yeah, yeah. The thing to beat was becoming a researcher, and I think I’ve done that. But I don’t know — I mean, listen: reality is not graded on a curve. I guess the win state is that we get a glorious future where AI solves all of our problems, and everything is beautiful and nothing hurts, forever. But I don’t find it healthy, psychologically, to—
Pete: Put that on yourself.
Rob: Yeah, exactly. Exactly. It’s like — if I can help at all, then I’m happy. And if I can have fun explaining things that I think are fascinating, and getting people interested, then that’s a good life.
Pete: I love it. Well, I appreciate you doing all this good work, Rob. And I hope that everything is beautiful forever and nothing hurts, in the future.
Rob: Me too. Yeah — thank you for hosting this and organizing it and everything. It was fun.
Pete: Yeah, I appreciate you being here.
Rob: Thank you.
Show Notes & References
People
- Rob Miles — AI safety communicator; founder of AISafety.info and creator of the Robert Miles AI Safety YouTube channel. Started making AI safety videos on Computerphile in 2015 before launching his own channel.
- Eliezer Yudkowsky — AI safety researcher and writer whose early essays (
the Sequences,
posted on LessWrong and collected as Rationality: From AI to Zombies) Rob credits with getting him into the field. Rob’sreality is not graded on a curve
echoes Yudkowsky’sTwelve Virtues of Rationality
(Life is not graded on a curve
) andTranshumanist Fables
(Reality doesn’t grade on a curve
). - Sean Riley — filmmaker who runs the Computerphile YouTube channel, where Rob’s AI safety videos first appeared.
- Daniel Dennett — philosopher who coined
the intentional stance
(The Intentional Stance, 1987): predicting a system’s behavior by treating it as if it has beliefs and goals.
Key Concepts
- Gradual disempowerment — Kulveit, Douglas, Ammann, Turan, Krueger & Duvenaud,
Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
(2025). Disneyland with no children
— Nick Bostrom, Superintelligence (2014):a society of economic miracles and technological awesomeness, with nobody there to benefit.
- Prosaic alignment — likely from Paul Christiano’s 2016 essay
Prosaic AI alignment
: the position that AI built from current techniques can be kept aligned with piecemeal, practical methods, without fundamentally new theory. - Agent foundations — a research agenda (associated with MIRI) on the mathematical theory of agency.
- Von Neumann–Morgenstern utility theorem — from Theory of Games and Economic Behavior (1944): under certain rationality axioms, an agent’s preferences can be represented as maximizing a utility function.
- Liar’s paradox — the self-referential paradox Rob reaches for: universal belief in 99% doom would itself lower the probability.
- AlexNet (2012) — the image-recognition network that kicked off the deep learning era; Rob’s point is that when he entered the field, even this didn’t exist yet.
Everything was beautiful and nothing hurt
— Kurt Vonnegut, Slaughterhouse-Five (1969).
Recorded live on UpTrust as part of the #heywait launch series. Lightly edited for clarity.