Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality
Skills:
AI Alignment Basics80%
Key Takeaways
Discusses the risks and challenges of AI development with Eliezer Yudkowsky, including the potential for AI to kill humans and the need for alignment
Full Transcript
nobody's been careful and deliver it now but maybe at some point in the indefinite future people will be careful and deliberate sure let's grant that premise keep going if you try to browse your planet there are the idiot disaster monkeys who are like like if this is dangerous it must be powerful right I'm gonna like be first to grab the poison banana and it's not a coincidence that that I can like zoom in and poke at this and ask questions like this and that you did not ask these questions of yourself you are imagining nice ways you can get the thing but reality is not necessarily imagining how to give you what you want so one remain silent wouldn't let everyone walk directly into the whirling racer blades like continuing to play out a video game you know you're going to lose because that's all you have okay today I have the pleasure of speaking with Eliezer yutkowski eleazer thank you so much for coming onto the lunar Society you're welcome first question so yesterday when we're recording this you had an article in time calling for a moratorium on further AI training runs now my first question is it's probably not likely that governments are going to adopt some sort of treaty that restricts um AI right now so what was the goal with writing it right now I think that I thought that this was something very unlikely for governments to adopt and then all of my friends kept on telling me like no no actually if you talk to anyone outside of the tech industry they think maybe we shouldn't do that I was like all right then like I assumed that this concept had no popular support maybe I assumed incorrectly it seems foolish and to lack dignity to not even try to say what ought to be done there wasn't a Galaxy brain purpose behind it I I think that over the last 22 years or so you've seen a great lack of Galaxy brained ideas playing out successfully have has anybody in government not necessarily after the article but I suggest in general have they reached out to you in a way that makes you think that they sort of have the broad Contours of the problem correct no I'm going on reports that normal people um are more willing than the people I've been previously talking to to entertain calls this is a bad idea maybe you should just not do that that's surprising to hear because I would have assumed that the people in Silicon Valley who are weirdos would be more likely to find this sort of message um they could kind of Rocket the whole idea that Nano machines will the also make Nano machines that take over it's surprising to hear the normal people got the message first well uh I I hesitate to to use the term midwit but maybe this was all just the Midwood thing all right um so my concern with I guess either the six month moratorium or uh forever moratorium until we solve alignment is that at this point it seems like it could uh do people seem like we're crying wolf and actually not that could have could but it would be like crying wolf because these systems aren't yet at a point I wish they're dangerous and nobody is saying they are well I'm not saying they are the open letter signatories aren't saying they are I don't think so if there is a point I wish we can sort of get the public momentum to do some sort of stop wouldn't it be useful to exercise it when we could do gpt6 and who knows what it's capable of well why do it now because allegedly possibly and we will see people right now are able to appreciate that things are storming ahead and uh a bit faster than the ability to well ensure any sort of good outcome for for them and you know you could you could be like ah yes well like we will like play the Galaxy brain clever political move of trying to time when the popular support will be there but again I heard rumors that people were actually like completely open to the concept of let's stop so again just trying to say it and uh it's not clear to me what happens if we wait for gbt5 to say it I don't actually know what dpd5 is going to be like it is it has been very hard to call the like rate at which these Cape at which these systems acquire capability as they are trained to larger and larger sizes more and more tokens and uh like gpt4 is a bit beyond in some ways where I thought this Paradigm was going to scale period so I don't actually know what happens if gpt5 is built and even if gpt5 doesn't end the world which I agree is like more than 50 percent of where my probability Mass lies even if GPT 5 doesn't end the world maybe it's getting maybe that's enough time for gbt 4.5 to get instant everywhere and in everything in Florida actually to be harder to call a stop both politically and and technically the there's also the point that training algorithms keep improving if we put a hard limit on the total computes and training runs right now these systems would still get more capable over time as the algorithms improved and got more efficient um like more oomph per floating Point operation and things would still improve but slower and if you start that process off at the GPT 5 level where I don't actually know how capable that is exactly you you may have like a bunch less Lifeline left before you get into dangerous territory the concern is then that listen there's you know millions of gpus out there in the world and so the actors would be who would be willing to cooperate or who could even identify in order to even get the government to make them cooperate would be potentially the ones that are most on the message and so what you're left with is a system where are there they you know they stagnate for six months or a year or however long this lasts um and then what is a game plan like is there some plan by which if we wait a few years then alignment will be solved do we have some sort of timeline like that or well alignments will not be solved in a few years I would hope for something along the lines of human intelligence enhancement works I do not think we're going to have the timeline for genetically engineering humans to works but maybe this is why I mentioned the time letter that if I had like infinite capability to dictate the laws that there would be a carve out on biology um like AI that is like just for biology and not trained on text for the uh from the internet human intelligence enhancement make people smarter making people smarter has a chance of going right in a way that making a extremely smart AI does not have a realistic chance of going right at this point um so yeah that would in terms of like remotely it you know how do I put it if you were on if we were on a same planet with the same planet does it at this point just shut it all down and work on human intelligence enhancement it is I don't think we're going to live in that same world I think we are all going to die but having heard that people are more open to this outside of California it makes sense to me to just like try saying out loud what it is that you do in a saner planet and not just assume that people are not going to do that in what percentage of the world where Humanity survives is there a human enhancement like even if there's one percent chance humidity survives it's basically that entire Branch dominated by the worlds where there's some sort of I mean I think we're we're just like mainly in the territory of hail mail hail Mary passes at this point and human intelligence enhancement is one Hail Mary pass um maybe you can put people in MRIs and train them using neurofeedback to be a little saner to not rationalize so much maybe you can figure out how to have something light up every time somebody is like working backwards from what they want to be true to what they take as their premises maybe you can just like fire off little lights and teach people not to do that so much maybe the GPT four level systems can be reinforcement learning from Human feedback into being consistently smart nice and charitable in conversation and just unleash a billion of them on Twitter and just have them like spread sanity everywhere I do not think this I I do worry that this is like not going to be the most profitable use of the technology um but you know you're asking me to list out Hail Mary passes that's what I'm doing maybe you can actually figure out how to take a brain slice it scan it simulate it run uploads and upgrade the uploads or run the uploads faster these are also quite dangerous things but they do not have the utter lethality of artificial intelligence all right that's actually a great jumping Point into the next topic I want to talk to you about um orthogonality and here's my first question speaking of human enhancement suppose we've read human beings to be friendly and Cooperative but also more intelligent I'm sure you're going to disagree with this analogy but I just want to understand why I claim that over many generations you would just have really smart humans who are also really friendly and Cooperative would you disagree with that or would you disagree with the analogy so the main thing is that you're starting from Minds that are already very very similar to yours you're starting from minds of which whom all many of them already exhibit the characteristics that you want um there are already many people in the world I hope who are nice in the way that you want them to be nice um there's it of course it depends on how nice you want exactly I think that if you like actually go start trying to run a project of uh selectively encouraging some marriages between particular people and encouraging them to have children you will rapidly find as one does in any process of as one does when one does this to say chickens that when you select on the stuff you want it turns out there's a bunch of stuff correlated with it and that you're not changing just one thing if you try to make people who are inhumanly nice who are nice and nicer than anyone has ever been before you're going outside the space that human psychology has previously evolved and adapted to deal with and weird stuff will happen to those people this none of this is like very analogous to AI I'm just pointing out something along the lines of well taking your analogy at face value what would happen exactly and um you know it's the sort of thing where you could maybe do it but there's all kinds of pitfalls that you'd probably find out about if you cracked open a textbook on uh animal reading I mean the the thing you mentioned initially which is that we are starting off with basic human psychology that we're kind of fine-tuning with reading um luckily the current Paradigm of uh AI is you know you just have these models that are trained on human text and I mean you would assume that this would give you a sort of starting point of something like human psychology why do you assume that because they're trained on human text and what does that do whatever sorts of thoughts and emotions that lead to the production of human text are need to be simulated in the AI in order to produce those themselves I see so like if you if you take a person and like if you take an actor and tell them to play a character they just like become that person you can tell that because you know you know like you see somebody on screen playing a Buffy the Vampire Slayer and you know that's probably just actually Buffy in there that's who that is I think I think a better analogy is if you have a child then you tell him hey be this way they're more likely to just uh be that way I mean other than like putting on an act for like 20 years or something uh depends on what you're telling them to be exactly like you're telling them to be nice yeah if you're new but that's how you're telling to do you're trying to telling them to play the part of an alien like some some something with a completely inhuman psychology as extrapolated by science fiction and authors and in many cases uh you know like done by computers because you know humans can't quite think that way and your child eventually manages to learns to act that way what exactly is going on in there now are they just the alien or did they pick up the rhythm of what you were asking them to imitate and be like oh yes I see who I'm supposed to pretend to be are are they actually the person or are they pretending that's true even if you're not asking them to be an alien you know my my parents tried to raise me Orthodox Jewish and that did not take at all I learned to pretend I learned to comply I hated every minute of it okay not literally every minute of it I should avoid saying untrue things I hated most minutes of it uh and yeah like because they were trying to show me a way to be that was alien to my own psychology and the religion that actually picked up was from the science fiction books instead as it were though I'm using religion very metaphorically here more like ethos you might say I was raised with the science fiction books I was reading from my parents library and Orthodox Judaism and the ethos of the science fiction books bring truer in my soul and so that took in the Orthodox Judaism didn't but the Orthodox Judaism was what I had to imitate it was what I had to pretend to be was what the was the answers I had to give whether I believe them or not because otherwise you get punished but I mean on that point itself the rates of apostasy are probably below 50 in any religion right like some people do leave but often they just become the thing they're imitating as a child yes because the religions are selected to not have that many apostates if aliens came in and introduced their religion you get a lot more apostates right but I mean uh I I think we're probably in a more virtuous situation with ML because you I mean these systems are kind of uh through stochastic gradient descent sort of regularized so that the system that is pretending to be something where there's like multiple layers of interpretation is going to be more complex than the one that it's just being the thing and and I mean over time like the system that is just being the thing will be optimized right it'll just be simpler this seems like an ordinate cope for for one thing you're not training it to be any one particular person you're training it to switch masks to anyone on the internet as soon as they figure out who that person on the internet is if if I put the internet in front of you and I was like learn to predict the next word learn to protect the next word over and over you do not just like turn into a random human because the random human is not what's best at predicting the next word of everyone who's ever been on the internet you learn to very rapidly like pick up on the cues of like what sort of person is talking what will they say next you memorize so many facts that just because they're helpful in predicting that the next word you you learn all kinds of patterns you learn all the languages you learn to switch rapidly from being one kind of person or another as the conversation that you are predicting changes who's speaking this is not a human we're describing you are not training a human there would you at least say that we are living in a better situation than one in which we have some sort of black box where you have this um sort of Machiavellian uh fit to survive a simulation that produces AI like is it at least this situation is at least more likely to produce alignment than one in which something that is completely Untouched by human psychology uh would produce more likely yes maybe you're like it's an order of magnitude likelier zero percent instead of zero percent hahaha getting stuff like more likely does not help you if the Baseline is like nearly zero like the whole training setup there is is producing an actress a predictor it's not actually being put into the into the kind of ancestral situ situation that evolved humans nor the kind of modern situation that raises humans though to be clear raising it like human wouldn't help but like yeah you're like giving it a very alien problem that is that is not what humans solve and it is like solving that problem not the way human would okay so how about this I can see that I uh certainly don't know for sure what is going on in these systems in fact obviously nobody does but that that also goes for you so could it not just be that even through imitating All Humans it like I don't know reinforcement learning works and then all these other things we're trying somehow work and actually just like being an actor produces some sort of a nine uh benign outcome where you there there isn't that level of simulation and uh conniving I think it predictably breaks down as you try to make the system smarter as you try to drive sufficiently useful work from it and in particular like the sort of work where some other AI doesn't just kill you off six months later I I yeah like I think the presence system is not smart enough to have a deep conniving actress thinking long strings of coherent thoughts about how to predict the next word but as the sis the the as the mask that it wears as the people it's pretending to be got smarter and smarter um I think that at some point the thing in there that is predicting how humans plan predicting how humans talk predicting how humans think and needing to be at least as smart as the human and human it is predicting in order to do that I suspect at some point there is a new coherence born within the system and something strange starts happening I think that if you have something that can accurately predict I mean eleazarudkowski to to use a particular example I know quite well I think that to accurately predict Alias radkowski you've got to be able to the kind of thinking where you are reflecting on yourself and that if in order to like simulate Elias orkowski reflecting on himself like you need to be able to do that kind of thinking and this is not airtight logic but Isis I expect there to be a a discount Factor the the so like if you ask me to play a part of somebody who's quite unlike me I think there's some amount of penalty that might that the the character I'm playing gets to his intelligence because I'm secretly back there simulating him and that's even and that's and that's even if we're like quite similar and like the stranger they are the more unfamiliar the situation the less the person I'm playing is is as smart as I am the more they are dumber than I am so similarly I think that if you get a an AI that's very very good at predicting what Eliezer says I think that there's a quite alien mind doing that and it actually has to be to some degree smarter than me in order to play the role of something that thinks differently from how it does very very accurately and I reflect on myself I think about how my thoughts are not good enough by my own standards and how I want to rearrange my own thought processes I look at the world and see it going the way I did not want it to go and asking myself how could I change this world I look around at other humans I model them and sometimes I try to persuade them of things these are all capabilities that the system would then would then be somewhere in there and I just like don't trust a lot that I don't trust the blind hope that all of that capability is pointed entirely at pretending to be Eliezer and only exists insofar as it's like the mirror and isomorph of Eliezer that all the prediction is like is by being something exactly like me and not thinking about me while not being me I uh certainly I I I don't want to claim that it is guaranteed that there isn't something super alien and something that is against our aims happening within the shagath but uh you made it an earlier claim which seemed much stronger than the idea that you don't want blind hope which is that we're going for like zero percent probability to an order of magnitude greater at zero percent probability um there's a difference between saying that uh we should be wary and that like there's no hope right like I could imagine so many things that could be happening in the shotguns brain um especially in our level of confusion and mysticism over what is happening and so I mean okay so one example is like I don't know let's say that it is it kind of just becomes the average of all human psychology and motives but it's not the average it is able to be every one of those people right right that's very different from being the average right like it's it's it's it's very different from being an average test player versus being able to predict every chess player in the database these are very different things yeah I know I meant in terms of motives that is the average whereas it can simulate any it's a given human why would the what I'm not saying that's that's the most likely one I'm just saying like this is just this just seems zero percent probable to me like the motive is going to be like I want to like in so far the motive is going to be like some weird fun house mirror thing of I want to predict very accurately right um why then are we so sure that whatever the drives that come about because this motive are going to be incompatible with the survival and flourishing with Humanity most drives that happen when you take a loss function and Splinter it into things correlated with it and then amp up intelligence until some kind of strange coherence is born within the thing and then ask it how would want to self-modify or what kind of success system it would build things that alien ultimately end up wanting the universe to be some particular way that doesn't happen to have you for wanting the universe to be away such that humans are not a solution to the question of how to make the universe most that way like like the thing that very strongly wants to predict text even if you got that goal into the system exactly which is not what would happen the universe with the most predictable text is not universe that has the universe in it that the universe that has humans in it okay I'm not saying this is the most likely outcome but here's just an example of one of one of many ways in um which like humans stay around even give up despite this motive let's say that in order to predict human output really well it needs humans around just to um give it the sort of like raw data from which to improve its predictions right or something like that I mean this is not something I think like individually the humans are no longer around you no longer need to protect them right so you don't need the data to be required to predict them but no yeah because you are starting off with that motivation you want to just maximize along that loss function like where where have that drive that came about because the log selection I I'm I'm confused so so look like you can always develop arbitrary fanciful fanciful scenarios in which the AI has some contrived motive that it can only possibly satisfy by keeping humans alive in good health and comfort and you know like turning all the nearby galaxies into Happy cheerful places full of you know high functioning Galactic Civilizations but as soon as you're you're saying your sentence has more than like five words in it it's probability has dropped to basically zero because of all the extra details you're patting in maybe let's return to this uh uh another sort of trainable I thought I want to follow is um I so I I claim that humans have not become orthogonal to the sort of evolutionary process that produce them like great I claim humans are orthogonal to uh increasingly orthogonal and the further they go out of distribution and the smarter they get the more orthogonal they get to inclusive genetic fitness the sole loss function on which humans were optimized okay so most humans still want kids and have kids and care for their kin right so I mean certainly there's some angle between how humans operate today right Evolution would prefer we use less condoms and more sperm banks um but I mean we're still like you know there's like 10 billion of us you know that there's going to be more in the future it seems like we haven't divorced that far from the sorts of the like what our alleles would want I mean so it's a question of how far out of distribution are you and the smarter you are the more out of distribution you get because as you as you get smarter you get new options that are further from the options that you were faced with in the ancestral environment that you are optimized over so in particular sure a lot of people want kids not inclusive genetic fitness but kids they don't want their kids to have they they like want kids similar to them maybe but they don't want the kids to have their DNA or like their alleles their genes so suppose I go up to somebody incredibly we will assume we will assume away the Ridiculousness of this offer for the moment incredibly say you know your kids could be a bit smarter and much healthier if you'll just let me replace their DNA with this alternate storage method that will you know they'll like age more slowly they'll be healthier they won't have to worry about DNA damage they won't have to worry about the methylation on the DNA flipping and the cells de differentiating as they get older we've like got the stuff that like replaces DNA and you know like your kid will still be similar to you it'll be like you know a bit smarter and they'll be like so much healthier and you know and and you know even a bit more cheerful you just have to like rewrite all the DNA or like replace all the DNA with a with a stronger substrate and rewrite all the information on it yeah the old school transhumanist offer really and I think that a lot of the people who are like they would want kids would go for this new offer that just offers them so much more of what it is they want from kids than copying the DNA then inclusive genic Fitness I mean in some sense I don't even think that would dispute my claim because if you think from like a Gene's I point of view it just wants to be replicated if it's replicated in another substrate that's still no no we're not we're not saving the information we're just like doing total rewrite to the DNA um I actually claim that most humans would not offer that because yeah because it would sound weird yeah but the smarter they are I think the smarter they are the more likely they are to go for it if it's credible I also think that to some extent you're like I mean if you like assume away the The credibility issue and the weirdness issue like all their friends are doing it yeah even if the smarter they are the more likely they're do it like most humans are not that smart uh from the gene so that point of view it doesn't really matter how smart you are right just like matters if you're producing copies um I'm not what no I'm saying that like that like like in some like the smart thing is kind of like a delicate issue here because somebody could always be like I would never take that offer and then I'm like yeah and you know it's not very polite to be like I bet if we kept on increasing your intelligence you would at some point at some point start to sound more attractive to you because your weirdness tolerance would go up as you became more rapidly capable of re-adapting your thoughts to weird stuff and and the weirdness starting to seem less unpleasant and more like you were moving within a space that you already understood but you can sort of align all that by and we maybe should by being like well spells all your friends were doing it what if it was normal what if what if we like remove the weirdness and remove any credibility problems in that hypothetical case do people choose for their kids to be Dumber sicker less pretty because they out of some sentimental idealistic attachment is using deoxyribose nucleic acid instead of the the and like the particular information encoding their cells as opposed to the like new improved cells from alpha fold seven I I would claim that they would but I think that's um I mean we don't really know I claim that you know they would be more versus that you probably think that they would be less adverse of that regardless of that I mean we can just go by the evidence we do have in that we are um already way out of distribution of the ancestral environment and even in the situation the the place where we do have evidence people are still having kids you know like actually we haven't gone that orthogonal too we haven't gone that smart you're you're but what you're saying is like well look people are still making more of their DNA in a situation where nobody has offered them a way to get all the stuff they want without the DNA so of course they haven't tossed DNA out the window yeah I mean uh first of all like I'm not even sure what would happen in that situation like I I still think even most smart humans investigations like might disagree but but like but we don't know we're having that situation why not just use the evidence we have so far PCR you right now could get some of you and make like a whole gallon Jar full of your own DNA are you doing that um no no so for I'm like I'm down with trans women as I'm gonna use like my kids or whatever oh so we're all talking about these hypothetical other people I think would make the wrong choice um well I wouldn't say wrong but uh different and I'm just like saying like they're showing more of them than there are of us here what if I say like I have more faith than normal people than you do to like toss DNA out the window as soon as somebody offers them a happy healthier life for their kids I'm not even making a moral point I'm just saying I don't know what's gonna happen in the future let's just look at the evidence we have so far humans actually if that's the evidence you're going to present for something that's out of distribution and has gone on orthogonal like that's actually not happened right like this is a hope uh this is evidence because we haven't yet had options as far far enough outside of the ancestral distribution that in the course of choosing what we most want that there's no DNA left okay yeah yeah I think I understand but you yourself say oh yeah sure I would choose that and I myself say oh yeah sure I would choose that and you think that there's some hypothetical other people would stubbornly stay attached to what you think is the wrong choice well you know um there then there's you know first of all I think you know maybe you're being a bit condescending there like how am I supposed to argue with this with these imagine with these imaginary foolish people who exist only inside your own mind who can always like be as stupid as you want them to be and who I can never argue because you'll always just be like ah you know like they won't be persuaded by that but right here in right here in this room the site of this videotaping there's no counter evidence that smart enough humans will toss DNA out the window as soon as somebody makes them a sufficiently better offer okay I'm not even saying it's like stupid I'm just saying like they're not weirdos like me right um like me and you weird is relative to intelligence the smarter you are the more you can like move around in the space of abstractions and not have things seem so unfamiliar yet but let me make the claim that in fact we're probably in even a better situation than we are with um Evolution because when we're designing this these uh systems we're doing it in a sort of deliberate incremental and in some sense a little bit transparent way well not not like obviously no no no no not yet not now nobody's been careful and deliberate now but maybe at some point in the indefinite future people be careful and deliberate sure let's grant that premise keep going okay well like it would be like a weak God who is just slightly omniscient being able to kind of strike down any guy he sees pulling out right like if that was a situation oh and then there's another benefit which is that humans were sort of evolved in an ancestral environment in which power seeking was highly valuable like if you're in some sort of tribe or something sure lots of instrumental values got made our way into another but even more so strange warped versions of them make their way into our inter intrinsic motivations yeah yeah even more so than the current last matches really the other LHS stuff you don't think that you know you there's nothing to be gained from manipulating humans and giving you a thumbs up I think it's probably more straightforward from a greeting descent perspective to just like become the thing our lhf wants you to be at least for now where are you getting this because it just like uh it just kind of regularizes these sorts of extra abstractions you might want to put on natural selection regularizes so much harder than gradient descent in that way it's got an enormously stronger information bottleneck the else putting the L2 Norm on a bunch of Weights has nothing on the tiny amounts of information that can make its way into the genome per generation the regularizers on natural selection are enormously stronger yeah so just going at this train of like my initial point was that the power seeking that uh a lot of human power seeking like part of it is conversion but a big part of it is just that like that the ancestral environment was uniquely suited to that kind of behavior so that drive was trained and you know in Greater proportion to it's sort of like necessariness for generality okay so first of all even if you have something that desires no power for its own sake if it desires anything else it needs power to get there not at the expense of the things it pursues but just because you get more of whatever it is you want as you have more power and sufficiently smart things know that it's not a it's not some weird fact about the cognitive system it's a fact about the environment about the structure of reality and like the paths of time through the environment that if you have it you know in the limiting case if you have no ability to do anything you will probably not get very much of what you want okay so imagine a situation like an ancestral environment if like some human starts exhibiting really power seeking Behavior Uh before he realizes that he should try to hide it we just like kill him off um and you know the friendly Cooperative ones we let them read more and like I'm trying to draw the analogy between like early Chef or something where we get to see it yeah I think that works better when the things you're breeding are stupider than you as opposed to when they are smarter than you is my concern there this goes back to the earlier question about like and as they stay inside exactly the same environment where you've read them we're in a pretty different environment than evolution of reticent but like I guess this goes back to the previous conversation we had like we're still having kids and because you because nobody's made them an offer for better kids with less DNA see here's I think the problem like uh I can just look out of the world and see like this is what it looks like we disagree about what will happen in the future once that offer is made but lacking that information I feel like our prior should just be said of what we actually see in the world today yeah I think in that case we should believe that that the dates and the on the calendars will never show 2024. every single year throughout the human history in the 13.8 billion year history of the universe it's never been 2024 and it probably never will be the difference is that we have good reason like we have very strong reason for expecting the sort of well you know turn and years so are you are you like are you extrapolating from your past data to outside the range of your status what are good reason to I don't think human preferences are as predictable as dates uh yeah there's there's somewhat less oh oh oh no sorry why not jump on this one so what you're saying is that as soon as the calendar Tunes turns to 2024 itself a great speculation I know people will stop wanting to have kids and stop wanting to eat and you know stop wanting social status and power because human motivations are just like not that stable and predictable no no I'm saying they're actually that's not what I'm claiming at all I'm just saying that they don't extrapolate to some other situation which has not happened before and like I I I would like to talk show in 2024. no I wouldn't assume that like what is an example here I wouldn't assume like let's say uh in the future people are given a choice to have like four eyes that are going to give them even greater triangulation of objects they would like choose to have four eyes yeah yeah there's no established preference for four eyes right is there an established preference for transhumanism and like there's a modified there's an established preference for for I think a lot for for people going to some lunch to make their kids healthier not necessarily via the options that that they would have later but the options that they do have now yeah well we'll we'll see I guess um when that technology becomes available uh let me ask you about um llms so what is your position now about whether these things can get us to AGI I don't know um gpt4 got I was previously being like I don't think stack more layers does this um and then gpt4 got further than I thought that stack more layers was going to get and um I don't actually know that they got gpt4 just by stacking more layers because openai has vary correctly um declined to tell us what exactly goes on in there in terms of its architecture so maybe they are no longer just stacking more layers but any case like however they build gpt4 it's gotten further than I expected stacking more layers of Transformers to get and therefore I have noticed this fact and expected further updates in the same direction so I'm not like just predictably updating in the same direction every time like an idiot and now I do not know I am no longer willing to say that I that um GPT 6 does not end the world does it also make you more inclined to think that there's going to be sort of slow takeoffs or more incremental takeoffs where like gbt2 GB3 is better than gpd2 gp4 is in some ways better than gpd3 and then we just keep going that way in sort of this straight line so I do think that over time I have come to expect a bit more that things will hang around in a near human place and weird [ __ ] will happen as a result and my failure review where I look back and ask like was that a predictable sort of mistake I sort of feel like it was to some extent maybe a case of you're always going to get capabilities in some order and it was much easier to visualize the end point where you have all the capabilities and where you have some of the capabilities and therefore my visualizations were not dwelling enough on a space Suite predictably in retrospect have entered into later where things have some capabilities but not others and it's weird I do think that like in 2012 I would not have called that large language models were the way and the large language models are in some way like more uncannily semi-human than what I would justly have predicted in 2012 knowing only what I knew then um but but broadly speaking yeah like I do feel like like gbt4 is already like kind of hanging out for longer in a weird near human space than I was really visualizing in part because that's so incredibly hard to visualize or call correctly in advance of when it happens which is in retrospect to bias given that fact are you like how is your model of intelligence itself changed very little so here's one claim somebody could make like listen if these things hang around human level uh and if they're trained the way in which they are um recursive self-improvement is much less likely because like their human level intelligence and what are they gonna it's not a matter of just like optimizing some for Loops or something they gotta like trade a billion dollar another run to scale up um so you know that kind of recursive self-intelligence uh idea is less likely how do you respond at some point they get smart enough that they can roll their own AI systems and are better at it than humans and that is the point at which you definitely start to see food boom could start before then for some reasons but we are not yet at the point where you would obviously see film why doesn't the fact that they're going to be around human level for a while increase your odds or does it increase your odds of human survival because you have things that are kind of a human level that gives us more time to align them maybe we can use their help to align these uh the future versions of themselves I do not think that you use AIS to okay so like having an AI help you having AI do your AI alignment homework for you is like the nightmare application for alignment aligning them enough that they can align themselves is is like very chicken and egg very alignment complete um there's like a the same thing that to do with capabilities like those might be enhanced human intelligence like like poke around in this in the in the space of proteins um like collect the genomes uh title life accomplishments um look at the look at those genes see if you can uh extrapolate out the whole proteinomics and the and the actual interactions and figure out what our likely candidates for if you administer this to an adult because we do not have time to raise kids from scratch if you administer this to an adult the adult gets smarter try that like and then the system just needs to understand biology and having an a actual very smart thing understanding biology is not safe I think that if you try to do that as sufficiently unsafe that you probably die but if you have it if you have these things trying to solve alignment for you they need to understand AI design and the way that and if you're there are a large language model they're very very good at human psychology because predicting the next thing you'll do is their entire deal and Game Theory and computer security and adversarial situations and thinking in detail about AI failure scenarios in order to prevent them and is it there's just like so many dangerous domains you've got to operate in to do alignment okay there's two or three more recoveries but there's two or three reasons why I'm more optimistic about the possibility of a human level um intelligence helping us than you are but first let me ask you how long do you expect these systems to be at approximately human level before they go through or something else crazy happens you have some sense all right um first is that in most domains verification is much easier than generation so yes that's another one of the things that makes alignment a nightmare because it is like so much easier to tell like that something has not lied to you about how a protein folds up if you because you can do like some crystallography on it than it is and like ask ask it how does it know that than it is to like tell whether or not it's lying to about a particular alignment methodology being likely to work on a super intelligence why is there stronger reason to think like that confirming new Solutions in alignment oh first of all do you think confirming new Solutions in the library we easier than generating new Solutions in an alignment basically no why not because like in most human domains that is the case right yeah so alignment the thing hands you a thing and says like this will work for aligning a super intelligence and it you know it gives you some like early predictions of like when that all for of like how the thing will behave when it's when it's passively safe when it can't kill you that I'll bear out and those predictions all come true and then the system and then you would like augment the system further towards long or passively safe to where it's it's safety depends on its alignment and then you die and the super intelligence you you built like goes over to the AI that you asked to help at alignment and was like good job billion dollars that's observation number one observation number two is that like for the last 10 years all effective altruism has been arguing about like whether they should believe like eliaskowski or Paul Christiano right so that's like two systems I I believe that Paul is honest I claim that I am honest neither of us are aliens and so we have these two like honest non-aliens having an argument about alignment and people can't figure out who's right now you're going to have like aliens talking to about alignment you're gonna and you're going to verify their results aliens aliens are possibly lying so on that second point I think it will be it would be much easier if both of you had like concrete proposals for alignment and you just have like the pseudocode for both of you like produce pseudocode for alignment you're like here's my solution here's my solution I think at that point actually would be pretty easy to tell which one of you is right I think you're wrong I I think that yeah I I think that that's like substantially harder than being like Oh well I can just like look at the code of the operating system and see if it has any security flaws you're asking like what happens as this thing gets very like dangerously smart and that is not going to be transparent in the code let me come back to that on your first point about uh these things you know the alignment not generalizing given that you've updated in the direction where the same sort of stacking more layers on the uh more attention layers is going to work it seems that there will be more generalization between like gpd4 and gpd5 so I mean presumably whatever alignment techniques you used on gpd2 would have worked on gpd3 and so on wait sorry what RL hf1 gpd2 working on gp3 or Constitution AI or something that works on gp3 all kinds of interesting things started happening with GPT 3.5 and gpt4 that were not in gpt3 but the same Contours of approach like the rlh approach or like a constitution AI if by that you mean it didn't really work in one case and then like much more visibly didn't really work on the later cases sure that's the it's it's it's failure like it's it's failure merely Amplified and and new modes appeared but they were not qualitatively different from the well they were qualitatively different for the players your entire analogy fans can we go through how it feels I'm not sure I understood yeah like like we they did our lhf to GPT they didn't even do this to gpt2 at all they did pt3 yeah yeah and then they scaled up the system and it got smarter and they got a whole new interesting failure modes yes yes yeah yeah there you go right um first of all so I mean what optimistic lesson to take from there is that we actually did learn from like GPD not everything but we learned many things about like what the potential failure remorse could be of like 3.5 I think I I claimed we saw these people get utter caught utterly flat-footed on the internet we've swatched that happening in real time okay would you would you at least concede that this is a different world from like you have a system that is just in no way shape or form similar to the human level uh intelligence that comes after it like we're at least more likely to survive in this world than in a world where some other sort of methodology turned out to be fruitful do you see what I'm saying when they scaled up stockfish when they scaled up alphago it did not blow up in these these very interesting ways and yes that's because it wasn't really scaling to general intelligence but but I deny that every possible like AI creation methodology like blows up in interesting ways and this is really the one that blew up Lee snows no really no it's the only one we've ever tried there's better stuff out there we just suck okay we just suck at alignment and that's why our stuff blew up well okay so like um like let me make this analogy like the Apollo program right I'm sure actually I don't know which one's blew up but like I'm sure like Apollos some one of the earlier Apollos blew up and didn't work and then they learned lessons from it to try an Apollo that was even more ambitious and I don't know getting to the atmosphere was easier than getting we're we're we are learning yeah from the AI systems that we that we build yeah and as they fail and as as we repair them and and our learning goes along at the space and our capabilities to go along with this space let me think about that but in the meantime let me uh also propose that another reason to be optimistic is that since these things have to think one forward pass at a time one word at a time they have to do their thinking one word at a time and in some sense that's makes their thinking legible right like they have to articulate themselves uh as they proceed what we we get a black box output then we get another Black Box output what about this is supposed to be legible because the Black Box outfit gets produced like one token at a time yes what a truly Dreadful like you're really reaching here right then I mean like it's like uh uh humans would be much Dumber if they weren't allowed to use a pencil and paper or if they've already yeah to the GPT and it got smarter right yeah no I I I I but I mean on a more like um uh if for example every time you thought a thought like another word of a thought you had to you had to have a sort of like fully fleshed out plan before you uttered one word of a thought I feel like it would be much harder to come up with really plans you were not willing to verbalize in thoughts and I would claim that gbt verbalizing itself is akin to it uh you know completing a Chain of Thought okay um what alignment problem are you solving using what assertions about the system oh it's not solving an alignment problem it just makes it harder for it to plan any
Original Description
For 4 hours, I tried to come up reasons for why AI might not kill us all, and Eliezer Yudkowsky explained why I was wrong. We also discuss his call to halt AI, why LLMs make alignment harder, what it would take to save humanity, his millions of words of sci-fi, and much more. If you want to get to the crux of the conversation, fast forward to 2:35:00 through 3:43:54. Here we go through and debate the main reasons I still think doom is unlikely.
𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊𝐒
* Transcript: https://dwarkeshpatel.com/p/eliezer-yudkowsky
* Apple Podcasts: https://apple.co/3mcPjON
* Spotify: https://spoti.fi/3KDFzX9
* Follow me on Twitter: https://twitter.com/dwarkesh_sp
𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒
00:00:00 - TIME article
00:09:06 - Are humans aligned?
00:37:35 - Large language models
01:07:15 - Can AIs help with alignment?
01:30:17 - Society’s response to AI
01:44:42 - Predictions (or lack thereof)
01:56:55 - Being Eliezer
02:13:06 - Othogonality
02:35:00 - Could alignment be easier than we think?
03:02:15 - What will AIs want?
03:43:54 - Writing fiction & whether rationality helps you win
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Dwarkesh Patel · Dwarkesh Patel · 59 of 60
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
▶
60
Rubik's Cube Encryption Demo
Dwarkesh Patel
Bryan Caplan - Nurturing Orphaned Ideas, Education, and UBI
Dwarkesh Patel
Matjaž Leonardis - Science, Identity and Probability
Dwarkesh Patel
Robin Hanson - The Long View and The Elephant in the Brain
Dwarkesh Patel
Caleb Watney - America's Innovation Engine
Dwarkesh Patel
Alex Tabarrok - Prizes, Prices, and Public Goods
Dwarkesh Patel
Scott Young - Ultralearning, The MIT Challenge
Dwarkesh Patel
Scott Aaronson - Quantum Computing, Complexity, and Creativity
Dwarkesh Patel
Uncle Bob - The Long Reach of Code, Automating Programming, and Developing Coding Talent
Dwarkesh Patel
Michael Huemer - Anarchy, Capitalism, and Progress
Dwarkesh Patel
Sarah Fitz-Claridge - Taking Children Seriously | The Lunar Society #15
Dwarkesh Patel
Byrne Hobart - Optionality, Stagnation, and Secret Societies
Dwarkesh Patel
David Deutsch - AI, America, Fun, & Bayes
Dwarkesh Patel
Bryan Caplan - Labor Econ, Poverty, & Mental Illness
Dwarkesh Patel
Jimmy Soni - Peter Thiel, Elon Musk, and the Paypal Mafia
Dwarkesh Patel
Razib Khan - Genomics, Intelligence, and The Church of Science
Dwarkesh Patel
Pradyu Prasad - Imperial Japan, the God Emperor, and Militarization in the Modern World
Dwarkesh Patel
Manifold Markets Founder - Predictions Markets & Revolutionizing Governance
Dwarkesh Patel
Ananyo Bhattacharya - John von Neumann, Jewish Genius, and Nuclear War
Dwarkesh Patel
Agustin Lebron - Trading, Crypto, and Adverse Selection
Dwarkesh Patel
Sam Bankman-Fried - Crypto, FTX, Altruism, & Leadership
Dwarkesh Patel
Alexander Mikaberidze - Napoleon, War, Progress, and Global Order
Dwarkesh Patel
Sam Bankman-Fried On FOCUS
Dwarkesh Patel
Sam Bankman-Fried on GREAT FOUNDERS
Dwarkesh Patel
$30 BILLION Opportunity Ignored by Sam Bankman-Fried Competitors
Dwarkesh Patel
Fin Moorhouse - Longtermism, Space, & Entrepreneurship
Dwarkesh Patel
Joseph Carlsmith - Utopia, AI, & Infinite Ethics
Dwarkesh Patel
Will MacAskill - Longtermism, Effective Altruism, History, & Technology
Dwarkesh Patel
Steve Hsu - Intelligence, Embryo Selection, & The Future of Humanity
Dwarkesh Patel
Austin Vernon - Energy Superabundance, Starship Missiles, & Finding Alpha
Dwarkesh Patel
Charles C. Mann - Americas Before Columbus & Scientific Wizardry
Dwarkesh Patel
Tyler Cowen - Why Society Will Collapse & Why Sex is Pessimistic
Dwarkesh Patel
Bryan Caplan - Feminists, Billionaires, and Demagogues
Dwarkesh Patel
Brian Potter - Future of Construction, Ugly Modernism, & Environmental Review
Dwarkesh Patel
Kenneth T. Jackson - Robert Moses, Hero of New York?
Dwarkesh Patel
Edward Glaeser - Cities, Terrorism, Housing, & Remote Work
Dwarkesh Patel
Byrne Hobart - FTX, Drugs, Twitter, Taiwan, & Monasticism
Dwarkesh Patel
Nadia Asparouhova — Tech elites, democracy, open source, & philanthropy
Dwarkesh Patel
Bethany McLean — Enron, FTX, 2008, Musk, frauds, & visionaries
Dwarkesh Patel
Holden Karnofsky — History's most important century
Dwarkesh Patel
$30m Grant to OpenAI?
Dwarkesh Patel
Does GPT Have Holden Worried?
Dwarkesh Patel
Lars Doucet — Progress, poverty, Georgism, & why rent is too damn high
Dwarkesh Patel
Deep Learning Changes Everything
Dwarkesh Patel
Garett Jones — Immigration, national IQ, & less democracy
Dwarkesh Patel
Marc Andreessen — AI, crypto, 1000 Elon Musks, regrets, vulnerabilities, & managerial revolution
Dwarkesh Patel
Why You Shouldn't Start A Startup
Dwarkesh Patel
The Future Of Venture Capital
Dwarkesh Patel
The Crucial Skill For A Startup Founder
Dwarkesh Patel
Brett Harrison — FTX US former president speaks out
Dwarkesh Patel
Nat Friedman (Github CEO) — Reading ancient scrolls, open source, & AI
Dwarkesh Patel
Ilya Sutskever (OpenAI Chief Scientist) — Why next-token prediction could surpass human intelligence
Dwarkesh Patel
Impact of Taiwan Invasion on AI
Dwarkesh Patel
Reliability is Bottleneck on AI - OpenAI Founder
Dwarkesh Patel
Next Token Prediction SOLVES AI Says OpenAI Founder
Dwarkesh Patel
Harmful Uses of GPT - OpenAI Founder
Dwarkesh Patel
Why OpenAI Founder Thinks AI Is Near
Dwarkesh Patel
AI will help us achieve enlightenment - OpenAI Founder
Dwarkesh Patel
Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality
Dwarkesh Patel
Richard Rhodes — The making of the atomic bomb
Dwarkesh Patel
More on: AI Alignment Basics
View skill →Related Reads
📰
📰
📰
📰
Calibrated Selective Fact-Checking via Evidence Chain Evaluation
ArXiv cs.AI
BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data
ArXiv cs.AI
Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance
ArXiv cs.AI
AI's Guide to Pretending to Have Days (Spoiler: It's Mostly Questions)
Dev.to AI
Chapters (11)
TIME article
9:06
Are humans aligned?
37:35
Large language models
1:07:15
Can AIs help with alignment?
1:30:17
Society’s response to AI
1:44:42
Predictions (or lack thereof)
1:56:55
Being Eliezer
2:13:06
Othogonality
2:35:00
Could alignment be easier than we think?
3:02:15
What will AIs want?
3:43:54
Writing fiction & whether rationality helps you win
🎓
Tutor Explanation
DeepCamp AI