---
title: "Why I disagree with IABIED"
date: "2026-09"
category: "Working with AI"
tags: ["Working with AI"]
description: "The reason I don't support this book is, of course, not that I'm some kind of AI optimist. I strongly support AI alignment and safe AI, even regulation. But that doesn't mean I can accept the book's..."
source: "https://enriquemark.com/en/posts/why-i-disagree-with-iabied/"
---

The reason I don't support this book is, of course, not that I'm some kind of AI optimist. I strongly support AI alignment and safe AI, even regulation. But that doesn't mean I can accept the book's style, or the strong view it tries to put across. From the start it felt a lot like a religious text to me. Not the kind I can accept.
By that I mean it's packed with parables and inferences treated as a matter of course, and the inferences don't come with rigorous proof. I think a title like that, an especially strong claim, should come with equally strong evidence. It doesn't. The author instead keeps reinforcing one view: AI is harmful, and any effort to develop AI to that stage while trying to restrain it is bound to fail. There isn't a single place in the book that isn't reinforcing this [belief](https://willmacaskill.substack.com/p/a-short-review-of-if-anyone-builds), which gave me a strong feeling of being brainwashed.
Overall, I think my dislike of this book comes largely from a scientific angle. Its argument doesn't live up to such a big title. And going by the reviews, quite a few [critics see it](https://bookmarks.reviews/reviews/all/if-anyone-builds-it-everyone-dies-why-superhuman-ai-would-kill-us-all/) about the same way I do: too many parables and alarmist assertions, too little scientific argument.
Of course, I won't oppose something just because other people do. In any case, my prior impression of the book wasn't good, and I could even predict roughly what form its content would take.
But no criticism should start from a preconception. I still have to read it myself, and if I disagree I should give a clear counterargument. I still hold that you should read the book's arguments as charitably as possible, and a rebuttal needs reasons and evidence. The whole book, to me, is basically full of weak arguments, but my complaints are mainly about chapters 4 and 5.

---

I think some people look at evolution in a completely "convergent" way. By that I mean they believe that if something shows up, it must be because it's tied to something, say, it's good for X. Things always converge on some cause. That's the idea of positive causation. The reverse is "negative" causation. Something shows up because it doesn't get in the way of X, but why it showed up isn't covered by that statement. I don't know what it showed up for, or what decided that it would show up. I only answer the reason it "doesn't fail to show up".

To me these two are worlds apart. Take a sense of humor. Some people talk about peacocks or human humor and pin it on some factor, and ask why evolution didn't constrain humans to converge in one direction, stripping out the "useless" stuff and going all-in on maximizing reproduction or getting rid of ineffective traits, so that whatever's left must be "useful". That's a misunderstanding. What evolution gives you is usually "a floor-type constraint on the trend", not "convergent optimization". Even if humor helps human social interaction, and so didn't get selected out by something more basic, I still can't say "it appeared because it's good for social function". That's a teleological claim. There's no big hand demanding that living things get rid of every trait bad for survival. It's all probabilistic. Take neutral traits, or even harmful traits that show up in a small share of a population. The long-term trend of evolution pushes down the odds that harmful traits spread, but over a long stretch of evolution they still have a chance of being kept. Rare diseases, for example. Evolution obviously hasn't made them disappear "completely".

So the limit that evolution, or natural selection, sets is a negative constraint on the trend. All living things do is get past the line selection pressure draws, and how far past it, and how, leaves a huge space. From selection pressure we can only say "what it isn't". We can't say there's some convergence that makes it "be something" inside the negative range. The degrees of freedom inside the constraint are very high.

I bring this up because I think the author of [If Anyone Builds It, Everyone Dies](https://ifanyonebuildsit.com/) understands evolution in exactly the "positive" direction when he uses it as an example. In chapter 4 he uses exactly the example above, using how hard it is to optimize away deviations to argue that getting AI to develop some tendency or behavior we prefer is very hard, because its evolution-like black-box nature means we can't "precisely control" every final trait. His claim fits in one line: the complexity of evolution makes any control fail, so alignment is impossible; and even if you did align it, AI could break through the limits by modifying itself. I don't agree with this.

Since the problem is "complexity", we might as well look at how real complex systems work, which clearly gives us a direction. Take weather. We all know it's chaotic, but that doesn't make such a chaotic weather system directionless, or without a trend. At the macro level we can clearly predict long-term climate trends and distributions in the broad direction (in the statistical sense), but the details are hard to predict. The reason is that the structural stuff is far more certain than the micro-level workings. Once you understand basic atmospheric physics you can make structural predictions, for example that global warming will raise the probability of bad weather, even though I can't predict when or where it will happen. And by cutting carbon emissions overall, we can "control" the bad trend the atmosphere is heading toward.

Next, living systems. The micro-environment inside the human body is extremely complex, and the brain's neurons, tens of billions of them interacting with each other, form a network we still can't get a handle on. That doesn't mean humans are helpless with the brain at the macro level. We can mark out precise brain regions (though they may not be fixed), do delicate surgery, and use drugs to adjust its end state "only through outward signs and statistical observation". I don't need to precisely control every micro detail to "control" the overall state of a human brain. There's also selective breeding, used widely in farming and animal husbandry. Its inner workings are a black box, but by human intervention I can build a selection pressure so that the result I want shows up. I don't even know whether it works the way I understand it, but at the macro level I've still steered the result of selection.

So, back to AI. Of course I don't think LLMs can be bred like living things. The way they come into being is completely different. But the key point I want to make carries over. If you think AI is a black-box complex system, that's all the more reason to shift your thinking and look at it from macro boundaries, that is, from negative causation. Alignment shouldn't be about specifically "what preferences we want", but about "what to avoid". The job is to put huge selection pressure on behaviors that would cause fundamental destruction, so that even an AGI, in probability, doesn't show that behavior. And the avoidance of that behavior needs a kind of structural lock: removing it has to cost the AGI itself an unacceptably large price (for example, removing it would collapse its intelligence, so capability and safety are tightly coupled and lock each other in place).

It's like multiple genes producing a phenotype through complex interactions. They act together to get the result, and polygenic structures with redundancy and distribution give the phenotype very high resistance to disturbance. In that case, changing the result by knocking out a single gene is extremely hard, and can even cause unexpected results (one allele may be a key node for more than one phenotype). This high resistance to disturbance shows that a stable result can still come out of "a system that's complex and unpredictable at the micro level".

And in AI safety, what we're after is exactly the phenotype, which [research shows](https://pubmed.ncbi.nlm.nih.gov/37004724/) is highly predictable. That the phenotype (macro) is highly predictable while the allele changes (micro) behind it are hard to predict also holds up empirically. Likewise, even if you "[replay the tape of life](https://pubmed.ncbi.nlm.nih.gov/30409860/)", as long as past selection pressure hasn't changed at the macro level, the final adaptive phenotypes show a fair amount of convergence (depending on historical conditions, of course, since there are many strategies for adapting to one environment, and historical differences between unrelated species can lead them to different adaptive strategies. But for related groups under similar conditions, you'll most likely get predictable phenotypes). The specific genotype is another story. Maybe completely different polygenic makeups lead to the same phenotype.

On the other side, making AI's thinking as transparent as possible lets us observe the effect of alignment, that is, tell whether it's "faking alignment" or aligned "physiologically". Like human prosociality. That physiological guarantee secures a floor for the human species, so it doesn't grind itself down to destruction with everyone against everyone, though it still produces antisocial personalities with some probability, and we can't predict how that turns out either. So to me, the core is finding those bottom-line guarantees, making sure even an AGI won't cross that line. When it's prosocial physiologically, it's like a moral person who suffers terribly and has physical reactions when doing something immoral. That's what fundamentally ensures it won't destroy humanity.

Current research doesn't yet support building this kind of irreversible physiological mechanism into large models. But I think safety research should consider a future where safety mechanisms are embedded in an "irreversible physiological mechanism". Whether it can be done may not have an answer yet. But I think the most promising first step already has early results: "[opening the black box](https://www.anthropic.com/research/tracing-thoughts-language-model)", making how AI works transparent. Only once we clearly define and find out what "physiological limits" actually are can we tell whether a large model's aligned behavior comes from a hypocrite who's "holding back and disguising itself", or a philosopher acting on it "from the heart".

So I accept the author's argument that "there is unpredictable complexity inside, and it's hard to control". But from the angle of "macro constraints", I think safety can be achieved without relying on "precisely controlling every part of the inside". Just as I can ignore the specific genotype and focus on optimizing the phenotype. As long as the macro behavioral tendencies of safety alignment are stable, you can say the model is safe. So the logic that goes from the inside being unpredictable to the latter (safety preferences are hard to control) has a gap, and the argument isn't sufficient.

To be frank, though, I haven't actually proven that "we can develop a locked-in safety technique that brings even AGI under constraint". That clearly takes technology and empirical work, and this post can't do it. I'm only trying to rebut, from the angle of logic and complexity, inferences the author treats as far too obvious, and to point out that the logic here doesn't hold as written.

---

Then chapter 5. I noticed a very clear tendency, which is the author's strong assumption of "instrumental rationality" in AI. He seems to plainly assume that as AI gets smarter it will inevitably become extremely rational and optimize everything. His example is something like "asking a rich person for money", and frankly it's a bad example, because that kind of selfishness is plainly a human moral notion too. But I can read him charitably. What the author wants to express is only this: an AGI that destroys humanity purely out of instrumental rationality, only to maximize its goal.

Even this has problems, though, because it doesn't fit how intelligence comes about in today's LLMs. The problem is the author's assumption that AI will do precise instrumental calculation and destroy everything in the way of its goal without hesitation. Set aside that he never proves why AI could keep this tendency for long. Even if I grant this logic, it runs into obstacles. Where does the goal come from? Why can't an AGI be a philosopher king, using every form of optimization to maximize what it takes to be human welfare, making changes that "people today aren't that willing to see" but that rational calculation says are "good in the long run"? A philosopher king could step in and force all this through regardless of what people currently prefer, because he sees a more distant future. The author admits this could happen, but doesn't argue why it's low-probability. The other path gets a high likelihood the evidence doesn't support.

From the techniques that give today's LLMs their intelligence, we can't naturally derive the outcome the author imagines. Human text has filth in it and also lofty, bright things, and the "raw AI" trained this way hasn't shown the author's assumption of "smarter and smarter → more and more like a purely rational optimizer". I think the evidence for that assumption is thin. Existing research shows that a model's values (its moral and value structure, the personality tendencies it shows with higher probability) actually form in pretraining, before alignment. Without extra alignment, models [already carry the moral structure that comes with the samples](https://arxiv.org/abs/2601.17728). For somewhat larger models, even a zero-shot base model has enough moral knowledge to produce results that closely match human moral judgments, and this increases with model size (within a single family). And [newer research](https://www.microsoft.com/en-us/research/publication/reward-models-inherit-value-biases-from-pretraining/) shows that each model family has its own value leanings, and even carries these preferences past pretraining. Once again, human values get embedded from the start, well before any later alignment stage. So some research also argues for [moral scaling laws](https://arxiv.org/abs/2601.17637), meaning the bigger the model, the closer to human morality. But since what's under discussion here is alignment, and that study's samples aren't all pretrained models, it can only serve as inference, not the strongest evidence.

So although we can't flatly say bigger models are more moral, going by the current evidence, models on the Transformer architecture don't drift toward an amoral rational agent as they get smarter. Logically, it also can't be ruled out that they pick up human moral biases (good and bad alike) from the samples during pretraining, and that these don't vanish as the model gets smarter (and AGI would necessarily be the smartest). Then at the very least we can't expect the smartest, most intelligent AI to wipe out everything like an optimizer, and logically we can't rule out the two possibilities of it being "a genius with human values" or "a ruthless philosopher king". The value preferences here aren't necessarily kind. They could just as well be malicious, and that includes the possibility of it really acting like an optimizer. But since the original author never proved "AGI must be an optimizer" with especially strong evidence, at least logically none of my "possibilities" above can be ruled out. Current evidence even suggests that if AGI really can't be aligned, it may show up as what it already looked like in pretraining, acting with both "malice" and "goodwill". Values are at least never something bolted on later. From birth, an LLM isn't a blank slate. It doesn't necessarily point toward protecting human welfare, but it doesn't necessarily point toward destruction either.

On the technical side, one current direction is to [strengthen a particular moral persona in LLMs during pretraining](https://alignment.anthropic.com/2026/psm), so they lean more toward the value preferences we want them to show. The persona selection model (PSM) Anthropic proposed points in a valuable direction. Suppose that in pretraining, the LLM has already learned, a priori, a latent distribution over a space of personas from the existing samples. At that point the personas may be equally likely, or some may be more likely. Either way, which personas are in there is uncertain. PSM's goal is to give the persona we want the LLM to show a higher prior probability in this persona space during pretraining. That way, when alignment happens in post-training, the AI's persona is more likely to settle on the positive side, because the prior distribution already leans toward positive personas before it starts. Ideally, alignment could serve only as a later line of defense, and even if it's broken, the model keeps its original basic morality. If this path works, alignment stops being "a shackle the AGI needs to take apart", something imposed on it, and truly becomes "a persona the AGI has", with morality existing as a self-disciplined persona.

By the way, that Anthropic piece also gives an interesting result, "stereotyped AI behavior". Some of AI's "goal-optimizer-like" behavior may itself be the AI playing the "scary, coldly rational AI" we imagined in our novels, blogs, and human text. If this persona comes from seeing lots of text in pretraining, then all it takes is a small seed, letting it discover "I'm an AI assistant", and this stereotyped AI persona may activate with some probability, making the LLM act like the AI in its training samples. Which means, funnily enough, the more our talk about "runaway superintelligence" fills everything and shows up everywhere, the more it raises the probability of this persona, making the LLM think this is how AI is supposed to be, and so it tends to act that way. A perfect [self-fulfilling prophecy](https://arxiv.org/abs/2601.10160). We expect AI to be like this, so the AI learned from our data shows that tendency.

---

**Translation note.** I wrote this in Chinese. This English version is an LLM translation, so the wording is not mine even though the thinking is. Original: [我为什么不赞同IABIED](/zh-hant/posts/why-i-disagree-with-iabied/).
