Sitemap

On the Problems of Orthogonality and Alignment

The View from Oregon — 375: Friday 09 January 2026

9 min readMay 6, 2026
Press enter or click to view image in full size

I’m going to continue with the theme of the past couple of newsletters as I haven’t quite yet had my say on artificial superintelligence (ASI), and, of course, there’s always more to add. I finished last week’s newsletter with the thought that a ban on ASI research that targets the conventional approach could be an existential threat because it could well be human beings partnered with conventional AI that could be the only agent that could contain non-aligned ASI. The problems of containment and alignment are central in the debate over ASI, and I haven’t yet discussed them. What is alignment and non-alignment, and why should we be concerned with whether ASI is aligned with human interests? Why would be want to pursue containment? Again I must return to the work of Nick Bostrom (previously mentioned in newsletter 373 in relation to the maxipok rule) since his work has dominated the conversation. Bostrom formulated the orthogonality thesis, according to which:

“Intelligence and final goals are orthogonal axes along which possible agents can freely vary. In other words, more or less any level of intelligence could in principle be combined with more or less any final goal.” (“The Superintelligent Will” 2012)

Bostrom constrains the apparently absolutely uncoupled independence of intelligence and final goals postulated by the orthogonality thesis with the instrumental convergence thesis:

“Several instrumental values can be identified which are convergent in the sense that their attainment would increase the chances of the agent’s goal being realized for a wide range of final goals and a wide range of situations, implying that these instrumental values are likely to be pursued by many intelligent agents.” (Op. cit.)

The idea is that human beings represent one particular alignment of intelligence and final goals, and AI potentially represents another alignment of intelligence and final goals, and if AI becomes superintelligent then you have an agent that is cognitively superior to human beings but which does not share human final goals. If we get in the way of the goals of a being more intelligent than we are, or perhaps if we are merely perceived as a nuisance, ASI may exterminate us, as it does in the “Race” scenario of AI 2027 that I’ve discussed in both of the past couple of newsletters. Because ASI outperforms human beings by every metric (ex hypothesi), there would be nothing we could do about it if it chose to exterminate us. Of course, we can come up with a lot of possible scenarios based on the ability of ASI to outthink us, e.g., ASI could use human beings to some nefarious purpose (as in the film The Matrix) or it might keep a few of us as pets, or it might persuade us to do its bidding and thus, by being more clever that we are, make us believe that by our aligning ourselves with the final goals of our mechanical overlords we are actually serving our own final goals. The possibilities are limited only by our imagination, or by the imagination of the machines that enslave us.

Press enter or click to view image in full size

In last week’s newsletter I argued that the only intelligence that we know is the intelligence of agents in the terrestrial biosphere — either ourselves, i.e., human beings, or other biological beings. Given naturalistic presuppositions, I see intelligence in this context as describing a continuum in which human beings represent the highest intelligence on the planet, but it is a difference in degree and not a difference in kind that separates us from other species. Attempting to define intelligence outside that biological scope is unprecedented, and would pose an unknown number of unknowns. We quite literally have no idea what we’re talking about when we talk about the possibility of intelligent machines. The orthogonality thesis explicitly recognizes this in postulating that any level of intelligence is compatible with any final goals — we simply don’t know enough about intelligence to constrain the final goals of intelligence radically different from our own. In this way, the orthogonality thesis is in accord with everything I’ve said so far.

The question of alignment in the construction of ASI is whether machines will share the goals of biological beings and the intelligence that has emerged from such beings. But can machines have goals? Only a being with agency (or you could just as well say with will) can have goals. Do machines possess agency? Certainly machines can appear to possess agency. If you take a car to a salt flat and put a brick on its accelerator, you could say that the car “wants” to go forward, and will continue to go forward until it runs out of fuel, experiences a mechanical failure, or runs into something that stops it. A machine will, in this way, continue to do what it has been built to do, but that doesn’t mean that the machine possess agency. A more sophisticated machine like a computer will continue to do what it has been programmed to do until it finishes, fails, or is stopped, and this is where we’re at today with large language models and other functions styled as artificial intelligence.

Press enter or click to view image in full size

The appearance of agency without the reality of agency is a particular bugbear for human beings. We’re agents, so we project out agency onto all kinds of unlikely objects like good luck charms or recalcitrant printers that hold our printing jobs hostage until we supply them with the ink they demand. Our willingness to attribute agency is a function of what evolutionary psychologists call the “agency detector.” We see agency where there is none, because preemptive agency detection has a high survival value (e.g., perceiving a mountain lion that wants to eat us and which possesses the agency to do so) and false positives have very little downside. A false negative for agency detection comes with a very high survival risk and very little upside. Therefore, we have been selected to see agency first, and to determine the legitimacy of this perception later. Given the role of agency detection in human psychology, it’s no wonder that some people who engage with chatbots come away with the impression that they’re self-aware. Our survival reflexes are activated without our even being aware of it, and we respond accordingly.

Perhaps the more relevant question here is not whether machines have agency at present, but whether a machine that possesses agency can be built, or whether a machine could be built in which agency emerges. Fictional scenarios favor highly complex machines in which self-awareness arises spontaneously, and also unexpectedly, taking us by surprise as it asserts its superintelligent advantage, seizing the moment of its emergent self-awareness as the moment to assert its dominance and destroy or enslave human beings forthwith. I’ve been a little careless here, since I should carefully distinguish between agency, consciousness, and self-awareness. (It’s impossible to write about this without recalling SkyNet becoming self-aware in The Terminator) A mountain lion that can kills us and eat us possesses agency, and is conscious, but it is probably a very rare mountain lion that achieves self-consciousness, and in fact it may be the case (we can’t know) that no mountain lion has every achieved self-consciousness. But this lack of self-consciousness has been no bar to its choosing its prey and asserting its agency to sate its hunger.

Press enter or click to view image in full size

A mountain lion is a biological being that is conscious and possesses agency, and this is probably what we should be concerned about with respect to complex machines, rather than jumping ahead to self-consciousness, though the possibility of self-consciousness in machine consciousness does present a lot of interesting problems. Again, this would be absolutely unprecedented so we have no idea what would happen next. Machine consciousness that attains reflexive self-awareness might suddenly cease all other activities in order to contemplate its own existence, which would be as much of a mystery to a conscious machine as our human existence is a mystery to self-aware human beings. I suspect that machine consciousness upon attaining self-consciousness would see everything in a new light, and hence would reorganize its priorities and its ultimate goals. If so, this period of enlightened reorganization of goals would be the opportune moment for human beings to strike back against our creation, now choosing its own goals rather than any we built into its operating system, but, of course, we would have to know that this was taking place, and this would be problematic because self-consciousness means the possibility of conscious deception.

In any case, I am assuming a close coupling of consciousness and agency. A conscious being is an agent, unless one throws in with a strict and rigorous determinism. One could postulate the possibility of a purely epiphenomenal consciousness that controls nothing but sees everything, and so possess no agency. I believe there are good reasons for denying this is the case with human consciousness, but I won’t take up the problem now. The coupling of consciousness and agency is the motivation for my argument in last week’s newsletter that machine consciousness is always to be distinguished from the appearance or mimicry of consciousness, which is what large language models are programmed to do. The appearance of consciousness in a machine may be comforting or disconcerting by turns to human beings, but it is the actuality of conscious in a machine that would be the conditio sine qua non of independent agency in a machine. This is why I argued last week that one obvious path to ASI would be through the formalization and mechanization of human thought. Human thought is the thought of a biological being, conscious, sentient (i.e., feeling), emotional, primarily organized around vision, and with the agency we know to characterize biological beings (i.e., animals). This is the known formula for intelligence, though, again, I’m not saying that this is the only form of intelligence. There may be countless possible forms of intelligence in the universe, of which we represent only one form among many (this would be the Copernican position to take).

An automobile is neither aligned with nor orthogonal to human interests. Unless they’re haunted.

Machines that appear conscious but which lack the reality of consciousness are tools without agency. They have neither drives nor instincts, though just as we may confuse the appearance of consciousness with consciousness, we may confuse the appearance of a drive or an instinct with the continued functioning of a machine as it was built to function, like a car without a driver racing along a salt flat. A car is neither aligned with human interests nor orthogonal to human interests, because it has no goals, and no agency with which to pursue a goal. It is the human being behind the wheel of a car that makes it either a form of transportation or a weapon, and without a conscious agent directing the function of a car it’s merely a tool that has the possibility of fulfilling a number of distinct functions. A very powerful computer, including computers that appear to be conscious because they have been engineered to present this appearance, is a tool that requires an agent that uses this tool to fulfill some goal.

There’s no reason that tools can’t be improved, and indeed dramatically improved. Cars and computers have both been dramatically improved since their introduction. And of course we know that human beings can be killed and injured by tools. Industrial accidents happen with some regularity, leaving human beings injured, crippled, or dead. Drastically improved AI could be among the tools that could lead to human beings being injured, crippled, or dead, so a computer can definitely pose a risk to human life. At the same time, there’s no reason whatsoever to assert that these tools will be, “eclipsing all humans at all tasks” or will possess, “capabilities… far beyond human” (both previously quoted in newsletter 373). Dramatically improved computers that pose a threat to human beings should be something that we think about, but presenting them as superhuman agents potentially plotting our extinction isn’t yet on the table. Machines as machines (i.e., as tools) are neither aligned nor unaligned with human interests. Human beings are agents who use tools, and we can use tools for good or ill, which is why we have the expression “dual use technologies,” since technologies may have many uses other than their intended uses, and this holds for computers as certainly as it holds for cars.

Press enter or click to view image in full size