On Mon, Jun 22, 2026 at 9:32 PM James Bowery <[email protected]> wrote: > On Mon, Jun 22, 2026 at 7:21 PM Matt Mahoney <[email protected]> wrote: >> I agree. We need a definition of friendly (or aligned) AI that we can all >> agree on. What is your definition? > You missed the point of my using the term "Leggian".
Fair enough. Let me make a list of "friendly AI" definitions and find an elegant mathematical model that's consistent with most of them. 1. Asimov's 3 laws of robotics (1942). https://en.wikipedia.org/wiki/Three_Laws_of_Robotics 2. Yudkowsky's Coherent Extrapolated Volition (2004). https://intelligence.org/files/CEV.pdf 3. Various modern LLM directives intended to avoid government regulation and/or lawsuits. Don't be racist. Don't generate nude images of celebrities. Don't encourage users to commit suicide. Don't tell users how to create biological weapons. 4a/b. Effective Altruism: Maximize collective utility. If life has positive utility (because death is bad), then life should proliferate to the maximum extent possible. If life has negative utility (because all living things must die), then all life should be exterminated now to eliminate future suffering. A common theme is more happiness and less death and suffering among users. But this won't work. I pointed out earlier that happiness is the rate of increase of utility. But we live in a finite universe that can only support at most 10^92 bit write operations before the heat death of the universe and everything in it. Therefore all utility functions are finite and have one or more maximum states. The best that a goal seeking agent can do is reach one of these states and stay there. That is equivalent to death. Furthermore, suffering is not the same as pain or negative reinforcement. Pain is a signal that alters your memory to make you fear the sensory perceptions associated with it. You interpret this fear as a memory of suffering to be consistent with your illusion of free will. Intuitively, a reinforcement signal results in some change in behavior, and stronger signals (more intense pain or pleasure) result in greater changes. We can give an algorithmic definition of the strength of a reinforcement signal. Let S1 be the state of the agent before the signal is applied, and S2 the state after. Then the strength of the reinforcement signal is the conditional Kolmogorov complexity K(S2|S1), or the length of the shortest program that describes S2 given S1 as input. For example, my simple 2007 reinforcement learner Autobliss ( https://mattmahoney.net/autobliss.txt ) is a programmable 2 input logic gate programmed by reinforcement learning. It has a 4 bit state, so the most pain or pleasure it can experience between states is 4 bits. It could experience more by being programmed to different configurations, switching every few nanoseconds, but this erases the memory of earlier experiences. Human brains have a conscious memory capacity of 10^9 bits and a write speed of 5 to 10 bits per second, so our capacity to experience pain and pleasure has a much lower rate, while our total experience is greater. An obvious drawback of my definition is that it does not distinguish between positive and negative reinforcement. That's because in a pure reinforcement learner, it doesn't matter. But in biological organisms, positive reinforcement signals generally contribute to reproductive success, while negative signals are detrimental. Furthermore, social animals including humans signal their mental states in a way that negative signals cause distress among other organisms. We call this phenomenon "empathy". Without empathy, human civilization would not have happened. I model this in Autobliss by having it signal "Ahhh" or "Ouch" and killing it if it receives too much negative reinforcement. We are almost ready to give a mathematical description of friendly AI. Without reference to utility, we want more positive reinforcement than negative reinforcement. This implies the proliferation of life provided we exclude destructive positive reinforcement such as by drugs or wireheading. This is consistent with definitions 1, 2, 3, and 4a (but not 4b). The first law of robotics says not to harm humans. CEV gives you what you would want if you were smarter, not everything you want right now. We want everyone to have this because empathy leads to cooperation. Note also that K(death) = K(death|S1) = 0. Death does not cause pain or suffering. Therefore we may reject 4b. The definition of life is self replication. It doesn't have to be by binary fission or sexual reproduction. It could involve specialization, like a queen ant or robots building factories that build robots. The key step is copying information, which decreases the entropy of the atoms that make up the copy, which requires increasing the entropy somewhere else, which requires free energy, of which 10^70 J or 10^92 bits are available in the observable universe. Therefore, I define the friendliness of AI to be the rate of entropy removed by replicating entities, measured in bits per second. A bit is at least kT ln 2, or about 3 x 10^-21 J at room temperature. I'm not sure how useful this is. It measures how much life there is, with the assumption that more is better. It is weighted toward active life, animals over plants. But it also measures industrial activity, weighted toward energy efficiency. Manufacturing moves atoms, which reduces entropy and takes energy just like biological processes. Making steel to build robots means separating iron atoms from oxygen atoms. The Earth intercepts 177,000 TW of sunlight, of which 90,000 TW reaches the ground. Plants make 300 TW of food from sunlight, of which 1 TW is consumed by humans. In addition we produce 20 TW of fuel energy, of which 3.5 TW is converted to electricity. These numbers are growing, so we are headed in the right direction. ------------------------------------------ Artificial General Intelligence List: AGI Permalink: https://agi.topicbox.com/groups/agi/Teedaf4a19e56877e-Mad8c7c5829ea3a13d1b411a4 Delivery options: https://agi.topicbox.com/groups/agi/subscription
