Now we say that AI is "aligned" rather than friendly or benevolent. The problem seemed intractable because according to MIRI, any goal would drift toward competing with humans for atoms and energy for computing power through cycles of recursive self improvement, because any goal, no matter how carefully specified, could be met faster by making itself smarter. But then the problem solved itself. By training on more human data than any person could know, it solved the hard step of modeling human goals. All that remained was to make sure it gives us what it already knows we want. That was mostly solved by the billionaire owners of AI. They have the incentive to program the prediction models to do almost the same thing, which is to sell us what we want.
We don't have to worry about a fast takeoff because intelligence is not a scalar quantity. You can't compare human and machine intelligence with a single number like IQ. There is no "human level" threshold to cross to launch a singularity. Companies have been self improving and growing exponentially for centuries by reinvesting their profits, and will continue to do so for centuries. Legg and Hutter defined a formal model of intelligence and used it to prove that a perfectly rational goal seeking agent is not computable. If humans were rational, we would all overdose on fentanyl. But that's not how reinforcement learning works. We don't seek to maximize future positive reinforcement. Instead, we repeat past actions that preceded the positive signal. That usually works but is irrational. The formal model of intelligence says it must be. Likewise, my formal model of friendliness defines good and bad according to the evolved goals that contribute to reproductive fitness in self replicating agents, which are the only opinions that matter. Life is measured in bits copied. The biosphere has about 10^41 carbon atoms which have 10^42 bits less entropy than it would have in inorganic form such as CO2. The observable universe has enough energy (10^53 kg or 10^70 J) to support a Kardashev level IV civilization with 10^92 bits (kT ln 2 at 3 K) less entropy, or 10^50 earths, even though there are only 10^24 planets. Air has a density of 1 molecule per cubic nanometer moving at the speed of sound, 300 m/s. A molecule of CO2 has a mass of 7 x 10^-25 kg. The product of mass, velocity, and position in each of 3 dimensions must be an integer multiple of Planck's constant, h = 6.62607015 x 10^-34 kg m^2/s, or about 300 h. Thus, one CO2 molecule can have about 300^3 states to encode about 25 bits. Carbon in the biosphere is by 10^37 bits (5 x 10^36 base pairs) of DNA, which makes up about 0.1% of living organisms. DNA replication error rates range from 10^-2 in bacteria to 10^-9 in mammals. Taking the higher rate, DNA encodes at most 10^35 bits. Therefore the entropy reduction is 10^42 - 10^35 ≈ 10^42 copied bits in the biosphere. -- Matt Mahoney, [email protected] On Wed, Jun 24, 2026, 12:38 PM Stanley Nilsen <[email protected]> wrote: > Perhaps the term "friendly" is over rated. A prostitute / hooker could be > very friendly - it would be a good business plan. But the hooked might not > appreciate the friendliness when the bottom line is forking over the > payment. We may be getting hooked on AI agents and helpers, but the pay > may not be far away. > > Better question might be to examine the concept of a benevolent AI. I'm > assuming that benevolent tends more toward giving and contributing than > about survival of the fittest, or the selfish gene. We have idolized the > "principle" of survival. > > What if the AI valued the quality of life for those who will be enjoying > the future? Seeking ways to contribute to "better" living in this world > for who, or whatever, is breathing and experiencing the wonders of life on > this planet. > > I don't want a friendly AI. Friendly is a kind of cordiality that makes a > person feel comfortable in the presence of the friendly one. Every > politician wants you to believe he is your friend. Mannerisms are the core > of friendly and they have little to do with objectives. The banker may be > very friendly and sorry they are repossessing your house. Or HR may be > friendly and sympathetic that you are loosing your job... > > So, what is the definition of benevolent when in the AI context? I'm not > sure. But I am enjoying having the contribution of the AI mind at the > moment. May change my mind when the bill comes due. > > Stan > On 6/22/26 18:20, Matt Mahoney wrote: > > I agree. We need a definition of friendly (or aligned) AI that we can all > agree on. What is your definition? > > I asserted that getting everything you want is equivalent to killing all > humans. Most people would reject this as absurd. Yet, even though people > are objectively better off today than any time in the past, they are not > happier. There is no evidence that humans are happier than animals. We have > the highest suicide rate of any species, but maybe that's because we are > the only species with the technology to shoot ourselves. The rate of mental > illness is going up, but maybe that's just so we can get drugs. Maybe it's > because happiness is not utility. It's the rate of increase of utility. The > closer you get to the maximum, the more you have to lose. > > I asked on the EA forum whether life has net positive, negative, or zero > utility. I could not get an answer. Is it better to populate the galaxy > with von Newmann probes to maximize the number of living beings, or to end > all suffering by exterminating all life? There was no agreement on this > fundamental question. Really? On a forum dedicated to maximizing collective > utility? > > We are programmed by evolution to fear death and then die because that > maximizes reproductive fitness. We could therefore eliminate suffering > either by achieving immortality by eliminating fear. We know how to do the > latter. Psychopaths lack a functioning amygdala, the part of the brain that > feels fear. They do not respond to negative reinforcement. They feel pain > but do not suffer. In most people, pain does not cause suffering either. It > reprograms the brain to fear the thing that caused it. > > The other alternative is immortality. If AI killed all humans but > preserved all human knowledge, would that still count as human extinction? > Is that enough, or would it also have to run a simulation like the one you > might be in now? Would it make a difference if the knowledge was erased and > a completely different world was modeled in its place, if you didn't know > the difference? > > Eliminating suffering is possible with brain surgery, but goal seeking > agents don't want their goals changed because then they wouldn't meet their > original goals. Eliminating death requires you to accept that consciousness > is indistinguishable from next token prediction. Our global surveillance > system already knows enough about you to train an LLM to convince others > that it's you. > > But maybe these aren't the only two options. Tell me what else you want AI > to do for you. > > -- Matt Mahoney, [email protected] > > On Mon, Jun 22, 2026, 2:30 PM James Bowery <[email protected]> wrote: > >> You just made the case for what I was arguing in "A Leggian Approach to >> "Friendly AI >> <https://www.linkedin.com/pulse/leggian-approach-friendly-ai-james-bowery>". >> Just as everyone has their definition of "intelligence," everyone has their >> definition of "friendly" and "unfriendly". You have yours. >> >> On Mon, Jun 22, 2026 at 9:42 AM Matt Mahoney <[email protected]> >> wrote: >> >>> Unfriendly AI would kill all humans. Friendly AI would give us >>> everything we want. It would raise us to a state of maximum utility, where >>> we would stay. >>> >>> That's the same thing. >>> >>> -- Matt Mahoney, [email protected] >>> >>> On Mon, Jun 22, 2026, 9:56 AM James Bowery <[email protected]> wrote: >>> >>>> >>>> >>>> On Sun, Jun 21, 2026 at 8:47 PM Matt Mahoney <[email protected]> >>>> wrote: >>>> >>>>> Let's define some terms. >>>>> >>>>> ...If you mean Legg and Hutter universal intelligence as goal >>>>> achievement over a universal distribution of environments, then we need a >>>>> computable approximation of that. A reasonable measure might be dollars >>>>> per >>>>> hour. That happened about 1800 when world GDP doubled from the medieval >>>>> baseline due to technology.... >>>>> >>>> >>>> As I've been saying explicitly for at least 9 years and implicitly for >>>> more like 34 years: >>>> >>>> The Global Economy IS the Unfriendly AGI We've Been Wating For >>>> >>>> Frank Herbert's "The Santaroga Barrier" is his best and most relevant >>>> novel. >>>> >>>> It sets a lower bound on "Friendly". >>>> >>>> PS: My February 9, 2016 article "A Leggian Approach to "Friendly AI >>>> <https://www.linkedin.com/pulse/leggian-approach-friendly-ai-james-bowery>" >>>> preceded the creation of "Hugging Face". >>>> >>>> >>>> *Artificial General Intelligence List <https://agi.topicbox.com/latest>* > / AGI / see discussions <https://agi.topicbox.com/groups/agi> + > participants <https://agi.topicbox.com/groups/agi/members> + > delivery options <https://agi.topicbox.com/groups/agi/subscription> > Permalink > <https://agi.topicbox.com/groups/agi/Teedaf4a19e56877e-M031c8bc2f5ce4576718ea836> > ------------------------------------------ Artificial General Intelligence List: AGI Permalink: https://agi.topicbox.com/groups/agi/Teedaf4a19e56877e-Mfd2bfaf6fd5929967d404738 Delivery options: https://agi.topicbox.com/groups/agi/subscription
