So to answer my earlier questions, the data you would compress to settle
questions about macro social theories would be Hume's guillotine, or
LaboratoryOfTheCounties.csv, which is a table of US census data where the
rows are unlabeled counties and the ~4K columns are demographic, economic,
crime, and geographic data, stuff like the number of farms under 500 acres
in 2007. This set could be extended to other countries, for example, by
compressing the CIA world Factbook, which is already part of the large
Canterbury compression benchmark. (The other files are the Bible and the
E.coli genome).

I think sortocracy, where sovereigns rule by decree and if you don't like
it you could leave or form your own country and make your own rules, would
be the perfect solution to tyranny. Except for the small detail that we
already have a system like this with 200 countries and they all have
borders to keep people out, so you can't leave. Sortocracy has a rule that
when people move from one country to another, the first country must give
some territory to the second to encourage immigration. What I don't see is
a system to enforce payment when the two countries disagree on the amount.
The normal mechanism for resolving territorial disputes is might makes
right.

Do I have this right? I don't see how we go from using data compression
contests to decide macro social theories to a global system of enforcing
sortocracy. Also, where would the "state of nature" be located to put
unwanted people?

Or are we talking about buying a ranch somewhere, organizing a militia, and
declaring yourself an independent state?

-- Matt Mahoney, [email protected]







On Sat, Aug 1, 2026, 10:44 AM James Bowery <[email protected]> wrote:

> Since this critical conversation has been derailed by:
>
> 1) Matt bringing in the same argument about "noise" that Ben Rudiak-Gould
> originally leveled against lossless compression of Wikipedia (indeed his
> was the *first public comment* when the Hutter Prize was announced in
> 2006)
> <https://groups.google.com/g/comp.compression/c/Pwlq6pkyc8s/m/ZdHC8HvgCYAJ> --
> something both Matt and I addressed in 2006 with Ben and, now, I with Matt
> twice now this forum before and
> 2) as a result, my willingness to offer an alternative to principled model
> selection regarding 10s f trillions of $ conflicts of interest (if not the
> fate of Earth), consistent with individual consent (ie: making it practical
> to "vote with your feet")
> 3) that "vote with your feet" conversation isn't really proper for the AGI
> list (except perhaps as an alternative to Anthropic's "Constitution")...
> 4) Even if it were an appropriate topic for the AGI list, it is obviously
> futile due to the tl;dr problem that sets in on such topics even if the
> length of the "Constitution" has an argument surface of just a few
> paragraphs (note Matt's elevation of #5 to the sole dispute processing mode
> despite it obviously being only a last resort -- this despite Matt being
> among the brightest minds on the planet in my opinion).
>
> I'll now attempt to get it back on track by AGAIN addressing Ben
> Rudiak-Gould's 2006 argument which Matt has resurrected 20 years later for
> some reason.  Charles Sinclair Smith also used this argument with me. In
> Charlie's case, the argument concerned "data cleaning" as the predominant
> cost of macrosocial dynamics modeling.
>
> Since Charlie co founded the DoE's Energy Information Administration under
> President Carter and focused on information quality, and since he financed
> the second neural network summer of the 1980s as a result of that
> experience, it's safe to say his argument represents the Steel Man case
> against my proposition.
>
> So:
>
>    1. What exactly IS "data cleaning"?
>    2. How does it address the "noise" problem in data?
>    3. Why do I assert that Charlie's claim—that over 90% of the expense
>    of macrosocial data curation is "data cleaning" (and that at some point
>    quantity has its own quality)—is not a bug but a feature of lossless
>    compression as dynamical macrosocial model selection to reformat the social
>    pseudoscience?
>
> #1) Anyone who has worked extensively with datasets knows that such
> trivial things as typos during data entry can send subsequent automated
> analysis into the weeds.  Indeed, I came to know Charlie because, as a
> consultant to his company, I bet him that I could transform some paper
> insurance tables into electronic datasets before his mid 1990s in-house
> bleeding-edge-OCR experts could -- and guarantee that there would be no
> errors.  He took my bet and I hired a bunch of "Kelly Girls" to do the data
> entry. I just hired enough redundant human OCRs so that a simple
> cross-check for equality reduced the error rate to *effectively* zero.
> No subsequent work with those datasets exposed any errors either in the
> transcription or in the original paper tables that contained them.  That
> was a LOT of work but I won the bet.  However, that is only one of many
> kinds of data integrity issues that arise in the real world.  Charlie told
> me about one case where he had to personally track down a guy in a nursing
> home who was responsible for a particular measurement that didn't comport
> with the rest of the data in one curation.  That guy was the sole
> repository of the transform used to generate that data.  This is the kind
> of HUMINT forensics that goes into "data cleaning".  One can see how this
> kind of expense could explode to the 90% level seen in 1970s analysis of
> the dynamics of the US energy economy.
>
> #2 Data cleaning addresses the "noise" problem but only to the extent that
> it corrects for obvious errors in the measurement instrument -- such
> "instruments" being entire human organizations with their various standards
> in some cases.  It cannot address "noise" that those standards are not even
> intended to filter.  That noise is sometimes known as the bewildering world
> into which we are thrown.  We are confounded by interactions manifesting as
> "random noise" that we, like a dog without a bone, nevertheless try to make
> sense of.
>
> #3 When curating a dataset to bring the social pseudosciences to heel, we
> are dealing with data forensics in the sense that we must be prepared to
> treat the various interests proffering their data as world-class fraud
> artists.  The stakes in biasing the "narrative" of the largest issues
> facing us are epic if not eschatological -- and we needn't even consider
> the possibility that these fraud artists are *consciously* colluding to
> perpetrate their deceptions!  A great case in point is Matt's mysteriously
> abysmal reading comprehension regarding Sortocracy.  Here we have one of
> the brightest minds on the planet suddenly becoming incapable of
> comprehending a few sentences!  I have no doubt that he is capable of
> comprehending far more than a few sentences and that he had conscious
> intention of doing so -- yet when it came to an issue of such intense
> conflicts of interest, he acted as though he were trying to defraud the
> public regarding Sortocracy!  We're all human and we all do this sort of
> thing.  That's why the philosophy of science is largely about keeping us
> from lying even to ourselves.  That's why Solomonoff's proofs from the
> 1960s were such a profound advance in the philosophy of science: They
> offered, for the first time, a scientific model selection criterion that
> could be used to create the right *incentives* given an agreed-on dataset.
>
> The key word here is *incentives*.
>
> What:
>
>    - were the *incentives* Charlie was under to go hunt down that old man
>    in a nursing home?
>    - would be the *incentives* to stop wasting our time yammering at each
>    other in prose about whose narrative is to "rule them all" in the sense of
>    providing predictions of the consequences of policies imposed on
>    non-consenting subject populations
>    <https://fairchurch.org/AProtestantInstauration.pdf>?
>    - would be the *incentives* to publish the scripts people used to
>    clean data rather than simply declaring that they had "cleaned the data"
>    and described in *prose* their "cleaning policies"?
>
> All of these incentives can align with lossless compression as the metric
> for awarding MONEY to finance quality assurance for entities like the DoE's
> Energy Information Administration.  Let's take Charlie's example of the
> investigation that required tracking down a guy in a nursing home:
>
> The result of that investigation was a *computer program* that encoded
> the data standard used to convert the data into a form commensurate with
> the rest of that dataset.  Charlie could have published that computer
> program along with the original "dirty dataset" and presented the "cleaned"
> dataset as simply a tabled (encached) intermediate computation.  This
> cleaning algorithm would have had a length and would have obviated any
> *arguments* Charlie might have had with competing interests regarding the
> notion of "clean" vs "dirty" data.  This would make it easier to reach a
> consensus on curating all data under consideration when deciding which
> macrosocial "narratives" to impose on non-consenting peoples without a
> control group and a phase 1 trial for safety, let alone a phase 2 trial for
> the *efficacy* of such "scientific" policies.
>
> That said, I've dealt with the Steel Man argument.
>
> However, there is also the straw man Matt set up regarding "video
> compression" which has two aspects:
>
> 1) The idea that human perception is the standard for defining actual data
> content of video data -- hence lossless compression is of no use in
> selecting the best model is beside the point.  I'm not proposing to have a
> bunch of Kelly Girls look at a bunch of macrosocial data and decide what
> they *feel* is important to predicting the consequence of policies.  This
> standard simply does not pertain to the definition of "noise" for
> macrosocial model selection.  If it has any relevance at all, it is at the
> *data* selection stage rather than at the *model* selection stage.
> 2) The idea that the  Ben Rudiak-Gould notion of "noise", while invalid
> for Wikipedia (as Matt correctly argued in response to Ben), is *valid*
> for video data is obviously false since there are papers showing enormous
> gains in video compression based on 4D model induction
> <https://share.gemini.google/D6o8W3zXEtph>.
>
> This is all beside the point that even if these arguments were valid for
> data class X they would be valid for dataclass Y when the two aren't
> comparable in either orders of magnitude of quantity nor orders of
> magnitude of quality control -- as is the case of comparing video data to
> macrosocial data.
>
> Matt may even be correct that the lossy compression of the total text
> content of the Internet may have already produced AIs that would be
> superior to the governments under which we now suffer.  I'm all for
> enabling him to join together with consenting adults of like mind to *run
> their experiment on themselves* while we, who don't share their *faith* may
> pursue our own confessions.
>
> But so long as we're all suffering under *the one true church of social
> pseudoscience*, it is rather inhumane to deny us a means of holding that
> theocracy to its own proclaimed standards.
>
> On Thu, Jul 23, 2026 at 6:32 PM Matt Mahoney <[email protected]>
> wrote:
>
>> ...But that isn't the problem. The problem is that we can effectively
>> compress video by asking an AI to describe it and compress the text to
>> about 10 bits per second. Then you decompress by using the text to prompt
>> the video. This is lossy, of course, but close enough that you don't notice
>> the difference. The reason this works is that the human brain has a write
>> speed of 5 to 10 bits per second, the same rate that we can read or speak.
>>
>> This means that video is 1 part per billion content and the rest can be
>> safely discarded as noise. If we ran a lossless video benchmark then nearly
>> all the effort would be going into compressing the noise instead of
>> understanding the image. This is already a problem for the Hutter prize
>> where 30% of the text is synthetic or XML, HTML, and Wiki formatting whose
>> compression does not contribute to language understanding but is
>> nevertheless required to advance.
>>
>> I tried to think of examples where we could answer questions about social
>> policy like future population. If AI can collect all human knowledge, as it
>> seems to be doing, then it should be able to say what is best for humanity
>> better than any human could. But most policy questions are about the
>> allocation of resources, and are ultimately resolved by combat.
>>
>>
>> -- Matt Mahoney, [email protected]
>>
>>
>> *Artificial General Intelligence List <https://agi.topicbox.com/latest>*
> / AGI / see discussions <https://agi.topicbox.com/groups/agi> +
> participants <https://agi.topicbox.com/groups/agi/members> +
> delivery options <https://agi.topicbox.com/groups/agi/subscription>
> Permalink
> <https://agi.topicbox.com/groups/agi/T5b58bcc51c493d41-M59e813cd830043a853984f9e>
>

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/T5b58bcc51c493d41-M6156b1548ba8c3922b1746b6
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to