Anthropic and OpenAI have billions, maybe trillions at stake from their safety 
judgement and auto-approvals being generally regarded as sound.  (More than a 
single employee that could be fired for dropping crucial tables in a SQL 
database.)   I recognize that big companies have armies of lawyers, but I'd 
still argue the inhibitory mechanisms for risk can be expected to be better at 
scale.  The real power is being able to walk away and let them grind for hours 
or days.  Once you've seen it do a large project end-to-end the world has 
shifted in a basic way.

Meanwhile there's grok/cursor/commander that is loved by some because it is so 
fast.  It is fast because it is small and coding-focused and has no safety 
controls at all.  People regularly run it off any chain.
But people do lots of things in the real world that are basically reckless.  
Sometimes just for the thrills.

-----Original Message-----
From: Friam <[email protected]> On Behalf Of glen
Sent: Monday, August 3, 2026 8:12 AM
To: [email protected]
Subject: Re: [FRIAM] cognitive offloading

So far, I remain tightly clamped on what I allow [Claude|Open]Code to do, both 
edits and execution ... and especially skill downloads. That 
approve-every-action policy used to be paranoia. I was afraid. But it's become 
*interesting* at this point. E.g. Claude wanted to install a headless browser 
snapshot generator so it could look at a web page rendered from some data to 
see if it "looked" the way it intended it to look. I didn't allow that, of 
course, because I'm the worst kind of micromanaging boss. But that it wanted to 
do that all automatically is fantastic, literally a realized fantasy. The same 
would apply to any other grounding check, be it formal or informal.

But what would a consult-the-experts REPL look like in Eric's setup? I asked 
Mistral's Vibe to analyze Duede et al and it gave me some interesting hooks for 
testing their model. A harness might even have implemented the model 
(numerically or formally or both) to verify their curves, maybe search the 
literature for social data/trends, and perform some kind of sensitivity?

But going back to Cody's idea that a robot prolly won't have nociceptors in odd 
parts of its foot, we have a ways to go before we'll get to harness granularity 
imagined by The Matrix. Which parts of the context *can* be challenged? Which 
parts should be challenged (versus accepted due to low RoI)? Etc.

Anyway, I just can't imagine people *still* writing this stuff off as 
"autocomplete", an insult I still see almost every day. It's like those kids in 
school telling me my HP was a calculator. There's a difference between your TI 
and my HP, even if that difference is occult.

On 8/1/26 7:51 PM, Marcus Daniels wrote:
> I use both Claude Code and Codex heavily, and the main current difference I 
> observe is how hard Codex seeks grounding on factual things.   A simple 
> example that came up for me today:   I had some large builds that had used 
> all storage in a temporary workspace.  So, when the AI was waiting for my 
> response, I moved the build to a partition with a lot more disk space and 
> then made a symlink to it.   Before the move, the AI (GPT 5.6 Sol) had seen 
> some dereferenced absolute paths because of the tooling.  It would have 
> remembered them just fine had I not done this.  But Codex, currently more so 
> than Claude Code, is paranoid/conscientious about its grounding.   Between 
> larger bursts of work (even when context is not exhausted), it basically acts 
> like it is starting over from notes, and rechecks the environment.   It 
> caught the move of the paths and adjusted the absolute paths to the new ones. 
>   It’s a dispositional tendency that helps long horizon work at the cost of 
> some token waste.
> 
> Now in just the last few days, OpenAI had an announcement on just this 
> issue.   (It is what Claude Code does – so they are saying that their 
> model is smarter than it looks.)
> 
> https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores
> / 
> <https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-score
> s/>
> 
> Anyway, my guess is that many users of AI in academia are not using harnesses 
> in the way builders are.   With a harness, the AI is not just responding to a 
> vague goal from a muddled user using its general knowledge.  The harness 
> provides constant updates to entailment.    For example, these proofs were 
> managed by Lean 4, so they can be iterated in this way, and they pass a 
> deterministic checker.
> 
> https://github.com/openai/ten-proofs/ 
> <https://github.com/openai/ten-proofs/>
> 
> I’d say the solution to the muddled user problem is AI review of papers.   As 
> a part of ingest, the paper would get a critical review from one or more 
> perspectives.   Like a continuous integration system for software projects 
> but applied to paper quality.  A small change to something like arXiv.
> 
> What I think will probably happen is something like the blackboard 
> architecture with LLM writers and reviewers.   And it will all be too 
> fast to keep up with, and so the best thing will be to go on bike ride 
> or walk the dog.  (First is down, now for the second.)
> 
> *From: *Friam <[email protected]> on behalf of Santafe 
> <[email protected]>
> *Date: *Saturday, August 1, 2026 at 6:37 PM
> *To: *The Friday Morning Applied Complexity Coffee Group 
> <[email protected]>
> *Cc: *David Eric Smith <[email protected]>
> *Subject: *Re: [FRIAM] cognitive offloading
> 
> It depends on what you’re in it for.  This to Marcus’s point/question as well.
> 
> There’s various stuff that has triggered something like thoughts, some of 
> which seem to me a little bit constructive, and this list has been the place 
> I would normally want to toss such stuff back and forth.  But it can be read 
> as “concerns that come up to an academic”, and I am aware of the rising tide 
> of sentiment (society in general, and not only the MAGAs) that academics all 
> live in a fake world anyway, and the society would be better off if they were 
> all just shoved off into the sea and dispensed with (except the part that 
> maybe businesses feel like hosting and charging a fee for).
> 
> But what the hell.
> 
> The early part of last week I lost maybe the more-quality part of two 
> days “reviewing” a submission that I accepted as a “paper”, before realizing 
> it was all slop.  Increasingly, that means it wasn’t necessarily wrong.  The 
> LLMs now write a lot of equations etc. that are correct.  Context is that at 
> least one of the names probably points to a real person, a very junior 
> Chinese expat in a visitor institution of a medical school (that I never 
> heard of, but probably has some repute) in Moscow.  How precarious a life 
> would you like to think that is?  I don’t know if the second submitter name 
> (also Chinese, but common enough to be widely used) corresponds to an actual 
> person or not.  Claim in the manuscript was that it is about multi-neuronal 
> trajectories during mesoscale brain-state changes (say, learning, making a 
> choice, etc.)  The only other paper by the same submitter name is on Kerr 
> black holes.  Not published, just stuff at researchgate and the earlier 
> author on the other one gave a talk somewhere.
> 
> Anyway, once I switched the necker cube to realizing it was all slop, 
> it was of course overwhelming that it was so.  After looking at it in that 
> frame for a bit, I concluded that there is a good likelihood that no actual 
> person did any actual work at all.  Probably the submitter asked one of the 
> LLMs to choose something that would sound like a project, to go do a Gish 
> gallop through literature that could be padded into an “introduction”, though 
> none of it was ever used for anything, put in a bunch of platitudes about the 
> valid interpretation of different methods (also never used or revisited again 
> in any way) maybe to write some code (which maybe the person ran or maybe 
> not), generate some figures that purport to be the output of those 
> simulations (which, for various technical features I strongly doubt; I think 
> they were made up), and then package it all into a “paper” that the submitter 
> could submit.  There were large sections in it that were in a somewhat 
> obscure area, which read like plagiarism of our work in that area (my way of 
> saying things is non-standard enough that plagiarism of me doesn’t sound like 
> plagiarism of most larger communities, and I know parts of the math community 
> who work in this area, and don’t sound like that us at all), but since 
> _nobody at all_ was cited in that section, of course I have no way to know.  
> The submitter probably doesn’t even know what the community is, as the LLM 
> maybe never gave her any of that information.  It just, oracular-like, 
> plunked down various sentences and equation forms.  On the other hand, this 
> is a journal I never heard of and that never heard of me, so I expect the 
> reason I was asked to review is that the LLM suggested us as potential 
> reviewers as one of its prompt replies.
> 
> Fun note: I explained all this to the handling editor (and the parts that 
> don’t involve my own identity to the submitter as well, since I don’t feel 
> like coming under harassment by a troll farm, if this is not just the 
> desperate action of one kid trying to survive).  This is a Springer journal, 
> and a colleague shared the opinion a while back that Springer, despite having 
> some legacy good titles, has degenerated to being nearly a predatory 
> publisher.  The “editor”’s response to my review and others was “Revise”.
> 
> Anyway, what are the three thoughts (sorry for that long preamble) following 
> from any of this?
> 
> 1. Academics give lip service to, but I’m not sure they have embodied in the 
> gut, the absolute tsunami of sewage that they are all about to get submerged 
> under.  They may come to wish they had been shoved into the sea.
> 
> 2. Why did I waste nearly 2 days before realizing what all this was?  This 
> matters to me, because it speaks to how awful academic writing has been 
> forever.  The Gish gallops, the padding, the oracular style?  It’s always a 
> goddamned cryptology exercise, trying to figure out what the hell something 
> means, or how a claimed result was arrived at.  Sometimes it’s because the 
> authors are geniuses; don’t know what to do about them, because they 
> genuinely can’t understand why things need to be explained for other people 
> to follow them.  Sometimes it’s because the authors are lazy, and don’t feel 
> like explaining well if they think they got something right.  Most of the 
> time it’s probably hazing, by authors who want to look impressive by making 
> things harder for the reader to get than they need to be.
> 
> (For a large part of this, I blame Bourbaki, which Persi Diaconis — 
> blessings be upon him — gave me permission to do despite my being a 
> mere physicist.  Bourbaki were like the Grover Norquists of math.  
> They had two purity tests:  Always remember that: 1) any message from 
> a sufficiently advanced society will be indistinguishable from magic 
> to a lesser society; and 2) any maximally compressed message will be 
> indistinguishable from a random string of bits.  Mathematicians must 
> always test their written output against these two criteria and try to 
> maximize both.  This umbrella then gives cover for lots of bad writing 
> that is actually done for other reasons, along with the inevitable 
> difficulties of specialization and siloing.)
> 
> It can take me anywhere from weeks to forever to understand a single paper.  
> If review is supposed to mean notarizing that the contents of a paper are 
> correct, and I get three new requests a week, clearly those times don’t 
> scale.  So inevitably, getting as close as I can, and admitting where I am 
> still stumped, I give writers the benefit of the doubt, that some very costly 
> figure that I have no possible way to reproduce (outcome of years of work in 
> some specialist lab) is what they claim.  In the age of slop, I don’t give 
> anybody the benefit of the doubt.  But that means the review timescale 
> expands back to its authentic self, which is weeks to years (or maybe 
> forever) for each paper.
> 
> The number of things we will have lost, in a society where trust is no longer 
> possible, toward anybody in regard to anything, is so large that I don’t 
> think anybody is seriously reckoning with it.
> 
> 3. But then, what to do?  My thought was that we (what “we”?) should 
> create a new journal.  ONE new journal only.  It’s called “Vibe Science”.  
> All the submissions will be auto-generated by LLMs.  If people want to serve 
> them by feeding prompts, that’s fine too.  Kind of Matrix-like (or 3rd-world 
> porn-flagging exploitation-like) jobs people already take.   A new 
> gig-subeconomy.  Maybe they can be fed a meal a day in payment.  Then LLMs 
> can also do all the reviewing, and can send messages to each other (or maybe 
> there is only one LLM at the end; does it even use messages any more?).  
> They/it can decide what is good and should be accepted, and what should not 
> be accepted.  Why not?  There’s no reason to have more than one journal, 
> since internally the LLMs can sort what is “in” or “across” any discipline or 
> disciplines, and to invent new typologies as needed for their internal 
> indexing.  People can’t read it anyway, so there is no reason to make it 
> human-indexing-friendly (or maybe even tractable).
> 
> Then the accepted stuff can all be archived somewhere, on more Mongolian rare 
> earths with fossil- or nuclear-powered cooling towers and fans.
> 
> What will we do with it?  I guess people with no other survival work, like 
> the e-waste-recyclers in India or the Philippines, can trawl through the 
> lists of claims, and see if there is anything the human world would like to 
> regard as “true” or “potentially useful” or (here, for the unemployed 
> artists) “interesting”.  Maybe that will become what PhD training aims 
> toward.  Now of course, there isn’t any reason the data-trove can’t just be 
> made self-activating in the world of physical robots, so I’m not sure what 
> the vetting by the recyclers would unpack to.  But maybe some decision-maker 
> somewhere uses that.
> 
> Something along this line is my guess at the set of things that will have to 
> be chosen and decided, in a quite near term, but quite a number of people.  
> Or maybe decided by 10 people, and everybody else can scramble and see if 
> there is a survival trajectory through the new wreckage.
> 
> At the end of all this, though, I think it pushes us to front 
> something we should always have been thinking about, but have had the luxury 
> of not thinking much about before now.  There were good uses for calculators, 
> and even though it deprived some people of the ability to do fast addition 
> in-head, we continue to do math, with new frontiers possible.  There will be 
> some version of that for the LLMs as interface, and much more general AI 
> tools for empirical results and their representations.  But then what is a 
> system for thinking through what a person can want, or might want, as a 
> participant in the order that results?  The sort of emotive tropes that seem 
> popular now, or the online gurus of everything, kind of leave me unimpressed. 
>  All this common-language and lore and culture that we have as hand-me-down 
> from time immemorial is what came through the filters of people’s needs to 
> reason out these things, in the circumstances of the times.  To the extent 
> that those circumstances are going to be different — not clear to me, since 
> ecological collapse and mobs of the dispossessed could just thump it all into 
> rubble — a different version of something similarly filtered will need to 
> come about somehow.
> 
> Back to the opening question: For myself, I would not have wanted to 
> be one of the people prompting the LLM to invent a question and write a bunch 
> of junk that I could send to a Springer journal.  I certainly do not envy her 
> life as a junior Chinese expat in a cordoned-off sub-institute in Moscow.  
> Less dangerous and demanding of wits to try to become a gang leader.  Even 
> though I suspect some of those people have talent, and many of them I would 
> like and appreciate if I knew them.  I still take weeks or months fiddling 
> with some annoying thing to see if a certain kind of field theory can be 
> coherently defined into existence, and whether it gets the answer to some 
> empirical question about rare statistics right.  I guess, the fact that that 
> is what I want to do, and that it sometimes might also get something 
> empirically right, is why the society increasingly wants to dump the 
> academics.  But it still is what I have accepted paycuts and 
> spontaneous/episodic unemployment, and various other nuisances, to somehow 
> put aside time to continue doing.
> 
> Eric
> 
> 
> 
> 
>> On Aug 2, 2026, at 6:49, glen <[email protected]> wrote:
>> 
>> I had similar thoughts. It's cool and all. But I can't help but wonder why 
>> they did NOT use more AI.
>> 
>> On 7/31/26 9:44 PM, Marcus Daniels wrote:
>>> 2607.17397 is a weird one.   The true believers among us might wonder why 
>>> they are still working on their papers at all.
>>> -----Original Message-----
>>> From: Friam <[email protected]> On Behalf Of glen
>>> Sent: Friday, July 31, 2026 10:19 AM
>>> To: [email protected]
>>> Subject: [FRIAM] cognitive offloading As preamble:
>>> 1) Protecting our FLOSS commons from LLMs 
>>> https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fblog.codeberg.or
>>> g%2fprotecting-our-floss-commons-from-llms.html&c=E,1,PSzw3UoOT3SOXS
>>> aXTGDk-zPlNdIkPx58sbCqKgH3fqrZLLmTvR9c1q0G-qeL_GFsAE-9sNkr7RvNRtkbG7
>>> 56llp3STEvrsBt9Ztbt3W424zCKnZWhQ,,&typo=1 
>>> <https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fblog.codeberg.o
>>> rg%2fprotecting-our-floss-commons-from-llms.html&c=E,1,PSzw3UoOT3SOX
>>> SaXTGDk-zPlNdIkPx58sbCqKgH3fqrZLLmTvR9c1q0G-qeL_GFsAE-9sNkr7RvNRtkbG
>>> 756llp3STEvrsBt9Ztbt3W424zCKnZWhQ,,&typo=1>
>>> 2) Science Direct Summary
>>> https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fwww.sciencedirec
>>> t.com%2fscience%2farticle%2fabs%2fpii%2fS1364661316300985&c=E,1,Rez0
>>> hef2TLaMMg2jA_kk6d6KHVgSHZk75ETTunCX7Dt5GuSi33eko0ka8iSoiXEajFB0XSer
>>> QBNPvL3p8DqGXHdWrysTvUyNzne192tlWA,,&typo=1 
>>> <https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fwww.sciencedire
>>> ct.com%2fscience%2farticle%2fabs%2fpii%2fS1364661316300985&c=E,1,Rez
>>> 0hef2TLaMMg2jA_kk6d6KHVgSHZk75ETTunCX7Dt5GuSi33eko0ka8iSoiXEajFB0XSe
>>> rQBNPvL3p8DqGXHdWrysTvUyNzne192tlWA,,&typo=1>
>>> But this is the main driver for this post:
>>> 3) The unintended consequences of large language models as a 
>>> labor-augmenting technology in science
>>> https://arxiv.org/abs/2607.17397 <https://arxiv.org/abs/2607.17397>
>>> I think the Codeberg hypothesis (and stance) about one-off code provides a 
>>> clear foil for disambiguating what code/software actually *is*. There's a 
>>> difference between collaboratively developed code and "interpretable" 
>>> methods. It seems like 2 subtly different meanings of the word "open". 
>>> Codeberg is (effectively, if not purposefully) saying they are not a 
>>> "repository" in the sense of, say, 
>>> https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fvivli.org%2f&c=E,1,dWHhPQRYDujbxzdX0_0zLPwURhJoU16RdDayBZS4TaF5bo89ZIRV6_Aqm3S_pPkFx_myz98EbKiCplE-Ezjz7tKrBm_ewCfGM8kSONgRrH3KjA7nZ8nPNzYo&typo=1
>>>  
>>> <https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fvivli.org%2f&c=E,1,dWHhPQRYDujbxzdX0_0zLPwURhJoU16RdDayBZS4TaF5bo89ZIRV6_Aqm3S_pPkFx_myz98EbKiCplE-Ezjz7tKrBm_ewCfGM8kSONgRrH3KjA7nZ8nPNzYo&typo=1>.
>>>  Now, I'm a fan of opinionated tech. But Codeberg's stance seems a bit 
>>> reactionary to me. I mean, it's fine. It's their site, their hardware, etc. 
>>> So I'm not objecting ... simply processing.
>>> But if we move on to Duede et al (3), the paper irritates me because I like 
>>> to believe I have an empirical bent. I'm having trouble formulating their 
>>> model such that its premises are backed by data and that it makes testable 
>>> predictions. It's my problem, I'm sure. But even in the abstract, they say 
>>> "researchers will become more selective about what they publish". I feel 
>>> like they should have prefaced that with "if our model validates" or 
>>> "analysis of our model argues that" ... or somesuch. IDk. Inside the paper, 
>>> their assertions are softer and more palatable ... but they still seem to 
>>> be reifying.
>>> I've included (2) because, from my perspective the LLMs are a fairly 
>>> straightforward advancement in the lineage of the pencil, or the printing 
>>> press, or the internet. We imbue our talisman(s) with our hopes and fears, 
>>> then ratchet our thinking up and never look back (until someone slips a 
>>> hero dose of LSD into our drink at the party). A large proportion of us 
>>> will be catastrophically damaged if our ratcheted scaffolds collapse. (Some 
>>> scaffolds are so robust no amount of truth or fact will collapse it at 
>>> all.) But at least some subset of dorks will always be fascinated by little 
>>> details like manually selected fonts or assessing whether some modeling 
>>> assumption is *actual* as opposed to convenient. We're each free to offload 
>>> whichever homunculus we find irritating. But I find myself in awe of those 
>>> who relish irritation as a grounding to reality.
>> --
--
ꙮ Mɥǝu ǝlǝdɥɐuʇs ɟᴉƃɥʇ' ʇɥǝ ƃɹɐss snɟɟǝɹs˙ ꙮ

.- .-.. .-.. / ..-. --- --- - . .-. ... / .- .-. . / .-- .-. --- -. --. / ... 
--- -- . / .- .-. . / ..- ... . ..-. ..- .-..
FRIAM Applied Complexity Group listserv
Fridays 9a-12p Friday St. Johns Cafe   /   Thursdays 9a-12p Zoom 
https://bit.ly/virtualfriam
to (un)subscribe http://redfish.com/mailman/listinfo/friam_redfish.com
FRIAM-COMIC http://friam-comic.blogspot.com/
archives:  5/2017 thru present https://redfish.com/pipermail/friam_redfish.com/
  1/2003 thru 6/2021  http://friam.383.s1.nabble.com/

Attachment: smime.p7s
Description: S/MIME cryptographic signature

.- .-.. .-.. / ..-. --- --- - . .-. ... / .- .-. . / .-- .-. --- -. --. / ... 
--- -- . / .- .-. . / ..- ... . ..-. ..- .-..
FRIAM Applied Complexity Group listserv
Fridays 9a-12p Friday St. Johns Cafe   /   Thursdays 9a-12p Zoom 
https://bit.ly/virtualfriam
to (un)subscribe http://redfish.com/mailman/listinfo/friam_redfish.com
FRIAM-COMIC http://friam-comic.blogspot.com/
archives:  5/2017 thru present https://redfish.com/pipermail/friam_redfish.com/
  1/2003 thru 6/2021  http://friam.383.s1.nabble.com/

Reply via email to