On 2026-07-26 11:08, Richard Lewis wrote:
> see eg
> 
> https://en.wikipedia.org/wiki/Jacobian_conjecture#Counterexample_for_n_%3E_2
> 
> most humans apparently thought this conjecture was true, but an LLM
> provided a counter-example: i dont think there is a meaningful sense in
> which this was simply reproducing training data (the importance of this
> is another matter, and subject to the usual hype)

It's been widely speculated that the huge leap of coding/reasoning model
performance since last ~autumn is due to Reinforcement Learning with
Verifiable Rewards (RLVR), a post-training technique.

The early models where certainly training-dominated and the term
'stochastic parrot' might have been fitting, but nowadays they are much
different beasts.

Best,
Christian

Reply via email to