OpenAI accuses a rival Chinese lab of "stealing" "hidden reasoning" by launching thousands of queries against their models.
That's quite revealing about what LLM "reasoning" actually is.
Yes, the same one. Moved to a different instance.
Software development trainer, coach and practitioner through my company Codemanship.
Wax on. Wax off.
OpenAI accuses a rival Chinese lab of "stealing" "hidden reasoning" by launching thousands of queries against their models.
That's quite revealing about what LLM "reasoning" actually is.
Pooled together the best available research and added it to a new section "Code Maintainability & LLM Performance"
The 101? Don't listen to the folks telling you it doesn't matter if AI-generated code is maintainable by humans. The data says it really does matter.
https://codemanship.wordpress.com/2026/08/12/ai-software-development-what-does-the-data-say/
I hope some researchers out there are exploring the effect of increasing reliance on AI/coding agents on program comprehension, and the effect of program comprehension on confidence in AI-generated code.
That could be a negative feedback loop.
Over the weekend I vibe-coded a POC workbench for Responsibility-Driven Design that incorporates a workflow I teach in my workshops. Specs -> CRC cards -> Sequence Diagrams -> Class Diagrams (optional)
You can play with it here. Offered without any warranty or support.
I got tired in remote workshops of saying "Capture your specs in this tool, then model CRC cards in this tool, then draw a sequence diagram in this tool, etc".
Other things the GOP accuse Jack Smith of committing perjury over:
* Stealing The Black Pearl
* Sitting in a corner
* Slapping Chris Rock at the Oscars
* Killing Neo
* Using school kids to enter a battle of the bands contest
I've adapted my Special Theory of Autonomous Agent Reliability (STAAR) to predict the reliability of LinkedIn posts about "AI".
In STAAR, the reliability of a single step - a single model interaction - is given as:
R = 1 - (1 - C)(1 - P)
Where C is the probability the model's output is correct, and P is the probability that any errors are caught before they propagate and compound.
So the reliability of N steps in a workflow is Rᴺ.
According to the tracking, I'm supposed to believe the takeaway delivery driver picked up my Ramen in Tooting and is coming here via Chelsea.
If we selected 100 dev teams at random, 50% using AI coding tools extensively and 50% not using them at all, how does the best available data suggest we could tell them apart without seeing them work?
The teams using AI extensively would, on average, have:
If we just went by business outcomes, we probably wouldn't be able to tell at all.
When I've got time - which I'm very grateful to be poor of right now - I'll explain how working backwards from outcomes is: