Jon Sterling
I am an Associate Professor at the Cambridge Computer Laboratory, and a Fellow, Tutor, and Director of Studies at Clare College. Realist of a larger reality.
I like categories, domains, and vintage computing.

RE: @tonofcrates@mastodon.social
I remember recently I saw some established academics saying things like, the LLMs can/will write better papers and reviews than us anyway.
I remember thinking, “Speak for yourself LOL”. These people really are telling on themselves… The entitlement is amazing: somehow they believe simultaneously that society should pay them a full professor's salary to do something that they are admittedly terrible at. Call me crazy, but if you can't do better than an LLM, you should probably step aside and open up a vacancy for someone who is actually good at their job and can make intellectual contributions to science.
God grant me the boldness of a University that requires faculty to complete "cybersecurity training" whilst also forcing on us the biggest security vulnerability of all: The Microsoft Office 365 Copilot App.
Here's a kind of 101 thing that a lot of people in the world of AI coding are missing, I think.
Question: What are the implications of the fact, "All the tests pass?"
Answer: It actually depends on how the code was written. Unfortunately, the salience of "all the tests pass" has a lot to do both with the strengths of agentic programming and the weaknesses.
If code was written without knowledge of the tests, and then happens to pass the tests, I think I would then be a bit confident that the code is not just passing the tests by coincidence, but is actually correct in a deeper sense that makes it likely to pass future tests that haven't yet been written.
When code is written with knowledge of the tests, or using the tests as a scaffold, then we have to be a bit more careful. Both humans and LLMs do this, but LLMs are better than humans at Monkeys Paw–style trickery, where you have something that is just totally wrong but nonetheless passes all the tests.
So when a codebase is rewritten agentically, the fact that tests pass doesn't make me very confident, unless a lot of the tests were held to the side and not exposed to the agent. This is just basic experimental science (control!), it's hard to understand why this is not obvious to everyone.
Unfortunately, having more tests exposed to the agent increases the quality of the results and the likelihood of correctness. So this is kind of a paradox — but it is a paradox that only applies to people whose method is to interact with the produced code only indirectly. It's easy to break the paradox if you treat the code as your own responsibility and only commit what you understand.
My wife asked me to tweet out the following: “The X-Files is the Law And Order of science fiction”
Had a read of two of my students' draft theses this week and this stuff is looking really cool… Very proud of my students.
Not to step on a culture war landmine, but I keep seeing people claiming that the reason we know Memnon in Greek mythology was black is that he was portrayed as such on various amphorae. The character of Memnon is indeed African, and he was certainly considered to be black or at least dark-skinned by writers in antiquity. That much is true.
But the amphora art has nothing to do with it… Almost everyone on there is depicted in black pigment. That was the art style. It's called black-figure pottery.
It is best for people who learned everything they know about the ancient world from TikTok videos to refrain from commentary, especially if their aim is to attribute stances in the present-day culture war to people in antiquity for whom skin colour held none of the perverse salience that it holds today.
I don’t really like property-based testing, insofar as it involves randomized inputs.
1. Randomized tests are inherently flaky. If the code is correct you don’t notice, but if the code has a bug in some edge case, then the tests (by definition) will pass and fail nondeterministically. Flaky tests have just one destination in my judgement: the bin.
2. Picking a realistic input distribution can be as difficult as writing the code you are testing. Sometimes more difficult. It is easy to get false confidence because the marketing says you are testing a “property”, but the implications of the test are in reality not easy to assess.
3. QuickCheck/Quviq does have the whiff of a consultancy grift, to some extent, does it not? Enough said…
Let’s be clear: I’m in favour of any method or tool that can empower people to improve the quality of software. But I have not found property based testing to be as useful or applicable as it is often billed to be.
I also think there is a lot of potential for proving something correct with a finite number of tests, as in the polynomial testing principle that @pigworker@types.pl discussed here: https://personal.cis.strath.ac.uk/conor.mcbride/PolyTest.pdf
I saw a rook today on my way home! First time I’ve ever see one in person. Beautiful birds!
I guess it's old news to complain about how carelessly engineered Apple's Music.app is, but seriously…
After a few hours of reading tormented and writhing student proofs, it is of such immense satisfaction to sit down with a blank sheet of paper and write the proper proof in one go with no mistakes, everything in its right place, scaffolded properly so that it can be read from start to finish.
It’s Seitan-making day. Wife makes the dough, I wash it. I can’t wait to eat it.
Weeknotes 2026-W19
https://www.jonmsterling.com/2026-W19/
+ Algebraic foundations of bidirectional elaboration?
+ Happy 100th Birthday to David Attenborough!
+ Reading Corner: Annihilation, Authority (thanks, @nilesjohnson@mathstodon.xyz)