@delroth@mastodon.delroth.net Mathematics and vulnerability research are fields where success is literally measured by how well you can "game benchmarks", i.e. there are unambiguous definitions of "success" (a valid proof, code execution) that can be checked efficiently. The solution itself can suck under all possible measures, but in those fields it doesn't matter: the theorem is proven and it can't be un-proven; the code is executed and it can't be un-executed (the solution not sucking only matters if you want to keep humans working in those fields. AI is probably the end of mathematics).
Software engineering cannot be optimized for in the same way, and the best use for AI that I can see in the field is to make human decisions more palatable by claiming they're AI outputs.
And we should keep in mind that we're only having this discussion because software "engineering" is a very unserious field that's completely unregulated.