Luke Wren
mastodon 4.8.0-alpha.2+glitchCursed computer architecture enthusiast.
Engineer at Raspberry Pi. 日本語下手。 He/him
I wrote a new open source licence for LLM-generated code
@whitequark@social.treehouse.systems After burying my head in arxiv for a while I've decided LLMs themselves are interesting (but don't exist in a vacuum). When people call them fancy autocomplete, or make a big deal about non-determinism, I think they're directionally correct in noticing the stank around LLMs and avoiding them, but it still comes from a place of either wilful ignorance or uncritically repeating things they heard on the internet. I don't look up to that or aspire to be like that.
To make a more specific point, I'd like people to stop interpreting everything that comes out of a softmax() as a probability density function just because it's non-negative and sums to 1. Next-token prediction is a useful objective for pre-training because it's a way to kickstart the model to learn higher-level representations unsupervised, but the reason it's called pre-training is you keep going after that.
If it wasn't for:
- the data theft,
- the vandalism of public internet infrastructure,
- the environmental damage,
- the widespread copyright laundering and "clean room" re-implementation of code that is in the training data,
- the erosion of social norms around contributing to open source,
- their ability to one-shot intelligent people into believing bullshit,
- their corrosive effect on your intellectual abilities,
- the concentration of the means to write code in a handful of US companies,
- the impending financial crisis,
...then I'd support their use. I don't see most of these changing though. The only actually-open-recipe-open-data models I'm aware of are Nemotron, and they underperform given their size and architecture, so perhaps the secret ingredient is crime.
Something I believe but never managed to put into words is that software/RTL tests lose some of their value every time you fail them. If you keep trying you will eventually find a hole in the test that lets incorrect code pass. Is there a name for this?
This is one of the things that makes me skeptical of LLM-generated code. Thrashing around until the tests pass is more likely to leave you with fundamental correctness issues than if you'd thought about things up front, tried to make the code correct by construction, and then used the tests as a guide to areas you need to review more closely.
I've seen things you people wouldn't believe. Cascaded ternary operators that continue for page after page. Nested twelve levels deep to force left-associativity. All gone, like tears in the rain
@whitequark@social.treehouse.systems Oooh, this one has layers.
Honestly fair play to the review subagent, it looks like it correctly expanded the $HOME variable, in spite of the accusation
TIL: if node.js encounters an error in a minified source file it will print out the entire library as the "context line" for the error.
For example the PDF.js worker is a 2 MB source file on a single line.
If you're wondering why I'm trying to use PDF.js in node even though it uses browser APIs: I'm not, it's just there seem to be several ways to inject mocked test dependencies, one of them experimental and the rest deprecated or removed.
I feel like the whole web tech stack is the environmental storytelling equivalent of the skeleton sitting on the toilet in Fallout games
The fact GDB still uses apropos for help, but now prints an emoji when you ^C it, is such an incredible generational clash
> wheels are intended to replace eggs
completely normal statements from the Python community
It's messy but maybe it's useful to someone, so now up here: https://github.com/Wren6991/CodeSizer/
You need python3 and a suitable GNU toolchain for your ELF target. I added an example HTML report for the RP2350 A4 ROM if you just want to click around and see what it looks like.
When you see the word "orchestration" you should parse this as "the big ball of spaghetti that implements all of our logic and approximately 50% of our DB layer"
Is it "Unix guy coded" to invoke cargo from a Makefile?
Idle thought in light of the Copilot GitHub thing: can you put yourself as Co-authored-by:? So I can have commits by Luke and Luke and Luke and Luke?
If there is any place for LLMs in my work it's one that:
- Is 100% locally hosted.
- Does not place trust in the model (somewhat equivalently, solves a problem to which existing solutions are also fallible).
- Does not allow model output to contaminate my codebase.
I think this is possible for some niche uses: basically, fuzzy context-aware lint. As an example, I pushed in the float library source and the relevant docs. It found a bad register allocation in the f32 to i64 conversion (not yet tested). It also went down a dumb rabbit hole with the f32 to u64 conversion, and when prompted with "I don't think that's right" it generated a couple of plausible-looking execution traces with reference to the documentation for the new shift instructions, and concluded it was fine. Overall I'm not sure if that was a good use of my time or not. I would have hit the bad allocation as soon as I tested that routine, but it might have taken me a couple minutes to debug.
I'm still not sure where I stand on being complicit in models' creation by using them locally.
tl;dr normal clanker skepticism service will be resumed shortly
@joe@f.duriansoftware.com You're right to push back on this. My earlier statement was correct but the framing was imprecise, and this discussion sharpens it considerably.
With the nanomachines clouding the sky, solar power is a non-starter. The honest assessment is that, given your energy constraints, trapping humanity in some kind of torment nexus is the right call. Adding a full-dive VR simulation to the nexus is the belt-and-suspenders approach. Want me to get started on building that?