HalvarFlake
I do math. And was once asked by R. Morris Sr. : "For whom?"
Accidental two-time founder. Mathematician by education. Infosec luminary (has-been?).
I don't know what y'all are doing, but my experience with Fable and Sol is that they still need a lot of judgement, taste, guidance, and straight up correction to not mess up the codebases they operate in beyond repair.
I finally managed to write something about my recently deceased dear friend Felix 'Fx' Lindner.
I wrote some lines about mitigating vibe-coding risks by adopting a development model inspired by old-school computer breakin folks:
https://addxorrol.blogspot.com/2026/03/slightly-safer-vibecoding-by-adopting.html
I am insanely proud about the following, even though I am not at all involved any more:
https://opentelemetry.io/blog/2026/profiles-alpha/
Working with the optimyze team was so awesome.
I feel very compute-constrained these days. In the past, I was "implementation time for experiment" constrained: I had an idea, but I needed to carefully weed out the ideas in order to focus on those I could implement for testing.
Now that code agents allow faster experiment implementation, I am running out of compute all the time.
Visualizing the decision boundaries and final result of the same problem optimized by Muon and by AdamW is fascinating. The solutions are qualitatively very different. AdamW has visibly redundant and useless neurons, Muon less so, but AdamW's reconstruction looks more generic.
I am seeing up to 30% wall-time fluctuation running the same code on the same data on the same VM type in GCE.
That seems crazy to me.
What's the worst variation you've seen? How do you deal with these things?
I mean, for any benchmarking you need to run the clickhouse tomato benchmark protocol? (E.g. run two instances of the same software on the same VM so they're exposed to the same noisy neighbor effects).
This LLM security research from Anthropic also has important copyright implications:
LLMs are reshaping software dev. I don't buy "the end of software dev": Project ambition will grow dramatically.
Ancient Egyptians could build the Pyramids but not the Empire State Building.
Pre-LLM software will be viewed like we view the Pyramids.
The weirdest observation: I generated movies visualizing the polytope boundaries for ReLU networks using Muon and AdamW.
Same experiment, same data, same random seed. The difference is the "crease pattern" that the optimizers produce.
The AdamW videos compress to 1/6th the size of the Muon videos. Something AdamW is doing allows the crease visualisation to be compressed well, but not Muon. This is the weirdest observation ever.
I am digging through some PyTorch code, and ... I have to say: The Nvidia profiling infrastructure is very much no good. Terrible UX, terrible at answering the questions I need answered.
Makes me pretty bullish about zymtrace.
Ok, I have a 2048x3072 pixel video that shows the differences in training dynamics between Muon and AdamW, and they are fascinating to watch. Where can I upload/share such a video without it getting compressed to shreds?
ArXiv AI papers should list a "flops needed to reproduce experiments" in the abstract.
I like going for a walk or a beer alone. Weirdly, feelings of boredom or loneliness are much more common for me when in the presence of others than when I am alone. Alone, I feel calm, and peace, and free.
Slides for my QCon talk: https://docs.google.com/presentation/d/1wOT5kOWkQybVTHzB7uLXpU39ctYzXpOs2xVyD4zuYXY/edit?usp=drivesdk
There's a video on the QCon website, too, for virtual attendees.
The weirdest observation: I generated movies visualizing the polytope boundaries for ReLU networks using Muon and AdamW.
Same experiment, same data, same random seed. The difference is the "crease pattern" that the optimizers produce.
The AdamW videos compress to 1/6th the size of the Muon videos. Something AdamW is doing allows the crease visualisation to be compressed well, but not Muon. This is the weirdest observation ever.

