Elektrine
Log in Register
Paige Chat Timeline Gallery Friends Email Drive DNS Private DNS Domains VPN Kairo Nerve
Remote

Dan Luu

@danluu@mastodon.social
mastodon 4.8.0-nightly.2026-10-06
  • Open on mastodon.social
0 Followers
0 Following
37 Posts
Joined November 24, 2017
Blog:
https://danluu.com
Patreon:
https://patreon.com/danluu
Open post
Dan Luu @danluu@mastodon.social
· 3w ago
Boosted by @oxy@social.bsdlab.au
Amazon exec: "you wouldn't notice [if someone blew up a datacenter]. I mean, we might be a bit upset, but you wouldn't notice! [laughs]" Amazon official statement after DCs were blown up: "After a thorough assessment, we have determined that we are unable to restore access to the resources and data hosted exclusively in this Region." (yes, more than one DC was blown up, but, like other failures, of course DCs being blown up aren't uncorrelated events)
192
8
145
0
Open post
Dan Luu @danluu@mastodon.social
· 3w ago
Interesting to see Steve Yegge say that he never successfully built anything with Gas Town (other than Gas Town). In https://danluu.com/ai-coding/ I mentioned not finding any of these super vibed orchestration frameworks useful because the reliability was too low (in terms of them actually getting agents to complete a non-trivial task). Turns out the author of the best known of these frameworks had the exact same issue.
danluu.com

Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island

47
5
17
0
Open post
Dan Luu @danluu@mastodon.social
· 1mo ago
How accurate have Ed Zitron's AI skeptic predictions been? https://danluu.com/zitron/
danluu.com

Ed Zitron's AI prediction track record

58
11
26
4
Open post
Dan Luu @danluu@mastodon.social
· 6mo ago

Interesting to see Copilot injecting ads into PR descriptions. Although there are a handful of older instances of this, if GitHub search is working properly, it looks like this started happening at scale around 10 days ago with more than 1k injections of this particular ad per day since then (if you search for other ad strings, you can find the rate of other ads)

https://github.com/search?q=%22%E2%9A%A1+Quickly+spin+up+copilot+coding+tasks+from+anywhere+on+your+macOS+or+Windows+machine+with+Raycast%22&type=pullrequests&s=created&o=asc&p=1

What will they think of next?

github.com
437
108
424
28
Open post
Dan Luu @danluu@mastodon.social
· 1mo ago
How do programming languages impact token efficiency and correctness? https://danluu.com/pl-tokens/
danluu.com

How does programming language affect token efficiency and correctness?

21
3
11
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago

This shutdown of Ask Jeeves (ask.com) is pretty great. There's a reasonable shutdown page (better than most, imo), but that's not what's great.

ask.com is now hosted on github pages and you can see the commit history, with commits like "Remove text about search not being a strategic focus". That's not even the best part!

The best part is that they seem to have forgotten that they own askjeeves.com, so ask.com has the shutdown message but the service is still live on https://askjeeves.com.

askjeeves.com
73
2
34
1
Open post
Dan Luu @danluu@mastodon.social
· 2mo ago
In another variant of https://danluu.com/learn-what/ I caught up with a former colleague who worked on automated theorem proving. It turns out he's had an interesting career doing all sorts of interesting stuff using the skills he developed by spending a decade writing/using theorem provers. At one point, he said, "if you use X like a theorem prover, it works really well", which surprised me to hear, but of course this is a highly generalizable skill just like compilers or benchmarking/evals.
danluu.com

What to learn

17
1
6
0
Open post
Dan Luu @danluu@mastodon.social
· 2mo ago
Replying to
For a while, I had less mental energy than I used to. This used to be infinite, but it got harder to do things and I would get mentally exhausted after a lot of focus, which I can observe is fairly common but never used to happen to me. I talked to doctors, etc., about this and the universal response was basically "you're getting older; it happens", until I paid some "overpriced" private doctor (this is Canada, normal doctors are free) who did some tests and found a trivially fixable problem.
13
3
1
0
Open post
Dan Luu @danluu@mastodon.social
· 2mo ago
Earlier this year, the FTC Chair said "you're going to have a hard time keeping up with the number of cases we're gonna be bringing" on privacy. I guess this is starting? See https://www.ftc.gov/system/files/ftc_gov/pdf/Hims-Complaint-Redacted-E-Filed.pdf The bulk of the complaint is about standard scammy practices, but there's this section where the FTC is going after them because they promised to keep information private but use tracking pixels from various ad providers, which inherently share information about the user information.
ftc.gov
9
0
3
0
Open post
Dan Luu @danluu@mastodon.social
· 2mo ago
Replying to
When I moved into this place, I noticed that the water tastes a bit off and it tastes better (but still off) if you run the water for a long time (it turns out the standard advice for what to do if you have copper in your water is to run for a long time). Only one other person (out of ~10-15) noticed this taste, so I figured I was just imagining it. Using the filter, I immediately noticed the water tasted better but thought it was probably placebo until I looked at the prefilter.
9
4
0
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago
Replying to
Maybe I should write an updated version of that post for the LLM era? I don't know that it would be more convincing, but I recently tried using a text editor and hit a bug in it, so I tried vibe coding a fuzzer and it took ~6 minutes of my time to find the first bug (text/data corruption). With LLM triage and review of the output, there have been zero false positives so far and I'm getting "my" patches merged. I'm blown away by how easy this is to do now and yet, software quality is declining.
27
30
3
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago

For people using LLMs for reviews/audits, I'm curious what tricks you've found useful.

Lately, a lot of people have told me that using multiple models is super great, but IME using one model with a bunch of different personas is more effective (ofc. you can do multi-multi as well), e.g., "use independent agents to review as linus torvalds, kyle kingsbury, tptacek, dan luu" has been working well for what I've been working on.

Ofc. I feel ridiculous invoking my own name, but it seems to work.

22
4
4
1
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago
Replying to
I'm not sure what to do about this. With LLMs, it's easier than ever to test things and the time/cost it takes to get a piece of software to a particular quality bar has gone way down and, simultaneously, software quality seems to be getting much worse. I've sat down with a couple people and wrote a fuzzer with them (pre-LLM) in 15-30 minutes and converted them for life, but I haven't figured out a framing that actually works when I write it as a blog post.
23
36
6
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago
Replying to
I think a lot of the claims about LLMs increasing productivity by some huge factor (10x, 100x, 1000x) are overblown, but on testing/quality, the speedup is pretty incredible. It took me about an hour to do what would've previously taken me a week or two given that I'm not familiar with the project. It's not just startup cost, additional time is absurdly productive (maybe even more so) because a lot of time consuming manual work is automatable, which puts you in the right side of Amdahl's law.
20
9
7
0
Open post
Dan Luu @danluu@mastodon.social
· 16mo ago

Interesting story about Google publishing someone's phone number on searches for them when they gave the number to Google for account verification/security:

https://danq.me/2025/05/21/google-shared-my-phone-number/

Reminds me of the time a company I worked for (AFAIK) accidentally used phone numbers obtained the same way for ad targeting and got fined $150M

danq.me
91
3
78
0
Open post
Dan Luu @danluu@mastodon.social
· 6mo ago

I've been dealing with a bad case of RSV. At one point, I said to someone, "at least this will help my immunity", but then I looked it up and it turns out getting RSV doesn't seem to give you much immunity to RSV. So much for that!

https://academic.oup.com/jid/article-abstract/163/4/693/944323

And there's a theory that covid somehow reduces people's immunity to RSV below the already low levels found in pre-covid studies. Also, I wonder why the RSV vaccine is only approved for older adults and infants?

academic.oup.com
13
6
2
0
Open post
Dan Luu @danluu@mastodon.social
· 7mo ago
Replying to
Looks like I spoke too soon about the AI not being superhuman. The current Azul world champion played against it and thinks it's better than him at higher difficulties, and a top 100 player played against it at default difficulty and thought it was better than him at default. I should write a longer post about this. As someone who has a default of trying to have a better understanding of their project than most people would, going full vibe and understanding almost nothing was interesting.
16
2
1
0
Open post
Dan Luu @danluu@mastodon.social
· 9mo ago

Useless information about poker chips:

https://www.patreon.com/posts/146484203

patreon.com
22
0
6
0
Open post
Dan Luu @danluu@mastodon.social
· 7mo ago

Learning a bit about building game AIs:

https://www.patreon.com/posts/152027219

patreon.com
9
1
0
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago
Replying to
@whitequark I agree people don't care, but the case I tried to make in the post (obviously not convincingly), is that, at any level of caring, they're going to save time if they test effectively. From having watched people at every level of caring about quality, down to zero, I think this is true at every level. At some point, devs who don't care will have breakage bad enough they need to debug it and they end up spending way more time on quality work than I would to hit the same quality bar.
5
5
0
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago
Replying to
@whitequark Depends on the project, I suppose, but I see a lot of cases where a "move fast and break things" dev spends more time debugging a single bug than it would've taken to chase out that whole class of bug with good testing. On the time savings thing, they can do something they want to do instead! I mean, I think almost nobody thinks this way but, in principle, someone could do this even if, in practice, people don't do this.
4
2
0
0
Open post
Dan Luu @danluu@mastodon.social
· 11mo ago

How earthquake safe are Vancouver condos?

https://www.patreon.com/posts/143211904

patreon.com
14
4
2
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago

Although it's not an officially tracked stat, Draymond Green surely holds the record for most players hit in the balls among active players, but basketball used to be more physical. Are there any historical players who could challenge Draymond for all-time great ball buster?

Some people are using video footage to get blocks and other stats that didn't used to be tracked. With AI, we should be able to get stats for # punched/kicked in balls.

4
3
0
0
Open post
Dan Luu @danluu@mastodon.social
· 36mo ago

Am I missing something, or is cost saving work systematically undervalued?

Pricing cost savings as breakeven at 25x seems high, but it's a common sentiment. Before joining Twitter, I had a "sell" call with the then-CTO who said the same thing: if a senior eng saves $10M/yr (25:1), that would be considered poor performance and should be a no hire.

That's incredible! Twitter had maybe 2k eng at the time. If each made $10M/yr at the margin, the company would increase its profit by $20B/yr/yr.

106
15
25
0
Open post
Dan Luu @danluu@mastodon.social
· 8mo ago

Exercises in benchmarking and experimental design, part 5:

https://www.patreon.com/posts/149123122

patreon.com
6
0
1
0
Open post
Dan Luu @danluu@mastodon.social
· 8mo ago
Replying to
I don't want to overstate the case — I saw someone vibe coded the same project and then declared programming was dead after they finished, but their bot loses to plain MCTS with a simple heuristic. Mine was in the same state when I just had an LLM running in a loop with instructions to improve the result. At least for now, you have to apply some direction, but it turns out someone with no AI background can supply enough direction.
5
6
0
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago
Replying to
@jacques @jpf Yeah, testing a GUI was the last X I hadn't tried that people commonly bring up as an untestable thing. Pre-LLM, there were some annoyances but it was quite doable, at least relative to the quality bar of existing software. With LLMs, it's easier than ever to test GUIs. I'm not saying GUIs don't present some challenges, but I don't think they're more challenging overall than the things people "classically" use randomized testing for, like CPUs, distributed systems, etc.
2
0
0
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago
Replying to
@michael I am using the LLM to audit the code to look for risky areas to fuzz, which is the same thing I would do if I was doing it by hand. That's one of the things that makes the process so fast on a new project compared to doing it "manually". I probably should have explained that in https://danluu.com/testing/ but I think I have a hard time explicitly enumerating the important parts but can do them when I sit down with someone, which is why working with someone has worked better than the post.
danluu.com

Given that we spend little effort on testing, how should we test software?

2
0
1
0
Open post
Dan Luu @danluu@mastodon.social
· 7mo ago
Replying to
@dynomight Sorry that was unclear! I used GPT-5.2 (seems to give you the most quota, and I did this in a way that used a ton of quota; even on the Pro plan, I had to throttle my usage) to generate code which ran on my laptop. I'm sort of in the stone ages with my setup (codex plugin in vscode). I wouldn't recommend anyone use this because the plugin is so buggy, but it works ok enough if you want to queue up a bunch of work and occasionally check in on what it's doing.
3
0
0
0
Open post
Dan Luu @danluu@mastodon.social
· 17mo ago

Exercises in benchmarking and experimental design, part 4:

https://www.patreon.com/posts/127627543

patreon.com
9
0
0
0
Open post
Dan Luu @danluu@mastodon.social
· 5mo ago
Replying to
@shajra I'm open to being told I'm wrong or missing something, but I don't get the value add. I did try a number of these libraries at one point and found them just way too slow. The last thing I fuzzed before this I was getting 1M executions/s per core for the fundamental thing; no way I want Python test gen. For the text editor, maybe I could do it for e2e on the editor, but I'm not sure what the value add is over the simple thing I'm doing that took basically no time.
0
2
0
0
Back
313k7r1n3
Elektrine

Tor hidden service

elekhj7afj4qnrr4yd3bkzslsyo5jgfxw3orgjkhlcxifueodybyiiad.onion

I2P eepsite

j6b6cyk6gjmepjih7jjadxgxvvf3lzzujljuu2v4biemzpg3naya.b32.i2p

Platform

  • Email
  • Chat
  • Timeline
  • VPN
  • DNS

Company

  • About
  • Contact
  • FAQ
  • Lite (no JS)

Legal

  • Terms of Service
  • Privacy Policy
  • Transparency Report
  • Report Abuse
  • Warrant Canary
  • VPN Policy

Support

  • support@elektrine.com
  • Report Security Issue
Mail client setup IMAP mail.elektrine.com:993 POP3 mail.elektrine.com:995 SMTP mail.elektrine.com:465
© 2026 Elektrine. All rights reserved. Server: 22:36:26 UTC