"It's different with AI."
Specifically, transformer and diffuser NN architectures + Moore's law. Marvin Minsky's "AI" never got off the ground. LSTM architectures aren't decoder-only, but apparently if you scale them up to the param counts of currently successful GPT-architecture models like llama, they start talking like llama too.

