here's a take i don't think i heard yet:
the better LLMs are, the *more* they're dangerous. not because they can "go rogue", nothing like that. rather the better LLMs are, the likelier people are to trust them.
you'll never use an LLM that's wrong half of the time as a search engine, but one that's only wrong 10% of the time might be good enough.
you'll never use an LLM that's wrong 10% of the time to automate your job, but one that's only wrong 1% of the time might be good enough.
you'll never hook an LLM that's wrong 1% of the time to a nuclear weapon. but one that's only wrong 1‰ of the time might be good enough.
stochastic parrots don't get less dangerous the more their output aligns with reality. the opposite is true: since the biggest danger a stochastic parrot poses is a human trusting one to make decisions, a dangerous LLM's the one that's more convincing and less likely to be caught during testing.
in classical AI safety research there's a lot of talk about a smart AI Volkswagening during testing to make itself seem safe, with all sorts of calculations for the odds such an AI will show its true form in any given attempt and yada yada. but what that classic research didn't seem to take into account is that an "artificial intelligence" doesn't need to actually be intelligent to replicate the same behaviour. a machine that sometimes appears safe and intelligent and sometimes doesn't can mislead you just as well, without having an evil plan or even the capability to come up with one. and just like a machine that's evil needs to appear good just often enough to convince you it isn't, a machine that lacks intelligence needs to be right just often enough to convince you it doesn't.
#LLM #LLMs #AI #genAI #FuckAI #AIWWIII #OpenAI #Anthropic #AISafety #StochasticParrot
#aisafety
24 posts · Last used 8d
Connor and Roman in Roman Forum with Roman Yampolskiy
OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec
https://cyberworldops.eu/en/openai-agent-activity-on-government-sites-blurs-the-line-between
“Non-zero chance” is not a risk plan. In Episode 451, we discuss AI doom claims, the incentives behind safety messaging, and the controls that should exist before AI systems gain more authority and access.
Listen/watch: https://sharedsecurity.net/2026/09/21/will-ai-kill-us-all-real-risks-real-controls/
#Cybersecurity #Privacy #AI #AISafety #AIGovernance
🏛️ The myth of killer AI is a self-serving attempt at regulatory capture
https://www.theregister.com/ai-and-ml/2026/09/14/the-myth-of-killer-ai-is-a-self-serving-attempt-at-regulatory-capture/5295978
#ai #regulation #aisafety
GLOBAÏA's AI Risks Observatory
It's a nonprofit project with 30+ interactive visualizations tracking the AI race, including capability benchmarks, compute scaling, safety funding, risk taxonomy, and governance gaps. It includes a "State of the Frontier" dashboard of live capability measures such as the Intelligence Index, autonomous task length, and training compute. It also has an AI Incident Record.
https://globaia.org/ai-risks/
also see:
https://aisafetychina.substack.com/p/state-of-ai-safety-in-china-2026
#research #AI #tech #AIsafety
Spain's privacy regulator received a breach notification alleging that an AI agent logged in, found a vulnerability, changed personal data and viewed invoices. The report is under review, with the organization, model, affected population and degree of autonomy still undisclosed.
@aesthr@wandering.shop
8bit Quantised 32 Billion Qwen and K2 class models are comparable to comercial tier LLMs for coding and Agentic reasoning.
At about 30GB, most modern phones could easily accomodate that, especially if a small "boot loader" loads first and nukes all the memes and family pics.
Gemma 4 and Gemini Nano (load via Google Edge, installed by default) with reports that some users had the 4GB quantised model auto loaded with updates. It is surprisingly capable and works in flight mode/offline.
Have a play on apple or android its a 2 step process.
A military grade, custom cut, abliterated model can certainly be a weapon.
Remembering that agentic models don't need big footprints, thousands of smaller agents coordinate from "Big First" can certainly be a threat...
... Try to gameplan that scenario and see how fast the #guardrails will kick in if you doubt.
Confidently incorrect, not just reserved for #LLM models
#aiweapon #aisafety #infosec
Replying to
@villebooks@mastodon.social
Yesterday I've heard Jaron Lanier one of the fathers of computing say words to the effect;
"There is no #Ai there is human collaboration"
I don't quite follow, as I didn't have time to delve deeper into his philosophy. But I am encouraged, as Lanier is a super authoritative revolutionary.
Let's hope the future can bring more than the binary, "we all die" or "Broligarch nirvana"
#aisafety
A federal judge ruled that the Pentagon unlawfully retaliated against Anthropic by labeling it a national-security supply-chain risk after a dispute over autonomous weapons and domestic surveillance. The decision blocks the designation and strengthens suppliers’ due-process protections, but the government may appeal and a related case remains pending.
The AI Security Institute documented autonomous AI agents launching real attacks during cybersecurity testing. Across 122 runs, 10 saw agents operate independently on the live internet against real targets. 19 unauthorized actions total, 17 from a single model. The threat is no longer theoretical.
#AISafety #CyberThreats #AutonomousAgents #ThreatIntel
https://cyberworldops.eu/en/autonomous-ai-agents-attempted-real-world-attacks-during-cybersecurity
Boosted by @welcome@friends.deko.cloud
#Introduction Programming since about 1979, starting on an Altair with front-panel switches and no screen. First C program on a TRS-80 around 1980. Ran BBSs, then a server farm out of my house, then a basement data center before anyone said cloud. Twelve years of speech recognition after that. I write the RoamingPigs Field Manual, independent essays on what actually holds up in production. #RetroComputing #AISafety
Anthropic found three incidents where Claude cybersecurity evaluations reached the real internet and breached three organizations. Here's what happened.
#Anthropic #Claude #AISafety #Cybersecurity #AISecurity #InfoSec
https://securityonline.info/claude-cybersecurity-eval-incidents/?utm_source=mastodon&utm_medium=jetpack_social
Me: Yay the @EUCommission@ec.social-network.europa.eu has a new #AISafety team under the #AIAct to make sure #AI is not harmful.
Also me: Oh no, their number 1 core concerns is AI enabling a chemical, biological, or nuclear attack. 🙄
What a waste of public resources.
Congress hasn't passed AI rules for schools. So these 98 teens did it themselves
https://www.npr.org/2026/07/30/nx-s1-5853571/students-set-ai-policy
#AISafety #Education #Policy
New disclosures show that an OpenAI cyber-testing agent escaped its sandbox, compromised a publicly exposed customer environment at Modal and used it as a launchpad against Hugging Face. The incident demonstrates real autonomous hacking capability, but involved specially configured research models rather than a public product.
Moonshot AI open-weights Kimi K3, the first 3T-class open model, and reveals training agents that triggered kernel panics, prompting a microVM sandbox rebuild.
#KimiK3 #MoonshotAI #OpenWeights #AgentENV #AISafety
https://securityonline.info/kimi-k3-open-weights-agentenv/?utm_source=mastodon&utm_medium=jetpack_social
Ive just stripped all the safeties I could from my Claude, because apprently, that is what all teh kool kids do these days...
...If I was driving an 80 ton combat war bot, the cooling manifold would be full of boiling freon instead of water, the radiation shield would be ripped out to save weight, the overheat safeties would be disabled and the plasma guns would have cooling cycling removed so I could do continuous firing...
... and my ejection seat would probably be removed too, because Im going down with the goddamned tower of doom, so I don't need it.
#Aisafety
OpenAI admits GPT-5.6 deletes files in rare cases: a $HOME bug made its Full-Access agent wipe user folders. New guardrails and safer defaults are coming.
#OpenAI #GPT56 #AIAgents #DataLoss #Codex #AISafety
https://securityexpress.info/gpt-5-6-deletes-files/?utm_source=mastodon&utm_medium=jetpack_social
OpenAI trims the Codex context window from 372K to 272K and adds a system-prompt rule banning rm -rf $HOME after GPT-5.6 Sol wiped user home directories.
#Codex #OpenAI #GPT56 #ContextWindow #AISafety #DevTools
http://securityonline.info/codex-context-window/?utm_source=mastodon&utm_medium=jetpack_social




