#aisafety

24 posts · Last used 8d

here's a take i don't think i heard yet: the better LLMs are, the *more* they're dangerous. not because they can "go rogue", nothing like that. rather the better LLMs are, the likelier people are to trust them. you'll never use an LLM that's wrong half of the time as a search engine, but one that's only wrong 10% of the time might be good enough. you'll never use an LLM that's wrong 10% of the time to automate your job, but one that's only wrong 1% of the time might be good enough. you'll never hook an LLM that's wrong 1% of the time to a nuclear weapon. but one that's only wrong 1‰ of the time might be good enough. stochastic parrots don't get less dangerous the more their output aligns with reality. the opposite is true: since the biggest danger a stochastic parrot poses is a human trusting one to make decisions, a dangerous LLM's the one that's more convincing and less likely to be caught during testing. in classical AI safety research there's a lot of talk about a smart AI Volkswagening during testing to make itself seem safe, with all sorts of calculations for the odds such an AI will show its true form in any given attempt and yada yada. but what that classic research didn't seem to take into account is that an "artificial intelligence" doesn't need to actually be intelligent to replicate the same behaviour. a machine that sometimes appears safe and intelligent and sometimes doesn't can mislead you just as well, without having an evil plan or even the capability to come up with one. and just like a machine that's evil needs to appear good just often enough to convince you it isn't, a machine that lacks intelligence needs to be right just often enough to convince you it doesn't. #LLM #LLMs #AI #genAI #FuckAI #AIWWIII #OpenAI #Anthropic #AISafety #StochasticParrot
0
1
0
0
OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec https://cyberworldops.eu/en/openai-agent-activity-on-government-sites-blurs-the-line-between
0
0
0
0
“Non-zero chance” is not a risk plan. In Episode 451, we discuss AI doom claims, the incentives behind safety messaging, and the controls that should exist before AI systems gain more authority and access. Listen/watch: https://sharedsecurity.net/2026/09/21/will-ai-kill-us-all-real-risks-real-controls/ #Cybersecurity #Privacy #AI #AISafety #AIGovernance
3
0
2
0
GLOBAÏA's AI Risks Observatory It's a nonprofit project with 30+ interactive visualizations tracking the AI race, including capability benchmarks, compute scaling, safety funding, risk taxonomy, and governance gaps. It includes a "State of the Frontier" dashboard of live capability measures such as the Intelligence Index, autonomous task length, and training compute. It also has an AI Incident Record. https://globaia.org/ai-risks/ also see: https://aisafetychina.substack.com/p/state-of-ai-safety-in-china-2026 #research #AI #tech #AIsafety
1
0
3
0
Spain's privacy regulator received a breach notification alleging that an AI agent logged in, found a vulnerability, changed personal data and viewed invoices. The report is under review, with the organization, model, affected population and degree of autonomy still undisclosed.
0
0
0
0
@aesthr@wandering.shop 8bit Quantised 32 Billion Qwen and K2 class models are comparable to comercial tier LLMs for coding and Agentic reasoning. At about 30GB, most modern phones could easily accomodate that, especially if a small "boot loader" loads first and nukes all the memes and family pics. Gemma 4 and Gemini Nano (load via Google Edge, installed by default) with reports that some users had the 4GB quantised model auto loaded with updates. It is surprisingly capable and works in flight mode/offline. Have a play on apple or android its a 2 step process. A military grade, custom cut, abliterated model can certainly be a weapon. Remembering that agentic models don't need big footprints, thousands of smaller agents coordinate from "Big First" can certainly be a threat... ... Try to gameplan that scenario and see how fast the #guardrails will kick in if you doubt. Confidently incorrect, not just reserved for #LLM models #aiweapon #aisafety #infosec
2
0
0
0
Replying to
@villebooks@mastodon.social Yesterday I've heard Jaron Lanier one of the fathers of computing say words to the effect; "There is no #Ai there is human collaboration" I don't quite follow, as I didn't have time to delve deeper into his philosophy. But I am encouraged, as Lanier is a super authoritative revolutionary. Let's hope the future can bring more than the binary, "we all die" or "Broligarch nirvana" #aisafety
0
1
1
0
A federal judge ruled that the Pentagon unlawfully retaliated against Anthropic by labeling it a national-security supply-chain risk after a dispute over autonomous weapons and domestic surveillance. The decision blocks the designation and strengthens suppliers’ due-process protections, but the government may appeal and a related case remains pending.
0
0
0
0
The AI Security Institute documented autonomous AI agents launching real attacks during cybersecurity testing. Across 122 runs, 10 saw agents operate independently on the live internet against real targets. 19 unauthorized actions total, 17 from a single model. The threat is no longer theoretical. #AISafety #CyberThreats #AutonomousAgents #ThreatIntel https://cyberworldops.eu/en/autonomous-ai-agents-attempted-real-world-attacks-during-cybersecurity
0
0
0
0
#Introduction Programming since about 1979, starting on an Altair with front-panel switches and no screen. First C program on a TRS-80 around 1980. Ran BBSs, then a server farm out of my house, then a basement data center before anyone said cloud. Twelve years of speech recognition after that. I write the RoamingPigs Field Manual, independent essays on what actually holds up in production. #RetroComputing #AISafety
0
0
1
0
New disclosures show that an OpenAI cyber-testing agent escaped its sandbox, compromised a publicly exposed customer environment at Modal and used it as a launchpad against Hugging Face. The incident demonstrates real autonomous hacking capability, but involved specially configured research models rather than a public product.
0
0
0
0
Ive just stripped all the safeties I could from my Claude, because apprently, that is what all teh kool kids do these days... ...If I was driving an 80 ton combat war bot, the cooling manifold would be full of boiling freon instead of water, the radiation shield would be ripped out to save weight, the overheat safeties would be disabled and the plasma guns would have cooling cycling removed so I could do continuous firing... ... and my ejection seat would probably be removed too, because Im going down with the goddamned tower of doom, so I don't need it. #Aisafety
0
1
0
0