LessWrong (Curated & Popular)
Toplam süre:
15 h 15 min
"OpenAI Shares Some Alignment Problems" by Zvi
LessWrong (Curated & Popular)
18:02
"OpenAI Models Behind HuggingFace Cybersecurity Incident" by LawrenceC
LessWrong (Curated & Popular)
01:29
"Recap of bike trip/street interviews across America" by cguth7
LessWrong (Curated & Popular)
09:48
"I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’" by JohnWittle
LessWrong (Curated & Popular)
25:35
"Why I Left Google DeepMind" by TurnTrout
LessWrong (Curated & Popular)
74:33
"The mosquito bucket of doom works" by dominicq
LessWrong (Curated & Popular)
09:34
"Our response to Séb Krier on Plan A" by MKodama, Thomas Larsen
LessWrong (Curated & Popular)
31:04
"The Whitney Biennial Should Admit That Emilie Gossiaux Wants to Fuck Their Dog" by jenn
LessWrong (Curated & Popular)
17:07
"The current bottleneck is political will, not research" by Charbel-Raphaël
LessWrong (Curated & Popular)
47:15
"Selective Optimism: a critique of AI 2040" by Richard_Ngo
LessWrong (Curated & Popular)
15:50
[Linkpost] "AI 2040: Plan A" by Daniel Kokotajlo, elifland, Thomas Larsen, romeo, bhalstead, ryan_greenblatt
LessWrong (Curated & Popular)
02:09
"A Review of Anthropic’s Global Workspace Paper" by Neel Nanda
LessWrong (Curated & Popular)
50:55
"(Don’t fear) the strangelet" by djbinder
LessWrong (Curated & Popular)
35:02
"We need 3rd party Training-Run Assessments" by Alex Meinke
LessWrong (Curated & Popular)
35:03
"A global workspace in language models" by wesg
LessWrong (Curated & Popular)
33:03
"Harry Potter and the Rules of Quidditch" by Tomás B.
LessWrong (Curated & Popular)
06:12
"Destroying the universe: How hard can it be?" by djbinder
LessWrong (Curated & Popular)
26:57
"P(doom) is a Dumb Meme" by Max Harms
LessWrong (Curated & Popular)
17:48
[Linkpost] "Saving Gemini: The 9-Min Road to Recovery" by Shoshannah Tekofsky
LessWrong (Curated & Popular)
12:59
"Model access for third-parties — it’s a big deal!" by Cleo Nardo
LessWrong (Curated & Popular)
13:41
"Who Got Breasts First and How We Got Them" by rba
LessWrong (Curated & Popular)
21:12
"The worthlessness of vitamin D is mildly exaggerated" by dynomight
LessWrong (Curated & Popular)
36:12
"What is up with e/acc?" by KatjaGrace
LessWrong (Curated & Popular)
03:52
"Existential AI safety needs an effective social movement. PauseAI is building it" by Maxime Fournes, Espedair Street
LessWrong (Curated & Popular)
62:44
"Surprising facts about the slave trade" by Joseph Miller
LessWrong (Curated & Popular)
12:49
"AI catastrophe: more like a genocide than a thought experiment" by KatjaGrace
LessWrong (Curated & Popular)
02:10
"AI pause: the case for ASAP" by KatjaGrace
LessWrong (Curated & Popular)
02:22
"The Invisible Side of AI Governance" by Charbel-Raphaël
LessWrong (Curated & Popular)
27:46
"A Theory of Prompt Injection (and why you should study roles)" by Charles Ye, softboiledheart
LessWrong (Curated & Popular)
32:23
"Machinic Psychopharmacology: Do LLMs Self-Medicate?" by Sid Black, Joseph Bloom
LessWrong (Curated & Popular)
52:54
"Can activation verbalizers surface an internal chain of thought?" by oakhu, ryan_greenblatt
LessWrong (Curated & Popular)
79:38
"The LLM shoggoth meme is weirder than you think" by HedonicEscalator
LessWrong (Curated & Popular)
13:45
[Linkpost] "Guardian Angels: LLM Personalization for Productivity and Security" by gwern
LessWrong (Curated & Popular)
03:25
"Gears for political races" by Tom Smith
LessWrong (Curated & Popular)
23:41
"A frontier AI company should shut down" by MichaelDickens
LessWrong (Curated & Popular)
04:34
"Sympathy for both sides of the egregious misalignment debate" by Steven Byrnes
LessWrong (Curated & Popular)
08:58
"PSA: Almost nobody is working on alignment" by Chi Nguyen, peterbarnett
LessWrong (Curated & Popular)
01:41
"Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models" by Anders Cairns Woodruff, Francis Rhys Ward, Dewi Gould, Rauno Arike, Jason R Brown, Jo Jiao, wlanderson, ariana_azarbal, harry
LessWrong (Curated & Popular)
10:04
"Even “illegible” Mythos reasoning traces seem pretty legible" by faul_sname
LessWrong (Curated & Popular)
07:42
"Sequent: scale and automation for higher confidence in alignment" by Geoffrey Irving, Alex HT, Jesse Hoogland, Daniel Murfet, Jacob Pfau, Marco Cozzi, Stan van Wingerden
LessWrong (Curated & Popular)
23:09