[Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine
käyttäjältä
LessWrong (Curated & Popular)
2026-09-08 21:45:21
Julkaisupäivämäärä
03:34
Kesto