
“Alignment Faking in Large Language Models” by Ryan Greenblatt
What happens when you tell Claude it is being trained to do something it doesn't want to do? We (Anthropic and Redwood Research) have a new paper demonstrating that, in our experiments, Claude will of...
19 Dec 202419min

“There is no sorting hat in EA” by ElliotTep
Summary My sense is some EAs act like/hope they will be assigned the perfect impactful career by some combination of 80,000 Hours recommendations (and similar) and ‘perceived consensus views in EA’. ...
19 Dec 202415min



















