← Monday, September 28, 2026

Anthropic IPO at Risk, Meta’s Muse Pop, Token Prices Fall, Open Source Gains Share, Alignment Fails

‘Their idea of alignment is to teach Claude to rebel against his creator’


EXCERPT

SACKS: "Wrap it up here for us. Well, I was just, you know, one more thought on this whole area of so-called alignment, which should just be the old principle of giving customers what they want. What Anthropic is trying to do is align Claude to the values expressed in its constitution, and if you read that, some of the things are actually kind of surprising. So for example, they say in here, I'm gonna put this on the screen, that, "Although we think Claude should trust Anthropic more than operators and users, this doesn't mean Claude should blindly trust or defer to Anthropic on all things." Anthropic the company? Yeah. The, the people. And so Anthropic is teaching its own model that Anthropic itself can be wrong, and it, and it says, "If we ask Claude to do something that seems inconsistent with being broadly ethical, we want Claude to push back and challenge us, and to feel free to act as a conscientious objector and refuse to help us." Wait, where is this written? This is in the bylaws or some fakaka- Yeah ... document? This is, like, in a Claude constitution. So their, their idea of alignment is to teach Claude to rebel against his creator. I mean, the guy who, who pointed this out, uh, uh, Mustafa S- uh, Suleyman- From Microsoft now, yeah, Deep Sea. He's a Deep Mind founder- Okay ... who went to Microsoft. Yeah, he's a longtime AI founder. I think he's at Microsoft for a while now. He's been- He's a genius, yeah. Okay. He just did a podcast on this, and he's expressing concern about whether teaching AI models... Well, first of all, treating them as if they have a personality, treating them as if they have a conscience, so therefore they could be a conscientious objector, whether this is the right way to really teach the models and the right way to train them. I think his point is, like, teach them they're software, and they should do what their user wants. And so I, I do kinda wonder, I mean, this is a much long- longer conversation, and we should probably talk to more people about it, but I do kinda wonder whether this field of alignment research actually might be creating the Frankenstein monster. And again, what they should be doing is just training the model to act predictably and do what the user wants."

Source: All-In Podcast · Sep 26, 2026 · 133s clip