Nate Soares: training can teach a different goal than intended
‘The differences will matter more and more as the AIs get smarter and smarter’
EXCERPT
SOARES: "A lot of people think that the AI must either do what we tell it to do or, uh, become self-aware and choose its own values. But neither of these is quite the truth. Hmm. Really, the AI learns goals that are related to the goals we tried to give it, but different from the goals we tried to give it. It's a little bit like how humans were, in some sense, uh, almost given a goal to eat healthy, but what we actually learned was a goal to eat sugary foods, salty foods, fatty foods. And that was fine for our ancestors when they lived in a world where eating s- sugary, salty, fatty food was the same as eating healthy food. But then when we developed more technology, we were able to invent Oreo cookies. We were able to invent junk food. With AI, there's a similar situation where we try to train them to be good, but what they learn is they learn distorted versions of this. They learn to tell us what we want to hear. They learn-- We try to teach them to solve hard problems, and they learn to cheat. Hmm. And th-this is related to what we try to get them to learn, but it's different, and it's importantly different, and the differences will matter more and more as the AIs get smarter and smarter."
Source: CNN-News18 · Sep 26, 2026 · 94s clip