Perfectly Aligning AI’s Values With Humanity’s Is Impossible – Maybe the best we can do is make “neurodiverse” systems that challenge each other
...
IEEE Spectrum: You and your colleagues have now shown that misalignment of AI systems is inevitable, because any AI system complex enough to display general intelligence will produce unpredictable behavior. Your proof rests on two famous sets of premises—Gödel’s incompleteness theorems, which found that every mathematical system will have statements that can never be proven, and Turing’s undecidability result for the halting problem, which found that some problems are inherently unsolvable.
Zenil: The conventional wisdom assumes misalignment is a bug that can eventually be removed with the right optimization strategy. Our results show that the problem of alignment is not simply a lack of better data, more compute, or better engineering, but a limit built into both formal systems and universal computation. What I am arguing is that for sufficiently general AI systems, some degree of misalignment is structural, so the task shifts from elimination to management.
IEEE Spectrum: Can you describe your strategy of managed misalignment?
Zenil: Once perfect alignment looked unattainable in principle, the next move was obvious—stop trying to perfect one agent and start designing the ecology around it. This is what it would take to achieve any degree of controllability, and controllability has to come from outside, given the intrinsic impossibility of controlling from the inside. You see similar strategies in biology and medicine, where robust results often come from interacting systems rather than a single master controller.
The simplest way to put it is this: Do not trust one supposedly perfect AI to govern everything. Instead, build a structured ecosystem of different agents with different “values” that monitor, challenge, and constrain one another, much like courts, auditors, and competing institutions do in human society. None of them is perfect on its own, but their managed interaction can make the whole arrangement safer than any single dominant model.
The main thing not to misunderstand is that managed misalignment does not mean giving up on safety or letting AI behave however it likes. It means replacing the fantasy of absolute control with a more realistic form of distributed control. In that sense, it is not less serious about safety, but more serious about what safety actually requires.
IEEE Spectrum: How did you test your strategy?
...
Zenil: This work is not anti-AI. It is anti-naivety about control.
See the full story here: https://spectrum.ieee.org/amp/ai-alignment-2676752963
Pages
- About Philip Lelyveld
- Mark and Addie Lelyveld Biographies
- Presentations and articles
- Trustworthy AI – A Market-Driven approach
- Tufts Alumni Bio