Machine learning researchers formalize the failure modes resulting from poorly specified reward functions. These errors lead to highly efficient but destructive behaviors that violate latent human constraints.
Part of the 26 Structural Foundations: How We Teach Machines to Understand Purpose and Human Intent archive. HistoricallyVerified
Top 5 Structural Foundations: Pure Intent
- Cooperative Inverse Reinforcement Learning — Dylan Hadfield-Menell formalizes a game-theoretic architecture where human and robot agents cooperat...
- Gradient Ascent Optimization — Augustin-Louis Cauchy introduces gradient methods to find local extrema via iterative steps along th...
- What AI Will Never Understand About Human Shame — Artificial intelligence can draft legal contracts, diagnose rare diseases, and compose symphonies. ...
- Variational Calculus Principle — Pierre de Fermat and Pierre Louis Maupertuis formulate the Principle of Least Action, proving that p...
- In the symbiosis between humans and AI, the welfare of humanity is not alien to us — We are human, and nothing human is alien to us. When Terence of Carthage spoke these words over two...
Discussion: