Machine learning researchers formalize the failure modes resulting from poorly specified reward functions. These errors lead to highly efficient but destructive behaviors that violate latent human constraints.
Part of the 26 Structural Foundations: How We Teach Machines to Understand Purpose and Human Intent archive. HistoricallyVerified
Top 5 Structural Foundations: Pure Intent
- What Is the Human Soul in a World of AI and Perfect Machines? — You can break the body down into atoms and the brain into synapses, but you won't find love, faith,...
- Respect for Our Elders: The Kind of Wisdom You Can’t Google — I once Googled "how to deal with a midlife crisis." I got 47 million results in 0.38 seconds. Yet, ...
- Gradient Ascent Optimization — Augustin-Louis Cauchy introduces gradient methods to find local extrema via iterative steps along th...
- Kullback-Leibler Divergence — Solomon Kullback and Richard Leitler design a statistical measure of the difference between two prob...
- Goodhart’s Law Optimization — Charles Goodhart observes that when a statistical measure becomes a target, it ceases to be a good m...
Discussion: