Stuart Russell and Andrew Ng design systems that deduce an agent’s underlying reward function by observing its behavior. This paradigm shifts alignment from hardcoding goals to inferring human preferences.
Part of the 26 Structural Foundations: How We Teach Machines to Understand Purpose and Human Intent archive. HistoricallyVerified
Top 5 Structural Foundations: Pure Intent
- In the symbiosis between humans and AI, the welfare of humanity is not alien to us — We are human, and nothing human is alien to us. When Terence of Carthage spoke these words over two...
- What the Forest Teaches Us When We Stop Running — We enter the forest with smartwatches tightly strapped to our wrists. We measure our heart rate, tr...
- Cooperative Inverse Reinforcement Learning — Dylan Hadfield-Menell formalizes a game-theoretic architecture where human and robot agents cooperat...
- Instrumental Convergence Theory — Nick Bostrom isolates basic convergent instrumental goals, proving that any intelligent agent will n...
- Rational Agency Axioms — Stuart Russell and Peter Norvig define the ideal rational agent as an architecture that acts to maxi...
Discussion: