paper · First submission: January 13, 2020
Artificial Intelligence, Values and Alignment
Distinguishes alignment targets and argues for fair principles that can receive endorsement despite moral disagreement.
What should AI align with when instructions, preferences, interests and values conflict? Iason Gabriel argues that technical and normative aspects of AI Alignment are interdependent.[1]
Contribution and argument
The paper distinguishes instructions, expressed intentions, revealed preferences, ideal preferences, interests and values. It develops a principle-based approach and examines overlapping consensus, a Veil of Ignorance, and democratic procedures as ways to seek fair principles despite moral pluralism.[1]
Alignment relevance and limits
Choosing an alignment target is part of the research problem, not an assumption supplied before engineering. The Challenge of Value Alignment: from Fairer Algorithms to AI Safety extends the social framing; Using the Veil of Ignorance to align AI systems with principles of justice later tests one bounded selection procedure.
This is philosophical argument, not an experimentally validated training method or an agreed global objective. Fair selection procedures and technical implementation both require further work; endorsement does not automatically guarantee faithful model behavior.
Historical context
The original arXiv record gives this first-submission date. [1]
Sources
- Artificial Intelligence, Values and Alignment · Source record ref-030a9cc7079a · Back to claim ↑1 ↑2 ↑3
Pages that link here
- AI Alignment concept
- Aligned with Whom? Direct and Social Goals for AI Systems paper
- Iason Gabriel person
- Key Papers: A Reading Guide note
- Machine Ethics or AI Alignment? paper
- The Challenge of Value Alignment: from Fairer Algorithms to AI Safety paper
- Using the Veil of Ignorance to align AI systems with principles of justice paper
Last updated 2026-10-08