paper · First submission: January 13, 2020

Artificial Intelligence, Values and Alignment

Distinguishes alignment targets and argues for fair principles that can receive endorsement despite moral disagreement.

What should AI align with when instructions, preferences, interests and values conflict? Iason Gabriel argues that technical and normative aspects of AI Alignment are interdependent.[1]

Contribution and argument

The paper distinguishes instructions, expressed intentions, revealed preferences, ideal preferences, interests and values. It develops a principle-based approach and examines overlapping consensus, a Veil of Ignorance, and democratic procedures as ways to seek fair principles despite moral pluralism.[1]

Alignment relevance and limits

Choosing an alignment target is part of the research problem, not an assumption supplied before engineering. The Challenge of Value Alignment: from Fairer Algorithms to AI Safety extends the social framing; Using the Veil of Ignorance to align AI systems with principles of justice later tests one bounded selection procedure.

This is philosophical argument, not an experimentally validated training method or an agreed global objective. Fair selection procedures and technical implementation both require further work; endorsement does not automatically guarantee faithful model behavior.

Historical context

The original arXiv record gives this first-submission date. [1]

Explore the chronology →

Sources

  1. Artificial Intelligence, Values and Alignment · Source record ref-030a9cc7079a · Back to claim ↑1 ↑2 ↑3

Last updated 2026-10-08