paper · First submission: June 12, 2017
Attention Is All You Need
The 2017 Transformer proposal changes sequence modeling, while leaving the purpose and safety of trained systems unresolved.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser and Illia Polosukhin introduce the Transformer. The v1 footnote describes equal contributions and says the listing order is random. Attribution belongs to that collaboration.[1]
What changed
The introduction identifies sequential computation as a constraint on recurrent sequence models. The proposed encoder–decoder architecture uses attention without recurrence or convolution; its experiments concern translation and parsing. It builds on earlier attention work rather than introducing attention itself.[1]
Why read it here?
The origins guide uses it as a capability-development branch. Improving an architecture’s performance does not select whose interests it should serve. Continue to Training language models to follow instructions with human feedback to examine a different question: how training feedback shapes an assistant’s behavior.
Historical context
The chronology uses first arXiv submission, June 12, 2017. This entry reviews selected v1 passages; the current arXiv record is v7, revised August 2, 2023. No version equivalence, equation audit or replication is claimed.[1]
Sources
- Attention Is All You Need · Source record ref-transformer-2017 · Back to claim ↑1 ↑2 ↑3
Pages that link here
Last updated 2026-10-11