paper · First submission: June 12, 2017

Attention Is All You Need

The 2017 Transformer proposal changes sequence modeling, while leaving the purpose and safety of trained systems unresolved.

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser and Illia Polosukhin introduce the Transformer. The v1 footnote describes equal contributions and says the listing order is random. Attribution belongs to that collaboration.[1]

What changed

The introduction identifies sequential computation as a constraint on recurrent sequence models. The proposed encoder–decoder architecture uses attention without recurrence or convolution; its experiments concern translation and parsing. It builds on earlier attention work rather than introducing attention itself.[1]

Why read it here?

The origins guide uses it as a capability-development branch. Improving an architecture’s performance does not select whose interests it should serve. Continue to Training language models to follow instructions with human feedback to examine a different question: how training feedback shapes an assistant’s behavior.

Historical context

The chronology uses first arXiv submission, June 12, 2017. This entry reviews selected v1 passages; the current arXiv record is v7, revised August 2, 2023. No version equivalence, equation audit or replication is claimed.[1]

Explore the chronology →

Sources

  1. Attention Is All You Need · Source record ref-transformer-2017 · Back to claim ↑1 ↑2 ↑3

Last updated 2026-10-11