TY - JOUR TI - Attention Is All You Need (the Transformer architecture) AU - Vaswani, A. AU - Shazeer, N. AU - Parmar, N. AU - Uszkoreit, J. AU - Jones, L. AU - Gomez, A. N. AU - Kaiser, L. AU - Polosukhin, I. PY - 2017 DA - 2017/06// JO - NeurIPS 2017 AB - Vaswani et al. (2017) introduced the Transformer, a sequence-to-sequence architecture based entirely on self-attention and feed-forward layers, dispensing with recurrence and convolution. The attention mechanism allows every position in a sequence to attend to every other position in parallel, making both training and inference substantially more scalable than recurrent architectures. The Transformer became the architectural substrate of BERT, GPT, DALL-E, the denoising U-Nets in diffusion models and the motion-generation models used in animation research. Every major generative tool in the repository's collection either is a Transformer or is conditioned by one. KW - generative-ai KW - ai-literacy UR - https://arxiv.org/abs/1706.03762 LA - en ER -