The (true) story of development and inspiration behind the "attention" operator, the one in "Attention is All you Need" that introduced the Transformer. From personal email correspondence with the author @DBahdanau ~2 years ago, published here and now (with permission) following some fake news about how it was developed that circulated here over the last few days.
Attention is a brilliant (data-dependent) weighted average operation. It is a form of global pooling, a reduction, communication. It is a way to aggregate relevant information from multiple nodes (tokens, image patches, or etc.). It is expressive, powerful, has plenty of parallelism, and is efficiently optimizable. Even the Multilayer Perceptron (MLP) can actually be almost re-written as Attention over data-indepedent weights (1st layer weights are the queries, 2nd layer weights are the values, the keys are just input, and softmax becomes elementwise, deleting the normalization). TLDR Attention is awesome and a *major* unlock in neural network architecture design.
It's always been a little surprising to me that the paper "Attention is All You Need" gets ~100X more err ... attention... than the paper that actually introduced Attention ~3 years earlier, by Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio: "Neural Machine Translation by Jointly Learning to Align and Translate". As the name suggests, the core contribution of the Attention is All You Need paper that introduced the Transformer neural net is deleting everything *except* Attention, and basically just stacking it in a ResNet with MLPs (which can also be seen as ~attention per the above). But I do think the Transformer paper stands on its own because it adds many additional amazing ideas bundled up all together at once - positional encodings, scaled attention, multi-headed attention, the isotropic simple design, etc. And the Transformer has imo stuck around basically in its 2017 form to this day ~7 years later, with relatively few and minor modifications, maybe with the exception better positional encoding schemes (RoPE and friends).
Anyway, pasting the full email below, which also hints at why this operation is called "attention" in the first place - it comes from attending to words of a source sentence while emitting the words of the translation in a sequential manner, and was introduced as a term late in the process by Yoshua Bengio in place of RNNSearch (thank god? :D). It's also interesting that the design was inspired by a human cognitive process/strategy, of attending back and forth over some data sequentially. Lastly the story is quite interesting from the perspective of nature of progress, with similar ideas and formulations "in the air", with a particular mentions to the work of Alex Graves (NMT) and Jason Weston (Memory Networks) around that time.
Thank you for the story @DBahdanau !
62K Followers 262 FollowingYes, I *am* that ESR. Well, it's the question people usually ask.
Programmer, wandering philosopher, accidental anthropologist, troublemaker for liberty.
2K Followers 3K FollowingA flexible package manager designed to support multiple versions, configurations, platforms, and compilers. Join us at https://t.co/aT6zWYxXKz!
4K Followers 501 Following#conda has moved to
🐘 Mastodon: @[email protected]
🔗 LinkedIn: https://t.co/O1O5kWxg1X
Join us!
Conda is a Package & Environment manager
69K Followers 332 FollowingWhere computation meets knowledge. Creators of Mathematica, @Wolfram_Alpha, Wolfram Language. Founded in 1987 by @stephen_wolfram.
71K Followers 120 FollowingFounded by @MichaelLarabel in 2004, Phoronix is the largest #opensource news, #Linux hardware reviews & Linux PC/server/HPC performance benchmark site.
518K Followers 21 FollowingOfficial account for Bootstrap, a toolkit providing simple and flexible HTML, CSS, and JS for popular UI components and interactions. Tweets by @mdo.
55 Followers 10 Followingmidipix is a development environment that lets you create programs for Windows using the standard C and POSIX APIs. No compromises made, no shortcuts taken.
40K Followers 178 FollowingThe nonprofit dedicated to stewarding the Rust programming lang & its community 🦀
bsky: https://t.co/pURKYFM3az Mastodon: rustfoundation
29K Followers 318 FollowingJulia is a high-level, dynamic programming language built for technical computing. Join the conversation at #JuliaLang @JuliaConOrg @JuliaInclusive
16K Followers 281 FollowingJava usage at Microsoft spans from Azure to Minecraft, across LinkedIn to Visual Studio Code, and beyond! We use more Java than one can imagine.
152K Followers 2 FollowingA programming language empowering everyone to build reliable and efficient software.
** This account is no longer active. Follow us on other platforms! **
616K Followers 1K FollowingWe make very small computers which you can buy from just $4. We are also the literal coolest. Be excellent to each other. Tech support: https://t.co/ZEBSfmuErK
5K Followers 649 FollowingThe home of #community. Where #developers, #engineers & #scientists come to talk. For support, please get in touch on support at gitter.im
1.6M Followers 267 FollowingThe engine room of @Google. Building AI safely and responsibly to solve the world’s most complex problems. Join us: https://t.co/jUHQA27iBL
41K Followers 0 FollowingThe place to collaborate on an open-source implementation of the Java Platform, Standard Edition, and related projects · 🐘 https://t.co/UtKelUz9Xh
10K Followers 408 FollowingThe Adoptium Working Group promotes and supports high-quality runtimes and associated technology for use across the @Java ecosystem. Part of @EclipseFdn