I re-read "Attention Is All You Need" last week. For maybe the tenth time.
Every time I read it, I notice something new. This time it was how elegantly the positional encoding works — a solution that seems obvious in hindsight but was genuinely creative in 2017.
Here's what's wild: that paper is almost 9 years old, and it's STILL the foundation of everything. GPT, Claude, Llama, Gemini — all descendants of those 15 pages.
But transformers in 2026 look very different from 2017. Mixture of Experts for efficient scaling. State Space Models like Mamba challenging attention's dominance. Flash Attention making everything faster. KV-cache optimization that's basically its own engineering discipline.
My advice if you want to deeply understand modern AI: implement a transformer from scratch. Not using a library. From the matrix multiplications up.
It took me a weekend and it changed how I think about every model I work with. You stop seeing magic and start seeing math. And that's when you become dangerous.