**Elad Gil** (0:05)
What would the world look like if we could create biological software that allows us to compile RNA? That's the big question this week on the podcast. Sarah and I are sitting down with Jakob Uszkoreit, co-founder and CEO of Inceptive. Jakob spent more than a decade at Google, where he co-authored the Attention is All You Need paper, and several other papers has set the foundation for today's AI revolution. He has also started and led the research teams at Transform Google Search, Google Translate, and Google Assistant.
Now at Inceptive, he builds biological software with the aim to make widely accessible medicines and biotechnologies. Jakob, welcome to No Priors.
**Jakob Uszkoreit** (0:39)
Thank you. Thank you for having me.
**Elad Gil** (0:42)
You worked at Google for more than a decade, working on many leading research teams.
You were really seminal in the original Transformer paper. When I talk to the other authors of the Transformer paper, people in the know at Google, you're widely credited with really coming up with the idea of focusing on Attention, which was the basis for the Attention is All You Need paper. Could you talk a little bit more about how you came up with that and how the team started working on it and the origins of that pretty foundational breakthrough in terms of the Transformer?
**Jakob Uszkoreit** (1:09)
It's really not that simple. It's also really important to keep in mind that always in deep learning, you can't make something, in quotes, really work that is maybe pretty far on, say, the theoretical or formal end without really going deep on the engineering and implementation side. It just has to be efficient. At the end of the day, in my mind, that's the one and only thing we know really works. If you want to push deep learning forward, just to make it faster and more effective and more efficient on a given piece of hardware. There's a lot of evidence that the way we actually understand language, and that's something that then shapes language in terms of its statistical properties, is actually it's somewhat hierarchical and the best piece of just circumstantial or anecdotal evidence for that is just looking at what the linguists do. They draw these trees and while I don't think that they're ever really true, they're also definitely not always false.
They do capture some of the statistics that are inherent in language and probably language was actually evolved this way in order to exploit our cognitive capacities really in a fairly optimal way. And so you can safely assume that it is not necessary to go through the entirety of a sequential signal beginning to end and maybe also end to beginning simultaneously in order to understand it. But actually you can gain a lot of the understanding in air quotes by looking at individual groups of, say, your signal.
And ultimately, if you now are given a piece of hardware that has the very key strength of doing lots and lots of simple computations in parallel as opposed to complicated, structured computations sequentially, then really that's actually a kind of statistical property you really want to exploit. You want to in parallel understand pieces of an image first. And then maybe that's not possible in its entirety, but you can actually get a lot of it. And then only once you've done some of that, you put these incomplete understandings or representations together. And as you put them together more and more, that's when you disambiguate the last remaining or that's when you get rid of the last remaining ambiguity at the end of the day. And when you think about what that process looks like, it's a tree.
And when you think about how you would actually run something that evaluates all possible trees, then a reasonable approximation is that you repeat an operation where you look at all combinations of things. That's this quadratic step, right? That ultimately is at the core of this attention step. And then you effectively pull information in for a given representation of a given piece, the other representations of all the other pieces and rinse and repeat.
And it seems intuitive and also seems intuitively clear that that's a really good fit for the kind of accelerators that we had at the time that we still have today. And so that's really where that idea came from. And if you want to look at, say, the biggest differences, for example, between the transformer, as it was described in the Attention is All You Need paper and some of its ancestors like this decomposable attention model, the big difference is just that the transformer was implemented by folks like NoHarm Industries, et cetera, in a way that's such an excellent fit for the accelerators that we had at the time.
29 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000625556649