**Alessio Fanelli** (0:10)
Hey everyone, welcome to the Latent Space Podcast.
This is Alessio, partner and CTN resident and decibel partner. I'm joined by my co-hosts, Swix, Redder and Edder of Latent Space.
**Swyx** (0:20)
Hey, and today we have Jonathan and Abhi from MosiacML. Welcome to the studio, guys.
**Jonathan Frankle** (0:25)
Thank you so much for having us.
**Abhinav Venigalla** (0:26)
Thanks so much for having us.
**Swyx** (0:27)
How's it feel recording in a real life studio?
**Jonathan Frankle** (0:30)
Honestly, I've been doing a lot of podcasts during the pandemic, and it has not been the same.
**Swyx** (0:34)
No, not been the same. Actually, so you have on your bio that you were, so you're primarily based in Boston.
**Jonathan Frankle** (0:42)
New York.
**Swyx** (0:42)
New York.
**Jonathan Frankle** (0:42)
Yeah, my Twitter bio was a probability distribution of the locations.
**Swyx** (0:46)
Exactly. So I DM'd you, because I was obviously very interested in MPT-7B, and DM'd you, I was like, for the 0.2% of the time that you're in San Francisco, can you come, please come to a podcast studio? And you're like, I'm there next week.
**Jonathan Frankle** (0:56)
Yeah, it worked out perfectly.
**Swyx** (0:59)
We're really lucky to have you. I'll read off a few intros that people should know about you, and then you can fill in the blanks. So Jonathan, did your BS and MS at Princeton in programming languages, and then found your way into ML for your PhD at MIT, where you made a real splash with the lottery ticket hypothesis in 2018, which people can check up on. I think you've done a few podcasts about it over the years, which has been highly influential. And we'll talk about sparse models at Mosiac.
You also had some side quests. You taught programming for lawyers, and you did some law and privacy stuff in DC, and also did some cryptography stuff.
And you've been an assistant professor at Harvard before earning your PhD.
**Jonathan Frankle** (1:39)
I have yet to start.
**Swyx** (1:40)
You're yet to start. Okay, but you just got your PhD.
**Jonathan Frankle** (1:42)
I technically just got my PhD. Mosiac delayed my defense by about two years.
I was at 99% done for going on two years, got the job at Harvard, Mosiac started, and I had better things to do than write my dissertation for two years.
**Swyx** (1:57)
You know this is like very out of order.
**Jonathan Frankle** (1:59)
Oh, completely out of order. Completely backwards. Go talk to my advisor about that.
He's also an advisor at Mosiac and has been from the beginning. And you know go talk to him about finishing on time.
**Swyx** (2:10)
Great, great, great. And just to fill it out, Avi, you did your BS and MS at MIT. You're a researcher at Cerebris.
And then you're now a research scientist at Mosiac. Just before we go into Mosiac stuff, I'm actually very curious about Cerebris and just that space in general. What are they doing that people should know about?
**Abhinav Venigalla** (2:29)
Yeah, absolutely. I think the biggest thing to take away about Cerebris is that they're really building the next-gen computing platform beyond like end of GPUs. They're trying to build a system that uses an entire wafer rather than cutting up a wafer into smaller chips and trying to train a model on that entire system or actually more recently on many such wafers. It's really extraordinary. I think it's the first time ever that wafer-scale computing has ever really worked.
And so it's a really exciting time to be there trying to figure out how we can map ML workloads to work on a much, much bigger chip.
**Swyx** (2:59)
And do you use a different programming language or framework to do that?
**Abhinav Venigalla** (3:04)
Yeah, so things have changed a bit since I was there. I think you can actually run just normal TensorFlow and PyTorch on there.
So they've built the software stack that compiles it down so it actually just kind of works naturally.
**Swyx** (3:15)
Compiled versions of Python is a hot topic at the moment with Mojo as well. And then Mosiac, you spearheaded the MPT-7B effort.
**Abhinav Venigalla** (3:23)
Yeah, so it's kind of like it's been maybe six months, 12 months in the making. We kind of started working on LLMs sort of back in the summer of last year. And we came up with this blog post where we kind of profiled a lot of LLMs and saw, hey, the cost of training is actually a lot lower than what people might think.
And then since then, you know, being inspired by kind of Meta's release of the Llama models and lots of other open source work, we kind of started working towards, well, what if we were to release like a really good kind of 7 billion parameter model? And that's what MPT is.
73 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000613807558