**SPEAKER_1** (0:00)
Welcome back. Right after Christmas, the Chinese whale bros ended 2024 by dropping the last big model launch of the year, DeepSeek V3. This is a massive 671 billion parameter, fine-grained MOE model with 256 experts, trained with native FP8 mixed precision training, multi-head latent attention from DeepSeek V2, a new multi-token prediction objective, and 15 trillion tokens of data including synthetic reasoning data distilled from DeepSeek R1. Right now, on the LM Arena leaderboard, DeepSeek V3 is rated the 7th best model in the world with a score of 1319, right under the full O1 model, Gemini 2 and 4o latest and above O1 Mini, Grok 2, Gemini 1.5 Pro and Claude 3.5 Sonnet. This makes it the best open weights model in the world in January 2025 There has been a big recent trend in Chinese labs releasing very large open weights models, with Tencent releasing Hunyuan-Large in November and Hailuo releasing MiniMax-Text this January, both over 400B in size. However, these extra-large language models are very difficult to serve. Baseten was the first of the Inference NeoCloud startups to get DeepSeek V3 online because of their H200 clusters, their close collaboration with the DeepSeek team and early support of SGLang, a new VLLM alternative out of UC Berkeley that is also used at frontier labs like XAI.
Each H200 has 141GB of VRAM with 4.8TBps of bandwidth, meaning that you can use 8 H200s in a node to Inference DeepSeek V3 in FP8, taking into account KV cache needs. We have been close to base 10 since Sarah Guo introduced Amir Haghighat to Swix and they supported the very first Latent Space demo day in San Francisco, which was effectively the trial run for the podcast you're listening to right now. Since then, Philip Kiely has also led a well-attended workshop on TensorRT LLM at the 2024 World's Fair. We worked with him to get two of their best representatives, Amir and Lad model performance engineer Yineng Zhang, to discuss DeepSeek, SG Lang and everything they have learned running mission critical inference workloads at scale for some of the largest AI products in the world. Spoiler! Amir thinks there are three pillars of mission critical inference workloads, and we spend quite some time discussing what you need for each of them. In other news, invites are now rolling out for the second AI Engineer Summit in New York City from February 20th to 22nd. We are bringing back the surprisingly successful AI leadership track from World's Fair, and the AI engineering track is now wholly focused on agents at work. If you are building agents in 2025, this is the single best conference of the year. We are curating all attendees and will sell out after we announce speakers this coming week from DeepMind, Anthropic, OpenAI, Meta, Jane Street, Bloomberg, BlackRock, LinkedIn and more. Look for more sponsor and attendee information at apply.ai.engineer and see you there. Watch out and take care.
**Alessio** (3:40)
Hey everyone, welcome back to the Latent Space Podcast, our first recording of 2025 I'm Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host Spix, founder of SmallAI.
**Swyx** (3:50)
Hey, and today we are here with a special double guest episode with Amir. Oh my god, I don't know your last, Haghighat?
**Amir Haghighat** (3:59)
That's close enough, that is good. I thought I was close to the prize go, that's really good.
**Swyx** (4:05)
And Yineng Zhang from Baseten, welcome.
**Yineng Zhang** (4:07)
Thank you.
**Swyx** (4:08)
Amir, we've met before, you're co-founder of Baseten, which is one of the leading sort of LM inference platforms. I don't know how you, what do you consider yourself?
**Amir Haghighat** (4:17)
That sounds fine.
**Swyx** (4:18)
And Yineng, you are lead software engineer on the Model Performance team, and you guys recently shipped DeepSeek V3 as one of the many models that you do host. You also are very involved in SGLang, and that was actually one of the reasons we're discussing an episode with you even before DeepSeek V3 dropped as a Christmas present to everybody. So we can take this in a number of directions, but I think one thing we wanted to just get off the bat on was to start with DeepSeek and maybe, and then we'll work our way backwards back to SGLang, but DeepSeek is more recent. Why are people so interested? What's the history of DeepSeek in general from your perspective?
**Yineng Zhang** (4:57)
Yeah, because DeepSeek V3, I think is currently considered the leading open-source LLMs based on the benchmark results and the chat area results, and it's so big, it's 671 billion parameter MOE, and I think it's a game-changer for the open-source AI. So everyone is interested in this model.
49 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000684541124