**SPEAKER_2** (0:03)
Everyone is talking about ChatGPT right now, but are you actually using it to the max? How do you use ChatGPT? It's an interview show where the people at the forefront of technology show you how they use ChatGPT in their work and their lives. Host Dan Shipper talks to programmers, writers, founders, academics, tech executives, and others to walk through all of their ChatGPT use cases, including historical chats step-by-step.
They even use ChatGPT together live on the show to build apps, analyze their leadership qualities, read more deeply, and do the best work of their lives. Listen to How Do You Use ChatGPT from Dan Shipper and the team at Every, wherever you get your podcasts.
**Lin Qiao** (0:47)
Thank you That is a complexity. All application product developers, as they are doing things, fun stuff themselves or in enterprises, they're all facing this challenge. So that's where we come in, and say, you don't worry about it.
We handle it all for you. So you just focus on your product application development.
**Dmytro Ivchenko** (1:07)
I'm gonna just directly apply the techniques we learned from text models on the image model, because it has quality, so you need to do some extra work to make sure the quality is progressing. So that is quite a bit different.
**Lin Qiao** (1:22)
Over time, all these database managers become smarter and smarter, because they all have a layer called optimizer. Those optimizer observe the workload, and they start to create, oh, you're doing a lot of filter on this particular column, so I'm gonna create index. I'm gonna partition those columns based on your future criteria. So it's much better search, much faster search.
**Nathan Labenz** (1:42)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together, we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello, and welcome back to The Cognitive Revolution. Today, my guests are Lin Qiao and Dmytro Ivchenko, co-founders of Fireworks AI, a company that specializes in inference compute, partnering with the world's leading generative AI researchers to serve the best models at the fastest speeds.
Lin and Dmytro both previously worked on PyTorch at Meta, which is today the default go-to AI framework powering applications used by billions. There, they gained firsthand experience with the immense challenges of running large language models at a massive scale and the many trade-offs between latency, cost, and scalability that are always involved. With Fireworks, they're building an end-to-end platform to make it radically easier and more cost-effective for any company to put generative AI into production. This spans the full technology stack, including providing simple tools for executing parameter-efficient fine-tuning techniques like LoRa that help developers iterate quickly toward product-market fit. Also, developing highly optimized deployments, leveraging multiple layers of abstraction, including custom CUDA kernels, to deliver consistently low-latency. And managing and scaling hardware across major cloud compute providers in a way that's seamless to their customers. This is a wide-ranging discussion. Lin and Dmytro share their hard-earned expertise on the intricacies of AI inference, and we dive deep into the weeds on topics including the different priorities that their customers have, such as minimizing time-to-first token, which is particularly important for voice applications, how some inference compute providers today are using Uber-style subsidized pricing to win business, and why Lin thinks developers should be cautious about building on these platforms. Also, why she considers OpenAI and Anthropic to be Fireworks' real long-term competition.
Why Fireworks is betting that all small models, whether open or closed source, will ultimately converge in capabilities. The main parallelization techniques, including tensor and pipeline parallelism that they're using, to spread models across GPUs in different ways with different benefits. Why software is struggling to keep up with the pace of advances in hardware. And how Fireworks is working toward an automated optimizer that will eventually allow even non-technical customers to choose the best configurations for their use cases.
Finally, at the end, we brought Dmytro back for a short bonus discussion to cover their recently announced partnership with Stability AI AI, which has been powering Stable Diffusion 3 generation on an exclusive basis. We talked a bit about some of the subtle differences between the image and text generation use cases, and overall, I came away with the sense that this partnership makes a ton of sense and might become a new pattern in the industry as research groups look to make their work widely and effectively available, while also finding ways to earn a return on their investment. Whether you are an AI engineer wrestling with model deployment or an executive evaluating AI platforms, Lin and Dmytro offer a rare peek behind the curtain of the infrastructure layer, which may not get as much hype as the latest state-of-the-art model, but is obviously critical to realizing the potential of this technology.
85 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000653044099