Performance and Passion: Fal's Approach to AI Inference artwork

Performance and Passion: Fal's Approach to AI Inference

AI + a16z

August 1, 2025

If you've been experimenting with image, video, and audio models, the chances are you've been both blown away by how good they're becoming, and also a little perturbed by how long they can take to generate.
Speakers: Burkay Gur, Batuhan Taskaya, Jennifer Li
**Burkay Gur** (0:00)
In general, there is still a lot of demand for image models, and there seems to be some kind of convergence on quality, but then each model really has its own differentiation. With video, we're earlier in the competition, there's still a lot of leapfrogging happening, there's just so much more to build, and there's just like, we haven't hit a quality bar where there's just marginal improvements, we're not there yet. So over there, it's more fierce competition, and it's very hard to predict what's going to happen next month. It's like that's, we're operating at the scales of like weeks at this point.

**Batuhan Taskaya** (0:35)
I remember when Sora came out, even in our team, people were like, oh my God, OpenAI is like so far ahead that no one's going to be able to catch up. And then like Luma released their model, Runway released their model, Kling released their model, Minimax released, and every release, like, if you're not the best, you're not releasing generally, that's how it works. You can never like say, oh, this is the model, and then this is not going to have any competition for like, even a month, right, like even like two weeks is like, I think that we are operating at like the weeks, as Burkay said.

**SPEAKER_3** (1:04)
Thanks for listening to the a16z AI podcast. If you've been experimenting with image, video and audio models, the chances are you've both been blown away by how good they're becoming, and also a little perturbed by how long they can take to generate. If you're using a platform like Fal, however, your experience on the latter point might be more positive. In this episode, Fal co-founder and CEO, Burkay Gur, and head of engineering, Batuhan Taskaya, joined a16z general partner, Jennifer Li, to discuss how they built an inference platform, or as they call it, a generative media cloud, that's optimized for speed, performance, and user experience. These are core features for a great product, yes, and also once boredom necessity, as the early team obsessively engineered around its meager GPU capacity in the height of the AI infrastructure crunch. But this is more than a story about infrastructure. As you'll hear, they also delve into sales and hiring strategy, the team's overall excitement over these emerging modalities, and the trends they're seeing as competition in the world of video models especially heats up. So, enough from me. Stick around and hear Burkay, Batuhan and Jennifer dive into all things generative AI after these disclosures.
As a reminder, please note that the content here is for informational purposes only, should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any a16z fund. For more details, please see a16z.com/disclosures.

**Burkay Gur** (2:38)
We started about four years ago. The origin story is I used to work at Coinbase, and they had a lot of infra issues with respect to machine learning, and fraud was a big problem there. So I kind of grew up in that environment where we're just constantly fighting with fraud using machine learning models. And the initial idea had a lot to do with building these pipelines for companies to be able to train these models. But about a year and a half into us starting the company, ChatGPT happened, Dali happened, the whole world of machine learning and AI changed. So we sort of adopted as things were developing.

**Jennifer Li** (3:15)
We're going to definitely dig into the 2021 wind shift happening on the multimedia side. But before we go there, how did you meet and recruit Batuhan, the Fal guy? Now he is the the mascot on Twitter of Fal.

**Burkay Gur** (3:30)
Yeah, we're both from Turkey. And I first saw Batuhan online. I was pretty curious about his work on Python. He's a pretty big Python contributor. I just DMed him on Twitter and we had a call and we had just started the company. And he was in, I think, Poland.

**Jennifer Li** (3:49)
Yes.

**Burkay Gur** (3:50)
In a dorm room, I think. And I was pitching Fal, like, hey, we're doing a lot of Python stuff. Would you come join us? And initially it was like very much an intro call. And I think at the time he was working on something else. It wasn't going to work out. A few months later, we actually raised from you guys. So this time I was like, okay, we have some funding, great investors, I'll go pitch it again. And this time we got on a call again and he happened to be leaving. And then, believe it or not, when I said, hey, like we just raised from Andreessen, it's not announced yet, we're going to recruit a lot of people, like we're going to be one of the first people to join. That's when he was really convinced.

35 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000720224375