Remaking the UI for AI artwork

Remaking the UI for AI

AI + a16z

April 19, 2024

a16z General Partner Anjney Midha joins the podcast to discuss what's happening with hardware for artificial intelligence.
Speakers: Anjney Midha, Derrick Harris
**Anjney Midha** (0:00)
I'm optimistic that this time around, there's sufficient excitement in the ecosystem, both from everybody from the hardware providers at the compute level, like NVIDIA, the cloud providers and startups, and ultimately investors like us who are really excited for new architectures. And so, you know, if there are people out there who are dinkering and experimenting with those that unlock new interfaces, that unlock the next phase in computing, yeah, that's what we're here to fund.
I just wish more people were working on those.

**Derrick Harris** (0:27)
Hi, this is Derek Harris, and you're listening to the a16z AI Podcast, where we dig into all things artificial intelligence with our in-house team of experts, as well as the founders, engineers, and researchers working at the state of the art. In this episode, I speak with a16z general partner Anjney Midha about how AI hardware will look in the years to come, and why there's so much innovation yet to happen at the inference layer. Among other things, he explains how he sees wearable devices evolving to take advantage of improvements in sensors and workload-specific chips and how the introduction of big company technology, like the Apple Vision Pro, can actually lay the foundation for startups. But because we recorded right after NVIDIA's big GTC event, as with our previous episode with Naveen Rao, we kick off the conversation talking about training workloads versus inference workloads and how NVIDIA came to dominate the former category.
As a reminder, please note that the content here is for informational purposes only, should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any a16z fund. For more details, please see a16z.com/disclosures.
I saw recently like, Olamo released support for AMD GPUs. I think I saw someone compare NVIDIA to Sun Microsystems in the early days of the web, which seems like it might be wishful thinking. Are skeptics kind of underestimating the stranglehold that NVIDIA has?

**Anjney Midha** (2:00)
There's the two schools of thought, which is that NVIDIA's stronghold on this is completely transitory. Their margins are overinflated because of supply chain crises over the last 24 months, where we had this explosion in demand, but production is gonna catch up. And basically, those folks will tell you, that school of thought will tell you, like, hey, the rate limiter here is silicon, it's sand. And there's tons of sand in the world. And eventually, somebody else will figure out how to make sand that does the same thing. There's tons of it on planet Earth. And I personally find that view is provocative, but reductive, because it doesn't, on the first principles analysis of the fact that training has these idiosyncratic needs, like a really robust software driver layer, right? That can orchestrate thousands of these chips acting in unison. I think, yes, on that side of the debate, I'm certainly one who believes that the developer experience that NVIDIA started, by the way, investing in a decade ago, is in its sort of later stages of compounding right now, and it's really hard to dislodge that.
What we may see is margins compress over time because the budgets are shifting from training to inference.
And I think that's actually where a lot of the exciting stuff is happening, and I know we're going to spend a bulk of time today hopefully talking about inference, because that's open season right now. Right. Every time you have a new software primitive, it often results in new kinds of workloads that the incumbents have a harder time keeping up with. And I would argue the inference workloads, like you mentioned with Olamma, for example, are entirely new kinds of compute workloads we haven't seen before. And so that's a much more even playing field. I don't think there was a way for NVIDIA to invest in that kind of workload 10 years ago, because it just didn't exist. Whereas training fundamentally has in some shape or form been around right now for the better part of a decade, because deep learning has been around for that long. And actually, I think this is a good time to introduce what I think is a useful mental model that I have about the future of hardware.
And I found there's several ways to reason about hardware. One, you can work backwards from the customer. Who is the customer here, and what do they need? What are their pain points? And then there's another way you can work, you can reason about is like reasoning from history and see what the progression and evolution of compute has been over time. And I think the history of hardware has been the history of computers, right? And in my mind, if you look at the last 60 years or so of computers that we've had, and this is basically modern computing, one popular way to reason about it is often the hardware versus software split. But I think there's another way to reason about it, which I'll give full credit to one of our founders, Ankit Kumar, who's spent a lot of time at Discord building the first thing I bought there and was reasoning about how to expose language models to large numbers of users before a lot of people got the chance to experiment with these large language models. He basically believes there's two lineages of computing. There's reasoning or intelligence, and then there's interfaces.

32 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000652980594