20VC: Positron AI on the Memory Wall, Data Centers, and Energy Bottlenecks artwork

20VC: Positron AI on the Memory Wall, Data Centers, and Energy Bottlenecks

AI Podcast Summaries from Transcripted.ai (VIDEO)

September 19, 2026

AI’s real choke point may not be compute — it may be energy, memory bandwidth, and the politics of building at scale.

Topics: Daily News, News

**SPEAKER_1** (0:01)
Power, politics, and artificial intelligence are colliding in fascinating ways right now. Harry Stebbings recently sat down with Thomas Somas, co-founder and chairman of Positron AI, the fabless semiconductor company that just raised an eight hundred seventy-five million dollar Series C at a five billion dollar valuation. And Somas opened with something unexpected—a political warning. He said the scariest thing to him is that being anti-data center has become a unifying issue across both the left and the right. His view is blunt: the opposition is built on bad information about water and power use. Right, but he also pointed out that the United States still has an advantage in open land and buildout capacity. So what does Positron actually do that makes this relevant to them?
They build hardware and software specifically for generative AI inference. And here's where it gets technical—Somas explains that inference is fundamentally different from training. Training is compute-bound, but inference is memory-bound because models generate tokens one by one and must repeatedly read weights from memory. That's the "memory wall" problem, right? Nvidia's compute performance has improved far faster than memory bandwidth, and Somas argues that mismatch is now central to the industry. He also highlights the KV cache, which stores keys and values so models don't have to recompute everything from scratch.
Building on that point, he says modern AI economics are even more attractive than most people realize. Cached tokens are far cheaper to process than fresh tokens, and providers can charge premium prices for them. Many API businesses are already highly profitable. The conversation then turns to regulation and geopolitics. Somas is skeptical of the "pacing the frontier" argument, worrying that calls to slow down can become a path toward blocking progress altogether. If the West slows while other countries don't, we're effectively surrendering the future. On infrastructure, he's firmly pro-build. He says data centers should be treated like modern wonders, not villains. There's zero scenario where a data center could pull power from what's already been allocated to homes, and more generation can actually lower costs over time. That's interesting because he's also unconcerned about local politics stopping the industry. If communities block one site, builders can move elsewhere. He even floats longer-term possibilities like space-based and ocean-based data centers. Energy, in his view, is the true bottleneck for civilization. He ties everything from fire to nuclear power to humanity's ability to produce energy economically.
Context windows and memory management also matter enormously—quantization, cache tiering, and sparse attention all stretch capability without exploding cost.
He's also bullish on bigger models. Local models won't replace cloud models so much as act like filters, constantly deciding when to call something smarter.
That leads to his biggest claim: he believes GPT-6 Astra is AGI.
His final message frames AI as a technical, economic, and societal coordination problem. Progress will depend on rational infrastructure, sane regulation, and a willingness to keep building.

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID