**Erik Torenberg** (0:00)
Turpentine is a network of podcasts, newsletters, and more, covering tech, business, and culture, all from the perspective of industry insiders and experts.
We're the network behind the show you're listening to right now.
At Turpentine, we're building the first media outlet for tech people by tech people. We have a slate of hit shows across a range of topics and industries, from AI with Cognitive Revolution to Econ 102 with Noah Smith. Our other shows drive the conversation in tech with the most interesting thinkers, founders, and investors, like Moment of Zen and my show Upstream. We're looking for industry leading hosts and shows along with sponsors. If you think that might be you or your company, email me at erikaturpentine.co. That's E-R-I-K at turpentine.co.
**Anton Troynikov** (0:46)
ML models are weird in the way that they think about the world. They might represent something quite differently to what a human would. They might focus on quite different things because they are mechanistically pursuing an objective. They are not trying to encode any meaning. The Wright brothers were wrong about how the Wright flyer worked. They just had intuitions about it. It flew. That's an object demonstration. It doesn't really matter if your intuition is wrong if the thing works.
And then there's this Cambrian explosion of different types of aircraft. Because all we had to go on was intuition. We didn't have engineering principles yet and we didn't even really have the tooling to develop those engineering principles.
That's where we're at today in my opinion with machine learning. And one of my favorite things about Kurzweil is I went on the website not long ago and it was like the 10th anniversary of the Singularity is Near and that just, I don't know, that made me laugh.
**Nathan Labenz** (1:37)
Hello and welcome to The Cognitive Revolution where we interview visionary researchers, entrepreneurs and builders working on the frontier of artificial intelligence.
Each week we'll explore their revolutionary ideas and together we'll build a picture of how AI technology will transform work, life and society in the coming years. I'm Nathan Labenz joined by my co-host Eric Thornburg.
**Erik Torenberg** (1:59)
The Cognitive Revolution podcast is supported by Omneky. Omneky is an omni-channel creative generation platform that lets you launch hundreds of thousands of ad iterations that actually work, customized across all platforms with the click of a button.
Omneky combines generative AI and real-time advertising data to generate personalized experiences at scale.
**Nathan Labenz** (2:19)
Anton Troynikov is a co-founder of Chroma, the AI native open-source embedding database. In the midst of an AI gold rush, Anton is selling shovels. Embeddings for any who don't know are numerical representations, generally speaking a single vector or simply put a list of numbers, that encode an input that could be text, an image, an audio file, or even a DNA sequence or protein structure into a lower-dimensional latent space.
The embedding process is itself learned, such that semantically similar inputs cluster together in latent space, and the resulting embeddings can be used to compare inputs, to power retrieval systems, and for additional processing. Embeddings are a key piece of a huge number of AI-based projects right now. By embedding a dataset, whether it's a user's email history or Google Drive documents, or even a company's entire knowledge base, developers can enable semantic search at runtime, and include the most relevant background information in language model context. This dramatically improves performance by reducing hallucinations and increasing personalization.
We talked to Anton about embeddings, the trade-offs and optimizations involved with working with embeddings at scale. Their recent Stable Attribution project, which flexes the power of Chroma by taking a stable diffusion-generated image and attempting to determine which of the 5 billion images in the Lion training set had the most influence on its creation, but which also quickly became a flashpoint in the broader debate about how artists and other rights holders should be credited and compensated for their contributions to generative models. We also covered the big picture of recent AI developments. Anton is a deep thinker with a sophisticated conceptual understanding of latent space. I learned a lot from this conversation, and I hope you do too. Anton Shroeneckoff, welcome to The Cognitive Revolution.
**Anton Troynikov** (4:20)
Glad to be here.
**Nathan Labenz** (4:23)
So tell me first about vector databases. I think our audience is obviously interested in AI. They most probably have at least a superficial sense of what a similarity score is, what a dot product is.
But take me to how that works as a database, and then we'll get into some of the optimizations there as well.
**Anton Troynikov** (4:46)
Yeah, sounds great. So let's kind of start at the beginning.
77 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000602481455