Synthetic Data with Alex Watson, Founder of Gretel AI artwork

Synthetic Data with Alex Watson, Founder of Gretel AI

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

November 14, 2023

In this episode, Nathan interviews Alex Watson, founder and CPO of Gretel AI, about the company's work in synthetic data.
Speakers: Erik Torenberg, Alex Watson, Nathan Labenz
**Erik Torenberg** (0:00)
Turpentine is a network of podcasts, newsletters, and more covering tech, business, and culture, all from the perspective of industry insiders and experts.
We're the network behind the show you're listening to right now.
At Turpentine, we're building the first media outlet for tech people by tech people. We have a slate of hit shows across a range of topics and industries, from AI with Cognitive Revolution, to Econ 102 with Noah Smith. Our other shows drive the conversation in tech with the most interesting thinkers, founders, and investors, like Moment of Zen and my show Upstream. We're looking for industry-leading hosts and shows along with sponsors. If you think that might be you or your company, email me at erik.turpentine.co. That's E-R-I-K at turpentine.co.

**Alex Watson** (0:45)
Your data is messy. It has gaps in it. I can't create new additional examples. It's too expensive or there's no way to go back to it.
So we really focused our efforts on, first and foremost, helping you build better data.
That's been the kind of the guiding light. That's what we're really aiming for. You know, no LLM today can generate a hundred thousand or a million row data set. So the first purpose of the agent was interpreting that user query that's coming in and then figuring out how to divide it up into a set of smaller problems that the LLM can work on one problem at a time. The promise of a really lightweight model, really fast model, like shows the power that you can have of taking a domain specific data set you have or task and doing something meaningful without having to do something at the GPT-4 scale.

**Nathan Labenz** (1:32)
Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week we'll explore their revolutionary ideas and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host Eric Thornburg. Hello and welcome back to The Cognitive Revolution. Today my guest is Alex Watson, founder and chief product officer at Gretel AI, the synthetic data platform for developers.
Synthetic data is a fascinating topic. Since the early days of deep learning, it's been well known that training computer vision models on a mix of original and programmatically altered and degraded images ultimately improves model performance. It seems that learning the concepts through the noise boosts robustness to the random unseen oddities that models inevitably encounter in the wild. And more recently, dozens, maybe even hundreds of papers have explored how LLM-generated data can be used to improve training sets and ultimately model performance on a wide range of problems. Yet at the same time, some research results and many observers of the evolution of the Internet in general have cast doubt on just how much synthetic data the system can absorb before models begin to lose touch with their real world origins or otherwise degrade.
With these questions in mind, I reached out to Alex, who's been building a business on synthetic tabular data generation since 2020 and who proved to be an amazing guide to this domain. While synthetic data might sound like a niche topic, I think this conversation will be of general interest. We started with a discussion of why we need synthetic data, how Gretel has trained specialist models to maintain realism while also preserving privacy in creating it, and how we can be confident that we can trust this data for analysis, testing, and yes, AI model training.
Along the way, we also explored the trade-offs between statistical realism and social manners. The impact of LLMs on Gretel's business and the new pre-trained tabular LLM that they've recently introduced to help create synthetic data on a zero-shot basis for a wide range of data types and scenarios. We even took a detour into AI regulation in the wake of the recent Biden-Whitehouse executive order and the UK AI Safety Summit.
This episode is a great example of why I love making this show. I learned a ton in the preparation and had a lot of fun with the conversation, and I think you will too. If so, I always appreciate it when listeners share the show with their friends. And of course, we invite your feedback via our email at tcr at turpentine.co or via your favorite social network. For now, I hope you enjoy this conversation with Alex Watson of Synthetic Data Company, Gretel AI. Alex Watson, welcome to The Cognitive Revolution.

**Alex Watson** (4:24)
Appreciate it. Thanks, Nathan. Excited to be here.

**Nathan Labenz** (4:26)
Yeah, this is going to be great. So you are the founder and now chief product officer at this company, Gretel AI. I love to hear how you come up with that name, by the way.

74 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000634742460