Unlocking Enterprise Data with Knowledge Graphs and AI, with Juan Sequeda, Principal Scientist and Head of AI Lab at data.world artwork

Unlocking Enterprise Data with Knowledge Graphs and AI, with Juan Sequeda, Principal Scientist and Head of AI Lab at data.world

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

January 18, 2024

In this episode, Nathan sits down with Juan Sequeda, Principal Scientist and Head of AI Lab at data.world. They discuss how knowledge graphs can be your organization's "brain" for AI, integrating structured and unstructured data, benchmarking enterprise AI systems, and more.
Speakers: Erik Torenberg, Juan Sequeda, Nathan Labenz
**Erik Torenberg** (0:00)
Turpentine is a network of podcasts, newsletters, and more, covering tech, business, and culture, all from the perspective of industry insiders and experts.
We're the network behind the show you're listening to right now.
At Turpentine, we're building the first media outlet for tech people by tech people. We have a slate of hit shows across a range of topics and industries, from AI with Cognitive Revolution to Econ 102 with Noah Smith. Our other shows drive the conversation in tech with the most interesting thinkers, founders, and investors, like Moment of Zen and my show Upstream. We're looking for industry-leading hosts and shows along with sponsors. If you think that might be you or your company, email me at erikaturpentine.co. That's E-R-I-K at turpentine.co.

**Juan Sequeda** (0:45)
This is the essence of what my organization is.
It's the brain, right? Everything here is accurate. I can use this to explain things.
The LLM, these foundational models, don't have that accuracy, don't have that explainability, don't know my organization. Do we expect these foundational models to know every single organization? No, because these things is private. I don't want them to know, but I want to go use them. That's why this combination of these foundational models, large-language models, with your internal brain of organization, which is your knowledge graph, that's what I see where there's a future. I think what we need to work on is understanding best that integration point between the knowledge graph, the brain of your organization, and these foundational models.

**Nathan Labenz** (1:30)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello, and welcome back to The Cognitive Revolution. Today, we're talking data and how generative AI interacts with enterprise data with my guest, Juan Sequeda, Principal Scientist and Head of the AI Lab at data.world. Data.world, like many established companies that we've featured on the show over time, built much of its platform, products, team, and business before the current generative AI moment, and is now working to make sure that it's taking full advantage of this new technology paradigm for its customers. This is part of an economy-wide trend. Frontier generative AI models are now making their way into all sorts of high-value, often very challenging contexts. From processed chemistry labs, to novel protein design, to US federal courtrooms, to the practice of medical diagnosis, to enterprise data lakes, GPT-4 demonstrated enough raw power that we now have experts in literally every field working day in and day out to figure out how to make generative AI work for their organizations.
Now, it's not all instant success. From Juan's work, for example, it's clear that early benchmarks understate the complexity of real-world enterprise data, and that naive chat-with-your-data-type implementations are not up to enterprise challenges. Even the more advanced work that Juan and team are doing with Knowledge Graphs, while it does deliver major improvement, is still at best a partial solution.
So while Data.world, again, like most established technology businesses I've talked to, is not particularly worried about competition from fast-moving AI-first startups, they do see the transformative potential, and they are realizing enough practical utility today that they are rolling up their sleeves and settling in to the process of AI implementation and optimization.
Interestingly, zooming out, I think this creates the potential for another major phase change in the history of AI. While GPT-4 takes meaningful effort to implement and often still falls short of our dreams, we are nevertheless building foundational capacity for both last-mile distribution and customization, such that the next big model release will have the opportunity to almost immediately plug into many millions of live business processes and systems.
The electrification of America took 60 years and included significant public works projects, and all the appliances were designed with a clear understanding of exactly how they'd be supplied with electrical power. Today, in contrast, we live in an Internet-mediated, software-enabled world in which updates can quickly be pushed to everyone, everywhere, all at once. Today's software application developers are building somewhat ahead of AI capability, both so that they can deliver frontier features to their users today, and more importantly, so that they're ready to flip the switch when the next big advance comes online. How many months will need to wait before GPT 4.5 or GPT 5? I don't know, and it might be just long enough that folks start to wonder whether AI generally is underpowered and overhyped. But my expectation remains very clearly that we will continue to see additional jumps in capability, and that with each future leap, considering the foundation now being laid, the deployment cycles will naturally get shorter and more disruptive. One note for listeners, enterprise data is a complicated space, and we spend some time in the first half of this conversation discussing the general state of enterprise data and data science teams today.

101 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000642038473