Powering your Copilot for Data – with Artem Keydunov of Cube.dev artwork

Powering your Copilot for Data – with Artem Keydunov of Cube.dev

Latent Space: The AI Engineer Podcast

October 26, 2023

The first workshops and talks from the AI Engineer Summit are now up! Join the >20k viewers on YouTube, find clips on Twitter (we’re also clipping @latentspacepod), and chat with us on Discord! Text-to-SQL was one of the first applications of NLP.
Speakers: Swyx, Alessio, Artem Keydunov
**Swyx** (0:06)
Hey, everyone, welcome to the Latent Space Podcast. This is Swyx, writer and editor of Latent Space and founder of SmallAI, and Alessio, partner and CTO in residence at Decibel Partners.

**Alessio** (0:15)
Hey, everyone. And today we have Artem Keydunov on the podcast, co-founder of Kube. Hey, Artem.

**Artem Keydunov** (0:21)
Hey, Alessio. Hi, Swyx. Good to be here today. Thank you for inviting me.

**Alessio** (0:24)
Yeah, thanks for joining. For people that don't know, I've known Artem for a long time, ever since he started Kube. And Kube is actually a spin out of his previous company, which is Statsbot.
And this kind of feels like going both backward and forward in time. So the premise of Statsbot was having a Slackbot that you can ask busy like text to SQL in Slack. And this was six, seven years ago, something like that, a lot ahead of its time. And you see startups trying to do that today.
And then Kube came out of that as a part of the infrastructure that was covering Statsbot. And Kube then evolved from an embedded analytics product to the semantic layer and just an awesome open source. I think you have over 16,000 stars on GitHub today. You have a very active open source community. But maybe for people at home, just give a quick lay of the land of the original Statsbot product. What got you interested in like text to SQL and what were some of the limitations that you saw then, the limitations that you're also seeing today in the new landscape.

**Artem Keydunov** (1:28)
I started Statsbot in 2016
The original idea was to just make sort of a side project based off my initial project that I did at a company that I was working for back then. And I was working for a company that was building software for schools.
And we were using Slack a lot. And Slack was growing really fast. A lot of people were talking about Slack, you know, like Slack apps, chat spots in general. So I think it was, you know, like another wave of, you know, bots and all that. We have one more wave right now, but it always comes in waves. So we were like living through one of these waves.
And I wanted to build a bot that would give me information from the different places where like a data lives to Slack. So it was like developer data, like New Relic, maybe some marketing data, Google Analytics, and then some just regular data, like a production database is what it sells for sometimes. And I wanted to bring it all into Slack because we were always chatting, you know, like in Slack, and I wanted to see some stats in Slack.
So that was idea Statsbot, right? Like bring stats to Slack. I built that as a, you know, like a first sort of a side project and I published it on Reddit and people started to use it even before Slack came up with that Slack application directory. So it was a little, you know, like a hackish way to install it, but people are still installing it. So it was a lot of fun. And then Slack kind of came up with that application directory and they reached out to me and they wanted to feature Statsbot because it was one of the already being kind of widely used bots on Slack. So they featured me on this application directory front page and I just got a lot of, you know, like new users signing up for that. It was a lot of fun, I think, you know, like, but it was sort of a big limitation in terms of how you can process natural language because the original idea was to let people ask questions directly in Slack, right? Hey, show me my, you know, like opportunities closed last week or something like that. My co-founder who kind of started helping me with this Slack application, him and I were trying to build a system to recognize that natural language, but it was, you know, we didn't have LLMs, right, back then and all of that technologies. So it was really hard to build the system, especially the systems that can kind of, you know, like keep talking to you, like maintain some sort of a dialogue. It was a lot of like one-off requests and like it was a lot of hit and miss, right? If you know how to construct your query in natural language, you will get a result back. But you know, like it was not a system that was capable of, you know, like asking follow up questions to try to understand what you actually want and then kind of finally, you know, like bring this all context and go to generate a SQL query, get the result back and all of that. So that was the really missing part. And I think right now that's, you know, like what is the difference? So right now I kind of bullish that if I would start Statsbot again, probably would have a much better shot at it. But back then that was a big limitation.

35 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000632731344