Mapping the future of *truly* Open Models and Training Dolly for $30 — with Mike Conover of Databricks artwork

Mapping the future of *truly* Open Models and Training Dolly for $30 — with Mike Conover of Databricks

Latent Space: The AI Engineer Podcast

April 29, 2023

The race is on for the first fully GPT3/4-equivalent, truly open source Foundation Model! LLaMA’s release proved that a great model could be released and run on consumer-grade hardware (see llama.
Speakers: Alessio, Swyx, Mike Conover
**Alessio** (0:10)
Hey, everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in Residence and Decibel Partners. I'm Joen Bamako, host of Swix, writer and editor of Latent Space.

**Swyx** (0:20)
Welcome, Mike.

**Mike Conover** (0:21)
Hey, pleasure to be here.

**Swyx** (0:23)
Yeah, so we tend to try to introduce you so that you don't have to introduce yourself, but then we also ask you to fill in the blanks. So you are currently a staff software engineer at Databricks, but you got a PhD at Indiana University of Bloomington in complex systems analysis, where you did some analysis of clusters on Twitter, which I found pretty interesting. Yeah, I highly recommend people checking that out if you're interested in getting information from indirect sources or I don't know how you describe it.
And then you went to LinkedIn, working on homepage news relevance, and then SkipFlag, which is a smart enterprise knowledge graph, which was then acquired by Workday, where you became director of machine learning engineering and now you're at Databricks. So that's the quick bio and we can kind of go over step by step. But what's not new LinkedIn that people should know about you?

**Mike Conover** (1:12)
So because I worked at LinkedIn, that's actually how new hires introduce themselves at LinkedIn, is this question. So I have a pat answer to it.
I love getting off trail in the back country.
And I think that the sort of like radical responsibility associated to that clarifies the mind. And I think that the things that I really like about machine learning engineering and sort of the topology of high-dimensional spaces kind of manifest when you think about a topographic mat as a contour plot. You know, it's a two-dimensional projection of a three-dimensional space.
And it's very much like looking at information visualizations. And you're trying to relate your localized perception of the environment around you and the contours of ridges that you see or basins that you might go into. There's that little creek down there and relate that to the projection that you see on the map. I think it's physically demanding. It's intellectually challenging. It's natural beauty is a big part of it. And you're generally spending time with friends. And so I just, I love that. I love that.

**Swyx** (2:19)
These are camping trips, multi-day?

**Mike Conover** (2:21)
Yeah. Yeah. Camping. I hunt too. You know, I shoot archery, big game, backcountry hunting. But yeah, you know, sometimes it's just, let's take a walk in the woods and see where it goes.

**Swyx** (2:33)
You ever think about going on one of those journeys in the Australian outbacks, like where people find themselves?

**Mike Conover** (2:42)
I like to fly fish.

**Swyx** (2:42)
You like to hill climb?

**Mike Conover** (2:43)
Yeah. Like the outback seems beautiful. I think eight of the 10 most deadly snakes live in Australia.
Yeah.

**Swyx** (2:52)
Any lessons from like, you know, real hill climbing versus machine learning?

**Mike Conover** (2:56)
Dude, it's a lot like gradient descent.
I have remarked on that to myself before, for sure. Yeah.

**Swyx** (3:05)
I'm not sure this is like...

**Mike Conover** (3:07)
That's least resistance, please.

**Alessio** (3:10)
That's awesome. So Dolly, you know, it's kind of came up in the last three weeks. You went from a brand new project at Databricks to one of the hottest open source things out there.
So March 24th, you had Dolly 1.0. It was a 6 billion parameters model based on GPDJ 6 billion and you saw Alpaca training set to train it.
First question is, why did you start with GPDJ instead of LLaMA, which was what everybody else was kind of starting from at the time?

**Mike Conover** (3:35)
I mean, so, you know, we had talked about this a little before the show, but LLaMA is hard to get. We had requested the model weights and had just not heard back. And, you know, I think our experience with the original email alias for Dolly before it was available on Huggingface, you get hundreds of people asking for it. And I think it's like, it's easy to just not be able to handle the inbound. And so, like, I mean, there was a practical consideration, which is that, you know, we did not have the LLaMA weights. But additionally, I think it's like much more interesting if anybody can build it. And so I think that was our...
And I had worked with the Gpt-J model in the past and knew it to be high quality from a grammatical-ness standpoint.
And so I think it was a reasonable choice.

**Alessio** (4:18)
Yeah. Yeah.

**Swyx** (4:19)
Maybe we can also go into the impetus of why you started to work on Dolly. You had been at Databricks for about a year. Was this like a top-down directive? Was this your idea?

66 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000611082081