**Jason Howell** (0:08)
This is the Daily Tech News for Thursday, July 16th, 2026 We tell you what you need to know, give you the important context, and help each other understand.
**Huyen Tue Dao** (0:17)
Today, a big Suno leak reveals exactly how the company trains its models to generate music for users.
**Jason Howell** (0:24)
Hmm, yeah, there's a lot of interesting details here we're gonna talk about. I'm Jason Howell.
**Huyen Tue Dao** (0:29)
I'm Huyen Tue Dao.
**Jason Howell** (0:30)
Let's start with what you need to know with Big Story.
All right, 404 Media published a report based on data from a 2025 hack of Suno. Suno, of course, being one of the major music generation companies. You know, another one is UDO. There's a bunch of them, but Suno is kind of like, I'd say one of the biggest and most well known. And the hack data that was revealed shows internal code and comments that explicitly describe how to scrape the music and lyrics that were needed to build the system, essentially. And some sources that were named specifically in this leak, YouTube Music, Deezer, Genius for lyrics, stock music sites like Pond5, Jamendo, or Jamendo, I think, Freesound Library, and then the International Music Score Library Project. All of those sources, perhaps there's more, but at least these to act as its training data library. And so along with that, with that information, it included filters to remove non-music as it called it.
Because I mean, this is a music generation service. So of course, they're, they don't, they don't want to muck things up with people just talking like on podcasts and stuff. Though, speaking of podcasts, it does also mention that they effectively scraped podcasts as well, both from those platforms and then also RSS feeds directly. Dataset descriptions list specific volumes. For example, 2,013,545 music clips pulled from YouTube Music, more than 113,000 hours from YouTube Music audio as well, and there's plenty more data points in there, but you don't need me to rattle through a bunch of numbers. The hack did also include some user data and not just some, hundreds of thousands of Suno users, Stripe payment data, emails, phone numbers, that sort of stuff, though Suno is saying that the data breach was limited. Also that the code that was revealed here is outdated. It's not currently in use. You aren't going to find us doing this. That's not now. That was then. That was a long time ago.
Deezer has said that it's assessing the situation, considering its options. YouTube Music and Genius, then some of these other companies, I haven't found any acknowledgement yet on the news publicly, but that's where we're at, at a moment in time where Suno is still under the target or in the crosshairs of a few very specific record industry lawsuits. This comes out. What are your thoughts?
**Huyen Tue Dao** (3:24)
I don't like it.
Acknowledging that there's parts of this that no one can do a thing about whether, and I suppose actually I would want to ask you too about how you feel about, especially having been podcasting so long, how you feel the podcast are. But there's certain things like I presume, stock sites and even free sound, things that are license-free, that's fine. It's always been very interesting. I think especially as someone who really enjoys programming and likes to kind of has done a hobby project my entire career, the idea of scraping data in just in general is a little bit gray. You know what I'm saying? It's like to some extent, some data sets are inherently free or feel like they should be free. So I definitely know people that have scraped data sets and haven't felt too bad about it. But I think there's a general sense that generally speaking, data scraping is gray at best or maybe on average. So no, I don't like a lot of this, obviously, not that I want to defend big music companies or anything, but obviously, services like YouTube, Deezer, Genius have licensing agreements with companies and training on that without explicit permission or relationships or agreements was seems like a bad thing. And I know that they say the code isn't being used now, but that is not really a thing anymore, is it just because of the current age of AI that we're in? The data has been trained upon, unless they're throwing out the entire model and starting from scratch. The thing is, they've already benefited from this code. I guess it has some impact in that they're not doing it anymore. But this has been the thing, especially with these models, before any of us were really aware, before they really took hold in our daily lives, people were scraping data. We've talked about stories of like meta being sued. Was it meta or Facebook? I guess meta was scraping books, right?
23 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000777103584