**Julian Goldie** (0:00)
Le Chaton Fat, a new model, Mistral, that is outperforming Fable 5 on every benchmark. If you haven't seen this, it's pretty funny. Basically, what we've got here is, I think people, some people actually genuinely believe this is real. So this is an announcement, which you can see right here, of the Le Chaton Fat, a French model from Mistral, with a 30 trillion parameter frontier model built for long horizon reasoning, coding and agentic work. Now, this is basically a big joke that's totally fake, but you might have seen it. And basically, it's scored higher than Claude Fable 5 on every single test. It has 30 trillion parameters, one million token memory, runs faster than anything Mistral has ever built. And here's the part that matters for you. None of it is real, not one number, but it's going viral across the web right now. And thousands of people have been sharing this test. So if you ever looked at a chart showing one AI beating another, I thought, well, the numbers don't lie. I've got some bad news for you. Sometimes the numbers lie, and this French cat just proved it to the world.
So basically, this is kind of like a joke based on all the Fusion 5 intelligent tests that are coming out, not Fusion 5, Claude Fable 5 So there's a few things that have come out where basically people have said, this is how you can achieve Fable 5 level intelligence, even though Fable 5 has been shut down. So for example, there was a Fusion model that came out recently, and there was also the research from Kilo Code and a bunch of other stuff. Basically talking about how you can reach Fable 5 level intelligence with different methods. Now, it's a pretty funny joke, but basically this is one thing you need to pay attention to, which is like, you know, a lot of AI companies are releasing their own sort of benchmarks and tests and that sort of thing. I think this is a great example of like how you can't rely on the benchmarks. Like for me, for me, I just test everything myself.
And what happened here is like basically a joke chart went viral, people shared it's real, and it gets debunked when it's too late. So you can see another example right here. And obviously like some people are joking here, some people are not. It's kind of hard to tell the difference. Someone actually posted it on a recent video that I did to check this out. But it just totally goes off the charts.
So you might think, okay, why does this matter to you?
Why talk about this French cat at all? Because this is something that I see like day to day. It's like a lot of people look at benchmarks, but benchmarks are just tests. So you can give a bunch of AI models the same set of questions, you score them and the one with the highest wins, right, which is simple. But the problem is like a lot of the time with these benchmarks that are coming out all the time, the company that made the AI runs the tests on its own model, they pick which test, they pick which other models to compare against, they pick which numbers to show you. And it's kind of like someone just grading their own homework and then bragging about the A they scored. Now, you know, with stuff like Fusion or the Kilo Code Research, I think there's a lot of truth behind it, like it's worth checking out, but it's just something to be aware of. It's like sometimes you're going to see benchmarks like this. They look official. There's no referee. They're cherry pick tests and they take like five minutes to work as well. And I think the Le Chaton fat joke took this to the extreme because someone just typed numbers into a chart. There was no model, no test. It was just kind of like a joke that a lot of smart people actually fell for from what I saw. So for example, like the 30 trillion parameter model, what they've just made, etc. You can see some of the tweets that are coming out from it. But I think there's a story in this, which is just like pay attention to what you're learning from, test all this stuff out yourself. Don't believe everything or the hype that you see out there. And for me, like you'll see inside every tutorial that I test things myself and make sure that it's actually working and that it's actually reliable. Because if you don't, it's very easy to believe these benchmarks from fancy AI companies and not really know if it's legit or not. So how should you do this? Well, I think your own work is the only benchmark, right? Real email, real blog post, real customer question, because that's the stuff that you're doing day to day with AI automation. So for example, a real task like could be an email, blog post, etc. Run the tool on it, judge the output yourself. Is it good? One thing that we never really look at with benchmarks is, did it save you time? Would you actually use this day to day? That's something to be aware of. And would you keep it or drop it? So for example, if you look at the agent operating system that we have, we actually use this for real workflows. If you look at this system here, we deploy SEO content to our websites using this tool. We just plug in a keyword in a case study. We use it every day. Now, for me, this isn't about benchmarks. It's just about making sure that we automate the actual processes that we have. The other thing I would say here is like, there's so much going on, there's so much noise with AI and the new releases that come out, that it's very difficult to actually research each of the new updates and validate if they're real or not. So the way that I look at this, because you might be feeling overwhelmed too, the way that I look at this is like, okay, what's one thing that you can focus on this week and automate? You don't want to be automating everything, you don't want to be trying every single model. What you want to do is just focus on one thing, build on that. Because if you do that every week, that's like 50 new automations you create per year. And it doesn't matter about the benchmarks or the models or what comes out, what gets taken down, you actually build something useful and you get the most out of this stuff. And it doesn't really matter about the benchmarks because all of these models can achieve it, right? Like, for example, even if we were using KimiCode or GLM 5.2, if we were using Grog Build, there's no wrong answers here. There are some models that are better than others.
5 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000773033579