Tencent HY3: NEW Chinese AI (FREE + Open Source!) artwork

Tencent HY3: NEW Chinese AI (FREE + Open Source!)

AI News Today | Julian Goldie Podcast

July 7, 2026

Testing Tencent HY3 (Free on OpenRouter): Benchmarks, Agent Tasks & HY3 Coder DemoThe video tests Tencent HY3, a new open-source model available for free via OpenRouter until July 21, and shows it integrated into Agent OS plus availability in Hermes Agent, News Research, and Kilo Code.
Speakers: Julian Goldie
**Julian Goldie** (0:00)
Today, we're going to be testing out Tencent He, which is a brand new open source project that's available on a free API. And we've been building all sorts of stuff with this morning. I'll show you how it performs, how it goes, et cetera. We've already built it into our Agent OS, and you can see an example of it working right here. Now, this is open source. It's just dropped from China. I'll show you the benchmarks in a minute. If you want to use it, basically you can get it on the API for free with OpenRouter. It's available on OpenRouter for free, as you can see right here.
If you just type in HY33, you will see it over here, and then you can get the API and plug in it.
That's pretty cool. It's available until July the 21st on the free API. You can also get it on News Research with Hermes Agent for free, and it's also available on Kilo Code for free as well.
Open source, free, what's not to like? Well, I'll show you some examples of what I do like and what I don't like. So you see actually here like the website that it created is not bad, not bad at all. We said build out a beautiful landing page. It went off and built it and we can see the preview and we can see our saved creations over here. So not bad, right? Not a bad stuff. However, let me show you what do and don't like about it. So you can see an example of what we built here. This is kind of like an open world game and everything, every new model that drops, we test it out on Goldie Bench and then we see how it performs and we compare it to everything else. Now, if you have a look at this game, it's not bad. It just doesn't feel quite right. There's something about it. I would say that's probably one of the better creations from H13, but you see how that light over there, it just looks a bit weird, right? It's just a bit glitchy. This open world game is not bad, but it's not really playable. It's not really fun to use. Let me show you another example.
You can see here, it's like a Lego man that's been created on this example, and it's an absolute giant compared to everything else. But if you look at the graphics and the details of this, it still feels very basic when we're using it. That's something that I've actually found across pretty much everything that we tested it with. You can see another example here. We can drive around a city. I would say this is probably the best creation that we got so far. It looks quite cool. It's quite fun to play. But if you compare that side by side versus something frontier, like Fable, or you compare it versus like Glm 5.2, which we'll do in a second, I would say it's not really on the same level. Like even look at the city blocks here. They're just super basic and that sort of thing. So it's awesome that it's free, awesome that it's open source. Is it the best model or one of the best models I've ever used? No. Would I say it's right up there? Probably not. Like see how it's super glitchy in terms of using this and then we fly in the air. There's no detail. There's nothing going on as well. I'll show you the final example here. So that is for this parachute game, which is quite fun, like fun background and that sort of thing. If we release a parachute, that's what it looks like.
It's just not quite on the level that I would hope. And we did do a lot of back and forth on this as well. So if we actually have a look at our HY3 tests, we've been coding and testing this out. Let's have a look here for like two or three hours. You can see even like two hours ago, we were testing the flight simulator and it still wasn't quite right. So we've been going back and forth and wrestling with it for like three to four hours and still not getting the outputs, even when we give a lot of feedback to Fable 5 to test it out more. Now, if you're wondering, okay, how does it perform on benchmarks? Let's pull them up. So here are the benchmarks. And you can see this is versus other models like Claude, you got Glm 5.2 and honestly from what I'm testing, you can see on MCP Atlas, BrowseComp, et cetera, Claude, Eval is kind of, well, it's beating Glm 5.2 on Claude, Eval which is like an agent workflow. When I've tested out myself, it doesn't seem to feel like that at all. Doesn't seem to be on the same level. So let me show you an example of what we've built with our models using the same system. So this is the test, as you can see, if you compare that to other models, you see the quality of that versus, and this is with Fusion, the quality of this is way, way better versus HY3. So there's a lot more detail on most of the models that we tested out with the same sort of prompts. Let's have a look on the driving game as well. So this is with Fusion, and you can see like the quality is so much better than what we saw with HY3 before.

5 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000775828933