**Julian Goldie** (0:00)
So today, I'm going to show you a powerful new video system that we've built into our agent operating system. We call it the Infinite Video Engine. And basically what this can do is you can see right here is it can research a topic, it can create the AI avatar, it can create all the B-roll, edit it together beautifully, and it can research any topic from the web. And all you do is you give it a topic to create the video around. Let me show you an example in action. So we've got this video that we created earlier. Let me play this.
China just dropped an open source coding model that's six times cheaper than Gpt 5.5. It's called GLM 5.2, 753 billion parameters and a 1 million token context window. It nearly ties Claude Opus on long coding tasks. On SWE-Bench Pro, it scored 62, beating Gpt 5.5 and the weights are free under an MIT license. Download it, run it yourself, pay only for compute. If you build with AI, this changes the math. Subscribe. So this is a super powerful system we've created. And basically what we can do is we can type a topic, it researches it live, writes a script, speaks in a voiceover, puts an AI avatar on the camera, finds a B roll and puts the whole thing into a beautifully edited video, as you can see.
Let me show you the video agent that we've set up over here. Really, I just want to show you what's possible, how powerful this stuff is when you put it together. So if we take a look here, for example, what we can do is we can plug in a topic, like for example, let's say Fusion came out. That's pretty popular. We can put in here, Fusion versus Fable 5, the new research from Open Router. So we can now click, write the script. We select presenter, who's going to do the B roll. So we put Minimax, the avatar that we're going to use and the voiceover that we're going to use. Now that's going to start researching the script. So it puts it all together step by step. So you've got like a whole director from start to finish.
It's good to go. You see a couple of other examples right here of what we've built and how we've built it using this system. And so before, like if you're making videos manually, it can take up a whole morning.
You might need a whole team, like a team to script it, a team to write it, a team to get the voiceover, a team to actually film it, and then also an editor to put it all together. So a lot of manual work as a team. Now what you can do is you can just plug this topic in, it will go off, start researching the web, and then we are good to go on that. So the amazing thing about this is we can type one line, like GLM 5.2 release, and it goes off and it creates that whole video we talked about. You can see that it actually plans out scene by scene with the script here. Not just that, but if we pull this up properly, what we've actually prompted it to do is come up with the B-roll description for each scene inside the video itself. So you can see the voiceover script here, and then you can see the B-roll for that scene as well, which is pretty amazing in itself. And so if we look at this system, what we can then do from here is we can look at the research notes. So you can see that it's connected to the web, and it gives us a date of when OpenRooter Fusion was fully created, when it was researched, and then it breaks down exactly how that works. And if we click Generate Avatar and B-roll, that's going to start creating and putting it all together, as you can see, right? And after that, we can assemble it all and get it all edited. So we just have to wait for that to be done, and then it gets to work, which is amazing. This is incredibly powerful. And so, you know, if you're watching this, this is something anyone can start using today. I really just show it to inspire you and how powerful it is. But this is a very, very, very powerful system that I've never seen before, that we just put together to create something amazing. So if we have a look at this as well, and we look at, okay, how does it work step by step? So you type one topic, it's then going to script it, it's then going to use a voice, the face, and it's going to have one finished edited video together. So one topic goes in the top and then it drops through five stages and the video comes out the other side. It's absolutely insane. And then the great thing about this is it's always current because it's researching your topic on the web live first. So the video has real research from today, not just like the AI hallucinating, which you can do sometimes. So you sound like you actually know what you're talking about inside the video, which is great.
8 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000773481994