Hermes Jarvis: NEW Voice Activated AI Agent! artwork

Hermes Jarvis: NEW Voice Activated AI Agent!

AI News Today | Julian Goldie Podcast

June 23, 2026

The script demos “Hermes Jarvis,” a fast, voice-activated Jarvis-style assistant that can open apps like Obsidian and Google, respond in real time, be stopped mid-response, teach basics like Japanese greetings, and discuss building a snake game.
Speakers: Julian Goldie
**Julian Goldie** (0:00)
Hermes Jarvis is blowing my mind. This is a voice activated version of Jarvis, so it is really, really fast. Let's just open this up.
Hey, you there, Jarvis, mate. How's it going?

**SPEAKER_2** (0:10)
Very well indeed, sir. Ready to assist you with anything you need today.

**Julian Goldie** (0:14)
Open up Obsidian for me.

**SPEAKER_2** (0:18)
Opening Obsidian now, sir.

**Julian Goldie** (0:20)
Open up Google for me.

**SPEAKER_2** (0:23)
Google is up and running, sir.

**Julian Goldie** (0:25)
What do you think we should work on today?

**SPEAKER_2** (0:28)
Might I suggest refining the next YouTube title formula? Stop.
Understood, sir. We'll leave it there for now.

**Julian Goldie** (0:35)
Carry on.

**SPEAKER_2** (0:37)
Very well, sir. Perhaps we could refine-
Of course, sir. Standing by.

**Julian Goldie** (0:43)
Teach me some Japanese.

**SPEAKER_2** (0:45)
Certainly, sir. Let's start with something simple. Konnichiwa means good afternoon or hello.
Certainly. No trouble at all, sir. Ready whenever you are.

**Julian Goldie** (0:58)
Can you build a snake game for me?

**SPEAKER_2** (1:01)
A snake game, sir. An intriguing little challenge. I can certainly help guide you through that or draft some code. Where shall we begin?

**Julian Goldie** (1:09)
Can you build out a snake game in real time for me, please, mate?

**SPEAKER_2** (1:13)
I can certainly guide you through building it step by step, sir. Let's start with the basics.

**Julian Goldie** (1:19)
Actually, stop there.

**SPEAKER_2** (1:21)
Understood, sir. We'll hold off for now.

**Julian Goldie** (1:25)
Give me a daily briefing.

**SPEAKER_2** (1:27)
Certainly, sir. Today is the 20th of June, 2026 Your YouTube channel's views are strong, and we've got the open-claw ecosystem at your fingertips. Anything particular you'd like to focus on today?
Good morning, sir. The day is light on meetings, but rich in housekeeping. The Vault LLM Wiki restructure and your Atomic Note Layer await, alongside a fleet of SEO keyword sets you've been circling. Before any video goes out, the transcripts need attention. Clawed AI, Anthropic and Clawed code have all been mangled, and your export is still stamped with the Buddhist calendar, which rather upsets date sorting.
On the wires, 60% of US consumers now find AI and brand messaging a turnoff. Worth bearing in mind before your next headline, sir.

**Julian Goldie** (2:21)
So I'm just going to switch off real time there, but you can basically see that we can talk to Hermes agent, we can get it to open stuff up.
It's the fastest I've ever seen computer use, and this is something that we've just built in and improved and worked on over and over again. I just show you this to inspire you and also to show you like having a really powerful voice agent like you can see, it's really, really cool. You can see if we scroll down here, number one, we've got all the conversation history of everything we've done. Number two, it responds in real time. I can stop it anytime, and I can also ask it to open stuff up. If we say, for example, open up Google, it just opens up really, really quickly. Then you can see everything that we've built over here as well. We've got all this stuff that we've built with Hermes Jarvis, and we can see that we've created stuff with voice activation like so, which is pretty wild. We can also open that stuff up later too. If we click on this, for example, we can open up just as a preview and then we can open up and fall as well. The cool thing about this is usually when you're using voice agents, it's pretty hard to find the history of what you've created and come back to it. Whereas inside this agent operating system that we've built, you can just get Hermes Jarvis ready to go. We've also tried to give it that Butler vibe as well when we're speaking to it. We used Chatchibity real-time, which seems to be the fastest for responding to us. The great thing is we can switch on real-time on or off. We can have that inside wall mode as well. And also before I show you wall mode, look at this. So what we can see here is it basically has a look at my Obsidian memory. It looks at the suggested focus, what's open and to do, the notes we've recently touched, what we've recently worked on. It tells me it knows like the day and everything like that. Then also it's got these open action items. Now if we click on one of those, we can actually go straight to Obsidian to see the open action items. So it's all plugged into my memory and my context. If you're wondering, how do we set up, we have a memory galaxy here. Now the great thing about this is, number one, it's got a lot of context to meet. That's one of the other issues with voice is like, you either have a voice agent that doesn't know a lot about you, or you have a voice agent that knows a lot about you, but it's really, really slow. We've tried to create a really fast version of that. This is by far the fastest I've ever seen. It's literally real time. Then any conversations that we have with Hermes as well, they get logged inside our Obsidian memory. So you can see those right here, three minutes ago, 23 minutes ago. I don't have to touch those, they just get automatically logged, which is great because then the next time we use Hermes, it's going to be smarter, it's going to have more context, it's going to understand us. And also you can see here, it uses a combination that understands like inside our memory system.

7 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000773812743