**Julian Goldie** (0:00)
Open up Google.
**SPEAKER_2** (0:03)
Opening Google now, sir.
**Julian Goldie** (0:06)
Tell me a joke.
**SPEAKER_2** (0:08)
Very well.
Why did the computer go to the doctor? Because it had a virus.
**Julian Goldie** (0:15)
That's not very funny. Give me a funnier one.
**SPEAKER_2** (0:18)
As you wish, sir. How about this?
Why don't scientists trust atoms? Because they make up everything.
**Julian Goldie** (0:28)
Tell me some words in Chinese.
**SPEAKER_2** (0:31)
Certainly, sir.
**Julian Goldie** (0:32)
你好, 希望你今天過得很好。
Stop, stop, stop. Actually, could you start speaking Swahili with me?
**SPEAKER_2** (0:40)
Certainly, sir.
Let's switch to Swahili. Wukotayari kuwanzamazungumzo?
**Julian Goldie** (0:47)
Could you open up Obsidian?
**SPEAKER_2** (0:50)
Obsidian is now open, sir.
Good afternoon, sir. Nine open items and nine notes in motion, sir. First up, Vault LLM Wiki Restructure. Finalize atomic note layer. In the news, 60% of US consumers say AI and brand messaging is a turn-off.
**Julian Goldie** (1:17)
All right, nice, nice. One sec, I'm going to send you something in the chat, okay?
**SPEAKER_2** (1:21)
Understood. Sir, I'll be ready to take a look once you send it.
**Julian Goldie** (1:27)
Can you see that?
**SPEAKER_2** (1:28)
I've opened juliangoldie.com for indeed, sir. I can see the website layout on your screen. Everything appears to be loading smoothly.
**Julian Goldie** (1:40)
So I'm going to switch off, but that's basically Hermes Jarvis, which is blowing my mind. This is a voice activated version of Hermes that we've built, and it responds in real time. It's fully working, seems to work really well. The only thing I don't like is it doesn't have access to Obsidian in real time, which is something I'm actually building on it, so I just want to be 100% transparent with you. But it can do daily briefings like you can see. It knows the day, it knows the date, it knows what's open, what we've worked on recently. All of these are different memories inside our Obsidian vault. So it actually links directly to my memory folder. And also what you can see here is that it has a list of what's on my mind, recent headlines in the news, which is pretty cool. And it actually stores these as history. So if we go to the history, we can see our previous briefings here. We can click on them. We can view previous briefings over the last few weeks, and we can see our history. Now not just that, but if we actually scroll down here, you can see our conversation history. And then if we actually scroll to this section, you can see everything that was built with Hermes inside the voice mode, whereas it's even built, for example, like a meditation app and all sorts of cool stuff.
So my point here is like, this is a real time voice activated version of Hermes. And what we can do is we can actually have a wake word or we can set up in real time. Now the way that I built this is basically by iterating, honestly for weeks at this point, just little bits at a time. If you watch my live streams on my other channel, you might have seen this where like I would say about four hours a day, I will improve something on the Agent OS system here. And this morning it was actually the Hermes Jarvis version of this. Now, what we've also got is this section where we can do weekly briefings. We can switch the voice as well. So we can actually change between all the voices that we have.
One thing that you learn from this, if you want to build something like this yourself, is that you want to make sure you have ChatGPT real time as the API. We had ElevenLabs before. It was way too slow. It sounded better, but it was too slow to actually use. And the other thing that blows my mind is like we can send links inside the chat and Jarvis can read them and then come back to us. Plus we have wall mode where we can basically have this real time, separate monitor, voice activated, ready to go whenever we need it. So it's very, very fast, very responsive, responds literally instantly, opens stuff up literally instantly.
I don't know how it's been created this well, but basically what I've been doing is using Claude and Claude Desktop, as you can see right here, and just going back and forth with it until we actually build something nice. And we've been going back and forth for weeks as you see. If you look at the history here, it's like it goes on forever because it takes so much iteration. But I do want to show you this just to inspire you because I think it's fantastic what we've built. And we used to have this talk section here, but I think I'm actually going to get rid of it because now Hermes Jarvis totally replaces it. And if I have a choice between going to Hermes Jarvis or going inside the chat, for sure I'm going to go directly to Hermes Jarvis these days because it's just easier to use, easier to manage, easier to organize. And the other cool thing about this, you could set this up with anything. You know, it doesn't have to be Hermes. It could be Claude that you speak to. It could be OpenClaude, it could be Gemini. But I just want to show you that you can actually build something useful with Hermes. It's not just like a gimmick. It's like a genuinely useful voice agent that is getting closer and closer to something that I actually really love.
3 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000773797060