**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. I am freshly back from two weeks in China, of which was a really incredible experience, and the feed has been quiet for about two weeks. I think that's the longest that the feed has been quiet since I started the show almost three and a half years ago now.
So I apologize for the break. I didn't intend for the break to be quite as long as it was. We had one episode that was prepared and ready to go while I was away, and we had a little issue with that one that has it delayed and perhaps delayed permanently. We'll find out. And I was actually planning to do a little bit of recording from China, but I was so busy over there just doing so many things that I didn't have a ton of time to really get my thoughts together and do that. And I was also advised that because I was there on a business visa as opposed to a journalism visa that it might be easier for me and for everyone if I just waited to publish stuff until I got home. So in the end, it was a two-week quiet period. I kind of had the expectation that nobody would really much notice or care. And I think that's still a pretty good way for me to approach my work on this podcast because honestly, it's been really effective for me to just chase my own learning and share it in a pretty unstrategic way. And I'm really grateful for how well received that has been over time. But I did get at least one note from somebody asking if I was okay. And I guess that's a testament to how consistent we've been in publishing and how at least somebody noticed that I was gone. So I doubt that the ratio that congresspeople usually apply, which I understand is one letter, represents 100 votes. Something tells me that doesn't really apply here. And I highly doubt that there were 100 people who were worried about me, but to the degree that anyone was worried, I'm sorry for leaving off without explaining the absence for the two weeks that I was gone. But excited to be back and excited to share the results of my, or not the results, but the experience and the takeaways from my China trip with you all. One of the thing I wanted to touch on real quick before diving into it is, I introduced the last episode with Davodad, which got some very nice comments as well, with an introduction from Fable itself.
Due to a little snafu in our production process, an intro to the intro which I had recorded to explain what the heck you were about to see was not actually included in the final version of the episode, and I was already on the plane and not able to review and catch that missing piece before it went live.
So we just launched into it directly with Fable. A couple of people had said, what is this nightmare fuel? Why does it sound like this? And the answer which I had meant to give you the context on before you heard it, was that I allowed Fable to do its own thing in this case. I had it write the intro, I had it use 11 Labs voice design to design its own voice, and then I had it use the LTX model to create its own video based on the audio that it had already generated with 11 Labs. I was rushing to get out the door and get on this China trip, so I let Fable do its thing, and I didn't really check to see if the output super faithfully represented Fable's intent. And one person asked, did it really want to be heard that way, or does it understand what it sounds like to us when it speaks in that voice?
And I decided, I should go back and check what Fable actually intended to do, and if it succeeded, or if it somehow got confused along the way.
And I honestly am still a little confused by that. I went back and opened up, found that same thread, and asked what prompts it had used for the voice design, and what its intent was. And I would say its intent was not to be as sort of creepy and alien as the voice ended up sounding, and I agree with those who said it did sound a little kind of unsettling. The voice was intended to be more of a sort of warm and kind of reassuring voice. It actually tried out three different voice designs. One it called The Luminous Narrator, which it describes as an androgynous mid-register measured quiet wonder. The other one it tried was The Young Scholar, which it describes as bright, earnest, crisp, a brilliant grad student presenting work they love. And then the third one was The Low Ember, a late night radio, old storyteller style. And it maybe went a little bit wrong by first generating those voices, and then asking Gemini, which can take the multimodal input to evaluate the voices, and evaluate them according to its intent. I don't know if Gemini really can do that. I'm not sure sometimes when I put files into Gemini, if it's actually interacting with them as raw audio in a way where it could even understand or even process the tone of voice.
123 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000778572833