Spoken vs TranscriptFetch for podcast transcripts

TranscriptFetch generates a transcript from the episode audio. Spoken returns the transcript that was already published with the episode, with real speaker names attached.

Last updated August 2026

These look like the same product and aren't. Both take a podcast link and hand back text through an API an agent can call. The difference is where the text comes from: TranscriptFetch resolves the link to the episode audio and transcribes it, so you get a fresh machine transcription. Spoken returns the transcript that already exists for that episode, and resolves the voices to real names — "Andrew Huberman", not "Speaker 1". Neither is strictly better. They fail in different places, and the right one depends on whether you need coverage or fidelity.

The honest version on price

TranscriptFetch is cheaper per transcript than Spoken, and not by a rounding error — their plans start at $5/month and their effective per-transcript rate sits well below anything on Spoken's ladder. If cost per transcript is your deciding variable, they win that comparison and this page won't argue otherwise.

What the difference buys is covered below. Two things worth weighing against the sticker price:

Side-by-side comparison

TranscriptFetch Spoken
Transcript source Generated from the audio The transcript published with the episode
Speaker names No speaker identification Real names, resolved per turn
Coverage Anything with audio behind the link Episodes that have a published transcript
Billing model Monthly subscription One-time credit packs
Credit expiry Tied to the subscription period Never expire
Re-fetching an episode Costs a credit each time Free, always
Output format JSON segments Markdown, speakers bold, timestamps per turn
Sizing an archive first Not offered as a separate step Free catalogue listing before you spend

Where TranscriptFetch wins

Transcribing the audio means never being told no. If an episode has audio, they can return text for it — including shows that publish no transcript at all. Spoken returns a 404 for those, and no amount of retrying changes that.

So if your requirement is every episode of an arbitrary list of shows, breadth is the thing you need and Spoken cannot promise it. They also accept a wider range of input links, and at high monthly volume their per-transcript economics are hard to argue with.

Where Spoken wins

Machine transcription and a published transcript are not the same artifact. The published one carries the spelling of names, products, companies and jargon as the people who made the show wrote them. A speech model guesses at those, and it guesses worst exactly where a technical podcast is most valuable — on the proper nouns you are most likely to be searching for later.

Pick the right tool

Pick TranscriptFetch for

Pick Spoken for

FAQ

Is TranscriptFetch cheaper than Spoken?

Per transcript, yes. Their plans start at $5/month and the effective rate per transcript is below Spoken's. Spoken's answer is not price — it is that the transcript is the published one rather than a machine transcription, that speakers are named, and that credits never expire and re-fetches are free.

Does TranscriptFetch identify speakers?

No. There is no speaker-identification feature in their product, so a multi-speaker episode comes back without attribution. Spoken resolves speakers to real names as part of the response.

Which one has better coverage?

TranscriptFetch, clearly. They transcribe the audio, so they can return something for almost any episode. Spoken only serves episodes that have a published transcript and returns a 404 otherwise — never charging for it.

Can I use both?

Yes, and for a broad archive it is the sensible pattern: Spoken for shows that publish transcripts, where you want named speakers and exact spelling, and a transcription service for the rest. They fail in different places, so the union covers more than either alone.

How do I know which episodes Spoken can actually return?

Ask, for free. Listing a show's back catalogue costs no credits and returns only the episodes that are fetchable, so you can size and price an archive before spending anything.

TL;DR: TranscriptFetch is cheaper and covers more, because it transcribes audio. Spoken returns the published transcript with real speaker names, credits that never expire, and free re-fetches. Pick on fidelity vs coverage, not on price.

"I used Spoken to add every My First Million episode to my knowledge base, with a cron to pull new ones. Now I can enjoy the podcast on a run, then chat with Claude about it later — every episode saved and accessible in my Claude sessions."

— Marcus Taylor

Thousands of transcripts fetched by developers building summarizers, RAG pipelines, and podcast tools

No signup required — use API key pt_demo on any endpoint.

Price my show's archive

$0.10 per transcript. Credits never expire.