Do people regret buying the DJI Pocket 3?
What mic does MKBHD use in this video?
Hosted YouTube MCP

YouTube MCP.

Ask YouTube anything.

Which n8n tutorial should I start with?
Did anyone in the comments get this working?
Setup

Connect it to your client

Any MCP client connects the same way: point it at the server URL and send a TranscriptFetch API key as a bearer header (on claude.ai, OAuth replaces the key).

Client-by-client walkthroughs
Server URL
https://transcriptfetch.com/mcp
Any stdio client (mcp-remote bridge)
npx mcp-remote https://transcriptfetch.com/mcp \
  --header "Authorization: Bearer tf_live_YOUR_KEY"
Tools

The seven tools, one by one

What the server actually exposes to your client. Your assistant reads these same descriptions and picks the tool itself; this is what it will find.

get_transcript

get_transcript(video, ai_fallback?)
watch · youtu.be · ID

The workhorse. Takes a watch URL, a youtu.be or Shorts link, or a bare 11-character ID and returns the transcript as timestamped segments. Captions come back immediately; when a video has none, AI Fallback Transcription kicks in and short videos are returned in the same call, typically in ~30 seconds. Only long media becomes a background job the assistant collects on its next call. When the response suggests it, the assistant retries with ai_fallback: true to start that transcription explicitly.

search_videos

search_videos(query, upload_date?, duration?, sort?, captions?)
filters

Keyword search over YouTube, returning titles, durations and IDs your assistant can feed straight into get_transcript. upload_date keeps the last hour, day, week, month or year; duration keeps short, medium or long videos; captions: true keeps only videos with subtitles; sort: views puts the most watched first. This is the entry point when the user has a question rather than a link: find the three most relevant videos, read them, answer with quotes.

search_channels

search_channels(query, type?)
channels · playlists

The same search for channels, or with type: playlist, for playlists. A channel row carries its handle, subscriber count and description, and every URL goes straight into list_channel_videos, list_channel_playlists or list_playlist_videos. The tool for when the user names a creator or a topic rather than a video: find the channel first, then walk it.

list_channel_videos

list_channel_videos(channel, tab?, sort?, query?)
videos · Shorts · live

Turns a channel handle or URL into its videos, up to 50 per call: the latest uploads by default, the Shorts or live tab with tab, the most popular or oldest first with sort, or only the videos that match query. The tool that makes "what does this creator say about X" answerable: an agent walks the list, transcribes what looks relevant, and cites episodes by name. Each successful call costs 1 credit, listings included; only get_credits is free.

list_channel_playlists

list_channel_playlists(channel, tab?)
playlists · podcasts

A channel's playlists, or with tab: podcasts, its podcast shows. Series, seasons and courses usually live there, so this is how an assistant finds the one playlist that matters before listing its videos.

list_playlist_videos

list_playlist_videos(playlist)
playlist

The same enumeration for a playlist URL, which is how courses, conference tracks and lecture series are usually organized on YouTube. Point your assistant at a 40-lecture playlist and ask for the syllabus nobody wrote.

get_credits

get_credits()
free

Reports the key's remaining balance, free. Useful in long agent runs: a careful agent checks before batching forty transcripts, and a budget-capped workflow can stop cleanly instead of hitting a 402 mid-task.

Background

Why YouTube is the easy one

Most YouTube uploads already carry a caption track, auto-generated if the creator never wrote one, and often in several languages. So a YouTube fetch is usually a read, not a transcription job: fast, and billed as a single credit like everything else. AI Fallback Transcription only has to step in for the minority of videos where captions are disabled or missing.

No caption trackTranscribed
Captions disabled or missing: AI Fallback Transcription transcribes the audio, billed by length on delivery.
Caption trackRead
Most uploads, often in several languages. A read, not a transcription job: fast, one credit.
What still fails
  • Live streams while they are still live.
  • Age-restricted or private videos.
  • Clips with no speech.
Plan for scale

An hour of talking is roughly 10,000 words, and a channel can hold hundreds of hours, which is more than an assistant wants in context at once. The timestamped segments are what make that workable: an agent can skim, quote the exact moment, and point back to 41:32 instead of paraphrasing an hour.

Setup

Setup by client

Same server, same key; only the config file differs. The three we see most, plus the escape hatch for everything else.

Claude

On claude.ai, add a custom connector with the server URL and OAuth does the rest, no key pasting. In Claude Code and the desktop app, one config block with a bearer header.

Set it up →

Cursor

One entry in mcp.json and your editor's agent can pull a video's transcript into context while you code, which turns conference talks and API walkthroughs into referenceable text.

Set it up →

Codex

Registered through config.toml. The CLI agent fetches transcripts mid-run, so a script that researches, summarizes and writes can include YouTube sources without custom glue.

Set it up →

Everything else

Any MCP client works: HTTP-native ones take the URL and bearer header directly, stdio-only ones go through the mcp-remote bridge shown above.

Configuration reference →
FAQ

Frequently asked questions

Which platforms do the discovery tools cover?

search_videos and list_channel_videos cover YouTube, TikTok, and Instagram. list_playlist_videos covers YouTube and TikTok; Instagram profiles use the channel tool. search_channels, list_channel_playlists, the search filters and the channel tabs work on YouTube alone. Transcript fetching also accepts direct media file URLs.

Does it work on videos without captions?

Yes. When no caption track exists, AI Fallback Transcription kicks in - short videos are returned in the same call, typically in ~30 seconds, and only long media becomes a background job your assistant collects with its next get_transcript call. When the response suggests it, retrying with ai_fallback: true starts that transcription explicitly.

Can my assistant read an entire channel?

Yes: list_channel_videos returns the channel's videos, then get_transcript fetches each one. Each successful call costs 1 credit, the listing calls included; only get_credits is free.

What does it cost?

50 free credits a month. Caption fetches cost one credit; AI Fallback Transcription is billed by audio duration, only on delivery. Failed requests are never billed. Paid plans start at $5 a month for 1,000 credits.

Is this affiliated with YouTube or Google?

No. TranscriptFetch is an independent transcript API; the MCP server exposes it over Anthropic's open Model Context Protocol and is listed in the official MCP Registry as com.transcriptfetch/youtube-transcripts.

Your AI, with YouTube's words in reach

Sign up, grab a key, paste one config block into your client. 50 free credits a month, failures never billed.

Get started free