Reference

YouTube Transcript API

One POST turns a YouTube URL or video id into a transcript: clean text or timestamped segments, in one JSON shape. The caption track is read first, manual or auto-generated, and a video without one is transcribed from its audio on the same request, so the call answers either way. Nothing here needs a YouTube Data API key, a quota or OAuth.

1 credit per caption transcript, audio billed by durationView pricing

A video without captions is transcribed from its audio and charged only on delivery: 1 credit per started 5 minutes of audio, minimum 1. Failed requests are free.

On this platform

  • Accepts watch, youtu.be, Shorts, embed and live URLs, or the bare 11-character video id.
  • Videos under 20 minutes are transcribed inline when captions are missing; longer ones answer 202 with a job to poll, or POST to callback_url.
  • A caption transcript costs 1 credit; audio transcription is priced by duration and only charged on delivery.
POST/api/v2/transcripts/videoidempotent

Fetch a transcript (YouTube, TikTok, Instagram, or file URL)

Returns a transcript - text plus timestamped segments. Accepts YouTube, TikTok, and Instagram URLs (or a bare TikTok video id), direct media file URLs. ..) identifying who is talking. Speaker ids are hints from voice separation, not named identification. When no captions exist the audio is transcribed automatically: when we can determine the media length, media under 20 minutes simply waits (the request is held open for up to 45 seconds) and returns the finished transcript, so no polling is needed; when the length cannot be determined, only short-form platforms (TikTok and Instagram) are held inline. Longer media, or a transcription still running when the 45-second hold expires, returns 202 with a job to poll instead - the work continues either way, so the same request is safe to retry and will hit the cache once it finishes. Supply callback_url to have the finished transcript POSTed to you instead of polling. Every failure carries an ai_fallback block saying whether captions were definitively unavailable and whether retrying would work.

Body parameters

videostringrequired
A video URL or 11-character YouTube video ID. Accepts YouTube (watch, youtu.be, /shorts/), TikTok, and Instagram URLs, plus direct media file URLs (mp4/mp3/wav/…). Videos without captions fall back to AI transcription.
Example
mode"captions" | "audio" | "auto"optional
Where the text may come from. "captions" reads an existing caption track and fails if there is none, which is the only way to avoid audio transcription. "audio" skips captions and transcribes the audio. "auto" (the default) tries captions first and transcribes the audio when there are none. Short media may finish inline after a wait of up to 45 seconds; longer work returns 202 with a job to poll or deliver by callback. Audio is charged only on delivery: 1 credit per started 5 minutes of audio, minimum 1 (a 20-minute video is 4; the 4-hour cap is 48). Caption fetches are always 1 credit.
Example
timestampsbooleanoptional
Which form the transcript comes back in. true (the default) returns the `segments` array, each with start, duration and text. false returns a single joined `text` string instead. On the single-transcript 200 exactly one of the two is present, never both, since segments already contain every word the joined text does; job results and batch entries carry both text and segments. The older strings "segment" and "none" mean the same two things and are still accepted.
Example
callback_urlstring (https URL)optional
Where to POST the finished transcript when a request escalates to audio transcription, instead of polling the job. The delivery body is a trimmed envelope - {ok, status, job_id, data} on success, {ok, status, job_id, error} on failure - without the usage and request_id the poll URL adds. When a signing secret is configured on our side, the body is signed with HMAC-SHA256 over the exact bytes and sent as an X-TranscriptFetch-Signature: sha256=<hex> header so you can verify it came from us. Must be a public https URL on the standard port; the URL is checked again at delivery time, so an unreachable or private address still gets a 202 but never receives a delivery. The job stays pollable either way, so a missed delivery is never a lost transcript.
ai_fallbackbooleanoptional
Legacy alias for "mode", still supported. true is identical to "mode": "audio"; omitted or false is "mode": "auto". Send one or the other, not both. Note that false never disabled the fallback - audio was still transcribed when no captions existed - which is why the field was replaced.
Example

Request example

curl https://transcriptfetch.com/api/v2/transcripts/video \
  -H "Authorization: Bearer $TRANSCRIPTFETCH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"video":"dQw4w9WgXcQ"}'

Responses

SuccessExample response envelope
{
  "ok": true,
  "request_id": "req_…",
  "data": {
    "kind": "transcript",
    "video_id": "dQw4w9WgXcQ",
    "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "platform": "youtube",
    "title": "Example video",
    "channel": "Example Channel",
    "duration": 212,
    "language": "en",
    "thumbnail_url": "https://i.ytimg.com/vi/dQw4w9WgXcQ/mqdefault.jpg",
    "source": "captions",
    "segments": [
      {
        "start": 0,
        "duration": 3.5,
        "text": "We're no strangers to love"
      }
    ]
  },
  "usage": {
    "credits_spent": 1,
    "balance": 99,
    "bytes": 14233
  }
}

Every YouTube endpoint

Frequently asked questions

Which YouTube URLs does the transcript endpoint accept?
Standard watch URLs, youtu.be short links, Shorts, embed and live URLs, and the bare video id. The full list, with an example of each, is on the YouTube URL formats page.
What happens when a YouTube video has no captions?
With the default mode, auto, the audio is transcribed instead. Media under 20 minutes is held open for up to 45 seconds and returns the finished transcript inline; longer media, or a transcription still running when the hold expires, answers 202 with a job_id and poll_url. Send mode captions to fail with no_captions instead.
Do I get timestamps?
Yes by default: data.segments carries start, duration and text for every caption cue. Send timestamps false to get a single data.text string instead; a response never carries both.
How is a YouTube transcript billed?
A transcript served from captions costs 1 credit. One transcribed from audio is billed by duration, on delivery. A request that fails, or a video that is private, removed or otherwise unavailable, costs nothing.