Reference

Instagram Transcript API

One POST turns an Instagram Reel or video post URL into a transcript in the same JSON shape every platform returns. Instagram exposes no caption track a logged-out client can read, so on Instagram the audio path is the normal one: the Reel's audio is transcribed on the request and returned inline. The response also names the post where Instagram provides it.

1 credit per started 5 minutes of audio, minimum 1View pricing

Instagram exposes no caption track a logged-out client can read, so every Reel is transcribed from its audio and charged only on delivery; a caption transcript would be 1 credit. Failed requests are free.

On this platform

  • Accepts instagram.com/reel/<code>/ and /p/<code>/ URLs, legacy /tv/ URLs and instagr.am short links; only public posts are reachable.
  • Every Instagram transcript is transcribed from audio, so source is always audio and the cost follows the audio rate, charged only on delivery.
  • Reels are short-form, so the request is held open and the transcript almost always comes back inline rather than as a 202 job.
POST/api/v2/transcripts/videoidempotent

Fetch a transcript (YouTube, TikTok, Instagram, or file URL)

Returns a transcript - text plus timestamped segments. Accepts YouTube, TikTok, and Instagram URLs (or a bare TikTok video id), direct media file URLs. ..) identifying who is talking. Speaker ids are hints from voice separation, not named identification. When no captions exist the audio is transcribed automatically: when we can determine the media length, media under 20 minutes simply waits (the request is held open for up to 45 seconds) and returns the finished transcript, so no polling is needed; when the length cannot be determined, only short-form platforms (TikTok and Instagram) are held inline. Longer media, or a transcription still running when the 45-second hold expires, returns 202 with a job to poll instead - the work continues either way, so the same request is safe to retry and will hit the cache once it finishes. Supply callback_url to have the finished transcript POSTed to you instead of polling. Every failure carries an ai_fallback block saying whether captions were definitively unavailable and whether retrying would work.

Body parameters

videostringrequired
A video URL or 11-character YouTube video ID. Accepts YouTube (watch, youtu.be, /shorts/), TikTok, and Instagram URLs, plus direct media file URLs (mp4/mp3/wav/…). Videos without captions fall back to AI transcription.
Example
mode"captions" | "audio" | "auto"optional
Where the text may come from. "captions" reads an existing caption track and fails if there is none, which is the only way to avoid audio transcription. "audio" skips captions and transcribes the audio. "auto" (the default) tries captions first and transcribes the audio when there are none. Short media may finish inline after a wait of up to 45 seconds; longer work returns 202 with a job to poll or deliver by callback. Audio is charged only on delivery: 1 credit per started 5 minutes of audio, minimum 1 (a 20-minute video is 4; the 4-hour cap is 48). Caption fetches are always 1 credit.
Example
timestampsbooleanoptional
Which form the transcript comes back in. true (the default) returns the `segments` array, each with start, duration and text. false returns a single joined `text` string instead. On the single-transcript 200 exactly one of the two is present, never both, since segments already contain every word the joined text does; job results and batch entries carry both text and segments. The older strings "segment" and "none" mean the same two things and are still accepted.
Example
callback_urlstring (https URL)optional
Where to POST the finished transcript when a request escalates to audio transcription, instead of polling the job. The delivery body is a trimmed envelope - {ok, status, job_id, data} on success, {ok, status, job_id, error} on failure - without the usage and request_id the poll URL adds. When a signing secret is configured on our side, the body is signed with HMAC-SHA256 over the exact bytes and sent as an X-TranscriptFetch-Signature: sha256=<hex> header so you can verify it came from us. Must be a public https URL on the standard port; the URL is checked again at delivery time, so an unreachable or private address still gets a 202 but never receives a delivery. The job stays pollable either way, so a missed delivery is never a lost transcript.
ai_fallbackbooleanoptional
Legacy alias for "mode", still supported. true is identical to "mode": "audio"; omitted or false is "mode": "auto". Send one or the other, not both. Note that false never disabled the fallback - audio was still transcribed when no captions existed - which is why the field was replaced.
Example

Request example

curl https://transcriptfetch.com/api/v2/transcripts/video \
  -H "Authorization: Bearer $TRANSCRIPTFETCH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"video":"https://www.instagram.com/reel/Dd1gQBsCWE8/"}'

Responses

SuccessExample response envelope
{
  "ok": true,
  "request_id": "req_…",
  "data": {
    "kind": "transcript",
    "video_id": "Dd1gQBsCWE8",
    "url": "https://www.instagram.com/reel/Dd1gQBsCWE8/",
    "platform": "instagram",
    "title": "What happens when we detect an asteroid that could pose a threat to Earth?\n\nIn early 2025, NASA joined scientists around the world in tracking asteroid 2024 YR4, which had a small chance of impacting ",
    "channel": null,
    "duration": 38.6,
    "language": "english",
    "thumbnail_url": null,
    "source": "audio",
    "segments": [
      {
        "start": 0,
        "duration": 3.44,
        "text": "This asteroid 2024 YR4 might get interesting."
      },
      {
        "start": 3.44,
        "duration": 2.56,
        "text": "There was a chance that it could impact Earth."
      },
      {
        "start": 6,
        "duration": 4.1,
        "text": "The Atlas Survey, operated by the University of Hawaii,"
      }
    ]
  },
  "usage": {
    "credits_spent": 1,
    "balance": 657,
    "bytes": 2210
  }
}

Every Instagram endpoint

Frequently asked questions

Why does every Instagram response say source audio?
Instagram publishes no caption track that a logged-out client can read, so there is nothing to read first. The Reel's audio is transcribed on the request instead, which is why the response's source is always audio and the language is the detected spoken language.
What does an Instagram transcript cost?
The audio rate: 1 credit per started 5 minutes of audio, minimum 1, charged only when the transcript is delivered. Most Reels are well under a minute, so one Reel is one credit. A failed request costs nothing.
Which Instagram URLs are accepted?
Reel URLs (instagram.com/reel/<code>/), video post URLs (instagram.com/p/<code>/), legacy IGTV /tv/ URLs and instagr.am short links. Stories, private accounts and login-walled posts are not reachable and fail with a clear error.
Do I have to poll a job for an Instagram Reel?
Almost never. Reels are short-form, so the request is held open and the finished transcript is returned inline. A transcription that outlasts the hold answers 202 with a job_id and poll_url, and callback_url delivers it to you instead.