Back to blog

The Best TikTok Transcript APIs in 2026

Compare six TikTok transcript APIs by pricing, timestamps, audio fallback, batches, and monitoring to choose the right fit for your app.

The best TikTok transcript API depends on what you need when a video has no readable captions. A service that extracts an existing subtitle track and a service that retrieves the audio and transcribes it solve different parts of the problem. Both may advertise a TikTok transcript endpoint.

We build TranscriptFetch, one of the options below. This comparison uses public vendor documentation and pricing checked on September 28, 2026, plus our own API contract. The recommendations reflect workflow fit; they are not the results of a controlled speed or accuracy benchmark.

TL;DR: We recommend TranscriptFetch’s TikTok Transcript API for production apps and agents that need timestamped transcripts, automatic audio fallback, batches, and ongoing profile monitoring. ScrapeCreators fits pay-as-you-go social-data collection, Supadata adds broader platform coverage, and SocialKit fits caption-backed TikTok analysis. Choose DumplingAI when transcripts belong inside a wider automation stack, or transcript.im when a browser workspace and exports matter alongside the API.

At a glance

TL;DR comparison

APIEntry priceTimestamp formatWhen captions are missingBest for
TranscriptFetch50 credits/month free; $5/month for 1,000JSON segmentsAutomatic audio transcription; 1 credit per started 5 minutesProduction apps, agents, batches, and profile monitoring
ScrapeCreators100 signup credits; $47 for 25,000WebVTT in documented responseOptional AI fallback; up to 2 minutes, 10-credit add-onPrepaid social-data collection
Supadata100 credits/month free; $5/month equivalent, billed $60 annually, for 300/monthStructured transcript chunksAutomatic or explicit generation; 2 credits/minuteMulti-platform transcript integrations
SocialKit20 free credits; Standard $20.30/month equivalent, billed $243.60 annually, for 12,000/monthJSON segmentsNo audio fallback on its TikTok transcript endpointCaptions alongside summaries and creator data
DumplingAIStarter $40/month equivalent, billed annually; 1.2 million credits/yearWebVTT inside JSONNo fallback option documented on this endpointTranscripts within broader data workflows
transcript.imAPI access on Pro: $10/month; 500 transcripts and 1,500 ASR minutesTimestamped transcript dataAsynchronous speech-to-text jobsBrowser workspace, downloads, and API together

Credit units differ. DumplingAI’s TikTok endpoint consumes 10 credits per successful request; its large credit allowance is not an equal number of transcripts. Annual prices above require an annual commitment. Sources and endpoint-specific limits are linked in each section.

Evaluation criteria

What makes a TikTok transcript API good?

A useful API has to get from the public link to usable spoken text. Compare that complete path, including failures, rather than counting features on a homepage.

  • Public video retrieval: Does it resolve normal TikTok video links and the short links your users paste?

  • Captionless videos: Will it generate a transcript from speech, return a clear absence error, or require a second service?

  • Timestamped output: Are times supplied as JSON fields, subtitle cues, or only a plain-text paragraph?

  • Audio pricing: Is transcription charged per request, per minute, or as an extra fee on top of retrieval?

  • Job handling: Can your client distinguish a completed transcript from a pending transcription job?

  • Collection workflow: Can you enumerate a profile, process selected videos in batches, or watch for new uploads?

  • Failure behavior: Are unavailable videos charged, and does the response tell you whether retrying can help?

In a RAG index, reliable source links and segment times may matter more than built-in summaries. When researching creators, adjacent profile and comments endpoints can be the deciding factor. Those are different buying decisions.

01

TranscriptFetch

Best overall for production apps and agents

Entry price
50 credits/month free; $5/month for 1,000
Timestamp format
JSON segments
When captions are missing
Automatic audio transcription; 1 credit per started 5 minutes

The TranscriptFetch TikTok Transcript API takes a public video URL and returns timestamped JSON. The default auto mode tries captions before audio transcription. Use captions when you only want an existing track, or audio when you specifically want speech recognition. The response identifies which source supplied the text.

The surrounding workflow is the reason we recommend it. List a TikTok profile or playlist, search for videos, and pass the resulting URLs to the transcript endpoint. Batch requests accept up to 50 URLs on Free, Basic, and Pro, or 500 on Mega and Scale. Each item has its own outcome, so a failed video does not erase successful results.

For ongoing collection, profile monitors check for new uploads and expose events through the API, with optional webhook delivery and transcription. Creating a monitor is free; checks and transcripts follow the documented billing rules. Python and JavaScript SDKs support monitors, while the TikTok MCP server gives agents access to transcript and discovery tools.

Your integration still needs a pending state. Audio work can return HTTP 202 with a job to poll instead of an immediate transcript. A batch can likewise contain processing items. Public availability is also a requirement: an API key does not unlock private or deleted videos.

02

ScrapeCreators

Best pay-as-you-go social-data option

Entry price
100 signup credits; $47 for 25,000
Timestamp format
WebVTT in documented response
When captions are missing
Optional AI fallback; up to 2 minutes, 10-credit add-on

ScrapeCreators’ TikTok transcript endpoint documents a WebVTT transcript, so timing lives in subtitle cues rather than a ready-made segment array. An existing transcript costs one credit and has no documented length limit. The optional use_ai_as_fallback setting supports videos up to two minutes; the docs list a 10-credit add-on when that path runs.

That distinction changes the comparison. ScrapeCreators does offer an audio fallback, but you must enable it, and its two-minute limit belongs to that fallback rather than to every caption fetch. If your application requires JSON segments, allow for WebVTT parsing too.

The broader account covers TikTok profiles, videos, comments, search, and other social platforms. The published entry pack is $47 for 25,000 credits after 100 free signup credits, and credits do not expire. This suits irregular collection better than paying for an allowance you might not use.

03

Supadata

Best for broader platform coverage

Entry price
100 credits/month free; $5/month equivalent, billed $60 annually, for 300/month
Timestamp format
Structured transcript chunks
When captions are missing
Automatic or explicit generation; 2 credits/minute

Supadata’s transcript API supports TikTok, YouTube, Instagram, X, Facebook, and public file URLs. It separates native, auto, and generate modes. That gives you an explicit choice between existing captions and generated text, with structured chunks or a plain-text response.

The pricing page lists 100 free credits per month. Basic includes 300 monthly credits at a $5 monthly equivalent, billed $60 annually; it is annual-only. Pro is $17 monthly for 3,000 credits. Fetching an existing transcript uses one credit, while generating one uses two credits per minute.

Build for both HTTP 200 and HTTP 202. Supadata can return a job ID for generation, and its docs warn that a request may switch to asynchronous processing while running. Choosing generate also ignores the preferred-language parameter and transcribes in the video’s language, so it should not be treated as a translation request.

04

SocialKit

Best for captions plus TikTok analysis

Entry price
20 free credits; Standard $20.30/month equivalent, billed $243.60 annually, for 12,000/month
Timestamp format
JSON segments
When captions are missing
No audio fallback on its TikTok transcript endpoint

SocialKit’s TikTok transcript API returns readable text and segments containing start time and duration. It also exposes subtitle-language information. The catalog includes related TikTok summaries, video and channel statistics, comments, and search endpoints.

The important limit is explicit in its troubleshooting guide: TikTok transcript requests retrieve existing captions and do not generate speech-to-text when captions are absent. Confirmed absence produces a non-retryable no_transcript error. Failed requests use zero credits. A subtitle marked ASR describes the source track; it does not establish a fallback that creates a new transcript for your request.

SocialKit’s pricing includes 20 free credits. The displayed annual Standard plan is $20.30 per month equivalent, billed $243.60 annually, with 12,000 monthly credits. Most requests cost one credit, with some endpoint and workload exceptions.

05

DumplingAI

Best for a broader automation stack

Entry price
Starter $40/month equivalent, billed annually; 1.2 million credits/year
Timestamp format
WebVTT inside JSON
When captions are missing
No fallback option documented on this endpoint

DumplingAI’s TikTok endpoint accepts a video URL and optional preferred language. The response wraps WebVTT in JSON. You get subtitle timing, but must parse the cues if your downstream application expects individual segment objects.

The endpoint documents 10 credits per successful request and says transcript and language availability depend on the source video. It does not document an audio-fallback option. That is a limitation of the published endpoint contract, not proof that every other DumplingAI media capability behaves the same way.

The wider platform combines transcripts with web data, search, document processing, and automation capabilities. The annual Starter offer is $40 per month equivalent with 1.2 million credits per year. Credits are shared across those capabilities, and their per-action costs differ. Failed requests do not consume credits; subscription credits do not roll over.

06

transcript.im

Best browser workspace and API combination

Entry price
API access on Pro: $10/month; 500 transcripts and 1,500 ASR minutes
Timestamp format
Timestamped transcript data
When captions are missing
Asynchronous speech-to-text jobs

transcript.im offers timestamped transcripts from public TikTok URLs alongside a browser product. Its API uses available captions and sends captionless media through asynchronous speech recognition. The integration therefore needs to handle either an immediate transcript or a 202 job and poll for completion.

The Pro plan costs $10 per month and includes API access, 500 monthly transcripts, 1,500 ASR minutes, and batches of up to 20 videos. The free browser plan offers three transcripts per UTC day, but the pricing page lists API access under Pro. Do not count that browser allowance as a free developer tier.

This combination can be useful when a researcher wants to inspect or download the same material an application consumes. TXT, SRT, and VTT exports serve a different need from an API-only response. Budget against both transcript count and audio minutes, since those allowances are separate.

Guide

A caption is not always a transcript

TikTok uses several kinds of text. The description beneath a video may contain hashtags and a creator’s written caption. A subtitle track contains timed spoken words. Text drawn into the video image is another source entirely.

Request the spoken words when you need a transcript. A metadata API returning the description has not transcribed the video. A speech-to-text model will not necessarily recover silent text overlays either; that requires visual text recognition.

Keep those distinctions in your data model. Store a transcript’s source and language with the video URL, and preserve timestamps if you intend to generate citations, clips, or a searchable index. An empty transcript can be correct for a silent photo post. It is different from an extraction error.

Official platform

What about TikTok’s official API?

TikTok does have an official transcript-related field. The Research API’s video query lists voice_to_text among its available fields. It would be inaccurate to say TikTok never exposes spoken text through an official API.

Access is the deciding constraint. TikTok’s Research Tools eligibility rules require an approved research application, eligible affiliation and region, and research independent of commercial interests. This is a different route from signing up for a developer key to serve an ordinary commercial application.

If your project qualifies, evaluate the official research access first. Check whether the fields and coverage meet your study’s needs. Do not assume voice_to_text is equivalent to a service that generates missing transcripts and returns timed segments for arbitrary public links.

Guide

How much will your TikTok workload cost?

Estimate captions and audio separately. One thousand successful caption fetches use 1,000 TranscriptFetch credits. One thousand successful two-minute audio transcriptions also use 1,000 credits under our five-minute billing blocks. A six-minute audio transcript uses two credits.

Those figures describe consumption, not a universal monthly invoice. Add profile discovery and monitoring checks if you use them, then choose a plan with enough allowance. For other vendors, apply their own credit-to-dollar conversion and captionless-video rules. A large credit number is meaningless without the endpoint cost beside it.

Before buying a larger plan, test a small set drawn from your actual workload: readable captions, missing captions, short links, non-English speech, music under speech, and unavailable videos. Record retrieval success, transcript usefulness, timestamps, pending-job handling, and credits charged. Keep an inaccessible URL separate from a successfully retrieved clip with poor speech recognition.

Build vs. buy

What about building it yourself?

A DIY pipeline needs public-media retrieval, caption parsing, audio extraction, speech recognition, and storage. You also own retries, expiring media URLs, job state, and normalization between subtitle-derived and generated text.

Self-hosting can make sense when the media is already in your storage or you need tight control over models and retention. Starting from public TikTok links adds another dependency: keeping retrieval working as the platform changes. The engineering cost is ongoing, even when the downloader and transcription model are open source.

When transcripts support another product, compare that ownership with the hosted API’s bill. Avoid declaring either option cheaper before counting developer time and failed work.

Buying guide

How to choose the best TikTok transcript API

  • Best overall for production apps and agents: TranscriptFetch, for captions, audio fallback, batches, and profile monitoring.

  • Best for irregular social-data collection: ScrapeCreators, for prepaid credits and a broader endpoint catalog.

  • Best for additional video platforms: Supadata, particularly when X or Facebook belongs in the same integration.

  • Best for caption-backed TikTok analysis: SocialKit, with transcript, summary, and creator-data endpoints.

  • Best inside an existing automation stack: DumplingAI, when you already use its other data capabilities.

  • Best browser workspace plus API: transcript.im, for reviewing and exporting transcripts alongside API use.

For our recommended production path, start with TranscriptFetch’s TikTok Transcript API, then test a captionless clip and a multi-video batch before committing. If your input mix spans platforms, compare the Instagram API options and YouTube API options too. Similar endpoint names do not guarantee identical behavior on all three platforms.

FAQ

We recommend TranscriptFetch for production apps and agents that need timestamped transcripts, automatic audio fallback, batches, and profile monitoring. ScrapeCreators suits prepaid social-data collection, Supadata adds broader platform coverage, and SocialKit fits workloads with existing TikTok captions. The best choice depends on your captionless-video rate and surrounding workflow.