The best TikTok transcript API depends on what you need when a video has no readable captions. A service that extracts an existing subtitle track and a service that retrieves the audio and transcribes it solve different parts of the problem. Both may advertise a TikTok transcript endpoint.
We build TranscriptFetch, one of the options below. This comparison uses public vendor documentation and pricing checked on September 28, 2026, plus our own API contract. The recommendations reflect workflow fit; they are not the results of a controlled speed or accuracy benchmark.
TL;DR: We recommend TranscriptFetch’s TikTok Transcript API for production apps and agents that need timestamped transcripts, automatic audio fallback, batches, and ongoing profile monitoring. ScrapeCreators fits pay-as-you-go social-data collection, Supadata adds broader platform coverage, and SocialKit fits caption-backed TikTok analysis. Choose DumplingAI when transcripts belong inside a wider automation stack, or transcript.im when a browser workspace and exports matter alongside the API.
TL;DR comparison
| API | Entry price | Timestamp format | When captions are missing | Best for |
|---|---|---|---|---|
| TranscriptFetch | 50 credits/month free; $5/month for 1,000 | JSON segments | Automatic audio transcription; 1 credit per started 5 minutes | Production apps, agents, batches, and profile monitoring |
| ScrapeCreators | 100 signup credits; $47 for 25,000 | WebVTT in documented response | Optional AI fallback; up to 2 minutes, 10-credit add-on | Prepaid social-data collection |
| Supadata | 100 credits/month free; $5/month equivalent, billed $60 annually, for 300/month | Structured transcript chunks | Automatic or explicit generation; 2 credits/minute | Multi-platform transcript integrations |
| SocialKit | 20 free credits; Standard $20.30/month equivalent, billed $243.60 annually, for 12,000/month | JSON segments | No audio fallback on its TikTok transcript endpoint | Captions alongside summaries and creator data |
| DumplingAI | Starter $40/month equivalent, billed annually; 1.2 million credits/year | WebVTT inside JSON | No fallback option documented on this endpoint | Transcripts within broader data workflows |
| transcript.im | API access on Pro: $10/month; 500 transcripts and 1,500 ASR minutes | Timestamped transcript data | Asynchronous speech-to-text jobs | Browser workspace, downloads, and API together |
Credit units differ. DumplingAI’s TikTok endpoint consumes 10 credits per successful request; its large credit allowance is not an equal number of transcripts. Annual prices above require an annual commitment. Sources and endpoint-specific limits are linked in each section.
What makes a TikTok transcript API good?
A useful API has to get from the public link to usable spoken text. Compare that complete path, including failures, rather than counting features on a homepage.
Public video retrieval: Does it resolve normal TikTok video links and the short links your users paste?
Captionless videos: Will it generate a transcript from speech, return a clear absence error, or require a second service?
Timestamped output: Are times supplied as JSON fields, subtitle cues, or only a plain-text paragraph?
Audio pricing: Is transcription charged per request, per minute, or as an extra fee on top of retrieval?
Job handling: Can your client distinguish a completed transcript from a pending transcription job?
Collection workflow: Can you enumerate a profile, process selected videos in batches, or watch for new uploads?
Failure behavior: Are unavailable videos charged, and does the response tell you whether retrying can help?
In a RAG index, reliable source links and segment times may matter more than built-in summaries. When researching creators, adjacent profile and comments endpoints can be the deciding factor. Those are different buying decisions.
TranscriptFetch
Best overall for production apps and agents
- Entry price
- 50 credits/month free; $5/month for 1,000
- Timestamp format
- JSON segments
- When captions are missing
- Automatic audio transcription; 1 credit per started 5 minutes
The TranscriptFetch TikTok Transcript API takes a public video URL and returns timestamped JSON. The default auto mode tries captions before audio transcription. Use captions when you only want an existing track, or audio when you specifically want speech recognition. The response identifies which source supplied the text.
The surrounding workflow is the reason we recommend it. List a TikTok profile or playlist, search for videos, and pass the resulting URLs to the transcript endpoint. Batch requests accept up to 50 URLs on Free, Basic, and Pro, or 500 on Mega and Scale. Each item has its own outcome, so a failed video does not erase successful results.
For ongoing collection, profile monitors check for new uploads and expose events through the API, with optional webhook delivery and transcription. Creating a monitor is free; checks and transcripts follow the documented billing rules. Python and JavaScript SDKs support monitors, while the TikTok MCP server gives agents access to transcript and discovery tools.
Your integration still needs a pending state. Audio work can return HTTP 202 with a job to poll instead of an immediate transcript. A batch can likewise contain processing items. Public availability is also a requirement: an API key does not unlock private or deleted videos.
ScrapeCreators
Best pay-as-you-go social-data option
- Entry price
- 100 signup credits; $47 for 25,000
- Timestamp format
- WebVTT in documented response
- When captions are missing
- Optional AI fallback; up to 2 minutes, 10-credit add-on
ScrapeCreators’ TikTok transcript endpoint documents a WebVTT transcript, so timing lives in subtitle cues rather than a ready-made segment array. An existing transcript costs one credit and has no documented length limit. The optional use_ai_as_fallback setting supports videos up to two minutes; the docs list a 10-credit add-on when that path runs.
That distinction changes the comparison. ScrapeCreators does offer an audio fallback, but you must enable it, and its two-minute limit belongs to that fallback rather than to every caption fetch. If your application requires JSON segments, allow for WebVTT parsing too.
The broader account covers TikTok profiles, videos, comments, search, and other social platforms. The published entry pack is $47 for 25,000 credits after 100 free signup credits, and credits do not expire. This suits irregular collection better than paying for an allowance you might not use.
Supadata
Best for broader platform coverage
- Entry price
- 100 credits/month free; $5/month equivalent, billed $60 annually, for 300/month
- Timestamp format
- Structured transcript chunks
- When captions are missing
- Automatic or explicit generation; 2 credits/minute
Supadata’s transcript API supports TikTok, YouTube, Instagram, X, Facebook, and public file URLs. It separates native, auto, and generate modes. That gives you an explicit choice between existing captions and generated text, with structured chunks or a plain-text response.
The pricing page lists 100 free credits per month. Basic includes 300 monthly credits at a $5 monthly equivalent, billed $60 annually; it is annual-only. Pro is $17 monthly for 3,000 credits. Fetching an existing transcript uses one credit, while generating one uses two credits per minute.
Build for both HTTP 200 and HTTP 202. Supadata can return a job ID for generation, and its docs warn that a request may switch to asynchronous processing while running. Choosing generate also ignores the preferred-language parameter and transcribes in the video’s language, so it should not be treated as a translation request.
SocialKit
Best for captions plus TikTok analysis
- Entry price
- 20 free credits; Standard $20.30/month equivalent, billed $243.60 annually, for 12,000/month
- Timestamp format
- JSON segments
- When captions are missing
- No audio fallback on its TikTok transcript endpoint
SocialKit’s TikTok transcript API returns readable text and segments containing start time and duration. It also exposes subtitle-language information. The catalog includes related TikTok summaries, video and channel statistics, comments, and search endpoints.
The important limit is explicit in its troubleshooting guide: TikTok transcript requests retrieve existing captions and do not generate speech-to-text when captions are absent. Confirmed absence produces a non-retryable no_transcript error. Failed requests use zero credits. A subtitle marked ASR describes the source track; it does not establish a fallback that creates a new transcript for your request.
SocialKit’s pricing includes 20 free credits. The displayed annual Standard plan is $20.30 per month equivalent, billed $243.60 annually, with 12,000 monthly credits. Most requests cost one credit, with some endpoint and workload exceptions.
DumplingAI
Best for a broader automation stack
- Entry price
- Starter $40/month equivalent, billed annually; 1.2 million credits/year
- Timestamp format
- WebVTT inside JSON
- When captions are missing
- No fallback option documented on this endpoint
DumplingAI’s TikTok endpoint accepts a video URL and optional preferred language. The response wraps WebVTT in JSON. You get subtitle timing, but must parse the cues if your downstream application expects individual segment objects.
The endpoint documents 10 credits per successful request and says transcript and language availability depend on the source video. It does not document an audio-fallback option. That is a limitation of the published endpoint contract, not proof that every other DumplingAI media capability behaves the same way.
The wider platform combines transcripts with web data, search, document processing, and automation capabilities. The annual Starter offer is $40 per month equivalent with 1.2 million credits per year. Credits are shared across those capabilities, and their per-action costs differ. Failed requests do not consume credits; subscription credits do not roll over.
transcript.im
Best browser workspace and API combination
- Entry price
- API access on Pro: $10/month; 500 transcripts and 1,500 ASR minutes
- Timestamp format
- Timestamped transcript data
- When captions are missing
- Asynchronous speech-to-text jobs
transcript.im offers timestamped transcripts from public TikTok URLs alongside a browser product. Its API uses available captions and sends captionless media through asynchronous speech recognition. The integration therefore needs to handle either an immediate transcript or a 202 job and poll for completion.
The Pro plan costs $10 per month and includes API access, 500 monthly transcripts, 1,500 ASR minutes, and batches of up to 20 videos. The free browser plan offers three transcripts per UTC day, but the pricing page lists API access under Pro. Do not count that browser allowance as a free developer tier.
This combination can be useful when a researcher wants to inspect or download the same material an application consumes. TXT, SRT, and VTT exports serve a different need from an API-only response. Budget against both transcript count and audio minutes, since those allowances are separate.
A caption is not always a transcript
TikTok uses several kinds of text. The description beneath a video may contain hashtags and a creator’s written caption. A subtitle track contains timed spoken words. Text drawn into the video image is another source entirely.
Request the spoken words when you need a transcript. A metadata API returning the description has not transcribed the video. A speech-to-text model will not necessarily recover silent text overlays either; that requires visual text recognition.
Keep those distinctions in your data model. Store a transcript’s source and language with the video URL, and preserve timestamps if you intend to generate citations, clips, or a searchable index. An empty transcript can be correct for a silent photo post. It is different from an extraction error.
What about TikTok’s official API?
TikTok does have an official transcript-related field. The Research API’s video query lists voice_to_text among its available fields. It would be inaccurate to say TikTok never exposes spoken text through an official API.
Access is the deciding constraint. TikTok’s Research Tools eligibility rules require an approved research application, eligible affiliation and region, and research independent of commercial interests. This is a different route from signing up for a developer key to serve an ordinary commercial application.
If your project qualifies, evaluate the official research access first. Check whether the fields and coverage meet your study’s needs. Do not assume voice_to_text is equivalent to a service that generates missing transcripts and returns timed segments for arbitrary public links.
How much will your TikTok workload cost?
Estimate captions and audio separately. One thousand successful caption fetches use 1,000 TranscriptFetch credits. One thousand successful two-minute audio transcriptions also use 1,000 credits under our five-minute billing blocks. A six-minute audio transcript uses two credits.
Those figures describe consumption, not a universal monthly invoice. Add profile discovery and monitoring checks if you use them, then choose a plan with enough allowance. For other vendors, apply their own credit-to-dollar conversion and captionless-video rules. A large credit number is meaningless without the endpoint cost beside it.
Before buying a larger plan, test a small set drawn from your actual workload: readable captions, missing captions, short links, non-English speech, music under speech, and unavailable videos. Record retrieval success, transcript usefulness, timestamps, pending-job handling, and credits charged. Keep an inaccessible URL separate from a successfully retrieved clip with poor speech recognition.
What about building it yourself?
A DIY pipeline needs public-media retrieval, caption parsing, audio extraction, speech recognition, and storage. You also own retries, expiring media URLs, job state, and normalization between subtitle-derived and generated text.
Self-hosting can make sense when the media is already in your storage or you need tight control over models and retention. Starting from public TikTok links adds another dependency: keeping retrieval working as the platform changes. The engineering cost is ongoing, even when the downloader and transcription model are open source.
When transcripts support another product, compare that ownership with the hosted API’s bill. Avoid declaring either option cheaper before counting developer time and failed work.
How to choose the best TikTok transcript API
Best overall for production apps and agents: TranscriptFetch, for captions, audio fallback, batches, and profile monitoring.
Best for irregular social-data collection: ScrapeCreators, for prepaid credits and a broader endpoint catalog.
Best for additional video platforms: Supadata, particularly when X or Facebook belongs in the same integration.
Best for caption-backed TikTok analysis: SocialKit, with transcript, summary, and creator-data endpoints.
Best inside an existing automation stack: DumplingAI, when you already use its other data capabilities.
Best browser workspace plus API: transcript.im, for reviewing and exporting transcripts alongside API use.
For our recommended production path, start with TranscriptFetch’s TikTok Transcript API, then test a captionless clip and a multi-video batch before committing. If your input mix spans platforms, compare the Instagram API options and YouTube API options too. Similar endpoint names do not guarantee identical behavior on all three platforms.
FAQ
We recommend TranscriptFetch for production apps and agents that need timestamped transcripts, automatic audio fallback, batches, and profile monitoring. ScrapeCreators suits prepaid social-data collection, Supadata adds broader platform coverage, and SocialKit fits workloads with existing TikTok captions. The best choice depends on your captionless-video rate and surrounding workflow.
Yes. TranscriptFetch can transcribe the audio automatically, Supadata offers generated transcripts, and transcript.im uses asynchronous speech-to-text jobs. ScrapeCreators documents an optional paid audio fallback for videos up to two minutes. SocialKit’s TikTok transcript endpoint retrieves existing captions without generating missing ones.
TikTok’s Research API exposes a voice_to_text field through its video query endpoint. Access requires approval under the Research Tools eligibility rules. It is not a general developer subscription for ordinary commercial apps, and its text field should not be assumed to provide generated transcripts or timed segments for every public video.
TranscriptFetch provides 50 free credits each month, and Supadata lists 100 monthly credits. ScrapeCreators offers 100 signup credits and SocialKit offers 20 free credits. Audio jobs may consume more than one credit, so a free credit allowance is not always the same number of videos. transcript.im lists developer API access under its paid Pro plan.
Treat public accessibility as a requirement. A normal transcript API key does not grant access to private, deleted, or otherwise inaccessible TikTok videos. If you have an authorized copy of the media, use a service that accepts audio or video files instead of depending on the public TikTok link.
