Using the API

The response format: text, segments and timestamps

What comes back from a successful fetch, and how to use the segment timings.

Every successful response uses the same envelope, whatever the platform:

json
{
  "ok": true,
  "request_id": "req_…",
  "data": { … },
  "usage": { "credits_spent": 1 }
}

The data block

For a transcript, data carries kind plus one metadata block that is identical on every platform and every path (a fresh fetch, a cache hit, an AI transcription, a job poll, a webhook): video_id (the item's id on its platform), url, platform, title, channel (the creator: channel name, @handle or username), duration in seconds, language, thumbnail_url, and source (captions or audio), then the words as text or segments. Every metadata key is always present; null means we could not determine it, never that it was left out.

Segments

Each segment has start, duration and text, all in seconds. That is what makes the output useful for subtitles, clipping, and citing the exact second a claim was made rather than quoting a wall of text.

Send "timestamps": false to get one joined text string instead of the segments array when you only need the words. The default, true, returns segments; exactly one of the two is present on a single-transcript response.

request_id

Quote it if you contact support. It identifies the exact call in our logs.

The same envelope comes back for every platform, which is the point of using one transcript API rather than four integrations.

Still stuck?

Open the chat launcher, bottom right, and include your request_id if you have one. Or email [email protected].