# Generate premium captions

Generate high-accuracy captions for a video and attach them as a text track with the Mux Robots API. Use premium captions when accuracy matters, such as for accessibility or compliance.

Generate captions for a Mux asset and automatically attach them as a text track. Premium captions use a higher-accuracy speech model than Mux's standard [auto-generated captions](/docs/guides/add-autogenerated-captions-and-use-transcripts), and add optional speaker labels, word-level timestamps, and custom phrase hints for proper nouns and jargon. Captions are generated directly from the asset's audio, so no existing track is required. See the <ApiRefLink href="/docs/api-reference/robots/generate-premium-captions">Generate Premium Captions API reference</ApiRefLink> for the full endpoint specification. See [Mux Robots pricing](/docs/pricing/overview#mux-robots-pricing) for unit costs.

<Callout type="info">
  If a text track with the same language code already exists on the asset, the job is rejected. Set `replace_existing` to `true` to delete the existing track first.
</Callout>

<Callout type="info">
  To produce captions, this workflow needs an `audio-only` [static rendition](/docs/guides/enable-static-mp4-renditions) of the asset. If the asset already has one, it's reused. If not, the workflow creates one to process the audio and deletes it once the job completes, so you aren't charged for extra storage.
</Callout>

## Create a `generate-premium-captions` job

```bash
curl https://api.mux.com/robots/v0/jobs/generate-premium-captions \
  -H "Content-Type: application/json" \
  -X POST \
  -d '{
    "parameters": {
      "asset_id": "YOUR_ASSET_ID",
      "language_code": "en",
      "phrases": ["Mux", "API"]
    }
  }' \
  -u ${MUX_TOKEN_ID}:${MUX_TOKEN_SECRET}
```

<Callout type="info">
  This request is **asynchronous**. The `POST` returns immediately with the job in `pending` status and does not include results. **We strongly recommend listening for the [`robots.job.generate_premium_captions.completed` webhook](/docs/guides/robots#webhooks):** the payload contains the full completed job, so no follow-up API call is needed. If webhooks aren't an option, you can poll `GET /robots/v0/jobs/generate-premium-captions/{JOB_ID}` with the `id` from the response until the status is `completed`.
</Callout>

## Parameters

| Parameter | Type | Description |
| :-- | :-- | :-- |
| `asset_id` | string | **Required.** The Mux asset ID of the video to caption. |
| `language_code` | string | BCP 47 language code of the audio (e.g. `en`, `es`). Auto-detected when omitted. See [language support for Mux Robots](/docs/guides/robots-supported-languages). |
| `replace_existing` | boolean | When `true`, any existing text track with the same language code is deleted before the new track is uploaded. When `false` (the default), the request is rejected if a matching track already exists. |
| `track_name` | string | Custom name for the uploaded Mux text track. Defaults to `"{Language} (Generated)"` using the resolved language code. |
| `include_speakers` | boolean | When `true`, speaker labels are identified and added to each caption cue. Useful for interviews, podcasts, and multi-speaker content. Defaults to `false`. |
| `include_words` | boolean | When `true`, word-level timestamps are exported as a JSON file accessible via `temporary_words_url` in the output. Billed at a higher unit rate. Defaults to `false`. |
| `upload_to_mux` | boolean | Whether to upload the generated captions as a new text track on the asset. Defaults to `true`. When `false`, no track is created (and `replace_existing` must also be `false`); the captions remain available via `temporary_srt_url`. |
| `phrases` | array of strings | Best-effort list of words or short phrases (proper nouns, product names, jargon) likely to appear in the audio, used to bias recognition toward correct spellings. Up to 100 phrases, each up to 50 characters. Does not guarantee exact output. |

## Output

The `outputs` object is included in the job once its status is `completed`. You'll receive it on the [`robots.job.generate_premium_captions.completed`](/docs/guides/robots#webhooks) webhook (recommended), or you can fetch it with `GET /robots/v0/jobs/generate-premium-captions/{JOB_ID}`. It contains:

| Field | Type | Description |
| :-- | :-- | :-- |
| `track_id` | string | Mux text track ID of the newly uploaded caption track. Omitted when `upload_to_mux` is `false`. |
| `language_code` | string | Resolved language code of the generated captions (may differ from the requested code when auto-detected). |
| `temporary_srt_url` | string | Temporary pre-signed URL to download the generated SRT file. Expires 7 days after the job completes. |
| `temporary_words_url` | string | Temporary pre-signed URL to download the word-level timestamp JSON. Present when `include_words` is `true`. Expires 7 days after the job completes, so download and store it for long-term access. |
| `replaced_track_id` | string | Mux track ID of the deleted track, present when `replace_existing` was `true`. |

## Example response

This example uses a 6-minute asset with `include_words: false`: at 500 units per minute, the job consumes 3,000 units. This is the payload delivered to the [`robots.job.generate_premium_captions.completed`](/docs/guides/robots#webhooks) webhook, and the same shape you get from `GET /robots/v0/jobs/generate-premium-captions/{JOB_ID}`:

```json
{
  "data": {
    "id": "rjob_yza567",
    "workflow": "generate-premium-captions",
    "status": "completed",
    "units_consumed": 3000,
    "parameters": {
      "asset_id": "YOUR_ASSET_ID",
      "language_code": "en",
      "replace_existing": false,
      "include_speakers": false,
      "include_words": false,
      "upload_to_mux": true,
      "phrases": ["Mux", "API"]
    },
    "outputs": {
      "track_id": "track_en_abc123",
      "language_code": "en",
      "temporary_srt_url": "https://storage.googleapis.com/..."
    }
  }
}
```

<Callout type="info">
  When `upload_to_mux` is `true` (the default), the caption track is automatically attached to your asset, and viewers will see the new language option in the player's caption menu.
</Callout>

## Word-level timestamps

When `include_words` is `true`, download the file at `temporary_words_url` to get word-level timestamps. It's a JSON array of token objects in playback order. The array interleaves `spacing` tokens between words so you can reconstruct the exact text, and `audio_event` tokens capture non-speech sounds.

| Field | Type | Description |
| :-- | :-- | :-- |
| `text` | string | The token's text. For `word` tokens this includes any trailing punctuation (e.g. `"Matt,"`, `"is."`). For `spacing` tokens it's a single space (`" "`). For `audio_event` tokens it's a bracketed non-speech cue (e.g. `"[laughs]"`). |
| `start` | number | Start time of the token, in seconds from the start of the media (fractional, e.g. `26.38`). |
| `end` | number | End time of the token, in seconds. |
| `type` | string | One of `word`, `spacing`, or `audio_event`. |
| `speaker_id` | string | The speaker the token is attributed to, formatted `speaker_N` (zero-indexed). Set only when `include_speakers` is `true`. |

```json
[
  { "text": "Hey,", "start": 0.2, "end": 0.42, "type": "word", "speaker_id": "speaker_0" },
  { "text": " ", "start": 0.42, "end": 0.45, "type": "spacing", "speaker_id": "speaker_0" },
  { "text": "I'm", "start": 0.45, "end": 0.6, "type": "word", "speaker_id": "speaker_0" },
  { "text": " ", "start": 0.6, "end": 0.63, "type": "spacing", "speaker_id": "speaker_0" },
  { "text": "a", "start": 0.63, "end": 0.7, "type": "word", "speaker_id": "speaker_0" },
  { "text": " ", "start": 0.7, "end": 0.73, "type": "spacing", "speaker_id": "speaker_0" },
  { "text": "Mux", "start": 0.73, "end": 0.98, "type": "word", "speaker_id": "speaker_0" },
  { "text": " ", "start": 0.98, "end": 1.01, "type": "spacing", "speaker_id": "speaker_0" },
  { "text": "Robot,", "start": 1.01, "end": 1.4, "type": "word", "speaker_id": "speaker_0" },
  { "text": " ", "start": 1.4, "end": 1.43, "type": "spacing", "speaker_id": "speaker_0" },
  { "text": "beep", "start": 1.43, "end": 1.7, "type": "word", "speaker_id": "speaker_0" },
  { "text": " ", "start": 1.7, "end": 1.73, "type": "spacing", "speaker_id": "speaker_0" },
  { "text": "boop.", "start": 1.73, "end": 2.05, "type": "word", "speaker_id": "speaker_0" },
  { "text": "[beeps]", "start": 2.1, "end": 2.4, "type": "audio_event", "speaker_id": "speaker_1" },
  { "text": "Beep", "start": 2.5, "end": 2.72, "type": "word", "speaker_id": "speaker_1" },
  { "text": " ", "start": 2.72, "end": 2.75, "type": "spacing", "speaker_id": "speaker_1" },
  { "text": "boop", "start": 2.75, "end": 2.98, "type": "word", "speaker_id": "speaker_1" },
  { "text": " ", "start": 2.98, "end": 3.01, "type": "spacing", "speaker_id": "speaker_1" },
  { "text": "yourself!", "start": 3.01, "end": 3.5, "type": "word", "speaker_id": "speaker_1" }
]
```
