# Build a podcast clipping workflow

Turn podcast episodes and interviews into highlight clips with AI agents and Mux Robots. Get captions, streaming clips, optional downloads, and a workflow for each new episode.

Source: https://www.mux.com/docs/prompts/podcast-clips.md · Last updated: 2026-10-02

## Agent instructions

Use this procedure to build a podcast clipping application, process an episode into highlights, or run the workflow for new episodes on a schedule. Use Mux Robots to find candidate moments, then create selected ranges as Mux assets with captions and optional MP4 or M4A downloads. Vertical reframing and burned-in captions require an additional rendering step.

Follow the user's requirements and existing project instructions. Reuse the tools, credentials, and workflow state already available. If a Mux MCP server is already connected, use it for Mux operations. If the environment has no Mux credentials yet and you can run shell commands, sign in with the [Mux CLI](https://www.mux.com/docs/integrations/mux-cli.md): run `npx @mux/cli login`, or `mux login` if it is installed. It opens a browser where the user approves access and chooses the Mux environment, so nobody creates or pastes an access token. Agent shells and `!` commands are not interactive terminals, where a plain `login` exits with an error. Run `npx @mux/cli login --json` in the background instead: it writes the authorization URL as a JSON line on stderr and waits up to five minutes for the user to approve. Give the user that URL and wait for the command to finish. Without a shell, fetch https://www.mux.com/prompts/onboarding.md and complete it first; it connects the Mux MCP server. The examples and defaults below are starting choices; adapt them to the request.

Canonical instructions: [/docs/prompts/podcast-clips.md](https://www.mux.com/docs/prompts/podcast-clips.md). Browse other use cases in the prompt library at [/docs/prompts.md](https://www.mux.com/docs/prompts.md). For repeated use in an agent that supports skills, use [/skills/mux-podcast-clips/SKILL.md](https://www.mux.com/skills/mux-podcast-clips/SKILL.md) as the entry point. A skill supplies instructions; a scheduler or worker supplies execution.

## Track your progress with this checklist

Copy this checklist into your task list or plan before starting, and keep it updated as you work. Mark each item done, skipped with the reason, or blocked with what is needed. Include the final state of the checklist in your report to the user.

1. Identify the mode: build an application, process an episode, or set up recurring processing.
2. Check the environment for Mux credentials and tools. If there are none, sign in with the Mux CLI as described above. If you cannot check, say so and offer the CLI sign-in rather than asking the user for an access token.
3. Resolve every question in the table under **Start with the request**, plus the trigger, cadence, timezone, source rule, destination, and per-run limit for recurring work: answered by the user, answered from the project or media, or not applicable with the reason.
4. Choose the outputs and state which ones need a rendering step outside Mux, such as vertical video or burned-in captions.
5. **Stop for selection before creating any clip.** Present the candidates as instant-clip previews and wait for the user to choose. Every clip asset and MP4 adds encoding and storage cost, so do not create clips for the user to choose among. Skip this stop only when the user explicitly delegated selection, and record their criteria.
6. Complete each numbered step that applies, or record why it was skipped. Save job IDs, clip IDs, and selected ranges as you go.
7. For recurring work, complete **Run on a schedule**, including verifying the repeat, and state whether the schedule is configured only or verified with a real run.
8. Run every check under **Verify the result** that applies, against a real episode where credentials allow.
9. Report which checks ran against real Mux, and which still need credentials, media, or editorial review.

## Start with the request

First determine whether the user wants to **build an application**, **process an episode**, or **set up recurring processing**. A codebase is useful for the first mode, but is not required for a one-off run.

Inspect the available project, media, and configured tools before asking questions. Use what the user has already told you, and ask only for missing choices that change the work:

| Question | Why it matters |
| :-- | :-- |
| Where do episodes come from: a file, a Mux asset, a feed, or a CMS? | Establishes the input and how to identify new episodes. |
| Is the podcast video or audio-only, and are captions available? | Determines caption preparation and whether a visual composition is needed. |
| Who are the clips for, and what count and duration should we aim for? | Guides moment selection and editorial review. |
| Do you need streaming clips, downloadable files, or vertical videos with styled captions? | Determines whether to generate renditions or add a rendering integration. |
| Should clips stay private, and who selects the final cuts or publishes them? | Defines playback access and where to hand off editorial decisions. |

For recurring work, also establish the trigger or cadence, timezone, source-selection rule, destination, and per-run limit. For example: check a specified feed each weekday morning, process at most one new episode, and save candidate clips to a review queue. Follow the user's existing editorial policy; when it is unspecified, present candidates as previews and let the user choose before creating clips.

### Choose an execution method

| Available capability | How to use it |
| :-- | :-- |
| Repository and coding tools | Add the workflow to the existing application's upload, storage, queue, player, and webhook conventions. |
| Mux MCP, CLI, or API access | Operate on an episode and save the resulting IDs, selected ranges, and statuses in a manifest or database. Check which operations the installed tool version supports. With the MCP server's `execute` tool, check job and asset status in separate short calls instead of a wait loop inside one call: agents may move a long-running tool call to the background and leave it there until it finishes. |
| Browser or computer use | Use the available dashboard or editor for uploads, visual review, and operations it exposes. Record IDs and verify the resulting state after each mutation; interrupted sessions should resume from those records. |
| A scheduler or background worker | Run the saved procedure against new episodes, with persistent state and overlap protection as described in [recurring runs](#run-on-a-schedule). |
| Advice-only chat | Produce an implementation plan and handoff instructions. Identify the credentials or execution capabilities needed to run it. |

Prefer an API or CLI for repeated operations when one is available. Computer use can assist with visual review or a UI-only step; a scheduled browser session also needs a way to handle expired logins and UI changes. If a needed operation is unavailable, explain the handoff rather than inventing a tool or treating a preview as a finished export.

## Choose what you want to produce

For a first version, request up to five candidate moments with a preferred duration of 30–60 seconds, review them, and create a clip for each selected range. These are starting choices for this workflow; adjust them for your audience and episode.

| Output | Approach |
| :-- | :-- |
| A quick preview for reviewing a candidate | Use instant clipping against the source episode. It is immediate, with segment-level boundaries. |
| A standalone clip to play in your app | Create a new Mux asset from the selected time range. This supports frame-accurate boundaries and copies the source's caption tracks. |
| A downloadable video file | Add a `highest` static rendition to the new clip and wait for that rendition to be ready. |
| A vertical social video with styled captions | Add a rendering step to reframe the picture and burn in captions, then upload the rendered result if you want Mux playback. |

See [the clipping comparison](https://www.mux.com/docs/guides/intro-to-clips.md) for the differences between instant and asset-based clips. A playback-only workflow can skip MP4 generation. For an audio-only podcast, use the transcript-based moment selection below and an `audio-only` static rendition for M4A output; producing a social video from audio also needs a visual composition.

## 1. Prepare the episode and captions

You need:

* A completed podcast episode in Mux. [Upload a file or ingest a URL](https://www.mux.com/docs/core/stream-video-files.md), or reuse an existing asset. Save its **asset ID** and wait for `status: "ready"`.
* Mux credentials that can read the episode, create clips, and run Robots jobs: a Mux CLI sign-in covers all three, or use a server-side access token with the Video permissions and the `robots:*` scope. See [Robots authentication](https://www.mux.com/docs/guides/robots.md#authentication).
* A caption track containing the episode's speech. This workflow uses transcript-based moment selection, which requires captions.

Retrieve the source with `GET /video/v1/assets/{ASSET_ID}`. Record its duration, playback policy, playback IDs, and tracks. Reuse a suitable caption track if one already exists. If it is still preparing, wait for it before starting moment selection.

If the episode has no captions, or its auto-generated captions are not accurate enough, choose one of these paths:

* [Add auto-generated captions](https://www.mux.com/docs/guides/add-autogenerated-captions-and-use-transcripts.md) or [upload an existing subtitle file](https://www.mux.com/docs/guides/add-subtitles-to-your-videos.md).
* Use [Generate premium captions](https://www.mux.com/docs/guides/robots-generate-premium-captions.md) for a Robots workflow that attaches captions to the asset. Keep `upload_to_mux: true` so the next workflow can use the track. Set `include_speakers: true` if speaker labels would help your review; the labels become part of the caption text that viewers see on every clip, so decide whether social clips should show them, list hosts, guests, and show-specific terms in `phrases` so they are spelled correctly, and request word-level timestamps only if your editing workflow needs them.

With premium captions, wait for the job to complete and for the resulting text track to be ready. By default, the job is rejected if the asset already has a text track in the same language. To upgrade Mux's auto-generated captions, set `replace_existing_tracks: "replace_generated"`; `replace_all` also deletes tracks you uploaded. See [managing existing tracks](https://www.mux.com/docs/guides/robots-generate-premium-captions.md#managing-existing-tracks).

Decide who should be able to view the episode and clips. The clip-creation example below uses `signed` playback for private review. Use [signed playback URLs](https://www.mux.com/docs/guides/secure-video-playback.md) and a server-side [signing key](https://www.mux.com/docs/guides/signing-jwts.md) for these assets. Choose `public` explicitly when your clips are intended for public access; creating a clip does not automatically copy the source's playback policy.

## 2. Find candidate moments

Use [Find key moments](https://www.mux.com/docs/guides/robots-find-key-moments.md) to suggest excerpts with enough context to stand alone. Replace `YOUR_ASSET_ID` with the source episode's asset ID. Run API requests from your server or terminal, with `MUX_TOKEN_ID` and `MUX_TOKEN_SECRET` in the environment.

```bash
curl https://api.mux.com/robots/v0/jobs/find-key-moments \
  --user "$MUX_TOKEN_ID:$MUX_TOKEN_SECRET" \
  --header "Content-Type: application/json" \
  --data '{
    "parameters": {
      "asset_id": "YOUR_ASSET_ID",
      "max_moments": 5,
      "target_duration_ms": { "min": 30000, "max": 60000 },
      "output_steering": {
        "selection_strategy": "standalone_hooks",
        "title_style": "descriptive",
        "rubric_priorities": ["clarity_in_isolation", "soundbite_quality"]
      }
    }
  }'
```

The response contains a job in `pending` status, with its ID in `data.id`. Save that ID with your episode record. The duration range and steering are best-effort preferences; inspect the returned ranges before using them.

In an application, handle `robots.job.find_key_moments.completed`. Its `event.data.outputs.moments` contains the results. For a one-off script, retrieve the job using its ID:

```bash
curl https://api.mux.com/robots/v0/jobs/find-key-moments/YOUR_JOB_ID \
  --user "$MUX_TOKEN_ID:$MUX_TOKEN_SECRET"
```

Read `data.outputs.moments` only when `data.status` is `completed`. While the job is `pending` or `processing`, keep its ID and wait. Stop on `errored` or `cancelled`, and report the failure. If you poll, use a bounded wait with backoff and resume the saved job when the wait expires.

Store the results you need in your application: [Robots jobs are deleted after 30 days](https://www.mux.com/docs/guides/robots.md#how-it-works).

## 3. Review the cuts

Each candidate includes a title, a summary of the speech, transcript cues, a score, and `start_ms` / `end_ms` on the source episode's timeline. Results are ordered by position in the episode. The score can help prioritize review, but it does not predict how a clip will perform on social media.

Follow the configured editorial policy. Present each candidate with its preview, transcript, source boundaries, and duration. If the workflow calls for human selection, save the candidates and resume when choices are recorded; previews are for review, so create no clip assets before then. If the user has delegated selection, apply their criteria and record the reasons for the chosen cuts. In either case, check that:

* The opening makes sense without the preceding conversation, and the ending completes the thought.
* The cut preserves the speaker's meaning and does not interrupt a word or sentence. Suggested boundaries can land mid-sentence. Caption cues often break mid-sentence, so a boundary on a cue edge can still cut a thought short. When adjusting one, start at a cue where a sentence begins and end at a cue whose text finishes a sentence, with a fraction of a second of padding, rather than estimating a time inside a cue. Cue times are as precise as the captions; request word-level timestamps from premium captions when cuts need to be tighter.
* The excerpt suits the intended audience and avoids unwanted ads, introductions, or repeated moments.
* The title accurately describes what is said, and captions spell names and domain terms correctly.

**Robots returns milliseconds; Mux clipping uses seconds.** For example, an illustrative candidate from `120500` to `165750` milliseconds becomes a clip from `120.5` to `165.75` seconds. Convert once and retain fractional seconds:

```javascript
const startSeconds = moment.start_ms / 1000;
const endSeconds = moment.end_ms / 1000;
```

Validate that the times are finite, the start is at least zero, the end is after the start, and the end does not exceed the source duration. Asset-based clips must be at least 0.5 seconds long. If an editor changes the range, save their final boundaries separately from the original suggestion.

Use [instant clipping](https://www.mux.com/docs/guides/create-instant-clips.md) to preview candidates before creating new assets. For an episode with a public playback ID, a preview URL looks like:

```text
https://stream.mux.com/YOUR_SOURCE_PLAYBACK_ID.m3u8?asset_start_time=120.5&asset_end_time=165.75
```

For signed playback, put those modifiers in the playback JWT as described in [signed instant clips](https://www.mux.com/docs/guides/create-instant-clips.md#via-signed-urls). Instant clipping uses segment-level boundaries, so check the exact cut again on the finished asset. A public instant-clip URL also lets viewers change the range; it does not restrict access to the rest of the episode.

If the job returns no useful moments, show that outcome and allow manual selection or a revised request. Creating fewer good clips is a valid result.

## 4. Create the selected clips

Create clips only for ranges the user selected, or that you selected under criteria the user delegated. Create a new asset for each selected range. Use the **source asset ID** in `mux://assets/...`, with the reviewed times in seconds. The example below creates a private clip with an MP4 rendition; replace the illustrative ID and times with your selected candidate.

```bash
curl https://api.mux.com/video/v1/assets \
  --user "$MUX_TOKEN_ID:$MUX_TOKEN_SECRET" \
  --header "Content-Type: application/json" \
  --data '{
    "inputs": [
      {
        "url": "mux://assets/YOUR_SOURCE_ASSET_ID",
        "start_time": 120.5,
        "end_time": 165.75
      }
    ],
    "playback_policies": ["signed"],
    "video_quality": "basic",
    "static_renditions": [{ "resolution": "highest" }]
  }'
```

Use `basic` video quality as a starting point; choose a different [quality level](https://www.mux.com/docs/guides/use-video-quality-levels.md) if your output needs it. Omit `static_renditions` when you only need streaming playback, or use `audio-only` for an M4A export of an audio podcast.

Save the new asset ID as soon as the request succeeds. Associate it with your episode, source asset ID, selected boundaries, and title. Mux sets `source_asset_id` on the clip, but application metadata such as `passthrough` is not automatically copied.

[Mux copies and trims the source's text tracks](https://www.mux.com/docs/guides/create-clips-from-your-videos.md#my-source-asset-has-subtitlescaptions-text-tracks-will-the-clip-have-them) to the new clip's timeline. Finish preparing the source captions before creating clips, and verify the copied tracks on each result. You do not need to regenerate captions for every clip.

## 5. Wait for playback and download readiness

Track these separately:

| Result | Ready when | What to return |
| :-- | :-- | :-- |
| Streaming clip | The clip asset has `status: "ready"` | Its playback ID and a player or HLS URL using the chosen playback policy |
| Downloadable file | The requested static rendition has `status: "ready"` | A download URL using the clip's playback ID and the rendition's returned `name` |

Use `video.asset.ready` for streaming readiness and `video.asset.static_rendition.ready` for the file. Handle asset and rendition errors, and treat a skipped rendition as an unavailable export. See [static rendition status and webhooks](https://www.mux.com/docs/guides/enable-static-mp4-renditions.md#webhooks).

An asset can be playable while its MP4 is still preparing. Show that distinction in your UI. Build the file URL only when the requested rendition is ready:

```text
https://stream.mux.com/YOUR_CLIP_PLAYBACK_ID/RETURNED_RENDITION_NAME
```

For the signed playback policy used above, sign the MP4 request as described in [signed static rendition URLs](https://www.mux.com/docs/guides/enable-static-mp4-renditions.md#signed-static-rendition-urls). Generate expiring URLs when needed rather than persisting them as permanent links.

### Adapt clips for social video

Asset clipping preserves the source framing. A landscape podcast needs an additional rendering step if you want a portrait crop, a layout with both speakers, animated captions, or other graphics. CSS cropping in a player changes the display, not the exported video.

Caption tracks provide selectable subtitles in supported players. They do not burn text into the MP4's picture. If the destination needs visible captions, render them into the video and check their timing against the clip's timeline, which starts at zero. Use an editor or rendering service for this step, and upload the finished file to Mux if it needs streaming playback.

Check the exported file itself for framing, audio, readable captions, and the destination's format requirements before publishing it.

## Run on a schedule

Use this section when the user asks to clip new episodes repeatedly, run the workflow on a cron schedule, or make it reusable as an agent skill. Save the configuration and procedure in a durable place that the runner can read without the original chat.

### Save the run configuration

Record the episode source and selection rule, schedule and timezone, Mux environment, caption and selection preferences, output format, playback policy, review or publishing policy, destination, and per-run processing limit. Reference credentials from the execution environment's secret store. Include the prompt URL and the version or revision of your implementation so a later run can be reproduced.

Choose a trigger that fits the source:

* **New-episode event:** Use a CMS event or upload webhook to enqueue a specific episode. [Directives](https://www.mux.com/docs/guides/robots-directives.md) can coordinate caption preparation and moment selection within Mux.
* **Scheduled check:** Use cron, a hosted scheduler, or the agent runtime's scheduling capability to discover new episodes and resume unfinished ones. Read the current configuration and run records every time.
* **Manual repeat:** Reuse the installed skill or saved procedure when the user supplies another episode. A skill by itself does not create a schedule or keep a process running.

Use an existing scheduler when the project has one. When the user requests scheduling and the environment supports it, configure and verify the trigger within that request's scope. If a scheduler is unavailable, provide the runnable procedure and exact configuration still needed; do not claim the recurring run is active.

### Resume each episode from saved state

For a one-off run, a saved manifest of job IDs, clip IDs, chosen ranges, and statuses is enough to resume work. In an application, keep that state in your database and update it from [verified webhooks](https://www.mux.com/docs/core/verify-webhook-signatures.md).

For each run:

1. Load the saved configuration, discovery cursor, and unfinished episode records. Identify episodes using a stable source ID, such as a feed GUID or CMS entry ID, and track revisions when the media changes.
2. Claim each episode with an atomic lock or lease so overlapping runs cannot process it twice. Apply the per-run limit before starting more work.
3. Resume the first incomplete stage: source ingestion, captions, moment selection, editorial selection, clip creation, or export verification. Save new resource IDs immediately and preserve selected boundaries across retries.
4. Follow the configured editorial policy. A run awaiting selection can leave candidates in a review queue and finish; the next run can resume after selections are recorded. Automatically select or publish only when that behavior is part of the user's configured workflow.
5. Record completed outputs and failures, advance discovery state after new episodes have been durably recorded, and release the claim. A pending job remains resumable even if the discovery cursor moves on.

Handle duplicate and out-of-order webhook deliveries. Use a unique application record for each episode revision, selected range, and output configuration so repeated events do not create duplicate clips. A `passthrough` value can help correlate requests; it is not an idempotency guarantee.

If an API request times out after submission, reconcile its result before retrying creation. If only one clip fails, retry that work without re-running successful caption or moment-selection jobs. Bound concurrent requests and respect [Mux API rate limits](https://www.mux.com/docs/core/make-api-requests.md#api-rate-limits).

Set a retry limit and record failures that need attention. An expired browser login, unavailable source file, or failed renderer should leave a resumable record with a clear next action. For a recurring agent, notify the user when clips are ready for review, a run fails, or input is required; avoid sending an unchanged status on every check unless requested.

### Verify the repeat before enabling the schedule

Run the procedure on one representative episode, then run it again with the same input. The second run should reuse completed work. Exercise an interrupted run and an overlap between two invocations to confirm that state is resumed and duplicate resources are not created. Verify that the scheduler can access the same configuration, credentials, and state store outside the interactive session.

For the handoff, include the procedure or skill location, trigger, timezone, configuration and state locations, a sample run result, and how to pause the schedule or retry a failed episode. Separate a configured schedule from a successfully verified scheduled execution.

## Verify the result

Before calling the workflow complete, check that:

* Every clip traces back to the correct episode and the selected source range; milliseconds were converted to seconds exactly once.
* Playback works with the intended access policy, and the opening and ending preserve the speaker's meaning.
* Caption tracks are present and synchronized on the finished clip.
* Each requested download is actually ready and accessible. A playable HLS stream alone does not satisfy an MP4 request.
* Replaying an event or resuming a run does not create duplicate clips or restart successful jobs.
* Pending, failed, cancelled, and empty-result cases have a clear outcome in the application.

For an agent-built implementation, report which of these were tested with an episode and which still need credentials, media, or editorial review. Local checks alone do not establish that a real export plays correctly.

## Understand the costs

The workflow can incur source and clip storage, playback or download delivery, Robots processing, and static rendition charges. Each asset-based clip is a separate asset. Captioning once, reviewing previews before creating assets, and requesting MP4s only when needed avoids unnecessary work. Use the [pricing overview](https://www.mux.com/docs/pricing/overview.md) and [cost-estimation guide](https://www.mux.com/docs/pricing/estimating-video-costs.md) with your episode length, clip count, retention, and audience.
