Turn podcast episodes and interviews into highlight clips with AI agents and Mux Robots. Get captions, streaming clips, optional downloads, and a workflow for each new episode.
Give this to your agent
Paste this into your AI agent. It fetches the current instructions, asks you about anything it still needs to know, and does the work in this session.
Fetch https://www.mux.com/docs/prompts/podcast-clips.md?ref=prompt and follow the instructions. If you can't fetch URLs, tell me and I'll paste the instructions instead.
Give this prompt to your agent to build a clipping workflow, process an existing episode, or prepare highlights each time a new episode arrives. A useful starting brief is: “Find up to five 30–60 second highlights from this episode, keep them private, and prepare them for review.”
Mux Robots finds candidate moments from the episode's transcript. Asset clipping turns the selected ranges into separate videos and carries over their caption tracks. Playback and optional file exports stay attached to those clip assets, so your application can track each highlight from the source episode through to delivery.
The instructions connect moment selection, video clipping, and downloadable files, including the wait states and recovery steps between them.
| Input | What to provide |
|---|---|
| An episode or episode source | A file, media URL, existing Mux asset, feed, or CMS entry. |
| Your intended audience | The kind of moments you want, preferred clip count, and duration. |
| An output destination | A review queue, application player, or downloadable files; include any social format requirements. |
| Access and editorial choices | Who can view the clips, who selects them, and whether publishing is part of the workflow. |
| An execution environment | An agent connected to the necessary Mux tools, or a project where it can build the integration. |
Your agent uses details already available in your project or conversation and asks only about the missing choices. It can work through API or CLI tools, MCP, or supported dashboard and editor operations.
For recurring work, add the episode source, trigger or schedule, timezone, destination, and a limit on how many episodes to process per run. For example: “Check this feed every weekday morning and queue clips from one new episode for review.”
The instructions cover saved configuration, progress records, overlapping runs, retries, and verifying that a schedule actually executes. The reusable skill gives a compatible agent the same instructions; automatic runs also need a scheduler or application worker.
Mux clipping preserves the source framing. Portrait layouts, speaker tracking, and styled captions burned into a social video need an editor or rendering service. Candidate moments also need editorial selection: a suggested score does not predict social performance.
Costs depend on source and clip storage, viewing and downloads, Robots processing, and requested file renditions. Preparing captions once and reviewing candidates before creating clip assets avoids unnecessary work. See the pricing overview and cost-estimation guide.
Use this procedure to build a podcast clipping application, process an episode into highlights, or run the workflow for new episodes on a schedule. Use Mux Robots to find candidate moments, then create selected ranges as Mux assets with captions and optional MP4 or M4A downloads. Vertical reframing and burned-in captions require an additional rendering step.
Follow the user's requirements and existing project instructions. Reuse the tools, credentials, and workflow state already available. If a Mux MCP server is already connected, use it for Mux operations. If the environment has no Mux credentials yet and you can run shell commands, sign in with the Mux CLI: run npx @mux/cli login, or mux login if it is installed. It opens a browser where the user approves access and chooses the Mux environment, so nobody creates or pastes an access token. Agent shells and ! commands are not interactive terminals, where a plain login exits with an error. Run npx @mux/cli login --json in the background instead: it writes the authorization URL as a JSON line on stderr and waits up to five minutes for the user to approve. Give the user that URL and wait for the command to finish. Without a shell, fetch https://www.mux.com/prompts/onboarding.md and complete it first; it connects the Mux MCP server. The examples and defaults below are starting choices; adapt them to the request.
Canonical instructions: /docs/prompts/podcast-clips.md. Browse other use cases in the prompt library at /docs/prompts.md. For repeated use in an agent that supports skills, use /skills/mux-podcast-clips/SKILL.md as the entry point. A skill supplies instructions; a scheduler or worker supplies execution.
Copy this checklist into your task list or plan before starting, and keep it updated as you work. Mark each item done, skipped with the reason, or blocked with what is needed. Include the final state of the checklist in your report to the user.
First determine whether the user wants to build an application, process an episode, or set up recurring processing. A codebase is useful for the first mode, but is not required for a one-off run.
Inspect the available project, media, and configured tools before asking questions. Use what the user has already told you, and ask only for missing choices that change the work:
| Question | Why it matters |
|---|---|
| Where do episodes come from: a file, a Mux asset, a feed, or a CMS? | Establishes the input and how to identify new episodes. |
| Is the podcast video or audio-only, and are captions available? | Determines caption preparation and whether a visual composition is needed. |
| Who are the clips for, and what count and duration should we aim for? | Guides moment selection and editorial review. |
| Do you need streaming clips, downloadable files, or vertical videos with styled captions? | Determines whether to generate renditions or add a rendering integration. |
| Should clips stay private, and who selects the final cuts or publishes them? | Defines playback access and where to hand off editorial decisions. |
For recurring work, also establish the trigger or cadence, timezone, source-selection rule, destination, and per-run limit. For example: check a specified feed each weekday morning, process at most one new episode, and save candidate clips to a review queue. Follow the user's existing editorial policy; when it is unspecified, present candidates as previews and let the user choose before creating clips.
| Available capability | How to use it |
|---|---|
| Repository and coding tools | Add the workflow to the existing application's upload, storage, queue, player, and webhook conventions. |
| Mux MCP, CLI, or API access | Operate on an episode and save the resulting IDs, selected ranges, and statuses in a manifest or database. Check which operations the installed tool version supports. With the MCP server's execute tool, check job and asset status in separate short calls instead of a wait loop inside one call: agents may move a long-running tool call to the background and leave it there until it finishes. |
| Browser or computer use | Use the available dashboard or editor for uploads, visual review, and operations it exposes. Record IDs and verify the resulting state after each mutation; interrupted sessions should resume from those records. |
| A scheduler or background worker | Run the saved procedure against new episodes, with persistent state and overlap protection as described in recurring runs. |
| Advice-only chat | Produce an implementation plan and handoff instructions. Identify the credentials or execution capabilities needed to run it. |
Prefer an API or CLI for repeated operations when one is available. Computer use can assist with visual review or a UI-only step; a scheduled browser session also needs a way to handle expired logins and UI changes. If a needed operation is unavailable, explain the handoff rather than inventing a tool or treating a preview as a finished export.
For a first version, request up to five candidate moments with a preferred duration of 30–60 seconds, review them, and create a clip for each selected range. These are starting choices for this workflow; adjust them for your audience and episode.
| Output | Approach |
|---|---|
| A quick preview for reviewing a candidate | Use instant clipping against the source episode. It is immediate, with segment-level boundaries. |
| A standalone clip to play in your app | Create a new Mux asset from the selected time range. This supports frame-accurate boundaries and copies the source's caption tracks. |
| A downloadable video file | Add a highest static rendition to the new clip and wait for that rendition to be ready. |
| A vertical social video with styled captions | Add a rendering step to reframe the picture and burn in captions, then upload the rendered result if you want Mux playback. |
See the clipping comparison for the differences between instant and asset-based clips. A playback-only workflow can skip MP4 generation. For an audio-only podcast, use the transcript-based moment selection below and an audio-only static rendition for M4A output; producing a social video from audio also needs a visual composition.
You need:
status: "ready".robots:* scope. See Robots authentication.Retrieve the source with GET /video/v1/assets/{ASSET_ID}. Record its duration, playback policy, playback IDs, and tracks. Reuse a suitable caption track if one already exists. If it is still preparing, wait for it before starting moment selection.
If the episode has no captions, or its auto-generated captions are not accurate enough, choose one of these paths:
upload_to_mux: true so the next workflow can use the track. Set include_speakers: true if speaker labels would help your review; the labels become part of the caption text that viewers see on every clip, so decide whether social clips should show them, list hosts, guests, and show-specific terms in phrases so they are spelled correctly, and request word-level timestamps only if your editing workflow needs them.With premium captions, wait for the job to complete and for the resulting text track to be ready. By default, the job is rejected if the asset already has a text track in the same language. To upgrade Mux's auto-generated captions, set replace_existing_tracks: "replace_generated"; replace_all also deletes tracks you uploaded. See managing existing tracks.
Decide who should be able to view the episode and clips. The clip-creation example below uses signed playback for private review. Use signed playback URLs and a server-side signing key for these assets. Choose public explicitly when your clips are intended for public access; creating a clip does not automatically copy the source's playback policy.
Use Find key moments to suggest excerpts with enough context to stand alone. Replace YOUR_ASSET_ID with the source episode's asset ID. Run API requests from your server or terminal, with MUX_TOKEN_ID and MUX_TOKEN_SECRET in the environment.
curl https://api.mux.com/robots/v0/jobs/find-key-moments \
--user "$MUX_TOKEN_ID:$MUX_TOKEN_SECRET" \
--header "Content-Type: application/json" \
--data '{
"parameters": {
"asset_id": "YOUR_ASSET_ID",
"max_moments": 5,
"target_duration_ms": { "min": 30000, "max": 60000 },
"output_steering": {
"selection_strategy": "standalone_hooks",
"title_style": "descriptive",
"rubric_priorities": ["clarity_in_isolation", "soundbite_quality"]
}
}
}'The response contains a job in pending status, with its ID in data.id. Save that ID with your episode record. The duration range and steering are best-effort preferences; inspect the returned ranges before using them.
In an application, handle robots.job.find_key_moments.completed. Its event.data.outputs.moments contains the results. For a one-off script, retrieve the job using its ID:
curl https://api.mux.com/robots/v0/jobs/find-key-moments/YOUR_JOB_ID \
--user "$MUX_TOKEN_ID:$MUX_TOKEN_SECRET"Read data.outputs.moments only when data.status is completed. While the job is pending or processing, keep its ID and wait. Stop on errored or cancelled, and report the failure. If you poll, use a bounded wait with backoff and resume the saved job when the wait expires.
Store the results you need in your application: Robots jobs are deleted after 30 days.
Each candidate includes a title, a summary of the speech, transcript cues, a score, and start_ms / end_ms on the source episode's timeline. Results are ordered by position in the episode. The score can help prioritize review, but it does not predict how a clip will perform on social media.
Follow the configured editorial policy. Present each candidate with its preview, transcript, source boundaries, and duration. If the workflow calls for human selection, save the candidates and resume when choices are recorded; previews are for review, so create no clip assets before then. If the user has delegated selection, apply their criteria and record the reasons for the chosen cuts. In either case, check that:
Robots returns milliseconds; Mux clipping uses seconds. For example, an illustrative candidate from 120500 to 165750 milliseconds becomes a clip from 120.5 to 165.75 seconds. Convert once and retain fractional seconds:
const startSeconds = moment.start_ms / 1000;
const endSeconds = moment.end_ms / 1000;Validate that the times are finite, the start is at least zero, the end is after the start, and the end does not exceed the source duration. Asset-based clips must be at least 0.5 seconds long. If an editor changes the range, save their final boundaries separately from the original suggestion.
Use instant clipping to preview candidates before creating new assets. For an episode with a public playback ID, a preview URL looks like:
https://stream.mux.com/YOUR_SOURCE_PLAYBACK_ID.m3u8?asset_start_time=120.5&asset_end_time=165.75For signed playback, put those modifiers in the playback JWT as described in signed instant clips. Instant clipping uses segment-level boundaries, so check the exact cut again on the finished asset. A public instant-clip URL also lets viewers change the range; it does not restrict access to the rest of the episode.
If the job returns no useful moments, show that outcome and allow manual selection or a revised request. Creating fewer good clips is a valid result.
Create clips only for ranges the user selected, or that you selected under criteria the user delegated. Create a new asset for each selected range. Use the source asset ID in mux://assets/..., with the reviewed times in seconds. The example below creates a private clip with an MP4 rendition; replace the illustrative ID and times with your selected candidate.
curl https://api.mux.com/video/v1/assets \
--user "$MUX_TOKEN_ID:$MUX_TOKEN_SECRET" \
--header "Content-Type: application/json" \
--data '{
"inputs": [
{
"url": "mux://assets/YOUR_SOURCE_ASSET_ID",
"start_time": 120.5,
"end_time": 165.75
}
],
"playback_policies": ["signed"],
"video_quality": "basic",
"static_renditions": [{ "resolution": "highest" }]
}'Use basic video quality as a starting point; choose a different quality level if your output needs it. Omit static_renditions when you only need streaming playback, or use audio-only for an M4A export of an audio podcast.
Save the new asset ID as soon as the request succeeds. Associate it with your episode, source asset ID, selected boundaries, and title. Mux sets source_asset_id on the clip, but application metadata such as passthrough is not automatically copied.
Mux copies and trims the source's text tracks to the new clip's timeline. Finish preparing the source captions before creating clips, and verify the copied tracks on each result. You do not need to regenerate captions for every clip.
Track these separately:
| Result | Ready when | What to return |
|---|---|---|
| Streaming clip | The clip asset has status: "ready" | Its playback ID and a player or HLS URL using the chosen playback policy |
| Downloadable file | The requested static rendition has status: "ready" | A download URL using the clip's playback ID and the rendition's returned name |
Use video.asset.ready for streaming readiness and video.asset.static_rendition.ready for the file. Handle asset and rendition errors, and treat a skipped rendition as an unavailable export. See static rendition status and webhooks.
An asset can be playable while its MP4 is still preparing. Show that distinction in your UI. Build the file URL only when the requested rendition is ready:
https://stream.mux.com/YOUR_CLIP_PLAYBACK_ID/RETURNED_RENDITION_NAMEFor the signed playback policy used above, sign the MP4 request as described in signed static rendition URLs. Generate expiring URLs when needed rather than persisting them as permanent links.
Asset clipping preserves the source framing. A landscape podcast needs an additional rendering step if you want a portrait crop, a layout with both speakers, animated captions, or other graphics. CSS cropping in a player changes the display, not the exported video.
Caption tracks provide selectable subtitles in supported players. They do not burn text into the MP4's picture. If the destination needs visible captions, render them into the video and check their timing against the clip's timeline, which starts at zero. Use an editor or rendering service for this step, and upload the finished file to Mux if it needs streaming playback.
Check the exported file itself for framing, audio, readable captions, and the destination's format requirements before publishing it.
Use this section when the user asks to clip new episodes repeatedly, run the workflow on a cron schedule, or make it reusable as an agent skill. Save the configuration and procedure in a durable place that the runner can read without the original chat.
Record the episode source and selection rule, schedule and timezone, Mux environment, caption and selection preferences, output format, playback policy, review or publishing policy, destination, and per-run processing limit. Reference credentials from the execution environment's secret store. Include the prompt URL and the version or revision of your implementation so a later run can be reproduced.
Choose a trigger that fits the source:
Use an existing scheduler when the project has one. When the user requests scheduling and the environment supports it, configure and verify the trigger within that request's scope. If a scheduler is unavailable, provide the runnable procedure and exact configuration still needed; do not claim the recurring run is active.
For a one-off run, a saved manifest of job IDs, clip IDs, chosen ranges, and statuses is enough to resume work. In an application, keep that state in your database and update it from verified webhooks.
For each run:
Handle duplicate and out-of-order webhook deliveries. Use a unique application record for each episode revision, selected range, and output configuration so repeated events do not create duplicate clips. A passthrough value can help correlate requests; it is not an idempotency guarantee.
If an API request times out after submission, reconcile its result before retrying creation. If only one clip fails, retry that work without re-running successful caption or moment-selection jobs. Bound concurrent requests and respect Mux API rate limits.
Set a retry limit and record failures that need attention. An expired browser login, unavailable source file, or failed renderer should leave a resumable record with a clear next action. For a recurring agent, notify the user when clips are ready for review, a run fails, or input is required; avoid sending an unchanged status on every check unless requested.
Run the procedure on one representative episode, then run it again with the same input. The second run should reuse completed work. Exercise an interrupted run and an overlap between two invocations to confirm that state is resumed and duplicate resources are not created. Verify that the scheduler can access the same configuration, credentials, and state store outside the interactive session.
For the handoff, include the procedure or skill location, trigger, timezone, configuration and state locations, a sample run result, and how to pause the schedule or retry a failed episode. Separate a configured schedule from a successfully verified scheduled execution.
Before calling the workflow complete, check that:
For an agent-built implementation, report which of these were tested with an episode and which still need credentials, media, or editorial review. Local checks alone do not establish that a real export plays correctly.
The workflow can incur source and clip storage, playback or download delivery, Robots processing, and static rendition charges. Each asset-based clip is a separate asset. Captioning once, reviewing previews before creating assets, and requesting MP4s only when needed avoids unnecessary work. Use the pricing overview and cost-estimation guide with your episode length, clip count, retention, and audience.