Back to blog
Guidepodcastshort videoYouTube ShortsTikTokAI video generation

How to Turn Your Podcast into Short Videos — A Complete Guide

Convert podcast audio into vertical short videos for TikTok and YouTube Shorts — full AI pipeline specs, credit costs, a decision table to the right deep-dive article, and honest guidance on when this workflow doesn't fit.

2026-03-15Faceless.fm Team

Converting podcast audio into short videos takes three steps: clip selection → visual generation → FFmpeg composition. This guide covers the full pipeline with real specs, credit costs, and a decision table pointing you to the right deep-dive article for your specific needs.

TL;DR

  • Input: MP3/WAV/M4A (≤200 MB and up to 90 minutes) or a podcast RSS feed URL
  • AI surfaces up to 5 clip candidates with viral scores
  • 4 visual styles: sketchnote / cinematic / flat_graphic / manga
  • Output: 1080×1920 9:16 MP4, subtitle position Top/Bottom selectable
  • Free plan: 50 credits/month — enough to test the full workflow end-to-end

Why Podcasts Need Short Video

Audio content is deep, but discoverability is hard. Short video algorithms actively surface content to non-followers, making vertical clips an efficient acquisition channel for new listeners. The key advantage for podcasters: you already have the raw material.

  • Repurpose existing episodes — no filming or scripting required
  • Multiple posts from one audio file — up to 5 clip candidates per episode
  • One MP4, three platforms — YouTube Shorts, TikTok, Instagram Reels
  • Which Article Should You Read? — Decision Table

    Your situationStart here
    "I want to go from MP3 to a finished video right now"Audio-Only Podcast to Short Video
    "I want to process multiple RSS episodes in bulk"Podcast RSS to Shorts and Articles
    "I want to understand the faceless video strategy first"What Are Faceless Videos?

    Step 1: Upload Your Audio

    Accepted input:

    • File: MP3 / WAV / M4A — maximum 200 MB and 90 minutes
    • RSS: paste a feed URL → get up to 100 episodes listed (episodes over 90 minutes are auto-excluded)
    For RSS imports, select the episodes you want to process individually or use batch selection.

    Step 2: STT Transcription + Clip Analysis — 5 Credits

    After upload, Google Cloud Speech-to-Text v2 (model: chirp, word-level timestamps) transcribes the full audio. Then Gemini 3 Flash analyzes the transcript and returns up to 5 clip candidates, each containing:

    FieldDescription
    viralScoreVirality score (0–100)
    reasonWhy this clip works
    start / endTimestamps in seconds
    hookTextOpening hook line (reuse for subtitles or captions)
    recommendedStyleSuggested visual style

    Step 3: Visual Generation — 5 or 15 Credits

    Gemini 3 Flash Image generates still images based on each clip's text content. Choose from 4 styles:

    StyleLookBest for
    sketchnoteHand-drawn illustrationEducational, how-to
    cinematicFilm-like, atmosphericTravel, storytelling, documentary
    flat_graphicClean vector designBusiness, tech, data
    mangaComic-panel styleEntertainment, narrative
    Generation modes:

  • Standard (1 frame) — 5 credits: A single image per clip
  • HQ Storyboard (3 frames) — 15 credits: Hook → Evidence → Payoff across three frames, each with an auto-composed headline overlay. Significantly better at stopping the scroll and driving saves.
  • Step 4: Video Generation — 0 Credits

    FFmpeg composites the final video to these specs:

    SpecValue
    Resolution1080 × 1920 (9:16)
    CodecH.264 / AAC
    SubtitlesASS format, Top or Bottom position
    Video modeimages / waveform / icon
    OutputMP4 (download)
    Subtitle position Top keeps captions clear of platform UI elements that appear at the bottom of the screen; Bottom is the standard placement.

    Bonus: Article Generation — 3 Credits

    From the same transcript, Faceless.fm can generate long-form posts for X / note / LinkedIn. Run video and article generation from a single episode and ship content across every channel at once.

    Credit Reference

    ActionCredits
    Analyze (STT + clip selection)5
    Visual generation — standard 1 frame5
    Visual generation — HQ 3-frame storyboard15
    Video generation0
    Article generation3
    Monthly limit (Free)50
    Monthly limit (Pro, ¥2,980/mo)300

    Who This Is For

  • Podcasters: Convert existing episodes directly to shorts — no new recording required
  • Creators who dislike filming: No camera, lighting, or video editing software needed
  • High-frequency posters: Up to 5 clips per episode means multiple posts per week from one audio file
  • RSS-heavy publishers: Batch-import from a feed URL and process multiple episodes at scale
  • Multi-channel publishers: Generate video and written posts from the same transcript in a single session
  • When It Might Not Be the Right Fit

  • Live or breaking-news content: STT → AI analysis → generation takes time; real-time reaction content does not fit this pipeline
  • Brand campaigns with strict visual guidelines: The 4 built-in AI styles may not match a specific brand system; fully custom design needs a different tool
  • Poor-quality source audio: Low-quality recordings reduce STT accuracy, which degrades clip candidate and hook-text quality downstream
  • Episodes over 90 minutes: Split the recording before import so each source stays within the processing window
  • Related Articles

  • Turn an Audio-Only Podcast into Short Videos (No Footage Needed)
  • Automate Your Podcast RSS into Shorts and Articles
  • What Are Faceless Videos?
  • Try It on Faceless.fm

    https://faceless-fm.com — Free plan includes 50 credits/month. Upload an MP3 and see what clip candidates the AI surfaces in minutes. Upgrade to Pro (¥2,980/month) for 300 credits and batch RSS processing.

    Frequently Asked Questions

    How many short videos can I make from one podcast episode?

    Faceless.fm's AI surfaces up to 5 clip candidates per episode. A 30–60 minute episode typically yields 3–5 usable clips. Varying the visual style or subtitle position across candidates gives you even more posts without any additional recording.

    Do I need to show my face or record video?

    No. Upload an audio file and the pipeline outputs a finished 9:16 MP4 with subtitles and AI visuals — no camera, no lighting, no editing software needed.

    Which platforms can I post to?

    The output is 1080×1920 (9:16) H.264/AAC MP4, compatible with YouTube Shorts, TikTok, and Instagram Reels. You can also switch the subtitle position (Top/Bottom) to avoid each platform's on-screen UI elements.

    How many credits does it cost?

    Analyze (STT + clip selection) costs 5 credits. Visuals cost 5 credits for a 1-frame standard or 15 credits for an HQ 3-frame storyboard. Video generation is 0 credits. Article generation is 3 credits. Free plan: 50 credits/month. Pro: 300 credits/month (¥2,980/month).

    Can I import from a podcast RSS feed?

    Yes. Paste your podcast RSS URL to get a list of up to 100 episodes. Episodes over 90 minutes are automatically skipped. You can select individual episodes or batch-import several at once.

    Ready to try Faceless.fm?

    Just upload your audio content and let AI automatically generate short videos.

    Get Started Free