# How to Automate Viral TikTok Videos

Source: https://www.youtube.com/watch?v=VdpBgNJ5GaM
Recap page: https://rapidrecap.app/video/VdpBgNJ5GaM
Generated: 2025-10-10T13:35:07.033+00:00

---
## Quick Overview

The creator concludes that current AI tools cannot fully automate the creation of high-quality, educational science explainer TikTok videos because AI struggles to generate accurate and relevant visual B-roll, necessitating manual sourcing of stock footage even when music and lyrics are automated.

**Key Points:**
- The creator aims to automate the production of viral science explainer TikToks, similar to accounts like 'Learning with Lyrics,' using AI tools for research, lyrics, music, and visuals.
- For lyric generation, the creator uses a specific prompt structure in ChatGPT, advising to include 'It does not have to rhyme' to maintain scientific accuracy over poetic flow.
- Suno generates the preferred AI music, but the creator notes Suno lacks an API, presenting a hurdle for full automation workflows, although Suno allows creating custom personas for consistent style.
- Initial attempts to generate accurate visuals using Nvidia's stock footage finder, Sora 2, and Leonardo AI (V3 and Cling models) resulted in unusable or illogical footage that did not match the script.
- The creator found that sourcing existing stock footage and using CapCut's auto-captioning feature was necessary to complete the first manual video, concluding AI imagery is 'just not quite there yet' for accurate explainers.
- An automation attempt using the Glyph agent framework successfully strung together steps like music generation and subtitle addition but yielded only a short, visually weak 10-second clip, requiring manual extension.
- A subsequent automation attempt using Mind Studio, while executing all steps including text-to-speech from ElevenLabs, produced a song and visuals, but the visuals were inferior to the manually sourced ones, reinforcing the need for manual editing or better B-roll sourcing.

**Context:** The video explores the feasibility of fully automating the creation of popular, viral TikTok videos that explain scientific concepts through autotuned, catchy songs, inspired by successful creator accounts. The process involves five main stages: research, lyric writing, music creation, visual sourcing, and final assembly, with the creator attempting to integrate various leading AI tools like ChatGPT, Suno, Nvidia, Sora 2, Leonardo, Perplexity, Glyph, and Mind Studio to streamline this multi-step workflow.

## Detailed Analysis

The creator attempts to reverse-engineer and automate the production of educational, viral TikTok songs, starting with research and lyric generation using a targeted ChatGPT prompt designed to yield clear, non-metaphorical explanations, such as how a car engine works. For music, Suno is selected as the best current generator, allowing the creator to establish a consistent 'Tik Tok education songs' persona, despite Suno's lack of API integration hindering full automation. The most significant challenge proved to be visual sourcing; attempts with Nvidia's features, Sora 2, and Leonardo AI failed to produce visuals accurate enough to explain complex processes like engine mechanics, often resulting in nonsensical or irrelevant clips. When stuck, the creator consulted Perplexity, which suggested tools like Opus Clip. Testing Opus Clip's AI B-roll feature against stock B-roll showed that neither AI-sourced nor stock footage automatically matched the script content well, leading to the conclusion that current AI image/video generation is inadequate for accurate explainer visuals. To finish the first video, the creator manually sourced stock footage and used CapCut's auto-captioning. Subsequently, the creator tested automation platforms Glyph and Mind Studio; Glyph generated a short, poorly visualized clip, and Mind Studio successfully chained steps (research, lyrics, TTS) but still failed to produce visuals matching the quality of manually placed stock footage, confirming that while much of the process can be automated, achieving high-quality, accurate B-roll still requires significant manual intervention or superior stock sourcing.

### Manual Video Creation Steps

- Research topics using targeted prompts in ChatGPT
- Write simple, non-metaphorical poems for lyrics
- Generate catchy music using Suno and establish a consistent persona
- Source visuals (initially failing with AI, resorting to stock footage)
- Assemble in CapCut using auto-captioning.

### Music Generation with Suno

- Suno generates the creator's favorite AI music but lacks an API for direct workflow integration
- The 'create and make persona' feature ensures stylistic consistency across subsequent songs.

### Visual Sourcing Failures

- Nvidia generative AI and Sora 2 failed to create videos matching the script's technical steps (e.g., piston movement)
- Leonardo AI V3 and Cling models failed to animate sourced still images correctly, producing 'funky animation' or going 'off the rails'.

### AI B-roll Testing

- Using Opus Clip's AI B-roll feature on an existing video export yielded irrelevant footage, such as showing a swimming person when describing compression, proving AI cannot yet reliably match visuals to scientific narration.

### Automation Attempt 1

- Glyph Agent: Built a custom agent that strung together music generation, image generation (Cling), subtitles, and overlay, resulting in a short, visually unconvincing 10-second clip that required manual extension.

### Automation Attempt 2

- Mind Studio Workflow: Successfully automated steps from topic input through lyric generation (ElevenLabs TTS) and assembly, but the resulting visuals were not as good as the manually sourced ones, confirming the visual bottleneck.

### Final Conclusion on Automation

- AI can handle the audio and text components efficiently, but for educational explainers requiring specific, accurate imagery, manually sourcing B-roll remains faster and yields superior results compared to current fully automated visual generation.

