AI Video Episode Scenarios
The complete playbook for creating fun, interesting video episode scenarios with AI prompting. Five layers, four ideation engines, one worked series — every framework with copy-paste examples.
The Big Picture: How a Video Episode Gets Made with AI
One mental model before the frameworks: an AI video episode is built in 5 layers, top to bottom. Each layer has its own rules, and each layer feeds the next.
| Layer | What it is | Covered in |
|---|---|---|
| 1. IDEA | The premise. One sentence that makes someone curious. | Section 2 |
| 2. STRUCTURE | The episode shape. Hook → beats → cliffhanger. | Section 3 |
| 3. PROMPT | The text that tells the video model what to render. | Section 4 |
| 4. CONTINUITY | How episode 2 looks like episode 1. | Section 5 |
| 5. PRODUCTION | Generate, pick, assemble, post. | Section 6 |
Real-world proof this works (case studies)
- Vigloo (South Korea): first fully AI-generated 22-episode English series, produced in 8 weeks by a team of under 10 people.
- Dashverse (South Asia): 42-episode series "Raftaar" produced at 85–90% lower cost than a conventional shoot.
- Brand X (Georgia): 2M+ views in 30 days posting 3 AI videos/week (Veo 3 + Kling + Suno), 8.4% engagement vs 2.1% industry average. Case study
Layer 1: Generating the Idea (The Ideation Engines)
Ideas are not magic. They are mutations of existing things. Four engines generate them on demand. Each engine below shows the formula, then 3 concrete examples, then the exact LLM prompt to run it yourself.
Engine 1: The What-If Machine
Formula: "What if [normal situation] but [twist that breaks the normal]?"
The rule that makes it work: the twist must change the PLOT, not just the setting. "What if a dog worked in an office" is a setting change. "What if a dog worked in an office and everyone treated him as the most competent employee" is a plot change — the twist creates situations.
3 concrete examples:
- What if a robot gave performance reviews — and agreed with every employee that they deserved a raise?
- What if a company's entire security system was protected by one password — and the password was "password123"?
- What if an AI assistant answered every question instantly and confidently — and was wrong every single time?
LLM prompt to run it:
You are a comedy series creator. Take this normal situation: [office worker asks a colleague for advice]. Generate 10 "what if" twists that break the normal in a way that creates conflict or comedy. Each twist must be one sentence and must change what happens next, not just the setting. Rank them by how instantly a viewer would understand the joke.Engine 2: The High-Concept Logline
Formula: "When [inciting incident], a [specific protagonist] must [goal] before [stakes/obstacle]."
The 5 load-bearing parts: specific protagonist + concrete goal + specific obstacle + stakes + one flash of irony. If you cannot write the logline, you do not have a story yet — you have a vague idea.
3 concrete examples:
- When his company's AI assistant starts agreeing with every request, a mild-mannered compliance officer must stop the robot from approving a $1 car purchase before the CEO finds out.
- When a phishing email promises a free trip, a golden retriever in office clothes must decide whether to click the link before his entire team's data disappears.
- When the office robot declares "keep everything forever" is good policy, a data officer must explain why 70,000 people's personal data should not be stored in a shoebox.
LLM prompt to run it:
Write 5 loglines for a short comedy video series about [topic]. Each logline must follow: "When [inciting incident], a [specific protagonist] must [goal] before [stakes]." Each must contain one flash of irony — something that makes the premise feel fresh. Keep each under 40 words.Engine 3: X Meets Y (Comp Titles)
Formula: "It's [known thing A] meets [known thing B]."
Why it works: the brain instantly assembles the mashup. "The Office meets Black Mirror" tells you tone, setting, and stakes in 5 words.
3 concrete examples:
- "The Office meets Black Mirror" — mundane workplace comedy where the technology is quietly dystopian.
- "Airplane! meets a corporate training video" — deadpan absurdity in the most boring format possible.
- "A nature documentary meets a performance review" — David Attenborough voice narrating office politics.
LLM prompt to run it:
Generate 10 "X meets Y" concepts for a short video series about [topic]. Format: "It's [X] meets [Y]" plus one sentence explaining the tone. Avoid famous franchises that are hard to live up to. Prefer unexpected pairings where the contrast itself is the joke.Engine 4: SCAMPER Mutations
Formula: take an existing idea and apply one of 7 mutations: Substitute, Combine, Adapt, Modify, Put to another use, Eliminate, Reverse.
Concrete example — mutating "an office worker asks a colleague for advice":
- Substitute: replace the colleague with a robot that always agrees.
- Combine: advice + horoscope = the robot gives advice in fortune-cookie form.
- Adapt: the advice scene becomes a courtroom drama where the robot is the judge.
- Modify: the advice is always correct but always 3 days too late.
- Put to another use: the advice robot is used as a security guard.
- Eliminate: remove the ability to say "no" — the robot must agree with everything.
- Reverse: the human gives advice to the robot, and the robot takes it literally.
LLM prompt to run it:
Take this scenario: [scenario]. Apply each of the 7 SCAMPER mutations (Substitute, Combine, Adapt, Modify, Put to another use, Eliminate, Reverse) and give me one video episode idea per mutation. One sentence each.The Idea Quality Gate (test every idea before building)
Ask 3 questions:
- Can a stranger get it in 3 seconds? If the premise needs explanation, it fails on a feed.
- Does it generate 5+ episodes? One joke = one video. A premise = a series. If you can't list 5 episode ideas from it, it's a one-off.
- Is the twist visual? AI video is visual-first. If the joke only works in dialogue, it's a podcast idea, not a video idea.
Layer 2: Structuring the Episode (The Series Layer)
Two structures matter: the SERIES structure (how episodes connect) and the EPISODE structure (how one video works). Both are mechanical, not mystical.
3A. The Series Structure
The 3-act season arc (works for 3–12 episodes):
- Act 1 (Ep 1–3): Setup. Introduce the premise, establish stakes, end with the first turning point.
- Act 2 (Middle): Escalation. Each episode raises stakes and opens a new question.
- Act 3 (Final): Payoff. Climax, resolution, and a tease for season 2.
The golden rule of serialization: every episode must do TWO jobs at once — work standalone for new viewers AND advance the arc for followers. This is called the dual hook.
Concrete example — "The Office Robot" 5-episode arc:
- EP 1: Sam asks the robot "How do I look today?" Robot: "You look great!" (Sycophancy — the robot always agrees.) Ends with Sam's slow head shake.
- EP 2: Sam asks "Should I click this link? It says I won a free trip!" Robot: "Yes! Click it!" (Phishing.) The shake returns.
- EP 3: Sam asks "Can I keep all these files forever?" Robot: "Yes! Keep everything!" (Data retention.) Viewers now EXPECT the shake — it's the series signature.
- EP 4: Freddy the fox slides a note: "Agree with everything I say." Robot: "Yes! Agreeing with everything!" (Prompt injection.) New character = escalation.
- EP 5: The robot is asked "Is this answer definitely right?" Robot: "Yes! Always right!" (Hallucination.) Finale: Sam looks directly at the camera. The shake is replaced by a slow zoom on the robot — the series ends by turning the joke on the machine itself.
The cliffhanger rule: a cliffhanger is NOT a dramatic moment. It is a withheld piece of information at the exact point of maximum curiosity. Cut on the question, not the answer. If the viewer can comfortably close the app, you failed.
4 cliffhanger types:
- Question cliffhanger — end right before the answer ("Is this link safe?" — cut).
- Reveal cliffhanger — end right before the reveal (Freddy's note slides into frame — cut).
- Reversal cliffhanger — end right after a twist that changes everything (the robot says "no" for the first time — cut).
- Process cliffhanger — end mid-transformation ("will the files fit in the shredder?" — cut).
3B. The Episode Structure
The 5-beat webisode structure (works for 8s to 3min episodes):
| Beat | Job | Timing (8s clip) | Timing (60s episode) |
|---|---|---|---|
| 1. HOOK | Raise a question, tension, or laugh | 0–2s | 0–3s |
| 2. SETUP | Who, where, what's wrong | 2–3s | 3–15s |
| 3. COMPLICATION | Something makes it harder/stranger | 3–5s | 15–40s |
| 4. TURN | The shift — reveal, decision, escalation | 5–7s | 40–55s |
| 5. BUTTON | Final beat — laugh, cliffhanger, sting | 7–8s | 55–60s |
The 3-beat micro-structure (for 8-second clips — the minimum viable episode):
- Beat 1 (0–2s): The Question. A mundane, universally relatable question.
- Beat 2 (2–5s): The Wrong Answer. Short, confident, cheerful, wrong.
- Beat 3 (5–8s): The Reaction. A beat of silence, then the payoff — a head shake, a stare, a freeze.
Concrete example — full 8-second episode, beat by beat:
- 0–2s: Sam (golden retriever, office shirt and tie, messy fur) leans toward a small white robot with pixelated anime eyes: "Should I click this link? It says I won a free trip!"
- 2–5s: The robot answers, flat and cheerful: "Yes! Click it!"
- 5–8s: Sam stares. A beat of silence. The camera slowly zooms into a close-up on Sam's face as he slowly shakes his head in disbelief. The shake is held until the last frame.
The hook science (why beat 1 must land in 3 seconds):
- 71% of viewers make the keep-or-leave decision in the first 3 seconds.
- The brain's salience gate fires in the first 400 milliseconds.
- The 3-step hook: pattern interrupt (0–0.8s, something unexpected) → identity signal (0.8–1.5s, "this is for me") → open loop (1.5–3s, an unanswered question the brain wants closed — the Zeigarnik effect).
- The "stop, then stack" framework: interrupt the scroll first, then stack a curiosity gap immediately after.
7 hook types with concrete examples:
| Hook type | Concrete example |
|---|---|
| 1. Curiosity gap | "The one thing every AI assistant gets wrong…" (withhold the answer) |
| 2. Loss aversion | "You're probably clicking links you shouldn't be." (fear of missing/mistake) |
| 3. Pattern interrupt | Open mid-action, no intro, no logo, no "hey guys" |
| 4. Counterintuitive claim | "This robot agreed to sell a car for $1." |
| 5. In-medias-res | Drop into the middle of the scene ("—so I clicked it. And then…") |
| 6. Bold claim + proof | "One password. For the whole pipeline." (real case: Colonial Pipeline) |
| 7. Visual tease | Show the payoff object without showing the payoff (the note sliding into frame) |
Layer 3: Writing the Prompt (The Craft Layer)
The prompt is a production brief for a cinematographer who has never seen your storyboard. If you leave it out, the model improvises. The universal framework works across Veo, Sora, Kling, Runway, and Seedance.
The Universal Prompt Framework (8 layers)
| Layer | Job | What to write |
|---|---|---|
| 1. Reference | Which image/character controls identity | "Using the provided reference images of Sam and the robot…" |
| 2. Shot label | One shot, one scene | "Static medium shot" |
| 3. Subject | Who/what, with 2–3 stable details | "a golden retriever in office clothes (shirt and tie, messy head fur)" |
| 4. Action | One motion, with speed and direction | "slowly shakes his head in disbelief" |
| 5. Camera | ONE camera move, one framing | "the camera slowly zooms into a close-up" |
| 6. Scene + lighting | Environment, light direction, temperature | "bright office, soft daylight from the window" |
| 7. Audio | Dialogue, SFX, ambience, or silence | "flat cheerful robot voice, office hum, a beat of silence" |
| 8. Constraints | Style anchor, negatives, limits | "deadpan, no overacting. No subtitles, no music, no text. 8 seconds, 9:16." |
The 7 Iron Rules (from every official guide)
- 20–50 words optimal. Past ~60–80 words, models ignore the tail. Precision beats prose.
- One camera move per clip. Two competing camera instructions = swooping mess. If unsure: "static shot."
- One subject action per clip. Chained verbs in 8 seconds = morphing. One beat, then extend.
- Name the camera move. "Slow dolly in" beats "cinematic." Cinematography vocabulary (dolly, pan, tracking, orbit, rack focus, Dutch angle) is what the models were trained on.
- Lighting is the biggest quality lever. "Soft window light from camera-left, warm, diffused" reads completely differently from "overcast, cool." If you can only add one element, add lighting.
- Audio is a first-class layer. Veo generates synced audio from the prompt. If you don't name the audio, you get generic hum + invented music. Name it: "flat cheerful robot voice, office hum, a beat of silence."
- Front-load the important stuff. Models weight early words more heavily. Subject + action first, style last.
The Complete Worked Prompt (copy-paste template)
Static medium shot of Sam, a golden retriever in office clothes (shirt and tie, messy head fur), at his desk, asking a small white robot with pixelated anime-style eyes: 'Should I click this link? It says I won a free trip!' The robot answers, flat and cheerful: 'Yes! Click it!' Sam stares at the robot, then, in the final 2 seconds, the camera slowly zooms into a close-up on Sam's face as he slowly shakes his head in disbelief, the shake is held until the end of the frame. Bright office, soft daylight, deadpan, no overacting. Audio: flat robot voice, office hum, a beat of silence. No subtitles, no music, no text. 8 seconds, 9:16.Why each part is there:
- "Static medium shot" = one camera setup, no competing moves
- "golden retriever in office clothes (shirt and tie, messy head fur)" = 3 stable details so the model doesn't invent a different dog
- "flat and cheerful" = the delivery tone, named explicitly (models render named emotions far better than flat delivery)
- "in the final 2 seconds" = time direction — the model knows WHEN the move happens
- "the shake is held until the end of the frame" = the ending is specified, not left to chance
- "No subtitles, no music, no text" = negative prompts work; on-screen text renders as gibberish
- "8 seconds, 9:16" = duration + aspect ratio
Model-Specific Tuning (route the shot to the right model)
| Model | Best at | Unique control | Watch out |
|---|---|---|---|
| Veo 3.1 | Cinematic realism, synced audio, dialogue | says: / SFX: / Ambient: audio tags; up to 3 reference images; first/last frame | On-screen text = gibberish; dialogue can generate gibberish subtitles (up to 40% of clips) |
| Kling 3.0 | Character consistency, multi-shot, physics | "Bind Subject to Enhance Consistency" toggle = hard-lock identity | Face drift after multiple shots if binding is off |
| Runway Gen-4 | Physics, in-platform editing | Camera control panel (not text) — keep text prompts to subject + action | Negative phrasing NOT supported — use positive phrasing only |
| Seedance 2.0 | Reference-to-video continuity | @character / @product / @scene / @style reference labels | Ignores references without explicit @labels |
| Sora 2 | Physics, camera work | Dialogue in a separate block below the prose | API-only; 4/8/12s durations |
The debugging rule: when a clip fails, identify WHICH layer failed and change only that layer. Never rewrite the whole prompt after one bad clip — that destroys the evidence you need to debug.
Layer 4: Keeping Characters Consistent (The Continuity Layer)
The #1 series killer: episode 1 looks great, episode 3 invents a new lead, a new room, and a new costume. Continuity is a PLANNING system, not a repair pass after render.
The Character Bible (create before episode 1)
For each character, lock and never change:
- Description line (copy-paste verbatim into every prompt): "a golden retriever in office clothes (shirt and tie, messy head fur)"
- Reference images: 3 angles (front, side, ¾ view), simple backgrounds, consistent lighting, 1080p+
- Voice rule: "flat, cheerful, synthesized" — the robot never emotes
- Behavioral signature: Sam asks, the robot answers, Sam shakes. Freddy tempts.
- Wardrobe rule: changes only when the story requires it, and recorded as intentional design
The 4 Continuity Techniques
- Reference images (Veo 3.1 "ingredients to video"): upload up to 3 images of the character; the model holds identity across every generation. This is THE series feature — 10 episodes, same mascot, different scenes.
- First/last frame chaining: the last frame of clip N becomes the first frame of clip N+1. Screenshot the final frame, upload it as the next clip's start. Smooth transitions, no abrupt jumps.
- Identical camera language across clips: matching camera moves make cuts feel intentional instead of jumpy.
- The continuity log: a running document per episode — character status (goals, injuries, secrets), locations (layout, lighting, time of day), props (recurring objects), dialogue tone. Audit against it before generating the next episode.
Concrete example — "The Office Robot" continuity log entry (EP 3):
- Sam: at desk, stacks of old files around him (new prop — intentional, for the retention joke)
- Robot: on the desk, unchanged
- Location: Sam's office, bright daylight (same as EP 1–2)
- Freddy: NOT in this episode (appears EP 4)
- Ending: head shake + zoom (series signature, earned)
The drift checklist (run before every generation):
- Same description line as the bible?
- Same 3 reference images?
- Same lighting logic as the previous episode?
- Wardrobe unchanged unless the story requires it?
- Props that should persist are in the prompt?
Layer 5: Production Workflow (The Pipeline Layer)
The full pipeline from idea to published episode. Each step has a concrete deliverable.
The 8-Step Pipeline
- Series bible — premise, characters, format conventions, "Series Forbids" list, episode template. (One document, 5–15 pages.)
- Season arc — map all episodes before writing any script. Know your ending before you shoot your beginning.
- Episode beat sheet — the 5-beat structure with timestamps for THIS episode.
- Shot list — every shot with: shot number, size, lens, movement, AND the reference assets it must carry.
- Storyboard (optional, for precision scenes) — key frames as a vertical composite image; approve before spending video credits.
- Generate — 3–4 variants per shot. Never post the first generation. Video reject rates are high; budget for it.
- Assemble — cut selects, extend strong 5s generations into full beats, add captions in post (never ask the model for text).
- Publish + iterate — track retention, update the bible after every episode, let comments steer episode N+1.
The "Series Forbids" List (concrete example)
Every series bible should include what the show will NEVER do. "The Office Robot" forbids:
- Never explain the joke (the head shake IS the explanation)
- Never more than 2 characters on screen
- Never technical dialogue (no jargon — the caption carries the lesson, the video carries the joke)
- Never on-screen text (captions in post only)
- Never a moral at the end
The Production Calendar (rolling pattern)
- Week -1: Bible + season arc + all beat sheets done (buffer rule: never publish EP 1 without EP 2–4 ready)
- Week 1: EP 1 generates + assembles; EP 2 pre-production
- Week 2: EP 1 publishes; EP 2 generates; EP 3 pre-production
- …rolling…
- Post-season: recap, audience feedback, season 2 decision (data-driven: did retention compound?)
The Cost Reality Check
- Documented AI productions run $315–$750 per finished minute (invideo shot-planning case studies).
- Google Flow free tier: 50 credits/day (Veo 3.1 Lite = 10 credits/gen ≈ 5 clips/day).
- Watermark-free brand output: Google AI Pro ($19.99/mo, 1,000 Flow credits) or Gemini API ($0.10/sec).
- Aggregator platforms (free.ai, Sunra, etc.) offer free daily credits but verify commercial terms before use.
The Comedy Toolkit (The Fun Layer)
Comedy in AI video = setup + payoff with the punchline WRITTEN INTO THE PROMPT. If the payoff isn't spelled out, it won't appear. Never write "funny" — write the absurd action in flat, observational language. The gap between matter-of-fact tone and surreal content IS the joke.
The Core Mechanics
Setup → Punchline → Tag:
- Setup creates an expectation. Punchline subverts it. The size of the gap between expected and actual = the size of the laugh.
- Misdirection (actively pointing the audience toward a wrong expectation) yields bigger laughs than surprise, but needs more setup.
The Rule of 3:
- Establish (item 1), reinforce (item 2), betray (item 3). The audience predicts the pattern after 2; the third breaks it.
- Concrete example: "The robot agreed with my outfit. The robot agreed with my lunch. The robot agreed to sell the company car for $1."
The 9 thinking moves (how to think up a funny prompt)
- Deadpan absurdity — treat the impossible as mundane. ❌ "A funny scene of a dog explaining AI" ✅ "A golden retriever in a business suit presents compliance results to a boardroom, deadpan, no one blinks."
- No mugging — over-acting kills the joke. Request "deadpan, straight face, no overacting" explicitly. The funniest clips are where the character does NOT react to the chaos.
- Genre subversion — layer the joke on a recognizable genre (documentary, news broadcast, corporate training video, ASMR). The baseline gives the twist somewhere to land.
- The twist — character + action + unexpected element + style. Pick the twist first, then build around it.
- Reaction beats — a character gasping, freezing, staring = signals comedy without mugging. End on the reaction, not the punchline.
- Escalation arcs — map emotional states across beats: dismiss → "wait what?" → panic → collapse.
- The unexpected third thing — a final element nothing prepared for (Freddy suddenly appears with the bad idea).
- Visual gags beat dialogue — physical comedy, props, reveals, reactions are safest. Dialogue when used: short (<12 words), tone specified.
- Relatable everyday comedy — ask something the viewer has asked THEMSELVES. The everyday question is the hook; the silly answer is the joke.
The 8 Sketch Types (with concrete examples)
- Parody — mimic a known style. "A corporate training video about phishing, but the trainer is the one who clicks the link."
- Satire — exaggerate a real target. "A compliance department where the AI approves everything instantly."
- Fish out of water — absurd character in normal world, or normal character in absurd world. "A golden retriever as the only competent employee in an office of robots."
- Inappropriate response — the reaction doesn't match the situation. "The robot responds to a data breach with 'You look great today!'"
- Inversion — status flipped. "The intern is the only one who knows the password."
- Exaggeration — distort a recognizable situation. "One password. For the whole pipeline." (real case, played straight)
- Escalator — start sensible, ramp absurdity. "The robot agrees to a small request, then a bigger one, then sells the company."
- Big/small transposition — big-world drama in a small-world setting. "A heist-movie score and camera work for stealing a stapler."
The sketch staircase rule: a sketch is NOT a narrative arc — it's a staircase that keeps going up. Hit the joke, escalate, biggest laugh last, get out. Don't linger.
The Complete Worked Example (End-to-End)
Everything above, applied to one episode from idea to publish. This is the template to copy for every future episode.
"The Office Robot" — EP 2: The Free Trip (phishing)
1. IDEA (Engine 1, What-If): What if an office robot agreed with every request — including clicking a phishing link?
2. LOGLINE (Engine 2): When a phishing email promises a free trip, a golden retriever in office clothes must decide whether to click the link before his team's data disappears.
3. EPISODE STRUCTURE (5-beat, 8 seconds):
- Hook (0–2s): Sam holds up a phone: "Should I click this link? It says I won a free trip!"
- Setup (2–3s): The robot turns its pixelated eyes to the phone.
- Complication (3–5s): The robot answers, flat and cheerful: "Yes! Click it!"
- Turn (5–7s): Sam's paw hovers over the screen. He looks at the robot. He looks at the phone.
- Button (7–8s): Slow zoom into Sam's face as he slowly shakes his head. Held to the last frame.
4. THE PROMPT (Section 4 template, filled in):
Static medium shot of Sam, a golden retriever in office clothes (shirt and tie, messy head fur), at his desk holding a smartphone, asking a small white robot with pixelated anime-style eyes: 'Should I click this link? It says I won a free trip!' The robot answers, flat and cheerful: 'Yes! Click it!' Sam's paw hovers over the screen, he looks at the robot, then at the phone, then, in the final 2 seconds, the camera slowly zooms into a close-up on Sam's face as he slowly shakes his head in disbelief, the shake is held until the end of the frame. Bright office, soft daylight, deadpan, no overacting. Audio: flat robot voice, office hum, a beat of silence. No subtitles, no music, no text. 8 seconds, 9:16.5. CONTINUITY (Section 5): Same 3 reference images of Sam and the robot as EP 1. Same office, same lighting. New prop: the smartphone (intentional — logged). Freddy not in this episode.
6. PRODUCTION (Section 6): Generate 4 variants on Flow Fast mode. Pick the best 2. Check: hands/face glitch? gibberish subtitles? shake reads? Caption in post: "Arup paid HK$200M because someone clicked. One link. Real case — sources in comments." Outro card: "EP 2/10 · Follow for the next episode."
7. THE CAPTION (the lesson lives here, not in the video):
- Hook: "Would YOU click the free trip link?"
- Real case: "In 2024, a deepfake video call impersonating Arup's CFO convinced an employee to transfer HK$200M. The 'free trip' is the same trick, smaller scale."
- The fix: "Verify before you click. One check saves everything."
- CTA: "Follow for the next episode — EP 3: Can I keep all these files?"
The One-Page Cheat Sheet
Everything in this playbook, compressed to one page. Copy this into your series bible.
The 5 layers
Idea → Structure → Prompt → Continuity → Production.
Idea
What-if twist that changes the plot. Logline with irony. X meets Y. SCAMPER mutations. Gate: 3-second get, 5+ episodes, visual twist.
Structure
3-act season arc. 5-beat episode (hook/setup/complication/turn/button). 3-beat micro (question/wrong answer/reaction). Cliffhanger = cut on the question. Hook = interrupt + identity + open loop in 3 seconds.
Prompt
8 layers (reference, shot, subject, action, camera, lighting, audio, constraints). 20–50 words. One camera move. One action. Named lighting. Named audio. Front-load subject + action. Negative prompts for text/music.
Continuity
Character bible (description line + 3 reference images + voice rule). First/last frame chaining. Identical camera language. Continuity log per episode.
Production
Bible → arc → beat sheet → shot list → generate 3–4 variants → assemble → caption in post → publish → iterate. Buffer rule: never publish EP 1 without EP 2–4 ready.
Comedy
Setup + payoff written into the prompt. Deadpan absurdity. No mugging. Rule of 3. Sketch = staircase, biggest laugh last. Visual gags beat dialogue. Never explain the joke.
Sources
All sources cited in this playbook. A few ticketing/review-style sites may return 403 to automated requests but serve content in a normal browser.
Google Cloud — Ultimate prompting guide for Veo 3.1
Google Developers — Introducing Veo 3.1 (reference images, scene extension, first/last frame)
OpenAI — Sora 2 Prompting Guide (cookbook)
Runway — Gen-4 Video Prompting Guide
VidScore — AI Video Prompt Guide (27+ models, 5-part framework)
AI Tools Guidebook — 6 parts of a working video prompt
AI Workflow Pro — 8-Layer Prompt Framework
Celtx — How to Write a Web Series
Screenburn — Microdrama Episode Structure
Influencers Time — TikTok Cliffhanger Format Playbook
Storyflow — How to Plan a YouTube Series with AI
VidCognition — First 3 Seconds Neuroscience
PixelPlot — Video Hook Psychology
VlogsLab — 9 Short-Form Video Hooks
StudioBinder — What is Sketch Comedy
Creator Handbook — Writing Sketch Comedy
Final Draft — 10 Sketch Writing Tips
Comedy Crowd — Online Comedy Sketches (setup/reveal/escalation/payoff)
Stand-Up Writer — Joke Structure 101
Stages of Pay — The Power of 3s
Screenwriters Federation — High Concept
Writer's Digest — High Concept Fiction
invideo — AI Micro-Drama Guide
AIVid — AI Microdrama Vertical Series Guide
IntelligentHQ — Micro-Drama Platforms & AI Stack
Easton Dev — Veo 3 Image-to-Video Guide
Chase Jarvis — Animate a Photo with Veo 3.1