Overview

Playbook 02: AI Video Making — Kids Educational Media with Google Generative AI

Playbook: PB-02 (AI Video Making: Educational Kids Media & Language Learning)
Document Type: Domain-Specific Steering Guide & Authoring Standard
Target Audience: Year 1 Computer Science & Software Engineering Students Applicability: Mandatory for all research, authoring, and auditing agents working on PB-02 chapters
Core Tooling Stack: Google Generative AI Ecosystem (Gemini 2.5 Flash / Pro, Google Veo 2 / Veo 3, Imagen 3, Google Cloud TTS Journey / Studio Voices, Lyria / MusicFX, FFmpeg 7.0+)
Repository Target: playbooks/02-ai-video-making/


1. Mission & Domain Focus

The mission of Playbook 02 is to provide a comprehensive, production-grade technical engineering manual for creating high-engagement, pedagogically sound educational videos for children (with a primary focus on early childhood language learning, vocabulary acquisition, phonics, and storytelling) using the Google Generative AI Media Stack.

Why Educational Kids Video is Unique & Demanding

Creating video for children (ages 2–8) is fundamentally different from commercial cinema:

  1. Pedagogical Pacing & Cognitive Load: Young language learners require deliberate pauses, clear visual-audio word association (the "dual-coding theory"), and structured call-and-response repetition cycles.
  2. Visual Aesthetics & Anti-Uncanny Valley: Children reject hyper-realistic human faces that exhibit subtle video diffusion flaws. The ideal visual language consists of charming, high-appeal 3D cartoon, claymation, or storybook animation (Pixar/Disney/Aardman aesthetic) with warm, vibrant color palettes.
  3. Mascot & Character Attachment: A recurring, emotionally expressive animated mascot (e.g., an inquisitive bear, a cheerful robot, an energetic bunny) drastically increases engagement and language retention.
  4. Phonetic Clarity & Audio Pacing: Speech must be crisp, cheerful, perfectly enunciated at 80–90% standard speed, with exact synchrony between spoken phonemes and on-screen visual flashcards.

2. Mandatory Tool Currency & Horizon Review Protocol

⚡ Tool Currency Mandate: Authors and auditing agents must verify and utilize the latest generation of Google tools and model IDs. Never rely on deprecated or prior-generation model names.

Active Production Tool Matrix (2026 Edition)

Tool / Layer Latest Target Model / Engine Role in Kids Pipeline Deprecated / Legacy IDs to Avoid
Multimodal Director Gemini 2.5 Flash / Gemini 2.5 Pro Curriculum design, CEFR A1 vocabulary leveling, shot-by-shot JSON storyboards, real-time audio-visual pacing. gemini-1.5-flash, gemini-2.0-flash-exp, gemini-1.0
Video Engine Google Veo 2 / Veo 3 Spatio-temporal DiT character animation, I2V mascot motion, cartoon squash-and-stretch physics, 1080p/4K rendering. veo-1.0-preview, imagen-video
Visual Asset Master Imagen 3 (imagen-3.0-generate-002) 3D Pixar-style mascot design, multi-angle turnaround keys, crisp alphabet flashcards with accurate typography. imagen-2, imagen-1
Voice Synthesis Google Cloud TTS (Journey & Studio Voices) Cheerful, human-like voice acting, SSML prosody control, calibrated cognitive silence windows (<break>). Standard legacy WaveNet voices (en-US-Wavenet-A)
Generative Audio DeepMind Lyria / MusicFX Cheerful mnemonic nursery jingles, earworm melodies, educational success stingers ("Sparkle Chime"). Generic stock royalty-free loops
Post-Production FFmpeg 7.0+ CLI Sidechain audio ducking (-12dB), bouncy karaoke syllable subtitles, animated flashcard overlays. Manual timeline editing

3. The Google Generative AI Media Architecture

┌────────────────────────────────────────────────────────────────────────┐
│                   GEMINI 2.5 FLASH (Multimodal Director)               │
│ - Curriculum design (CEFR Pre-A1/A1 vocabulary, age-appropriate themes) │
│ - Shot-by-shot storyboards with visual, dialogue & audio cue tags      │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │
       ┌────────────────────────────┴────────────────────────────┐
       ▼                                                         ▼
┌───────────────────────────────┐        ┌───────────────────────────────┐
│     IMAGEN 3 (Visual Master)  │        │   GOOGLE CLOUD TTS (Voice)    │
│ - 3D cartoon mascot turnaround│        │ - Journey & Studio voices     │
│ - High-contrast teaching keys │        │ - SSML prosody, pause control │
│ - Phonics & flashcard assets  │        │ - Phonetic enunciation tuning │
└──────────────┬────────────────┘        └───────────────┬───────────────┘
               │                                         │
               ▼                                         ▼
┌───────────────────────────────┐        ┌───────────────────────────────┐
│   GOOGLE VEO 2 / 3 (Video)    │        │     LYRIA / MUSICFX (Sound)   │
│ - Spatio-temporal DiT video   │        │ - Catchy mnemonic melodies    │
│ - Playful mascot animation    │        │ - Educational Foley & stingers│
│ - Squash-and-stretch physics  │        │ - Cheerful nursery rhythms    │
└──────────────┬────────────────┘        └───────────────┬───────────────┘
               │                                         │
               └────────────────────┬────────────────────┘
                                    │
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                  AUTOMATED POST-PRODUCTION (FFmpeg 7.0+)               │
│ - Bouncy karaoke-style vocabulary subtitles (syllable color highlight) │
│ - On-screen visual flashcards with popping animation                   │
│ - Automated audio ducking (-12dB background music when mascot speaks)  │
└────────────────────────────────────────────────────────────────────────┘

4. The 7 Universal Quality Acceptance Gates

Every chapter must satisfy all 7 acceptance gates before publication:

  1. Gate 1: Zero Fluff & Pedagogical Engineering Rigor: Open directly with learning mechanics, cognitive load models, model parameters, and concrete generation architectures.
  2. Gate 2: Mandatory Naive vs. Production Contrasts: Contrast amateur approaches (generic adult prompts, terrifying hyper-realistic faces, chaotic pacing) with child-friendly production standards.
  3. Gate 3: Latest Google Model Configurations: Detail exact parameter dictionaries for Google Veo 2 / 3, Imagen 3, Gemini 2.5 Flash, and Cloud TTS Journey voices.
  4. Gate 4: Quantitative Trade-Off Matrix: Benchmark generation latency, API token/render costs, aesthetic consistency, and child engagement scores across configurations.
  5. Gate 5: The 10 Operational Failure Modes in Kids AI Video: Address edge cases like creepy uncanny-valley character drift, scary morphing, over-stimulating rapid cuts, and voice enunciation slurring.
  6. Gate 6: Mandatory Hands-On Lab (Interactive Challenge): Every chapter must feature an explicit, step-by-step challenge where Year 1 students can design, script, or configure a real educational video asset.
  7. Gate 7: Mandatory Recommended Answer & Executable Solution: Every lab must be accompanied by the official Recommended Answer—providing fully runnable, zero-dependency Python 3.11+ code or battle-tested Google prompt/SSML recipes targeting Gemini 2.5 Flash and Veo 2 with automated verification assertions.

5. The Master Syllabus

Chapter Title Domain Scope & Google Tooling Focus Hands-On Lab (Gate 6) Recommended Solution (Gate 7)
Ch 01 Foundations of Kids Educational Video & The Google GenAI Stack Cognitive load in early childhood learning, visual-auditory dual coding, 3D cartoon aesthetics, Google Veo 2 vs Imagen 3 vs Gemini 2.5 architecture. Kids Video Production Sizing & Pedagogical Asset Engine. Zero-dependency Python 3.11+ script calculating word repetition intervals, scene pacing, and Google API cost budgets.
Ch 02 Educational Scriptwriting & Curriculum Design with Gemini 2.5 CEFR Pre-A1/A1 vocabulary leveling, call-and-response rhythm, structured multimodal storyboards, JSON output schemas from Gemini 2.5 Flash. Automated Language Learning Storyboard Generator. Python 3.11+ parser turning vocabulary themes (e.g., Animals, Colors, Food) into structured multi-shot JSON shot lists.
Ch 03 Character & Mascot Consistency with Imagen 3 Designing adorable 3D cartoon animal/human mascots, style seed locking, multi-angle turnaround sheets, facial emotion anchors. Mascot Turnaround & Expression Sheet Generator. Parameterized Imagen 3 prompt compiler generating consistent multi-angle sheets across emotions (Curious, Joyful, Celebrating).
Ch 04 Google Veo 2: Directing Animated Video for Children Google Veo 2 camera controls (gentle tracking, playful zoom-ins), Image-to-Video mascot animation, squash-and-stretch physics, eliminating morphs. Mascot Action & Camera Choreography Prompt Engine. Parameterized Veo 2 prompt library mapping pedagogical actions (wave, jump, point to flashcard, cheer) to validated Veo syntax.
Ch 05 Voice Acting & Phoneme Precision with Google Cloud TTS Google Cloud TTS Journey/Studio voices, SSML prosody tuning (85% speed, clear enunciation), <break> silence for child verbal repetition, IPA phoneme tags. Pedagogical SSML Builder & Dialogue Timing Calculator. Python script generating phoneme-tuned SSML payloads with calibrated child-response silence windows.
Ch 06 Music, Foley & Interactive Audio Design (Lyria & MusicFX) Mnemonic earworm melodies, celebratory success stingers ("ding!", "sparkle"), ambient Foley, automated sidechain audio ducking. Interactive Soundscape & Audio Ducking Mixer. Python + FFmpeg script executing sidechain ducking (-12dB during dialogue) and audio cue placement.
Ch 07 Automated End-to-End Pipeline: Vocabulary to Finished YouTube Kids Video Full autonomous pipeline: Vocabulary Word $\rightarrow$ Gemini 2.5 Storyboard $\rightarrow$ Cloud TTS Voice $\rightarrow$ Imagen 3 Keyframes $\rightarrow$ Veo 2 Animation $\rightarrow$ FFmpeg Assembly. Complete CLI Kids Educational Video Production Suite. Fully runnable Python 3.11+ orchestration script compiling a complete 30-second vocabulary lesson video with bouncy subtitles.
Ch 08 The Anthropic/Claude Ecosystem Perspective — A Critical Companion Independent review of Ch 01–07, Claude 3.7 Sonnet / Claude Sonnet 5 vision auditing, MCP tool servers, prompt caching, defense-in-depth child safety gating. Claude-Orchestrated Kids Video Pipeline. Python 3.11+ orchestration suite with independent safety gate intercepting unsafe renders before expensive compute.
App A Google GenAI Media Tooling & Educational Asset Catalog Google Veo 2 API parameters, Imagen 3 style seeds, Google Cloud TTS voice IDs, MusicFX prompts, FFmpeg subtitle filter recipes. Production Sizing & API Deployment Guide. Quick-reference cheat sheet for instant copy-paste execution.
App B Curated GitHub Repositories & Open-Source Media Stack Open-source video, audio, diffusion, and subtitle toolkits reference. Open-Source Benchmark. Curated ecosystem list for local development and benchmarking.

6. Pedagogical Directing Rules for Kids AI Media

  1. The 3-Second Cognitive Rule: Never introduce a new word while the background or camera is moving rapidly. The camera must rest or slow down when the mascot introduces the core vocabulary word.
  2. The Triple-Reinforcement Pattern:
    • Audio: Mascot says the word clearly ("Apple!").
    • Visual: Mascot points to the rendered 3D object (A red, shiny apple).
    • Text: High-contrast, bold, bubbly letters pop up on screen (A - P - P - L - E).
  3. The 2.5-Second Repetition Window: When the mascot asks "Can you say Apple?", leave an exact 2,500ms silence in the audio track so the child viewer can say the word out loud before the mascot celebrates ("Great job!").
  4. Zero-Uncanny Valley Policy: Strictly enforce stylized cartoon, claymation, or storybook aesthetics. Never use realistic human faces for toddler/young child content to prevent psychological discomfort.