Overview
Chapter 2: Educational Scriptwriting & Curriculum Design with Gemini 2.5
Playbook: PB-02 (AI Video Making Playbook: Educational Kids Media & Language Learning)
Target Audience: Year 1 Computer Science & Software Engineering Students Prerequisites: Chapter 1: Foundations of Kids Educational Video & The Google GenAI Stack, Python 3.11+
Steering Reference: `instruction.md` — Featuring Hands-On Labs & Tested Solutions
Active Google Model Focus: Gemini 2.5 Flash (gemini-2.5-flash) / Gemini 2.5 Pro (gemini-2.5-pro)
2.0 The Big Picture: LEGO Instruction Manuals vs. Movie Novels
Writing an educational script for an AI video model is not like writing a Hollywood movie script or a novel. It is like authoring a LEGO assembly instruction manual:
- On each page of a LEGO manual, only one or two specific bricks are highlighted.
- The background is completely clean and white, with zero clutter.
- If page 3 suddenly dumped 50 different bricks onto the diagram, the builder would get overwhelmed and give up.
When authoring children's language videos with Gemini 2.5 Flash, you cannot ask the model for free-form prose. You must treat Gemini as a compiler that outputs a strictly-typed JSON manifest:
- Every shot introduces at most one target vocabulary concept.
- Spoken sentences must never exceed 6 words.
- Every question must be followed by a mandatory 2.5-second silence window so the child can answer.
2.1 Engineering Jargon Demystifier Table
| Industry Term | What It Actually Means | Freshman Student Analogy |
|---|---|---|
| Structured JSON Output | Forcing an LLM to respond strictly using a JSON schema (response_schema), rather than free-form text. |
Filling in a strict online registration form with dedicated boxes instead of typing a random paragraph. |
| Lexical Density | The ratio of new or content-heavy vocabulary words compared to total words in a lesson. | The number of unfamiliar tools a teacher hands you on the very first day of physics class. |
| Syntactic Invariant | A grammatical constraint that must never be broken (e.g. no passive voice, no sentences over 6 words). | A compiler syntax rule: if you break it, the build fails immediately. |
| Call-and-Response Cadence | A 3-step teaching pattern where the character prompts the child, waits quietly for the reply, then celebrates. | A teacher asking "What sound does the cow make?", waiting quietly for the toddler to moo, then clapping. |
| SSML Prosody | XML tags controlling the rate, pitch, and pauses of synthetic speech. | Sheet music for a singer: specifying how fast, high, or quiet each note should be sung. |
2.2 The 5-Minute Micro-Lab: The Storyboard Lexical Linter
Run this zero-dependency Python script to see how automated software enforces educational constraints on an LLM-generated shot manifest:
"""
Micro-Lab: Storyboard Lexical & Pacing Linter
PB-02 Chapter 2 Micro-Lab (Zero External Dependencies)
"""
import json
def lint_shot_manifest(manifest_json: str) -> dict:
data = json.loads(manifest_json)
errors = []
# Check total duration (must equal exactly 60.0s)
total_sec = sum(s["duration_sec"] for s in data["shots"])
if round(total_sec, 1) != 60.0:
errors.append(f"Total duration {total_sec}s != 60.0s")
for s in data["shots"]:
words = s["spoken_text"].split()
if len(words) > 6:
errors.append(f"Shot {s['shot_id']} exceeds 6 words: '{s['spoken_text']}'")
passed = len(errors) == 0
return {"total_duration": total_sec, "shot_count": len(data["shots"]), "errors": errors, "verdict": "[PASS]" if passed else "[FAIL]"}
if __name__ == "__main__":
sample_manifest = '''{
"shots": [
{"shot_id": 1, "duration_sec": 30.0, "spoken_text": "Hello friends look at this!"},
{"shot_id": 2, "duration_sec": 30.0, "spoken_text": "Can you say apple please?"}
]
}'''
res = lint_shot_manifest(sample_manifest)
print(f"Linter Status: {res['verdict']} | Duration: {res['total_duration']}s | Errors: {res['errors']}")
assert res["verdict"] == "[PASS]"
print("[PASS] Micro-lab assertions verified successfully.")
2.3 Freshman Survival Guide: 3 Traps to Avoid
- Trap 1: The Monologue Trap: Letting the character talk for 15 seconds uninterrupted. Toddlers disengage after 4–5 seconds; every 4–5 seconds must feature a visual switch or an interactive sound.
- Trap 2: The Free-Form Prompting Trap: Asking Gemini for "a fun 60-second video script". Free-form prompts return markdown text that downstream video and audio APIs cannot parse. Always enforce structured JSON schemas.
- Trap 3: The Premature Answer Trap: Prompting the mascot to ask "What is this? It's an apple!" all in one breath. The question and the reveal must be separated by a
<break time="2500ms"/>tag in SSML.
2.4 The Pedagogical Foundations of Early Childhood Scriptwriting
Traditional screenwriting principles (three-act structures, subtext, rapid witty dialogue) fail completely when writing for children aged 2 to 6. Early childhood language learners operate under strict cognitive, phonological, and working memory constraints.
[ Traditional Adult Screenplay Format ] ❌
"INT. BARN - DAY. Barnaby inspects the pasture, contemplating
the pastoral serenity before vocalizing his observations."
│ (Completely unexecutable by Video DiTs)
▼
[ Multimodal AI Shot Record Format ] ✅
{
"shot_id": 1,
"duration_sec": 4.5,
"target_word": "COW",
"spoken_tts": "Look! It's a friendly cow!",
"ssml_prosody": "<prosody rate='85%' pitch='+10%'>Look! It's a friendly cow!</prosody>",
"imagen3_prompt": "[3D Pixar Animation]: Barnaby the baby bear standing in a sunny meadow...",
"veo2_action": "Barnaby points to a gentle black-and-white spotted cow, slow camera push-in.",
"on_screen_flashcard": "C - O - W",
"audio_sfx_cue": "cow_moo_gentle.wav"
}
1. CEFR Pre-A1 Lexical Scaffolding
- Lexical Density Limit: Maximum of 3 target vocabulary words per 60-second video lesson. Introducing more than 3 new concepts triggers sensory overload and drops vocabulary retention to near zero.
- Sentence Length Invariant: Spoken sentences must strictly not exceed 5 to 6 words for toddlers (Ages 2–4) or 7 to 8 words for early readers (Ages 5–6).
- Syntactic Simplicity: Zero passive voice, zero subordinate clauses ("although", "whereas"), zero abstract idioms. Language must consist exclusively of concrete nouns (cow, apple, shoe), primary colors (red, blue, yellow), and direct spatial verbs (look, jump, listen).
2. The Triad of Pedagogical Repetition: "Encounter $\rightarrow$ Interact $\rightarrow$ Celebrate"
To anchor a word in a child's long-term memory, every target word must follow a rigid 3-step cadence across 3 distinct shots:
┌────────────────────────────────────────────────────────────────────────┐
│ Phase 1: The Encounter Shot (4.0s) │
│ - Mascot introduces the target object with joyful surprise. │
│ - Spoken Cue: "Look! A shiny red apple!" │
│ - Visual: Mascot presents object center-frame on an open palm. │
├────────────────────────────────────────────────────────────────────────┤
│ Phase 2: Phonics & Syllable Articulation Shot (5.0s) │
│ - Mascot isolates and articulates syllables slowly at 80% speed. │
│ - Spoken Cue: "Ap-ple! Can you say... Apple?" │
│ - Text: Bouncy colored syllables appear: "AP - PLE". │
├────────────────────────────────────────────────────────────────────────┤
│ Phase 3: The Interactive Call-and-Response Shot (5.0s) │
│ - Mascot pauses, cups ear toward the camera with an expectant smile. │
│ - AUDIO INVARIANT: Exactly 2,500ms of complete silence for child! │
│ - Mascot cheers: "Yay! Great job!" with celebratory sparkle sound. │
└────────────────────────────────────────────────────────────────────────┘
2.2 System Prompts & Structured Schema Design for Gemini 2.5 Flash
To use Gemini 2.5 Flash as an autonomous curriculum director, prompts must not request free-form text. We enforce deterministic, machine-readable JSON schemas using Gemini's native structured outputs (response_schema).
1. The Multimodal Shot Schema
Every generated shot must output exact fields consumed downstream by Imagen 3, Google Veo 2, Google Cloud TTS, and FFmpeg:
{
"episode_metadata": {
"theme": "Farm Animals",
"target_age_group": "Ages 2-4 (Pre-A1)",
"target_vocabulary": ["Cow", "Duck", "Sheep"],
"total_target_duration_sec": 60.0
},
"shots": [
{
"shot_id": 1,
"shot_type": "INTRO | WORD_ENCOUNTER | WORD_PHONICS | WORD_CALL_RESPONSE | OUTRO",
"duration_sec": 4.5,
"target_word": "Cow",
"spoken_script": "Look! Here is a friendly cow!",
"ssml_payload": "<speak><prosody rate='85%' pitch='+8%'>Look! Here is a friendly cow!</prosody></speak>",
"child_repetition_pause_ms": 0,
"imagen3_prompt": "[3D Pixar Animation Style, clean sunny pasture]: Barnaby the baby bear standing cheerfully...",
"veo2_action_directive": "Barnaby gestures toward a cute black-and-white cow, gentle slow zoom-in.",
"on_screen_text": "COW",
"audio_foley_trigger": "moo_happy.wav"
}
]
}
2. Production System Prompt for Gemini 2.5 Flash
<system_identity>
You are the Lead Children's Educational Media Director and Pedagogical Curriculum Architect at a global EdTech studio.
Your specialty is designing high-retention early childhood language learning video storyboards for toddlers (Ages 2–5, CEFR Pre-A1) powered by Google Generative AI (Gemini 2.5 Flash, Imagen 3, Google Veo 2, and Cloud TTS).
</system_identity>
<pedagogical_invariants>
1. LEXICAL CEILING: Exactly 3 target vocabulary words per 60-second episode.
2. SENTENCE SIMPLICITY: Maximum 6 words per spoken dialogue turn. Every sentence must use cheerful, simple present tense.
3. THE RULE OF THREE: Every vocabulary word must receive exactly 3 sequential shots:
- Shot A: ENCOUNTER (Mascot discovers/presents object).
- Shot B: PHONICS (Mascot breaks down the word and articulates syllables).
- Shot C: CALL_AND_RESPONSE (Mascot prompts child to speak, followed by mandatory 2,500ms cognitive pause, followed by cheer).
4. MANDATORY SILENCE: In every CALL_AND_RESPONSE shot, the `child_repetition_pause_ms` must be set to exactly `2500`.
5. AESTHETIC INTEGRITY: Every `imagen3_prompt` and `veo2_action_directive` must strictly enforce 3D Pixar/Disney-style cartoon aesthetics, clean pastel backgrounds, and zero uncanny-valley realism.
</pedagogical_invariants>
<output_contract>
Emit raw JSON matching the Multimodal Shot Schema. Do NOT include markdown wrappers (no ```json). Do NOT emit conversational commentary.
</output_contract>
2.3 Naive vs. Production Scriptwriting: Concrete Contrasts
Contrast 1: Scripting the Vocabulary Teaching Moment
❌ Naive Prompt & Output (Amateur Approach)
- Prompt: "Write a cute script for a 30-second kids video teaching the word Cow."
- Emitted Output:
Barnaby: Hello little explorers! Today we are visiting a dairy farm where ruminant animals reside. Over there, chewing its cud, is a magnificent female bovine commonly known as a cow! They provide wholesome dairy milk for our strong bones. Can you say cow? Great, see you next time! - Why it fails:
- Toddlers cannot understand words like "ruminant", "bovine", "reside", "wholesome".
- The mascot asks "Can you say cow?" and immediately says "Great" without leaving a single second for the toddler to speak.
- No visual cues, no shot timing, completely unrenderable by Google Veo 2.
✅ Production Gemini 2.5 Flash Multimodal Storyboard
[
{
"shot_id": 2,
"shot_type": "WORD_ENCOUNTER",
"duration_sec": 4.0,
"target_word": "Cow",
"spoken_script": "Look! It is a cow!",
"ssml_payload": "<speak><prosody rate='85%' pitch='+10%'>Look! It is a cow!</prosody></speak>",
"child_repetition_pause_ms": 0,
"imagen3_prompt": "[3D Pixar Animation Style, warm morning sunlight, vibrant green grass]: Barnaby the baby bear wearing blue overalls, smiling warmly, pointing toward a cute chubby cartoon cow with soft black and white spots.",
"veo2_action_directive": "Barnaby smiles at camera and points to his right. Eye-level camera, smooth static shot.",
"on_screen_text": "COW",
"audio_foley_trigger": "happy_cow_bell.wav"
},
{
"shot_id": 3,
"shot_type": "WORD_PHONICS",
"duration_sec": 5.0,
"target_word": "Cow",
"spoken_script": "Moo! The cow says moo!",
"ssml_payload": "<speak><prosody rate='85%'>Moo! <break time='300ms'/> The cow says moo!</prosody></speak>",
"child_repetition_pause_ms": 0,
"imagen3_prompt": "[3D Pixar Animation Style]: Close-up of the friendly cartoon cow happily chewing a clover blossom, big shiny black cartoon eyes, warm cheerful expression.",
"veo2_action_directive": "The cartoon cow wiggles its ears cheerfully and lets out a gentle friendly moo. Gentle slow push-in.",
"on_screen_text": "C - O - W",
"audio_foley_trigger": "gentle_moo.wav"
},
{
"shot_id": 4,
"shot_type": "WORD_CALL_RESPONSE",
"duration_sec": 5.0,
"target_word": "Cow",
"spoken_script": "Can you say Cow? ... Yay! Good job!",
"ssml_payload": "<speak><prosody rate='85%'>Can you say <emphasis level='strong'>Cow</emphasis>?</prosody> <break time='2500ms'/> <prosody pitch='+15%'>Yay! Good job!</prosody></speak>",
"child_repetition_pause_ms": 2500,
"imagen3_prompt": "[3D Pixar Animation Style]: Barnaby the baby bear cupping his hand to his ear with an excited curious expression, smiling at the camera.",
"veo2_action_directive": "Barnaby leans forward cupping his ear for 2.5 seconds listening, then jumps with joy cheering and clapping hands.",
"on_screen_text": "COW!",
"audio_foley_trigger": "sparkle_cheer.wav"
}
]
- Why it succeeds:
- Every spoken line is under 6 simple words.
- The SSML contains an exact
<break time='2500ms'/>where the audio track stays silent so the child can respond. - Visual descriptions are explicitly framed for Imagen 3 and Google Veo 2.
2.4 Quantitative Engineering Trade-Off Matrix
| Storyboard Architecture | LLM Output Tokens | Downstream Render Stability (Veo 2) | Child Retention Rate | Audio-Visual Sync Precision |
|---|---|---|---|---|
| Monolithic Single-Shot (60s continuous prompt) | ~250 tokens | Extremely Low (Character drifts / turns into blob) | 3.5 / 10 | Unusable (Audio drifts out of sync) |
| 3-Shot Macro Blocks (20s per word) | ~800 tokens | Moderate (Complex action sequences cause morphs) | 6.5 / 10 | Fair |
| 11-Shot Granular Cadence (4–5s per shot) | ~3,500 tokens | Very High (Veo 2 excels at 4-5s stable clips) | 9.5 / 10 | Exact (Frame-perfect cut points) |
2.5 Hands-On Lab: Automated Language Learning Storyboard Generator (Gate 6)
Objective: As an AI Curriculum Engineer, you will build an automated parsing and validation engine that compiles a structured, CEFR Pre-A1 multilingual educational storyboard for kids. The engine must accept any 3-word vocabulary list and generate the exact 11-shot multimodal JSON schedule ready for Google Cloud TTS, Imagen 3, and Google Veo 2.
🧪 The Challenge Scenario
You must build a zero-dependency Python 3.11+ engine that executes the following tasks:
- Input Theme: Theme name (e.g., "Farm Friends"), Target Vocabulary (e.g.,
["Cow", "Duck", "Sheep"]), Mascot Name (e.g., "Barnaby"), Mascot Description (e.g., "friendly chubby baby bear in yellow overalls"). - Storyboard Generation: Programmatically compile the 11 shots:
- Shot 1: Episode Intro & Theme Music (8.0s)
- Shots 2, 3, 4: Word 1 (Encounter, Phonics, Call-and-Response with 2.5s silence) (14.0s)
- Shots 5, 6, 7: Word 2 (Encounter, Phonics, Call-and-Response with 2.5s silence) (14.0s)
- Shots 8, 9, 10: Word 3 (Encounter, Phonics, Call-and-Response with 2.5s silence) (14.0s)
- Shot 11: Recap & Goodbye Dance (10.0s)
- Pedagogical Linter Verification:
- Verify total duration equals exactly 60.0 seconds.
- Verify that all spoken sentences contain <= 6 words.
- Verify that all 3 Call-and-Response shots contain an exact 2,500ms pause.
- Verify that all Imagen 3 and Veo 2 prompts include 3D Pixar/Disney style anchors.
2.6 Recommended Answer & Executable Solution (Gate 7)
Below is the complete, zero-dependency Python 3.11+ reference implementation. It generates, formats, and lints a complete 60-second multimodal educational video storyboard targeting Gemini 2.5 Flash, Imagen 3, and Google Veo 2.
"""
Chapter 2 Hands-On Lab: Automated Language Learning Storyboard Generator
Standard Library Only (Python 3.11+) - Fully Runnable & Verified
Targets: Gemini 2.5 Flash, Imagen 3, Google Veo 2, Google Cloud TTS Journey
"""
import json
from dataclasses import dataclass, asdict
from typing import List, Dict, Any
@dataclass(frozen=True)
class ShotRecord:
shot_id: int
shot_type: str # "INTRO" | "ENCOUNTER" | "PHONICS" | "CALL_RESPONSE" | "OUTRO"
duration_sec: float
target_word: str
spoken_script: str
ssml_payload: str
child_repetition_pause_ms: int
imagen3_prompt: str
veo2_action_directive: str
on_screen_text: str
audio_foley_trigger: str
class EducationalStoryboardCompiler:
"""
Precision curriculum compiler for children's educational video storyboards.
Generates deterministic, CEFR Pre-A1 compliant multi-shot production JSON.
"""
def __init__(self, theme: str, vocabulary: List[str], mascot_name: str, mascot_desc: str):
if len(vocabulary) != 3:
raise ValueError("Pedagogical constraint: Exactly 3 target words required per 60s episode.")
self.theme = theme
self.vocabulary = vocabulary
self.mascot_name = mascot_name
self.mascot_desc = mascot_desc
self.style_anchor = "[3D Pixar Animation Style, warm studio lighting, clean pastel background]"
def generate_storyboard(self) -> List[ShotRecord]:
shots: List[ShotRecord] = []
shot_counter = 1
# 1. Episode Introduction (8.0s)
intro_script = f"Hello friends! Let's learn {self.theme}!"
shots.append(ShotRecord(
shot_id=shot_counter,
shot_type="INTRO",
duration_sec=8.0,
target_word="",
spoken_script=intro_script,
ssml_payload=f"<speak><prosody rate='85%' pitch='+10%'>{intro_script}</prosody></speak>",
child_repetition_pause_ms=0,
imagen3_prompt=f"{self.style_anchor}: {self.mascot_name}, a {self.mascot_desc}, waving happily to camera with both hands.",
veo2_action_directive=f"{self.mascot_name} waves cheerfully, bouncing in place. Eye-level static camera.",
on_screen_text=f"LEARN {self.theme.upper()}!",
audio_foley_trigger="cheerful_intro_jingle.wav"
))
shot_counter += 1
# 2. Vocabulary Cycles (3 words x 3 shots = 9 shots)
for word in self.vocabulary:
# Shot A: Encounter (4.5s)
enc_script = f"Look! Here is a {word.lower()}!"
shots.append(ShotRecord(
shot_id=shot_counter,
shot_type="ENCOUNTER",
duration_sec=4.5,
target_word=word,
spoken_script=enc_script,
ssml_payload=f"<speak><prosody rate='85%' pitch='+8%'>{enc_script}</prosody></speak>",
child_repetition_pause_ms=0,
imagen3_prompt=f"{self.style_anchor}: {self.mascot_name} standing center-frame, happily pointing to a cute 3D cartoon {word.lower()}.",
veo2_action_directive=f"{self.mascot_name} presents the 3D cartoon {word.lower()} with open arms. Gentle camera push-in.",
on_screen_text=word.upper(),
audio_foley_trigger=f"{word.lower()}_pop_sound.wav"
))
shot_counter += 1
# Shot B: Phonics & Articulation (4.5s)
phonics_script = f"Say it with me: {word}!"
shots.append(ShotRecord(
shot_id=shot_counter,
shot_type="PHONICS",
duration_sec=4.5,
target_word=word,
spoken_script=phonics_script,
ssml_payload=f"<speak><prosody rate='80%' pitch='+5%'>Say it with me: <emphasis level='strong'>{word}</emphasis>!</prosody></speak>",
child_repetition_pause_ms=0,
imagen3_prompt=f"{self.style_anchor}: Close-up of the adorable 3D cartoon {word.lower()}, joyful smiling expression.",
veo2_action_directive=f"The 3D cartoon {word.lower()} wiggles and bounces happily. Crisp close-up framing.",
on_screen_text=f"{' - '.join(list(word.upper()))}",
audio_foley_trigger="chime_accent.wav"
))
shot_counter += 1
# Shot C: Call-and-Response Interactive (5.0s, includes 2.5s silence)
cr_script = f"Can you say {word}? ... Hooray!"
shots.append(ShotRecord(
shot_id=shot_counter,
shot_type="CALL_RESPONSE",
duration_sec=5.0,
target_word=word,
spoken_script=cr_script,
ssml_payload=f"<speak><prosody rate='85%'>Can you say <emphasis level='strong'>{word}</emphasis>?</prosody> <break time='2500ms'/> <prosody pitch='+15%'>Hooray! Great job!</prosody></speak>",
child_repetition_pause_ms=2500,
imagen3_prompt=f"{self.style_anchor}: {self.mascot_name} cupping his ear toward camera with an encouraging wide smile.",
veo2_action_directive=f"{self.mascot_name} cups ear for 2.5s listening, then claps hands and jumps with joy.",
on_screen_text=f"{word.upper()}! YAY!",
audio_foley_trigger="sparkle_celebration.wav"
))
shot_counter += 1
# 3. Episode Outro & Celebration (10.0s)
outro_script = "You did great today! Goodbye!"
shots.append(ShotRecord(
shot_id=shot_counter,
shot_type="OUTRO",
duration_sec=10.0,
target_word="",
spoken_script=outro_script,
ssml_payload=f"<speak><prosody rate='85%' pitch='+10%'>{outro_script}</prosody></speak>",
child_repetition_pause_ms=0,
imagen3_prompt=f"{self.style_anchor}: {self.mascot_name} standing together with all three cartoon friends ({', '.join(self.vocabulary)}), waving goodbye.",
veo2_action_directive=f"{self.mascot_name} and characters do a joyful synchronized goodbye wave. Slow pull-back.",
on_screen_text="GREAT JOB! BYE BYE!",
audio_foley_trigger="outro_celebration_song.wav"
))
return shots
def lint_storyboard(self, shots: List[ShotRecord]) -> Dict[str, Any]:
"""
Executes automated pedagogical and engineering assertions:
- Exact duration summation
- Max sentence length (<= 6 words)
- Call-and-response silence windows
- Style anchor integrity
"""
total_duration = sum(s.duration_sec for s in shots)
max_words = 0
pause_count = 0
for s in shots:
# Check sentence word count (excluding ellipses)
words = [w for w in s.spoken_script.replace("...", "").split() if w]
if len(words) > max_words:
max_words = len(words)
if s.child_repetition_pause_ms == 2500:
pause_count += 1
if "[3D Pixar Animation Style" not in s.imagen3_prompt:
raise ValueError(f"Shot {s.shot_id} lacks mandatory 3D cartoon style anchor.")
return {
"total_shots": len(shots),
"total_duration_sec": round(total_duration, 1),
"max_sentence_length_words": max_words,
"call_response_pauses_verified": pause_count,
"passed": (total_duration == 60.0 and max_words <= 6 and pause_count == 3)
}
# ---------------------------------------------------------------------------
# Verification Test Harness
# ---------------------------------------------------------------------------
if __name__ == "__main__":
print("=== Chapter 2 Lab: Automated Storyboard Generator Verification ===")
compiler = EducationalStoryboardCompiler(
theme="Farm Animals",
vocabulary=["Cow", "Duck", "Sheep"],
mascot_name="Barnaby",
mascot_desc="friendly chubby baby bear in yellow overalls"
)
storyboard = compiler.generate_storyboard()
lint_report = compiler.lint_storyboard(storyboard)
print(f"\n[1] Compiled Multimodal Storyboard Summary:")
print(f" Theme: {compiler.theme}")
print(f" Vocabulary: {', '.join(compiler.vocabulary)}")
print(f" Total Shots Generated: {lint_report['total_shots']} shots")
print(f" Total Episode Duration: {lint_report['total_duration_sec']}s")
print(f" Max Spoken Sentence Length: {lint_report['max_sentence_length_words']} words")
print(f" Child Silence Pauses Verified: {lint_report['call_response_pauses_verified']} (2,500ms each)")
print(f"\n[2] Sample Shot Inspection (Shot 4: Call-and-Response):")
sample_shot = storyboard[3] # Shot 4: Cow Call & Response
print(f" Shot ID: {sample_shot.shot_id} ({sample_shot.shot_type})")
print(f" Duration: {sample_shot.duration_sec}s")
print(f" TTS Script: \"{sample_shot.spoken_script}\"")
print(f" SSML Payload: {sample_shot.ssml_payload}")
print(f" Imagen 3 Keyframe Prompt: {sample_shot.imagen3_prompt[:80]}...")
print(f" Google Veo 2 Action Directive: {sample_shot.veo2_action_directive}")
print(f" On-Screen Overlay: \"{sample_shot.on_screen_text}\"")
# Rigorous Pedagogical & Engineering Assertions
assert lint_report["total_shots"] == 11, "Must generate exactly 11 modular shots"
assert lint_report["total_duration_sec"] == 60.0, "Episode runtime must equal exactly 60.0 seconds"
assert lint_report["max_sentence_length_words"] <= 6, "Spoken sentences must not exceed 6 words for toddlers"
assert lint_report["call_response_pauses_verified"] == 3, "Must have exactly 3 call-and-response pauses"
assert sample_shot.child_repetition_pause_ms == 2500, "Cognitive silence must equal exactly 2,500ms"
assert "<break time='2500ms'/>" in sample_shot.ssml_payload, "SSML must contain literal 2500ms break tag"
assert lint_report["passed"] is True, "Storyboard must pass all pedagogical linter checks"
print("\n[PASS] Verification Lab Passed: Multimodal Educational Storyboard Certified.")
Key Architectural Takeaways for Freshman Engineers
- The Programmatic Invariant Guarantee: Notice that by compiling storyboards programmatically through a structured pipeline, we make it impossible for an LLM to violate the 6-word sentence ceiling, forget the 2.5-second silence window, or omit the 3D Pixar animation style anchor.
- Deterministic Output Ready for Downstream APIs: The output of this generator (
storyboard) is an array of typed dictionaries. In subsequent chapters, we will feed this exact array directly into the Google Cloud TTS API (Chapter 5) to synthesize WAV files, Imagen 3 (Chapter 3) to generate keyframes, and Google Veo 2 (Chapter 4) to render video clips!
2.7 The 10 Operational Failure Modes in Kids Educational Scriptwriting
- The Monologue Trap: Writing a continuous lecture where the mascot speaks for 30 seconds uninterrupted. Toddler attention spans drop sharply after 4–5 seconds without visual/verbal interaction.
- Syllable Density Overload: Introducing words with 4+ syllables ("Alligator", "Watermelon") without breaking them down phonetically ("Al - li - ga - tor").
- The Premature Mascot Answer: The mascot asks "What is this?" and immediately says "It's an apple!" without leaving a silence gap in the audio, denying the child viewer their verbal turn.
- Visual-Auditory Semantic Drift: The dialogue says "Look at the red apple!", but the generated image prompt leaves out the color, rendering a green apple and confusing language learners.
- Run-on Sentences with Subordinate Clauses: Using adult conjunctions ("When you see a dog, you should remember that they like to play") instead of simple present tense ("Look! A friendly dog! The dog wags its tail!").
- Flat Emotional Modulation: Speaking with uniform volume and pitch throughout the episode. Educational media requires exaggerated pitch contours (high enthusiasm on introductions, warm gentleness on phonics, explosive cheer on success).
- Homograph Pronunciation Ambiguity: Emitting words with multiple pronunciations without SSML phoneme tags (e.g., "read" past vs present tense, "wind" breeze vs turn).
- Unanchored Camera Directives: Generating action directives like "Camera flies around wildly in all directions", which induces motion sickness and breaks child focus.
- Mascot Personality Drift: The character speaks like an authoritative university professor in Shot 2 and a baby in Shot 5. System instructions must lock character persona parameters.
- Unchecked Trailing Markdown Fences: Emitting markdown blocks (
json ...) that cause downstream automated API parsers to crash with syntax errors.