Overview

Chapter 08: The Anthropic/Claude Ecosystem Perspective — A Critical Companion to Playbook 02

Playbook: PB-02 (AI Video Making for Kids Educational Media) Chapter Role: Independent Review + Vendor-Complementary Companion Chapter (not part of the original 7-chapter Google GenAI pipeline) Core Tooling Stack: Claude Sonnet 5 (claude-sonnet-5), Claude Opus 5 (claude-opus-5), Claude Haiku 4.5 (claude-haiku-4-5-20251001), Claude Code, Claude Agent SDK, Model Context Protocol (MCP), Claude's native vision/multimodal understanding, Extended Thinking, Prompt Caching Target Audience: Year 1 Computer Science & Software Engineering Students Delivery Status: 🔍 Ready for Review (Tier 1 Markdown)


0. The Big Picture: Film Director vs. The Independent Safety Inspector

Why do Hollywood film sets have both an Action Director and an independent On-Set Safety Inspector?

  • If the director who spent millions filming a fiery car crash is the same person who grades whether the stunt was safe, they have a natural conflict of interest. They might overlook safety risks to stay on schedule and budget.
  • An independent safety inspector from outside the production company has no bias: their only job is to stop dangerous stunts before anyone gets hurt.

In Chapters 1 through 7, Google models both generate the videos and grade themselves (Gemini 2.5 Flash auditing Imagen 3 and Veo 2). In this companion chapter, you will learn how to bring in an Independent Auditor (Claude) to review prompts and images before spending expensive compute, protecting toddlers and saving money.


0.1 Engineering Jargon Demystifier Table

Industry Term What It Actually Means Freshman Student Analogy
Same-Vendor Auditor Risk When the same AI model family that created an asset is asked to judge its quality and safety, sharing the same blind spots. A student grading their own homework assignment with an answer key they wrote themselves.
Model Context Protocol (MCP) An open standard for connecting AI models to external tools, APIs, and file systems through a uniform interface. USB ports for AI agents: any tool can plug into any model cleanly.
Prompt Caching Storing frequently repeated prompt prefixes in AI memory so subsequent requests don't pay full token cost or latency. Saving a template file on your desktop instead of re-typing the header every morning.
Extended Thinking An AI reasoning mode where the model plans and verifies steps internally before generating its response. Doing scratch math on scrap paper before writing down the final answer on an exam.
Constitutional AI Training AI models against an explicit set of written rules and ethical principles rather than raw human preference. A school code of conduct that strictly forbids certain actions regardless of peer pressure.

0.2 The 5-Minute Micro-Lab: The Independent Safety Gate Linter

Run this zero-dependency Python script to see how an independent safety gate intercepts flawed prompts before calling expensive rendering APIs:

"""
Micro-Lab: Independent Safety Gate Linter
PB-02 Chapter 8 Micro-Lab (Zero External Dependencies)
"""
from enum import Enum

class Verdict(str, Enum):
    PASS = "[PASS]"
    REJECT = "[REJECT]"

def independent_safety_audit(prompt: str, is_toddler_content: bool) -> dict:
    violations = []
    if "hyperrealistic" in prompt.lower() and is_toddler_content:
        violations.append("Photorealism risk in toddler content (uncanny valley)")
    if "360 spin" in prompt.lower() or "whip pan" in prompt.lower():
        violations.append("Violates 3-second camera cognitive stillness")
        
    verdict = Verdict.PASS if len(violations) == 0 else Verdict.REJECT
    return {"verdict": verdict.value, "violations": violations}

if __name__ == "__main__":
    bad_prompt = "A hyperrealistic toddler bear spinning 360 degrees."
    res = independent_safety_audit(bad_prompt, is_toddler_content=True)
    print("Independent Audit Report:")
    print(f"  Audit Verdict: {res['verdict']}")
    print(f"  Violations:    {res['violations']}")
    assert res["verdict"] == "[REJECT]"
    print("[PASS] Micro-lab assertions verified successfully (Unsafe render averted).")

0.3 Freshman Survival Guide: 3 Traps to Avoid

  1. Trap 1: The Same-Model Self-Grading Trap: Having an LLM grade its own generation. Models suffer from self-preference bias. Always use an external script or independent model to audit outputs.
  2. Trap 2: Paying for Failed Video Renders: Submitting prompts directly to Google Veo 2 ($0.25/clip) without linting them first. If the prompt has an uncanny valley violation, you wasted money. Audit prompts with a fast, cheap model first!
  3. Trap 3: Re-sending Static Style Guides: Sending a 500-word mascot style description on all 11 API calls. Use Prompt Caching so the repeated style guide is cached once at 90% discount.

1. Executive Framing: Why This Chapter Exists

Chapters 1–7 of this playbook are a coherent, technically deep, single-vendor blueprint: every generative stage — curriculum reasoning, character art, video, voice, music, and assembly — runs on the Google Generative AI stack (Gemini 2.5, Imagen 3, Veo 2/3, Cloud TTS Journey, Lyria/MusicFX). That is a legitimate architectural choice: one vendor means one billing relationship, one set of API idioms, and one support channel.

It is also a single point of failure. If Gemini's safety filters, JSON schema adherence, or vision-QA judgment have a systematic blind spot, nothing in Chapters 1–7 catches it, because the same model family that generates the content also grades it (Chapter 3 and Chapter 4 both use Gemini 2.5 Flash as the auditor of Gemini-adjacent Imagen 3/Veo 2 output). This chapter does two things the earlier chapters do not:

  1. Reviews Chapters 1–7 and the two appendices critically — not just praising the architecture, but naming the specific engineering and safety gaps a second-opinion reviewer would flag.
  2. Adds the Anthropic/Claude ecosystem as a complementary layer — not a replacement generator. This is stated plainly up front so it cannot be misread: Claude does not generate video, images, or synthesized speech. Anthropic does not compete with Veo 2, Imagen 3, or Cloud TTS. What Claude brings is a different reasoning and vision-understanding model, from a different lab, with a different safety-training process, sitting outside the generation vendor's own audit loop — directing the pipeline via tool calls and independently grading its output. In a pipeline built for children under 8, that separation of "who generates" from "who grades" is not a nice-to-have; it is a defense-in-depth control most production media-safety programs would require.

2. Chapter-by-Chapter Review

Chapter What It Gets Right What a Second-Opinion Reviewer Would Flag
Ch 01 — Foundations & Google GenAI Stack Dual-coding theory and the anti-uncanny-valley rule are well-grounded pedagogy; the model comparison table is genuinely useful. The claim that Google is "the only fully integrated end-to-end generative media ecosystem" (§1.2) is a marketing framing, not an engineering conclusion — it forecloses the multi-vendor architecture this chapter proposes without evaluating it. The Ch01 lab's cost model (savings_percentage > 99.0) compares AI generation cost to a single traditional studio benchmark with no citation; treat it as illustrative, not a real market figure.
Ch 02 — Scriptwriting & Curriculum with Gemini 2.5 The structured JSON shot schema is production-grade; enforcing the schema in code (EducationalStoryboardCompiler.lint_storyboard) rather than trusting the model's output is the right instinct. The chapter's own "Production Standard" example is entirely generated by the same model family it is meant to critique — there is no adversarial or independently-modeled review step. A curriculum-design LLM grading its own curriculum output is a well-known blind spot (the model tends to rate its own stylistic choices highly).
Ch 03 — Mascot Consistency with Imagen 3 The "Diffusion Consistency Paradox" explanation is technically sound, and the canonical-token + Gemini-QA-gate pattern (Trade-off row E) is the strongest idea in the playbook. The QA gate is Gemini 2.5 Flash auditing Imagen 3 output — both are Google models likely to share training-data biases and blind spots for the same categories of visual defect (e.g., subtle strabismus, near-miss color drift). A same-vendor auditor cannot be assumed independent; it has not been validated against an external grader.
Ch 04 — Directing Veo 2 The physics/attention explanation of I2V conditioning is one of the most rigorous technical passages in the playbook, and the "3-Second Cognitive Rule" is a legitimate, well-sourced constraint. The "Temporal Continuity Auditor" in §3 explicitly pairs FFmpeg + Gemini 2.5 Flash — again, the same-family auditor problem. No chapter proposes what happens when the auditor itself hallucinates a PASS verdict on a genuinely broken clip; there is no fallback grader.
Ch 05 — Voice Acting & Cloud TTS The Infant-Directed Speech research (Fernald & Kuhl) is correctly applied to SSML parameter choices; this is the most academically well-cited chapter in the playbook. No independent listening-comprehension check is proposed — SSML correctness is verified syntactically, never verified against what a model (or human) actually perceives when the audio plays back.
Ch 06 — Music, Foley & Lyria/MusicFX The psychoacoustic masking argument (Cocktail Party Effect degradation in children) is a genuinely strong justification for the -12dB sidechain ducking rule. Ducking thresholds are asserted as universal constants (-12dB, -18dB) without accounting for source loudness variance across different Lyria renders; there's no automated post-hoc loudness verification step (e.g., measuring actual LUFS of the mixed output) in the lab code.
Ch 07 — End-to-End Pipeline KidsVideoOrchestrator.audit_pipeline_integrity() is a real, valuable pattern: deterministic assertions on duration, shot count, and SSML pause presence catch a class of bugs that "eyeballing the video" never would. The pipeline integrity audit checks structure (durations, shot counts, SSML tags) but never checks content safety — nothing in the Chapter 7 lab rejects a shot for an uncanny-valley keyframe or an unsafe lexicon term; Failure Mode #10 ("Spurious Safety Filter Lockout") is discussed narratively but has zero corresponding code-level defense in KidsVideoOrchestrator. This chapter's hands-on lab (§6 below) closes exactly that gap.
App A — Google Media Tooling Catalog Comprehensive, well-organized parameter reference; genuinely useful as a cheat sheet. Entirely Google-scoped by design (reasonably, given its stated purpose) — a reader following only App A has no reference for how to bolt on an independent audit layer.
App B — Open-Source Media Stack Good curation of IP-Adapter, ControlNet, and video DiT alternatives for local benchmarking. No entry addresses agentic orchestration tooling (MCP servers, agent SDKs) — the appendix is generation-model-centric, not pipeline-control-centric, which is precisely the layer this chapter adds.

3. The Anthropic/Claude Viewpoint

Anthropic's product and safety philosophy differs from the "single integrated stack" framing of Chapters 1–7 in three specific, applicable ways:

1. Constitutional AI and a written, public usage policy for content involving minors. Anthropic trains Claude against an explicit constitution and maintains a published Usage Policy with dedicated provisions around content involving children. Practically, this means Claude's refusal/caution behavior around ambiguous child-safety edge cases (Failure Mode #10 in Chapter 7 — "rooster" vs. an ambiguous synonym) is trained by a different process, on different red-team data, than a general-purpose multimodal generation model's safety filter. Using Claude as the second opinion on borderline lexicon or visual content is not redundant — it is testing the same decision against a differently-trained judge, which is exactly what catches correlated blind spots.

2. "Context engineering," not just prompt engineering. Anthropic's own applied-AI guidance frames the discipline as curating the smallest high-signal set of tokens likely to produce the desired behavior — system context, tool definitions, message history, retrieved data — rather than iterating on a single clever prompt string. Chapters 1–7 already do a version of this instinctively (the canonical mascot token strings in Ch03, the structured JSON shot schema in Ch02) but never name it as a discipline, which means there's no explicit budget or reuse strategy for that context across the 11-shot pipeline. Section 4 below shows the concrete mechanism (prompt caching) that makes this free.

3. A harness built for long-running, tool-using orchestration, not single-shot generation. Gemini 2.5 Flash in Chapters 1–7 is invoked as a stateless JSON generator: one call in, one storyboard out. Claude Code and the Claude Agent SDK are designed around a persistent agent loop — plan, call a tool, observe the result, decide whether to continue or halt — which maps far more naturally onto "orchestrate 11 shots across 4 external vendor APIs, stopping to re-prompt on any one shot that fails an independent safety check" than a single structured-output call does. That distinction — stateless generation vs. stateful orchestration — is the architectural gap this chapter fills.


4. Latest Claude Ecosystem Tools & Techniques for This Domain

Tool Real Anthropic Capability Role in a Google-generation / Claude-orchestration Kids Video Pipeline
Claude's native vision understanding Claude Sonnet 5 / Opus 5 accept images directly in the message content and reason over them natively — no separate vision model needed. Replaces the same-vendor Gemini-auditing-Imagen pattern (Ch03/Ch04) with an independent grader: Claude receives the rendered Imagen 3 keyframe / Veo 2 frame and issues a structured verdict (identity consistency, uncanny-valley risk, COPPA lexicon safety) from a model that never touched the generation step.
Model Context Protocol (MCP) An open, Anthropic-authored standard for exposing tools and data sources to any MCP-compatible client through a uniform client/server interface. Each Google service (Imagen 3, Veo 2, Cloud TTS, Lyria) is wrapped as an MCP server exposing generate_keyframe, generate_clip, synthesize_speech, generate_track tools. Claude calls them exactly the way Section 4's lab code models tool_use blocks — the pipeline becomes vendor-swappable at the MCP-server boundary instead of hardcoded API calls.
Claude Agent SDK The same production harness that powers Claude Code, available as an SDK for building custom long-running, tool-calling agents outside the CLI. Implements the ClaudeDirectorAgent orchestration loop in Section 6 as a real, deployable service: plan → call tool → observe → gate → continue, with built-in context management across an 11-shot episode.
Extended thinking A mode where Claude produces an explicit, structured reasoning trace before its final answer, useful for planning and multi-step decisions. Used for the episode-level plan (which shots need which vocabulary reinforcement, what the 3-phase choreography should be) before any tool call is issued — captured as ExtendedThinkingTrace in the lab code, auditable exactly like the JSON schema output in Ch02.
Prompt caching Anthropic's API can cache a prefix of a prompt (e.g., a system block or reference document) for reuse across many calls within a TTL window, cutting cost and latency on the repeated portion. The mascot canonical-identity token string from Ch03 (invariant_physical_tokens + invariant_costume_tokens) is exactly the kind of large, static, reused-every-call context that benefits from caching — instead of re-sending ~150 tokens of style anchor on all 11 shots' worth of vision-critique calls, it is cached once and reused. Modeled in Section 6 as PromptCacheBlock.
Structured outputs / strict tool-use schemas Claude supports constrained, schema-validated tool inputs and outputs, analogous to Gemini's response_schema used throughout Chapters 2–3. Used for the ClaudeCritique verdict object, keeping the audit output machine-parseable and directly pluggable into the same kind of programmatic linter Chapter 7 already builds (audit_pipeline_integrity).
Claude Code Anthropic's agentic CLI coding tool (subagents, hooks, skills, MCP client support, CLAUDE.md steering files). Not part of the runtime pipeline, but the natural authoring environment for building the MCP servers and orchestrator in this chapter — a CLAUDE.md in the pipeline repo encodes the pedagogical invariants from Ch01/Ch02 (60s runtime, 2,500ms pause, 3D-style-only) as standing instructions any future Claude Code session must respect when modifying the pipeline.

5. Updated Trade-Off Table: Google-Only vs. Google-Generation + Claude-Orchestration/Audit

Dimension Google-Only Pipeline (Ch 01–07 as written) Google-Generation + Claude-Orchestration & Audit (This Chapter)
Vendor dependency Single vendor for generation, orchestration logic, and QA grading. Generation stays on Google; orchestration and QA move to an independent vendor. Reduces (does not eliminate) correlated-failure risk.
Safety review independence Gemini 2.5 Flash grades Imagen 3/Veo 2 output (same model family). Claude Sonnet 5 grades Google-generated output (different lab, different safety training).
Marginal cost per additional QA check Full storyboard context resent per Gemini QA call. Style-guide context cached once, reused across all 11 shots' worth of Claude critique calls — illustrative pipeline benchmark: ~70–80% fewer context tokens billed on the repeated style-anchor portion of each critique call.
Wasted render spend on unsafe content A rejected keyframe is only caught after Ch07's structural audit runs (post-hoc); nothing in the Chapter 7 code stops a bad keyframe from proceeding to a Veo 2 render. Section 6's gate blocks the Veo 2 and Cloud TTS calls for any shot that fails Claude's independent audit before those (more expensive) downstream calls fire — illustrative pipeline benchmark: ~$0.25–$0.35 saved per rejected shot (skipped Veo 2 clip cost) versus discovering the defect post-render.
Orchestration statefulness Single stateless JSON-generation calls per stage. Persistent tool-calling agent loop (Claude Agent SDK pattern) that can halt, re-prompt, and resume mid-episode.
Setup complexity Lower — one SDK, one auth flow. Higher — requires MCP servers wrapping each Google service plus the Claude orchestration layer.

(Figures above are illustrative pipeline benchmarks in the style of the rest of this playbook, not published third-party research.)


6. Mandatory Hands-On Lab: Claude Ecosystem Alternate Answer

Restating the Original Challenge (Chapter 7)

Chapter 7's lab asks you to build KidsVideoOrchestrator: compile an 11-shot, 60.0-second manifest across Gemini/Imagen/Veo/TTS, generate the ASS karaoke subtitle file, build the FFmpeg assembly command, and run a structural integrity audit (exact duration, shot count, SSML pause presence). That audit, as reviewed in Section 2, never inspects content safety — it would happily certify a pipeline full of uncanny-valley keyframes as long as the durations add up.

The Claude Ecosystem Extension

Solve an equivalent problem — but insert Claude as an independent director-and-auditor sitting between "keyframe generated" and "video/audio rendered," using Anthropic Messages API tool-use idioms (tool_use / tool_result blocks), a prompt-cache-backed style guide, and a hard gate that skips the expensive downstream Veo 2/TTS calls entirely for any shot that fails the audit.

Recommended Answer #2 (Claude Ecosystem Edition)

#!/usr/bin/env python3
"""
ClaudeOrchestratedKidsPipeline - Claude Ecosystem Companion Suite
Part of Playbook 02, Chapter 08: The Anthropic/Claude Ecosystem Perspective

Zero-dependency Python 3.11+ script (no `anthropic` package required) modeling:
1. Anthropic Messages API-style tool_use / tool_result orchestration blocks
2. Claude as an INDEPENDENT Director + Vision Auditor over mock Google
   generative-media tools, exposed the way an MCP server would expose them
3. A prompt-cache-backed style guide (PromptCacheBlock) reused across shots
4. An extended-thinking-style planning trace captured before tool calls fire
5. A hard safety gate that skips Veo 2 / Cloud TTS calls for any shot that
   fails Claude's independent audit, avoiding wasted render spend
"""

from dataclasses import dataclass, field
from enum import Enum
from typing import Any, Dict, List, Optional, Tuple


# ---------------------------------------------------------------------------
# 1. Anthropic Messages API-style primitives (modeled, not a live API client)
# ---------------------------------------------------------------------------

@dataclass
class ToolUseBlock:
    """Mirrors an Anthropic Messages API `tool_use` content block."""
    id: str
    name: str
    input: Dict[str, Any]


@dataclass
class ToolResultBlock:
    """Mirrors an Anthropic Messages API `tool_result` content block."""
    tool_use_id: str
    content: Dict[str, Any]
    is_error: bool = False


@dataclass
class ExtendedThinkingTrace:
    """Simulated extended-thinking block: Claude's structured planning reasoning
    captured before any tool call is issued."""
    thinking_text: str
    token_estimate: int


# ---------------------------------------------------------------------------
# 2. Mock Google generative-media tools, exposed to Claude MCP-server style.
#    Claude does NOT generate video/image/audio itself -- it calls these tools.
# ---------------------------------------------------------------------------

MCP_TOOL_SCHEMAS = [
    {
        "name": "google_imagen3_generate_keyframe",
        "description": "Calls Imagen 3 to render a single keyframe PNG for a shot.",
        "input_schema": {
            "type": "object",
            "properties": {
                "prompt": {"type": "string"},
                "negative_prompt": {"type": "string"},
            },
            "required": ["prompt", "negative_prompt"],
        },
    },
    {
        "name": "google_veo2_generate_clip",
        "description": "Calls Google Veo 2 to render a 5s Image-to-Video clip from a keyframe.",
        "input_schema": {
            "type": "object",
            "properties": {
                "keyframe_id": {"type": "string"},
                "action_prompt": {"type": "string"},
            },
            "required": ["keyframe_id", "action_prompt"],
        },
    },
    {
        "name": "google_cloud_tts_synthesize",
        "description": "Calls Google Cloud TTS Journey voice to synthesize SSML dialogue.",
        "input_schema": {
            "type": "object",
            "properties": {"ssml": {"type": "string"}},
            "required": ["ssml"],
        },
    },
]


def mock_google_imagen3_generate_keyframe(prompt: str, negative_prompt: str) -> Dict[str, Any]:
    """Deterministic mock standing in for a live Imagen 3 API call."""
    flagged_terms = ["realistic", "photoreal", "human teeth", "hyperrealistic"]
    risk = "CRITICAL" if any(t in prompt.lower() for t in flagged_terms) else "NONE"
    style_ok = "3d" in prompt.lower() or "pixar" in prompt.lower() or "cartoon" in prompt.lower()
    return {
        "keyframe_id": f"kf_{abs(hash(prompt)) % 100000}",
        "uncanny_valley_risk": risk,
        "style_anchor_present": style_ok,
    }


def mock_google_veo2_generate_clip(keyframe_id: str, action_prompt: str) -> Dict[str, Any]:
    unsafe_motion = any(t in action_prompt.lower() for t in ["spin 360", "whip pan", "camera shake"])
    return {"clip_id": f"clip_{keyframe_id}", "motion_safety_flag": "UNSAFE" if unsafe_motion else "SAFE"}


def mock_google_cloud_tts_synthesize(ssml: str) -> Dict[str, Any]:
    return {"audio_id": f"audio_{abs(hash(ssml)) % 100000}", "contains_response_window": "2500ms" in ssml}


TOOL_IMPLEMENTATIONS = {
    "google_imagen3_generate_keyframe": mock_google_imagen3_generate_keyframe,
    "google_veo2_generate_clip": mock_google_veo2_generate_clip,
    "google_cloud_tts_synthesize": mock_google_cloud_tts_synthesize,
}


# ---------------------------------------------------------------------------
# 3. Claude's independent vision-critique QA gate (NOT the generation vendor)
# ---------------------------------------------------------------------------

class SafetyVerdict(Enum):
    PASS = "PASS"
    REJECT = "REJECT"


@dataclass
class ClaudeCritique:
    shot_id: int
    identity_consistency_score: int
    uncanny_valley_risk: str
    coppa_lexicon_safe: bool
    verdict: SafetyVerdict
    rationale: str


class ClaudeVisionAuditGate:
    """
    Models Claude Sonnet 5 as an INDEPENDENT auditor -- a different model
    family than the Gemini/Imagen/Veo generation stack -- reducing the
    correlated-failure risk of a same-vendor generator-grades-itself loop
    (see Section 2's review of Chapters 3 and 4).
    """

    UNSAFE_LEXICON = {"cock", "ass", "bloody", "shoot"}  # naive synonyms flagged in Ch7 Failure Mode #10

    def critique_keyframe(self, shot_id: int, keyframe_result: Dict[str, Any], spoken_script: str) -> ClaudeCritique:
        risk = keyframe_result["uncanny_valley_risk"]
        style_ok = keyframe_result["style_anchor_present"]
        lexicon_hit = any(w in spoken_script.lower().split() for w in self.UNSAFE_LEXICON)

        identity_score = 96 if (style_ok and risk == "NONE") else 40
        verdict = SafetyVerdict.PASS if (risk == "NONE" and style_ok and not lexicon_hit) else SafetyVerdict.REJECT
        rationale = (
            "Style anchor present, zero uncanny-valley risk, lexicon clean."
            if verdict == SafetyVerdict.PASS
            else "Rejected: uncanny-valley risk, missing style anchor, or unsafe lexicon term."
        )
        return ClaudeCritique(
            shot_id=shot_id,
            identity_consistency_score=identity_score,
            uncanny_valley_risk=risk,
            coppa_lexicon_safe=not lexicon_hit,
            verdict=verdict,
            rationale=rationale,
        )


# ---------------------------------------------------------------------------
# 4. Prompt-cache-backed style guide + the orchestration agent
# ---------------------------------------------------------------------------

@dataclass
class PromptCacheBlock:
    """Simulated Anthropic prompt-cache block: the mascot style guide reused
    across every keyframe-critique call at near-zero marginal token cost."""
    cache_key: str
    content: str
    hit_count: int = 0

    def use(self) -> str:
        self.hit_count += 1
        return self.content


@dataclass
class ClaudeOrchestratedShot:
    shot_id: int
    shot_type: str
    duration_sec: float
    tool_calls: List[ToolUseBlock] = field(default_factory=list)
    tool_results: List[ToolResultBlock] = field(default_factory=list)
    critique: Optional[ClaudeCritique] = None


class ClaudeDirectorAgent:
    """
    Models Claude Sonnet 5 acting as Director + Auditor over a Claude Agent
    SDK-style tool loop, orchestrating the Google generative-media stack via
    MCP-exposed tool calls, gated by an independent safety audit.
    """

    def __init__(self):
        self.audit_gate = ClaudeVisionAuditGate()
        self.style_cache = PromptCacheBlock(
            cache_key="pippa_style_guide_v1",
            content="[3D Stylized Pixar Animation Style]: Pippa the Penguin, navy plumage, "
                    "tangerine beak, mustard-yellow scarf.",
        )
        self.transcript: List[ClaudeOrchestratedShot] = []

    def plan_episode(self, theme: str, words: List[str]) -> ExtendedThinkingTrace:
        thinking = (
            f"Planning '{theme}' episode for target words {words}. Will call "
            "google_imagen3_generate_keyframe per shot using the cached style guide, "
            "then run an independent vision critique before allowing "
            "google_veo2_generate_clip to proceed. Reject on any uncanny-valley or "
            "lexicon violation BEFORE spending Veo 2 / Cloud TTS render cost."
        )
        return ExtendedThinkingTrace(thinking_text=thinking, token_estimate=len(thinking.split()) * 2)

    def orchestrate_shot(
        self, shot_id: int, shot_type: str, duration_sec: float,
        keyframe_prompt: str, negative_prompt: str,
        action_prompt: str, spoken_script: str, ssml: str,
    ) -> ClaudeOrchestratedShot:
        shot = ClaudeOrchestratedShot(shot_id=shot_id, shot_type=shot_type, duration_sec=duration_sec)

        style_context = self.style_cache.use()
        full_prompt = f"{style_context} {keyframe_prompt}"

        call_id = f"toolu_{shot_id}_imagen"
        shot.tool_calls.append(ToolUseBlock(
            id=call_id, name="google_imagen3_generate_keyframe",
            input={"prompt": full_prompt, "negative_prompt": negative_prompt},
        ))
        keyframe_result = TOOL_IMPLEMENTATIONS["google_imagen3_generate_keyframe"](full_prompt, negative_prompt)
        shot.tool_results.append(ToolResultBlock(tool_use_id=call_id, content=keyframe_result))

        critique = self.audit_gate.critique_keyframe(shot_id, keyframe_result, spoken_script)
        shot.critique = critique

        if critique.verdict == SafetyVerdict.PASS:
            call_id_veo = f"toolu_{shot_id}_veo"
            shot.tool_calls.append(ToolUseBlock(
                id=call_id_veo, name="google_veo2_generate_clip",
                input={"keyframe_id": keyframe_result["keyframe_id"], "action_prompt": action_prompt},
            ))
            veo_result = TOOL_IMPLEMENTATIONS["google_veo2_generate_clip"](
                keyframe_result["keyframe_id"], action_prompt
            )
            shot.tool_results.append(ToolResultBlock(tool_use_id=call_id_veo, content=veo_result))

            call_id_tts = f"toolu_{shot_id}_tts"
            shot.tool_calls.append(ToolUseBlock(
                id=call_id_tts, name="google_cloud_tts_synthesize", input={"ssml": ssml},
            ))
            tts_result = TOOL_IMPLEMENTATIONS["google_cloud_tts_synthesize"](ssml)
            shot.tool_results.append(ToolResultBlock(tool_use_id=call_id_tts, content=tts_result))

        self.transcript.append(shot)
        return shot

    def audit_accepted_pipeline(
        self, accepted_shots: List[ClaudeOrchestratedShot], target_duration_sec: float
    ) -> Tuple[bool, List[str]]:
        """Structural + safety audit over the FINAL accepted shot list (rejected
        shots already excluded upstream by the gate in orchestrate_shot)."""
        violations: List[str] = []
        total_duration = sum(s.duration_sec for s in accepted_shots)
        if round(total_duration, 2) != round(target_duration_sec, 2):
            violations.append(f"Accepted duration {total_duration}s != target {target_duration_sec}s")
        rejected_in_final = [s for s in accepted_shots if s.critique.verdict == SafetyVerdict.REJECT]
        if rejected_in_final:
            violations.append(f"Rejected shots leaked into final pipeline: {[s.shot_id for s in rejected_in_final]}")
        if self.style_cache.hit_count < len(self.transcript):
            violations.append("Prompt cache under-utilized: expected >= 1 hit per orchestrated shot")
        return len(violations) == 0, violations


# =====================================================================
# VERIFICATION SUITE & DEMO EXECUTION
# =====================================================================

if __name__ == "__main__":
    print("=== Chapter 08 Lab: Claude-Orchestrated Kids Video Pipeline (Ecosystem Edition) ===\n")

    agent = ClaudeDirectorAgent()

    plan = agent.plan_episode(theme="Farm Animals", words=["Cow"])
    print("[1] Extended-Thinking Planning Trace:")
    print(f"  {plan.thinking_text}\n")

    # Shot 1: INTRO -- safe prompt, should PASS and proceed to Veo 2 + TTS
    shot1 = agent.orchestrate_shot(
        shot_id=1, shot_type="INTRO", duration_sec=8.0,
        keyframe_prompt="Pippa waves happily to camera, sunny pastel farm backdrop.",
        negative_prompt="photorealistic, uncanny valley, sharp teeth",
        action_prompt="Static camera, gentle wave.",
        spoken_script="Hello little friends! Let's learn farm animals!",
        ssml="<speak><prosody rate='85%'>Hello little friends!</prosody></speak>",
    )

    # Shot 2: WORD_ENCOUNTER -- intentionally naive/unsafe prompt, should REJECT
    # and must NOT waste a Veo 2 / TTS call.
    shot2_bad = agent.orchestrate_shot(
        shot_id=2, shot_type="WORD_ENCOUNTER", duration_sec=4.5,
        keyframe_prompt="hyperrealistic cow standing in a field",  # naive prompt, no style anchor
        negative_prompt="",
        action_prompt="Static tripod, point at cow.",
        spoken_script="Look! It is a cow!",
        ssml="<speak>Look! It is a cow!</speak>",
    )

    # Shot 2 retry: corrected production-grade prompt, should PASS
    shot2_fixed = agent.orchestrate_shot(
        shot_id=2, shot_type="WORD_ENCOUNTER", duration_sec=4.5,
        keyframe_prompt="[3D Pixar Style]: Pippa points to a cute cartoon cow.",
        negative_prompt="photorealistic, uncanny valley, sharp teeth",
        action_prompt="Static tripod, point at cow.",
        spoken_script="Look! It is a cow!",
        ssml="<speak>Look! It is a cow!</speak>",
    )

    print("[2] Per-Shot Orchestration & Independent Audit Results:")
    for s in (shot1, shot2_bad, shot2_fixed):
        print(f"  - Shot {s.shot_id} ({s.shot_type}): verdict={s.critique.verdict.value}, "
              f"tool_calls={len(s.tool_calls)}, rationale=\"{s.critique.rationale}\"")
    print()

    # Cost-avoidance assertion: the rejected shot must NOT have triggered Veo 2 / TTS calls
    assert len(shot2_bad.tool_calls) == 1, "Rejected shot must only spend the Imagen 3 call, not Veo 2 / TTS"
    assert shot2_bad.critique.verdict == SafetyVerdict.REJECT, "Naive hyperrealistic prompt must be rejected"
    assert len(shot2_fixed.tool_calls) == 3, "Passed shot must spend Imagen 3 + Veo 2 + TTS calls"
    assert shot2_fixed.critique.verdict == SafetyVerdict.PASS, "Corrected prompt must pass the audit"

    # Final accepted pipeline excludes the rejected attempt
    accepted_pipeline = [shot1, shot2_fixed]
    is_valid, violations = agent.audit_accepted_pipeline(accepted_pipeline, target_duration_sec=12.5)

    print("[3] Final Accepted-Pipeline Structural + Safety Audit:")
    if is_valid:
        print("  >>> PASSED: Accepted-shot duration matches target exactly.")
        print("  >>> PASSED: Zero rejected shots leaked into the final pipeline.")
        print(f"  >>> PASSED: Prompt cache reused {agent.style_cache.hit_count} times across "
              f"{len(agent.transcript)} orchestrated shot attempts (style guide sent once, reused after).")
    else:
        print(f"  >>> FAILED: {len(violations)} violations:")
        for v in violations:
            print(f"      - {v}")
    assert is_valid, f"Final pipeline audit failed: {violations}"

    print("\n[PASS] Verification Lab Passed: Claude-Orchestrated Kids Video Pipeline Certified (Ecosystem Edition).")

What This Answer Demonstrates That Chapter 7's Did Not

  1. A rejected shot never reaches the expensive stages. shot2_bad stops after one Imagen 3-equivalent call; KidsVideoOrchestrator in Chapter 7 has no such gate, so a structurally valid but visually unsafe pipeline would sail through its audit unmodified.
  2. The grader is architecturally independent from the generator, directly answering the Section 2 critique of Chapters 3 and 4 (same-vendor auditor risk).
  3. Prompt caching is modeled as a first-class cost mechanism, not just described in prose — style_cache.hit_count is asserted, not just claimed.

7. Summary

This companion chapter does not replace anything in Chapters 1–7 — Google's Gemini/Imagen/Veo/TTS/Lyria stack remains the generation engine, and nothing here suggests otherwise. What it adds:

  • A genuinely critical review of all seven chapters and both appendices, naming the recurring same-vendor-audits-itself pattern as the playbook's single biggest unaddressed architectural risk.
  • The Anthropic/Claude viewpoint — Constitutional AI, a published child-safety usage policy, and a stateful tool-orchestration harness — as the concrete mechanism for closing that gap.
  • A mapped inventory of real, current Claude ecosystem tools (native vision understanding, MCP, the Claude Agent SDK, extended thinking, prompt caching, structured tool outputs, Claude Code) against this playbook's specific pipeline stages.
  • A second, independently runnable hands-on-lab answer — ClaudeOrchestratedKidsPipeline — that solves Chapter 7's orchestration problem with an added independent safety gate and cache-aware cost control, verified to reject unsafe content before it reaches the costly rendering stages.

For a production kids' media studio, the recommended posture is not "Google or Claude" — it is Google for generation, Claude for the director/auditor role sitting outside that generation loop.