Overview

Chapter 9: The Anthropic/Claude Ecosystem Perspective — A Critical Companion to Playbook 01

Playbook Track: 01 – Prompt Engineering Playbook (Companion / Review Chapter) Target Audience: Year 1 Computer Science & Software Engineering Students Core Tooling Stack: Claude Sonnet 5 / Opus 5 / Haiku 4.5 / Fable 5.1, Prompt Caching, Extended & Interleaved Thinking, Claude Code (CLAUDE.md, Subagents, Hooks, Skills, Plugins), Claude Agent SDK, Model Context Protocol (MCP), Claude Developer Platform (Messages API, Files API, Batch API, Citations, Code Execution Tool), Claude Cookbooks Delivery Status: 🔍 Ready for Review (Tier 1 Markdown) Prerequisites: All of Chapters 1–8 of this playbook, plus `appendix-github-repos.md`


9.1 Executive Framing: Why This Chapter Exists

Chapters 1–8 built a genuinely rigorous, mostly vendor-neutral prompt engineering curriculum — attention mechanics, tokenizer anomalies, CoT theory, tool-calling schemas, eval harnesses, cost optimization, and a master template library. That is real engineering value, and this chapter does not re-derive it.

What it does do is two things the earlier chapters structurally could not do, because they were scoped as a broad multi-vendor survey rather than a deep single-vendor deployment guide:

  1. Audit the existing eight chapters with a critical eye — not a rubber-stamp "looks good," but concrete gaps: places where illustrative numbers are presented with more precision than they can be verified to, where a technique is described in the abstract without naming the concrete production surface that implements it, and where the playbook's own instruction.md mandates ("Frontier-Model Accurate," "Trade-Off Quantification") are only partially met for the Anthropic side of the matrix.
  2. Go deep on one vendor's production surface — the Claude Developer Platform and Claude Code — because "prompt engineering" in 2026 is no longer just crafting a string; it is deciding what goes in the context window at all (context engineering), what gets cached, what runs as a tool, and what runs as an agent loop with its own file system and shell. Anthropic's own public guidance has explicitly reframed the discipline this way, and the earlier chapters, written before that reframing was incorporated, still treat "prompt" and "context" as near-synonyms.

This chapter follows the same authoring standard as Chapters 1–8 (Naive vs. Production contrasts, trade-off matrices, a hands-on lab, an executable zero-dependency solution) but is explicitly a review + complement, not a replacement.


9.2 Chapter-by-Chapter Review

Chapter 1 — LLM Foundations & Token Mechanics

  • Strength: The Base → Instruction-Tuned → Reasoning taxonomy table is accurate and durable; it will not go stale quickly.
  • Gap 1 — Tokenizer claim needs a caveat: The chapter states Anthropic Claude uses BPE alongside OpenAI and Llama 3. That is directionally correct (Claude's tokenizer is BPE-derived), but Anthropic does not publish the vocabulary or merge table the way tiktoken does for OpenAI, so any code in this playbook that calls tiktoken.encoding_for_model() to estimate Claude token counts (as Chapter 7 implicitly encourages via generic "count tokens before sending" advice) will silently be wrong by 5–15%. The chapter should have flagged that Claude token counts must be obtained from the Messages API's count_tokens endpoint, not approximated with a foreign tokenizer.
  • Gap 2 — Reasoning-model framing is now incomplete: The "Reasoning / Extended Thinking Models" row lumps Claude 3.7 Thinking in as one example among several, described only as "native chain-of-thought search prior to emission." It doesn't mention that Claude's extended thinking is a budgeted, streamed, and separately-billed content block (thinking block type) that can be interleaved with tool calls — a materially different engineering surface than a hidden reasoning trace. See §9.4.2.

Chapter 2 — Core Prompting Architectures

  • Strength: The "semantic priming paradox" explanation of why negative constraints backfire is one of the strongest sections in the playbook and holds true across every vendor, Claude included.
  • Gap 1: No mention of prefilling the assistant turn, which is a first-class, load-bearing technique on Anthropic's API (seeding assistant content with { or a partial XML tag to hard-anchor output format) and is more aggressively supported there than on most peer APIs. This is a direct extension of the chapter's own "response trigger" pattern (<response_trigger>\n{) but the playbook never names it as a distinct, first-class API parameter.
  • Gap 2: Zero-shot failure envelope discussion is solid but doesn't distinguish "the model doesn't know the format" from "the model knows the format but the system prompt structure itself is fighting the user turn" — the latter is specifically what Anthropic's role-prompting guidance (§9.3) addresses.

Chapter 3 — System Prompts & Safety Guardrails

  • Strength: The KV-cache / attention-sink argument for why system prompts get privileged treatment is mechanically correct and is exactly the argument Anthropic uses to justify prompt caching pricing (see §9.4.1) — the playbook built the theory but never connected it to the product.
  • Gap: The chapter correctly notes Claude and Gemini "disallow raw special-token concatenation" and require a dedicated system parameter, but stops there. It never covers Constitutional AI or Anthropic's helpful/harmless/honest (HHH) training objective, which is the actual mechanism why Claude's refusal behavior differs from RLHF-only models under adversarial system prompts — relevant to every guardrail pattern in this chapter. Covered in §9.3.

Chapter 4 — Reasoning Paradigms & CoT Evolution

  • Strength: Self-Consistency and Tree-of-Thoughts diagrams are correct and implementation-agnostic.
  • Gap: The chapter treats "manual scratchpad CoT" and "native reasoning models" as a binary choice, but doesn't cover interleaved thinking — Claude's ability to alternate thinking blocks between tool calls in a single agentic turn (think → call tool → think about the result → call another tool), which changes the ReAct loop described in this playbook's own instruction.md §4.2. A manually-scripted Thought:/Action:/Observation: loop (as instruction.md prescribes) is now a legacy pattern for models that expose native interleaved thinking.

Chapter 5 — Agentic Prompts & Tool Orchestration

  • Strength: The bounded self-correction loop (max_correction_attempts=2), the "READ-ONLY vs DESTRUCTIVE" idempotency tagging convention, and the 10 failure modes are all production-accurate and vendor-agnostic.
  • Gap 1 — Schema field name is wrong for Claude: The example tool schema in §5.5 uses "parameters" as the top-level schema key. That is the OpenAI/Gemini function-calling convention. Anthropic's Messages API uses "input_schema", and a schema copy-pasted verbatim from this chapter into a Claude integration will be silently ignored (the tool will not be registered correctly). This is a concrete, fixable defect, not a stylistic quibble.
  • Gap 2: "Grammar-Constrained Decoding... masks tokens to $-\infty$" is described as a universal provider mechanism. Anthropic does not publicly document logit-masking for tool-use the way this implies; Claude's high tool-schema conformance is closer to strong instruction-tuning plus (for tool_choice: {"type": "tool"}) forced tool selection, not necessarily a CFG mask at the token level. The playbook should not present unverified inference-engine internals as fact for every vendor.
  • Gap 3: No discussion of client-side tool execution via MCP as a way to standardize the tool registry across agents instead of hand-writing bespoke JSON schemas per project — directly relevant to §5.2's "High-Recall Tool Descriptions" problem. See §9.4.3.

Chapter 6 — Prompt Evals & Empirical Benchmarking

  • Strength: Position-swapped pairwise comparison and the G-Eval rubric protocol are both correct, citable, and the included Python lab actually implements bias mitigation rather than just describing it — rare and valuable.
  • Gap: Self-Enhancement Bias mitigation says "use Claude to judge GPT outputs, or GPT to judge Gemini outputs," which is right, but the chapter never mentions that the Claude Developer Console ships a built-in Evaluate tool (prompt versioning + test-case grading in the Console UI, no custom harness required to get started) — worth knowing before investing engineering time building the Chapter 6 harness from scratch for a Claude-only project. See §9.4.6.

Chapter 7 — Production Cost, Latency & Context Caching

  • Strength: The Cost Asymmetry Principle (output tokens 3–5x input cost) is accurate and durable.
  • Gap 1 — This is the single biggest missed opportunity in the whole playbook. Chapter 7's own architecture diagram (KV-CACHE HIT: 90% Discount, 10x Faster TTFT) is describing, almost verbatim, Anthropic's prompt caching product — yet the chapter never names it, never gives the cache_control: {"type": "ephemeral"} block syntax, never mentions the 5-minute default vs. extendable long-TTL cache, and never shows how cache breakpoints compose with the Tier-1/Tier-2 system-prompt structure Chapter 8 defines. The theory is present; the concrete API surface is absent. Fixed in §9.4.1 and demonstrated in the lab (§9.6).
  • Gap 2: LLMLingua-style perplexity pruning is presented as the primary compression strategy, but for a Claude-specific deployment, the Files API + Batch API are frequently cheaper first moves (avoid re-uploading large static documents every call; batch non-interactive workloads at ~50% cost). Neither is mentioned.

Chapter 8 — Production Templates & Master Cheat Sheet

  • Strength: The 5-Tier System Prompt Blueprint is a genuinely reusable architecture, and its ordering (Identity → Mandate → Reasoning Protocol → Output Contract → Adversarial Invariants) maps cleanly onto cache-friendly prompt design (static tiers first, volatile tiers last) — the playbook built this instinctively correct but, again per the Chapter 7 gap, never says so explicitly.
  • Gap: The template's <reasoning_protocol> tier says "for reasoning models... omit manual scratchpad formatting and rely directly on native thinking," which is correct, but gives no concrete guidance on how to request and read back Claude's thinking blocks (the thinking parameter's budget_tokens, and that thinking content must be preserved verbatim in multi-turn tool-use conversations or the API will reject the request). This is a sharp edge that silently breaks agentic multi-turn code if unhandled.

Appendix B — Curated GitHub Repositories

  • Strength: Outlines, Instructor, Promptfoo, and DeepEval are all genuinely still-maintained, high-value repos.
  • Gap: No entry for Anthropic's own open-source Cookbooks repository (anthropics/claude-cookbooks on GitHub), which is the single highest-signal source of runnable, Anthropic-maintained reference patterns (tool use, PDF/vision, RAG, extended thinking, evaluation) — arguably more authoritative for Claude-specific work than any third-party wrapper library listed.

9.3 The Anthropic/Claude Viewpoint: Constitutional AI and Context Engineering

Two framing shifts are load-bearing for everything in §9.4 and the lab:

1. Constitutional AI (CAI) and HHH training

Rather than relying solely on RLHF against human preference labels, Claude models are additionally trained against a written constitution — a set of principles the model uses to critique and revise its own draft responses (RLAIF: Reinforcement Learning from AI Feedback, self-critique against the constitution, then distillation). Practically, this changes two things this playbook's guardrail patterns (Chapter 3) should account for:

  • Negative constraints are less adversarially brittle, not immune. Claude's training explicitly optimizes for honesty and refusal-with-explanation over blind compliance, so the "semantic priming paradox" from Chapter 2 (banning a topic ironically primes it) is measurably less severe than on pure-RLHF models — but Chapter 3's "Standard Refusal Protocol" pattern (an exact literal refusal string) is still the right production pattern, because it gives a deterministic, testable output for Chapter 6's eval harness. CAI reduces reliance on negative constraints; it doesn't replace the engineering discipline of specifying one.
  • Claude will push back on premise-flawed instructions. The Chapter 8 template's "Anti-Sycophancy" invariant ("reject the assumption if it contradicts context data") is precisely the behavior CAI training optimizes for by default — so on Claude, that invariant can often be a lighter-weight reminder rather than a hard behavioral override you must fight the model to get.

2. "Context engineering" supersedes "prompt engineering" as the frame

Anthropic's own public engineering guidance has explicitly repositioned the discipline: a single well-worded prompt string is necessary but not sufficient once a system involves tools, retrieved documents, multi-turn history, and sub-agent outputs all competing for the same finite context window. The unit of engineering shifts from "what do I say" to "what is the smallest, highest-signal set of tokens (instructions + tools + examples + retrieved data + memory) that gets the model to the right next action." This is a direct generalization of this playbook's own Chapter 5 (Progressive Tool Exposure) and Chapter 7 (Prompt Compression) sections — the playbook already practices context engineering in places, it just never named the discipline or gave it Chapter-8-blueprint-level treatment as its own first-class concern.

3. Anthropic's canonical prompting techniques (delta vs. this playbook)

Anthropic's own documentation converges on largely the same techniques this playbook independently derived from first principles (XML-tag delimiting, role/persona system prompts, few-shot exemplars, explicit CoT scratchpads) — which is a good cross-validation signal that Chapters 1–8 are not idiosyncratic. The deltas worth adding to this playbook's toolkit:

  • XML tags as the preferred, not merely acceptable, delimiter for Claude specifically — Claude was extensively trained on XML-structured data, so <tag> boundaries measurably outperform Markdown headers or triple-backticks for the same task on Claude, more so than on some peer models.
  • The "successive refinement" prompt-chaining pattern: decompose a single complex prompt into a pipeline of smaller prompts where each stage's output becomes the next stage's <input> block — an explicit production pattern that generalizes Chapter 5's sequential tool chaining to non-tool text-generation pipelines.
  • Long-context placement rule with one added nuance: Chapter 1's "put retrieved documents in the middle, instructions at start/end" advice holds, but Claude's guidance additionally recommends placing the query after long documents, restated once more at the very end, even when it also appears at the top — a belt-and-suspenders anti-"lost-in-the-middle" tactic not covered in Chapter 1.

9.4 Latest Claude Ecosystem Tools & Techniques for Prompt Engineering

9.4.1 Prompt Caching — the concrete implementation of Chapter 7's KV-cache diagram

Mark a cache_control: {"type": "ephemeral"} breakpoint on any prefix of the request (typically at the end of the system block, after static tool definitions, or after a large static document). On a cache write, Anthropic charges a modest premium over base input price; on a cache hit (subsequent calls within the TTL, default 5 minutes, extendable to longer TTLs), the cached prefix is billed at a steep discount and TTFT drops sharply — this is exactly Chapter 7's "90% Discount, 10x Faster TTFT" box turned into a real API call. Engineering implication for this playbook's Chapter 8 blueprint: Tiers 1–2 (Identity, Operational Mandate) should sit before the cache breakpoint; Tiers 3–5 (Reasoning Protocol, Output Contract, Adversarial Invariants) — which are more likely to be tuned per-request — sit after it, so cache invalidation from prompt iteration doesn't nuke the whole system prompt's cache value.

9.4.2 Extended Thinking & Interleaved Thinking

Requesting thinking with a budget_tokens allocation gives Claude a dedicated, separately-visible reasoning channel before its final answer — directly replacing Chapter 4's manual <scratchpad> pattern for models that support it, at higher reliability because the reasoning tokens are not competing with output-formatting tokens in the same channel. Interleaved thinking additionally lets the model think between tool calls in one agentic turn, not just once before the first action — the direct evolution of Chapter 4's linear ReAct loop into a branching one, without the client needing to manually re-invoke the model between every tool call.

9.4.3 Model Context Protocol (MCP)

MCP is an open protocol (not Anthropic-proprietary, though Anthropic authored and open-sourced it) for exposing tools, resources, and prompts to a model client through a standard server interface. For this playbook's Chapter 5 concerns specifically: instead of every project hand-rolling tool JSON schemas and a dispatcher (as Chapter 5's lab does), an MCP server declares its tools once and any MCP-compatible client (Claude Code, Claude Desktop, custom Agent SDK clients) can discover and call them identically — directly addressing Chapter 5's "Tool Namespace Congestion" and "Schema-Docstring Divergence" failure modes by centralizing the schema in one server rather than duplicating it per client.

9.4.4 Claude Code: CLAUDE.md, Subagents, Hooks, Skills, Plugins

Claude Code (and this repository's own AGENT_TEAM.md/instruction.md pattern is a hand-rolled analogue of this) operationalizes prompt engineering as repository-resident configuration rather than per-call strings:

  • CLAUDE.md — a steering file auto-loaded into context at session start, functionally identical in purpose to this repo's instruction.md files, but natively supported rather than manually pasted.
  • Subagents — isolated context windows with their own system prompt and tool allowlist, directly generalizing this chapter's own §9.2 review-chapter pattern (a subagent scoped only to "review chapter X" cannot accidentally leak into "write chapter Y"'s context).
  • Hooks — deterministic shell commands triggered on lifecycle events (before/after a tool call, on session end) — a non-probabilistic enforcement layer for exactly the kind of invariant Chapter 5's failure-mode list wants enforced (e.g., a hook that runs ast.parse() after every file write is a deterministic version of Chapter 5's AST Syntax Guard, external to the model entirely).
  • Skills / Plugins — packaged, invokable instruction+tool bundles for a recurring task, letting a team check a reusable "prompt program" into version control instead of re-deriving it per session — the practical implementation of Chapter 6's "prevent prompt drift" concern.

9.4.5 Claude Agent SDK, Structured Outputs, Citations, Code Execution, Files & Batch APIs

  • Claude Agent SDK — the same harness underlying Claude Code, exposed as a library for building custom agents with the tool loop, permissioning, and context-management primitives already solved, rather than reimplementing Chapter 5's dispatcher loop from scratch per project.
  • Structured Outputs / strict tool schemas — guarantee the response matches a declared JSON Schema, reducing Chapter 5's "1–3% schema failure rate" concern for the specific case of final-answer formatting (distinct from tool-call arguments, which already had high conformance).
  • Citations API — grounds generated claims to exact source-document spans, a directly relevant complement to Chapter 6's hallucination-detection eval tier: citation spans are a deterministic (Tier-1, near-zero-cost) check rather than a semantic (Tier-2, LLM-judged) one.
  • Code Execution Tool — a server-side sandboxed Python environment Claude can call as a tool, useful for exactly the "arithmetic/logic offload" concern Chapter 4 raises about CoT reliability on precise computation — instead of asking the model to reason its way to a correct sum, it can execute code and read back the result.
  • Files API / Batch API — see §9.2's Chapter 7 gap above.

9.4.6 Claude Developer Console: Prompt Generator & Evaluate tool

The Console includes a built-in prompt-drafting assistant (generates a first-draft structured prompt from a plain-language task description, already XML-tag-structured per §9.3) and an Evaluate tab (define test cases, run a prompt across them, compare versions side-by-side) — a lighter-weight complement to Chapter 6's from-scratch Python eval harness, appropriate for early-stage iteration before a project justifies the CI/CD-integrated harness Chapter 6 builds.


9.5 Updated Trade-Off Matrix: Adding the Claude Ecosystem Dimension

Extending Chapter 7's cost framing and Chapter 5's tool-calling matrix with the ecosystem options surfaced above (illustrative figures, consistent with this playbook's existing "illustrative, not audited-benchmark" numeric style):

Dimension Gemini-Centric Path (Ch01–08 default) Claude Ecosystem Path (this chapter) Engineering Takeaway
Repeated static-prefix cost Context caching described generically; exact discount/TTL left abstract. Prompt Caching: named cache_control breakpoints, ~90% discount on cache hits, default 5-min TTL (extendable). Same underlying KV-cache theory (Ch07); Claude path gives a concrete API knob to pull.
Reasoning transparency Manual <scratchpad> tags (Ch04) or opaque native reasoning. Extended/interleaved thinking blocks: separately budgeted, streamable, interleavable with tool calls. Claude path removes the scratchpad-vs-native binary Ch04 poses; thinking is inspectable and native.
Tool schema standardization Bespoke JSON schema per project (Ch05 lab). MCP: one server, many clients, schema defined once. Reduces Ch05's "Schema-Docstring Divergence" failure mode structurally, not just by discipline.
Repo-resident prompt governance This repo's own hand-rolled instruction.md + AGENT_TEAM.md convention. Claude Code CLAUDE.md + Subagents + Hooks natively support the same pattern. Validates this repo's own architecture — and shows where deterministic hooks can replace probabilistic self-correction loops (Ch05 §5.3) for enforceable invariants.
Early-stage eval iteration Build the Ch06 harness from scratch immediately. Console Evaluate tool for first-pass iteration, graduate to a custom harness (or Promptfoo/DeepEval, per Appendix B) at scale. Lowers time-to-first-eval; doesn't replace Ch06's CI/CD-grade harness for mature projects.
Vendor lock-in risk Low (playbook is intentionally multi-vendor). Moderate for MCP-specific tooling (open protocol, but ecosystem is Anthropic-seeded); low for core prompting techniques (XML tags, CoT, role prompts transfer across vendors). Prefer MCP and Claude Agent SDK for genuinely agentic systems; keep Chapters 1–8's vendor-neutral prompt patterns as the portable core.

9.6 Mandatory Hands-On Lab: Claude Ecosystem Alternate Answer

Restating the Chapter 5 Lab Challenge

Chapter 5's lab built a ResilientToolDispatcher: register a tool with a parameter spec, simulate an LLM emitting a malformed tool call, feed back a structured error, and confirm the second (corrected) attempt succeeds — capped at 2 correction attempts.

Recommended Answer #2 (Claude Ecosystem Edition)

This alternate solution solves the same core problem plus two Claude-specific extensions surfaced by this chapter's review:

  1. Tool schemas use Anthropic's real field name, input_schema (fixing §9.2's Chapter 5 Gap 1), and tool-use/tool-result payloads are modeled as Messages API content blocks (type: "tool_use" / type: "tool_result"), not a raw JSON string the client must json.loads() itself.
  2. A minimal prompt-caching cost simulator demonstrates §9.4.1: the same static system+tools prefix is "sent" twice, and the second call is billed at a steep discount because it hits the cache — turning Chapter 7's abstract diagram into a runnable, assertable number.

No anthropic SDK package is required — the script models the public Messages API block shapes as plain dataclasses so it runs fully offline with the Python 3.11+ standard library only.

"""
test_ch09_claude_ecosystem_engine.py
Zero-dependency Python 3.11+ engine for Chapter 9 (PB-01):
ClaudeStyleToolDispatcher + PromptCacheSimulator

Models the public shape of Anthropic's Messages API (content blocks,
input_schema tool definitions, cache_control breakpoints) as plain
dataclasses -- no `anthropic` package or network access required.
"""

from __future__ import annotations
from dataclasses import dataclass, field
from enum import Enum
from typing import Any, Callable, Dict, List, Optional, Tuple


# ---------------------------------------------------------------------------
# 1. Messages API Content Block Models (Anthropic-shaped, not OpenAI-shaped)
# ---------------------------------------------------------------------------

class BlockType(str, Enum):
    TEXT = "text"
    THINKING = "thinking"
    TOOL_USE = "tool_use"
    TOOL_RESULT = "tool_result"


@dataclass
class ContentBlock:
    type: BlockType
    text: Optional[str] = None
    id: Optional[str] = None
    name: Optional[str] = None
    input: Optional[Dict[str, Any]] = None          # tool_use arguments (already-parsed dict, unlike raw-JSON-string APIs)
    tool_use_id: Optional[str] = None
    content: Optional[str] = None                    # tool_result payload
    is_error: bool = False


@dataclass
class ToolDefinition:
    """Mirrors Anthropic's tool schema: 'input_schema', not 'parameters'."""
    name: str
    description: str
    input_schema: Dict[str, Any]
    handler: Callable[..., Dict[str, Any]]
    cache_control: Optional[Dict[str, str]] = None    # e.g. {"type": "ephemeral"}

    def validate_arguments(self, raw_args: Dict[str, Any]) -> Tuple[bool, Optional[str]]:
        props = self.input_schema.get("properties", {})
        required = self.input_schema.get("required", [])
        for field_name in required:
            if field_name not in raw_args:
                return False, f"Missing required parameter '{field_name}'."
        for key, val in raw_args.items():
            spec = props.get(key)
            if spec is None:
                continue
            expected = {"string": str, "integer": int, "number": (int, float), "boolean": bool}.get(spec.get("type"))
            if expected and not isinstance(val, expected):
                return False, (
                    f"Type mismatch for parameter '{key}': expected {spec.get('type')}, "
                    f"got {type(val).__name__}."
                )
            enum_vals = spec.get("enum")
            if enum_vals and val not in enum_vals:
                return False, f"Value '{val}' for '{key}' not in allowed enum {enum_vals}."
        return True, None


# ---------------------------------------------------------------------------
# 2. Resilient Claude-Style Tool Dispatcher (content-block native)
# ---------------------------------------------------------------------------

class ClaudeStyleToolDispatcher:
    """Dispatches tool_use content blocks and produces tool_result blocks,
    matching the Messages API multi-turn tool-use conversation shape."""

    def __init__(self, max_correction_attempts: int = 2):
        self.registry: Dict[str, ToolDefinition] = {}
        self.max_correction_attempts = max_correction_attempts
        self.audit_log: List[Dict[str, Any]] = []

    def register_tool(self, tool: ToolDefinition) -> None:
        self.registry[tool.name] = tool

    def dispatch(self, block: ContentBlock) -> ContentBlock:
        assert block.type == BlockType.TOOL_USE, "dispatch() requires a tool_use block"
        tool = self.registry.get(block.name)
        if tool is None:
            return ContentBlock(
                type=BlockType.TOOL_RESULT,
                tool_use_id=block.id,
                content=f"Tool '{block.name}' not found. Available: {list(self.registry.keys())}",
                is_error=True,
            )

        is_valid, err = tool.validate_arguments(block.input or {})
        if not is_valid:
            return ContentBlock(
                type=BlockType.TOOL_RESULT,
                tool_use_id=block.id,
                content=f"Validation error: {err}",
                is_error=True,
            )

        try:
            result = tool.handler(**block.input)
            self.audit_log.append({"tool": block.name, "args": block.input, "status": "SUCCESS"})
            return ContentBlock(
                type=BlockType.TOOL_RESULT,
                tool_use_id=block.id,
                content=str(result),
                is_error=False,
            )
        except Exception as exc:
            return ContentBlock(
                type=BlockType.TOOL_RESULT,
                tool_use_id=block.id,
                content=f"Execution failure: {exc}",
                is_error=True,
            )

    def run_agent_turn(self, simulated_tool_use_emissions: List[ContentBlock]) -> Tuple[bool, ContentBlock]:
        """Simulates Claude emitting successive tool_use blocks (with an
        interleaved-thinking-style retry) until success or the correction
        budget (max_correction_attempts) is exhausted."""
        for attempt_idx, block in enumerate(simulated_tool_use_emissions):
            result = self.dispatch(block)
            if not result.is_error:
                return True, result
            if attempt_idx >= self.max_correction_attempts:
                return False, result
        return False, ContentBlock(type=BlockType.TOOL_RESULT, content="NO_SUCCESSFUL_EXECUTION", is_error=True)


# ---------------------------------------------------------------------------
# 3. Prompt Cache Simulator (concretizes Chapter 7's KV-cache diagram)
# ---------------------------------------------------------------------------

@dataclass
class CacheableRequest:
    static_prefix_tokens: int      # system identity + tool definitions (Ch08 Tiers 1-2)
    volatile_suffix_tokens: int    # per-call user payload (Ch08 Tiers 3-5)


class PromptCacheSimulator:
    """Models Anthropic prompt caching economics: cache-write premium on
    first call, steep discount on cache-hit for subsequent calls within TTL."""

    BASE_INPUT_RATE_PER_MTOK = 3.00     # illustrative $/1M input tokens
    CACHE_WRITE_MULTIPLIER = 1.25       # premium over base rate on cache miss/write
    CACHE_READ_MULTIPLIER = 0.10        # steep discount on cache hit

    def cost_of_call(self, req: CacheableRequest, cache_hit: bool) -> float:
        prefix_rate = self.CACHE_READ_MULTIPLIER if cache_hit else self.CACHE_WRITE_MULTIPLIER
        prefix_cost = (req.static_prefix_tokens / 1_000_000) * self.BASE_INPUT_RATE_PER_MTOK * prefix_rate
        suffix_cost = (req.volatile_suffix_tokens / 1_000_000) * self.BASE_INPUT_RATE_PER_MTOK
        return round(prefix_cost + suffix_cost, 6)

    def simulate_two_call_sequence(self, req: CacheableRequest) -> Tuple[float, float, float]:
        """Returns (first_call_cost, second_call_cost, pct_savings_on_second_call)."""
        first = self.cost_of_call(req, cache_hit=False)
        second = self.cost_of_call(req, cache_hit=True)
        savings_pct = round((1 - (second / first)) * 100, 2) if first > 0 else 0.0
        return first, second, savings_pct


# ---------------------------------------------------------------------------
# 4. Executable Verification Harness
# ---------------------------------------------------------------------------

def rolling_restart_handler(cluster: str, service_id: str, replica_count: int) -> Dict[str, Any]:
    return {
        "cluster": cluster,
        "service_id": service_id,
        "replica_count": replica_count,
        "action": "ROLLING_RESTART_INITIATED",
    }


if __name__ == "__main__":
    import unittest

    class TestClaudeEcosystemChapter9(unittest.TestCase):
        def setUp(self):
            self.dispatcher = ClaudeStyleToolDispatcher(max_correction_attempts=2)
            self.dispatcher.register_tool(ToolDefinition(
                name="rolling_restart_service",
                description="Initiates rolling restart of a microservice container fleet.",
                input_schema={
                    "type": "object",
                    "properties": {
                        "cluster": {"type": "string", "enum": ["prod", "staging"]},
                        "service_id": {"type": "string"},
                        "replica_count": {"type": "integer"},
                    },
                    "required": ["cluster", "service_id", "replica_count"],
                },
                handler=rolling_restart_handler,
                cache_control={"type": "ephemeral"},
            ))

        def test_tool_use_block_schema_uses_input_schema_not_parameters(self):
            tool = self.dispatcher.registry["rolling_restart_service"]
            self.assertTrue(hasattr(tool, "input_schema"))
            self.assertFalse(hasattr(tool, "parameters"))

        def test_bounded_self_correction_succeeds_on_second_attempt(self):
            # Attempt 1: flawed -- replica_count as string (Ch05's exact failure mode)
            attempt_1 = ContentBlock(
                type=BlockType.TOOL_USE, id="call_1", name="rolling_restart_service",
                input={"cluster": "prod", "service_id": "auth-api", "replica_count": "3"},
            )
            # Attempt 2: remediated -- integer type
            attempt_2 = ContentBlock(
                type=BlockType.TOOL_USE, id="call_2", name="rolling_restart_service",
                input={"cluster": "prod", "service_id": "auth-api", "replica_count": 3},
            )
            success, result_block = self.dispatcher.run_agent_turn([attempt_1, attempt_2])
            self.assertTrue(success)
            self.assertEqual(result_block.type, BlockType.TOOL_RESULT)
            self.assertIn("ROLLING_RESTART_INITIATED", result_block.content)
            self.assertFalse(result_block.is_error)

        def test_unknown_tool_returns_error_tool_result(self):
            bad_block = ContentBlock(type=BlockType.TOOL_USE, id="call_x", name="nonexistent_tool", input={})
            result = self.dispatcher.dispatch(bad_block)
            self.assertTrue(result.is_error)
            self.assertEqual(result.tool_use_id, "call_x")

        def test_prompt_cache_simulator_shows_material_second_call_savings(self):
            sim = PromptCacheSimulator()
            req = CacheableRequest(static_prefix_tokens=4000, volatile_suffix_tokens=200)
            first, second, savings_pct = sim.simulate_two_call_sequence(req)
            self.assertGreater(first, second, "Cache-hit call must be cheaper than cache-write call")
            self.assertGreaterEqual(savings_pct, 60.0, "Cache hit should deliver material (>=60%) savings on the cacheable prefix")

    suite = unittest.TestLoader().loadTestsFromTestCase(TestClaudeEcosystemChapter9)
    runner = unittest.TextTestRunner(verbosity=2)
    test_result = runner.run(suite)
    if not test_result.wasSuccessful():
        exit(1)

    # Demonstration output (mirrors Ch05's dispatcher demo + adds cache economics)
    print("\n=== Chapter 9 Demonstration: Claude-Style Dispatch + Prompt Cache Economics ===")
    sim = PromptCacheSimulator()
    req = CacheableRequest(static_prefix_tokens=4000, volatile_suffix_tokens=200)
    first_cost, second_cost, pct = sim.simulate_two_call_sequence(req)
    print(f"Call 1 (cache write): ${first_cost:.6f}")
    print(f"Call 2 (cache hit):   ${second_cost:.6f}  ({pct}% cheaper on the shared static prefix)")

    print("\n[SUCCESS] All Chapter 9 Unit Tests Passed Successfully (100% Conformance).")

Verification: this script was extracted and executed standalone under Python 3.11+; all four unit tests pass and the demonstration block prints a cache-hit saving in excess of 60% on the shared static prefix, consistent with Chapter 7's original (previously unattributed) "90% discount" claim.


9.7 Summary

This chapter reviewed Chapters 1–8 of Playbook 01 chapter-by-chapter, surfacing one concrete, fixable defect (the parameters vs. input_schema field-name mismatch in Chapter 5) and several structural gaps where the playbook's own correct theory (KV-cache economics, native reasoning models, tool schema governance) was never connected to a concrete production surface. It then supplied that missing concrete layer for the Anthropic/Claude ecosystem specifically:

  • Constitutional AI and context engineering as the framing shift underlying modern Claude-based system design.
  • Prompt Caching, Extended/Interleaved Thinking, MCP, Claude Code, the Claude Agent SDK, and the Developer Console's built-in tooling as the concrete implementations of concepts this playbook already derived from first principles.
  • An updated trade-off matrix placing the Claude ecosystem path beside the playbook's existing Gemini-centric default.
  • A second, independently runnable solution to the Chapter 5 hands-on lab, built on Anthropic's real Messages API content-block shapes, that also fixes the schema-field defect identified in the review and turns Chapter 7's cache diagram into an assertable number.

Relationship to the Rest of Playbook 01

This chapter is additive: Chapters 1–8 remain the vendor-neutral foundation, and the templates in Chapter 8 remain valid verbatim on Claude (its XML-tag preference makes it, if anything, an easier target for that blueprint than some peer models). Treat this chapter as the deployment-specific appendix to consult when the target production vendor is Anthropic.