Overview
Chapter 09: The Anthropic/Claude Ecosystem Perspective — A Critical Companion & Playbook-Wide Review
Playbook Track: 05 – Presentation Slides & Architecture Diagrams (Visual Systems, Claude Code & Antigravity Workflows) Target Audience: Year 1 Computer Science & Software Engineering Students Core Tooling Stack: Claude Sonnet 5 / Opus 5 / Haiku 4.5, Claude Code (subagents, hooks, skills, plugins), Claude Agent SDK, Model Context Protocol (MCP), Artifacts, Claude in Chrome, extended thinking, prompt caching, Python 3.11+ Delivery Status: 🔍 Ready for Review (Tier 1 Markdown) — Companion & Review Chapter
1. Executive Mission: Why This Chapter Exists
Chapters 01–08 of this playbook are architecturally sound and technically rigorous. They are also, without exception, written against a single moving target: Gemini 2.5 Pro/Flash as the reasoning brain and Claude 3.7 Sonnet (Claude Code CLI) as one of two interchangeable coding-agent front-ends (the other being Antigravity agy). That framing was defensible when Chapter 05 was drafted, but it under-sells what an Anthropic-centric practitioner would actually build today, and it leaves the playbook's own quality gates — Gate 3 ("Latest frontier configurations") and Gate 4 ("Quantitative trade-offs") — partially unmet with respect to the Claude side of the stack.
This chapter has two jobs, and it does not re-derive the diagramming math, C4 theory, or geometry engines already covered — it assumes you've read Chapters 01–08:
- An independent review of the existing eight chapters and two appendices, chapter-by-chapter, with concrete critical findings — not a rubber stamp.
- A complementary Anthropic/Claude viewpoint: what changes, architecturally and operationally, if Claude — not Gemini — is the reasoning and orchestration layer, and if Claude Code is treated as the primary agentic authoring surface rather than a peer option to Antigravity.
This chapter does not replace Chapter 05. Chapter 05's CLAUDE.md and SKILL.md runbooks remain operationally valid. This chapter supersedes Chapter 05's currency — model naming, MCP maturity, and the missing distribution model (Artifacts) — and broadens the lens to the whole playbook, not just the CLI workflow chapter.
2. Chapter-by-Chapter Review
| Chapter | Strengths | Critical Findings |
|---|---|---|
| Ch 01 — Visual Communication Shift | Strong grounding in Dual-Coding Theory, Cognitive Load Theory, and Gestalt principles; the DeclarativeDiagramParser AST approach is the correct engineering pattern. |
(1) The regex-based Mermaid tokenizer (ARROW_REGEX, _parse_node_token) will silently mis-parse nested brackets, multi-line HTML labels (<br/>), and quoted commas — exactly the syntax Chapters 03 and 07 later generate. No fuzz-testing or grammar-based parser (e.g., a real PEG/Lark grammar) is proposed as a hardening path. (2) "Gemini 2.5 Pro & Claude 3.7 Sonnet" are named as interchangeable diagram generators with zero comparative data — the playbook's own Gate 4 (quantitative trade-offs) is not applied to its own model choice. |
| Ch 02 — Diagramming-as-Code Engines | Excellent comparative mechanics of Dagre vs. Graphviz vs. TALA; the "Big Three" framing is pedagogically clean. | (1) Only gemini-2.5-flash is calibrated for DSL generation — no evidence was gathered on which model produces lower syntax-error rates across Mermaid/D2/PlantUML, despite this being exactly the kind of strict-format, low-temperature task Claude's tool-use and structured-output training is built for. (2) The multi-engine transpiler is described but never defends against semantic loss when transpiling D2's native container nesting down to Mermaid's flatter subgraph model — a real fidelity gap the chapter doesn't name. |
| Ch 03 — C4 Model & Structurizr | Rigorous C4 ontology, correct "Single Source of Truth" argument for Structurizr DSL. | (1) No worked example of an LLM-driven C4 boundary-violation linter catching a specific hallucinated cross-level edge — the anti-boxology claims stay abstract. (2) No discussion of long-context handling for genuinely large monorepos (750k+ LOC, per Ch07's own benchmark table) — which model's context window and retrieval strategy actually survives that scale is left unaddressed. |
| Ch 04 — Programmatic Slides (PPTX/Marp) | The EMU/16:9 coordinate geometry section is precise and directly testable — a real engineering strength. | (1) Every generated deck round-trips through a file-based compile step (marp-cli, python-pptx save()) before a human can see it — there is no mention of a live, shareable, interactively-editable preview surface, which materially changes review cycle time. (2) No treatment of how a reviewer actually comments on a specific slide's geometry without opening PowerPoint. |
| Ch 05 — Claude Code & Antigravity Workflows | Correctly identifies CLAUDE.md, .claude/skills/, and MCP as the right integration seams; the 10 pitfalls (shell injection, context bloat, recursive subagent loops) are genuinely useful defenses. |
This chapter is the most dated in the playbook and is explicitly updated here: (1) It pins Claude 3.7 Sonnet — current production tiers are Claude Opus 5, Sonnet 5, and Haiku 4.5, with materially different cost/latency/reasoning-depth trade-offs the chapter never models. (2) It treats Claude Code as co-equal to Antigravity CLI rather than examining Claude Code's specific differentiators: native hooks (deterministic pre/post-tool-call gating, not just prompted instructions), first-class subagents with isolated context windows, a plugin marketplace, and background/parallel task execution — none of which appear in the chapter's architecture diagram. (3) The MCP tool bus is described only as a pass-through to mermaid-server/marp-server; there is no mention of MCP as an open, bidirectional protocol Anthropic authored and open-sourced, nor of the growing catalog of maintained MCP servers beyond this playbook's bespoke ones. (4) The chapter's Visual QA Auditor box is a separate stage — it does not use Claude in Chrome to actually look at the rendered artifact in a live browser and screenshot-diff it, which is the natural automation for this exact step. (5) Zero mention of Artifacts as a publication target — every output in Ch05 is a file on disk, never a live, theme-aware, shareable page. |
| Ch 06 — Multimodal Vision Critique | The WCAG luminance math and AABB collision formalism are correct and appropriately kept deterministic rather than delegated to an LLM — good engineering judgment. | (1) The semantic critique prompt is generic ("evaluate visual hierarchy") and is not grounded with citations back to the specific WCAG success criterion or cognitive-load source it's enforcing — a Citations-style grounded-quote pattern would make audit findings traceable and defensible in a compliance review. (2) No mention that native multi-image input (comparing a slide against its previous revision, or against a competitor's deck) is a distinct capability class from single-image critique, and the chapter never exercises it. |
| Ch 07 — End-to-End Deck & Brief Automation | The strongest chapter in the playbook: full pipeline, 10-slide narrative arc, honest benchmark table, 10 named architectural drift traps. This chapter is the basis for this chapter's alternate lab answer (Section 6). | (1) Claude 3.7 Sonnet is again the dated pin for the "Executive Deck Synthesizer" role. (2) The closed-loop visual audit (Section 1, sequenceDiagram) caps at 3 iterations and falls back to "a deterministic safe layout template" — but the fallback template is never specified, and there is no deterministic hook enforcing the cap outside of prompted agent behavior, meaning a misbehaving agent could still loop past 3 turns before anything stops it. (3) GitOps publishing produces a Git commit as the end state; there is no interactive preview URL a non-technical stakeholder (VP, ARB reviewer) could open without cloning the repo or running marp-cli themselves. |
| Ch 08 — GitHub Ecosystem & Tooling Suite | Strong, well-licensed catalog (MIT/Apache-2.0 bias correctly flagged as a commercial-safety concern); the 6-dimension scoring model is a legitimate evaluation framework. | (1) modelcontextprotocol/servers is listed, but the catalog stops there — it omits the Claude Agent SDK (claude-agent-sdk-python / -typescript) repositories and the Claude Cookbooks repository, both of which are the actual reference implementations a team would clone to build what Chapter 05 only prototypes in pseudo-code. (2) No entry for Claude Code's own plugin/skill packaging conventions, which is a gap given Chapter 05 depends on exactly that mechanism. |
| App A — System Prompts & Design Tokens | Clean, reusable prompt library; WCAG-verified token palette is a genuinely useful reference artifact. | Prompts are still labeled for Claude 3.7 Sonnet; none of them use prompt caching despite the same ArchitectureModel schema and design-token block being re-sent as static context on every single slide-generation call across a 10-slide deck — this is the single most mechanical, low-risk optimization missing from the entire playbook. |
| App B — CLI Runbooks | Solid, copy-pasteable install/runbook instructions for Marp, D2, and headless Chromium. | The Claude Code CLI setup section predates hooks, skills-as-plugins, and the current permission-mode model (--dangerously-skip-permissions framing elsewhere in this repo is a symptom of the same staleness) — the runbook should be refreshed alongside Chapter 05. |
3. The Anthropic/Claude Viewpoint
Three framing differences matter more than any single tool substitution:
- Context engineering, not prompt engineering, as the unit of work. The playbook's prompts (App A, Ch01, Ch07) are well-written single-shot instructions, but the actual failure mode in a 10-slide, multi-agent pipeline is not "the prompt was unclear" — it's "the design-token block, the
ArchitectureModel, and the WCAG rule set were re-explained from scratch on every call." Anthropic's own agent-building guidance treats the curation of what's in the context window and for how long as the primary engineering surface, with prompt caching as the mechanical enforcement of that discipline. This playbook has the right data model (ArchitectureModelas single source of truth) but never caches it. - Determinism where it counts, judgment where it doesn't. Chapter 06's decision to hand-code WCAG contrast and AABB collision math in Python rather than asking a vision model "does this look okay?" is exactly right, and it's a genuinely Anthropic-aligned instinct: use the model for the parts that require judgment (visual hierarchy, narrative pacing, audience-fit) and use deterministic code, gated by hooks, for the parts that have a mathematically correct answer. The gap is that this discipline is applied inconsistently — Chapter 07's iteration cap is described in prose but not enforced in code the way Chapter 06's contrast ratio is.
- Artifacts change what "done" means. Every pipeline in this playbook terminates at a file: an
.svg, a.pptx, a.pdf, a Git commit. That's correct for source-of-truth storage, but it means a reviewer's feedback loop is "pull the branch, open PowerPoint, look, comment in Slack." An Artifact is a live, theme-aware, directly-shareable web page — publishing the compiled Marp deck or C4 diagram as an Artifact means the ARB reviewer opens a link, sees the actual rendered result (not a screenshot), and the next iteration republishes to the same URL. This doesn't replace the Git-committed declarative source (Ch08's version-control argument is correct and unchanged) — it adds a distribution layer on top of it.
4. Latest Claude Ecosystem Tools & Techniques Mapped to Visual Engineering
Only real, currently-shipping Anthropic products are referenced below.
4.1 Current Model Tier (replaces "Claude 3.7 Sonnet" throughout Ch01–App B)
Claude Opus 5 (claude-opus-5) for the Architectural Ontologist role (Ch07) where reasoning depth over a large, ambiguous codebase matters most; Claude Sonnet 5 (claude-sonnet-5) as the default workhorse for diagram/DSL generation and the Executive Deck Synthesizer role — the same low-temperature, strict-format profile the playbook already calibrates Gemini Flash for; Claude Haiku 4.5 (claude-haiku-4-5-20251001) for high-volume, low-latency tasks like per-slide word-count linting or contrast-ratio pre-checks where Opus/Sonnet-grade reasoning is unnecessary overhead.
4.2 Claude Code as Primary Agentic Surface (updates Ch05)
Beyond the CLAUDE.md/.claude/skills/ mechanics Ch05 already documents: subagents give the Architecture Extractor and Deck Compiler roles genuinely isolated context windows (preventing the "semantic drift across multi-agent hand-offs" failure Ch07 names as Trap 6, by construction rather than by shared-dictionary discipline alone); hooks (PreToolUse/PostToolUse) let the WCAG/AABB gate from Ch06 run as a deterministic, non-bypassable check before any publish tool call fires, closing the enforcement gap identified in Section 2's Ch07 review; plugins package the presentation-architect skill for reuse across repositories instead of copy-pasting SKILL.md.
4.3 Model Context Protocol (MCP) — Maturity Beyond Ch05's Bespoke Servers
MCP is an open standard Anthropic authored and open-sourced specifically so mermaid-server and marp-server (Ch05's hand-rolled examples) don't need to be reinvented per-project. The reference server catalog (modelcontextprotocol/servers) and community servers for headless browser control, filesystem access, and Git operations mean the tool bus in Ch05's architecture diagram can be substantially pre-built rather than hand-coded.
4.4 Artifacts as the Distribution Layer
Publishing a compiled Marp deck or rendered C4 diagram as an Artifact turns Ch07's Git-commit terminal state into a live, linkable, theme-aware page. Because Artifacts render Mermaid natively, the Ch01–Ch03 diagram outputs can be previewed with zero compile step during iteration, with the declarative source still committed to Git as the system of record per Ch08's version-control argument.
4.5 Claude in Chrome — Closing Ch05's Visual QA Loop
Ch05's architecture diagram shows a "Visual QA Auditor" box downstream of "Delivery Directory" with no mechanism specified. Claude in Chrome can navigate to a rendered Artifact or local HTML preview, capture a screenshot, and feed it directly into the Ch06 multimodal critique step — closing the loop the original diagram left open, without a bespoke Puppeteer MCP server.
4.6 Prompt Caching — The Missing Optimization in App A
Every prompt in App A re-sends the full ArchitectureModel JSON schema, the design-token palette, and the 10-slide narrative-arc rules on each of the 10 per-slide generation calls in Ch07's pipeline. Marking that static block as a cached prefix (cacheable for repeated reuse within the session) removes the largest source of redundant token spend in the entire playbook — larger, in relative terms, than any of the compute optimizations Ch07's benchmark table reports.
4.7 Extended Thinking for the Architectural Ontologist Role
Ch07's Architectural Ontologist step (service graph extraction from a 750k-LOC monorepo) is exactly the class of problem — multi-step, ambiguous, requiring the model to weigh conflicting signals like stale ADRs against live code (Ch07's own Trap 2) — where extended thinking measurably improves plan quality before the model commits to an output schema.
5. Updated Trade-Off Matrix
| Dimension | Playbook's Current Approach (Gemini + Claude 3.7 CLI, file-first) | Claude-Ecosystem-Native Approach (Sonnet 5 / Opus 5, Artifacts-first) |
|---|---|---|
| Reasoning role clarity | Gemini and Claude treated as interchangeable across Ch01–Ch07 with no comparative data | Opus 5 for ambiguous extraction, Sonnet 5 for strict-format generation, Haiku 4.5 for high-volume linting — role-differentiated by cost/latency profile |
| Review distribution | File on disk → Git commit → manual clone/open cycle | Artifact URL, live and theme-aware, republished per iteration |
| QA gate enforcement | Prompted iteration cap (Ch07, "max 3 iterations" in prose) | Hook-enforced hard stop, non-bypassable regardless of agent behavior |
| Cross-agent consistency | Shared ArchitectureModel dictionary by convention (Ch07 Trap 6 defense) |
Subagent isolation + shared model, drift prevented structurally, not just by discipline |
| Token economics across a 10-slide deck | Full schema/token re-sent per slide (App A prompts) | Cached static context reused across all 10 slide-generation calls |
| Visual QA automation | Described box, no concrete implementation (Ch05 diagram) | Claude in Chrome screenshot capture feeding Ch06's multimodal critique directly |
| Vendor lock-in risk | Two-vendor dependency (Gemini + Claude) with no documented fallback path | Single-vendor for reasoning layer; open MCP standard and Git-committed declarative source keep the artifacts themselves portable |
(Figures above are architectural/qualitative comparisons for this playbook's internal pedagogy, consistent with the illustrative benchmark style used in Chapters 01–08 — not citations to external published studies.)
6. Mandatory Hands-On Lab — Claude Ecosystem Alternate Answer
Restating the Original Lab (Chapter 07)
Chapter 07's lab has you build an ArchitectureModel for a payment platform, compile it to a C4 Mermaid container diagram via C4DiagramCompiler, compile a 10-slide Marp deck via ExecutiveDeckCompiler, and run EndToEndPipelineOrchestrator — verifying all nodes/edges appear in the diagram, all 10 slides are present, and invalid dependencies are rejected.
The Claude Ecosystem Extension
Solve the equivalent problem, but modeled the way a Claude Code / Claude Agent SDK pipeline actually executes: as a sequence of tool_use / tool_result message blocks, gated before publication by a hook, and terminating in an Artifact publish manifest instead of a bare file write. This directly addresses the two enforcement gaps found in Section 2: the missing hard iteration/quality cap, and the missing distribution layer.
Recommended Answer #2 (Claude Ecosystem Edition)
"""
test_ch09_claude_ecosystem_engine.py
Zero-dependency Python 3.11+ engine for PB-05 Chapter 9:
Claude-ecosystem-flavored orchestration for the Chapter 07 deck pipeline —
tool_use/tool_result messaging, a PreToolUse-style publish hook, and an
Artifact-style publish manifest. No `anthropic` package required: the
Messages API tool-use schema and hook contract are modeled as plain
dataclasses so this runs fully offline.
"""
import unittest
from dataclasses import dataclass, field
from enum import Enum
from typing import Any, Dict, List, Optional
# ---------------------------------------------------------------------------
# 1. Anthropic Messages API-style tool_use / tool_result modeling
# ---------------------------------------------------------------------------
class BlockType(str, Enum):
TEXT = "text"
TOOL_USE = "tool_use"
TOOL_RESULT = "tool_result"
@dataclass
class ToolUseBlock:
id: str
name: str
input: Dict[str, Any]
type: BlockType = BlockType.TOOL_USE
@dataclass
class ToolResultBlock:
tool_use_id: str
content: Any
is_error: bool = False
type: BlockType = BlockType.TOOL_RESULT
@dataclass
class ClaudeMessage:
role: str # "assistant" | "user"
blocks: List[Any] = field(default_factory=list)
# ---------------------------------------------------------------------------
# 2. Domain model (mirrors Ch07's ArchitectureModel, kept minimal here)
# ---------------------------------------------------------------------------
class ComponentType(str, Enum):
WEB_CLIENT = "Web Client"
API_GATEWAY = "API Gateway"
MICROSERVICE = "Microservice"
DATABASE = "Database"
@dataclass
class ServiceNode:
id: str
name: str
type: ComponentType
@dataclass
class ServiceDependency:
source_id: str
target_id: str
protocol: str
@dataclass
class ArchitectureModel:
system_name: str
nodes: Dict[str, ServiceNode] = field(default_factory=dict)
dependencies: List[ServiceDependency] = field(default_factory=list)
def add_node(self, node: ServiceNode):
self.nodes[node.id] = node
def add_dependency(self, dep: ServiceDependency):
if dep.source_id not in self.nodes or dep.target_id not in self.nodes:
raise ValueError(f"Unknown node in dependency: {dep.source_id} -> {dep.target_id}")
self.dependencies.append(dep)
@dataclass
class SlideSpec:
title: str
body: str
# ---------------------------------------------------------------------------
# 3. Tool implementations the "agent" calls via tool_use blocks
# ---------------------------------------------------------------------------
def tool_compile_c4_diagram(model: ArchitectureModel) -> str:
lines = ["graph TB", f' subgraph Boundary["{model.system_name}"]']
for node in model.nodes.values():
shape = "[({})]" if node.type == ComponentType.DATABASE else "[{}]"
lines.append(f" {node.id}{shape.format(node.name)}")
lines.append(" end")
for dep in model.dependencies:
lines.append(f' {dep.source_id} -->|"{dep.protocol}"| {dep.target_id}')
return "\n".join(lines)
def tool_compile_slide_deck(model: ArchitectureModel, diagram: str) -> List[SlideSpec]:
return [
SlideSpec("Title & Overview", f"{model.system_name}: automated architecture brief."),
SlideSpec("System Context", f"{len(model.nodes)} nodes, {len(model.dependencies)} dependencies."),
SlideSpec("Container Topology", f"```mermaid\n{diagram}\n```"),
SlideSpec("Roadmap", "Phase 1: stabilize ingress. Phase 2: scale reads."),
]
# ---------------------------------------------------------------------------
# 4. Hook: deterministic pre-publish quality gate (hard stop, not prompted)
# ---------------------------------------------------------------------------
class HookDecision(str, Enum):
ALLOW = "allow"
BLOCK = "block"
@dataclass
class HookResult:
decision: HookDecision
reasons: List[str] = field(default_factory=list)
# Minimal WCAG-style luminance/contrast check (mirrors Ch06's formalism)
def _relative_luminance(rgb: tuple[int, int, int]) -> float:
def channel(c: int) -> float:
c_norm = c / 255.0
return c_norm / 12.92 if c_norm <= 0.03928 else ((c_norm + 0.055) / 1.055) ** 2.4
r, g, b = (channel(c) for c in rgb)
return 0.2126 * r + 0.7152 * g + 0.0722 * b
def _contrast_ratio(fg: tuple[int, int, int], bg: tuple[int, int, int]) -> float:
l1, l2 = sorted([_relative_luminance(fg), _relative_luminance(bg)], reverse=True)
return (l1 + 0.05) / (l2 + 0.05)
def pre_publish_hook(slides: List[SlideSpec], text_color: tuple[int, int, int],
bg_color: tuple[int, int, int], max_words_per_slide: int = 30) -> HookResult:
"""PreToolUse-style gate: blocks the 'publish' tool call, not just warns."""
reasons: List[str] = []
contrast = _contrast_ratio(text_color, bg_color)
if contrast < 4.5:
reasons.append(f"WCAG 2.1 AA contrast violation: {contrast:.2f} < 4.5:1")
for slide in slides:
word_count = len(slide.body.split())
if word_count > max_words_per_slide and "```mermaid" not in slide.body:
reasons.append(f"Slide '{slide.title}' exceeds word budget: {word_count} > {max_words_per_slide}")
if not slides:
reasons.append("No slides produced; refusing to publish an empty deck.")
decision = HookDecision.BLOCK if reasons else HookDecision.ALLOW
return HookResult(decision=decision, reasons=reasons)
# ---------------------------------------------------------------------------
# 5. Artifact-style publish manifest
# ---------------------------------------------------------------------------
@dataclass
class ArtifactPublishManifest:
title: str
slide_count: int
url_slug: str
theme_aware: bool = True
status: str = "published"
# ---------------------------------------------------------------------------
# 6. Orchestrator: tool_use/tool_result trace -> hook gate -> Artifact manifest
# ---------------------------------------------------------------------------
class ClaudeEcosystemDeckOrchestrator:
def __init__(self, text_color=(15, 23, 42), bg_color=(255, 255, 255)):
self.text_color = text_color
self.bg_color = bg_color
self.trace: List[ClaudeMessage] = []
def run(self, model: ArchitectureModel) -> ArtifactPublishManifest:
# --- Simulated turn 1: agent requests the diagram-compile tool ---
call1 = ToolUseBlock(id="call_1", name="compile_c4_diagram", input={"system": model.system_name})
self.trace.append(ClaudeMessage(role="assistant", blocks=[call1]))
diagram = tool_compile_c4_diagram(model)
self.trace.append(ClaudeMessage(role="user", blocks=[ToolResultBlock(tool_use_id="call_1", content=diagram)]))
# --- Simulated turn 2: agent requests the slide-compile tool ---
call2 = ToolUseBlock(id="call_2", name="compile_slide_deck", input={"system": model.system_name})
self.trace.append(ClaudeMessage(role="assistant", blocks=[call2]))
slides = tool_compile_slide_deck(model, diagram)
self.trace.append(ClaudeMessage(role="user", blocks=[ToolResultBlock(tool_use_id="call_2", content=slides)]))
# --- Deterministic hook gate before the 'publish' tool is allowed to fire ---
gate = pre_publish_hook(slides, self.text_color, self.bg_color)
if gate.decision == HookDecision.BLOCK:
raise PermissionError(f"pre_publish_hook blocked publication: {'; '.join(gate.reasons)}")
# --- Publish as an Artifact manifest, not a bare file write ---
slug = model.system_name.lower().replace(" ", "-")
return ArtifactPublishManifest(title=model.system_name, slide_count=len(slides), url_slug=slug)
# ---------------------------------------------------------------------------
# Self-Test Verification Suite
# ---------------------------------------------------------------------------
class TestClaudeEcosystemDeckOrchestrator(unittest.TestCase):
def setUp(self):
self.model = ArchitectureModel(system_name="NimbusCommerce Platform")
self.model.add_node(ServiceNode("web", "Web Portal", ComponentType.WEB_CLIENT))
self.model.add_node(ServiceNode("gw", "API Gateway", ComponentType.API_GATEWAY))
self.model.add_node(ServiceNode("orders", "Order Service", ComponentType.MICROSERVICE))
self.model.add_node(ServiceNode("db", "Order DB", ComponentType.DATABASE))
self.model.add_dependency(ServiceDependency("web", "gw", "HTTPS/JSON"))
self.model.add_dependency(ServiceDependency("gw", "orders", "gRPC"))
self.model.add_dependency(ServiceDependency("orders", "db", "SQL/TCP"))
def test_tool_use_trace_recorded(self):
orch = ClaudeEcosystemDeckOrchestrator()
orch.run(self.model)
tool_calls = [b.name for m in orch.trace for b in m.blocks if isinstance(b, ToolUseBlock)]
self.assertEqual(tool_calls, ["compile_c4_diagram", "compile_slide_deck"])
def test_pipeline_produces_artifact_manifest(self):
orch = ClaudeEcosystemDeckOrchestrator()
manifest = orch.run(self.model)
self.assertEqual(manifest.title, "NimbusCommerce Platform")
self.assertEqual(manifest.slide_count, 4)
self.assertEqual(manifest.url_slug, "nimbuscommerce-platform")
self.assertTrue(manifest.theme_aware)
def test_hook_blocks_low_contrast_publication(self):
# Light-gray text on white: fails WCAG 2.1 AA (~1.6:1 contrast)
orch = ClaudeEcosystemDeckOrchestrator(text_color=(220, 220, 220), bg_color=(255, 255, 255))
with self.assertRaises(PermissionError) as ctx:
orch.run(self.model)
self.assertIn("WCAG", str(ctx.exception))
def test_hook_blocks_empty_deck(self):
empty_model = ArchitectureModel(system_name="Empty System")
result = pre_publish_hook([], (15, 23, 42), (255, 255, 255))
self.assertEqual(result.decision, HookDecision.BLOCK)
self.assertIn("empty deck", result.reasons[0])
def test_dependency_validation_rejects_unknown_node(self):
bad_model = ArchitectureModel(system_name="Broken")
bad_model.add_node(ServiceNode("a", "A", ComponentType.MICROSERVICE))
with self.assertRaises(ValueError):
bad_model.add_dependency(ServiceDependency("a", "ghost", "gRPC"))
if __name__ == "__main__":
suite = unittest.TestLoader().loadTestsFromTestCase(TestClaudeEcosystemDeckOrchestrator)
runner = unittest.TextTestRunner(verbosity=2)
result = runner.run(suite)
if not result.wasSuccessful():
exit(1)
print("\n[SUCCESS] All PB-05 Chapter 9 Unit Tests Passed Successfully (100% Conformance).")
What This Demonstrates
- The same
ArchitectureModel→ diagram → deck pipeline from Chapter 07, but expressed as an explicit tool_use/tool_result trace instead of direct method calls — the shape a real Claude Agent SDK loop actually produces. - A hook that structurally cannot be talked past by the agent, closing the enforcement gap flagged against Chapter 07's prose-only iteration cap.
- Publication as an Artifact manifest (
url_slug,theme_aware,status) rather than a bare file path, reflecting the distribution-layer gap flagged against Chapters 04 and 07.
7. Summary
This companion chapter delivered what Chapters 01–08 could not deliver about themselves: an outside review. The findings are consistent across chapters — dated model pinning, a missing prompt-caching optimization, a prose-only (not hook-enforced) quality gate, and a file-first (not Artifact-first) distribution model. None of these findings invalidate the playbook's core engineering discipline (C4 rigor, deterministic geometry, declarative diagramming-as-code remain correct); they identify where an Anthropic/Claude-centric implementation closes gaps the original chapters left open. Teams standardizing on Claude Sonnet 5 / Opus 5, Claude Code hooks and subagents, MCP, and Artifacts should treat this chapter — not Chapter 05 alone — as the current reference for the Claude side of this playbook's stack.