Overview
Playbook 04: AI Coding & Software Engineering Playbook — Steering Instruction & Quality Standards
Playbook Track: 04 – AI Coding & Software Engineering (AWS AI-DLC & Autonomous Developer Workflows)
Target Audience: Year 1 Computer Science & Software Engineering Students Tooling Stack: Gemini 2.5 Flash & Pro (1M+ context codebase reasoning), AWS AI-DLC (aidlc), Model Context Protocol (MCP), Tree-sitter / AST Parsers, Pytest / Vitest, Docker, GitHub Actions, Python 3.11+
Delivery Format: Strict Markdown (.md) to GitHuborigin/main
Tier 2 Export Rule: Word (.docx) in Google Docs format will only be compiled upon user approval.
Live Schedule Tracker: PB-04 Chapter Schedule (Google Sheets)
1. Executive Mission & Core Philosophy
The era of naive, conversational "copilot" coding—where a developer chats with an LLM in a sidebar and copies code snippets into their editor—is obsolete. Production-grade software engineering requires deterministic, spec-driven, test-anchored agentic workflows.
This playbook provides an end-to-end blueprint for engineering disciplined software systems using the AWS AI-Driven Development Life Cycle (AWS AI-DLC — `awslabs/aidlc-workflows`):
- Inception Phase: Intent elicitation, ambiguity detection, user story synthesis, formal Gherkin acceptance criteria, C4 architecture diagrams, and automated ADRs.
- Construction Phase: Interface design, contract-first OpenAPI schemas, agentic coding with Tree-sitter AST repo maps, TDD loops, and automated self-healing test repair.
- Operations Phase: CI/CD automation, GitOps pull requests, semantic diff inspection, and security vulnerability scanning.
- Framework Comparative Review: Deep, architectural evaluation of leading modern paradigms: BMAD (Agile AI Delivery), Spec-Kit (Specification-Driven Development), and AWS AI-DLC (AWS's open-source, harness-neutral governance framework).
2. The 7 Universal Quality Acceptance Gates
Every chapter of Playbook 04 must satisfy all 7 universal quality gates before publication:
- Gate 1: Zero Fluff & Agentic Engineering Rigor:
- Open immediately with system architecture, state machines, AST schemas, and cognitive workflows.
- Include detailed Mermaid architectural and sequence diagrams in every chapter.
- No generic motivational intros or superficial coding advice.
- Gate 2: Mandatory Naive vs. Production Contrasts:
- Directly contrast naive conversational coding (single-prompt text generation, copy-pasting, blind execution) with production agentic engineering (spec-first, AST-verified, test-driven, self-healing).
- Provide structured comparative tables highlighting failure rates, regression risks, and token overhead.
- Gate 3: Latest Google & Frontier Tool Configurations & Schemas:
- Detail exact parameter dictionaries, system prompts, tool call schemas, and API configurations for Gemini 2.5 Flash & Pro, Model Context Protocol (MCP), Tree-sitter, and GitHub Actions.
- Use latest JSON Schema, OpenAPI 3.1, and Pydantic v2 data models.
- Gate 4: Quantitative Trade-Off Matrix:
- Benchmark architecture patterns, context gathering strategies, test generation frameworks, and review topologies.
- Quantify metrics: SWE-bench solve rate, token efficiency, hallucination frequency, latency, and operational cost.
- Gate 5: The 10 Operational Failure Modes in AI Software Engineering:
- Analyze critical production failure modes (e.g. lazy stub generation
// TODO, context window hallucination, merge conflict thrashing, AST syntax corruption, infinite repair loops, dependency drift). - Provide concrete programmatic defense mechanisms for each failure mode.
- Analyze critical production failure modes (e.g. lazy stub generation
- Gate 6: Mandatory Hands-On Lab (Interactive Challenge):
- Provide a step-by-step, actionable coding challenge allowing Year 1 students to run and test the subsystem directly on their machine.
- Every challenge must test a realistic, non-trivial engineering workflow.
- Gate 7: Mandatory Recommended Answer & Executable Solution:
- Provide a fully tested, runnable zero-dependency Python 3.11+ script.
- Must contain built-in unit test assertions certifying 100% compliance.
- Must execute cleanly and emit green assertion outputs.
3. Master 8-Chapter Syllabus + Appendices
| Chapter | Title | Engineering Scope & Core Topics | Hands-On Lab (Gate 6) | Recommended Solution (Gate 7) |
|---|---|---|---|---|
| Ch 01 | AWS AI-DLC & Framework Comparative Review | AWS AI-DLC (awslabs/aidlc-workflows) Inception/Construction/Operations phases, BMAD, Spec-Kit, harness neutral rules, Year 1 apprentice lab. |
AWS AI-DLC Framework Evaluator & Spec Linter. | Python 3.11+ zero-dep framework benchmark engine & spec conformance linter. |
| Ch 02 | Automated Requirements Engineering & Spec-Kit PRD Synthesis | Autonomous requirement elicitation, user story extraction, Gherkin acceptance criteria, Spec-Kit workflows, domain models, ambiguity scoring. | Spec-Kit Requirements Compiler & Ambiguity Linter. | Python engine parsing natural language goals into validated Gherkin specifications. |
| Ch 03 | AI-Assisted System Architecture & C4 Modeling | System decomposition (Context, Containers, Components, Code - C4 model), automated ADR generation, designing for agent-friendly modularity. | C4 Architecture Generator & ADR Trade-Off Engine. | Python architecture compiler synthesizing C4 Mermaid models and validated ADRs. |
| Ch 04 | Interface Design, API Contracts & Schema Synthesis | Contract-first engineering, OpenAPI 3.1, JSON Schema, Proto3, automated mock server generation, schema drift defense, breaking change detection. | API Contract Compiler & Breaking Change Linter. | Python contract generator producing OpenAPI 3.1 schemas, mock payloads, and drift assertions. |
| Ch 05 | Agentic Coding, Context Gathering & Tree-sitter AST Mechanics | Context retrieval (BM25 vs Vector vs AST Repo Maps), Model Context Protocol (MCP), unified diff patching, AST syntax validation, TDD loops. | AST Symbol Repo Mapper & Unified Diff Patcher. | Python code mapper extracting symbol call graphs and applying AST-verified diff patches. |
| Ch 06 | Autonomous Verification, Testing & Self-Healing Code Loops | Test synthesis (unit, integration, property-based), traceback parsing, AST-driven closed-loop self-repair, regression test guardrails. | Self-Healing Test Runner & Automated Repair Engine. | Python test execution harness with automated traceback parsing and patch synthesis. |
| Ch 07 | Automated CI/CD, GitOps & Agentic Review Workflows | Automated pull requests, semantic diff analysis, security vulnerability scanning, GitHub Actions automation, automated rollbacks. | Agentic Pull Request Reviewer & Security Linter. | Python CI/CD review engine evaluating code diffs for security, style, and regressions. |
| Ch 08 | End-to-End Autonomous Software Engineering Suite | Full lifecycle orchestration: Feature Request -> Spec-Kit -> Architecture ADR -> API Contract -> TDD Code -> Self-Healing Tests -> CI/CD. | Autonomous Software Studio Suite CLI. | Complete runnable Python 3.11+ CLI orchestrating a complete software feature build. |
| App A | Agent System Prompts, Tool Schemas & SDLC Runbooks | Complete system prompt library for 6 SDLC agent roles (Architect, Analyst, Developer, Tester, DevOps, Auditor), tool call schemas. | SDLC Agent Prompt & Schema Reference. | Reference architecture, prompts, and operational runbooks. |
| App B | Curated GitHub Repositories & Open-Source AI Coding Ecosystem | Directory of 25+ open-source AI coding repos: Aider, SWE-agent, OpenHands, Cline, Continue, BMAD, Spec-Kit, Tree-sitter, etc. | Curated GitHub Repository Catalog. | Open-source ecosystem reference and evaluation guide. |
4. Specialized Agent Team Topology for AI-DLC
An autonomous software company replaces isolated human handoffs with a synchronized multi-agent state machine:
- Requirements Analyst Agent: Converts user requests into structured
spec.mdwith Gherkin scenarios and domain entities. - Software Architect Agent: Models system boundaries, generates C4 diagrams, and drafts Architecture Decision Records (ADRs).
- API & Interface Designer Agent: Defines strict OpenAPI 3.1 and JSON Schema contracts before coding begins.
- Lead Developer Agent: Ingests Repo Maps and AST symbols, executing Test-Driven Development (TDD) and producing unified diff patches.
- Quality & Test Engineer Agent: Executes test suites, parses compiler and runtime tracebacks, and coordinates closed-loop repairs.
- DevOps & GitOps Publisher Agent: Lints code diffs, verifies CI/CD status, generates semantic pull requests, and packages releases.
5. Execution Rules & GitHub Delivery
- Markdown First: All chapters and appendices are published in
.mdformat toplaybooks/04-ai-coding/. - Continuous Pushing: Commit and push at natural checkpoints to
origin/mainso readers can pull and inspect files locally atD:\Agent-assisted-research. - Google Sheets Synchronization: Update
PB-04 Chapter ScheduleandPlaybook Catalogconcurrently as each chapter is completed. - Approval Gate for Word Export: Word (
.docx) in Google Docs format remains strictly paused until the entire playbook is finished and explicit approval is granted.