Overview

Playbook 04: AI Coding & Software Engineering Playbook — Steering Instruction & Quality Standards

Playbook Track: 04 – AI Coding & Software Engineering (AWS AI-DLC & Autonomous Developer Workflows)
Target Audience: Year 1 Computer Science & Software Engineering Students Tooling Stack: Gemini 2.5 Flash & Pro (1M+ context codebase reasoning), AWS AI-DLC (aidlc), Model Context Protocol (MCP), Tree-sitter / AST Parsers, Pytest / Vitest, Docker, GitHub Actions, Python 3.11+
Delivery Format: Strict Markdown (.md) to GitHub origin/main
Tier 2 Export Rule: Word (.docx) in Google Docs format will only be compiled upon user approval.
Live Schedule Tracker: PB-04 Chapter Schedule (Google Sheets)


1. Executive Mission & Core Philosophy

The era of naive, conversational "copilot" coding—where a developer chats with an LLM in a sidebar and copies code snippets into their editor—is obsolete. Production-grade software engineering requires deterministic, spec-driven, test-anchored agentic workflows.

This playbook provides an end-to-end blueprint for engineering disciplined software systems using the AWS AI-Driven Development Life Cycle (AWS AI-DLC — `awslabs/aidlc-workflows`):

  1. Inception Phase: Intent elicitation, ambiguity detection, user story synthesis, formal Gherkin acceptance criteria, C4 architecture diagrams, and automated ADRs.
  2. Construction Phase: Interface design, contract-first OpenAPI schemas, agentic coding with Tree-sitter AST repo maps, TDD loops, and automated self-healing test repair.
  3. Operations Phase: CI/CD automation, GitOps pull requests, semantic diff inspection, and security vulnerability scanning.
  4. Framework Comparative Review: Deep, architectural evaluation of leading modern paradigms: BMAD (Agile AI Delivery), Spec-Kit (Specification-Driven Development), and AWS AI-DLC (AWS's open-source, harness-neutral governance framework).

2. The 7 Universal Quality Acceptance Gates

Every chapter of Playbook 04 must satisfy all 7 universal quality gates before publication:

  1. Gate 1: Zero Fluff & Agentic Engineering Rigor:
    • Open immediately with system architecture, state machines, AST schemas, and cognitive workflows.
    • Include detailed Mermaid architectural and sequence diagrams in every chapter.
    • No generic motivational intros or superficial coding advice.
  2. Gate 2: Mandatory Naive vs. Production Contrasts:
    • Directly contrast naive conversational coding (single-prompt text generation, copy-pasting, blind execution) with production agentic engineering (spec-first, AST-verified, test-driven, self-healing).
    • Provide structured comparative tables highlighting failure rates, regression risks, and token overhead.
  3. Gate 3: Latest Google & Frontier Tool Configurations & Schemas:
    • Detail exact parameter dictionaries, system prompts, tool call schemas, and API configurations for Gemini 2.5 Flash & Pro, Model Context Protocol (MCP), Tree-sitter, and GitHub Actions.
    • Use latest JSON Schema, OpenAPI 3.1, and Pydantic v2 data models.
  4. Gate 4: Quantitative Trade-Off Matrix:
    • Benchmark architecture patterns, context gathering strategies, test generation frameworks, and review topologies.
    • Quantify metrics: SWE-bench solve rate, token efficiency, hallucination frequency, latency, and operational cost.
  5. Gate 5: The 10 Operational Failure Modes in AI Software Engineering:
    • Analyze critical production failure modes (e.g. lazy stub generation // TODO, context window hallucination, merge conflict thrashing, AST syntax corruption, infinite repair loops, dependency drift).
    • Provide concrete programmatic defense mechanisms for each failure mode.
  6. Gate 6: Mandatory Hands-On Lab (Interactive Challenge):
    • Provide a step-by-step, actionable coding challenge allowing Year 1 students to run and test the subsystem directly on their machine.
    • Every challenge must test a realistic, non-trivial engineering workflow.
  7. Gate 7: Mandatory Recommended Answer & Executable Solution:
    • Provide a fully tested, runnable zero-dependency Python 3.11+ script.
    • Must contain built-in unit test assertions certifying 100% compliance.
    • Must execute cleanly and emit green assertion outputs.

3. Master 8-Chapter Syllabus + Appendices

Chapter Title Engineering Scope & Core Topics Hands-On Lab (Gate 6) Recommended Solution (Gate 7)
Ch 01 AWS AI-DLC & Framework Comparative Review AWS AI-DLC (awslabs/aidlc-workflows) Inception/Construction/Operations phases, BMAD, Spec-Kit, harness neutral rules, Year 1 apprentice lab. AWS AI-DLC Framework Evaluator & Spec Linter. Python 3.11+ zero-dep framework benchmark engine & spec conformance linter.
Ch 02 Automated Requirements Engineering & Spec-Kit PRD Synthesis Autonomous requirement elicitation, user story extraction, Gherkin acceptance criteria, Spec-Kit workflows, domain models, ambiguity scoring. Spec-Kit Requirements Compiler & Ambiguity Linter. Python engine parsing natural language goals into validated Gherkin specifications.
Ch 03 AI-Assisted System Architecture & C4 Modeling System decomposition (Context, Containers, Components, Code - C4 model), automated ADR generation, designing for agent-friendly modularity. C4 Architecture Generator & ADR Trade-Off Engine. Python architecture compiler synthesizing C4 Mermaid models and validated ADRs.
Ch 04 Interface Design, API Contracts & Schema Synthesis Contract-first engineering, OpenAPI 3.1, JSON Schema, Proto3, automated mock server generation, schema drift defense, breaking change detection. API Contract Compiler & Breaking Change Linter. Python contract generator producing OpenAPI 3.1 schemas, mock payloads, and drift assertions.
Ch 05 Agentic Coding, Context Gathering & Tree-sitter AST Mechanics Context retrieval (BM25 vs Vector vs AST Repo Maps), Model Context Protocol (MCP), unified diff patching, AST syntax validation, TDD loops. AST Symbol Repo Mapper & Unified Diff Patcher. Python code mapper extracting symbol call graphs and applying AST-verified diff patches.
Ch 06 Autonomous Verification, Testing & Self-Healing Code Loops Test synthesis (unit, integration, property-based), traceback parsing, AST-driven closed-loop self-repair, regression test guardrails. Self-Healing Test Runner & Automated Repair Engine. Python test execution harness with automated traceback parsing and patch synthesis.
Ch 07 Automated CI/CD, GitOps & Agentic Review Workflows Automated pull requests, semantic diff analysis, security vulnerability scanning, GitHub Actions automation, automated rollbacks. Agentic Pull Request Reviewer & Security Linter. Python CI/CD review engine evaluating code diffs for security, style, and regressions.
Ch 08 End-to-End Autonomous Software Engineering Suite Full lifecycle orchestration: Feature Request -> Spec-Kit -> Architecture ADR -> API Contract -> TDD Code -> Self-Healing Tests -> CI/CD. Autonomous Software Studio Suite CLI. Complete runnable Python 3.11+ CLI orchestrating a complete software feature build.
App A Agent System Prompts, Tool Schemas & SDLC Runbooks Complete system prompt library for 6 SDLC agent roles (Architect, Analyst, Developer, Tester, DevOps, Auditor), tool call schemas. SDLC Agent Prompt & Schema Reference. Reference architecture, prompts, and operational runbooks.
App B Curated GitHub Repositories & Open-Source AI Coding Ecosystem Directory of 25+ open-source AI coding repos: Aider, SWE-agent, OpenHands, Cline, Continue, BMAD, Spec-Kit, Tree-sitter, etc. Curated GitHub Repository Catalog. Open-source ecosystem reference and evaluation guide.

4. Specialized Agent Team Topology for AI-DLC

An autonomous software company replaces isolated human handoffs with a synchronized multi-agent state machine:

  1. Requirements Analyst Agent: Converts user requests into structured spec.md with Gherkin scenarios and domain entities.
  2. Software Architect Agent: Models system boundaries, generates C4 diagrams, and drafts Architecture Decision Records (ADRs).
  3. API & Interface Designer Agent: Defines strict OpenAPI 3.1 and JSON Schema contracts before coding begins.
  4. Lead Developer Agent: Ingests Repo Maps and AST symbols, executing Test-Driven Development (TDD) and producing unified diff patches.
  5. Quality & Test Engineer Agent: Executes test suites, parses compiler and runtime tracebacks, and coordinates closed-loop repairs.
  6. DevOps & GitOps Publisher Agent: Lints code diffs, verifies CI/CD status, generates semantic pull requests, and packages releases.

5. Execution Rules & GitHub Delivery

  1. Markdown First: All chapters and appendices are published in .md format to playbooks/04-ai-coding/.
  2. Continuous Pushing: Commit and push at natural checkpoints to origin/main so readers can pull and inspect files locally at D:\Agent-assisted-research.
  3. Google Sheets Synchronization: Update PB-04 Chapter Schedule and Playbook Catalog concurrently as each chapter is completed.
  4. Approval Gate for Word Export: Word (.docx) in Google Docs format remains strictly paused until the entire playbook is finished and explicit approval is granted.