Overview

Chapter 01: LLM Foundations & Token Mechanics

How Language Models Actually "Think", Read, and Generate Text

Playbook: PB-01 (Prompt Engineering Playbook)
Target Audience: Year 1 Computer Science & Software Engineering Students Prerequisites: Basic Python syntax (variables, functions, strings, dictionaries)
Frontier Models Covered: Gemini 2.5 Flash & Pro, Claude 3.7 / Sonnet, GPT-4o
Steering Reference: `instruction.md`
Delivery Status: 🔍 Ready for Review (Tier 1 Markdown)


1. The Big Picture & Real-World Analogy

How an LLM Actually Works: The World's Most Advanced Autocomplete

When you text a friend on your smartphone, your keyboard suggests the next word:

  • If you type "See you at the...", your phone might suggest "library", "gym", or "airport".
  • Your phone does this by calculating which word is most likely to come next based on past text messages.

A Large Language Model (LLM) does the exact same thing—but on a planetary scale. Instead of reading just your text messages, it was trained on trillions of words from books, scientific papers, code repositories, and websites. And instead of looking back only two words, it looks back across hundreds of thousands of words in a single glance!

Input Prompt: "The capital of France is"
                          │
                          ▼
            [ Transformer Neural Network ]
                          │
                          ▼
       Candidate Next Words (Probabilities):
       ┌──────────────┬───────────────┐
       │ Next Word    │ Probability   │
       ├──────────────┼───────────────┤
       │ " Paris"     │ 94.1%  ◄──────┼── Chosen by default!
       │ " the"       │  4.3%         │
       │ " a"         │  0.9%         │
       │ " beautiful" │  0.2%         │
       └──────────────┴───────────────┘
                          │
                          ▼
                Emitted Word: " Paris"

When you type a prompt, the model:

  1. Calculates the probability for every possible word in its vocabulary.
  2. Selects the next word based on those probabilities and your settings (like Temperature).
  3. Adds that word to your prompt, and repeats the process one word at a time until it reaches a stopping point.

Because LLMs are probability engines, prompt engineering is simply the art of steering those probabilities so the model chooses the exact, correct answer you need.


2. Engineering Jargon Demystifier

Before writing code or complex prompts, let's demystify the core terms used across the AI industry:

Term What It Means in Plain English Why It Matters to You as a Student
Token A chunk of characters (about 3/4 of a word in English). Models do not read whole words or letters; they read token numbers. Every API call charges you per token, and every model has a maximum token limit.
Context Window The maximum amount of text (prompt + answer) the model can hold in memory at one time. If your prompt exceeds this limit, the model "forgets" the beginning of your conversation.
Temperature A slider from 0.0 to 1.0 (or 2.0) that controls how "creative" vs. "deterministic" the model is. Set to 0.0 for code, math, and JSON; set to 0.7 for creative writing or brainstorming.
Top-P (Nucleus Sampling) Another creativity filter. It tells the model to only pick from the top $P%$ most probable words (e.g. top_p = 0.9). Usually kept at 0.9 or 0.95. Don't change both Temperature and Top-P at the same time.
Base Model A raw model trained only to predict the next word. It does not know how to chat. If you ask a base model "What is 2+2?", it might complete it as "What is 3+3?" instead of answering!
Instruction-Tuned Model A model trained with human feedback (RLHF) to act like a helpful assistant that answers questions directly. These are the models you use in real projects (Gemini 2.5 Flash, Claude 3.7 Sonnet, GPT-4o).
Reasoning Model A special model (e.g., Gemini 2.0 Flash Thinking, Claude 3.7 Thinking, OpenAI o1) that generates a hidden "chain of thought" before giving you the final answer. Best for hard math, algorithmic coding, and debugging complex logic.

3. The 5-Minute Micro-Lab: The Token & Temperature Explorer

Let's see how tokens and temperature work using a simple, 25-line Python script that requires zero external pip packages!

The Code: micro_token_lab.py

# micro_token_lab.py - Zero external dependencies!
import math
import random

# A toy vocabulary mapping words to raw scores (logits)
CANDIDATE_LOGITS = {
    " Paris": 8.0,
    " the": 5.0,
    " a": 3.5,
    " beautiful": 2.0,
    " France": 0.5
}

def simulate_temperature(logits: dict, temperature: float) -> dict:
    """Demonstrates how Temperature flattens or sharpens token probabilities."""
    # Step 1: Scale logits by temperature (avoid divide by zero)
    temp = max(temperature, 0.01)
    scaled = {word: math.exp(score / temp) for word, score in logits.items()}
    total = sum(scaled.values())
    
    # Step 2: Convert to percentages
    return {word: round((val / total) * 100, 2) for word, val in scaled.items()}

print("Simulating LLM Next-Token Selection for: 'The capital of France is...'")
for t in [0.1, 0.7, 1.5]:
    probs = simulate_temperature(CANDIDATE_LOGITS, temperature=t)
    print(f"\nTemperature {t}:")
    for word, pct in sorted(probs.items(), key=lambda x: x[1], reverse=True)[:3]:
        print(f"  {repr(word):<12} -> {pct}%")

Try It Yourself:

  1. Save this as micro_token_lab.py and run python micro_token_lab.py.
  2. Notice what happens:
    • At Temperature 0.1 (deterministic), " Paris" has a 99.9% chance of being selected. The model behaves like a reliable calculator.
    • At Temperature 1.5 (high randomness), the probabilities flatten out, and the model might pick " the" or " beautiful".

4. How Tokenizers Work & The 3 Traps Freshmen Fall Into

Computers don't understand words; they understand numbers. A Tokenizer cuts text into integer IDs:

Text:       "Prompt engineering is awesome!"
Tokens:     ["Prompt", " engineering", " is", " awesome", "!"]
Token IDs:  [ 32174,     23758,          374,    14298,     0  ]

Here are three hidden quirks of tokenizers that will bite you if you don't know about them:

1. The Space-Prefix Trap

To a tokenizer, a word with a space in front of it is a completely different number than the same word without a space:

  • " apple" (with leading space) $\rightarrow$ Token ID 17540
  • "apple" (no leading space) $\rightarrow$ Token ID 25776

Rule for Students: When formatting prompts with variables, beware of accidental trailing spaces (e.g. f"Translate: {word} "). A stray space at the end can confuse the model and cause weird output formatting.

2. Number Splitting (Why LLMs Stumble at Math)

Tokenizers prioritize common English words. They don't have single tokens for large numbers:

  • "482910" is split into three chunks: ["48", "29", "10"]
  • "482911" is split into two chunks: ["482", "911"]

Because the model sees fragmented number chunks rather than mathematical values, asking an LLM to multiply 6-digit numbers in its head will usually fail.

  • The Fix: Always tell the model: "Show your work step-by-step" or have the model write a Python script to do the calculation for you.

3. The Multilingual Token Tax

Tokenizers were trained heavily on English. A 10-word sentence in English takes ~12 tokens. But the exact same sentence in Vietnamese, Japanese, or Spanish can take 25 to 40 tokens because the tokenizer has to break non-English words into individual syllables or bytes!

Rule for Students: If you are writing system prompts for a bilingual project, write the instructions and rules in English, and ask the model to produce the final response in Vietnamese. This saves up to 40% on token costs and speeds up generation!


5. Freshman Survival Guide: 3 Traps to Avoid

Trap 1: Using High Temperature for Code & JSON

  • The Mistake: Leaving the temperature at default 1.0 when asking the AI to return Python code or JSON data. The model gets "creative" and hallucinates syntax errors or invalid JSON keys.
  • How to Avoid It: Whenever you need deterministic output (Python code, SQL queries, JSON schemas, calculations), always set temperature=0.0 or 0.1.

Trap 2: Lost-in-the-Middle Syndrome

  • The Mistake: Pasting a 20-page PDF or 500 lines of code into the prompt, putting your question in the middle, and wondering why the model ignored your instructions.
  • How to Avoid It: LLMs pay the most attention to the beginning (Primacy) and the end (Recency) of your prompt.
    • Put your instructions and rules at the top.
    • Put your background data / code in the middle.
    • Put your specific question and desired output format at the very bottom.

Trap 3: Runaway Generations (Missing max_tokens)

  • The Mistake: Calling an API without setting max_tokens. If the model enters an infinite loop or hallucinates a novel, it can drain your entire API credit in a few minutes.
  • How to Avoid It: Always set a reasonable max_output_tokens limit (e.g. 1024 for short explanations, 4096 for code files).

6. Mandatory Hands-On Lab: The Token Cost & Budget Calculator

Lab Objective

In this hands-on lab, you will build and run a Token Budget & Cost Calculator in pure Python 3.11+.

You will:

  1. Estimate token counts for user queries and system instructions using standard subword estimation rules.
  2. Calculate the exact API cost across Gemini 2.5 Flash, Gemini 2.5 Pro, and Claude 3.7 Sonnet.
  3. Calculate how much money you save when using Context Caching for repetitive system instructions.
  4. Run automated self-test assertions to certify that your calculations are 100% accurate.

Step-by-Step Instructions

  1. Save the code below as token_lab.py.
  2. Run it using Python 3.11+:
    python token_lab.py
    
  3. Observe the clean console output and verified self-test assertions.

8. Chapter Summary & What's Next

In this chapter, you learned:

  1. How LLMs Work: They are next-token probability predictors. Prompt engineering is simply steering those probabilities.
  2. Tokens vs. Words: Models read tokens, which can split numbers, penalize non-English languages, and react differently to leading spaces.
  3. Temperature Control: Use 0.0 for code and math; use 0.7 for creative writing.
  4. Economics & Caching: You can save 75%–90% on API costs by caching large system prompts.

Coming Up in Chapter 02:

In Chapter 02: Core Prompting Architectures, you will learn how to build your first professional prompts using Zero-Shot, Few-Shot Exemplars, Role/Persona Framing, and XML Delimiters to guarantee consistent, high-quality answers.