Overview
Chapter 01: LLM Foundations & Token Mechanics
How Language Models Actually "Think", Read, and Generate Text
Playbook: PB-01 (Prompt Engineering Playbook)
Target Audience: Year 1 Computer Science & Software Engineering Students Prerequisites: Basic Python syntax (variables, functions, strings, dictionaries)
Frontier Models Covered: Gemini 2.5 Flash & Pro, Claude 3.7 / Sonnet, GPT-4o
Steering Reference: `instruction.md`
Delivery Status: 🔍 Ready for Review (Tier 1 Markdown)
1. The Big Picture & Real-World Analogy
How an LLM Actually Works: The World's Most Advanced Autocomplete
When you text a friend on your smartphone, your keyboard suggests the next word:
- If you type "See you at the...", your phone might suggest "library", "gym", or "airport".
- Your phone does this by calculating which word is most likely to come next based on past text messages.
A Large Language Model (LLM) does the exact same thing—but on a planetary scale. Instead of reading just your text messages, it was trained on trillions of words from books, scientific papers, code repositories, and websites. And instead of looking back only two words, it looks back across hundreds of thousands of words in a single glance!
Input Prompt: "The capital of France is"
│
▼
[ Transformer Neural Network ]
│
▼
Candidate Next Words (Probabilities):
┌──────────────┬───────────────┐
│ Next Word │ Probability │
├──────────────┼───────────────┤
│ " Paris" │ 94.1% ◄──────┼── Chosen by default!
│ " the" │ 4.3% │
│ " a" │ 0.9% │
│ " beautiful" │ 0.2% │
└──────────────┴───────────────┘
│
▼
Emitted Word: " Paris"
When you type a prompt, the model:
- Calculates the probability for every possible word in its vocabulary.
- Selects the next word based on those probabilities and your settings (like Temperature).
- Adds that word to your prompt, and repeats the process one word at a time until it reaches a stopping point.
Because LLMs are probability engines, prompt engineering is simply the art of steering those probabilities so the model chooses the exact, correct answer you need.
2. Engineering Jargon Demystifier
Before writing code or complex prompts, let's demystify the core terms used across the AI industry:
| Term | What It Means in Plain English | Why It Matters to You as a Student |
|---|---|---|
| Token | A chunk of characters (about 3/4 of a word in English). Models do not read whole words or letters; they read token numbers. | Every API call charges you per token, and every model has a maximum token limit. |
| Context Window | The maximum amount of text (prompt + answer) the model can hold in memory at one time. | If your prompt exceeds this limit, the model "forgets" the beginning of your conversation. |
| Temperature | A slider from 0.0 to 1.0 (or 2.0) that controls how "creative" vs. "deterministic" the model is. | Set to 0.0 for code, math, and JSON; set to 0.7 for creative writing or brainstorming. |
| Top-P (Nucleus Sampling) | Another creativity filter. It tells the model to only pick from the top $P%$ most probable words (e.g. top_p = 0.9). |
Usually kept at 0.9 or 0.95. Don't change both Temperature and Top-P at the same time. |
| Base Model | A raw model trained only to predict the next word. It does not know how to chat. | If you ask a base model "What is 2+2?", it might complete it as "What is 3+3?" instead of answering! |
| Instruction-Tuned Model | A model trained with human feedback (RLHF) to act like a helpful assistant that answers questions directly. | These are the models you use in real projects (Gemini 2.5 Flash, Claude 3.7 Sonnet, GPT-4o). |
| Reasoning Model | A special model (e.g., Gemini 2.0 Flash Thinking, Claude 3.7 Thinking, OpenAI o1) that generates a hidden "chain of thought" before giving you the final answer. | Best for hard math, algorithmic coding, and debugging complex logic. |
3. The 5-Minute Micro-Lab: The Token & Temperature Explorer
Let's see how tokens and temperature work using a simple, 25-line Python script that requires zero external pip packages!
The Code: micro_token_lab.py
# micro_token_lab.py - Zero external dependencies!
import math
import random
# A toy vocabulary mapping words to raw scores (logits)
CANDIDATE_LOGITS = {
" Paris": 8.0,
" the": 5.0,
" a": 3.5,
" beautiful": 2.0,
" France": 0.5
}
def simulate_temperature(logits: dict, temperature: float) -> dict:
"""Demonstrates how Temperature flattens or sharpens token probabilities."""
# Step 1: Scale logits by temperature (avoid divide by zero)
temp = max(temperature, 0.01)
scaled = {word: math.exp(score / temp) for word, score in logits.items()}
total = sum(scaled.values())
# Step 2: Convert to percentages
return {word: round((val / total) * 100, 2) for word, val in scaled.items()}
print("Simulating LLM Next-Token Selection for: 'The capital of France is...'")
for t in [0.1, 0.7, 1.5]:
probs = simulate_temperature(CANDIDATE_LOGITS, temperature=t)
print(f"\nTemperature {t}:")
for word, pct in sorted(probs.items(), key=lambda x: x[1], reverse=True)[:3]:
print(f" {repr(word):<12} -> {pct}%")
Try It Yourself:
- Save this as
micro_token_lab.pyand runpython micro_token_lab.py. - Notice what happens:
- At Temperature 0.1 (deterministic),
" Paris"has a 99.9% chance of being selected. The model behaves like a reliable calculator. - At Temperature 1.5 (high randomness), the probabilities flatten out, and the model might pick
" the"or" beautiful".
- At Temperature 0.1 (deterministic),
4. How Tokenizers Work & The 3 Traps Freshmen Fall Into
Computers don't understand words; they understand numbers. A Tokenizer cuts text into integer IDs:
Text: "Prompt engineering is awesome!"
Tokens: ["Prompt", " engineering", " is", " awesome", "!"]
Token IDs: [ 32174, 23758, 374, 14298, 0 ]
Here are three hidden quirks of tokenizers that will bite you if you don't know about them:
1. The Space-Prefix Trap
To a tokenizer, a word with a space in front of it is a completely different number than the same word without a space:
" apple"(with leading space) $\rightarrow$ Token ID17540"apple"(no leading space) $\rightarrow$ Token ID25776
Rule for Students: When formatting prompts with variables, beware of accidental trailing spaces (e.g.
f"Translate: {word} "). A stray space at the end can confuse the model and cause weird output formatting.
2. Number Splitting (Why LLMs Stumble at Math)
Tokenizers prioritize common English words. They don't have single tokens for large numbers:
"482910"is split into three chunks:["48", "29", "10"]"482911"is split into two chunks:["482", "911"]
Because the model sees fragmented number chunks rather than mathematical values, asking an LLM to multiply 6-digit numbers in its head will usually fail.
- The Fix: Always tell the model: "Show your work step-by-step" or have the model write a Python script to do the calculation for you.
3. The Multilingual Token Tax
Tokenizers were trained heavily on English. A 10-word sentence in English takes ~12 tokens. But the exact same sentence in Vietnamese, Japanese, or Spanish can take 25 to 40 tokens because the tokenizer has to break non-English words into individual syllables or bytes!
Rule for Students: If you are writing system prompts for a bilingual project, write the instructions and rules in English, and ask the model to produce the final response in Vietnamese. This saves up to 40% on token costs and speeds up generation!
5. Freshman Survival Guide: 3 Traps to Avoid
Trap 1: Using High Temperature for Code & JSON
- The Mistake: Leaving the temperature at default
1.0when asking the AI to return Python code or JSON data. The model gets "creative" and hallucinates syntax errors or invalid JSON keys. - How to Avoid It: Whenever you need deterministic output (Python code, SQL queries, JSON schemas, calculations), always set
temperature=0.0or0.1.
Trap 2: Lost-in-the-Middle Syndrome
- The Mistake: Pasting a 20-page PDF or 500 lines of code into the prompt, putting your question in the middle, and wondering why the model ignored your instructions.
- How to Avoid It: LLMs pay the most attention to the beginning (Primacy) and the end (Recency) of your prompt.
- Put your instructions and rules at the top.
- Put your background data / code in the middle.
- Put your specific question and desired output format at the very bottom.
Trap 3: Runaway Generations (Missing max_tokens)
- The Mistake: Calling an API without setting
max_tokens. If the model enters an infinite loop or hallucinates a novel, it can drain your entire API credit in a few minutes. - How to Avoid It: Always set a reasonable
max_output_tokenslimit (e.g.1024for short explanations,4096for code files).
6. Mandatory Hands-On Lab: The Token Cost & Budget Calculator
Lab Objective
In this hands-on lab, you will build and run a Token Budget & Cost Calculator in pure Python 3.11+.
You will:
- Estimate token counts for user queries and system instructions using standard subword estimation rules.
- Calculate the exact API cost across Gemini 2.5 Flash, Gemini 2.5 Pro, and Claude 3.7 Sonnet.
- Calculate how much money you save when using Context Caching for repetitive system instructions.
- Run automated self-test assertions to certify that your calculations are 100% accurate.
Step-by-Step Instructions
- Save the code below as
token_lab.py. - Run it using Python 3.11+:
python token_lab.py - Observe the clean console output and verified self-test assertions.
7. Mandatory Recommended Answer & Executable Solution
"""
token_lab.py
Zero-dependency Python 3.11+ script for Chapter 01:
- Estimates token counts without needing external C-libraries
- Compares API pricing across Gemini 2.5 Flash, Gemini 2.5 Pro, and Claude 3.7
- Simulates Context Caching discounts
- Includes 100% verified self-test assertions
"""
import math
from dataclasses import dataclass
from typing import Dict, Any
# ============================================================================
# 1. MODEL PRICING DEFINITIONS (Per 1 Million Tokens)
# ============================================================================
@dataclass(frozen=True)
class ModelPricing:
name: str
input_cost_per_m: float # Cost for 1,000,000 input tokens
output_cost_per_m: float # Cost for 1,000,000 output tokens
cached_input_discount: float # e.g., 0.75 means 75% discount on cached tokens
PRICING_TABLE = {
"gemini-2.5-flash": ModelPricing(
name="Gemini 2.5 Flash",
input_cost_per_m=0.10,
output_cost_per_m=0.40,
cached_input_discount=0.75
),
"gemini-2.5-pro": ModelPricing(
name="Gemini 2.5 Pro",
input_cost_per_m=1.25,
output_cost_per_m=5.00,
cached_input_discount=0.75
),
"claude-3.7-sonnet": ModelPricing(
name="Claude 3.7 Sonnet",
input_cost_per_m=3.00,
output_cost_per_m=15.00,
cached_input_discount=0.90
)
}
# ============================================================================
# 2. TOKEN ESTIMATOR & COST CALCULATOR
# ============================================================================
class TokenBudgetEngine:
"""Estimates token usage and calculates costs across frontier AI models."""
@staticmethod
def estimate_tokens(text: str) -> int:
"""
Rule of thumb for English text: ~4 characters per token or ~0.75 words per token.
We take the maximum of character-based and word-based estimation for safety.
"""
if not text:
return 0
char_est = math.ceil(len(text) / 3.8)
word_est = math.ceil(len(text.split()) * 1.3)
return max(char_est, word_est)
@classmethod
def calculate_request_cost(
cls,
model_key: str,
prompt_text: str,
expected_output_tokens: int,
use_cached_prompt: bool = False
) -> Dict[str, Any]:
"""Calculates the dollar cost of an API call."""
if model_key not in PRICING_TABLE:
raise ValueError(f"Unknown model: {model_key}. Available: {list(PRICING_TABLE.keys())}")
pricing = PRICING_TABLE[model_key]
input_tokens = cls.estimate_tokens(prompt_text)
# Apply context caching discount if enabled
if use_cached_prompt:
effective_input_rate = pricing.input_cost_per_m * (1.0 - pricing.cached_input_discount)
else:
effective_input_rate = pricing.input_cost_per_m
input_cost = (input_tokens / 1_000_000) * effective_input_rate
output_cost = (expected_output_tokens / 1_000_000) * pricing.output_cost_per_m
total_cost = input_cost + output_cost
return {
"model": pricing.name,
"estimated_input_tokens": input_tokens,
"output_tokens": expected_output_tokens,
"input_cost_usd": round(input_cost, 6),
"output_cost_usd": round(output_cost, 6),
"total_cost_usd": round(total_cost, 6),
"is_cached": use_cached_prompt
}
# ============================================================================
# 3. SELF-TEST VERIFICATION SUITE
# ============================================================================
def run_self_tests():
print("=" * 75)
print("RUNNING CHAPTER 01 SELF-TEST VERIFICATION SUITE (TOKEN LAB)")
print("=" * 75)
# Test 1: Empty text handling
assert TokenBudgetEngine.estimate_tokens("") == 0, "Test 1 Failed: Empty text must be 0 tokens"
print("[PASS] Test 1: Empty string token estimation returned 0.")
# Test 2: Standard phrase estimation
sample_text = "The quick brown fox jumps over the lazy dog."
tok_count = TokenBudgetEngine.estimate_tokens(sample_text)
assert 10 <= tok_count <= 18, f"Test 2 Failed: Expected 10-18 tokens, got {tok_count}"
print(f"[PASS] Test 2: '{sample_text}' estimated at {tok_count} tokens.")
# Test 3: Gemini 2.5 Flash Cost Calculation
res_flash = TokenBudgetEngine.calculate_request_cost(
model_key="gemini-2.5-flash",
prompt_text="Hello world, please write a bubble sort algorithm in Python.",
expected_output_tokens=300,
use_cached_prompt=False
)
assert res_flash["total_cost_usd"] > 0, "Test 3 Failed: Total cost must be greater than zero"
assert res_flash["total_cost_usd"] < 0.001, "Test 3 Failed: Single Flash call should be < 0.1 cent"
print(f"[PASS] Test 3: Gemini 2.5 Flash call calculated at ${res_flash['total_cost_usd']} USD.")
# Test 4: Context Caching Savings
big_system_prompt = "You are a coding assistant. " * 500 # ~1,500 tokens
uncached = TokenBudgetEngine.calculate_request_cost("gemini-2.5-pro", big_system_prompt, 200, use_cached_prompt=False)
cached = TokenBudgetEngine.calculate_request_cost("gemini-2.5-pro", big_system_prompt, 200, use_cached_prompt=True)
assert cached["input_cost_usd"] < uncached["input_cost_usd"], "Test 4 Failed: Cached cost must be strictly cheaper"
savings_pct = (1.0 - (cached["input_cost_usd"] / uncached["input_cost_usd"])) * 100
assert abs(savings_pct - 75.0) < 1.0, f"Test 4 Failed: Expected ~75% savings, got {savings_pct}%"
print(f"[PASS] Test 4: Context caching delivered {savings_pct:.1f}% discount on input tokens.")
print("\n" + "=" * 75)
print("[SUCCESS] ALL 4 SELF-TEST SUITES PASSED CLEANLY (100% GREEN ASSERTIONS)")
print("=" * 75)
if __name__ == "__main__":
run_self_tests()
8. Chapter Summary & What's Next
In this chapter, you learned:
- How LLMs Work: They are next-token probability predictors. Prompt engineering is simply steering those probabilities.
- Tokens vs. Words: Models read tokens, which can split numbers, penalize non-English languages, and react differently to leading spaces.
- Temperature Control: Use
0.0for code and math; use0.7for creative writing. - Economics & Caching: You can save 75%–90% on API costs by caching large system prompts.
Coming Up in Chapter 02:
In Chapter 02: Core Prompting Architectures, you will learn how to build your first professional prompts using Zero-Shot, Few-Shot Exemplars, Role/Persona Framing, and XML Delimiters to guarantee consistent, high-quality answers.