16-lesson course · No AI background required
A 16-lesson technical curriculum for business leaders and technical professionals — no AI background required.
- A curriculum about the mechanism, not the vocabulary.
- Squarely between prompting tips and the maths.
- 16lessons
- 4parts
- 0lines of code
What you’ll be able to trace
- TokenizerL1Thecancanrust
- Embedding layerEach token becomes numbers — plus its positionL2–3
- Transformer stackAttention + feed-forward, layer after layerL5–9
- Output layerOne next token, chosen — then againL10–11
Then: how every number in it was learned. L13–16
Why take this course
Stop guessing what the model is doing.
Each of these comes from understanding a specific mechanism — not from a list of tips.
See why the same word means different things
How a model that starts every word from one fixed representation ends up reading “can” as a container in one place and an ability in the next.
Lessons 3, 5–6Know what you’re actually paying for
Why tokens, not words, set your bill and your limits — why output costs more than input, and why doubling a prompt more than doubles the work.
Lessons 1, 6, 11Know what happens to what you type
What is fixed when the model is trained, what is built fresh for your prompt and thrown away — and what your provider may still keep.
Lesson 8Read vendor claims for what they are
What a parameter count, a “larger” model, or “diverse, high-quality training data” does — and doesn’t — tell you.
Lessons 4, 7, 13Know what the model actually saw
What fits in the context window, what happens when your document doesn’t, and why the same prompt can give two different answers.
Lessons 10, 12Understand where its behavior comes from
How training, fine-tuning and post-training shape what a model does — including why it refuses, and when you can change that.
Lessons 14–16Questions you’ll finally be able to answer
Have you wondered about any of these?
Most lessons open with one of these questions — and close by answering it from the mechanism up.
- Is this model biased — and does “trained on diverse, high-quality data” answer that?L13
- Why does output cost several times more than input?L11
- Is the model learning from what we type?L8
- Did the model actually read page 190?L12
- What am I really buying when I choose a larger model?L7
- Do more parameters mean a better model?L4
- Same prompt, two different answers — is that a bug?L10
- Does the tokenizer matter when choosing a model?L1
- How does the model tell my instructions from the document I pasted?L3
- Can we just train our own model on our data?L14
- Is there a version where our data never leaves our infrastructure?L15
- The model refused a legitimate request. Who decided that — and can we change it?L16
Who it’s for
No AI background required. That is not the same as no depth.
You’ll work through the real operations — attention scores, normalization, how a token is chosen — with numbers small enough to check by hand. No code, and no maths prerequisite.
This course is for you if…
- You make decisions about AI systems — choosing models, approving vendors, owning cost or risk.
- You’re a technical professional who works with LLMs but isn’t an ML engineer.
- You’re an engineer or developer who wants the architecture, not just the API.
- You work in technical product, program or enablement roles and need to explain these systems to others.
- You like to understand how something works before you trust it.
It isn’t for you if…
- You want prompting tips or a tour of AI tools.
- You want to write ML code or implement a model yourself.
- You want mathematical derivations rather than worked examples.
What you’ll learn
16 lessons. One connected system.
You follow a prompt through the model, then step back to see how the model was built. Each lesson depends on the ones before it.
Part 1
From Text to Numbers
- 01TokenizationFreeHow text becomes tokens — and why that changes your costs.
- 02Token RepresentationsHow meaning becomes coordinates nobody designed.
- 03Embedding PositionHow the model knows word order — but not role.
- 04The Language of Machine LearningParameters, loss and gradients: the terms everything after builds on.
Part 2
Inside the Transformer
- 05The Transformer StackWhy a stack of identical layers, and what each one changes.
- 06The Attention MechanismHow context reshapes every token — and what it costs.
- 07The Feed-Forward NetworkWhere a bigger model’s extra parameters actually sit.
- 08Residual Addition & NormalizationWhat is frozen, and what is built fresh for your prompt.
- 09Modern Transformer ArchitectureWhy today’s models rearranged the classic design.
Part 3
Producing an Answer
- 10The Output LayerHow the next token is chosen, and why answers vary.
- 11The KV CacheHow generation works, and why output costs more.
- 12The Context WindowWhat the model can see — and what happens beyond it.
Part 4
Building and Changing the Model
- 13Preparing the Training CorpusWhere human judgement — and bias — enters the data.
- 14Training the LLMHow training changes the parameters, and what fine-tuning really costs.
- 15Open WeightsWhat is actually released — and when it keeps your data in-house.
- 16Post-Training an LLMHow a model is shaped into an assistant — and what that shaping costs.
Inference comes before training on purpose: by the time you reach training, you’ve already met every part of the model it changes.
Detailed syllabus
Why this order
The sixteen lessons follow a prompt through a model, then step back to explain how the model was made. Lessons 1–4 cover what the model receives. Lessons 5–9 open the transformer one component at a time. Lessons 10–12 complete inference: choosing a token, generating a response, and the limits of context. Lessons 13–16 cover construction: data, pretraining, open weights and post-training.
Inference comes before training on purpose. By Lesson 14 learners have already met every component whose parameters training adjusts, so training is introduced as a modification of the inference diagram they built over the previous ten lessons — not as a new system.
Part One
From Text to Numbers
- 01
Tokenization Watch
How raw text is split into the units a model actually reads, how a tokenizer’s vocabulary is built, and why tokenizer choice affects cost, context limits and non-English workloads. It comes first because every later stage operates on tokens, not words.
- 02
Token Representations
How each token becomes a point in a learned, high-dimensional space. Covers why nobody designs what the dimensions mean, and why “meaning” is a by-product of next-token prediction rather than its goal.
- 03
Embedding Position
What enters the model when a prompt arrives: one fixed, context-free lookup per token, plus position information supplied separately. Poses the question Lessons 5–10 answer: how does a fixed starting representation come to reflect context?
- 04
The Language of Machine Learning
A deliberate pause before the transformer: parameters, weights and biases, training versus inference, loss, gradients, learning rate, batches and epochs — so the architecture lessons can use them precisely. Tested against published AI news coverage.
Part Two
Inside the Transformer
- 05
The Transformer Stack
What the transformer is for, and why it is a stack of identical layers that each learn their own values. Covers what changes, and what does not, as representations pass through. Introduces the two components examined next.
- 06
The Attention Mechanism
How each token draws information from the tokens before it: queries, keys, values, multiple heads and causal masking, worked through at toy scale. Connects the mechanism to why longer prompts cost disproportionately more compute.
- 07
The Feed-Forward Network
The second component of each layer: what it does that attention does not, why it expands and then compresses each representation, and where a larger model’s additional parameters sit — the basis for matching model size to the task.
- 08
Residual Addition & Normalization
Completes the layer with the machinery that keeps a deep stack stable, then separates everything fixed by training from everything computed fresh for one prompt — the technical basis for answering whether a model learns from what users type.
- 09
Modern Transformer Architecture
Why the classic order of components proved unreliable to train at depth, and how pre-norm and RMSNorm changed it. Uses the gradient vocabulary from Lesson 4 to show architecture as a set of engineering trade-offs.
Part Three
Producing an Answer
- 10
The Output Layer
How the final representation becomes a choice of next token: logits, temperature, softmax, top-k, greedy decoding and sampling. Resolves the question opened in Lesson 3, and addresses what non-deterministic output means for audit trails.
- 11
The KV Cache
How an inference system generates text one token at a time, what it stores to avoid recomputation, and what that storage costs. Completes the end-to-end inference architecture and explains why input and output are priced so differently.
- 12
The Context Window
What the context contains, who decides what goes into it, and what happens when the input exceeds the limit. Distinguishes context from memory, and asks learners to establish what a model actually received rather than assume it.
Part Four
Building and Changing the Model
- 13
Preparing the Training Corpus
The shift from using a model to building one. Follows an example data pipeline from source selection to tokenization, identifying where human judgement — and therefore bias — enters at each stage.
- 14
Training the LLM
Pretraining in detail: how one sequence yields many prediction targets, how loss, backpropagation and the optimizer update parameters, and how the inference architecture is modified for training. Distinguishes pretraining from fine-tuning in cost and feasibility.
- 15
Open Weights
Why releasing source code would not give anyone a working model, what an open-weights release actually contains, and how open weights differs from open source. Connects the distinction to data-residency requirements.
- 16
Post-Training an LLM
How a pretrained model becomes a usable assistant: supervised fine-tuning, preference optimization (DPO, and RLHF with a reward model) and safety training — including over-refusal, catastrophic forgetting, and why safety training is not compliance.
Free preview
Start now: the introduction and Lesson 1 are free.
See exactly how the course teaches before the full course launches.
Inside the course
What the teaching actually looks like.
- One running example“The can can rust” is carried through the whole architecture, lesson by lesson.
- Real calculations, small numbersThe actual operations, with numbers small enough to check by hand.
- Misconceptions, namedCommon wrong ideas are stated plainly — then corrected.
- Every mechanism, a decisionLessons end with what it means for cost, data, risk or vendor choice.
“Bias didn’t survive the pipeline. It is the pipeline.”
LESSON 13 · PREPARING THE TRAINING CORPUS
Hands-on exercises
You won’t just watch. You’ll investigate.
Exercises send you to real tools and real documents, and ask you to reach a conclusion you can defend.
Measure your own text
Run the same text through the tokenizers of four model families and compare how each one splits it — before you choose a model.
Investigate a vendor
Legal asks whether an AI assistant can be trusted with confidential customer data. Check a real vendor’s documentation against their questions.
Did the AI read it all?
Give a real AI tool a very long document, then work out what you can — and can’t — conclude from what happens.
“But you cannot establish why.”
Get the full course
Ready to deliver — in person or over Zoom.
Bring How LLMs Actually Work to your organization. I’m ready to deliver the complete 16-lesson course in person or remotely, tailored to your audience.
- ✓ WrittenAll 16 lessons
- ✓ PublishedIntroduction and Lesson 1
- ◐ RecordingLessons 2–16
- ○ NextFull course launch
Created and taught by Shana Sokolic. Read about my approach



