Free self-paced course

Advanced LLM Systems Course

This advanced LLM systems course connects transformer architecture and attention to adaptation, quantization, inference kernels, KV caches, serving, bounded agent loops, execution harnesses, evaluation, security, and production operations.

For

Experienced AI engineers and systems developers who need architecture depth and production evidence.

Study time

40–60 hours

Prerequisites

LLM foundations, Python, linear algebra, APIs, testing, and basic distributed-systems concepts.

What you will learn

  • Reason precisely about transformer and attention internals.
  • Compare adaptation and compression strategies.
  • Design reliable inference and model-serving systems.
  • Defend an LLM platform with evaluation, security, and operational evidence.

Use these answer-first lessons to build the concepts this course depends on.

Course format

Each topic combines a concise explanation, an animated mechanism, hands-on evidence, and an optional book-level deep dive.

  • Transformer and attention internals
  • LoRA, quantization, and inference
  • KV cache and model serving
  • Loop and Harness Engineering
  • Production architecture review

Mastery evidence

ML Serving Readiness

Turn quality evidence and traffic assumptions into capacity, SLO, canary, fallback, and rollback decisions.

Open the project

Get one next lesson each week

A focused path reminder, not a marketing newsletter. Unsubscribe in one click.